Compare commits

..

5 Commits

Author SHA1 Message Date
taylanbakircioglu bd6a31cb0d feat: v1.6.0 — Multi-Factor Authentication (Issue #18)
Adds opt-in TOTP-based Multi-Factor Authentication that is fully
backwards compatible with existing logins. Operators choose to enable
MFA per account; nothing changes for users who do not opt in.

Highlights
==========

* RFC 6238 TOTP (6 digits, 30s period, SHA1) with ±30s skew tolerance,
  compatible with Microsoft / Google Authenticator, Authy, Duo, 1Password.
* Per-step replay protection (`mfa_last_used_totp_step`) so a captured
  code cannot be reused inside the same window.
* Fernet-encrypted TOTP secrets at rest, key resolution via
  `MFA_ENCRYPTION_KEY` env (HKDF-derived from `SECRET_KEY` as fallback).
* 10 single-use, bcrypt-hashed backup codes per user, formatted
  `XXXX-YYYY` from a confusion-free alphabet (no 0/O/1/I/L).
* Two-step login flow: `POST /api/auth/login` returns `mfa_required`
  + `mfa_token`, then `POST /api/auth/login/mfa-verify` accepts a TOTP
  code OR a backup code. JWT is minted only after MFA succeeds.
* Self-service: users enable / disable MFA from their own row in the
  Users page; admins reset (single user or bulk) but never enable on
  behalf of someone else (matches AWS IAM / GitHub / Google Workspace).
* Bulk emergency reset CLI: `scripts/admin-mfa-reset-all.sh`.

Security hardening
==================

* Atomic transactions with `SELECT … FOR UPDATE` on `mfa_pending_logins`
  and `users` rows so concurrent verify / enroll calls cannot race.
* `/api/mfa/enroll/start` refuses re-enrollment when MFA is already on
  (prevents silent secret rotation via a stolen JWT).
* Pydantic `ValidationError` messages are sanitized before reaching the
  audit log so request bodies (TOTP / backup codes in flight) never
  appear in plaintext.
* Slowapi rate limits are per-USER, not per-IP, with a trusted-proxy
  XFF strategy so a single ingress address cannot exhaust the bucket
  for thousands of operators (`MFA_TRUSTED_PROXY_CIDRS`,
  `MFA_RATE_LIMIT_*` env-overridable).
* Login query now scopes to `is_active = TRUE` so a soft-deleted row
  with the same username can no longer occlude the active user
  (also closes a small account-enumeration side channel).

Database
========

Additive migrations (idempotent `ADD COLUMN IF NOT EXISTS`,
`CREATE TABLE IF NOT EXISTS`):

  - users: mfa_enabled, mfa_method, mfa_secret_encrypted,
    mfa_enrolled_at, mfa_last_used_at, mfa_last_used_totp_step
  - mfa_backup_codes (user_id ON DELETE CASCADE)
  - mfa_pending_logins (user_id ON DELETE CASCADE, challenge_token,
    attempts, expires_at)
  - mfa_pending_enrollments (user_id ON DELETE CASCADE)

Frontend
========

* Login page becomes a 3-phase state machine
  (credentials → MFA → submitting); legacy single-step login is
  preserved for users who haven't enrolled.
* New MFAEnrollModal (3-step wizard: QR + secret → verify → backup
  codes) using `qrcode.react`.
* Users page shows MFA column + per-row enable/disable/reset actions.
  Admins viewing other users with MFA off see a non-actionable info
  icon explaining that only the user themselves can enable MFA.

Deployment
==========

* `MFA_ENCRYPTION_KEY` is added to `k8s/manifests/03-secrets.yaml` as
  a placeholder; `SECRET_KEY` is also placeholder-ized so both are
  injected by the existing pipeline pattern (sed-replace + apply).
* No new build-time env vars are required for the frontend. The SPA
  uses `window.location.host` for `/api/*` and is routed by the
  existing nginx ingress configuration.
* `frontend/.dockerignore` ensures host `.env*` files cannot bleed
  into the production bundle.

Tests
=====

* New unit suites:
  - `test_mfa_service.py` (TOTP, encryption, backup codes)
  - `test_mfa_backwards_compat.py` (regression — non-MFA flow unchanged)
  - `test_mfa_rate_limits.py` (env override + dataclass immutability)
  - `test_mfa_rate_limit_key.py` (JWT key, trusted-proxy XFF, fallbacks)
* All existing 1000+ unit tests continue to pass.

Documentation
=============

* README MFA section (overview, day-to-day operations, emergency
  reset CLI, env variables, rate-limit tuning).
* `scripts/README.md` documents the bulk reset script.

Issue: #18
2026-05-19 04:35:16 +03:00
dependabot[bot] 445639d202 chore(deps-dev): bump @babel/plugin-transform-modules-systemjs (#15)
Bumps [@babel/plugin-transform-modules-systemjs](https://github.com/babel/babel/tree/HEAD/packages/babel-plugin-transform-modules-systemjs) from 7.29.0 to 7.29.4.
- [Release notes](https://github.com/babel/babel/releases)
- [Changelog](https://github.com/babel/babel/blob/main/CHANGELOG.md)
- [Commits](https://github.com/babel/babel/commits/v7.29.4/packages/babel-plugin-transform-modules-systemjs)

---
updated-dependencies:
- dependency-name: "@babel/plugin-transform-modules-systemjs"
  dependency-version: 7.29.4
  dependency-type: indirect
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2026-05-14 00:41:27 +03:00
taylanbakircioglu f7e0df15e3 ci(release): stage version.json into backend build context (drift fix)
The backend Docker image is built with `context: ./backend`, so the
repo-root `version.json` is outside the build context and never
reaches the container. Backend `main.py` falls back to a compile-
time `_version_info` constant when `/app/version.json` is missing.

In practice this produced a real production drift: a successful
redeploy of the v1.5.2 tree silently reported `"v1.5.0"` in
`/api/version` for a window of releases because the constant in
main.py had not been bumped in lockstep with `version.json`, and
the canonical file was never available to read inside the
container.

Fix is workflow-only:
  * New "stage version.json into backend build context" step
    (between `read product version` and `set up qemu`) that runs
    `cp version.json backend/version.json` so the next
    `docker buildx build` includes it.
  * `.gitignore` entry for `backend/version.json` keeps `git status`
    clean for developers (the canonical file remains at repo root;
    `backend/version.json` is a transient CI artefact).

Backend reading logic is unchanged: the loop in `main.py` first
tries `/app/version.json`, then falls back to the constant.
Post-fix, the first path WILL find the file and produce the
correct response; the constant becomes a pure defensive fallback
(rather than the production hot path it accidentally became).

No code or test changes needed: existing tests assert against the
`_version_info` dict regardless of whether it was populated from
JSON or the fallback constant.
2026-05-14 00:08:18 +03:00
taylanbakircioglu d7208528f7 fix: v1.5.2 — ACME Diagnostics Panel Hardening + AGPL-3.0 relicense (Bulgu #94/#95/#96)
A focused hardening pass on the v1.5.0 ACME Diagnostic Panel
surface, exercised against a live production deployment (Round-25
+ Round-26 audits) and supplemented by an AGPL-3.0 relicense.

------------------------------------------------------------------
LICENSE — Relicense to AGPL-3.0-or-later
------------------------------------------------------------------
Effective v1.5.2 the project is licensed under the **GNU Affero
General Public License v3.0 (or later)**. v1.5.0 and v1.5.1
remain under the prior MIT terms.

The relicense is consistent with the project's intent as a
community-operated HAProxy management surface: forks that run
HAProxy OpenManager as a network service for third parties are
now required to publish their modifications under the same
license (AGPL §13). Day-to-day single-tenant deployments,
internal corporate use, and ordinary forks-for-fixes are
unaffected.

Changes:
  * LICENSE replaced with full AGPL-3.0 text.
  * README "## License" section rewritten with the AGPL summary
    + the network-service obligation.
  * frontend/package.json gains `"license": "AGPL-3.0-or-later"`.

------------------------------------------------------------------
BULGU #94 / #95 — Diagnostic Panel Must Never Opaque-500
------------------------------------------------------------------
Live exercise of the v1.5.0 Diagnostic Panel against a deployed
build surfaced two opaque-500 paths. The panel exists to make
ACME failures legible; producing an opaque HTTP 500 defeats the
entire feature. Fix shape: every endpoint now returns either a
canonical 4xx (auth / not-found / rate-limit) or an HTTP 200
"structured failure envelope" that the React UI knows how to
render — never a 500 for an in-suite failure.

Affected paths:

POST /api/letsencrypt/orders/{order_id}/diagnostics
  Pre-fix: a UndefinedColumnError or DB-connectivity failure
  inside `run_checks` bubbled out of the bare try/finally and
  surfaced as a generic 500 with no operator-actionable detail.
  Post-fix: setup-stage and run-stage failures are caught
  separately and converted to a `status: diagnostics_unavailable`
  envelope carrying `error_stage`, `error_type`, `error_message`,
  and a `correlation_id` that the operator can grep in the
  backend log. Individual checks are wrapped in `_safe_check`
  so one broken check (e.g. DNS lookup timeout) never crashes
  the suite — the failing check shows up as `status: "fail"`
  with its message, the others still run.

GET /api/letsencrypt/orders/{order_id}/events
  Pre-fix: the SQL `SELECT … status FROM user_activity_logs`
  referenced a column that did not exist in the canonical
  migration; every diagnostic-panel open against an order with
  any user-activity-log correlation got an `UndefinedColumnError`
  500. Post-fix: the endpoint now introspects
  `information_schema.columns` and projects only the columns
  actually present. Partial failures (one source dies, the
  other works) are reported via `meta.errors[]` rather than
  collapsing the whole timeline.

POST /api/letsencrypt/orders/{order_id}/diagnostics/{check_id}/rerun
  Same structured-envelope contract as the full-suite POST,
  scoped to a single check row.

Frontend (`frontend/src/components/ACMEAutomation.js`):
  * Distinct `diagRunError` / `diagEventsError` / `diagMeta`
    states so the modal can render the cause inline (Antd Alert)
    instead of a silent dropdown.
  * Event-log auto-tail polling backs off after 3 consecutive
    failures so the Network tab does not get spammed with 500s
    every 5s.
  * Correlation IDs visible in every error banner.

------------------------------------------------------------------
BULGU #96 — Clean 404 for Out-Of-Range order_id
------------------------------------------------------------------
A live exercise of the post-#94 diagnostic panel against the
deployed build surfaced one remaining contract gap. A path-
param `order_id` outside the Postgres int4 range
(e.g. > 2_147_483_647) caused `_load_order` to raise
`asyncpg.exceptions.DataError: invalid input for query
argument $1: 2147483648 (value out of int32 range)`. Round-25
correctly surfaced this in a `diagnostics_unavailable`
envelope — but that envelope leaked SQL implementation detail
("query argument $1", "int32 range", DataError class name)
into the operator-facing response body.

Semantically an out-of-range integer can never reference a
real order — it's just "not found". `_load_order` now catches
`asyncpg.exceptions.DataError` and re-raises a canonical
`HTTPException(404, "Order {id} not found")`. Because all
three endpoints re-raise `HTTPException` from their outer
try/except (the Round-25 envelope only fires for non-
HTTPException crashes), the canonical 404 path now wins
end-to-end across /diagnostics, /events, and /rerun.

------------------------------------------------------------------
TEST / LINT / LIVE VERIFICATION
------------------------------------------------------------------
  * Backend pytest 1104/1104 (the +20 vs v1.5.1's 1084 are the
    Round-25 and #96 contract pins; see
    test_acme_diagnostics_router_round25.py).
  * Live prod-canary verification: every endpoint return shape
    confirmed against the deployed build — int4 overflow returns
    clean 404 with no SQL leak, normal paths return Round-25
    envelopes, HTTP method matrix returns 405 on wrong verbs,
    no auth returns 401, invalid `check_id` returns 400, and
    `meta.correlation_id` is present on every diagnostic
    response.

------------------------------------------------------------------
COMPATIBILITY
------------------------------------------------------------------
  * No breaking API contract changes: `status` field on the
    diagnostic response can now be `"diagnostics_unavailable"`
    in addition to the existing pass-through of the
    underlying order status (`pending` / `valid` / `invalid` /
    `cancelled` / …) — older UIs that only switch on the
    existing values render the `diagnostics_unavailable`
    case as "unknown status" rather than crashing.
  * Frontend handles the new envelope shape AND the legacy
    HTTP 4xx/5xx paths.
2026-05-14 00:07:00 +03:00
taylanbakircioglu 2e7db4d99f fix: v1.5.1 — Round-23 + Round-24 audit follow-ups (Bulgu #83 → #93)
A live-deployment audit pass over the v1.5.0 Site Wizard + ACME
Diagnostic Panel surface. Two adversarial review rounds (R23, R24)
each capped by an end-to-end smoke test against a multi-cluster
staging deployment.

Bulgu #83 — Frontend Management page warned about stale data
without a clear retry CTA. The toast now carries an in-place
"Reload" action and the page-level Empty state surfaces the same
recovery affordance, so operators never get stuck on a stale-data
view without an obvious way out.

Bulgu #84 — ACME diagnostics ran with the wrong "last_heartbeat"
column reference against the agents table. Aligned the SELECT
with the actual schema column (`last_seen`); pinned by an idempotent
regression test in `test_acme_diagnostics.py`.

Bulgu #85 — ACME order error_detail rendering could leak the raw
asyncpg/SQL exception class name when humanize_error_detail
encountered an unhandled CA response shape. Added a backwards-
compatible fallback branch that emits an "ACME error (raw)" panel
without exposing parse_error class name to the user.

Bulgu #86 — Multi-cluster apply with concurrent rejects could
leave wizard_staged orders dangling without their parent draft.
Pinned via reject_order_with_cluster_orphan test.

Bulgu #87 — Frontend Management page list virtualization
mis-keyed during a re-sort + stale-row replace race; fixed by
keying rows on `id + version` so React reconciler does not reuse
DOM for a logically different row.

Bulgu #88 — Site Wizard "Cancel" mid-flow now surfaces an
unsaved-draft prompt with explicit Save / Discard buttons (and
the same prompt on browser tab close), so the operator never
loses 5 steps of input to an accidental ESC.

Bulgu #89 — Existing-cert SSL mode showed an empty dropdown when
the cluster had >100 certs because the listing endpoint
default-limited results. Endpoint now exposes pagination AND
the wizard switches to client-side filtering above 50 rows.

Bulgu #90 — ACME pre-check on the wizard preview path did NOT
re-validate the account against `letsencrypt_accounts` if the
operator stepped Back/Forward between SSL and Review. Added a
debounced re-validation on Review entry.

Bulgu #93 — Site Wizard hsts_enabled toggle in HTTPS frontend
was idempotent-by-name (the generated `http-response set-header
Strict-Transport-Security` line could duplicate across a Save +
Apply cycle). The renderer now upserts the header in place.

Cumulative outcome: backend pytest 1084/1084, frontend lint
clean, and a 6-hour live-deployment smoke session against staging
with no regressions reported.
2026-05-14 00:06:02 +03:00
45 changed files with 6915 additions and 1145 deletions
+14
View File
@@ -26,6 +26,20 @@ jobs:
fi
echo "VERSION=$VERSION" >> $GITHUB_OUTPUT
# The backend image is built with `context: ./backend`, so the
# repo-root version.json is OUTSIDE the build context and never
# reaches the container. Backend `main.py` falls back to a
# compile-time constant when /app/version.json is missing, which
# caused a real production drift: a redeploy of the v1.5.2 tree
# silently still reported "v1.5.0" in `/api/version` because the
# constant in main.py had been bumped but the file was not
# available to read. Stage version.json into the backend
# context here so the canonical file IS shipped and the
# constant only serves as a defensive fallback. The staged file
# is gitignored to keep `git status` clean for developers.
- name: stage version.json into backend build context
run: cp version.json backend/version.json
- name: set up qemu
uses: docker/setup-qemu-action@v3
+4 -2
View File
@@ -39,8 +39,10 @@ venv.bak/
*.sqlite
*.sqlite3
# Docker
.dockerignore
# Build-time staged version.json (CI `cp version.json backend/`).
# The canonical file lives at repo root; this path is a transient
# copy for the backend Docker build context.
backend/version.json
# IDE
.vscode/
+674 -17
View File
@@ -1,22 +1,679 @@
MIT License
HAProxy OpenManager
Copyright (C) 2025-2026 Taylan Bakırcıoğlu and HAProxy OpenManager Contributors
Copyright (c) 2025 HAProxy OpenManager Contributors
This program is free software: you can redistribute it and/or modify
it under the terms of the GNU Affero General Public License as published
by the Free Software Foundation, either version 3 of the License, or
(at your option) any later version.
Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
in the Software without restriction, including without limitation the rights
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
copies of the Software, and to permit persons to whom the Software is
furnished to do so, subject to the following conditions:
This program is distributed in the hope that it will be useful,
but WITHOUT ANY WARRANTY; without even the implied warranty of
MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
GNU Affero General Public License for more details.
The above copyright notice and this permission notice shall be included in all
copies or substantial portions of the Software.
You should have received a copy of the GNU Affero General Public License
along with this program. If not, see <https://www.gnu.org/licenses/>.
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
SOFTWARE.
---
GNU AFFERO GENERAL PUBLIC LICENSE
Version 3, 19 November 2007
Copyright (C) 2007 Free Software Foundation, Inc. <https://fsf.org/>
Everyone is permitted to copy and distribute verbatim copies
of this license document, but changing it is not allowed.
Preamble
The GNU Affero General Public License is a free, copyleft license for
software and other kinds of works, specifically designed to ensure
cooperation with the community in the case of network server software.
The licenses for most software and other practical works are designed
to take away your freedom to share and change the works. By contrast,
our General Public Licenses are intended to guarantee your freedom to
share and change all versions of a program--to make sure it remains free
software for all its users.
When we speak of free software, we are referring to freedom, not
price. Our General Public Licenses are designed to make sure that you
have the freedom to distribute copies of free software (and charge for
them if you wish), that you receive source code or can get it if you
want it, that you can change the software or use pieces of it in new
free programs, and that you know you can do these things.
Developers that use our General Public Licenses protect your rights
with two steps: (1) assert copyright on the software, and (2) offer
you this License which gives you legal permission to copy, distribute
and/or modify the software.
A secondary benefit of defending all users' freedom is that
improvements made in alternate versions of the program, if they
receive widespread use, become available for other developers to
incorporate. Many developers of free software are heartened and
encouraged by the resulting cooperation. However, in the case of
software used on network servers, this result may fail to come about.
The GNU General Public License permits making a modified version and
letting the public access it on a server without ever releasing its
source code to the public.
The GNU Affero General Public License is designed specifically to
ensure that, in such cases, the modified source code becomes available
to the community. It requires the operator of a network server to
provide the source code of the modified version running there to the
users of that server. Therefore, public use of a modified version, on
a publicly accessible server, gives the public access to the source
code of the modified version.
An older license, called the Affero General Public License and
published by Affero, was designed to accomplish similar goals. This is
a different license, not a version of the Affero GPL, but Affero has
released a new version of the Affero GPL which permits relicensing under
this license.
The precise terms and conditions for copying, distribution and
modification follow.
TERMS AND CONDITIONS
0. Definitions.
"This License" refers to version 3 of the GNU Affero General Public License.
"Copyright" also means copyright-like laws that apply to other kinds of
works, such as semiconductor masks.
"The Program" refers to any copyrightable work licensed under this
License. Each licensee is addressed as "you". "Licensees" and
"recipients" may be individuals or organizations.
To "modify" a work means to copy from or adapt all or part of the work
in a fashion requiring copyright permission, other than the making of an
exact copy. The resulting work is called a "modified version" of the
earlier work or a work "based on" the earlier work.
A "covered work" means either the unmodified Program or a work based
on the Program.
To "propagate" a work means to do anything with it that, without
permission, would make you directly or secondarily liable for
infringement under applicable copyright law, except executing it on a
computer or modifying a private copy. Propagation includes copying,
distribution (with or without modification), making available to the
public, and in some countries other activities as well.
To "convey" a work means any kind of propagation that enables other
parties to make or receive copies. Mere interaction with a user through
a computer network, with no transfer of a copy, is not conveying.
An interactive user interface displays "Appropriate Legal Notices"
to the extent that it includes a convenient and prominently visible
feature that (1) displays an appropriate copyright notice, and (2)
tells the user that there is no warranty for the work (except to the
extent that warranties are provided), that licensees may convey the
work under this License, and how to view a copy of this License. If
the interface presents a list of user commands or options, such as a
menu, a prominent item in the list meets this criterion.
1. Source Code.
The "source code" for a work means the preferred form of the work
for making modifications to it. "Object code" means any non-source
form of a work.
A "Standard Interface" means an interface that either is an official
standard defined by a recognized standards body, or, in the case of
interfaces specified for a particular programming language, one that
is widely used among developers working in that language.
The "System Libraries" of an executable work include anything, other
than the work as a whole, that (a) is included in the normal form of
packaging a Major Component, but which is not part of that Major
Component, and (b) serves only to enable use of the work with that
Major Component, or to implement a Standard Interface for which an
implementation is available to the public in source code form. A
"Major Component", in this context, means a major essential component
(kernel, window system, and so on) of the specific operating system
(if any) on which the executable work runs, or a compiler used to
produce the work, or an object code interpreter used to run it.
The "Corresponding Source" for a work in object code form means all
the source code needed to generate, install, and (for an executable
work) run the object code and to modify the work, including scripts to
control those activities. However, it does not include the work's
System Libraries, or general-purpose tools or generally available free
programs which are used unmodified in performing those activities but
which are not part of the work. For example, Corresponding Source
includes interface definition files associated with source files for
the work, and the source code for shared libraries and dynamically
linked subprograms that the work is specifically designed to require,
such as by intimate data communication or control flow between those
subprograms and other parts of the work.
The Corresponding Source need not include anything that users
can regenerate automatically from other parts of the Corresponding
Source.
The Corresponding Source for a work in source code form is that
same work.
2. Basic Permissions.
All rights granted under this License are granted for the term of
copyright on the Program, and are irrevocable provided the stated
conditions are met. This License explicitly affirms your unlimited
permission to run the unmodified Program. The output from running a
covered work is covered by this License only if the output, given its
content, constitutes a covered work. This License acknowledges your
rights of fair use or other equivalent, as provided by copyright law.
You may make, run and propagate covered works that you do not
convey, without conditions so long as your license otherwise remains
in force. You may convey covered works to others for the sole purpose
of having them make modifications exclusively for you, or provide you
with facilities for running those works, provided that you comply with
the terms of this License in conveying all material for which you do
not control copyright. Those thus making or running the covered works
for you must do so exclusively on your behalf, under your direction
and control, on terms that prohibit them from making any copies of
your copyrighted material outside their relationship with you.
Conveying under any other circumstances is permitted solely under
the conditions stated below. Sublicensing is not allowed; section 10
makes it unnecessary.
3. Protecting Users' Legal Rights From Anti-Circumvention Law.
No covered work shall be deemed part of an effective technological
measure under any applicable law fulfilling obligations under article
11 of the WIPO copyright treaty adopted on 20 December 1996, or
similar laws prohibiting or restricting circumvention of such
measures.
When you convey a covered work, you waive any legal power to forbid
circumvention of technological measures to the extent such circumvention
is effected by exercising rights under this License with respect to
the covered work, and you disclaim any intention to limit operation or
modification of the work as a means of enforcing, against the work's
users, your or third parties' legal rights to forbid circumvention of
technological measures.
4. Conveying Verbatim Copies.
You may convey verbatim copies of the Program's source code as you
receive it, in any medium, provided that you conspicuously and
appropriately publish on each copy an appropriate copyright notice;
keep intact all notices stating that this License and any
non-permissive terms added in accord with section 7 apply to the code;
keep intact all notices of the absence of any warranty; and give all
recipients a copy of this License along with the Program.
You may charge any price or no price for each copy that you convey,
and you may offer support or warranty protection for a fee.
5. Conveying Modified Source Versions.
You may convey a work based on the Program, or the modifications to
produce it from the Program, in the form of source code under the
terms of section 4, provided that you also meet all of these conditions:
a) The work must carry prominent notices stating that you modified
it, and giving a relevant date.
b) The work must carry prominent notices stating that it is
released under this License and any conditions added under section
7. This requirement modifies the requirement in section 4 to
"keep intact all notices".
c) You must license the entire work, as a whole, under this
License to anyone who comes into possession of a copy. This
License will therefore apply, along with any applicable section 7
additional terms, to the whole of the work, and all its parts,
regardless of how they are packaged. This License gives no
permission to license the work in any other way, but it does not
invalidate such permission if you have separately received it.
d) If the work has interactive user interfaces, each must display
Appropriate Legal Notices; however, if the Program has interactive
interfaces that do not display Appropriate Legal Notices, your
work need not make them do so.
A compilation of a covered work with other separate and independent
works, which are not by their nature extensions of the covered work,
and which are not combined with it such as to form a larger program,
in or on a volume of a storage or distribution medium, is called an
"aggregate" if the compilation and its resulting copyright are not
used to limit the access or legal rights of the compilation's users
beyond what the individual works permit. Inclusion of a covered work
in an aggregate does not cause this License to apply to the other
parts of the aggregate.
6. Conveying Non-Source Forms.
You may convey a covered work in object code form under the terms
of sections 4 and 5, provided that you also convey the
machine-readable Corresponding Source under the terms of this License,
in one of these ways:
a) Convey the object code in, or embodied in, a physical product
(including a physical distribution medium), accompanied by the
Corresponding Source fixed on a durable physical medium
customarily used for software interchange.
b) Convey the object code in, or embodied in, a physical product
(including a physical distribution medium), accompanied by a
written offer, valid for at least three years and valid for as
long as you offer spare parts or customer support for that product
model, to give anyone who possesses the object code either (1) a
copy of the Corresponding Source for all the software in the
product that is covered by this License, on a durable physical
medium customarily used for software interchange, for a price no
more than your reasonable cost of physically performing this
conveying of source, or (2) access to copy the
Corresponding Source from a network server at no charge.
c) Convey individual copies of the object code with a copy of the
written offer to provide the Corresponding Source. This
alternative is allowed only occasionally and noncommercially, and
only if you received the object code with such an offer, in accord
with subsection 6b.
d) Convey the object code by offering access from a designated
place (gratis or for a charge), and offer equivalent access to the
Corresponding Source in the same way through the same place at no
further charge. You need not require recipients to copy the
Corresponding Source along with the object code. If the place to
copy the object code is a network server, the Corresponding Source
may be on a different server (operated by you or a third party)
that supports equivalent copying facilities, provided you maintain
clear directions next to the object code saying where to find the
Corresponding Source. Regardless of what server hosts the
Corresponding Source, you remain obligated to ensure that it is
available for as long as needed to satisfy these requirements.
e) Convey the object code using peer-to-peer transmission, provided
you inform other peers where the object code and Corresponding
Source of the work are being offered to the general public at no
charge under subsection 6d.
A separable portion of the object code, whose source code is excluded
from the Corresponding Source as a System Library, need not be
included in conveying the object code work.
A "User Product" is either (1) a "consumer product", which means any
tangible personal property which is normally used for personal, family,
or household purposes, or (2) anything designed or sold for incorporation
into a dwelling. In determining whether a product is a consumer product,
doubtful cases shall be resolved in favor of coverage. For a particular
product received by a particular user, "normally used" refers to a
typical or common use of that class of product, regardless of the status
of the particular user or of the way in which the particular user
actually uses, or expects or is expected to use, the product. A product
is a consumer product regardless of whether the product has substantial
commercial, industrial or non-consumer uses, unless such uses represent
the only significant mode of use of the product.
"Installation Information" for a User Product means any methods,
procedures, authorization keys, or other information required to install
and execute modified versions of a covered work in that User Product from
a modified version of its Corresponding Source. The information must
suffice to ensure that the continued functioning of the modified object
code is in no case prevented or interfered with solely because
modification has been made.
If you convey an object code work under this section in, or with, or
specifically for use in, a User Product, and the conveying occurs as
part of a transaction in which the right of possession and use of the
User Product is transferred to the recipient in perpetuity or for a
fixed term (regardless of how the transaction is characterized), the
Corresponding Source conveyed under this section must be accompanied
by the Installation Information. But this requirement does not apply
if neither you nor any third party retains the ability to install
modified object code on the User Product (for example, the work has
been installed in ROM).
The requirement to provide Installation Information does not include a
requirement to continue to provide support service, warranty, or updates
for a work that has been modified or installed by the recipient, or for
the User Product in which it has been modified or installed. Access to a
network may be denied when the modification itself materially and
adversely affects the operation of the network or violates the rules and
protocols for communication across the network.
Corresponding Source conveyed, and Installation Information provided,
in accord with this section must be in a format that is publicly
documented (and with an implementation available to the public in
source code form), and must require no special password or key for
unpacking, reading or copying.
7. Additional Terms.
"Additional permissions" are terms that supplement the terms of this
License by making exceptions from one or more of its conditions.
Additional permissions that are applicable to the entire Program shall
be treated as though they were included in this License, to the extent
that they are valid under applicable law. If additional permissions
apply only to part of the Program, that part may be used separately
under those permissions, but the entire Program remains governed by
this License without regard to the additional permissions.
When you convey a copy of a covered work, you may at your option
remove any additional permissions from that copy, or from any part of
it. (Additional permissions may be written to require their own
removal in certain cases when you modify the work.) You may place
additional permissions on material, added by you to a covered work,
for which you have or can give appropriate copyright permission.
Notwithstanding any other provision of this License, for material you
add to a covered work, you may (if authorized by the copyright holders of
that material) supplement the terms of this License with terms:
a) Disclaiming warranty or limiting liability differently from the
terms of sections 15 and 16 of this License; or
b) Requiring preservation of specified reasonable legal notices or
author attributions in that material or in the Appropriate Legal
Notices displayed by works containing it; or
c) Prohibiting misrepresentation of the origin of that material, or
requiring that modified versions of such material be marked in
reasonable ways as different from the original version; or
d) Limiting the use for publicity purposes of names of licensors or
authors of the material; or
e) Declining to grant rights under trademark law for use of some
trade names, trademarks, or service marks; or
f) Requiring indemnification of licensors and authors of that
material by anyone who conveys the material (or modified versions of
it) with contractual assumptions of liability to the recipient, for
any liability that these contractual assumptions directly impose on
those licensors and authors.
All other non-permissive additional terms are considered "further
restrictions" within the meaning of section 10. If the Program as you
received it, or any part of it, contains a notice stating that it is
governed by this License along with a term that is a further
restriction, you may remove that term. If a license document contains
a further restriction but permits relicensing or conveying under this
License, you may add to a covered work material governed by the terms
of that license document, provided that the further restriction does
not survive such relicensing or conveying.
If you add terms to a covered work in accord with this section, you
must place, in the relevant source files, a statement of the
additional terms that apply to those files, or a notice indicating
where to find the applicable terms.
Additional terms, permissive or non-permissive, may be stated in the
form of a separately written license, or stated as exceptions;
the above requirements apply either way.
8. Termination.
You may not propagate or modify a covered work except as expressly
provided under this License. Any attempt otherwise to propagate or
modify it is void, and will automatically terminate your rights under
this License (including any patent licenses granted under the third
paragraph of section 11).
However, if you cease all violation of this License, then your
license from a particular copyright holder is reinstated (a)
provisionally, unless and until the copyright holder explicitly and
finally terminates your license, and (b) permanently, if the copyright
holder fails to notify you of the violation by some reasonable means
prior to 60 days after the cessation.
Moreover, your license from a particular copyright holder is
reinstated permanently if the copyright holder notifies you of the
violation by some reasonable means, this is the first time you have
received notice of violation of this License (for any work) from that
copyright holder, and you cure the violation prior to 30 days after
your receipt of the notice.
Termination of your rights under this section does not terminate the
licenses of parties who have received copies or rights from you under
this License. If your rights have been terminated and not permanently
reinstated, you do not qualify to receive new licenses for the same
material under section 10.
9. Acceptance Not Required for Having Copies.
You are not required to accept this License in order to receive or
run a copy of the Program. Ancillary propagation of a covered work
occurring solely as a consequence of using peer-to-peer transmission
to receive a copy likewise does not require acceptance. However,
nothing other than this License grants you permission to propagate or
modify any covered work. These actions infringe copyright if you do
not accept this License. Therefore, by modifying or propagating a
covered work, you indicate your acceptance of this License to do so.
10. Automatic Licensing of Downstream Recipients.
Each time you convey a covered work, the recipient automatically
receives a license from the original licensors, to run, modify and
propagate that work, subject to this License. You are not responsible
for enforcing compliance by third parties with this License.
An "entity transaction" is a transaction transferring control of an
organization, or substantially all assets of one, or subdividing an
organization, or merging organizations. If propagation of a covered
work results from an entity transaction, each party to that
transaction who receives a copy of the work also receives whatever
licenses to the work the party's predecessor in interest had or could
give under the previous paragraph, plus a right to possession of the
Corresponding Source of the work from the predecessor in interest, if
the predecessor has it or can get it with reasonable efforts.
You may not impose any further restrictions on the exercise of the
rights granted or affirmed under this License. For example, you may
not impose a license fee, royalty, or other charge for exercise of
rights granted under this License, and you may not initiate litigation
(including a cross-claim or counterclaim in a lawsuit) alleging that
any patent claim is infringed by making, using, selling, offering for
sale, or importing the Program or any portion of it.
11. Patents.
A "contributor" is a copyright holder who authorizes use under this
License of the Program or a work on which the Program is based. The
work thus licensed is called the contributor's "contributor version".
A contributor's "essential patent claims" are all patent claims
owned or controlled by the contributor, whether already acquired or
hereafter acquired, that would be infringed by some manner, permitted
by this License, of making, using, or selling its contributor version,
but do not include claims that would be infringed only as a
consequence of further modification of the contributor version. For
purposes of this definition, "control" includes the right to grant
patent sublicenses in a manner consistent with the requirements of
this License.
Each contributor grants you a non-exclusive, worldwide, royalty-free
patent license under the contributor's essential patent claims, to
make, use, sell, offer for sale, import and otherwise run, modify and
propagate the contents of its contributor version.
In the following three paragraphs, a "patent license" is any express
agreement or commitment, however denominated, not to enforce a patent
(such as an express permission to practice a patent or covenant not to
sue for patent infringement). To "grant" such a patent license to a
party means to make such an agreement or commitment not to enforce a
patent against the party.
If you convey a covered work, knowingly relying on a patent license,
and the Corresponding Source of the work is not available for anyone
to copy, free of charge and under the terms of this License, through a
publicly available network server or other readily accessible means,
then you must either (1) cause the Corresponding Source to be so
available, or (2) arrange to deprive yourself of the benefit of the
patent license for this particular work, or (3) arrange, in a manner
consistent with the requirements of this License, to extend the patent
license to downstream recipients. "Knowingly relying" means you have
actual knowledge that, but for the patent license, your conveying the
covered work in a country, or your recipient's use of the covered work
in a country, would infringe one or more identifiable patents in that
country that you have reason to believe are valid.
If, pursuant to or in connection with a single transaction or
arrangement, you convey, or propagate by procuring conveyance of, a
covered work, and grant a patent license to some of the parties
receiving the covered work authorizing them to use, propagate, modify
or convey a specific copy of the covered work, then the patent license
you grant is automatically extended to all recipients of the covered
work and works based on it.
A patent license is "discriminatory" if it does not include within
the scope of its coverage, prohibits the exercise of, or is
conditioned on the non-exercise of one or more of the rights that are
specifically granted under this License. You may not convey a covered
work if you are a party to an arrangement with a third party that is
in the business of distributing software, under which you make payment
to the third party based on the extent of your activity of conveying
the work, and under which the third party grants, to any of the
parties who would receive the covered work from you, a discriminatory
patent license (a) in connection with copies of the covered work
conveyed by you (or copies made from those copies), or (b) primarily
for and in connection with specific products or compilations that
contain the covered work, unless you entered into that arrangement,
or that patent license was granted, prior to 28 March 2007.
Nothing in this License shall be construed as excluding or limiting
any implied license or other defenses to infringement that may
otherwise be available to you under applicable patent law.
12. No Surrender of Others' Freedom.
If conditions are imposed on you (whether by court order, agreement or
otherwise) that contradict the conditions of this License, they do not
excuse you from the conditions of this License. If you cannot convey a
covered work so as to satisfy simultaneously your obligations under this
License and any other pertinent obligations, then as a consequence you may
not convey it at all. For example, if you agree to terms that obligate you
to collect a royalty for further conveying from those to whom you convey
the Program, the only way you could satisfy both those terms and this
License would be to refrain entirely from conveying the Program.
13. Remote Network Interaction; Use with the GNU General Public License.
Notwithstanding any other provision of this License, if you modify the
Program, your modified version must prominently offer all users
interacting with it remotely through a computer network (if your version
supports such interaction) an opportunity to receive the Corresponding
Source of your version by providing access to the Corresponding Source
from a network server at no charge, through some standard or customary
means of facilitating copying of software. This Corresponding Source
shall include the Corresponding Source for any work covered by version 3
of the GNU General Public License that is incorporated pursuant to the
following paragraph.
Notwithstanding any other provision of this License, you have
permission to link or combine any covered work with a work licensed
under version 3 of the GNU General Public License into a single
combined work, and to convey the resulting work. The terms of this
License will continue to apply to the part which is the covered work,
but the work with which it is combined will remain governed by version
3 of the GNU General Public License.
14. Revised Versions of this License.
The Free Software Foundation may publish revised and/or new versions of
the GNU Affero General Public License from time to time. Such new versions
will be similar in spirit to the present version, but may differ in detail to
address new problems or concerns.
Each version is given a distinguishing version number. If the
Program specifies that a certain numbered version of the GNU Affero General
Public License "or any later version" applies to it, you have the
option of following the terms and conditions either of that numbered
version or of any later version published by the Free Software
Foundation. If the Program does not specify a version number of the
GNU Affero General Public License, you may choose any version ever published
by the Free Software Foundation.
If the Program specifies that a proxy can decide which future
versions of the GNU Affero General Public License can be used, that proxy's
public statement of acceptance of a version permanently authorizes you
to choose that version for the Program.
Later license versions may give you additional or different
permissions. However, no additional obligations are imposed on any
author or copyright holder as a result of your choosing to follow a
later version.
15. Disclaimer of Warranty.
THERE IS NO WARRANTY FOR THE PROGRAM, TO THE EXTENT PERMITTED BY
APPLICABLE LAW. EXCEPT WHEN OTHERWISE STATED IN WRITING THE COPYRIGHT
HOLDERS AND/OR OTHER PARTIES PROVIDE THE PROGRAM "AS IS" WITHOUT WARRANTY
OF ANY KIND, EITHER EXPRESSED OR IMPLIED, INCLUDING, BUT NOT LIMITED TO,
THE IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR
PURPOSE. THE ENTIRE RISK AS TO THE QUALITY AND PERFORMANCE OF THE PROGRAM
IS WITH YOU. SHOULD THE PROGRAM PROVE DEFECTIVE, YOU ASSUME THE COST OF
ALL NECESSARY SERVICING, REPAIR OR CORRECTION.
16. Limitation of Liability.
IN NO EVENT UNLESS REQUIRED BY APPLICABLE LAW OR AGREED TO IN WRITING
WILL ANY COPYRIGHT HOLDER, OR ANY OTHER PARTY WHO MODIFIES AND/OR CONVEYS
THE PROGRAM AS PERMITTED ABOVE, BE LIABLE TO YOU FOR DAMAGES, INCLUDING ANY
GENERAL, SPECIAL, INCIDENTAL OR CONSEQUENTIAL DAMAGES ARISING OUT OF THE
USE OR INABILITY TO USE THE PROGRAM (INCLUDING BUT NOT LIMITED TO LOSS OF
DATA OR DATA BEING RENDERED INACCURATE OR LOSSES SUSTAINED BY YOU OR THIRD
PARTIES OR A FAILURE OF THE PROGRAM TO OPERATE WITH ANY OTHER PROGRAMS),
EVEN IF SUCH HOLDER OR OTHER PARTY HAS BEEN ADVISED OF THE POSSIBILITY OF
SUCH DAMAGES.
17. Interpretation of Sections 15 and 16.
If the disclaimer of warranty and limitation of liability provided
above cannot be given local legal effect according to their terms,
reviewing courts shall apply local law that most closely approximates
an absolute waiver of all civil liability in connection with the
Program, unless a warranty or assumption of liability accompanies a
copy of the Program in return for a fee.
END OF TERMS AND CONDITIONS
How to Apply These Terms to Your New Programs
If you develop a new program, and you want it to be of the greatest
possible use to the public, the best way to achieve this is to make it
free software which everyone can redistribute and change under these terms.
To do so, attach the following notices to the program. It is safest
to attach them to the start of each source file to most effectively
state the exclusion of warranty; and each file should have at least
the "copyright" line and a pointer to where the full notice is found.
<one line to give the program's name and a brief idea of what it does.>
Copyright (C) <year> <name of author>
This program is free software: you can redistribute it and/or modify
it under the terms of the GNU Affero General Public License as published by
the Free Software Foundation, either version 3 of the License, or
(at your option) any later version.
This program is distributed in the hope that it will be useful,
but WITHOUT ANY WARRANTY; without even the implied warranty of
MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
GNU Affero General Public License for more details.
You should have received a copy of the GNU Affero General Public License
along with this program. If not, see <https://www.gnu.org/licenses/>.
Also add information on how to contact you by electronic and paper mail.
If your software can interact with users remotely through a computer
network, you should also make sure that it provides a way for users to
get its source. For example, if your program is a web application, its
interface could display a "Source" link that leads users to an archive
of the code. There are many ways you could offer source, and different
solutions will be better for different programs; see section 13 for the
specific requirements.
You should also get your employer (if you work as a programmer) or school,
if any, to sign a "copyright disclaimer" for the program, if necessary.
For more information on this, and how to apply and follow the GNU AGPL, see
<https://www.gnu.org/licenses/>.
+175 -744
View File
@@ -17,6 +17,7 @@ Modern, web-based management interface for HAProxy load balancers with multi-clu
- [Apply Management & Version Control](#apply-management--version-control)
- [Security & Certificate Management](#security--certificate-management)
- [ACME Automation](#acme-automation---automated-ssl-certificates)
- [Site Wizard - Guided Multi-Step Host Setup](#site-wizard---guided-multi-step-host-setup)
- [WAF Management](#waf-management)
- [IP Inventory](#ip-inventory---cross-cluster-ip-search)
3. [Key Capabilities](#key-capabilities)
@@ -36,6 +37,7 @@ Modern, web-based management interface for HAProxy load balancers with multi-clu
- [Agent Management](#agent-management---haproxy-agent-deployment--monitoring)
- [Dashboard](#dashboard---main-overview--real-time-monitoring)
- [Frontend Management](#frontend-management---virtual-host--routing-configuration)
- [New Site Wizard](#new-site-wizard---guided-multi-step-host-setup)
- [Backend Servers](#backend-servers---server-pool-management)
- [Configuration](#configuration---haproxy-config-file-management)
- [Apply Management](#apply-management---change-tracking--deployment)
@@ -97,11 +99,13 @@ This architecture provides better security (no inbound connections to HAProxy se
✅ **Agent-Based Pull Architecture** - Secure, scalable management without inbound connections
✅ **Multi-Cluster & Pool Management** - Organize and manage multiple HAProxy clusters from one interface
✅ **Frontend/Backend/Server CRUD** - Complete entity management with visual UI
✅ **New Site Wizard** - Step-by-step guided setup for a complete proxied host (frontend + backend + servers + SSL/ACME) in one consolidated atomic apply
✅ **Bulk Config Import** - Import existing `haproxy.cfg` files with smart SSL auto-assignment
✅ **Version Control & Rollback** - Every change versioned with one-click restore capability
✅ **Real-Time Monitoring** - Live stats, health checks, and performance dashboards
✅ **SSL Certificate Management** - Centralized SSL with expiration tracking
✅ **ACME Auto SSL (Let's Encrypt)** - Automated certificate issuance, renewal, and deployment via ACME protocol
✅ **ACME Certificate Diagnostic Panel** - Automated preflight that checks agent readiness, DNS resolution, port 80 reachability, and ACME challenge ACL before issuing certificates
✅ **WAF Rules** - Web Application Firewall management and deployment
✅ **Agent Script Versioning** - Update agents via UI (Monaco editor) with auto-upgrade
✅ **Token-Based Agent Auth** - Secure token management with revoke/renew
@@ -192,6 +196,11 @@ This architecture provides better security (no inbound connections to HAProxy se
**ACME Settings** - ACME/SSL Automation configuration with provider selection, staging mode, and auto-renewal settings
![ACME Settings](docs/screenshots/acme-settings.png)
### Site Wizard - Guided Multi-Step Host Setup
**New Site Wizard** - Multi-step guided form for creating a complete proxied host (domains + backend pool + servers + frontend + SSL/ACME) with live HAProxy config validation and atomic apply
![New Site Wizard](docs/screenshots/site-wizard.png)
### WAF Management
**WAF Management** - WAF rule configuration with request filtering, rate limiting, and advanced options
@@ -638,6 +647,23 @@ The dashboard displays comprehensive real-time metrics collected by agents from
- **Advanced Options**: Connection limits, timeouts, and performance tuning
- **Search & Filter**: Real-time search and status-based filtering
### New Site Wizard - Guided Multi-Step Host Setup
The New Site Wizard is the recommended entry point for creating a complete proxied host. It collects every piece of configuration a host needs across a five-step guided form and creates the matching frontend + backend + servers + SSL binding as a single atomic apply.
- **5-step guided flow**:
1. **Domains** - host header / SNI / wildcard validation, IDN/Punycode normalization, port + bind-address collision check
2. **Backend** - pool name, load-balance algorithm, health check method/URI, cookie-based persistence (RFC 6265 and HAProxy-parser-safe character validation)
3. **Servers** - per-server address/port/weight, optional CA bundle reference (cluster-RBAC enforced), backup-server + cookie-value validators
4. **SSL** - three modes: existing certificate (cluster-RBAC), PEM upload (chain validation + SAN/CN match), or ACME order (account binding + preflight)
5. **Review + Dry-Run** - live `POST /api/sites/preview` runs the proposed config through `haproxy -c` and surfaces every WARNING/ERROR with a marker comment that pinpoints the offending block before anything touches the database
- **Draft persistence**: every step auto-saves to a server-side draft (30-day TTL, 50 drafts per user cap, cluster-scoped, PEM material stripped at rest)
- **Resume across sessions**: drafts can be reopened from a different browser; cluster swap mid-wizard surfaces a confirmation prompt to prevent cross-cluster contamination
- **Atomic apply**: frontend + backend + servers + SSL binding land as a single PENDING change reviewable from Apply Management - accept and apply, or reject as a whole
- **Cluster RBAC**: every step honours `user_pool_access` and cluster RBAC; SSL certificate references are validated against the target cluster at both preview AND create time
- **Hardening**: IDN/Punycode safe slug generation, cookie/header injection guards, path-traversal protection on SSL filenames, per-user-per-minute rate limit on create + preview, advisory locks against concurrent creation
- **Backwards compatible URLs**: legacy `/proxied-hosts/new` route continues to resolve to the new wizard
### Backend Servers - Server Pool Management
- **Server Management**: Add, edit, remove, and configure backend servers
- **Health Checks**: HTTP/TCP health check configuration and monitoring
@@ -837,6 +863,20 @@ sequenceDiagram
**Key behavior**: Auto-renewed certificates follow the **exact same Apply pipeline** as manual SSL updates. The cert transitions `PENDING -> APPLIED` automatically, agents are notified, and cross-cluster propagation works identically to manual Apply. No separate deployment mechanism is used.
#### ACME Certificate Diagnostic Panel
Available from the ACME Automation list (`Diagnose` button on each order, or by clicking the order's status tag), the Diagnostic Panel runs a full preflight checklist before a certificate is issued or renewed and surfaces every blocker in a single operator-friendly view:
- **Agent reachability** - verifies that at least one agent in each target cluster is online and within the last-seen window
- **DNS resolution** - resolves every domain on the order (A / AAAA / CNAME) and flags wildcards and non-resolvable names with the exact upstream resolver error
- **Port 80 reachability** - checks that the HTTP-01 challenge port is reachable from the public internet via the agent's egress path
- **ACME challenge ACL preview** - renders the `acl1 !acl1` + `http-request return` block that will be injected at apply time, so the operator can see exactly what HAProxy will receive
- **HSTS / rate-limit policy collision check** - warns if existing frontend rules would short-circuit the challenge route
- **Account binding check** - confirms a valid ACME account exists for the selected provider + cluster combination
- **Event Log timeline** - merged view of typed `acme_order_events` rows and correlated `user_activity_logs` entries; auto-tails every 5s while the order is in `pending`/`processing`
- **Humanized error display** - covers 11+ RFC8555 problem types (`badNonce`, `caa`, `connection`, `rateLimited`, `unauthorized`, ...) with full backwards compatibility for legacy plain-string error details
- **Operator-friendly error envelope** - every failure returns `{ correlation_id, field, cause, remediation }` so root-causing is one API call away
#### ACME Features
| Feature | Description |
@@ -1025,6 +1065,115 @@ The IP Inventory page provides a unified view of all IP addresses across every c
- **API Keys**: User API key generation and management
- **Role Assignment**: Dynamic role assignment and permission updates
#### Multi-Factor Authentication (MFA) — v1.6.0 (Issue #18)
MFA is **optional per account** and **default OFF**. Existing users keep their
single-factor (username/password) login unless they choose to enable it. The
feature is fully additive: nothing changes for accounts that don't opt in.
**For end-users**
- Open **Users** → find your own row → click **Enable MFA**.
- A wizard opens with three steps:
1. **Set up** — scan the QR code with Google Authenticator / Authy /
1Password / Microsoft Authenticator, or paste the displayed secret manually.
2. **Verify** — enter the current 6-digit code from your app.
3. **Backup codes** — save the 10 single-use recovery codes (format
`XXXX-YYYY`). They are shown only once. Use the **Copy all** /
**Download .txt** buttons.
- After enrollment your sign-in becomes two-step: username/password →
6-digit TOTP (or a backup code).
- To turn MFA off again, open your row's **Disable MFA** action and enter a
current TOTP or backup code.
**For admins**
- On any user with MFA enabled, the **Reset MFA** action wipes the user's
TOTP secret, backup codes, and pending challenges. A required *reason* is
written to the audit log. After reset the user logs in with their password
and may re-enroll.
**Emergency: reset MFA for every user**
Use the CLI helper when an authenticator outage / mass key loss happens.
Double confirmation is required; the action is irreversible.
```bash
API_URL=https://hap.example.com ADMIN_TOKEN=eyJ... \
./scripts/admin-mfa-reset-all.sh
# Prompts ask for: 'yes' → 'RESET ALL MFA' → reason
# Audit log: action='mfa.disabled.admin_bulk_reset'
```
**Configuration**
- Backend env var `MFA_ENCRYPTION_KEY` — a 44-char Fernet key used to encrypt
TOTP secrets at rest. Generate with:
```bash
python3 -c "from cryptography.fernet import Fernet; print(Fernet.generate_key().decode())"
```
On Kubernetes the key lives in the `backend-secret` Secret
(`k8s/manifests/03-secrets.yaml`). The shipped manifest uses the placeholder
`mfa_encryption_key_replace_me`; replace it via your CI/CD pipeline (e.g.
`sed` step) before `kubectl apply`.
- Optional `MFA_ACCOUNT_LABEL_DOMAIN` — overrides the per-user otpauth label
domain so QR codes show e.g. `alice@hap.example.com` instead of the request
hostname.
**API endpoints (all additive)**
| Method | Path | Notes |
|---|---|---|
| `POST` | `/api/auth/login` | Returns `mfa_required:true`+`mfa_token` for MFA users; legacy shape otherwise |
| `POST` | `/api/auth/login/mfa-verify` | Submits TOTP or backup code; returns JWT |
| `GET` | `/api/mfa/status` | Self status (enabled / method / backup codes remaining) |
| `POST` | `/api/mfa/enroll/start` | Begin TOTP enrollment |
| `POST` | `/api/mfa/enroll/confirm` | Confirm enrollment, returns 10 backup codes once |
| `POST` | `/api/mfa/disable` | Self-disable (TOTP or backup required) |
| `POST` | `/api/mfa/backup-codes/regenerate` | Issue 10 fresh backup codes (TOTP only) |
| `GET` | `/api/mfa/admin/status/{id}` | Admin: any user's MFA status |
| `POST` | `/api/mfa/admin-reset/{id}` | Admin: reset a single user's MFA |
| `POST` | `/api/mfa/admin-reset-all` | Admin: emergency reset for all users |
**Rate limits (per-user, ingress-aware, operationally tunable)**
MFA endpoints are rate-limited via slowapi+Redis with a **user-aware key
function** (`backend/middleware/mfa_rate_limit_key.py`):
1. If the request carries a valid Bearer JWT → bucket is `user:<id>`.
Each operator gets an isolated bucket; an org-wide MFA rollout is no
longer bottlenecked by the shared ingress IP.
2. Else, if the TCP peer is in `MFA_TRUSTED_PROXY_CIDRS` → bucket is the
first `X-Forwarded-For` hop (real client IP behind the ingress).
3. Else → bucket is the TCP peer (slowapi default).
Defaults live in `backend/middleware/mfa_rate_limits.py` and are sized for
**enterprise-scale** rollouts (thousands of operators). Each one is
overridable via env var; an invalid string logs a `WARNING` and falls back
to the default without crashing.
| Endpoint | Env var | Default | Bucket |
|---|---|---|---|
| `POST /api/mfa/enroll/start` | `MFA_RATE_LIMIT_ENROLL_START` | `10/minute` | per user |
| `POST /api/mfa/enroll/confirm` | `MFA_RATE_LIMIT_ENROLL_CONFIRM` | `10/minute` | per user |
| `POST /api/mfa/disable` | `MFA_RATE_LIMIT_DISABLE` | `10/minute` | per user |
| `POST /api/mfa/backup-codes/regenerate` | `MFA_RATE_LIMIT_REGENERATE_BACKUP_CODES` | `5/hour` | per user |
| `POST /api/mfa/admin-reset/{id}` | `MFA_RATE_LIMIT_ADMIN_RESET` | `60/hour` | per admin |
| `POST /api/mfa/admin-reset-all` | `MFA_RATE_LIMIT_ADMIN_RESET_ALL` | `1/day` | per admin |
Limit string format follows slowapi: `<count>/<second|minute|hour|day>`.
**Trusted-proxy configuration**
`MFA_TRUSTED_PROXY_CIDRS` — comma-separated CIDR list, e.g.
`10.0.0.0/8,172.16.0.0/12,192.168.0.0/16`. Empty (default) disables XFF
parsing — XFF from any peer is then ignored, which is the safe choice when
the topology is unknown. Set this when your backend sits behind a known
ingress / load balancer so anonymous flows still get per-real-IP buckets.
The pre-existing `/api/auth/login` rate-limiting policy is unchanged. On
Kubernetes, see commented overrides in
`k8s/manifests/07-configmaps.yaml::backend-config`.
### Settings - System Configuration
- **Theme Settings**: Light/dark mode toggle and UI customization
- **ACME / SSL Automation**: Configure ACME provider, directory URL, staging mode, auto-renewal, EAB credentials, and test CA connectivity
@@ -1484,12 +1633,25 @@ AGENT_CONFIG_SYNC_INTERVAL_SECONDS=30
```
#### Frontend Configuration
```bash
# API Endpoint (auto-detected if empty)
# Leave empty in production to use same-origin (window.location)
REACT_APP_API_URL="" # For development: "http://localhost:8000"
# Environment
The frontend is a Create-React-App single-page app served as a static bundle
(`serve -s build`). It uses **same-origin** (`window.location.host`) for all
`/api/*` calls — no env vars are needed in production. Routing is handled
entirely by the nginx reverse proxy in front of the frontend pod (Kubernetes
ingress + `nginx-config` ConfigMap, or `nginx/nginx.conf` for Docker Compose).
> ⚠️ **Do NOT set `REACT_APP_API_URL` in your CI/CD pipeline.** CRA inlines
> `REACT_APP_*` values into the bundle at **build time**, so any value baked
> in there overrides the runtime same-origin detection and breaks every
> deployment whose URL does not match the inlined string. Leave the variable
> unset; the bundle will resolve to whatever host the user is browsing.
```bash
# Optional, only when you intentionally need a cross-origin API
# (then CORS_ORIGINS on the backend must include the SPA's origin):
# REACT_APP_API_URL="https://api.example.com"
# Build settings
NODE_ENV="production"
GENERATE_SOURCEMAP="false"
```
@@ -2169,7 +2331,13 @@ journalctl -u keepalived
## License
This project is licensed under the MIT License - see the [LICENSE](LICENSE) file for details.
This project is licensed under the **GNU Affero General Public License v3.0 (or later)** — see the [LICENSE](LICENSE) file for the full text.
A short summary:
- You are free to use, modify, and distribute this software.
- If you run a modified version as a network service (e.g., SaaS), you must make the modified source available to its users (AGPL §13).
- Any redistribution or derivative work must remain under AGPL-3.0-or-later.
## Author
@@ -2199,744 +2367,7 @@ Developed with ❤️ for the HAProxy community
## Release Notes
### v1.5.0 — ACME Diagnostics & Site Wizard
#### Highlights
- **Issue #13: ACME Diagnostic Panel.** From the ACME Automation list, click the new `Diagnose` button (or the order's status tag) to launch a Modal with three Tabs:
1. **Pre-flight Checks** — DNS resolution, port-80 reachability, HAProxy routing, ACME account status, agent health. Each check has its own `Re-run` button.
2. **Event Log** — Merged timeline of typed `acme_order_events` rows (added in v1.5.0) and correlated `user_activity_logs` entries; auto-tails every 5s while the order is in `pending`/`processing`.
3. **Raw Error** — Humanized error display covering 11+ RFC8555 problem types (`badNonce`, `caa`, `connection`, `rateLimited`, `unauthorized`, ...) with full backwards compatibility for legacy plain-string `error_detail`.
- **Issue #14: New Site Setup Wizard.** A single guided flow (`/sites/new`) that creates a Backend + Servers + HTTP Frontend (and optional HTTPS Frontend) in one atomic transaction. SSL choice supports four modes:
- `acme` — defers HTTPS frontend creation to a deferred `post_completion_actions` block on a wizard-staged ACME order; the order is promoted to a real LE call only after the agent confirms the gating `bulk-site-create-{ts}` config version (legacy `bulk-proxied-host-create-{ts}` is still recognised by the reject path for historical APPLIED versions).
- `upload` — uploads PEM cert+key in the same transaction.
- `existing` — reuses an admin-uploaded cert and ensures the cluster junction is set.
- `none` — HTTP-only host.
The wizard ships with the same visual ACL rule builder used by the standalone Frontend Management page (Routing & ACLs section on the Frontend step) so operators define `acl` / `use_backend` / `redirect` rules from cluster-scoped backend dropdowns instead of free-text HAProxy directives. Drafts persist for 30 days with PEM material stripped at rest. Reject of the wizard's PENDING version cleanly rolls back ALL wizard entities (backends, servers, frontend(s), SSL row if any, AND the wizard-staged `letsencrypt_orders` row).
#### Migration Release Notes
This release adds **idempotent** migrations only — no destructive schema changes:
- New columns on `letsencrypt_orders`:
- `post_completion_actions JSONB` (deferred actions for wizard ACME mode)
- `wizard_staged_until TIMESTAMPTZ` (24h timeout for wizard-staged orders)
- `pending_apply_version_name VARCHAR(255)` + partial index `WHERE status='wizard_staged'`
- `created_by INTEGER REFERENCES users(id) ON DELETE SET NULL`
- New tables: `acme_order_events` (typed event log, 90d retention), `wizard_drafts` (30d retention).
- New composite index `idx_user_activity_logs_user_action_time` for the per-user-per-minute rate-limit COUNT(*) used by both new features.
The `letsencrypt_orders.status` column has no CHECK constraint; the new `wizard_staged` value coexists with all existing statuses (`pending`, `ready`, `processing`, `valid`, `invalid`, ...).
The reject path's force-delete fallback now also covers `entity_type='letsencrypt_order'` snapshots so wizard-staged ACME orders are removed when their parent PENDING version is rejected.
#### Rollback Considerations
- **Forward compatibility (v1.5.0 → future).** All new columns/tables are additive; older code paths that do not know about them are unaffected.
- **Backward rollback (v1.5.0 → v1.4.0).** The new columns/tables remain in the database harmlessly; v1.4.0 simply ignores them. Wizard-staged ACME orders that were never promoted to `pending` will not progress on v1.4.0 (the v1.4.0 background task does not select `status='wizard_staged'`); admins can either:
1. Wait for the 24h `wizard_staged_until` timeout to fire (v1.5.0 only) — only relevant if rolling back temporarily, OR
2. Manually `DELETE FROM letsencrypt_orders WHERE status='wizard_staged'` and re-run the wizard once you re-deploy v1.5.0.
- **Wizard ACME failure scenarios.** If the agent never confirms the gating config version (e.g. agent down), the wizard-staged ACME order will time out and transition to `status='invalid'` after 24h with `error_detail='wizard staged timeout (>24h with no agent confirm)'` — surfaced in the new Diagnostic Panel.
- **Drafts.** PEM material is server-side stripped from `wizard_drafts.payload`; rolling back will not leak keys at rest.
#### v1.5.x — Site Wizard module rename + ACL UX parity
A non-breaking follow-up to v1.5.0 that retires the internal "Proxied Host" namespace in favour of "Site" everywhere it used to leak into operators' workflow:
- **Module file rename.** `backend/routers/proxied_host.py` and `backend/models/proxied_host.py` are now `site_wizard.py`. The Pydantic class `ProxiedHostCreate` (and its sibling `ProxiedHostPreflightAcme` / `ProxiedHostDraftCreate`) was renamed to `SiteCreate` etc. with a module-level alias `ProxiedHostCreate = SiteCreate` so existing imports keep working.
- **API URL prefix rename.** The wizard now mounts at `/api/sites/*` (canonical). The legacy `/api/proxied-hosts/*` slug is preserved as a hidden `308 Permanent Redirect` alias on `main.py`, so external integrators keep working through the redirect during the transition window. The frontend axios calls all target `/api/sites/*` directly.
- **Audit-log version-name rename.** Wizard-applied versions now carry the prefix `bulk-site-create-{ts}`. The cluster reject path on `cluster.py` recognises BOTH the new prefix and the legacy `bulk-proxied-host-create-{ts}` so historical APPLIED versions still clean up.
- **Activity-log action + resource_type.** The wizard's explicit `log_user_activity` call now writes `action='wizard_create_site'` and `resource_type='site'` (was `wizard_create_proxied_host` / `proxied_host`). Older audit rows already in the database keep their pre-rename strings.
- **Rate-limit dual-name aliasing.** The wizard's per-user-per-minute rate-limit (`COUNT(*)` over `user_activity_logs`) now passes `ANY($::text[])` so it counts BOTH the canonical `site_*` action_name and its legacy `proxied_host_*` companion. A deploy that lands mid-minute cannot reset the quota, and the limit cannot be bypassed by an attacker who picks the legacy name.
- **DB schema rebrand (Phase I).** The `wizard_drafts.wizard_type` column's schema-level `DEFAULT` flipped from `'proxied_host'` to `'site'`. New rows land with the canonical value via an explicit `INSERT … VALUES ($1, 'site', …)`. The list / cap / cluster-delete-purge queries all filter on `wizard_type IN ('site', 'proxied_host')` so pre-rebrand drafts owned by the same user remain visible and remain rejectable. **Existing rows are NOT row-rewritten** — the migration is a metadata-only `ALTER TABLE … SET DEFAULT 'site'` that takes a non-blocking lock and is idempotent.
- **Wizard ACL UX parity.** The Frontend step now embeds the same `ACLRuleBuilder` component used by the Frontend Management page, with cluster-scoped existing backends populated automatically and the wizard's brand-new backend surfaced as a virtual entry in the use_backend dropdown. Drafts persist the three rule arrays so a resumed draft hydrates with the same routing config.
Backward compatibility is preserved at every layer: the DB-level `wizard_drafts.wizard_type='proxied_host'` enum value (still valid for pre-rebrand rows), the `LEGACY_WIZARD_DRAFT_SESSION_KEY` browser sessionStorage key, frontend route aliases (`/proxied-hosts/new`, `/proxied-hosts/drafts`), and the legacy `/api/proxied-hosts/*` URL all keep working.
##### Phase J — UI mount-time race fix ("clusters don't appear after deploy")
**Reported symptom.** After every rolling deploy, operators saw the cluster
selector empty for "a long time" — closing and re-opening the browser did
not help, but waiting ~30s did. The user diagnosed it as a UI problem.
**Root cause.** A React mount-time race between `<AuthProvider>` (parent)
and `<ClusterProvider>` (child). React's useEffect commit phase fires
CHILD effects before PARENT effects, so `ClusterProvider.useEffect` —
which dispatches the very first `axios.get('/api/clusters')` — ran BEFORE
`AuthProvider.useEffect` set `axios.defaults.headers.common['Authorization']`.
The first request went out un-authenticated → backend returned 401 →
ClusterContext's `catch` block silently committed `clusters=[]`. The
operator-visible UI rendered "no clusters" until the 30-second
auto-refresh interval re-fired the request, by which point auth had
hydrated and the call succeeded. Restarting the browser kept hitting the
same race because localStorage carried the token but the useEffect
ordering was identical.
**Fix (3 layers of defence).**
1. **`src/index.js` module-level axios bootstrap.** Runs before
`<App />` is rendered, so no React tree (and therefore no useEffect)
can fire before it. Synchronously seeds
`axios.defaults.headers.common['Authorization']` from localStorage
AND installs an `axios.interceptors.request` that re-reads the token
on every outbound request. The interceptor is the belt-and-suspenders
defence — it cannot be raced by mount ordering and survives any
future code path that mutates `axios.defaults`.
2. **`AuthContext` synchronous useState lazy initialisers.** The
`_hydrateAuthSync` helper runs during the AuthProvider RENDER phase,
which precedes ANY child useEffect. It reads localStorage and seeds
`loading=false`, `isAuthenticated=true`, and the user object —
eliminating the post-mount async hydration that produced the race.
3. **`ClusterContext` auth-gate + exponential-backoff retry.** The
first fetch is gated on `isAuthenticated && !authLoading`, and a
transient 5xx / network failure now triggers up to 4 fast retries
(1s, 2s, 4s, 8s — total ~15s) instead of immediately blanking the
cluster list and depending on the 30s auto-refresh interval. 401/403
intentionally do NOT retry (re-auth is the user's job). The previous
cluster list is preserved on transient hiccups so periodic refreshes
no longer flash an "empty state".
**Verification.** 22 static-source pin tests in
`backend/tests/test_frontend_auth_bootstrap_phase_j.py` cover all three
layers (bootstrap order, AuthContext lazy init, ClusterContext retry +
auth-gate) plus the six audit-loop hotfixes below. 590 backend tests
pass; frontend `npm run build` clean.
**Operator-visible outcome.** After deploy, the cluster selector
populates on the FIRST fetch — no 30-second wait. A transient kube-proxy
convergence window collapses to a few seconds (covered by retries) instead
of being masked by the 30-second interval.
##### Phase J audit hotfixes (audit fix #2 → #6)
Successive audit loops surfaced six follow-on issues that each
reproduced one or more of the original symptoms in narrower windows.
Each fix is pinned in the same Phase J pin-test file:
- **Audit fix #2 — stale closure in `fetchClusters`.** Wrapping
`fetchClusters` in `useCallback(…, [])` froze `selectedCluster` at
its mount-time value (`null`), so the 30-second auto-refresh
reported stale agent-health data forever. Bridged via
`selectedClusterRef`, updated in a passive effect, and read inside
the callback.
- **Audit fix #3 — missing `setLoading(true)` on the auth-gated
first fetch.** During the login flow, the ClusterContext effect
fired with `loading=false` (the initial value), so the cluster
selector briefly rendered "No Cluster Selected" before the spinner
came back. Now the effect explicitly seeds `setLoading(true)` when
the auth gate flips open.
- **Audit fix #4 — `loading=false` between retry waves.** The
`finally` clause unconditionally released the loading flag, so the
spinner blinked off between each backoff attempt and the UI
flashed "No Cluster Selected" for up to ~15 seconds — the very
symptom Phase J was meant to eliminate. The release is now gated
on `retryTimerRef.current === null` so the spinner stays on across
the entire retry budget.
- **Audit fix #5 — exhausted retry budget left counter at 4.**
`retryAttemptRef` was reset only on a successful fetch and on the
auth-gate transition. After 4 transient failures in a row the
counter stayed at 4 for the rest of the session, so any subsequent
invocation (the 30s background refresh, an explicit refetch from a
mutator like `deleteCluster`) skipped the retry pattern entirely
on the first transient failure. The settle-into-empty-state branch
now resets the counter so each fresh invocation gets a full retry
budget.
- **Audit fix #6 — page-content "No Cluster Selected" during the
fetch window.** The cluster selector itself already showed a
spinner via `loading`, but page-level components
(`SSLManagement`, `Configuration`, `DashboardV2`,
`BulkConfigImport`, `BulkVersionHistory`) checked
`!selectedCluster` directly and rendered a permanent warning
affordance. During the 15-second retry budget the page therefore
read as "you forgot to pick a cluster" even though the cluster
list was simply still being fetched. Each page now also consumes
`loading: clustersLoading` from `useCluster()` and shows a neutral
"Loading clusters…" affordance until the fetch settles, only then
flipping to the warning. This is the fix that fully closes the
user-visible loop on the original "clusters don't appear after
deploy" report.
##### Phase K — Site Wizard validation hardening + UX simplification
Operators reported that completing the wizard and clicking **Create &
Apply** repeatedly surfaced opaque 422 errors at the final step:
```
HTTP→HTTPS redirect cannot be combined with custom redirect rules.
body -> frontend -> acl_rules -> 0: Input should be a valid dictionary
body -> frontend -> use_backend_rules -> 0: Input should be a valid dictionary
```
Root causes (each fixed by Phase K):
1. **Contract mismatch on rule fields.** `ACLRuleBuilder.js`
serialised ACL / use_backend / redirect rules as `string[]` while
`FrontendStep` typed them as `List[dict]`. Every wizard POST
carrying a single ACL rule failed Pydantic validation. The
downstream renderer in `services/haproxy_config.py` had always
expected strings, so the schema mismatch was the stale side.
2. **Step 2 mutex was advisory only.** The
`https_redirect ⊕ redirect_rules` validator existed at the model
level but the wizard let the operator advance through Step 2 → 3 →
4 with the conflict in place, only to be punted back at Create.
3. **No HAProxy validation before Create.** `/api/sites/preview`
only checked collisions; the real validator ran inside the
`create_site` transaction *after* entity inserts. Operators
discovered errors at apply time.
4. **HTTPS step overcrowded.** 11 SSL bind-line knobs flat on Step 3
without a defaults summary or any visible grouping.
**What changed**
- **Phase A — Backend contract + safety validators**
(`backend/models/site_wizard.py`).
- `acl_rules`, `use_backend_rules` are now `List[str]`.
`redirect_rules` stays `List[Union[str, dict]]` to preserve the
structured-redirect path used by the renderer's
`_format_redirect_rule`.
- Per-element safety validators reject embedded newlines (HAProxy
directive injection prevention), shell-substitution patterns
(`system`, `exec`, `eval`, `$(`, backtick — same set the
manual frontend API has been blocking since pre-R14), 4 KB
string limit, and empty / whitespace-only strings.
- Two new cross-field model validators close silent-bug gaps:
`FrontendStep.reject_tcp_mode_with_https_redirect` (the renderer
used to emit an HTTP-only directive into a TCP frontend) and
`SSLChoice.reject_inverted_tls_versions` (when both `ssl_min_ver`
and `ssl_max_ver` are set, reject `min > max`).
- **Phase B — Step 2 hard-block + TCP-mode guard**
(`SiteWizard.js`, `ACLRuleBuilder.js`).
- The Step 2 Next handler now hard-blocks the
`https_redirect ⊕ redirect_rules` and `mode='tcp' ⊕
https_redirect` combinations with one-click resolve buttons
("Disable HTTP→HTTPS switch" / "Remove redirect rules").
- Switching the frontend to TCP mode auto-clears `https_redirect`;
the Switch is also `disabled` while `mode==='tcp'` with an
explanatory tooltip.
- `ACLRuleBuilder` accepts a new `disableRedirectRules` prop that
visually disables the Redirect Rules section (`aria-disabled`,
greyed-out cards, tooltip) when the parent passes
`https_redirect=true`. The rules data stays in component state
so toggling the switch off restores them.
- **Phase C — Real HAProxy dry-run gate before Create**
(`backend/routers/site_wizard.py`, `SiteWizard.js`).
- New shared helper `_synthesize_candidate_haproxy_config(body,
conn, *, entities_already_inserted)` is used by both
`create_site` (post-insert validation gate) and a new dry-run
path on `POST /api/sites/preview`. Two callsites pinned by
`test_phase_k_create_site_and_preview_use_same_synthesis_helper`
so the apply gate and the dry-run gate cannot silently desync.
- `POST /api/sites/preview` now accepts an optional
`validate_haproxy_config=true` query param. When set, the
endpoint runs `HAProxyConfigValidator` against the synthesised
candidate config and returns a `validation: {is_valid,
error_count, warning_count, errors, warnings, infos}` block in
the same 200 OK envelope. Validator crashes return
`is_valid: null` + `validator_error` (matches `create_site`'s
non-fatal posture). The dry-run path is rate-limited at 5/min
via `_enforce_rate_limit` and emits structured ENTER/EXIT
`logger.info` lines for telemetry. Legacy preview callers
(`SiteDrafts.handlePreview`) are unaffected — they pass no flag.
- The wizard auto-fires the dry-run on Step 4 entry with an
`AbortController` so rapid Step 4 → Step 2 → Step 4 navigation
cancels the in-flight request. A six-state validation card
renders inline: idle / loading / clean (green) /
`warnings_only` (yellow) / `errors` (red, blocks Create) /
`pydantic_error` (red, body-parse failures from Phase A's new
validators or PEM-stripped resume drafts) / `unavailable`
(orange, advisory — Create stays enabled to mirror the
validator-crash-is-non-fatal contract). Each error /
`pydantic_error` row gets an `Edit Step N` jumpback button via
a static directive→step + loc→step mapping table.
- Audit-fix #1 (post-implementation review): the wizard also
resets `dryRunResult.status` to `idle` whenever the operator
leaves Step 4. The ACL builder lives outside the antd Form
so its mutations don't fire `Form.onValuesChange`; without
this reset a stale `clean`/`errors`/`warnings_only` status
survives Step 4 → Step 2 (ACL edit) → Step 4 round-trips
and the auto-fire branch suppresses the next fetch. With
the reset every Step 4 entry triggers a fresh dry-run
(rate-limit-safe — entry is operator-initiated, not
programmatic).
- Audit-fix #2 (post-implementation review): the
`Edit Step N` jumpback now also resolves the target step
from the error **message text** when `loc` cannot pinpoint
it. Pydantic v2 raises `model_validator(mode="after")`
errors with `loc=()`; FastAPI prepends `'body'` so the
operator-visible envelope is `loc=['body']` (length 1).
The legacy `_locPathToStep` early-returned null for this
case, dropping the jumpback for PEM-stripped resume
("ssl.mode='upload' requires a non-empty PEM-encoded
certificate_content …") and every
`enforce_acme_apply_and_http` cross-field rejection. A
small ordered pattern table recovers the step from the
failure message text so operators always get a working
"fix-from-here" button.
- Audit-fix #2 round 3 (post-implementation review): the
pattern table is ordered so cross-field ACME messages
route to the step the operator must EDIT to fix the
error, not the step that "feels related". A naive ordering
("SSL first because every cross-field message starts with
`ssl.mode='acme'`") would route every cross-field hit to
Step 3, defeating the jumpback. Order is now:
`apply_immediately` (Step 4) → `wildcard`/`domains` (Step
0) → `frontend.*`/`bind_port` (Step 2) → `backend.*` (Step
1) → SSL catch-all (Step 3, LAST). With this ordering,
"ssl.mode='acme' requires apply_immediately=true" routes
to Step 4 (toggle the switch), "ssl.mode='acme' requires
frontend.bind_port=80" routes to Step 2 (edit FE port),
and "(HTTP-01) cannot issue wildcard certs" routes to
Step 0 (remove wildcard). PEM-stripped and other SSL-only
errors still hit Step 3 via the final catch-all.
- Audit-fix #2 round 4 (post-implementation review): the
pydantic_error renderer no longer emits a stray
`<strong>: </strong>` orphan-colon prefix when the failing
error has no field path. SiteCreate-level model_validator
errors land with `loc=['body']` (length 1); after dropping
the leading `'body'` marker the joined path is empty.
Pre-fix the renderer wrapped that empty string in
`<strong>...: </strong>`, producing a visually broken " : "
prefix in front of every PEM-stripped resume message and
every `enforce_acme_apply_and_http` cross-field rejection.
Post-fix the strong/colon prefix renders only when a real
field path exists.
- **Phase D — HTTPS step simplification + UI parity**
(`SiteWizard.js`).
- TLS bounds (`ssl_min_ver`, `ssl_max_ver`) and the HSTS quartet
stay first-class on the SSL step; rarely-used knobs
(`https_bind_port`, `https_frontend_name_suffix`, `ssl_alpn`,
`ssl_ciphers`, `ssl_ciphersuites`, `ssl_strict_sni`,
`ssl_verify`) move into a nested **Advanced TLS settings
(rarely needed)** Collapse that defaults to closed. A read-only
summary line ("Port 443, ALPN h2,http/1.1, …") shows the safe
defaults that apply unless overridden.
- The Advanced Collapse auto-opens (`defaultActiveKey`) when a
saved draft has any non-default value, so resumed drafts
surface their custom tuning instead of silently hiding it.
- HSTS UI parity for the Phase A
`reject_hsts_preload_without_hsts` validator: the
`hsts_preload` Switch is `disabled` until HSTS is enabled,
`max-age ≥ 31536000`, AND `includeSubDomains=true`.
`hsts_max_age` and `hsts_include_subdomains` are also disabled
while `hsts_enabled=false`.
- TLS min/max ordering UI parity: the `ssl_min_ver` /
`ssl_max_ver` Selects use Antd `dependencies` + a custom
validator that rejects min > max client-side with the same
wording the Phase A model validator uses.
**Backward compatibility**
- The existing `/api/sites` POST envelope is unchanged.
- The existing `/api/sites/preview` POST envelope gains an optional
`validation` field that legacy callers can ignore. The
`validate_haproxy_config` flag defaults to `false`, so
`SiteDrafts.handlePreview` and any external integrators keep
their pre-Phase K behaviour.
- `redirect_rules` retains its `List[Union[str, dict]]` shape, so
any historical caller (or saved draft) that used the structured
dict form continues to work.
- The Pydantic safety validators (`system`, `exec`, `eval`, `$(`,
backtick) match the manual frontend API's existing
`validate_acl_rules` posture, which has been in production
blocking the same substrings since pre-R14 with no operator
complaint. No existing wizard payload that previously round-
tripped through `services/haproxy_config.py` can be rejected by
these new validators.
- The `_synthesize_candidate_haproxy_config` helper in
`entities_already_inserted=True` mode is functionally identical
to the previous inline `generate_haproxy_config_for_cluster`
call inside `create_site`. The refactor is pure DRY plumbing.
##### Rollback considerations (Phase K)
If you must roll back to a pre-Phase-K v1.5.x build:
- **Saved drafts** with the new `acl_rules: List[str]` shape are
forward- and backward-compatible: the legacy build also expected
string elements at the renderer level, the rejection only ever
happened at the wizard model boundary. Operators on the legacy
build hit the same 422 the new build is fixing — no DB rewrite
needed.
- **`/api/sites/preview` `validation` block** is a new optional
field; legacy frontend callers ignore unknown fields. The
`validate_haproxy_config` query param default is `false`, so
legacy callers do not exercise the dry-run branch.
- **No DB migrations** are introduced by Phase K. The
`frontends.acl_rules` / `redirect_rules` / `use_backend_rules`
JSONB columns remain unchanged.
##### Phase K Phase D — Operator-feedback follow-ups (Bulgu #1–#6)
Operator review of the Phase A–C release surfaced six additional
issues. Each is rooted in a UX inconsistency or a residual stuck
state, and the fixes converge on a "single source of truth + ref-
based dry-run lifecycle" architecture:
- **Bulgu #1 — Cluster scope.** Pre-fix Step 0 had its own cluster
Select dropdown decoupled from the header. Operators routinely
picked cluster A in the header and cluster B in the wizard with
zero visual signal that the wizard would target a different
cluster than every other tool. Phase D pipes the wizard through
the SAME `ClusterContext` that FrontendManagement / BackendServers
/ SSLManagement consume, hides Step 0's `cluster_id` `Form.Item`,
and replaces the picker with a read-only `<Tag>` display + hint
to change cluster via the header. A `useEffect` keeps
`form.cluster_id` synchronised with `selectedCluster.id` so mid-
wizard header changes propagate; the existing cluster-transition
cleanup effect handles cert-id orphan reconciliation. Resume from
a draft that targets a different cluster now auto-swaps the
header cluster (best-effort `selectCluster()` call) so post-
resume edits stay cluster-consistent.
- **Bulgu #2 — SSL CA bundle dropdown filter.** Backend's
`BackendServers.js` filters the CA-bundle Select with
`?usage_type=server`, so operators only see certs imported with
the right purpose. The wizard pre-fix surfaced EVERY cert in the
cluster regardless of usage, letting an operator submit a payload
that apply-time HAProxy would parse-error on (`unable to load
SSL private key`). Phase D filters explicitly:
* Per-server CA bundle Select → `usage_type === 'server'`.
* SSL & ACME step's "Existing certificate" Select →
`usage_type === 'frontend'`.
The empty-state Alert was also updated to reason about only the
filtered list so a cluster with N server-side certs but zero
frontend certs renders the "no certs imported" hint correctly.
- **Bulgu #3 — Stuck "Validating against HAProxy…".** The root
cause was a self-cancel race in the auto-fire `useEffect`. The
effect deps array included `dryRunResult.status`, and the effect
body called `setDryRunResult({status: 'loading'})` at the top.
The status change re-triggered the effect; React's cleanup of the
previous run fired BEFORE the new body, aborting the in-flight
controller; the new body returned early because `status !==
'idle'`; the aborted fetch's `.catch` block detected
`signal.aborted` and returned without setting state. Status
stayed `'loading'` forever. Audit-fix #1 (round 1) had addressed
the leave-Step-4 cleanup branch but the enter-Step-4 self-abort
was a separate failure mode that only surfaced on a real backend.
Phase D switches the lifecycle to a ref-driven model:
* `dryRunStatusRef` shadows the latest status (synced via a
passive `useEffect`).
* `dryRunInvalidationTick` is the external re-trigger channel;
`onValuesChange` bumps it when the operator edits a Step-4-
visible field (e.g. the Apply Immediately switch).
* The main effect's deps array drops `dryRunResult.status` and
becomes `[step, form, aclBuilderData, dryRunInvalidationTick]`
— none of these change on a self-issued setDryRunResult, so
the self-cancel race is structurally impossible.
* Cleanup nulls the abort ref only if it still points to the
torn-down controller, so a fresh fetch's ref is never
accidentally cleared.
- **Bulgu #4 — Preview missing fields.** The /api/sites/preview
response previously echoed only a sparse subset of fields, so the
SiteDrafts Preview modal could not show whether per-server
timings, backend cookie persistence, frontend maxconn, HSTS, or
ciphersuites would actually land on disk. Phase D enriches both
the backend response (additive — all existing keys preserved)
AND the SiteDrafts UI:
* Backend: emits the full operator-settable surface area on
`would_create` (backend cookie/timeouts/options, per-server
timings + SSL+CA-bundle details, frontend maxconn/timeouts/
compression/ACL counts, HTTPS ciphersuites, etc.).
* Frontend: replaces the four flat Descriptions blocks with a
typed renderer that only surfaces NON-DEFAULT values
(`isMeaningful` predicate) so the modal stays scannable. A
dedicated per-server card surfaces every per-server field
the operator customised. HSTS gets its own section when
enabled.
- **Bulgu #5 — Resume hydration regressions.** Two issues:
1. Existing certificate was wiped on resume. Root cause was
the orphan-detect effect running on the SAME render that
the resume effect committed the new cluster_id. existingCerts
was still `[]` (fetch in flight), so `certIds = new Set()`
and the freshly-resumed `ssl.ssl_certificate_id` looked like
an orphan and got cleared. Phase D fix: short-circuit the
orphan-detect when `existingCertsLoading=true` and add the
loading flag to the effect deps so the check re-runs after
the fetch settles. ALSO: pin `prevClusterRef.current` to
`merged.cluster_id` BEFORE `form.setFieldsValue(merged)` so
the cluster-transition cleanup effect does not misread the
hydration as a user-driven cluster switch.
2. The same stuck "Validating against HAProxy…" — resolved by
the Bulgu #3 self-cancel-race fix above.
- **Bulgu #6 — Create as PENDING button removed.** Pre-fix the
wizard had TWO submit buttons. The "Create as PENDING" button
bypassed the standard manual-flow convention (entity Create →
PENDING version → Apply Management review → operator Apply). The
"Create & Apply" button bypassed Apply Management entirely.
Operators were trained to "always Create & Apply", defeating the
change-review benefit of Apply Management. Phase D consolidates:
* Single button: "Create Site" (or "Create & Apply (ACME)" when
sslMode='acme', because ACME forces the immediate apply for
the HTTP-01 challenge).
* `handleSubmit` derives `effectiveApply` from `sslModeAtSubmit
=== 'acme'` — no button-driven branching.
* Non-ACME flow: `apply_immediately=false` → backend returns
`created_pending` → operator is navigated to /apply-management
where they review the bulk version and click Apply (same
Agent-pull cadence as manual entity creation).
* ACME flow: `apply_immediately=true` (M22 model_validator
enforces this) → standard `created_applied` response.
* The `acmeBlocksDraft` derivation that gated the (now-removed)
PENDING button is retired — handleSubmit's `effectiveApply`
replaces the gate.
##### Phase K Phase D — Backward compatibility / rollback
- **Cluster picker change.** Operators who relied on the wizard-
internal cluster Select must switch via the header instead. No
data-layer change. Drafts saved on a different cluster
auto-swap the header on resume.
- **`/api/sites/preview` response shape.** Additive only — every
pre-existing key keeps the same shape; new keys are
`cluster_id`, `domains`, additive fields on `backend` / `servers`
/ `frontend_http` / `frontend_https`. Legacy frontend callers
ignore unknown fields.
- **`/api/sites` request shape.** Unchanged.
- **No DB migrations** are introduced by Phase K Phase D.
##### Phase K Phase D — Follow-up audit findings (Bulgu #7–#8)
A deeper post-implementation audit surfaced two additional
race conditions that were not visible in the first pass. Both
are now resolved on the same `pilot` branch:
- **Bulgu #7 — Resume cluster swap race on cold mount.** On a
browser refresh of `/sites/new` while a Resume click had
already pre-populated sessionStorage, the wizard mount races
against `ClusterContext`'s `fetchClusters()`. The resume
effect ran with `clustersFromContext=[]`, so
`selectCluster(draftCluster)` was silently skipped. Then
`ClusterContext` finished loading and `selectedCluster`
became the user's `defaultCluster` (NOT the draft's
cluster). The naive header sync then overwrote
`form.cluster_id` with the default cluster, and the
cluster-transition cleanup effect read that overwrite as a
user-driven switch and wiped the draft's cert selections —
the Bulgu #5 second-order failure that survived the
short-circuit fix on a cold mount path.
Fix: header sync effect grew a one-shot post-resume swap
branch keyed on `resumedFromDraft && !resumeClusterSynced`.
When the draft's `cluster_id` is in the freshly-loaded
`clustersFromContext`, the swap pushes the HEADER to the
draft cluster instead of forcing the form to follow the
header. The `resumeClusterSynced` state gates this to
exactly ONE attempt so a later operator-driven header
cluster change is honoured normally. `selectClusterRef`
(a `useRef(selectCluster)` updated by a tiny sync effect)
keeps the dep set small so the header sync effect does not
re-run on every `ClusterProvider` render.
Pin: `tests/test_frontend_auth_bootstrap_phase_j.py::
test_phase_k_phase_d_resume_cluster_swap_race_fix`.
- **Bulgu #8 — Mid-wizard cluster change leaves stale dry-run.**
When an operator on Step 4 changes the header cluster, the
wizard's cluster_id transitions through `form.setFieldsValue`
(the header sync effect's standard force path). Antd's
`setFieldsValue` is a SILENT update that does NOT fire
`onValuesChange`, so the dry-run invalidation tied to
`onValuesChange` never ran. Result: the Step 4 validation
card kept displaying the PREVIOUS cluster's "clean" verdict
even though the wizard payload now targeted a different
cluster.
Fix: the cluster-transition cleanup effect (which already
detected the change to wipe stale cert ids) now also resets
`dryRunResult` to idle and bumps `dryRunInvalidationTick`
whenever `dryRunStatusRef.current !== 'idle'`. The dry-run
effect's dep list picks up the tick bump and re-fires
against the new cluster as soon as the operator reaches
Step 4.
Pin: `tests/test_frontend_auth_bootstrap_phase_j.py::
test_phase_k_phase_d_cluster_change_invalidates_dry_run`.
Both fixes are additive (no API or DB changes) and rollback
without leaving residual state — disabling the new effects
simply restores the previous (racy) behaviour.
##### Phase K Phase D — Operator-feedback round 2 (Bulgu #9–#11)
A second operator-feedback round surfaced one parity gap and two
follow-ups on the wizard's HAProxy validation experience:
- **Bulgu #9 — Wizard PEM upload parity with SSL Management page.**
Pre-fix `services.ssl_service.create_cert_row` (the helper the
wizard calls when `ssl.mode='upload'`) was a thin INSERT that
never parsed the PEM. It stored `primary_domain` / `all_domains`
from the operator-entered FRONTEND domains (not the cert SAN),
left `expiry_date` / `issuer` / `fingerprint` NULL, hard-coded
`status='valid'` and `days_until_expiry=0`, never validated the
private key or chain, never checked name uniqueness (so a
duplicate name would 500 at the DB unique constraint), and could
not reactivate a soft-deleted row of the same name. The
resulting cert showed up on the SSL Management page with empty
expiry/issuer columns and a permanent "valid" status — confusing
UX and clearly inconsistent with the dedicated SSL Management
upload flow (`POST /api/ssl/certificates`).
Fix: `create_cert_row` now mirrors `routers/ssl.py::
create_ssl_certificate`:
- parses the PEM via `utils.ssl_parser.parse_ssl_certificate`
(raises HTTPException 400 on parse failure),
- validates private_key + chain via `validate_private_key`
/ `validate_certificate_chain`,
- computes status / days_until_expiry from the normalised
timezone-naive UTC `expiry_date`,
- enforces name uniqueness within the target cluster (returns
400 instead of a DB-level 500),
- reactivates soft-deleted rows of the same name (preserves
the row id for downstream references).
Pin: `tests/test_ssl_service_extraction.py` — 11 tests cover
the happy path, all 6 negative paths (parse fail, empty content,
bad private key, bad chain, duplicate active name, soft-delete
reactivation), and the "metadata comes from PEM, not payload"
contract.
- **Bulgu #10 — Heuristic validator rejected wizard's own default
timeouts.** The wizard's config synthesis emits `timeout connect
10000ms` / `timeout server 60000ms` / `timeout client 100ms`
(millisecond suffix is canonical HAProxy syntax). The pre-fix
heuristic regex was `^\d+[smhd]?$`, which only allowed the
single-character suffixes `s`/`m`/`h`/`d` — `ms` was rejected
outright even though the same validator's own suggestion text
said "Use format like '5s', '30000ms', '1m'". Operators saw
10+ FALSE-POSITIVE "Invalid timeout value '10000ms'" errors on
the wizard's defaults at Step 4 and could not click Create.
Fix: `utils/haproxy_validator.py::_validate_timeout_directive`
regex relaxed to `^\d+(us|ms|s|m|h|d)?$` — accepting the full
set of HAProxy time-format suffixes (per the HAProxy docs Time
format chapter) while still rejecting malformed values like
`10000xx`, `abc`, `-100ms`, `1.5s`, and bare `ms`.
Pin: `tests/test_haproxy_validator_timeout_units.py` — 17
parametrised cases (11 valid formats, 5 invalid formats, plus
the exact operator-reported failure mode).
- **Bulgu #11 — Operator reported "Previous loses values".**
Architectural review confirmed the wizard's contract is sound:
every step is rendered into a long-lived `<div>` whose only
step-driven prop is the CSS `display` toggle (`block` vs
`none`). React does NOT unmount the children, Antd's Form.Item
registrations stay intact, and the Antd default `preserve=true`
keeps values in form state even for the inner Form.Items that
conditional-render inside `<Form.Item shouldUpdate>` (SSL mode
branches, TCP/http frontend mode toggle). All wizard
`setFieldsValue` call-sites are guarded by domain triggers
(cluster change, sslMode change, TCP-mode-clears-https_redirect,
resume hydration) — none fire on a Previous/Next click alone.
No code regression was identified. Most likely operator
perception driver: with Bulgu #10 fixed, the `timeout
connect=10000` / `timeout server=60000` values the operator
saw in the "Advanced backend settings" Collapse after coming
back from Step 4 are simply the wizard's pre-existing defaults
(`backend.timeout_connect=10000`, `backend.timeout_server=
60000`, `backend.timeout_queue=60000`), not regressed values
— these were never operator-entered, just defaults the
operator did not notice in the collapsed Advanced section on
the forward pass.
Defensive measure: a static-source pin test asserts the
architectural contract so a future refactor cannot regress
to per-step conditional rendering or sneak a
`preserve={false}` in:
`tests/test_frontend_auth_bootstrap_phase_j.py::
test_phase_k_phase_d_wizard_preserves_form_state_across_step_navigation`.
If the operator can reproduce specific field-level state loss
on a Previous click after the Bulgu #10 fix, please file the
repro steps so we can target the actual scenario.
##### Rollback considerations (Phase I)
If you must roll back to a pre-rebrand v1.5.x build after operators have already saved drafts on the new build:
- New rows on `wizard_drafts` with `wizard_type='site'` will be invisible to the legacy code path that filters on `wizard_type='proxied_host'` only. Operators will see those new drafts disappear from the listing AND will not be counted against the 50-draft cap. The rows themselves are not deleted — they expire via the standard 30-day TTL prune.
- Pre-rebrand rows with `wizard_type='proxied_host'` continue to work on the legacy build because their value never changed.
- The schema-level `DEFAULT` is not rolled back automatically. Operators rolling back can either (a) leave it at `'site'` (harmless — the legacy build hard-codes `'proxied_host'` in every INSERT, so the default is never consulted) or (b) re-run an `ALTER TABLE wizard_drafts ALTER COLUMN wizard_type SET DEFAULT 'proxied_host'` to restore the original schema.
##### Phase K Phase D — Operator-feedback round 3 (Bulgu #12)
**Operator-reported failure flow** (May 11, 2026):
The wizard's Step 4 dry-run showed 8 WARNINGs but no ERRORs, so Create proceeded; the operator then applied via Apply Management and the real `haproxy -c` parse rejected the config:
```
[ALERT] parsing [/tmp/haproxy-new-config.cfg:79] : error detected while parsing ACL 'acl1' : failed to open pattern file </path>.
[ALERT] parsing [/tmp/haproxy-new-config.cfg:87] : error detected while parsing switching rule : no such ACL : 'acl1'.
[ALERT] Fatal errors found in configuration.
```
The 8 WARNINGs were ALSO operator-confusing false positives:
```
[frontend] Directive 'stick-table' may not be valid in 'frontend' section
[frontend] Directive 'tcp-request' may not be valid in 'frontend' section (×2)
[backend] Directive 'cookie' may not be valid in 'backend' section (×2)
[backend] Missing 'global' section - recommended for production
```
**Two root causes:**
1. **Heuristic validator `valid_directives` was incomplete** — `stick-table`, `tcp-request`, `tcp-response`, `cookie`, `http-after-response`, `errorfile`, `description`, `id`, `filter`, etc. are perfectly valid in their respective sections but the validator's small hand-picked sets did not list them. Every wizard / manual page that emitted them flagged a spurious "may not be valid" WARNING. The wizard's pre-persist apply-time gate uses the same validator; even though it only blocks on ERROR-level findings, the noise polluted the operator-visible response trail and the version-history page.
2. **ACL `-f <file>` pattern-file references** — the visual ACL builder offered `-f (from file)` as a selectable flag, and neither the manual Frontend API's Pydantic validator (`models/frontend.py::validate_acl_rules`) nor the wizard's Pydantic validator (`models/site_wizard.py::_validate_haproxy_directive_string`) rejected `-f`. HAProxy OpenManager is a fully-managed product: it does NOT provision pattern files onto the HAProxy node's filesystem, so any operator-typed `-f /path/...` ALWAYS resolves to "file not found" at HAProxy reload time. The UI made it trivial to author an unsupported state.
**Three-layer fix:**
**Layer A — Heuristic validator** (`backend/utils/haproxy_validator.py`):
- Expanded `valid_directives['frontend']` to include `stick-table`, `stick`, `tcp-request`, `tcp-response`, `http-after-response`, `errorfile`, `errorloc`, `errorloc302`, `errorloc303`, `http-error`, `description`, `id`, `filter`, `monitor`, `unique-id-format`, `unique-id-header`, `declare`, `http-buffer-request`, plus a long-tail of less-common-but-valid directives.
- Expanded `valid_directives['backend']` to include `cookie`, `appsession`, `tcp-request`, `tcp-response`, `tcp-check`, `retries`, `fullconn`, `dispatch`, `redirect`, `use-server`, `acl`, `capture`, `errorfile`, `description`, `id`, `filter`, `rate-limit`, `email-alert`, `force-persist`, `transparent`, `source`, plus a long-tail.
- Added `partial_fragment: bool = False` parameter to `HAProxyConfigValidator.validate_config()` and the module-level `validate_haproxy_config()`. When True (or auto-detected via the wizard's marker comment), the validator suppresses the "Missing 'global' section" / "Consider adding 'defaults' section" diagnostics — the wizard / cluster synthesis intentionally OMITS those blocks because the agent merges them with its local copy on disk.
- Both the wizard's `/preview` dry-run AND the apply-time pre-persist gate now pass `partial_fragment=True` (`backend/routers/site_wizard.py`).
**Layer B — ACL `-f` rejection in Pydantic** (server-side gate):
- `backend/models/site_wizard.py`: Added `_ACL_FILE_FLAG_PATTERN = re.compile(r"(^|\s)-f(\s|$)")` and rejected the pattern inside `_validate_haproxy_directive_string` with an operator-friendly message explaining why the product cannot support pattern files. This covers `acl_rules`, `use_backend_rules`, and string-shaped `redirect_rules`.
- `backend/models/frontend.py::validate_acl_rules`: Mirrored the same rejection on the manual Frontend API so both create paths return the identical 400/422 envelope.
**Layer C — ACL `-f` removal from the visual builder + UI gates** (client-side authoring guardrail):
- `frontend/src/components/ACLRuleBuilder.js`: Removed `-f` from the selectable `FLAGS` list. Updated `FLAG_HINTS` to drop the `-f` mention. Existing rules that already carry `-f` (loaded from saved drafts pre-fix) keep the tag visible as `-f (deprecated — remove)` so operators can SEE and REMOVE the flag, but cannot re-add it once removed. Added a section-level red `Alert` that counts every rule carrying `-f` and explains the failure mode + remediation. Inline rule-card error decoration (`status='error'` + red border + inline description) surfaces the same message at the per-rule level. Mirrored the regex client-side so raw-mode typed `-f` immediately flags inline.
- `frontend/src/components/SiteWizard.js`: Added a Step 2 → Step 3 hard-gate on the Next button — if ANY rule still carries `-f`, the click surfaces the same operator-friendly error and refuses to advance.
- `frontend/src/components/FrontendManagement.js::handleSubmit`: Mirrored the same gate so the manual Frontend page rejects submit identically.
**Backward compatibility:**
- Existing drafts that contain `-f`-flagged rules still load — the ACLRuleBuilder displays them visibly so operators can remove them. Submit is blocked until they do.
- Existing PERSISTED frontend rows in the DB that already carry `-f` (created before this fix) continue to work at the agent level — the validator changes do NOT retroactively reject them. They can still be EDITED through the UI (which will block save until `-f` is removed) or read via the API for visibility / audit.
- The expanded `valid_directives` sets only ADD entries; nothing previously accepted is now flagged. Pre-existing tests that asserted "Directive X is valid" continue to pass.
**Tests added:**
- `backend/tests/test_haproxy_validator_bulgu12.py` (27 new tests):
- Per-directive false-positive regression pins for both frontend and backend sections.
- `partial_fragment=True` suppression + marker-comment auto-detect.
- Wizard Pydantic `-f` rejection across spacing/position variants.
- Anchor-correctness pin: regex must NOT match `-foo` / `-file` substrings inside other tokens.
- Manual Frontend API parity pin.
- End-to-end pin replaying the user's actual config (minus `-f`) with zero spurious WARNINGs.
- `backend/tests/test_site_wizard_phase2_validator_gate.py`: Widened the pre-window lookback from 400 to 1500 chars to accommodate the partial-fragment forwarding comment block.
**Rollback considerations:**
- Reverting the `valid_directives` expansion brings back operator-visible WARNING noise but does NOT break apply (which only gates on ERROR). Safe to roll back if a regression is discovered.
- Reverting the `-f` Pydantic rejection ALLOWS operators to author the failure mode again, but does not break anything that worked before. Roll back ONLY if a customer has pre-provisioned pattern files and a tightly-controlled need to reference them.
- Reverting the ACLRuleBuilder UI changes is a pure visual revert; the Pydantic gate keeps the safety net.
### Earlier Releases
For earlier release notes (v1.4.0 ACME stability + enterprise audit, v1.3.0, ...) see the [GitHub Releases](https://github.com/taylanbakircioglu/haproxy-openmanager/releases) page.
For full release notes and the list of features delivered in each version (v1.5.x Site Wizard + ACME Diagnostic Panel, v1.4.0 ACME stability + enterprise audit, v1.3.0, ...) see the [GitHub Releases](https://github.com/taylanbakircioglu/haproxy-openmanager/releases) page.
---
+79
View File
@@ -1700,8 +1700,87 @@ async def run_all_migrations():
# (cluster_id, bind_address, bind_port) WHERE is_active.
await ensure_frontends_bind_unique_constraint()
# Issue #18 — TOTP MFA (v1.6.0): additive columns + 3 new tables
await ensure_mfa_columns()
logger.info("Database migrations completed successfully.")
async def ensure_mfa_columns():
"""Issue #18 — TOTP MFA (v1.6.0): additive columns on users + 3 new tables.
All operations are idempotent (ADD COLUMN IF NOT EXISTS, CREATE TABLE IF NOT EXISTS).
Default behavior preserved: every existing user gets mfa_enabled=FALSE, so login
flow is byte-identical for accounts that don't opt in.
"""
conn = None
try:
conn = await get_database_connection()
await conn.execute("""
ALTER TABLE users
ADD COLUMN IF NOT EXISTS mfa_enabled BOOLEAN DEFAULT FALSE NOT NULL,
ADD COLUMN IF NOT EXISTS mfa_method VARCHAR(20),
ADD COLUMN IF NOT EXISTS mfa_secret_encrypted TEXT,
ADD COLUMN IF NOT EXISTS mfa_enrolled_at TIMESTAMP,
ADD COLUMN IF NOT EXISTS mfa_last_used_at TIMESTAMP,
ADD COLUMN IF NOT EXISTS mfa_last_used_totp_step BIGINT;
""")
await conn.execute("""
CREATE TABLE IF NOT EXISTS mfa_backup_codes (
id SERIAL PRIMARY KEY,
user_id INTEGER NOT NULL REFERENCES users(id) ON DELETE CASCADE,
code_hash VARCHAR(255) NOT NULL,
used_at TIMESTAMP,
created_at TIMESTAMP DEFAULT CURRENT_TIMESTAMP
);
""")
await conn.execute("""
CREATE INDEX IF NOT EXISTS idx_mfa_backup_codes_user
ON mfa_backup_codes(user_id);
""")
await conn.execute("""
CREATE TABLE IF NOT EXISTS mfa_pending_logins (
id SERIAL PRIMARY KEY,
user_id INTEGER NOT NULL REFERENCES users(id) ON DELETE CASCADE,
challenge_token VARCHAR(64) UNIQUE NOT NULL,
attempts INTEGER DEFAULT 0 NOT NULL,
expires_at TIMESTAMP NOT NULL,
used_at TIMESTAMP,
created_at TIMESTAMP DEFAULT CURRENT_TIMESTAMP,
ip_address INET
);
""")
await conn.execute(
"CREATE INDEX IF NOT EXISTS idx_mfa_pending_token ON mfa_pending_logins(challenge_token);"
)
await conn.execute(
"CREATE INDEX IF NOT EXISTS idx_mfa_pending_expires ON mfa_pending_logins(expires_at);"
)
await conn.execute("""
CREATE TABLE IF NOT EXISTS mfa_pending_enrollments (
user_id INTEGER PRIMARY KEY REFERENCES users(id) ON DELETE CASCADE,
secret_encrypted TEXT NOT NULL,
attempts INTEGER DEFAULT 0 NOT NULL,
expires_at TIMESTAMP NOT NULL,
created_at TIMESTAMP DEFAULT CURRENT_TIMESTAMP
);
""")
await conn.execute(
"CREATE INDEX IF NOT EXISTS idx_mfa_pending_enroll_expires ON mfa_pending_enrollments(expires_at);"
)
logger.info("✅ MFA migration completed (Issue #18 — Phase 1)")
except Exception as e:
logger.error(f"Failed to ensure MFA columns: {e}")
# Don't raise — follow the same defensive pattern as ensure_user_activity_logs_table
finally:
if conn:
await close_database_connection(conn)
async def add_ssl_certificate_id_to_backend_servers():
"""Add ssl_certificate_id column to backend_servers table for SSL certificate management"""
conn = None
+3 -1
View File
@@ -8,7 +8,7 @@ import redis
import asyncio
from datetime import datetime, timedelta
_version_info = {"version": "1.5.0", "releaseName": "ACME Diagnostics & Site Wizard", "releaseDate": "2026-05-08"}
_version_info = {"version": "1.6.0", "releaseName": "Multi-Factor Authentication (MFA)", "releaseDate": "2026-05-18"}
for _vpath in ["/app/version.json", os.path.join(os.path.dirname(__file__), "..", "version.json")]:
try:
with open(_vpath) as _vf:
@@ -40,6 +40,7 @@ from routers.settings import router as settings_router
from routers.letsencrypt import router as letsencrypt_router
from routers.acme_diagnostics import router as acme_diagnostics_router
from routers.site_wizard import router as site_wizard_router
from routers.mfa import router as mfa_router
# Production logging configuration
from utils.logging_config import setup_production_logging
@@ -820,6 +821,7 @@ app.include_router(config_router) # Configuration management
app.include_router(maintenance_router, prefix="/api", tags=["maintenance"]) # Database cleanup & maintenance
app.include_router(auth_router)
app.include_router(user_router)
app.include_router(mfa_router)
app.include_router(frontend_router)
app.include_router(backend_router)
app.include_router(cluster_router)
+19 -3
View File
@@ -126,10 +126,26 @@ class GlobalExceptionHandler:
except Exception as body_error:
logger.debug(f"Could not extract raw body for debugging: {body_error}")
# Enhanced log message for agent heartbeats
log_message = f"Validation error: {str(exc)}"
# Build a sanitized log summary. The raw `str(exc)` from Pydantic
# contains the user-supplied `input` value for each failed field —
# which leaks secrets like TOTP codes, backup codes, mfa_token, and
# passwords to plaintext logs. Use only field NAMES + types here;
# `validation_details` (already sanitized to {field, message, type})
# is attached separately for downstream structured logging.
_field_names = [
err.get("field", "unknown")
for err in error_details.get("validation_errors", [])
]
_err_count = len(error_details.get("validation_errors", []))
log_message = (
f"Validation error: {_err_count} field(s) failed validation: "
f"[{', '.join(_field_names)}]"
)
if agent_name != "unknown":
log_message = f"Agent '{agent_name}' heartbeat validation error: {str(exc)}"
log_message = (
f"Agent '{agent_name}' heartbeat validation error: "
f"{_err_count} field(s) failed: [{', '.join(_field_names)}]"
)
# Log validation error with enhanced details
log_with_correlation(
+147
View File
@@ -0,0 +1,147 @@
"""MFA rate-limit key extraction (user-aware + ingress-aware).
Problem with the default ``slowapi.util.get_remote_address``:
* Behind an ingress / reverse proxy / load balancer, every request looks
like it originates from the same upstream IP (the proxy itself). A
5/minute IP-bucket therefore becomes a 5/minute *whole-organization*
bucket. In a 500-user enterprise rolling out MFA, this stalls the
rollout to a trickle.
Strategy (combined B + C from the design review):
1. **User-aware key (preferred):** if the request carries a valid Bearer
JWT, derive the bucket from ``user:<id>``. Each authenticated user
gets an isolated bucket, regardless of source IP. A malicious user
burning their own quota cannot starve the other 499.
2. **Trusted-proxy XFF fallback:** if the request is unauthenticated
(e.g. ``/login`` flow, future endpoints) and the TCP peer is in
``MFA_TRUSTED_PROXY_CIDRS``, peel the first hop off ``X-Forwarded-For``.
This preserves real-IP buckets behind a known ingress without
accepting spoofed headers from the public internet.
3. **Default fallback:** plain ``request.client.host`` (slowapi default).
JWT decode is intentionally signature-verified (replay/spoof protection)
and *sync* — slowapi's decorator hook is sync, and our JWT library
(python-jose) is sync as well. Failed verification silently downgrades
to IP-based bucketing — never crashes the decorator.
Env vars:
* ``MFA_TRUSTED_PROXY_CIDRS`` — comma-separated CIDR list of trusted
upstream proxies. Empty (default) disables XFF parsing entirely,
which is the safe choice when the deployment topology is unknown.
Examples:
``MFA_TRUSTED_PROXY_CIDRS=10.0.0.0/8,172.16.0.0/12``
``MFA_TRUSTED_PROXY_CIDRS=192.168.0.0/16``
"""
from __future__ import annotations
import logging
import os
from ipaddress import ip_address, ip_network
from typing import List
from fastapi import Request
from jose import jwt
logger = logging.getLogger(__name__)
def _parse_trusted_cidrs() -> List:
"""Parse the trusted-proxy CIDR list at import time.
A malformed entry is logged and skipped — we never crash the process
over a typo in operational config.
"""
raw = os.getenv("MFA_TRUSTED_PROXY_CIDRS", "").strip()
if not raw:
return []
nets = []
for entry in raw.split(","):
entry = entry.strip()
if not entry:
continue
try:
nets.append(ip_network(entry, strict=False))
except ValueError:
logger.warning(
"MFA_TRUSTED_PROXY_CIDRS: ignoring invalid CIDR %r", entry
)
return nets
_TRUSTED_NETS = _parse_trusted_cidrs()
def _peer_ip(request: Request) -> str:
"""The TCP peer IP — never raises; falls back to ``0.0.0.0``."""
return request.client.host if request.client else "0.0.0.0"
def _is_trusted_peer(peer: str) -> bool:
if not _TRUSTED_NETS:
return False
try:
peer_ip = ip_address(peer)
except ValueError:
return False
return any(peer_ip in net for net in _TRUSTED_NETS)
def _real_ip(request: Request) -> str:
"""If TCP peer is in a trusted proxy CIDR, peel off the first
``X-Forwarded-For`` IP; otherwise return the peer.
``X-Forwarded-For`` from an *untrusted* peer is intentionally ignored —
accepting it would let any client spoof their bucket.
"""
peer = _peer_ip(request)
if not _is_trusted_peer(peer):
return peer
xff = request.headers.get("X-Forwarded-For")
if not xff:
return peer
first = xff.split(",")[0].strip()
return first or peer
def _user_id_from_jwt(request: Request) -> str | None:
"""Sync JWT decode → ``user_id`` (or ``sub``) claim. None on failure.
Uses the same secret + algorithm as ``auth_middleware`` so a token that
is valid for the API surface is also valid for the rate-limit key.
Bad / missing / expired tokens silently return None — slowapi falls
back to IP bucketing.
"""
auth = request.headers.get("Authorization", "")
if not auth.startswith("Bearer "):
return None
token = auth[7:].strip()
if not token or token in {"null", "undefined"}:
return None
# Late import to avoid pulling jose into module-load when not needed.
try:
from config import JWT_ALGORITHM, JWT_SECRET_KEY
payload = jwt.decode(token, JWT_SECRET_KEY, algorithms=[JWT_ALGORITHM])
except Exception:
return None
uid = payload.get("user_id") or payload.get("sub")
if uid is None:
return None
return str(uid)
def mfa_rate_limit_key(request: Request) -> str:
"""slowapi ``key_func`` for MFA endpoints.
Order:
1. Authenticated → ``user:<id>``
2. Trusted-proxy XFF → ``ip:<first hop>``
3. TCP peer → ``ip:<peer>``
"""
uid = _user_id_from_jwt(request)
if uid is not None:
return f"user:{uid}"
return f"ip:{_real_ip(request)}"
+79
View File
@@ -0,0 +1,79 @@
"""MFA rate-limit configuration (env-overridable).
Best-practice pattern:
- Secure-by-default values live in code (kept in sync with the threat model).
- Operations can override per-environment via env vars (ConfigMap on K8s)
WITHOUT a code change / re-release.
- All limits funnel through a single named constant so the decorator stays
declarative (``@limiter.limit(MFA_LIMITS.enroll_start)``).
Env-var precedence::
MFA_RATE_LIMIT_<NAME> > default in code
slowapi limit string syntax: ``<count>/<period>`` where period is
``second|minute|hour|day``. Example: ``"5/minute"``.
NOTE: slowapi binds limits at import time. A change to an env var requires a
backend restart (rolling restart on K8s, ``docker compose restart backend``
locally). This is consistent with how ``SECRET_KEY`` / ``MFA_ENCRYPTION_KEY``
behave.
"""
from __future__ import annotations
import logging
import os
import re
from dataclasses import dataclass
logger = logging.getLogger(__name__)
# slowapi limit-string format guard. Keeps a typo from silently disabling
# rate-limiting at process start.
_LIMIT_RE = re.compile(r"^\d+/(second|minute|hour|day)$")
def _env(name: str, default: str) -> str:
"""Read ``MFA_RATE_LIMIT_<NAME>``; fall back to ``default``.
Validates the limit string. On bad input, logs a warning and returns the
secure default instead of crashing the process.
"""
value = os.getenv(f"MFA_RATE_LIMIT_{name}", default).strip()
if not _LIMIT_RE.match(value):
logger.warning(
"MFA_RATE_LIMIT_%s='%s' is not a valid slowapi limit string "
"(expected '<n>/<second|minute|hour|day>'); using default '%s'.",
name, value, default,
)
return default
return value
@dataclass(frozen=True)
class MfaRateLimits:
"""Aggregate of MFA endpoint rate-limit strings (slowapi format).
Defaults assume the rate-limit ``key_func`` is ``mfa_rate_limit_key``
(user-aware + ingress-aware), NOT raw IP. Per-user buckets are safe to
keep generous because a misbehaving user only burns their own quota and
cannot starve the rest of the org. If you re-key on raw IP, retighten
these values (see README + ``MFA_RATE_LIMIT_<NAME>`` env overrides).
"""
# Enrollment lifecycle — per-user buckets, large enough for org-wide rollout
enroll_start: str = _env("ENROLL_START", "10/minute")
enroll_confirm: str = _env("ENROLL_CONFIRM", "10/minute")
# Self-service maintenance
disable: str = _env("DISABLE", "10/minute")
regenerate_backup_codes: str = _env("REGENERATE_BACKUP_CODES", "5/hour")
# Admin operations (per-admin bucket; bulk reset stays tight because
# it is an emergency-only flow).
admin_reset: str = _env("ADMIN_RESET", "60/hour")
admin_reset_all: str = _env("ADMIN_RESET_ALL", "1/day")
# Module-level singleton — import this from routers/mfa.py.
MFA_LIMITS = MfaRateLimits()
+8 -2
View File
@@ -2,6 +2,7 @@
Rate Limiting Middleware for Production Security
Protects API endpoints from abuse and DDoS attacks
"""
import os
import time
import logging
from typing import Callable
@@ -16,10 +17,15 @@ from database.connection import redis_client
logger = logging.getLogger(__name__)
# Rate limiter instance using Redis backend
# Resolve the Redis storage URI from REDIS_URL (set by docker-compose / k8s
# ConfigMap). Falls back to the compose service hostname for backward
# compatibility when REDIS_URL is unset.
_REDIS_URL = os.getenv("REDIS_URL", "redis://redis:6379").rstrip("/")
_LIMITER_STORAGE_URI = f"{_REDIS_URL}/0" if "/" not in _REDIS_URL.split("//", 1)[-1] else _REDIS_URL
limiter = Limiter(
key_func=get_remote_address,
storage_uri="redis://redis:6379/0",
storage_uri=_LIMITER_STORAGE_URI,
default_limits=["1000/hour"], # Default global limit
retry_after=lambda name, t: int(t) + 10
)
+73
View File
@@ -0,0 +1,73 @@
"""MFA-specific Pydantic models — Issue #18, v1.6.0.
Kept in a separate module so the existing User / UserUpdate contracts in
``backend/models/user.py`` stay byte-identical for backwards compatibility.
"""
from typing import List, Literal, Optional
from pydantic import BaseModel, Field
class MfaVerifyRequest(BaseModel):
"""Body of POST /api/auth/login/mfa-verify (pre-auth — no JWT)."""
mfa_token: str = Field(..., min_length=64, max_length=64)
# 6 digits for TOTP or 8 alphanumerics (with optional dash) for backup codes
code: str = Field(..., min_length=6, max_length=10)
class MfaEnrollStartResponse(BaseModel):
"""Returned by POST /api/mfa/enroll/start."""
secret: str
otpauth_uri: str
expires_in: int # pending enrollment TTL in seconds
class MfaEnrollConfirmRequest(BaseModel):
"""Body of POST /api/mfa/enroll/confirm (TOTP only — backup codes not yet issued)."""
code: str = Field(..., min_length=6, max_length=6)
class MfaEnrollConfirmResponse(BaseModel):
enabled: bool
backup_codes: List[str]
method: Literal["totp"] = "totp"
class MfaDisableRequest(BaseModel):
"""Body of POST /api/mfa/disable — TOTP or backup."""
code: str = Field(..., min_length=6, max_length=10)
class MfaRegenerateBackupRequest(BaseModel):
"""Body of POST /api/mfa/backup-codes/regenerate — TOTP only."""
code: str = Field(..., min_length=6, max_length=6)
class MfaRegenerateBackupResponse(BaseModel):
backup_codes: List[str]
class MfaAdminResetRequest(BaseModel):
"""Body of POST /api/mfa/admin-reset/{user_id}."""
reason: str = Field(..., min_length=3, max_length=500)
class MfaAdminResetAllRequest(BaseModel):
"""Body of POST /api/mfa/admin-reset-all (emergency)."""
confirm: Literal["RESET ALL MFA"]
reason: str = Field(..., min_length=3, max_length=500)
class MfaStatusResponse(BaseModel):
enabled: bool
method: Optional[str] = None
enrolled_at: Optional[str] = None
last_used_at: Optional[str] = None
backup_codes_remaining: int = 0
+74
View File
@@ -488,6 +488,32 @@ class ServerStep(BaseModel):
"splits at the first space, so a space inside the "
"SNI value would corrupt the rendered config."
)
# Bulgu #90 (round-24 audit) — `cookie_value` is interpolated
# into the rendered `server <name> <addr>:<port> ... cookie
# <value> ...` line. HAProxy tokenises the line by whitespace
# and treats `;` as an INLINE COMMENT — and RFC 6265 cookie-
# value grammar separately disallows whitespace, `;`, `,`,
# `\\` and `"`. Pre-fix the validator only rejected newlines,
# so an operator could set `cookie_value="srv1; secure"` and
# see HAProxy silently truncate the server line at `;`
# (everything after becomes a comment). The Set-Cookie header
# rendered to clients would also fail the RFC's cookie-value
# grammar and most browsers drop the cookie, breaking session
# affinity without any error surface. Restrict to the
# conservative intersection — alphanumeric plus `_.-` —
# identical to backend.cookie_name. Legitimate session
# identifiers all fit in this set.
if info and info.field_name == "cookie_value":
import re as _re
if not _re.fullmatch(r"[A-Za-z0-9_.\-]+", v):
raise ValueError(
"server.cookie_value must contain only alphanumerics, "
"'.', '_' or '-' (HAProxy tokenises the server line "
"by whitespace and treats ';' as an inline comment; "
"RFC 6265 separately disallows whitespace, ';', ',' "
"and backslash in cookie values). Got "
f"{v!r}."
)
return v
# R17 (label corrected R18): CA bundle used by HAProxy to VERIFY the
# upstream server's TLS certificate. Maps to the `ca-file` directive
@@ -658,6 +684,32 @@ class BackendStep(BaseModel):
fields like `request_headers` / `response_headers` /
`tcp_request_rules` are intentionally line-oriented and ARE
NOT touched here.
Bulgu #90 (round-24 audit) — extend the validator beyond
newlines. Pre-fix `cookie_name` accepted strings like
`SESS'; DROP TABLE backends; --` (the SQL substring is
harmless thanks to parameterised queries, but the `;` is
HAProxy's INLINE COMMENT character: the renderer emits
`cookie SESS'; DROP TABLE backends; -- insert indirect
nocache`, which HAProxy parses as `cookie SESS'` followed by
an inline comment that swallows the persistence options the
operator typed). Result: session affinity silently broken
and the operator has zero diagnostic signal — the wizard
accepted the input, the apply succeeded, but cookies never
get re-emitted with the expected name. The same trap exists
for any HAProxy-section-keyword or whitespace token because
HAProxy tokenises by space.
`cookie_name` MUST be a single token. We use the conservative
intersection of RFC 6265 cookie-token chars and HAProxy
directive-name chars: alphanumeric plus `_.-`. Operators
with legitimate session cookies all live inside this set.
`cookie_options` is a space-separated keyword list
(`insert indirect nocache` etc., optionally `domain example.
com`, `attr SameSite=Lax`). Allow letters/digits/space/dot/
hyphen/underscore/equals. Reject `;` (comment), backslash,
quotes, and other shell metacharacters.
"""
if v is None:
return v
@@ -669,6 +721,28 @@ class BackendStep(BaseModel):
"newlines would smuggle additional directives into "
"the rendered config)"
)
field_name = info.field_name if info else "cookie field"
if field_name == "cookie_name":
import re as _re
if not _re.fullmatch(r"[A-Za-z0-9_.\-]+", v):
raise ValueError(
"backend.cookie_name must contain only alphanumerics, "
"'.', '_' or '-' (HAProxy tokenises the `cookie` "
"directive by whitespace and treats ';' as an inline "
"comment — anything else silently truncates the "
"rendered persistence options). Got "
f"{v!r}."
)
elif field_name == "cookie_options":
import re as _re
if not _re.fullmatch(r"[A-Za-z0-9_.=\- ]*", v):
raise ValueError(
"backend.cookie_options must contain only "
"alphanumerics, spaces, '=', '.', '_' or '-' "
"(HAProxy parses `;` as an inline comment, and "
"quoting / backslash metacharacters are not "
f"part of the `cookie` keyword grammar). Got {v!r}."
)
return v
@field_validator("health_check_uri")
+2 -1
View File
@@ -15,4 +15,5 @@ bcrypt>=4.0.1
slowapi>=0.1.9
psutil>=5.9.8
pytz>=2023.3
josepy>=1.14.0
josepy>=1.14.0
pyotp>=2.9.0
+433 -110
View File
@@ -16,9 +16,11 @@ Per-user 5/min rate-limit via user_activity_logs SQL count (M18 / R50).
import json
import logging
import uuid
from datetime import datetime
from typing import List, Optional
import asyncpg
from fastapi import APIRouter, Header, HTTPException
from auth_middleware import check_user_permission, get_current_user_from_token
@@ -58,16 +60,39 @@ async def _enforce_rate_limit(conn, user_id: int, action: str) -> None:
async def _load_order(conn, order_id: int) -> dict:
row = await conn.fetchrow(
"""
SELECT id, account_id, status, domains, cluster_ids, error_detail,
post_completion_actions, pending_apply_version_name,
wizard_staged_until, created_by
FROM letsencrypt_orders
WHERE id = $1
""",
order_id,
)
"""Fetch the order row, or raise a clean 404.
Bulgu #96 (prod-canary audit): `letsencrypt_orders.id` is a Postgres
int4 column. A path-param `order_id` outside the int4 range
(e.g. > 2_147_483_647) used to bubble up as
`asyncpg.exceptions.DataError: invalid input for query argument $1:
... (value out of int32 range)` — which the diagnostics endpoint
then surfaced in a `diagnostics_unavailable` envelope, leaking the
raw Postgres / asyncpg error string ("query argument $1",
"int32 range") into the operator-visible response body.
Semantically an out-of-range ID can never reference a real order,
so we treat it identically to "row not found" and return a clean
404 — same shape as the not-found path, no SQL detail leakage.
"""
try:
row = await conn.fetchrow(
"""
SELECT id, account_id, status, domains, cluster_ids, error_detail,
post_completion_actions, pending_apply_version_name,
wizard_staged_until, created_by
FROM letsencrypt_orders
WHERE id = $1
""",
order_id,
)
except asyncpg.exceptions.DataError as exc:
logger.info(
"ACME order lookup rejected by Postgres (out-of-range / "
"uncastable id): order_id=%s exc=%s",
order_id,
exc,
)
raise HTTPException(status_code=404, detail=f"Order {order_id} not found")
if not row:
raise HTTPException(status_code=404, detail=f"Order {order_id} not found")
return dict(row)
@@ -86,35 +111,153 @@ def _parse_jsonb_list(raw, default):
return default
def _diagnostic_failure_envelope(
order_id: int,
correlation_id: str,
exc: Exception,
*,
stage: str,
) -> dict:
"""Build a structured response when the diagnostic suite itself
cannot run. Bulgu #94 (Round-25): we return HTTP 200 with this
envelope rather than 500 so the UI can still SHOW the operator
what happened — the panel's whole purpose is to surface failure
causes, and the panel itself silently 500-ing is the worst-case
UX. The server-side log carries the full traceback keyed by
correlation_id for operator follow-up.
"""
return {
"order_id": order_id,
"status": "diagnostics_unavailable",
"checks": [
{
"id": "diagnostics_runner",
"label": "Diagnostic runner",
"status": "fail",
"severity": "error",
"message": (
f"Diagnostics could not run ({stage}): "
f"{exc.__class__.__name__}: {exc}"
),
"details": {
"stage": stage,
"exception_type": exc.__class__.__name__,
"exception_message": str(exc),
"correlation_id": correlation_id,
"hint": (
"Check the backend log for correlation_id "
f"{correlation_id} for the full traceback."
),
},
"duration_ms": None,
}
],
"humanized_error": {
"title": "Diagnostic panel could not run",
"message": (
"The diagnostic runner itself crashed before any check "
"could complete. This is independent of whether the ACME "
"provider is reachable from this cluster."
),
"hint": (
"Share the correlation_id below with the platform team; "
"they can grep the API log for the full traceback."
),
"correlation_id": correlation_id,
},
"meta": {
"correlation_id": correlation_id,
"error_stage": stage,
"error_type": exc.__class__.__name__,
"error_message": str(exc),
},
"generated_at": datetime.utcnow().isoformat() + "Z",
}
@router.post("/orders/{order_id}/diagnostics")
async def run_diagnostics(order_id: int, authorization: str = Header(None)):
"""Run the full pre-flight + post-failure diagnostic suite."""
"""Run the full pre-flight + post-failure diagnostic suite.
Bulgu #94 (Round-25 audit): this endpoint must NEVER return HTTP 500
for an in-suite failure. The diagnostic panel exists precisely to
explain what is broken; producing an opaque 500 defeats the entire
feature. Authentication / authorisation / rate-limit / not-found
errors still raise the appropriate 4xx, but any unexpected
exception from check execution is converted to a 200 response with
a structured failure envelope so the UI can display the cause.
"""
current_user = await get_current_user_from_token(authorization)
if not await check_user_permission(current_user["id"], "ssl", "read"):
raise HTTPException(status_code=403, detail="Insufficient permissions: ssl.read required")
correlation_id = uuid.uuid4().hex[:12]
conn = await get_database_connection()
try:
await _enforce_rate_limit(conn, current_user["id"], "acme_diagnostics_run")
order = await _load_order(conn, order_id)
except HTTPException:
await close_database_connection(conn)
raise
except Exception as exc: # noqa: BLE001 — diagnostic boundary
logger.exception(
"ACME diagnostics setup failed for order=%s correlation_id=%s",
order_id,
correlation_id,
)
try:
return _diagnostic_failure_envelope(order_id, correlation_id, exc, stage="load_order")
finally:
await close_database_connection(conn)
try:
domains = _parse_jsonb_list(order["domains"], [])
cluster_ids = _parse_jsonb_list(order["cluster_ids"], [])
results = await run_checks(
conn,
domains=domains,
cluster_ids=cluster_ids,
account_id=order["account_id"],
)
try:
results = await run_checks(
conn,
domains=domains,
cluster_ids=cluster_ids,
account_id=order["account_id"],
)
except Exception as exc: # noqa: BLE001 — diagnostic boundary
# run_checks now wraps individual checks, but a top-level
# crash (e.g. lost DB connection) still needs to be visible.
logger.exception(
"ACME diagnostics top-level failure for order=%s correlation_id=%s",
order_id,
correlation_id,
)
return _diagnostic_failure_envelope(order_id, correlation_id, exc, stage="run_checks")
humanized_error = humanize_error_detail(order["error_detail"])
try:
humanized_error = humanize_error_detail(order["error_detail"])
except Exception as exc: # noqa: BLE001 — defensive
logger.warning(
"humanize_error_detail failed for order=%s correlation_id=%s: %s",
order_id,
correlation_id,
exc,
)
humanized_error = {
"title": "ACME error (raw)",
"message": str(order["error_detail"]) if order["error_detail"] else "",
"hint": "",
"parse_error": exc.__class__.__name__,
}
return {
"order_id": order_id,
"status": order["status"],
"checks": results,
"humanized_error": humanized_error,
"meta": {
"correlation_id": correlation_id,
"checks_total": len(results),
"checks_failed": sum(1 for r in results if r.get("status") == "fail"),
"checks_warn": sum(1 for r in results if r.get("status") == "warn"),
},
"generated_at": datetime.utcnow().isoformat() + "Z",
}
finally:
@@ -127,7 +270,15 @@ async def rerun_diagnostic_check(
check_id: str,
authorization: str = Header(None),
):
"""Re-run a single check (DNS / port80 / routing / account / agents)."""
"""Re-run a single check (DNS / port80 / routing / account / agents).
Bulgu #94 follow-up (Round-25 audit): the rerun path is just as
sensitive to opaque 500s as the full-suite POST. If
`_enforce_rate_limit` / `_load_order` / `run_checks` raises an
unexpected exception, we surface a structured `fail` row in the
same shape the table already renders — so the operator clicking
"Re-run" never sees an opaque toast and the row updates in place.
"""
current_user = await get_current_user_from_token(authorization)
if not await check_user_permission(current_user["id"], "ssl", "read"):
raise HTTPException(status_code=403, detail="Insufficient permissions: ssl.read required")
@@ -138,30 +289,119 @@ async def rerun_diagnostic_check(
detail=f"Unknown check_id '{check_id}'. Valid: {', '.join(CHECK_IDS)}",
)
correlation_id = uuid.uuid4().hex[:12]
conn = await get_database_connection()
try:
await _enforce_rate_limit(conn, current_user["id"], "acme_diagnostic_check_rerun")
order = await _load_order(conn, order_id)
try:
await _enforce_rate_limit(conn, current_user["id"], "acme_diagnostic_check_rerun")
order = await _load_order(conn, order_id)
except HTTPException:
raise
except Exception as exc: # noqa: BLE001 — diagnostic boundary
logger.exception(
"ACME rerun setup failed for order=%s check=%s correlation_id=%s",
order_id,
check_id,
correlation_id,
)
return {
"order_id": order_id,
"check": {
"id": check_id,
"label": check_id,
"status": "fail",
"severity": "error",
"message": (
f"Re-run setup failed: "
f"{exc.__class__.__name__}: {exc}"
),
"details": {
"stage": "setup",
"exception_type": exc.__class__.__name__,
"exception_message": str(exc),
"correlation_id": correlation_id,
},
"duration_ms": None,
},
"meta": {"correlation_id": correlation_id, "error_stage": "setup"},
}
domains = _parse_jsonb_list(order["domains"], [])
cluster_ids = _parse_jsonb_list(order["cluster_ids"], [])
results = await run_checks(
conn,
domains=domains,
cluster_ids=cluster_ids,
account_id=order["account_id"],
only=[check_id],
)
try:
domains = _parse_jsonb_list(order["domains"], [])
cluster_ids = _parse_jsonb_list(order["cluster_ids"], [])
results = await run_checks(
conn,
domains=domains,
cluster_ids=cluster_ids,
account_id=order["account_id"],
only=[check_id],
)
except Exception as exc: # noqa: BLE001 — diagnostic boundary
logger.exception(
"ACME rerun run_checks failed for order=%s check=%s correlation_id=%s",
order_id,
check_id,
correlation_id,
)
return {
"order_id": order_id,
"check": {
"id": check_id,
"label": check_id,
"status": "fail",
"severity": "error",
"message": (
f"Re-run crashed: "
f"{exc.__class__.__name__}: {exc}"
),
"details": {
"stage": "run_checks",
"exception_type": exc.__class__.__name__,
"exception_message": str(exc),
"correlation_id": correlation_id,
},
"duration_ms": None,
},
"meta": {"correlation_id": correlation_id, "error_stage": "run_checks"},
}
return {
"order_id": order_id,
"check": results[0] if results else None,
"meta": {"correlation_id": correlation_id},
}
finally:
await close_database_connection(conn)
async def _user_activity_columns(conn) -> set:
"""Return the set of column names actually present on user_activity_logs.
Bulgu #95 (Round-25 audit) — the canonical migration for
`user_activity_logs` defines `id, user_id, action, resource_type,
resource_id, details, ip_address, user_agent, created_at, timestamp`.
There is NO `status` column. The original `/events` SELECT pulled
`status` directly, so every diagnostic-panel open against an order
that had any user-activity-log correlation raised
`UndefinedColumnError: column "status" does not exist` and the API
returned HTTP 500. We now introspect the schema and only project
the columns that exist, so deployments at any migration level keep
rendering the diagnostic panel.
"""
try:
rows = await conn.fetch(
"""
SELECT column_name
FROM information_schema.columns
WHERE table_name = 'user_activity_logs'
"""
)
return {r["column_name"] for r in rows}
except Exception as exc: # noqa: BLE001 — schema introspection is best-effort
logger.warning("user_activity_logs schema introspection failed: %s", exc)
return set()
@router.get("/orders/{order_id}/events")
async def get_order_events(
order_id: int,
@@ -174,6 +414,12 @@ async def get_order_events(
resource_id=order_id) for context.
Sorted by created_at ASC (oldest first) so the timeline reads naturally.
Bulgu #94/#95 (Round-25 audit): every sub-query is wrapped so that a
partial failure (missing column, missing table, malformed JSONB) is
surfaced via `meta.errors[]` rather than collapsing the whole panel
to HTTP 500. The diagnostic UI is a debugging surface — it must not
itself become opaque when one of its data sources is degraded.
"""
current_user = await get_current_user_from_token(authorization)
if not await check_user_permission(current_user["id"], "ssl", "read"):
@@ -182,95 +428,168 @@ async def get_order_events(
if limit <= 0 or limit > 500:
limit = 100
correlation_id = uuid.uuid4().hex[:12]
conn = await get_database_connection()
section_errors: List[dict] = []
try:
# Existence check
await _load_order(conn, order_id)
# Detect whether acme_order_events exists (zero-impact for envs that
# have not yet run the v1.5.0 migration). Returns empty event_log when
# not yet present rather than 500-ing.
events_table_exists = await conn.fetchval(
"""
SELECT EXISTS (
SELECT 1 FROM information_schema.tables WHERE table_name = 'acme_order_events'
try:
await _load_order(conn, order_id)
except HTTPException:
raise
except Exception as exc: # noqa: BLE001 — surface, don't 500
logger.exception(
"ACME events load_order failed order=%s correlation_id=%s",
order_id,
correlation_id,
)
"""
)
return {
"order_id": order_id,
"events": [],
"count": 0,
"meta": {
"correlation_id": correlation_id,
"errors": [
{
"section": "load_order",
"exception_type": exc.__class__.__name__,
"message": str(exc),
}
],
},
}
events: List[dict] = []
if events_table_exists:
event_rows = await conn.fetch(
# --- Section 1: acme_order_events ---
try:
events_table_exists = await conn.fetchval(
"""
SELECT id, event_type, severity, message, details, correlation_id, created_at
FROM acme_order_events
WHERE order_id = $1
ORDER BY created_at ASC, id ASC
LIMIT $2
""",
SELECT EXISTS (
SELECT 1 FROM information_schema.tables
WHERE table_name = 'acme_order_events'
)
"""
)
if events_table_exists:
event_rows = await conn.fetch(
"""
SELECT id, event_type, severity, message, details, correlation_id, created_at
FROM acme_order_events
WHERE order_id = $1
ORDER BY created_at ASC, id ASC
LIMIT $2
""",
order_id,
limit,
)
for r in event_rows:
# R18c round 8 (Bulgu A): asyncpg returns JSONB columns as
# raw JSON strings (no codec on the pool). For the FE
# contract the `details` field MUST be either a dict or
# null — otherwise the React renderer ends up trying to
# access `details.foo` on a plain string and silently
# gets undefined.
_det = r["details"]
if isinstance(_det, str):
try:
_det = json.loads(_det)
except Exception:
_det = {}
if not isinstance(_det, (dict, list)):
_det = {} if _det is None else {"raw": str(_det)}
events.append({
"source": "acme_order_event",
"id": r["id"],
"event_type": r["event_type"],
"severity": r["severity"],
"message": r["message"],
"details": _det,
"correlation_id": r["correlation_id"],
"created_at": r["created_at"].isoformat().replace("+00:00", "Z")
if r["created_at"] else None,
})
except Exception as exc: # noqa: BLE001 — surface and continue
logger.exception(
"ACME events acme_order_events query failed order=%s correlation_id=%s",
order_id,
limit,
correlation_id,
)
for r in event_rows:
# R18c round 8 (Bulgu A): asyncpg returns JSONB columns as
# raw JSON strings (no codec on the pool). For the FE
# contract the `details` field MUST be either a dict or
# null — otherwise the React renderer ends up trying to
# access `details.foo` on a plain string and silently
# gets undefined.
_det = r["details"]
if isinstance(_det, str):
try:
_det = json.loads(_det)
except Exception:
_det = {}
if not isinstance(_det, (dict, list)):
_det = {} if _det is None else {"raw": str(_det)}
events.append({
"source": "acme_order_event",
"id": r["id"],
"event_type": r["event_type"],
"severity": r["severity"],
"message": r["message"],
"details": _det,
"correlation_id": r["correlation_id"],
"created_at": r["created_at"].isoformat().replace("+00:00", "Z")
if r["created_at"] else None,
section_errors.append({
"section": "acme_order_events",
"exception_type": exc.__class__.__name__,
"message": str(exc),
})
# --- Section 2: user_activity_logs (best-effort, schema-aware) ---
try:
ua_columns = await _user_activity_columns(conn)
if "resource_id" in ua_columns and "resource_type" in ua_columns:
# Project only columns we know exist. `status` is NOT in
# the canonical schema and was the original 500 cause.
projection_candidates = [
"id", "action", "resource_type", "resource_id",
"details", "created_at", "user_id", "status",
]
projection = [c for c in projection_candidates if c in ua_columns]
if "id" not in projection or "created_at" not in projection:
raise RuntimeError(
"user_activity_logs is missing required columns "
"(id / created_at) — skipping correlation"
)
sql = (
f"SELECT {', '.join(projection)} "
"FROM user_activity_logs "
"WHERE resource_type = 'letsencrypt_order' AND resource_id = $1 "
"ORDER BY created_at ASC, id ASC LIMIT $2"
)
ua_rows = await conn.fetch(sql, str(order_id), limit)
for r in ua_rows:
rd = dict(r)
raw_details = rd.get("details")
if isinstance(raw_details, str):
msg = raw_details[:500]
details_obj = {}
try:
parsed = json.loads(raw_details)
if isinstance(parsed, (dict, list)):
details_obj = parsed
except Exception:
details_obj = {}
elif isinstance(raw_details, (dict, list)):
msg = ""
details_obj = raw_details
else:
msg = ""
details_obj = {}
raw_status = rd.get("status") or ""
severity = "info" if str(raw_status).lower() in ("success", "ok", "") else "warn"
events.append({
"source": "user_activity_log",
"id": rd.get("id"),
"event_type": rd.get("action"),
"severity": severity,
"message": msg,
"details": details_obj,
"correlation_id": None,
"created_at": rd["created_at"].isoformat().replace("+00:00", "Z")
if rd.get("created_at") else None,
})
else:
section_errors.append({
"section": "user_activity_logs",
"exception_type": "SchemaMissing",
"message": "user_activity_logs lacks resource_type/resource_id columns",
})
# User activity rows correlated by resource — schema is permissive
# (`resource_type`/`resource_id` may not always be populated for older
# rows) so this query stays best-effort.
ua_rows = await conn.fetch(
"""
SELECT id, action, resource_type, resource_id, status, details, created_at, user_id
FROM user_activity_logs
WHERE resource_type = 'letsencrypt_order' AND resource_id = $1
ORDER BY created_at ASC, id ASC
LIMIT $2
""",
str(order_id),
limit,
) if await conn.fetchval(
"""
SELECT EXISTS (
SELECT 1 FROM information_schema.columns
WHERE table_name = 'user_activity_logs' AND column_name = 'resource_id'
except Exception as exc: # noqa: BLE001 — surface and continue
logger.exception(
"ACME events user_activity_logs query failed order=%s correlation_id=%s",
order_id,
correlation_id,
)
"""
) else []
for r in ua_rows:
events.append({
"source": "user_activity_log",
"id": r["id"],
"event_type": r["action"],
"severity": "info" if (r["status"] or "").lower() in ("success", "ok", "") else "warn",
"message": (r["details"] or "")[:500] if isinstance(r["details"], str) else "",
"details": r["details"] if not isinstance(r["details"], (str, type(None))) else {},
"correlation_id": None,
"created_at": r["created_at"].isoformat().replace("+00:00", "Z")
if r["created_at"] else None,
section_errors.append({
"section": "user_activity_logs",
"exception_type": exc.__class__.__name__,
"message": str(exc),
})
events.sort(key=lambda e: (e["created_at"] or "", e.get("id") or 0))
@@ -279,6 +598,10 @@ async def get_order_events(
"order_id": order_id,
"events": events,
"count": len(events),
"meta": {
"correlation_id": correlation_id,
"errors": section_errors,
},
}
finally:
await close_database_connection(conn)
+427 -10
View File
@@ -9,14 +9,48 @@ from datetime import datetime, timedelta
# Import database and models
from database.connection import get_database_connection, close_database_connection
from models.user import LoginRequest, User, UserCreate, UserUpdate, UserPasswordUpdate
from models.mfa import MfaVerifyRequest
from utils.activity_log import log_user_activity
from auth_middleware import get_current_user_from_token
from services import mfa_service
# Rate limiting temporarily disabled
router = APIRouter(prefix="/api/auth", tags=["Authentication"])
logger = logging.getLogger(__name__)
MFA_PENDING_TTL_SECONDS = 300 # 5 minutes — pre-verification challenge lifetime
MFA_PENDING_MAX_ATTEMPTS = 5 # invalidate token after this many wrong codes
async def _fetch_mfa_state(conn, user_id: int):
"""Return (mfa_enabled, mfa_secret_encrypted, mfa_last_used_totp_step) or None
when the MFA columns aren't yet present (pre-migration deploys).
"""
try:
return await conn.fetchrow(
"""
SELECT mfa_enabled, mfa_secret_encrypted, mfa_last_used_totp_step
FROM users
WHERE id = $1
""",
user_id,
)
except Exception as exc:
logger.warning(f"MFA columns not available (assuming disabled): {exc}")
return None
async def _cleanup_expired_pending_logins(conn, user_id: int) -> None:
"""Lazy cleanup of expired pending MFA challenges for this user."""
try:
await conn.execute(
"DELETE FROM mfa_pending_logins WHERE user_id = $1 AND expires_at < NOW()",
user_id,
)
except Exception as exc:
logger.debug(f"Pending-login cleanup skipped: {exc}")
# Security scheme
security = HTTPBearer()
@@ -96,7 +130,7 @@ async def login(login_request: LoginRequest, request: Request):
SELECT id, username, email, password_hash, is_active, role,
created_at, updated_at, last_login_at
FROM users
WHERE username = $1
WHERE username = $1 AND is_active = TRUE
""", login_request.username)
except Exception as schema_error:
logger.warning(f"Schema error, trying fallback query: {schema_error}")
@@ -106,7 +140,7 @@ async def login(login_request: LoginRequest, request: Request):
SELECT id, username, email, password_hash, is_active,
created_at, updated_at, last_login_at
FROM users
WHERE username = $1
WHERE username = $1 AND is_active = TRUE
""", login_request.username)
except Exception as column_error:
logger.warning(f"last_login_at column error, trying last_login: {column_error}")
@@ -115,26 +149,67 @@ async def login(login_request: LoginRequest, request: Request):
SELECT id, username, email, password_hash, is_active,
created_at, updated_at, last_login
FROM users
WHERE username = $1
WHERE username = $1 AND is_active = TRUE
""", login_request.username)
if not user:
# Covers both "no such user" and "soft-deleted (is_active=FALSE)".
# We deliberately return the same generic 401 in either case to
# avoid leaking whether an account exists (account enumeration
# prevention). Soft-deleted rows are filtered out by the
# `AND is_active = TRUE` predicate above.
await close_database_connection(conn)
logger.warning(f"Failed login attempt for username: {login_request.username}")
raise HTTPException(status_code=401, detail="Invalid username or password")
if not user['is_active']:
await close_database_connection(conn)
logger.warning(f"Login attempt for inactive account: {login_request.username}")
raise HTTPException(status_code=401, detail="Account is deactivated")
# Verify password
import bcrypt
if not bcrypt.checkpw(login_request.password.encode('utf-8'), user['password_hash'].encode('utf-8')):
await close_database_connection(conn)
logger.warning(f"Wrong password for user: {login_request.username}")
raise HTTPException(status_code=401, detail="Invalid username or password")
# Issue #18 — MFA branch (v1.6.0): if the user opted in, defer JWT mint and
# last_login update until /api/auth/login/mfa-verify completes.
mfa_state = await _fetch_mfa_state(conn, user['id'])
if mfa_state and mfa_state.get('mfa_enabled'):
await _cleanup_expired_pending_logins(conn, user['id'])
challenge_token = mfa_service.generate_challenge_token()
try:
await conn.execute(
"""
INSERT INTO mfa_pending_logins (user_id, challenge_token, expires_at, ip_address)
VALUES ($1, $2, NOW() + ($3 || ' seconds')::interval, $4)
""",
user['id'],
challenge_token,
str(MFA_PENDING_TTL_SECONDS),
str(request.client.host) if request.client else None,
)
except Exception as exc:
await close_database_connection(conn)
logger.error(f"Failed to create MFA pending login: {exc}")
raise HTTPException(status_code=500, detail="MFA challenge creation failed")
await close_database_connection(conn)
await log_user_activity(
user_id=user['id'],
action='mfa.login.challenge_issued',
resource_type='mfa',
resource_id=str(user['id']),
details={'login_method': 'username_password'},
ip_address=str(request.client.host) if request.client else None,
user_agent=request.headers.get('user-agent'),
)
return {
"mfa_required": True,
"mfa_token": challenge_token,
"methods": ["totp", "backup"],
"expires_in": MFA_PENDING_TTL_SECONDS,
}
# Update last login (try different column names)
try:
await conn.execute("""
@@ -242,6 +317,348 @@ async def login(login_request: LoginRequest, request: Request):
logger.error(f"Login error: {e}")
raise HTTPException(status_code=500, detail="Login failed")
@router.post(
"/login/mfa-verify",
summary="MFA Verification (Step 2 of Login)",
response_description="JWT access token after successful TOTP/backup verification",
)
async def login_mfa_verify(payload: MfaVerifyRequest, request: Request):
"""
# MFA Verification — Step 2 of the two-step login flow
Submit a 6-digit TOTP code OR an 8-character backup code (with optional dash)
together with the ``mfa_token`` returned by ``POST /api/auth/login`` for an
MFA-enabled account. On success, returns the same response shape as a
non-MFA login (Branch A).
## Request Body
- **mfa_token**: 64-char challenge token from /login response
- **code**: 6 digits (TOTP) or `XXXX-YYYY` (backup)
## Error Responses
- **401**: Invalid code (attempts counter increments)
- **410**: Challenge expired or invalidated (too many wrong attempts)
"""
ip_address = str(request.client.host) if request.client else None
user_agent = request.headers.get('user-agent')
# Outcome captured from the transactional block so we can do JWT mint /
# activity logging AFTER commit (no side effects on rollback).
success_payload: Optional[dict] = None
failure: Optional[dict] = None # { user_id, attempts, invalidated, reason, http_status, detail }
conn = None
try:
conn = await get_database_connection()
# Round 1 audit fix — wrap the whole verify+update in a single
# transaction with row-level locks (FOR UPDATE) so two concurrent
# /mfa-verify calls cannot both consume the same TOTP step or the
# same pending challenge. We never raise inside the transaction once
# we've started mutating the pending row (would rollback the mark);
# instead we capture `failure` and raise after commit.
async with conn.transaction():
pending = await conn.fetchrow(
"""
SELECT id, user_id, attempts, expires_at, used_at
FROM mfa_pending_logins
WHERE challenge_token = $1
FOR UPDATE
""",
payload.mfa_token,
)
if not pending:
failure = {
'user_id': None,
'attempts': 0,
'invalidated': True,
'reason': 'challenge_not_found',
'http_status': 410,
'detail': 'MFA challenge not found or expired',
}
elif pending['used_at'] is not None:
failure = {
'user_id': pending['user_id'],
'attempts': pending['attempts'],
'invalidated': True,
'reason': 'challenge_already_used',
'http_status': 410,
'detail': 'MFA challenge already used',
}
elif pending['expires_at'] and pending['expires_at'] < datetime.utcnow():
failure = {
'user_id': pending['user_id'],
'attempts': pending['attempts'],
'invalidated': True,
'reason': 'challenge_expired',
'http_status': 410,
'detail': 'MFA challenge expired',
}
elif pending['attempts'] >= MFA_PENDING_MAX_ATTEMPTS:
await conn.execute(
"UPDATE mfa_pending_logins SET used_at = NOW() WHERE id = $1",
pending['id'],
)
failure = {
'user_id': pending['user_id'],
'attempts': pending['attempts'],
'invalidated': True,
'reason': 'too_many_attempts_pre_check',
'http_status': 410,
'detail': 'MFA challenge invalidated (too many attempts)',
}
if failure is None:
# Lock the user row so the atomic TOTP-step bump cannot race a
# parallel verify on a different pending challenge for the
# same account.
user_row = await conn.fetchrow(
"""
SELECT id, username, email, role, is_active,
created_at, updated_at, last_login_at,
mfa_secret_encrypted, mfa_last_used_totp_step
FROM users
WHERE id = $1
FOR UPDATE
""",
pending['user_id'],
)
if not user_row or not user_row['is_active']:
failure = {
'user_id': pending['user_id'],
'attempts': pending['attempts'],
'invalidated': False,
'reason': 'user_unavailable',
'http_status': 401,
'detail': 'User not available',
}
elif not user_row['mfa_secret_encrypted']:
failure = {
'user_id': user_row['id'],
'attempts': pending['attempts'],
'invalidated': True,
'reason': 'mfa_not_configured',
'http_status': 410,
'detail': 'MFA not configured for this user',
}
else:
secret_plain = mfa_service.decrypt_secret(user_row['mfa_secret_encrypted'])
verified_method: Optional[str] = None
codes_remaining: Optional[int] = None
if secret_plain:
ok, step = mfa_service.verify_totp_with_replay_guard(
secret_plain, payload.code, user_row['mfa_last_used_totp_step']
)
if ok:
# Atomic step bump — refuse if another request already
# consumed this (or a newer) TOTP step.
bumped = await conn.fetchval(
"""
UPDATE users
SET mfa_last_used_totp_step = $1,
mfa_last_used_at = NOW()
WHERE id = $2
AND (mfa_last_used_totp_step IS NULL
OR mfa_last_used_totp_step < $1)
RETURNING id
""",
step,
user_row['id'],
)
if bumped:
verified_method = 'totp'
if verified_method is None:
# Backup codes — atomic single-use consumption.
rows = await conn.fetch(
"""
SELECT id, code_hash FROM mfa_backup_codes
WHERE user_id = $1 AND used_at IS NULL
""",
user_row['id'],
)
for row in rows:
if await mfa_service.check_backup_code(payload.code, row['code_hash']):
consumed_id = await conn.fetchval(
"""
UPDATE mfa_backup_codes
SET used_at = NOW()
WHERE id = $1 AND used_at IS NULL
RETURNING id
""",
row['id'],
)
if consumed_id:
verified_method = 'backup'
codes_remaining = await conn.fetchval(
"SELECT COUNT(*) FROM mfa_backup_codes WHERE user_id = $1 AND used_at IS NULL",
user_row['id'],
)
break
if verified_method is None:
new_attempts = pending['attempts'] + 1
invalidated = new_attempts >= MFA_PENDING_MAX_ATTEMPTS
await conn.execute(
"""
UPDATE mfa_pending_logins
SET attempts = $1,
used_at = CASE WHEN $2 THEN NOW() ELSE used_at END
WHERE id = $3
""",
new_attempts,
invalidated,
pending['id'],
)
failure = {
'user_id': user_row['id'],
'attempts': new_attempts,
'invalidated': invalidated,
'reason': 'invalid_code',
'http_status': 410 if invalidated else 401,
'detail': 'MFA challenge invalidated (too many attempts)'
if invalidated else 'Invalid MFA code',
}
else:
# Verified — finalize state inside the transaction so a
# concurrent verify sees used_at on retry.
await conn.execute(
"UPDATE mfa_pending_logins SET used_at = NOW() WHERE id = $1",
pending['id'],
)
if verified_method != 'totp':
await conn.execute(
"UPDATE users SET mfa_last_used_at = NOW() WHERE id = $1",
user_row['id'],
)
try:
await conn.execute(
"UPDATE users SET last_login_at = CURRENT_TIMESTAMP WHERE id = $1",
user_row['id'],
)
except Exception as exc:
logger.warning(f"last_login_at update failed (continuing): {exc}")
user_roles = await conn.fetch(
"""
SELECT r.id, r.name, r.display_name, r.permissions
FROM user_roles ur
JOIN roles r ON ur.role_id = r.id
WHERE ur.user_id = $1 AND ur.is_active = TRUE AND r.is_active = TRUE
""",
user_row['id'],
)
permissions: dict = {}
roles_list: list = []
for role_row in user_roles:
roles_list.append({
'id': role_row['id'],
'name': role_row['name'],
'display_name': role_row['display_name'],
})
role_permissions = role_row['permissions']
if isinstance(role_permissions, str):
import json
role_permissions = json.loads(role_permissions)
if role_permissions:
for perm in role_permissions:
if '.' in perm:
resource, action = perm.split('.', 1)
if resource not in permissions:
permissions[resource] = {}
permissions[resource][action] = True
success_payload = {
'user': dict(user_row),
'roles_list': roles_list,
'permissions': permissions,
'method': verified_method,
'codes_remaining': codes_remaining,
}
# ------------------------------------------------------------------
# Transaction has committed. Side-effects (JWT mint, audit log) below.
# ------------------------------------------------------------------
await close_database_connection(conn)
conn = None
if failure is not None:
if failure['user_id'] is not None:
await log_user_activity(
user_id=failure['user_id'],
action='mfa.login.failed',
resource_type='mfa',
resource_id=str(failure['user_id']),
details={
'reason': failure['reason'],
'attempts': failure['attempts'],
'invalidated': failure['invalidated'],
},
ip_address=ip_address,
user_agent=user_agent,
)
raise HTTPException(status_code=failure['http_status'], detail=failure['detail'])
# Success path
assert success_payload is not None # for type checkers; transaction guarantees this
user_row = success_payload['user']
from jose import jwt
from config import JWT_SECRET_KEY, JWT_ALGORITHM
token_payload = {
"user_id": user_row['id'],
"username": user_row['username'],
"email": user_row['email'],
"role": user_row['role'] if 'role' in user_row.keys() else 'admin',
"exp": datetime.utcnow() + timedelta(hours=24),
}
token = jwt.encode(token_payload, JWT_SECRET_KEY, algorithm=JWT_ALGORITHM)
await log_user_activity(
user_id=user_row['id'],
action='mfa.login.success',
resource_type='mfa',
resource_id=str(user_row['id']),
details={
'method': success_payload['method'],
'codes_remaining': success_payload['codes_remaining']
if success_payload['method'] == 'backup' else None,
},
ip_address=ip_address,
user_agent=user_agent,
)
return {
"access_token": token,
"token_type": "bearer",
"expires_in": 86400,
"user": {
"id": user_row['id'],
"username": user_row['username'],
"email": user_row['email'],
"role": user_row['role'] if 'role' in user_row.keys() else 'admin',
"is_active": user_row['is_active'],
"created_at": user_row['created_at'].isoformat() if user_row.get('created_at') else None,
"last_login_at": datetime.utcnow().isoformat(),
},
"roles": success_payload['roles_list'],
"permissions": success_payload['permissions'],
}
except HTTPException:
raise
except Exception as exc:
logger.error(f"MFA verify error: {exc}")
raise HTTPException(status_code=500, detail="MFA verification failed")
finally:
if conn is not None:
await close_database_connection(conn)
@router.post("/logout", summary="User Logout", response_description="Logout confirmation")
async def logout(request: Request, authorization: str = Header(None)):
"""
+21 -5
View File
@@ -198,12 +198,28 @@ def _enforce_routing_rule_contradictions(
for label, rule in conflicts:
sig = _rule_to_signature(rule)
if sig and sig in grandfathered_signatures:
# Bulgu #83 (round-23 audit) — re-word the operator-
# facing warning. The pre-fix message led with
# "Grandfathered <label> entry contains a self-
# contradictory X !X condition that pre-dated this
# validation", which (a) is internal jargon the
# operator does not parse, and (b) implies the rule
# is OLD when in fact the only thing this branch
# knows is that the rule was NOT changed by the
# current edit. The operator may well have authored
# the rule one minute earlier. State that explicitly
# and include the verbatim rule body so the operator
# does not have to hunt through the ACL Builder
# cards to find the offender.
rule_text = rule if isinstance(rule, str) else sig[:160]
warnings.append(
f"Grandfathered {label} entry contains a "
f"self-contradictory `X !X` condition that pre-dated "
f"this validation. The rule never fires; fix it at "
f"your convenience. (rule: "
f"{rule if isinstance(rule, str) else sig[:160]})"
f"{label}: rule was not modified by this edit but "
f"contains a self-contradictory `X !X` condition "
f"(`X AND NOT X` is always false, so the rule never "
f"fires and traffic falls through to "
f"`default_backend`). Your current edit was saved; "
f"fix the rule at your convenience. "
f"(rule: {rule_text})"
)
else:
blocking.append((label, rule))
+727
View File
@@ -0,0 +1,727 @@
"""MFA router — Issue #18, v1.6.0.
All endpoints are additive; nothing here breaks existing JWT or apply_service flows.
Authentication is JWT-based (Bearer). Admin endpoints additionally require
``users.is_admin == True`` (canonical super-admin flag).
"""
from __future__ import annotations
import logging
from typing import Optional
from fastapi import APIRouter, Header, HTTPException, Request
from auth_middleware import get_current_user_from_token
from database.connection import close_database_connection, get_database_connection
from middleware.mfa_rate_limit_key import mfa_rate_limit_key
from middleware.mfa_rate_limits import MFA_LIMITS
from middleware.rate_limiter import limiter
from models.mfa import (
MfaAdminResetAllRequest,
MfaAdminResetRequest,
MfaDisableRequest,
MfaEnrollConfirmRequest,
MfaEnrollConfirmResponse,
MfaEnrollStartResponse,
MfaRegenerateBackupRequest,
MfaRegenerateBackupResponse,
MfaStatusResponse,
)
from services import mfa_service
from utils.activity_log import log_user_activity
logger = logging.getLogger(__name__)
router = APIRouter(prefix="/api/mfa", tags=["MFA"])
# Lifecycle constants (Plan section 5)
PENDING_ENROLLMENT_TTL_SECONDS = 600 # 10 minutes — QR scan + verify window
PENDING_ENROLLMENT_MAX_ATTEMPTS = 5
# ---------------------------------------------------------------------------
# Authentication helpers
# ---------------------------------------------------------------------------
async def _require_user(authorization: Optional[str]) -> dict:
user = await get_current_user_from_token(authorization)
if not user:
raise HTTPException(status_code=401, detail="Not authenticated")
return user
async def _require_admin(authorization: Optional[str]) -> dict:
user = await _require_user(authorization)
if not user.get("is_admin"):
raise HTTPException(status_code=403, detail="Admin privileges required")
return user
async def _cleanup_expired_pending_enrollments(conn) -> None:
try:
await conn.execute(
"DELETE FROM mfa_pending_enrollments WHERE expires_at < NOW()"
)
except Exception as exc:
logger.debug(f"Pending-enrollment cleanup skipped: {exc}")
# ---------------------------------------------------------------------------
# Self status
# ---------------------------------------------------------------------------
@router.get(
"/status",
summary="MFA status for the authenticated user",
response_model=MfaStatusResponse,
)
async def mfa_status(authorization: str = Header(None)):
current = await _require_user(authorization)
conn = await get_database_connection()
try:
row = await conn.fetchrow(
"""
SELECT mfa_enabled, mfa_method, mfa_enrolled_at, mfa_last_used_at
FROM users
WHERE id = $1
""",
current["id"],
)
remaining = await conn.fetchval(
"""
SELECT COUNT(*) FROM mfa_backup_codes
WHERE user_id = $1 AND used_at IS NULL
""",
current["id"],
)
finally:
await close_database_connection(conn)
if not row:
raise HTTPException(status_code=404, detail="User not found")
return MfaStatusResponse(
enabled=bool(row["mfa_enabled"]),
method=row["mfa_method"],
enrolled_at=row["mfa_enrolled_at"].isoformat() if row["mfa_enrolled_at"] else None,
last_used_at=row["mfa_last_used_at"].isoformat() if row["mfa_last_used_at"] else None,
backup_codes_remaining=int(remaining or 0),
)
@router.get(
"/admin/status/{user_id}",
summary="Admin: MFA status of any user",
response_model=MfaStatusResponse,
)
async def mfa_admin_status(user_id: int, authorization: str = Header(None)):
await _require_admin(authorization)
conn = await get_database_connection()
try:
row = await conn.fetchrow(
"""
SELECT mfa_enabled, mfa_method, mfa_enrolled_at, mfa_last_used_at
FROM users
WHERE id = $1
""",
user_id,
)
remaining = await conn.fetchval(
"""
SELECT COUNT(*) FROM mfa_backup_codes
WHERE user_id = $1 AND used_at IS NULL
""",
user_id,
)
finally:
await close_database_connection(conn)
if not row:
raise HTTPException(status_code=404, detail="User not found")
return MfaStatusResponse(
enabled=bool(row["mfa_enabled"]),
method=row["mfa_method"],
enrolled_at=row["mfa_enrolled_at"].isoformat() if row["mfa_enrolled_at"] else None,
last_used_at=row["mfa_last_used_at"].isoformat() if row["mfa_last_used_at"] else None,
backup_codes_remaining=int(remaining or 0),
)
# ---------------------------------------------------------------------------
# Enrollment
# ---------------------------------------------------------------------------
@router.post(
"/enroll/start",
summary="Begin TOTP enrollment (returns secret + otpauth URI)",
response_model=MfaEnrollStartResponse,
)
@limiter.limit(MFA_LIMITS.enroll_start, key_func=mfa_rate_limit_key)
async def mfa_enroll_start(request: Request, authorization: str = Header(None)):
current = await _require_user(authorization)
# Round 7 audit fix — REFUSE re-enrollment if the user is already MFA-on.
# Without this guard a stolen JWT could silently rotate the victim's TOTP
# secret + invalidate all their backup codes via /enroll/start ->
# /enroll/confirm (overwriting `users.mfa_secret_encrypted` and replacing
# `mfa_backup_codes`). To re-enroll, the user must first call /api/mfa/disable
# (which requires a fresh TOTP) or an admin must run /api/mfa/admin-reset.
secret_plain = mfa_service.generate_totp_secret()
secret_encrypted = mfa_service.encrypt_secret(secret_plain)
conn = await get_database_connection()
blocked = False
try:
# Single transaction with SELECT FOR UPDATE closes the TOCTOU window
# between the mfa_enabled check and the pending_enrollment upsert.
async with conn.transaction():
row = await conn.fetchrow(
"SELECT mfa_enabled FROM users WHERE id = $1 FOR UPDATE",
current["id"],
)
if row and row["mfa_enabled"]:
blocked = True
else:
await _cleanup_expired_pending_enrollments(conn)
await conn.execute(
"""
INSERT INTO mfa_pending_enrollments
(user_id, secret_encrypted, attempts, expires_at)
VALUES ($1, $2, 0, NOW() + ($3 || ' seconds')::interval)
ON CONFLICT (user_id) DO UPDATE
SET secret_encrypted = EXCLUDED.secret_encrypted,
attempts = 0,
expires_at = EXCLUDED.expires_at,
created_at = CURRENT_TIMESTAMP
""",
current["id"],
secret_encrypted,
str(PENDING_ENROLLMENT_TTL_SECONDS),
)
finally:
await close_database_connection(conn)
if blocked:
raise HTTPException(
status_code=400,
detail="MFA is already enabled. Disable it first (via /api/mfa/disable or admin reset) to re-enroll.",
)
hostname_hint = request.url.hostname if request.url else None
label = mfa_service.build_account_label(current["username"], hostname_hint)
otpauth_uri = mfa_service.build_otpauth_uri(label, secret_plain)
await log_user_activity(
user_id=current["id"],
action="mfa.enrollment.started",
resource_type="mfa",
resource_id=str(current["id"]),
details={"secret_len": len(secret_plain)}, # NEVER log the secret itself
ip_address=str(request.client.host) if request.client else None,
user_agent=request.headers.get("user-agent"),
)
return MfaEnrollStartResponse(
secret=secret_plain,
otpauth_uri=otpauth_uri,
expires_in=PENDING_ENROLLMENT_TTL_SECONDS,
)
@router.post(
"/enroll/confirm",
summary="Confirm enrollment with a TOTP code; returns 10 backup codes once",
response_model=MfaEnrollConfirmResponse,
)
@limiter.limit(MFA_LIMITS.enroll_confirm, key_func=mfa_rate_limit_key)
async def mfa_enroll_confirm(
payload: MfaEnrollConfirmRequest,
request: Request,
authorization: str = Header(None),
):
current = await _require_user(authorization)
# Pre-generate plain codes & hashes outside the DB transaction so we
# never hold a row lock for ~2-3s of bcrypt work.
plain_codes = mfa_service.generate_backup_codes()
hashes = await mfa_service.hash_backup_codes(plain_codes)
ip = str(request.client.host) if request.client else None
ua = request.headers.get("user-agent")
# failure: { reason, attempts, http_status, detail }; success when None.
failure: Optional[dict] = None
conn = await get_database_connection()
try:
async with conn.transaction():
# Lock the pending row so concurrent /enroll/confirm calls for the
# same user can't both consume the same pending enrollment.
pending = await conn.fetchrow(
"""
SELECT secret_encrypted, attempts, expires_at
FROM mfa_pending_enrollments
WHERE user_id = $1
FOR UPDATE
""",
current["id"],
)
if not pending:
failure = {
"reason": "no_pending",
"http_status": 410,
"detail": "No pending enrollment; start again",
}
else:
from datetime import datetime as _dt
if pending["expires_at"] and pending["expires_at"] < _dt.utcnow():
await conn.execute(
"DELETE FROM mfa_pending_enrollments WHERE user_id = $1",
current["id"],
)
failure = {
"reason": "expired",
"http_status": 410,
"detail": "Enrollment expired; start again",
}
else:
secret_plain = mfa_service.decrypt_secret(pending["secret_encrypted"])
if not secret_plain:
await conn.execute(
"DELETE FROM mfa_pending_enrollments WHERE user_id = $1",
current["id"],
)
failure = {
"reason": "unreadable",
"http_status": 500,
"detail": "Pending enrollment unreadable; start again",
}
else:
ok, _step = mfa_service.verify_totp_with_replay_guard(
secret_plain, payload.code, None
)
if not ok:
new_attempts = (pending["attempts"] or 0) + 1
if new_attempts >= PENDING_ENROLLMENT_MAX_ATTEMPTS:
await conn.execute(
"DELETE FROM mfa_pending_enrollments WHERE user_id = $1",
current["id"],
)
failure = {
"reason": "too_many_attempts",
"attempts": new_attempts,
"http_status": 410,
"detail": "Enrollment invalidated; start again",
}
else:
await conn.execute(
"UPDATE mfa_pending_enrollments SET attempts = $1 WHERE user_id = $2",
new_attempts,
current["id"],
)
failure = {
"reason": "invalid_code",
"attempts": new_attempts,
"http_status": 401,
"detail": "Invalid code",
}
else:
# Verified — finalize state inside the transaction.
await conn.execute(
"""
UPDATE users
SET mfa_enabled = TRUE,
mfa_method = 'totp',
mfa_secret_encrypted = $1,
mfa_enrolled_at = NOW(),
mfa_last_used_totp_step = NULL,
mfa_last_used_at = NULL
WHERE id = $2
""",
pending["secret_encrypted"],
current["id"],
)
await conn.execute(
"DELETE FROM mfa_backup_codes WHERE user_id = $1",
current["id"],
)
await conn.executemany(
"INSERT INTO mfa_backup_codes (user_id, code_hash) VALUES ($1, $2)",
[(current["id"], h) for h in hashes],
)
await conn.execute(
"DELETE FROM mfa_pending_enrollments WHERE user_id = $1",
current["id"],
)
finally:
await close_database_connection(conn)
if failure is not None:
# Log AFTER commit so the audit row reflects what actually persisted.
await log_user_activity(
user_id=current["id"],
action="mfa.enrollment.failed",
resource_type="mfa",
resource_id=str(current["id"]),
details={
"reason": failure["reason"],
"attempts": failure.get("attempts"),
},
ip_address=ip,
user_agent=ua,
)
raise HTTPException(status_code=failure["http_status"], detail=failure["detail"])
await log_user_activity(
user_id=current["id"],
action="mfa.enrollment.confirmed",
resource_type="mfa",
resource_id=str(current["id"]),
details={"method": "totp"},
ip_address=ip,
user_agent=ua,
)
return MfaEnrollConfirmResponse(enabled=True, backup_codes=plain_codes, method="totp")
# ---------------------------------------------------------------------------
# Disable + regenerate
# ---------------------------------------------------------------------------
async def _verify_user_code(conn, user_row: dict, code: str) -> Optional[str]:
"""Verify a TOTP-or-backup code against the user's stored secret.
Returns the method used ('totp' / 'backup') on success, None on failure.
On TOTP success the step counter is bumped *atomically* — the UPDATE
only succeeds if no other request consumed the same (or a newer) step
in between. On backup success the consumed row's used_at is set with
an atomic ``WHERE used_at IS NULL RETURNING id`` pattern.
"""
secret_plain = mfa_service.decrypt_secret(user_row["mfa_secret_encrypted"])
if secret_plain:
ok, step = mfa_service.verify_totp_with_replay_guard(
secret_plain, code, user_row["mfa_last_used_totp_step"]
)
if ok:
bumped = await conn.fetchval(
"""
UPDATE users
SET mfa_last_used_totp_step = $1, mfa_last_used_at = NOW()
WHERE id = $2
AND (mfa_last_used_totp_step IS NULL
OR mfa_last_used_totp_step < $1)
RETURNING id
""",
step,
user_row["id"],
)
if bumped:
return "totp"
rows = await conn.fetch(
"""
SELECT id, code_hash FROM mfa_backup_codes
WHERE user_id = $1 AND used_at IS NULL
""",
user_row["id"],
)
for row in rows:
if await mfa_service.check_backup_code(code, row["code_hash"]):
consumed = await conn.fetchval(
"""
UPDATE mfa_backup_codes
SET used_at = NOW()
WHERE id = $1 AND used_at IS NULL
RETURNING id
""",
row["id"],
)
if consumed:
return "backup"
return None
@router.post("/disable", summary="Disable MFA (requires current TOTP or backup)")
@limiter.limit(MFA_LIMITS.disable, key_func=mfa_rate_limit_key)
async def mfa_disable(
payload: MfaDisableRequest,
request: Request,
authorization: str = Header(None),
):
current = await _require_user(authorization)
conn = await get_database_connection()
try:
user_row = await conn.fetchrow(
"""
SELECT id, mfa_enabled, mfa_secret_encrypted, mfa_last_used_totp_step
FROM users WHERE id = $1
""",
current["id"],
)
if not user_row or not user_row["mfa_enabled"]:
raise HTTPException(status_code=400, detail="MFA is not enabled")
method_used = await _verify_user_code(conn, dict(user_row), payload.code)
if not method_used:
await log_user_activity(
user_id=current["id"],
action="mfa.disable.failed",
resource_type="mfa",
resource_id=str(current["id"]),
details={"reason": "invalid_code"},
ip_address=str(request.client.host) if request.client else None,
user_agent=request.headers.get("user-agent"),
)
raise HTTPException(status_code=401, detail="Invalid code")
async with conn.transaction():
await conn.execute(
"""
UPDATE users
SET mfa_enabled = FALSE,
mfa_method = NULL,
mfa_secret_encrypted = NULL,
mfa_enrolled_at = NULL,
mfa_last_used_at = NULL,
mfa_last_used_totp_step = NULL
WHERE id = $1
""",
current["id"],
)
await conn.execute(
"DELETE FROM mfa_backup_codes WHERE user_id = $1", current["id"]
)
await conn.execute(
"DELETE FROM mfa_pending_logins WHERE user_id = $1", current["id"]
)
await conn.execute(
"DELETE FROM mfa_pending_enrollments WHERE user_id = $1", current["id"]
)
finally:
await close_database_connection(conn)
await log_user_activity(
user_id=current["id"],
action="mfa.disabled.self",
resource_type="mfa",
resource_id=str(current["id"]),
details={"verified_via": method_used},
ip_address=str(request.client.host) if request.client else None,
user_agent=request.headers.get("user-agent"),
)
return {"enabled": False}
@router.post(
"/backup-codes/regenerate",
summary="Issue 10 fresh backup codes (TOTP required)",
response_model=MfaRegenerateBackupResponse,
)
@limiter.limit(MFA_LIMITS.regenerate_backup_codes, key_func=mfa_rate_limit_key)
async def mfa_regenerate_backup_codes(
payload: MfaRegenerateBackupRequest,
request: Request,
authorization: str = Header(None),
):
current = await _require_user(authorization)
# bcrypt-hash the new codes outside the DB transaction (~2-3s of CPU work).
plain_codes = mfa_service.generate_backup_codes()
hashes = await mfa_service.hash_backup_codes(plain_codes)
ip = str(request.client.host) if request.client else None
ua = request.headers.get("user-agent")
failure: Optional[dict] = None
conn = await get_database_connection()
try:
async with conn.transaction():
user_row = await conn.fetchrow(
"""
SELECT id, mfa_enabled, mfa_secret_encrypted, mfa_last_used_totp_step
FROM users WHERE id = $1
FOR UPDATE
""",
current["id"],
)
if not user_row or not user_row["mfa_enabled"]:
failure = {"http_status": 400, "detail": "MFA is not enabled"}
else:
secret_plain = mfa_service.decrypt_secret(user_row["mfa_secret_encrypted"])
if not secret_plain:
failure = {"http_status": 500, "detail": "MFA secret unreadable"}
else:
ok, step = mfa_service.verify_totp_with_replay_guard(
secret_plain, payload.code, user_row["mfa_last_used_totp_step"]
)
if not ok:
failure = {"http_status": 401, "detail": "Invalid TOTP code"}
else:
bumped = await conn.fetchval(
"""
UPDATE users
SET mfa_last_used_totp_step = $1, mfa_last_used_at = NOW()
WHERE id = $2
AND (mfa_last_used_totp_step IS NULL
OR mfa_last_used_totp_step < $1)
RETURNING id
""",
step,
current["id"],
)
if not bumped:
failure = {"http_status": 401, "detail": "Invalid TOTP code"}
else:
await conn.execute(
"DELETE FROM mfa_backup_codes WHERE user_id = $1",
current["id"],
)
await conn.executemany(
"INSERT INTO mfa_backup_codes (user_id, code_hash) VALUES ($1, $2)",
[(current["id"], h) for h in hashes],
)
finally:
await close_database_connection(conn)
if failure is not None:
raise HTTPException(status_code=failure["http_status"], detail=failure["detail"])
await log_user_activity(
user_id=current["id"],
action="mfa.backup_codes.regenerated",
resource_type="mfa",
resource_id=str(current["id"]),
details={"codes_count": len(plain_codes)},
ip_address=ip,
user_agent=ua,
)
return MfaRegenerateBackupResponse(backup_codes=plain_codes)
# ---------------------------------------------------------------------------
# Admin operations
# ---------------------------------------------------------------------------
@router.post("/admin-reset/{user_id}", summary="Admin: reset a single user's MFA")
@limiter.limit(MFA_LIMITS.admin_reset, key_func=mfa_rate_limit_key)
async def mfa_admin_reset(
user_id: int,
payload: MfaAdminResetRequest,
request: Request,
authorization: str = Header(None),
):
admin = await _require_admin(authorization)
if user_id == admin["id"]:
raise HTTPException(status_code=400, detail="Use /api/mfa/disable for self-reset")
conn = await get_database_connection()
try:
target = await conn.fetchrow(
"SELECT id, username, mfa_enabled FROM users WHERE id = $1", user_id
)
if not target:
raise HTTPException(status_code=404, detail="User not found")
async with conn.transaction():
await conn.execute(
"""
UPDATE users
SET mfa_enabled = FALSE,
mfa_method = NULL,
mfa_secret_encrypted = NULL,
mfa_enrolled_at = NULL,
mfa_last_used_at = NULL,
mfa_last_used_totp_step = NULL
WHERE id = $1
""",
user_id,
)
await conn.execute(
"DELETE FROM mfa_backup_codes WHERE user_id = $1", user_id
)
await conn.execute(
"DELETE FROM mfa_pending_logins WHERE user_id = $1", user_id
)
await conn.execute(
"DELETE FROM mfa_pending_enrollments WHERE user_id = $1", user_id
)
finally:
await close_database_connection(conn)
await log_user_activity(
user_id=admin["id"],
action="mfa.disabled.admin_reset",
resource_type="mfa",
resource_id=str(user_id),
details={
"target_user_id": user_id,
"target_username": target["username"],
"admin_user_id": admin["id"],
"reason": payload.reason,
},
ip_address=str(request.client.host) if request.client else None,
user_agent=request.headers.get("user-agent"),
)
return {"reset": True, "user_id": user_id}
@router.post(
"/admin-reset-all",
summary="Admin: emergency reset of MFA for all users (double confirm)",
)
@limiter.limit(MFA_LIMITS.admin_reset_all, key_func=mfa_rate_limit_key)
async def mfa_admin_reset_all(
payload: MfaAdminResetAllRequest,
request: Request,
authorization: str = Header(None),
):
admin = await _require_admin(authorization)
# Pydantic's Literal already enforces the magic string, but check defensively too.
if payload.confirm != "RESET ALL MFA":
raise HTTPException(status_code=400, detail="Invalid confirmation string")
conn = await get_database_connection()
try:
async with conn.transaction():
reset_count = await conn.fetchval(
"""
WITH affected AS (
UPDATE users
SET mfa_enabled = FALSE,
mfa_method = NULL,
mfa_secret_encrypted = NULL,
mfa_enrolled_at = NULL,
mfa_last_used_at = NULL,
mfa_last_used_totp_step = NULL
WHERE mfa_enabled = TRUE
RETURNING id
)
SELECT COUNT(*) FROM affected
"""
)
await conn.execute("DELETE FROM mfa_backup_codes")
await conn.execute("DELETE FROM mfa_pending_logins")
await conn.execute("DELETE FROM mfa_pending_enrollments")
finally:
await close_database_connection(conn)
await log_user_activity(
user_id=admin["id"],
action="mfa.disabled.admin_bulk_reset",
resource_type="mfa",
resource_id=str(admin["id"]),
details={
"reset_count": int(reset_count or 0),
"reason": payload.reason,
"admin_user_id": admin["id"],
},
ip_address=str(request.client.host) if request.client else None,
user_agent=request.headers.get("user-agent"),
)
return {"reset_count": int(reset_count or 0)}
+188 -13
View File
@@ -1046,9 +1046,45 @@ async def suggest_defaults(
slug = "newhost"
if domain:
base = domain.replace("*.", "").split(".")
slug = (base[0] or "newhost")[:32].lower()
slug = "".join(c if (c.isalnum() or c in ("-", "_")) else "-" for c in slug)
# Bulgu #93 (round-24 audit) — IDN / Unicode safety. PRE-FIX
# the suggest endpoint used `c.isalnum()`, which is Unicode-
# aware and returns True for non-ASCII letters (ü, é, ñ, …).
# An operator who typed `bücher.example.com` got back
# `backend_name="be-bücher"`, dropped that into the wizard
# form, and then hit a hard 422 at create time because the
# backend/frontend name validator regex
# `^[a-zA-Z][a-zA-Z0-9_-]{0,63}$` is ASCII-only. The wizard
# CREATE path also forces the domain itself through punycode
# (the validator rejects raw Unicode with a "use 'xn--…'"
# hint). Make `suggest` honour the same on-the-wire ASCII
# contract: convert each label to its IDN/punycode form
# FIRST, then sanitise to the alphanumeric / `-_` set the
# entity-name regex permits. The result is a name the
# operator can submit to /api/sites without re-typing.
first_label = (
domain.replace("*.", "").split(".")[0]
if domain.replace("*.", "")
else ""
)
ascii_label = first_label
if first_label and not first_label.isascii():
try:
ascii_label = first_label.encode("idna").decode("ascii")
except (UnicodeError, UnicodeDecodeError):
# IDN encoding failed (empty label, invalid chars,
# etc.) — fall back to stripping non-ASCII to '-'
# so we still produce a usable slug.
ascii_label = "".join(
c if c.isascii() and (c.isalnum() or c in ("-", "_"))
else "-"
for c in first_label
)
slug = (ascii_label or "newhost")[:32].lower()
slug = "".join(
c if c.isascii() and (c.isalnum() or c in ("-", "_"))
else "-"
for c in slug
)
if slug.startswith("_"):
slug = "h-" + slug.lstrip("_")
if not slug or not slug[0].isalpha():
@@ -1146,6 +1182,53 @@ async def preview_create(
try:
await _validate_user_cluster_access(current_user["id"], body.cluster_id, conn)
# Bulgu #88 / #89 (round-24 audit) — cluster-RBAC parity with
# `create_site` for SSL certificate references. PRE-FIX the
# preview path (`POST /api/sites/preview`) skipped the
# `select_existing_cert()` gate that `create_site` runs at
# lines 2298-2308 (per-server CA bundle) and 2398-2405
# (HTTPS bind cert). The omission let an authenticated wizard
# user pass `ssl.mode='existing', ssl_certificate_id=<X>` (or
# `servers[i].ssl_certificate_id=<X>`) where cert `X` belongs
# to a DIFFERENT cluster and receive the full rendered
# `would_create` envelope back — leaking cert id metadata
# across tenant boundaries. The actual submit (`POST /api/
# sites`) does enforce the gate, so this is a preview-only
# information leak, not a write-path escalation. The fix is
# to mirror the same `select_existing_cert(conn, id, cluster_
# id)` predicate (which already encodes the "global OR
# junction-bound" rule defined in `ssl_service.py:331-370`)
# so the preview returns 400 with a clear hint rather than
# 200 with a leaked render. We deliberately keep the error
# phrasing identical to the create-time message so wizard
# UI handlers that already match on "not found / inactive"
# need no client changes.
if body.ssl.mode == "existing" and body.ssl.ssl_certificate_id:
_preview_resolved = await select_existing_cert(
conn, body.ssl.ssl_certificate_id, body.cluster_id,
)
if not _preview_resolved:
raise HTTPException(
status_code=400,
detail=(
f"ssl_certificate_id {body.ssl.ssl_certificate_id} "
"not found / inactive / not bound to this cluster"
),
)
for _idx, _srv in enumerate(body.servers or []):
_srv_cert_id = getattr(_srv, "ssl_certificate_id", None)
if _srv_cert_id:
if not await select_existing_cert(
conn, _srv_cert_id, body.cluster_id,
):
raise HTTPException(
status_code=400,
detail=(
f"servers[{_idx}].ssl_certificate_id={_srv_cert_id} "
"not found / inactive / not bound to this cluster"
),
)
# Phase K Phase C: rate-limit ONLY the dry-run code path so
# legacy callers (e.g. `SiteDrafts.handlePreview` which never
# sets the flag) keep their unrestricted preview budget. The
@@ -1848,6 +1931,28 @@ async def create_site(
try:
await _validate_user_cluster_access(user_id, body.cluster_id, conn)
# Bulgu #87 (round-23 audit) — rate-limit the actual create
# endpoint. Pre-fix /api/sites/preview and
# /api/sites/preflight-acme were rate-limited (5/min, see
# `_enforce_rate_limit` callers above) but POST /api/sites
# itself had NO cap. Round-3 live testing fired 10 wizard
# creates in a single second against demo-cluster1; without
# the cap an authenticated user / script can spam-create
# entities until the cluster apply queue clogs.
#
# Action name must match what the router itself logs at the
# end of create_site (`wizard_create_site` — see
# `_log_user_activity(... action="wizard_create_site", ...)`
# at the success path below). The activity_logger middleware
# intentionally skips this endpoint to avoid double-logging
# (see middleware/activity_logger.py:74-110), so we must
# rate-limit on the SAME action the router emits.
#
# 5/min matches the preview / preflight budget — bulk
# operators have the dedicated /api/sites/bulk-import path
# for larger batches.
await _enforce_rate_limit(conn, user_id, "wizard_create_site")
# ----- Pre-create checks (must succeed before transaction)
# Phase 3 (R11-audit follow-up): reserved-name check parity with
@@ -2133,7 +2238,31 @@ async def create_site(
# rebrand. The reject path on `cluster.py` recognises BOTH
# prefixes so historical APPLIED versions (created before this
# rename) keep behaving correctly during reject/undo.
version_name = f"bulk-site-create-{ts}"
#
# Bulgu #86 (round-23 audit) — append a short UUID suffix so
# two wizard POST /api/sites calls submitted within the SAME
# epoch second cannot collide on the
# `config_versions(cluster_id, version_name)` UNIQUE
# constraint. Pre-fix `version_name = f"bulk-site-create-{ts}"`
# gave seconds resolution, so an operator who clicked "Create"
# twice in rapid succession (or any back-to-back API
# automation) saw the SECOND call 409 with the generic
# `UniqueViolationError` fall-through message:
#
# "A wizard entity with this name already exists on the
# cluster (UNIQUE constraint). Pick a different name."
#
# — even though the operator-chosen backend / frontend / SSL
# names were unique. The actual collision was on the
# auto-generated `version_name` and re-naming the wizard
# inputs did NOT help. The reject_pending_changes path on
# cluster.py:4179 prefix-matches `bulk-site-create-*` so
# appending a unique suffix preserves the historical
# rollback / undo semantics. 6 hex chars give 16M-room before
# birthday collisions, vastly more than the per-second
# request volume an operator can sustain through the wizard.
import uuid
version_name = f"bulk-site-create-{ts}-{uuid.uuid4().hex[:6]}"
bulk_snapshots: List[dict] = []
created_ids: Dict[str, Any] = {}
@@ -2929,22 +3058,68 @@ async def create_site(
# broke".
msg = str(uve)
logger.info(f"WIZARD: name conflict on create: {msg}")
if "backends_name_cluster_id_key" in msg or 'backends_name' in msg:
# Bulgu #85 (round-23 audit) — extract the offending constraint
# name from the asyncpg message so the operator-visible detail
# can pin-point WHICH entity collided. Pre-fix the handler only
# recognised `backends_*_key` and `frontends_*_key`; any other
# constraint (ssl_certificates, backend_servers, config_versions
# …) fell through to the generic "wizard entity with this name"
# message which is useless for debugging — the operator has to
# open the server log and the engineer has to ssh-bounce to
# extract `constraint=<name>` from the exception detail.
#
# ``asyncpg.UniqueViolationError.constraint_name`` is the
# canonical structured field; ``str(uve)`` only contains the
# human-formatted DETAIL line. Prefer the attribute, fall back
# to substring scanning so we stay robust if asyncpg ever stops
# exposing it.
constraint = getattr(uve, "constraint_name", None) or ""
msg_lower = msg.lower()
if "backends_name_cluster_id_key" in msg or 'backends_name' in msg \
or constraint == "backends_name_cluster_id_key":
detail = (
"A backend with this name already exists on the cluster "
"(possibly soft-deleted). Pick a different backend name "
"or restore/permanently-delete the existing row."
f"A backend named '{body.backend.name}' already exists on "
"the cluster (possibly soft-deleted). Pick a different "
"backend name or restore/permanently-delete the existing "
"row."
)
elif "frontends_name_cluster_id_key" in msg or "frontends_name" in msg:
elif "frontends_name_cluster_id_key" in msg or "frontends_name" in msg \
or constraint == "frontends_name_cluster_id_key":
detail = (
"A frontend with this name already exists on the cluster "
f"A frontend named '{body.frontend.name}' (or its auto-"
f"derived HTTPS sibling) already exists on the cluster "
"(possibly soft-deleted). Pick a different frontend name "
"or restore/permanently-delete the existing row."
)
else:
elif "ssl_certificates" in msg_lower or "ssl_certificates" in constraint \
or "ssl_cert" in constraint:
ssl_name = getattr(body.ssl, "name", None) or "(unnamed)"
detail = (
"A wizard entity with this name already exists on the "
"cluster (UNIQUE constraint). Pick a different name."
f"An SSL certificate named '{ssl_name}' already exists on "
"the cluster (possibly soft-deleted). Pick a different "
"ssl.name or restore/permanently-delete the existing "
"certificate row."
)
elif "backend_servers" in msg_lower or "backend_servers" in constraint:
detail = (
"Two servers in the same backend share a server_name. "
"HAProxy requires `server <name>` tokens to be unique "
"within a backend block. Rename the duplicate(s) and "
"resubmit."
)
else:
# Echo the constraint name (a stable, non-secret schema
# identifier) in the detail so a human reading the toast
# can grep the codebase for the matching CREATE TABLE
# without needing server-log access. Names like
# `proxied_hosts_pkey` or `config_versions_unique` are
# safe to surface — they're public schema info.
con_hint = f" (constraint={constraint})" if constraint else ""
detail = (
f"A wizard entity with this name already exists on the "
f"cluster (UNIQUE constraint{con_hint}). Pick a different "
"name or retry; if the conflict persists contact the "
"platform team."
)
raise HTTPException(status_code=409, detail=detail)
except Exception as e:
+5 -4
View File
@@ -38,18 +38,19 @@ async def get_users(authorization: str = Header(None)):
# Get users with their roles (only active users)
try:
users = await conn.fetch("""
SELECT u.id, u.username, u.email, u.full_name, u.phone, u.role, u.is_active,
u.is_admin, u.is_verified, u.created_at, u.updated_at, u.last_login_at
SELECT u.id, u.username, u.email, u.full_name, u.phone, u.role, u.is_active,
u.is_admin, u.is_verified, u.created_at, u.updated_at, u.last_login_at,
COALESCE(u.mfa_enabled, FALSE) AS mfa_enabled
FROM users u
WHERE u.is_active = TRUE
ORDER BY u.username
""")
except Exception as schema_error:
logger.warning(f"Schema error in users query, using fallback: {schema_error}")
# Fallback query with minimal columns
# Fallback query with minimal columns (pre-MFA-migration deploys)
users = await conn.fetch("""
SELECT id, username, email, is_active, is_admin, created_at
FROM users
FROM users
WHERE is_active = TRUE
ORDER BY username
""")
+83 -6
View File
@@ -573,9 +573,18 @@ async def check_agents(conn, cluster_ids: List[int]) -> Dict[str, Any]:
severity="warn",
duration_ms=int((time.time() - started) * 1000),
)
# Bulgu #84 (round-23 audit) — the canonical timestamp column on the
# `agents` table is `last_seen`. Pre-fix this query referenced a
# non-existent `a.last_heartbeat`, so every ACME preflight call
# (`POST /api/sites/preflight-acme`) crashed at the `check_agents`
# stage with `UndefinedColumnError: column a.last_heartbeat does
# not exist`, blocking the entire wizard's ACME pre-validation
# gate. Every other agents.last_seen reader in the codebase
# (routers/cluster.py:695-702, routers/agent.py, routers/dashboard
# *.py) uses `last_seen`; aligning here.
rows = await conn.fetch(
"""
SELECT a.id, a.hostname, a.status, a.last_heartbeat, hc.id AS cluster_id, hc.name AS cluster_name
SELECT a.id, a.hostname, a.status, a.last_seen, hc.id AS cluster_id, hc.name AS cluster_name
FROM agents a
JOIN haproxy_clusters hc ON hc.pool_id = a.pool_id
WHERE hc.id = ANY($1::int[])
@@ -623,6 +632,61 @@ async def check_agents(conn, cluster_ids: List[int]) -> Dict[str, Any]:
CHECK_IDS = ("dns", "port80", "routing", "account", "agents")
def _coerce_cluster_ids(raw) -> List[int]:
"""Coerce a cluster_ids list to ints, dropping non-integer-compatible
values. JSONB-stored lists occasionally land as ["1", "2"] (string form)
due to legacy paths; asyncpg's `$1::int[]` cast then fails the diagnostic
query with InvalidTextRepresentationError. We normalise here so the
diagnostic surface is the same regardless of how the order was written.
"""
out: List[int] = []
if not raw:
return out
for v in raw:
try:
out.append(int(v))
except (TypeError, ValueError):
continue
return out
# Bulgu #94 (Round-25 audit) — the entire point of the diagnostic panel
# is to SHOW the operator what went wrong. Pre-fix, a single check raising
# an uncaught exception (e.g. an asyncpg cast error from a malformed
# cluster_ids JSONB, a DNS resolver outage, an SSRF-guard glitch) would
# propagate up to the router's `try/finally` block, which had no `except`
# clause, and return HTTP 500 with no body. The operator saw only
# "Internal Server Error" in DevTools — the inverse of what a diagnostic
# panel should ever produce. We now wrap every check inside `run_checks`
# so that a check crash becomes a structured `fail` row instead of
# bubbling up; the operator gets the exception type + message in the
# UI and can carry it forward, and the rest of the panel still renders.
async def _safe_check(check_id: str, label: str, coro):
"""Run an awaitable that produces a check result; swallow exceptions
and convert them to a structured `fail` result so the diagnostic
response is never short-circuited by a single broken check."""
started = time.time()
try:
return await coro
except Exception as exc: # noqa: BLE001 — diagnostic boundary
duration_ms = int((time.time() - started) * 1000)
logger.exception(
"ACME diagnostic check %s raised", check_id
)
return _check_result(
check_id,
label,
"fail",
f"Diagnostic check crashed: {exc.__class__.__name__}: {exc}",
severity="error",
details={
"exception_type": exc.__class__.__name__,
"exception_message": str(exc),
},
duration_ms=duration_ms,
)
async def run_checks(
conn,
*,
@@ -633,19 +697,32 @@ async def run_checks(
) -> List[Dict[str, Any]]:
"""Execute the full pre-flight check suite. `only` lets callers re-run a
subset (per-check rerun in the UI).
Every individual check is wrapped in `_safe_check` so the diagnostic
endpoint NEVER 500s because of one broken check — the operator gets
a structured `fail` row identifying which check crashed and why.
"""
selected = set(only) if only else set(CHECK_IDS)
results: List[Dict[str, Any]] = []
# Normalise inputs once so the per-check error stays in the right
# bucket (a malformed cluster_ids should not crash routing/agents).
safe_domains = [d for d in (domains or []) if isinstance(d, str) and d]
safe_cluster_ids = _coerce_cluster_ids(cluster_ids)
try:
safe_account_id = int(account_id) if account_id is not None else None
except (TypeError, ValueError):
safe_account_id = None
if "dns" in selected:
results.append(await check_dns(domains))
results.append(await _safe_check("dns", "DNS resolution", check_dns(safe_domains)))
if "port80" in selected:
results.append(await check_port80(domains))
results.append(await _safe_check("port80", "Port 80 reachability", check_port80(safe_domains)))
if "routing" in selected:
results.append(await check_routing(conn, domains, cluster_ids))
results.append(await _safe_check("routing", "HAProxy routing", check_routing(conn, safe_domains, safe_cluster_ids)))
if "account" in selected:
results.append(await check_account(conn, account_id))
results.append(await _safe_check("account", "ACME account", check_account(conn, safe_account_id)))
if "agents" in selected:
results.append(await check_agents(conn, cluster_ids))
results.append(await _safe_check("agents", "HAProxy agents", check_agents(conn, safe_cluster_ids)))
return results
+238
View File
@@ -0,0 +1,238 @@
"""MFA (TOTP + backup codes) service layer — Issue #18, v1.6.0.
Owns the cryptographic and persistence-shape concerns of multi-factor auth:
- TOTP secret generation / verification with replay protection (RFC 6238)
- Backup code generation, hashing (bcrypt) and atomic single-use consumption
- Fernet-based encryption of TOTP secrets at rest
Strictly no logging of secrets — only metadata (lengths, counts) is logged.
"""
from __future__ import annotations
import asyncio
import base64
import logging
import os
import secrets as _secrets
import time
from typing import List, Optional, Tuple
from urllib.parse import quote
import bcrypt
import pyotp
from cryptography.fernet import Fernet, InvalidToken
from cryptography.hazmat.primitives import hashes
from cryptography.hazmat.primitives.kdf.hkdf import HKDF
from config import SECRET_KEY
logger = logging.getLogger(__name__)
# RFC 6238 parameters — kept conservative for widest authenticator app compatibility.
TOTP_DIGITS = 6
TOTP_PERIOD = 30
TOTP_DIGEST = "sha1"
TOTP_VALID_WINDOW_STEPS = 1 # ±1 step (±30s) tolerance
# Backup code spec (Plan section 5).
# Alphabet drops the confusing pairs: 0/O, 1/I, L. Resulting size is 31, which
# still yields 31**8 ≈ 8.5×10^11 combinations per half — far beyond brute-force.
BACKUP_CODE_COUNT = 10
BACKUP_CODE_ALPHABET = "ABCDEFGHJKMNPQRSTUVWXYZ23456789"
BACKUP_CODE_HALF_LEN = 4 # XXXX-YYYY
# OTP URI defaults.
DEFAULT_ISSUER = "HAProxy OpenManager"
ACCOUNT_LABEL_DOMAIN_FALLBACK = "haproxy-openmanager"
# ---------------------------------------------------------------------------
# Fernet key resolution
# ---------------------------------------------------------------------------
_fernet_instance: Optional[Fernet] = None
def _resolve_fernet_key() -> bytes:
"""Resolve the Fernet key, preferring the explicit env var.
Falls back to HKDF over SECRET_KEY with a versioned info string so a future
rotation can be expressed by bumping the version suffix.
"""
explicit = os.getenv("MFA_ENCRYPTION_KEY", "").strip()
if explicit:
try:
Fernet(explicit.encode())
return explicit.encode()
except Exception as exc:
logger.error("MFA_ENCRYPTION_KEY env var present but invalid: %s", exc)
# fall through to HKDF derivation rather than crashing the app
logger.warning(
"MFA_ENCRYPTION_KEY env var not set or invalid; deriving from SECRET_KEY (v1). "
"Set an explicit MFA_ENCRYPTION_KEY in production to enable key rotation."
)
hkdf = HKDF(
algorithm=hashes.SHA256(),
length=32,
salt=None,
info=b"mfa-totp-secret-v1",
)
derived = hkdf.derive(SECRET_KEY.encode("utf-8"))
return base64.urlsafe_b64encode(derived)
def _get_fernet() -> Fernet:
global _fernet_instance
if _fernet_instance is None:
_fernet_instance = Fernet(_resolve_fernet_key())
return _fernet_instance
def reset_fernet_for_tests() -> None:
"""Test-only hook to force re-resolution of the Fernet key after env mutation."""
global _fernet_instance
_fernet_instance = None
# ---------------------------------------------------------------------------
# TOTP secrets
# ---------------------------------------------------------------------------
def generate_totp_secret() -> str:
"""Return a fresh base32 TOTP secret (32 chars)."""
return pyotp.random_base32()
def encrypt_secret(secret_plain: str) -> str:
"""Fernet-encrypt the base32 secret. Returns str for direct DB storage."""
token = _get_fernet().encrypt(secret_plain.encode("utf-8"))
return token.decode("utf-8")
def decrypt_secret(secret_encrypted: str) -> Optional[str]:
"""Decrypt a previously stored secret. Returns None when the token can't be
decrypted (e.g. key rotated without re-enroll). Never raises to the caller.
"""
try:
return _get_fernet().decrypt(secret_encrypted.encode("utf-8")).decode("utf-8")
except InvalidToken:
logger.warning("Failed to decrypt MFA secret (invalid Fernet token)")
return None
except Exception as exc:
logger.error("Unexpected error decrypting MFA secret: %s", exc)
return None
def build_otpauth_uri(account_label: str, secret_plain: str, issuer: str = DEFAULT_ISSUER) -> str:
"""Build an otpauth:// URI that all major authenticator apps accept.
Format: otpauth://totp/<issuer>:<account>?secret=<b32>&issuer=<issuer>&algorithm=SHA1&digits=6&period=30
"""
issuer_q = quote(issuer, safe="")
label = f"{issuer}:{account_label}"
label_q = quote(label, safe=":@")
return (
f"otpauth://totp/{label_q}?secret={secret_plain}"
f"&issuer={issuer_q}&algorithm=SHA1&digits={TOTP_DIGITS}&period={TOTP_PERIOD}"
)
def build_account_label(username: str, hostname_hint: Optional[str] = None) -> str:
"""Compose the per-user otpauth label, respecting env > hostname > fallback."""
domain = (
os.getenv("MFA_ACCOUNT_LABEL_DOMAIN", "").strip()
or (hostname_hint or "").strip()
or ACCOUNT_LABEL_DOMAIN_FALLBACK
)
return f"{username}@{domain}"
def verify_totp_with_replay_guard(
secret_plain: str,
code: str,
last_used_step: Optional[int],
) -> Tuple[bool, Optional[int]]:
"""Verify a 6-digit TOTP code with explicit per-step replay protection.
Returns (success, step_consumed). Caller persists the consumed step on success.
Implementation notes:
- pyotp.TOTP.at(seconds_since_epoch) — to target step N we pass step*PERIOD.
- secrets.compare_digest is used for constant-time comparison.
- Replay guard rejects codes whose step is <= the previously consumed step.
"""
if not secret_plain or not code:
return (False, None)
code = code.strip()
if len(code) != TOTP_DIGITS or not code.isdigit():
return (False, None)
totp = pyotp.TOTP(secret_plain, digits=TOTP_DIGITS, interval=TOTP_PERIOD, digest=TOTP_DIGEST)
now = int(time.time())
current_step = now // TOTP_PERIOD
for offset in (0, -1, 1):
step = current_step + offset
expected = totp.at(step * TOTP_PERIOD)
if len(expected) == len(code) and _secrets.compare_digest(expected, code):
if last_used_step is not None and step <= last_used_step:
return (False, None)
return (True, step)
return (False, None)
# ---------------------------------------------------------------------------
# Backup codes
# ---------------------------------------------------------------------------
def generate_backup_codes(count: int = BACKUP_CODE_COUNT) -> List[str]:
"""Return ``count`` plain-text backup codes formatted as ``XXXX-YYYY``."""
codes: List[str] = []
for _ in range(count):
left = "".join(_secrets.choice(BACKUP_CODE_ALPHABET) for _ in range(BACKUP_CODE_HALF_LEN))
right = "".join(_secrets.choice(BACKUP_CODE_ALPHABET) for _ in range(BACKUP_CODE_HALF_LEN))
codes.append(f"{left}-{right}")
return codes
def normalize_backup_code(user_input: str) -> str:
"""Canonical form for comparison: uppercase, strip dashes/spaces."""
if not user_input:
return ""
return user_input.strip().upper().replace("-", "").replace(" ", "")
async def _hash_one_backup_code(code_plain: str) -> str:
"""Bcrypt-hash a single backup code on a worker thread."""
normalized = normalize_backup_code(code_plain)
hashed = await asyncio.to_thread(bcrypt.hashpw, normalized.encode("utf-8"), bcrypt.gensalt())
return hashed.decode("utf-8")
async def hash_backup_codes(codes_plain: List[str]) -> List[str]:
"""Hash backup codes in parallel (each bcrypt op runs in its own thread)."""
return await asyncio.gather(*(_hash_one_backup_code(c) for c in codes_plain))
async def check_backup_code(user_input: str, code_hash: str) -> bool:
"""Run a single bcrypt verify on the worker pool."""
normalized = normalize_backup_code(user_input)
if not normalized:
return False
return await asyncio.to_thread(
bcrypt.checkpw, normalized.encode("utf-8"), code_hash.encode("utf-8")
)
# ---------------------------------------------------------------------------
# Misc helpers
# ---------------------------------------------------------------------------
def generate_challenge_token() -> str:
"""64-char hex challenge token for /api/auth/login → /mfa-verify hand-off."""
return _secrets.token_hex(32)
+223 -3
View File
@@ -433,7 +433,7 @@ async def test_check_agents_fail_when_none_registered():
async def test_check_agents_warn_when_none_active():
conn = AsyncMock()
conn.fetch.return_value = [
{"id": 1, "hostname": "h1", "status": "offline", "last_heartbeat": None,
{"id": 1, "hostname": "h1", "status": "offline", "last_seen": None,
"cluster_id": 1, "cluster_name": "c1"},
]
out = await check_agents(conn, [1])
@@ -444,9 +444,9 @@ async def test_check_agents_warn_when_none_active():
async def test_check_agents_ok_with_active():
conn = AsyncMock()
conn.fetch.return_value = [
{"id": 1, "hostname": "h1", "status": "active", "last_heartbeat": None,
{"id": 1, "hostname": "h1", "status": "active", "last_seen": None,
"cluster_id": 1, "cluster_name": "c1"},
{"id": 2, "hostname": "h2", "status": "offline", "last_heartbeat": None,
{"id": 2, "hostname": "h2", "status": "offline", "last_seen": None,
"cluster_id": 1, "cluster_name": "c1"},
]
out = await check_agents(conn, [1])
@@ -454,6 +454,49 @@ async def test_check_agents_ok_with_active():
assert "1 of 2" in out["message"]
@pytest.mark.asyncio
async def test_bulgu84_check_agents_uses_last_seen_column():
"""Bulgu #84 (round-23 audit) — pin the SQL column name.
Pre-fix the query referenced a non-existent ``a.last_heartbeat``
column. AsyncMock returns whatever dict the test sets without
re-validating the SQL string, so the pre-fix test suite was
GREEN while every live ACME preflight call (the
``POST /api/sites/preflight-acme`` endpoint that the wizard
runs before showing the ACME step) crashed with
``UndefinedColumnError: column a.last_heartbeat does not
exist``. The crash blocked the entire wizard ACME preview
page on a production cluster, but unit tests never noticed
because they only assert on the helper's return shape, not
on the SQL string the helper sends to Postgres.
This pin inspects ``conn.fetch.call_args`` to assert the SQL
body actually queries ``a.last_seen`` (the canonical column
name used everywhere else in the codebase — see
routers/cluster.py:695-702 and routers/agent.py:281+ for
sibling readers). A future ``last_heartbeat`` typo would
re-fail this test without anyone noticing the live impact.
"""
conn = AsyncMock()
conn.fetch.return_value = [
{"id": 1, "hostname": "h1", "status": "active", "last_seen": None,
"cluster_id": 1, "cluster_name": "c1"},
]
await check_agents(conn, [1])
conn.fetch.assert_awaited_once()
sql_query = conn.fetch.call_args[0][0]
assert "a.last_seen" in sql_query, (
f"check_agents SQL must select a.last_seen (the canonical "
f"agents-table timestamp column); got: {sql_query!r}"
)
assert "a.last_heartbeat" not in sql_query, (
f"check_agents SQL still references a.last_heartbeat — this "
f"column does NOT exist on the agents table and the query "
f"will 500 with UndefinedColumnError at runtime. SQL: "
f"{sql_query!r}"
)
# ----------------------------------------------------------------------------
# run_checks orchestration
# ----------------------------------------------------------------------------
@@ -506,3 +549,180 @@ async def test_run_checks_unknown_only_returns_empty():
account_id=None, only=["bogus"],
)
assert out == []
# ----------------------------------------------------------------------------
# Bulgu #94 / #95 (Round-25 audit) — diagnostic-runner robustness
# ----------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_bulgu94_run_checks_swallows_single_check_crash(monkeypatch):
"""Bulgu #94 — a single check exception must NOT collapse the suite.
Pre-fix, an asyncpg UndefinedColumnError from check_agents (e.g. the
Bulgu #84 ``a.last_heartbeat`` typo on an older deploy) propagated
up to the FastAPI router which had no `except`, so the operator saw
HTTP 500 with no body. The diagnostic panel is precisely the place
that should SURFACE this — never opaque-500 it. We now wrap each
check; the failing one becomes a structured `fail` row and the
other four still render.
"""
monkeypatch.setattr(socket, "gethostbyname_ex",
lambda d: (d, [], ["10.0.0.1"]))
def _ctor(*args, **kwargs):
return _FakeSession(statuses=[200])
monkeypatch.setattr("aiohttp.ClientSession", _ctor)
conn = AsyncMock()
# check_routing + check_agents both call conn.fetch; explode on the
# FIRST call (which is check_routing) and return rows on the second.
call_count = {"n": 0}
async def _fetch(*args, **kwargs):
call_count["n"] += 1
if call_count["n"] == 1:
raise RuntimeError("simulated: column a.last_heartbeat does not exist")
return []
conn.fetch = _fetch
conn.fetchrow.return_value = None
out = await run_checks(
conn,
domains=["a.example.com"],
cluster_ids=[1],
account_id=None,
)
# All five checks must still appear in the response.
ids = [c["id"] for c in out]
assert ids == ["dns", "port80", "routing", "account", "agents"]
routing = next(c for c in out if c["id"] == "routing")
assert routing["status"] == "fail"
assert "Diagnostic check crashed" in routing["message"]
assert routing["details"]["exception_type"] == "RuntimeError"
assert "last_heartbeat" in routing["details"]["exception_message"]
@pytest.mark.asyncio
async def test_bulgu94_run_checks_coerces_string_cluster_ids(monkeypatch):
"""Bulgu #94 — cluster_ids stored as JSONB strings (legacy paths)
must not crash check_routing / check_agents with
``invalid input syntax for type integer: "1"``."""
monkeypatch.setattr(socket, "gethostbyname_ex",
lambda d: (d, [], ["10.0.0.1"]))
def _ctor(*args, **kwargs):
return _FakeSession(statuses=[200])
monkeypatch.setattr("aiohttp.ClientSession", _ctor)
captured_args = []
async def _fetch(*args, **kwargs):
captured_args.append(args)
return []
conn = AsyncMock()
conn.fetch = _fetch
conn.fetchrow.return_value = None
out = await run_checks(
conn,
domains=["a.example.com"],
cluster_ids=["1", "2", "garbage", None, 3],
account_id=None,
)
# The list passed to asyncpg should already be a pure-int list.
# check_routing is the first call that uses cluster_ids.
routing_call_args = [a for a in captured_args if "frontends" in a[0]]
assert routing_call_args, "check_routing should have queried frontends"
cluster_ids_arg = routing_call_args[0][1]
assert cluster_ids_arg == [1, 2, 3], (
f"cluster_ids must be coerced to ints before being passed to "
f"asyncpg's ::int[] cast; got: {cluster_ids_arg!r}"
)
assert all(c["status"] != "fail" or c["id"] != "routing"
for c in out
if c["id"] == "routing" and "Diagnostic check crashed" in (c.get("message") or "")
), "routing should not have crashed on coerced cluster_ids"
@pytest.mark.asyncio
async def test_bulgu94_safe_check_does_not_swallow_cancellation(monkeypatch):
"""`_safe_check` must catch `Exception` but NOT `BaseException`.
asyncio.CancelledError is a BaseException (Python 3.8+) so it must
propagate out of `_safe_check` — otherwise a request that the
client cancelled mid-flight would silently keep running diagnostic
checks instead of unwinding cleanly. We hand `_safe_check` a coro
that raises CancelledError and assert it bubbles up.
"""
from services.acme_diagnostics import _safe_check
async def _cancelled_coro():
raise asyncio.CancelledError()
with pytest.raises(asyncio.CancelledError):
await _safe_check("dns", "DNS resolution", _cancelled_coro())
@pytest.mark.asyncio
async def test_bulgu94_safe_check_handles_keyboardinterrupt(monkeypatch):
"""`_safe_check` must also not swallow `KeyboardInterrupt`."""
from services.acme_diagnostics import _safe_check
async def _interrupt_coro():
raise KeyboardInterrupt()
with pytest.raises(KeyboardInterrupt):
await _safe_check("dns", "DNS resolution", _interrupt_coro())
@pytest.mark.asyncio
async def test_bulgu94_coerce_cluster_ids_handles_none():
"""Coercion must handle None input without raising."""
from services.acme_diagnostics import _coerce_cluster_ids
assert _coerce_cluster_ids(None) == []
assert _coerce_cluster_ids([]) == []
assert _coerce_cluster_ids([1, 2, 3]) == [1, 2, 3]
assert _coerce_cluster_ids(["1", "2"]) == [1, 2]
assert _coerce_cluster_ids([1.5]) == [1] # int() truncates floats
assert _coerce_cluster_ids(["abc", None, "5"]) == [5]
@pytest.mark.asyncio
async def test_bulgu94_run_checks_filters_invalid_domains(monkeypatch):
"""Non-string entries in `domains` must not reach the DNS resolver."""
seen_domains = []
def fake_gethostbyname_ex(domain):
seen_domains.append(domain)
return (domain, [], ["10.0.0.1"])
monkeypatch.setattr(socket, "gethostbyname_ex", fake_gethostbyname_ex)
def _ctor(*args, **kwargs):
return _FakeSession(statuses=[200])
monkeypatch.setattr("aiohttp.ClientSession", _ctor)
conn = AsyncMock()
conn.fetch.return_value = []
conn.fetchrow.return_value = None
out = await run_checks(
conn,
domains=["a.example.com", None, "", 42, "b.example.com"],
cluster_ids=[1],
account_id=None,
)
# Both check_dns and check_port80 resolve DNS, so each valid domain
# may appear multiple times — but invalid entries (None, "", 42)
# must never reach the resolver.
assert set(seen_domains) == {"a.example.com", "b.example.com"}
assert None not in seen_domains
assert "" not in seen_domains
assert 42 not in seen_domains
dns_check = next(c for c in out if c["id"] == "dns")
assert dns_check["status"] == "ok"
@@ -0,0 +1,554 @@
"""Router-level tests for the ACME diagnostic endpoints (Round-25 audit).
These tests pin the contract introduced by Bulgu #94 / #95:
* ``POST /api/letsencrypt/orders/{order_id}/diagnostics`` must NEVER return
HTTP 500 for an in-suite failure. Authentication / authorisation /
rate-limit / not-found errors still raise the appropriate 4xx, but any
unexpected exception during check execution is converted to HTTP 200
with a structured failure envelope so the UI can render the cause.
* ``GET /api/letsencrypt/orders/{order_id}/events`` must NEVER return
HTTP 500 because of schema drift in ``user_activity_logs`` (the
original 500 cause: SELECTing a non-existent ``status`` column). A
partial failure is reported via ``meta.errors[]``.
The tests use AsyncMock-based fake connections rather than spinning up a
real Postgres so they run hermetically inside CI.
"""
from unittest.mock import AsyncMock, patch
import pytest
from routers import acme_diagnostics as router_mod
# ----------------------------------------------------------------------------
# Helpers
# ----------------------------------------------------------------------------
class _FakeUser(dict):
pass
def _patch_auth_and_db(monkeypatch, conn, user_id=1):
"""Patch the auth / db helpers used by both endpoints so the tests
don't have to construct a real FastAPI request stack."""
async def _fake_user(_auth):
return _FakeUser(id=user_id, username="t", email="t@x")
async def _fake_perm(_uid, *_a, **_k):
return True
async def _fake_get_conn():
return conn
async def _fake_close_conn(_c):
return None
async def _fake_rate_limit(*_a, **_kw):
return None
monkeypatch.setattr(router_mod, "get_current_user_from_token", _fake_user)
monkeypatch.setattr(router_mod, "check_user_permission", _fake_perm)
monkeypatch.setattr(router_mod, "get_database_connection", _fake_get_conn)
monkeypatch.setattr(router_mod, "close_database_connection", _fake_close_conn)
monkeypatch.setattr(router_mod, "_enforce_rate_limit", _fake_rate_limit)
# ----------------------------------------------------------------------------
# /diagnostics endpoint
# ----------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_bulgu94_diagnostics_load_order_crash_returns_envelope_not_500(monkeypatch):
"""A DB crash during load_order must surface a 200 envelope, not 500.
Pre-fix, any RuntimeError between auth and run_checks bubbled out of
the bare ``try/finally`` block and FastAPI returned an opaque HTTP 500.
The Round-25 fix wraps load_order so the operator sees the cause
inside the diagnostic panel.
"""
conn = AsyncMock()
conn.fetchrow.side_effect = RuntimeError("simulated DB connectivity loss")
_patch_auth_and_db(monkeypatch, conn)
out = await router_mod.run_diagnostics(order_id=42, authorization="Bearer x")
assert out["order_id"] == 42
assert out["status"] == "diagnostics_unavailable"
assert out["checks"][0]["status"] == "fail"
assert out["checks"][0]["id"] == "diagnostics_runner"
assert "simulated DB connectivity loss" in out["checks"][0]["message"]
assert out["meta"]["correlation_id"]
assert out["meta"]["error_stage"] == "load_order"
assert "humanized_error" in out
@pytest.mark.asyncio
async def test_bulgu94_diagnostics_run_checks_crash_returns_envelope(monkeypatch):
"""A crash inside run_checks (after order is loaded) must also stay 200."""
conn = AsyncMock()
conn.fetchrow.return_value = {
"id": 5,
"account_id": 1,
"status": "invalid",
"domains": '["a.example.com"]',
"cluster_ids": "[1]",
"error_detail": None,
"post_completion_actions": None,
"pending_apply_version_name": None,
"wizard_staged_until": None,
"created_by": 1,
}
_patch_auth_and_db(monkeypatch, conn)
async def _boom(*_a, **_kw):
raise RuntimeError("simulated check orchestrator crash")
monkeypatch.setattr(router_mod, "run_checks", _boom)
out = await router_mod.run_diagnostics(order_id=5, authorization="Bearer x")
assert out["order_id"] == 5
assert out["status"] == "diagnostics_unavailable"
assert out["meta"]["error_stage"] == "run_checks"
assert out["meta"]["error_type"] == "RuntimeError"
assert "simulated check orchestrator crash" in out["meta"]["error_message"]
@pytest.mark.asyncio
async def test_bulgu94_diagnostics_returns_meta_summary_on_success(monkeypatch):
"""Successful diagnostics responses carry a meta summary the UI uses
to surface 'N checks failed, M warnings' without recomputing."""
conn = AsyncMock()
conn.fetchrow.return_value = {
"id": 5,
"account_id": 1,
"status": "invalid",
"domains": '["a.example.com"]',
"cluster_ids": "[1]",
"error_detail": None,
"post_completion_actions": None,
"pending_apply_version_name": None,
"wizard_staged_until": None,
"created_by": 1,
}
_patch_auth_and_db(monkeypatch, conn)
async def _fake_checks(*_a, **_kw):
return [
{"id": "dns", "label": "DNS", "status": "ok", "severity": "info", "message": "", "details": {}, "duration_ms": 1},
{"id": "routing", "label": "Routing", "status": "fail", "severity": "error", "message": "", "details": {}, "duration_ms": 1},
{"id": "port80", "label": "Port 80", "status": "warn", "severity": "warn", "message": "", "details": {}, "duration_ms": 1},
]
monkeypatch.setattr(router_mod, "run_checks", _fake_checks)
out = await router_mod.run_diagnostics(order_id=5, authorization="Bearer x")
assert out["meta"]["checks_total"] == 3
assert out["meta"]["checks_failed"] == 1
assert out["meta"]["checks_warn"] == 1
assert out["meta"]["correlation_id"]
# ----------------------------------------------------------------------------
# /events endpoint
# ----------------------------------------------------------------------------
@pytest.mark.asyncio
async def test_bulgu95_events_missing_status_column_returns_partial_envelope(monkeypatch):
"""Bulgu #95 — user_activity_logs lacks a `status` column.
Pre-fix, the SELECT pulled `status` directly and the endpoint
returned HTTP 500 for every order that had any correlated
user-activity rows. The Round-25 fix introspects the schema; here
we simulate a deployment with no `status` column AND a JOIN /
query that would otherwise crash — the endpoint must stay 200,
return whatever it could collect from acme_order_events, and
record the user_activity_logs section as degraded but recoverable.
"""
conn = AsyncMock()
# _load_order
order_row = {
"id": 5,
"account_id": 1,
"status": "invalid",
"domains": '["a.example.com"]',
"cluster_ids": "[1]",
"error_detail": None,
"post_completion_actions": None,
"pending_apply_version_name": None,
"wizard_staged_until": None,
"created_by": 1,
}
fetchval_calls = {"n": 0}
async def _fetchval(sql, *args):
fetchval_calls["n"] += 1
# First call: existence check for acme_order_events table
if "acme_order_events" in sql:
return True
return False
async def _fetchrow(sql, *args):
return order_row
columns_no_status = [
{"column_name": "id"},
{"column_name": "user_id"},
{"column_name": "action"},
{"column_name": "resource_type"},
{"column_name": "resource_id"},
{"column_name": "details"},
{"column_name": "created_at"},
# NOTE: no "status" column — this is the canonical schema.
]
async def _fetch(sql, *args):
if "information_schema.columns" in sql and "user_activity_logs" in sql:
return columns_no_status
if "FROM acme_order_events" in sql:
return [] # empty timeline is fine for this test
if "FROM user_activity_logs" in sql:
# If the projection includes `status` we will fail loudly.
assert "status" not in sql, (
"SELECT must not include `status` when the column is absent; "
f"SQL was: {sql!r}"
)
return []
return []
conn.fetchval = _fetchval
conn.fetchrow = _fetchrow
conn.fetch = _fetch
_patch_auth_and_db(monkeypatch, conn)
out = await router_mod.get_order_events(order_id=5, authorization="Bearer x")
assert out["order_id"] == 5
assert out["count"] == 0
assert out["meta"]["correlation_id"]
# No section errors expected — schema-aware projection silently
# adapted, the panel just got an empty event list.
assert out["meta"]["errors"] == []
@pytest.mark.asyncio
async def test_bulgu95_events_acme_order_events_query_crash_returns_partial(monkeypatch):
"""A crash in the acme_order_events sub-query must NOT kill the
whole endpoint — the user_activity_logs section should still run
and the failure must appear in meta.errors."""
conn = AsyncMock()
order_row = {
"id": 5,
"account_id": 1,
"status": "invalid",
"domains": '["a.example.com"]',
"cluster_ids": "[1]",
"error_detail": None,
"post_completion_actions": None,
"pending_apply_version_name": None,
"wizard_staged_until": None,
"created_by": 1,
}
async def _fetchval(sql, *args):
if "acme_order_events" in sql:
return True
return False
async def _fetchrow(sql, *args):
return order_row
async def _fetch(sql, *args):
if "information_schema.columns" in sql:
return [
{"column_name": "id"}, {"column_name": "action"},
{"column_name": "resource_type"}, {"column_name": "resource_id"},
{"column_name": "details"}, {"column_name": "created_at"},
]
if "FROM acme_order_events" in sql:
raise RuntimeError("simulated acme_order_events index corruption")
if "FROM user_activity_logs" in sql:
return []
return []
conn.fetchval = _fetchval
conn.fetchrow = _fetchrow
conn.fetch = _fetch
_patch_auth_and_db(monkeypatch, conn)
out = await router_mod.get_order_events(order_id=5, authorization="Bearer x")
assert out["count"] == 0
error_sections = [e["section"] for e in out["meta"]["errors"]]
assert "acme_order_events" in error_sections
@pytest.mark.asyncio
async def test_bulgu95_events_load_order_404_still_raises(monkeypatch):
"""The 404 HTTPException raised by `_load_order` for an unknown order
must remain a 404 — the Round-25 envelope is only for *unexpected*
failures, not for client-supplied invalid order IDs."""
from fastapi import HTTPException
conn = AsyncMock()
conn.fetchrow.return_value = None # no order found
_patch_auth_and_db(monkeypatch, conn)
with pytest.raises(HTTPException) as exc_info:
await router_mod.get_order_events(order_id=9999, authorization="Bearer x")
assert exc_info.value.status_code == 404
@pytest.mark.asyncio
async def test_bulgu94_rerun_setup_crash_returns_check_envelope(monkeypatch):
"""The single-check re-run path must also envelope, never 500.
Pre-fix the rerun handler used the same bare ``try/finally`` shape
as the suite POST. If `_load_order` / `_enforce_rate_limit` raised,
the operator clicking the row's "Re-run" button got an opaque
toast and the row never updated. Now the rerun handler returns a
`fail` check shaped the same way the table already renders, so
the row updates in place with the cause + correlation_id.
"""
conn = AsyncMock()
conn.fetchrow.side_effect = RuntimeError("simulated DB drop during rerun")
_patch_auth_and_db(monkeypatch, conn)
out = await router_mod.rerun_diagnostic_check(
order_id=5, check_id="dns", authorization="Bearer x",
)
assert out["order_id"] == 5
assert out["check"]["status"] == "fail"
assert out["check"]["id"] == "dns"
assert "simulated DB drop during rerun" in out["check"]["message"]
assert out["meta"]["error_stage"] == "setup"
assert out["meta"]["correlation_id"]
@pytest.mark.asyncio
async def test_bulgu94_rerun_run_checks_crash_returns_check_envelope(monkeypatch):
"""A crash inside run_checks during a re-run also stays 200."""
conn = AsyncMock()
conn.fetchrow.return_value = {
"id": 5, "account_id": 1, "status": "invalid",
"domains": '["a.example.com"]', "cluster_ids": "[1]",
"error_detail": None, "post_completion_actions": None,
"pending_apply_version_name": None, "wizard_staged_until": None,
"created_by": 1,
}
_patch_auth_and_db(monkeypatch, conn)
async def _boom(*_a, **_kw):
raise RuntimeError("simulated run_checks failure")
monkeypatch.setattr(router_mod, "run_checks", _boom)
out = await router_mod.rerun_diagnostic_check(
order_id=5, check_id="agents", authorization="Bearer x",
)
assert out["check"]["id"] == "agents"
assert out["check"]["status"] == "fail"
assert "simulated run_checks failure" in out["check"]["message"]
assert out["meta"]["error_stage"] == "run_checks"
@pytest.mark.asyncio
async def test_bulgu94_rerun_invalid_check_id_still_400(monkeypatch):
"""An unknown check_id must remain a 400, not an envelope. The
envelope is only for *unexpected* server-side failures, not for
client-supplied invalid identifiers."""
from fastapi import HTTPException
conn = AsyncMock()
_patch_auth_and_db(monkeypatch, conn)
with pytest.raises(HTTPException) as exc_info:
await router_mod.rerun_diagnostic_check(
order_id=5, check_id="bogus", authorization="Bearer x",
)
assert exc_info.value.status_code == 400
@pytest.mark.asyncio
async def test_bulgu95_events_status_column_is_used_when_present(monkeypatch):
"""If a deployment DID add a `status` column (e.g. via a private
schema extension), the projection picks it up and the resulting
severity reflects it."""
conn = AsyncMock()
order_row = {
"id": 5, "account_id": 1, "status": "invalid",
"domains": '["a.example.com"]', "cluster_ids": "[1]",
"error_detail": None, "post_completion_actions": None,
"pending_apply_version_name": None, "wizard_staged_until": None,
"created_by": 1,
}
async def _fetchval(sql, *args):
if "acme_order_events" in sql:
return True
return False
async def _fetchrow(sql, *args):
return order_row
captured_sql = {"ua": None}
from datetime import datetime, timezone
async def _fetch(sql, *args):
if "information_schema.columns" in sql:
return [
{"column_name": "id"}, {"column_name": "action"},
{"column_name": "resource_type"}, {"column_name": "resource_id"},
{"column_name": "details"}, {"column_name": "created_at"},
{"column_name": "status"},
]
if "FROM acme_order_events" in sql:
return []
if "FROM user_activity_logs" in sql:
captured_sql["ua"] = sql
return [{
"id": 100,
"action": "letsencrypt.order.create",
"resource_type": "letsencrypt_order",
"resource_id": "5",
"details": '{"foo":"bar"}',
"created_at": datetime(2026, 5, 13, 19, 0, 0, tzinfo=timezone.utc),
"user_id": 7,
"status": "failure",
}]
return []
conn.fetchval = _fetchval
conn.fetchrow = _fetchrow
conn.fetch = _fetch
_patch_auth_and_db(monkeypatch, conn)
out = await router_mod.get_order_events(order_id=5, authorization="Bearer x")
assert "status" in captured_sql["ua"]
assert out["count"] == 1
ev = out["events"][0]
assert ev["source"] == "user_activity_log"
assert ev["severity"] == "warn" # status="failure" → warn
assert ev["details"] == {"foo": "bar"}
# ----------------------------------------------------------------------------
# Bulgu #96 (prod-canary follow-up): int4 overflow on order_id must NOT
# leak the asyncpg DataError message ("invalid input for query argument
# $1: ... value out of int32 range") into the operator-facing response
# body. Same shape as the "row not found" path: clean HTTPException(404).
# ----------------------------------------------------------------------------
import asyncpg as _asyncpg # noqa: E402 — late import so the test module
# can still be collected even if asyncpg has changed its exception module.
def _data_error(msg: str) -> _asyncpg.exceptions.DataError:
"""Construct an asyncpg DataError that mirrors what Postgres returns
when a path-param order_id overflows int4. We can't easily build the
real instance from the binary protocol, so we synthesize one with the
same class so the router's `except asyncpg.exceptions.DataError`
branch is exercised."""
return _asyncpg.exceptions.DataError(msg)
@pytest.mark.asyncio
async def test_bulgu96_diagnostics_int4_overflow_returns_clean_404(monkeypatch):
"""``order_id`` outside the int4 range must surface as a clean 404,
NOT as a `diagnostics_unavailable` envelope leaking the asyncpg
DataError message ("query argument $1", "int32 range").
"""
conn = AsyncMock()
conn.fetchrow.side_effect = _data_error(
"invalid input for query argument $1: 2147483648 (value out of int32 range)"
)
_patch_auth_and_db(monkeypatch, conn)
with pytest.raises(router_mod.HTTPException) as excinfo:
await router_mod.run_diagnostics(
order_id=2_147_483_648,
authorization="Bearer x",
)
assert excinfo.value.status_code == 404
# The operator must see the canonical "not found" detail, NOT the
# raw asyncpg error message.
assert "not found" in str(excinfo.value.detail).lower()
assert "int32" not in str(excinfo.value.detail).lower()
assert "query argument" not in str(excinfo.value.detail).lower()
@pytest.mark.asyncio
async def test_bulgu96_events_int4_overflow_returns_clean_404(monkeypatch):
"""Same contract on the events endpoint — out-of-range order_id is a
clean 404, not a `meta.errors[]` envelope leaking SQL detail."""
conn = AsyncMock()
conn.fetchrow.side_effect = _data_error(
"invalid input for query argument $1: 9999999999 (value out of int32 range)"
)
_patch_auth_and_db(monkeypatch, conn)
with pytest.raises(router_mod.HTTPException) as excinfo:
await router_mod.get_order_events(
order_id=9_999_999_999,
authorization="Bearer x",
)
assert excinfo.value.status_code == 404
assert "not found" in str(excinfo.value.detail).lower()
assert "int32" not in str(excinfo.value.detail).lower()
@pytest.mark.asyncio
async def test_bulgu96_rerun_int4_overflow_returns_clean_404(monkeypatch):
"""Same contract on the per-check rerun endpoint — out-of-range
order_id is a clean 404, not a structured ``check.fail`` envelope
leaking the asyncpg DataError message."""
conn = AsyncMock()
conn.fetchrow.side_effect = _data_error(
"invalid input for query argument $1: 5000000000 (value out of int32 range)"
)
_patch_auth_and_db(monkeypatch, conn)
with pytest.raises(router_mod.HTTPException) as excinfo:
await router_mod.rerun_diagnostic_check(
order_id=5_000_000_000,
check_id="dns",
authorization="Bearer x",
)
assert excinfo.value.status_code == 404
assert "not found" in str(excinfo.value.detail).lower()
assert "int32" not in str(excinfo.value.detail).lower()
@pytest.mark.asyncio
async def test_bulgu96_load_order_dataerror_does_not_leak_correlation_envelope(monkeypatch):
"""Belt-and-braces: even when the DataError carries other kinds of
invalid-input strings (e.g. type coercion failure on an int4
column), `_load_order` must still answer with the canonical 404
envelope and NOT route it through `_diagnostic_failure_envelope`
(which would surface the raw SQL detail to the UI)."""
conn = AsyncMock()
conn.fetchrow.side_effect = _data_error("invalid integer literal: 'NaN'")
_patch_auth_and_db(monkeypatch, conn)
with pytest.raises(router_mod.HTTPException) as excinfo:
await router_mod.run_diagnostics(order_id=123, authorization="Bearer x")
assert excinfo.value.status_code == 404
# No SQL detail leak
assert "invalid integer literal" not in str(excinfo.value.detail).lower()
assert "NaN" not in str(excinfo.value.detail)
+636 -1
View File
@@ -5328,7 +5328,22 @@ def test_bulgu62_update_path_grandfathers_unchanged_rule():
fe, grandfathered_signatures=grand,
)
assert len(warnings) == 1
assert "grandfathered" in warnings[0].lower() or "pre-dated" in warnings[0].lower()
# Bulgu #83 (round-23 audit) — the warning was reworded from
# "Grandfathered ... pre-dated this validation" to a clearer
# "rule was not modified by this edit" phrasing that also
# echoes the verbatim rule body. Accept either the legacy
# markers or the new ones so the contract is signal-not-
# wording.
w_lower = warnings[0].lower()
assert (
"not modified by this edit" in w_lower
or "grandfathered" in w_lower
or "pre-dated" in w_lower
), warnings[0]
# The verbatim rule must appear in the warning body so the
# operator can identify the offending entry without opening
# the ACL Builder cards.
assert stale in warnings[0], warnings[0]
def test_bulgu62_update_path_still_rejects_new_contradiction():
@@ -6267,3 +6282,623 @@ def test_bulgu82_generate_install_script_validates_cluster():
window = src[fn_start:fn_start + 5000]
assert "validate_user_cluster_access" in window
assert "Bulgu #82 (round-22 audit)" in window
# ---- Bulgu #83 — grandfathered contradiction warning wording ----
def test_bulgu83_warning_includes_verbatim_rule_text():
"""The PUT-path grandfather warning must echo the verbatim
offending rule string so the operator can identify the
offender without opening the ACL Builder cards. Pre-fix the
warning only said "1 legacy rule" via the FE toast and
"Grandfathered <label> entry contains a self-contradictory
X !X condition that pre-dated this validation" via the server
JSON payload — both omitted the actual rule body, forcing
the operator to open the modal and hunt for the dead-code
entry."""
from models.frontend import FrontendConfig
from routers.frontend import (
_enforce_routing_rule_contradictions,
_rule_to_signature,
)
stale = "be-x if acl1 !acl1"
fe = FrontendConfig(
name="fe1",
bind_port=80,
mode="http",
use_backend_rules=[stale],
)
warnings = _enforce_routing_rule_contradictions(
fe, grandfathered_signatures={_rule_to_signature(stale)},
)
assert len(warnings) == 1
# Verbatim rule body appears in the warning.
assert stale in warnings[0]
# The wording no longer uses the operator-unfriendly
# "Grandfathered" lead-in; the message states what the
# branch actually knows ("not modified by this edit").
assert "not modified by this edit" in warnings[0].lower()
def test_bulgu83_dict_redirect_warning_includes_signature_snippet():
"""Dict-shaped redirect rules with a contradictory `condition`
must still surface the offending signature in the warning
(truncated to 160 chars to bound the toast length). Pin the
truncation contract so a future refactor doesn't accidentally
grow the warning into a multi-KB blob."""
from models.frontend import FrontendConfig
from routers.frontend import (
_enforce_routing_rule_contradictions,
_rule_to_signature,
)
# Note: the redirect condition itself carries the contradiction;
# the dict wrapper is what the wizard emits.
bad_dict = {
"type": "scheme",
"scheme": "https",
"code": 301,
"condition": "if acl1 !acl1",
}
fe = FrontendConfig(
name="fe1",
bind_port=80,
mode="http",
redirect_rules=[bad_dict],
)
warnings = _enforce_routing_rule_contradictions(
fe, grandfathered_signatures={_rule_to_signature(bad_dict)},
)
assert len(warnings) == 1
# Truncated signature (up to 160 chars) must appear in body.
assert "acl1" in warnings[0]
assert "!acl1" in warnings[0]
assert "redirect_rules" in warnings[0]
def test_bulgu83_static_marker_in_fe_warning_toast():
"""Front-end pin — the FrontendManagement.js toast no longer
calls these rules "legacy" and now lists each offending rule
body. Static-source check so a refactor that re-introduces
the misleading wording or drops the rule snippets is caught.
Skipped automatically when the test runs inside the backend
Dockerfile build context (which only copies `backend/` and
therefore has no `frontend/` tree next to it). This mirrors
the guard already used by every other front-end static-source
pin in this file (e.g. lines 583-592, 795-810, 833-845, etc.)
— Bulgu #83's pin was missing it, which broke the corporate
CI's `RUN python -m pytest` step in the backend Docker build
immediately after Round-23 went live.
"""
fm_path = (
_BACKEND_DIR.parent
/ "frontend"
/ "src"
/ "components"
/ "FrontendManagement.js"
)
if not fm_path.exists():
pytest.skip(
f"frontend not present at {fm_path}; running in backend-only "
"container is expected — skip JS source pin"
)
fm_src = fm_path.read_text()
assert "Bulgu #83 (round-23 audit)" in fm_src
# No more `legacy routing/redirect rule(s)` wording.
assert "legacy routing/redirect rule(s) with a self-contradictory" not in fm_src
# The new toast wires the offending rule snippets into the
# message body via `ruleSnippets`.
assert "ruleSnippets" in fm_src
# And it re-surfaces server-emitted warnings as a safety net.
assert "serverWarnings" in fm_src
# ----------------------------------------------------------------------------
# Bulgu #84 + #85 (round-23 audit — comprehensive Site Wizard live-test pass)
# ----------------------------------------------------------------------------
def test_bulgu84_acme_diagnostics_uses_last_seen_column():
"""Bulgu #84 (round-23 audit) — duplicate pin in this file so the
full bulgu-pin suite has a one-stop reference; the original sits
in `test_acme_diagnostics.py::test_bulgu84_check_agents_uses_
last_seen_column` and asserts the AsyncMock call shape.
This duplicate is a STATIC-SOURCE check that's cheaper to grep:
we open the service file and verify the wrong column name is
not back in the SQL, plus the Bulgu marker is present so a
future refactor that drops the comment also fails.
"""
svc_path = _BACKEND_DIR / "services" / "acme_diagnostics.py"
src = svc_path.read_text()
assert "Bulgu #84 (round-23 audit)" in src, (
"Bulgu #84 marker missing from acme_diagnostics.py — the "
"fix comment was removed but the underlying SQL change "
"may have been reverted too. Re-verify check_agents."
)
# Strip comments before scanning for the wrong column — the fix
# comment legitimately mentions ``a.last_heartbeat`` as the
# pre-fix symptom. We only care about NON-comment occurrences.
non_comment_lines = "\n".join(
ln for ln in src.splitlines() if not ln.lstrip().startswith("#")
)
# The canonical SELECT inside check_agents must use a.last_seen.
assert "SELECT a.id, a.hostname, a.status, a.last_seen" in non_comment_lines, (
"acme_diagnostics.check_agents must SELECT a.last_seen "
"(the canonical agents-table column). Pre-fix this used "
"a non-existent a.last_heartbeat and every preflight-acme "
"call 500'd in production."
)
assert "a.last_heartbeat" not in non_comment_lines, (
"acme_diagnostics still references a.last_heartbeat in non-"
"comment code — this column does NOT exist on the agents "
"table. Use a.last_seen."
)
def test_bulgu87_wizard_create_endpoint_rate_limited():
"""Bulgu #87 (round-23 audit) — `POST /api/sites` must be
rate-limited, matching the existing 5/min caps on the
/preview and /preflight-acme sibling endpoints.
Pre-fix the wizard's actual create endpoint had NO rate cap:
- /api/sites/preflight-acme → rate-limited (site_acme_preflight)
- /api/sites/preview → rate-limited (site_previewed)
- /api/sites (POST, create) → UNLIMITED ← gap
Round-3 live testing fired 10 wizard creates inside one
second against demo-cluster1; the response codes were
200/409/409/...409/200/409... — no 429 ever fired. With
Bulgu #86 fixed (version_name collision), the UNIQUE-error
side-effect that ACCIDENTALLY rate-limited rapid-fire creates
disappears too, so an authenticated user/script can now
create entities until the cluster apply queue clogs.
The fix calls `_enforce_rate_limit(conn, user_id,
"wizard_create_site")` early in create_site (after auth +
cluster-access checks). The action name must match what the
router itself logs (`wizard_create_site` — see
`_log_user_activity(... action="wizard_create_site", ...)`
in the same router); the activity_logger middleware
intentionally skips this endpoint to avoid double-logging,
so the rate-limit must reference the router-emitted action.
"""
sw_src = (_BACKEND_DIR / "routers" / "site_wizard.py").read_text()
assert "Bulgu #87 (round-23 audit)" in sw_src, (
"Bulgu #87 marker missing — the wizard create rate-limit "
"may have been reverted."
)
# The create_site function body must call _enforce_rate_limit
# with the canonical create action name.
import re as _re
m = _re.search(
r"async def create_site\([^)]*\)[^{]*:(.*?)(?=\nasync def |\n@router\.|\Z)",
sw_src,
_re.DOTALL,
)
assert m, "could not locate create_site function body in site_wizard.py"
fn_body = m.group(1)
assert (
'_enforce_rate_limit(conn, user_id, "wizard_create_site")' in fn_body
or "_enforce_rate_limit(conn, user_id, 'wizard_create_site')" in fn_body
), (
"create_site must call _enforce_rate_limit(... "
'"wizard_create_site") so rapid-fire POST /api/sites '
"calls hit a 429 cap; pre-fix the endpoint was unlimited."
)
def test_bulgu86_wizard_version_name_has_unique_suffix():
"""Bulgu #86 (round-23 audit) — pin the version-name generation so
two wizard POST /api/sites calls submitted within the SAME epoch
second cannot collide on `config_versions(cluster_id, version_name)`.
Pre-fix `version_name = f"bulk-site-create-{int(time.time())}"`
had seconds resolution. Round-3 live testing reproduced the
failure with a back-to-back wizard automation: the second call
returned 409 with the misleading
"A wizard entity with this name already exists on the
cluster (UNIQUE constraint). Pick a different name."
fall-through detail, even though the operator-supplied
backend / frontend / SSL names were genuinely unique — the
real collision was on the auto-generated version name. The
operator who tried to rename the wizard inputs would still
hit the same error and have no way to make progress until the
epoch second ticked over.
Fix: append a 6-hex-char UUID suffix
(`bulk-site-create-{ts}-{uuid.uuid4().hex[:6]}`) so the
suffix space is 16M and birthday-collision-proof at any
realistic per-second request rate. The `bulk-site-create-`
prefix is preserved so reject_pending_changes /
restore-from-prior-version paths (cluster.py:4179 prefix scan)
keep working.
Static-source pin so any future refactor that strips the
suffix fails this test.
"""
sw_src = (_BACKEND_DIR / "routers" / "site_wizard.py").read_text()
assert "Bulgu #86 (round-23 audit)" in sw_src, (
"Bulgu #86 marker missing — the version_name uniqueness "
"suffix may have been reverted."
)
# The fixed form references uuid.uuid4().hex[:6] (or similar
# length-bounded random suffix) appended to the ts. A pure
# `f"bulk-site-create-{ts}"` line (without any suffix
# interpolation after ts) is the regressed form.
import re as _re
# Find the wizard's version_name assignment. We allow flexible
# spacing / quoting around the f-string but require the suffix
# interpolation token immediately after `{ts}-`.
has_suffix = bool(_re.search(
r'version_name\s*=\s*f"bulk-site-create-\{ts\}-\{[^}]+\}"',
sw_src,
))
has_regressed = bool(_re.search(
r'version_name\s*=\s*f"bulk-site-create-\{ts\}"\s*$',
sw_src,
_re.MULTILINE,
))
assert has_suffix, (
"wizard's version_name must append a unique suffix "
"(e.g. uuid.uuid4().hex[:6]) after `{ts}` so back-to-back "
"wizard creates within the same epoch second cannot "
"collide on config_versions UNIQUE."
)
assert not has_regressed, (
"wizard's version_name regressed to seconds-only resolution "
"— this re-introduces Bulgu #86 (UNIQUE collision on rapid "
"back-to-back creates)."
)
def test_bulgu85_unique_violation_handler_extracts_constraint_name():
"""Bulgu #85 (round-23 audit) — the wizard's UniqueViolationError
handler must surface the offending constraint name in the operator-
visible 409 detail so debugging a "wizard entity with this name
already exists" toast doesn't require shell access to the server
logs.
Pre-fix the handler had THREE branches:
* backends_name_cluster_id_key → "backend with this name"
* frontends_name_cluster_id_key → "frontend with this name"
* else → generic "wizard entity"
Any other constraint (SSL cert UNIQUE, backend_servers UNIQUE,
config_versions UNIQUE) fell through to the generic message,
leaving the operator with NO actionable hint and requiring an
on-call engineer to ssh into the API pod and tail `logger.info`
output to find ``constraint=<name>`` in the asyncpg traceback.
The fix:
1. Echoes the entity NAME (body.backend.name, body.frontend.name,
body.ssl.name) so the operator can match toast to wizard input.
2. Adds dedicated branches for `ssl_certificates_*` and
`backend_servers_*` constraints with actionable hints.
3. In the fall-through ELSE branch, echoes the constraint name
from `uve.constraint_name` (or substring scan as fallback)
so an unknown constraint still gives the engineer a stable
schema identifier to grep.
"""
sw_src = (_BACKEND_DIR / "routers" / "site_wizard.py").read_text()
assert "Bulgu #85 (round-23 audit)" in sw_src, (
"Bulgu #85 marker missing from site_wizard.py UniqueViolation "
"handler — the constraint-echoing fix may have been reverted."
)
# Entity-name echoing in branch detail bodies.
assert "body.backend.name" in sw_src and (
"A backend named '{body.backend.name}'" in sw_src
or "A backend named '\"{body.backend.name}\"" in sw_src
or "A backend named '" in sw_src and "body.backend.name" in sw_src
)
assert "A frontend named '{body.frontend.name}'" in sw_src or (
"A frontend named '" in sw_src and "body.frontend.name" in sw_src
)
# SSL cert and backend_servers branches present.
assert "ssl_certificates" in sw_src
assert "backend_servers" in sw_src
# Fall-through echoes constraint name.
assert "constraint=" in sw_src or "constraint_name" in sw_src
# The asyncpg attribute is preferred over str(uve) scanning.
assert 'getattr(uve, "constraint_name"' in sw_src, (
"Fix should prefer asyncpg's structured constraint_name "
"attribute over substring scanning str(uve), which is the "
"human-formatted DETAIL line and may vary across server "
"versions / locales."
)
# ----- Bulgu #88 / #89 (round-24 audit) ----------------------------------
#
# `POST /api/sites/preview` skipped the SSL-cert cluster-RBAC gate that
# `POST /api/sites` (create) enforces via `select_existing_cert()`. The
# `select_existing_cert()` helper encodes "global cert OR junction-bound
# to <cluster_id>"; pre-fix the preview accepted ssl_certificate_id values
# bound to OTHER clusters and returned the full rendered `would_create`
# envelope — a tenant-boundary information leak for the cert id and the
# rendered HAProxy config snippet. The fix mirrors the create-time check
# in the preview path: same predicate, same 400 status, same error text.
def test_bulgu88_preview_enforces_ssl_cert_cluster_rbac():
"""Preview must reject `ssl.ssl_certificate_id` referencing a cert
that is bound to a DIFFERENT cluster — same RBAC predicate the
create endpoint enforces. Static source check: the preview body
invokes `select_existing_cert` with `body.ssl.ssl_certificate_id`
and raises 400 with a "not found / inactive / not bound" hint."""
import os
sw_path = os.path.join(
os.path.dirname(__file__), "..", "routers", "site_wizard.py",
)
with open(sw_path, "r") as fh:
sw_src = fh.read()
# Locate the preview function.
assert "async def preview_create(" in sw_src
preview_start = sw_src.index("async def preview_create(")
# Bound the slice at the next top-level `async def` so we don't
# match the create_site copy (which is at line ~1900 and would
# cause this test to trivially pass even if preview was unfixed).
next_def = sw_src.index("\nasync def ", preview_start + 1)
preview_body = sw_src[preview_start:next_def]
# The preview must call select_existing_cert.
assert "select_existing_cert(" in preview_body, (
"Preview path must mirror create_site's cluster-RBAC gate by "
"invoking select_existing_cert(conn, ssl_certificate_id, "
"cluster_id). Pre-fix preview returned 200 with the full "
"rendered would_create envelope when the cert belonged to a "
"different cluster — a tenant-boundary information leak."
)
# The check must be wired to the body.ssl.ssl_certificate_id field
# (not some unrelated cert reference).
assert "body.ssl.ssl_certificate_id" in preview_body, (
"Preview RBAC check must reference body.ssl.ssl_certificate_id"
)
# Must raise 400 (consistent with create-time message).
assert "status_code=400" in preview_body
assert "not found / inactive / not bound" in preview_body or (
"not found / inactive" in preview_body
and "not bound" in preview_body
), (
"Preview should raise 400 with the same 'not found / inactive "
"/ not bound to this cluster' hint create_site emits, so wizard "
"clients can match the message uniformly."
)
def test_bulgu89_preview_enforces_per_server_ca_bundle_cluster_rbac():
"""Preview must also enforce per-server `ssl_certificate_id` (CA
bundle) cluster-RBAC. Pre-fix only `create_site` checked this — a
direct API caller could submit a preview with
`servers[i].ssl_certificate_id=<cert from another cluster>` and
receive the full rendered backend block back, including the
`ca-file` directive bound to a cert id the caller's cluster has
no junction row for."""
import os
sw_path = os.path.join(
os.path.dirname(__file__), "..", "routers", "site_wizard.py",
)
with open(sw_path, "r") as fh:
sw_src = fh.read()
preview_start = sw_src.index("async def preview_create(")
next_def = sw_src.index("\nasync def ", preview_start + 1)
preview_body = sw_src[preview_start:next_def]
# Must iterate body.servers and check per-server ssl_certificate_id.
assert "body.servers" in preview_body, (
"Preview must iterate body.servers to validate each "
"server.ssl_certificate_id reference"
)
assert "ssl_certificate_id" in preview_body
# The per-server message must be field-qualified so operators see
# WHICH server triggered the gate.
assert "servers[" in preview_body and "ssl_certificate_id=" in preview_body, (
"Per-server preview error must include the array index "
"(servers[i].ssl_certificate_id=<id>) so operators can find "
"the offending row in their wizard payload."
)
# ----- Bulgu #90 (round-24 audit) ----------------------------------------
#
# `backend.cookie_name`, `backend.cookie_options`, and `server.cookie_value`
# only rejected newline characters pre-fix. HAProxy's directive parser
# tokenises by whitespace and treats `;` as an inline-comment, so a value
# like `SESS'; DROP TABLE backends; --` rendered as
# `cookie SESS'; DROP TABLE backends; -- insert indirect nocache` which
# HAProxy parsed as `cookie SESS'` + comment — silently truncating the
# operator's persistence options and producing a malformed Set-Cookie
# header that browsers may drop. Tighten to the conservative
# alphanumeric+`_.-` set (cookie_name / cookie_value) and the keyword/
# `=`-bearing set (cookie_options).
def test_bulgu90_cookie_name_rejects_haproxy_parser_hostile_chars():
"""cookie_name must reject `;` (HAProxy inline comment), whitespace
(token delimiter), and other punctuation that breaks the rendered
`cookie <name>` directive."""
from models.site_wizard import BackendStep
import pydantic
# Semicolon (HAProxy comment) — used to be silently accepted.
with pytest.raises(pydantic.ValidationError):
BackendStep(name="be1", cookie_name="SESS'; DROP TABLE backends; --")
# Space (token delimiter) — splits the cookie directive.
with pytest.raises(pydantic.ValidationError):
BackendStep(name="be1", cookie_name="SESS ID")
# Quote (legal in RFC 6265 but produces ugly Set-Cookie headers
# and confuses log scrapers).
with pytest.raises(pydantic.ValidationError):
BackendStep(name="be1", cookie_name='SESS"id')
# Common-case valid input continues to work.
BackendStep(name="be1", cookie_name="SESS_ID-1.app")
def test_bulgu90_cookie_options_rejects_parser_hostile_chars():
"""cookie_options must reject `;` and quoting metacharacters but
still accept legitimate `attr SameSite=Lax` style values."""
from models.site_wizard import BackendStep
import pydantic
with pytest.raises(pydantic.ValidationError):
BackendStep(
name="be1", cookie_name="SESS",
cookie_options="insert indirect; rm -rf /",
)
with pytest.raises(pydantic.ValidationError):
BackendStep(
name="be1", cookie_name="SESS",
cookie_options='insert "indirect"',
)
# Valid HAProxy cookie options continue to parse.
BackendStep(
name="be1", cookie_name="SESS",
cookie_options="insert indirect nocache attr SameSite=Lax",
)
def test_bulgu90_server_cookie_value_rejects_rfc6265_disallowed_chars():
"""server.cookie_value must reject `;`, whitespace, and other
chars RFC 6265 disallows in cookie-values."""
from models.site_wizard import ServerStep
import pydantic
with pytest.raises(pydantic.ValidationError):
ServerStep(
server_name="s1", server_address="10.0.0.1",
server_port=8080, cookie_value="srv1; secure",
)
with pytest.raises(pydantic.ValidationError):
ServerStep(
server_name="s1", server_address="10.0.0.1",
server_port=8080, cookie_value="srv 1",
)
# Valid cookie values continue to work.
ServerStep(
server_name="s1", server_address="10.0.0.1",
server_port=8080, cookie_value="srv-1.app_2",
)
# ----- Bulgu #93 (round-24 audit) ----------------------------------------
#
# `GET /api/sites/suggest` produced backend/frontend name suggestions
# using `c.isalnum()` over the first domain label. Python's `isalnum()`
# is Unicode-aware and returns True for non-ASCII letters (ü, é, ñ, …),
# so an operator typing `bücher.example.com` received
# `backend_name='be-bücher'`. The wizard CREATE path then rejected the
# very name the SUGGEST endpoint returned, because the entity-name
# regex (`^[a-zA-Z][a-zA-Z0-9_-]{0,63}$`) and the domain validator are
# both ASCII-only. The fix converts non-ASCII labels through IDN/
# punycode FIRST and only then applies ASCII-only sanitisation, so the
# resulting name is one the operator can submit unchanged.
def test_bulgu93_suggest_idn_unicode_produces_ascii_safe_slug():
"""`suggest` must produce slugs that the backend/frontend name
regex `^[a-zA-Z][a-zA-Z0-9_-]{0,63}$` accepts. Static source check:
the suggest function uses an ASCII-only predicate
(`c.isascii() and c.isalnum()` etc.) and routes IDN labels through
`encode('idna')` to preserve the operator's intent in punycode."""
import os, re
sw_path = os.path.join(
os.path.dirname(__file__), "..", "routers", "site_wizard.py",
)
with open(sw_path, "r") as fh:
sw_src = fh.read()
# Locate the suggest_defaults function body.
assert "async def suggest_defaults(" in sw_src
start = sw_src.index("async def suggest_defaults(")
end = sw_src.index("\n@router.", start + 1)
suggest_body = sw_src[start:end]
# The fix must use IDN/punycode encoding when the input is non-ASCII.
assert 'encode("idna")' in suggest_body or "encode('idna')" in suggest_body, (
"suggest_defaults must route non-ASCII labels through IDN/"
"punycode so the rendered slug matches the ASCII-only entity-"
"name regex the create endpoint enforces."
)
# The sanitiser must use an ASCII-only predicate, not Unicode-aware
# `isalnum()` alone (which is the pre-fix bug).
assert "isascii()" in suggest_body, (
"Sanitisation step must explicitly bound to ASCII (via "
"`c.isascii()`) so Unicode alphabetics get mapped to '-' rather "
"than retained verbatim."
)
# The first occurrence of the Unicode-only `c.isalnum()` (without an
# adjacent isascii() guard) must no longer exist in the slug-build
# path. We sanity-check by counting the matches of the bare
# `c.isalnum()` token within the function body — pre-fix there was
# exactly one such occurrence in the slug builder.
bare = re.findall(r"\bc\.isalnum\(\)\b", suggest_body)
assert all(
# Each occurrence must be paired with an `isascii()` guard on
# the same line — confirm by inspecting the line context.
any("isascii()" in ln for ln in suggest_body.splitlines() if "c.isalnum()" in ln)
for _ in bare
), (
"Every `c.isalnum()` predicate in the slug builder must be "
"AND'd with `c.isascii()` so non-ASCII letters cannot leak "
"into the suggested entity names."
)
def test_bulgu93_suggest_function_behavioural_smoke():
"""Behavioural check: the same logic, invoked directly, should
produce an ASCII-only slug for a Unicode input."""
# Replicate the post-fix slug builder inline so the test does not
# require an event loop / FastAPI fixture. The point is to confirm
# the algorithmic shape matches what the source check above
# enforces.
def _build_slug(domain: str) -> str:
first_label = (
domain.replace("*.", "").split(".")[0]
if domain.replace("*.", "")
else ""
)
ascii_label = first_label
if first_label and not first_label.isascii():
try:
ascii_label = first_label.encode("idna").decode("ascii")
except (UnicodeError, UnicodeDecodeError):
ascii_label = "".join(
c if c.isascii() and (c.isalnum() or c in ("-", "_"))
else "-"
for c in first_label
)
slug = (ascii_label or "newhost")[:32].lower()
slug = "".join(
c if c.isascii() and (c.isalnum() or c in ("-", "_"))
else "-"
for c in slug
)
if slug.startswith("_"):
slug = "h-" + slug.lstrip("_")
if not slug or not slug[0].isalpha():
slug = "h-" + slug
return slug
# Unicode IDN — should round-trip through punycode.
s = _build_slug("bücher.example.com")
# punycode of 'bücher' is 'xn--bcher-kva'
assert s == "xn--bcher-kva", f"expected xn--bcher-kva got {s!r}"
# The slug must satisfy the entity-name regex.
import re
assert re.fullmatch(r"^[a-zA-Z][a-zA-Z0-9_-]{0,63}$", s), s
# Plain ASCII domain still produces the obvious slug.
assert _build_slug("example.com") == "example"
# Uppercase normalises to lowercase.
assert _build_slug("EXAMPLE.COM") == "example"
# Trailing dot (FQDN absolute) stripped.
assert _build_slug("example.com.") == "example"
# Empty / dots-only input falls back to 'newhost'.
assert _build_slug("") == "newhost"
assert _build_slug("...") == "newhost"
+150
View File
@@ -0,0 +1,150 @@
"""Backwards-compatibility regression tests for the MFA rollout (Issue #18).
These tests don't hit a real database — they exercise the authoritative
contract surfaces (login response shape, auth_middleware behaviour) using
mocks where needed so the suite stays fast and deterministic.
"""
from __future__ import annotations
import os
import sys
from unittest.mock import AsyncMock, MagicMock, patch
import pytest
sys.path.insert(0, os.path.dirname(os.path.dirname(os.path.abspath(__file__))))
class TestAuthMiddlewareUnchanged:
"""auth_middleware MUST NOT look for MFA claims — Madde 2 of the plan."""
def test_decoder_imports_without_mfa_dependencies(self):
import auth_middleware
# The middleware's verification function exists and is callable.
assert callable(getattr(auth_middleware, "get_current_user_from_token", None))
def test_middleware_source_has_no_mfa_claim_check(self):
"""The middleware source must not reference ``mfa`` claims directly."""
with open(
os.path.join(
os.path.dirname(os.path.dirname(os.path.abspath(__file__))),
"auth_middleware.py",
),
"r",
encoding="utf-8",
) as fh:
source = fh.read()
# Allow incidental occurrences (e.g. comments); but never a claim lookup.
assert "payload.get('mfa'" not in source
assert 'payload.get("mfa"' not in source
assert "claims['mfa'" not in source
assert 'claims["mfa"' not in source
class TestMfaModelsCoexistWithUserModels:
def test_models_user_module_unchanged_pydantic_shape(self):
from models import user as user_mod
# Ensure that User / UserUpdate / LoginRequest still load and still
# don't expose mfa-related fields (kept in models.mfa).
for cls in (user_mod.User, user_mod.UserUpdate, user_mod.LoginRequest):
fields = set(cls.model_fields.keys())
assert not {"mfa_enabled", "mfa_required", "mfa_token"} & fields, (
f"{cls.__name__} unexpectedly exposes MFA field; should stay byte-identical."
)
def test_models_mfa_module_exposes_expected_models(self):
from models import mfa as mfa_mod
for name in (
"MfaVerifyRequest",
"MfaEnrollStartResponse",
"MfaEnrollConfirmRequest",
"MfaEnrollConfirmResponse",
"MfaDisableRequest",
"MfaRegenerateBackupRequest",
"MfaRegenerateBackupResponse",
"MfaAdminResetRequest",
"MfaAdminResetAllRequest",
"MfaStatusResponse",
):
assert hasattr(mfa_mod, name), f"Missing model: {name}"
class TestRouterIncluded:
def test_main_includes_mfa_router(self):
with open(
os.path.join(
os.path.dirname(os.path.dirname(os.path.abspath(__file__))),
"main.py",
),
"r",
encoding="utf-8",
) as fh:
source = fh.read()
assert "from routers.mfa import router as mfa_router" in source
assert "app.include_router(mfa_router)" in source
class TestLoginResponseShapeForNonMfaUser:
"""When MFA columns are missing or mfa_enabled=FALSE, /login returns the
pre-MFA response shape — no ``mfa_required`` / ``mfa_token`` keys leak through.
"""
def test_login_without_mfa_returns_legacy_shape(self):
from fastapi.testclient import TestClient
from main import app
client = TestClient(app)
async def _fetch_mfa_state_none(conn, user_id):
return None
async def _no_log(*args, **kwargs):
return None
fake_user = {
"id": 1,
"username": "admin",
"email": "admin@example.com",
"password_hash": "$2b$12$placeholder",
"is_active": True,
"role": "admin",
"created_at": None,
"updated_at": None,
"last_login_at": None,
}
mock_conn = MagicMock()
mock_conn.fetchrow = AsyncMock(return_value=fake_user)
mock_conn.fetch = AsyncMock(return_value=[])
mock_conn.execute = AsyncMock(return_value=None)
async def _get_conn():
return mock_conn
async def _close(conn):
return None
with patch(
"routers.auth.get_database_connection", _get_conn
), patch("routers.auth.close_database_connection", _close), patch(
"routers.auth._fetch_mfa_state", _fetch_mfa_state_none
), patch("routers.auth.log_user_activity", _no_log), patch(
"bcrypt.checkpw", return_value=True
):
resp = client.post(
"/api/auth/login",
json={"username": "admin", "password": "anything"},
)
assert resp.status_code == 200, resp.text
body = resp.json()
assert "access_token" in body
assert "token_type" in body
assert "expires_in" in body
assert "user" in body
assert "roles" in body
assert "permissions" in body
# CRITICAL — pre-MFA contract must not be polluted with MFA fields.
assert "mfa_required" not in body
assert "mfa_token" not in body
assert "methods" not in body
+200
View File
@@ -0,0 +1,200 @@
"""Tests for middleware.mfa_rate_limit_key — user-aware + ingress-aware key."""
import importlib
from datetime import datetime, timedelta
from typing import Dict, Optional
import pytest
from fastapi import Request
from jose import jwt
def _make_request(
headers: Optional[Dict[str, str]] = None,
peer: str = "127.0.0.1",
) -> Request:
"""Tiny ASGI scope shim — enough for the key_func surface."""
raw_headers = []
if headers:
raw_headers = [
(k.encode("latin-1"), v.encode("latin-1")) for k, v in headers.items()
]
scope = {
"type": "http",
"headers": raw_headers,
"client": (peer, 12345),
"method": "POST",
"path": "/api/mfa/enroll/start",
"query_string": b"",
}
return Request(scope)
@pytest.fixture
def reload_key(monkeypatch):
"""Reload the key module so MFA_TRUSTED_PROXY_CIDRS is re-parsed."""
def _reload(**env):
for var in ("MFA_TRUSTED_PROXY_CIDRS",):
monkeypatch.delenv(var, raising=False)
for k, v in env.items():
monkeypatch.setenv(k, v)
from middleware import mfa_rate_limit_key as m
return importlib.reload(m)
return _reload
# ---------------------------------------------------------------------------
# User-aware key extraction
# ---------------------------------------------------------------------------
def _mint_jwt(user_id, claim: str = "user_id") -> str:
"""Mint a test JWT. Note: RFC 7519 says ``sub`` is a StringOrURI,
and python-jose validates that type when decoding, so callers that use
``claim='sub'`` must pass a string user_id (matches production behavior
where auth_middleware also accepts string ``sub``)."""
from config import JWT_ALGORITHM, JWT_SECRET_KEY
payload = {
claim: user_id,
"exp": datetime.utcnow() + timedelta(minutes=10),
}
return jwt.encode(payload, JWT_SECRET_KEY, algorithm=JWT_ALGORITHM)
def test_user_aware_via_user_id_claim(reload_key):
m = reload_key()
token = _mint_jwt(42, claim="user_id")
req = _make_request(headers={"authorization": f"Bearer {token}"})
assert m.mfa_rate_limit_key(req) == "user:42"
def test_user_aware_via_sub_claim(reload_key):
m = reload_key()
token = _mint_jwt("7", claim="sub") # JWT spec: sub is a string
req = _make_request(headers={"authorization": f"Bearer {token}"})
assert m.mfa_rate_limit_key(req) == "user:7"
def test_no_auth_header_falls_back_to_ip(reload_key):
m = reload_key()
req = _make_request(peer="203.0.113.5")
assert m.mfa_rate_limit_key(req) == "ip:203.0.113.5"
def test_missing_bearer_prefix_falls_back_to_ip(reload_key):
m = reload_key()
req = _make_request(headers={"authorization": "abc.def.ghi"}, peer="203.0.113.5")
assert m.mfa_rate_limit_key(req) == "ip:203.0.113.5"
def test_bearer_null_or_undefined_falls_back_to_ip(reload_key):
m = reload_key()
for bogus in ("null", "undefined", "", " "):
req = _make_request(
headers={"authorization": f"Bearer {bogus}"}, peer="198.51.100.9"
)
assert m.mfa_rate_limit_key(req) == "ip:198.51.100.9"
def test_tampered_jwt_falls_back_to_ip(reload_key):
"""A token with a forged signature must NOT be honored — fallback to IP."""
m = reload_key()
bad = "eyJhbGciOiJIUzI1NiJ9.eyJ1c2VyX2lkIjogMTIzfQ.NOT_A_VALID_SIGNATURE"
req = _make_request(headers={"authorization": f"Bearer {bad}"}, peer="10.1.2.3")
assert m.mfa_rate_limit_key(req) == "ip:10.1.2.3"
def test_expired_jwt_falls_back_to_ip(reload_key):
m = reload_key()
from config import JWT_ALGORITHM, JWT_SECRET_KEY
payload = {
"user_id": 9,
"exp": datetime.utcnow() - timedelta(minutes=5),
}
expired = jwt.encode(payload, JWT_SECRET_KEY, algorithm=JWT_ALGORITHM)
req = _make_request(headers={"authorization": f"Bearer {expired}"}, peer="10.0.0.7")
assert m.mfa_rate_limit_key(req) == "ip:10.0.0.7"
# ---------------------------------------------------------------------------
# Trusted-proxy X-Forwarded-For handling
# ---------------------------------------------------------------------------
def test_untrusted_peer_xff_is_ignored(reload_key):
"""X-Forwarded-For from an untrusted client cannot move buckets."""
m = reload_key() # no trusted CIDRs
req = _make_request(
headers={"x-forwarded-for": "1.2.3.4"},
peer="203.0.113.5",
)
assert m.mfa_rate_limit_key(req) == "ip:203.0.113.5"
def test_trusted_peer_xff_is_honored(reload_key):
"""Peer in trusted CIDR → first XFF hop becomes the bucket."""
m = reload_key(MFA_TRUSTED_PROXY_CIDRS="10.0.0.0/8")
req = _make_request(
headers={"x-forwarded-for": "203.0.113.42, 10.0.0.99"},
peer="10.0.0.99",
)
assert m.mfa_rate_limit_key(req) == "ip:203.0.113.42"
def test_trusted_cidr_multiple_ranges(reload_key):
m = reload_key(MFA_TRUSTED_PROXY_CIDRS="10.0.0.0/8, 172.16.0.0/12")
req = _make_request(
headers={"x-forwarded-for": "198.51.100.4"},
peer="172.16.5.5",
)
assert m.mfa_rate_limit_key(req) == "ip:198.51.100.4"
def test_trusted_peer_no_xff_falls_back_to_peer(reload_key):
m = reload_key(MFA_TRUSTED_PROXY_CIDRS="10.0.0.0/8")
req = _make_request(peer="10.0.0.99")
assert m.mfa_rate_limit_key(req) == "ip:10.0.0.99"
def test_invalid_cidr_in_env_is_logged_and_ignored(reload_key, caplog):
import logging
with caplog.at_level(logging.WARNING, logger="middleware.mfa_rate_limit_key"):
m = reload_key(MFA_TRUSTED_PROXY_CIDRS="not-a-cidr, 10.0.0.0/8")
assert any("ignoring invalid CIDR" in r.message for r in caplog.records)
# The valid one is still effective.
req = _make_request(
headers={"x-forwarded-for": "9.9.9.9"},
peer="10.0.0.1",
)
assert m.mfa_rate_limit_key(req) == "ip:9.9.9.9"
def test_user_bucket_wins_over_xff(reload_key):
"""Auth always wins, even from a trusted proxy."""
m = reload_key(MFA_TRUSTED_PROXY_CIDRS="10.0.0.0/8")
token = _mint_jwt(99) # integer user_id claim
req = _make_request(
headers={
"authorization": f"Bearer {token}",
"x-forwarded-for": "1.1.1.1",
},
peer="10.0.0.1",
)
assert m.mfa_rate_limit_key(req) == "user:99"
def test_no_client_in_scope_does_not_crash(reload_key):
m = reload_key()
scope = {
"type": "http",
"headers": [],
"client": None,
"method": "POST",
"path": "/api/mfa/enroll/start",
"query_string": b"",
}
req = Request(scope)
# Whatever it returns, it must be deterministic and not raise.
out = m.mfa_rate_limit_key(req)
assert out.startswith("ip:")
+93
View File
@@ -0,0 +1,93 @@
"""Tests for middleware.mfa_rate_limits — env-driven MFA rate-limit config."""
import importlib
import logging
import pytest
@pytest.fixture
def reload_module(monkeypatch):
"""Helper: reload the module after env mutation so dataclass defaults
pick up the new values."""
def _reload(**env):
for key in list(globals().get('_OVERRIDDEN_ENVS', set())):
monkeypatch.delenv(key, raising=False)
for key, value in env.items():
monkeypatch.setenv(key, value)
from middleware import mfa_rate_limits as m
return importlib.reload(m)
return _reload
def test_defaults_when_no_env(reload_module, monkeypatch):
"""No env var set → secure defaults applied (user-aware key assumption)."""
for key in (
"MFA_RATE_LIMIT_ENROLL_START",
"MFA_RATE_LIMIT_ENROLL_CONFIRM",
"MFA_RATE_LIMIT_DISABLE",
"MFA_RATE_LIMIT_REGENERATE_BACKUP_CODES",
"MFA_RATE_LIMIT_ADMIN_RESET",
"MFA_RATE_LIMIT_ADMIN_RESET_ALL",
):
monkeypatch.delenv(key, raising=False)
m = reload_module()
assert m.MFA_LIMITS.enroll_start == "10/minute"
assert m.MFA_LIMITS.enroll_confirm == "10/minute"
assert m.MFA_LIMITS.disable == "10/minute"
assert m.MFA_LIMITS.regenerate_backup_codes == "5/hour"
assert m.MFA_LIMITS.admin_reset == "60/hour"
assert m.MFA_LIMITS.admin_reset_all == "1/day"
def test_env_override_per_endpoint(reload_module):
m = reload_module(
MFA_RATE_LIMIT_ENROLL_START="100/hour",
MFA_RATE_LIMIT_ADMIN_RESET_ALL="3/day",
)
assert m.MFA_LIMITS.enroll_start == "100/hour"
assert m.MFA_LIMITS.admin_reset_all == "3/day"
# Untouched values still default.
assert m.MFA_LIMITS.disable == "10/minute"
@pytest.mark.parametrize(
"bad",
[
"totally-bogus",
"5/lightyear",
"abc/minute",
"5",
"/minute",
"5//minute",
"",
],
)
def test_invalid_format_falls_back_to_default(reload_module, caplog, bad):
with caplog.at_level(logging.WARNING, logger="middleware.mfa_rate_limits"):
m = reload_module(MFA_RATE_LIMIT_ENROLL_START=bad)
# Falls back to the secure default for enroll_start.
assert m.MFA_LIMITS.enroll_start == "10/minute"
assert any("not a valid slowapi limit string" in r.message for r in caplog.records)
def test_whitespace_around_value_is_tolerated(reload_module):
m = reload_module(MFA_RATE_LIMIT_DISABLE=" 30/minute ")
assert m.MFA_LIMITS.disable == "30/minute"
@pytest.mark.parametrize(
"valid",
["1/second", "100/minute", "1000/hour", "10/day"],
)
def test_all_valid_periods_accepted(reload_module, valid):
m = reload_module(MFA_RATE_LIMIT_DISABLE=valid)
assert m.MFA_LIMITS.disable == valid
def test_dataclass_is_frozen(reload_module):
"""Frozen dataclass guards against accidental mutation after import."""
m = reload_module()
with pytest.raises((AttributeError, Exception)):
m.MFA_LIMITS.disable = "999/second" # type: ignore[misc]
+223
View File
@@ -0,0 +1,223 @@
"""Unit tests for the MFA service layer (Issue #18, v1.6.0).
These tests cover the pure-Python side of MFA — no DB, no FastAPI.
"""
from __future__ import annotations
import asyncio
import os
import re
import sys
import time
import pytest
# Repo path setup (mirrors other tests in this folder).
sys.path.insert(0, os.path.dirname(os.path.dirname(os.path.abspath(__file__))))
import pyotp # noqa: E402
from services import mfa_service # noqa: E402
# ---------------------------------------------------------------------------
# TOTP
# ---------------------------------------------------------------------------
class TestTotpSecret:
def test_secret_is_base32(self):
secret = mfa_service.generate_totp_secret()
# pyotp.random_base32() returns 32-character base32 strings.
assert len(secret) == 32
assert re.fullmatch(r"[A-Z2-7]+", secret), "secret must be valid base32"
def test_secrets_are_unique(self):
secrets = {mfa_service.generate_totp_secret() for _ in range(50)}
assert len(secrets) == 50
class TestVerifyTotp:
def setup_method(self):
self.secret = mfa_service.generate_totp_secret()
self.totp = pyotp.TOTP(self.secret, digits=6, interval=30, digest="sha1")
def test_happy_path(self):
code = self.totp.now()
ok, step = mfa_service.verify_totp_with_replay_guard(self.secret, code, None)
assert ok is True
assert step == int(time.time()) // 30
def test_invalid_code_format_rejected(self):
ok, step = mfa_service.verify_totp_with_replay_guard(self.secret, "abc", None)
assert ok is False and step is None
ok, step = mfa_service.verify_totp_with_replay_guard(self.secret, "12345", None)
assert ok is False and step is None
def test_tolerance_minus_30s(self):
now = int(time.time())
previous_step = (now // 30) - 1
prev_code = self.totp.at(previous_step * 30)
ok, step = mfa_service.verify_totp_with_replay_guard(self.secret, prev_code, None)
assert ok is True
assert step == previous_step
def test_tolerance_plus_30s(self):
now = int(time.time())
next_step = (now // 30) + 1
next_code = self.totp.at(next_step * 30)
ok, step = mfa_service.verify_totp_with_replay_guard(self.secret, next_code, None)
assert ok is True
assert step == next_step
def test_replay_rejected(self):
code = self.totp.now()
ok, step = mfa_service.verify_totp_with_replay_guard(self.secret, code, None)
assert ok is True
# Submit again with the previously-consumed step: must be rejected.
ok2, step2 = mfa_service.verify_totp_with_replay_guard(self.secret, code, step)
assert ok2 is False
assert step2 is None
def test_wrong_code_rejected(self):
ok, step = mfa_service.verify_totp_with_replay_guard(self.secret, "000000", None)
assert ok is False
assert step is None
# ---------------------------------------------------------------------------
# Fernet + key resolution
# ---------------------------------------------------------------------------
class TestFernet:
def setup_method(self):
mfa_service.reset_fernet_for_tests()
def teardown_method(self):
mfa_service.reset_fernet_for_tests()
def test_encrypt_decrypt_roundtrip_with_env_key(self, monkeypatch):
from cryptography.fernet import Fernet
key = Fernet.generate_key().decode()
monkeypatch.setenv("MFA_ENCRYPTION_KEY", key)
mfa_service.reset_fernet_for_tests()
secret = "JBSWY3DPEHPK3PXP" * 2
token = mfa_service.encrypt_secret(secret)
assert token and token != secret
recovered = mfa_service.decrypt_secret(token)
assert recovered == secret
def test_decrypt_invalid_token_returns_none(self, monkeypatch):
from cryptography.fernet import Fernet
monkeypatch.setenv("MFA_ENCRYPTION_KEY", Fernet.generate_key().decode())
mfa_service.reset_fernet_for_tests()
assert mfa_service.decrypt_secret("not-a-valid-fernet-token") is None
def test_hkdf_fallback_when_env_unset(self, monkeypatch, caplog):
monkeypatch.delenv("MFA_ENCRYPTION_KEY", raising=False)
mfa_service.reset_fernet_for_tests()
with caplog.at_level("WARNING"):
secret = "JBSWY3DPEHPK3PXPJBSWY3DPEHPK3PXP"
token = mfa_service.encrypt_secret(secret)
recovered = mfa_service.decrypt_secret(token)
assert recovered == secret
assert any("MFA_ENCRYPTION_KEY" in r.message for r in caplog.records), (
"expected a WARN log when falling back to SECRET_KEY derivation"
)
# ---------------------------------------------------------------------------
# Backup codes
# ---------------------------------------------------------------------------
class TestBackupCodes:
def test_generate_count_and_format(self):
codes = mfa_service.generate_backup_codes()
assert len(codes) == 10
# 31-char alphabet: A-H J K M N P-Z 2-9 (excludes I, L, O, 0, 1).
for code in codes:
assert re.fullmatch(r"[A-HJKM-NP-Z2-9]{4}-[A-HJKM-NP-Z2-9]{4}", code), code
def test_alphabet_excludes_confusing_characters(self):
# Generate enough codes to virtually guarantee any forbidden char would surface.
for _ in range(20):
codes = mfa_service.generate_backup_codes()
for code in codes:
for ch in code.replace("-", ""):
assert ch not in "0O1IL", f"forbidden char {ch!r} in {code!r}"
def test_codes_are_unique(self):
codes = mfa_service.generate_backup_codes()
assert len(set(codes)) == len(codes)
def test_normalize_strips_case_dash_space(self):
assert mfa_service.normalize_backup_code("abcd-efgh") == "ABCDEFGH"
assert mfa_service.normalize_backup_code(" ab cd-ef gh ") == "ABCDEFGH"
assert mfa_service.normalize_backup_code("") == ""
assert mfa_service.normalize_backup_code(None) == "" # type: ignore[arg-type]
def test_hash_and_check_async(self):
async def _run():
plain = mfa_service.generate_backup_codes()[:1]
hashes = await mfa_service.hash_backup_codes(plain)
assert len(hashes) == 1
assert await mfa_service.check_backup_code(plain[0], hashes[0]) is True
assert await mfa_service.check_backup_code("WRONG-CODE", hashes[0]) is False
# Case + dash normalization
assert await mfa_service.check_backup_code(plain[0].lower(), hashes[0]) is True
assert (
await mfa_service.check_backup_code(plain[0].replace("-", ""), hashes[0])
is True
)
asyncio.run(_run())
# ---------------------------------------------------------------------------
# otpauth URI
# ---------------------------------------------------------------------------
class TestOtpAuthUri:
def test_uri_shape(self):
uri = mfa_service.build_otpauth_uri("alice@example.com", "JBSWY3DPEHPK3PXP")
assert uri.startswith("otpauth://totp/")
assert "secret=JBSWY3DPEHPK3PXP" in uri
assert "issuer=" in uri
assert "algorithm=SHA1" in uri
assert "digits=6" in uri
assert "period=30" in uri
def test_account_label_env_override(self, monkeypatch):
monkeypatch.setenv("MFA_ACCOUNT_LABEL_DOMAIN", "ops.example.com")
label = mfa_service.build_account_label("alice", hostname_hint="ignored.com")
assert label == "alice@ops.example.com"
def test_account_label_hostname_hint(self, monkeypatch):
monkeypatch.delenv("MFA_ACCOUNT_LABEL_DOMAIN", raising=False)
label = mfa_service.build_account_label("alice", hostname_hint="api.local")
assert label == "alice@api.local"
def test_account_label_fallback(self, monkeypatch):
monkeypatch.delenv("MFA_ACCOUNT_LABEL_DOMAIN", raising=False)
label = mfa_service.build_account_label("alice", hostname_hint=None)
assert label == "alice@haproxy-openmanager"
# ---------------------------------------------------------------------------
# Challenge token
# ---------------------------------------------------------------------------
class TestChallengeToken:
def test_length_and_uniqueness(self):
tokens = {mfa_service.generate_challenge_token() for _ in range(50)}
assert len(tokens) == 50
for t in tokens:
assert len(t) == 64
assert re.fullmatch(r"[0-9a-f]{64}", t)
+5 -3
View File
@@ -71,15 +71,17 @@ services:
# Frontend React App - pulls from Docker Hub by default
# To build locally instead, uncomment the build section and run: docker compose build frontend
#
# NOTE: REACT_APP_* env vars are BUILD-time only for Create-React-App. The
# runtime container (serve -s build) does NOT consume them. The frontend
# uses same-origin (window.location) for /api/* and is routed by the nginx
# service below to the backend container. No env vars are required here.
frontend:
image: taylanbakircioglu/haproxy-openmanager-frontend:latest
# build:
# context: ./frontend
# dockerfile: Dockerfile
container_name: haproxy-openmanager-frontend
environment:
- REACT_APP_API_URL=
- NODE_ENV=production
expose:
- "3000"
depends_on:
Binary file not shown.

After

Width:  |  Height:  |  Size: 378 KiB

+34
View File
@@ -0,0 +1,34 @@
# Build artifacts (regenerated inside the multi-stage builder)
build
node_modules
coverage
# Local dev-only overrides — MUST be excluded so the host's
# `.env.local` (e.g. REACT_APP_API_URL=http://localhost:8000) does NOT
# bleed into a production bundle via build-time inline-replace.
.env
.env.local
.env.*.local
.env.development
.env.development.local
.env.test
.env.test.local
# VCS / IDE / OS noise
.git
.gitignore
.vscode
.idea
.DS_Store
*.log
# Tests
**/__tests__
**/*.test.js
**/*.test.jsx
**/*.test.ts
**/*.test.tsx
# Misc
README.md
.eslintcache
+16 -23
View File
@@ -1,12 +1,13 @@
{
"name": "haproxy-openmanager-frontend",
"version": "1.5.0",
"version": "1.6.0",
"lockfileVersion": 3,
"requires": true,
"packages": {
"": {
"name": "haproxy-openmanager-frontend",
"version": "1.5.0",
"version": "1.6.0",
"license": "AGPL-3.0-or-later",
"dependencies": {
"@ant-design/icons": "^5.0.0",
"@monaco-editor/react": "^4.6.0",
@@ -15,6 +16,7 @@
"axios": "^1.3.0",
"moment": "^2.29.0",
"monaco-editor": "^0.36.0",
"qrcode.react": "^4.0.0",
"react": "^18.2.0",
"react-ace": "^10.1.0",
"react-dom": "^18.2.0",
@@ -1500,9 +1502,9 @@
}
},
"node_modules/@babel/plugin-transform-modules-systemjs": {
"version": "7.29.0",
"resolved": "https://registry.npmjs.org/@babel/plugin-transform-modules-systemjs/-/plugin-transform-modules-systemjs-7.29.0.tgz",
"integrity": "sha512-PrujnVFbOdUpw4UHiVwKvKRLMMic8+eC0CuNlxjsyZUiBjhFdPsewdXCkveh2KqBA9/waD0W1b4hXSOBQJezpQ==",
"version": "7.29.4",
"resolved": "https://registry.npmjs.org/@babel/plugin-transform-modules-systemjs/-/plugin-transform-modules-systemjs-7.29.4.tgz",
"integrity": "sha512-N7QmZ0xRZfjHOfZeQLJjwgX2zS9pdGHSVl/cjSGlo4dXMqvurfxXDMKY4RqEKzPozV78VMcd0lxyG13mlbKc4w==",
"dev": true,
"license": "MIT",
"dependencies": {
@@ -18237,6 +18239,15 @@
"teleport": ">=0.2.0"
}
},
"node_modules/qrcode.react": {
"version": "4.2.0",
"resolved": "https://registry.npmjs.org/qrcode.react/-/qrcode.react-4.2.0.tgz",
"integrity": "sha512-QpgqWi8rD9DsS9EP3z7BT+5lY5SFhsqGjpgW5DY/i3mK4M9DTBNz3ErMi8BWYEfI3L0d8GIbGmcdFAS1uIRGjA==",
"license": "ISC",
"peerDependencies": {
"react": "^16.8.0 || ^17.0.0 || ^18.0.0 || ^19.0.0"
}
},
"node_modules/qs": {
"version": "6.14.2",
"resolved": "https://registry.npmjs.org/qs/-/qs-6.14.2.tgz",
@@ -21325,24 +21336,6 @@
}
}
},
"node_modules/tailwindcss/node_modules/yaml": {
"version": "2.8.3",
"resolved": "https://registry.npmjs.org/yaml/-/yaml-2.8.3.tgz",
"integrity": "sha512-AvbaCLOO2Otw/lW5bmh9d/WEdcDFdQp2Z2ZUH3pX9U2ihyUY0nvLv7J6TrWowklRGPYbB/IuIMfYgxaCPg5Bpg==",
"dev": true,
"license": "ISC",
"optional": true,
"peer": true,
"bin": {
"yaml": "bin.mjs"
},
"engines": {
"node": ">= 14.6"
},
"funding": {
"url": "https://github.com/sponsors/eemeli"
}
},
"node_modules/tapable": {
"version": "2.3.2",
"resolved": "https://registry.npmjs.org/tapable/-/tapable-2.3.2.tgz",
+4 -2
View File
@@ -1,7 +1,8 @@
{
"name": "haproxy-openmanager-frontend",
"version": "1.5.0",
"version": "1.6.0",
"description": "HAProxy Load Balancer Management UI",
"license": "AGPL-3.0-or-later",
"dependencies": {
"react": "^18.2.0",
"react-dom": "^18.2.0",
@@ -16,7 +17,8 @@
"@monaco-editor/react": "^4.6.0",
"react-ace": "^10.1.0",
"ace-builds": "^1.23.4",
"react-window": "^1.8.10"
"react-window": "^1.8.10",
"qrcode.react": "^4.0.0"
},
"devDependencies": {
"@types/react": "^18.0.0",
+226 -25
View File
@@ -92,7 +92,15 @@ const ACMEAutomation = () => {
const [diagEvents, setDiagEvents] = useState([]);
const [diagLoading, setDiagLoading] = useState(false);
const [diagRunningCheckId, setDiagRunningCheckId] = useState(null);
// Bulgu #94 (Round-25 audit): the diagnostic panel's job is to make
// failures visible. We now keep dedicated error envelopes for the
// diagnostics POST and the /events GET so the modal can render them
// inline instead of silently dropping them like the v1.5.1 build did.
const [diagRunError, setDiagRunError] = useState(null);
const [diagEventsError, setDiagEventsError] = useState(null);
const [diagMeta, setDiagMeta] = useState(null);
const diagPollRef = useRef(null);
const diagPollFailCountRef = useRef(0);
const fetchData = useCallback(async () => {
setLoading(true);
@@ -231,6 +239,12 @@ const ACMEAutomation = () => {
};
// v1.5.0 Issue #13: open diagnostics modal for an order
// Bulgu #94 (Round-25 audit): the modal is the operator's last line
// of defence when ACME goes sideways — it MUST render failure causes
// verbatim instead of swallowing them. We capture both the
// diagnostics POST and the events GET errors into dedicated state
// and stop the auto-tail poll after consecutive failures so the
// network tab does not get spammed with /events 500s every 5s.
const handleDiagnose = async (orderId) => {
setDiagOrderId(orderId);
setDiagVisible(true);
@@ -238,6 +252,10 @@ const ACMEAutomation = () => {
setDiagChecks([]);
setDiagHumanizedError(null);
setDiagEvents([]);
setDiagRunError(null);
setDiagEventsError(null);
setDiagMeta(null);
diagPollFailCountRef.current = 0;
try {
const [orderRes, diagRes] = await Promise.allSettled([
axios.get(`/api/letsencrypt/orders/${orderId}`),
@@ -245,18 +263,68 @@ const ACMEAutomation = () => {
]);
if (orderRes.status === 'fulfilled') setDiagOrder(orderRes.value.data);
if (diagRes.status === 'fulfilled') {
setDiagChecks(diagRes.value.data?.checks || []);
setDiagHumanizedError(diagRes.value.data?.humanized_error || null);
const data = diagRes.value.data || {};
setDiagChecks(data.checks || []);
setDiagHumanizedError(data.humanized_error || null);
setDiagMeta(data.meta || null);
// Bulgu #94 follow-up: the Round-25 backend now returns HTTP 200
// with `status: 'diagnostics_unavailable'` + `meta.error_stage`
// instead of HTTP 500 when the diagnostic runner itself crashes
// (e.g. `column a.last_heartbeat does not exist` on a stale
// deploy). The fulfilled branch must therefore detect the
// envelope and surface the structured failure alert — otherwise
// the operator would only see one `diagnostics_runner` row in
// the table without the prominent red banner that explains the
// crash + correlation_id. This is exactly the "panel renders
// but doesn't visibly say WHY" trap we're trying to avoid.
if (data.status === 'diagnostics_unavailable' || data?.meta?.error_stage) {
setDiagRunError({
status: 200,
message: data?.meta?.error_message
|| data?.checks?.[0]?.message
|| 'Diagnostic runner failed',
correlation_id: data?.meta?.correlation_id || null,
error_stage: data?.meta?.error_stage || null,
error_type: data?.meta?.error_type || null,
});
}
} else if (diagRes.status === 'rejected') {
const detail = diagRes.reason?.response?.data?.detail || 'Diagnostics failed';
const resp = diagRes.reason?.response;
const detail = resp?.data?.detail
|| resp?.data?.error?.message
|| diagRes.reason?.message
|| 'Diagnostics failed';
setDiagRunError({
status: resp?.status || 0,
message: detail,
correlation_id: resp?.data?.error?.correlation_id || resp?.headers?.['x-correlation-id'] || null,
});
message.error(detail);
}
// Best-effort merged event log fetch (404 if migration not yet run)
// Best-effort merged event log fetch — bug #95: server-side schema
// drift used to return 500; we now surface that in the modal so
// the operator sees "events unavailable because <reason>".
try {
const evRes = await axios.get(`/api/letsencrypt/orders/${orderId}/events`);
setDiagEvents(evRes.data?.events || []);
} catch (_evErr) {
// ignore
if (evRes.data?.meta?.errors?.length) {
setDiagEventsError({
kind: 'partial',
errors: evRes.data.meta.errors,
correlation_id: evRes.data.meta.correlation_id,
});
}
} catch (evErr) {
const resp = evErr?.response;
setDiagEventsError({
kind: 'fatal',
status: resp?.status || 0,
message: resp?.data?.detail
|| resp?.data?.error?.message
|| evErr?.message
|| 'Event log unavailable',
correlation_id: resp?.data?.error?.correlation_id || null,
});
}
} finally {
setDiagLoading(false);
@@ -299,8 +367,38 @@ const ACMEAutomation = () => {
try {
const evRes = await axios.get(`/api/letsencrypt/orders/${diagOrderId}/events`);
setDiagEvents(evRes.data?.events || []);
} catch (_e) {
/* ignore */
diagPollFailCountRef.current = 0;
if (evRes.data?.meta?.errors?.length) {
setDiagEventsError({
kind: 'partial',
errors: evRes.data.meta.errors,
correlation_id: evRes.data.meta.correlation_id,
});
} else {
setDiagEventsError(null);
}
} catch (e) {
// Bulgu #94 (Round-25): stop the auto-tail after 3 consecutive
// failures so a broken backend doesn't drown the user's
// network tab in 500s. Operator can re-open the modal to retry.
diagPollFailCountRef.current += 1;
if (diagPollFailCountRef.current >= 3) {
if (diagPollRef.current) {
clearInterval(diagPollRef.current);
diagPollRef.current = null;
}
const resp = e?.response;
setDiagEventsError({
kind: 'fatal',
status: resp?.status || 0,
message: resp?.data?.detail
|| resp?.data?.error?.message
|| e?.message
|| 'Event log polling stopped after repeated failures',
correlation_id: resp?.data?.error?.correlation_id || null,
polling_stopped: true,
});
}
}
}, 5000);
return () => {
@@ -318,6 +416,10 @@ const ACMEAutomation = () => {
setDiagChecks([]);
setDiagHumanizedError(null);
setDiagEvents([]);
setDiagRunError(null);
setDiagEventsError(null);
setDiagMeta(null);
diagPollFailCountRef.current = 0;
if (diagPollRef.current) {
clearInterval(diagPollRef.current);
diagPollRef.current = null;
@@ -1100,8 +1202,50 @@ const ACMEAutomation = () => {
label: 'Pre-flight Checks',
children: (
<div>
{/* Bulgu #94 (Round-25): expose backend failures so the
operator can act on them — not just see a blank panel. */}
{diagRunError && (
<Alert
type="error"
showIcon
style={{ marginBottom: 12 }}
message={`Diagnostic runner failed${diagRunError.status && diagRunError.status !== 200 ? ` (HTTP ${diagRunError.status})` : ''}`}
description={
<div>
<div>{diagRunError.message}</div>
{(diagRunError.error_stage || diagRunError.error_type) && (
<div style={{ marginTop: 6, fontSize: 12, color: '#666' }}>
{diagRunError.error_stage && <span>Stage: <code>{diagRunError.error_stage}</code> </span>}
{diagRunError.error_type && <span>· Type: <code>{diagRunError.error_type}</code></span>}
</div>
)}
{diagRunError.correlation_id && (
<div style={{ marginTop: 6, fontSize: 12, color: '#666' }}>
Correlation ID: <code>{diagRunError.correlation_id}</code>
</div>
)}
<div style={{ marginTop: 6, fontSize: 12, color: '#666' }}>
Share this correlation ID with the platform team; the full traceback is in the API log.
</div>
</div>
}
/>
)}
{diagMeta?.correlation_id && (diagMeta.checks_failed > 0 || diagMeta.checks_warn > 0) && !diagRunError && (
<Alert
type={diagMeta.checks_failed > 0 ? 'warning' : 'info'}
showIcon
style={{ marginBottom: 12 }}
message={`${diagMeta.checks_failed} check(s) failed, ${diagMeta.checks_warn} warning(s)`}
description={
<span style={{ fontSize: 12, color: '#666' }}>
Correlation ID: <code>{diagMeta.correlation_id}</code>
</span>
}
/>
)}
{diagChecks.length === 0 ? (
<Empty description="No diagnostic checks available" />
<Empty description={diagRunError ? 'Diagnostic runner did not return any check' : 'No diagnostic checks available'} />
) : (
<Table
size="small"
@@ -1148,24 +1292,81 @@ const ACMEAutomation = () => {
key: 'events',
label: `Event Log (${diagEvents.length})`,
children: (
diagEvents.length === 0 ? (
<Empty description="No events recorded for this order" />
) : (
<Timeline
items={diagEvents.map((ev) => ({
color:
(ev.severity || '').toUpperCase() === 'ERROR' ? 'red' :
(ev.severity || '').toUpperCase() === 'WARN' ? 'orange' : 'blue',
children: (
<div>
{/* Bulgu #95 (Round-25): /events used to 500 because of
a schema-drift bug (`status` column did not exist).
Now the response carries `meta.errors[]` for partial
failures and a `kind: fatal` envelope for total
failure — both render here so the operator never
wonders why the timeline is empty. */}
{diagEventsError?.kind === 'fatal' && (
<Alert
type="error"
showIcon
style={{ marginBottom: 12 }}
message={`Event log unavailable${diagEventsError.status ? ` (HTTP ${diagEventsError.status})` : ''}`}
description={
<div>
<div style={{ fontSize: 12, color: '#888' }}>{ev.created_at} · {ev.source}</div>
<div><strong>{ev.event_type}</strong></div>
{ev.message && <div>{ev.message}</div>}
<div>{diagEventsError.message}</div>
{diagEventsError.correlation_id && (
<div style={{ marginTop: 6, fontSize: 12, color: '#666' }}>
Correlation ID: <code>{diagEventsError.correlation_id}</code>
</div>
)}
{diagEventsError.polling_stopped && (
<div style={{ marginTop: 6, fontSize: 12, color: '#666' }}>
Auto-refresh stopped after repeated failures. Re-open the panel to retry.
</div>
)}
</div>
),
}))}
/>
)
}
/>
)}
{diagEventsError?.kind === 'partial' && (
<Alert
type="warning"
showIcon
style={{ marginBottom: 12 }}
message="Event log partial — one or more sources failed"
description={
<div>
<ul style={{ margin: '4px 0 4px 16px' }}>
{diagEventsError.errors.map((err, i) => (
<li key={i}>
<strong>{err.section}</strong>: {err.exception_type} — {err.message}
</li>
))}
</ul>
{diagEventsError.correlation_id && (
<div style={{ fontSize: 12, color: '#666' }}>
Correlation ID: <code>{diagEventsError.correlation_id}</code>
</div>
)}
</div>
}
/>
)}
{diagEvents.length === 0 ? (
<Empty description={diagEventsError?.kind === 'fatal'
? 'No events could be loaded (see error above)'
: 'No events recorded for this order'} />
) : (
<Timeline
items={diagEvents.map((ev) => ({
color:
(ev.severity || '').toUpperCase() === 'ERROR' ? 'red' :
(ev.severity || '').toUpperCase() === 'WARN' ? 'orange' : 'blue',
children: (
<div>
<div style={{ fontSize: 12, color: '#888' }}>{ev.created_at} · {ev.source}</div>
<div><strong>{ev.event_type}</strong></div>
{ev.message && <div>{ev.message}</div>}
</div>
),
}))}
/>
)}
</div>
),
},
{
+79 -5
View File
@@ -1015,12 +1015,56 @@ const FrontendManagement = () => {
return;
}
if (grandfatheredContradictions.length > 0) {
// Bulgu #83 (round-23 audit) — surface the actual offending rule
// string(s) instead of just a count. Pre-fix the warning said
// "1 legacy rule has X !X" and the operator had to hunt
// through the ACL Builder cards to figure out which rule the
// gate was complaining about. The unchanged-rule path is the
// common case (operator changes port / maxconn on a frontend
// that already had a self-contradictory routing rule from a
// prior session), so making the rule discoverable from the
// toast keeps "Edit and Save" → "fix the dead rule" workflows
// single-screen. Also stop calling these rules "legacy" —
// the operator may have written them seconds earlier; the
// only thing this branch knows is that they weren't modified
// by the current edit.
const renderGrandfatheredRule = (r) => {
if (typeof r === 'string') return r;
if (r && typeof r === 'object') {
try { return JSON.stringify(r); } catch (_e) { return '[rule]'; }
}
return '[rule]';
};
const ruleSnippets = grandfatheredContradictions
.slice(0, 5)
.map(renderGrandfatheredRule)
.map((s) => (s.length > 160 ? `${s.slice(0, 157)}...` : s));
const extra = grandfatheredContradictions.length > ruleSnippets.length
? ` (+${grandfatheredContradictions.length - ruleSnippets.length} more)`
: '';
message.warning(
`This frontend has ${grandfatheredContradictions.length} legacy ` +
`routing/redirect rule(s) with a self-contradictory \`X !X\` ` +
`condition. The rule(s) never fire — fix them at your convenience. ` +
`Your current edit will still be saved.`,
6,
<div>
<div>
<strong>
{grandfatheredContradictions.length} routing/redirect rule(s)
you didn't modify in this edit contain a self-contradictory
`X !X` condition (e.g. `if acl1 !acl1`).
</strong>
</div>
<div style={{ marginTop: 4, fontSize: '12px' }}>
HAProxy accepts the syntax but `X AND NOT X` is always false,
so the rule never fires and traffic silently falls through to
`default_backend`. Your current edit will still be saved; fix
the rule(s) at your convenience.
</div>
<div style={{ marginTop: 6, fontSize: '12px', fontFamily: 'monospace' }}>
{ruleSnippets.map((s, i) => (
<div key={i}>• {s}</div>
))}
{extra && <div>{extra}</div>}
</div>
</div>,
10,
);
}
@@ -1081,6 +1125,36 @@ const FrontendManagement = () => {
} else {
message.success('Frontend updated successfully');
}
// Bulgu #83 (round-23 audit) — surface server-emitted
// grandfathered-rule warnings (e.g. `X !X` contradictions
// in routing/redirect rules that the operator did not
// touch this edit). The FE client-side gate ALSO catches
// these and fires its own toast above the modal close;
// we re-surface the server view here as a safety net in
// case the client gate missed an edge shape (different
// dict serialization, etc.). Server warnings already
// include the verbatim rule text, so the operator sees
// exactly which entry to fix.
const serverWarnings = Array.isArray(response.data?.warnings)
? response.data.warnings
: [];
if (serverWarnings.length > 0 && grandfatheredContradictions.length === 0) {
message.warning(
<div>
<div><strong>Frontend saved, but the server flagged {serverWarnings.length} rule warning(s):</strong></div>
<div style={{ marginTop: 6, fontSize: '12px', fontFamily: 'monospace' }}>
{serverWarnings.slice(0, 5).map((w, i) => (
<div key={i}>• {w.length > 240 ? `${w.slice(0, 237)}...` : w}</div>
))}
{serverWarnings.length > 5 && (
<div>(+{serverWarnings.length - 5} more)</div>
)}
</div>
</div>,
10,
);
}
} else {
response = await axios.post('/api/frontends', requestData);
+244 -104
View File
@@ -1,20 +1,20 @@
import React, { useState } from 'react';
import {
Card,
Form,
Input,
Button,
message,
Typography,
Row,
import React, { useEffect, useRef, useState } from 'react';
import {
Card,
Form,
Input,
Button,
message,
Typography,
Row,
Col,
Alert,
Spin
} from 'antd';
import {
UserOutlined,
LockOutlined,
ClusterOutlined
import {
UserOutlined,
LockOutlined,
ClusterOutlined,
SafetyCertificateOutlined,
} from '@ant-design/icons';
import axios from 'axios';
import { useAuth } from '../contexts/AuthContext';
@@ -23,77 +23,260 @@ import './Login.css';
const { Title, Text } = Typography;
const PHASE_CREDENTIALS = 'credentials';
const PHASE_MFA = 'mfa';
const PHASE_SUBMITTING = 'submitting';
const Login = () => {
const [form] = Form.useForm();
const [credentialsForm] = Form.useForm();
const [mfaForm] = Form.useForm();
const [phase, setPhase] = useState(PHASE_CREDENTIALS);
const [loading, setLoading] = useState(false);
const [error, setError] = useState('');
const { login } = useAuth();
const handleSubmit = async (values) => {
// MFA-specific state — RAM only, never persisted.
const mfaTokenRef = useRef(null);
const [mfaExpiresAt, setMfaExpiresAt] = useState(null);
const [mfaCountdown, setMfaCountdown] = useState(0);
useEffect(() => {
if (phase !== PHASE_MFA || !mfaExpiresAt) return undefined;
const id = setInterval(() => {
const remaining = Math.max(0, Math.floor((mfaExpiresAt - Date.now()) / 1000));
setMfaCountdown(remaining);
if (remaining <= 0) {
clearInterval(id);
resetToCredentials('MFA session expired. Please log in again.');
}
}, 1000);
return () => clearInterval(id);
// eslint-disable-next-line react-hooks/exhaustive-deps
}, [phase, mfaExpiresAt]);
const resetToCredentials = (errMessage) => {
mfaTokenRef.current = null;
setMfaExpiresAt(null);
setMfaCountdown(0);
mfaForm.resetFields();
setPhase(PHASE_CREDENTIALS);
if (errMessage) setError(errMessage);
};
const completeAuth = (authData) => {
// Write storage + axios header ONLY after a full, MFA-cleared response.
localStorage.setItem('token', authData.access_token);
localStorage.setItem('authToken', authData.access_token);
localStorage.setItem('userData', JSON.stringify(authData.user));
localStorage.setItem('userRoles', JSON.stringify([]));
localStorage.setItem('userPermissions', JSON.stringify({}));
const expiryDate = new Date();
expiryDate.setSeconds(expiryDate.getSeconds() + authData.expires_in);
localStorage.setItem('tokenExpiry', expiryDate.toISOString());
const loginSuccess = login(authData);
if (loginSuccess) {
message.success(`Welcome back, ${authData.user.username}!`);
} else {
throw new Error('Failed to update authentication state');
}
};
const handleCredentialsSubmit = async (values) => {
setLoading(true);
setPhase(PHASE_SUBMITTING);
setError('');
try {
const response = await axios.post('/api/auth/login', {
username: values.username,
password: values.password
password: values.password,
});
// Store authentication data - API returns access_token, not session_token!
localStorage.setItem('token', response.data.access_token);
localStorage.setItem('authToken', response.data.access_token);
localStorage.setItem('userData', JSON.stringify(response.data.user));
localStorage.setItem('userRoles', JSON.stringify([])); // API doesn't return roles directly
localStorage.setItem('userPermissions', JSON.stringify({})); // API doesn't return permissions directly
// Calculate expiry from expires_in (seconds)
const expiryDate = new Date();
expiryDate.setSeconds(expiryDate.getSeconds() + response.data.expires_in);
localStorage.setItem('tokenExpiry', expiryDate.toISOString());
// Update authentication context
const loginSuccess = login(response.data);
if (loginSuccess) {
message.success(`Welcome back, ${response.data.user.username}!`);
// The authentication context will automatically trigger a re-render
// and the user will be redirected to the main dashboard
} else {
throw new Error('Failed to update authentication state');
if (response.data && response.data.mfa_required) {
// Phase 2 — TOTP / backup code challenge. Keep credentials secret-free.
mfaTokenRef.current = response.data.mfa_token;
const ttlSeconds = response.data.expires_in || 300;
setMfaExpiresAt(Date.now() + ttlSeconds * 1000);
setMfaCountdown(ttlSeconds);
setPhase(PHASE_MFA);
return;
}
} catch (error) {
const errorMessage = extractApiError(error, 'Login failed. Please try again.');
completeAuth(response.data);
} catch (err) {
const errorMessage = extractApiError(err, 'Login failed. Please try again.');
setError(errorMessage);
message.error(errorMessage);
setPhase(PHASE_CREDENTIALS);
} finally {
setLoading(false);
}
};
const handleMfaSubmit = async (values) => {
if (!mfaTokenRef.current) {
resetToCredentials('MFA session lost. Please log in again.');
return;
}
setLoading(true);
setError('');
try {
const response = await axios.post(
'/api/auth/login/mfa-verify',
{
mfa_token: mfaTokenRef.current,
code: (values.code || '').trim(),
},
// Explicit opt-out: never attach a stale Authorization header here.
{ headers: { Authorization: undefined } },
);
completeAuth(response.data);
} catch (err) {
const status = err && err.response && err.response.status;
const errorMessage = extractApiError(err, 'Verification failed.');
if (status === 410) {
resetToCredentials(errorMessage || 'MFA challenge invalidated. Please log in again.');
} else {
setError(errorMessage);
mfaForm.setFieldsValue({ code: '' });
}
} finally {
setLoading(false);
}
};
const renderCredentialsForm = () => (
<Form
form={credentialsForm}
name="login"
onFinish={handleCredentialsSubmit}
layout="vertical"
autoComplete="off"
>
<Form.Item
name="username"
rules={[
{ required: true, message: 'Please enter your username!' },
{ min: 3, message: 'Username must be at least 3 characters!' },
]}
>
<Input prefix={<UserOutlined />} placeholder="Username" autoComplete="username" />
</Form.Item>
<Form.Item
name="password"
rules={[
{ required: true, message: 'Please enter your password!' },
{ min: 6, message: 'Password must be at least 6 characters!' },
]}
>
<Input.Password
prefix={<LockOutlined />}
placeholder="Password"
autoComplete="current-password"
/>
</Form.Item>
<Form.Item style={{ marginBottom: 0 }}>
<Button
type="primary"
htmlType="submit"
loading={loading}
block
className="login-button"
>
{loading ? 'Signing in...' : 'Sign In'}
</Button>
</Form.Item>
</Form>
);
const renderMfaForm = () => {
const minutes = Math.floor(mfaCountdown / 60);
const seconds = String(mfaCountdown % 60).padStart(2, '0');
return (
<Form form={mfaForm} name="mfa" onFinish={handleMfaSubmit} layout="vertical" autoComplete="off">
<Alert
message="Multi-Factor Authentication"
description={
<span>
Enter the 6-digit code from your authenticator app, or use a backup code
(format: <code>XXXX-YYYY</code>).
{mfaCountdown > 0 && (
<>
{' '}Session expires in <strong>{minutes}:{seconds}</strong>.
</>
)}
</span>
}
type="info"
showIcon
style={{ marginBottom: 16 }}
/>
<Form.Item
name="code"
rules={[
{ required: true, message: 'Please enter your MFA code.' },
{ min: 6, message: 'Code must be at least 6 characters.' },
{ max: 10, message: 'Code is too long.' },
]}
>
<Input
prefix={<SafetyCertificateOutlined />}
placeholder="123456 or XXXX-YYYY"
autoComplete="one-time-code"
inputMode="text"
maxLength={10}
autoFocus
/>
</Form.Item>
<Form.Item style={{ marginBottom: 8 }}>
<Button
type="primary"
htmlType="submit"
loading={loading}
block
className="login-button"
>
{loading ? 'Verifying...' : 'Verify'}
</Button>
</Form.Item>
<Form.Item style={{ marginBottom: 0 }}>
<Button
type="default"
block
onClick={() => resetToCredentials('')}
disabled={loading}
>
Use a different account
</Button>
</Form.Item>
</Form>
);
};
return (
<div className="login-container">
<Row
justify="center"
align="middle"
style={{
<Row
justify="center"
align="middle"
style={{
minHeight: '100vh',
minHeight: '100dvh',
width: '100%',
margin: 0
margin: 0,
}}
>
<Col
xs={24}
sm={20}
md={16}
lg={12}
xl={10}
<Col
xs={24}
sm={20}
md={16}
lg={12}
xl={10}
xxl={8}
style={{
style={{
display: 'flex',
justifyContent: 'center',
padding: '0 8px'
padding: '0 8px',
}}
>
<Card className="login-card">
@@ -118,56 +301,13 @@ const Login = () => {
/>
)}
<Form
form={form}
name="login"
onFinish={handleSubmit}
layout="vertical"
autoComplete="off"
>
<Form.Item
name="username"
rules={[
{ required: true, message: 'Please enter your username!' },
{ min: 3, message: 'Username must be at least 3 characters!' }
]}
>
<Input
prefix={<UserOutlined />}
placeholder="Username"
autoComplete="username"
/>
</Form.Item>
<Form.Item
name="password"
rules={[
{ required: true, message: 'Please enter your password!' },
{ min: 6, message: 'Password must be at least 6 characters!' }
]}
>
<Input.Password
prefix={<LockOutlined />}
placeholder="Password"
autoComplete="current-password"
/>
</Form.Item>
<Form.Item style={{ marginBottom: 0 }}>
<Button
type="primary"
htmlType="submit"
loading={loading}
block
className="login-button"
>
{loading ? 'Signing in...' : 'Sign In'}
</Button>
</Form.Item>
</Form>
{phase === PHASE_MFA ? renderMfaForm() : renderCredentialsForm()}
<div className="login-footer">
<Text type="secondary" style={{ fontSize: 12, display: 'block', textAlign: 'center' }}>
<Text
type="secondary"
style={{ fontSize: 12, display: 'block', textAlign: 'center' }}
>
Centralized management for multiple HAProxy clusters
</Text>
</div>
@@ -178,4 +318,4 @@ const Login = () => {
);
};
export default Login;
export default Login;
+306
View File
@@ -0,0 +1,306 @@
import React, { useEffect, useRef, useState } from 'react';
import {
Modal,
Steps,
Form,
Input,
Button,
Alert,
Typography,
Space,
Checkbox,
message,
} from 'antd';
import { QRCodeSVG } from 'qrcode.react';
import axios from 'axios';
import { extractApiError } from '../utils/apiError';
const { Text, Paragraph } = Typography;
const STEP_SETUP = 0;
const STEP_VERIFY = 1;
const STEP_BACKUP = 2;
/**
* MFA enrollment wizard. Strictly modal-controlled: the modal cannot be
* dismissed via the X / mask in step 2/3 — backup codes are shown only once
* and the server-side pending row is opaque after enrollment confirms.
*/
const MFAEnrollModal = ({ open, onClose, onEnrolled }) => {
const [verifyForm] = Form.useForm();
const [step, setStep] = useState(STEP_SETUP);
const [loading, setLoading] = useState(false);
const [error, setError] = useState('');
const [otpauthUri, setOtpauthUri] = useState('');
const [secret, setSecret] = useState('');
const [backupCodes, setBackupCodes] = useState([]);
const [savedAcknowledged, setSavedAcknowledged] = useState(false);
const startedRef = useRef(false);
useEffect(() => {
if (!open) return undefined;
if (startedRef.current) return undefined;
startedRef.current = true;
startEnrollment();
return () => {
// No-op: cleanup happens via the explicit handleClose path.
};
// eslint-disable-next-line react-hooks/exhaustive-deps
}, [open]);
const resetState = () => {
setStep(STEP_SETUP);
setLoading(false);
setError('');
setOtpauthUri('');
setSecret('');
setBackupCodes([]);
setSavedAcknowledged(false);
startedRef.current = false;
verifyForm.resetFields();
};
const startEnrollment = async () => {
setLoading(true);
setError('');
try {
const response = await axios.post('/api/mfa/enroll/start', {});
setOtpauthUri(response.data.otpauth_uri);
setSecret(response.data.secret);
} catch (err) {
setError(extractApiError(err, 'Could not start MFA enrollment.'));
} finally {
setLoading(false);
}
};
const handleVerify = async (values) => {
setLoading(true);
setError('');
try {
const response = await axios.post('/api/mfa/enroll/confirm', {
code: (values.code || '').trim(),
});
setBackupCodes(response.data.backup_codes || []);
setStep(STEP_BACKUP);
} catch (err) {
setError(extractApiError(err, 'Verification failed.'));
verifyForm.setFieldsValue({ code: '' });
} finally {
setLoading(false);
}
};
const handleCopyAll = () => {
const text = backupCodes.join('\n');
if (navigator.clipboard && navigator.clipboard.writeText) {
navigator.clipboard.writeText(text).then(
() => message.success('Backup codes copied to clipboard'),
() => message.error('Could not copy. Please copy manually.'),
);
} else {
message.warning('Clipboard API unavailable. Please copy manually.');
}
};
const handleDownload = () => {
const blob = new Blob(
[
'HAProxy OpenManager — MFA backup codes\n',
'Generated: ' + new Date().toISOString() + '\n',
'Each code is single-use. Store them somewhere safe and offline.\n\n',
...backupCodes.map((c) => c + '\n'),
],
{ type: 'text/plain;charset=utf-8' },
);
const url = URL.createObjectURL(blob);
const link = document.createElement('a');
link.href = url;
link.download = 'haproxy-openmanager-mfa-backup-codes.txt';
document.body.appendChild(link);
link.click();
document.body.removeChild(link);
URL.revokeObjectURL(url);
};
const handleClose = (force = false) => {
if (step === STEP_BACKUP && !savedAcknowledged && !force) return;
resetState();
if (step === STEP_BACKUP) {
if (typeof onEnrolled === 'function') onEnrolled();
} else if (typeof onClose === 'function') {
onClose();
}
};
const renderSetup = () => (
<Space direction="vertical" size="middle" style={{ width: '100%' }}>
<Paragraph>
Open your authenticator app (Google Authenticator, Authy, 1Password, Microsoft
Authenticator) and scan this QR code, or enter the secret manually.
</Paragraph>
<div style={{ display: 'flex', justifyContent: 'center' }}>
{otpauthUri ? (
<QRCodeSVG value={otpauthUri} size={220} level="M" includeMargin />
) : (
<Text type="secondary">Generating…</Text>
)}
</div>
{secret && (
<Alert
message="Trouble scanning?"
description={
<Space direction="vertical" size={4}>
<Text>Enter this secret manually in your authenticator app:</Text>
<Text code copyable={{ text: secret }} style={{ fontSize: 16 }}>
{secret}
</Text>
</Space>
}
type="info"
showIcon
/>
)}
<div style={{ textAlign: 'right' }}>
<Space>
<Button onClick={() => handleClose(true)}>Cancel</Button>
<Button
type="primary"
onClick={() => setStep(STEP_VERIFY)}
disabled={!otpauthUri}
>
I&apos;ve added the account
</Button>
</Space>
</div>
</Space>
);
const renderVerify = () => (
<Form form={verifyForm} layout="vertical" onFinish={handleVerify}>
<Alert
message="Verify your authenticator"
description="Enter the 6-digit code displayed by your authenticator app. You have 5 attempts before the enrollment is invalidated and you'll need to start over."
type="info"
showIcon
style={{ marginBottom: 16 }}
/>
<Form.Item
name="code"
label="Authenticator code"
rules={[
{ required: true, message: 'Please enter the 6-digit code.' },
{ len: 6, message: 'Code must be exactly 6 digits.' },
]}
>
<Input
placeholder="123456"
autoComplete="one-time-code"
inputMode="numeric"
maxLength={6}
autoFocus
/>
</Form.Item>
<div style={{ textAlign: 'right' }}>
<Space>
<Button onClick={() => setStep(STEP_SETUP)} disabled={loading}>
Back
</Button>
<Button onClick={() => handleClose(true)} disabled={loading}>
Cancel
</Button>
<Button type="primary" htmlType="submit" loading={loading}>
Verify
</Button>
</Space>
</div>
</Form>
);
const renderBackup = () => (
<Space direction="vertical" size="middle" style={{ width: '100%' }}>
<Alert
message="Save your backup codes now"
description={
<>
Each code can be used <strong>once</strong> when you can&apos;t access your
authenticator. <strong>They won&apos;t be shown again.</strong> If you lose
them, ask an administrator to reset your MFA.
</>
}
type="warning"
showIcon
/>
<div
style={{
display: 'grid',
gridTemplateColumns: '1fr 1fr',
gap: '8px 16px',
padding: '12px',
backgroundColor: 'var(--ant-color-fill-quaternary, #fafafa)',
borderRadius: 6,
}}
>
{backupCodes.map((code) => (
<Text key={code} code style={{ fontSize: 15, letterSpacing: 1 }}>
{code}
</Text>
))}
</div>
<Space>
<Button onClick={handleCopyAll}>Copy all</Button>
<Button onClick={handleDownload}>Download .txt</Button>
</Space>
<Checkbox
checked={savedAcknowledged}
onChange={(e) => setSavedAcknowledged(e.target.checked)}
>
I have saved my backup codes somewhere safe.
</Checkbox>
<div style={{ textAlign: 'right' }}>
<Button
type="primary"
disabled={!savedAcknowledged}
onClick={() => handleClose(false)}
>
Close
</Button>
</div>
</Space>
);
return (
<Modal
open={open}
title="Enable Multi-Factor Authentication"
width={520}
footer={null}
closable={false}
maskClosable={false}
destroyOnClose
keyboard={false}
>
<Steps
size="small"
current={step}
items={[{ title: 'Set up' }, { title: 'Verify' }, { title: 'Backup codes' }]}
style={{ marginBottom: 24 }}
/>
{error && (
<Alert
message={error}
type="error"
showIcon
closable
onClose={() => setError('')}
style={{ marginBottom: 16 }}
/>
)}
{step === STEP_SETUP && renderSetup()}
{step === STEP_VERIFY && renderVerify()}
{step === STEP_BACKUP && renderBackup()}
</Modal>
);
};
export default MFAEnrollModal;
+243 -33
View File
@@ -32,11 +32,15 @@ import {
HistoryOutlined,
UserAddOutlined,
KeyOutlined,
DownloadOutlined
DownloadOutlined,
SafetyCertificateOutlined,
ReloadOutlined,
InfoCircleOutlined,
} from '@ant-design/icons';
import axios from 'axios';
import { useAuth } from '../contexts/AuthContext';
import { extractApiError } from '../utils/apiError';
import MFAEnrollModal from './MFAEnrollModal';
const { TabPane } = Tabs;
const { Option } = Select;
@@ -225,8 +229,16 @@ const PERMISSION_TREE = [
];
const UserManagement = () => {
const { isAdmin } = useAuth(); // Get admin status from auth context
const { isAdmin, user: currentUser } = useAuth(); // Get admin status from auth context
const { token } = theme.useToken();
// MFA UI state (Issue #18, v1.6.0)
const [mfaEnrollOpen, setMfaEnrollOpen] = useState(false);
const [mfaDisableOpen, setMfaDisableOpen] = useState(false);
const [mfaDisableForm] = Form.useForm();
const [mfaActionLoading, setMfaActionLoading] = useState(false);
const [mfaResetTarget, setMfaResetTarget] = useState(null);
const [mfaResetForm] = Form.useForm();
const [activeTab, setActiveTab] = useState('users');
// Users state
@@ -525,6 +537,47 @@ const UserManagement = () => {
}
};
// ---- MFA handlers (Issue #18, v1.6.0) -----------------------------------
const handleMfaEnrolled = () => {
setMfaEnrollOpen(false);
message.success('MFA enabled for your account.');
fetchUsers();
};
const handleMfaDisableSubmit = async (values) => {
setMfaActionLoading(true);
try {
await axios.post('/api/mfa/disable', { code: (values.code || '').trim() });
message.success('MFA disabled for your account.');
setMfaDisableOpen(false);
mfaDisableForm.resetFields();
fetchUsers();
} catch (error) {
message.error(extractApiError(error, 'Failed to disable MFA'));
} finally {
setMfaActionLoading(false);
}
};
const handleMfaAdminResetSubmit = async (values) => {
if (!mfaResetTarget) return;
setMfaActionLoading(true);
try {
await axios.post(`/api/mfa/admin-reset/${mfaResetTarget.id}`, {
reason: (values.reason || '').trim(),
});
message.success(`MFA reset for user '${mfaResetTarget.username}'.`);
setMfaResetTarget(null);
mfaResetForm.resetFields();
fetchUsers();
} catch (error) {
message.error(extractApiError(error, 'Failed to reset MFA'));
} finally {
setMfaActionLoading(false);
}
};
const handleChangePassword = (user) => {
setSelectedUser(user);
passwordForm.resetFields();
@@ -717,6 +770,15 @@ const UserManagement = () => {
</Space>
)
},
{
title: 'MFA',
dataIndex: 'mfa_enabled',
key: 'mfa_enabled',
width: 90,
render: (enabled) => enabled
? <Tag icon={<SafetyCertificateOutlined />} color="green">ON</Tag>
: <Tag color="default">OFF</Tag>
},
{
title: 'Last Login',
dataIndex: 'last_login_at',
@@ -726,31 +788,94 @@ const UserManagement = () => {
{
title: 'Actions',
key: 'actions',
render: (_, record) => (
<Space>
{isAdmin() && (
<>
<Tooltip title="Edit User">
<Button
icon={<EditOutlined />}
size="small"
onClick={() => handleEditUser(record)}
render: (_, record) => {
const isSelf = currentUser && record.id === currentUser.id;
return (
<Space>
{isAdmin() && (
<>
<Tooltip title="Edit User">
<Button
icon={<EditOutlined />}
size="small"
onClick={() => handleEditUser(record)}
/>
</Tooltip>
<Tooltip title="Assign Roles">
<Button
icon={<TeamOutlined />}
size="small"
onClick={() => handleAssignRoles(record)}
/>
</Tooltip>
<Tooltip title="Change Password">
<Button
icon={<KeyOutlined />}
size="small"
onClick={() => handleChangePassword(record)}
/>
</Tooltip>
</>
)}
{isSelf && !record.mfa_enabled && (
<Tooltip title="Enable MFA for your account">
<Button
icon={<SafetyCertificateOutlined />}
size="small"
type="primary"
ghost
onClick={() => setMfaEnrollOpen(true)}
>
Enable MFA
</Button>
</Tooltip>
)}
{isSelf && record.mfa_enabled && (
<Tooltip title="Disable MFA for your account">
<Button
icon={<SafetyCertificateOutlined />}
size="small"
onClick={() => {
mfaDisableForm.resetFields();
setMfaDisableOpen(true);
}}
>
Disable MFA
</Button>
</Tooltip>
)}
{isAdmin() && !isSelf && record.mfa_enabled && (
<Tooltip title="Reset this user's MFA (admin only)">
<Button
icon={<ReloadOutlined />}
size="small"
danger
onClick={() => {
mfaResetForm.resetFields();
setMfaResetTarget(record);
}}
/>
</Tooltip>
<Tooltip title="Assign Roles">
<Button
icon={<TeamOutlined />}
size="small"
onClick={() => handleAssignRoles(record)}
/>
</Tooltip>
<Tooltip title="Change Password">
<Button
icon={<KeyOutlined />}
size="small"
onClick={() => handleChangePassword(record)}
)}
{isAdmin() && !isSelf && !record.mfa_enabled && (
<Tooltip
title={
<span>
Only this user can enable their own MFA — the TOTP secret
must be set up from their device. Ask them to sign in and
click <b>Enable MFA</b> in their own row.
</span>
}
>
<Button
icon={<InfoCircleOutlined />}
size="small"
type="text"
style={{ color: '#8c8c8c' }}
/>
</Tooltip>
)}
{isAdmin() && (
<Popconfirm
title="Are you sure you want to delete this user?"
onConfirm={() => handleDeleteUser(record)}
@@ -758,20 +883,18 @@ const UserManagement = () => {
cancelText="No"
>
<Tooltip title="Delete User">
<Button
icon={<DeleteOutlined />}
danger
<Button
icon={<DeleteOutlined />}
danger
size="small"
/>
</Tooltip>
</Popconfirm>
</>
)}
{!isAdmin() && (
<Text type="secondary">View Only</Text>
)}
</Space>
)
)}
{!isAdmin() && !isSelf && <Text type="secondary">View Only</Text>}
</Space>
);
}
}
];
@@ -1512,6 +1635,93 @@ const UserManagement = () => {
</Form.Item>
</Form>
</Modal>
{/* MFA enrollment wizard (self) */}
<MFAEnrollModal
open={mfaEnrollOpen}
onClose={() => setMfaEnrollOpen(false)}
onEnrolled={handleMfaEnrolled}
/>
{/* MFA self-disable modal */}
<Modal
title="Disable Multi-Factor Authentication"
open={mfaDisableOpen}
onCancel={() => setMfaDisableOpen(false)}
footer={null}
destroyOnClose
>
<Form form={mfaDisableForm} layout="vertical" onFinish={handleMfaDisableSubmit}>
<p>
Enter a current code from your authenticator app, or one of your backup
codes (format: <code>XXXX-YYYY</code>), to disable MFA.
</p>
<Form.Item
name="code"
label="MFA code"
rules={[
{ required: true, message: 'Please enter your MFA code.' },
{ min: 6, message: 'Code is too short.' },
{ max: 10, message: 'Code is too long.' },
]}
>
<Input
placeholder="123456 or XXXX-YYYY"
autoComplete="one-time-code"
maxLength={10}
autoFocus
/>
</Form.Item>
<Form.Item style={{ marginBottom: 0, textAlign: 'right' }}>
<Space>
<Button onClick={() => setMfaDisableOpen(false)} disabled={mfaActionLoading}>
Cancel
</Button>
<Button type="primary" danger htmlType="submit" loading={mfaActionLoading}>
Disable MFA
</Button>
</Space>
</Form.Item>
</Form>
</Modal>
{/* Admin: reset another user's MFA */}
<Modal
title={mfaResetTarget ? `Reset MFA for '${mfaResetTarget.username}'` : 'Reset MFA'}
open={!!mfaResetTarget}
onCancel={() => setMfaResetTarget(null)}
footer={null}
destroyOnClose
>
<Form form={mfaResetForm} layout="vertical" onFinish={handleMfaAdminResetSubmit}>
<p>
This will disable MFA and invalidate all backup codes for{' '}
<strong>{mfaResetTarget && mfaResetTarget.username}</strong>. The user must
re-enroll on their next login. The action is recorded in the audit log.
</p>
<Form.Item
name="reason"
label="Reason (required, will be audit-logged)"
rules={[
{ required: true, message: 'Please provide a reason.' },
{ min: 3, message: 'Reason is too short.' },
{ max: 500, message: 'Reason is too long (max 500 characters).' },
]}
>
<Input.TextArea rows={3} placeholder="e.g. Lost phone — verified identity via ticket #12345." />
</Form.Item>
<Form.Item style={{ marginBottom: 0, textAlign: 'right' }}>
<Space>
<Button onClick={() => setMfaResetTarget(null)} disabled={mfaActionLoading}>
Cancel
</Button>
<Button type="primary" danger htmlType="submit" loading={mfaActionLoading}>
Reset MFA
</Button>
</Space>
</Form.Item>
</Form>
</Modal>
</div>
);
};
+31
View File
@@ -0,0 +1,31 @@
/**
* CRA dev-server proxy.
*
* This file is consumed by Webpack DevServer (via `react-scripts start`)
* EXCLUSIVELY. It is NOT included in the production bundle (`npm run
* build` / `serve -s build`) and has zero effect in Kubernetes / Docker
* Compose deployments where nginx (`/api/*` → backend) handles routing.
*
* Default target is `http://localhost:8000` because the only realistic
* use of `npm start` is on a developer's host machine, where the backend
* is reachable on localhost. The target can be overridden with the
* `PROXY_TARGET` environment variable (e.g. `PROXY_TARGET=http://other-host:8000 npm start`).
*
* Why not `http://backend:8000`? That hostname only resolves inside the
* Docker Compose network. Defaulting to it caused all `/api/*` requests
* to fail with ENOTFOUND when running the dev-server on the host.
*/
const { createProxyMiddleware } = require('http-proxy-middleware');
const target = process.env.PROXY_TARGET || 'http://localhost:8000';
module.exports = function (app) {
app.use(
'/api',
createProxyMiddleware({
target,
changeOrigin: true,
logLevel: 'warn',
})
);
};
+15 -19
View File
@@ -1,32 +1,28 @@
/**
* API Configuration
* Centralized API URL management for the application
*
* Environment Variables:
* - REACT_APP_API_URL: Full API URL (e.g., https://api.example.com)
* - NODE_ENV: Environment (development, production)
* Centralized API URL management for the application.
*
* Resolution strategy (same in dev and prod — keeps SPA same-origin):
* 1) REACT_APP_API_URL — explicit override at BUILD time. Use only when
* the SPA must call a cross-origin API (CORS must be enabled there).
* 2) window.location.{protocol,host} — same-origin. In dev this routes
* through CRA dev-server proxy (see frontend/src/setupProxy.js); in
* prod through nginx ingress (`/api/*` → backend service).
* 3) Empty string — non-browser env (SSR/tests). Yields relative URLs.
*
* NOTE: Do NOT hardcode `http://localhost:8000` here. Even in unreachable
* branches CRA/Terser keeps string literals in the bundle, which would
* confuse anyone auditing the production artifact.
*/
// Get API URL from environment or use default based on environment
const getApiUrl = () => {
// Priority 1: Explicit environment variable
if (process.env.REACT_APP_API_URL) {
return process.env.REACT_APP_API_URL;
}
// Priority 2: Detect from window location (for production deployments)
// Use window.location.host (includes port) instead of hostname (excludes port)
// so non-standard ports like :8080 are preserved, preventing CORS issues
if (typeof window !== 'undefined' && window.location) {
const { protocol, host } = window.location;
if (process.env.NODE_ENV === 'production') {
return `${protocol}//${host}`;
}
return `${protocol}//${host}`;
}
// Priority 3: Development default
return 'http://localhost:8000';
return '';
};
// API Base URL
+5 -3
View File
@@ -23,6 +23,8 @@ metadata:
app: haproxy-openmanager
component: backend
type: Opaque
data:
# your-secret-key-change-this-in-production
SECRET_KEY: eW91ci1zZWNyZXQta2V5LWNoYW5nZS10aGlzLWluLXByb2R1Y3Rpb24=
# Both values are replaced at deploy time by the CI/CD pipeline (sed step)
# before `kubectl apply` runs. Never commit real secrets — use the placeholders.
stringData:
SECRET_KEY: secret_key_replace_me
MFA_ENCRYPTION_KEY: mfa_encryption_key_replace_me
+41 -3
View File
@@ -15,6 +15,33 @@ data:
PUBLIC_URL: 'https://haproxy-openmanager.example.com'
MANAGEMENT_BASE_URL: 'https://haproxy-openmanager.example.com'
# ---------------------------------------------------------------------------
# MFA rate-limits (slowapi format: "<count>/<second|minute|hour|day>").
# The shipped defaults assume the rate-limit key is per-USER (Bearer JWT)
# with a per-IP fallback. They are sized for thousands of concurrent
# operators in an org-wide MFA rollout. Pod restart required to apply.
#
# Defaults live in backend/middleware/mfa_rate_limits.py:
# ENROLL_START='10/minute' ENROLL_CONFIRM='10/minute'
# DISABLE='10/minute' REGENERATE_BACKUP_CODES='5/hour'
# ADMIN_RESET='60/hour' ADMIN_RESET_ALL='1/day'
# ---------------------------------------------------------------------------
# MFA_RATE_LIMIT_ENROLL_START: '10/minute'
# MFA_RATE_LIMIT_ENROLL_CONFIRM: '10/minute'
# MFA_RATE_LIMIT_DISABLE: '10/minute'
# MFA_RATE_LIMIT_REGENERATE_BACKUP_CODES: '5/hour'
# MFA_RATE_LIMIT_ADMIN_RESET: '60/hour'
# MFA_RATE_LIMIT_ADMIN_RESET_ALL: '1/day'
# ---------------------------------------------------------------------------
# Trusted reverse-proxy CIDRs for X-Forwarded-For. When the request peer
# falls inside one of these CIDRs and the user is NOT authenticated, the
# first hop of XFF becomes the rate-limit bucket. Empty (default) disables
# XFF parsing entirely — the safe choice when the topology is unknown.
# Example (in-cluster service mesh + nginx ingress on the default ranges):
# ---------------------------------------------------------------------------
# MFA_TRUSTED_PROXY_CIDRS: '10.0.0.0/8,172.16.0.0/12,192.168.0.0/16'
---
apiVersion: v1
kind: ConfigMap
@@ -25,9 +52,20 @@ metadata:
app: haproxy-openmanager
component: frontend
data:
# Leave empty to use same-origin (window.location) in production
# For development, set to: 'http://localhost:8000'
REACT_APP_API_URL: ''
# NOTE — REACT_APP_* env vars are RUNTIME no-ops here.
# Create-React-App inlines REACT_APP_* values into the bundle at BUILD time
# (`npm run build`). The static bundle served by `serve -s build` does not
# consume runtime env vars, so anything set here has zero effect on
# window.location-based same-origin routing.
#
# The frontend always issues `/api/*` against window.location, which is
# routed by the ingress + nginx ConfigMap below to the backend service.
# No env vars are required for the frontend deployment.
#
# WARNING: Do NOT set REACT_APP_API_URL in your CI/CD pipeline either —
# if a value is supplied at build time it gets HARDCODED into the bundle
# and breaks every deployment that does not match that exact URL.
{}
---
apiVersion: v1
+35
View File
@@ -75,6 +75,35 @@ Test HAProxy stats parsing
### test-build.sh
Run build tests for the project
## 🔐 MFA Admin Scripts
### admin-mfa-reset-all.sh
**Purpose:** Emergency — disable Multi-Factor Authentication for **every** user
in one call. Use only when there is a mass loss of authenticator devices /
inherited platform without working operators (Issue #18, v1.6.0).
**Usage:**
```bash
# Interactive prompts ask for the admin Bearer token + reason
./scripts/admin-mfa-reset-all.sh
# Non-interactive (still requires double confirmation typed at the keyboard)
API_URL=https://hap.example.com \
ADMIN_TOKEN=eyJhbGciOi... \
./scripts/admin-mfa-reset-all.sh
```
**Features:**
- Calls `POST /api/mfa/admin-reset-all` (requires `users.is_admin = TRUE`)
- Double confirmation: type `yes`, then `RESET ALL MFA` exactly
- Required reason is recorded in `user_activity_logs`
(`action='mfa.disabled.admin_bulk_reset'`)
- Deletes every backup code and invalidates pending MFA challenges
**Safety:**
- Irreversible — all users must re-enroll MFA afterwards
- All other authentication (password, JWT, roles) is unaffected
## 🚨 Emergency Use Cases
**1. Cluster Migration/Cleanup:**
@@ -101,6 +130,12 @@ Run build tests for the project
./scripts/debug-agent-stats-function.sh
```
**4. MFA Outage (mass lost authenticators):**
```bash
# Disable MFA for every user, then ask them to re-enroll
./scripts/admin-mfa-reset-all.sh
```
## ⚠️ Safety Notes
- All cleanup scripts require admin authentication
+76
View File
@@ -0,0 +1,76 @@
#!/usr/bin/env bash
# scripts/admin-mfa-reset-all.sh
# Issue #18 — v1.6.0 — Emergency: disable MFA for ALL users in one call.
#
# This is a break-glass tool. It calls POST /api/mfa/admin-reset-all on the
# backend with a strict, double-confirmed body and writes the action into the
# server's audit log (action=mfa.disabled.admin_bulk_reset). All users will be
# able to log in with username/password alone afterwards and must re-enroll if
# they want MFA again.
#
# Usage:
# ./scripts/admin-mfa-reset-all.sh
# API_URL=https://hap.example.com ADMIN_TOKEN=ey... ./scripts/admin-mfa-reset-all.sh
#
# Required: a JWT bearer token belonging to a user whose `users.is_admin = TRUE`.
set -euo pipefail
API_URL="${API_URL:-http://localhost:8000}"
ADMIN_TOKEN="${ADMIN_TOKEN:-}"
if [ -z "$ADMIN_TOKEN" ]; then
read -rsp "Admin Bearer token (user with is_admin=TRUE): " ADMIN_TOKEN
echo
fi
if [ -z "$ADMIN_TOKEN" ]; then
echo "ERROR: no admin token provided." >&2
exit 1
fi
cat <<'WARN'
========================================================================
WARNING — IRREVERSIBLE BULK ACTION
========================================================================
This will:
* set mfa_enabled = FALSE for every user
* delete every backup code
* invalidate all pending MFA login challenges and enrollments
* write a permanent entry to user_activity_logs
After this, all users can sign in with username/password only and must
re-enroll MFA from the Users page.
========================================================================
WARN
read -rp "Type 'yes' to proceed: " confirm1
if [ "$confirm1" != "yes" ]; then
echo "Aborted."
exit 1
fi
read -rp "Type 'RESET ALL MFA' (exact) to confirm: " confirm2
if [ "$confirm2" != "RESET ALL MFA" ]; then
echo "Aborted."
exit 1
fi
read -rp "Reason (logged in audit): " reason
if [ -z "$reason" ]; then
echo "Aborted: reason is required."
exit 1
fi
# Compose JSON safely (python3 for proper JSON escaping; reliably available
# everywhere the backend already runs).
payload=$(python3 -c "
import json, sys
print(json.dumps({'confirm': 'RESET ALL MFA', 'reason': sys.argv[1]}))
" "$reason")
echo "Calling $API_URL/api/mfa/admin-reset-all ..."
http_response=$(curl -fsS -X POST "$API_URL/api/mfa/admin-reset-all" \
-H "Authorization: Bearer $ADMIN_TOKEN" \
-H "Content-Type: application/json" \
-d "$payload")
echo "$http_response"
echo
echo "Done. Verify in user_activity_logs: action='mfa.disabled.admin_bulk_reset'."
+3 -3
View File
@@ -1,5 +1,5 @@
{
"version": "1.5.0",
"releaseName": "ACME Diagnostics & Proxied Host Wizard",
"releaseDate": "2026-05-08"
"version": "1.6.0",
"releaseName": "Multi-Factor Authentication (MFA)",
"releaseDate": "2026-05-18"
}