mirror of
https://github.com/taylanbakircioglu/haproxy-openmanager.git
synced 2026-09-23 19:06:25 +00:00
main
9 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
5f5c7f1c75 |
docs(v1.11.0): document what actually ships, with the measurements behind it
The release notes inherited from the feature branch described the version it was written against, not the one going out. - `SCHEMA_VERSION` is 11 -> 12, not 10 -> 11, and the upgrade notes now say why: 11 was taken by v1.10.4 while this was in review, and the version gate would have skipped the migration entirely on every existing install. Includes the no-op recovery path for anyone running a pre-release build that recorded 11. - Successful agent polls are not logged by default, with the measured table behind it: 2 424 bytes/row on PostgreSQL 15 against the real schema and all nine indexes, ~9 792 logged calls/day/agent, and what that means at 20, 200 and 500 nodes both ways. The point is not the disk, it is that the row cap holds by DELETING, so without this the configured 7-day/30-day retention quietly becomes a few hours for everything in the table. - Runtime cost stated as measured numbers rather than adjectives: 27.7 us per request, 1.4 us on an excluded path, 18.8 us per row on the writer, 0.096 % of one core at 500 nodes. - REQUEST_LOG_QUEUE_MAX_BYTES documented in .env.template and CONFIG.md, with the reason it exists: the row count alone does not bound memory when max_body_bytes is operator-editable to 256 KB. - The old "raise REQUEST_LOG_QUEUE_MAX if you see drops" advice is corrected - following it could OOM the worker. Lower max_body_bytes or sample_rate first; if you do raise the queue, raise its byte ceiling with it. - Two behaviours that used to be silent are now written down: sink counters are per worker, and clearing the exclude-path list falls back to the shipped defaults rather than logging everything. - The `operator` role's visibility of agent rows is documented, including what it deliberately does NOT extend to (anonymous traffic and the usernames in failed logins). The v1.10.4 through v1.10.14 notes are unchanged and still above this in both files. |
||
|
|
4c84596215 |
feat(logging): unified request/response log with configurable retention
Applies PR #59 by Mustafa Ulukaya (github.com/taylanbakircioglu/haproxy-openmanager/pull/59,
head
|
||
|
|
bb774141d4 |
fix(acme): make the challenge backend fixable from the panel
Correcting a wrong ACME challenge backend was impossible without a shell, and
even with one the correction did not reach the nodes.
The mint gate only fired when `acme_enabled` flipped. `acme_backend_url` was
written to the DB and minted nothing, so Apply answered "No pending changes to
apply" and the nodes kept the old address forever. It is now decided by
comparing the rendered `server _acme_mgmt` line against the active version —
the one line that answers "would the nodes talk to a different address?".
Comparing whole configs would flag every unrelated pending edit.
The field had no UI at all. Added to the cluster form with validation that
mirrors the backend rules, and keyed on `model_fields_set` so clearing it
reverts to the global setting — with a plain `is not None` test an empty box is
indistinguishable from "not submitted", so a value could never be removed.
Validation is asymmetric on purpose (utils/acme_backend_url):
- at the write boundary, reject what cannot express a reachable target —
including the two silent traps: a scheme-less value became `localhost`, and
an out-of-range port raised inside the generator and destroyed the config
- at render time, never reject. The shipped defaults are themselves loopback,
so refusing to render would make every acme_enabled cluster unappliable,
including for changes unrelated to ACME. Problems are logged and surfaced.
The port-less default stays 8080 rather than moving to HTTP's 80: the bundled
compose publishes nginx on 8080, so installs relying on it work today and the
first sign of breaking them would be the unattended renewal loop months later.
The omission is warned about instead.
RFC1918 is allowed and is usually the right answer here, and no DNS resolution
is performed — both deliberate departures from utils/ssrf_guard, whose policy
is the opposite of what this address needs. What the management host can
resolve says nothing about what the HAProxy node can reach.
Diagnostics stop reporting success on a dead path:
- check_port80 uses GET instead of HEAD and classifies the body. A proxy that
has lost its /.well-known/acme-challenge/ location serves its SPA with HTTP
200, which `status in (200, 404)` accepted as healthy. Warnings also surface
when other domains pass, which previously hid the most diagnostic outcome.
- check_routing filters `mode`, joins `acme_enabled` and reads the APPLIED
config instead of counting database rows, and reports a loopback target.
- every new condition is `warn`, never `fail`: the site wizard blocks submit on
any fail, so a new failing condition would lock every install on upgrade day.
Also: normalise `frontends.mode` once per frontend. It is nullable, and the
raw value was interpolated into `mode {}`, emitting a literal `mode None` that
HAProxy rejects — taking down the whole cluster config. The ACME gate and the
backend-mode check now read the same normalised value.
And stop hardcoding PUBLIC_URL / MANAGEMENT_BASE_URL in docker-compose, which
silently ignored the operator's .env and made the wrong default load-bearing.
|
||
|
|
ef26860df9 |
feat(logging): unified request/response log with configurable retention (v1.11.0)
Until now the only record of what happened was `user_activity_logs`, which stores non-GET 2xx operations with no bodies. When something failed you could see that a counter went up, never what was sent or what came back. This adds one queryable timeline covering both directions: - inbound: every API call, including GETs and including 4xx/5xx, with the user, client IP, status, duration and — redacted, size-capped — the request and response bodies. - outbound: every HTTP call the backend makes, tagged with who it went to (ACME/Let's Encrypt, Cloudflare, GoDaddy, HAProxy stats, agents, the ACME diagnostics probe). Outbound rows inherit the inbound request's id, so one operator action and the CA/DNS calls it triggered read as a single trace: opening a failed "Request Certificate" shows the exact POST /acme/new-order and the CA's 429 underneath. Implementation notes: - Capture is a pure-ASGI middleware that TEES the request and response streams rather than draining them. `await request.body()` inside a BaseHTTPMiddleware would consume the receive channel and break the raw-body agent heartbeat handler. Registered last so it is outermost: it then sees the final client-visible response and seeds correlation_id_context before the error handler reads it. - Rows are written by a batching background writer with a bounded queue, so the request path never awaits the database and a saturated logger drops rows visibly (surfaced on the page) instead of blocking. Redaction runs on the writer, off the request coroutine. - Secrets never land: headers are an allowlist with Authorization/Cookie kept only as a presence marker; body keys and value shapes are redacted (passwords, tokens, api_token, API keys, private-key PEMs, JWTs); the ACME JWS request body is never stored, because a stored protected+signature pair is a replayable credential — a summary is logged instead; DNS-provider errors record only the exception type; the ACME HTTP-01 challenge endpoint is excluded so key_authorization is never captured. - Retention is operator-configurable in Settings -> Request Log: separate day counts for successful and failed rows (7 / 30) plus a hard row cap (500k), whichever is reached first. Pruned in batches under a Postgres advisory lock, with the day counts bound as parameters, never interpolated. - New permissions requestlog.read / requestlog.manage. super_admin and security_admin get both, operator gets read, viewer gets neither. Schema: one new table (request_logs) plus its settings seed, SCHEMA_VERSION 10 -> 11, auto-migrated. No existing table altered, no agent or rendered-config change. Kill switches: REQUEST_LOG_ENABLED=false (middleware never registered) or the `enabled` toggle in Settings. Tests: 245 new (7 backend files + 1 frontend), full suite 1655 backend + 17 frontend passing. |
||
|
|
eee0a4716a |
feat(ssl): encrypt the pending CSR private key at rest (v1.10.1, closes #53)
Closes the follow-up filed during the v1.9.0 CSR review. The private key of a
PENDING CSR is now Fernet-encrypted in the database instead of being stored as
a raw PEM.
Why this key specifically: it is the one key in the system that sits idle. It
is generated at CSR creation, waits for an external CA to sign the request
(days to weeks), and is destroyed the moment the signed certificate is
imported. It is never transmitted to an agent and never leaves the server.
ssl_certificates.private_key_content and the ACME order keys are deliberately
NOT covered, because agents must receive those in plaintext on every poll, so
encrypting them at rest buys nothing without an end-to-end redesign.
Implementation follows the pattern already used for the VRRP secret, TOTP
secrets and DNS provider credentials: a new utils/csr_key_crypto.py with its
own CSR_ENCRYPTION_KEY env var and its own HKDF info string
("csr-private-key-v1"), so rotating one secret class never affects another.
No schema change and deliberately NO SCHEMA_VERSION bump: the Fernet token
replaces the PEM inside the existing ssl_csrs.private_key_pem TEXT column. A
bump would re-run the migration sequence and re-seed the four built-in roles to
their defaults, which is a needless side effect for a storage-format change.
Backward compatible with no data migration. Rows written before this release
hold a raw PEM and are still read unchanged; the discriminator is exact rather
than a heuristic, since a Fernet token is base64url and can never contain the
"-----BEGIN" marker. Legacy rows drain naturally because a CSR's key copy is
NULLed on import.
A key that cannot be decrypted (SECRET_KEY rotated while CSR_ENCRYPTION_KEY was
unset) now fails with an explicit "delete this CSR and create a new one" error.
Previously that situation would have surfaced as the far more confusing
"certificate does not match this CSR's private key".
Also documents all four per-purpose encryption keys in .env.template. Only
VIP_ENCRYPTION_KEY was listed; MFA_ENCRYPTION_KEY and
DNS_PROVIDER_ENCRYPTION_KEY had been missing since v1.6.0 and v1.8.0.
Verified before release, on a corporate pre-production environment and locally:
- Full backend suite 1234 -> 1243 passed (+9 new tests), 0 failed.
- Against a real Postgres: a CSR created through the API stores a Fernet token
with no PEM header in the column, and imports successfully.
- Full 1.10.0 -> 1.10.1 -> 1.10.0 drill on one database volume. The upgrade
logs "Schema already at version 10 (>= 10); skipping migration run", so no
migration executes and the built-in roles are not re-seeded. A CSR created on
1.10.0 with a plaintext key imports successfully after the upgrade, which is
the backward-compatibility guarantee proven against a real row rather than a
mock.
- rsa-2048, rsa-4096 and ecdsa-p384 all round-trip through create, encrypt,
decrypt and import.
- Key derivation is stable across processes: two independent containers sharing
SECRET_KEY decrypt each other's tokens (required for UVICORN_WORKERS > 1 and
multi-replica deployments), while a different SECRET_KEY yields None rather
than a wrong key or an exception.
- Downgrade behaviour was measured, not assumed: 1.10.0 cannot parse the token
and fails with HTTP 500 "key parse failed (encrypted?)" rather than pairing a
wrong key. The rollback note states the measured behaviour.
- No CSR endpoint returns the key in any form: list and detail responses
contain neither a PEM nor a Fernet token.
Not changed here, from the issue's "worth folding in" list: the create rate
limit is not a concurrency guard, create_csr holds a pooled connection across
RSA key generation, detail=str(e) echoes internal error text (a repo-wide
convention), and is_global skips cluster validation in both routers/ssl.py and
routers/csr.py. None are storage concerns and each is a separate change.
|
||
|
|
23257b02cf |
perf(api): opt-in uvicorn workers + heartbeat query consolidation (v1.8.6, Issue #35)
A user running the API on a 2-core/4GB host reported slow-feeling API responses (Issue #35 follow-up). Review of the hot paths found no pathological defect; the dominant factors are the single uvicorn worker (one core serves all requests) and the constant agent-poll baseline (4 requests per agent every 30s). Two zero-risk improvements: - backend/Dockerfile: CMD now honors UVICORN_WORKERS, falling back to WEB_CONCURRENCY and then 1. Flagless uvicorn natively honors WEB_CONCURRENCY, so the fallback keeps any deployment that relied on it byte-for-byte compatible; with the final default of 1 worker uvicorn runs in-process exactly as before. >1 enables the multiprocess supervisor so multi-core hosts can use all cores. Background tasks are already multi-replica safe (FOR UPDATE SKIP LOCKED / advisory locks), as exercised by the k8s HPA deployment (2-10 replicas). `exec` keeps uvicorn as PID 1 (clean SIGTERM, verified ~1s docker stop with 2 workers). docker-compose.yml passes UVICORN_WORKERS through as empty-when-unset so a user-set WEB_CONCURRENCY is never overridden; .env.template documents it. - routers/agent.py heartbeat (by-name endpoint): the agent's status/version/upgrade_status were read with three separate single-column SELECTs against the same row; now one SELECT. Identical values and None semantics (single consistent snapshot instead of three reads); saves two round-trips per heartbeat per agent every 30s. The legacy by-id heartbeat endpoint is untouched; the heartbeat API contract is unchanged for agents of every version. - README: new "Performance Tuning" section (worker/replica scaling, and how to use the X-Response-Time header plus "Slow request detected" logs to pinpoint slow endpoints). Verification: full backend suite in docker green (1063 passed, 151 skipped; also re-run by the runtime image build); worker-count expansion matrix (unset->1, UVICORN_WORKERS=2->2, WEB_CONCURRENCY=3->3, both->UVICORN_WORKERS, empty->fallback) all correct; default run confirmed single-process with uvicorn as PID 1 and healthy API; UVICORN_WORKERS=2 confirmed parent + 2 workers, healthy API, clean shutdown; live heartbeats verified for register + existing-agent paths AND degraded agents (no stats socket / haproxy stopped / garbage stats CSV / unknown backend in server_statuses): all return 200, agent row updates correctly, zero backend errors. No schema, API, or agent changes. |
||
|
|
a1192e602d |
feat: HA / VIP (Keepalived) management from the UI (#27)
Manage highly-available virtual IPs backed by Keepalived (VRRP) directly from the OpenManager
UI — no more SSHing into nodes to install/configure Keepalived by hand. Builds on the agent
pull-architecture: define the VIP centrally, click Apply, and the agents converge.
Highlights:
- New "HA / VIP" tab: create a virtual IP, pick a per-node interface, select which pool nodes
participate (MASTER/BACKUP roles + priorities); live MASTER/BACKUP/FAULT per node.
- On Apply, agents install & configure Keepalived (unicast VRRP, cloud-safe default) across the
major distros (Debian/Ubuntu, RHEL/CentOS/Alma/Rocky, Fedora, SUSE/openSUSE, Alpine) with a
HAProxy health-check, so the VIP fails over automatically when HAProxy drops.
- Single-node (a managed floating IP without failover) and multi-node VRRP failover both work.
- VIP changes ride the standard Apply Management flow with the standard "View Change" diff.
- Approval-gated deletion (safety): deleting a running VIP is staged for approval and the VIP
keeps running, untouched, until you approve it — an agent never tears a VIP down without an
explicit human approval. Per-VIP Diagnostics view; opt-in package uninstall (only on nodes
where OpenManager installed it). A node already running a hand-managed Keepalived is detected
and never overwritten ("externally managed").
- Fully opt-in and backward compatible: nodes/clusters without a VIP are unaffected. Adds
vip_instances + vip_members tables (idempotent SCHEMA_VERSION bump; existing data and
passwords unaffected) and a `vip` RBAC permission group.
- Also includes a HAProxy config-generator robustness fix: auto-inject a stick-table when a
frontend uses a stick counter (track-sc / sc_*_rate) but declares none.
On-prem / L2 (VRRP) scope; the UI notes the cloud caveat.
|
||
|
|
13179279d6 |
fix: port conflict validation ignoring bind_address and CORS on non-standard ports
- Fix frontend port conflict validation to consider bind_address+port combination instead of port-only. HAProxy allows same port on different bind addresses (e.g., bind 10.0.0.1:443 vs bind 10.0.0.2:443). This was blocking frontend edit/save in multi-VIP environments. - Add form field dependency so port re-validates when bind_address changes. - Fix API URL construction using window.location.host instead of hostname to preserve non-standard ports (e.g., :8080), preventing CORS errors in BulkConfigImport. - Make CORS_ORIGINS configurable via environment variable. |
||
|
|
6aae0f4309 | Initial commit |