6 Commits

Author SHA1 Message Date
taylanbakircioglu eee0a4716a feat(ssl): encrypt the pending CSR private key at rest (v1.10.1, closes #53)
Closes the follow-up filed during the v1.9.0 CSR review. The private key of a
PENDING CSR is now Fernet-encrypted in the database instead of being stored as
a raw PEM.

Why this key specifically: it is the one key in the system that sits idle. It
is generated at CSR creation, waits for an external CA to sign the request
(days to weeks), and is destroyed the moment the signed certificate is
imported. It is never transmitted to an agent and never leaves the server.
ssl_certificates.private_key_content and the ACME order keys are deliberately
NOT covered, because agents must receive those in plaintext on every poll, so
encrypting them at rest buys nothing without an end-to-end redesign.

Implementation follows the pattern already used for the VRRP secret, TOTP
secrets and DNS provider credentials: a new utils/csr_key_crypto.py with its
own CSR_ENCRYPTION_KEY env var and its own HKDF info string
("csr-private-key-v1"), so rotating one secret class never affects another.

No schema change and deliberately NO SCHEMA_VERSION bump: the Fernet token
replaces the PEM inside the existing ssl_csrs.private_key_pem TEXT column. A
bump would re-run the migration sequence and re-seed the four built-in roles to
their defaults, which is a needless side effect for a storage-format change.

Backward compatible with no data migration. Rows written before this release
hold a raw PEM and are still read unchanged; the discriminator is exact rather
than a heuristic, since a Fernet token is base64url and can never contain the
"-----BEGIN" marker. Legacy rows drain naturally because a CSR's key copy is
NULLed on import.

A key that cannot be decrypted (SECRET_KEY rotated while CSR_ENCRYPTION_KEY was
unset) now fails with an explicit "delete this CSR and create a new one" error.
Previously that situation would have surfaced as the far more confusing
"certificate does not match this CSR's private key".

Also documents all four per-purpose encryption keys in .env.template. Only
VIP_ENCRYPTION_KEY was listed; MFA_ENCRYPTION_KEY and
DNS_PROVIDER_ENCRYPTION_KEY had been missing since v1.6.0 and v1.8.0.

Verified before release, on a corporate pre-production environment and locally:
- Full backend suite 1234 -> 1243 passed (+9 new tests), 0 failed.
- Against a real Postgres: a CSR created through the API stores a Fernet token
  with no PEM header in the column, and imports successfully.
- Full 1.10.0 -> 1.10.1 -> 1.10.0 drill on one database volume. The upgrade
  logs "Schema already at version 10 (>= 10); skipping migration run", so no
  migration executes and the built-in roles are not re-seeded. A CSR created on
  1.10.0 with a plaintext key imports successfully after the upgrade, which is
  the backward-compatibility guarantee proven against a real row rather than a
  mock.
- rsa-2048, rsa-4096 and ecdsa-p384 all round-trip through create, encrypt,
  decrypt and import.
- Key derivation is stable across processes: two independent containers sharing
  SECRET_KEY decrypt each other's tokens (required for UVICORN_WORKERS > 1 and
  multi-replica deployments), while a different SECRET_KEY yields None rather
  than a wrong key or an exception.
- Downgrade behaviour was measured, not assumed: 1.10.0 cannot parse the token
  and fails with HTTP 500 "key parse failed (encrypted?)" rather than pairing a
  wrong key. The rollback note states the measured behaviour.
- No CSR endpoint returns the key in any form: list and detail responses
  contain neither a PEM nor a Fernet token.

Not changed here, from the issue's "worth folding in" list: the create rate
limit is not a concurrency guard, create_csr holds a pooled connection across
RSA key generation, detail=str(e) echoes internal error text (a repo-wide
convention), and is_global skips cluster validation in both routers/ssl.py and
routers/csr.py. None are storage concerns and each is a separate change.
2026-08-08 01:44:32 +03:00
mustafa.ulukaya 81ab674072 docs(v1.10.0): document the GoDaddy DNS provider and upgrade notes
README: add GoDaddy to the two feature bullets and to the DNS-01 provider
catalog, spelling out that the API Key must be a Production key (the first key
the developer dashboard issues is an OTE/test key and is rejected), that the
zone must be in the same account, that the account needs a registered domain
before GoDaddy permits DNS API access, and that a Personal Access Token works
with the Secret left blank. Note that publishing is automatic for GoDaddy as
well as Cloudflare, and add the release-notes entry.

UPGRADE_GUIDE: new section stating there is no SCHEMA_VERSION bump, so the
built-in-role re-seed warning from v1.9.0 does not apply, and no new
environment variable, API-shape or agent change. Two limits are stated
explicitly rather than glossed: the credential check is a read, so a token
with read but not write scope saves successfully and only fails at the first
publish; and downgrading after adopting GoDaddy is not a no-op, because an
unknown provider name degrades DNS-01 orders to the manual-confirm path and
leaves published TXT records marked cleaned without being removed.
2026-08-07 09:01:35 +03:00
taylanbakircioglu 71786200dd docs(v1.9.0): correct upgrade notes on built-in role re-seed + backfill 1.8.8-1.8.10 release notes
Found during a v1.8.10 to v1.9.0 upgrade drill on a populated database
(schema v9 to v10) before releasing. Documentation only, no code change.

1. The v1.9.0 upgrade notes said "custom roles need no changes", which reads
   as "role data is untouched". It is not: because the SCHEMA_VERSION bump
   re-runs the whole idempotent sequence, update_system_roles_to_enterprise_rbac()
   issues an unconditional UPDATE roles SET ... permissions = <defaults> for
   the four BUILT-IN roles. In the drill an `operator` role that had been
   narrowed by removing apply.execute and config.bulk_import came back with
   both restored (57 to 59 permissions). Operator-created roles are NOT
   affected; the re-seed matches the four built-in names only.

   This is pre-existing behaviour of every SCHEMA_VERSION bump and is
   documented as intentional in migrations.py, so it is not introduced by the
   CSR feature. The v1.7.0 upgrade notes carried this caveat and it was not
   carried forward. Restored, with an export and re-apply procedure.

   Also clarified why the admin password is safe: the default-user seeding is
   guarded by an existence check ("safer than ON CONFLICT"), not an upsert,
   so an operator-changed password survives.

2. README release notes jumped from v1.8.7 straight to v1.9.0 because
   v1.8.8, v1.8.9 and v1.8.10 were never backfilled. Added all three.
2026-08-05 23:43:14 +03:00
mustafa.ulukaya 69e12f7459 chore(version): bump to 1.9.0 - CSR creation
- backend/version.json + frontend package version to 1.9.0
- README: feature list entry, CSR workflow section, SSL CSR API reference,
  v1.9.0 release notes
- UPGRADE_GUIDE: v1.9.0 section (additive ssl_csrs table, SCHEMA_VERSION
  9 -> 10, no RBAC changes, zero agent impact, rollback note)
2026-08-04 21:24:34 +03:00
taylanbakircioglu a1192e602d feat: HA / VIP (Keepalived) management from the UI (#27)
Manage highly-available virtual IPs backed by Keepalived (VRRP) directly from the OpenManager
UI — no more SSHing into nodes to install/configure Keepalived by hand. Builds on the agent
pull-architecture: define the VIP centrally, click Apply, and the agents converge.

Highlights:
- New "HA / VIP" tab: create a virtual IP, pick a per-node interface, select which pool nodes
  participate (MASTER/BACKUP roles + priorities); live MASTER/BACKUP/FAULT per node.
- On Apply, agents install & configure Keepalived (unicast VRRP, cloud-safe default) across the
  major distros (Debian/Ubuntu, RHEL/CentOS/Alma/Rocky, Fedora, SUSE/openSUSE, Alpine) with a
  HAProxy health-check, so the VIP fails over automatically when HAProxy drops.
- Single-node (a managed floating IP without failover) and multi-node VRRP failover both work.
- VIP changes ride the standard Apply Management flow with the standard "View Change" diff.
- Approval-gated deletion (safety): deleting a running VIP is staged for approval and the VIP
  keeps running, untouched, until you approve it — an agent never tears a VIP down without an
  explicit human approval. Per-VIP Diagnostics view; opt-in package uninstall (only on nodes
  where OpenManager installed it). A node already running a hand-managed Keepalived is detected
  and never overwritten ("externally managed").
- Fully opt-in and backward compatible: nodes/clusters without a VIP are unaffected. Adds
  vip_instances + vip_members tables (idempotent SCHEMA_VERSION bump; existing data and
  passwords unaffected) and a `vip` RBAC permission group.
- Also includes a HAProxy config-generator robustness fix: auto-inject a stick-table when a
  frontend uses a stick counter (track-sc / sc_*_rate) but declares none.

On-prem / L2 (VRRP) scope; the UI notes the cloud caveat.
2026-06-07 01:52:20 +03:00
taylanbakircioglu 73bed7811f docs: Add CONFIG.md and UPGRADE_GUIDE.md with public-friendly URLs
- Add configuration guide for environment variables
- Add agent upgrade guide
- Use example.com instead of company-specific URLs
- No sensitive information included
2025-11-12 14:15:51 +03:00