94 Commits

Author SHA1 Message Date
taylanbakircioglu d7208528f7 fix: v1.5.2 — ACME Diagnostics Panel Hardening + AGPL-3.0 relicense (Bulgu #94/#95/#96)
A focused hardening pass on the v1.5.0 ACME Diagnostic Panel
surface, exercised against a live production deployment (Round-25
+ Round-26 audits) and supplemented by an AGPL-3.0 relicense.

------------------------------------------------------------------
LICENSE — Relicense to AGPL-3.0-or-later
------------------------------------------------------------------
Effective v1.5.2 the project is licensed under the **GNU Affero
General Public License v3.0 (or later)**. v1.5.0 and v1.5.1
remain under the prior MIT terms.

The relicense is consistent with the project's intent as a
community-operated HAProxy management surface: forks that run
HAProxy OpenManager as a network service for third parties are
now required to publish their modifications under the same
license (AGPL §13). Day-to-day single-tenant deployments,
internal corporate use, and ordinary forks-for-fixes are
unaffected.

Changes:
  * LICENSE replaced with full AGPL-3.0 text.
  * README "## License" section rewritten with the AGPL summary
    + the network-service obligation.
  * frontend/package.json gains `"license": "AGPL-3.0-or-later"`.

------------------------------------------------------------------
BULGU #94 / #95 — Diagnostic Panel Must Never Opaque-500
------------------------------------------------------------------
Live exercise of the v1.5.0 Diagnostic Panel against a deployed
build surfaced two opaque-500 paths. The panel exists to make
ACME failures legible; producing an opaque HTTP 500 defeats the
entire feature. Fix shape: every endpoint now returns either a
canonical 4xx (auth / not-found / rate-limit) or an HTTP 200
"structured failure envelope" that the React UI knows how to
render — never a 500 for an in-suite failure.

Affected paths:

POST /api/letsencrypt/orders/{order_id}/diagnostics
  Pre-fix: a UndefinedColumnError or DB-connectivity failure
  inside `run_checks` bubbled out of the bare try/finally and
  surfaced as a generic 500 with no operator-actionable detail.
  Post-fix: setup-stage and run-stage failures are caught
  separately and converted to a `status: diagnostics_unavailable`
  envelope carrying `error_stage`, `error_type`, `error_message`,
  and a `correlation_id` that the operator can grep in the
  backend log. Individual checks are wrapped in `_safe_check`
  so one broken check (e.g. DNS lookup timeout) never crashes
  the suite — the failing check shows up as `status: "fail"`
  with its message, the others still run.

GET /api/letsencrypt/orders/{order_id}/events
  Pre-fix: the SQL `SELECT … status FROM user_activity_logs`
  referenced a column that did not exist in the canonical
  migration; every diagnostic-panel open against an order with
  any user-activity-log correlation got an `UndefinedColumnError`
  500. Post-fix: the endpoint now introspects
  `information_schema.columns` and projects only the columns
  actually present. Partial failures (one source dies, the
  other works) are reported via `meta.errors[]` rather than
  collapsing the whole timeline.

POST /api/letsencrypt/orders/{order_id}/diagnostics/{check_id}/rerun
  Same structured-envelope contract as the full-suite POST,
  scoped to a single check row.

Frontend (`frontend/src/components/ACMEAutomation.js`):
  * Distinct `diagRunError` / `diagEventsError` / `diagMeta`
    states so the modal can render the cause inline (Antd Alert)
    instead of a silent dropdown.
  * Event-log auto-tail polling backs off after 3 consecutive
    failures so the Network tab does not get spammed with 500s
    every 5s.
  * Correlation IDs visible in every error banner.

------------------------------------------------------------------
BULGU #96 — Clean 404 for Out-Of-Range order_id
------------------------------------------------------------------
A live exercise of the post-#94 diagnostic panel against the
deployed build surfaced one remaining contract gap. A path-
param `order_id` outside the Postgres int4 range
(e.g. > 2_147_483_647) caused `_load_order` to raise
`asyncpg.exceptions.DataError: invalid input for query
argument $1: 2147483648 (value out of int32 range)`. Round-25
correctly surfaced this in a `diagnostics_unavailable`
envelope — but that envelope leaked SQL implementation detail
("query argument $1", "int32 range", DataError class name)
into the operator-facing response body.

Semantically an out-of-range integer can never reference a
real order — it's just "not found". `_load_order` now catches
`asyncpg.exceptions.DataError` and re-raises a canonical
`HTTPException(404, "Order {id} not found")`. Because all
three endpoints re-raise `HTTPException` from their outer
try/except (the Round-25 envelope only fires for non-
HTTPException crashes), the canonical 404 path now wins
end-to-end across /diagnostics, /events, and /rerun.

------------------------------------------------------------------
TEST / LINT / LIVE VERIFICATION
------------------------------------------------------------------
  * Backend pytest 1104/1104 (the +20 vs v1.5.1's 1084 are the
    Round-25 and #96 contract pins; see
    test_acme_diagnostics_router_round25.py).
  * Live prod-canary verification: every endpoint return shape
    confirmed against the deployed build — int4 overflow returns
    clean 404 with no SQL leak, normal paths return Round-25
    envelopes, HTTP method matrix returns 405 on wrong verbs,
    no auth returns 401, invalid `check_id` returns 400, and
    `meta.correlation_id` is present on every diagnostic
    response.

------------------------------------------------------------------
COMPATIBILITY
------------------------------------------------------------------
  * No breaking API contract changes: `status` field on the
    diagnostic response can now be `"diagnostics_unavailable"`
    in addition to the existing pass-through of the
    underlying order status (`pending` / `valid` / `invalid` /
    `cancelled` / …) — older UIs that only switch on the
    existing values render the `diagnostics_unavailable`
    case as "unknown status" rather than crashing.
  * Frontend handles the new envelope shape AND the legacy
    HTTP 4xx/5xx paths.
2026-05-14 00:07:00 +03:00
taylanbakircioglu 2e7db4d99f fix: v1.5.1 — Round-23 + Round-24 audit follow-ups (Bulgu #83#93)
A live-deployment audit pass over the v1.5.0 Site Wizard + ACME
Diagnostic Panel surface. Two adversarial review rounds (R23, R24)
each capped by an end-to-end smoke test against a multi-cluster
staging deployment.

Bulgu #83 — Frontend Management page warned about stale data
without a clear retry CTA. The toast now carries an in-place
"Reload" action and the page-level Empty state surfaces the same
recovery affordance, so operators never get stuck on a stale-data
view without an obvious way out.

Bulgu #84 — ACME diagnostics ran with the wrong "last_heartbeat"
column reference against the agents table. Aligned the SELECT
with the actual schema column (`last_seen`); pinned by an idempotent
regression test in `test_acme_diagnostics.py`.

Bulgu #85 — ACME order error_detail rendering could leak the raw
asyncpg/SQL exception class name when humanize_error_detail
encountered an unhandled CA response shape. Added a backwards-
compatible fallback branch that emits an "ACME error (raw)" panel
without exposing parse_error class name to the user.

Bulgu #86 — Multi-cluster apply with concurrent rejects could
leave wizard_staged orders dangling without their parent draft.
Pinned via reject_order_with_cluster_orphan test.

Bulgu #87 — Frontend Management page list virtualization
mis-keyed during a re-sort + stale-row replace race; fixed by
keying rows on `id + version` so React reconciler does not reuse
DOM for a logically different row.

Bulgu #88 — Site Wizard "Cancel" mid-flow now surfaces an
unsaved-draft prompt with explicit Save / Discard buttons (and
the same prompt on browser tab close), so the operator never
loses 5 steps of input to an accidental ESC.

Bulgu #89 — Existing-cert SSL mode showed an empty dropdown when
the cluster had >100 certs because the listing endpoint
default-limited results. Endpoint now exposes pagination AND
the wizard switches to client-side filtering above 50 rows.

Bulgu #90 — ACME pre-check on the wizard preview path did NOT
re-validate the account against `letsencrypt_accounts` if the
operator stepped Back/Forward between SSL and Review. Added a
debounced re-validation on Review entry.

Bulgu #93 — Site Wizard hsts_enabled toggle in HTTPS frontend
was idempotent-by-name (the generated `http-response set-header
Strict-Transport-Security` line could duplicate across a Save +
Apply cycle). The renderer now upserts the header in place.

Cumulative outcome: backend pytest 1084/1084, frontend lint
clean, and a 6-hour live-deployment smoke session against staging
with no regressions reported.
2026-05-14 00:06:02 +03:00
taylanbakircioglu 02b1cb2bca feat: v1.5.0 — Site Wizard (Issue #14) + ACME Diagnostic Panel (Issue #13)
Closes #13, Closes #14.

This release squashes the v1.4.0 → v1.5.0 development line. v1.4.0
shipped the ACME stability & enterprise audit (Issues #10/#11/#12).
v1.5.0 builds on that foundation with two co-equal headline features
plus a 22-round audit campaign hardening the prior configuration
surface. License remains MIT for v1.5.0 (relicense to AGPL-3.0
lands in v1.5.2).

------------------------------------------------------------------
HEADLINE FEATURE A — ACME Diagnostic Panel (Issue #13)
------------------------------------------------------------------
A live pre-flight + post-failure diagnostic surface for every ACME
order, reachable from the ACME Automation page. The panel exists
to make ACME failures legible to operators who do NOT have shell
access to the API host.

Endpoints (`backend/routers/acme_diagnostics.py`):
  POST /api/letsencrypt/orders/{order_id}/diagnostics
       Run the full 5-check suite (DNS / port-80 / routing /
       account / agents) and humanize the order's `error_detail`
       (>=11 RFC-8555 problem types, backwards compatible with
       legacy plain-string failures).
  POST /api/letsencrypt/orders/{order_id}/diagnostics/
                                {check_id}/rerun
       Re-run a single check in place — used by the "Re-run"
       button on every row of the modal's pre-flight table.
  GET  /api/letsencrypt/orders/{order_id}/events
       Merged event timeline combining the typed
       `acme_order_events` rows with correlated
       `user_activity_logs` entries (resource_type =
       'letsencrypt_order' AND resource_id = order_id). The
       diagnostic modal auto-tails this timeline every 5 seconds
       while open.

Service-level checks (`backend/services/acme_diagnostics.py`):
  * DNS resolution via stdlib socket.gethostbyname_ex through
    run_in_executor (intentionally avoiding an aiodns runtime
    dep for v1.5.0).
  * Port-80 HEAD probe, target locked to the order's domains,
    success on HTTP 200 OR 404, warns on egress timeout
    (corp egress policies routinely blackhole outbound 80 —
    fail-hard would be too noisy).
  * SSRF guard: probe refuses non-public IPs and surfaces the
    skip in the diagnostic result; IPv4-mapped IPv6 normalisation
    closes the `::ffff:169.254.169.254` cloud-metadata vector.
  * HAProxy routing presence check: matches the order's
    cluster_ids to a port-80 HTTP frontend.
  * ACME account validity check against `letsencrypt_accounts`.
  * Agent presence check (>=1 active agent in target cluster).
  * Every sub-check wrapped in a wall-clock timeout to bound
    impact on the API event loop.

RBAC: ssl.read for run, ssl.read for events. Per-user 5/min rate
limit on both run and rerun, backed by the (user_id, action,
created_at DESC) composite index.

Frontend (`frontend/src/components/ACMEAutomation.js`):
  * "Diagnose" button on every order row + the existing
    "stuck order" warning row.
  * Modal with two tabs:
    - Pre-flight Checks (Antd Table with status pills + Re-run
      buttons + humanized error banner)
    - Event Log (Antd Timeline with auto-tail polling, scroll-
      to-bottom, pause-on-hover)
  * Correlation IDs surfaced in error banners and individual
    check fail details for backend-log lookup.

------------------------------------------------------------------
HEADLINE FEATURE B — Site Setup Wizard (Issue #14)
------------------------------------------------------------------
A single guided flow that creates a Backend + Servers + HTTP
Frontend (and optional HTTPS Frontend) in one atomic transaction.

Endpoints (`backend/routers/site_wizard.py`):
  POST /api/site-wizard/preview     — diff-preview the changeset
  POST /api/site-wizard/create      — atomic execute
  POST /api/site-wizard/reject      — clean rollback (including
                                       any wizard_staged ACME
                                       orders)
  GET  /api/site-wizard/drafts      — draft persistence
  PUT  /api/site-wizard/drafts/{id} — save/update
  DELETE /api/site-wizard/drafts/{id}

Feature surface:
  * One screen captures both backend (mode + servers) AND
    frontend (http + optional https + SSL mode) inputs.
  * SSL modes: ACME (new order, HTTP-01 only for v1.5.0),
    Upload (existing PEM), Existing (link to a stored cert),
    or None.
  * ACME-staged path: wizard_staged_until watermark on the
    `letsencrypt_orders` row defers finalisation until agent
    confirmation; per-mode reject cleanly cancels and rolls
    back the staged order.
  * Live diff preview against the cluster's current generated
    config (renderer-evolution noise stripped — track-sc<N>
    dedup, per-server cookie strip, defaults-cookie
    inheritance, listen-block flattening).
  * Draft persistence with PEM stripped at save time (private
    keys never round-trip through the drafts table).
  * Per-cluster multi-tenancy: drafts and wizard_staged orders
    are isolated to the creating user's cluster scope.

Frontend (`frontend/src/components/SiteWizard.js`):
  * 4-step Antd Steps flow: Backend → Frontend → SSL → Review.
  * Render the live diff preview inline before commit.
  * Antd Form-level validation mirrors backend Pydantic
    validators (numeric bounds, HAProxy reserved keywords, ALPN
    consistency, IPv6 scope-id, domain regex, server name
    dedup).

------------------------------------------------------------------
AUDIT CAMPAIGN — Rounds 1 → 22 (Bulgu #1#82)
------------------------------------------------------------------
v1.5.0 includes 22 adversarial review passes. Each round produced
its own commit set in the corporate development line; this squash
collapses those into the v1.5.0 release artefact. Highlights:

  Round 1-4   Site Wizard core: dry-run parity, single-line
              value injection guard, ACL -f pattern-file block,
              SSL parity, timeout regex, form-state pin.
  Round 5-7   defaults-cookie inheritance, server-named-cookie
              guard, fe/be mode mismatch, duplicate server
              names, health_check_uri + server_address
              validators.
  Round 8-10  cookie_name / cookie_options newline-injection
              guard, dry-run parity (round 9), TCP-mode HTTP-only
              feature blockers.
  Round 11    SSL name path traversal + health-check >= 1.
  Round 12-13 SSL & ACME deep dive (Bulgu #23-#32).
  Round 14    single-line value injection (Bulgu #33).
  Round 15-17 ACME multi-tenant UX, numeric bounds, HAProxy
              reserved keywords, ALPN/TLS consistency,
              all-backup, multi-domain & multi-user enterprise
              edges, drain/HSTS/post-completion (Bulgu
              #34-#53).
  Round 18-21 concurrency, agent state, TCP-mode HTTP-only,
              list size caps, IPv6 scope-id, preview account
              validation, TCP backend + balance uri reject
              (Bulgu #54-#61).
  Round 22    FE error visibility + 3x stale-data lockouts,
              referential integrity + cascade safety,
              authentication & authorization, multi-cluster
              isolation, apply_pending_changes concurrency,
              script injection + bulk import multi-tenancy,
              prefix-stripped signature comparison
              (Bulgu #62-#82).

------------------------------------------------------------------
NO CORPORATE-SPECIFIC ARTIFACTS
------------------------------------------------------------------
This squash deliberately sanitises corporate hostnames, container
registry references, and TLS secret names into generic
placeholders (`your-registry.example.com/your-org`,
`haproxy-openmanager*.example.com`, `wildcard-tls`,
`taylanbakircioglu/haproxy-openmanager-*`) so the public artefact
contains no internal infrastructure detail. Pilot / development
history that retained those values stays in the corporate fork
and is NOT part of this commit.
2026-05-14 00:04:19 +03:00
taylanbakircioglu 07942a82e8 feat: ACME stability & enterprise audit (v1.4.0) — fixes #10 #11 #12
Issue #10 — Silent timezone failure on ACME certificate save
- Normalize tz-aware expiry_date to UTC tz-naive before INSERT/UPDATE in
  ssl_certificates (TIMESTAMP WITHOUT TIME ZONE) — restores ACME download path

Issue #11 — Duplicate _acme_challenge_backend in generated config
- Generator guard: skip auto-append when backend already rendered
- Agent-sync filter: _should_sync_backend() drops system-managed backends and
  detaches their server rows to prevent orphans
- Restore filter: IGNORED_BACKENDS skips reserved names during cluster restore
- Parser warning: reserved_backend_names blocks accidental manual import
- Cleanup migration: removes orphan rows + cascading server entries (idempotent)

Issue #12 — Validated orders required manual completion
- New 60s background task complete_pending_acme_orders, flag-independent,
  multi-replica safe via FOR UPDATE SKIP LOCKED + 30s updated_at watermark
- Per-order pg_advisory_lock(0x41434D45, order_id) serializes UI-Complete and
  auto-task races; idempotency guard returns existing certificate cleanly
- retry_order endpoint reports in_progress: true within 30s window so the UI
  surfaces an info toast instead of duplicating CA requests
- ACMEAutomation surfaces stuck orders (status=valid && !ssl_certificate_id)
  with a one-click Complete action and Cancel fallback; conditional 30s polling

Other hardening
- ACME state machine error_detail persisted as structured JSON across challenge,
  finalize, download stages for actionable post-mortems
- CertificateRequest Pydantic model: domain regex + min_length/max_length and
  cluster_ids defaulting to all ACME-enabled clusters when "global" is selected
- Renewal cluster fallback now requires acme_enabled=TRUE in addition to active
- _complete_certificate preserves manual cluster assignments on renewal,
  surfaces cluster_errors, raises explicit error on missing private key
- Audit logging covers acme_certificate_requested/revoked, ca_chain_imported,
  account created/deactivated/purged, order retried/cancelled
- Settings UI exposes acme.staging_url_override for private test CAs (Pebble)
- Schema additions: acme_challenges.attempts (default 0) and last_attempt_at,
  index idx_letsencrypt_orders_status_updated; all migrations idempotent

Tests (41/41 passing)
- test_acme_expiry_normalize, test_acme_duplicate_backend,
  test_acme_state_machine, test_acme_pydantic_validation,
  test_acme_audit_logging, test_acme_concurrency

CI / packaging
- docker-build.yml reads version.json and pushes additional product-version
  tag (e.g. 1.4.0) alongside latest and timestamp build id

Closes #10
Closes #11
Closes #12
2026-05-06 23:53:43 +03:00
taylanbakircioglu cd4a94beb1 feat: agent IP/VIP live update + script update detection banner
- Agent scripts now detect and send ip_address in DAEMON heartbeat (Linux: ip route, macOS: ifconfig)
- Backend validates agent-reported IPs via ipaddress stdlib, COALESCE preserves existing on NULL
- IP/VIP change logging (non-critical, try/except wrapped) for operational visibility
- New source_file_hash column on agent_script_templates for reliable update detection
- Migration changed to ON CONFLICT DO NOTHING to prevent overwriting UI-customized scripts on restart
- GET /versions returns script_update_available flag (disk hash vs DB hash comparison with fallback)
- Frontend Alert banner warns users of new agent script versions and directs to Reset to Defaults
- Reset to Defaults and Popconfirm modals explicitly warn about custom script edit loss
- Full backward compatibility: old agents without ip_address field continue working unchanged

Made-with: Cursor
2026-04-17 10:50:51 +03:00
taylanbakircioglu 61024e24cb fix: strip rate-limit and WAF auto-generated content from bulk import comparison
Made-with: Cursor
2026-04-17 10:45:12 +03:00
taylanbakircioglu 6a349e22c6 fix: strip auto-generated ACME content from bulk import preview
Made-with: Cursor
2026-04-17 10:43:30 +03:00
taylanbakircioglu f37f3afd71 feat: bulk import change detection, multi-select delete, auto-content stripping
- Server-level change detection in bulk import (field-by-field comparison
  for 17 server attributes with UPDATE/NO CHANGES status and tooltip)
- Multi-select delete for backends and frontends with dependency checks
- Dashboard "Backends Summary" address column for servers
- Fix unique constraint violation on bulk-create for existing servers
  (natural key lookup matching DB constraint instead of backend_id FK)
- ORDER BY is_active DESC on all entity lookups to prefer active records
- Strip auto-generated content (ACME, rate-limit, WAF) from bulk import
  comparison to eliminate false positive changes on re-import
- Fix toolbar overflow with Space wrap prop
- Frontend bulk delete modal clarity (selected vs deletable count)
- Version bump to 1.3.0

Made-with: Cursor
2026-04-14 20:21:07 +03:00
taylanbakircioglu 7fffaddaf8 fix: bulk import should not inject hardcoded timeout defaults
When a backend block in haproxy.cfg does not specify explicit timeout
values, the bulk import was injecting hardcoded defaults (connect 10s,
server 60s, queue 60s) into the database. These then appeared in the
generated config and overrode the agent's defaults section. Now, only
explicitly declared timeouts are stored; omitted ones remain NULL so
the agent's existing defaults section stays in effect.
2026-04-13 05:52:31 +03:00
taylanbakircioglu 36fdba52bc fix: ACME setup guide accuracy, reject rollback, and UX improvements
- Step 3 (Enable ACME on Cluster) now shows a process icon instead of
  a misleading green checkmark when ACME is enabled but not yet applied.
  Per-cluster "(pending apply)" annotation for multi-cluster setups.
- Step 4 button and all /apply-management navigation buttons now say
  "Apply Changes" instead of "Configure" for clearer guidance.
- Setup Guide auto-selects the correct cluster before navigating to
  Apply Management, showing pending cluster names in alerts.
- Pending ACME disable changes are now correctly detected in Step 4
  even when acme_enabled is already FALSE in the database.
- Entity snapshot rollback for cluster ACME settings: reject correctly
  restores acme_enabled/acme_backend_url to pre-change values.
- Deduplication logic prevents "last wins" bug when multiple ACME
  toggles are rejected in sequence.
- Connection leak prevention with try/finally around conn2 in ACME
  config version creation.
- Step 4 branching uses boolean has_enabled instead of fragile string
  truthiness check.

Made-with: Cursor
2026-04-04 16:57:57 +03:00
taylanbakircioglu deb784bb65 fix: ACME certificate issuance improvements, null-safe hardening, guided setup UX, and order list enhancements (Issue #9)
- Fix critical cascading NULL status bug in ACME challenge flow that could cause 404s
- Add defense-in-depth NULL handling across all ACME service methods
- Add new GET /api/letsencrypt/prerequisites endpoint for configuration checks
- Add interactive ACME Setup Guide with step-by-step navigation links
- Add URL-based tab navigation in Settings and SSL Management pages
- Harden retry flow: return clear 409 errors for invalid/cancelled orders
- Allow cancellation of invalid orders (backend + frontend)
- Improve error message extraction with consistent getErrorMsg helper
- Add status filter tabs (Active/Completed/Failed/All) for order list
- Add visual dimming for cancelled/invalid orders
- Add enhanced pagination with size changer and total count
- Add comprehensive diagnostic logging with ACME: prefix
- Update README with ACME architecture docs, quick start guide, and troubleshooting

Resolves #9

Made-with: Cursor
2026-04-04 15:15:05 +03:00
taylanbakircioglu 2fc2603d1a fix: ACME account removal, status display, and timeline icon clipping
- Add permanent delete endpoint for deactivated ACME accounts (DELETE /accounts/{id}/permanent)
- Add "Remove from database" button in ACME Accounts modal for deactivated accounts
- Fix ACME Account stat card showing "Active" for deactivated accounts
- Fix Timeline dot/icon clipping in Applied/Rejected/Pending sections

Made-with: Cursor
2026-04-02 00:36:56 +03:00
taylanbakircioglu 93d7ad8fdb feat: add ACME Auto SSL with Let's Encrypt integration (v1.1.0)
Add automated SSL certificate management via ACME protocol (RFC 8555):
- Full ACME client implementation (account registration, HTTP-01 challenges, certificate issuance/renewal)
- Configurable ACME providers (Let's Encrypt, ZeroSSL, custom CA) via Settings UI
- Auto-renewal scheduler with PENDING -> Apply -> APPLIED flow alignment
- ACME account management (register, deactivate) from UI
- Certificate request wizard with domain validation and cluster targeting
- Zero changes to HAProxy agent scripts - challenges routed through existing architecture
- Comprehensive security hardening (no private key exposure in API responses)
- Full backward compatibility with existing SSL, Apply, Rollback, and Restore workflows
- Updated README, API documentation, and Kubernetes deployment notes
- Version management embedded in code (v1.1.0)
- UI messaging improvements for agent-pull architecture accuracy

Made-with: Cursor
2026-04-02 00:25:50 +03:00
taylanbakircioglu 922bc2ce6d improve: SSL certificate listing reliability, error visibility, and ARM64 support
Fixes #6

- Fix NameError in soft-deleted certificate reactivation path by
  reordering variable extraction before DB operations
- Replace silent empty-array returns with HTTP 500 on SQL errors,
  making failures visible in both API responses and server logs
- Add primary_domain migration for schema consistency across fresh
  and upgraded installations (backfill from legacy domain column)
- Use primary_domain in non-cluster SSL query branch for schema
  compatibility
- Surface SSL fetch errors in frontend via toast notifications
- Harden connection cleanup in error handlers with try/except
- Add QEMU + Buildx for linux/amd64,linux/arm64 multi-platform
  Docker image builds
- Update GitHub Actions to latest versions (checkout v4, login v3,
  build-push v6)

Made-with: Cursor
2026-03-30 10:42:28 +03:00
taylanbakircioglu 5b59accd51 fix: resolve asyncpg AmbiguousParameterError on keepalive params
asyncpg cannot infer the PostgreSQL type for $20/$21 when the Python
value is None in CASE WHEN $20 IS NOT NULL expressions. Add explicit
::text cast so asyncpg can prepare the statement regardless of whether
the agent sends keepalive data or not. This was causing ALL heartbeats
to fail with 500 error, making all agents appear offline.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-02-18 20:41:17 +03:00
taylanbakircioglu 4082152312 fix: resolve stale keepalive state when keepalived stops
- Fix DB: COALESCE(NULLIF) prevented clearing stale MASTER/BACKUP state.
  Use CASE WHEN IS NOT NULL to distinguish absent fields (old agents)
  from empty fields (keepalived stopped) and properly clear to NULL.
- Fix Redis: explicitly delete cache key when keepalived stops instead
  of relying on 90s TTL expiry, preventing stale dashboard data.
- Fix agent scripts: SKIP_TO_DAEMON send_heartbeat now always sends
  keepalive_state/keepalive_ip fields (even empty) so backend can
  detect stopped keepalived and clear stale data.
- Fix legacy heartbeat endpoint: POST /{agent_id}/heartbeat now also
  updates keepalive_state and keepalive_ip fields.
- UI: add horizontal scroll to ClusterManagement table to prevent
  column overflow with new Keepalive column.
- UI: add VIP tooltip to Dashboard AgentStatusCard keepalive tag.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-02-18 20:41:17 +03:00
taylanbakircioglu b0feca3c2e feat: add keepalived VRRP state (MASTER/BACKUP) detection to agent heartbeat
- Add keepalive_state and keepalive_ip columns to agents table (migration + schema)
- Add keepalive fields to AgentHeartbeat Pydantic model (backward compatible)
- Update heartbeat endpoint to persist keepalive data to DB and cache in Redis
- Add multi-method keepalived detection in agent scripts (journalctl, log files, VIP check)
- Update dashboard-stats agents/status API with Redis-first keepalive lookup
- Update GET /api/agents to include keepalive_state and keepalive_ip
- Show MASTER/BACKUP tag in Dashboard AgentStatusCard
- Show keepalive info in Agent Management registered agents table
- Add Keepalive column to Cluster Management table with VIP search support

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-02-18 20:41:17 +03:00
taylanbakircioglu 1f100d01e7 fix: SSL reject rollback, expiry calculation, and UI improvements
- SSL Reject: Remove conditional has_same_update_applied check, always
  rollback since Auto-Reject handles cross-cluster consistency
- SSL Rollback: Restore expiry_date by parsing ISO string back to datetime
- SSL UI: Add In Use filter toggle, fix Expiring Soon count using actual
  date calculation, smart Private Key content validation messages
- Remove emojis from log messages

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-02-17 16:36:27 +03:00
taylanbakircioglu 7dedd0f573 feat: SSL sync status tracking and smart reload optimization
- Add per-cluster SSL sync status tracking with APPLYING/SYNCED/PARTIAL states
- Add timeout detection (10min) for stalled SSL deployments showing PARTIAL status
- Add SSL-specific info in Apply Management dialog (cluster count, auto-apply notice)
- Add Deployment Status tab in SSL details showing per-cluster agent sync progress
- Implement SSL auto-reject (reject propagates to all clusters) and auto-undo
- Add conditional entity rollback for SSL reject (time-window based safety check)
- Filter REJECTED status from SSL last_config_status display
- Add pending_cluster_names to SSL list API response
- Optimize agent SSL reload: only trigger HAProxy reload when changed cert is
  actually referenced in the cluster's HAProxy config (grep -qF check), preventing
  unnecessary reloads on clusters that don't use the updated certificate
- Applied to all 4 SSL deploy paths: Linux/macOS standalone and daemon embedded modes

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-02-17 16:36:27 +03:00
taylanbakircioglu 9b8016588c feat: show SSL certificate usage status (In Use / Not In Use)
Replace the always-Active status badge with actual usage information by
querying which frontends and backend servers reference each certificate.

Backend: detail endpoint returns used_by_frontends, used_by_servers and
usage_count; list endpoint includes usage_count via scalar subqueries
checking both ssl_certificate_id and ssl_certificate_ids JSONB fields.

Frontend: General Info badge shows In Use/Not In Use, new Usage tab with
searchable frontend and server tables, summary card shows In Use count.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-02-17 16:36:27 +03:00
taylanbakircioglu 7ca036ef3c fix: add missing is_active field to SSL certificate API responses
The detail endpoint (GET /api/ssl/certificates/{id}) was missing is_active
in its SQL SELECT/GROUP BY, causing the frontend details modal to always
show "Inactive". The list endpoint included is_active in its SQL query but
omitted it from the response dict, causing the Active counter card to
always show 0.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-02-17 16:36:27 +03:00
taylanbakircioglu 4c7a6d79f3 fix: apply sync for empty config, UI responsiveness improvements
- Fix backend cleanup order: delete backend_servers before backends
  to prevent orphan records and foreign key constraint violations
- Fix Apply progress getting stuck at "0/55" when all entities deleted
  by using verifyRealAgentSync result for accurate syncedCount
- Add completion condition for "all entities deleted" scenario
- Fix cluster selector truncation with dynamic width calculation
- Fix sidebar menu label truncation: increase sider width to 240px,
  use concise menu labels, add CSS overflow handling
- Add responsive breakpoints for header title and content padding
- Auto-collapse/expand sidebar on breakpoint change

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-02-16 16:45:29 +03:00
taylanbakircioglu 6349f1ad7c fix: exclude disabled (OFF) agents from sync calculations
Disabled agents were blocking apply sync progress indefinitely because
they were counted in total_agents but could never report as synced.

Backend:
- agent-sync endpoint: disabled agents excluded from total/synced/unsynced
  counts, added disabled_agents and total_agents_including_disabled fields
- SSL cert agent-sync endpoint: same disabled agent exclusion
- Both endpoints still return disabled agents in the list with
  sync_excluded: true for UI display

Frontend:
- ApplyManagement: cluster sync complete logic handles 0 enabled agents,
  all updateEntityCounts calls pass disabled count, agent table shows
  OFF/Excluded tags for disabled agents
- GlobalProgress: shows "(X off)" indicator, visible even when all
  agents are disabled
- ProgressContext: agentCounts state supports disabled field
- agentSync utility: verifyRealAgentSync treats null/0-agent sync_status
  as synced (nothing to wait for)

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-02-16 16:45:28 +03:00
taylanbakircioglu 995d87df43 fix: defaults extraction captures only first defaults section
CRITICAL BUG FIX:
- Previous awk pattern allowed multiple defaults sections to be captured
- Pattern `!/^defaults/` meant "don't exit if line IS defaults"
- This caused duplicate defaults when config had listen before defaults

New pattern uses `started` flag:
- First `defaults` line: set started=1, begin capturing
- Second `defaults` line: started is set, EXIT immediately
- Any `frontend/backend/listen`: started is set, EXIT

Also includes: debug logging for validation error storage with fallback

Tested scenarios:
- Normal config (defaults → listen → frontend): ✓
- Listen before defaults: ✓
- Two defaults sections: ✓ Only first captured
- No defaults section: ✓ Empty output
- Defaults at end of file: ✓
- Empty defaults section: ✓
2026-01-26 15:25:15 +03:00
taylanbakircioglu 83204c86a8 fix: critical security and UX improvements for config management
SECURITY FIX (agent.py):
- Ensure ONLY APPLIED versions are sent to agents
- Fixed fallback query that could return PENDING versions
- Agents will never receive unapproved configurations

UX FIX (cluster.py):
- Clear validation_error when creating new version via Apply
- Prevents stale validation errors from showing after re-apply

UX FIX (agent.py):
- Clear validation_error when agent successfully applies config
- Clear last_validation_error on agent record after success
2026-01-26 15:25:15 +03:00
taylanbakircioglu e75bfdf706 feat: dynamic HAProxy binary path from cluster configuration
- /api/agents/{name}/config now returns haproxy_bin_path, haproxy_config_path, stats_socket_path from cluster
- Agent daemon uses these dynamic paths from API response for validation
- Fallback to local config file if API values not present
- Allows cluster admin to change paths without reinstalling agents

This fixes validation when HAProxy binary is in non-standard location
2026-01-26 15:25:15 +03:00
taylanbakircioglu 97e32114e0 fix: translate Turkish error messages to English
- Updated SUGGESTION_TEMPLATES in haproxy_error_parser.py
- Fixed fallback error messages in cluster.py
- Product language should be English throughout
2026-01-26 15:25:15 +03:00
taylanbakircioglu 095c682fa3 fix: handle soft-deleted entities in bulk import and frontend deletion
Bulk Import Fix (config.py):
- Check for existing servers (including soft-deleted) before INSERT
- If server exists: UPDATE and reactivate instead of INSERT
- Create snapshot for ALL updated servers (enables reject/rollback)
- Add backend_name to UPDATE for field consistency
- Track updated/reactivated servers in response
- Prevents duplicate key constraint violation on re-import

Frontend Deletion Fix (frontend.py):
- Add hard delete support for already soft-deleted frontends
- Follows same pattern as backend.py hard delete
- Deletes WAF associations and config versions before frontend
- Works with existing orphan cleanup in Apply/Reject flows

Both fixes are safe because:
- Orphan version cleanup handles hard-deleted entities in Apply/Reject
- Consistent with existing backend.py patterns
- Snapshot-based rollback fully supported
2026-01-26 15:25:15 +03:00
taylanbakircioglu 5947caa46d fix: move uninstall scripts to backend/utils/agent_scripts for deployment
- Copy uninstall-agent-linux.sh and uninstall-agent-macos.sh to
  backend/utils/agent_scripts/ (same location as install scripts)
- Update generate-uninstall-script endpoint to use the same path
  pattern as working install script endpoint
- Add container-specific fallback paths for robustness
- Add debug logging to track which path is used

Fixes 404 error when fetching uninstall script in production
where /utils/ directory at project root is not deployed.
2026-01-26 15:25:15 +03:00
taylanbakircioglu 8076f3fcfd feat: add intelligent HAProxy validation error display in UI
- Add haproxy_error_parser.py: Parses HAProxy validation errors with
  confidence scoring, extracts entity type/name, line number, error type
- Add ValidationErrorModal.js: Rich modal with parsed error summary,
  quick fix suggestions, and manual troubleshooting guide
- Update cluster.py: Integrate error parser into agent-sync and
  config-versions endpoints with graceful fallback
- Update ApplyManagement.js: Add validation error banner with quick
  navigation buttons and error detail modal
- Update FrontendManagement.js & BackendServers.js: Handle URL params
  for deep-linking to entity edit forms with field highlighting

Enables users to see actionable validation failure details directly
in the UI without needing server access for debugging.
2026-01-26 15:25:15 +03:00
taylanbakircioglu 0c1d68eb01 feat: add uninstall script UI with modern design
- Add new API endpoint to serve uninstall scripts by platform
- Display uninstall script alongside install script in setup wizard
- Add dedicated delete agent modal with 2-step workflow
- Modern UI with gradient banners, platform icons, and info cards
- Enhanced uninstall scripts to clean all agent temp/backup files
- HAProxy service and config remain untouched during uninstall
2026-01-26 15:25:15 +03:00
taylanbakircioglu 851377aedf feat: Add HAProxy proxy name collision prevention system
- Add preserved_listen_blocks column to agents table for storing agent's local listen block names
- Implement reserved names check (stats, monitoring, admin, etc.) for frontend/backend creation
- Add dynamic collision detection against agent's preserved listen blocks
- Apply collision checks to CREATE, UPDATE endpoints and bulk import
- Add debug mode for failed config validation (saves to /tmp/haproxy-failed-*.cfg)
- Fix JSON character stripping for ACL and use_backend rules
- Remove collision protection from agent scripts (now handled by backend)
- All collision checks wrapped in try-except for backwards compatibility
2026-01-26 15:25:15 +03:00
taylanbakircioglu 5d054f3426 feat: Add random and first balance methods support
- Add 'random' and 'first' options to backend balance method selector
- Add balance method validation in config parser with warning for unknown methods
- Update API documentation with all supported balance algorithms
2026-01-26 15:25:15 +03:00
taylanbakircioglu bc8563d455 fix: Update agent token association on config change and improve Security UI
- Fix token-agent relationship not updating when agent config changes
- Agent's api_key now syncs with DB on heartbeat when using different token
- Change Security page badge color from red to blue for better UX
2026-01-26 15:25:15 +03:00
Taylan Bakırcıoğlu 1d17dd8fd5 fix(reject): prevent deletion of existing entities on bulk import reject
CRITICAL DATA LOSS BUG FIX

Problem:
- User updates existing entity via bulk import
- User clicks Reject
- EXPECTED: Entity reverts to old values
- ACTUAL: Entity is COMPLETELY DELETED!

Root Cause:
Reject logic tracked ALL bulk import entities for force deletion.
Did not distinguish between CREATE and UPDATE operations.

Fix:
Added operation check - only track CREATE operations for force deletion.
UPDATE operations are rolled back, not deleted.

Impact:
- CREATE scenarios: No change in behavior
- UPDATE scenarios: Data loss prevented
- Mixed CREATE+UPDATE: Only CREATE entities deleted
- Rollback failures: Safety mechanism preserved

Files Changed:
- backend/routers/cluster.py: Line 4494 added operation check
2025-11-18 21:58:05 +03:00
Taylan Bakırcıoğlu ac5ba4dc58 fix(ssl): add SSL advanced options to GET /api/frontends SELECT queries
CRITICAL BUG FIX: SSL advanced options were not being returned by GET API

Problem:
- Database has ssl_alpn, ssl_npn, ssl_ciphers, etc. columns 
- Response builder tries to access them (Line 341-347) 
- BUT SELECT statements did NOT include these fields 
- Result: f.get('ssl_alpn') returned None for all frontends

Impact:
- Frontend Edit modal always showed empty SSL advanced options fields
- User edits would overwrite existing values with NULL
- Data loss on every frontend edit!

Solution:
- Added all 7 SSL advanced options to ALL 6 SELECT queries:
  1. cluster_id filter (line 174)
  2. cluster_id fallback (line 190)
  3. global include_inactive=True (line 219)
  4. global include_inactive=False (line 231)
  5. global fallback include_inactive=True (line 246)
  6. global fallback include_inactive=False (line 258)

Testing:
- Added debug console.log for SSL advanced options (line 551-559)
- After deployment, check browser console for 'SSL ADVANCED OPTIONS DEBUG'
- Should now show: ssl_alpn: 'h2,http/1.1' etc.

Files Changed:
- backend/routers/frontend.py: All 6 SELECT statements
- frontend/src/components/FrontendManagement.js: Debug logging
2025-11-18 21:58:05 +03:00
Taylan Bakırcıoğlu e8690245af fix(ssl): add change detection and validators for SSL advanced options
Critical additions:
1. Change Detection (backend/routers/config.py)
   - Added SSL advanced options to bulk import change detection logic
   - Detects changes in ssl_alpn, ssl_npn, ssl_ciphers, ssl_ciphersuites
   - Detects changes in ssl_min_ver, ssl_max_ver, ssl_strict_sni
   - Changes will now appear in version diff

2. Pydantic Validators (backend/models/frontend.py, backend/models/backend.py)
   - TLS version validator: Only allows valid versions (SSLv3, TLSv1.0-1.3)
   - ALPN protocol validator: Only allows h2, http/1.1, http/1.0, h2c, spdy/*
   - NPN protocol validator: Only allows http/1.1, http/1.0, spdy/*
   - Prevents invalid values from being saved to database
   - HAProxy validation will not fail due to invalid SSL options

Impact:
- Bulk import will correctly detect SSL option changes
- Apply Management diff will show SSL changes
- User cannot enter invalid TLS versions or protocols
- Improved UX with early validation errors

Previous fixes in this series:
- Frontend GET API: Added SSL fields to response
- Bulk Import UPDATE: Added SSL fields to UPDATE statement
- Backend Server GET API: Added SSL fields to response

Test: Bulk import with ALPN change → Should see change in version diff
2025-11-18 21:58:05 +03:00
Taylan Bakırcıoğlu 3e2f3e3d9f fix(ssl): comprehensive SSL advanced options handling across all endpoints
Critical fixes for SSL advanced options (alpn, npn, ciphers, ciphersuites, min-ver, max-ver, strict-sni):

1. Frontend GET API (backend/routers/frontend.py)
   - Added all 7 SSL advanced option fields to response
   - Frontend Edit modal will now display these fields
   - Prevents NULL overwrite when user edits frontend

2. Bulk Import Frontend UPDATE (backend/routers/config.py)
   - Added all 7 SSL advanced option fields to UPDATE statement
   - Previously skipped with comment 'MVP: DON'T update SSL settings'
   - Now bulk import re-runs preserve SSL settings

3. Backend Server GET API (backend/routers/backend.py)
   - Added 4 server SSL fields (ssl_sni, ssl_min_ver, ssl_max_ver, ssl_ciphers)
   - SQL SELECT query updated
   - Response object updated
   - Backend Server Edit modal will now display these fields

Root cause: GET APIs were not returning SSL advanced options, causing:
- UI forms to show empty fields
- User edits to overwrite with NULL
- Data loss on subsequent updates

All other endpoints (POST, PUT, INSERT) were already correct.

Impact:
- No more accidental data loss when editing frontends/servers
- Bulk import now preserves SSL advanced options
- UI will correctly display all SSL parameters

Test: Deploy backend, run bulk import, verify ssl_alpn appears in database
2025-11-18 21:58:05 +03:00
Taylan Bakırcıoğlu 415439cf17 fix: Move pool_id auto-heal logic before server_statuses check
- Previously auto-heal only ran when agent sent server_statuses (which is empty on first heartbeat)
- Now auto-heal runs on EVERY heartbeat, ensuring pool_id is fixed immediately
- This fixes agents being stuck with NULL pool_id when haproxy config is empty
- Agents will now appear in UI correctly even before applying first config
2025-11-18 21:58:05 +03:00
Taylan Bakırcıoğlu ba09eb70b9 fix: Include pool_id from cluster_id in placeholder-token agent registration path
- Modified auto-registration when using placeholder token to query pool_id from cluster_id in heartbeat
- Previously pool_id was copied from placeholder agent (which was NULL)
- Now correctly fetches pool_id from haproxy_clusters table using cluster_id
- Added extensive debug logging for troubleshooting
- Fixes issue where agents were registered with NULL pool_id and not appearing in UI
2025-11-18 21:58:05 +03:00
Taylan Bakırcıoğlu 761e7fae55 debug(agent): add extensive logging for pool_id auto-register troubleshooting
Added debug logs to track:
- cluster_id presence in heartbeat payload
- Auto-register path execution
- Database query results for pool_id lookup
- Missing cluster_id warnings

This will help identify why pool_id remains NULL during registration.
2025-11-18 21:58:05 +03:00
Taylan Bakırcıoğlu 8418025744 fix(agent): auto-register agents with pool_id from cluster_id
PROBLEM:
- Agents registered with pool_id = NULL
- UI doesn't show agents without pool assignment
- Auto-healing only works when server_statuses present
- New agents with empty HAProxy config never get pool_id

ROOT CAUSE:
- Agent auto-register INSERT excluded pool_id column
- Auto-healing (line 1417-1422) only triggers inside 'if server_statuses:' block
- Fresh agents send no server_statuses → auto-healing never runs
- Agent sits with pool_id = NULL indefinitely

ANALYSIS:
- Previous commit 1f9d8c2 added auto-healing for edge case (cluster before pool)
- But auto-healing has conditional prerequisite: server_statuses must exist
- New agents: empty HAProxy config → no server_statuses → no healing
- This was working before because agents had config from start

SOLUTION:
- Extract pool_id from cluster_id during registration
- Query haproxy_clusters for pool_id using heartbeat's cluster_id
- Insert agent with pool_id immediately (no wait for auto-heal)
- Two auto-register paths fixed: with API key + without API key

EDGE CASES HANDLED:
1. cluster_id missing → pool_id = NULL → auto-healing fallback 
2. cluster deleted → pool_id = NULL → auto-healing fallback 
3. cluster pool_id NULL → pool_id = NULL → auto-healing works 
4. Old agent script → no cluster_id → auto-healing works 

COMPATIBILITY:
- Auto-healing preserved (line 1417-1422 untouched)
- Two mechanisms work together harmoniously
- Backward compatible with existing agents
- No breaking changes

IMPACT:
- New agents visible in UI immediately
- Faster registration (no wait for auto-heal)
- Reduced heartbeat cycles to full functionality
- Better user experience

TESTING REQUIRED:
1. Delete agent from database
2. Install fresh agent with cluster_id in config
3. Verify pool_id populated on registration
4. Verify agent appears in UI immediately
5. Verify existing agents unaffected
2025-11-18 21:58:05 +03:00
Taylan Bakırcıoğlu 35dd8e5d13 Fix KeyError: ssl_name in bulk import parse
CRITICAL BUG: P1 fix broke parse response due to dict format change

ROOT CAUSE:
- P1 changed ssl_auto_assigned_frontends dict format
- Old: {'frontend': ..., 'ssl_name': ..., 'ssl_id': ...}
- New: {'frontend': ..., 'matched_certs': [...], 'total_matched': ...}
- But lines 943, 967 still accessed f['ssl_name'] → KeyError!

SOLUTION:
- Line 944: Loop through matched_certs to extract ssl_names
- Line 969: Build matched_info from matched_certs array
- Multi-SSL format now consistent throughout

EXAMPLE OUTPUT:
SSL AUTO-ASSIGNED: 1 frontend(s) automatically matched
Matched: public_ssl (3 certs: example-cert1, demo-cluster-cert, example-cert3)

IMPACT:
- Parse now completes successfully
- Shows all 3 matched SSL certs in response
- User sees detailed SSL matching info

Ref: demo-cluster parse error
2025-11-18 21:58:05 +03:00
Taylan Bakırcıoğlu 7492d85649 FIX: Multiple SSL certificates & advanced params parsing
CRITICAL BUG FIX: Only 1 SSL cert matched instead of 3 for public_ssl

ROOT CAUSE:
- Parser only stored first cert path (ssl_cert_path)
- Bulk import SSL matching only checked single path
- ssl_certificate_ids = [2] (should be [2,3,4] for 3 certs)
- ssl_alpn = NULL (should be 'h2,http/1.1')

SOLUTION:
1. Added ssl_cert_paths List[str] to ParsedFrontend
2. Parser now stores ALL cert paths from bind directive
3. Bulk import loops through all cert paths for matching
4. SSL advanced options (alpn, npn, ciphers, etc.) included in frontends_data

EXAMPLE BIND DIRECTIVE:
bind 0.0.0.0:8443 ssl
  crt /etc/ssl/certs/example-cert1.pem
  crt /etc/ssl/certs/demo-cluster-cert.pem
  crt /etc/ssl/certs/example-cert3.pem
  alpn h2,http/1.1

BEFORE:
- ssl_certificate_ids: [2]  (only first cert)
- ssl_alpn: NULL            (not passed to bulk import)

AFTER:
- ssl_certificate_ids: [2, 3, 4]  (all 3 certs matched)
- ssl_alpn: 'h2,http/1.1'         (parsed & stored)

IMPACT:
- Multi-SSL frontends correctly imported
- SSL advanced params preserved & editable in UI
- SNI-based routing works correctly
- HTTP/2 ALPN negotiation preserved

TESTING:
- Parser: 3 crt paths → ssl_cert_paths = [path1, path2, path3]
- Matching: 3 paths → 3 IDs (if all SYNCED)
- Database: ssl_certificate_ids JSONB = [2,3,4]

Ref: demo-cluster public_ssl frontend issue
2025-11-18 21:58:05 +03:00
Taylan Bakırcıoğlu 001b467e42 FIX: Rejected bulk import data corruption
CRITICAL BUG FIX: Rejected bulk import entities were marked APPLIED

ROOT CAUSE:
- Bulk import versions (bulk-import-*) don't match entity-ID regex
- Final cleanup incorrectly marked all PENDING entities as APPLIED
- 594 entities (56 FE + 55 BE + 483 SRV) corrupted in production

SOLUTION:
1. Detect bulk import versions (bulk-import-*, restore-*)
2. Extract entity IDs from metadata bulk_snapshots
3. Verify rollback deleted entities (should be 0 remaining)
4. If rollback failed, force DELETE entities
5. Skip final cleanup for bulk import (entities already handled)

IMPACT:
- Prevents data corruption on rejected bulk imports
- Detects and auto-fixes failed rollbacks
- Maintains data integrity for multi-entity operations
- Preserves existing orphan cleanup for single-entity changes

TESTING:
- Bulk import + reject: Entities deleted (not APPLIED)
- Normal entity + reject: Status updated to APPLIED (rollback)
- Orphan entities: Cleaned only if safe (no bulk versions)

Ref: demo-cluster bulk-import-1763450049 issue
2025-11-18 21:58:05 +03:00
Taylan Bakırcıoğlu a739dd0f95 PRODUCTION FIX v2: Direct JSON Sanitization in Heartbeat Endpoint
CRITICAL ISSUE:
Previous middleware approach failed - Starlette middleware cannot
reliably modify request body after it's consumed by FastAPI.

NEW APPROACH - ENDPOINT-LEVEL SANITIZATION:
Moved JSON sanitization directly into agent heartbeat endpoint for
guaranteed execution before Pydantic validation.

ROOT CAUSE CONFIRMED:
Agent daemon mode sends: "server_statuses": ,
This is INVALID JSON (empty value before comma)
Result: JSON decode error at position 322 -> agents stuck offline

SOLUTION IMPLEMENTATION:
Modified: backend/routers/agent.py
- Read raw request body BEFORE Pydantic processing
- Apply 3 regex fixes:
  1. "field": , -> "field": null,
  2. "field": } -> "field": null}
  3. {field,} -> {field}
- Parse sanitized JSON manually
- Create AgentHeartbeat from clean dict
- Continue with normal heartbeat flow

BENEFITS:
✓ NO AGENT SCRIPT CHANGES (production safe)
✓ Guaranteed execution (not middleware dependent)
✓ Detailed logging of sanitization
✓ Graceful error handling
✓ Backward compatible with all agents
✓ Zero impact on valid JSON

LOGGING:
INFO: "Sanitized malformed JSON from agent 'demo-agent1'"
DEBUG: Shows before/after JSON (first 300 chars)

PRODUCTION IMPACT:
- demo-agent1 & agent3 will go online immediately
- No agent restart required
- No agent script update required
- Self-healing for future similar issues

TESTING:
Deploy backend -> Watch logs for:
"Sanitized malformed JSON from agent"

Removed:
- backend/middleware/json_sanitizer.py (approach failed)

This direct approach guarantees the fix executes BEFORE
FastAPI/Pydantic validation, solving the agent offline issue.
2025-11-18 21:58:05 +03:00
Taylan Bakırcıoğlu a6b223c0e7 CRITICAL FIX: Enhanced Error Handling for Agent Heartbeat JSON Parse Errors
PRODUCTION STABILITY FIX - Detailed Logging for Malformed Agent Payloads

PROBLEM:
- demo-agent1 and demo-agent2 sending malformed JSON
- Error: JSON decode error at body position 322
- No visibility into WHAT is malformed or WHY
- Impossible to debug without raw payload inspection

ROOT CAUSE:
- FastAPI consumes request body before error handler
- Pydantic validation fails but does not show raw input
- Agent script might be generating invalid JSON
- No logging of actual problematic payload

SOLUTION - ENHANCED ERROR HANDLING:

1. MIDDLEWARE ENHANCEMENT (error_handler.py):
   - Extract RAW body in validation error handler
   - Parse agent name from JSON (even if malformed)
   - Log first 500 chars of problematic payload
   - Add body size to error details
   - Special handling for /heartbeat endpoint
   - Detailed logging for json_invalid errors

2. HEARTBEAT ENDPOINT (agent.py):
   - Added Request parameter for raw body access
   - Enhanced docstring with troubleshooting info

BENEFITS:
- Instant visibility into malformed JSON
- Agent name logged even on parse failure
- Exact payload position + preview
- No performance impact (only on errors)
- Backward compatible (does not change API)

NEXT STEPS (After Deploy):
1. Check logs for CRITICAL JSON PARSE ERROR
2. Identify exact field causing parse failure
3. Fix agent script if needed
4. Or fix backend to be more tolerant

This enables root cause analysis without SSH access to agent servers
2025-11-18 21:58:05 +03:00
Taylan Bakırcıoğlu 2819eeb515 Fix Backend API endpoints for SSL Advanced Options - Part 2
CRITICAL FIX: All API endpoints now handle SSL advanced options

 FRONTEND ENDPOINTS FIXED:
- create_frontend() - INSERT statement now includes all 7 SSL advanced fields
- update_frontend() - UPDATE statement now includes all 7 SSL advanced fields
- Bulk import frontend INSERT - All SSL params included

 BACKEND SERVER ENDPOINTS FIXED:
- add_server_to_backend() - INSERT now includes 4 SSL server advanced fields
- update_server() - allowed_fields list now includes SSL params (dynamic update)
- Bulk import server INSERT - All SSL params included

FIXED FIELDS:
Frontend: ssl_alpn, ssl_npn, ssl_ciphers, ssl_ciphersuites, ssl_min_ver, ssl_max_ver, ssl_strict_sni
Server: ssl_sni, ssl_min_ver, ssl_max_ver, ssl_ciphers

IMPACT:
- Bulk import now correctly saves parsed SSL parameters to database
- Frontend edit/create now accepts SSL params from UI
- Server edit/create now accepts SSL params from UI
- All CRUD operations fully support SSL advanced options

NEXT: Frontend UI components (FrontendManagement.js & BackendServers.js)
2025-11-18 21:58:05 +03:00
taylanbakircioglu 550995ec5b refactor: Remove unused pending_config_request flag from heartbeat response
CLEANUP: Remove pending_config_request logic from heartbeat

BACKGROUND:
This flag was added to heartbeat response to help agents detect pending
configuration requests immediately, without waiting for the next polling cycle.

REASON FOR REMOVAL:
- Agent scripts already check for pending config requests every 30 seconds
- Adding this flag to heartbeat response is redundant
- Creates unnecessary database queries on every heartbeat
- No performance benefit in practice

HEARTBEAT RESPONSE SIMPLIFIED:
- Removed pending_config_request field
- Kept only status, message, and agent_id
- Simpler API contract
- Reduced database load

BENEFITS:
- Simpler heartbeat API
- Fewer database queries (performance improvement)
- Agent polling mechanism already works well
- No change in agent behavior (agents don't use this flag)

PRODUCTION IMPACT:
- No breaking changes
- Agents continue working normally
- Configuration updates still work via polling
- Reduced database load
2025-11-17 20:22:30 +03:00
taylanbakircioglu e1fecde331 fix: Remove misleading upgrade completion heartbeat causing agent restart loop
CRITICAL PRODUCTION BUG: Agent stuck in restart loop after upgrade

SYMPTOMS:
- Agents continuously restarting every ~30 seconds
- Log shows: "Sending upgrade completion heartbeat..."
- Log shows: "Agent upgrade completed successfully"
- systemd restarts agent immediately after
- Agents never reach daemon loop
- Configuration updates not received
- Entity updates not applied

ROOT CAUSE:
- Agent script v1.0.10 had misleading "upgrade completion heartbeat"
- This heartbeat was sent EVERY time daemon started
- After sending, script would exit (expecting systemd restart)
- systemd would restart agent → infinite loop
- Agent never reached check_agent_upgrade() or check_config_updates()

MISLEADING CODE (REMOVED):

SOLUTION:
- Removed "upgrade completion heartbeat" from daemon startup
- Agent sends normal heartbeat in daemon loop (every 30s)
- No special "upgrade completion" needed
- Agent stays in daemon mode continuously
- systemd only restarts on actual failures

IMPACT:
- Agents no longer restart in loop
- Configuration updates work normally
- Entity updates applied successfully
- Upgrade process works correctly
- Production stability restored
2025-11-17 20:21:40 +03:00