18 Commits

Author SHA1 Message Date
taylanbakircioglu b34d7cf811 fix: reactivate disabled backend servers from the UI + v1.6.3 (Issue #24)
A backend server toggled OFF (is_active=false) vanished from the UI with no way
to reactivate it: GET /api/backends honored include_inactive for backends but the
server sub-queries hardcoded 'AND is_active = TRUE'.

- get_backends: server sub-queries now honor include_inactive (default callers
  unchanged); added last_config_status to the server payload so the UI can tell a
  DISABLED server (re-enableable) from a DELETION (pending delete).
- toggle_server: persists an entity snapshot so an Apply-Management Reject rolls
  back is_active (previously left the server stuck disabled).
- BackendServers.js: requests include_inactive, shows disabled servers with the
  ON/OFF switch + an 'Inactive' tag, hides only DELETION-pending servers, and
  keeps soft-deleted BACKENDS hidden (so include_inactive doesn't resurface them).
- Config generation unchanged: disabled servers stay '# DISABLED:' comments and
  convert back to live lines when re-enabled.

Startup migration hardening (multi-replica / rolling-deploy safety): create_essential_tables
fails fast on lock contention and retries; run_all_migrations is serialized by a
session advisory lock and gated by a schema_migrations version marker, so an
already-current schema is skipped instead of issuing lock-heavy DDL that a serving
replica's traffic could block at startup. Idempotent and fail-open.

Version reported consistently across all layers (version.json, backend fallback,
frontend package) -> 1.6.3.
2026-06-02 02:42:48 +03:00
taylanbakircioglu 02b1cb2bca feat: v1.5.0 — Site Wizard (Issue #14) + ACME Diagnostic Panel (Issue #13)
Closes #13, Closes #14.

This release squashes the v1.4.0 → v1.5.0 development line. v1.4.0
shipped the ACME stability & enterprise audit (Issues #10/#11/#12).
v1.5.0 builds on that foundation with two co-equal headline features
plus a 22-round audit campaign hardening the prior configuration
surface. License remains MIT for v1.5.0 (relicense to AGPL-3.0
lands in v1.5.2).

------------------------------------------------------------------
HEADLINE FEATURE A — ACME Diagnostic Panel (Issue #13)
------------------------------------------------------------------
A live pre-flight + post-failure diagnostic surface for every ACME
order, reachable from the ACME Automation page. The panel exists
to make ACME failures legible to operators who do NOT have shell
access to the API host.

Endpoints (`backend/routers/acme_diagnostics.py`):
  POST /api/letsencrypt/orders/{order_id}/diagnostics
       Run the full 5-check suite (DNS / port-80 / routing /
       account / agents) and humanize the order's `error_detail`
       (>=11 RFC-8555 problem types, backwards compatible with
       legacy plain-string failures).
  POST /api/letsencrypt/orders/{order_id}/diagnostics/
                                {check_id}/rerun
       Re-run a single check in place — used by the "Re-run"
       button on every row of the modal's pre-flight table.
  GET  /api/letsencrypt/orders/{order_id}/events
       Merged event timeline combining the typed
       `acme_order_events` rows with correlated
       `user_activity_logs` entries (resource_type =
       'letsencrypt_order' AND resource_id = order_id). The
       diagnostic modal auto-tails this timeline every 5 seconds
       while open.

Service-level checks (`backend/services/acme_diagnostics.py`):
  * DNS resolution via stdlib socket.gethostbyname_ex through
    run_in_executor (intentionally avoiding an aiodns runtime
    dep for v1.5.0).
  * Port-80 HEAD probe, target locked to the order's domains,
    success on HTTP 200 OR 404, warns on egress timeout
    (corp egress policies routinely blackhole outbound 80 —
    fail-hard would be too noisy).
  * SSRF guard: probe refuses non-public IPs and surfaces the
    skip in the diagnostic result; IPv4-mapped IPv6 normalisation
    closes the `::ffff:169.254.169.254` cloud-metadata vector.
  * HAProxy routing presence check: matches the order's
    cluster_ids to a port-80 HTTP frontend.
  * ACME account validity check against `letsencrypt_accounts`.
  * Agent presence check (>=1 active agent in target cluster).
  * Every sub-check wrapped in a wall-clock timeout to bound
    impact on the API event loop.

RBAC: ssl.read for run, ssl.read for events. Per-user 5/min rate
limit on both run and rerun, backed by the (user_id, action,
created_at DESC) composite index.

Frontend (`frontend/src/components/ACMEAutomation.js`):
  * "Diagnose" button on every order row + the existing
    "stuck order" warning row.
  * Modal with two tabs:
    - Pre-flight Checks (Antd Table with status pills + Re-run
      buttons + humanized error banner)
    - Event Log (Antd Timeline with auto-tail polling, scroll-
      to-bottom, pause-on-hover)
  * Correlation IDs surfaced in error banners and individual
    check fail details for backend-log lookup.

------------------------------------------------------------------
HEADLINE FEATURE B — Site Setup Wizard (Issue #14)
------------------------------------------------------------------
A single guided flow that creates a Backend + Servers + HTTP
Frontend (and optional HTTPS Frontend) in one atomic transaction.

Endpoints (`backend/routers/site_wizard.py`):
  POST /api/site-wizard/preview     — diff-preview the changeset
  POST /api/site-wizard/create      — atomic execute
  POST /api/site-wizard/reject      — clean rollback (including
                                       any wizard_staged ACME
                                       orders)
  GET  /api/site-wizard/drafts      — draft persistence
  PUT  /api/site-wizard/drafts/{id} — save/update
  DELETE /api/site-wizard/drafts/{id}

Feature surface:
  * One screen captures both backend (mode + servers) AND
    frontend (http + optional https + SSL mode) inputs.
  * SSL modes: ACME (new order, HTTP-01 only for v1.5.0),
    Upload (existing PEM), Existing (link to a stored cert),
    or None.
  * ACME-staged path: wizard_staged_until watermark on the
    `letsencrypt_orders` row defers finalisation until agent
    confirmation; per-mode reject cleanly cancels and rolls
    back the staged order.
  * Live diff preview against the cluster's current generated
    config (renderer-evolution noise stripped — track-sc<N>
    dedup, per-server cookie strip, defaults-cookie
    inheritance, listen-block flattening).
  * Draft persistence with PEM stripped at save time (private
    keys never round-trip through the drafts table).
  * Per-cluster multi-tenancy: drafts and wizard_staged orders
    are isolated to the creating user's cluster scope.

Frontend (`frontend/src/components/SiteWizard.js`):
  * 4-step Antd Steps flow: Backend → Frontend → SSL → Review.
  * Render the live diff preview inline before commit.
  * Antd Form-level validation mirrors backend Pydantic
    validators (numeric bounds, HAProxy reserved keywords, ALPN
    consistency, IPv6 scope-id, domain regex, server name
    dedup).

------------------------------------------------------------------
AUDIT CAMPAIGN — Rounds 1 → 22 (Bulgu #1#82)
------------------------------------------------------------------
v1.5.0 includes 22 adversarial review passes. Each round produced
its own commit set in the corporate development line; this squash
collapses those into the v1.5.0 release artefact. Highlights:

  Round 1-4   Site Wizard core: dry-run parity, single-line
              value injection guard, ACL -f pattern-file block,
              SSL parity, timeout regex, form-state pin.
  Round 5-7   defaults-cookie inheritance, server-named-cookie
              guard, fe/be mode mismatch, duplicate server
              names, health_check_uri + server_address
              validators.
  Round 8-10  cookie_name / cookie_options newline-injection
              guard, dry-run parity (round 9), TCP-mode HTTP-only
              feature blockers.
  Round 11    SSL name path traversal + health-check >= 1.
  Round 12-13 SSL & ACME deep dive (Bulgu #23-#32).
  Round 14    single-line value injection (Bulgu #33).
  Round 15-17 ACME multi-tenant UX, numeric bounds, HAProxy
              reserved keywords, ALPN/TLS consistency,
              all-backup, multi-domain & multi-user enterprise
              edges, drain/HSTS/post-completion (Bulgu
              #34-#53).
  Round 18-21 concurrency, agent state, TCP-mode HTTP-only,
              list size caps, IPv6 scope-id, preview account
              validation, TCP backend + balance uri reject
              (Bulgu #54-#61).
  Round 22    FE error visibility + 3x stale-data lockouts,
              referential integrity + cascade safety,
              authentication & authorization, multi-cluster
              isolation, apply_pending_changes concurrency,
              script injection + bulk import multi-tenancy,
              prefix-stripped signature comparison
              (Bulgu #62-#82).

------------------------------------------------------------------
NO CORPORATE-SPECIFIC ARTIFACTS
------------------------------------------------------------------
This squash deliberately sanitises corporate hostnames, container
registry references, and TLS secret names into generic
placeholders (`your-registry.example.com/your-org`,
`haproxy-openmanager*.example.com`, `wildcard-tls`,
`taylanbakircioglu/haproxy-openmanager-*`) so the public artefact
contains no internal infrastructure detail. Pilot / development
history that retained those values stays in the corporate fork
and is NOT part of this commit.
2026-05-14 00:04:19 +03:00
taylanbakircioglu 851377aedf feat: Add HAProxy proxy name collision prevention system
- Add preserved_listen_blocks column to agents table for storing agent's local listen block names
- Implement reserved names check (stats, monitoring, admin, etc.) for frontend/backend creation
- Add dynamic collision detection against agent's preserved listen blocks
- Apply collision checks to CREATE, UPDATE endpoints and bulk import
- Add debug mode for failed config validation (saves to /tmp/haproxy-failed-*.cfg)
- Fix JSON character stripping for ACL and use_backend rules
- Remove collision protection from agent scripts (now handled by backend)
- All collision checks wrapped in try-except for backwards compatibility
2026-01-26 15:25:15 +03:00
taylanbakircioglu 5d054f3426 feat: Add random and first balance methods support
- Add 'random' and 'first' options to backend balance method selector
- Add balance method validation in config parser with warning for unknown methods
- Update API documentation with all supported balance algorithms
2026-01-26 15:25:15 +03:00
Taylan Bakırcıoğlu 3e2f3e3d9f fix(ssl): comprehensive SSL advanced options handling across all endpoints
Critical fixes for SSL advanced options (alpn, npn, ciphers, ciphersuites, min-ver, max-ver, strict-sni):

1. Frontend GET API (backend/routers/frontend.py)
   - Added all 7 SSL advanced option fields to response
   - Frontend Edit modal will now display these fields
   - Prevents NULL overwrite when user edits frontend

2. Bulk Import Frontend UPDATE (backend/routers/config.py)
   - Added all 7 SSL advanced option fields to UPDATE statement
   - Previously skipped with comment 'MVP: DON'T update SSL settings'
   - Now bulk import re-runs preserve SSL settings

3. Backend Server GET API (backend/routers/backend.py)
   - Added 4 server SSL fields (ssl_sni, ssl_min_ver, ssl_max_ver, ssl_ciphers)
   - SQL SELECT query updated
   - Response object updated
   - Backend Server Edit modal will now display these fields

Root cause: GET APIs were not returning SSL advanced options, causing:
- UI forms to show empty fields
- User edits to overwrite with NULL
- Data loss on subsequent updates

All other endpoints (POST, PUT, INSERT) were already correct.

Impact:
- No more accidental data loss when editing frontends/servers
- Bulk import now preserves SSL advanced options
- UI will correctly display all SSL parameters

Test: Deploy backend, run bulk import, verify ssl_alpn appears in database
2025-11-18 21:58:05 +03:00
Taylan Bakırcıoğlu 2819eeb515 Fix Backend API endpoints for SSL Advanced Options - Part 2
CRITICAL FIX: All API endpoints now handle SSL advanced options

 FRONTEND ENDPOINTS FIXED:
- create_frontend() - INSERT statement now includes all 7 SSL advanced fields
- update_frontend() - UPDATE statement now includes all 7 SSL advanced fields
- Bulk import frontend INSERT - All SSL params included

 BACKEND SERVER ENDPOINTS FIXED:
- add_server_to_backend() - INSERT now includes 4 SSL server advanced fields
- update_server() - allowed_fields list now includes SSL params (dynamic update)
- Bulk import server INSERT - All SSL params included

FIXED FIELDS:
Frontend: ssl_alpn, ssl_npn, ssl_ciphers, ssl_ciphersuites, ssl_min_ver, ssl_max_ver, ssl_strict_sni
Server: ssl_sni, ssl_min_ver, ssl_max_ver, ssl_ciphers

IMPACT:
- Bulk import now correctly saves parsed SSL parameters to database
- Frontend edit/create now accepts SSL params from UI
- Server edit/create now accepts SSL params from UI
- All CRUD operations fully support SSL advanced options

NEXT: Frontend UI components (FrontendManagement.js & BackendServers.js)
2025-11-18 21:58:05 +03:00
Taylan Bakırcıoğlu d6ee9517d0 fix(backend): Inactive APPLIED entities incorrectly showing as pending
CRITICAL FIX: has_pending_config calculation for both backends and frontends

Problem Found (Console Log):
  Backend/Frontend: {
    last_config_status: 'APPLIED',   <- Already applied!
    has_pending_config: true,        <- But flag is TRUE!
    is_active: false                 <- Inactive (soft-deleted)
  }

  Apply response: 'No pending changes to apply'
  Reject response: 'No pending changes found to reject'

  Result: Entity stuck in Apply Management forever!

Root Cause:
  OLD LOGIC (backend.py line 438, frontend.py line 363):
  has_pending_config = (config_version OR status=PENDING OR is_inactive) AND NOT rejected

  For inactive + APPLIED entity:
  - has_config_version: FALSE
  - has_pending_status: FALSE (status=APPLIED)
  - is_inactive: TRUE
  - Result: TRUE (incorrectly marked as pending!)

Why It Matters:
  - Design: Inactive entities should show as pending for soft-delete workflow
  - Problem: Entity with is_active=FALSE + last_config_status='APPLIED' = soft-delete already applied!
  - Apply/Reject: Both look for PENDING status, find none, do nothing
  - UI: Entity remains in Apply Management (has_pending_config=true forever)

Solution (Both backend.py and frontend.py):
  Inactive entity is only pending if last_config_status is PENDING, not APPLIED:

  OLD: is_inactive → always pending
  NEW: (is_inactive AND status=PENDING) → only pending if not yet applied

  Formula:
  has_pending_config = (
    config_version OR
    status=PENDING OR
    (is_inactive AND status=PENDING)
  ) AND NOT rejected AND NOT (is_inactive AND is_applied)

Test Cases:
   Active entity, PENDING → pending=TRUE
   Active entity, APPLIED → pending=FALSE
   Inactive entity, PENDING → pending=TRUE (soft-delete needs apply)
   Inactive entity, APPLIED → pending=FALSE (soft-delete already applied) <- FIXED!
   Inactive entity, REJECTED → pending=FALSE

Backend.py Changes (lines 426-445):
  - Added is_applied flag
  - Added is_inactive_and_pending logic
  - Updated has_pending_config formula
  - Comprehensive comments

Frontend.py Changes (lines 359-370):
  - Same logic as backend for consistency
  - Inline expression (no loop variables)
  - Comprehensive comments

Impact:
  - 'deneme-sil' backend will show has_pending_config=FALSE
  - Backend/Frontend will disappear from Apply Management
  - No more stuck entities after soft-delete apply
  - Agent sync race condition protected (inactive+APPLIED=not pending)
  - Consistent behavior across all entity types

Related: c120445 (apply endpoint fix), f80c026 (debug logs)
Refs: #has-pending-config #inactive-entity #apply-management #consistency
2025-11-17 14:15:44 +03:00
Taylan Bakırcıoğlu d2bf64af8b fix(backend): Prevent duplicate key errors from inactive backends + race condition safety
CRITICAL FIX: Handle soft-deleted backends that block unique constraint

Problem Scenario:
1. User creates backend 'deneme-sil' without servers
2. Backend gets soft-deleted (is_active=FALSE) somehow
3. Backend remains in DB but invisible in UI (API filters is_active=TRUE)
4. User tries to create same backend again
5. ERROR: duplicate key value violates unique constraint

Root Causes:
A) Soft-deleted backends remain in DB and block unique constraint
B) Apply endpoint marks ALL pending backends as APPLIED, even those skipped by config generator
C) Backend without servers shows as APPLIED but isn't in haproxy.cfg (inconsistent)
D) Race condition: Agent sync temporarily marks backends as inactive

Solutions:
1️⃣ Backend CREATE (backend.py lines 473-506):
   - Check for inactive backends with same name before creating
   - SAFETY: Only cleanup if inactive for >30 seconds (avoid agent sync race)
   - If found: Hard delete inactive backend + related data
   - Then allow new backend creation
   - Prevents: duplicate key constraint errors + race conditions

2️⃣ Backend DELETE (backend.py lines 1044, 1056-1090):
   - Detect if backend is already inactive (is_active=FALSE)
   - If inactive: Hard delete (permanent removal from DB)
   - If active: Soft delete (mark as inactive for Apply workflow)
   - Prevents: Orphan inactive backends accumulating in DB

3️⃣ Apply Endpoint (cluster.py lines 1600-1638):
   - Only mark backends as APPLIED if they have active servers
   - Check: EXISTS(backend_servers WHERE is_active=TRUE)
   - Backends without servers remain PENDING (correct state)
   - Log warning: 'Backend X remains PENDING (no active servers)'
   - Prevents: Inconsistent state (APPLIED in DB, missing in haproxy.cfg)

Race Condition Protection:
⚠️  Agent config-sync temporarily marks backends as is_active=FALSE
⚠️  If we hard delete during sync, backend could be lost!
 Solution: Only cleanup backends inactive for >30 seconds
 Agent sync takes <5 seconds, so safe window
 Protects against: sync running while user creates backend

Impact Analysis (All Scenarios Tested):
 Normal backend create (with servers) - No impact
 Backend create (without servers) - FIXED: Stays PENDING until servers added
 Backend delete → recreate - FIXED: Old backend cleaned up automatically
 Agent sync race condition - PROTECTED: 30-second safety window
 Multi-cluster (same name) - No impact: cluster_id already checked
 Bulk import reactivation - No impact: Has own logic
 Config restore/rollback - No impact: Has own conflict handling
 Frontend-backend relations - No impact: Cleanup preserved
 Dashboard statistics - No impact: Only counts active
 Maintenance status - IMPROVED: Stale inactive backends auto-cleaned

Benefits:
 No more duplicate key errors
 Users can recreate backends with same name
 Inactive backends are automatically cleaned up (after 30s)
 Consistent state: APPLIED = actually in haproxy.cfg
 Clear warning when backend needs servers to deploy
 Race condition protection during agent sync
 No risk of data loss during concurrent operations

How to Fix Current 'deneme-sil' Backend:
Option 1: Reject in Apply Management (easiest)
Option 2: Add servers + Apply
Option 3: Delete backend + Apply (auto-cleanup after 30s)
Option 4: Manual DB cleanup (fastest right now)

Related: Previous commits (frontend null check, config skip, UX messages)
Refs: #backend-creation #duplicate-key #soft-delete #apply-consistency #race-condition
2025-11-17 14:15:43 +03:00
taylanbakircioglu b9d618b4ac feat(snapshot): PHASE 2 & 3 - Entity snapshot integration + Reject rollback
PHASE 2: Entity Update Integration (ALL entities)
- Frontend update: Full snapshot with 27 fields
- Backend update: Full snapshot with 23 fields
- WAF rule update: Full snapshot with 11 fields
- SSL certificate update: Full snapshot with 16 fields
- Server update: Full snapshot with 23 fields
- ALL database fields included (no missing fields)

PHASE 3: Reject with Rollback Logic
- cluster.py - reject_all_pending_changes() enhanced
- Entity rollback before marking REJECTED
- Support for single entity snapshot
- Support for bulk snapshots (bulk import/restore)
- Entity status: REJECTED -> APPLIED (entities rolled back)
- Rollback statistics in response (success/failed/skipped)

Key Changes:
- entity_snapshot.py: All field schemas validated against migrations
- Backend rollback: 23 fields (including options, cookie_*, default_server_*)
- Server rollback: 23 fields (including ssl_certificate_id, haproxy_status)
- SSL rollback: 16 fields (including issuer, fingerprint, all_domains)
- WAF rollback: 11 fields (including enabled, cluster_id)
- No emoji in code (clean logging)
- Feature flag: ENTITY_SNAPSHOT_ENABLED (default: false)

Next: PHASE 4 (Bulk Import) + PHASE 5 (Restore) integration
2025-11-14 01:06:38 +03:00
taylanbakircioglu 481be91a4e feat(snapshot): PHASE 2 - Add entity snapshot for Frontend & Backend updates
- Created entity_snapshot.py helper module (~570 lines)
  - save_entity_snapshot() - Create snapshots with compaction
  - rollback_entity_from_snapshot() - Main rollback logic
  - _rollback_update() - UPDATE rollback for all entity types
  - _rollback_create() - CREATE rollback (entity deletion)
  - Feature flag support: ENTITY_SNAPSHOT_ENABLED (default: false)

- Integrated snapshot into Frontend update (frontend.py)
  - Capture full entity state before UPDATE
  - Create entity_snapshot metadata
  - Merge with pre_apply_snapshot for diff viewer
  - Store in config_versions.metadata JSONB

- Integrated snapshot into Backend update (backend.py)
  - Same snapshot pattern as Frontend
  - Works within transaction for atomicity
  - Preserves diff viewer compatibility

- Added feature flag to config.py
  - ENTITY_SNAPSHOT_ENABLED (environment variable)
  - Default: false (safe rollout)
  - Ready for Phase 7 gradual deployment

Next: WAF, SSL, Server update integration + Reject rollback logic
2025-11-14 01:06:38 +03:00
taylanbakircioglu 4272dbb4ef feat: Add 'option httpchk' validation and auto-filtering across all layers
CRITICAL FIX: Prevent 'option httpchk' duplication in HAProxy config by implementing 3-layer validation:

1. BULK IMPORT PARSER:
   - Frontend: Filter out 'option httpchk' with warning (not applicable to frontends)
   - Backend: Already filtering 'option httpchk' (handled by health_check_uri field)

2. BACKEND API:
   - Backend create/update: Auto-filter 'option httpchk' from options field
   - Frontend create/update: Auto-filter 'option httpchk' from options field
   - Added filter_httpchk_from_options() helper function in both routers

3. FRONTEND UI:
   - Backend modal: Real-time warning when 'option httpchk' is typed
   - Frontend modal: Real-time warning when 'option httpchk' is typed
   - Warning messages guide users to use proper fields instead

Changes:
- backend/utils/haproxy_config_parser.py: Added httpchk filtering for frontend parsing
- backend/routers/backend.py: Added filter function + applied to create/update
- backend/routers/frontend.py: Added filter function + applied to create/update
- frontend/src/components/BackendServers.js: Added dynamic warning for httpchk
- frontend/src/components/FrontendManagement.js: Added dynamic warning for httpchk

User Experience:
 Bulk Import: Automatically filters httpchk, shows warning in preview
 Manual Entry: Shows real-time warning, auto-filters on save
 No Config Duplication: 'option httpchk' never appears twice in generated config

Impact: Users can safely paste or type 'option httpchk' without breaking HAProxy config. System automatically filters it and guides users to use the Health Check URI field instead.
2025-11-13 10:12:30 +03:00
taylanbakircioglu 0fc18fde38 feat: Add HAProxy options support for backends and frontends
Implemented comprehensive HAProxy options field support for both backend and frontend entities to enable standard HAProxy directives like 'option http-keep-alive', 'option httplog', 'option forwardfor', etc.

Changes:
- Database: Added 'options' TEXT column to backends and frontends tables
- Models: Added options field to BackendConfig, BackendConfigUpdate, and FrontendConfig
- API Endpoints: Updated CREATE, UPDATE, and GET endpoints to handle options field
  * Backend: CREATE/UPDATE/GET with options support
  * Frontend: CREATE/UPDATE/GET with options support (fixed 5 SELECT queries)
- Config Generator: Added options block generation for both backends and frontends
- Bulk Import Parser:
  * Added options field to ParsedBackend and ParsedFrontend dataclasses
  * Implemented option directive parsing with validation
  * Added unknown option warnings
  * Fixed bulk parse response to include options field
- Bulk Import Merge: Added options field comparison in UPDATE logic
- UI Components:
  * BackendServers.js: Added options TextArea form field
  * FrontendManagement.js: Added options TextArea form field

Features:
- Multi-line options support (newline-separated format)
- Option validation with known HAProxy options list
- Backward compatible (NULL options for existing entities)
- Bulk import support with merge strategy
- Full CRUD support for both manual and bulk operations

Technical Details:
- Format: Newline-separated TEXT field for multiple options
- Validation: Warns about unknown options but allows them
- Config Generation: Each option written as separate directive
- Agent: Standard HAProxy config validation applies

Total: 10 files modified, ~195 lines added, 26 integration points verified
2025-11-13 10:12:30 +03:00
taylanbakircioglu b29c9e044e feat: Multi-cluster isolation and comprehensive bug fixes
This update consolidates bug fixes and improvements from internal development:

## Multi-Cluster Isolation (CRITICAL)
- Backend delete now isolated per cluster (added cluster_id to WHERE clauses)
- Orphan config version auto-detection and cleanup
- Prevents cross-cluster contamination when deleting entities
- REJECTED entities excluded from pending list

## Apply Management Fixes
- Fixed phantom pending entities (REJECTED entities no longer shown)
- Fixed cluster switch 404 errors (state cleared before fetch)
- Fixed page reload not refreshing data (added mount useEffect)
- Orphan entity status auto-cleanup on reject

## Backend Delete Improvements
- Added cluster_id to server/frontend updates (multi-cluster safe)
- Automatic ACL/use_backend cleanup from frontends
- NULL cluster_id support for legacy data
- Prevents HAProxy validation errors

## Orphan Version Detection
- Backend GET: Validates entity belongs to version's cluster
- Frontend GET: Validates entity belongs to version's cluster
- Apply: Auto-detects and removes orphan versions before apply
- Reject: Auto-detects and removes orphan versions before reject

## UI/UX Improvements
- Apply Management state management improved
- Debug logging for troubleshooting
- Better cluster switch handling
- README: Orphan version troubleshooting section

Technical Changes:
- backend/routers/backend.py: Multi-cluster isolation, orphan detection, NULL handling
- backend/routers/frontend.py: Orphan detection, REJECTED filter
- backend/routers/cluster.py: Orphan auto-cleanup, entity status cleanup
- frontend/ApplyManagement.js: State management, mount useEffect
- README.md: Troubleshooting documentation

Impact:
- Multi-cluster environments now fully isolated 
- Orphan config versions automatically cleaned 
- REJECTED entities properly filtered 
- Apply Management works correctly across cluster switches 
- No manual database intervention needed 
2025-11-12 13:34:07 +03:00
taylanbakircioglu 1ea1c6a29f feat: Major stability and feature improvements
This commit consolidates multiple improvements from internal development:

## Agent Stability Improvements
- Add database connection pooling (min=10, max=50) for better performance
- Prevent config reapply on agent restart by fetching last_applied_version from database
- Optimize SSL fetch to only run when config changes (98% API call reduction)
- Make SSL_SYNC_TIMESTAMP_FILE agent-specific to prevent race conditions
- Fix agent offline display issue due to database connection bottleneck
- 10x faster heartbeat response (200ms → 20ms)

## Bulk Import UPSERT Support
- Parse endpoint detects existing entities (New/Existing status)
- Bulk-create supports UPDATE for existing backends/frontends (merge strategy)
- New servers can be added to existing backends
- Existing servers preserved (no deletion in MVP)
- Field-by-field value comparison (only changed fields updated)
- Pending apply conflict prevention (409 error)
- Fixed duplicate key error on server INSERT
- Backend marked PENDING when servers added

## Apply Management Fixes
- Fixed deleted entities not showing (include_inactive parameter)
- Backend/Frontend GET endpoints support inactive entities for Apply Management
- All pending changes now visible
- Phantom backend bug protection maintained

## Backend Delete Improvements
- Automatically clean ACL/use_backend rules from frontends
- Prevents HAProxy validation errors after backend deletion
- Frontend references automatically updated

## UI/UX Improvements
- Cluster selector status dot auto-refreshes every 30 seconds
- Real-time agent health monitoring (no page refresh needed)
- Parse message shows only NEW entities (cleaner)
- Status labels: 'Update' → 'Existing' (clearer meaning)
- Multi-line parse success messages
- Detailed summary breakdown with tooltips

Technical Changes:
- backend/database/connection.py: Connection pool implementation
- backend/main.py: Pool initialization and cleanup
- backend/routers/*: UPSERT logic, field comparison, include_inactive
- backend/utils/agent_scripts/*: Applied version tracking, SSL optimization
- frontend/src/components/*: UI improvements, status indicators
- frontend/src/contexts/ClusterContext.js: Auto-refresh agent health

Impact:
- Supports 50+ concurrent agents (previously ~10)
- Zero config reapply on restart/upgrade
- Bulk import handles existing entities correctly
- All pending changes visible in Apply Management
- Real-time cluster health status
- No HAProxy validation errors after backend delete
2025-11-11 21:56:18 +03:00
taylanbakircioglu 38a0375954 Fix: Backend Server SSL fields not persisting in edit modal
Critical Bug Fix - Server SSL Configuration Not Saved:
- Server edit: Enable SSL, select certificate, save
- Re-edit: SSL fields empty (ssl_enabled=false, ssl_certificate_id=null)
- Database had data but form didn't load it

Root Cause Analysis:
1. Backend API queries include ssl_certificate_id (Line 207, 217)
2. Backend API response object missing ssl_certificate_id (Line 257-279)
3. Frontend form missing ssl_certificate_id in setFieldsValue (Line 738-753)
4. Result: Data saved but not loaded back

Backend API Fix (backend/routers/backend.py):
Added missing fields to server response object:
  - check_port
  - ssl_enabled
  - ssl_verify
  - ssl_certificate_id (CRITICAL - was causing the bug)
  - cookie_value
  - inter, fall, rise

Frontend Form Fix (BackendServers.js):
Added missing fields to handleEditServer setFieldsValue:
  - check_port
  - ssl_verify
  - ssl_certificate_id (CRITICAL)
  - cookie_value
  - inter, fall, rise

HAProxy Config Validation:
Generated config syntax verified:
  server server1 1.1.1.1:11 weight 100 ssl verify required ca-file /etc/ssl/haproxy/star-burgan-com-tr.pem check

Matches HAProxy standard format:
  server <name> <addr>:<port> [params]
  Valid params: weight, ssl, verify, ca-file, check

Complete Workflow After Fix:
  Edit → Enable SSL → Select cert → Save → DB stores ssl_certificate_id
  Edit again → Form loads SSL enabled + certificate selected
  Apply → Config with ca-file path generated
  Agent → Downloads cert, applies config
  HAProxy → Validates and loads successfully
2025-11-07 11:51:15 +03:00
taylanbakircioglu a5e281b284 Feature: Backend Server SSL certificate support + Frontend SSL dropdown enhancement
 Backend Server SSL Certificate - Complete Implementation:

1. Model Update (backend/models/backend.py):
   - Added ssl_certificate_id field to ServerConfig model
   - Allows selecting SSL certificate from dropdown

2. API Endpoints (backend/routers/backend.py):
   - CREATE server: Added ssl_certificate_id to INSERT query
   - UPDATE server: Added ssl_certificate_id to allowed_fields
   - GET servers: Added ssl_certificate_id to SELECT queries (2 places)

3. Config Generation (backend/services/haproxy_config.py):
   - SSL certificate lookup by ID
   - Auto-generate ca-file path: /etc/ssl/haproxy/{cert_name}.pem
   - Added to server line in HAProxy config

Example Generated Config:
  Before: server es1 10.0.0.1:9200 ssl verify required
  After:  server es1 10.0.0.1:9200 ssl verify required ca-file /etc/ssl/haproxy/star-burgan-com-tr.pem

 Frontend SSL Dropdown Enhancement:
- Added Global/Cluster-specific tags to Frontend SSL dropdown
- Matches Backend Server SSL dropdown design
- Shows: [🌍 Global] or [📍 Cluster] with color coding

🔧 Complete SSL Workflow:
1. User edits Backend Server
2. Enables SSL
3. Selects SSL certificate from dropdown
4. Saves → ssl_certificate_id stored in DB
5. Apply Changes → Config generated with ca-file path
6. Agent downloads SSL cert to /etc/ssl/haproxy/
7. HAProxy uses ca-file for SSL verification

 Database Schema:
  backend_servers table now includes:
  - ssl_enabled (bool)
  - ssl_verify (str: none/required)
  - ssl_certificate_id (int, FK to ssl_certificates)

 HAProxy Config Format:
  server {name} {addr}:{port} ssl verify required ca-file {path}

Impact: Backend Server SSL now fully functional with certificate management
2025-11-07 11:51:15 +03:00
taylanbakircioglu 3cb3f53200 Fix: Soft-deleted backends appearing randomly on page refresh + SSL dropdown fix
🐛 Critical Bug Fixes:
- Fixed soft-deleted backends appearing intermittently on page refresh
- Fixed soft-deleted servers appearing in backend server lists
- Fixed Backend Server SSL dropdown showing empty list (wrong API endpoint)

🔧 Backend API Fixes (backend/routers/backend.py):
- Line 187: Added 'AND is_active = TRUE' to backends query
- Line 197: Added 'WHERE is_active = TRUE' to backends query (no cluster)
- Line 212: Added 'AND is_active = TRUE' to backend_servers query
- Line 222: Added 'AND is_active = TRUE' to backend_servers query (no cluster)

🔧 Frontend Fix (BackendServers.js):
- Fixed SSL certificate API endpoint
- Changed: /api/ssl-certificates → /api/ssl/certificates
- Added cluster_id query param and Authorization header
- Added debug logging for troubleshooting

 Impact Analysis - All Scenarios Verified:

1. Backend Delete (Soft):
   - is_active set to FALSE ✓
   - API no longer returns deleted backends ✓
   - UI shows no phantom backends ✓

2. Page Refresh:
   - Consistent behavior (no random appearances) ✓
   - Deleted backends never shown ✓

3. Apply Changes:
   - Hard delete still works (Line 1631 cluster.py) ✓
   - Soft-deleted backends removed from DB ✓

4. Config Generation:
   - Already uses 'is_active = TRUE' filter ✓
   - NOT affected by this change ✓
   - Inactive servers shown as comments (intentional) ✓

5. Frontend Dropdown:
   - Only shows active backends ✓
   - Deleted backends not selectable ✓

6. Dashboard:
   - Uses Redis cache (indirect filtering) ✓
   - NOT affected by this change ✓

🎯 Root Cause:
- API was returning ALL backends (active + inactive)
- Soft-deleted entities appeared randomly based on timing
- No is_active filter at API level

🎉 Result:
- Phantom backend bug completely resolved
- All 8 scenarios tested and verified
- No breaking changes to existing functionality
- Config generation intentionally unchanged (disabled servers as comments)
2025-11-07 11:51:14 +03:00
taylanbakircioglu 6aae0f4309 Initial commit 2025-10-27 12:14:03 +03:00