Manage highly-available virtual IPs backed by Keepalived (VRRP) directly from the OpenManager
UI — no more SSHing into nodes to install/configure Keepalived by hand. Builds on the agent
pull-architecture: define the VIP centrally, click Apply, and the agents converge.
Highlights:
- New "HA / VIP" tab: create a virtual IP, pick a per-node interface, select which pool nodes
participate (MASTER/BACKUP roles + priorities); live MASTER/BACKUP/FAULT per node.
- On Apply, agents install & configure Keepalived (unicast VRRP, cloud-safe default) across the
major distros (Debian/Ubuntu, RHEL/CentOS/Alma/Rocky, Fedora, SUSE/openSUSE, Alpine) with a
HAProxy health-check, so the VIP fails over automatically when HAProxy drops.
- Single-node (a managed floating IP without failover) and multi-node VRRP failover both work.
- VIP changes ride the standard Apply Management flow with the standard "View Change" diff.
- Approval-gated deletion (safety): deleting a running VIP is staged for approval and the VIP
keeps running, untouched, until you approve it — an agent never tears a VIP down without an
explicit human approval. Per-VIP Diagnostics view; opt-in package uninstall (only on nodes
where OpenManager installed it). A node already running a hand-managed Keepalived is detected
and never overwritten ("externally managed").
- Fully opt-in and backward compatible: nodes/clusters without a VIP are unaffected. Adds
vip_instances + vip_members tables (idempotent SCHEMA_VERSION bump; existing data and
passwords unaffected) and a `vip` RBAC permission group.
- Also includes a HAProxy config-generator robustness fix: auto-inject a stick-table when a
frontend uses a stick counter (track-sc / sc_*_rate) but declares none.
On-prem / L2 (VRRP) scope; the UI notes the cloud caveat.
6.8 KiB
Upgrade Notes — v1.7.0 (HA / VIP Keepalived management, Issue #27)
Backward compatible & opt-in. Upgrading to v1.7.0 changes nothing for existing clusters/agents until you create a VIP:
- Schema:
SCHEMA_VERSIONbumps to3, so on first start the (idempotent) migration sequence re-runs once and adds two new tables (vip_instances,vip_members) plus an additivevip_instances.applied_snapshotcolumn (enables rejecting a pending VIP change and restoring the previous applied state). No existing table is altered. Existing rows and the admin password are not reset (default users are create-if-missing). The only data effect is that the four built-in system roles (super_admin/operator/security_admin/viewer) are re-seeded to their canonical permission sets plus the newvip.*permissions — this is the long-standing behavior of the role seeder; custom roles are untouched. - Agents: the agent script gains an opt-in keepalived deploy that is a no-op on
any node without an applied VIP, and it never overwrites a hand-managed
/etc/keepalived/keepalived.conf(it reports "externally managed" instead). - Scope: VRRP VIPs target bare-metal / VMware / on-prem L2 networks. On AWS/Azure/GCP the cloud fabric doesn't honor VRRP/gratuitous-ARP; the UI surfaces this. Ensure host firewalls permit VRRP (IP protocol 112).
- Optional env:
VIP_ENCRYPTION_KEY(see.env.template) — if unset, the VRRP secret encryption key is derived fromSECRET_KEY(like MFA).
No rollback steps are required to disable the feature: simply don't create VIPs (or delete them — agents tear down their managed keepalived on the next poll).
Agent Upgrade Guide - Dashboard Stats Fix
Problem
Dashboard showing 0 metrics after agent auto-upgrade due to missing environment variables (SOCAT_BIN, STATS_SOCKET_PATH) when agent restarts in daemon mode.
Solution
Implemented lazy initialization in both get_haproxy_stats_csv() and get_server_statuses() functions. These functions now initialize their dependencies on first call, making them completely independent of global variable initialization.
Deployment Steps
Step 1: Wait for Pipeline ⏳
# Pipeline is currently running after git push
# Check status: https://[your-azure-devops]/pipeline
# Wait for deployment to complete (~3-5 minutes)
Step 2: Verify Backend Deployment ✅
# Check backend logs for successful deployment
kubectl logs -f deployment/backend -n haproxy-manager | head -20
# Expected: New pod started with latest code
Step 3: Update Agent Version in UI 🔄
Option A: Automatic Script Sync (Recommended)
# Get admin token
TOKEN=$(curl -k -s -X POST "https://haproxy-manager.example.com/api/auth/login" \
-H "Content-Type: application/json" \
-d '{"username": "admin", "password": "admin123"}' | jq -r '.access_token')
# Sync scripts from files to database (creates version 1.0.0)
curl -k -X POST "https://haproxy-manager.example.com/api/agents/sync-scripts-from-files" \
-H "Authorization: Bearer $TOKEN" | jq .
# Expected: {"status": "success", "synced": ["linux", "macos"]}
Option B: Manual UI Update
- Go to Agent Management page
- Click Settings → Agent Versions
- For Linux platform:
- Click Edit Script
- Version: Keep current or increment (e.g., 1.0.3)
- Changelog: Add "Fixed stats collection after upgrade"
- Click Save (this syncs file content to database)
- Repeat for macOS platform
Step 4: Upgrade Agents 🚀
Option A: UI (Single Agent)
- Go to Agent Management page
- Select agent (e.g.,
demo-agent) - Click Upgrade button
- Wait 30 seconds for agent to restart
Option B: API (Batch Upgrade)
# Get all agents
curl -k -s -X GET "https://haproxy-manager.example.com/api/agents" \
-H "Authorization: Bearer $TOKEN" | jq '.agents[] | {id, name, version}'
# Upgrade specific agent (replace {agent_id})
curl -k -X POST "https://haproxy-manager.example.com/api/agents/{agent_id}/upgrade" \
-H "Authorization: Bearer $TOKEN" | jq .
Step 5: Verify Stats Collection 📊
Backend Logs (30 seconds after upgrade):
kubectl logs -f deployment/backend -n haproxy-manager | grep -E "haproxy_stats_csv|demo-agent"
# Expected logs:
# ✅ "Has haproxy_stats_csv: True"
# ✅ "CSV preview: IyBweG..." (base64 data)
# ✅ "STATS: Parsed 15 rows for cluster demo-cluster1"
Agent Logs (on agent server):
sudo tail -f /var/log/haproxy-agent/agent.log | grep STATS
# Expected logs:
# ✅ "STATS: Initialized socat: /usr/bin/socat"
# ✅ "STATS: Using default socket path: /var/run/haproxy/admin.sock"
# ✅ "STATS: Fetched CSV: 2121 bytes, 9 lines, base64: 2828 chars"
Dashboard UI:
- Open Dashboard page
- Refresh page (F5)
- Check metrics:
- ✅ Frontend/Backend filters populated
- ✅ Overview metrics showing real data (not 0)
- ✅ Charts showing data points
- ✅ "Waiting for Agent Data" warning gone
Troubleshooting
Issue: Backend still shows "Has haproxy_stats_csv: False"
Cause: Agent hasn't upgraded yet or using old script version
Solution:
# Check agent version on agent server
grep "AGENT_VERSION" /usr/local/bin/haproxy-agent | head -1
# Force agent restart
sudo systemctl restart haproxy-agent
# Check logs immediately
sudo tail -20 /var/log/haproxy-agent/agent.log
Issue: "STATS: socat not available"
Cause: socat not installed
Solution:
# Install socat
sudo yum install -y socat # RHEL/CentOS
sudo apt install -y socat # Debian/Ubuntu
# Restart agent
sudo systemctl restart haproxy-agent
Issue: "STATS: Socket not found: /var/run/haproxy/admin.sock"
Cause: HAProxy stats socket not configured
Solution:
# Check HAProxy config for stats socket
grep "stats socket" /etc/haproxy/haproxy.cfg
# Add if missing (in global section):
# stats socket /var/run/haproxy/admin.sock mode 666 level admin
# Reload HAProxy
sudo systemctl reload haproxy
Verification Checklist
- Pipeline completed successfully
- Backend pod restarted with new code
- Agent scripts synced to database (version visible in UI)
- Agents upgraded to new version
- Backend logs show "Has haproxy_stats_csv: True"
- Agent logs show STATS messages
- Dashboard showing real metrics
- Charts populated with data
- No "Waiting for Agent Data" warning
Next Upgrades
This fix is permanent. Future agent upgrades will NOT break stats collection because:
- ✅ Functions are self-contained with lazy initialization
- ✅ No dependency on global variable initialization order
- ✅ Works in any restart scenario (systemd, daemon mode, upgrade)
- ✅ Backward compatible with existing agents
No manual intervention needed for future upgrades! 🎉