mirror of
https://github.com/rcourtman/Pulse.git
synced 2026-10-04 05:07:30 +00:00
5fd05efa83
A Proxmox host wedged on a ZFS deadlock yesterday took the cluster API poll with it (context deadline exceeded). The unified connections aggregator flipped the Connection from active to stale to unreachable, and the Settings / Infrastructure page rendered the right badges, but no top-nav alert ever fired because nothing was actively notifying off that derived state. Patrol's deterministic triage flagged it every minute, but its LLM investigation stage has been broken since 2026-02-26 so flags never escalated into user-visible findings. Result: a 3 hour outage I only noticed because I happened to open Settings. This wires an active notification off the same connection state the Settings badges already use: - internal/alerts/connection.go: new CheckConnection + clearConnectionDegradedAlert that fire connection-degraded after three consecutive stale or unreachable observations. Severity scales: stale warning, unreachable / unauthorized critical. Clear runs through the same recovery-confirmation gate as clearNodeOfflineAlert so a single flap back to active doesn't silently resolve a real outage. Paused, disabled, and non-platform connections are no-ops. - internal/api/connections_alerts.go: snapshot translator that turns api.Connection into the narrow alerts.ConnectionSnapshot view. Keeping the snapshot type inside the alerts package preserves the existing api -> monitoring import direction; the monitor would have cycled if it called back into api directly. - internal/monitoring: new SetConnectionsSnapshotLister hook + a per-tick checkConnectionAlerts call in the main poll loop, alongside the existing evaluate*Agents passes. - internal/api/router.go: register the lister closure on r.monitor so the alerts loop sees the same Connection rows the HTTP handler does. - internal/alerts/specs/types.go: add "connection" to the migration bridge list of accepted ResourceTypes, alongside node / docker-host / proxmox-disk / etc. The connection concept doesn't have a canonical unified resource type yet; this matches the existing pattern for alert-keyed resources that aren't first-class canonical. Test coverage in internal/alerts/connection_test.go covers active never fires, three stale observations escalate from pending to warning, unreachable escalates warning to critical, unauthorized fires critical cold, paused / disabled / agent never fire, recovery confirmation gate, and a stale flap during recovery resets the gate. TestResourceAlertSpecValidateAllowsConnectionMigrationBridgeType mirrors the existing migration-bridge proof tests for the new type.