mirror of
https://github.com/Studio-Saelix/sencho.git
synced 2026-08-10 18:56:53 +00:00
fix(dashboard): harden real-time dashboard with bug fixes and design compliance (#517)
* fix(dashboard): harden real-time dashboard with bug fixes and design compliance - Fix host alert spam: add 5-minute cooldown for CPU/RAM/disk threshold alerts, preventing duplicate notifications every 30s during sustained breaches. Extract shared dispatchWithCooldown helper (also used by Docker janitor alerts). - Fix memory metric inflation: subtract filesystem cache from stored memory_mb values, matching the existing calculateMemoryPercent logic. - Fix crash detection reliability: replace fragile 'seconds ago' string matching with a tracked Set of alerted container IDs. Containers are only alerted once per crash event, with automatic cleanup when they start running again or after a 1-hour TTL. - Fix health status bar: exited containers now trigger 'degraded' state independently of unread error notifications. - Fix CPU chart Y-axis: auto-scale when aggregate container CPU exceeds 100% instead of silently clipping at the hardcoded domain ceiling. - Fix grammar: 'actives' to 'active' in container count label. - Add shadow-card-bevel to all dashboard cards per design system. - Update dashboard docs to reflect revised health status thresholds. * test(dashboard): update monitor service tests for new alert signatures - Add container Id fields to crash detection test fixtures - Update host alert assertions to match dispatchWithCooldown 3-arg call - Fix unhealthy container test to use State: 'unhealthy' instead of State: 'running' (running containers are now skipped in crash detect)
This commit is contained in:
@@ -16,8 +16,8 @@ The top bar provides an at-a-glance health assessment for the active node. Sench
|
||||
| Status | Meaning |
|
||||
|--------|---------|
|
||||
| **Healthy** | All systems nominal. No resources above warning thresholds, no unread errors. |
|
||||
| **Degraded** | At least one resource is above 80%, or there are unread error alerts. |
|
||||
| **Critical** | At least one resource is above 90%, or there are exited containers with unread errors. |
|
||||
| **Degraded** | At least one resource is above 80%, there are exited containers, or there are unread error alerts. |
|
||||
| **Critical** | At least one resource is above 90%, or there are exited containers combined with unread errors. |
|
||||
|
||||
The bar also shows the active node name, the number of running containers, and the current alert count.
|
||||
|
||||
|
||||
Reference in New Issue
Block a user