fix(monitor): collapse repeated host-metric alerts into per-window summary (F-11) (#1175)

* fix(monitor): collapse repeated host-metric alerts into per-window summary (F-11)

A host metric over threshold previously dispatched one notification every 5
minutes for the duration of the breach, producing 7+ identical messages
in 35 minutes and spamming Discord/Slack routes. Replace the hardcoded
5-minute cooldown for CPU/RAM/disk with a per-metric suppression window
(default 60 min, configurable via host_alert_suppression_mins). The first
breach fires immediately; subsequent cycles within the window are silently
counted; the next dispatch after the window elapses carries a summary
suffix listing how many cycles were suppressed and when the breach first
crossed threshold. Recovery clears the counter so re-breach fires fresh.

The pattern mirrors PolicyEnforcement.notifyTrivyMissingOnce: module-scope
Map, in-memory only, in-cycle dedup, with a test-reset helper. The
existing system_state row keeps post-restart re-fires bounded.

Janitor and per-stack alert rules are unchanged; they already have
adequate cadence and per-rule cooldown respectively.

* fix(ci): restore backend and frontend checks

* fix(e2e): remove create button timing race

* fix(e2e): harden create double-click test

* fix(monitor): clear persisted F-11 timestamp on recovery + clamp suppression window

Independent audit on the previous commit surfaced two issues.

1. clearHostMetricSuppression early-returned on missing in-memory state,
   leaving a stale system_state.last_host_*_alert_ts row alive after a
   process restart. Scenario: breach fires + persists timestamp, process
   restarts, metric recovers before another evaluate cycle re-seeds the
   in-memory Map, recovery cleanup early-returns. Next re-breach inside
   the original window hits the restart-survivability branch and is
   silently suppressed instead of firing fresh. Fix: read persisted
   state in clearHostMetricSuppression and reset to '0' independently
   of in-memory presence. The read-before-write also skips redundant
   writes when the row is already cleared.

2. host_alert_suppression_mins is validated by zod on the bulk PATCH
   path but the single-key POST /api/settings path accepts allowlisted
   keys without re-validation. A 999999999-minute value would silence
   host alerts for centuries. Add MAX_HOST_ALERT_SUPPRESSION_MIN = 1440
   mirroring the zod max, and clamp via Math.min in evaluateGlobalSettings.

Two new vitest cases (restart-then-recovery-then-rebreach; the 1440
clamp) confirmed failing before the fix, passing after. The existing
"metric drop" case updated to use a mock-backed persistence pattern
consistent with the new restart-scenario tests. 73/73 monitor-service
tests green; full backend suite 2507/2510 (same pre-existing Windows
EBUSY flake on filesystem-backup.test.ts as baseline).
This commit is contained in:
Anso
2026-05-23 06:28:01 -04:00
committed by GitHub
parent e46f6980f8
commit fcff8e9047
8 changed files with 560 additions and 30 deletions
+33 -16
View File
@@ -58,17 +58,35 @@ test.describe('Stack management', () => {
await loginAs(page);
await waitForStacksLoaded(page);
// Count POST /api/stacks calls and hold the response briefly so the
// disabled-button window is observable from the test. Awaiting
// route.continue() avoids the fire-and-forget pattern that can subtly
// stall a request on a busy CI runner.
let postCount = 0;
let releaseCreate: () => void = () => undefined;
let markFirstPostSeen: () => void = () => undefined;
let releasedCreate = false;
const createMayContinue = new Promise<void>((resolve) => {
releaseCreate = () => {
if (releasedCreate) return;
releasedCreate = true;
resolve();
};
});
const firstPostSeen = new Promise<void>((resolve) => {
markFirstPostSeen = resolve;
});
// Count POST /api/stacks calls and hold the first response until after a
// second click has been attempted. Continuing the original route after an
// assertion failure can race Playwright cleanup, so fetch and fulfill it
// explicitly once the test is ready to let the request complete.
await page.route('**/api/stacks', async (route) => {
if (route.request().method() === 'POST') {
postCount += 1;
await new Promise((resolve) => setTimeout(resolve, 400));
markFirstPostSeen();
await createMayContinue;
const response = await route.fetch();
await route.fulfill({ response });
return;
}
await route.continue();
await route.fallback();
});
try {
@@ -78,22 +96,21 @@ test.describe('Stack management', () => {
const createBtn = page.locator('[role="dialog"]').getByRole('button', { name: /^Create/ });
await createBtn.click();
await firstPostSeen;
// The synchronous useRef guard plus React's disabled re-render must
// surface a disabled Create button while the POST is in flight. If this
// assertion ever times out, the busy-state path has regressed.
await expect(createBtn).toBeDisabled({ timeout: 1_500 });
// Best-effort second click against the now-disabled native button. A real
// user mash translates to a no-op at the browser level on a disabled
// submit button; the synchronous ref guard catches it even if Playwright
// managed to dispatch the click anyway.
// Best-effort second click while the first POST is still in flight. A
// real user mash translates to a no-op once the button disables; the
// synchronous ref guard catches it even if Playwright manages to dispatch
// the click before React commits the disabled state.
await createBtn.click({ force: true, timeout: 500 }).catch(() => undefined);
await page.waitForTimeout(100);
expect(postCount).toBe(1);
releaseCreate();
await expect(page.getByRole('dialog')).toBeHidden({ timeout: 8_000 });
await expect(page.getByText(`Stack "${stackName}" created.`)).toBeVisible({ timeout: 5_000 });
expect(postCount).toBe(1);
} finally {
releaseCreate();
await page.unroute('**/api/stacks');
// Cleanup
await page.evaluate(async (name) => {