diff --git a/docs/qualification/PATROL_ASSISTANT_CUSTOMER_JOURNEY.md b/docs/qualification/PATROL_ASSISTANT_CUSTOMER_JOURNEY.md index d17fda6e5..151103333 100644 --- a/docs/qualification/PATROL_ASSISTANT_CUSTOMER_JOURNEY.md +++ b/docs/qualification/PATROL_ASSISTANT_CUSTOMER_JOURNEY.md @@ -1965,3 +1965,52 @@ The autonomous provider refusal remains enforced. Integration receipts are in `tmp/patrol-history-integration/`. The full integration hook gates its merge commit and push. All previously recorded model and wider-readiness gaps remain open. + +## Legacy recorder retirement, 2026-09-06 + +The disconnected incident coordinator, five-second cached-metrics sampler, +pre-incident buffers, archive writer and unused adapters are removed. They had +no production alert trigger. Canonical resource history remains the primary +incident evidence for Assistant. No replacement diagnosis or scheduling policy +was added. + +Explicit archive lookup now requires exact organization, resource and window +binding. The reader is lazy and read-only. Saved file contents, modification +time, mode, old observations, metadata and summary values survive reads. Missing +archives, malformed files and missing windows remain distinct outcomes. Legacy +`recording` status is historical, and the response discloses that the old +`summary.duration_ms` field contains nanoseconds. An old file is never rewritten +to make its evidence appear current. + +The incidents API now reports `active_count: null` with +`active_count_status: not_measured`. The retired coordinator's empty map never +established a measured zero. Its legacy incident-memory listing still needs a +canonical query design covering aliases, canonical-only events, honest bounds +and propagated projection-read errors. This is recorded as an open modernization +residual rather than treating that listing as complete. + +Final-source registered archive-tool receipts pass Playwright at `/patrol`, +1440x1000, 900x1000 and 390x1000. The five cases cover a saved observation, +unavailable/malformed archives, the wrong resource and a missing window. +Verification includes hover/focus, Enter expansion, exact tool input/output, +deep scrolling, Space collapse, Escape, reload and controlled session reopening. +Actual pixels were inspected. Incoming main alert dispatch wording also passes +its isolated real Overview browser script at all three widths. These controlled +responses prove rendering, not model diagnosis, installed delivery or server +persistence. + +The final worker Pro binary is +`bd29e6f27be7b3ad4cfbc37842f4da90f08c6a48c9fc23b12c9c597b67346c9b`. +After the managed local restart, canonical and legacy queries still return the +same seven retained homelab records with the original fired/resolved events +unchanged. The live incidents API reports an unmeasured count. Cached provider +refusal remains enforced. No model request, paid spend, provider retry, +production collector change or fault injection occurred in this slice. + +Archive, tools, chat, AI runtime and targeted API/race checks passed on the +worker. One full API run as root invalidated its mode-bit persistence-failure +fixture. That fixture passes unchanged under the normal worker account. The +full API rerun passes under that account (286.960s), as do the incoming +startup-replay and legacy-boundary source checks. Frontend type checks and all +29 incoming alert tests pass. The exact staged hook gates landing. +Private receipts are under `tmp/patrol-archive-retirement/` in the workspace. diff --git a/docs/release-control/v6/internal/status.json b/docs/release-control/v6/internal/status.json index 4c33faa3a..3796c3558 100644 --- a/docs/release-control/v6/internal/status.json +++ b/docs/release-control/v6/internal/status.json @@ -10201,7 +10201,7 @@ }, { "id": "patrol-assistant-customer-outcome-qualification", - "summary": "The redesign goal remains open, with its executable plan and detailed receipts in docs/qualification/PATROL_ASSISTANT_CUSTOMER_JOURNEY.md. Patrol owns investigation and Assistant continues the same issue. Observations, hypotheses, accepted proposals, execution and independently verified outcomes remain distinct. The recorded baseline contains 127 paid installations, 71 Patrol-enabled, 23 with Assistant calls and fourteen verified resolutions from one installation. Schema17 outcome/provider/cost fields had no adoption. These do not establish representative success, false-alarm or missed-problem rates. Shared provenance/history/risk fixes and removal of proposal-as-proof and proxy completion policy landed through PR1928/1929. PR1934 merged canonical tool/transcript identity. PR1935 contains measurement-presence, capacity, canonical config-read, command-connectivity and tmpfs collection/query corrections through 6e18777d30f30b498def39d30016a697cabc4ea7. Exact worker hooks and named browser matrices passed, with remote landing pending. Real ordinary storage diagnosis failed by ruling out pressure without filesystem capacity evidence. Same-session recovery correctly identified present health and alert resolution but overstated continuous control health. Independent fault-intact, recovery, two-pass cleanup and persisted transcript checks passed. The current incident history correction replaces the primary legacy recorder read with the existing organization-pinned canonical resource timeline, preserving source/time semantics, bounded history, empty coverage and failed-read distinctions. Explicit legacy archive reads remain resource-bound. Its registered-tool SQLite, full tools package, focused race and final-source browser proofs pass. The primary history read was pushed as 580a246 in PR1935. The shared monitor/store identity correction now joins exact full Docker references with canonical container history through an organization-scoped history-only alias index. Original event content and approval/operator authority remain unchanged. Callback, removal/restart, replay, tenant/control and registered-tool tests, full affected packages, race checks and scoped lookup benchmarks passed. Final worker-build API proof returns the same seven retained homelab records under canonical and legacy queries, preserving the original fired/resolved events exactly. Six-case Playwright proof at desktop/intermediate/mobile widths passed. The exact thirteen-file worker hook passed and 919331d5 was pushed through PR1935. Integration with main 11a8cc2180 preserves both contract changes and passes affected backend, race, frontend and final-build browser/API proofs. The full integration hook gates the merge commit. The unused legacy recorder/coordinator startup needs retirement while preserving saved archives. Installed tmpfs collector and real-model interpretation remain unqualified. Other named residuals include unsupported filters, typed compatibility canonical-ID lookup, legacy direction availability, Docker-host history and responsive mount details. Claude Max explicitly refused autonomous Patrol readiness. Cached refusal and API409 enforcement remain enforced. Separate paid-provider approval is pending, with no paid request or policy bypass. Ordinary Assistant is not autonomous qualification. Required local work still includes reliable interpretation, config-read/model retest, storage/backup and approved/rejected action outcomes. Independent volunteered Pro environments remain a separate wider-readiness gate.", + "summary": "The redesign goal remains open, with its executable plan and detailed receipts in docs/qualification/PATROL_ASSISTANT_CUSTOMER_JOURNEY.md. Patrol owns investigation and Assistant continues the same issue. Observations, hypotheses, accepted proposals, execution and independently verified outcomes remain distinct. The recorded baseline contains 127 paid installations, 71 Patrol-enabled, 23 with Assistant calls and fourteen verified resolutions from one installation. Schema17 outcome/provider/cost fields had no adoption. These do not establish representative success, false-alarm or missed-problem rates. Shared provenance/history/risk fixes and removal of proposal-as-proof and proxy completion policy landed through PR1928/1929. PR1934 merged canonical tool/transcript identity. PR1935 contains measurement-presence, capacity, canonical config-read, command-connectivity and tmpfs collection/query corrections through 6e18777d30f30b498def39d30016a697cabc4ea7. Exact worker hooks and named browser matrices passed, with remote landing pending. Real ordinary storage diagnosis failed by ruling out pressure without filesystem capacity evidence. Same-session recovery correctly identified present health and alert resolution but overstated continuous control health. Independent fault-intact, recovery, two-pass cleanup and persisted transcript checks passed. The current incident history correction replaces the primary legacy recorder read with the existing organization-pinned canonical resource timeline, preserving source/time semantics, bounded history, empty coverage and failed-read distinctions. Explicit legacy archive reads remain resource-bound. Its registered-tool SQLite, full tools package, focused race and final-source browser proofs pass. The primary history read was pushed as 580a246 in PR1935. The shared monitor/store identity correction now joins exact full Docker references with canonical container history through an organization-scoped history-only alias index. Original event content and approval/operator authority remain unchanged. Callback, removal/restart, replay, tenant/control and registered-tool tests, full affected packages, race checks and scoped lookup benchmarks passed. Final worker-build API proof returns the same seven retained homelab records under canonical and legacy queries, preserving the original fired/resolved events exactly. Six-case Playwright proof at desktop/intermediate/mobile widths passed. The exact thirteen-file worker hook passed and 919331d5 was pushed through PR1935. Integration with main 11a8cc2180 preserves both contract changes and passes affected backend, race, frontend and final-build browser/API proofs. The full integration hook gates the merge commit. The disconnected legacy recorder/coordinator and its cached-metrics adapter are retired. Explicit read-only archive lookup preserves recorded values and discloses legacy duration units. The old incidents API live count is explicitly unmeasured. Final archive/history/API and race checks pass, including the full API suite under the normal worker account after identifying a root-only permission-fixture failure. Final worker-build browser proof covers five archive outcomes across desktop/intermediate/mobile widths. Live canonical and legacy queries preserve the same seven original homelab records, and the API count is unmeasured. The exact staged hook gates landing. The legacy incident-memory listing still needs alias-aware canonical queries, canonical-only events, honest bounds and propagated projection-read failures. Installed tmpfs collector and real-model interpretation remain unqualified. Other named residuals include unsupported filters, typed compatibility canonical-ID lookup, legacy direction availability, Docker-host history and responsive mount details. Claude Max explicitly refused autonomous Patrol readiness. Cached refusal and API409 enforcement remain enforced. Separate paid-provider approval is pending, with no paid request or policy bypass. Ordinary Assistant is not autonomous qualification. Required local work still includes reliable interpretation, config-read/model retest, storage/backup and approved/rejected action outcomes. Independent volunteered Pro environments remain a separate wider-readiness gate.", "owner": "project-owner", "status": "planned", "recorded_at": "2026-09-05", diff --git a/docs/release-control/v6/internal/subsystems/agent-lifecycle.md b/docs/release-control/v6/internal/subsystems/agent-lifecycle.md index 6ff13db8a..2567589c5 100644 --- a/docs/release-control/v6/internal/subsystems/agent-lifecycle.md +++ b/docs/release-control/v6/internal/subsystems/agent-lifecycle.md @@ -7907,3 +7907,9 @@ positive matching evidence, provider replacement/return, write failure and automatic versus unknown-provenance cleanup. State and config tests cover atomic replacement and preservation of lifecycle evidence. These are synthetic local proofs, not reporter confirmation or installed-release resolution of #1930. + +Historical incident archives are explicit reads, not a collector lifecycle. +`internal/api/router.go` no longer starts an incident coordinator or a cached +metrics sampling loop. Organization teardown drops the archive reference without +saving or deleting recordings. Alert observation and recovery continue through +the alert manager and canonical resource timeline. diff --git a/docs/release-control/v6/internal/subsystems/ai-runtime.md b/docs/release-control/v6/internal/subsystems/ai-runtime.md index 7e4b62e2b..cbddc5956 100644 --- a/docs/release-control/v6/internal/subsystems/ai-runtime.md +++ b/docs/release-control/v6/internal/subsystems/ai-runtime.md @@ -8117,3 +8117,30 @@ unchanged pre-existing inventory. The ordinary package run skips live Docker work unless explicitly enabled. A direct fixture restart is teardown and must never be counted as a Pulse approval, execution, rejection or outcome. Model-led and canonical-action qualification remain separate required evidence. + +### Legacy incident recording retirement + +`internal/metrics/incident_archive.go` owns the historical recording format and +explicit, organization-pinned archive reads. `IncidentArchiveProvider` exposes +only a resource-bound window lookup with an error result. The primary +`pulse_knowledge` incident action reads the canonical resource timeline. An +explicit `window_id` reads saved legacy observations and labels their timestamp +and historical-status limits. It must never start a recorder or infer source +freshness from a recording timestamp. The disconnected incident coordinator, +fleet sampling adapter, timer loop, retention writer and duplicate tool adapter +are retired. There is no replacement incident scheduler or diagnosis policy. + +The archive reader preserves saved timestamps, resource labels, metadata and +summary values, including records older than the former retention period. The response explicitly +identifies the nanosecond encoding of the historical `summary.duration_ms` field. It +distinguishes unavailable archives, failed reads and absent exact resource/window +pairs. Proof lives in `internal/metrics/incident_archive_test.go` and the +registered-tool cases in `internal/ai/tools/incident_history_test.go`. + +The legacy incidents listing still exposes incident memory and is not a complete +canonical incident query. Its old sampler-derived `active_count` is now null with +`active_count_status=not_measured`. Canonical-only events, alias-aware listing, +query bounds and projection-read errors remain an explicit modernization gap in +`patrol-assistant-customer-outcome-qualification`. The shared resource timeline +remains the evidence owner. Do not invent another incident lifecycle to repair +this listing. diff --git a/docs/release-control/v6/internal/subsystems/alerts.md b/docs/release-control/v6/internal/subsystems/alerts.md index ed5a3b0e3..a47e21132 100644 --- a/docs/release-control/v6/internal/subsystems/alerts.md +++ b/docs/release-control/v6/internal/subsystems/alerts.md @@ -2647,3 +2647,20 @@ The hook and destinations caller regressions in `useAlertDestinationsTabState.test.tsx` pin ordering and loading ownership. `scripts/check-delivery-health-ordering.mjs` exercises the real caller and card in Chromium with scripted API completions; it is not installed delivery proof. + +### Alert status distinguishes dispatch from destination evidence + +The active alert card renders a valid diagnosis `lastNotified` timestamp as +“Dispatch requested”, never “Notified”: the alert manager records this field +before invoking delivery callbacks. A cooldown's `nextEligibleAt` is labelled +“next eligible”, not a promised send time. Neither field proves destination +acceptance or recipient receipt; that evidence must not be inferred from the +muted presentation tone. Missing or invalid timestamps retain the existing +pending/cooldown fallback; acknowledged alerts retain their badge without a +second status line. No API field, notification policy or shared primitive changes. + +The existing wrapping status text must remain readable at desktop and phone +widths despite the longer labels. The presentation and Overview delivery-status +tests cover the evidence boundary; `scripts/check-alert-dispatch-copy.mjs` +qualifies the real Overview with scripted API data in Chromium, not installed +notification delivery. diff --git a/docs/release-control/v6/internal/subsystems/api-contracts.md b/docs/release-control/v6/internal/subsystems/api-contracts.md index fcb1d3e1f..25cb1e0a1 100644 --- a/docs/release-control/v6/internal/subsystems/api-contracts.md +++ b/docs/release-control/v6/internal/subsystems/api-contracts.md @@ -10670,3 +10670,18 @@ reasoning and real remediation in `docs/qualification/PATROL_ASSISTANT_CUSTOMER_JOURNEY.md`. The repeatable browser proof is `scripts/check-patrol-assistant-journey.mjs`. A passing scripted response does not establish a useful customer outcome or model qualification. + +### Explicit historical incident archives + +`internal/api/router.go` pins a read-only legacy archive to each organization's +Assistant service. Its native `IncidentArchiveProvider` capability requires both +resource and window identifiers and propagates read errors. No archive setup or +shutdown writes files or starts sampling. The tool preserves stored metadata and +anomalies and marks historical recording status as historical. + +`GET /api/ai/incidents` retains the `active_count` key as null and adds +`active_count_status=not_measured` in every response. Incident memory, an empty +result and unavailable services cannot establish a current count. The old +coordinator never received production alert callbacks, so its zero was not a +measurement. The legacy listing's broader canonical query and read-error +modernization remains open under the customer-outcome qualification gap. diff --git a/docs/release-control/v6/internal/subsystems/frontend-primitives.md b/docs/release-control/v6/internal/subsystems/frontend-primitives.md index df846828d..d383c79f3 100644 --- a/docs/release-control/v6/internal/subsystems/frontend-primitives.md +++ b/docs/release-control/v6/internal/subsystems/frontend-primitives.md @@ -7162,3 +7162,20 @@ confirmation and wrapping controls remain unchanged; no new primitive is added. The focused hook/caller tests and `scripts/check-delivery-health-ordering.mjs` cover this dependency at desktop and narrow widths using scripted health and queue-action responses, without claiming backend notification delivery. + +### Alert status distinguishes dispatch from destination evidence + +The active alert card renders a valid diagnosis `lastNotified` timestamp as +“Dispatch requested”, never “Notified”: the alert manager records this field +before invoking delivery callbacks. A cooldown's `nextEligibleAt` is labelled +“next eligible”, not a promised send time. Neither field proves destination +acceptance or recipient receipt; that evidence must not be inferred from the +muted presentation tone. Missing or invalid timestamps retain the existing +pending/cooldown fallback; acknowledged alerts retain their badge without a +second status line. No API field, notification policy or shared primitive changes. + +The existing wrapping status text must remain readable at desktop and phone +widths despite the longer labels. The presentation and Overview delivery-status +tests cover the evidence boundary; `scripts/check-alert-dispatch-copy.mjs` +qualifies the real Overview with scripted API data in Chromium, not installed +notification delivery. diff --git a/docs/release-control/v6/internal/subsystems/performance-and-scalability.md b/docs/release-control/v6/internal/subsystems/performance-and-scalability.md index 0b215b09d..9556fd2d8 100644 --- a/docs/release-control/v6/internal/subsystems/performance-and-scalability.md +++ b/docs/release-control/v6/internal/subsystems/performance-and-scalability.md @@ -3130,3 +3130,9 @@ candidate on the same worker with alternating samples, and retains full-route and middleware controls. Identical source or instruction sequences alone do not prove identical timing. The recorded final ten-pair check passes the unchanged time/bytes/allocation gate with no adjacent request-path regression. + +The disconnected incident recorder and its fleet metrics adapter are retired. +Router initialization no longer launches their five-second cached-metrics loop +or allocates per-resource pre-incident buffers. Explicit legacy archive reads +are lazy and bounded to the existing 16 MiB file limit. Canonical resource +history supplies current diagnostic evidence without a second sampling loop. diff --git a/docs/release-control/v6/internal/subsystems/registry.json b/docs/release-control/v6/internal/subsystems/registry.json index 5532bfd99..d510933ef 100644 --- a/docs/release-control/v6/internal/subsystems/registry.json +++ b/docs/release-control/v6/internal/subsystems/registry.json @@ -2074,6 +2074,7 @@ "internal/api/ai_intelligence_handlers.go", "internal/config/ai.go", "internal/config/patrol_autopilot_persistence.go", + "internal/metrics/incident_archive.go", "pkg/aicontracts/action_broker.go", "pkg/aicontracts/fix_execution.go", "pkg/aicontracts/investigation.go", @@ -2088,6 +2089,20 @@ "exact_files": [], "require_explicit_path_policy_coverage": true, "path_policies": [ + { + "id": "legacy-incident-archive", + "label": "Explicit read-only legacy recording archive proof", + "match_prefixes": [], + "match_files": [ + "internal/metrics/incident_archive.go" + ], + "allow_same_subsystem_tests": false, + "test_prefixes": [], + "exact_files": [ + "internal/ai/tools/incident_history_test.go", + "internal/metrics/incident_archive_test.go" + ] + }, { "id": "retained-metric-evidence", "label": "retained metric evidence and observation coverage proof", @@ -2190,6 +2205,7 @@ "internal/api/ai_handlers_more_test.go", "internal/api/ai_handlers_patrol_actions_additional_test.go", "internal/api/ai_handlers_test.go", + "internal/api/ai_intelligence_handlers_remediation_additional_test.go", "internal/api/ai_intelligence_handlers_test.go", "internal/api/issue1640_readiness_gate_test.go", "internal/api/issue1640_readiness_transport_test.go", diff --git a/docs/release-control/v6/internal/subsystems/security-privacy.md b/docs/release-control/v6/internal/subsystems/security-privacy.md index bec73958c..3b34d5660 100644 --- a/docs/release-control/v6/internal/subsystems/security-privacy.md +++ b/docs/release-control/v6/internal/subsystems/security-privacy.md @@ -2809,3 +2809,10 @@ optional link to the configured public URL. It never includes resource names, finding text, commands, evidence, or model names, and it uses the tenant's existing email configuration and recipients under the admin-only report schedule routes. + +Explicit legacy incident archive reads use the organization's pinned data path +and exact resource/window identifiers. They retain bounded regular-file and +symlink checks and expose no enumeration, sampling or writing capability. +Removing the disconnected recorder does not alter alert, action approval or +operator authority. An unrelated organization receives no default archive +fallback. diff --git a/docs/release-control/v6/internal/subsystems/storage-recovery.md b/docs/release-control/v6/internal/subsystems/storage-recovery.md index dc956e1c8..65aff7bc4 100644 --- a/docs/release-control/v6/internal/subsystems/storage-recovery.md +++ b/docs/release-control/v6/internal/subsystems/storage-recovery.md @@ -6001,3 +6001,11 @@ Unlike resource reports, a `patrol_digest` schedule run produces no generated file under the tenant `reports` directory and never calls the retention prune; the email body is the only artifact. Recovery and retention state are unaffected. + +Legacy `incident_windows.json` recovery is read-only through +`internal/metrics/incident_archive.go`. Construction does not read the file, and +explicit reads neither expire records nor rewrite contents or permissions. +Malformed, oversized, symlink and non-regular archive paths fail visibly. Missing +archives remain distinguishable from a valid archive with no matching window. +The original recording times and historical status must never establish current +source freshness or active recording. diff --git a/frontend-modern/browser-verification.json b/frontend-modern/browser-verification.json index c9ddf6dc6..379a1e717 100644 --- a/frontend-modern/browser-verification.json +++ b/frontend-modern/browser-verification.json @@ -1,35 +1,39 @@ { "version": 1, - "base_sha": "919331d5b3f6076f8616b07eb8e9ca611f26ee52", - "verified_at": "2026-09-06T15:39:31.729Z", + "base_sha": "ab6d21400074393379a78350249ba4f5412941c0", + "verified_at": "2026-09-06T16:30:47.586Z", "result": "passed", "changed_paths": [ - "frontend-modern/src/features/alerts/useNotificationDeliveryLog.ts" + "frontend-modern/src/features/alerts/deliveryDiagnosisPresentation.ts" ], "content_sha256": { - "frontend-modern/src/features/alerts/useNotificationDeliveryLog.ts": "be919497714d749fcbe9a48a325aa82282b7afa8d6866d84eb1cc09ab52c14a6" + "frontend-modern/src/features/alerts/deliveryDiagnosisPresentation.ts": "bc3eadf8f790517b377430a42eb67e8fdfcadb53bb13b87cd971d6a9ab91d607" }, "backend_content_sha256": { - "internal/ai/tools/executor.go": "c476d6aec1f7925467f1547f0219e6c6265639a62b397649adc82a8ed87b7e4a", - "internal/ai/tools/tools_knowledge.go": "d98ea1bab481acce5eee4988d74d364dc73c4068ca33725d843b54117cc9cc85", - "internal/config/host_continuity.go": "641e7593d1581d6e230c76fd7845b06e648ce24f94b3def6291ec4a38f505259", - "internal/models/models.go": "39684f7535cece1a635648c521cc5057a5ddac764dc6e6ae332eb4287ade9e71", - "internal/monitoring/monitor_agents.go": "52291b9f8521d5794f464147ea61990be654bfc19b9b5da35b6adaf02fdf70e5", - "internal/monitoring/monitor_pbs_pmg.go": "a1986a20cf8c535a13b87ff2d5cbff730df89b74253262b96793d3c44edbf2f7", - "internal/unifiedresources/history_identity.go": "3934ac49acce29b16dd671769719112208be053a8658c678f798645a6a38395d", - "internal/unifiedresources/monitor_adapter.go": "5b51089bac8e85e7c776956c17471599ec15add4e5e8502cd8687c4f788fcdb1", - "internal/unifiedresources/store.go": "018176a4762b06adf52c78128f8351ca93819cbfa10bf2242535fef3d7c59191" + "internal/ai/adapters/adapters.go": "487245a52e1ae85ffecd38c6e4f006bfd6c8efff6568c92c19688a8db56d7ecc", + "internal/ai/chat/service.go": "4f130864717c5141ce974b8a367aeca19d49baa02a4906ecb65591d8b4c02bd6", + "internal/ai/tools/executor.go": "3a30f7720d30be6fdd75f3224d2da937881e77fee76fb819347f83e247fccb65", + "internal/ai/tools/tools_knowledge.go": "12c16ed7893dd49c18a7db1dd3fe3224c266102e9a23c37bc6b86fe5b33b9e4c", + "internal/api/ai_handler.go": "a407a5e55320b8e5d41d121193dc654ef02f70bd7b4566948bf4cd0a5763470a", + "internal/api/ai_handlers.go": "e3963c5a564432e9ee590f3c3db4a2220abc435cd78f6b14213559423a0283f6", + "internal/api/ai_intelligence_handlers.go": "f3864f8a53a1adaa2f64983279c3349dfff1e3095fe2ee8843ea22afe3dab577", + "internal/api/router.go": "cb5a99f8d12a7b576bf606f76fb4dfe6d1db888305249879268e6464efcb7c3f", + "internal/metrics/incident_archive.go": "2843483fca0bbafa4c6ff14b419bd599f6aeac7e20025b604cadbf16bddeed7a" }, - "binary_sha256": "234f625cb74be3d300facfed1bb17b17e20c44c06037a1f8f4fffc1c4f49d621", + "removed_backend_paths": [ + "internal/ai/incident_coordinator.go", + "internal/metrics/incident_recorder.go" + ], + "binary_sha256": "bd29e6f27be7b3ad4cfbc37842f4da90f08c6a48c9fc23b12c9c597b67346c9b", "rendering_content_sha256": { - "frontend-modern/src/components/AI/Chat/hooks/useChat.ts": "0b56b7a56e35d51ca96f0e126dd493b3164aa9e0ad4d8ae24bcf3af7a574b97c", "frontend-modern/src/components/AI/Chat/ChatMessages.tsx": "9672f7608d1e3a531c73cba20fd4a78752316783212afd0c292ddfd11d2bf371", "frontend-modern/src/components/AI/Chat/ToolExecutionBlock.tsx": "cc7bd548a418c3863486f0fe987c5c3110c2f6cdfa70b630b2ec05fb794d60d6", - "frontend-modern/src/features/alerts/useNotificationDeliveryLog.ts": "be919497714d749fcbe9a48a325aa82282b7afa8d6866d84eb1cc09ab52c14a6" + "frontend-modern/src/components/AI/Chat/hooks/useChat.ts": "0b56b7a56e35d51ca96f0e126dd493b3164aa9e0ad4d8ae24bcf3af7a574b97c", + "frontend-modern/src/features/alerts/deliveryDiagnosisPresentation.ts": "bc3eadf8f790517b377430a42eb67e8fdfcadb53bb13b87cd971d6a9ab91d607" }, "routes": [ "/patrol", - "http://127.0.0.1:5197/qualification (isolated delivery-log component)" + "http://127.0.0.1:5198/qualification (isolated real OverviewTab)" ], "viewports": [ { @@ -43,26 +47,16 @@ { "width": 390, "height": 1000 - }, - { - "width": 1440, - "height": 900 - }, - { - "width": 390, - "height": 900 } ], "states": [ - "Six captured registered-tool results: migrated legacy Docker fired event plus canonical recovery, retained lifecycle, bounded history, empty history, unavailable store and failed read.", - "Exact event IDs, original resource references, observation/occurrence times and source metadata survive expansion. Empty coverage and failed tool states remain distinct.", - "Actual removed homelab container canonical and legacy timeline APIs return the same seven records, preserving both previously captured fired/resolved records exactly. Provider refusal remains enforced.", - "Incoming delivery-log integration: latest evidence survives older success/failure, held-event requests do not hold the attempt spinner, newest pending work remains pending, and latest failure stays unavailable." + "Five captured registered-tool archive cases: saved observation with historical status and nanosecond duration disclosure, unavailable archive, malformed archive, wrong resource and missing window. Failed reads remain failed.", + "Incoming Overview dispatch wording: ready/cooldown records show Dispatch requested, cooldown identifies next eligibility, and missing dispatch timestamp stays Notification pending. No delivery success is inferred.", + "Live canonical and legacy homelab timeline queries retain the same seven records and preserve the exact original alert fired/resolved records. The incidents endpoint returns active_count=null and active_count_status=not_measured. Provider refusal remains enforced." ], "interactions": [ - "Hover/focus, Enter expansion, exact input/output comparison, deepest scrolling, Space collapse and Escape at /patrol, 1440x1000, 900x1000 and 390x1000. Actual pixels inspected at all three widths.", - "Full reload, controlled session selection, reopening all six results and exact retained input/output and completed/failed comparison.", - "Final merged-source worker Pro build installed locally and restarted healthy. Source/binary hashes fixed before and after. Private proof: tmp/patrol-history-integration/browser/receipt.json and live-history-proof.json. Captured tool/session rendering does not qualify real-model diagnosis or server persistence. No provider call or fault injection.", - "Incoming delivery component at 1440x900 and 390x900: Configuration Retry, delayed old/new attempt and held responses, enabled/disabled refresh state, failure state and pixel inspection. scripts/check-delivery-log-ordering.mjs passed with scripted APIs. This is not full-tab, installed service or external notification qualification." + "At /patrol, hover/focus and Enter expansion, exact input/output comparison, deepest output scrolling, Space collapse, Escape, full reload, controlled session selection and reopening each result. Actual pixels inspected for successful and failed outcomes at desktop, intermediate and narrow widths.", + "The incoming real Overview component was exercised through scripts/check-alert-dispatch-copy.mjs at all three widths. Status text stays within the viewport and no page errors occurred. Actual pixels inspected. No delivery action invoked.", + "Final worker Pro build installed locally and managed restart recovered healthy. Private source/binary binding and receipts: tmp/patrol-archive-retirement/runtime-binding.json, browser/receipt.json and live-history-proof.json. Zero provider calls or faults in this proof. Controlled tool/session replay proves rendering only, not model diagnosis or server persistence." ] } diff --git a/frontend-modern/src/features/alerts/__tests__/OverviewTab.deliverystatus.test.tsx b/frontend-modern/src/features/alerts/__tests__/OverviewTab.deliverystatus.test.tsx index 995a149a4..631ae1590 100644 --- a/frontend-modern/src/features/alerts/__tests__/OverviewTab.deliverystatus.test.tsx +++ b/frontend-modern/src/features/alerts/__tests__/OverviewTab.deliverystatus.test.tsx @@ -111,6 +111,19 @@ describe('OverviewTab delivery status line', () => { expect(getDeliveryDiagnoses).toHaveBeenCalled(); }); + it.each([ + { status: 'would_send', reason: 'ready' }, + { status: 'suppressed', reason: 'cooldown', nextEligibleAt: '2026-08-26T10:20:00Z' }, + ] as const)('does not turn dispatch evidence into receipt for $reason', async (state) => { + getDeliveryDiagnoses.mockResolvedValue([ + makeDiagnosis('a1', { ...state, lastNotified: '2026-08-26T10:15:00Z' }), + ]); + render(() => ); + await waitFor(() => expect(screen.getByText(/^Dispatch requested /)).toBeTruthy()); + expect(screen.queryByText(/^Notified /)).toBeNull(); + if (state.reason === 'cooldown') expect(screen.getByText(/next eligible/)).toBeTruthy(); + }); + it('renders no delivery line when the diagnosis fetch fails', async () => { const activeAlerts: Record = { a1: makeAlert('a1') }; getDeliveryDiagnoses.mockRejectedValue(new Error('boom')); diff --git a/frontend-modern/src/features/alerts/__tests__/OverviewTab.total24h.test.tsx b/frontend-modern/src/features/alerts/__tests__/OverviewTab.total24h.test.tsx index dedd034d3..1309b153d 100644 --- a/frontend-modern/src/features/alerts/__tests__/OverviewTab.total24h.test.tsx +++ b/frontend-modern/src/features/alerts/__tests__/OverviewTab.total24h.test.tsx @@ -1,5 +1,5 @@ import { afterEach, beforeEach, describe, expect, it, vi } from 'vitest'; -import { cleanup, render, screen } from '@solidjs/testing-library'; +import { cleanup, render, screen, waitFor } from '@solidjs/testing-library'; import { DEFAULT_LOCALE, setActiveLocale } from '@/i18n'; import type { Alert, AlertDeliveryDiagnosis } from '@/types/api'; @@ -75,6 +75,37 @@ describe('OverviewTab Last 24 Hours stat', () => { setActiveLocale(DEFAULT_LOCALE); }); + it('keeps dispatch evidence separate from the triggered count', async () => { + vi.useRealTimers(); + const now = Date.now(); + getDeliveryDiagnoses.mockResolvedValue([ + { + alertIdentifier: 'old', + alertId: 'old', + status: 'would_send', + reason: 'ready', + lastNotified: new Date(now).toISOString(), + } as AlertDeliveryDiagnosis, + ]); + render(() => ( + + )); + await waitFor(() => expect(screen.getByText(/^Dispatch requested /)).toBeTruthy()); + expect(screen.queryByText(/^Notified /)).toBeNull(); + expect( + screen + .getByText('Triggered (24h)') + .closest('tr') + ?.querySelector('[data-testid="alert-overview-stat-value"]')?.textContent, + ).toBe('0'); + }); + it('counts only alerts with startTime within the last 24 hours', () => { const now = Date.now(); const oneHourAgo = new Date(now - 3_600_000).toISOString(); diff --git a/frontend-modern/src/features/alerts/__tests__/deliveryDiagnosisPresentation.test.ts b/frontend-modern/src/features/alerts/__tests__/deliveryDiagnosisPresentation.test.ts index ecc7be09c..c5c4f53fd 100644 --- a/frontend-modern/src/features/alerts/__tests__/deliveryDiagnosisPresentation.test.ts +++ b/frontend-modern/src/features/alerts/__tests__/deliveryDiagnosisPresentation.test.ts @@ -35,13 +35,45 @@ describe('describeAlertDeliveryStatus', () => { expect(describeAlertDeliveryStatus(diagnosis, true)).toBeNull(); }); - it('shows the notified time when eligible and already notified', () => { + it('shows dispatch evidence without claiming destination success', () => { const diagnosis = baseDiagnosis({ lastNotified: '2026-08-26T10:15:00Z' }); const line = describeAlertDeliveryStatus(diagnosis, false); expect(line?.tone).toBe('muted'); - expect(line?.label).toMatch(/^Notified /); + expect(line?.label).toMatch(/^Dispatch requested /); }); + it('shows dispatch without promising another send when cooldown has no next time', () => { + const line = describeAlertDeliveryStatus( + baseDiagnosis({ + status: 'suppressed', + reason: 'cooldown', + lastNotified: '2026-08-26T10:15:00Z', + }), + false, + ); + expect(line?.label).toMatch(/^Dispatch requested /); + expect(line?.label).not.toContain('next'); + }); + + it.each([undefined, '', 'invalid'])( + 'does not invent dispatch for timestamp %s', + (lastNotified) => { + expect(describeAlertDeliveryStatus(baseDiagnosis({ lastNotified }), false)?.label).toBe( + 'Notification pending', + ); + expect( + describeAlertDeliveryStatus( + baseDiagnosis({ + status: 'suppressed', + reason: 'cooldown', + lastNotified, + }), + false, + )?.label, + ).toBe('Waiting for cooldown'); + }, + ); + it('shows pending when eligible but never notified', () => { const line = describeAlertDeliveryStatus(baseDiagnosis({}), false); expect(line).toEqual({ label: 'Notification pending', tone: 'muted' }); @@ -67,7 +99,7 @@ describe('describeAlertDeliveryStatus', () => { }); const line = describeAlertDeliveryStatus(diagnosis, false); expect(line?.tone).toBe('muted'); - expect(line?.label).toMatch(/^Notified .* — next /); + expect(line?.label).toMatch(/^Dispatch requested .* — next eligible /); }); it.each([ diff --git a/frontend-modern/src/features/alerts/deliveryDiagnosisPresentation.ts b/frontend-modern/src/features/alerts/deliveryDiagnosisPresentation.ts index b5ca69e95..33db25b55 100644 --- a/frontend-modern/src/features/alerts/deliveryDiagnosisPresentation.ts +++ b/frontend-modern/src/features/alerts/deliveryDiagnosisPresentation.ts @@ -54,9 +54,11 @@ export const describeAlertDeliveryStatus = ( const reason = (diagnosis.reason || '').split(':')[0]; + // lastNotified is recorded before delivery callbacks; it is dispatch evidence, + // not confirmation that a destination accepted or a person received anything. if (diagnosis.status === 'would_send') { const notifiedAt = formatShortTime(diagnosis.lastNotified); - if (notifiedAt) return { label: `Notified ${notifiedAt}`, tone: 'muted' }; + if (notifiedAt) return { label: `Dispatch requested ${notifiedAt}`, tone: 'muted' }; return { label: 'Notification pending', tone: 'muted' }; } @@ -75,9 +77,12 @@ export const describeAlertDeliveryStatus = ( const notifiedAt = formatShortTime(diagnosis.lastNotified); const nextAt = formatShortTime(diagnosis.nextEligibleAt); if (notifiedAt && nextAt) { - return { label: `Notified ${notifiedAt} — next ${nextAt}`, tone: 'muted' }; + return { + label: `Dispatch requested ${notifiedAt} — next eligible ${nextAt}`, + tone: 'muted', + }; } - if (notifiedAt) return { label: `Notified ${notifiedAt}`, tone: 'muted' }; + if (notifiedAt) return { label: `Dispatch requested ${notifiedAt}`, tone: 'muted' }; return { label: 'Waiting for cooldown', tone: 'muted' }; } case 'rate_limited': diff --git a/internal/ai/adapters/adapters.go b/internal/ai/adapters/adapters.go index 85e160a9a..a7e7987f8 100644 --- a/internal/ai/adapters/adapters.go +++ b/internal/ai/adapters/adapters.go @@ -8,7 +8,6 @@ import ( "fmt" "os" "path/filepath" - "strconv" "strings" "sync" "time" @@ -73,165 +72,6 @@ func (a *ForecastDataAdapter) GetMetricHistory(resourceID, metric string, from, return result, nil } -// MetricsAdapter provides current metrics for resources. -// It implements metrics.MetricsProvider for the incident recorder. -// Uses ReadState as the sole data source (SRC-03m migration). -type MetricsAdapter struct { - readState unifiedresources.ReadState -} - -// NewMetricsAdapter creates a new adapter for current metrics. -// ReadState is the sole data source for both GetMonitoredResourceIDs and -// GetCurrentMetrics. Returns nil if readState is nil. -func NewMetricsAdapter(readState unifiedresources.ReadState) *MetricsAdapter { - if readState == nil { - return nil - } - return &MetricsAdapter{readState: readState} -} - -// GetMonitoredResourceIDs returns all resource IDs currently being monitored. -// This is used by the incident recorder to maintain pre-incident buffers for all resources. -// Returns both unified IDs and Proxmox source IDs so that pre-incident buffers -// are keyed by both (alert-triggered recordings use source IDs). -func (a *MetricsAdapter) GetMonitoredResourceIDs() []string { - var ids []string - for _, vm := range a.readState.VMs() { - ids = append(ids, vm.ID()) - if sid := vm.SourceID(); sid != "" && sid != vm.ID() { - ids = append(ids, sid) - } - } - for _, ct := range a.readState.Containers() { - ids = append(ids, ct.ID()) - if sid := ct.SourceID(); sid != "" && sid != ct.ID() { - ids = append(ids, sid) - } - } - for _, node := range a.readState.Nodes() { - ids = append(ids, node.ID()) - if sid := node.SourceID(); sid != "" && sid != node.ID() { - ids = append(ids, sid) - } - } - return ids -} - -// GetCurrentMetricsBatch returns current metrics for every resource -// GetCurrentMetrics can resolve, in one pass over the views, keyed by every -// ID form GetCurrentMetrics matches (unified ID, source ID, VMID string, -// name). The per-ID method scans all views per call, so a sampler asking for -// thousands of resources per tick must use this instead: first-key-wins -// mirrors the per-ID method's VM -> container -> node -> storage precedence. -func (a *MetricsAdapter) GetCurrentMetricsBatch() map[string]map[string]float64 { - out := make(map[string]map[string]float64) - put := func(metrics map[string]float64, keys ...string) { - for _, key := range keys { - if key == "" { - continue - } - if _, exists := out[key]; !exists { - out[key] = metrics - } - } - } - - for _, vm := range a.readState.VMs() { - put(map[string]float64{ - "cpu": vm.CPUPercent(), - "memory": vm.MemoryPercent(), - "disk": vm.DiskPercent(), - "netin": vm.NetIn(), - "netout": vm.NetOut(), - "diskread": vm.DiskRead(), - "diskwrite": vm.DiskWrite(), - }, vm.ID(), vm.SourceID(), strconv.Itoa(vm.VMID())) - } - for _, ct := range a.readState.Containers() { - put(map[string]float64{ - "cpu": ct.CPUPercent(), - "memory": ct.MemoryPercent(), - "disk": ct.DiskPercent(), - "netin": ct.NetIn(), - "netout": ct.NetOut(), - "diskread": ct.DiskRead(), - "diskwrite": ct.DiskWrite(), - }, ct.ID(), ct.SourceID(), strconv.Itoa(ct.VMID())) - } - for _, node := range a.readState.Nodes() { - put(map[string]float64{ - "cpu": node.CPUPercent(), - "memory": node.MemoryPercent(), - "disk": node.DiskPercent(), - }, node.ID(), node.SourceID(), node.Name()) - } - for _, sp := range a.readState.StoragePools() { - put(map[string]float64{ - "disk": sp.DiskPercent(), - "used": float64(sp.DiskUsed()), - "total": float64(sp.DiskTotal()), - }, sp.ID(), sp.SourceID(), sp.Name()) - } - return out -} - -// GetCurrentMetrics returns current metrics for a resource. -// Matches by unified ID, Proxmox source ID, VMID string, or name. -// CPU/memory/disk values are normalized to 0-100 percentage scale. -func (a *MetricsAdapter) GetCurrentMetrics(resourceID string) (map[string]float64, error) { - metrics := make(map[string]float64) - - // Check VMs - for _, vm := range a.readState.VMs() { - if vm.ID() == resourceID || vm.SourceID() == resourceID || strconv.Itoa(vm.VMID()) == resourceID { - metrics["cpu"] = vm.CPUPercent() - metrics["memory"] = vm.MemoryPercent() - metrics["disk"] = vm.DiskPercent() - metrics["netin"] = vm.NetIn() - metrics["netout"] = vm.NetOut() - metrics["diskread"] = vm.DiskRead() - metrics["diskwrite"] = vm.DiskWrite() - return metrics, nil - } - } - - // Check containers - for _, ct := range a.readState.Containers() { - if ct.ID() == resourceID || ct.SourceID() == resourceID || strconv.Itoa(ct.VMID()) == resourceID { - metrics["cpu"] = ct.CPUPercent() - metrics["memory"] = ct.MemoryPercent() - metrics["disk"] = ct.DiskPercent() - metrics["netin"] = ct.NetIn() - metrics["netout"] = ct.NetOut() - metrics["diskread"] = ct.DiskRead() - metrics["diskwrite"] = ct.DiskWrite() - return metrics, nil - } - } - - // Check nodes - for _, node := range a.readState.Nodes() { - if node.ID() == resourceID || node.SourceID() == resourceID || node.Name() == resourceID { - metrics["cpu"] = node.CPUPercent() - metrics["memory"] = node.MemoryPercent() - metrics["disk"] = node.DiskPercent() - return metrics, nil - } - } - - // Check storage - for _, sp := range a.readState.StoragePools() { - if sp.ID() == resourceID || sp.SourceID() == resourceID || sp.Name() == resourceID { - metrics["disk"] = sp.DiskPercent() - metrics["used"] = float64(sp.DiskUsed()) - metrics["total"] = float64(sp.DiskTotal()) - return metrics, nil - } - } - - return metrics, nil -} - // CommandExecutorAdapter adapts the agent execution system to remediation.CommandExecutor. // This allows the remediation engine to execute commands on targets. type CommandExecutorAdapter struct { @@ -267,69 +107,6 @@ func (e *CommandExecutionDisabledError) Error() string { return "command execution is disabled - commands must be run manually" } -// IncidentRecorderToolAdapter adapts metrics.IncidentRecorder to tools.IncidentRecorderProvider -type IncidentRecorderToolAdapter struct { - recorder IncidentRecorderSource -} - -// IncidentRecorderSource defines what we need from an incident recorder -type IncidentRecorderSource interface { - GetWindowsForResource(resourceID string, limit int) []*IncidentWindowData - GetWindow(windowID string) *IncidentWindowData -} - -// IncidentWindowData represents incident window data -type IncidentWindowData struct { - ID string - ResourceID string - ResourceName string - ResourceType string - TriggerType string - TriggerID string - StartTime time.Time - EndTime *time.Time - Status string - DataPoints []IncidentDataPointData - Summary *IncidentSummaryData -} - -// IncidentDataPointData represents a single data point -type IncidentDataPointData struct { - Timestamp time.Time - Metrics map[string]float64 -} - -// IncidentSummaryData provides summary statistics -type IncidentSummaryData struct { - Duration time.Duration - DataPoints int - Peaks map[string]float64 - Lows map[string]float64 - Averages map[string]float64 - Changes map[string]float64 -} - -// NewIncidentRecorderToolAdapter creates a new incident recorder adapter -func NewIncidentRecorderToolAdapter(recorder IncidentRecorderSource) *IncidentRecorderToolAdapter { - return &IncidentRecorderToolAdapter{recorder: recorder} -} - -// GetWindowsForResource returns incident windows for a resource -func (a *IncidentRecorderToolAdapter) GetWindowsForResource(resourceID string, limit int) []*IncidentWindowData { - if a.recorder == nil { - return nil - } - return a.recorder.GetWindowsForResource(resourceID, limit) -} - -// GetWindow returns a specific incident window -func (a *IncidentRecorderToolAdapter) GetWindow(windowID string) *IncidentWindowData { - if a.recorder == nil { - return nil - } - return a.recorder.GetWindow(windowID) -} - // EventCorrelatorToolAdapter adapts proxmox.EventCorrelator to tools.EventCorrelatorProvider type EventCorrelatorToolAdapter struct { correlator EventCorrelatorSource diff --git a/internal/ai/adapters/adapters_additional_test.go b/internal/ai/adapters/adapters_additional_test.go index b8454ea24..5f20f33cf 100644 --- a/internal/ai/adapters/adapters_additional_test.go +++ b/internal/ai/adapters/adapters_additional_test.go @@ -6,23 +6,9 @@ import ( "testing" "time" - "github.com/rcourtman/pulse-go-rewrite/internal/models" "github.com/rcourtman/pulse-go-rewrite/internal/monitoring" ) -type stubIncidentRecorder struct { - windows []*IncidentWindowData - window *IncidentWindowData -} - -func (s *stubIncidentRecorder) GetWindowsForResource(resourceID string, limit int) []*IncidentWindowData { - return s.windows -} - -func (s *stubIncidentRecorder) GetWindow(windowID string) *IncidentWindowData { - return s.window -} - type stubEventCorrelator struct { correlations []EventCorrelationData events []ProxmoxEventData @@ -69,65 +55,6 @@ func TestForecastDataAdapter_GetMetricHistory(t *testing.T) { } } -func TestMetricsAdapter_GetMonitoredResourceIDs(t *testing.T) { - state := models.StateSnapshot{ - Nodes: []models.Node{{ID: "node/pve1", Name: "pve1", Instance: "inst1"}}, - VMs: []models.VM{{ID: "qemu/100", VMID: 100, Name: "vm-1", Node: "pve1", Instance: "inst1"}}, - Containers: []models.Container{{ID: "lxc/200", VMID: 200, Name: "ct-1", Node: "pve1", Instance: "inst1"}}, - } - adapter := NewMetricsAdapter(readStateFromSnapshot(state)) - ids := adapter.GetMonitoredResourceIDs() - - // Should include both unified IDs and source IDs (3 resources × 2 IDs each = 6) - if len(ids) < 3 { - t.Fatalf("expected at least 3 IDs, got %d: %v", len(ids), ids) - } - - // Verify no empty IDs - for _, id := range ids { - if id == "" { - t.Fatalf("unexpected empty ID in %v", ids) - } - } - - // Verify source IDs are present (for pre-incident buffer compatibility) - idSet := make(map[string]bool, len(ids)) - for _, id := range ids { - idSet[id] = true - } - if !idSet["qemu/100"] { - t.Errorf("expected source ID 'qemu/100' in monitored IDs, got %v", ids) - } - if !idSet["lxc/200"] { - t.Errorf("expected source ID 'lxc/200' in monitored IDs, got %v", ids) - } - if !idSet["node/pve1"] { - t.Errorf("expected source ID 'node/pve1' in monitored IDs, got %v", ids) - } -} - -func TestIncidentRecorderToolAdapter(t *testing.T) { - adapter := NewIncidentRecorderToolAdapter(nil) - if adapter.GetWindowsForResource("res", 1) != nil { - t.Fatalf("expected nil windows for nil recorder") - } - if adapter.GetWindow("id") != nil { - t.Fatalf("expected nil window for nil recorder") - } - - recorder := &stubIncidentRecorder{ - windows: []*IncidentWindowData{{ID: "w1"}}, - window: &IncidentWindowData{ID: "w1"}, - } - adapter = NewIncidentRecorderToolAdapter(recorder) - if len(adapter.GetWindowsForResource("res", 1)) != 1 { - t.Fatalf("expected windows from recorder") - } - if adapter.GetWindow("w1") == nil { - t.Fatalf("expected window from recorder") - } -} - func TestEventCorrelatorToolAdapter(t *testing.T) { adapter := NewEventCorrelatorToolAdapter(nil) if adapter.GetCorrelationsForResource("res", time.Minute) != nil { diff --git a/internal/ai/adapters/adapters_batch_test.go b/internal/ai/adapters/adapters_batch_test.go deleted file mode 100644 index ef9b23479..000000000 --- a/internal/ai/adapters/adapters_batch_test.go +++ /dev/null @@ -1,80 +0,0 @@ -package adapters - -import ( - "reflect" - "strconv" - "testing" - - "github.com/rcourtman/pulse-go-rewrite/internal/models" -) - -// The batch method must return exactly what per-ID lookups return for every -// ID form the per-ID method matches, including colliding VMID strings where -// the VM -> container precedence decides the winner. -func TestGetCurrentMetricsBatchMatchesPerIDLookups(t *testing.T) { - state := models.StateSnapshot{ - Nodes: []models.Node{ - {ID: "node/pve1", Name: "pve1", Instance: "inst1", CPU: 0.35, Memory: models.Memory{Usage: 60}}, - }, - VMs: []models.VM{ - { - ID: "qemu/100", VMID: 100, Name: "webserver", Node: "pve1", Instance: "inst1", - CPU: 45.5, Memory: models.Memory{Usage: 72.3}, Disk: models.Disk{Usage: 55}, - NetworkIn: 1024, NetworkOut: 512, DiskRead: 2048, DiskWrite: 1024, - }, - }, - Containers: []models.Container{ - { - ID: "lxc/104", VMID: 104, Name: "auth", Node: "pve1", Instance: "inst1", - CPU: 12.5, Memory: models.Memory{Usage: 30}, Disk: models.Disk{Usage: 20}, - }, - }, - Storage: []models.Storage{ - {ID: "storage/local", Name: "local", Node: "pve1", Instance: "inst1", Usage: 41, Used: 41, Total: 100}, - }, - } - - adapter := NewMetricsAdapter(readStateFromSnapshot(state)) - batch := adapter.GetCurrentMetricsBatch() - if len(batch) == 0 { - t.Fatal("batch returned no entries for a populated state") - } - - for id := range batch { - perID, err := adapter.GetCurrentMetrics(id) - if err != nil { - t.Fatalf("GetCurrentMetrics(%q) error: %v", id, err) - } - if !reflect.DeepEqual(batch[id], perID) { - t.Fatalf("metrics diverged for %q:\nbatch: %+v\nper-ID: %+v", id, batch[id], perID) - } - } - - // Every monitored ID must be resolvable through the batch. - for _, id := range adapter.GetMonitoredResourceIDs() { - if _, ok := batch[id]; !ok { - t.Fatalf("monitored ID %q missing from batch", id) - } - } -} - -func TestGetCurrentMetricsBatchVMIDCollisionPrefersVM(t *testing.T) { - state := models.StateSnapshot{ - Nodes: []models.Node{ - {ID: "node/pve1", Name: "pve1", Instance: "inst1"}, - }, - VMs: []models.VM{ - {ID: "qemu/200", VMID: 200, Name: "vm-two-hundred", Node: "pve1", Instance: "inst1", CPU: 80}, - }, - Containers: []models.Container{ - {ID: "lxc/200", VMID: 200, Name: "ct-two-hundred", Node: "pve1", Instance: "inst1", CPU: 10}, - }, - } - - adapter := NewMetricsAdapter(readStateFromSnapshot(state)) - batch := adapter.GetCurrentMetricsBatch() - perID, _ := adapter.GetCurrentMetrics(strconv.Itoa(200)) - if !reflect.DeepEqual(batch["200"], perID) { - t.Fatalf("VMID collision winner diverged:\nbatch: %+v\nper-ID: %+v", batch["200"], perID) - } -} diff --git a/internal/ai/adapters/adapters_test.go b/internal/ai/adapters/adapters_test.go index 6112fc783..c48344c25 100644 --- a/internal/ai/adapters/adapters_test.go +++ b/internal/ai/adapters/adapters_test.go @@ -2,21 +2,10 @@ package adapters import ( "context" - "fmt" "testing" "time" - - "github.com/rcourtman/pulse-go-rewrite/internal/models" - "github.com/rcourtman/pulse-go-rewrite/internal/unifiedresources" ) -// readStateFromSnapshot creates a ReadState from a models.StateSnapshot for testing. -func readStateFromSnapshot(snapshot models.StateSnapshot) unifiedresources.ReadState { - rr := unifiedresources.NewRegistry(nil) - rr.IngestSnapshot(snapshot) - return rr -} - func TestForecastDataAdapter_NilHistory(t *testing.T) { adapter := NewForecastDataAdapter(nil) if adapter != nil { @@ -24,295 +13,6 @@ func TestForecastDataAdapter_NilHistory(t *testing.T) { } } -func TestMetricsAdapter_GetCurrentMetrics_VM(t *testing.T) { - state := models.StateSnapshot{ - Nodes: []models.Node{ - {ID: "node/pve1", Name: "pve1", Instance: "inst1"}, - }, - VMs: []models.VM{ - { - ID: "qemu/100", - VMID: 100, - Name: "webserver", - Node: "pve1", - Instance: "inst1", - CPU: 45.5, - Memory: models.Memory{ - Usage: 72.3, - Used: 1024, - Total: 2048, - }, - Disk: models.Disk{ - Usage: 55.0, - Used: 5000, - Total: 10000, - }, - NetworkIn: 1024000, - NetworkOut: 512000, - DiskRead: 2048000, - DiskWrite: 1024000, - }, - }, - } - - rs := readStateFromSnapshot(state) - adapter := NewMetricsAdapter(rs) - - // Get the unified resource ID from ReadState - vms := rs.VMs() - if len(vms) != 1 { - t.Fatalf("expected 1 VM, got %d", len(vms)) - } - vmID := vms[0].ID() - - metrics, err := adapter.GetCurrentMetrics(vmID) - if err != nil { - t.Errorf("Unexpected error: %v", err) - return - } - - if metrics["cpu"] != 45.5 { - t.Errorf("Expected CPU 45.5, got %f", metrics["cpu"]) - } - if metrics["memory"] != 72.3 { - t.Errorf("Expected memory 72.3, got %f", metrics["memory"]) - } - if metrics["disk"] != 55.0 { - t.Errorf("Expected disk 55.0, got %f", metrics["disk"]) - } - if metrics["netin"] != 1024000 { - t.Errorf("Expected netin 1024000, got %f", metrics["netin"]) - } -} - -func TestMetricsAdapter_GetCurrentMetrics_Container(t *testing.T) { - state := models.StateSnapshot{ - Nodes: []models.Node{ - {ID: "node/pve1", Name: "pve1", Instance: "inst1"}, - }, - Containers: []models.Container{ - { - ID: "lxc/101", - VMID: 101, - Name: "container1", - Node: "pve1", - Instance: "inst1", - CPU: 25.0, - Memory: models.Memory{ - Usage: 45.0, - Used: 512, - Total: 1024, - }, - Disk: models.Disk{ - Usage: 30.0, - Used: 3000, - Total: 10000, - }, - NetworkIn: 500000, - NetworkOut: 250000, - DiskRead: 1000000, - DiskWrite: 500000, - }, - }, - } - - rs := readStateFromSnapshot(state) - adapter := NewMetricsAdapter(rs) - - containers := rs.Containers() - if len(containers) != 1 { - t.Fatalf("expected 1 container, got %d", len(containers)) - } - ctID := containers[0].ID() - - metrics, err := adapter.GetCurrentMetrics(ctID) - if err != nil { - t.Errorf("Unexpected error: %v", err) - return - } - - if metrics["cpu"] != 25.0 { - t.Errorf("Expected CPU 25.0, got %f", metrics["cpu"]) - } - if metrics["memory"] != 45.0 { - t.Errorf("Expected memory 45.0, got %f", metrics["memory"]) - } -} - -func TestMetricsAdapter_GetCurrentMetrics_Node(t *testing.T) { - state := models.StateSnapshot{ - Nodes: []models.Node{ - { - ID: "node/pve1", - Name: "pve1", - Instance: "inst1", - CPU: 25.5, - Memory: models.Memory{ - Usage: 65.0, - Used: 6500, - Total: 10000, - }, - Disk: models.Disk{ - Usage: 40.0, - Used: 4000, - Total: 10000, - }, - }, - }, - } - - rs := readStateFromSnapshot(state) - adapter := NewMetricsAdapter(rs) - - nodes := rs.Nodes() - if len(nodes) != 1 { - t.Fatalf("expected 1 node, got %d", len(nodes)) - } - nodeID := nodes[0].ID() - - metrics, err := adapter.GetCurrentMetrics(nodeID) - if err != nil { - t.Errorf("Unexpected error: %v", err) - return - } - - if metrics["cpu"] != 25.5 { - t.Errorf("Expected CPU 25.5, got %f", metrics["cpu"]) - } - if metrics["memory"] != 65.0 { - t.Errorf("Expected memory 65.0, got %f", metrics["memory"]) - } - if metrics["disk"] != 40.0 { - t.Errorf("Expected disk 40.0, got %f", metrics["disk"]) - } -} - -func TestMetricsAdapter_GetCurrentMetrics_NodeByName(t *testing.T) { - state := models.StateSnapshot{ - Nodes: []models.Node{ - { - ID: "node/pve1", - Name: "pve1", - Instance: "inst1", - CPU: 25.5, - Memory: models.Memory{ - Usage: 65.0, - Used: 6500, - Total: 10000, - }, - Disk: models.Disk{ - Usage: 40.0, - Used: 4000, - Total: 10000, - }, - }, - }, - } - - adapter := NewMetricsAdapter(readStateFromSnapshot(state)) - - // Node lookup by name should still work - metrics, err := adapter.GetCurrentMetrics("pve1") - if err != nil { - t.Errorf("Unexpected error: %v", err) - return - } - - if metrics["cpu"] != 25.5 { - t.Errorf("Expected CPU 25.5 when matching by name, got %f", metrics["cpu"]) - } -} - -func TestMetricsAdapter_GetCurrentMetrics_Storage(t *testing.T) { - state := models.StateSnapshot{ - Nodes: []models.Node{ - {ID: "node/pve1", Name: "pve1", Instance: "inst1"}, - }, - Storage: []models.Storage{ - { - ID: "storage/local-zfs", - Name: "local-zfs", - Node: "pve1", - Instance: "inst1", - Used: 50000000000, - Total: 100000000000, - Usage: 50.0, - }, - }, - } - - rs := readStateFromSnapshot(state) - adapter := NewMetricsAdapter(rs) - - pools := rs.StoragePools() - if len(pools) != 1 { - t.Fatalf("expected 1 storage pool, got %d", len(pools)) - } - storageID := pools[0].ID() - - metrics, err := adapter.GetCurrentMetrics(storageID) - if err != nil { - t.Errorf("Unexpected error: %v", err) - return - } - - if metrics["disk"] != 50.0 { - t.Errorf("Expected disk 50.0, got %f", metrics["disk"]) - } - if metrics["used"] != 50000000000 { - t.Errorf("Expected used 50000000000, got %f", metrics["used"]) - } - if metrics["total"] != 100000000000 { - t.Errorf("Expected total 100000000000, got %f", metrics["total"]) - } -} - -func TestMetricsAdapter_GetCurrentMetrics_StorageByName(t *testing.T) { - state := models.StateSnapshot{ - Nodes: []models.Node{ - {ID: "node/pve1", Name: "pve1", Instance: "inst1"}, - }, - Storage: []models.Storage{ - { - ID: "storage/local-zfs", - Name: "local-zfs", - Node: "pve1", - Instance: "inst1", - Used: 50000000000, - Total: 100000000000, - Usage: 50.0, - }, - }, - } - - adapter := NewMetricsAdapter(readStateFromSnapshot(state)) - - // Storage lookup by name should still work - metrics, err := adapter.GetCurrentMetrics("local-zfs") - if err != nil { - t.Errorf("Unexpected error: %v", err) - return - } - - if metrics["disk"] != 50.0 { - t.Errorf("Expected disk 50.0, got %f", metrics["disk"]) - } -} - -func TestMetricsAdapter_GetCurrentMetrics_NotFound(t *testing.T) { - state := models.StateSnapshot{} - - adapter := NewMetricsAdapter(readStateFromSnapshot(state)) - - metrics, err := adapter.GetCurrentMetrics("nonexistent") - if err != nil { - t.Errorf("Unexpected error: %v", err) - } - if len(metrics) != 0 { - t.Errorf("Expected empty metrics, got %d entries", len(metrics)) - } -} - func TestCommandExecutorAdapter_Disabled(t *testing.T) { adapter := NewCommandExecutorAdapter() @@ -349,131 +49,3 @@ func TestCommandExecutionDisabledError_Message(t *testing.T) { t.Errorf("Unexpected error message: %s", msg) } } - -func TestMetricsAdapter_NilReadState(t *testing.T) { - adapter := NewMetricsAdapter(nil) - if adapter != nil { - t.Error("Expected nil adapter for nil ReadState") - } -} - -func TestMetricsAdapter_VMIDMatch(t *testing.T) { - state := models.StateSnapshot{ - Nodes: []models.Node{ - {ID: "node/pve1", Name: "pve1", Instance: "inst1"}, - }, - VMs: []models.VM{ - { - ID: "qemu/100", - VMID: 100, - Name: "webserver", - Node: "pve1", - Instance: "inst1", - CPU: 45.5, - Memory: models.Memory{ - Usage: 72.3, - Used: 1024, - Total: 2048, - }, - Disk: models.Disk{ - Usage: 55.0, - Used: 5000, - Total: 10000, - }, - }, - }, - } - - adapter := NewMetricsAdapter(readStateFromSnapshot(state)) - - // Test lookup by VMID string - metrics, err := adapter.GetCurrentMetrics(fmt.Sprintf("%d", 100)) - if err != nil { - t.Errorf("Unexpected error: %v", err) - return - } - - if metrics["cpu"] != 45.5 { - t.Errorf("Expected CPU 45.5 when matching by VMID, got %f", metrics["cpu"]) - } -} - -func TestMetricsAdapter_IDConsistency(t *testing.T) { - // Verify that GetMonitoredResourceIDs returns IDs that work with GetCurrentMetrics - state := models.StateSnapshot{ - Nodes: []models.Node{ - {ID: "node/pve1", Name: "pve1", Instance: "inst1"}, - }, - VMs: []models.VM{ - { - ID: "qemu/100", VMID: 100, Name: "vm1", Node: "pve1", Instance: "inst1", - CPU: 50.0, Memory: models.Memory{Usage: 60.0, Used: 600, Total: 1000}, - Disk: models.Disk{Usage: 70.0, Used: 700, Total: 1000}, - }, - }, - } - - adapter := NewMetricsAdapter(readStateFromSnapshot(state)) - ids := adapter.GetMonitoredResourceIDs() - if len(ids) < 1 { - t.Fatalf("expected at least 1 ID, got %d", len(ids)) - } - - // Each ID from GetMonitoredResourceIDs should be usable with GetCurrentMetrics - foundVM := false - for _, id := range ids { - metrics, err := adapter.GetCurrentMetrics(id) - if err != nil { - t.Errorf("GetCurrentMetrics(%q) error: %v", id, err) - continue - } - if cpu, ok := metrics["cpu"]; ok && cpu == 50.0 { - foundVM = true - } - } - if !foundVM { - t.Error("Expected to find VM metrics via GetMonitoredResourceIDs() IDs") - } -} - -func TestMetricsAdapter_SourceIDMatch(t *testing.T) { - // Verify that GetCurrentMetrics can find resources by their Proxmox source ID - state := models.StateSnapshot{ - Nodes: []models.Node{ - {ID: "node/pve1", Name: "pve1", Instance: "inst1"}, - }, - VMs: []models.VM{ - { - ID: "qemu/100", VMID: 100, Name: "webserver", Node: "pve1", Instance: "inst1", - CPU: 45.5, Memory: models.Memory{Usage: 72.3, Used: 1024, Total: 2048}, - Disk: models.Disk{Usage: 55.0, Used: 5000, Total: 10000}, - }, - }, - Storage: []models.Storage{ - { - ID: "storage/local-zfs", Name: "local-zfs", Node: "pve1", Instance: "inst1", - Used: 50000000000, Total: 100000000000, Usage: 50.0, - }, - }, - } - - adapter := NewMetricsAdapter(readStateFromSnapshot(state)) - - // Lookup VM by Proxmox source ID - metrics, err := adapter.GetCurrentMetrics("qemu/100") - if err != nil { - t.Fatalf("Unexpected error: %v", err) - } - if metrics["cpu"] != 45.5 { - t.Errorf("Expected CPU 45.5 via source ID, got %f", metrics["cpu"]) - } - - // Lookup storage by Proxmox source ID - metrics, err = adapter.GetCurrentMetrics("storage/local-zfs") - if err != nil { - t.Fatalf("Unexpected error: %v", err) - } - if metrics["disk"] != 50.0 { - t.Errorf("Expected disk 50.0 via source ID, got %f", metrics["disk"]) - } -} diff --git a/internal/ai/chat/service.go b/internal/ai/chat/service.go index c800e6bae..ad92cdc60 100644 --- a/internal/ai/chat/service.go +++ b/internal/ai/chat/service.go @@ -60,7 +60,7 @@ type ( AgentProfileManager = tools.AgentProfileManager FindingsManager = tools.FindingsManager MetadataUpdater = tools.MetadataUpdater - IncidentRecorderProvider = tools.IncidentRecorderProvider + IncidentArchiveProvider = tools.IncidentArchiveProvider EventCorrelatorProvider = tools.EventCorrelatorProvider KnowledgeStoreProvider = tools.KnowledgeStoreProvider AssistantDiscoveryProvider = tools.DiscoveryProvider @@ -3598,11 +3598,11 @@ func (s *Service) SetMetadataUpdater(updater MetadataUpdater) { } } -func (s *Service) SetIncidentRecorderProvider(provider IncidentRecorderProvider) { +func (s *Service) SetIncidentArchiveProvider(provider IncidentArchiveProvider) { s.mu.Lock() defer s.mu.Unlock() if s.executor != nil { - s.executor.SetIncidentRecorderProvider(provider) + s.executor.SetIncidentArchiveProvider(provider) } } diff --git a/internal/ai/chat/service_additional_test.go b/internal/ai/chat/service_additional_test.go index 6e8505d38..4a3ed1c01 100644 --- a/internal/ai/chat/service_additional_test.go +++ b/internal/ai/chat/service_additional_test.go @@ -16,7 +16,7 @@ func TestServiceSettersAndAutonomousMode(t *testing.T) { agenticLoop: loop, } - service.SetIncidentRecorderProvider(nil) + service.SetIncidentArchiveProvider(nil) service.SetEventCorrelatorProvider(nil) service.SetKnowledgeStoreProvider(nil) diff --git a/internal/ai/incident_coordinator.go b/internal/ai/incident_coordinator.go deleted file mode 100644 index 3ab75b7a2..000000000 --- a/internal/ai/incident_coordinator.go +++ /dev/null @@ -1,325 +0,0 @@ -// Package ai provides AI-powered infrastructure analysis. -package ai - -import ( - "sync" - "time" - - "github.com/rcourtman/pulse-go-rewrite/internal/ai/memory" - "github.com/rcourtman/pulse-go-rewrite/internal/alerts" - "github.com/rcourtman/pulse-go-rewrite/internal/metrics" - "github.com/rs/zerolog/log" -) - -// IncidentCoordinatorConfig configures the incident coordinator -type IncidentCoordinatorConfig struct { - PreBuffer time.Duration // History to capture before incident (default: 5 min) - PostDuration time.Duration // How long to record after trigger (default: 10 min) - MaxConcurrent int // Maximum concurrent incident recordings (default: 50) - EnableRecorder bool // Whether to enable high-frequency recording -} - -// DefaultIncidentCoordinatorConfig returns sensible defaults -func DefaultIncidentCoordinatorConfig() IncidentCoordinatorConfig { - return IncidentCoordinatorConfig{ - PreBuffer: 5 * time.Minute, - PostDuration: 10 * time.Minute, - MaxConcurrent: 50, - EnableRecorder: true, - } -} - -// IncidentCoordinator coordinates incident recording between the metrics.IncidentRecorder -// (for high-frequency data capture) and memory.IncidentStore (for incident timeline tracking). -type IncidentCoordinator struct { - mu sync.RWMutex - - config IncidentCoordinatorConfig - - // Components - recorder *metrics.IncidentRecorder // High-frequency metrics capture - incidentStore *memory.IncidentStore // Incident timeline tracking - - // Active incidents - maps alert ID to window ID - activeIncidents map[string]activeIncident - - // Control - running bool -} - -type activeIncident struct { - windowID string - resourceID string - startedAt time.Time - stopTimer *time.Timer // Timer to auto-stop recording after post-duration -} - -// NewIncidentCoordinator creates a new incident coordinator -func NewIncidentCoordinator(cfg IncidentCoordinatorConfig) *IncidentCoordinator { - if cfg.PreBuffer <= 0 { - cfg.PreBuffer = 5 * time.Minute - } - if cfg.PostDuration <= 0 { - cfg.PostDuration = 10 * time.Minute - } - if cfg.MaxConcurrent <= 0 { - cfg.MaxConcurrent = 50 - } - - return &IncidentCoordinator{ - config: cfg, - activeIncidents: make(map[string]activeIncident), - } -} - -// SetRecorder sets the metrics incident recorder -func (c *IncidentCoordinator) SetRecorder(recorder *metrics.IncidentRecorder) { - c.mu.Lock() - defer c.mu.Unlock() - c.recorder = recorder -} - -// SetIncidentStore sets the incident timeline store -func (c *IncidentCoordinator) SetIncidentStore(store *memory.IncidentStore) { - c.mu.Lock() - defer c.mu.Unlock() - c.incidentStore = store -} - -// Start starts the incident coordinator -func (c *IncidentCoordinator) Start() { - c.mu.Lock() - defer c.mu.Unlock() - if c.running { - return - } - c.running = true - log.Info().Msg("incident coordinator started") -} - -// Stop stops the incident coordinator -func (c *IncidentCoordinator) Stop() { - c.mu.Lock() - defer c.mu.Unlock() - if !c.running { - return - } - c.running = false - - // Stop all active timers - for _, inc := range c.activeIncidents { - if inc.stopTimer != nil { - inc.stopTimer.Stop() - } - } - c.activeIncidents = make(map[string]activeIncident) - - log.Info().Msg("incident coordinator stopped") -} - -// OnAlertFired is called when an alert fires - starts incident recording -func (c *IncidentCoordinator) OnAlertFired(alert *alerts.Alert) { - if alert == nil { - return - } - - c.mu.Lock() - defer c.mu.Unlock() - - if !c.running { - return - } - - // Check if we already have an active incident for this alert - if _, exists := c.activeIncidents[alert.ID]; exists { - log.Debug(). - Str("alert_identifier", alert.ID). - Msg("Incident already being recorded for this alert") - return - } - - // Check concurrent limit - if len(c.activeIncidents) >= c.config.MaxConcurrent { - log.Warn(). - Str("alert_identifier", alert.ID). - Int("active_count", len(c.activeIncidents)). - Msg("Incident coordinator at capacity, skipping new incident") - return - } - - // Start high-frequency recording if enabled and recorder available - var windowID string - if c.config.EnableRecorder && c.recorder != nil { - windowID = c.recorder.StartRecording( - alert.ResourceID, - alert.ResourceName, - "", // resourceType - we don't always have this - "alert", - alert.ID, - ) - } - - // Record in incident store - if c.incidentStore != nil { - c.incidentStore.RecordAlertFired(alert) - } - - // Track the active incident - inc := activeIncident{ - windowID: windowID, - resourceID: alert.ResourceID, - startedAt: time.Now(), - } - - c.activeIncidents[alert.ID] = inc - - log.Info(). - Str("alert_identifier", alert.ID). - Str("resource_id", alert.ResourceID). - Str("window_id", windowID). - Msg("Incident coordinator: Started incident recording") -} - -// OnAlertCleared is called when an alert clears - schedules recording stop -func (c *IncidentCoordinator) OnAlertCleared(alert *alerts.Alert) { - if alert == nil { - return - } - - c.mu.Lock() - - inc, exists := c.activeIncidents[alert.ID] - if !exists { - c.mu.Unlock() - return - } - - // Record resolution in incident store - if c.incidentStore != nil { - c.incidentStore.RecordAlertResolved(alert, time.Now()) - } - - // If no recorder or no window, just clean up immediately - if c.recorder == nil || inc.windowID == "" { - delete(c.activeIncidents, alert.ID) - c.mu.Unlock() - return - } - - // Schedule stop after post-duration to capture post-incident data - alertID := alert.ID - timer := time.AfterFunc(c.config.PostDuration, func() { - c.stopIncidentRecording(alertID) - }) - - inc.stopTimer = timer - c.activeIncidents[alert.ID] = inc - c.mu.Unlock() - - log.Info(). - Str("alert_identifier", alert.ID). - Str("window_id", inc.windowID). - Dur("post_duration", c.config.PostDuration). - Msg("Incident coordinator: Alert cleared, scheduled recording stop") -} - -// stopIncidentRecording stops recording for a specific alert -func (c *IncidentCoordinator) stopIncidentRecording(alertID string) { - c.mu.Lock() - defer c.mu.Unlock() - - inc, exists := c.activeIncidents[alertID] - if !exists { - return - } - - // Stop the recorder - if c.recorder != nil && inc.windowID != "" { - c.recorder.StopRecording(inc.windowID) - } - - // Clean up - if inc.stopTimer != nil { - inc.stopTimer.Stop() - } - delete(c.activeIncidents, alertID) - - log.Info(). - Str("alert_identifier", alertID). - Str("window_id", inc.windowID). - Msg("Incident coordinator: Stopped incident recording") -} - -// OnAnomalyDetected is called when an anomaly is detected - starts focused recording -func (c *IncidentCoordinator) OnAnomalyDetected(resourceID, resourceType, metric string, severity string) { - c.mu.Lock() - defer c.mu.Unlock() - - if !c.running || !c.config.EnableRecorder || c.recorder == nil { - return - } - - // Create a pseudo-alert ID for the anomaly - anomalyID := "anomaly-" + resourceID + "-" + metric - - // Check if we already have an active incident for this - if _, exists := c.activeIncidents[anomalyID]; exists { - return - } - - // Check concurrent limit - if len(c.activeIncidents) >= c.config.MaxConcurrent { - return - } - - // Start recording - windowID := c.recorder.StartRecording( - resourceID, - "", // name - resourceType, - "anomaly", - anomalyID, - ) - - // Track the active incident - c.activeIncidents[anomalyID] = activeIncident{ - windowID: windowID, - resourceID: resourceID, - startedAt: time.Now(), - } - - // Schedule auto-stop after post-duration (anomalies don't have "clear" events) - timer := time.AfterFunc(c.config.PostDuration, func() { - c.stopIncidentRecording(anomalyID) - }) - c.activeIncidents[anomalyID] = activeIncident{ - windowID: windowID, - resourceID: resourceID, - startedAt: time.Now(), - stopTimer: timer, - } - - log.Info(). - Str("resource_id", resourceID). - Str("metric", metric). - Str("severity", severity). - Str("window_id", windowID). - Msg("Incident coordinator: Started anomaly recording") -} - -// GetActiveIncidentCount returns the number of active incidents being recorded -func (c *IncidentCoordinator) GetActiveIncidentCount() int { - c.mu.RLock() - defer c.mu.RUnlock() - return len(c.activeIncidents) -} - -// GetRecordingWindowID returns the recording window ID for an alert -func (c *IncidentCoordinator) GetRecordingWindowID(alertID string) string { - c.mu.RLock() - defer c.mu.RUnlock() - if inc, exists := c.activeIncidents[alertID]; exists { - return inc.windowID - } - return "" -} diff --git a/internal/ai/incident_coordinator_additional_test.go b/internal/ai/incident_coordinator_additional_test.go deleted file mode 100644 index 11b71b39f..000000000 --- a/internal/ai/incident_coordinator_additional_test.go +++ /dev/null @@ -1,179 +0,0 @@ -package ai - -import ( - "testing" - "time" - - "github.com/rcourtman/pulse-go-rewrite/internal/ai/memory" - "github.com/rcourtman/pulse-go-rewrite/internal/alerts" - "github.com/rcourtman/pulse-go-rewrite/internal/metrics" -) - -func TestNewIncidentCoordinator_DefaultFallbacks(t *testing.T) { - cfg := IncidentCoordinatorConfig{ - PreBuffer: -1 * time.Second, - PostDuration: 0, - MaxConcurrent: 0, - } - - coord := NewIncidentCoordinator(cfg) - defaults := DefaultIncidentCoordinatorConfig() - - if coord.config.PreBuffer != defaults.PreBuffer { - t.Fatalf("expected default pre-buffer %v, got %v", defaults.PreBuffer, coord.config.PreBuffer) - } - if coord.config.PostDuration != defaults.PostDuration { - t.Fatalf("expected default post-duration %v, got %v", defaults.PostDuration, coord.config.PostDuration) - } - if coord.config.MaxConcurrent != defaults.MaxConcurrent { - t.Fatalf("expected default max concurrent %d, got %d", defaults.MaxConcurrent, coord.config.MaxConcurrent) - } - if coord.activeIncidents == nil { - t.Fatal("expected active incident map to be initialized") - } -} - -func TestIncidentCoordinator_OnAlertFired_RequiresRunning(t *testing.T) { - coord := NewIncidentCoordinator(DefaultIncidentCoordinatorConfig()) - store := memory.NewIncidentStore(memory.IncidentStoreConfig{}) - coord.SetIncidentStore(store) - - alert := &alerts.Alert{ - ID: "alert-requires-running", - ResourceID: "resource-requires-running", - ResourceName: "resource-requires-running", - } - - coord.OnAlertFired(nil) - coord.OnAlertFired(alert) - - if got := coord.GetActiveIncidentCount(); got != 0 { - t.Fatalf("expected no active incidents while coordinator is stopped, got %d", got) - } - if got := len(store.ListIncidentsByResource(alert.ResourceID, 0)); got != 0 { - t.Fatalf("expected no incident-store records while stopped, got %d", got) - } - - coord.Start() - coord.OnAlertFired(alert) - - if got := coord.GetActiveIncidentCount(); got != 1 { - t.Fatalf("expected one active incident after start, got %d", got) - } - if got := len(store.ListIncidentsByResource(alert.ResourceID, 0)); got != 1 { - t.Fatalf("expected one incident-store record after start, got %d", got) - } -} - -func TestIncidentCoordinator_OnAlertCleared_ImmediateCleanupWithoutRecorder(t *testing.T) { - coord := NewIncidentCoordinator(DefaultIncidentCoordinatorConfig()) - store := memory.NewIncidentStore(memory.IncidentStoreConfig{}) - coord.SetIncidentStore(store) - coord.Start() - - alert := &alerts.Alert{ - ID: "alert-no-recorder", - ResourceID: "resource-no-recorder", - ResourceName: "resource-no-recorder", - } - - coord.OnAlertFired(alert) - if got := coord.GetActiveIncidentCount(); got != 1 { - t.Fatalf("expected one active incident after fire, got %d", got) - } - - coord.OnAlertCleared(nil) - coord.OnAlertCleared(&alerts.Alert{ID: "missing"}) - if got := coord.GetActiveIncidentCount(); got != 1 { - t.Fatalf("expected active incident to remain after nil/missing clears, got %d", got) - } - - coord.OnAlertCleared(alert) - if got := coord.GetActiveIncidentCount(); got != 0 { - t.Fatalf("expected incident to be cleaned up immediately without recorder, got %d", got) - } - - incidents := store.ListIncidentsByResource(alert.ResourceID, 0) - if len(incidents) != 1 { - t.Fatalf("expected exactly one stored incident, got %d", len(incidents)) - } - if incidents[0].Status != memory.IncidentStatusResolved { - t.Fatalf("expected incident status %q, got %q", memory.IncidentStatusResolved, incidents[0].Status) - } - if incidents[0].ClosedAt == nil { - t.Fatal("expected incident ClosedAt to be set on clear") - } -} - -func TestIncidentCoordinator_OnAnomalyDetected_DuplicateAndCapacityAndStop(t *testing.T) { - cfg := DefaultIncidentCoordinatorConfig() - cfg.MaxConcurrent = 1 - cfg.PostDuration = time.Minute - - coord := NewIncidentCoordinator(cfg) - - recCfg := metrics.DefaultIncidentRecorderConfig() - recCfg.SampleInterval = 10 * time.Millisecond - recorder := metrics.NewIncidentRecorder(recCfg) - recorder.SetMetricsProvider(&MockMetricsProvider{data: map[string]map[string]float64{ - "resource-anomaly": {"cpu": 95}, - }}) - recorder.Start() - defer recorder.Stop() - - coord.SetRecorder(recorder) - coord.Start() - - coord.OnAnomalyDetected("resource-anomaly", "agent", "cpu", "critical") - if got := coord.GetActiveIncidentCount(); got != 1 { - t.Fatalf("expected one anomaly incident, got %d", got) - } - - coord.OnAnomalyDetected("resource-anomaly", "agent", "cpu", "critical") - if got := coord.GetActiveIncidentCount(); got != 1 { - t.Fatalf("expected duplicate anomaly to be ignored, got %d", got) - } - - coord.OnAnomalyDetected("resource-anomaly", "agent", "memory", "warning") - if got := coord.GetActiveIncidentCount(); got != 1 { - t.Fatalf("expected anomaly over capacity to be ignored, got %d", got) - } - - coord.Stop() - if got := coord.GetActiveIncidentCount(); got != 0 { - t.Fatalf("expected active incidents to be cleared on stop, got %d", got) - } - - coord.OnAnomalyDetected("resource-anomaly", "agent", "cpu", "critical") - if got := coord.GetActiveIncidentCount(); got != 0 { - t.Fatalf("expected anomalies to be ignored while stopped, got %d", got) - } -} - -func TestIncidentCoordinator_OnAnomalyDetected_CanonicalizesLegacyHostAlias(t *testing.T) { - cfg := DefaultIncidentCoordinatorConfig() - cfg.PostDuration = time.Minute - - coord := NewIncidentCoordinator(cfg) - - recCfg := metrics.DefaultIncidentRecorderConfig() - recorder := metrics.NewIncidentRecorder(recCfg) - recorder.SetMetricsProvider(&MockMetricsProvider{data: map[string]map[string]float64{ - "resource-anomaly": {"cpu": 95}, - }}) - recorder.Start() - defer recorder.Stop() - - coord.SetRecorder(recorder) - coord.Start() - - coord.OnAnomalyDetected("resource-anomaly", "host", "cpu", "critical") - - windows := recorder.GetWindowsForResource("resource-anomaly", 0) - if len(windows) != 1 { - t.Fatalf("expected one anomaly window, got %d", len(windows)) - } - if windows[0].ResourceType != "agent" { - t.Fatalf("expected anomaly recording resource type to be canonicalized to agent, got %q", windows[0].ResourceType) - } -} diff --git a/internal/ai/incident_coordinator_test.go b/internal/ai/incident_coordinator_test.go deleted file mode 100644 index 9bb4a9d0d..000000000 --- a/internal/ai/incident_coordinator_test.go +++ /dev/null @@ -1,219 +0,0 @@ -package ai - -import ( - "sync" - "testing" - "time" - - "github.com/rcourtman/pulse-go-rewrite/internal/alerts" - "github.com/rcourtman/pulse-go-rewrite/internal/metrics" -) - -// MockMetricsProvider for the real IncidentRecorder -type MockMetricsProvider struct { - mu sync.Mutex - data map[string]map[string]float64 -} - -func (m *MockMetricsProvider) GetCurrentMetrics(resourceID string) (map[string]float64, error) { - m.mu.Lock() - defer m.mu.Unlock() - if val, ok := m.data[resourceID]; ok { - return val, nil - } - return map[string]float64{"cpu": 10.0}, nil -} - -func (m *MockMetricsProvider) GetMonitoredResourceIDs() []string { - m.mu.Lock() - defer m.mu.Unlock() - keys := make([]string, 0, len(m.data)) - for k := range m.data { - keys = append(keys, k) - } - return keys -} - -func TestIncidentCoordinator_Lifecycle(t *testing.T) { - cfg := DefaultIncidentCoordinatorConfig() - coord := NewIncidentCoordinator(cfg) - - if coord.running { - t.Error("Coordinator should not be running initially") - } - - coord.Start() - if !coord.running { - t.Error("Coordinator should be running after Start()") - } - - // Double start should be safe - coord.Start() - - coord.Stop() - if coord.running { - t.Error("Coordinator should not be running after Stop()") - } - - // Double stop should be safe - coord.Stop() -} - -func TestIncidentCoordinator_OnAlertFired(t *testing.T) { - cfg := DefaultIncidentCoordinatorConfig() - coord := NewIncidentCoordinator(cfg) - - // Create real recorder with mock provider - recCfg := metrics.DefaultIncidentRecorderConfig() - recCfg.SampleInterval = 50 * time.Millisecond // fast sampling - recorder := metrics.NewIncidentRecorder(recCfg) - provider := &MockMetricsProvider{data: make(map[string]map[string]float64)} - recorder.SetMetricsProvider(provider) - recorder.Start() // Recorder must be started - defer recorder.Stop() - - // Inject recorder into coordinator - coord.SetRecorder(recorder) - coord.Start() - - alert := &alerts.Alert{ - ID: "alert-1", - ResourceID: "res-1", - } - - coord.OnAlertFired(alert) - - if coord.GetActiveIncidentCount() != 1 { - t.Errorf("Expected 1 active incident, got %d", coord.GetActiveIncidentCount()) - } - - // Verify recorder has active window - wid := coord.GetRecordingWindowID("alert-1") - if wid == "" { - t.Error("Expected valid window ID") - } - - // Fire same alert again - should be ignored - coord.OnAlertFired(alert) - if coord.GetActiveIncidentCount() != 1 { - t.Errorf("Expected count to remain 1, got %d", coord.GetActiveIncidentCount()) - } - - // Fire another alert - alert2 := &alerts.Alert{ - ID: "alert-2", - ResourceID: "res-2", - } - coord.OnAlertFired(alert2) - if coord.GetActiveIncidentCount() != 2 { - t.Errorf("Expected 2 active incidents, got %d", coord.GetActiveIncidentCount()) - } -} - -func TestIncidentCoordinator_OnAlertCleared(t *testing.T) { - cfg := DefaultIncidentCoordinatorConfig() - cfg.PostDuration = 50 * time.Millisecond // fast for testing - coord := NewIncidentCoordinator(cfg) - - // Create real recorder - recCfg := metrics.DefaultIncidentRecorderConfig() - recorder := metrics.NewIncidentRecorder(recCfg) - provider := &MockMetricsProvider{data: make(map[string]map[string]float64)} - recorder.SetMetricsProvider(provider) - recorder.Start() - defer recorder.Stop() - - coord.SetRecorder(recorder) - coord.Start() - - alert := &alerts.Alert{ID: "alert-1", ResourceID: "res-1"} - coord.OnAlertFired(alert) - - if coord.GetActiveIncidentCount() != 1 { - t.Fatal("Failed to start incident") - } - - // Clear alert - coord.OnAlertCleared(alert) - - // Since we have a recorder and postDuration is 50ms, it should NOT be removed immediately - if coord.GetActiveIncidentCount() != 1 { - t.Error("Incident should NOT be removed immediately when recorder is active") - } - - // Wait for post duration - time.Sleep(100 * time.Millisecond) - - // Now it should be removed (via time.AfterFunc callback) - if coord.GetActiveIncidentCount() != 0 { - t.Errorf("Incident should be removed after post duration, count=%d", coord.GetActiveIncidentCount()) - } -} - -func TestIncidentCoordinator_MaxConcurrent(t *testing.T) { - cfg := DefaultIncidentCoordinatorConfig() - cfg.MaxConcurrent = 1 - coord := NewIncidentCoordinator(cfg) - coord.Start() - - coord.OnAlertFired(&alerts.Alert{ID: "alert-1", ResourceID: "res-1"}) - if coord.GetActiveIncidentCount() != 1 { - t.Fatal("Should accept first incident") - } - - coord.OnAlertFired(&alerts.Alert{ID: "alert-2", ResourceID: "res-2"}) - if coord.GetActiveIncidentCount() != 1 { - t.Error("Should ignore second incident due to cap") - } -} - -func TestIncidentCoordinator_OnAnomalyDetected(t *testing.T) { - cfg := DefaultIncidentCoordinatorConfig() - coord := NewIncidentCoordinator(cfg) - - recCfg := metrics.DefaultIncidentRecorderConfig() - recorder := metrics.NewIncidentRecorder(recCfg) - provider := &MockMetricsProvider{data: make(map[string]map[string]float64)} - recorder.SetMetricsProvider(provider) - recorder.Start() - defer recorder.Stop() - - coord.SetRecorder(recorder) - coord.Start() - - // Start anomaly recording - coord.OnAnomalyDetected("res-1", "agent", "cpu", "critical") - - if coord.GetActiveIncidentCount() != 1 { - t.Errorf("Should start incident for anomaly, got %d", coord.GetActiveIncidentCount()) - } - - // Anomaly ID format check (internal detail, but verify implicitly via count) -} - -func TestIncidentCoordinator_GetRecordingWindowID(t *testing.T) { - cfg := DefaultIncidentCoordinatorConfig() - coord := NewIncidentCoordinator(cfg) - - recCfg := metrics.DefaultIncidentRecorderConfig() - recorder := metrics.NewIncidentRecorder(recCfg) - recorder.Start() - defer recorder.Stop() - - coord.SetRecorder(recorder) - coord.Start() - - alert := &alerts.Alert{ID: "alert-1", ResourceID: "res-1"} - coord.OnAlertFired(alert) - - // With recorder, windowID should be present (start with 'iw-') - wid := coord.GetRecordingWindowID("alert-1") - if wid == "" { - t.Error("Expected valid window ID") - } - - widMissing := coord.GetRecordingWindowID("missing") - if widMissing != "" { - t.Error("Expected empty window ID for missing alert") - } -} diff --git a/internal/ai/tools/executor.go b/internal/ai/tools/executor.go index bf045aca2..48f0468b8 100644 --- a/internal/ai/tools/executor.go +++ b/internal/ai/tools/executor.go @@ -559,9 +559,9 @@ type ExecutorConfig struct { AgentProfileManager AgentProfileManager // Optional providers - intelligence - IncidentRecorderProvider IncidentRecorderProvider - EventCorrelatorProvider EventCorrelatorProvider - KnowledgeStoreProvider KnowledgeStoreProvider + IncidentArchiveProvider IncidentArchiveProvider + EventCorrelatorProvider EventCorrelatorProvider + KnowledgeStoreProvider KnowledgeStoreProvider // Optional providers - discovery DiscoveryProvider DiscoveryProvider @@ -627,9 +627,9 @@ type PulseToolExecutor struct { agentProfileManager AgentProfileManager // Intelligence providers - incidentRecorderProvider IncidentRecorderProvider - eventCorrelatorProvider EventCorrelatorProvider - knowledgeStoreProvider KnowledgeStoreProvider + incidentArchiveProvider IncidentArchiveProvider + eventCorrelatorProvider EventCorrelatorProvider + knowledgeStoreProvider KnowledgeStoreProvider // Discovery provider discoveryProvider DiscoveryProvider @@ -750,7 +750,7 @@ func NewPulseToolExecutor(cfg ExecutorConfig) *PulseToolExecutor { metadataUpdater: cfg.MetadataUpdater, findingsManager: cfg.FindingsManager, agentProfileManager: cfg.AgentProfileManager, - incidentRecorderProvider: cfg.IncidentRecorderProvider, + incidentArchiveProvider: cfg.IncidentArchiveProvider, eventCorrelatorProvider: cfg.EventCorrelatorProvider, knowledgeStoreProvider: cfg.KnowledgeStoreProvider, discoveryProvider: cfg.DiscoveryProvider, @@ -827,7 +827,7 @@ func (e *PulseToolExecutor) Clone() *PulseToolExecutor { metadataUpdater: e.metadataUpdater, findingsManager: e.findingsManager, agentProfileManager: e.agentProfileManager, - incidentRecorderProvider: e.incidentRecorderProvider, + incidentArchiveProvider: e.incidentArchiveProvider, eventCorrelatorProvider: e.eventCorrelatorProvider, knowledgeStoreProvider: e.knowledgeStoreProvider, discoveryProvider: e.discoveryProvider, @@ -1019,9 +1019,9 @@ func (e *PulseToolExecutor) SetUpdatesProvider(provider UpdatesProvider) { e.updatesProvider = provider } -// SetIncidentRecorderProvider sets the incident recorder provider -func (e *PulseToolExecutor) SetIncidentRecorderProvider(provider IncidentRecorderProvider) { - e.incidentRecorderProvider = provider +// SetIncidentArchiveProvider sets the read-only legacy incident archive provider +func (e *PulseToolExecutor) SetIncidentArchiveProvider(provider IncidentArchiveProvider) { + e.incidentArchiveProvider = provider } // SetEventCorrelatorProvider sets the event correlator provider @@ -1221,7 +1221,7 @@ func (e *PulseToolExecutor) isToolAvailable(name string) bool { case agentcapabilities.PulseDiscoveryToolName: return e.discoveryProvider != nil case agentcapabilities.PulseKnowledgeToolName: - return e.actionAuditStore != nil || e.knowledgeStoreProvider != nil || e.incidentRecorderProvider != nil || e.eventCorrelatorProvider != nil + return e.actionAuditStore != nil || e.knowledgeStoreProvider != nil || e.incidentArchiveProvider != nil || e.eventCorrelatorProvider != nil case agentcapabilities.PulsePMGToolName: return e.hasReadState() case agentcapabilities.PulseSummarizeToolName: diff --git a/internal/ai/tools/incident_history_test.go b/internal/ai/tools/incident_history_test.go index 0b8ff365f..cc21e2f80 100644 --- a/internal/ai/tools/incident_history_test.go +++ b/internal/ai/tools/incident_history_test.go @@ -4,11 +4,14 @@ import ( "context" "encoding/json" "errors" + "os" + "path/filepath" "strings" "testing" "time" "github.com/rcourtman/pulse-go-rewrite/internal/agentcapabilities" + "github.com/rcourtman/pulse-go-rewrite/internal/metrics" "github.com/rcourtman/pulse-go-rewrite/internal/unifiedresources" "github.com/stretchr/testify/require" ) @@ -51,11 +54,10 @@ func (failedIncidentHistoryStore) GetRecentChanges(string, time.Time, int) ([]un return nil, errors.New("history store unavailable") } -type incidentArchiveFixture struct{ window *IncidentWindow } +type incidentArchiveFixture struct{ window *metrics.IncidentWindow } -func (s incidentArchiveFixture) GetWindow(string) *IncidentWindow { return s.window } -func (s incidentArchiveFixture) GetWindowsForResource(string, int) []*IncidentWindow { - panic("canonical incident reads must not enumerate legacy recordings") +func (s incidentArchiveFixture) GetWindow(string, string) (*metrics.IncidentWindow, error) { + return s.window, nil } func TestIncidentHistoryRetainsCanonicalEvidence(t *testing.T) { @@ -76,7 +78,7 @@ func TestIncidentHistoryRetainsCanonicalEvidence(t *testing.T) { } // No live inventory or incident recorder is necessary to read a retained // event for a container that has since been removed. - exec := NewPulseToolExecutor(ExecutorConfig{ActionAuditStore: store, IncidentRecorderProvider: incidentArchiveFixture{}}) + exec := NewPulseToolExecutor(ExecutorConfig{ActionAuditStore: store, IncidentArchiveProvider: incidentArchiveFixture{}}) require.True(t, exec.isToolAvailable(agentcapabilities.PulseKnowledgeToolName)) for _, tc := range []struct { name string @@ -137,7 +139,7 @@ func TestIncidentHistoryUnavailableAndInvalid(t *testing.T) { {"large limit", unifiedresources.NewMemoryStore(), map[string]interface{}{"resource_id": "app-container-1", "limit": float64(201)}}, } { t.Run(tc.name, func(t *testing.T) { - exec := NewPulseToolExecutor(ExecutorConfig{ActionAuditStore: tc.store, IncidentRecorderProvider: incidentArchiveFixture{}}) + exec := NewPulseToolExecutor(ExecutorConfig{ActionAuditStore: tc.store, IncidentArchiveProvider: incidentArchiveFixture{}}) tc.input["action"] = "incidents" result, err := exec.registry.Execute(context.Background(), exec, agentcapabilities.PulseKnowledgeToolName, tc.input) require.NoError(t, err) @@ -150,8 +152,8 @@ func TestIncidentHistoryUnavailableAndInvalid(t *testing.T) { } func TestIncidentHistoryLegacyArchiveIsResourceBound(t *testing.T) { - window := &IncidentWindow{ID: "archive-1", ResourceID: "vm-1"} - exec := NewPulseToolExecutor(ExecutorConfig{IncidentRecorderProvider: incidentArchiveFixture{window}}) + window := &metrics.IncidentWindow{ID: "archive-1", ResourceID: "vm-1"} + exec := NewPulseToolExecutor(ExecutorConfig{IncidentArchiveProvider: incidentArchiveFixture{window}}) for _, id := range []string{"vm-1", "vm-2"} { result, err := exec.executeGetIncidentWindow(context.Background(), map[string]interface{}{"resource_id": id, "window_id": window.ID}) require.NoError(t, err) @@ -164,4 +166,56 @@ func TestIncidentHistoryLegacyArchiveIsResourceBound(t *testing.T) { require.NotContains(t, result.Content[0].Text, "vm-1") } } + result, err := exec.executeGetIncidentWindow(context.Background(), map[string]interface{}{"resource_id": window.ResourceID, "window_id": "wrong-window"}) + require.NoError(t, err) + require.True(t, result.IsError, "a provider cannot substitute another archived window") + +} + +func TestIncidentHistoryArchiveReadOutcomes(t *testing.T) { + dir := t.TempDir() + archivePath := filepath.Join(dir, "incident_windows.json") + archive := metrics.NewIncidentArchive(dir) + exec := NewPulseToolExecutor(ExecutorConfig{IncidentArchiveProvider: archive}) + original := `{"completed_windows":[{"id":"saved-window","resource_id":"vm-archive","status":"recording","start_time":"2020-01-02T03:04:05Z","data_points":[{"timestamp":"2020-01-02T03:04:06Z","metrics":{"cpu":12.5},"metadata":{"source":"cached"}}],"summary":{"duration_ms":60000000000,"anomalies":["stored observation"]}}]}` + for _, tc := range []struct { + name, raw, resource, window string + wantError bool + }{ + {"archive unavailable", "", "vm-archive", "saved-window", true}, + {"archive malformed", "{", "vm-archive", "saved-window", true}, + {"archive success", original, "vm-archive", "saved-window", false}, + {"archive wrong resource", original, "vm-other", "saved-window", true}, + {"archive missing window", original, "vm-archive", "missing", true}, + } { + t.Run(tc.name, func(t *testing.T) { + if tc.raw != "" { + require.NoError(t, os.WriteFile(archivePath, []byte(tc.raw), 0600)) + } + input := map[string]interface{}{"action": "incidents", "resource_id": tc.resource, "window_id": tc.window} + result, err := exec.registry.Execute(context.Background(), exec, agentcapabilities.PulseKnowledgeToolName, input) + require.NoError(t, err) + require.Equal(t, tc.wantError, result.IsError, result.Content) + if !tc.wantError { + var got struct { + Window *metrics.IncidentWindow `json:"window"` + ReadOnly bool `json:"archive_read_only"` + DurationUnit string `json:"summary_duration_unit"` + } + require.NoError(t, json.Unmarshal([]byte(result.Content[0].Text), &got)) + require.True(t, got.ReadOnly) + require.Equal(t, "nanoseconds", got.DurationUnit) + require.Equal(t, time.Minute, got.Window.Summary.Duration) + require.Equal(t, metrics.IncidentWindowStatusRecording, got.Window.Status) + require.Equal(t, "cached", got.Window.DataPoints[0].Metadata["source"]) + require.Equal(t, []string{"stored observation"}, got.Window.Summary.Anomalies) + require.Contains(t, result.Content[0].Text, "does not mean recording is active") + } else if tc.name == "archive wrong resource" { + require.NotContains(t, result.Content[0].Text, "stored observation") + } + capture, err := json.Marshal(map[string]any{"case": tc.name, "input": input, "result": result}) + require.NoError(t, err) + t.Logf("INCIDENT_EVIDENCE %s", capture) + }) + } } diff --git a/internal/ai/tools/tools_knowledge.go b/internal/ai/tools/tools_knowledge.go index 7af911beb..a32c544d8 100644 --- a/internal/ai/tools/tools_knowledge.go +++ b/internal/ai/tools/tools_knowledge.go @@ -7,44 +7,14 @@ import ( "time" "github.com/rcourtman/pulse-go-rewrite/internal/agentcapabilities" + "github.com/rcourtman/pulse-go-rewrite/internal/metrics" "github.com/rcourtman/pulse-go-rewrite/internal/unifiedresources" ) -// IncidentRecorderProvider provides access to incident recording data -type IncidentRecorderProvider interface { - GetWindowsForResource(resourceID string, limit int) []*IncidentWindow - GetWindow(windowID string) *IncidentWindow -} - -// IncidentWindow represents a high-frequency recording window during an incident -type IncidentWindow struct { - ID string `json:"id"` - ResourceID string `json:"resource_id"` - ResourceName string `json:"resource_name,omitempty"` - ResourceType string `json:"resource_type,omitempty"` - TriggerType string `json:"trigger_type"` - TriggerID string `json:"trigger_id,omitempty"` - StartTime time.Time `json:"start_time"` - EndTime *time.Time `json:"end_time,omitempty"` - Status string `json:"status"` - DataPoints []IncidentDataPoint `json:"data_points"` - Summary *IncidentSummary `json:"summary,omitempty"` -} - -// IncidentDataPoint represents a single data point in an incident window -type IncidentDataPoint struct { - Timestamp time.Time `json:"timestamp"` - Metrics map[string]float64 `json:"metrics"` -} - -// IncidentSummary provides computed statistics about an incident window -type IncidentSummary struct { - Duration time.Duration `json:"duration_ms"` - DataPoints int `json:"data_points"` - Peaks map[string]float64 `json:"peaks"` - Lows map[string]float64 `json:"lows"` - Averages map[string]float64 `json:"averages"` - Changes map[string]float64 `json:"changes"` +// IncidentArchiveProvider provides explicit, resource-bound reads of saved +// legacy recordings. Live incident evidence comes from the canonical timeline. +type IncidentArchiveProvider interface { + GetWindow(resourceID, windowID string) (*metrics.IncidentWindow, error) } // EventCorrelatorProvider provides access to correlated events @@ -184,17 +154,22 @@ func (e *PulseToolExecutor) executeGetIncidentWindow(_ context.Context, args map // Isolate legacy recordings from the canonical timeline. Their sample times // are recorder timestamps, not verified source observation timestamps. if windowID != "" { - if e.incidentRecorderProvider == nil { + if e.incidentArchiveProvider == nil { return NewErrorResult(fmt.Errorf("legacy incident recording archive is unavailable")), nil } - window := e.incidentRecorderProvider.GetWindow(windowID) - if window == nil || window.ResourceID != resourceID { + window, err := e.incidentArchiveProvider.GetWindow(resourceID, windowID) + if err != nil { + return NewErrorResult(fmt.Errorf("read legacy incident recording archive: %w", err)), nil + } + if window == nil || window.ResourceID != resourceID || window.ID != windowID { return NewErrorResult(fmt.Errorf("legacy incident recording not found for the requested resource")), nil } return NewJSONResult(map[string]interface{}{ - "source": "legacy_incident_recording", - "window": window, - "evidence_limit": "Recording timestamps do not establish when the source measured each value. Repeated values may be cached observations. This archive is not the canonical incident timeline.", + "source": "legacy_incident_recording", + "archive_read_only": true, + "summary_duration_unit": "nanoseconds", + "window": window, + "evidence_limit": "Recording timestamps do not establish when the source measured each value. Repeated values may be cached observations. The legacy summary.duration_ms field contains nanoseconds. Stored recording status is historical and does not mean recording is active. This archive is not the canonical incident timeline.", }), nil } diff --git a/internal/api/ai_handler.go b/internal/api/ai_handler.go index f7ec08655..aff2273af 100644 --- a/internal/api/ai_handler.go +++ b/internal/api/ai_handler.go @@ -78,7 +78,7 @@ type AIService interface { SetFindingsManager(manager chat.FindingsManager) SetMetadataUpdater(updater chat.MetadataUpdater) SetKnowledgeStoreProvider(provider chat.KnowledgeStoreProvider) - SetIncidentRecorderProvider(provider chat.IncidentRecorderProvider) + SetIncidentArchiveProvider(provider chat.IncidentArchiveProvider) SetEventCorrelatorProvider(provider chat.EventCorrelatorProvider) SetDiscoveryProvider(provider chat.AssistantDiscoveryProvider) SetUnifiedResourceProvider(provider chat.AssistantUnifiedResourceProvider) diff --git a/internal/api/ai_handler_recovery_wiring_test.go b/internal/api/ai_handler_recovery_wiring_test.go index 8a15ba9c3..c04cd38bf 100644 --- a/internal/api/ai_handler_recovery_wiring_test.go +++ b/internal/api/ai_handler_recovery_wiring_test.go @@ -88,15 +88,15 @@ func (s *capturingAIService) SetGuestConfigProvider(provider chat.AssistantGuest func (s *capturingAIService) SetAppContainerConfigProvider(provider chat.AssistantAppContainerConfigProvider) { s.appContainerConfigProvider = provider } -func (s *capturingAIService) SetBackupProvider(provider chat.AssistantBackupProvider) {} -func (s *capturingAIService) SetDiskHealthProvider(provider chat.AssistantDiskHealthProvider) {} -func (s *capturingAIService) SetUpdatesProvider(provider chat.AssistantUpdatesProvider) {} -func (s *capturingAIService) SetFindingsManager(manager chat.FindingsManager) {} -func (s *capturingAIService) SetMetadataUpdater(updater chat.MetadataUpdater) {} -func (s *capturingAIService) SetKnowledgeStoreProvider(provider chat.KnowledgeStoreProvider) {} -func (s *capturingAIService) SetIncidentRecorderProvider(provider chat.IncidentRecorderProvider) {} -func (s *capturingAIService) SetEventCorrelatorProvider(provider chat.EventCorrelatorProvider) {} -func (s *capturingAIService) SetDiscoveryProvider(provider chat.AssistantDiscoveryProvider) {} +func (s *capturingAIService) SetBackupProvider(provider chat.AssistantBackupProvider) {} +func (s *capturingAIService) SetDiskHealthProvider(provider chat.AssistantDiskHealthProvider) {} +func (s *capturingAIService) SetUpdatesProvider(provider chat.AssistantUpdatesProvider) {} +func (s *capturingAIService) SetFindingsManager(manager chat.FindingsManager) {} +func (s *capturingAIService) SetMetadataUpdater(updater chat.MetadataUpdater) {} +func (s *capturingAIService) SetKnowledgeStoreProvider(provider chat.KnowledgeStoreProvider) {} +func (s *capturingAIService) SetIncidentArchiveProvider(provider chat.IncidentArchiveProvider) {} +func (s *capturingAIService) SetEventCorrelatorProvider(provider chat.EventCorrelatorProvider) {} +func (s *capturingAIService) SetDiscoveryProvider(provider chat.AssistantDiscoveryProvider) {} func (s *capturingAIService) SetUnifiedResourceProvider(provider chat.AssistantUnifiedResourceProvider) { } func (s *capturingAIService) SetAppContainerActionProvider(provider chat.AssistantAppContainerActionProvider) { diff --git a/internal/api/ai_handler_test.go b/internal/api/ai_handler_test.go index 5cd14e0cc..ad572870c 100644 --- a/internal/api/ai_handler_test.go +++ b/internal/api/ai_handler_test.go @@ -281,7 +281,7 @@ func (m *MockAIService) SetMetadataUpdater(updater chat.MetadataUpdater) { m.Cal func (m *MockAIService) SetKnowledgeStoreProvider(provider chat.KnowledgeStoreProvider) { m.Called(provider) } -func (m *MockAIService) SetIncidentRecorderProvider(provider chat.IncidentRecorderProvider) { +func (m *MockAIService) SetIncidentArchiveProvider(provider chat.IncidentArchiveProvider) { m.Called(provider) } func (m *MockAIService) SetEventCorrelatorProvider(provider chat.EventCorrelatorProvider) { diff --git a/internal/api/ai_handlers.go b/internal/api/ai_handlers.go index 2a89407cd..43d2f2678 100644 --- a/internal/api/ai_handlers.go +++ b/internal/api/ai_handlers.go @@ -92,22 +92,20 @@ type AISettingsHandler struct { alertBridge *unified.AlertBridge // Bridge between alerts and unified store // Event-driven patrol (Phase 7) - triggerManager *ai.TriggerManager // Event-driven patrol trigger manager - incidentCoordinator *ai.IncidentCoordinator // Incident recording coordinator - incidentRecorder *metrics.IncidentRecorder // High-frequency incident recorder - intelligenceMu sync.RWMutex - proxmoxCorrelators map[string]*proxmox.EventCorrelator - learningStores map[string]*learning.LearningStore - forecastServices map[string]*forecast.Service - remediationEngines map[string]aicontracts.RemediationEngine - incidentStores map[string]*memory.IncidentStore - circuitBreakers map[string]*circuit.Breaker - discoveryStores map[string]*servicediscovery.Store - unifiedStores map[string]*unified.UnifiedStore - alertBridges map[string]*unified.AlertBridge - triggerManagers map[string]*ai.TriggerManager - incidentCoordinators map[string]*ai.IncidentCoordinator - incidentRecorders map[string]*metrics.IncidentRecorder + triggerManager *ai.TriggerManager // Event-driven patrol trigger manager + incidentArchive *metrics.IncidentArchive // Read-only legacy incident archive + intelligenceMu sync.RWMutex + proxmoxCorrelators map[string]*proxmox.EventCorrelator + learningStores map[string]*learning.LearningStore + forecastServices map[string]*forecast.Service + remediationEngines map[string]aicontracts.RemediationEngine + incidentStores map[string]*memory.IncidentStore + circuitBreakers map[string]*circuit.Breaker + discoveryStores map[string]*servicediscovery.Store + unifiedStores map[string]*unified.UnifiedStore + alertBridges map[string]*unified.AlertBridge + triggerManagers map[string]*ai.TriggerManager + incidentArchives map[string]*metrics.IncidentArchive // Investigation orchestration (Patrol Autonomy) chatHandler *AIHandler // Chat service handler for investigations @@ -318,25 +316,24 @@ func NewAISettingsHandler(mtp *config.MultiTenantPersistence, mtm *monitoring.Mu } handler := &AISettingsHandler{ - mtPersistence: mtp, - mtMonitor: mtm, - defaultConfig: defaultConfig, - defaultPersistence: defaultPersistence, - hostedMode: hostedMode, - aiServices: make(map[string]*ai.Service), - agentServer: agentServer, - proxmoxCorrelators: make(map[string]*proxmox.EventCorrelator), - learningStores: make(map[string]*learning.LearningStore), - forecastServices: make(map[string]*forecast.Service), - remediationEngines: make(map[string]aicontracts.RemediationEngine), - incidentStores: make(map[string]*memory.IncidentStore), - circuitBreakers: make(map[string]*circuit.Breaker), - discoveryStores: make(map[string]*servicediscovery.Store), - unifiedStores: make(map[string]*unified.UnifiedStore), - alertBridges: make(map[string]*unified.AlertBridge), - triggerManagers: make(map[string]*ai.TriggerManager), - incidentCoordinators: make(map[string]*ai.IncidentCoordinator), - incidentRecorders: make(map[string]*metrics.IncidentRecorder), + mtPersistence: mtp, + mtMonitor: mtm, + defaultConfig: defaultConfig, + defaultPersistence: defaultPersistence, + hostedMode: hostedMode, + aiServices: make(map[string]*ai.Service), + agentServer: agentServer, + proxmoxCorrelators: make(map[string]*proxmox.EventCorrelator), + learningStores: make(map[string]*learning.LearningStore), + forecastServices: make(map[string]*forecast.Service), + remediationEngines: make(map[string]aicontracts.RemediationEngine), + incidentStores: make(map[string]*memory.IncidentStore), + circuitBreakers: make(map[string]*circuit.Breaker), + discoveryStores: make(map[string]*servicediscovery.Store), + unifiedStores: make(map[string]*unified.UnifiedStore), + alertBridges: make(map[string]*unified.AlertBridge), + triggerManagers: make(map[string]*ai.TriggerManager), + incidentArchives: make(map[string]*metrics.IncidentArchive), } defaultAIService = ai.NewService(defaultPersistence, tenantAgentServerForOrganization(agentServer, "default")) @@ -1198,11 +1195,8 @@ func (h *AISettingsHandler) ensureIntelligenceMapsLocked() { if h.triggerManagers == nil { h.triggerManagers = make(map[string]*ai.TriggerManager) } - if h.incidentCoordinators == nil { - h.incidentCoordinators = make(map[string]*ai.IncidentCoordinator) - } - if h.incidentRecorders == nil { - h.incidentRecorders = make(map[string]*metrics.IncidentRecorder) + if h.incidentArchives == nil { + h.incidentArchives = make(map[string]*metrics.IncidentArchive) } } @@ -1505,96 +1499,49 @@ func (h *AISettingsHandler) GetTriggerManagerForOrg(orgID string) *ai.TriggerMan return nil } -// SetIncidentCoordinator sets the incident recording coordinator -func (h *AISettingsHandler) SetIncidentCoordinator(coordinator *ai.IncidentCoordinator) { - h.SetIncidentCoordinatorForOrg("default", coordinator) +// SetIncidentArchive sets the read-only legacy incident archive +func (h *AISettingsHandler) SetIncidentArchive(archive *metrics.IncidentArchive) { + h.SetIncidentArchiveForOrg("default", archive) } -// SetIncidentCoordinatorForOrg sets the incident recording coordinator for an org. -func (h *AISettingsHandler) SetIncidentCoordinatorForOrg(orgID string, coordinator *ai.IncidentCoordinator) { +// SetIncidentArchiveForOrg sets the read-only legacy incident archive for an org. +func (h *AISettingsHandler) SetIncidentArchiveForOrg(orgID string, archive *metrics.IncidentArchive) { if h == nil { return } orgID = normalizeAIIntelligenceOrgID(orgID) h.intelligenceMu.Lock() h.ensureIntelligenceMapsLocked() - if coordinator == nil { - delete(h.incidentCoordinators, orgID) + if archive == nil { + delete(h.incidentArchives, orgID) } else { - h.incidentCoordinators[orgID] = coordinator + h.incidentArchives[orgID] = archive } h.intelligenceMu.Unlock() if orgID == "default" { - h.incidentCoordinator = coordinator + h.incidentArchive = archive } } -// GetIncidentCoordinator returns the incident recording coordinator -func (h *AISettingsHandler) GetIncidentCoordinator() *ai.IncidentCoordinator { - return h.GetIncidentCoordinatorForOrg("default") +// GetIncidentArchive returns the read-only legacy incident archive +func (h *AISettingsHandler) GetIncidentArchive() *metrics.IncidentArchive { + return h.GetIncidentArchiveForOrg("default") } -// GetIncidentCoordinatorForOrg returns the incident recording coordinator for an org. -func (h *AISettingsHandler) GetIncidentCoordinatorForOrg(orgID string) *ai.IncidentCoordinator { +// GetIncidentArchiveForOrg returns the read-only legacy incident archive for an org. +func (h *AISettingsHandler) GetIncidentArchiveForOrg(orgID string) *metrics.IncidentArchive { if h == nil { return nil } orgID = normalizeAIIntelligenceOrgID(orgID) h.intelligenceMu.RLock() - if coordinator := h.incidentCoordinators[orgID]; coordinator != nil { + if archive := h.incidentArchives[orgID]; archive != nil { h.intelligenceMu.RUnlock() - return coordinator + return archive } h.intelligenceMu.RUnlock() if orgID == "default" { - return h.incidentCoordinator - } - return nil -} - -// SetIncidentRecorder sets the high-frequency incident recorder -func (h *AISettingsHandler) SetIncidentRecorder(recorder *metrics.IncidentRecorder) { - h.SetIncidentRecorderForOrg("default", recorder) -} - -// SetIncidentRecorderForOrg sets the high-frequency incident recorder for an org. -func (h *AISettingsHandler) SetIncidentRecorderForOrg(orgID string, recorder *metrics.IncidentRecorder) { - if h == nil { - return - } - orgID = normalizeAIIntelligenceOrgID(orgID) - h.intelligenceMu.Lock() - h.ensureIntelligenceMapsLocked() - if recorder == nil { - delete(h.incidentRecorders, orgID) - } else { - h.incidentRecorders[orgID] = recorder - } - h.intelligenceMu.Unlock() - if orgID == "default" { - h.incidentRecorder = recorder - } -} - -// GetIncidentRecorder returns the high-frequency incident recorder -func (h *AISettingsHandler) GetIncidentRecorder() *metrics.IncidentRecorder { - return h.GetIncidentRecorderForOrg("default") -} - -// GetIncidentRecorderForOrg returns the high-frequency incident recorder for an org. -func (h *AISettingsHandler) GetIncidentRecorderForOrg(orgID string) *metrics.IncidentRecorder { - if h == nil { - return nil - } - orgID = normalizeAIIntelligenceOrgID(orgID) - h.intelligenceMu.RLock() - if recorder := h.incidentRecorders[orgID]; recorder != nil { - h.intelligenceMu.RUnlock() - return recorder - } - h.intelligenceMu.RUnlock() - if orgID == "default" { - return h.incidentRecorder + return h.incidentArchive } return nil } @@ -1656,44 +1603,6 @@ func (h *AISettingsHandler) ListTriggerManagers() map[string]*ai.TriggerManager return out } -// ListIncidentCoordinators returns incident coordinators keyed by org. -func (h *AISettingsHandler) ListIncidentCoordinators() map[string]*ai.IncidentCoordinator { - out := make(map[string]*ai.IncidentCoordinator) - if h == nil { - return out - } - h.intelligenceMu.RLock() - for orgID, coordinator := range h.incidentCoordinators { - if coordinator != nil { - out[orgID] = coordinator - } - } - h.intelligenceMu.RUnlock() - if _, ok := out["default"]; !ok && h.incidentCoordinator != nil { - out["default"] = h.incidentCoordinator - } - return out -} - -// ListIncidentRecorders returns incident recorders keyed by org. -func (h *AISettingsHandler) ListIncidentRecorders() map[string]*metrics.IncidentRecorder { - out := make(map[string]*metrics.IncidentRecorder) - if h == nil { - return out - } - h.intelligenceMu.RLock() - for orgID, recorder := range h.incidentRecorders { - if recorder != nil { - out[orgID] = recorder - } - } - h.intelligenceMu.RUnlock() - if _, ok := out["default"]; !ok && h.incidentRecorder != nil { - out["default"] = h.incidentRecorder - } - return out -} - // StopPatrol stops the background AI patrol service func (h *AISettingsHandler) StopPatrol() { if h.defaultAIService != nil { @@ -1760,10 +1669,8 @@ func (h *AISettingsHandler) RemoveTenantIntelligence(orgID string) { } var ( - bridge *unified.AlertBridge - trigger *ai.TriggerManager - coordinator *ai.IncidentCoordinator - recorder *metrics.IncidentRecorder + bridge *unified.AlertBridge + trigger *ai.TriggerManager ) h.intelligenceMu.Lock() @@ -1775,13 +1682,10 @@ func (h *AISettingsHandler) RemoveTenantIntelligence(orgID string) { delete(h.discoveryStores, orgID) bridge = h.alertBridges[orgID] trigger = h.triggerManagers[orgID] - coordinator = h.incidentCoordinators[orgID] - recorder = h.incidentRecorders[orgID] delete(h.unifiedStores, orgID) delete(h.alertBridges, orgID) delete(h.triggerManagers, orgID) - delete(h.incidentCoordinators, orgID) - delete(h.incidentRecorders, orgID) + delete(h.incidentArchives, orgID) delete(h.proxmoxCorrelators, orgID) h.intelligenceMu.Unlock() @@ -1791,12 +1695,6 @@ func (h *AISettingsHandler) RemoveTenantIntelligence(orgID string) { if trigger != nil { trigger.Stop() } - if coordinator != nil { - coordinator.Stop() - } - if recorder != nil { - recorder.Stop() - } } // GetAlertTriggeredAnalyzer returns the alert-triggered analyzer for wiring into alert callbacks diff --git a/internal/api/ai_handlers_setters_additional_test.go b/internal/api/ai_handlers_setters_additional_test.go index 6095aee26..6b9d9e286 100644 --- a/internal/api/ai_handlers_setters_additional_test.go +++ b/internal/api/ai_handlers_setters_additional_test.go @@ -50,16 +50,10 @@ func TestAISettingsHandler_SettersAndGetters(t *testing.T) { t.Fatalf("GetTriggerManager returned unexpected manager") } - coordinator := ai.NewIncidentCoordinator(ai.IncidentCoordinatorConfig{}) - handler.SetIncidentCoordinator(coordinator) - if handler.GetIncidentCoordinator() != coordinator { - t.Fatalf("GetIncidentCoordinator returned unexpected coordinator") - } - - recorder := &metrics.IncidentRecorder{} - handler.SetIncidentRecorder(recorder) - if handler.GetIncidentRecorder() != recorder { - t.Fatalf("GetIncidentRecorder returned unexpected recorder") + recorder := &metrics.IncidentArchive{} + handler.SetIncidentArchive(recorder) + if handler.GetIncidentArchive() != recorder { + t.Fatalf("GetIncidentArchive returned unexpected recorder") } handler.WireOrchestratorAfterChatStart() @@ -82,15 +76,19 @@ func TestAISettingsHandler_IntelligenceServicesAreOrgScoped(t *testing.T) { t.Fatalf("expected nil correlator for unrelated org, got %#v", got) } - defaultRecorder := &metrics.IncidentRecorder{} - tenantRecorder := &metrics.IncidentRecorder{} - handler.SetIncidentRecorderForOrg("default", defaultRecorder) - handler.SetIncidentRecorderForOrg("acme", tenantRecorder) - if got := handler.GetIncidentRecorder(); got != defaultRecorder { - t.Fatalf("expected default recorder, got %#v", got) + defaultArchive := &metrics.IncidentArchive{} + tenantArchive := &metrics.IncidentArchive{} + handler.SetIncidentArchiveForOrg("default", defaultArchive) + handler.SetIncidentArchiveForOrg("acme", tenantArchive) + if got := handler.GetIncidentArchive(); got != defaultArchive { + t.Fatalf("expected default archive, got %#v", got) } - if got := handler.GetIncidentRecorderForOrg("acme"); got != tenantRecorder { - t.Fatalf("expected tenant recorder, got %#v", got) + if got := handler.GetIncidentArchiveForOrg("acme"); got != tenantArchive { + t.Fatalf("expected tenant archive, got %#v", got) + } + + if got := handler.GetIncidentArchiveForOrg("unrelated"); got != nil { + t.Fatalf("unrelated org received an archive: %#v", got) } defaultLearningStore := learning.NewLearningStore(learning.LearningStoreConfig{}) @@ -201,13 +199,12 @@ func TestAISettingsHandler_IntelligenceServicesAreOrgScoped(t *testing.T) { func TestAISettingsHandler_RemoveTenantService_TrimsOrgID(t *testing.T) { handler := &AISettingsHandler{ - aiServices: map[string]*ai.Service{"acme": nil}, - investigationStores: map[string]aicontracts.InvestigationStore{"acme": nil}, - proxmoxCorrelators: map[string]*proxmox.EventCorrelator{"acme": nil}, - alertBridges: map[string]*unified.AlertBridge{"acme": nil}, - triggerManagers: map[string]*ai.TriggerManager{"acme": nil}, - incidentCoordinators: map[string]*ai.IncidentCoordinator{"acme": nil}, - incidentRecorders: map[string]*metrics.IncidentRecorder{"acme": nil}, + aiServices: map[string]*ai.Service{"acme": nil}, + investigationStores: map[string]aicontracts.InvestigationStore{"acme": nil}, + proxmoxCorrelators: map[string]*proxmox.EventCorrelator{"acme": nil}, + alertBridges: map[string]*unified.AlertBridge{"acme": nil}, + triggerManagers: map[string]*ai.TriggerManager{"acme": nil}, + incidentArchives: map[string]*metrics.IncidentArchive{"acme": nil}, } handler.RemoveTenantService(" acme ") @@ -227,10 +224,7 @@ func TestAISettingsHandler_RemoveTenantService_TrimsOrgID(t *testing.T) { if _, ok := handler.triggerManagers["acme"]; ok { t.Fatalf("expected trigger manager entry to be removed") } - if _, ok := handler.incidentCoordinators["acme"]; ok { - t.Fatalf("expected incident coordinator entry to be removed") - } - if _, ok := handler.incidentRecorders["acme"]; ok { + if _, ok := handler.incidentArchives["acme"]; ok { t.Fatalf("expected incident recorder entry to be removed") } } diff --git a/internal/api/ai_intelligence_handlers.go b/internal/api/ai_intelligence_handlers.go index 84311f122..7eba442a4 100644 --- a/internal/api/ai_intelligence_handlers.go +++ b/internal/api/ai_intelligence_handlers.go @@ -1420,20 +1420,14 @@ func (h *AISettingsHandler) HandleGetRecentIncidents(w http.ResponseWriter, r *h } } - // Get coordinator status - coordinator := h.GetIncidentCoordinatorForOrg(GetOrgID(r.Context())) - var activeCount int - if coordinator != nil { - activeCount = coordinator.GetActiveIncidentCount() - } - // Get incident data from patrol service svc := h.GetAIService(r.Context()) if svc == nil { if err := utils.WriteJSONResponse(w, map[string]interface{}{ - "incidents": []interface{}{}, - "active_count": activeCount, - "message": "Pulse Patrol service not available", + "incidents": []interface{}{}, + "active_count": nil, + "active_count_status": "not_measured", + "message": "Pulse Patrol service not available", }); err != nil { log.Error().Err(err).Msg("Failed to write incidents response") } @@ -1443,9 +1437,10 @@ func (h *AISettingsHandler) HandleGetRecentIncidents(w http.ResponseWriter, r *h patrol := svc.GetPatrolService() if patrol == nil { if err := utils.WriteJSONResponse(w, map[string]interface{}{ - "incidents": []interface{}{}, - "active_count": activeCount, - "message": "Patrol service not available", + "incidents": []interface{}{}, + "active_count": nil, + "active_count_status": "not_measured", + "message": "Patrol service not available", }); err != nil { log.Error().Err(err).Msg("Failed to write incidents response") } @@ -1456,9 +1451,10 @@ func (h *AISettingsHandler) HandleGetRecentIncidents(w http.ResponseWriter, r *h incidentStore := patrol.GetIncidentStore() if incidentStore == nil { if err := utils.WriteJSONResponse(w, map[string]interface{}{ - "incidents": []interface{}{}, - "active_count": activeCount, - "message": "Incident store not available", + "incidents": []interface{}{}, + "active_count": nil, + "active_count_status": "not_measured", + "message": "Incident store not available", }); err != nil { log.Error().Err(err).Msg("Failed to write incidents response") } @@ -1476,9 +1472,10 @@ func (h *AISettingsHandler) HandleGetRecentIncidents(w http.ResponseWriter, r *h // This is a limitation - we may want to add ListRecentIncidents to the store incidentSummary := incidentStore.FormatForPatrol(limit) if err := utils.WriteJSONResponse(w, map[string]interface{}{ - "incidents": []interface{}{}, - "incident_summary": incidentSummary, - "active_count": activeCount, + "incidents": []interface{}{}, + "incident_summary": incidentSummary, + "active_count": nil, + "active_count_status": "not_measured", }); err != nil { log.Error().Err(err).Msg("Failed to write incidents response") } @@ -1486,8 +1483,9 @@ func (h *AISettingsHandler) HandleGetRecentIncidents(w http.ResponseWriter, r *h } if err := utils.WriteJSONResponse(w, map[string]interface{}{ - "incidents": incidents, - "active_count": activeCount, + "incidents": incidents, + "active_count": nil, + "active_count_status": "not_measured", }); err != nil { log.Error().Err(err).Msg("Failed to write incidents response") } diff --git a/internal/api/ai_intelligence_handlers_remediation_additional_test.go b/internal/api/ai_intelligence_handlers_remediation_additional_test.go index e30d01154..243b31d14 100644 --- a/internal/api/ai_intelligence_handlers_remediation_additional_test.go +++ b/internal/api/ai_intelligence_handlers_remediation_additional_test.go @@ -9,7 +9,6 @@ import ( "testing" "time" - "github.com/rcourtman/pulse-go-rewrite/internal/ai" "github.com/rcourtman/pulse-go-rewrite/internal/ai/circuit" "github.com/rcourtman/pulse-go-rewrite/internal/ai/memory" "github.com/rcourtman/pulse-go-rewrite/internal/alerts" @@ -72,11 +71,6 @@ func setupIncidentHandler(t *testing.T) (*AISettingsHandler, *memory.IncidentSto store := memory.NewIncidentStore(memory.IncidentStoreConfig{DataDir: ""}) patrol.SetIncidentStore(store) - coordinator := ai.NewIncidentCoordinator(ai.IncidentCoordinatorConfig{EnableRecorder: false}) - coordinator.SetIncidentStore(store) - coordinator.Start() - handler.SetIncidentCoordinator(coordinator) - alert := &alerts.Alert{ ID: "alert-1", Type: "cpu", @@ -86,7 +80,7 @@ func setupIncidentHandler(t *testing.T) (*AISettingsHandler, *memory.IncidentSto StartTime: time.Now(), LastSeen: time.Now(), } - coordinator.OnAlertFired(alert) + store.RecordAlertFired(alert) return handler, store } @@ -110,8 +104,8 @@ func TestHandleGetRecentIncidents(t *testing.T) { if len(incidents) != 1 { t.Fatalf("expected 1 incident, got %d", len(incidents)) } - if resp["active_count"].(float64) < 1 { - t.Fatalf("expected active_count >= 1") + if resp["active_count"] != nil || resp["active_count_status"] != "not_measured" { + t.Fatalf("saved incident context must not imply a measured live count: %#v", resp) } } @@ -173,3 +167,30 @@ func TestHandleGetIncidentData(t *testing.T) { t.Fatalf("expected formatted_context to be populated") } } + +func TestHandleGetRecentIncidentsCountIsNotMeasured(t *testing.T) { + withMemory, _ := setupIncidentHandler(t) + for _, tc := range []struct { + name string + handler *AISettingsHandler + query string + }{ + {"unavailable", &AISettingsHandler{}, ""}, + {"fleet context", withMemory, ""}, + {"resource context", withMemory, "?resource_id=res-1"}, + {"empty resource context", withMemory, "?resource_id=absent"}, + } { + t.Run(tc.name, func(t *testing.T) { + rec := httptest.NewRecorder() + tc.handler.HandleGetRecentIncidents(rec, httptest.NewRequest(http.MethodGet, "/api/ai/incidents"+tc.query, nil)) + var body map[string]interface{} + if err := json.Unmarshal(rec.Body.Bytes(), &body); err != nil { + t.Fatal(err) + } + count, present := body["active_count"] + if !present || count != nil || body["active_count_status"] != "not_measured" { + t.Fatalf("archive/context presence cannot establish live count: %#v", body) + } + }) + } +} diff --git a/internal/api/contract_test.go b/internal/api/contract_test.go index bbeb5fb28..b04e57829 100644 --- a/internal/api/contract_test.go +++ b/internal/api/contract_test.go @@ -18650,8 +18650,6 @@ func TestContract_AssistantProviderSeamsDoNotUseMCPTerminology(t *testing.T) { } intelligenceAdapters := string(intelligenceAdaptersSource) for _, fragment := range []string{ - `type IncidentRecorderToolAdapter struct`, - `func NewIncidentRecorderToolAdapter(recorder IncidentRecorderSource) *IncidentRecorderToolAdapter`, `type EventCorrelatorToolAdapter struct`, `func NewEventCorrelatorToolAdapter(correlator EventCorrelatorSource) *EventCorrelatorToolAdapter`, } { diff --git a/internal/api/router.go b/internal/api/router.go index 6b7157532..6cb59249f 100644 --- a/internal/api/router.go +++ b/internal/api/router.go @@ -2873,41 +2873,8 @@ func (r *Router) initializeAIIntelligenceServices(ctx context.Context, orgID, da log.Info().Msg("AI Intelligence: Event-driven trigger manager initialized and started") } - // 12. Initialize incident coordinator for high-frequency recording - if patrol != nil { - incidentCoordinator := ai.NewIncidentCoordinator(ai.DefaultIncidentCoordinatorConfig()) - - // Wire the incident store if available - if incidentStore := patrol.GetIncidentStore(); incidentStore != nil { - incidentCoordinator.SetIncidentStore(incidentStore) - } - - // Create metrics adapter for incident recorder (ReadState is sole source since SRC-03m) - var metricsAdapter *adapters.MetricsAdapter - if monitor != nil { - metricsAdapter = adapters.NewMetricsAdapter(monitor.GetUnifiedReadState()) - } - - // Initialize and wire the incident recorder (high-frequency metrics) - if metricsAdapter != nil { - recorderCfg := metrics.DefaultIncidentRecorderConfig() - recorderCfg.DataDir = dataDir - recorder := metrics.NewIncidentRecorder(recorderCfg) - recorder.SetMetricsProvider(metricsAdapter) - recorder.Start() - incidentCoordinator.SetRecorder(recorder) - r.aiSettingsHandler.SetIncidentRecorderForOrg(orgID, recorder) - log.Info().Msg("AI Intelligence: Incident recorder initialized and started") - } - - // Start the coordinator - incidentCoordinator.Start() - - // Store reference - r.aiSettingsHandler.SetIncidentCoordinatorForOrg(orgID, incidentCoordinator) - - log.Info().Msg("AI Intelligence: Incident coordinator initialized and started") - } + // Legacy recordings are available only through explicit archive lookup. + r.aiSettingsHandler.SetIncidentArchiveForOrg(orgID, metrics.NewIncidentArchive(dataDir)) log.Info().Msg("AI Intelligence: All Phase 6 & 7 services initialized successfully") } @@ -2965,24 +2932,6 @@ func (r *Router) ShutdownAIIntelligence() { log.Debug().Str("org_id", orgID).Msg("AI Intelligence: Trigger manager stopped") } - // 4. Stop incident coordinators (stop high-frequency recording) - for orgID, incidentCoordinator := range r.aiSettingsHandler.ListIncidentCoordinators() { - if incidentCoordinator == nil { - continue - } - incidentCoordinator.Stop() - log.Debug().Str("org_id", orgID).Msg("AI Intelligence: Incident coordinator stopped") - } - - // 4b. Stop incident recorders (stop background sampling) - for orgID, incidentRecorder := range r.aiSettingsHandler.ListIncidentRecorders() { - if incidentRecorder == nil { - continue - } - incidentRecorder.Stop() - log.Debug().Str("org_id", orgID).Msg("AI Intelligence: Incident recorder stopped") - } - // 5. Cleanup learning stores (removes old records, persists if data dir configured) for orgID, learningStore := range r.aiSettingsHandler.ListLearningStores() { if learningStore == nil { @@ -3353,15 +3302,15 @@ func (r *Router) wireAIChatDependenciesForService(ctx context.Context, service A } // Wire intelligence providers for Assistant tools. - // - IncidentRecorderProvider: high-frequency incident data (pulse_get_incident_window) + // - IncidentArchiveProvider: explicit reads of saved legacy recordings // - EventCorrelatorProvider: Proxmox events (pulse_correlate_events) // - KnowledgeStoreProvider: notes (pulse_remember, pulse_recall) - // Wire incident recorder provider (high-frequency incident data) + // Wire the org-pinned archive reader without creating or sampling data. if r.aiSettingsHandler != nil { - if recorder := r.aiSettingsHandler.GetIncidentRecorderForOrg(orgID); recorder != nil { - service.SetIncidentRecorderProvider(&incidentRecorderProviderWrapper{recorder: recorder}) - log.Debug().Msg("AI chat: Incident recorder provider wired") + if archive := r.aiSettingsHandler.GetIncidentArchiveForOrg(orgID); archive != nil { + service.SetIncidentArchiveProvider(archive) + log.Debug().Msg("AI chat: Incident archive provider wired") } } @@ -3476,82 +3425,6 @@ func (w *forecastResourceIterator) ForecastStoragePools() []forecast.ResourceInf return result } -// incidentRecorderProviderWrapper adapts metrics.IncidentRecorder to tools.IncidentRecorderProvider. -type incidentRecorderProviderWrapper struct { - recorder *metrics.IncidentRecorder -} - -func (w *incidentRecorderProviderWrapper) GetWindowsForResource(resourceID string, limit int) []*tools.IncidentWindow { - if w.recorder == nil { - return nil - } - - windows := w.recorder.GetWindowsForResource(resourceID, limit) - if len(windows) == 0 { - return nil - } - - result := make([]*tools.IncidentWindow, 0, len(windows)) - for _, window := range windows { - if window == nil { - continue - } - result = append(result, convertIncidentWindow(window)) - } - return result -} - -func (w *incidentRecorderProviderWrapper) GetWindow(windowID string) *tools.IncidentWindow { - if w.recorder == nil { - return nil - } - window := w.recorder.GetWindow(windowID) - if window == nil { - return nil - } - return convertIncidentWindow(window) -} - -func convertIncidentWindow(window *metrics.IncidentWindow) *tools.IncidentWindow { - if window == nil { - return nil - } - - points := make([]tools.IncidentDataPoint, 0, len(window.DataPoints)) - for _, point := range window.DataPoints { - points = append(points, tools.IncidentDataPoint{ - Timestamp: point.Timestamp, - Metrics: point.Metrics, - }) - } - - var summary *tools.IncidentSummary - if window.Summary != nil { - summary = &tools.IncidentSummary{ - Duration: window.Summary.Duration, - DataPoints: window.Summary.DataPoints, - Peaks: window.Summary.Peaks, - Lows: window.Summary.Lows, - Averages: window.Summary.Averages, - Changes: window.Summary.Changes, - } - } - - return &tools.IncidentWindow{ - ID: window.ID, - ResourceID: window.ResourceID, - ResourceName: window.ResourceName, - ResourceType: window.ResourceType, - TriggerType: window.TriggerType, - TriggerID: window.TriggerID, - StartTime: window.StartTime, - EndTime: window.EndTime, - Status: string(window.Status), - DataPoints: points, - Summary: summary, - } -} - func (r *Router) publishActionCompletedAgentEvent(broadcaster *AgentEventBroadcaster, record unifiedresources.ActionAuditRecord) { if broadcaster == nil { return diff --git a/internal/api/router_wrappers_additional_test.go b/internal/api/router_wrappers_additional_test.go index d52422847..73d9488a8 100644 --- a/internal/api/router_wrappers_additional_test.go +++ b/internal/api/router_wrappers_additional_test.go @@ -8,7 +8,6 @@ import ( "github.com/rcourtman/pulse-go-rewrite/internal/ai/patterns" "github.com/rcourtman/pulse-go-rewrite/internal/ai/proxmox" "github.com/rcourtman/pulse-go-rewrite/internal/config" - "github.com/rcourtman/pulse-go-rewrite/internal/metrics" "github.com/rcourtman/pulse-go-rewrite/internal/models" "github.com/rcourtman/pulse-go-rewrite/internal/monitoring" unifiedresources "github.com/rcourtman/pulse-go-rewrite/internal/unifiedresources" @@ -52,84 +51,6 @@ func TestForecastResourceIterator_NilReadState(t *testing.T) { } } -func TestIncidentRecorderProviderWrapper(t *testing.T) { - now := time.Now().UTC() - end := now.Add(5 * time.Minute) - - activeWindow := &metrics.IncidentWindow{ - ID: "win-active", - ResourceID: "res-1", - ResourceName: "Resource", - ResourceType: "vm", - TriggerType: "alert", - TriggerID: "alert-1", - StartTime: now, - EndTime: &end, - Status: metrics.IncidentWindowStatusRecording, - DataPoints: []metrics.IncidentDataPoint{ - {Timestamp: now, Metrics: map[string]float64{"cpu": 10}}, - }, - Summary: &metrics.IncidentSummary{ - Duration: 5 * time.Minute, - DataPoints: 1, - Peaks: map[string]float64{"cpu": 10}, - Lows: map[string]float64{"cpu": 5}, - Averages: map[string]float64{"cpu": 7}, - Changes: map[string]float64{"cpu": 2}, - }, - } - - completedWindow := &metrics.IncidentWindow{ - ID: "win-complete", - ResourceID: "res-1", - ResourceName: "Resource", - ResourceType: "vm", - TriggerType: "alert", - TriggerID: "alert-2", - StartTime: now.Add(-time.Hour), - Status: metrics.IncidentWindowStatusComplete, - DataPoints: []metrics.IncidentDataPoint{ - {Timestamp: now.Add(-time.Hour), Metrics: map[string]float64{"cpu": 20}}, - }, - } - - recorder := metrics.NewIncidentRecorder(metrics.DefaultIncidentRecorderConfig()) - setUnexportedField(t, recorder, "activeWindows", map[string]*metrics.IncidentWindow{"win-active": activeWindow}) - setUnexportedField(t, recorder, "completedWindows", []*metrics.IncidentWindow{completedWindow}) - - wrapper := &incidentRecorderProviderWrapper{recorder: recorder} - windows := wrapper.GetWindowsForResource("res-1", 10) - if len(windows) != 2 { - t.Fatalf("expected 2 windows, got %d", len(windows)) - } - - ids := []string{windows[0].ID, windows[1].ID} - if !containsStringSlice(ids, "win-active") || !containsStringSlice(ids, "win-complete") { - t.Fatalf("unexpected window ids %v", ids) - } - - window := wrapper.GetWindow("win-active") - if window == nil || window.ResourceID != "res-1" || window.Status == "" { - t.Fatalf("unexpected window: %#v", window) - } -} - -func TestIncidentRecorderProviderWrapper_NilRecorder(t *testing.T) { - wrapper := &incidentRecorderProviderWrapper{} - if got := wrapper.GetWindowsForResource("res-1", 5); got != nil { - t.Fatalf("expected nil windows, got %#v", got) - } - if got := wrapper.GetWindow("win-1"); got != nil { - t.Fatalf("expected nil window, got %#v", got) - } -} - -func TestConvertIncidentWindowNil(t *testing.T) { - if got := convertIncidentWindow(nil); got != nil { - t.Fatalf("expected nil window, got %#v", got) - } -} - func TestEventCorrelatorProviderWrapper(t *testing.T) { now := time.Now().UTC() corr := proxmox.EventCorrelation{ diff --git a/internal/metrics/incident_archive.go b/internal/metrics/incident_archive.go new file mode 100644 index 000000000..5d8e0d243 --- /dev/null +++ b/internal/metrics/incident_archive.go @@ -0,0 +1,170 @@ +// Package metrics preserves access to legacy incident recording archives. +package metrics + +import ( + "encoding/json" + "errors" + "fmt" + "io" + "os" + "path/filepath" + "strings" + "time" +) + +// IncidentWindow represents a saved legacy recording window. Status and sample timestamps are historical +type IncidentWindow struct { + ID string `json:"id"` + ResourceID string `json:"resource_id"` + ResourceName string `json:"resource_name,omitempty"` + ResourceType string `json:"resource_type,omitempty"` + TriggerType string `json:"trigger_type"` // "alert", "anomaly", "focus", "manual" + TriggerID string `json:"trigger_id,omitempty"` + StartTime time.Time `json:"start_time"` + EndTime *time.Time `json:"end_time,omitempty"` + Status IncidentWindowStatus `json:"status"` + DataPoints []IncidentDataPoint `json:"data_points"` + Summary *IncidentSummary `json:"summary,omitempty"` +} + +// IncidentWindowStatus represents the status of an incident window +type IncidentWindowStatus string + +const ( + IncidentWindowStatusRecording IncidentWindowStatus = "recording" + IncidentWindowStatusComplete IncidentWindowStatus = "complete" + IncidentWindowStatusTruncated IncidentWindowStatus = "truncated" // Stopped due to limits + + maxIncidentWindowsFileSize = 16 << 20 // 16 MiB +) + +var errUnsafeIncidentArchivePath = errors.New("unsafe incident archive path") + +// IncidentDataPoint represents a single data point in an incident window +type IncidentDataPoint struct { + Timestamp time.Time `json:"timestamp"` + Metrics map[string]float64 `json:"metrics"` // cpu, memory, disk, etc. + Metadata map[string]interface{} `json:"metadata,omitempty"` +} + +// IncidentSummary provides computed statistics about an incident window +type IncidentSummary struct { + Duration time.Duration `json:"duration_ms"` + DataPoints int `json:"data_points"` + Peaks map[string]float64 `json:"peaks"` // Maximum values + Lows map[string]float64 `json:"lows"` // Minimum values + Averages map[string]float64 `json:"averages"` // Average values + Changes map[string]float64 `json:"changes"` // Change from start to end + Anomalies []string `json:"anomalies,omitempty"` // Detected anomalies +} + +// IncidentArchive reads saved recordings on explicit request. It never starts +// collectors or rewrites, expires, creates or changes permissions on archives. +type IncidentArchive struct{ filePath string } + +var ErrIncidentArchiveUnavailable = errors.New("legacy incident recording archive is unavailable") + +func NewIncidentArchive(dataDir string) *IncidentArchive { + if strings.TrimSpace(dataDir) == "" { + return &IncidentArchive{} + } + return &IncidentArchive{filePath: filepath.Join(dataDir, "incident_windows.json")} +} + +// GetWindow requires the exact resource and window identifiers in this org's +// archive. Historical names and aliases cannot authorize an archive lookup. +func (a *IncidentArchive) GetWindow(resourceID, windowID string) (*IncidentWindow, error) { + if a == nil || a.filePath == "" { + return nil, ErrIncidentArchiveUnavailable + } + if strings.TrimSpace(resourceID) == "" || strings.TrimSpace(windowID) == "" { + return nil, errors.New("resource_id and window_id are required") + } + data, err := readBoundedRegularFile(a.filePath, maxIncidentWindowsFileSize) + if errors.Is(err, os.ErrNotExist) { + return nil, ErrIncidentArchiveUnavailable + } + if err != nil { + return nil, fmt.Errorf("read legacy incident archive: %w", err) + } + var saved struct { + CompletedWindows json.RawMessage `json:"completed_windows"` + } + if err := json.Unmarshal(data, &saved); err != nil { + return nil, fmt.Errorf("decode legacy incident archive: %w", err) + } + if len(saved.CompletedWindows) == 0 { + return nil, errors.New("legacy incident archive has no completed_windows field") + } + var windows []*IncidentWindow + if err := json.Unmarshal(saved.CompletedWindows, &windows); err != nil { + return nil, fmt.Errorf("decode legacy incident windows: %w", err) + } + var match *IncidentWindow + for _, window := range windows { + if window == nil || window.ID != windowID || window.ResourceID != resourceID { + continue + } + if match != nil { + return nil, errors.New("legacy incident archive contains duplicate resource/window identifiers") + } + match = window + } + return match, nil +} + +func validateRegularFilePath(path string, info os.FileInfo) error { + if info.Mode()&os.ModeSymlink != 0 { + return fmt.Errorf("%w: refusing symlink path %q", errUnsafeIncidentArchivePath, path) + } + if !info.Mode().IsRegular() { + return fmt.Errorf("%w: non-regular path %q", errUnsafeIncidentArchivePath, path) + } + return nil +} + +func readBoundedRegularFile(path string, maxSize int64) ([]byte, error) { + initialInfo, err := os.Lstat(path) + if err != nil { + return nil, err + } + if err := validateRegularFilePath(path, initialInfo); err != nil { + return nil, err + } + if maxSize > 0 && initialInfo.Size() > maxSize { + return nil, fmt.Errorf("%w: file %q exceeds size limit (%d bytes)", errUnsafeIncidentArchivePath, path, initialInfo.Size()) + } + + file, err := os.Open(path) + if err != nil { + return nil, err + } + defer func() { + _ = file.Close() + }() + + openInfo, err := file.Stat() + if err != nil { + return nil, err + } + if err := validateRegularFilePath(path, openInfo); err != nil { + return nil, err + } + if !os.SameFile(initialInfo, openInfo) { + return nil, fmt.Errorf("%w: file %q changed during read", errUnsafeIncidentArchivePath, path) + } + + reader := io.Reader(file) + if maxSize > 0 { + reader = io.LimitReader(file, maxSize+1) + } + + data, err := io.ReadAll(reader) + if err != nil { + return nil, err + } + if maxSize > 0 && int64(len(data)) > maxSize { + return nil, fmt.Errorf("%w: file %q exceeded size limit while reading", errUnsafeIncidentArchivePath, path) + } + return data, nil +} diff --git a/internal/metrics/incident_archive_test.go b/internal/metrics/incident_archive_test.go new file mode 100644 index 000000000..c162eb787 --- /dev/null +++ b/internal/metrics/incident_archive_test.go @@ -0,0 +1,117 @@ +package metrics + +import ( + "encoding/json" + "os" + "path/filepath" + "testing" + "time" + + "github.com/stretchr/testify/require" +) + +func TestIncidentArchivePreservesHistoricalFileAndValues(t *testing.T) { + dir := t.TempDir() + path := filepath.Join(dir, "incident_windows.json") + // Older than the former retention window, with data that the old adapter dropped. + original := []byte(`{"completed_windows":[null,{"id":"old-window","resource_id":"docker:old-host/full-id","resource_type":"host","status":"recording","start_time":"2020-01-02T03:04:05Z","data_points":[{"timestamp":"2020-01-02T03:04:06Z","metrics":{"cpu":12.5},"metadata":{"source":"cached","nested":{"retained":true}}}],"summary":{"duration_ms":60000000000,"data_points":1,"anomalies":["stored observation"]}}]}`) + require.NoError(t, os.WriteFile(path, original, 0640)) + require.NoError(t, os.Chmod(path, 0640)) + oldTime := time.Date(2020, 1, 2, 3, 4, 5, 0, time.UTC) + require.NoError(t, os.Chtimes(path, oldTime, oldTime)) + before, err := os.Stat(path) + require.NoError(t, err) + archive := NewIncidentArchive(dir) + window, err := archive.GetWindow("docker:old-host/full-id", "old-window") + require.NoError(t, err) + require.NotNil(t, window) + require.Equal(t, IncidentWindowStatusRecording, window.Status) + require.Equal(t, "host", window.ResourceType) + require.Equal(t, oldTime, window.StartTime) + require.Equal(t, time.Minute, window.Summary.Duration) + require.Equal(t, []string{"stored observation"}, window.Summary.Anomalies) + require.Equal(t, map[string]interface{}{"retained": true}, window.DataPoints[0].Metadata["nested"]) + // Each read decodes independently, so a caller cannot alter later evidence. + window.DataPoints[0].Metrics["cpu"] = 99 + again, err := archive.GetWindow("docker:old-host/full-id", "old-window") + require.NoError(t, err) + require.Equal(t, 12.5, again.DataPoints[0].Metrics["cpu"]) + for _, key := range [][2]string{{"other-resource", "old-window"}, {"docker:old-host/full-id", "absent"}} { + got, err := archive.GetWindow(key[0], key[1]) + require.NoError(t, err) + require.Nil(t, got) + } + after, err := os.Stat(path) + require.NoError(t, err) + got, err := os.ReadFile(path) + require.NoError(t, err) + require.Equal(t, original, got) + require.Equal(t, before.Mode(), after.Mode()) + require.Equal(t, before.ModTime(), after.ModTime()) +} + +func TestIncidentArchiveReadsOnlyOnRequestAndReportsFailures(t *testing.T) { + dir := t.TempDir() + path := filepath.Join(dir, "incident_windows.json") + archive := NewIncidentArchive(dir) + _, err := os.Stat(path) + require.True(t, os.IsNotExist(err)) + _, err = archive.GetWindow("r", "w") + require.ErrorIs(t, err, ErrIncidentArchiveUnavailable) + // A repaired or newly restored file is visible on the next explicit read. + for _, raw := range []string{`{"completed_windows":[]}`, `{"completed_windows":null}`} { + require.NoError(t, os.WriteFile(path, []byte(raw), 0600)) + got, err := archive.GetWindow("r", "w") + require.NoError(t, err) + require.Nil(t, got) + } + for _, raw := range []string{`{`, `{}`, `null`, `{"completed_windows":{}}`, `{"completed_windows":[{"id":"w","resource_id":"r"},{"id":"w","resource_id":"r"}]}`} { + require.NoError(t, os.WriteFile(path, []byte(raw), 0600)) + got, err := archive.GetWindow("r", "w") + require.Error(t, err) + require.Nil(t, got) + } + require.NoError(t, os.Remove(path)) + require.NoError(t, os.Mkdir(path, 0700)) + _, err = archive.GetWindow("r", "w") + require.ErrorIs(t, err, errUnsafeIncidentArchivePath) + require.NoError(t, os.Remove(path)) + target := filepath.Join(dir, "other.json") + require.NoError(t, os.WriteFile(target, []byte(`{"completed_windows":[]}`), 0600)) + require.NoError(t, os.Symlink(target, path)) + _, err = archive.GetWindow("r", "w") + require.ErrorIs(t, err, errUnsafeIncidentArchivePath) + require.NoError(t, os.Remove(path)) + f, err := os.Create(path) + require.NoError(t, err) + require.NoError(t, f.Truncate(maxIncidentWindowsFileSize+1)) + require.NoError(t, f.Close()) + _, err = archive.GetWindow("r", "w") + require.ErrorIs(t, err, errUnsafeIncidentArchivePath) + for _, a := range []*IncidentArchive{nil, NewIncidentArchive("")} { + _, err = a.GetWindow("r", "w") + require.ErrorIs(t, err, ErrIncidentArchiveUnavailable) + } +} + +func TestIncidentArchiveRequiresExactResourceAndOrg(t *testing.T) { + aDir, bDir := t.TempDir(), t.TempDir() + for _, org := range []struct{ dir, name string }{{aDir, "tenant-a"}, {bDir, "tenant-b"}} { + raw, err := json.Marshal(map[string]interface{}{"completed_windows": []*IncidentWindow{{ID: "same-window", ResourceID: "same-resource", ResourceName: org.name}}}) + require.NoError(t, err) + require.NoError(t, os.WriteFile(filepath.Join(org.dir, "incident_windows.json"), raw, 0600)) + } + for _, org := range []struct{ dir, name string }{{aDir, "tenant-a"}, {bDir, "tenant-b"}} { + a := NewIncidentArchive(org.dir) + got, err := a.GetWindow("same-resource", "same-window") + require.NoError(t, err) + require.Equal(t, org.name, got.ResourceName) + for _, keys := range [][2]string{{"same-resource ", "same-window"}, {"same-resource", "same-window "}} { + got, err = a.GetWindow(keys[0], keys[1]) + require.NoError(t, err) + require.Nil(t, got) + } + _, err = a.GetWindow("", "same-window") + require.Error(t, err) + } +} diff --git a/internal/metrics/incident_recorder.go b/internal/metrics/incident_recorder.go deleted file mode 100644 index 24dc091e9..000000000 --- a/internal/metrics/incident_recorder.go +++ /dev/null @@ -1,948 +0,0 @@ -// Package metrics provides metrics collection and incident recording functionality. -package metrics - -import ( - "encoding/json" - "errors" - "fmt" - "io" - "os" - "path/filepath" - "strings" - "sync" - "sync/atomic" - "time" - - "github.com/rcourtman/pulse-go-rewrite/internal/unifiedresources" - "github.com/rs/zerolog/log" -) - -// IncidentWindow represents a high-frequency recording window during an incident -type IncidentWindow struct { - ID string `json:"id"` - ResourceID string `json:"resource_id"` - ResourceName string `json:"resource_name,omitempty"` - ResourceType string `json:"resource_type,omitempty"` - TriggerType string `json:"trigger_type"` // "alert", "anomaly", "focus", "manual" - TriggerID string `json:"trigger_id,omitempty"` - StartTime time.Time `json:"start_time"` - EndTime *time.Time `json:"end_time,omitempty"` - Status IncidentWindowStatus `json:"status"` - DataPoints []IncidentDataPoint `json:"data_points"` - Summary *IncidentSummary `json:"summary,omitempty"` -} - -// IncidentWindowStatus represents the status of an incident window -type IncidentWindowStatus string - -const ( - IncidentWindowStatusRecording IncidentWindowStatus = "recording" - IncidentWindowStatusComplete IncidentWindowStatus = "complete" - IncidentWindowStatusTruncated IncidentWindowStatus = "truncated" // Stopped due to limits - - incidentRecorderDirPerm = 0o700 - incidentRecorderFilePerm = 0o600 - maxIncidentWindowsFileSize = 16 << 20 // 16 MiB - maxWindowIDResourceSegment = 64 - unknownWindowResourceSegment = "unknown" -) - -var errUnsafeIncidentPersistencePath = errors.New("unsafe incident recorder persistence path") - -// IncidentDataPoint represents a single data point in an incident window -type IncidentDataPoint struct { - Timestamp time.Time `json:"timestamp"` - Metrics map[string]float64 `json:"metrics"` // cpu, memory, disk, etc. - Metadata map[string]interface{} `json:"metadata,omitempty"` -} - -// IncidentSummary provides computed statistics about an incident window -type IncidentSummary struct { - Duration time.Duration `json:"duration_ms"` - DataPoints int `json:"data_points"` - Peaks map[string]float64 `json:"peaks"` // Maximum values - Lows map[string]float64 `json:"lows"` // Minimum values - Averages map[string]float64 `json:"averages"` // Average values - Changes map[string]float64 `json:"changes"` // Change from start to end - Anomalies []string `json:"anomalies,omitempty"` // Detected anomalies -} - -// IncidentRecorderConfig configures the incident recorder -type IncidentRecorderConfig struct { - // Recording settings - SampleInterval time.Duration // How often to record data points (default: 5s) - PreIncidentWindow time.Duration // How much data to capture before incident (default: 5min) - PostIncidentWindow time.Duration // How much data to capture after incident (default: 10min) - MaxDataPointsPerWindow int // Maximum data points per window (default: 500) - - // Storage settings - DataDir string - MaxWindows int // Maximum number of windows to keep (default: 100) - RetentionDuration time.Duration // How long to keep windows (default: 24h) -} - -// DefaultIncidentRecorderConfig returns sensible defaults -func DefaultIncidentRecorderConfig() IncidentRecorderConfig { - return IncidentRecorderConfig{ - SampleInterval: 5 * time.Second, - PreIncidentWindow: 5 * time.Minute, - PostIncidentWindow: 10 * time.Minute, - MaxDataPointsPerWindow: 500, - MaxWindows: 100, - RetentionDuration: 24 * time.Hour, - } -} - -// MetricsProvider provides current metrics for a resource -type MetricsProvider interface { - GetCurrentMetrics(resourceID string) (map[string]float64, error) - GetMonitoredResourceIDs() []string // Returns all resource IDs being monitored -} - -// BatchMetricsProvider is an optional MetricsProvider extension. recordSample -// asks for metrics once per monitored resource every tick; a provider whose -// per-ID lookup scans all resources turns that into an O(n^2) tick, so when -// this is implemented the recorder fetches the whole batch in one pass. -type BatchMetricsProvider interface { - GetCurrentMetricsBatch() map[string]map[string]float64 -} - -// IncidentRecorder captures high-frequency metrics during incidents -type IncidentRecorder struct { - mu sync.RWMutex - - config IncidentRecorderConfig - provider MetricsProvider - - // Active recordings - activeWindows map[string]*IncidentWindow // keyed by window ID - - // Completed recordings (ring buffer) - completedWindows []*IncidentWindow - - // Background recording for pre-incident buffer - preIncidentBuffer map[string][]IncidentDataPoint // keyed by resource ID - - // Persistence - dataDir string - filePath string - saveIOMu sync.Mutex - - // Control - stopCh chan struct{} - loopDone chan struct{} - running bool - - // Async save coordination - saveMu sync.Mutex - saveCond *sync.Cond - saveInProgress bool - saveRequested bool -} - -// NewIncidentRecorder creates a new incident recorder -func NewIncidentRecorder(cfg IncidentRecorderConfig) *IncidentRecorder { - if cfg.SampleInterval <= 0 { - cfg.SampleInterval = 5 * time.Second - } - if cfg.PreIncidentWindow <= 0 { - cfg.PreIncidentWindow = 5 * time.Minute - } - if cfg.PostIncidentWindow <= 0 { - cfg.PostIncidentWindow = 10 * time.Minute - } - if cfg.MaxDataPointsPerWindow <= 0 { - cfg.MaxDataPointsPerWindow = 500 - } - if cfg.MaxWindows <= 0 { - cfg.MaxWindows = 100 - } - if cfg.RetentionDuration <= 0 { - cfg.RetentionDuration = 24 * time.Hour - } - if cfg.DataDir != "" { - trimmed := strings.TrimSpace(cfg.DataDir) - if trimmed == "" { - log.Warn().Msg("Ignoring incident recorder data dir: blank after trimming whitespace") - cfg.DataDir = "" - } else { - cfg.DataDir = filepath.Clean(trimmed) - } - } - - dataDir := strings.TrimSpace(cfg.DataDir) - if dataDir != "" { - dataDir = filepath.Clean(dataDir) - } - - recorder := &IncidentRecorder{ - config: cfg, - activeWindows: make(map[string]*IncidentWindow), - completedWindows: make([]*IncidentWindow, 0), - preIncidentBuffer: make(map[string][]IncidentDataPoint), - dataDir: dataDir, - stopCh: make(chan struct{}), - loopDone: make(chan struct{}), - } - recorder.saveCond = sync.NewCond(&recorder.saveMu) - - if dataDir != "" { - recorder.filePath = filepath.Join(dataDir, "incident_windows.json") - if err := recorder.loadFromDisk(); err != nil { - log.Warn(). - Str("file_path", recorder.filePath). - Err(err). - Msg("Failed to load incident windows from disk") - } - } - - return recorder -} - -// SetMetricsProvider sets the metrics provider for recording -func (r *IncidentRecorder) SetMetricsProvider(provider MetricsProvider) { - r.mu.Lock() - defer r.mu.Unlock() - r.provider = provider -} - -// Start begins background recording for pre-incident buffer -func (r *IncidentRecorder) Start() { - r.mu.Lock() - if r.running { - r.mu.Unlock() - return - } - stopCh := make(chan struct{}) - loopDone := make(chan struct{}) - r.running = true - r.stopCh = stopCh - r.loopDone = loopDone - r.mu.Unlock() - - go r.recordingLoop(stopCh, loopDone) - log.Info(). - Dur("sample_interval", r.config.SampleInterval). - Dur("pre_incident_window", r.config.PreIncidentWindow). - Dur("post_incident_window", r.config.PostIncidentWindow). - Int("max_data_points_per_window", r.config.MaxDataPointsPerWindow). - Msg("Incident recorder started") -} - -// Stop stops the incident recorder -func (r *IncidentRecorder) Stop() { - r.mu.Lock() - if !r.running { - r.mu.Unlock() - r.waitForPendingSaves() - return - } - r.running = false - close(r.stopCh) - loopDone := r.loopDone - r.mu.Unlock() - - if loopDone != nil { - <-loopDone - } - - // Flush any async save goroutine triggered before shutdown so TempDir cleanup - // and final persistence do not race a background rename/write. - r.waitForPendingSaves() - - // Save to disk - if err := r.saveToDisk(); err != nil { - log.Warn(). - Str("file_path", r.filePath). - Err(err). - Msg("Failed to save incident windows on stop") - } - log.Info().Msg("incident recorder stopped") -} - -// recordingLoop runs in the background to maintain pre-incident buffers and active windows -func (r *IncidentRecorder) recordingLoop(stopCh <-chan struct{}, done chan<- struct{}) { - defer close(done) - - ticker := time.NewTicker(r.config.SampleInterval) - defer ticker.Stop() - - for { - select { - case <-stopCh: - return - case <-ticker.C: - r.recordSample() - } - } -} - -// recordSample captures a data point for all active windows and buffers -func (r *IncidentRecorder) recordSample() { - r.mu.Lock() - if r.provider == nil { - r.mu.Unlock() - return - } - - now := time.Now() - shouldSave := false - - // One pass over the provider when it supports batching; the per-ID - // fallback preserves behavior for providers that do not. A missing batch - // entry maps to the empty metrics the per-ID method returns for unknown - // IDs. - var metricsBatch map[string]map[string]float64 - if batchProvider, ok := r.provider.(BatchMetricsProvider); ok { - metricsBatch = batchProvider.GetCurrentMetricsBatch() - } - currentMetrics := func(resourceID string) (map[string]float64, error) { - if metricsBatch != nil { - if metrics, ok := metricsBatch[resourceID]; ok { - return metrics, nil - } - } - // A batch that lacks the ID must not change per-provider semantics: - // fall through so providers that error or synthesize for unknown IDs - // keep doing exactly that. Misses are the rare path, so the batch - // still absorbs the per-tick fan-out. - return r.provider.GetCurrentMetrics(resourceID) - } - - // Record for active windows - for _, window := range r.activeWindows { - if window.Status != IncidentWindowStatusRecording { - continue - } - - // Check if we've exceeded the post-incident window - if window.EndTime != nil && now.After(*window.EndTime) { - r.completeWindowLocked(window) - shouldSave = true - continue - } - - // Check if we've exceeded max data points - if len(window.DataPoints) >= r.config.MaxDataPointsPerWindow { - window.Status = IncidentWindowStatusTruncated - log.Warn(). - Str("window_id", window.ID). - Str("resource_id", window.ResourceID). - Int("max_data_points_per_window", r.config.MaxDataPointsPerWindow). - Msg("Truncating incident window after reaching max data points") - r.completeWindowLocked(window) - continue - } - - // Get metrics - metrics, err := currentMetrics(window.ResourceID) - if err != nil { - log.Debug(). - Str("window_id", window.ID). - Str("resource_id", window.ResourceID). - Err(err). - Msg("failed to get metrics for incident window") - continue - } - - window.DataPoints = append(window.DataPoints, IncidentDataPoint{ - Timestamp: now, - Metrics: copyMetrics(metrics), - }) - } - - // Continuously buffer ALL monitored resources for pre-incident data - // This ensures we have history when an alert fires on any resource - monitoredResources := r.provider.GetMonitoredResourceIDs() - bufferCutoff := now.Add(-r.config.PreIncidentWindow) - - for _, resourceID := range monitoredResources { - metrics, err := currentMetrics(resourceID) - if err != nil { - log.Debug(). - Str("resource_id", resourceID). - Err(err). - Msg("Failed to get metrics for pre-incident buffer") - continue - } - - // Add to pre-incident buffer - buffer := r.preIncidentBuffer[resourceID] - buffer = append(buffer, IncidentDataPoint{ - Timestamp: now, - Metrics: copyMetrics(metrics), - }) - - // Keep only last PreIncidentWindow duration - kept := make([]IncidentDataPoint, 0, len(buffer)) - for _, dp := range buffer { - if dp.Timestamp.After(bufferCutoff) { - kept = append(kept, dp) - } - } - r.preIncidentBuffer[resourceID] = kept - } - - // Clean up buffers for resources no longer monitored - monitoredSet := make(map[string]bool, len(monitoredResources)) - for _, resourceID := range monitoredResources { - monitoredSet[resourceID] = true - } - for resourceID := range r.preIncidentBuffer { - if !monitoredSet[resourceID] { - delete(r.preIncidentBuffer, resourceID) - } - } - r.mu.Unlock() - - if shouldSave { - if err := r.saveToDisk(); err != nil { - log.Warn().Err(err).Msg("Failed to save incident windows") - } - } -} - -// StartRecording begins recording an incident window -func (r *IncidentRecorder) StartRecording(resourceID, resourceName, resourceType, triggerType, triggerID string) string { - r.mu.Lock() - defer r.mu.Unlock() - - // Check if we already have an active window for this resource - for _, window := range r.activeWindows { - if window.ResourceID == resourceID && window.Status == IncidentWindowStatusRecording { - // Extend existing window - endTime := time.Now().Add(r.config.PostIncidentWindow) - window.EndTime = &endTime - return window.ID - } - } - - // Create new window - windowID := generateWindowID(resourceID) - now := time.Now() - endTime := now.Add(r.config.PostIncidentWindow) - normalizedResourceType := normalizeIncidentResourceType(resourceType) - - window := &IncidentWindow{ - ID: windowID, - ResourceID: resourceID, - ResourceName: resourceName, - ResourceType: normalizedResourceType, - TriggerType: triggerType, - TriggerID: triggerID, - StartTime: now.Add(-r.config.PreIncidentWindow), // Include pre-incident data - EndTime: &endTime, - Status: IncidentWindowStatusRecording, - DataPoints: make([]IncidentDataPoint, 0), - } - - // Copy pre-incident buffer if available - if preBuffer, ok := r.preIncidentBuffer[resourceID]; ok { - window.DataPoints = append(window.DataPoints, copyDataPoints(preBuffer)...) - } - - r.activeWindows[windowID] = window - - log.Info(). - Str("window_id", windowID). - Str("resource_id", resourceID). - Str("trigger_type", triggerType). - Msg("started incident recording") - - return windowID -} - -func normalizeIncidentResourceType(resourceType string) string { - normalized := strings.ToLower(strings.TrimSpace(resourceType)) - if canonical, ok := unifiedresources.CanonicalizeLegacyResourceTypeAlias(normalized); ok { - return canonical - } - return normalized -} - -// StopRecording stops recording for a specific window -func (r *IncidentRecorder) StopRecording(windowID string) { - r.mu.Lock() - shouldSave := false - if window, ok := r.activeWindows[windowID]; ok { - r.completeWindowLocked(window) - shouldSave = true - } - r.mu.Unlock() - - if shouldSave { - if err := r.saveToDisk(); err != nil { - log.Warn().Err(err).Msg("Failed to save incident windows") - } - } -} - -// completeWindowLocked finalizes a recording window. -// Caller must hold r.mu. -func (r *IncidentRecorder) completeWindowLocked(window *IncidentWindow) { - if window.Status != IncidentWindowStatusRecording && window.Status != IncidentWindowStatusTruncated { - return - } - - now := time.Now() - if window.Status == IncidentWindowStatusRecording { - window.Status = IncidentWindowStatusComplete - } - window.EndTime = &now - - // Compute summary - window.Summary = r.computeSummary(window) - - // Move to completed - r.completedWindows = append(r.completedWindows, window) - delete(r.activeWindows, window.ID) - - // Trim completed windows - r.trimCompletedWindows() - - log.Info(). - Str("window_id", window.ID). - Str("resource_id", window.ResourceID). - Str("status", string(window.Status)). - Int("data_points", len(window.DataPoints)). - Msg("completed incident recording") - - // Save asynchronously. - r.requestAsyncSave() -} - -// computeSummary computes statistics for a window -func (r *IncidentRecorder) computeSummary(window *IncidentWindow) *IncidentSummary { - if len(window.DataPoints) == 0 { - return nil - } - - summary := &IncidentSummary{ - DataPoints: len(window.DataPoints), - Peaks: make(map[string]float64), - Lows: make(map[string]float64), - Averages: make(map[string]float64), - Changes: make(map[string]float64), - } - - // Calculate duration - if len(window.DataPoints) > 1 { - first := window.DataPoints[0].Timestamp - last := window.DataPoints[len(window.DataPoints)-1].Timestamp - summary.Duration = last.Sub(first) - } - - // Track sums for averages - sums := make(map[string]float64) - counts := make(map[string]int) - - // First and last values for change calculation - firstValues := make(map[string]float64) - lastValues := make(map[string]float64) - - for i, dp := range window.DataPoints { - for metric, value := range dp.Metrics { - // Track first value - if i == 0 { - firstValues[metric] = value - summary.Peaks[metric] = value - summary.Lows[metric] = value - } - - // Track last value - lastValues[metric] = value - - // Track peaks and lows - if value > summary.Peaks[metric] { - summary.Peaks[metric] = value - } - if value < summary.Lows[metric] { - summary.Lows[metric] = value - } - - // Track sums for average - sums[metric] += value - counts[metric]++ - } - } - - // Calculate averages and changes - for metric, sum := range sums { - if counts[metric] > 0 { - summary.Averages[metric] = sum / float64(counts[metric]) - } - if first, ok := firstValues[metric]; ok { - if last, ok := lastValues[metric]; ok { - summary.Changes[metric] = last - first - } - } - } - - return summary -} - -// trimCompletedWindows removes old windows -func (r *IncidentRecorder) trimCompletedWindows() { - // Remove by retention duration - cutoff := time.Now().Add(-r.config.RetentionDuration) - kept := make([]*IncidentWindow, 0, len(r.completedWindows)) - for _, w := range r.completedWindows { - if w.EndTime != nil && w.EndTime.After(cutoff) { - kept = append(kept, w) - } - } - r.completedWindows = kept - - // Remove by max windows - if len(r.completedWindows) > r.config.MaxWindows { - r.completedWindows = r.completedWindows[len(r.completedWindows)-r.config.MaxWindows:] - } -} - -// GetWindow returns a specific incident window -func (r *IncidentRecorder) GetWindow(windowID string) *IncidentWindow { - r.mu.RLock() - defer r.mu.RUnlock() - - // Check active windows - if window, ok := r.activeWindows[windowID]; ok { - return copyWindow(window) - } - - // Check completed windows - for _, window := range r.completedWindows { - if window.ID == windowID { - return copyWindow(window) - } - } - - return nil -} - -// GetWindowsForResource returns all incident windows for a resource -func (r *IncidentRecorder) GetWindowsForResource(resourceID string, limit int) []*IncidentWindow { - r.mu.RLock() - defer r.mu.RUnlock() - - var result []*IncidentWindow - - // Check active windows - for _, window := range r.activeWindows { - if window.ResourceID == resourceID { - result = append(result, copyWindow(window)) - } - } - - // Check completed windows (in reverse order for most recent first) - for i := len(r.completedWindows) - 1; i >= 0; i-- { - if r.completedWindows[i].ResourceID == resourceID { - result = append(result, copyWindow(r.completedWindows[i])) - if limit > 0 && len(result) >= limit { - break - } - } - } - - return result -} - -// saveToDisk persists completed windows -func (r *IncidentRecorder) saveToDisk() error { - if r.filePath == "" { - return nil - } - r.saveIOMu.Lock() - defer r.saveIOMu.Unlock() - - data := struct { - CompletedWindows []*IncidentWindow `json:"completed_windows"` - }{ - CompletedWindows: r.snapshotCompletedWindows(), - } - - jsonData, err := json.MarshalIndent(data, "", " ") - if err != nil { - return fmt.Errorf("incident recorder save: marshal completed windows: %w", err) - } - - if err := ensureOwnerOnlyDir(r.dataDir); err != nil { - return err - } - - if info, err := os.Lstat(r.filePath); err == nil { - if err := validateRegularFilePath(r.filePath, info); err != nil { - return err - } - } else if !errors.Is(err, os.ErrNotExist) { - return err - } - - tmpFile, err := os.CreateTemp(r.dataDir, filepath.Base(r.filePath)+".*.tmp") - if err != nil { - return err - } - tmpPath := tmpFile.Name() - cleanup := true - defer func() { - if cleanup { - _ = os.Remove(tmpPath) - } - }() - - if err := tmpFile.Chmod(incidentRecorderFilePerm); err != nil { - _ = tmpFile.Close() - return err - } - if _, err := tmpFile.Write(jsonData); err != nil { - _ = tmpFile.Close() - return err - } - if err := tmpFile.Close(); err != nil { - return err - } - if err := os.Rename(tmpPath, r.filePath); err != nil { - return err - } - cleanup = false - return os.Chmod(r.filePath, incidentRecorderFilePerm) -} - -func (r *IncidentRecorder) requestAsyncSave() { - if r.filePath == "" { - return - } - - r.saveMu.Lock() - if r.saveInProgress { - r.saveRequested = true - r.saveMu.Unlock() - return - } - r.saveInProgress = true - r.saveMu.Unlock() - - go r.runAsyncSaves() -} - -func (r *IncidentRecorder) runAsyncSaves() { - for { - if err := r.saveToDisk(); err != nil { - log.Warn(). - Str("file_path", r.filePath). - Err(err). - Msg("Failed to save incident windows") - } - - r.saveMu.Lock() - if !r.saveRequested { - r.saveInProgress = false - r.saveCond.Broadcast() - r.saveMu.Unlock() - return - } - r.saveRequested = false - r.saveMu.Unlock() - } -} - -func (r *IncidentRecorder) snapshotCompletedWindows() []*IncidentWindow { - r.mu.RLock() - defer r.mu.RUnlock() - - snapshot := make([]*IncidentWindow, len(r.completedWindows)) - for i, window := range r.completedWindows { - snapshot[i] = copyWindow(window) - } - return snapshot -} - -func (r *IncidentRecorder) waitForPendingSaves() { - if r.filePath == "" { - return - } - - r.saveMu.Lock() - for r.saveInProgress || r.saveRequested { - r.saveCond.Wait() - } - r.saveMu.Unlock() - - r.saveIOMu.Lock() - r.saveIOMu.Unlock() -} - -// loadFromDisk loads completed windows -func (r *IncidentRecorder) loadFromDisk() error { - if r.filePath == "" { - return nil - } - - jsonData, err := readBoundedRegularFile(r.filePath, maxIncidentWindowsFileSize) - if err != nil { - if errors.Is(err, os.ErrNotExist) { - return nil - } - return fmt.Errorf("incident recorder load: read file %q: %w", r.filePath, err) - } - - var data struct { - CompletedWindows []*IncidentWindow `json:"completed_windows"` - } - - if err := json.Unmarshal(jsonData, &data); err != nil { - return fmt.Errorf("incident recorder load: parse file %q: %w", r.filePath, err) - } - - r.completedWindows = make([]*IncidentWindow, 0, len(data.CompletedWindows)) - for _, window := range data.CompletedWindows { - if window == nil { - continue - } - r.completedWindows = append(r.completedWindows, window) - } - r.trimCompletedWindows() - - return os.Chmod(r.filePath, incidentRecorderFilePerm) -} - -// Helper functions - -func ensureOwnerOnlyDir(dir string) error { - if err := os.MkdirAll(dir, incidentRecorderDirPerm); err != nil { - return err - } - return os.Chmod(dir, incidentRecorderDirPerm) -} - -func validateRegularFilePath(path string, info os.FileInfo) error { - if info.Mode()&os.ModeSymlink != 0 { - return fmt.Errorf("%w: refusing symlink path %q", errUnsafeIncidentPersistencePath, path) - } - if !info.Mode().IsRegular() { - return fmt.Errorf("%w: non-regular path %q", errUnsafeIncidentPersistencePath, path) - } - return nil -} - -func readBoundedRegularFile(path string, maxSize int64) ([]byte, error) { - initialInfo, err := os.Lstat(path) - if err != nil { - return nil, err - } - if err := validateRegularFilePath(path, initialInfo); err != nil { - return nil, err - } - if maxSize > 0 && initialInfo.Size() > maxSize { - return nil, fmt.Errorf("%w: file %q exceeds size limit (%d bytes)", errUnsafeIncidentPersistencePath, path, initialInfo.Size()) - } - - file, err := os.Open(path) - if err != nil { - return nil, err - } - defer func() { - _ = file.Close() - }() - - openInfo, err := file.Stat() - if err != nil { - return nil, err - } - if err := validateRegularFilePath(path, openInfo); err != nil { - return nil, err - } - if !os.SameFile(initialInfo, openInfo) { - return nil, fmt.Errorf("%w: file %q changed during read", errUnsafeIncidentPersistencePath, path) - } - - reader := io.Reader(file) - if maxSize > 0 { - reader = io.LimitReader(file, maxSize+1) - } - - data, err := io.ReadAll(reader) - if err != nil { - return nil, err - } - if maxSize > 0 && int64(len(data)) > maxSize { - return nil, fmt.Errorf("%w: file %q exceeded size limit while reading", errUnsafeIncidentPersistencePath, path) - } - return data, nil -} - -func copyWindow(w *IncidentWindow) *IncidentWindow { - if w == nil { - return nil - } - windowCopy := *w - if w.EndTime != nil { - t := *w.EndTime - windowCopy.EndTime = &t - } - windowCopy.DataPoints = copyDataPoints(w.DataPoints) - if w.Summary != nil { - s := *w.Summary - s.Peaks = copyMetrics(w.Summary.Peaks) - s.Lows = copyMetrics(w.Summary.Lows) - s.Averages = copyMetrics(w.Summary.Averages) - s.Changes = copyMetrics(w.Summary.Changes) - if w.Summary.Anomalies != nil { - s.Anomalies = append([]string(nil), w.Summary.Anomalies...) - } - windowCopy.Summary = &s - } - return &windowCopy -} - -var windowCounter int64 - -func generateWindowID(resourceID string) string { - counter := atomic.AddInt64(&windowCounter, 1) - return "iw-" + resourceID + "-" + time.Now().Format("20060102150405") + "-" + intToString(int(counter)) -} - -func copyDataPoints(points []IncidentDataPoint) []IncidentDataPoint { - copied := make([]IncidentDataPoint, len(points)) - for i, dp := range points { - copied[i] = dp - copied[i].Metrics = copyMetrics(dp.Metrics) - if dp.Metadata != nil { - copied[i].Metadata = make(map[string]interface{}, len(dp.Metadata)) - for k, v := range dp.Metadata { - copied[i].Metadata[k] = v - } - } - } - return copied -} - -func copyMetrics(metrics map[string]float64) map[string]float64 { - if metrics == nil { - return nil - } - - copied := make(map[string]float64, len(metrics)) - for k, v := range metrics { - copied[k] = v - } - return copied -} - -func intToString(n int) string { - if n == 0 { - return "0" - } - negative := n < 0 - if negative { - n = -n - } - var result string - for n > 0 { - result = string(rune('0'+n%10)) + result - n /= 10 - } - if negative { - result = "-" + result - } - return result -} diff --git a/internal/metrics/incident_recorder_additional_test.go b/internal/metrics/incident_recorder_additional_test.go deleted file mode 100644 index 9c26c98a6..000000000 --- a/internal/metrics/incident_recorder_additional_test.go +++ /dev/null @@ -1,133 +0,0 @@ -package metrics - -import ( - "sync/atomic" - "testing" - "time" -) - -type countingProvider struct { - ids []string - metrics map[string]map[string]float64 - calls int32 -} - -func (c *countingProvider) GetCurrentMetrics(resourceID string) (map[string]float64, error) { - atomic.AddInt32(&c.calls, 1) - metrics, ok := c.metrics[resourceID] - if !ok { - return nil, errNoMetrics(resourceID) - } - copied := make(map[string]float64, len(metrics)) - for k, v := range metrics { - copied[k] = v - } - return copied, nil -} - -func (c *countingProvider) GetMonitoredResourceIDs() []string { - return append([]string{}, c.ids...) -} - -func waitForCalls(t *testing.T, provider *countingProvider, timeout time.Duration) { - t.Helper() - deadline := time.Now().Add(timeout) - for time.Now().Before(deadline) { - if atomic.LoadInt32(&provider.calls) > 0 { - return - } - time.Sleep(5 * time.Millisecond) - } - t.Fatal("expected provider to be called") -} - -func TestDefaultIncidentRecorderConfig(t *testing.T) { - cfg := DefaultIncidentRecorderConfig() - if cfg.SampleInterval == 0 || cfg.PreIncidentWindow == 0 || cfg.PostIncidentWindow == 0 { - t.Fatalf("default config should be non-zero, got %+v", cfg) - } - if cfg.MaxDataPointsPerWindow == 0 || cfg.MaxWindows == 0 || cfg.RetentionDuration == 0 { - t.Fatalf("default config should be non-zero, got %+v", cfg) - } -} - -func TestIncidentRecorderStartStop(t *testing.T) { - recorder := NewIncidentRecorder(IncidentRecorderConfig{ - SampleInterval: 5 * time.Millisecond, - PreIncidentWindow: 10 * time.Millisecond, - PostIncidentWindow: 10 * time.Millisecond, - MaxDataPointsPerWindow: 5, - }) - - provider := &countingProvider{ - ids: []string{"res-1"}, - metrics: map[string]map[string]float64{ - "res-1": {"cpu": 1}, - }, - } - recorder.SetMetricsProvider(provider) - - recorder.Start() - waitForCalls(t, provider, 200*time.Millisecond) - recorder.Stop() - - if recorder.running { - t.Fatalf("expected recorder to be stopped") - } -} - -func TestGetWindowsForResource(t *testing.T) { - recorder := NewIncidentRecorder(IncidentRecorderConfig{}) - - active := &IncidentWindow{ID: "active-1", ResourceID: "res-1"} - recorder.activeWindows["active-1"] = active - - got := recorder.GetWindowsForResource("res-1", 0) - if len(got) != 1 || got[0].ID != "active-1" { - t.Fatalf("expected active window, got %+v", got) - } - - recorder.activeWindows = map[string]*IncidentWindow{} - recorder.completedWindows = []*IncidentWindow{ - {ID: "old", ResourceID: "res-1"}, - {ID: "new", ResourceID: "res-1"}, - } - - limited := recorder.GetWindowsForResource("res-1", 1) - if len(limited) != 1 || limited[0].ID != "new" { - t.Fatalf("expected most recent completed window, got %+v", limited) - } -} - -func TestRecordSampleSkipsPreIncidentBufferOnMetricsError(t *testing.T) { - recorder := NewIncidentRecorder(IncidentRecorderConfig{ - PreIncidentWindow: time.Minute, - PostIncidentWindow: time.Minute, - MaxDataPointsPerWindow: 10, - }) - - provider := &stubMetricsProvider{ - metricsByID: map[string]map[string]float64{ - "res-ok": {"cpu": 1}, - }, - ids: []string{"res-ok", "res-missing"}, - } - recorder.SetMetricsProvider(provider) - - windowID := recorder.StartRecording("res-ok", "db", "agent", "alert", "alert-1") - recorder.recordSample() - - window := recorder.activeWindows[windowID] - if window == nil { - t.Fatalf("expected active window %s", windowID) - } - if len(window.DataPoints) != 1 { - t.Fatalf("expected active window sample to be captured, got %d", len(window.DataPoints)) - } - if len(recorder.preIncidentBuffer["res-ok"]) == 0 { - t.Fatalf("expected pre-incident buffer for res-ok") - } - if _, ok := recorder.preIncidentBuffer["res-missing"]; ok { - t.Fatalf("expected no pre-incident buffer for res-missing when metrics collection fails") - } -} diff --git a/internal/metrics/incident_recorder_batch_test.go b/internal/metrics/incident_recorder_batch_test.go deleted file mode 100644 index 741b7ebb9..000000000 --- a/internal/metrics/incident_recorder_batch_test.go +++ /dev/null @@ -1,84 +0,0 @@ -package metrics - -import ( - "errors" - "testing" - "time" -) - -// gappyBatchProvider implements both MetricsProvider and BatchMetricsProvider -// but omits one monitored ID from the batch. The recorder must fall through -// to the per-ID method for that ID so provider semantics are preserved. -type gappyBatchProvider struct { - batchCalls int - perIDCalls map[string]int - perIDResults map[string]map[string]float64 - perIDErr map[string]error -} - -func (p *gappyBatchProvider) GetMonitoredResourceIDs() []string { - return []string{"in-batch", "not-in-batch", "erroring"} -} - -func (p *gappyBatchProvider) GetCurrentMetricsBatch() map[string]map[string]float64 { - p.batchCalls++ - return map[string]map[string]float64{ - "in-batch": {"cpu": 42}, - } -} - -func (p *gappyBatchProvider) GetCurrentMetrics(resourceID string) (map[string]float64, error) { - if p.perIDCalls == nil { - p.perIDCalls = map[string]int{} - } - p.perIDCalls[resourceID]++ - if err := p.perIDErr[resourceID]; err != nil { - return nil, err - } - return p.perIDResults[resourceID], nil -} - -func TestRecordSampleBatchMissFallsThroughToPerID(t *testing.T) { - recorder := NewIncidentRecorder(IncidentRecorderConfig{ - SampleInterval: time.Hour, // ticks driven manually - PreIncidentWindow: time.Minute, - PostIncidentWindow: time.Minute, - MaxDataPointsPerWindow: 5, - }) - provider := &gappyBatchProvider{ - perIDResults: map[string]map[string]float64{ - "not-in-batch": {"cpu": 7}, - }, - perIDErr: map[string]error{ - "erroring": errors.New("unknown resource"), - }, - } - recorder.SetMetricsProvider(provider) - - recorder.recordSample() - - if provider.batchCalls != 1 { - t.Fatalf("batch calls = %d, want 1", provider.batchCalls) - } - if provider.perIDCalls["in-batch"] != 0 { - t.Fatalf("batch hit %q still took the per-ID path", "in-batch") - } - if provider.perIDCalls["not-in-batch"] != 1 || provider.perIDCalls["erroring"] != 1 { - t.Fatalf("batch misses did not fall through per-ID: %+v", provider.perIDCalls) - } - - recorder.mu.RLock() - defer recorder.mu.RUnlock() - if got := len(recorder.preIncidentBuffer["not-in-batch"]); got != 1 { - t.Fatalf("fallthrough metrics not buffered: %d points", got) - } - if got := recorder.preIncidentBuffer["not-in-batch"][0].Metrics["cpu"]; got != 7 { - t.Fatalf("fallthrough buffered cpu = %v, want 7 (per-ID value)", got) - } - if got := len(recorder.preIncidentBuffer["erroring"]); got != 0 { - t.Fatalf("erroring ID gained %d buffered points, want 0 (per-ID error must skip)", got) - } - if got := len(recorder.preIncidentBuffer["in-batch"]); got != 1 { - t.Fatalf("batch-served ID not buffered: %d points", got) - } -} diff --git a/internal/metrics/incident_recorder_concurrency_test.go b/internal/metrics/incident_recorder_concurrency_test.go deleted file mode 100644 index d56ebec6e..000000000 --- a/internal/metrics/incident_recorder_concurrency_test.go +++ /dev/null @@ -1,86 +0,0 @@ -package metrics - -import ( - "sync" - "testing" - "time" -) - -func TestGenerateWindowIDConcurrentUnique(t *testing.T) { - t.Parallel() - - const ( - workers = 16 - idsPerWork = 128 - ) - - ids := make(chan string, workers*idsPerWork) - start := make(chan struct{}) - - var wg sync.WaitGroup - for i := 0; i < workers; i++ { - wg.Add(1) - go func() { - defer wg.Done() - <-start - for j := 0; j < idsPerWork; j++ { - ids <- generateWindowID("res-1") - } - }() - } - - close(start) - wg.Wait() - close(ids) - - seen := make(map[string]struct{}, workers*idsPerWork) - for id := range ids { - if _, exists := seen[id]; exists { - t.Fatalf("duplicate window ID generated: %s", id) - } - seen[id] = struct{}{} - } -} - -func TestIncidentRecorderConcurrentStartStopAndFlush(t *testing.T) { - recorder := NewIncidentRecorder(IncidentRecorderConfig{ - SampleInterval: time.Millisecond, - PreIncidentWindow: 10 * time.Millisecond, - PostIncidentWindow: 10 * time.Millisecond, - MaxDataPointsPerWindow: 10, - DataDir: t.TempDir(), - }) - provider := &stubMetricsProvider{ - metricsByID: map[string]map[string]float64{ - "res-1": {"cpu": 1}, - }, - ids: []string{"res-1"}, - } - recorder.SetMetricsProvider(provider) - - const goroutines = 8 - const iterations = 15 - start := make(chan struct{}) - - var wg sync.WaitGroup - for i := 0; i < goroutines; i++ { - wg.Add(1) - go func() { - defer wg.Done() - <-start - for j := 0; j < iterations; j++ { - recorder.Start() - windowID := recorder.StartRecording("res-1", "db", "agent", "alert", "a-1") - recorder.recordSample() - recorder.StopRecording(windowID) - recorder.Stop() - } - }() - } - - close(start) - wg.Wait() - - // Final stop should remain idempotent after concurrent shutdowns. - recorder.Stop() -} diff --git a/internal/metrics/incident_recorder_coverage_test.go b/internal/metrics/incident_recorder_coverage_test.go deleted file mode 100644 index f365531ae..000000000 --- a/internal/metrics/incident_recorder_coverage_test.go +++ /dev/null @@ -1,431 +0,0 @@ -package metrics - -import ( - "bytes" - "math" - "os" - "path/filepath" - "sync/atomic" - "testing" - "time" - - "github.com/rs/zerolog" - "github.com/rs/zerolog/log" -) - -func TestNewIncidentRecorderLoadFromDiskInvalidJSON(t *testing.T) { - dir := t.TempDir() - path := filepath.Join(dir, "incident_windows.json") - if err := os.WriteFile(path, []byte("{invalid"), 0600); err != nil { - t.Fatalf("write invalid json: %v", err) - } - - recorder := NewIncidentRecorder(IncidentRecorderConfig{DataDir: dir}) - if recorder == nil { - t.Fatal("expected recorder") - } - if len(recorder.completedWindows) != 0 { - t.Fatalf("expected no completed windows on invalid json, got %d", len(recorder.completedWindows)) - } -} - -func TestIncidentRecorderStartStopIdempotentGuards(t *testing.T) { - recorder := NewIncidentRecorder(IncidentRecorderConfig{ - SampleInterval: 10 * time.Millisecond, - }) - - recorder.Start() - firstStopCh := recorder.stopCh - - recorder.Start() - if recorder.stopCh != firstStopCh { - t.Fatal("expected second Start call to be a no-op while running") - } - - recorder.Stop() - if recorder.running { - t.Fatal("expected recorder to be stopped") - } - - // Should be a no-op and should not panic. - recorder.Stop() -} - -func TestIncidentRecorderStopWaitsForPendingSavesWhenAlreadyStopped(t *testing.T) { - recorder := NewIncidentRecorder(IncidentRecorderConfig{ - DataDir: t.TempDir(), - }) - - recorder.saveMu.Lock() - recorder.saveInProgress = true - recorder.saveMu.Unlock() - - done := make(chan struct{}) - go func() { - recorder.Stop() - close(done) - }() - - select { - case <-done: - t.Fatal("expected Stop to wait for pending saves even when recorder is already stopped") - case <-time.After(20 * time.Millisecond): - } - - recorder.saveMu.Lock() - recorder.saveInProgress = false - recorder.saveCond.Broadcast() - recorder.saveMu.Unlock() - - select { - case <-done: - case <-time.After(time.Second): - t.Fatal("timed out waiting for Stop to return after pending saves were cleared") - } -} - -func TestRecordSampleNoProviderNoop(t *testing.T) { - recorder := NewIncidentRecorder(IncidentRecorderConfig{}) - recorder.recordSample() -} - -func TestRecordSampleCoversActiveWindowBranches(t *testing.T) { - recorder := NewIncidentRecorder(IncidentRecorderConfig{ - PreIncidentWindow: time.Second, - PostIncidentWindow: time.Second, - MaxDataPointsPerWindow: 1, - MaxWindows: 10, - RetentionDuration: time.Hour, - }) - provider := &stubMetricsProvider{ - metricsByID: map[string]map[string]float64{ - "res-ok": {"cpu": 1}, - "res-active-ok": {"cpu": 2}, - }, - ids: []string{"res-ok", "res-buffer-missing"}, - } - recorder.SetMetricsProvider(provider) - - now := time.Now() - past := now.Add(-time.Millisecond) - recorder.activeWindows["skip-non-recording"] = &IncidentWindow{ - ID: "skip-non-recording", - ResourceID: "res-active-ok", - Status: IncidentWindowStatusComplete, - EndTime: &past, - } - recorder.activeWindows["expire-now"] = &IncidentWindow{ - ID: "expire-now", - ResourceID: "res-active-ok", - Status: IncidentWindowStatusRecording, - EndTime: &past, - } - recorder.activeWindows["truncate-now"] = &IncidentWindow{ - ID: "truncate-now", - ResourceID: "res-active-ok", - Status: IncidentWindowStatusRecording, - DataPoints: []IncidentDataPoint{ - {Timestamp: now, Metrics: map[string]float64{"cpu": 7}}, - }, - } - recorder.activeWindows["metrics-error"] = &IncidentWindow{ - ID: "metrics-error", - ResourceID: "res-active-missing", - Status: IncidentWindowStatusRecording, - } - - recorder.preIncidentBuffer["res-ok"] = []IncidentDataPoint{ - {Timestamp: now.Add(-2 * time.Second), Metrics: map[string]float64{"cpu": 0.5}}, - } - recorder.preIncidentBuffer["stale-resource"] = []IncidentDataPoint{ - {Timestamp: now, Metrics: map[string]float64{"cpu": 9}}, - } - - recorder.recordSample() - - if _, ok := recorder.activeWindows["expire-now"]; ok { - t.Fatal("expected expired window to complete") - } - if _, ok := recorder.activeWindows["truncate-now"]; ok { - t.Fatal("expected truncated window to complete") - } - - metricsErrWindow, ok := recorder.activeWindows["metrics-error"] - if !ok { - t.Fatal("expected metrics-error window to remain active") - } - if len(metricsErrWindow.DataPoints) != 0 { - t.Fatalf("expected metrics-error window to skip append, got %d points", len(metricsErrWindow.DataPoints)) - } - - if _, ok := recorder.preIncidentBuffer["stale-resource"]; ok { - t.Fatal("expected stale pre-incident buffer to be removed") - } - if got := len(recorder.preIncidentBuffer["res-ok"]); got != 1 { - t.Fatalf("expected pre-incident buffer trim to keep 1 point, got %d", got) - } - - foundTruncated := false - foundCompleted := false - for _, w := range recorder.completedWindows { - if w.ID == "truncate-now" && w.Status == IncidentWindowStatusTruncated { - foundTruncated = true - } - if w.ID == "expire-now" && w.Status == IncidentWindowStatusComplete { - foundCompleted = true - } - } - if !foundTruncated { - t.Fatal("expected truncated completed window") - } - if !foundCompleted { - t.Fatal("expected completed expired window") - } -} - -func TestStartRecordingCopiesPreIncidentBuffer(t *testing.T) { - recorder := NewIncidentRecorder(IncidentRecorderConfig{ - PreIncidentWindow: time.Minute, - PostIncidentWindow: time.Minute, - }) - recorder.preIncidentBuffer["res-1"] = []IncidentDataPoint{ - {Timestamp: time.Now().Add(-30 * time.Second), Metrics: map[string]float64{"cpu": 1}}, - } - - windowID := recorder.StartRecording("res-1", "db", "agent", "alert", "alert-1") - window := recorder.activeWindows[windowID] - if window == nil { - t.Fatalf("expected active window %s", windowID) - } - if len(window.DataPoints) != 1 { - t.Fatalf("expected pre-incident points to be copied, got %d", len(window.DataPoints)) - } -} - -func TestCompleteWindowNoopForNonRecordingStatus(t *testing.T) { - recorder := NewIncidentRecorder(IncidentRecorderConfig{}) - window := &IncidentWindow{ - ID: "already-complete", - ResourceID: "res-1", - Status: IncidentWindowStatusComplete, - } - recorder.activeWindows[window.ID] = window - - recorder.completeWindowLocked(window) - - if len(recorder.completedWindows) != 0 { - t.Fatalf("expected no completed windows to be appended, got %d", len(recorder.completedWindows)) - } - if _, ok := recorder.activeWindows[window.ID]; !ok { - t.Fatal("expected window to remain in active map when completion is skipped") - } -} - -func TestComputeSummaryNoDataPoints(t *testing.T) { - recorder := NewIncidentRecorder(IncidentRecorderConfig{}) - if summary := recorder.computeSummary(&IncidentWindow{}); summary != nil { - t.Fatal("expected nil summary when there are no data points") - } -} - -func TestTrimCompletedWindowsEnforcesMaxWindows(t *testing.T) { - recorder := NewIncidentRecorder(IncidentRecorderConfig{ - MaxWindows: 2, - RetentionDuration: time.Hour, - }) - now := time.Now() - recorder.completedWindows = []*IncidentWindow{ - {ID: "w1", EndTime: &now}, - {ID: "w2", EndTime: &now}, - {ID: "w3", EndTime: &now}, - } - - recorder.trimCompletedWindows() - - if len(recorder.completedWindows) != 2 { - t.Fatalf("expected 2 windows after trim, got %d", len(recorder.completedWindows)) - } - if recorder.completedWindows[0].ID != "w2" || recorder.completedWindows[1].ID != "w3" { - t.Fatalf("expected newest windows to be retained, got %s and %s", recorder.completedWindows[0].ID, recorder.completedWindows[1].ID) - } -} - -func TestGetWindowActiveAndMissing(t *testing.T) { - recorder := NewIncidentRecorder(IncidentRecorderConfig{}) - recorder.activeWindows["active-1"] = &IncidentWindow{ID: "active-1", ResourceID: "res-1"} - - if got := recorder.GetWindow("active-1"); got == nil { - t.Fatal("expected active window to be returned") - } - if got := recorder.GetWindow("does-not-exist"); got != nil { - t.Fatal("expected nil for missing window") - } -} - -func TestStopHandlesSaveErrorPath(t *testing.T) { - recorder := NewIncidentRecorder(IncidentRecorderConfig{ - DataDir: t.TempDir(), - SampleInterval: 50 * time.Millisecond, - PreIncidentWindow: 10 * time.Millisecond, - }) - now := time.Now() - recorder.completedWindows = []*IncidentWindow{ - { - ID: "nan-window", - EndTime: &now, - DataPoints: []IncidentDataPoint{ - {Timestamp: now, Metrics: map[string]float64{"cpu": math.NaN()}}, - }, - }, - } - - recorder.Start() - recorder.Stop() -} - -type logSignalWriter struct { - hit atomic.Bool -} - -func (w *logSignalWriter) Write(p []byte) (int, error) { - if bytes.Contains(p, []byte("Failed to save incident windows")) { - w.hit.Store(true) - } - return len(p), nil -} - -func TestCompleteWindowAsyncSaveErrorPath(t *testing.T) { - base := t.TempDir() - fileAsDir := filepath.Join(base, "file-instead-of-dir") - if err := os.WriteFile(fileAsDir, []byte("x"), 0600); err != nil { - t.Fatalf("write setup file: %v", err) - } - - writer := &logSignalWriter{} - originalLogger := log.Logger - log.Logger = zerolog.New(writer).Level(zerolog.WarnLevel) - t.Cleanup(func() { - log.Logger = originalLogger - }) - - recorder := NewIncidentRecorder(IncidentRecorderConfig{ - DataDir: fileAsDir, - RetentionDuration: time.Hour, - MaxWindows: 10, - }) - window := &IncidentWindow{ - ID: "async-save-error", - ResourceID: "res-1", - Status: IncidentWindowStatusRecording, - DataPoints: []IncidentDataPoint{ - {Timestamp: time.Now(), Metrics: map[string]float64{"cpu": 42}}, - }, - } - recorder.activeWindows[window.ID] = window - - recorder.completeWindowLocked(window) - recorder.waitForPendingSaves() - - if !writer.hit.Load() { - t.Fatal("expected async save error warning to be logged") - } -} - -func TestSaveToDiskErrorPaths(t *testing.T) { - t.Run("marshal error", func(t *testing.T) { - recorder := NewIncidentRecorder(IncidentRecorderConfig{DataDir: t.TempDir()}) - now := time.Now() - recorder.completedWindows = []*IncidentWindow{ - { - ID: "marshal-fail", - EndTime: &now, - DataPoints: []IncidentDataPoint{ - {Timestamp: now, Metrics: map[string]float64{"cpu": math.NaN()}}, - }, - }, - } - - if err := recorder.saveToDisk(); err == nil { - t.Fatal("expected marshal error") - } - }) - - t.Run("mkdir error", func(t *testing.T) { - base := t.TempDir() - fileAsDir := filepath.Join(base, "file-instead-of-dir") - if err := os.WriteFile(fileAsDir, []byte("x"), 0600); err != nil { - t.Fatalf("write setup file: %v", err) - } - - recorder := NewIncidentRecorder(IncidentRecorderConfig{DataDir: fileAsDir}) - now := time.Now() - recorder.completedWindows = []*IncidentWindow{{ID: "w1", EndTime: &now}} - - if err := recorder.saveToDisk(); err == nil { - t.Fatal("expected mkdir error") - } - }) - - t.Run("write temp file error", func(t *testing.T) { - dir := t.TempDir() - recorder := NewIncidentRecorder(IncidentRecorderConfig{DataDir: dir}) - now := time.Now() - recorder.completedWindows = []*IncidentWindow{{ID: "w2", EndTime: &now}} - recorder.filePath = filepath.Join(dir, "missing-subdir", "incident_windows.json") - - if err := recorder.saveToDisk(); err == nil { - t.Fatal("expected write temp file error") - } - }) -} - -func TestLoadFromDiskErrorPaths(t *testing.T) { - t.Run("empty file path", func(t *testing.T) { - recorder := NewIncidentRecorder(IncidentRecorderConfig{}) - recorder.filePath = "" - if err := recorder.loadFromDisk(); err != nil { - t.Fatalf("expected nil error for empty file path, got %v", err) - } - }) - - t.Run("missing file", func(t *testing.T) { - recorder := NewIncidentRecorder(IncidentRecorderConfig{DataDir: t.TempDir()}) - if err := recorder.loadFromDisk(); err != nil { - t.Fatalf("expected nil error for missing file, got %v", err) - } - }) - - t.Run("read error", func(t *testing.T) { - recorder := NewIncidentRecorder(IncidentRecorderConfig{}) - recorder.filePath = t.TempDir() - if err := recorder.loadFromDisk(); err == nil { - t.Fatal("expected read error") - } - }) - - t.Run("unmarshal error", func(t *testing.T) { - path := filepath.Join(t.TempDir(), "incident_windows.json") - if err := os.WriteFile(path, []byte("{"), 0600); err != nil { - t.Fatalf("write invalid json: %v", err) - } - - recorder := NewIncidentRecorder(IncidentRecorderConfig{}) - recorder.filePath = path - if err := recorder.loadFromDisk(); err == nil { - t.Fatal("expected unmarshal error") - } - }) -} - -func TestCopyWindowNilAndIntToStringEdges(t *testing.T) { - if copyWindow(nil) != nil { - t.Fatal("expected nil copy for nil input") - } - - if got := intToString(0); got != "0" { - t.Fatalf("expected 0, got %s", got) - } - if got := intToString(-42); got != "-42" { - t.Fatalf("expected -42, got %s", got) - } -} diff --git a/internal/metrics/incident_recorder_test.go b/internal/metrics/incident_recorder_test.go deleted file mode 100644 index 2d48ffe84..000000000 --- a/internal/metrics/incident_recorder_test.go +++ /dev/null @@ -1,438 +0,0 @@ -package metrics - -import ( - "os" - "path/filepath" - "strings" - "testing" - "time" -) - -type stubMetricsProvider struct { - metricsByID map[string]map[string]float64 - ids []string -} - -func (s *stubMetricsProvider) GetCurrentMetrics(resourceID string) (map[string]float64, error) { - metrics, ok := s.metricsByID[resourceID] - if !ok { - return nil, errNoMetrics(resourceID) - } - copied := make(map[string]float64, len(metrics)) - for k, v := range metrics { - copied[k] = v - } - return copied, nil -} - -func (s *stubMetricsProvider) GetMonitoredResourceIDs() []string { - return append([]string{}, s.ids...) -} - -type errNoMetrics string - -func (e errNoMetrics) Error() string { - return "no metrics for " + string(e) -} - -func TestNewIncidentRecorderDefaults(t *testing.T) { - recorder := NewIncidentRecorder(IncidentRecorderConfig{}) - - if recorder.config.SampleInterval != 5*time.Second { - t.Fatalf("expected default sample interval, got %s", recorder.config.SampleInterval) - } - if recorder.config.PreIncidentWindow != 5*time.Minute { - t.Fatalf("expected default pre-incident window, got %s", recorder.config.PreIncidentWindow) - } - if recorder.config.PostIncidentWindow != 10*time.Minute { - t.Fatalf("expected default post-incident window, got %s", recorder.config.PostIncidentWindow) - } - if recorder.config.MaxDataPointsPerWindow != 500 { - t.Fatalf("expected default max data points, got %d", recorder.config.MaxDataPointsPerWindow) - } - if recorder.config.MaxWindows != 100 { - t.Fatalf("expected default max windows, got %d", recorder.config.MaxWindows) - } - if recorder.config.RetentionDuration != 24*time.Hour { - t.Fatalf("expected default retention, got %s", recorder.config.RetentionDuration) - } -} - -func TestNewIncidentRecorderTrimsDataDir(t *testing.T) { - dir := t.TempDir() - - recorder := NewIncidentRecorder(IncidentRecorderConfig{ - DataDir: " " + dir + " ", - }) - - if recorder.config.DataDir != dir { - t.Fatalf("expected trimmed data dir %q, got %q", dir, recorder.config.DataDir) - } - if recorder.dataDir != dir { - t.Fatalf("expected recorder data dir %q, got %q", dir, recorder.dataDir) - } - wantPath := filepath.Join(dir, "incident_windows.json") - if recorder.filePath != wantPath { - t.Fatalf("expected file path %q, got %q", wantPath, recorder.filePath) - } -} - -func TestNewIncidentRecorderWhitespaceOnlyDataDirDisablesPersistence(t *testing.T) { - recorder := NewIncidentRecorder(IncidentRecorderConfig{ - DataDir: " ", - }) - - if recorder.config.DataDir != "" { - t.Fatalf("expected empty config data dir, got %q", recorder.config.DataDir) - } - if recorder.dataDir != "" { - t.Fatalf("expected empty recorder data dir, got %q", recorder.dataDir) - } - if recorder.filePath != "" { - t.Fatalf("expected empty file path, got %q", recorder.filePath) - } -} - -func TestStartRecordingExtendsWindow(t *testing.T) { - recorder := NewIncidentRecorder(IncidentRecorderConfig{ - PreIncidentWindow: time.Minute, - PostIncidentWindow: time.Minute, - }) - - firstID := recorder.StartRecording("res-1", "db", "agent", "alert", "alert-1") - firstWindow := recorder.activeWindows[firstID] - if firstWindow == nil { - t.Fatalf("expected window for %s", firstID) - } - firstEnd := *firstWindow.EndTime - - secondID := recorder.StartRecording("res-1", "db", "agent", "alert", "alert-2") - if secondID != firstID { - t.Fatalf("expected same window ID, got %s and %s", firstID, secondID) - } - secondWindow := recorder.activeWindows[secondID] - if secondWindow.EndTime.Before(firstEnd) { - t.Fatalf("expected end time to extend or remain, got %s before %s", secondWindow.EndTime, firstEnd) - } -} - -func TestStartRecordingCanonicalizesLegacyHostAlias(t *testing.T) { - recorder := NewIncidentRecorder(IncidentRecorderConfig{ - PreIncidentWindow: time.Minute, - PostIncidentWindow: time.Minute, - }) - - windowID := recorder.StartRecording("res-1", "db", "host", "alert", "alert-1") - window := recorder.activeWindows[windowID] - if window == nil { - t.Fatalf("expected window for %s", windowID) - } - if window.ResourceType != "agent" { - t.Fatalf("expected legacy host alias to canonicalize to agent, got %q", window.ResourceType) - } -} - -func TestRecordSampleBuffersAndCleansUp(t *testing.T) { - recorder := NewIncidentRecorder(IncidentRecorderConfig{ - PreIncidentWindow: time.Minute, - PostIncidentWindow: time.Minute, - MaxDataPointsPerWindow: 10, - }) - - provider := &stubMetricsProvider{ - metricsByID: map[string]map[string]float64{ - "res-1": {"cpu": 1}, - "res-2": {"cpu": 2}, - }, - ids: []string{"res-1", "res-2"}, - } - recorder.SetMetricsProvider(provider) - - recorder.preIncidentBuffer["gone"] = []IncidentDataPoint{ - {Timestamp: time.Now().Add(-time.Minute), Metrics: map[string]float64{"cpu": 0.5}}, - } - - windowID := recorder.StartRecording("res-1", "db", "agent", "alert", "alert-1") - recorder.recordSample() - - window := recorder.activeWindows[windowID] - if window == nil { - t.Fatalf("expected active window %s", windowID) - } - if len(window.DataPoints) != 1 { - t.Fatalf("expected 1 data point, got %d", len(window.DataPoints)) - } - - if len(recorder.preIncidentBuffer["res-1"]) == 0 { - t.Fatalf("expected pre-incident buffer for res-1") - } - if len(recorder.preIncidentBuffer["res-2"]) == 0 { - t.Fatalf("expected pre-incident buffer for res-2") - } - if _, ok := recorder.preIncidentBuffer["gone"]; ok { - t.Fatalf("expected cleanup of unmonitored resource buffer") - } -} - -func TestStopRecordingCompletesWindow(t *testing.T) { - recorder := NewIncidentRecorder(IncidentRecorderConfig{ - PreIncidentWindow: time.Minute, - PostIncidentWindow: time.Minute, - }) - provider := &stubMetricsProvider{ - metricsByID: map[string]map[string]float64{ - "res-1": {"cpu": 1}, - }, - ids: []string{"res-1"}, - } - recorder.SetMetricsProvider(provider) - - windowID := recorder.StartRecording("res-1", "db", "agent", "alert", "alert-1") - recorder.recordSample() - recorder.StopRecording(windowID) - - if _, ok := recorder.activeWindows[windowID]; ok { - t.Fatalf("expected window %s to be removed from active windows", windowID) - } - if len(recorder.completedWindows) != 1 { - t.Fatalf("expected 1 completed window, got %d", len(recorder.completedWindows)) - } - if recorder.completedWindows[0].Status != IncidentWindowStatusComplete { - t.Fatalf("expected completed status, got %s", recorder.completedWindows[0].Status) - } - if recorder.completedWindows[0].Summary == nil { - t.Fatalf("expected summary to be computed") - } -} - -func TestComputeSummary(t *testing.T) { - recorder := NewIncidentRecorder(IncidentRecorderConfig{}) - start := time.Now().Add(-time.Second) - end := start.Add(time.Second) - window := &IncidentWindow{ - DataPoints: []IncidentDataPoint{ - {Timestamp: start, Metrics: map[string]float64{"cpu": 1, "mem": 4}}, - {Timestamp: end, Metrics: map[string]float64{"cpu": 3, "mem": 2}}, - }, - } - - summary := recorder.computeSummary(window) - if summary == nil { - t.Fatalf("expected summary") - } - if summary.DataPoints != 2 { - t.Fatalf("expected 2 data points, got %d", summary.DataPoints) - } - if summary.Peaks["cpu"] != 3 || summary.Lows["cpu"] != 1 { - t.Fatalf("unexpected cpu stats: peaks=%v lows=%v", summary.Peaks["cpu"], summary.Lows["cpu"]) - } - if summary.Peaks["mem"] != 4 || summary.Lows["mem"] != 2 { - t.Fatalf("unexpected mem stats: peaks=%v lows=%v", summary.Peaks["mem"], summary.Lows["mem"]) - } - if summary.Averages["cpu"] != 2 { - t.Fatalf("unexpected cpu average: %v", summary.Averages["cpu"]) - } - if summary.Averages["mem"] != 3 { - t.Fatalf("unexpected mem average: %v", summary.Averages["mem"]) - } - if summary.Changes["cpu"] != 2 || summary.Changes["mem"] != -2 { - t.Fatalf("unexpected changes: cpu=%v mem=%v", summary.Changes["cpu"], summary.Changes["mem"]) - } - if summary.Duration != time.Second { - t.Fatalf("unexpected duration: %s", summary.Duration) - } -} - -func TestCopyWindowDeepCopy(t *testing.T) { - now := time.Now() - end := now.Add(time.Second) - window := &IncidentWindow{ - ID: "window-1", - EndTime: &end, - DataPoints: []IncidentDataPoint{ - { - Timestamp: now, - Metrics: map[string]float64{"cpu": 1}, - Metadata: map[string]interface{}{"host": "node-1"}, - }, - }, - Summary: &IncidentSummary{ - Peaks: map[string]float64{"cpu": 1}, - Lows: map[string]float64{"cpu": 1}, - Averages: map[string]float64{"cpu": 1}, - Changes: map[string]float64{"cpu": 0}, - Anomalies: []string{"initial"}, - }, - } - - clone := copyWindow(window) - if clone == nil || clone == window { - t.Fatalf("expected deep copy") - } - if clone.Summary == window.Summary { - t.Fatalf("expected summary to be copied") - } - - window.DataPoints[0].Metrics["cpu"] = 9 - window.DataPoints[0].Metadata["host"] = "mutated" - window.Summary.Peaks["cpu"] = 9 - window.Summary.Anomalies[0] = "mutated" - *window.EndTime = end.Add(5 * time.Second) - window.Summary.Peaks["cpu"] = 9 - - if clone.DataPoints[0].Metrics["cpu"] != 1 { - t.Fatalf("expected data points to be copied") - } - if clone.DataPoints[0].Metadata["host"] != "node-1" { - t.Fatalf("expected metadata to be copied") - } - if clone.EndTime.Equal(*window.EndTime) { - t.Fatalf("expected end time to be copied") - } - if clone.Summary.Peaks["cpu"] != 1 { - t.Fatalf("expected summary maps to be copied") - } - if clone.Summary.Anomalies[0] != "initial" { - t.Fatalf("expected summary anomalies to be copied") - } -} - -func TestSaveAndLoad(t *testing.T) { - dir := t.TempDir() - recorder := NewIncidentRecorder(IncidentRecorderConfig{DataDir: dir}) - - end := time.Now() - recorder.completedWindows = []*IncidentWindow{ - { - ID: "window-1", - EndTime: &end, - Status: IncidentWindowStatusComplete, - DataPoints: []IncidentDataPoint{{Timestamp: end, Metrics: map[string]float64{"cpu": 1}}}, - }, - } - - if err := recorder.saveToDisk(); err != nil { - t.Fatalf("save failed: %v", err) - } - - loaded := NewIncidentRecorder(IncidentRecorderConfig{DataDir: dir}) - window := loaded.GetWindow("window-1") - if window == nil { - t.Fatalf("expected window to load from disk") - } - if window.Status != IncidentWindowStatusComplete { - t.Fatalf("expected status to persist, got %s", window.Status) - } -} - -func TestSaveToDiskSecuresPermissions(t *testing.T) { - dir := t.TempDir() - if err := os.Chmod(dir, 0o755); err != nil { - t.Fatalf("chmod dir failed: %v", err) - } - - recorder := NewIncidentRecorder(IncidentRecorderConfig{DataDir: dir}) - recorder.completedWindows = []*IncidentWindow{{ID: "window-1", EndTime: ptrTime(time.Now())}} - - if err := recorder.saveToDisk(); err != nil { - t.Fatalf("save failed: %v", err) - } - - dirInfo, err := os.Stat(dir) - if err != nil { - t.Fatalf("stat dir failed: %v", err) - } - if got := dirInfo.Mode().Perm(); got != 0o700 { - t.Fatalf("expected dir permissions 0700, got %o", got) - } - - fileInfo, err := os.Stat(filepath.Join(dir, "incident_windows.json")) - if err != nil { - t.Fatalf("stat file failed: %v", err) - } - if got := fileInfo.Mode().Perm(); got != 0o600 { - t.Fatalf("expected file permissions 0600, got %o", got) - } -} - -func TestLoadFromDiskRejectsSymlink(t *testing.T) { - dir := t.TempDir() - target := filepath.Join(dir, "target.json") - if err := os.WriteFile(target, []byte("{}"), 0o600); err != nil { - t.Fatalf("write target failed: %v", err) - } - link := filepath.Join(dir, "incident_windows.json") - requireSymlinkOrSkip(t, target, link) - - recorder := &IncidentRecorder{ - config: DefaultIncidentRecorderConfig(), - dataDir: dir, - filePath: link, - } - err := recorder.loadFromDisk() - if err == nil { - t.Fatal("expected symlink path to be rejected") - } - if !strings.Contains(err.Error(), "symlink") { - t.Fatalf("expected symlink error, got: %v", err) - } -} - -func TestLoadFromDiskRejectsOversizedFile(t *testing.T) { - dir := t.TempDir() - path := filepath.Join(dir, "incident_windows.json") - tooLarge := make([]byte, maxIncidentWindowsFileSize+1) - if err := os.WriteFile(path, tooLarge, 0o600); err != nil { - t.Fatalf("write oversized file failed: %v", err) - } - - recorder := &IncidentRecorder{ - config: DefaultIncidentRecorderConfig(), - dataDir: dir, - filePath: path, - } - err := recorder.loadFromDisk() - if err == nil { - t.Fatal("expected oversized file to be rejected") - } - if !strings.Contains(err.Error(), "exceeds size limit") { - t.Fatalf("expected size-limit error, got: %v", err) - } -} - -func TestSaveToDiskRejectsSymlinkDestination(t *testing.T) { - dir := t.TempDir() - target := filepath.Join(dir, "target.json") - if err := os.WriteFile(target, []byte("secret"), 0o600); err != nil { - t.Fatalf("write target failed: %v", err) - } - link := filepath.Join(dir, "incident_windows.json") - requireSymlinkOrSkip(t, target, link) - - recorder := &IncidentRecorder{ - config: DefaultIncidentRecorderConfig(), - dataDir: dir, - filePath: link, - completedWindows: []*IncidentWindow{ - {ID: "window-1", EndTime: ptrTime(time.Now())}, - }, - } - err := recorder.saveToDisk() - if err == nil { - t.Fatal("expected symlink destination to be rejected") - } - if !strings.Contains(err.Error(), "symlink") { - t.Fatalf("expected symlink error, got: %v", err) - } -} - -func ptrTime(v time.Time) *time.Time { - return &v -} - -func requireSymlinkOrSkip(t *testing.T, target, link string) { - t.Helper() - if err := os.Symlink(target, link); err != nil { - t.Skipf("symlink not supported in this environment: %v", err) - } -} diff --git a/internal/monitoring/monitor_alert_handling_test.go b/internal/monitoring/monitor_alert_handling_test.go index ea2c40375..945a5c3ea 100644 --- a/internal/monitoring/monitor_alert_handling_test.go +++ b/internal/monitoring/monitor_alert_handling_test.go @@ -414,8 +414,28 @@ func TestLifecycleReplayMaterializesImportedHistoryTimeline(t *testing.T) { monitor := &Monitor{alertManager: manager, incidentStore: incidentStore} resourceStore := unifiedresources.NewMemoryStore() adapter := unifiedresources.NewMonitorAdapter(unifiedresources.NewRegistry(resourceStore)) - monitor.SetResourceStore(adapter) - monitor.SetResourceStore(adapter) + // Hold replay at its serialization boundary. Router construction attaches + // this store, so attachment must return even while history repair cannot + // make progress. Eventual timeline assertions alone miss a synchronous + // replay regression that stalls startup on an upgrade backlog. + monitor.alertProjectionReplayMu.Lock() + attached := make(chan struct{}) + go func() { + monitor.SetResourceStore(adapter) + monitor.SetResourceStore(adapter) + close(attached) + }() + select { + case <-attached: + monitor.alertProjectionReplayMu.Unlock() + case <-time.After(2 * time.Second): + // Release the probe before failing, including for a synchronous-replay + // negative control, so no goroutine retains the test's stores. + monitor.alertProjectionReplayMu.Unlock() + <-attached + monitor.alertProjectionWG.Wait() + t.Fatal("resource-store attachment waited for lifecycle replay") + } monitor.alertProjectionWG.Wait() timeline := incidentStore.GetTimelineByAlertAt(snapshot.ID, snapshot.StartTime) diff --git a/internal/unifiedresources/code_standards_test.go b/internal/unifiedresources/code_standards_test.go index dd48d807e..f9c85fa78 100644 --- a/internal/unifiedresources/code_standards_test.go +++ b/internal/unifiedresources/code_standards_test.go @@ -2515,13 +2515,6 @@ func TestV6DirectHostAliasValidatorCoverage(t *testing.T) { `[]string{"host", "guest", "docker", "container", "lxc", "qemu", "docker_container", "docker_service"}`, }, }, - { - path: filepath.Join(repoRoot, "internal", "ai", "incident_coordinator_additional_test.go"), - requiredSnippets: []string{ - `TestIncidentCoordinator_OnAnomalyDetected_CanonicalizesLegacyHostAlias`, - `expected anomaly recording resource type to be canonicalized to agent`, - }, - }, { path: filepath.Join(repoRoot, "internal", "ai", "tools", "tools_metrics_alerts_test.go"), requiredSnippets: []string{ @@ -2557,13 +2550,6 @@ func TestV6DirectHostAliasValidatorCoverage(t *testing.T) { `legacy k8s alias rejected`, }, }, - { - path: filepath.Join(repoRoot, "internal", "metrics", "incident_recorder_test.go"), - requiredSnippets: []string{ - `TestStartRecordingCanonicalizesLegacyHostAlias`, - `expected legacy host alias to canonicalize to agent`, - }, - }, { path: filepath.Join(repoRoot, "internal", "api", "resourceapi", "resources_test.go"), requiredSnippets: []string{ diff --git a/scripts/check-alert-dispatch-copy.mjs b/scripts/check-alert-dispatch-copy.mjs new file mode 100644 index 000000000..c3daecccb --- /dev/null +++ b/scripts/check-alert-dispatch-copy.mjs @@ -0,0 +1,123 @@ +// Isolated real-browser component qualification; no installed backend or delivery claim. +import { createServer } from "../frontend-modern/node_modules/vite/dist/node/index.js"; +import solid from "../frontend-modern/node_modules/vite-plugin-solid/dist/esm/index.mjs"; +import { chromium } from "@playwright/test"; +import { resolve } from "node:path"; +import { mkdirSync } from "node:fs"; +import assert from "node:assert/strict"; +const root = resolve("frontend-modern"); +process.chdir(root); +const fixture = ` +import { render } from 'solid-js/web'; +import { Router, Route } from '@solidjs/router'; +import { AlertsAPI } from '/src/api/alerts'; +import { NotificationsAPI } from '/src/api/notifications'; +import { OverviewTab } from '/src/features/alerts/OverviewTab'; +import '/src/index.css'; +const ids = ['ready','cooldown','pending']; +AlertsAPI.getDeliveryDiagnoses = async () => ids.map(id => ({ +alertIdentifier:id, alertId:id, trackingKey:id, status:id==='cooldown'?'suppressed':'would_send', +reason:id==='cooldown'?'cooldown':'ready', lastNotified:id==='pending'?undefined:'2026-08-26T10:15:00Z', +nextEligibleAt:id==='cooldown'?'2026-08-26T10:20:00Z':undefined +})); +AlertsAPI.getEvents = async () => []; +NotificationsAPI.getHealth = async () => ({queue:{status:'healthy'}}); +const alerts = Object.fromEntries(ids.map(id => [id, {id,resourceId:id,resourceName:'VM '+id, +type:'cpu',level:'warning',message:'High CPU on '+id,startTime:'2026-08-26T10:00:00Z',acknowledged:false,node:'node1'}])); +function Fixture() { return
{}} showQuickTip={()=>false} dismissQuickTip={()=>{}} showAcknowledged={()=>true} +setShowAcknowledged={()=>{}} alertsDisabled={()=>false}/>
; } +render(()=>,document.getElementById('root')); +`; +const server = await createServer({ + root, + configFile: false, + optimizeDeps: { + noDiscovery: true, + entries: [], + esbuildOptions: { target: "esnext" }, + }, + esbuild: { target: "esnext" }, + plugins: [ + solid(), + { + name: "dispatch-fixture", + configureServer(s) { + s.middlewares.use((req, res, next) => { + if (req.url === "/qualification") { + res.setHeader("Content-Type", "text/html"); + res.end( + '
', + ); + } else next(); + }); + }, + resolveId(id) { + if (id === "/dispatch-fixture.tsx") return id; + }, + load(id) { + if (id === "/dispatch-fixture.tsx") return fixture; + }, + }, + ], + resolve: { alias: { "@": resolve(root, "src") } }, + server: { host: "127.0.0.1", port: 5198, strictPort: true }, +}); +let browser; +try { + await server.listen(); + browser = await chromium.launch({ headless: true }); + mkdirSync("/tmp/pulse-alert-dispatch", { recursive: true }); + for (const width of [1440, 900, 390]) { + const page = await browser.newPage({ viewport: { width, height: 1000 } }); + const errors = []; + page.on("pageerror", (e) => { + errors.push(e.message); + console.error(e.message); + }); + page.on("console", (m) => { + if (m.type() === "error") console.error(m.text()); + }); + await page.route("http://127.0.0.1:5198/api/**", (route) => + route.fulfill({ json: [] }), + ); + await page.goto("http://127.0.0.1:5198/qualification"); + await page.getByText(/^Dispatch requested .*next eligible/).waitFor(); + assert.equal(await page.getByText(/^Dispatch requested /).count(), 2); + assert.equal(await page.getByText(/^Notified /).count(), 0); + assert.equal( + await page.getByText("Notification pending", { exact: true }).count(), + 1, + ); + for (const label of await page.getByText(/^Dispatch requested /).all()) { + assert.equal( + await label.evaluate((el) => { + const r = document.createRange(); + r.selectNodeContents(el); + return [...r.getClientRects()].every( + (b) => b.left >= 0 && b.right <= innerWidth, + ); + }), + true, + "status text must fit viewport", + ); + } + assert.deepEqual(errors, []); + await page.screenshot({ + path: "/tmp/pulse-alert-dispatch/" + width + ".png", + fullPage: true, + }); + await page.close(); + } + console.log( + JSON.stringify({ + result: "passed", + viewports: [1440, 900, 390], + scope: + "Real Overview and Chromium; scripted diagnoses, not installed delivery or receipt", + }), + ); +} finally { + await browser?.close(); + await server.close(); +}