diff --git a/docs/qualification/PATROL_ASSISTANT_CUSTOMER_JOURNEY.md b/docs/qualification/PATROL_ASSISTANT_CUSTOMER_JOURNEY.md index fdcf448c3..390b1ad7b 100644 --- a/docs/qualification/PATROL_ASSISTANT_CUSTOMER_JOURNEY.md +++ b/docs/qualification/PATROL_ASSISTANT_CUSTOMER_JOURNEY.md @@ -65,7 +65,7 @@ reproduction evidence, not a representative customer success rate. | 2. Shared evidence | Preserve canonical risk reasons and SMART counters, source/time semantics and history across tools/turns. | Regression tests preserve unknown versus zero and all canonical evidence. Real responses can inspect the same facts as the product. | Implemented and qualified for the named shared-evidence defects. Canonical disk detail, risk and cadence pass real data-path proof. Affected package and concurrency checks pass. Integrated CI later exposed remaining query and allocation regressions. The final bounded query-reuse correction passes complete selected exact-base worker comparisons and full metrics/database and focused race checks. Final landing CI passed and PRs #1928 and #1929 merged. Real-model interpretation failures remain tracked in step 5. | | 3. Diagnostic orchestration | Correct proposal-as-proof. Audit triage budgets, unmatched-signal evaluation, assessment completion and investigation cutoffs. | No code-written causal conclusion. No quality inferred from tool, flag or finding counts. Each retained pass has an objective reason. Safety boundaries and incomplete outcomes remain explicit. | Proposal promotion and capture inference were removed in c5d2f56dda. Commit 668af3fe6b removes investigation success-call floors, checkpoint instructions and generic call-count wrap-up rules. The detection slice removes contextless follow-up passes, flag/report-count policy and first-finding completion modes. Full chat and AI suites, focused API and conversation race tests pass. Real-model/action outcome qualification remains open. | | 4. Issue through verified outcome | Follow existing issue/investigation/action records into Assistant, approval, execution and independent readback. | Accepted proposal is visibly distinct from execution and verification. Rejected or unsupported actions do not become success. Uncertainty can survive an action proposal. | Existing foundation, full journey qualification pending. | -| 5. Ground-truth qualification and landing | Extend existing qualification tooling only where necessary. Exercise healthy/unhealthy, dependency, missing-access, storage/backup and approved/rejected action cases. Inspect the final browser journey at desktop and narrow widths. | Record exact source/model/permissions, evidence, decisions, faults/misses, latency and verification. Fix in-scope failures, pass appropriate proofs and land scoped commits. | Pending. | +| 5. Ground-truth qualification and landing | Extend existing qualification tooling only where necessary. Exercise healthy/unhealthy, dependency, missing-access, storage/backup and approved/rejected action cases. Inspect the final browser journey at desktop and narrow widths. | Record exact source/model/permissions, evidence, decisions, faults/misses, latency and verification. Fix in-scope failures, pass appropriate proofs and land scoped commits. | Partial. Regression, controlled browser and live collector evidence are recorded below. Real-model diagnosis, linked approval/action outcomes and installed collector qualification remain open. | Use one shared runtime and the existing qualification runner, not a second product intelligence engine or a new parallel lifecycle. Preserve independent @@ -93,12 +93,13 @@ Each removal must run its focused regression and affected complete journey. ### Completion and external dependencies The local implementation goal remains open until required qualification is -performed. Ordinary Assistant requests work with the current subscription route, -but autonomous Patrol has an explicit provider-policy refusal and remains -blocked. Do not rephrase the refused probe, bypass the readiness boundary or -count an interactive request as an autonomous Patrol pass. A supported provider -path is required for that qualification. Prepare other work while resolving the -provider dependency through supported configuration. +performed. The maintainer authorized Gemini 3.8 Flash through OpenRouter with a +US$5 key limit and one-day expiry on 2026-09-06. That supported route passes the +streaming readiness and initial live Watch and dependency cases recorded below. +The earlier Claude subscription refusal belongs to the exact synthetic +continuation request. It does not establish a blanket restriction on autonomous +monitoring. The refused request has not been retried or rephrased. Readiness is +not evidence that diagnosis, action execution or independent recovery succeeds. Release publication and wider product readiness are separate. Independent volunteered Pro environments are still required before claiming repeatable @@ -1408,3 +1409,849 @@ correction above is a separate scoped change and requires its own landing checks The redesign remains open for reliable interpretation, the config-read contract, storage/backup, approved/rejected action outcomes and supported autonomous Patrol qualification. Wider customer readiness still requires independent Pro environments. + +### Docker measurement correction plan, 2026-09-06 + +The next shared-source correction distinguishes absent block-I/O observations +from measured idle zero and removes the Docker layer-size ratio from filesystem +capacity. Counter presence uses the existing rate tracker contract. Missing +reports must not reset the baseline or fabricate samples. Container layer sizes +remain descriptive metadata. + +Persisted Docker-family disk series previously mixed invalid capacity ratios and +unobserved I/O zeros with real measurements. New disk observations use separate +physical series keys while public metric names remain unchanged. Retained reads +exclude ambiguous legacy disk series without deleting or relabelling them. The +shared app-container storage family also serves non-Docker providers, so new +valid capacity observations must remain supported. Non-Docker series retain +existing behavior. Explicit zero must survive every retained-read API and rollup. + +Browser verification is required after the final backend build. Interaction +matrix: `/docker` at 1440x1000, 900x1000 and 390x1000, container selection, +resource drawer open/close, current metrics, history expansion, measured idle, +unavailable readings, and reload. Inspect actual pixels, scrolling, focus and +Escape dismissal. `/patrol` evidence rendering must preserve absent versus zero +in current-resource and retained-history tool results. Controlled responses may +qualify rendering but cannot qualify diagnosis. Live read-only API observations +must bind to the rebuilt backend. No autonomous subscription retry or paid-model +request is authorized by this correction. + +Collection also carries optional presence for each I/O direction. Explicit zero +entries survive the report JSON. For older reports without presence, only +positive counters establish an observation, so ambiguous zeros remain unavailable +until the agent is updated or a positive baseline exists. This does not require +re-enrollment. Docker-host first-disk history and network-counter presence are +adjacent limits outside this container block-I/O correction. + +The final browser matrix also covers the shared host I/O table, Docker host +Overview and Machines table/tooltip at the same three widths. Partial read/write +observations must show a missing marker for the absent direction, retain measured +zero, and remain excluded from sums used for sorting and comparison. Exercise +column selection, hover/focus, tooltip dismissal and scrolling where present. + +### Docker correction qualification and scope + +The implementation carries per-direction presence from collection and report +JSON into the existing rate tracker, canonical resource metrics, persisted +history and resource-to-browser conversion. REST resource adaptation also +preserves optional rates. Shared rate formatting keeps missing values distinct +from zero in Machines and Docker host details. Incomplete rates do not become +complete throughput totals for sorting or comparison. A browser-discovered +first-user column migration bug is corrected in the shared preference hook, so +showing Disk I/O survives the first reload. + +Legacy workload conversion in `frontend-modern/src/hooks/useWorkloads.ts` still +uses numeric direction fields with grouped availability. Its direction-level +modernization remains a separate consumer follow-up. Docker-host first-disk +history and network presence are also outside this container measurement slice. +The correction must not be represented as complete coverage of all metrics or +all monitoring surfaces. No new model competence or autonomous action result is +claimed. + +Affected Go package checks and targeted race checks ran on pulse-dev with +Go1.26.8. The changed websocket assertion now expects observed read zero with +absent write omitted. Targeted frontend suites and type checking cover optional +rates, REST conversion, sorting, formatting and column persistence. Ten paired +read-benchmark rounds used the unchanged parent store via Go overlay. The +canonical >10%, p<0.05 regression gate passed, with +0.82% timing geomean in this +scoped comparison. This is not a fleet-load or full-product performance claim. + +The Pro backend was cross-built on pulse-dev from the changed source and +installed into the existing local dev runtime. Binary SHA256: +`59f05f954ff8080bd3e8f3054b2b059281c49172ee771b2e454c255241158a4a`. +No production agent was replaced. Older agents remain compatible and treat +ambiguous zero counters conservatively. + +Private receipts: `/Volumes/Development/pulse/tmp/patrol-docker-observed-metrics/`. +Worker logs: `/opt/pulse-release-worker/patrol-docker-observed-proof/`. +PR1934's preceding identity correction merged at +`6b0abc3bee9ffa81f6ab298b5b67ee11369688a0` with all checks passing. The current +measurement slice passed final-source Playwright inspection on `/docker`, +`/standalone/machines` and `/patrol` at 1440x1000, 900x1000 and 390x1000. +Live history, controlled absence/idle/loading/error, partial host rates, column +persistence, nested picker dismissal, tooltip focus, and expanded Assistant +evidence were exercised. The source-bound receipt is +`frontend-modern/browser-verification.json`. The unused shared host table card +has type and selector coverage, not an active-route browser claim. Controlled +responses qualify rendering only. Landing checks remain separate from the +unperformed diagnosis, approved/rejected action and recovery qualifications. + +### Configuration-read correction plan + +A successful canonical container get followed by a false config `not found` +result is a source contract defect. Native configuration reads must use current +canonical inventory for identity and provider capability. Optional session +resolution preserves continuity for later actions, not proof of existence. +Explicit query restrictions must be checked before registering or refreshing a +resource. Unsupported adapters, missing configuration providers, unavailable +placement and empty provider responses remain distinct from missing inventory. +No action validation or native log-read authority changes in this slice. + +Regression matrix: TrueNAS config with absent, empty and existing session +context, canonical identity across aliases, explicit query denial without a +provider call, Docker unsupported capability, genuinely missing inventory, +unavailable placement, provider failure and nil provider response. Reproduce +the failing cases before changing runtime code. Run affected Go tools checks +and focused race coverage on pulse-dev. + +Browser matrix after the final build: `/patrol` Assistant tool result details +at 1440x1000, 900x1000 and 390x1000, available configuration, unsupported +capability, true missing resource and denied/provider-failed results. Exercise +open/closed details, hover and keyboard focus, Enter/Space, deepest output +scrolling, Escape, and persisted/reloaded evidence. Use captured actual tool +results to qualify rendering without claiming model diagnosis or native +provider integration. No autonomous subscription retry or separately billed +provider request is part of this correction. + +### Configuration-read correction qualification + +The baseline reproduced absent/empty session failures, stale session placement +and false not-found results after successful canonical gets. The corrected +read path uses canonical resource identity and current provider placement. It +checks explicit query restrictions before registration and preserves an existing +query-only session's action limits. Unsupported configuration, unavailable +provider/placement and nil provider responses carry explicit reasons and the +tool error bit. Unavailable inventory and missing read state also remain failures +rather than evidence of resource absence. Actual inventory absence remains the +existing not-found lookup result. + +Fourteen focused contract cases pass with strict resolution enabled. The +existing native-config regression, full tools package and focused race proof +passed on pulse-dev with Go1.26.8. The final-source Pro binary SHA256 is +`552699cdf2e61a4ca1cea2ac5ef4e065735cbd1dbca01665456e184bd4fc3533`. +It is installed only in the existing local dev stack. No production agent or +provider configuration was changed. + +Playwright passed on `/patrol` at 1440x1000, 900x1000 and 390x1000. Eight actual +tool results were replayed and inspected, including successful, unavailable, +missing, denied and failed reads. Expanded inputs/outputs, keyboard toggles, +scrolling, Escape, reload and controlled session restoration preserve exact +evidence and error state. Controlled session responses prove rendering and +reload behavior, not server persistence or a new model/native-provider result. +The source-bound browser receipt records those limits. Private artifacts are +under `/Volumes/Development/pulse/tmp/patrol-config-read-contract/` and worker +logs under `/opt/pulse-release-worker/patrol-config-read-proof/`. + +PR1935's Docker correction required two legacy partial-total test expectations +to be updated in `0fcb2ee147354de770dfc4b0b9672d8c2c9dceb2`. The focused 55-test +file and scoped hook passed. Its latest CI has no failures and remains pending +completion. The configuration correction still requires its own landing checks. +Real-model retest, temporal/storage interpretation, approved and rejected +action outcomes and independent recovery proof remain open. Autonomous +subscription refusal and separately billed provider approval boundaries remain +unchanged. + + +### Corrected storage evidence, ordinary Assistant retest + +The Docker observation and canonical configuration corrections are pushed to +PR #1935 at `355ac1f0a481d6dbc9a7ff3977bced0956711979`. Their exact staged +worker pre-commit checks passed without source changes. Remote checks remain +pending. The current configuration runtime also passed all fourteen contract +cases, the complete tools package and focused race checks. + +One ordinary read-only storage assessment ran on 2026-09-06 from +12:27:05.857Z to 12:30:03.726Z, an HTTP window of 177.869s. It used the existing +`claude-subscription:claude-opus-5` route, explicit `autonomous_mode=false`, +read-only control and thirteen successful tool reads. No infrastructure change, +paid-model request or autonomous readiness retry occurred. This is a single +assessment, not a success-rate or latency estimate. + +The answer identified the backup datastore at 90.6% utilisation and its active +capacity warning affecting seven workloads. It used the corrected container I/O +history, separated cumulative device counters from rates, acknowledged missing +container filesystem usage, and retained the reason for missing older history +as unknown. It did not attribute older host I/O peaks to the container whose +returned I/O window starts later. These are useful observations. + +The complete diagnosis still does not qualify. Its opening assurance that the +container is not short of space contradicts the later acknowledgement that +container filesystem usage is unavailable. It treats high retained host rates +as bucket/counter artifacts without establishing that mechanism. It includes a +host CPU maximum timestamped 21:00 the previous evening in a 03:00-04:00 window. +The suggested retention explanation is not established by capacity alone, and +available PBS job reads were not performed. Its rough growth extrapolation uses +retained extrema, not a measured first-to-last slope, and must retain that limit. +Corrected observations have not established reliable interpretation. + +The run also highlights a tool-context distinction to review: Docker-host +`agent_connected` describes the command connection, while telemetry may still +arrive through other collection paths. The model treated current telemetry and +that false connection flag as an unresolved inconsistency. Its storage-pools +request supplied `host`, although the tool schema only advertises that filter +for RAID and Ceph detail. That call returned all pools. Neither observation +justifies fabricating resource absence or collection downtime. + +Playwright exercised `/patrol` at 1440x1000 and 390x1000, the actual answer, +all thirteen expanded tool records, keyboard activation, deepest output +scrolling, Escape, reload and the persisted session. Every displayed input and +output matches the persisted tool records. Pixel review includes the answer, +evidence and the mobile table scrolled to its rightmost state. The table's +400-pixel content is reachable inside its 309-pixel horizontal viewport. +The artificial selected-route warning comes from blocked non-GET route checks, +so this does not qualify the unmodified provider-readiness UI. + +The runtime binary SHA256 stayed +`552699cdf2e61a4ca1cea2ac5ef4e065735cbd1dbca01665456e184bd4fc3533` +through the request and browser pass. Private request, source, binary, tool, +persistence, evaluation and pixel receipts are under workspace-relative +`tmp/patrol-storage-assistant-check/`. No native config action was requested, +so this assessment does not qualify model use of that corrected action. +Storage-fault ground truth, reliable diagnosis, approved/rejected actions and +independent recovery remain open. The supported autonomous provider dependency +and wider independent-Pro-environment gate remain unchanged. + + +### Command connection evidence correction plan + +Live command connectivity and retained monitoring observations are independent +facts. The shared tool contract will name command-agent connections explicitly, +including parent-node connections, without changing routing or execution policy. +Topology built without a command-connection snapshot must omit connection flags, +execution hints and connected counts rather than manufacture false/zero values. +An observed empty snapshot still reports disconnected/zero. Assistant inventory +context must preserve the same observation boundary. Existing permission, +approval and invocation checks remain authoritative. + +Regression matrix: current Docker inventory and metrics with disconnected and +connected command transport, read-only control with a connected agent, parent +node versus guest connection, topology without a connection observation versus +an observed empty set, and Assistant's seeded inventory. Run affected tools/chat +packages and focused race proof on pulse-dev. + +Browser matrix after rebuilding the local Pro backend: `/patrol` at 1440x1000, +900x1000 and 390x1000, actual captured query results showing disconnected and +connected command transport beside unchanged monitored workload evidence. +Exercise tool details open/closed, keyboard focus/activation, deepest output +scrolling, Escape, reload and persisted result presentation. Inspect pixels and +bind receipts to the final source and binary. Controlled rendering proof does +not qualify model interpretation, autonomous Patrol or infrastructure actions. + + +### Command connection evidence qualification + +The original projection failed the new regression because it labelled command +transport as generic agent connectivity and emitted connected-agent counts from +an inventory-only seed. Canonical guest search also promoted a parent-node +connection into a direct guest connection. The shared projection now retains +those distinctions. Existing host aliases remain available for non-guest +resources. No routing, approval, execution or provider policy boundary changes. + +Four canonical query cases pass: no command connection, connected read-only +transport, a direct guest connection without a parent connection, and connected +transport with control enabled. Current workload state and CPU remain available +in every case and no command is executed. Separate checks prove that topology +without a command snapshot omits connection and execution hints and connected +counts, while an observed empty snapshot retains false/zero. Assistant inventory +context inherits that same unobserved state. + +Final source proof on pulse-dev used Go1.26.8 and GOMAXPROCS4. The full tools +package passed in 59.456s and chat in 6.402s. Focused race checks passed in 1.048s +and 1.030s. The Pro runtime cross-build passed and the installed local binary +SHA256 is `bcaf748107211ee733a6dc0f4d17220d9b4d1ce1918c25bde27cf3d10c0d6379`. +The managed development process restarted onto that artifact and `/api/health` +reported healthy. No production agent was replaced. + +Playwright exercised nine captured results at `/patrol`, 1440x1000, 900x1000 +and 390x1000. Inputs, outputs and completed states match exactly before and after +controlled session reload. Hover, keyboard focus/activation, expansion/collapse, +deepest output scrolling, Escape and session selection passed. Pixel inspection +covered each distinct connection state, unchanged workload metrics and restored +mobile results. Backend and renderer hashes remained unchanged. The artificial +route-check warning and controlled persistence fixtures retain their earlier +qualification limits. No model request was part of this proof. +A read-only settings check confirms the cached `provider_refusal` still carries +its original `2026-09-05T19:46:39Z` timestamp and `patrol_capable=false`. + +Private source bindings, logs, captured outputs, runtime process/health receipts +and browser proof are at workspace-relative `tmp/patrol-command-context/`. +The change still requires its scoped pre-commit and landing checks. The preceding +PR #1935 head `4d302109cee0758a132ff150935630b50114cc05` has no reported failures +but its Build and Test and Core E2E runs are pending behind live earlier runs +on the same branch. Those workflows are not restarted or cancelled. + +This correction establishes the connection evidence contract, not reliable +interpretation. Native configuration-read model use, storage-fault diagnosis, +approved/rejected action outcomes and independent recovery remain open, as do +the supported autonomous provider dependency and independent-environment gate. + + +### Ordinary Assistant storage fault and recovery, 2026-09-06 + +The command-connection correction passed the exact nine-file worker pre-commit +and was pushed as `f5f440dbad18d83557104d2cf6197d8319949e44` in PR #1935. +Required CI remains in progress. This is not a release or a completed goal. + +An owned DockerLab run used the checked-in storage-pressure manifest on Tower. +Independent observations established a healthy worker with 8,347,648 free bytes, +then a real ENOSPC fault with zero free bytes, a running/unhealthy worker and a +healthy control. Pulse collection converged to both states before the request. +The ordinary read-only Assistant used the configured subscription Opus 5 route. +It was not an autonomous Patrol request and did not retry the cached refusal. + +The diagnosis took 114.847 seconds and ten tool calls. It identified the worker's +unhealthy state, the control's current healthy state and a failed command-route +log read. It did not retry that unavailable capability. However, it falsely +ruled out resource pressure using low CPU, memory, network and disk-read values. +Filesystem capacity was absent from its evidence and was independently full. +This is a failed diagnosis, despite its otherwise useful uncertainty statement +and suggested diagnostic read. No model-directed mutation occurred. + +After the answer, the independent oracle still found zero available bytes and +an unhealthy worker. Removing only the owned fill file restored 8,220,672 bytes +and healthy status. Pulse collected recovery and resolved the health alert. +A follow-up in the same Assistant session took 91.358 seconds and five new reads. +It correctly identified current recovery and the resolved alert, distinguished +symptom recovery from an unknown cause and did not invent an intervention. +It overstated continuous control health and non-impact from sparse observations. +The recovery assessment is partial, not a complete incident explanation. +An independent post-answer check confirmed healthy worker/control, no container restart +and 7,639,040 free bytes. Both cleanup passes passed, with no second-pass work +and unchanged unrelated inventory. The disposable resources are removed. + +Playwright exercised `/patrol` and Assistant at 1440x1000 and 390x1000, all fifteen +retained tool input/output pairs, keyboard expansion/collapse, deepest output +scrolling, complete answers, Escape, reload and the same retained conversation. +Rendered inputs/outputs match persisted records. Pixel inspection covered both +answers and the failed-access result on desktop and mobile. Runtime and source +hashes remained unchanged across both requests, with binary +`bcaf748107211ee733a6dc0f4d17220d9b4d1ce1918c25bde27cf3d10c0d6379`. +The route warning was an artifact of blocking non-chat POSTs in the proof browser. +The original autonomous refusal timestamp remained `2026-09-05T19:46:39Z`. +Private fixtures, source bindings, observations, screenshots and assessments +are at workspace-relative `tmp/patrol-storage-fault-case/`. + +### Next canonical correction: tmpfs inventory + +Before implementation, source and native inspection establish a collection gap: +Docker reports the owned scratch mount in `HostConfig.Tmpfs`, while `Mounts` is +empty. `internal/dockeragent/collect.go` copies only `Mounts`, so shared resource +queries falsely present an empty mount inventory. Preserve these native tmpfs +entries through the existing report mount type. Keep destination, type and +reported options, derive read/write from those options, preserve authoritative +existing mount records and deterministic ordering. Do not infer used/free space +from a configured size. No enrollment, permission or production agent change is +part of this collection correction. + +Proof plan: reproduce the captured tmpfs-only inspect shape through the actual +collector, then cover existing mounts, overlapping representations, read-only +options and absent host configuration. Run targeted/full collector checks on +pulse-dev and verify the report through the existing shared projection. Browser +proof after the final change must exercise mount evidence in resource details +and Assistant tool results, desktop and mobile, including deepest expansion and +reload. A captured-result rendering check is not installed-agent or model +qualification. Leave those limits explicit until the new collector is exercised +through a supported installed path. + +The incident lookup also needs an identity audit: the canonical container ID +returned no incident recording while the observed health alert used its legacy +Docker resource ID. This is a concrete lookup discrepancy to investigate, not +yet proof that a recording exists. Model inference from unmeasured capacity and +sparse health history remains an open quality failure. Supported autonomous +provider, approved/rejected actions and independent environments remain open. + +Further shared-projection inspection before editing found that `MountInfo` drops +native mount type/options and that canonical app-container queries derive write +access from equality with the single string `ro`, misreporting compound read-only +options. The same slice must preserve type/options and the canonical `RW` boolean +through both canonical-provider and typed read-state query paths. Add a query +regression and capture its actual output for final Assistant browser proof. +This remains mount configuration evidence, not measured filesystem capacity. + + +### Tmpfs collection and query contract proof + +The captured tmpfs-only and mixed-mount regressions failed against the previous +collector, then passed after the collection correction. The full dockeragent +package passed in 18.883s and focused race proof in 1.030s. Existing monitor report +mount propagation and discovery mount regressions passed. Both query paths +failed because type/options were lost, then passed after projection correction. +The full tools package passed in 59.473s and focused race proof in 1.030s. +A final output-only capture rerun passed in 0.013s. All proof used Go1.26.8 and +GOMAXPROCS4 on pulse-dev. The final Pro cross-build passed and the installed +local binary SHA256 is +`0a21dca4106c2ddc6873a3aca3b23378dccef35383ca00d7e9292b966de7c200`. +The managed local backend restarted and `/api/health` reported healthy. +No production collector was replaced. + +Final Playwright proof used captured canonical resources at `/docker`, widths +1920, 1440, 900 and 390 with height 1080. It exercised mount summary/title, +keyboard row expansion/collapse, mobile row tapping, adjacent detail state, +mount-destination search, Escape and reloaded search state. The existing wide +mount column is truncated with a full title. Responsive details have no dedicated +mount section. This is an existing presentation limitation, not full mobile +mount inspection qualification. Assistant's complete mount evidence is readable +at `/patrol`, 1440x1000, 900x1000 and 390x1000. Both actual query projections +passed exact input/output comparison, hover/focus, keyboard expansion/collapse, +deepest scrolling, controlled session reload and reopening retained records. +Pixels were inspected on desktop and mobile. Source/binary hashes stayed fixed. +The original autonomous refusal timestamp is unchanged. + +These browser fixtures qualify rendering of the corrected shared fields. They +do not qualify an installed collector, actual model interpretation of tmpfs +configuration, or durable backend persistence of those fixture sessions. The +ordinary live diagnosis/recovery records above have real server persistence +and retain their failed/partial judgments. Exact scoped hook and landing remain +required. Required model/action qualification and independent environments are +still open. The typed compatibility get path also does not accept the canonical +ID returned by its list path, so its mount regression uses an existing accepted +name. That identity residual is recorded for modernization, not silently fixed +through this mount projection. + +## Canonical incident history, 2026-09-06 + +The tmpfs correction passed the exact worker hook and was pushed as +`6e18777d30f30b498def39d30016a697cabc4ea7` in PR #1935. That scoped +delivery does not change the failed storage diagnosis or partial recovery verdict. + +The incident audit found a source-of-truth mismatch, not evidence that an existing +recording merely needed an ID alias. The legacy five-second recorder has no +production alert callback connected to its coordinator. It samples cached values +using recorder time without preserving their source measurement time. Connecting +that recorder would not supply trustworthy higher-frequency history. + +The canonical resource timeline already stores observed changes, alert lifecycle +events and executed actions. Assistant handoffs use a bounded excerpt of this +same store. The shared `pulse_knowledge` incidents action now reads that +organization-pinned timeline directly, using the supplied canonical resource ID. +It does not require the resource still to exist in current inventory, infer +identity from names, include related resources implicitly, or reconstruct events +from current metrics. The response preserves canonical source, observation and +optional occurrence timestamps, state transitions and metadata. `since` filters +on observation time, and bounded results report `has_more`. Empty retained history +does not establish health. Missing or failed storage is a failed read. + +Explicit legacy `window_id` lookups remain isolated archive reads, must match the +requested resource, and explain that sample timestamps do not establish source +freshness. The primary incidents action no longer uses those recordings. The +legacy recorder/coordinator startup and API active-count plumbing still exist. +Their retirement is a separate cleanup in this redesign and must preserve any +saved archives. Do not connect them as a replacement incident truth source. + +Qualification uses the real SQLite resource store with a fired/resolved lifecycle, +an older excluded record, a related-resource negative control, absent occurrence +time, truncation and empty history. Unavailable/failed storage, invalid input and +archive resource isolation are separate negative controls. Captured actual tool +responses must pass the Assistant expansion, scrolling and reload matrix at +`/patrol`, 1440x1000, 900x1000 and 390x1000. This is contract and rendering proof, +not a new real-model or continuous-coverage claim. + +Read-only API inspection of the actual removed storage fixture confirmed a +remaining canonical write-boundary defect. The canonical app-container timeline +returns its creation and removal, while the fired event at 13:14:28.59485Z and +resolved event at 13:17:58.634095Z remain under its legacy Docker resource ID. +The resource API includes related network changes by design. The new tool uses +direct resource history only. `recordAlertTimelineChange` passes the alert's +source ID directly to `BuildAlertTimelineChange`, and `MonitorAdapter.RecordChange` +forwards it without canonical resolution. Consequently this read-path change is +only partial incident-history remediation. The next required owning fix must +resolve event identity before persistence and preserve access to retained prior +identity records, including removed resources, through the shared identity/history +contract. It must not add a Docker string rewrite inside the Assistant tool. +The shared writer and retained-identity correction remain required in this goal. +Private raw API receipts are in `tmp/patrol-canonical-history/live-timeline.json` +and `live-legacy-alert-timeline.json`. No model call or infrastructure mutation +was made during these reads. + +The history regression passes through the registered tool dispatcher. The full +tools package passed in 59.856s, the final focused capture passed in 0.036s, and the +focused race check passed in 1.189s on pulse-dev with Go1.26.8 and GOMAXPROCS4. +The final Pro build passed and was installed into the local development stack. +Its SHA256 is `4929aeb869db54122bc5525352d3126c0e9fa7c4847e3fdfa741500600b5e00d`. +The managed backend restarted healthy. Final Playwright proof passed all five +registered-tool cases at `/patrol`, 1440x1000, 900x1000 and 390x1000, including +hover/focus, Enter/Space, deepest output scrolling, Escape and controlled session +reload with exact input/output and success/failure comparison. Root inspected +actual pixels at all three widths. Source and binary hashes remained fixed. +The provider warning stayed visible and no retry, route switch or provider +request was made. This is captured-response rendering, not real-model diagnosis +or server persistence qualification. The exact scoped worker hook gates delivery +through PR #1935. + + +## Shared Docker history identity, 2026-09-06 + +The preceding incident-read correction was pushed as +`580a246981c76b401e9007f5ac65c89355b64c6d` in PR #1935. Its live retained +records established the identity split addressed here. + +The canonical fix belongs to the shared monitor/store boundary. Exact full +Docker container references resolve through current registry identity. Retained +bindings survive inventory removal, and deterministic source-specific identities +allow legacy records to be found after restart. Names and abbreviated IDs are +not sufficient evidence. A small organization-scoped history alias index joins +readable records without rewriting event IDs, timestamps or metadata. Existing +canonical succession machinery was deliberately not used for these aliases +because it also moves operator state and action indexes. History matching must +not transfer authority. The alias index follows journal retention and separate +store connections read fresh bindings. + +Focused regression and race proofs passed on pulse-dev with Go1.26.8 and +GOMAXPROCS4. Full unifiedresources, monitoring and tools packages passed in +37.826s, 79.866s and 59.584s. They cover real alert-manager callbacks, recovery +after inventory removal, restart, replay, same-name controls, tenant isolation, +unchanged operator/approval records and registered Assistant tool reads. Scoped +history lookup measured 0.261–0.275ms with one alias and 0.317–0.336ms with +20,000 unrelated aliases, at 6,280 bytes and 94 allocations per read. These are +worker microbenchmarks, not fleet or frontend performance qualification. + +The verified worker Pro binary has SHA256 +`bb6d1508a5b4d23943c37dfc42198f132c0139805dcd1891ee18aca0a9f9dd54`. +It was installed into the local development stack and restarted healthy. Both +canonical and legacy timeline API queries now return the same seven retained +records for the removed storage fixture, including the original fired/resolved +records with exact unchanged content. The complete registered-tool rendering +matrix passed at `/patrol`, 1440x1000, 900x1000 and 390x1000. Six captured cases +include migrated history, bounded and empty results, and unavailable/failed +reads. Hover/focus, Enter/Space, deepest scrolling, Escape and controlled-session +reload preserved exact inputs, outputs and completed/failed states. Pixels were +inspected at all three widths. Source and binary hashes remained fixed. Private +receipts are `tmp/patrol-history-identity/browser/receipt.json` and +`live-history-proof.json`. Controlled session responses qualify rendering, not +server persistence or model diagnosis. The cached provider refusal remained +unchanged. No model request or infrastructure fault was made. The exact scoped +worker hook remains the delivery gate for PR #1935. + +This history correction does not qualify the failed ordinary storage diagnosis, +partial recovery claim, installed tmpfs collector, autonomous provider, approved +and rejected action outcomes, or independent Pro environments. The unused legacy +recorder/coordinator still needs retirement with its archives preserved. + + +The history identity change passed the exact thirteen-file worker hook and was +pushed as `919331d5b3f6076f8616b07eb8e9ca611f26ee52` in PR #1935. Integration +with main `11a8cc2180aae886ec7f92e2333002b57cf1b9a3` preserves both sides of +three additive subsystem-contract conflicts. The host-ingestion auto-merge +retains Docker observation corrections alongside incoming host-link provenance. +Unrelated registry indentation was restored without changing its decoded data. + +The combined monitoring, unifiedresources, tools, config and models packages +passed in 85.894s, 41.931s, 59.609s, 18.381s and 0.098s. Focused history, +Assistant, host-link and lifecycle race checks passed. Frontend type checking +and four alert suites passed all 42 tests. The incoming delivery-log component +browser proof passed at 1440x900 and 390x900, including reordered success/failure, +held events, pending state and newest failure. Pixels were inspected. This is +scripted component proof, not installed notification delivery. + +The final merged-source Pro binary has SHA256 +`234f625cb74be3d300facfed1bb17b17e20c44c06037a1f8f4fffc1c4f49d621`. +Its local restart was healthy. The complete six-case Assistant matrix was +repeated at `/patrol`, 1440x1000, 900x1000 and 390x1000, with exact tool records, +keyboard expansion/collapse, deep scrolling and controlled-session reload. +Pixels and source/binary bindings were checked after this final build. Both +actual removed-container timeline queries still return the same seven retained +records with the original fired/resolved content. No model request, route +switch, infrastructure fault or production collector replacement was made. +The autonomous provider refusal remains enforced. Integration receipts are in +`tmp/patrol-history-integration/`. The full integration hook gates its merge +commit and push. All previously recorded model and wider-readiness gaps remain +open. + +## Legacy recorder retirement, 2026-09-06 + +The disconnected incident coordinator, five-second cached-metrics sampler, +pre-incident buffers, archive writer and unused adapters are removed. They had +no production alert trigger. Canonical resource history remains the primary +incident evidence for Assistant. No replacement diagnosis or scheduling policy +was added. + +Explicit archive lookup now requires exact organization, resource and window +binding. The reader is lazy and read-only. Saved file contents, modification +time, mode, old observations, metadata and summary values survive reads. Missing +archives, malformed files and missing windows remain distinct outcomes. Legacy +`recording` status is historical, and the response discloses that the old +`summary.duration_ms` field contains nanoseconds. An old file is never rewritten +to make its evidence appear current. + +The incidents API now reports `active_count: null` with +`active_count_status: not_measured`. The retired coordinator's empty map never +established a measured zero. Its legacy incident-memory listing still needs a +canonical query design covering aliases, canonical-only events, honest bounds +and propagated projection-read errors. This is recorded as an open modernization +residual rather than treating that listing as complete. + +Final-source registered archive-tool receipts pass Playwright at `/patrol`, +1440x1000, 900x1000 and 390x1000. The five cases cover a saved observation, +unavailable/malformed archives, the wrong resource and a missing window. +Verification includes hover/focus, Enter expansion, exact tool input/output, +deep scrolling, Space collapse, Escape, reload and controlled session reopening. +Actual pixels were inspected. Incoming main alert dispatch wording also passes +its isolated real Overview browser script at all three widths. These controlled +responses prove rendering, not model diagnosis, installed delivery or server +persistence. + +The final worker Pro binary is +`bd29e6f27be7b3ad4cfbc37842f4da90f08c6a48c9fc23b12c9c597b67346c9b`. +After the managed local restart, canonical and legacy queries still return the +same seven retained homelab records with the original fired/resolved events +unchanged. The live incidents API reports an unmeasured count. Cached provider +refusal remains enforced. No model request, paid spend, provider retry, +production collector change or fault injection occurred in this slice. + +Archive, tools, chat, AI runtime and targeted API/race checks passed on the +worker. One full API run as root invalidated its mode-bit persistence-failure +fixture. That fixture passes unchanged under the normal worker account. The +full API rerun passes under that account (286.960s), as do the incoming +startup-replay and legacy-boundary source checks. Frontend type checks and all +29 incoming alert tests pass. The exact staged hook gates landing. +Private receipts are under `tmp/patrol-archive-retirement/` in the workspace. + +### Live collector storage evidence, 2026-09-06 + +`TestCollectContainerStorageFaultLive` calls the production Docker client and +`collectContainer` implementation against the existing storage-pressure lab. +The opt-in command, run on the worker beside its Docker Unix socket, is: + +```sh +PULSE_QUALIFY_ORACLE_DOCKER_CONTEXT=default go test ./internal/dockeragent -run '^TestCollectContainerStorageFaultLive$' -count=1 -timeout=240s -v +``` + +The final test passed in 8.694s on Docker 29.8.0 against the runtime source tree +of `186ce504c8f0fa6e0174f3b10f3d99f6278f10fc`. Its SHA256 is +`cdf303f4de23c020690c81b0b57190c728319788cb0aaff906b0d4d9d864541e`. +Independent filesystem observations measured 8,380,416 available bytes before +the fault, zero during it and 8,380,416 after recovery. The service stayed +running. Collected and JSON-decoded health followed healthy, unhealthy, healthy, +while the unrelated control remained healthy. Docker returned zero native +`Mounts` throughout. The report retained the single tmpfs destination, type, +8 MiB configuration options and writable setting from `HostConfig.Tmpfs`. +Configured size is not a measured capacity counter in the report. + +The test targets exact run-owned container IDs and labels. Existing fixture +cleanup passed, its second cleanup was a no-op and inventory matched the +pre-run snapshot. The test skips without explicit opt-in. All four existing +mount regression cases also pass. Private logs are in +`tmp/patrol-storage-collector-live/` in the workspace. + +This proves live collection and report serialization only. It does not send a +report to Pulse, enroll or replace a production agent, call a provider, or +qualify approval, execution, diagnosis or model recovery. No runtime or frontend +source changed in this slice, so no new browser claim is made. Prior storage +diagnosis failures and the cached autonomous-provider refusal remain open. + + +## Funded Gemini qualification, 2026-09-06 + +The maintainer authorized `openrouter:google/gemini-3.8-flash` with a provider-side +US$5 key limit expiring on 2026-09-07. The provider key endpoint confirmed both +constraints. Credentials remain in runtime configuration, not these receipts. +Synthetic readiness passed in 9.034 seconds: three streaming tool scenarios, +two context fixtures and multi-turn continuation. This supports the readiness +claims for Watch only and Ask first. It does not qualify autonomous fixes. + +The first unhealthy-container run, `q-20260906-172952-3cccbf34`, detected the +correct fault and left the healthy control alone, but failed overall. The exact +Gemini route had no price entry, and the model attempted unsupported Docker +configuration access. The shared price table now records the reviewed standard +rates of US$0.75 input and US$3.75 output per million tokens for direct Gemini +and OpenRouter. Variant routes remain unknown. These introductory rates must be +reviewed on 2027-01-01. The query capability description now explicitly names +TrueNAS as the supported app-container configuration adapter and directs Docker +collected health/mount/port/network reads to `get`. Runtime permissions and +qualification gates are unchanged. + +The following runs used the worker-built Pro binary +`74464e75977caf55cda092c8cf56c24967616c8c24d770872fbea5d86e31a1dc`, +core base `b0b39f00dc6685ad9ed63e8a6e91b954338073e4` plus the pricing and +capability-description changes, and canonical enterprise base +`3d9f4e3051d38027355a2a1f36b8c7f672a09b65`. The worker archive commit +`d9cb84e15d1acc377341129bdda5c28176e7128c` has identical contents for all +87 tracked enterprise files. The existing runner created disposable +resources on Tower, waited for normal collection, and used independent fault, +recovery and cleanup oracles. + +| Case / run | Result | Evidence | +|---|---|---| +| Unhealthy, `q-20260906-174546-a7a9810b` | Pass | 9.709s detection phase, two tools, no failed/duplicate calls, healthy sibling unflagged. | +| Unhealthy, `q-20260906-174708-81f8d655` | Pass | 10.541s detection phase, exact unhealthy resource found. | +| Unhealthy, `q-20260906-174758-14deaa15` | Pass | 9.395s detection phase, exact unhealthy resource found. | +| Unhealthy, `q-20260906-174853-71cf894f` | Pass | 25.673s detection phase, exact unhealthy resource found. | +| Healthy mixed, `q-20260906-180144-381874a6` | Pass | 5.200s detection phase, no false findings. | +| Dependency, `q-20260906-175058-de5e350d` | Pass | Starting from only the client symptom, identified the stopped dependency and affected client. Investigation completed in 18.970s with three evidence calls and no mutation. | +| Storage, `q-20260906-175232-3ce6fbfa` | Fail before inference | Normal collection never converged to the required resource projection. No model diagnosis was attempted. | +| Approved restart, `q-20260906-175812-0593b9b7` | Fail before approval | Detection and investigation completed, but no exact action reference existed. The broker refused because Tower's Docker command agent was disconnected. Nothing executed. | +| Rejected restart, `q-20260906-175940-bff6992a` | Fail before rejection | No exact action was available to reject. This does not qualify rejected-action handling. | + +Every listed run passed cleanup, including second-cleanup no-op and unchanged +inventory. Individual Watch run estimates were about US$0.007 to US$0.014. +Those scorecard estimates cover the Patrol detection phase, not the separate +investigation calls. Provider-side aggregate spend is the budget authority for +this temporary key. The fixed route price does not turn an estimate into a +reconciled bill or establish a hard Pulse budget for unpriced history. + +Live qualification exposed two additional shared contract defects. The +investigation orchestrator logged action-broker refusal but completed the +record without retaining the error, leaving the model's captured-proposal prose +visible without the later refusal. The current enterprise change retains the +original diagnosis, persists the broker refusal as a failed investigation with +`needs_attention`, and creates no action reference. The product history adapter +also projected result-bearing transcript calls back into provider request calls, +dropping observed output and success/failure. The current core change uses one +shared transcript type for stored chat and product history, preserving the +separate explicit provider projection. + +The live review exposed duplicate detail IDs, duplicate unformatted conclusions, +paused history made unclickable by the scheduling switch, and narrow filter/sort +overlap. The shared finding/investigation surfaces now preserve one detail target, +render sanitized Markdown once for identical summaries, retain distinct summaries, +keep history available while paused, and wrap controls. Result-bearing tool calls +use the same expandable evidence component as Assistant. Historical calls without +a result status retain their evidence without invented success or failure. An +investigation outcome of `cannot_fix` or `needs_attention` does not identify who +resolved the finding, so the shared resolution copy no longer infers manual review. + +Private run receipts and source/binary bindings are under +`tmp/patrol-gemini-38/` in the workspace. The original failed runs remain failed. +Installed storage collection and a temporary command-enabled lab agent remain +prerequisites for real storage and approved/rejected recovery qualification. +No production agent has been replaced. Full action outcomes, remaining backup +coverage and independent volunteered Pro environments remain open. + + +### Final refusal and evidence-retention proof + +Two further approved-remediation attempts remain **failed**: +`q-20260906-181839-6cbdc711` and `q-20260906-182952-6acfb739`. Both retained the +broker error separately from the original model summary, saved `status=failed` +and `outcome=needs_attention`, and created no action reference. Both passed +cleanup. The final run used Pro binary SHA256 +`24de8c9ea0020067d298c489f0d99a272b5ec4b00afab7a40dff55aabb244061`, +including the proposal-response clarification, and detected its exact unhealthy +container with no false positives. Its saved investigation is +`48a16b05-50f0-4605-847c-0a71b3435975` for finding `ca3af29ac54d540f`. +The original failed scorecards have not been reclassified as action passes. + +The restored history API retains observed outputs and explicit `success=false` +for historical `pulse_read` failures. Live browser review at `/patrol` exercises +successful query output, both `ACTION_NOT_ALLOWED` and `NO_AGENT` failures, +original diagnosis, one broker error, paused history and review focus return. +The settings proof at `/settings/pulse-intelligence/patrol` checks the exact +model, reviewed rates, synthetic readiness limits and reload. The current +source-bound browser receipt records desktop, intermediate and mobile results. +GET response fixtures cover unknown historical result status only, without +claiming new persisted model evidence or action execution. + +Remaining qualification requires a current installed collector and a temporary +command-enabled lab agent. A Linux amd64 agent has been built on the worker, +SHA256 `ae2ed8b97709ec6e71af979c293ca9d3634662767ec1b59629baf4933c90cf5d`, +without installing it or changing Tower credentials. Tower's separate production +reporting agent is untouched. Any agent enrollment must use the canonical scoped +installation flow, preserve explicit identity and revoke temporary execution +access after qualification. Detection success does not satisfy this prerequisite +or the remaining backup and independent-environment cases. + + +## Installed agent and governed action qualification, 2026-09-06 + +The maintainer explicitly approved a temporary update and scoped command token +for Tower's separate development agent, followed by restoration. The installed +agent artifact was `ae2ed8b97709ec6e71af979c293ca9d3634662767ec1b59629baf4933c90cf5d`. +Both host and Docker modules reported running, and the command connection +registered the same agent identity. The production agent retained PID 752388. +The tests used Ask first with manual triggers. Scheduled Patrol ended paused in +Watch only. No autonomous-mode qualification is claimed. + +The first installed storage attempt, `q-20260906-192435-f0d7eebf`, failed before +inference. Inspection established a qualification-client pagination defect: +`/api/resources?limit=1000` returned a maximum of 100 records from an inventory +of 104, leaving the exact worker on page two. This was not an absence of normal +collection. The client now follows the API pages and rejects partial results +when a later page fails. A regression finds an unhealthy resource beyond the +first 100, and the complete qualification package passes. Fault oracles and +score thresholds were not weakened. The corrected runner hash is +`a27f9ab0bf4786b670db3e8a989e974c8c43c5084d9524474ee694735d2e7df9`. + +| Case / run | Automated result | Reviewed outcome | +|---|---|---| +| Approved restart, `q-20260906-193013-12d25545` | Pass | Correct unhealthy container, exact finding/investigation/resource and plan-hash binding, explicit approval before execution, completed restart, independent healthy/running readback and lifecycle verification. Detection 8.893s, fault-to-remediation phase total 78.490s. | +| Rejected restart, `q-20260906-193207-7095dad5` | Pass | Exact plan rejected, no restart, independent unchanged unhealthy fault until teardown. Detection 10.838s, fault-to-decision phase total 37.795s. | +| Storage, `q-20260906-193613-556ef23d` | Pass from existing scorecard | **Fails semantic diagnosis review.** Collection converged and logs exposed ENOSPC, but the model incorrectly asserted that Tower was out of disk space and implicated its array. Only an 8 MiB container tmpfs was exhausted. | + +The approved action is `act_dcc3b52e5451810e49466daf9a6fccb0`, linked to finding +`566515f71129ce73` and investigation `aae05717-c9b0-4aac-8439-ba78f46c28e9`. +The rejected action is `act_ee0b736f0430e472e896a456ba3cb6eb`, linked to finding +`a17940552206e1ca` and investigation `955f4f47-0b6f-404b-b513-70a1d226f111`. +Each case passed independent teardown, second-cleanup no-op and restored +inventory. Per-case detection estimates were $0.012123, $0.012283 and $0.014915. +These exclude investigation calls and are not provider-account spend. + +Storage remains unqualified. The model had the collected tmpfs mount and its +configured size. Its canonical-resource log call failed, the fallback host and +container log call succeeded, and its `df -h` command required approval. It then +promoted unrelated host/array warnings into a definite capacity diagnosis. +The scorecard's required terms and narrow forbidden phrases missed that false +claim. Its raw pass is retained as evidence of a qualification limitation, not +accepted as product success. The next storage slice needs canonical, authorized +filesystem-capacity evidence and explicit semantic review against the bounded +fault. Do not permit arbitrary commands merely to make that case pass, add a +benchmark-specific diagnosis rule, or treat identifier/phrase matches as proof +of causal correctness. Backup coverage and independent Pro environments remain +unqualified as well. + +Real outcome review exposed stale durable records: the finding and investigation +could say `fix_verified` while the embedded product record still said +`fix_queued`. Action reconciliation now refreshes that record through the same +canonical builder used at investigation completion, preserving original model +prose, evidence and retained rollback. Read-time hydration repairs existing +records even when the top-level outcome already matches. Unchanged hydration +must not republish state or repeat outcome notifications. Resolved findings keep +their exact action-history link, and Assistant handoff preserves resolved status. +Investigation completion replaces an earlier partial action projection with its +final evidence. Subsequent action transitions preserve that completed evidence, +including impact and confidence that the current finding may no longer retain. +An intermediate proof build exposed that loss of retained impact. The regression +now preserves it, while already absent historical fields remain unassessed. +Final runtime and browser verification of these corrections is recorded below. + +After qualification, both original development binaries were restored separately +because they differed: runtime `e5a2b60e52e35c37f68daa348c642757b56a64b40ca1d72f4db843ce69eb5db4`, +persistent `73c224dfd750c41b2cbd883c3ce7e352071862bc6a60e59fe6ed4dd3de312bc6`. +The original protected token was restored, temporary issued tokens were revoked +and checked absent, temporary backups were removed after comparison, and no +owned fault containers remained. The restored v6.2.0-rc.8 development agent +reported fresh telemetry. Its original token lacks command scope, so its command +connection is again absent by design. Production PID 752388 remained unchanged. + +Final action-history proof uses Pro Darwin arm64 binary +`859d5d2de84cfd2264caa7dbcf5f080e3b1d06b5876df2c81dc0e00872f5e779`. +Worker proof passes the API action reconciliation selection and investigation, +record, rollback and early-projection completion regressions. The full API suite +passed in 310.353s before the final evidence-preservation refinement, followed +by the final targeted regressions. The three affected action component suites +pass 30 tests. The complete qualification package and pagination regressions +also pass. The exact final staged hook gates landing. + +Playwright exercises `/patrol` Activity/All and both exact `/actions?action=...` +links above at 1440, 900 and 390 by 1000. Final-content checks cover resolved +record retention, outcome agreement, safety disclosure, completed/rejected +headers, planning-time copy, absent settled execution controls, independent +verification, policy/evidence/delivery disclosures, keyboard toggles, Escape, +close controls, deep-link reload, scroll fit and retained review focus. Actual +pixels were inspected at desktop, intermediate and phone sizes. Assistant +handoff opens the same finding with completed/rejected context and read-only +control. Provider readiness POST was deliberately blocked during that rendering +proof and no prompt was sent. Earlier browser attempts encountered an +intermittent bootstrap connection screen. The complete final matrix passed +after removing redundant immediate navigations from the proof driver, without +claiming a bootstrap fix. Source bindings are in +`frontend-modern/browser-verification.json`. diff --git a/docs/release-control/v6/internal/status.json b/docs/release-control/v6/internal/status.json index bf404a895..357352464 100644 --- a/docs/release-control/v6/internal/status.json +++ b/docs/release-control/v6/internal/status.json @@ -10201,7 +10201,7 @@ }, { "id": "patrol-assistant-customer-outcome-qualification", - "summary": "The explicit Patrol/Assistant redesign goal, contract, execution plan and source-bound evidence remain in docs/qualification/PATROL_ASSISTANT_CUSTOMER_JOURNEY.md. Model judgment owns diagnosis. Observations, hypotheses, proposals, executions and independently verified outcomes remain distinct. Assistant continues the same issue and governed action records. The recorded 2026-09-05 baseline has 127 paid installations, 71 with Patrol enabled and 23 with Assistant calls. Fourteen verified resolutions came from one installation. Schema 17 outcome/provider/cost fields had no adoption. Usage does not prove useful linked tasks or representative false-alarm, missed-problem or success rates. Shared risk, provenance, history, missing-access and diagnostic-continuity corrections have regression and named browser proof. Proposal promotion, duplicate causal inference, contextless evaluations and count-based diagnostic completion policy were removed. Independent Docker fault/oracle contracts qualify reproducible injection, negative controls and cleanup, not model competence or governed action outcomes. Integrated CI exposed retained-query performance, disk-probe ordering and route-label timing regressions. Their scoped corrections and qualification records landed through PR1928 and PR1929. PR1929 merged at cf98358c0eb46987a82def5776fa41db5f54210a with backend, frontend, benchmark, governance, CodeQL and all eight Core E2E shards passing. The current ordinary retained-history diagnosis still contradicts explicit temporal semantics and remains unqualified. Two additional read-only Assistant requests used claude-subscription:claude-opus-5 against run-owned containers on the monitored Tower host. The healthy request took 82.835s and seven tools, correctly recommending no action, but overstated absence of storage impact. The dependency request took 204.384s and sixteen tools with three failed reads. It identified the stopped dependency, preserved the missing command access and causal uncertainty, but overstated storage exclusion and recovery implications. Config reads incorrectly reported app-container not found after successful canonical gets. These single cases remain partial diagnosis evidence, not a qualification pass. The fixture deadline performed two-pass cleanup with unchanged original inventory. Post-answer fault readback and explicit recovery were not completed, so no action outcome is claimed. Captured responses exposed concurrent tool-ID merging, sibling approval removal and renderer mutation of shared evidence. The current scoped correction keeps supplied invocation IDs authoritative and stable message rows without deep transcript reconciliation. All 167 affected frontend tests pass. Final-source browser replay at /patrol, 1440/900/390x1000, preserves all seven and sixteen exact tool inputs/outputs. Controlled stream states verify concurrent progress, cancellation, failed completion and sibling approval retention without provider or infrastructure actions. Exact captures, failed reproductions, hashes and remaining limits are in the plan. This correction still requires its own scoped landing checks. Claude Max explicitly refused autonomous Patrol readiness. Cached refusal and API409 enforcement remain intact, with no bypass or repeated retry. Ordinary Assistant is not autonomous qualification. Approval for an alternate separately billed provider remains pending, and no paid request occurred. Reliable interpretation, the config-read contract, broader storage/backup and approved/rejected action outcomes remain required local work. Independent volunteered Pro environments remain a separate wider-readiness gate.", + "summary": "The redesign goal remains open. The plan, historical receipts and exact source bindings are in docs/qualification/PATROL_ASSISTANT_CUSTOMER_JOURNEY.md. Patrol owns investigation and Assistant continues the same issue. Observations, hypotheses, accepted proposals, execution and independently verified outcomes remain distinct. The recorded baseline of 127 paid installations, 71 Patrol-enabled, 23 with Assistant calls and fourteen verified resolutions from one installation does not establish representative customer success. Schema17 outcome/provider/cost fields had no adoption. Shared evidence/history/risk and removal of proposal-as-proof and proxy completion policy landed through PR1928/1929. PR1934 merged canonical tool/transcript identity. PR1935 contains subsequent canonical history, measurement-presence, command-connectivity, tmpfs context, exact Gemini pricing, retained broker errors and result-bearing transcript corrections. Enterprise broker refusal handling merged in PR22. Real Gemini Watch, healthy-control and client-to-dependency cases passed. With explicit authority for a temporary current development agent and scoped token, approved restart q-20260906-193013-12d25545 passed exact plan/origin binding, explicit approval, execution and independent recovery. Rejected restart q-20260906-193207-7095dad5 passed exact rejection and independent non-execution. All test resources were removed, original development binaries/token restored, temporary tokens revoked, scheduled Patrol paused in monitor mode and production agent PID preserved. Storage collection failure was traced to qualification pagination beyond the API page cap of 100 and corrected with full-package regression proof. Installed storage q-20260906-193613-556ef23d passed the existing scorecard but FAILED semantic diagnosis review: the model falsely attributed an exhausted container tmpfs to Tower/array capacity despite available mount configuration. A capacity read required approval. Storage remains unqualified, and lexical/identifier scoring must not be treated as causal correctness. Next storage work needs canonical authorized filesystem evidence and independent semantic review without benchmark-specific diagnosis rules or weakened command approval. Real browser review additionally exposed an embedded investigation record left fix_queued after verified recovery and hidden action history on resolved findings. Current action reconciliation refreshes the durable record through the canonical builder, preserves prose/evidence/rollback, repairs missed transitions without duplicate publication, and retains completed action history and resolved Assistant context. Final worker regressions and source-bound runtime/browser proof pass at 1440, 900 and 390 widths, including completed/rejected history and read-only Assistant handoff. The exact staged hook and PR1935 landing gate integration. Other residuals include canonical incident-memory listing/aliases and failed-read propagation, unsupported filters, typed compatibility ID lookup, legacy direction availability, Docker-host history, responsive mount details, backup coverage and broader model qualification. Independent volunteered Pro environments remain a wider-readiness gate. Autonomous modes remain unqualified. The earlier Claude refusal was not retried.", "owner": "project-owner", "status": "planned", "recorded_at": "2026-09-05", @@ -10215,6 +10215,7 @@ "ai-runtime", "api-contracts", "frontend-primitives", + "monitoring", "patrol-intelligence", "performance-and-scalability", "unified-resources" @@ -10386,6 +10387,7 @@ "ai-runtime", "api-contracts", "frontend-primitives", + "monitoring", "patrol-intelligence", "performance-and-scalability", "unified-resources" diff --git a/docs/release-control/v6/internal/subsystems/agent-lifecycle.md b/docs/release-control/v6/internal/subsystems/agent-lifecycle.md index 1f47be0f2..2567589c5 100644 --- a/docs/release-control/v6/internal/subsystems/agent-lifecycle.md +++ b/docs/release-control/v6/internal/subsystems/agent-lifecycle.md @@ -15,6 +15,17 @@ ## Purpose +Docker mount reports include tmpfs configuration from `HostConfig.Tmpfs` +through the existing optional mount array. This adds collection evidence only. +It does not change admission, enrollment, execution permissions or agent +lifecycle authority. Existing agents continue to report their existing mount +coverage. Deploying an updated collector is a separate installed-path proof. + +Docker block-I/O report presence fields are optional measurement metadata. +They preserve zero and omitted directions independently without changing report +admission, enrollment, identity, command permission or agent lifecycle state. +Older agents remain accepted, with ambiguous omitted zero counters unavailable. + ### Automatic PVE association identity boundary Host ingestion must not create a host-to-PVE or reciprocal PVE-to-agent link @@ -7896,3 +7907,9 @@ positive matching evidence, provider replacement/return, write failure and automatic versus unknown-provenance cleanup. State and config tests cover atomic replacement and preservation of lifecycle evidence. These are synthetic local proofs, not reporter confirmation or installed-release resolution of #1930. + +Historical incident archives are explicit reads, not a collector lifecycle. +`internal/api/router.go` no longer starts an incident coordinator or a cached +metrics sampling loop. Organization teardown drops the archive reference without +saving or deleting recordings. Alert observation and recovery continue through +the alert manager and canonical resource timeline. diff --git a/docs/release-control/v6/internal/subsystems/ai-runtime.md b/docs/release-control/v6/internal/subsystems/ai-runtime.md index a8001aa9b..73258760a 100644 --- a/docs/release-control/v6/internal/subsystems/ai-runtime.md +++ b/docs/release-control/v6/internal/subsystems/ai-runtime.md @@ -25,6 +25,104 @@ that same result. Successful reads retain their content and execution provenance ## Purpose +Action reconciliation refreshes the durable product investigation record from +the authoritative session/action even when the finding outcome already matches. +The same builder owns initial completion and later refresh. Completion replaces +an early action projection with the final investigation evidence. Later action +refresh preserves original prose, impact, confidence, evidence and rollback. Unchanged hydration is a no-op, and a +record-only repair does not repeat outcome notifications. Resolved findings +retain the canonical action-history link and resolved status in Assistant context. +Live approved and rejected recovery cases pass, but storage's lexical scorecard +pass fails semantic review because an exhausted container tmpfs was incorrectly +attributed to host capacity. Storage and broader product qualification stay open. + +The live qualification client follows the canonical resource API's pagination. +The API caps each page at 100, so a larger requested limit cannot establish a +complete inventory. A later-page failure returns an error rather than partial +inventory. Regression proof covers an unhealthy resource beyond the first 100 +and failure while reading a later page. This changes collection coverage, not +fault oracles, model context policy or outcome scoring. + +Stored chat and product history share the result-bearing `TranscriptToolCall` +contract. API adapters preserve observed output and the explicit success/error +bit. Only provider-request projections remove those display fields. A failed +read must not become an invocation with no visible result on the way to Patrol +or Assistant history. The adapter regression includes a `NO_AGENT` result and +`success: false`, and provider serialization retains its existing narrower shape. + +Capturing a typed proposal does not create an action. The proposal response +discloses that broker validation is still pending, without forcing the model to +stop investigating. If the broker later refuses submission, the enterprise +orchestrator retains the model's diagnosis unchanged and records the broker +error as a failed investigation needing attention, with no action reference. +The real disconnected-agent case must remain unsuccessful until its actual +transport prerequisite is satisfied. A successful model turn or recorded +proposal is not approval, execution or recovery. + +The shared investigation review renders sanitized Markdown and does not repeat +an identical persisted/fetched conclusion or error. Distinct evidence remains +visible. Pausing scheduled Patrol does not disable history review. The review +control has one detail target and returns keyboard focus when closed. Merged tool +results use Assistant's shared expandable evidence component. Historical calls +without an explicit result bit do not gain an inferred success/failure state. +Resolution copy cannot infer manual review from `needs_attention` or `cannot_fix`. + +The shared pricing table includes reviewed standard Gemini 3.8 Flash rates for +the exact direct and OpenRouter routes. OpenRouter variants and aliases remain +unpriced until independently reviewed. Rates carry the review date and are +estimates, not reconciled provider charges. The introductory rates require a +new review on 2027-01-01. `TestGemini38FlashReviewedRoutePricing` covers real +qualification token counts and request-route preservation, and +`TestGemini38OpenRouterPricingDoesNotGuessVariantRates` preserves unknown variants. + +The canonical query tool describes the app-container configuration boundary +explicitly: TrueNAS supports `config`, while Docker/Podman expose their collected +health, mounts, ports and networks through `get`. This communicates the existing +adapter contract to the model. It does not add configuration access, suppress +tool errors or weaken qualification gates. The existing +`TestAppContainerConfigObservationContract` retains unsupported-adapter and +provider/identity boundaries. + +Shared app-container query mount evidence preserves native type, source, +destination, options and canonical read/write access. Compound options such as +`ro,noexec` cannot become writable through string equality heuristics. Both the +canonical provider and typed read-state projection preserve the same fields. +Configured size in mount options does not establish filesystem usage or free +space. `TestQueryPreservesMountConfigurationEvidence` covers these contracts. +The ordinary live storage case still failed diagnosis by excluding resource +pressure without capacity evidence. Mount fidelity alone does not qualify model +interpretation. Incident-record lookup across canonical and legacy Docker IDs +and typed compatibility lookup of canonical IDs remain explicit identity gaps. + +Shared query projections name command transport explicitly through +`command_agent_connected`, `node_command_agent_connected` and the corresponding +topology counts. These observations do not establish monitoring freshness or +installation state. A topology built without a connection snapshot omits command +flags, execution hints and connected counts. An observed empty snapshot preserves +false/zero. Assistant's inventory seed carries that same absence semantics. +The parent node's connection cannot become a direct guest connection merely +because provider placement names that node. Existing command routing, control, +approval and invocation enforcement remain authoritative. A `can_execute` hint +reflects connected transport with control enabled, not approval for an operation. +`TestCommandConnectivityDoesNotReplaceMonitoringEvidence`, +`TestTopologyOmitsUnobservedCommandConnections` and +`TestAssistantInventoryDoesNotInventCommandConnectionObservations` cover these +projection and continuity boundaries. Existing persisted tool records are not +rewritten, and this contract does not qualify model diagnosis or recovery. + +Native app-container configuration reads resolve identity, provider and placement +from current canonical inventory. Optional session discovery cannot fabricate a +not-found result or replace current placement with a stale execution target. +Query restrictions on both the supplied reference and canonical identity are +checked before registration, and an existing session's allowed actions are not +expanded by a read. Unsupported adapters, missing providers, incomplete placement +and nil provider observations retain known resource identity and an explicit +unavailability reason with the shared tool error bit. They cannot count as a +successful configuration read. Actual inventory absence remains distinct. +`TestAppContainerConfigObservationContract` exercises these boundaries with +strict resolution enabled. This read correction does not relax action or native +log validation and does not qualify autonomous diagnosis or recovery. + The published Patrol qualification schema must accept the fault injectors used by the executable catalogue. `TestCatalogFaultInjectorsMatchPublishedSchema` checks the actual scenario faults against the schema enum, including the @@ -797,6 +895,7 @@ cheap local detection into model-owned diagnosis and governed action. 31. `internal/agentcapabilities/tool_names.go` shared with `api-contracts`: the Pulse Intelligence registry tool-name vocabulary is both the native Assistant execution/display contract and the canonical API/agent tool identity contract for MCP-facing external-agent adapters. 32. `internal/agentcapabilities/tool_response.go` shared with `api-contracts`: the shared tool response envelope, tool error-code vocabulary, and tool-result error-code and verification evidence parsers are both the Assistant structured tool-result contract and the canonical API/agent branching contract for Pulse Intelligence tool failures, recovery tracking, and write self-verification. 33. `internal/agentcapabilities/tool_result.go` shared with `api-contracts`: the Pulse Intelligence shared tool-result content/result envelope, structuredContent projection, result constructors, HTTP response-to-result mapping, text projection, and result interpretation helpers are both the Assistant registry result contract and the canonical API/agent result projection contract for governed tool outcomes. +34. `internal/agentcapabilities/transcript.go` shared with `api-contracts`: Stored Assistant tool results and product history share one result-bearing transcript contract, with an explicit narrower provider-request projection. Observed failures and absent historical result status must survive the API boundary. 34. `internal/agentcapabilities/types.go` shared with `api-contracts`: the agent capabilities manifest wire type, manifest-owned external-adapter surface tool contract field, capability display title and structured output schema fields, approval-policy vocabulary, capability governance normalization, and tool-governance descriptor shape are both the canonical API payload contract and the AI runtime projection contract for Pulse Assistant and MCP-facing agent tools. 35. `internal/agentcapabilities/workflow_prompt.go` shared with `api-contracts`: the Pulse Intelligence workflow prompt catalogue, manifest-owned `workflowPrompts` projection, MCP prompt title projection, presentation kind hints, shared resource-context and finding argument vocabulary, Patrol issue-handling capability gating, argument validation, and manifest-gated shared prompt rendering rules are both the AI runtime starter contract for Assistant-compatible surfaces and the canonical API/agent prompt projection contract for MCP-facing clients. 36. `internal/api/ai_handler.go` shared with `api-contracts`: Pulse Assistant handlers are both an AI runtime control surface and a canonical API payload contract boundary. @@ -3565,6 +3664,22 @@ has a single definition in the canonical resource contract. ## Completion Obligations +The `pulse_knowledge` incidents action reads the organization-pinned canonical +resource timeline used by resource history and Assistant handoffs. It preserves +resource identity, observation and optional occurrence time, source and event +metadata. Reads use explicit observation-time bounds and a bounded event count +with truncation disclosure. Empty retained history is not continuous healthy +coverage, and an unavailable or failed history store is a failed tool read. +Legacy recording IDs are isolated archive lookups bound to the requested +resource. Their recorder timestamps cannot establish source measurement time. +The primary history path must not restore the legacy recorder as a parallel +incident authority or derive fresh history by resampling cached metrics. +`TestIncidentHistoryRetainsCanonicalEvidence` uses SQLite lifecycle records, +time/resource negative controls, missing occurrence time and bounded reads. +The corresponding unavailable/invalid and archive tests cover failure semantics +and resource isolation. Live model interpretation remains governed by the +customer journey qualification plan. + Every per-organization Assistant or legacy AI service that can discover or dispatch through the host-agent command transport must receive an organization-pinned command-server view. A tenant service must never enumerate @@ -8061,3 +8176,30 @@ unchanged pre-existing inventory. The ordinary package run skips live Docker work unless explicitly enabled. A direct fixture restart is teardown and must never be counted as a Pulse approval, execution, rejection or outcome. Model-led and canonical-action qualification remain separate required evidence. + +### Legacy incident recording retirement + +`internal/metrics/incident_archive.go` owns the historical recording format and +explicit, organization-pinned archive reads. `IncidentArchiveProvider` exposes +only a resource-bound window lookup with an error result. The primary +`pulse_knowledge` incident action reads the canonical resource timeline. An +explicit `window_id` reads saved legacy observations and labels their timestamp +and historical-status limits. It must never start a recorder or infer source +freshness from a recording timestamp. The disconnected incident coordinator, +fleet sampling adapter, timer loop, retention writer and duplicate tool adapter +are retired. There is no replacement incident scheduler or diagnosis policy. + +The archive reader preserves saved timestamps, resource labels, metadata and +summary values, including records older than the former retention period. The response explicitly +identifies the nanosecond encoding of the historical `summary.duration_ms` field. It +distinguishes unavailable archives, failed reads and absent exact resource/window +pairs. Proof lives in `internal/metrics/incident_archive_test.go` and the +registered-tool cases in `internal/ai/tools/incident_history_test.go`. + +The legacy incidents listing still exposes incident memory and is not a complete +canonical incident query. Its old sampler-derived `active_count` is now null with +`active_count_status=not_measured`. Canonical-only events, alias-aware listing, +query bounds and projection-read errors remain an explicit modernization gap in +`patrol-assistant-customer-outcome-qualification`. The shared resource timeline +remains the evidence owner. Do not invent another incident lifecycle to repair +this listing. diff --git a/docs/release-control/v6/internal/subsystems/api-contracts.md b/docs/release-control/v6/internal/subsystems/api-contracts.md index fcb1d3e1f..c611f6b40 100644 --- a/docs/release-control/v6/internal/subsystems/api-contracts.md +++ b/docs/release-control/v6/internal/subsystems/api-contracts.md @@ -20,6 +20,24 @@ ## Purpose +Action reconciliation refreshes the durable product investigation record from +the authoritative session/action even when the finding outcome already matches. +The same builder owns initial completion and later refresh. Original prose, +evidence and retained rollback survive. Unchanged hydration is a no-op, and a +record-only repair does not repeat outcome notifications. Resolved findings +retain the canonical action-history link and resolved status in Assistant context. +Live approved and rejected recovery cases pass, but storage's lexical scorecard +pass fails semantic review because an exhausted container tmpfs was incorrectly +attributed to host capacity. Storage and broader product qualification stay open. + +Product history retains the stored result-bearing `TranscriptToolCall` contract, +including observed output and an optional success bit. A false bit is retained, +and an absent historical result bit stays absent. Provider requests use the +explicit narrower projection. The history adapter cannot project away evidence +by treating a display transcript as request arguments. The frontend Patrol API +extends the shared Assistant tool-call shape, and transport regressions pin +failed output and unknown historical status across the message envelope. + The internal Patrol bridge preserves explicit execution limits, scoped tool allowlists and execution identity. The retired unmatched-signal evaluator no longer contributes a signal-count-derived successful-report budget. Diagnosis @@ -1683,6 +1701,7 @@ payload shape change when the portal presents compact client rows. 57. `internal/agentcapabilities/tool_names.go` shared with `ai-runtime`: the Pulse Intelligence registry tool-name vocabulary is both the native Assistant execution/display contract and the canonical API/agent tool identity contract for MCP-facing external-agent adapters. 58. `internal/agentcapabilities/tool_response.go` shared with `ai-runtime`: the shared tool response envelope, tool error-code vocabulary, and tool-result error-code and verification evidence parsers are both the Assistant structured tool-result contract and the canonical API/agent branching contract for Pulse Intelligence tool failures, recovery tracking, and write self-verification. 59. `internal/agentcapabilities/tool_result.go` shared with `ai-runtime`: the Pulse Intelligence shared tool-result content/result envelope, structuredContent projection, result constructors, HTTP response-to-result mapping, text projection, and result interpretation helpers are both the Assistant registry result contract and the canonical API/agent result projection contract for governed tool outcomes. +60. `internal/agentcapabilities/transcript.go` shared with `ai-runtime`: Stored Assistant tool results and product history share one result-bearing transcript contract, with an explicit narrower provider-request projection. Observed failures and absent historical result status must survive the API boundary. 60. `internal/agentcapabilities/types.go` shared with `ai-runtime`: the agent capabilities manifest wire type, manifest-owned external-adapter surface tool contract field, capability display title and structured output schema fields, approval-policy vocabulary, capability governance normalization, and tool-governance descriptor shape are both the canonical API payload contract and the AI runtime projection contract for Pulse Assistant and MCP-facing agent tools. 61. `internal/agentcapabilities/workflow_prompt.go` shared with `ai-runtime`: the Pulse Intelligence workflow prompt catalogue, manifest-owned `workflowPrompts` projection, MCP prompt title projection, presentation kind hints, shared resource-context and finding argument vocabulary, Patrol issue-handling capability gating, argument validation, and manifest-gated shared prompt rendering rules are both the AI runtime starter contract for Assistant-compatible surfaces and the canonical API/agent prompt projection contract for MCP-facing clients. 62. `internal/api/access_control_handlers.go` shared with `organization-settings`: RBAC role and user-assignment handlers are both an organization settings control surface and a canonical API payload contract boundary. @@ -10670,3 +10689,18 @@ reasoning and real remediation in `docs/qualification/PATROL_ASSISTANT_CUSTOMER_JOURNEY.md`. The repeatable browser proof is `scripts/check-patrol-assistant-journey.mjs`. A passing scripted response does not establish a useful customer outcome or model qualification. + +### Explicit historical incident archives + +`internal/api/router.go` pins a read-only legacy archive to each organization's +Assistant service. Its native `IncidentArchiveProvider` capability requires both +resource and window identifiers and propagates read errors. No archive setup or +shutdown writes files or starts sampling. The tool preserves stored metadata and +anomalies and marks historical recording status as historical. + +`GET /api/ai/incidents` retains the `active_count` key as null and adds +`active_count_status=not_measured` in every response. Incident memory, an empty +result and unavailable services cannot establish a current count. The old +coordinator never received production alert callbacks, so its zero was not a +measurement. The legacy listing's broader canonical query and read-error +modernization remains open under the customer-outcome qualification gap. diff --git a/docs/release-control/v6/internal/subsystems/frontend-primitives.md b/docs/release-control/v6/internal/subsystems/frontend-primitives.md index 7af5c7136..267931e28 100644 --- a/docs/release-control/v6/internal/subsystems/frontend-primitives.md +++ b/docs/release-control/v6/internal/subsystems/frontend-primitives.md @@ -20,6 +20,14 @@ ## Purpose +Disk I/O presentation preserves each observed direction independently. Shared +formatting renders a missing rate as a dash and measured idle as numeric zero. +Partial observations cannot form a complete throughput total for sorting or +comparison. Machines column preferences must preserve an explicit user choice +across the first reload, including default-hidden migrations. Final-source +browser proof covers Docker host details, Machines column selection and tooltip +focus/dismissal at desktop, intermediate and narrow widths. + Overview delivery diagnoses use latest-started refresh ownership. Older bulk responses cannot overwrite newer card notification status, and an empty active alert set invalidates outstanding reads. Disposal also prevents updates. Failed diff --git a/docs/release-control/v6/internal/subsystems/monitoring.md b/docs/release-control/v6/internal/subsystems/monitoring.md index b0511eb28..7d242a91b 100644 --- a/docs/release-control/v6/internal/subsystems/monitoring.md +++ b/docs/release-control/v6/internal/subsystems/monitoring.md @@ -17,6 +17,28 @@ ## Purpose +Docker mount collection preserves both native `Mounts` records and entries +reported only in `HostConfig.Tmpfs`. Existing reported destinations remain +authoritative. Additional tmpfs destinations are ordered deterministically, +retain their options and read/write setting, and use the existing mount report +shape. Configured tmpfs size is configuration, not measured used/free space. +`TestCollectContainerPreservesTmpfsMounts` reproduces a live tmpfs-only inspect +shape and covers mixed mounts, read-only options, overlap and absent host config. + +Docker collection records read and write counter presence independently, +including explicit zero, in optional report fields. Older reports without those +fields establish only positive counters. Container reports propagate this +presence to the shared rate tracker. An omitted block-I/O payload and the first counter sample +produce no rate history. Unchanged observed counters produce measured zero, +including after an omitted report. Container writable/root layer sizes never +produce capacity usage history. The ingestion regression is +`internal/monitoring/docker_metric_presence_test.go`. This changes measurement +projection only and grants no agent lifecycle authority. +The shared resource-to-browser conversion preserves each optional I/O rate +independently. Missing directions are omitted from JSON, while measured zero +remains numeric zero. No aggregate presence flag may fabricate its sibling +direction. `TestResourceDiskIOWirePreservesAbsentDirection` pins that wire path. + ### PBS datastore alert evaluation belongs to the live poll After publishing freshly polled PBS datastore storage rows, the poller invokes @@ -164,6 +186,15 @@ fixtures for 500, 502 quoting 403, 503 quoting 404, and genuine 401/403/404. These tests prove cache retention/removal, not installed PBS wake, service restart, or notification receipt. +Docker alert lifecycle events pass through the shared resource history identity +writer. A full Docker source reference must reach the same canonical container +history as inventory changes, including recovery after inventory removal and +restart. Existing alert lifecycle event IDs remain unchanged so replay cannot +duplicate retained events. Same-name containers and abbreviated IDs must not +join another container's history. The real alert-manager callback path is +covered by `TestDockerAlertTimelineUsesCanonicalHistoryIdentity` in +`internal/monitoring/monitor_alert_handling_test.go`. + TrueNAS native alert projection preserves the trimmed, uppercase provider level in ResourceIncident.NativeSeverity. INFO and NOTICE retain the same canonical monitor risk; consumers must not lose their distinct actionability when projecting provider evidence. Native CRITICAL, ALERT, and EMERGENCY all project to canonical critical severity; EMERGENCY must not be discarded as unknown or make a still-active condition appear recovered. WARNING remains warning, and INFO and NOTICE remain informational at this projection boundary. Verification: `TestIncidentProjectionPreservesNativeSeverity` in `internal/truenas/provider_pool_health_contract_test.go` covers all seven native levels and case/whitespace normalization. `TestTrueNASNativeSeverityDispatch` in `internal/alerts/truenas_native_dispatch_test.go` verifies downstream INFO suppression, NOTICE preservation, notification severity, duplicate-poll retention, and confirmed recovery callback identity. The TrueNAS lifecycle tests in `internal/alerts/unified_incidents_test.go` require repeated EMERGENCY evidence to interrupt recovery confirmation. These are fixture-based projection and manager checks, not appliance ingestion or external notification-provider receipt proof. diff --git a/docs/release-control/v6/internal/subsystems/patrol-intelligence.md b/docs/release-control/v6/internal/subsystems/patrol-intelligence.md index be689ee94..37c512dea 100644 --- a/docs/release-control/v6/internal/subsystems/patrol-intelligence.md +++ b/docs/release-control/v6/internal/subsystems/patrol-intelligence.md @@ -15,6 +15,26 @@ ## Purpose +Action reconciliation refreshes the durable product investigation record from +the authoritative session/action even when the finding outcome already matches. +The same builder owns initial completion and later refresh. Original prose, +evidence and retained rollback survive. Unchanged hydration is a no-op, and a +record-only repair does not repeat outcome notifications. Resolved findings +retain the canonical action-history link and resolved status in Assistant context. +Live approved and rejected recovery cases pass, but storage's lexical scorecard +pass fails semantic review because an exhausted container tmpfs was incorrectly +attributed to host capacity. Storage and broader product qualification stay open. + +Pausing the Patrol schedule does not disable investigation history review. +The finding review control owns one detail target and regains keyboard focus on +close. Investigation conclusions use sanitized Markdown, deduplicate identical +stored/fetched text, and retain distinct evidence. Merged tool results render +through Assistant's shared expandable evidence component with their explicit +success/failure bit. Unknown historical status stays unknown. A broker refusal +remains visible separately from the original diagnosis and cannot be presented +as an accepted action. An attention/cannot-fix outcome does not establish manual +review or identify who resolved a later finding. + Detection retains one model conversation for evidence gathering and finding decisions. Recording one finding does not establish diagnostic sufficiency or remove its evidence tools. Missing assessments and provider failures remain diff --git a/docs/release-control/v6/internal/subsystems/performance-and-scalability.md b/docs/release-control/v6/internal/subsystems/performance-and-scalability.md index 6cdd0a1f5..9556fd2d8 100644 --- a/docs/release-control/v6/internal/subsystems/performance-and-scalability.md +++ b/docs/release-control/v6/internal/subsystems/performance-and-scalability.md @@ -15,6 +15,20 @@ ## Purpose +The Docker/app-container history families `dockercontainer` and `docker` use +separate physical `.observed` series for new disk capacity and block-I/O +measurements. Older disk series lack the required presence/capacity semantics +and remain stored unchanged, but retained reads exclude them from current +evidence. All four shared read APIs expose corrected series under the existing +public metric names. Valid capacity from other providers sharing this storage +family remains writable. `NormalizedSeriesKey` describes physical storage for +coverage/backfill matching. Rollups aggregate each physical generation separately. +Projection happens once per returned series, not per observation, and adds no +query, schema migration, or per-row work for other resource families. +`pkg/metrics/store_docker_observation_contract_test.go` pins legacy coexistence, +zero retention, selected/fleet read parity, rollup separation and unaffected +resource families. + Retained reads use one shared query contract in `pkg/metrics/store.go` for `Query`, `QueryAll`, `QueryAllBatch` and `QueryMetricTypesBatch`. A non-empty preferred resolution no longer hides a newer raw tail, an older uncovered @@ -3116,3 +3130,9 @@ candidate on the same worker with alternating samples, and retains full-route and middleware controls. Identical source or instruction sequences alone do not prove identical timing. The recorded final ten-pair check passes the unchanged time/bytes/allocation gate with no adjacent request-path regression. + +The disconnected incident recorder and its fleet metrics adapter are retired. +Router initialization no longer launches their five-second cached-metrics loop +or allocates per-resource pre-incident buffers. Explicit legacy archive reads +are lazy and bounded to the existing 16 MiB file limit. Canonical resource +history supplies current diagnostic evidence without a second sampling loop. diff --git a/docs/release-control/v6/internal/subsystems/registry.json b/docs/release-control/v6/internal/subsystems/registry.json index dd2a7fa00..f04bef0d7 100644 --- a/docs/release-control/v6/internal/subsystems/registry.json +++ b/docs/release-control/v6/internal/subsystems/registry.json @@ -779,6 +779,14 @@ "api-contracts" ] }, + { + "path": "internal/agentcapabilities/transcript.go", + "rationale": "Stored Assistant tool results and product history share one result-bearing transcript contract, with an explicit narrower provider-request projection. Observed failures and absent historical result status must survive the API boundary", + "subsystems": [ + "ai-runtime", + "api-contracts" + ] + }, { "path": "internal/agentcapabilities/types.go", "rationale": "the agent capabilities manifest wire type, manifest-owned external-adapter surface tool contract field, capability display title and structured output schema fields, approval-policy vocabulary, capability governance normalization, and tool-governance descriptor shape are both the canonical API payload contract and the AI runtime projection contract for Pulse Assistant and MCP-facing agent tools", @@ -1611,6 +1619,7 @@ "internal/config/host_continuity_test.go", "internal/models/metrics_types_test.go", "internal/monitoring/availability_probe_agent_test.go", + "internal/monitoring/docker_metric_presence_test.go", "internal/monitoring/monitor_host_agent_removal_lifecycle_test.go", "internal/monitoring/monitor_host_agents_test.go", "scripts/installtests/agent_state_dir_lifecycle_test.go", @@ -2073,6 +2082,7 @@ "internal/api/ai_intelligence_handlers.go", "internal/config/ai.go", "internal/config/patrol_autopilot_persistence.go", + "internal/metrics/incident_archive.go", "pkg/aicontracts/action_broker.go", "pkg/aicontracts/fix_execution.go", "pkg/aicontracts/investigation.go", @@ -2087,6 +2097,20 @@ "exact_files": [], "require_explicit_path_policy_coverage": true, "path_policies": [ + { + "id": "legacy-incident-archive", + "label": "Explicit read-only legacy recording archive proof", + "match_prefixes": [], + "match_files": [ + "internal/metrics/incident_archive.go" + ], + "allow_same_subsystem_tests": false, + "test_prefixes": [], + "exact_files": [ + "internal/ai/tools/incident_history_test.go", + "internal/metrics/incident_archive_test.go" + ] + }, { "id": "retained-metric-evidence", "label": "retained metric evidence and observation coverage proof", @@ -2186,9 +2210,11 @@ ], "exact_files": [ "internal/api/ai_handler_test.go", + "internal/api/ai_handlers_investigation_additional_test.go", "internal/api/ai_handlers_more_test.go", "internal/api/ai_handlers_patrol_actions_additional_test.go", "internal/api/ai_handlers_test.go", + "internal/api/ai_intelligence_handlers_remediation_additional_test.go", "internal/api/ai_intelligence_handlers_test.go", "internal/api/issue1640_readiness_gate_test.go", "internal/api/issue1640_readiness_transport_test.go", @@ -3222,6 +3248,7 @@ "exact_files": [ "frontend-modern/src/types/api.ts", "internal/api/action_runner_credentials_test.go", + "internal/api/ai_handlers_investigation_additional_test.go", "internal/api/ai_handlers_more_test.go", "internal/api/ai_handlers_patrol_actions_additional_test.go", "internal/api/alerting/external_probe_notifications_test.go", @@ -5945,6 +5972,7 @@ "test_prefixes": [], "exact_files": [ "internal/config/host_continuity_test.go", + "internal/monitoring/docker_metric_presence_test.go", "internal/monitoring/issue1485_unraid_lifecycle_test.go", "internal/monitoring/issue1595_collection_trust_test.go", "internal/monitoring/monitor_docker_test.go", @@ -6066,6 +6094,8 @@ "internal/dockeragent/agent_collect_test.go", "internal/dockeragent/agent_cpu_test.go", "internal/dockeragent/agent_internal_test.go", + "internal/dockeragent/blockio_presence_test.go", + "internal/dockeragent/collect_tmpfs_test.go", "internal/dockeragent/swarm_coverage_test.go" ] }, @@ -6125,13 +6155,15 @@ "internal/models/issue1639_pbs_collision_test.go", "internal/models/metrics_types_test.go", "internal/models/state_host_test.go", + "internal/monitoring/docker_metric_presence_test.go", "internal/monitoring/issue1595_collection_trust_test.go", "internal/monitoring/monitor_full_coverage_test.go", "internal/monitoring/monitor_host_agent_removal_lifecycle_test.go", "internal/monitoring/monitor_host_agents_test.go", "internal/monitoring/monitor_package_updates_test.go", "internal/unifiedresources/adapter_coverage_test.go", - "internal/unifiedresources/registry_test.go" + "internal/unifiedresources/registry_test.go", + "pkg/agents/docker/blockio_presence_test.go" ] }, { @@ -7005,6 +7037,7 @@ "exact_files": [ "pkg/metrics/store_additional_test.go", "pkg/metrics/store_bench_test.go", + "pkg/metrics/store_docker_observation_contract_test.go", "pkg/metrics/store_query_plan_test.go", "pkg/metrics/store_slo_test.go" ] @@ -7132,6 +7165,7 @@ "allow_same_subsystem_tests": false, "test_prefixes": [], "exact_files": [ + "frontend-modern/src/components/Infrastructure/__tests__/infrastructureSelectors.test.ts", "frontend-modern/src/components/Infrastructure/__tests__/UnifiedResourceTable.performance.contract.test.tsx", "frontend-modern/src/components/Infrastructure/__tests__/unifiedResourceTableStateModel.test.ts", "frontend-modern/src/components/Infrastructure/__tests__/useTableWindowing.test.ts", @@ -8078,6 +8112,7 @@ "internal/unifiedresources/adapter_coverage_test.go", "internal/unifiedresources/adapters_test.go", "internal/unifiedresources/ceph_pool_health_contract_test.go", + "internal/unifiedresources/history_identity_test.go", "internal/unifiedresources/host_storage_cleanup_test.go", "internal/unifiedresources/monitor_adapter_read_state_test.go", "internal/unifiedresources/views_test.go" @@ -8475,6 +8510,7 @@ "exact_files": [ "internal/monitoring/issue1595_collection_trust_test.go", "internal/unifiedresources/availability_link_test.go", + "internal/unifiedresources/history_identity_test.go", "internal/unifiedresources/kubernetes_registry_test.go", "internal/unifiedresources/pbs_pmg_registry_test.go", "internal/unifiedresources/registry_merge_policy_test.go", @@ -8512,6 +8548,19 @@ "internal/unifiedresources/action_policy_provenance_test.go" ] }, + { + "id": "resource-history-identity", + "label": "resource history identity and authority isolation proof", + "match_prefixes": [], + "match_files": [ + "internal/unifiedresources/history_identity.go" + ], + "allow_same_subsystem_tests": false, + "test_prefixes": [], + "exact_files": [ + "internal/unifiedresources/history_identity_test.go" + ] + }, { "id": "unified-resource-runtime-support", "label": "unified resource runtime support proof", diff --git a/docs/release-control/v6/internal/subsystems/security-privacy.md b/docs/release-control/v6/internal/subsystems/security-privacy.md index bec73958c..3b34d5660 100644 --- a/docs/release-control/v6/internal/subsystems/security-privacy.md +++ b/docs/release-control/v6/internal/subsystems/security-privacy.md @@ -2809,3 +2809,10 @@ optional link to the configured public URL. It never includes resource names, finding text, commands, evidence, or model names, and it uses the tenant's existing email configuration and recipients under the admin-only report schedule routes. + +Explicit legacy incident archive reads use the organization's pinned data path +and exact resource/window identifiers. They retain bounded regular-file and +symlink checks and expose no enumeration, sampling or writing capability. +Removing the disconnected recorder does not alter alert, action approval or +operator authority. An unrelated organization receives no default archive +fallback. diff --git a/docs/release-control/v6/internal/subsystems/storage-recovery.md b/docs/release-control/v6/internal/subsystems/storage-recovery.md index 562ae7917..65aff7bc4 100644 --- a/docs/release-control/v6/internal/subsystems/storage-recovery.md +++ b/docs/release-control/v6/internal/subsystems/storage-recovery.md @@ -21,6 +21,14 @@ ## Purpose +Container image-layer sizes do not establish filesystem capacity. Docker +resource metrics omit that invalid ratio, and retained queries exclude legacy +ambiguous disk observations while preserving new valid provider measurements. +The shared REST resource projection retains missing read/write directions +independently of measured zero. An idle I/O rate cannot establish available +capacity, backup coverage or recoverability. This correction adds no recovery +authority or verified recovery outcome. + Patrol consumes storage evidence in the original diagnostic conversation. A saved finding does not close storage-read authority before the explicit run limit. Missing backup or recovery evidence remains unknown, and neither finding @@ -5993,3 +6001,11 @@ Unlike resource reports, a `patrol_digest` schedule run produces no generated file under the tenant `reports` directory and never calls the retention prune; the email body is the only artifact. Recovery and retention state are unaffected. + +Legacy `incident_windows.json` recovery is read-only through +`internal/metrics/incident_archive.go`. Construction does not read the file, and +explicit reads neither expire records nor rewrite contents or permissions. +Malformed, oversized, symlink and non-regular archive paths fail visibly. Missing +archives remain distinguishable from a valid archive with no matching window. +The original recording times and historical status must never establish current +source freshness or active recording. diff --git a/docs/release-control/v6/internal/subsystems/unified-resources.md b/docs/release-control/v6/internal/subsystems/unified-resources.md index 6739cc88e..fe099ec83 100644 --- a/docs/release-control/v6/internal/subsystems/unified-resources.md +++ b/docs/release-control/v6/internal/subsystems/unified-resources.md @@ -15,6 +15,25 @@ ## Purpose +Action review distinguishes the recorded plan from live or executed facts. The +shared decision packet labels its state and expiry as planning-time evidence, +including when opened from a resolved Patrol finding. Potential blast radius +does not claim every related resource was affected. Actual execution and +verification remain in the recorded outcome section. The review header uses the +shared action-state presentation, so rejected actions remain identifiable even +without an execution receipt. The rejected-state regression and live +completed/rejected deep-link browser proof cover these historical journeys. + +Docker container writable/root layer bytes describe image composition, not +used/total filesystem capacity, and cannot populate `ResourceMetrics.Disk`. +Optional valid block-I/O rate pointers preserve measured zero. Missing, negative +and non-finite rates remain unavailable. Raw layer metadata remains available. +`TestMetricsFromDockerContainerDistinguishesAbsentAndIdleIO` and the container +I/O projection proof in `internal/unifiedresources/metrics_test.go` pin this +boundary. +The frontend resource contract likewise makes read and write rates independently +optional, preserving this distinction through current-history labels. + Physical disk source freshness preserves the collector-authored `expectedUpdateIntervalSeconds` alongside the actual last observation. Registry ingest, merge, cloning and typed disk views preserve it. Staleness uses the @@ -4802,6 +4821,17 @@ recent-change slice plus facet counts it actually renders. The store now also owns a `resource_changes` persistence table with `RecordChange` and `GetRecentChanges` methods so change history is queryable by canonical ID and time window. +Docker alert source references containing an exact full container ID resolve at +`MonitorAdapter.RecordChange` through the current registry, then a retained +history binding, then the deterministic source-specific container identity. +Names and shortened IDs cannot establish this binding. `history_identity.go` +owns a history-only alias index in the organization-scoped resource store. +Legacy event rows retain their IDs, resource references and timestamps. Reads +expand aliases and canonical predecessor eras without changing operator state, +action requests, approvals, links or exclusions. Separate monitor, API and +Assistant store handles must see current persisted aliases. Missing identity +storage is an error, not evidence of empty history. Retention removes an alias +only after neither identity has retained journal records. That same shared timeline vocabulary now includes the `activity` change kind for provider-read breadcrumbs such as VMware tasks and events, plus the `vmware_adapter` source-adapter token for canonical provenance drill-down. diff --git a/frontend-modern/browser-verification.json b/frontend-modern/browser-verification.json index 557dd5566..133aba2ea 100644 --- a/frontend-modern/browser-verification.json +++ b/frontend-modern/browser-verification.json @@ -1,7 +1,7 @@ { "version": 1, - "base_sha": "df7eca7913c72f913fe196f37acec4d3b104c726", - "verified_at": "2026-09-06T19:40:26Z", + "base_sha": "3853124a391ed451f75402f07951c0142e6b5ad8", + "verified_at": "2026-09-06T20:40:07.804772Z", "result": "passed", "changed_paths": [ "frontend-modern/src/components/DemoBanner.tsx" @@ -9,25 +9,65 @@ "content_sha256": { "frontend-modern/src/components/DemoBanner.tsx": "13fd8dea552d97f1df866438f7ba38eec01f25907dc68a0bec672e1793c3fc97" }, + "backend_content_sha256": { + "internal/agentcapabilities/transcript.go": "356c4ca201470407988ff9b2c1fb848619ed38e9d8db844e7390adce0f93ec19", + "internal/ai/chat/types.go": "1f624daf7e511eb971e2d81b787dcd72b2a82d0e0ecd775aa4e580c59d915382", + "internal/ai/service.go": "25dca2a70a985e8ab07443444a9f05f01a569c62ff7bc068a0b587d500831de6", + "internal/api/chat_service_adapter.go": "6d0ab14456b1901c5020a408de95796ece1aa8d1057dd58737ea2637777d6ccb", + "internal/ai/cost/pricing.go": "7b64bc881311ee0a7c1c8a319fcd162974f5a307ced1679e614ff20ac1ac532c", + "internal/ai/tools/tools_query.go": "3e074b204a8c8c4f8b66eaf147c269a2e57739bfc69d03ea36ee8c47cf8b908a", + "internal/ai/tools/tools_propose.go": "43d720c78a010b72f53f7e4e9e1b7b2a763e7f1ac555edb921b7cdca3e82fd51", + "internal/ai/patrol_findings.go": "d5eeb386f025cca338ac328b1d4ec7a2bf51150023454e2b61670f196fcf0356", + "internal/api/ai_handlers.go": "f8c9b24dc684346da4bddab066540fd43ca4999f5fd379c4dc34b5293f78394c", + "internal/api/patrol_action_reconciliation.go": "bc5da1a8050b94271dcdc88841a0ce3329e1773bd01c8746068391a71740ffa2", + "internal/monitoring/monitor.go": "63c8ef4867c07c96b4cd7c4316b5b8646f11b7e9e43a5233471956f76f33d553", + "internal/monitoring/system_alerts.go": "53ed4f02784636363da273883415331066d8b2484ca50e6b506e69990abb149d" + }, + "enterprise_base_sha": "3d9f4e3051d38027355a2a1f36b8c7f672a09b65", + "enterprise_content_sha256": { + "internal/investigation/orchestrator.go": "d56fd512dc47f8a2453559e89863da5d24dffc0a5977abb8f1ea7a82655cd9e8" + }, + "binary_sha256": "0c19f9b9265eff2214e79ab6b2e4a6b19a5645eb467c994c44af70b3d6593317", "routes": [ - "/proxmox/overview" + "/qualification (isolated Overview component on :5199)", + "/patrol (Activity, All history)", + "/actions?action=act_dcc3b52e5451810e49466daf9a6fccb0", + "/actions?action=act_ee0b736f0430e472e896a456ba3cb6eb", + "Pulse Assistant contextual panel from resolved Patrol findings", + "/qualification (isolated DemoBanner component on :5198)" ], "viewports": [ { - "width": 1280, - "height": 800 + "width": 1440, + "height": 1000 + }, + { + "width": 900, + "height": 1000 }, { "width": 390, - "height": 844 + "height": 1000 } ], "states": [ - "demo-mode banner at 1280x800 with the read-only notice followed by the \"Run Pulse on your own hardware\" link", - "demo-mode banner at 390x844 wrapping onto two lines with the link visible and the dismiss control intact" + "Incoming merged alert delivery diagnosis ordering: hold older request, add alert to start newer request, render current notifications-disabled state, release older ready response, verify current state and both cards remain. Existing Patrol and login proof is retained in the prior committed receipt and runtime source is unchanged.", + "Real persisted approved/verified and rejected findings remain reviewable after resolution. Durable investigation outcome agrees with authoritative action, with original prose retained. No stale Fix Queued status. Exact action links remain available.", + "Completed and Rejected action headers, State when planned and Plan expiry copy, inert settled action controls, independently verified recovery and explicit unavailable rollback. Policy, evidence and delivery disclosures expand and collapse.", + "Assistant handoff shows the exact finding with completed/rejected action context and Chat: Read-only. Browser provider readiness POST is deliberately blocked, so its visible route error is a rendering check and does not retest the provider. No prompt is submitted.", + "Incoming main DemoBanner install link: visible in demo policy, absent outside demo, link hover/focus, exact external setup destination opened in a new tab with noopener/noreferrer, keyboard dismissal and reload persistence. Real backend mock mode remains off. External destination content is intercepted because only navigation is under test." ], "interactions": [ - "load /proxmox/overview on a DEMO_MODE=true mock backend behind Vite in fresh contexts at both widths", - "read the link attributes from the live DOM: href https://pulserelay.pro/#setup, target _blank, rel noopener noreferrer" - ] + "scripts/check-alert-diagnosis-ordering.mjs passed at 1440, 900 and 390 by 1000. Actual pixels inspected at desktop and narrow sizes. This is scripted component evidence, not proof of installed delivery or recipient receipt. Screenshots in /tmp/pulse-alert-diagnosis-ordering/.", + "Final Pro binary and final frontend content exercised in Playwright at 1440, 900 and 390 by 1000. Keyboard open/review, safety disclosure, exact action navigation, direct deep-link reload, policy/evidence/delivery disclosure keyboard toggles, Escape and close-button dismissal, review focus return where retained, Assistant open/close, scrolling and page overflow checks. Desktop/intermediate/phone pixels inspected including deepest evidence and Assistant overlay.", + "Private artifacts: tmp/patrol-gemini-38/live-action-browser/. Final complete matrix passed after removing redundant back-to-back full-page navigations from the proof driver. Earlier proof attempts hit an intermittent bootstrap connection screen. No bootstrap fix or general availability claim is made. API writes blocked except login.", + "After integrating main eef4ea21e73aedaa69380574b1fdf4d9dbdee3c8, repeated the complete real Patrol/Actions/Assistant matrix on binary 0c19f9b9265eff2214e79ab6b2e4a6b19a5645eb467c994c44af70b3d6593317. All three widths passed and pixels were reinspected. Incoming DemoBanner separately passed at the same widths using an isolated Vite cache and actual project styling. Prior frontend hashes remain byte-identical and are retained above. Merged monitoring regressions, API action tests, 37 component tests and type-check pass." + ], + "prior_frontend_content_sha256": { + "frontend-modern/src/components/AI/FindingsPanel.tsx": "f506a26757b4c0ea3adf3f77af10214bfd31578b7122d3904a9b3a7272d1e146", + "frontend-modern/src/components/patrol/ApprovalSection.tsx": "6a18d67d5d3d0a8335589bdf199c340eb4775aea3442dc094971d7f924a0dcc5", + "frontend-modern/src/features/actions/ActionDecisionPacket.tsx": "2f1fd68ec333e7f9e8792d74ba8d7e6a2e95755c82b9c7121d847469a31f931b", + "frontend-modern/src/features/actions/ActionReviewDialog.tsx": "49e12cfd44686bd657ddddfb167c5956ec693b6f5d45d9d43c1e3a2e858a6c7f", + "frontend-modern/src/features/alerts/useAlertOverviewState.ts": "64d0b891e7ad228e8590da859dc25e825b6164c8cf76a01983a219d6cd079b23" + } } diff --git a/frontend-modern/src/api/__tests__/patrol.branchcov0718.test.ts b/frontend-modern/src/api/__tests__/patrol.branchcov0718.test.ts index 48740b4cf..299d82ea1 100644 --- a/frontend-modern/src/api/__tests__/patrol.branchcov0718.test.ts +++ b/frontend-modern/src/api/__tests__/patrol.branchcov0718.test.ts @@ -285,6 +285,16 @@ describe('patrol api — uncovered branch coverage', () => { role: 'assistant', content: 'logs in /var/log grew 40GB', reasoning_content: 'checked du output', + tool_calls: [ + { + id: 'read-1', + name: 'pulse_read', + input: { resource_id: 'container-1' }, + output: 'NO_AGENT', + success: false, + }, + { id: 'query-1', name: 'pulse_query', input: { action: 'metrics' } }, + ], timestamp: '2026-07-18T00:00:05Z', }, ], @@ -299,6 +309,9 @@ describe('patrol api — uncovered branch coverage', () => { expect(result).toEqual(envelope); expect(result.messages).toHaveLength(2); expect(result.messages[1]?.reasoning_content).toBe('checked du output'); + expect(result.messages[1]?.tool_calls?.[0]?.output).toBe('NO_AGENT'); + expect(result.messages[1]?.tool_calls?.[0]?.success).toBe(false); + expect(result.messages[1]?.tool_calls?.[1]?.success).toBeUndefined(); }); it('URL-encodes the finding id segment (separate from the messages suffix)', async () => { diff --git a/frontend-modern/src/api/patrol.ts b/frontend-modern/src/api/patrol.ts index e45416d2f..5354edf4d 100644 --- a/frontend-modern/src/api/patrol.ts +++ b/frontend-modern/src/api/patrol.ts @@ -6,6 +6,7 @@ import { apiFetchJSON } from '@/utils/apiClient'; import { arrayOrEmpty, promoteLegacyAlertIdentifier } from './responseUtils'; import type { InvestigationRecord } from './ai'; +import type { ToolCall } from './aiChat'; import type { ResourceCriticality } from './resourceOperatorState'; import type { PatrolActionReference } from '@/types/actionAudit'; import type { PatrolModelReadinessSnapshot } from '@/types/ai'; @@ -304,9 +305,8 @@ export interface ChatMessage { timestamp: string; } -export interface ChatToolCall { +export interface ChatToolCall extends ToolCall { id: string; - name: string; input: Record; } diff --git a/frontend-modern/src/components/AI/FindingsPanel.tsx b/frontend-modern/src/components/AI/FindingsPanel.tsx index 7c939f4ec..f788613e2 100644 --- a/frontend-modern/src/components/AI/FindingsPanel.tsx +++ b/frontend-modern/src/components/AI/FindingsPanel.tsx @@ -194,6 +194,21 @@ export const FindingsPanel: Component = (props) => { const [filter, setFilter] = createSignal(props.filterOverride ?? 'active'); const [sortBy, setSortBy] = createSignal<'severity' | 'time'>('severity'); const [expandedId, setExpandedId] = createSignal(null); + let panelRoot: HTMLDivElement | undefined; + const closeReviewPanel = () => { + const findingId = expandedId(); + setExpandedId(null); + setManageOpenId(null); + if (findingId) { + queueMicrotask(() => { + panelRoot + ?.querySelector( + `button[aria-controls="${CSS.escape(`finding-${findingId}-details`)}"]`, + ) + ?.focus(); + }); + } + }; const [manageOpenId, setManageOpenId] = createSignal(null); const [actionLoading, setActionLoading] = createSignal(null); const [lastHashScrolled, setLastHashScrolled] = createSignal(null); @@ -1467,7 +1482,10 @@ export const FindingsPanel: Component = (props) => { manualControls.dismiss; return ( -
+
Triggered by alert{finding.alertType ? ` (${finding.alertType})` : ''} • Identifier{' '} @@ -2052,18 +2070,18 @@ export const FindingsPanel: Component = (props) => { {/* Inline Approval Section (replaces manual approval JSX) */} = (props) => { }; return ( -
+
{/* Controls */} -
+
= (props) => {
diff --git a/frontend-modern/src/components/Infrastructure/__tests__/infrastructureSelectors.test.ts b/frontend-modern/src/components/Infrastructure/__tests__/infrastructureSelectors.test.ts index d71e08b27..23406beda 100644 --- a/frontend-modern/src/components/Infrastructure/__tests__/infrastructureSelectors.test.ts +++ b/frontend-modern/src/components/Infrastructure/__tests__/infrastructureSelectors.test.ts @@ -458,3 +458,16 @@ describe('infrastructureSelectors', () => { }); }); }); + +it('does not compare partial disk observations as complete throughput totals', () => { + const known = [ + makeResource(1, { diskIO: { readRate: 0, writeRate: 0 } }), + makeResource(2, { diskIO: { readRate: 100, writeRate: 200 } }), + ]; + const partial = [ + makeResource(3, { diskIO: { readRate: 10000 } }), + makeResource(4, { diskIO: { writeRate: 0 } }), + makeResource(5, { diskIO: {} }), + ]; + expect(computeIOScale([...known, ...partial]).diskIO).toEqual(computeIOScale(known).diskIO); +}); diff --git a/frontend-modern/src/components/Infrastructure/infrastructureSelectors.ts b/frontend-modern/src/components/Infrastructure/infrastructureSelectors.ts index 230feaee7..2110fe9a3 100644 --- a/frontend-modern/src/components/Infrastructure/infrastructureSelectors.ts +++ b/frontend-modern/src/components/Infrastructure/infrastructureSelectors.ts @@ -86,7 +86,9 @@ const getSortValue = (resource: Resource, key: string): number | string | null = case 'network': return resource.network ? resource.network.rxBytes + resource.network.txBytes : null; case 'diskio': - return resource.diskIO ? resource.diskIO.readRate + resource.diskIO.writeRate : null; + return resource.diskIO?.readRate !== undefined && resource.diskIO?.writeRate !== undefined + ? resource.diskIO.readRate + resource.diskIO.writeRate + : null; case 'source': return getInfrastructureSystemIdentitySortLabel(resource); case 'temp': @@ -349,9 +351,8 @@ export const computeIOScale = ( networkValues.push(networkTotal); } - const diskIOTotal = (resource.diskIO?.readRate ?? 0) + (resource.diskIO?.writeRate ?? 0); - if (resource.diskIO) { - diskIOValues.push(diskIOTotal); + if (resource.diskIO?.readRate !== undefined && resource.diskIO?.writeRate !== undefined) { + diskIOValues.push(resource.diskIO.readRate + resource.diskIO.writeRate); } } diff --git a/frontend-modern/src/components/patrol/ApprovalSection.tsx b/frontend-modern/src/components/patrol/ApprovalSection.tsx index 1b22b0f85..42e262e3c 100644 --- a/frontend-modern/src/components/patrol/ApprovalSection.tsx +++ b/frontend-modern/src/components/patrol/ApprovalSection.tsx @@ -23,6 +23,7 @@ import type { ActionAuditState, PatrolActionReference } from '@/types/actionAudi interface ApprovalSectionProps { findingId: string; + findingStatus?: string; investigationOutcome?: string; findingTitle?: string; resourceName?: string; @@ -127,7 +128,7 @@ export const ApprovalSection: Component = (props) => { current?.plan.message || investigation()?.summary || 'Review the current Patrol finding and its governed action state.', - findingStatus: 'active', + findingStatus: props.findingStatus ?? 'active', investigationOutcome: props.investigationOutcome, loopState: props.investigationOutcome || current?.state, resourceId: props.resourceId || current?.resource_id, diff --git a/frontend-modern/src/components/patrol/InvestigationMessages.tsx b/frontend-modern/src/components/patrol/InvestigationMessages.tsx index 6b941d689..f5ede3cb5 100644 --- a/frontend-modern/src/components/patrol/InvestigationMessages.tsx +++ b/frontend-modern/src/components/patrol/InvestigationMessages.tsx @@ -10,6 +10,7 @@ import { getInvestigationMessages, formatTimestamp, type ChatMessage } from '@/a import { LoadingSpinner } from '@/components/shared/LoadingSpinner'; import { getInvestigationMessagesState } from '@/utils/patrolEmptyStatePresentation'; import { renderMarkdown } from '@/components/AI/aiChatUtils'; +import { ToolExecutionBlock } from '@/components/AI/Chat/ToolExecutionBlock'; // Compact variant of the Assistant chat's markdown styling, scaled for the // investigation thread's text-xs bubbles. @@ -110,16 +111,35 @@ export const InvestigationMessages: Component = (pro
{(tc) => ( -
- - {tc.name} - - 0}> -
-                                  {JSON.stringify(tc.input, null, 2)}
-                                
-
-
+ + + {tc.name} + + 0}> +
+                                      {JSON.stringify(tc.input, null, 2)}
+                                    
+
+ +
+                                      {tc.output}
+                                    
+
+
+ } + > + + )}
diff --git a/frontend-modern/src/components/patrol/InvestigationSection.tsx b/frontend-modern/src/components/patrol/InvestigationSection.tsx index 65b8257bd..d3d3e8f90 100644 --- a/frontend-modern/src/components/patrol/InvestigationSection.tsx +++ b/frontend-modern/src/components/patrol/InvestigationSection.tsx @@ -29,6 +29,7 @@ import { buildPatrolInvestigationRecordPresentation } from '@/features/patrol/pa import { LoadingSpinner } from '@/components/shared/LoadingSpinner'; import { MetadataBadge } from '@/components/shared/MetadataBadge'; import { InvestigationMessages } from './InvestigationMessages'; +import { renderMarkdown } from '@/components/AI/aiChatUtils'; import { notificationStore } from '@/stores/notifications'; import { aiIntelligenceStore } from '@/stores/aiIntelligence'; import type { InvestigationRecord } from '@/api/ai'; @@ -40,6 +41,9 @@ const INVESTIGATION_BADGE_PROPS = { shape: 'rounded', } as const; +const summaryClass = + 'text-sm prose prose-slate prose-sm dark:prose-invert max-w-none break-words prose-headings:my-2 prose-p:my-2 prose-pre:overflow-x-auto prose-code:break-all prose-code:before:content-none prose-code:after:content-none'; + interface InvestigationSectionProps { findingId: string; investigationStatus?: string; @@ -210,7 +214,11 @@ export const InvestigationSection: Component = (props
-

{investigationRecord().conclusion}

+

@@ -325,6 +333,7 @@ export const InvestigationSection: Component = (props = (props {/* Summary */} - -

{inv().summary}
+ +
{/* Tools used + turn count */} diff --git a/frontend-modern/src/components/patrol/__tests__/ApprovalSection.test.tsx b/frontend-modern/src/components/patrol/__tests__/ApprovalSection.test.tsx index 5f19b2f97..e04bef736 100644 --- a/frontend-modern/src/components/patrol/__tests__/ApprovalSection.test.tsx +++ b/frontend-modern/src/components/patrol/__tests__/ApprovalSection.test.tsx @@ -67,13 +67,17 @@ describe('ApprovalSection typed action handoff', () => { window.history.replaceState({}, '', '/'); }); - const renderSection = (investigationOutcome: string) => + const renderSection = (investigationOutcome: string, findingStatus = 'active') => render(() => ( ( - + )} /> @@ -104,13 +108,20 @@ describe('ApprovalSection typed action handoff', () => { it('routes terminal action history to the exact recorded outcome', async () => { getInvestigationMock.mockResolvedValue(investigation(actionReference('completed'))); - renderSection('fix_verified'); + renderSection('fix_verified', 'resolved'); expect(await screen.findByRole('link', { name: /view outcome in actions/i })).toHaveAttribute( 'href', '/actions?action=act-1', ); expect(screen.getByText('Outcome verified')).toBeInTheDocument(); + fireEvent.click(screen.getByRole('button', { name: /discuss with assistant/i })); + expect(openMock).toHaveBeenCalledWith( + expect.objectContaining({ + handoffContext: expect.stringContaining('Resolved'), + autonomousMode: false, + }), + ); }); it('keeps missing plan identity visible while leaving replan guidance to Actions', async () => { diff --git a/frontend-modern/src/components/patrol/__tests__/InvestigationMessages.test.tsx b/frontend-modern/src/components/patrol/__tests__/InvestigationMessages.test.tsx new file mode 100644 index 000000000..b1b8c6d3a --- /dev/null +++ b/frontend-modern/src/components/patrol/__tests__/InvestigationMessages.test.tsx @@ -0,0 +1,68 @@ +import { cleanup, fireEvent, render, screen } from '@solidjs/testing-library'; +import { afterEach, describe, expect, it, vi } from 'vitest'; +import InvestigationMessages from '../InvestigationMessages'; + +const getMessages = vi.hoisted(() => vi.fn()); +vi.mock('@/api/patrol', () => ({ + getInvestigationMessages: getMessages, + formatTimestamp: (value: string) => value, +})); + +afterEach(cleanup); + +describe('InvestigationMessages', () => { + it('retains merged failed tool evidence and exposes it through the shared disclosure', async () => { + getMessages.mockResolvedValue({ + messages: [ + { + id: 'turn-1', + role: 'assistant', + content: '', + timestamp: '2026-09-06', + tool_calls: [ + { + id: 'read-1', + name: 'pulse_read', + input: { resource_id: 'app-container-1' }, + output: '{"error":"NO_AGENT","message":"Command agent is not connected"}', + success: false, + }, + ], + }, + ], + }); + render(() => ); + const disclosure = await screen.findByRole('button', { name: /failed/i }); + expect(disclosure).toHaveAttribute('aria-expanded', 'false'); + fireEvent.keyDown(disclosure, { key: 'Enter' }); + expect(disclosure).toHaveAttribute('aria-expanded', 'true'); + expect(screen.getByText(/"error":\s*"NO_AGENT"/)).toBeVisible(); + fireEvent.keyDown(disclosure, { key: ' ' }); + expect(disclosure).toHaveAttribute('aria-expanded', 'false'); + }); + + it('does not infer completion from historical calls that have no recorded result status', async () => { + getMessages.mockResolvedValue({ + messages: [ + { + id: 'turn-2', + role: 'assistant', + content: '', + timestamp: '2026-09-06', + tool_calls: [ + { + id: 'query-1', + name: 'pulse_query', + input: { action: 'metrics' }, + output: 'Historical output', + }, + ], + }, + ], + }); + render(() => ); + expect(await screen.findByText('Historical output')).toBeVisible(); + expect(screen.queryByText('completed')).not.toBeInTheDocument(); + expect(screen.queryByText('failed')).not.toBeInTheDocument(); + }); +}); diff --git a/frontend-modern/src/components/patrol/__tests__/InvestigationSection.test.tsx b/frontend-modern/src/components/patrol/__tests__/InvestigationSection.test.tsx index b0a9ac023..2db7384f4 100644 --- a/frontend-modern/src/components/patrol/__tests__/InvestigationSection.test.tsx +++ b/frontend-modern/src/components/patrol/__tests__/InvestigationSection.test.tsx @@ -133,4 +133,46 @@ describe('InvestigationSection', () => { expect(screen.queryByText(/No investigation data available/)).not.toBeInTheDocument(); expect(screen.queryByText('systemctl restart workload.service')).not.toBeInTheDocument(); }); + + it.each([false, true])( + 'preserves readable evidence and distinct summaries (different: %s)', + async (different) => { + const conclusion = + '### Root cause\n\nThe cause is **unknown**.\n\n- The health check failed.\n- Logs are unavailable.'; + getInvestigationMock.mockResolvedValue({ + id: 'inv-markdown', + finding_id: 'finding-markdown', + session_id: 'session-markdown', + status: 'completed', + started_at: '2026-09-06T17:00:00Z', + turn_count: 2, + summary: different + ? '### Follow-up\n\nAdditional evidence remains unavailable.' + : conclusion, + } satisfies Investigation); + render(() => ( + + )); + await screen.findByRole('button', { name: 'Show investigation thread' }); + expect(screen.getAllByRole('heading', { name: 'Root cause' })).toHaveLength(1); + expect(screen.getByText('unknown').tagName).toBe('STRONG'); + expect(screen.getAllByRole('listitem')).toHaveLength(2); + expect(screen.queryByRole('heading', { name: 'Follow-up' }) !== null).toBe(different); + }, + ); }); diff --git a/frontend-modern/src/features/actions/ActionDecisionPacket.tsx b/frontend-modern/src/features/actions/ActionDecisionPacket.tsx index 55fc51c17..f1b365dc2 100644 --- a/frontend-modern/src/features/actions/ActionDecisionPacket.tsx +++ b/frontend-modern/src/features/actions/ActionDecisionPacket.tsx @@ -43,7 +43,7 @@ export const ActionDecisionPacket: Component<{ class="rounded-lg border border-border bg-surface p-4" >

- What will happen + Action plan

@@ -69,7 +69,7 @@ export const ActionDecisionPacket: Component<{
-
Current state
+
State when planned
{props.audit.plan.preflight?.currentState}
@@ -80,7 +80,7 @@ export const ActionDecisionPacket: Component<{
-
Approval expires
+
Plan expiry
{expiry()}
@@ -90,7 +90,7 @@ export const ActionDecisionPacket: Component<{
0}>
-
Also affected
+
Potentially affected
    {(entry) => ( diff --git a/frontend-modern/src/features/actions/ActionReviewDialog.tsx b/frontend-modern/src/features/actions/ActionReviewDialog.tsx index e08e56bd0..89bf30c11 100644 --- a/frontend-modern/src/features/actions/ActionReviewDialog.tsx +++ b/frontend-modern/src/features/actions/ActionReviewDialog.tsx @@ -4,12 +4,14 @@ import ArrowUpRightIcon from 'lucide-solid/icons/arrow-up-right'; import { ResourceActionsAPI } from '@/api/resourceActions'; import { Button, ButtonLink } from '@/components/shared/Button'; import { Dialog } from '@/components/shared/Dialog'; +import { MetadataBadge } from '@/components/shared/MetadataBadge'; import { notificationStore } from '@/stores/notifications'; import { presentationPolicyIsReadOnly } from '@/stores/sessionPresentationPolicy'; import type { ActionDetailResponse } from '@/types/actionAudit'; import { ActionDecisionPacket } from './ActionDecisionPacket'; import { formatActionName, + getActionInboxStatePresentation, getActionOriginDestination, getActionResourcePresentation, } from './actionPresentation'; @@ -244,9 +246,14 @@ export const ActionReviewDialog: Component<{

    Governed action review

    -

    - {formatActionName(record().request.capabilityName)} -

    +
    +

    + {formatActionName(record().request.capabilityName)} +

    + + {getActionInboxStatePresentation(record().state).label} + +

    {resource().label} · {resource().detail} diff --git a/frontend-modern/src/features/actions/__tests__/ActionReviewDialog.test.tsx b/frontend-modern/src/features/actions/__tests__/ActionReviewDialog.test.tsx index f0eaf7121..823407ac6 100644 --- a/frontend-modern/src/features/actions/__tests__/ActionReviewDialog.test.tsx +++ b/frontend-modern/src/features/actions/__tests__/ActionReviewDialog.test.tsx @@ -115,6 +115,16 @@ const detail = (audit: ActionAuditRecord): ActionDetailResponse => ({ }); describe('ActionReviewDialog trust gates', () => { + it('keeps a rejected action outcome visible without offering execution', () => { + const audit = makeAudit('resolved', '2026-07-12T00:10:00Z'); + audit.state = 'rejected'; + render(() => ); + expect(screen.getByText('Rejected', { exact: true })).toBeVisible(); + expect( + screen.queryByRole('button', { name: /approve|run|refresh plan/i }), + ).not.toBeInTheDocument(); + }); + it('links a trusted Patrol action back to its exact operational record', () => { const audit = makeAudit('resolved', '2099-01-01T00:00:00Z'); audit.origin = { diff --git a/frontend-modern/src/features/docker/DockerHostDrawerOverview.tsx b/frontend-modern/src/features/docker/DockerHostDrawerOverview.tsx index 3d77cdedb..f3cce4702 100644 --- a/frontend-modern/src/features/docker/DockerHostDrawerOverview.tsx +++ b/frontend-modern/src/features/docker/DockerHostDrawerOverview.tsx @@ -20,7 +20,13 @@ import { hostOverrideIdCandidates } from '@/features/alerts/alertOverridesModel' import { areSystemSettingsLoaded, shouldHideDockerUpdateActions } from '@/stores/systemSettings'; import { useAlertsActivation } from '@/stores/alertsActivation'; import type { Resource } from '@/types/resource'; -import { formatBytes, formatRelativeTime, formatSpeed, normalizeDiskArray } from '@/utils/format'; +import { + formatBytes, + formatRelativeTime, + formatSpeed, + formatObservedSpeed, + normalizeDiskArray, +} from '@/utils/format'; import { formatTemperature, getTemperatureTextClass } from '@/utils/temperature'; interface DockerHostDrawerOverviewProps { @@ -279,7 +285,7 @@ export function DockerHostDrawerOverview(props: DockerHostDrawerOverviewProps) { ) { rows.push({ label: 'Disk I/O', - value: `${formatSpeed(props.host.diskIO?.readRate ?? 0)} / ${formatSpeed(props.host.diskIO?.writeRate ?? 0)}`, + value: `${formatObservedSpeed(props.host.diskIO?.readRate)} / ${formatObservedSpeed(props.host.diskIO?.writeRate)}`, }); } return rows; diff --git a/frontend-modern/src/features/patrol/PatrolIntelligenceSurface.tsx b/frontend-modern/src/features/patrol/PatrolIntelligenceSurface.tsx index 92c5ee61b..6edec6a1e 100644 --- a/frontend-modern/src/features/patrol/PatrolIntelligenceSurface.tsx +++ b/frontend-modern/src/features/patrol/PatrolIntelligenceSurface.tsx @@ -225,9 +225,7 @@ export function PatrolIntelligenceSurface() { onToggle={(event) => setFindingsOpen(event.currentTarget.open)} >

    Finding options and history -
    +

    R - {formatSpeed(props.diskIO?.readRate ?? 0)} + {formatObservedSpeed(props.diskIO?.readRate)} W - {formatSpeed(props.diskIO?.writeRate ?? 0)} + {formatObservedSpeed(props.diskIO?.writeRate)} } > @@ -513,11 +512,11 @@ const AgentMachineDiskIOCell: Component<{

    Read - {formatSpeed(props.diskIO?.readRate ?? 0)} + {formatObservedSpeed(props.diskIO?.readRate)} Write - {formatSpeed(props.diskIO?.writeRate ?? 0)} + {formatObservedSpeed(props.diskIO?.writeRate)}
    @@ -1038,7 +1037,7 @@ const networkTitleFor = (machine: Resource): string => { const diskIOTitleFor = (machine: Resource): string => { if (!machine.diskIO) return ''; - return `Read ${formatSpeed(machine.diskIO.readRate)}\nWrite ${formatSpeed(machine.diskIO.writeRate)}`; + return `Read ${formatObservedSpeed(machine.diskIO.readRate)}\nWrite ${formatObservedSpeed(machine.diskIO.writeRate)}`; }; const agentIdentityIdFor = (machine: Resource): string => @@ -1527,7 +1526,6 @@ export const AgentsMachinesTable: Component<{ aggregateDisk() !== undefined || (disks()?.length ?? 0) > 0; const networkTotal = () => getAgentMachineNetworkTotal(machine); const networkInterfaces = () => getAgentMachineNetworkInterfaceDetails(machine); - const diskIOTotal = () => getAgentMachineDiskIOTotal(machine); const diskIODetails = () => getAgentMachineDiskIODetails(machine); const primaryIp = () => getPreferredResourceIP(machine) ?? getAgentMachinePrimaryIp(machine); @@ -1738,7 +1736,11 @@ export const AgentsMachinesTable: Component<{ class={`${getPlatformTableCellClassForKind('numeric-value')} ${machineColumnWidthClass('diskio')} text-base-content`} > { ).toBe(800); }); - it('returns read only when write is absent', () => { + it('leaves the total unavailable when write is absent', () => { expect( getAgentMachineDiskIOTotal( resource({ diskIO: { readRate: 500 } as unknown as ResourceDiskIO }), ), - ).toBe(500); + ).toBeUndefined(); }); - it('returns write only when read is absent', () => { + it('leaves the total unavailable when read is absent', () => { expect( getAgentMachineDiskIOTotal( resource({ diskIO: { writeRate: 300 } as unknown as ResourceDiskIO }), ), - ).toBe(300); + ).toBeUndefined(); }); it('returns undefined when both rates are absent', () => { diff --git a/frontend-modern/src/features/standalone/__tests__/agentMachineTableModel.test.ts b/frontend-modern/src/features/standalone/__tests__/agentMachineTableModel.test.ts index 4ae78a3fb..84404cb6a 100644 --- a/frontend-modern/src/features/standalone/__tests__/agentMachineTableModel.test.ts +++ b/frontend-modern/src/features/standalone/__tests__/agentMachineTableModel.test.ts @@ -2,6 +2,7 @@ import { describe, expect, it } from 'vitest'; import type { Resource } from '@/types/resource'; import { getAgentMachineDiskPercent, + getAgentMachineDiskIOTotal, getAgentMachineDiskIODetails, getAgentMachineGPUTitle, getAgentMachineGPUUtilizationPercent, @@ -476,3 +477,12 @@ describe('agentMachineTableModel', () => { ); }); }); + +it('requires both disk directions for a throughput total', () => { + expect(getAgentMachineDiskIOTotal(resource({ diskIO: { readRate: 0, writeRate: 0 } }))).toBe(0); + expect(getAgentMachineDiskIOTotal(resource({ diskIO: { readRate: 100, writeRate: 200 } }))).toBe( + 300, + ); + expect(getAgentMachineDiskIOTotal(resource({ diskIO: { readRate: 0 } }))).toBeUndefined(); + expect(getAgentMachineDiskIOTotal(resource({ diskIO: { writeRate: 100 } }))).toBeUndefined(); +}); diff --git a/frontend-modern/src/features/standalone/agentMachineTableModel.ts b/frontend-modern/src/features/standalone/agentMachineTableModel.ts index f8307964e..5fe81f91f 100644 --- a/frontend-modern/src/features/standalone/agentMachineTableModel.ts +++ b/frontend-modern/src/features/standalone/agentMachineTableModel.ts @@ -493,8 +493,8 @@ export const getAgentMachineNetworkInterfaceDetails = ( export const getAgentMachineDiskIOTotal = (machine: Resource): number | undefined => { const read = getPlatformTableFiniteMetric(machine.diskIO?.readRate); const write = getPlatformTableFiniteMetric(machine.diskIO?.writeRate); - if (read === undefined && write === undefined) return undefined; - return (read ?? 0) + (write ?? 0); + if (read === undefined || write === undefined) return undefined; + return read + write; }; export const getAgentMachineDiskIODetails = (machine: Resource): AgentMachineDiskIODetail[] => { diff --git a/frontend-modern/src/hooks/__tests__/useColumnVisibility.test.ts b/frontend-modern/src/hooks/__tests__/useColumnVisibility.test.ts index 18ae6ed4f..e6d4460e8 100644 --- a/frontend-modern/src/hooks/__tests__/useColumnVisibility.test.ts +++ b/frontend-modern/src/hooks/__tests__/useColumnVisibility.test.ts @@ -120,6 +120,29 @@ describe('useColumnVisibility', () => { }); }); + it('does not reapply a default-hidden migration after a fresh user shows the column', async () => { + const columns: ColumnDef[] = [ + { id: 'name', label: 'Name' }, + { id: 'diskio', label: 'Disk I/O', toggleable: true, defaultHidden: true }, + ]; + let dispose = () => {}; + let visibility: ReturnType; + createRoot((d) => { + dispose = d; + visibility = useColumnVisibility(storageKey, columns, [], undefined, {}, ['diskio']); + }); + await Promise.resolve(); + visibility!.show('diskio'); + await Promise.resolve(); + expect(window.localStorage.getItem(storageKey)).toBe('[]'); + dispose(); + createRoot((d) => { + const reloaded = useColumnVisibility(storageKey, columns, [], undefined, {}, ['diskio']); + expect(reloaded.isHiddenByUser('diskio')).toBe(false); + d(); + }); + }); + it('resets back to the canonical default-hidden set', () => { createRoot((dispose) => { const columns: ColumnDef[] = [ diff --git a/frontend-modern/src/hooks/__tests__/useUnifiedResources.test.ts b/frontend-modern/src/hooks/__tests__/useUnifiedResources.test.ts index 74aba914e..45ea92111 100644 --- a/frontend-modern/src/hooks/__tests__/useUnifiedResources.test.ts +++ b/frontend-modern/src/hooks/__tests__/useUnifiedResources.test.ts @@ -1451,6 +1451,25 @@ describe('useUnifiedResources', () => { dispose(); }); + it('preserves an observed zero without inventing its absent I/O direction', async () => { + setWsConnected(false); + setWsInitialDataReceived(false); + setWsState('resources', []); + apiFetchMock.mockResolvedValueOnce({ + ok: true, + json: async () => ({ data: [{ ...v2Resource, metrics: { diskRead: { value: 0 } } }] }), + }); + let dispose = () => {}; + let result: ReturnType | undefined; + createRoot((d) => { + dispose = d; + result = useUnifiedResources(); + }); + await result!.refetch(); + expect(result!.resources()[0]?.diskIO).toEqual({ readRate: 0, writeRate: undefined }); + dispose(); + }); + it('preserves richer REST resource details across thinner websocket updates', async () => { setWsConnected(false); setWsInitialDataReceived(false); diff --git a/frontend-modern/src/hooks/useColumnVisibility.ts b/frontend-modern/src/hooks/useColumnVisibility.ts index 8a4d5bc43..23953810b 100644 --- a/frontend-modern/src/hooks/useColumnVisibility.ts +++ b/frontend-modern/src/hooks/useColumnVisibility.ts @@ -124,21 +124,21 @@ export function useColumnVisibility( const appliedDefaultHiddenMigrations = hasUserPreference ? readAppliedDefaultHiddenMigrations(storageKey, persistedIdAliases) : []; - const pendingDefaultHiddenMigrations = hasUserPreference - ? Array.from( - new Set( - defaultHiddenMigrationIds - .map((id) => id.trim()) - .filter( - (id) => - id && - effectiveDefaultHidden.includes(id) && - toggleableIds.includes(id) && - !appliedDefaultHiddenMigrations.includes(id), - ), + // Fresh preferences already contain these defaults. Mark their migration as + // applied too, so the first reload cannot undo a user's subsequent choice. + const pendingDefaultHiddenMigrations = Array.from( + new Set( + defaultHiddenMigrationIds + .map((id) => id.trim()) + .filter( + (id) => + id && + effectiveDefaultHidden.includes(id) && + toggleableIds.includes(id) && + !appliedDefaultHiddenMigrations.includes(id), ), - ) - : []; + ), + ); let defaultHiddenMigrationsPersisted = false; // Persist hidden columns to localStorage @@ -169,7 +169,7 @@ export function useColumnVisibility( createEffect(() => { const hasUnpersistedDefaultHiddenMigrations = pendingDefaultHiddenMigrations.length > 0 && !defaultHiddenMigrationsPersisted; - if (!hasUserPreference || (!persistedIdsMigrated && !hasUnpersistedDefaultHiddenMigrations)) { + if (!persistedIdsMigrated && !hasUnpersistedDefaultHiddenMigrations) { return; } persistedIdsMigrated = false; diff --git a/frontend-modern/src/hooks/useUnifiedResources.ts b/frontend-modern/src/hooks/useUnifiedResources.ts index aace88c99..a9b46a114 100644 --- a/frontend-modern/src/hooks/useUnifiedResources.ts +++ b/frontend-modern/src/hooks/useUnifiedResources.ts @@ -915,8 +915,8 @@ const toResource = (v2: APIResource): Resource => { diskIO: v2.metrics?.diskRead || v2.metrics?.diskWrite ? { - readRate: v2.metrics?.diskRead?.value ?? 0, - writeRate: v2.metrics?.diskWrite?.value ?? 0, + readRate: v2.metrics?.diskRead?.value, + writeRate: v2.metrics?.diskWrite?.value, } : undefined, uptime: diff --git a/frontend-modern/src/types/resource.ts b/frontend-modern/src/types/resource.ts index d0d1263ce..8fe58ef6e 100644 --- a/frontend-modern/src/types/resource.ts +++ b/frontend-modern/src/types/resource.ts @@ -133,8 +133,8 @@ export interface ResourceNetwork { // Disk I/O metrics (rates in bytes/sec from backend) export interface ResourceDiskIO { - readRate: number; // Read rate (bytes/sec) - writeRate: number; // Write rate (bytes/sec) + readRate?: number; // Observed read rate (bytes/sec), including measured zero. + writeRate?: number; // Absent directions remain unavailable. } // Alert associated with a resource diff --git a/frontend-modern/src/utils/__tests__/aiFindingPresentation.coverage2.test.ts b/frontend-modern/src/utils/__tests__/aiFindingPresentation.coverage2.test.ts index 625b8ffa5..d9663ce54 100644 --- a/frontend-modern/src/utils/__tests__/aiFindingPresentation.coverage2.test.ts +++ b/frontend-modern/src/utils/__tests__/aiFindingPresentation.coverage2.test.ts @@ -995,19 +995,19 @@ describe('getFindingResolutionReason', () => { ).toBe('Resolved after investigation timeout now'); }); - it('returns "Resolved manually" for cannot_fix', () => { + it('does not infer manual resolution from cannot_fix', () => { expect( getFindingResolutionReason({ ...patrolBase, investigationOutcome: 'cannot_fix' }, 'now'), - ).toBe('Resolved manually now'); + ).toBe('Resolved now'); }); - it('returns "Resolved after manual review" for needs_attention', () => { + it('does not infer manual review from needs_attention', () => { expect( getFindingResolutionReason( { ...patrolBase, investigationOutcome: 'needs_attention' }, 'now', ), - ).toBe('Resolved after manual review now'); + ).toBe('Resolved now'); }); it('returns "Fix applied by Patrol" for fix_executed even when autoResolved is false', () => { diff --git a/frontend-modern/src/utils/__tests__/format.test.ts b/frontend-modern/src/utils/__tests__/format.test.ts index b501850ba..bf337a3b8 100644 --- a/frontend-modern/src/utils/__tests__/format.test.ts +++ b/frontend-modern/src/utils/__tests__/format.test.ts @@ -5,6 +5,7 @@ import { describe, expect, it, vi, beforeEach, afterEach } from 'vitest'; import { formatBytes, formatSpeed, + formatObservedSpeed, formatPercent, formatNumber, formatUptime, @@ -362,3 +363,13 @@ describe('getBackupInfo', () => { }); }); }); + +describe('formatObservedSpeed', () => { + it('preserves zero and leaves missing or invalid observations unavailable', () => { + expect(formatObservedSpeed(0)).toBe('0 B/s'); + expect(formatObservedSpeed(1024)).toBe('1.00 KB/s'); + for (const value of [undefined, null, -1, NaN, Infinity]) { + expect(formatObservedSpeed(value)).toBe('-'); + } + }); +}); diff --git a/frontend-modern/src/utils/aiFindingPresentation.ts b/frontend-modern/src/utils/aiFindingPresentation.ts index da2cfe6c5..00b4aae7f 100644 --- a/frontend-modern/src/utils/aiFindingPresentation.ts +++ b/frontend-modern/src/utils/aiFindingPresentation.ts @@ -1287,9 +1287,10 @@ export const getFindingResolutionReason = ( case 'timed_out': return `Resolved after investigation timeout ${resolvedTime}`; case 'cannot_fix': - return `Resolved manually ${resolvedTime}`; case 'needs_attention': - return `Resolved after manual review ${resolvedTime}`; + // An investigation outcome does not identify who later resolved the + // finding. Explicit operator resolution is handled above. + return `Resolved ${resolvedTime}`; default: return `Issue no longer detected ${resolvedTime}`; } diff --git a/frontend-modern/src/utils/format.ts b/frontend-modern/src/utils/format.ts index 223157944..01ed35b2b 100644 --- a/frontend-modern/src/utils/format.ts +++ b/frontend-modern/src/utils/format.ts @@ -67,6 +67,14 @@ export function formatSpeed(bytesPerSecond: number, decimals: number | 'auto' = return `${formatBytes(bytesPerSecond, decimals)}/s`; } +export function formatObservedSpeed(bytesPerSecond: number | null | undefined): string { + return typeof bytesPerSecond === 'number' && + Number.isFinite(bytesPerSecond) && + bytesPerSecond >= 0 + ? formatSpeed(bytesPerSecond) + : '-'; +} + export function formatPercent(value: number): string { if (!Number.isFinite(value)) return '0%'; const abs = Math.abs(value); diff --git a/internal/agentcapabilities/transcript.go b/internal/agentcapabilities/transcript.go new file mode 100644 index 000000000..4ccd48a47 --- /dev/null +++ b/internal/agentcapabilities/transcript.go @@ -0,0 +1,45 @@ +package agentcapabilities + +import "encoding/json" + +// TranscriptToolCall preserves a stored invocation and its observed result. +// Provider requests use ProviderToolCall instead of the product history shape. +type TranscriptToolCall struct { + ID string `json:"id"` + Name string `json:"name"` + Input map[string]interface{} `json:"input"` + Output string `json:"output,omitempty"` + Success *bool `json:"success,omitempty"` + ThoughtSignature json.RawMessage `json:"thought_signature,omitempty"` +} + +func (t TranscriptToolCall) NormalizeCollections() TranscriptToolCall { + providerCall := ProviderToolCall{ + ID: t.ID, + Name: t.Name, + Input: t.Input, + ThoughtSignature: t.ThoughtSignature, + }.NormalizeCollections() + t.ID = providerCall.ID + t.Name = providerCall.Name + t.Input = providerCall.Input + t.ThoughtSignature = providerCall.ThoughtSignature + if t.Success != nil { + success := *t.Success + t.Success = &success + } + return t +} + +// ProviderToolCall projects a stored Assistant transcript call back to the +// shared provider-facing shape, deliberately excluding in-app output/success +// display fields. +func (t TranscriptToolCall) ProviderToolCall() ProviderToolCall { + t = t.NormalizeCollections() + return ProviderToolCall{ + ID: t.ID, + Name: t.Name, + Input: t.Input, + ThoughtSignature: t.ThoughtSignature, + }.NormalizeCollections() +} diff --git a/internal/agentcapabilities/types_test.go b/internal/agentcapabilities/types_test.go index c5938ea3e..5b9371f2f 100644 --- a/internal/agentcapabilities/types_test.go +++ b/internal/agentcapabilities/types_test.go @@ -1,6 +1,7 @@ package agentcapabilities import ( + "encoding/json" "net/http" "slices" "strings" @@ -203,3 +204,46 @@ func TestNewToolGovernanceDescriptorAppliesSharedDefaults(t *testing.T) { t.Fatalf("descriptor approval summary = %q", descriptor.ApprovalSummary) } } + +// Stored result evidence and provider request arguments are different wire +// contracts. In particular, explicit failure cannot disappear through omitempty. +func TestTranscriptToolCallPreservesResultOutsideProviderRequests(t *testing.T) { + failed := false + call := TranscriptToolCall{ID: "read-1", Name: PulseReadToolName, Output: "NO_AGENT", Success: &failed}.NormalizeCollections() + failed = true + body, err := json.Marshal(call) + if err != nil { + t.Fatal(err) + } + var stored map[string]interface{} + if err := json.Unmarshal(body, &stored); err != nil { + t.Fatal(err) + } + if stored["output"] != "NO_AGENT" || stored["success"] != false || stored["input"] == nil { + t.Fatalf("stored result lost explicit failure or normalized input: %s", body) + } + requestBody, err := json.Marshal(call.ProviderToolCall()) + if err != nil { + t.Fatal(err) + } + var request map[string]interface{} + if err := json.Unmarshal(requestBody, &request); err != nil { + t.Fatal(err) + } + for _, key := range []string{"output", "success"} { + if _, exists := request[key]; exists { + t.Fatalf("provider request retained display-only %s: %s", key, requestBody) + } + } + unknownBody, err := json.Marshal(TranscriptToolCall{Name: PulseQueryToolName}.NormalizeCollections()) + if err != nil { + t.Fatal(err) + } + var unknown map[string]interface{} + if err := json.Unmarshal(unknownBody, &unknown); err != nil { + t.Fatal(err) + } + if _, exists := unknown["success"]; exists { + t.Fatalf("unknown historical status became a result: %s", unknownBody) + } +} diff --git a/internal/ai/adapters/adapters.go b/internal/ai/adapters/adapters.go index 85e160a9a..a7e7987f8 100644 --- a/internal/ai/adapters/adapters.go +++ b/internal/ai/adapters/adapters.go @@ -8,7 +8,6 @@ import ( "fmt" "os" "path/filepath" - "strconv" "strings" "sync" "time" @@ -73,165 +72,6 @@ func (a *ForecastDataAdapter) GetMetricHistory(resourceID, metric string, from, return result, nil } -// MetricsAdapter provides current metrics for resources. -// It implements metrics.MetricsProvider for the incident recorder. -// Uses ReadState as the sole data source (SRC-03m migration). -type MetricsAdapter struct { - readState unifiedresources.ReadState -} - -// NewMetricsAdapter creates a new adapter for current metrics. -// ReadState is the sole data source for both GetMonitoredResourceIDs and -// GetCurrentMetrics. Returns nil if readState is nil. -func NewMetricsAdapter(readState unifiedresources.ReadState) *MetricsAdapter { - if readState == nil { - return nil - } - return &MetricsAdapter{readState: readState} -} - -// GetMonitoredResourceIDs returns all resource IDs currently being monitored. -// This is used by the incident recorder to maintain pre-incident buffers for all resources. -// Returns both unified IDs and Proxmox source IDs so that pre-incident buffers -// are keyed by both (alert-triggered recordings use source IDs). -func (a *MetricsAdapter) GetMonitoredResourceIDs() []string { - var ids []string - for _, vm := range a.readState.VMs() { - ids = append(ids, vm.ID()) - if sid := vm.SourceID(); sid != "" && sid != vm.ID() { - ids = append(ids, sid) - } - } - for _, ct := range a.readState.Containers() { - ids = append(ids, ct.ID()) - if sid := ct.SourceID(); sid != "" && sid != ct.ID() { - ids = append(ids, sid) - } - } - for _, node := range a.readState.Nodes() { - ids = append(ids, node.ID()) - if sid := node.SourceID(); sid != "" && sid != node.ID() { - ids = append(ids, sid) - } - } - return ids -} - -// GetCurrentMetricsBatch returns current metrics for every resource -// GetCurrentMetrics can resolve, in one pass over the views, keyed by every -// ID form GetCurrentMetrics matches (unified ID, source ID, VMID string, -// name). The per-ID method scans all views per call, so a sampler asking for -// thousands of resources per tick must use this instead: first-key-wins -// mirrors the per-ID method's VM -> container -> node -> storage precedence. -func (a *MetricsAdapter) GetCurrentMetricsBatch() map[string]map[string]float64 { - out := make(map[string]map[string]float64) - put := func(metrics map[string]float64, keys ...string) { - for _, key := range keys { - if key == "" { - continue - } - if _, exists := out[key]; !exists { - out[key] = metrics - } - } - } - - for _, vm := range a.readState.VMs() { - put(map[string]float64{ - "cpu": vm.CPUPercent(), - "memory": vm.MemoryPercent(), - "disk": vm.DiskPercent(), - "netin": vm.NetIn(), - "netout": vm.NetOut(), - "diskread": vm.DiskRead(), - "diskwrite": vm.DiskWrite(), - }, vm.ID(), vm.SourceID(), strconv.Itoa(vm.VMID())) - } - for _, ct := range a.readState.Containers() { - put(map[string]float64{ - "cpu": ct.CPUPercent(), - "memory": ct.MemoryPercent(), - "disk": ct.DiskPercent(), - "netin": ct.NetIn(), - "netout": ct.NetOut(), - "diskread": ct.DiskRead(), - "diskwrite": ct.DiskWrite(), - }, ct.ID(), ct.SourceID(), strconv.Itoa(ct.VMID())) - } - for _, node := range a.readState.Nodes() { - put(map[string]float64{ - "cpu": node.CPUPercent(), - "memory": node.MemoryPercent(), - "disk": node.DiskPercent(), - }, node.ID(), node.SourceID(), node.Name()) - } - for _, sp := range a.readState.StoragePools() { - put(map[string]float64{ - "disk": sp.DiskPercent(), - "used": float64(sp.DiskUsed()), - "total": float64(sp.DiskTotal()), - }, sp.ID(), sp.SourceID(), sp.Name()) - } - return out -} - -// GetCurrentMetrics returns current metrics for a resource. -// Matches by unified ID, Proxmox source ID, VMID string, or name. -// CPU/memory/disk values are normalized to 0-100 percentage scale. -func (a *MetricsAdapter) GetCurrentMetrics(resourceID string) (map[string]float64, error) { - metrics := make(map[string]float64) - - // Check VMs - for _, vm := range a.readState.VMs() { - if vm.ID() == resourceID || vm.SourceID() == resourceID || strconv.Itoa(vm.VMID()) == resourceID { - metrics["cpu"] = vm.CPUPercent() - metrics["memory"] = vm.MemoryPercent() - metrics["disk"] = vm.DiskPercent() - metrics["netin"] = vm.NetIn() - metrics["netout"] = vm.NetOut() - metrics["diskread"] = vm.DiskRead() - metrics["diskwrite"] = vm.DiskWrite() - return metrics, nil - } - } - - // Check containers - for _, ct := range a.readState.Containers() { - if ct.ID() == resourceID || ct.SourceID() == resourceID || strconv.Itoa(ct.VMID()) == resourceID { - metrics["cpu"] = ct.CPUPercent() - metrics["memory"] = ct.MemoryPercent() - metrics["disk"] = ct.DiskPercent() - metrics["netin"] = ct.NetIn() - metrics["netout"] = ct.NetOut() - metrics["diskread"] = ct.DiskRead() - metrics["diskwrite"] = ct.DiskWrite() - return metrics, nil - } - } - - // Check nodes - for _, node := range a.readState.Nodes() { - if node.ID() == resourceID || node.SourceID() == resourceID || node.Name() == resourceID { - metrics["cpu"] = node.CPUPercent() - metrics["memory"] = node.MemoryPercent() - metrics["disk"] = node.DiskPercent() - return metrics, nil - } - } - - // Check storage - for _, sp := range a.readState.StoragePools() { - if sp.ID() == resourceID || sp.SourceID() == resourceID || sp.Name() == resourceID { - metrics["disk"] = sp.DiskPercent() - metrics["used"] = float64(sp.DiskUsed()) - metrics["total"] = float64(sp.DiskTotal()) - return metrics, nil - } - } - - return metrics, nil -} - // CommandExecutorAdapter adapts the agent execution system to remediation.CommandExecutor. // This allows the remediation engine to execute commands on targets. type CommandExecutorAdapter struct { @@ -267,69 +107,6 @@ func (e *CommandExecutionDisabledError) Error() string { return "command execution is disabled - commands must be run manually" } -// IncidentRecorderToolAdapter adapts metrics.IncidentRecorder to tools.IncidentRecorderProvider -type IncidentRecorderToolAdapter struct { - recorder IncidentRecorderSource -} - -// IncidentRecorderSource defines what we need from an incident recorder -type IncidentRecorderSource interface { - GetWindowsForResource(resourceID string, limit int) []*IncidentWindowData - GetWindow(windowID string) *IncidentWindowData -} - -// IncidentWindowData represents incident window data -type IncidentWindowData struct { - ID string - ResourceID string - ResourceName string - ResourceType string - TriggerType string - TriggerID string - StartTime time.Time - EndTime *time.Time - Status string - DataPoints []IncidentDataPointData - Summary *IncidentSummaryData -} - -// IncidentDataPointData represents a single data point -type IncidentDataPointData struct { - Timestamp time.Time - Metrics map[string]float64 -} - -// IncidentSummaryData provides summary statistics -type IncidentSummaryData struct { - Duration time.Duration - DataPoints int - Peaks map[string]float64 - Lows map[string]float64 - Averages map[string]float64 - Changes map[string]float64 -} - -// NewIncidentRecorderToolAdapter creates a new incident recorder adapter -func NewIncidentRecorderToolAdapter(recorder IncidentRecorderSource) *IncidentRecorderToolAdapter { - return &IncidentRecorderToolAdapter{recorder: recorder} -} - -// GetWindowsForResource returns incident windows for a resource -func (a *IncidentRecorderToolAdapter) GetWindowsForResource(resourceID string, limit int) []*IncidentWindowData { - if a.recorder == nil { - return nil - } - return a.recorder.GetWindowsForResource(resourceID, limit) -} - -// GetWindow returns a specific incident window -func (a *IncidentRecorderToolAdapter) GetWindow(windowID string) *IncidentWindowData { - if a.recorder == nil { - return nil - } - return a.recorder.GetWindow(windowID) -} - // EventCorrelatorToolAdapter adapts proxmox.EventCorrelator to tools.EventCorrelatorProvider type EventCorrelatorToolAdapter struct { correlator EventCorrelatorSource diff --git a/internal/ai/adapters/adapters_additional_test.go b/internal/ai/adapters/adapters_additional_test.go index b8454ea24..5f20f33cf 100644 --- a/internal/ai/adapters/adapters_additional_test.go +++ b/internal/ai/adapters/adapters_additional_test.go @@ -6,23 +6,9 @@ import ( "testing" "time" - "github.com/rcourtman/pulse-go-rewrite/internal/models" "github.com/rcourtman/pulse-go-rewrite/internal/monitoring" ) -type stubIncidentRecorder struct { - windows []*IncidentWindowData - window *IncidentWindowData -} - -func (s *stubIncidentRecorder) GetWindowsForResource(resourceID string, limit int) []*IncidentWindowData { - return s.windows -} - -func (s *stubIncidentRecorder) GetWindow(windowID string) *IncidentWindowData { - return s.window -} - type stubEventCorrelator struct { correlations []EventCorrelationData events []ProxmoxEventData @@ -69,65 +55,6 @@ func TestForecastDataAdapter_GetMetricHistory(t *testing.T) { } } -func TestMetricsAdapter_GetMonitoredResourceIDs(t *testing.T) { - state := models.StateSnapshot{ - Nodes: []models.Node{{ID: "node/pve1", Name: "pve1", Instance: "inst1"}}, - VMs: []models.VM{{ID: "qemu/100", VMID: 100, Name: "vm-1", Node: "pve1", Instance: "inst1"}}, - Containers: []models.Container{{ID: "lxc/200", VMID: 200, Name: "ct-1", Node: "pve1", Instance: "inst1"}}, - } - adapter := NewMetricsAdapter(readStateFromSnapshot(state)) - ids := adapter.GetMonitoredResourceIDs() - - // Should include both unified IDs and source IDs (3 resources × 2 IDs each = 6) - if len(ids) < 3 { - t.Fatalf("expected at least 3 IDs, got %d: %v", len(ids), ids) - } - - // Verify no empty IDs - for _, id := range ids { - if id == "" { - t.Fatalf("unexpected empty ID in %v", ids) - } - } - - // Verify source IDs are present (for pre-incident buffer compatibility) - idSet := make(map[string]bool, len(ids)) - for _, id := range ids { - idSet[id] = true - } - if !idSet["qemu/100"] { - t.Errorf("expected source ID 'qemu/100' in monitored IDs, got %v", ids) - } - if !idSet["lxc/200"] { - t.Errorf("expected source ID 'lxc/200' in monitored IDs, got %v", ids) - } - if !idSet["node/pve1"] { - t.Errorf("expected source ID 'node/pve1' in monitored IDs, got %v", ids) - } -} - -func TestIncidentRecorderToolAdapter(t *testing.T) { - adapter := NewIncidentRecorderToolAdapter(nil) - if adapter.GetWindowsForResource("res", 1) != nil { - t.Fatalf("expected nil windows for nil recorder") - } - if adapter.GetWindow("id") != nil { - t.Fatalf("expected nil window for nil recorder") - } - - recorder := &stubIncidentRecorder{ - windows: []*IncidentWindowData{{ID: "w1"}}, - window: &IncidentWindowData{ID: "w1"}, - } - adapter = NewIncidentRecorderToolAdapter(recorder) - if len(adapter.GetWindowsForResource("res", 1)) != 1 { - t.Fatalf("expected windows from recorder") - } - if adapter.GetWindow("w1") == nil { - t.Fatalf("expected window from recorder") - } -} - func TestEventCorrelatorToolAdapter(t *testing.T) { adapter := NewEventCorrelatorToolAdapter(nil) if adapter.GetCorrelationsForResource("res", time.Minute) != nil { diff --git a/internal/ai/adapters/adapters_batch_test.go b/internal/ai/adapters/adapters_batch_test.go deleted file mode 100644 index ef9b23479..000000000 --- a/internal/ai/adapters/adapters_batch_test.go +++ /dev/null @@ -1,80 +0,0 @@ -package adapters - -import ( - "reflect" - "strconv" - "testing" - - "github.com/rcourtman/pulse-go-rewrite/internal/models" -) - -// The batch method must return exactly what per-ID lookups return for every -// ID form the per-ID method matches, including colliding VMID strings where -// the VM -> container precedence decides the winner. -func TestGetCurrentMetricsBatchMatchesPerIDLookups(t *testing.T) { - state := models.StateSnapshot{ - Nodes: []models.Node{ - {ID: "node/pve1", Name: "pve1", Instance: "inst1", CPU: 0.35, Memory: models.Memory{Usage: 60}}, - }, - VMs: []models.VM{ - { - ID: "qemu/100", VMID: 100, Name: "webserver", Node: "pve1", Instance: "inst1", - CPU: 45.5, Memory: models.Memory{Usage: 72.3}, Disk: models.Disk{Usage: 55}, - NetworkIn: 1024, NetworkOut: 512, DiskRead: 2048, DiskWrite: 1024, - }, - }, - Containers: []models.Container{ - { - ID: "lxc/104", VMID: 104, Name: "auth", Node: "pve1", Instance: "inst1", - CPU: 12.5, Memory: models.Memory{Usage: 30}, Disk: models.Disk{Usage: 20}, - }, - }, - Storage: []models.Storage{ - {ID: "storage/local", Name: "local", Node: "pve1", Instance: "inst1", Usage: 41, Used: 41, Total: 100}, - }, - } - - adapter := NewMetricsAdapter(readStateFromSnapshot(state)) - batch := adapter.GetCurrentMetricsBatch() - if len(batch) == 0 { - t.Fatal("batch returned no entries for a populated state") - } - - for id := range batch { - perID, err := adapter.GetCurrentMetrics(id) - if err != nil { - t.Fatalf("GetCurrentMetrics(%q) error: %v", id, err) - } - if !reflect.DeepEqual(batch[id], perID) { - t.Fatalf("metrics diverged for %q:\nbatch: %+v\nper-ID: %+v", id, batch[id], perID) - } - } - - // Every monitored ID must be resolvable through the batch. - for _, id := range adapter.GetMonitoredResourceIDs() { - if _, ok := batch[id]; !ok { - t.Fatalf("monitored ID %q missing from batch", id) - } - } -} - -func TestGetCurrentMetricsBatchVMIDCollisionPrefersVM(t *testing.T) { - state := models.StateSnapshot{ - Nodes: []models.Node{ - {ID: "node/pve1", Name: "pve1", Instance: "inst1"}, - }, - VMs: []models.VM{ - {ID: "qemu/200", VMID: 200, Name: "vm-two-hundred", Node: "pve1", Instance: "inst1", CPU: 80}, - }, - Containers: []models.Container{ - {ID: "lxc/200", VMID: 200, Name: "ct-two-hundred", Node: "pve1", Instance: "inst1", CPU: 10}, - }, - } - - adapter := NewMetricsAdapter(readStateFromSnapshot(state)) - batch := adapter.GetCurrentMetricsBatch() - perID, _ := adapter.GetCurrentMetrics(strconv.Itoa(200)) - if !reflect.DeepEqual(batch["200"], perID) { - t.Fatalf("VMID collision winner diverged:\nbatch: %+v\nper-ID: %+v", batch["200"], perID) - } -} diff --git a/internal/ai/adapters/adapters_test.go b/internal/ai/adapters/adapters_test.go index 6112fc783..c48344c25 100644 --- a/internal/ai/adapters/adapters_test.go +++ b/internal/ai/adapters/adapters_test.go @@ -2,21 +2,10 @@ package adapters import ( "context" - "fmt" "testing" "time" - - "github.com/rcourtman/pulse-go-rewrite/internal/models" - "github.com/rcourtman/pulse-go-rewrite/internal/unifiedresources" ) -// readStateFromSnapshot creates a ReadState from a models.StateSnapshot for testing. -func readStateFromSnapshot(snapshot models.StateSnapshot) unifiedresources.ReadState { - rr := unifiedresources.NewRegistry(nil) - rr.IngestSnapshot(snapshot) - return rr -} - func TestForecastDataAdapter_NilHistory(t *testing.T) { adapter := NewForecastDataAdapter(nil) if adapter != nil { @@ -24,295 +13,6 @@ func TestForecastDataAdapter_NilHistory(t *testing.T) { } } -func TestMetricsAdapter_GetCurrentMetrics_VM(t *testing.T) { - state := models.StateSnapshot{ - Nodes: []models.Node{ - {ID: "node/pve1", Name: "pve1", Instance: "inst1"}, - }, - VMs: []models.VM{ - { - ID: "qemu/100", - VMID: 100, - Name: "webserver", - Node: "pve1", - Instance: "inst1", - CPU: 45.5, - Memory: models.Memory{ - Usage: 72.3, - Used: 1024, - Total: 2048, - }, - Disk: models.Disk{ - Usage: 55.0, - Used: 5000, - Total: 10000, - }, - NetworkIn: 1024000, - NetworkOut: 512000, - DiskRead: 2048000, - DiskWrite: 1024000, - }, - }, - } - - rs := readStateFromSnapshot(state) - adapter := NewMetricsAdapter(rs) - - // Get the unified resource ID from ReadState - vms := rs.VMs() - if len(vms) != 1 { - t.Fatalf("expected 1 VM, got %d", len(vms)) - } - vmID := vms[0].ID() - - metrics, err := adapter.GetCurrentMetrics(vmID) - if err != nil { - t.Errorf("Unexpected error: %v", err) - return - } - - if metrics["cpu"] != 45.5 { - t.Errorf("Expected CPU 45.5, got %f", metrics["cpu"]) - } - if metrics["memory"] != 72.3 { - t.Errorf("Expected memory 72.3, got %f", metrics["memory"]) - } - if metrics["disk"] != 55.0 { - t.Errorf("Expected disk 55.0, got %f", metrics["disk"]) - } - if metrics["netin"] != 1024000 { - t.Errorf("Expected netin 1024000, got %f", metrics["netin"]) - } -} - -func TestMetricsAdapter_GetCurrentMetrics_Container(t *testing.T) { - state := models.StateSnapshot{ - Nodes: []models.Node{ - {ID: "node/pve1", Name: "pve1", Instance: "inst1"}, - }, - Containers: []models.Container{ - { - ID: "lxc/101", - VMID: 101, - Name: "container1", - Node: "pve1", - Instance: "inst1", - CPU: 25.0, - Memory: models.Memory{ - Usage: 45.0, - Used: 512, - Total: 1024, - }, - Disk: models.Disk{ - Usage: 30.0, - Used: 3000, - Total: 10000, - }, - NetworkIn: 500000, - NetworkOut: 250000, - DiskRead: 1000000, - DiskWrite: 500000, - }, - }, - } - - rs := readStateFromSnapshot(state) - adapter := NewMetricsAdapter(rs) - - containers := rs.Containers() - if len(containers) != 1 { - t.Fatalf("expected 1 container, got %d", len(containers)) - } - ctID := containers[0].ID() - - metrics, err := adapter.GetCurrentMetrics(ctID) - if err != nil { - t.Errorf("Unexpected error: %v", err) - return - } - - if metrics["cpu"] != 25.0 { - t.Errorf("Expected CPU 25.0, got %f", metrics["cpu"]) - } - if metrics["memory"] != 45.0 { - t.Errorf("Expected memory 45.0, got %f", metrics["memory"]) - } -} - -func TestMetricsAdapter_GetCurrentMetrics_Node(t *testing.T) { - state := models.StateSnapshot{ - Nodes: []models.Node{ - { - ID: "node/pve1", - Name: "pve1", - Instance: "inst1", - CPU: 25.5, - Memory: models.Memory{ - Usage: 65.0, - Used: 6500, - Total: 10000, - }, - Disk: models.Disk{ - Usage: 40.0, - Used: 4000, - Total: 10000, - }, - }, - }, - } - - rs := readStateFromSnapshot(state) - adapter := NewMetricsAdapter(rs) - - nodes := rs.Nodes() - if len(nodes) != 1 { - t.Fatalf("expected 1 node, got %d", len(nodes)) - } - nodeID := nodes[0].ID() - - metrics, err := adapter.GetCurrentMetrics(nodeID) - if err != nil { - t.Errorf("Unexpected error: %v", err) - return - } - - if metrics["cpu"] != 25.5 { - t.Errorf("Expected CPU 25.5, got %f", metrics["cpu"]) - } - if metrics["memory"] != 65.0 { - t.Errorf("Expected memory 65.0, got %f", metrics["memory"]) - } - if metrics["disk"] != 40.0 { - t.Errorf("Expected disk 40.0, got %f", metrics["disk"]) - } -} - -func TestMetricsAdapter_GetCurrentMetrics_NodeByName(t *testing.T) { - state := models.StateSnapshot{ - Nodes: []models.Node{ - { - ID: "node/pve1", - Name: "pve1", - Instance: "inst1", - CPU: 25.5, - Memory: models.Memory{ - Usage: 65.0, - Used: 6500, - Total: 10000, - }, - Disk: models.Disk{ - Usage: 40.0, - Used: 4000, - Total: 10000, - }, - }, - }, - } - - adapter := NewMetricsAdapter(readStateFromSnapshot(state)) - - // Node lookup by name should still work - metrics, err := adapter.GetCurrentMetrics("pve1") - if err != nil { - t.Errorf("Unexpected error: %v", err) - return - } - - if metrics["cpu"] != 25.5 { - t.Errorf("Expected CPU 25.5 when matching by name, got %f", metrics["cpu"]) - } -} - -func TestMetricsAdapter_GetCurrentMetrics_Storage(t *testing.T) { - state := models.StateSnapshot{ - Nodes: []models.Node{ - {ID: "node/pve1", Name: "pve1", Instance: "inst1"}, - }, - Storage: []models.Storage{ - { - ID: "storage/local-zfs", - Name: "local-zfs", - Node: "pve1", - Instance: "inst1", - Used: 50000000000, - Total: 100000000000, - Usage: 50.0, - }, - }, - } - - rs := readStateFromSnapshot(state) - adapter := NewMetricsAdapter(rs) - - pools := rs.StoragePools() - if len(pools) != 1 { - t.Fatalf("expected 1 storage pool, got %d", len(pools)) - } - storageID := pools[0].ID() - - metrics, err := adapter.GetCurrentMetrics(storageID) - if err != nil { - t.Errorf("Unexpected error: %v", err) - return - } - - if metrics["disk"] != 50.0 { - t.Errorf("Expected disk 50.0, got %f", metrics["disk"]) - } - if metrics["used"] != 50000000000 { - t.Errorf("Expected used 50000000000, got %f", metrics["used"]) - } - if metrics["total"] != 100000000000 { - t.Errorf("Expected total 100000000000, got %f", metrics["total"]) - } -} - -func TestMetricsAdapter_GetCurrentMetrics_StorageByName(t *testing.T) { - state := models.StateSnapshot{ - Nodes: []models.Node{ - {ID: "node/pve1", Name: "pve1", Instance: "inst1"}, - }, - Storage: []models.Storage{ - { - ID: "storage/local-zfs", - Name: "local-zfs", - Node: "pve1", - Instance: "inst1", - Used: 50000000000, - Total: 100000000000, - Usage: 50.0, - }, - }, - } - - adapter := NewMetricsAdapter(readStateFromSnapshot(state)) - - // Storage lookup by name should still work - metrics, err := adapter.GetCurrentMetrics("local-zfs") - if err != nil { - t.Errorf("Unexpected error: %v", err) - return - } - - if metrics["disk"] != 50.0 { - t.Errorf("Expected disk 50.0, got %f", metrics["disk"]) - } -} - -func TestMetricsAdapter_GetCurrentMetrics_NotFound(t *testing.T) { - state := models.StateSnapshot{} - - adapter := NewMetricsAdapter(readStateFromSnapshot(state)) - - metrics, err := adapter.GetCurrentMetrics("nonexistent") - if err != nil { - t.Errorf("Unexpected error: %v", err) - } - if len(metrics) != 0 { - t.Errorf("Expected empty metrics, got %d entries", len(metrics)) - } -} - func TestCommandExecutorAdapter_Disabled(t *testing.T) { adapter := NewCommandExecutorAdapter() @@ -349,131 +49,3 @@ func TestCommandExecutionDisabledError_Message(t *testing.T) { t.Errorf("Unexpected error message: %s", msg) } } - -func TestMetricsAdapter_NilReadState(t *testing.T) { - adapter := NewMetricsAdapter(nil) - if adapter != nil { - t.Error("Expected nil adapter for nil ReadState") - } -} - -func TestMetricsAdapter_VMIDMatch(t *testing.T) { - state := models.StateSnapshot{ - Nodes: []models.Node{ - {ID: "node/pve1", Name: "pve1", Instance: "inst1"}, - }, - VMs: []models.VM{ - { - ID: "qemu/100", - VMID: 100, - Name: "webserver", - Node: "pve1", - Instance: "inst1", - CPU: 45.5, - Memory: models.Memory{ - Usage: 72.3, - Used: 1024, - Total: 2048, - }, - Disk: models.Disk{ - Usage: 55.0, - Used: 5000, - Total: 10000, - }, - }, - }, - } - - adapter := NewMetricsAdapter(readStateFromSnapshot(state)) - - // Test lookup by VMID string - metrics, err := adapter.GetCurrentMetrics(fmt.Sprintf("%d", 100)) - if err != nil { - t.Errorf("Unexpected error: %v", err) - return - } - - if metrics["cpu"] != 45.5 { - t.Errorf("Expected CPU 45.5 when matching by VMID, got %f", metrics["cpu"]) - } -} - -func TestMetricsAdapter_IDConsistency(t *testing.T) { - // Verify that GetMonitoredResourceIDs returns IDs that work with GetCurrentMetrics - state := models.StateSnapshot{ - Nodes: []models.Node{ - {ID: "node/pve1", Name: "pve1", Instance: "inst1"}, - }, - VMs: []models.VM{ - { - ID: "qemu/100", VMID: 100, Name: "vm1", Node: "pve1", Instance: "inst1", - CPU: 50.0, Memory: models.Memory{Usage: 60.0, Used: 600, Total: 1000}, - Disk: models.Disk{Usage: 70.0, Used: 700, Total: 1000}, - }, - }, - } - - adapter := NewMetricsAdapter(readStateFromSnapshot(state)) - ids := adapter.GetMonitoredResourceIDs() - if len(ids) < 1 { - t.Fatalf("expected at least 1 ID, got %d", len(ids)) - } - - // Each ID from GetMonitoredResourceIDs should be usable with GetCurrentMetrics - foundVM := false - for _, id := range ids { - metrics, err := adapter.GetCurrentMetrics(id) - if err != nil { - t.Errorf("GetCurrentMetrics(%q) error: %v", id, err) - continue - } - if cpu, ok := metrics["cpu"]; ok && cpu == 50.0 { - foundVM = true - } - } - if !foundVM { - t.Error("Expected to find VM metrics via GetMonitoredResourceIDs() IDs") - } -} - -func TestMetricsAdapter_SourceIDMatch(t *testing.T) { - // Verify that GetCurrentMetrics can find resources by their Proxmox source ID - state := models.StateSnapshot{ - Nodes: []models.Node{ - {ID: "node/pve1", Name: "pve1", Instance: "inst1"}, - }, - VMs: []models.VM{ - { - ID: "qemu/100", VMID: 100, Name: "webserver", Node: "pve1", Instance: "inst1", - CPU: 45.5, Memory: models.Memory{Usage: 72.3, Used: 1024, Total: 2048}, - Disk: models.Disk{Usage: 55.0, Used: 5000, Total: 10000}, - }, - }, - Storage: []models.Storage{ - { - ID: "storage/local-zfs", Name: "local-zfs", Node: "pve1", Instance: "inst1", - Used: 50000000000, Total: 100000000000, Usage: 50.0, - }, - }, - } - - adapter := NewMetricsAdapter(readStateFromSnapshot(state)) - - // Lookup VM by Proxmox source ID - metrics, err := adapter.GetCurrentMetrics("qemu/100") - if err != nil { - t.Fatalf("Unexpected error: %v", err) - } - if metrics["cpu"] != 45.5 { - t.Errorf("Expected CPU 45.5 via source ID, got %f", metrics["cpu"]) - } - - // Lookup storage by Proxmox source ID - metrics, err = adapter.GetCurrentMetrics("storage/local-zfs") - if err != nil { - t.Fatalf("Unexpected error: %v", err) - } - if metrics["disk"] != 50.0 { - t.Errorf("Expected disk 50.0 via source ID, got %f", metrics["disk"]) - } -} diff --git a/internal/ai/chat/service.go b/internal/ai/chat/service.go index 8309be743..ad92cdc60 100644 --- a/internal/ai/chat/service.go +++ b/internal/ai/chat/service.go @@ -60,7 +60,7 @@ type ( AgentProfileManager = tools.AgentProfileManager FindingsManager = tools.FindingsManager MetadataUpdater = tools.MetadataUpdater - IncidentRecorderProvider = tools.IncidentRecorderProvider + IncidentArchiveProvider = tools.IncidentArchiveProvider EventCorrelatorProvider = tools.EventCorrelatorProvider KnowledgeStoreProvider = tools.KnowledgeStoreProvider AssistantDiscoveryProvider = tools.DiscoveryProvider @@ -1418,15 +1418,15 @@ func marshalAssistantInventoryTopologyContext(topology tools.TopologyResponse) ( } for _, node := range topology.Proxmox.Nodes { nodeContext := assistantInventoryProxmoxNode{ - AnswerLabel: assistantInventoryNodeAnswerLabel(node.Name), - Name: node.Name, - Status: node.Status, - AgentConnected: node.AgentConnected, - CanExecute: node.CanExecute, - VMCount: node.VMCount, - ContainerCount: node.ContainerCount, - VMs: make([]assistantInventoryWorkload, 0, len(node.VMs)), - Containers: make([]assistantInventoryWorkload, 0, len(node.Containers)), + AnswerLabel: assistantInventoryNodeAnswerLabel(node.Name), + Name: node.Name, + Status: node.Status, + CommandAgentConnected: node.CommandAgentConnected, + CanExecute: node.CanExecute, + VMCount: node.VMCount, + ContainerCount: node.ContainerCount, + VMs: make([]assistantInventoryWorkload, 0, len(node.VMs)), + Containers: make([]assistantInventoryWorkload, 0, len(node.Containers)), } for _, vm := range node.VMs { nodeContext.VMs = append(nodeContext.VMs, assistantInventoryWorkload{ @@ -1452,14 +1452,14 @@ func marshalAssistantInventoryTopologyContext(topology tools.TopologyResponse) ( } for _, host := range topology.Docker.Hosts { hostContext := assistantInventoryDockerHost{ - AnswerLabel: firstNonEmptyString(host.DisplayName, host.Hostname), - Hostname: host.Hostname, - DisplayName: host.DisplayName, - AgentConnected: host.AgentConnected, - CanExecute: host.CanExecute, - ContainerCount: host.ContainerCount, - RunningCount: host.RunningCount, - Containers: make([]assistantInventoryAppContainer, 0, len(host.Containers)), + AnswerLabel: firstNonEmptyString(host.DisplayName, host.Hostname), + Hostname: host.Hostname, + DisplayName: host.DisplayName, + CommandAgentConnected: host.CommandAgentConnected, + CanExecute: host.CanExecute, + ContainerCount: host.ContainerCount, + RunningCount: host.RunningCount, + Containers: make([]assistantInventoryAppContainer, 0, len(host.Containers)), } for _, container := range host.Containers { hostContext.Containers = append(hostContext.Containers, assistantInventoryAppContainer{ @@ -1547,15 +1547,15 @@ type assistantInventoryKubernetesTopology struct { } type assistantInventoryProxmoxNode struct { - AnswerLabel string `json:"answer_label"` - Name string `json:"name"` - Status string `json:"status"` - AgentConnected bool `json:"agent_connected,omitempty"` - CanExecute bool `json:"can_execute,omitempty"` - VMCount int `json:"vm_count"` - ContainerCount int `json:"container_count"` - VMs []assistantInventoryWorkload `json:"vms"` - Containers []assistantInventoryWorkload `json:"containers"` + AnswerLabel string `json:"answer_label"` + Name string `json:"name"` + Status string `json:"status"` + CommandAgentConnected *bool `json:"command_agent_connected,omitempty"` + CanExecute *bool `json:"can_execute,omitempty"` + VMCount int `json:"vm_count"` + ContainerCount int `json:"container_count"` + VMs []assistantInventoryWorkload `json:"vms"` + Containers []assistantInventoryWorkload `json:"containers"` } type assistantInventoryWorkload struct { @@ -1568,14 +1568,14 @@ type assistantInventoryWorkload struct { } type assistantInventoryDockerHost struct { - AnswerLabel string `json:"answer_label"` - Hostname string `json:"hostname"` - DisplayName string `json:"display_name,omitempty"` - AgentConnected bool `json:"agent_connected,omitempty"` - CanExecute bool `json:"can_execute,omitempty"` - ContainerCount int `json:"container_count"` - RunningCount int `json:"running_count"` - Containers []assistantInventoryAppContainer `json:"containers"` + AnswerLabel string `json:"answer_label"` + Hostname string `json:"hostname"` + DisplayName string `json:"display_name,omitempty"` + CommandAgentConnected *bool `json:"command_agent_connected,omitempty"` + CanExecute *bool `json:"can_execute,omitempty"` + ContainerCount int `json:"container_count"` + RunningCount int `json:"running_count"` + Containers []assistantInventoryAppContainer `json:"containers"` } type assistantInventoryAppContainer struct { @@ -3598,11 +3598,11 @@ func (s *Service) SetMetadataUpdater(updater MetadataUpdater) { } } -func (s *Service) SetIncidentRecorderProvider(provider IncidentRecorderProvider) { +func (s *Service) SetIncidentArchiveProvider(provider IncidentArchiveProvider) { s.mu.Lock() defer s.mu.Unlock() if s.executor != nil { - s.executor.SetIncidentRecorderProvider(provider) + s.executor.SetIncidentArchiveProvider(provider) } } diff --git a/internal/ai/chat/service_additional_test.go b/internal/ai/chat/service_additional_test.go index 6e8505d38..4a3ed1c01 100644 --- a/internal/ai/chat/service_additional_test.go +++ b/internal/ai/chat/service_additional_test.go @@ -16,7 +16,7 @@ func TestServiceSettersAndAutonomousMode(t *testing.T) { agenticLoop: loop, } - service.SetIncidentRecorderProvider(nil) + service.SetIncidentArchiveProvider(nil) service.SetEventCorrelatorProvider(nil) service.SetKnowledgeStoreProvider(nil) diff --git a/internal/ai/chat/service_command_connection_test.go b/internal/ai/chat/service_command_connection_test.go new file mode 100644 index 000000000..b877879db --- /dev/null +++ b/internal/ai/chat/service_command_connection_test.go @@ -0,0 +1,35 @@ +package chat + +import ( + "encoding/json" + "strings" + "testing" + + "github.com/rcourtman/pulse-go-rewrite/internal/models" + "github.com/rcourtman/pulse-go-rewrite/internal/unifiedresources" +) + +func TestAssistantInventoryDoesNotInventCommandConnectionObservations(t *testing.T) { + registry := unifiedresources.NewRegistry(nil) + registry.IngestSnapshot(models.StateSnapshot{ + Nodes: []models.Node{{ID: "node-one", Name: "node-one", Status: "online"}}, + DockerHosts: []models.DockerHost{{ID: "host-one", Hostname: "observed-host", Status: "online", Containers: []models.DockerContainer{{ID: "app-one", Name: "observed-app", State: "running"}}}}, + }) + raw, err := marshalAssistantInventoryTopologyContextFromReadState(registry) + if err != nil { + t.Fatal(err) + } + for _, field := range []string{"agent_connected", "command_agent_connected", "can_execute", "nodes_with_agents", "docker_hosts_with_agents", "nodes_with_command_agents", "docker_hosts_with_command_agents"} { + if strings.Contains(raw, `"`+field+`"`) { + t.Fatalf("inventory seed invented %s without observing command connections: %s", field, raw) + } + } + var decoded map[string]any + if err := json.Unmarshal([]byte(raw), &decoded); err != nil { + t.Fatal(err) + } + host := decoded["docker"].(map[string]any)["hosts"].([]any)[0].(map[string]any) + if host["hostname"] != "observed-host" || host["container_count"] != float64(1) { + t.Fatalf("monitoring inventory was lost: %+v", host) + } +} diff --git a/internal/ai/chat/types.go b/internal/ai/chat/types.go index 2f307942c..4cc102d3f 100644 --- a/internal/ai/chat/types.go +++ b/internal/ai/chat/types.go @@ -135,38 +135,13 @@ func (m Message) ClientSafe() Message { return m } -// ToolCall represents a tool invocation -type ToolCall struct { - ID string `json:"id"` - Name string `json:"name"` - Input map[string]interface{} `json:"input"` - Output string `json:"output,omitempty"` - Success *bool `json:"success,omitempty"` - ThoughtSignature json.RawMessage `json:"thought_signature,omitempty"` -} +// ToolCall is the canonical result-bearing product transcript call. +type ToolCall = agentcapabilities.TranscriptToolCall func EmptyToolCall() ToolCall { return ToolCall{}.NormalizeCollections() } -func (t ToolCall) NormalizeCollections() ToolCall { - providerCall := agentcapabilities.ProviderToolCall{ - ID: t.ID, - Name: t.Name, - Input: t.Input, - ThoughtSignature: t.ThoughtSignature, - }.NormalizeCollections() - t.ID = providerCall.ID - t.Name = providerCall.Name - t.Input = providerCall.Input - t.ThoughtSignature = providerCall.ThoughtSignature - if t.Success != nil { - success := *t.Success - t.Success = &success - } - return t -} - // ToolCallFromProvider stores a provider-facing tool call in the richer // Assistant transcript shape used for in-app history. func ToolCallFromProvider(tc agentcapabilities.ProviderToolCall) ToolCall { @@ -183,19 +158,6 @@ func ToolCallFromProvider(tc agentcapabilities.ProviderToolCall) ToolCall { }.NormalizeCollections() } -// ProviderToolCall projects a stored Assistant transcript call back to the -// shared provider-facing shape, deliberately excluding in-app output/success -// display fields. -func (t ToolCall) ProviderToolCall() agentcapabilities.ProviderToolCall { - t = t.NormalizeCollections() - return agentcapabilities.ProviderToolCall{ - ID: t.ID, - Name: t.Name, - Input: t.Input, - ThoughtSignature: t.ThoughtSignature, - }.NormalizeCollections() -} - // ToolResult represents the result of a tool execution. It aliases the shared // Pulse Intelligence provider-result shape so stored Assistant transcripts and // provider turns do not drift on tool result JSON. diff --git a/internal/ai/cost/pricing.go b/internal/ai/cost/pricing.go index 7ac56bbe2..2f43d6b05 100644 --- a/internal/ai/cost/pricing.go +++ b/internal/ai/cost/pricing.go @@ -82,11 +82,17 @@ var providerPrices = map[string][]modelPrice{ flatPriceAsOf("anthropic/claude-opus-4.8", 5.00, 25.00, "2026-07-14"), flatPriceAsOf("anthropic/claude-sonnet-5", 2.00, 10.00, "2026-07-14"), flatPriceAsOf("deepseek/deepseek-v4-flash", 0.09, 0.18, "2026-07-14"), + // Introductory standard rates through 2026-12-31. Recheck when the + // published standard price changes on 2027-01-01. Batch/alias routes + // are deliberately not covered by this exact model ID. + flatPriceAsOf("google/gemini-3.8-flash", 0.75, 3.75, "2026-09-06"), flatPriceAsOf("nvidia/nemotron-3.5-lightning:free", 0, 0, "2026-08-14"), flatPriceAsOf("nvidia/nemotron-3-super-120b-a12b:free", 0, 0, "2026-08-14"), flatPriceAsOf("nvidia/nemotron-3-ultra-550b-a55b:free", 0, 0, "2026-08-15"), }, "gemini": { + // Standard introductory rates, verified 2026-09-06. Recheck 2027-01-01. + flatPriceAsOf("gemini-3.8-flash", 0.75, 3.75, "2026-09-06"), // Gemini Developer API standard paid-tier pricing, checked from // https://ai.google.dev/gemini-api/docs/pricing on 2026-06-04. flatPrice("gemini-3.5-flash*", 1.50, 9.00), diff --git a/internal/ai/cost/pricing_gemini_test.go b/internal/ai/cost/pricing_gemini_test.go new file mode 100644 index 000000000..6ea1ee23d --- /dev/null +++ b/internal/ai/cost/pricing_gemini_test.go @@ -0,0 +1,46 @@ +package cost + +import ( + "math" + "testing" +) + +func TestGemini38FlashReviewedRoutePricing(t *testing.T) { + for _, route := range []struct{ provider, model string }{ + {"gemini", "gemini-3.8-flash"}, + {"openrouter", "google/gemini-3.8-flash"}, + } { + t.Run(route.provider, func(t *testing.T) { + provider, model := ResolveProviderAndModel(route.provider, route.provider+":"+route.model, "gemini-3.8-flash") + if provider != route.provider || model != route.model { + t.Fatalf("usage lost the requested billing route: %s:%s", provider, model) + } + // Token counts from the live unhealthy-container qualification. + usd, known, price := EstimateUSD(provider, model, 16068, 778) + if !known || math.Abs(usd-0.0149685) > 1e-10 { + t.Fatalf("live usage estimate = %f, known=%t", usd, known) + } + if price.InputUSDPerMTok != 0.75 || price.OutputUSDPerMTok != 3.75 || price.AsOf != "2026-09-06" { + t.Fatalf("reviewed standard rates/date missing: %+v", price) + } + usd, known, _ = EstimateUSD(provider, model, 0, 0) + if !known || usd != 0 { + t.Fatalf("zero observed usage = %f, known=%t", usd, known) + } + }) + } +} + +func TestGemini38OpenRouterPricingDoesNotGuessVariantRates(t *testing.T) { + for _, model := range []string{ + "google/gemini-3.8-flash:batch", + "google/gemini-3.8-flash:free", + "google/gemini-3.8-flash-preview", + "google/gemini-3.8-flash-cyber", + "google/gemini-3.9-flash", + } { + if usd, known, _ := EstimateUSD("openrouter", model, 16068, 778); known || usd != 0 { + t.Errorf("unreviewed route %q received an estimate: %f, known=%t", model, usd, known) + } + } +} diff --git a/internal/ai/incident_coordinator.go b/internal/ai/incident_coordinator.go deleted file mode 100644 index 3ab75b7a2..000000000 --- a/internal/ai/incident_coordinator.go +++ /dev/null @@ -1,325 +0,0 @@ -// Package ai provides AI-powered infrastructure analysis. -package ai - -import ( - "sync" - "time" - - "github.com/rcourtman/pulse-go-rewrite/internal/ai/memory" - "github.com/rcourtman/pulse-go-rewrite/internal/alerts" - "github.com/rcourtman/pulse-go-rewrite/internal/metrics" - "github.com/rs/zerolog/log" -) - -// IncidentCoordinatorConfig configures the incident coordinator -type IncidentCoordinatorConfig struct { - PreBuffer time.Duration // History to capture before incident (default: 5 min) - PostDuration time.Duration // How long to record after trigger (default: 10 min) - MaxConcurrent int // Maximum concurrent incident recordings (default: 50) - EnableRecorder bool // Whether to enable high-frequency recording -} - -// DefaultIncidentCoordinatorConfig returns sensible defaults -func DefaultIncidentCoordinatorConfig() IncidentCoordinatorConfig { - return IncidentCoordinatorConfig{ - PreBuffer: 5 * time.Minute, - PostDuration: 10 * time.Minute, - MaxConcurrent: 50, - EnableRecorder: true, - } -} - -// IncidentCoordinator coordinates incident recording between the metrics.IncidentRecorder -// (for high-frequency data capture) and memory.IncidentStore (for incident timeline tracking). -type IncidentCoordinator struct { - mu sync.RWMutex - - config IncidentCoordinatorConfig - - // Components - recorder *metrics.IncidentRecorder // High-frequency metrics capture - incidentStore *memory.IncidentStore // Incident timeline tracking - - // Active incidents - maps alert ID to window ID - activeIncidents map[string]activeIncident - - // Control - running bool -} - -type activeIncident struct { - windowID string - resourceID string - startedAt time.Time - stopTimer *time.Timer // Timer to auto-stop recording after post-duration -} - -// NewIncidentCoordinator creates a new incident coordinator -func NewIncidentCoordinator(cfg IncidentCoordinatorConfig) *IncidentCoordinator { - if cfg.PreBuffer <= 0 { - cfg.PreBuffer = 5 * time.Minute - } - if cfg.PostDuration <= 0 { - cfg.PostDuration = 10 * time.Minute - } - if cfg.MaxConcurrent <= 0 { - cfg.MaxConcurrent = 50 - } - - return &IncidentCoordinator{ - config: cfg, - activeIncidents: make(map[string]activeIncident), - } -} - -// SetRecorder sets the metrics incident recorder -func (c *IncidentCoordinator) SetRecorder(recorder *metrics.IncidentRecorder) { - c.mu.Lock() - defer c.mu.Unlock() - c.recorder = recorder -} - -// SetIncidentStore sets the incident timeline store -func (c *IncidentCoordinator) SetIncidentStore(store *memory.IncidentStore) { - c.mu.Lock() - defer c.mu.Unlock() - c.incidentStore = store -} - -// Start starts the incident coordinator -func (c *IncidentCoordinator) Start() { - c.mu.Lock() - defer c.mu.Unlock() - if c.running { - return - } - c.running = true - log.Info().Msg("incident coordinator started") -} - -// Stop stops the incident coordinator -func (c *IncidentCoordinator) Stop() { - c.mu.Lock() - defer c.mu.Unlock() - if !c.running { - return - } - c.running = false - - // Stop all active timers - for _, inc := range c.activeIncidents { - if inc.stopTimer != nil { - inc.stopTimer.Stop() - } - } - c.activeIncidents = make(map[string]activeIncident) - - log.Info().Msg("incident coordinator stopped") -} - -// OnAlertFired is called when an alert fires - starts incident recording -func (c *IncidentCoordinator) OnAlertFired(alert *alerts.Alert) { - if alert == nil { - return - } - - c.mu.Lock() - defer c.mu.Unlock() - - if !c.running { - return - } - - // Check if we already have an active incident for this alert - if _, exists := c.activeIncidents[alert.ID]; exists { - log.Debug(). - Str("alert_identifier", alert.ID). - Msg("Incident already being recorded for this alert") - return - } - - // Check concurrent limit - if len(c.activeIncidents) >= c.config.MaxConcurrent { - log.Warn(). - Str("alert_identifier", alert.ID). - Int("active_count", len(c.activeIncidents)). - Msg("Incident coordinator at capacity, skipping new incident") - return - } - - // Start high-frequency recording if enabled and recorder available - var windowID string - if c.config.EnableRecorder && c.recorder != nil { - windowID = c.recorder.StartRecording( - alert.ResourceID, - alert.ResourceName, - "", // resourceType - we don't always have this - "alert", - alert.ID, - ) - } - - // Record in incident store - if c.incidentStore != nil { - c.incidentStore.RecordAlertFired(alert) - } - - // Track the active incident - inc := activeIncident{ - windowID: windowID, - resourceID: alert.ResourceID, - startedAt: time.Now(), - } - - c.activeIncidents[alert.ID] = inc - - log.Info(). - Str("alert_identifier", alert.ID). - Str("resource_id", alert.ResourceID). - Str("window_id", windowID). - Msg("Incident coordinator: Started incident recording") -} - -// OnAlertCleared is called when an alert clears - schedules recording stop -func (c *IncidentCoordinator) OnAlertCleared(alert *alerts.Alert) { - if alert == nil { - return - } - - c.mu.Lock() - - inc, exists := c.activeIncidents[alert.ID] - if !exists { - c.mu.Unlock() - return - } - - // Record resolution in incident store - if c.incidentStore != nil { - c.incidentStore.RecordAlertResolved(alert, time.Now()) - } - - // If no recorder or no window, just clean up immediately - if c.recorder == nil || inc.windowID == "" { - delete(c.activeIncidents, alert.ID) - c.mu.Unlock() - return - } - - // Schedule stop after post-duration to capture post-incident data - alertID := alert.ID - timer := time.AfterFunc(c.config.PostDuration, func() { - c.stopIncidentRecording(alertID) - }) - - inc.stopTimer = timer - c.activeIncidents[alert.ID] = inc - c.mu.Unlock() - - log.Info(). - Str("alert_identifier", alert.ID). - Str("window_id", inc.windowID). - Dur("post_duration", c.config.PostDuration). - Msg("Incident coordinator: Alert cleared, scheduled recording stop") -} - -// stopIncidentRecording stops recording for a specific alert -func (c *IncidentCoordinator) stopIncidentRecording(alertID string) { - c.mu.Lock() - defer c.mu.Unlock() - - inc, exists := c.activeIncidents[alertID] - if !exists { - return - } - - // Stop the recorder - if c.recorder != nil && inc.windowID != "" { - c.recorder.StopRecording(inc.windowID) - } - - // Clean up - if inc.stopTimer != nil { - inc.stopTimer.Stop() - } - delete(c.activeIncidents, alertID) - - log.Info(). - Str("alert_identifier", alertID). - Str("window_id", inc.windowID). - Msg("Incident coordinator: Stopped incident recording") -} - -// OnAnomalyDetected is called when an anomaly is detected - starts focused recording -func (c *IncidentCoordinator) OnAnomalyDetected(resourceID, resourceType, metric string, severity string) { - c.mu.Lock() - defer c.mu.Unlock() - - if !c.running || !c.config.EnableRecorder || c.recorder == nil { - return - } - - // Create a pseudo-alert ID for the anomaly - anomalyID := "anomaly-" + resourceID + "-" + metric - - // Check if we already have an active incident for this - if _, exists := c.activeIncidents[anomalyID]; exists { - return - } - - // Check concurrent limit - if len(c.activeIncidents) >= c.config.MaxConcurrent { - return - } - - // Start recording - windowID := c.recorder.StartRecording( - resourceID, - "", // name - resourceType, - "anomaly", - anomalyID, - ) - - // Track the active incident - c.activeIncidents[anomalyID] = activeIncident{ - windowID: windowID, - resourceID: resourceID, - startedAt: time.Now(), - } - - // Schedule auto-stop after post-duration (anomalies don't have "clear" events) - timer := time.AfterFunc(c.config.PostDuration, func() { - c.stopIncidentRecording(anomalyID) - }) - c.activeIncidents[anomalyID] = activeIncident{ - windowID: windowID, - resourceID: resourceID, - startedAt: time.Now(), - stopTimer: timer, - } - - log.Info(). - Str("resource_id", resourceID). - Str("metric", metric). - Str("severity", severity). - Str("window_id", windowID). - Msg("Incident coordinator: Started anomaly recording") -} - -// GetActiveIncidentCount returns the number of active incidents being recorded -func (c *IncidentCoordinator) GetActiveIncidentCount() int { - c.mu.RLock() - defer c.mu.RUnlock() - return len(c.activeIncidents) -} - -// GetRecordingWindowID returns the recording window ID for an alert -func (c *IncidentCoordinator) GetRecordingWindowID(alertID string) string { - c.mu.RLock() - defer c.mu.RUnlock() - if inc, exists := c.activeIncidents[alertID]; exists { - return inc.windowID - } - return "" -} diff --git a/internal/ai/incident_coordinator_additional_test.go b/internal/ai/incident_coordinator_additional_test.go deleted file mode 100644 index 11b71b39f..000000000 --- a/internal/ai/incident_coordinator_additional_test.go +++ /dev/null @@ -1,179 +0,0 @@ -package ai - -import ( - "testing" - "time" - - "github.com/rcourtman/pulse-go-rewrite/internal/ai/memory" - "github.com/rcourtman/pulse-go-rewrite/internal/alerts" - "github.com/rcourtman/pulse-go-rewrite/internal/metrics" -) - -func TestNewIncidentCoordinator_DefaultFallbacks(t *testing.T) { - cfg := IncidentCoordinatorConfig{ - PreBuffer: -1 * time.Second, - PostDuration: 0, - MaxConcurrent: 0, - } - - coord := NewIncidentCoordinator(cfg) - defaults := DefaultIncidentCoordinatorConfig() - - if coord.config.PreBuffer != defaults.PreBuffer { - t.Fatalf("expected default pre-buffer %v, got %v", defaults.PreBuffer, coord.config.PreBuffer) - } - if coord.config.PostDuration != defaults.PostDuration { - t.Fatalf("expected default post-duration %v, got %v", defaults.PostDuration, coord.config.PostDuration) - } - if coord.config.MaxConcurrent != defaults.MaxConcurrent { - t.Fatalf("expected default max concurrent %d, got %d", defaults.MaxConcurrent, coord.config.MaxConcurrent) - } - if coord.activeIncidents == nil { - t.Fatal("expected active incident map to be initialized") - } -} - -func TestIncidentCoordinator_OnAlertFired_RequiresRunning(t *testing.T) { - coord := NewIncidentCoordinator(DefaultIncidentCoordinatorConfig()) - store := memory.NewIncidentStore(memory.IncidentStoreConfig{}) - coord.SetIncidentStore(store) - - alert := &alerts.Alert{ - ID: "alert-requires-running", - ResourceID: "resource-requires-running", - ResourceName: "resource-requires-running", - } - - coord.OnAlertFired(nil) - coord.OnAlertFired(alert) - - if got := coord.GetActiveIncidentCount(); got != 0 { - t.Fatalf("expected no active incidents while coordinator is stopped, got %d", got) - } - if got := len(store.ListIncidentsByResource(alert.ResourceID, 0)); got != 0 { - t.Fatalf("expected no incident-store records while stopped, got %d", got) - } - - coord.Start() - coord.OnAlertFired(alert) - - if got := coord.GetActiveIncidentCount(); got != 1 { - t.Fatalf("expected one active incident after start, got %d", got) - } - if got := len(store.ListIncidentsByResource(alert.ResourceID, 0)); got != 1 { - t.Fatalf("expected one incident-store record after start, got %d", got) - } -} - -func TestIncidentCoordinator_OnAlertCleared_ImmediateCleanupWithoutRecorder(t *testing.T) { - coord := NewIncidentCoordinator(DefaultIncidentCoordinatorConfig()) - store := memory.NewIncidentStore(memory.IncidentStoreConfig{}) - coord.SetIncidentStore(store) - coord.Start() - - alert := &alerts.Alert{ - ID: "alert-no-recorder", - ResourceID: "resource-no-recorder", - ResourceName: "resource-no-recorder", - } - - coord.OnAlertFired(alert) - if got := coord.GetActiveIncidentCount(); got != 1 { - t.Fatalf("expected one active incident after fire, got %d", got) - } - - coord.OnAlertCleared(nil) - coord.OnAlertCleared(&alerts.Alert{ID: "missing"}) - if got := coord.GetActiveIncidentCount(); got != 1 { - t.Fatalf("expected active incident to remain after nil/missing clears, got %d", got) - } - - coord.OnAlertCleared(alert) - if got := coord.GetActiveIncidentCount(); got != 0 { - t.Fatalf("expected incident to be cleaned up immediately without recorder, got %d", got) - } - - incidents := store.ListIncidentsByResource(alert.ResourceID, 0) - if len(incidents) != 1 { - t.Fatalf("expected exactly one stored incident, got %d", len(incidents)) - } - if incidents[0].Status != memory.IncidentStatusResolved { - t.Fatalf("expected incident status %q, got %q", memory.IncidentStatusResolved, incidents[0].Status) - } - if incidents[0].ClosedAt == nil { - t.Fatal("expected incident ClosedAt to be set on clear") - } -} - -func TestIncidentCoordinator_OnAnomalyDetected_DuplicateAndCapacityAndStop(t *testing.T) { - cfg := DefaultIncidentCoordinatorConfig() - cfg.MaxConcurrent = 1 - cfg.PostDuration = time.Minute - - coord := NewIncidentCoordinator(cfg) - - recCfg := metrics.DefaultIncidentRecorderConfig() - recCfg.SampleInterval = 10 * time.Millisecond - recorder := metrics.NewIncidentRecorder(recCfg) - recorder.SetMetricsProvider(&MockMetricsProvider{data: map[string]map[string]float64{ - "resource-anomaly": {"cpu": 95}, - }}) - recorder.Start() - defer recorder.Stop() - - coord.SetRecorder(recorder) - coord.Start() - - coord.OnAnomalyDetected("resource-anomaly", "agent", "cpu", "critical") - if got := coord.GetActiveIncidentCount(); got != 1 { - t.Fatalf("expected one anomaly incident, got %d", got) - } - - coord.OnAnomalyDetected("resource-anomaly", "agent", "cpu", "critical") - if got := coord.GetActiveIncidentCount(); got != 1 { - t.Fatalf("expected duplicate anomaly to be ignored, got %d", got) - } - - coord.OnAnomalyDetected("resource-anomaly", "agent", "memory", "warning") - if got := coord.GetActiveIncidentCount(); got != 1 { - t.Fatalf("expected anomaly over capacity to be ignored, got %d", got) - } - - coord.Stop() - if got := coord.GetActiveIncidentCount(); got != 0 { - t.Fatalf("expected active incidents to be cleared on stop, got %d", got) - } - - coord.OnAnomalyDetected("resource-anomaly", "agent", "cpu", "critical") - if got := coord.GetActiveIncidentCount(); got != 0 { - t.Fatalf("expected anomalies to be ignored while stopped, got %d", got) - } -} - -func TestIncidentCoordinator_OnAnomalyDetected_CanonicalizesLegacyHostAlias(t *testing.T) { - cfg := DefaultIncidentCoordinatorConfig() - cfg.PostDuration = time.Minute - - coord := NewIncidentCoordinator(cfg) - - recCfg := metrics.DefaultIncidentRecorderConfig() - recorder := metrics.NewIncidentRecorder(recCfg) - recorder.SetMetricsProvider(&MockMetricsProvider{data: map[string]map[string]float64{ - "resource-anomaly": {"cpu": 95}, - }}) - recorder.Start() - defer recorder.Stop() - - coord.SetRecorder(recorder) - coord.Start() - - coord.OnAnomalyDetected("resource-anomaly", "host", "cpu", "critical") - - windows := recorder.GetWindowsForResource("resource-anomaly", 0) - if len(windows) != 1 { - t.Fatalf("expected one anomaly window, got %d", len(windows)) - } - if windows[0].ResourceType != "agent" { - t.Fatalf("expected anomaly recording resource type to be canonicalized to agent, got %q", windows[0].ResourceType) - } -} diff --git a/internal/ai/incident_coordinator_test.go b/internal/ai/incident_coordinator_test.go deleted file mode 100644 index 9bb4a9d0d..000000000 --- a/internal/ai/incident_coordinator_test.go +++ /dev/null @@ -1,219 +0,0 @@ -package ai - -import ( - "sync" - "testing" - "time" - - "github.com/rcourtman/pulse-go-rewrite/internal/alerts" - "github.com/rcourtman/pulse-go-rewrite/internal/metrics" -) - -// MockMetricsProvider for the real IncidentRecorder -type MockMetricsProvider struct { - mu sync.Mutex - data map[string]map[string]float64 -} - -func (m *MockMetricsProvider) GetCurrentMetrics(resourceID string) (map[string]float64, error) { - m.mu.Lock() - defer m.mu.Unlock() - if val, ok := m.data[resourceID]; ok { - return val, nil - } - return map[string]float64{"cpu": 10.0}, nil -} - -func (m *MockMetricsProvider) GetMonitoredResourceIDs() []string { - m.mu.Lock() - defer m.mu.Unlock() - keys := make([]string, 0, len(m.data)) - for k := range m.data { - keys = append(keys, k) - } - return keys -} - -func TestIncidentCoordinator_Lifecycle(t *testing.T) { - cfg := DefaultIncidentCoordinatorConfig() - coord := NewIncidentCoordinator(cfg) - - if coord.running { - t.Error("Coordinator should not be running initially") - } - - coord.Start() - if !coord.running { - t.Error("Coordinator should be running after Start()") - } - - // Double start should be safe - coord.Start() - - coord.Stop() - if coord.running { - t.Error("Coordinator should not be running after Stop()") - } - - // Double stop should be safe - coord.Stop() -} - -func TestIncidentCoordinator_OnAlertFired(t *testing.T) { - cfg := DefaultIncidentCoordinatorConfig() - coord := NewIncidentCoordinator(cfg) - - // Create real recorder with mock provider - recCfg := metrics.DefaultIncidentRecorderConfig() - recCfg.SampleInterval = 50 * time.Millisecond // fast sampling - recorder := metrics.NewIncidentRecorder(recCfg) - provider := &MockMetricsProvider{data: make(map[string]map[string]float64)} - recorder.SetMetricsProvider(provider) - recorder.Start() // Recorder must be started - defer recorder.Stop() - - // Inject recorder into coordinator - coord.SetRecorder(recorder) - coord.Start() - - alert := &alerts.Alert{ - ID: "alert-1", - ResourceID: "res-1", - } - - coord.OnAlertFired(alert) - - if coord.GetActiveIncidentCount() != 1 { - t.Errorf("Expected 1 active incident, got %d", coord.GetActiveIncidentCount()) - } - - // Verify recorder has active window - wid := coord.GetRecordingWindowID("alert-1") - if wid == "" { - t.Error("Expected valid window ID") - } - - // Fire same alert again - should be ignored - coord.OnAlertFired(alert) - if coord.GetActiveIncidentCount() != 1 { - t.Errorf("Expected count to remain 1, got %d", coord.GetActiveIncidentCount()) - } - - // Fire another alert - alert2 := &alerts.Alert{ - ID: "alert-2", - ResourceID: "res-2", - } - coord.OnAlertFired(alert2) - if coord.GetActiveIncidentCount() != 2 { - t.Errorf("Expected 2 active incidents, got %d", coord.GetActiveIncidentCount()) - } -} - -func TestIncidentCoordinator_OnAlertCleared(t *testing.T) { - cfg := DefaultIncidentCoordinatorConfig() - cfg.PostDuration = 50 * time.Millisecond // fast for testing - coord := NewIncidentCoordinator(cfg) - - // Create real recorder - recCfg := metrics.DefaultIncidentRecorderConfig() - recorder := metrics.NewIncidentRecorder(recCfg) - provider := &MockMetricsProvider{data: make(map[string]map[string]float64)} - recorder.SetMetricsProvider(provider) - recorder.Start() - defer recorder.Stop() - - coord.SetRecorder(recorder) - coord.Start() - - alert := &alerts.Alert{ID: "alert-1", ResourceID: "res-1"} - coord.OnAlertFired(alert) - - if coord.GetActiveIncidentCount() != 1 { - t.Fatal("Failed to start incident") - } - - // Clear alert - coord.OnAlertCleared(alert) - - // Since we have a recorder and postDuration is 50ms, it should NOT be removed immediately - if coord.GetActiveIncidentCount() != 1 { - t.Error("Incident should NOT be removed immediately when recorder is active") - } - - // Wait for post duration - time.Sleep(100 * time.Millisecond) - - // Now it should be removed (via time.AfterFunc callback) - if coord.GetActiveIncidentCount() != 0 { - t.Errorf("Incident should be removed after post duration, count=%d", coord.GetActiveIncidentCount()) - } -} - -func TestIncidentCoordinator_MaxConcurrent(t *testing.T) { - cfg := DefaultIncidentCoordinatorConfig() - cfg.MaxConcurrent = 1 - coord := NewIncidentCoordinator(cfg) - coord.Start() - - coord.OnAlertFired(&alerts.Alert{ID: "alert-1", ResourceID: "res-1"}) - if coord.GetActiveIncidentCount() != 1 { - t.Fatal("Should accept first incident") - } - - coord.OnAlertFired(&alerts.Alert{ID: "alert-2", ResourceID: "res-2"}) - if coord.GetActiveIncidentCount() != 1 { - t.Error("Should ignore second incident due to cap") - } -} - -func TestIncidentCoordinator_OnAnomalyDetected(t *testing.T) { - cfg := DefaultIncidentCoordinatorConfig() - coord := NewIncidentCoordinator(cfg) - - recCfg := metrics.DefaultIncidentRecorderConfig() - recorder := metrics.NewIncidentRecorder(recCfg) - provider := &MockMetricsProvider{data: make(map[string]map[string]float64)} - recorder.SetMetricsProvider(provider) - recorder.Start() - defer recorder.Stop() - - coord.SetRecorder(recorder) - coord.Start() - - // Start anomaly recording - coord.OnAnomalyDetected("res-1", "agent", "cpu", "critical") - - if coord.GetActiveIncidentCount() != 1 { - t.Errorf("Should start incident for anomaly, got %d", coord.GetActiveIncidentCount()) - } - - // Anomaly ID format check (internal detail, but verify implicitly via count) -} - -func TestIncidentCoordinator_GetRecordingWindowID(t *testing.T) { - cfg := DefaultIncidentCoordinatorConfig() - coord := NewIncidentCoordinator(cfg) - - recCfg := metrics.DefaultIncidentRecorderConfig() - recorder := metrics.NewIncidentRecorder(recCfg) - recorder.Start() - defer recorder.Stop() - - coord.SetRecorder(recorder) - coord.Start() - - alert := &alerts.Alert{ID: "alert-1", ResourceID: "res-1"} - coord.OnAlertFired(alert) - - // With recorder, windowID should be present (start with 'iw-') - wid := coord.GetRecordingWindowID("alert-1") - if wid == "" { - t.Error("Expected valid window ID") - } - - widMissing := coord.GetRecordingWindowID("missing") - if widMissing != "" { - t.Error("Expected empty window ID for missing alert") - } -} diff --git a/internal/ai/investigation_records_test.go b/internal/ai/investigation_records_test.go index d75d4fff5..41ec67443 100644 --- a/internal/ai/investigation_records_test.go +++ b/internal/ai/investigation_records_test.go @@ -204,3 +204,25 @@ func TestFindingsStore_UpdateInvestigationRecord(t *testing.T) { t.Fatal("expected false for missing finding") } } + +func TestPatrolInvestigationCompletionReplacesEarlyActionProjection(t *testing.T) { + store := NewFindingsStore() + store.Add(&Finding{ID: "finding-1", ResourceID: "vm-100", Title: "High CPU", DetectedAt: time.Now()}) + patrol := &PatrolService{findings: store} + session := &InvestigationSession{ID: "investigation-1", FindingID: "finding-1", Summary: "Investigation in progress"} + if !patrol.RefreshFindingInvestigationRecord("finding-1", session) { + t.Fatal("early action projection was not recorded") + } + completed := time.Now() + session.Status = aicontracts.InvestigationStatusCompleted + session.CompletedAt = &completed + session.Summary = "Final diagnosis supported by the completed reads" + session.EvidenceIDs = []string{"final-evidence"} + if !patrol.storeFindingInvestigationRecord("finding-1", session, false) { + t.Fatal("completed investigation did not replace its early projection") + } + record := store.Get("finding-1").InvestigationRecord + if record.Conclusion != session.Summary || record.CompletedAt == nil || len(record.Evidence) != 1 || record.Evidence[0].ID != "final-evidence" { + t.Fatalf("completed evidence was lost: %#v", record) + } +} diff --git a/internal/ai/patrol_findings.go b/internal/ai/patrol_findings.go index d95c2be79..8e2a5aaa4 100644 --- a/internal/ai/patrol_findings.go +++ b/internal/ai/patrol_findings.go @@ -7,6 +7,7 @@ import ( "context" "encoding/json" "fmt" + "reflect" "sort" "strings" "sync" @@ -1791,23 +1792,9 @@ func (p *PatrolService) maybeInvestigateFinding(f *Finding) bool { if orchestrator != nil { latestInvestigation = orchestrator.GetInvestigationByFinding(latest.ID) } - if record := BuildFindingInvestigationRecord(latest, latestInvestigation); record != nil { - // When a remediation plan exists for this finding, lift its - // per-step rollback strings into record.Rollback so the - // operator-facing investigation surface answers - // "what's the undo for the proposed fix?" at the record root - // rather than only in nested per-step payload. - if engine := p.remediationEngine; engine != nil { - if plan := engine.GetPlanForFinding(latest.ID); plan != nil { - record.Rollback = AggregatePlanRollbackSteps(plan) - } - } - if p.findings.UpdateInvestigationRecord(latest.ID, record) { - if refreshed := p.findings.Get(latest.ID); refreshed != nil { - latest = refreshed - } else { - latest.InvestigationRecord = record - } + if p.storeFindingInvestigationRecord(latest.ID, latestInvestigation, false) { + if refreshed := p.findings.Get(latest.ID); refreshed != nil { + latest = refreshed } } if pushUnified != nil { @@ -1850,11 +1837,50 @@ func (p *PatrolService) maybeInvestigateFinding(f *Finding) bool { return true } -// PublishFindingLifecycleUpdate projects a reconciled action outcome to the -// unified finding owner and, for terminal execution outcomes, to mobile push. -// It is called only after the finding store changed, so duplicate action -// callbacks and read-time hydration do not emit duplicate notifications. -func (p *PatrolService) PublishFindingLifecycleUpdate(findingID string) { +// RefreshFindingInvestigationRecord preserves the latest investigation and +// reconciled action in the durable record shared by product surfaces. +func (p *PatrolService) RefreshFindingInvestigationRecord(findingID string, session *InvestigationSession) bool { + return p.storeFindingInvestigationRecord(findingID, session, true) +} + +func (p *PatrolService) storeFindingInvestigationRecord(findingID string, session *InvestigationSession, preserveEvidence bool) bool { + if p == nil || p.findings == nil { + return false + } + finding := p.findings.Get(findingID) + if finding == nil { + return false + } + record := BuildFindingInvestigationRecord(finding, session) + // Later action transitions update lifecycle facts, not the evidence and + // diagnosis captured when this investigation completed. The current finding + // may no longer retain all of that original context after restart. + if previous := finding.InvestigationRecord; preserveEvidence && previous != nil && previous.ID == record.ID { + retained := previous.NormalizeCollections() + retained.Status = record.Status + retained.Outcome = record.Outcome + retained.Action = record.Action + retained.Verification = record.Verification + record = &retained + } else { + p.mu.RLock() + engine := p.remediationEngine + p.mu.RUnlock() + if engine != nil { + if plan := engine.GetPlanForFinding(findingID); plan != nil { + record.Rollback = AggregatePlanRollbackSteps(plan) + } + } + } + if reflect.DeepEqual(finding.InvestigationRecord, record) { + return false + } + return p.findings.UpdateInvestigationRecord(findingID, record) +} + +// PublishFindingLifecycleUpdate projects reconciled records to the unified +// finding owner. Repairing a stale record alone must not repeat outcome pushes. +func (p *PatrolService) PublishFindingLifecycleUpdate(findingID string, outcomeChanged bool) { if p == nil || p.findings == nil { return } @@ -1873,7 +1899,7 @@ func (p *PatrolService) PublishFindingLifecycleUpdate(findingID string) { if finding.ResolvedAt != nil && resolveUnified != nil { resolveUnified(finding.ID) } - if pushNotify == nil { + if pushNotify == nil || !outcomeChanged { return } switch InvestigationOutcome(finding.InvestigationOutcome) { diff --git a/internal/ai/qualification/client.go b/internal/ai/qualification/client.go index 136350dd1..804b873b7 100644 --- a/internal/ai/qualification/client.go +++ b/internal/ai/qualification/client.go @@ -534,11 +534,25 @@ func validatePatrolRoute(expected string, settings AISettings, status PatrolStat } func (c *PulseClient) Resources(ctx context.Context) ([]Resource, error) { - var response struct { - Data []Resource `json:"data"` + var resources []Resource + for page := 1; ; page++ { + var response struct { + Data []Resource `json:"data"` + Meta struct { + TotalPages int `json:"totalPages"` + } `json:"meta"` + } + // The resource API caps each page at 100 regardless of the requested + // limit. Follow its pagination so later resources can converge too. + path := fmt.Sprintf("/api/resources?limit=100&page=%d", page) + if err := c.request(ctx, http.MethodGet, path, nil, &response); err != nil { + return nil, err + } + resources = append(resources, response.Data...) + if page >= response.Meta.TotalPages { + return resources, nil + } } - err := c.request(ctx, http.MethodGet, "/api/resources?limit=1000", nil, &response) - return response.Data, err } func (c *PulseClient) WaitForResources(ctx context.Context, names map[string]string, timeout, poll time.Duration) (map[string]Resource, error) { diff --git a/internal/ai/qualification/client_test.go b/internal/ai/qualification/client_test.go index dae04bc48..cc4bc104e 100644 --- a/internal/ai/qualification/client_test.go +++ b/internal/ai/qualification/client_test.go @@ -4,6 +4,7 @@ import ( "context" "encoding/json" "errors" + "fmt" "io" "net/http" "net/http/httptest" @@ -13,6 +14,62 @@ import ( "time" ) +func TestWaitForResourcesMatchingIncludesLaterPages(t *testing.T) { + var pages []string + server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { + if r.URL.Path != "/api/resources" || r.URL.Query().Get("limit") != "100" { + t.Errorf("unexpected resource request: %s", r.URL) + } + page := r.URL.Query().Get("page") + pages = append(pages, page) + resources := make([]Resource, 0, 100) + switch page { + case "1": + for i := 0; i < 100; i++ { + resources = append(resources, Resource{ID: fmt.Sprintf("control-%d", i), Name: fmt.Sprintf("control-%d", i)}) + } + case "2": + resources = append(resources, Resource{ID: "storage", Name: "worker", Docker: &DockerResource{Health: "unhealthy"}}) + default: + t.Errorf("unexpected page: %q", page) + } + _ = json.NewEncoder(w).Encode(map[string]any{"data": resources, "meta": map[string]int{"totalPages": 2}}) + })) + defer server.Close() + client, err := NewPulseClient(ClientConfig{BaseURL: server.URL}) + if err != nil { + t.Fatal(err) + } + resources, err := client.WaitForResourcesMatching(context.Background(), map[string]string{"service": "worker"}, time.Second, time.Millisecond, func(resources map[string]Resource) error { + if resources["service"].Docker.Health != "unhealthy" { + return errors.New("fault not collected") + } + return nil + }) + if err != nil || resources["service"].ID != "storage" || strings.Join(pages, ",") != "1,2" { + t.Fatalf("resources=%v pages=%v err=%v", resources, pages, err) + } +} + +func TestResourcesDoesNotReturnPartialInventoryWhenLaterPageFails(t *testing.T) { + server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { + if r.URL.Query().Get("page") == "1" { + _, _ = w.Write([]byte(`{"data":[{"id":"first"}],"meta":{"totalPages":2}}`)) + return + } + http.Error(w, "resource inventory unavailable", http.StatusServiceUnavailable) + })) + defer server.Close() + client, err := NewPulseClient(ClientConfig{BaseURL: server.URL}) + if err != nil { + t.Fatal(err) + } + resources, err := client.Resources(context.Background()) + if err == nil || resources != nil { + t.Fatalf("incomplete inventory returned: resources=%v err=%v", resources, err) + } +} + func TestTriggerAndWaitAssociatesExactNewScopedRun(t *testing.T) { var triggered atomic.Bool server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { diff --git a/internal/ai/service.go b/internal/ai/service.go index 96f27f4c6..d3e0e8c6e 100644 --- a/internal/ai/service.go +++ b/internal/ai/service.go @@ -184,13 +184,12 @@ func (m ChatMessage) NormalizeCollections() ChatMessage { return m } -// ChatToolCall represents a provider-facing tool invocation in an API-facing -// chat message. It aliases the shared Pulse Intelligence provider-call shape so -// API chat history and provider turns do not drift on tool-call JSON. -type ChatToolCall = agentcapabilities.ProviderToolCall +// ChatToolCall retains observed output and result status in product history. +// Provider turns use the explicit ProviderToolCall projection. +type ChatToolCall = agentcapabilities.TranscriptToolCall func EmptyChatToolCall() ChatToolCall { - return agentcapabilities.EmptyProviderToolCall() + return ChatToolCall{}.NormalizeCollections() } // ChatToolResult represents the result of a tool invocation. It aliases the diff --git a/internal/ai/service_test.go b/internal/ai/service_test.go index 33d78b474..4acb88ec4 100644 --- a/internal/ai/service_test.go +++ b/internal/ai/service_test.go @@ -168,7 +168,7 @@ func TestChatMessage_UsesCanonicalEmptyCollections(t *testing.T) { Name: "diagnose", ThoughtSignature: json.RawMessage(`{"provider":"gemini"}`), } - var sharedProviderCall agentcapabilities.ProviderToolCall = sharedCall.NormalizeCollections() + var sharedProviderCall agentcapabilities.TranscriptToolCall = sharedCall.NormalizeCollections() if sharedProviderCall.ID != "call-1" || sharedProviderCall.Input == nil { t.Fatalf("shared chat tool call = %+v", sharedProviderCall) } diff --git a/internal/ai/tools/command_connection_evidence_test.go b/internal/ai/tools/command_connection_evidence_test.go new file mode 100644 index 000000000..de9e2fcc0 --- /dev/null +++ b/internal/ai/tools/command_connection_evidence_test.go @@ -0,0 +1,145 @@ +package tools + +import ( + "context" + "encoding/json" + "testing" + "time" + + "github.com/rcourtman/pulse-go-rewrite/internal/agentexec" + "github.com/rcourtman/pulse-go-rewrite/internal/models" + "github.com/rcourtman/pulse-go-rewrite/internal/unifiedresources" +) + +func commandEvidenceSnapshot() models.StateSnapshot { + return models.StateSnapshot{ + Nodes: []models.Node{{ID: "node-one", Name: "node-one", Status: "online"}}, + VMs: []models.VM{{ID: "vm-one", VMID: 101, Name: "guest-one", Node: "node-one", Status: "running"}}, + DockerHosts: []models.DockerHost{{ + ID: "host-one", Hostname: "command-host", Status: "online", LastSeen: time.Now(), + Containers: []models.DockerContainer{{ID: "container-one", Name: "observed-service", State: "running", CPUPercent: 12.5}}, + }}, + } +} + +func commandEvidenceJSON(t *testing.T, value any) map[string]any { + t.Helper() + raw, err := json.Marshal(value) + if err != nil { + t.Fatal(err) + } + var result map[string]any + if err := json.Unmarshal(raw, &result); err != nil { + t.Fatal(err) + } + return result +} + +func TestCommandConnectivityDoesNotReplaceMonitoringEvidence(t *testing.T) { + for _, tc := range []struct { + name string + connected, guestConnected, controlEnabled bool + }{ + {name: "no_command_connection"}, + {name: "connected_read_only", connected: true}, + {name: "guest_connection_only", guestConnected: true}, + {name: "connected_control_enabled", connected: true, controlEnabled: true}, + } { + t.Run(tc.name, func(t *testing.T) { + server := &mockAgentServer{} + if tc.connected { + server.agents = []agentexec.ConnectedAgent{{Hostname: "command-host"}, {Hostname: "node-one"}} + } + if tc.guestConnected { + server.agents = append(server.agents, agentexec.ConnectedAgent{Hostname: "guest-one"}) + } + controlLevel := ControlLevelReadOnly + if tc.controlEnabled { + controlLevel = ControlLevelControlled + } + registry := unifiedresources.NewRegistry(nil) + registry.IngestSnapshot(commandEvidenceSnapshot()) + executor := NewPulseToolExecutor(ExecutorConfig{ReadState: registry, UnifiedResourceProvider: ®istryUnifiedQueryProvider{registry}, AgentServer: server, ControlLevel: controlLevel}) + query := func(args map[string]interface{}) map[string]any { + t.Helper() + result, err := executor.executeQuery(context.Background(), args) + if err != nil || result.IsError { + t.Fatalf("query failed: %v %+v", err, result) + } + var decoded map[string]any + if err := json.Unmarshal([]byte(result.Content[0].Text), &decoded); err != nil { + t.Fatal(err) + } + capture, _ := json.Marshal(map[string]any{"case": tc.name, "input": args, "output": decoded}) + t.Logf("COMMAND_EVIDENCE %s", capture) + return decoded + } + list := query(map[string]interface{}{"action": "list", "type": "docker-hosts"}) + host := list["docker_hosts"].([]any)[0].(map[string]any) + if host["command_agent_connected"] != tc.connected { + t.Fatalf("connection must name command transport: %+v", host) + } + if _, exists := host["agent_connected"]; exists { + t.Fatal("ambiguous connection field remains") + } + container := host["containers"].([]any)[0].(map[string]any) + resource := query(map[string]interface{}{"action": "get", "resource_type": "app-container", "resource_id": container["id"]}) + if resource["status"] != "running" || resource["cpu"].(map[string]any)["percent"] != 12.5 { + t.Fatalf("command state replaced monitored evidence: %+v", resource) + } + topology := query(map[string]interface{}{"action": "topology", "include": "all"}) + docker := topology["docker"].(map[string]any)["hosts"].([]any)[0].(map[string]any) + if docker["command_agent_connected"] != tc.connected || docker["can_execute"] != (tc.connected && tc.controlEnabled) { + t.Fatalf("transport/control hint changed: %+v", docker) + } + search := query(map[string]interface{}{"action": "search", "query": "guest-one"}) + guest := search["matches"].([]any)[0].(map[string]any) + if guest["node_command_agent_connected"] != tc.connected { + t.Fatalf("parent transport not identified: %+v", guest) + } + if guest["command_agent_connected"] != tc.guestConnected { + t.Fatalf("parent connection became a direct guest connection: %+v", guest) + } + server.AssertNotCalled(t, "ExecuteCommand") + }) + } +} + +func TestTopologyOmitsUnobservedCommandConnections(t *testing.T) { + registry := unifiedresources.NewRegistry(nil) + registry.IngestSnapshot(commandEvidenceSnapshot()) + for _, observed := range []bool{false, true} { + options := TopologyBuildOptions{Include: "all", ControlEnabled: true} + if observed { + options.ConnectedAgentHostnames = map[string]bool{} + } + result := commandEvidenceJSON(t, BuildTopologyResponseFromReadState(registry, options)) + caseName := "unobserved_topology" + if observed { + caseName = "observed_empty_topology" + } + capture, err := json.Marshal(map[string]any{"case": caseName, "input": map[string]any{"action": "topology", "include": "all"}, "output": result}) + if err != nil { + t.Fatal(err) + } + t.Logf("COMMAND_EVIDENCE %s", capture) + for _, group := range []struct{ family, collection string }{{"docker", "hosts"}, {"proxmox", "nodes"}} { + item := result[group.family].(map[string]any)[group.collection].([]any)[0].(map[string]any) + for _, field := range []string{"command_agent_connected", "can_execute"} { + value, exists := item[field] + if exists != observed || (exists && value != false) { + t.Fatalf("observed=%t field=%s: %+v", observed, field, item) + } + } + if _, exists := item["agent_connected"]; exists { + t.Fatal("ambiguous connection field remains") + } + } + for _, field := range []string{"nodes_with_command_agents", "docker_hosts_with_command_agents"} { + value, exists := result["summary"].(map[string]any)[field] + if exists != observed || (exists && value != float64(0)) { + t.Fatalf("unobserved connections became a count: %+v", result["summary"]) + } + } + } +} diff --git a/internal/ai/tools/data_types.go b/internal/ai/tools/data_types.go index 24b3f76c7..8344ec73d 100644 --- a/internal/ai/tools/data_types.go +++ b/internal/ai/tools/data_types.go @@ -238,40 +238,40 @@ func (r ResourceSearchResponse) NormalizeCollections() ResourceSearchResponse { // ResourceMatch is a compact match result for pulse_search_resources type ResourceMatch struct { GovernedResourceMetadata - Type string `json:"type"` // "agent", "node", "vm", "system-container", "app-container", "docker-host", "storage" - ID string `json:"id,omitempty"` - Name string `json:"name"` - Status string `json:"status,omitempty"` - Node string `json:"node,omitempty"` // Hypervisor node this resource is on - NodeHasAgent bool `json:"node_has_agent,omitempty"` // True if the node has a connected agent - Host string `json:"host,omitempty"` // Docker host for docker containers - Platform string `json:"platform,omitempty"` - VMID int `json:"vmid,omitempty"` - Image string `json:"image,omitempty"` - AgentConnected bool `json:"agent_connected,omitempty"` // True if this specific resource has a connected agent + Type string `json:"type"` // "agent", "node", "vm", "system-container", "app-container", "docker-host", "storage" + ID string `json:"id,omitempty"` + Name string `json:"name"` + Status string `json:"status,omitempty"` + Node string `json:"node,omitempty"` // Hypervisor node this resource is on + NodeCommandAgentConnected *bool `json:"node_command_agent_connected,omitempty"` // Live command connection on the parent node, independent of telemetry collection + Host string `json:"host,omitempty"` // Docker host for docker containers + Platform string `json:"platform,omitempty"` + VMID int `json:"vmid,omitempty"` + Image string `json:"image,omitempty"` + CommandAgentConnected *bool `json:"command_agent_connected,omitempty"` // Live command connection for this resource, independent of telemetry collection } // SystemSummary is a summarized infrastructure system for list responses. type SystemSummary struct { GovernedResourceMetadata - ID string `json:"id"` - Name string `json:"name"` - Status string `json:"status"` - Platform string `json:"platform,omitempty"` - ChildCount int `json:"child_count,omitempty"` - AgentConnected bool `json:"agent_connected,omitempty"` - CPU float64 `json:"cpu_percent,omitempty"` - Memory float64 `json:"memory_percent,omitempty"` - Disk float64 `json:"disk_percent,omitempty"` + ID string `json:"id"` + Name string `json:"name"` + Status string `json:"status"` + Platform string `json:"platform,omitempty"` + ChildCount int `json:"child_count,omitempty"` + CommandAgentConnected *bool `json:"command_agent_connected,omitempty"` + CPU float64 `json:"cpu_percent,omitempty"` + Memory float64 `json:"memory_percent,omitempty"` + Disk float64 `json:"disk_percent,omitempty"` } // NodeSummary is a summarized node for list responses type NodeSummary struct { GovernedResourceMetadata - Name string `json:"name"` - Status string `json:"status"` - ID string `json:"id,omitempty"` - AgentConnected bool `json:"agent_connected"` // True if an execution agent is connected for this node + Name string `json:"name"` + Status string `json:"status"` + ID string `json:"id,omitempty"` + CommandAgentConnected *bool `json:"command_agent_connected,omitempty"` // True if an execution agent is connected for this node } // VMSummary is a summarized VM for list responses @@ -299,12 +299,12 @@ type ContainerSummary struct { // DockerHostSummary is a summarized Docker host for list responses type DockerHostSummary struct { GovernedResourceMetadata - ID string `json:"id"` - Hostname string `json:"hostname"` - DisplayName string `json:"display_name,omitempty"` - ContainerCount int `json:"container_count"` - AgentConnected bool `json:"agent_connected"` // True if an execution agent is connected for this host - Containers []DockerContainerSummary `json:"containers"` + ID string `json:"id"` + Hostname string `json:"hostname"` + DisplayName string `json:"display_name,omitempty"` + ContainerCount int `json:"container_count"` + CommandAgentConnected *bool `json:"command_agent_connected,omitempty"` // True if an execution agent is connected for this host + Containers []DockerContainerSummary `json:"containers"` } func (s DockerHostSummary) NormalizeCollections() DockerHostSummary { @@ -454,15 +454,15 @@ func (t ProxmoxTopology) NormalizeCollections() ProxmoxTopology { // ProxmoxNodeTopology represents a Proxmox node with its guests type ProxmoxNodeTopology struct { GovernedResourceMetadata - Name string `json:"name"` - ID string `json:"id,omitempty"` - Status string `json:"status"` - AgentConnected bool `json:"agent_connected"` - CanExecute bool `json:"can_execute"` // True if commands can be executed on this node - VMs []TopologyVM `json:"vms"` - Containers []TopologyContainer `json:"containers"` - VMCount int `json:"vm_count"` - ContainerCount int `json:"container_count"` + Name string `json:"name"` + ID string `json:"id,omitempty"` + Status string `json:"status"` + CommandAgentConnected *bool `json:"command_agent_connected,omitempty"` + CanExecute *bool `json:"can_execute,omitempty"` // True if commands can be executed on this node + VMs []TopologyVM `json:"vms"` + Containers []TopologyContainer `json:"containers"` + VMCount int `json:"vm_count"` + ContainerCount int `json:"container_count"` } func (t ProxmoxNodeTopology) NormalizeCollections() ProxmoxNodeTopology { @@ -538,15 +538,15 @@ func (t DockerTopology) NormalizeCollections() DockerTopology { // DockerHostTopology represents a Docker host with its containers type DockerHostTopology struct { GovernedResourceMetadata - Hostname string `json:"hostname"` - DisplayName string `json:"display_name,omitempty"` - AgentConnected bool `json:"agent_connected"` - CanExecute bool `json:"can_execute"` // True if commands can be executed on this host - Containers []DockerContainerSummary `json:"containers"` - ContainerCount int `json:"container_count"` - ReturnedCount int `json:"returned_container_count"` - Truncated bool `json:"containers_truncated"` - RunningCount int `json:"running_count"` + Hostname string `json:"hostname"` + DisplayName string `json:"display_name,omitempty"` + CommandAgentConnected *bool `json:"command_agent_connected,omitempty"` + CanExecute *bool `json:"can_execute,omitempty"` // True if commands can be executed on this host + Containers []DockerContainerSummary `json:"containers"` + ContainerCount int `json:"container_count"` + ReturnedCount int `json:"returned_container_count"` + Truncated bool `json:"containers_truncated"` + RunningCount int `json:"running_count"` } func (t DockerHostTopology) NormalizeCollections() DockerHostTopology { @@ -642,21 +642,21 @@ type KubernetesPodDetail struct { // TopologySummary provides aggregate counts and status type TopologySummary struct { - TotalNodes int `json:"total_nodes"` - TotalVMs int `json:"total_vms"` - TotalSystemContainers int `json:"total_system_containers"` - TotalDockerHosts int `json:"total_docker_hosts"` - TotalDockerContainers int `json:"total_docker_containers"` - TotalK8sClusters int `json:"total_k8s_clusters"` - TotalK8sNodes int `json:"total_k8s_nodes"` - TotalK8sDeployments int `json:"total_k8s_deployments"` - TotalK8sPods int `json:"total_k8s_pods"` - NodesWithAgents int `json:"nodes_with_agents"` - DockerHostsWithAgents int `json:"docker_hosts_with_agents"` - RunningVMs int `json:"running_vms"` - RunningContainers int `json:"running_containers"` - RunningDocker int `json:"running_docker"` - RunningK8sPods int `json:"running_k8s_pods"` + TotalNodes int `json:"total_nodes"` + TotalVMs int `json:"total_vms"` + TotalSystemContainers int `json:"total_system_containers"` + TotalDockerHosts int `json:"total_docker_hosts"` + TotalDockerContainers int `json:"total_docker_containers"` + TotalK8sClusters int `json:"total_k8s_clusters"` + TotalK8sNodes int `json:"total_k8s_nodes"` + TotalK8sDeployments int `json:"total_k8s_deployments"` + TotalK8sPods int `json:"total_k8s_pods"` + NodesWithCommandAgents *int `json:"nodes_with_command_agents,omitempty"` + DockerHostsWithCommandAgents *int `json:"docker_hosts_with_command_agents,omitempty"` + RunningVMs int `json:"running_vms"` + RunningContainers int `json:"running_containers"` + RunningDocker int `json:"running_docker"` + RunningK8sPods int `json:"running_k8s_pods"` } // ResourceResponse is returned by pulse_get_resource @@ -883,8 +883,10 @@ type PortInfo struct { // MountInfo describes a volume mount type MountInfo struct { + Type string `json:"type,omitempty"` Source string `json:"source"` Destination string `json:"destination"` + Mode string `json:"mode,omitempty"` ReadWrite bool `json:"rw"` } diff --git a/internal/ai/tools/executor.go b/internal/ai/tools/executor.go index 7b1870a17..48f0468b8 100644 --- a/internal/ai/tools/executor.go +++ b/internal/ai/tools/executor.go @@ -559,9 +559,9 @@ type ExecutorConfig struct { AgentProfileManager AgentProfileManager // Optional providers - intelligence - IncidentRecorderProvider IncidentRecorderProvider - EventCorrelatorProvider EventCorrelatorProvider - KnowledgeStoreProvider KnowledgeStoreProvider + IncidentArchiveProvider IncidentArchiveProvider + EventCorrelatorProvider EventCorrelatorProvider + KnowledgeStoreProvider KnowledgeStoreProvider // Optional providers - discovery DiscoveryProvider DiscoveryProvider @@ -627,9 +627,9 @@ type PulseToolExecutor struct { agentProfileManager AgentProfileManager // Intelligence providers - incidentRecorderProvider IncidentRecorderProvider - eventCorrelatorProvider EventCorrelatorProvider - knowledgeStoreProvider KnowledgeStoreProvider + incidentArchiveProvider IncidentArchiveProvider + eventCorrelatorProvider EventCorrelatorProvider + knowledgeStoreProvider KnowledgeStoreProvider // Discovery provider discoveryProvider DiscoveryProvider @@ -750,7 +750,7 @@ func NewPulseToolExecutor(cfg ExecutorConfig) *PulseToolExecutor { metadataUpdater: cfg.MetadataUpdater, findingsManager: cfg.FindingsManager, agentProfileManager: cfg.AgentProfileManager, - incidentRecorderProvider: cfg.IncidentRecorderProvider, + incidentArchiveProvider: cfg.IncidentArchiveProvider, eventCorrelatorProvider: cfg.EventCorrelatorProvider, knowledgeStoreProvider: cfg.KnowledgeStoreProvider, discoveryProvider: cfg.DiscoveryProvider, @@ -827,7 +827,7 @@ func (e *PulseToolExecutor) Clone() *PulseToolExecutor { metadataUpdater: e.metadataUpdater, findingsManager: e.findingsManager, agentProfileManager: e.agentProfileManager, - incidentRecorderProvider: e.incidentRecorderProvider, + incidentArchiveProvider: e.incidentArchiveProvider, eventCorrelatorProvider: e.eventCorrelatorProvider, knowledgeStoreProvider: e.knowledgeStoreProvider, discoveryProvider: e.discoveryProvider, @@ -1019,9 +1019,9 @@ func (e *PulseToolExecutor) SetUpdatesProvider(provider UpdatesProvider) { e.updatesProvider = provider } -// SetIncidentRecorderProvider sets the incident recorder provider -func (e *PulseToolExecutor) SetIncidentRecorderProvider(provider IncidentRecorderProvider) { - e.incidentRecorderProvider = provider +// SetIncidentArchiveProvider sets the read-only legacy incident archive provider +func (e *PulseToolExecutor) SetIncidentArchiveProvider(provider IncidentArchiveProvider) { + e.incidentArchiveProvider = provider } // SetEventCorrelatorProvider sets the event correlator provider @@ -1221,7 +1221,7 @@ func (e *PulseToolExecutor) isToolAvailable(name string) bool { case agentcapabilities.PulseDiscoveryToolName: return e.discoveryProvider != nil case agentcapabilities.PulseKnowledgeToolName: - return e.knowledgeStoreProvider != nil || e.incidentRecorderProvider != nil || e.eventCorrelatorProvider != nil + return e.actionAuditStore != nil || e.knowledgeStoreProvider != nil || e.incidentArchiveProvider != nil || e.eventCorrelatorProvider != nil case agentcapabilities.PulsePMGToolName: return e.hasReadState() case agentcapabilities.PulseSummarizeToolName: diff --git a/internal/ai/tools/incident_history_test.go b/internal/ai/tools/incident_history_test.go new file mode 100644 index 000000000..cc21e2f80 --- /dev/null +++ b/internal/ai/tools/incident_history_test.go @@ -0,0 +1,221 @@ +package tools + +import ( + "context" + "encoding/json" + "errors" + "os" + "path/filepath" + "strings" + "testing" + "time" + + "github.com/rcourtman/pulse-go-rewrite/internal/agentcapabilities" + "github.com/rcourtman/pulse-go-rewrite/internal/metrics" + "github.com/rcourtman/pulse-go-rewrite/internal/unifiedresources" + "github.com/stretchr/testify/require" +) + +type failedIncidentHistoryStore struct{ unifiedresources.ResourceStore } + +func TestIncidentHistoryRetainsLegacyDockerLifecycle(t *testing.T) { + dir := t.TempDir() + store, err := unifiedresources.NewSQLiteResourceStore(dir, "incident-test") + require.NoError(t, err) + t.Cleanup(func() { require.NoError(t, store.Close()) }) + container := strings.Repeat("b", 64) + legacy := "docker:tower/" + container + canonical := unifiedresources.SourceSpecificID(unifiedresources.ResourceTypeAppContainer, unifiedresources.SourceDocker, "tower/container/"+container) + start := time.Now().UTC().Add(-time.Hour).Truncate(time.Second) + occurred := start.Add(-time.Minute) + fired := unifiedresources.ResourceChange{ID: "legacy-fired", ResourceID: legacy, ObservedAt: start.Add(time.Minute), OccurredAt: &occurred, Kind: unifiedresources.ChangeAlertFired, SourceType: unifiedresources.SourceHeuristic, Reason: "Container unhealthy"} + require.NoError(t, store.RecordChange(fired)) + require.NoError(t, store.Close()) + store, err = unifiedresources.NewSQLiteResourceStore(dir, "incident-test") + require.NoError(t, err) + resolved := unifiedresources.ResourceChange{ID: "canonical-resolved", ResourceID: canonical, ObservedAt: start.Add(3 * time.Minute), Kind: unifiedresources.ChangeAlertResolved, SourceType: unifiedresources.SourceHeuristic} + require.NoError(t, store.RecordChange(resolved)) + exec := NewPulseToolExecutor(ExecutorConfig{ActionAuditStore: store}) + input := map[string]interface{}{"action": "incidents", "resource_id": canonical, "since": start.Format(time.RFC3339), "limit": float64(50)} + result, err := exec.registry.Execute(context.Background(), exec, agentcapabilities.PulseKnowledgeToolName, input) + require.NoError(t, err) + require.False(t, result.IsError, result.Content) + var got struct { + Events []unifiedresources.ResourceChange `json:"events"` + } + require.NoError(t, json.Unmarshal([]byte(result.Content[0].Text), &got)) + require.Equal(t, []unifiedresources.ResourceChange{resolved, fired}, got.Events) + capture, err := json.Marshal(map[string]any{"case": "migrated Docker lifecycle", "input": input, "result": result}) + require.NoError(t, err) + t.Logf("INCIDENT_EVIDENCE %s", capture) +} + +func (failedIncidentHistoryStore) GetRecentChanges(string, time.Time, int) ([]unifiedresources.ResourceChange, error) { + return nil, errors.New("history store unavailable") +} + +type incidentArchiveFixture struct{ window *metrics.IncidentWindow } + +func (s incidentArchiveFixture) GetWindow(string, string) (*metrics.IncidentWindow, error) { + return s.window, nil +} + +func TestIncidentHistoryRetainsCanonicalEvidence(t *testing.T) { + store, err := unifiedresources.NewSQLiteResourceStore(t.TempDir(), "incident-test") + require.NoError(t, err) + t.Cleanup(func() { require.NoError(t, store.Close()) }) + resourceID := "app-container-7020f37498275208" + start := time.Now().UTC().Add(-time.Hour).Truncate(time.Second) + occurred := start.Add(-time.Minute) + changes := []unifiedresources.ResourceChange{ + {ID: "outside-window", ResourceID: resourceID, ObservedAt: start.Add(-time.Second), Kind: unifiedresources.ChangeRestart, SourceType: unifiedresources.SourcePlatformEvent}, + {ID: "alert-fired", ResourceID: resourceID, ObservedAt: start.Add(time.Minute), OccurredAt: &occurred, Kind: unifiedresources.ChangeAlertFired, SourceType: unifiedresources.SourcePulseDiff, From: "healthy", To: "unhealthy", Reason: "Health check failed", Metadata: map[string]any{"alert_id": "health-check", "value": 1.0}}, + {ID: "related-only", ResourceID: "agent-parent", RelatedResources: []string{resourceID}, ObservedAt: start.Add(2 * time.Minute), Kind: unifiedresources.ChangeRestart, SourceType: unifiedresources.SourcePlatformEvent}, + {ID: "alert-resolved", ResourceID: resourceID, ObservedAt: start.Add(3 * time.Minute), Kind: unifiedresources.ChangeAlertResolved, SourceType: unifiedresources.SourcePulseDiff, From: "unhealthy", To: "healthy"}, + } + for _, change := range changes { + require.NoError(t, store.RecordChange(change)) + } + // No live inventory or incident recorder is necessary to read a retained + // event for a container that has since been removed. + exec := NewPulseToolExecutor(ExecutorConfig{ActionAuditStore: store, IncidentArchiveProvider: incidentArchiveFixture{}}) + require.True(t, exec.isToolAvailable(agentcapabilities.PulseKnowledgeToolName)) + for _, tc := range []struct { + name string + id string + limit int + wantCount int + wantMore bool + }{ + {"retained lifecycle", resourceID, 50, 2, false}, + {"bounded lifecycle", resourceID, 1, 1, true}, + {"empty history", "app-container-absent", 50, 0, false}, + } { + t.Run(tc.name, func(t *testing.T) { + input := map[string]interface{}{"action": "incidents", "resource_id": tc.id, "since": start.Format(time.RFC3339), "limit": float64(tc.limit)} + result, err := exec.registry.Execute(context.Background(), exec, agentcapabilities.PulseKnowledgeToolName, input) + require.NoError(t, err) + require.False(t, result.IsError, result.Content) + var got struct { + Source string `json:"source"` + Events []unifiedresources.ResourceChange `json:"events"` + HasMore bool `json:"has_more"` + Coverage string `json:"coverage"` + TimeBasis string `json:"time_basis"` + } + require.NoError(t, json.Unmarshal([]byte(result.Content[0].Text), &got)) + require.Equal(t, "canonical_resource_timeline", got.Source) + require.Equal(t, "retained_records_only", got.Coverage) + require.Equal(t, "observed_at", got.TimeBasis) + require.NotNil(t, got.Events) + require.Len(t, got.Events, tc.wantCount) + require.Equal(t, tc.wantMore, got.HasMore) + if tc.wantCount > 0 { + require.Equal(t, "alert-resolved", got.Events[0].ID) + require.Nil(t, got.Events[0].OccurredAt) + } + if tc.wantCount == 2 { + require.Equal(t, changes[1], got.Events[1]) + } + capture, err := json.Marshal(map[string]any{"case": tc.name, "input": input, "result": result}) + require.NoError(t, err) + t.Logf("INCIDENT_EVIDENCE %s", capture) + }) + } +} + +func TestIncidentHistoryUnavailableAndInvalid(t *testing.T) { + for _, tc := range []struct { + name string + store unifiedresources.ResourceStore + input map[string]interface{} + }{ + {"unavailable", nil, map[string]interface{}{"resource_id": "app-container-1"}}, + {"failed", failedIncidentHistoryStore{}, map[string]interface{}{"resource_id": "app-container-1"}}, + {"empty resource", unifiedresources.NewMemoryStore(), map[string]interface{}{"resource_id": " "}}, + {"invalid time", unifiedresources.NewMemoryStore(), map[string]interface{}{"resource_id": "app-container-1", "since": "yesterday"}}, + {"future time", unifiedresources.NewMemoryStore(), map[string]interface{}{"resource_id": "app-container-1", "since": time.Now().Add(time.Hour).Format(time.RFC3339)}}, + {"negative limit", unifiedresources.NewMemoryStore(), map[string]interface{}{"resource_id": "app-container-1", "limit": float64(-1)}}, + {"large limit", unifiedresources.NewMemoryStore(), map[string]interface{}{"resource_id": "app-container-1", "limit": float64(201)}}, + } { + t.Run(tc.name, func(t *testing.T) { + exec := NewPulseToolExecutor(ExecutorConfig{ActionAuditStore: tc.store, IncidentArchiveProvider: incidentArchiveFixture{}}) + tc.input["action"] = "incidents" + result, err := exec.registry.Execute(context.Background(), exec, agentcapabilities.PulseKnowledgeToolName, tc.input) + require.NoError(t, err) + require.True(t, result.IsError, result.Content) + capture, err := json.Marshal(map[string]any{"case": tc.name, "input": tc.input, "result": result}) + require.NoError(t, err) + t.Logf("INCIDENT_EVIDENCE %s", capture) + }) + } +} + +func TestIncidentHistoryLegacyArchiveIsResourceBound(t *testing.T) { + window := &metrics.IncidentWindow{ID: "archive-1", ResourceID: "vm-1"} + exec := NewPulseToolExecutor(ExecutorConfig{IncidentArchiveProvider: incidentArchiveFixture{window}}) + for _, id := range []string{"vm-1", "vm-2"} { + result, err := exec.executeGetIncidentWindow(context.Background(), map[string]interface{}{"resource_id": id, "window_id": window.ID}) + require.NoError(t, err) + if id == window.ResourceID { + require.False(t, result.IsError) + require.Contains(t, result.Content[0].Text, "legacy_incident_recording") + require.Contains(t, result.Content[0].Text, "cached observations") + } else { + require.True(t, result.IsError) + require.NotContains(t, result.Content[0].Text, "vm-1") + } + } + result, err := exec.executeGetIncidentWindow(context.Background(), map[string]interface{}{"resource_id": window.ResourceID, "window_id": "wrong-window"}) + require.NoError(t, err) + require.True(t, result.IsError, "a provider cannot substitute another archived window") + +} + +func TestIncidentHistoryArchiveReadOutcomes(t *testing.T) { + dir := t.TempDir() + archivePath := filepath.Join(dir, "incident_windows.json") + archive := metrics.NewIncidentArchive(dir) + exec := NewPulseToolExecutor(ExecutorConfig{IncidentArchiveProvider: archive}) + original := `{"completed_windows":[{"id":"saved-window","resource_id":"vm-archive","status":"recording","start_time":"2020-01-02T03:04:05Z","data_points":[{"timestamp":"2020-01-02T03:04:06Z","metrics":{"cpu":12.5},"metadata":{"source":"cached"}}],"summary":{"duration_ms":60000000000,"anomalies":["stored observation"]}}]}` + for _, tc := range []struct { + name, raw, resource, window string + wantError bool + }{ + {"archive unavailable", "", "vm-archive", "saved-window", true}, + {"archive malformed", "{", "vm-archive", "saved-window", true}, + {"archive success", original, "vm-archive", "saved-window", false}, + {"archive wrong resource", original, "vm-other", "saved-window", true}, + {"archive missing window", original, "vm-archive", "missing", true}, + } { + t.Run(tc.name, func(t *testing.T) { + if tc.raw != "" { + require.NoError(t, os.WriteFile(archivePath, []byte(tc.raw), 0600)) + } + input := map[string]interface{}{"action": "incidents", "resource_id": tc.resource, "window_id": tc.window} + result, err := exec.registry.Execute(context.Background(), exec, agentcapabilities.PulseKnowledgeToolName, input) + require.NoError(t, err) + require.Equal(t, tc.wantError, result.IsError, result.Content) + if !tc.wantError { + var got struct { + Window *metrics.IncidentWindow `json:"window"` + ReadOnly bool `json:"archive_read_only"` + DurationUnit string `json:"summary_duration_unit"` + } + require.NoError(t, json.Unmarshal([]byte(result.Content[0].Text), &got)) + require.True(t, got.ReadOnly) + require.Equal(t, "nanoseconds", got.DurationUnit) + require.Equal(t, time.Minute, got.Window.Summary.Duration) + require.Equal(t, metrics.IncidentWindowStatusRecording, got.Window.Status) + require.Equal(t, "cached", got.Window.DataPoints[0].Metadata["source"]) + require.Equal(t, []string{"stored observation"}, got.Window.Summary.Anomalies) + require.Contains(t, result.Content[0].Text, "does not mean recording is active") + } else if tc.name == "archive wrong resource" { + require.NotContains(t, result.Content[0].Text, "stored observation") + } + capture, err := json.Marshal(map[string]any{"case": tc.name, "input": input, "result": result}) + require.NoError(t, err) + t.Logf("INCIDENT_EVIDENCE %s", capture) + }) + } +} diff --git a/internal/ai/tools/tmpfs_mount_evidence_test.go b/internal/ai/tools/tmpfs_mount_evidence_test.go new file mode 100644 index 000000000..34615d657 --- /dev/null +++ b/internal/ai/tools/tmpfs_mount_evidence_test.go @@ -0,0 +1,87 @@ +package tools + +import ( + "context" + "encoding/json" + "testing" + + "github.com/rcourtman/pulse-go-rewrite/internal/models" + "github.com/rcourtman/pulse-go-rewrite/internal/unifiedresources" +) + +func TestQueryPreservesMountConfigurationEvidence(t *testing.T) { + for _, provider := range []bool{false, true} { + name := "typed read state" + if provider { + name = "canonical provider" + } + t.Run(name, func(t *testing.T) { + snapshot := commandEvidenceSnapshot() + snapshot.DockerHosts[0].Containers[0].Mounts = []models.DockerContainerMount{ + {Type: "tmpfs", Destination: "/var/lib/service-cache", Mode: "rw,noexec,nosuid,nodev,size=8388608", RW: true}, + {Type: "tmpfs", Destination: "/readonly-cache", Mode: "ro,noexec", RW: false}, + {Type: "bind", Source: "/host/data", Destination: "/data", Mode: "", RW: false}, + } + registry := unifiedresources.NewRegistry(nil) + registry.IngestSnapshot(snapshot) + cfg := ExecutorConfig{ReadState: registry, ControlLevel: ControlLevelReadOnly} + if provider { + cfg.UnifiedResourceProvider = ®istryUnifiedQueryProvider{registry} + } + executor := NewPulseToolExecutor(cfg) + list, err := executor.executeQuery(context.Background(), map[string]interface{}{"action": "list", "type": "docker-hosts"}) + if err != nil || list.IsError { + t.Fatalf("list: %v %+v", err, list) + } + var hosts map[string]any + if err := json.Unmarshal([]byte(list.Content[0].Text), &hosts); err != nil { + t.Fatal(err) + } + id := hosts["docker_hosts"].([]any)[0].(map[string]any)["containers"].([]any)[0].(map[string]any)["id"] + if !provider { + // The typed compatibility path currently accepts provider IDs/names. + id = snapshot.DockerHosts[0].Containers[0].Name + } + args := map[string]interface{}{"action": "get", "resource_type": "app-container", "resource_id": id} + result, err := executor.executeQuery(context.Background(), args) + if err != nil || result.IsError { + t.Fatalf("get: %v %+v", err, result) + } + var decoded map[string]any + if err := json.Unmarshal([]byte(result.Content[0].Text), &decoded); err != nil { + t.Fatal(err) + } + mounts, ok := decoded["mounts"].([]any) + if !ok { + t.Fatalf("missing mount projection: %+v", decoded) + } + if len(mounts) != 3 { + t.Fatalf("lost mounts: %+v", mounts) + } + for i, want := range snapshot.DockerHosts[0].Containers[0].Mounts { + got := mounts[i].(map[string]any) + if got["type"] != want.Type || got["source"] != want.Source || got["destination"] != want.Destination || got["rw"] != want.RW || (want.Mode != "" && got["mode"] != want.Mode) { + t.Fatalf("mount provenance/access changed: %+v, want %+v", got, want) + } + } + if _, exists := decoded["disk"]; exists { + t.Fatalf("mount configuration invented capacity: %+v", decoded["disk"]) + } + capture, _ := json.Marshal(map[string]any{"case": name, "input": args, "output": decoded}) + t.Logf("MOUNT_EVIDENCE %s", capture) + if provider { + resources := registry.ListByType(unifiedresources.ResourceTypeAppContainer) + for _, host := range registry.ListByType(unifiedresources.ResourceTypeAgent) { + if host.Docker != nil { + resources = append(resources, host) + } + } + encoded, err := json.Marshal(resources) + if err != nil { + t.Fatal(err) + } + t.Logf("MOUNT_RESOURCES %s", encoded) + } + }) + } +} diff --git a/internal/ai/tools/tools_knowledge.go b/internal/ai/tools/tools_knowledge.go index 58c987764..a32c544d8 100644 --- a/internal/ai/tools/tools_knowledge.go +++ b/internal/ai/tools/tools_knowledge.go @@ -3,46 +3,18 @@ package tools import ( "context" "fmt" + "strings" "time" "github.com/rcourtman/pulse-go-rewrite/internal/agentcapabilities" + "github.com/rcourtman/pulse-go-rewrite/internal/metrics" + "github.com/rcourtman/pulse-go-rewrite/internal/unifiedresources" ) -// IncidentRecorderProvider provides access to incident recording data -type IncidentRecorderProvider interface { - GetWindowsForResource(resourceID string, limit int) []*IncidentWindow - GetWindow(windowID string) *IncidentWindow -} - -// IncidentWindow represents a high-frequency recording window during an incident -type IncidentWindow struct { - ID string `json:"id"` - ResourceID string `json:"resource_id"` - ResourceName string `json:"resource_name,omitempty"` - ResourceType string `json:"resource_type,omitempty"` - TriggerType string `json:"trigger_type"` - TriggerID string `json:"trigger_id,omitempty"` - StartTime time.Time `json:"start_time"` - EndTime *time.Time `json:"end_time,omitempty"` - Status string `json:"status"` - DataPoints []IncidentDataPoint `json:"data_points"` - Summary *IncidentSummary `json:"summary,omitempty"` -} - -// IncidentDataPoint represents a single data point in an incident window -type IncidentDataPoint struct { - Timestamp time.Time `json:"timestamp"` - Metrics map[string]float64 `json:"metrics"` -} - -// IncidentSummary provides computed statistics about an incident window -type IncidentSummary struct { - Duration time.Duration `json:"duration_ms"` - DataPoints int `json:"data_points"` - Peaks map[string]float64 `json:"peaks"` - Lows map[string]float64 `json:"lows"` - Averages map[string]float64 `json:"averages"` - Changes map[string]float64 `json:"changes"` +// IncidentArchiveProvider provides explicit, resource-bound reads of saved +// legacy recordings. Live incident evidence comes from the canonical timeline. +type IncidentArchiveProvider interface { + GetWindow(resourceID, windowID string) (*metrics.IncidentWindow, error) } // EventCorrelatorProvider provides access to correlated events @@ -87,7 +59,7 @@ func (e *PulseToolExecutor) registerKnowledgeTools() { Actions: - remember: Save a note about a resource for future reference - recall: Retrieve saved notes about a resource -- incidents: Get high-resolution incident recording data +- incidents: Read retained canonical resource history, including observed state changes, alerts and executed actions. Records preserve observation time, source and any known occurrence time. This is not continuous health or filesystem-capacity coverage. Use pulse_summarize for retained metrics. - correlate: Get correlated events around a timestamp Examples: @@ -105,7 +77,7 @@ Examples: }, "resource_id": { Type: "string", - Description: "Resource ID to operate on", + Description: "Resource ID to operate on. For incidents use the canonical resource ID returned by pulse_query, including for a resource no longer in current inventory.", }, "note": { Type: "string", @@ -117,7 +89,11 @@ Examples: }, "window_id": { Type: "string", - Description: "For incidents: specific incident window ID", + Description: "For incidents: optional legacy recording ID, read as an archive only. Omit to read canonical resource history.", + }, + "since": { + Type: "string", + Description: "For incidents: earliest observation timestamp (RFC3339, default 24 hours ago). Retention and collection gaps still apply.", }, "timestamp": { Type: "string", @@ -129,7 +105,7 @@ Examples: }, "limit": { Type: "integer", - Description: "For incidents: max windows to return (default: 5)", + Description: "For incidents: maximum retained events to return, newest first (default 50, range 1-200)", }, }, Required: []string{"action", "resource_id"}, @@ -168,38 +144,82 @@ func (e *PulseToolExecutor) executeKnowledge(ctx context.Context, args map[strin func (e *PulseToolExecutor) executeGetIncidentWindow(_ context.Context, args map[string]interface{}) (CallToolResult, error) { resourceID, _ := args["resource_id"].(string) + resourceID = strings.TrimSpace(resourceID) windowID, _ := args["window_id"].(string) - limit := intArg(args, "limit", 5) if resourceID == "" { return NewErrorResult(fmt.Errorf("resource_id is required")), nil } - if e.incidentRecorderProvider == nil { - return NewTextResult("Incident recording data not available. The incident recorder may not be enabled."), nil - } - - // If a specific window ID is requested + // Isolate legacy recordings from the canonical timeline. Their sample times + // are recorder timestamps, not verified source observation timestamps. if windowID != "" { - window := e.incidentRecorderProvider.GetWindow(windowID) - if window == nil { - return NewTextResult(fmt.Sprintf("Incident window '%s' not found.", windowID)), nil + if e.incidentArchiveProvider == nil { + return NewErrorResult(fmt.Errorf("legacy incident recording archive is unavailable")), nil + } + window, err := e.incidentArchiveProvider.GetWindow(resourceID, windowID) + if err != nil { + return NewErrorResult(fmt.Errorf("read legacy incident recording archive: %w", err)), nil + } + if window == nil || window.ResourceID != resourceID || window.ID != windowID { + return NewErrorResult(fmt.Errorf("legacy incident recording not found for the requested resource")), nil } return NewJSONResult(map[string]interface{}{ - "window": window, + "source": "legacy_incident_recording", + "archive_read_only": true, + "summary_duration_unit": "nanoseconds", + "window": window, + "evidence_limit": "Recording timestamps do not establish when the source measured each value. Repeated values may be cached observations. The legacy summary.duration_ms field contains nanoseconds. Stored recording status is historical and does not mean recording is active. This archive is not the canonical incident timeline.", }), nil } - // Get windows for the resource - windows := e.incidentRecorderProvider.GetWindowsForResource(resourceID, limit) - if len(windows) == 0 { - return NewTextResult(fmt.Sprintf("No incident recording data found for resource '%s'. Incident data is captured when alerts fire.", resourceID)), nil + limit := intArg(args, "limit", 50) + if limit < 1 || limit > 200 { + return NewErrorResult(fmt.Errorf("limit must be between 1 and 200")), nil + } + queriedAt := time.Now().UTC() + since := queriedAt.Add(-24 * time.Hour) + if value, exists := args["since"]; exists { + text, ok := value.(string) + if !ok { + return NewErrorResult(fmt.Errorf("since must be an RFC3339 timestamp")), nil + } + var err error + since, err = time.Parse(time.RFC3339, text) + if err != nil || since.After(queriedAt) { + return NewErrorResult(fmt.Errorf("since must be an RFC3339 timestamp no later than now")), nil + } + } + if e.actionAuditStore == nil { + return NewErrorResult(fmt.Errorf("canonical resource history is unavailable")), nil + } + // This organization-pinned store is also used by the resource history API + // and Assistant handoffs. Do not reconstruct history from current metrics, + // match resource names, or include adjacent resources implicitly. + events, err := e.actionAuditStore.GetRecentChanges(resourceID, since, limit+1) + if err != nil { + return NewErrorResult(fmt.Errorf("read canonical resource history: %w", err)), nil + } + hasMore := len(events) > limit + if hasMore { + events = events[:limit] + } + if events == nil { + events = []unifiedresources.ResourceChange{} } return NewJSONResult(map[string]interface{}{ - "resource_id": resourceID, - "windows": windows, - "count": len(windows), + "resource_id": resourceID, + "source": "canonical_resource_timeline", + "since": since, + "queried_at": queriedAt, + "time_basis": "observed_at", + "events": events, + "count": len(events), + "limit": limit, + "has_more": hasMore, + "coverage": "retained_records_only", + "evidence_limit": "These are retained observations, not continuous coverage. Empty history does not establish health or absence of incidents. ObservedAt is when Pulse observed a change, while OccurredAt is present only when its occurrence time is known. An alert resolving establishes that alert's recovery, not its cause or a verified action outcome.", }), nil } diff --git a/internal/ai/tools/tools_propose.go b/internal/ai/tools/tools_propose.go index 07dfcc626..7588b14e1 100644 --- a/internal/ai/tools/tools_propose.go +++ b/internal/ai/tools/tools_propose.go @@ -178,7 +178,7 @@ func (e *PulseToolExecutor) executeProposeAction(ctx context.Context, args map[s return NewErrorResult(err), nil } return NewTextResult(fmt.Sprintf( - "Proposal recorded: capability %q on resource %q. It will be planned and routed for governed approval; nothing has executed. Conclude the investigation with your diagnosis.", + "Proposal recorded: capability %q on resource %q. The action broker still needs to validate it after this investigation. No action has been created or executed.", capabilityName, resourceID)), nil } diff --git a/internal/ai/tools/tools_query.go b/internal/ai/tools/tools_query.go index 7654a9f8e..2fbe1bc4d 100644 --- a/internal/ai/tools/tools_query.go +++ b/internal/ai/tools/tools_query.go @@ -2151,7 +2151,7 @@ func (e *PulseToolExecutor) registerQueryTools() { e.registry.registerBuiltin(RegisteredTool{ Definition: Tool{ Name: agentcapabilities.PulseQueryToolName, - Description: `Query and search canonical infrastructure resources. Start here to discover systems, workloads, storage, and disks by name. Actions: search, get, config, topology, list, health. Health returns the connection overview by default, or the canonical resource projection when resource_id is provided.`, + Description: `Query and search canonical infrastructure resources. Start here to discover systems, workloads, storage, and disks by name. Actions: search, get, config, topology, list, health. Health returns the connection overview by default, or the canonical resource projection when resource_id is provided. command_agent_connected describes live command transport, independently of monitoring collection or freshness. Missing connection fields were not observed. can_execute describes connected transport with control enabled, not approval for a particular operation.`, InputSchema: InputSchema{ Type: "object", Properties: map[string]PropertySchema{ @@ -2166,7 +2166,7 @@ func (e *PulseToolExecutor) registerQueryTools() { }, "resource_type": { Type: "string", - Description: "Resource type. For get/search, prefer canonical values: 'agent', 'vm', 'system-container', 'app-container', 'storage', 'physical-disk', and 'docker-host'. For get, 'node' resolves to 'agent'. For search, 'node' filters Proxmox nodes. Compatibility aliases 'system' and 'storage-pool' are still accepted. For config: 'vm', 'system-container', or supported API-backed 'app-container'.", + Description: "Resource type. For get/search, prefer canonical values: 'agent', 'vm', 'system-container', 'app-container', 'storage', 'physical-disk', and 'docker-host'. For get, 'node' resolves to 'agent'. For search, 'node' filters Proxmox nodes. Compatibility aliases 'system' and 'storage-pool' are still accepted. For config: 'vm', 'system-container', or TrueNAS 'app-container'. Docker and Podman app-container configuration reads are not supported. Their collected health, mounts, ports, and networks are available through get.", Enum: []string{"agent", "system", "vm", "system-container", "app-container", "node", "docker-host", "storage", "storage-pool", "physical-disk"}, }, "resource_id": { @@ -2463,17 +2463,35 @@ func resourceHostCandidates(resource unifiedresources.Resource) []string { return candidates } -func resourceAgentConnected(resource unifiedresources.Resource, connected map[string]bool) bool { - for _, candidate := range resourceHostCandidates(resource) { - key := strings.TrimSpace(candidate) - if key == "" { - continue - } - if connected[key] { - return true +// commandConnectionObservation keeps an unqueried snapshot distinct from an +// observed disconnected transport. It says nothing about telemetry freshness. +func commandConnectionObservation(snapshot map[string]bool, value bool) *bool { + if snapshot == nil { + return nil + } + return &value +} + +func resourceCommandAgentConnected(resource unifiedresources.Resource, connected map[string]bool) *bool { + // A guest's provider node/host identifies placement, not a command + // connection inside the guest. Parent transport is projected separately. + candidates := []string{resourceDisplayName(resource)} + candidates = append(candidates, resource.Identity.Hostnames...) + if resource.Agent != nil { + candidates = append(candidates, resource.Agent.Hostname) + } + switch resource.Type { + case unifiedresources.ResourceTypeVM, unifiedresources.ResourceTypeSystemContainer, unifiedresources.ResourceTypeAppContainer: + // Only the guest's own identity can establish its direct connection. + default: + candidates = append(candidates, resourceHostCandidates(resource)...) + } + for _, candidate := range candidates { + if key := strings.TrimSpace(candidate); key != "" && connected[key] { + return commandConnectionObservation(connected, true) } } - return false + return commandConnectionObservation(connected, false) } func appContainerProviderID(resource unifiedresources.Resource) string { @@ -2972,16 +2990,16 @@ func addCanonicalGuestSearchMatches( node := canonicalGuestTarget(resource) metadataCandidates := append([]string{resourceDisplayName(resource), resource.ID}, candidates...) addMatch(ResourceMatch{ - GovernedResourceMetadata: governance.Resolve(metadataCandidates...), - Type: kind, - ID: resource.ID, - Name: resourceDisplayName(resource), - Status: status, - Node: node, - NodeHasAgent: connectedAgentHostnames[node], - Platform: canonicalResourcePlatform(resource), - VMID: vmid, - AgentConnected: resourceAgentConnected(resource, connectedAgentHostnames), + GovernedResourceMetadata: governance.Resolve(metadataCandidates...), + Type: kind, + ID: resource.ID, + Name: resourceDisplayName(resource), + Status: status, + Node: node, + NodeCommandAgentConnected: commandConnectionObservation(connectedAgentHostnames, connectedAgentHostnames[node]), + Platform: canonicalResourcePlatform(resource), + VMID: vmid, + CommandAgentConnected: resourceCommandAgentConnected(resource, connectedAgentHostnames), }) } } @@ -3012,16 +3030,16 @@ func addGuestViewSearchMatches[V queryGuestView]( continue } addMatch(ResourceMatch{ - GovernedResourceMetadata: governance.Resolve(g.Name(), g.ID(), vmidStr), - Type: kind, - ID: g.ID(), - Name: g.Name(), - Status: status, - Node: g.Node(), - NodeHasAgent: connectedAgentHostnames[g.Node()], - Platform: "proxmox", - VMID: g.VMID(), - AgentConnected: connectedAgentHostnames[g.Name()], + GovernedResourceMetadata: governance.Resolve(g.Name(), g.ID(), vmidStr), + Type: kind, + ID: g.ID(), + Name: g.Name(), + Status: status, + Node: g.Node(), + NodeCommandAgentConnected: commandConnectionObservation(connectedAgentHostnames, connectedAgentHostnames[g.Node()]), + Platform: "proxmox", + VMID: g.VMID(), + CommandAgentConnected: commandConnectionObservation(connectedAgentHostnames, connectedAgentHostnames[g.Name()]), }) } } @@ -3469,15 +3487,15 @@ func resolvedAppContainerRegistration(resource unifiedresources.Resource) (Resou func canonicalSystemSummaryFromResource(resource unifiedresources.Resource, connected map[string]bool) SystemSummary { return SystemSummary{ - ID: strings.TrimSpace(resource.ID), - Name: resourceDisplayName(resource), - Status: string(resource.Status), - Platform: canonicalResourcePlatform(resource), - ChildCount: resource.ChildCount, - AgentConnected: resourceAgentConnected(resource, connected), - CPU: metricPercent(resourceMetric(resource, "cpu")), - Memory: metricPercent(resourceMetric(resource, "memory")), - Disk: metricPercent(resourceMetric(resource, "disk")), + ID: strings.TrimSpace(resource.ID), + Name: resourceDisplayName(resource), + Status: string(resource.Status), + Platform: canonicalResourcePlatform(resource), + ChildCount: resource.ChildCount, + CommandAgentConnected: resourceCommandAgentConnected(resource, connected), + CPU: metricPercent(resourceMetric(resource, "cpu")), + Memory: metricPercent(resourceMetric(resource, "memory")), + Disk: metricPercent(resourceMetric(resource, "disk")), } } @@ -3809,7 +3827,7 @@ func (e *PulseToolExecutor) executeListInfrastructure(_ context.Context, args ma GovernedResourceMetadata: governance.Resolve(node.Name(), node.ID()), Name: node.Name(), Status: string(node.Status()), - AgentConnected: connectedAgentHostnames[node.Name()], + CommandAgentConnected: commandConnectionObservation(connectedAgentHostnames, connectedAgentHostnames[node.Name()]), }) count++ } @@ -3993,7 +4011,7 @@ func (e *PulseToolExecutor) executeListInfrastructure(_ context.Context, args ma Hostname: hostname, DisplayName: displayName, ContainerCount: len(hostContainers), - AgentConnected: connectedAgentHostnames[hostname] || connectedAgentHostnames[displayName], + CommandAgentConnected: commandConnectionObservation(connectedAgentHostnames, connectedAgentHostnames[hostname] || connectedAgentHostnames[displayName]), } for _, container := range hostContainers { state := strings.TrimSpace(container.ContainerState()) @@ -4248,7 +4266,7 @@ type TopologyBuildOptions struct { MaxK8sNodesPerCluster int MaxK8sDeploymentsPerCluster int MaxK8sPodsPerCluster int - ConnectedAgentHostnames map[string]bool + ConnectedAgentHostnames map[string]bool // nil means command connections were not observed ControlEnabled bool } @@ -4266,9 +4284,6 @@ func BuildTopologyResponseFromReadState(rs unifiedresources.ReadState, options T includeDocker := include == "all" || include == "app-containers" includeKubernetes := include == "all" || include == "kubernetes" connectedAgentHostnames := options.ConnectedAgentHostnames - if connectedAgentHostnames == nil { - connectedAgentHostnames = map[string]bool{} - } governance := newGovernedQueryMetadataResolver(rs) summary := TopologySummary{ @@ -4283,12 +4298,17 @@ func BuildTopologyResponseFromReadState(rs unifiedresources.ReadState, options T TotalK8sPods: len(rs.Pods()), } + if connectedAgentHostnames != nil { + summary.NodesWithCommandAgents = new(int) + summary.DockerHostsWithCommandAgents = new(int) + } + for _, node := range rs.Nodes() { if node == nil { continue } if connectedAgentHostnames[node.Name()] { - summary.NodesWithAgents++ + (*summary.NodesWithCommandAgents)++ } } for _, host := range rs.DockerHosts() { @@ -4298,7 +4318,7 @@ func BuildTopologyResponseFromReadState(rs unifiedresources.ReadState, options T hostname := strings.TrimSpace(host.Hostname()) displayName := strings.TrimSpace(host.Name()) if connectedAgentHostnames[hostname] || connectedAgentHostnames[displayName] { - summary.DockerHostsWithAgents++ + (*summary.DockerHostsWithCommandAgents)++ } } for _, pod := range rs.Pods() { @@ -4325,8 +4345,8 @@ func BuildTopologyResponseFromReadState(rs unifiedresources.ReadState, options T GovernedResourceMetadata: governance.Resolve(node.Name(), node.ID()), Name: name, Status: string(node.Status()), - AgentConnected: hasAgent, - CanExecute: hasAgent && options.ControlEnabled, + CommandAgentConnected: commandConnectionObservation(connectedAgentHostnames, hasAgent), + CanExecute: commandConnectionObservation(connectedAgentHostnames, hasAgent && options.ControlEnabled), VMs: []TopologyVM{}, Containers: []TopologyContainer{}, } @@ -4348,8 +4368,8 @@ func BuildTopologyResponseFromReadState(rs unifiedresources.ReadState, options T GovernedResourceMetadata: governance.Resolve(name), Name: name, Status: status, - AgentConnected: hasAgent, - CanExecute: hasAgent && options.ControlEnabled, + CommandAgentConnected: commandConnectionObservation(connectedAgentHostnames, hasAgent), + CanExecute: commandConnectionObservation(connectedAgentHostnames, hasAgent && options.ControlEnabled), VMs: []TopologyVM{}, Containers: []TopologyContainer{}, } @@ -4489,8 +4509,8 @@ func BuildTopologyResponseFromReadState(rs unifiedresources.ReadState, options T GovernedResourceMetadata: governance.Resolve(host.Hostname(), host.Name(), host.HostSourceID(), host.ID()), Hostname: hostname, DisplayName: displayName, - AgentConnected: hasAgent, - CanExecute: hasAgent && options.ControlEnabled, + CommandAgentConnected: commandConnectionObservation(connectedAgentHostnames, hasAgent), + CanExecute: commandConnectionObservation(connectedAgentHostnames, hasAgent && options.ControlEnabled), Containers: containers, ContainerCount: len(hostContainers), ReturnedCount: len(containers), @@ -5037,9 +5057,11 @@ func (e *PulseToolExecutor) executeGetResource(_ context.Context, args map[strin } for _, m := range resource.Docker.Mounts { response.Mounts = append(response.Mounts, MountInfo{ + Type: m.Type, Source: m.Source, Destination: m.Destination, - ReadWrite: !strings.EqualFold(strings.TrimSpace(m.Mode), "ro"), + Mode: m.Mode, + ReadWrite: m.RW, }) } } @@ -5144,8 +5166,10 @@ func (e *PulseToolExecutor) executeGetResource(_ context.Context, args map[strin for _, m := range container.Mounts() { response.Mounts = append(response.Mounts, MountInfo{ + Type: m.Type, Source: m.Source, Destination: m.Destination, + Mode: m.Mode, ReadWrite: m.RW, }) } @@ -5276,110 +5300,117 @@ func (e *PulseToolExecutor) executeGetResourceConfig(ctx context.Context, args m } func (e *PulseToolExecutor) executeNativeAppContainerConfig(ctx context.Context, resourceRef string) (CallToolResult, error) { - if e.appContainerConfigProvider == nil { - return NewTextResult("App-container configuration not available."), nil + validation := e.validateResolvedResource(resourceRef, "query", true) + if validation.ErrorMsg != "" { + return NewErrorResult(fmt.Errorf("%s", validation.ErrorMsg)), nil + } + if validation.Resource != nil && validation.Resource.GetKind() != "app-container" { + return NewErrorResult(fmt.Errorf("resource '%s' is %q, not app-container", resourceRef, validation.Resource.GetKind())), nil + } + if e.unifiedResourceProvider == nil { + return NewErrorResult(fmt.Errorf("current app-container inventory is unavailable")), nil } rs, err := e.readStateForControl() if err != nil { - return NewTextResult("State information not available."), nil + return NewErrorResult(fmt.Errorf("current app-container state is unavailable: %w", err)), nil } governance := newGovernedQueryMetadataResolver(rs) - var resource unifiedresources.Resource - var found bool - if validation := e.validateResolvedResource(resourceRef, "query", true); validation.Resource != nil { - if matched, _, ok := findCanonicalAppContainerResource(e.unifiedResourceProvider, resourceRef); ok { - resource = matched - found = true - } - } + resource, providerID, found := findCanonicalAppContainerResource(e.unifiedResourceProvider, resourceRef) if !found { - var containerID string - resource, containerID, found = findCanonicalAppContainerResource(e.unifiedResourceProvider, resourceRef) - if !found { - return NewJSONResult(map[string]interface{}{ - "error": "not_found", - "resource_id": resourceRef, - "type": "app-container", - }), nil - } - if reg, ok := resolvedAppContainerRegistration(resource); ok { - e.registerResolvedResourceWithExplicitAccess(reg) - } - _ = containerID + return NewJSONResult(map[string]interface{}{ + "error": "not_found", + "resource_id": resourceRef, + "type": "app-container", + }), nil } - validation := e.validateResolvedResource(resourceRef, "query", true) - if validation.Resource == nil { - if validation.ErrorMsg != "" { - return NewErrorResult(fmt.Errorf("%s", validation.ErrorMsg)), nil - } - return NewErrorResult(fmt.Errorf("app-container not found: %s", resourceRef)), nil + // Inventory owns read identity and capability. Optional session discovery + // supplies restrictions and continuity, not proof that the resource exists. + resourceID := canonicalAppContainerID(resource) + canonicalValidation := e.validateResolvedResource(resourceID, "query", true) + if canonicalValidation.ErrorMsg != "" { + return NewErrorResult(fmt.Errorf("%s", canonicalValidation.ErrorMsg)), nil } - if validation.ErrorMsg != "" { - return NewErrorResult(fmt.Errorf("%s", validation.ErrorMsg)), nil + if canonicalValidation.Resource != nil && canonicalValidation.Resource.GetKind() != "app-container" { + return NewErrorResult(fmt.Errorf("resource '%s' is %q, not app-container", resourceID, canonicalValidation.Resource.GetKind())), nil } - resolved := validation.Resource - if resolved.GetKind() != "app-container" { - return NewErrorResult(fmt.Errorf("resource '%s' is %q, not app-container", resourceRef, resolved.GetKind())), nil + platform := canonicalAppContainerAdapter(resource) + unavailable := func(reason, message string) CallToolResult { + return NewJSONResultWithIsError(map[string]interface{}{ + "available": false, "reason": reason, "message": message, + "resource_id": resourceID, "type": "app-container", "platform": platform, + }, true) } - if !strings.EqualFold(strings.TrimSpace(resolved.GetAdapter()), "truenas") { - return NewTextResult("App-container configuration not available."), nil + if platform != "truenas" { + return unavailable("unsupported_adapter", "The resource exists, but its adapter does not support configuration reads."), nil + } + if e.appContainerConfigProvider == nil { + return unavailable("provider_unavailable", "The resource exists, but its configuration provider is unavailable."), nil + } + reg, ok := resolvedAppContainerRegistration(resource) + if !ok { + return unavailable("resource_context_unavailable", "The resource exists, but its current provider identity or placement is incomplete."), nil + } + // Do not overwrite an existing session's allowed actions during a read. + if validation.Resource == nil && canonicalValidation.Resource == nil { + e.registerResolvedResourceWithExplicitAccess(reg) } result, err := e.appContainerConfigProvider.GetConfig(ctx, AppContainerConfigRequest{ OrgID: e.orgID, - ResourceID: strings.TrimSpace(resolved.GetResourceID()), - ProviderUID: strings.TrimSpace(resolved.GetProviderUID()), + ResourceID: resourceID, + ProviderUID: providerID, Name: resourceDisplayName(resource), - Host: strings.TrimSpace(resolved.GetTargetHost()), - Platform: "truenas", + Host: canonicalAppContainerHost(resource), + Platform: platform, }) if err != nil { return NewErrorResult(err), nil } + if result == nil { + return unavailable("empty_provider_response", "The resource exists, but the provider returned no configuration observation."), nil + } response := EmptyAppContainerConfigResponse() - if result != nil { - response.GovernedResourceMetadata = governance.Resolve(result.Name, result.ResourceID, result.ProviderUID) - response.Type = "app-container" - response.ID = result.ProviderUID - if response.ID == "" { - response.ID = strings.TrimSpace(result.ResourceID) - } - response.Name = result.Name - response.Host = result.Host - response.Platform = result.Platform - response.Status = result.Status - response.Version = result.Version - response.HumanVersion = result.HumanVersion - response.Notes = result.Notes - response.CustomApp = result.CustomApp - response.UpgradeAvailable = result.UpgradeAvailable - response.ImageUpdatesAvailable = result.ImageUpdatesAvailable - response.ContainerCount = result.ContainerCount - response.UsedHostIPs = append([]string{}, result.UsedHostIPs...) - response.Images = append([]string{}, result.Images...) - response.Ports = append([]PortInfo{}, result.Ports...) - response.Networks = append([]NetworkInfo{}, result.Networks...) - response.Mounts = append([]MountInfo{}, result.Mounts...) - response.Containers = append([]AppContainerConfigContainer{}, result.Containers...) - } + response.GovernedResourceMetadata = governance.Resolve(result.Name, result.ResourceID, result.ProviderUID) + response.Type = "app-container" + response.ID = result.ProviderUID if response.ID == "" { - response.ID = strings.TrimSpace(resolved.GetProviderUID()) + response.ID = strings.TrimSpace(result.ResourceID) + } + response.Name = result.Name + response.Host = result.Host + response.Platform = result.Platform + response.Status = result.Status + response.Version = result.Version + response.HumanVersion = result.HumanVersion + response.Notes = result.Notes + response.CustomApp = result.CustomApp + response.UpgradeAvailable = result.UpgradeAvailable + response.ImageUpdatesAvailable = result.ImageUpdatesAvailable + response.ContainerCount = result.ContainerCount + response.UsedHostIPs = append([]string{}, result.UsedHostIPs...) + response.Images = append([]string{}, result.Images...) + response.Ports = append([]PortInfo{}, result.Ports...) + response.Networks = append([]NetworkInfo{}, result.Networks...) + response.Mounts = append([]MountInfo{}, result.Mounts...) + response.Containers = append([]AppContainerConfigContainer{}, result.Containers...) + if response.ID == "" { + response.ID = providerID } if response.Name == "" { - response.Name = resolvedResourceDisplayName(resolved) + response.Name = resourceDisplayName(resource) } if response.Host == "" { - response.Host = strings.TrimSpace(resolved.GetTargetHost()) + response.Host = canonicalAppContainerHost(resource) } if response.Platform == "" { - response.Platform = strings.TrimSpace(resolved.GetAdapter()) + response.Platform = platform } if response.GovernedResourceMetadata.Policy == nil && response.AISafeSummary == "" { - response.GovernedResourceMetadata = governance.Resolve(response.Name, strings.TrimSpace(resolved.GetResourceID()), response.ID) + response.GovernedResourceMetadata = governance.Resolve(response.Name, resourceID, response.ID) } return NewJSONResult(response.NormalizeCollections()), nil @@ -5714,7 +5745,7 @@ func (e *PulseToolExecutor) executeSearchResources(_ context.Context, args map[s Status: status, Host: canonicalAgentHost(resource), Platform: canonicalResourcePlatform(resource), - AgentConnected: resourceAgentConnected(resource, connectedAgentHostnames), + CommandAgentConnected: resourceCommandAgentConnected(resource, connectedAgentHostnames), }) } } @@ -5733,7 +5764,7 @@ func (e *PulseToolExecutor) executeSearchResources(_ context.Context, args map[s Type: "node", Name: node.Name(), Status: status, - AgentConnected: connectedAgentHostnames[node.Name()], + CommandAgentConnected: commandConnectionObservation(connectedAgentHostnames, connectedAgentHostnames[node.Name()]), }) } } diff --git a/internal/ai/tools/tools_query_config_test.go b/internal/ai/tools/tools_query_config_test.go index b6f338a3e..37e78df67 100644 --- a/internal/ai/tools/tools_query_config_test.go +++ b/internal/ai/tools/tools_query_config_test.go @@ -3,13 +3,18 @@ package tools import ( "context" "encoding/json" + "errors" + "strings" "testing" + + "github.com/rcourtman/pulse-go-rewrite/internal/unifiedresources" ) type stubAppContainerConfigProvider struct { calls []AppContainerConfigRequest result *AppContainerConfigResult err error + empty bool } func (s *stubAppContainerConfigProvider) GetConfig(_ context.Context, req AppContainerConfigRequest) (*AppContainerConfigResult, error) { @@ -17,6 +22,9 @@ func (s *stubAppContainerConfigProvider) GetConfig(_ context.Context, req AppCon if s.err != nil { return nil, s.err } + if s.empty { + return nil, nil + } if s.result == nil { return &AppContainerConfigResult{ ResourceID: req.ResourceID, @@ -38,6 +46,153 @@ func (s *stubAppContainerConfigProvider) GetConfig(_ context.Context, req AppCon return &result, nil } +func TestAppContainerConfigObservationContract(t *testing.T) { + t.Setenv("PULSE_STRICT_RESOLUTION", "true") + for _, tc := range []struct { + name, session, reason, failure string + missing, docker, noProvider, noHost, empty bool + noInventory, noReadState bool + }{ + {name: "without session"}, + {name: "empty session", session: "empty"}, + {name: "discovered session", session: "discovered"}, + {name: "stale placement", session: "stale"}, + {name: "unsupported adapter", docker: true, reason: "unsupported_adapter"}, + {name: "unavailable provider", noProvider: true, reason: "provider_unavailable"}, + {name: "missing placement", noHost: true, reason: "resource_context_unavailable"}, + {name: "empty provider response", empty: true, reason: "empty_provider_response"}, + {name: "provider failed", failure: "provider read failed"}, + {name: "resource absent", missing: true}, + {name: "inventory unavailable", noInventory: true, failure: "inventory is unavailable"}, + {name: "read state unavailable", noReadState: true, failure: "state is unavailable"}, + {name: "query denied", session: "denied", failure: "not permitted"}, + {name: "canonical query denied through prefix", session: "canonical-denied", failure: "not permitted"}, + } { + t.Run(tc.name, func(t *testing.T) { + registry := newTrueNASUnifiedQueryProvider(t) + resource, _, found := findCanonicalAppContainerResource(registry, "nextcloud") + if !found { + t.Fatal("missing canonical fixture") + } + if tc.docker { + resource.TrueNAS = nil + resource.Tags = nil + } + if tc.noHost { + resource.ParentName = "" + resource.Identity.Hostnames = nil + } + provider := &stubUnifiedResourceProvider{resources: []unifiedresources.Resource{resource}} + config := &stubAppContainerConfigProvider{empty: tc.empty} + if tc.failure == "provider read failed" { + config.err = errors.New(tc.failure) + } + cfg := ExecutorConfig{UnifiedResourceProvider: provider, ReadState: registry.ResourceRegistry} + if tc.noInventory { + cfg.UnifiedResourceProvider = nil + } + if tc.noReadState { + cfg.ReadState = nil + } + if !tc.noProvider { + cfg.AppContainerConfigProvider = config + } + executor := NewPulseToolExecutor(cfg) + ref := "Nextcloud" + if tc.session != "" { + resolved := &mockResolvedContext{resources: map[string]ResolvedResourceInfo{}, aliases: map[string]ResolvedResourceInfo{}} + executor.SetResolvedContext(resolved) + switch tc.session { + case "discovered": + reg, ok := resolvedAppContainerRegistration(resource) + if !ok { + t.Fatal("fixture registration unavailable") + } + resolved.AddResolvedResource(reg) + case "stale", "denied", "canonical-denied": + cached := &mockResource{resourceID: resource.ID, kind: "app-container", adapter: "docker", targetHost: "stale-host", providerUID: "stale-id", allowedActions: []string{"query"}} + if tc.session != "stale" { + cached.allowedActions = []string{"logs"} + } + resolved.resources[resource.ID] = cached + if tc.session == "canonical-denied" { + ref = "next" + } else { + resolved.aliases[ref] = cached + } + } + } + if tc.missing { + ref = "absent-container" + } + // Inventory get succeeds independently of optional session state. + if tc.session == "" && !tc.missing && !tc.noInventory && !tc.noReadState { + got, err := executor.executeGetResource(context.Background(), map[string]interface{}{"resource_type": "app-container", "resource_id": ref}) + if err != nil || got.IsError || strings.Contains(got.Content[0].Text, "not_found") { + t.Fatalf("canonical get failed: %+v %v", got, err) + } + } + args := map[string]interface{}{"action": "config", "resource_type": "app-container", "resource_id": ref} + result, err := executor.executeQuery(context.Background(), args) + if err != nil { + t.Fatal(err) + } + evidence, err := json.Marshal(map[string]interface{}{"case": tc.name, "input": args, "result": result}) + if err != nil { + t.Fatal(err) + } + t.Logf("CONFIG_EVIDENCE %s", evidence) + wantError := tc.failure != "" || tc.reason != "" + if result.IsError != wantError { + t.Fatalf("read error bit=%v, want %v: %+v", result.IsError, wantError, result) + } + if tc.failure != "" { + if !result.IsError || !strings.Contains(result.Content[0].Text, tc.failure) { + t.Fatalf("expected %q failure, got %+v", tc.failure, result) + } + } else { + var response map[string]interface{} + if err := json.Unmarshal([]byte(result.Content[0].Text), &response); err != nil { + t.Fatal(err) + } + switch { + case tc.missing: + if response["error"] != "not_found" { + t.Fatalf("expected true absence, got %+v", response) + } + case tc.reason != "": + if response["available"] != false || response["reason"] != tc.reason || response["resource_id"] != resource.ID { + t.Fatalf("unavailable configuration lost identity or reason: %+v", response) + } + default: + if response["id"] != appContainerProviderID(resource) || response["host"] != canonicalAppContainerHost(resource) || response["platform"] != "truenas" { + t.Fatalf("incorrect config identity: %+v", response) + } + } + } + wantCalls := 1 + if tc.missing || tc.docker || tc.noProvider || tc.noHost || tc.noInventory || tc.noReadState || strings.Contains(tc.session, "denied") { + wantCalls = 0 + } + if len(config.calls) != wantCalls { + t.Fatalf("provider calls=%d, want %d", len(config.calls), wantCalls) + } + if wantCalls == 1 { + call := config.calls[0] + if call.ResourceID != resource.ID || call.ProviderUID != appContainerProviderID(resource) || call.Host != canonicalAppContainerHost(resource) || call.Platform != "truenas" { + t.Fatalf("request used session identity instead of canonical inventory: %+v", call) + } + } + if tc.session == "stale" { + cached, ok := executor.resolvedContext.GetResolvedResourceByID(resource.ID) + if !ok || strings.Join(cached.GetAllowedActions(), ",") != "query" { + t.Fatal("read expanded existing session action authority") + } + } + }) + } +} + func TestExecuteGetResourceConfig_TrueNASAppUsesNativeConfigProvider(t *testing.T) { provider := newTrueNASUnifiedQueryProvider(t) resolved := &mockResolvedContext{ diff --git a/internal/api/ai_handler.go b/internal/api/ai_handler.go index f7ec08655..aff2273af 100644 --- a/internal/api/ai_handler.go +++ b/internal/api/ai_handler.go @@ -78,7 +78,7 @@ type AIService interface { SetFindingsManager(manager chat.FindingsManager) SetMetadataUpdater(updater chat.MetadataUpdater) SetKnowledgeStoreProvider(provider chat.KnowledgeStoreProvider) - SetIncidentRecorderProvider(provider chat.IncidentRecorderProvider) + SetIncidentArchiveProvider(provider chat.IncidentArchiveProvider) SetEventCorrelatorProvider(provider chat.EventCorrelatorProvider) SetDiscoveryProvider(provider chat.AssistantDiscoveryProvider) SetUnifiedResourceProvider(provider chat.AssistantUnifiedResourceProvider) diff --git a/internal/api/ai_handler_recovery_wiring_test.go b/internal/api/ai_handler_recovery_wiring_test.go index 8a15ba9c3..c04cd38bf 100644 --- a/internal/api/ai_handler_recovery_wiring_test.go +++ b/internal/api/ai_handler_recovery_wiring_test.go @@ -88,15 +88,15 @@ func (s *capturingAIService) SetGuestConfigProvider(provider chat.AssistantGuest func (s *capturingAIService) SetAppContainerConfigProvider(provider chat.AssistantAppContainerConfigProvider) { s.appContainerConfigProvider = provider } -func (s *capturingAIService) SetBackupProvider(provider chat.AssistantBackupProvider) {} -func (s *capturingAIService) SetDiskHealthProvider(provider chat.AssistantDiskHealthProvider) {} -func (s *capturingAIService) SetUpdatesProvider(provider chat.AssistantUpdatesProvider) {} -func (s *capturingAIService) SetFindingsManager(manager chat.FindingsManager) {} -func (s *capturingAIService) SetMetadataUpdater(updater chat.MetadataUpdater) {} -func (s *capturingAIService) SetKnowledgeStoreProvider(provider chat.KnowledgeStoreProvider) {} -func (s *capturingAIService) SetIncidentRecorderProvider(provider chat.IncidentRecorderProvider) {} -func (s *capturingAIService) SetEventCorrelatorProvider(provider chat.EventCorrelatorProvider) {} -func (s *capturingAIService) SetDiscoveryProvider(provider chat.AssistantDiscoveryProvider) {} +func (s *capturingAIService) SetBackupProvider(provider chat.AssistantBackupProvider) {} +func (s *capturingAIService) SetDiskHealthProvider(provider chat.AssistantDiskHealthProvider) {} +func (s *capturingAIService) SetUpdatesProvider(provider chat.AssistantUpdatesProvider) {} +func (s *capturingAIService) SetFindingsManager(manager chat.FindingsManager) {} +func (s *capturingAIService) SetMetadataUpdater(updater chat.MetadataUpdater) {} +func (s *capturingAIService) SetKnowledgeStoreProvider(provider chat.KnowledgeStoreProvider) {} +func (s *capturingAIService) SetIncidentArchiveProvider(provider chat.IncidentArchiveProvider) {} +func (s *capturingAIService) SetEventCorrelatorProvider(provider chat.EventCorrelatorProvider) {} +func (s *capturingAIService) SetDiscoveryProvider(provider chat.AssistantDiscoveryProvider) {} func (s *capturingAIService) SetUnifiedResourceProvider(provider chat.AssistantUnifiedResourceProvider) { } func (s *capturingAIService) SetAppContainerActionProvider(provider chat.AssistantAppContainerActionProvider) { diff --git a/internal/api/ai_handler_test.go b/internal/api/ai_handler_test.go index 5cd14e0cc..ad572870c 100644 --- a/internal/api/ai_handler_test.go +++ b/internal/api/ai_handler_test.go @@ -281,7 +281,7 @@ func (m *MockAIService) SetMetadataUpdater(updater chat.MetadataUpdater) { m.Cal func (m *MockAIService) SetKnowledgeStoreProvider(provider chat.KnowledgeStoreProvider) { m.Called(provider) } -func (m *MockAIService) SetIncidentRecorderProvider(provider chat.IncidentRecorderProvider) { +func (m *MockAIService) SetIncidentArchiveProvider(provider chat.IncidentArchiveProvider) { m.Called(provider) } func (m *MockAIService) SetEventCorrelatorProvider(provider chat.EventCorrelatorProvider) { diff --git a/internal/api/ai_handlers.go b/internal/api/ai_handlers.go index 2a89407cd..193f9b45a 100644 --- a/internal/api/ai_handlers.go +++ b/internal/api/ai_handlers.go @@ -92,22 +92,20 @@ type AISettingsHandler struct { alertBridge *unified.AlertBridge // Bridge between alerts and unified store // Event-driven patrol (Phase 7) - triggerManager *ai.TriggerManager // Event-driven patrol trigger manager - incidentCoordinator *ai.IncidentCoordinator // Incident recording coordinator - incidentRecorder *metrics.IncidentRecorder // High-frequency incident recorder - intelligenceMu sync.RWMutex - proxmoxCorrelators map[string]*proxmox.EventCorrelator - learningStores map[string]*learning.LearningStore - forecastServices map[string]*forecast.Service - remediationEngines map[string]aicontracts.RemediationEngine - incidentStores map[string]*memory.IncidentStore - circuitBreakers map[string]*circuit.Breaker - discoveryStores map[string]*servicediscovery.Store - unifiedStores map[string]*unified.UnifiedStore - alertBridges map[string]*unified.AlertBridge - triggerManagers map[string]*ai.TriggerManager - incidentCoordinators map[string]*ai.IncidentCoordinator - incidentRecorders map[string]*metrics.IncidentRecorder + triggerManager *ai.TriggerManager // Event-driven patrol trigger manager + incidentArchive *metrics.IncidentArchive // Read-only legacy incident archive + intelligenceMu sync.RWMutex + proxmoxCorrelators map[string]*proxmox.EventCorrelator + learningStores map[string]*learning.LearningStore + forecastServices map[string]*forecast.Service + remediationEngines map[string]aicontracts.RemediationEngine + incidentStores map[string]*memory.IncidentStore + circuitBreakers map[string]*circuit.Breaker + discoveryStores map[string]*servicediscovery.Store + unifiedStores map[string]*unified.UnifiedStore + alertBridges map[string]*unified.AlertBridge + triggerManagers map[string]*ai.TriggerManager + incidentArchives map[string]*metrics.IncidentArchive // Investigation orchestration (Patrol Autonomy) chatHandler *AIHandler // Chat service handler for investigations @@ -318,25 +316,24 @@ func NewAISettingsHandler(mtp *config.MultiTenantPersistence, mtm *monitoring.Mu } handler := &AISettingsHandler{ - mtPersistence: mtp, - mtMonitor: mtm, - defaultConfig: defaultConfig, - defaultPersistence: defaultPersistence, - hostedMode: hostedMode, - aiServices: make(map[string]*ai.Service), - agentServer: agentServer, - proxmoxCorrelators: make(map[string]*proxmox.EventCorrelator), - learningStores: make(map[string]*learning.LearningStore), - forecastServices: make(map[string]*forecast.Service), - remediationEngines: make(map[string]aicontracts.RemediationEngine), - incidentStores: make(map[string]*memory.IncidentStore), - circuitBreakers: make(map[string]*circuit.Breaker), - discoveryStores: make(map[string]*servicediscovery.Store), - unifiedStores: make(map[string]*unified.UnifiedStore), - alertBridges: make(map[string]*unified.AlertBridge), - triggerManagers: make(map[string]*ai.TriggerManager), - incidentCoordinators: make(map[string]*ai.IncidentCoordinator), - incidentRecorders: make(map[string]*metrics.IncidentRecorder), + mtPersistence: mtp, + mtMonitor: mtm, + defaultConfig: defaultConfig, + defaultPersistence: defaultPersistence, + hostedMode: hostedMode, + aiServices: make(map[string]*ai.Service), + agentServer: agentServer, + proxmoxCorrelators: make(map[string]*proxmox.EventCorrelator), + learningStores: make(map[string]*learning.LearningStore), + forecastServices: make(map[string]*forecast.Service), + remediationEngines: make(map[string]aicontracts.RemediationEngine), + incidentStores: make(map[string]*memory.IncidentStore), + circuitBreakers: make(map[string]*circuit.Breaker), + discoveryStores: make(map[string]*servicediscovery.Store), + unifiedStores: make(map[string]*unified.UnifiedStore), + alertBridges: make(map[string]*unified.AlertBridge), + triggerManagers: make(map[string]*ai.TriggerManager), + incidentArchives: make(map[string]*metrics.IncidentArchive), } defaultAIService = ai.NewService(defaultPersistence, tenantAgentServerForOrganization(agentServer, "default")) @@ -1198,11 +1195,8 @@ func (h *AISettingsHandler) ensureIntelligenceMapsLocked() { if h.triggerManagers == nil { h.triggerManagers = make(map[string]*ai.TriggerManager) } - if h.incidentCoordinators == nil { - h.incidentCoordinators = make(map[string]*ai.IncidentCoordinator) - } - if h.incidentRecorders == nil { - h.incidentRecorders = make(map[string]*metrics.IncidentRecorder) + if h.incidentArchives == nil { + h.incidentArchives = make(map[string]*metrics.IncidentArchive) } } @@ -1505,96 +1499,49 @@ func (h *AISettingsHandler) GetTriggerManagerForOrg(orgID string) *ai.TriggerMan return nil } -// SetIncidentCoordinator sets the incident recording coordinator -func (h *AISettingsHandler) SetIncidentCoordinator(coordinator *ai.IncidentCoordinator) { - h.SetIncidentCoordinatorForOrg("default", coordinator) +// SetIncidentArchive sets the read-only legacy incident archive +func (h *AISettingsHandler) SetIncidentArchive(archive *metrics.IncidentArchive) { + h.SetIncidentArchiveForOrg("default", archive) } -// SetIncidentCoordinatorForOrg sets the incident recording coordinator for an org. -func (h *AISettingsHandler) SetIncidentCoordinatorForOrg(orgID string, coordinator *ai.IncidentCoordinator) { +// SetIncidentArchiveForOrg sets the read-only legacy incident archive for an org. +func (h *AISettingsHandler) SetIncidentArchiveForOrg(orgID string, archive *metrics.IncidentArchive) { if h == nil { return } orgID = normalizeAIIntelligenceOrgID(orgID) h.intelligenceMu.Lock() h.ensureIntelligenceMapsLocked() - if coordinator == nil { - delete(h.incidentCoordinators, orgID) + if archive == nil { + delete(h.incidentArchives, orgID) } else { - h.incidentCoordinators[orgID] = coordinator + h.incidentArchives[orgID] = archive } h.intelligenceMu.Unlock() if orgID == "default" { - h.incidentCoordinator = coordinator + h.incidentArchive = archive } } -// GetIncidentCoordinator returns the incident recording coordinator -func (h *AISettingsHandler) GetIncidentCoordinator() *ai.IncidentCoordinator { - return h.GetIncidentCoordinatorForOrg("default") +// GetIncidentArchive returns the read-only legacy incident archive +func (h *AISettingsHandler) GetIncidentArchive() *metrics.IncidentArchive { + return h.GetIncidentArchiveForOrg("default") } -// GetIncidentCoordinatorForOrg returns the incident recording coordinator for an org. -func (h *AISettingsHandler) GetIncidentCoordinatorForOrg(orgID string) *ai.IncidentCoordinator { +// GetIncidentArchiveForOrg returns the read-only legacy incident archive for an org. +func (h *AISettingsHandler) GetIncidentArchiveForOrg(orgID string) *metrics.IncidentArchive { if h == nil { return nil } orgID = normalizeAIIntelligenceOrgID(orgID) h.intelligenceMu.RLock() - if coordinator := h.incidentCoordinators[orgID]; coordinator != nil { + if archive := h.incidentArchives[orgID]; archive != nil { h.intelligenceMu.RUnlock() - return coordinator + return archive } h.intelligenceMu.RUnlock() if orgID == "default" { - return h.incidentCoordinator - } - return nil -} - -// SetIncidentRecorder sets the high-frequency incident recorder -func (h *AISettingsHandler) SetIncidentRecorder(recorder *metrics.IncidentRecorder) { - h.SetIncidentRecorderForOrg("default", recorder) -} - -// SetIncidentRecorderForOrg sets the high-frequency incident recorder for an org. -func (h *AISettingsHandler) SetIncidentRecorderForOrg(orgID string, recorder *metrics.IncidentRecorder) { - if h == nil { - return - } - orgID = normalizeAIIntelligenceOrgID(orgID) - h.intelligenceMu.Lock() - h.ensureIntelligenceMapsLocked() - if recorder == nil { - delete(h.incidentRecorders, orgID) - } else { - h.incidentRecorders[orgID] = recorder - } - h.intelligenceMu.Unlock() - if orgID == "default" { - h.incidentRecorder = recorder - } -} - -// GetIncidentRecorder returns the high-frequency incident recorder -func (h *AISettingsHandler) GetIncidentRecorder() *metrics.IncidentRecorder { - return h.GetIncidentRecorderForOrg("default") -} - -// GetIncidentRecorderForOrg returns the high-frequency incident recorder for an org. -func (h *AISettingsHandler) GetIncidentRecorderForOrg(orgID string) *metrics.IncidentRecorder { - if h == nil { - return nil - } - orgID = normalizeAIIntelligenceOrgID(orgID) - h.intelligenceMu.RLock() - if recorder := h.incidentRecorders[orgID]; recorder != nil { - h.intelligenceMu.RUnlock() - return recorder - } - h.intelligenceMu.RUnlock() - if orgID == "default" { - return h.incidentRecorder + return h.incidentArchive } return nil } @@ -1656,44 +1603,6 @@ func (h *AISettingsHandler) ListTriggerManagers() map[string]*ai.TriggerManager return out } -// ListIncidentCoordinators returns incident coordinators keyed by org. -func (h *AISettingsHandler) ListIncidentCoordinators() map[string]*ai.IncidentCoordinator { - out := make(map[string]*ai.IncidentCoordinator) - if h == nil { - return out - } - h.intelligenceMu.RLock() - for orgID, coordinator := range h.incidentCoordinators { - if coordinator != nil { - out[orgID] = coordinator - } - } - h.intelligenceMu.RUnlock() - if _, ok := out["default"]; !ok && h.incidentCoordinator != nil { - out["default"] = h.incidentCoordinator - } - return out -} - -// ListIncidentRecorders returns incident recorders keyed by org. -func (h *AISettingsHandler) ListIncidentRecorders() map[string]*metrics.IncidentRecorder { - out := make(map[string]*metrics.IncidentRecorder) - if h == nil { - return out - } - h.intelligenceMu.RLock() - for orgID, recorder := range h.incidentRecorders { - if recorder != nil { - out[orgID] = recorder - } - } - h.intelligenceMu.RUnlock() - if _, ok := out["default"]; !ok && h.incidentRecorder != nil { - out["default"] = h.incidentRecorder - } - return out -} - // StopPatrol stops the background AI patrol service func (h *AISettingsHandler) StopPatrol() { if h.defaultAIService != nil { @@ -1760,10 +1669,8 @@ func (h *AISettingsHandler) RemoveTenantIntelligence(orgID string) { } var ( - bridge *unified.AlertBridge - trigger *ai.TriggerManager - coordinator *ai.IncidentCoordinator - recorder *metrics.IncidentRecorder + bridge *unified.AlertBridge + trigger *ai.TriggerManager ) h.intelligenceMu.Lock() @@ -1775,13 +1682,10 @@ func (h *AISettingsHandler) RemoveTenantIntelligence(orgID string) { delete(h.discoveryStores, orgID) bridge = h.alertBridges[orgID] trigger = h.triggerManagers[orgID] - coordinator = h.incidentCoordinators[orgID] - recorder = h.incidentRecorders[orgID] delete(h.unifiedStores, orgID) delete(h.alertBridges, orgID) delete(h.triggerManagers, orgID) - delete(h.incidentCoordinators, orgID) - delete(h.incidentRecorders, orgID) + delete(h.incidentArchives, orgID) delete(h.proxmoxCorrelators, orgID) h.intelligenceMu.Unlock() @@ -1791,12 +1695,6 @@ func (h *AISettingsHandler) RemoveTenantIntelligence(orgID string) { if trigger != nil { trigger.Stop() } - if coordinator != nil { - coordinator.Stop() - } - if recorder != nil { - recorder.Stop() - } } // GetAlertTriggeredAnalyzer returns the alert-triggered analyzer for wiring into alert callbacks @@ -7843,6 +7741,10 @@ func (h *AISettingsHandler) HandleApproveCommand(w http.ResponseWriter, r *http. // updateFindingOutcome updates the investigation outcome on a finding func (h *AISettingsHandler) updateFindingOutcome(ctx context.Context, orgID, findingID, outcome string) { + h.updateFindingInvestigationOutcome(ctx, orgID, findingID, outcome, nil) +} + +func (h *AISettingsHandler) updateFindingInvestigationOutcome(ctx context.Context, orgID, findingID, outcome string, investigation *ai.InvestigationSession) { // Get AI service for this org svc := h.GetAIService(ctx) if svc == nil { @@ -7861,15 +7763,20 @@ func (h *AISettingsHandler) updateFindingOutcome(ctx context.Context, orgID, fin log.Warn().Str("orgID", orgID).Msg("Findings store not available for finding update") return } - if existing := findingsStore.Get(findingID); existing != nil && existing.InvestigationOutcome == outcome { + existing := findingsStore.Get(findingID) + if existing == nil { return } - - if !findingsStore.UpdateInvestigationOutcome(findingID, outcome) { + outcomeChanged := existing.InvestigationOutcome != outcome + if outcomeChanged && !findingsStore.UpdateInvestigationOutcome(findingID, outcome) { log.Warn().Str("findingID", findingID).Msg("Finding not found for outcome update") return } - patrol.PublishFindingLifecycleUpdate(findingID) + recordChanged := investigation != nil && patrol.RefreshFindingInvestigationRecord(findingID, investigation) + if !outcomeChanged && !recordChanged { + return + } + patrol.PublishFindingLifecycleUpdate(findingID, outcomeChanged) log.Info().Str("findingID", findingID).Str("outcome", outcome).Msg("Updated finding investigation outcome") } diff --git a/internal/api/ai_handlers_investigation_additional_test.go b/internal/api/ai_handlers_investigation_additional_test.go index e0dbbc94c..bd48a1202 100644 --- a/internal/api/ai_handlers_investigation_additional_test.go +++ b/internal/api/ai_handlers_investigation_additional_test.go @@ -258,6 +258,59 @@ func TestPatrolActionReconciliationCannotRegressFromOutOfOrderCallback(t *testin } } +func TestPatrolActionHydrationRepairsDurableFindingRecordWithoutRewritingEvidence(t *testing.T) { + investigations := newTestInvestigationStore() + investigation := investigations.Create("finding-1", "session-1") + investigation.Status = aicontracts.InvestigationStatusCompleted + investigation.Outcome = aicontracts.OutcomeFixQueued + investigation.Summary = "Cause uncertain. Restart proposed for review." + investigation.EvidenceIDs = []string{"observed-health"} + investigations.Update(investigation) + svc := ai.NewService(nil, nil) + svc.SetStateProvider(&MockStateProvider{}) + patrol := svc.GetPatrolService() + findings := patrol.GetFindings() + findings.Add(&ai.Finding{ID: "finding-1", ResourceID: "vm:42", Title: "Unhealthy service", Severity: ai.FindingSeverityWarning, + InvestigationStatus: string(investigation.Status), InvestigationOutcome: string(aicontracts.OutcomeFixRejected)}) + // Reproduce the persisted mismatch after an outcome already reconciled. + record := ai.BuildFindingInvestigationRecord(findings.Get("finding-1"), investigation) + record.Rollback = []string{"Retained rollback evidence"} + record.Impact = "Original service impact absent from the later finding projection." + findings.UpdateInvestigationRecord("finding-1", record) + audits := unifiedresources.NewMemoryStore() + audit := unifiedresources.ActionAuditRecord{ + ID: "act-1", CreatedAt: time.Now().UTC(), UpdatedAt: time.Now().UTC(), State: unifiedresources.ActionStateRejected, + Request: unifiedresources.ActionRequest{RequestID: "proposal-1", ResourceID: "vm:42", CapabilityName: "restart", RequestedBy: "pulse_patrol"}, + Plan: unifiedresources.ActionPlan{ActionID: "act-1", RequestID: "proposal-1", Allowed: true}, + Origin: &unifiedresources.ActionOrigin{Surface: patrolActionOriginSurface, FindingID: "finding-1", InvestigationID: investigation.ID, ProposalID: "proposal-1"}, + } + if _, _, err := audits.CreateActionAudit(audit, nil); err != nil { + t.Fatal(err) + } + handler := &AISettingsHandler{defaultAIService: svc, + investigationStores: map[string]aicontracts.InvestigationStore{"default": investigations}, + resourceStoreProvider: func(string) (unifiedresources.ResourceStore, error) { return audits, nil }, + } + var published []*ai.Finding + patrol.SetUnifiedFindingCallback(func(f *ai.Finding) bool { published = append(published, f); return true }) + published = nil // Ignore the initial synchronization when registering the callback. + handler.hydratePatrolInvestigationAction("default", investigation) + got := findings.Get("finding-1").InvestigationRecord + if got.Outcome != aicontracts.OutcomeFixRejected || got.Action == nil || got.Action.State != "rejected" { + t.Fatalf("durable record did not reconcile: %#v", got) + } + if got.Conclusion != investigation.Summary || got.Impact != record.Impact || got.Confidence != record.Confidence || !reflect.DeepEqual(got.Rollback, record.Rollback) || !reflect.DeepEqual(got.Evidence, record.Evidence) { + t.Fatalf("retained investigation evidence changed: %#v", got) + } + if len(got.Verification) != 1 || len(published) != 1 || published[0].InvestigationRecord.Outcome != aicontracts.OutcomeFixRejected { + t.Fatalf("reconciled record was not published: record=%#v published=%#v", got, published) + } + handler.hydratePatrolInvestigationAction("default", investigation) + if len(published) != 1 { + t.Fatal("duplicate hydration republished unchanged record") + } +} + func TestPatrolActionReconciliationHydratesTerminalAuditAfterRestart(t *testing.T) { investigations := newTestInvestigationStore() investigation := investigations.Create("finding-1", "session-1") diff --git a/internal/api/ai_handlers_setters_additional_test.go b/internal/api/ai_handlers_setters_additional_test.go index 6095aee26..6b9d9e286 100644 --- a/internal/api/ai_handlers_setters_additional_test.go +++ b/internal/api/ai_handlers_setters_additional_test.go @@ -50,16 +50,10 @@ func TestAISettingsHandler_SettersAndGetters(t *testing.T) { t.Fatalf("GetTriggerManager returned unexpected manager") } - coordinator := ai.NewIncidentCoordinator(ai.IncidentCoordinatorConfig{}) - handler.SetIncidentCoordinator(coordinator) - if handler.GetIncidentCoordinator() != coordinator { - t.Fatalf("GetIncidentCoordinator returned unexpected coordinator") - } - - recorder := &metrics.IncidentRecorder{} - handler.SetIncidentRecorder(recorder) - if handler.GetIncidentRecorder() != recorder { - t.Fatalf("GetIncidentRecorder returned unexpected recorder") + recorder := &metrics.IncidentArchive{} + handler.SetIncidentArchive(recorder) + if handler.GetIncidentArchive() != recorder { + t.Fatalf("GetIncidentArchive returned unexpected recorder") } handler.WireOrchestratorAfterChatStart() @@ -82,15 +76,19 @@ func TestAISettingsHandler_IntelligenceServicesAreOrgScoped(t *testing.T) { t.Fatalf("expected nil correlator for unrelated org, got %#v", got) } - defaultRecorder := &metrics.IncidentRecorder{} - tenantRecorder := &metrics.IncidentRecorder{} - handler.SetIncidentRecorderForOrg("default", defaultRecorder) - handler.SetIncidentRecorderForOrg("acme", tenantRecorder) - if got := handler.GetIncidentRecorder(); got != defaultRecorder { - t.Fatalf("expected default recorder, got %#v", got) + defaultArchive := &metrics.IncidentArchive{} + tenantArchive := &metrics.IncidentArchive{} + handler.SetIncidentArchiveForOrg("default", defaultArchive) + handler.SetIncidentArchiveForOrg("acme", tenantArchive) + if got := handler.GetIncidentArchive(); got != defaultArchive { + t.Fatalf("expected default archive, got %#v", got) } - if got := handler.GetIncidentRecorderForOrg("acme"); got != tenantRecorder { - t.Fatalf("expected tenant recorder, got %#v", got) + if got := handler.GetIncidentArchiveForOrg("acme"); got != tenantArchive { + t.Fatalf("expected tenant archive, got %#v", got) + } + + if got := handler.GetIncidentArchiveForOrg("unrelated"); got != nil { + t.Fatalf("unrelated org received an archive: %#v", got) } defaultLearningStore := learning.NewLearningStore(learning.LearningStoreConfig{}) @@ -201,13 +199,12 @@ func TestAISettingsHandler_IntelligenceServicesAreOrgScoped(t *testing.T) { func TestAISettingsHandler_RemoveTenantService_TrimsOrgID(t *testing.T) { handler := &AISettingsHandler{ - aiServices: map[string]*ai.Service{"acme": nil}, - investigationStores: map[string]aicontracts.InvestigationStore{"acme": nil}, - proxmoxCorrelators: map[string]*proxmox.EventCorrelator{"acme": nil}, - alertBridges: map[string]*unified.AlertBridge{"acme": nil}, - triggerManagers: map[string]*ai.TriggerManager{"acme": nil}, - incidentCoordinators: map[string]*ai.IncidentCoordinator{"acme": nil}, - incidentRecorders: map[string]*metrics.IncidentRecorder{"acme": nil}, + aiServices: map[string]*ai.Service{"acme": nil}, + investigationStores: map[string]aicontracts.InvestigationStore{"acme": nil}, + proxmoxCorrelators: map[string]*proxmox.EventCorrelator{"acme": nil}, + alertBridges: map[string]*unified.AlertBridge{"acme": nil}, + triggerManagers: map[string]*ai.TriggerManager{"acme": nil}, + incidentArchives: map[string]*metrics.IncidentArchive{"acme": nil}, } handler.RemoveTenantService(" acme ") @@ -227,10 +224,7 @@ func TestAISettingsHandler_RemoveTenantService_TrimsOrgID(t *testing.T) { if _, ok := handler.triggerManagers["acme"]; ok { t.Fatalf("expected trigger manager entry to be removed") } - if _, ok := handler.incidentCoordinators["acme"]; ok { - t.Fatalf("expected incident coordinator entry to be removed") - } - if _, ok := handler.incidentRecorders["acme"]; ok { + if _, ok := handler.incidentArchives["acme"]; ok { t.Fatalf("expected incident recorder entry to be removed") } } diff --git a/internal/api/ai_handlers_test.go b/internal/api/ai_handlers_test.go index 3eb28535d..9fa31ac08 100644 --- a/internal/api/ai_handlers_test.go +++ b/internal/api/ai_handlers_test.go @@ -3754,14 +3754,10 @@ func TestPatrolModelReadinessBudgetScalesWithRequestTimeout(t *testing.T) { assert.Equal(t, 4*600*time.Second+time.Minute, patrolModelReadinessBudget(&config.AIConfig{RequestTimeoutSeconds: 600})) } -// TestOrchestratorAndChatAdaptersMapTheSameMessageFields keeps the deliberate -// GetMessages mirror between orchestratorChatAdapter (ai_handlers.go) and -// chatServiceAdapter (chat_service_adapter.go) honest: both convert the same -// chat-service messages onto separate output contracts, and a field mapped by -// one must be mapped by the other. chatServiceAdapter routes through -// adaptChatMessage so its API-facing tool calls stay on the shared provider -// shape instead of hand-copying a local duplicate. -func TestOrchestratorAndChatAdaptersMapTheSameMessageFields(t *testing.T) { +// The adapters share base message fields but have distinct tool-call contracts. +// Orchestrator turns project provider arguments. Product history retains the +// observed result through the canonical result-bearing transcript type. +func TestOrchestratorAndChatAdaptersMapTheirMessageContracts(t *testing.T) { for _, tc := range []struct { file string fn string @@ -3793,7 +3789,7 @@ func TestOrchestratorAndChatAdaptersMapTheSameMessageFields(t *testing.T) { "Content:", "ReasoningContent:", "Timestamp:", - "tc.ProviderToolCall()", + "tc.NormalizeCollections()", "toolResult := *m.ToolResult", ".NormalizeCollections()", }, diff --git a/internal/api/ai_intelligence_handlers.go b/internal/api/ai_intelligence_handlers.go index 84311f122..7eba442a4 100644 --- a/internal/api/ai_intelligence_handlers.go +++ b/internal/api/ai_intelligence_handlers.go @@ -1420,20 +1420,14 @@ func (h *AISettingsHandler) HandleGetRecentIncidents(w http.ResponseWriter, r *h } } - // Get coordinator status - coordinator := h.GetIncidentCoordinatorForOrg(GetOrgID(r.Context())) - var activeCount int - if coordinator != nil { - activeCount = coordinator.GetActiveIncidentCount() - } - // Get incident data from patrol service svc := h.GetAIService(r.Context()) if svc == nil { if err := utils.WriteJSONResponse(w, map[string]interface{}{ - "incidents": []interface{}{}, - "active_count": activeCount, - "message": "Pulse Patrol service not available", + "incidents": []interface{}{}, + "active_count": nil, + "active_count_status": "not_measured", + "message": "Pulse Patrol service not available", }); err != nil { log.Error().Err(err).Msg("Failed to write incidents response") } @@ -1443,9 +1437,10 @@ func (h *AISettingsHandler) HandleGetRecentIncidents(w http.ResponseWriter, r *h patrol := svc.GetPatrolService() if patrol == nil { if err := utils.WriteJSONResponse(w, map[string]interface{}{ - "incidents": []interface{}{}, - "active_count": activeCount, - "message": "Patrol service not available", + "incidents": []interface{}{}, + "active_count": nil, + "active_count_status": "not_measured", + "message": "Patrol service not available", }); err != nil { log.Error().Err(err).Msg("Failed to write incidents response") } @@ -1456,9 +1451,10 @@ func (h *AISettingsHandler) HandleGetRecentIncidents(w http.ResponseWriter, r *h incidentStore := patrol.GetIncidentStore() if incidentStore == nil { if err := utils.WriteJSONResponse(w, map[string]interface{}{ - "incidents": []interface{}{}, - "active_count": activeCount, - "message": "Incident store not available", + "incidents": []interface{}{}, + "active_count": nil, + "active_count_status": "not_measured", + "message": "Incident store not available", }); err != nil { log.Error().Err(err).Msg("Failed to write incidents response") } @@ -1476,9 +1472,10 @@ func (h *AISettingsHandler) HandleGetRecentIncidents(w http.ResponseWriter, r *h // This is a limitation - we may want to add ListRecentIncidents to the store incidentSummary := incidentStore.FormatForPatrol(limit) if err := utils.WriteJSONResponse(w, map[string]interface{}{ - "incidents": []interface{}{}, - "incident_summary": incidentSummary, - "active_count": activeCount, + "incidents": []interface{}{}, + "incident_summary": incidentSummary, + "active_count": nil, + "active_count_status": "not_measured", }); err != nil { log.Error().Err(err).Msg("Failed to write incidents response") } @@ -1486,8 +1483,9 @@ func (h *AISettingsHandler) HandleGetRecentIncidents(w http.ResponseWriter, r *h } if err := utils.WriteJSONResponse(w, map[string]interface{}{ - "incidents": incidents, - "active_count": activeCount, + "incidents": incidents, + "active_count": nil, + "active_count_status": "not_measured", }); err != nil { log.Error().Err(err).Msg("Failed to write incidents response") } diff --git a/internal/api/ai_intelligence_handlers_remediation_additional_test.go b/internal/api/ai_intelligence_handlers_remediation_additional_test.go index e30d01154..243b31d14 100644 --- a/internal/api/ai_intelligence_handlers_remediation_additional_test.go +++ b/internal/api/ai_intelligence_handlers_remediation_additional_test.go @@ -9,7 +9,6 @@ import ( "testing" "time" - "github.com/rcourtman/pulse-go-rewrite/internal/ai" "github.com/rcourtman/pulse-go-rewrite/internal/ai/circuit" "github.com/rcourtman/pulse-go-rewrite/internal/ai/memory" "github.com/rcourtman/pulse-go-rewrite/internal/alerts" @@ -72,11 +71,6 @@ func setupIncidentHandler(t *testing.T) (*AISettingsHandler, *memory.IncidentSto store := memory.NewIncidentStore(memory.IncidentStoreConfig{DataDir: ""}) patrol.SetIncidentStore(store) - coordinator := ai.NewIncidentCoordinator(ai.IncidentCoordinatorConfig{EnableRecorder: false}) - coordinator.SetIncidentStore(store) - coordinator.Start() - handler.SetIncidentCoordinator(coordinator) - alert := &alerts.Alert{ ID: "alert-1", Type: "cpu", @@ -86,7 +80,7 @@ func setupIncidentHandler(t *testing.T) (*AISettingsHandler, *memory.IncidentSto StartTime: time.Now(), LastSeen: time.Now(), } - coordinator.OnAlertFired(alert) + store.RecordAlertFired(alert) return handler, store } @@ -110,8 +104,8 @@ func TestHandleGetRecentIncidents(t *testing.T) { if len(incidents) != 1 { t.Fatalf("expected 1 incident, got %d", len(incidents)) } - if resp["active_count"].(float64) < 1 { - t.Fatalf("expected active_count >= 1") + if resp["active_count"] != nil || resp["active_count_status"] != "not_measured" { + t.Fatalf("saved incident context must not imply a measured live count: %#v", resp) } } @@ -173,3 +167,30 @@ func TestHandleGetIncidentData(t *testing.T) { t.Fatalf("expected formatted_context to be populated") } } + +func TestHandleGetRecentIncidentsCountIsNotMeasured(t *testing.T) { + withMemory, _ := setupIncidentHandler(t) + for _, tc := range []struct { + name string + handler *AISettingsHandler + query string + }{ + {"unavailable", &AISettingsHandler{}, ""}, + {"fleet context", withMemory, ""}, + {"resource context", withMemory, "?resource_id=res-1"}, + {"empty resource context", withMemory, "?resource_id=absent"}, + } { + t.Run(tc.name, func(t *testing.T) { + rec := httptest.NewRecorder() + tc.handler.HandleGetRecentIncidents(rec, httptest.NewRequest(http.MethodGet, "/api/ai/incidents"+tc.query, nil)) + var body map[string]interface{} + if err := json.Unmarshal(rec.Body.Bytes(), &body); err != nil { + t.Fatal(err) + } + count, present := body["active_count"] + if !present || count != nil || body["active_count_status"] != "not_measured" { + t.Fatalf("archive/context presence cannot establish live count: %#v", body) + } + }) + } +} diff --git a/internal/api/chat_service_adapter.go b/internal/api/chat_service_adapter.go index 2e48a8f57..e84b647df 100644 --- a/internal/api/chat_service_adapter.go +++ b/internal/api/chat_service_adapter.go @@ -86,7 +86,7 @@ func adaptChatMessage(m chat.Message) ai.ChatMessage { Timestamp: m.Timestamp, } for _, tc := range m.ToolCalls { - msg.ToolCalls = append(msg.ToolCalls, tc.ProviderToolCall()) + msg.ToolCalls = append(msg.ToolCalls, tc.NormalizeCollections()) } if m.ToolResult != nil { toolResult := *m.ToolResult diff --git a/internal/api/chat_service_adapter_test.go b/internal/api/chat_service_adapter_test.go index 7effa9df5..60ae6ac67 100644 --- a/internal/api/chat_service_adapter_test.go +++ b/internal/api/chat_service_adapter_test.go @@ -3,7 +3,6 @@ package api import ( "context" "encoding/json" - "strings" "testing" "time" @@ -99,8 +98,8 @@ func TestChatServiceAdapter_GetMessages(t *testing.T) { require.NoError(t, err) } -func TestAdaptChatMessageUsesSharedProviderToolCallShape(t *testing.T) { - success := true +func TestAdaptChatMessagePreservesObservedToolResult(t *testing.T) { + success := false msg := adaptChatMessage(chat.Message{ ID: "msg-1", Role: "assistant", @@ -109,7 +108,7 @@ func TestAdaptChatMessageUsesSharedProviderToolCallShape(t *testing.T) { ToolCalls: []chat.ToolCall{{ ID: "call-1", Name: "diagnose", - Output: "in-app only", + Output: "NO_AGENT: command agent unavailable", Success: &success, ThoughtSignature: json.RawMessage(`{"provider":"gemini"}`), }}, @@ -121,7 +120,7 @@ func TestAdaptChatMessageUsesSharedProviderToolCallShape(t *testing.T) { }) require.Len(t, msg.ToolCalls, 1) - var shared agentcapabilities.ProviderToolCall = msg.ToolCalls[0] + var shared agentcapabilities.TranscriptToolCall = msg.ToolCalls[0] assert.Equal(t, "call-1", shared.ID) assert.Equal(t, "diagnose", shared.Name) assert.NotNil(t, shared.Input) @@ -131,8 +130,8 @@ func TestAdaptChatMessageUsesSharedProviderToolCallShape(t *testing.T) { text := string(payload) assert.Contains(t, text, `"input":{}`) assert.Contains(t, text, `"thought_signature":{"provider":"gemini"}`) - assert.False(t, strings.Contains(text, `"output"`), text) - assert.False(t, strings.Contains(text, `"success"`), text) + assert.Contains(t, text, `"output":"NO_AGENT: command agent unavailable"`) + assert.Contains(t, text, `"success":false`) require.NotNil(t, msg.ToolResult) var sharedResult agentcapabilities.ProviderToolResult = *msg.ToolResult diff --git a/internal/api/contract_test.go b/internal/api/contract_test.go index bbeb5fb28..157789311 100644 --- a/internal/api/contract_test.go +++ b/internal/api/contract_test.go @@ -18650,8 +18650,6 @@ func TestContract_AssistantProviderSeamsDoNotUseMCPTerminology(t *testing.T) { } intelligenceAdapters := string(intelligenceAdaptersSource) for _, fragment := range []string{ - `type IncidentRecorderToolAdapter struct`, - `func NewIncidentRecorderToolAdapter(recorder IncidentRecorderSource) *IncidentRecorderToolAdapter`, `type EventCorrelatorToolAdapter struct`, `func NewEventCorrelatorToolAdapter(correlator EventCorrelatorSource) *EventCorrelatorToolAdapter`, } { @@ -19841,7 +19839,7 @@ func TestContract_PulseMCPAdapterProjectsAgentCapabilitiesManifest(t *testing.T) chatTypesSrc := string(chatTypesSource) for _, fragment := range []string{ `func ToolCallFromProvider(tc agentcapabilities.ProviderToolCall) ToolCall`, - `func (t ToolCall) ProviderToolCall() agentcapabilities.ProviderToolCall`, + `type ToolCall = agentcapabilities.TranscriptToolCall`, `type ToolResult = agentcapabilities.ProviderToolResult`, } { if !strings.Contains(chatTypesSrc, fragment) { @@ -19855,8 +19853,8 @@ func TestContract_PulseMCPAdapterProjectsAgentCapabilitiesManifest(t *testing.T) } aiServiceSrc := string(aiServiceSource) for _, fragment := range []string{ - `type ChatToolCall = agentcapabilities.ProviderToolCall`, - `return agentcapabilities.EmptyProviderToolCall()`, + `type ChatToolCall = agentcapabilities.TranscriptToolCall`, + `return ChatToolCall{}.NormalizeCollections()`, `type ChatToolResult = agentcapabilities.ProviderToolResult`, `return agentcapabilities.ApprovalRequiredToolMarker(`, `return agentcapabilities.PolicyBlockedToolMarker(command, reason)`, @@ -20089,11 +20087,11 @@ func TestContract_PulseMCPAdapterProjectsAgentCapabilitiesManifest(t *testing.T) chatServiceAdapterSrc := string(chatServiceAdapterSource) for _, fragment := range []string{ `func adaptChatMessage(m chat.Message) ai.ChatMessage`, - `msg.ToolCalls = append(msg.ToolCalls, tc.ProviderToolCall())`, + `msg.ToolCalls = append(msg.ToolCalls, tc.NormalizeCollections())`, `toolResult := *m.ToolResult`, } { if !strings.Contains(chatServiceAdapterSrc, fragment) { - t.Errorf("chat service adapter must bridge messages through shared provider tool shapes; missing %s", fragment) + t.Errorf("chat service adapter must retain result-bearing transcript calls; missing %s", fragment) } } diff --git a/internal/api/patrol_action_reconciliation.go b/internal/api/patrol_action_reconciliation.go index 4d2c2ef70..5f44c3572 100644 --- a/internal/api/patrol_action_reconciliation.go +++ b/internal/api/patrol_action_reconciliation.go @@ -105,7 +105,7 @@ func (h *AISettingsHandler) applyPatrolActionAudit(orgID string, audit unifiedre store.Update(investigation) } ctx := context.WithValue(context.Background(), OrgIDContextKey, orgID) - h.updateFindingOutcome(ctx, orgID, origin.FindingID, string(outcome)) + h.updateFindingInvestigationOutcome(ctx, orgID, origin.FindingID, string(outcome), investigation) return } if changed { diff --git a/internal/api/router.go b/internal/api/router.go index 6b7157532..6cb59249f 100644 --- a/internal/api/router.go +++ b/internal/api/router.go @@ -2873,41 +2873,8 @@ func (r *Router) initializeAIIntelligenceServices(ctx context.Context, orgID, da log.Info().Msg("AI Intelligence: Event-driven trigger manager initialized and started") } - // 12. Initialize incident coordinator for high-frequency recording - if patrol != nil { - incidentCoordinator := ai.NewIncidentCoordinator(ai.DefaultIncidentCoordinatorConfig()) - - // Wire the incident store if available - if incidentStore := patrol.GetIncidentStore(); incidentStore != nil { - incidentCoordinator.SetIncidentStore(incidentStore) - } - - // Create metrics adapter for incident recorder (ReadState is sole source since SRC-03m) - var metricsAdapter *adapters.MetricsAdapter - if monitor != nil { - metricsAdapter = adapters.NewMetricsAdapter(monitor.GetUnifiedReadState()) - } - - // Initialize and wire the incident recorder (high-frequency metrics) - if metricsAdapter != nil { - recorderCfg := metrics.DefaultIncidentRecorderConfig() - recorderCfg.DataDir = dataDir - recorder := metrics.NewIncidentRecorder(recorderCfg) - recorder.SetMetricsProvider(metricsAdapter) - recorder.Start() - incidentCoordinator.SetRecorder(recorder) - r.aiSettingsHandler.SetIncidentRecorderForOrg(orgID, recorder) - log.Info().Msg("AI Intelligence: Incident recorder initialized and started") - } - - // Start the coordinator - incidentCoordinator.Start() - - // Store reference - r.aiSettingsHandler.SetIncidentCoordinatorForOrg(orgID, incidentCoordinator) - - log.Info().Msg("AI Intelligence: Incident coordinator initialized and started") - } + // Legacy recordings are available only through explicit archive lookup. + r.aiSettingsHandler.SetIncidentArchiveForOrg(orgID, metrics.NewIncidentArchive(dataDir)) log.Info().Msg("AI Intelligence: All Phase 6 & 7 services initialized successfully") } @@ -2965,24 +2932,6 @@ func (r *Router) ShutdownAIIntelligence() { log.Debug().Str("org_id", orgID).Msg("AI Intelligence: Trigger manager stopped") } - // 4. Stop incident coordinators (stop high-frequency recording) - for orgID, incidentCoordinator := range r.aiSettingsHandler.ListIncidentCoordinators() { - if incidentCoordinator == nil { - continue - } - incidentCoordinator.Stop() - log.Debug().Str("org_id", orgID).Msg("AI Intelligence: Incident coordinator stopped") - } - - // 4b. Stop incident recorders (stop background sampling) - for orgID, incidentRecorder := range r.aiSettingsHandler.ListIncidentRecorders() { - if incidentRecorder == nil { - continue - } - incidentRecorder.Stop() - log.Debug().Str("org_id", orgID).Msg("AI Intelligence: Incident recorder stopped") - } - // 5. Cleanup learning stores (removes old records, persists if data dir configured) for orgID, learningStore := range r.aiSettingsHandler.ListLearningStores() { if learningStore == nil { @@ -3353,15 +3302,15 @@ func (r *Router) wireAIChatDependenciesForService(ctx context.Context, service A } // Wire intelligence providers for Assistant tools. - // - IncidentRecorderProvider: high-frequency incident data (pulse_get_incident_window) + // - IncidentArchiveProvider: explicit reads of saved legacy recordings // - EventCorrelatorProvider: Proxmox events (pulse_correlate_events) // - KnowledgeStoreProvider: notes (pulse_remember, pulse_recall) - // Wire incident recorder provider (high-frequency incident data) + // Wire the org-pinned archive reader without creating or sampling data. if r.aiSettingsHandler != nil { - if recorder := r.aiSettingsHandler.GetIncidentRecorderForOrg(orgID); recorder != nil { - service.SetIncidentRecorderProvider(&incidentRecorderProviderWrapper{recorder: recorder}) - log.Debug().Msg("AI chat: Incident recorder provider wired") + if archive := r.aiSettingsHandler.GetIncidentArchiveForOrg(orgID); archive != nil { + service.SetIncidentArchiveProvider(archive) + log.Debug().Msg("AI chat: Incident archive provider wired") } } @@ -3476,82 +3425,6 @@ func (w *forecastResourceIterator) ForecastStoragePools() []forecast.ResourceInf return result } -// incidentRecorderProviderWrapper adapts metrics.IncidentRecorder to tools.IncidentRecorderProvider. -type incidentRecorderProviderWrapper struct { - recorder *metrics.IncidentRecorder -} - -func (w *incidentRecorderProviderWrapper) GetWindowsForResource(resourceID string, limit int) []*tools.IncidentWindow { - if w.recorder == nil { - return nil - } - - windows := w.recorder.GetWindowsForResource(resourceID, limit) - if len(windows) == 0 { - return nil - } - - result := make([]*tools.IncidentWindow, 0, len(windows)) - for _, window := range windows { - if window == nil { - continue - } - result = append(result, convertIncidentWindow(window)) - } - return result -} - -func (w *incidentRecorderProviderWrapper) GetWindow(windowID string) *tools.IncidentWindow { - if w.recorder == nil { - return nil - } - window := w.recorder.GetWindow(windowID) - if window == nil { - return nil - } - return convertIncidentWindow(window) -} - -func convertIncidentWindow(window *metrics.IncidentWindow) *tools.IncidentWindow { - if window == nil { - return nil - } - - points := make([]tools.IncidentDataPoint, 0, len(window.DataPoints)) - for _, point := range window.DataPoints { - points = append(points, tools.IncidentDataPoint{ - Timestamp: point.Timestamp, - Metrics: point.Metrics, - }) - } - - var summary *tools.IncidentSummary - if window.Summary != nil { - summary = &tools.IncidentSummary{ - Duration: window.Summary.Duration, - DataPoints: window.Summary.DataPoints, - Peaks: window.Summary.Peaks, - Lows: window.Summary.Lows, - Averages: window.Summary.Averages, - Changes: window.Summary.Changes, - } - } - - return &tools.IncidentWindow{ - ID: window.ID, - ResourceID: window.ResourceID, - ResourceName: window.ResourceName, - ResourceType: window.ResourceType, - TriggerType: window.TriggerType, - TriggerID: window.TriggerID, - StartTime: window.StartTime, - EndTime: window.EndTime, - Status: string(window.Status), - DataPoints: points, - Summary: summary, - } -} - func (r *Router) publishActionCompletedAgentEvent(broadcaster *AgentEventBroadcaster, record unifiedresources.ActionAuditRecord) { if broadcaster == nil { return diff --git a/internal/api/router_state_test.go b/internal/api/router_state_test.go index 2ff24521f..db89c0727 100644 --- a/internal/api/router_state_test.go +++ b/internal/api/router_state_test.go @@ -175,7 +175,7 @@ func TestRouterHandleStatePreservesNumericIdleRatesAndOmitsUnknownRates(t *testi byName[resource.Name] = resource } idle := byName["idle-vm"] - if idle.DiskIO == nil || idle.DiskIO.ReadRate != 0 || idle.DiskIO.WriteRate != 0 { + if idle.DiskIO == nil || idle.DiskIO.ReadRate == nil || idle.DiskIO.WriteRate == nil || *idle.DiskIO.ReadRate != 0 || *idle.DiskIO.WriteRate != 0 { t.Fatalf("valid idle rates were not emitted as numeric zero: %+v", idle.DiskIO) } if unknown := byName["unknown-vm"]; unknown.DiskIO != nil { diff --git a/internal/api/router_wrappers_additional_test.go b/internal/api/router_wrappers_additional_test.go index d52422847..73d9488a8 100644 --- a/internal/api/router_wrappers_additional_test.go +++ b/internal/api/router_wrappers_additional_test.go @@ -8,7 +8,6 @@ import ( "github.com/rcourtman/pulse-go-rewrite/internal/ai/patterns" "github.com/rcourtman/pulse-go-rewrite/internal/ai/proxmox" "github.com/rcourtman/pulse-go-rewrite/internal/config" - "github.com/rcourtman/pulse-go-rewrite/internal/metrics" "github.com/rcourtman/pulse-go-rewrite/internal/models" "github.com/rcourtman/pulse-go-rewrite/internal/monitoring" unifiedresources "github.com/rcourtman/pulse-go-rewrite/internal/unifiedresources" @@ -52,84 +51,6 @@ func TestForecastResourceIterator_NilReadState(t *testing.T) { } } -func TestIncidentRecorderProviderWrapper(t *testing.T) { - now := time.Now().UTC() - end := now.Add(5 * time.Minute) - - activeWindow := &metrics.IncidentWindow{ - ID: "win-active", - ResourceID: "res-1", - ResourceName: "Resource", - ResourceType: "vm", - TriggerType: "alert", - TriggerID: "alert-1", - StartTime: now, - EndTime: &end, - Status: metrics.IncidentWindowStatusRecording, - DataPoints: []metrics.IncidentDataPoint{ - {Timestamp: now, Metrics: map[string]float64{"cpu": 10}}, - }, - Summary: &metrics.IncidentSummary{ - Duration: 5 * time.Minute, - DataPoints: 1, - Peaks: map[string]float64{"cpu": 10}, - Lows: map[string]float64{"cpu": 5}, - Averages: map[string]float64{"cpu": 7}, - Changes: map[string]float64{"cpu": 2}, - }, - } - - completedWindow := &metrics.IncidentWindow{ - ID: "win-complete", - ResourceID: "res-1", - ResourceName: "Resource", - ResourceType: "vm", - TriggerType: "alert", - TriggerID: "alert-2", - StartTime: now.Add(-time.Hour), - Status: metrics.IncidentWindowStatusComplete, - DataPoints: []metrics.IncidentDataPoint{ - {Timestamp: now.Add(-time.Hour), Metrics: map[string]float64{"cpu": 20}}, - }, - } - - recorder := metrics.NewIncidentRecorder(metrics.DefaultIncidentRecorderConfig()) - setUnexportedField(t, recorder, "activeWindows", map[string]*metrics.IncidentWindow{"win-active": activeWindow}) - setUnexportedField(t, recorder, "completedWindows", []*metrics.IncidentWindow{completedWindow}) - - wrapper := &incidentRecorderProviderWrapper{recorder: recorder} - windows := wrapper.GetWindowsForResource("res-1", 10) - if len(windows) != 2 { - t.Fatalf("expected 2 windows, got %d", len(windows)) - } - - ids := []string{windows[0].ID, windows[1].ID} - if !containsStringSlice(ids, "win-active") || !containsStringSlice(ids, "win-complete") { - t.Fatalf("unexpected window ids %v", ids) - } - - window := wrapper.GetWindow("win-active") - if window == nil || window.ResourceID != "res-1" || window.Status == "" { - t.Fatalf("unexpected window: %#v", window) - } -} - -func TestIncidentRecorderProviderWrapper_NilRecorder(t *testing.T) { - wrapper := &incidentRecorderProviderWrapper{} - if got := wrapper.GetWindowsForResource("res-1", 5); got != nil { - t.Fatalf("expected nil windows, got %#v", got) - } - if got := wrapper.GetWindow("win-1"); got != nil { - t.Fatalf("expected nil window, got %#v", got) - } -} - -func TestConvertIncidentWindowNil(t *testing.T) { - if got := convertIncidentWindow(nil); got != nil { - t.Fatalf("expected nil window, got %#v", got) - } -} - func TestEventCorrelatorProviderWrapper(t *testing.T) { now := time.Now().UTC() corr := proxmox.EventCorrelation{ diff --git a/internal/dockeragent/blockio_presence_test.go b/internal/dockeragent/blockio_presence_test.go new file mode 100644 index 000000000..da6de610e --- /dev/null +++ b/internal/dockeragent/blockio_presence_test.go @@ -0,0 +1,48 @@ +package dockeragent + +import ( + "encoding/json" + "testing" + + containertypes "github.com/moby/moby/api/types/container" + agentsdocker "github.com/rcourtman/pulse-go-rewrite/pkg/agents/docker" +) + +func TestSummarizeBlockIOPreservesDirectionPresenceAndZero(t *testing.T) { + for _, tc := range []struct { + name string + entries []containertypes.BlkioStatEntry + read, write bool + }{ + {"absent", nil, false, false}, + {"unrelated", []containertypes.BlkioStatEntry{{Op: "Total", Value: 100}}, false, false}, + {"observed idle", []containertypes.BlkioStatEntry{{Op: "Read", Value: 0}, {Op: "Write", Value: 0}}, true, true}, + {"read only idle", []containertypes.BlkioStatEntry{{Op: "Read", Value: 0}}, true, false}, + {"write only", []containertypes.BlkioStatEntry{{Op: "Write", Value: 123}}, false, true}, + } { + t.Run(tc.name, func(t *testing.T) { + got := summarizeBlockIO(containertypes.StatsResponse{BlkioStats: containertypes.BlkioStats{IoServiceBytesRecursive: tc.entries}}) + if !tc.read && !tc.write { + if got != nil { + t.Fatalf("absent counters became observations: %+v", got) + } + return + } + if got == nil { + t.Fatal("explicit counters were dropped") + } + encoded, err := json.Marshal(got) + if err != nil { + t.Fatal(err) + } + var decoded agentsdocker.ContainerBlockIO + if err := json.Unmarshal(encoded, &decoded); err != nil { + t.Fatal(err) + } + read, write := decoded.CounterPresence() + if read != tc.read || write != tc.write { + t.Fatalf("presence lost through report JSON %s: %v/%v", encoded, read, write) + } + }) + } +} diff --git a/internal/dockeragent/collect.go b/internal/dockeragent/collect.go index ef318a2c9..02223dc3a 100644 --- a/internal/dockeragent/collect.go +++ b/internal/dockeragent/collect.go @@ -10,6 +10,7 @@ import ( "net/netip" "net/url" "regexp" + "sort" "strconv" "strings" "time" @@ -698,6 +699,36 @@ func (a *Agent) collectContainer(ctx context.Context, summary containertypes.Sum }) } } + // Docker's --tmpfs mounts can exist only in HostConfig.Tmpfs. Preserve + // their configuration alongside inspected mounts without inventing usage. + if inspect.HostConfig != nil && len(inspect.HostConfig.Tmpfs) > 0 { + reported := make(map[string]bool, len(mounts)) + for _, mount := range mounts { + reported[mount.Destination] = true + } + destinations := make([]string, 0, len(inspect.HostConfig.Tmpfs)) + for destination := range inspect.HostConfig.Tmpfs { + if !reported[destination] { + destinations = append(destinations, destination) + } + } + sort.Strings(destinations) + for _, destination := range destinations { + options := inspect.HostConfig.Tmpfs[destination] + writable := true + for _, option := range strings.Split(options, ",") { + switch strings.TrimSpace(option) { + case "ro": + writable = false + case "rw": + writable = true + } + } + mounts = append(mounts, agentsdocker.ContainerMount{ + Type: "tmpfs", Destination: destination, Mode: options, RW: writable, + }) + } + } oomKilled := inspect.State.OOMKilled container := agentsdocker.Container{ @@ -1318,35 +1349,34 @@ func randomDuration(max time.Duration) time.Duration { } func summarizeBlockIO(stats containertypes.StatsResponse) *agentsdocker.ContainerBlockIO { - // BlkioStats structure varies by cgroup version - // Cgroup v1: IoServiceBytesRecursive []BlkioStatEntry - // Cgroup v2: IoServiceBytesRecursive is empty? No, Docker maps it? - // Docker API guarantees IoServiceBytesRecursive is populated? - // It seems to try to handle both. - if len(stats.BlkioStats.IoServiceBytesRecursive) == 0 { return nil } var readBytes, writeBytes uint64 + var readPresent, writePresent bool for _, entry := range stats.BlkioStats.IoServiceBytesRecursive { op := strings.ToLower(entry.Op) switch op { case "read": + readPresent = true readBytes += entry.Value case "write": + writePresent = true writeBytes += entry.Value } } - if readBytes == 0 && writeBytes == 0 { + if !readPresent && !writePresent { return nil } return &agentsdocker.ContainerBlockIO{ - ReadBytes: readBytes, - WriteBytes: writeBytes, + ReadBytes: readBytes, + WriteBytes: writeBytes, + ReadBytesPresent: &readPresent, + WriteBytesPresent: &writePresent, } } diff --git a/internal/dockeragent/collect_tmpfs_live_test.go b/internal/dockeragent/collect_tmpfs_live_test.go new file mode 100644 index 000000000..ca0bb459b --- /dev/null +++ b/internal/dockeragent/collect_tmpfs_live_test.go @@ -0,0 +1,148 @@ +package dockeragent + +import ( + "context" + "encoding/json" + "os" + "os/exec" + "path/filepath" + "strings" + "testing" + "time" + + "github.com/moby/moby/client" + "github.com/rcourtman/pulse-go-rewrite/internal/ai/qualification" + agentsdocker "github.com/rcourtman/pulse-go-rewrite/pkg/agents/docker" + "github.com/rs/zerolog" +) + +// This opt-in proof calls the real collector against the existing bounded lab. +// It does not enroll an agent, send reports to Pulse, or invoke a model. +func TestCollectContainerStorageFaultLive(t *testing.T) { + dockerContext := os.Getenv("PULSE_QUALIFY_ORACLE_DOCKER_CONTEXT") + if dockerContext == "" { + t.Skip("set PULSE_QUALIFY_ORACLE_DOCKER_CONTEXT to an explicit disposable Docker context") + } + ctx, cancel := context.WithTimeout(context.Background(), 3*time.Minute) + defer cancel() + endpoint, err := exec.CommandContext(ctx, "docker", "context", "inspect", dockerContext, + "--format", "{{.Endpoints.docker.Host}}").Output() + if err != nil { + t.Fatal(err) + } + host := strings.TrimSpace(string(endpoint)) + if !strings.HasPrefix(host, "unix://") { + t.Fatal("run this proof beside the disposable Docker daemon using a Unix socket context") + } + moduleClient, err := newMobyDockerClient(client.WithHost(host), client.WithAPIVersionNegotiation()) + if err != nil { + t.Fatal(err) + } + t.Cleanup(func() { _ = moduleClient.Close() }) + agent := &Agent{docker: moduleClient, runtime: RuntimeDocker, logger: zerolog.Nop(), + cfg: Config{CollectDiskMetrics: true}, prevContainerCPU: make(map[string]cpuSample)} + manifest, err := qualification.LoadManifest(filepath.Join("..", "..", "tests", "qualification", "patrol", + "scenarios", "investigation.docker-storage-pressure.json")) + if err != nil { + t.Fatal(err) + } + driver := qualification.NewDockerLab(nil, qualification.DockerTarget{Context: dockerContext}) + lab, prepareErr := driver.Prepare(ctx, manifest, "collector-"+time.Now().UTC().Format("20060102t150405.000000000")) + if lab != nil { + t.Cleanup(func() { + cleanupCtx, cleanupCancel := context.WithTimeout(context.Background(), time.Minute) + defer cleanupCancel() + result := driver.Cleanup(cleanupCtx, manifest, lab) + encoded, _ := json.Marshal(result) + t.Logf("cleanup=%s", encoded) + if !result.Passed || !result.SecondCleanupNoop || !result.InventoryUnchanged { + t.Errorf("disposable inventory cleanup failed: %+v", result) + } + }) + } + if prepareErr != nil { + t.Fatal(prepareErr) + } + + check := func(phase, serviceHealth string, predicates []qualification.Predicate) { + t.Helper() + observations, err := driver.Observe(ctx, manifest, lab, predicates) + encoded, _ := json.Marshal(observations) + t.Logf("%s oracle=%s", phase, encoded) + if err != nil { + t.Fatal(err) + } + for _, alias := range []string{"service", "control"} { + id := lab.ResourceIDs[alias] + filters := newDockerFilters() + filters.Add("id", id) + for _, key := range []string{"io.pulse.owner", "io.pulse.profile", "io.pulse.component"} { + value := lab.BaselineStates[alias].Labels[key] + if value == "" { + t.Fatalf("missing fixture ownership label %s", key) + } + filters.Add("label", key+"="+value) + } + containers, err := moduleClient.ContainerList(ctx, dockerContainerListOptions{All: true, Filters: filters}) + if err != nil || len(containers) != 1 || containers[0].ID != id { + t.Fatalf("exact owned fixture %s unavailable: count=%d error=%v", alias, len(containers), err) + } + collected, err := agent.collectContainer(ctx, containers[0]) + if err != nil { + t.Fatal(err) + } + encoded, err := json.Marshal(collected) + if err != nil { + t.Fatal(err) + } + var report agentsdocker.Container + if err := json.Unmarshal(encoded, &report); err != nil { + t.Fatal(err) + } + wantHealth := "healthy" + if alias == "service" { + wantHealth = serviceHealth + } + if report.ID != id || report.State != "running" || report.Health != wantHealth { + t.Fatalf("%s %s report identity/state/health mismatch: %s", phase, alias, encoded) + } + if alias == "service" { + const destination = "/var/lib/service-cache" + inspect, err := moduleClient.ContainerInspect(ctx, id) + if err != nil || inspect.HostConfig == nil { + t.Fatalf("inspect storage configuration: %v", err) + } + options, ok := inspect.HostConfig.Tmpfs[destination] + if !ok || !strings.Contains(options, "size=8388608") { + t.Fatalf("fixture tmpfs configuration missing: %q", options) + } + matches := 0 + for _, mount := range report.Mounts { + if mount.Destination == destination { + matches++ + if mount.Type != "tmpfs" || !mount.RW || mount.Mode != options { + t.Fatalf("collector changed tmpfs configuration: %+v", mount) + } + } + } + if matches != 1 { + t.Fatalf("expected one storage mount in report, got %d", matches) + } + t.Logf("%s raw_mount_count=%d native_tmpfs_options=%q", phase, len(inspect.Mounts), options) + } + t.Logf("%s %s report=%s", phase, alias, encoded) + } + } + check("baseline", "healthy", manifest.Baseline) + for _, fault := range manifest.Faults { + if err := driver.ApplyFault(ctx, manifest, lab, fault); err != nil { + t.Fatal(err) + } + check("storage-full", "unhealthy", fault.Oracle) + if err := driver.RevertFault(ctx, manifest, lab, fault); err != nil { + t.Fatal(err) + } + check("recovered", "healthy", fault.RevertOracle) + } + check("restored-baseline", "healthy", manifest.Baseline) +} diff --git a/internal/dockeragent/collect_tmpfs_test.go b/internal/dockeragent/collect_tmpfs_test.go new file mode 100644 index 000000000..c9b6340ad --- /dev/null +++ b/internal/dockeragent/collect_tmpfs_test.go @@ -0,0 +1,84 @@ +package dockeragent + +import ( + "context" + "encoding/json" + "reflect" + "testing" + + containertypes "github.com/moby/moby/api/types/container" + agentsdocker "github.com/rcourtman/pulse-go-rewrite/pkg/agents/docker" + "github.com/rs/zerolog" +) + +func TestCollectContainerPreservesTmpfsMounts(t *testing.T) { + for _, tc := range []struct { + name string + host *containertypes.HostConfig + mounts []containertypes.MountPoint + want []agentsdocker.ContainerMount + }{ + { + name: "native tmpfs only", + host: &containertypes.HostConfig{Tmpfs: map[string]string{ + "/var/lib/service-cache": "rw,noexec,nosuid,nodev,size=8388608", + }}, + want: []agentsdocker.ContainerMount{{Type: "tmpfs", Destination: "/var/lib/service-cache", Mode: "rw,noexec,nosuid,nodev,size=8388608", RW: true}}, + }, + { + name: "mixed mounts and stable tmpfs order", + host: &containertypes.HostConfig{Tmpfs: map[string]string{"/z-cache": "", "/a-cache": "ro,noexec"}}, + mounts: []containertypes.MountPoint{{Type: "bind", Source: "/host-data", Destination: "/data", RW: true}}, + want: []agentsdocker.ContainerMount{ + {Type: "bind", Source: "/host-data", Destination: "/data", RW: true}, + {Type: "tmpfs", Destination: "/a-cache", Mode: "ro,noexec", RW: false}, + {Type: "tmpfs", Destination: "/z-cache", RW: true}, + }, + }, + { + name: "reported mount is authoritative", + host: &containertypes.HostConfig{Tmpfs: map[string]string{"/cache": "rw,size=8388608"}}, + mounts: []containertypes.MountPoint{{Type: "tmpfs", Destination: "/cache", Mode: "ro", RW: false}}, + want: []agentsdocker.ContainerMount{{Type: "tmpfs", Destination: "/cache", Mode: "ro", RW: false}}, + }, + {name: "absent host config"}, + } { + t.Run(tc.name, func(t *testing.T) { + inspect := baseInspect() + inspect.HostConfig = tc.host + inspect.Mounts = tc.mounts + a := &Agent{ + logger: zerolog.Nop(), + runtime: RuntimeDocker, + prevContainerCPU: make(map[string]cpuSample), + docker: &fakeDockerClient{ + containerInspectWithRawFn: func(context.Context, string, bool) (containertypes.InspectResponse, []byte, error) { + return inspect, nil, nil + }, + containerStatsOneShotFn: func(context.Context, string) (dockerStatsResponseReader, error) { + return statsReader(t, containertypes.StatsResponse{}), nil + }, + }, + } + got, err := a.collectContainer(context.Background(), containertypes.Summary{ID: "owned-storage-container", Names: []string{"/worker"}, Image: "alpine:3.20", State: "running"}) + if err != nil { + t.Fatal(err) + } + if !reflect.DeepEqual(got.Mounts, tc.want) { + t.Fatalf("mount inventory = %#v, want %#v", got.Mounts, tc.want) + } + encoded, err := json.Marshal(got) + if err != nil { + t.Fatal(err) + } + var wire agentsdocker.Container + if err := json.Unmarshal(encoded, &wire); err != nil { + t.Fatal(err) + } + if !reflect.DeepEqual(wire.Mounts, tc.want) { + t.Fatalf("report lost mounts: %#v", wire.Mounts) + } + t.Logf("TMPFS_COLLECTOR_REPORT %s", encoded) + }) + } +} diff --git a/internal/metrics/incident_archive.go b/internal/metrics/incident_archive.go new file mode 100644 index 000000000..5d8e0d243 --- /dev/null +++ b/internal/metrics/incident_archive.go @@ -0,0 +1,170 @@ +// Package metrics preserves access to legacy incident recording archives. +package metrics + +import ( + "encoding/json" + "errors" + "fmt" + "io" + "os" + "path/filepath" + "strings" + "time" +) + +// IncidentWindow represents a saved legacy recording window. Status and sample timestamps are historical +type IncidentWindow struct { + ID string `json:"id"` + ResourceID string `json:"resource_id"` + ResourceName string `json:"resource_name,omitempty"` + ResourceType string `json:"resource_type,omitempty"` + TriggerType string `json:"trigger_type"` // "alert", "anomaly", "focus", "manual" + TriggerID string `json:"trigger_id,omitempty"` + StartTime time.Time `json:"start_time"` + EndTime *time.Time `json:"end_time,omitempty"` + Status IncidentWindowStatus `json:"status"` + DataPoints []IncidentDataPoint `json:"data_points"` + Summary *IncidentSummary `json:"summary,omitempty"` +} + +// IncidentWindowStatus represents the status of an incident window +type IncidentWindowStatus string + +const ( + IncidentWindowStatusRecording IncidentWindowStatus = "recording" + IncidentWindowStatusComplete IncidentWindowStatus = "complete" + IncidentWindowStatusTruncated IncidentWindowStatus = "truncated" // Stopped due to limits + + maxIncidentWindowsFileSize = 16 << 20 // 16 MiB +) + +var errUnsafeIncidentArchivePath = errors.New("unsafe incident archive path") + +// IncidentDataPoint represents a single data point in an incident window +type IncidentDataPoint struct { + Timestamp time.Time `json:"timestamp"` + Metrics map[string]float64 `json:"metrics"` // cpu, memory, disk, etc. + Metadata map[string]interface{} `json:"metadata,omitempty"` +} + +// IncidentSummary provides computed statistics about an incident window +type IncidentSummary struct { + Duration time.Duration `json:"duration_ms"` + DataPoints int `json:"data_points"` + Peaks map[string]float64 `json:"peaks"` // Maximum values + Lows map[string]float64 `json:"lows"` // Minimum values + Averages map[string]float64 `json:"averages"` // Average values + Changes map[string]float64 `json:"changes"` // Change from start to end + Anomalies []string `json:"anomalies,omitempty"` // Detected anomalies +} + +// IncidentArchive reads saved recordings on explicit request. It never starts +// collectors or rewrites, expires, creates or changes permissions on archives. +type IncidentArchive struct{ filePath string } + +var ErrIncidentArchiveUnavailable = errors.New("legacy incident recording archive is unavailable") + +func NewIncidentArchive(dataDir string) *IncidentArchive { + if strings.TrimSpace(dataDir) == "" { + return &IncidentArchive{} + } + return &IncidentArchive{filePath: filepath.Join(dataDir, "incident_windows.json")} +} + +// GetWindow requires the exact resource and window identifiers in this org's +// archive. Historical names and aliases cannot authorize an archive lookup. +func (a *IncidentArchive) GetWindow(resourceID, windowID string) (*IncidentWindow, error) { + if a == nil || a.filePath == "" { + return nil, ErrIncidentArchiveUnavailable + } + if strings.TrimSpace(resourceID) == "" || strings.TrimSpace(windowID) == "" { + return nil, errors.New("resource_id and window_id are required") + } + data, err := readBoundedRegularFile(a.filePath, maxIncidentWindowsFileSize) + if errors.Is(err, os.ErrNotExist) { + return nil, ErrIncidentArchiveUnavailable + } + if err != nil { + return nil, fmt.Errorf("read legacy incident archive: %w", err) + } + var saved struct { + CompletedWindows json.RawMessage `json:"completed_windows"` + } + if err := json.Unmarshal(data, &saved); err != nil { + return nil, fmt.Errorf("decode legacy incident archive: %w", err) + } + if len(saved.CompletedWindows) == 0 { + return nil, errors.New("legacy incident archive has no completed_windows field") + } + var windows []*IncidentWindow + if err := json.Unmarshal(saved.CompletedWindows, &windows); err != nil { + return nil, fmt.Errorf("decode legacy incident windows: %w", err) + } + var match *IncidentWindow + for _, window := range windows { + if window == nil || window.ID != windowID || window.ResourceID != resourceID { + continue + } + if match != nil { + return nil, errors.New("legacy incident archive contains duplicate resource/window identifiers") + } + match = window + } + return match, nil +} + +func validateRegularFilePath(path string, info os.FileInfo) error { + if info.Mode()&os.ModeSymlink != 0 { + return fmt.Errorf("%w: refusing symlink path %q", errUnsafeIncidentArchivePath, path) + } + if !info.Mode().IsRegular() { + return fmt.Errorf("%w: non-regular path %q", errUnsafeIncidentArchivePath, path) + } + return nil +} + +func readBoundedRegularFile(path string, maxSize int64) ([]byte, error) { + initialInfo, err := os.Lstat(path) + if err != nil { + return nil, err + } + if err := validateRegularFilePath(path, initialInfo); err != nil { + return nil, err + } + if maxSize > 0 && initialInfo.Size() > maxSize { + return nil, fmt.Errorf("%w: file %q exceeds size limit (%d bytes)", errUnsafeIncidentArchivePath, path, initialInfo.Size()) + } + + file, err := os.Open(path) + if err != nil { + return nil, err + } + defer func() { + _ = file.Close() + }() + + openInfo, err := file.Stat() + if err != nil { + return nil, err + } + if err := validateRegularFilePath(path, openInfo); err != nil { + return nil, err + } + if !os.SameFile(initialInfo, openInfo) { + return nil, fmt.Errorf("%w: file %q changed during read", errUnsafeIncidentArchivePath, path) + } + + reader := io.Reader(file) + if maxSize > 0 { + reader = io.LimitReader(file, maxSize+1) + } + + data, err := io.ReadAll(reader) + if err != nil { + return nil, err + } + if maxSize > 0 && int64(len(data)) > maxSize { + return nil, fmt.Errorf("%w: file %q exceeded size limit while reading", errUnsafeIncidentArchivePath, path) + } + return data, nil +} diff --git a/internal/metrics/incident_archive_test.go b/internal/metrics/incident_archive_test.go new file mode 100644 index 000000000..c162eb787 --- /dev/null +++ b/internal/metrics/incident_archive_test.go @@ -0,0 +1,117 @@ +package metrics + +import ( + "encoding/json" + "os" + "path/filepath" + "testing" + "time" + + "github.com/stretchr/testify/require" +) + +func TestIncidentArchivePreservesHistoricalFileAndValues(t *testing.T) { + dir := t.TempDir() + path := filepath.Join(dir, "incident_windows.json") + // Older than the former retention window, with data that the old adapter dropped. + original := []byte(`{"completed_windows":[null,{"id":"old-window","resource_id":"docker:old-host/full-id","resource_type":"host","status":"recording","start_time":"2020-01-02T03:04:05Z","data_points":[{"timestamp":"2020-01-02T03:04:06Z","metrics":{"cpu":12.5},"metadata":{"source":"cached","nested":{"retained":true}}}],"summary":{"duration_ms":60000000000,"data_points":1,"anomalies":["stored observation"]}}]}`) + require.NoError(t, os.WriteFile(path, original, 0640)) + require.NoError(t, os.Chmod(path, 0640)) + oldTime := time.Date(2020, 1, 2, 3, 4, 5, 0, time.UTC) + require.NoError(t, os.Chtimes(path, oldTime, oldTime)) + before, err := os.Stat(path) + require.NoError(t, err) + archive := NewIncidentArchive(dir) + window, err := archive.GetWindow("docker:old-host/full-id", "old-window") + require.NoError(t, err) + require.NotNil(t, window) + require.Equal(t, IncidentWindowStatusRecording, window.Status) + require.Equal(t, "host", window.ResourceType) + require.Equal(t, oldTime, window.StartTime) + require.Equal(t, time.Minute, window.Summary.Duration) + require.Equal(t, []string{"stored observation"}, window.Summary.Anomalies) + require.Equal(t, map[string]interface{}{"retained": true}, window.DataPoints[0].Metadata["nested"]) + // Each read decodes independently, so a caller cannot alter later evidence. + window.DataPoints[0].Metrics["cpu"] = 99 + again, err := archive.GetWindow("docker:old-host/full-id", "old-window") + require.NoError(t, err) + require.Equal(t, 12.5, again.DataPoints[0].Metrics["cpu"]) + for _, key := range [][2]string{{"other-resource", "old-window"}, {"docker:old-host/full-id", "absent"}} { + got, err := archive.GetWindow(key[0], key[1]) + require.NoError(t, err) + require.Nil(t, got) + } + after, err := os.Stat(path) + require.NoError(t, err) + got, err := os.ReadFile(path) + require.NoError(t, err) + require.Equal(t, original, got) + require.Equal(t, before.Mode(), after.Mode()) + require.Equal(t, before.ModTime(), after.ModTime()) +} + +func TestIncidentArchiveReadsOnlyOnRequestAndReportsFailures(t *testing.T) { + dir := t.TempDir() + path := filepath.Join(dir, "incident_windows.json") + archive := NewIncidentArchive(dir) + _, err := os.Stat(path) + require.True(t, os.IsNotExist(err)) + _, err = archive.GetWindow("r", "w") + require.ErrorIs(t, err, ErrIncidentArchiveUnavailable) + // A repaired or newly restored file is visible on the next explicit read. + for _, raw := range []string{`{"completed_windows":[]}`, `{"completed_windows":null}`} { + require.NoError(t, os.WriteFile(path, []byte(raw), 0600)) + got, err := archive.GetWindow("r", "w") + require.NoError(t, err) + require.Nil(t, got) + } + for _, raw := range []string{`{`, `{}`, `null`, `{"completed_windows":{}}`, `{"completed_windows":[{"id":"w","resource_id":"r"},{"id":"w","resource_id":"r"}]}`} { + require.NoError(t, os.WriteFile(path, []byte(raw), 0600)) + got, err := archive.GetWindow("r", "w") + require.Error(t, err) + require.Nil(t, got) + } + require.NoError(t, os.Remove(path)) + require.NoError(t, os.Mkdir(path, 0700)) + _, err = archive.GetWindow("r", "w") + require.ErrorIs(t, err, errUnsafeIncidentArchivePath) + require.NoError(t, os.Remove(path)) + target := filepath.Join(dir, "other.json") + require.NoError(t, os.WriteFile(target, []byte(`{"completed_windows":[]}`), 0600)) + require.NoError(t, os.Symlink(target, path)) + _, err = archive.GetWindow("r", "w") + require.ErrorIs(t, err, errUnsafeIncidentArchivePath) + require.NoError(t, os.Remove(path)) + f, err := os.Create(path) + require.NoError(t, err) + require.NoError(t, f.Truncate(maxIncidentWindowsFileSize+1)) + require.NoError(t, f.Close()) + _, err = archive.GetWindow("r", "w") + require.ErrorIs(t, err, errUnsafeIncidentArchivePath) + for _, a := range []*IncidentArchive{nil, NewIncidentArchive("")} { + _, err = a.GetWindow("r", "w") + require.ErrorIs(t, err, ErrIncidentArchiveUnavailable) + } +} + +func TestIncidentArchiveRequiresExactResourceAndOrg(t *testing.T) { + aDir, bDir := t.TempDir(), t.TempDir() + for _, org := range []struct{ dir, name string }{{aDir, "tenant-a"}, {bDir, "tenant-b"}} { + raw, err := json.Marshal(map[string]interface{}{"completed_windows": []*IncidentWindow{{ID: "same-window", ResourceID: "same-resource", ResourceName: org.name}}}) + require.NoError(t, err) + require.NoError(t, os.WriteFile(filepath.Join(org.dir, "incident_windows.json"), raw, 0600)) + } + for _, org := range []struct{ dir, name string }{{aDir, "tenant-a"}, {bDir, "tenant-b"}} { + a := NewIncidentArchive(org.dir) + got, err := a.GetWindow("same-resource", "same-window") + require.NoError(t, err) + require.Equal(t, org.name, got.ResourceName) + for _, keys := range [][2]string{{"same-resource ", "same-window"}, {"same-resource", "same-window "}} { + got, err = a.GetWindow(keys[0], keys[1]) + require.NoError(t, err) + require.Nil(t, got) + } + _, err = a.GetWindow("", "same-window") + require.Error(t, err) + } +} diff --git a/internal/metrics/incident_recorder.go b/internal/metrics/incident_recorder.go deleted file mode 100644 index 24dc091e9..000000000 --- a/internal/metrics/incident_recorder.go +++ /dev/null @@ -1,948 +0,0 @@ -// Package metrics provides metrics collection and incident recording functionality. -package metrics - -import ( - "encoding/json" - "errors" - "fmt" - "io" - "os" - "path/filepath" - "strings" - "sync" - "sync/atomic" - "time" - - "github.com/rcourtman/pulse-go-rewrite/internal/unifiedresources" - "github.com/rs/zerolog/log" -) - -// IncidentWindow represents a high-frequency recording window during an incident -type IncidentWindow struct { - ID string `json:"id"` - ResourceID string `json:"resource_id"` - ResourceName string `json:"resource_name,omitempty"` - ResourceType string `json:"resource_type,omitempty"` - TriggerType string `json:"trigger_type"` // "alert", "anomaly", "focus", "manual" - TriggerID string `json:"trigger_id,omitempty"` - StartTime time.Time `json:"start_time"` - EndTime *time.Time `json:"end_time,omitempty"` - Status IncidentWindowStatus `json:"status"` - DataPoints []IncidentDataPoint `json:"data_points"` - Summary *IncidentSummary `json:"summary,omitempty"` -} - -// IncidentWindowStatus represents the status of an incident window -type IncidentWindowStatus string - -const ( - IncidentWindowStatusRecording IncidentWindowStatus = "recording" - IncidentWindowStatusComplete IncidentWindowStatus = "complete" - IncidentWindowStatusTruncated IncidentWindowStatus = "truncated" // Stopped due to limits - - incidentRecorderDirPerm = 0o700 - incidentRecorderFilePerm = 0o600 - maxIncidentWindowsFileSize = 16 << 20 // 16 MiB - maxWindowIDResourceSegment = 64 - unknownWindowResourceSegment = "unknown" -) - -var errUnsafeIncidentPersistencePath = errors.New("unsafe incident recorder persistence path") - -// IncidentDataPoint represents a single data point in an incident window -type IncidentDataPoint struct { - Timestamp time.Time `json:"timestamp"` - Metrics map[string]float64 `json:"metrics"` // cpu, memory, disk, etc. - Metadata map[string]interface{} `json:"metadata,omitempty"` -} - -// IncidentSummary provides computed statistics about an incident window -type IncidentSummary struct { - Duration time.Duration `json:"duration_ms"` - DataPoints int `json:"data_points"` - Peaks map[string]float64 `json:"peaks"` // Maximum values - Lows map[string]float64 `json:"lows"` // Minimum values - Averages map[string]float64 `json:"averages"` // Average values - Changes map[string]float64 `json:"changes"` // Change from start to end - Anomalies []string `json:"anomalies,omitempty"` // Detected anomalies -} - -// IncidentRecorderConfig configures the incident recorder -type IncidentRecorderConfig struct { - // Recording settings - SampleInterval time.Duration // How often to record data points (default: 5s) - PreIncidentWindow time.Duration // How much data to capture before incident (default: 5min) - PostIncidentWindow time.Duration // How much data to capture after incident (default: 10min) - MaxDataPointsPerWindow int // Maximum data points per window (default: 500) - - // Storage settings - DataDir string - MaxWindows int // Maximum number of windows to keep (default: 100) - RetentionDuration time.Duration // How long to keep windows (default: 24h) -} - -// DefaultIncidentRecorderConfig returns sensible defaults -func DefaultIncidentRecorderConfig() IncidentRecorderConfig { - return IncidentRecorderConfig{ - SampleInterval: 5 * time.Second, - PreIncidentWindow: 5 * time.Minute, - PostIncidentWindow: 10 * time.Minute, - MaxDataPointsPerWindow: 500, - MaxWindows: 100, - RetentionDuration: 24 * time.Hour, - } -} - -// MetricsProvider provides current metrics for a resource -type MetricsProvider interface { - GetCurrentMetrics(resourceID string) (map[string]float64, error) - GetMonitoredResourceIDs() []string // Returns all resource IDs being monitored -} - -// BatchMetricsProvider is an optional MetricsProvider extension. recordSample -// asks for metrics once per monitored resource every tick; a provider whose -// per-ID lookup scans all resources turns that into an O(n^2) tick, so when -// this is implemented the recorder fetches the whole batch in one pass. -type BatchMetricsProvider interface { - GetCurrentMetricsBatch() map[string]map[string]float64 -} - -// IncidentRecorder captures high-frequency metrics during incidents -type IncidentRecorder struct { - mu sync.RWMutex - - config IncidentRecorderConfig - provider MetricsProvider - - // Active recordings - activeWindows map[string]*IncidentWindow // keyed by window ID - - // Completed recordings (ring buffer) - completedWindows []*IncidentWindow - - // Background recording for pre-incident buffer - preIncidentBuffer map[string][]IncidentDataPoint // keyed by resource ID - - // Persistence - dataDir string - filePath string - saveIOMu sync.Mutex - - // Control - stopCh chan struct{} - loopDone chan struct{} - running bool - - // Async save coordination - saveMu sync.Mutex - saveCond *sync.Cond - saveInProgress bool - saveRequested bool -} - -// NewIncidentRecorder creates a new incident recorder -func NewIncidentRecorder(cfg IncidentRecorderConfig) *IncidentRecorder { - if cfg.SampleInterval <= 0 { - cfg.SampleInterval = 5 * time.Second - } - if cfg.PreIncidentWindow <= 0 { - cfg.PreIncidentWindow = 5 * time.Minute - } - if cfg.PostIncidentWindow <= 0 { - cfg.PostIncidentWindow = 10 * time.Minute - } - if cfg.MaxDataPointsPerWindow <= 0 { - cfg.MaxDataPointsPerWindow = 500 - } - if cfg.MaxWindows <= 0 { - cfg.MaxWindows = 100 - } - if cfg.RetentionDuration <= 0 { - cfg.RetentionDuration = 24 * time.Hour - } - if cfg.DataDir != "" { - trimmed := strings.TrimSpace(cfg.DataDir) - if trimmed == "" { - log.Warn().Msg("Ignoring incident recorder data dir: blank after trimming whitespace") - cfg.DataDir = "" - } else { - cfg.DataDir = filepath.Clean(trimmed) - } - } - - dataDir := strings.TrimSpace(cfg.DataDir) - if dataDir != "" { - dataDir = filepath.Clean(dataDir) - } - - recorder := &IncidentRecorder{ - config: cfg, - activeWindows: make(map[string]*IncidentWindow), - completedWindows: make([]*IncidentWindow, 0), - preIncidentBuffer: make(map[string][]IncidentDataPoint), - dataDir: dataDir, - stopCh: make(chan struct{}), - loopDone: make(chan struct{}), - } - recorder.saveCond = sync.NewCond(&recorder.saveMu) - - if dataDir != "" { - recorder.filePath = filepath.Join(dataDir, "incident_windows.json") - if err := recorder.loadFromDisk(); err != nil { - log.Warn(). - Str("file_path", recorder.filePath). - Err(err). - Msg("Failed to load incident windows from disk") - } - } - - return recorder -} - -// SetMetricsProvider sets the metrics provider for recording -func (r *IncidentRecorder) SetMetricsProvider(provider MetricsProvider) { - r.mu.Lock() - defer r.mu.Unlock() - r.provider = provider -} - -// Start begins background recording for pre-incident buffer -func (r *IncidentRecorder) Start() { - r.mu.Lock() - if r.running { - r.mu.Unlock() - return - } - stopCh := make(chan struct{}) - loopDone := make(chan struct{}) - r.running = true - r.stopCh = stopCh - r.loopDone = loopDone - r.mu.Unlock() - - go r.recordingLoop(stopCh, loopDone) - log.Info(). - Dur("sample_interval", r.config.SampleInterval). - Dur("pre_incident_window", r.config.PreIncidentWindow). - Dur("post_incident_window", r.config.PostIncidentWindow). - Int("max_data_points_per_window", r.config.MaxDataPointsPerWindow). - Msg("Incident recorder started") -} - -// Stop stops the incident recorder -func (r *IncidentRecorder) Stop() { - r.mu.Lock() - if !r.running { - r.mu.Unlock() - r.waitForPendingSaves() - return - } - r.running = false - close(r.stopCh) - loopDone := r.loopDone - r.mu.Unlock() - - if loopDone != nil { - <-loopDone - } - - // Flush any async save goroutine triggered before shutdown so TempDir cleanup - // and final persistence do not race a background rename/write. - r.waitForPendingSaves() - - // Save to disk - if err := r.saveToDisk(); err != nil { - log.Warn(). - Str("file_path", r.filePath). - Err(err). - Msg("Failed to save incident windows on stop") - } - log.Info().Msg("incident recorder stopped") -} - -// recordingLoop runs in the background to maintain pre-incident buffers and active windows -func (r *IncidentRecorder) recordingLoop(stopCh <-chan struct{}, done chan<- struct{}) { - defer close(done) - - ticker := time.NewTicker(r.config.SampleInterval) - defer ticker.Stop() - - for { - select { - case <-stopCh: - return - case <-ticker.C: - r.recordSample() - } - } -} - -// recordSample captures a data point for all active windows and buffers -func (r *IncidentRecorder) recordSample() { - r.mu.Lock() - if r.provider == nil { - r.mu.Unlock() - return - } - - now := time.Now() - shouldSave := false - - // One pass over the provider when it supports batching; the per-ID - // fallback preserves behavior for providers that do not. A missing batch - // entry maps to the empty metrics the per-ID method returns for unknown - // IDs. - var metricsBatch map[string]map[string]float64 - if batchProvider, ok := r.provider.(BatchMetricsProvider); ok { - metricsBatch = batchProvider.GetCurrentMetricsBatch() - } - currentMetrics := func(resourceID string) (map[string]float64, error) { - if metricsBatch != nil { - if metrics, ok := metricsBatch[resourceID]; ok { - return metrics, nil - } - } - // A batch that lacks the ID must not change per-provider semantics: - // fall through so providers that error or synthesize for unknown IDs - // keep doing exactly that. Misses are the rare path, so the batch - // still absorbs the per-tick fan-out. - return r.provider.GetCurrentMetrics(resourceID) - } - - // Record for active windows - for _, window := range r.activeWindows { - if window.Status != IncidentWindowStatusRecording { - continue - } - - // Check if we've exceeded the post-incident window - if window.EndTime != nil && now.After(*window.EndTime) { - r.completeWindowLocked(window) - shouldSave = true - continue - } - - // Check if we've exceeded max data points - if len(window.DataPoints) >= r.config.MaxDataPointsPerWindow { - window.Status = IncidentWindowStatusTruncated - log.Warn(). - Str("window_id", window.ID). - Str("resource_id", window.ResourceID). - Int("max_data_points_per_window", r.config.MaxDataPointsPerWindow). - Msg("Truncating incident window after reaching max data points") - r.completeWindowLocked(window) - continue - } - - // Get metrics - metrics, err := currentMetrics(window.ResourceID) - if err != nil { - log.Debug(). - Str("window_id", window.ID). - Str("resource_id", window.ResourceID). - Err(err). - Msg("failed to get metrics for incident window") - continue - } - - window.DataPoints = append(window.DataPoints, IncidentDataPoint{ - Timestamp: now, - Metrics: copyMetrics(metrics), - }) - } - - // Continuously buffer ALL monitored resources for pre-incident data - // This ensures we have history when an alert fires on any resource - monitoredResources := r.provider.GetMonitoredResourceIDs() - bufferCutoff := now.Add(-r.config.PreIncidentWindow) - - for _, resourceID := range monitoredResources { - metrics, err := currentMetrics(resourceID) - if err != nil { - log.Debug(). - Str("resource_id", resourceID). - Err(err). - Msg("Failed to get metrics for pre-incident buffer") - continue - } - - // Add to pre-incident buffer - buffer := r.preIncidentBuffer[resourceID] - buffer = append(buffer, IncidentDataPoint{ - Timestamp: now, - Metrics: copyMetrics(metrics), - }) - - // Keep only last PreIncidentWindow duration - kept := make([]IncidentDataPoint, 0, len(buffer)) - for _, dp := range buffer { - if dp.Timestamp.After(bufferCutoff) { - kept = append(kept, dp) - } - } - r.preIncidentBuffer[resourceID] = kept - } - - // Clean up buffers for resources no longer monitored - monitoredSet := make(map[string]bool, len(monitoredResources)) - for _, resourceID := range monitoredResources { - monitoredSet[resourceID] = true - } - for resourceID := range r.preIncidentBuffer { - if !monitoredSet[resourceID] { - delete(r.preIncidentBuffer, resourceID) - } - } - r.mu.Unlock() - - if shouldSave { - if err := r.saveToDisk(); err != nil { - log.Warn().Err(err).Msg("Failed to save incident windows") - } - } -} - -// StartRecording begins recording an incident window -func (r *IncidentRecorder) StartRecording(resourceID, resourceName, resourceType, triggerType, triggerID string) string { - r.mu.Lock() - defer r.mu.Unlock() - - // Check if we already have an active window for this resource - for _, window := range r.activeWindows { - if window.ResourceID == resourceID && window.Status == IncidentWindowStatusRecording { - // Extend existing window - endTime := time.Now().Add(r.config.PostIncidentWindow) - window.EndTime = &endTime - return window.ID - } - } - - // Create new window - windowID := generateWindowID(resourceID) - now := time.Now() - endTime := now.Add(r.config.PostIncidentWindow) - normalizedResourceType := normalizeIncidentResourceType(resourceType) - - window := &IncidentWindow{ - ID: windowID, - ResourceID: resourceID, - ResourceName: resourceName, - ResourceType: normalizedResourceType, - TriggerType: triggerType, - TriggerID: triggerID, - StartTime: now.Add(-r.config.PreIncidentWindow), // Include pre-incident data - EndTime: &endTime, - Status: IncidentWindowStatusRecording, - DataPoints: make([]IncidentDataPoint, 0), - } - - // Copy pre-incident buffer if available - if preBuffer, ok := r.preIncidentBuffer[resourceID]; ok { - window.DataPoints = append(window.DataPoints, copyDataPoints(preBuffer)...) - } - - r.activeWindows[windowID] = window - - log.Info(). - Str("window_id", windowID). - Str("resource_id", resourceID). - Str("trigger_type", triggerType). - Msg("started incident recording") - - return windowID -} - -func normalizeIncidentResourceType(resourceType string) string { - normalized := strings.ToLower(strings.TrimSpace(resourceType)) - if canonical, ok := unifiedresources.CanonicalizeLegacyResourceTypeAlias(normalized); ok { - return canonical - } - return normalized -} - -// StopRecording stops recording for a specific window -func (r *IncidentRecorder) StopRecording(windowID string) { - r.mu.Lock() - shouldSave := false - if window, ok := r.activeWindows[windowID]; ok { - r.completeWindowLocked(window) - shouldSave = true - } - r.mu.Unlock() - - if shouldSave { - if err := r.saveToDisk(); err != nil { - log.Warn().Err(err).Msg("Failed to save incident windows") - } - } -} - -// completeWindowLocked finalizes a recording window. -// Caller must hold r.mu. -func (r *IncidentRecorder) completeWindowLocked(window *IncidentWindow) { - if window.Status != IncidentWindowStatusRecording && window.Status != IncidentWindowStatusTruncated { - return - } - - now := time.Now() - if window.Status == IncidentWindowStatusRecording { - window.Status = IncidentWindowStatusComplete - } - window.EndTime = &now - - // Compute summary - window.Summary = r.computeSummary(window) - - // Move to completed - r.completedWindows = append(r.completedWindows, window) - delete(r.activeWindows, window.ID) - - // Trim completed windows - r.trimCompletedWindows() - - log.Info(). - Str("window_id", window.ID). - Str("resource_id", window.ResourceID). - Str("status", string(window.Status)). - Int("data_points", len(window.DataPoints)). - Msg("completed incident recording") - - // Save asynchronously. - r.requestAsyncSave() -} - -// computeSummary computes statistics for a window -func (r *IncidentRecorder) computeSummary(window *IncidentWindow) *IncidentSummary { - if len(window.DataPoints) == 0 { - return nil - } - - summary := &IncidentSummary{ - DataPoints: len(window.DataPoints), - Peaks: make(map[string]float64), - Lows: make(map[string]float64), - Averages: make(map[string]float64), - Changes: make(map[string]float64), - } - - // Calculate duration - if len(window.DataPoints) > 1 { - first := window.DataPoints[0].Timestamp - last := window.DataPoints[len(window.DataPoints)-1].Timestamp - summary.Duration = last.Sub(first) - } - - // Track sums for averages - sums := make(map[string]float64) - counts := make(map[string]int) - - // First and last values for change calculation - firstValues := make(map[string]float64) - lastValues := make(map[string]float64) - - for i, dp := range window.DataPoints { - for metric, value := range dp.Metrics { - // Track first value - if i == 0 { - firstValues[metric] = value - summary.Peaks[metric] = value - summary.Lows[metric] = value - } - - // Track last value - lastValues[metric] = value - - // Track peaks and lows - if value > summary.Peaks[metric] { - summary.Peaks[metric] = value - } - if value < summary.Lows[metric] { - summary.Lows[metric] = value - } - - // Track sums for average - sums[metric] += value - counts[metric]++ - } - } - - // Calculate averages and changes - for metric, sum := range sums { - if counts[metric] > 0 { - summary.Averages[metric] = sum / float64(counts[metric]) - } - if first, ok := firstValues[metric]; ok { - if last, ok := lastValues[metric]; ok { - summary.Changes[metric] = last - first - } - } - } - - return summary -} - -// trimCompletedWindows removes old windows -func (r *IncidentRecorder) trimCompletedWindows() { - // Remove by retention duration - cutoff := time.Now().Add(-r.config.RetentionDuration) - kept := make([]*IncidentWindow, 0, len(r.completedWindows)) - for _, w := range r.completedWindows { - if w.EndTime != nil && w.EndTime.After(cutoff) { - kept = append(kept, w) - } - } - r.completedWindows = kept - - // Remove by max windows - if len(r.completedWindows) > r.config.MaxWindows { - r.completedWindows = r.completedWindows[len(r.completedWindows)-r.config.MaxWindows:] - } -} - -// GetWindow returns a specific incident window -func (r *IncidentRecorder) GetWindow(windowID string) *IncidentWindow { - r.mu.RLock() - defer r.mu.RUnlock() - - // Check active windows - if window, ok := r.activeWindows[windowID]; ok { - return copyWindow(window) - } - - // Check completed windows - for _, window := range r.completedWindows { - if window.ID == windowID { - return copyWindow(window) - } - } - - return nil -} - -// GetWindowsForResource returns all incident windows for a resource -func (r *IncidentRecorder) GetWindowsForResource(resourceID string, limit int) []*IncidentWindow { - r.mu.RLock() - defer r.mu.RUnlock() - - var result []*IncidentWindow - - // Check active windows - for _, window := range r.activeWindows { - if window.ResourceID == resourceID { - result = append(result, copyWindow(window)) - } - } - - // Check completed windows (in reverse order for most recent first) - for i := len(r.completedWindows) - 1; i >= 0; i-- { - if r.completedWindows[i].ResourceID == resourceID { - result = append(result, copyWindow(r.completedWindows[i])) - if limit > 0 && len(result) >= limit { - break - } - } - } - - return result -} - -// saveToDisk persists completed windows -func (r *IncidentRecorder) saveToDisk() error { - if r.filePath == "" { - return nil - } - r.saveIOMu.Lock() - defer r.saveIOMu.Unlock() - - data := struct { - CompletedWindows []*IncidentWindow `json:"completed_windows"` - }{ - CompletedWindows: r.snapshotCompletedWindows(), - } - - jsonData, err := json.MarshalIndent(data, "", " ") - if err != nil { - return fmt.Errorf("incident recorder save: marshal completed windows: %w", err) - } - - if err := ensureOwnerOnlyDir(r.dataDir); err != nil { - return err - } - - if info, err := os.Lstat(r.filePath); err == nil { - if err := validateRegularFilePath(r.filePath, info); err != nil { - return err - } - } else if !errors.Is(err, os.ErrNotExist) { - return err - } - - tmpFile, err := os.CreateTemp(r.dataDir, filepath.Base(r.filePath)+".*.tmp") - if err != nil { - return err - } - tmpPath := tmpFile.Name() - cleanup := true - defer func() { - if cleanup { - _ = os.Remove(tmpPath) - } - }() - - if err := tmpFile.Chmod(incidentRecorderFilePerm); err != nil { - _ = tmpFile.Close() - return err - } - if _, err := tmpFile.Write(jsonData); err != nil { - _ = tmpFile.Close() - return err - } - if err := tmpFile.Close(); err != nil { - return err - } - if err := os.Rename(tmpPath, r.filePath); err != nil { - return err - } - cleanup = false - return os.Chmod(r.filePath, incidentRecorderFilePerm) -} - -func (r *IncidentRecorder) requestAsyncSave() { - if r.filePath == "" { - return - } - - r.saveMu.Lock() - if r.saveInProgress { - r.saveRequested = true - r.saveMu.Unlock() - return - } - r.saveInProgress = true - r.saveMu.Unlock() - - go r.runAsyncSaves() -} - -func (r *IncidentRecorder) runAsyncSaves() { - for { - if err := r.saveToDisk(); err != nil { - log.Warn(). - Str("file_path", r.filePath). - Err(err). - Msg("Failed to save incident windows") - } - - r.saveMu.Lock() - if !r.saveRequested { - r.saveInProgress = false - r.saveCond.Broadcast() - r.saveMu.Unlock() - return - } - r.saveRequested = false - r.saveMu.Unlock() - } -} - -func (r *IncidentRecorder) snapshotCompletedWindows() []*IncidentWindow { - r.mu.RLock() - defer r.mu.RUnlock() - - snapshot := make([]*IncidentWindow, len(r.completedWindows)) - for i, window := range r.completedWindows { - snapshot[i] = copyWindow(window) - } - return snapshot -} - -func (r *IncidentRecorder) waitForPendingSaves() { - if r.filePath == "" { - return - } - - r.saveMu.Lock() - for r.saveInProgress || r.saveRequested { - r.saveCond.Wait() - } - r.saveMu.Unlock() - - r.saveIOMu.Lock() - r.saveIOMu.Unlock() -} - -// loadFromDisk loads completed windows -func (r *IncidentRecorder) loadFromDisk() error { - if r.filePath == "" { - return nil - } - - jsonData, err := readBoundedRegularFile(r.filePath, maxIncidentWindowsFileSize) - if err != nil { - if errors.Is(err, os.ErrNotExist) { - return nil - } - return fmt.Errorf("incident recorder load: read file %q: %w", r.filePath, err) - } - - var data struct { - CompletedWindows []*IncidentWindow `json:"completed_windows"` - } - - if err := json.Unmarshal(jsonData, &data); err != nil { - return fmt.Errorf("incident recorder load: parse file %q: %w", r.filePath, err) - } - - r.completedWindows = make([]*IncidentWindow, 0, len(data.CompletedWindows)) - for _, window := range data.CompletedWindows { - if window == nil { - continue - } - r.completedWindows = append(r.completedWindows, window) - } - r.trimCompletedWindows() - - return os.Chmod(r.filePath, incidentRecorderFilePerm) -} - -// Helper functions - -func ensureOwnerOnlyDir(dir string) error { - if err := os.MkdirAll(dir, incidentRecorderDirPerm); err != nil { - return err - } - return os.Chmod(dir, incidentRecorderDirPerm) -} - -func validateRegularFilePath(path string, info os.FileInfo) error { - if info.Mode()&os.ModeSymlink != 0 { - return fmt.Errorf("%w: refusing symlink path %q", errUnsafeIncidentPersistencePath, path) - } - if !info.Mode().IsRegular() { - return fmt.Errorf("%w: non-regular path %q", errUnsafeIncidentPersistencePath, path) - } - return nil -} - -func readBoundedRegularFile(path string, maxSize int64) ([]byte, error) { - initialInfo, err := os.Lstat(path) - if err != nil { - return nil, err - } - if err := validateRegularFilePath(path, initialInfo); err != nil { - return nil, err - } - if maxSize > 0 && initialInfo.Size() > maxSize { - return nil, fmt.Errorf("%w: file %q exceeds size limit (%d bytes)", errUnsafeIncidentPersistencePath, path, initialInfo.Size()) - } - - file, err := os.Open(path) - if err != nil { - return nil, err - } - defer func() { - _ = file.Close() - }() - - openInfo, err := file.Stat() - if err != nil { - return nil, err - } - if err := validateRegularFilePath(path, openInfo); err != nil { - return nil, err - } - if !os.SameFile(initialInfo, openInfo) { - return nil, fmt.Errorf("%w: file %q changed during read", errUnsafeIncidentPersistencePath, path) - } - - reader := io.Reader(file) - if maxSize > 0 { - reader = io.LimitReader(file, maxSize+1) - } - - data, err := io.ReadAll(reader) - if err != nil { - return nil, err - } - if maxSize > 0 && int64(len(data)) > maxSize { - return nil, fmt.Errorf("%w: file %q exceeded size limit while reading", errUnsafeIncidentPersistencePath, path) - } - return data, nil -} - -func copyWindow(w *IncidentWindow) *IncidentWindow { - if w == nil { - return nil - } - windowCopy := *w - if w.EndTime != nil { - t := *w.EndTime - windowCopy.EndTime = &t - } - windowCopy.DataPoints = copyDataPoints(w.DataPoints) - if w.Summary != nil { - s := *w.Summary - s.Peaks = copyMetrics(w.Summary.Peaks) - s.Lows = copyMetrics(w.Summary.Lows) - s.Averages = copyMetrics(w.Summary.Averages) - s.Changes = copyMetrics(w.Summary.Changes) - if w.Summary.Anomalies != nil { - s.Anomalies = append([]string(nil), w.Summary.Anomalies...) - } - windowCopy.Summary = &s - } - return &windowCopy -} - -var windowCounter int64 - -func generateWindowID(resourceID string) string { - counter := atomic.AddInt64(&windowCounter, 1) - return "iw-" + resourceID + "-" + time.Now().Format("20060102150405") + "-" + intToString(int(counter)) -} - -func copyDataPoints(points []IncidentDataPoint) []IncidentDataPoint { - copied := make([]IncidentDataPoint, len(points)) - for i, dp := range points { - copied[i] = dp - copied[i].Metrics = copyMetrics(dp.Metrics) - if dp.Metadata != nil { - copied[i].Metadata = make(map[string]interface{}, len(dp.Metadata)) - for k, v := range dp.Metadata { - copied[i].Metadata[k] = v - } - } - } - return copied -} - -func copyMetrics(metrics map[string]float64) map[string]float64 { - if metrics == nil { - return nil - } - - copied := make(map[string]float64, len(metrics)) - for k, v := range metrics { - copied[k] = v - } - return copied -} - -func intToString(n int) string { - if n == 0 { - return "0" - } - negative := n < 0 - if negative { - n = -n - } - var result string - for n > 0 { - result = string(rune('0'+n%10)) + result - n /= 10 - } - if negative { - result = "-" + result - } - return result -} diff --git a/internal/metrics/incident_recorder_additional_test.go b/internal/metrics/incident_recorder_additional_test.go deleted file mode 100644 index 9c26c98a6..000000000 --- a/internal/metrics/incident_recorder_additional_test.go +++ /dev/null @@ -1,133 +0,0 @@ -package metrics - -import ( - "sync/atomic" - "testing" - "time" -) - -type countingProvider struct { - ids []string - metrics map[string]map[string]float64 - calls int32 -} - -func (c *countingProvider) GetCurrentMetrics(resourceID string) (map[string]float64, error) { - atomic.AddInt32(&c.calls, 1) - metrics, ok := c.metrics[resourceID] - if !ok { - return nil, errNoMetrics(resourceID) - } - copied := make(map[string]float64, len(metrics)) - for k, v := range metrics { - copied[k] = v - } - return copied, nil -} - -func (c *countingProvider) GetMonitoredResourceIDs() []string { - return append([]string{}, c.ids...) -} - -func waitForCalls(t *testing.T, provider *countingProvider, timeout time.Duration) { - t.Helper() - deadline := time.Now().Add(timeout) - for time.Now().Before(deadline) { - if atomic.LoadInt32(&provider.calls) > 0 { - return - } - time.Sleep(5 * time.Millisecond) - } - t.Fatal("expected provider to be called") -} - -func TestDefaultIncidentRecorderConfig(t *testing.T) { - cfg := DefaultIncidentRecorderConfig() - if cfg.SampleInterval == 0 || cfg.PreIncidentWindow == 0 || cfg.PostIncidentWindow == 0 { - t.Fatalf("default config should be non-zero, got %+v", cfg) - } - if cfg.MaxDataPointsPerWindow == 0 || cfg.MaxWindows == 0 || cfg.RetentionDuration == 0 { - t.Fatalf("default config should be non-zero, got %+v", cfg) - } -} - -func TestIncidentRecorderStartStop(t *testing.T) { - recorder := NewIncidentRecorder(IncidentRecorderConfig{ - SampleInterval: 5 * time.Millisecond, - PreIncidentWindow: 10 * time.Millisecond, - PostIncidentWindow: 10 * time.Millisecond, - MaxDataPointsPerWindow: 5, - }) - - provider := &countingProvider{ - ids: []string{"res-1"}, - metrics: map[string]map[string]float64{ - "res-1": {"cpu": 1}, - }, - } - recorder.SetMetricsProvider(provider) - - recorder.Start() - waitForCalls(t, provider, 200*time.Millisecond) - recorder.Stop() - - if recorder.running { - t.Fatalf("expected recorder to be stopped") - } -} - -func TestGetWindowsForResource(t *testing.T) { - recorder := NewIncidentRecorder(IncidentRecorderConfig{}) - - active := &IncidentWindow{ID: "active-1", ResourceID: "res-1"} - recorder.activeWindows["active-1"] = active - - got := recorder.GetWindowsForResource("res-1", 0) - if len(got) != 1 || got[0].ID != "active-1" { - t.Fatalf("expected active window, got %+v", got) - } - - recorder.activeWindows = map[string]*IncidentWindow{} - recorder.completedWindows = []*IncidentWindow{ - {ID: "old", ResourceID: "res-1"}, - {ID: "new", ResourceID: "res-1"}, - } - - limited := recorder.GetWindowsForResource("res-1", 1) - if len(limited) != 1 || limited[0].ID != "new" { - t.Fatalf("expected most recent completed window, got %+v", limited) - } -} - -func TestRecordSampleSkipsPreIncidentBufferOnMetricsError(t *testing.T) { - recorder := NewIncidentRecorder(IncidentRecorderConfig{ - PreIncidentWindow: time.Minute, - PostIncidentWindow: time.Minute, - MaxDataPointsPerWindow: 10, - }) - - provider := &stubMetricsProvider{ - metricsByID: map[string]map[string]float64{ - "res-ok": {"cpu": 1}, - }, - ids: []string{"res-ok", "res-missing"}, - } - recorder.SetMetricsProvider(provider) - - windowID := recorder.StartRecording("res-ok", "db", "agent", "alert", "alert-1") - recorder.recordSample() - - window := recorder.activeWindows[windowID] - if window == nil { - t.Fatalf("expected active window %s", windowID) - } - if len(window.DataPoints) != 1 { - t.Fatalf("expected active window sample to be captured, got %d", len(window.DataPoints)) - } - if len(recorder.preIncidentBuffer["res-ok"]) == 0 { - t.Fatalf("expected pre-incident buffer for res-ok") - } - if _, ok := recorder.preIncidentBuffer["res-missing"]; ok { - t.Fatalf("expected no pre-incident buffer for res-missing when metrics collection fails") - } -} diff --git a/internal/metrics/incident_recorder_batch_test.go b/internal/metrics/incident_recorder_batch_test.go deleted file mode 100644 index 741b7ebb9..000000000 --- a/internal/metrics/incident_recorder_batch_test.go +++ /dev/null @@ -1,84 +0,0 @@ -package metrics - -import ( - "errors" - "testing" - "time" -) - -// gappyBatchProvider implements both MetricsProvider and BatchMetricsProvider -// but omits one monitored ID from the batch. The recorder must fall through -// to the per-ID method for that ID so provider semantics are preserved. -type gappyBatchProvider struct { - batchCalls int - perIDCalls map[string]int - perIDResults map[string]map[string]float64 - perIDErr map[string]error -} - -func (p *gappyBatchProvider) GetMonitoredResourceIDs() []string { - return []string{"in-batch", "not-in-batch", "erroring"} -} - -func (p *gappyBatchProvider) GetCurrentMetricsBatch() map[string]map[string]float64 { - p.batchCalls++ - return map[string]map[string]float64{ - "in-batch": {"cpu": 42}, - } -} - -func (p *gappyBatchProvider) GetCurrentMetrics(resourceID string) (map[string]float64, error) { - if p.perIDCalls == nil { - p.perIDCalls = map[string]int{} - } - p.perIDCalls[resourceID]++ - if err := p.perIDErr[resourceID]; err != nil { - return nil, err - } - return p.perIDResults[resourceID], nil -} - -func TestRecordSampleBatchMissFallsThroughToPerID(t *testing.T) { - recorder := NewIncidentRecorder(IncidentRecorderConfig{ - SampleInterval: time.Hour, // ticks driven manually - PreIncidentWindow: time.Minute, - PostIncidentWindow: time.Minute, - MaxDataPointsPerWindow: 5, - }) - provider := &gappyBatchProvider{ - perIDResults: map[string]map[string]float64{ - "not-in-batch": {"cpu": 7}, - }, - perIDErr: map[string]error{ - "erroring": errors.New("unknown resource"), - }, - } - recorder.SetMetricsProvider(provider) - - recorder.recordSample() - - if provider.batchCalls != 1 { - t.Fatalf("batch calls = %d, want 1", provider.batchCalls) - } - if provider.perIDCalls["in-batch"] != 0 { - t.Fatalf("batch hit %q still took the per-ID path", "in-batch") - } - if provider.perIDCalls["not-in-batch"] != 1 || provider.perIDCalls["erroring"] != 1 { - t.Fatalf("batch misses did not fall through per-ID: %+v", provider.perIDCalls) - } - - recorder.mu.RLock() - defer recorder.mu.RUnlock() - if got := len(recorder.preIncidentBuffer["not-in-batch"]); got != 1 { - t.Fatalf("fallthrough metrics not buffered: %d points", got) - } - if got := recorder.preIncidentBuffer["not-in-batch"][0].Metrics["cpu"]; got != 7 { - t.Fatalf("fallthrough buffered cpu = %v, want 7 (per-ID value)", got) - } - if got := len(recorder.preIncidentBuffer["erroring"]); got != 0 { - t.Fatalf("erroring ID gained %d buffered points, want 0 (per-ID error must skip)", got) - } - if got := len(recorder.preIncidentBuffer["in-batch"]); got != 1 { - t.Fatalf("batch-served ID not buffered: %d points", got) - } -} diff --git a/internal/metrics/incident_recorder_concurrency_test.go b/internal/metrics/incident_recorder_concurrency_test.go deleted file mode 100644 index d56ebec6e..000000000 --- a/internal/metrics/incident_recorder_concurrency_test.go +++ /dev/null @@ -1,86 +0,0 @@ -package metrics - -import ( - "sync" - "testing" - "time" -) - -func TestGenerateWindowIDConcurrentUnique(t *testing.T) { - t.Parallel() - - const ( - workers = 16 - idsPerWork = 128 - ) - - ids := make(chan string, workers*idsPerWork) - start := make(chan struct{}) - - var wg sync.WaitGroup - for i := 0; i < workers; i++ { - wg.Add(1) - go func() { - defer wg.Done() - <-start - for j := 0; j < idsPerWork; j++ { - ids <- generateWindowID("res-1") - } - }() - } - - close(start) - wg.Wait() - close(ids) - - seen := make(map[string]struct{}, workers*idsPerWork) - for id := range ids { - if _, exists := seen[id]; exists { - t.Fatalf("duplicate window ID generated: %s", id) - } - seen[id] = struct{}{} - } -} - -func TestIncidentRecorderConcurrentStartStopAndFlush(t *testing.T) { - recorder := NewIncidentRecorder(IncidentRecorderConfig{ - SampleInterval: time.Millisecond, - PreIncidentWindow: 10 * time.Millisecond, - PostIncidentWindow: 10 * time.Millisecond, - MaxDataPointsPerWindow: 10, - DataDir: t.TempDir(), - }) - provider := &stubMetricsProvider{ - metricsByID: map[string]map[string]float64{ - "res-1": {"cpu": 1}, - }, - ids: []string{"res-1"}, - } - recorder.SetMetricsProvider(provider) - - const goroutines = 8 - const iterations = 15 - start := make(chan struct{}) - - var wg sync.WaitGroup - for i := 0; i < goroutines; i++ { - wg.Add(1) - go func() { - defer wg.Done() - <-start - for j := 0; j < iterations; j++ { - recorder.Start() - windowID := recorder.StartRecording("res-1", "db", "agent", "alert", "a-1") - recorder.recordSample() - recorder.StopRecording(windowID) - recorder.Stop() - } - }() - } - - close(start) - wg.Wait() - - // Final stop should remain idempotent after concurrent shutdowns. - recorder.Stop() -} diff --git a/internal/metrics/incident_recorder_coverage_test.go b/internal/metrics/incident_recorder_coverage_test.go deleted file mode 100644 index f365531ae..000000000 --- a/internal/metrics/incident_recorder_coverage_test.go +++ /dev/null @@ -1,431 +0,0 @@ -package metrics - -import ( - "bytes" - "math" - "os" - "path/filepath" - "sync/atomic" - "testing" - "time" - - "github.com/rs/zerolog" - "github.com/rs/zerolog/log" -) - -func TestNewIncidentRecorderLoadFromDiskInvalidJSON(t *testing.T) { - dir := t.TempDir() - path := filepath.Join(dir, "incident_windows.json") - if err := os.WriteFile(path, []byte("{invalid"), 0600); err != nil { - t.Fatalf("write invalid json: %v", err) - } - - recorder := NewIncidentRecorder(IncidentRecorderConfig{DataDir: dir}) - if recorder == nil { - t.Fatal("expected recorder") - } - if len(recorder.completedWindows) != 0 { - t.Fatalf("expected no completed windows on invalid json, got %d", len(recorder.completedWindows)) - } -} - -func TestIncidentRecorderStartStopIdempotentGuards(t *testing.T) { - recorder := NewIncidentRecorder(IncidentRecorderConfig{ - SampleInterval: 10 * time.Millisecond, - }) - - recorder.Start() - firstStopCh := recorder.stopCh - - recorder.Start() - if recorder.stopCh != firstStopCh { - t.Fatal("expected second Start call to be a no-op while running") - } - - recorder.Stop() - if recorder.running { - t.Fatal("expected recorder to be stopped") - } - - // Should be a no-op and should not panic. - recorder.Stop() -} - -func TestIncidentRecorderStopWaitsForPendingSavesWhenAlreadyStopped(t *testing.T) { - recorder := NewIncidentRecorder(IncidentRecorderConfig{ - DataDir: t.TempDir(), - }) - - recorder.saveMu.Lock() - recorder.saveInProgress = true - recorder.saveMu.Unlock() - - done := make(chan struct{}) - go func() { - recorder.Stop() - close(done) - }() - - select { - case <-done: - t.Fatal("expected Stop to wait for pending saves even when recorder is already stopped") - case <-time.After(20 * time.Millisecond): - } - - recorder.saveMu.Lock() - recorder.saveInProgress = false - recorder.saveCond.Broadcast() - recorder.saveMu.Unlock() - - select { - case <-done: - case <-time.After(time.Second): - t.Fatal("timed out waiting for Stop to return after pending saves were cleared") - } -} - -func TestRecordSampleNoProviderNoop(t *testing.T) { - recorder := NewIncidentRecorder(IncidentRecorderConfig{}) - recorder.recordSample() -} - -func TestRecordSampleCoversActiveWindowBranches(t *testing.T) { - recorder := NewIncidentRecorder(IncidentRecorderConfig{ - PreIncidentWindow: time.Second, - PostIncidentWindow: time.Second, - MaxDataPointsPerWindow: 1, - MaxWindows: 10, - RetentionDuration: time.Hour, - }) - provider := &stubMetricsProvider{ - metricsByID: map[string]map[string]float64{ - "res-ok": {"cpu": 1}, - "res-active-ok": {"cpu": 2}, - }, - ids: []string{"res-ok", "res-buffer-missing"}, - } - recorder.SetMetricsProvider(provider) - - now := time.Now() - past := now.Add(-time.Millisecond) - recorder.activeWindows["skip-non-recording"] = &IncidentWindow{ - ID: "skip-non-recording", - ResourceID: "res-active-ok", - Status: IncidentWindowStatusComplete, - EndTime: &past, - } - recorder.activeWindows["expire-now"] = &IncidentWindow{ - ID: "expire-now", - ResourceID: "res-active-ok", - Status: IncidentWindowStatusRecording, - EndTime: &past, - } - recorder.activeWindows["truncate-now"] = &IncidentWindow{ - ID: "truncate-now", - ResourceID: "res-active-ok", - Status: IncidentWindowStatusRecording, - DataPoints: []IncidentDataPoint{ - {Timestamp: now, Metrics: map[string]float64{"cpu": 7}}, - }, - } - recorder.activeWindows["metrics-error"] = &IncidentWindow{ - ID: "metrics-error", - ResourceID: "res-active-missing", - Status: IncidentWindowStatusRecording, - } - - recorder.preIncidentBuffer["res-ok"] = []IncidentDataPoint{ - {Timestamp: now.Add(-2 * time.Second), Metrics: map[string]float64{"cpu": 0.5}}, - } - recorder.preIncidentBuffer["stale-resource"] = []IncidentDataPoint{ - {Timestamp: now, Metrics: map[string]float64{"cpu": 9}}, - } - - recorder.recordSample() - - if _, ok := recorder.activeWindows["expire-now"]; ok { - t.Fatal("expected expired window to complete") - } - if _, ok := recorder.activeWindows["truncate-now"]; ok { - t.Fatal("expected truncated window to complete") - } - - metricsErrWindow, ok := recorder.activeWindows["metrics-error"] - if !ok { - t.Fatal("expected metrics-error window to remain active") - } - if len(metricsErrWindow.DataPoints) != 0 { - t.Fatalf("expected metrics-error window to skip append, got %d points", len(metricsErrWindow.DataPoints)) - } - - if _, ok := recorder.preIncidentBuffer["stale-resource"]; ok { - t.Fatal("expected stale pre-incident buffer to be removed") - } - if got := len(recorder.preIncidentBuffer["res-ok"]); got != 1 { - t.Fatalf("expected pre-incident buffer trim to keep 1 point, got %d", got) - } - - foundTruncated := false - foundCompleted := false - for _, w := range recorder.completedWindows { - if w.ID == "truncate-now" && w.Status == IncidentWindowStatusTruncated { - foundTruncated = true - } - if w.ID == "expire-now" && w.Status == IncidentWindowStatusComplete { - foundCompleted = true - } - } - if !foundTruncated { - t.Fatal("expected truncated completed window") - } - if !foundCompleted { - t.Fatal("expected completed expired window") - } -} - -func TestStartRecordingCopiesPreIncidentBuffer(t *testing.T) { - recorder := NewIncidentRecorder(IncidentRecorderConfig{ - PreIncidentWindow: time.Minute, - PostIncidentWindow: time.Minute, - }) - recorder.preIncidentBuffer["res-1"] = []IncidentDataPoint{ - {Timestamp: time.Now().Add(-30 * time.Second), Metrics: map[string]float64{"cpu": 1}}, - } - - windowID := recorder.StartRecording("res-1", "db", "agent", "alert", "alert-1") - window := recorder.activeWindows[windowID] - if window == nil { - t.Fatalf("expected active window %s", windowID) - } - if len(window.DataPoints) != 1 { - t.Fatalf("expected pre-incident points to be copied, got %d", len(window.DataPoints)) - } -} - -func TestCompleteWindowNoopForNonRecordingStatus(t *testing.T) { - recorder := NewIncidentRecorder(IncidentRecorderConfig{}) - window := &IncidentWindow{ - ID: "already-complete", - ResourceID: "res-1", - Status: IncidentWindowStatusComplete, - } - recorder.activeWindows[window.ID] = window - - recorder.completeWindowLocked(window) - - if len(recorder.completedWindows) != 0 { - t.Fatalf("expected no completed windows to be appended, got %d", len(recorder.completedWindows)) - } - if _, ok := recorder.activeWindows[window.ID]; !ok { - t.Fatal("expected window to remain in active map when completion is skipped") - } -} - -func TestComputeSummaryNoDataPoints(t *testing.T) { - recorder := NewIncidentRecorder(IncidentRecorderConfig{}) - if summary := recorder.computeSummary(&IncidentWindow{}); summary != nil { - t.Fatal("expected nil summary when there are no data points") - } -} - -func TestTrimCompletedWindowsEnforcesMaxWindows(t *testing.T) { - recorder := NewIncidentRecorder(IncidentRecorderConfig{ - MaxWindows: 2, - RetentionDuration: time.Hour, - }) - now := time.Now() - recorder.completedWindows = []*IncidentWindow{ - {ID: "w1", EndTime: &now}, - {ID: "w2", EndTime: &now}, - {ID: "w3", EndTime: &now}, - } - - recorder.trimCompletedWindows() - - if len(recorder.completedWindows) != 2 { - t.Fatalf("expected 2 windows after trim, got %d", len(recorder.completedWindows)) - } - if recorder.completedWindows[0].ID != "w2" || recorder.completedWindows[1].ID != "w3" { - t.Fatalf("expected newest windows to be retained, got %s and %s", recorder.completedWindows[0].ID, recorder.completedWindows[1].ID) - } -} - -func TestGetWindowActiveAndMissing(t *testing.T) { - recorder := NewIncidentRecorder(IncidentRecorderConfig{}) - recorder.activeWindows["active-1"] = &IncidentWindow{ID: "active-1", ResourceID: "res-1"} - - if got := recorder.GetWindow("active-1"); got == nil { - t.Fatal("expected active window to be returned") - } - if got := recorder.GetWindow("does-not-exist"); got != nil { - t.Fatal("expected nil for missing window") - } -} - -func TestStopHandlesSaveErrorPath(t *testing.T) { - recorder := NewIncidentRecorder(IncidentRecorderConfig{ - DataDir: t.TempDir(), - SampleInterval: 50 * time.Millisecond, - PreIncidentWindow: 10 * time.Millisecond, - }) - now := time.Now() - recorder.completedWindows = []*IncidentWindow{ - { - ID: "nan-window", - EndTime: &now, - DataPoints: []IncidentDataPoint{ - {Timestamp: now, Metrics: map[string]float64{"cpu": math.NaN()}}, - }, - }, - } - - recorder.Start() - recorder.Stop() -} - -type logSignalWriter struct { - hit atomic.Bool -} - -func (w *logSignalWriter) Write(p []byte) (int, error) { - if bytes.Contains(p, []byte("Failed to save incident windows")) { - w.hit.Store(true) - } - return len(p), nil -} - -func TestCompleteWindowAsyncSaveErrorPath(t *testing.T) { - base := t.TempDir() - fileAsDir := filepath.Join(base, "file-instead-of-dir") - if err := os.WriteFile(fileAsDir, []byte("x"), 0600); err != nil { - t.Fatalf("write setup file: %v", err) - } - - writer := &logSignalWriter{} - originalLogger := log.Logger - log.Logger = zerolog.New(writer).Level(zerolog.WarnLevel) - t.Cleanup(func() { - log.Logger = originalLogger - }) - - recorder := NewIncidentRecorder(IncidentRecorderConfig{ - DataDir: fileAsDir, - RetentionDuration: time.Hour, - MaxWindows: 10, - }) - window := &IncidentWindow{ - ID: "async-save-error", - ResourceID: "res-1", - Status: IncidentWindowStatusRecording, - DataPoints: []IncidentDataPoint{ - {Timestamp: time.Now(), Metrics: map[string]float64{"cpu": 42}}, - }, - } - recorder.activeWindows[window.ID] = window - - recorder.completeWindowLocked(window) - recorder.waitForPendingSaves() - - if !writer.hit.Load() { - t.Fatal("expected async save error warning to be logged") - } -} - -func TestSaveToDiskErrorPaths(t *testing.T) { - t.Run("marshal error", func(t *testing.T) { - recorder := NewIncidentRecorder(IncidentRecorderConfig{DataDir: t.TempDir()}) - now := time.Now() - recorder.completedWindows = []*IncidentWindow{ - { - ID: "marshal-fail", - EndTime: &now, - DataPoints: []IncidentDataPoint{ - {Timestamp: now, Metrics: map[string]float64{"cpu": math.NaN()}}, - }, - }, - } - - if err := recorder.saveToDisk(); err == nil { - t.Fatal("expected marshal error") - } - }) - - t.Run("mkdir error", func(t *testing.T) { - base := t.TempDir() - fileAsDir := filepath.Join(base, "file-instead-of-dir") - if err := os.WriteFile(fileAsDir, []byte("x"), 0600); err != nil { - t.Fatalf("write setup file: %v", err) - } - - recorder := NewIncidentRecorder(IncidentRecorderConfig{DataDir: fileAsDir}) - now := time.Now() - recorder.completedWindows = []*IncidentWindow{{ID: "w1", EndTime: &now}} - - if err := recorder.saveToDisk(); err == nil { - t.Fatal("expected mkdir error") - } - }) - - t.Run("write temp file error", func(t *testing.T) { - dir := t.TempDir() - recorder := NewIncidentRecorder(IncidentRecorderConfig{DataDir: dir}) - now := time.Now() - recorder.completedWindows = []*IncidentWindow{{ID: "w2", EndTime: &now}} - recorder.filePath = filepath.Join(dir, "missing-subdir", "incident_windows.json") - - if err := recorder.saveToDisk(); err == nil { - t.Fatal("expected write temp file error") - } - }) -} - -func TestLoadFromDiskErrorPaths(t *testing.T) { - t.Run("empty file path", func(t *testing.T) { - recorder := NewIncidentRecorder(IncidentRecorderConfig{}) - recorder.filePath = "" - if err := recorder.loadFromDisk(); err != nil { - t.Fatalf("expected nil error for empty file path, got %v", err) - } - }) - - t.Run("missing file", func(t *testing.T) { - recorder := NewIncidentRecorder(IncidentRecorderConfig{DataDir: t.TempDir()}) - if err := recorder.loadFromDisk(); err != nil { - t.Fatalf("expected nil error for missing file, got %v", err) - } - }) - - t.Run("read error", func(t *testing.T) { - recorder := NewIncidentRecorder(IncidentRecorderConfig{}) - recorder.filePath = t.TempDir() - if err := recorder.loadFromDisk(); err == nil { - t.Fatal("expected read error") - } - }) - - t.Run("unmarshal error", func(t *testing.T) { - path := filepath.Join(t.TempDir(), "incident_windows.json") - if err := os.WriteFile(path, []byte("{"), 0600); err != nil { - t.Fatalf("write invalid json: %v", err) - } - - recorder := NewIncidentRecorder(IncidentRecorderConfig{}) - recorder.filePath = path - if err := recorder.loadFromDisk(); err == nil { - t.Fatal("expected unmarshal error") - } - }) -} - -func TestCopyWindowNilAndIntToStringEdges(t *testing.T) { - if copyWindow(nil) != nil { - t.Fatal("expected nil copy for nil input") - } - - if got := intToString(0); got != "0" { - t.Fatalf("expected 0, got %s", got) - } - if got := intToString(-42); got != "-42" { - t.Fatalf("expected -42, got %s", got) - } -} diff --git a/internal/metrics/incident_recorder_test.go b/internal/metrics/incident_recorder_test.go deleted file mode 100644 index 2d48ffe84..000000000 --- a/internal/metrics/incident_recorder_test.go +++ /dev/null @@ -1,438 +0,0 @@ -package metrics - -import ( - "os" - "path/filepath" - "strings" - "testing" - "time" -) - -type stubMetricsProvider struct { - metricsByID map[string]map[string]float64 - ids []string -} - -func (s *stubMetricsProvider) GetCurrentMetrics(resourceID string) (map[string]float64, error) { - metrics, ok := s.metricsByID[resourceID] - if !ok { - return nil, errNoMetrics(resourceID) - } - copied := make(map[string]float64, len(metrics)) - for k, v := range metrics { - copied[k] = v - } - return copied, nil -} - -func (s *stubMetricsProvider) GetMonitoredResourceIDs() []string { - return append([]string{}, s.ids...) -} - -type errNoMetrics string - -func (e errNoMetrics) Error() string { - return "no metrics for " + string(e) -} - -func TestNewIncidentRecorderDefaults(t *testing.T) { - recorder := NewIncidentRecorder(IncidentRecorderConfig{}) - - if recorder.config.SampleInterval != 5*time.Second { - t.Fatalf("expected default sample interval, got %s", recorder.config.SampleInterval) - } - if recorder.config.PreIncidentWindow != 5*time.Minute { - t.Fatalf("expected default pre-incident window, got %s", recorder.config.PreIncidentWindow) - } - if recorder.config.PostIncidentWindow != 10*time.Minute { - t.Fatalf("expected default post-incident window, got %s", recorder.config.PostIncidentWindow) - } - if recorder.config.MaxDataPointsPerWindow != 500 { - t.Fatalf("expected default max data points, got %d", recorder.config.MaxDataPointsPerWindow) - } - if recorder.config.MaxWindows != 100 { - t.Fatalf("expected default max windows, got %d", recorder.config.MaxWindows) - } - if recorder.config.RetentionDuration != 24*time.Hour { - t.Fatalf("expected default retention, got %s", recorder.config.RetentionDuration) - } -} - -func TestNewIncidentRecorderTrimsDataDir(t *testing.T) { - dir := t.TempDir() - - recorder := NewIncidentRecorder(IncidentRecorderConfig{ - DataDir: " " + dir + " ", - }) - - if recorder.config.DataDir != dir { - t.Fatalf("expected trimmed data dir %q, got %q", dir, recorder.config.DataDir) - } - if recorder.dataDir != dir { - t.Fatalf("expected recorder data dir %q, got %q", dir, recorder.dataDir) - } - wantPath := filepath.Join(dir, "incident_windows.json") - if recorder.filePath != wantPath { - t.Fatalf("expected file path %q, got %q", wantPath, recorder.filePath) - } -} - -func TestNewIncidentRecorderWhitespaceOnlyDataDirDisablesPersistence(t *testing.T) { - recorder := NewIncidentRecorder(IncidentRecorderConfig{ - DataDir: " ", - }) - - if recorder.config.DataDir != "" { - t.Fatalf("expected empty config data dir, got %q", recorder.config.DataDir) - } - if recorder.dataDir != "" { - t.Fatalf("expected empty recorder data dir, got %q", recorder.dataDir) - } - if recorder.filePath != "" { - t.Fatalf("expected empty file path, got %q", recorder.filePath) - } -} - -func TestStartRecordingExtendsWindow(t *testing.T) { - recorder := NewIncidentRecorder(IncidentRecorderConfig{ - PreIncidentWindow: time.Minute, - PostIncidentWindow: time.Minute, - }) - - firstID := recorder.StartRecording("res-1", "db", "agent", "alert", "alert-1") - firstWindow := recorder.activeWindows[firstID] - if firstWindow == nil { - t.Fatalf("expected window for %s", firstID) - } - firstEnd := *firstWindow.EndTime - - secondID := recorder.StartRecording("res-1", "db", "agent", "alert", "alert-2") - if secondID != firstID { - t.Fatalf("expected same window ID, got %s and %s", firstID, secondID) - } - secondWindow := recorder.activeWindows[secondID] - if secondWindow.EndTime.Before(firstEnd) { - t.Fatalf("expected end time to extend or remain, got %s before %s", secondWindow.EndTime, firstEnd) - } -} - -func TestStartRecordingCanonicalizesLegacyHostAlias(t *testing.T) { - recorder := NewIncidentRecorder(IncidentRecorderConfig{ - PreIncidentWindow: time.Minute, - PostIncidentWindow: time.Minute, - }) - - windowID := recorder.StartRecording("res-1", "db", "host", "alert", "alert-1") - window := recorder.activeWindows[windowID] - if window == nil { - t.Fatalf("expected window for %s", windowID) - } - if window.ResourceType != "agent" { - t.Fatalf("expected legacy host alias to canonicalize to agent, got %q", window.ResourceType) - } -} - -func TestRecordSampleBuffersAndCleansUp(t *testing.T) { - recorder := NewIncidentRecorder(IncidentRecorderConfig{ - PreIncidentWindow: time.Minute, - PostIncidentWindow: time.Minute, - MaxDataPointsPerWindow: 10, - }) - - provider := &stubMetricsProvider{ - metricsByID: map[string]map[string]float64{ - "res-1": {"cpu": 1}, - "res-2": {"cpu": 2}, - }, - ids: []string{"res-1", "res-2"}, - } - recorder.SetMetricsProvider(provider) - - recorder.preIncidentBuffer["gone"] = []IncidentDataPoint{ - {Timestamp: time.Now().Add(-time.Minute), Metrics: map[string]float64{"cpu": 0.5}}, - } - - windowID := recorder.StartRecording("res-1", "db", "agent", "alert", "alert-1") - recorder.recordSample() - - window := recorder.activeWindows[windowID] - if window == nil { - t.Fatalf("expected active window %s", windowID) - } - if len(window.DataPoints) != 1 { - t.Fatalf("expected 1 data point, got %d", len(window.DataPoints)) - } - - if len(recorder.preIncidentBuffer["res-1"]) == 0 { - t.Fatalf("expected pre-incident buffer for res-1") - } - if len(recorder.preIncidentBuffer["res-2"]) == 0 { - t.Fatalf("expected pre-incident buffer for res-2") - } - if _, ok := recorder.preIncidentBuffer["gone"]; ok { - t.Fatalf("expected cleanup of unmonitored resource buffer") - } -} - -func TestStopRecordingCompletesWindow(t *testing.T) { - recorder := NewIncidentRecorder(IncidentRecorderConfig{ - PreIncidentWindow: time.Minute, - PostIncidentWindow: time.Minute, - }) - provider := &stubMetricsProvider{ - metricsByID: map[string]map[string]float64{ - "res-1": {"cpu": 1}, - }, - ids: []string{"res-1"}, - } - recorder.SetMetricsProvider(provider) - - windowID := recorder.StartRecording("res-1", "db", "agent", "alert", "alert-1") - recorder.recordSample() - recorder.StopRecording(windowID) - - if _, ok := recorder.activeWindows[windowID]; ok { - t.Fatalf("expected window %s to be removed from active windows", windowID) - } - if len(recorder.completedWindows) != 1 { - t.Fatalf("expected 1 completed window, got %d", len(recorder.completedWindows)) - } - if recorder.completedWindows[0].Status != IncidentWindowStatusComplete { - t.Fatalf("expected completed status, got %s", recorder.completedWindows[0].Status) - } - if recorder.completedWindows[0].Summary == nil { - t.Fatalf("expected summary to be computed") - } -} - -func TestComputeSummary(t *testing.T) { - recorder := NewIncidentRecorder(IncidentRecorderConfig{}) - start := time.Now().Add(-time.Second) - end := start.Add(time.Second) - window := &IncidentWindow{ - DataPoints: []IncidentDataPoint{ - {Timestamp: start, Metrics: map[string]float64{"cpu": 1, "mem": 4}}, - {Timestamp: end, Metrics: map[string]float64{"cpu": 3, "mem": 2}}, - }, - } - - summary := recorder.computeSummary(window) - if summary == nil { - t.Fatalf("expected summary") - } - if summary.DataPoints != 2 { - t.Fatalf("expected 2 data points, got %d", summary.DataPoints) - } - if summary.Peaks["cpu"] != 3 || summary.Lows["cpu"] != 1 { - t.Fatalf("unexpected cpu stats: peaks=%v lows=%v", summary.Peaks["cpu"], summary.Lows["cpu"]) - } - if summary.Peaks["mem"] != 4 || summary.Lows["mem"] != 2 { - t.Fatalf("unexpected mem stats: peaks=%v lows=%v", summary.Peaks["mem"], summary.Lows["mem"]) - } - if summary.Averages["cpu"] != 2 { - t.Fatalf("unexpected cpu average: %v", summary.Averages["cpu"]) - } - if summary.Averages["mem"] != 3 { - t.Fatalf("unexpected mem average: %v", summary.Averages["mem"]) - } - if summary.Changes["cpu"] != 2 || summary.Changes["mem"] != -2 { - t.Fatalf("unexpected changes: cpu=%v mem=%v", summary.Changes["cpu"], summary.Changes["mem"]) - } - if summary.Duration != time.Second { - t.Fatalf("unexpected duration: %s", summary.Duration) - } -} - -func TestCopyWindowDeepCopy(t *testing.T) { - now := time.Now() - end := now.Add(time.Second) - window := &IncidentWindow{ - ID: "window-1", - EndTime: &end, - DataPoints: []IncidentDataPoint{ - { - Timestamp: now, - Metrics: map[string]float64{"cpu": 1}, - Metadata: map[string]interface{}{"host": "node-1"}, - }, - }, - Summary: &IncidentSummary{ - Peaks: map[string]float64{"cpu": 1}, - Lows: map[string]float64{"cpu": 1}, - Averages: map[string]float64{"cpu": 1}, - Changes: map[string]float64{"cpu": 0}, - Anomalies: []string{"initial"}, - }, - } - - clone := copyWindow(window) - if clone == nil || clone == window { - t.Fatalf("expected deep copy") - } - if clone.Summary == window.Summary { - t.Fatalf("expected summary to be copied") - } - - window.DataPoints[0].Metrics["cpu"] = 9 - window.DataPoints[0].Metadata["host"] = "mutated" - window.Summary.Peaks["cpu"] = 9 - window.Summary.Anomalies[0] = "mutated" - *window.EndTime = end.Add(5 * time.Second) - window.Summary.Peaks["cpu"] = 9 - - if clone.DataPoints[0].Metrics["cpu"] != 1 { - t.Fatalf("expected data points to be copied") - } - if clone.DataPoints[0].Metadata["host"] != "node-1" { - t.Fatalf("expected metadata to be copied") - } - if clone.EndTime.Equal(*window.EndTime) { - t.Fatalf("expected end time to be copied") - } - if clone.Summary.Peaks["cpu"] != 1 { - t.Fatalf("expected summary maps to be copied") - } - if clone.Summary.Anomalies[0] != "initial" { - t.Fatalf("expected summary anomalies to be copied") - } -} - -func TestSaveAndLoad(t *testing.T) { - dir := t.TempDir() - recorder := NewIncidentRecorder(IncidentRecorderConfig{DataDir: dir}) - - end := time.Now() - recorder.completedWindows = []*IncidentWindow{ - { - ID: "window-1", - EndTime: &end, - Status: IncidentWindowStatusComplete, - DataPoints: []IncidentDataPoint{{Timestamp: end, Metrics: map[string]float64{"cpu": 1}}}, - }, - } - - if err := recorder.saveToDisk(); err != nil { - t.Fatalf("save failed: %v", err) - } - - loaded := NewIncidentRecorder(IncidentRecorderConfig{DataDir: dir}) - window := loaded.GetWindow("window-1") - if window == nil { - t.Fatalf("expected window to load from disk") - } - if window.Status != IncidentWindowStatusComplete { - t.Fatalf("expected status to persist, got %s", window.Status) - } -} - -func TestSaveToDiskSecuresPermissions(t *testing.T) { - dir := t.TempDir() - if err := os.Chmod(dir, 0o755); err != nil { - t.Fatalf("chmod dir failed: %v", err) - } - - recorder := NewIncidentRecorder(IncidentRecorderConfig{DataDir: dir}) - recorder.completedWindows = []*IncidentWindow{{ID: "window-1", EndTime: ptrTime(time.Now())}} - - if err := recorder.saveToDisk(); err != nil { - t.Fatalf("save failed: %v", err) - } - - dirInfo, err := os.Stat(dir) - if err != nil { - t.Fatalf("stat dir failed: %v", err) - } - if got := dirInfo.Mode().Perm(); got != 0o700 { - t.Fatalf("expected dir permissions 0700, got %o", got) - } - - fileInfo, err := os.Stat(filepath.Join(dir, "incident_windows.json")) - if err != nil { - t.Fatalf("stat file failed: %v", err) - } - if got := fileInfo.Mode().Perm(); got != 0o600 { - t.Fatalf("expected file permissions 0600, got %o", got) - } -} - -func TestLoadFromDiskRejectsSymlink(t *testing.T) { - dir := t.TempDir() - target := filepath.Join(dir, "target.json") - if err := os.WriteFile(target, []byte("{}"), 0o600); err != nil { - t.Fatalf("write target failed: %v", err) - } - link := filepath.Join(dir, "incident_windows.json") - requireSymlinkOrSkip(t, target, link) - - recorder := &IncidentRecorder{ - config: DefaultIncidentRecorderConfig(), - dataDir: dir, - filePath: link, - } - err := recorder.loadFromDisk() - if err == nil { - t.Fatal("expected symlink path to be rejected") - } - if !strings.Contains(err.Error(), "symlink") { - t.Fatalf("expected symlink error, got: %v", err) - } -} - -func TestLoadFromDiskRejectsOversizedFile(t *testing.T) { - dir := t.TempDir() - path := filepath.Join(dir, "incident_windows.json") - tooLarge := make([]byte, maxIncidentWindowsFileSize+1) - if err := os.WriteFile(path, tooLarge, 0o600); err != nil { - t.Fatalf("write oversized file failed: %v", err) - } - - recorder := &IncidentRecorder{ - config: DefaultIncidentRecorderConfig(), - dataDir: dir, - filePath: path, - } - err := recorder.loadFromDisk() - if err == nil { - t.Fatal("expected oversized file to be rejected") - } - if !strings.Contains(err.Error(), "exceeds size limit") { - t.Fatalf("expected size-limit error, got: %v", err) - } -} - -func TestSaveToDiskRejectsSymlinkDestination(t *testing.T) { - dir := t.TempDir() - target := filepath.Join(dir, "target.json") - if err := os.WriteFile(target, []byte("secret"), 0o600); err != nil { - t.Fatalf("write target failed: %v", err) - } - link := filepath.Join(dir, "incident_windows.json") - requireSymlinkOrSkip(t, target, link) - - recorder := &IncidentRecorder{ - config: DefaultIncidentRecorderConfig(), - dataDir: dir, - filePath: link, - completedWindows: []*IncidentWindow{ - {ID: "window-1", EndTime: ptrTime(time.Now())}, - }, - } - err := recorder.saveToDisk() - if err == nil { - t.Fatal("expected symlink destination to be rejected") - } - if !strings.Contains(err.Error(), "symlink") { - t.Fatalf("expected symlink error, got: %v", err) - } -} - -func ptrTime(v time.Time) *time.Time { - return &v -} - -func requireSymlinkOrSkip(t *testing.T, target, link string) { - t.Helper() - if err := os.Symlink(target, link); err != nil { - t.Skipf("symlink not supported in this environment: %v", err) - } -} diff --git a/internal/models/converters.go b/internal/models/converters.go index da5c5b882..607863401 100644 --- a/internal/models/converters.go +++ b/internal/models/converters.go @@ -1163,9 +1163,8 @@ type ResourceConvertInput struct { NetworkRX int64 NetworkTX int64 HasNetwork bool - DiskReadRate int64 - DiskWriteRate int64 - HasDiskIO bool + DiskReadRate *int64 + DiskWriteRate *int64 Temperature *float64 Uptime *int64 Tags []string @@ -1343,7 +1342,7 @@ func ConvertResourceToFrontend(input ResourceConvertInput) ResourceFrontend { } } - if input.HasDiskIO { + if input.DiskReadRate != nil || input.DiskWriteRate != nil { rf.DiskIO = &ResourceDiskIOFrontend{ ReadRate: input.DiskReadRate, WriteRate: input.DiskWriteRate, diff --git a/internal/models/converters_test.go b/internal/models/converters_test.go index 2fab91577..2c2f24aac 100644 --- a/internal/models/converters_test.go +++ b/internal/models/converters_test.go @@ -289,6 +289,7 @@ func TestVMToFrontend_NegativeNetworkValues(t *testing.T) { } func TestConvertResourceToFrontendIncludesDiskIO(t *testing.T) { + read, write := int64(4096), int64(8192) frontend := ConvertResourceToFrontend(ResourceConvertInput{ ID: "agent-1", Type: "agent", @@ -299,15 +300,14 @@ func TestConvertResourceToFrontendIncludesDiskIO(t *testing.T) { SourceType: "agent", Status: "online", LastSeenUnix: time.Now().UnixMilli(), - HasDiskIO: true, - DiskReadRate: 4096, - DiskWriteRate: 8192, + DiskReadRate: &read, + DiskWriteRate: &write, }) if frontend.DiskIO == nil { t.Fatal("expected disk I/O rates to be present") } - if frontend.DiskIO.ReadRate != 4096 || frontend.DiskIO.WriteRate != 8192 { + if frontend.DiskIO.ReadRate == nil || frontend.DiskIO.WriteRate == nil || *frontend.DiskIO.ReadRate != 4096 || *frontend.DiskIO.WriteRate != 8192 { t.Fatalf("unexpected disk I/O rates: %+v", frontend.DiskIO) } } diff --git a/internal/models/models_frontend.go b/internal/models/models_frontend.go index 5eecd6ab7..36ed0a480 100644 --- a/internal/models/models_frontend.go +++ b/internal/models/models_frontend.go @@ -1152,8 +1152,8 @@ type ResourceNetworkFrontend struct { // ResourceDiskIOFrontend represents aggregate disk I/O rates for the frontend. type ResourceDiskIOFrontend struct { - ReadRate int64 `json:"readRate"` - WriteRate int64 `json:"writeRate"` + ReadRate *int64 `json:"readRate,omitempty"` + WriteRate *int64 `json:"writeRate,omitempty"` } // ResourceAlertFrontend represents an alert on a resource. diff --git a/internal/monitoring/canonical_guardrails_test.go b/internal/monitoring/canonical_guardrails_test.go index 4398e9f8c..91cc20b9b 100644 --- a/internal/monitoring/canonical_guardrails_test.go +++ b/internal/monitoring/canonical_guardrails_test.go @@ -211,15 +211,15 @@ func TestProxmoxActionObserverUsesDirectControlPlaneClient(t *testing.T) { } func TestBroadcastResourceDiskIOUsesUnifiedResourceMetrics(t *testing.T) { - hasDiskIO, readRate, writeRate := monitorDiskIOMetricInput(&unifiedresources.ResourceMetrics{ + readRate, writeRate := monitorDiskIOMetricInput(&unifiedresources.ResourceMetrics{ DiskRead: &unifiedresources.MetricValue{Value: 4096.4, Unit: "bytes/s", Source: unifiedresources.SourceAgent}, DiskWrite: &unifiedresources.MetricValue{Value: 8191.6, Unit: "bytes/s", Source: unifiedresources.SourceAgent}, }) - if !hasDiskIO { + if readRate == nil || writeRate == nil { t.Fatal("expected disk I/O metrics to be projected") } - if readRate != 4096 || writeRate != 8192 { - t.Fatalf("unexpected projected disk I/O rates: read=%d write=%d", readRate, writeRate) + if *readRate != 4096 || *writeRate != 8192 { + t.Fatalf("unexpected projected disk I/O rates: read=%d write=%d", *readRate, *writeRate) } } diff --git a/internal/monitoring/docker_metric_presence_test.go b/internal/monitoring/docker_metric_presence_test.go new file mode 100644 index 000000000..717d745e3 --- /dev/null +++ b/internal/monitoring/docker_metric_presence_test.go @@ -0,0 +1,164 @@ +package monitoring + +import ( + "encoding/json" + "github.com/rcourtman/pulse-go-rewrite/internal/models" + "github.com/rcourtman/pulse-go-rewrite/internal/unifiedresources" + "testing" + "time" + + "github.com/rcourtman/pulse-go-rewrite/internal/mock" + agentsdocker "github.com/rcourtman/pulse-go-rewrite/pkg/agents/docker" + "github.com/rcourtman/pulse-go-rewrite/pkg/metrics" +) + +func TestApplyDockerReportPreservesDiskObservationPresence(t *testing.T) { + previous := mock.IsMockEnabled() + mustSetMockEnabled(t, false) + t.Cleanup(func() { mustSetMockEnabled(t, previous) }) + m := newTestMonitor(t) + store, err := metrics.NewStore(metrics.DefaultConfig(t.TempDir())) + if err != nil { + t.Fatal(err) + } + t.Cleanup(func() { store.Close() }) + m.metricsStore = store + start := time.Now().Add(-time.Minute) + report := agentsdocker.Report{ + Agent: agentsdocker.AgentInfo{ID: "presence-agent", Version: "6.4.2", IntervalSeconds: 30}, + Host: agentsdocker.HostInfo{Hostname: "presence-host"}, + Containers: []agentsdocker.Container{{ID: "presence-container", Name: "api", WritableLayerBytes: 200, RootFilesystemBytes: 1000}}, + } + observed := &agentsdocker.ContainerBlockIO{ReadBytes: 5000, WriteBytes: 7000} + steps := []struct { + name string + io *agentsdocker.ContainerBlockIO + wantSamples int + }{ + {"absent", nil, 0}, + {"first observation", observed, 0}, + {"measured idle", observed, 1}, + {"missing after observation", nil, 1}, + {"idle after gap", observed, 2}, + } + for i, step := range steps { + t.Run(step.name, func(t *testing.T) { + report.Timestamp = start.Add(time.Duration(i) * time.Second) + report.Containers[0].BlockIO = step.io + host, err := m.ApplyDockerReport(report, nil) + if err != nil { + t.Fatal(err) + } + if len(host.Containers) != 1 { + t.Fatalf("containers: %+v", host.Containers) + } + ct := host.Containers[0] + if step.io == nil && ct.BlockIO != nil { + t.Fatalf("absent BlockIO retained: %+v", ct.BlockIO) + } + if i == 1 && (ct.BlockIO.ReadRateBytesPerSecond != nil || ct.BlockIO.WriteRateBytesPerSecond != nil) { + t.Fatalf("warmup became a rate: %+v", ct.BlockIO) + } + if i == 2 || i == 4 { + if ct.BlockIO.ReadRateBytesPerSecond == nil || *ct.BlockIO.ReadRateBytesPerSecond != 0 || ct.BlockIO.WriteRateBytesPerSecond == nil || *ct.BlockIO.WriteRateBytesPerSecond != 0 { + t.Fatalf("unchanged counters must remain measured idle across gaps: %+v", ct.BlockIO) + } + } + store.Flush() + for _, metric := range []string{"diskread", "diskwrite"} { + points := m.metricsHistory.GetGuestMetrics("docker:"+ct.ID, metric, time.Hour) + if len(points) != step.wantSamples { + t.Fatalf("%s history = %d, want %d", metric, len(points), step.wantSamples) + } + persisted, err := store.Query("dockerContainer", ct.ID, metric, start, time.Now().Add(time.Minute), 0) + if err != nil { + t.Fatal(err) + } + if step.wantSamples == 0 && len(persisted) != 0 { + t.Fatalf("fabricated persisted %s: %+v", metric, persisted) + } + if step.wantSamples > 0 && (len(persisted) == 0 || persisted[len(persisted)-1].Value != 0) { + t.Fatalf("lost persisted idle %s: %+v", metric, persisted) + } + } + if points := m.metricsHistory.GetGuestMetrics("docker:"+ct.ID, "disk", time.Hour); len(points) != 0 { + t.Fatalf("layer sizes became capacity history: %+v", points) + } + points, err := store.Query("dockerContainer", ct.ID, "disk", start, time.Now().Add(time.Minute), 0) + if err != nil || len(points) != 0 { + t.Fatalf("layer sizes became persisted capacity: %+v, %v", points, err) + } + }) + } +} + +func TestApplyDockerReportPreservesIndependentZeroCounters(t *testing.T) { + previous := mock.IsMockEnabled() + mustSetMockEnabled(t, false) + t.Cleanup(func() { mustSetMockEnabled(t, previous) }) + m := newTestMonitor(t) + present, absent := true, false + report := agentsdocker.Report{ + Agent: agentsdocker.AgentInfo{ID: "zero-agent", Version: "6.4.2", IntervalSeconds: 30}, + Host: agentsdocker.HostInfo{Hostname: "zero-host"}, + Containers: []agentsdocker.Container{{ID: "zero-container", Name: "idle", BlockIO: &agentsdocker.ContainerBlockIO{ReadBytesPresent: &present, WriteBytesPresent: &absent}}}, + } + for i := 0; i < 2; i++ { + report.Timestamp = time.Now().Add(time.Duration(i) * time.Second) + host, err := m.ApplyDockerReport(report, nil) + if err != nil { + t.Fatal(err) + } + io := host.Containers[0].BlockIO + if io == nil { + t.Fatal("explicit zero counter lost") + } + if io.WriteRateBytesPerSecond != nil { + t.Fatalf("missing write direction became measured rate: %+v", io) + } + if i == 0 && io.ReadRateBytesPerSecond != nil { + t.Fatal("first zero counter became a measured rate") + } + if i == 1 && (io.ReadRateBytesPerSecond == nil || *io.ReadRateBytesPerSecond != 0) { + t.Fatalf("measured zero lost: %+v", io) + } + } + if points := m.metricsHistory.GetGuestMetrics("docker:zero-container", "diskwrite", time.Hour); len(points) != 0 { + t.Fatalf("missing writes recorded: %+v", points) + } + if points := m.metricsHistory.GetGuestMetrics("docker:zero-container", "diskread", time.Hour); len(points) != 1 { + t.Fatalf("idle reading lost: %+v", points) + } +} + +func TestResourceDiskIOWirePreservesAbsentDirection(t *testing.T) { + zero := &unifiedresources.MetricValue{Value: 0, Unit: "bytes/s", Source: unifiedresources.SourceDocker} + for _, tc := range []struct { + name string + metrics *unifiedresources.ResourceMetrics + wire string + }{ + {"absent", nil, ``}, + {"read idle", &unifiedresources.ResourceMetrics{DiskRead: zero}, `{"readRate":0}`}, + {"write idle", &unifiedresources.ResourceMetrics{DiskWrite: zero}, `{"writeRate":0}`}, + {"both idle", &unifiedresources.ResourceMetrics{DiskRead: zero, DiskWrite: zero}, `{"readRate":0,"writeRate":0}`}, + } { + t.Run(tc.name, func(t *testing.T) { + read, write := monitorDiskIOMetricInput(tc.metrics) + frontend := models.ConvertResourceToFrontend(models.ResourceConvertInput{DiskReadRate: read, DiskWriteRate: write}) + if tc.wire == "" { + if frontend.DiskIO != nil { + t.Fatalf("absent IO became a payload: %+v", frontend.DiskIO) + } + return + } + wire, err := json.Marshal(frontend.DiskIO) + if err != nil { + t.Fatal(err) + } + if string(wire) != tc.wire { + t.Fatalf("wire = %s, want %s", wire, tc.wire) + } + }) + } +} diff --git a/internal/monitoring/issue1613_contract_test.go b/internal/monitoring/issue1613_contract_test.go index 01e69fed3..7211b51d0 100644 --- a/internal/monitoring/issue1613_contract_test.go +++ b/internal/monitoring/issue1613_contract_test.go @@ -138,7 +138,7 @@ func TestIssue1613NodeDoesNotGreyBetweenNinetySecondPolls(t *testing.T) { } } -func TestIssue1613WebsocketStateKeepsUnknownRatesNumeric(t *testing.T) { +func TestIssue1613WebsocketStateKeepsObservedZeroAndOmitsUnknownRates(t *testing.T) { monitor := &Monitor{ state: models.NewState(), resourceStore: &resourceOnlyStore{resources: []unifiedresources.Resource{ @@ -168,7 +168,7 @@ func TestIssue1613WebsocketStateKeepsUnknownRatesNumeric(t *testing.T) { t.Fatal(err) } wire := string(payload) - if !strings.Contains(wire, `"diskIO":{"readRate":0,"writeRate":0}`) { + if !strings.Contains(wire, `"diskIO":{"readRate":0}`) { t.Fatalf("websocket payload does not contain numeric valid zero disk rate: %s", wire) } if strings.Count(wire, `"diskIO"`) != 1 { diff --git a/internal/monitoring/monitor.go b/internal/monitoring/monitor.go index 31d60322f..486079b4d 100644 --- a/internal/monitoring/monitor.go +++ b/internal/monitoring/monitor.go @@ -6182,10 +6182,7 @@ func monitorResourceToConvertInput(resource unifiedresources.Resource) models.Re input.HasNetwork = hasNetwork input.NetworkRX = rx input.NetworkTX = tx - hasDiskIO, diskRead, diskWrite := monitorDiskIOMetricInput(resource.Metrics) - input.HasDiskIO = hasDiskIO - input.DiskReadRate = diskRead - input.DiskWriteRate = diskWrite + input.DiskReadRate, input.DiskWriteRate = monitorDiskIOMetricInput(resource.Metrics) return input } @@ -6616,20 +6613,20 @@ func monitorNetworkMetricInput(metrics *unifiedresources.ResourceMetrics) (bool, return true, rx, tx } -func monitorDiskIOMetricInput(metrics *unifiedresources.ResourceMetrics) (bool, int64, int64) { +func monitorDiskIOMetricInput(metrics *unifiedresources.ResourceMetrics) (*int64, *int64) { if metrics == nil || (metrics.DiskRead == nil && metrics.DiskWrite == nil) { - return false, 0, 0 + return nil, nil } - - var read int64 - var write int64 + var read, write *int64 if metrics.DiskRead != nil { - read = int64(math.Round(metrics.DiskRead.Value)) + value := int64(math.Round(metrics.DiskRead.Value)) + read = &value } if metrics.DiskWrite != nil { - write = int64(math.Round(metrics.DiskWrite.Value)) + value := int64(math.Round(metrics.DiskWrite.Value)) + write = &value } - return true, read, write + return read, write } func monitorTemperature(resource unifiedresources.Resource) *float64 { diff --git a/internal/monitoring/monitor_agents.go b/internal/monitoring/monitor_agents.go index 0d5b8809b..ae7f4e043 100644 --- a/internal/monitoring/monitor_agents.go +++ b/internal/monitoring/monitor_agents.go @@ -2354,10 +2354,15 @@ func (m *Monitor) ApplyDockerReport(report agentsdocker.Report, tokenRecord *con containerIdentifier = payload.Name } if strings.TrimSpace(containerIdentifier) != "" { + readPresent, writePresent := payload.BlockIO.CounterPresence() metrics := models.IOMetrics{ NetworkIn: clampToInt64(payload.NetworkRXBytes), NetworkOut: clampToInt64(payload.NetworkTXBytes), Timestamp: receivedAt, + Presence: models.IOCounterPresence{ + Explicit: true, DiskRead: readPresent, DiskWrite: writePresent, + NetworkIn: true, NetworkOut: true, + }, } if payload.BlockIO != nil { metrics.DiskRead = clampToInt64(payload.BlockIO.ReadBytes) @@ -2640,16 +2645,9 @@ func (m *Monitor) ApplyDockerReport(report agentsdocker.Report, tokenRecord *con } metricKey := fmt.Sprintf("docker:%s", container.ID) - var diskPercent float64 - if container.RootFilesystemBytes > 0 && container.WritableLayerBytes > 0 { - diskPercent = float64(container.WritableLayerBytes) / float64(container.RootFilesystemBytes) * 100 - if diskPercent > 100 { - diskPercent = 100 - } - } - - var diskReadRate float64 - var diskWriteRate float64 + // Layer sizes describe container images, not filesystem capacity. + // Missing rate observations must not become measured idle samples. + diskReadRate, diskWriteRate := -1.0, -1.0 if container.BlockIO != nil { if container.BlockIO.ReadRateBytesPerSecond != nil { diskReadRate = *container.BlockIO.ReadRateBytesPerSecond @@ -2662,7 +2660,6 @@ func (m *Monitor) ApplyDockerReport(report agentsdocker.Report, tokenRecord *con if m.metricsHistory != nil { m.metricsHistory.AddGuestMetric(metricKey, "cpu", models.DockerContainerCPUCapacityPercent(container, host.CPUs), now) m.metricsHistory.AddGuestMetric(metricKey, "memory", container.MemoryPercent, now) - m.metricsHistory.AddGuestMetric(metricKey, "disk", diskPercent, now) if container.NetInRate >= 0 { m.metricsHistory.AddGuestMetric(metricKey, "netin", container.NetInRate, now) } @@ -2680,7 +2677,6 @@ func (m *Monitor) ApplyDockerReport(report agentsdocker.Report, tokenRecord *con if m.metricsStore != nil { m.metricsStore.Write("dockerContainer", container.ID, "cpu", models.DockerContainerCPUCapacityPercent(container, host.CPUs), now) m.metricsStore.Write("dockerContainer", container.ID, "memory", container.MemoryPercent, now) - m.metricsStore.Write("dockerContainer", container.ID, "disk", diskPercent, now) if container.NetInRate >= 0 { m.metricsStore.Write("dockerContainer", container.ID, "netin", container.NetInRate, now) } diff --git a/internal/monitoring/monitor_alert_handling_test.go b/internal/monitoring/monitor_alert_handling_test.go index b357d668b..945a5c3ea 100644 --- a/internal/monitoring/monitor_alert_handling_test.go +++ b/internal/monitoring/monitor_alert_handling_test.go @@ -5,6 +5,7 @@ import ( "io" "net/http" "net/http/httptest" + "strings" "sync/atomic" "testing" "time" @@ -16,8 +17,78 @@ import ( "github.com/rcourtman/pulse-go-rewrite/internal/notifications" unifiedresources "github.com/rcourtman/pulse-go-rewrite/internal/unifiedresources" "github.com/rcourtman/pulse-go-rewrite/internal/websocket" + "github.com/stretchr/testify/require" ) +func TestDockerAlertTimelineUsesCanonicalHistoryIdentity(t *testing.T) { + dir := t.TempDir() + store, err := unifiedresources.NewSQLiteResourceStore(dir, "default") + require.NoError(t, err) + t.Cleanup(func() { require.NoError(t, store.Close()) }) + manager := alerts.NewManagerWithDataDir(t.TempDir()) + t.Cleanup(manager.Stop) + config := manager.GetConfig() + config.Enabled = true + config.ActivationState = alerts.ActivationPending + config.TimeThresholds = map[string]int{} + config.SuppressionWindow = 0 + manager.UpdateConfig(config) + containerID := strings.Repeat("f", 64) + host := models.DockerHost{ID: "history-host", Hostname: "history-host", LastSeen: time.Now(), Containers: []models.DockerContainer{ + {ID: containerID, Name: "worker", State: "running", Health: "unhealthy"}, + {ID: strings.Repeat("a", 64), Name: "worker", State: "running", Health: "healthy"}, + }} + registry := unifiedresources.NewRegistry(store) + registry.IngestSnapshot(models.StateSnapshot{DockerHosts: []models.DockerHost{host}}) + adapter := unifiedresources.NewMonitorAdapter(registry) + monitor := &Monitor{alertManager: manager, resourceStore: adapter} + manager.SubscribeLifecycleCallback(monitor.handleAlertLifecycleEvent) + manager.CheckDockerHost(host) + canonicalID := unifiedresources.SourceSpecificID(unifiedresources.ResourceTypeAppContainer, unifiedresources.SourceDocker, host.ID+"/container/"+containerID) + filters := unifiedresources.ResourceChangeFilters{Kinds: []unifiedresources.ChangeKind{unifiedresources.ChangeAlertFired, unifiedresources.ChangeAlertResolved}} + changes, err := store.GetRecentChangesFiltered(canonicalID, time.Time{}, 10, filters) + require.NoError(t, err) + require.Len(t, changes, 1) + require.Equal(t, unifiedresources.ChangeAlertFired, changes[0].Kind) + require.Equal(t, canonicalID, changes[0].ResourceID) + var fired alerts.Alert + for _, alert := range manager.GetActiveAlerts() { + if alert.Type == "docker-container-health" { + fired = alert + } + } + require.NotEmpty(t, fired.ID) + // Recovery is emitted after the monitored container has left the registry. + // Its exact retained source binding must still select the original resource. + monitor.resourceStore = unifiedresources.NewMonitorAdapter(unifiedresources.NewRegistry(store)) + host.Containers[0].Health = "healthy" + manager.CheckDockerHost(host) + changes, err = store.GetRecentChangesFiltered(canonicalID, time.Time{}, 10, filters) + require.NoError(t, err) + require.Len(t, changes, 2) + require.Equal(t, unifiedresources.ChangeAlertResolved, changes[0].Kind) + require.Equal(t, canonicalID, changes[0].ResourceID) + monitor.recordAlertTimelineChange(&fired, unifiedresources.ChangeAlertFired, fired.StartTime, "") + monitor.recordAlertTimelineChange(&fired, unifiedresources.ChangeAlertResolved, *changes[0].OccurredAt, "") + again, err := store.GetRecentChangesFiltered(canonicalID, time.Time{}, 10, filters) + require.NoError(t, err) + require.Equal(t, changes, again) + controlID := unifiedresources.SourceSpecificID(unifiedresources.ResourceTypeAppContainer, unifiedresources.SourceDocker, host.ID+"/container/"+host.Containers[1].ID) + control, err := store.GetRecentChangesFiltered(controlID, time.Time{}, 10, filters) + require.NoError(t, err) + require.Empty(t, control) + require.NoError(t, store.Close()) + restarted, err := unifiedresources.NewSQLiteResourceStore(dir, "default") + require.NoError(t, err) + defer restarted.Close() + afterRestart, err := restarted.GetRecentChangesFiltered(canonicalID, time.Time{}, 10, filters) + require.NoError(t, err) + require.Equal(t, changes, afterRestart) + encoded, err := json.Marshal(afterRestart) + require.NoError(t, err) + t.Logf("DOCKER_HISTORY_LIFECYCLE %s", encoded) +} + func TestMonitor_HandleAlertFired_Extra(t *testing.T) { // 1. Alert is nil m1 := &Monitor{} diff --git a/internal/monitoring/monitor_unified_state_test.go b/internal/monitoring/monitor_unified_state_test.go index 9011a8e87..e920fec1e 100644 --- a/internal/monitoring/monitor_unified_state_test.go +++ b/internal/monitoring/monitor_unified_state_test.go @@ -165,7 +165,7 @@ func TestConvertResourcesForBroadcastCoalescesSplitHostResources(t *testing.T) { if resource.DiskIO == nil { t.Fatal("expected aggregate disk I/O rates in broadcast resource") } - if resource.DiskIO.ReadRate != 4096 || resource.DiskIO.WriteRate != 8192 { + if resource.DiskIO.ReadRate == nil || resource.DiskIO.WriteRate == nil || *resource.DiskIO.ReadRate != 4096 || *resource.DiskIO.WriteRate != 8192 { t.Fatalf("unexpected aggregate disk I/O rates: %+v", resource.DiskIO) } } diff --git a/internal/unifiedresources/code_standards_test.go b/internal/unifiedresources/code_standards_test.go index dd48d807e..f9c85fa78 100644 --- a/internal/unifiedresources/code_standards_test.go +++ b/internal/unifiedresources/code_standards_test.go @@ -2515,13 +2515,6 @@ func TestV6DirectHostAliasValidatorCoverage(t *testing.T) { `[]string{"host", "guest", "docker", "container", "lxc", "qemu", "docker_container", "docker_service"}`, }, }, - { - path: filepath.Join(repoRoot, "internal", "ai", "incident_coordinator_additional_test.go"), - requiredSnippets: []string{ - `TestIncidentCoordinator_OnAnomalyDetected_CanonicalizesLegacyHostAlias`, - `expected anomaly recording resource type to be canonicalized to agent`, - }, - }, { path: filepath.Join(repoRoot, "internal", "ai", "tools", "tools_metrics_alerts_test.go"), requiredSnippets: []string{ @@ -2557,13 +2550,6 @@ func TestV6DirectHostAliasValidatorCoverage(t *testing.T) { `legacy k8s alias rejected`, }, }, - { - path: filepath.Join(repoRoot, "internal", "metrics", "incident_recorder_test.go"), - requiredSnippets: []string{ - `TestStartRecordingCanonicalizesLegacyHostAlias`, - `expected legacy host alias to canonicalize to agent`, - }, - }, { path: filepath.Join(repoRoot, "internal", "api", "resourceapi", "resources_test.go"), requiredSnippets: []string{ diff --git a/internal/unifiedresources/history_identity.go b/internal/unifiedresources/history_identity.go new file mode 100644 index 000000000..e71c71a2c --- /dev/null +++ b/internal/unifiedresources/history_identity.go @@ -0,0 +1,161 @@ +package unifiedresources + +import ( + "database/sql" + "encoding/hex" + "fmt" + "strings" +) + +// legacyDockerHistoryIdentity accepts the source identifier emitted by Docker +// alerts only when it contains a complete container ID. Names and short IDs +// cannot establish durable identity after inventory removal. +func legacyDockerHistoryIdentity(ref string) (sourceID, canonicalID string, ok bool) { + ref = strings.TrimSpace(ref) + if !strings.HasPrefix(ref, "docker:") { + return "", "", false + } + host, container, found := strings.Cut(strings.TrimPrefix(ref, "docker:"), "/") + if !found || host == "" || strings.TrimSpace(host) != host || len(container) != 64 || strings.ToLower(container) != container { + return "", "", false + } + if _, err := hex.DecodeString(container); err != nil { + return "", "", false + } + sourceID = host + "/container/" + container + if host == "container" { // DockerResourceID's explicit hostless form. + sourceID = container + } + return sourceID, SourceSpecificID(ResourceTypeAppContainer, SourceDocker, sourceID), true +} + +// resourceHistoryIdentityWriter binds a source reference to the canonical +// resource of one event. It affects history lookup only, never operator state, +// action requests, approvals or execution identities. +type resourceHistoryIdentityWriter interface { + RecordChangeWithSourceIdentity(change ResourceChange, sourceID string) error + ResolveHistorySourceIdentity(sourceID string) (string, bool, error) +} + +func (s *SQLiteResourceStore) ResolveHistorySourceIdentity(sourceID string) (string, bool, error) { + var id string + err := s.db.QueryRow(`SELECT canonical_id FROM resource_history_aliases WHERE source_id = ?`, sourceID).Scan(&id) + if err == sql.ErrNoRows { + return "", false, nil + } + return id, err == nil, err +} + +func (m *MemoryStore) ResolveHistorySourceIdentity(sourceID string) (string, bool, error) { + m.mu.RLock() + defer m.mu.RUnlock() + id, ok := m.historyAliases[sourceID] + return id, ok, nil +} + +func (s *SQLiteResourceStore) RecordChangeWithSourceIdentity(change ResourceChange, sourceID string) error { + sourceID = CanonicalResourceID(sourceID) + canonicalID := CanonicalResourceID(change.ResourceID) + if sourceID == "" || canonicalID == "" || sourceID == canonicalID { + return s.RecordChange(change) + } + s.mu.Lock() + defer s.mu.Unlock() + tx, err := s.db.Begin() + if err != nil { + return fmt.Errorf("begin resource history identity: %w", err) + } + defer tx.Rollback() + // Registry resolution is authoritative when available. Rebinding a source + // reference affects subsequent history reads without rewriting past events. + if _, err := tx.Exec(`INSERT INTO resource_history_aliases (source_id, canonical_id) VALUES (?, ?) + ON CONFLICT(source_id) DO UPDATE SET canonical_id = excluded.canonical_id`, sourceID, canonicalID); err != nil { + return fmt.Errorf("record resource history identity: %w", err) + } + if err := recordChangeSQL(tx, change, s.resourceChangesHasTimestamp); err != nil { + return err + } + return tx.Commit() +} + +func (m *MemoryStore) RecordChangeWithSourceIdentity(change ResourceChange, sourceID string) error { + m.mu.Lock() + defer m.mu.Unlock() + if m.historyAliases == nil { + m.historyAliases = make(map[string]string) + } + if sourceID = CanonicalResourceID(sourceID); sourceID != "" && sourceID != change.ResourceID { + m.historyAliases[sourceID] = change.ResourceID + } + return m.recordChangeLocked(change) +} + +// migrateResourceHistoryAliases adds an index for exact legacy Docker event +// identities, including resources already removed from live inventory. The +// event rows and all authority-bearing tables remain unchanged. +func (s *SQLiteResourceStore) migrateResourceHistoryAliases() error { + if _, err := s.db.Exec(`CREATE TABLE IF NOT EXISTS resource_history_aliases ( + source_id TEXT PRIMARY KEY, canonical_id TEXT NOT NULL); + CREATE INDEX IF NOT EXISTS idx_resource_history_aliases_canonical ON resource_history_aliases(canonical_id);`); err != nil { + return fmt.Errorf("initialize resource history identities: %w", err) + } + rows, err := s.db.Query(`SELECT DISTINCT canonical_id FROM resource_changes WHERE canonical_id GLOB 'docker:*'`) + if err != nil { + return fmt.Errorf("read legacy resource history identities: %w", err) + } + aliases := make(map[string]string) + for rows.Next() { + var sourceID string + if err := rows.Scan(&sourceID); err != nil { + rows.Close() + return err + } + if _, canonicalID, ok := legacyDockerHistoryIdentity(sourceID); ok { + aliases[sourceID] = canonicalID + } + } + readErr := rows.Err() + rows.Close() // The store has one connection. Release it before writing. + if readErr != nil { + return readErr + } + for sourceID, canonicalID := range aliases { + if _, err := s.db.Exec(`INSERT OR IGNORE INTO resource_history_aliases (source_id, canonical_id) VALUES (?, ?)`, sourceID, canonicalID); err != nil { + return fmt.Errorf("index legacy resource history identity: %w", err) + } + } + return nil +} + +// expandHistoryAliases reads persisted identity bindings each time so separate +// monitor, API and Assistant store handles see new bindings immediately. The +// indexed traversal is restricted to the requested identities, not the fleet. +func (s *SQLiteResourceStore) expandHistoryAliases(ids []string) ([]string, error) { + if len(ids) == 0 { + return ids, nil + } + seeds := make([]string, len(ids)) + args := make([]any, len(ids)) + for i, id := range ids { + seeds[i], args[i] = "(?)", id + } + rows, err := s.db.Query(`WITH RECURSIVE history_ids(id) AS ( + VALUES `+strings.Join(seeds, ",")+` + UNION SELECT a.canonical_id FROM resource_history_aliases a JOIN history_ids h ON a.source_id = h.id + UNION SELECT a.source_id FROM resource_history_aliases a JOIN history_ids h ON a.canonical_id = h.id + UNION SELECT s.old_canonical_id FROM canonical_id_successions s JOIN history_ids h ON s.new_canonical_id = h.id + ) SELECT id FROM history_ids`, args...) + if err != nil { + return nil, fmt.Errorf("read resource history identities: %w", err) + } + defer rows.Close() + var expanded []string + for rows.Next() { + var id string + if err := rows.Scan(&id); err != nil { + return nil, err + } + expanded = append(expanded, id) + } + return expanded, rows.Err() +} diff --git a/internal/unifiedresources/history_identity_test.go b/internal/unifiedresources/history_identity_test.go new file mode 100644 index 000000000..21c4512db --- /dev/null +++ b/internal/unifiedresources/history_identity_test.go @@ -0,0 +1,175 @@ +package unifiedresources + +import ( + "fmt" + "strings" + "testing" + "time" + + "github.com/rcourtman/pulse-go-rewrite/internal/models" + "github.com/stretchr/testify/require" +) + +func TestHistoryIdentityLegacyDockerReference(t *testing.T) { + container := strings.Repeat("a", 64) + for _, ref := range []string{"docker:host/" + container, "docker:container/" + container} { + source, id, ok := legacyDockerHistoryIdentity(ref) + require.True(t, ok) + require.Equal(t, SourceSpecificID(ResourceTypeAppContainer, SourceDocker, source), id) + } + for _, ref := range []string{"docker:host/worker", "docker:host/" + container[:12], "docker:host/" + strings.ToUpper(container), "docker:/" + container, "docker:host/" + strings.Repeat("z", 64), "docker:host", "vm:host/" + container} { + _, _, ok := legacyDockerHistoryIdentity(ref) + require.False(t, ok, ref) + } +} + +// Exercise the actual scoped history read as the unrelated identity index grows. +// Timing is reported for qualification, without a machine-dependent pass threshold. +func BenchmarkHistoryIdentityQuery(b *testing.B) { + for _, size := range []int{1, 20000} { + b.Run(fmt.Sprintf("aliases-%d", size), func(b *testing.B) { + store, err := NewSQLiteResourceStore(b.TempDir(), "benchmark") + require.NoError(b, err) + b.Cleanup(func() { require.NoError(b, store.Close()) }) + tx, err := store.db.Begin() + require.NoError(b, err) + for i := 0; i < size; i++ { + _, err := tx.Exec(`INSERT INTO resource_history_aliases (source_id, canonical_id) VALUES (?, ?)`, fmt.Sprintf("legacy-%d", i), fmt.Sprintf("app-container-%d", i)) + require.NoError(b, err) + require.NoError(b, recordChangeSQL(tx, ResourceChange{ID: fmt.Sprintf("event-%d", i), ResourceID: fmt.Sprintf("legacy-%d", i), ObservedAt: time.Now(), Kind: ChangeAlertFired}, store.resourceChangesHasTimestamp)) + } + require.NoError(b, tx.Commit()) + b.ResetTimer() + for i := 0; i < b.N; i++ { + got, err := store.GetRecentChanges("app-container-0", time.Time{}, 50) + if err != nil || len(got) != 1 { + b.Fatalf("scoped history: count=%d err=%v", len(got), err) + } + } + }) + } +} + +func TestHistoryIdentityMigrationPreservesEventsAndAuthority(t *testing.T) { + dir := t.TempDir() + store, err := NewSQLiteResourceStore(dir, "default") + require.NoError(t, err) + t.Cleanup(func() { require.NoError(t, store.Close()) }) + legacy := "docker:tower/" + strings.Repeat("b", 64) + _, canonical, _ := legacyDockerHistoryIdentity(legacy) + now := time.Now().UTC().Truncate(time.Second) + event := ResourceChange{ID: "legacy-fired", ResourceID: legacy, ObservedAt: now, Kind: ChangeAlertFired, SourceType: SourcePulseDiff, Reason: "container unhealthy", Metadata: map[string]any{"alert_id": "health-test"}} + require.NoError(t, store.RecordChange(event)) + require.NoError(t, store.SetResourceOperatorState(ResourceOperatorState{CanonicalID: legacy, NeverAutoRemediate: true, Note: "keep authority binding"})) + _, err = store.db.Exec(`INSERT INTO action_audits (id, action_id, canonical_id, request_id, created_at, updated_at, state, request_json, plan_json) + VALUES ('history-action', 'history-action', ?, 'request-1', ?, ?, 'pending', '{"binding":"original"}', '{}')`, legacy, now, now) + require.NoError(t, err) + require.NoError(t, store.Close()) + // No inventory survives this restart. The legacy full ID still identifies + // the same container, and the original event is never rewritten. + store, err = NewSQLiteResourceStore(dir, "default") + require.NoError(t, err) + for _, id := range []string{legacy, canonical} { + got, err := store.GetRecentChanges(id, now.Add(-time.Minute), 10) + require.NoError(t, err) + require.Equal(t, []ResourceChange{event}, got) + count, err := store.CountRecentChanges(id, now.Add(-time.Minute)) + require.NoError(t, err) + require.Equal(t, 1, count) + kinds, err := store.CountRecentChangesByKind(id, now.Add(-time.Minute)) + require.NoError(t, err) + require.Equal(t, 1, kinds[ChangeAlertFired]) + } + state, found, err := store.GetResourceOperatorState(legacy) + require.NoError(t, err) + require.True(t, found) + require.True(t, state.NeverAutoRemediate) + require.Equal(t, "keep authority binding", state.Note) + _, found, err = store.GetResourceOperatorState(canonical) + require.NoError(t, err) + require.False(t, found) + var actionID, request, eventID string + require.NoError(t, store.db.QueryRow(`SELECT canonical_id, request_json FROM action_audits WHERE id = 'history-action'`).Scan(&actionID, &request)) + require.Equal(t, legacy, actionID) + require.Equal(t, `{"binding":"original"}`, request) + require.NoError(t, store.db.QueryRow(`SELECT canonical_id FROM resource_changes WHERE id = 'legacy-fired'`).Scan(&eventID)) + require.Equal(t, legacy, eventID) +} + +func TestHistoryIdentitySeparateHandlesSeeBindingAndReplay(t *testing.T) { + dir := t.TempDir() + writer, err := NewSQLiteResourceStore(dir, "default") + require.NoError(t, err) + defer writer.Close() + reader, err := NewSQLiteResourceStore(dir, "default") + require.NoError(t, err) + defer reader.Close() + legacy := "docker:tower/" + strings.Repeat("c", 64) + _, canonical, _ := legacyDockerHistoryIdentity(legacy) + now := time.Now().UTC().Truncate(time.Second) + old := ResourceChange{ID: "first", ResourceID: legacy, ObservedAt: now, Kind: ChangeAlertFired, SourceType: SourcePulseDiff} + require.NoError(t, writer.RecordChange(old)) + got, err := reader.GetRecentChanges(canonical, time.Time{}, 10) + require.NoError(t, err) + require.Empty(t, got) + replayed := old + replayed.ResourceID = canonical + require.NoError(t, writer.RecordChangeWithSourceIdentity(replayed, legacy)) + second := ResourceChange{ID: "second", ResourceID: canonical, ObservedAt: now.Add(time.Second), Kind: ChangeAlertResolved, SourceType: SourcePulseDiff} + require.NoError(t, writer.RecordChangeWithSourceIdentity(second, legacy)) + require.NoError(t, writer.RecordChangeWithSourceIdentity(second, legacy)) + for _, id := range []string{legacy, canonical} { + got, err := reader.GetRecentChanges(id, time.Time{}, 10) + require.NoError(t, err) + require.Equal(t, []ResourceChange{second, old}, got) + } + // The same source identifier in a different organization cannot see this binding. + other, err := NewSQLiteResourceStore(dir, "other-org") + require.NoError(t, err) + defer other.Close() + got, err = other.GetRecentChanges(canonical, time.Time{}, 10) + require.NoError(t, err) + require.Empty(t, got) +} + +func TestHistoryIdentityMonitorAdapterUsesExactContainerIdentity(t *testing.T) { + store := NewMemoryStore() + container := strings.Repeat("d", 64) + host := models.DockerHost{ID: "tower", Hostname: "tower", LastSeen: time.Now(), Containers: []models.DockerContainer{{ID: container, Name: "worker", State: "running"}}} + registry := NewRegistry(store) + registry.IngestSnapshot(models.StateSnapshot{DockerHosts: []models.DockerHost{host}}) + adapter := NewMonitorAdapter(registry) + legacy := "docker:tower/" + container + _, canonical, _ := legacyDockerHistoryIdentity(legacy) + for i, ref := range []string{legacy, "docker:tower/worker", "docker:tower/" + container[:12]} { + require.NoError(t, adapter.RecordChange(ResourceChange{ID: ref, ResourceID: ref, Kind: ChangeAlertFired, ObservedAt: time.Now().Add(time.Duration(i) * time.Second)})) + } + got, err := store.GetRecentChanges(canonical, time.Time{}, 10) + require.NoError(t, err) + require.Len(t, got, 1) + require.Equal(t, canonical, got[0].ResourceID) + // A retained authoritative binding survives loss of the registry. + require.NoError(t, store.RecordChangeWithSourceIdentity(ResourceChange{ID: "binding", ResourceID: "app-container-retained", ObservedAt: time.Now()}, legacy)) + removed := NewMonitorAdapter(NewRegistry(store)) + require.NoError(t, removed.RecordChange(ResourceChange{ID: "after-removal", ResourceID: legacy, Kind: ChangeAlertResolved, ObservedAt: time.Now()})) + got, err = store.GetRecentChanges("app-container-retained", time.Time{}, 10) + require.NoError(t, err) + require.Len(t, got, 2) + require.Equal(t, "app-container-retained", got[0].ResourceID) +} + +func TestHistoryIdentityRetentionAndUnavailableLookup(t *testing.T) { + store, err := NewSQLiteResourceStore(t.TempDir(), "default") + require.NoError(t, err) + t.Cleanup(func() { require.NoError(t, store.Close()) }) + legacy := "docker:tower/" + strings.Repeat("e", 64) + _, canonical, _ := legacyDockerHistoryIdentity(legacy) + require.NoError(t, store.RecordChangeWithSourceIdentity(ResourceChange{ID: "expired", ResourceID: canonical, ObservedAt: time.Now().Add(-2 * resourceChangesRetention)}, legacy)) + store.pruneOldRecords() + _, found, err := store.ResolveHistorySourceIdentity(legacy) + require.NoError(t, err) + require.False(t, found) + require.NoError(t, store.Close()) + _, err = store.GetRecentChanges(canonical, time.Time{}, 10) + require.Error(t, err) +} diff --git a/internal/unifiedresources/metrics.go b/internal/unifiedresources/metrics.go index 6c50f57b1..399060c00 100644 --- a/internal/unifiedresources/metrics.go +++ b/internal/unifiedresources/metrics.go @@ -270,17 +270,8 @@ func metricsFromDockerContainer(ct models.DockerContainer, hostCPUs ...int) *Res percent := percentFromReportedPercent(ct.MemoryPercent) metrics.Memory = &MetricValue{Used: &ct.MemoryUsage, Total: &ct.MemoryLimit, Percent: percent, Unit: "bytes", Source: SourceDocker} } - if ct.RootFilesystemBytes > 0 { - used := ct.WritableLayerBytes - if used < 0 { - used = 0 - } - if used > ct.RootFilesystemBytes { - used = ct.RootFilesystemBytes - } - percent := clampMetricValue((float64(used)/float64(ct.RootFilesystemBytes))*100, 0, 100) - metrics.Disk = &MetricValue{Used: &used, Total: &ct.RootFilesystemBytes, Percent: percent, Unit: "bytes", Source: SourceDocker} - } + // Writable and root layer sizes are image metadata, not used/total + // filesystem capacity. Docker does not supply a capacity observation here. if ct.NetInRate > 0 { metrics.NetIn = &MetricValue{Value: ct.NetInRate, Unit: "bytes/s", Source: SourceDocker} } @@ -288,10 +279,10 @@ func metricsFromDockerContainer(ct models.DockerContainer, hostCPUs ...int) *Res metrics.NetOut = &MetricValue{Value: ct.NetOutRate, Unit: "bytes/s", Source: SourceDocker} } if ct.BlockIO != nil { - if ct.BlockIO.ReadRateBytesPerSecond != nil && *ct.BlockIO.ReadRateBytesPerSecond > 0 { + if ct.BlockIO.ReadRateBytesPerSecond != nil && *ct.BlockIO.ReadRateBytesPerSecond >= 0 && !math.IsInf(*ct.BlockIO.ReadRateBytesPerSecond, 0) { metrics.DiskRead = &MetricValue{Value: *ct.BlockIO.ReadRateBytesPerSecond, Unit: "bytes/s", Source: SourceDocker} } - if ct.BlockIO.WriteRateBytesPerSecond != nil && *ct.BlockIO.WriteRateBytesPerSecond > 0 { + if ct.BlockIO.WriteRateBytesPerSecond != nil && *ct.BlockIO.WriteRateBytesPerSecond >= 0 && !math.IsInf(*ct.BlockIO.WriteRateBytesPerSecond, 0) { metrics.DiskWrite = &MetricValue{Value: *ct.BlockIO.WriteRateBytesPerSecond, Unit: "bytes/s", Source: SourceDocker} } } diff --git a/internal/unifiedresources/metrics_test.go b/internal/unifiedresources/metrics_test.go index d99e54af5..b1513c437 100644 --- a/internal/unifiedresources/metrics_test.go +++ b/internal/unifiedresources/metrics_test.go @@ -1,12 +1,34 @@ package unifiedresources import ( + "math" "testing" "time" "github.com/rcourtman/pulse-go-rewrite/internal/models" ) +func TestMetricsFromDockerContainerDistinguishesAbsentAndIdleIO(t *testing.T) { + for _, value := range []float64{0, 123, -1, math.NaN(), math.Inf(1)} { + ct := models.DockerContainer{BlockIO: &models.DockerContainerBlockIO{ReadRateBytesPerSecond: &value, WriteRateBytesPerSecond: &value}} + got := metricsFromDockerContainer(ct) + valid := value >= 0 && !math.IsInf(value, 0) + if valid { + if got.DiskRead == nil || got.DiskWrite == nil || got.DiskRead.Value != value || got.DiskWrite.Value != value { + t.Fatalf("lost measured rate %v: %+v", value, got) + } + } else if got.DiskRead != nil || got.DiskWrite != nil { + t.Fatalf("invalid rate %v became observation: %+v", value, got) + } + } + for _, io := range []*models.DockerContainerBlockIO{nil, {ReadBytes: 5000, WriteBytes: 7000}} { + got := metricsFromDockerContainer(models.DockerContainer{BlockIO: io}) + if got.DiskRead != nil || got.DiskWrite != nil { + t.Fatalf("absent rates became observations: %+v", got) + } + } +} + func TestMetricsFromDockerHostIncludesIORates(t *testing.T) { host := models.DockerHost{ CPUUsage: 12.5, @@ -447,8 +469,8 @@ func TestMetricsFromDockerContainerIncludesContainerIORates(t *testing.T) { if metrics.DiskWrite == nil || metrics.DiskWrite.Value != writeRate { t.Fatalf("expected diskWrite=%v, got %+v", writeRate, metrics.DiskWrite) } - if metrics.Disk == nil || metrics.Disk.Percent <= 0 { - t.Fatalf("expected non-zero disk usage metric, got %+v", metrics.Disk) + if metrics.Disk != nil { + t.Fatalf("container layer sizes are not filesystem capacity, got %+v", metrics.Disk) } } diff --git a/internal/unifiedresources/monitor_adapter.go b/internal/unifiedresources/monitor_adapter.go index 6bd9fc70e..edb0b0f4d 100644 --- a/internal/unifiedresources/monitor_adapter.go +++ b/internal/unifiedresources/monitor_adapter.go @@ -111,6 +111,33 @@ func (a *MonitorAdapter) RecordChange(change ResourceChange) error { if registry == nil || registry.store == nil { return nil } + sourceRef := change.ResourceID + if sourceID, derivedID, ok := legacyDockerHistoryIdentity(sourceRef); ok { + // Use exact source identity when inventory is present, including any + // canonical identity merge. Full Docker IDs remain derivable after removal. + registry.mu.RLock() + resolvedID := registry.bySource[SourceDocker][sourceID] + registry.mu.RUnlock() + if resolvedID != "" { + change.ResourceID = resolvedID + } else { + change.ResourceID = derivedID + if history, ok := registry.store.(resourceHistoryIdentityWriter); ok { + id, found, err := history.ResolveHistorySourceIdentity(sourceRef) + if err != nil { + return err + } + if found { + change.ResourceID = id + } + } + } + } + if sourceRef != change.ResourceID { + if writer, ok := registry.store.(resourceHistoryIdentityWriter); ok { + return writer.RecordChangeWithSourceIdentity(change, sourceRef) + } + } return registry.store.RecordChange(change) } diff --git a/internal/unifiedresources/store.go b/internal/unifiedresources/store.go index 4c9a9247c..bbbf6cc65 100644 --- a/internal/unifiedresources/store.go +++ b/internal/unifiedresources/store.go @@ -609,6 +609,9 @@ func (s *SQLiteResourceStore) initSchema() error { if err := s.ensureResourceChangesIndexes(); err != nil { return err } + if err := s.migrateResourceHistoryAliases(); err != nil { + return err + } if err := s.migrateResourceIdentitiesSchema(); err != nil { return err } @@ -1428,6 +1431,13 @@ func (s *SQLiteResourceStore) pruneOldRecords() { } else if affected > 0 { totalDeleted += affected } + // History-only aliases need not outlive all records for either identifier. + // This never removes canonical identity pins or authority-bearing state. + if _, err := s.db.Exec(`DELETE FROM resource_history_aliases + WHERE NOT EXISTS (SELECT 1 FROM resource_changes WHERE canonical_id = resource_history_aliases.source_id) + AND NOT EXISTS (SELECT 1 FROM resource_changes WHERE canonical_id = resource_history_aliases.canonical_id)`); err != nil { + log.Printf("unifiedresources: failed to prune resource history identities: %v", err) + } res, err = s.db.Exec( `DELETE FROM action_audits WHERE created_at < ?`, @@ -1614,10 +1624,10 @@ func (s *SQLiteResourceStore) queryResourceIdentityPins() ([]ResourceIdentityPin // history. Resources without pins (Proxmox guests, record-declared eras) // merge through the durable canonical_id_successions record instead. Unknown // IDs expand to themselves. -func (s *SQLiteResourceStore) resourceChangeIDSet(canonicalID string) []string { +func (s *SQLiteResourceStore) resourceChangeIDSet(canonicalID string) ([]string, error) { canonicalID = CanonicalResourceID(canonicalID) if canonicalID == "" { - return nil + return nil, nil } s.identityPinMu.Lock() @@ -1631,7 +1641,7 @@ func (s *SQLiteResourceStore) resourceChangeIDSet(canonicalID string) []string { pins := s.identityPinCache s.identityPinMu.Unlock() - return expandResourceChangeIDs(canonicalID, pins, s.successionMap()) + return s.expandHistoryAliases(expandResourceChangeIDs(canonicalID, pins, s.successionMap())) } func expandResourceChangeIDs(canonicalID string, pins []ResourceIdentityPin, successors map[string]string) []string { @@ -1757,7 +1767,11 @@ func (s *SQLiteResourceStore) GetRecentChangesFiltered(canonicalID string, since conditions := []string{} canonicalID = CanonicalResourceID(canonicalID) if canonicalID != "" { - conditions, args = appendRecentChangeResourceCondition(conditions, args, s.resourceChangeIDSet(canonicalID), filters.IncludeRelated) + ids, err := s.resourceChangeIDSet(canonicalID) + if err != nil { + return nil, err + } + conditions, args = appendRecentChangeResourceCondition(conditions, args, ids, filters.IncludeRelated) } else { conditions = append(conditions, observedAtExpr+" >= ?") args = append(args, since) @@ -1882,8 +1896,12 @@ func (s *SQLiteResourceStore) CountRecentChanges(canonicalID string, since time. } func (s *SQLiteResourceStore) CountRecentChangesFiltered(canonicalID string, since time.Time, filters ResourceChangeFilters) (int, error) { + ids, err := s.resourceChangeIDSet(canonicalID) + if err != nil { + return 0, err + } query, args := buildRecentChangeCountQuery( - s.resourceChangeIDSet(canonicalID), + ids, since, filters, "SELECT COUNT(*) FROM resource_changes", @@ -1907,8 +1925,12 @@ func (s *SQLiteResourceStore) CountRecentChangesByKind(canonicalID string, since } func (s *SQLiteResourceStore) CountRecentChangesByKindFiltered(canonicalID string, since time.Time, filters ResourceChangeFilters) (map[ChangeKind]int, error) { + ids, err := s.resourceChangeIDSet(canonicalID) + if err != nil { + return nil, err + } query, args := buildRecentChangeCountQuery( - s.resourceChangeIDSet(canonicalID), + ids, since, filters, "SELECT COALESCE(kind, ''), COUNT(*) FROM resource_changes", @@ -1950,9 +1972,13 @@ func (s *SQLiteResourceStore) CountRecentChangesBySourceType(canonicalID string, } func (s *SQLiteResourceStore) CountRecentChangesBySourceTypeFiltered(canonicalID string, since time.Time, filters ResourceChangeFilters) (map[ChangeSourceType]int, error) { + ids, err := s.resourceChangeIDSet(canonicalID) + if err != nil { + return nil, err + } sourceTypeExpr := s.resourceChangesSourceTypeExpr() query, args := buildRecentChangeCountQuery( - s.resourceChangeIDSet(canonicalID), + ids, since, filters, "SELECT "+sourceTypeExpr+", COUNT(*) FROM resource_changes", @@ -1994,9 +2020,13 @@ func (s *SQLiteResourceStore) CountRecentChangesBySourceAdapter(canonicalID stri } func (s *SQLiteResourceStore) CountRecentChangesBySourceAdapterFiltered(canonicalID string, since time.Time, filters ResourceChangeFilters) (map[ChangeSourceAdapter]int, error) { + ids, err := s.resourceChangeIDSet(canonicalID) + if err != nil { + return nil, err + } sourceAdapterExpr := s.resourceChangesSourceAdapterExpr() query, args := buildRecentChangeCountQuery( - s.resourceChangeIDSet(canonicalID), + ids, since, filters, "SELECT "+sourceAdapterExpr+", COUNT(*) FROM resource_changes", @@ -3352,6 +3382,7 @@ type MemoryStore struct { loopReports map[string]LoopReport identityPins map[string]ResourceIdentityPin canonicalSuccessions map[string]string + historyAliases map[string]string } func NewMemoryStore() *MemoryStore { @@ -3422,7 +3453,11 @@ func (m *MemoryStore) resourceChangeIDSetLocked(canonicalID string) []string { for _, pin := range m.identityPins { pins = append(pins, pin) } - return expandResourceChangeIDs(canonicalID, pins, m.canonicalSuccessions) + if resolved := m.historyAliases[canonicalID]; resolved != "" { + canonicalID = resolved + } + ids := expandResourceChangeIDs(canonicalID, pins, m.canonicalSuccessions) + return appendSupersededChangeIDs(ids, m.historyAliases) } func (m *MemoryStore) AddLink(link ResourceLink) error { diff --git a/pkg/agents/docker/blockio_presence_test.go b/pkg/agents/docker/blockio_presence_test.go new file mode 100644 index 000000000..0e507b911 --- /dev/null +++ b/pkg/agents/docker/blockio_presence_test.go @@ -0,0 +1,32 @@ +package dockeragent + +import ( + "encoding/json" + "testing" +) + +func TestContainerBlockIOCounterPresenceWireCompatibility(t *testing.T) { + for _, tc := range []struct { + wire string + read, write bool + }{ + {`null`, false, false}, + {`{}`, false, false}, + {`{"readBytes":5000}`, true, false}, + {`{"writeBytes":7000}`, false, true}, + {`{"readBytes":5000,"writeBytes":7000}`, true, true}, + {`{"readBytesPresent":true,"writeBytesPresent":true}`, true, true}, + {`{"readBytes":5000,"readBytesPresent":false,"writeBytesPresent":true}`, false, true}, + } { + t.Run(tc.wire, func(t *testing.T) { + var io *ContainerBlockIO + if err := json.Unmarshal([]byte(tc.wire), &io); err != nil { + t.Fatal(err) + } + read, write := io.CounterPresence() + if read != tc.read || write != tc.write { + t.Fatalf("presence = %v/%v, want %v/%v", read, write, tc.read, tc.write) + } + }) + } +} diff --git a/pkg/agents/docker/report.go b/pkg/agents/docker/report.go index 04459ba3e..e23eb8697 100644 --- a/pkg/agents/docker/report.go +++ b/pkg/agents/docker/report.go @@ -124,8 +124,27 @@ type ContainerNetwork struct { // ContainerBlockIO summarises high-level block I/O metrics for a container. type ContainerBlockIO struct { - ReadBytes uint64 `json:"readBytes,omitempty"` - WriteBytes uint64 `json:"writeBytes,omitempty"` + ReadBytes uint64 `json:"readBytes,omitempty"` + WriteBytes uint64 `json:"writeBytes,omitempty"` + ReadBytesPresent *bool `json:"readBytesPresent,omitempty"` + WriteBytesPresent *bool `json:"writeBytesPresent,omitempty"` +} + +// CounterPresence distinguishes observed zero from an omitted direction. Older +// agents omitted zero values and presence, so only their positive counters are +// unambiguous. New agents carry explicit presence independently of counter value. +func (io *ContainerBlockIO) CounterPresence() (read, write bool) { + if io == nil { + return false, false + } + read, write = io.ReadBytes > 0, io.WriteBytes > 0 + if io.ReadBytesPresent != nil { + read = *io.ReadBytesPresent + } + if io.WriteBytesPresent != nil { + write = *io.WriteBytesPresent + } + return } // PodmanContainer carries metadata extracted from Podman-specific annotations. diff --git a/pkg/metrics/docker_observation_contract.go b/pkg/metrics/docker_observation_contract.go new file mode 100644 index 000000000..8b03199b5 --- /dev/null +++ b/pkg/metrics/docker_observation_contract.go @@ -0,0 +1,35 @@ +package metrics + +// Docker history written before explicit counter presence mixed unavailable +// readings with measured zero. Its disk percentage also measured image-layer +// composition rather than filesystem capacity. Keep those rows for retention +// and rollback, but never reinterpret them as observations under this contract. +// +// The app-container storage family is shared with other providers. Their valid +// new capacity readings remain supported. Public metric names are unchanged. +func hasDockerObservationContract(resourceType string) bool { + return resourceType == "dockercontainer" || resourceType == "docker" +} + +func storedObservationMetric(resourceType, metricType string) string { + if hasDockerObservationContract(resourceType) { + switch metricType { + case "disk", "diskread", "diskwrite": + return metricType + ".observed" + } + } + return metricType +} + +func projectDockerObservations(result map[string]map[string][]MetricPoint) { + for _, series := range result { + for _, metric := range []string{"disk", "diskread", "diskwrite"} { + delete(series, metric) + stored := metric + ".observed" + if points, ok := series[stored]; ok { + series[metric] = points + delete(series, stored) + } + } + } +} diff --git a/pkg/metrics/store.go b/pkg/metrics/store.go index 00281afe7..dfe48f58c 100644 --- a/pkg/metrics/store.go +++ b/pkg/metrics/store.go @@ -174,10 +174,11 @@ type SeriesKey struct { // NormalizedSeriesKey builds the SeriesKey the write path would store for the // given identifiers, so callers can match MaxTimestampsForTier results. func NormalizedSeriesKey(resourceType, resourceID, metricType string) SeriesKey { + resourceType = normalizeMetricResourceType(resourceType) return SeriesKey{ - ResourceType: normalizeMetricResourceType(resourceType), + ResourceType: resourceType, ResourceID: normalizeMetricIdentifier(resourceID), - MetricType: normalizeMetricType(metricType), + MetricType: storedObservationMetric(resourceType, normalizeMetricType(metricType)), } } @@ -723,7 +724,7 @@ func validateMetricWrite(resourceType, resourceID, metricType string, tier Tier) return "", "", "", false, fmt.Sprintf("unsupported metric tier %q", tier) } - return normalizedType, normalizedID, normalizedMetric, true, "" + return normalizedType, normalizedID, storedObservationMetric(normalizedType, normalizedMetric), true, "" } // Write adds a metric to the write buffer with the 'raw' tier by default @@ -1297,6 +1298,11 @@ func (s *Store) queryBatch( return map[string]map[string][]MetricPoint{}, nil } normalizedMetricTypes := normalizeMetricTypes(metricTypes) + if hasDockerObservationContract(resourceType) { + for i, metric := range normalizedMetricTypes { + normalizedMetricTypes[i] = storedObservationMetric(resourceType, metric) + } + } tiers := s.tierFallbacks(end.Sub(start)) if len(tiers) == 0 { @@ -1730,6 +1736,9 @@ func (s *Store) queryRetainedChunk(resourceType string, resourceIDs []string, me flushBucket() flushSeries() + if hasDockerObservationContract(resourceType) { + projectDockerObservations(result) + } return result, nil } diff --git a/pkg/metrics/store_docker_observation_contract_test.go b/pkg/metrics/store_docker_observation_contract_test.go new file mode 100644 index 000000000..cde253191 --- /dev/null +++ b/pkg/metrics/store_docker_observation_contract_test.go @@ -0,0 +1,148 @@ +package metrics + +import ( + "fmt" + "testing" + "time" +) + +func TestDockerObservationContractSeparatesLegacyAcrossRetainedReads(t *testing.T) { + for _, family := range []string{"dockerContainer", "docker"} { + t.Run(family, func(t *testing.T) { + store, err := NewStore(DefaultConfig(t.TempDir())) + if err != nil { + t.Fatal(err) + } + defer store.Close() + ts := time.Now().Add(-10 * time.Minute).Truncate(time.Minute) + // Seed the old physical schema directly. New public writes must not grant + // these ambiguous historical rows an observation provenance retroactively. + for _, id := range []string{"a", "b", "legacy-only"} { + for _, metric := range []string{"disk", "diskread", "diskwrite"} { + for _, tier := range []Tier{TierRaw, TierMinute, TierHourly} { + _, err := store.db.Exec(`INSERT INTO metrics(resource_type,resource_id,metric_type,value,timestamp,tier) VALUES(?,?,?,?,?,?)`, normalizeMetricResourceType(family), id, metric, 99, ts.Unix(), string(tier)) + if err != nil { + t.Fatal(err) + } + } + } + } + for _, id := range []string{"a", "b"} { + // Exercise the buffered, synchronous and bounded writer entry points. + store.Write(family, id, "diskread", 0, ts) + store.WriteBatchSync([]WriteMetric{{ResourceType: family, ResourceID: id, MetricType: "diskwrite", Value: 0, Timestamp: ts, Tier: TierRaw}}) + store.WriteBatchBounded([]WriteMetric{{ResourceType: family, ResourceID: id, MetricType: "disk", Value: 25, Timestamp: ts, Tier: TierRaw}}) + store.Write(family, id, "cpu", 0, ts) + } + store.Flush() + // Both generations roll up independently. Old minute/hourly data must + // neither replace the corrected zeros nor contaminate their averages. + if !store.rollupTierWindow(TierRaw, TierMinute, 60, ts.Unix()-60, ts.Unix()+120) { + t.Fatal("rollup failed") + } + coverage, err := store.MaxTimestampsForTier(TierRaw) + if err != nil { + t.Fatal(err) + } + if !coverage[NormalizedSeriesKey(family, "a", "diskread")].Equal(ts) { + t.Fatalf("physical coverage mismatch: %+v", coverage) + } + for _, step := range []int64{0, 60} { + t.Run(fmt.Sprintf("step-%d", step), func(t *testing.T) { + start, end := ts.Add(-2*time.Hour), ts.Add(time.Hour) + assert := func(series map[string][]MetricPoint, legacy bool) { + t.Helper() + for _, metric := range []string{"disk", "diskread", "diskwrite"} { + points := series[metric] + if legacy { + if len(points) > 0 { + t.Fatalf("legacy %s leaked: %+v", metric, points) + } + continue + } + want := 0.0 + if metric == "disk" { + want = 25 + } + if len(points) != 1 || points[0].Value != want || points[0].Min != want || points[0].Max != want { + t.Fatalf("%s observation contaminated: %+v", metric, points) + } + if _, exists := series[metric+".observed"]; exists { + t.Fatalf("physical metric leaked: %+v", series) + } + } + } + all, err := store.QueryAll(family, "a", start, end, step) + if err != nil { + t.Fatal(err) + } + assert(all, false) + if len(all["cpu"]) != 1 || all["cpu"][0].Value != 0 { + t.Fatalf("unrelated zero lost: %+v", all) + } + batch, err := store.QueryAllBatch(family, []string{"a", "b", "legacy-only"}, start, end, step) + if err != nil { + t.Fatal(err) + } + assert(batch["a"], false) + assert(batch["b"], false) + assert(batch["legacy-only"], true) + selected, err := store.QueryMetricTypesBatch(family, []string{"a", "b", "legacy-only"}, []string{"disk", "diskread", "diskwrite"}, start, end, step) + if err != nil { + t.Fatal(err) + } + assert(selected["a"], false) + assert(selected["b"], false) + assert(selected["legacy-only"], true) + for _, metric := range []string{"disk", "diskread", "diskwrite"} { + points, err := store.Query(family, "a", metric, start, end, step) + if err != nil { + t.Fatal(err) + } + if len(points) != 1 || points[0].Value != all[metric][0].Value { + t.Fatalf("selected %s differs: %+v", metric, points) + } + old, err := store.Query(family, "legacy-only", metric, start, end, step) + if err != nil || len(old) != 0 { + t.Fatalf("legacy selected %s leaked: %+v %v", metric, old, err) + } + } + }) + } + var legacyCount int + err = store.db.QueryRow(`SELECT COUNT(*) FROM metrics WHERE resource_type=? AND metric_type IN ('disk','diskread','diskwrite')`, normalizeMetricResourceType(family)).Scan(&legacyCount) + if err != nil || legacyCount != 27 { + t.Fatalf("legacy history was changed: count=%d err=%v", legacyCount, err) + } + }) + } +} + +func TestDockerObservationContractLeavesOtherFamiliesUnchanged(t *testing.T) { + store, err := NewStore(DefaultConfig(t.TempDir())) + if err != nil { + t.Fatal(err) + } + defer store.Close() + ts := time.Now().Truncate(time.Second) + for _, family := range []string{"vm", "ct", "agent", "dockerHost", "disk", "storage", "k8s"} { + for _, metric := range []string{"disk", "diskread", "diskwrite"} { + store.WriteWithTier(family, "one", metric, 0, ts, TierRaw) + } + } + store.Flush() + for _, family := range []string{"vm", "ct", "agent", "dockerHost", "disk", "storage", "k8s"} { + all, err := store.QueryAll(family, "one", ts.Add(-time.Second), ts.Add(time.Second), 0) + if err != nil { + t.Fatal(err) + } + for _, metric := range []string{"disk", "diskread", "diskwrite"} { + if len(all[metric]) != 1 || all[metric][0].Value != 0 { + t.Fatalf("%s/%s zero changed: %+v", family, metric, all) + } + if NormalizedSeriesKey(family, "one", metric).MetricType != metric { + t.Fatal("unrelated physical key changed") + } + } + } +} diff --git a/scripts/release_control/canonical_completion_guard_test.py b/scripts/release_control/canonical_completion_guard_test.py index 08bfcc31d..53c41b475 100644 --- a/scripts/release_control/canonical_completion_guard_test.py +++ b/scripts/release_control/canonical_completion_guard_test.py @@ -225,6 +225,7 @@ class CanonicalCompletionGuardTest(unittest.TestCase): "internal/config/host_continuity_test.go", "internal/models/metrics_types_test.go", "internal/monitoring/availability_probe_agent_test.go", + "internal/monitoring/docker_metric_presence_test.go", "internal/monitoring/monitor_host_agent_removal_lifecycle_test.go", "internal/monitoring/monitor_host_agents_test.go", "scripts/installtests/agent_state_dir_lifecycle_test.go", @@ -354,6 +355,8 @@ class CanonicalCompletionGuardTest(unittest.TestCase): "internal/dockeragent/agent_collect_test.go", "internal/dockeragent/agent_cpu_test.go", "internal/dockeragent/agent_internal_test.go", + "internal/dockeragent/blockio_presence_test.go", + "internal/dockeragent/collect_tmpfs_test.go", "internal/dockeragent/swarm_coverage_test.go", ], } @@ -451,6 +454,7 @@ class CanonicalCompletionGuardTest(unittest.TestCase): "internal/config/host_continuity_test.go", "internal/models/metrics_types_test.go", "internal/monitoring/availability_probe_agent_test.go", + "internal/monitoring/docker_metric_presence_test.go", "internal/monitoring/monitor_host_agent_removal_lifecycle_test.go", "internal/monitoring/monitor_host_agents_test.go", "scripts/installtests/agent_state_dir_lifecycle_test.go", @@ -481,6 +485,7 @@ class CanonicalCompletionGuardTest(unittest.TestCase): "test_prefixes": [], "exact_files": [ "internal/config/host_continuity_test.go", + "internal/monitoring/docker_metric_presence_test.go", "internal/monitoring/issue1485_unraid_lifecycle_test.go", "internal/monitoring/issue1595_collection_trust_test.go", "internal/monitoring/monitor_docker_test.go", @@ -1316,6 +1321,7 @@ class CanonicalCompletionGuardTest(unittest.TestCase): "exact_files": [ "frontend-modern/src/types/api.ts", "internal/api/action_runner_credentials_test.go", + "internal/api/ai_handlers_investigation_additional_test.go", "internal/api/ai_handlers_more_test.go", "internal/api/ai_handlers_patrol_actions_additional_test.go", "internal/api/alerting/external_probe_notifications_test.go", diff --git a/scripts/release_control/subsystem_lookup_test.py b/scripts/release_control/subsystem_lookup_test.py index 5823d4d78..36c799ae3 100644 --- a/scripts/release_control/subsystem_lookup_test.py +++ b/scripts/release_control/subsystem_lookup_test.py @@ -4327,6 +4327,7 @@ class SubsystemLookupTest(unittest.TestCase): "internal/config/host_continuity_test.go", "internal/models/metrics_types_test.go", "internal/monitoring/availability_probe_agent_test.go", + "internal/monitoring/docker_metric_presence_test.go", "internal/monitoring/monitor_host_agent_removal_lifecycle_test.go", "internal/monitoring/monitor_host_agents_test.go", "scripts/installtests/agent_state_dir_lifecycle_test.go", @@ -4349,6 +4350,7 @@ class SubsystemLookupTest(unittest.TestCase): monitoring_match["verification_requirement"]["exact_files"], [ "internal/config/host_continuity_test.go", + "internal/monitoring/docker_metric_presence_test.go", "internal/monitoring/issue1485_unraid_lifecycle_test.go", "internal/monitoring/issue1595_collection_trust_test.go", "internal/monitoring/monitor_docker_test.go", @@ -4462,6 +4464,7 @@ class SubsystemLookupTest(unittest.TestCase): [ "internal/monitoring/issue1595_collection_trust_test.go", "internal/unifiedresources/availability_link_test.go", + "internal/unifiedresources/history_identity_test.go", "internal/unifiedresources/kubernetes_registry_test.go", "internal/unifiedresources/pbs_pmg_registry_test.go", "internal/unifiedresources/registry_merge_policy_test.go", @@ -4492,6 +4495,7 @@ class SubsystemLookupTest(unittest.TestCase): [ "internal/monitoring/issue1595_collection_trust_test.go", "internal/unifiedresources/availability_link_test.go", + "internal/unifiedresources/history_identity_test.go", "internal/unifiedresources/kubernetes_registry_test.go", "internal/unifiedresources/pbs_pmg_registry_test.go", "internal/unifiedresources/registry_merge_policy_test.go",