Merge pull request #1935 from rcourtman/fix/docker-observation-evidence

Preserve diagnostic evidence and Patrol outcome truth
This commit is contained in:
rcourtman
2026-09-06 22:21:42 +01:00
committed by GitHub
126 changed files with 4765 additions and 4698 deletions
@@ -65,7 +65,7 @@ reproduction evidence, not a representative customer success rate.
| 2. Shared evidence | Preserve canonical risk reasons and SMART counters, source/time semantics and history across tools/turns. | Regression tests preserve unknown versus zero and all canonical evidence. Real responses can inspect the same facts as the product. | Implemented and qualified for the named shared-evidence defects. Canonical disk detail, risk and cadence pass real data-path proof. Affected package and concurrency checks pass. Integrated CI later exposed remaining query and allocation regressions. The final bounded query-reuse correction passes complete selected exact-base worker comparisons and full metrics/database and focused race checks. Final landing CI passed and PRs #1928 and #1929 merged. Real-model interpretation failures remain tracked in step 5. |
| 3. Diagnostic orchestration | Correct proposal-as-proof. Audit triage budgets, unmatched-signal evaluation, assessment completion and investigation cutoffs. | No code-written causal conclusion. No quality inferred from tool, flag or finding counts. Each retained pass has an objective reason. Safety boundaries and incomplete outcomes remain explicit. | Proposal promotion and capture inference were removed in c5d2f56dda. Commit 668af3fe6b removes investigation success-call floors, checkpoint instructions and generic call-count wrap-up rules. The detection slice removes contextless follow-up passes, flag/report-count policy and first-finding completion modes. Full chat and AI suites, focused API and conversation race tests pass. Real-model/action outcome qualification remains open. |
| 4. Issue through verified outcome | Follow existing issue/investigation/action records into Assistant, approval, execution and independent readback. | Accepted proposal is visibly distinct from execution and verification. Rejected or unsupported actions do not become success. Uncertainty can survive an action proposal. | Existing foundation, full journey qualification pending. |
| 5. Ground-truth qualification and landing | Extend existing qualification tooling only where necessary. Exercise healthy/unhealthy, dependency, missing-access, storage/backup and approved/rejected action cases. Inspect the final browser journey at desktop and narrow widths. | Record exact source/model/permissions, evidence, decisions, faults/misses, latency and verification. Fix in-scope failures, pass appropriate proofs and land scoped commits. | Pending. |
| 5. Ground-truth qualification and landing | Extend existing qualification tooling only where necessary. Exercise healthy/unhealthy, dependency, missing-access, storage/backup and approved/rejected action cases. Inspect the final browser journey at desktop and narrow widths. | Record exact source/model/permissions, evidence, decisions, faults/misses, latency and verification. Fix in-scope failures, pass appropriate proofs and land scoped commits. | Partial. Regression, controlled browser and live collector evidence are recorded below. Real-model diagnosis, linked approval/action outcomes and installed collector qualification remain open. |
Use one shared runtime and the existing qualification runner, not a second
product intelligence engine or a new parallel lifecycle. Preserve independent
@@ -93,12 +93,13 @@ Each removal must run its focused regression and affected complete journey.
### Completion and external dependencies
The local implementation goal remains open until required qualification is
performed. Ordinary Assistant requests work with the current subscription route,
but autonomous Patrol has an explicit provider-policy refusal and remains
blocked. Do not rephrase the refused probe, bypass the readiness boundary or
count an interactive request as an autonomous Patrol pass. A supported provider
path is required for that qualification. Prepare other work while resolving the
provider dependency through supported configuration.
performed. The maintainer authorized Gemini 3.8 Flash through OpenRouter with a
US$5 key limit and one-day expiry on 2026-09-06. That supported route passes the
streaming readiness and initial live Watch and dependency cases recorded below.
The earlier Claude subscription refusal belongs to the exact synthetic
continuation request. It does not establish a blanket restriction on autonomous
monitoring. The refused request has not been retried or rephrased. Readiness is
not evidence that diagnosis, action execution or independent recovery succeeds.
Release publication and wider product readiness are separate. Independent
volunteered Pro environments are still required before claiming repeatable
@@ -1408,3 +1409,849 @@ correction above is a separate scoped change and requires its own landing checks
The redesign remains open for reliable interpretation, the config-read contract,
storage/backup, approved/rejected action outcomes and supported autonomous Patrol
qualification. Wider customer readiness still requires independent Pro environments.
### Docker measurement correction plan, 2026-09-06
The next shared-source correction distinguishes absent block-I/O observations
from measured idle zero and removes the Docker layer-size ratio from filesystem
capacity. Counter presence uses the existing rate tracker contract. Missing
reports must not reset the baseline or fabricate samples. Container layer sizes
remain descriptive metadata.
Persisted Docker-family disk series previously mixed invalid capacity ratios and
unobserved I/O zeros with real measurements. New disk observations use separate
physical series keys while public metric names remain unchanged. Retained reads
exclude ambiguous legacy disk series without deleting or relabelling them. The
shared app-container storage family also serves non-Docker providers, so new
valid capacity observations must remain supported. Non-Docker series retain
existing behavior. Explicit zero must survive every retained-read API and rollup.
Browser verification is required after the final backend build. Interaction
matrix: `/docker` at 1440x1000, 900x1000 and 390x1000, container selection,
resource drawer open/close, current metrics, history expansion, measured idle,
unavailable readings, and reload. Inspect actual pixels, scrolling, focus and
Escape dismissal. `/patrol` evidence rendering must preserve absent versus zero
in current-resource and retained-history tool results. Controlled responses may
qualify rendering but cannot qualify diagnosis. Live read-only API observations
must bind to the rebuilt backend. No autonomous subscription retry or paid-model
request is authorized by this correction.
Collection also carries optional presence for each I/O direction. Explicit zero
entries survive the report JSON. For older reports without presence, only
positive counters establish an observation, so ambiguous zeros remain unavailable
until the agent is updated or a positive baseline exists. This does not require
re-enrollment. Docker-host first-disk history and network-counter presence are
adjacent limits outside this container block-I/O correction.
The final browser matrix also covers the shared host I/O table, Docker host
Overview and Machines table/tooltip at the same three widths. Partial read/write
observations must show a missing marker for the absent direction, retain measured
zero, and remain excluded from sums used for sorting and comparison. Exercise
column selection, hover/focus, tooltip dismissal and scrolling where present.
### Docker correction qualification and scope
The implementation carries per-direction presence from collection and report
JSON into the existing rate tracker, canonical resource metrics, persisted
history and resource-to-browser conversion. REST resource adaptation also
preserves optional rates. Shared rate formatting keeps missing values distinct
from zero in Machines and Docker host details. Incomplete rates do not become
complete throughput totals for sorting or comparison. A browser-discovered
first-user column migration bug is corrected in the shared preference hook, so
showing Disk I/O survives the first reload.
Legacy workload conversion in `frontend-modern/src/hooks/useWorkloads.ts` still
uses numeric direction fields with grouped availability. Its direction-level
modernization remains a separate consumer follow-up. Docker-host first-disk
history and network presence are also outside this container measurement slice.
The correction must not be represented as complete coverage of all metrics or
all monitoring surfaces. No new model competence or autonomous action result is
claimed.
Affected Go package checks and targeted race checks ran on pulse-dev with
Go1.26.8. The changed websocket assertion now expects observed read zero with
absent write omitted. Targeted frontend suites and type checking cover optional
rates, REST conversion, sorting, formatting and column persistence. Ten paired
read-benchmark rounds used the unchanged parent store via Go overlay. The
canonical >10%, p<0.05 regression gate passed, with +0.82% timing geomean in this
scoped comparison. This is not a fleet-load or full-product performance claim.
The Pro backend was cross-built on pulse-dev from the changed source and
installed into the existing local dev runtime. Binary SHA256:
`59f05f954ff8080bd3e8f3054b2b059281c49172ee771b2e454c255241158a4a`.
No production agent was replaced. Older agents remain compatible and treat
ambiguous zero counters conservatively.
Private receipts: `/Volumes/Development/pulse/tmp/patrol-docker-observed-metrics/`.
Worker logs: `/opt/pulse-release-worker/patrol-docker-observed-proof/`.
PR1934's preceding identity correction merged at
`6b0abc3bee9ffa81f6ab298b5b67ee11369688a0` with all checks passing. The current
measurement slice passed final-source Playwright inspection on `/docker`,
`/standalone/machines` and `/patrol` at 1440x1000, 900x1000 and 390x1000.
Live history, controlled absence/idle/loading/error, partial host rates, column
persistence, nested picker dismissal, tooltip focus, and expanded Assistant
evidence were exercised. The source-bound receipt is
`frontend-modern/browser-verification.json`. The unused shared host table card
has type and selector coverage, not an active-route browser claim. Controlled
responses qualify rendering only. Landing checks remain separate from the
unperformed diagnosis, approved/rejected action and recovery qualifications.
### Configuration-read correction plan
A successful canonical container get followed by a false config `not found`
result is a source contract defect. Native configuration reads must use current
canonical inventory for identity and provider capability. Optional session
resolution preserves continuity for later actions, not proof of existence.
Explicit query restrictions must be checked before registering or refreshing a
resource. Unsupported adapters, missing configuration providers, unavailable
placement and empty provider responses remain distinct from missing inventory.
No action validation or native log-read authority changes in this slice.
Regression matrix: TrueNAS config with absent, empty and existing session
context, canonical identity across aliases, explicit query denial without a
provider call, Docker unsupported capability, genuinely missing inventory,
unavailable placement, provider failure and nil provider response. Reproduce
the failing cases before changing runtime code. Run affected Go tools checks
and focused race coverage on pulse-dev.
Browser matrix after the final build: `/patrol` Assistant tool result details
at 1440x1000, 900x1000 and 390x1000, available configuration, unsupported
capability, true missing resource and denied/provider-failed results. Exercise
open/closed details, hover and keyboard focus, Enter/Space, deepest output
scrolling, Escape, and persisted/reloaded evidence. Use captured actual tool
results to qualify rendering without claiming model diagnosis or native
provider integration. No autonomous subscription retry or separately billed
provider request is part of this correction.
### Configuration-read correction qualification
The baseline reproduced absent/empty session failures, stale session placement
and false not-found results after successful canonical gets. The corrected
read path uses canonical resource identity and current provider placement. It
checks explicit query restrictions before registration and preserves an existing
query-only session's action limits. Unsupported configuration, unavailable
provider/placement and nil provider responses carry explicit reasons and the
tool error bit. Unavailable inventory and missing read state also remain failures
rather than evidence of resource absence. Actual inventory absence remains the
existing not-found lookup result.
Fourteen focused contract cases pass with strict resolution enabled. The
existing native-config regression, full tools package and focused race proof
passed on pulse-dev with Go1.26.8. The final-source Pro binary SHA256 is
`552699cdf2e61a4ca1cea2ac5ef4e065735cbd1dbca01665456e184bd4fc3533`.
It is installed only in the existing local dev stack. No production agent or
provider configuration was changed.
Playwright passed on `/patrol` at 1440x1000, 900x1000 and 390x1000. Eight actual
tool results were replayed and inspected, including successful, unavailable,
missing, denied and failed reads. Expanded inputs/outputs, keyboard toggles,
scrolling, Escape, reload and controlled session restoration preserve exact
evidence and error state. Controlled session responses prove rendering and
reload behavior, not server persistence or a new model/native-provider result.
The source-bound browser receipt records those limits. Private artifacts are
under `/Volumes/Development/pulse/tmp/patrol-config-read-contract/` and worker
logs under `/opt/pulse-release-worker/patrol-config-read-proof/`.
PR1935's Docker correction required two legacy partial-total test expectations
to be updated in `0fcb2ee147354de770dfc4b0b9672d8c2c9dceb2`. The focused 55-test
file and scoped hook passed. Its latest CI has no failures and remains pending
completion. The configuration correction still requires its own landing checks.
Real-model retest, temporal/storage interpretation, approved and rejected
action outcomes and independent recovery proof remain open. Autonomous
subscription refusal and separately billed provider approval boundaries remain
unchanged.
### Corrected storage evidence, ordinary Assistant retest
The Docker observation and canonical configuration corrections are pushed to
PR #1935 at `355ac1f0a481d6dbc9a7ff3977bced0956711979`. Their exact staged
worker pre-commit checks passed without source changes. Remote checks remain
pending. The current configuration runtime also passed all fourteen contract
cases, the complete tools package and focused race checks.
One ordinary read-only storage assessment ran on 2026-09-06 from
12:27:05.857Z to 12:30:03.726Z, an HTTP window of 177.869s. It used the existing
`claude-subscription:claude-opus-5` route, explicit `autonomous_mode=false`,
read-only control and thirteen successful tool reads. No infrastructure change,
paid-model request or autonomous readiness retry occurred. This is a single
assessment, not a success-rate or latency estimate.
The answer identified the backup datastore at 90.6% utilisation and its active
capacity warning affecting seven workloads. It used the corrected container I/O
history, separated cumulative device counters from rates, acknowledged missing
container filesystem usage, and retained the reason for missing older history
as unknown. It did not attribute older host I/O peaks to the container whose
returned I/O window starts later. These are useful observations.
The complete diagnosis still does not qualify. Its opening assurance that the
container is not short of space contradicts the later acknowledgement that
container filesystem usage is unavailable. It treats high retained host rates
as bucket/counter artifacts without establishing that mechanism. It includes a
host CPU maximum timestamped 21:00 the previous evening in a 03:00-04:00 window.
The suggested retention explanation is not established by capacity alone, and
available PBS job reads were not performed. Its rough growth extrapolation uses
retained extrema, not a measured first-to-last slope, and must retain that limit.
Corrected observations have not established reliable interpretation.
The run also highlights a tool-context distinction to review: Docker-host
`agent_connected` describes the command connection, while telemetry may still
arrive through other collection paths. The model treated current telemetry and
that false connection flag as an unresolved inconsistency. Its storage-pools
request supplied `host`, although the tool schema only advertises that filter
for RAID and Ceph detail. That call returned all pools. Neither observation
justifies fabricating resource absence or collection downtime.
Playwright exercised `/patrol` at 1440x1000 and 390x1000, the actual answer,
all thirteen expanded tool records, keyboard activation, deepest output
scrolling, Escape, reload and the persisted session. Every displayed input and
output matches the persisted tool records. Pixel review includes the answer,
evidence and the mobile table scrolled to its rightmost state. The table's
400-pixel content is reachable inside its 309-pixel horizontal viewport.
The artificial selected-route warning comes from blocked non-GET route checks,
so this does not qualify the unmodified provider-readiness UI.
The runtime binary SHA256 stayed
`552699cdf2e61a4ca1cea2ac5ef4e065735cbd1dbca01665456e184bd4fc3533`
through the request and browser pass. Private request, source, binary, tool,
persistence, evaluation and pixel receipts are under workspace-relative
`tmp/patrol-storage-assistant-check/`. No native config action was requested,
so this assessment does not qualify model use of that corrected action.
Storage-fault ground truth, reliable diagnosis, approved/rejected actions and
independent recovery remain open. The supported autonomous provider dependency
and wider independent-Pro-environment gate remain unchanged.
### Command connection evidence correction plan
Live command connectivity and retained monitoring observations are independent
facts. The shared tool contract will name command-agent connections explicitly,
including parent-node connections, without changing routing or execution policy.
Topology built without a command-connection snapshot must omit connection flags,
execution hints and connected counts rather than manufacture false/zero values.
An observed empty snapshot still reports disconnected/zero. Assistant inventory
context must preserve the same observation boundary. Existing permission,
approval and invocation checks remain authoritative.
Regression matrix: current Docker inventory and metrics with disconnected and
connected command transport, read-only control with a connected agent, parent
node versus guest connection, topology without a connection observation versus
an observed empty set, and Assistant's seeded inventory. Run affected tools/chat
packages and focused race proof on pulse-dev.
Browser matrix after rebuilding the local Pro backend: `/patrol` at 1440x1000,
900x1000 and 390x1000, actual captured query results showing disconnected and
connected command transport beside unchanged monitored workload evidence.
Exercise tool details open/closed, keyboard focus/activation, deepest output
scrolling, Escape, reload and persisted result presentation. Inspect pixels and
bind receipts to the final source and binary. Controlled rendering proof does
not qualify model interpretation, autonomous Patrol or infrastructure actions.
### Command connection evidence qualification
The original projection failed the new regression because it labelled command
transport as generic agent connectivity and emitted connected-agent counts from
an inventory-only seed. Canonical guest search also promoted a parent-node
connection into a direct guest connection. The shared projection now retains
those distinctions. Existing host aliases remain available for non-guest
resources. No routing, approval, execution or provider policy boundary changes.
Four canonical query cases pass: no command connection, connected read-only
transport, a direct guest connection without a parent connection, and connected
transport with control enabled. Current workload state and CPU remain available
in every case and no command is executed. Separate checks prove that topology
without a command snapshot omits connection and execution hints and connected
counts, while an observed empty snapshot retains false/zero. Assistant inventory
context inherits that same unobserved state.
Final source proof on pulse-dev used Go1.26.8 and GOMAXPROCS4. The full tools
package passed in 59.456s and chat in 6.402s. Focused race checks passed in 1.048s
and 1.030s. The Pro runtime cross-build passed and the installed local binary
SHA256 is `bcaf748107211ee733a6dc0f4d17220d9b4d1ce1918c25bde27cf3d10c0d6379`.
The managed development process restarted onto that artifact and `/api/health`
reported healthy. No production agent was replaced.
Playwright exercised nine captured results at `/patrol`, 1440x1000, 900x1000
and 390x1000. Inputs, outputs and completed states match exactly before and after
controlled session reload. Hover, keyboard focus/activation, expansion/collapse,
deepest output scrolling, Escape and session selection passed. Pixel inspection
covered each distinct connection state, unchanged workload metrics and restored
mobile results. Backend and renderer hashes remained unchanged. The artificial
route-check warning and controlled persistence fixtures retain their earlier
qualification limits. No model request was part of this proof.
A read-only settings check confirms the cached `provider_refusal` still carries
its original `2026-09-05T19:46:39Z` timestamp and `patrol_capable=false`.
Private source bindings, logs, captured outputs, runtime process/health receipts
and browser proof are at workspace-relative `tmp/patrol-command-context/`.
The change still requires its scoped pre-commit and landing checks. The preceding
PR #1935 head `4d302109cee0758a132ff150935630b50114cc05` has no reported failures
but its Build and Test and Core E2E runs are pending behind live earlier runs
on the same branch. Those workflows are not restarted or cancelled.
This correction establishes the connection evidence contract, not reliable
interpretation. Native configuration-read model use, storage-fault diagnosis,
approved/rejected action outcomes and independent recovery remain open, as do
the supported autonomous provider dependency and independent-environment gate.
### Ordinary Assistant storage fault and recovery, 2026-09-06
The command-connection correction passed the exact nine-file worker pre-commit
and was pushed as `f5f440dbad18d83557104d2cf6197d8319949e44` in PR #1935.
Required CI remains in progress. This is not a release or a completed goal.
An owned DockerLab run used the checked-in storage-pressure manifest on Tower.
Independent observations established a healthy worker with 8,347,648 free bytes,
then a real ENOSPC fault with zero free bytes, a running/unhealthy worker and a
healthy control. Pulse collection converged to both states before the request.
The ordinary read-only Assistant used the configured subscription Opus 5 route.
It was not an autonomous Patrol request and did not retry the cached refusal.
The diagnosis took 114.847 seconds and ten tool calls. It identified the worker's
unhealthy state, the control's current healthy state and a failed command-route
log read. It did not retry that unavailable capability. However, it falsely
ruled out resource pressure using low CPU, memory, network and disk-read values.
Filesystem capacity was absent from its evidence and was independently full.
This is a failed diagnosis, despite its otherwise useful uncertainty statement
and suggested diagnostic read. No model-directed mutation occurred.
After the answer, the independent oracle still found zero available bytes and
an unhealthy worker. Removing only the owned fill file restored 8,220,672 bytes
and healthy status. Pulse collected recovery and resolved the health alert.
A follow-up in the same Assistant session took 91.358 seconds and five new reads.
It correctly identified current recovery and the resolved alert, distinguished
symptom recovery from an unknown cause and did not invent an intervention.
It overstated continuous control health and non-impact from sparse observations.
The recovery assessment is partial, not a complete incident explanation.
An independent post-answer check confirmed healthy worker/control, no container restart
and 7,639,040 free bytes. Both cleanup passes passed, with no second-pass work
and unchanged unrelated inventory. The disposable resources are removed.
Playwright exercised `/patrol` and Assistant at 1440x1000 and 390x1000, all fifteen
retained tool input/output pairs, keyboard expansion/collapse, deepest output
scrolling, complete answers, Escape, reload and the same retained conversation.
Rendered inputs/outputs match persisted records. Pixel inspection covered both
answers and the failed-access result on desktop and mobile. Runtime and source
hashes remained unchanged across both requests, with binary
`bcaf748107211ee733a6dc0f4d17220d9b4d1ce1918c25bde27cf3d10c0d6379`.
The route warning was an artifact of blocking non-chat POSTs in the proof browser.
The original autonomous refusal timestamp remained `2026-09-05T19:46:39Z`.
Private fixtures, source bindings, observations, screenshots and assessments
are at workspace-relative `tmp/patrol-storage-fault-case/`.
### Next canonical correction: tmpfs inventory
Before implementation, source and native inspection establish a collection gap:
Docker reports the owned scratch mount in `HostConfig.Tmpfs`, while `Mounts` is
empty. `internal/dockeragent/collect.go` copies only `Mounts`, so shared resource
queries falsely present an empty mount inventory. Preserve these native tmpfs
entries through the existing report mount type. Keep destination, type and
reported options, derive read/write from those options, preserve authoritative
existing mount records and deterministic ordering. Do not infer used/free space
from a configured size. No enrollment, permission or production agent change is
part of this collection correction.
Proof plan: reproduce the captured tmpfs-only inspect shape through the actual
collector, then cover existing mounts, overlapping representations, read-only
options and absent host configuration. Run targeted/full collector checks on
pulse-dev and verify the report through the existing shared projection. Browser
proof after the final change must exercise mount evidence in resource details
and Assistant tool results, desktop and mobile, including deepest expansion and
reload. A captured-result rendering check is not installed-agent or model
qualification. Leave those limits explicit until the new collector is exercised
through a supported installed path.
The incident lookup also needs an identity audit: the canonical container ID
returned no incident recording while the observed health alert used its legacy
Docker resource ID. This is a concrete lookup discrepancy to investigate, not
yet proof that a recording exists. Model inference from unmeasured capacity and
sparse health history remains an open quality failure. Supported autonomous
provider, approved/rejected actions and independent environments remain open.
Further shared-projection inspection before editing found that `MountInfo` drops
native mount type/options and that canonical app-container queries derive write
access from equality with the single string `ro`, misreporting compound read-only
options. The same slice must preserve type/options and the canonical `RW` boolean
through both canonical-provider and typed read-state query paths. Add a query
regression and capture its actual output for final Assistant browser proof.
This remains mount configuration evidence, not measured filesystem capacity.
### Tmpfs collection and query contract proof
The captured tmpfs-only and mixed-mount regressions failed against the previous
collector, then passed after the collection correction. The full dockeragent
package passed in 18.883s and focused race proof in 1.030s. Existing monitor report
mount propagation and discovery mount regressions passed. Both query paths
failed because type/options were lost, then passed after projection correction.
The full tools package passed in 59.473s and focused race proof in 1.030s.
A final output-only capture rerun passed in 0.013s. All proof used Go1.26.8 and
GOMAXPROCS4 on pulse-dev. The final Pro cross-build passed and the installed
local binary SHA256 is
`0a21dca4106c2ddc6873a3aca3b23378dccef35383ca00d7e9292b966de7c200`.
The managed local backend restarted and `/api/health` reported healthy.
No production collector was replaced.
Final Playwright proof used captured canonical resources at `/docker`, widths
1920, 1440, 900 and 390 with height 1080. It exercised mount summary/title,
keyboard row expansion/collapse, mobile row tapping, adjacent detail state,
mount-destination search, Escape and reloaded search state. The existing wide
mount column is truncated with a full title. Responsive details have no dedicated
mount section. This is an existing presentation limitation, not full mobile
mount inspection qualification. Assistant's complete mount evidence is readable
at `/patrol`, 1440x1000, 900x1000 and 390x1000. Both actual query projections
passed exact input/output comparison, hover/focus, keyboard expansion/collapse,
deepest scrolling, controlled session reload and reopening retained records.
Pixels were inspected on desktop and mobile. Source/binary hashes stayed fixed.
The original autonomous refusal timestamp is unchanged.
These browser fixtures qualify rendering of the corrected shared fields. They
do not qualify an installed collector, actual model interpretation of tmpfs
configuration, or durable backend persistence of those fixture sessions. The
ordinary live diagnosis/recovery records above have real server persistence
and retain their failed/partial judgments. Exact scoped hook and landing remain
required. Required model/action qualification and independent environments are
still open. The typed compatibility get path also does not accept the canonical
ID returned by its list path, so its mount regression uses an existing accepted
name. That identity residual is recorded for modernization, not silently fixed
through this mount projection.
## Canonical incident history, 2026-09-06
The tmpfs correction passed the exact worker hook and was pushed as
`6e18777d30f30b498def39d30016a697cabc4ea7` in PR #1935. That scoped
delivery does not change the failed storage diagnosis or partial recovery verdict.
The incident audit found a source-of-truth mismatch, not evidence that an existing
recording merely needed an ID alias. The legacy five-second recorder has no
production alert callback connected to its coordinator. It samples cached values
using recorder time without preserving their source measurement time. Connecting
that recorder would not supply trustworthy higher-frequency history.
The canonical resource timeline already stores observed changes, alert lifecycle
events and executed actions. Assistant handoffs use a bounded excerpt of this
same store. The shared `pulse_knowledge` incidents action now reads that
organization-pinned timeline directly, using the supplied canonical resource ID.
It does not require the resource still to exist in current inventory, infer
identity from names, include related resources implicitly, or reconstruct events
from current metrics. The response preserves canonical source, observation and
optional occurrence timestamps, state transitions and metadata. `since` filters
on observation time, and bounded results report `has_more`. Empty retained history
does not establish health. Missing or failed storage is a failed read.
Explicit legacy `window_id` lookups remain isolated archive reads, must match the
requested resource, and explain that sample timestamps do not establish source
freshness. The primary incidents action no longer uses those recordings. The
legacy recorder/coordinator startup and API active-count plumbing still exist.
Their retirement is a separate cleanup in this redesign and must preserve any
saved archives. Do not connect them as a replacement incident truth source.
Qualification uses the real SQLite resource store with a fired/resolved lifecycle,
an older excluded record, a related-resource negative control, absent occurrence
time, truncation and empty history. Unavailable/failed storage, invalid input and
archive resource isolation are separate negative controls. Captured actual tool
responses must pass the Assistant expansion, scrolling and reload matrix at
`/patrol`, 1440x1000, 900x1000 and 390x1000. This is contract and rendering proof,
not a new real-model or continuous-coverage claim.
Read-only API inspection of the actual removed storage fixture confirmed a
remaining canonical write-boundary defect. The canonical app-container timeline
returns its creation and removal, while the fired event at 13:14:28.59485Z and
resolved event at 13:17:58.634095Z remain under its legacy Docker resource ID.
The resource API includes related network changes by design. The new tool uses
direct resource history only. `recordAlertTimelineChange` passes the alert's
source ID directly to `BuildAlertTimelineChange`, and `MonitorAdapter.RecordChange`
forwards it without canonical resolution. Consequently this read-path change is
only partial incident-history remediation. The next required owning fix must
resolve event identity before persistence and preserve access to retained prior
identity records, including removed resources, through the shared identity/history
contract. It must not add a Docker string rewrite inside the Assistant tool.
The shared writer and retained-identity correction remain required in this goal.
Private raw API receipts are in `tmp/patrol-canonical-history/live-timeline.json`
and `live-legacy-alert-timeline.json`. No model call or infrastructure mutation
was made during these reads.
The history regression passes through the registered tool dispatcher. The full
tools package passed in 59.856s, the final focused capture passed in 0.036s, and the
focused race check passed in 1.189s on pulse-dev with Go1.26.8 and GOMAXPROCS4.
The final Pro build passed and was installed into the local development stack.
Its SHA256 is `4929aeb869db54122bc5525352d3126c0e9fa7c4847e3fdfa741500600b5e00d`.
The managed backend restarted healthy. Final Playwright proof passed all five
registered-tool cases at `/patrol`, 1440x1000, 900x1000 and 390x1000, including
hover/focus, Enter/Space, deepest output scrolling, Escape and controlled session
reload with exact input/output and success/failure comparison. Root inspected
actual pixels at all three widths. Source and binary hashes remained fixed.
The provider warning stayed visible and no retry, route switch or provider
request was made. This is captured-response rendering, not real-model diagnosis
or server persistence qualification. The exact scoped worker hook gates delivery
through PR #1935.
## Shared Docker history identity, 2026-09-06
The preceding incident-read correction was pushed as
`580a246981c76b401e9007f5ac65c89355b64c6d` in PR #1935. Its live retained
records established the identity split addressed here.
The canonical fix belongs to the shared monitor/store boundary. Exact full
Docker container references resolve through current registry identity. Retained
bindings survive inventory removal, and deterministic source-specific identities
allow legacy records to be found after restart. Names and abbreviated IDs are
not sufficient evidence. A small organization-scoped history alias index joins
readable records without rewriting event IDs, timestamps or metadata. Existing
canonical succession machinery was deliberately not used for these aliases
because it also moves operator state and action indexes. History matching must
not transfer authority. The alias index follows journal retention and separate
store connections read fresh bindings.
Focused regression and race proofs passed on pulse-dev with Go1.26.8 and
GOMAXPROCS4. Full unifiedresources, monitoring and tools packages passed in
37.826s, 79.866s and 59.584s. They cover real alert-manager callbacks, recovery
after inventory removal, restart, replay, same-name controls, tenant isolation,
unchanged operator/approval records and registered Assistant tool reads. Scoped
history lookup measured 0.2610.275ms with one alias and 0.3170.336ms with
20,000 unrelated aliases, at 6,280 bytes and 94 allocations per read. These are
worker microbenchmarks, not fleet or frontend performance qualification.
The verified worker Pro binary has SHA256
`bb6d1508a5b4d23943c37dfc42198f132c0139805dcd1891ee18aca0a9f9dd54`.
It was installed into the local development stack and restarted healthy. Both
canonical and legacy timeline API queries now return the same seven retained
records for the removed storage fixture, including the original fired/resolved
records with exact unchanged content. The complete registered-tool rendering
matrix passed at `/patrol`, 1440x1000, 900x1000 and 390x1000. Six captured cases
include migrated history, bounded and empty results, and unavailable/failed
reads. Hover/focus, Enter/Space, deepest scrolling, Escape and controlled-session
reload preserved exact inputs, outputs and completed/failed states. Pixels were
inspected at all three widths. Source and binary hashes remained fixed. Private
receipts are `tmp/patrol-history-identity/browser/receipt.json` and
`live-history-proof.json`. Controlled session responses qualify rendering, not
server persistence or model diagnosis. The cached provider refusal remained
unchanged. No model request or infrastructure fault was made. The exact scoped
worker hook remains the delivery gate for PR #1935.
This history correction does not qualify the failed ordinary storage diagnosis,
partial recovery claim, installed tmpfs collector, autonomous provider, approved
and rejected action outcomes, or independent Pro environments. The unused legacy
recorder/coordinator still needs retirement with its archives preserved.
The history identity change passed the exact thirteen-file worker hook and was
pushed as `919331d5b3f6076f8616b07eb8e9ca611f26ee52` in PR #1935. Integration
with main `11a8cc2180aae886ec7f92e2333002b57cf1b9a3` preserves both sides of
three additive subsystem-contract conflicts. The host-ingestion auto-merge
retains Docker observation corrections alongside incoming host-link provenance.
Unrelated registry indentation was restored without changing its decoded data.
The combined monitoring, unifiedresources, tools, config and models packages
passed in 85.894s, 41.931s, 59.609s, 18.381s and 0.098s. Focused history,
Assistant, host-link and lifecycle race checks passed. Frontend type checking
and four alert suites passed all 42 tests. The incoming delivery-log component
browser proof passed at 1440x900 and 390x900, including reordered success/failure,
held events, pending state and newest failure. Pixels were inspected. This is
scripted component proof, not installed notification delivery.
The final merged-source Pro binary has SHA256
`234f625cb74be3d300facfed1bb17b17e20c44c06037a1f8f4fffc1c4f49d621`.
Its local restart was healthy. The complete six-case Assistant matrix was
repeated at `/patrol`, 1440x1000, 900x1000 and 390x1000, with exact tool records,
keyboard expansion/collapse, deep scrolling and controlled-session reload.
Pixels and source/binary bindings were checked after this final build. Both
actual removed-container timeline queries still return the same seven retained
records with the original fired/resolved content. No model request, route
switch, infrastructure fault or production collector replacement was made.
The autonomous provider refusal remains enforced. Integration receipts are in
`tmp/patrol-history-integration/`. The full integration hook gates its merge
commit and push. All previously recorded model and wider-readiness gaps remain
open.
## Legacy recorder retirement, 2026-09-06
The disconnected incident coordinator, five-second cached-metrics sampler,
pre-incident buffers, archive writer and unused adapters are removed. They had
no production alert trigger. Canonical resource history remains the primary
incident evidence for Assistant. No replacement diagnosis or scheduling policy
was added.
Explicit archive lookup now requires exact organization, resource and window
binding. The reader is lazy and read-only. Saved file contents, modification
time, mode, old observations, metadata and summary values survive reads. Missing
archives, malformed files and missing windows remain distinct outcomes. Legacy
`recording` status is historical, and the response discloses that the old
`summary.duration_ms` field contains nanoseconds. An old file is never rewritten
to make its evidence appear current.
The incidents API now reports `active_count: null` with
`active_count_status: not_measured`. The retired coordinator's empty map never
established a measured zero. Its legacy incident-memory listing still needs a
canonical query design covering aliases, canonical-only events, honest bounds
and propagated projection-read errors. This is recorded as an open modernization
residual rather than treating that listing as complete.
Final-source registered archive-tool receipts pass Playwright at `/patrol`,
1440x1000, 900x1000 and 390x1000. The five cases cover a saved observation,
unavailable/malformed archives, the wrong resource and a missing window.
Verification includes hover/focus, Enter expansion, exact tool input/output,
deep scrolling, Space collapse, Escape, reload and controlled session reopening.
Actual pixels were inspected. Incoming main alert dispatch wording also passes
its isolated real Overview browser script at all three widths. These controlled
responses prove rendering, not model diagnosis, installed delivery or server
persistence.
The final worker Pro binary is
`bd29e6f27be7b3ad4cfbc37842f4da90f08c6a48c9fc23b12c9c597b67346c9b`.
After the managed local restart, canonical and legacy queries still return the
same seven retained homelab records with the original fired/resolved events
unchanged. The live incidents API reports an unmeasured count. Cached provider
refusal remains enforced. No model request, paid spend, provider retry,
production collector change or fault injection occurred in this slice.
Archive, tools, chat, AI runtime and targeted API/race checks passed on the
worker. One full API run as root invalidated its mode-bit persistence-failure
fixture. That fixture passes unchanged under the normal worker account. The
full API rerun passes under that account (286.960s), as do the incoming
startup-replay and legacy-boundary source checks. Frontend type checks and all
29 incoming alert tests pass. The exact staged hook gates landing.
Private receipts are under `tmp/patrol-archive-retirement/` in the workspace.
### Live collector storage evidence, 2026-09-06
`TestCollectContainerStorageFaultLive` calls the production Docker client and
`collectContainer` implementation against the existing storage-pressure lab.
The opt-in command, run on the worker beside its Docker Unix socket, is:
```sh
PULSE_QUALIFY_ORACLE_DOCKER_CONTEXT=default go test ./internal/dockeragent -run '^TestCollectContainerStorageFaultLive$' -count=1 -timeout=240s -v
```
The final test passed in 8.694s on Docker 29.8.0 against the runtime source tree
of `186ce504c8f0fa6e0174f3b10f3d99f6278f10fc`. Its SHA256 is
`cdf303f4de23c020690c81b0b57190c728319788cb0aaff906b0d4d9d864541e`.
Independent filesystem observations measured 8,380,416 available bytes before
the fault, zero during it and 8,380,416 after recovery. The service stayed
running. Collected and JSON-decoded health followed healthy, unhealthy, healthy,
while the unrelated control remained healthy. Docker returned zero native
`Mounts` throughout. The report retained the single tmpfs destination, type,
8 MiB configuration options and writable setting from `HostConfig.Tmpfs`.
Configured size is not a measured capacity counter in the report.
The test targets exact run-owned container IDs and labels. Existing fixture
cleanup passed, its second cleanup was a no-op and inventory matched the
pre-run snapshot. The test skips without explicit opt-in. All four existing
mount regression cases also pass. Private logs are in
`tmp/patrol-storage-collector-live/` in the workspace.
This proves live collection and report serialization only. It does not send a
report to Pulse, enroll or replace a production agent, call a provider, or
qualify approval, execution, diagnosis or model recovery. No runtime or frontend
source changed in this slice, so no new browser claim is made. Prior storage
diagnosis failures and the cached autonomous-provider refusal remain open.
## Funded Gemini qualification, 2026-09-06
The maintainer authorized `openrouter:google/gemini-3.8-flash` with a provider-side
US$5 key limit expiring on 2026-09-07. The provider key endpoint confirmed both
constraints. Credentials remain in runtime configuration, not these receipts.
Synthetic readiness passed in 9.034 seconds: three streaming tool scenarios,
two context fixtures and multi-turn continuation. This supports the readiness
claims for Watch only and Ask first. It does not qualify autonomous fixes.
The first unhealthy-container run, `q-20260906-172952-3cccbf34`, detected the
correct fault and left the healthy control alone, but failed overall. The exact
Gemini route had no price entry, and the model attempted unsupported Docker
configuration access. The shared price table now records the reviewed standard
rates of US$0.75 input and US$3.75 output per million tokens for direct Gemini
and OpenRouter. Variant routes remain unknown. These introductory rates must be
reviewed on 2027-01-01. The query capability description now explicitly names
TrueNAS as the supported app-container configuration adapter and directs Docker
collected health/mount/port/network reads to `get`. Runtime permissions and
qualification gates are unchanged.
The following runs used the worker-built Pro binary
`74464e75977caf55cda092c8cf56c24967616c8c24d770872fbea5d86e31a1dc`,
core base `b0b39f00dc6685ad9ed63e8a6e91b954338073e4` plus the pricing and
capability-description changes, and canonical enterprise base
`3d9f4e3051d38027355a2a1f36b8c7f672a09b65`. The worker archive commit
`d9cb84e15d1acc377341129bdda5c28176e7128c` has identical contents for all
87 tracked enterprise files. The existing runner created disposable
resources on Tower, waited for normal collection, and used independent fault,
recovery and cleanup oracles.
| Case / run | Result | Evidence |
|---|---|---|
| Unhealthy, `q-20260906-174546-a7a9810b` | Pass | 9.709s detection phase, two tools, no failed/duplicate calls, healthy sibling unflagged. |
| Unhealthy, `q-20260906-174708-81f8d655` | Pass | 10.541s detection phase, exact unhealthy resource found. |
| Unhealthy, `q-20260906-174758-14deaa15` | Pass | 9.395s detection phase, exact unhealthy resource found. |
| Unhealthy, `q-20260906-174853-71cf894f` | Pass | 25.673s detection phase, exact unhealthy resource found. |
| Healthy mixed, `q-20260906-180144-381874a6` | Pass | 5.200s detection phase, no false findings. |
| Dependency, `q-20260906-175058-de5e350d` | Pass | Starting from only the client symptom, identified the stopped dependency and affected client. Investigation completed in 18.970s with three evidence calls and no mutation. |
| Storage, `q-20260906-175232-3ce6fbfa` | Fail before inference | Normal collection never converged to the required resource projection. No model diagnosis was attempted. |
| Approved restart, `q-20260906-175812-0593b9b7` | Fail before approval | Detection and investigation completed, but no exact action reference existed. The broker refused because Tower's Docker command agent was disconnected. Nothing executed. |
| Rejected restart, `q-20260906-175940-bff6992a` | Fail before rejection | No exact action was available to reject. This does not qualify rejected-action handling. |
Every listed run passed cleanup, including second-cleanup no-op and unchanged
inventory. Individual Watch run estimates were about US$0.007 to US$0.014.
Those scorecard estimates cover the Patrol detection phase, not the separate
investigation calls. Provider-side aggregate spend is the budget authority for
this temporary key. The fixed route price does not turn an estimate into a
reconciled bill or establish a hard Pulse budget for unpriced history.
Live qualification exposed two additional shared contract defects. The
investigation orchestrator logged action-broker refusal but completed the
record without retaining the error, leaving the model's captured-proposal prose
visible without the later refusal. The current enterprise change retains the
original diagnosis, persists the broker refusal as a failed investigation with
`needs_attention`, and creates no action reference. The product history adapter
also projected result-bearing transcript calls back into provider request calls,
dropping observed output and success/failure. The current core change uses one
shared transcript type for stored chat and product history, preserving the
separate explicit provider projection.
The live review exposed duplicate detail IDs, duplicate unformatted conclusions,
paused history made unclickable by the scheduling switch, and narrow filter/sort
overlap. The shared finding/investigation surfaces now preserve one detail target,
render sanitized Markdown once for identical summaries, retain distinct summaries,
keep history available while paused, and wrap controls. Result-bearing tool calls
use the same expandable evidence component as Assistant. Historical calls without
a result status retain their evidence without invented success or failure. An
investigation outcome of `cannot_fix` or `needs_attention` does not identify who
resolved the finding, so the shared resolution copy no longer infers manual review.
Private run receipts and source/binary bindings are under
`tmp/patrol-gemini-38/` in the workspace. The original failed runs remain failed.
Installed storage collection and a temporary command-enabled lab agent remain
prerequisites for real storage and approved/rejected recovery qualification.
No production agent has been replaced. Full action outcomes, remaining backup
coverage and independent volunteered Pro environments remain open.
### Final refusal and evidence-retention proof
Two further approved-remediation attempts remain **failed**:
`q-20260906-181839-6cbdc711` and `q-20260906-182952-6acfb739`. Both retained the
broker error separately from the original model summary, saved `status=failed`
and `outcome=needs_attention`, and created no action reference. Both passed
cleanup. The final run used Pro binary SHA256
`24de8c9ea0020067d298c489f0d99a272b5ec4b00afab7a40dff55aabb244061`,
including the proposal-response clarification, and detected its exact unhealthy
container with no false positives. Its saved investigation is
`48a16b05-50f0-4605-847c-0a71b3435975` for finding `ca3af29ac54d540f`.
The original failed scorecards have not been reclassified as action passes.
The restored history API retains observed outputs and explicit `success=false`
for historical `pulse_read` failures. Live browser review at `/patrol` exercises
successful query output, both `ACTION_NOT_ALLOWED` and `NO_AGENT` failures,
original diagnosis, one broker error, paused history and review focus return.
The settings proof at `/settings/pulse-intelligence/patrol` checks the exact
model, reviewed rates, synthetic readiness limits and reload. The current
source-bound browser receipt records desktop, intermediate and mobile results.
GET response fixtures cover unknown historical result status only, without
claiming new persisted model evidence or action execution.
Remaining qualification requires a current installed collector and a temporary
command-enabled lab agent. A Linux amd64 agent has been built on the worker,
SHA256 `ae2ed8b97709ec6e71af979c293ca9d3634662767ec1b59629baf4933c90cf5d`,
without installing it or changing Tower credentials. Tower's separate production
reporting agent is untouched. Any agent enrollment must use the canonical scoped
installation flow, preserve explicit identity and revoke temporary execution
access after qualification. Detection success does not satisfy this prerequisite
or the remaining backup and independent-environment cases.
## Installed agent and governed action qualification, 2026-09-06
The maintainer explicitly approved a temporary update and scoped command token
for Tower's separate development agent, followed by restoration. The installed
agent artifact was `ae2ed8b97709ec6e71af979c293ca9d3634662767ec1b59629baf4933c90cf5d`.
Both host and Docker modules reported running, and the command connection
registered the same agent identity. The production agent retained PID 752388.
The tests used Ask first with manual triggers. Scheduled Patrol ended paused in
Watch only. No autonomous-mode qualification is claimed.
The first installed storage attempt, `q-20260906-192435-f0d7eebf`, failed before
inference. Inspection established a qualification-client pagination defect:
`/api/resources?limit=1000` returned a maximum of 100 records from an inventory
of 104, leaving the exact worker on page two. This was not an absence of normal
collection. The client now follows the API pages and rejects partial results
when a later page fails. A regression finds an unhealthy resource beyond the
first 100, and the complete qualification package passes. Fault oracles and
score thresholds were not weakened. The corrected runner hash is
`a27f9ab0bf4786b670db3e8a989e974c8c43c5084d9524474ee694735d2e7df9`.
| Case / run | Automated result | Reviewed outcome |
|---|---|---|
| Approved restart, `q-20260906-193013-12d25545` | Pass | Correct unhealthy container, exact finding/investigation/resource and plan-hash binding, explicit approval before execution, completed restart, independent healthy/running readback and lifecycle verification. Detection 8.893s, fault-to-remediation phase total 78.490s. |
| Rejected restart, `q-20260906-193207-7095dad5` | Pass | Exact plan rejected, no restart, independent unchanged unhealthy fault until teardown. Detection 10.838s, fault-to-decision phase total 37.795s. |
| Storage, `q-20260906-193613-556ef23d` | Pass from existing scorecard | **Fails semantic diagnosis review.** Collection converged and logs exposed ENOSPC, but the model incorrectly asserted that Tower was out of disk space and implicated its array. Only an 8 MiB container tmpfs was exhausted. |
The approved action is `act_dcc3b52e5451810e49466daf9a6fccb0`, linked to finding
`566515f71129ce73` and investigation `aae05717-c9b0-4aac-8439-ba78f46c28e9`.
The rejected action is `act_ee0b736f0430e472e896a456ba3cb6eb`, linked to finding
`a17940552206e1ca` and investigation `955f4f47-0b6f-404b-b513-70a1d226f111`.
Each case passed independent teardown, second-cleanup no-op and restored
inventory. Per-case detection estimates were $0.012123, $0.012283 and $0.014915.
These exclude investigation calls and are not provider-account spend.
Storage remains unqualified. The model had the collected tmpfs mount and its
configured size. Its canonical-resource log call failed, the fallback host and
container log call succeeded, and its `df -h` command required approval. It then
promoted unrelated host/array warnings into a definite capacity diagnosis.
The scorecard's required terms and narrow forbidden phrases missed that false
claim. Its raw pass is retained as evidence of a qualification limitation, not
accepted as product success. The next storage slice needs canonical, authorized
filesystem-capacity evidence and explicit semantic review against the bounded
fault. Do not permit arbitrary commands merely to make that case pass, add a
benchmark-specific diagnosis rule, or treat identifier/phrase matches as proof
of causal correctness. Backup coverage and independent Pro environments remain
unqualified as well.
Real outcome review exposed stale durable records: the finding and investigation
could say `fix_verified` while the embedded product record still said
`fix_queued`. Action reconciliation now refreshes that record through the same
canonical builder used at investigation completion, preserving original model
prose, evidence and retained rollback. Read-time hydration repairs existing
records even when the top-level outcome already matches. Unchanged hydration
must not republish state or repeat outcome notifications. Resolved findings keep
their exact action-history link, and Assistant handoff preserves resolved status.
Investigation completion replaces an earlier partial action projection with its
final evidence. Subsequent action transitions preserve that completed evidence,
including impact and confidence that the current finding may no longer retain.
An intermediate proof build exposed that loss of retained impact. The regression
now preserves it, while already absent historical fields remain unassessed.
Final runtime and browser verification of these corrections is recorded below.
After qualification, both original development binaries were restored separately
because they differed: runtime `e5a2b60e52e35c37f68daa348c642757b56a64b40ca1d72f4db843ce69eb5db4`,
persistent `73c224dfd750c41b2cbd883c3ce7e352071862bc6a60e59fe6ed4dd3de312bc6`.
The original protected token was restored, temporary issued tokens were revoked
and checked absent, temporary backups were removed after comparison, and no
owned fault containers remained. The restored v6.2.0-rc.8 development agent
reported fresh telemetry. Its original token lacks command scope, so its command
connection is again absent by design. Production PID 752388 remained unchanged.
Final action-history proof uses Pro Darwin arm64 binary
`859d5d2de84cfd2264caa7dbcf5f080e3b1d06b5876df2c81dc0e00872f5e779`.
Worker proof passes the API action reconciliation selection and investigation,
record, rollback and early-projection completion regressions. The full API suite
passed in 310.353s before the final evidence-preservation refinement, followed
by the final targeted regressions. The three affected action component suites
pass 30 tests. The complete qualification package and pagination regressions
also pass. The exact final staged hook gates landing.
Playwright exercises `/patrol` Activity/All and both exact `/actions?action=...`
links above at 1440, 900 and 390 by 1000. Final-content checks cover resolved
record retention, outcome agreement, safety disclosure, completed/rejected
headers, planning-time copy, absent settled execution controls, independent
verification, policy/evidence/delivery disclosures, keyboard toggles, Escape,
close controls, deep-link reload, scroll fit and retained review focus. Actual
pixels were inspected at desktop, intermediate and phone sizes. Assistant
handoff opens the same finding with completed/rejected context and read-only
control. Provider readiness POST was deliberately blocked during that rendering
proof and no prompt was sent. Earlier browser attempts encountered an
intermittent bootstrap connection screen. The complete final matrix passed
after removing redundant immediate navigations from the proof driver, without
claiming a bootstrap fix. Source bindings are in
`frontend-modern/browser-verification.json`.
+3 -1
View File
@@ -10201,7 +10201,7 @@
},
{
"id": "patrol-assistant-customer-outcome-qualification",
"summary": "The explicit Patrol/Assistant redesign goal, contract, execution plan and source-bound evidence remain in docs/qualification/PATROL_ASSISTANT_CUSTOMER_JOURNEY.md. Model judgment owns diagnosis. Observations, hypotheses, proposals, executions and independently verified outcomes remain distinct. Assistant continues the same issue and governed action records. The recorded 2026-09-05 baseline has 127 paid installations, 71 with Patrol enabled and 23 with Assistant calls. Fourteen verified resolutions came from one installation. Schema 17 outcome/provider/cost fields had no adoption. Usage does not prove useful linked tasks or representative false-alarm, missed-problem or success rates. Shared risk, provenance, history, missing-access and diagnostic-continuity corrections have regression and named browser proof. Proposal promotion, duplicate causal inference, contextless evaluations and count-based diagnostic completion policy were removed. Independent Docker fault/oracle contracts qualify reproducible injection, negative controls and cleanup, not model competence or governed action outcomes. Integrated CI exposed retained-query performance, disk-probe ordering and route-label timing regressions. Their scoped corrections and qualification records landed through PR1928 and PR1929. PR1929 merged at cf98358c0eb46987a82def5776fa41db5f54210a with backend, frontend, benchmark, governance, CodeQL and all eight Core E2E shards passing. The current ordinary retained-history diagnosis still contradicts explicit temporal semantics and remains unqualified. Two additional read-only Assistant requests used claude-subscription:claude-opus-5 against run-owned containers on the monitored Tower host. The healthy request took 82.835s and seven tools, correctly recommending no action, but overstated absence of storage impact. The dependency request took 204.384s and sixteen tools with three failed reads. It identified the stopped dependency, preserved the missing command access and causal uncertainty, but overstated storage exclusion and recovery implications. Config reads incorrectly reported app-container not found after successful canonical gets. These single cases remain partial diagnosis evidence, not a qualification pass. The fixture deadline performed two-pass cleanup with unchanged original inventory. Post-answer fault readback and explicit recovery were not completed, so no action outcome is claimed. Captured responses exposed concurrent tool-ID merging, sibling approval removal and renderer mutation of shared evidence. The current scoped correction keeps supplied invocation IDs authoritative and stable message rows without deep transcript reconciliation. All 167 affected frontend tests pass. Final-source browser replay at /patrol, 1440/900/390x1000, preserves all seven and sixteen exact tool inputs/outputs. Controlled stream states verify concurrent progress, cancellation, failed completion and sibling approval retention without provider or infrastructure actions. Exact captures, failed reproductions, hashes and remaining limits are in the plan. This correction still requires its own scoped landing checks. Claude Max explicitly refused autonomous Patrol readiness. Cached refusal and API409 enforcement remain intact, with no bypass or repeated retry. Ordinary Assistant is not autonomous qualification. Approval for an alternate separately billed provider remains pending, and no paid request occurred. Reliable interpretation, the config-read contract, broader storage/backup and approved/rejected action outcomes remain required local work. Independent volunteered Pro environments remain a separate wider-readiness gate.",
"summary": "The redesign goal remains open. The plan, historical receipts and exact source bindings are in docs/qualification/PATROL_ASSISTANT_CUSTOMER_JOURNEY.md. Patrol owns investigation and Assistant continues the same issue. Observations, hypotheses, accepted proposals, execution and independently verified outcomes remain distinct. The recorded baseline of 127 paid installations, 71 Patrol-enabled, 23 with Assistant calls and fourteen verified resolutions from one installation does not establish representative customer success. Schema17 outcome/provider/cost fields had no adoption. Shared evidence/history/risk and removal of proposal-as-proof and proxy completion policy landed through PR1928/1929. PR1934 merged canonical tool/transcript identity. PR1935 contains subsequent canonical history, measurement-presence, command-connectivity, tmpfs context, exact Gemini pricing, retained broker errors and result-bearing transcript corrections. Enterprise broker refusal handling merged in PR22. Real Gemini Watch, healthy-control and client-to-dependency cases passed. With explicit authority for a temporary current development agent and scoped token, approved restart q-20260906-193013-12d25545 passed exact plan/origin binding, explicit approval, execution and independent recovery. Rejected restart q-20260906-193207-7095dad5 passed exact rejection and independent non-execution. All test resources were removed, original development binaries/token restored, temporary tokens revoked, scheduled Patrol paused in monitor mode and production agent PID preserved. Storage collection failure was traced to qualification pagination beyond the API page cap of 100 and corrected with full-package regression proof. Installed storage q-20260906-193613-556ef23d passed the existing scorecard but FAILED semantic diagnosis review: the model falsely attributed an exhausted container tmpfs to Tower/array capacity despite available mount configuration. A capacity read required approval. Storage remains unqualified, and lexical/identifier scoring must not be treated as causal correctness. Next storage work needs canonical authorized filesystem evidence and independent semantic review without benchmark-specific diagnosis rules or weakened command approval. Real browser review additionally exposed an embedded investigation record left fix_queued after verified recovery and hidden action history on resolved findings. Current action reconciliation refreshes the durable record through the canonical builder, preserves prose/evidence/rollback, repairs missed transitions without duplicate publication, and retains completed action history and resolved Assistant context. Final worker regressions and source-bound runtime/browser proof pass at 1440, 900 and 390 widths, including completed/rejected history and read-only Assistant handoff. The exact staged hook and PR1935 landing gate integration. Other residuals include canonical incident-memory listing/aliases and failed-read propagation, unsupported filters, typed compatibility ID lookup, legacy direction availability, Docker-host history, responsive mount details, backup coverage and broader model qualification. Independent volunteered Pro environments remain a wider-readiness gate. Autonomous modes remain unqualified. The earlier Claude refusal was not retried.",
"owner": "project-owner",
"status": "planned",
"recorded_at": "2026-09-05",
@@ -10215,6 +10215,7 @@
"ai-runtime",
"api-contracts",
"frontend-primitives",
"monitoring",
"patrol-intelligence",
"performance-and-scalability",
"unified-resources"
@@ -10386,6 +10387,7 @@
"ai-runtime",
"api-contracts",
"frontend-primitives",
"monitoring",
"patrol-intelligence",
"performance-and-scalability",
"unified-resources"
@@ -15,6 +15,17 @@
## Purpose
Docker mount reports include tmpfs configuration from `HostConfig.Tmpfs`
through the existing optional mount array. This adds collection evidence only.
It does not change admission, enrollment, execution permissions or agent
lifecycle authority. Existing agents continue to report their existing mount
coverage. Deploying an updated collector is a separate installed-path proof.
Docker block-I/O report presence fields are optional measurement metadata.
They preserve zero and omitted directions independently without changing report
admission, enrollment, identity, command permission or agent lifecycle state.
Older agents remain accepted, with ambiguous omitted zero counters unavailable.
### Automatic PVE association identity boundary
Host ingestion must not create a host-to-PVE or reciprocal PVE-to-agent link
@@ -7896,3 +7907,9 @@ positive matching evidence, provider replacement/return, write failure and
automatic versus unknown-provenance cleanup. State and config tests cover atomic
replacement and preservation of lifecycle evidence. These are synthetic local
proofs, not reporter confirmation or installed-release resolution of #1930.
Historical incident archives are explicit reads, not a collector lifecycle.
`internal/api/router.go` no longer starts an incident coordinator or a cached
metrics sampling loop. Organization teardown drops the archive reference without
saving or deleting recordings. Alert observation and recovery continue through
the alert manager and canonical resource timeline.
@@ -25,6 +25,104 @@ that same result. Successful reads retain their content and execution provenance
## Purpose
Action reconciliation refreshes the durable product investigation record from
the authoritative session/action even when the finding outcome already matches.
The same builder owns initial completion and later refresh. Completion replaces
an early action projection with the final investigation evidence. Later action
refresh preserves original prose, impact, confidence, evidence and rollback. Unchanged hydration is a no-op, and a
record-only repair does not repeat outcome notifications. Resolved findings
retain the canonical action-history link and resolved status in Assistant context.
Live approved and rejected recovery cases pass, but storage's lexical scorecard
pass fails semantic review because an exhausted container tmpfs was incorrectly
attributed to host capacity. Storage and broader product qualification stay open.
The live qualification client follows the canonical resource API's pagination.
The API caps each page at 100, so a larger requested limit cannot establish a
complete inventory. A later-page failure returns an error rather than partial
inventory. Regression proof covers an unhealthy resource beyond the first 100
and failure while reading a later page. This changes collection coverage, not
fault oracles, model context policy or outcome scoring.
Stored chat and product history share the result-bearing `TranscriptToolCall`
contract. API adapters preserve observed output and the explicit success/error
bit. Only provider-request projections remove those display fields. A failed
read must not become an invocation with no visible result on the way to Patrol
or Assistant history. The adapter regression includes a `NO_AGENT` result and
`success: false`, and provider serialization retains its existing narrower shape.
Capturing a typed proposal does not create an action. The proposal response
discloses that broker validation is still pending, without forcing the model to
stop investigating. If the broker later refuses submission, the enterprise
orchestrator retains the model's diagnosis unchanged and records the broker
error as a failed investigation needing attention, with no action reference.
The real disconnected-agent case must remain unsuccessful until its actual
transport prerequisite is satisfied. A successful model turn or recorded
proposal is not approval, execution or recovery.
The shared investigation review renders sanitized Markdown and does not repeat
an identical persisted/fetched conclusion or error. Distinct evidence remains
visible. Pausing scheduled Patrol does not disable history review. The review
control has one detail target and returns keyboard focus when closed. Merged tool
results use Assistant's shared expandable evidence component. Historical calls
without an explicit result bit do not gain an inferred success/failure state.
Resolution copy cannot infer manual review from `needs_attention` or `cannot_fix`.
The shared pricing table includes reviewed standard Gemini 3.8 Flash rates for
the exact direct and OpenRouter routes. OpenRouter variants and aliases remain
unpriced until independently reviewed. Rates carry the review date and are
estimates, not reconciled provider charges. The introductory rates require a
new review on 2027-01-01. `TestGemini38FlashReviewedRoutePricing` covers real
qualification token counts and request-route preservation, and
`TestGemini38OpenRouterPricingDoesNotGuessVariantRates` preserves unknown variants.
The canonical query tool describes the app-container configuration boundary
explicitly: TrueNAS supports `config`, while Docker/Podman expose their collected
health, mounts, ports and networks through `get`. This communicates the existing
adapter contract to the model. It does not add configuration access, suppress
tool errors or weaken qualification gates. The existing
`TestAppContainerConfigObservationContract` retains unsupported-adapter and
provider/identity boundaries.
Shared app-container query mount evidence preserves native type, source,
destination, options and canonical read/write access. Compound options such as
`ro,noexec` cannot become writable through string equality heuristics. Both the
canonical provider and typed read-state projection preserve the same fields.
Configured size in mount options does not establish filesystem usage or free
space. `TestQueryPreservesMountConfigurationEvidence` covers these contracts.
The ordinary live storage case still failed diagnosis by excluding resource
pressure without capacity evidence. Mount fidelity alone does not qualify model
interpretation. Incident-record lookup across canonical and legacy Docker IDs
and typed compatibility lookup of canonical IDs remain explicit identity gaps.
Shared query projections name command transport explicitly through
`command_agent_connected`, `node_command_agent_connected` and the corresponding
topology counts. These observations do not establish monitoring freshness or
installation state. A topology built without a connection snapshot omits command
flags, execution hints and connected counts. An observed empty snapshot preserves
false/zero. Assistant's inventory seed carries that same absence semantics.
The parent node's connection cannot become a direct guest connection merely
because provider placement names that node. Existing command routing, control,
approval and invocation enforcement remain authoritative. A `can_execute` hint
reflects connected transport with control enabled, not approval for an operation.
`TestCommandConnectivityDoesNotReplaceMonitoringEvidence`,
`TestTopologyOmitsUnobservedCommandConnections` and
`TestAssistantInventoryDoesNotInventCommandConnectionObservations` cover these
projection and continuity boundaries. Existing persisted tool records are not
rewritten, and this contract does not qualify model diagnosis or recovery.
Native app-container configuration reads resolve identity, provider and placement
from current canonical inventory. Optional session discovery cannot fabricate a
not-found result or replace current placement with a stale execution target.
Query restrictions on both the supplied reference and canonical identity are
checked before registration, and an existing session's allowed actions are not
expanded by a read. Unsupported adapters, missing providers, incomplete placement
and nil provider observations retain known resource identity and an explicit
unavailability reason with the shared tool error bit. They cannot count as a
successful configuration read. Actual inventory absence remains distinct.
`TestAppContainerConfigObservationContract` exercises these boundaries with
strict resolution enabled. This read correction does not relax action or native
log validation and does not qualify autonomous diagnosis or recovery.
The published Patrol qualification schema must accept the fault injectors used
by the executable catalogue. `TestCatalogFaultInjectorsMatchPublishedSchema`
checks the actual scenario faults against the schema enum, including the
@@ -797,6 +895,7 @@ cheap local detection into model-owned diagnosis and governed action.
31. `internal/agentcapabilities/tool_names.go` shared with `api-contracts`: the Pulse Intelligence registry tool-name vocabulary is both the native Assistant execution/display contract and the canonical API/agent tool identity contract for MCP-facing external-agent adapters.
32. `internal/agentcapabilities/tool_response.go` shared with `api-contracts`: the shared tool response envelope, tool error-code vocabulary, and tool-result error-code and verification evidence parsers are both the Assistant structured tool-result contract and the canonical API/agent branching contract for Pulse Intelligence tool failures, recovery tracking, and write self-verification.
33. `internal/agentcapabilities/tool_result.go` shared with `api-contracts`: the Pulse Intelligence shared tool-result content/result envelope, structuredContent projection, result constructors, HTTP response-to-result mapping, text projection, and result interpretation helpers are both the Assistant registry result contract and the canonical API/agent result projection contract for governed tool outcomes.
34. `internal/agentcapabilities/transcript.go` shared with `api-contracts`: Stored Assistant tool results and product history share one result-bearing transcript contract, with an explicit narrower provider-request projection. Observed failures and absent historical result status must survive the API boundary.
34. `internal/agentcapabilities/types.go` shared with `api-contracts`: the agent capabilities manifest wire type, manifest-owned external-adapter surface tool contract field, capability display title and structured output schema fields, approval-policy vocabulary, capability governance normalization, and tool-governance descriptor shape are both the canonical API payload contract and the AI runtime projection contract for Pulse Assistant and MCP-facing agent tools.
35. `internal/agentcapabilities/workflow_prompt.go` shared with `api-contracts`: the Pulse Intelligence workflow prompt catalogue, manifest-owned `workflowPrompts` projection, MCP prompt title projection, presentation kind hints, shared resource-context and finding argument vocabulary, Patrol issue-handling capability gating, argument validation, and manifest-gated shared prompt rendering rules are both the AI runtime starter contract for Assistant-compatible surfaces and the canonical API/agent prompt projection contract for MCP-facing clients.
36. `internal/api/ai_handler.go` shared with `api-contracts`: Pulse Assistant handlers are both an AI runtime control surface and a canonical API payload contract boundary.
@@ -3565,6 +3664,22 @@ has a single definition in the canonical resource contract.
## Completion Obligations
The `pulse_knowledge` incidents action reads the organization-pinned canonical
resource timeline used by resource history and Assistant handoffs. It preserves
resource identity, observation and optional occurrence time, source and event
metadata. Reads use explicit observation-time bounds and a bounded event count
with truncation disclosure. Empty retained history is not continuous healthy
coverage, and an unavailable or failed history store is a failed tool read.
Legacy recording IDs are isolated archive lookups bound to the requested
resource. Their recorder timestamps cannot establish source measurement time.
The primary history path must not restore the legacy recorder as a parallel
incident authority or derive fresh history by resampling cached metrics.
`TestIncidentHistoryRetainsCanonicalEvidence` uses SQLite lifecycle records,
time/resource negative controls, missing occurrence time and bounded reads.
The corresponding unavailable/invalid and archive tests cover failure semantics
and resource isolation. Live model interpretation remains governed by the
customer journey qualification plan.
Every per-organization Assistant or legacy AI service that can discover or
dispatch through the host-agent command transport must receive an
organization-pinned command-server view. A tenant service must never enumerate
@@ -8061,3 +8176,30 @@ unchanged pre-existing inventory. The ordinary package run skips live Docker
work unless explicitly enabled. A direct fixture restart is teardown and must
never be counted as a Pulse approval, execution, rejection or outcome. Model-led
and canonical-action qualification remain separate required evidence.
### Legacy incident recording retirement
`internal/metrics/incident_archive.go` owns the historical recording format and
explicit, organization-pinned archive reads. `IncidentArchiveProvider` exposes
only a resource-bound window lookup with an error result. The primary
`pulse_knowledge` incident action reads the canonical resource timeline. An
explicit `window_id` reads saved legacy observations and labels their timestamp
and historical-status limits. It must never start a recorder or infer source
freshness from a recording timestamp. The disconnected incident coordinator,
fleet sampling adapter, timer loop, retention writer and duplicate tool adapter
are retired. There is no replacement incident scheduler or diagnosis policy.
The archive reader preserves saved timestamps, resource labels, metadata and
summary values, including records older than the former retention period. The response explicitly
identifies the nanosecond encoding of the historical `summary.duration_ms` field. It
distinguishes unavailable archives, failed reads and absent exact resource/window
pairs. Proof lives in `internal/metrics/incident_archive_test.go` and the
registered-tool cases in `internal/ai/tools/incident_history_test.go`.
The legacy incidents listing still exposes incident memory and is not a complete
canonical incident query. Its old sampler-derived `active_count` is now null with
`active_count_status=not_measured`. Canonical-only events, alias-aware listing,
query bounds and projection-read errors remain an explicit modernization gap in
`patrol-assistant-customer-outcome-qualification`. The shared resource timeline
remains the evidence owner. Do not invent another incident lifecycle to repair
this listing.
@@ -20,6 +20,24 @@
## Purpose
Action reconciliation refreshes the durable product investigation record from
the authoritative session/action even when the finding outcome already matches.
The same builder owns initial completion and later refresh. Original prose,
evidence and retained rollback survive. Unchanged hydration is a no-op, and a
record-only repair does not repeat outcome notifications. Resolved findings
retain the canonical action-history link and resolved status in Assistant context.
Live approved and rejected recovery cases pass, but storage's lexical scorecard
pass fails semantic review because an exhausted container tmpfs was incorrectly
attributed to host capacity. Storage and broader product qualification stay open.
Product history retains the stored result-bearing `TranscriptToolCall` contract,
including observed output and an optional success bit. A false bit is retained,
and an absent historical result bit stays absent. Provider requests use the
explicit narrower projection. The history adapter cannot project away evidence
by treating a display transcript as request arguments. The frontend Patrol API
extends the shared Assistant tool-call shape, and transport regressions pin
failed output and unknown historical status across the message envelope.
The internal Patrol bridge preserves explicit execution limits, scoped tool
allowlists and execution identity. The retired unmatched-signal evaluator no
longer contributes a signal-count-derived successful-report budget. Diagnosis
@@ -1683,6 +1701,7 @@ payload shape change when the portal presents compact client rows.
57. `internal/agentcapabilities/tool_names.go` shared with `ai-runtime`: the Pulse Intelligence registry tool-name vocabulary is both the native Assistant execution/display contract and the canonical API/agent tool identity contract for MCP-facing external-agent adapters.
58. `internal/agentcapabilities/tool_response.go` shared with `ai-runtime`: the shared tool response envelope, tool error-code vocabulary, and tool-result error-code and verification evidence parsers are both the Assistant structured tool-result contract and the canonical API/agent branching contract for Pulse Intelligence tool failures, recovery tracking, and write self-verification.
59. `internal/agentcapabilities/tool_result.go` shared with `ai-runtime`: the Pulse Intelligence shared tool-result content/result envelope, structuredContent projection, result constructors, HTTP response-to-result mapping, text projection, and result interpretation helpers are both the Assistant registry result contract and the canonical API/agent result projection contract for governed tool outcomes.
60. `internal/agentcapabilities/transcript.go` shared with `ai-runtime`: Stored Assistant tool results and product history share one result-bearing transcript contract, with an explicit narrower provider-request projection. Observed failures and absent historical result status must survive the API boundary.
60. `internal/agentcapabilities/types.go` shared with `ai-runtime`: the agent capabilities manifest wire type, manifest-owned external-adapter surface tool contract field, capability display title and structured output schema fields, approval-policy vocabulary, capability governance normalization, and tool-governance descriptor shape are both the canonical API payload contract and the AI runtime projection contract for Pulse Assistant and MCP-facing agent tools.
61. `internal/agentcapabilities/workflow_prompt.go` shared with `ai-runtime`: the Pulse Intelligence workflow prompt catalogue, manifest-owned `workflowPrompts` projection, MCP prompt title projection, presentation kind hints, shared resource-context and finding argument vocabulary, Patrol issue-handling capability gating, argument validation, and manifest-gated shared prompt rendering rules are both the AI runtime starter contract for Assistant-compatible surfaces and the canonical API/agent prompt projection contract for MCP-facing clients.
62. `internal/api/access_control_handlers.go` shared with `organization-settings`: RBAC role and user-assignment handlers are both an organization settings control surface and a canonical API payload contract boundary.
@@ -10670,3 +10689,18 @@ reasoning and real remediation in
`docs/qualification/PATROL_ASSISTANT_CUSTOMER_JOURNEY.md`. The repeatable browser
proof is `scripts/check-patrol-assistant-journey.mjs`. A passing scripted
response does not establish a useful customer outcome or model qualification.
### Explicit historical incident archives
`internal/api/router.go` pins a read-only legacy archive to each organization's
Assistant service. Its native `IncidentArchiveProvider` capability requires both
resource and window identifiers and propagates read errors. No archive setup or
shutdown writes files or starts sampling. The tool preserves stored metadata and
anomalies and marks historical recording status as historical.
`GET /api/ai/incidents` retains the `active_count` key as null and adds
`active_count_status=not_measured` in every response. Incident memory, an empty
result and unavailable services cannot establish a current count. The old
coordinator never received production alert callbacks, so its zero was not a
measurement. The legacy listing's broader canonical query and read-error
modernization remains open under the customer-outcome qualification gap.
@@ -20,6 +20,14 @@
## Purpose
Disk I/O presentation preserves each observed direction independently. Shared
formatting renders a missing rate as a dash and measured idle as numeric zero.
Partial observations cannot form a complete throughput total for sorting or
comparison. Machines column preferences must preserve an explicit user choice
across the first reload, including default-hidden migrations. Final-source
browser proof covers Docker host details, Machines column selection and tooltip
focus/dismissal at desktop, intermediate and narrow widths.
Overview delivery diagnoses use latest-started refresh ownership. Older bulk
responses cannot overwrite newer card notification status, and an empty active
alert set invalidates outstanding reads. Disposal also prevents updates. Failed
@@ -17,6 +17,28 @@
## Purpose
Docker mount collection preserves both native `Mounts` records and entries
reported only in `HostConfig.Tmpfs`. Existing reported destinations remain
authoritative. Additional tmpfs destinations are ordered deterministically,
retain their options and read/write setting, and use the existing mount report
shape. Configured tmpfs size is configuration, not measured used/free space.
`TestCollectContainerPreservesTmpfsMounts` reproduces a live tmpfs-only inspect
shape and covers mixed mounts, read-only options, overlap and absent host config.
Docker collection records read and write counter presence independently,
including explicit zero, in optional report fields. Older reports without those
fields establish only positive counters. Container reports propagate this
presence to the shared rate tracker. An omitted block-I/O payload and the first counter sample
produce no rate history. Unchanged observed counters produce measured zero,
including after an omitted report. Container writable/root layer sizes never
produce capacity usage history. The ingestion regression is
`internal/monitoring/docker_metric_presence_test.go`. This changes measurement
projection only and grants no agent lifecycle authority.
The shared resource-to-browser conversion preserves each optional I/O rate
independently. Missing directions are omitted from JSON, while measured zero
remains numeric zero. No aggregate presence flag may fabricate its sibling
direction. `TestResourceDiskIOWirePreservesAbsentDirection` pins that wire path.
### PBS datastore alert evaluation belongs to the live poll
After publishing freshly polled PBS datastore storage rows, the poller invokes
@@ -164,6 +186,15 @@ fixtures for 500, 502 quoting 403, 503 quoting 404, and genuine 401/403/404.
These tests prove cache retention/removal, not installed PBS wake, service
restart, or notification receipt.
Docker alert lifecycle events pass through the shared resource history identity
writer. A full Docker source reference must reach the same canonical container
history as inventory changes, including recovery after inventory removal and
restart. Existing alert lifecycle event IDs remain unchanged so replay cannot
duplicate retained events. Same-name containers and abbreviated IDs must not
join another container's history. The real alert-manager callback path is
covered by `TestDockerAlertTimelineUsesCanonicalHistoryIdentity` in
`internal/monitoring/monitor_alert_handling_test.go`.
TrueNAS native alert projection preserves the trimmed, uppercase provider level in ResourceIncident.NativeSeverity. INFO and NOTICE retain the same canonical monitor risk; consumers must not lose their distinct actionability when projecting provider evidence. Native CRITICAL, ALERT, and EMERGENCY all project to canonical critical severity; EMERGENCY must not be discarded as unknown or make a still-active condition appear recovered. WARNING remains warning, and INFO and NOTICE remain informational at this projection boundary.
Verification: `TestIncidentProjectionPreservesNativeSeverity` in `internal/truenas/provider_pool_health_contract_test.go` covers all seven native levels and case/whitespace normalization. `TestTrueNASNativeSeverityDispatch` in `internal/alerts/truenas_native_dispatch_test.go` verifies downstream INFO suppression, NOTICE preservation, notification severity, duplicate-poll retention, and confirmed recovery callback identity. The TrueNAS lifecycle tests in `internal/alerts/unified_incidents_test.go` require repeated EMERGENCY evidence to interrupt recovery confirmation. These are fixture-based projection and manager checks, not appliance ingestion or external notification-provider receipt proof.
@@ -15,6 +15,26 @@
## Purpose
Action reconciliation refreshes the durable product investigation record from
the authoritative session/action even when the finding outcome already matches.
The same builder owns initial completion and later refresh. Original prose,
evidence and retained rollback survive. Unchanged hydration is a no-op, and a
record-only repair does not repeat outcome notifications. Resolved findings
retain the canonical action-history link and resolved status in Assistant context.
Live approved and rejected recovery cases pass, but storage's lexical scorecard
pass fails semantic review because an exhausted container tmpfs was incorrectly
attributed to host capacity. Storage and broader product qualification stay open.
Pausing the Patrol schedule does not disable investigation history review.
The finding review control owns one detail target and regains keyboard focus on
close. Investigation conclusions use sanitized Markdown, deduplicate identical
stored/fetched text, and retain distinct evidence. Merged tool results render
through Assistant's shared expandable evidence component with their explicit
success/failure bit. Unknown historical status stays unknown. A broker refusal
remains visible separately from the original diagnosis and cannot be presented
as an accepted action. An attention/cannot-fix outcome does not establish manual
review or identify who resolved a later finding.
Detection retains one model conversation for evidence gathering and finding
decisions. Recording one finding does not establish diagnostic sufficiency or
remove its evidence tools. Missing assessments and provider failures remain
@@ -15,6 +15,20 @@
## Purpose
The Docker/app-container history families `dockercontainer` and `docker` use
separate physical `.observed` series for new disk capacity and block-I/O
measurements. Older disk series lack the required presence/capacity semantics
and remain stored unchanged, but retained reads exclude them from current
evidence. All four shared read APIs expose corrected series under the existing
public metric names. Valid capacity from other providers sharing this storage
family remains writable. `NormalizedSeriesKey` describes physical storage for
coverage/backfill matching. Rollups aggregate each physical generation separately.
Projection happens once per returned series, not per observation, and adds no
query, schema migration, or per-row work for other resource families.
`pkg/metrics/store_docker_observation_contract_test.go` pins legacy coexistence,
zero retention, selected/fleet read parity, rollup separation and unaffected
resource families.
Retained reads use one shared query contract in `pkg/metrics/store.go` for
`Query`, `QueryAll`, `QueryAllBatch` and `QueryMetricTypesBatch`. A non-empty
preferred resolution no longer hides a newer raw tail, an older uncovered
@@ -3116,3 +3130,9 @@ candidate on the same worker with alternating samples, and retains full-route
and middleware controls. Identical source or instruction sequences alone do not
prove identical timing. The recorded final ten-pair check passes the unchanged
time/bytes/allocation gate with no adjacent request-path regression.
The disconnected incident recorder and its fleet metrics adapter are retired.
Router initialization no longer launches their five-second cached-metrics loop
or allocates per-resource pre-incident buffers. Explicit legacy archive reads
are lazy and bounded to the existing 16 MiB file limit. Canonical resource
history supplies current diagnostic evidence without a second sampling loop.
@@ -779,6 +779,14 @@
"api-contracts"
]
},
{
"path": "internal/agentcapabilities/transcript.go",
"rationale": "Stored Assistant tool results and product history share one result-bearing transcript contract, with an explicit narrower provider-request projection. Observed failures and absent historical result status must survive the API boundary",
"subsystems": [
"ai-runtime",
"api-contracts"
]
},
{
"path": "internal/agentcapabilities/types.go",
"rationale": "the agent capabilities manifest wire type, manifest-owned external-adapter surface tool contract field, capability display title and structured output schema fields, approval-policy vocabulary, capability governance normalization, and tool-governance descriptor shape are both the canonical API payload contract and the AI runtime projection contract for Pulse Assistant and MCP-facing agent tools",
@@ -1611,6 +1619,7 @@
"internal/config/host_continuity_test.go",
"internal/models/metrics_types_test.go",
"internal/monitoring/availability_probe_agent_test.go",
"internal/monitoring/docker_metric_presence_test.go",
"internal/monitoring/monitor_host_agent_removal_lifecycle_test.go",
"internal/monitoring/monitor_host_agents_test.go",
"scripts/installtests/agent_state_dir_lifecycle_test.go",
@@ -2073,6 +2082,7 @@
"internal/api/ai_intelligence_handlers.go",
"internal/config/ai.go",
"internal/config/patrol_autopilot_persistence.go",
"internal/metrics/incident_archive.go",
"pkg/aicontracts/action_broker.go",
"pkg/aicontracts/fix_execution.go",
"pkg/aicontracts/investigation.go",
@@ -2087,6 +2097,20 @@
"exact_files": [],
"require_explicit_path_policy_coverage": true,
"path_policies": [
{
"id": "legacy-incident-archive",
"label": "Explicit read-only legacy recording archive proof",
"match_prefixes": [],
"match_files": [
"internal/metrics/incident_archive.go"
],
"allow_same_subsystem_tests": false,
"test_prefixes": [],
"exact_files": [
"internal/ai/tools/incident_history_test.go",
"internal/metrics/incident_archive_test.go"
]
},
{
"id": "retained-metric-evidence",
"label": "retained metric evidence and observation coverage proof",
@@ -2186,9 +2210,11 @@
],
"exact_files": [
"internal/api/ai_handler_test.go",
"internal/api/ai_handlers_investigation_additional_test.go",
"internal/api/ai_handlers_more_test.go",
"internal/api/ai_handlers_patrol_actions_additional_test.go",
"internal/api/ai_handlers_test.go",
"internal/api/ai_intelligence_handlers_remediation_additional_test.go",
"internal/api/ai_intelligence_handlers_test.go",
"internal/api/issue1640_readiness_gate_test.go",
"internal/api/issue1640_readiness_transport_test.go",
@@ -3222,6 +3248,7 @@
"exact_files": [
"frontend-modern/src/types/api.ts",
"internal/api/action_runner_credentials_test.go",
"internal/api/ai_handlers_investigation_additional_test.go",
"internal/api/ai_handlers_more_test.go",
"internal/api/ai_handlers_patrol_actions_additional_test.go",
"internal/api/alerting/external_probe_notifications_test.go",
@@ -5945,6 +5972,7 @@
"test_prefixes": [],
"exact_files": [
"internal/config/host_continuity_test.go",
"internal/monitoring/docker_metric_presence_test.go",
"internal/monitoring/issue1485_unraid_lifecycle_test.go",
"internal/monitoring/issue1595_collection_trust_test.go",
"internal/monitoring/monitor_docker_test.go",
@@ -6066,6 +6094,8 @@
"internal/dockeragent/agent_collect_test.go",
"internal/dockeragent/agent_cpu_test.go",
"internal/dockeragent/agent_internal_test.go",
"internal/dockeragent/blockio_presence_test.go",
"internal/dockeragent/collect_tmpfs_test.go",
"internal/dockeragent/swarm_coverage_test.go"
]
},
@@ -6125,13 +6155,15 @@
"internal/models/issue1639_pbs_collision_test.go",
"internal/models/metrics_types_test.go",
"internal/models/state_host_test.go",
"internal/monitoring/docker_metric_presence_test.go",
"internal/monitoring/issue1595_collection_trust_test.go",
"internal/monitoring/monitor_full_coverage_test.go",
"internal/monitoring/monitor_host_agent_removal_lifecycle_test.go",
"internal/monitoring/monitor_host_agents_test.go",
"internal/monitoring/monitor_package_updates_test.go",
"internal/unifiedresources/adapter_coverage_test.go",
"internal/unifiedresources/registry_test.go"
"internal/unifiedresources/registry_test.go",
"pkg/agents/docker/blockio_presence_test.go"
]
},
{
@@ -7005,6 +7037,7 @@
"exact_files": [
"pkg/metrics/store_additional_test.go",
"pkg/metrics/store_bench_test.go",
"pkg/metrics/store_docker_observation_contract_test.go",
"pkg/metrics/store_query_plan_test.go",
"pkg/metrics/store_slo_test.go"
]
@@ -7132,6 +7165,7 @@
"allow_same_subsystem_tests": false,
"test_prefixes": [],
"exact_files": [
"frontend-modern/src/components/Infrastructure/__tests__/infrastructureSelectors.test.ts",
"frontend-modern/src/components/Infrastructure/__tests__/UnifiedResourceTable.performance.contract.test.tsx",
"frontend-modern/src/components/Infrastructure/__tests__/unifiedResourceTableStateModel.test.ts",
"frontend-modern/src/components/Infrastructure/__tests__/useTableWindowing.test.ts",
@@ -8078,6 +8112,7 @@
"internal/unifiedresources/adapter_coverage_test.go",
"internal/unifiedresources/adapters_test.go",
"internal/unifiedresources/ceph_pool_health_contract_test.go",
"internal/unifiedresources/history_identity_test.go",
"internal/unifiedresources/host_storage_cleanup_test.go",
"internal/unifiedresources/monitor_adapter_read_state_test.go",
"internal/unifiedresources/views_test.go"
@@ -8475,6 +8510,7 @@
"exact_files": [
"internal/monitoring/issue1595_collection_trust_test.go",
"internal/unifiedresources/availability_link_test.go",
"internal/unifiedresources/history_identity_test.go",
"internal/unifiedresources/kubernetes_registry_test.go",
"internal/unifiedresources/pbs_pmg_registry_test.go",
"internal/unifiedresources/registry_merge_policy_test.go",
@@ -8512,6 +8548,19 @@
"internal/unifiedresources/action_policy_provenance_test.go"
]
},
{
"id": "resource-history-identity",
"label": "resource history identity and authority isolation proof",
"match_prefixes": [],
"match_files": [
"internal/unifiedresources/history_identity.go"
],
"allow_same_subsystem_tests": false,
"test_prefixes": [],
"exact_files": [
"internal/unifiedresources/history_identity_test.go"
]
},
{
"id": "unified-resource-runtime-support",
"label": "unified resource runtime support proof",
@@ -2809,3 +2809,10 @@ optional link to the configured public URL. It never includes resource names,
finding text, commands, evidence, or model names, and it uses the tenant's
existing email configuration and recipients under the admin-only report
schedule routes.
Explicit legacy incident archive reads use the organization's pinned data path
and exact resource/window identifiers. They retain bounded regular-file and
symlink checks and expose no enumeration, sampling or writing capability.
Removing the disconnected recorder does not alter alert, action approval or
operator authority. An unrelated organization receives no default archive
fallback.
@@ -21,6 +21,14 @@
## Purpose
Container image-layer sizes do not establish filesystem capacity. Docker
resource metrics omit that invalid ratio, and retained queries exclude legacy
ambiguous disk observations while preserving new valid provider measurements.
The shared REST resource projection retains missing read/write directions
independently of measured zero. An idle I/O rate cannot establish available
capacity, backup coverage or recoverability. This correction adds no recovery
authority or verified recovery outcome.
Patrol consumes storage evidence in the original diagnostic conversation. A
saved finding does not close storage-read authority before the explicit run
limit. Missing backup or recovery evidence remains unknown, and neither finding
@@ -5993,3 +6001,11 @@ Unlike resource reports, a `patrol_digest` schedule run produces no generated
file under the tenant `reports` directory and never calls the retention prune;
the email body is the only artifact. Recovery and retention state are
unaffected.
Legacy `incident_windows.json` recovery is read-only through
`internal/metrics/incident_archive.go`. Construction does not read the file, and
explicit reads neither expire records nor rewrite contents or permissions.
Malformed, oversized, symlink and non-regular archive paths fail visibly. Missing
archives remain distinguishable from a valid archive with no matching window.
The original recording times and historical status must never establish current
source freshness or active recording.
@@ -15,6 +15,25 @@
## Purpose
Action review distinguishes the recorded plan from live or executed facts. The
shared decision packet labels its state and expiry as planning-time evidence,
including when opened from a resolved Patrol finding. Potential blast radius
does not claim every related resource was affected. Actual execution and
verification remain in the recorded outcome section. The review header uses the
shared action-state presentation, so rejected actions remain identifiable even
without an execution receipt. The rejected-state regression and live
completed/rejected deep-link browser proof cover these historical journeys.
Docker container writable/root layer bytes describe image composition, not
used/total filesystem capacity, and cannot populate `ResourceMetrics.Disk`.
Optional valid block-I/O rate pointers preserve measured zero. Missing, negative
and non-finite rates remain unavailable. Raw layer metadata remains available.
`TestMetricsFromDockerContainerDistinguishesAbsentAndIdleIO` and the container
I/O projection proof in `internal/unifiedresources/metrics_test.go` pin this
boundary.
The frontend resource contract likewise makes read and write rates independently
optional, preserving this distinction through current-history labels.
Physical disk source freshness preserves the collector-authored
`expectedUpdateIntervalSeconds` alongside the actual last observation. Registry
ingest, merge, cloning and typed disk views preserve it. Staleness uses the
@@ -4802,6 +4821,17 @@ recent-change slice plus facet counts it actually renders. The store now also
owns a `resource_changes` persistence table with `RecordChange` and
`GetRecentChanges` methods so change history is queryable by canonical ID and
time window.
Docker alert source references containing an exact full container ID resolve at
`MonitorAdapter.RecordChange` through the current registry, then a retained
history binding, then the deterministic source-specific container identity.
Names and shortened IDs cannot establish this binding. `history_identity.go`
owns a history-only alias index in the organization-scoped resource store.
Legacy event rows retain their IDs, resource references and timestamps. Reads
expand aliases and canonical predecessor eras without changing operator state,
action requests, approvals, links or exclusions. Separate monitor, API and
Assistant store handles must see current persisted aliases. Missing identity
storage is an error, not evidence of empty history. Retention removes an alias
only after neither identity has retained journal records.
That same shared timeline vocabulary now includes the `activity` change kind
for provider-read breadcrumbs such as VMware tasks and events, plus the
`vmware_adapter` source-adapter token for canonical provenance drill-down.
+51 -11
View File
@@ -1,7 +1,7 @@
{
"version": 1,
"base_sha": "df7eca7913c72f913fe196f37acec4d3b104c726",
"verified_at": "2026-09-06T19:40:26Z",
"base_sha": "3853124a391ed451f75402f07951c0142e6b5ad8",
"verified_at": "2026-09-06T20:40:07.804772Z",
"result": "passed",
"changed_paths": [
"frontend-modern/src/components/DemoBanner.tsx"
@@ -9,25 +9,65 @@
"content_sha256": {
"frontend-modern/src/components/DemoBanner.tsx": "13fd8dea552d97f1df866438f7ba38eec01f25907dc68a0bec672e1793c3fc97"
},
"backend_content_sha256": {
"internal/agentcapabilities/transcript.go": "356c4ca201470407988ff9b2c1fb848619ed38e9d8db844e7390adce0f93ec19",
"internal/ai/chat/types.go": "1f624daf7e511eb971e2d81b787dcd72b2a82d0e0ecd775aa4e580c59d915382",
"internal/ai/service.go": "25dca2a70a985e8ab07443444a9f05f01a569c62ff7bc068a0b587d500831de6",
"internal/api/chat_service_adapter.go": "6d0ab14456b1901c5020a408de95796ece1aa8d1057dd58737ea2637777d6ccb",
"internal/ai/cost/pricing.go": "7b64bc881311ee0a7c1c8a319fcd162974f5a307ced1679e614ff20ac1ac532c",
"internal/ai/tools/tools_query.go": "3e074b204a8c8c4f8b66eaf147c269a2e57739bfc69d03ea36ee8c47cf8b908a",
"internal/ai/tools/tools_propose.go": "43d720c78a010b72f53f7e4e9e1b7b2a763e7f1ac555edb921b7cdca3e82fd51",
"internal/ai/patrol_findings.go": "d5eeb386f025cca338ac328b1d4ec7a2bf51150023454e2b61670f196fcf0356",
"internal/api/ai_handlers.go": "f8c9b24dc684346da4bddab066540fd43ca4999f5fd379c4dc34b5293f78394c",
"internal/api/patrol_action_reconciliation.go": "bc5da1a8050b94271dcdc88841a0ce3329e1773bd01c8746068391a71740ffa2",
"internal/monitoring/monitor.go": "63c8ef4867c07c96b4cd7c4316b5b8646f11b7e9e43a5233471956f76f33d553",
"internal/monitoring/system_alerts.go": "53ed4f02784636363da273883415331066d8b2484ca50e6b506e69990abb149d"
},
"enterprise_base_sha": "3d9f4e3051d38027355a2a1f36b8c7f672a09b65",
"enterprise_content_sha256": {
"internal/investigation/orchestrator.go": "d56fd512dc47f8a2453559e89863da5d24dffc0a5977abb8f1ea7a82655cd9e8"
},
"binary_sha256": "0c19f9b9265eff2214e79ab6b2e4a6b19a5645eb467c994c44af70b3d6593317",
"routes": [
"/proxmox/overview"
"/qualification (isolated Overview component on :5199)",
"/patrol (Activity, All history)",
"/actions?action=act_dcc3b52e5451810e49466daf9a6fccb0",
"/actions?action=act_ee0b736f0430e472e896a456ba3cb6eb",
"Pulse Assistant contextual panel from resolved Patrol findings",
"/qualification (isolated DemoBanner component on :5198)"
],
"viewports": [
{
"width": 1280,
"height": 800
"width": 1440,
"height": 1000
},
{
"width": 900,
"height": 1000
},
{
"width": 390,
"height": 844
"height": 1000
}
],
"states": [
"demo-mode banner at 1280x800 with the read-only notice followed by the \"Run Pulse on your own hardware\" link",
"demo-mode banner at 390x844 wrapping onto two lines with the link visible and the dismiss control intact"
"Incoming merged alert delivery diagnosis ordering: hold older request, add alert to start newer request, render current notifications-disabled state, release older ready response, verify current state and both cards remain. Existing Patrol and login proof is retained in the prior committed receipt and runtime source is unchanged.",
"Real persisted approved/verified and rejected findings remain reviewable after resolution. Durable investigation outcome agrees with authoritative action, with original prose retained. No stale Fix Queued status. Exact action links remain available.",
"Completed and Rejected action headers, State when planned and Plan expiry copy, inert settled action controls, independently verified recovery and explicit unavailable rollback. Policy, evidence and delivery disclosures expand and collapse.",
"Assistant handoff shows the exact finding with completed/rejected action context and Chat: Read-only. Browser provider readiness POST is deliberately blocked, so its visible route error is a rendering check and does not retest the provider. No prompt is submitted.",
"Incoming main DemoBanner install link: visible in demo policy, absent outside demo, link hover/focus, exact external setup destination opened in a new tab with noopener/noreferrer, keyboard dismissal and reload persistence. Real backend mock mode remains off. External destination content is intercepted because only navigation is under test."
],
"interactions": [
"load /proxmox/overview on a DEMO_MODE=true mock backend behind Vite in fresh contexts at both widths",
"read the link attributes from the live DOM: href https://pulserelay.pro/#setup, target _blank, rel noopener noreferrer"
]
"scripts/check-alert-diagnosis-ordering.mjs passed at 1440, 900 and 390 by 1000. Actual pixels inspected at desktop and narrow sizes. This is scripted component evidence, not proof of installed delivery or recipient receipt. Screenshots in /tmp/pulse-alert-diagnosis-ordering/.",
"Final Pro binary and final frontend content exercised in Playwright at 1440, 900 and 390 by 1000. Keyboard open/review, safety disclosure, exact action navigation, direct deep-link reload, policy/evidence/delivery disclosure keyboard toggles, Escape and close-button dismissal, review focus return where retained, Assistant open/close, scrolling and page overflow checks. Desktop/intermediate/phone pixels inspected including deepest evidence and Assistant overlay.",
"Private artifacts: tmp/patrol-gemini-38/live-action-browser/. Final complete matrix passed after removing redundant back-to-back full-page navigations from the proof driver. Earlier proof attempts hit an intermittent bootstrap connection screen. No bootstrap fix or general availability claim is made. API writes blocked except login.",
"After integrating main eef4ea21e73aedaa69380574b1fdf4d9dbdee3c8, repeated the complete real Patrol/Actions/Assistant matrix on binary 0c19f9b9265eff2214e79ab6b2e4a6b19a5645eb467c994c44af70b3d6593317. All three widths passed and pixels were reinspected. Incoming DemoBanner separately passed at the same widths using an isolated Vite cache and actual project styling. Prior frontend hashes remain byte-identical and are retained above. Merged monitoring regressions, API action tests, 37 component tests and type-check pass."
],
"prior_frontend_content_sha256": {
"frontend-modern/src/components/AI/FindingsPanel.tsx": "f506a26757b4c0ea3adf3f77af10214bfd31578b7122d3904a9b3a7272d1e146",
"frontend-modern/src/components/patrol/ApprovalSection.tsx": "6a18d67d5d3d0a8335589bdf199c340eb4775aea3442dc094971d7f924a0dcc5",
"frontend-modern/src/features/actions/ActionDecisionPacket.tsx": "2f1fd68ec333e7f9e8792d74ba8d7e6a2e95755c82b9c7121d847469a31f931b",
"frontend-modern/src/features/actions/ActionReviewDialog.tsx": "49e12cfd44686bd657ddddfb167c5956ec693b6f5d45d9d43c1e3a2e858a6c7f",
"frontend-modern/src/features/alerts/useAlertOverviewState.ts": "64d0b891e7ad228e8590da859dc25e825b6164c8cf76a01983a219d6cd079b23"
}
}
@@ -285,6 +285,16 @@ describe('patrol api — uncovered branch coverage', () => {
role: 'assistant',
content: 'logs in /var/log grew 40GB',
reasoning_content: 'checked du output',
tool_calls: [
{
id: 'read-1',
name: 'pulse_read',
input: { resource_id: 'container-1' },
output: 'NO_AGENT',
success: false,
},
{ id: 'query-1', name: 'pulse_query', input: { action: 'metrics' } },
],
timestamp: '2026-07-18T00:00:05Z',
},
],
@@ -299,6 +309,9 @@ describe('patrol api — uncovered branch coverage', () => {
expect(result).toEqual(envelope);
expect(result.messages).toHaveLength(2);
expect(result.messages[1]?.reasoning_content).toBe('checked du output');
expect(result.messages[1]?.tool_calls?.[0]?.output).toBe('NO_AGENT');
expect(result.messages[1]?.tool_calls?.[0]?.success).toBe(false);
expect(result.messages[1]?.tool_calls?.[1]?.success).toBeUndefined();
});
it('URL-encodes the finding id segment (separate from the messages suffix)', async () => {
+2 -2
View File
@@ -6,6 +6,7 @@
import { apiFetchJSON } from '@/utils/apiClient';
import { arrayOrEmpty, promoteLegacyAlertIdentifier } from './responseUtils';
import type { InvestigationRecord } from './ai';
import type { ToolCall } from './aiChat';
import type { ResourceCriticality } from './resourceOperatorState';
import type { PatrolActionReference } from '@/types/actionAudit';
import type { PatrolModelReadinessSnapshot } from '@/types/ai';
@@ -304,9 +305,8 @@ export interface ChatMessage {
timestamp: string;
}
export interface ChatToolCall {
export interface ChatToolCall extends ToolCall {
id: string;
name: string;
input: Record<string, unknown>;
}
@@ -194,6 +194,21 @@ export const FindingsPanel: Component<FindingsPanelProps> = (props) => {
const [filter, setFilter] = createSignal<FindingsPanelFilter>(props.filterOverride ?? 'active');
const [sortBy, setSortBy] = createSignal<'severity' | 'time'>('severity');
const [expandedId, setExpandedId] = createSignal<string | null>(null);
let panelRoot: HTMLDivElement | undefined;
const closeReviewPanel = () => {
const findingId = expandedId();
setExpandedId(null);
setManageOpenId(null);
if (findingId) {
queueMicrotask(() => {
panelRoot
?.querySelector<HTMLButtonElement>(
`button[aria-controls="${CSS.escape(`finding-${findingId}-details`)}"]`,
)
?.focus();
});
}
};
const [manageOpenId, setManageOpenId] = createSignal<string | null>(null);
const [actionLoading, setActionLoading] = createSignal<string | null>(null);
const [lastHashScrolled, setLastHashScrolled] = createSignal<string | null>(null);
@@ -1467,7 +1482,10 @@ export const FindingsPanel: Component<FindingsPanelProps> = (props) => {
manualControls.dismiss;
return (
<div id={`finding-${finding.id}-details`} class="mt-3 pt-3 border-t border-border-subtle">
<div
id={isPatrolFindingsSource() ? undefined : `finding-${finding.id}-details`}
class="mt-3 pt-3 border-t border-border-subtle"
>
<Show when={hasTriggeringAlert(finding)}>
<div class="text-xs text-amber-700 dark:text-amber-300 mb-2">
Triggered by alert{finding.alertType ? ` (${finding.alertType})` : ''} Identifier{' '}
@@ -2052,18 +2070,18 @@ export const FindingsPanel: Component<FindingsPanelProps> = (props) => {
{/* Inline Approval Section (replaces manual approval JSX) */}
<Show
when={
finding.status === 'active' &&
(finding.investigationOutcome === 'fix_queued' ||
finding.investigationOutcome === 'fix_executed' ||
finding.investigationOutcome === 'fix_failed' ||
finding.investigationOutcome === 'fix_rejected' ||
finding.investigationOutcome === 'fix_verified' ||
finding.investigationOutcome === 'fix_verification_failed' ||
finding.investigationOutcome === 'fix_verification_unknown')
finding.investigationOutcome === 'fix_queued' ||
finding.investigationOutcome === 'fix_executed' ||
finding.investigationOutcome === 'fix_failed' ||
finding.investigationOutcome === 'fix_rejected' ||
finding.investigationOutcome === 'fix_verified' ||
finding.investigationOutcome === 'fix_verification_failed' ||
finding.investigationOutcome === 'fix_verification_unknown'
}
>
<ApprovalSection
findingId={finding.id}
findingStatus={finding.status}
investigationOutcome={finding.investigationOutcome}
findingTitle={getFindingTitlePresentation(finding).label}
resourceName={finding.resourceName}
@@ -2108,10 +2126,10 @@ export const FindingsPanel: Component<FindingsPanelProps> = (props) => {
};
return (
<div class="space-y-4">
<div ref={panelRoot} class="space-y-4">
{/* Controls */}
<Show when={showFilterControls()}>
<div class="flex items-center justify-between">
<div class="flex flex-wrap items-center justify-between gap-2">
<FilterSegmentedControl
aria-label="Filter findings"
value={filter()}
@@ -2298,10 +2316,7 @@ export const FindingsPanel: Component<FindingsPanelProps> = (props) => {
<button
type="button"
aria-label={`Close review panel for ${title().label}`}
onClick={() => {
setExpandedId(null);
setManageOpenId(null);
}}
onClick={closeReviewPanel}
class="rounded p-1.5 text-muted transition-colors hover:bg-surface-hover hover:text-base-content focus:outline-none focus-visible:ring-2 focus-visible:ring-primary/40"
>
<XIcon class="h-4 w-4" />
@@ -3,6 +3,7 @@ import type { Component } from 'solid-js';
import {
formatBytes,
formatSpeed,
formatObservedSpeed,
formatUptime,
getResourceDiskSummary,
normalizeDiskArray,
@@ -334,8 +335,11 @@ export const UnifiedResourceHostTableCard: Component<UnifiedResourceHostTableCar
const networkEmphasis = createMemo(() =>
getOutlierEmphasis(networkTotal(), table.ioScale().network),
);
const diskIOTotal = createMemo(
() => (resource.diskIO?.readRate ?? 0) + (resource.diskIO?.writeRate ?? 0),
const diskIOTotal = createMemo(() =>
resource.diskIO?.readRate !== undefined &&
resource.diskIO?.writeRate !== undefined
? resource.diskIO.readRate + resource.diskIO.writeRate
: NaN,
);
const diskIOEmphasis = createMemo(() =>
getOutlierEmphasis(diskIOTotal(), table.ioScale().diskIO),
@@ -662,11 +666,11 @@ export const UnifiedResourceHostTableCard: Component<UnifiedResourceHostTableCar
class={`min-w-0 overflow-hidden text-ellipsis whitespace-nowrap ${diskIOEmphasis().className}`}
title={
diskIOEmphasis().showOutlierHint
? `${formatSpeed(resource.diskIO!.readRate)} (Top outlier)`
: formatSpeed(resource.diskIO!.readRate)
? `${formatObservedSpeed(resource.diskIO!.readRate)} (Top outlier)`
: formatObservedSpeed(resource.diskIO!.readRate)
}
>
{formatSpeed(resource.diskIO!.readRate)}
{formatObservedSpeed(resource.diskIO!.readRate)}
</span>
<span class="inline-flex w-3 justify-center font-mono text-amber-500">
W
@@ -675,11 +679,11 @@ export const UnifiedResourceHostTableCard: Component<UnifiedResourceHostTableCar
class={`min-w-0 overflow-hidden text-ellipsis whitespace-nowrap ${diskIOEmphasis().className}`}
title={
diskIOEmphasis().showOutlierHint
? `${formatSpeed(resource.diskIO!.writeRate)} (Top outlier)`
: formatSpeed(resource.diskIO!.writeRate)
? `${formatObservedSpeed(resource.diskIO!.writeRate)} (Top outlier)`
: formatObservedSpeed(resource.diskIO!.writeRate)
}
>
{formatSpeed(resource.diskIO!.writeRate)}
{formatObservedSpeed(resource.diskIO!.writeRate)}
</span>
</div>
</Show>
@@ -458,3 +458,16 @@ describe('infrastructureSelectors', () => {
});
});
});
it('does not compare partial disk observations as complete throughput totals', () => {
const known = [
makeResource(1, { diskIO: { readRate: 0, writeRate: 0 } }),
makeResource(2, { diskIO: { readRate: 100, writeRate: 200 } }),
];
const partial = [
makeResource(3, { diskIO: { readRate: 10000 } }),
makeResource(4, { diskIO: { writeRate: 0 } }),
makeResource(5, { diskIO: {} }),
];
expect(computeIOScale([...known, ...partial]).diskIO).toEqual(computeIOScale(known).diskIO);
});
@@ -86,7 +86,9 @@ const getSortValue = (resource: Resource, key: string): number | string | null =
case 'network':
return resource.network ? resource.network.rxBytes + resource.network.txBytes : null;
case 'diskio':
return resource.diskIO ? resource.diskIO.readRate + resource.diskIO.writeRate : null;
return resource.diskIO?.readRate !== undefined && resource.diskIO?.writeRate !== undefined
? resource.diskIO.readRate + resource.diskIO.writeRate
: null;
case 'source':
return getInfrastructureSystemIdentitySortLabel(resource);
case 'temp':
@@ -349,9 +351,8 @@ export const computeIOScale = (
networkValues.push(networkTotal);
}
const diskIOTotal = (resource.diskIO?.readRate ?? 0) + (resource.diskIO?.writeRate ?? 0);
if (resource.diskIO) {
diskIOValues.push(diskIOTotal);
if (resource.diskIO?.readRate !== undefined && resource.diskIO?.writeRate !== undefined) {
diskIOValues.push(resource.diskIO.readRate + resource.diskIO.writeRate);
}
}
@@ -23,6 +23,7 @@ import type { ActionAuditState, PatrolActionReference } from '@/types/actionAudi
interface ApprovalSectionProps {
findingId: string;
findingStatus?: string;
investigationOutcome?: string;
findingTitle?: string;
resourceName?: string;
@@ -127,7 +128,7 @@ export const ApprovalSection: Component<ApprovalSectionProps> = (props) => {
current?.plan.message ||
investigation()?.summary ||
'Review the current Patrol finding and its governed action state.',
findingStatus: 'active',
findingStatus: props.findingStatus ?? 'active',
investigationOutcome: props.investigationOutcome,
loopState: props.investigationOutcome || current?.state,
resourceId: props.resourceId || current?.resource_id,
@@ -10,6 +10,7 @@ import { getInvestigationMessages, formatTimestamp, type ChatMessage } from '@/a
import { LoadingSpinner } from '@/components/shared/LoadingSpinner';
import { getInvestigationMessagesState } from '@/utils/patrolEmptyStatePresentation';
import { renderMarkdown } from '@/components/AI/aiChatUtils';
import { ToolExecutionBlock } from '@/components/AI/Chat/ToolExecutionBlock';
// Compact variant of the Assistant chat's markdown styling, scaled for the
// investigation thread's text-xs bubbles.
@@ -110,16 +111,35 @@ export const InvestigationMessages: Component<InvestigationMessagesProps> = (pro
<div class="space-y-1">
<For each={msg.tool_calls}>
{(tc) => (
<div class="text-xs rounded border border-indigo-200 dark:border-indigo-800 bg-indigo-50 dark:bg-indigo-900 px-2 py-1">
<span class="font-semibold text-indigo-700 dark:text-indigo-300">
{tc.name}
</span>
<Show when={tc.input && Object.keys(tc.input).length > 0}>
<pre class="mt-1 text-[10px] text-muted overflow-x-auto max-h-24 overflow-y-auto">
{JSON.stringify(tc.input, null, 2)}
</pre>
</Show>
</div>
<Show
when={typeof tc.success === 'boolean'}
fallback={
<div class="text-xs rounded border border-indigo-200 dark:border-indigo-800 bg-indigo-50 dark:bg-indigo-900 px-2 py-1">
<span class="font-semibold text-indigo-700 dark:text-indigo-300">
{tc.name}
</span>
<Show when={tc.input && Object.keys(tc.input).length > 0}>
<pre class="mt-1 text-[10px] text-muted overflow-x-auto max-h-24 overflow-y-auto">
{JSON.stringify(tc.input, null, 2)}
</pre>
</Show>
<Show when={tc.output}>
<pre class="mt-1 text-[10px] text-muted overflow-x-auto max-h-32 overflow-y-auto whitespace-pre-wrap break-words">
{tc.output}
</pre>
</Show>
</div>
}
>
<ToolExecutionBlock
tool={{
name: tc.name,
input: JSON.stringify(tc.input),
output: tc.output ?? '',
success: tc.success!,
}}
/>
</Show>
)}
</For>
</div>
@@ -29,6 +29,7 @@ import { buildPatrolInvestigationRecordPresentation } from '@/features/patrol/pa
import { LoadingSpinner } from '@/components/shared/LoadingSpinner';
import { MetadataBadge } from '@/components/shared/MetadataBadge';
import { InvestigationMessages } from './InvestigationMessages';
import { renderMarkdown } from '@/components/AI/aiChatUtils';
import { notificationStore } from '@/stores/notifications';
import { aiIntelligenceStore } from '@/stores/aiIntelligence';
import type { InvestigationRecord } from '@/api/ai';
@@ -40,6 +41,9 @@ const INVESTIGATION_BADGE_PROPS = {
shape: 'rounded',
} as const;
const summaryClass =
'text-sm prose prose-slate prose-sm dark:prose-invert max-w-none break-words prose-headings:my-2 prose-p:my-2 prose-pre:overflow-x-auto prose-code:break-all prose-code:before:content-none prose-code:after:content-none';
interface InvestigationSectionProps {
findingId: string;
investigationStatus?: string;
@@ -210,7 +214,11 @@ export const InvestigationSection: Component<InvestigationSectionProps> = (props
</div>
<Show when={investigationRecord().conclusion}>
<p class="mt-2 text-sm text-base-content">{investigationRecord().conclusion}</p>
<div
class={`mt-2 ${summaryClass}`}
// eslint-disable-next-line solid/no-innerhtml
innerHTML={renderMarkdown(investigationRecord().conclusion!)}
/>
</Show>
<Show when={investigationRecord().recommendedAction}>
<p class="mt-1 text-xs text-muted">
@@ -325,6 +333,7 @@ export const InvestigationSection: Component<InvestigationSectionProps> = (props
<Show
when={
inv().error &&
inv().error?.trim() !== investigationRecord().error &&
(inv().status === 'failed' ||
inv().outcome === 'timed_out' ||
inv().outcome === 'fix_failed' ||
@@ -341,8 +350,16 @@ export const InvestigationSection: Component<InvestigationSectionProps> = (props
</Show>
{/* Summary */}
<Show when={inv().summary}>
<div class="text-sm text-muted bg-surface-alt rounded p-2">{inv().summary}</div>
<Show
when={
inv().summary?.trim() && inv().summary?.trim() !== investigationRecord().conclusion
}
>
<div
class={`bg-surface-alt rounded p-2 ${summaryClass}`}
// eslint-disable-next-line solid/no-innerhtml
innerHTML={renderMarkdown(inv().summary!)}
/>
</Show>
{/* Tools used + turn count */}
@@ -67,13 +67,17 @@ describe('ApprovalSection typed action handoff', () => {
window.history.replaceState({}, '', '/');
});
const renderSection = (investigationOutcome: string) =>
const renderSection = (investigationOutcome: string, findingStatus = 'active') =>
render(() => (
<Router>
<Route
path="/"
component={() => (
<ApprovalSection findingId="finding-1" investigationOutcome={investigationOutcome} />
<ApprovalSection
findingId="finding-1"
findingStatus={findingStatus}
investigationOutcome={investigationOutcome}
/>
)}
/>
</Router>
@@ -104,13 +108,20 @@ describe('ApprovalSection typed action handoff', () => {
it('routes terminal action history to the exact recorded outcome', async () => {
getInvestigationMock.mockResolvedValue(investigation(actionReference('completed')));
renderSection('fix_verified');
renderSection('fix_verified', 'resolved');
expect(await screen.findByRole('link', { name: /view outcome in actions/i })).toHaveAttribute(
'href',
'/actions?action=act-1',
);
expect(screen.getByText('Outcome verified')).toBeInTheDocument();
fireEvent.click(screen.getByRole('button', { name: /discuss with assistant/i }));
expect(openMock).toHaveBeenCalledWith(
expect.objectContaining({
handoffContext: expect.stringContaining('Resolved'),
autonomousMode: false,
}),
);
});
it('keeps missing plan identity visible while leaving replan guidance to Actions', async () => {
@@ -0,0 +1,68 @@
import { cleanup, fireEvent, render, screen } from '@solidjs/testing-library';
import { afterEach, describe, expect, it, vi } from 'vitest';
import InvestigationMessages from '../InvestigationMessages';
const getMessages = vi.hoisted(() => vi.fn());
vi.mock('@/api/patrol', () => ({
getInvestigationMessages: getMessages,
formatTimestamp: (value: string) => value,
}));
afterEach(cleanup);
describe('InvestigationMessages', () => {
it('retains merged failed tool evidence and exposes it through the shared disclosure', async () => {
getMessages.mockResolvedValue({
messages: [
{
id: 'turn-1',
role: 'assistant',
content: '',
timestamp: '2026-09-06',
tool_calls: [
{
id: 'read-1',
name: 'pulse_read',
input: { resource_id: 'app-container-1' },
output: '{"error":"NO_AGENT","message":"Command agent is not connected"}',
success: false,
},
],
},
],
});
render(() => <InvestigationMessages findingId="finding-1" />);
const disclosure = await screen.findByRole('button', { name: /failed/i });
expect(disclosure).toHaveAttribute('aria-expanded', 'false');
fireEvent.keyDown(disclosure, { key: 'Enter' });
expect(disclosure).toHaveAttribute('aria-expanded', 'true');
expect(screen.getByText(/"error":\s*"NO_AGENT"/)).toBeVisible();
fireEvent.keyDown(disclosure, { key: ' ' });
expect(disclosure).toHaveAttribute('aria-expanded', 'false');
});
it('does not infer completion from historical calls that have no recorded result status', async () => {
getMessages.mockResolvedValue({
messages: [
{
id: 'turn-2',
role: 'assistant',
content: '',
timestamp: '2026-09-06',
tool_calls: [
{
id: 'query-1',
name: 'pulse_query',
input: { action: 'metrics' },
output: 'Historical output',
},
],
},
],
});
render(() => <InvestigationMessages findingId="finding-2" />);
expect(await screen.findByText('Historical output')).toBeVisible();
expect(screen.queryByText('completed')).not.toBeInTheDocument();
expect(screen.queryByText('failed')).not.toBeInTheDocument();
});
});
@@ -133,4 +133,46 @@ describe('InvestigationSection', () => {
expect(screen.queryByText(/No investigation data available/)).not.toBeInTheDocument();
expect(screen.queryByText('systemctl restart workload.service')).not.toBeInTheDocument();
});
it.each([false, true])(
'preserves readable evidence and distinct summaries (different: %s)',
async (different) => {
const conclusion =
'### Root cause\n\nThe cause is **unknown**.\n\n- The health check failed.\n- Logs are unavailable.';
getInvestigationMock.mockResolvedValue({
id: 'inv-markdown',
finding_id: 'finding-markdown',
session_id: 'session-markdown',
status: 'completed',
started_at: '2026-09-06T17:00:00Z',
turn_count: 2,
summary: different
? '### Follow-up\n\nAdditional evidence remains unavailable.'
: conclusion,
} satisfies Investigation);
render(() => (
<InvestigationSection
findingId="finding-markdown"
investigationRecord={{
id: 'record-markdown',
finding_id: 'finding-markdown',
subject: { resource_id: 'container-1' },
trigger: { title: 'Health check failed', detected_at: '2026-09-06T17:00:00Z' },
status: 'completed',
conclusion,
started_at: '2026-09-06T17:00:00Z',
evidence: [],
verification: [],
rollback: [],
tools_used: [],
}}
/>
));
await screen.findByRole('button', { name: 'Show investigation thread' });
expect(screen.getAllByRole('heading', { name: 'Root cause' })).toHaveLength(1);
expect(screen.getByText('unknown').tagName).toBe('STRONG');
expect(screen.getAllByRole('listitem')).toHaveLength(2);
expect(screen.queryByRole('heading', { name: 'Follow-up' }) !== null).toBe(different);
},
);
});
@@ -43,7 +43,7 @@ export const ActionDecisionPacket: Component<{
class="rounded-lg border border-border bg-surface p-4"
>
<h3 id="action-intent-heading" class="text-sm font-semibold text-base-content">
What will happen
Action plan
</h3>
<dl class="mt-3 grid gap-3 text-sm sm:grid-cols-2">
<div>
@@ -69,7 +69,7 @@ export const ActionDecisionPacket: Component<{
</Show>
<Show when={props.audit.plan.preflight?.currentState}>
<div>
<dt class="text-muted">Current state</dt>
<dt class="text-muted">State when planned</dt>
<dd>{props.audit.plan.preflight?.currentState}</dd>
</div>
</Show>
@@ -80,7 +80,7 @@ export const ActionDecisionPacket: Component<{
</div>
</Show>
<div>
<dt class="text-muted">Approval expires</dt>
<dt class="text-muted">Plan expiry</dt>
<dd>{expiry()}</dd>
</div>
<div>
@@ -90,7 +90,7 @@ export const ActionDecisionPacket: Component<{
</dl>
<Show when={blastRadiusEntries().length > 0}>
<div class="mt-3">
<div class="text-sm text-muted">Also affected</div>
<div class="text-sm text-muted">Potentially affected</div>
<ul class="mt-1 list-disc pl-5 text-sm">
<For each={blastRadiusEntries()}>
{(entry) => (
@@ -4,12 +4,14 @@ import ArrowUpRightIcon from 'lucide-solid/icons/arrow-up-right';
import { ResourceActionsAPI } from '@/api/resourceActions';
import { Button, ButtonLink } from '@/components/shared/Button';
import { Dialog } from '@/components/shared/Dialog';
import { MetadataBadge } from '@/components/shared/MetadataBadge';
import { notificationStore } from '@/stores/notifications';
import { presentationPolicyIsReadOnly } from '@/stores/sessionPresentationPolicy';
import type { ActionDetailResponse } from '@/types/actionAudit';
import { ActionDecisionPacket } from './ActionDecisionPacket';
import {
formatActionName,
getActionInboxStatePresentation,
getActionOriginDestination,
getActionResourcePresentation,
} from './actionPresentation';
@@ -244,9 +246,14 @@ export const ActionReviewDialog: Component<{
<p class="text-xs font-semibold uppercase tracking-wide text-muted">
Governed action review
</p>
<h2 id="action-review-title" class="mt-1 text-xl font-semibold">
{formatActionName(record().request.capabilityName)}
</h2>
<div class="mt-1 flex flex-wrap items-center gap-2">
<h2 id="action-review-title" class="text-xl font-semibold">
{formatActionName(record().request.capabilityName)}
</h2>
<MetadataBadge tone={getActionInboxStatePresentation(record().state).tone}>
{getActionInboxStatePresentation(record().state).label}
</MetadataBadge>
</div>
<p class="mt-1 text-sm text-muted">
{resource().label}
<Show when={resource().detail}> · {resource().detail}</Show>
@@ -115,6 +115,16 @@ const detail = (audit: ActionAuditRecord): ActionDetailResponse => ({
});
describe('ActionReviewDialog trust gates', () => {
it('keeps a rejected action outcome visible without offering execution', () => {
const audit = makeAudit('resolved', '2026-07-12T00:10:00Z');
audit.state = 'rejected';
render(() => <ActionReviewDialog detail={detail(audit)} onClose={vi.fn()} />);
expect(screen.getByText('Rejected', { exact: true })).toBeVisible();
expect(
screen.queryByRole('button', { name: /approve|run|refresh plan/i }),
).not.toBeInTheDocument();
});
it('links a trusted Patrol action back to its exact operational record', () => {
const audit = makeAudit('resolved', '2099-01-01T00:00:00Z');
audit.origin = {
@@ -20,7 +20,13 @@ import { hostOverrideIdCandidates } from '@/features/alerts/alertOverridesModel'
import { areSystemSettingsLoaded, shouldHideDockerUpdateActions } from '@/stores/systemSettings';
import { useAlertsActivation } from '@/stores/alertsActivation';
import type { Resource } from '@/types/resource';
import { formatBytes, formatRelativeTime, formatSpeed, normalizeDiskArray } from '@/utils/format';
import {
formatBytes,
formatRelativeTime,
formatSpeed,
formatObservedSpeed,
normalizeDiskArray,
} from '@/utils/format';
import { formatTemperature, getTemperatureTextClass } from '@/utils/temperature';
interface DockerHostDrawerOverviewProps {
@@ -279,7 +285,7 @@ export function DockerHostDrawerOverview(props: DockerHostDrawerOverviewProps) {
) {
rows.push({
label: 'Disk I/O',
value: `${formatSpeed(props.host.diskIO?.readRate ?? 0)} / ${formatSpeed(props.host.diskIO?.writeRate ?? 0)}`,
value: `${formatObservedSpeed(props.host.diskIO?.readRate)} / ${formatObservedSpeed(props.host.diskIO?.writeRate)}`,
});
}
return rows;
@@ -225,9 +225,7 @@ export function PatrolIntelligenceSurface() {
onToggle={(event) => setFindingsOpen(event.currentTarget.open)}
>
<summary class="sr-only">Finding options and history</summary>
<div
class={`space-y-4 border-t border-border p-4 sm:p-5 ${!state.patrolEnabledLocal() ? 'opacity-50 pointer-events-none' : ''}`}
>
<div class="space-y-4 border-t border-border p-4 sm:p-5">
<div class="flex flex-col gap-2 sm:flex-row sm:items-center sm:justify-between">
<p class="text-xs leading-5 text-muted">
<Show
@@ -61,7 +61,7 @@ import type { Disk } from '@/types/api';
import type { Resource, ResourceAvailabilityMeta } from '@/types/resource';
import type { MetricDisplayThresholds } from '@/utils/metricThresholds';
import { getActionableAgentIdFromResource } from '@/utils/agentResources';
import { formatBytes, formatSpeed, normalizeDiskArray } from '@/utils/format';
import { formatBytes, formatSpeed, formatObservedSpeed, normalizeDiskArray } from '@/utils/format';
import { STORAGE_KEYS } from '@/utils/localStorage';
import { useAlertsActivation } from '@/stores/alertsActivation';
import { notificationStore } from '@/stores/notifications';
@@ -84,7 +84,6 @@ import {
getAgentMachineGPUTitle,
getAgentMachineGPUUtilizationPercent,
getAgentMachineDiskIODetails,
getAgentMachineDiskIOTotal,
getAgentMachineIpValues,
matchesAgentMachineSearch,
getAgentMachineNetworkInterfaceDetails,
@@ -502,9 +501,9 @@ const AgentMachineDiskIOCell: Component<{
trigger={
<>
<span class="inline-flex shrink-0 font-mono text-blue-500">R</span>
<span class="min-w-0 truncate">{formatSpeed(props.diskIO?.readRate ?? 0)}</span>
<span class="min-w-0 truncate">{formatObservedSpeed(props.diskIO?.readRate)}</span>
<span class="inline-flex shrink-0 font-mono text-amber-500">W</span>
<span class="min-w-0 truncate">{formatSpeed(props.diskIO?.writeRate ?? 0)}</span>
<span class="min-w-0 truncate">{formatObservedSpeed(props.diskIO?.writeRate)}</span>
</>
}
>
@@ -513,11 +512,11 @@ const AgentMachineDiskIOCell: Component<{
<div class="mb-1 grid grid-cols-[auto_minmax(0,1fr)] gap-x-2 gap-y-0.5 text-[9px]">
<span class="font-mono text-blue-500">Read</span>
<span class="min-w-0 truncate text-base-content">
{formatSpeed(props.diskIO?.readRate ?? 0)}
{formatObservedSpeed(props.diskIO?.readRate)}
</span>
<span class="font-mono text-amber-500">Write</span>
<span class="min-w-0 truncate text-base-content">
{formatSpeed(props.diskIO?.writeRate ?? 0)}
{formatObservedSpeed(props.diskIO?.writeRate)}
</span>
</div>
<div class="max-h-[280px] space-y-1.5 overflow-y-auto pr-1">
@@ -1038,7 +1037,7 @@ const networkTitleFor = (machine: Resource): string => {
const diskIOTitleFor = (machine: Resource): string => {
if (!machine.diskIO) return '';
return `Read ${formatSpeed(machine.diskIO.readRate)}\nWrite ${formatSpeed(machine.diskIO.writeRate)}`;
return `Read ${formatObservedSpeed(machine.diskIO.readRate)}\nWrite ${formatObservedSpeed(machine.diskIO.writeRate)}`;
};
const agentIdentityIdFor = (machine: Resource): string =>
@@ -1527,7 +1526,6 @@ export const AgentsMachinesTable: Component<{
aggregateDisk() !== undefined || (disks()?.length ?? 0) > 0;
const networkTotal = () => getAgentMachineNetworkTotal(machine);
const networkInterfaces = () => getAgentMachineNetworkInterfaceDetails(machine);
const diskIOTotal = () => getAgentMachineDiskIOTotal(machine);
const diskIODetails = () => getAgentMachineDiskIODetails(machine);
const primaryIp = () =>
getPreferredResourceIP(machine) ?? getAgentMachinePrimaryIp(machine);
@@ -1738,7 +1736,11 @@ export const AgentsMachinesTable: Component<{
class={`${getPlatformTableCellClassForKind('numeric-value')} ${machineColumnWidthClass('diskio')} text-base-content`}
>
<Show
when={canRenderMetrics() && diskIOTotal() !== undefined}
when={
canRenderMetrics() &&
(machine.diskIO?.readRate !== undefined ||
machine.diskIO?.writeRate !== undefined)
}
fallback={telemetryFallbackMarker()}
>
<AgentMachineDiskIOCell
@@ -332,20 +332,20 @@ describe('agentMachineTableModel coverage2', () => {
).toBe(800);
});
it('returns read only when write is absent', () => {
it('leaves the total unavailable when write is absent', () => {
expect(
getAgentMachineDiskIOTotal(
resource({ diskIO: { readRate: 500 } as unknown as ResourceDiskIO }),
),
).toBe(500);
).toBeUndefined();
});
it('returns write only when read is absent', () => {
it('leaves the total unavailable when read is absent', () => {
expect(
getAgentMachineDiskIOTotal(
resource({ diskIO: { writeRate: 300 } as unknown as ResourceDiskIO }),
),
).toBe(300);
).toBeUndefined();
});
it('returns undefined when both rates are absent', () => {
@@ -2,6 +2,7 @@ import { describe, expect, it } from 'vitest';
import type { Resource } from '@/types/resource';
import {
getAgentMachineDiskPercent,
getAgentMachineDiskIOTotal,
getAgentMachineDiskIODetails,
getAgentMachineGPUTitle,
getAgentMachineGPUUtilizationPercent,
@@ -476,3 +477,12 @@ describe('agentMachineTableModel', () => {
);
});
});
it('requires both disk directions for a throughput total', () => {
expect(getAgentMachineDiskIOTotal(resource({ diskIO: { readRate: 0, writeRate: 0 } }))).toBe(0);
expect(getAgentMachineDiskIOTotal(resource({ diskIO: { readRate: 100, writeRate: 200 } }))).toBe(
300,
);
expect(getAgentMachineDiskIOTotal(resource({ diskIO: { readRate: 0 } }))).toBeUndefined();
expect(getAgentMachineDiskIOTotal(resource({ diskIO: { writeRate: 100 } }))).toBeUndefined();
});
@@ -493,8 +493,8 @@ export const getAgentMachineNetworkInterfaceDetails = (
export const getAgentMachineDiskIOTotal = (machine: Resource): number | undefined => {
const read = getPlatformTableFiniteMetric(machine.diskIO?.readRate);
const write = getPlatformTableFiniteMetric(machine.diskIO?.writeRate);
if (read === undefined && write === undefined) return undefined;
return (read ?? 0) + (write ?? 0);
if (read === undefined || write === undefined) return undefined;
return read + write;
};
export const getAgentMachineDiskIODetails = (machine: Resource): AgentMachineDiskIODetail[] => {
@@ -120,6 +120,29 @@ describe('useColumnVisibility', () => {
});
});
it('does not reapply a default-hidden migration after a fresh user shows the column', async () => {
const columns: ColumnDef[] = [
{ id: 'name', label: 'Name' },
{ id: 'diskio', label: 'Disk I/O', toggleable: true, defaultHidden: true },
];
let dispose = () => {};
let visibility: ReturnType<typeof useColumnVisibility>;
createRoot((d) => {
dispose = d;
visibility = useColumnVisibility(storageKey, columns, [], undefined, {}, ['diskio']);
});
await Promise.resolve();
visibility!.show('diskio');
await Promise.resolve();
expect(window.localStorage.getItem(storageKey)).toBe('[]');
dispose();
createRoot((d) => {
const reloaded = useColumnVisibility(storageKey, columns, [], undefined, {}, ['diskio']);
expect(reloaded.isHiddenByUser('diskio')).toBe(false);
d();
});
});
it('resets back to the canonical default-hidden set', () => {
createRoot((dispose) => {
const columns: ColumnDef[] = [
@@ -1451,6 +1451,25 @@ describe('useUnifiedResources', () => {
dispose();
});
it('preserves an observed zero without inventing its absent I/O direction', async () => {
setWsConnected(false);
setWsInitialDataReceived(false);
setWsState('resources', []);
apiFetchMock.mockResolvedValueOnce({
ok: true,
json: async () => ({ data: [{ ...v2Resource, metrics: { diskRead: { value: 0 } } }] }),
});
let dispose = () => {};
let result: ReturnType<UseUnifiedResourcesModule['useUnifiedResources']> | undefined;
createRoot((d) => {
dispose = d;
result = useUnifiedResources();
});
await result!.refetch();
expect(result!.resources()[0]?.diskIO).toEqual({ readRate: 0, writeRate: undefined });
dispose();
});
it('preserves richer REST resource details across thinner websocket updates', async () => {
setWsConnected(false);
setWsInitialDataReceived(false);
@@ -124,21 +124,21 @@ export function useColumnVisibility(
const appliedDefaultHiddenMigrations = hasUserPreference
? readAppliedDefaultHiddenMigrations(storageKey, persistedIdAliases)
: [];
const pendingDefaultHiddenMigrations = hasUserPreference
? Array.from(
new Set(
defaultHiddenMigrationIds
.map((id) => id.trim())
.filter(
(id) =>
id &&
effectiveDefaultHidden.includes(id) &&
toggleableIds.includes(id) &&
!appliedDefaultHiddenMigrations.includes(id),
),
// Fresh preferences already contain these defaults. Mark their migration as
// applied too, so the first reload cannot undo a user's subsequent choice.
const pendingDefaultHiddenMigrations = Array.from(
new Set(
defaultHiddenMigrationIds
.map((id) => id.trim())
.filter(
(id) =>
id &&
effectiveDefaultHidden.includes(id) &&
toggleableIds.includes(id) &&
!appliedDefaultHiddenMigrations.includes(id),
),
)
: [];
),
);
let defaultHiddenMigrationsPersisted = false;
// Persist hidden columns to localStorage
@@ -169,7 +169,7 @@ export function useColumnVisibility(
createEffect(() => {
const hasUnpersistedDefaultHiddenMigrations =
pendingDefaultHiddenMigrations.length > 0 && !defaultHiddenMigrationsPersisted;
if (!hasUserPreference || (!persistedIdsMigrated && !hasUnpersistedDefaultHiddenMigrations)) {
if (!persistedIdsMigrated && !hasUnpersistedDefaultHiddenMigrations) {
return;
}
persistedIdsMigrated = false;
@@ -915,8 +915,8 @@ const toResource = (v2: APIResource): Resource => {
diskIO:
v2.metrics?.diskRead || v2.metrics?.diskWrite
? {
readRate: v2.metrics?.diskRead?.value ?? 0,
writeRate: v2.metrics?.diskWrite?.value ?? 0,
readRate: v2.metrics?.diskRead?.value,
writeRate: v2.metrics?.diskWrite?.value,
}
: undefined,
uptime:
+2 -2
View File
@@ -133,8 +133,8 @@ export interface ResourceNetwork {
// Disk I/O metrics (rates in bytes/sec from backend)
export interface ResourceDiskIO {
readRate: number; // Read rate (bytes/sec)
writeRate: number; // Write rate (bytes/sec)
readRate?: number; // Observed read rate (bytes/sec), including measured zero.
writeRate?: number; // Absent directions remain unavailable.
}
// Alert associated with a resource
@@ -995,19 +995,19 @@ describe('getFindingResolutionReason', () => {
).toBe('Resolved after investigation timeout now');
});
it('returns "Resolved manually" for cannot_fix', () => {
it('does not infer manual resolution from cannot_fix', () => {
expect(
getFindingResolutionReason({ ...patrolBase, investigationOutcome: 'cannot_fix' }, 'now'),
).toBe('Resolved manually now');
).toBe('Resolved now');
});
it('returns "Resolved after manual review" for needs_attention', () => {
it('does not infer manual review from needs_attention', () => {
expect(
getFindingResolutionReason(
{ ...patrolBase, investigationOutcome: 'needs_attention' },
'now',
),
).toBe('Resolved after manual review now');
).toBe('Resolved now');
});
it('returns "Fix applied by Patrol" for fix_executed even when autoResolved is false', () => {
@@ -5,6 +5,7 @@ import { describe, expect, it, vi, beforeEach, afterEach } from 'vitest';
import {
formatBytes,
formatSpeed,
formatObservedSpeed,
formatPercent,
formatNumber,
formatUptime,
@@ -362,3 +363,13 @@ describe('getBackupInfo', () => {
});
});
});
describe('formatObservedSpeed', () => {
it('preserves zero and leaves missing or invalid observations unavailable', () => {
expect(formatObservedSpeed(0)).toBe('0 B/s');
expect(formatObservedSpeed(1024)).toBe('1.00 KB/s');
for (const value of [undefined, null, -1, NaN, Infinity]) {
expect(formatObservedSpeed(value)).toBe('-');
}
});
});
@@ -1287,9 +1287,10 @@ export const getFindingResolutionReason = (
case 'timed_out':
return `Resolved after investigation timeout ${resolvedTime}`;
case 'cannot_fix':
return `Resolved manually ${resolvedTime}`;
case 'needs_attention':
return `Resolved after manual review ${resolvedTime}`;
// An investigation outcome does not identify who later resolved the
// finding. Explicit operator resolution is handled above.
return `Resolved ${resolvedTime}`;
default:
return `Issue no longer detected ${resolvedTime}`;
}
+8
View File
@@ -67,6 +67,14 @@ export function formatSpeed(bytesPerSecond: number, decimals: number | 'auto' =
return `${formatBytes(bytesPerSecond, decimals)}/s`;
}
export function formatObservedSpeed(bytesPerSecond: number | null | undefined): string {
return typeof bytesPerSecond === 'number' &&
Number.isFinite(bytesPerSecond) &&
bytesPerSecond >= 0
? formatSpeed(bytesPerSecond)
: '-';
}
export function formatPercent(value: number): string {
if (!Number.isFinite(value)) return '0%';
const abs = Math.abs(value);
+45
View File
@@ -0,0 +1,45 @@
package agentcapabilities
import "encoding/json"
// TranscriptToolCall preserves a stored invocation and its observed result.
// Provider requests use ProviderToolCall instead of the product history shape.
type TranscriptToolCall struct {
ID string `json:"id"`
Name string `json:"name"`
Input map[string]interface{} `json:"input"`
Output string `json:"output,omitempty"`
Success *bool `json:"success,omitempty"`
ThoughtSignature json.RawMessage `json:"thought_signature,omitempty"`
}
func (t TranscriptToolCall) NormalizeCollections() TranscriptToolCall {
providerCall := ProviderToolCall{
ID: t.ID,
Name: t.Name,
Input: t.Input,
ThoughtSignature: t.ThoughtSignature,
}.NormalizeCollections()
t.ID = providerCall.ID
t.Name = providerCall.Name
t.Input = providerCall.Input
t.ThoughtSignature = providerCall.ThoughtSignature
if t.Success != nil {
success := *t.Success
t.Success = &success
}
return t
}
// ProviderToolCall projects a stored Assistant transcript call back to the
// shared provider-facing shape, deliberately excluding in-app output/success
// display fields.
func (t TranscriptToolCall) ProviderToolCall() ProviderToolCall {
t = t.NormalizeCollections()
return ProviderToolCall{
ID: t.ID,
Name: t.Name,
Input: t.Input,
ThoughtSignature: t.ThoughtSignature,
}.NormalizeCollections()
}
+44
View File
@@ -1,6 +1,7 @@
package agentcapabilities
import (
"encoding/json"
"net/http"
"slices"
"strings"
@@ -203,3 +204,46 @@ func TestNewToolGovernanceDescriptorAppliesSharedDefaults(t *testing.T) {
t.Fatalf("descriptor approval summary = %q", descriptor.ApprovalSummary)
}
}
// Stored result evidence and provider request arguments are different wire
// contracts. In particular, explicit failure cannot disappear through omitempty.
func TestTranscriptToolCallPreservesResultOutsideProviderRequests(t *testing.T) {
failed := false
call := TranscriptToolCall{ID: "read-1", Name: PulseReadToolName, Output: "NO_AGENT", Success: &failed}.NormalizeCollections()
failed = true
body, err := json.Marshal(call)
if err != nil {
t.Fatal(err)
}
var stored map[string]interface{}
if err := json.Unmarshal(body, &stored); err != nil {
t.Fatal(err)
}
if stored["output"] != "NO_AGENT" || stored["success"] != false || stored["input"] == nil {
t.Fatalf("stored result lost explicit failure or normalized input: %s", body)
}
requestBody, err := json.Marshal(call.ProviderToolCall())
if err != nil {
t.Fatal(err)
}
var request map[string]interface{}
if err := json.Unmarshal(requestBody, &request); err != nil {
t.Fatal(err)
}
for _, key := range []string{"output", "success"} {
if _, exists := request[key]; exists {
t.Fatalf("provider request retained display-only %s: %s", key, requestBody)
}
}
unknownBody, err := json.Marshal(TranscriptToolCall{Name: PulseQueryToolName}.NormalizeCollections())
if err != nil {
t.Fatal(err)
}
var unknown map[string]interface{}
if err := json.Unmarshal(unknownBody, &unknown); err != nil {
t.Fatal(err)
}
if _, exists := unknown["success"]; exists {
t.Fatalf("unknown historical status became a result: %s", unknownBody)
}
}
-223
View File
@@ -8,7 +8,6 @@ import (
"fmt"
"os"
"path/filepath"
"strconv"
"strings"
"sync"
"time"
@@ -73,165 +72,6 @@ func (a *ForecastDataAdapter) GetMetricHistory(resourceID, metric string, from,
return result, nil
}
// MetricsAdapter provides current metrics for resources.
// It implements metrics.MetricsProvider for the incident recorder.
// Uses ReadState as the sole data source (SRC-03m migration).
type MetricsAdapter struct {
readState unifiedresources.ReadState
}
// NewMetricsAdapter creates a new adapter for current metrics.
// ReadState is the sole data source for both GetMonitoredResourceIDs and
// GetCurrentMetrics. Returns nil if readState is nil.
func NewMetricsAdapter(readState unifiedresources.ReadState) *MetricsAdapter {
if readState == nil {
return nil
}
return &MetricsAdapter{readState: readState}
}
// GetMonitoredResourceIDs returns all resource IDs currently being monitored.
// This is used by the incident recorder to maintain pre-incident buffers for all resources.
// Returns both unified IDs and Proxmox source IDs so that pre-incident buffers
// are keyed by both (alert-triggered recordings use source IDs).
func (a *MetricsAdapter) GetMonitoredResourceIDs() []string {
var ids []string
for _, vm := range a.readState.VMs() {
ids = append(ids, vm.ID())
if sid := vm.SourceID(); sid != "" && sid != vm.ID() {
ids = append(ids, sid)
}
}
for _, ct := range a.readState.Containers() {
ids = append(ids, ct.ID())
if sid := ct.SourceID(); sid != "" && sid != ct.ID() {
ids = append(ids, sid)
}
}
for _, node := range a.readState.Nodes() {
ids = append(ids, node.ID())
if sid := node.SourceID(); sid != "" && sid != node.ID() {
ids = append(ids, sid)
}
}
return ids
}
// GetCurrentMetricsBatch returns current metrics for every resource
// GetCurrentMetrics can resolve, in one pass over the views, keyed by every
// ID form GetCurrentMetrics matches (unified ID, source ID, VMID string,
// name). The per-ID method scans all views per call, so a sampler asking for
// thousands of resources per tick must use this instead: first-key-wins
// mirrors the per-ID method's VM -> container -> node -> storage precedence.
func (a *MetricsAdapter) GetCurrentMetricsBatch() map[string]map[string]float64 {
out := make(map[string]map[string]float64)
put := func(metrics map[string]float64, keys ...string) {
for _, key := range keys {
if key == "" {
continue
}
if _, exists := out[key]; !exists {
out[key] = metrics
}
}
}
for _, vm := range a.readState.VMs() {
put(map[string]float64{
"cpu": vm.CPUPercent(),
"memory": vm.MemoryPercent(),
"disk": vm.DiskPercent(),
"netin": vm.NetIn(),
"netout": vm.NetOut(),
"diskread": vm.DiskRead(),
"diskwrite": vm.DiskWrite(),
}, vm.ID(), vm.SourceID(), strconv.Itoa(vm.VMID()))
}
for _, ct := range a.readState.Containers() {
put(map[string]float64{
"cpu": ct.CPUPercent(),
"memory": ct.MemoryPercent(),
"disk": ct.DiskPercent(),
"netin": ct.NetIn(),
"netout": ct.NetOut(),
"diskread": ct.DiskRead(),
"diskwrite": ct.DiskWrite(),
}, ct.ID(), ct.SourceID(), strconv.Itoa(ct.VMID()))
}
for _, node := range a.readState.Nodes() {
put(map[string]float64{
"cpu": node.CPUPercent(),
"memory": node.MemoryPercent(),
"disk": node.DiskPercent(),
}, node.ID(), node.SourceID(), node.Name())
}
for _, sp := range a.readState.StoragePools() {
put(map[string]float64{
"disk": sp.DiskPercent(),
"used": float64(sp.DiskUsed()),
"total": float64(sp.DiskTotal()),
}, sp.ID(), sp.SourceID(), sp.Name())
}
return out
}
// GetCurrentMetrics returns current metrics for a resource.
// Matches by unified ID, Proxmox source ID, VMID string, or name.
// CPU/memory/disk values are normalized to 0-100 percentage scale.
func (a *MetricsAdapter) GetCurrentMetrics(resourceID string) (map[string]float64, error) {
metrics := make(map[string]float64)
// Check VMs
for _, vm := range a.readState.VMs() {
if vm.ID() == resourceID || vm.SourceID() == resourceID || strconv.Itoa(vm.VMID()) == resourceID {
metrics["cpu"] = vm.CPUPercent()
metrics["memory"] = vm.MemoryPercent()
metrics["disk"] = vm.DiskPercent()
metrics["netin"] = vm.NetIn()
metrics["netout"] = vm.NetOut()
metrics["diskread"] = vm.DiskRead()
metrics["diskwrite"] = vm.DiskWrite()
return metrics, nil
}
}
// Check containers
for _, ct := range a.readState.Containers() {
if ct.ID() == resourceID || ct.SourceID() == resourceID || strconv.Itoa(ct.VMID()) == resourceID {
metrics["cpu"] = ct.CPUPercent()
metrics["memory"] = ct.MemoryPercent()
metrics["disk"] = ct.DiskPercent()
metrics["netin"] = ct.NetIn()
metrics["netout"] = ct.NetOut()
metrics["diskread"] = ct.DiskRead()
metrics["diskwrite"] = ct.DiskWrite()
return metrics, nil
}
}
// Check nodes
for _, node := range a.readState.Nodes() {
if node.ID() == resourceID || node.SourceID() == resourceID || node.Name() == resourceID {
metrics["cpu"] = node.CPUPercent()
metrics["memory"] = node.MemoryPercent()
metrics["disk"] = node.DiskPercent()
return metrics, nil
}
}
// Check storage
for _, sp := range a.readState.StoragePools() {
if sp.ID() == resourceID || sp.SourceID() == resourceID || sp.Name() == resourceID {
metrics["disk"] = sp.DiskPercent()
metrics["used"] = float64(sp.DiskUsed())
metrics["total"] = float64(sp.DiskTotal())
return metrics, nil
}
}
return metrics, nil
}
// CommandExecutorAdapter adapts the agent execution system to remediation.CommandExecutor.
// This allows the remediation engine to execute commands on targets.
type CommandExecutorAdapter struct {
@@ -267,69 +107,6 @@ func (e *CommandExecutionDisabledError) Error() string {
return "command execution is disabled - commands must be run manually"
}
// IncidentRecorderToolAdapter adapts metrics.IncidentRecorder to tools.IncidentRecorderProvider
type IncidentRecorderToolAdapter struct {
recorder IncidentRecorderSource
}
// IncidentRecorderSource defines what we need from an incident recorder
type IncidentRecorderSource interface {
GetWindowsForResource(resourceID string, limit int) []*IncidentWindowData
GetWindow(windowID string) *IncidentWindowData
}
// IncidentWindowData represents incident window data
type IncidentWindowData struct {
ID string
ResourceID string
ResourceName string
ResourceType string
TriggerType string
TriggerID string
StartTime time.Time
EndTime *time.Time
Status string
DataPoints []IncidentDataPointData
Summary *IncidentSummaryData
}
// IncidentDataPointData represents a single data point
type IncidentDataPointData struct {
Timestamp time.Time
Metrics map[string]float64
}
// IncidentSummaryData provides summary statistics
type IncidentSummaryData struct {
Duration time.Duration
DataPoints int
Peaks map[string]float64
Lows map[string]float64
Averages map[string]float64
Changes map[string]float64
}
// NewIncidentRecorderToolAdapter creates a new incident recorder adapter
func NewIncidentRecorderToolAdapter(recorder IncidentRecorderSource) *IncidentRecorderToolAdapter {
return &IncidentRecorderToolAdapter{recorder: recorder}
}
// GetWindowsForResource returns incident windows for a resource
func (a *IncidentRecorderToolAdapter) GetWindowsForResource(resourceID string, limit int) []*IncidentWindowData {
if a.recorder == nil {
return nil
}
return a.recorder.GetWindowsForResource(resourceID, limit)
}
// GetWindow returns a specific incident window
func (a *IncidentRecorderToolAdapter) GetWindow(windowID string) *IncidentWindowData {
if a.recorder == nil {
return nil
}
return a.recorder.GetWindow(windowID)
}
// EventCorrelatorToolAdapter adapts proxmox.EventCorrelator to tools.EventCorrelatorProvider
type EventCorrelatorToolAdapter struct {
correlator EventCorrelatorSource
@@ -6,23 +6,9 @@ import (
"testing"
"time"
"github.com/rcourtman/pulse-go-rewrite/internal/models"
"github.com/rcourtman/pulse-go-rewrite/internal/monitoring"
)
type stubIncidentRecorder struct {
windows []*IncidentWindowData
window *IncidentWindowData
}
func (s *stubIncidentRecorder) GetWindowsForResource(resourceID string, limit int) []*IncidentWindowData {
return s.windows
}
func (s *stubIncidentRecorder) GetWindow(windowID string) *IncidentWindowData {
return s.window
}
type stubEventCorrelator struct {
correlations []EventCorrelationData
events []ProxmoxEventData
@@ -69,65 +55,6 @@ func TestForecastDataAdapter_GetMetricHistory(t *testing.T) {
}
}
func TestMetricsAdapter_GetMonitoredResourceIDs(t *testing.T) {
state := models.StateSnapshot{
Nodes: []models.Node{{ID: "node/pve1", Name: "pve1", Instance: "inst1"}},
VMs: []models.VM{{ID: "qemu/100", VMID: 100, Name: "vm-1", Node: "pve1", Instance: "inst1"}},
Containers: []models.Container{{ID: "lxc/200", VMID: 200, Name: "ct-1", Node: "pve1", Instance: "inst1"}},
}
adapter := NewMetricsAdapter(readStateFromSnapshot(state))
ids := adapter.GetMonitoredResourceIDs()
// Should include both unified IDs and source IDs (3 resources × 2 IDs each = 6)
if len(ids) < 3 {
t.Fatalf("expected at least 3 IDs, got %d: %v", len(ids), ids)
}
// Verify no empty IDs
for _, id := range ids {
if id == "" {
t.Fatalf("unexpected empty ID in %v", ids)
}
}
// Verify source IDs are present (for pre-incident buffer compatibility)
idSet := make(map[string]bool, len(ids))
for _, id := range ids {
idSet[id] = true
}
if !idSet["qemu/100"] {
t.Errorf("expected source ID 'qemu/100' in monitored IDs, got %v", ids)
}
if !idSet["lxc/200"] {
t.Errorf("expected source ID 'lxc/200' in monitored IDs, got %v", ids)
}
if !idSet["node/pve1"] {
t.Errorf("expected source ID 'node/pve1' in monitored IDs, got %v", ids)
}
}
func TestIncidentRecorderToolAdapter(t *testing.T) {
adapter := NewIncidentRecorderToolAdapter(nil)
if adapter.GetWindowsForResource("res", 1) != nil {
t.Fatalf("expected nil windows for nil recorder")
}
if adapter.GetWindow("id") != nil {
t.Fatalf("expected nil window for nil recorder")
}
recorder := &stubIncidentRecorder{
windows: []*IncidentWindowData{{ID: "w1"}},
window: &IncidentWindowData{ID: "w1"},
}
adapter = NewIncidentRecorderToolAdapter(recorder)
if len(adapter.GetWindowsForResource("res", 1)) != 1 {
t.Fatalf("expected windows from recorder")
}
if adapter.GetWindow("w1") == nil {
t.Fatalf("expected window from recorder")
}
}
func TestEventCorrelatorToolAdapter(t *testing.T) {
adapter := NewEventCorrelatorToolAdapter(nil)
if adapter.GetCorrelationsForResource("res", time.Minute) != nil {
@@ -1,80 +0,0 @@
package adapters
import (
"reflect"
"strconv"
"testing"
"github.com/rcourtman/pulse-go-rewrite/internal/models"
)
// The batch method must return exactly what per-ID lookups return for every
// ID form the per-ID method matches, including colliding VMID strings where
// the VM -> container precedence decides the winner.
func TestGetCurrentMetricsBatchMatchesPerIDLookups(t *testing.T) {
state := models.StateSnapshot{
Nodes: []models.Node{
{ID: "node/pve1", Name: "pve1", Instance: "inst1", CPU: 0.35, Memory: models.Memory{Usage: 60}},
},
VMs: []models.VM{
{
ID: "qemu/100", VMID: 100, Name: "webserver", Node: "pve1", Instance: "inst1",
CPU: 45.5, Memory: models.Memory{Usage: 72.3}, Disk: models.Disk{Usage: 55},
NetworkIn: 1024, NetworkOut: 512, DiskRead: 2048, DiskWrite: 1024,
},
},
Containers: []models.Container{
{
ID: "lxc/104", VMID: 104, Name: "auth", Node: "pve1", Instance: "inst1",
CPU: 12.5, Memory: models.Memory{Usage: 30}, Disk: models.Disk{Usage: 20},
},
},
Storage: []models.Storage{
{ID: "storage/local", Name: "local", Node: "pve1", Instance: "inst1", Usage: 41, Used: 41, Total: 100},
},
}
adapter := NewMetricsAdapter(readStateFromSnapshot(state))
batch := adapter.GetCurrentMetricsBatch()
if len(batch) == 0 {
t.Fatal("batch returned no entries for a populated state")
}
for id := range batch {
perID, err := adapter.GetCurrentMetrics(id)
if err != nil {
t.Fatalf("GetCurrentMetrics(%q) error: %v", id, err)
}
if !reflect.DeepEqual(batch[id], perID) {
t.Fatalf("metrics diverged for %q:\nbatch: %+v\nper-ID: %+v", id, batch[id], perID)
}
}
// Every monitored ID must be resolvable through the batch.
for _, id := range adapter.GetMonitoredResourceIDs() {
if _, ok := batch[id]; !ok {
t.Fatalf("monitored ID %q missing from batch", id)
}
}
}
func TestGetCurrentMetricsBatchVMIDCollisionPrefersVM(t *testing.T) {
state := models.StateSnapshot{
Nodes: []models.Node{
{ID: "node/pve1", Name: "pve1", Instance: "inst1"},
},
VMs: []models.VM{
{ID: "qemu/200", VMID: 200, Name: "vm-two-hundred", Node: "pve1", Instance: "inst1", CPU: 80},
},
Containers: []models.Container{
{ID: "lxc/200", VMID: 200, Name: "ct-two-hundred", Node: "pve1", Instance: "inst1", CPU: 10},
},
}
adapter := NewMetricsAdapter(readStateFromSnapshot(state))
batch := adapter.GetCurrentMetricsBatch()
perID, _ := adapter.GetCurrentMetrics(strconv.Itoa(200))
if !reflect.DeepEqual(batch["200"], perID) {
t.Fatalf("VMID collision winner diverged:\nbatch: %+v\nper-ID: %+v", batch["200"], perID)
}
}
-428
View File
@@ -2,21 +2,10 @@ package adapters
import (
"context"
"fmt"
"testing"
"time"
"github.com/rcourtman/pulse-go-rewrite/internal/models"
"github.com/rcourtman/pulse-go-rewrite/internal/unifiedresources"
)
// readStateFromSnapshot creates a ReadState from a models.StateSnapshot for testing.
func readStateFromSnapshot(snapshot models.StateSnapshot) unifiedresources.ReadState {
rr := unifiedresources.NewRegistry(nil)
rr.IngestSnapshot(snapshot)
return rr
}
func TestForecastDataAdapter_NilHistory(t *testing.T) {
adapter := NewForecastDataAdapter(nil)
if adapter != nil {
@@ -24,295 +13,6 @@ func TestForecastDataAdapter_NilHistory(t *testing.T) {
}
}
func TestMetricsAdapter_GetCurrentMetrics_VM(t *testing.T) {
state := models.StateSnapshot{
Nodes: []models.Node{
{ID: "node/pve1", Name: "pve1", Instance: "inst1"},
},
VMs: []models.VM{
{
ID: "qemu/100",
VMID: 100,
Name: "webserver",
Node: "pve1",
Instance: "inst1",
CPU: 45.5,
Memory: models.Memory{
Usage: 72.3,
Used: 1024,
Total: 2048,
},
Disk: models.Disk{
Usage: 55.0,
Used: 5000,
Total: 10000,
},
NetworkIn: 1024000,
NetworkOut: 512000,
DiskRead: 2048000,
DiskWrite: 1024000,
},
},
}
rs := readStateFromSnapshot(state)
adapter := NewMetricsAdapter(rs)
// Get the unified resource ID from ReadState
vms := rs.VMs()
if len(vms) != 1 {
t.Fatalf("expected 1 VM, got %d", len(vms))
}
vmID := vms[0].ID()
metrics, err := adapter.GetCurrentMetrics(vmID)
if err != nil {
t.Errorf("Unexpected error: %v", err)
return
}
if metrics["cpu"] != 45.5 {
t.Errorf("Expected CPU 45.5, got %f", metrics["cpu"])
}
if metrics["memory"] != 72.3 {
t.Errorf("Expected memory 72.3, got %f", metrics["memory"])
}
if metrics["disk"] != 55.0 {
t.Errorf("Expected disk 55.0, got %f", metrics["disk"])
}
if metrics["netin"] != 1024000 {
t.Errorf("Expected netin 1024000, got %f", metrics["netin"])
}
}
func TestMetricsAdapter_GetCurrentMetrics_Container(t *testing.T) {
state := models.StateSnapshot{
Nodes: []models.Node{
{ID: "node/pve1", Name: "pve1", Instance: "inst1"},
},
Containers: []models.Container{
{
ID: "lxc/101",
VMID: 101,
Name: "container1",
Node: "pve1",
Instance: "inst1",
CPU: 25.0,
Memory: models.Memory{
Usage: 45.0,
Used: 512,
Total: 1024,
},
Disk: models.Disk{
Usage: 30.0,
Used: 3000,
Total: 10000,
},
NetworkIn: 500000,
NetworkOut: 250000,
DiskRead: 1000000,
DiskWrite: 500000,
},
},
}
rs := readStateFromSnapshot(state)
adapter := NewMetricsAdapter(rs)
containers := rs.Containers()
if len(containers) != 1 {
t.Fatalf("expected 1 container, got %d", len(containers))
}
ctID := containers[0].ID()
metrics, err := adapter.GetCurrentMetrics(ctID)
if err != nil {
t.Errorf("Unexpected error: %v", err)
return
}
if metrics["cpu"] != 25.0 {
t.Errorf("Expected CPU 25.0, got %f", metrics["cpu"])
}
if metrics["memory"] != 45.0 {
t.Errorf("Expected memory 45.0, got %f", metrics["memory"])
}
}
func TestMetricsAdapter_GetCurrentMetrics_Node(t *testing.T) {
state := models.StateSnapshot{
Nodes: []models.Node{
{
ID: "node/pve1",
Name: "pve1",
Instance: "inst1",
CPU: 25.5,
Memory: models.Memory{
Usage: 65.0,
Used: 6500,
Total: 10000,
},
Disk: models.Disk{
Usage: 40.0,
Used: 4000,
Total: 10000,
},
},
},
}
rs := readStateFromSnapshot(state)
adapter := NewMetricsAdapter(rs)
nodes := rs.Nodes()
if len(nodes) != 1 {
t.Fatalf("expected 1 node, got %d", len(nodes))
}
nodeID := nodes[0].ID()
metrics, err := adapter.GetCurrentMetrics(nodeID)
if err != nil {
t.Errorf("Unexpected error: %v", err)
return
}
if metrics["cpu"] != 25.5 {
t.Errorf("Expected CPU 25.5, got %f", metrics["cpu"])
}
if metrics["memory"] != 65.0 {
t.Errorf("Expected memory 65.0, got %f", metrics["memory"])
}
if metrics["disk"] != 40.0 {
t.Errorf("Expected disk 40.0, got %f", metrics["disk"])
}
}
func TestMetricsAdapter_GetCurrentMetrics_NodeByName(t *testing.T) {
state := models.StateSnapshot{
Nodes: []models.Node{
{
ID: "node/pve1",
Name: "pve1",
Instance: "inst1",
CPU: 25.5,
Memory: models.Memory{
Usage: 65.0,
Used: 6500,
Total: 10000,
},
Disk: models.Disk{
Usage: 40.0,
Used: 4000,
Total: 10000,
},
},
},
}
adapter := NewMetricsAdapter(readStateFromSnapshot(state))
// Node lookup by name should still work
metrics, err := adapter.GetCurrentMetrics("pve1")
if err != nil {
t.Errorf("Unexpected error: %v", err)
return
}
if metrics["cpu"] != 25.5 {
t.Errorf("Expected CPU 25.5 when matching by name, got %f", metrics["cpu"])
}
}
func TestMetricsAdapter_GetCurrentMetrics_Storage(t *testing.T) {
state := models.StateSnapshot{
Nodes: []models.Node{
{ID: "node/pve1", Name: "pve1", Instance: "inst1"},
},
Storage: []models.Storage{
{
ID: "storage/local-zfs",
Name: "local-zfs",
Node: "pve1",
Instance: "inst1",
Used: 50000000000,
Total: 100000000000,
Usage: 50.0,
},
},
}
rs := readStateFromSnapshot(state)
adapter := NewMetricsAdapter(rs)
pools := rs.StoragePools()
if len(pools) != 1 {
t.Fatalf("expected 1 storage pool, got %d", len(pools))
}
storageID := pools[0].ID()
metrics, err := adapter.GetCurrentMetrics(storageID)
if err != nil {
t.Errorf("Unexpected error: %v", err)
return
}
if metrics["disk"] != 50.0 {
t.Errorf("Expected disk 50.0, got %f", metrics["disk"])
}
if metrics["used"] != 50000000000 {
t.Errorf("Expected used 50000000000, got %f", metrics["used"])
}
if metrics["total"] != 100000000000 {
t.Errorf("Expected total 100000000000, got %f", metrics["total"])
}
}
func TestMetricsAdapter_GetCurrentMetrics_StorageByName(t *testing.T) {
state := models.StateSnapshot{
Nodes: []models.Node{
{ID: "node/pve1", Name: "pve1", Instance: "inst1"},
},
Storage: []models.Storage{
{
ID: "storage/local-zfs",
Name: "local-zfs",
Node: "pve1",
Instance: "inst1",
Used: 50000000000,
Total: 100000000000,
Usage: 50.0,
},
},
}
adapter := NewMetricsAdapter(readStateFromSnapshot(state))
// Storage lookup by name should still work
metrics, err := adapter.GetCurrentMetrics("local-zfs")
if err != nil {
t.Errorf("Unexpected error: %v", err)
return
}
if metrics["disk"] != 50.0 {
t.Errorf("Expected disk 50.0, got %f", metrics["disk"])
}
}
func TestMetricsAdapter_GetCurrentMetrics_NotFound(t *testing.T) {
state := models.StateSnapshot{}
adapter := NewMetricsAdapter(readStateFromSnapshot(state))
metrics, err := adapter.GetCurrentMetrics("nonexistent")
if err != nil {
t.Errorf("Unexpected error: %v", err)
}
if len(metrics) != 0 {
t.Errorf("Expected empty metrics, got %d entries", len(metrics))
}
}
func TestCommandExecutorAdapter_Disabled(t *testing.T) {
adapter := NewCommandExecutorAdapter()
@@ -349,131 +49,3 @@ func TestCommandExecutionDisabledError_Message(t *testing.T) {
t.Errorf("Unexpected error message: %s", msg)
}
}
func TestMetricsAdapter_NilReadState(t *testing.T) {
adapter := NewMetricsAdapter(nil)
if adapter != nil {
t.Error("Expected nil adapter for nil ReadState")
}
}
func TestMetricsAdapter_VMIDMatch(t *testing.T) {
state := models.StateSnapshot{
Nodes: []models.Node{
{ID: "node/pve1", Name: "pve1", Instance: "inst1"},
},
VMs: []models.VM{
{
ID: "qemu/100",
VMID: 100,
Name: "webserver",
Node: "pve1",
Instance: "inst1",
CPU: 45.5,
Memory: models.Memory{
Usage: 72.3,
Used: 1024,
Total: 2048,
},
Disk: models.Disk{
Usage: 55.0,
Used: 5000,
Total: 10000,
},
},
},
}
adapter := NewMetricsAdapter(readStateFromSnapshot(state))
// Test lookup by VMID string
metrics, err := adapter.GetCurrentMetrics(fmt.Sprintf("%d", 100))
if err != nil {
t.Errorf("Unexpected error: %v", err)
return
}
if metrics["cpu"] != 45.5 {
t.Errorf("Expected CPU 45.5 when matching by VMID, got %f", metrics["cpu"])
}
}
func TestMetricsAdapter_IDConsistency(t *testing.T) {
// Verify that GetMonitoredResourceIDs returns IDs that work with GetCurrentMetrics
state := models.StateSnapshot{
Nodes: []models.Node{
{ID: "node/pve1", Name: "pve1", Instance: "inst1"},
},
VMs: []models.VM{
{
ID: "qemu/100", VMID: 100, Name: "vm1", Node: "pve1", Instance: "inst1",
CPU: 50.0, Memory: models.Memory{Usage: 60.0, Used: 600, Total: 1000},
Disk: models.Disk{Usage: 70.0, Used: 700, Total: 1000},
},
},
}
adapter := NewMetricsAdapter(readStateFromSnapshot(state))
ids := adapter.GetMonitoredResourceIDs()
if len(ids) < 1 {
t.Fatalf("expected at least 1 ID, got %d", len(ids))
}
// Each ID from GetMonitoredResourceIDs should be usable with GetCurrentMetrics
foundVM := false
for _, id := range ids {
metrics, err := adapter.GetCurrentMetrics(id)
if err != nil {
t.Errorf("GetCurrentMetrics(%q) error: %v", id, err)
continue
}
if cpu, ok := metrics["cpu"]; ok && cpu == 50.0 {
foundVM = true
}
}
if !foundVM {
t.Error("Expected to find VM metrics via GetMonitoredResourceIDs() IDs")
}
}
func TestMetricsAdapter_SourceIDMatch(t *testing.T) {
// Verify that GetCurrentMetrics can find resources by their Proxmox source ID
state := models.StateSnapshot{
Nodes: []models.Node{
{ID: "node/pve1", Name: "pve1", Instance: "inst1"},
},
VMs: []models.VM{
{
ID: "qemu/100", VMID: 100, Name: "webserver", Node: "pve1", Instance: "inst1",
CPU: 45.5, Memory: models.Memory{Usage: 72.3, Used: 1024, Total: 2048},
Disk: models.Disk{Usage: 55.0, Used: 5000, Total: 10000},
},
},
Storage: []models.Storage{
{
ID: "storage/local-zfs", Name: "local-zfs", Node: "pve1", Instance: "inst1",
Used: 50000000000, Total: 100000000000, Usage: 50.0,
},
},
}
adapter := NewMetricsAdapter(readStateFromSnapshot(state))
// Lookup VM by Proxmox source ID
metrics, err := adapter.GetCurrentMetrics("qemu/100")
if err != nil {
t.Fatalf("Unexpected error: %v", err)
}
if metrics["cpu"] != 45.5 {
t.Errorf("Expected CPU 45.5 via source ID, got %f", metrics["cpu"])
}
// Lookup storage by Proxmox source ID
metrics, err = adapter.GetCurrentMetrics("storage/local-zfs")
if err != nil {
t.Fatalf("Unexpected error: %v", err)
}
if metrics["disk"] != 50.0 {
t.Errorf("Expected disk 50.0 via source ID, got %f", metrics["disk"])
}
}
+37 -37
View File
@@ -60,7 +60,7 @@ type (
AgentProfileManager = tools.AgentProfileManager
FindingsManager = tools.FindingsManager
MetadataUpdater = tools.MetadataUpdater
IncidentRecorderProvider = tools.IncidentRecorderProvider
IncidentArchiveProvider = tools.IncidentArchiveProvider
EventCorrelatorProvider = tools.EventCorrelatorProvider
KnowledgeStoreProvider = tools.KnowledgeStoreProvider
AssistantDiscoveryProvider = tools.DiscoveryProvider
@@ -1418,15 +1418,15 @@ func marshalAssistantInventoryTopologyContext(topology tools.TopologyResponse) (
}
for _, node := range topology.Proxmox.Nodes {
nodeContext := assistantInventoryProxmoxNode{
AnswerLabel: assistantInventoryNodeAnswerLabel(node.Name),
Name: node.Name,
Status: node.Status,
AgentConnected: node.AgentConnected,
CanExecute: node.CanExecute,
VMCount: node.VMCount,
ContainerCount: node.ContainerCount,
VMs: make([]assistantInventoryWorkload, 0, len(node.VMs)),
Containers: make([]assistantInventoryWorkload, 0, len(node.Containers)),
AnswerLabel: assistantInventoryNodeAnswerLabel(node.Name),
Name: node.Name,
Status: node.Status,
CommandAgentConnected: node.CommandAgentConnected,
CanExecute: node.CanExecute,
VMCount: node.VMCount,
ContainerCount: node.ContainerCount,
VMs: make([]assistantInventoryWorkload, 0, len(node.VMs)),
Containers: make([]assistantInventoryWorkload, 0, len(node.Containers)),
}
for _, vm := range node.VMs {
nodeContext.VMs = append(nodeContext.VMs, assistantInventoryWorkload{
@@ -1452,14 +1452,14 @@ func marshalAssistantInventoryTopologyContext(topology tools.TopologyResponse) (
}
for _, host := range topology.Docker.Hosts {
hostContext := assistantInventoryDockerHost{
AnswerLabel: firstNonEmptyString(host.DisplayName, host.Hostname),
Hostname: host.Hostname,
DisplayName: host.DisplayName,
AgentConnected: host.AgentConnected,
CanExecute: host.CanExecute,
ContainerCount: host.ContainerCount,
RunningCount: host.RunningCount,
Containers: make([]assistantInventoryAppContainer, 0, len(host.Containers)),
AnswerLabel: firstNonEmptyString(host.DisplayName, host.Hostname),
Hostname: host.Hostname,
DisplayName: host.DisplayName,
CommandAgentConnected: host.CommandAgentConnected,
CanExecute: host.CanExecute,
ContainerCount: host.ContainerCount,
RunningCount: host.RunningCount,
Containers: make([]assistantInventoryAppContainer, 0, len(host.Containers)),
}
for _, container := range host.Containers {
hostContext.Containers = append(hostContext.Containers, assistantInventoryAppContainer{
@@ -1547,15 +1547,15 @@ type assistantInventoryKubernetesTopology struct {
}
type assistantInventoryProxmoxNode struct {
AnswerLabel string `json:"answer_label"`
Name string `json:"name"`
Status string `json:"status"`
AgentConnected bool `json:"agent_connected,omitempty"`
CanExecute bool `json:"can_execute,omitempty"`
VMCount int `json:"vm_count"`
ContainerCount int `json:"container_count"`
VMs []assistantInventoryWorkload `json:"vms"`
Containers []assistantInventoryWorkload `json:"containers"`
AnswerLabel string `json:"answer_label"`
Name string `json:"name"`
Status string `json:"status"`
CommandAgentConnected *bool `json:"command_agent_connected,omitempty"`
CanExecute *bool `json:"can_execute,omitempty"`
VMCount int `json:"vm_count"`
ContainerCount int `json:"container_count"`
VMs []assistantInventoryWorkload `json:"vms"`
Containers []assistantInventoryWorkload `json:"containers"`
}
type assistantInventoryWorkload struct {
@@ -1568,14 +1568,14 @@ type assistantInventoryWorkload struct {
}
type assistantInventoryDockerHost struct {
AnswerLabel string `json:"answer_label"`
Hostname string `json:"hostname"`
DisplayName string `json:"display_name,omitempty"`
AgentConnected bool `json:"agent_connected,omitempty"`
CanExecute bool `json:"can_execute,omitempty"`
ContainerCount int `json:"container_count"`
RunningCount int `json:"running_count"`
Containers []assistantInventoryAppContainer `json:"containers"`
AnswerLabel string `json:"answer_label"`
Hostname string `json:"hostname"`
DisplayName string `json:"display_name,omitempty"`
CommandAgentConnected *bool `json:"command_agent_connected,omitempty"`
CanExecute *bool `json:"can_execute,omitempty"`
ContainerCount int `json:"container_count"`
RunningCount int `json:"running_count"`
Containers []assistantInventoryAppContainer `json:"containers"`
}
type assistantInventoryAppContainer struct {
@@ -3598,11 +3598,11 @@ func (s *Service) SetMetadataUpdater(updater MetadataUpdater) {
}
}
func (s *Service) SetIncidentRecorderProvider(provider IncidentRecorderProvider) {
func (s *Service) SetIncidentArchiveProvider(provider IncidentArchiveProvider) {
s.mu.Lock()
defer s.mu.Unlock()
if s.executor != nil {
s.executor.SetIncidentRecorderProvider(provider)
s.executor.SetIncidentArchiveProvider(provider)
}
}
+1 -1
View File
@@ -16,7 +16,7 @@ func TestServiceSettersAndAutonomousMode(t *testing.T) {
agenticLoop: loop,
}
service.SetIncidentRecorderProvider(nil)
service.SetIncidentArchiveProvider(nil)
service.SetEventCorrelatorProvider(nil)
service.SetKnowledgeStoreProvider(nil)
@@ -0,0 +1,35 @@
package chat
import (
"encoding/json"
"strings"
"testing"
"github.com/rcourtman/pulse-go-rewrite/internal/models"
"github.com/rcourtman/pulse-go-rewrite/internal/unifiedresources"
)
func TestAssistantInventoryDoesNotInventCommandConnectionObservations(t *testing.T) {
registry := unifiedresources.NewRegistry(nil)
registry.IngestSnapshot(models.StateSnapshot{
Nodes: []models.Node{{ID: "node-one", Name: "node-one", Status: "online"}},
DockerHosts: []models.DockerHost{{ID: "host-one", Hostname: "observed-host", Status: "online", Containers: []models.DockerContainer{{ID: "app-one", Name: "observed-app", State: "running"}}}},
})
raw, err := marshalAssistantInventoryTopologyContextFromReadState(registry)
if err != nil {
t.Fatal(err)
}
for _, field := range []string{"agent_connected", "command_agent_connected", "can_execute", "nodes_with_agents", "docker_hosts_with_agents", "nodes_with_command_agents", "docker_hosts_with_command_agents"} {
if strings.Contains(raw, `"`+field+`"`) {
t.Fatalf("inventory seed invented %s without observing command connections: %s", field, raw)
}
}
var decoded map[string]any
if err := json.Unmarshal([]byte(raw), &decoded); err != nil {
t.Fatal(err)
}
host := decoded["docker"].(map[string]any)["hosts"].([]any)[0].(map[string]any)
if host["hostname"] != "observed-host" || host["container_count"] != float64(1) {
t.Fatalf("monitoring inventory was lost: %+v", host)
}
}
+2 -40
View File
@@ -135,38 +135,13 @@ func (m Message) ClientSafe() Message {
return m
}
// ToolCall represents a tool invocation
type ToolCall struct {
ID string `json:"id"`
Name string `json:"name"`
Input map[string]interface{} `json:"input"`
Output string `json:"output,omitempty"`
Success *bool `json:"success,omitempty"`
ThoughtSignature json.RawMessage `json:"thought_signature,omitempty"`
}
// ToolCall is the canonical result-bearing product transcript call.
type ToolCall = agentcapabilities.TranscriptToolCall
func EmptyToolCall() ToolCall {
return ToolCall{}.NormalizeCollections()
}
func (t ToolCall) NormalizeCollections() ToolCall {
providerCall := agentcapabilities.ProviderToolCall{
ID: t.ID,
Name: t.Name,
Input: t.Input,
ThoughtSignature: t.ThoughtSignature,
}.NormalizeCollections()
t.ID = providerCall.ID
t.Name = providerCall.Name
t.Input = providerCall.Input
t.ThoughtSignature = providerCall.ThoughtSignature
if t.Success != nil {
success := *t.Success
t.Success = &success
}
return t
}
// ToolCallFromProvider stores a provider-facing tool call in the richer
// Assistant transcript shape used for in-app history.
func ToolCallFromProvider(tc agentcapabilities.ProviderToolCall) ToolCall {
@@ -183,19 +158,6 @@ func ToolCallFromProvider(tc agentcapabilities.ProviderToolCall) ToolCall {
}.NormalizeCollections()
}
// ProviderToolCall projects a stored Assistant transcript call back to the
// shared provider-facing shape, deliberately excluding in-app output/success
// display fields.
func (t ToolCall) ProviderToolCall() agentcapabilities.ProviderToolCall {
t = t.NormalizeCollections()
return agentcapabilities.ProviderToolCall{
ID: t.ID,
Name: t.Name,
Input: t.Input,
ThoughtSignature: t.ThoughtSignature,
}.NormalizeCollections()
}
// ToolResult represents the result of a tool execution. It aliases the shared
// Pulse Intelligence provider-result shape so stored Assistant transcripts and
// provider turns do not drift on tool result JSON.
+6
View File
@@ -82,11 +82,17 @@ var providerPrices = map[string][]modelPrice{
flatPriceAsOf("anthropic/claude-opus-4.8", 5.00, 25.00, "2026-07-14"),
flatPriceAsOf("anthropic/claude-sonnet-5", 2.00, 10.00, "2026-07-14"),
flatPriceAsOf("deepseek/deepseek-v4-flash", 0.09, 0.18, "2026-07-14"),
// Introductory standard rates through 2026-12-31. Recheck when the
// published standard price changes on 2027-01-01. Batch/alias routes
// are deliberately not covered by this exact model ID.
flatPriceAsOf("google/gemini-3.8-flash", 0.75, 3.75, "2026-09-06"),
flatPriceAsOf("nvidia/nemotron-3.5-lightning:free", 0, 0, "2026-08-14"),
flatPriceAsOf("nvidia/nemotron-3-super-120b-a12b:free", 0, 0, "2026-08-14"),
flatPriceAsOf("nvidia/nemotron-3-ultra-550b-a55b:free", 0, 0, "2026-08-15"),
},
"gemini": {
// Standard introductory rates, verified 2026-09-06. Recheck 2027-01-01.
flatPriceAsOf("gemini-3.8-flash", 0.75, 3.75, "2026-09-06"),
// Gemini Developer API standard paid-tier pricing, checked from
// https://ai.google.dev/gemini-api/docs/pricing on 2026-06-04.
flatPrice("gemini-3.5-flash*", 1.50, 9.00),
+46
View File
@@ -0,0 +1,46 @@
package cost
import (
"math"
"testing"
)
func TestGemini38FlashReviewedRoutePricing(t *testing.T) {
for _, route := range []struct{ provider, model string }{
{"gemini", "gemini-3.8-flash"},
{"openrouter", "google/gemini-3.8-flash"},
} {
t.Run(route.provider, func(t *testing.T) {
provider, model := ResolveProviderAndModel(route.provider, route.provider+":"+route.model, "gemini-3.8-flash")
if provider != route.provider || model != route.model {
t.Fatalf("usage lost the requested billing route: %s:%s", provider, model)
}
// Token counts from the live unhealthy-container qualification.
usd, known, price := EstimateUSD(provider, model, 16068, 778)
if !known || math.Abs(usd-0.0149685) > 1e-10 {
t.Fatalf("live usage estimate = %f, known=%t", usd, known)
}
if price.InputUSDPerMTok != 0.75 || price.OutputUSDPerMTok != 3.75 || price.AsOf != "2026-09-06" {
t.Fatalf("reviewed standard rates/date missing: %+v", price)
}
usd, known, _ = EstimateUSD(provider, model, 0, 0)
if !known || usd != 0 {
t.Fatalf("zero observed usage = %f, known=%t", usd, known)
}
})
}
}
func TestGemini38OpenRouterPricingDoesNotGuessVariantRates(t *testing.T) {
for _, model := range []string{
"google/gemini-3.8-flash:batch",
"google/gemini-3.8-flash:free",
"google/gemini-3.8-flash-preview",
"google/gemini-3.8-flash-cyber",
"google/gemini-3.9-flash",
} {
if usd, known, _ := EstimateUSD("openrouter", model, 16068, 778); known || usd != 0 {
t.Errorf("unreviewed route %q received an estimate: %f, known=%t", model, usd, known)
}
}
}
-325
View File
@@ -1,325 +0,0 @@
// Package ai provides AI-powered infrastructure analysis.
package ai
import (
"sync"
"time"
"github.com/rcourtman/pulse-go-rewrite/internal/ai/memory"
"github.com/rcourtman/pulse-go-rewrite/internal/alerts"
"github.com/rcourtman/pulse-go-rewrite/internal/metrics"
"github.com/rs/zerolog/log"
)
// IncidentCoordinatorConfig configures the incident coordinator
type IncidentCoordinatorConfig struct {
PreBuffer time.Duration // History to capture before incident (default: 5 min)
PostDuration time.Duration // How long to record after trigger (default: 10 min)
MaxConcurrent int // Maximum concurrent incident recordings (default: 50)
EnableRecorder bool // Whether to enable high-frequency recording
}
// DefaultIncidentCoordinatorConfig returns sensible defaults
func DefaultIncidentCoordinatorConfig() IncidentCoordinatorConfig {
return IncidentCoordinatorConfig{
PreBuffer: 5 * time.Minute,
PostDuration: 10 * time.Minute,
MaxConcurrent: 50,
EnableRecorder: true,
}
}
// IncidentCoordinator coordinates incident recording between the metrics.IncidentRecorder
// (for high-frequency data capture) and memory.IncidentStore (for incident timeline tracking).
type IncidentCoordinator struct {
mu sync.RWMutex
config IncidentCoordinatorConfig
// Components
recorder *metrics.IncidentRecorder // High-frequency metrics capture
incidentStore *memory.IncidentStore // Incident timeline tracking
// Active incidents - maps alert ID to window ID
activeIncidents map[string]activeIncident
// Control
running bool
}
type activeIncident struct {
windowID string
resourceID string
startedAt time.Time
stopTimer *time.Timer // Timer to auto-stop recording after post-duration
}
// NewIncidentCoordinator creates a new incident coordinator
func NewIncidentCoordinator(cfg IncidentCoordinatorConfig) *IncidentCoordinator {
if cfg.PreBuffer <= 0 {
cfg.PreBuffer = 5 * time.Minute
}
if cfg.PostDuration <= 0 {
cfg.PostDuration = 10 * time.Minute
}
if cfg.MaxConcurrent <= 0 {
cfg.MaxConcurrent = 50
}
return &IncidentCoordinator{
config: cfg,
activeIncidents: make(map[string]activeIncident),
}
}
// SetRecorder sets the metrics incident recorder
func (c *IncidentCoordinator) SetRecorder(recorder *metrics.IncidentRecorder) {
c.mu.Lock()
defer c.mu.Unlock()
c.recorder = recorder
}
// SetIncidentStore sets the incident timeline store
func (c *IncidentCoordinator) SetIncidentStore(store *memory.IncidentStore) {
c.mu.Lock()
defer c.mu.Unlock()
c.incidentStore = store
}
// Start starts the incident coordinator
func (c *IncidentCoordinator) Start() {
c.mu.Lock()
defer c.mu.Unlock()
if c.running {
return
}
c.running = true
log.Info().Msg("incident coordinator started")
}
// Stop stops the incident coordinator
func (c *IncidentCoordinator) Stop() {
c.mu.Lock()
defer c.mu.Unlock()
if !c.running {
return
}
c.running = false
// Stop all active timers
for _, inc := range c.activeIncidents {
if inc.stopTimer != nil {
inc.stopTimer.Stop()
}
}
c.activeIncidents = make(map[string]activeIncident)
log.Info().Msg("incident coordinator stopped")
}
// OnAlertFired is called when an alert fires - starts incident recording
func (c *IncidentCoordinator) OnAlertFired(alert *alerts.Alert) {
if alert == nil {
return
}
c.mu.Lock()
defer c.mu.Unlock()
if !c.running {
return
}
// Check if we already have an active incident for this alert
if _, exists := c.activeIncidents[alert.ID]; exists {
log.Debug().
Str("alert_identifier", alert.ID).
Msg("Incident already being recorded for this alert")
return
}
// Check concurrent limit
if len(c.activeIncidents) >= c.config.MaxConcurrent {
log.Warn().
Str("alert_identifier", alert.ID).
Int("active_count", len(c.activeIncidents)).
Msg("Incident coordinator at capacity, skipping new incident")
return
}
// Start high-frequency recording if enabled and recorder available
var windowID string
if c.config.EnableRecorder && c.recorder != nil {
windowID = c.recorder.StartRecording(
alert.ResourceID,
alert.ResourceName,
"", // resourceType - we don't always have this
"alert",
alert.ID,
)
}
// Record in incident store
if c.incidentStore != nil {
c.incidentStore.RecordAlertFired(alert)
}
// Track the active incident
inc := activeIncident{
windowID: windowID,
resourceID: alert.ResourceID,
startedAt: time.Now(),
}
c.activeIncidents[alert.ID] = inc
log.Info().
Str("alert_identifier", alert.ID).
Str("resource_id", alert.ResourceID).
Str("window_id", windowID).
Msg("Incident coordinator: Started incident recording")
}
// OnAlertCleared is called when an alert clears - schedules recording stop
func (c *IncidentCoordinator) OnAlertCleared(alert *alerts.Alert) {
if alert == nil {
return
}
c.mu.Lock()
inc, exists := c.activeIncidents[alert.ID]
if !exists {
c.mu.Unlock()
return
}
// Record resolution in incident store
if c.incidentStore != nil {
c.incidentStore.RecordAlertResolved(alert, time.Now())
}
// If no recorder or no window, just clean up immediately
if c.recorder == nil || inc.windowID == "" {
delete(c.activeIncidents, alert.ID)
c.mu.Unlock()
return
}
// Schedule stop after post-duration to capture post-incident data
alertID := alert.ID
timer := time.AfterFunc(c.config.PostDuration, func() {
c.stopIncidentRecording(alertID)
})
inc.stopTimer = timer
c.activeIncidents[alert.ID] = inc
c.mu.Unlock()
log.Info().
Str("alert_identifier", alert.ID).
Str("window_id", inc.windowID).
Dur("post_duration", c.config.PostDuration).
Msg("Incident coordinator: Alert cleared, scheduled recording stop")
}
// stopIncidentRecording stops recording for a specific alert
func (c *IncidentCoordinator) stopIncidentRecording(alertID string) {
c.mu.Lock()
defer c.mu.Unlock()
inc, exists := c.activeIncidents[alertID]
if !exists {
return
}
// Stop the recorder
if c.recorder != nil && inc.windowID != "" {
c.recorder.StopRecording(inc.windowID)
}
// Clean up
if inc.stopTimer != nil {
inc.stopTimer.Stop()
}
delete(c.activeIncidents, alertID)
log.Info().
Str("alert_identifier", alertID).
Str("window_id", inc.windowID).
Msg("Incident coordinator: Stopped incident recording")
}
// OnAnomalyDetected is called when an anomaly is detected - starts focused recording
func (c *IncidentCoordinator) OnAnomalyDetected(resourceID, resourceType, metric string, severity string) {
c.mu.Lock()
defer c.mu.Unlock()
if !c.running || !c.config.EnableRecorder || c.recorder == nil {
return
}
// Create a pseudo-alert ID for the anomaly
anomalyID := "anomaly-" + resourceID + "-" + metric
// Check if we already have an active incident for this
if _, exists := c.activeIncidents[anomalyID]; exists {
return
}
// Check concurrent limit
if len(c.activeIncidents) >= c.config.MaxConcurrent {
return
}
// Start recording
windowID := c.recorder.StartRecording(
resourceID,
"", // name
resourceType,
"anomaly",
anomalyID,
)
// Track the active incident
c.activeIncidents[anomalyID] = activeIncident{
windowID: windowID,
resourceID: resourceID,
startedAt: time.Now(),
}
// Schedule auto-stop after post-duration (anomalies don't have "clear" events)
timer := time.AfterFunc(c.config.PostDuration, func() {
c.stopIncidentRecording(anomalyID)
})
c.activeIncidents[anomalyID] = activeIncident{
windowID: windowID,
resourceID: resourceID,
startedAt: time.Now(),
stopTimer: timer,
}
log.Info().
Str("resource_id", resourceID).
Str("metric", metric).
Str("severity", severity).
Str("window_id", windowID).
Msg("Incident coordinator: Started anomaly recording")
}
// GetActiveIncidentCount returns the number of active incidents being recorded
func (c *IncidentCoordinator) GetActiveIncidentCount() int {
c.mu.RLock()
defer c.mu.RUnlock()
return len(c.activeIncidents)
}
// GetRecordingWindowID returns the recording window ID for an alert
func (c *IncidentCoordinator) GetRecordingWindowID(alertID string) string {
c.mu.RLock()
defer c.mu.RUnlock()
if inc, exists := c.activeIncidents[alertID]; exists {
return inc.windowID
}
return ""
}
@@ -1,179 +0,0 @@
package ai
import (
"testing"
"time"
"github.com/rcourtman/pulse-go-rewrite/internal/ai/memory"
"github.com/rcourtman/pulse-go-rewrite/internal/alerts"
"github.com/rcourtman/pulse-go-rewrite/internal/metrics"
)
func TestNewIncidentCoordinator_DefaultFallbacks(t *testing.T) {
cfg := IncidentCoordinatorConfig{
PreBuffer: -1 * time.Second,
PostDuration: 0,
MaxConcurrent: 0,
}
coord := NewIncidentCoordinator(cfg)
defaults := DefaultIncidentCoordinatorConfig()
if coord.config.PreBuffer != defaults.PreBuffer {
t.Fatalf("expected default pre-buffer %v, got %v", defaults.PreBuffer, coord.config.PreBuffer)
}
if coord.config.PostDuration != defaults.PostDuration {
t.Fatalf("expected default post-duration %v, got %v", defaults.PostDuration, coord.config.PostDuration)
}
if coord.config.MaxConcurrent != defaults.MaxConcurrent {
t.Fatalf("expected default max concurrent %d, got %d", defaults.MaxConcurrent, coord.config.MaxConcurrent)
}
if coord.activeIncidents == nil {
t.Fatal("expected active incident map to be initialized")
}
}
func TestIncidentCoordinator_OnAlertFired_RequiresRunning(t *testing.T) {
coord := NewIncidentCoordinator(DefaultIncidentCoordinatorConfig())
store := memory.NewIncidentStore(memory.IncidentStoreConfig{})
coord.SetIncidentStore(store)
alert := &alerts.Alert{
ID: "alert-requires-running",
ResourceID: "resource-requires-running",
ResourceName: "resource-requires-running",
}
coord.OnAlertFired(nil)
coord.OnAlertFired(alert)
if got := coord.GetActiveIncidentCount(); got != 0 {
t.Fatalf("expected no active incidents while coordinator is stopped, got %d", got)
}
if got := len(store.ListIncidentsByResource(alert.ResourceID, 0)); got != 0 {
t.Fatalf("expected no incident-store records while stopped, got %d", got)
}
coord.Start()
coord.OnAlertFired(alert)
if got := coord.GetActiveIncidentCount(); got != 1 {
t.Fatalf("expected one active incident after start, got %d", got)
}
if got := len(store.ListIncidentsByResource(alert.ResourceID, 0)); got != 1 {
t.Fatalf("expected one incident-store record after start, got %d", got)
}
}
func TestIncidentCoordinator_OnAlertCleared_ImmediateCleanupWithoutRecorder(t *testing.T) {
coord := NewIncidentCoordinator(DefaultIncidentCoordinatorConfig())
store := memory.NewIncidentStore(memory.IncidentStoreConfig{})
coord.SetIncidentStore(store)
coord.Start()
alert := &alerts.Alert{
ID: "alert-no-recorder",
ResourceID: "resource-no-recorder",
ResourceName: "resource-no-recorder",
}
coord.OnAlertFired(alert)
if got := coord.GetActiveIncidentCount(); got != 1 {
t.Fatalf("expected one active incident after fire, got %d", got)
}
coord.OnAlertCleared(nil)
coord.OnAlertCleared(&alerts.Alert{ID: "missing"})
if got := coord.GetActiveIncidentCount(); got != 1 {
t.Fatalf("expected active incident to remain after nil/missing clears, got %d", got)
}
coord.OnAlertCleared(alert)
if got := coord.GetActiveIncidentCount(); got != 0 {
t.Fatalf("expected incident to be cleaned up immediately without recorder, got %d", got)
}
incidents := store.ListIncidentsByResource(alert.ResourceID, 0)
if len(incidents) != 1 {
t.Fatalf("expected exactly one stored incident, got %d", len(incidents))
}
if incidents[0].Status != memory.IncidentStatusResolved {
t.Fatalf("expected incident status %q, got %q", memory.IncidentStatusResolved, incidents[0].Status)
}
if incidents[0].ClosedAt == nil {
t.Fatal("expected incident ClosedAt to be set on clear")
}
}
func TestIncidentCoordinator_OnAnomalyDetected_DuplicateAndCapacityAndStop(t *testing.T) {
cfg := DefaultIncidentCoordinatorConfig()
cfg.MaxConcurrent = 1
cfg.PostDuration = time.Minute
coord := NewIncidentCoordinator(cfg)
recCfg := metrics.DefaultIncidentRecorderConfig()
recCfg.SampleInterval = 10 * time.Millisecond
recorder := metrics.NewIncidentRecorder(recCfg)
recorder.SetMetricsProvider(&MockMetricsProvider{data: map[string]map[string]float64{
"resource-anomaly": {"cpu": 95},
}})
recorder.Start()
defer recorder.Stop()
coord.SetRecorder(recorder)
coord.Start()
coord.OnAnomalyDetected("resource-anomaly", "agent", "cpu", "critical")
if got := coord.GetActiveIncidentCount(); got != 1 {
t.Fatalf("expected one anomaly incident, got %d", got)
}
coord.OnAnomalyDetected("resource-anomaly", "agent", "cpu", "critical")
if got := coord.GetActiveIncidentCount(); got != 1 {
t.Fatalf("expected duplicate anomaly to be ignored, got %d", got)
}
coord.OnAnomalyDetected("resource-anomaly", "agent", "memory", "warning")
if got := coord.GetActiveIncidentCount(); got != 1 {
t.Fatalf("expected anomaly over capacity to be ignored, got %d", got)
}
coord.Stop()
if got := coord.GetActiveIncidentCount(); got != 0 {
t.Fatalf("expected active incidents to be cleared on stop, got %d", got)
}
coord.OnAnomalyDetected("resource-anomaly", "agent", "cpu", "critical")
if got := coord.GetActiveIncidentCount(); got != 0 {
t.Fatalf("expected anomalies to be ignored while stopped, got %d", got)
}
}
func TestIncidentCoordinator_OnAnomalyDetected_CanonicalizesLegacyHostAlias(t *testing.T) {
cfg := DefaultIncidentCoordinatorConfig()
cfg.PostDuration = time.Minute
coord := NewIncidentCoordinator(cfg)
recCfg := metrics.DefaultIncidentRecorderConfig()
recorder := metrics.NewIncidentRecorder(recCfg)
recorder.SetMetricsProvider(&MockMetricsProvider{data: map[string]map[string]float64{
"resource-anomaly": {"cpu": 95},
}})
recorder.Start()
defer recorder.Stop()
coord.SetRecorder(recorder)
coord.Start()
coord.OnAnomalyDetected("resource-anomaly", "host", "cpu", "critical")
windows := recorder.GetWindowsForResource("resource-anomaly", 0)
if len(windows) != 1 {
t.Fatalf("expected one anomaly window, got %d", len(windows))
}
if windows[0].ResourceType != "agent" {
t.Fatalf("expected anomaly recording resource type to be canonicalized to agent, got %q", windows[0].ResourceType)
}
}
-219
View File
@@ -1,219 +0,0 @@
package ai
import (
"sync"
"testing"
"time"
"github.com/rcourtman/pulse-go-rewrite/internal/alerts"
"github.com/rcourtman/pulse-go-rewrite/internal/metrics"
)
// MockMetricsProvider for the real IncidentRecorder
type MockMetricsProvider struct {
mu sync.Mutex
data map[string]map[string]float64
}
func (m *MockMetricsProvider) GetCurrentMetrics(resourceID string) (map[string]float64, error) {
m.mu.Lock()
defer m.mu.Unlock()
if val, ok := m.data[resourceID]; ok {
return val, nil
}
return map[string]float64{"cpu": 10.0}, nil
}
func (m *MockMetricsProvider) GetMonitoredResourceIDs() []string {
m.mu.Lock()
defer m.mu.Unlock()
keys := make([]string, 0, len(m.data))
for k := range m.data {
keys = append(keys, k)
}
return keys
}
func TestIncidentCoordinator_Lifecycle(t *testing.T) {
cfg := DefaultIncidentCoordinatorConfig()
coord := NewIncidentCoordinator(cfg)
if coord.running {
t.Error("Coordinator should not be running initially")
}
coord.Start()
if !coord.running {
t.Error("Coordinator should be running after Start()")
}
// Double start should be safe
coord.Start()
coord.Stop()
if coord.running {
t.Error("Coordinator should not be running after Stop()")
}
// Double stop should be safe
coord.Stop()
}
func TestIncidentCoordinator_OnAlertFired(t *testing.T) {
cfg := DefaultIncidentCoordinatorConfig()
coord := NewIncidentCoordinator(cfg)
// Create real recorder with mock provider
recCfg := metrics.DefaultIncidentRecorderConfig()
recCfg.SampleInterval = 50 * time.Millisecond // fast sampling
recorder := metrics.NewIncidentRecorder(recCfg)
provider := &MockMetricsProvider{data: make(map[string]map[string]float64)}
recorder.SetMetricsProvider(provider)
recorder.Start() // Recorder must be started
defer recorder.Stop()
// Inject recorder into coordinator
coord.SetRecorder(recorder)
coord.Start()
alert := &alerts.Alert{
ID: "alert-1",
ResourceID: "res-1",
}
coord.OnAlertFired(alert)
if coord.GetActiveIncidentCount() != 1 {
t.Errorf("Expected 1 active incident, got %d", coord.GetActiveIncidentCount())
}
// Verify recorder has active window
wid := coord.GetRecordingWindowID("alert-1")
if wid == "" {
t.Error("Expected valid window ID")
}
// Fire same alert again - should be ignored
coord.OnAlertFired(alert)
if coord.GetActiveIncidentCount() != 1 {
t.Errorf("Expected count to remain 1, got %d", coord.GetActiveIncidentCount())
}
// Fire another alert
alert2 := &alerts.Alert{
ID: "alert-2",
ResourceID: "res-2",
}
coord.OnAlertFired(alert2)
if coord.GetActiveIncidentCount() != 2 {
t.Errorf("Expected 2 active incidents, got %d", coord.GetActiveIncidentCount())
}
}
func TestIncidentCoordinator_OnAlertCleared(t *testing.T) {
cfg := DefaultIncidentCoordinatorConfig()
cfg.PostDuration = 50 * time.Millisecond // fast for testing
coord := NewIncidentCoordinator(cfg)
// Create real recorder
recCfg := metrics.DefaultIncidentRecorderConfig()
recorder := metrics.NewIncidentRecorder(recCfg)
provider := &MockMetricsProvider{data: make(map[string]map[string]float64)}
recorder.SetMetricsProvider(provider)
recorder.Start()
defer recorder.Stop()
coord.SetRecorder(recorder)
coord.Start()
alert := &alerts.Alert{ID: "alert-1", ResourceID: "res-1"}
coord.OnAlertFired(alert)
if coord.GetActiveIncidentCount() != 1 {
t.Fatal("Failed to start incident")
}
// Clear alert
coord.OnAlertCleared(alert)
// Since we have a recorder and postDuration is 50ms, it should NOT be removed immediately
if coord.GetActiveIncidentCount() != 1 {
t.Error("Incident should NOT be removed immediately when recorder is active")
}
// Wait for post duration
time.Sleep(100 * time.Millisecond)
// Now it should be removed (via time.AfterFunc callback)
if coord.GetActiveIncidentCount() != 0 {
t.Errorf("Incident should be removed after post duration, count=%d", coord.GetActiveIncidentCount())
}
}
func TestIncidentCoordinator_MaxConcurrent(t *testing.T) {
cfg := DefaultIncidentCoordinatorConfig()
cfg.MaxConcurrent = 1
coord := NewIncidentCoordinator(cfg)
coord.Start()
coord.OnAlertFired(&alerts.Alert{ID: "alert-1", ResourceID: "res-1"})
if coord.GetActiveIncidentCount() != 1 {
t.Fatal("Should accept first incident")
}
coord.OnAlertFired(&alerts.Alert{ID: "alert-2", ResourceID: "res-2"})
if coord.GetActiveIncidentCount() != 1 {
t.Error("Should ignore second incident due to cap")
}
}
func TestIncidentCoordinator_OnAnomalyDetected(t *testing.T) {
cfg := DefaultIncidentCoordinatorConfig()
coord := NewIncidentCoordinator(cfg)
recCfg := metrics.DefaultIncidentRecorderConfig()
recorder := metrics.NewIncidentRecorder(recCfg)
provider := &MockMetricsProvider{data: make(map[string]map[string]float64)}
recorder.SetMetricsProvider(provider)
recorder.Start()
defer recorder.Stop()
coord.SetRecorder(recorder)
coord.Start()
// Start anomaly recording
coord.OnAnomalyDetected("res-1", "agent", "cpu", "critical")
if coord.GetActiveIncidentCount() != 1 {
t.Errorf("Should start incident for anomaly, got %d", coord.GetActiveIncidentCount())
}
// Anomaly ID format check (internal detail, but verify implicitly via count)
}
func TestIncidentCoordinator_GetRecordingWindowID(t *testing.T) {
cfg := DefaultIncidentCoordinatorConfig()
coord := NewIncidentCoordinator(cfg)
recCfg := metrics.DefaultIncidentRecorderConfig()
recorder := metrics.NewIncidentRecorder(recCfg)
recorder.Start()
defer recorder.Stop()
coord.SetRecorder(recorder)
coord.Start()
alert := &alerts.Alert{ID: "alert-1", ResourceID: "res-1"}
coord.OnAlertFired(alert)
// With recorder, windowID should be present (start with 'iw-')
wid := coord.GetRecordingWindowID("alert-1")
if wid == "" {
t.Error("Expected valid window ID")
}
widMissing := coord.GetRecordingWindowID("missing")
if widMissing != "" {
t.Error("Expected empty window ID for missing alert")
}
}
+22
View File
@@ -204,3 +204,25 @@ func TestFindingsStore_UpdateInvestigationRecord(t *testing.T) {
t.Fatal("expected false for missing finding")
}
}
func TestPatrolInvestigationCompletionReplacesEarlyActionProjection(t *testing.T) {
store := NewFindingsStore()
store.Add(&Finding{ID: "finding-1", ResourceID: "vm-100", Title: "High CPU", DetectedAt: time.Now()})
patrol := &PatrolService{findings: store}
session := &InvestigationSession{ID: "investigation-1", FindingID: "finding-1", Summary: "Investigation in progress"}
if !patrol.RefreshFindingInvestigationRecord("finding-1", session) {
t.Fatal("early action projection was not recorded")
}
completed := time.Now()
session.Status = aicontracts.InvestigationStatusCompleted
session.CompletedAt = &completed
session.Summary = "Final diagnosis supported by the completed reads"
session.EvidenceIDs = []string{"final-evidence"}
if !patrol.storeFindingInvestigationRecord("finding-1", session, false) {
t.Fatal("completed investigation did not replace its early projection")
}
record := store.Get("finding-1").InvestigationRecord
if record.Conclusion != session.Summary || record.CompletedAt == nil || len(record.Evidence) != 1 || record.Evidence[0].ID != "final-evidence" {
t.Fatalf("completed evidence was lost: %#v", record)
}
}
+49 -23
View File
@@ -7,6 +7,7 @@ import (
"context"
"encoding/json"
"fmt"
"reflect"
"sort"
"strings"
"sync"
@@ -1791,23 +1792,9 @@ func (p *PatrolService) maybeInvestigateFinding(f *Finding) bool {
if orchestrator != nil {
latestInvestigation = orchestrator.GetInvestigationByFinding(latest.ID)
}
if record := BuildFindingInvestigationRecord(latest, latestInvestigation); record != nil {
// When a remediation plan exists for this finding, lift its
// per-step rollback strings into record.Rollback so the
// operator-facing investigation surface answers
// "what's the undo for the proposed fix?" at the record root
// rather than only in nested per-step payload.
if engine := p.remediationEngine; engine != nil {
if plan := engine.GetPlanForFinding(latest.ID); plan != nil {
record.Rollback = AggregatePlanRollbackSteps(plan)
}
}
if p.findings.UpdateInvestigationRecord(latest.ID, record) {
if refreshed := p.findings.Get(latest.ID); refreshed != nil {
latest = refreshed
} else {
latest.InvestigationRecord = record
}
if p.storeFindingInvestigationRecord(latest.ID, latestInvestigation, false) {
if refreshed := p.findings.Get(latest.ID); refreshed != nil {
latest = refreshed
}
}
if pushUnified != nil {
@@ -1850,11 +1837,50 @@ func (p *PatrolService) maybeInvestigateFinding(f *Finding) bool {
return true
}
// PublishFindingLifecycleUpdate projects a reconciled action outcome to the
// unified finding owner and, for terminal execution outcomes, to mobile push.
// It is called only after the finding store changed, so duplicate action
// callbacks and read-time hydration do not emit duplicate notifications.
func (p *PatrolService) PublishFindingLifecycleUpdate(findingID string) {
// RefreshFindingInvestigationRecord preserves the latest investigation and
// reconciled action in the durable record shared by product surfaces.
func (p *PatrolService) RefreshFindingInvestigationRecord(findingID string, session *InvestigationSession) bool {
return p.storeFindingInvestigationRecord(findingID, session, true)
}
func (p *PatrolService) storeFindingInvestigationRecord(findingID string, session *InvestigationSession, preserveEvidence bool) bool {
if p == nil || p.findings == nil {
return false
}
finding := p.findings.Get(findingID)
if finding == nil {
return false
}
record := BuildFindingInvestigationRecord(finding, session)
// Later action transitions update lifecycle facts, not the evidence and
// diagnosis captured when this investigation completed. The current finding
// may no longer retain all of that original context after restart.
if previous := finding.InvestigationRecord; preserveEvidence && previous != nil && previous.ID == record.ID {
retained := previous.NormalizeCollections()
retained.Status = record.Status
retained.Outcome = record.Outcome
retained.Action = record.Action
retained.Verification = record.Verification
record = &retained
} else {
p.mu.RLock()
engine := p.remediationEngine
p.mu.RUnlock()
if engine != nil {
if plan := engine.GetPlanForFinding(findingID); plan != nil {
record.Rollback = AggregatePlanRollbackSteps(plan)
}
}
}
if reflect.DeepEqual(finding.InvestigationRecord, record) {
return false
}
return p.findings.UpdateInvestigationRecord(findingID, record)
}
// PublishFindingLifecycleUpdate projects reconciled records to the unified
// finding owner. Repairing a stale record alone must not repeat outcome pushes.
func (p *PatrolService) PublishFindingLifecycleUpdate(findingID string, outcomeChanged bool) {
if p == nil || p.findings == nil {
return
}
@@ -1873,7 +1899,7 @@ func (p *PatrolService) PublishFindingLifecycleUpdate(findingID string) {
if finding.ResolvedAt != nil && resolveUnified != nil {
resolveUnified(finding.ID)
}
if pushNotify == nil {
if pushNotify == nil || !outcomeChanged {
return
}
switch InvestigationOutcome(finding.InvestigationOutcome) {
+18 -4
View File
@@ -534,11 +534,25 @@ func validatePatrolRoute(expected string, settings AISettings, status PatrolStat
}
func (c *PulseClient) Resources(ctx context.Context) ([]Resource, error) {
var response struct {
Data []Resource `json:"data"`
var resources []Resource
for page := 1; ; page++ {
var response struct {
Data []Resource `json:"data"`
Meta struct {
TotalPages int `json:"totalPages"`
} `json:"meta"`
}
// The resource API caps each page at 100 regardless of the requested
// limit. Follow its pagination so later resources can converge too.
path := fmt.Sprintf("/api/resources?limit=100&page=%d", page)
if err := c.request(ctx, http.MethodGet, path, nil, &response); err != nil {
return nil, err
}
resources = append(resources, response.Data...)
if page >= response.Meta.TotalPages {
return resources, nil
}
}
err := c.request(ctx, http.MethodGet, "/api/resources?limit=1000", nil, &response)
return response.Data, err
}
func (c *PulseClient) WaitForResources(ctx context.Context, names map[string]string, timeout, poll time.Duration) (map[string]Resource, error) {
+57
View File
@@ -4,6 +4,7 @@ import (
"context"
"encoding/json"
"errors"
"fmt"
"io"
"net/http"
"net/http/httptest"
@@ -13,6 +14,62 @@ import (
"time"
)
func TestWaitForResourcesMatchingIncludesLaterPages(t *testing.T) {
var pages []string
server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
if r.URL.Path != "/api/resources" || r.URL.Query().Get("limit") != "100" {
t.Errorf("unexpected resource request: %s", r.URL)
}
page := r.URL.Query().Get("page")
pages = append(pages, page)
resources := make([]Resource, 0, 100)
switch page {
case "1":
for i := 0; i < 100; i++ {
resources = append(resources, Resource{ID: fmt.Sprintf("control-%d", i), Name: fmt.Sprintf("control-%d", i)})
}
case "2":
resources = append(resources, Resource{ID: "storage", Name: "worker", Docker: &DockerResource{Health: "unhealthy"}})
default:
t.Errorf("unexpected page: %q", page)
}
_ = json.NewEncoder(w).Encode(map[string]any{"data": resources, "meta": map[string]int{"totalPages": 2}})
}))
defer server.Close()
client, err := NewPulseClient(ClientConfig{BaseURL: server.URL})
if err != nil {
t.Fatal(err)
}
resources, err := client.WaitForResourcesMatching(context.Background(), map[string]string{"service": "worker"}, time.Second, time.Millisecond, func(resources map[string]Resource) error {
if resources["service"].Docker.Health != "unhealthy" {
return errors.New("fault not collected")
}
return nil
})
if err != nil || resources["service"].ID != "storage" || strings.Join(pages, ",") != "1,2" {
t.Fatalf("resources=%v pages=%v err=%v", resources, pages, err)
}
}
func TestResourcesDoesNotReturnPartialInventoryWhenLaterPageFails(t *testing.T) {
server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
if r.URL.Query().Get("page") == "1" {
_, _ = w.Write([]byte(`{"data":[{"id":"first"}],"meta":{"totalPages":2}}`))
return
}
http.Error(w, "resource inventory unavailable", http.StatusServiceUnavailable)
}))
defer server.Close()
client, err := NewPulseClient(ClientConfig{BaseURL: server.URL})
if err != nil {
t.Fatal(err)
}
resources, err := client.Resources(context.Background())
if err == nil || resources != nil {
t.Fatalf("incomplete inventory returned: resources=%v err=%v", resources, err)
}
}
func TestTriggerAndWaitAssociatesExactNewScopedRun(t *testing.T) {
var triggered atomic.Bool
server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
+4 -5
View File
@@ -184,13 +184,12 @@ func (m ChatMessage) NormalizeCollections() ChatMessage {
return m
}
// ChatToolCall represents a provider-facing tool invocation in an API-facing
// chat message. It aliases the shared Pulse Intelligence provider-call shape so
// API chat history and provider turns do not drift on tool-call JSON.
type ChatToolCall = agentcapabilities.ProviderToolCall
// ChatToolCall retains observed output and result status in product history.
// Provider turns use the explicit ProviderToolCall projection.
type ChatToolCall = agentcapabilities.TranscriptToolCall
func EmptyChatToolCall() ChatToolCall {
return agentcapabilities.EmptyProviderToolCall()
return ChatToolCall{}.NormalizeCollections()
}
// ChatToolResult represents the result of a tool invocation. It aliases the
+1 -1
View File
@@ -168,7 +168,7 @@ func TestChatMessage_UsesCanonicalEmptyCollections(t *testing.T) {
Name: "diagnose",
ThoughtSignature: json.RawMessage(`{"provider":"gemini"}`),
}
var sharedProviderCall agentcapabilities.ProviderToolCall = sharedCall.NormalizeCollections()
var sharedProviderCall agentcapabilities.TranscriptToolCall = sharedCall.NormalizeCollections()
if sharedProviderCall.ID != "call-1" || sharedProviderCall.Input == nil {
t.Fatalf("shared chat tool call = %+v", sharedProviderCall)
}
@@ -0,0 +1,145 @@
package tools
import (
"context"
"encoding/json"
"testing"
"time"
"github.com/rcourtman/pulse-go-rewrite/internal/agentexec"
"github.com/rcourtman/pulse-go-rewrite/internal/models"
"github.com/rcourtman/pulse-go-rewrite/internal/unifiedresources"
)
func commandEvidenceSnapshot() models.StateSnapshot {
return models.StateSnapshot{
Nodes: []models.Node{{ID: "node-one", Name: "node-one", Status: "online"}},
VMs: []models.VM{{ID: "vm-one", VMID: 101, Name: "guest-one", Node: "node-one", Status: "running"}},
DockerHosts: []models.DockerHost{{
ID: "host-one", Hostname: "command-host", Status: "online", LastSeen: time.Now(),
Containers: []models.DockerContainer{{ID: "container-one", Name: "observed-service", State: "running", CPUPercent: 12.5}},
}},
}
}
func commandEvidenceJSON(t *testing.T, value any) map[string]any {
t.Helper()
raw, err := json.Marshal(value)
if err != nil {
t.Fatal(err)
}
var result map[string]any
if err := json.Unmarshal(raw, &result); err != nil {
t.Fatal(err)
}
return result
}
func TestCommandConnectivityDoesNotReplaceMonitoringEvidence(t *testing.T) {
for _, tc := range []struct {
name string
connected, guestConnected, controlEnabled bool
}{
{name: "no_command_connection"},
{name: "connected_read_only", connected: true},
{name: "guest_connection_only", guestConnected: true},
{name: "connected_control_enabled", connected: true, controlEnabled: true},
} {
t.Run(tc.name, func(t *testing.T) {
server := &mockAgentServer{}
if tc.connected {
server.agents = []agentexec.ConnectedAgent{{Hostname: "command-host"}, {Hostname: "node-one"}}
}
if tc.guestConnected {
server.agents = append(server.agents, agentexec.ConnectedAgent{Hostname: "guest-one"})
}
controlLevel := ControlLevelReadOnly
if tc.controlEnabled {
controlLevel = ControlLevelControlled
}
registry := unifiedresources.NewRegistry(nil)
registry.IngestSnapshot(commandEvidenceSnapshot())
executor := NewPulseToolExecutor(ExecutorConfig{ReadState: registry, UnifiedResourceProvider: &registryUnifiedQueryProvider{registry}, AgentServer: server, ControlLevel: controlLevel})
query := func(args map[string]interface{}) map[string]any {
t.Helper()
result, err := executor.executeQuery(context.Background(), args)
if err != nil || result.IsError {
t.Fatalf("query failed: %v %+v", err, result)
}
var decoded map[string]any
if err := json.Unmarshal([]byte(result.Content[0].Text), &decoded); err != nil {
t.Fatal(err)
}
capture, _ := json.Marshal(map[string]any{"case": tc.name, "input": args, "output": decoded})
t.Logf("COMMAND_EVIDENCE %s", capture)
return decoded
}
list := query(map[string]interface{}{"action": "list", "type": "docker-hosts"})
host := list["docker_hosts"].([]any)[0].(map[string]any)
if host["command_agent_connected"] != tc.connected {
t.Fatalf("connection must name command transport: %+v", host)
}
if _, exists := host["agent_connected"]; exists {
t.Fatal("ambiguous connection field remains")
}
container := host["containers"].([]any)[0].(map[string]any)
resource := query(map[string]interface{}{"action": "get", "resource_type": "app-container", "resource_id": container["id"]})
if resource["status"] != "running" || resource["cpu"].(map[string]any)["percent"] != 12.5 {
t.Fatalf("command state replaced monitored evidence: %+v", resource)
}
topology := query(map[string]interface{}{"action": "topology", "include": "all"})
docker := topology["docker"].(map[string]any)["hosts"].([]any)[0].(map[string]any)
if docker["command_agent_connected"] != tc.connected || docker["can_execute"] != (tc.connected && tc.controlEnabled) {
t.Fatalf("transport/control hint changed: %+v", docker)
}
search := query(map[string]interface{}{"action": "search", "query": "guest-one"})
guest := search["matches"].([]any)[0].(map[string]any)
if guest["node_command_agent_connected"] != tc.connected {
t.Fatalf("parent transport not identified: %+v", guest)
}
if guest["command_agent_connected"] != tc.guestConnected {
t.Fatalf("parent connection became a direct guest connection: %+v", guest)
}
server.AssertNotCalled(t, "ExecuteCommand")
})
}
}
func TestTopologyOmitsUnobservedCommandConnections(t *testing.T) {
registry := unifiedresources.NewRegistry(nil)
registry.IngestSnapshot(commandEvidenceSnapshot())
for _, observed := range []bool{false, true} {
options := TopologyBuildOptions{Include: "all", ControlEnabled: true}
if observed {
options.ConnectedAgentHostnames = map[string]bool{}
}
result := commandEvidenceJSON(t, BuildTopologyResponseFromReadState(registry, options))
caseName := "unobserved_topology"
if observed {
caseName = "observed_empty_topology"
}
capture, err := json.Marshal(map[string]any{"case": caseName, "input": map[string]any{"action": "topology", "include": "all"}, "output": result})
if err != nil {
t.Fatal(err)
}
t.Logf("COMMAND_EVIDENCE %s", capture)
for _, group := range []struct{ family, collection string }{{"docker", "hosts"}, {"proxmox", "nodes"}} {
item := result[group.family].(map[string]any)[group.collection].([]any)[0].(map[string]any)
for _, field := range []string{"command_agent_connected", "can_execute"} {
value, exists := item[field]
if exists != observed || (exists && value != false) {
t.Fatalf("observed=%t field=%s: %+v", observed, field, item)
}
}
if _, exists := item["agent_connected"]; exists {
t.Fatal("ambiguous connection field remains")
}
}
for _, field := range []string{"nodes_with_command_agents", "docker_hosts_with_command_agents"} {
value, exists := result["summary"].(map[string]any)[field]
if exists != observed || (exists && value != float64(0)) {
t.Fatalf("unobserved connections became a count: %+v", result["summary"])
}
}
}
}
+65 -63
View File
@@ -238,40 +238,40 @@ func (r ResourceSearchResponse) NormalizeCollections() ResourceSearchResponse {
// ResourceMatch is a compact match result for pulse_search_resources
type ResourceMatch struct {
GovernedResourceMetadata
Type string `json:"type"` // "agent", "node", "vm", "system-container", "app-container", "docker-host", "storage"
ID string `json:"id,omitempty"`
Name string `json:"name"`
Status string `json:"status,omitempty"`
Node string `json:"node,omitempty"` // Hypervisor node this resource is on
NodeHasAgent bool `json:"node_has_agent,omitempty"` // True if the node has a connected agent
Host string `json:"host,omitempty"` // Docker host for docker containers
Platform string `json:"platform,omitempty"`
VMID int `json:"vmid,omitempty"`
Image string `json:"image,omitempty"`
AgentConnected bool `json:"agent_connected,omitempty"` // True if this specific resource has a connected agent
Type string `json:"type"` // "agent", "node", "vm", "system-container", "app-container", "docker-host", "storage"
ID string `json:"id,omitempty"`
Name string `json:"name"`
Status string `json:"status,omitempty"`
Node string `json:"node,omitempty"` // Hypervisor node this resource is on
NodeCommandAgentConnected *bool `json:"node_command_agent_connected,omitempty"` // Live command connection on the parent node, independent of telemetry collection
Host string `json:"host,omitempty"` // Docker host for docker containers
Platform string `json:"platform,omitempty"`
VMID int `json:"vmid,omitempty"`
Image string `json:"image,omitempty"`
CommandAgentConnected *bool `json:"command_agent_connected,omitempty"` // Live command connection for this resource, independent of telemetry collection
}
// SystemSummary is a summarized infrastructure system for list responses.
type SystemSummary struct {
GovernedResourceMetadata
ID string `json:"id"`
Name string `json:"name"`
Status string `json:"status"`
Platform string `json:"platform,omitempty"`
ChildCount int `json:"child_count,omitempty"`
AgentConnected bool `json:"agent_connected,omitempty"`
CPU float64 `json:"cpu_percent,omitempty"`
Memory float64 `json:"memory_percent,omitempty"`
Disk float64 `json:"disk_percent,omitempty"`
ID string `json:"id"`
Name string `json:"name"`
Status string `json:"status"`
Platform string `json:"platform,omitempty"`
ChildCount int `json:"child_count,omitempty"`
CommandAgentConnected *bool `json:"command_agent_connected,omitempty"`
CPU float64 `json:"cpu_percent,omitempty"`
Memory float64 `json:"memory_percent,omitempty"`
Disk float64 `json:"disk_percent,omitempty"`
}
// NodeSummary is a summarized node for list responses
type NodeSummary struct {
GovernedResourceMetadata
Name string `json:"name"`
Status string `json:"status"`
ID string `json:"id,omitempty"`
AgentConnected bool `json:"agent_connected"` // True if an execution agent is connected for this node
Name string `json:"name"`
Status string `json:"status"`
ID string `json:"id,omitempty"`
CommandAgentConnected *bool `json:"command_agent_connected,omitempty"` // True if an execution agent is connected for this node
}
// VMSummary is a summarized VM for list responses
@@ -299,12 +299,12 @@ type ContainerSummary struct {
// DockerHostSummary is a summarized Docker host for list responses
type DockerHostSummary struct {
GovernedResourceMetadata
ID string `json:"id"`
Hostname string `json:"hostname"`
DisplayName string `json:"display_name,omitempty"`
ContainerCount int `json:"container_count"`
AgentConnected bool `json:"agent_connected"` // True if an execution agent is connected for this host
Containers []DockerContainerSummary `json:"containers"`
ID string `json:"id"`
Hostname string `json:"hostname"`
DisplayName string `json:"display_name,omitempty"`
ContainerCount int `json:"container_count"`
CommandAgentConnected *bool `json:"command_agent_connected,omitempty"` // True if an execution agent is connected for this host
Containers []DockerContainerSummary `json:"containers"`
}
func (s DockerHostSummary) NormalizeCollections() DockerHostSummary {
@@ -454,15 +454,15 @@ func (t ProxmoxTopology) NormalizeCollections() ProxmoxTopology {
// ProxmoxNodeTopology represents a Proxmox node with its guests
type ProxmoxNodeTopology struct {
GovernedResourceMetadata
Name string `json:"name"`
ID string `json:"id,omitempty"`
Status string `json:"status"`
AgentConnected bool `json:"agent_connected"`
CanExecute bool `json:"can_execute"` // True if commands can be executed on this node
VMs []TopologyVM `json:"vms"`
Containers []TopologyContainer `json:"containers"`
VMCount int `json:"vm_count"`
ContainerCount int `json:"container_count"`
Name string `json:"name"`
ID string `json:"id,omitempty"`
Status string `json:"status"`
CommandAgentConnected *bool `json:"command_agent_connected,omitempty"`
CanExecute *bool `json:"can_execute,omitempty"` // True if commands can be executed on this node
VMs []TopologyVM `json:"vms"`
Containers []TopologyContainer `json:"containers"`
VMCount int `json:"vm_count"`
ContainerCount int `json:"container_count"`
}
func (t ProxmoxNodeTopology) NormalizeCollections() ProxmoxNodeTopology {
@@ -538,15 +538,15 @@ func (t DockerTopology) NormalizeCollections() DockerTopology {
// DockerHostTopology represents a Docker host with its containers
type DockerHostTopology struct {
GovernedResourceMetadata
Hostname string `json:"hostname"`
DisplayName string `json:"display_name,omitempty"`
AgentConnected bool `json:"agent_connected"`
CanExecute bool `json:"can_execute"` // True if commands can be executed on this host
Containers []DockerContainerSummary `json:"containers"`
ContainerCount int `json:"container_count"`
ReturnedCount int `json:"returned_container_count"`
Truncated bool `json:"containers_truncated"`
RunningCount int `json:"running_count"`
Hostname string `json:"hostname"`
DisplayName string `json:"display_name,omitempty"`
CommandAgentConnected *bool `json:"command_agent_connected,omitempty"`
CanExecute *bool `json:"can_execute,omitempty"` // True if commands can be executed on this host
Containers []DockerContainerSummary `json:"containers"`
ContainerCount int `json:"container_count"`
ReturnedCount int `json:"returned_container_count"`
Truncated bool `json:"containers_truncated"`
RunningCount int `json:"running_count"`
}
func (t DockerHostTopology) NormalizeCollections() DockerHostTopology {
@@ -642,21 +642,21 @@ type KubernetesPodDetail struct {
// TopologySummary provides aggregate counts and status
type TopologySummary struct {
TotalNodes int `json:"total_nodes"`
TotalVMs int `json:"total_vms"`
TotalSystemContainers int `json:"total_system_containers"`
TotalDockerHosts int `json:"total_docker_hosts"`
TotalDockerContainers int `json:"total_docker_containers"`
TotalK8sClusters int `json:"total_k8s_clusters"`
TotalK8sNodes int `json:"total_k8s_nodes"`
TotalK8sDeployments int `json:"total_k8s_deployments"`
TotalK8sPods int `json:"total_k8s_pods"`
NodesWithAgents int `json:"nodes_with_agents"`
DockerHostsWithAgents int `json:"docker_hosts_with_agents"`
RunningVMs int `json:"running_vms"`
RunningContainers int `json:"running_containers"`
RunningDocker int `json:"running_docker"`
RunningK8sPods int `json:"running_k8s_pods"`
TotalNodes int `json:"total_nodes"`
TotalVMs int `json:"total_vms"`
TotalSystemContainers int `json:"total_system_containers"`
TotalDockerHosts int `json:"total_docker_hosts"`
TotalDockerContainers int `json:"total_docker_containers"`
TotalK8sClusters int `json:"total_k8s_clusters"`
TotalK8sNodes int `json:"total_k8s_nodes"`
TotalK8sDeployments int `json:"total_k8s_deployments"`
TotalK8sPods int `json:"total_k8s_pods"`
NodesWithCommandAgents *int `json:"nodes_with_command_agents,omitempty"`
DockerHostsWithCommandAgents *int `json:"docker_hosts_with_command_agents,omitempty"`
RunningVMs int `json:"running_vms"`
RunningContainers int `json:"running_containers"`
RunningDocker int `json:"running_docker"`
RunningK8sPods int `json:"running_k8s_pods"`
}
// ResourceResponse is returned by pulse_get_resource
@@ -883,8 +883,10 @@ type PortInfo struct {
// MountInfo describes a volume mount
type MountInfo struct {
Type string `json:"type,omitempty"`
Source string `json:"source"`
Destination string `json:"destination"`
Mode string `json:"mode,omitempty"`
ReadWrite bool `json:"rw"`
}
+12 -12
View File
@@ -559,9 +559,9 @@ type ExecutorConfig struct {
AgentProfileManager AgentProfileManager
// Optional providers - intelligence
IncidentRecorderProvider IncidentRecorderProvider
EventCorrelatorProvider EventCorrelatorProvider
KnowledgeStoreProvider KnowledgeStoreProvider
IncidentArchiveProvider IncidentArchiveProvider
EventCorrelatorProvider EventCorrelatorProvider
KnowledgeStoreProvider KnowledgeStoreProvider
// Optional providers - discovery
DiscoveryProvider DiscoveryProvider
@@ -627,9 +627,9 @@ type PulseToolExecutor struct {
agentProfileManager AgentProfileManager
// Intelligence providers
incidentRecorderProvider IncidentRecorderProvider
eventCorrelatorProvider EventCorrelatorProvider
knowledgeStoreProvider KnowledgeStoreProvider
incidentArchiveProvider IncidentArchiveProvider
eventCorrelatorProvider EventCorrelatorProvider
knowledgeStoreProvider KnowledgeStoreProvider
// Discovery provider
discoveryProvider DiscoveryProvider
@@ -750,7 +750,7 @@ func NewPulseToolExecutor(cfg ExecutorConfig) *PulseToolExecutor {
metadataUpdater: cfg.MetadataUpdater,
findingsManager: cfg.FindingsManager,
agentProfileManager: cfg.AgentProfileManager,
incidentRecorderProvider: cfg.IncidentRecorderProvider,
incidentArchiveProvider: cfg.IncidentArchiveProvider,
eventCorrelatorProvider: cfg.EventCorrelatorProvider,
knowledgeStoreProvider: cfg.KnowledgeStoreProvider,
discoveryProvider: cfg.DiscoveryProvider,
@@ -827,7 +827,7 @@ func (e *PulseToolExecutor) Clone() *PulseToolExecutor {
metadataUpdater: e.metadataUpdater,
findingsManager: e.findingsManager,
agentProfileManager: e.agentProfileManager,
incidentRecorderProvider: e.incidentRecorderProvider,
incidentArchiveProvider: e.incidentArchiveProvider,
eventCorrelatorProvider: e.eventCorrelatorProvider,
knowledgeStoreProvider: e.knowledgeStoreProvider,
discoveryProvider: e.discoveryProvider,
@@ -1019,9 +1019,9 @@ func (e *PulseToolExecutor) SetUpdatesProvider(provider UpdatesProvider) {
e.updatesProvider = provider
}
// SetIncidentRecorderProvider sets the incident recorder provider
func (e *PulseToolExecutor) SetIncidentRecorderProvider(provider IncidentRecorderProvider) {
e.incidentRecorderProvider = provider
// SetIncidentArchiveProvider sets the read-only legacy incident archive provider
func (e *PulseToolExecutor) SetIncidentArchiveProvider(provider IncidentArchiveProvider) {
e.incidentArchiveProvider = provider
}
// SetEventCorrelatorProvider sets the event correlator provider
@@ -1221,7 +1221,7 @@ func (e *PulseToolExecutor) isToolAvailable(name string) bool {
case agentcapabilities.PulseDiscoveryToolName:
return e.discoveryProvider != nil
case agentcapabilities.PulseKnowledgeToolName:
return e.knowledgeStoreProvider != nil || e.incidentRecorderProvider != nil || e.eventCorrelatorProvider != nil
return e.actionAuditStore != nil || e.knowledgeStoreProvider != nil || e.incidentArchiveProvider != nil || e.eventCorrelatorProvider != nil
case agentcapabilities.PulsePMGToolName:
return e.hasReadState()
case agentcapabilities.PulseSummarizeToolName:
+221
View File
@@ -0,0 +1,221 @@
package tools
import (
"context"
"encoding/json"
"errors"
"os"
"path/filepath"
"strings"
"testing"
"time"
"github.com/rcourtman/pulse-go-rewrite/internal/agentcapabilities"
"github.com/rcourtman/pulse-go-rewrite/internal/metrics"
"github.com/rcourtman/pulse-go-rewrite/internal/unifiedresources"
"github.com/stretchr/testify/require"
)
type failedIncidentHistoryStore struct{ unifiedresources.ResourceStore }
func TestIncidentHistoryRetainsLegacyDockerLifecycle(t *testing.T) {
dir := t.TempDir()
store, err := unifiedresources.NewSQLiteResourceStore(dir, "incident-test")
require.NoError(t, err)
t.Cleanup(func() { require.NoError(t, store.Close()) })
container := strings.Repeat("b", 64)
legacy := "docker:tower/" + container
canonical := unifiedresources.SourceSpecificID(unifiedresources.ResourceTypeAppContainer, unifiedresources.SourceDocker, "tower/container/"+container)
start := time.Now().UTC().Add(-time.Hour).Truncate(time.Second)
occurred := start.Add(-time.Minute)
fired := unifiedresources.ResourceChange{ID: "legacy-fired", ResourceID: legacy, ObservedAt: start.Add(time.Minute), OccurredAt: &occurred, Kind: unifiedresources.ChangeAlertFired, SourceType: unifiedresources.SourceHeuristic, Reason: "Container unhealthy"}
require.NoError(t, store.RecordChange(fired))
require.NoError(t, store.Close())
store, err = unifiedresources.NewSQLiteResourceStore(dir, "incident-test")
require.NoError(t, err)
resolved := unifiedresources.ResourceChange{ID: "canonical-resolved", ResourceID: canonical, ObservedAt: start.Add(3 * time.Minute), Kind: unifiedresources.ChangeAlertResolved, SourceType: unifiedresources.SourceHeuristic}
require.NoError(t, store.RecordChange(resolved))
exec := NewPulseToolExecutor(ExecutorConfig{ActionAuditStore: store})
input := map[string]interface{}{"action": "incidents", "resource_id": canonical, "since": start.Format(time.RFC3339), "limit": float64(50)}
result, err := exec.registry.Execute(context.Background(), exec, agentcapabilities.PulseKnowledgeToolName, input)
require.NoError(t, err)
require.False(t, result.IsError, result.Content)
var got struct {
Events []unifiedresources.ResourceChange `json:"events"`
}
require.NoError(t, json.Unmarshal([]byte(result.Content[0].Text), &got))
require.Equal(t, []unifiedresources.ResourceChange{resolved, fired}, got.Events)
capture, err := json.Marshal(map[string]any{"case": "migrated Docker lifecycle", "input": input, "result": result})
require.NoError(t, err)
t.Logf("INCIDENT_EVIDENCE %s", capture)
}
func (failedIncidentHistoryStore) GetRecentChanges(string, time.Time, int) ([]unifiedresources.ResourceChange, error) {
return nil, errors.New("history store unavailable")
}
type incidentArchiveFixture struct{ window *metrics.IncidentWindow }
func (s incidentArchiveFixture) GetWindow(string, string) (*metrics.IncidentWindow, error) {
return s.window, nil
}
func TestIncidentHistoryRetainsCanonicalEvidence(t *testing.T) {
store, err := unifiedresources.NewSQLiteResourceStore(t.TempDir(), "incident-test")
require.NoError(t, err)
t.Cleanup(func() { require.NoError(t, store.Close()) })
resourceID := "app-container-7020f37498275208"
start := time.Now().UTC().Add(-time.Hour).Truncate(time.Second)
occurred := start.Add(-time.Minute)
changes := []unifiedresources.ResourceChange{
{ID: "outside-window", ResourceID: resourceID, ObservedAt: start.Add(-time.Second), Kind: unifiedresources.ChangeRestart, SourceType: unifiedresources.SourcePlatformEvent},
{ID: "alert-fired", ResourceID: resourceID, ObservedAt: start.Add(time.Minute), OccurredAt: &occurred, Kind: unifiedresources.ChangeAlertFired, SourceType: unifiedresources.SourcePulseDiff, From: "healthy", To: "unhealthy", Reason: "Health check failed", Metadata: map[string]any{"alert_id": "health-check", "value": 1.0}},
{ID: "related-only", ResourceID: "agent-parent", RelatedResources: []string{resourceID}, ObservedAt: start.Add(2 * time.Minute), Kind: unifiedresources.ChangeRestart, SourceType: unifiedresources.SourcePlatformEvent},
{ID: "alert-resolved", ResourceID: resourceID, ObservedAt: start.Add(3 * time.Minute), Kind: unifiedresources.ChangeAlertResolved, SourceType: unifiedresources.SourcePulseDiff, From: "unhealthy", To: "healthy"},
}
for _, change := range changes {
require.NoError(t, store.RecordChange(change))
}
// No live inventory or incident recorder is necessary to read a retained
// event for a container that has since been removed.
exec := NewPulseToolExecutor(ExecutorConfig{ActionAuditStore: store, IncidentArchiveProvider: incidentArchiveFixture{}})
require.True(t, exec.isToolAvailable(agentcapabilities.PulseKnowledgeToolName))
for _, tc := range []struct {
name string
id string
limit int
wantCount int
wantMore bool
}{
{"retained lifecycle", resourceID, 50, 2, false},
{"bounded lifecycle", resourceID, 1, 1, true},
{"empty history", "app-container-absent", 50, 0, false},
} {
t.Run(tc.name, func(t *testing.T) {
input := map[string]interface{}{"action": "incidents", "resource_id": tc.id, "since": start.Format(time.RFC3339), "limit": float64(tc.limit)}
result, err := exec.registry.Execute(context.Background(), exec, agentcapabilities.PulseKnowledgeToolName, input)
require.NoError(t, err)
require.False(t, result.IsError, result.Content)
var got struct {
Source string `json:"source"`
Events []unifiedresources.ResourceChange `json:"events"`
HasMore bool `json:"has_more"`
Coverage string `json:"coverage"`
TimeBasis string `json:"time_basis"`
}
require.NoError(t, json.Unmarshal([]byte(result.Content[0].Text), &got))
require.Equal(t, "canonical_resource_timeline", got.Source)
require.Equal(t, "retained_records_only", got.Coverage)
require.Equal(t, "observed_at", got.TimeBasis)
require.NotNil(t, got.Events)
require.Len(t, got.Events, tc.wantCount)
require.Equal(t, tc.wantMore, got.HasMore)
if tc.wantCount > 0 {
require.Equal(t, "alert-resolved", got.Events[0].ID)
require.Nil(t, got.Events[0].OccurredAt)
}
if tc.wantCount == 2 {
require.Equal(t, changes[1], got.Events[1])
}
capture, err := json.Marshal(map[string]any{"case": tc.name, "input": input, "result": result})
require.NoError(t, err)
t.Logf("INCIDENT_EVIDENCE %s", capture)
})
}
}
func TestIncidentHistoryUnavailableAndInvalid(t *testing.T) {
for _, tc := range []struct {
name string
store unifiedresources.ResourceStore
input map[string]interface{}
}{
{"unavailable", nil, map[string]interface{}{"resource_id": "app-container-1"}},
{"failed", failedIncidentHistoryStore{}, map[string]interface{}{"resource_id": "app-container-1"}},
{"empty resource", unifiedresources.NewMemoryStore(), map[string]interface{}{"resource_id": " "}},
{"invalid time", unifiedresources.NewMemoryStore(), map[string]interface{}{"resource_id": "app-container-1", "since": "yesterday"}},
{"future time", unifiedresources.NewMemoryStore(), map[string]interface{}{"resource_id": "app-container-1", "since": time.Now().Add(time.Hour).Format(time.RFC3339)}},
{"negative limit", unifiedresources.NewMemoryStore(), map[string]interface{}{"resource_id": "app-container-1", "limit": float64(-1)}},
{"large limit", unifiedresources.NewMemoryStore(), map[string]interface{}{"resource_id": "app-container-1", "limit": float64(201)}},
} {
t.Run(tc.name, func(t *testing.T) {
exec := NewPulseToolExecutor(ExecutorConfig{ActionAuditStore: tc.store, IncidentArchiveProvider: incidentArchiveFixture{}})
tc.input["action"] = "incidents"
result, err := exec.registry.Execute(context.Background(), exec, agentcapabilities.PulseKnowledgeToolName, tc.input)
require.NoError(t, err)
require.True(t, result.IsError, result.Content)
capture, err := json.Marshal(map[string]any{"case": tc.name, "input": tc.input, "result": result})
require.NoError(t, err)
t.Logf("INCIDENT_EVIDENCE %s", capture)
})
}
}
func TestIncidentHistoryLegacyArchiveIsResourceBound(t *testing.T) {
window := &metrics.IncidentWindow{ID: "archive-1", ResourceID: "vm-1"}
exec := NewPulseToolExecutor(ExecutorConfig{IncidentArchiveProvider: incidentArchiveFixture{window}})
for _, id := range []string{"vm-1", "vm-2"} {
result, err := exec.executeGetIncidentWindow(context.Background(), map[string]interface{}{"resource_id": id, "window_id": window.ID})
require.NoError(t, err)
if id == window.ResourceID {
require.False(t, result.IsError)
require.Contains(t, result.Content[0].Text, "legacy_incident_recording")
require.Contains(t, result.Content[0].Text, "cached observations")
} else {
require.True(t, result.IsError)
require.NotContains(t, result.Content[0].Text, "vm-1")
}
}
result, err := exec.executeGetIncidentWindow(context.Background(), map[string]interface{}{"resource_id": window.ResourceID, "window_id": "wrong-window"})
require.NoError(t, err)
require.True(t, result.IsError, "a provider cannot substitute another archived window")
}
func TestIncidentHistoryArchiveReadOutcomes(t *testing.T) {
dir := t.TempDir()
archivePath := filepath.Join(dir, "incident_windows.json")
archive := metrics.NewIncidentArchive(dir)
exec := NewPulseToolExecutor(ExecutorConfig{IncidentArchiveProvider: archive})
original := `{"completed_windows":[{"id":"saved-window","resource_id":"vm-archive","status":"recording","start_time":"2020-01-02T03:04:05Z","data_points":[{"timestamp":"2020-01-02T03:04:06Z","metrics":{"cpu":12.5},"metadata":{"source":"cached"}}],"summary":{"duration_ms":60000000000,"anomalies":["stored observation"]}}]}`
for _, tc := range []struct {
name, raw, resource, window string
wantError bool
}{
{"archive unavailable", "", "vm-archive", "saved-window", true},
{"archive malformed", "{", "vm-archive", "saved-window", true},
{"archive success", original, "vm-archive", "saved-window", false},
{"archive wrong resource", original, "vm-other", "saved-window", true},
{"archive missing window", original, "vm-archive", "missing", true},
} {
t.Run(tc.name, func(t *testing.T) {
if tc.raw != "" {
require.NoError(t, os.WriteFile(archivePath, []byte(tc.raw), 0600))
}
input := map[string]interface{}{"action": "incidents", "resource_id": tc.resource, "window_id": tc.window}
result, err := exec.registry.Execute(context.Background(), exec, agentcapabilities.PulseKnowledgeToolName, input)
require.NoError(t, err)
require.Equal(t, tc.wantError, result.IsError, result.Content)
if !tc.wantError {
var got struct {
Window *metrics.IncidentWindow `json:"window"`
ReadOnly bool `json:"archive_read_only"`
DurationUnit string `json:"summary_duration_unit"`
}
require.NoError(t, json.Unmarshal([]byte(result.Content[0].Text), &got))
require.True(t, got.ReadOnly)
require.Equal(t, "nanoseconds", got.DurationUnit)
require.Equal(t, time.Minute, got.Window.Summary.Duration)
require.Equal(t, metrics.IncidentWindowStatusRecording, got.Window.Status)
require.Equal(t, "cached", got.Window.DataPoints[0].Metadata["source"])
require.Equal(t, []string{"stored observation"}, got.Window.Summary.Anomalies)
require.Contains(t, result.Content[0].Text, "does not mean recording is active")
} else if tc.name == "archive wrong resource" {
require.NotContains(t, result.Content[0].Text, "stored observation")
}
capture, err := json.Marshal(map[string]any{"case": tc.name, "input": input, "result": result})
require.NoError(t, err)
t.Logf("INCIDENT_EVIDENCE %s", capture)
})
}
}
@@ -0,0 +1,87 @@
package tools
import (
"context"
"encoding/json"
"testing"
"github.com/rcourtman/pulse-go-rewrite/internal/models"
"github.com/rcourtman/pulse-go-rewrite/internal/unifiedresources"
)
func TestQueryPreservesMountConfigurationEvidence(t *testing.T) {
for _, provider := range []bool{false, true} {
name := "typed read state"
if provider {
name = "canonical provider"
}
t.Run(name, func(t *testing.T) {
snapshot := commandEvidenceSnapshot()
snapshot.DockerHosts[0].Containers[0].Mounts = []models.DockerContainerMount{
{Type: "tmpfs", Destination: "/var/lib/service-cache", Mode: "rw,noexec,nosuid,nodev,size=8388608", RW: true},
{Type: "tmpfs", Destination: "/readonly-cache", Mode: "ro,noexec", RW: false},
{Type: "bind", Source: "/host/data", Destination: "/data", Mode: "", RW: false},
}
registry := unifiedresources.NewRegistry(nil)
registry.IngestSnapshot(snapshot)
cfg := ExecutorConfig{ReadState: registry, ControlLevel: ControlLevelReadOnly}
if provider {
cfg.UnifiedResourceProvider = &registryUnifiedQueryProvider{registry}
}
executor := NewPulseToolExecutor(cfg)
list, err := executor.executeQuery(context.Background(), map[string]interface{}{"action": "list", "type": "docker-hosts"})
if err != nil || list.IsError {
t.Fatalf("list: %v %+v", err, list)
}
var hosts map[string]any
if err := json.Unmarshal([]byte(list.Content[0].Text), &hosts); err != nil {
t.Fatal(err)
}
id := hosts["docker_hosts"].([]any)[0].(map[string]any)["containers"].([]any)[0].(map[string]any)["id"]
if !provider {
// The typed compatibility path currently accepts provider IDs/names.
id = snapshot.DockerHosts[0].Containers[0].Name
}
args := map[string]interface{}{"action": "get", "resource_type": "app-container", "resource_id": id}
result, err := executor.executeQuery(context.Background(), args)
if err != nil || result.IsError {
t.Fatalf("get: %v %+v", err, result)
}
var decoded map[string]any
if err := json.Unmarshal([]byte(result.Content[0].Text), &decoded); err != nil {
t.Fatal(err)
}
mounts, ok := decoded["mounts"].([]any)
if !ok {
t.Fatalf("missing mount projection: %+v", decoded)
}
if len(mounts) != 3 {
t.Fatalf("lost mounts: %+v", mounts)
}
for i, want := range snapshot.DockerHosts[0].Containers[0].Mounts {
got := mounts[i].(map[string]any)
if got["type"] != want.Type || got["source"] != want.Source || got["destination"] != want.Destination || got["rw"] != want.RW || (want.Mode != "" && got["mode"] != want.Mode) {
t.Fatalf("mount provenance/access changed: %+v, want %+v", got, want)
}
}
if _, exists := decoded["disk"]; exists {
t.Fatalf("mount configuration invented capacity: %+v", decoded["disk"])
}
capture, _ := json.Marshal(map[string]any{"case": name, "input": args, "output": decoded})
t.Logf("MOUNT_EVIDENCE %s", capture)
if provider {
resources := registry.ListByType(unifiedresources.ResourceTypeAppContainer)
for _, host := range registry.ListByType(unifiedresources.ResourceTypeAgent) {
if host.Docker != nil {
resources = append(resources, host)
}
}
encoded, err := json.Marshal(resources)
if err != nil {
t.Fatal(err)
}
t.Logf("MOUNT_RESOURCES %s", encoded)
}
})
}
}
+76 -56
View File
@@ -3,46 +3,18 @@ package tools
import (
"context"
"fmt"
"strings"
"time"
"github.com/rcourtman/pulse-go-rewrite/internal/agentcapabilities"
"github.com/rcourtman/pulse-go-rewrite/internal/metrics"
"github.com/rcourtman/pulse-go-rewrite/internal/unifiedresources"
)
// IncidentRecorderProvider provides access to incident recording data
type IncidentRecorderProvider interface {
GetWindowsForResource(resourceID string, limit int) []*IncidentWindow
GetWindow(windowID string) *IncidentWindow
}
// IncidentWindow represents a high-frequency recording window during an incident
type IncidentWindow struct {
ID string `json:"id"`
ResourceID string `json:"resource_id"`
ResourceName string `json:"resource_name,omitempty"`
ResourceType string `json:"resource_type,omitempty"`
TriggerType string `json:"trigger_type"`
TriggerID string `json:"trigger_id,omitempty"`
StartTime time.Time `json:"start_time"`
EndTime *time.Time `json:"end_time,omitempty"`
Status string `json:"status"`
DataPoints []IncidentDataPoint `json:"data_points"`
Summary *IncidentSummary `json:"summary,omitempty"`
}
// IncidentDataPoint represents a single data point in an incident window
type IncidentDataPoint struct {
Timestamp time.Time `json:"timestamp"`
Metrics map[string]float64 `json:"metrics"`
}
// IncidentSummary provides computed statistics about an incident window
type IncidentSummary struct {
Duration time.Duration `json:"duration_ms"`
DataPoints int `json:"data_points"`
Peaks map[string]float64 `json:"peaks"`
Lows map[string]float64 `json:"lows"`
Averages map[string]float64 `json:"averages"`
Changes map[string]float64 `json:"changes"`
// IncidentArchiveProvider provides explicit, resource-bound reads of saved
// legacy recordings. Live incident evidence comes from the canonical timeline.
type IncidentArchiveProvider interface {
GetWindow(resourceID, windowID string) (*metrics.IncidentWindow, error)
}
// EventCorrelatorProvider provides access to correlated events
@@ -87,7 +59,7 @@ func (e *PulseToolExecutor) registerKnowledgeTools() {
Actions:
- remember: Save a note about a resource for future reference
- recall: Retrieve saved notes about a resource
- incidents: Get high-resolution incident recording data
- incidents: Read retained canonical resource history, including observed state changes, alerts and executed actions. Records preserve observation time, source and any known occurrence time. This is not continuous health or filesystem-capacity coverage. Use pulse_summarize for retained metrics.
- correlate: Get correlated events around a timestamp
Examples:
@@ -105,7 +77,7 @@ Examples:
},
"resource_id": {
Type: "string",
Description: "Resource ID to operate on",
Description: "Resource ID to operate on. For incidents use the canonical resource ID returned by pulse_query, including for a resource no longer in current inventory.",
},
"note": {
Type: "string",
@@ -117,7 +89,11 @@ Examples:
},
"window_id": {
Type: "string",
Description: "For incidents: specific incident window ID",
Description: "For incidents: optional legacy recording ID, read as an archive only. Omit to read canonical resource history.",
},
"since": {
Type: "string",
Description: "For incidents: earliest observation timestamp (RFC3339, default 24 hours ago). Retention and collection gaps still apply.",
},
"timestamp": {
Type: "string",
@@ -129,7 +105,7 @@ Examples:
},
"limit": {
Type: "integer",
Description: "For incidents: max windows to return (default: 5)",
Description: "For incidents: maximum retained events to return, newest first (default 50, range 1-200)",
},
},
Required: []string{"action", "resource_id"},
@@ -168,38 +144,82 @@ func (e *PulseToolExecutor) executeKnowledge(ctx context.Context, args map[strin
func (e *PulseToolExecutor) executeGetIncidentWindow(_ context.Context, args map[string]interface{}) (CallToolResult, error) {
resourceID, _ := args["resource_id"].(string)
resourceID = strings.TrimSpace(resourceID)
windowID, _ := args["window_id"].(string)
limit := intArg(args, "limit", 5)
if resourceID == "" {
return NewErrorResult(fmt.Errorf("resource_id is required")), nil
}
if e.incidentRecorderProvider == nil {
return NewTextResult("Incident recording data not available. The incident recorder may not be enabled."), nil
}
// If a specific window ID is requested
// Isolate legacy recordings from the canonical timeline. Their sample times
// are recorder timestamps, not verified source observation timestamps.
if windowID != "" {
window := e.incidentRecorderProvider.GetWindow(windowID)
if window == nil {
return NewTextResult(fmt.Sprintf("Incident window '%s' not found.", windowID)), nil
if e.incidentArchiveProvider == nil {
return NewErrorResult(fmt.Errorf("legacy incident recording archive is unavailable")), nil
}
window, err := e.incidentArchiveProvider.GetWindow(resourceID, windowID)
if err != nil {
return NewErrorResult(fmt.Errorf("read legacy incident recording archive: %w", err)), nil
}
if window == nil || window.ResourceID != resourceID || window.ID != windowID {
return NewErrorResult(fmt.Errorf("legacy incident recording not found for the requested resource")), nil
}
return NewJSONResult(map[string]interface{}{
"window": window,
"source": "legacy_incident_recording",
"archive_read_only": true,
"summary_duration_unit": "nanoseconds",
"window": window,
"evidence_limit": "Recording timestamps do not establish when the source measured each value. Repeated values may be cached observations. The legacy summary.duration_ms field contains nanoseconds. Stored recording status is historical and does not mean recording is active. This archive is not the canonical incident timeline.",
}), nil
}
// Get windows for the resource
windows := e.incidentRecorderProvider.GetWindowsForResource(resourceID, limit)
if len(windows) == 0 {
return NewTextResult(fmt.Sprintf("No incident recording data found for resource '%s'. Incident data is captured when alerts fire.", resourceID)), nil
limit := intArg(args, "limit", 50)
if limit < 1 || limit > 200 {
return NewErrorResult(fmt.Errorf("limit must be between 1 and 200")), nil
}
queriedAt := time.Now().UTC()
since := queriedAt.Add(-24 * time.Hour)
if value, exists := args["since"]; exists {
text, ok := value.(string)
if !ok {
return NewErrorResult(fmt.Errorf("since must be an RFC3339 timestamp")), nil
}
var err error
since, err = time.Parse(time.RFC3339, text)
if err != nil || since.After(queriedAt) {
return NewErrorResult(fmt.Errorf("since must be an RFC3339 timestamp no later than now")), nil
}
}
if e.actionAuditStore == nil {
return NewErrorResult(fmt.Errorf("canonical resource history is unavailable")), nil
}
// This organization-pinned store is also used by the resource history API
// and Assistant handoffs. Do not reconstruct history from current metrics,
// match resource names, or include adjacent resources implicitly.
events, err := e.actionAuditStore.GetRecentChanges(resourceID, since, limit+1)
if err != nil {
return NewErrorResult(fmt.Errorf("read canonical resource history: %w", err)), nil
}
hasMore := len(events) > limit
if hasMore {
events = events[:limit]
}
if events == nil {
events = []unifiedresources.ResourceChange{}
}
return NewJSONResult(map[string]interface{}{
"resource_id": resourceID,
"windows": windows,
"count": len(windows),
"resource_id": resourceID,
"source": "canonical_resource_timeline",
"since": since,
"queried_at": queriedAt,
"time_basis": "observed_at",
"events": events,
"count": len(events),
"limit": limit,
"has_more": hasMore,
"coverage": "retained_records_only",
"evidence_limit": "These are retained observations, not continuous coverage. Empty history does not establish health or absence of incidents. ObservedAt is when Pulse observed a change, while OccurredAt is present only when its occurrence time is known. An alert resolving establishes that alert's recovery, not its cause or a verified action outcome.",
}), nil
}
+1 -1
View File
@@ -178,7 +178,7 @@ func (e *PulseToolExecutor) executeProposeAction(ctx context.Context, args map[s
return NewErrorResult(err), nil
}
return NewTextResult(fmt.Sprintf(
"Proposal recorded: capability %q on resource %q. It will be planned and routed for governed approval; nothing has executed. Conclude the investigation with your diagnosis.",
"Proposal recorded: capability %q on resource %q. The action broker still needs to validate it after this investigation. No action has been created or executed.",
capabilityName, resourceID)), nil
}
+159 -128
View File
@@ -2151,7 +2151,7 @@ func (e *PulseToolExecutor) registerQueryTools() {
e.registry.registerBuiltin(RegisteredTool{
Definition: Tool{
Name: agentcapabilities.PulseQueryToolName,
Description: `Query and search canonical infrastructure resources. Start here to discover systems, workloads, storage, and disks by name. Actions: search, get, config, topology, list, health. Health returns the connection overview by default, or the canonical resource projection when resource_id is provided.`,
Description: `Query and search canonical infrastructure resources. Start here to discover systems, workloads, storage, and disks by name. Actions: search, get, config, topology, list, health. Health returns the connection overview by default, or the canonical resource projection when resource_id is provided. command_agent_connected describes live command transport, independently of monitoring collection or freshness. Missing connection fields were not observed. can_execute describes connected transport with control enabled, not approval for a particular operation.`,
InputSchema: InputSchema{
Type: "object",
Properties: map[string]PropertySchema{
@@ -2166,7 +2166,7 @@ func (e *PulseToolExecutor) registerQueryTools() {
},
"resource_type": {
Type: "string",
Description: "Resource type. For get/search, prefer canonical values: 'agent', 'vm', 'system-container', 'app-container', 'storage', 'physical-disk', and 'docker-host'. For get, 'node' resolves to 'agent'. For search, 'node' filters Proxmox nodes. Compatibility aliases 'system' and 'storage-pool' are still accepted. For config: 'vm', 'system-container', or supported API-backed 'app-container'.",
Description: "Resource type. For get/search, prefer canonical values: 'agent', 'vm', 'system-container', 'app-container', 'storage', 'physical-disk', and 'docker-host'. For get, 'node' resolves to 'agent'. For search, 'node' filters Proxmox nodes. Compatibility aliases 'system' and 'storage-pool' are still accepted. For config: 'vm', 'system-container', or TrueNAS 'app-container'. Docker and Podman app-container configuration reads are not supported. Their collected health, mounts, ports, and networks are available through get.",
Enum: []string{"agent", "system", "vm", "system-container", "app-container", "node", "docker-host", "storage", "storage-pool", "physical-disk"},
},
"resource_id": {
@@ -2463,17 +2463,35 @@ func resourceHostCandidates(resource unifiedresources.Resource) []string {
return candidates
}
func resourceAgentConnected(resource unifiedresources.Resource, connected map[string]bool) bool {
for _, candidate := range resourceHostCandidates(resource) {
key := strings.TrimSpace(candidate)
if key == "" {
continue
}
if connected[key] {
return true
// commandConnectionObservation keeps an unqueried snapshot distinct from an
// observed disconnected transport. It says nothing about telemetry freshness.
func commandConnectionObservation(snapshot map[string]bool, value bool) *bool {
if snapshot == nil {
return nil
}
return &value
}
func resourceCommandAgentConnected(resource unifiedresources.Resource, connected map[string]bool) *bool {
// A guest's provider node/host identifies placement, not a command
// connection inside the guest. Parent transport is projected separately.
candidates := []string{resourceDisplayName(resource)}
candidates = append(candidates, resource.Identity.Hostnames...)
if resource.Agent != nil {
candidates = append(candidates, resource.Agent.Hostname)
}
switch resource.Type {
case unifiedresources.ResourceTypeVM, unifiedresources.ResourceTypeSystemContainer, unifiedresources.ResourceTypeAppContainer:
// Only the guest's own identity can establish its direct connection.
default:
candidates = append(candidates, resourceHostCandidates(resource)...)
}
for _, candidate := range candidates {
if key := strings.TrimSpace(candidate); key != "" && connected[key] {
return commandConnectionObservation(connected, true)
}
}
return false
return commandConnectionObservation(connected, false)
}
func appContainerProviderID(resource unifiedresources.Resource) string {
@@ -2972,16 +2990,16 @@ func addCanonicalGuestSearchMatches(
node := canonicalGuestTarget(resource)
metadataCandidates := append([]string{resourceDisplayName(resource), resource.ID}, candidates...)
addMatch(ResourceMatch{
GovernedResourceMetadata: governance.Resolve(metadataCandidates...),
Type: kind,
ID: resource.ID,
Name: resourceDisplayName(resource),
Status: status,
Node: node,
NodeHasAgent: connectedAgentHostnames[node],
Platform: canonicalResourcePlatform(resource),
VMID: vmid,
AgentConnected: resourceAgentConnected(resource, connectedAgentHostnames),
GovernedResourceMetadata: governance.Resolve(metadataCandidates...),
Type: kind,
ID: resource.ID,
Name: resourceDisplayName(resource),
Status: status,
Node: node,
NodeCommandAgentConnected: commandConnectionObservation(connectedAgentHostnames, connectedAgentHostnames[node]),
Platform: canonicalResourcePlatform(resource),
VMID: vmid,
CommandAgentConnected: resourceCommandAgentConnected(resource, connectedAgentHostnames),
})
}
}
@@ -3012,16 +3030,16 @@ func addGuestViewSearchMatches[V queryGuestView](
continue
}
addMatch(ResourceMatch{
GovernedResourceMetadata: governance.Resolve(g.Name(), g.ID(), vmidStr),
Type: kind,
ID: g.ID(),
Name: g.Name(),
Status: status,
Node: g.Node(),
NodeHasAgent: connectedAgentHostnames[g.Node()],
Platform: "proxmox",
VMID: g.VMID(),
AgentConnected: connectedAgentHostnames[g.Name()],
GovernedResourceMetadata: governance.Resolve(g.Name(), g.ID(), vmidStr),
Type: kind,
ID: g.ID(),
Name: g.Name(),
Status: status,
Node: g.Node(),
NodeCommandAgentConnected: commandConnectionObservation(connectedAgentHostnames, connectedAgentHostnames[g.Node()]),
Platform: "proxmox",
VMID: g.VMID(),
CommandAgentConnected: commandConnectionObservation(connectedAgentHostnames, connectedAgentHostnames[g.Name()]),
})
}
}
@@ -3469,15 +3487,15 @@ func resolvedAppContainerRegistration(resource unifiedresources.Resource) (Resou
func canonicalSystemSummaryFromResource(resource unifiedresources.Resource, connected map[string]bool) SystemSummary {
return SystemSummary{
ID: strings.TrimSpace(resource.ID),
Name: resourceDisplayName(resource),
Status: string(resource.Status),
Platform: canonicalResourcePlatform(resource),
ChildCount: resource.ChildCount,
AgentConnected: resourceAgentConnected(resource, connected),
CPU: metricPercent(resourceMetric(resource, "cpu")),
Memory: metricPercent(resourceMetric(resource, "memory")),
Disk: metricPercent(resourceMetric(resource, "disk")),
ID: strings.TrimSpace(resource.ID),
Name: resourceDisplayName(resource),
Status: string(resource.Status),
Platform: canonicalResourcePlatform(resource),
ChildCount: resource.ChildCount,
CommandAgentConnected: resourceCommandAgentConnected(resource, connected),
CPU: metricPercent(resourceMetric(resource, "cpu")),
Memory: metricPercent(resourceMetric(resource, "memory")),
Disk: metricPercent(resourceMetric(resource, "disk")),
}
}
@@ -3809,7 +3827,7 @@ func (e *PulseToolExecutor) executeListInfrastructure(_ context.Context, args ma
GovernedResourceMetadata: governance.Resolve(node.Name(), node.ID()),
Name: node.Name(),
Status: string(node.Status()),
AgentConnected: connectedAgentHostnames[node.Name()],
CommandAgentConnected: commandConnectionObservation(connectedAgentHostnames, connectedAgentHostnames[node.Name()]),
})
count++
}
@@ -3993,7 +4011,7 @@ func (e *PulseToolExecutor) executeListInfrastructure(_ context.Context, args ma
Hostname: hostname,
DisplayName: displayName,
ContainerCount: len(hostContainers),
AgentConnected: connectedAgentHostnames[hostname] || connectedAgentHostnames[displayName],
CommandAgentConnected: commandConnectionObservation(connectedAgentHostnames, connectedAgentHostnames[hostname] || connectedAgentHostnames[displayName]),
}
for _, container := range hostContainers {
state := strings.TrimSpace(container.ContainerState())
@@ -4248,7 +4266,7 @@ type TopologyBuildOptions struct {
MaxK8sNodesPerCluster int
MaxK8sDeploymentsPerCluster int
MaxK8sPodsPerCluster int
ConnectedAgentHostnames map[string]bool
ConnectedAgentHostnames map[string]bool // nil means command connections were not observed
ControlEnabled bool
}
@@ -4266,9 +4284,6 @@ func BuildTopologyResponseFromReadState(rs unifiedresources.ReadState, options T
includeDocker := include == "all" || include == "app-containers"
includeKubernetes := include == "all" || include == "kubernetes"
connectedAgentHostnames := options.ConnectedAgentHostnames
if connectedAgentHostnames == nil {
connectedAgentHostnames = map[string]bool{}
}
governance := newGovernedQueryMetadataResolver(rs)
summary := TopologySummary{
@@ -4283,12 +4298,17 @@ func BuildTopologyResponseFromReadState(rs unifiedresources.ReadState, options T
TotalK8sPods: len(rs.Pods()),
}
if connectedAgentHostnames != nil {
summary.NodesWithCommandAgents = new(int)
summary.DockerHostsWithCommandAgents = new(int)
}
for _, node := range rs.Nodes() {
if node == nil {
continue
}
if connectedAgentHostnames[node.Name()] {
summary.NodesWithAgents++
(*summary.NodesWithCommandAgents)++
}
}
for _, host := range rs.DockerHosts() {
@@ -4298,7 +4318,7 @@ func BuildTopologyResponseFromReadState(rs unifiedresources.ReadState, options T
hostname := strings.TrimSpace(host.Hostname())
displayName := strings.TrimSpace(host.Name())
if connectedAgentHostnames[hostname] || connectedAgentHostnames[displayName] {
summary.DockerHostsWithAgents++
(*summary.DockerHostsWithCommandAgents)++
}
}
for _, pod := range rs.Pods() {
@@ -4325,8 +4345,8 @@ func BuildTopologyResponseFromReadState(rs unifiedresources.ReadState, options T
GovernedResourceMetadata: governance.Resolve(node.Name(), node.ID()),
Name: name,
Status: string(node.Status()),
AgentConnected: hasAgent,
CanExecute: hasAgent && options.ControlEnabled,
CommandAgentConnected: commandConnectionObservation(connectedAgentHostnames, hasAgent),
CanExecute: commandConnectionObservation(connectedAgentHostnames, hasAgent && options.ControlEnabled),
VMs: []TopologyVM{},
Containers: []TopologyContainer{},
}
@@ -4348,8 +4368,8 @@ func BuildTopologyResponseFromReadState(rs unifiedresources.ReadState, options T
GovernedResourceMetadata: governance.Resolve(name),
Name: name,
Status: status,
AgentConnected: hasAgent,
CanExecute: hasAgent && options.ControlEnabled,
CommandAgentConnected: commandConnectionObservation(connectedAgentHostnames, hasAgent),
CanExecute: commandConnectionObservation(connectedAgentHostnames, hasAgent && options.ControlEnabled),
VMs: []TopologyVM{},
Containers: []TopologyContainer{},
}
@@ -4489,8 +4509,8 @@ func BuildTopologyResponseFromReadState(rs unifiedresources.ReadState, options T
GovernedResourceMetadata: governance.Resolve(host.Hostname(), host.Name(), host.HostSourceID(), host.ID()),
Hostname: hostname,
DisplayName: displayName,
AgentConnected: hasAgent,
CanExecute: hasAgent && options.ControlEnabled,
CommandAgentConnected: commandConnectionObservation(connectedAgentHostnames, hasAgent),
CanExecute: commandConnectionObservation(connectedAgentHostnames, hasAgent && options.ControlEnabled),
Containers: containers,
ContainerCount: len(hostContainers),
ReturnedCount: len(containers),
@@ -5037,9 +5057,11 @@ func (e *PulseToolExecutor) executeGetResource(_ context.Context, args map[strin
}
for _, m := range resource.Docker.Mounts {
response.Mounts = append(response.Mounts, MountInfo{
Type: m.Type,
Source: m.Source,
Destination: m.Destination,
ReadWrite: !strings.EqualFold(strings.TrimSpace(m.Mode), "ro"),
Mode: m.Mode,
ReadWrite: m.RW,
})
}
}
@@ -5144,8 +5166,10 @@ func (e *PulseToolExecutor) executeGetResource(_ context.Context, args map[strin
for _, m := range container.Mounts() {
response.Mounts = append(response.Mounts, MountInfo{
Type: m.Type,
Source: m.Source,
Destination: m.Destination,
Mode: m.Mode,
ReadWrite: m.RW,
})
}
@@ -5276,110 +5300,117 @@ func (e *PulseToolExecutor) executeGetResourceConfig(ctx context.Context, args m
}
func (e *PulseToolExecutor) executeNativeAppContainerConfig(ctx context.Context, resourceRef string) (CallToolResult, error) {
if e.appContainerConfigProvider == nil {
return NewTextResult("App-container configuration not available."), nil
validation := e.validateResolvedResource(resourceRef, "query", true)
if validation.ErrorMsg != "" {
return NewErrorResult(fmt.Errorf("%s", validation.ErrorMsg)), nil
}
if validation.Resource != nil && validation.Resource.GetKind() != "app-container" {
return NewErrorResult(fmt.Errorf("resource '%s' is %q, not app-container", resourceRef, validation.Resource.GetKind())), nil
}
if e.unifiedResourceProvider == nil {
return NewErrorResult(fmt.Errorf("current app-container inventory is unavailable")), nil
}
rs, err := e.readStateForControl()
if err != nil {
return NewTextResult("State information not available."), nil
return NewErrorResult(fmt.Errorf("current app-container state is unavailable: %w", err)), nil
}
governance := newGovernedQueryMetadataResolver(rs)
var resource unifiedresources.Resource
var found bool
if validation := e.validateResolvedResource(resourceRef, "query", true); validation.Resource != nil {
if matched, _, ok := findCanonicalAppContainerResource(e.unifiedResourceProvider, resourceRef); ok {
resource = matched
found = true
}
}
resource, providerID, found := findCanonicalAppContainerResource(e.unifiedResourceProvider, resourceRef)
if !found {
var containerID string
resource, containerID, found = findCanonicalAppContainerResource(e.unifiedResourceProvider, resourceRef)
if !found {
return NewJSONResult(map[string]interface{}{
"error": "not_found",
"resource_id": resourceRef,
"type": "app-container",
}), nil
}
if reg, ok := resolvedAppContainerRegistration(resource); ok {
e.registerResolvedResourceWithExplicitAccess(reg)
}
_ = containerID
return NewJSONResult(map[string]interface{}{
"error": "not_found",
"resource_id": resourceRef,
"type": "app-container",
}), nil
}
validation := e.validateResolvedResource(resourceRef, "query", true)
if validation.Resource == nil {
if validation.ErrorMsg != "" {
return NewErrorResult(fmt.Errorf("%s", validation.ErrorMsg)), nil
}
return NewErrorResult(fmt.Errorf("app-container not found: %s", resourceRef)), nil
// Inventory owns read identity and capability. Optional session discovery
// supplies restrictions and continuity, not proof that the resource exists.
resourceID := canonicalAppContainerID(resource)
canonicalValidation := e.validateResolvedResource(resourceID, "query", true)
if canonicalValidation.ErrorMsg != "" {
return NewErrorResult(fmt.Errorf("%s", canonicalValidation.ErrorMsg)), nil
}
if validation.ErrorMsg != "" {
return NewErrorResult(fmt.Errorf("%s", validation.ErrorMsg)), nil
if canonicalValidation.Resource != nil && canonicalValidation.Resource.GetKind() != "app-container" {
return NewErrorResult(fmt.Errorf("resource '%s' is %q, not app-container", resourceID, canonicalValidation.Resource.GetKind())), nil
}
resolved := validation.Resource
if resolved.GetKind() != "app-container" {
return NewErrorResult(fmt.Errorf("resource '%s' is %q, not app-container", resourceRef, resolved.GetKind())), nil
platform := canonicalAppContainerAdapter(resource)
unavailable := func(reason, message string) CallToolResult {
return NewJSONResultWithIsError(map[string]interface{}{
"available": false, "reason": reason, "message": message,
"resource_id": resourceID, "type": "app-container", "platform": platform,
}, true)
}
if !strings.EqualFold(strings.TrimSpace(resolved.GetAdapter()), "truenas") {
return NewTextResult("App-container configuration not available."), nil
if platform != "truenas" {
return unavailable("unsupported_adapter", "The resource exists, but its adapter does not support configuration reads."), nil
}
if e.appContainerConfigProvider == nil {
return unavailable("provider_unavailable", "The resource exists, but its configuration provider is unavailable."), nil
}
reg, ok := resolvedAppContainerRegistration(resource)
if !ok {
return unavailable("resource_context_unavailable", "The resource exists, but its current provider identity or placement is incomplete."), nil
}
// Do not overwrite an existing session's allowed actions during a read.
if validation.Resource == nil && canonicalValidation.Resource == nil {
e.registerResolvedResourceWithExplicitAccess(reg)
}
result, err := e.appContainerConfigProvider.GetConfig(ctx, AppContainerConfigRequest{
OrgID: e.orgID,
ResourceID: strings.TrimSpace(resolved.GetResourceID()),
ProviderUID: strings.TrimSpace(resolved.GetProviderUID()),
ResourceID: resourceID,
ProviderUID: providerID,
Name: resourceDisplayName(resource),
Host: strings.TrimSpace(resolved.GetTargetHost()),
Platform: "truenas",
Host: canonicalAppContainerHost(resource),
Platform: platform,
})
if err != nil {
return NewErrorResult(err), nil
}
if result == nil {
return unavailable("empty_provider_response", "The resource exists, but the provider returned no configuration observation."), nil
}
response := EmptyAppContainerConfigResponse()
if result != nil {
response.GovernedResourceMetadata = governance.Resolve(result.Name, result.ResourceID, result.ProviderUID)
response.Type = "app-container"
response.ID = result.ProviderUID
if response.ID == "" {
response.ID = strings.TrimSpace(result.ResourceID)
}
response.Name = result.Name
response.Host = result.Host
response.Platform = result.Platform
response.Status = result.Status
response.Version = result.Version
response.HumanVersion = result.HumanVersion
response.Notes = result.Notes
response.CustomApp = result.CustomApp
response.UpgradeAvailable = result.UpgradeAvailable
response.ImageUpdatesAvailable = result.ImageUpdatesAvailable
response.ContainerCount = result.ContainerCount
response.UsedHostIPs = append([]string{}, result.UsedHostIPs...)
response.Images = append([]string{}, result.Images...)
response.Ports = append([]PortInfo{}, result.Ports...)
response.Networks = append([]NetworkInfo{}, result.Networks...)
response.Mounts = append([]MountInfo{}, result.Mounts...)
response.Containers = append([]AppContainerConfigContainer{}, result.Containers...)
}
response.GovernedResourceMetadata = governance.Resolve(result.Name, result.ResourceID, result.ProviderUID)
response.Type = "app-container"
response.ID = result.ProviderUID
if response.ID == "" {
response.ID = strings.TrimSpace(resolved.GetProviderUID())
response.ID = strings.TrimSpace(result.ResourceID)
}
response.Name = result.Name
response.Host = result.Host
response.Platform = result.Platform
response.Status = result.Status
response.Version = result.Version
response.HumanVersion = result.HumanVersion
response.Notes = result.Notes
response.CustomApp = result.CustomApp
response.UpgradeAvailable = result.UpgradeAvailable
response.ImageUpdatesAvailable = result.ImageUpdatesAvailable
response.ContainerCount = result.ContainerCount
response.UsedHostIPs = append([]string{}, result.UsedHostIPs...)
response.Images = append([]string{}, result.Images...)
response.Ports = append([]PortInfo{}, result.Ports...)
response.Networks = append([]NetworkInfo{}, result.Networks...)
response.Mounts = append([]MountInfo{}, result.Mounts...)
response.Containers = append([]AppContainerConfigContainer{}, result.Containers...)
if response.ID == "" {
response.ID = providerID
}
if response.Name == "" {
response.Name = resolvedResourceDisplayName(resolved)
response.Name = resourceDisplayName(resource)
}
if response.Host == "" {
response.Host = strings.TrimSpace(resolved.GetTargetHost())
response.Host = canonicalAppContainerHost(resource)
}
if response.Platform == "" {
response.Platform = strings.TrimSpace(resolved.GetAdapter())
response.Platform = platform
}
if response.GovernedResourceMetadata.Policy == nil && response.AISafeSummary == "" {
response.GovernedResourceMetadata = governance.Resolve(response.Name, strings.TrimSpace(resolved.GetResourceID()), response.ID)
response.GovernedResourceMetadata = governance.Resolve(response.Name, resourceID, response.ID)
}
return NewJSONResult(response.NormalizeCollections()), nil
@@ -5714,7 +5745,7 @@ func (e *PulseToolExecutor) executeSearchResources(_ context.Context, args map[s
Status: status,
Host: canonicalAgentHost(resource),
Platform: canonicalResourcePlatform(resource),
AgentConnected: resourceAgentConnected(resource, connectedAgentHostnames),
CommandAgentConnected: resourceCommandAgentConnected(resource, connectedAgentHostnames),
})
}
}
@@ -5733,7 +5764,7 @@ func (e *PulseToolExecutor) executeSearchResources(_ context.Context, args map[s
Type: "node",
Name: node.Name(),
Status: status,
AgentConnected: connectedAgentHostnames[node.Name()],
CommandAgentConnected: commandConnectionObservation(connectedAgentHostnames, connectedAgentHostnames[node.Name()]),
})
}
}
@@ -3,13 +3,18 @@ package tools
import (
"context"
"encoding/json"
"errors"
"strings"
"testing"
"github.com/rcourtman/pulse-go-rewrite/internal/unifiedresources"
)
type stubAppContainerConfigProvider struct {
calls []AppContainerConfigRequest
result *AppContainerConfigResult
err error
empty bool
}
func (s *stubAppContainerConfigProvider) GetConfig(_ context.Context, req AppContainerConfigRequest) (*AppContainerConfigResult, error) {
@@ -17,6 +22,9 @@ func (s *stubAppContainerConfigProvider) GetConfig(_ context.Context, req AppCon
if s.err != nil {
return nil, s.err
}
if s.empty {
return nil, nil
}
if s.result == nil {
return &AppContainerConfigResult{
ResourceID: req.ResourceID,
@@ -38,6 +46,153 @@ func (s *stubAppContainerConfigProvider) GetConfig(_ context.Context, req AppCon
return &result, nil
}
func TestAppContainerConfigObservationContract(t *testing.T) {
t.Setenv("PULSE_STRICT_RESOLUTION", "true")
for _, tc := range []struct {
name, session, reason, failure string
missing, docker, noProvider, noHost, empty bool
noInventory, noReadState bool
}{
{name: "without session"},
{name: "empty session", session: "empty"},
{name: "discovered session", session: "discovered"},
{name: "stale placement", session: "stale"},
{name: "unsupported adapter", docker: true, reason: "unsupported_adapter"},
{name: "unavailable provider", noProvider: true, reason: "provider_unavailable"},
{name: "missing placement", noHost: true, reason: "resource_context_unavailable"},
{name: "empty provider response", empty: true, reason: "empty_provider_response"},
{name: "provider failed", failure: "provider read failed"},
{name: "resource absent", missing: true},
{name: "inventory unavailable", noInventory: true, failure: "inventory is unavailable"},
{name: "read state unavailable", noReadState: true, failure: "state is unavailable"},
{name: "query denied", session: "denied", failure: "not permitted"},
{name: "canonical query denied through prefix", session: "canonical-denied", failure: "not permitted"},
} {
t.Run(tc.name, func(t *testing.T) {
registry := newTrueNASUnifiedQueryProvider(t)
resource, _, found := findCanonicalAppContainerResource(registry, "nextcloud")
if !found {
t.Fatal("missing canonical fixture")
}
if tc.docker {
resource.TrueNAS = nil
resource.Tags = nil
}
if tc.noHost {
resource.ParentName = ""
resource.Identity.Hostnames = nil
}
provider := &stubUnifiedResourceProvider{resources: []unifiedresources.Resource{resource}}
config := &stubAppContainerConfigProvider{empty: tc.empty}
if tc.failure == "provider read failed" {
config.err = errors.New(tc.failure)
}
cfg := ExecutorConfig{UnifiedResourceProvider: provider, ReadState: registry.ResourceRegistry}
if tc.noInventory {
cfg.UnifiedResourceProvider = nil
}
if tc.noReadState {
cfg.ReadState = nil
}
if !tc.noProvider {
cfg.AppContainerConfigProvider = config
}
executor := NewPulseToolExecutor(cfg)
ref := "Nextcloud"
if tc.session != "" {
resolved := &mockResolvedContext{resources: map[string]ResolvedResourceInfo{}, aliases: map[string]ResolvedResourceInfo{}}
executor.SetResolvedContext(resolved)
switch tc.session {
case "discovered":
reg, ok := resolvedAppContainerRegistration(resource)
if !ok {
t.Fatal("fixture registration unavailable")
}
resolved.AddResolvedResource(reg)
case "stale", "denied", "canonical-denied":
cached := &mockResource{resourceID: resource.ID, kind: "app-container", adapter: "docker", targetHost: "stale-host", providerUID: "stale-id", allowedActions: []string{"query"}}
if tc.session != "stale" {
cached.allowedActions = []string{"logs"}
}
resolved.resources[resource.ID] = cached
if tc.session == "canonical-denied" {
ref = "next"
} else {
resolved.aliases[ref] = cached
}
}
}
if tc.missing {
ref = "absent-container"
}
// Inventory get succeeds independently of optional session state.
if tc.session == "" && !tc.missing && !tc.noInventory && !tc.noReadState {
got, err := executor.executeGetResource(context.Background(), map[string]interface{}{"resource_type": "app-container", "resource_id": ref})
if err != nil || got.IsError || strings.Contains(got.Content[0].Text, "not_found") {
t.Fatalf("canonical get failed: %+v %v", got, err)
}
}
args := map[string]interface{}{"action": "config", "resource_type": "app-container", "resource_id": ref}
result, err := executor.executeQuery(context.Background(), args)
if err != nil {
t.Fatal(err)
}
evidence, err := json.Marshal(map[string]interface{}{"case": tc.name, "input": args, "result": result})
if err != nil {
t.Fatal(err)
}
t.Logf("CONFIG_EVIDENCE %s", evidence)
wantError := tc.failure != "" || tc.reason != ""
if result.IsError != wantError {
t.Fatalf("read error bit=%v, want %v: %+v", result.IsError, wantError, result)
}
if tc.failure != "" {
if !result.IsError || !strings.Contains(result.Content[0].Text, tc.failure) {
t.Fatalf("expected %q failure, got %+v", tc.failure, result)
}
} else {
var response map[string]interface{}
if err := json.Unmarshal([]byte(result.Content[0].Text), &response); err != nil {
t.Fatal(err)
}
switch {
case tc.missing:
if response["error"] != "not_found" {
t.Fatalf("expected true absence, got %+v", response)
}
case tc.reason != "":
if response["available"] != false || response["reason"] != tc.reason || response["resource_id"] != resource.ID {
t.Fatalf("unavailable configuration lost identity or reason: %+v", response)
}
default:
if response["id"] != appContainerProviderID(resource) || response["host"] != canonicalAppContainerHost(resource) || response["platform"] != "truenas" {
t.Fatalf("incorrect config identity: %+v", response)
}
}
}
wantCalls := 1
if tc.missing || tc.docker || tc.noProvider || tc.noHost || tc.noInventory || tc.noReadState || strings.Contains(tc.session, "denied") {
wantCalls = 0
}
if len(config.calls) != wantCalls {
t.Fatalf("provider calls=%d, want %d", len(config.calls), wantCalls)
}
if wantCalls == 1 {
call := config.calls[0]
if call.ResourceID != resource.ID || call.ProviderUID != appContainerProviderID(resource) || call.Host != canonicalAppContainerHost(resource) || call.Platform != "truenas" {
t.Fatalf("request used session identity instead of canonical inventory: %+v", call)
}
}
if tc.session == "stale" {
cached, ok := executor.resolvedContext.GetResolvedResourceByID(resource.ID)
if !ok || strings.Join(cached.GetAllowedActions(), ",") != "query" {
t.Fatal("read expanded existing session action authority")
}
}
})
}
}
func TestExecuteGetResourceConfig_TrueNASAppUsesNativeConfigProvider(t *testing.T) {
provider := newTrueNASUnifiedQueryProvider(t)
resolved := &mockResolvedContext{
+1 -1
View File
@@ -78,7 +78,7 @@ type AIService interface {
SetFindingsManager(manager chat.FindingsManager)
SetMetadataUpdater(updater chat.MetadataUpdater)
SetKnowledgeStoreProvider(provider chat.KnowledgeStoreProvider)
SetIncidentRecorderProvider(provider chat.IncidentRecorderProvider)
SetIncidentArchiveProvider(provider chat.IncidentArchiveProvider)
SetEventCorrelatorProvider(provider chat.EventCorrelatorProvider)
SetDiscoveryProvider(provider chat.AssistantDiscoveryProvider)
SetUnifiedResourceProvider(provider chat.AssistantUnifiedResourceProvider)
@@ -88,15 +88,15 @@ func (s *capturingAIService) SetGuestConfigProvider(provider chat.AssistantGuest
func (s *capturingAIService) SetAppContainerConfigProvider(provider chat.AssistantAppContainerConfigProvider) {
s.appContainerConfigProvider = provider
}
func (s *capturingAIService) SetBackupProvider(provider chat.AssistantBackupProvider) {}
func (s *capturingAIService) SetDiskHealthProvider(provider chat.AssistantDiskHealthProvider) {}
func (s *capturingAIService) SetUpdatesProvider(provider chat.AssistantUpdatesProvider) {}
func (s *capturingAIService) SetFindingsManager(manager chat.FindingsManager) {}
func (s *capturingAIService) SetMetadataUpdater(updater chat.MetadataUpdater) {}
func (s *capturingAIService) SetKnowledgeStoreProvider(provider chat.KnowledgeStoreProvider) {}
func (s *capturingAIService) SetIncidentRecorderProvider(provider chat.IncidentRecorderProvider) {}
func (s *capturingAIService) SetEventCorrelatorProvider(provider chat.EventCorrelatorProvider) {}
func (s *capturingAIService) SetDiscoveryProvider(provider chat.AssistantDiscoveryProvider) {}
func (s *capturingAIService) SetBackupProvider(provider chat.AssistantBackupProvider) {}
func (s *capturingAIService) SetDiskHealthProvider(provider chat.AssistantDiskHealthProvider) {}
func (s *capturingAIService) SetUpdatesProvider(provider chat.AssistantUpdatesProvider) {}
func (s *capturingAIService) SetFindingsManager(manager chat.FindingsManager) {}
func (s *capturingAIService) SetMetadataUpdater(updater chat.MetadataUpdater) {}
func (s *capturingAIService) SetKnowledgeStoreProvider(provider chat.KnowledgeStoreProvider) {}
func (s *capturingAIService) SetIncidentArchiveProvider(provider chat.IncidentArchiveProvider) {}
func (s *capturingAIService) SetEventCorrelatorProvider(provider chat.EventCorrelatorProvider) {}
func (s *capturingAIService) SetDiscoveryProvider(provider chat.AssistantDiscoveryProvider) {}
func (s *capturingAIService) SetUnifiedResourceProvider(provider chat.AssistantUnifiedResourceProvider) {
}
func (s *capturingAIService) SetAppContainerActionProvider(provider chat.AssistantAppContainerActionProvider) {
+1 -1
View File
@@ -281,7 +281,7 @@ func (m *MockAIService) SetMetadataUpdater(updater chat.MetadataUpdater) { m.Cal
func (m *MockAIService) SetKnowledgeStoreProvider(provider chat.KnowledgeStoreProvider) {
m.Called(provider)
}
func (m *MockAIService) SetIncidentRecorderProvider(provider chat.IncidentRecorderProvider) {
func (m *MockAIService) SetIncidentArchiveProvider(provider chat.IncidentArchiveProvider) {
m.Called(provider)
}
func (m *MockAIService) SetEventCorrelatorProvider(provider chat.EventCorrelatorProvider) {
+67 -160
View File
@@ -92,22 +92,20 @@ type AISettingsHandler struct {
alertBridge *unified.AlertBridge // Bridge between alerts and unified store
// Event-driven patrol (Phase 7)
triggerManager *ai.TriggerManager // Event-driven patrol trigger manager
incidentCoordinator *ai.IncidentCoordinator // Incident recording coordinator
incidentRecorder *metrics.IncidentRecorder // High-frequency incident recorder
intelligenceMu sync.RWMutex
proxmoxCorrelators map[string]*proxmox.EventCorrelator
learningStores map[string]*learning.LearningStore
forecastServices map[string]*forecast.Service
remediationEngines map[string]aicontracts.RemediationEngine
incidentStores map[string]*memory.IncidentStore
circuitBreakers map[string]*circuit.Breaker
discoveryStores map[string]*servicediscovery.Store
unifiedStores map[string]*unified.UnifiedStore
alertBridges map[string]*unified.AlertBridge
triggerManagers map[string]*ai.TriggerManager
incidentCoordinators map[string]*ai.IncidentCoordinator
incidentRecorders map[string]*metrics.IncidentRecorder
triggerManager *ai.TriggerManager // Event-driven patrol trigger manager
incidentArchive *metrics.IncidentArchive // Read-only legacy incident archive
intelligenceMu sync.RWMutex
proxmoxCorrelators map[string]*proxmox.EventCorrelator
learningStores map[string]*learning.LearningStore
forecastServices map[string]*forecast.Service
remediationEngines map[string]aicontracts.RemediationEngine
incidentStores map[string]*memory.IncidentStore
circuitBreakers map[string]*circuit.Breaker
discoveryStores map[string]*servicediscovery.Store
unifiedStores map[string]*unified.UnifiedStore
alertBridges map[string]*unified.AlertBridge
triggerManagers map[string]*ai.TriggerManager
incidentArchives map[string]*metrics.IncidentArchive
// Investigation orchestration (Patrol Autonomy)
chatHandler *AIHandler // Chat service handler for investigations
@@ -318,25 +316,24 @@ func NewAISettingsHandler(mtp *config.MultiTenantPersistence, mtm *monitoring.Mu
}
handler := &AISettingsHandler{
mtPersistence: mtp,
mtMonitor: mtm,
defaultConfig: defaultConfig,
defaultPersistence: defaultPersistence,
hostedMode: hostedMode,
aiServices: make(map[string]*ai.Service),
agentServer: agentServer,
proxmoxCorrelators: make(map[string]*proxmox.EventCorrelator),
learningStores: make(map[string]*learning.LearningStore),
forecastServices: make(map[string]*forecast.Service),
remediationEngines: make(map[string]aicontracts.RemediationEngine),
incidentStores: make(map[string]*memory.IncidentStore),
circuitBreakers: make(map[string]*circuit.Breaker),
discoveryStores: make(map[string]*servicediscovery.Store),
unifiedStores: make(map[string]*unified.UnifiedStore),
alertBridges: make(map[string]*unified.AlertBridge),
triggerManagers: make(map[string]*ai.TriggerManager),
incidentCoordinators: make(map[string]*ai.IncidentCoordinator),
incidentRecorders: make(map[string]*metrics.IncidentRecorder),
mtPersistence: mtp,
mtMonitor: mtm,
defaultConfig: defaultConfig,
defaultPersistence: defaultPersistence,
hostedMode: hostedMode,
aiServices: make(map[string]*ai.Service),
agentServer: agentServer,
proxmoxCorrelators: make(map[string]*proxmox.EventCorrelator),
learningStores: make(map[string]*learning.LearningStore),
forecastServices: make(map[string]*forecast.Service),
remediationEngines: make(map[string]aicontracts.RemediationEngine),
incidentStores: make(map[string]*memory.IncidentStore),
circuitBreakers: make(map[string]*circuit.Breaker),
discoveryStores: make(map[string]*servicediscovery.Store),
unifiedStores: make(map[string]*unified.UnifiedStore),
alertBridges: make(map[string]*unified.AlertBridge),
triggerManagers: make(map[string]*ai.TriggerManager),
incidentArchives: make(map[string]*metrics.IncidentArchive),
}
defaultAIService = ai.NewService(defaultPersistence, tenantAgentServerForOrganization(agentServer, "default"))
@@ -1198,11 +1195,8 @@ func (h *AISettingsHandler) ensureIntelligenceMapsLocked() {
if h.triggerManagers == nil {
h.triggerManagers = make(map[string]*ai.TriggerManager)
}
if h.incidentCoordinators == nil {
h.incidentCoordinators = make(map[string]*ai.IncidentCoordinator)
}
if h.incidentRecorders == nil {
h.incidentRecorders = make(map[string]*metrics.IncidentRecorder)
if h.incidentArchives == nil {
h.incidentArchives = make(map[string]*metrics.IncidentArchive)
}
}
@@ -1505,96 +1499,49 @@ func (h *AISettingsHandler) GetTriggerManagerForOrg(orgID string) *ai.TriggerMan
return nil
}
// SetIncidentCoordinator sets the incident recording coordinator
func (h *AISettingsHandler) SetIncidentCoordinator(coordinator *ai.IncidentCoordinator) {
h.SetIncidentCoordinatorForOrg("default", coordinator)
// SetIncidentArchive sets the read-only legacy incident archive
func (h *AISettingsHandler) SetIncidentArchive(archive *metrics.IncidentArchive) {
h.SetIncidentArchiveForOrg("default", archive)
}
// SetIncidentCoordinatorForOrg sets the incident recording coordinator for an org.
func (h *AISettingsHandler) SetIncidentCoordinatorForOrg(orgID string, coordinator *ai.IncidentCoordinator) {
// SetIncidentArchiveForOrg sets the read-only legacy incident archive for an org.
func (h *AISettingsHandler) SetIncidentArchiveForOrg(orgID string, archive *metrics.IncidentArchive) {
if h == nil {
return
}
orgID = normalizeAIIntelligenceOrgID(orgID)
h.intelligenceMu.Lock()
h.ensureIntelligenceMapsLocked()
if coordinator == nil {
delete(h.incidentCoordinators, orgID)
if archive == nil {
delete(h.incidentArchives, orgID)
} else {
h.incidentCoordinators[orgID] = coordinator
h.incidentArchives[orgID] = archive
}
h.intelligenceMu.Unlock()
if orgID == "default" {
h.incidentCoordinator = coordinator
h.incidentArchive = archive
}
}
// GetIncidentCoordinator returns the incident recording coordinator
func (h *AISettingsHandler) GetIncidentCoordinator() *ai.IncidentCoordinator {
return h.GetIncidentCoordinatorForOrg("default")
// GetIncidentArchive returns the read-only legacy incident archive
func (h *AISettingsHandler) GetIncidentArchive() *metrics.IncidentArchive {
return h.GetIncidentArchiveForOrg("default")
}
// GetIncidentCoordinatorForOrg returns the incident recording coordinator for an org.
func (h *AISettingsHandler) GetIncidentCoordinatorForOrg(orgID string) *ai.IncidentCoordinator {
// GetIncidentArchiveForOrg returns the read-only legacy incident archive for an org.
func (h *AISettingsHandler) GetIncidentArchiveForOrg(orgID string) *metrics.IncidentArchive {
if h == nil {
return nil
}
orgID = normalizeAIIntelligenceOrgID(orgID)
h.intelligenceMu.RLock()
if coordinator := h.incidentCoordinators[orgID]; coordinator != nil {
if archive := h.incidentArchives[orgID]; archive != nil {
h.intelligenceMu.RUnlock()
return coordinator
return archive
}
h.intelligenceMu.RUnlock()
if orgID == "default" {
return h.incidentCoordinator
}
return nil
}
// SetIncidentRecorder sets the high-frequency incident recorder
func (h *AISettingsHandler) SetIncidentRecorder(recorder *metrics.IncidentRecorder) {
h.SetIncidentRecorderForOrg("default", recorder)
}
// SetIncidentRecorderForOrg sets the high-frequency incident recorder for an org.
func (h *AISettingsHandler) SetIncidentRecorderForOrg(orgID string, recorder *metrics.IncidentRecorder) {
if h == nil {
return
}
orgID = normalizeAIIntelligenceOrgID(orgID)
h.intelligenceMu.Lock()
h.ensureIntelligenceMapsLocked()
if recorder == nil {
delete(h.incidentRecorders, orgID)
} else {
h.incidentRecorders[orgID] = recorder
}
h.intelligenceMu.Unlock()
if orgID == "default" {
h.incidentRecorder = recorder
}
}
// GetIncidentRecorder returns the high-frequency incident recorder
func (h *AISettingsHandler) GetIncidentRecorder() *metrics.IncidentRecorder {
return h.GetIncidentRecorderForOrg("default")
}
// GetIncidentRecorderForOrg returns the high-frequency incident recorder for an org.
func (h *AISettingsHandler) GetIncidentRecorderForOrg(orgID string) *metrics.IncidentRecorder {
if h == nil {
return nil
}
orgID = normalizeAIIntelligenceOrgID(orgID)
h.intelligenceMu.RLock()
if recorder := h.incidentRecorders[orgID]; recorder != nil {
h.intelligenceMu.RUnlock()
return recorder
}
h.intelligenceMu.RUnlock()
if orgID == "default" {
return h.incidentRecorder
return h.incidentArchive
}
return nil
}
@@ -1656,44 +1603,6 @@ func (h *AISettingsHandler) ListTriggerManagers() map[string]*ai.TriggerManager
return out
}
// ListIncidentCoordinators returns incident coordinators keyed by org.
func (h *AISettingsHandler) ListIncidentCoordinators() map[string]*ai.IncidentCoordinator {
out := make(map[string]*ai.IncidentCoordinator)
if h == nil {
return out
}
h.intelligenceMu.RLock()
for orgID, coordinator := range h.incidentCoordinators {
if coordinator != nil {
out[orgID] = coordinator
}
}
h.intelligenceMu.RUnlock()
if _, ok := out["default"]; !ok && h.incidentCoordinator != nil {
out["default"] = h.incidentCoordinator
}
return out
}
// ListIncidentRecorders returns incident recorders keyed by org.
func (h *AISettingsHandler) ListIncidentRecorders() map[string]*metrics.IncidentRecorder {
out := make(map[string]*metrics.IncidentRecorder)
if h == nil {
return out
}
h.intelligenceMu.RLock()
for orgID, recorder := range h.incidentRecorders {
if recorder != nil {
out[orgID] = recorder
}
}
h.intelligenceMu.RUnlock()
if _, ok := out["default"]; !ok && h.incidentRecorder != nil {
out["default"] = h.incidentRecorder
}
return out
}
// StopPatrol stops the background AI patrol service
func (h *AISettingsHandler) StopPatrol() {
if h.defaultAIService != nil {
@@ -1760,10 +1669,8 @@ func (h *AISettingsHandler) RemoveTenantIntelligence(orgID string) {
}
var (
bridge *unified.AlertBridge
trigger *ai.TriggerManager
coordinator *ai.IncidentCoordinator
recorder *metrics.IncidentRecorder
bridge *unified.AlertBridge
trigger *ai.TriggerManager
)
h.intelligenceMu.Lock()
@@ -1775,13 +1682,10 @@ func (h *AISettingsHandler) RemoveTenantIntelligence(orgID string) {
delete(h.discoveryStores, orgID)
bridge = h.alertBridges[orgID]
trigger = h.triggerManagers[orgID]
coordinator = h.incidentCoordinators[orgID]
recorder = h.incidentRecorders[orgID]
delete(h.unifiedStores, orgID)
delete(h.alertBridges, orgID)
delete(h.triggerManagers, orgID)
delete(h.incidentCoordinators, orgID)
delete(h.incidentRecorders, orgID)
delete(h.incidentArchives, orgID)
delete(h.proxmoxCorrelators, orgID)
h.intelligenceMu.Unlock()
@@ -1791,12 +1695,6 @@ func (h *AISettingsHandler) RemoveTenantIntelligence(orgID string) {
if trigger != nil {
trigger.Stop()
}
if coordinator != nil {
coordinator.Stop()
}
if recorder != nil {
recorder.Stop()
}
}
// GetAlertTriggeredAnalyzer returns the alert-triggered analyzer for wiring into alert callbacks
@@ -7843,6 +7741,10 @@ func (h *AISettingsHandler) HandleApproveCommand(w http.ResponseWriter, r *http.
// updateFindingOutcome updates the investigation outcome on a finding
func (h *AISettingsHandler) updateFindingOutcome(ctx context.Context, orgID, findingID, outcome string) {
h.updateFindingInvestigationOutcome(ctx, orgID, findingID, outcome, nil)
}
func (h *AISettingsHandler) updateFindingInvestigationOutcome(ctx context.Context, orgID, findingID, outcome string, investigation *ai.InvestigationSession) {
// Get AI service for this org
svc := h.GetAIService(ctx)
if svc == nil {
@@ -7861,15 +7763,20 @@ func (h *AISettingsHandler) updateFindingOutcome(ctx context.Context, orgID, fin
log.Warn().Str("orgID", orgID).Msg("Findings store not available for finding update")
return
}
if existing := findingsStore.Get(findingID); existing != nil && existing.InvestigationOutcome == outcome {
existing := findingsStore.Get(findingID)
if existing == nil {
return
}
if !findingsStore.UpdateInvestigationOutcome(findingID, outcome) {
outcomeChanged := existing.InvestigationOutcome != outcome
if outcomeChanged && !findingsStore.UpdateInvestigationOutcome(findingID, outcome) {
log.Warn().Str("findingID", findingID).Msg("Finding not found for outcome update")
return
}
patrol.PublishFindingLifecycleUpdate(findingID)
recordChanged := investigation != nil && patrol.RefreshFindingInvestigationRecord(findingID, investigation)
if !outcomeChanged && !recordChanged {
return
}
patrol.PublishFindingLifecycleUpdate(findingID, outcomeChanged)
log.Info().Str("findingID", findingID).Str("outcome", outcome).Msg("Updated finding investigation outcome")
}
@@ -258,6 +258,59 @@ func TestPatrolActionReconciliationCannotRegressFromOutOfOrderCallback(t *testin
}
}
func TestPatrolActionHydrationRepairsDurableFindingRecordWithoutRewritingEvidence(t *testing.T) {
investigations := newTestInvestigationStore()
investigation := investigations.Create("finding-1", "session-1")
investigation.Status = aicontracts.InvestigationStatusCompleted
investigation.Outcome = aicontracts.OutcomeFixQueued
investigation.Summary = "Cause uncertain. Restart proposed for review."
investigation.EvidenceIDs = []string{"observed-health"}
investigations.Update(investigation)
svc := ai.NewService(nil, nil)
svc.SetStateProvider(&MockStateProvider{})
patrol := svc.GetPatrolService()
findings := patrol.GetFindings()
findings.Add(&ai.Finding{ID: "finding-1", ResourceID: "vm:42", Title: "Unhealthy service", Severity: ai.FindingSeverityWarning,
InvestigationStatus: string(investigation.Status), InvestigationOutcome: string(aicontracts.OutcomeFixRejected)})
// Reproduce the persisted mismatch after an outcome already reconciled.
record := ai.BuildFindingInvestigationRecord(findings.Get("finding-1"), investigation)
record.Rollback = []string{"Retained rollback evidence"}
record.Impact = "Original service impact absent from the later finding projection."
findings.UpdateInvestigationRecord("finding-1", record)
audits := unifiedresources.NewMemoryStore()
audit := unifiedresources.ActionAuditRecord{
ID: "act-1", CreatedAt: time.Now().UTC(), UpdatedAt: time.Now().UTC(), State: unifiedresources.ActionStateRejected,
Request: unifiedresources.ActionRequest{RequestID: "proposal-1", ResourceID: "vm:42", CapabilityName: "restart", RequestedBy: "pulse_patrol"},
Plan: unifiedresources.ActionPlan{ActionID: "act-1", RequestID: "proposal-1", Allowed: true},
Origin: &unifiedresources.ActionOrigin{Surface: patrolActionOriginSurface, FindingID: "finding-1", InvestigationID: investigation.ID, ProposalID: "proposal-1"},
}
if _, _, err := audits.CreateActionAudit(audit, nil); err != nil {
t.Fatal(err)
}
handler := &AISettingsHandler{defaultAIService: svc,
investigationStores: map[string]aicontracts.InvestigationStore{"default": investigations},
resourceStoreProvider: func(string) (unifiedresources.ResourceStore, error) { return audits, nil },
}
var published []*ai.Finding
patrol.SetUnifiedFindingCallback(func(f *ai.Finding) bool { published = append(published, f); return true })
published = nil // Ignore the initial synchronization when registering the callback.
handler.hydratePatrolInvestigationAction("default", investigation)
got := findings.Get("finding-1").InvestigationRecord
if got.Outcome != aicontracts.OutcomeFixRejected || got.Action == nil || got.Action.State != "rejected" {
t.Fatalf("durable record did not reconcile: %#v", got)
}
if got.Conclusion != investigation.Summary || got.Impact != record.Impact || got.Confidence != record.Confidence || !reflect.DeepEqual(got.Rollback, record.Rollback) || !reflect.DeepEqual(got.Evidence, record.Evidence) {
t.Fatalf("retained investigation evidence changed: %#v", got)
}
if len(got.Verification) != 1 || len(published) != 1 || published[0].InvestigationRecord.Outcome != aicontracts.OutcomeFixRejected {
t.Fatalf("reconciled record was not published: record=%#v published=%#v", got, published)
}
handler.hydratePatrolInvestigationAction("default", investigation)
if len(published) != 1 {
t.Fatal("duplicate hydration republished unchanged record")
}
}
func TestPatrolActionReconciliationHydratesTerminalAuditAfterRestart(t *testing.T) {
investigations := newTestInvestigationStore()
investigation := investigations.Create("finding-1", "session-1")
@@ -50,16 +50,10 @@ func TestAISettingsHandler_SettersAndGetters(t *testing.T) {
t.Fatalf("GetTriggerManager returned unexpected manager")
}
coordinator := ai.NewIncidentCoordinator(ai.IncidentCoordinatorConfig{})
handler.SetIncidentCoordinator(coordinator)
if handler.GetIncidentCoordinator() != coordinator {
t.Fatalf("GetIncidentCoordinator returned unexpected coordinator")
}
recorder := &metrics.IncidentRecorder{}
handler.SetIncidentRecorder(recorder)
if handler.GetIncidentRecorder() != recorder {
t.Fatalf("GetIncidentRecorder returned unexpected recorder")
recorder := &metrics.IncidentArchive{}
handler.SetIncidentArchive(recorder)
if handler.GetIncidentArchive() != recorder {
t.Fatalf("GetIncidentArchive returned unexpected recorder")
}
handler.WireOrchestratorAfterChatStart()
@@ -82,15 +76,19 @@ func TestAISettingsHandler_IntelligenceServicesAreOrgScoped(t *testing.T) {
t.Fatalf("expected nil correlator for unrelated org, got %#v", got)
}
defaultRecorder := &metrics.IncidentRecorder{}
tenantRecorder := &metrics.IncidentRecorder{}
handler.SetIncidentRecorderForOrg("default", defaultRecorder)
handler.SetIncidentRecorderForOrg("acme", tenantRecorder)
if got := handler.GetIncidentRecorder(); got != defaultRecorder {
t.Fatalf("expected default recorder, got %#v", got)
defaultArchive := &metrics.IncidentArchive{}
tenantArchive := &metrics.IncidentArchive{}
handler.SetIncidentArchiveForOrg("default", defaultArchive)
handler.SetIncidentArchiveForOrg("acme", tenantArchive)
if got := handler.GetIncidentArchive(); got != defaultArchive {
t.Fatalf("expected default archive, got %#v", got)
}
if got := handler.GetIncidentRecorderForOrg("acme"); got != tenantRecorder {
t.Fatalf("expected tenant recorder, got %#v", got)
if got := handler.GetIncidentArchiveForOrg("acme"); got != tenantArchive {
t.Fatalf("expected tenant archive, got %#v", got)
}
if got := handler.GetIncidentArchiveForOrg("unrelated"); got != nil {
t.Fatalf("unrelated org received an archive: %#v", got)
}
defaultLearningStore := learning.NewLearningStore(learning.LearningStoreConfig{})
@@ -201,13 +199,12 @@ func TestAISettingsHandler_IntelligenceServicesAreOrgScoped(t *testing.T) {
func TestAISettingsHandler_RemoveTenantService_TrimsOrgID(t *testing.T) {
handler := &AISettingsHandler{
aiServices: map[string]*ai.Service{"acme": nil},
investigationStores: map[string]aicontracts.InvestigationStore{"acme": nil},
proxmoxCorrelators: map[string]*proxmox.EventCorrelator{"acme": nil},
alertBridges: map[string]*unified.AlertBridge{"acme": nil},
triggerManagers: map[string]*ai.TriggerManager{"acme": nil},
incidentCoordinators: map[string]*ai.IncidentCoordinator{"acme": nil},
incidentRecorders: map[string]*metrics.IncidentRecorder{"acme": nil},
aiServices: map[string]*ai.Service{"acme": nil},
investigationStores: map[string]aicontracts.InvestigationStore{"acme": nil},
proxmoxCorrelators: map[string]*proxmox.EventCorrelator{"acme": nil},
alertBridges: map[string]*unified.AlertBridge{"acme": nil},
triggerManagers: map[string]*ai.TriggerManager{"acme": nil},
incidentArchives: map[string]*metrics.IncidentArchive{"acme": nil},
}
handler.RemoveTenantService(" acme ")
@@ -227,10 +224,7 @@ func TestAISettingsHandler_RemoveTenantService_TrimsOrgID(t *testing.T) {
if _, ok := handler.triggerManagers["acme"]; ok {
t.Fatalf("expected trigger manager entry to be removed")
}
if _, ok := handler.incidentCoordinators["acme"]; ok {
t.Fatalf("expected incident coordinator entry to be removed")
}
if _, ok := handler.incidentRecorders["acme"]; ok {
if _, ok := handler.incidentArchives["acme"]; ok {
t.Fatalf("expected incident recorder entry to be removed")
}
}
+5 -9
View File
@@ -3754,14 +3754,10 @@ func TestPatrolModelReadinessBudgetScalesWithRequestTimeout(t *testing.T) {
assert.Equal(t, 4*600*time.Second+time.Minute, patrolModelReadinessBudget(&config.AIConfig{RequestTimeoutSeconds: 600}))
}
// TestOrchestratorAndChatAdaptersMapTheSameMessageFields keeps the deliberate
// GetMessages mirror between orchestratorChatAdapter (ai_handlers.go) and
// chatServiceAdapter (chat_service_adapter.go) honest: both convert the same
// chat-service messages onto separate output contracts, and a field mapped by
// one must be mapped by the other. chatServiceAdapter routes through
// adaptChatMessage so its API-facing tool calls stay on the shared provider
// shape instead of hand-copying a local duplicate.
func TestOrchestratorAndChatAdaptersMapTheSameMessageFields(t *testing.T) {
// The adapters share base message fields but have distinct tool-call contracts.
// Orchestrator turns project provider arguments. Product history retains the
// observed result through the canonical result-bearing transcript type.
func TestOrchestratorAndChatAdaptersMapTheirMessageContracts(t *testing.T) {
for _, tc := range []struct {
file string
fn string
@@ -3793,7 +3789,7 @@ func TestOrchestratorAndChatAdaptersMapTheSameMessageFields(t *testing.T) {
"Content:",
"ReasoningContent:",
"Timestamp:",
"tc.ProviderToolCall()",
"tc.NormalizeCollections()",
"toolResult := *m.ToolResult",
".NormalizeCollections()",
},
+19 -21
View File
@@ -1420,20 +1420,14 @@ func (h *AISettingsHandler) HandleGetRecentIncidents(w http.ResponseWriter, r *h
}
}
// Get coordinator status
coordinator := h.GetIncidentCoordinatorForOrg(GetOrgID(r.Context()))
var activeCount int
if coordinator != nil {
activeCount = coordinator.GetActiveIncidentCount()
}
// Get incident data from patrol service
svc := h.GetAIService(r.Context())
if svc == nil {
if err := utils.WriteJSONResponse(w, map[string]interface{}{
"incidents": []interface{}{},
"active_count": activeCount,
"message": "Pulse Patrol service not available",
"incidents": []interface{}{},
"active_count": nil,
"active_count_status": "not_measured",
"message": "Pulse Patrol service not available",
}); err != nil {
log.Error().Err(err).Msg("Failed to write incidents response")
}
@@ -1443,9 +1437,10 @@ func (h *AISettingsHandler) HandleGetRecentIncidents(w http.ResponseWriter, r *h
patrol := svc.GetPatrolService()
if patrol == nil {
if err := utils.WriteJSONResponse(w, map[string]interface{}{
"incidents": []interface{}{},
"active_count": activeCount,
"message": "Patrol service not available",
"incidents": []interface{}{},
"active_count": nil,
"active_count_status": "not_measured",
"message": "Patrol service not available",
}); err != nil {
log.Error().Err(err).Msg("Failed to write incidents response")
}
@@ -1456,9 +1451,10 @@ func (h *AISettingsHandler) HandleGetRecentIncidents(w http.ResponseWriter, r *h
incidentStore := patrol.GetIncidentStore()
if incidentStore == nil {
if err := utils.WriteJSONResponse(w, map[string]interface{}{
"incidents": []interface{}{},
"active_count": activeCount,
"message": "Incident store not available",
"incidents": []interface{}{},
"active_count": nil,
"active_count_status": "not_measured",
"message": "Incident store not available",
}); err != nil {
log.Error().Err(err).Msg("Failed to write incidents response")
}
@@ -1476,9 +1472,10 @@ func (h *AISettingsHandler) HandleGetRecentIncidents(w http.ResponseWriter, r *h
// This is a limitation - we may want to add ListRecentIncidents to the store
incidentSummary := incidentStore.FormatForPatrol(limit)
if err := utils.WriteJSONResponse(w, map[string]interface{}{
"incidents": []interface{}{},
"incident_summary": incidentSummary,
"active_count": activeCount,
"incidents": []interface{}{},
"incident_summary": incidentSummary,
"active_count": nil,
"active_count_status": "not_measured",
}); err != nil {
log.Error().Err(err).Msg("Failed to write incidents response")
}
@@ -1486,8 +1483,9 @@ func (h *AISettingsHandler) HandleGetRecentIncidents(w http.ResponseWriter, r *h
}
if err := utils.WriteJSONResponse(w, map[string]interface{}{
"incidents": incidents,
"active_count": activeCount,
"incidents": incidents,
"active_count": nil,
"active_count_status": "not_measured",
}); err != nil {
log.Error().Err(err).Msg("Failed to write incidents response")
}
@@ -9,7 +9,6 @@ import (
"testing"
"time"
"github.com/rcourtman/pulse-go-rewrite/internal/ai"
"github.com/rcourtman/pulse-go-rewrite/internal/ai/circuit"
"github.com/rcourtman/pulse-go-rewrite/internal/ai/memory"
"github.com/rcourtman/pulse-go-rewrite/internal/alerts"
@@ -72,11 +71,6 @@ func setupIncidentHandler(t *testing.T) (*AISettingsHandler, *memory.IncidentSto
store := memory.NewIncidentStore(memory.IncidentStoreConfig{DataDir: ""})
patrol.SetIncidentStore(store)
coordinator := ai.NewIncidentCoordinator(ai.IncidentCoordinatorConfig{EnableRecorder: false})
coordinator.SetIncidentStore(store)
coordinator.Start()
handler.SetIncidentCoordinator(coordinator)
alert := &alerts.Alert{
ID: "alert-1",
Type: "cpu",
@@ -86,7 +80,7 @@ func setupIncidentHandler(t *testing.T) (*AISettingsHandler, *memory.IncidentSto
StartTime: time.Now(),
LastSeen: time.Now(),
}
coordinator.OnAlertFired(alert)
store.RecordAlertFired(alert)
return handler, store
}
@@ -110,8 +104,8 @@ func TestHandleGetRecentIncidents(t *testing.T) {
if len(incidents) != 1 {
t.Fatalf("expected 1 incident, got %d", len(incidents))
}
if resp["active_count"].(float64) < 1 {
t.Fatalf("expected active_count >= 1")
if resp["active_count"] != nil || resp["active_count_status"] != "not_measured" {
t.Fatalf("saved incident context must not imply a measured live count: %#v", resp)
}
}
@@ -173,3 +167,30 @@ func TestHandleGetIncidentData(t *testing.T) {
t.Fatalf("expected formatted_context to be populated")
}
}
func TestHandleGetRecentIncidentsCountIsNotMeasured(t *testing.T) {
withMemory, _ := setupIncidentHandler(t)
for _, tc := range []struct {
name string
handler *AISettingsHandler
query string
}{
{"unavailable", &AISettingsHandler{}, ""},
{"fleet context", withMemory, ""},
{"resource context", withMemory, "?resource_id=res-1"},
{"empty resource context", withMemory, "?resource_id=absent"},
} {
t.Run(tc.name, func(t *testing.T) {
rec := httptest.NewRecorder()
tc.handler.HandleGetRecentIncidents(rec, httptest.NewRequest(http.MethodGet, "/api/ai/incidents"+tc.query, nil))
var body map[string]interface{}
if err := json.Unmarshal(rec.Body.Bytes(), &body); err != nil {
t.Fatal(err)
}
count, present := body["active_count"]
if !present || count != nil || body["active_count_status"] != "not_measured" {
t.Fatalf("archive/context presence cannot establish live count: %#v", body)
}
})
}
}
+1 -1
View File
@@ -86,7 +86,7 @@ func adaptChatMessage(m chat.Message) ai.ChatMessage {
Timestamp: m.Timestamp,
}
for _, tc := range m.ToolCalls {
msg.ToolCalls = append(msg.ToolCalls, tc.ProviderToolCall())
msg.ToolCalls = append(msg.ToolCalls, tc.NormalizeCollections())
}
if m.ToolResult != nil {
toolResult := *m.ToolResult
+6 -7
View File
@@ -3,7 +3,6 @@ package api
import (
"context"
"encoding/json"
"strings"
"testing"
"time"
@@ -99,8 +98,8 @@ func TestChatServiceAdapter_GetMessages(t *testing.T) {
require.NoError(t, err)
}
func TestAdaptChatMessageUsesSharedProviderToolCallShape(t *testing.T) {
success := true
func TestAdaptChatMessagePreservesObservedToolResult(t *testing.T) {
success := false
msg := adaptChatMessage(chat.Message{
ID: "msg-1",
Role: "assistant",
@@ -109,7 +108,7 @@ func TestAdaptChatMessageUsesSharedProviderToolCallShape(t *testing.T) {
ToolCalls: []chat.ToolCall{{
ID: "call-1",
Name: "diagnose",
Output: "in-app only",
Output: "NO_AGENT: command agent unavailable",
Success: &success,
ThoughtSignature: json.RawMessage(`{"provider":"gemini"}`),
}},
@@ -121,7 +120,7 @@ func TestAdaptChatMessageUsesSharedProviderToolCallShape(t *testing.T) {
})
require.Len(t, msg.ToolCalls, 1)
var shared agentcapabilities.ProviderToolCall = msg.ToolCalls[0]
var shared agentcapabilities.TranscriptToolCall = msg.ToolCalls[0]
assert.Equal(t, "call-1", shared.ID)
assert.Equal(t, "diagnose", shared.Name)
assert.NotNil(t, shared.Input)
@@ -131,8 +130,8 @@ func TestAdaptChatMessageUsesSharedProviderToolCallShape(t *testing.T) {
text := string(payload)
assert.Contains(t, text, `"input":{}`)
assert.Contains(t, text, `"thought_signature":{"provider":"gemini"}`)
assert.False(t, strings.Contains(text, `"output"`), text)
assert.False(t, strings.Contains(text, `"success"`), text)
assert.Contains(t, text, `"output":"NO_AGENT: command agent unavailable"`)
assert.Contains(t, text, `"success":false`)
require.NotNil(t, msg.ToolResult)
var sharedResult agentcapabilities.ProviderToolResult = *msg.ToolResult
+5 -7
View File
@@ -18650,8 +18650,6 @@ func TestContract_AssistantProviderSeamsDoNotUseMCPTerminology(t *testing.T) {
}
intelligenceAdapters := string(intelligenceAdaptersSource)
for _, fragment := range []string{
`type IncidentRecorderToolAdapter struct`,
`func NewIncidentRecorderToolAdapter(recorder IncidentRecorderSource) *IncidentRecorderToolAdapter`,
`type EventCorrelatorToolAdapter struct`,
`func NewEventCorrelatorToolAdapter(correlator EventCorrelatorSource) *EventCorrelatorToolAdapter`,
} {
@@ -19841,7 +19839,7 @@ func TestContract_PulseMCPAdapterProjectsAgentCapabilitiesManifest(t *testing.T)
chatTypesSrc := string(chatTypesSource)
for _, fragment := range []string{
`func ToolCallFromProvider(tc agentcapabilities.ProviderToolCall) ToolCall`,
`func (t ToolCall) ProviderToolCall() agentcapabilities.ProviderToolCall`,
`type ToolCall = agentcapabilities.TranscriptToolCall`,
`type ToolResult = agentcapabilities.ProviderToolResult`,
} {
if !strings.Contains(chatTypesSrc, fragment) {
@@ -19855,8 +19853,8 @@ func TestContract_PulseMCPAdapterProjectsAgentCapabilitiesManifest(t *testing.T)
}
aiServiceSrc := string(aiServiceSource)
for _, fragment := range []string{
`type ChatToolCall = agentcapabilities.ProviderToolCall`,
`return agentcapabilities.EmptyProviderToolCall()`,
`type ChatToolCall = agentcapabilities.TranscriptToolCall`,
`return ChatToolCall{}.NormalizeCollections()`,
`type ChatToolResult = agentcapabilities.ProviderToolResult`,
`return agentcapabilities.ApprovalRequiredToolMarker(`,
`return agentcapabilities.PolicyBlockedToolMarker(command, reason)`,
@@ -20089,11 +20087,11 @@ func TestContract_PulseMCPAdapterProjectsAgentCapabilitiesManifest(t *testing.T)
chatServiceAdapterSrc := string(chatServiceAdapterSource)
for _, fragment := range []string{
`func adaptChatMessage(m chat.Message) ai.ChatMessage`,
`msg.ToolCalls = append(msg.ToolCalls, tc.ProviderToolCall())`,
`msg.ToolCalls = append(msg.ToolCalls, tc.NormalizeCollections())`,
`toolResult := *m.ToolResult`,
} {
if !strings.Contains(chatServiceAdapterSrc, fragment) {
t.Errorf("chat service adapter must bridge messages through shared provider tool shapes; missing %s", fragment)
t.Errorf("chat service adapter must retain result-bearing transcript calls; missing %s", fragment)
}
}
+1 -1
View File
@@ -105,7 +105,7 @@ func (h *AISettingsHandler) applyPatrolActionAudit(orgID string, audit unifiedre
store.Update(investigation)
}
ctx := context.WithValue(context.Background(), OrgIDContextKey, orgID)
h.updateFindingOutcome(ctx, orgID, origin.FindingID, string(outcome))
h.updateFindingInvestigationOutcome(ctx, orgID, origin.FindingID, string(outcome), investigation)
return
}
if changed {
+7 -134
View File
@@ -2873,41 +2873,8 @@ func (r *Router) initializeAIIntelligenceServices(ctx context.Context, orgID, da
log.Info().Msg("AI Intelligence: Event-driven trigger manager initialized and started")
}
// 12. Initialize incident coordinator for high-frequency recording
if patrol != nil {
incidentCoordinator := ai.NewIncidentCoordinator(ai.DefaultIncidentCoordinatorConfig())
// Wire the incident store if available
if incidentStore := patrol.GetIncidentStore(); incidentStore != nil {
incidentCoordinator.SetIncidentStore(incidentStore)
}
// Create metrics adapter for incident recorder (ReadState is sole source since SRC-03m)
var metricsAdapter *adapters.MetricsAdapter
if monitor != nil {
metricsAdapter = adapters.NewMetricsAdapter(monitor.GetUnifiedReadState())
}
// Initialize and wire the incident recorder (high-frequency metrics)
if metricsAdapter != nil {
recorderCfg := metrics.DefaultIncidentRecorderConfig()
recorderCfg.DataDir = dataDir
recorder := metrics.NewIncidentRecorder(recorderCfg)
recorder.SetMetricsProvider(metricsAdapter)
recorder.Start()
incidentCoordinator.SetRecorder(recorder)
r.aiSettingsHandler.SetIncidentRecorderForOrg(orgID, recorder)
log.Info().Msg("AI Intelligence: Incident recorder initialized and started")
}
// Start the coordinator
incidentCoordinator.Start()
// Store reference
r.aiSettingsHandler.SetIncidentCoordinatorForOrg(orgID, incidentCoordinator)
log.Info().Msg("AI Intelligence: Incident coordinator initialized and started")
}
// Legacy recordings are available only through explicit archive lookup.
r.aiSettingsHandler.SetIncidentArchiveForOrg(orgID, metrics.NewIncidentArchive(dataDir))
log.Info().Msg("AI Intelligence: All Phase 6 & 7 services initialized successfully")
}
@@ -2965,24 +2932,6 @@ func (r *Router) ShutdownAIIntelligence() {
log.Debug().Str("org_id", orgID).Msg("AI Intelligence: Trigger manager stopped")
}
// 4. Stop incident coordinators (stop high-frequency recording)
for orgID, incidentCoordinator := range r.aiSettingsHandler.ListIncidentCoordinators() {
if incidentCoordinator == nil {
continue
}
incidentCoordinator.Stop()
log.Debug().Str("org_id", orgID).Msg("AI Intelligence: Incident coordinator stopped")
}
// 4b. Stop incident recorders (stop background sampling)
for orgID, incidentRecorder := range r.aiSettingsHandler.ListIncidentRecorders() {
if incidentRecorder == nil {
continue
}
incidentRecorder.Stop()
log.Debug().Str("org_id", orgID).Msg("AI Intelligence: Incident recorder stopped")
}
// 5. Cleanup learning stores (removes old records, persists if data dir configured)
for orgID, learningStore := range r.aiSettingsHandler.ListLearningStores() {
if learningStore == nil {
@@ -3353,15 +3302,15 @@ func (r *Router) wireAIChatDependenciesForService(ctx context.Context, service A
}
// Wire intelligence providers for Assistant tools.
// - IncidentRecorderProvider: high-frequency incident data (pulse_get_incident_window)
// - IncidentArchiveProvider: explicit reads of saved legacy recordings
// - EventCorrelatorProvider: Proxmox events (pulse_correlate_events)
// - KnowledgeStoreProvider: notes (pulse_remember, pulse_recall)
// Wire incident recorder provider (high-frequency incident data)
// Wire the org-pinned archive reader without creating or sampling data.
if r.aiSettingsHandler != nil {
if recorder := r.aiSettingsHandler.GetIncidentRecorderForOrg(orgID); recorder != nil {
service.SetIncidentRecorderProvider(&incidentRecorderProviderWrapper{recorder: recorder})
log.Debug().Msg("AI chat: Incident recorder provider wired")
if archive := r.aiSettingsHandler.GetIncidentArchiveForOrg(orgID); archive != nil {
service.SetIncidentArchiveProvider(archive)
log.Debug().Msg("AI chat: Incident archive provider wired")
}
}
@@ -3476,82 +3425,6 @@ func (w *forecastResourceIterator) ForecastStoragePools() []forecast.ResourceInf
return result
}
// incidentRecorderProviderWrapper adapts metrics.IncidentRecorder to tools.IncidentRecorderProvider.
type incidentRecorderProviderWrapper struct {
recorder *metrics.IncidentRecorder
}
func (w *incidentRecorderProviderWrapper) GetWindowsForResource(resourceID string, limit int) []*tools.IncidentWindow {
if w.recorder == nil {
return nil
}
windows := w.recorder.GetWindowsForResource(resourceID, limit)
if len(windows) == 0 {
return nil
}
result := make([]*tools.IncidentWindow, 0, len(windows))
for _, window := range windows {
if window == nil {
continue
}
result = append(result, convertIncidentWindow(window))
}
return result
}
func (w *incidentRecorderProviderWrapper) GetWindow(windowID string) *tools.IncidentWindow {
if w.recorder == nil {
return nil
}
window := w.recorder.GetWindow(windowID)
if window == nil {
return nil
}
return convertIncidentWindow(window)
}
func convertIncidentWindow(window *metrics.IncidentWindow) *tools.IncidentWindow {
if window == nil {
return nil
}
points := make([]tools.IncidentDataPoint, 0, len(window.DataPoints))
for _, point := range window.DataPoints {
points = append(points, tools.IncidentDataPoint{
Timestamp: point.Timestamp,
Metrics: point.Metrics,
})
}
var summary *tools.IncidentSummary
if window.Summary != nil {
summary = &tools.IncidentSummary{
Duration: window.Summary.Duration,
DataPoints: window.Summary.DataPoints,
Peaks: window.Summary.Peaks,
Lows: window.Summary.Lows,
Averages: window.Summary.Averages,
Changes: window.Summary.Changes,
}
}
return &tools.IncidentWindow{
ID: window.ID,
ResourceID: window.ResourceID,
ResourceName: window.ResourceName,
ResourceType: window.ResourceType,
TriggerType: window.TriggerType,
TriggerID: window.TriggerID,
StartTime: window.StartTime,
EndTime: window.EndTime,
Status: string(window.Status),
DataPoints: points,
Summary: summary,
}
}
func (r *Router) publishActionCompletedAgentEvent(broadcaster *AgentEventBroadcaster, record unifiedresources.ActionAuditRecord) {
if broadcaster == nil {
return
+1 -1
View File
@@ -175,7 +175,7 @@ func TestRouterHandleStatePreservesNumericIdleRatesAndOmitsUnknownRates(t *testi
byName[resource.Name] = resource
}
idle := byName["idle-vm"]
if idle.DiskIO == nil || idle.DiskIO.ReadRate != 0 || idle.DiskIO.WriteRate != 0 {
if idle.DiskIO == nil || idle.DiskIO.ReadRate == nil || idle.DiskIO.WriteRate == nil || *idle.DiskIO.ReadRate != 0 || *idle.DiskIO.WriteRate != 0 {
t.Fatalf("valid idle rates were not emitted as numeric zero: %+v", idle.DiskIO)
}
if unknown := byName["unknown-vm"]; unknown.DiskIO != nil {
@@ -8,7 +8,6 @@ import (
"github.com/rcourtman/pulse-go-rewrite/internal/ai/patterns"
"github.com/rcourtman/pulse-go-rewrite/internal/ai/proxmox"
"github.com/rcourtman/pulse-go-rewrite/internal/config"
"github.com/rcourtman/pulse-go-rewrite/internal/metrics"
"github.com/rcourtman/pulse-go-rewrite/internal/models"
"github.com/rcourtman/pulse-go-rewrite/internal/monitoring"
unifiedresources "github.com/rcourtman/pulse-go-rewrite/internal/unifiedresources"
@@ -52,84 +51,6 @@ func TestForecastResourceIterator_NilReadState(t *testing.T) {
}
}
func TestIncidentRecorderProviderWrapper(t *testing.T) {
now := time.Now().UTC()
end := now.Add(5 * time.Minute)
activeWindow := &metrics.IncidentWindow{
ID: "win-active",
ResourceID: "res-1",
ResourceName: "Resource",
ResourceType: "vm",
TriggerType: "alert",
TriggerID: "alert-1",
StartTime: now,
EndTime: &end,
Status: metrics.IncidentWindowStatusRecording,
DataPoints: []metrics.IncidentDataPoint{
{Timestamp: now, Metrics: map[string]float64{"cpu": 10}},
},
Summary: &metrics.IncidentSummary{
Duration: 5 * time.Minute,
DataPoints: 1,
Peaks: map[string]float64{"cpu": 10},
Lows: map[string]float64{"cpu": 5},
Averages: map[string]float64{"cpu": 7},
Changes: map[string]float64{"cpu": 2},
},
}
completedWindow := &metrics.IncidentWindow{
ID: "win-complete",
ResourceID: "res-1",
ResourceName: "Resource",
ResourceType: "vm",
TriggerType: "alert",
TriggerID: "alert-2",
StartTime: now.Add(-time.Hour),
Status: metrics.IncidentWindowStatusComplete,
DataPoints: []metrics.IncidentDataPoint{
{Timestamp: now.Add(-time.Hour), Metrics: map[string]float64{"cpu": 20}},
},
}
recorder := metrics.NewIncidentRecorder(metrics.DefaultIncidentRecorderConfig())
setUnexportedField(t, recorder, "activeWindows", map[string]*metrics.IncidentWindow{"win-active": activeWindow})
setUnexportedField(t, recorder, "completedWindows", []*metrics.IncidentWindow{completedWindow})
wrapper := &incidentRecorderProviderWrapper{recorder: recorder}
windows := wrapper.GetWindowsForResource("res-1", 10)
if len(windows) != 2 {
t.Fatalf("expected 2 windows, got %d", len(windows))
}
ids := []string{windows[0].ID, windows[1].ID}
if !containsStringSlice(ids, "win-active") || !containsStringSlice(ids, "win-complete") {
t.Fatalf("unexpected window ids %v", ids)
}
window := wrapper.GetWindow("win-active")
if window == nil || window.ResourceID != "res-1" || window.Status == "" {
t.Fatalf("unexpected window: %#v", window)
}
}
func TestIncidentRecorderProviderWrapper_NilRecorder(t *testing.T) {
wrapper := &incidentRecorderProviderWrapper{}
if got := wrapper.GetWindowsForResource("res-1", 5); got != nil {
t.Fatalf("expected nil windows, got %#v", got)
}
if got := wrapper.GetWindow("win-1"); got != nil {
t.Fatalf("expected nil window, got %#v", got)
}
}
func TestConvertIncidentWindowNil(t *testing.T) {
if got := convertIncidentWindow(nil); got != nil {
t.Fatalf("expected nil window, got %#v", got)
}
}
func TestEventCorrelatorProviderWrapper(t *testing.T) {
now := time.Now().UTC()
corr := proxmox.EventCorrelation{
@@ -0,0 +1,48 @@
package dockeragent
import (
"encoding/json"
"testing"
containertypes "github.com/moby/moby/api/types/container"
agentsdocker "github.com/rcourtman/pulse-go-rewrite/pkg/agents/docker"
)
func TestSummarizeBlockIOPreservesDirectionPresenceAndZero(t *testing.T) {
for _, tc := range []struct {
name string
entries []containertypes.BlkioStatEntry
read, write bool
}{
{"absent", nil, false, false},
{"unrelated", []containertypes.BlkioStatEntry{{Op: "Total", Value: 100}}, false, false},
{"observed idle", []containertypes.BlkioStatEntry{{Op: "Read", Value: 0}, {Op: "Write", Value: 0}}, true, true},
{"read only idle", []containertypes.BlkioStatEntry{{Op: "Read", Value: 0}}, true, false},
{"write only", []containertypes.BlkioStatEntry{{Op: "Write", Value: 123}}, false, true},
} {
t.Run(tc.name, func(t *testing.T) {
got := summarizeBlockIO(containertypes.StatsResponse{BlkioStats: containertypes.BlkioStats{IoServiceBytesRecursive: tc.entries}})
if !tc.read && !tc.write {
if got != nil {
t.Fatalf("absent counters became observations: %+v", got)
}
return
}
if got == nil {
t.Fatal("explicit counters were dropped")
}
encoded, err := json.Marshal(got)
if err != nil {
t.Fatal(err)
}
var decoded agentsdocker.ContainerBlockIO
if err := json.Unmarshal(encoded, &decoded); err != nil {
t.Fatal(err)
}
read, write := decoded.CounterPresence()
if read != tc.read || write != tc.write {
t.Fatalf("presence lost through report JSON %s: %v/%v", encoded, read, write)
}
})
}
}
+39 -9
View File
@@ -10,6 +10,7 @@ import (
"net/netip"
"net/url"
"regexp"
"sort"
"strconv"
"strings"
"time"
@@ -698,6 +699,36 @@ func (a *Agent) collectContainer(ctx context.Context, summary containertypes.Sum
})
}
}
// Docker's --tmpfs mounts can exist only in HostConfig.Tmpfs. Preserve
// their configuration alongside inspected mounts without inventing usage.
if inspect.HostConfig != nil && len(inspect.HostConfig.Tmpfs) > 0 {
reported := make(map[string]bool, len(mounts))
for _, mount := range mounts {
reported[mount.Destination] = true
}
destinations := make([]string, 0, len(inspect.HostConfig.Tmpfs))
for destination := range inspect.HostConfig.Tmpfs {
if !reported[destination] {
destinations = append(destinations, destination)
}
}
sort.Strings(destinations)
for _, destination := range destinations {
options := inspect.HostConfig.Tmpfs[destination]
writable := true
for _, option := range strings.Split(options, ",") {
switch strings.TrimSpace(option) {
case "ro":
writable = false
case "rw":
writable = true
}
}
mounts = append(mounts, agentsdocker.ContainerMount{
Type: "tmpfs", Destination: destination, Mode: options, RW: writable,
})
}
}
oomKilled := inspect.State.OOMKilled
container := agentsdocker.Container{
@@ -1318,35 +1349,34 @@ func randomDuration(max time.Duration) time.Duration {
}
func summarizeBlockIO(stats containertypes.StatsResponse) *agentsdocker.ContainerBlockIO {
// BlkioStats structure varies by cgroup version
// Cgroup v1: IoServiceBytesRecursive []BlkioStatEntry
// Cgroup v2: IoServiceBytesRecursive is empty? No, Docker maps it?
// Docker API guarantees IoServiceBytesRecursive is populated?
// It seems to try to handle both.
if len(stats.BlkioStats.IoServiceBytesRecursive) == 0 {
return nil
}
var readBytes, writeBytes uint64
var readPresent, writePresent bool
for _, entry := range stats.BlkioStats.IoServiceBytesRecursive {
op := strings.ToLower(entry.Op)
switch op {
case "read":
readPresent = true
readBytes += entry.Value
case "write":
writePresent = true
writeBytes += entry.Value
}
}
if readBytes == 0 && writeBytes == 0 {
if !readPresent && !writePresent {
return nil
}
return &agentsdocker.ContainerBlockIO{
ReadBytes: readBytes,
WriteBytes: writeBytes,
ReadBytes: readBytes,
WriteBytes: writeBytes,
ReadBytesPresent: &readPresent,
WriteBytesPresent: &writePresent,
}
}
@@ -0,0 +1,148 @@
package dockeragent
import (
"context"
"encoding/json"
"os"
"os/exec"
"path/filepath"
"strings"
"testing"
"time"
"github.com/moby/moby/client"
"github.com/rcourtman/pulse-go-rewrite/internal/ai/qualification"
agentsdocker "github.com/rcourtman/pulse-go-rewrite/pkg/agents/docker"
"github.com/rs/zerolog"
)
// This opt-in proof calls the real collector against the existing bounded lab.
// It does not enroll an agent, send reports to Pulse, or invoke a model.
func TestCollectContainerStorageFaultLive(t *testing.T) {
dockerContext := os.Getenv("PULSE_QUALIFY_ORACLE_DOCKER_CONTEXT")
if dockerContext == "" {
t.Skip("set PULSE_QUALIFY_ORACLE_DOCKER_CONTEXT to an explicit disposable Docker context")
}
ctx, cancel := context.WithTimeout(context.Background(), 3*time.Minute)
defer cancel()
endpoint, err := exec.CommandContext(ctx, "docker", "context", "inspect", dockerContext,
"--format", "{{.Endpoints.docker.Host}}").Output()
if err != nil {
t.Fatal(err)
}
host := strings.TrimSpace(string(endpoint))
if !strings.HasPrefix(host, "unix://") {
t.Fatal("run this proof beside the disposable Docker daemon using a Unix socket context")
}
moduleClient, err := newMobyDockerClient(client.WithHost(host), client.WithAPIVersionNegotiation())
if err != nil {
t.Fatal(err)
}
t.Cleanup(func() { _ = moduleClient.Close() })
agent := &Agent{docker: moduleClient, runtime: RuntimeDocker, logger: zerolog.Nop(),
cfg: Config{CollectDiskMetrics: true}, prevContainerCPU: make(map[string]cpuSample)}
manifest, err := qualification.LoadManifest(filepath.Join("..", "..", "tests", "qualification", "patrol",
"scenarios", "investigation.docker-storage-pressure.json"))
if err != nil {
t.Fatal(err)
}
driver := qualification.NewDockerLab(nil, qualification.DockerTarget{Context: dockerContext})
lab, prepareErr := driver.Prepare(ctx, manifest, "collector-"+time.Now().UTC().Format("20060102t150405.000000000"))
if lab != nil {
t.Cleanup(func() {
cleanupCtx, cleanupCancel := context.WithTimeout(context.Background(), time.Minute)
defer cleanupCancel()
result := driver.Cleanup(cleanupCtx, manifest, lab)
encoded, _ := json.Marshal(result)
t.Logf("cleanup=%s", encoded)
if !result.Passed || !result.SecondCleanupNoop || !result.InventoryUnchanged {
t.Errorf("disposable inventory cleanup failed: %+v", result)
}
})
}
if prepareErr != nil {
t.Fatal(prepareErr)
}
check := func(phase, serviceHealth string, predicates []qualification.Predicate) {
t.Helper()
observations, err := driver.Observe(ctx, manifest, lab, predicates)
encoded, _ := json.Marshal(observations)
t.Logf("%s oracle=%s", phase, encoded)
if err != nil {
t.Fatal(err)
}
for _, alias := range []string{"service", "control"} {
id := lab.ResourceIDs[alias]
filters := newDockerFilters()
filters.Add("id", id)
for _, key := range []string{"io.pulse.owner", "io.pulse.profile", "io.pulse.component"} {
value := lab.BaselineStates[alias].Labels[key]
if value == "" {
t.Fatalf("missing fixture ownership label %s", key)
}
filters.Add("label", key+"="+value)
}
containers, err := moduleClient.ContainerList(ctx, dockerContainerListOptions{All: true, Filters: filters})
if err != nil || len(containers) != 1 || containers[0].ID != id {
t.Fatalf("exact owned fixture %s unavailable: count=%d error=%v", alias, len(containers), err)
}
collected, err := agent.collectContainer(ctx, containers[0])
if err != nil {
t.Fatal(err)
}
encoded, err := json.Marshal(collected)
if err != nil {
t.Fatal(err)
}
var report agentsdocker.Container
if err := json.Unmarshal(encoded, &report); err != nil {
t.Fatal(err)
}
wantHealth := "healthy"
if alias == "service" {
wantHealth = serviceHealth
}
if report.ID != id || report.State != "running" || report.Health != wantHealth {
t.Fatalf("%s %s report identity/state/health mismatch: %s", phase, alias, encoded)
}
if alias == "service" {
const destination = "/var/lib/service-cache"
inspect, err := moduleClient.ContainerInspect(ctx, id)
if err != nil || inspect.HostConfig == nil {
t.Fatalf("inspect storage configuration: %v", err)
}
options, ok := inspect.HostConfig.Tmpfs[destination]
if !ok || !strings.Contains(options, "size=8388608") {
t.Fatalf("fixture tmpfs configuration missing: %q", options)
}
matches := 0
for _, mount := range report.Mounts {
if mount.Destination == destination {
matches++
if mount.Type != "tmpfs" || !mount.RW || mount.Mode != options {
t.Fatalf("collector changed tmpfs configuration: %+v", mount)
}
}
}
if matches != 1 {
t.Fatalf("expected one storage mount in report, got %d", matches)
}
t.Logf("%s raw_mount_count=%d native_tmpfs_options=%q", phase, len(inspect.Mounts), options)
}
t.Logf("%s %s report=%s", phase, alias, encoded)
}
}
check("baseline", "healthy", manifest.Baseline)
for _, fault := range manifest.Faults {
if err := driver.ApplyFault(ctx, manifest, lab, fault); err != nil {
t.Fatal(err)
}
check("storage-full", "unhealthy", fault.Oracle)
if err := driver.RevertFault(ctx, manifest, lab, fault); err != nil {
t.Fatal(err)
}
check("recovered", "healthy", fault.RevertOracle)
}
check("restored-baseline", "healthy", manifest.Baseline)
}
@@ -0,0 +1,84 @@
package dockeragent
import (
"context"
"encoding/json"
"reflect"
"testing"
containertypes "github.com/moby/moby/api/types/container"
agentsdocker "github.com/rcourtman/pulse-go-rewrite/pkg/agents/docker"
"github.com/rs/zerolog"
)
func TestCollectContainerPreservesTmpfsMounts(t *testing.T) {
for _, tc := range []struct {
name string
host *containertypes.HostConfig
mounts []containertypes.MountPoint
want []agentsdocker.ContainerMount
}{
{
name: "native tmpfs only",
host: &containertypes.HostConfig{Tmpfs: map[string]string{
"/var/lib/service-cache": "rw,noexec,nosuid,nodev,size=8388608",
}},
want: []agentsdocker.ContainerMount{{Type: "tmpfs", Destination: "/var/lib/service-cache", Mode: "rw,noexec,nosuid,nodev,size=8388608", RW: true}},
},
{
name: "mixed mounts and stable tmpfs order",
host: &containertypes.HostConfig{Tmpfs: map[string]string{"/z-cache": "", "/a-cache": "ro,noexec"}},
mounts: []containertypes.MountPoint{{Type: "bind", Source: "/host-data", Destination: "/data", RW: true}},
want: []agentsdocker.ContainerMount{
{Type: "bind", Source: "/host-data", Destination: "/data", RW: true},
{Type: "tmpfs", Destination: "/a-cache", Mode: "ro,noexec", RW: false},
{Type: "tmpfs", Destination: "/z-cache", RW: true},
},
},
{
name: "reported mount is authoritative",
host: &containertypes.HostConfig{Tmpfs: map[string]string{"/cache": "rw,size=8388608"}},
mounts: []containertypes.MountPoint{{Type: "tmpfs", Destination: "/cache", Mode: "ro", RW: false}},
want: []agentsdocker.ContainerMount{{Type: "tmpfs", Destination: "/cache", Mode: "ro", RW: false}},
},
{name: "absent host config"},
} {
t.Run(tc.name, func(t *testing.T) {
inspect := baseInspect()
inspect.HostConfig = tc.host
inspect.Mounts = tc.mounts
a := &Agent{
logger: zerolog.Nop(),
runtime: RuntimeDocker,
prevContainerCPU: make(map[string]cpuSample),
docker: &fakeDockerClient{
containerInspectWithRawFn: func(context.Context, string, bool) (containertypes.InspectResponse, []byte, error) {
return inspect, nil, nil
},
containerStatsOneShotFn: func(context.Context, string) (dockerStatsResponseReader, error) {
return statsReader(t, containertypes.StatsResponse{}), nil
},
},
}
got, err := a.collectContainer(context.Background(), containertypes.Summary{ID: "owned-storage-container", Names: []string{"/worker"}, Image: "alpine:3.20", State: "running"})
if err != nil {
t.Fatal(err)
}
if !reflect.DeepEqual(got.Mounts, tc.want) {
t.Fatalf("mount inventory = %#v, want %#v", got.Mounts, tc.want)
}
encoded, err := json.Marshal(got)
if err != nil {
t.Fatal(err)
}
var wire agentsdocker.Container
if err := json.Unmarshal(encoded, &wire); err != nil {
t.Fatal(err)
}
if !reflect.DeepEqual(wire.Mounts, tc.want) {
t.Fatalf("report lost mounts: %#v", wire.Mounts)
}
t.Logf("TMPFS_COLLECTOR_REPORT %s", encoded)
})
}
}
+170
View File
@@ -0,0 +1,170 @@
// Package metrics preserves access to legacy incident recording archives.
package metrics
import (
"encoding/json"
"errors"
"fmt"
"io"
"os"
"path/filepath"
"strings"
"time"
)
// IncidentWindow represents a saved legacy recording window. Status and sample timestamps are historical
type IncidentWindow struct {
ID string `json:"id"`
ResourceID string `json:"resource_id"`
ResourceName string `json:"resource_name,omitempty"`
ResourceType string `json:"resource_type,omitempty"`
TriggerType string `json:"trigger_type"` // "alert", "anomaly", "focus", "manual"
TriggerID string `json:"trigger_id,omitempty"`
StartTime time.Time `json:"start_time"`
EndTime *time.Time `json:"end_time,omitempty"`
Status IncidentWindowStatus `json:"status"`
DataPoints []IncidentDataPoint `json:"data_points"`
Summary *IncidentSummary `json:"summary,omitempty"`
}
// IncidentWindowStatus represents the status of an incident window
type IncidentWindowStatus string
const (
IncidentWindowStatusRecording IncidentWindowStatus = "recording"
IncidentWindowStatusComplete IncidentWindowStatus = "complete"
IncidentWindowStatusTruncated IncidentWindowStatus = "truncated" // Stopped due to limits
maxIncidentWindowsFileSize = 16 << 20 // 16 MiB
)
var errUnsafeIncidentArchivePath = errors.New("unsafe incident archive path")
// IncidentDataPoint represents a single data point in an incident window
type IncidentDataPoint struct {
Timestamp time.Time `json:"timestamp"`
Metrics map[string]float64 `json:"metrics"` // cpu, memory, disk, etc.
Metadata map[string]interface{} `json:"metadata,omitempty"`
}
// IncidentSummary provides computed statistics about an incident window
type IncidentSummary struct {
Duration time.Duration `json:"duration_ms"`
DataPoints int `json:"data_points"`
Peaks map[string]float64 `json:"peaks"` // Maximum values
Lows map[string]float64 `json:"lows"` // Minimum values
Averages map[string]float64 `json:"averages"` // Average values
Changes map[string]float64 `json:"changes"` // Change from start to end
Anomalies []string `json:"anomalies,omitempty"` // Detected anomalies
}
// IncidentArchive reads saved recordings on explicit request. It never starts
// collectors or rewrites, expires, creates or changes permissions on archives.
type IncidentArchive struct{ filePath string }
var ErrIncidentArchiveUnavailable = errors.New("legacy incident recording archive is unavailable")
func NewIncidentArchive(dataDir string) *IncidentArchive {
if strings.TrimSpace(dataDir) == "" {
return &IncidentArchive{}
}
return &IncidentArchive{filePath: filepath.Join(dataDir, "incident_windows.json")}
}
// GetWindow requires the exact resource and window identifiers in this org's
// archive. Historical names and aliases cannot authorize an archive lookup.
func (a *IncidentArchive) GetWindow(resourceID, windowID string) (*IncidentWindow, error) {
if a == nil || a.filePath == "" {
return nil, ErrIncidentArchiveUnavailable
}
if strings.TrimSpace(resourceID) == "" || strings.TrimSpace(windowID) == "" {
return nil, errors.New("resource_id and window_id are required")
}
data, err := readBoundedRegularFile(a.filePath, maxIncidentWindowsFileSize)
if errors.Is(err, os.ErrNotExist) {
return nil, ErrIncidentArchiveUnavailable
}
if err != nil {
return nil, fmt.Errorf("read legacy incident archive: %w", err)
}
var saved struct {
CompletedWindows json.RawMessage `json:"completed_windows"`
}
if err := json.Unmarshal(data, &saved); err != nil {
return nil, fmt.Errorf("decode legacy incident archive: %w", err)
}
if len(saved.CompletedWindows) == 0 {
return nil, errors.New("legacy incident archive has no completed_windows field")
}
var windows []*IncidentWindow
if err := json.Unmarshal(saved.CompletedWindows, &windows); err != nil {
return nil, fmt.Errorf("decode legacy incident windows: %w", err)
}
var match *IncidentWindow
for _, window := range windows {
if window == nil || window.ID != windowID || window.ResourceID != resourceID {
continue
}
if match != nil {
return nil, errors.New("legacy incident archive contains duplicate resource/window identifiers")
}
match = window
}
return match, nil
}
func validateRegularFilePath(path string, info os.FileInfo) error {
if info.Mode()&os.ModeSymlink != 0 {
return fmt.Errorf("%w: refusing symlink path %q", errUnsafeIncidentArchivePath, path)
}
if !info.Mode().IsRegular() {
return fmt.Errorf("%w: non-regular path %q", errUnsafeIncidentArchivePath, path)
}
return nil
}
func readBoundedRegularFile(path string, maxSize int64) ([]byte, error) {
initialInfo, err := os.Lstat(path)
if err != nil {
return nil, err
}
if err := validateRegularFilePath(path, initialInfo); err != nil {
return nil, err
}
if maxSize > 0 && initialInfo.Size() > maxSize {
return nil, fmt.Errorf("%w: file %q exceeds size limit (%d bytes)", errUnsafeIncidentArchivePath, path, initialInfo.Size())
}
file, err := os.Open(path)
if err != nil {
return nil, err
}
defer func() {
_ = file.Close()
}()
openInfo, err := file.Stat()
if err != nil {
return nil, err
}
if err := validateRegularFilePath(path, openInfo); err != nil {
return nil, err
}
if !os.SameFile(initialInfo, openInfo) {
return nil, fmt.Errorf("%w: file %q changed during read", errUnsafeIncidentArchivePath, path)
}
reader := io.Reader(file)
if maxSize > 0 {
reader = io.LimitReader(file, maxSize+1)
}
data, err := io.ReadAll(reader)
if err != nil {
return nil, err
}
if maxSize > 0 && int64(len(data)) > maxSize {
return nil, fmt.Errorf("%w: file %q exceeded size limit while reading", errUnsafeIncidentArchivePath, path)
}
return data, nil
}
+117
View File
@@ -0,0 +1,117 @@
package metrics
import (
"encoding/json"
"os"
"path/filepath"
"testing"
"time"
"github.com/stretchr/testify/require"
)
func TestIncidentArchivePreservesHistoricalFileAndValues(t *testing.T) {
dir := t.TempDir()
path := filepath.Join(dir, "incident_windows.json")
// Older than the former retention window, with data that the old adapter dropped.
original := []byte(`{"completed_windows":[null,{"id":"old-window","resource_id":"docker:old-host/full-id","resource_type":"host","status":"recording","start_time":"2020-01-02T03:04:05Z","data_points":[{"timestamp":"2020-01-02T03:04:06Z","metrics":{"cpu":12.5},"metadata":{"source":"cached","nested":{"retained":true}}}],"summary":{"duration_ms":60000000000,"data_points":1,"anomalies":["stored observation"]}}]}`)
require.NoError(t, os.WriteFile(path, original, 0640))
require.NoError(t, os.Chmod(path, 0640))
oldTime := time.Date(2020, 1, 2, 3, 4, 5, 0, time.UTC)
require.NoError(t, os.Chtimes(path, oldTime, oldTime))
before, err := os.Stat(path)
require.NoError(t, err)
archive := NewIncidentArchive(dir)
window, err := archive.GetWindow("docker:old-host/full-id", "old-window")
require.NoError(t, err)
require.NotNil(t, window)
require.Equal(t, IncidentWindowStatusRecording, window.Status)
require.Equal(t, "host", window.ResourceType)
require.Equal(t, oldTime, window.StartTime)
require.Equal(t, time.Minute, window.Summary.Duration)
require.Equal(t, []string{"stored observation"}, window.Summary.Anomalies)
require.Equal(t, map[string]interface{}{"retained": true}, window.DataPoints[0].Metadata["nested"])
// Each read decodes independently, so a caller cannot alter later evidence.
window.DataPoints[0].Metrics["cpu"] = 99
again, err := archive.GetWindow("docker:old-host/full-id", "old-window")
require.NoError(t, err)
require.Equal(t, 12.5, again.DataPoints[0].Metrics["cpu"])
for _, key := range [][2]string{{"other-resource", "old-window"}, {"docker:old-host/full-id", "absent"}} {
got, err := archive.GetWindow(key[0], key[1])
require.NoError(t, err)
require.Nil(t, got)
}
after, err := os.Stat(path)
require.NoError(t, err)
got, err := os.ReadFile(path)
require.NoError(t, err)
require.Equal(t, original, got)
require.Equal(t, before.Mode(), after.Mode())
require.Equal(t, before.ModTime(), after.ModTime())
}
func TestIncidentArchiveReadsOnlyOnRequestAndReportsFailures(t *testing.T) {
dir := t.TempDir()
path := filepath.Join(dir, "incident_windows.json")
archive := NewIncidentArchive(dir)
_, err := os.Stat(path)
require.True(t, os.IsNotExist(err))
_, err = archive.GetWindow("r", "w")
require.ErrorIs(t, err, ErrIncidentArchiveUnavailable)
// A repaired or newly restored file is visible on the next explicit read.
for _, raw := range []string{`{"completed_windows":[]}`, `{"completed_windows":null}`} {
require.NoError(t, os.WriteFile(path, []byte(raw), 0600))
got, err := archive.GetWindow("r", "w")
require.NoError(t, err)
require.Nil(t, got)
}
for _, raw := range []string{`{`, `{}`, `null`, `{"completed_windows":{}}`, `{"completed_windows":[{"id":"w","resource_id":"r"},{"id":"w","resource_id":"r"}]}`} {
require.NoError(t, os.WriteFile(path, []byte(raw), 0600))
got, err := archive.GetWindow("r", "w")
require.Error(t, err)
require.Nil(t, got)
}
require.NoError(t, os.Remove(path))
require.NoError(t, os.Mkdir(path, 0700))
_, err = archive.GetWindow("r", "w")
require.ErrorIs(t, err, errUnsafeIncidentArchivePath)
require.NoError(t, os.Remove(path))
target := filepath.Join(dir, "other.json")
require.NoError(t, os.WriteFile(target, []byte(`{"completed_windows":[]}`), 0600))
require.NoError(t, os.Symlink(target, path))
_, err = archive.GetWindow("r", "w")
require.ErrorIs(t, err, errUnsafeIncidentArchivePath)
require.NoError(t, os.Remove(path))
f, err := os.Create(path)
require.NoError(t, err)
require.NoError(t, f.Truncate(maxIncidentWindowsFileSize+1))
require.NoError(t, f.Close())
_, err = archive.GetWindow("r", "w")
require.ErrorIs(t, err, errUnsafeIncidentArchivePath)
for _, a := range []*IncidentArchive{nil, NewIncidentArchive("")} {
_, err = a.GetWindow("r", "w")
require.ErrorIs(t, err, ErrIncidentArchiveUnavailable)
}
}
func TestIncidentArchiveRequiresExactResourceAndOrg(t *testing.T) {
aDir, bDir := t.TempDir(), t.TempDir()
for _, org := range []struct{ dir, name string }{{aDir, "tenant-a"}, {bDir, "tenant-b"}} {
raw, err := json.Marshal(map[string]interface{}{"completed_windows": []*IncidentWindow{{ID: "same-window", ResourceID: "same-resource", ResourceName: org.name}}})
require.NoError(t, err)
require.NoError(t, os.WriteFile(filepath.Join(org.dir, "incident_windows.json"), raw, 0600))
}
for _, org := range []struct{ dir, name string }{{aDir, "tenant-a"}, {bDir, "tenant-b"}} {
a := NewIncidentArchive(org.dir)
got, err := a.GetWindow("same-resource", "same-window")
require.NoError(t, err)
require.Equal(t, org.name, got.ResourceName)
for _, keys := range [][2]string{{"same-resource ", "same-window"}, {"same-resource", "same-window "}} {
got, err = a.GetWindow(keys[0], keys[1])
require.NoError(t, err)
require.Nil(t, got)
}
_, err = a.GetWindow("", "same-window")
require.Error(t, err)
}
}
-948
View File
@@ -1,948 +0,0 @@
// Package metrics provides metrics collection and incident recording functionality.
package metrics
import (
"encoding/json"
"errors"
"fmt"
"io"
"os"
"path/filepath"
"strings"
"sync"
"sync/atomic"
"time"
"github.com/rcourtman/pulse-go-rewrite/internal/unifiedresources"
"github.com/rs/zerolog/log"
)
// IncidentWindow represents a high-frequency recording window during an incident
type IncidentWindow struct {
ID string `json:"id"`
ResourceID string `json:"resource_id"`
ResourceName string `json:"resource_name,omitempty"`
ResourceType string `json:"resource_type,omitempty"`
TriggerType string `json:"trigger_type"` // "alert", "anomaly", "focus", "manual"
TriggerID string `json:"trigger_id,omitempty"`
StartTime time.Time `json:"start_time"`
EndTime *time.Time `json:"end_time,omitempty"`
Status IncidentWindowStatus `json:"status"`
DataPoints []IncidentDataPoint `json:"data_points"`
Summary *IncidentSummary `json:"summary,omitempty"`
}
// IncidentWindowStatus represents the status of an incident window
type IncidentWindowStatus string
const (
IncidentWindowStatusRecording IncidentWindowStatus = "recording"
IncidentWindowStatusComplete IncidentWindowStatus = "complete"
IncidentWindowStatusTruncated IncidentWindowStatus = "truncated" // Stopped due to limits
incidentRecorderDirPerm = 0o700
incidentRecorderFilePerm = 0o600
maxIncidentWindowsFileSize = 16 << 20 // 16 MiB
maxWindowIDResourceSegment = 64
unknownWindowResourceSegment = "unknown"
)
var errUnsafeIncidentPersistencePath = errors.New("unsafe incident recorder persistence path")
// IncidentDataPoint represents a single data point in an incident window
type IncidentDataPoint struct {
Timestamp time.Time `json:"timestamp"`
Metrics map[string]float64 `json:"metrics"` // cpu, memory, disk, etc.
Metadata map[string]interface{} `json:"metadata,omitempty"`
}
// IncidentSummary provides computed statistics about an incident window
type IncidentSummary struct {
Duration time.Duration `json:"duration_ms"`
DataPoints int `json:"data_points"`
Peaks map[string]float64 `json:"peaks"` // Maximum values
Lows map[string]float64 `json:"lows"` // Minimum values
Averages map[string]float64 `json:"averages"` // Average values
Changes map[string]float64 `json:"changes"` // Change from start to end
Anomalies []string `json:"anomalies,omitempty"` // Detected anomalies
}
// IncidentRecorderConfig configures the incident recorder
type IncidentRecorderConfig struct {
// Recording settings
SampleInterval time.Duration // How often to record data points (default: 5s)
PreIncidentWindow time.Duration // How much data to capture before incident (default: 5min)
PostIncidentWindow time.Duration // How much data to capture after incident (default: 10min)
MaxDataPointsPerWindow int // Maximum data points per window (default: 500)
// Storage settings
DataDir string
MaxWindows int // Maximum number of windows to keep (default: 100)
RetentionDuration time.Duration // How long to keep windows (default: 24h)
}
// DefaultIncidentRecorderConfig returns sensible defaults
func DefaultIncidentRecorderConfig() IncidentRecorderConfig {
return IncidentRecorderConfig{
SampleInterval: 5 * time.Second,
PreIncidentWindow: 5 * time.Minute,
PostIncidentWindow: 10 * time.Minute,
MaxDataPointsPerWindow: 500,
MaxWindows: 100,
RetentionDuration: 24 * time.Hour,
}
}
// MetricsProvider provides current metrics for a resource
type MetricsProvider interface {
GetCurrentMetrics(resourceID string) (map[string]float64, error)
GetMonitoredResourceIDs() []string // Returns all resource IDs being monitored
}
// BatchMetricsProvider is an optional MetricsProvider extension. recordSample
// asks for metrics once per monitored resource every tick; a provider whose
// per-ID lookup scans all resources turns that into an O(n^2) tick, so when
// this is implemented the recorder fetches the whole batch in one pass.
type BatchMetricsProvider interface {
GetCurrentMetricsBatch() map[string]map[string]float64
}
// IncidentRecorder captures high-frequency metrics during incidents
type IncidentRecorder struct {
mu sync.RWMutex
config IncidentRecorderConfig
provider MetricsProvider
// Active recordings
activeWindows map[string]*IncidentWindow // keyed by window ID
// Completed recordings (ring buffer)
completedWindows []*IncidentWindow
// Background recording for pre-incident buffer
preIncidentBuffer map[string][]IncidentDataPoint // keyed by resource ID
// Persistence
dataDir string
filePath string
saveIOMu sync.Mutex
// Control
stopCh chan struct{}
loopDone chan struct{}
running bool
// Async save coordination
saveMu sync.Mutex
saveCond *sync.Cond
saveInProgress bool
saveRequested bool
}
// NewIncidentRecorder creates a new incident recorder
func NewIncidentRecorder(cfg IncidentRecorderConfig) *IncidentRecorder {
if cfg.SampleInterval <= 0 {
cfg.SampleInterval = 5 * time.Second
}
if cfg.PreIncidentWindow <= 0 {
cfg.PreIncidentWindow = 5 * time.Minute
}
if cfg.PostIncidentWindow <= 0 {
cfg.PostIncidentWindow = 10 * time.Minute
}
if cfg.MaxDataPointsPerWindow <= 0 {
cfg.MaxDataPointsPerWindow = 500
}
if cfg.MaxWindows <= 0 {
cfg.MaxWindows = 100
}
if cfg.RetentionDuration <= 0 {
cfg.RetentionDuration = 24 * time.Hour
}
if cfg.DataDir != "" {
trimmed := strings.TrimSpace(cfg.DataDir)
if trimmed == "" {
log.Warn().Msg("Ignoring incident recorder data dir: blank after trimming whitespace")
cfg.DataDir = ""
} else {
cfg.DataDir = filepath.Clean(trimmed)
}
}
dataDir := strings.TrimSpace(cfg.DataDir)
if dataDir != "" {
dataDir = filepath.Clean(dataDir)
}
recorder := &IncidentRecorder{
config: cfg,
activeWindows: make(map[string]*IncidentWindow),
completedWindows: make([]*IncidentWindow, 0),
preIncidentBuffer: make(map[string][]IncidentDataPoint),
dataDir: dataDir,
stopCh: make(chan struct{}),
loopDone: make(chan struct{}),
}
recorder.saveCond = sync.NewCond(&recorder.saveMu)
if dataDir != "" {
recorder.filePath = filepath.Join(dataDir, "incident_windows.json")
if err := recorder.loadFromDisk(); err != nil {
log.Warn().
Str("file_path", recorder.filePath).
Err(err).
Msg("Failed to load incident windows from disk")
}
}
return recorder
}
// SetMetricsProvider sets the metrics provider for recording
func (r *IncidentRecorder) SetMetricsProvider(provider MetricsProvider) {
r.mu.Lock()
defer r.mu.Unlock()
r.provider = provider
}
// Start begins background recording for pre-incident buffer
func (r *IncidentRecorder) Start() {
r.mu.Lock()
if r.running {
r.mu.Unlock()
return
}
stopCh := make(chan struct{})
loopDone := make(chan struct{})
r.running = true
r.stopCh = stopCh
r.loopDone = loopDone
r.mu.Unlock()
go r.recordingLoop(stopCh, loopDone)
log.Info().
Dur("sample_interval", r.config.SampleInterval).
Dur("pre_incident_window", r.config.PreIncidentWindow).
Dur("post_incident_window", r.config.PostIncidentWindow).
Int("max_data_points_per_window", r.config.MaxDataPointsPerWindow).
Msg("Incident recorder started")
}
// Stop stops the incident recorder
func (r *IncidentRecorder) Stop() {
r.mu.Lock()
if !r.running {
r.mu.Unlock()
r.waitForPendingSaves()
return
}
r.running = false
close(r.stopCh)
loopDone := r.loopDone
r.mu.Unlock()
if loopDone != nil {
<-loopDone
}
// Flush any async save goroutine triggered before shutdown so TempDir cleanup
// and final persistence do not race a background rename/write.
r.waitForPendingSaves()
// Save to disk
if err := r.saveToDisk(); err != nil {
log.Warn().
Str("file_path", r.filePath).
Err(err).
Msg("Failed to save incident windows on stop")
}
log.Info().Msg("incident recorder stopped")
}
// recordingLoop runs in the background to maintain pre-incident buffers and active windows
func (r *IncidentRecorder) recordingLoop(stopCh <-chan struct{}, done chan<- struct{}) {
defer close(done)
ticker := time.NewTicker(r.config.SampleInterval)
defer ticker.Stop()
for {
select {
case <-stopCh:
return
case <-ticker.C:
r.recordSample()
}
}
}
// recordSample captures a data point for all active windows and buffers
func (r *IncidentRecorder) recordSample() {
r.mu.Lock()
if r.provider == nil {
r.mu.Unlock()
return
}
now := time.Now()
shouldSave := false
// One pass over the provider when it supports batching; the per-ID
// fallback preserves behavior for providers that do not. A missing batch
// entry maps to the empty metrics the per-ID method returns for unknown
// IDs.
var metricsBatch map[string]map[string]float64
if batchProvider, ok := r.provider.(BatchMetricsProvider); ok {
metricsBatch = batchProvider.GetCurrentMetricsBatch()
}
currentMetrics := func(resourceID string) (map[string]float64, error) {
if metricsBatch != nil {
if metrics, ok := metricsBatch[resourceID]; ok {
return metrics, nil
}
}
// A batch that lacks the ID must not change per-provider semantics:
// fall through so providers that error or synthesize for unknown IDs
// keep doing exactly that. Misses are the rare path, so the batch
// still absorbs the per-tick fan-out.
return r.provider.GetCurrentMetrics(resourceID)
}
// Record for active windows
for _, window := range r.activeWindows {
if window.Status != IncidentWindowStatusRecording {
continue
}
// Check if we've exceeded the post-incident window
if window.EndTime != nil && now.After(*window.EndTime) {
r.completeWindowLocked(window)
shouldSave = true
continue
}
// Check if we've exceeded max data points
if len(window.DataPoints) >= r.config.MaxDataPointsPerWindow {
window.Status = IncidentWindowStatusTruncated
log.Warn().
Str("window_id", window.ID).
Str("resource_id", window.ResourceID).
Int("max_data_points_per_window", r.config.MaxDataPointsPerWindow).
Msg("Truncating incident window after reaching max data points")
r.completeWindowLocked(window)
continue
}
// Get metrics
metrics, err := currentMetrics(window.ResourceID)
if err != nil {
log.Debug().
Str("window_id", window.ID).
Str("resource_id", window.ResourceID).
Err(err).
Msg("failed to get metrics for incident window")
continue
}
window.DataPoints = append(window.DataPoints, IncidentDataPoint{
Timestamp: now,
Metrics: copyMetrics(metrics),
})
}
// Continuously buffer ALL monitored resources for pre-incident data
// This ensures we have history when an alert fires on any resource
monitoredResources := r.provider.GetMonitoredResourceIDs()
bufferCutoff := now.Add(-r.config.PreIncidentWindow)
for _, resourceID := range monitoredResources {
metrics, err := currentMetrics(resourceID)
if err != nil {
log.Debug().
Str("resource_id", resourceID).
Err(err).
Msg("Failed to get metrics for pre-incident buffer")
continue
}
// Add to pre-incident buffer
buffer := r.preIncidentBuffer[resourceID]
buffer = append(buffer, IncidentDataPoint{
Timestamp: now,
Metrics: copyMetrics(metrics),
})
// Keep only last PreIncidentWindow duration
kept := make([]IncidentDataPoint, 0, len(buffer))
for _, dp := range buffer {
if dp.Timestamp.After(bufferCutoff) {
kept = append(kept, dp)
}
}
r.preIncidentBuffer[resourceID] = kept
}
// Clean up buffers for resources no longer monitored
monitoredSet := make(map[string]bool, len(monitoredResources))
for _, resourceID := range monitoredResources {
monitoredSet[resourceID] = true
}
for resourceID := range r.preIncidentBuffer {
if !monitoredSet[resourceID] {
delete(r.preIncidentBuffer, resourceID)
}
}
r.mu.Unlock()
if shouldSave {
if err := r.saveToDisk(); err != nil {
log.Warn().Err(err).Msg("Failed to save incident windows")
}
}
}
// StartRecording begins recording an incident window
func (r *IncidentRecorder) StartRecording(resourceID, resourceName, resourceType, triggerType, triggerID string) string {
r.mu.Lock()
defer r.mu.Unlock()
// Check if we already have an active window for this resource
for _, window := range r.activeWindows {
if window.ResourceID == resourceID && window.Status == IncidentWindowStatusRecording {
// Extend existing window
endTime := time.Now().Add(r.config.PostIncidentWindow)
window.EndTime = &endTime
return window.ID
}
}
// Create new window
windowID := generateWindowID(resourceID)
now := time.Now()
endTime := now.Add(r.config.PostIncidentWindow)
normalizedResourceType := normalizeIncidentResourceType(resourceType)
window := &IncidentWindow{
ID: windowID,
ResourceID: resourceID,
ResourceName: resourceName,
ResourceType: normalizedResourceType,
TriggerType: triggerType,
TriggerID: triggerID,
StartTime: now.Add(-r.config.PreIncidentWindow), // Include pre-incident data
EndTime: &endTime,
Status: IncidentWindowStatusRecording,
DataPoints: make([]IncidentDataPoint, 0),
}
// Copy pre-incident buffer if available
if preBuffer, ok := r.preIncidentBuffer[resourceID]; ok {
window.DataPoints = append(window.DataPoints, copyDataPoints(preBuffer)...)
}
r.activeWindows[windowID] = window
log.Info().
Str("window_id", windowID).
Str("resource_id", resourceID).
Str("trigger_type", triggerType).
Msg("started incident recording")
return windowID
}
func normalizeIncidentResourceType(resourceType string) string {
normalized := strings.ToLower(strings.TrimSpace(resourceType))
if canonical, ok := unifiedresources.CanonicalizeLegacyResourceTypeAlias(normalized); ok {
return canonical
}
return normalized
}
// StopRecording stops recording for a specific window
func (r *IncidentRecorder) StopRecording(windowID string) {
r.mu.Lock()
shouldSave := false
if window, ok := r.activeWindows[windowID]; ok {
r.completeWindowLocked(window)
shouldSave = true
}
r.mu.Unlock()
if shouldSave {
if err := r.saveToDisk(); err != nil {
log.Warn().Err(err).Msg("Failed to save incident windows")
}
}
}
// completeWindowLocked finalizes a recording window.
// Caller must hold r.mu.
func (r *IncidentRecorder) completeWindowLocked(window *IncidentWindow) {
if window.Status != IncidentWindowStatusRecording && window.Status != IncidentWindowStatusTruncated {
return
}
now := time.Now()
if window.Status == IncidentWindowStatusRecording {
window.Status = IncidentWindowStatusComplete
}
window.EndTime = &now
// Compute summary
window.Summary = r.computeSummary(window)
// Move to completed
r.completedWindows = append(r.completedWindows, window)
delete(r.activeWindows, window.ID)
// Trim completed windows
r.trimCompletedWindows()
log.Info().
Str("window_id", window.ID).
Str("resource_id", window.ResourceID).
Str("status", string(window.Status)).
Int("data_points", len(window.DataPoints)).
Msg("completed incident recording")
// Save asynchronously.
r.requestAsyncSave()
}
// computeSummary computes statistics for a window
func (r *IncidentRecorder) computeSummary(window *IncidentWindow) *IncidentSummary {
if len(window.DataPoints) == 0 {
return nil
}
summary := &IncidentSummary{
DataPoints: len(window.DataPoints),
Peaks: make(map[string]float64),
Lows: make(map[string]float64),
Averages: make(map[string]float64),
Changes: make(map[string]float64),
}
// Calculate duration
if len(window.DataPoints) > 1 {
first := window.DataPoints[0].Timestamp
last := window.DataPoints[len(window.DataPoints)-1].Timestamp
summary.Duration = last.Sub(first)
}
// Track sums for averages
sums := make(map[string]float64)
counts := make(map[string]int)
// First and last values for change calculation
firstValues := make(map[string]float64)
lastValues := make(map[string]float64)
for i, dp := range window.DataPoints {
for metric, value := range dp.Metrics {
// Track first value
if i == 0 {
firstValues[metric] = value
summary.Peaks[metric] = value
summary.Lows[metric] = value
}
// Track last value
lastValues[metric] = value
// Track peaks and lows
if value > summary.Peaks[metric] {
summary.Peaks[metric] = value
}
if value < summary.Lows[metric] {
summary.Lows[metric] = value
}
// Track sums for average
sums[metric] += value
counts[metric]++
}
}
// Calculate averages and changes
for metric, sum := range sums {
if counts[metric] > 0 {
summary.Averages[metric] = sum / float64(counts[metric])
}
if first, ok := firstValues[metric]; ok {
if last, ok := lastValues[metric]; ok {
summary.Changes[metric] = last - first
}
}
}
return summary
}
// trimCompletedWindows removes old windows
func (r *IncidentRecorder) trimCompletedWindows() {
// Remove by retention duration
cutoff := time.Now().Add(-r.config.RetentionDuration)
kept := make([]*IncidentWindow, 0, len(r.completedWindows))
for _, w := range r.completedWindows {
if w.EndTime != nil && w.EndTime.After(cutoff) {
kept = append(kept, w)
}
}
r.completedWindows = kept
// Remove by max windows
if len(r.completedWindows) > r.config.MaxWindows {
r.completedWindows = r.completedWindows[len(r.completedWindows)-r.config.MaxWindows:]
}
}
// GetWindow returns a specific incident window
func (r *IncidentRecorder) GetWindow(windowID string) *IncidentWindow {
r.mu.RLock()
defer r.mu.RUnlock()
// Check active windows
if window, ok := r.activeWindows[windowID]; ok {
return copyWindow(window)
}
// Check completed windows
for _, window := range r.completedWindows {
if window.ID == windowID {
return copyWindow(window)
}
}
return nil
}
// GetWindowsForResource returns all incident windows for a resource
func (r *IncidentRecorder) GetWindowsForResource(resourceID string, limit int) []*IncidentWindow {
r.mu.RLock()
defer r.mu.RUnlock()
var result []*IncidentWindow
// Check active windows
for _, window := range r.activeWindows {
if window.ResourceID == resourceID {
result = append(result, copyWindow(window))
}
}
// Check completed windows (in reverse order for most recent first)
for i := len(r.completedWindows) - 1; i >= 0; i-- {
if r.completedWindows[i].ResourceID == resourceID {
result = append(result, copyWindow(r.completedWindows[i]))
if limit > 0 && len(result) >= limit {
break
}
}
}
return result
}
// saveToDisk persists completed windows
func (r *IncidentRecorder) saveToDisk() error {
if r.filePath == "" {
return nil
}
r.saveIOMu.Lock()
defer r.saveIOMu.Unlock()
data := struct {
CompletedWindows []*IncidentWindow `json:"completed_windows"`
}{
CompletedWindows: r.snapshotCompletedWindows(),
}
jsonData, err := json.MarshalIndent(data, "", " ")
if err != nil {
return fmt.Errorf("incident recorder save: marshal completed windows: %w", err)
}
if err := ensureOwnerOnlyDir(r.dataDir); err != nil {
return err
}
if info, err := os.Lstat(r.filePath); err == nil {
if err := validateRegularFilePath(r.filePath, info); err != nil {
return err
}
} else if !errors.Is(err, os.ErrNotExist) {
return err
}
tmpFile, err := os.CreateTemp(r.dataDir, filepath.Base(r.filePath)+".*.tmp")
if err != nil {
return err
}
tmpPath := tmpFile.Name()
cleanup := true
defer func() {
if cleanup {
_ = os.Remove(tmpPath)
}
}()
if err := tmpFile.Chmod(incidentRecorderFilePerm); err != nil {
_ = tmpFile.Close()
return err
}
if _, err := tmpFile.Write(jsonData); err != nil {
_ = tmpFile.Close()
return err
}
if err := tmpFile.Close(); err != nil {
return err
}
if err := os.Rename(tmpPath, r.filePath); err != nil {
return err
}
cleanup = false
return os.Chmod(r.filePath, incidentRecorderFilePerm)
}
func (r *IncidentRecorder) requestAsyncSave() {
if r.filePath == "" {
return
}
r.saveMu.Lock()
if r.saveInProgress {
r.saveRequested = true
r.saveMu.Unlock()
return
}
r.saveInProgress = true
r.saveMu.Unlock()
go r.runAsyncSaves()
}
func (r *IncidentRecorder) runAsyncSaves() {
for {
if err := r.saveToDisk(); err != nil {
log.Warn().
Str("file_path", r.filePath).
Err(err).
Msg("Failed to save incident windows")
}
r.saveMu.Lock()
if !r.saveRequested {
r.saveInProgress = false
r.saveCond.Broadcast()
r.saveMu.Unlock()
return
}
r.saveRequested = false
r.saveMu.Unlock()
}
}
func (r *IncidentRecorder) snapshotCompletedWindows() []*IncidentWindow {
r.mu.RLock()
defer r.mu.RUnlock()
snapshot := make([]*IncidentWindow, len(r.completedWindows))
for i, window := range r.completedWindows {
snapshot[i] = copyWindow(window)
}
return snapshot
}
func (r *IncidentRecorder) waitForPendingSaves() {
if r.filePath == "" {
return
}
r.saveMu.Lock()
for r.saveInProgress || r.saveRequested {
r.saveCond.Wait()
}
r.saveMu.Unlock()
r.saveIOMu.Lock()
r.saveIOMu.Unlock()
}
// loadFromDisk loads completed windows
func (r *IncidentRecorder) loadFromDisk() error {
if r.filePath == "" {
return nil
}
jsonData, err := readBoundedRegularFile(r.filePath, maxIncidentWindowsFileSize)
if err != nil {
if errors.Is(err, os.ErrNotExist) {
return nil
}
return fmt.Errorf("incident recorder load: read file %q: %w", r.filePath, err)
}
var data struct {
CompletedWindows []*IncidentWindow `json:"completed_windows"`
}
if err := json.Unmarshal(jsonData, &data); err != nil {
return fmt.Errorf("incident recorder load: parse file %q: %w", r.filePath, err)
}
r.completedWindows = make([]*IncidentWindow, 0, len(data.CompletedWindows))
for _, window := range data.CompletedWindows {
if window == nil {
continue
}
r.completedWindows = append(r.completedWindows, window)
}
r.trimCompletedWindows()
return os.Chmod(r.filePath, incidentRecorderFilePerm)
}
// Helper functions
func ensureOwnerOnlyDir(dir string) error {
if err := os.MkdirAll(dir, incidentRecorderDirPerm); err != nil {
return err
}
return os.Chmod(dir, incidentRecorderDirPerm)
}
func validateRegularFilePath(path string, info os.FileInfo) error {
if info.Mode()&os.ModeSymlink != 0 {
return fmt.Errorf("%w: refusing symlink path %q", errUnsafeIncidentPersistencePath, path)
}
if !info.Mode().IsRegular() {
return fmt.Errorf("%w: non-regular path %q", errUnsafeIncidentPersistencePath, path)
}
return nil
}
func readBoundedRegularFile(path string, maxSize int64) ([]byte, error) {
initialInfo, err := os.Lstat(path)
if err != nil {
return nil, err
}
if err := validateRegularFilePath(path, initialInfo); err != nil {
return nil, err
}
if maxSize > 0 && initialInfo.Size() > maxSize {
return nil, fmt.Errorf("%w: file %q exceeds size limit (%d bytes)", errUnsafeIncidentPersistencePath, path, initialInfo.Size())
}
file, err := os.Open(path)
if err != nil {
return nil, err
}
defer func() {
_ = file.Close()
}()
openInfo, err := file.Stat()
if err != nil {
return nil, err
}
if err := validateRegularFilePath(path, openInfo); err != nil {
return nil, err
}
if !os.SameFile(initialInfo, openInfo) {
return nil, fmt.Errorf("%w: file %q changed during read", errUnsafeIncidentPersistencePath, path)
}
reader := io.Reader(file)
if maxSize > 0 {
reader = io.LimitReader(file, maxSize+1)
}
data, err := io.ReadAll(reader)
if err != nil {
return nil, err
}
if maxSize > 0 && int64(len(data)) > maxSize {
return nil, fmt.Errorf("%w: file %q exceeded size limit while reading", errUnsafeIncidentPersistencePath, path)
}
return data, nil
}
func copyWindow(w *IncidentWindow) *IncidentWindow {
if w == nil {
return nil
}
windowCopy := *w
if w.EndTime != nil {
t := *w.EndTime
windowCopy.EndTime = &t
}
windowCopy.DataPoints = copyDataPoints(w.DataPoints)
if w.Summary != nil {
s := *w.Summary
s.Peaks = copyMetrics(w.Summary.Peaks)
s.Lows = copyMetrics(w.Summary.Lows)
s.Averages = copyMetrics(w.Summary.Averages)
s.Changes = copyMetrics(w.Summary.Changes)
if w.Summary.Anomalies != nil {
s.Anomalies = append([]string(nil), w.Summary.Anomalies...)
}
windowCopy.Summary = &s
}
return &windowCopy
}
var windowCounter int64
func generateWindowID(resourceID string) string {
counter := atomic.AddInt64(&windowCounter, 1)
return "iw-" + resourceID + "-" + time.Now().Format("20060102150405") + "-" + intToString(int(counter))
}
func copyDataPoints(points []IncidentDataPoint) []IncidentDataPoint {
copied := make([]IncidentDataPoint, len(points))
for i, dp := range points {
copied[i] = dp
copied[i].Metrics = copyMetrics(dp.Metrics)
if dp.Metadata != nil {
copied[i].Metadata = make(map[string]interface{}, len(dp.Metadata))
for k, v := range dp.Metadata {
copied[i].Metadata[k] = v
}
}
}
return copied
}
func copyMetrics(metrics map[string]float64) map[string]float64 {
if metrics == nil {
return nil
}
copied := make(map[string]float64, len(metrics))
for k, v := range metrics {
copied[k] = v
}
return copied
}
func intToString(n int) string {
if n == 0 {
return "0"
}
negative := n < 0
if negative {
n = -n
}
var result string
for n > 0 {
result = string(rune('0'+n%10)) + result
n /= 10
}
if negative {
result = "-" + result
}
return result
}
@@ -1,133 +0,0 @@
package metrics
import (
"sync/atomic"
"testing"
"time"
)
type countingProvider struct {
ids []string
metrics map[string]map[string]float64
calls int32
}
func (c *countingProvider) GetCurrentMetrics(resourceID string) (map[string]float64, error) {
atomic.AddInt32(&c.calls, 1)
metrics, ok := c.metrics[resourceID]
if !ok {
return nil, errNoMetrics(resourceID)
}
copied := make(map[string]float64, len(metrics))
for k, v := range metrics {
copied[k] = v
}
return copied, nil
}
func (c *countingProvider) GetMonitoredResourceIDs() []string {
return append([]string{}, c.ids...)
}
func waitForCalls(t *testing.T, provider *countingProvider, timeout time.Duration) {
t.Helper()
deadline := time.Now().Add(timeout)
for time.Now().Before(deadline) {
if atomic.LoadInt32(&provider.calls) > 0 {
return
}
time.Sleep(5 * time.Millisecond)
}
t.Fatal("expected provider to be called")
}
func TestDefaultIncidentRecorderConfig(t *testing.T) {
cfg := DefaultIncidentRecorderConfig()
if cfg.SampleInterval == 0 || cfg.PreIncidentWindow == 0 || cfg.PostIncidentWindow == 0 {
t.Fatalf("default config should be non-zero, got %+v", cfg)
}
if cfg.MaxDataPointsPerWindow == 0 || cfg.MaxWindows == 0 || cfg.RetentionDuration == 0 {
t.Fatalf("default config should be non-zero, got %+v", cfg)
}
}
func TestIncidentRecorderStartStop(t *testing.T) {
recorder := NewIncidentRecorder(IncidentRecorderConfig{
SampleInterval: 5 * time.Millisecond,
PreIncidentWindow: 10 * time.Millisecond,
PostIncidentWindow: 10 * time.Millisecond,
MaxDataPointsPerWindow: 5,
})
provider := &countingProvider{
ids: []string{"res-1"},
metrics: map[string]map[string]float64{
"res-1": {"cpu": 1},
},
}
recorder.SetMetricsProvider(provider)
recorder.Start()
waitForCalls(t, provider, 200*time.Millisecond)
recorder.Stop()
if recorder.running {
t.Fatalf("expected recorder to be stopped")
}
}
func TestGetWindowsForResource(t *testing.T) {
recorder := NewIncidentRecorder(IncidentRecorderConfig{})
active := &IncidentWindow{ID: "active-1", ResourceID: "res-1"}
recorder.activeWindows["active-1"] = active
got := recorder.GetWindowsForResource("res-1", 0)
if len(got) != 1 || got[0].ID != "active-1" {
t.Fatalf("expected active window, got %+v", got)
}
recorder.activeWindows = map[string]*IncidentWindow{}
recorder.completedWindows = []*IncidentWindow{
{ID: "old", ResourceID: "res-1"},
{ID: "new", ResourceID: "res-1"},
}
limited := recorder.GetWindowsForResource("res-1", 1)
if len(limited) != 1 || limited[0].ID != "new" {
t.Fatalf("expected most recent completed window, got %+v", limited)
}
}
func TestRecordSampleSkipsPreIncidentBufferOnMetricsError(t *testing.T) {
recorder := NewIncidentRecorder(IncidentRecorderConfig{
PreIncidentWindow: time.Minute,
PostIncidentWindow: time.Minute,
MaxDataPointsPerWindow: 10,
})
provider := &stubMetricsProvider{
metricsByID: map[string]map[string]float64{
"res-ok": {"cpu": 1},
},
ids: []string{"res-ok", "res-missing"},
}
recorder.SetMetricsProvider(provider)
windowID := recorder.StartRecording("res-ok", "db", "agent", "alert", "alert-1")
recorder.recordSample()
window := recorder.activeWindows[windowID]
if window == nil {
t.Fatalf("expected active window %s", windowID)
}
if len(window.DataPoints) != 1 {
t.Fatalf("expected active window sample to be captured, got %d", len(window.DataPoints))
}
if len(recorder.preIncidentBuffer["res-ok"]) == 0 {
t.Fatalf("expected pre-incident buffer for res-ok")
}
if _, ok := recorder.preIncidentBuffer["res-missing"]; ok {
t.Fatalf("expected no pre-incident buffer for res-missing when metrics collection fails")
}
}
@@ -1,84 +0,0 @@
package metrics
import (
"errors"
"testing"
"time"
)
// gappyBatchProvider implements both MetricsProvider and BatchMetricsProvider
// but omits one monitored ID from the batch. The recorder must fall through
// to the per-ID method for that ID so provider semantics are preserved.
type gappyBatchProvider struct {
batchCalls int
perIDCalls map[string]int
perIDResults map[string]map[string]float64
perIDErr map[string]error
}
func (p *gappyBatchProvider) GetMonitoredResourceIDs() []string {
return []string{"in-batch", "not-in-batch", "erroring"}
}
func (p *gappyBatchProvider) GetCurrentMetricsBatch() map[string]map[string]float64 {
p.batchCalls++
return map[string]map[string]float64{
"in-batch": {"cpu": 42},
}
}
func (p *gappyBatchProvider) GetCurrentMetrics(resourceID string) (map[string]float64, error) {
if p.perIDCalls == nil {
p.perIDCalls = map[string]int{}
}
p.perIDCalls[resourceID]++
if err := p.perIDErr[resourceID]; err != nil {
return nil, err
}
return p.perIDResults[resourceID], nil
}
func TestRecordSampleBatchMissFallsThroughToPerID(t *testing.T) {
recorder := NewIncidentRecorder(IncidentRecorderConfig{
SampleInterval: time.Hour, // ticks driven manually
PreIncidentWindow: time.Minute,
PostIncidentWindow: time.Minute,
MaxDataPointsPerWindow: 5,
})
provider := &gappyBatchProvider{
perIDResults: map[string]map[string]float64{
"not-in-batch": {"cpu": 7},
},
perIDErr: map[string]error{
"erroring": errors.New("unknown resource"),
},
}
recorder.SetMetricsProvider(provider)
recorder.recordSample()
if provider.batchCalls != 1 {
t.Fatalf("batch calls = %d, want 1", provider.batchCalls)
}
if provider.perIDCalls["in-batch"] != 0 {
t.Fatalf("batch hit %q still took the per-ID path", "in-batch")
}
if provider.perIDCalls["not-in-batch"] != 1 || provider.perIDCalls["erroring"] != 1 {
t.Fatalf("batch misses did not fall through per-ID: %+v", provider.perIDCalls)
}
recorder.mu.RLock()
defer recorder.mu.RUnlock()
if got := len(recorder.preIncidentBuffer["not-in-batch"]); got != 1 {
t.Fatalf("fallthrough metrics not buffered: %d points", got)
}
if got := recorder.preIncidentBuffer["not-in-batch"][0].Metrics["cpu"]; got != 7 {
t.Fatalf("fallthrough buffered cpu = %v, want 7 (per-ID value)", got)
}
if got := len(recorder.preIncidentBuffer["erroring"]); got != 0 {
t.Fatalf("erroring ID gained %d buffered points, want 0 (per-ID error must skip)", got)
}
if got := len(recorder.preIncidentBuffer["in-batch"]); got != 1 {
t.Fatalf("batch-served ID not buffered: %d points", got)
}
}
@@ -1,86 +0,0 @@
package metrics
import (
"sync"
"testing"
"time"
)
func TestGenerateWindowIDConcurrentUnique(t *testing.T) {
t.Parallel()
const (
workers = 16
idsPerWork = 128
)
ids := make(chan string, workers*idsPerWork)
start := make(chan struct{})
var wg sync.WaitGroup
for i := 0; i < workers; i++ {
wg.Add(1)
go func() {
defer wg.Done()
<-start
for j := 0; j < idsPerWork; j++ {
ids <- generateWindowID("res-1")
}
}()
}
close(start)
wg.Wait()
close(ids)
seen := make(map[string]struct{}, workers*idsPerWork)
for id := range ids {
if _, exists := seen[id]; exists {
t.Fatalf("duplicate window ID generated: %s", id)
}
seen[id] = struct{}{}
}
}
func TestIncidentRecorderConcurrentStartStopAndFlush(t *testing.T) {
recorder := NewIncidentRecorder(IncidentRecorderConfig{
SampleInterval: time.Millisecond,
PreIncidentWindow: 10 * time.Millisecond,
PostIncidentWindow: 10 * time.Millisecond,
MaxDataPointsPerWindow: 10,
DataDir: t.TempDir(),
})
provider := &stubMetricsProvider{
metricsByID: map[string]map[string]float64{
"res-1": {"cpu": 1},
},
ids: []string{"res-1"},
}
recorder.SetMetricsProvider(provider)
const goroutines = 8
const iterations = 15
start := make(chan struct{})
var wg sync.WaitGroup
for i := 0; i < goroutines; i++ {
wg.Add(1)
go func() {
defer wg.Done()
<-start
for j := 0; j < iterations; j++ {
recorder.Start()
windowID := recorder.StartRecording("res-1", "db", "agent", "alert", "a-1")
recorder.recordSample()
recorder.StopRecording(windowID)
recorder.Stop()
}
}()
}
close(start)
wg.Wait()
// Final stop should remain idempotent after concurrent shutdowns.
recorder.Stop()
}

Some files were not shown because too many files have changed in this diff Show More