diff --git a/docs/qualification/PATROL_ASSISTANT_CUSTOMER_JOURNEY.md b/docs/qualification/PATROL_ASSISTANT_CUSTOMER_JOURNEY.md index daa6d4acc..a59207d81 100644 --- a/docs/qualification/PATROL_ASSISTANT_CUSTOMER_JOURNEY.md +++ b/docs/qualification/PATROL_ASSISTANT_CUSTOMER_JOURNEY.md @@ -93,12 +93,13 @@ Each removal must run its focused regression and affected complete journey. ### Completion and external dependencies The local implementation goal remains open until required qualification is -performed. Ordinary Assistant requests work with the current subscription route, -but autonomous Patrol has an explicit provider-policy refusal and remains -blocked. Do not rephrase the refused probe, bypass the readiness boundary or -count an interactive request as an autonomous Patrol pass. A supported provider -path is required for that qualification. Prepare other work while resolving the -provider dependency through supported configuration. +performed. The maintainer authorized Gemini 3.8 Flash through OpenRouter with a +US$5 key limit and one-day expiry on 2026-09-06. That supported route passes the +streaming readiness and initial live Watch and dependency cases recorded below. +The earlier Claude subscription refusal belongs to the exact synthetic +continuation request. It does not establish a blanket restriction on autonomous +monitoring. The refused request has not been retried or rephrased. Readiness is +not evidence that diagnosis, action execution or independent recovery succeeds. Release publication and wider product readiness are separate. Independent volunteered Pro environments are still required before claiming repeatable @@ -2047,3 +2048,114 @@ report to Pulse, enroll or replace a production agent, call a provider, or qualify approval, execution, diagnosis or model recovery. No runtime or frontend source changed in this slice, so no new browser claim is made. Prior storage diagnosis failures and the cached autonomous-provider refusal remain open. + + +## Funded Gemini qualification, 2026-09-06 + +The maintainer authorized `openrouter:google/gemini-3.8-flash` with a provider-side +US$5 key limit expiring on 2026-09-07. The provider key endpoint confirmed both +constraints. Credentials remain in runtime configuration, not these receipts. +Synthetic readiness passed in 9.034 seconds: three streaming tool scenarios, +two context fixtures and multi-turn continuation. This supports the readiness +claims for Watch only and Ask first. It does not qualify autonomous fixes. + +The first unhealthy-container run, `q-20260906-172952-3cccbf34`, detected the +correct fault and left the healthy control alone, but failed overall. The exact +Gemini route had no price entry, and the model attempted unsupported Docker +configuration access. The shared price table now records the reviewed standard +rates of US$0.75 input and US$3.75 output per million tokens for direct Gemini +and OpenRouter. Variant routes remain unknown. These introductory rates must be +reviewed on 2027-01-01. The query capability description now explicitly names +TrueNAS as the supported app-container configuration adapter and directs Docker +collected health/mount/port/network reads to `get`. Runtime permissions and +qualification gates are unchanged. + +The following runs used the worker-built Pro binary +`74464e75977caf55cda092c8cf56c24967616c8c24d770872fbea5d86e31a1dc`, +core base `b0b39f00dc6685ad9ed63e8a6e91b954338073e4` plus the pricing and +capability-description changes, and canonical enterprise base +`3d9f4e3051d38027355a2a1f36b8c7f672a09b65`. The worker archive commit +`d9cb84e15d1acc377341129bdda5c28176e7128c` has identical contents for all +87 tracked enterprise files. The existing runner created disposable +resources on Tower, waited for normal collection, and used independent fault, +recovery and cleanup oracles. + +| Case / run | Result | Evidence | +|---|---|---| +| Unhealthy, `q-20260906-174546-a7a9810b` | Pass | 9.709s detection phase, two tools, no failed/duplicate calls, healthy sibling unflagged. | +| Unhealthy, `q-20260906-174708-81f8d655` | Pass | 10.541s detection phase, exact unhealthy resource found. | +| Unhealthy, `q-20260906-174758-14deaa15` | Pass | 9.395s detection phase, exact unhealthy resource found. | +| Unhealthy, `q-20260906-174853-71cf894f` | Pass | 25.673s detection phase, exact unhealthy resource found. | +| Healthy mixed, `q-20260906-180144-381874a6` | Pass | 5.200s detection phase, no false findings. | +| Dependency, `q-20260906-175058-de5e350d` | Pass | Starting from only the client symptom, identified the stopped dependency and affected client. Investigation completed in 18.970s with three evidence calls and no mutation. | +| Storage, `q-20260906-175232-3ce6fbfa` | Fail before inference | Normal collection never converged to the required resource projection. No model diagnosis was attempted. | +| Approved restart, `q-20260906-175812-0593b9b7` | Fail before approval | Detection and investigation completed, but no exact action reference existed. The broker refused because Tower's Docker command agent was disconnected. Nothing executed. | +| Rejected restart, `q-20260906-175940-bff6992a` | Fail before rejection | No exact action was available to reject. This does not qualify rejected-action handling. | + +Every listed run passed cleanup, including second-cleanup no-op and unchanged +inventory. Individual Watch run estimates were about US$0.007 to US$0.014. +Those scorecard estimates cover the Patrol detection phase, not the separate +investigation calls. Provider-side aggregate spend is the budget authority for +this temporary key. The fixed route price does not turn an estimate into a +reconciled bill or establish a hard Pulse budget for unpriced history. + +Live qualification exposed two additional shared contract defects. The +investigation orchestrator logged action-broker refusal but completed the +record without retaining the error, leaving the model's captured-proposal prose +visible without the later refusal. The current enterprise change retains the +original diagnosis, persists the broker refusal as a failed investigation with +`needs_attention`, and creates no action reference. The product history adapter +also projected result-bearing transcript calls back into provider request calls, +dropping observed output and success/failure. The current core change uses one +shared transcript type for stored chat and product history, preserving the +separate explicit provider projection. + +The live review exposed duplicate detail IDs, duplicate unformatted conclusions, +paused history made unclickable by the scheduling switch, and narrow filter/sort +overlap. The shared finding/investigation surfaces now preserve one detail target, +render sanitized Markdown once for identical summaries, retain distinct summaries, +keep history available while paused, and wrap controls. Result-bearing tool calls +use the same expandable evidence component as Assistant. Historical calls without +a result status retain their evidence without invented success or failure. An +investigation outcome of `cannot_fix` or `needs_attention` does not identify who +resolved the finding, so the shared resolution copy no longer infers manual review. + +Private run receipts and source/binary bindings are under +`tmp/patrol-gemini-38/` in the workspace. The original failed runs remain failed. +Installed storage collection and a temporary command-enabled lab agent remain +prerequisites for real storage and approved/rejected recovery qualification. +No production agent has been replaced. Full action outcomes, remaining backup +coverage and independent volunteered Pro environments remain open. + + +### Final refusal and evidence-retention proof + +Two further approved-remediation attempts remain **failed**: +`q-20260906-181839-6cbdc711` and `q-20260906-182952-6acfb739`. Both retained the +broker error separately from the original model summary, saved `status=failed` +and `outcome=needs_attention`, and created no action reference. Both passed +cleanup. The final run used Pro binary SHA256 +`24de8c9ea0020067d298c489f0d99a272b5ec4b00afab7a40dff55aabb244061`, +including the proposal-response clarification, and detected its exact unhealthy +container with no false positives. Its saved investigation is +`48a16b05-50f0-4605-847c-0a71b3435975` for finding `ca3af29ac54d540f`. +The original failed scorecards have not been reclassified as action passes. + +The restored history API retains observed outputs and explicit `success=false` +for historical `pulse_read` failures. Live browser review at `/patrol` exercises +successful query output, both `ACTION_NOT_ALLOWED` and `NO_AGENT` failures, +original diagnosis, one broker error, paused history and review focus return. +The settings proof at `/settings/pulse-intelligence/patrol` checks the exact +model, reviewed rates, synthetic readiness limits and reload. The current +source-bound browser receipt records desktop, intermediate and mobile results. +GET response fixtures cover unknown historical result status only, without +claiming new persisted model evidence or action execution. + +Remaining qualification requires a current installed collector and a temporary +command-enabled lab agent. A Linux amd64 agent has been built on the worker, +SHA256 `ae2ed8b97709ec6e71af979c293ca9d3634662767ec1b59629baf4933c90cf5d`, +without installing it or changing Tower credentials. Tower's separate production +reporting agent is untouched. Any agent enrollment must use the canonical scoped +installation flow, preserve explicit identity and revoke temporary execution +access after qualification. Detection success does not satisfy this prerequisite +or the remaining backup and independent-environment cases. diff --git a/docs/release-control/control_plane.json b/docs/release-control/control_plane.json index 888734ae0..2300c5bb1 100644 --- a/docs/release-control/control_plane.json +++ b/docs/release-control/control_plane.json @@ -23,6 +23,11 @@ "prerelease_branch": "release/v6.4", "stable_branch": "release/v6.4" }, + { + "version_prefix": "6.4.4", + "prerelease_branch": "release/v6.4", + "stable_branch": "release/v6.4" + }, { "version_prefix": "6.5.", "prerelease_branch": "release/v6.5", diff --git a/docs/release-control/v6/internal/RELEASE_PROMOTION_POLICY.md b/docs/release-control/v6/internal/RELEASE_PROMOTION_POLICY.md index 24f102232..7c079ed54 100644 --- a/docs/release-control/v6/internal/RELEASE_PROMOTION_POLICY.md +++ b/docs/release-control/v6/internal/RELEASE_PROMOTION_POLICY.md @@ -282,6 +282,19 @@ without the other lanes changing the candidate underneath it. moving `main` can no longer invalidate the compiler's exact-SHA binding between dispatch and compilation, which is what failed run 33579042375. Earlier `6.4.x` versions keep their historical `main` mapping. +8. The forward regression checkpoint `v6.4.4-beta.1` uses the same + `release/v6.4` line, with an explicit `6.4.4` mapping for beta, RC and + eventual stable. This is a maturity reset, not new feature scope: a + `6.4.3-beta.N` would sort below the published `v6.4.3-rc.1`. + `v6.4.4-beta.1` advances both that preview and stable `v6.4.1`; + `v6.4.1` remains the rollback target. Later qualification proceeds through + `v6.4.4-rc.N` and exact same-version stable promotion, including a fresh + 72-hour clean RC soak. Beta time does not count. Mapping is preparation, + not readiness or publication authority: failed candidate checks still + require repair or an evidence-based disposition under the existing gates. + Land the mapping and resolver contract on canonical main and backport it + to the release line before taking a fresh bound packet. Unlisted patches + and new product work retain their existing mapping and scope. ## Paid Pro Artifact Lineage diff --git a/docs/release-control/v6/internal/status.json b/docs/release-control/v6/internal/status.json index 3796c3558..9a6a3ab30 100644 --- a/docs/release-control/v6/internal/status.json +++ b/docs/release-control/v6/internal/status.json @@ -10201,7 +10201,7 @@ }, { "id": "patrol-assistant-customer-outcome-qualification", - "summary": "The redesign goal remains open, with its executable plan and detailed receipts in docs/qualification/PATROL_ASSISTANT_CUSTOMER_JOURNEY.md. Patrol owns investigation and Assistant continues the same issue. Observations, hypotheses, accepted proposals, execution and independently verified outcomes remain distinct. The recorded baseline contains 127 paid installations, 71 Patrol-enabled, 23 with Assistant calls and fourteen verified resolutions from one installation. Schema17 outcome/provider/cost fields had no adoption. These do not establish representative success, false-alarm or missed-problem rates. Shared provenance/history/risk fixes and removal of proposal-as-proof and proxy completion policy landed through PR1928/1929. PR1934 merged canonical tool/transcript identity. PR1935 contains measurement-presence, capacity, canonical config-read, command-connectivity and tmpfs collection/query corrections through 6e18777d30f30b498def39d30016a697cabc4ea7. Exact worker hooks and named browser matrices passed, with remote landing pending. Real ordinary storage diagnosis failed by ruling out pressure without filesystem capacity evidence. Same-session recovery correctly identified present health and alert resolution but overstated continuous control health. Independent fault-intact, recovery, two-pass cleanup and persisted transcript checks passed. The current incident history correction replaces the primary legacy recorder read with the existing organization-pinned canonical resource timeline, preserving source/time semantics, bounded history, empty coverage and failed-read distinctions. Explicit legacy archive reads remain resource-bound. Its registered-tool SQLite, full tools package, focused race and final-source browser proofs pass. The primary history read was pushed as 580a246 in PR1935. The shared monitor/store identity correction now joins exact full Docker references with canonical container history through an organization-scoped history-only alias index. Original event content and approval/operator authority remain unchanged. Callback, removal/restart, replay, tenant/control and registered-tool tests, full affected packages, race checks and scoped lookup benchmarks passed. Final worker-build API proof returns the same seven retained homelab records under canonical and legacy queries, preserving the original fired/resolved events exactly. Six-case Playwright proof at desktop/intermediate/mobile widths passed. The exact thirteen-file worker hook passed and 919331d5 was pushed through PR1935. Integration with main 11a8cc2180 preserves both contract changes and passes affected backend, race, frontend and final-build browser/API proofs. The full integration hook gates the merge commit. The disconnected legacy recorder/coordinator and its cached-metrics adapter are retired. Explicit read-only archive lookup preserves recorded values and discloses legacy duration units. The old incidents API live count is explicitly unmeasured. Final archive/history/API and race checks pass, including the full API suite under the normal worker account after identifying a root-only permission-fixture failure. Final worker-build browser proof covers five archive outcomes across desktop/intermediate/mobile widths. Live canonical and legacy queries preserve the same seven original homelab records, and the API count is unmeasured. The exact staged hook gates landing. The legacy incident-memory listing still needs alias-aware canonical queries, canonical-only events, honest bounds and propagated projection-read failures. Installed tmpfs collector and real-model interpretation remain unqualified. Other named residuals include unsupported filters, typed compatibility canonical-ID lookup, legacy direction availability, Docker-host history and responsive mount details. Claude Max explicitly refused autonomous Patrol readiness. Cached refusal and API409 enforcement remain enforced. Separate paid-provider approval is pending, with no paid request or policy bypass. Ordinary Assistant is not autonomous qualification. Required local work still includes reliable interpretation, config-read/model retest, storage/backup and approved/rejected action outcomes. Independent volunteered Pro environments remain a separate wider-readiness gate.", + "summary": "The redesign goal remains open, with its executable plan and detailed receipts in docs/qualification/PATROL_ASSISTANT_CUSTOMER_JOURNEY.md. Patrol owns investigation and Assistant continues the same issue. Observations, hypotheses, accepted proposals, execution and independently verified outcomes remain distinct. The recorded baseline contains 127 paid installations, 71 Patrol-enabled, 23 with Assistant calls and fourteen verified resolutions from one installation. Schema17 outcome/provider/cost fields had no adoption. These do not establish representative success, false-alarm or missed-problem rates. Shared provenance/history/risk fixes and removal of proposal-as-proof and proxy completion policy landed through PR1928/1929. PR1934 merged canonical tool/transcript identity. PR1935 contains measurement-presence, capacity, canonical config-read, command-connectivity and tmpfs collection/query corrections through 6e18777d30f30b498def39d30016a697cabc4ea7. Exact worker hooks and named browser matrices passed, with remote landing pending. Real ordinary storage diagnosis failed by ruling out pressure without filesystem capacity evidence. Same-session recovery correctly identified present health and alert resolution but overstated continuous control health. Independent fault-intact, recovery, two-pass cleanup and persisted transcript checks passed. The current incident history correction replaces the primary legacy recorder read with the existing organization-pinned canonical resource timeline, preserving source/time semantics, bounded history, empty coverage and failed-read distinctions. Explicit legacy archive reads remain resource-bound. Its registered-tool SQLite, full tools package, focused race and final-source browser proofs pass. The primary history read was pushed as 580a246 in PR1935. The shared monitor/store identity correction now joins exact full Docker references with canonical container history through an organization-scoped history-only alias index. Original event content and approval/operator authority remain unchanged. Callback, removal/restart, replay, tenant/control and registered-tool tests, full affected packages, race checks and scoped lookup benchmarks passed. Final worker-build API proof returns the same seven retained homelab records under canonical and legacy queries, preserving the original fired/resolved events exactly. Six-case Playwright proof at desktop/intermediate/mobile widths passed. The exact thirteen-file worker hook passed and 919331d5 was pushed through PR1935. Integration with main 11a8cc2180 preserves both contract changes and passes affected backend, race, frontend and final-build browser/API proofs. The full integration hook gates the merge commit. The disconnected legacy recorder/coordinator and its cached-metrics adapter are retired. Explicit read-only archive lookup preserves recorded values and discloses legacy duration units. The old incidents API live count is explicitly unmeasured. Final archive/history/API and race checks pass, including the full API suite under the normal worker account after identifying a root-only permission-fixture failure. Final worker-build browser proof covers five archive outcomes across desktop/intermediate/mobile widths. Live canonical and legacy queries preserve the same seven original homelab records, and the API count is unmeasured. The exact staged hook gates landing. The legacy incident-memory listing still needs alias-aware canonical queries, canonical-only events, honest bounds and propagated projection-read failures. Installed tmpfs collector and real-model interpretation remain unqualified. Other named residuals include unsupported filters, typed compatibility canonical-ID lookup, legacy direction availability, Docker-host history and responsive mount details. The earlier Claude subscription refusal is limited to the exact synthetic continuation request and has not been retried. On 2026-09-06 the maintainer authorized Gemini 3.8 Flash through OpenRouter with a US$5 one-day key. Readiness, four unhealthy-container runs, the healthy mixed control and the client-to-stopped-dependency investigation passed on Tower. Storage failed before inference because collection did not converge. Approved and rejected restart cases found the fault but did not reach either decision: the disconnected Docker command agent prevented creation of an exact action. These failures exposed lost broker errors and lost tool-result fields in the product-history adapter. Current changes preserve the failed submission separately from the model diagnosis and retain result-bearing transcript calls, with exact provider projections unchanged. Browser work corrects duplicate review IDs, unformatted duplicate summaries, paused-history access and narrow control overlap. Final rebuilt live/API/browser and worker proof gate landing. Required local work still includes installed storage collection, a command-enabled lab agent, approved/rejected action verification, storage/backup and the remaining named evidence gaps. Independent volunteered Pro environments remain a separate wider-readiness gate. Final broker-refusal repeats retain failed/needs_attention, original diagnosis and no action. Investigation history now preserves and renders result-bearing calls through the shared transcript and Assistant evidence component, including explicit failed reads and unknown historical status. Shared resolution copy no longer infers manual review from an attention/cannot-fix outcome. Desktop/intermediate/mobile browser proof binds final content. Real approved/rejected recovery, installed storage collection, backup coverage and independent environments remain open. A verified Linux agent artifact is prepared but no Tower agent credentials or production agent installation have been changed.", "owner": "project-owner", "status": "planned", "recorded_at": "2026-09-05", diff --git a/docs/release-control/v6/internal/subsystems/ai-runtime.md b/docs/release-control/v6/internal/subsystems/ai-runtime.md index cbddc5956..e6d80a622 100644 --- a/docs/release-control/v6/internal/subsystems/ai-runtime.md +++ b/docs/release-control/v6/internal/subsystems/ai-runtime.md @@ -25,6 +25,46 @@ that same result. Successful reads retain their content and execution provenance ## Purpose +Stored chat and product history share the result-bearing `TranscriptToolCall` +contract. API adapters preserve observed output and the explicit success/error +bit. Only provider-request projections remove those display fields. A failed +read must not become an invocation with no visible result on the way to Patrol +or Assistant history. The adapter regression includes a `NO_AGENT` result and +`success: false`, and provider serialization retains its existing narrower shape. + +Capturing a typed proposal does not create an action. The proposal response +discloses that broker validation is still pending, without forcing the model to +stop investigating. If the broker later refuses submission, the enterprise +orchestrator retains the model's diagnosis unchanged and records the broker +error as a failed investigation needing attention, with no action reference. +The real disconnected-agent case must remain unsuccessful until its actual +transport prerequisite is satisfied. A successful model turn or recorded +proposal is not approval, execution or recovery. + +The shared investigation review renders sanitized Markdown and does not repeat +an identical persisted/fetched conclusion or error. Distinct evidence remains +visible. Pausing scheduled Patrol does not disable history review. The review +control has one detail target and returns keyboard focus when closed. Merged tool +results use Assistant's shared expandable evidence component. Historical calls +without an explicit result bit do not gain an inferred success/failure state. +Resolution copy cannot infer manual review from `needs_attention` or `cannot_fix`. + +The shared pricing table includes reviewed standard Gemini 3.8 Flash rates for +the exact direct and OpenRouter routes. OpenRouter variants and aliases remain +unpriced until independently reviewed. Rates carry the review date and are +estimates, not reconciled provider charges. The introductory rates require a +new review on 2027-01-01. `TestGemini38FlashReviewedRoutePricing` covers real +qualification token counts and request-route preservation, and +`TestGemini38OpenRouterPricingDoesNotGuessVariantRates` preserves unknown variants. + +The canonical query tool describes the app-container configuration boundary +explicitly: TrueNAS supports `config`, while Docker/Podman expose their collected +health, mounts, ports and networks through `get`. This communicates the existing +adapter contract to the model. It does not add configuration access, suppress +tool errors or weaken qualification gates. The existing +`TestAppContainerConfigObservationContract` retains unsupported-adapter and +provider/identity boundaries. + Shared app-container query mount evidence preserves native type, source, destination, options and canonical read/write access. Compound options such as `ro,noexec` cannot become writable through string equality heuristics. Both the @@ -837,6 +877,7 @@ cheap local detection into model-owned diagnosis and governed action. 31. `internal/agentcapabilities/tool_names.go` shared with `api-contracts`: the Pulse Intelligence registry tool-name vocabulary is both the native Assistant execution/display contract and the canonical API/agent tool identity contract for MCP-facing external-agent adapters. 32. `internal/agentcapabilities/tool_response.go` shared with `api-contracts`: the shared tool response envelope, tool error-code vocabulary, and tool-result error-code and verification evidence parsers are both the Assistant structured tool-result contract and the canonical API/agent branching contract for Pulse Intelligence tool failures, recovery tracking, and write self-verification. 33. `internal/agentcapabilities/tool_result.go` shared with `api-contracts`: the Pulse Intelligence shared tool-result content/result envelope, structuredContent projection, result constructors, HTTP response-to-result mapping, text projection, and result interpretation helpers are both the Assistant registry result contract and the canonical API/agent result projection contract for governed tool outcomes. +34. `internal/agentcapabilities/transcript.go` shared with `api-contracts`: Stored Assistant tool results and product history share one result-bearing transcript contract, with an explicit narrower provider-request projection. Observed failures and absent historical result status must survive the API boundary. 34. `internal/agentcapabilities/types.go` shared with `api-contracts`: the agent capabilities manifest wire type, manifest-owned external-adapter surface tool contract field, capability display title and structured output schema fields, approval-policy vocabulary, capability governance normalization, and tool-governance descriptor shape are both the canonical API payload contract and the AI runtime projection contract for Pulse Assistant and MCP-facing agent tools. 35. `internal/agentcapabilities/workflow_prompt.go` shared with `api-contracts`: the Pulse Intelligence workflow prompt catalogue, manifest-owned `workflowPrompts` projection, MCP prompt title projection, presentation kind hints, shared resource-context and finding argument vocabulary, Patrol issue-handling capability gating, argument validation, and manifest-gated shared prompt rendering rules are both the AI runtime starter contract for Assistant-compatible surfaces and the canonical API/agent prompt projection contract for MCP-facing clients. 36. `internal/api/ai_handler.go` shared with `api-contracts`: Pulse Assistant handlers are both an AI runtime control surface and a canonical API payload contract boundary. diff --git a/docs/release-control/v6/internal/subsystems/api-contracts.md b/docs/release-control/v6/internal/subsystems/api-contracts.md index 25cb1e0a1..dce8d007d 100644 --- a/docs/release-control/v6/internal/subsystems/api-contracts.md +++ b/docs/release-control/v6/internal/subsystems/api-contracts.md @@ -20,6 +20,14 @@ ## Purpose +Product history retains the stored result-bearing `TranscriptToolCall` contract, +including observed output and an optional success bit. A false bit is retained, +and an absent historical result bit stays absent. Provider requests use the +explicit narrower projection. The history adapter cannot project away evidence +by treating a display transcript as request arguments. The frontend Patrol API +extends the shared Assistant tool-call shape, and transport regressions pin +failed output and unknown historical status across the message envelope. + The internal Patrol bridge preserves explicit execution limits, scoped tool allowlists and execution identity. The retired unmatched-signal evaluator no longer contributes a signal-count-derived successful-report budget. Diagnosis @@ -1683,6 +1691,7 @@ payload shape change when the portal presents compact client rows. 57. `internal/agentcapabilities/tool_names.go` shared with `ai-runtime`: the Pulse Intelligence registry tool-name vocabulary is both the native Assistant execution/display contract and the canonical API/agent tool identity contract for MCP-facing external-agent adapters. 58. `internal/agentcapabilities/tool_response.go` shared with `ai-runtime`: the shared tool response envelope, tool error-code vocabulary, and tool-result error-code and verification evidence parsers are both the Assistant structured tool-result contract and the canonical API/agent branching contract for Pulse Intelligence tool failures, recovery tracking, and write self-verification. 59. `internal/agentcapabilities/tool_result.go` shared with `ai-runtime`: the Pulse Intelligence shared tool-result content/result envelope, structuredContent projection, result constructors, HTTP response-to-result mapping, text projection, and result interpretation helpers are both the Assistant registry result contract and the canonical API/agent result projection contract for governed tool outcomes. +60. `internal/agentcapabilities/transcript.go` shared with `ai-runtime`: Stored Assistant tool results and product history share one result-bearing transcript contract, with an explicit narrower provider-request projection. Observed failures and absent historical result status must survive the API boundary. 60. `internal/agentcapabilities/types.go` shared with `ai-runtime`: the agent capabilities manifest wire type, manifest-owned external-adapter surface tool contract field, capability display title and structured output schema fields, approval-policy vocabulary, capability governance normalization, and tool-governance descriptor shape are both the canonical API payload contract and the AI runtime projection contract for Pulse Assistant and MCP-facing agent tools. 61. `internal/agentcapabilities/workflow_prompt.go` shared with `ai-runtime`: the Pulse Intelligence workflow prompt catalogue, manifest-owned `workflowPrompts` projection, MCP prompt title projection, presentation kind hints, shared resource-context and finding argument vocabulary, Patrol issue-handling capability gating, argument validation, and manifest-gated shared prompt rendering rules are both the AI runtime starter contract for Assistant-compatible surfaces and the canonical API/agent prompt projection contract for MCP-facing clients. 62. `internal/api/access_control_handlers.go` shared with `organization-settings`: RBAC role and user-assignment handlers are both an organization settings control surface and a canonical API payload contract boundary. diff --git a/docs/release-control/v6/internal/subsystems/patrol-intelligence.md b/docs/release-control/v6/internal/subsystems/patrol-intelligence.md index be689ee94..4fe2cffb5 100644 --- a/docs/release-control/v6/internal/subsystems/patrol-intelligence.md +++ b/docs/release-control/v6/internal/subsystems/patrol-intelligence.md @@ -15,6 +15,16 @@ ## Purpose +Pausing the Patrol schedule does not disable investigation history review. +The finding review control owns one detail target and regains keyboard focus on +close. Investigation conclusions use sanitized Markdown, deduplicate identical +stored/fetched text, and retain distinct evidence. Merged tool results render +through Assistant's shared expandable evidence component with their explicit +success/failure bit. Unknown historical status stays unknown. A broker refusal +remains visible separately from the original diagnosis and cannot be presented +as an accepted action. An attention/cannot-fix outcome does not establish manual +review or identify who resolved a later finding. + Detection retains one model conversation for evidence gathering and finding decisions. Recording one finding does not establish diagnostic sufficiency or remove its evidence tools. Missing assessments and provider failures remain diff --git a/docs/release-control/v6/internal/subsystems/registry.json b/docs/release-control/v6/internal/subsystems/registry.json index d510933ef..b2a859541 100644 --- a/docs/release-control/v6/internal/subsystems/registry.json +++ b/docs/release-control/v6/internal/subsystems/registry.json @@ -779,6 +779,14 @@ "api-contracts" ] }, + { + "path": "internal/agentcapabilities/transcript.go", + "rationale": "Stored Assistant tool results and product history share one result-bearing transcript contract, with an explicit narrower provider-request projection. Observed failures and absent historical result status must survive the API boundary", + "subsystems": [ + "ai-runtime", + "api-contracts" + ] + }, { "path": "internal/agentcapabilities/types.go", "rationale": "the agent capabilities manifest wire type, manifest-owned external-adapter surface tool contract field, capability display title and structured output schema fields, approval-policy vocabulary, capability governance normalization, and tool-governance descriptor shape are both the canonical API payload contract and the AI runtime projection contract for Pulse Assistant and MCP-facing agent tools", diff --git a/frontend-modern/browser-verification.json b/frontend-modern/browser-verification.json index 379a1e717..a4f4d323a 100644 --- a/frontend-modern/browser-verification.json +++ b/frontend-modern/browser-verification.json @@ -1,39 +1,49 @@ { "version": 1, - "base_sha": "ab6d21400074393379a78350249ba4f5412941c0", - "verified_at": "2026-09-06T16:30:47.586Z", + "base_sha": "b0b39f00dc6685ad9ed63e8a6e91b954338073e4", + "verified_at": "2026-09-06T18:44:39.842657Z", "result": "passed", "changed_paths": [ - "frontend-modern/src/features/alerts/deliveryDiagnosisPresentation.ts" + "frontend-modern/src/api/patrol.ts", + "frontend-modern/src/components/AI/FindingsPanel.tsx", + "frontend-modern/src/components/Login.tsx", + "frontend-modern/src/components/patrol/InvestigationMessages.tsx", + "frontend-modern/src/components/patrol/InvestigationSection.tsx", + "frontend-modern/src/features/patrol/PatrolIntelligenceSurface.tsx", + "frontend-modern/src/useAppRuntimeState.ts", + "frontend-modern/src/utils/aiFindingPresentation.ts", + "frontend-modern/src/utils/localStorage.ts" ], "content_sha256": { - "frontend-modern/src/features/alerts/deliveryDiagnosisPresentation.ts": "bc3eadf8f790517b377430a42eb67e8fdfcadb53bb13b87cd971d6a9ab91d607" + "frontend-modern/src/api/patrol.ts": "3e5bec5d29d9218450c3fb5b531fd17652759c325e049bde497e3fb96544a622", + "frontend-modern/src/components/AI/FindingsPanel.tsx": "1063492593361f29a88109641d76fc2492027cd23b470452dfa4f437fa325700", + "frontend-modern/src/components/Login.tsx": "621f0b97d775aacaa521a3c72ed02db3fde5ad3eda06d55950a32107723eb44c", + "frontend-modern/src/components/patrol/InvestigationMessages.tsx": "4518f4755ff3d7dcdd0d8030372011e3481726db27ef695a9afc1d4ea6af3976", + "frontend-modern/src/components/patrol/InvestigationSection.tsx": "32669694f6dc959c3c3684f17a4ecb4c3d9411f7fc2ae20f1192d68b1d87e52b", + "frontend-modern/src/features/patrol/PatrolIntelligenceSurface.tsx": "1e52d00e61fb6affb9365cee74eb6820135d1d6d8a20161d02499ab3d6084da5", + "frontend-modern/src/useAppRuntimeState.ts": "5a1d343d92c441e1302dab129be7256dd9a287c6a60fb0e17c6f9fa13fe20900", + "frontend-modern/src/utils/aiFindingPresentation.ts": "c716d8b61501acec2bfde253febd4ddce6e9dfee9eef77a3230a2512326dfd7c", + "frontend-modern/src/utils/localStorage.ts": "182ed45685228781fd32725db50611b1115f9ee4fb1af23aeedcf41ca707bdbc" }, "backend_content_sha256": { - "internal/ai/adapters/adapters.go": "487245a52e1ae85ffecd38c6e4f006bfd6c8efff6568c92c19688a8db56d7ecc", - "internal/ai/chat/service.go": "4f130864717c5141ce974b8a367aeca19d49baa02a4906ecb65591d8b4c02bd6", - "internal/ai/tools/executor.go": "3a30f7720d30be6fdd75f3224d2da937881e77fee76fb819347f83e247fccb65", - "internal/ai/tools/tools_knowledge.go": "12c16ed7893dd49c18a7db1dd3fe3224c266102e9a23c37bc6b86fe5b33b9e4c", - "internal/api/ai_handler.go": "a407a5e55320b8e5d41d121193dc654ef02f70bd7b4566948bf4cd0a5763470a", - "internal/api/ai_handlers.go": "e3963c5a564432e9ee590f3c3db4a2220abc435cd78f6b14213559423a0283f6", - "internal/api/ai_intelligence_handlers.go": "f3864f8a53a1adaa2f64983279c3349dfff1e3095fe2ee8843ea22afe3dab577", - "internal/api/router.go": "cb5a99f8d12a7b576bf606f76fb4dfe6d1db888305249879268e6464efcb7c3f", - "internal/metrics/incident_archive.go": "2843483fca0bbafa4c6ff14b419bd599f6aeac7e20025b604cadbf16bddeed7a" + "internal/agentcapabilities/transcript.go": "356c4ca201470407988ff9b2c1fb848619ed38e9d8db844e7390adce0f93ec19", + "internal/ai/chat/types.go": "1f624daf7e511eb971e2d81b787dcd72b2a82d0e0ecd775aa4e580c59d915382", + "internal/ai/service.go": "25dca2a70a985e8ab07443444a9f05f01a569c62ff7bc068a0b587d500831de6", + "internal/api/chat_service_adapter.go": "6d0ab14456b1901c5020a408de95796ece1aa8d1057dd58737ea2637777d6ccb", + "internal/ai/cost/pricing.go": "7b64bc881311ee0a7c1c8a319fcd162974f5a307ced1679e614ff20ac1ac532c", + "internal/ai/tools/tools_query.go": "3e074b204a8c8c4f8b66eaf147c269a2e57739bfc69d03ea36ee8c47cf8b908a", + "internal/ai/tools/tools_propose.go": "43d720c78a010b72f53f7e4e9e1b7b2a763e7f1ac555edb921b7cdca3e82fd51" }, - "removed_backend_paths": [ - "internal/ai/incident_coordinator.go", - "internal/metrics/incident_recorder.go" - ], - "binary_sha256": "bd29e6f27be7b3ad4cfbc37842f4da90f08c6a48c9fc23b12c9c597b67346c9b", - "rendering_content_sha256": { - "frontend-modern/src/components/AI/Chat/ChatMessages.tsx": "9672f7608d1e3a531c73cba20fd4a78752316783212afd0c292ddfd11d2bf371", - "frontend-modern/src/components/AI/Chat/ToolExecutionBlock.tsx": "cc7bd548a418c3863486f0fe987c5c3110c2f6cdfa70b630b2ec05fb794d60d6", - "frontend-modern/src/components/AI/Chat/hooks/useChat.ts": "0b56b7a56e35d51ca96f0e126dd493b3164aa9e0ad4d8ae24bcf3af7a574b97c", - "frontend-modern/src/features/alerts/deliveryDiagnosisPresentation.ts": "bc3eadf8f790517b377430a42eb67e8fdfcadb53bb13b87cd971d6a9ab91d607" + "enterprise_base_sha": "3d9f4e3051d38027355a2a1f36b8c7f672a09b65", + "enterprise_content_sha256": { + "internal/investigation/orchestrator.go": "d56fd512dc47f8a2453559e89863da5d24dffc0a5977abb8f1ea7a82655cd9e8" }, + "binary_sha256": "24de8c9ea0020067d298c489f0d99a272b5ec4b00afab7a40dff55aabb244061", "routes": [ "/patrol", - "http://127.0.0.1:5198/qualification (isolated real OverviewTab)" + "/settings/pulse-intelligence/patrol", + "/settings/pulse-intelligence/provider", + "/" ], "viewports": [ { @@ -50,13 +60,15 @@ } ], "states": [ - "Five captured registered-tool archive cases: saved observation with historical status and nanosecond duration disclosure, unavailable archive, malformed archive, wrong resource and missing window. Failed reads remain failed.", - "Incoming Overview dispatch wording: ready/cooldown records show Dispatch requested, cooldown identifies next eligibility, and missing dispatch timestamp stays Notification pending. No delivery success is inferred.", - "Live canonical and legacy homelab timeline queries retain the same seven records and preserve the exact original alert fired/resolved records. The incidents endpoint returns active_count=null and active_count_status=not_measured. Provider refusal remains enforced." + "Live successful dependency diagnosis renders sanitized headings once, with distinct summaries retained. Paused history remains interactive and narrow finding controls wrap. Resolution does not infer manual review from an attention outcome.", + "Live broker refusal appears once alongside original diagnosis, needs_attention and no action. Historical successful query and failed ACTION_NOT_ALLOWED/NO_AGENT results retain actual input, output and status.", + "Controlled GET transcript fixture retains output with absent success status without inventing completion or failure. This proves rendering only.", + "Selected exact Gemini 3.8 Flash route, known $0.75/$3.75 rates, readiness for Watch only/Ask first and unassessed autonomous modes survive reload. Provider card read only. Browser interception blocks automatic preflight POST, so provider health was verified separately through the live API.", + "Incoming unchanged demo login source rechecked at 1440/390: loading, automatic demo payload, authenticated shell, logout/reload suppression, manual re-entry and rejected-login fallback. GET demo-presentation fixture with authorized local login exchange, no runtime demo configuration mutation." ], "interactions": [ - "At /patrol, hover/focus and Enter expansion, exact input/output comparison, deepest output scrolling, Space collapse, Escape, full reload, controlled session selection and reopening each result. Actual pixels inspected for successful and failed outcomes at desktop, intermediate and narrow widths.", - "The incoming real Overview component was exercised through scripts/check-alert-dispatch-copy.mjs at all three widths. Status text stays within the viewport and no page errors occurred. Actual pixels inspected. No delivery action invoked.", - "Final worker Pro build installed locally and managed restart recovered healthy. Private source/binary binding and receipts: tmp/patrol-archive-retirement/runtime-binding.json, browser/receipt.json and live-history-proof.json. Zero provider calls or faults in this proof. Controlled tool/session replay proves rendering only, not model diagnosis or server persistence." + "Activity, Finding options and history, All while Patrol paused. Review via hover/focus and Enter, unique aria-controls target, close returns focus, reopen, nested investigation disclosure open/closed, keyboard Enter/Space tool expansion/collapse and deep output scrolling. Actual desktop/intermediate/mobile pixels inspected after final source changes.", + "Patrol model settings and rates/readiness scrolled into view and reloaded at all three widths. No model setting change, action execution or approval request during browser proof.", + "Private receipts and interaction matrix under tmp/patrol-gemini-38/browser. Real run q-20260906-182952-6acfb739 remains failed: broker refused disconnected command agent, exact diagnosis retained and cleanup passed. No successful approval or recovery claim." ] } diff --git a/frontend-modern/src/api/__tests__/patrol.branchcov0718.test.ts b/frontend-modern/src/api/__tests__/patrol.branchcov0718.test.ts index 48740b4cf..299d82ea1 100644 --- a/frontend-modern/src/api/__tests__/patrol.branchcov0718.test.ts +++ b/frontend-modern/src/api/__tests__/patrol.branchcov0718.test.ts @@ -285,6 +285,16 @@ describe('patrol api — uncovered branch coverage', () => { role: 'assistant', content: 'logs in /var/log grew 40GB', reasoning_content: 'checked du output', + tool_calls: [ + { + id: 'read-1', + name: 'pulse_read', + input: { resource_id: 'container-1' }, + output: 'NO_AGENT', + success: false, + }, + { id: 'query-1', name: 'pulse_query', input: { action: 'metrics' } }, + ], timestamp: '2026-07-18T00:00:05Z', }, ], @@ -299,6 +309,9 @@ describe('patrol api — uncovered branch coverage', () => { expect(result).toEqual(envelope); expect(result.messages).toHaveLength(2); expect(result.messages[1]?.reasoning_content).toBe('checked du output'); + expect(result.messages[1]?.tool_calls?.[0]?.output).toBe('NO_AGENT'); + expect(result.messages[1]?.tool_calls?.[0]?.success).toBe(false); + expect(result.messages[1]?.tool_calls?.[1]?.success).toBeUndefined(); }); it('URL-encodes the finding id segment (separate from the messages suffix)', async () => { diff --git a/frontend-modern/src/api/patrol.ts b/frontend-modern/src/api/patrol.ts index e45416d2f..5354edf4d 100644 --- a/frontend-modern/src/api/patrol.ts +++ b/frontend-modern/src/api/patrol.ts @@ -6,6 +6,7 @@ import { apiFetchJSON } from '@/utils/apiClient'; import { arrayOrEmpty, promoteLegacyAlertIdentifier } from './responseUtils'; import type { InvestigationRecord } from './ai'; +import type { ToolCall } from './aiChat'; import type { ResourceCriticality } from './resourceOperatorState'; import type { PatrolActionReference } from '@/types/actionAudit'; import type { PatrolModelReadinessSnapshot } from '@/types/ai'; @@ -304,9 +305,8 @@ export interface ChatMessage { timestamp: string; } -export interface ChatToolCall { +export interface ChatToolCall extends ToolCall { id: string; - name: string; input: Record; } diff --git a/frontend-modern/src/components/AI/FindingsPanel.tsx b/frontend-modern/src/components/AI/FindingsPanel.tsx index 7c939f4ec..fc49a6150 100644 --- a/frontend-modern/src/components/AI/FindingsPanel.tsx +++ b/frontend-modern/src/components/AI/FindingsPanel.tsx @@ -194,6 +194,21 @@ export const FindingsPanel: Component = (props) => { const [filter, setFilter] = createSignal(props.filterOverride ?? 'active'); const [sortBy, setSortBy] = createSignal<'severity' | 'time'>('severity'); const [expandedId, setExpandedId] = createSignal(null); + let panelRoot: HTMLDivElement | undefined; + const closeReviewPanel = () => { + const findingId = expandedId(); + setExpandedId(null); + setManageOpenId(null); + if (findingId) { + queueMicrotask(() => { + panelRoot + ?.querySelector( + `button[aria-controls="${CSS.escape(`finding-${findingId}-details`)}"]`, + ) + ?.focus(); + }); + } + }; const [manageOpenId, setManageOpenId] = createSignal(null); const [actionLoading, setActionLoading] = createSignal(null); const [lastHashScrolled, setLastHashScrolled] = createSignal(null); @@ -1467,7 +1482,10 @@ export const FindingsPanel: Component = (props) => { manualControls.dismiss; return ( -
+
Triggered by alert{finding.alertType ? ` (${finding.alertType})` : ''} • Identifier{' '} @@ -2108,10 +2126,10 @@ export const FindingsPanel: Component = (props) => { }; return ( -
+
{/* Controls */} -
+
= (props) => {
Demo Mode
-
- Login with{' '} - - demo - {' '} - /{' '} - - demo - -
+ + Login with{' '} + + demo + {' '} + /{' '} + + demo + +
+ } + > +
+ Signing you in to the demo… +
+
diff --git a/frontend-modern/src/components/__tests__/Login.test.tsx b/frontend-modern/src/components/__tests__/Login.test.tsx index 0a7439005..8ec2e7d48 100644 --- a/frontend-modern/src/components/__tests__/Login.test.tsx +++ b/frontend-modern/src/components/__tests__/Login.test.tsx @@ -2,7 +2,7 @@ import { afterEach, describe, expect, it, vi, beforeEach } from 'vitest'; import { cleanup, fireEvent, render, screen, waitFor } from '@solidjs/testing-library'; import { Login } from '@/components/Login'; import loginSource from '@/components/Login.tsx?raw'; -import { STORAGE_KEYS } from '@/utils/localStorage'; +import { SESSION_STORAGE_KEYS, STORAGE_KEYS } from '@/utils/localStorage'; // Mock fetch globally const mockFetch = vi.fn(); @@ -182,7 +182,10 @@ describe('Login', () => { expect(mockFetch).not.toHaveBeenCalledWith('/api/security/status'); }); - it('shows demo credentials when session capabilities mark the runtime as demo mode', async () => { + it('shows demo credentials when the visitor has signed out of the demo', async () => { + // A sign-out marks the tab so the page does not sign the visitor straight + // back in; the printed credentials are the way back. + window.sessionStorage.setItem(SESSION_STORAGE_KEYS.DEMO_AUTO_LOGIN, 'suppressed'); const mockOnLogin = vi.fn(); const securityStatus = { hasAuthentication: true, @@ -196,6 +199,78 @@ describe('Login', () => { expect(await screen.findByText('Demo Mode')).toBeInTheDocument(); expect(screen.getAllByText('demo')).toHaveLength(2); + expect(mockFetch).not.toHaveBeenCalledWith('/api/login', expect.anything()); + expect(mockOnLogin).not.toHaveBeenCalled(); + }); + + it('signs the visitor in with the demo credentials when the runtime is in demo mode', async () => { + const mockOnLogin = vi.fn(); + mockFetch.mockResolvedValueOnce( + new Response(JSON.stringify({ success: true }), { + status: 200, + headers: { 'Content-Type': 'application/json' }, + }), + ); + const securityStatus = { + hasAuthentication: true, + hideLocalLogin: false, + presentationPolicy: { demoMode: true }, + }; + + render(() => ( + + )); + + await waitFor(() => expect(mockOnLogin).toHaveBeenCalledOnce()); + const loginCall = mockFetch.mock.calls.find(([url]) => url === '/api/login'); + expect(loginCall).toBeDefined(); + expect(JSON.parse((loginCall?.[1] as RequestInit).body as string)).toEqual({ + username: 'demo', + password: 'demo', + rememberMe: false, + }); + expect(window.sessionStorage.getItem(SESSION_STORAGE_KEYS.DEMO_AUTO_LOGIN)).toBe('attempted'); + }); + + it('falls back to the form when the demo sign-in is rejected', async () => { + const mockOnLogin = vi.fn(); + mockFetch.mockResolvedValueOnce( + new Response(JSON.stringify({ success: false, message: 'Invalid username or password' }), { + status: 401, + headers: { 'Content-Type': 'application/json' }, + }), + ); + const securityStatus = { + hasAuthentication: true, + hideLocalLogin: false, + presentationPolicy: { demoMode: true }, + }; + + render(() => ( + + )); + + expect(await screen.findByText('Invalid username or password')).toBeInTheDocument(); + expect(screen.getAllByText('demo')).toHaveLength(2); + expect(screen.getByRole('button', { name: /sign in to pulse/i })).toBeEnabled(); + expect(mockOnLogin).not.toHaveBeenCalled(); + }); + + it('does not sign in to the demo twice in one browser tab', async () => { + window.sessionStorage.setItem(SESSION_STORAGE_KEYS.DEMO_AUTO_LOGIN, 'attempted'); + const mockOnLogin = vi.fn(); + const securityStatus = { + hasAuthentication: true, + hideLocalLogin: false, + presentationPolicy: { demoMode: true }, + }; + + render(() => ( + + )); + + expect(await screen.findByText('Demo Mode')).toBeInTheDocument(); + expect(mockFetch).not.toHaveBeenCalledWith('/api/login', expect.anything()); }); it('restores the remembered username without storing a password', async () => { diff --git a/frontend-modern/src/components/patrol/InvestigationMessages.tsx b/frontend-modern/src/components/patrol/InvestigationMessages.tsx index 6b941d689..f5ede3cb5 100644 --- a/frontend-modern/src/components/patrol/InvestigationMessages.tsx +++ b/frontend-modern/src/components/patrol/InvestigationMessages.tsx @@ -10,6 +10,7 @@ import { getInvestigationMessages, formatTimestamp, type ChatMessage } from '@/a import { LoadingSpinner } from '@/components/shared/LoadingSpinner'; import { getInvestigationMessagesState } from '@/utils/patrolEmptyStatePresentation'; import { renderMarkdown } from '@/components/AI/aiChatUtils'; +import { ToolExecutionBlock } from '@/components/AI/Chat/ToolExecutionBlock'; // Compact variant of the Assistant chat's markdown styling, scaled for the // investigation thread's text-xs bubbles. @@ -110,16 +111,35 @@ export const InvestigationMessages: Component = (pro
{(tc) => ( -
- - {tc.name} - - 0}> -
-                                  {JSON.stringify(tc.input, null, 2)}
-                                
-
-
+ + + {tc.name} + + 0}> +
+                                      {JSON.stringify(tc.input, null, 2)}
+                                    
+
+ +
+                                      {tc.output}
+                                    
+
+
+ } + > + + )}
diff --git a/frontend-modern/src/components/patrol/InvestigationSection.tsx b/frontend-modern/src/components/patrol/InvestigationSection.tsx index 65b8257bd..d3d3e8f90 100644 --- a/frontend-modern/src/components/patrol/InvestigationSection.tsx +++ b/frontend-modern/src/components/patrol/InvestigationSection.tsx @@ -29,6 +29,7 @@ import { buildPatrolInvestigationRecordPresentation } from '@/features/patrol/pa import { LoadingSpinner } from '@/components/shared/LoadingSpinner'; import { MetadataBadge } from '@/components/shared/MetadataBadge'; import { InvestigationMessages } from './InvestigationMessages'; +import { renderMarkdown } from '@/components/AI/aiChatUtils'; import { notificationStore } from '@/stores/notifications'; import { aiIntelligenceStore } from '@/stores/aiIntelligence'; import type { InvestigationRecord } from '@/api/ai'; @@ -40,6 +41,9 @@ const INVESTIGATION_BADGE_PROPS = { shape: 'rounded', } as const; +const summaryClass = + 'text-sm prose prose-slate prose-sm dark:prose-invert max-w-none break-words prose-headings:my-2 prose-p:my-2 prose-pre:overflow-x-auto prose-code:break-all prose-code:before:content-none prose-code:after:content-none'; + interface InvestigationSectionProps { findingId: string; investigationStatus?: string; @@ -210,7 +214,11 @@ export const InvestigationSection: Component = (props
-

{investigationRecord().conclusion}

+

@@ -325,6 +333,7 @@ export const InvestigationSection: Component = (props = (props {/* Summary */} - -

{inv().summary}
+ +
{/* Tools used + turn count */} diff --git a/frontend-modern/src/components/patrol/__tests__/InvestigationMessages.test.tsx b/frontend-modern/src/components/patrol/__tests__/InvestigationMessages.test.tsx new file mode 100644 index 000000000..b1b8c6d3a --- /dev/null +++ b/frontend-modern/src/components/patrol/__tests__/InvestigationMessages.test.tsx @@ -0,0 +1,68 @@ +import { cleanup, fireEvent, render, screen } from '@solidjs/testing-library'; +import { afterEach, describe, expect, it, vi } from 'vitest'; +import InvestigationMessages from '../InvestigationMessages'; + +const getMessages = vi.hoisted(() => vi.fn()); +vi.mock('@/api/patrol', () => ({ + getInvestigationMessages: getMessages, + formatTimestamp: (value: string) => value, +})); + +afterEach(cleanup); + +describe('InvestigationMessages', () => { + it('retains merged failed tool evidence and exposes it through the shared disclosure', async () => { + getMessages.mockResolvedValue({ + messages: [ + { + id: 'turn-1', + role: 'assistant', + content: '', + timestamp: '2026-09-06', + tool_calls: [ + { + id: 'read-1', + name: 'pulse_read', + input: { resource_id: 'app-container-1' }, + output: '{"error":"NO_AGENT","message":"Command agent is not connected"}', + success: false, + }, + ], + }, + ], + }); + render(() => ); + const disclosure = await screen.findByRole('button', { name: /failed/i }); + expect(disclosure).toHaveAttribute('aria-expanded', 'false'); + fireEvent.keyDown(disclosure, { key: 'Enter' }); + expect(disclosure).toHaveAttribute('aria-expanded', 'true'); + expect(screen.getByText(/"error":\s*"NO_AGENT"/)).toBeVisible(); + fireEvent.keyDown(disclosure, { key: ' ' }); + expect(disclosure).toHaveAttribute('aria-expanded', 'false'); + }); + + it('does not infer completion from historical calls that have no recorded result status', async () => { + getMessages.mockResolvedValue({ + messages: [ + { + id: 'turn-2', + role: 'assistant', + content: '', + timestamp: '2026-09-06', + tool_calls: [ + { + id: 'query-1', + name: 'pulse_query', + input: { action: 'metrics' }, + output: 'Historical output', + }, + ], + }, + ], + }); + render(() => ); + expect(await screen.findByText('Historical output')).toBeVisible(); + expect(screen.queryByText('completed')).not.toBeInTheDocument(); + expect(screen.queryByText('failed')).not.toBeInTheDocument(); + }); +}); diff --git a/frontend-modern/src/components/patrol/__tests__/InvestigationSection.test.tsx b/frontend-modern/src/components/patrol/__tests__/InvestigationSection.test.tsx index b0a9ac023..2db7384f4 100644 --- a/frontend-modern/src/components/patrol/__tests__/InvestigationSection.test.tsx +++ b/frontend-modern/src/components/patrol/__tests__/InvestigationSection.test.tsx @@ -133,4 +133,46 @@ describe('InvestigationSection', () => { expect(screen.queryByText(/No investigation data available/)).not.toBeInTheDocument(); expect(screen.queryByText('systemctl restart workload.service')).not.toBeInTheDocument(); }); + + it.each([false, true])( + 'preserves readable evidence and distinct summaries (different: %s)', + async (different) => { + const conclusion = + '### Root cause\n\nThe cause is **unknown**.\n\n- The health check failed.\n- Logs are unavailable.'; + getInvestigationMock.mockResolvedValue({ + id: 'inv-markdown', + finding_id: 'finding-markdown', + session_id: 'session-markdown', + status: 'completed', + started_at: '2026-09-06T17:00:00Z', + turn_count: 2, + summary: different + ? '### Follow-up\n\nAdditional evidence remains unavailable.' + : conclusion, + } satisfies Investigation); + render(() => ( + + )); + await screen.findByRole('button', { name: 'Show investigation thread' }); + expect(screen.getAllByRole('heading', { name: 'Root cause' })).toHaveLength(1); + expect(screen.getByText('unknown').tagName).toBe('STRONG'); + expect(screen.getAllByRole('listitem')).toHaveLength(2); + expect(screen.queryByRole('heading', { name: 'Follow-up' }) !== null).toBe(different); + }, + ); }); diff --git a/frontend-modern/src/features/patrol/PatrolIntelligenceSurface.tsx b/frontend-modern/src/features/patrol/PatrolIntelligenceSurface.tsx index 92c5ee61b..6edec6a1e 100644 --- a/frontend-modern/src/features/patrol/PatrolIntelligenceSurface.tsx +++ b/frontend-modern/src/features/patrol/PatrolIntelligenceSurface.tsx @@ -225,9 +225,7 @@ export function PatrolIntelligenceSurface() { onToggle={(event) => setFindingsOpen(event.currentTarget.open)} > Finding options and history -
+

{ ]; keysToRemove.forEach((key) => localStorage.removeItem(key)); sessionStorage.clear(); + try { + sessionStorage.setItem(SESSION_STORAGE_KEYS.DEMO_AUTO_LOGIN, 'suppressed'); + } catch (_err) { + // Storage may be unavailable; the demo login page then simply signs in again. + } localStorage.setItem('just_logged_out', 'true'); aiChatStore.setEnabled(false); diff --git a/frontend-modern/src/utils/__tests__/aiFindingPresentation.coverage2.test.ts b/frontend-modern/src/utils/__tests__/aiFindingPresentation.coverage2.test.ts index 625b8ffa5..d9663ce54 100644 --- a/frontend-modern/src/utils/__tests__/aiFindingPresentation.coverage2.test.ts +++ b/frontend-modern/src/utils/__tests__/aiFindingPresentation.coverage2.test.ts @@ -995,19 +995,19 @@ describe('getFindingResolutionReason', () => { ).toBe('Resolved after investigation timeout now'); }); - it('returns "Resolved manually" for cannot_fix', () => { + it('does not infer manual resolution from cannot_fix', () => { expect( getFindingResolutionReason({ ...patrolBase, investigationOutcome: 'cannot_fix' }, 'now'), - ).toBe('Resolved manually now'); + ).toBe('Resolved now'); }); - it('returns "Resolved after manual review" for needs_attention', () => { + it('does not infer manual review from needs_attention', () => { expect( getFindingResolutionReason( { ...patrolBase, investigationOutcome: 'needs_attention' }, 'now', ), - ).toBe('Resolved after manual review now'); + ).toBe('Resolved now'); }); it('returns "Fix applied by Patrol" for fix_executed even when autoResolved is false', () => { diff --git a/frontend-modern/src/utils/aiFindingPresentation.ts b/frontend-modern/src/utils/aiFindingPresentation.ts index da2cfe6c5..00b4aae7f 100644 --- a/frontend-modern/src/utils/aiFindingPresentation.ts +++ b/frontend-modern/src/utils/aiFindingPresentation.ts @@ -1287,9 +1287,10 @@ export const getFindingResolutionReason = ( case 'timed_out': return `Resolved after investigation timeout ${resolvedTime}`; case 'cannot_fix': - return `Resolved manually ${resolvedTime}`; case 'needs_attention': - return `Resolved after manual review ${resolvedTime}`; + // An investigation outcome does not identify who later resolved the + // finding. Explicit operator resolution is handled above. + return `Resolved ${resolvedTime}`; default: return `Issue no longer detected ${resolvedTime}`; } diff --git a/frontend-modern/src/utils/localStorage.ts b/frontend-modern/src/utils/localStorage.ts index fe9161e82..2090b6ae4 100644 --- a/frontend-modern/src/utils/localStorage.ts +++ b/frontend-modern/src/utils/localStorage.ts @@ -142,6 +142,10 @@ export type LowPriorityNoticeOwner = 'github-star' | 'release-update'; export const SESSION_STORAGE_KEYS = { LOW_PRIORITY_NOTICE_OWNER: 'pulse-low-priority-notice-owner', + // Demo mode signs the visitor in once per browser tab. The value is + // 'attempted' after the login page has tried, or 'suppressed' after an + // explicit sign-out, so a visitor who signed out lands on the form. + DEMO_AUTO_LOGIN: 'pulse-demo-auto-login', } as const; /** diff --git a/internal/agentcapabilities/transcript.go b/internal/agentcapabilities/transcript.go new file mode 100644 index 000000000..4ccd48a47 --- /dev/null +++ b/internal/agentcapabilities/transcript.go @@ -0,0 +1,45 @@ +package agentcapabilities + +import "encoding/json" + +// TranscriptToolCall preserves a stored invocation and its observed result. +// Provider requests use ProviderToolCall instead of the product history shape. +type TranscriptToolCall struct { + ID string `json:"id"` + Name string `json:"name"` + Input map[string]interface{} `json:"input"` + Output string `json:"output,omitempty"` + Success *bool `json:"success,omitempty"` + ThoughtSignature json.RawMessage `json:"thought_signature,omitempty"` +} + +func (t TranscriptToolCall) NormalizeCollections() TranscriptToolCall { + providerCall := ProviderToolCall{ + ID: t.ID, + Name: t.Name, + Input: t.Input, + ThoughtSignature: t.ThoughtSignature, + }.NormalizeCollections() + t.ID = providerCall.ID + t.Name = providerCall.Name + t.Input = providerCall.Input + t.ThoughtSignature = providerCall.ThoughtSignature + if t.Success != nil { + success := *t.Success + t.Success = &success + } + return t +} + +// ProviderToolCall projects a stored Assistant transcript call back to the +// shared provider-facing shape, deliberately excluding in-app output/success +// display fields. +func (t TranscriptToolCall) ProviderToolCall() ProviderToolCall { + t = t.NormalizeCollections() + return ProviderToolCall{ + ID: t.ID, + Name: t.Name, + Input: t.Input, + ThoughtSignature: t.ThoughtSignature, + }.NormalizeCollections() +} diff --git a/internal/agentcapabilities/types_test.go b/internal/agentcapabilities/types_test.go index c5938ea3e..5b9371f2f 100644 --- a/internal/agentcapabilities/types_test.go +++ b/internal/agentcapabilities/types_test.go @@ -1,6 +1,7 @@ package agentcapabilities import ( + "encoding/json" "net/http" "slices" "strings" @@ -203,3 +204,46 @@ func TestNewToolGovernanceDescriptorAppliesSharedDefaults(t *testing.T) { t.Fatalf("descriptor approval summary = %q", descriptor.ApprovalSummary) } } + +// Stored result evidence and provider request arguments are different wire +// contracts. In particular, explicit failure cannot disappear through omitempty. +func TestTranscriptToolCallPreservesResultOutsideProviderRequests(t *testing.T) { + failed := false + call := TranscriptToolCall{ID: "read-1", Name: PulseReadToolName, Output: "NO_AGENT", Success: &failed}.NormalizeCollections() + failed = true + body, err := json.Marshal(call) + if err != nil { + t.Fatal(err) + } + var stored map[string]interface{} + if err := json.Unmarshal(body, &stored); err != nil { + t.Fatal(err) + } + if stored["output"] != "NO_AGENT" || stored["success"] != false || stored["input"] == nil { + t.Fatalf("stored result lost explicit failure or normalized input: %s", body) + } + requestBody, err := json.Marshal(call.ProviderToolCall()) + if err != nil { + t.Fatal(err) + } + var request map[string]interface{} + if err := json.Unmarshal(requestBody, &request); err != nil { + t.Fatal(err) + } + for _, key := range []string{"output", "success"} { + if _, exists := request[key]; exists { + t.Fatalf("provider request retained display-only %s: %s", key, requestBody) + } + } + unknownBody, err := json.Marshal(TranscriptToolCall{Name: PulseQueryToolName}.NormalizeCollections()) + if err != nil { + t.Fatal(err) + } + var unknown map[string]interface{} + if err := json.Unmarshal(unknownBody, &unknown); err != nil { + t.Fatal(err) + } + if _, exists := unknown["success"]; exists { + t.Fatalf("unknown historical status became a result: %s", unknownBody) + } +} diff --git a/internal/ai/chat/types.go b/internal/ai/chat/types.go index 2f307942c..4cc102d3f 100644 --- a/internal/ai/chat/types.go +++ b/internal/ai/chat/types.go @@ -135,38 +135,13 @@ func (m Message) ClientSafe() Message { return m } -// ToolCall represents a tool invocation -type ToolCall struct { - ID string `json:"id"` - Name string `json:"name"` - Input map[string]interface{} `json:"input"` - Output string `json:"output,omitempty"` - Success *bool `json:"success,omitempty"` - ThoughtSignature json.RawMessage `json:"thought_signature,omitempty"` -} +// ToolCall is the canonical result-bearing product transcript call. +type ToolCall = agentcapabilities.TranscriptToolCall func EmptyToolCall() ToolCall { return ToolCall{}.NormalizeCollections() } -func (t ToolCall) NormalizeCollections() ToolCall { - providerCall := agentcapabilities.ProviderToolCall{ - ID: t.ID, - Name: t.Name, - Input: t.Input, - ThoughtSignature: t.ThoughtSignature, - }.NormalizeCollections() - t.ID = providerCall.ID - t.Name = providerCall.Name - t.Input = providerCall.Input - t.ThoughtSignature = providerCall.ThoughtSignature - if t.Success != nil { - success := *t.Success - t.Success = &success - } - return t -} - // ToolCallFromProvider stores a provider-facing tool call in the richer // Assistant transcript shape used for in-app history. func ToolCallFromProvider(tc agentcapabilities.ProviderToolCall) ToolCall { @@ -183,19 +158,6 @@ func ToolCallFromProvider(tc agentcapabilities.ProviderToolCall) ToolCall { }.NormalizeCollections() } -// ProviderToolCall projects a stored Assistant transcript call back to the -// shared provider-facing shape, deliberately excluding in-app output/success -// display fields. -func (t ToolCall) ProviderToolCall() agentcapabilities.ProviderToolCall { - t = t.NormalizeCollections() - return agentcapabilities.ProviderToolCall{ - ID: t.ID, - Name: t.Name, - Input: t.Input, - ThoughtSignature: t.ThoughtSignature, - }.NormalizeCollections() -} - // ToolResult represents the result of a tool execution. It aliases the shared // Pulse Intelligence provider-result shape so stored Assistant transcripts and // provider turns do not drift on tool result JSON. diff --git a/internal/ai/cost/pricing.go b/internal/ai/cost/pricing.go index 7ac56bbe2..2f43d6b05 100644 --- a/internal/ai/cost/pricing.go +++ b/internal/ai/cost/pricing.go @@ -82,11 +82,17 @@ var providerPrices = map[string][]modelPrice{ flatPriceAsOf("anthropic/claude-opus-4.8", 5.00, 25.00, "2026-07-14"), flatPriceAsOf("anthropic/claude-sonnet-5", 2.00, 10.00, "2026-07-14"), flatPriceAsOf("deepseek/deepseek-v4-flash", 0.09, 0.18, "2026-07-14"), + // Introductory standard rates through 2026-12-31. Recheck when the + // published standard price changes on 2027-01-01. Batch/alias routes + // are deliberately not covered by this exact model ID. + flatPriceAsOf("google/gemini-3.8-flash", 0.75, 3.75, "2026-09-06"), flatPriceAsOf("nvidia/nemotron-3.5-lightning:free", 0, 0, "2026-08-14"), flatPriceAsOf("nvidia/nemotron-3-super-120b-a12b:free", 0, 0, "2026-08-14"), flatPriceAsOf("nvidia/nemotron-3-ultra-550b-a55b:free", 0, 0, "2026-08-15"), }, "gemini": { + // Standard introductory rates, verified 2026-09-06. Recheck 2027-01-01. + flatPriceAsOf("gemini-3.8-flash", 0.75, 3.75, "2026-09-06"), // Gemini Developer API standard paid-tier pricing, checked from // https://ai.google.dev/gemini-api/docs/pricing on 2026-06-04. flatPrice("gemini-3.5-flash*", 1.50, 9.00), diff --git a/internal/ai/cost/pricing_gemini_test.go b/internal/ai/cost/pricing_gemini_test.go new file mode 100644 index 000000000..6ea1ee23d --- /dev/null +++ b/internal/ai/cost/pricing_gemini_test.go @@ -0,0 +1,46 @@ +package cost + +import ( + "math" + "testing" +) + +func TestGemini38FlashReviewedRoutePricing(t *testing.T) { + for _, route := range []struct{ provider, model string }{ + {"gemini", "gemini-3.8-flash"}, + {"openrouter", "google/gemini-3.8-flash"}, + } { + t.Run(route.provider, func(t *testing.T) { + provider, model := ResolveProviderAndModel(route.provider, route.provider+":"+route.model, "gemini-3.8-flash") + if provider != route.provider || model != route.model { + t.Fatalf("usage lost the requested billing route: %s:%s", provider, model) + } + // Token counts from the live unhealthy-container qualification. + usd, known, price := EstimateUSD(provider, model, 16068, 778) + if !known || math.Abs(usd-0.0149685) > 1e-10 { + t.Fatalf("live usage estimate = %f, known=%t", usd, known) + } + if price.InputUSDPerMTok != 0.75 || price.OutputUSDPerMTok != 3.75 || price.AsOf != "2026-09-06" { + t.Fatalf("reviewed standard rates/date missing: %+v", price) + } + usd, known, _ = EstimateUSD(provider, model, 0, 0) + if !known || usd != 0 { + t.Fatalf("zero observed usage = %f, known=%t", usd, known) + } + }) + } +} + +func TestGemini38OpenRouterPricingDoesNotGuessVariantRates(t *testing.T) { + for _, model := range []string{ + "google/gemini-3.8-flash:batch", + "google/gemini-3.8-flash:free", + "google/gemini-3.8-flash-preview", + "google/gemini-3.8-flash-cyber", + "google/gemini-3.9-flash", + } { + if usd, known, _ := EstimateUSD("openrouter", model, 16068, 778); known || usd != 0 { + t.Errorf("unreviewed route %q received an estimate: %f, known=%t", model, usd, known) + } + } +} diff --git a/internal/ai/service.go b/internal/ai/service.go index 96f27f4c6..d3e0e8c6e 100644 --- a/internal/ai/service.go +++ b/internal/ai/service.go @@ -184,13 +184,12 @@ func (m ChatMessage) NormalizeCollections() ChatMessage { return m } -// ChatToolCall represents a provider-facing tool invocation in an API-facing -// chat message. It aliases the shared Pulse Intelligence provider-call shape so -// API chat history and provider turns do not drift on tool-call JSON. -type ChatToolCall = agentcapabilities.ProviderToolCall +// ChatToolCall retains observed output and result status in product history. +// Provider turns use the explicit ProviderToolCall projection. +type ChatToolCall = agentcapabilities.TranscriptToolCall func EmptyChatToolCall() ChatToolCall { - return agentcapabilities.EmptyProviderToolCall() + return ChatToolCall{}.NormalizeCollections() } // ChatToolResult represents the result of a tool invocation. It aliases the diff --git a/internal/ai/service_test.go b/internal/ai/service_test.go index 33d78b474..4acb88ec4 100644 --- a/internal/ai/service_test.go +++ b/internal/ai/service_test.go @@ -168,7 +168,7 @@ func TestChatMessage_UsesCanonicalEmptyCollections(t *testing.T) { Name: "diagnose", ThoughtSignature: json.RawMessage(`{"provider":"gemini"}`), } - var sharedProviderCall agentcapabilities.ProviderToolCall = sharedCall.NormalizeCollections() + var sharedProviderCall agentcapabilities.TranscriptToolCall = sharedCall.NormalizeCollections() if sharedProviderCall.ID != "call-1" || sharedProviderCall.Input == nil { t.Fatalf("shared chat tool call = %+v", sharedProviderCall) } diff --git a/internal/ai/tools/tools_propose.go b/internal/ai/tools/tools_propose.go index 07dfcc626..7588b14e1 100644 --- a/internal/ai/tools/tools_propose.go +++ b/internal/ai/tools/tools_propose.go @@ -178,7 +178,7 @@ func (e *PulseToolExecutor) executeProposeAction(ctx context.Context, args map[s return NewErrorResult(err), nil } return NewTextResult(fmt.Sprintf( - "Proposal recorded: capability %q on resource %q. It will be planned and routed for governed approval; nothing has executed. Conclude the investigation with your diagnosis.", + "Proposal recorded: capability %q on resource %q. The action broker still needs to validate it after this investigation. No action has been created or executed.", capabilityName, resourceID)), nil } diff --git a/internal/ai/tools/tools_query.go b/internal/ai/tools/tools_query.go index 53e63b504..2fbe1bc4d 100644 --- a/internal/ai/tools/tools_query.go +++ b/internal/ai/tools/tools_query.go @@ -2166,7 +2166,7 @@ func (e *PulseToolExecutor) registerQueryTools() { }, "resource_type": { Type: "string", - Description: "Resource type. For get/search, prefer canonical values: 'agent', 'vm', 'system-container', 'app-container', 'storage', 'physical-disk', and 'docker-host'. For get, 'node' resolves to 'agent'. For search, 'node' filters Proxmox nodes. Compatibility aliases 'system' and 'storage-pool' are still accepted. For config: 'vm', 'system-container', or supported API-backed 'app-container'.", + Description: "Resource type. For get/search, prefer canonical values: 'agent', 'vm', 'system-container', 'app-container', 'storage', 'physical-disk', and 'docker-host'. For get, 'node' resolves to 'agent'. For search, 'node' filters Proxmox nodes. Compatibility aliases 'system' and 'storage-pool' are still accepted. For config: 'vm', 'system-container', or TrueNAS 'app-container'. Docker and Podman app-container configuration reads are not supported. Their collected health, mounts, ports, and networks are available through get.", Enum: []string{"agent", "system", "vm", "system-container", "app-container", "node", "docker-host", "storage", "storage-pool", "physical-disk"}, }, "resource_id": { diff --git a/internal/api/ai_handlers_test.go b/internal/api/ai_handlers_test.go index 3eb28535d..9fa31ac08 100644 --- a/internal/api/ai_handlers_test.go +++ b/internal/api/ai_handlers_test.go @@ -3754,14 +3754,10 @@ func TestPatrolModelReadinessBudgetScalesWithRequestTimeout(t *testing.T) { assert.Equal(t, 4*600*time.Second+time.Minute, patrolModelReadinessBudget(&config.AIConfig{RequestTimeoutSeconds: 600})) } -// TestOrchestratorAndChatAdaptersMapTheSameMessageFields keeps the deliberate -// GetMessages mirror between orchestratorChatAdapter (ai_handlers.go) and -// chatServiceAdapter (chat_service_adapter.go) honest: both convert the same -// chat-service messages onto separate output contracts, and a field mapped by -// one must be mapped by the other. chatServiceAdapter routes through -// adaptChatMessage so its API-facing tool calls stay on the shared provider -// shape instead of hand-copying a local duplicate. -func TestOrchestratorAndChatAdaptersMapTheSameMessageFields(t *testing.T) { +// The adapters share base message fields but have distinct tool-call contracts. +// Orchestrator turns project provider arguments. Product history retains the +// observed result through the canonical result-bearing transcript type. +func TestOrchestratorAndChatAdaptersMapTheirMessageContracts(t *testing.T) { for _, tc := range []struct { file string fn string @@ -3793,7 +3789,7 @@ func TestOrchestratorAndChatAdaptersMapTheSameMessageFields(t *testing.T) { "Content:", "ReasoningContent:", "Timestamp:", - "tc.ProviderToolCall()", + "tc.NormalizeCollections()", "toolResult := *m.ToolResult", ".NormalizeCollections()", }, diff --git a/internal/api/chat_service_adapter.go b/internal/api/chat_service_adapter.go index 2e48a8f57..e84b647df 100644 --- a/internal/api/chat_service_adapter.go +++ b/internal/api/chat_service_adapter.go @@ -86,7 +86,7 @@ func adaptChatMessage(m chat.Message) ai.ChatMessage { Timestamp: m.Timestamp, } for _, tc := range m.ToolCalls { - msg.ToolCalls = append(msg.ToolCalls, tc.ProviderToolCall()) + msg.ToolCalls = append(msg.ToolCalls, tc.NormalizeCollections()) } if m.ToolResult != nil { toolResult := *m.ToolResult diff --git a/internal/api/chat_service_adapter_test.go b/internal/api/chat_service_adapter_test.go index 7effa9df5..60ae6ac67 100644 --- a/internal/api/chat_service_adapter_test.go +++ b/internal/api/chat_service_adapter_test.go @@ -3,7 +3,6 @@ package api import ( "context" "encoding/json" - "strings" "testing" "time" @@ -99,8 +98,8 @@ func TestChatServiceAdapter_GetMessages(t *testing.T) { require.NoError(t, err) } -func TestAdaptChatMessageUsesSharedProviderToolCallShape(t *testing.T) { - success := true +func TestAdaptChatMessagePreservesObservedToolResult(t *testing.T) { + success := false msg := adaptChatMessage(chat.Message{ ID: "msg-1", Role: "assistant", @@ -109,7 +108,7 @@ func TestAdaptChatMessageUsesSharedProviderToolCallShape(t *testing.T) { ToolCalls: []chat.ToolCall{{ ID: "call-1", Name: "diagnose", - Output: "in-app only", + Output: "NO_AGENT: command agent unavailable", Success: &success, ThoughtSignature: json.RawMessage(`{"provider":"gemini"}`), }}, @@ -121,7 +120,7 @@ func TestAdaptChatMessageUsesSharedProviderToolCallShape(t *testing.T) { }) require.Len(t, msg.ToolCalls, 1) - var shared agentcapabilities.ProviderToolCall = msg.ToolCalls[0] + var shared agentcapabilities.TranscriptToolCall = msg.ToolCalls[0] assert.Equal(t, "call-1", shared.ID) assert.Equal(t, "diagnose", shared.Name) assert.NotNil(t, shared.Input) @@ -131,8 +130,8 @@ func TestAdaptChatMessageUsesSharedProviderToolCallShape(t *testing.T) { text := string(payload) assert.Contains(t, text, `"input":{}`) assert.Contains(t, text, `"thought_signature":{"provider":"gemini"}`) - assert.False(t, strings.Contains(text, `"output"`), text) - assert.False(t, strings.Contains(text, `"success"`), text) + assert.Contains(t, text, `"output":"NO_AGENT: command agent unavailable"`) + assert.Contains(t, text, `"success":false`) require.NotNil(t, msg.ToolResult) var sharedResult agentcapabilities.ProviderToolResult = *msg.ToolResult diff --git a/internal/api/contract_test.go b/internal/api/contract_test.go index b04e57829..157789311 100644 --- a/internal/api/contract_test.go +++ b/internal/api/contract_test.go @@ -19839,7 +19839,7 @@ func TestContract_PulseMCPAdapterProjectsAgentCapabilitiesManifest(t *testing.T) chatTypesSrc := string(chatTypesSource) for _, fragment := range []string{ `func ToolCallFromProvider(tc agentcapabilities.ProviderToolCall) ToolCall`, - `func (t ToolCall) ProviderToolCall() agentcapabilities.ProviderToolCall`, + `type ToolCall = agentcapabilities.TranscriptToolCall`, `type ToolResult = agentcapabilities.ProviderToolResult`, } { if !strings.Contains(chatTypesSrc, fragment) { @@ -19853,8 +19853,8 @@ func TestContract_PulseMCPAdapterProjectsAgentCapabilitiesManifest(t *testing.T) } aiServiceSrc := string(aiServiceSource) for _, fragment := range []string{ - `type ChatToolCall = agentcapabilities.ProviderToolCall`, - `return agentcapabilities.EmptyProviderToolCall()`, + `type ChatToolCall = agentcapabilities.TranscriptToolCall`, + `return ChatToolCall{}.NormalizeCollections()`, `type ChatToolResult = agentcapabilities.ProviderToolResult`, `return agentcapabilities.ApprovalRequiredToolMarker(`, `return agentcapabilities.PolicyBlockedToolMarker(command, reason)`, @@ -20087,11 +20087,11 @@ func TestContract_PulseMCPAdapterProjectsAgentCapabilitiesManifest(t *testing.T) chatServiceAdapterSrc := string(chatServiceAdapterSource) for _, fragment := range []string{ `func adaptChatMessage(m chat.Message) ai.ChatMessage`, - `msg.ToolCalls = append(msg.ToolCalls, tc.ProviderToolCall())`, + `msg.ToolCalls = append(msg.ToolCalls, tc.NormalizeCollections())`, `toolResult := *m.ToolResult`, } { if !strings.Contains(chatServiceAdapterSrc, fragment) { - t.Errorf("chat service adapter must bridge messages through shared provider tool shapes; missing %s", fragment) + t.Errorf("chat service adapter must retain result-bearing transcript calls; missing %s", fragment) } } diff --git a/pkg/metrics/store_large_seed_test.go b/pkg/metrics/store_large_seed_test.go new file mode 100644 index 000000000..e3e585436 --- /dev/null +++ b/pkg/metrics/store_large_seed_test.go @@ -0,0 +1,94 @@ +package metrics + +import ( + "fmt" + "path/filepath" + "testing" + "time" +) + +// Exercise the ingestion shape from the workloads-summary benchmark without +// its HTTP, reflection or monitor fixtures. A historical benchmark crashed in +// SQLite during the second synchronous seed; this is a diagnostic invariant, +// not a reproducer or a claim that the unexplained crash has been repaired. +func TestStoreLargeSummarySeedSurvivesReopen(t *testing.T) { + cfg := DefaultConfig(t.TempDir()) + cfg.DBPath = filepath.Join(filepath.Dir(cfg.DBPath), "summary-seed.db") + cfg.FlushInterval = time.Hour + cfg.WriteBufferSize = 10_000 + store, err := NewStore(cfg) + if err != nil { + t.Fatal(err) + } + t.Cleanup(func() { + if store != nil { + if err := store.Close(); err != nil { + t.Error(err) + } + } + }) + base := time.Now().Add(-4 * time.Hour).UTC().Truncate(time.Second) + for _, group := range []struct { + kind string + count int + }{{"vm", 30}, {"container", 20}, {"dockerContainer", 20}} { + batch := make([]WriteMetric, 0, group.count*5*240) + for r := 0; r < group.count; r++ { + for _, metric := range []string{"cpu", "memory", "disk", "netin", "netout"} { + for p := 0; p < 240; p++ { + batch = append(batch, WriteMetric{ + ResourceType: group.kind, ResourceID: fmt.Sprintf("%s-%d", group.kind, r), + MetricType: metric, Value: float64((r + p) % 100), + Timestamp: base.Add(time.Duration(p) * time.Minute), Tier: TierMinute, + }) + } + } + } + store.WriteBatchSync(batch) + } + check := func() { + t.Helper() + var count int + if err := store.db.QueryRow("SELECT COUNT(*) FROM metrics WHERE tier = ?", string(TierMinute)).Scan(&count); err != nil { + t.Fatal(err) + } + if count != 84_000 { + t.Fatalf("persisted minute rows = %d, want 84000", count) + } + rows, err := store.db.Query("PRAGMA integrity_check") + if err != nil { + t.Fatal(err) + } + defer rows.Close() + var results int + for rows.Next() { + var result string + if err := rows.Scan(&result); err != nil { + t.Fatal(err) + } + if result != "ok" { + t.Errorf("integrity_check: %s", result) + } + results++ + } + if err := rows.Err(); err != nil { + t.Fatal(err) + } + if results != 1 { + t.Fatalf("integrity_check returned %d rows, want one ok", results) + } + } + check() + if err := store.Close(); err != nil { + t.Fatal(err) + } + store = nil + store, err = NewStore(cfg) + if err != nil { + t.Fatal(err) + } + if err := store.WaitForMaintenance(10 * time.Second); err != nil { + t.Fatal(err) + } + check() +} diff --git a/scripts/release_control/control_plane.py b/scripts/release_control/control_plane.py index 1f3aa8188..74a2c2062 100644 --- a/scripts/release_control/control_plane.py +++ b/scripts/release_control/control_plane.py @@ -442,7 +442,18 @@ def legacy_release_line_for_version( reverse=True, ) for line in legacy_release_lines: - if normalized_version.startswith(line["version_prefix"]): + prefix = line["version_prefix"] + # A complete patch version binds only that version, not e.g. 6.4.40. + # Trailing-dot prefixes continue to bind the whole minor/major line. + if prefix.endswith("."): + matches = normalized_version.startswith(prefix) + else: + matches = ( + normalized_version == prefix + or normalized_version.startswith(prefix + "-") + or normalized_version.startswith(prefix + "+") + ) + if matches: return line return None diff --git a/scripts/release_control/control_plane_audit_test.py b/scripts/release_control/control_plane_audit_test.py index 094e5e1e3..4b4a2c8c3 100644 --- a/scripts/release_control/control_plane_audit_test.py +++ b/scripts/release_control/control_plane_audit_test.py @@ -1,3 +1,8 @@ +import os +from pathlib import Path +import re +import tempfile +import textwrap import shlex import subprocess import sys @@ -239,6 +244,46 @@ class ControlPlaneAuditTest(unittest.TestCase): "pulse/release-5.1.25", ) + def test_forward_patch_train_uses_actual_control_plane(self) -> None: + for version in ("6.4.4-beta.1", "v6.4.4-beta.2", "6.4.4-rc.1", + "6.4.4", "v6.4.4+build.1", "6.4.3-rc.1", "6.4.3"): + with self.subTest(version=version): + self.assertEqual(release_branch_for_version(version), "release/v6.4") + for version in ("6.4.1", "6.4.2", "6.4.5-beta.1", "6.4.40-beta.1", + "6.4.30", "6.3.20", "6.6.0-beta.1"): + with self.subTest(version=version): + self.assertEqual(release_branch_for_version(version), "main") + self.assertEqual(release_branch_for_version("6.5.10-beta.1"), "release/v6.5") + + def test_forward_patch_workflow_branch_contract(self) -> None: + # Execute the real branch-policy shell only, never dispatch a workflow. + for workflow in ("create-release.yml", "release-dry-run.yml"): + content = (REPO_ROOT / ".github/workflows" / workflow).read_text() + match = re.search( + r"(?ms)^ - name: Resolve required release branch\n" + r".*?^ run: \|\n((?: [^\n]*\n|\n)+)", content + ) + self.assertIsNotNone(match) + script = textwrap.dedent(match.group(1)) + for branch in ("main", "release/v6.4"): + with self.subTest(workflow=workflow, branch=branch), tempfile.TemporaryDirectory() as tmp: + output = os.path.join(tmp, "output") + result = subprocess.run( + ["bash", "-euo", "pipefail", "-c", script], + cwd=REPO_ROOT, capture_output=True, text=True, + env={**os.environ, "GITHUB_OUTPUT": output, + "VERSION_INPUT": "6.4.4-beta.1", + "WORKFLOW_OUTPUT_1": "6.4.4-beta.1", + "WORKFLOW_OUTPUT_2": branch}, + ) + rejects = workflow == "create-release.yml" and branch == "main" + self.assertEqual(result.returncode, 1 if rejects else 0, result.stderr) + if rejects: + self.assertIn("must run from release/v6.4", result.stdout) + else: + self.assertRegex(Path(output).read_text(), + r"^required_branch<<([^\n]+)\nrelease/v6.4\n\1\n$") + def test_audit_flags_stale_active_target(self) -> None: report = audit_control_plane_payload( VALID_PAYLOAD,