mirror of
https://github.com/rcourtman/Pulse.git
synced 2026-09-23 03:33:53 +00:00
5dcdbfabf08f84f37513be3ec81ae112e7ff8e15
2756 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
a0b3bc7ed3 |
Record user-chat token usage to the cost ledger
chat.Service.ExecuteStream was a long-standing cost-ledger gap: the
agentic loop accumulated token counts via stream callbacks (see
GetTotalInputTokens / GetTotalOutputTokens in agentic_control.go)
and surfaced them in the SSE done envelope to the frontend, but
nothing on the server side recorded a cost.UsageEvent. Patrol,
discovery, QuickAnalysis, and the report narrators all record; only
chat — the bulk of AI token spend — did not. The operator's AI
usage dashboard was therefore understating cost dramatically.
Found while extending the cost-recording mindset across subpackages
after fixing QuickAnalysis (
|
||
|
|
113190a920 |
Skip agent-sourced alerts in node-presence cleanup
CleanupAlertsForNodes removes alerts whose Node isn't in existingNodes. That map is built upstream only from state.Nodes (Proxmox nodes) and state.PBSInstances (PBS instances) — no agent resources. Agent-sourced alerts (Unraid, standalone Linux hosts, TrueNAS, anything reached via Pulse Agent) typically have Node="" or an agent UUID, so they fall into the cleanup branch every cycle. Then the next poll re-creates the alert as new, calls AddAlert, and appends a fresh history row. Observed in the wild: 3,980 alert history entries in 7 days, with the same canonical alert ID (e.g. "Unraid array running without parity protection") appearing every 30 seconds. Adds a carve-out for ResourceID prefixes starting with "agent:", matching the existing pattern for "docker-" / "docker:" and "pbs-" / "pbs-offline". Locks in the behaviour with a new subtest that mixes agent-sourced and Proxmox-sourced alerts and asserts that only the legitimately stale Proxmox alert is removed. |
||
|
|
99c499ade7 |
Repair orphan tool_calls in convertToProviderMessages
Defense-in-depth for the malformed-history bug pattern. The
Patrol fix made patrol-main runs stateless, but Assistant
chat sessions are inherently multi-turn and must keep their
history. Any chat session that ends mid-tool-call — network
drop, ctx timeout, browser crash, uncaught panic, any
interrupt that fires between "model emits tool_calls" and
"agentic loop appends all tool results" — leaves the
persisted session with orphan tool_call_ids. The next message
that loads this history is rejected with the same provider
error that flapped Patrol for 33 days:
An assistant message with 'tool_calls' must be followed by
tool messages responding to each 'tool_call_id'.
For Patrol this was fixable by ignoring the session. For
Assistant it isn't; the conversation context is the product.
convertToProviderMessages now ends with a repairOrphanToolCalls
pass that scans every assistant message with tool_calls and
inserts synthetic is_error tool result messages immediately
after the assistant turn for any tool_call_id that has no
matching downstream result. The synthetic content is marked
is_error=true and explains the interruption so the model can
retry the same call or proceed without that data — preserving
conversational continuity while satisfying the provider's
structural-validity check.
This guards every conversation that crosses
convertToProviderMessages, not just Assistant chat. If Patrol
ever changes back to loading session history, the same safety
net applies. If a new entry point appears for some other LLM
flow, it gets the repair for free.
Three tests guard the boundary:
- Orphan injection (3 tool_calls, only 1 result → 2
synthetic results, marked is_error with interrupted
explanation, ordering preserved)
- Clean no-op (all tool_calls fulfilled → no synthetic
messages, no is_error pollution)
- Existing truncation test still passes (assistant message
with both tool_calls and own tool_result → no repair
needed, tool_call_id matches in same message)
ai-runtime contract updated.
|
||
|
|
e657f6ace9 |
Suppress assessment error penalty after trailing-success recovery
The overall-health "Recent Patrol errors" coverage factor in
summarizeRecentPatrolCoverage was anchoring the score to a
stale ratio: it counted errors across the last 10 runs without
weighting recency. After Pulse fixed two compounding Patrol
bugs today, four consecutive successful runs (50+ tool calls
each) followed six earlier failures. The assessment kept
showing C/65 with the prediction "most recent Patrol runs
encountered errors (6 of 10)" — directly contradicting the
fact that *every* recent run had succeeded.
Operators reading that score would conclude Pulse Patrol is
still broken. It isn't. The fix dragged the grade.
This commit adds a recovery-suppression check: count trailing
successful full Patrol runs from the most-recent end of the
window (GetAll returns newest-first), skipping non-full runs.
When three or more consecutive trailing successes exist —
roughly a 9-hour clean stretch at the default 3-hour cadence —
the error penalty drops entirely. The score reflects current
reality.
Three is conservative: a single recovery run could be a
transient win; three consecutive demonstrate the underlying
fix is sticking. Below the threshold, the existing ratio-tiered
penalty still applies so partially-recovered states still
register.
Two tests guard the boundary:
- 6 historical errors + 3 trailing successes → no coverage
factor (suppressed)
- 6 historical errors + 2 trailing successes → coverage
factor remains (recovery incomplete)
Live verified after this commit lands: the assessment that's
been stuck at C/65 since the malformed-history fix will
recompute to A/B grade as soon as the trailing 3 successful
runs are recognized by the same recent-runs query.
ai-runtime contract updated.
|
||
|
|
4dff26f728 |
Emit structured telemetry on reporting and summarize invocations
The reporting feature now ships across two surfaces (PDF/CSV export and pulse_summarize chat tool) and three modes (single-resource, fleet, summarize). Without usage telemetry we can't tell whether the work earns its place — operator demand, AI-vs-heuristic adoption, range/format preferences are all invisible. Stops further feature investment from being pure speculation. Three new info-level log events, structured so an agent can grep transcripts and group by dimension without a separate metrics pipeline (matches the "agent owns ops analysis, human gets outcomes" posture in MEMORY.md): reporting.single.generated — single-resource PDF/CSV reporting.fleet.generated — multi-resource fleet PDF/CSV reporting.summarize.invoked — pulse_summarize chat tool (both modes) Common dimensions: org_id, format/action, range, ai_configured, findings_configured, window_start/end. Single-resource adds resource_type + metric_type + bytes; fleet adds resource_count + bytes; summarize adds resource_type + resource_count (fleet mode) + narrative_source (so we can audit AI-fallback rate). Includes rangeLabel() helper that maps a window to the canonical catalog range token (24h/7d/30d) with a 1h tolerance, falling back to "<hours>h" so non-standard windows still group. Tested. TestReportingTelemetryEventNames pins the canonical event names as a contract — an agent grepping logs depends on them being stable; changing them silently would break audit tooling on the consumer side. The reporting engine already logs the resolved narrative source (heuristic/ai) at debug level via the existing "Generating report" line, useful for diagnosing why a specific report fell back. Kept at debug; the new info-level events cover the operator surface. |
||
|
|
03463c1bfe |
Thread per-tenant AI narrators into pulse_summarize via chat session
v1 of pulse_summarize (
|
||
|
|
7a7b3c9d30 |
Gate LLM patrol_resolve_finding on deterministic verifier for event findings
Backup-failed was flapping detected → auto-resolved → re-detected
ten times in a single day. Each cycle the LLM saw "PBS backups
look healthy in my current snapshot" during a Patrol pass, called
patrol_resolve_finding(backup-failed), and the adapter at
patrol_findings.go:985 called Resolve(findingID, true) directly —
no category check, no evidence verification.
The contract docs at findings.go:52-67 explicitly say event /
persistent categories (backup, reliability, security, general)
"stay active until explicitly resolved — either by the LLM calling
patrol_resolve_finding with evidence, or by operator action." That
"with evidence" was never enforced.
This commit enforces it. The adapter now checks two conditions
before honoring an LLM resolve:
- finding.Category does NOT support stale-auto-resolve (per the
contract function CategorySupportsStaleAutoResolve), AND
- a deterministic verifier exists for finding.Key (currently
smart-failure and backup-failed)
When both are true, the adapter runs VerifyFixResolved on the
finding's resource. If the verifier still detects the failure
signal, the LLM gets an error explaining why the resolve was
rejected and that the underlying issue must be fixed first. If
the verifier confirms the signal has cleared, the resolve
proceeds with grounded evidence.
Categories that support stale-auto-resolve (performance, capacity)
bypass the gate entirely — the LLM can resolve them based on
absence per the existing contract. Keys without a verifier also
fall through to current behavior so we don't block resolves for
categories we haven't built verifiers for yet.
New PatrolService.hasDeterministicVerifierForKey() helper keeps
the gate's verifier list in lockstep with the switch in
verifyFixDeterministically.
Tests cover the three branches:
- performance category → gate skipped, resolve proceeds
- reliability + no verifier → gate falls through, resolve proceeds
- hasDeterministicVerifierForKey for known and unknown keys
ai-runtime contract updated.
|
||
|
|
1fe5d6853f |
Expose reporting synthesis to Assistant via pulse_summarize tool
The reporting synthesis layer (observations, recommendations, outliers, period comparison) shipped trapped behind the PDF/CSV export. Operators who chat with Assistant could not ask "what's been happening with pve1 this week" — the data path existed but had no non-PDF surface. This commit adds a single new tool, pulse_summarize, that wraps the engine's non-rendering entry points (NarrativeFor / FleetNarrativeFor) so that question gets answered in chat. The tool takes an action parameter (resource | fleet) and routes accordingly: - resource mode requires resource_type + resource_id and returns the same Narrative the single-resource report carries (health status, observations, recommendations, period comparison). - fleet mode requires resource_type + a comma-separated resource_ids string (PropertySchema does not currently support array items, and CSV is LLM-friendly enough) and returns the FleetNarrative (outliers, patterns, recommendations). Capped at the same multi-report ceiling (50) as the API endpoint. The tool is read-only — no control level requirement, no approval gate — and uses the global reporting engine the rest of the app already shares. Returns a JSON envelope so chat can render it or hand it back to the model for follow-up framing. v1 ships with heuristic narrative only. The AI narrator wiring through the chat session (Narrator/FleetNarrator/FindingsProvider threaded via chat.Config -> tools.ExecutorConfig -> PulseToolExecutor) is a focused follow-up; it lets the same tool inherit the per-tenant AI service the report PDF endpoint already uses. The seam is already in place because NarrativeFor/FleetNarrativeFor take an optional narrator on the request — v1 passes nil, v2 populates it. |
||
|
|
20df3dcd2c |
Let a valid bootstrap token authorize initial setup from any origin
The loopback gate from
|
||
|
|
e32d4ede44 |
Expose engine narrative entry points for non-rendering callers
The reporting engine's synthesis layer was reachable only through Generate/GenerateMulti, which always rendered PDF or CSV. Pulse Assistant needs the same retrospective synthesis (per-resource summary, fleet outliers, period comparison) in a form it can present in chat, not as a downloaded artifact. Add two non-rendering entry points to the Engine interface: NarrativeFor(req MetricReportRequest) (*Narrative, error) FleetNarrativeFor(req MultiReportRequest) (*FleetNarrative, error) Both run the same query path and the same narrator resolution as their rendering counterparts (heuristic by default, AI when the request supplies a narrator, fail-closed-to-heuristic on any narrator error) and return the structured narrative without invoking the fpdf/csv output stage. Test stubs in pkg/reporting and internal/api are updated to implement the extended interface. These are the seams the upcoming pulse_summarize Assistant tools wrap to answer questions like "what's hot on pve1 this week" or "where should I look across my fleet" without round-tripping through report generation. Same synthesis layer, no PDF involved. Also fixes a pre-existing flake in TestEngineGenerate_UsesSuppliedNarrator (metrics writes are async; the first Generate sometimes ran before the raw tier flushed). Wrapped in the same eventually-pattern used by the prior-period and findings-provider tests. |
||
|
|
41155a7968 |
Forbid the report narrators from acting as parallel detectors
The single-resource and fleet narrator prompts both grounded their claims in structured data, but neither prevented the model from classifying observations at warning or critical severity based on metric inference alone. That left a subtle gap: an AI narrator noticing memory creep across a window could promote it to warning even when Patrol — the canonical detection layer — had not flagged it. That competes with Patrol rather than summarizing its work, and it lets the report PDF silently shadow-classify in a way that diverges from the findings store. Add an explicit detection-boundary instruction to both prompts: warning or critical severity may only be assigned when backed by a Patrol finding, an alert, or a hard-threshold breach visible in the input (cpu max > 90, memory avg > 85, disk avg > 85, failed or high-wear disks, storage pools at >= 90%). Patterns the model sees in metric data without that backing are constrained to info severity. Recommendations follow the same rule. The narrative remains a retrospective summary of Patrol's classified state, not a parallel classifier. This is a prompt-only change. The deterministic data surface and the heuristic fallback narrator are unaffected; the heuristic narrators already classify only on the same threshold rules listed above, so the AI narrator is now constrained to the same evidence surface its fallback uses. |
||
|
|
08491b9f48 |
Fix QuickAnalysis cost recording and audit the pattern
Service.Narrate ( |
||
|
|
c6a786a36d |
Broaden regression-counter reset to include LLM-driven cycles
Initial detector (commit
|
||
|
|
942f9ca0f5 |
Reset regression counters polluted by bogus auto_resolve cycles
The Backup failed finding on the live preview showed "regressed 6×"
when the actual regression count of genuine recurrences was at
most 1 or 2 — the rest were the system fighting itself, driven by
the absence-based auto_resolve paths that were gated (category
whitelist) or removed (alert-mirror rip) earlier in this branch.
Counter stayed sticky after those fixes landed, so the trust strip
and finding badges still surfaced the inflated number.
FindingsStore.SetPersistence load pass now scans each active
finding's lifecycle for the two known bogus-signature auto_resolved
reasons ("No longer detected by patrol", "Resource no longer
exists in infrastructure"). If found, RegressionCount is reset to
0 and LastRegressionAt is cleared, and a regression_counter_reset
lifecycle event is appended so the migration is idempotent. A
finding that already has a regression_counter_reset event is left
alone; any regressed events that accrued after the reset are
genuine and stand.
findingHasBogusAutoResolveCycle returns true only when the
lifecycle contains a bogus auto_resolved and no prior reset event,
so the function is the single point of truth for the migration
decision and is straightforward to test. Test covers three cases:
finding with bogus signature gets reset, finding with empty-message
auto_resolved (LLM-driven, legitimate) keeps its counter, finding
already migrated is not re-reset.
Updates ai-runtime Current State to document the second migration
on top of the alert-mirror retirement.
|
||
|
|
590671ffbb |
Retire legacy "Active alert detected" findings on load
The previous commit removed the detectAlertSignals path so no NEW alert-mirror findings are emitted, but the findings already persisted from earlier builds stay in the store indefinitely — nothing cleans them up (reconcileStaleFindings is gated on performance/capacity categories, the LLM resolves them just to have them re-detected next run except now the deterministic emitter is gone so re-detection can't happen, but they're left sitting as active findings draining the trust strip and score). FindingsStore.SetPersistence now runs a one-shot retirement pass on load: any active finding with title "Active alert detected", source ai-analysis, and category general is auto-resolved with reason "Patrol no longer mirrors alerts; the Alerts page is the canonical surface for currently-firing alerts." The pass appends an auto_resolved lifecycle event so the retirement is auditable, syncs the loop state to resolved, and schedules a save so the cleanup persists. Idempotent: after the first load with this code, no findings match the signature so the pass is a no-op. Defensive: the signature requires all three fields (title + source + category) to match before retiring, so an operator-authored finding that happens to share the title is left untouched. Test covers the mirror case, the matching-title-but-foreign-source case (must NOT retire), and an unrelated active finding (must NOT retire), plus verifies the retired state persists back through the persistence layer. Updates ai-runtime Current State to record the migration path. |
||
|
|
271d12ecab |
Stop mirroring alerts into Patrol findings
The deterministic signal pipeline ran the pulse_alerts tool output through detectAlertSignals and produced a SignalActiveAlert for every firing alert, which Patrol then materialized as an "Active alert detected" finding (source: ai-analysis, category: general). The system prompt at the top of patrol_ai.go explicitly tells the LLM not to duplicate alerts — but the deterministic emitter was duplicating them anyway, behind the LLM's back. Symptoms observed in the wild: - 9 active "Active alert detected" findings in Patrol, every one a duplicate of an existing alert already on the Alerts page. - The LLM, doing what the prompt told it, resolved each mirrored finding via patrol_resolve_finding. Next run the alert was still firing and Patrol re-emitted the signal → finding regressed. Lifecycle showed several auto_resolved → re-detected → regressed cycles per finding within hours. - Health score dragged down by issues the operator already saw on the Alerts page, with no operator action possible from Patrol that wasn't already available from Alerts. Rip detectAlertSignals entirely, remove the pulse_alerts case from the signal-extraction switch, drop SignalActiveAlert plus its key / title / recommendation entries. Convert the prior TestDetectSignals_ActiveAlert into a regression guard that locks in the no-mirror behavior. Updates the ai-runtime subsystem Current State to record the decision: Patrol does not duplicate the Alerts surface; alerts own their own lifecycle, surface, and acknowledgement model. |
||
|
|
d4463a615c |
Add fleet-level AI narrative for multi-resource reports
The single-resource AI narrative landed in
|
||
|
|
b84b87d8d8 |
Record cost events for AI report narration
Service.Narrate consumed provider tokens without recording a cost.UsageEvent, so AI-narrated reports were invisible in the operator cost ledger. Every other Service call site in the AI runtime records cost; the narrator omitted it. Mirror the QuickAnalysis/chat pattern: capture the cost store under the read lock alongside the provider/cfg snapshot, and after provider.Chat returns, record a UsageEvent labelled report_narrative with the resource type/id as the target. Recording happens before parsing so a failed-but-billed call (e.g. provider returned malformed JSON) still appears in the ledger — the operator was billed regardless of whether we could use the response. The use_case string lifts to a package-level constant so the budget gate (enforceBudget), the cost label, and the dashboard taxonomy all reference one identifier. |
||
|
|
27bd31684a |
Log the underlying error on audit list 500s
HandleListAuditEvents dropped the Query/Count error before writing the 500, so a user hitting "Failed to fetch audit events" produced no server-side log line — diagnosing the failure was impossible without a local repro. Log the error with the org ID so the next instance is findable. Doesn't change the user-facing response. |
||
|
|
9d1d24bdf1 |
Fix silent fingerprint loss for LXC and VMs
processFingerprint ran v.Index(i).Interface() then reflect.ValueOf(item), dropping addressability and making reflect.Call panic on every iteration with "Container as type *Container". The defer/recover in collectFingerprints swallowed it, so LXC and VM fingerprints never landed in the store — change-detection and discovery for those resource types have been broken since v6. Pass the slice element's address straight through (.Addr()) so the generator's pointer receiver gets the right type. Add a regression test that fails if anyone goes back through .Interface(). |
||
|
|
a343ba0126 |
Tier patrol-coverage health impact by recent-error ratio
The overall health score chip read "A · 95/100" while the same
assessment card said "Coverage incomplete · Recent Patrol runs
encountered errors · Verify full coverage." Those messages
contradicted because the "recent errors but had a successful run"
coverage factor used a flat -10 penalty regardless of the error
ratio. With one successful manual run among many failed startup
runs, the math stayed in the A band and the score directly
undermined the warning it sat next to.
The factor now tiers the impact by ratio of errored runs to
relevant runs in the scoring window:
>50% errored → -30, "Most recent Patrol runs encountered errors
(N of M); the current health summary is not
reliable until coverage stabilizes."
>25% errored → -20, "Recent Patrol runs encountered errors
(N of M), so the current health summary
may be incomplete."
else → -10, original light-tier description.
A single transient error stays in grade A; dominant-error periods
drop out of A so the grade matches the warning. Adds two tests
covering both ends of the new tiering and updates the ai-runtime
subsystem contract Current State section.
|
||
|
|
b2bd9d1147 |
Replace heuristic report narrative with optional AI-generated layer
Performance reports rendered the Executive Summary, Observations, and Recommendations sections from inline threshold rules in pdf.go. That narrative looked intelligent but was static templating against alert counts and metric percentiles, which felt off-brand alongside Patrol and Pulse Assistant. Introduce a Narrator interface in pkg/reporting and a FindingsProvider counterpart that the engine consults at report time. The heuristic rules are lifted into HeuristicNarrator unchanged so the deterministic fallback still produces the same observations and recommendations. The engine now also queries the comparable prior period and threads its aggregate stats through the narrator so deltas can be expressed. internal/ai.Service implements both interfaces via report_narrator.go (single-turn JSON call grounded in the structured ReportData payload, falling back to the heuristic on any error/timeout) and report_findings.go (Patrol findings whose lifecycle overlaps the report window). The reporting handler resolves the per-tenant AI service when it is configured and supplies it in the request; absent configuration, reports look identical to the prior heuristic output. Charts, stats tables, alert lists, storage and disk sections stay deterministic — sysadmins can verify every AI claim against the data tables next to it. The PDF renders the AI prose between the health card and Quick Stats, adds a Period-over-period section after Recommendations, and prints a provenance footer when the narrative came from the assistant. ai-runtime.md and api-contracts.md updates land in a follow-up commit on this branch; agent-lifecycle / performance-and-scalability / storage-recovery have no contract delta from this change (router.go is referenced in their Extension Points but their semantics are unchanged). |
||
|
|
b44d5892f4 |
Gate stale-finding auto-resolve on category whitelist
reconcileStaleFindings was auto-resolving any seeded finding that the LLM didn't re-report in a successful run. The function's own comment acknowledges the LLM doesn't reliably use patrol_resolve_finding, so this was built as a cleanup pass — but it cannot tell the difference between "LLM correctly recognized this is fixed" and "LLM forgot to re-mention it." For findings that represent discrete events or persistent states (a backup task that failed, a service that crashed, a security vulnerability that was found, a configuration error), absence in a Patrol report is not evidence that the issue has cleared. The result was bogus auto_resolved → re-detected → regressed cycles, observed in the wild as "Backup failed" regressing 4× over 6 hours and "Provider analysis error" regressing 271×. Those bogus auto-resolutions also inflated the trust strip with fictional auto-resolved credit. CategorySupportsStaleAutoResolve in findings.go gates the cleanup: only `performance` and `capacity` findings — continuous current-state metric thresholds — may be auto-resolved from absence. The other four categories (reliability, backup, security, general) stay active until explicitly resolved. Updates the ai-runtime subsystem contract Current State section with the whitelist and the adjacent lifecycle dedup rules already landed. Adds TestReconcileStaleFindings_SkipsNonCurrentStateCategories with table-driven subtests for all four event/persistent categories, and TestCategorySupportsStaleAutoResolve to lock in the whitelist. |
||
|
|
a1a9003dfd |
Drop duplicate loop_state lifecycle events for every transition
syncLoopStateLocked was emitting a generic "loop_state" lifecycle event on every successful transition, duplicating the semantic event the caller had just emitted. A finding that auto-resolved showed two adjacent rows in the Lifecycle drawer: Auto-resolved (detected -> resolved) Loop state changed (detected -> resolved) Same from/to, same timestamp, no extra information. Every transition was paired with a duplicate. Removed the generic loop_state emission. Every caller of syncLoopStateLocked already emits the semantic event for the transition it caused (auto_resolved, regressed, dismissed, acknowledged, snoozed, suppression_lifted, reminded, etc.). The loop_transition_violation branch stays — that's the only signal that an invalid transition was rejected, not a duplicate. Adds TestFindingsStore_TransitionDoesNotAlsoEmitGenericLoopStateEvent to lock in the behavior. |
||
|
|
43760fb0d0 |
Patrol runs are now stateless — drop prior session history
Patrol's "patrol-main" session was reused across every scheduled run, so ExecutePatrolStream loaded the full session history into the agentic loop's input. When any prior run ended after the model emitted tool_calls but before all tool results landed (provider error, timeout, context cancellation), the orphan tool_calls were persisted and every subsequent run inherited them. The provider then rejected the conversation with: An assistant message with 'tool_calls' must be followed by tool messages responding to each 'tool_call_id'. (insufficient tool messages following tool_calls message) Patrol failed 271 consecutive runs with this error before it was diagnosed. Each new run added another user prompt on top of the broken structure, so the message slice grew to 33+ messages with one assistant turn at position 23 holding 4 orphan tool_call_ids and 9 user prompts stacked after it. The "patrol runs need a clean slate" comment at line 1898 documents that the knowledge accumulator is freshened per run; the conversation history was the matching gap. ExecutePatrolStream now passes only this run's user prompt to the agentic loop. The session is still written to for audit/forensics, just no longer fed back into the model. Live verified: Patrol now completes successfully (18 tool calls, 3m23s) on a session that previously failed every run with the malformed-history error. The runtime-failure finding auto-resolved on this same successful run. Adds a classifier bucket for "insufficient tool messages" errors so any future regression in this area surfaces with a meaningful diagnostic instead of the generic "Provider analysis error" fallback. New PatrolFailureCauseMalformedToolHistory cause; predicate patrolMalformedToolHistory matches DeepSeek's exact phrasing and OpenAI's similar variants. ai-runtime contract updated. |
||
|
|
e087eff00e |
Stop polluting Patrol finding lifecycle with no-op heartbeats
Every Patrol scan that re-detected an already-active finding was
appending a "detected (same_state -> same_state)" lifecycle event
with the message "Detected by Pulse Patrol". A finding active for
6 scans rendered as 6 stacked rows reading
"Detected Detected by Pulse Patrol (detected -> detected)".
Backend: drop the unconditional re-detection lifecycle append in
findings.go. The lifecycle should record state transitions, not
heartbeats — TimesRaised and LastSeenAt already track recurrence,
and the genuine transition events ("regressed", "reminded",
"suppression_lifted") are emitted from their own branches upstream.
Frontend: defensively hide (from -> to) spans where from === to
and strip the type-label prefix from the message so already-
persisted polluted lifecycle entries also render cleanly until
they age out of the per-finding event cap.
Adds a test that locks in the new backend behavior.
|
||
|
|
297556fb65 |
Surface cached preflight in Patrol tools readiness check
Previously the "Patrol tools" readiness check was static
model-name pattern matching: it told the operator whether
the selected model is on Pulse's tool-capable allowlist,
not whether tools have actually been verified to work.
After today's preflight cache, that's strictly less
information than what we already know.
resolvePatrolToolsCheck now consults
aiService.CachedPatrolPreflight() and grounds the check
in real evidence when a result for the configured
provider+model exists:
- cached green (success + tool_call_observed) →
"Tool calling verified <age> against <model>." (ready)
- cached failure → classified summary + "(last preflight
<age>)." (not_ready)
- cached soft warning (model_tool_support_unverified) →
same with warning status
- no cache or model mismatch → static fallback
formatPatrolPreflightAge produces stable English ("just
now", "5m ago", "2h ago", "3d ago") with full unit
coverage (11 cases).
HandleGetAISettings now also includes patrol_readiness in
its response — previously only the PUT response carried it,
so the Patrol page only got augmented readiness after a
save. The frontend already had patrol_readiness typed and
read it from useAISettingsState.
ai-runtime, api-contracts, agent-lifecycle (dep), and
storage-recovery (dep) contracts updated.
|
||
|
|
f74add8271 |
Auto-seed Patrol preflight cache on Pulse startup
Closes the cold-start gap in the preflight observability layer: every Pulse restart blanked the cached "last verified" indicator until the next save or manual click, which meant operators saw "never verified" on every upgrade or process restart even when the configured Patrol model was working fine. NewAISettingsHandler now reuses the existing aiSettingsUpdateRequiresPatrolPreflight predicate with a nil "prior config" — semantically "no in-memory cache yet, just booted." When the loaded config has assistant enabled and a Patrol model, the handler dispatches the same async TriggerPatrolPreflightAsync the save path uses. Routine boots where assistant is disabled (or no Patrol model is selected) skip the dispatch so we never write a misleading "Pulse Assistant is not enabled" entry into the cache. Live verified: after Pulse restart, /api/settings/ai surfaces a fresh patrol_preflight with success=true within ~6s of boot, no operator action required. Predicate test extended with two named cases that document the dual-purpose use (startup seed + skip-when-disabled). ai-runtime, api-contracts, agent-lifecycle (dep), and storage-recovery (dep) contracts updated. |
||
|
|
33b6bf97d1 |
Auto-trigger Patrol preflight when settings save moves Patrol transport
Closes the resilience gap: today an operator picks a Patrol
model, saves, and Patrol silently fails on its first
scheduled run because nothing verified the model actually
calls tools. Now the save handler dispatches a one-shot
preflight in the background whenever the change actually
moved Patrol's transport, and the cached result surfaces on
the next /api/settings/ai poll via patrol_preflight.
Trigger conditions (aiSettingsUpdateRequiresPatrolPreflight):
- new config has assistant enabled AND a Patrol model
- AND any of:
* no prior config (first save)
* assistant was disabled, now enabled
* Patrol model changed
* API key for the new patrol model's provider changed
Routine saves that don't touch Patrol transport (theme,
control level, discovery toggles, unrelated provider keys)
skip preflight entirely so they don't burn provider tokens
or add 5-10s latency to every save.
Service-level TriggerPatrolPreflightAsync runs the call in a
goroutine with a 30s timeout. Detection helper has full
unit coverage including the negative paths.
|
||
|
|
728c42e47b |
Bring action endpoints onto the agent surface with the agent-stable envelope
Closes the last known gap in the agent substrate. The three
action endpoints (POST /api/actions/plan, /api/actions/{id}/decision,
/api/actions/{id}/execute) previously emitted the platform-wide
APIError shape (stable code under "code", human under "error").
The agent surface uses the inverted shape (stable code under
"error", human under "message"), so adding action capabilities
to the manifest as-is would have forced agents to remember which
envelope each capability uses.
The slice refactors actions.go to emit the agent-stable envelope
across all 42 writeErrorResponse call sites. writeJSONError gains
a writeJSONErrorWithDetails sibling so the 13 calls that pass
field-level reasons (validation failures) preserve that
information under a new optional `details` field. The action
endpoints' JSON shape becomes:
{"error": "<stable_code>", "message": "<human>",
"details"?: {"<field>": "<reason>"}}
Frontend impact: zero. Verified that no frontend code consumes
the three action endpoints; the refactor is API-only.
Three new manifest entries (plan_action, decide_action,
execute_action) under a new "action" category, with their
declared error codes pinned per capability. Internal-failure 5xx
codes (audit-store outages, encode failures) are not declared
per capability; agents branch on 5xx generically.
TestContract_AgentSurfaceErrorCodesMatchManifestDeclarations now
audits actions.go alongside the existing two handler files, with
a documented internal-only allowlist for the 5xx codes.
The TestAgentSubstrate_ActionEndpointsEmitAgentStableEnvelope e2e
test exercises one error path through each endpoint via the actual
HTTP boundary, asserting the agent-stable envelope reaches the
wire and the legacy APIError fields (code, status_code, timestamp)
do NOT — drift back would mean the refactor regressed.
The TestContract_ActionDryRunOnlyExecutionErrorJSONSnapshot pin
is updated to match the new envelope shape; the manifest's
category allowlist gains "action".
api-contracts.md documents the new envelope (with details map),
the action governance loop's place in the substrate, the
ai:execute scope distinction from monitoring:write, and the
"manifest projection has a footnote" trade-off: bringing an
existing endpoint into the agent surface may require migrating
its error envelope, but the substrate keeps a single envelope
contract rather than carrying a translation wrapper layer.
agent-lifecycle.md and storage-recovery.md document the action
surface joining the agent paradigm and its zero-new-persistence
posture respectively. AGENT_SUBSTRATE.md's "what it does not do
yet" no longer lists the action surface; it now reflects the
real outstanding items (consumer feedback, an in-Pulse agent
integrations panel, a distribution path for pulse-mcp).
|
||
|
|
add1096ec8 |
Cache Patrol preflight outcome and hydrate UI on settings load
The Verify Patrol button reset its result to empty on every page load — the operator had to re-click to see the verified state, even though nothing had changed. This commit adds the observability layer of the auto-preflight plan: every RunPatrolToolPreflight result is now cached on the AI Service and surfaced through /api/settings/ai as patrol_preflight, so the inline result panel rehydrates on page load with the most-recent outcome and a "last verified Xs ago" indicator. Backend: patrolPreflightCache (mutex-guarded) on Service with defensive-copy CachedPatrolPreflight() accessor; every RunPatrolToolPreflight branch (success, soft warning, classified failure, validation early-return) records into the cache. PatrolPreflightSnapshot projects the cached result onto the AI settings response. Tests cover both success-then-failure supersession and the defensive-copy invariant. Frontend: PatrolPreflightSnapshot type mirrors the wire shape; hydratePatrolPreflightFromSettings(data) projects the snapshot into the same response shape the manual button writes; loadSettings and updateSettings flows call it. The result panel renders a "last verified Xs ago" line under the provider/model row when recorded_at_unix is present. End-to-end smoke verified against deepseek-v4-flash: panel rehydrates as green "Tool calling verified · last verified just now" after page reload. Auto-preflight on save (the trigger half of the resilience plan) follows in the next commit. Contracts: ai-runtime, api-contracts, agent-lifecycle (dep), storage-recovery (dep), frontend-primitives all updated to reflect the new patrol_preflight surface and hydration contract. Verification artifacts: settingsArchitecture + patrolPreflight client tests. |
||
|
|
404f87854e |
Pin cross-org and cross-resource isolation on the bundle's pending approvals
The AgentApprovalsProvider closure in router.go applied the
BelongsToOrg and CanonicalResourceID filters inline, which made
the substrate's tenant-isolation property impossible to test
without booting the full router. Drift in the closure (e.g.
swapping BelongsToOrg for a hardcoded "default" or dropping the
resource-id check) would let an agent with one org's token see
approvals targeting another org's infrastructure, but no test
sat right next to that logic to catch it.
Extracts the body into a named function in agent_resource_context.go
(pendingApprovalsForResourceFromStore) behind a minimal
approvalsPendingProvider interface. The closure in router.go
now delegates to it. Four unit tests pin the substrate's
isolation property:
- FiltersByOrg: same resource id, two orgs, each query returns
only its own org's approval.
- FiltersByResource: same org, two resource ids, each query
returns only its own resource's approval.
- LegacyEmptyOrgIsDefaultOnly: approvals without OrgID are
treated as default-org per BelongsToOrg's documented
semantics; legacy approvals do not leak into a non-default
org's bundle.
- EmptyInputsReturnNil: defensive shape on nil store, empty
resource id, and empty store.
The existing TestContract_AgentResourceContextWiresApprovalsProvider
pin is updated to follow the extraction. Both halves of the
wire-up are now pinned: router.go installs the closure with the
correct delegation, and agent_resource_context.go owns the
filter logic with both safety checks present.
This is the test the substrate was missing: nothing else proved
that an agent with one org's token cannot see another org's
pending approvals at the bundle layer.
Contract-neutral commit: no wire shape, manifest entry, or error
code changed. The refactor preserves identical behaviour;
PULSE_ALLOW_CONTRACT_NEUTRAL_COMMIT is set with a documented
reason since three of the four contract docs the canonical-shape
guard would normally demand are actively mid-edit by another
agent on patrol-preflight work, and trampling them would create
a collision the protocol explicitly forbids.
|
||
|
|
e26a57a157 |
Add POST /api/ai/patrol/preflight tool-call verification
The existing per-provider /api/ai/test endpoints only call ListModels — they pass for every provider that returns a catalog, even when Patrol fails 100% of runs because tools aren't actually wired up. That gap is what let the DeepSeek tool_choice rejection silently fail Patrol for 33 days before the recent fix landed. POST /api/ai/patrol/preflight runs a one-shot tool-call round-trip with the configured (or overridden) Patrol provider+model and a minimal verify_pulse_patrol tool. Failures route through ClassifyPatrolRuntimeFailure so the new tool_choice_rejected and no_tool_capable_endpoint causes surface here too. A successful provider call where the model returned plain text (no tool call) is reported as a soft warning (model_tool_support_unverified): Patrol may still work but the operator should run a real pass to confirm. The endpoint bypasses the chat service so cost recording isn't charged for verification, and uses ScopeSettingsWrite to align with the existing /api/ai/test gating. Backend + typed frontend client (runPatrolPreflight); UI button on Assistant & Patrol settings follows. Contracts updated: - ai-runtime: completion obligation extended to cover the new verification surface - api-contracts: payload shape (tool_call_observed, duration_ms) noted in obligations - agent-lifecycle, storage-recovery: dependent-extension acknowledgment that ai-runtime owns the new route despite it living under internal/api/ |
||
|
|
f2d9d2aba8 |
Split overgreedy "tools not supported" classifier into three causes
The Patrol runtime classifier collapsed three distinct upstream
conditions into one misleading "Selected model does not support
Patrol tools" message:
1. Provider rejected the *value* Pulse sent for tool selection
(e.g. DeepSeek's "deepseek-reasoner does not support this
tool_choice" — the model accepts tools, just not the forced
coercion). The DeepSeek fix in
|
||
|
|
46145df925 |
Coerce DeepSeek tool_choice to "auto" so Patrol stops failing
DeepSeek's API server-side aliases deepseek-v4-flash and
deepseek-v4-pro to deepseek-reasoner, which rejects forced
tool_choice with HTTP 400 ("deepseek-reasoner does not support
this tool_choice"). Pulse's classifier then surfaced this as
"Selected model does not support Patrol tools," misdirecting
diagnosis to the model rather than the request shape.
supportsForcedToolChoice now returns false for any DeepSeek
client, so every DeepSeek model falls back to tool_choice
"auto" regardless of how DeepSeek routes the requested ID.
The ai-runtime contract is updated to match: the
provider-transport boundary now coerces forced tool_choice for
every direct DeepSeek model ID, not only unknown ones.
Patrol verified end-to-end: 20 tool calls, 9 findings, prior
runtime failure auto-resolved.
|
||
|
|
2f0468a87b |
Verify SSHSIG on in-app update artifacts
The unattended timer (scripts/pulse-auto-update.sh) and the public bootstrap (scripts/install.sh, /install.sh) all verify the .sshsig sidecar against the pinned pulse-installer ed25519 key before trusting a release artifact. The in-app updater verified SHA256 only — same artifact, same root execution context, lower trust bar. Closing the asymmetry: the in-app tarball download in ApplyUpdate, adapter_installsh.go's install.sh download (piped into bash as root), and the rollback binary download now fetch and verify the .sshsig sidecar against the same pinned key, fail-closed. The signing infrastructure (release_asset_common.sh, validate-release.sh, backfill-release-assets.sh) already produces and validates these signatures for every release; this teaches the Go updater to honor what the shell paths have always required. ssh-keygen is shelled out to so the in-app updater shares the exact trust path used by the unattended path, with a package-level function variable for test injection so unit tests don't require ssh-keygen on the build host. Extends the deployment-installability contract's release-trust-fail-closed invariant to cover the in-app updater paths. |
||
|
|
eeb2975d22 |
Stability sweep on the agent-substrate arc
Three things landed:
1. /api/agent/capabilities was missing from publicPathsAllowlist
in router_public_paths_inventory_test.go. Slice 47 added the
path to publicPaths in router.go and to publicRouteAllowlist
in route_inventory_test.go but missed this second mirror,
which scans publicPaths via go/ast. The test was failing on
origin; this commit closes the gap.
2. The error-envelope paragraph in api-contracts.md now
distinguishes capability-specific stable codes (the closed
set declared per capability in the manifest) from
cross-cutting codes the multi-tenant / auth middleware
emits universally (invalid_org, org_suspended, access_denied).
The previous wording implied all stable codes lived in
per-capability errorCodes lists, which would have forced
duplication on every capability or misled agents about which
codes to expect.
3. New contract pin TestContract_AgentSurfaceErrorCodesMatch-
ManifestDeclarations enforces the symmetry both directions:
every code emitted by an agent-surface handler must be either
declared in the matching capability or be one of the three
cross-cutting codes; every manifest-declared code must have a
matching emission. Drift either way is a contract regression.
Pin verified clean against the current handler set.
Stale forward-reference fixed: the capabilities paragraph no
longer says "future MCP-server slices read the manifest" — slice
51 already shipped that adapter.
Sweep also surfaced two failures in internal/mock/ from
unrelated platform-support drift (unraid token set added in
|
||
|
|
8aa22d0605 |
Surface action verification on the action.completed SSE payload
Closes the certainty loop for agents watching the substrate's push
channel. The action audit's read-after-write probe outcome was
already persisted on the audit record, but agents watching
action.completed only learned "the action ran" — they had to fetch
/api/actions/{id} to know whether the read-back probe confirmed
the intended state. That defeated the substrate's
push-notification guarantee for dispatch certainty.
The new agent-stable AgentResourceActionVerification projection
(ran, success, command, note, ranAt — output stays in the audit
record, deliberately omitted from events to keep payloads small)
is now carried on both:
- the action.completed SSE payload, projected from
record.Result.Verification by the router-side bridge in
wireAIChatDependenciesForService, and
- the resource-context bundle's recentActions surface, via the
same shared projectAgentResourceVerification helper
so the bundle (depth) and the doorbell (push) speak the same
vocabulary. Refused-before-dispatch failures omit verification
(the probe never runs) so agents branch on field presence to
distinguish "no probe attempted" from "probe ran with empty
result". Three contract pins lock the symmetry: payload field
present, router bridge populates it, bundle parallels.
The capabilities manifest's subscribe_events description now
mentions the verification block so external agents discover the
field through the same path they already use to learn the rest
of the agent surface.
|
||
|
|
5156c03eed |
End-to-end test the operator-state write loop through HTTP
Closes the e2e contract proof on the write side. The only write capability the manifest declares is the operator-state intent loop (set / get / clear), and this test boots the full router stack to walk every state of that loop through the actual HTTP boundary — proving the manifest's declared error codes for set_operator_state and get_operator_state reach the wire from the handlers, the URL canonical id authoritatively wins over body-supplied ids (no scope-confusion writes), and SetAt/SetBy are server-populated so attribution cannot be spoofed. The flow exercised: GET unset → 404 operator_state_not_set PUT valid → 200 with persisted state + server SetAt GET → round-trips PUT invalid criticality → 400 operator_state_invalid DELETE → 204 GET → 404 operator_state_not_set (loop closed) DELETE again → 204 (idempotent) Two contract pins lock the audit-honesty and error-token contracts so a future refactor of the handler can't silently regress either: SetAt/SetBy populated server-side, URL-id wins over body-id, and the validator's domain error maps to the stable wire token via errors.Is rather than message-matching. Together with the read-side e2e (slice 47), the agent surface — read, write, push — has now been exercised end-to-end as one substrate. |
||
|
|
8cf15fe639 |
End-to-end test the agent substrate's discovery → triage → depth flow
The unit tests cover each piece in isolation; this test boots the
full router stack and proves the discovery → triage → depth chain
works as one substrate through the actual HTTP boundary an
external agent would hit. It found two real bugs slice 40
introduced and slice 45/46 didn't surface:
- /api/agent/capabilities was documented as unauthenticated but
was missing from the router's publicPaths list, so the global
auth middleware was 401'ing the discovery manifest. Fixed by
adding the path to publicPaths and pinning the contract so it
cannot regress.
- The error-envelope shape across the agent surface is
{"error": "<stable_code>", "message": "<human>"}, written via
writeJSONError — not the {"code": ...} shape I had assumed in
the docs. Pinned the wire shape on api-contracts.md so the
documented error contract matches what writeJSONError actually
writes.
The e2e test exercises capabilities discovery, triage via
fleet-context, and depth via resource-context with an unknown id
to confirm the resource_not_found stable error code reaches the
wire under the canonical "error" key. The subscribe_events SSE
path is probed unauthenticated to confirm it's gated (401) rather
than 404 — discovery's claim is honest.
|
||
|
|
a168215f6a |
Add /api/agent/fleet-context for org-wide triage in one read
The substrate had a per-resource bundle but no fleet view, so
"where do I focus?" forced agents to walk every resource id and
bundle each — O(N) round trips that scale with fleet size. The
fleet endpoint returns a thin per-resource rollup in a single
read: identity, operator-intent flags (intentionallyOffline,
neverAutoRemediate, maintenanceWindowActive), per-severity
finding counts, and pending-approval count.
Same auth scope and same provider wiring as the per-resource
bundle — operator-state via the canonical unified store, findings
via AgentFindingsProvider, approvals via AgentApprovalsProvider —
so the fleet sweep is the per-resource bundle's wiring multiplied
by N with no new dependencies. Audit reads are deliberately
omitted from the rollup; agents that want depth on a flagged
resource follow up via /api/agent/resource-context/{id}.
The capabilities manifest declares get_fleet_context with
AgentFleetContext as the response shape so external agents
discover the triage entry point through the same path they
already use to learn the rest of the agent surface.
|
||
|
|
d8f6b1e508 |
Bundle pending approvals into the agent resource-context endpoint
The substrate's "everything an agent needs in one read" guarantee covered identity, operator state, findings, and recent actions but forced a separate /api/approvals call for pending governance state. AgentResourceContext now carries pendingApprovals as a lightweight AgentResourceApprovalSummary projection — same vocabulary as approval.pending SSE events, so the doorbell and the bundle agree on shape. AgentApprovalsProvider is the parallel seam to AgentFindingsProvider; the router wires a closure that resolves approval.GetStore() at request time, scopes via BelongsToOrg, and filters by CanonicalResourceID so cross-tenant or cross-resource pending requests don't leak. Empty arrays preserve the iteration-safe contract the existing sections already follow. |
||
|
|
7fe9b1c492 |
Use cursor-help on TagBadges hover-only +N indicator
The "+N" overflow indicator on TagBadges was styled with `cursor-pointer`, which signals a clickable affordance — but the element only listens for mouseenter/mouseleave to show a tooltip and has no click handler. Switch to `cursor-help` so the cursor matches the actual interaction (hover for more info), avoiding a phantom click expectation. |
||
|
|
52669128e6 |
Drop redundant policy gates in resource-link routing
Tail of the operator-local-UI redaction sweep ( |
||
|
|
51c5d344ce |
Plumb operator-state and operational memory into investigation findings
Closes the "has context vs uses context" gap that defines Pulse's agent-paradigm differentiation. The orchestrator (in pulse-pro) used to receive a Finding with no awareness of the operator's commitments — Patrol could investigate a resource the operator had marked never-auto-remediate and propose a restart fix that the action broker would refuse downstream. The proposal shouldn't have happened in the first place. Adds two optional fields to aicontracts.Finding: - OperatorContext: intentionally offline, never auto-remediate, maintenance window with computed active flag, criticality, note. Populated in MaybeInvestigateFinding from the same operator-state projection the suppression hot path consumes, so investigation reasoning and suppression behavior cannot drift apart. - OperationalMemory: regression count, previous resolved fix summary, last regression timestamp, times raised. Populated in ToCoreFinding from fields the internal Finding already carries. ResourceOperatorStateProjection grew a NeverAutoRemediate field — the investigation read path needs it (so the orchestrator can avoid proposing fixes the broker would refuse) even though the suppression hot path doesn't. Same projection serves both reads. Both fields are nil when there's no signal (fresh finding, no operator state) so the orchestrator branches on absence rather than parsing zero-valued structs. The pulse-pro orchestrator consumes the fields in a separate slice; this slice ships the in-repo half of the data path. |
||
|
|
94bfd48a9d |
Add /api/agent/events SSE stream for real-time agent notifications
Third slice on the agent-paradigm pivot, closing the substrate triangle (discovery + bundled reads + push). Agents subscribe once to a long-lived SSE connection and receive real-time events instead of polling: finding.created when a new finding is raised, heartbeat every 15 seconds for keepalive. Each event carries a monotonic ID so agents can dedupe and reason about ordering across reconnects. The broadcaster fan-outs to multiple subscribers and drops events for slow consumers rather than blocking the publish path — publishers cannot stall on consumer slowness. The findings-runtime hook in router.go publishes finding.created when the finding is new AND not auto-dismissed by operator-state suppression (operator already said to stay quiet about that resource); patrol-cycle re-detection of existing findings doesn't fire the event. Capabilities manifest declares the stream under subscribe_events so external agents discover it through the same channel as the REST surface. SSE chosen over WebSocket because it's simpler, works through every HTTP proxy without special-casing, and matches the existing deploy_handlers pattern; agents that need bidirectional comms call REST endpoints in parallel. Tests pin the broadcaster's pub/sub semantics (fan-out, unsubscribe, slow-consumer drop, monotonic IDs), the SSE handler's stream contract (text/event-stream, no-cache, X-Accel-Buffering=no), and the connected/published-event delivery via httptest.NewServer. A contract test pins the publish-gate semantics so operator-state suppression and stream notifications stay aligned. |
||
|
|
71797f9b21 |
Add /api/agent/capabilities discovery manifest for agent integrations
Second slice on the agent-paradigm pivot: the discovery document any external agent (Claude Code, custom integrations, future MCP servers) needs to learn what Pulse exposes. Each capability declares its agent-stable name (snake_case), description, category, REST surface, required scope, response shape, and the closed set of stable error codes the response may carry. Agents branch on the codes (operator_state_invalid, resource_not_found, etc.) rather than parsing human messages. The manifest is hand-authored, not auto-generated, because the contract decisions (what's agent-stable, which categories, which error codes) are product-shaping and must not drift behind code changes. Adding a capability is a deliberate "this is part of the agent surface" commitment. v1 surface includes: get_resource_context (substrate from slice 39), get/set/clear_operator_state (slice 30), and the finding-lifecycle actions (acknowledge, snooze, dismiss, resolve). Action-broker capabilities are not in v1 because they go through approval flow, not direct dispatch — those need their own contract design. Tests pin: stable shape, version contract, unique-and-snake_case names, every capability has method/path/scope, closed category set, required error codes for the most consequential capabilities. The manifest is unauthenticated and cacheable (5min); the underlying capabilities keep their own auth scopes. |
||
|
|
14f9270a5e |
Add /api/agent/resource-context/{id} substrate endpoint for agents
First slice on the agent-paradigm pivot: instead of building more human-glance UI, expose substrate that any agent (in-process Patrol, external Claude Code, future MCP-driven setups) can consume in one read. The endpoint returns the full situated picture of a resource — identity, operator-set state with server-computed maintenanceWindowActive flag, active findings as a lightweight seven-question-schema projection, and recent action audits with refusal tokens (resource_remediation_locked:, plan_drift:) preserved verbatim for agent branching. Substrate is the right shape here: an agent reasoning about a resource gets everything it needs without chaining four or five calls, and the projection types decouple agent-stable wire shape from internal type evolution. Active findings flow through an AgentFindingsProvider adapter wired in router.go from the patrol service, keeping the api package free of an internal/ai import. Always-array fields (activeFindings, recentActions) and omitempty-on-absent (operatorState) give agents stable iteration and clean field-presence branching. AgentContextHandler owns the agent surface as its own type so it evolves independently of resource CRUD. Each test pins a specific contract: identity round-trip, operator-state projection with computed flag, empty-state shape, refusal-token preservation, 404 shape, method gating. |
||
|
|
eae8ca2a68 |
Wake operator-state-suppressed findings when the suppression lifts
Real product gap exposed by closing the operator-state feature: a finding auto-dismissed while a maintenance window covered `now` would stay dismissed forever even after the window ended. Same for findings auto-dismissed under IntentionallyOffline once the operator cleared the flag. The time-bounded suppression silently became permanent. Adds a third wake condition to the dismissed-branch in FindingsStore.Add: when a finding's most recent dismissed lifecycle event carries operator_state_cause metadata, AND the provider reports no current suppression for the resource, clear the dismissal and emit a suppression_lifted lifecycle event naming the previous cause. Manual operator dismissals (no operator_state_cause) are unaffected — the findOperatorStateDismissCause helper stops at the first dismissed event when scanning newest first, so a manual dismissal that supersedes an earlier auto-dismiss is not falsely re-awakened. Tests cover both signal types, the manual-dismissal isolation, and the helper's newest-first scan order. |
||
|
|
b822ef17e6 |
Pin: operator-state-suppressed findings skip autonomous investigation
Cross-slice contract worth making explicit: when slices 31/32 auto-dismiss a finding because the operator's per-resource state suppresses it, that finding must not also burn investigation budget. The existing chain already delivers this — findings.Add sets DismissedReason="expected_behavior", and ShouldInvestigate gates on DismissedReason != "" — but the relationship was implicit. Without a test, a future refactor of either branch could silently start investigating operator-suppressed findings again. Pins the contract with a table-driven test covering both signals (intentionally_offline and maintenance_window) at every autonomy level (approval/assisted/full), plus the lifecycle-cause metadata attribution. No runtime change — only the test and a contract paragraph naming the dependency. |