Commit Graph

2839 Commits

Author SHA1 Message Date
rcourtman 321f563a52 Surface blocked workload inventory sources
Show workload-capable source failures on Workloads and keep matching Proxmox host agents attached to their API source when inventory collection is blocked.
2026-05-14 00:22:18 +01:00
rcourtman 2be14562ee Preserve infrastructure continuity on first login
Ensure unified resource snapshots include recent standalone host-agent continuity so Infrastructure does not briefly undercount connected systems after login or restart.
2026-05-13 23:36:17 +01:00
rcourtman 823bd3fbb1 Fix fleet command policy convergence 2026-05-13 22:18:47 +01:00
rcourtman d3670d7882 Redact legacy action audit verification details 2026-05-13 21:50:56 +01:00
rcourtman 3a7efda5f2 Redact action verification audit details 2026-05-13 21:20:59 +01:00
rcourtman ceb9b87cfb Correct fleet config drift truth 2026-05-13 20:38:26 +01:00
rcourtman 17253d27fd Surface connection rollout posture 2026-05-13 20:16:35 +01:00
rcourtman 427b7ff5c6 fix: route patrol update safety through read state 2026-05-13 19:24:25 +01:00
rcourtman adc46e2d0a Harden action audit verification redaction 2026-05-13 19:10:17 +01:00
rcourtman e8b3c7fcf7 Restore remote config signature compatibility
Keep desired config fingerprints as response metadata derived from the signed command and settings payload.

Use merged agent profile settings when building remote config fingerprints.
2026-05-13 19:00:02 +01:00
rcourtman 554158c575 Add desired config fingerprint metadata 2026-05-13 18:51:24 +01:00
rcourtman dbe31bd8d6 Fix action verification read projections 2026-05-13 18:36:00 +01:00
rcourtman f024d3b560 Align action audit verification projection 2026-05-13 18:36:00 +01:00
rcourtman d55888fb7f Correct PBS job health evidence boundaries 2026-05-13 17:05:03 +01:00
rcourtman dd8c3a78ce Add PBS job health evidence ledger 2026-05-13 17:05:03 +01:00
rcourtman da2537e4ab Add self-hosted commercial continuity proof 2026-05-13 16:44:26 +01:00
rcourtman f5aa2b9cd1 Fix pre-commit sibling roots in worktrees 2026-05-13 15:56:19 +01:00
rcourtman fbbf1daf4c Harden action start refusal audit 2026-05-13 14:25:04 +01:00
rcourtman c1db220e3a Make maintenance evidence writes atomic 2026-05-13 14:19:59 +01:00
rcourtman 14d5cd4b3b Record maintenance evidence in resource timelines 2026-05-13 14:19:59 +01:00
rcourtman eb78bf9f37 Persist permanent action refusal outcomes 2026-05-13 14:04:14 +01:00
rcourtman 1a47c03b2b Fix auto-register refresh notifications 2026-05-13 13:59:11 +01:00
rcourtman 4c4d433666 Record VMware provider build failures 2026-05-13 13:27:25 +01:00
rcourtman 6e9c360301 Propagate PBS datastore health into monitoring 2026-05-13 13:15:10 +01:00
rcourtman 0fbcc90184 Harden external model prompt-secret sanitation 2026-05-13 13:03:40 +01:00
rcourtman 31331b5451 Fix grouped notification cancellation 2026-05-13 13:00:43 +01:00
rcourtman 1836e2a3cc Fix mixed quiet-hours notification replay queueing 2026-05-13 12:14:56 +01:00
rcourtman 5513781193 Ensure cloud emails use support reply-to 2026-05-13 12:10:02 +01:00
rcourtman 6112fcd5ae Replay quiet-hours alert notifications 2026-05-13 11:54:56 +01:00
rcourtman 7827f2b077 Add PDM HTTP alert source 2026-05-13 11:35:24 +01:00
rcourtman cfef1b67ae Fix action execution dry-run guard and verification bridge
Ensure dry-run-only actions fail with action_dry_run_only before executor availability checks, and bridge action completion verification projection through router payload mapping.
2026-05-13 11:16:28 +01:00
rcourtman 8197fe6b1e Use disk type thresholds for SMART temperatures 2026-05-13 10:44:11 +01:00
rcourtman 3cecf9576d Seed disk temperature type defaults 2026-05-13 10:44:11 +01:00
rcourtman becee760dd Add disk temperature by type normalization 2026-05-13 10:44:11 +01:00
rcourtman 950bd187a9 seed DiskFillByType defaults and normalize on load
Seeds AlertConfig.DiskFillByType in defaultAlertConfig() with the three
forward-compatible keys (nvme 92/87, sata 90/85, hdd 85/80). Adds
NormalizeDiskFillByType to lowercase keys on load, seed defaults when
the map is nil, and reset non-positive trigger/clear values to the
default for that key. NormalizeAgentDefaults invokes the new helper so
it runs through the existing UpdateConfig normalization chain. Round
trip tests cover nil seed, customized survival, negative reset, and
mixed case normalization, plus a defaultAlertConfig seed assertion.

Only the "nvme" key is consulted by the host evaluation today (the
inference helper only matches /dev/nvme*); "sata" and "hdd" are
forward-compatible placeholders for when the agent protocol carries
hardware type explicitly.
2026-05-13 07:16:07 +01:00
rcourtman 85fbbd4ef6 consult DiskFillByType in host disk fill alert evaluation
Inserts a per-type threshold lookup before the global AgentDefaults.Disk
fallback in CheckHost's disk evaluation. inferDiskHardwareType maps the
disk device path to "nvme" (only NVMe is inferable today); when the
DiskFillByType map carries a matching key, that hysteresis threshold is
used, otherwise evaluation falls through to the existing global default.
Storage-type unified_eval branch is unchanged.

Tests prove: (a) NVMe device at 91% does not alert when DiskFillByType
sets the nvme trigger to 92; (b) the same device at 93% alerts at the
nvme threshold; (c) /dev/sda1 falls back to the global threshold; (d)
the storage-type alert branch is not regressed.
2026-05-13 07:16:07 +01:00
rcourtman 005873bdd5 add disk-type-aware threshold field and device inference helper
Adds DiskFillByType map to AlertConfig (keys: nvme, sata, hdd ->
HysteresisThreshold). Adds inferDiskHardwareType helper that returns
"nvme" for /dev/nvme* device paths (case-insensitive) and "" for
anything else, since sata vs hdd cannot be reliably inferred from
device paths today. Unit tests cover the four input shapes (nvme,
sata, empty, mixed case).

Substrate only. No wire-in to host alert evaluation yet.
2026-05-13 07:13:46 +01:00
rcourtman 99d0edb04a wire PDM alert bridge into PatrolService and seed demo 2026-05-13 05:42:25 +01:00
rcourtman b9844f22c1 prove PDM alert bridge emits and resolves through FindingsStore.Add 2026-05-13 05:42:25 +01:00
rcourtman 8aad44174b add PDM alert bridge substrate in internal/ai 2026-05-13 05:42:25 +01:00
rcourtman 4097ddcb9c handle repeated update-safety digest changes 2026-05-13 04:18:27 +01:00
rcourtman 5821b6c810 wire update-safety watcher into PatrolService and seed demo 2026-05-13 04:16:35 +01:00
rcourtman 2e87e374eb prove update-safety watcher emits and resolves through FindingsStore.Add 2026-05-13 04:16:35 +01:00
rcourtman ab71410ccb add update-safety watcher substrate in internal/ai 2026-05-13 04:16:35 +01:00
rcourtman 178c961723 wire storm throttler into PatrolService from the constructor
The substrate has been in place for two commits but no live code path
constructs a throttler, so production behaviour is unchanged. This
commit closes the loop: `NewPatrolService` instantiates a
`findingStormThrottler` and calls
`p.findings.SetStormThrottler(p.stormThrottler)` immediately after
the `FindingsStore` is constructed, so every patrol emission flows
through the observer from the first run onward.

Constructor-internal wire-up keeps the throttler intrinsic to the
service rather than a deployment-time concern. No `router.go` seam,
no contract test pin, no goroutine to stop; the throttler holds
in-memory state and is garbage-collected when the PatrolService is
replaced.

The field sits next to `unifiedResourceProvider` to match the brief's
location and keep adjacent observer/provider seams visually grouped.
2026-05-13 02:41:45 +01:00
rcourtman 38a471120d prove storm throttler emits and resolves through FindingsStore.Add
Commit 1 landed the substrate but only exercised it via direct calls
on the throttler. This commit drives a real `FindingsStore.Add` end
to end with the throttler wired via `SetStormThrottler`, so the hook
inside `Add`'s new-finding branch is on the test path.

Covers the four behaviours the brief calls out:

* Two distinct findings on the same resource within the window keep
  the storm finding absent.
* The third distinct finding within the window emits a single storm
  finding (severity warning, category reliability,
  Source="finding-storm", IsActive).
* A fourth distinct finding within the same window does NOT spawn a
  duplicate — the stable ID lands the re-entry on the existing-
  finding branch and bumps TimesRaised.
* Cycle guard: emissions whose own Source matches the storm-finding
  source flow through Add without tripping a storm.
* Filter: emissions against the synthetic patrol-runtime resource
  ("ai-service") flow through Add without tripping a storm.
* Resolve: after a synthetic quiet window the throttler returns a
  sentinel that `ResolveWithReason` converts into an auto-resolved
  storm finding with the expected reason string.

The resolve case is exercised by calling `observeLocked` directly
with a synthetic future time (then dispatching the sentinel through
`ResolveWithReason`) to avoid a 120s wall-clock wait in the test
binary; the production hook in `Add` follows the same path on the
next real emission against the cluster.

No production-code changes in this commit.
2026-05-13 02:41:45 +01:00
rcourtman 8f9ff92701 add storm throttler substrate in FindingsStore.Add
A single noisy resource that trips several patrol detectors in a tight
window currently emits one finding per symptom. The findings panel
fans those out, which is the right surface for the detectors but the
wrong surface for the operator: a flapping VM with CPU, swap, and
disk pressure should read as one "this resource is in a bad state"
event, not three concurrent rows competing for attention.

This commit lands the throttler substrate. The new
`findingStormThrottler` is invoked from `(*FindingsStore).Add`'s
new-finding branch (while `s.mu` is held), clusters emissions by
`Finding.ResourceID`, and when a cluster crosses
`stormThreshold` (3) within `stormWindow` (60s) returns a storm
finding for the caller to re-enter through `s.Add`. The stable ID
`finding-storm:<resourceID>` dedups subsequent emissions through the
existing-finding branch, and the cycle guard on `Source ==
"finding-storm"` keeps the storm finding's own emission from
recursing back into the observer.

Auto-resolve is lazy by design: on the next observe trip where a
cluster has fallen below threshold AND the storm finding has been
quiet for at least `2 * stormWindow`, the observer returns a
sentinel that the caller routes through `ResolveWithReason`. No
sweep goroutine, no extra wiring surface — the substrate is
contained in `internal/ai/` and exposes a single nil-safe setter
(`SetStormThrottler`).

This commit is observer-only: nothing constructs a throttler yet, so
in-tree behavior is unchanged. Wire-up lands in a follow-up.
2026-05-13 02:41:45 +01:00
rcourtman d4f84e69e0 add hourly sweep for overdue will_fix_later commitments 2026-05-13 01:13:38 +01:00
rcourtman 183bf5b617 persist remind_at and remind_count on will_fix_later findings 2026-05-13 01:13:38 +01:00
rcourtman 0e0c90da53 emit reliability finding when an alert starts flapping
Wire the alerts manager's new flapping-detected callback in the AI
intelligence initialization path. Two things happen on each first
transition into the flapping cooldown window for a tracking key:

1. A reliability-category finding is written directly to the findings
   store via emitFlappingPostmortemFinding. Path B from the lane brief:
   the finding is durable without depending on patrol synthesis, so
   the operator sees the diagnosis the moment Pulse decides to
   suppress. The finding ID is derived from the canonical tracking
   key ("alert-flapping:<trackingKey>") so re-detection inside the
   cooldown window folds into the existing record via the same-ID
   branch of FindingsStore.Add -- one finding per flapping condition,
   not one per dispatch.

2. A scoped FlappingPostmortemPatrolScope is enqueued on the trigger
   manager so an actual patrol run can enrich the finding with deeper
   context once it lands.

The finding body names the flapping threshold, window, and cooldown
the manager is currently configured with, plus an action hint
(widen threshold, raise cooldown, or stabilise the resource). That
turns the suppressed alert from silence into a closable item on the
FindingsPanel.

FindingCategoryReliability is reused; no new category, no parent/
child finding structure -- those are deferred per the lane brief.
2026-05-13 00:28:24 +01:00