Show workload-capable source failures on Workloads and keep matching Proxmox host agents attached to their API source when inventory collection is blocked.
Ensure unified resource snapshots include recent standalone host-agent continuity so Infrastructure does not briefly undercount connected systems after login or restart.
Keep desired config fingerprints as response metadata derived from the signed command and settings payload.
Use merged agent profile settings when building remote config fingerprints.
Ensure dry-run-only actions fail with action_dry_run_only before executor availability checks, and bridge action completion verification projection through router payload mapping.
Seeds AlertConfig.DiskFillByType in defaultAlertConfig() with the three
forward-compatible keys (nvme 92/87, sata 90/85, hdd 85/80). Adds
NormalizeDiskFillByType to lowercase keys on load, seed defaults when
the map is nil, and reset non-positive trigger/clear values to the
default for that key. NormalizeAgentDefaults invokes the new helper so
it runs through the existing UpdateConfig normalization chain. Round
trip tests cover nil seed, customized survival, negative reset, and
mixed case normalization, plus a defaultAlertConfig seed assertion.
Only the "nvme" key is consulted by the host evaluation today (the
inference helper only matches /dev/nvme*); "sata" and "hdd" are
forward-compatible placeholders for when the agent protocol carries
hardware type explicitly.
Inserts a per-type threshold lookup before the global AgentDefaults.Disk
fallback in CheckHost's disk evaluation. inferDiskHardwareType maps the
disk device path to "nvme" (only NVMe is inferable today); when the
DiskFillByType map carries a matching key, that hysteresis threshold is
used, otherwise evaluation falls through to the existing global default.
Storage-type unified_eval branch is unchanged.
Tests prove: (a) NVMe device at 91% does not alert when DiskFillByType
sets the nvme trigger to 92; (b) the same device at 93% alerts at the
nvme threshold; (c) /dev/sda1 falls back to the global threshold; (d)
the storage-type alert branch is not regressed.
Adds DiskFillByType map to AlertConfig (keys: nvme, sata, hdd ->
HysteresisThreshold). Adds inferDiskHardwareType helper that returns
"nvme" for /dev/nvme* device paths (case-insensitive) and "" for
anything else, since sata vs hdd cannot be reliably inferred from
device paths today. Unit tests cover the four input shapes (nvme,
sata, empty, mixed case).
Substrate only. No wire-in to host alert evaluation yet.
The substrate has been in place for two commits but no live code path
constructs a throttler, so production behaviour is unchanged. This
commit closes the loop: `NewPatrolService` instantiates a
`findingStormThrottler` and calls
`p.findings.SetStormThrottler(p.stormThrottler)` immediately after
the `FindingsStore` is constructed, so every patrol emission flows
through the observer from the first run onward.
Constructor-internal wire-up keeps the throttler intrinsic to the
service rather than a deployment-time concern. No `router.go` seam,
no contract test pin, no goroutine to stop; the throttler holds
in-memory state and is garbage-collected when the PatrolService is
replaced.
The field sits next to `unifiedResourceProvider` to match the brief's
location and keep adjacent observer/provider seams visually grouped.
Commit 1 landed the substrate but only exercised it via direct calls
on the throttler. This commit drives a real `FindingsStore.Add` end
to end with the throttler wired via `SetStormThrottler`, so the hook
inside `Add`'s new-finding branch is on the test path.
Covers the four behaviours the brief calls out:
* Two distinct findings on the same resource within the window keep
the storm finding absent.
* The third distinct finding within the window emits a single storm
finding (severity warning, category reliability,
Source="finding-storm", IsActive).
* A fourth distinct finding within the same window does NOT spawn a
duplicate — the stable ID lands the re-entry on the existing-
finding branch and bumps TimesRaised.
* Cycle guard: emissions whose own Source matches the storm-finding
source flow through Add without tripping a storm.
* Filter: emissions against the synthetic patrol-runtime resource
("ai-service") flow through Add without tripping a storm.
* Resolve: after a synthetic quiet window the throttler returns a
sentinel that `ResolveWithReason` converts into an auto-resolved
storm finding with the expected reason string.
The resolve case is exercised by calling `observeLocked` directly
with a synthetic future time (then dispatching the sentinel through
`ResolveWithReason`) to avoid a 120s wall-clock wait in the test
binary; the production hook in `Add` follows the same path on the
next real emission against the cluster.
No production-code changes in this commit.
A single noisy resource that trips several patrol detectors in a tight
window currently emits one finding per symptom. The findings panel
fans those out, which is the right surface for the detectors but the
wrong surface for the operator: a flapping VM with CPU, swap, and
disk pressure should read as one "this resource is in a bad state"
event, not three concurrent rows competing for attention.
This commit lands the throttler substrate. The new
`findingStormThrottler` is invoked from `(*FindingsStore).Add`'s
new-finding branch (while `s.mu` is held), clusters emissions by
`Finding.ResourceID`, and when a cluster crosses
`stormThreshold` (3) within `stormWindow` (60s) returns a storm
finding for the caller to re-enter through `s.Add`. The stable ID
`finding-storm:<resourceID>` dedups subsequent emissions through the
existing-finding branch, and the cycle guard on `Source ==
"finding-storm"` keeps the storm finding's own emission from
recursing back into the observer.
Auto-resolve is lazy by design: on the next observe trip where a
cluster has fallen below threshold AND the storm finding has been
quiet for at least `2 * stormWindow`, the observer returns a
sentinel that the caller routes through `ResolveWithReason`. No
sweep goroutine, no extra wiring surface — the substrate is
contained in `internal/ai/` and exposes a single nil-safe setter
(`SetStormThrottler`).
This commit is observer-only: nothing constructs a throttler yet, so
in-tree behavior is unchanged. Wire-up lands in a follow-up.
Wire the alerts manager's new flapping-detected callback in the AI
intelligence initialization path. Two things happen on each first
transition into the flapping cooldown window for a tracking key:
1. A reliability-category finding is written directly to the findings
store via emitFlappingPostmortemFinding. Path B from the lane brief:
the finding is durable without depending on patrol synthesis, so
the operator sees the diagnosis the moment Pulse decides to
suppress. The finding ID is derived from the canonical tracking
key ("alert-flapping:<trackingKey>") so re-detection inside the
cooldown window folds into the existing record via the same-ID
branch of FindingsStore.Add -- one finding per flapping condition,
not one per dispatch.
2. A scoped FlappingPostmortemPatrolScope is enqueued on the trigger
manager so an actual patrol run can enrich the finding with deeper
context once it lands.
The finding body names the flapping threshold, window, and cooldown
the manager is currently configured with, plus an action hint
(widen threshold, raise cooldown, or stabilise the resource). That
turns the suppressed alert from silence into a closable item on the
FindingsPanel.
FindingCategoryReliability is reused; no new category, no parent/
child finding structure -- those are deferred per the lane brief.