Seeds AlertConfig.DiskFillByType in defaultAlertConfig() with the three
forward-compatible keys (nvme 92/87, sata 90/85, hdd 85/80). Adds
NormalizeDiskFillByType to lowercase keys on load, seed defaults when
the map is nil, and reset non-positive trigger/clear values to the
default for that key. NormalizeAgentDefaults invokes the new helper so
it runs through the existing UpdateConfig normalization chain. Round
trip tests cover nil seed, customized survival, negative reset, and
mixed case normalization, plus a defaultAlertConfig seed assertion.
Only the "nvme" key is consulted by the host evaluation today (the
inference helper only matches /dev/nvme*); "sata" and "hdd" are
forward-compatible placeholders for when the agent protocol carries
hardware type explicitly.
Inserts a per-type threshold lookup before the global AgentDefaults.Disk
fallback in CheckHost's disk evaluation. inferDiskHardwareType maps the
disk device path to "nvme" (only NVMe is inferable today); when the
DiskFillByType map carries a matching key, that hysteresis threshold is
used, otherwise evaluation falls through to the existing global default.
Storage-type unified_eval branch is unchanged.
Tests prove: (a) NVMe device at 91% does not alert when DiskFillByType
sets the nvme trigger to 92; (b) the same device at 93% alerts at the
nvme threshold; (c) /dev/sda1 falls back to the global threshold; (d)
the storage-type alert branch is not regressed.
Adds DiskFillByType map to AlertConfig (keys: nvme, sata, hdd ->
HysteresisThreshold). Adds inferDiskHardwareType helper that returns
"nvme" for /dev/nvme* device paths (case-insensitive) and "" for
anything else, since sata vs hdd cannot be reliably inferred from
device paths today. Unit tests cover the four input shapes (nvme,
sata, empty, mixed case).
Substrate only. No wire-in to host alert evaluation yet.
The substrate has been in place for two commits but no live code path
constructs a throttler, so production behaviour is unchanged. This
commit closes the loop: `NewPatrolService` instantiates a
`findingStormThrottler` and calls
`p.findings.SetStormThrottler(p.stormThrottler)` immediately after
the `FindingsStore` is constructed, so every patrol emission flows
through the observer from the first run onward.
Constructor-internal wire-up keeps the throttler intrinsic to the
service rather than a deployment-time concern. No `router.go` seam,
no contract test pin, no goroutine to stop; the throttler holds
in-memory state and is garbage-collected when the PatrolService is
replaced.
The field sits next to `unifiedResourceProvider` to match the brief's
location and keep adjacent observer/provider seams visually grouped.
Commit 1 landed the substrate but only exercised it via direct calls
on the throttler. This commit drives a real `FindingsStore.Add` end
to end with the throttler wired via `SetStormThrottler`, so the hook
inside `Add`'s new-finding branch is on the test path.
Covers the four behaviours the brief calls out:
* Two distinct findings on the same resource within the window keep
the storm finding absent.
* The third distinct finding within the window emits a single storm
finding (severity warning, category reliability,
Source="finding-storm", IsActive).
* A fourth distinct finding within the same window does NOT spawn a
duplicate — the stable ID lands the re-entry on the existing-
finding branch and bumps TimesRaised.
* Cycle guard: emissions whose own Source matches the storm-finding
source flow through Add without tripping a storm.
* Filter: emissions against the synthetic patrol-runtime resource
("ai-service") flow through Add without tripping a storm.
* Resolve: after a synthetic quiet window the throttler returns a
sentinel that `ResolveWithReason` converts into an auto-resolved
storm finding with the expected reason string.
The resolve case is exercised by calling `observeLocked` directly
with a synthetic future time (then dispatching the sentinel through
`ResolveWithReason`) to avoid a 120s wall-clock wait in the test
binary; the production hook in `Add` follows the same path on the
next real emission against the cluster.
No production-code changes in this commit.
A single noisy resource that trips several patrol detectors in a tight
window currently emits one finding per symptom. The findings panel
fans those out, which is the right surface for the detectors but the
wrong surface for the operator: a flapping VM with CPU, swap, and
disk pressure should read as one "this resource is in a bad state"
event, not three concurrent rows competing for attention.
This commit lands the throttler substrate. The new
`findingStormThrottler` is invoked from `(*FindingsStore).Add`'s
new-finding branch (while `s.mu` is held), clusters emissions by
`Finding.ResourceID`, and when a cluster crosses
`stormThreshold` (3) within `stormWindow` (60s) returns a storm
finding for the caller to re-enter through `s.Add`. The stable ID
`finding-storm:<resourceID>` dedups subsequent emissions through the
existing-finding branch, and the cycle guard on `Source ==
"finding-storm"` keeps the storm finding's own emission from
recursing back into the observer.
Auto-resolve is lazy by design: on the next observe trip where a
cluster has fallen below threshold AND the storm finding has been
quiet for at least `2 * stormWindow`, the observer returns a
sentinel that the caller routes through `ResolveWithReason`. No
sweep goroutine, no extra wiring surface — the substrate is
contained in `internal/ai/` and exposes a single nil-safe setter
(`SetStormThrottler`).
This commit is observer-only: nothing constructs a throttler yet, so
in-tree behavior is unchanged. Wire-up lands in a follow-up.
Wire the alerts manager's new flapping-detected callback in the AI
intelligence initialization path. Two things happen on each first
transition into the flapping cooldown window for a tracking key:
1. A reliability-category finding is written directly to the findings
store via emitFlappingPostmortemFinding. Path B from the lane brief:
the finding is durable without depending on patrol synthesis, so
the operator sees the diagnosis the moment Pulse decides to
suppress. The finding ID is derived from the canonical tracking
key ("alert-flapping:<trackingKey>") so re-detection inside the
cooldown window folds into the existing record via the same-ID
branch of FindingsStore.Add -- one finding per flapping condition,
not one per dispatch.
2. A scoped FlappingPostmortemPatrolScope is enqueued on the trigger
manager so an actual patrol run can enrich the finding with deeper
context once it lands.
The finding body names the flapping threshold, window, and cooldown
the manager is currently configured with, plus an action hint
(widen threshold, raise cooldown, or stabilise the resource). That
turns the suppressed alert from silence into a closable item on the
FindingsPanel.
FindingCategoryReliability is reused; no new category, no parent/
child finding structure -- those are deferred per the lane brief.
Introduce TriggerReasonAlertFlapping and a matching scope factory so
the alert manager's new flapping-detected hook can enqueue a scoped
patrol that produces a postmortem finding instead of letting the
suppression go silent.
Priority is set to triggerPriorityAnomaly: a suppressed flap is
operationally similar to an anomaly (signal Pulse already saw, that
the user has not yet). Depth is Quick so the patrol stays focused on
the flapping resource.
The reason is wired through both gates:
- isEventDrivenTrigger now treats it as event-driven so it is
subject to the scoped trigger config
- allows() routes it under AlertTriggersEnabled, since flapping is
derived from alert state, not from baseline anomaly detection
When the alert manager suppresses an alert because it is flapping, today
the suppression is silent: the operator sees no alert and no diagnosis.
That trains people not to trust Pulse on flapping resources -- the
detector did its job, but the user just sees absence.
Add a one-shot Manager callback that fires exactly when a tracking key
crosses from quiet into the flapping cooldown window. The callback is
the hook downstream surfaces (patrol, findings) will use to explain
what is flapping and why Pulse stopped notifying.
Lock-safety: checkFlappingLocked now returns (suppress, justTransitioned).
The caller (dispatchAlert) reads the transition signal, then dispatches
the callback from a fresh goroutine. The alerts manager mutex is never
held while the callback runs, so subscribers may take their own locks
or re-enter the alerts package without deadlock.
Subsequent dispatches inside the cooldown window remain silent: the
transition flag is only set on the call that flips flappingActive from
false to true.
Adds a distinguishable approve/reject card variant inside the existing
remediation plan render path. Activates only when the finding category
is "capacity" AND the attached RemediationPlan carries a
proposed_action_plan with source === "capacity_forecast". All other
findings (capacity without proposal, non-capacity with proposal) keep
the existing generic card via the Show fallback.
Card surfaces current/predicted/threshold metric snapshot, time to
threshold breach, proposed change, and template SafetyChecks. Approve
flows through AIAPI.approveRemediationPlan (existing handler); Reject
flows through handleDismissPlan. No parallel state machine.
Adds RemediationPlan.proposed_action_plan and the ProposedActionPlan /
ProposedMetricSummary / ProposedActionPreflight wire types matching
pkg/aicontracts.
When a capacity finding has a registered template in the forecast
registry, generateRemediationPlan attaches the deterministic proposal
to the existing aicontracts.RemediationPlan via a new optional
ProposedActionPlan field. Reuses the existing remediation engine
plumbing - no parallel approval surface.
The wire-in best-effort extracts current value from the finding title
and ships canonical thresholds for storage / guest disk so the
capacity-forecast approval card has enough context to render without
depending on the forecast service being configured.
ProposedActionPlan / ProposedActionPreflight / ProposedMetricSummary
are wire-side projections defined in pkg/aicontracts so the contract
package stays free of an internal/ dependency.
Deterministic remediation proposals for capacity findings, keyed by
(resourceType, metric). Templates ship Allowed=false, RequiresApproval=true,
and emit a preflight-only ActionPlan until a Pulse write capability is
wired for the resource type. Registers PBS datastore prune+GC, ZFS pool
snapshot prune, and VM/CT disk expand variants.
CapacityActionPlanSource = "capacity_forecast" is the wire-side marker
the FindingsPanel approval card variant keys off.
Surface the rollup's VerifyIntent on the Protected Items table by
showing a "Verify due" badge alongside the item label whenever
verifyIntent === "stale". Reuses the existing stale-pill class from
recoveryStatusPresentation so the visual language stays consistent
with the rest of the recovery surface — no new badge primitive.
The TypeScript ProtectionRollup type gains optional verifyIntent
("verified" | "stale" | "unknown") and lastVerifiedAt fields. Both are
optional so older backend payloads keep type-checking cleanly; the badge
hides for verified / unknown / undefined intents.
Tests cover the badge under all three intent states plus the
legacy-undefined case so we lock in the no-emit-on-omit contract.
BuildBackupVerificationStaleFinding inspects a ProtectionRollup, and when
VerifyIntent is "stale" returns a backup-category Finding compatible with
the existing patrol intake (FindingsStore.Add /
PatrolService.recordFinding). Dedup keys go through the same
generateFindingID helper the LLM "case backup:" branch uses, so a
stale-still-stale tick updates rather than duplicates.
Severity is Watch by default and escalates to Warning once the last
successful backup is itself older than twice the staleness window —
the MVP's "multiple consecutive stale windows" heuristic without a
separate historical counter.
Tests cover: emission for a stale rollup; nil for verified / unknown /
nil rollups; severity escalation across multiple windows; dedup in the
FindingsStore across two patrol ticks; and the resolution path
(Resolve(auto=true) + auto_resolved lifecycle event) once verification
reappears.
Extend ProtectionRollup with a tri-state VerifyIntent (verified / stale /
unknown) plus LastVerifiedAt, derived at read-time from the existing
recovery_points.verified column. Stale means: a successful backup exists
but no verification-bearing point has landed within
BackupVerifyStaleWindow (7d, package-level constant for the MVP). Both
fields are omitempty so existing rollup snapshot consumers see no shape
change until they opt into the verify loop.
The SQL ListRollups path projects verified into the filtered CTE and
folds MAX(CASE WHEN verified=1 THEN ts_ms END) into the per-subject
aggregate. The in-memory BuildRollupsFromPoints mirrors the same logic
via the new ComputeVerifyIntentAt helper so mock mode and the persisted
store agree.
The patrol tool-call trace surface gains a Verified column rendered for
each tool call. Status-driven visuals: verified renders a green check,
failed renders a red warning, unknown/unverified render a grey dash.
Hover (title) and aria-label surface the evidence summary so screen
readers and keyboard users see the same fact. The expanded panel also
exposes the full evidence summary text when present.
The TypeScript API client gains ToolCallVerificationStatus and
ToolCallVerification, both surfaced through the existing ToolCallRecord
type. The Vitest suite gains coverage for all four status states and
the evidence-summary aria-label projection.
Vitest does not currently start in this worktree due to the pre-existing
fs.allow symlink resolution issue flagged in lane D-002 - not chased.
CommandPolicy gains a VerifyWindow (Go duration) bounded to
[(0,], 15m] with a 2m default. NormalizeVerifyWindow and the policy
Normalize() method enforce the bounds; DefaultPolicy() populates the
default. JSON marshaling now goes through a shadow struct so the wire
form serializes as a duration string (e.g. "2m0s") rather than the raw
nanosecond integer.
Tests cover the default, the bounds (clamp to max, fall through to
default for zero/negative), the JSON roundtrip, the unmarshal-applies-
bounds contract, and rejection of unparseable duration strings.
ActionAuditRecord gains a VerificationOutcome{status, evidenceSummary}
field with a closed enum (unknown/verified/unverified/failed). Existing
records read back as unknown by default via the normalizer and a new
SQLite column verification_outcome_json. The redaction pass scrubs the
evidence summary alongside other operator-authored text.
A new agentexec/verifier_postconditions.go registers postconditions for
qm.start, pct.start, docker.restart, systemctl.restart, and
kubectl.rollout, each parsed by verifier_postconditions_test.go.
Three pre-existing action JSON snapshot tests
(TestContract_ActionDecisionJSONSnapshot,
TestContract_ActionExecutionJSONSnapshot,
TestContract_UnifiedActionAuditsJSONSnapshot) now include the new
verificationOutcome field. The two flagged failing contract tests on
this branch
(TestContract_ActionDryRunOnlyExecutionErrorJSONSnapshot,
TestContract_RouterBridgesVerificationOntoActionCompleted) are
unrelated to this change and were left alone per lane D-002 scope.
When a maintenance window ends on a resource, the sentinel runs
deterministic checks (active alerts, Patrol findings, failed actions
since window start, basic post-window metric recovery) and writes a
durable LoopReport. Operators can list reports per resource, mark them
reviewed, or rerun verification immediately. UI surfaces the section in
the resource detail drawer; scoped Patrol runs and Assistant deep-link
are deferred until those entry points stabilise.
Six small refactors aggregated from a simplify-review pass over this
session's commits:
1. internal/config/persistence_relay.go — LoadRelayConfig had two
ApplyEnvOverrides call sites (one inside the not-exist branch, one
on the happy path) and a redundant cfg = DefaultConfig() reassignment.
Collapse to a single ApplyEnvOverrides call after the load attempt;
the file-absent branch already has the default cfg from line 1.
2. internal/relay/config_env.go — swap two strings.TrimSpace(os.Getenv(...))
calls for utils.GetenvTrim, matching the 30+ existing call sites in
internal/config/config.go. Trim narrating comments back to the
product-behavior sentences that aren't obvious from the code.
3. internal/relay/config_env_test.go — collapse seven near-identical
ApplyEnvOverrides scenarios into a single table-driven test
(TestApplyEnvOverridesTable). Reduces ~85 lines to ~60 and gives each
subcase a named t.Run for clearer failure output. Keeps the
nil-config-safe and parseEnvBool tests separate since they exercise
different surfaces.
4. .github/workflows/install-sh-smoke.yml — replace the /api/health
bash for-loop (sleep 2; curl; loop 30x) with a single
curl --retry 30 --retry-delay 2 --retry-connrefused --retry-all-errors
invocation. Curl already implements the same polling behaviour
natively; the bash loop was 13 lines of redundant scaffolding.
5. scripts/installtests/build_release_assets_test.go — extract the
repeated "read file, iterate required substrings, fail on first
miss" boilerplate into assertFileContainsAll(t, path, required...).
Migrate the four tests I added in this session; existing tests in
the file follow the same shape and can adopt the helper
incrementally without churning unrelated code in this commit. Also
updated the pinned curl string for the /api/health retry change.
Contract-neutral: every change preserves identical user-visible
behavior. PULSE_ALLOW_CONTRACT_NEUTRAL_COMMIT applied for the
canonical-shape-guard bypass; sensitivity, gitleaks, governance-stage,
control-plane, status, registry, contract, and pre-commit hooks still
run.
Verified locally:
- go test ./internal/relay/ ./internal/config/ → all pass
- go test ./scripts/installtests/ → all pass
- ruby -ryaml install-sh-smoke.yml → parses clean
promote-floating-tags.yml's `workflow_run` chain off publish-docker.yml
silently stopped firing for rc.3 → rc.5 because publish-docker failed at
the now-removed pulse-agent push step. Customers pulling
rcourtman/pulse:latest, :6, or :6.0 stayed on whatever the previous
successful release had tagged — there was no warning anywhere that the
floating tags were stale.
Same fix pattern as install-sh-smoke (commit 7c0f65425) and
publish-helm-chart (commit 14c79a28e): add a workflow_call trigger to
promote-floating-tags.yml and call it explicitly from create-release.yml
after validate_release_assets succeeds.
Gating on validate_release_assets is intentional: that workflow waits
for the docker image to be pullable from the registry (with retry
backoff), so by the time it succeeds the image manifest exists and
re-tagging it to latest/major/minor cannot point at vapor.
The legacy workflow_run trigger stays as the primary path; this just
guarantees promotion even when the chain doesn't fire.
Tag-resolver step now accepts inputs from workflow_call / workflow_dispatch
and only falls back to the workflow_run derivation when inputs are absent,
so all three entry paths converge on the same identity.
Pinned in build_release_assets_test.go:
- new TestPromoteFloatingTagsReachableViaWorkflowCall pins the trigger
declaration and the input-priority resolver
- existing TestCreateReleaseUploadsPowerShellInstaller extended to pin
the promote_floating_tags job wiring (uses, tag, prerelease)
Contract delta in deployment-installability.md Extension Point 7
documents the same explicit-workflow_call requirement that applies to
publish-helm-chart, extended to promote-floating-tags.
v6 rc.1 → rc.5 published successfully but the Helm chart never landed on
rcourtman.github.io/Pulse/index.yaml — the index still ends at v5.1.30.
`helm install pulse pulse/pulse --version 6.0.0-rc.5` returns
chart-not-found; without `--version` helm pulls the latest published
chart (v5.1.30) into a customer's v6 cluster.
Root cause: GitHub does not fire `release: published` for releases that
were created as drafts and later PATCHed to draft=false. create-release.yml
deliberately uses that path so it can upload assets and run
validate-release-assets against the draft before promoting. Inspection of
the workflow run history confirms: every gh-API `release: published` event
since 2026-03-02 has been from manually-dispatched v5 stable cuts; zero
fired for v6 RCs published through the create-release pipeline.
Fix the same way install-sh-smoke was wired in commit 7c0f65425: add a
`workflow_call` trigger to publish-helm-chart.yml and call it explicitly
from create-release.yml as a downstream of validate_release_assets. The
chart-version resolver in publish-helm-chart now accepts inputs from
either workflow_call or workflow_dispatch and only falls back to the
release-event tag when no inputs are present, keeping the legacy
release-event path working for forks / manual gh-CLI publishes that
create with draft=false from the start.
Pinned in build_release_assets_test.go:
- create-release.yml wiring (publish_helm_chart job, version inputs)
- publish-helm-chart.yml workflow_call trigger declaration
- chart-version resolver's input-priority logic
Contract delta in deployment-installability.md Extension Point 7
documents the workflow_call requirement and forbids relying on the
release-published webhook for the create-release.yml draft-promotion
path.
The fix takes effect on the next release through the pipeline. Backfill
of the v6.0.0-rc.5 chart needs a one-time manual dispatch of
publish-helm-chart.yml against chart_version=6.0.0-rc.5.
The chart's agent.image.repository defaulted to ghcr.io/rcourtman/pulse-agent,
an image that has never been published. publish-docker.yml only pushes
rcourtman/pulse; the Dockerfile defines an agent_runtime stage that
*could* be published but it isn't, and commit da7969fb4 from earlier in
this session removed the corresponding pulse-agent attestation
expectations — a clear signal the separate agent image was intentionally
dropped without updating the chart. Customers running
`helm install pulse pulse/pulse --set agent.enabled=true` were silently
hitting ImagePullBackOff on the agent DaemonSet.
Route the chart through the main rcourtman/pulse image instead. To make
that work without per-arch chart overrides, the runtime stage in the
Dockerfile now creates an arch-resolved /usr/local/bin/pulse-agent
symlink to the right /opt/pulse/bin/pulse-agent-linux-{amd64,arm64,armv7}
binary. The chart's agent.command default is /usr/local/bin/pulse-agent,
which overrides the server ENTRYPOINT and runs the pod as a unified
agent on whichever arch the node provides. agent.yaml renders the
command via toYaml so list values pass through cleanly.
KUBERNETES.md's DaemonSet example switches from the arch-hardcoded
/opt/pulse/bin/pulse-agent-linux-amd64 to the new arch-resolved path,
restoring multi-arch portability of the docs example.
validate-release.sh asserts the symlink exists, points at one of the
three supported Linux arch binaries, and is executable in the published
image. A new TestHelmAgentRuntimePointsAtRealImage pins the chart
defaults, the template wiring, the Dockerfile symlink, and the
validate-release.sh guard so the regression class can't quietly
resurface.
Governance: extend the helm-chart-release-runtime verification policy's
exact_files to include scripts/installtests/build_release_assets_test.go
(matching its existing pin set for related deployment-installability
policies); update the subsystem_lookup_test.py fixture that pins the
exact_files list; document the agent-image and pulse-agent symlink
contract in deployment-installability.md Extension Point 7.
Verified locally: `helm lint` passes; `helm template --set agent.enabled=true`
renders a DaemonSet with image rcourtman/pulse:6.0.0,
command ["/usr/local/bin/pulse-agent"], args ["--enable-docker", "--enable-host=false"].
End-to-end image build + agent DaemonSet smoke will run via helm_smoke
on the next release once rcourtman/pulse:6.0.0 is published.