Commit Graph

6114 Commits

Author SHA1 Message Date
rcourtman 8197fe6b1e Use disk type thresholds for SMART temperatures 2026-05-13 10:44:11 +01:00
rcourtman 3cecf9576d Seed disk temperature type defaults 2026-05-13 10:44:11 +01:00
rcourtman becee760dd Add disk temperature by type normalization 2026-05-13 10:44:11 +01:00
rcourtman 950bd187a9 seed DiskFillByType defaults and normalize on load
Seeds AlertConfig.DiskFillByType in defaultAlertConfig() with the three
forward-compatible keys (nvme 92/87, sata 90/85, hdd 85/80). Adds
NormalizeDiskFillByType to lowercase keys on load, seed defaults when
the map is nil, and reset non-positive trigger/clear values to the
default for that key. NormalizeAgentDefaults invokes the new helper so
it runs through the existing UpdateConfig normalization chain. Round
trip tests cover nil seed, customized survival, negative reset, and
mixed case normalization, plus a defaultAlertConfig seed assertion.

Only the "nvme" key is consulted by the host evaluation today (the
inference helper only matches /dev/nvme*); "sata" and "hdd" are
forward-compatible placeholders for when the agent protocol carries
hardware type explicitly.
2026-05-13 07:16:07 +01:00
rcourtman 85fbbd4ef6 consult DiskFillByType in host disk fill alert evaluation
Inserts a per-type threshold lookup before the global AgentDefaults.Disk
fallback in CheckHost's disk evaluation. inferDiskHardwareType maps the
disk device path to "nvme" (only NVMe is inferable today); when the
DiskFillByType map carries a matching key, that hysteresis threshold is
used, otherwise evaluation falls through to the existing global default.
Storage-type unified_eval branch is unchanged.

Tests prove: (a) NVMe device at 91% does not alert when DiskFillByType
sets the nvme trigger to 92; (b) the same device at 93% alerts at the
nvme threshold; (c) /dev/sda1 falls back to the global threshold; (d)
the storage-type alert branch is not regressed.
2026-05-13 07:16:07 +01:00
rcourtman 005873bdd5 add disk-type-aware threshold field and device inference helper
Adds DiskFillByType map to AlertConfig (keys: nvme, sata, hdd ->
HysteresisThreshold). Adds inferDiskHardwareType helper that returns
"nvme" for /dev/nvme* device paths (case-insensitive) and "" for
anything else, since sata vs hdd cannot be reliably inferred from
device paths today. Unit tests cover the four input shapes (nvme,
sata, empty, mixed case).

Substrate only. No wire-in to host alert evaluation yet.
2026-05-13 07:13:46 +01:00
rcourtman 99d0edb04a wire PDM alert bridge into PatrolService and seed demo 2026-05-13 05:42:25 +01:00
rcourtman b9844f22c1 prove PDM alert bridge emits and resolves through FindingsStore.Add 2026-05-13 05:42:25 +01:00
rcourtman 8aad44174b add PDM alert bridge substrate in internal/ai 2026-05-13 05:42:25 +01:00
rcourtman 4097ddcb9c handle repeated update-safety digest changes 2026-05-13 04:18:27 +01:00
rcourtman 5821b6c810 wire update-safety watcher into PatrolService and seed demo 2026-05-13 04:16:35 +01:00
rcourtman 2e87e374eb prove update-safety watcher emits and resolves through FindingsStore.Add 2026-05-13 04:16:35 +01:00
rcourtman ab71410ccb add update-safety watcher substrate in internal/ai 2026-05-13 04:16:35 +01:00
rcourtman 178c961723 wire storm throttler into PatrolService from the constructor
The substrate has been in place for two commits but no live code path
constructs a throttler, so production behaviour is unchanged. This
commit closes the loop: `NewPatrolService` instantiates a
`findingStormThrottler` and calls
`p.findings.SetStormThrottler(p.stormThrottler)` immediately after
the `FindingsStore` is constructed, so every patrol emission flows
through the observer from the first run onward.

Constructor-internal wire-up keeps the throttler intrinsic to the
service rather than a deployment-time concern. No `router.go` seam,
no contract test pin, no goroutine to stop; the throttler holds
in-memory state and is garbage-collected when the PatrolService is
replaced.

The field sits next to `unifiedResourceProvider` to match the brief's
location and keep adjacent observer/provider seams visually grouped.
2026-05-13 02:41:45 +01:00
rcourtman 38a471120d prove storm throttler emits and resolves through FindingsStore.Add
Commit 1 landed the substrate but only exercised it via direct calls
on the throttler. This commit drives a real `FindingsStore.Add` end
to end with the throttler wired via `SetStormThrottler`, so the hook
inside `Add`'s new-finding branch is on the test path.

Covers the four behaviours the brief calls out:

* Two distinct findings on the same resource within the window keep
  the storm finding absent.
* The third distinct finding within the window emits a single storm
  finding (severity warning, category reliability,
  Source="finding-storm", IsActive).
* A fourth distinct finding within the same window does NOT spawn a
  duplicate — the stable ID lands the re-entry on the existing-
  finding branch and bumps TimesRaised.
* Cycle guard: emissions whose own Source matches the storm-finding
  source flow through Add without tripping a storm.
* Filter: emissions against the synthetic patrol-runtime resource
  ("ai-service") flow through Add without tripping a storm.
* Resolve: after a synthetic quiet window the throttler returns a
  sentinel that `ResolveWithReason` converts into an auto-resolved
  storm finding with the expected reason string.

The resolve case is exercised by calling `observeLocked` directly
with a synthetic future time (then dispatching the sentinel through
`ResolveWithReason`) to avoid a 120s wall-clock wait in the test
binary; the production hook in `Add` follows the same path on the
next real emission against the cluster.

No production-code changes in this commit.
2026-05-13 02:41:45 +01:00
rcourtman 8f9ff92701 add storm throttler substrate in FindingsStore.Add
A single noisy resource that trips several patrol detectors in a tight
window currently emits one finding per symptom. The findings panel
fans those out, which is the right surface for the detectors but the
wrong surface for the operator: a flapping VM with CPU, swap, and
disk pressure should read as one "this resource is in a bad state"
event, not three concurrent rows competing for attention.

This commit lands the throttler substrate. The new
`findingStormThrottler` is invoked from `(*FindingsStore).Add`'s
new-finding branch (while `s.mu` is held), clusters emissions by
`Finding.ResourceID`, and when a cluster crosses
`stormThreshold` (3) within `stormWindow` (60s) returns a storm
finding for the caller to re-enter through `s.Add`. The stable ID
`finding-storm:<resourceID>` dedups subsequent emissions through the
existing-finding branch, and the cycle guard on `Source ==
"finding-storm"` keeps the storm finding's own emission from
recursing back into the observer.

Auto-resolve is lazy by design: on the next observe trip where a
cluster has fallen below threshold AND the storm finding has been
quiet for at least `2 * stormWindow`, the observer returns a
sentinel that the caller routes through `ResolveWithReason`. No
sweep goroutine, no extra wiring surface — the substrate is
contained in `internal/ai/` and exposes a single nil-safe setter
(`SetStormThrottler`).

This commit is observer-only: nothing constructs a throttler yet, so
in-tree behavior is unchanged. Wire-up lands in a follow-up.
2026-05-13 02:41:45 +01:00
rcourtman e84afb140e render Overdue commitments filter in FindingsPanel 2026-05-13 01:13:38 +01:00
rcourtman d4f84e69e0 add hourly sweep for overdue will_fix_later commitments 2026-05-13 01:13:38 +01:00
rcourtman 183bf5b617 persist remind_at and remind_count on will_fix_later findings 2026-05-13 01:13:38 +01:00
rcourtman 0e0c90da53 emit reliability finding when an alert starts flapping
Wire the alerts manager's new flapping-detected callback in the AI
intelligence initialization path. Two things happen on each first
transition into the flapping cooldown window for a tracking key:

1. A reliability-category finding is written directly to the findings
   store via emitFlappingPostmortemFinding. Path B from the lane brief:
   the finding is durable without depending on patrol synthesis, so
   the operator sees the diagnosis the moment Pulse decides to
   suppress. The finding ID is derived from the canonical tracking
   key ("alert-flapping:<trackingKey>") so re-detection inside the
   cooldown window folds into the existing record via the same-ID
   branch of FindingsStore.Add -- one finding per flapping condition,
   not one per dispatch.

2. A scoped FlappingPostmortemPatrolScope is enqueued on the trigger
   manager so an actual patrol run can enrich the finding with deeper
   context once it lands.

The finding body names the flapping threshold, window, and cooldown
the manager is currently configured with, plus an action hint
(widen threshold, raise cooldown, or stabilise the resource). That
turns the suppressed alert from silence into a closable item on the
FindingsPanel.

FindingCategoryReliability is reused; no new category, no parent/
child finding structure -- those are deferred per the lane brief.
2026-05-13 00:28:24 +01:00
rcourtman fb05c38ec7 add patrol scope for flapping postmortem
Introduce TriggerReasonAlertFlapping and a matching scope factory so
the alert manager's new flapping-detected hook can enqueue a scoped
patrol that produces a postmortem finding instead of letting the
suppression go silent.

Priority is set to triggerPriorityAnomaly: a suppressed flap is
operationally similar to an anomaly (signal Pulse already saw, that
the user has not yet). Depth is Quick so the patrol stays focused on
the flapping resource.

The reason is wired through both gates:
  - isEventDrivenTrigger now treats it as event-driven so it is
    subject to the scoped trigger config
  - allows() routes it under AlertTriggersEnabled, since flapping is
    derived from alert state, not from baseline anomaly detection
2026-05-13 00:28:24 +01:00
rcourtman fa9cc5d691 fire flapping-detected callback on first transition
When the alert manager suppresses an alert because it is flapping, today
the suppression is silent: the operator sees no alert and no diagnosis.
That trains people not to trust Pulse on flapping resources -- the
detector did its job, but the user just sees absence.

Add a one-shot Manager callback that fires exactly when a tracking key
crosses from quiet into the flapping cooldown window. The callback is
the hook downstream surfaces (patrol, findings) will use to explain
what is flapping and why Pulse stopped notifying.

Lock-safety: checkFlappingLocked now returns (suppress, justTransitioned).
The caller (dispatchAlert) reads the transition signal, then dispatches
the callback from a fresh goroutine. The alerts manager mutex is never
held while the callback runs, so subscribers may take their own locks
or re-enter the alerts package without deadlock.

Subsequent dispatches inside the cooldown window remain silent: the
transition flag is only set on the call that flips flappingActive from
false to true.
2026-05-13 00:28:24 +01:00
rcourtman 2c81815f5f render capacity-forecast approval card in FindingsPanel
Adds a distinguishable approve/reject card variant inside the existing
remediation plan render path. Activates only when the finding category
is "capacity" AND the attached RemediationPlan carries a
proposed_action_plan with source === "capacity_forecast". All other
findings (capacity without proposal, non-capacity with proposal) keep
the existing generic card via the Show fallback.

Card surfaces current/predicted/threshold metric snapshot, time to
threshold breach, proposed change, and template SafetyChecks. Approve
flows through AIAPI.approveRemediationPlan (existing handler); Reject
flows through handleDismissPlan. No parallel state machine.

Adds RemediationPlan.proposed_action_plan and the ProposedActionPlan /
ProposedMetricSummary / ProposedActionPreflight wire types matching
pkg/aicontracts.
2026-05-12 23:41:55 +01:00
rcourtman 2e064e6c6c attach forecast proposals onto patrol RemediationPlan
When a capacity finding has a registered template in the forecast
registry, generateRemediationPlan attaches the deterministic proposal
to the existing aicontracts.RemediationPlan via a new optional
ProposedActionPlan field. Reuses the existing remediation engine
plumbing - no parallel approval surface.

The wire-in best-effort extracts current value from the finding title
and ships canonical thresholds for storage / guest disk so the
capacity-forecast approval card has enough context to render without
depending on the forecast service being configured.

ProposedActionPlan / ProposedActionPreflight / ProposedMetricSummary
are wire-side projections defined in pkg/aicontracts so the contract
package stays free of an internal/ dependency.
2026-05-12 23:41:55 +01:00
rcourtman a5629d5701 add capacity-forecast action template registry
Deterministic remediation proposals for capacity findings, keyed by
(resourceType, metric). Templates ship Allowed=false, RequiresApproval=true,
and emit a preflight-only ActionPlan until a Pulse write capability is
wired for the resource type. Registers PBS datastore prune+GC, ZFS pool
snapshot prune, and VM/CT disk expand variants.

CapacityActionPlanSource = "capacity_forecast" is the wire-side marker
the FindingsPanel approval card variant keys off.
2026-05-12 23:41:55 +01:00
rcourtman b2e437b198 Separate Patrol finding controls from disclosure 2026-05-12 22:52:36 +01:00
rcourtman b213b59f7a Prioritize Patrol findings in Assistant handoffs 2026-05-12 22:46:53 +01:00
rcourtman 1e33d3089c Bound Patrol refresh state 2026-05-12 22:35:50 +01:00
rcourtman 19beb6399e render Verify due badge on Protected Items rows
Surface the rollup's VerifyIntent on the Protected Items table by
showing a "Verify due" badge alongside the item label whenever
verifyIntent === "stale". Reuses the existing stale-pill class from
recoveryStatusPresentation so the visual language stays consistent
with the rest of the recovery surface — no new badge primitive.

The TypeScript ProtectionRollup type gains optional verifyIntent
("verified" | "stale" | "unknown") and lastVerifiedAt fields. Both are
optional so older backend payloads keep type-checking cleanly; the badge
hides for verified / unknown / undefined intents.

Tests cover the badge under all three intent states plus the
legacy-undefined case so we lock in the no-emit-on-omit contract.
2026-05-12 22:26:46 +01:00
rcourtman e69633daf8 emit backup_verification_stale finding from VerifyIntent
BuildBackupVerificationStaleFinding inspects a ProtectionRollup, and when
VerifyIntent is "stale" returns a backup-category Finding compatible with
the existing patrol intake (FindingsStore.Add /
PatrolService.recordFinding). Dedup keys go through the same
generateFindingID helper the LLM "case backup:" branch uses, so a
stale-still-stale tick updates rather than duplicates.

Severity is Watch by default and escalates to Warning once the last
successful backup is itself older than twice the staleness window —
the MVP's "multiple consecutive stale windows" heuristic without a
separate historical counter.

Tests cover: emission for a stale rollup; nil for verified / unknown /
nil rollups; severity escalation across multiple windows; dedup in the
FindingsStore across two patrol ticks; and the resolution path
(Resolve(auto=true) + auto_resolved lifecycle event) once verification
reappears.
2026-05-12 22:23:55 +01:00
rcourtman bd62b6d8ad add VerifyIntent rollup substrate
Extend ProtectionRollup with a tri-state VerifyIntent (verified / stale /
unknown) plus LastVerifiedAt, derived at read-time from the existing
recovery_points.verified column. Stale means: a successful backup exists
but no verification-bearing point has landed within
BackupVerifyStaleWindow (7d, package-level constant for the MVP). Both
fields are omitempty so existing rollup snapshot consumers see no shape
change until they opt into the verify loop.

The SQL ListRollups path projects verified into the filtered CTE and
folds MAX(CASE WHEN verified=1 THEN ts_ms END) into the per-subject
aggregate. The in-memory BuildRollupsFromPoints mirrors the same logic
via the new ComputeVerifyIntentAt helper so mock mode and the persisted
store agree.
2026-05-12 22:23:55 +01:00
rcourtman bb04bf9042 Keep Patrol findings visible during refresh 2026-05-12 22:13:19 +01:00
rcourtman 5465852be0 Use Patrol source loading for findings panel 2026-05-12 22:06:09 +01:00
rcourtman d01c896e2c add verified column to RunToolCallTrace
The patrol tool-call trace surface gains a Verified column rendered for
each tool call. Status-driven visuals: verified renders a green check,
failed renders a red warning, unknown/unverified render a grey dash.
Hover (title) and aria-label surface the evidence summary so screen
readers and keyboard users see the same fact. The expanded panel also
exposes the full evidence summary text when present.

The TypeScript API client gains ToolCallVerificationStatus and
ToolCallVerification, both surfaced through the existing ToolCallRecord
type. The Vitest suite gains coverage for all four status states and
the evidence-summary aria-label projection.

Vitest does not currently start in this worktree due to the pre-existing
fs.allow symlink resolution issue flagged in lane D-002 - not chased.
2026-05-12 21:53:49 +01:00
rcourtman 7f38a3d8bf add verify_window to agentexec command policy
CommandPolicy gains a VerifyWindow (Go duration) bounded to
[(0,], 15m] with a 2m default. NormalizeVerifyWindow and the policy
Normalize() method enforce the bounds; DefaultPolicy() populates the
default. JSON marshaling now goes through a shadow struct so the wire
form serializes as a duration string (e.g. "2m0s") rather than the raw
nanosecond integer.

Tests cover the default, the bounds (clamp to max, fall through to
default for zero/negative), the JSON roundtrip, the unmarshal-applies-
bounds contract, and rejection of unparseable duration strings.
2026-05-12 21:53:49 +01:00
rcourtman 006821327f add verification outcome and capability postcondition substrate
ActionAuditRecord gains a VerificationOutcome{status, evidenceSummary}
field with a closed enum (unknown/verified/unverified/failed). Existing
records read back as unknown by default via the normalizer and a new
SQLite column verification_outcome_json. The redaction pass scrubs the
evidence summary alongside other operator-authored text.

A new agentexec/verifier_postconditions.go registers postconditions for
qm.start, pct.start, docker.restart, systemctl.restart, and
kubectl.rollout, each parsed by verifier_postconditions_test.go.

Three pre-existing action JSON snapshot tests
(TestContract_ActionDecisionJSONSnapshot,
TestContract_ActionExecutionJSONSnapshot,
TestContract_UnifiedActionAuditsJSONSnapshot) now include the new
verificationOutcome field. The two flagged failing contract tests on
this branch
(TestContract_ActionDryRunOnlyExecutionErrorJSONSnapshot,
TestContract_RouterBridgesVerificationOntoActionCompleted) are
unrelated to this change and were left alone per lane D-002 scope.
2026-05-12 21:53:49 +01:00
rcourtman 3cbd49cb54 Align Patrol coverage caveats with full runs 2026-05-12 21:33:44 +01:00
rcourtman 015e7f6555 Add maintenance verification reports
When a maintenance window ends on a resource, the sentinel runs
deterministic checks (active alerts, Patrol findings, failed actions
since window start, basic post-window metric recovery) and writes a
durable LoopReport. Operators can list reports per resource, mark them
reviewed, or rerun verification immediately. UI surfaces the section in
the resource detail drawer; scoped Patrol runs and Assistant deep-link
are deferred until those entry points stabilise.
2026-05-12 21:10:58 +01:00
rcourtman 5b18438528 Use checked wording for limited Patrol recency 2026-05-12 17:53:53 +01:00
rcourtman 80c061cd45 Separate Patrol runtime issues from findings copy 2026-05-12 17:43:28 +01:00
rcourtman 450e7516f0 Keep Patrol findings from showing all-clear copy 2026-05-12 17:34:30 +01:00
rcourtman 93c62e691a Aggregate simplify-review cleanups (no behavior change)
Six small refactors aggregated from a simplify-review pass over this
session's commits:

1. internal/config/persistence_relay.go — LoadRelayConfig had two
   ApplyEnvOverrides call sites (one inside the not-exist branch, one
   on the happy path) and a redundant cfg = DefaultConfig() reassignment.
   Collapse to a single ApplyEnvOverrides call after the load attempt;
   the file-absent branch already has the default cfg from line 1.

2. internal/relay/config_env.go — swap two strings.TrimSpace(os.Getenv(...))
   calls for utils.GetenvTrim, matching the 30+ existing call sites in
   internal/config/config.go. Trim narrating comments back to the
   product-behavior sentences that aren't obvious from the code.

3. internal/relay/config_env_test.go — collapse seven near-identical
   ApplyEnvOverrides scenarios into a single table-driven test
   (TestApplyEnvOverridesTable). Reduces ~85 lines to ~60 and gives each
   subcase a named t.Run for clearer failure output. Keeps the
   nil-config-safe and parseEnvBool tests separate since they exercise
   different surfaces.

4. .github/workflows/install-sh-smoke.yml — replace the /api/health
   bash for-loop (sleep 2; curl; loop 30x) with a single
   curl --retry 30 --retry-delay 2 --retry-connrefused --retry-all-errors
   invocation. Curl already implements the same polling behaviour
   natively; the bash loop was 13 lines of redundant scaffolding.

5. scripts/installtests/build_release_assets_test.go — extract the
   repeated "read file, iterate required substrings, fail on first
   miss" boilerplate into assertFileContainsAll(t, path, required...).
   Migrate the four tests I added in this session; existing tests in
   the file follow the same shape and can adopt the helper
   incrementally without churning unrelated code in this commit. Also
   updated the pinned curl string for the /api/health retry change.

Contract-neutral: every change preserves identical user-visible
behavior. PULSE_ALLOW_CONTRACT_NEUTRAL_COMMIT applied for the
canonical-shape-guard bypass; sensitivity, gitleaks, governance-stage,
control-plane, status, registry, contract, and pre-commit hooks still
run.

Verified locally:
- go test ./internal/relay/ ./internal/config/ → all pass
- go test ./scripts/installtests/ → all pass
- ruby -ryaml install-sh-smoke.yml → parses clean
2026-05-12 17:32:11 +01:00
rcourtman 22a94f47d9 Skip release publish downstreams for drafts 2026-05-12 17:32:11 +01:00
rcourtman 8b0f3564f6 Fail closed on stale API action plans 2026-05-12 17:32:11 +01:00
rcourtman 29a815ef2a Fail closed on stale API action plans 2026-05-12 16:55:51 +01:00
rcourtman 3566a4d61d Drive promote-floating-tags via workflow_call from create-release
promote-floating-tags.yml's `workflow_run` chain off publish-docker.yml
silently stopped firing for rc.3 → rc.5 because publish-docker failed at
the now-removed pulse-agent push step. Customers pulling
rcourtman/pulse:latest, :6, or :6.0 stayed on whatever the previous
successful release had tagged — there was no warning anywhere that the
floating tags were stale.

Same fix pattern as install-sh-smoke (commit 7c0f65425) and
publish-helm-chart (commit 14c79a28e): add a workflow_call trigger to
promote-floating-tags.yml and call it explicitly from create-release.yml
after validate_release_assets succeeds.

Gating on validate_release_assets is intentional: that workflow waits
for the docker image to be pullable from the registry (with retry
backoff), so by the time it succeeds the image manifest exists and
re-tagging it to latest/major/minor cannot point at vapor.

The legacy workflow_run trigger stays as the primary path; this just
guarantees promotion even when the chain doesn't fire.

Tag-resolver step now accepts inputs from workflow_call / workflow_dispatch
and only falls back to the workflow_run derivation when inputs are absent,
so all three entry paths converge on the same identity.

Pinned in build_release_assets_test.go:
- new TestPromoteFloatingTagsReachableViaWorkflowCall pins the trigger
  declaration and the input-priority resolver
- existing TestCreateReleaseUploadsPowerShellInstaller extended to pin
  the promote_floating_tags job wiring (uses, tag, prerelease)

Contract delta in deployment-installability.md Extension Point 7
documents the same explicit-workflow_call requirement that applies to
publish-helm-chart, extended to promote-floating-tags.
2026-05-12 16:47:51 +01:00
rcourtman 14c79a28e7 Trigger publish-helm-chart via workflow_call from create-release
v6 rc.1 → rc.5 published successfully but the Helm chart never landed on
rcourtman.github.io/Pulse/index.yaml — the index still ends at v5.1.30.
`helm install pulse pulse/pulse --version 6.0.0-rc.5` returns
chart-not-found; without `--version` helm pulls the latest published
chart (v5.1.30) into a customer's v6 cluster.

Root cause: GitHub does not fire `release: published` for releases that
were created as drafts and later PATCHed to draft=false. create-release.yml
deliberately uses that path so it can upload assets and run
validate-release-assets against the draft before promoting. Inspection of
the workflow run history confirms: every gh-API `release: published` event
since 2026-03-02 has been from manually-dispatched v5 stable cuts; zero
fired for v6 RCs published through the create-release pipeline.

Fix the same way install-sh-smoke was wired in commit 7c0f65425: add a
`workflow_call` trigger to publish-helm-chart.yml and call it explicitly
from create-release.yml as a downstream of validate_release_assets. The
chart-version resolver in publish-helm-chart now accepts inputs from
either workflow_call or workflow_dispatch and only falls back to the
release-event tag when no inputs are present, keeping the legacy
release-event path working for forks / manual gh-CLI publishes that
create with draft=false from the start.

Pinned in build_release_assets_test.go:
- create-release.yml wiring (publish_helm_chart job, version inputs)
- publish-helm-chart.yml workflow_call trigger declaration
- chart-version resolver's input-priority logic

Contract delta in deployment-installability.md Extension Point 7
documents the workflow_call requirement and forbids relying on the
release-published webhook for the create-release.yml draft-promotion
path.

The fix takes effect on the next release through the pipeline. Backfill
of the v6.0.0-rc.5 chart needs a one-time manual dispatch of
publish-helm-chart.yml against chart_version=6.0.0-rc.5.
2026-05-12 16:30:53 +01:00
rcourtman c6d5c4590a Keep agent heartbeats stream local 2026-05-12 16:14:28 +01:00
rcourtman ab62b46c1f Fix helm chart agent.enabled by routing through main pulse image
The chart's agent.image.repository defaulted to ghcr.io/rcourtman/pulse-agent,
an image that has never been published. publish-docker.yml only pushes
rcourtman/pulse; the Dockerfile defines an agent_runtime stage that
*could* be published but it isn't, and commit da7969fb4 from earlier in
this session removed the corresponding pulse-agent attestation
expectations — a clear signal the separate agent image was intentionally
dropped without updating the chart. Customers running
`helm install pulse pulse/pulse --set agent.enabled=true` were silently
hitting ImagePullBackOff on the agent DaemonSet.

Route the chart through the main rcourtman/pulse image instead. To make
that work without per-arch chart overrides, the runtime stage in the
Dockerfile now creates an arch-resolved /usr/local/bin/pulse-agent
symlink to the right /opt/pulse/bin/pulse-agent-linux-{amd64,arm64,armv7}
binary. The chart's agent.command default is /usr/local/bin/pulse-agent,
which overrides the server ENTRYPOINT and runs the pod as a unified
agent on whichever arch the node provides. agent.yaml renders the
command via toYaml so list values pass through cleanly.

KUBERNETES.md's DaemonSet example switches from the arch-hardcoded
/opt/pulse/bin/pulse-agent-linux-amd64 to the new arch-resolved path,
restoring multi-arch portability of the docs example.
validate-release.sh asserts the symlink exists, points at one of the
three supported Linux arch binaries, and is executable in the published
image. A new TestHelmAgentRuntimePointsAtRealImage pins the chart
defaults, the template wiring, the Dockerfile symlink, and the
validate-release.sh guard so the regression class can't quietly
resurface.

Governance: extend the helm-chart-release-runtime verification policy's
exact_files to include scripts/installtests/build_release_assets_test.go
(matching its existing pin set for related deployment-installability
policies); update the subsystem_lookup_test.py fixture that pins the
exact_files list; document the agent-image and pulse-agent symlink
contract in deployment-installability.md Extension Point 7.

Verified locally: `helm lint` passes; `helm template --set agent.enabled=true`
renders a DaemonSet with image rcourtman/pulse:6.0.0,
command ["/usr/local/bin/pulse-agent"], args ["--enable-docker", "--enable-host=false"].
End-to-end image build + agent DaemonSet smoke will run via helm_smoke
on the next release once rcourtman/pulse:6.0.0 is published.
2026-05-12 16:11:56 +01:00
rcourtman 0b98cded45 Bulk count agent fleet approvals 2026-05-12 16:00:31 +01:00