Commit Graph

9188 Commits

Author SHA1 Message Date
rcourtman e31fc37983 Gate Patrol actions on agent preflight 2026-08-14 01:12:49 +01:00
rcourtman ff15c91274 Run validated local Patrol observers 2026-08-14 00:45:22 +01:00
rcourtman 71a0fa37a1 Add model-authored Patrol observer proposals 2026-08-14 00:24:52 +01:00
rcourtman ff2bfbb320 Make platform telemetry presentation evidence-aware 2026-08-14 00:18:23 +01:00
rcourtman 903b579f8a Add durable Patrol objectives 2026-08-13 23:59:33 +01:00
rcourtman 64d8b04d36 Keep card-layout backup and snapshot columns out of the numeric editor
The guests table falls back to the card layout below its 14-column width
minimum, and there the Backup and Snapshot columns rendered the generic
numeric metric editor. Saving wrote {trigger, clear} into the backup and
snapshot blocks, which the backend reads as an all-zero disabled config,
so a typed warning-day value silently disabled that guest's backup
alerts. The card layout now renders the same on/off toggle badges as the
desktop rows, and the save paths drop backup/snapshot keys from staged
threshold maps so no editor surface can persist that shape again.

Refs #1126

Contract-Neutral: behavioral bug fix refs #1126: card-layout backup/snapshot columns rendered numeric editor and corrupted override blocks; no public contract delta
2026-08-13 23:56:49 +01:00
rcourtman a7646e5f86 Fix Patrol autonomy and refusal reporting 2026-08-13 23:26:08 +01:00
rcourtman d27abf6fda Fix vSphere backup status presentation 2026-08-13 23:22:30 +01:00
rcourtman 14a3b8b5af Fix thermal history for agent machines 2026-08-13 22:58:29 +01:00
rcourtman 8165eb1066 Format patrol attention workbench test
Prettier's format check failed on main after bcabda085, run 31743135892.
Formatting only, the suite passes unchanged.
2026-08-13 22:12:22 +01:00
rcourtman 3981ce552b Honor explicit cluster member address overrides and surface recovery failures
An explicit connection address override on a cluster member was silently
discarded when VerifySSL was enabled and the member had no per-endpoint
fingerprint: the hostname-for-TLS preference displaced the operator's
address, so overriding an undialable discovered hostname changed nothing.
The override now wins in every TLS mode, and the hostname preference
applies only to auto-discovered addresses.

Failed endpoint recovery attempts also logged their cause at debug level
only, leaving the recurring 'No endpoints recovered' warning without a
reason. The warning now carries per-endpoint failure reasons and the
sanitized error is stored so the UI health status shows it too.

Refs #1665

Contract-Neutral: behavioral bug fix refs #1665: cluster member IPOverride honored in all TLS modes, recovery failure reasons surfaced; no public contract delta
2026-08-13 22:06:29 +01:00
Martin cd02d9f3c0 Honor the configured request timeout for discovery AI analysis (#1713)
* Honor the configured request timeout for discovery AI analysis

Both discovery services accepted an AIAnalysisTimeout config field, but
neither production construction site ever set it, so every AI analysis
call ran under the hardcoded 45s fallback. RequestTimeoutSeconds reached
the provider HTTP client but never the per-analysis context deadline, so
raising it could not help a slow local model: the outer 45s deadline
always fired first and surfaced "AI analysis timed out after 45s".

Add GetDiscoveryAIAnalysisTimeout, which returns the configured request
timeout when one is set and keeps the tighter 45s discovery default
otherwise, and wire it into both service constructors. Both services also
gain SetAIAnalysisTimeout so a settings change applies live instead of
requiring a restart; the timeout is now read through a mutex-guarded
accessor since it can be mutated while scans are running.

* Inherit the configured request timeout for discovery analysis

Returning 45 seconds when RequestTimeoutSeconds is zero left the
discovery deadline disagreeing with what operators are shown. Zero is
omitted from the settings response, the frontend displays 300 via
request_timeout_seconds ?? 300 and states that value applies to
discovery, and it only submits the field when the form value differs
from the displayed current value. An installation showing 300 seconds
therefore still timed out discovery after 45, and saving the visible
default did not correct it.

GetDiscoveryAIAnalysisTimeout now delegates to GetRequestTimeout so the
persisted default canonicalizes to 300 seconds. The discovery services
keep their own 45-second fallback for construction without an AI
configuration.
2026-08-13 21:56:53 +01:00
rcourtman bcabda085f Point patrol attention lifecycle at the durable threshold and finding controls
Two users independently concluded the permanent 'Remember as expected'
dismissals were removed (discussions #1623, #1699) because the Needs
attention workbench only offers acknowledge and temporary suppression,
while the finding-level Manage menu sits behind a collapsed disclosure
below it.

The attention detail's Lifecycle section now states that its controls
cover only the current occurrence and links the two durable paths: the
alert thresholds page, and the Patrol findings panel via a control that
expands the disclosure and scrolls to it. The disclosure summary also
carries an open-work count badge so it no longer reads as an empty
archive.
2026-08-13 21:53:56 +01:00
rcourtman 929c03b490 Prepare v6.2.2-rc.2 release v6.2.2-rc.2 helm-chart-6.2.2-rc.2 2026-08-13 10:54:36 +01:00
rcourtman a39935182d Parse smartctl 7.5 power_mode object on guarded probes
smartmontools 7.5 emits power_mode as an {ata_value, name} object whenever
the -n guard runs CHECK POWER MODE, which is every rotational-disk probe.
The parser declared the field as a string, so json.Unmarshal failed for the
whole document and healthy spinning disks degraded to the lossy text
fallback, surfacing as no usable SMART data while guard-free SSD probes
kept working. That is the exact rotational-only failure split in the
discussion 1690 debug log (SAS3224 HBA, smartctl 7.5). Decode both shapes
the field has used, key standby on the reported name, and never fail the
document over this field. Also broaden the text fallback standby match to
the EPC names (STANDBY_Y, STANDBY (OS), SLEEP) that the mode-suffixed
match missed.

Refs #1690

Contract-Neutral: smartctl 7.5 power_mode JSON parse fix, internal decoder only, no DiskSMART payload or subsystem contract delta (discussion #1690)
2026-08-13 10:25:28 +01:00
rcourtman 3355f7a671 Read host Docker credentials for private registry update checks
Container update detection only ever negotiated anonymous pull tokens, so
containers from registries that reject anonymous digest HEADs pinned a
permanent "authentication required" badge (#1706). The agent already runs
on the Docker host, so the checker now resolves the same credential store
docker pull uses - config.json auths entries, credsStore/credHelpers
credential helpers (docker-credential-<name> get), and Podman's auth.json -
and presents the stored login: Basic auth on Bearer token negotiation and
on the hardcoded Docker Hub / ghcr.io token endpoints, direct answers to
Basic challenges, and the refresh-token grant for identity-token logins
such as Azure ACR.

Credentials never leave the host: they are only presented to the registry
or its token endpoint, helper output stays out of reported check errors,
and lookups are cached in memory for five minutes. Helper names are
validated before exec, and a stale login falls back to the anonymous path
so checks that used to work keep working. Set
PULSE_DISABLE_REGISTRY_CREDENTIALS=true (--disable-registry-credentials)
to keep detection anonymous-only. The agent-lifecycle and security-privacy
subsystem contracts pin the host-local credential boundary.
2026-08-13 10:06:26 +01:00
rcourtman faf13973c4 Keep release race builds off WSL tmpfs v6.2.2-rc.1 helm-chart-6.2.2-rc.1 2026-08-12 17:56:11 +01:00
rcourtman 647f3e6a6c Keep shipped upgrade guide in release sync 2026-08-12 17:43:09 +01:00
rcourtman 97c231462f Fix Linux installer-test stub recursion 2026-08-12 17:30:12 +01:00
rcourtman 2c48696ad7 Prepare v6.2.2-rc.1 release 2026-08-12 17:22:10 +01:00
rcourtman 31b626d73c Fix release worker Go toolchain parity check 2026-08-12 17:11:04 +01:00
rcourtman 1bb9ba4208 Accelerate release qualification with exact-SHA worker 2026-08-12 17:07:12 +01:00
rcourtman a44dc8cfc8 Clarify that TERMS.md covers only the commercial Pro distribution
The README License section read as if parts of this repository were
dual-licensed. State explicitly that everything in the repository is MIT
and that the Pulse Pro Terms of Service govern only the separately built
and distributed commercial product. Add the same scope statement to the
top of TERMS.md and its in-app copy.
2026-08-12 16:01:46 +01:00
rcourtman d6cf090737 Apply agent LXC filesystem data on the cluster/resources poll path
The host agent's node-local pct df inventory was only applied by the
per-node container poll fallback. Cluster instances are served by the
efficient cluster/resources path, which never called the enrichment, so
installs on that path showed rootfs-only container filesystems (and,
since the config-mount restoration, config-listed mounts with unknown
usage) no matter how healthy the linked agent was. Reported on #1477
after the reporter installed agents specifically to get per-mount usage.

Reproduced end to end against a live PVE cluster: the agent shipped the
full inventory, the server linked the agent and populated the
filesystem cache every report, and the poll path never read it. With
the enrichment applied after metadata enrichment, mirroring the
per-node path's ordering, the same rig surfaces real per-mount usage.

Refs #1477

Contract-Neutral: behavioral fix: cluster/resources poll path now applies agent pct df enrichment like the per-node path (#1477), no public contract delta
2026-08-12 14:45:38 +01:00
rcourtman b8d61b9e12 Sync shipped PROXY_AUTH.md with the repo doc
The proxy-auth documentation commits this morning updated
docs/PROXY_AUTH.md without refreshing the copy shipped in
frontend-modern/public/docs, so the docsLinks shipped-content sync test
went red on the next push that ran the frontend suite.
2026-08-12 12:16:24 +01:00
rcourtman 5632ee10e8 Keep the pending planned poll slot instead of compounding it
Every poll tick rebuilds the adaptive plan, and BuildPlan anchored the
next run on the previously planned NextRun by unconditionally adding the
selected interval. While a planned slot has not elapsed, each planning
pass therefore pushed it another full interval into the future. With the
default 10s tick that compounds NextRun ahead of wall-clock as soon as
the adaptive interval stretches beyond the tick cadence, so an instance
whose data was fresh at startup was never due again: API polling starved
permanently, the connection dropped to Agent and stale exactly when the
stretch began, and no error was ever logged. Matches the fourth #1437
reproduction (standalone degrades at ~2 minutes, cluster at ~4, stays
degraded), whose bundle shows a clean log with no polls after startup.

A pending future slot is now kept as planned, and it tightens to
now+interval when a staleness-driven interval shrink justifies an
earlier run. An elapsed slot still advances exactly one interval, so
ordinary cadence is unchanged.

Refs #1437

Contract-Neutral: behavioral fix: pending planned poll slot no longer compounds NextRun past wall-clock (#1437), no public contract delta
2026-08-12 11:22:29 +01:00
rcourtman 4eb156661e Document the 5.x proxy-auth mitigation and its end-of-life status
The 5.x line is end-of-life and will not be patched for the proxy-auth admin
gating defect, so the docs have to carry its mitigation: a 5.x operator has
nowhere else to look, and the security advisory is a moment in time while this
page is where they actually land.

5.x needs two changes, and either alone leaves admin open. It never applies the
documented "admin" default, so PROXY_AUTH_ADMIN_ROLE must be set explicitly.
It also only evaluates roles when the header carries a value — absent or empty
leaves isAdmin at its true default — so the proxy has to send a placeholder for
users with no groups, which is exactly the case an IdP tends to produce for the
least privileged account. 6.x fixes both: gating activates on the role header
alone, and a blank header resolves to non-admin.

Includes the negative test both ways round, since "it returns 200" is the only
symptom an operator can actually observe.
2026-08-12 10:21:46 +01:00
rcourtman abf5e45b70 Clarify proxy-auth upgrade requirements 2026-08-12 09:51:29 +01:00
rcourtman fbc584554c Add v6.2.2 proxy-auth release note 2026-08-12 09:48:19 +01:00
rcourtman 46f11fffc4 Document the proxy auth header trust boundary
The proxy auth docs never told operators that the identity and role headers
must be replaced rather than appended, nor that Pulse has to be unreachable
except through the proxy. Both are prerequisites for the scheme being safe at
all, and neither is enforceable from inside Pulse.

The append case is the sharp edge: Pulse reads the first value of a repeated
header, so a client-supplied role header that arrives ahead of the proxy's
value decides the admin verdict. The client does not need the shared secret to
do it, because the proxy attaches the secret itself. Verified against a scratch
instance: sending "X-Proxy-Roles: user" then "admin" yields 403 while "admin"
then "user" yields 200 on /api/system/settings.

Documented rather than fixed in code on purpose. Matching any value instead of
the first would make injection strictly easier, and rejecting repeated headers
outright would break identity providers that legitimately emit one header per
group. The trust boundary is the proxy's to hold.
2026-08-12 09:47:55 +01:00
rcourtman 0c3f475dea Warn when proxy role gating falls back to the default admin role
Configuring PROXY_AUTH_ROLE_HEADER without PROXY_AUTH_ADMIN_ROLE now gates
admin access on an exact-case match against the "admin" default. That closes
the fail-open, but it also means an IdP sending "Admins", "Admin", or
"authentik Admins" grants nobody admin — and a proxy-auth-only deployment has
no local credential to fall back on. The change is recoverable by setting
PROXY_AUTH_ADMIN_ROLE and restarting, but only if the operator can tell that
is what happened.

Log a startup warning naming the header, the effective admin role, the
case-sensitivity, and both ways out. The warning is scoped to deployments
that actually configured a role header; defaulting the value with no role
header is inert and stays at info. The troubleshooting entry in
docs/PROXY_AUTH.md now leads with the exact-match rule, since that is where
a locked-out operator looks first.
2026-08-12 09:43:52 +01:00
rcourtman 11a8aa3b2f Fail closed when proxy auth configures a role header but no admin role
A reverse-proxy deployment that set PROXY_AUTH_ROLE_HEADER without also
setting PROXY_AUTH_ADMIN_ROLE granted every proxy-authenticated user full
administrator access. CheckProxyAuth only evaluated roles when both values
were non-empty, so the half-configuration skipped role gating entirely and
returned isAdmin=true. docs/PROXY_AUTH.md has always documented an `admin`
default for that variable, but the Config struct's envconfig `default` tags
are legacy and never applied (config.go), so nothing ever populated it.

CheckProxyAuth is the single admin verdict all 20+ proxy-auth gates consume,
so the fail-open reached every one of them. Verified on a scratch instance
with PROXY_AUTH_ROLE_HEADER set and no admin role: a request carrying only
`X-Proxy-Roles: user` received HTTP 200 and the full admin payload from
GET /api/system/settings, HTTP 200 from POST /api/system/settings/update,
and proxyAuthIsAdmin=true from /api/security/status. All three now return
403 / false, while `X-Proxy-Roles: admin` still passes.

Resolve the documented default in both layers that can produce the verdict:
config load populates ProxyAuthAdminRole when proxy auth is configured, and
CheckProxyAuth now keys role gating on the role header alone, resolving an
empty admin role through config.DefaultProxyAuthAdminRole. Configuring a
role header is the operator's signal that admin access is role-gated;
leaving the admin role unset must not switch that off.

Deployments that intentionally treat every proxied user as an admin are
unaffected: that is still expressed by leaving the role header unset.
2026-08-12 09:38:02 +01:00
rcourtman 34194e57be Show the real monitoring cadence to non-admin sessions
Non-admin sessions cannot read GET /api/system/settings, so the Settings
General Monitoring Cadence card fell back to the Realtime (10s) preset
regardless of the configured interval; an issue #1601 reporter read that
as the server polling faster for non-admins. Publish the effective
pvePollingInterval on the authenticated runtime-display projection
(runtime config first, persisted value only as fallback, matching the
admin route's precedence), consume it in the viewer fallback of the
settings state, and run that initialization for sessions without
infrastructureRead too, whose ungated General panel previously never
initialized presentation state at all.
2026-08-12 09:20:01 +01:00
rcourtman 3adeb77d60 Secure configuration transfer authorization (#1714)
Co-authored-by: Pulse Autonomous Maintainer <rcourtman@users.noreply.github.com>
2026-08-12 07:32:50 +01:00
rcourtman 359a372cdf Make per-guest backup/snapshot toggles inherit global thresholds
The per-guest Backup/Snapshot toggle persisted a full copy of the current
global defaults just to flip enabled, freezing the threshold values into
the override. Later global edits (a 32-day backup warning) then silently
never applied to toggled guests, which kept firing at the frozen factory
7-day warning. Reported twice in discussion #1126.

- Guest overrides now resolve against the globals at evaluation time:
  zero-valued fields inherit the global value, explicit values still win,
  and hand-written sparse overrides stop decoding as accidental zeros.
- The toggles write enabled-only overrides instead of freezing a copy.
- Normalization rewrites stored overrides whose thresholds exactly match
  the current globals into sparse form, which is behavior-preserving at
  migration time and un-freezes existing installs.
- The Backups/Snapshots global editors reconcile warning/critical pairs
  by adjusting the untouched field, so typing a 32-day warning no longer
  silently snaps back to the 14-day critical default.

Refs #1126
2026-08-12 06:38:25 +01:00
rcourtman 1243cde497 Merge pull request #1712 from rcourtman/agent/2d25585f4d694195-audit-hmac
Harden audit signatures against boundary forgery
2026-08-12 05:19:43 +01:00
Pulse Autonomous Maintainer 1dcb414167 Harden audit signatures against boundary forgery 2026-08-12 05:05:11 +01:00
rcourtman a1b5f085b3 docs: update remote builder hostname 2026-08-12 00:38:21 +01:00
rcourtman 887b85bc83 Validate Core E2E tiering against run evidence (#1711)
Contract-Neutral: E2E test tier metadata and validation only; no deployment runtime or public contract change.

Co-authored-by: Pulse Autonomous Maintainer <rcourtman@users.noreply.github.com>
2026-08-11 23:18:14 +01:00
rcourtman d9df2b9422 Merge pull request #1709 from rcourtman/agent/6f81a6a0604241fb-autonomous-maintenance-sweep-for-pulse
Restore frontend formatting gate
2026-08-11 21:39:24 +01:00
Pulse Autonomous Maintainer 0512d07626 Format stacked disk branch coverage test 2026-08-11 21:24:50 +01:00
rcourtman d40c02742b Merge pull request #1708 from rcourtman/agent/6f81a6a0604241fb-autonomous-maintenance-sweep-for-pulse
Refresh auth recovery visual baselines
2026-08-11 21:22:38 +01:00
rcourtman e851253f3e Pin the config-only LXC mount wire shape end to end
Integration test for the exact production path: stock PVE answers the
LXC status query with an empty diskinfo map, enrichContainerMetadata
discovers the mpX mount from the container config alone, and the unified
resource projection serializes it with capacity, the -1 unknown-usage
sentinel, and an omitted used field while the live rootfs row survives.

Related to #1477
2026-08-11 21:03:27 +01:00
Pulse Autonomous Maintainer 3f1bf8c329 Refresh auth visual baselines 2026-08-11 21:01:36 +01:00
rcourtman 63f1a14f31 Show LXC mount points from container config on the API path again
Stock Proxmox reports no per-mount LXC usage through the status API, so
v6's API-polled containers listed only rootfs. The v5.1.32 fallback that
synthesized mount rows from the container config never crossed to the v6
line, and the v6.2.0 pct-df agent path only covers nodes running the
unified agent. Restore the fallback and improve it: parse size= so
config-only rows carry capacity, mark live usage unknown with the -1
sentinel, and merge without displacing the aggregate-seeded rootfs row.

Frontend consumers stop fabricating percents for sentinel rows: the
workloads row bar and summary math exclude them (tooltip lists them with
capacity), the drawer Filesystems block renders ?/<size> with no percent,
disk normalization preserves the sentinel, and per-machine max-disk
derivations skip them. Mock mode seeds one running container in this
exact shape so the surfaces stay exercised.

Related to #1477
2026-08-11 20:46:11 +01:00
rcourtman cea33b8ee2 Merge pull request #1707 from rcourtman/agent/47596ee9a4f54c87-autonomous-maintenance-sweep-for-pulse
Reflect global settings for non-admin viewers
2026-08-11 20:22:08 +01:00
Pulse Autonomous Maintainer 986a281006 Reflect global settings for non-admin viewers 2026-08-11 19:57:47 +01:00
courtmanr@gmail.com 5015755245 Refresh API access bundle budget
Contract-Neutral: Records the intentional shipped UI size without changing runtime behavior.
2026-08-11 17:13:22 +01:00
rcourtman 56262c6368 Fix provider MSP evaluation setup flow 2026-08-11 16:51:15 +01:00
courtmanr@gmail.com c76f07e5ec docs(agents): add customer support communication guidance 2026-08-11 16:39:36 +01:00