Commit Graph

9167 Commits

Author SHA1 Message Date
rcourtman 1bb9ba4208 Accelerate release qualification with exact-SHA worker 2026-08-12 17:07:12 +01:00
rcourtman a44dc8cfc8 Clarify that TERMS.md covers only the commercial Pro distribution
The README License section read as if parts of this repository were
dual-licensed. State explicitly that everything in the repository is MIT
and that the Pulse Pro Terms of Service govern only the separately built
and distributed commercial product. Add the same scope statement to the
top of TERMS.md and its in-app copy.
2026-08-12 16:01:46 +01:00
rcourtman d6cf090737 Apply agent LXC filesystem data on the cluster/resources poll path
The host agent's node-local pct df inventory was only applied by the
per-node container poll fallback. Cluster instances are served by the
efficient cluster/resources path, which never called the enrichment, so
installs on that path showed rootfs-only container filesystems (and,
since the config-mount restoration, config-listed mounts with unknown
usage) no matter how healthy the linked agent was. Reported on #1477
after the reporter installed agents specifically to get per-mount usage.

Reproduced end to end against a live PVE cluster: the agent shipped the
full inventory, the server linked the agent and populated the
filesystem cache every report, and the poll path never read it. With
the enrichment applied after metadata enrichment, mirroring the
per-node path's ordering, the same rig surfaces real per-mount usage.

Refs #1477

Contract-Neutral: behavioral fix: cluster/resources poll path now applies agent pct df enrichment like the per-node path (#1477), no public contract delta
2026-08-12 14:45:38 +01:00
rcourtman b8d61b9e12 Sync shipped PROXY_AUTH.md with the repo doc
The proxy-auth documentation commits this morning updated
docs/PROXY_AUTH.md without refreshing the copy shipped in
frontend-modern/public/docs, so the docsLinks shipped-content sync test
went red on the next push that ran the frontend suite.
2026-08-12 12:16:24 +01:00
rcourtman 5632ee10e8 Keep the pending planned poll slot instead of compounding it
Every poll tick rebuilds the adaptive plan, and BuildPlan anchored the
next run on the previously planned NextRun by unconditionally adding the
selected interval. While a planned slot has not elapsed, each planning
pass therefore pushed it another full interval into the future. With the
default 10s tick that compounds NextRun ahead of wall-clock as soon as
the adaptive interval stretches beyond the tick cadence, so an instance
whose data was fresh at startup was never due again: API polling starved
permanently, the connection dropped to Agent and stale exactly when the
stretch began, and no error was ever logged. Matches the fourth #1437
reproduction (standalone degrades at ~2 minutes, cluster at ~4, stays
degraded), whose bundle shows a clean log with no polls after startup.

A pending future slot is now kept as planned, and it tightens to
now+interval when a staleness-driven interval shrink justifies an
earlier run. An elapsed slot still advances exactly one interval, so
ordinary cadence is unchanged.

Refs #1437

Contract-Neutral: behavioral fix: pending planned poll slot no longer compounds NextRun past wall-clock (#1437), no public contract delta
2026-08-12 11:22:29 +01:00
rcourtman 4eb156661e Document the 5.x proxy-auth mitigation and its end-of-life status
The 5.x line is end-of-life and will not be patched for the proxy-auth admin
gating defect, so the docs have to carry its mitigation: a 5.x operator has
nowhere else to look, and the security advisory is a moment in time while this
page is where they actually land.

5.x needs two changes, and either alone leaves admin open. It never applies the
documented "admin" default, so PROXY_AUTH_ADMIN_ROLE must be set explicitly.
It also only evaluates roles when the header carries a value — absent or empty
leaves isAdmin at its true default — so the proxy has to send a placeholder for
users with no groups, which is exactly the case an IdP tends to produce for the
least privileged account. 6.x fixes both: gating activates on the role header
alone, and a blank header resolves to non-admin.

Includes the negative test both ways round, since "it returns 200" is the only
symptom an operator can actually observe.
2026-08-12 10:21:46 +01:00
rcourtman abf5e45b70 Clarify proxy-auth upgrade requirements 2026-08-12 09:51:29 +01:00
rcourtman fbc584554c Add v6.2.2 proxy-auth release note 2026-08-12 09:48:19 +01:00
rcourtman 46f11fffc4 Document the proxy auth header trust boundary
The proxy auth docs never told operators that the identity and role headers
must be replaced rather than appended, nor that Pulse has to be unreachable
except through the proxy. Both are prerequisites for the scheme being safe at
all, and neither is enforceable from inside Pulse.

The append case is the sharp edge: Pulse reads the first value of a repeated
header, so a client-supplied role header that arrives ahead of the proxy's
value decides the admin verdict. The client does not need the shared secret to
do it, because the proxy attaches the secret itself. Verified against a scratch
instance: sending "X-Proxy-Roles: user" then "admin" yields 403 while "admin"
then "user" yields 200 on /api/system/settings.

Documented rather than fixed in code on purpose. Matching any value instead of
the first would make injection strictly easier, and rejecting repeated headers
outright would break identity providers that legitimately emit one header per
group. The trust boundary is the proxy's to hold.
2026-08-12 09:47:55 +01:00
rcourtman 0c3f475dea Warn when proxy role gating falls back to the default admin role
Configuring PROXY_AUTH_ROLE_HEADER without PROXY_AUTH_ADMIN_ROLE now gates
admin access on an exact-case match against the "admin" default. That closes
the fail-open, but it also means an IdP sending "Admins", "Admin", or
"authentik Admins" grants nobody admin — and a proxy-auth-only deployment has
no local credential to fall back on. The change is recoverable by setting
PROXY_AUTH_ADMIN_ROLE and restarting, but only if the operator can tell that
is what happened.

Log a startup warning naming the header, the effective admin role, the
case-sensitivity, and both ways out. The warning is scoped to deployments
that actually configured a role header; defaulting the value with no role
header is inert and stays at info. The troubleshooting entry in
docs/PROXY_AUTH.md now leads with the exact-match rule, since that is where
a locked-out operator looks first.
2026-08-12 09:43:52 +01:00
rcourtman 11a8aa3b2f Fail closed when proxy auth configures a role header but no admin role
A reverse-proxy deployment that set PROXY_AUTH_ROLE_HEADER without also
setting PROXY_AUTH_ADMIN_ROLE granted every proxy-authenticated user full
administrator access. CheckProxyAuth only evaluated roles when both values
were non-empty, so the half-configuration skipped role gating entirely and
returned isAdmin=true. docs/PROXY_AUTH.md has always documented an `admin`
default for that variable, but the Config struct's envconfig `default` tags
are legacy and never applied (config.go), so nothing ever populated it.

CheckProxyAuth is the single admin verdict all 20+ proxy-auth gates consume,
so the fail-open reached every one of them. Verified on a scratch instance
with PROXY_AUTH_ROLE_HEADER set and no admin role: a request carrying only
`X-Proxy-Roles: user` received HTTP 200 and the full admin payload from
GET /api/system/settings, HTTP 200 from POST /api/system/settings/update,
and proxyAuthIsAdmin=true from /api/security/status. All three now return
403 / false, while `X-Proxy-Roles: admin` still passes.

Resolve the documented default in both layers that can produce the verdict:
config load populates ProxyAuthAdminRole when proxy auth is configured, and
CheckProxyAuth now keys role gating on the role header alone, resolving an
empty admin role through config.DefaultProxyAuthAdminRole. Configuring a
role header is the operator's signal that admin access is role-gated;
leaving the admin role unset must not switch that off.

Deployments that intentionally treat every proxied user as an admin are
unaffected: that is still expressed by leaving the role header unset.
2026-08-12 09:38:02 +01:00
rcourtman 34194e57be Show the real monitoring cadence to non-admin sessions
Non-admin sessions cannot read GET /api/system/settings, so the Settings
General Monitoring Cadence card fell back to the Realtime (10s) preset
regardless of the configured interval; an issue #1601 reporter read that
as the server polling faster for non-admins. Publish the effective
pvePollingInterval on the authenticated runtime-display projection
(runtime config first, persisted value only as fallback, matching the
admin route's precedence), consume it in the viewer fallback of the
settings state, and run that initialization for sessions without
infrastructureRead too, whose ungated General panel previously never
initialized presentation state at all.
2026-08-12 09:20:01 +01:00
rcourtman 3adeb77d60 Secure configuration transfer authorization (#1714)
Co-authored-by: Pulse Autonomous Maintainer <rcourtman@users.noreply.github.com>
2026-08-12 07:32:50 +01:00
rcourtman 359a372cdf Make per-guest backup/snapshot toggles inherit global thresholds
The per-guest Backup/Snapshot toggle persisted a full copy of the current
global defaults just to flip enabled, freezing the threshold values into
the override. Later global edits (a 32-day backup warning) then silently
never applied to toggled guests, which kept firing at the frozen factory
7-day warning. Reported twice in discussion #1126.

- Guest overrides now resolve against the globals at evaluation time:
  zero-valued fields inherit the global value, explicit values still win,
  and hand-written sparse overrides stop decoding as accidental zeros.
- The toggles write enabled-only overrides instead of freezing a copy.
- Normalization rewrites stored overrides whose thresholds exactly match
  the current globals into sparse form, which is behavior-preserving at
  migration time and un-freezes existing installs.
- The Backups/Snapshots global editors reconcile warning/critical pairs
  by adjusting the untouched field, so typing a 32-day warning no longer
  silently snaps back to the 14-day critical default.

Refs #1126
2026-08-12 06:38:25 +01:00
rcourtman 1243cde497 Merge pull request #1712 from rcourtman/agent/2d25585f4d694195-audit-hmac
Harden audit signatures against boundary forgery
2026-08-12 05:19:43 +01:00
Pulse Autonomous Maintainer 1dcb414167 Harden audit signatures against boundary forgery 2026-08-12 05:05:11 +01:00
rcourtman a1b5f085b3 docs: update remote builder hostname 2026-08-12 00:38:21 +01:00
rcourtman 887b85bc83 Validate Core E2E tiering against run evidence (#1711)
Contract-Neutral: E2E test tier metadata and validation only; no deployment runtime or public contract change.

Co-authored-by: Pulse Autonomous Maintainer <rcourtman@users.noreply.github.com>
2026-08-11 23:18:14 +01:00
rcourtman d9df2b9422 Merge pull request #1709 from rcourtman/agent/6f81a6a0604241fb-autonomous-maintenance-sweep-for-pulse
Restore frontend formatting gate
2026-08-11 21:39:24 +01:00
Pulse Autonomous Maintainer 0512d07626 Format stacked disk branch coverage test 2026-08-11 21:24:50 +01:00
rcourtman d40c02742b Merge pull request #1708 from rcourtman/agent/6f81a6a0604241fb-autonomous-maintenance-sweep-for-pulse
Refresh auth recovery visual baselines
2026-08-11 21:22:38 +01:00
rcourtman e851253f3e Pin the config-only LXC mount wire shape end to end
Integration test for the exact production path: stock PVE answers the
LXC status query with an empty diskinfo map, enrichContainerMetadata
discovers the mpX mount from the container config alone, and the unified
resource projection serializes it with capacity, the -1 unknown-usage
sentinel, and an omitted used field while the live rootfs row survives.

Related to #1477
2026-08-11 21:03:27 +01:00
Pulse Autonomous Maintainer 3f1bf8c329 Refresh auth visual baselines 2026-08-11 21:01:36 +01:00
rcourtman 63f1a14f31 Show LXC mount points from container config on the API path again
Stock Proxmox reports no per-mount LXC usage through the status API, so
v6's API-polled containers listed only rootfs. The v5.1.32 fallback that
synthesized mount rows from the container config never crossed to the v6
line, and the v6.2.0 pct-df agent path only covers nodes running the
unified agent. Restore the fallback and improve it: parse size= so
config-only rows carry capacity, mark live usage unknown with the -1
sentinel, and merge without displacing the aggregate-seeded rootfs row.

Frontend consumers stop fabricating percents for sentinel rows: the
workloads row bar and summary math exclude them (tooltip lists them with
capacity), the drawer Filesystems block renders ?/<size> with no percent,
disk normalization preserves the sentinel, and per-machine max-disk
derivations skip them. Mock mode seeds one running container in this
exact shape so the surfaces stay exercised.

Related to #1477
2026-08-11 20:46:11 +01:00
rcourtman cea33b8ee2 Merge pull request #1707 from rcourtman/agent/47596ee9a4f54c87-autonomous-maintenance-sweep-for-pulse
Reflect global settings for non-admin viewers
2026-08-11 20:22:08 +01:00
Pulse Autonomous Maintainer 986a281006 Reflect global settings for non-admin viewers 2026-08-11 19:57:47 +01:00
courtmanr@gmail.com 5015755245 Refresh API access bundle budget
Contract-Neutral: Records the intentional shipped UI size without changing runtime behavior.
2026-08-11 17:13:22 +01:00
rcourtman 56262c6368 Fix provider MSP evaluation setup flow 2026-08-11 16:51:15 +01:00
courtmanr@gmail.com c76f07e5ec docs(agents): add customer support communication guidance 2026-08-11 16:39:36 +01:00
courtmanr@gmail.com 1549216b6e Handle registry responses without digest headers 2026-08-11 16:39:36 +01:00
courtmanr@gmail.com 1b0cf176bd Explain retained alert delivery notices 2026-08-11 16:39:36 +01:00
courtmanr@gmail.com 791a2f86bf Route Pulse images through product update checks 2026-08-11 16:39:12 +01:00
courtmanr@gmail.com a49e016246 Hide duplicate PBS profile identities 2026-08-11 16:38:46 +01:00
courtmanr@gmail.com 66d8e90c0c Show Patrol findings in alert emails 2026-08-11 16:38:21 +01:00
courtmanr@gmail.com 16f9dbc709 Add push preferences to relay connect frames 2026-08-11 16:38:21 +01:00
courtmanr@gmail.com e21a3c4434 Classify normal WebSocket shutdowns as informational 2026-08-11 16:38:21 +01:00
courtmanr@gmail.com b3fdab4cae Persist first-run auth for systemd installs 2026-08-11 16:38:21 +01:00
courtmanr@gmail.com 1d0a0b7f7e Allow existing API tokens to be renamed 2026-08-11 16:38:21 +01:00
courtmanr@gmail.com 83dc17bf71 Avoid PVE linkage warnings for PBS agent profiles 2026-08-11 16:37:49 +01:00
courtmanr@gmail.com 0ac02e43c5 Allow insecure TLS for Proxmox agent installs 2026-08-11 16:37:49 +01:00
courtmanr@gmail.com 70469660b3 Clear Docker alerts when removing unified hosts 2026-08-11 16:37:37 +01:00
courtmanr@gmail.com 4dac4dd163 Allow agents to include filtered disk mounts 2026-08-11 16:37:37 +01:00
courtmanr@gmail.com df606ee81a Describe External Probes server-side alerting accurately 2026-08-11 16:37:37 +01:00
rcourtman 2897cee8c5 Merge pull request #1705 from rcourtman/agent/2cb3679640b14b39-autonomous-maintenance-sweep-for-pulse
Protect shared metrics database directories
2026-08-11 16:36:28 +01:00
rcourtman eddd8a1988 fix(release): show categorized changelog after updates 2026-08-11 16:02:40 +01:00
Pulse Autonomous Maintainer f445b7fa29 Protect shared metrics database directories 2026-08-11 15:57:17 +01:00
rcourtman 3826316eec Split Windows signing submission from approval-bound collection
Production SignPath signing requests require manual approval in the
SignPath UI, so the previous single-job flow (submit with
wait-for-completion inside a 40-minute window) let approval latency fail
the Windows build, and any re-run rebuilt the binaries and submitted a
second request needing a second approval.

The Windows lane is now two jobs: sign-windows-agent builds the unsigned
executables, submits the SignPath request without waiting, and uploads a
7-day signing-request record; collect-windows-signing absorbs approval
latency by polling the recorded request, downloads the signed artifact
by request id, and keeps the existing verification and evidence steps.
If approval outlasts the 115-minute polling window, the collection job
fails with re-run guidance and "Re-run failed jobs" collects the same
recorded request - no rebuild, no resubmission. The legacy PFX
break-glass backend rides the same two-job shape via an artifact
hand-off. Workflow output wiring, artifact names, and evidence content
are unchanged for downstream consumers.

The shape test now pins the async invariants (no wait-for-completion:
true in the candidate workflow), and the code signing policy plus the
deployment-installability contract describe the two-phase flow.
2026-08-11 15:50:25 +01:00
rcourtman d2e25dcb63 Reword evaluation availability clause clear of the telemetry guardrail
TestRepositoryDoesNotClaimTelemetryIsAnonymous flags any 'anonymous'
within 120 characters of 'telemetry' on one line. The MSP evaluation
clause landed in c17664b3d legitimately ends one sentence with
'telemetry' and starts the next with 'Anonymous evaluation', tripping
the scan without claiming telemetry is anonymous. State the same
requirement without the word collision.

Contract-Neutral: wording-only: same clause meaning, avoids telemetry-anonymous guardrail false positive
2026-08-11 15:18:12 +01:00
rcourtman 32d1fcbe42 Update canonical write-path guardrails for WriteBatchBounded
2bc4ed725 moved the unified metrics sync sites to WriteBatchBounded but
left the three source guardrails pinning the literal WriteBatchSync
call, so they failed. The guardrails' intent is that pipeline writes go
through the canonical batched ingestion path, which WriteBatchBounded
is; pin the new name.

Contract-Neutral: test-only: guardrail pins follow the WriteBatchBounded pipeline rename from 2bc4ed725
2026-08-11 14:58:09 +01:00
rcourtman 98f0711504 Sample mock seed assertions through the seeder's graph sampler
TestSeedMockMetricsHistory_SeedsVMwareMetricsStore and the TrueNAS
variant compared seeded storage values against the package-level
mock.SampleMetric, which resolves resource roles from the global
registry that only other tests populate. The assertions therefore
passed or failed depending on which tests ran earlier in the process:
isolated runs failed deterministically, and today's CI reshard flipped
the rest-0 shard red for commits that never touched the mock layer.

Both tests now build the same graph-aware sampler the seeder uses, so
the expectation is self-contained and order-independent.

Contract-Neutral: test-only: seed assertions sample via the seeder's graph sampler, removes cross-test registry dependence
2026-08-11 14:58:09 +01:00