Commit Graph

9433 Commits

Author SHA1 Message Date
rcourtman fa156e0bb1 Record chart and resource qualification 2026-08-21 18:53:27 +01:00
rcourtman 08827bb887 Extract chart and resource query services 2026-08-21 18:36:21 +01:00
rcourtman dca06991c2 Parallelize exact-version Docker publication 2026-08-21 18:25:26 +01:00
rcourtman 00157de00a Require two-shard backend memory admission 2026-08-21 18:14:33 +01:00
rcourtman e5389e2130 Parallelize inert release artifact staging 2026-08-21 18:08:22 +01:00
rcourtman 0027c82f84 Record rc.5 release convergence evidence 2026-08-21 17:56:23 +01:00
rcourtman 8869753b1e Require tailnet access for every paid runtime proof 2026-08-21 17:52:19 +01:00
rcourtman b16b8e5242 Bind Helm convergence to the release repository 2026-08-21 17:46:55 +01:00
rcourtman 1ef8797d28 Repair release activation recovery contracts 2026-08-21 17:43:01 +01:00
rcourtman 1327dddad5 Compress release backend test selectors v6.3.0-rc.5 helm-chart-6.3.0-rc.5 2026-08-21 17:04:18 +01:00
rcourtman 418402bf9e Use complete PVE release worktree 2026-08-21 16:19:11 +01:00
rcourtman d45ecd9a24 Include release helpers in PVE checkout 2026-08-21 16:14:21 +01:00
rcourtman d6591da900 Support PVE sparse release checkout 2026-08-21 16:11:17 +01:00
rcourtman d9e634a4eb Run release preparation on PVE 2026-08-21 16:07:33 +01:00
rcourtman 5bfcd3005a Isolate chat startup failure test data 2026-08-21 16:03:18 +01:00
rcourtman b261e42901 Record rc.5 acceleration qualification 2026-08-21 15:50:06 +01:00
rcourtman ae27ad7511 Bound release API test batch arguments 2026-08-21 15:47:12 +01:00
rcourtman 4217f72bdc Prepare v6.3.0-rc.5 release 2026-08-21 15:14:23 +01:00
rcourtman 8aca821590 Update API decomposition governance fixtures 2026-08-21 15:13:03 +01:00
rcourtman d505fce29d Record API decomposition qualification 2026-08-21 14:56:11 +01:00
rcourtman e61715462a Canonicalize install security helpers 2026-08-21 14:56:11 +01:00
rcourtman c68d5dd3d8 Extract configuration API runtime package 2026-08-21 14:56:07 +01:00
rcourtman 58bf77c1ae Extract alert delivery API package 2026-08-21 14:56:04 +01:00
rcourtman f76a8279a0 Parallelize release archive validation 2026-08-21 14:52:06 +01:00
rcourtman 74bd06a953 Promote qualified release payloads directly 2026-08-21 14:47:52 +01:00
rcourtman 962df694e9 Accept architecture-bound server signatures 2026-08-21 14:31:42 +01:00
rcourtman c661bdab57 Overlap inert release staging and qualification 2026-08-21 14:25:42 +01:00
rcourtman ae317c96bb Publish images from exact candidate payloads 2026-08-21 14:17:12 +01:00
rcourtman 97cf30aed6 Qualify containers in candidate workflow 2026-08-21 14:03:25 +01:00
rcourtman 51f4c64322 Qualify exact candidate containers on PVE 2026-08-21 13:56:52 +01:00
rcourtman c6bf50b455 Stage private release assets during qualification 2026-08-21 13:30:24 +01:00
rcourtman dcc67662b6 Stage release archives directly in parallel 2026-08-21 13:23:44 +01:00
rcourtman a5893e916c Use persistent caches on PVE release runners 2026-08-21 13:14:33 +01:00
rcourtman 312aa2c7e1 Govern private PVE release compilation 2026-08-21 13:09:40 +01:00
rcourtman d44c15a4fd Compile release payloads in parallel on PVE 2026-08-21 12:59:32 +01:00
rcourtman 0293f67635 Preserve API test order in release shards 2026-08-21 12:46:50 +01:00
rcourtman 9dac68fd62 Parallelize release qualification on PVE runners 2026-08-21 12:24:47 +01:00
rcourtman f4930cdd34 Synchronize the shipped transparency document
Keep the in-app documentation byte-for-byte aligned with the canonical disclosure after the support-email clarification.
v6.3.0-rc.4 helm-chart-6.3.0-rc.4
2026-08-21 09:42:19 +01:00
rcourtman 3c97e94094 Restore the rc.4 release-note contract
The publish-body condensation removed exact operator-facing statements pinned by the prerelease packet test. Restore those statements within the three-highlight limit, record the packet contract, and strengthen the proof to require the complete publish-safe sentences.
2026-08-21 09:36:04 +01:00
rcourtman 547fa77892 Describe autonomous support email in the AI transparency statement 2026-08-21 09:04:23 +01:00
rcourtman ae5df19b62 Give release backend tests hosted-runner headroom
The internal/api race suite now routinely exceeds the old 20-minute package timeout on hosted runners while passing. Set a governed 30-minute package timeout and 40-minute release job ceiling, pin the relationship with contract tests, and refresh the rc.4 packet with the fixes landed since preparation.
2026-08-21 08:49:05 +01:00
rcourtman ad91258d37 Pin the existing agent identity in the Unix credential repair command
The Repair Authentication command generated for Unix agents carried
--update and a fresh token but no identity, while the Windows path pins
PULSE_AGENT_ID and PULSE_HOSTNAME. A repair reinstall without the pin
can register a fresh suffixed agent identity for the same machine
instead of converging on the one being repaired, which is exactly what
a field report documented across repeated repair attempts.

Pass --agent-id and --hostname from the canonical connection the same
way the Windows command does.

Refs discussion #1748

Contract-Neutral: Discussion #1748: unix repair command now pins the existing agent identity, parity with the windows path; command-string generation only, no API or payload change
2026-08-21 06:48:44 +01:00
rcourtman af6d482515 Keep an unresponsive mount from freezing host disk collection
A hard-mounted network filesystem with an unreachable server blocks
statfs in an uninterruptible kernel wait. The collector issued that
syscall inline for every mount, so one dead NFS mount silently froze
the whole reporting cycle and stretched agent shutdown into the kernel
retry window, and it paid that price for mounts the fstype filter was
going to discard anyway.

Decide type- and mountpoint-based skips before the usage syscall, so
network filesystems are never probed unless explicitly included, and
bound every remaining usage call with a timeout that leaves at most one
in-flight probe per mountpoint. A mount whose call never returned is
skipped on later cycles and re-included when the stalled call answers.

Refs discussion #1747

Contract-Neutral: Discussion #1747: behavioral bugfix in host disk collection; no payload or API shape change, filtered mounts were never reported
2026-08-21 06:43:31 +01:00
rcourtman 5c13befcac Run the QNAP agent from the data volume instead of the RAM-backed root
The installer staged the download in /tmp and installed the runtime
binary to /usr/local/bin, both on the small RAM-backed QTS/QuTS hero
root, and the boot wrapper copied 34MiB back onto that root at every
boot. Roots without ~50MiB of headroom could not install at all, and
setting TMPDIR only moved the staging half of the requirement.

QNAP's own QPKG packages execute from the data volume, so do the same:
relocate the install dir to the data volume's state dir before the
preflight and download, default TMPDIR there too, skip the boot-time
self-copy when the stored and runtime binaries are one file, and remove
a pre-relocation runtime copy from /usr/local/bin to give that space
back. Split layouts with an operator-supplied state dir keep the copy
semantics. The rendered wrapper is exercised in both layouts by the
installer tests.

Refs #1617

Contract-Neutral: Refs #1617: QNAP installer layout fix with its deployment-installability contract clause staged in this commit; residual proof policies for unrelated boundaries do not apply to this shell-only change
2026-08-21 06:35:52 +01:00
rcourtman b549e232a5 Keep each alert recurrence as its own history row
Since the operational-trust records shipped in v6.2.0, every new firing
of a previously resolved alert folded into that alert's first history
row: setActiveAlertNoLock stamped the new occurrence's open record onto
the previous occurrence's resolved row, and the history dedup then
treated any two open records with the same identity as one incident
regardless of how far apart they were. Alert history therefore froze at
the upgrade date while notifications kept flowing, which is exactly how
users reported it.

Resolved rows now keep their final record unless the update belongs to
the same occurrence, and two open records no longer merge on identity
alone. The observation-gap window still coalesces genuine flapping, and
the five-minute refire continuity path is unchanged.

Refs #1497

Contract-Neutral: Refs #1497: behavioral bugfix in alert history occurrence dedup; no API shape change, history row schema unchanged, no subsystem contract names sameHistoryIncident
2026-08-21 06:30:16 +01:00
rcourtman 6bad0bd884 Give notification retries a horizon that survives a destination restart
The default attempt budget was three, and with the 1s/2s/4s backoff a
destination that was unreachable for about ten seconds had its
notifications dead-lettered permanently. A webhook receiver rebooting
alongside the infrastructure it monitors is routine, not terminal.
Eight attempts under the same doubling schedule span roughly three
minutes before dead-lettering. Dead-letter semantics are unchanged.

Refs #1721

Contract-Neutral: Refs #1721: raises the default delivery attempt budget only; dead-letter semantics, retention, and outcome vocabulary unchanged, notifications.md pins no attempt count
2026-08-21 06:21:33 +01:00
rcourtman b54b4c81c7 Stop stamping powered-off severity onto overrides that never set it
normalizeOverrides ran every override through NormalizePoweredOffSeverity,
which maps an unset severity to an explicit warning. Any guest with any
per-guest override (a disk tweak, a note) therefore had its powered-off
alerts silently downgraded from a global critical default to warning, and
the stamped value also round-tripped back to the UI as if the user had
chosen it. Leave unset severities empty so the merge keeps following the
global default, and normalize only values the user actually set.

Refs #1738

Contract-Neutral: Refs #1738: behavioral bugfix in alerts override normalization; poweredOffSeverity stays an optional field, no subsystem contract names it, no API shape change
2026-08-21 06:19:12 +01:00
rcourtman 7da385cf28 Prepare v6.3.0-rc.4 release 2026-08-21 00:00:37 +01:00
rcourtman 4c7b1a2434 Fix Docker-in-LXC probe storm against slow Proxmox hosts
The guest Docker socket probe hung minipc hard enough to need a power
cycle (2026-08-20): ~100 orphaned pct exec children, load 133, sshd and
pveproxy starved. Three bugs chained, each fixed here:

1. Dispatcher re-issued a probe while the previous one was still
   executing. The poll cycle's enrichment context had expired, so
   ExecuteCommand dispatched, returned the context error 50ms later,
   and the next 3s cycle sent the identical command again — unbounded
   concurrency against a host that was slow to begin with. The
   monitoring dispatcher now takes a per-guest in-flight claim before
   dispatching probe or inventory commands (completed probes release
   it; abandoned ones hold it for a 2-minute window), and both dispatch
   paths bail out under a dead context.

2. The host agent never got the July process-leak fix: 45480a5cc
   landed only on pulse/v6-release, so main-line agents killed just the
   direct shell on timeout, orphaning pct exec → lxc-attach children
   and blocking Wait on their inherited pipes (10s timeouts reported as
   300s+ durations). Port it: run each command in its own process
   group, SIGKILL the group on cancel, bound Wait with WaitDelay, and
   treat ErrWaitDelay after a clean exit as success.

3. Server-side abandonment never reached the agent. ExecuteCommand and
   ReadFile now refuse to dispatch under an already-expired context,
   and send a best-effort cancel_command when they stop waiting; the
   agent cancels the in-flight execution (killing its process group)
   and reports "command canceled". Older agents ignore the unknown
   message type.

Also add a per-node circuit breaker: three consecutive command failures
on one node suspend all Docker probe/inventory dispatch to it on the
existing 1m→30m backoff schedule, so a host-level stall (NFS flapping)
stops the probing entirely instead of failing guest by guest.

Regression tests simulate the storm without hardware: a never-returning
executor is not re-issued across poll cycles, an expired context
dispatches nothing and records no failure, abandoned probes hold their
claim, the breaker blocks new guests on a failing node, and the agent
kills the whole process group on timeout and on server-issued cancel.

Contract-Neutral: monitor.go delta is three private struct fields holding Docker probe dispatch state; host-agent deletion/re-enrollment lifecycle untouched — contracts and all other proofs are staged
2026-08-20 23:33:36 +01:00
rcourtman 5ed053dfc2 Count the alert history severity chips from the list's own predicate
The alert history FilterBar rendered bare Critical/Warning chips while
the platform tables' status chips now carry counts. Expose
countForSeverity from useAlertHistoryState, built on the same
filterAlertHistoryItems predicate the list renders through, so each
severity chip shows the row count its selection yields for the fetched
period and current search. Counts respect the shared Inventory totals
visibility preference; the Period facet stays uncounted because it is a
time scope, and only the currently fetched range is available to count.
2026-08-20 22:08:43 +01:00