Concurrent timer and queue callbacks could apply an older health snapshot after a newer one, hiding a new delivery failure or resurrecting a dismissed warning. Serialize the complete read/apply operation without holding the monitor or queue mutex across alert updates.
Add isolated channel-controlled stale-clear and stale-raise regression cases. Removing the lock fails both final-state assertions; restored code passes 100 race-enabled focused repetitions. This does not qualify the integrated release candidate or clear unrelated adverse evidence.
Change-source: pulse-maintainer
Contract-Neutral: Restores monitoring contract extension point 21 immediate canonical delivery-warning reconciliation by serializing existing read/apply operations; no public API, verdict, throttle, alert identity, or agent-lifecycle contract changes. Existing isolated ordering regression tests cover both stale-clear and stale-raise outcomes.
GetMonitor starts polling concurrently, so monitor.mu does not protect the fixture host slice from State.GetSnapshot. Use the state-owned UpsertHost setter to match the reader lock while retaining all canonical-token diagnostics assertions. Addresses the fixture race reported in PR1943 rest-1; no production behaviour changes.
Change-source: pulse-maintainer
Hold the replay serialization lock while attaching the resource store to prove history catch-up stays off the router construction path. Retain the imported-history and idempotence assertions after releasing replay so availability cannot be achieved by dropping repair.
Change-source: pulse-maintainer
Callback and queue restart tests cover separate boundaries. Exercise the monitor handlers and ordinary dispatcher together so missing capacity cannot silently produce a recovery, restart cannot replace incident identity, and recurrence still reaches a local HTTP receiver.
Change-source: pulse-maintainer
The advanced proposal's complete backend shard exposed a shared PBS fixture that returned an array from the datastore-status endpoint. Supply the healthy datastore observation those node-metric lifecycle tests intend, preventing the live datastore evaluator from fabricating an unrelated connectivity incident.
Change-source: pulse-maintainer
Cover disabled alerting, storage suppression and canonical datastore overrides before the polled capacity lifecycle. Removing the conversion alias makes the disabled-override regression case fail.
Change-source: pulse-maintainer
Exercise mirror and SQLite restoration through missing connectivity, independent healthy capacity, confirmed recovery and a second restart. Negative control fails in both modes with the old missing-status recovery behaviour.
Change-source: pulse-maintainer
The storage persistence regression stopped after recovery, leaving a later occurrence untested. Extend both recovery-mirror and SQLite paths through renewed pressure, restart, missing capacity and measured recovery to protect occurrence identity without claiming notification receipt.
Change-source: pulse-maintainer
Verify failed node-link persistence rolls back both the displaced owner and replacement, preserves report watermarks and reservations, and permits a durable retry on the same store. A rollback-omission negative control reproduces the assertion failure; no runtime behaviour changes.
Change-source: pulse-maintainer
Commit link and unlink journal updates before publishing state, preserve manual selections across report and provider refresh boundaries, and reserve dormant owners across restart. Re-evaluate known automatic associations using provider names while retaining legacy unknown links.
Change-source: pulse-maintainer
Record reproduced restart gaps and legacy provenance ambiguity; do not enable destructive cleanup.\n\nChange-source: pulse-maintainer
Change-source: pulse-maintainer
Symmetric bridge fixtures can pass when one address inventory loses its filter. Exercise repeated ingestion with one-sided Docker bridges and valid custom management bridges, checking both link directions. Mutation checks reject loss of either filter and the earlier broad bridge exclusion.
Change-source: pulse-maintainer
Limit automatic Docker bridge filtering to the generated br-<12 hex> convention so custom management bridges such as br-mgmt remain eligible for host association. Keep docker-prefixed and non-global-unicast exclusions intact.
Change-source: pulse-maintainer
A Docker bridge or link-local address can be unique among monitored PVE nodes but also exist on an unrelated NAS. Counting PVE owners alone then creates a false reciprocal agent association. Exclude host-local IPs and known Docker bridge interfaces from automatic network evidence, preserving management bridges and explicit unicast report IPs.
Seven synthetic negative cases fail before this repair and pass afterwards. Focused matcher and host-report tests pass under the race detector. This prevents a reproduced backend misassociation; it does not establish the cause or resolution of issue #1930, whose diagnostic payload remains unavailable.
Change-source: pulse-maintainer
Preserve the reviewed alert and TrueNAS work while incorporating the canonical API metrics optimization and patrol qualification record.
Change-source: pulse-maintainer
Use the normalized connectivity status consistently when deciding whether storage capacity is actionable, preserving the existing offline suppression rule for case and whitespace variants.
Change-source: pulse-maintainer
Do not count empty or unknown storage status as recovery evidence. Normalise status spelling for connectivity checks while leaving capacity evaluation independent and preserving existing inactive/disabled storage behaviour.
Change-source: pulse-maintainer
Route labels classify ASCII digits and hexadecimal UUID bytes. Scan those
bytes directly while preserving numeric precedence, Unicode names and
invalid UTF-8 handling. This reduces the shared normaliser overhead exposed
by paired landing benchmarks without changing the benchmark gate.
Refs #1928
Contract-Neutral: ASCII route-label optimization preserves label values, identifier precedence, Unicode and invalid UTF-8 behavior, and agent lifecycle authority.
Existing HTTP receipt coverage exercised uninterrupted operation while restart coverage ended at internal callbacks. Recreate both disk-backed managers during a missing-observation incident and after recovery to verify local transport delivery, incident identity, retained history and recurrence together. This does not qualify an installed binary or external provider.
Change-source: pulse-maintainer
Validate the checked-in dependency and three service restart fault contracts
against disposable Docker resources before using them to assess model output.
Verify baseline, injected fault, refused duplicate injection, explicit fixture
recovery, and two-pass cleanup with unchanged pre-existing inventory.
These opt-in tests make no Pulse or model request. Fixture recovery is teardown
and does not count as an approval, rejection, execution or customer outcome.
Record exact live, owning-package and source-bound proof in the redesign plan.
Resolve monitored topology before checking command connections so unavailable
inspection retains the known resource and parent node. Prevent known targets
from falling through to a colliding agent ID, and preserve the single-agent
requirement when no target is supplied.
Use one failed tool envelope for diagnostic reads and file mutations. Missing
connections neither prove an installation problem nor count as successful
writes. Hypervisor lifecycle authority remains on its canonical action path.
Verify disconnected and unknown targets, collision isolation, all affected
tool handlers, token/WebSocket scope boundaries, and the linked Patrol and
Assistant failure journey at desktop and narrow widths.
A completed probe remained discoverable after its first caller returned,
allowing the next collection to reuse stale filesystem measurements.
Remove it and publish completion within one registry critical section.
Keep in-flight sharing, cancellation and timeout behaviour intact.
The controlled regression fails before this change. Twenty full package
runs and three race runs pass on the worker.
Refs #1928
Add a disposable service-storage fault with an independent filesystem
oracle, bounded tmpfs writes, identity checks and verified recovery.
Exercise overwrite and symlink refusal without contacting a model.
Align the published schema with supported summary-term groups and validate
the complete catalogue in CI. Record the exact proof and remaining model
and missing-access qualification limits in the customer-journey plan.
Canonical tier reconciliation rebuilt SQL and probed absent preferred tiers
for each fallback point, regressing batch reads and allocation costs. Reuse
bounded query shapes with current bindings and snapshot-scoped absence
checks, then append consecutive points directly to their output series.
Preserve coverage and ordering semantics and verify fresh bindings after
new preferred observations arrive. Integrate current main test additions.
Extend abrupt-exit resolution coverage through the real notification processor and local HTTP receiver. Assert surviving grouped members and the resolved event, retaining terminal cancellation and dispatch assertions. This does not establish installed provider receipt or exactly-once crash delivery.
Validation: restart/crash tests passed ten race-enabled repetitions; broader receipt/restart selection passed three. Negative control erasing the recovery event fails the HTTP event/member assertion.
Change-source: pulse-maintainer
The runtime and three restart scenarios use health_process_stop, but the
published schema rejected it. Accept that implemented injector and check
actual catalogue fault types against the schema to prevent recurrence.
Record the remaining missing-access and storage qualification gaps.
Integrate the latest alert, delivery and action-result changes with the
Patrol evidence conversation. Replace the conflicted browser receipt
with current source-bound qualification and fix shared warning-card
wrapping exposed by the intermediate-width check.
Real diagnostic and autonomous action outcome qualification stays open.
Manager callback tests do not establish notification transport receipt. Exercise real monitor callbacks and the queued generic webhook path through firing, missing metrics, recovery and a distinct renewed breach, protecting against false recovery messages when PBS observations disappear.
Validation: ten focused race-enabled repetitions passed; three paired repetitions with the existing guest recovery transport test passed. Removing the PBS missing-metrics guard makes this test fail on a false recovery webhook. No runtime behaviour changes.
Change-source: pulse-maintainer
Remove contextless evaluator and assessment passes, signal-count budgets,
and post-finding prompt replacement. Keep evidence tools available until
explicit run limits and retain incomplete assessments and provider errors
alongside accepted decisions. Failed file reads now preserve error status
through the model, telemetry and saved Assistant history.
Full chat, AI and tools packages, focused Patrol API and race tests pass.
Real read-only and scripted browser checks preserve failed reads and linked
uncertainty. Real-model and verified action outcome qualification remain open.
SQLite remains authoritative when the JSON recovery mirror cannot be renamed. Exercise firing and resolved snapshots across restart so a failed mirror cannot silently lose or resurrect an incident. Close the event store before shutdown to prevent a second checkpoint from masking failure; also verify error reporting and temporary-file cleanup.
Change-source: pulse-maintainer
Missing capacity and a measured empty datastore both report zero usage. Protect the existing distinction across durable manager restarts so missing observations cannot send a false recovery, while genuine recovery and subsequent high usage still dispatch callbacks. Assert incident identity and persisted lifecycle event counts as well as the callback boundary.
Change-source: pulse-maintainer
Tool-call totals do not establish diagnostic sufficiency. Preserve seed-only
and failed-read conclusions, remove count-based completion instructions from
evidence, and retain configured limits and authority checks.
Keep findings grouped under alerts selectable in the shared review panel so
their investigations and access limits remain available to Assistant.
Main CI job 101407611069 failed when the pending-start test restored unrelated active alerts from the shared data directory. A synthetic persisted alert reproduces the same failure locally. Give all three legacy threshold fixtures their own temporary directory and stop their workers during cleanup, retaining every threshold and pending-start assertion.
Validation: seeded-directory focused race tests pass 30 repetitions after failing before the repair. No production alert behaviour changes.
Change-source: pulse-maintainer
Keep canonical disk risk, source freshness and retained history intact when
Assistant and Patrol gather evidence. Proposal acceptance validates an action
contract and must not rewrite uncertain conclusions as established root cause.
Preserve complete subscription tool batches without exposing routing envelopes
as answers. Keep wide answer tables readable and keyboard-scrollable on mobile.
Optimize retained tier reconciliation without discarding gaps or newer samples.
Record failed real-model diagnoses and outstanding autonomous qualification
separately from passing data-path and interface checks.
Retain storage policy aliases in durable metric and forecast metadata and use them during active-alert re-evaluation. Recover exact legacy PBS aliases from the recorded instance and datastore identity. Prevent a configuration reload from fabricating recovery against global defaults after restart.
Change-source: pulse-maintainer
Successful ntfy transitions alone do not prove a rejected recovery can be retried after restart. Exercise real HTTP 503/202 responses and SQLite reopen, retaining the firing receipt on failure and clearing it only after recovery succeeds. Assert failed and successful audits remain truthful without replaying the firing notification.
Change-source: pulse-maintainer
Exercise the real queue processor alongside direct delivery for warning, critical, recovery and same-identity refiring. Require HTTP payload/header receipt and committed sent state, using an isolated temporary queue.
Change-source: pulse-maintainer
The workspace-scoped managed relay proof still supplied an HTTP revocation origin after the canonical relay began enforcing HTTPS-only credential transport. Use a local TLS server and pass its certificate to the child process so the combined canonical repositories exercise the production boundary rather than failing during startup.
Change-source: pulse-maintainer
Resolution must retire obsolete queue rows without permanently muting a resource condition. Extend the restart/retry regression to require delivery of a new incident sharing the resolved alert ID, alongside recovery and surviving grouped alerts.
Change-source: pulse-maintainer
Exercise pending, interrupted, failed and dead-lettered grouped alerts without closing SQLite before restart. Verify that resolution survives and retry preserves only live members and genuine recovery. Also fix the disabled-delivery test race found by repeated race testing: queue construction already starts workers, so use the locked processor setter and await reconciliation.
Change-source: pulse-maintainer
Distinguish policy skips from provider success so suppressed jobs do not create false sent rows or successful audit entries. Reconcile cancelled queue health after releasing alert gates, preserve real attempt history, and cover all three providers for firing/recovery and global/destination disablement.
Change-source: pulse-maintainer
Preserve the reviewed notification retry finality fix and its exact ancestry after incorporating the batch-start upstream main.
Change-source: pulse-maintainer