Record incident history qualification and delivery

Record the landed implementation, passing final CI and exact local proof.
Retain the original benchmark failure and API timeouts, and keep model
reliability and independent-environment readiness limits explicit.

Refs #1782
This commit is contained in:
rcourtman
2026-09-08 03:20:20 +01:00
parent 783571bb35
commit 56087f534b
2 changed files with 98 additions and 25 deletions
@@ -3666,9 +3666,11 @@ on integrated main `977afdd9559c0e9d5859f4c79bcc48e889381bba` plus the scoped
changes. Affected regressions and final browser interaction/pixel checks pass.
Three funded Astra explanations passed, with both earlier Gemini failures
retained. The actual note survives reload and its text reaches Assistant.
Repository commit checks and scoped landing remain pending. Both the old
candidate and matching unchanged base timed out in the full API race suite,
so no passing whole-suite receipt is claimed. The subscription-provider refusal
Worker commit and push hooks passed. Implementation commit
`e2b6fe3b1658de784ee8f579831b7777dc211d00` landed through PR1973 as
`783571bb35c42c4813415d4981673f65c069ee0f`, after all 28 reported checks passed. The final CI API suite passed. The earlier old
candidate and matching unchanged base timed out in their full API race runs,
which remain failed historical receipts. The subscription-provider refusal
is preserved without retry or bypass, and wider readiness remains open.
### Incident-history candidate r1: implementation and proof in progress
@@ -4192,12 +4194,11 @@ remain recorded above. The final helpers are executable at
Their receipt directories are `browser`, `browser-states`, `browser-resize` and
`browser-saved-explanation` under that task directory.
Repository commit checks and scoped landing are the remaining local delivery
steps. The broader customer-outcome gap stays open. These receipts do not
Worker commit and push hooks passed. PR1973 landed the scoped continuation
after all 28 reported checks passed. The broader customer-outcome gap stays open. These receipts do not
qualify independent customer environments, unattended autonomy, backup restore,
population false-alarm/miss rates or latency SLOs. The failed broad API race
runs remain an explicit test limit, despite passing affected package and handler
proof. The final history explanation qualifies the funded Astra route and does
population false-alarm/miss rates or latency SLOs. The earlier failed broad API race
runs remain recorded alongside the passing final CI API suite. The final history explanation qualifies the funded Astra route and does
not erase the two Gemini failures or transfer the earlier named lab matrix to
another model.
@@ -4236,3 +4237,89 @@ Its audit and contract audit pass. The full hook then caught one expected-file
fixture missing the newly registered monitoring regression. Updating that fixture
to include the actual proof file preserves the guard's exact mapping assertion.
All 130 completion-helper tests passed in 2.645 s. This is test-only scope.
### Incident-history delivery evidence
Implementation commit `e2b6fe3b1658de784ee8f579831b7777dc211d00` is in
[PR1973](https://github.com/rcourtman/Pulse/pull/1973). All 28 final reported checks passed after the controlled benchmark rerun.
It merged at 02:11:45 UTC on 2026-09-08 as
`783571bb35c42c4813415d4981673f65c069ee0f`. No release has been published.
The complete configured pre-commit and pre-push hooks passed on the worker,
as rcourtman through its normal allocator, with Go 1.26.8 and GOMAXPROCS=4.
The staged tree remained `e31e2e077100922d229a25bc041e63a9311e6dfd` before and
after the hooks. The raw hook log SHA256 is
`e9d31b0347eb8464f87a4db5cbeb67b61f0cbc398d8748033003f38fc00557ea`.
This includes governance audits, 130 completion-helper tests, 163 lookup tests,
frontend lint/audits and TypeScript checking. The repoctl tests passed in
1.053 s. The worker did not have golangci-lint, whose optional hook block was
not run. CI remains responsible for its configured Go checks.
Local commit and push used task-local hooks that checked the exact tree against
the successful worker receipt. The commit-message hook and no-attribution check
still ran locally. Heavy checks were relocated, with no runtime source changes
between worker acceptance and the commit. The final complete memory race suite
passed in 2.344 s, and focused adapter query race tests passed in 2.071 s.
Raw logs remain under `/opt/pulse-release-worker/patrol-incident-history/`,
including `hook-final.log` and `final-memory-race.log`. Failed environment and
fixture attempts are retained separately. Final browser source, routes, widths,
interactions, funded outcomes and limits remain as recorded above. These local
results do not close the independent-environment rollout gate.
Model acceptance remains route-specific. The restored Gemini default has not
received a passing replacement receipt for its two unsupported action-absence
claims. Passing Astra history explanations do not repair or qualify Gemini's
judgment. No default-model change or production model-readiness claim is made.
Further default-route reliability work remains explicit alongside the separate
independent-environment rollout gate.
The first CI benchmark job 101904533464 failed only `NormalizeRoute/root`,
2.184 ns versus 2.497 ns, +14.31% (p=0.000, n=10), on AMD EPYC 7763.
Auto-merge was disabled on the failure. The exact tested synthetic merge was
`059179018ff6e650b489cdc6eeb6b619d70e8fbc`, against base
`9b57f7cacdb77c2d532a009295a3b34220281111`. Neither the normalizer nor its
benchmark source changed. In the worker-built test binaries, the benchmark loop
is identical and the normalizer instruction sequence differs only in relocated
addresses, which does
not establish equal microarchitectural timing.
A targeted worker comparison used those exact revisions, Go 1.26.8,
GOMAXPROCS=4 and ten alternating pairs at each of 100 ms and 1 s on an Intel
i5-10600. Results were 1.689 ns versus 1.714 ns (p=0.971) and 1.666 ns versus
1.714 ns (p=0.190), with zero allocations in both trees. The CI failure did not
reproduce in this sample. Load rose from 3.10/2.84/1.65 to 6.46/4.36/2.35 on the
eight-vCPU worker. A separate release test workload was observed afterward, so
this is not an idle-host receipt and does not establish the cause of the CI
failure. No further worker benchmark was started while that workload ran.
The unchanged source/loop and this comparison justify one controlled rerun of
the failed CI job with the same threshold and ten-sample policy. No parser
padding, unrelated optimization or threshold exception was introduced.
The executable worker recipe is
`/opt/pulse-release-worker/patrol-incident-history/benchmark1973.sh`, with raw
pairs and disassembly in its `benchmark1973/` directory. The original CI artifact
is retained in workspace `tmp/patrol-incident-history/bench-results-1973/`.
The complete final CI API job 101904533527 passed, from 01:20:38 to 01:51:05
UTC on 2026-09-08. The API package itself passed in 1721.252 s, recorded in
workspace `tmp/patrol-incident-history/ci-api1973.log`. All eight hosted Playwright shards and both remaining backend
shards also passed. This current receipt qualifies the tested merge candidate
and does not rewrite the earlier worker API timeouts as passes.
The one controlled CI benchmark repeat, job 101911401066, passed on AMD EPYC
9V74. Its exact base and candidate merge SHAs were unchanged. NormalizeRoute/root
measured 2.419 ns versus 2.466 ns, +1.94% (p=0.001, n=10), below the unchanged
>10% gate. Allocations remained zero. This is a measured increase, not identical
timing, and the different runner does not establish why the first measurement
failed. No further retry was requested. Raw log:
workspace `tmp/patrol-incident-history/ci-bench1973-retry.log`.
The final [Build and Test run](https://github.com/rcourtman/Pulse/actions/runs/34175674423)
and [Core E2E run](https://github.com/rcourtman/Pulse/actions/runs/34175674586)
passed. The final 28-check snapshot is retained in workspace
`tmp/patrol-incident-history/pr1973-final-checks.json`. Local main was fast-forwarded
to the merge. Its only additional source difference from the implementation
commit is the upstream deadman test. No runtime or accepted frontend content
changed during landing. This is scoped implementation delivery, with no release
publication or wider readiness assertion. The customer-outcome gap and proposed
candidate stay open for route reliability and independent-environment evidence.
+3 -17
View File
@@ -10201,7 +10201,7 @@
},
{
"id": "patrol-assistant-customer-outcome-qualification",
"summary": "The wider readiness gate remains open. The executable contract, honest baseline, exact runs and residuals are in docs/qualification/PATROL_ASSISTANT_CUSTOMER_JOURNEY.md. Earlier evidence/history/risk and typed-runner changes landed through core PR1928/1929/1934/1935/1951/1955 and enterprise PR22. PR1957 auto-merged as a66b8e11d7ca9ed5660ffd8725a1461661ca2fdf despite its recorded benchmark failure, which is not counted as a pass. The current slice makes canonical planning acceptance/refusal available inside the model turn, retains accepted actions across provider failure, enforces persisted actor/request identity, preserves complete approval/risk/outcome context and removes proposal/tool-count diagnosis proxies. Shared disk unknowns, streamed whitespace and replayed resolution timestamps are corrected. Affected package race suites and final frontend checks pass. Source-bound r28 Gemini qualification performed healthy, unhealthy, dependency, missing-access, storage-capacity, approved and rejected cases, including independently observed Docker recovery and VM110 start/stop. Earlier semantic failures remain recorded. Current r34 runtime passes final linked-history and Assistant Playwright at 1440, 900 and 390 widths, including fresh approved explanation, saved missing-access/VM/rejected/storage continuation, deep evidence, exact action states and scrolling. The local named implementation matrix is performed, with source-equivalence and exact limits recorded. Verified local delivery landed through core PR1960 (a42e3800d9a1b5ea185823469e891f44bc698c2c) and enterprise PR23 (b9fa43dcf0ee743652b20d1a866da8ca9c82cdbd). Their final CI and source-bound local proofs pass. Production-wide readiness remains a separate release_gate. Temporary collectors, runners, tokens and fixtures were cleaned and production processes preserved. Storage capacity does not qualify backup/restore. Legacy incident-memory/compatibility residuals, all-filter coverage, unattended autonomy and latency SLOs remain unqualified. Adoption of 127 paid installations, 71 Patrol-enabled installations and 23 Assistant users does not measure effectiveness. Fourteen verified resolutions from one installation do not establish population useful-diagnosis, false-alarm, missed-problem or latency rates, which remain unknown. Independent volunteered customer environments are the separate wider rollout gate. Preserve the explicit subscription-provider refusal without retry or bypass. Do not close this gap or candidate while wider readiness evidence is missing. The incident-history continuation implements canonical filtered occurrence queries, provenance, bounded history and explicit read errors, then repairs duplicate saved shells and mobile Assistant focus ownership found in live inspection. R4 browser checks pass, but its funded history explanation made unsupported action-record claims. Its initial related-resource criticism was withdrawn after auditing automatically attached history. R5 narrowed the shared history tool read scope, but its second Gemini explanation still asserted absent action records without reading actions. R6 repairs canonical alert target/risk ownership over legacy shells and adds reviewed cost estimates for funded GPT-6 Astra history qualification. Both funded Astra history and unavailable-read explanations passed factual review, with both Gemini failures retained and exact costs and context limits recorded. The actual note survived desktop reload, but resizing exposed a hidden mobile timeline. R7 projects shared expansion state into the phone drawer, and seven mobile-list tests pass. R7 also retains operator note text in the shared model context. Integrated main preserves upstream replay and checkpoint fixes, with a cross-resource occurrence-boundary regression corrected in the shared query. Final r10 build and source-bound browser acceptance pass, including saved notes, long-text wrapping, modal search ownership and keeping Assistant visible when resizing the destination. The additional funded history-and-note explanation passed. Five turns and titles cost US$0.36895475 estimated, with both earlier Gemini failures retained. Runtime defaults were restored and the subscription refusal preserved. Repository commit checks and scoped landing remain pending. Both the earlier candidate and its exact unchanged base timed out in the full API race suite at 30 minutes, with different active tests. Neither is a passing whole-suite receipt, and the comparison does not qualify the integrated candidate. These continuation requirements remain open.",
"summary": "The executable redesign plan, product contract, honest baseline and exact qualification receipts are in docs/qualification/PATROL_ASSISTANT_CUSTOMER_JOURNEY.md. Model-owned investigation and the named healthy, unhealthy, dependency, missing-access, storage-capacity, approved and rejected local matrix landed through core PR1960 and enterprise PR23. The shared incident-history continuation landed through PR1973 as 783571bb35c42c4813415d4981673f65c069ee0f, preserving canonical provenance, risk, bounded history, operator notes and the linked Assistant journey. Final worker hooks, affected race suites, all eight hosted Playwright shards, frontend, API and remaining CI tests pass. All 28 reported checks pass after one controlled benchmark repeat. Its first failure and loaded-worker comparison remain recorded, as do the earlier API timeouts. Final source-bound local browser proof covers /alerts at 1440, 900, 768, 767 and 390 widths. Five funded history turns cost US$0.36895475 estimated, with three Astra passes and two Gemini failures. The restored local Gemini route remains unqualified for its unsupported action-absence claims. Passing Astra history explanations do not qualify Gemini or transfer the earlier lab matrix between models. Production-wide readiness requires independent volunteered customer environments and known-condition outcome review beyond this maintainer homelab. Adoption and reported resolutions do not establish population useful-diagnosis, false-alarm, missed-problem or latency rates, which remain unknown. Storage-capacity proof does not qualify backup restore, unattended autonomy or every compatibility filter. Preserve the subscription-provider refusal without retry or bypass. Keep this gap and candidate open while these evidence limits remain.",
"owner": "project-owner",
"status": "planned",
"recorded_at": "2026-09-05",
@@ -10371,7 +10371,7 @@
{
"id": "patrol-assistant-customer-outcomes",
"name": "Patrol and Assistant Customer Outcomes",
"summary": "Simplify Patrol around model-owned investigation and one issue-to-verified-outcome journey shared with Assistant. Preserve canonical evidence and remove proposal-as-proof and proxy-driven diagnostic policy while retaining deterministic authority and independent verification. Execute the ordered redesign plan in docs/qualification/PATROL_ASSISTANT_CUSTOMER_JOURNEY.md, qualify unhealthy service diagnosis, backup or capacity risk and supported VM/LXC actions with negative controls, and record real model quality, latency and cost limits. Independent volunteered Pro qualification remains required for wider readiness. Local implementation and the named qualification matrix are delivered through core PR1960 and enterprise PR23. The candidate remains open for the independent-environment release gate. The subsequent canonical incident-history continuation is still in qualification, including a recorded failed model explanation. Do not treat that continuation as delivered until its final evidence and landing are recorded.",
"summary": "Simplify Patrol around capable-model judgment and one issue-to-verified-outcome journey shared with Assistant. Preserve canonical evidence, mechanical authority and independent verification. The executable redesign plan and exact limits are in docs/qualification/PATROL_ASSISTANT_CUSTOMER_JOURNEY.md. Local implementation and the named lab matrix landed through core PR1960 and enterprise PR23. PR1973 delivered the shared incident-history continuation with final source-bound browser proof and passing CI. Keep this candidate proposed for unresolved Gemini history reliability and the independent volunteered-environment rollout gate. Local implementation proof does not establish production-wide effectiveness, backup restore or unattended autonomy.",
"status": "proposed",
"recorded_at": "2026-09-05",
"target_id": "v6-product-lane-expansion",
@@ -10401,21 +10401,7 @@
]
}
],
"work_claims": [
{
"id": "patrol-incident-history-continuation-coverage-gap-patrol-assistant-customer-outcome-qualification",
"agent_id": "patrol-incident-history-continuation",
"summary": "Canonical incident history and final linked Assistant qualification",
"target_id": "v6-product-lane-expansion",
"claimed_at": "2026-09-07T20:00:54Z",
"heartbeat_at": "2026-09-08T00:40:59Z",
"expires_at": "2026-09-08T02:40:59Z",
"work_item": {
"kind": "coverage-gap",
"id": "patrol-assistant-customer-outcome-qualification"
}
}
],
"work_claims": [],
"open_decisions": [],
"source_of_truth_file": "docs/release-control/v6/internal/SOURCE_OF_TRUTH.md",
"resolved_decisions": [