291 Commits

Author SHA1 Message Date
xarmian a357b9f609 fix(mcp): restrict the remote transport to the protocol era pad can serve (TASK-2977) (#1310)
* fix(mcp): restrict the remote transport to the era pad can serve (TASK-2977)

mcp-go 1.0 implements the stateless protocol core from 2026-07-28 — no
handshake, no sessions, per-request identity in _meta — and its
Streamable HTTP transport advertises EVERY revision it implements by
default, serving both eras concurrently on one endpoint and deciding the
era per request. pad's construction site passed no version restriction,
so the library bump alone had main answering modern-era traffic through
server/discover while pad://_meta/version still published 2025-11-25 as
the maximum revision this server can negotiate.

NOT A PRODUCTION DEFECT, checked rather than assumed: app.getpad.dev and
mcp.getpad.dev both report commit 0e2cb06a built 2026-08-31, nine days
before the mcp-go 1.0 merge, and 0.58 has no 2026-07-28 constant at all.
The gap is in main, ahead of a hand deploy, so this lands before it can
become real.

Beyond the mismatch: pad_set_workspace pins a session default workspace
and the stateless era has no session for that pin to live in. So the
modern era is not something pad happens not to advertise, it is
something pad is not known to be able to serve. Establishing what it
would take is step 2 of this item; the honest advertisement meanwhile is
the era pad was built and tested against.

The set is DERIVED from mcp.LegacyProtocolVersions(), the SDK's own
answer to "which revisions use the handshake", so a future SDK adding a
legacy revision includes it and one adding a modern revision excludes
it, with no edit here. A hand-written list would silently mean the wrong
thing after either bump — the same shape as the defect being closed.

The option set moved into mcpserver.NewRemoteTransport so a test can
drive what cmd/pad actually constructs. An advertised set is only
correct if the option is PASSED, and a test building its own transport
would vouch for the option and not for the binding (CONVE-19).

Four tests, and the second is what makes the first mean anything:

- a well-formed 2026-07-28 server/discover against pad's transport is
  refused with code -32022, data.requested naming the version and
  data.supported carrying exactly the legacy four. Asserting "an error
  came back" would also pass on a transport that had simply broken.
- the identical bytes against an UNRESTRICTED mcp-go transport are
  SERVED, with 2026-07-28 among supportedVersions. Negative control: it
  is what says the refusal comes from pad's option rather than from a
  malformed request or a changed library default. Both requests carry
  the Mcp-Method header the modern era requires, so a refusal cannot be
  about headers.
- initialize still negotiates 2025-11-25 — the restriction must not
  break the era pad actually serves.
- AdvertisedMCPProtocolVersion equals the newest served revision, which
  is the half TestAdvertisedProtocolVersion structurally cannot see: it
  pins the literal to what the HANDSHAKE answers, and the modern era has
  no handshake.

Sweep: meta.go's two comments described the handshake cap as the whole
story. One construction site only — pad-cloud is an OAuth/billing layer
and builds no transport, so the pad binary in cloud mode is the single
place this is decided.

Claude-Session: https://claude.ai/code/session_01GqaEDuCtRiSJfa7eppWecn

* test(mcp): the handshake test measured less than its comment claimed (TASK-2977)

A mutation found this in my own instrument. Restricting the advertised
list to a version that EXCLUDES 2025-11-25 leaves the legacy-handshake
test green, so that test cannot be evidence that the restriction
preserved the era pad serves — which is exactly what its comment said it
was.

The mechanism, read rather than inferred from the green: initialize is
answered by MCPServer through mcp.NegotiateLegacyVersion, which consults
LATEST_LEGACY_PROTOCOL_VERSION and never the transport's list. The two
are independent.

So the comment now says what the test measures (the legacy path works)
and what it does not (that the restriction preserved it), and the
independence is pinned as its own subtest: a transport advertising only
2025-06-18 still answers initialize with 2025-11-25. If that ever fails,
the handshake has become coupled to the advertised list and the
restriction has become able to refuse legacy clients — the moment this
file needs a different test.

Nothing about the fix changes; the claim about the evidence does.

Claude-Session: https://claude.ai/code/session_01GqaEDuCtRiSJfa7eppWecn
2026-09-09 19:22:58 -04:00
David Barkhausen f262449b18 feat(cli): pad token create/list/revoke — CLI mint path for API tokens (#1237)
Contributed by b4rk13 (#879 follow-up). Reviewed under DECIS-212-style read: sits on the existing user-scoped /auth/tokens endpoints, no server changes; create requires a login session per #1267 and answers 403 session_required under PAD_TOKEN, list/revoke stay PAT-reachable.

Claude-Session: https://claude.ai/code/session_01W71Y4K5hGbbqqAhbVFnjB4
2026-09-09 16:09:29 -04:00
xarmian 4618876e3e fix(cli): pad server stop signals only a process it can prove is ours (BUG-2969) (#1299)
fix(cli): `pad server stop` signals only a process it can prove is ours (BUG-2969)

Measured on the merged binary before this change: a `sleep 600` whose pid had
been written into the PID file was SIGTERMed, and stop printed "Server stopped."
No pad server was running anywhere near that config.

Three things had to be true at once for that. os.FindProcess succeeds for ANY
pid on Unix. Nothing asked whether the pid belonged to a pad server. And the
confirmation loop polled the PORT — which is unhealthy from the first poll when
nothing was ever serving, so the success check was satisfied by the failure
case.

Liveness is the wrong question, and this is the trap the obvious fix falls into:
the stranger WAS alive. The question is whether the pid is OUR server.

## The discriminator

Unix takes an advisory flock on the PID file, held for the server's lifetime.
`stop` probes it non-blockingly: acquiring it proves nobody holds the file, so
the record is stale whatever the pid now names; failing to acquire proves a live
pad server holds THIS file. One implementation for Linux and macOS, no new
dependency, and the same primitive session_lock_unix.go has used since
TASK-2767.

Windows has no flock in that pattern, so it compares the process creation time
from GetProcessTimes against the one recorded at start — the attribute that
survives pid reuse, since a reused pid belongs to a process that started later.

The lead first ruled start-time comparison on every platform; I objected with
the cost (three implementations — /proc, a macOS sysctl promoting x/sys to a
direct dependency, and GetProcessTimes) and the ruling changed to this hybrid.
The cost table is on the item so the next reader sees why the shape moved.

The PID file gains a fingerprint on both platforms — pid, start time, executable
path — as JSON, with the legacy bare-integer form still parsed. A legacy record
carries no proof, which reads as UNPROVABLE, and unprovable means nothing is
signalled.

## Three races, each found by codex and each the same shape

1. Reading the record and checking ownership were separate steps, so a successor
   could claim the file between them: the lock then reported "held" — truthfully,
   about the successor — while the pid handed back was the predecessor's.
   pidFileOwner now returns the record it read from the descriptor it probed.
2. Removing the PID file after a successful stop could delete a fast successor's
   live record. It no longer removes at all there: the server removes its own on
   the way down, and a file left by a crash is handled by the next stop.
3. Removing a STALE file after the probe released the lock had the same window.
   The removal now happens inside the ownership check, while the lock is held —
   the only moment at which no replacement can have claimed the path. A claim
   arriving during that instant retries for half a second rather than losing its
   claim for the life of the process.

Windows deliberately does NOT delete a stale file: with no atomic primitive, a
check-then-remove would race a successor, and a stale file that the next start
overwrites is recoverable where a wrongly deleted record is not.

## Verified

Negative control, and it is the literal one: with the ownership check bypassed,
`go test` reports `signal: terminated` — the test binary is SIGTERMed by the
code under test, because the stale record names the test process itself.

Live, in throwaway HOMEs: a stale record naming a live `sleep` is refused and the
sleep survives (it was killed before this change); a stale record with a HEALTHY
port answering is still refused, nothing signalled, and both the stranger and the
real server survive; a server stopped through its own held record stops, and its
file is gone.

The CI smoke on windows-latest now stops the server with `pad server stop`
instead of Stop-Process, because that is the only place the Windows ownership
check runs — a smoke that killed the process directly would leave the
GetProcessTimes path unexercised on every platform.

make lint, make test green; codex CLEAN in round 4.

Claude-Session: https://claude.ai/code/session_01HeChkgZVYb3NTgTcckF5KR
2026-09-08 19:33:15 -04:00
xarmian 5ec17a7c92 fix(cli): pad server stop stops the server that is running, or says it is (BUG-2965) (#1298)
fix(cli): `pad server stop` stops the server that is running, or says it is (BUG-2965)

`StopServer` read the PID file and, on any read error, answered "server not
running (no PID file)" — without asking whether anything was listening. The file
was written in exactly one place, EnsureServer's auto-start branch, so a server
started any other way held the port with no file to find: a service unit, a
human running `pad server start`, the seats' refresh recipe relaunching with the
killed process's argv. A stop command that says "not running" about a running
process leaves the caller believing they stopped something, and the next thing
they do rests on that belief.

Two halves, per the item's property and corollary:

  - A missing PID file now asks the port. Only an unhealthy address earns "not
    running"; a healthy one earns a message naming the address, the missing
    file, and what to do instead. Deliberately NOT "find the listener and kill
    it" — resolving a pid from a port is platform-specific, and the process
    holding it may not be ours. A stop that kills by port can kill a stranger.
  - `pad server start` claims the PID file itself, so the file exists for every
    start path rather than only the auto-started one.

The second half took four codex rounds to get right, and each round found the
previous shape reintroducing the defect it was fixing:

  1. Write-then-defer-remove let a duplicate start overwrite a running server's
     entry and then delete it on the way out, leaving a healthy server
     unaddressable.
  2. Refusing to replace a live pid fixed that and opened its mirror: the start
     that LOST the port could still own the file, so the winner was unaddressable.
     The fix is ordering, not arbitration — BIND FIRST, then claim, so the file
     always names the process that owns the address. internal/server grows
     Listen and Serve for that; ListenAndServe is now the two together.
  3. With the bind first, EnsureServer's parent-side write became the stale
     mechanism (it records a child that may never bind) and the live-pid refusal
     became actively wrong (no live process can be serving an address we just
     bound). Both removed, along with processIsAlive, whose only remaining
     callers were its own tests.
  4. Cleanup is a read-then-remove, so running it AFTER the listener closes let
     a successor bind and claim between the two steps and lose its file to us.
     It now runs before the listener closes, while nothing else can legitimately
     own the file. The cost is a drain-window where a healthy server has no PID
     file and `stop` says so — a true message in place of a silent wrong one.

Verified live against the built binary, in a throwaway HOME, in both shapes:
start writes the file naming the serving process; a second start against the
held port fails at bind and leaves the first server's file intact; stop then
stops it and removes the file; a further stop reports "not running". The first
live run also caught a flaw in my own method — `stop` reads the config's port,
so the probe answered about 127.0.0.1:7777 (this box's dev server) until it was
re-run with PAD_PORT set. Re-checked after the restructure.

Mutants: the health check removed, the health branch still answering "not
running", an empty PID file, a cleanup that does not remove, and a cleanup that
removes a successor's file are each killed by a named test. The call site itself
is wiring a unit test cannot vouch for (CONVE-19) — that is what the live runs
cover, and the Listen/Serve split is pinned in internal/server.

make lint, make test green.

Claude-Session: https://claude.ai/code/session_01HeChkgZVYb3NTgTcckF5KR
2026-09-08 18:26:32 -04:00
Matt Faltyn 43d94a35d5 fix(cli): pass PostgreSQL connection strings to backup clients (#1289) 2026-09-08 12:30:45 -04:00
xarmian 367aae8e18 fix: one prefix grammar, and the ref parser widens to it (BUG-2943) (#1286)
* fix(collections): a DERIVED prefix is A-Z only (BUG-2943)

DerivePrefix took the first BYTE of each word, so a collection named
'TEMP Rook A 2870' got the prefix 'TRA2'. parseItemRef resolves a
PREFIX-NUMBER ref only when every prefix character is A-Z and otherwise falls
through to a slug lookup, so every item in that collection printed an issue ID
the CLI then refused: 'pad item show TRA2-2942' answered 'item not found'
while the slug resolved fine. Two functions, each locally reasonable,
disagreeing about what a prefix may contain — and the generator was the
permissive one, so the failure surfaced at read time on an identifier the
product itself minted and printed.

The first-BYTE bug had a second half: a word starting with a multi-byte rune
contributed a UTF-8 lead byte, so a collection named in most non-Latin
scripts produced a prefix that is not even valid text.

Non-letters are SKIPPED rather than mapped — there is no honest A-Z
substitute for '2' or 'Omega', and inventing one puts a character in the ID
that is in nobody's collection name. A name with no ASCII letters yields the
empty string, which store.CreateCollection already turns into its ITEM
fallback.

SCOPE, stated because the first draft of this message overstated it (codex
round 1 [P2]): DERIVED prefixes are safe now; the INVARIANT IS NOT ENFORCED.
Three other doors store a prefix verbatim and unvalidated — CreateCollection
with an explicit input.Prefix, UpdateCollection, and workspace import — so
the same unresolvable-ID defect is still reachable through the API, the
--prefix flag and a restore. Named on the trail with their call sites, held
for a ruling rather than swept into this commit, because the import door
wants a different answer from the other two: refusing a restore is not
obviously right.

The parity test lives in internal/store, where parseItemRef is: it asserts
the generator against the RESOLVER rather than against a restatement of the
resolver's rule, which is how these two drifted apart in the first place.

Claude-Session: https://claude.ai/code/session_01HeChkgZVYb3NTgTcckF5KR

* fix(store): one prefix grammar at all four doors, and the parser widens to it (BUG-2943)

The ruled shape, which dissolves the import dilemma rather than choosing a
side of it: collections.IsValidPrefix is the single definition — an uppercase
letter followed by uppercase letters or digits — and parseItemRef now asks it
instead of carrying its own stricter A-Z rule.

Because the PARSER widened, a workspace already carrying a prefix like AB1
resolves every item by its printed ID the moment this ships. No migration, no
rewrite of an identifier a user's other records may reference.

The four doors:

- derive: unchanged from the previous commit, still letters-only, still
  within the grammar;
- create with an explicit prefix: REFUSED if outside the grammar, with a
  message naming the rule. The caller typed it, so a refusal is actionable;
- update: same, and it matters more here — update is the door someone reaches
  for to FIX a bad prefix, so it must not accept another one;
- import: the most permissive door that can still be honest. Anything the
  parser resolves is accepted (which now includes digits); only a prefix NO
  surface could resolve is refused, naming the collection and saying the
  export can be edited. Carrying that verbatim would restore a workspace
  whose items print IDs the CLI answers 'not found' to, which is this
  item's defect rather than a compatibility owed.

An ABSENT prefix on import is not an unresolvable one. Old exports and every
fixture in the suite carry "", and the first version of this check refused
them — turning a fix for unresolvable IDs into one that cannot restore an old
bundle at all (caught by three server tests). It now takes the same
derive-then-ITEM fallback CreateCollection applies, which also upgrades it: an
empty prefix is itself unresolvable, since the ref would begin with a dash.

A prefix accepted only because the parser widened is logged at WARN, so an
operator can see an id-space that would have been rejected before rather than
inferring it from a resolve failure that no longer happens.

Tests for each door follow in the next commit.

Claude-Session: https://claude.ai/code/session_01HeChkgZVYb3NTgTcckF5KR

* test(store): one test per prefix door, plus the parser round trip (BUG-2943)

Each door asserted separately: 'they all call the same helper' is a claim
about the code, not about behaviour, and the bug was two definitions
disagreeing.

- create with an explicit prefix: AB1 accepted and resolves; ab1, 1AB, 'A B',
  A-B, A!, a non-Latin letter and a bare digit refused, with the rule named;
- update: AB1 accepted, a bad replacement refused AND the stored prefix
  unchanged after the refusal — update is the door someone uses to FIX a bad
  prefix, so it must not swap one unresolvable id-space for the next;
- import: a digit-bearing prefix restores unrewritten and resolves; one no
  surface can resolve is refused naming the collection and the export; an
  ABSENT prefix takes the create-path fallback and comes back resolvable;
- the parser: every prefix the doors accept round-trips, and 1AB / 9 / 'A B' /
  A! / a trailing dash / a bare prefix stay refused.

One correction: my first version of the parser test asserted that 'ab1-42' is
refused. It is not, and the code is right — parseItemRef upper-cases before
splitting, which is what makes "pad item show task-5" work. Case-insensitivity
is now PINNED rather than mis-asserted, because a later reader working from
the grammar comment alone would otherwise 'fix' it.

Claude-Session: https://claude.ai/code/session_01HeChkgZVYb3NTgTcckF5KR

(This message was rewritten once: the sentence above lost its example because
the original was written with backticks inside a double-quoted shell string,
which the shell EXECUTED and replaced with the command's empty output. The
span was blank in the commit as first written.)

* fix: the widened grammar reaches its consumers too (BUG-2943)

Codex round 2. Widening parseItemRef without widening what CONSUMES a ref
would have left the same two-definitions bug this unit is about, introduced
by its own fix:

- cmd/pad/cmd_github.go matched [A-Z]+-\d+, so 'pad github link' on a branch
  carrying a digit-bearing ref silently found nothing;
- web localSearch's palette Enter fast-path could not recognise one either.

Both now match collections.IsValidPrefix.

Tests strengthened, both on codex's reading:

- the import fallback pinned the VALUE, not just resolvability — asserting
  'non-empty and parseable' passes an implementation that stamps ITEM on
  every absent prefix, giving every collection in a restored workspace the
  same id-space. Two legs now: an ordinary name derives TASK, a letterless
  name falls through to ITEM, which is what makes it 'derive, THEN ITEM';
- the WARN the ruling asked for had no test, so it was a line nobody would
  notice was gone. Now asserted, with a control that an ordinary prefix does
  NOT warn — a log everything trips is a log an operator learns to skip.

Three comments still described the parser's A-Z rule as current, including
one in the file that changed it.

STILL OPEN, on the trail for a ruling: web paneTarget.ts keeps the narrow
grammar on PURPOSE — its comment argues a digit-permitting shape would
misclassify a slug like 'roadmap2-5' as a ref — and that argument cited the
server rule this unit just widened. Whether the guard follows or stays is a
question about the widening's blast radius, not a line to change quietly.

Claude-Session: https://claude.ai/code/session_01HeChkgZVYb3NTgTcckF5KR

* fix(web,store): the pane guard follows the server, and the precedence is written down (BUG-2943)

Both lead-ruled after codex round 2 surfaced them.

paneTarget's REF_SHAPE kept a LETTERS-ONLY grammar deliberately, citing the
server's A-Z loop as its warrant. The server dropped that rule, so the guard
was holding a grammar nothing else holds — which does not avoid a wrong
answer, it produces a different one. It now matches IsValidPrefix.

The cost is real and is stated in the test rather than buried: an HREF whose
last segment is ref-shaped under the wider grammar is compared by NUMBER with
the prefix discarded, so a genuine slug like 'roadmap2-5' now counts as the
same pane target as TASK-5. The existing test pinned the opposite and is
REPLACED, naming what changed and why. The prefix is dropped because a moved
item keeps a stale one (the server's own number-only fallback) and
PaneGuardItem carries no prefix to compare; tightening that means widening
that type and its callers, which is a separate change and is on the trail.

The SLUG-channel leg is kept as its own test: provenance, not grammar, is
what protects it — a target naming an item by slug is judged only as a slug.

ResolveItem's ref-before-slug precedence is now documented on the function
and pinned in both directions: 'ab1-42' resolves as a SLUG when no AB1-42
exists, and a live ref wins when it does (case-insensitively). The widening
made more strings ref-shaped, so 'is my slug still findable' needed an answer
that does not depend on reading the resolver.

Web unit tests run here via a node_modules SYMLINK to the main checkout,
which CLAUDE.md permits; npm ci was not run and must not be.

Claude-Session: https://claude.ai/code/session_01HeChkgZVYb3NTgTcckF5KR

* fix(web): the self-pane guard was the fourth copy of the ref grammar (BUG-2943)

Codex round 3 [P1]. The route-level refNumber() still matched [A-Za-z]+, so a
master whose ref is R2-1 parsed as null, its item_number fell back to 0, and
the same-item guard stopped recognising ?item=R2-1 as the master — mounting a
second provider for the item already on screen.

That is the FOURTH consumer found carrying its own copy of this grammar
(github branch extraction, the search palette, the pane target guard, and now
this). Four independent copies is the argument for the shared definition
rather than for four careful edits, and it is why the widening had to be
swept rather than applied where it was noticed.

Also from round 3: comments saying these client regexes 'match
collections.IsValidPrefix' were imprecise — the validator accepts uppercase
only, while the client patterns accept either case on purpose, because a user
types a ref however they like and the server upper-cases before splitting.
They mirror the ref GRAMMAR, and now say so.

Web gates run here through a node_modules symlink to the main checkout
(permitted; npm ci is not): vitest 2195 passed, svelte-check 0 errors.

Claude-Session: https://claude.ai/code/session_01HeChkgZVYb3NTgTcckF5KR

* fix(collections,e2e): the fifth and sixth copies, and the docs that taught the old rule (BUG-2943)

Codex round 4, after I claimed the sweep was complete twice.

- Two SEEDED PLAYBOOK BODIES carry their own ref grammar and instruct agents
  with it: playbook_library_plan.go and templates_sdd_spec.go both said a ref
  'matches ^[A-Z]+-\d+$'. An agent following those literally would refuse to
  treat AB1-42 as a ref — a grammar copy that lives in PROSE and is executed
  by a reader rather than a regexp engine, which is why two sweeps of the
  code missed it.
- Three e2e comments taught the defect as a rule: one of them carries the
  empirical confirmation ('GET /items/BS1-10 404'd while the slug worked'),
  which is precisely this bug. They now say the by-ref 404 is fixed and that
  the explicit prefix those suites pass buys DETERMINISM rather than dodging
  it.

Counting honestly: six live copies of one grammar, found in four rounds of
review, two of which I opened by asserting there were no more. The shared
definition is the fix; every one of these was a place that had quietly made
its own.

Gates: go test ./... 0, make lint 0 issues, vitest 2195 passed,
svelte-check 0 errors (web run through a node_modules symlink to the main
checkout — permitted; npm ci is not).

Claude-Session: https://claude.ai/code/session_01HeChkgZVYb3NTgTcckF5KR
2026-09-07 17:32:36 -04:00
xarmian 3a00ea9c8c fix(cli,mcp): item update --parent "" is refused, not silently ignored (BUG-2941) (#1285)
* fix(cli): item update --parent "" is refused, not silently ignored (BUG-2941)

It exited 0 and printed the updated item while doing nothing: hasFieldChanges
tests parentRef != "", so the empty value built no patch and the key the
server's clear-path needs (parent present, empty) never reached the wire. The
two representations of the link then disagreed — parent_id read null while
parent_ref and the child listing still named the parent, and the parent still
could not be closed for open_children.

BUG-2078 shipped --clear-parent as the working route and left this one looking
like it worked. Refusing rather than aliasing it: two spellings for one
operation is what produced the confusion, and naming the flag that does the
job is the actionable answer.

UPDATE only. On create an empty --parent expresses nothing to ignore and
--parent "$MAYBE_EMPTY" is a normal shell idiom; a test pins that asymmetry
as a decision rather than a gap.

Compat note for review: a script passing --parent "$P" with P empty gets a
loud failure where it used to get a silent no-op. That is the point of the
change, but it is a real behaviour change for callers who were relying on the
no-op, and it is the one thing here worth a second opinion.

Claude-Session: https://claude.ai/code/session_01HeChkgZVYb3NTgTcckF5KR

* fix(cli,mcp): refuse before the request, and keep MCP's empty-parent convention (BUG-2941)

Codex round 1 found two P1s in the first shape of this fix.

1. THE REFUSAL WAS TOO LATE. It sat beside the parent handling, which is
   after the item fetch, so a refused call still made a GET. My test only
   watched for writes, so it passed a version that refuses after fetching —
   the test was not an instrument for the claim it was named for. The guard
   now runs first in RunE, before the client exists, and the test counts
   EVERY request rather than ignoring GETs.

2. THE FIX WOULD HAVE SPLIT THE TRANSPORTS. Stdio MCP shells out to the CLI
   and BuildCLIArgs emits a flag for any key that is PRESENT, so the
   catalog's `parent: ""` — documented inert since v0.19 — became
   `--parent ""` and would now be refused on stdio while the remote door
   went on ignoring it. A transport divergence created by a fix for a
   transport-independent bug, which is the class BUG-2870 exists to close.
   dropInertEmptyParent removes the key before dispatch.

Dropped there, refused at the CLI: same input, opposite dispositions,
because the two surfaces have opposite conventions about what an empty
declared string means. At the CLI a human typing it means "detach"; in the
catalog it means "not provided", and `clear_parent` is the documented way.

Claude-Session: https://claude.ai/code/session_01HeChkgZVYb3NTgTcckF5KR

* test(mcp): the empty-parent drop is about parent, not about emptiness (BUG-2941)

Codex round 2 [P2]: the test could not tell 'drops an empty parent' from
'drops every empty-valued key', which would be a much larger and undiscussed
change to the tool's input handling. It now carries an empty `comment`
alongside and asserts that one still reaches the CLI.

Claude-Session: https://claude.ai/code/session_01HeChkgZVYb3NTgTcckF5KR

* docs: two wording corrections from codex round 2 (BUG-2941)

BuildCLIArgs emits THIS STRING FLAG whenever its key is present — booleans
and hidden flags are handled differently, so the broader claim was wrong
even though it held for the case at hand. And '--parent "" reads as detach'
described the caller's intent as though it were the code's behaviour; it
never was, which is the whole bug.

Claude-Session: https://claude.ai/code/session_01HeChkgZVYb3NTgTcckF5KR
2026-09-07 16:18:39 -04:00
xarmian ee0d945863 fix(cli,mcp): one --field key=value entry means one thing at every door (BUG-2870) (#1283)
* feat(items): one shared parse for a --field key=value entry (BUG-2870)

Six sites parsed that entry independently — item create, list, update, move
and copy in cmd/pad, plus ingestFieldKVP on the remote /mcp door — in four
spellings, and they disagreed about what it meant. The CLI sites used both
halves verbatim, so `--field " effort=l"` stored an undeclared field named
" effort" and left the declared `effort` untouched; the remote door trimmed
both halves and wrote `effort`. Same call, two stored keys, decided by which
transport the caller was on.

This is the helper only; the call sites move over in the commits that follow.

Two rules, deliberately asymmetric, per the day-60 ruling:

- a KEY whose trimmed form differs from what was written is REFUSED at every
  door, rather than silently retargeted to a different field;
- a VALUE is carried VERBATIM at every door, because trimming reinterprets a
  caller's bytes and on a text field the space is content. A padded value
  against a typed field is refused one layer down by validation, naming the
  field — measured, not assumed.

ErrFieldEntryMalformed is returned rather than handled because the six sites
deliberately disagree about a malformed entry (four skip it, copy hard-errors)
and unifying that is a separate decision.

Claude-Session: https://claude.ai/code/session_01HeChkgZVYb3NTgTcckF5KR

* fix(cli,mcp): all six --field parse sites go through the one helper (BUG-2870)

item create, list, update, move and copy in cmd/pad, plus ingestFieldKVP on
the remote /mcp door, now call items.SplitFieldEntry instead of each rolling
its own split. A padded key is refused at every door; a value reaches every
door verbatim.

Two sites keep something specific to them, both documented in place:

- `item list` is a READ filter, and it takes the same key rule deliberately:
  a padded key there filters on a field nobody declared and returns empty,
  which is indistinguishable from "no rows match".
- `item move` gets KEY normalisation only. Its values stay strings because
  the server types a declared field on that path too, so a clean
  `--field n=3` already stores the number 3 — measured before the change.

Each site keeps its historical disposition toward a MALFORMED entry (four
skip silently, copy hard-errors), which is why the helper classifies that
case rather than deciding it.

NOT YET EVIDENCE: ./internal/mcp, ./cmd/pad and ./internal/items all pass,
and that green does not show the divergence closed — the three BUG-2850
pinned tests exercise the catalog conflict pass, which never reaches
ingestFieldKVP. The door-level test and the re-grounding of that pass are
the next commits.

Claude-Session: https://claude.ai/code/session_01HeChkgZVYb3NTgTcckF5KR

* test(cli,mcp): pin the door-parity claim at both doors (BUG-2870)

Nothing in the suite asserted what the remote door STORED for a padded entry
— the three BUG-2850 tests that cite its trimming all exercise the catalog
conflict pass, which never reaches ingestFieldKVP. So the previous commit's
green was not evidence for the thing it changed.

Three files now hold the claim: internal/items pins the rule, internal/mcp
pins the remote door, cmd/pad pins the CLI door, and each cites the other
two. Padded key refused at both; padded value carried verbatim at both; a
refusal aborts the call rather than dropping one entry, and on the CLI it
happens before any request reaches the server.

Claude-Session: https://claude.ai/code/session_01HeChkgZVYb3NTgTcckF5KR

* fix(mcp): re-ground the conflict pass on the new door behaviour (BUG-2870)

The pass's rules were derived from ingestFieldKVP trimming, so changing the
door without changing the layer built on it would have been the same
one-door lapse a level up.

- parseFieldArray splits through items.SplitFieldEntry: a padded key is
  REFUSED before dispatch on both transports, and values are indexed RAW,
  because raw is now what both doors write.
- Both comparison sites compare raw for the same reason. The round-19
  "COMPARED TRIMMED" rule is superseded and its comment says so.
- detectFieldConflicts PROPAGATES the parse refusal instead of returning nil.
  It swallowed it as "the caller owns this error surface", which was true
  when the only possible error was a shape error — reshapeItemFields returns
  early with no `fields` object, so on the no-`fields` path (this bug's path)
  nobody owned it and a padded entry turned back into a success.
- A padded entry is refused in the pass rather than skipped. Skipping dropped
  it from conflict detection entirely, turning four existing refusals into
  successes.

The last two were caught by the BUG-2850 tests, not by reasoning: the first
shape of this commit passed a full package build and turned four guards off.

Seven tests still fail. They assert the OLD door behaviour and are the
specification being changed; each gets read on its own next, and is either
kept because the behaviour survives or replaced by a test stating the new
behaviour that cites the old name.

Claude-Session: https://claude.ai/code/session_01HeChkgZVYb3NTgTcckF5KR

* test(mcp): restate the seven BUG-2850 pins on the new rule (BUG-2870)

Each was read on its own and either kept or replaced; every replacement
names the test it replaces and why the old assertion was right at the time,
so the deletion is traceable rather than a green that appeared.

- padded value is not a conflict → IS a disagreement now that no door trims
  (" done" and "done" are two values), with an equal-values control leg.
- padded entries still caught (hierarchy) → refused EARLIER, by the padded-key
  rule, before the alias pass observes both keys. The alias guard keeps its
  three unpadded cases, which is what stops this being a hole.
- PaddedEqualDuplicateIsCanonicalized → IsRefused, plus a canonical control
  that still emits --field exactly once.
- MixedCanonicalAndPaddedDuplicatesCollapse → Refused. The round-8 finding
  survives: one canonical entry still does not make its padded sibling
  harmless, it is refused rather than swallowed.
- PaddedEntryAloneIsUntouched → IsRefused. That test pinned a DEFERRAL, in
  its own words "BUG-2870's business, not this PR's". This is that business.
- "fields carries the key — canonicalized, so accepted" → still refused,
  since nothing canonicalizes now; the per-key question it defended is still
  tested by the two legs beside it, and a canonical control was added.
- ReEmittedValueKeepsItsWhitespace → the re-emission path is gone, so it
  becomes a refusal test that also asserts the ADVISED form is accepted with
  its value untouched. The property it defended is pinned at both doors.

Claude-Session: https://claude.ai/code/session_01HeChkgZVYb3NTgTcckF5KR

* test(server): pin that a move override is typed server-side (BUG-2870)

The fact the ruling turned on, and the easiest one in this unit to lose: it
is invisible from cmd/pad, where moveCmd plainly sends a string.

- a declared number field given the STRING "3" through field_overrides ends
  up as the NUMBER 3, which is why move needs the shared KEY parse and no
  client-side typing;
- a padded " 3" is REFUSED with a 400 and the item does not move, which is
  the answer the remote door will now give too instead of trimming and
  succeeding.

t.Parallel per CONVE-2086 — both build their own server through testServer,
so each has its own database, limiter and bus.

Claude-Session: https://claude.ai/code/session_01HeChkgZVYb3NTgTcckF5KR

* chore(mcp): bump tool surface to 0.30 and sync the docs the guards enforce (BUG-2870)

Claude-Session: https://claude.ai/code/session_01HeChkgZVYb3NTgTcckF5KR

* refactor(mcp): remove the canonicalization the door change made unreachable (BUG-2870)

Two mechanisms existed to make a padded entry reach both doors as the same
write: the nonCanonical conflict guard (round 16) and the re-emission path
that rewrote a padded entry to canonical form (rounds 7/8). Both are dead
now — items.SplitFieldEntry refuses a padded key, so every entry that parses
satisfies `entry == key + "=" + value` BY CONSTRUCTION.

Removing each changed no test. That is consistent with "dead" and with
"untested" alike, so the construction argument above is what settles it —
recorded in the comments that replace them, along with what the removed
guard was defending and where that premise is enforced now.

Rewriting a caller's key was also the behaviour this bug is about, applied
by us rather than by a door: canonicalization silently changed the key the
caller wrote. Refusing says so instead.

Claude-Session: https://claude.ai/code/session_01HeChkgZVYb3NTgTcckF5KR

* docs(mcp): put the trimming narrations in the past tense (BUG-2870, CONVE-23)

Six comments described the old door behaviour in the present tense ("HTTP
trims and writes effort"), which reads as a claim about the code as it
stands. The rounds they narrate still explain why the surrounding rules
exist, so they are re-tensed rather than deleted.

Two references were checked and left alone because they are still true:
ingestFieldKVP does still store every field value as a STRING (coerce.go's
BUG-2850 note, and the github_pr hint in dispatch_http.go). This change
stopped it TRIMMING, not stringifying.

Claude-Session: https://claude.ai/code/session_01HeChkgZVYb3NTgTcckF5KR

* docs: sync CLAUDE.md to tool surface 0.30 (BUG-2870)

The drift guards cover instructions.md and README.md but not this file, and
its own 0.27 entry records the consequence: 'This entry was missing from
CLAUDE.md — the 0.27 unit swept instructions.md and README.md and not this
file.' The unit that makes a version line stale is the unit that owes it.

Both markers updated, and the entry states the two behaviour changes in the
terms they were ruled: /mcp refuses what it silently accepted, and the
swallowed parseFieldArray refusal that was landing four refusals as
successes on the no-fields path.

Claude-Session: https://claude.ai/code/session_01HeChkgZVYb3NTgTcckF5KR

* fix(mcp): finish the removal, and correct a claim I made twice (BUG-2870)

Codex round 1: no P1/P2, two nits, both real.

1. The re-emission removal was incomplete. `reEmitFields` and the branch
   that appended its entries survived with nothing populating the map, and
   two comments still described canonical re-emission as something this code
   does. Unreachable, but my own commit message had said the path was
   removed, so the code contradicted the claim. Removed, and the round-16/17
   paragraphs that decided WHEN to canonicalize go with it — they answered a
   question that no longer arises.

2. "The only behaviour change is /mcp refusing what it silently accepted" is
   WRONG, and it was in version.go, README.md and CLAUDE.md. Every door
   refuses a padded key now; they were merely accepting it differently —
   /mcp trimmed it and wrote the declared field, the CLI stored a ghost field
   beside it. What is /mcp-only is the VALUE half. Corrected in all three,
   with the correction itself recorded in the version.go entry so the next
   reader sees the claim was checked rather than a sentence that quietly
   changed shape.

Claude-Session: https://claude.ai/code/session_01HeChkgZVYb3NTgTcckF5KR

* docs(mcp): rename the predicate to the question it asks (BUG-2870)

Codex round 2: no P1/P2, three nits, all naming and prose.

- `canonicalized` is renamed `coveredByFieldsObject`. Nothing canonicalizes
  anything any more, and the only thing that predicate ever asked was
  whether the `fields` object carries THIS key — it kept the old name only
  because the guard it used to feed had been removed a commit earlier.
- parseFieldKVP's doc said invalid entries are skipped silently. True of a
  MALFORMED entry, false of a padded key, which now aborts the call.
- Three test comments still described re-emission as live, and version.go
  described this door's trimming in the present tense.

Nothing in these two rounds was a defect in the change itself; both rounds
found prose describing a version of the code that stopped existing partway
through the unit, which is the failure mode a re-grounding pass invites.

Claude-Session: https://claude.ai/code/session_01HeChkgZVYb3NTgTcckF5KR

* docs(mcp): last of the prose that outlived the code (BUG-2870)

Codex round 3: no P1/P2, prose only.

- the predicate's own comment still asked 'will anything canonicalize THIS
  key'; it asks whether the fields object carries the key, and always did;
- two test comments described re-emission and trimmed comparison as current.
  Both tests are kept — what they pin is narrower now and still worth
  pinning — with the change in what they mean written down.

Deliberately NOT changed: the comments and replacement-test names that cite
the OLD test names. Codex reads them as stale terminology; they are the
traceability the restatement commit was asked for, so a reader can find what
each replacement replaced.

Claude-Session: https://claude.ai/code/session_01HeChkgZVYb3NTgTcckF5KR

* docs(items): the coercion note names what the door does now (BUG-2870)

Codex round 4. The paragraph described ingestFieldKVP as doing
`dst[key] = val` unconditionally. Its CLAIM — every value arrives at the
server as a string — is still true and is the reason this file exists; the
description of the line is not, since that door now parses through
items.SplitFieldEntry. Restated so the still-true part is not carried by a
sentence a reader can falsify.

Claude-Session: https://claude.ai/code/session_01HeChkgZVYb3NTgTcckF5KR
2026-09-07 15:33:05 -04:00
xarmian 5d29b7715c fix(cli): item show --format markdown emits the body verbatim (IDEA-2937) (#1281)
* fix(cli): item show --format markdown emits the body verbatim (IDEA-2937)

`pad item show <ref> --format markdown` printed the body with fmt.Println,
adding a trailing newline the stored body did not have. The write half,
`pad item update <ref> --stdin`, sends what it is given and the server
stores it verbatim (measured live: a body with no trailing newline and one
with three, both stored exactly as sent). So the read-modify-write shape
every agent reaches for was not a fixed point — writing back what `show`
emitted appended one newline per cycle, without bound: 518, 519, 520, 521,
522 bytes on the reporter's fixture, reproduced here as 18 → 22 over four
cycles. With the fix, four cycles hold at 18 bytes with an identical sha.

Scope, stated because the trail records a loss as well as a growth: the
`$(pad item show ... --format markdown)` capture that LOSES a byte is NOT
this defect and is NOT fixed here. Command substitution strips every
trailing newline from whatever it captures — `C=$(cat file)` on the same
18-byte file yields 17 too — so a body ending in a newline loses that
newline through `$()` before and after this change. Tools that need a
lossless read must redirect or pipe, not capture. Whether the server should
normalize a stored body to end with exactly one newline (which would make
even the `$()` path a fixed point, at the cost of destroying deliberate
blank lines at the end of a body) is a separate question, left on the trail.

Tests assert equality with the body, not `Contains`, on both halves: what
`show` emits, and what `--stdin` sends. Both were run against the unfixed
line and fail there — restoring the Println kills all five show subtests,
and a TrimRight on the stdin read kills three update subtests.

Claude-Session: https://claude.ai/code/session_01HeChkgZVYb3NTgTcckF5KR

* fix(cli): close the test stdin pipe; the comment overstated what was measured

Codex round 1 returned CLEAN with two remarks, both real:

- feedStdin restored os.Stdin but never closed the read end, leaking one
  file descriptor per subtest.
- The code comment said "Both directions were this one line", claiming the
  `$(...)` byte LOSS as this defect too. It is not: command substitution
  strips every trailing newline from whatever it captures — `C=$(cat file)`
  on the same 18-byte file yields 17 — so that loss is identical before and
  after this change. The commit message was already corrected before the
  review; the comment was not, and a comment is what the next reader has.

`go test ./cmd/pad/ -race` exit 0 after both.

Claude-Session: https://claude.ai/code/session_01HeChkgZVYb3NTgTcckF5KR

* test(cli): captureStdout closes its pipe read end

Codex round 2 nit. The helper predates this branch, but the new round-trip
tests call it ten times, so the leak is ten descriptors per run rather than
a few. Closing after ReadFrom completes the cleanup the helper already
started with w.Close().

Claude-Session: https://claude.ai/code/session_01HeChkgZVYb3NTgTcckF5KR
2026-09-07 14:10:07 -04:00
xarmian bb8ec04ef1 fix(server,store): both workspace mint doors enforce their preconditions from one place (BUG-2809) (#1268)
handleCreateWorkspace and handleImportWorkspace mint the same thing
through the same store.CreateWorkspace, and enforced preconditions in two
places. Two had already diverged and been fixed one at a time, each found
by a reviewer rather than by the door that lacked it: the OAuth consent
grant (IDEA-2756) and the user-scoped plan limit (BUG-2793). A third was
live.

The shared place is internal/server/workspace_mint.go, split by WHEN a
precondition can run, and the split is load-bearing rather than tidy:

  beginWorkspaceMint  — everything that does not need the body (consent,
    plan limit, and the owner/source attributions). Runs before the body
    read, so a refused caller never uploads a bundle and a refusal cannot
    be probed by body shape; on the import route it sits above the
    Content-Type dispatch, so one line covers both body shapes.

  validateWorkspaceMintPayload — the payload-shaped rules. Returns an
    error rather than writing one, because the JSON doors answer 400
    bad_request and the bundle door answers 400 bad_bundle through
    importStatusError. The rule is shared; the envelope stays each door's.

Callers: handleCreateWorkspace, handleImportWorkspace, and importBundle.
The mint context reaches the bundle path as an ARGUMENT rather than on the
Server, because it is per-request state and the two things it carries are
exactly what two concurrent requests would differ on.

THE LIVE DEFECT. Import accepted an empty workspace name. Measured before
the fix: it created a workspace with name="" and slug="", and a second
such import landed on slug "-2" -- the first had taken the empty slug,
globally, and a slug is a routing key. Both import doors now refuse it,
checking the EFFECTIVE name (the ?name= override when given, the bundle's
own otherwise) because that is what becomes the slug. A control leg covers
the override, or the rule would be indistinguishable from "reject any
bundle whose payload name is empty" and would break rename-on-import.

SETTINGS: the item's premise was wrong and this corrects it rather than
fixing it. Malformed settings never reached the store unnormalized --
createWorkspaceQ calls NormalizeWorkspaceSettings itself and refuses. What
diverged was the STATUS: create answers 400, import answered 500
import_failed because handleImportWorkspace maps every store error that
way. Validating in the shared payload step makes both 400. Context stays
create-only: an export carries none, so applying it on import would be
inventing input.

SOURCE: imported workspaces got no attribution at all (BUG-1557).
store.ImportWorkspace now takes a source parameter, derived by the caller
from the request's auth shape exactly as create derives it -- a parameter
rather than an export field, because a bundle says what the workspace WAS
and where this copy is minted from is a fact about this request. The
operator path (pad db migrate-to-pg) passes "": it is a copy, not a
creation surface, and inventing "cli" would relabel every migrated
workspace's origin.

Userless callers (the inventory's fourth item) are deliberately unchanged.
beginWorkspaceMint preserves the userID != "" guard exactly as both doors
had it rather than changing behaviour under cover of a refactor; the
measurement and the ruling are on BUG-2914.

Five mutants, each verified to COMPILE first and each detected by its own
leg: either import door skipping the payload check, the create door
skipping it, checking the payload name instead of the effective name, and
passing "" for source. Two of them initially did not compile, and go test
answers a build failure with FAIL <pkg> [build failed], which in a
filtered run reads exactly like detection -- a false DETECTED, the mirror
of the false SURVIVED. Re-run with the orphaned variable kept alive.

Claude-Session: https://claude.ai/code/session_01HeChkgZVYb3NTgTcckF5KR
2026-09-06 23:36:58 -04:00
xarmian 7659ad3cd3 feat(server,web,cli): say when a relation's copy target is unusable, instead of offering a picker that cannot answer (IDEA-2899) (#1262)
* feat(server): the copy preflight says when a relation's target is not usable (IDEA-2899)

TASK-2869 made a `needs_value` relation row collectable as soon as it names
a target collection. Naming one is not having one: the slug can name a
collection that has been DELETED, or one this caller cannot READ. The dialog
then mounts a picker that can return nothing and, because the row is not
blocked, Confirm stays disabled carrying only the generic required-field
message — the user is told a value is missing and never told that no value is
reachable.

`collection_unavailable` on the needs_value row is the server saying so.

THE CLIENT CANNOT COMPUTE THIS, which is why it belongs here. The dialog's
destination collection list is filtered through `canEditCollection`, because
it drives the copy-INTO picker; a relation TARGET needs only READ access, so
a perfectly usable target routinely does not appear in that list. Testing
against it would refuse rows the user could have filled in — over-blocking,
which is the worse failure and invisible to whoever hits it.

`visibleCollectionIDs` is the read-scoped view, and its NAV-LENIENT shape is
right here rather than merely tolerable: it includes a collection reachable
only through an item-level grant, and the question is "could a picker here
return anything at all". One granted item is a picker with one row.

DELETED and UNREADABLE are deliberately not distinguished. Same consequence,
no client branch would differ — and separating them would tell a caller who
cannot read a collection that it nonetheless exists.

`omitempty` on a BOOL drops `false`, so the field is phrased NEGATIVELY.
Present-and-true means the server checked and the target is unusable; ABSENT
means available, or a server that does not report. A client must block only
on an explicit true, so absence stays "no information" rather than becoming a
value — the rule `access_epoch` follows on the item doors, and the one whose
violation cost two review rounds on IDEA-2898 this morning.

Costs nothing on the common path: a destination schema declaring no relation
field runs no query at all.

Claude-Session: https://claude.ai/code/session_01Xk9M5UVPdc84xL5E1mZkm8

* test(server): pin the type gate on collection_unavailable (IDEA-2899)

Found by a surviving mutant rather than by inspection: dropping the
`def.Type == "relation"` gate left every other test in the file green.

Nothing stops a schema declaring `collection` on a field of another type — the
validator does not police keys it has no use for — and such a field would then
pick up a flag whose meaning is defined only for relations. The dialog would
block a perfectly collectable `select` because some relation elsewhere in the
same schema points at a collection that happens to be gone.

The fixture is the discriminating one: ONE deleted collection, TWO required
rows that name it, and only one of them means anything by it.

Six mutants on this half, all killed: flag never set, flag always set, deleted
target not flagged, unreadable target not flagged, type gate dropped, and the
nil-visible-set case (an admin's "no filtering" read as "nothing visible",
which would flag every target for the callers who can see everything).

Claude-Session: https://claude.ai/code/session_01Xk9M5UVPdc84xL5E1mZkm8

* feat(web): block a relation whose target is unavailable, and stop advising a command that cannot work (IDEA-2899)

The client half. `isCollectable` now refuses a relation row the server has
flagged, so the row lands in `blockedFields`, Confirm is disabled with a
reason, and no picker mounts that could only come back empty.

`collection_unavailable !== true` is STRICT on purpose. The field is absent
when the target is fine and absent from a server that predates it, so absence
must read as "no information". (Over the domain the type admits — `boolean |
undefined` — the truthiness spelling is EQUIVALENT and a mutant swapping it in
survives; that is recorded in the source rather than papered over with an
off-contract fixture. The strict form is kept because it states the contract
where the next edit will read it, and the inverse spelling would block every
row against an older server.)

THE PART THAT IS NOT WIRING: the existing blocked-field notice said the field
"is a required <type> field. This dialog can't collect a value for that type
safely" and then printed `pad item copy … --field key=value`. Both halves are
FALSE here. The type is perfectly collectable; the TARGET is gone. And the CLI
runs as the same user against the same referent validation, so the command it
prints is refused for exactly the reason the user is already stuck — advice
that sends someone to do work that cannot succeed is worse than no advice.

So the message branches on `uncollectableReason`, names the collection and the
destination workspace, and the CLI line is now gated on `cliFillableField` —
the first blocked row the CLI can ACTUALLY fill. `blockedFields[0]` was
correct while every blocked row was type-shaped; with an unavailable relation
sorted first it named the one field `--field` cannot set either.

Eleven unit tests on `copyNeedsValue`, plus a source pin on the dialog whose
own measured limit is in its docblock. Client mutants: 7 real, 6 killed, 1
recorded as equivalent with the domain argument that makes it equivalent.

Claude-Session: https://claude.ai/code/session_01Xk9M5UVPdc84xL5E1mZkm8

* feat(cli): the copy preview marks an unavailable relation target and stops suggesting it (IDEA-2899)

Caught by `TestItemCopyMirrorsMatchServerShapes`, not by me. The CLI keeps a
mirror of the preflight response, and adding a field server-side without
mirroring it fails that test by design — a mirror that silently lags is a
mirror that lies. Working exactly as intended, and the reason this half exists
at all.

Mirroring the field turned out to be the smaller part. The CLI already prints
`target collection: people` for a relation row, and it builds an
`Add: --field owner_ref=<value>` suggestion from every unsupplied row. Both
are wrong when the target is unavailable: the first sends a user looking for a
ref in a collection they cannot read, and the second hands them a command the
referent validation refuses for exactly the reason they are already stuck.

So the target line is marked NOT AVAILABLE, and the row is excluded from the
suggestion with a sentence saying why — modelled on the empty-key branch,
which was written for the identical reason (a `--field =<value>` nobody can
run) and is three lines away.

That the same defect had to be fixed in two places is the shape worth naming:
the dialog and the CLI independently built "here is how to supply it" from
"here is a field needing a value", and neither had a notion of a field that
CANNOT be supplied. The empty-key case was the first instance and was fixed
locally; this is the second.

Five mutants on this half, all killed: suppression removed, suppression
applied to everything, the unavailable label dropped, the explanation dropped,
and the mirror field ignored. The available-target control leg is a separate
test so the omitempty contract is exercised on this surface too.

Claude-Session: https://claude.ai/code/session_01Xk9M5UVPdc84xL5E1mZkm8

* fix(cli): route all three "how to supply it" sites through one predicate (IDEA-2899)

Review found the fix applied at one door and not its siblings — my own
recurring shape, arriving again.

THREE places tell a CLI user how to resolve an unsatisfied field: the detailed
`renderItemCopyNeedsValue`, the `--dry-run` summary, and the error the command
returns. The first commit fixed the render. The other two went on printing
`--field key=value` at someone for whom no value exists — and the ERROR is the
line a script or a hurried reader actually sees, so it was the worst of the
three to leave.

`itemCopyUnfillable` is now the single definition all three consult. Not
because three call sites are tidier than one, but because three sites
independently answering "how do I supply this" is exactly how they diverged in
the first place.

The dry-run summary branches three ways rather than two, because the MIXED
case is the one a boolean gets wrong: some fields can be supplied and some
cannot, and collapsing that either suppresses advice the user needs or offers
advice they cannot use. The error hint is suppressed only when NO field can be
supplied — with one fillable field left, `--field key=value` is still true.

Also pins the BOUNDARY the same review probed: a target collection that is live
and readable but EMPTY is deliberately not flagged. The symptom looks
identical — an empty picker — but the cases differ where it matters. An
unavailable target is unfixable from inside the dialog, so blocking costs the
user nothing they had; an empty collection is resolved by creating the item and
retrying, and blocking would refuse a copy they were about to complete. It
would also cost a live-visible-item count per relation target on a dry run the
UI calls on every keystroke. The weaker case — an empty picker that says
nothing about WHY — is filed as IDEA-2905 and belongs to the picker.

Ten mutants across this round, all killed, including both directions on the
error hint and both directions on the dry-run branch.

Claude-Session: https://claude.ai/code/session_01Xk9M5UVPdc84xL5E1mZkm8

* fix: unfillable means EITHER reason, and a select never names a relation target (IDEA-2899)

Review round 2, two findings, both real and both about a rule stated in one
place and enforced in another.

**"Unfillable" answered for one of two reasons.** An EMPTY KEY cannot be
supplied either — `--field =value` is rejected by this command's own parser,
and the detailed render has explained that since Codex round 6. Only that
render knew: the --dry-run summary and the returned error went on advising
`--field` for those rows, because the predicate I extracted last commit covered
the relation reason alone. A predicate named "unfillable" that answers for half
its name is a worse trap than no predicate — right at the site that defined it,
wrong everywhere it was reused, which is precisely what extracting it was meant
to prevent.

Two functions now: `itemCopyUnfillable` (either reason — advice), and
`itemCopyUnavailableTarget` (the relation half — the render's own sentence,
since the two explanations are not interchangeable to a reader).

`itemCopyUnavailableTarget` deliberately does NOT also exclude empty keys,
though my first version did. A row can carry both faults, and a mutant removing
that exclusion survived every test — correctly, because all it changes is
printing two sentences that are both TRUE about such a row. The guard was
tidiness dressed as a rule; a condition nothing can distinguish is one the next
reader has to re-derive.

**`Collection` was emitted for non-relation fields**, while its own doc said it
is empty for every other type. That was a claim about the schemas people write,
not a property of the code: a `select` carrying `"collection": "people"` is
storable — field validation has no use for the key and does not police it — and
the value was copied straight through, so the CLI printed "target collection:
people" beneath a select. A relation fact asserted about a field that has none.
`relationTargetSlug` makes the documented contract true at the only place that
can make it true; my own type-gate test had created exactly that shape and
asserted only the FLAG, not the slug.

Three mutants on these fixes, all killed.

Claude-Session: https://claude.ai/code/session_01Xk9M5UVPdc84xL5E1mZkm8

* fix(cli): the explanation now names the reason that actually applies (IDEA-2899)

Review round 3, and the sharpest miss of this unit — my own, one commit old.

Broadening what a predicate ACTS on silently broadened what a sentence SAYS.
Once `itemCopyUnfillable` counted empty keys as well as unavailable relation
targets, a set of empty-key rows selected the all-unfillable branch and was
explained as "the relation target is not available to you" — a false statement
about rows that contain no relation at all. Same in the returned error, which
is the line a script sees.

The tell was there to be read: a sentence that was TRUE while the predicate was
narrower is a sentence to re-read the moment it widens. I broadened the
predicate deliberately, wrote a commit message about how a half-answering
predicate is a trap, and left the sentence describing the half.

`itemCopyUnfillableWhy` names the reasons actually present — relation targets,
empty keys, or both — and the two one-sentence sites consult it. The detailed
render is unchanged: it explains each reason where the row is printed, which is
why it uses the narrower count.

Four mutants, all killed, including the two that matter: the explanation always
saying "relation" (the defect) and never saying it (the same defect pointing the
other way). The test carries a mixed-reason leg, because a sentence that picks
one of two true reasons is the failure a single-reason fixture cannot see.

Also corrected: three comments claiming `itemCopyUnfillable` is relation-only
or that the detailed render consults it. Both stopped being true last commit.

Claude-Session: https://claude.ai/code/session_01Xk9M5UVPdc84xL5E1mZkm8

* fix: one row can carry both faults, and four docs said this was simpler than it is (IDEA-2899)

Review round 4. Four findings, no P1s, and the first is the one worth the round.

**A `continue` between the two counts.** `itemCopyUnfillableWhy` counted a row
as an unavailable relation target and then skipped the empty-key check, so ONE
row carrying both faults reported only the first. My mixed-case test used TWO
rows with one fault each — a different input, and the only one it exercised.
Two rows with one fault each and one row with two are not the same fixture, and
I built the weaker one while writing a commit message about fixtures that
cannot discriminate.

**The dialog could still print `--field =value`.** `cliFillableField` excluded
unavailable relation targets and not empty keys, so a required `json` field the
destination reported with no key was type-shaped, blocked, and still offered a
command the CLI's own parser rejects. The CLI has refused those since Codex
round 6; the web side had never learned it. Same defect, other surface —
which is the third time this unit has fixed one door and not its sibling.

**Cardinality.** "no --field can supply it" for several fields, and "reported
them with an empty key" for one. Both sites now agree with their counts, and
the empty-key phrase is neutral on number so it reads correctly after either.

**Four documents claimed every needs_value row is resolvable with an override**
— the CLI renderer's docblock, the server's `NeedsValue` field, the CLI mirror
type, and the dialog's collectability comment. That was true when each was
written and this unit falsified all four; a reader following any of them would
conclude the CLI had simply forgotten to print a flag.

Two mutants on the fixes, both killed: the `continue` restored, and the
dialog's empty-key exclusion removed.

Claude-Session: https://claude.ai/code/session_01Xk9M5UVPdc84xL5E1mZkm8

* refactor(cli): one tally, because the rounds said the branching was the problem (IDEA-2899)

Four review rounds returned 2, 2, 2 and 4 findings. The counts looked like slow
convergence; the DISTRIBUTION was the finding. Every defect after round 1 lived
in this one layer — how the CLI and the dialog say "here is how to supply it" —
while the server half that computes availability stayed clean throughout.

The layer had accreted exactly the way IDEA-2898's cold path did: a count, then
a second count for the other reason, then a phrase function, then a `continue`
between two counters that made a dual-fault row report half of itself. Round 4
fixed something round 3 introduced to fix something round 2 introduced. That is
not a run of bad luck, it is a shape.

So this round removes branches instead of adding a seventh guard.
`itemCopyTally` walks the rows once and returns what every caller needs;
`AllUnfillable()` is the condition both one-sentence sites test, and `Why()` is
the phrase both interpolate. Three helpers become one type. There is no second
definition of "unfillable" to drift from the first, and no sentence describing
a subset of what a predicate counts, because the sentence and the count come
from the same walk.

`Unfillable` is deliberately NOT `UnavailableTarget + EmptyKey`: one row can
carry both, and double-counting makes `Unfillable == Total` false for a set
that is entirely unfillable — the comparison every caller makes. A mutant does
the addition and dies.

Five mutants, all killed. The last needed a new test rather than a new fixture:
`AllUnfillable`'s `Total > 0` guard is unreachable from both current callers,
so a mutant removing it survived every command-level test. Keeping an
unreachable guard and calling it defence is how a promise becomes a lie, so the
tally is now unit-tested directly — an empty set is not "entirely unfillable",
and a future caller outside the `len() > 0` gate would otherwise be told
silently that nothing can be supplied.

Claude-Session: https://claude.ai/code/session_01Xk9M5UVPdc84xL5E1mZkm8
2026-09-06 15:48:22 -04:00
xarmian 07b2e439f2 TASK-2869 (U2b): the preflight names a relation's target collection, and the copy dialog scopes its picker to the destination (#1258)
* feat: the preflight's needs-value row names its relation target, and the copy dialog scopes its picker to the destination (TASK-2869)

U2b, per the day-55 ruling. Both blockers (U1 referent validation, U2 the
FieldEditor branch + picker) are in.

THE DEFECT. `ItemCopyPreflightNeedsValue` carried `type: "relation"` and no
target. FieldEditor gates its relation branch on `wsSlug` AND
`field.collection`, so a required relation in the destination reached the copy
dialog as a field it knew was a relation with no idea what to point at, and
rendered as FREE TEXT. Before U1 the copy stored whatever was typed.

SERVER. `collection` is added to the needs-value row, populated from the
DESTINATION schema's `def.Collection` — the only place it is known, since the
row is built from that schema. Additive and `omitempty`: a client that does not
read it is unaffected, and a row for a non-relation field is byte-identical to
before. Not a wire-version question, for the same reason
`models.ItemWriteWarnings` was not.

CLIENT. `toFieldDef` carries the collection through, and the FieldEditor call
passes `wsSlug={destWs}` — the DESTINATION, never the source. A relation
resolves at the destination, so the picker must list items the copy can
actually point at; that is same-workspace resolution AT the destination, not
the cross-workspace case PLAN-2857 rules out.

`relation` becomes collectable ONLY IF THE ROW NAMES ITS TARGET. Without a
collection, FieldEditor's gate renders the non-editable state, so offering the
row would produce a control that cannot be filled and a Confirm that cannot be
satisfied. Such a row now lands in the blocked list and the user is told which
field and why — the same disposition `multi_select` gets, for the same reason:
a control that silently cannot do its job is worse than an honest refusal.

TWO THINGS THIS UNIT TAUGHT ME THAT ARE NOT IN THE RULING.

1. THE WEB MUTANT SURVIVED, AND THAT IS WHY `isCollectable` MOVED. My first
   version left the predicate inline in `CopyItemDialog.svelte`. Making
   `relation` unconditionally collectable — the exact defect the negative leg
   of the proving test is about — passed EVERY suite in the repo. That is
   IDEA-2894's lesson arriving one unit later in the same file, so the
   predicate now lives in `$lib/items/copyNeedsValue` with tests. Two mutants
   die there: relation-always-collectable, and `multi_select` slipped into the
   collectable set.

   The Go half was pinned from the start (drop `Collection: def.Collection` ->
   FAIL naming the empty value and the expected slug). Only the client half was
   unpinned, and only because of where the code lived.

2. A U2-ERA TEST ASSERTED THE ABSENCE THIS UNIT CLOSES, and asserted it
   CORRECTLY. `fieldEditorRelationCallers.test.ts` required that the dialog
   pass no `wsSlug` and build its FieldDef from a shape with no `collection` —
   which was the behaviour, and withholding `wsSlug` was what kept an unscoped
   picker out. It also named this task by ref and told its successor to revisit
   the gate WITH the change rather than let it drift. Inverted here: the block
   now asserts the destination slug is passed, that the collection reaches the
   FieldDef, and — the half that is easy to lose — that the dialog still
   DELEGATES the collectability decision, so a future inlined predicate would
   pass the unit tests and fail this.

   Worth keeping: a test that pins a temporary absence should name what would
   make it wrong. This one did, and that is the only reason its inversion was a
   five-minute job instead of an argument about whether it was load-bearing.

Gates: `internal/server` ok 158.335s, `go vet` and `gofmt` clean,
`npm run check` 0 errors (6 pre-existing warnings), `make web-test` 127 files /
2123 tests. Postgres and CI are owed on this tip.

Claude-Session: https://claude.ai/code/session_01Xk9M5UVPdc84xL5E1mZkm8

* fix(cli): mirror the needs-value collection, and name the relation target in the CLI (TASK-2869)

TWO CONSUMERS I DID NOT SWEEP. The previous commit added `collection` to the
preflight's needs-value row and updated the server struct and the TypeScript
type — "both sides", as its own message put it. There are THREE sides.
`internal/cli` keeps a mirror of the preflight shape and
`TestItemCopyMirrorsMatchServerShapes` requires it to match the server field
for field. It failed in the Postgres gate, in a package the change did not
touch.

That is the third time this session a producer change broke a consumer I had
not enumerated, and the shape is always the same: I name the surfaces I edited
and call that the population. The instrument that would have caught it is not
"run more tests" but "grep for the type's name before claiming the sweep is
done" — `ItemCopyPreflightNeedsValue` appears in exactly three files and I
looked at two.

Mirrored, with a comment saying why it exists: a mirror that silently lags is a
mirror that lies, and the CLI renders these rows.

AND THE CLI NOW NAMES THE TARGET COLLECTION, which is the point of the unit on
the surface that has no picker at all. A row reading

    owner_ref            (Owner, relation) required — required, with no value…

tells a user a value is needed and nothing about what kind of value exists.
The dialog answers that with a scoped picker; the CLI had no answer. It now
prints the relation analogue of the `options:` line a select already gets:

    owner_ref            (Owner, relation) required — …
                           target collection: people

Test asserts the line appears for the relation row, appears EXACTLY ONCE with a
select row rendered alongside — so it cannot pass by printing unconditionally —
and that the select's own `options:` line still renders, so this did not
displace it.

Gates: `internal/cli` and `cmd/pad` green, build and gofmt clean. The full
Postgres run and CI are owed on this tip; the earlier PG run is the one that
caught the mirror and is superseded.

Claude-Session: https://claude.ai/code/session_01Xk9M5UVPdc84xL5E1mZkm8

* fix(web): the relation picker searches the preflight's canonical destination slug (TASK-2869)

Codex review, finding 3 of 3, and the only one of the three that belongs in
this unit.

`wsSlug={destWs}` handed `FieldEditor` a value that is NOT always a workspace
slug. An item can be opened through a workspace-UUID URL, the route parameter
is passed straight through as `sourceWsSlug`, and a same-workspace copy then
puts that UUID in `destWs`. `/search` resolves a workspace by SLUG only, so a
picker handed a UUID searches nothing and returns no results — a control that
looks usable, is not, and says nothing about why.

Now `pickerWsSlug`, which is the preflight response's own
`destination.workspace_slug`. The preflight IS the canonicalising round-trip:
the server resolved whatever it was given and answered with the real slug.
Falls back to `destWs` only before the first preflight returns, at which point
no needs-value row is rendered anyway.

The caller test asserts the prop AND the derivation, because asserting only the
prop would pass against a `pickerWsSlug` that was just `destWs` renamed.

THE OTHER TWO FINDINGS ARE REAL AND ARE FILED, NOT FIXED HERE.

IDEA-2898 — `ItemPicker` serves warm local-index results without re-authorising
them, so a collection whose access was revoked can still be listed. The cold
`/search` path is visibility-filtered and correct; the warm path is not. This
is PRE-EXISTING and applies to every caller of the picker, `ItemDetail`
included — last touched by TASK-2877, not by this unit. U2b widened the
exposure by adding a caller; it did not create the defect, and rewriting the
picker's cache-authorisation model inside a feature branch would be an
unrelated change riding along. The fix needs a decision about where the client
learns its access set from, which no current signal provides.

IDEA-2899 — a relation row can name a target collection that is DELETED or
UNREADABLE, so `isCollectable` says yes on a non-empty string and the user
meets a picker with nothing in it. I could not fix this correctly here, and the
reason is measured rather than assumed: the obvious test is to check the target
against `destCollections`, and that list is filtered by `canEditCollection` —
it is what the user may copy INTO. A relation TARGET needs only READ access, so
a perfectly usable target routinely is not in it. Using it would OVER-BLOCK,
refusing rows the user could have filled, which is a worse failure than the one
being fixed and invisible to whoever hits it. The right shape is probably the
server reporting the target's availability on the row it already builds — an
additive field on the same row this unit just changed, worth doing deliberately
rather than bolted on at the end of a branch.

Gates: `npm run check` 0 errors (6 pre-existing warnings), `make web-test` 127
files / 2123 tests. Postgres is running on this tip; CI is owed.

Claude-Session: https://claude.ai/code/session_01Xk9M5UVPdc84xL5E1mZkm8
2026-09-05 21:37:44 -04:00
xarmian 02846a6785 fix(cli): the session's registered agent is the name its writes carry (BUG-2882) (#1248)
* fix(cli): the session's registered agent is the name its writes carry (BUG-2882)

Two seats booted under one name; one re-registered under the right one
with `pad session register --agent`; every write it made afterwards still
carried the wrong name. The registry row and $PAD_AGENT were two
self-declarations of "the same value" — the row could be rewritten, the
environment could not, and nothing reconciled them or warned. Night 10
read both live seats inverted against each other.

ResolveAgentName now consults the registry record for the session that
owns this process FIRST: a stat-and-read of one file, no MkdirAll, no
lock, ignored when malformed, legacy, or carrying a different
process-start token (pid reuse). A non-empty registered name wins over
.pad.toml and $PAD_AGENT; an anonymous row leaves an environment name in
force. `pad session register` without --agent keeps the current name, as
before, because its default is the resolver. Help text, the record's doc
comment and README's precedence list say what is now true.

Test: register as rook under PAD_AGENT=wren and a .pad.toml name → rook;
default re-register → rook; anonymous → the .pad.toml name; a record for
this pid with another process's start token → ignored. The registry-step
mutant fails the first two assertions.

Fixes BUG-2882

* fix(cli): a registry record names this session only when it is verifiably this session's, and the identity tests stop reading the real registry

Codex round 1 on #1248. (1) registeredAgentForThisSession compared
process-start tokens only when both sides had one, so a stale row with no
token under a reused pid — a dead session's — would have named a live
one. Fail closed: when this process can read a token, the record must
carry the same one; and the record must pass the same OwnerLiveness
verdict `pad session list` applies. The token-less case is now a test
row, with a positive control after it. (2) The pre-existing resolver and
header tests cleared PAD_AGENT/CLAUDECODE but not HOME or the session-pid
variables, so run inside a registered seat they would have read that
seat's row as the resolver's first answer. They now run from a scratch
HOME with no session identity.

Refs BUG-2882

* fix(cli): a registry row names this session only if its owner is this process or an ancestor, where that can be checked

Codex round 2 on #1248. (1) TestPushItemSendsResolvedAgentHeader was the
one identity test round 1's hermeticity fix missed; it now isolates HOME
and the session env like its neighbours. (2) The registry step accepted a
row whose owner pid was alive and token-matched but NOT this process or
an ancestor — a misconfigured CLAUDE_PID pointing at a sibling session
could borrow that session's name. Refused where the platform can walk
ancestry. Not gated on PIDVerified: CaptureSessionOwner records
"cannot check" and "checked and wrong" as the same false, and the flag
alone would have disabled the step on every non-Linux platform. The
check is re-run and only a checked-and-wrong answer refuses. Test uses a
live non-ancestor child as the claimed owner; the refusal-dropped mutant
lets it name us.

Refs BUG-2882

* docs(cli): session list help — a row's name outranks PAD_AGENT; change it with --agent, not by re-registering (BUG-2882, codex round 3)
2026-09-04 12:50:50 -04:00
xarmian b437cc582d feat: item reminders — the fire-at-an-instant primitive, and one overdue rule for all four surfaces (IDEA-2641, closes #1010) (#1244)
* feat(store): item reminders — the fire-at-an-instant primitive (IDEA-2641)

Adds the storage, the scheduler tick, and the canonical event for one-shot
item reminders (GitHub #1010). Nothing in Pad acted at a target time before
this: a due_date makes an item show up as overdue once somebody asks the
dashboard, so "revisit TASK-X on the 1st" had to live in an external cron.

A TABLE, NOT A SCHEMA-FIELD ANNOTATION. The design sketch proposed marking
schema date fields with a `reminds: true` key on models.FieldDef; recon
overturned it. Such a key does not survive an ordinary collection edit, two
independent ways: the web editor destructures each field into an EditableField
and rebuilds a fresh definition key-by-key on save, so unknown keys are
dropped (`pattern` and `unique_scope` survive only because two lines were
hand-added for them), and models.CollectionSchema has fixed fields with no
catch-all, so any Go unmarshal+marshal round-trip strips unknown properties —
the hazard retargetRelationFieldsTx mutates raw JSON to avoid. Both failures
are silent and both disarm a whole collection's reminders at once. It is the
same defect class that moved traits out of the schema column in TASK-2657.

The table also gives the lifecycle a home. A reminder is armed, then fired,
then acknowledged, and a re-arm returns it to armed — per-reminder state a
field definition has nowhere to keep.

remind_at is an RFC3339 UTC instant, deliberately not a `date` schema value:
those admit both YYYY-MM-DD and full RFC3339 and are compared against the
SERVER'S LOCAL calendar day. A fire-at time cannot carry that ambiguity. The
remaining timezone question for due_date is filed separately.

Firing is one transaction per reminder carrying BOTH the fired_at write and
the outbox insert. That pairing is the point: a fired_at committed without its
event is a reminder that silently notifies nobody and can never be retried,
because the row has left the armed set; an event without fired_at fires every
tick forever. The UPDATE's own `fired_at IS NULL` predicate is the arbiter, so
two instances ticking at once produce exactly one winner.

item.reminder_due is admitted to the closed events/1 set as v1.2, with a new
PayloadReminder family and no SSE name. The subject is the REMINDER, not the
item: two reminders can be armed on one item, so an item-subject event could
not say which fired, and the reminder id is what an acknowledgement addresses.
A new payload family rather than reusing the item snapshot for the same
reason — a snapshot would validate and still not answer the only question the
event exists to answer. No SSE name in v1 because the poll surface is the
contract; adding one later is additive, removing one is not.

Ack is explicit and nothing else acks. An item reaching a terminal status
deliberately does NOT ack: that would make every status write a reminder
mutation, and it would silently consume a reminder set to fire after the work
was done.

* feat(server): reminder surfaces, and one shared overdue rule for all four

Second half of IDEA-2641: the HTTP surface, the scheduler tick's wiring, and
the fix for the finding that justified the unit — `ready` / `next` did no date
handling at all.

OVERDUE NOW HAS ONE IMPLEMENTATION. It used to live inline in the dashboard's
attention loop, which meant `pad project stale` inherited it (it filters that
very list) and the recommendation surface never saw it. So a deadline reached
the two surfaces that REPORT on work and never the one an agent PULLS from.
overdue.go is now the only place that decides, and all four call it.

Two behaviour changes fall out, both deliberate:

  - An overdue item bypasses the orphan branch's high/critical priority gate.
    That gate was where a deadline quietly stopped: a low-priority item three
    weeks late was reported by `stale` and never suggested by `next`.

  - Overdue sorts above in-progress. The list is capped at three, so a rank
    below in-progress would not merely order the deadline lower — on any
    workspace with three things in flight it would keep an overdue item off
    the surface entirely, which is indistinguishable from not shipping this.

The server-local-today comparison is UNCHANGED and known to be wrong for
multi-timezone deployments; it is filed as its own item with the cloud case
stated. Changing what "overdue" means on every existing instance inside a
change about where the rule LIVES is the kind of behaviour change nobody
reviews.

Fired reminders reach `next` / `ready` two ways, from one filtered list:
PendingReminders is the addressable form (it carries the id an ack needs), and
a prepended suggestion is the rendered form. They are prepended AFTER the cap
rather than entered as ranking candidates — a reminder is not a task competing
on priority, and whether it appeared should not depend on how busy the
workspace is.

Terminal-item reminders are FILTERED from the surface, never acked. Acking on
terminal status would couple every status write to reminder state and would
consume a reminder armed to fire after the work was done. The row stays
exactly as the user left it; the distinction is observable, and asserted.

Three guard tests caught this change and each was answered rather than
silenced:

  - The request-body reader guard was right: the handlers now go through
    decodeJSON, inheriting the NUL refusal and the size cap.

  - The canonical-events guard was right: item.reminder_due is admitted to the
    duplicated contract table as SPEC-3 v1.7, with the reminder subject kind
    and the new payload family. SPEC-3's own text owes the same amendment.

  - The NUL census asked for a decision on eight new columns. None carries
    caller text: ids and FKs are server-generated, four are the server clock,
    and remind_at is now re-parsed and re-formatted in the STORE as well as at
    the edge — so the stored value is always machine-produced from a parsed
    time and no caller bytes reach the column. The doc comment that used to
    say "the caller normalizes" protected nothing.

    Regenerating the baseline also found that GEN_NUL_BASELINE=1, which the
    test's own instructions name, was never implemented — the flag did
    nothing, so the documented path was hand-editing the file. Implemented, so
    the next reader gets the mechanism the instructions promise.

* test(reminders): the lifecycle, the four surfaces, and 22 killed mutants

Every test here was designed against a specific mutation and the mutation was
RUN. A green suite proves nothing about a suite nobody tried to break, and
three of the mutants I first wrote were not experiments at all.

Store (10 mutants, all killed): candidate predicate <= flipped to >=; the
event emission lifted out of the fire transaction; the fire UPDATE's
`fired_at IS NULL` arbiter removed; the RowsAffected check ignored; re-arm
clearing fired_at but not acked_at; ack losing `fired_at IS NOT NULL`; the
poll surface losing `acked_at IS NULL`; normalizeRemindAt no longer refusing;
it dropping .UTC(); GetReminder losing its workspace scope.

Surfaces (12, all killed): the priority gate no longer bypassing on overdue;
the sort no longer ranking overdue first; attention leaving the shared helper;
the reason losing its OVERDUE prefix; the comparison flipped to >; terminal
items no longer skipped; terminal reminders no longer filtered; the filter
ACKING instead of hiding; reminders appended instead of prepended; the tick
running on a far-future clock; ack answering 200 for an unfired reminder;
parseRemindAt accepting a bare date.

THREE MUTANTS DID NOT COUNT ON THE FIRST PASS and were rewritten. Two failed
to compile (`if false` orphaned a variable; deleting a parse orphaned an
import) and one had an anchor matching two call sites. A non-compiling mutant
emits zero FAIL lines and reads exactly like a surviving one — it invents a
hole that is not there — so the harness reports BUILD-FAIL and ANCHOR-BAD as
outcomes distinct from SURVIVED. It also restores files from an in-memory copy
rather than `git checkout`, which would delete uncommitted work in the tree.

ONE MUTANT GENUINELY SURVIVED and the test was at fault, not the mutant:
appending rather than prepending reminder suggestions was undetectable because
the fixture had a single item, so the reminder sat at index 0 either way. The
fixture now fills the three-item cap with in-progress work, where an appended
reminder lands fourth and vanishes. Faithful mutant, weak test — checked in
that order.

The same lesson shapes the four-surface fixture: it is a LOW-priority open
orphan, because that is the case the old code handled worst. A high-priority
task would have made the ready/next leg pass against the unfixed tree, which
is a green that measures nothing.

Negative controls throughout: a future deadline is not overdue and does not
reach the gate bypass; a tick with nothing due fires nothing; a completed item
is neither overdue nor suggested. Without them a helper that reported every
date, or a tick that fired everything, would satisfy every positive leg.

The lead's pin is asserted in both directions: a fired reminder on a done item
is ABSENT from the surface and PRESENT and still unacknowledged in the table.
Asserting only the absence would pass against an implementation that consumed
the row, which is the behaviour the pin exists to forbid.

* feat(mcp): pad_item.remind + ack-reminder, ToolSurfaceVersion 0.28

An agent that can RECEIVE a reminder but not set one has half the primitive.
The poll surface is pad_project.next / ready, both long exposed, so reminders
already reached agents — what was missing is the other half: deferring a piece
of work is exactly the moment an agent knows when it wants to be asked again,
and it had no way to say so.

Two additive actions, two optional params. Nothing existing moved, so a v0.27
consumer enumerating neither is unaffected — the v0.13 / v0.11 / v0.8
disposition, which likewise wired existing CLI verbs onto the catalog.

remind_at REFUSES a bare date rather than reading it as midnight. Worth
stating because the `date` schema type accepts YYYY-MM-DD and a caller will
reasonably try it here: a bare date names a 24-hour span, and choosing an hour
inside it would fire at a time nobody picked.

Re-arm and disarm stay CLI-only. Both address a reminder by an id the agent
would have to list first, and no listing action exists on this surface — a
door with no handle. Adding them later is additive.

Five guards had to be taught, and each was answered on its merits rather than
excluded: the HTTP parity test (route mappers added, so the actions work on
the remote transport rather than being advertised and unrouted), the
read-only catalog's cmdhelp fixture and expected cmdPath map, the field-
conflict classifier (remind_at / reminder_id are NOT field writers — a
reminder is a row in its own table addressed by its own id, so listing them
as classified sources would have pointed detectFieldConflicts at something
that is not a field source), and the instructions.md / README action tables.

That machinery is why the version bump is safe to make now, and it earned its
keep on this change: every one of the five failed on the first build after the
catalog entry landed.

CONVE-23 sweep for prose this falsifies:

  - SPEC-3 (DOC-2653) amended to v1.7 in the room, recording item.reminder_due
    with its new subject kind and payload family — the first canonical event
    with no user mutation behind it, since a scheduler tick produces it.
  - CLAUDE.md gains the reminder routes, the CLI verbs, and the v0.28 entry.
    It was also stale at 0.26 with NO v0.27 entry at all: the 0.27 unit swept
    instructions.md and README.md and missed this file. Both added.
  - skills/pad/SKILL.md gains the verbs and a routing entry, including the two
    things an agent will get wrong — the time is an instant, so ask for a time
    of day rather than picking one, and finishing the item does not
    acknowledge the reminder.

* fix(reminders): codex round 1 — four findings, all real, all with a pin

Round 1 found four defects and refuted none of them. Each fix carries a test
that fails against the code as it was, and each of those was mutation-checked.

**P1 — pending reminders bypassed item-level visibility.** Every other
dashboard section reads `allItems`, which the store already scoped to the
caller's collections AND their granted item ids. The pending-reminder list is
a direct workspace-wide query and inherited none of that, so a guest holding a
grant on ONE item could read the refs and titles of every other item in the
collection through its reminders — an item-level leak wearing a
notification's clothes. Now filtered with the same `isItemVisibleToGuest` call
the sibling sections use. The test's two items share a COLLECTION on purpose:
a collection-level filter was already applied, so separate collections would
have made it pass against the unfixed code.

**P1 — soft-deleted items could starve the queue permanently.** Candidate
selection ignored `deleted_at`, and `fireOneReminder` rolls back when it finds
the item gone — which leaves the reminder ARMED and therefore a candidate
again on the next pass. Candidates are ordered oldest-first and bounded by a
limit, so enough archived reminders fill every batch and no live reminder ever
fires. Silent, too: the tick reports zero fired and looks idle. Excluded in
the candidate query rather than skipped downstream, so those rows never occupy
a slot; the reminders themselves are kept, so restoring an item restores its
reminder with it — asserted, because a fix that reaped them would pass the
starvation test alone.

**P2 — the pass stopped at the first failing reminder.** The per-reminder
transaction exists precisely so one unfireable row cannot hold back the rest,
and `return fired, err` made that comment false — with candidates
oldest-first, one persistently broken old reminder blocks every newer one
forever. Now continues and joins the errors, so a pass that fired seven and
failed three reports both halves rather than reading as clean. The loop is
split behind an injected seam because a real mid-transaction failure is not
reachable from outside: the database refuses the corrupt rows that would cause
one (verified — invalid JSON in items.fields is rejected by the schema).

**P2 — suggestions dropped the reminder id.** The docs tell an agent to
acknowledge what it sees in next/ready, and the payload carried no handle: a
stateless poller could read the reminder and had no way to retire it, so it
would be shown the same item forever. `DashboardSuggestion` now carries
`reminder_id` (omitempty), `pad project next` prints the exact ack command,
and the test acks with the id the surface handed out rather than merely
checking the field is populated — a wrong-but-present id satisfies equality
with itself.

Four mutants, four killed; one was rewritten first because its anchor matched
two call sites and was therefore not an experiment.

* fix(reminders): codex round 2 — four findings, all real

**`--rearm` was unusable.** `ExactArgs(1)` forced an item ref that the rearm
branch then ignored, so the flag could not be reached without supplying a ref
that was silently discarded. Now `MaximumNArgs(1)`, with each mode checked
explicitly: a ref is required to arm, and a ref supplied ALONGSIDE `--rearm`
is refused rather than ignored — it names an item the reminder may not even
belong to, and quietly dropping it is how a user learns nothing about the
reminder they just moved.

**`unremind --format json` emitted plain text**, breaking the parseable-output
contract every sibling command honours.

**The MCP `ref` param did not list `remind`.** Agents read that flat
description to decide what to send, so an action missing from it is an
invalid call waiting to happen. It now also says what `ack-reminder` takes
instead, and why: a reminder is addressed by its own id because an item can
carry several.

**Fractional seconds fired early.** `time.Parse` accepts `09:00:00.900Z` and
`Format(RFC3339)` drops the fraction, so it was stored as `09:00:00Z` and
fired 900ms BEFORE the moment the caller named — silently, having rewritten
their value on the way in. Seconds are genuinely the stored resolution (the
column is compared as a string against a whole-second clock, and the tick runs
every 30s), so the only question was which way to resolve it, and truncation
resolved it the wrong way. `NormalizeInstant` now rounds UP: at most a second
of lateness, in exchange for a guarantee that can be stated — a reminder never
fires before the instant it was set for. Late is a reminder; early is a wrong
answer. Whole seconds round-trip exactly, which is asserted, because an
implementation that added a second unconditionally would otherwise pass.

Three mutants for this round, three killed (round-up→truncate,
round-up→unconditional-add, MaximumNArgs→ExactArgs). Thirty across the unit.

Two fixes carry no dedicated test and it is worth being explicit rather than
implying coverage: the `--format json` branch on `unremind` is a one-line
output change with no server-free way to drive it, and the MCP `ref`
description is prose the drift tests do not read — they assert an action is
DOCUMENTED, not that a param's sentence lists it.

* docs(reminders): the ack id is on the surface an agent polls, not only on the arm response

CONVE-23 follow-through on the round-1 fix. Both agent-facing docs told a
caller to acknowledge a reminder with the id "returned when you armed it" —
true, and useless to the caller that matters: a poller reading next/ready
never armed anything. The suggestion now carries reminder_id and `pad project
next` prints the exact ack command, so the docs say that instead.

The prose was written before the fix existed, which is exactly the case
CONVE-23 is about: a change that makes an instruction stale without touching
the file the instruction lives in.

* test(reminders): bind the tick LOOP to the work, not just the pass (CONVE-19)

Every other test in this file calls runReminderTick directly. That vouches for
the component and says nothing about whether anything ever calls it — a tick
that is never started is indistinguishable, from those tests, from one that
is. It is the convention's exact case, and the failure I recorded on my own
identity doc three times in one unit: I test the component and not the
binding.

Driven through the injectable tick channel so the assertion pins a SPECIFIC
pass instead of racing a 30-second ticker, and polled to a bounded deadline so
a loop that never runs FAILS rather than hanging the suite.

Mutant: drop `s.runReminderTick()` from the select and this goes red while
every direct-call test stays green. Killed.

The idempotence leg exists because a second Start spawning a second loop would
leave one running after Stop, making the BUG-842 drain invariant false for
this sweeper specifically — the one property a copied lifecycle is most likely
to get right by accident and least likely to be checked.

The cmd/pad call site (cmd_server.go, alongside StartTokenReaper) stays
verified by inspection: a source-scanning guard for it would be an instrument
asserting facts about source, which is code with an adversary and not worth it
for one line that sits in the middle of five identical neighbours.

* fix(reminders): codex round 3 — a deferred reminder fired anyway, and the poll surface was unbounded

**A re-arm mid-pass did not stop the fire.** The candidate scan selects an id;
before the UPDATE runs, a `--rearm` can move that reminder into the future.
Re-arm clears `fired_at`, so a predicate checking only `fired_at IS NULL`
still matched — the pass fired a reminder the user had just deferred and
emitted its event. The re-arm cannot undo that: it can clear the mark, but the
event is already on the outbox and at-least-once means a consumer has seen it.

The fire UPDATE now revalidates `remind_at <= nowTS` against the SAME nowTS the
candidate scan used. Same-value deliberately: the arbiter and the scan must
agree about when this pass is, or a reminder could pass one and fail the other
for no reason but clock drift inside a single pass.

**The poll surface was unbounded.** Every fired-and-unacknowledged reminder was
loaded and turned into a suggestion prepended to a list that is otherwise
capped at three, so a workspace with five hundred unacknowledged reminders
returned five hundred suggestions — in the dashboard response, the hottest read
in the product, growing until somebody acknowledged them.

Two bounds, because they are two different guarantees: the query takes a window
(default 50, oldest-fired first, so it holds what has waited longest), and the
prepended suggestions are capped at 5 so `suggested_next` stays a
recommendation rather than a second inbox. The full set stays addressable in
`pending_reminders`.

Truncation is REPORTED as a boolean, not a count. A count would have to be
post-visibility-filter to be true for the caller reading it, and the store
cannot compute that — the filter runs per item, above. "There are more than you
can see here" is the strongest claim the data supports, so it is the one made.

Four mutants; two killed outright, two survived and were run down under
CONVE-28:

- **Uncapped suggestions survived because the fixture had ONE reminder** —
  capped and uncapped are the same list at n=1. That is the SECOND time a
  single-item fixture hid a count-or-order property in this file. Fixture now
  arms eight; it also asserts all eight remain in `pending_reminders`, so the
  cap is pinned to the recommendation and not to the data.

- **Removing the SQL LIMIT survived, correctly, and the test comment now says
  so.** The Go slice cap bounds the PAYLOAD; the SQL LIMIT bounds the
  DATABASE'S work. Only the first is observable at this level — with the LIMIT
  gone the response is still bounded, while the query silently goes back to
  materialising every pending row before discarding most of them. That is a
  memory and I/O property with no assertion available here, so it is stated as
  a coverage boundary rather than papered over with a green that would not have
  measured it.

* docs(reminders): the fire predicate arbitrates against two actors, not one

CONVE-23 inside the file the round-3 fix touched. The comment described the
UPDATE as an arbiter for concurrent TICKS, which is what it was written for
and is why I did not re-read it when asked whether a user edit could race the
pass. It now says what it actually defends against, and names the general
shape: an arbiter is only an arbiter with respect to the writers it can see.

* fix(reminders): codex round 4 — the round-3 bound recreated the round-1 starvation

Round 3 bounded the poll surface. Round 4 caught what that bound did: the
query took the first N rows and the dashboard then discarded the ones it could
not show — hidden items, unauthorised items, completed items — so N such rows
hide a visible reminder behind them indefinitely, with no continuation to
reach it.

That is the SAME defect I had removed from the fire path one round earlier,
reintroduced in the read path within the hour. The general form is worth
stating because I clearly did not hold it: **a bounded window is only safe
when the discarding happens BEFORE the bound.** Filtering above a limit is a
starvation every time, and it does not matter what the filter is for.

Two halves, because the two filters are not the same kind of thing:

**Visibility is now scoped IN SQL**, using the same collection-id / item-id
sets every other dashboard section gets through `allItems` — the same
three-way shape as ItemListParams, where holding both collection grants and
item grants is an OR. Invisible rows no longer occupy the window at all, which
is strictly better than filtering them out afterwards and is what the sibling
sections have always done.

**Terminality is paged**, because SQL cannot evaluate it — a collection's
schema defines which statuses are terminal. The collector refills from the
next page when a page comes back short, bounded by a max scan so a workspace
full of completed items cannot turn a dashboard read into a table scan. The
bound is 10x the window: the common shape fills on the first page, and the
pathological shape terminates in a fixed number of indexed reads. Stopping at
the scan bound reports truncation, which is honest — there may be more, and we
did not look.

The empty-scope case is a THIRD state that reads like the second: nil
CollectionIDs means unrestricted, a non-nil EMPTY slice means this caller sees
no collections. Without an explicit guard they collapse, because the switch
matches none of its cases at length zero and adds no clause at all — so
"nothing visible" would return the whole workspace.

Three mutants, one survived: the empty-scope guard, because no dashboard-level
test produces that state (callers that would are refused earlier by workspace
access). Faithful mutant, missing test — it now has a direct one, with a
sanity leg so a build returning nothing cannot pass it by accident. A guard for
a state nothing exercises is exactly the one that rots.

* fix(reminders): codex round 5 — the MCP action I shipped did not work over stdio

**P1: local stdio MCP `remind` was unusable.** cmdhelp derives positionals by
regex from a command's `Use` string, and `<instant>` inside
`remind <ref> --remind-at <instant>` matched — it became a second REQUIRED
positional, so dispatch failed with `missing required argument "instant"`. The
action was advertised on a transport where it could not run.

**The MCP catalog's own tests did not catch it, and the reason is the finding.**
That suite builds its cmdhelp document BY HAND: I wrote `Args: mkArgs("ref")`
in it, so the fixture agreed with what I meant rather than with what the CLI
says. Five parity and drift tests passed against a document I authored to
match my own intention — the "a test that agrees with whatever the table says
is not a test of the table" shape, which the canonical-events test warns about
in its own comment two packages away. The new test reads the REAL command tree
via cmdhelp.Build, which is the only thing in this repo that can disagree with
me about what the CLI declares.

**P2: `pad project ready` withheld the ack handle** that `next` prints.
Showing a fired reminder on the surface an agent polls while withholding the
id it needs to retire it means the same entry comes back on every poll,
forever.

**P2: suggestions asserted a collection they did not have.** The orphan branch
admits ANY collection — its own comment claimed it gated on tasks "mirroring
the active-plan branch", and that comment was simply false — while the output
hardcoded `Collection: "tasks"` and the reason said "Open task". Pre-existing
for high-priority items since BUG-1082; my overdue bypass widened it to any
overdue item, which is how it surfaced.

Fixed by carrying the item's REAL collection rather than by narrowing the
branch: narrowing would silently drop the non-task items this has surfaced for
a year, and the defect is the mislabelling, not the inclusion. The false
comment is replaced with what the code actually does.

The first version of that test used an overdue IDEA and SKIPPED — ideas use
`new`, and the branch requires `open` or an active status, so it never became
a candidate. A test that cannot fire is a failed reconstruction, not a pass;
the fixture is now a bug-like collection whose vocabulary contains `open`,
which is the population the defect can actually reach.

Three mutants, three killed. Forty-one across the unit.

* fix(reminders): codex round 6 — reminders fired from soft-deleted workspaces

**P1, and the only defect in this unit whose consequence leaves the process.**
Workspace soft-delete deliberately keeps items for the 30-day restore window,
so the candidate query's filter on the ITEM's deleted_at found nothing wrong —
and the tick kept firing, emitting outbound webhook events for a workspace
whose owner had deleted it, possibly while deleting their account.

Both queries now join workspaces and require `w.deleted_at IS NULL`. Nothing
is destroyed: a restored workspace resumes firing, which the test asserts,
because "stops firing" and "is destroyed" are very different answers to
someone who restores a workspace and only one of them is right.

That test first failed for the WRONG REASON and the fixture was at fault: it
counted every outbox row in the workspace, and item creation writes its own,
so the assertion was satisfiable by the fixture itself and discriminated
nothing. Scoped to the reminder event type.

**`Use: "remind <ref>"` declared a requirement the command contradicts.**
cmdhelp derives the machine-readable arg spec from that string, and `--rearm`
takes no ref — so the published contract said "required" for something
optional. The requirement is CONDITIONAL, which cmdhelp cannot express, so the
honest declaration is `[ref]` plus the explicit check that names both call
shapes. The round-5 test grew a `required` column, which is what makes this
observable at all: asserting only the arg NAMES would have passed.

**The pad_item tool description omitted both new actions.** The params were
declared and the actions dispatched, but the prose an agent reads to decide
what a tool can do did not mention them — discoverable only by someone who
already knew to look. It now describes both, including the two things an agent
gets wrong: remind_at is an instant, and nothing but an explicit ack retires a
fired reminder.

Three mutants, three killed. Forty-four across the unit.

* fix(reminders): codex round 7 — one predicate for the scan and the arbiter

Third instance of one class, so this fixes the SHAPE rather than the instance.

The class: the candidate scan filters on something the fire transaction does
not revalidate, so a change committed between them fires a reminder that no
longer qualifies. Round 3 was a re-armed instant. Round 1's soft-deleted item
was the same thing caught from the other side. Round 7 is a workspace deleted
between the scan and the fire — the round-6 fix added the condition to the
SCAN only, and the arbiter went on not knowing about it.

Fixing those one at a time is what let the third happen. `reminderFireable` is
now a single string that both sites reference: the scan asks it and the fire
UPDATE re-asks it, so they cannot disagree, and a fourth condition is one edit
in one place rather than two edits someone has to remember are paired.

Written as a correlated EXISTS on item_reminders.item_id rather than a JOIN
precisely so the identical text is valid in both a SELECT and an UPDATE, and
the scan drops its table alias so the two uses are the same characters.

What deliberately stays outside it: `fired_at IS NULL` and `remind_at <= ?`
live on the reminder row itself, are already spelled identically at both
sites, and folding them in would need a parameter order the shared form cannot
express. Said in the comment so the omission reads as a decision.

Both directions are now tested at the arbiter — a workspace deleted mid-pass
and an item deleted mid-pass — because the item case previously relied on the
item load coming back nil, and someone simplifying the EXISTS down to the
workspace check alone would otherwise still see green.

Three mutants, three killed: the arbiter dropping the shared predicate, and
the predicate dropping each of its two halves. Forty-seven across the unit.

* fix(reminders): codex round 8 — workspace export silently dropped every reminder

WorkspaceExport is a hand-maintained field list, so a new table joins it only
if someone remembers. Reminders did not: a backup/restore, or a
SQLite→Postgres migration via `pad db migrate-to-pg`, dropped every pending
reminder with nothing in the destination to show anything had gone.

The line that list has always drawn is item-scoped workspace CONTENT
(comments, links, versions — exported) versus per-user state (stars, watches —
not). A reminder has no user column and hangs off an item, which puts it on
the exported side. Stating the rule rather than just adding the field, because
the next person adding a table needs to know which side they are on.

LIFECYCLE MARKS ARE CARRIED, not reset. A fired-and-unacknowledged reminder is
still owed to whoever armed it, so it arrives pending; an armed one whose
instant has passed fires once on the destination's first tick, which is what
would have happened had the workspace never moved. Re-arming everything on
import would invent a schedule the user did not set. NULL rather than empty
string for the unset marks — the lifecycle is defined by NULL-ness, and ""
would make a never-fired reminder read as fired at "".

TestMigratedTablesCoversTheExport caught the second half, which I would have
missed: `pad db migrate-to-pg`'s NUL preflight decides what to REFUSE on from
MigratedTables, so a table the migration copies and the preflight does not
know about is a gap in exactly the guard that exists to prevent one. Added
there too, with the reason it can never actually fire — every column is
machine-produced, so it is listed for coverage rather than expectation — and
the "six tables" prose it falsified is now seven.

Two mutants, two killed: export dropping the block, and import discarding the
marks. Forty-nine across the unit.

* test(reminders): state the fire-path invariant and pin it from the invariant

The lead's read on why rounds 4 and 7 were the same class: the fire path had
no stated invariant, so each fix defended an instance. This states it, and
derives the pin from the paragraph rather than from the bug history.

THE INVARIANT: the candidate scan is a hint and may be assumed to prove
nothing. Every condition that made a row a candidate is re-asserted inside the
transaction that marks it fired, in the same statement that does the marking,
so checking and writing are one atomic act.

Worded as "the scan proves nothing" rather than as a list on purpose — a list
invites the next person to add a condition to the scan and stop, which is
exactly what happened four times here.

TestFirePathInvariant is the pin: one table, one row per scan-side condition,
each invalidating that condition in the window between the scan and the fire
and asserting the same three things — nothing fires, no event leaves, the
reminder is not consumed. The earlier per-defect tests are folded in as rows;
they said the same thing one instance at a time, which is how four of these
shipped. Adding a fifth condition to the scan without a row here should feel
like an omission. It carries a positive control, because four cases that all
assert nothing happens would pass against a build that never fires at all.

The matrix immediately falsified a claim in the paragraph I had just written.
I wrote that the item load inside the transaction is "for the payload, not for
the check"; removing the item half of reminderFireable alone changes no
observable behaviour, because the load then returns nil and the deferred
rollback undoes the write. Item liveness is defended TWICE and a single-mutant
experiment cannot say which guard is carrying it — removing both is what kills
the test. Both are kept, the predicate is named as primary (the row never
matches, so no write happens at all), and the asymmetry is stated: workspace
liveness has no second line, which is why dropping ITS half does fail the pin.

Six mutants: five singles plus the pair. Five killed alone; the item single
survives by design and is documented as such rather than left as an unexplained
green. Fifty-five across the unit.

* fix(reminders): codex round 9 — one legacy row could hide every reminder

**P1: items.item_number is NULLABLE and I scanned it into an int.** Migration
006 added the column to existing rows, so a pre-numbering item still carries
NULL — and scanning NULL into an int fails the Scan, which fails the QUERY,
which degrades the whole pending-reminder section. One old row, and the
feature is dark for everyone in that workspace.

ListWatchesForUser, which this query was modelled on, uses sql.NullInt64 for
exactly this column. I copied its shape and dropped the part that handles the
column's actual nullability — the same way of being wrong as the round-5
cmdhelp fixture: borrowing a form without borrowing what it knows. The legacy
row now carries no ref rather than a fabricated "PREFIX-0", which would name a
different item.

**P1: export shipped reminders that import could only discard.** The items
section filters on deleted_at IS NULL, so a soft-deleted item is not in the
bundle and its reminder can never be reunited with it. My comment claimed the
item_links rationale — round-trip the raw graph so a restore reunites them —
which is true for links and false here, because links keep soft-deleted
endpoints in the bundle and items do not. A link is a row ABOUT two items; a
reminder whose item is absent is a dangling schedule.

**P2: import wrote remind_at raw.** Import is a writer, and a bundle is not
necessarily one this server produced — hand-edited, or from another instance.
A local offset or a bare date would land in the one column every comparison
downstream treats as a UTC instant, firing early, late, or never. It now
normalizes like every other door. An unparseable value is SKIPPED with a
warning rather than failing the restore, matching the lenient import-side
precedent already in this file, and the raw value's LENGTH is logged rather
than its content.

Three mutants, three killed; two needed rewriting because the single-line form
did not compile — reverting the nullable scan also requires reverting the
render, and dropping the normalization orphans a variable.

PROCESS FAULT, recorded because it makes this round's findings weaker than
they look: I edited the tree while this review was reading it — committed the
invariant work and ran five mutation experiments, which write and restore
source, over the same files. A review binds to the tree it read and I moved it
underneath. Every finding above was re-verified against the current tree
before being acted on, and the next round runs with no concurrent edits.

* fix(reminders): codex round 10 — one orphaned item aborted a whole restore

An ORPHANED item — one whose collection is missing from the bundle — still
gets an itemMap entry. It has to: the entry is written before the skip because
parent resolution inside the same loop reads the map for items it has not
reached yet. So `itemMap[x] != ""` is satisfied by an id that names no row,
and inserting a foreign key to it fails (SQLite enforces FKs here via the
DSN's `_pragma=foreign_keys(on)`; Postgres always does).

The pre-existing mapping is the sharp edge. The aggravating half was mine:
this loop treated a failed reminder insert as FATAL, where item_links and
item_versions both skip, so one orphaned item carrying a reminder rolled back
an entire 900-item workspace restore. A reminder is the least critical thing
in a bundle and it had the strictest failure handling in the file.

Both halves fixed: the loop gates on items that actually landed, and a failed
insert warns and skips like its siblings.

TWO GUARDS THAT ONLY DIE TOGETHER, and this is measured rather than assumed.
Reverting either alone leaves the test green — with the map gate restored the
skip survives the FK failure, and with the fatal return restored the gate
means the insert never fails. Removing both is what fails it. They are kept as
a pair because they defend the same failure at different depths (prevent the
bad write / survive a bad write arriving some other way), and the pair is
recorded in the code so a future reader does not delete one as dead after
watching its mutant survive. Second time this shape appeared today; the first
was item liveness on the fire path.

The bundle in the test is hand-built, because ExportWorkspace cannot produce
an orphan — which is the reason it needed a test. That shape only arrives from
a hand-edited or foreign bundle, and surviving those is what import is for.

Three mutants: two singles that survive by design, plus the pair that kills.
Sixty-one across the unit.

* fix(reminders): codex round 11 — four contract slips, one of them another unit's

**suggested_next returned up to eight entries against a cap of three.** Round 3
prepended reminders PAST the list's own cap, reasoning they should not compete
for slots. Every consumer — the web dashboard, `pad project next`, `pad project
ready` — is written for three.

Worse, it silently falsified a decision recorded elsewhere: BootstrapDashboard
deliberately has no suggested_next_overflow_count BECAUSE this list is capped
at three upstream, and its comment names raising that cap as the moment to add
one. My change made another unit's reasoning wrong in a file I never opened.
The combined list is now trimmed back to three, reminders still leading — a
reminder can push a task suggestion out, which is the right way round, and the
full set stays addressable in pending_reminders.

My first version of that trim used `limit`, which is REASSIGNED above to
len(candidates) — so on a workspace whose only entries are reminders it would
have truncated to zero, killing precisely the case the surface exists for.
Caught by reading the surrounding lines before running anything; it has its own
test now.

**pending_reminders was uncapped in the bootstrap projection.**
BootstrapDashboard embeds *DashboardResponse, so every new field joins the boot
payload automatically — here, a window of up to 50, which is the budget
PLAN-1410 spent a unit trimming. Capped at 5 with an overflow count, under its
own constant rather than borrowing bootstrapAttentionCap: they answer different
questions and a future change to one must not silently move the other.

**Truncation was reported from the wrong question.** The collector used the
store's `more` flag, which answers "is there another PAGE", not "did I read all
of THIS one" — so a window filling part way through the final page reported
that the caller had seen everything while unread rows sat behind the fill
point. The paging bounds are now injectable so the case is testable at all:
building it with a window of 50 needs ~75 rows in a specific pattern, with a
window of 3 it is four.

**Import accepted acked-without-fired**, which is not one of the lifecycle's
three states. Such a row fires, is excluded from the pending surface because it
is already acked, and can never be acknowledged because AckReminder requires
acked_at IS NULL — an event emitted into permanent invisibility. The
acknowledgement is dropped and the schedule kept, since an ack of something
that never fired means nothing.

Five mutants, five killed (one rewritten — removing the flag orphans a
variable). Sixty-six across the unit.

* fix(reminders): codex round 12 — a read is not a hold; scope the arm; ack from the ack

Four P2s from round 12 (two independent runs, both landing on the same
line of the fire path), each closed at the layer where it lives:

- fireOneReminder pins the item and workspace rows FOR NO KEY UPDATE on
  Postgres before the arbiter UPDATE. reminderFireable re-asserted
  liveness at the predicate's instant and nothing held it to the commit
  instant; under READ COMMITTED an archival could commit in between and
  the event left the process about a deleted resource. Same idiom and
  same lock strength as CreateAttachmentForLiveItem; SQLite is excluded
  by its BEGIN IMMEDIATE, not skipped for convenience. Two PG-only pins
  verify "blocked" in pg_stat_activity, not by elapsed time; the
  pin-removed mutant fails both.
- CreateReminder asserts "live item of THIS workspace" in the INSERT's
  own SELECT and returns ErrReminderItemGone otherwise. The table had an
  FK and no same-workspace constraint; a mismatched pair fed another
  workspace's title to this one's dashboard and webhooks. Handler maps
  it to 404.
- AckReminder matches every fired row (COALESCE keeps the first ack,
  updated_at moves only when acked_at does), so a no-match means exactly
  "not fired at the instant of the ack". The handler no longer decides
  409-vs-200 from the row it read before the UPDATE.
- The invariant paragraph gains its missing sentence: "at that instant"
  means the commit instant, and the pin is what makes the predicate's
  instant and the commit instant the same one.

Round-12 caveat carried: both runs were static reads (sandbox blocked
Go's build cache), so "four" is a floor, not a measurement.

Refs IDEA-2641

* fix(reminders): codex round 13 — a reminder's workspace must agree with its item's, at every read

Every reader scoped by r.workspace_id and then joined the item without
asserting the two agree. No door writes a disagreeing row today
(CreateReminder derives the pair from the item; import maps within the
workspace), and the table has nothing that forbids one — so a hand-edited
bundle, a future move door, or a direct write would carry one
workspace's item into another's dashboard, export, and webhooks.

The identity goes into reminderFireable (scan + arbiter), the Postgres
row pin, ListPendingReminders and the export query. One test writes the
row raw — the only way one can exist — and asserts it is inert at each
site; the predicate-removed mutant scans and fires it.

Refs IDEA-2641

* fix(reminders): codex round 14 — the by-id and by-item reads assert the same identity as every other read

GetReminder scoped by the row's own workspace_id and ListRemindersForItem
by item_id alone, so a row whose two columns disagree — the class rounds
12 and 13 closed at the scan, the arbiter, the pin, the pending surface
and the export — was still readable through the two reads that reach a
single row. reminderOwned is that identity on its own, without the
liveness half those two reads must not have (a fired reminder on an
archived item is history worth showing). The write paths reach a row
only through GetReminder, so scoping it scopes them; a row no door can
write needs no door to delete it. ListRemindersForItem now takes the
workspace its caller already resolved the item in.

The raw-row test asserts both reads refuse the row from both sides; the
reminderOwned-removed mutant surfaces it through GetReminder.

Refs IDEA-2641

* fix(reminders): codex round 16 — an archived item's reminders are readable, and its verbs say "archived"

The doors resolved the item live. Listing an archived item's reminders
answered 409 from a GET, and ack/re-arm/delete answered a bare 404 for a
reminder that exists on an item that exists — while the store, since
round 14, deliberately keeps that history readable. The API already has a
posture for archived items: GET reads them, mutations answer 409
"archived … restore it before editing" (writeItemResolveError). The list
now follows handleGetItem; the lifecycle verbs load the item
include-deleted, run the visibility check first, and then answer the same
409 every other item mutation does. One test walks archive → list 200 /
ack 409 / arm 409 → restore → ack 200 on the same rows.

Refs IDEA-2641

* fix(reminders): codex round 17 — one suggestion per item, the archived 409 by slug, and the door courtesy named

Three findings on the server pass. (1) An item that was both a fired
reminder and an ordinary candidate appeared in suggested_next twice; the
ordinary entry is dropped, the reminder entry (which carries the ack id)
stays, and two reminders on one item remain two entries. (2) Round 16's
409 for an archived item's reminder was written by re-resolving item.Ref,
which is derived and empty for a legacy item with no item_number — so the
class most likely to be legacy fell through to a bare 404. The slug is
handed over instead. (3) The archived check in resolveReminderForWrite is
check-then-write, and an archive landing in between lets the verb through:
accepted and documented — it is the posture of every item mutation here
(UpdateItem's UPDATE has no liveness clause), the outcome is benign, and
putting liveness in AckReminder's WHERE would re-create the no-match
ambiguity round 12 removed.

Refs IDEA-2641
2026-09-04 11:36:11 -04:00
xarmian 49e533d478 test(mcp,cli): pin same-name duplicate precedence on both doors (BUG-2850)
The lead's condition on the round-7 boundary. checkHierarchyAliasAmbiguity
refuses parent+plan — two NAMES for one target, which a caller can
collide without knowing — but deliberately does NOT refuse a same-name
duplicate (`--status A --field status=B`), because those are visibly
duplicates and both doors resolve them identically.

"Both doors resolve them identically" is the load-bearing half of that
argument and nothing enforced it. Two tests now do, one per door,
asserting the SAME outcome: the `field` entry overlays the named param,
because cmd_item.go and dispatch_http_advanced.go both apply named flags
first and overlay --field after.

Per-door mutation matrix, run this turn from file backups:

  make the named param win on the HTTP door -> only the mcp test fails
  make the named flag win on the CLI door   -> only the cmd/pad test fails

Neither mutant reddens the other door's test, which is the property
worth having: the doors cannot drift apart again without exactly one of
these going red and the boundary getting re-examined rather than
silently becoming untrue.

Also filed, per the lead's ruling: BUG-2870, the padded-`field`-key
divergence with NO `fields` object (`--field " effort=l"` stores an
undeclared " effort" key on the CLI door and writes `effort` on the
remote one). Out of scope here — it predates this PR's claim rather than
defending it — and its fix is a policy call on the CLI's input contract,
so it wants a ruling, not a quick patch.

gofmt clean · go vet clean · go test ./... green (29 packages)
2026-09-03 16:22:00 +00:00
xarmian 2cf9f0035a fix(mcp,server,cli): three codex round-2 findings (BUG-2850)
1. [P1] The structured-value refusal was in the wrong place and killed the
   fix. It went into BuildCLIArgs, which env.Dispatch runs for BOTH
   transports before handing off to whichever Dispatcher is configured — so
   it blocked the remote /mcp door too, and the native-field handling that
   is the whole point of this change was never reached.

   Moved into ExecDispatcher, which IS the stdio door. My own test could not
   see this: it called mapItemCreate directly, so it vouched for the mapper
   and not for the path that reaches it — CONVE-19's exact shape, in a unit
   where I had already written binding tests for the other half.

   The tests are now split along the two claims the first version conflated:
   nested values REACH the dispatcher (the remote door is unblocked), and
   refuseStructuredFieldsOverCLI refuses them at the CLI door naming the
   transport.

2. [P2] The CLI warning sat after the `--format json` early return, so the
   caller most likely to have sent a mistyped key — one piping stdout into a
   parser — was the one caller who never saw it. Moved above the return, and
   out of the `ref != ""` branch it was also trapped in. Still stderr.

3. [P2] A nil value in fields_patch DELETES the key (store/items.go), so
   reporting it as an undeclared field told the caller a field was stored
   that the same request removed. Filtered at the patch site, not inside
   UndeclaredFieldKeys, because nil means "store JSON null" on the
   full-fields path where reporting it is correct.

Gates: gofmt clean, go vet clean, go test ./... 29 packages ok.

Claude-Session: https://claude.ai/code/session_011Q4b1iHtJtSyMs7BA2ySxo
2026-09-03 01:07:45 +00:00
xarmian dc3fc2d50e feat(server,cli): name undeclared field keys on the write response (BUG-2850)
Undeclared keys are ACCEPTED — the census found 168 live values under 14 such
keys, and refusing them would break read-modify-write on items nobody edited
wrongly. But once stored, a typo and a deliberate extra field are
indistinguishable, so the write now says which keys it did not recognize.

- models.Item gains `Warnings *ItemWriteWarnings` with `undeclared_fields`,
  omitempty and additive. NEW API SURFACE: item write responses carried no
  warnings element before. Wrapping the response as {item, warnings} was the
  alternative and would have broken every existing parser; a clean write is
  byte-identical to before.
- items.UndeclaredFieldKeys consults models.IsReservedItemField rather than
  re-listing the reserved set — that set exists so callers ask, and its doc
  comment records what re-listing cost last time. So a write carrying
  implementation_notes or github_pr reports nothing.
- fields_patch reports only the PATCHED keys. A stray key already on the item
  is not something this write introduced, and naming it on every touch would
  train the reader to ignore the field.
- The CLI prints one line to STDERR. Never stdout: `--format json` output is
  piped into scripts, and a warning there would corrupt the JSON they parse.
- CLAUDE.md documents the element as new surface.

Controls: never attaching the warnings fails the pin; reverting the HTTP
mapper's native overlay fails the remote-door type test; dropping the
reserved-key exclusion fails its own test.

Two coverage gaps the controls FOUND rather than confirmed, both now closed:
the remote door's native overlay was covered by no MCP test at all (a revert
left the package green), and the reserved-key exclusion had no test either.
Both were written after the control survived, which is the only reason they
exist.

Gates: gofmt clean, go vet clean, go test ./... 29 packages ok.

Claude-Session: https://claude.ai/code/session_011Q4b1iHtJtSyMs7BA2ySxo
2026-09-03 00:54:56 +00:00
xarmian e4415ddd04 fix(cli): name the skipped-table suspects instead of counting them (BUG-2810)
Codex round 12, polish rather than a defect. The advisory reported how many
values in non-migrated tables mention a NUL escape, and then made the operator
run `pad db scan-nul` to learn which — when the rows were already in hand.

They are listed now, in the same shape as every other row this command prints.
The test asserts the table.column appears rather than only the surrounding
phrase, so a regression to a bare count fails it.
2026-09-02 18:42:36 +00:00
xarmian d9f3fe3881 fix(cli): filtering suspects out of the check also filtered them out of the report (BUG-2810)
Codex round 11. Round 10 stopped probing suspects from tables the migration
does not copy, which was right — but it also dropped them from the output,
while the comment two lines below still claimed "the others are still
REPORTED". A legacy shadowed-NUL in activities.metadata produced no warning at
all.

They are now COUNTED and named, pointing at `pad db scan-nul` for detail.
Counted rather than probed on purpose: whether one is actually fatal can only
be answered by the destination, and asking would put them back inside the
fail-closed rule the filter exists to keep them out of.

The test captures stderr and asserts the advisory appears, and is
mutation-verified: suppressing the notice fails it with "the suspect was
filtered out of the check AND out of the report".

THIS IS THE THIRD TIME on this branch that the same shape has appeared — a
filter that is right about what to ACT on quietly becoming a filter on what to
SAY. The first was the scan dropping suspects entirely; the second was the
preflight refusing on tables it does not copy and then, fixing that, going
silent about them. Each fix was correct about the action and wrong about the
reporting, and each time the comment stayed true while the code stopped being.
Worth naming as the pattern rather than as three unrelated defects.
2026-09-02 18:29:55 +00:00
xarmian 39a8366060 fix(cli): filter suspects by table BEFORE asking the destination (BUG-2810)
Codex round 10, and it is round 9's over-refusal reintroduced through the other
path. The fail-closed rule refuses on a suspect that could not be VERIFIED, and
it ran over every suspect before MigratedTables was applied — so an
unverifiable row in `users`, `sessions` or `activities` blocked a copy that
would never have touched it. Filtering first also stops the oracle making round
trips whose answer cannot matter.

Two things learned writing the test, both worth more than the fix.

**The first version of it proved nothing.** It used a nil destination, but the
fail-closed branch only runs once there is something to ask, so it passed with
and without the fix. The real fixture needs a live destination AND a genuinely
unverifiable row: a NULL primary key, which SQLite permits in a declared TEXT
PRIMARY KEY and no other engine does. Mutation-verified in the new shape — with
the filter back in its old position the test fails with the reported symptom,
"1 suspect value(s) could not be checked; nothing was migrated".

**Layer B is STRICTER than the shared predicate for this shape.** The fixture
would not insert until the triggers were dropped: SQLite's json_tree walks
tokens rather than building a map, so it sees the NUL in the shadowed member
that our Go predicate cannot. That narrows how such a row can exist at all — it
must be legacy data written before the triggers, which is exactly the
population BUG-2810 is about. Recorded in the fixture rather than left as a
surprise for the next person whose insert is refused.
2026-09-02 18:17:40 +00:00
xarmian e89c8c8ab6 fix(cli): two ways the preflight and its remedy disagreed with the migration (BUG-2810)
Codex round 9, both confirmed against the code rather than reasoned about.

**PAD_DATABASE_URL was treated as proof of a PostgreSQL deployment**, so the
flow this unit prescribes broke on itself. cmd_server.go opens PostgreSQL only
when PAD_DB_DRIVER=postgres; PAD_DATABASE_URL is ALSO migrate-to-pg's target,
and its default. An operator who follows the preflight — refused, told to run
`pad db repair-nul`, with the target URL still exported in their shell — got
"This deployment is PostgreSQL ... Nothing to scan or repair" and exit 0. The
remedy the refusal names did nothing, which is the failure mode this unit has
now produced three separate ways. PAD_DB_DRIVER alone decides. Verified by
running the real command with the target exported.

**The preflight refused on tables the migration does not copy.**
ExportWorkspace / ImportWorkspace read six tables, and migrate-to-pg's own help
says users, platform settings and auth data are not migrated — so a NUL in
users.name blocked a copy that would never touch it, demanding the operator
rewrite content unrelated to the migration they asked for.

Refusal is now filtered to store.MigratedTables(). Those rows are still
REPORTED: `pad db scan-nul` lists them, they are real, and going quiet about a
broken row because this command does not care about it would be the
information-discarding the preflight was already corrected for once.

The table set is pinned by REFLECTION over models.WorkspaceExport's shape, not
by a regex over ExportWorkspace's SQL — TASK-2825 already established that
multi-line and Sprintf-composed SQL are invisible to any source-level
instrument. It fails in both directions: a new export section with no entry
(a miss, ending in a half-finished migration) and a spurious entry (an
over-refusal).

One residual, stated rather than hidden: the export also skips SOFT-DELETED
collections and items, and this filter is per-table. A NUL in a soft-deleted
item still blocks. Narrowing it needs a per-row deleted_at check at every
candidate, which costs more than the remaining over-refusal — the operator's
way out is the same single command either way.
2026-09-02 18:01:37 +00:00
xarmian 0363c139a9 fix(store,cli): the oracle failed open, and it over-refuses one column (BUG-2810)
Three findings from codex round 5, all real; the third corrected a claim I had
made about the design.

**The suspect path could leave data unrepaired and exit 0.** The CLI printed
SuspectsFailed and then returned nil, checking only the violation bucket. A
script sees success; an operator who trusts the status moves on. Both buckets
now decide the exit code, and the decision is extracted into
nulRepairExitError so it is testable without a database — the bug was in the
decision, not in the repair, and a test that needs a fixture to reach it is a
test nobody writes.

**The destination oracle failed open.** Connection failures, timeouts and
read-back errors were bucketed with "the destination answered, about something
else" — reported and not refused on. So an UNVERIFIED suspect passed the
preflight, which is the defect the suspect class was added to correct arriving
by a different route.

There are now three outcomes rather than two: the server answered with a NUL
code (refuse), the server answered with another complaint about the value
(report, because a NUL preflight that quietly grew into a general one would
block migrations unrelated to this bug), and the server never answered
(REFUSE). ErrDestinationCheckUnavailable carries the third, and
TestDestinationOracleFailsClosedOnAnUnusableConnection pins it against a real
closed pool — with an open-pool control first, since a classifier that answered
"unavailable" for everything would satisfy the assertion and refuse every
migration.

**The oracle is not a perfect model of the migration, and I said it was.**
Codex claimed workspaces.settings is normalised on import, so the cast
over-refuses there. Measured rather than argued, by importing the same
shadowed-duplicate value into three columns against a real server:

	workspaces.settings  -> import SUCCEEDS, stored as {"a": "clean"}
	items.fields         -> import FAILS, SQLSTATE 22P05
	collections.schema   -> import FAILS, SQLSTATE 22P05

CreateWorkspace runs models.NormalizeWorkspaceSettings, a map round-trip that
drops the shadowed member. So the claim was right, and my own runtime demo
earlier on this branch — which used workspaces.settings — was showing a
spurious refusal.

The cast STAYS. That row is a value Layer B refuses on every write today and
exists only because it predates enforcement, so surviving the migration is an
accident of one column's normaliser rather than a property worth preserving,
and repair-nul clears it in one command. Deriving "would this column's writer
normalise it" is a per-column enumeration, which is the shape this cluster
keeps proving unmaintainable.

What changed is the CLAIM. The file header no longer says the oracle is "exact
in both directions" — it is exact about the VALUE and is not a model of the
MIGRATION; the refusal no longer tells an operator PostgreSQL would reject the
row, only that the value carries a NUL jsonb refuses; and the measurement and
the over-refusal are written into CheckJSONBAcceptable's doc comment and
docs/backup.md, which also now states that the check errs toward refusing.

The disposition is flagged to the lead rather than settled here: skipping
normalised columns is a scope call, not mine.
2026-09-02 17:05:13 +00:00
xarmian 57b7ca5f48 feat(store,cli): ask the destination about suspects instead of dropping them (BUG-2810)
Day-54 lead ruling on PR #1233, and the ruling names the defect precisely: the
scan's own SQL pre-filter already surfaces the shadowed-duplicate row as a
candidate, and `ParameterRefused` then drops it. So the preflight was
discarding information it was holding and going on to promise the migration
would go through. I had recorded that as an accepted residual on the grounds
that closing it would violate DOC-2823's one-layer rule — but that rule is
about what the enforcement layers REFUSE. It says nothing about a preflight
throwing away a candidate it had in hand.

**The SUSPECT class.** A pre-filter hit the predicate does not refuse. Most are
doubled-backslash literals — text that writes ABOUT the escape, which is the
false positive this whole predicate family exists to avoid. One member is not:
a NUL in a value shadowed by a LITERAL duplicate key, which a map-model decode
drops and PostgreSQL refuses. Nothing here can tell them apart, so nothing here
tries: `pad db scan-nul` lists them under their own heading, apart from the
violations, with what resolves them.

**The destination is the oracle.** `pad db migrate-to-pg` casts each suspect on
the TARGET connection — `SELECT $1::jsonb`, side-effect-free, and the very cast
an INSERT performs — and refuses on 22P05 / 22021. That is exact in both
directions precisely because it is not a fourth opinion of ours. Measured
against a real server: the literal is ACCEPTED, the shadowed duplicate is
REFUSED with 22P05, and a non-JSON value fails for a reason that is reported
rather than refused on, because a NUL preflight that quietly grew into a
general one would block migrations unrelated to this bug.

**The repair had to be measured, not assumed, and the answer changed the
design.** `textguard.Repair` leaves the shadowed value completely untouched:
its scanner is gated on DocumentDecodesNULAnyShape, a map-model question that
answers false for exactly this shape, so it never runs. A preflight that
refused the row and printed `pad db repair-nul` would have been printing a
command that does nothing to it — a remedy nobody ran (PATTE-135). So the
repair reaches the class through the token-level scanner, exported for this,
which rewrites the shadowed escape and still leaves the literal byte-identical
because it consumes escapes in order.

Suspects get their own buckets in the repair report rather than being folded
into Repaired, so the dry run's promise and the run's result stay the same
number.

**Nothing about what any layer REFUSES changed.** textguard.KnownGaps and its
pin are untouched, and TestScanNULInheritsTheRecordedKnownGaps still asserts
the scan does NOT detect the shape. TestSuspectsCollapseWhenBUG2812Lands fails
when the token-walk makes that false, and names every file to delete — the
suspect path is a second mechanism that exists only while the predicate is
blind.

**One defect this found that no test did.** Running the real command against a
real Postgres, the refusal announced "0 stored value(s) carry a NUL; nothing
was migrated" while listing one — the count used the violations only, and the
tests asserted the message CONTAINED "nothing was migrated" without reading the
number. Fixed, and the assertion now reads the count. The whole loop is now
verified end to end: preflight refuses, `repair-nul` fixes, the migration
completes.

My own prose from earlier on this branch is corrected with it. ScanNUL's doc
comment, the preflight's, and docs/backup.md all said this shape passes the
preflight and fails mid-copy, which the same commit makes false.
2026-09-02 16:45:38 +00:00
xarmian a65252dab1 fix(cli): print the NUL report on stdout so it can be captured (BUG-2810)
`pad db backup` and `pad db restore` keep their progress on stderr because
stdout may carry the backup itself. These two commands emit no data at all,
and their REPORT is the whole point — an operator piping
`pad db scan-nul > affected.txt` was getting an empty file and the list on
the terminal, which is the opposite of what the command is for.

Report to stdout; the confirmation prompt, its warning and the Postgres
not-applicable notice stay on stderr, where a prompt belongs.

Verified by running the real command with 2>/dev/null and reading the list.
2026-09-02 16:28:42 +00:00
xarmian 178b6b5010 fix(server,store): two more from codex rounds 3 and 4 (BUG-2810)
**The import repair could silently change what gets imported.** It decodes into
map[string]any, where a repeated object member keeps only the LAST value. The
TYPED decode that runs next does not agree: encoding/json unmarshals members in
order into the same struct field, so two `"workspace"` objects MERGE there and
collapse here. A body with duplicate members would therefore import differently
with --repair-nul than without, which is outside what a flag by that name may
do.

It now DECLINES such a body: returns it untouched, lets the gate judge it
exactly as it would without the flag, and says why in the refusal — "the
payload repeats the member X, and repairing it would change which value is
imported". Detection is a token walk, because a decode is what loses the
information: by the time there is a map the duplicate is gone. The detector's
own test carries the false positive that matters — the same member name in
SIBLING objects is not a duplicate, and a single shared set of names would
decline every real export, since items all carry `id`, `title`, `slug`.

Rewriting such a body faithfully wants a token-preserving pass, which is
BUG-2812's token-walk and not a rider on this. A real export cannot contain
duplicate members (json.Marshal does not emit them), so declining costs nothing
an operator meets by accident.

The tally now owns the repair — decodeJSONRepairingNUL takes it and calls
Apply — so the count and the declined reason come back through one object
instead of a return value a caller has to remember to record. That is the same
mistake this branch already made once, when the JSON path dropped the count and
the header reported 0 for an import that had rewritten a value.

**A row the repair could not address was reported as a failure.** A NUL in a
key column the list does not protect, on a row whose violation is elsewhere,
makes the address unbindable: Layer A inspects every bound parameter, including
a WHERE clause's, so the lookup is refused before SQLite is asked to find the
row. It landed in Failed carrying "invalid text parameter: parameter 2" — the
same information phrased as a fault in the repair rather than a property of the
row. Now detected up front and reported as a skip with the reason, alongside
the two skips that already existed.

**One finding NOT fixed, deliberately, and recorded instead.** Round 3 raised
that the scan misses a NUL in a value shadowed by a LITERAL duplicate key, so
such a row passes the migrate-to-pg preflight and then fails during the copy —
the exact failure the preflight replaces, surviving for one shape. That is
textguard.KnownGaps: a blind spot every layer shares on purpose, which DOC-2823
forbids closing in one layer alone, because layers disagreeing about one value
is the defect this cluster is made of. So it is named in ScanNUL's doc comment,
in the preflight's, and in docs/backup.md for the operator, and
TestScanNULInheritsTheRecordedKnownGaps pins the miss and FAILS when it stops
being one — the notification that BUG-2812 has landed and those three prose
sites need updating. The consequence is recorded on BUG-2812's trail.

Round 2's single finding was refuted rather than fixed: it predicted
TestRepairFlagReachesTheNestedAndObliqueForms would fail, on a mechanism that
describes the raw-byte scanner this branch had already replaced. The test
passes; the outer decode resolves the oblique spelling before the walk sees it.
2026-09-02 16:24:25 +00:00
xarmian 49bd342e4c fix(store,server,cli): three defects from codex round 1 (BUG-2810)
**The import flag could not repair the column it exists for.** `--repair-nul`
scanned the RAW body for a live escape, which is right for a value the gate
reads at the top level and wrong for the one that actually matters. An item's
`fields` blob travels through an export as a STRING: a NUL escape in the stored
blob marshals into the body with a DOUBLED backslash, which a raw scan must
leave alone because at that layer it is literal text — while the gate refuses it
anyway, since it decodes the body and re-parses that string as the document it
is.

So the repair now walks the DECODED body with the same classing bodyDecodesNUL
uses, one verb changed: where the gate asks textguard whether a value decodes to
a NUL, this asks textguard to repair it. Two walks of one shape in one package
is a real risk, and the mitigation is that they are measured against the same
corpus in both directions rather than reviewed for similarity —
TestBodyRepairMirrorsTheGateOverTheCorpus drives every case through the body
shape and asserts refused-becomes-accepted and accepted-stays-byte-identical.

Two consequences worth stating. The walk also reaches the OBLIQUE spelling — the
backslash written as its own escape, so the six characters never appear in the
raw bytes at all — which the scanner could not, so the test that pinned that
limit is replaced by one asserting the capability. And re-encoding is now
possible, so it is bounded: UseNumber, so an integer wider than float64 is not
silently re-emitted in scientific notation; SetEscapeHTML(false); and a body
with nothing to repair is returned byte-identical rather than round-tripped. The
mutation that removes UseNumber turns 9007199254740993 into ...992, and a test
says so.

The header is now X-Pad-Repaired-NUL-Values, because at the decoded layer an
escape is not a thing that exists any more and one nested document may have
carried several.

**The scan could not run on the databases it exists for.** Several protected
tables carry a NULLABLE workspace_id — activities, api_tokens, mcp_audit_log —
and the scan selected it into a plain *string, which fails with "converting NULL
to string is unsupported" and takes the scan, the repair and the migrate-to-pg
preflight down with it. Every column is now scanned as sql.NullString: SQLite
also permits NULL in a declared PRIMARY KEY that is neither INTEGER PRIMARY KEY
nor NOT NULL, which no other engine does, and a NULL key cannot address a row
for an UPDATE — such rows are reported and skipped with the reason rather than
handed a WHERE that matches nothing. Verified against the unfixed code: the scan
returned `scan activities.actor row: sql: Scan error ... converting NULL to
string`. It needed a VIOLATING row in such a table, which is why every fixture
that planted its rows in `items` missed it.

**--force by accident.** The repair skipped the running-server check whenever
--from was given — and the most natural --from an operator types is the path
`pad db scan-nul` just printed, which IS the live database. The check is now on
the resolved path (Abs + EvalSymlinks, so a symlinked data directory or a
relative path still matches), and a --from naming an unrelated backup stays
unguarded, which is correct: nothing is writing it.

The ordering moved with it. `store.New` runs pending migrations, so the refusal
now happens BEFORE the database is opened; opening first and refusing second
made the guard arrive after the thing it guards against.
2026-09-02 15:15:11 +00:00
xarmian 63da2f4f5f feat(store,server,cli): count and repair the legacy NUL population (BUG-2810)
Layers A and B stop the value being written. Neither makes a row that
already carries one go away, and BUG-2810's filing is what that costs: an
affected workspace exports with a 200 and re-imports with a 400, so a
self-hoster restoring their own backup is blocked with no path forward in
the product, and `pad db migrate-to-pg` fails partway through the copy
against PostgreSQL's jsonb parser rather than up front.

This is DOC-2823's S3, on Dave's day-54 rulings: U+FFFD as the replacement,
repair standalone only with a migrate-to-pg preflight that refuses and
prints the command, `--repair-nul` on import shipping default-strict.

ONE REPAIR, beside the one predicate. textguard.Repair lives next to
ParameterRefused because four layers that agree about what is REFUSED and
disagree about what a repair PRODUCES is this bug family arriving one step
later. Its contract is a property over the same corpus, in both directions:
every refused value becomes one all four layers accept, and every accepted
value comes back IDENTICAL. The second half is the load-bearing one — a
repair that tidies values nobody complained about rewrites
`{"a":"x\\u0000y"}`, six literal characters after a doubled backslash, and
corrupts it.

The JSON arm is a string-literal SCANNER, not decode-walk-remarshal, which
is what the recon write-up proposed before it was written. Re-marshalling
changes four things nobody asked to change — object key order, insignificant
whitespace, integers wider than float64, HTML-ish characters — and silently
drops one of a document's LITERAL duplicate keys, which is a gap BUG-2812
owns and the last thing a repair should do. Scanning copies every byte it
does not deliberately rewrite, so an untouched document is byte-identical
without that having to be argued. A substring replace is not equivalent and
the test that proves it took a mutation to find: a doubled-backslash literal
ALONE never reaches the scanner, so the discriminating fixture is one
document carrying a live escape AND a literal.

THE COUNT IS COMPUTED IN GO. Measured on the read path in this worktree: a
row planted with `bad<NUL>name` reads back into a Go string with all 8 bytes
and the NUL intact, while `length(name)` in the same database answers 3.
TASK-2824 found that C-truncation and concluded no DB-side REPAIR could be
trusted; the same measurement on the read path says no DB-side COUNT can be
either. SQL narrows — `instr(col, char(0))`, plus the escape prefix on
JSON-classed columns, which is textguard's own pre-filter — and never
decides. The decision stays ParameterRefused with isJSON from the shared
86-column list, i.e. Layer B's classing.

Row addressing is read from the live schema rather than a hand-kept map:
39 tables carry protected columns, one (item_wiki_links) declares no primary
key and is addressed by rowid, two have composite keys, and five have a
single key that is not `id`. The repair checks RowsAffected because an
address that stopped selecting its row would otherwise commit an UPDATE that
touched nothing and report it as repaired — the one failure an operator
cannot see in the output.

`email_optouts(email)` is both a protected column and its own primary key.
Repairing it changes the row's identity and can collide with an existing
row, which in that table means somebody starts receiving mail again. It is
reported and skipped, with the reason.

The import flag is NOT an exemption from the gate. `--repair-nul` buys the
body one repair attempt and then runs the same `bodyDecodesNUL` on the
repaired bytes, which still decides — a decode path that skipped the check
is the door BUG-2803 spent thirty rounds closing, on the endpoint carrying
the largest attacker-controlled body in the product. Only the ESCAPE form is
repaired: a raw NUL byte makes the document invalid JSON, and widening what
parses is not this flag's job. Both doors are covered, JSON and tar.gz,
because giving them different answers is how one of them keeps being
forgotten.

Postgres is settled with evidence rather than sent up as a ruling: it cannot
hold either defect (22021, 22P05) and the four-way differential test already
pins that, so the scan reports not-applicable WITH the reason rather than
returning a zero a reader could mistake for a clean database.

Spellings settled here, per the dispatch: `pad db scan-nul` and
`pad db repair-nul` as siblings rather than `repair --nul`, matching
`migrate-to-pg`'s hyphenated compound — a repair verb that errors when given
no flag is a worse shape, and there is no second repair to share it with.
scan-nul IS the dry run, so repair-nul grows no --dry-run. It refuses while
the server is running unless --force, on the `pad db restore` precedent: the
report is a claim about a database, and one somebody else is concurrently
writing makes it a claim about a moment that has passed.

docs/backup.md's section on this is rewritten. It still said the rule lives
in the binary and not the database, which S2 made false, and it pointed at
this item for a preflight and a repair that now exist. Its import examples
also showed `pad workspace import < file`, which has never worked — the file
is an argument.

Closes BUG-2810.
2026-09-02 14:53:31 +00:00
xarmian ba1255881d fix(server): refuse a decoded NUL in a JSON request body (BUG-2803) (#1220)
* fix(server): refuse a decoded NUL in a JSON request body (BUG-2803)

The body half of BUG-2782 (path) and BUG-2784 (query). A caller-supplied
string reached a Postgres text parameter, Postgres refused it, and the
handler answered 500 — the honest answer is 400.

WHY THE TRANSPORT RULE CANNOT BE EXTENDED, which is the whole reason this
is a different fix rather than a wider middleware. ValidateQuery works
because a decoded query value is a substring of the raw query with ASCII
substitutions: the bad byte in the raw text IS the bad byte in the value.
That property fails for a JSON body — the reachable NUL arrives as the
six-character escape, all ordinary ASCII — so no request middleware can
find it without decoding the body, which is the handler's job.

MECHANISM, each premise measured against encoding/json rather than
reasoned about:

  raw NUL inside a string  -> decode ERR (invalid character in string literal)
  raw NUL after the value  -> decode ERR
  the escape in a value    -> decodes to a string CONTAINING a NUL
  the escape in a KEY      -> same
  a DOUBLED backslash      -> decodes to literal text, NO NUL
  the uppercase spelling   -> not a JSON escape at all

So the escape is the only vector and its substring is a sound FAST PATH
(absent -> no NUL possible), but not a sufficient test: a doubled
backslash carries the same six characters and decodes to text. In this
product that is not hypothetical — items and documents store markdown,
and a document about JSON escapes is an ordinary thing to write. The
exact step is json.Decoder.Token(), which returns DECODED strings, covers
object keys and arbitrary nesting (an item's fields blob), and needs no
knowledge of the destination type.

NOT REFLECTION over the decoded value, the other obvious design: it sees
[]byte fields AFTER base64 decoding, so a body carrying legitimate binary
({"b":"AQAC"} -> bytes 01 00 02) would be refused for a NUL that is not
text. A token walk sees the base64 characters. No request struct has such
a field today (searched: []byte with a json tag in internal/server and
internal/models, non-test — only models.YjsUpdate.UpdateData, which no
handler decodes from a body); the token walk is chosen so adding one
later cannot silently start rejecting valid requests.

BUFFERING IS NOT A COST. json.Decoder.Decode already holds the whole
top-level value in memory — refill accumulates into dec.buf and grows it
by doubling (encoding/json/stream.go) — so streaming never avoided the
copy. Measured on the 64 MiB workspace-import shape, total allocation:
stream+Decode 354.7 MiB, ReadAll+Unmarshal 256.5 MiB, ReadAll+Decode
512.5 MiB. Peak heap is order-dependent and does not discriminate; the
first run of that measurement showed a 0.77x peak win that vanished when
the legs were swapped, so only the allocation figure is claimed.

POPULATION, measured on Postgres 17 through the real router with a
control leg on every endpoint (92 mutating routes enumerated via
chi.Walk; 13 probed):

  before: 12 of 13 DOOR (control 201 / NUL 500, SQLSTATE 22021)
  after:  0 of 13 — every NUL leg 400, every control leg unchanged

Confirmed doors: workspace name, collection name, item title, item
content, item fields value, item title via PATCH, comment body, agent
role name, view name, document title, webhook secret, workspace import.
workspace-token name is UNMEASURED, not clean — its control leg 500s on
an unrelated FK in this fixture. The other 79 routes are unprobed, not
claimed clean; the completeness argument is structural instead, and
enforced by a test rather than asserted.

SECOND DEFECT, named rather than slipped in: the six handlers that
decoded straight off r.Body had no http.MaxBytesReader either — the cap
decodeJSON has always applied — so each was an unbounded body read.
Routing them through decodeJSON closes that too.

COMPATIBILITY: json.Unmarshal refuses trailing non-whitespace after the
JSON value where Decode ignored it. Deliberate, same direction as this
fix, and the only behaviour change beyond the refusal. Trailing
whitespace still passes. An EMPTY body still returns a wrapped io.EOF,
because handlers_playbooks.go reads errors.Is(err, io.EOF) as "no
arguments supplied" — caught by TestPlaybookRunAcceptsEmptyBody, which is
exactly the wiring a helper-level change is blind to.

No call site changed for the refusal itself: all 65 decodeJSON callers
already turn a decode error into a 400 carrying err.Error().

Release note: a NUL character in a JSON request body now returns 400
instead of 500 on Postgres deployments.

Claude-Session: https://claude.ai/code/session_011T365kP1N9V88y15HxL4YN

* fix(server): follow the NUL refusal into JSON-encoded string fields (BUG-2803)

Codex round 1 on #1220: the check scanned ONE JSON layer, and several
fields cross the wire as JSON-ENCODED STRINGS rather than nested objects
— an item's fields, a collection's schema, a workspace's settings. The
OUTER decode of {"fields":"{...}"} yields the inner document as literal
text, in which the escape is still six ordinary characters and no NUL
exists, so the single-layer token walk passed it.

MEASURED on Postgres 17 with a control leg on each, after the
single-layer check was already in place:

  item.fields  as a JSON-encoded string   500   control 201
  collection.schema  as a string          500   control 201
  workspace.settings as a string          500   control 201

The error is DIFFERENT from the rest of this family, which is why it is
worth reading rather than assuming:

  insert collection: ERROR: unsupported Unicode escape sequence (SQLSTATE 22P05)

22P05, not the 22021 the path and query halves produce. The outer string
is pure ASCII so it never trips the text-encoding check; this is
Postgres's own JSON parser refusing the escape inside a document bound
for jsonb, which cannot represent a NUL. After this change all three
answer 400 with their control legs unchanged.

THE FIX: when a decoded string is itself a complete JSON object or array
— the class this API re-parses downstream — walk it too, to a depth
bound of 8. Recursion terminates on its own (each level is a strict
substring of the one above); the bound keeps a hostile body from buying
many full re-parses, and AT the bound the body is refused rather than
passed uninspected, since the escape is known to be present and the walk
has stopped looking.

WHAT THIS OVER-REFUSES, by design and pinned by a test: the rule is
structural, not destination-typed, so a plain TEXT field whose ENTIRE
value is a valid JSON document carrying the escape is refused too, even
though its column would have stored it. Prose ABOUT a JSON escape does
not parse as a bare document, so the case is narrow, and a value of that
shape breaks any consumer that parses it. The destination-typed
alternative — an allow-list of the fields that arrive JSON-encoded — is
exactly correct and goes stale in silence, which is the failure mode
ValidateQuery's comment rejects when it explains why per-site query
validators could not be written.

Tests: nested documents (fields/schema/settings/array/twice-encoded),
with controls for ordinary content, a doubled backslash INSIDE the
nested document, a string that starts like JSON but does not parse, and
prose that merely mentions the escape; the over-refusal pinned as a
decision rather than left as an accident; the depth bound; and a wiring
leg through the real router on SQLite, where the write would otherwise
SUCCEED so a green cannot be the database doing the work. Fixtures build
their JSON-encoded strings with encoding/json rather than hand-written
backslashes, since the escaping rules are the subject under test.

All four new tests fail with the recursion removed.

Claude-Session: https://claude.ai/code/session_011T365kP1N9V88y15HxL4YN

* test(server): build the NUL-bearing timeline fixture through the store (BUG-2803)

TestTimeline_NeverEmitsACursorItWouldRefuse built its fixture through the
API on a premise its own comment stated: "a structured id comes from the
item's fields blob, which nothing validates on write". BUG-2803 made that
false — decodeJSON now refuses a body whose strings decode to a NUL,
including one nested inside a JSON-encoded `fields` string — so the API
can no longer produce the row and the test 400'd on its fixture.

Repaired rather than deleted, because the DEFENCE it covers is still
live: rows in this shape can predate the rule, and the store has no such
check of its own, so a migration, an import or any future non-HTTP writer
can still produce one. The timeline must keep refusing to hand out a
cursor it would then reject.

The fixture now writes the blob directly, injecting the six-character
JSON escape rather than a raw NUL — the blob is JSON text and both
backends reject a raw NUL in it; the NUL comes into existence when Go
DECODES the blob, which is exactly how the timeline ends up with one
inside an entry id. The test is not vacuous under the change: it asserts
the NUL-bearing id took the positional fallback, so an injection that
failed to produce a NUL fails the test rather than passing quietly.

This is the CONVE-23 case — a change that falsifies existing prose owes a
sweep for that prose. The stale sentence was found by the test failing,
not by the sweep, which is the weaker of the two ways to find it.

Claude-Session: https://claude.ai/code/session_011T365kP1N9V88y15HxL4YN

* docs(server): correct the timeline comment BUG-2803 falsified (CONVE-23)

The entryID fallback's comment said note and decision ids "come from the
item's fields blob and nothing validates them on write". BUG-2803 made
that half false: the HTTP API now refuses a request body whose strings
decode to a NUL, including one nested inside a JSON-encoded `fields`
string. The sentence was true when written and nothing in this branch's
diff pointed at it.

The fallback still has to exist, and the corrected comment says why:
the STORE has no such check, so rows predating the rule — and anything
writing a blob by another path, a migration, an import, a future
non-HTTP writer — can still carry one.

SWEPT AND DELIBERATELY LEFT: two nearby comments
(handlers_timeline_id_collision_test.go, handlers_timeline_structured_test.go)
also say "nothing validates them on write". Both are about id FORMAT and
DUPLICATION — an imported artifact carrying a UUID-shaped id, a
hand-written blob repeating one — and this change validates neither. In
context those sentences remain true, so they are left alone rather than
edited into noise.

Sweep command: grep -rniE "nothing validates|not validated on write|no
validation on write|unvalidated" --include=*.go internal/ cmd/ — six
further hits, all about other subjects (github_pr raw writes, terminal
schema keys, push payload format, decodeJSON's size bound).

Claude-Session: https://claude.ai/code/session_011T365kP1N9V88y15HxL4YN

* fix(server): scope the nested-NUL walk to JSON-encoded fields (BUG-2803)

Codex round 2 on #1220. The nesting check from the previous commit
recursed into ANY string that parsed as a JSON document, on the argument
that a structural test beats a destination-typed one. That argument was
wrong in a way I had written down as an accepted trade and should have
weighed as a defect: a plain-text `content` value holding a JSON snippet
that merely MENTIONS the escape was accepted before this branch, is
stored in a text column that has no problem with it, and was newly
refused — including on RE-IMPORT of an export carrying it.

Refusing input the server itself produced is a worse failure than the
door the unscoped recursion was closing. Measured before the fix: a
workspace whose item content held such a snippet exported 200 and
re-imported 400.

The walk now descends only under keys whose STRING value is a JSON
document something downstream re-parses: config, events, fields,
metadata, phase_data, plan_overrides, schema, settings, tags, traits.

WHY A LIST IS SAFE HERE, when ValidateQuery's comment rejects exactly
this shape for query parameters: there the set of names is unbounded by
design (parseItemListParams turns any unrecognised parameter into a field
filter), so no list could be complete. Here the set is a closed property
of the wire model — a field is JSON-encoded because a Go struct declares
it as a string holding JSON — and
TestJSONEncodedFieldKeysCoversTheModels derives it from internal/models
and fails when a new one appears. The list cannot go stale in silence.

Over-inclusion is the safe direction and the list takes it: a listed key
that is not really JSON-encoded costs one parse attempt and can only
refuse a complete JSON document carrying the escape, while a missing key
reopens a door. `traits` is listed for that reason — it carries JSON but
its declaration has no comment saying so, which is exactly how the
derivation test would have missed it, so the test asserts coverage in one
direction only and the list is allowed to be a superset.

The walk also changed shape: decoding into `any` and walking the value,
rather than a token stream, because key context is needed to know which
subtree is JSON-encoded. The []byte reasoning is unchanged and still
holds — decoding into `any` never produces a []byte, so a base64 field is
seen as its ASCII text rather than as decoded bytes that might contain a
legitimate 0x00.

Tests: text fields carrying a JSON document are ACCEPTED (five keys),
with a leg proving the same document under a JSON-encoded key is still
refused, so the pair differs only in the key; the derivation test; and
the depth-bound fixture now nests under a JSON-encoded key at every
level, since nesting under an ordinary key would never start the
recursion and would have passed for the wrong reason.

STILL OPEN, and the lead holds it: a LEGACY row whose stored fields blob
already carries the escape still exports 200 and re-imports 400. That is
data this fix cannot make importable without weakening the write-side
refusal, and the disposition (repair sweep, flagged import, or documented
acceptance) is a product ruling. Recorded on the item.

Claude-Session: https://claude.ai/code/session_011T365kP1N9V88y15HxL4YN

* fix(server): close the three body doors codex round 3 found (BUG-2803)

All three verified before fixing, none taken on the reviewer's word.

1. BUNDLE IMPORT BYPASSED THE REFUSAL (P1). handlers_import_bundle.go
parses pad-export.json itself rather than through decodeJSON, so the
SAME workspace import — reached with Content-Type application/gzip
instead of application/json — walked straight past the NUL check into
Postgres. The bundle's export blob is now checked with bodyDecodesNUL
before ImportWorkspace, answering the same 400. Test drives a real
tar.gz through the router with a clean-bundle control leg, because this
path answers 400 for a dozen unrelated reasons (bad gzip, out-of-order
tar, duplicate entries) and a bare 400 would prove nothing.

2. ONE CALLER SWALLOWED THE NEW ERROR (P2). handlers_admin.go's
test-email endpoint read `if err := decodeJSON(...); err != nil ||
input.To == ""` and fell back to the admin's own address, so a body
carrying a NUL answered 200. An ABSENT body legitimately means "send it
to me"; a body that is present and REFUSED is a different thing, and
collapsing the two turns a validation error into a success. The two
cases are now separated on errors.Is(err, io.EOF).

3. THE COMPLETENESS TEST COULD NOT SEE PAST TWO CALL SHAPES (P2). It
scanned for json.NewDecoder(r.Body) and io.ReadAll(r.Body), so it was
blind to io.ReadAll(io.LimitReader(r.Body, n)) — a shape ALREADY in the
package — and to any alias or helper. A completeness test that misses a
live example is worse than none, because it reads as coverage. It now
scans for the thing that cannot be spelled around, a reference to the
request body at all, and requires every FILE touching one to be
accounted for with a written reason. Both directions are asserted: an
unaccounted file fails because a door may have opened, and an accounted
file that no longer touches a body ALSO fails, so the list cannot rot
into stale excuses that quietly cover a future reader. Verified with a
positive control (an added body reference in an unlisted file fails) and
a negative one (a stale entry fails).

FOUND BY THAT WIDENED SWEEP, and fixed here rather than filed: the raw
artifact import (POST /workspaces/{ws}/import-artifact) takes TEXT, not
JSON, so it never went through decodeJSON and inherited neither the NUL
refusal nor the path/query rule — a body is neither. A raw NUL or
invalid UTF-8 reached the store and Postgres answered 22021, which the
handler turned into a 500 for what is a client error. It now applies
bindableText, the same predicate ValidatePath and ValidateQuery use, and
answers 400 invalid_body. Note the shape difference from the JSON half:
there the ESCAPE is the vector because a decoder rejects a raw NUL;
here the RAW BYTE is, because nothing is in the way.

Each fix has a mutation run against it: disabling the bundle guard fails
the bundle test, disabling the artifact guard fails the artifact test,
and both controls still pass.

Claude-Session: https://claude.ai/code/session_011T365kP1N9V88y15HxL4YN

* fix(server): the escape gate was unsound, and YAML has its own (BUG-2803)

Codex round 4, two P1s, both reproduced before fixing.

1. THE FAST PATH LET A REAL NUL THROUGH. bodyDecodesNUL gated on "does
the raw body contain the six-character escape". That is unsound: the
BACKSLASH itself can be written as an escape, so a body carrying
\u0000 contains no literal six-character sequence anywhere in its
raw bytes, while the OUTER decode manufactures one inside the string —
and if that string is re-parsed as a JSON document (jsonEncodedFieldKeys)
the second parse turns it into a real NUL.

Measured through the real router before the fix: the oblique spelling
answered 201 where the direct one answered 400.

The mistake was applying a fact about how a NUL is spelled INSIDE a
decoded string to the RAW BYTES, where the backslash can itself be an
escape. That is the same layer-confusion this whole bug is made of, for
the third round running.

The gate is now a BACKSLASH. Every JSON escape mechanism requires one, so
a body with no backslash has decoded strings byte-identical to its raw
bytes, and a raw NUL cannot survive the decoder — no backslash therefore
means no NUL, at any depth, however spelled. Bodies WITH one pay for an
exact answer, a larger set than before (any nested JSON carries a
backslash-quote), which is the cost of being correct. The same
correction applies to the per-string pre-filter one level down.

2. YAML HAS ITS OWN ESCAPE VOCABULARY. The raw bindableText check added
last commit passes a double-quoted scalar `title: "a\0b"` — no NUL in
the request bytes — and the YAML decode manufactures one. Measured
before the fix: that artifact imported 201 with a NUL in the item title.
The decoded artifact is now checked too: title, body, and every
frontmatter field value, walked because a playbook's `arguments` is a
nested structure rather than a scalar. Keys are checked as well as
values, on the same precautionary grounds ValidateQuery states for
query parameter names.

Same shape as the JSON half in both cases: a value that is harmless
until a SECOND parse, checked at the layer that can see it.

Tests: the oblique spelling joins the nested-document table, and the
YAML escape joins the artifact table. Each is mutation-verified —
reverting the gate to the substring fails the oblique case only, and
disabling the post-decode artifact check fails the YAML case only, with
the raw-byte cases still killed by the raw check. That per-leg
discrimination is the point: it shows each check earns its own keep
rather than being covered by its neighbour.

Prose corrected where this falsified it: jsonNULEscape's "it is the ONLY
spelling" is true of the escape and was being used to justify a filter on
the raw bytes, which is a different claim. Both now say so explicitly.

Claude-Session: https://claude.ai/code/session_011T365kP1N9V88y15HxL4YN

* fix(server): multipart text fields and the bundle manifest (BUG-2803)

Codex round 4's two P2s. Both are the same shape as the rest: a
caller-supplied string reaching a text comparison through a door the
earlier fixes did not cover.

1. MULTIPART TEXT FIELDS. The multipart body is deliberately exempt from
the JSON rule — its payload is binary blob content and must not be
scanned for text validity — but its TEXT fields are a different thing.
`item_id` goes to ResolveItem and into a database comparison exactly as
the query-string channel does, and that channel has been validated at
the transport since BUG-2784; the form channel was not. multipartValues
now drops values that are not bindable text, which makes an unusable
value indistinguishable from an absent one — the disposition
resolveUploadItemID already applies to empty values.

The uploaded FILENAME gets the same predicate, with a fallback to a
generic name rather than a refusal: the bytes are fine, only the label
is unusable.

A NEGATIVE RESULT worth recording, because it changed the test: a RAW
NUL in the multipart header is NOT the vector. Go's multipart reader
refuses it as a malformed MIME header line before any handler sees it
(measured: 400, "malformed MIME header line"). The reachable spelling is
the RFC 5987 encoded form, filename*=UTF-8''sh%00ot.png, which the
header parser accepts and percent-decodes afterwards. The first version
of this test used the raw form and was testing a vector that does not
exist.

2. THE BUNDLE ATTACHMENT MANIFEST. A second JSON document inside the
tar.gz, parsed directly like pad-export.json was, so it needed the same
check. Without it a NUL in a manifest string reached
rehydrateAttachment, whose failure is logged and SKIPPED — so the import
reported success while silently dropping the attachment. The
skip-on-failure behaviour is pre-existing and deliberate (a partial
restore beats none); refusing the bad INPUT is what stops it being
reached this way. Left as it is, and named rather than quietly changed.

A VACUOUS ASSERTION THE MUTATION CAUGHT, recorded because the test would
otherwise have shipped as coverage: the filename leg first asserted
`!strings.ContainsRune(body, 0)` on the RESPONSE, which is JSON — a NUL
in the filename comes back as the six-character escape, not as a 0x00,
so the check passed whether or not the fix was present. It did pass with
the fallback disabled. Now it decodes the response and asserts the
replacement name. The item_id leg had the mirror-image weakness: it
asserted "not a 500", which is the Postgres-only symptom, so on SQLite it
would have passed either way; it now asserts the request behaves exactly
like the no-value control.

Every fix in this commit has a mutation against it, and each kills only
its own leg.

Claude-Session: https://claude.ai/code/session_011T365kP1N9V88y15HxL4YN

* refactor(server): drop the now-unused escape constant (BUG-2803)

The gate became a backslash check, which was the last production use of
jsonNULEscape; golangci-lint's unused check failed on the next run. Its
documentation was load-bearing, so the explanation moved into
bodyDecodesNUL's comment rather than being deleted with the variable —
including the distinction that made the old gate wrong (the escape has
one spelling INSIDE a decoded string, which is not a claim about the raw
bytes).

Caught by re-running lint on the tip after the previous commit rather
than trusting the run from the tip before it.

Claude-Session: https://claude.ai/code/session_011T365kP1N9V88y15HxL4YN

* fix(server): rune-safe truncation and User-Agent sanitising (BUG-2803)

Codex round 5 was asked for the POPULATION rather than a confirmation —
"enumerate every remaining way a caller-supplied string can reach a
database text or jsonb parameter without passing a validity check" — and
returned three residual classes with their sinks. Two are fixed here;
the third is filed, because measuring it needs a fixture this unit
should not grow.

1. TRUNCATION CAN UNDO THE VALIDATION. Four sites cut a caller string
with a plain byte slice (name[:120], input.Name[:200]). If the boundary
lands inside a multi-byte rune the result ends in a partial sequence and
is no longer valid UTF-8 — so a value that PASSED the body check a few
frames earlier arrives at the store unbindable, and Postgres answers
22021 for a request the server already accepted.

This is the interesting one, because no input-side round could have
found it: the defect is downstream of validation, and it is invisible
with ASCII fixtures, which is what every test in that area used.
truncateBindableText walks back off continuation bytes and drops the
straddling rune. Tested with 2-, 3- and 4-byte runes so an off-by-one
walk-back cannot pass them all, and with a counterfactual leg asserting
the naive slice really does produce unbindable output for the same
input — without it the cases would pass against an implementation that
did nothing.

2. USER-AGENT REACHES TEXT COLUMNS. It lands in activities.user_agent
(three document paths, the connected-apps revoke) and
sessions.user_agent (three login paths), and no rule here sees a header.
The disposition is SANITISE, not refuse, and that is deliberate: a
header is metadata this server chose to record, not something the caller
asked for, so a malformed one must not turn an otherwise fine request
into a 400. The two sites that HASH the header are left alone — sha256
over arbitrary bytes is well defined, and changing what is hashed would
invalidate every stored UAHash.

The filing's own earlier probe had recorded User-Agent as NOT
reproducing on the item-create path. That was true and did not
generalise; these are different sinks.

3. NOT FIXED, FILED: the OAuth form-encoded bodies
(/oauth/token, /oauth/authorize/decide, /oauth/revoke,
/oauth/introspect) parse url-encoded form data outside the shared body
validator, with connection_name reaching oauth_connections.name and
client_id reaching the oauth_clients.id lookup. This was the ORIGINAL
subject of BUG-2803 before the filing was re-scoped, and it was recorded
then as unreachable without a fosite-backed fixture. That is still true,
and round 5's sink list is far more than the filing had. Filed rather
than guessed at.

Claude-Session: https://claude.ai/code/session_011T365kP1N9V88y15HxL4YN

* fix(server): narrow the gate, stop refusing natural-shape fields (BUG-2803)

Codex round 6 plus one measurement of my own. Three changes, one of them
a revert of something I got wrong in the previous commit.

1. THE GATE COST TOO MUCH, so it is narrower and still sound. The
previous commit gated the walk on "does the raw body contain a
backslash", which is correct but catches every body carrying nested JSON
(each `\"` is a backslash). Measured on a ~377 KB import-shaped body:
60106 allocs/op with that gate versus 30073 with the walk disabled — the
walk was running on ordinary traffic.

The gate is now the four bytes that begin any \u escape for a character
below U+0100. The argument: to manufacture the six-character NUL escape
inside a decoded string, each of its characters arrives either literally
from the raw bytes — in which case the raw contains the escape, which
begins with that prefix — or from a \u escape of its own, and the three
characters involved (backslash U+005C, 'u' U+0075, '0' U+0030) all sit
below U+0100, so those escapes begin with it too. Back to 30073
allocs/op, identical to the walk-disabled build.

That argument is the same KIND of reasoning that was wrong two rounds
ago, so it does not stand on its own: a differential test runs the gated
function against an UNGATED walk over a corpus built to attack it —
oblique backslash, upper-case hex, an escaped 'u', an escaped '0', a
doubled backslash — and fails on any disagreement. It also asserts the
corpus contains both answers, since agreement over a one-sided corpus
would be vacuous. Reverting the gate to the old substring fails it.

2. THE CHECK REFUSED THE NATURAL SHAPE OF ITS OWN FIELDS. `tags` and
`fields` accept both a JSON-encoded STRING and their natural array/object
form, and the walk propagated "this subtree is JSON-encoded" into
containers — so a free-form tag whose whole value happened to be a JSON
document was refused, though nothing re-parses it. Measured: refused
before, accepted now, while the JSON-encoded spelling of the same field
is still refused. The flag now marks only a direct STRING child of a
listed key.

3. REVERTED: I wired the three LOGIN paths to the User-Agent sanitiser
last commit, before reading store.CreateSession. It HASHES the header
and stores no text — the round-5 enumeration named "sessions.user_agent"
and I took the name for a column. The change would have been actively
harmful: login would store sha256(sanitised) while middleware_auth still
compares sha256(RAW), so every session from a client with a non-UTF-8
User-Agent would fail validation. A sink named in a review is a pointer
to verify, not a finding. The real sink is activities.user_agent, from
three document paths and the connected-apps revoke.

4. And the wiring leg codex asked for, on that real sink: a request
through the router with a malformed header, reading the STORED value out
of the activities row, with a control asserting an ordinary header is
kept VERBATIM. Unwiring the production call site fails it; the helper's
unit test does not notice.

Claude-Session: https://claude.ai/code/session_011T365kP1N9V88y15HxL4YN

* fix(server): apply the key rule at every level, not once (BUG-2803)

Codex round 7, both findings, the first confirmed by measurement.

1. THE RECURSION WENT ONE LEVEL TOO DEEP. Once the walk descended into
a JSON-encoded string it treated the WHOLE subtree below as
JSON-encoded, so a value nested two levels down — an ordinary string
inside a `fields` blob that happens to hold JSON text — was refused.

That is a false rejection, and the measurement says so plainly. With the
depth-2 check disabled, on Postgres 17:

  depth 1 (the fields blob itself)     -> 400   (correct: Postgres parses it)
  depth 2 (a string INSIDE the blob)   -> 201   (accepted, no error)
  control                              -> 201

The handler parses `fields` ONCE. The inner text is re-escaped when the
blob is written, so what Postgres receives has a doubled backslash and no
escape at all. Only the document Postgres itself parses can carry a fatal
one.

The nested call now passes false rather than true, which makes this a KEY
RULE APPLIED AT EVERY LEVEL rather than a depth limit: a JSON-encoded key
INSIDE a document still recurses (pinned by a test), an ordinary one does
not. Same correction as round 6's natural-shape fix, one level further in
— I fixed the sibling case and left this one, which is CONVE-18's lesson
about my own enumeration being a sample too.

I checked whether anything re-parses a value inside the blob before
loosening this, rather than assuming: `arguments` was the candidate, and
parsePlaybookArguments asserts it is a native ARRAY (raw.([]any)) rather
than a JSON string, so it is covered by the natural-shape rule and needs
no second parse.

2. AN ERROR MESSAGE THAT SENT CLIENTS THE WRONG WAY. The OAuth dynamic
client registration handler prefixed every decode failure with "Request
body must be JSON". A body carrying a NUL is valid JSON, so that message
sends a client hunting a syntax error it does not have. The two failures
are now distinguished.

Round 7 also reports no break in normal CLI, MCP or web-client request
generation — they marshal JSON and encode paths and query parameters —
which is the first thing any round has said about the client surface.

Claude-Session: https://claude.ai/code/session_011T365kP1N9V88y15HxL4YN

* fix(server): complete the artifact check, make the walk path-aware (BUG-2803)

Codex round 8. It confirmed round 7's two fixes, then found two real
defects and two inaccurate comments — the comment half being the angle
the round was asked for.

1. THE ARTIFACT CHECK MISSED TWO REACHABLE FIELDS. artifactIsBindableText
walked the decoded artifact by TYPE, so it never covered Provenance —
whose strings are rendered into a Markdown footer appended to the stored
content — and never matched Arguments, declared []map[string]any, a
concrete slice type the walk's []any case does not match. A YAML NUL
escape in either reached storage.

It now MARSHALS the artifact and searches the output for the escape
encoding/json produces. A type switch over a struct that grows is a list
that goes stale in silence; marshalling covers every exported field,
including ones added later. The one thing it cannot see is invalid UTF-8
(which marshals to U+FFFD), and it does not need to: step 2 rejects that
in the request bytes, and YAML cannot manufacture it from valid input —
its escapes name code points, where \0 names a NUL.

Both new cases fail with the check disabled; the raw-byte cases still
pass, killed by the raw check, so each leg is discriminating.

2. THE WALK WAS NOT PATH-AWARE. A collection may declare a user field
literally named `schema` or `tags`. The walk consulted the wire-key list
at every level, so `{"fields":{"schema":"..."}}` treated a user field
name as a wire key and refused valid text holding a JSON example.

The key list is now consulted only OUTSIDE caller data — not under a
natural `fields` object, not inside an element of a `tags` array, not
inside a re-parsed document. Combined with round 7's fix that makes the
descent exactly one level deep BY CONSTRUCTION, which is why the depth
counter is gone: with the flag no longer inherited, a bound could never
fire, and dead protection reads as protection. The depth-bound test is
replaced by one that pins the property directly — an escape IN the
parsed document is refused, one BELOW it is accepted, and a
wire-key-shaped user field does not restart the descent.

3. THREE COMMENTS CORRECTED, all mine, all of the kind a reader would
believe without checking:

- MaxBytesReader: Close FORWARDS to the underlying body rather than
  being a no-op, and with a nil writer there is no automatic 413 — the
  cap surfaces as a read error the callers turn into 400. Behaviour
  unchanged; only the claim was wrong.
- parseArtifactRequest said "three checks" while implementing five, and
  its returns list omitted ErrArtifactUnbindableText. Both added by this
  branch, which is exactly the prose a change is most likely to falsify
  (CONVE-23).
- errJSONBodyNUL claimed all 65 callers surface its message. The STATUS
  is uniform; the wording is not — several substitute a generic string.

4. And one in a test: the timeline fixture said both backends hold a
CHECK constraint a raw NUL violates. items.fields is a plain TEXT column
with no CHECK on SQLite. What was OBSERVED is "SQL logic error:
malformed JSON"; the likely source is an expression index over
json_extract, and that attribution is recorded as NOT verified rather
than asserted.

Claude-Session: https://claude.ai/code/session_011T365kP1N9V88y15HxL4YN

* fix(server): a regression this branch introduced, and the same trap again (BUG-2803)

Codex round 9. Both findings are mine, one of them a regression from the
round-8 restructure two commits ago.

1. THE ROUND-8 RESTRUCTURE REOPENED THE ORIGINAL DOOR. Taking the
JSON-encoded branch for a listed key skipped the plain "does this string
contain a NUL" check and asked only "does the document this string
carries hold an escape". Those are different questions. So
{"fields":"a<NUL escape>b"} — a direct NUL in the fields value, the very
first case this whole change closed — was accepted again.

Both checks now run. The test pins all three legs: a direct NUL in the
fields string, an escape inside the fields document, and an ordinary
fields string that must still be accepted, so the first two cannot pass
merely because everything under a listed key is refused.

2. THE ARTIFACT CHECK FELL INTO THE TRAP IT WAS WRITTEN AGAINST. It
searched the MARSHALLED bytes for the escape sequence, and a value
holding the six LITERAL characters marshals to a doubled backslash which
still contains that sequence as a substring — so valid content was
refused. Artifacts are documentation; text about a JSON escape is
exactly what one carries.

Worse than the bug: the comment I wrote asserted the ambiguity "cannot
arise here". It was the same doubled-backslash case bodyDecodesNUL exists
to resolve, one function away, and I wrote a sentence explaining why it
did not apply instead of checking. The marshalled form is now decoded
again and walked with the same machinery — the round trip is what makes
every field reachable without a type switch, the walk is what makes the
answer exact.

Its test asserts literal escape TEXT is accepted in title, body and a
field value, with a counterfactual leg asserting a real NUL in each of
those places is still refused, so acceptance cannot come from the check
doing nothing.

Reverting either fix fails its test and only its test.

Claude-Session: https://claude.ai/code/session_011T365kP1N9V88y15HxL4YN

* docs(backup): the one case where an export is not importable (BUG-2803)

Codex round 12, an operational pass. It found no migration or config
requirement, and two documentation gaps.

docs/backup.md promises that application-level export/import is portable
across SQLite and PostgreSQL. Since BUG-2803 that has one exception: a
workspace whose stored data contains a NUL exports fine and is refused on
import. It can only affect data written before the rule existed and only
on SQLite, which accepted it — a PostgreSQL instance never stored one.

`pad db migrate-to-pg` has the SAME problem and reports it worse: it
copies rows directly and never passes through the import guard, so a
legacy row fails against PostgreSQL's JSONB parser partway through the
copy rather than being refused up front. That is the likelier way an
operator meets this, since it is the operation that puts an entire old
SQLite database in front of PostgreSQL for the first time. Recorded on
BUG-2810, which owns the preflight and repair.

Round 12's other finding — that the PR's stated release note covered the
JSON 500-to-400 change and none of the rest — is fixed in the PR body
rather than in the tree.

Claude-Session: https://claude.ai/code/session_011T365kP1N9V88y15HxL4YN

* test(server): close two blind spots the tests themselves had (BUG-2803)

Codex round 13, asked whether the new TESTS are sound. Five findings;
these are the two that were self-contained. The other three are recorded
on the item with what each needs.

1. THE COMPLETENESS SCAN WAS BLIND TO FORM BODIES. It matched only
`.Body`, so FormValue / ParseForm / MultipartForm — which read the
request body just as surely — were invisible. It therefore reported full
coverage while the OAuth form-encoded handlers were entirely outside its
view. Widened, and it immediately failed on handlers_oauth.go, which is
the instrument working.

That file is now ACCOUNTED FOR AS A KNOWN GAP rather than as safe: the
OAuth handlers read form-encoded bodies that no rule in this family
covers (the transport rules see the query half of r.Form, not the body
half), tracked as BUG-2811 and needing a fosite-backed fixture to
measure. The test now STATES the gap instead of being blind to it, which
is the difference between a completeness claim and a completeness
appearance.

2. THE TRUNCATION TEST ADMITTED AN IMPLEMENTATION THAT RETURNED "". Its
assertions were: within the limit, bindable text, a prefix of the input.
An empty string satisfies all three. It now also asserts that an input
fitting the limit comes back UNCHANGED, and that no more than one rune
(4 bytes) is lost to the boundary — so a truncator that drops too much
fails, not just one that keeps too much.

Both were found by asking whether a broken implementation would pass,
which is the question CONVE-12 is about and which I had applied to the
production code and not to these two tests.

Claude-Session: https://claude.ai/code/session_011T365kP1N9V88y15HxL4YN

* test(server): the three remaining round-13 gaps (BUG-2803)

Codex round 13's other three findings, all of the same shape: a test
that would stay green with the production change reverted.

1. THE MANIFEST CHECK WAS UNTESTED. The bundle test built archives
containing only pad-export.json, so disabling the INDEPENDENT attachment-
manifest check left the suite green. The new test builds a bundle with
both entries, differing only in the manifest, so a refusal cannot come
from the export half. Verified by disabling each check separately: only
the matching test fails, so the two are independently covered.

2. THE TEST-EMAIL CHANGE HAD NO HANDLER-LEVEL TEST. Every existing leg
exercised decodeJSON, so reverting handlers_admin.go to default EVERY
decode failure to the admin's own address passed them all. The new test
drives the real endpoint with a wired mock sender and pins the
distinction that used to collapse: an ABSENT body still means "send it
to me" (control), an ordinary body still sends (control), and a body that
is present and refused answers 400 rather than being reinterpreted as
the default recipient.

3. THE MULTIPART LEG CHECKED ONE BYTE CLASS. A filter rejecting NULs
while letting malformed UTF-8 through would have passed it. It now drives
both, which matters because invalid UTF-8 is the class that reaches
Postgres as 22021 on a UTF8 database.

Round 13 was asked whether the new TESTS are sound — deterministic,
order-independent, and failing on broken code. It reported the fixtures
isolated and found five ways they were not discriminating. Two were
fixed in the previous commit; these are the rest.

Claude-Session: https://claude.ai/code/session_011T365kP1N9V88y15HxL4YN

* test(server): pin the wiring at every call site, not one (BUG-2803)

Codex round 14 confirmed round 13's five, then found the same shape one
level out: reverting a SINGLE call site back to the unsafe form left the
whole suite green, because the surviving fixtures are ASCII and a
helper's unit test does not care who calls it.

TestTextSafeHelpersAreUsedAtEveryCallSite asserts the wiring STATICALLY
rather than adding a fixture per site (an OAuth connection, a cloud
login, four audit paths). A byte-slice truncation of a caller string
fails it, and so does a raw User-Agent read outside the exempt set. Both
directions are checked: finding none of the SAFE form also fails, so a
scan that silently matched nothing cannot pass forever.

The User-Agent exemptions carry counts rather than being blanket, so a
NEW raw read in an exempt file still fails. All four reads in
handlers_auth.go are exempt because they feed a HASH — CreateSession
hashes the header and stores no text — and sanitising before hashing
would be actively harmful: login would store sha256(sanitised) while the
session check still hashes the RAW header, failing validation for every
client with a non-UTF-8 User-Agent. middleware_request_text.go's one raw
read is requestUserAgent itself.

Verified by reverting one truncation call site and one User-Agent call
site independently; each fails the test.

Round 14's third finding is fixed behaviourally rather than statically,
because the static scan cannot see it — handlers_oauth.go is already
listed for its form-body reads. TestOAuthRegisterRefusesNULBody drives
the real dynamic-registration endpoint with cloud mode and an OAuth
server wired, with a control leg registering successfully, and pins both
the refusal and the message split: the body IS valid JSON, so the answer
must not send a client hunting a syntax error.

Claude-Session: https://claude.ai/code/session_011T365kP1N9V88y15HxL4YN

* fix(server): match wire keys the way the decoder does (BUG-2803)

Codex round 16, asked whether this change is consistent with its siblings
in the same file and extensible by someone who did not write it. It found
a live bypass instead.

encoding/json matches an incoming key to a struct field by an exact match
first and a CASE-INSENSITIVE one otherwise, so {"Fields":...} and
{"FIELDS":...} land in ItemCreate.Fields exactly as {"fields":...} does.
The walk looked the key up case-SENSITIVELY, so it skipped the nested
document for a body the handler went on to accept, and the database
answered the original 500.

Measured before the fix: `fields` refused, `Fields` and `FIELDS`
accepted.

This is the same defect shape as everything else in this unit — a check
that agrees with one layer's rules while the layer that actually consumes
the value uses different ones — which is why the fix is a PREDICATE
rather than a wider map: the map is the vocabulary, and the matching RULE
belongs to the consumer. Someone adding a key should not also have to
remember to add its spellings.

The test drives six spellings including mixed case, with a control
asserting an unlisted key stays caller data in any casing, so this is
case-insensitive matching rather than matching everything.

Claude-Session: https://claude.ai/code/session_011T365kP1N9V88y15HxL4YN

* fix(server): fold keys the way encoding/json folds them (BUG-2803)

Codex round 17, first of five findings. The previous commit fixed the
ASCII half of key matching and left the Unicode half, which is this
bug's own pattern one more time.

encoding/json matches with Unicode SIMPLE FOLDING, not lower-casing.
U+017F LATIN SMALL LETTER LONG S folds to 's', so "ſchema" reaches the
`schema` struct field while strings.ToLower("ſchema") is unchanged and
missed the allowlist — a nested NUL under that spelling reached the
handler undetected.

Matching is now strings.EqualFold against each canonical key. The test
carries both a lower-case fold spelling and an upper-case one alongside
the ASCII cases, and keeps its control asserting an unlisted key stays
caller data in any casing.

The other four round-17 findings are recorded on the item rather than
patched here: they are genuine layer disagreements (duplicate keys
merging differently in a typed decode than in a map, a scan-failure
disposition on inputs the typed decode tolerates, and unknown-field
policy) whose fixes are design decisions rather than corrections, and
this seat is near its context bar. Each is written up with the
measurement it needs.

Claude-Session: https://claude.ai/code/session_011T365kP1N9V88y15HxL4YN

* fix(server): pin that a NUL-bearing manifest refusal keeps the partial workspace (BUG-2803)

Codex round 18. The comment on the manifest NUL branch said refusing the
input "stops it from being reached this way" and stopped there, which
reads as though the refusal undoes the import. It does not.

A plain error with a non-nil workspace keeps the partial workspace,
exactly as every other manifest failure in this loop does — the rollback
branch fires only for *importStatusError, and mid-stream manifest
failures intentionally keep what was imported (TASK-896). Returning a
rollback-shaped error here would give NUL-bearing manifests different
semantics from malformed ones, which is a change to the bundle-import
contract rather than a fix to this bug.

So the behaviour is unchanged and now DELIBERATE: the comment states it,
and the test asserts the persisted state rather than only the HTTP
answer. Mutation: routing the branch through *importStatusError makes
the refusal roll back, and the new assertion fails naming the release
note it would falsify. The pre-existing status/body assertions do not
notice.

Claude-Session: https://claude.ai/code/session_01AUvLoXsKdS5sdpYju6rj4p

* docs(server): record the four map-model disagreements as dispositions, and pin them (BUG-2803)

Lead ruling day-68 is land-and-follow: this branch lands on its measured
commits, and the token-stream rewrite is the BUG-2812 unit's spec rather
than a late restructure of an 18-commit branch under review pressure.
That makes the four open findings from rounds 16-17 something to WRITE
DOWN precisely, not something to leave in a trail comment.

The doc comment on bodyDecodesNUL now carries all four, with the one
root cause named: this scan decodes into map[string]any and the typed
decode does not agree with that model about keys. Two under-refuse
(duplicate-key merge; scan-failure passthrough) and are BUG-2812's spec
- both dissolve under a walk that never builds values. Two over-refuse
(unknown fields; case-variant duplicates) and are ACCEPTED, because
refusing is the safe direction. The asymmetry is stated rather than
smoothed over: within the map model, (1) and (4) are one defect seen
from two sides and only one of them fails safe.

Finding (3) is an observable compatibility change - a forward-compatible
field carrying a NUL escape now gets a 400 where it got a 200 - so it
goes in the release note as well as here. A qualification only protects
where the actor meets it.

All four are pinned by a test, measured on this tip rather than carried
over from the round-16/17 write-up. The two known-gap legs assert the
WRONG answer on purpose: when BUG-2812 lands they FAIL, naming the doc
comment and the release note as what to update. Both gap legs carry a
premise assertion - the same bodies with the disagreement mechanism
removed ARE detected - without which they would pass against a check
that detected nothing.

Claude-Session: https://claude.ai/code/session_01AUvLoXsKdS5sdpYju6rj4p

* test(server): wire release-note item 10 to the router, with its before-state measured (BUG-2803)

The disposition test proves bodyDecodesNUL RETURNS true for an unknown
field carrying a NUL escape. The release note claims the API answers
400. Those are different claims and only the second one is what an
operator or client author reads - CONVE-19, my own convention: a
direct-call test vouches for the component, not its binding.

Two legs, and the control is the load-bearing one. An unknown field with
an ordinary value must still be ACCEPTED, so this pins "refused for the
NUL" rather than "refused for being unknown". The handler does not
reject unknown fields; if it ever started to, the note's explanation
would be wrong while its status code stayed right, and no
status-code-only assertion could see that.

The before-state is measured rather than asserted from memory. Disabling
the check makes the same request answer 201 - which is main's behaviour,
since decodeJSONWithLimit there unmarshals straight into the typed value
and the key is dropped. So "answers 400 where it answered 200" is a
measurement in both directions, not a recollection of one.

Claude-Session: https://claude.ai/code/session_01AUvLoXsKdS5sdpYju6rj4p

* docs(backup): the NUL rule lives in the binary, not the database (BUG-2803, BUG-2813)

Codex round 19, the fresh-angle deploy/rollback/mixed-version pass.

docs/backup.md said a NUL-bearing row "can only affect data written
before that rule existed, and only on SQLite". The second half is true.
The first half is false, and the reason is the interesting part: the
guard is in decodeJSONWithLimit, so the invariant is a property of the
running BINARY, not of the database.

On SQLite any window where an older binary serves the same database can
still write one - a rollback after upgrading, a staged rollout with an
old and a new instance sharing a database, a second older instance on
the same file. The window closes, the guard returns, and the rows are
already stored, behaving exactly like genuinely old ones. A rollback is
an ordinary operational move, so this is not an exotic path.

The doc now states the binary-version dependence, says which dialect is
affected and why PostgreSQL is not (it refuses a NUL itself, at every
version), and gives the operational answer: drain writes from older
binaries before the new one serves, or roll forward rather than back.

Store-layer enforcement - so the running build stops mattering - is
filed as BUG-2813 rather than added here. It is a dialect-level change
and the day-68 ruling on this unit is land-and-follow.

The same false implication was carried by the PR's release note calling
such a workspace "legacy"; corrected there too.

Claude-Session: https://claude.ai/code/session_01AUvLoXsKdS5sdpYju6rj4p

* docs(server): cite the ruling in house style, not the team-room day counter (BUG-2803)

"lead ruling day-68" is the internal day counter, which means nothing to
anyone reading this repo and is inconsistent with every other citation
in it - the codebase cites a lead ruling by DATE or by BUG ref, never by
day-N. Replaced with the bug ref, which is the part a reader can
actually follow.

Claude-Session: https://claude.ai/code/session_01AUvLoXsKdS5sdpYju6rj4p

* docs(server): drop a commit count I had already measured as wrong, and stop asserting a cause I borrowed (BUG-2803)

Two defects in a comment I wrote an hour ago, both of the kind this
unit's trail keeps recording.

"an 18-commit branch" - the branch was 20 commits at b0192871 when I
counted it this session, and is more now. 18 came from the previous
checkpoint's own miscount, which I had ALREADY identified and written up
before I typed it again here. A number that arrives inside a sentence
about something else does not feel like a claim, which is exactly why it
survives. The count is incidental to the argument, so it is gone rather
than corrected - a figure that has to be maintained to stay true is a
liability in a doc comment.

"this branch's one regression came from exactly that" - the ruling's
reasoning, restated by me as a verified fact. The regression I know
about came from wiring a fix off a reviewer-named sink list without
reading the mechanism, which is adjacent to "restructuring late under
review pressure" but is not the same mechanism, and I did not check
whether it is the one the ruling meant. Now attributed to the ruling and
stated as its reasoning, with the part I can defend - the review loop
finding something in nearly every round indicates a design problem -
carrying the argument.

Claude-Session: https://claude.ai/code/session_01AUvLoXsKdS5sdpYju6rj4p

* fix(server): sanitise the MCP audit tool_name, and correct three claims wider than their evidence (BUG-2803)

Codex round 20, asked for a POPULATION rather than a confirmation
(CONVE-24). It returned a covered list AND four findings; this commit
carries the two that belong to this unit plus the doc corrections.

## The door: MCP audit is a second reader, not a pass-through

parseMCPRequestBody runs its OWN json.Unmarshal and binds the decoded
method / params.name to mcp_audit_log.tool_name, TEXT NOT NULL. A
six-character NUL escape therefore arrives as a real NUL: PostgreSQL
refuses the audit INSERT with 22021 - the exact symptom this unit exists
to remove - and SQLite stores an unprintable tool name. Nothing upstream
catches it; the /mcp transport decodes the JSON-RPC envelope itself
rather than through decodeJSON, so the body rule never sees the request.

Measured before fixing: the decoded name reached the column intact.

This unit's own completeness map had CERTIFIED that reader as safe, on
the grounds that "decoding still happens in the MCP dispatcher". That is
true and it does not bear on what this middleware persists - a correct
description of a mechanism, with no question asked about what it does,
sitting in the one artifact whose job is to say the population is
covered. Corrected there too.

Disposition is SANITISE, not refuse, following the User-Agent precedent
from earlier in this unit, and the rule now lives in one extracted
helper (sanitiseStoredText) with the reasoning attached: the body rule
refuses because the caller asked to store that value; this serves
metadata the SERVER elected to record, where failing the write would
lose the audit row for precisely the request most worth auditing.

Both caller-derived returns are cleaned inside parseMCPRequestBody, so
both call sites - the ok path and the denied path - are covered at the
choke point rather than at either caller. Both are tested: params.name
AND the method path. Mutations un-sanitising each one compile and kill
only their own leg.

## Three claims corrected, all wider than their evidence

- "all 65 call sites" in server.go: measured 70. Removed rather than
  corrected, because the number has to be maintained to stay true and
  says nothing the sentence needs.
- docs/backup.md said a NUL "cannot be stored in a text or JSON column"
  absolutely, two paragraphs above my own text explaining that SQLite
  accepts one. Now stated as what it is: an application rule Pad
  enforces on both dialects, which is exactly why it has to be enforced.
- artifact_import.go said such a value "cannot be stored under any
  encoding this product supports". Refuses, not cannot - stating a
  policy as a capability tells the next reader SQLite enforces
  something it does not.

## Filed, not fixed

BUG-2814 - guarded writes re-emit at-rest NULs (move/copy/restore/
fields-patch), propagating a legacy value to rows that never had one.
Distinct from BUG-2813: that one is about writing a NUL while an old
binary serves, this is the fixed binary SPREADING one already present.
Both dissolve under the same store-layer enforcement, so they are filed
to be designed together rather than patched at each of a long and moving
list of re-emit sites.

Declined: round 20 also reported the release-note assertions as
unsupported. They live in the PR body, which a read-only sandbox cannot
see - the claim is about the reviewer's visibility, not the diff.

Claude-Session: https://claude.ai/code/session_01AUvLoXsKdS5sdpYju6rj4p

* fix(server): sanitise before testing for emptiness, so the audit fallback survives (BUG-2803)

Codex round 21 ranked this the most dangerous un-probed lens, and it is
a boundary my own round-20 fix created.

parseMCPRequestBody tested env.Method == "" and p.Name == "" BEFORE
sanitising. A value made entirely of NUL escapes is non-empty as
decoded and empty once cleaned, so it passed over the fallback and was
then blanked - storing an empty tool_name in a TEXT NOT NULL column.
That is exactly the silent drop the "(unknown)" / "tools/call"
fallbacks exist to prevent; the function's own doc comment says so.

Measured before fixing: both shapes returned an empty tool_name.

Fixed by ordering rather than by adding guards - clean first, then test
- so the invariant is structural instead of something each return has
to remember. Same by-construction preference as the symmetric-gate fix
earlier in this unit.

Worth recording that my first patch was WRONG in a way that compiled:
I put the sanitise above the json.Unmarshal that populates env, so the
method would always have been empty. Caught by printing the patched
function and reading it, not by trusting the script saying "patched".

Claude-Session: https://claude.ai/code/session_01AUvLoXsKdS5sdpYju6rj4p

* fix(server): classify MCP audit on the raw method, and trim only JSON whitespace (BUG-2803)

Codex round 22. Two P2s, both measured before fixing.

## A forgeable audit row - my own regression from the round-21 fix

The round-21 change reordered sanitise-before-compare so the fallback
would survive an all-NUL value. That reorder made the CLASSIFICATION
read the sanitised method, so "tools/<NUL>call" cleaned up INTO the
literal "tools/call" and the parser then lifted params.name and hashed
the arguments for a method that was never tools/call.

Measured: tool_name="pad_item" with a full 64-character args_hash - an
audit row indistinguishable from a genuine pad_item call, mintable by
anyone who can send a request. Worse than the review described it.

Fixed by splitting the two jobs, which were never the same job:
dispatch decisions read what the client actually SENT; sanitising is
for the value that gets STORED. The round-21 boundary is preserved -
a method empty only after cleaning still falls back to "(unknown)".

Fixing one boundary and creating another in the same function is worth
naming: the reorder was correct for the case it addressed and I did not
ask what else read that value.

## Go whitespace is not JSON whitespace

The empty-body shortcut used bytes.TrimSpace, i.e. unicode.IsSpace,
which strips \v, \f, U+00A0 and more. encoding/json accepts none of
them. So a body of just \v trimmed to empty, returned io.EOF, and an
EOF-tolerant caller - playbook run treats errors.Is(err, io.EOF) as "no
arguments supplied" and runs anyway - took a syntactically invalid body
for an ABSENT one.

Now trims exactly the four bytes JSON calls whitespace. The test drives
both directions, because only the pair discriminates: real JSON
whitespace must still shortcut to EOF or the playbook contract breaks,
and non-JSON whitespace must not or the divergence survives. Reverting
to TrimSpace compiles and fails three legs.

Claude-Session: https://claude.ai/code/session_01AUvLoXsKdS5sdpYju6rj4p

* test(server): give the walker an independent oracle, not one that shares its code (BUG-2803)

Codex round 22, finding 3. TestBodyDecodesNULGateAgreesWithAnUngatedWalk
compares the gated function against an "ungated" reference that calls
the SAME production valueDecodesNUL. That is valid for what the test
claims - it pins the raw-prefix GATE - but it structurally cannot see a
defect in the WALKER, because such a defect is present identically on
both sides and cancels.

That matters here specifically: every walker defect this unit has had
lived in traversal, descent, or key matching (rounds 1, 2, 4, 16, 17),
which is exactly the part the differential cannot check.

Added a second implementation of the contract, written in the test and
deliberately not calling the production walker. It shares encoding/json
and jsonEncodedFieldKeys; it does NOT share traversal, descent, or
key-matching. It is iterative with an explicit stack rather than
recursive, so a recursion-shaped bug cannot reproduce in it by accident.

Demonstrated rather than argued. With the nested-document descent
removed from the production walker - a mutant that reopens the exact
door this unit exists to close, and which compiles:

  differential (gate vs ungated)   ok      <- blind, as the finding said
  independent oracle               FAIL    <- catches it

The corpus is also asserted to contain BOTH answers, since two walkers
that always answer false agree perfectly.

Claude-Session: https://claude.ai/code/session_01AUvLoXsKdS5sdpYju6rj4p

* test(server): make the body-reader inventory type-aware, and state what it still cannot see (BUG-2803)

Codex round 22, finding 4. The inventory that claims every request-body
reader is accounted for was lexical, and wrong in three ways - all in
the direction that matters for a test whose job is to say nothing is
invisible:

  - it recognised only the variable names r and req, so a handler
    holding its request as httpReq or orig was INVISIBLE;
  - it matched inside COMMENTS, so prose could make a file look scanned;
  - the manually-listed traits field was already evidence of the
    model-regex blind spot.

My first fix broadened the pattern to any identifier. That was worse,
and worth recording: it matched every unrelated .Body field - input.Body
in comments, fetched.Body in url import, comment.Body, art.Body,
sidecarErr.Body - flagging five files that read no request body at all.
The only route to green would have been listing those five as
accounted, and an accounting entry HIDES future readers in its file. A
false entry is worse than a missing one, so I abandoned that approach
rather than tuning the regex.

Now keyed on the TYPE via go/ast: collect identifiers declared
*http.Request in a function signature, then find reader selectors on
exactly those identifiers. Names stop mattering, comments are not in the
AST, and .Body on anything else is not a match.

Positive control, run rather than argued: a handler taking httpReq
*http.Request and reading httpReq.Body is FLAGGED by the new scan, and
matched zero times by the old regex.

Two limits now stated in the test, because an unqualified completeness
claim is exactly how the MCP audit reader got certified safe while
persisting a decoded NUL: accounting is per FILE rather than per call
site, and only signature-declared requests are seen - one stashed in a
struct field or captured by a closure is not a parameter.

Claude-Session: https://claude.ai/code/session_01AUvLoXsKdS5sdpYju6rj4p

* fix(server): mark a cleaned audit identity, and repair two vacuous tests of my own (BUG-2803)

Codex round 23, plus a defect in my own instruments that the mutation
matrix found and the tests hid.

## Cleaning is lossy, so a cleaned identity was forgeable

Round 22 closed the coarse version: sanitising before classifying let
"tools/<NUL>call" become a genuine tools/call. Classifying on the raw
method fixed that. But sanitising still COLLAPSES distinct inputs onto
one output, so "pad_<NUL>item" stored exactly what "pad_item" stores -
same tool_name, same args_hash - and anyone able to send a request could
mint an audit row and a Prometheus label attributed to a real call.

Cleaning and identity are different jobs. sanitiseStoredTextChanged now
reports whether anything was removed, and an identity that only became
well-formed by cleaning is marked. The cleaned text is kept, so the row
stays diagnosable; the marker keeps it distinguishable. Descriptive text
(User-Agent) keeps the unmarked helper - nothing decides anything on it.

The parenthesised form is what this file already uses for a synthesised
value, and a real method or tool name does not begin with "(", so the
marker cannot itself be forged by choosing a clever name.

## Two of my own tests were vacuous, found by a surviving mutant

I wrote nul := "\u0000" in the round-21 and round-23 tests, which in Go
is the NUL CHARACTER, not the six-character escape text. Those bodies
were malformed JSON that encoding/json rejected, so neither test ever
reached the path it named. The comment on the line said "the escape, not
the character"; the code did the opposite, and the correct form was
already three lines away in the round-20 test.

Nothing in the test output showed this. It surfaced only because the
marker mutation SURVIVED, and because a surviving mutant was treated as
a question - does the test not discriminate, or did it not run - rather
than as either answer.

Both repaired and both now kill their mutants: removing the marker fails
with tool_name="pad_item" and a matching 64-character hash; removing
the emptiness guard fails with "(sanitised) " instead of "(unknown)".

Correction for the record: the round-21 checkpoint said that fix was
measured failing before the fix. That measurement used the broken
literal. The finding was real and the fix is right, but it is only
properly established as of this commit.

Claude-Session: https://claude.ai/code/session_01AUvLoXsKdS5sdpYju6rj4p

* test(server): use the canonical escNULLiteral helper, not a local literal (BUG-2803)

The helper is assembled from bytes precisely so this escape cannot decay
into the NUL character it describes, and its comment says so: written as
a Go literal it is one backslash away from being the NUL itself.

I rolled a local one in three tests anyway, and two of them decayed
exactly as that comment predicted - vacuous until the mutation matrix
caught them. The safeguard existed, was documented, and I walked past
it; using it is the only version of this fix that cannot recur.

Claude-Session: https://claude.ai/code/session_01AUvLoXsKdS5sdpYju6rj4p

* fix(metrics): bound the cleaned-identity marker as a metric label, and correct a false cardinality claim (BUG-2803, BUG-2817)

Codex round 24, which enumerated CONSUMERS of the values this unit
changed rather than asking again whether the guard is right. Most of
that enumeration came back FINE, which is the useful half; three
findings did not.

## The marker must not reach Prometheus as part of a name

The cleaned-identity marker is right for the audit ROW - an operator
reading one row needs to know which tool it resembles. It is wrong for
a metric SERIES: "(sanitised) pad_item" and "pad_item" would be two
series per user and per status, for a distinction no aggregate query
asks. metricsToolLabel collapses the marked form to the bare marker, so
it costs exactly ONE extra label value in total and that value is a
constant rather than anything a caller supplies.

Two tests, and the second exists because the first is not enough. The
direct-call test proves the collapse function collapses. The WIRING test
proves the emit path calls it - CONVE-19, my own convention. Measured:
with the call removed from recordMCPCallMetrics, the direct-call test
stays green and the wiring test fails naming the leaked label.

## A cardinality claim that was never true

internal/metrics documented the tool label as "bounded by the catalog
(~7 tools today)" with arithmetic resting on that. The value is
whatever the caller put in params.name, recorded even for requests that
dispatch later rejects, so an authenticated caller can mint a series per
request. The comment now says so and points at BUG-2817, filed with the
fix shape and the two wrinkles it has to decide - the catalog lives in
internal/mcp, and legitimate JSON-RPC methods are not catalog tools.

That unboundedness is PRE-EXISTING and not this unit's to fix; bounding
the marker's own contribution is, which is why the collapse is here and
the rest is filed.

Also corrected: I wrote BUG-2815 into two comments before filing, and
the filing came back BUG-2817. Predicting an identifier is the same
class of claim as predicting a count.

Claude-Session: https://claude.ai/code/session_01AUvLoXsKdS5sdpYju6rj4p

* fix: keep a storable extension in the filename fallback, and sync the rename draft (BUG-2803)

Codex round 24, the two remaining consumer findings. Both trace to this
unit, and both are cases where a value was made SAFE without asking what
reads it.

## The filename fallback was lossier than its sibling

An unstorable upload name became a bare "upload" - no extension - while
the empty-name fallback two lines below has always produced
"upload.bin". The unusable part of "sh<NUL>ot.png" is the STEM; ".png"
is ordinary text, and it is what consumers dispatch on:
Content-Disposition, the web download anchor, bundle export naming, and
, whose documented contract is handing a path to
something that opens files by extension. That command was measurably
affected - it treats any non-empty stored name as authoritative, so its
MIME-based extension fallback never ran and the temp file was
extensionless.

Fixed at the source: a storable extension survives the fallback,
bounded to 16 bytes so a hostile name cannot smuggle a long tail
through. The CLI keeps a defensive extension fallback for any
extensionless stored name, which also covers rows written before this.

Both directions are tested: "sh<NUL>ot.png" now stores "upload.png",
and "shot.p<NUL>ng" - where the EXTENSION is the unusable part - still
stores bare "upload". Without the second leg, "keep the extension"
could quietly become "keep whatever trails the last dot" and reintroduce
the value the fallback exists to remove. Dropping the extension again
fails the first leg.

## A rename that could never come clean

saveName replaced the app object but never updated the draft, so when
the server normalised the name the draft stayed as typed, the equality
check never matched, Save stayed enabled, and each press re-sent the
same request. The server caps at 120 BYTES via rune-safe truncation
while the input allows 120 CHARACTERS, so any multibyte name near the
limit diverges.

The draft is now assigned the value the server actually STORED rather
than compared for length, which stays correct for any future
normalisation. svelte-check: 0 errors.

Claude-Session: https://claude.ai/code/session_01AUvLoXsKdS5sdpYju6rj4p

* fix: reserve the synthesised-value namespace, and stop a fallback carrying an unvetted extension (BUG-2803, BUG-2818, BUG-2819)

Codex round 25, which probed whether the values this unit SYNTHESISES
can themselves be attacked. Earlier rounds asked whether the guard
refuses bad input; this asked what the substitutes are worth.

## A fallback must not carry an extension the product would refuse

Preserving a storable extension was right; bindableText was the wrong
bar for it. Control characters are valid UTF-8 and not NUL, so they are
storable - and they are STRIPPED when the name is written into
Content-Disposition. So ".s<VT>vg" passes the extension blocklist, which
sees no known extension, and reaches the client as ".svg".

attachments.SafeFallbackExtension now requires a KNOWN, ALLOWED
extension, so a synthesised name can only carry a suffix the product
already accepts on the ordinary path. Tested both ways: an obfuscated
.svg and an unknown .foo are both dropped to bare "upload", while
.png still survives.

That divergence is PRE-EXISTING on the ordinary path, where the caller's
name is stored as given and no fallback is involved - filed as BUG-2818
with the fix shape. This change only declines to add a second door.

## A mutation exposed a guard that could not fire

I first wrote an explicit alphanumeric loop in that predicate as well.
Removing it changed nothing: no key in extMIMEMap contains a
non-alphanumeric character, so the map lookup already excluded every
obfuscated suffix. Keeping an unreachable guard whose comment claims it
stops control characters would have misdescribed which line does the
work - so the loop is gone, and TestExtMIMEMapKeysArePlain enforces the
property it was relying on. A guard that survives its own mutation is a
question, not a clearance.

## The marker was forgeable, so the namespace is reserved

Marking only what cleaning changed was not enough. A caller may name a
tool "(unknown)" - what the parser returns for a malformed body - or
"(sanitised) pad_item", and a genuine request then records the same
identity as a substituted one. The older sentinels always had this;
the new marker inherited it.

A leading "(" is now reserved for values this server synthesises, and
any caller value entering that namespace is marked too, so the two
never collide. Cost stated: an MCP tool genuinely named with a leading
"(" is recorded marked; tool names are identifiers in every catalog
this server knows.

The principled fix for the whole class is a provenance FIELD rather than
sentinel strings in a caller-controlled namespace. That is BUG-2819 - it
is a migration on two tables, and the same trick cannot rescue attachment
filenames, which are legitimately named with parentheses.

Claude-Session: https://claude.ai/code/session_01AUvLoXsKdS5sdpYju6rj4p

* test(server): fix the independent oracle, which was wrong in a branch its corpus omitted (BUG-2803)

Codex round 26, finding 5, and it lands on the instrument I introduced
two rounds ago to check the walker.

The oracle descended into a listed key's JSON document whenever it met
one - including when that key appeared INSIDE a natural object that was
itself under a listed key. Production does not: a natural object or
array under a listed key is USER DATA, because the server marshals it
and nothing re-parses it, so a listed key appearing inside it is an
ordinary field name rather than a document marker.

Measured on {"fields":{"schema":"<escape text>"}}: production=false,
oracle=true. Production is RIGHT and the oracle was wrong, so had that
body been in the corpus the test would have failed and pointed at the
production walker.

It was not in the corpus. That is the part worth keeping: the test
already asserted its corpus was not one-sided - that BOTH answers
appear - and that check passed while a whole branch of the contract went
unexercised. Both answers appearing is not the same property as every
branch being covered, and I had treated it as though it were.

Fixed by giving the oracle the same user-data rule, and both bodies are
now in the corpus - the natural-object case that must answer false, and
its string-valued counterpart that must answer true.

Re-verified that the correction did not blunt it: with the production
nested descent removed, the oracle still fails.

Claude-Session: https://claude.ai/code/session_01AUvLoXsKdS5sdpYju6rj4p

* fix: refuse path-component filenames, and single-source the MIME extension table (BUG-2803)

Codex round 26, findings 2 and 3.

## ".." is not a filename, it is a path component

The server guard listed "", "." and "/" but not "..", which survives
bindableText. filepath.Ext("..") is "." - non-empty - so an extension
check waves it through too, and a consumer joining it onto a directory
gets that directory's PARENT. The CLI builds its temp path exactly that
way.

Both ends fixed, deliberately independently. The server now rejects any
name that is only dots or carries a separator, checked on the trimmed
form so "..." and "./" do not each need a case. The CLI sanitises the
name it receives regardless: a client that builds a local path out of a
remote string should not depend on the remote end having sanitised it,
and this CLI talks to whatever instance it is pointed at.

Tested with "..", "...", "./" and "a/b", with an ordinary name as the
premise leg. Restoring the old narrow guard fails it.

## Two tables for one relationship

The CLI kept its own MIME-to-extension table and it had drifted: images
and video but not gzip, tar, XML, YAML, TOML, HTML, JavaScript or
several documents the server has always allowed. So the extension
fallback added in round 24 silently did nothing for exactly the types
whose viewers most depend on it.

The CLI now delegates to attachments.ExtensionForMIME, and the second
table is gone. Measured after: gzip .gz, tar .tar, html .html, js .js,
pdf .pdf.

The reverse map needs one choice per type where several extensions
share one, and those preferences are asserted to name types the forward
map actually uses - because the first version listed "text/yaml", which
this map does not use (it says application/yaml), so that preference
could never fire. Same class as the alphanumeric guard removed in the
previous commit, caught the same way.

Also recorded against myself: I destroyed both new functions mid-edit by
running git checkout on a file with uncommitted work, to "revert an
approach". That is a documented trap I have hit before and had written
down. The committed function survived; the uncommitted ones did not.

Claude-Session: https://claude.ai/code/session_01AUvLoXsKdS5sdpYju6rj4p

* test(server): make the body-reader scan scope-aware, wrong in both directions before (BUG-2803)

Codex round 26, finding 4. The scan used ONE flat name-set per top-level
function, which is wrong in both directions at once:

  - a function literal inside a handler was scanned with the OUTER
    function's request names, so an unrelated inner variable that
    happened to be called r was FALSELY flagged;
  - a request arriving only as a function literal's own parameter was
    INVISIBLE, because literals were never given names of their own.

A false flag in this test is not harmless. The only way to green is to
add the file to the accounted list, and an accounting entry HIDES every
future reader in that file - so a false positive here converts directly
into a blind spot later. That is the same trap that made me abandon the
broadened regex two commits ago.

Now walks a SCOPE at a time. Each scope inherits its parent's request
names, drops any it shadows with a parameter of a different type, and
adds its own. Local aliases (req := r) are picked up as well, since that
is an ordinary thing for a handler to do and the alias reads the same
body.

Three controls, run rather than argued:

  closure parameter reader   -> FLAGGED
  local alias reader         -> FLAGGED
  shadowed inner variable    -> not flagged

The first two were invisible to the previous scanner, which never gave
literals their own names, and the third is the false positive it
produced - both by reading the code this replaces.

Limits restated honestly rather than left as they were, since two of
them are now closed. Still invisible: a request in a struct field, one
from a context, and one whose type reaches http.Request through an alias
or embedded field. This matches the literal spelling rather than
resolving types; closing those means the type checker, not the parser.

Claude-Session: https://claude.ai/code/session_01AUvLoXsKdS5sdpYju6rj4p

* fix: stop the reverse MIME map emitting BLOCKED extensions, and close four instrument gaps (BUG-2803)

Codex round 27 returned "do not merge yet" with three P1s. All of them
are mine, from the previous two commits.

## The reverse map turned a refusal list into a source of extensions

extMIMEMap is the FORWARD table used to REFUSE uploads - it deliberately
lists .svg, .exe, .com so those extensions can be recognised and
rejected. Reversing it wholesale meant ExtensionForMIME("image/svg+xml")
answered ".svg", where the old CLI table answered nothing, and
 names a local file with that.

So I closed an SVG door two commits ago and reopened one through the
MIME helper. Blocked types now get no reverse mapping at all, and the
test asserts it with a premise leg (the map must CONTAIN a blocked type,
or the assertion never runs). Removing the exclusion fails naming .svg,
.com and .msi.

## The oracle was closer, not identical

Production descends only into a JSON DOCUMENT - a string whose trimmed
form starts with { or [. The oracle unmarshalled any valid JSON, so a
SCALAR under a listed key made it answer true where production answers
false. Closer to production is not a usable oracle; only identical is.
Aligned, and the scalar case is in the corpus.

## The scan was still not scope-aware, and could now MISS a reader

A nested block shared the enclosing name-set, so
{ r := &http.Response{}; r.Body.Read(nil) } was FALSELY flagged. And the
shadowing rule deleted a name rebound to http.Request BY VALUE - which
still shares the Body, since it is an interface holding the same reader
- so that read became invisible. Blocks are now their own scope and a
value request counts.

## The controls I claimed were not in the suite

Round 27 was right: I had run them as throwaway probes and deleted them,
so nothing held the scanner to them. The scanner is now a package-level
helper and TestBodyReaderScanDiscriminates drives it over ten synthetic
files - six that must be detected (plain, unconventional name, closure
parameter, alias, value copy, form reader) and four that must not
(no request, shadowed by a closure parameter, rebound in a nested block,
mentioned only in a comment).

## And an over-refusal of my own making

The filename guard rejected any dot-only name and anything containing a
separator. Only "." and ".." are path components; "..." is an ordinary
POSIX filename, and filepath.Base has already reduced "a/b" to "b", so
the separator test was dead on this platform and removed rather than
left looking load-bearing. Preservation controls now pin that
legitimate names survive.

Also removed: an unused id parameter on safeLocalFilename.

Claude-Session: https://claude.ai/code/session_01AUvLoXsKdS5sdpYju6rj4p

* test(server): only a DEFINE can rebind a name, and cover the idiomatic reassignment (BUG-2803)

Found by probing my own previous commit rather than by a review round -
the first time in this sequence I have caught the adjacent breakage
before the next round did.

The scope rules deleted a request name on ANY assignment whose right
side was not a request identifier. That is wrong in the dangerous
direction, and it fires on the most idiomatic line in Go HTTP code:

    r = r.WithContext(ctx)
    io.ReadAll(r.Body)      // <- invisible to the scan

WithContext is a call, so the name was dropped and every later read went
unseen. Measured before the fix: MISSED.

The correct rule is type-sound. Go is statically typed, so a plain
cannot change a variable's type: if it held a request before, it holds
one after. Only a DEFINE introduces a new binding that can be something
else. So the delete is now gated on token.DEFINE, which is both more
correct and simpler than what it replaces.

Three controls added, and the two that would have caught this are the
ones I had not written: a WithContext reassignment, and readers inside
an if body and a for body - the last two because making every nested
block its own scope is exactly the kind of change that could have
started missing them. Thirteen controls now, six negative.

Reverting to delete-on-any-assignment compiles and fails the
WithContext leg.

Claude-Session: https://claude.ai/code/session_01AUvLoXsKdS5sdpYju6rj4p

* fix: make the body-reader scan conservative by design, and reduce filenames cross-platform (BUG-2803)

Codex round 28. Two findings, and the first is the fifth consecutive
round to find a FALSE NEGATIVE in the same instrument.

## Stop modelling scopes; change the error direction instead

Rounds 24 through 28 each found another way the scope-modelling scan
missed a real reader: a value-copied request, a plain
r = r.WithContext(ctx), a mixed r, ok := ... that reuses an existing
variable, and if/for/switch initialisers and case clauses whose scopes
it did not model. Each fix closed one case and left another. That is a
design telling me something, not a run of bad luck.

The two error directions are not symmetric here. A false NEGATIVE hides
a body reader, which is the entire thing this test exists to prevent. A
false POSITIVE costs one human review and an accounting entry with a
reason attached. So the scanner now OVER-APPROXIMATES on purpose: any
name bound to an http.Request anywhere in the file counts for the whole
file, aliases are followed to a fixed point, and names are never
un-bound. Every scope-shaped false negative becomes structurally
impossible.

The cost is real and is now asserted rather than discovered: two
controls that previously expected "not flagged" - a name shadowed by a
closure parameter, and one rebound in a nested block - now assert
CONSERVATIVELY FLAGGED, so the bias is on the record. Three of round
28's named misses are added as controls and pass: mixed short
declaration, switch case, if-initialiser shadow. Sixteen controls, and
the accounting test still passes against the real package - so the
over-approximation costs nothing today.

Exactness needs go/types with a real package load, which is a bigger
instrument than this test warrants. The comment says so, and names the
signal that would justify building it: an accounted entry whose reason
is "the scan over-flagged".

## A filename safe on this OS is not safe on the consumer's

filepath.Base is platform-specific, so on Unix it leaves a backslash
alone - and the stored name is consumed cross-platform. A Windows client
joining a stored "..\evil.png" onto a directory traverses upward.

Reduced to the leaf under BOTH separator conventions. This normalises
rather than refuses, which is less lossy than replacing the whole name
and keeps round 27's point that a backslash is legitimate on Unix.
Removing the reduction compiles and fails the new test.

Claude-Session: https://claude.ai/code/session_01AUvLoXsKdS5sdpYju6rj4p

* test(server): count readers per accounted file, add two missing reader methods, correct a false reason (BUG-2803)

Codex round 29, and its central point was aimed at my REASONING, not my
code. It was right.

## The over-approximation argument was wrong for per-file accounting

I justified a deliberately conservative scanner by saying a false
positive costs one review and one accounting entry. That is not what it
costs. Once a file is listed, a NEW reader added to it is covered by the
existing entry and the test stays green - so a false positive does not
cost a review, it permanently blinds the list for that file. My own
comment already recorded that hazard two commits earlier, and I argued
past it anyway.

The fix is to make the entry carry a COUNT of reader expressions rather
than a yes/no. Adding a reader to an accounted file now changes the
number and fails, so the entry must be re-read and its reason
re-justified. It churns exactly when a body reader is added or removed,
which is when a human should look.

Demonstrated: inserting r.PostFormValue into handlers_tokens.go - an
already-accounted file - is FLAGGED. Before this it was absorbed
silently.

The conservative bias stays, because the false-negative classes it
eliminates are real and the count now removes the reason it was
expensive.

## Two real reader methods were missing

MultipartReader STREAMS the body and FormFile triggers multipart parsing
of it. Neither was in the selector list, and FormFile is used in
production in handlers_attachments.go - so the list was incomplete
against code that exists, not hypothetically.

## A reason in the list was simply false

handlers_tokens.go was accounted as "a nil/ContentLength check only - it
never reads the body". It guards on those and then calls decodeJSON. A
wrong reason is the same defect as a missing entry: both let a reader
pass as reviewed.

## And I guessed the counts

I wrote plausible numbers for the per-file counts and every one was
wrong; the test reported the real ones on its first run. Same habit this
branch keeps catching - a figure written from expectation reads exactly
like a figure that was counted. They are measured now and the comment
says so.

Also recorded: my first verification script failed to apply its mutation
and still printed a verdict, which I nearly banked. It now aborts unless
the mutation is present in the file.

Claude-Session: https://claude.ai/code/session_01AUvLoXsKdS5sdpYju6rj4p

* fix(attachments,cli): close the closing round's two view defects, correct two overclaiming comments (BUG-2803)

The closing enumeration (successor seat, per the lead's convergence ruling)
returned two real attachment-view defects and two comments claiming more than
their code delivers. Fixed here; the round's two design-scale findings are
filed instead (BUG-2820 scanner precision via go/types, BUG-2822
Windows-unstorable filename forms).

- Reverse MIME map: four ALLOWED spellings (text/xml, text/yaml,
  application/javascript, audio/webm) had no reverse extension because no
  extMIMEMap entry uses them as its value — `pad attachment view` wrote an
  extensionless temp file for exactly the types the delegation was built to
  fix. Population measured against the whole allowlist: these four, no more.
  An alias table closes them; TestEveryAllowedMIMEHasAnExtension asserts the
  class property over the allowlist (a future allowlist entry with no reverse
  extension fails), plus alias hygiene (allowed keys only, no forward-derived
  collisions, alias extensions must map to ALLOWED types so the table can
  never mint a refused extension). Mutation-verified: removing the alias
  application fails the test on all four types.

- safeLocalFilename: a trailing dot survived every check and
  filepath.Ext("photo.") is "." — non-empty — so the MIME-extension fallback
  never fired and the temp file dispatched on no extension. Trailing dots are
  now stripped (cannot empty the name; dots-only names already returned
  early). The CLI guard also gains its first direct tests, including the
  backslash and traversal refusals that previously rode untested.
  Mutation-verified: removing the TrimRight fails both trailing-dot cases.

- Two comment corrections, same defect class the accounting list itself
  names (a wrong reason reads as review): the handlers_cloud.go entry said
  bodyHasCloudSecret "restores" the body — it restores the first 64 KiB and
  drops the tail, a bound that file documents; and the accounting test's
  header said its scan "cannot be spelled around" while its own KNOWN LIMITS
  block lists the spellings that get around it (struct field, context value,
  type alias). The header now matches the limits block.

* fix(server,attachments,cli): close closing-round-2's scanner blind spot and four stale comments (BUG-2803)

Closing round 2 (successor seat) found no product defects; all four findings
were in instruments and comments. Each verified against the code, then fixed:

- The alias fixed-point resolved only identifier RHS (`req := r`), so a
  dereferenced copy (`c := *r; io.ReadAll(c.Body)`) was an invisible body
  reader — and unlike the disclosed type-level classes, this one was not in
  the KNOWN LIMITS block. The copy shares the Body (an interface holding the
  same reader). StarExpr operands now join the alias set; a new control pins
  the case. Mutation-verified: reverting the StarExpr handling fails the
  control. The type-level classes (struct field, context value, type alias)
  remain disclosed and are BUG-2820's territory.

- The KNOWN LIMITS block said shadowed request names are "correctly
  ignored" while the controls deliberately assert they are conservatively
  OVER-FLAGGED — stale prose from the scope-aware era, falsified by the
  round-28 conservative flip that never touched those lines.

- The reverse-map stability loop compared only PREFERRED entries across
  rebuilds; it now compares the entire map (sizes and every mapping) against
  the first build. Boundary stated honestly: every multi-spelling type today
  is preference-pinned, so the full-map comparison discriminates only when a
  future non-preferred multi-spelling entry appears — that future entry is
  what it guards.

- Three orphaned/wrong comments: a `mimeForExt` doc block glued above
  ExtensionForMIME (the function it described is gone); the old hardcoded
  extension-table doc glued above safeLocalFilename (falsified by the
  delegation it predates); and two "120 chars" claims where the cap is 120
  BYTES rune-safe via truncateBindableText — the consent form's
  maxlength=120 counts characters, so a multibyte name passes the client
  and is still truncated server-side, which is now what the comments say.

* fix(cli,server,docs): close the attachment-view path escape, and closing-round-3's instrument and prose findings (BUG-2803)

Closing round 3 found the branch's first product defect since round 17, in
BRANCH-ADJACENT code the round-24 fallback extension work made reachable: the
`pad attachment view` id fallback joined the RAW id onto its temp dir, and the
client sent the id into the URL path UNESCAPED. An id is a CLI argument, but
the documented agent flow harvests it from item content ("pad-attachment:"
refs other workspace members write), so a traversal-shaped "id" could
re-route the HEAD/GET to a different endpoint whose 200 then vouched for it,
and the write escaped the temp dir. Both halves fixed and both
mutation-verified through a new command-level test: reverting the fallback
sanitize demonstrably wrote OUTSIDE the sandboxed TMPDIR; reverting the
PathEscape put a raw "../../" on the recorded wire.

- internal/cli: url.PathEscape(attachmentID) at both id-bearing client sites
  (HeadAttachment, DownloadAttachment — the enumerated population).
- cmd/pad: the id fallback runs through safeLocalFilename, generic
  "attachment" when nothing survives; view's long help no longer claims the
  filename is used "without rewriting the extension" — it describes the
  reduction and the MIME-extension append, and says why the CLI is stricter
  than the server (the name is written to YOUR filesystem).
- cmd/pad: attachmentViewCmd gets its first command-level test (CONVE-19 —
  the helper tests vouched for the component, not its wiring): disposition
  name, extensionless+MIME append, id fallback, traversal containment with a
  wire-escaping control, generic fallback.

Instrument and prose findings, each verified before fixing:

- The KNOWN LIMITS disclosure now names the ordinary alias forms the
  fixed-point does not walk (var-spec, call-derived, named results, range
  bindings) — they were in BUG-2820's filing but not in the in-file
  disclosure, which is what let the round read them as unfiled. The scanner
  itself deliberately does NOT grow another parser patch; go/types is the
  filed fix.
- TestTextSafeHelpersAreUsedAtEveryCallSite pins EXACT occurrence counts
  (measured: 1 declaration + 4 call sites each) instead of a >=4 floor a
  removed call site could hide under.
- middleware_mcp_audit: two stacked comment copies rested non-forgeability
  on "real names do not begin with (" — the exact reasoning round 25
  retired; the const doc now points at auditLabel's namespace-reservation
  rule, which is what actually makes the marker non-forgeable.
- docs/backup.md said repair is needed before "the export or migration" goes
  through, contradicting its own "exports fine" three paragraphs up — it is
  the IMPORT or migration that fails; the export succeeds either way.

* fix(server,cli): decode chunked watch bodies, refuse dot-segment attachment ids, correct two texts (BUG-2803)

Closing round 4 found one PRE-EXISTING product defect and one residue of the
round-3 fix, plus two wrong texts. Each verified before fixing:

- Watch creation gated its body decode on `ContentLength > 0`, so a CHUNKED
  request (ContentLength == -1) had its body silently DROPPED — the caller's
  predicate ignored, an unconditional watch created, 200 returned. The
  population of ContentLength gates in the package is exactly two:
  handlers_tokens.go already used the `!= 0` form, watches now matches it,
  with io.EOF tolerated so the documented no-body-is-valid contract holds
  for an empty chunked body too. Three handler-level tests discriminate the
  cases; the mutation (condition back to `> 0`) fails the two it should and
  passes the empty-body control. The accounting instrument then flagged the
  new `r.Body != nil` reference in the file — its exact job — and the file
  is now accounted with a measured reader count of 1.

- url.PathEscape leaves exact "." and ".." UNCHANGED, so those two ids still
  reached the wire as live dot segments for a proxy or server to normalize —
  the escaping added in round 3 did not cover them. Both id-bearing client
  sites now share attachmentIDPathSegment, which refuses exactly those two
  values before any request (a real id is a UUID; the refusal cannot fire on
  one). Mutation-verified: removing the refusal fails the new subtest, which
  also asserts zero requests reach a recording stub.

- The artifact rejection text said "NUL byte"; the same refusal fires for a
  NUL manufactured by a YAML escape during parsing, where no raw NUL byte
  exists — now "NUL character", in the handler message and the error var.

- A test comment claimed the User-Agent reaches sessions.user_agent as
  text; sessions store only ua_hash, as the accounting list's own exemption
  states two hundred lines up. The sentence now agrees with it.

* fix(attachments,server): remove a can't-fire MIME preference and a stale filename-guard sentence (BUG-2803)

Closing round 5 is down to two P3 comment defects; both verified and fixed:

- preferredExtensions "preferred" .md over a .markdown that has never been
  in the forward map — a line that cannot fire, the exact class this
  branch's own instruments hunt (the alphanumeric guard, the text/yaml
  preference, the charset loop). Entry removed; shortest-wins picks .md as
  the only candidate, unchanged. The preference-hygiene test now asserts
  every entry has a real competitor (>= 2 forward-map spellings), and the
  counterfactual — re-adding the entry — fails it.

- The upload filename guard still carried round 26's "checking the trimmed
  form rather than listing spellings" sentence directly above round 27's
  code that does the opposite (exact "." / ".." comparisons, longer dot
  runs deliberately preserved). The stale layer is gone; the surviving
  paragraph already records why.

* fix(server,docs): drop a dead test fixture, stop claiming the failing row is named (BUG-2803)

Closing round 6 returned one P3 — TestDecodeJSONTrimsOnlyJSONWhitespace
booted a full testServer it never used (`_ = srv`), dressing a direct
decodeJSON test in router coverage it does not have. Removed.

Its enumeration also re-read docs/backup.md against the code: "the failing
row is named in the error" is true of neither leg — the import answers 400
naming the RULE it refused on (the NUL check is body-wide and knows no row),
and `pad db migrate-to-pg` reports which WORKSPACE's copy failed. The doc
now says exactly that, and that locating the value is manual until
BUG-2810's preflight lands.

* fix(server,cli): retire a stale byte-search claim, close two instrument gaps from closing round 7 (BUG-2803)

Round 7 found no production defects; three instrument/comment findings:

- artifactIsBindableText's doc comment still asserted the round-8 byte-search
  approach and that "the ambiguity cannot arise here" — directly above the
  round-9 body comment recording that assertion as simply wrong and doing the
  round-trip walk instead. The doc paragraph now describes the round trip
  and points at the body's history.

- TestBodyReaderScanDiscriminates listed MultipartReader and FormFile in the
  scanner's selector set but had no control for either, so their removal
  from that list was undetectable. Two controls added.

- The attachment-view test proved nothing about the MIME delegation: every
  case used image/png, which the OLD hand-rolled table also knew, so a stale
  local table passed. Two cases added — application/gzip (a type round 26
  found missing from that table) must gain .gz, and blocked image/svg+xml
  must gain nothing. Mutation-verified: a stale-table mutant that answers
  only for png fails the gzip case. (First mutant attempt didn't build —
  unused import — and was not counted as a detection.)

The per-file same-count substitution gap round 7 restated is declined as
filed, not fixed: BUG-2820's filing already specifies per-call-site
accounting via go/types as the fix that retires the per-file count
workaround; the KNOWN LIMITS closing line now carries that ref.

* fix(attachments,server): sweep two pre-BUG-2413 disposition comments, pin the manifest refusal to 400 (BUG-2803)

Closing round 8 found no production defects; two evidence findings, verified
then fixed:

- Two comments still described the PRE-BUG-2413 disposition policy: the
  RenderChip mode doc said the HTTP layer serves every chip inline, and the
  read-path doc derived Content-Disposition from RenderMode. The live policy
  is the explicit fail-closed ServeInline allowlist — most chip types are
  served as "attachment". Both now say so and record the history.

- TestImportBundle_RefusesNULInManifest accepted any status >= 400, so the
  documented 400 could decay into a 500 unnoticed. Pinned to
  http.StatusBadRequest.
2026-08-30 21:30:19 -04:00
xarmian 91d92f184f feat(server): refuse workspace creation without may_create_workspaces consent (IDEA-2756) (#1212)
* feat(server): refuse workspace creation without may_create_workspaces consent (IDEA-2756)

The OAuth consent screen's "Let this app create new workspaces" checkbox
gated only the post-creation auto-add. A connection whose user left it
unticked could still create workspaces; it simply could not then see
them. A permission that does not prevent the action it names is a
consent mismatch.

Dave ruled it: the checkbox is a permission on whether the connected
token may CREATE, and it has to be true to what a user would honestly
expect from the option. The behaviour-change-for-existing-connections
argument loses to honest consent semantics.

Adds Server.requireWorkspaceCreationConsent, a shared gate at the top of
both endpoints that mint a workspace under the caller's account:

  POST /api/v1/workspaces         handleCreateWorkspace
  POST /api/v1/workspaces/import  handleImportWorkspace

Import reaches CreateWorkspace via store.ImportWorkspace, so it is the
same permission at a second door — lead-ruled as an application of the
same rationale, not a new decision. The gate sits above the Content-Type
dispatch, so it covers the tar.gz bundle path (whose only route is that
handler) and refuses before the 64 MiB body read.

Refusal is a 403, mirroring handleAuditLog's consent refusal (BUG-2102):
a hard decline rather than a narrowed response, because there is no
narrower version of creating a workspace.

Three non-refusal cases and one refusal, all but the last with a test:

  - not an OAuth grant (PAT, CLI session, local stdio) — creation rides
    on ordinary account authority
  - ErrOAuthConnectionNotFound (pre-Phase-C grant) — ALLOW, matching the
    backfill's may_create_workspaces=ON default. Deliberately asymmetric
    with maybeAutoAddCreatorConnection's not-found branch, which declines
    a convenience where this one would invent a refusal
  - flag set — proceeds; the auto-add is unchanged
  - a store I/O error — REFUSED, failing closed with a 500, because
    allowing the create when the deciding state could not be read grants
    a declined permission on the strength of a database blip. This is
    the one branch with no test: injecting a store read failure needs a
    fault-injecting store the package does not have, so it is reasoned
    rather than measured

Population enumerated before the fix (CONVE-18): five CreateWorkspace
call sites, two of them HTTP endpoints reachable by an OAuth token (both
gated). Excluded with reasons: autoCreateWorkspace (signup-time, no
connection in context), workspace restore (un-deletes an existing
workspace), /oauth/claim (grants access, does not mint), cmd_db.go
(local store copy, no HTTP). Search boundary: the sweep traced
Store.CreateWorkspace callers and did not look for a path that inserts a
workspace by raw SQL.

Ten tests, all driving the real router rather than calling handlers
directly (CONVE-19). Every refusal leg asserts that no workspace of that
name exists afterwards, not merely the status code (CONVE-12) — a guard
that 403s after the write passes a status-only assertion. Seven mutants,
seven detected, including both guard-placement mutations.

MCP tool surface 0.25 -> 0.26. Behaviour bump on the v0.9/v0.16/v0.25
grounds: no tool name, action enum or param shape changed, but
pad_workspace.create now refuses a call it used to permit. Closest
precedent is v0.10; unlike v0.10 there is deliberately no escape-hatch
param, because the gate encodes a decision the USER made at consent time
and a bypass flag would be the app overriding its own grant.

CONVE-23 sweep for prose the change falsified: instructions.md told
agents the create still succeeds and to use the claim flow (it would
have sent them to claim something that was never created); the
TASK-2753 allow-list guard entry asserted the same and posed IDEA-2756
as open; the MCP catalog and CLI help described only the flag=true path;
maybeAutoAddCreatorConnection's flag-off branch is now unreachable from
its sole caller and is documented as dead code kept for contract, to be
deleted only with the guard. CLAUDE.md was already stale at v0.24 (v0.25
bumped the constant without it) — brought to v0.26 with a backfilled
v0.25 line.

The consent screen and console copy are unchanged: they were the
misleading half of this bug, and the fix makes them true.

* docs(server): state the import gate's reachability precisely (IDEA-2756)

The import-side gate is correct but currently unexercised in production,
and the first framing of this change did not say so.

WithMCPTokenIdentity is stashed by exactly one middleware, MCPBearerAuth,
mounted on /mcp alone. An OAuth connection reaches an /api/v1 handler
only through the in-process MCP dispatcher, and that dispatcher's route
table has a workspace create action but no workspace import. So no
OAuth-bound caller can reach handleImportWorkspace today.

The gate stays, and the comment now says why: adding that action later
must not silently reopen the door, which is the state a create-only fix
would have left armed.

Found on a verify pass reading the middleware mount points, not by the
tests — they synthesize the OAuth identity into the request context, so
they prove the handler's behaviour GIVEN an identity and have no opinion
about which routes supply one (CONVE-19). Codex round 2 reached the same
conclusion independently.

* fix(server): correct five overstated claims from Codex round 3 (IDEA-2756)

All five were mine, all P2, none changing the gate's behaviour — four are
claims that were broader than the code, one is a test that proved less
than its name.

1. "Only re-authorization lifts it" was wrong in five places (version.go,
   README, CLAUDE.md, the MCP catalog description, CLI help). A user can
   also enable the flag on the EXISTING connection via
   PATCH /connected-apps/{id}/flags, which the console page drives —
   instructions.md said so and contradicted the others. All five now name
   both remedies, and both are still the user's, which is the part that
   matters: neither is reachable by the app.

2. "This branch is UNREACHABLE ... it is dead code" on
   maybeAutoAddCreatorConnection's flag-off branch was false. The gate
   reads the connection and that function reads it AGAIN after creation;
   a user revoking creation power from the console between those two
   reads lands exactly there. It is a real second check across a real
   TOCTOU window, failing in the safe direction. The claim was written
   from the call graph, which cannot see a concurrent write between two
   reads.

3. handlers_import_bundle.go's "Auth: any authenticated user" was made
   false by this change and the concept sweep never had a chance at it —
   it greps may_create / auto-add / creation power, and that sentence
   contains none of them. Corrected in place.

4. The two NonOAuthCallerUnaffected tests claimed PAT, CLI session and
   local stdio; each drives one PAT. The comments now state the fixture's
   real scope and why one caller stands for the class (the guard branches
   on an identity only MCPBearerAuth sets, so callers that skipped it are
   indistinguishable) rather than implying three fixtures.

5. The JSON import refusal leg would have passed with the gate below
   decodeJSONWithLimit — only the bundle leg pinned placement, and only
   for gzip. Adds TestImportWorkspace_ConsentRefusalPrecedesBodyDecode
   (malformed body: 400 if the gate is late, 403 if it is early),
   mirroring the create-side ordering legs.

Mutation matrix now 9 mutants, 9 detected. M8 (guard below the JSON
decode) is killed by the bundle leg too, so it shows the new test is
covered rather than necessary; M9 gates the bundle path and moves only
the JSON path's guard, and dies to the new leg ALONE. That is the mutant
that justifies the test.

* docs(server): the second consent check narrows the race, it does not close it (BUG-2792)

Round 3 caught me calling maybeAutoAddCreatorConnection's flag-off
branch dead code. The replacement comment then claimed the branch means
a revoked grant cannot silently gain a workspace — which is more safety
than the code delivers, and round 4 caught that.

The read and the AddConnectionWorkspace insert below it are separate
unconditional statements, so a revocation landing BETWEEN them still
adds the workspace. The check narrows the window; it does not close it.

Filed as BUG-2792 rather than folded in: the race is pre-existing and
unchanged by IDEA-2756, and closing it needs an atomic check-and-insert
at the store layer, written and gated for both dialects — materially
more diff and risk than this handler-level guard.

Both mistakes were the same shape in opposite directions: a claim about
concurrency derived from reading the call graph, which cannot see a
concurrent write between two reads.

* style(server): gofmt the doc comment (IDEA-2756)

gofmt wants blank lines between list items once one item spans multiple
paragraphs, which the BUG-2792 note made true.

My error, and worth naming exactly: I ran build, vet and the targeted
tests on this commit but not lint, because lint had passed on the
PREVIOUS commit and the change was 'only a comment'. The gate has to run
on the tree being pushed, not on an earlier one that resembles it. CI's
golangci-lint is pinned to the same v2.11.4 the Makefile installs, so
there was no version skew to blame — the local gate would have caught
this in 51 seconds.

* docs(server): correct ten overstated prose claims from Codex round 8 (IDEA-2756)

Round 8 reviewed only the prose this change adds. Ten claims were
broader than the code. All ten are mine; none changes behaviour. Rounds
3, 4 and 7 each caught one of these, which is why round 8 was pointed at
the class rather than at a new dimension.

The substantive ones:

- "gates every endpoint that MINTS a workspace" — autoCreateWorkspace
  mints from registration, bootstrap and oauth-login and is deliberately
  outside this gate. The helper doc and the test header now name the two
  callers and the exclusion instead of claiming universality.

- "the agent was handed a workspace it could not then see" (version.go,
  README, CLAUDE.md) — only true for a connection with an EXPLICIT
  allow-list. An all_current_workspaces=true connection is not gated per
  slug and could see what it made. The consent mismatch is the constant;
  the invisibility was its most visible symptom, not its definition.

- "ErrOAuthConnectionNotFound — a pre-Phase-C grant" asserted a cause the
  code cannot know: ANY missing row takes that branch. Now stated as the
  expected cause, with the limit of what the code can tell.

- "above the 64 MiB body read" conflated the two import paths. 64 MiB is
  the JSON decode's bound; the bundle path has its own, much larger. The
  gate precedes both, which is the property that actually matters.

- "the request context is decorated AFTER TokenAuth runs" was false, and
  inherited verbatim from the sibling helper this was modelled on
  (handlers_oauth_claim_test.go's doClaim), where it is also false. The
  wrapper sets the identity BEFORE ServeHTTP; it survives because
  nothing on the /api/v1 chain writes that key.

- "lets CreateWorkspace normalize it" — CreateWorkspace slugifies only
  when the supplied slug is EMPTY, and import supplies a non-empty one,
  so an imported workspace keeps the ?name= value verbatim.

- "The PAT needs a workspace to bind to" — CreateAPIToken takes
  WorkspaceID as optional.

And one where the first fix was worse than the finding:

- "Every refusal leg asserts no workspace exists afterwards" was false —
  the two ordering legs assert status only. My first correction ADDED
  those assertions, which is the trap the finding was pointing at: a
  malformed body and an empty name are rejected before creation under
  every guard placement, so "no such workspace exists" is true of broken
  and working code alike. Reverted; the header now states which legs
  carry the counterfactual, and why the ordering legs discriminate on
  status instead.

Gates re-run on the tree being pushed, not an earlier one: gofmt clean,
lint 0 issues, internal/server and internal/mcp green, mutation matrix
still 9/9.

* ci: re-trigger CI after a GitHub startup_failure (IDEA-2756)

No code change. The Go job on cb47c763 failed on BUG-2786 (the recurring
internal/events subscribe-confirm guard, which fails by asserting its own
premise: 'the acknowledgement never landed before the mark; this test could
not have discriminated'). CONVE-11 owes that failure a re-run before it can
be called a flake.

rerun-failed-jobs produced attempt 2 = startup_failure with the Go job stuck
in 'queued' — a GitHub infrastructure fault, not a test result — after which
the run refuses further retries ('This workflow run cannot be retried'). The
CI workflow has no workflow_dispatch trigger, so a push is the only way to
get a fresh run.

Evidence the failure is unrelated to this branch, gathered before re-running
rather than after: the branch touches 0 files under internal/events (11 files
total, none in that package), and the parent tip c6818500 had Go: SUCCESS with
the only non-comment Go difference being one added assertion in this branch's
own test file. Go (PostgreSQL) also passed on cb47c763, exercising the same
package.
2026-08-26 13:23:56 -04:00
xarmian e747a1610c feat(session): registry keyed on the harness session, carrying the agent name; pad session list / prune (TASK-2767) (#1200)
## Summary

TASK-2767 (IDEA-2750 part 2, with part 3 riding along — the keying fix and the reaping are one mechanism).

The local session registry (`~/.pad/sessions`) was keyed on the pid of the `pad session register` subprocess, which is dead before anyone reads the file. One session left a new file per call and its own pid appeared in none of them; the only live identifier was the harness pid a reader could parse out of the socket path's basename. In practice nothing wrote it (zero callers in `plugin/`, `skills/`, or hooks) and nothing read it.

Now:

- **One record per session, keyed on the harness session pid** — `$PAD_SESSION_PID` (harness-agnostic override), else `$CLAUDE_PID` (verified present in both the tool shell and a live plugin monitor's `/proc/<pid>/environ`), else the calling process. A set-but-invalid value is an error, not a silent fall-through.
- **The record carries the agent name** the session's writes are attributed to (`ResolveAgentName`: `.pad.toml agent_name` → `$PAD_AGENT` → detected runtime; `--agent` overrides, `--agent ""` is anonymous), the harness session id, and the messaging socket's identity (inode/device/mtime — the same binding the arm-state file uses).
- **One owner-identity type, one verdict.** `internal/cli/session_owner.go`: `SessionOwner` + tri-state `OwnerLiveness` (`alive` / `dead` / `unknown`). `armStateOwnerAlive` is now `OwnerLiveness(...) == alive` with its file contract preserved (socket identity else mtime; headless pid + start token; fail closed). The registry pruner takes the opposite posture on `unknown`: on Windows `pidAlive` reports dead for every pid, and a reaper built on that would delete every live session's record.
- **Verbs:** `pad session register [--agent]` (writes/refreshes; prunes dead records), `pad session list [--agent] [--cwd] [--all] [--format json]` (liveness per row, newest first; dead hidden unless `--all`), `pad session prune [--older-than DUR]` (dead always; unknown only under an explicit bound; alive never). Nothing on MCP — host-local filesystem state.
- **Who registers:** `plugin/scripts/pad-monitor.sh` runs `pad session register` on start, BEFORE the consent gate — presence is a fact, consent is a grant, and the record is local/0600/never on the wire.
- **Legacy v1 files** list as `legacy` rows: owner = socket-basename pid (else registrar pid), liveness by pid only (v1 recorded no socket identity, and the socket-without-identity rule would have judged every legacy record dead while its session ran). A legacy row can say a session exists, never who it is.

Lead rulings on the four open decisions, all as built: `agent`/`--agent` vocabulary; no server-presence merge in `list`; register from the monitor script before the gate; wire follow-on (agent name on the stream) filed separately as IDEA-2750 part 2b.

One ordering change from the plan's section A: pid precedence is `PAD_SESSION_PID` > `CLAUDE_PID` > self (explicit override beats detection, mirroring `PAD_AGENT` over runtime detection); the plan listed `CLAUDE_PID` first.

## Behaviour changes for existing users of `~/.pad/sessions` / `pad session register`

- Registry files are keyed on the **harness session pid** (`PAD_SESSION_PID` → `CLAUDE_PID` → self), not the `pad` command's pid; repeated registrations overwrite one record instead of accumulating.
- `pad session register` records the agent name, harness session id and socket identity; stores the **real path** of the cwd; prints a different text line and a different JSON shape (the full `SessionRecord`); and **rejects** an invalid `PAD_SESSION_PID` / `CLAUDE_PID` instead of silently keying on itself.
- Existing v1 files are read as `legacy` rows (owner = socket-basename pid, no agent name) and dead ones are pruned by the next register.
- The plugin monitor now registers (and prunes) on every start, before the consent gate.
- `armStateOwnerAlive` now delegates to the shared `OwnerLiveness`; the consent gate's observable behaviour is unchanged on every platform and key type (codex round 4 traced every caller; matrix M29 pins the socket-keyed mapping).

https://claude.ai/code/session_016zc6oxBvpax6Z3iQMsAJno
2026-08-25 15:31:16 -04:00
xarmian c73584088f fix(watchevents): detect a half-open Redis connection with a bus heartbeat (BUG-2769) (#1199)
* fix(watchevents): detect a half-open Redis connection with a bus heartbeat (BUG-2769)

internal/watchevents had the same defect as internal/events did, by the same
mechanism: ChannelWithSubscriptions on a connection whose go-redis health check
only writes. PubSub.Ping calls writeCmd and returns without reading a reply
(v9.22.0), so a route that stops carrying traffic without closing is invisible —
the instance blocks on a read forever while its replay buffer goes on looking
complete.

Named as a class sweep in BUG-2738's filing and deferred there. It became
load-bearing when that unit shipped: docs/deployment.md told operators the gap
was "closed on the activity stream and still open on the watch stream". This
diff falsifies that, which is why the prose sweep is part of it.

THE PORT IS SMALLER THAN THE ORIGINAL BY DESIGN. This bus holds ONE
process-wide subscription created in its constructor, off any request path, so
none of BUG-2747's establishment machinery exists to interact with: no
per-workspace map, no establishment record, no single-establisher wall, no
concurrency cap, no bounded-parallel recovery, and no per-workspace cycle
scoping. Cost is flat too — one frame per instance per interval regardless of
workspace count.

NO COMPANION COUNTER, and that was CHECKED rather than inherited.
internal/events needs pad_event_subscription_cycled_total because its
dropWorkspaceCoverage returns early when a workspace has no buffer, so the reset
reason under-reports the early-wedge case. dropCoverage here has no such branch:
it replaces the buffer and reports unconditionally, so idle_timeout is a
complete count on its own and a second metric would be a number needing to be
explained against its neighbour for no signal.

THE RECEIVE LOOP NOW OWNS ITS SUBSCRIPTION AND CONTEXT. A cycle replaces the
subscription under a running bus, and the loop reading the old one must tell "I
was replaced" from "the client died" — the second logs an ERROR and moves a
counter documented to mean the instance has gone deaf. The cycle cancels that
loop's own context before closing its PubSub, so it leaves by the quiet door.
Its own test.

I PORTED A FLAW ALONG WITH THE STRUCTURE, and the wiring test caught it: both
maintenance halves shared one kick channel, so whichever goroutine was waiting
consumed it and the other stayed on the stale cadence. internal/events' mutation
matrix found exactly that (M11c) and fixed it; the fix did not come across. That
is the contamination hazard this port's grounding warned about, in its literal
form, caught by the CONVE-19 test rather than by review.

Two more found by mutation, both missing tests rather than missing code: nothing
asserted that ordinary traffic keeps the instance alive (removing the per-frame
stamp survived, because every other test drives idleness through the clock), and
nothing asked for a SECOND detection (a replacement inheriting stale stamps
gives a detector that works exactly once, which is worse than one that never
runs because it looks like it works). The second needed a direct assertion on
the install stamps, because the behavioural route re-stamps the field it was
meant to be testing.

Trio in one commit as required: reason enumeration, the
pad_watchevents_sequence_resets_total Help string, and docs/deployment.md — plus
the two BUG-2738 sentences this falsifies and a new section explaining how the
watch bus differs from the activity one.

Claude-Session: https://claude.ai/code/session_01JVDBKbgn3Xt7ndW1YoYd8X

* fix(watchevents): fence stragglers and re-validate before the drop (codex r1)

Three findings, and two of them are BUG-2738 fixes I again failed to bring
across with the structure. That is now three times in one port: the shared kick
channel, the stale idle decision, and the missing generation. The mechanism is
the same each time — I ported what the code DOES and not what its review
history taught it, and each was caught by a test or a reviewer rather than by
me reading the source I was copying.

STALE IDLE DECISION. cycleIfIdle decided under one lock and tore down under
another; a heartbeat or notification arriving between them left a demonstrably
alive subscription being dropped and every client on the instance resynced for
nothing. BUG-2738 fixed exactly this at its round 11. Re-validated immediately
before the drop, with a positional seam so a test can land the recovery inside
the window rather than racing it.

NO GENERATION FENCE. Cancelling a receive loop and closing its PubSub does not
JOIN the goroutine, and go-redis's channel is buffered, so a frame from a
replaced subscription could still stamp the replacement's liveness, append to
its buffer, or drop its coverage. On a wedged route that is the worst
direction: the dead connection's buffered tail suppressing the detector for its
successor. One check at the top of the frame handler covers all three, because
the three must agree about whether a frame belongs to the live subscription.
The probe stamp is fenced separately, since a slow publish can outlive the
subscription it was sent for.

A COPIED COST PARAGRAPH THAT CONTRADICTED ITS OWN SECTION. The activity bus's
"each workspace has its own subscription, N frames per interval" text sat below
the new watch-specific section saying the opposite. Retitled and moved above it.

FOUR INSTRUMENT DEFECTS ON THE WAY, all found by mutation:

- Nothing asserted ordinary traffic keeps the instance alive — every other test
  drives idleness through the clock, so removing the per-frame stamp survived.
- Nothing asked for a SECOND detection, so a replacement inheriting stale stamps
  gave a detector that works exactly once — worse than one that never runs,
  because it looks like it works. Needed a direct assertion on the install
  stamps, since the behavioural route re-stamps the field under test.
- The generation tests asserted the PREDICATE, not that the loop calls it.
- And that wiring test could not discriminate on a frozen clock, where a stamp
  writes the value already there. It advances the clock first now.

Claude-Session: https://claude.ai/code/session_01JVDBKbgn3Xt7ndW1YoYd8X

* fix(watchevents): make the generation fence atomic with what it guards (r2)

Two P1s, both mine, both the same shape: a check in one lock acquisition and the
mutation it guards in another.

THE FENCE WAS NOT ATOMIC WITH ITS MUTATIONS. One check at the top of the frame
handler read well and guarded nothing reliably — a replacement between that
check and stampLastSeen / fanOutFromRedis / dropCoverage let a straggler through
to any of them. The generation now travels TO each mutation and is re-checked
under the same lock that mutates. A stale notification entering the
replacement's buffer is the worst of the three: it makes the instance vouch for
a span it never received, which is the false coverage claim this whole family
exists to remove.

THE OLD GENERATION STAYED CURRENT ACROSS THE REPLACEMENT. subGen was
incremented only after the new subscription was confirmed, leaving the cancel,
the close, the dial and a round trip during which the OLD generation still
passed every fence. Retired at teardown now, so during resubscribe NO generation
is current and a late frame is ignored everywhere. That also makes the failure
path honest: the "no notifications until restarted" log was false — no
generation is current, so the next idle tick tries again.

Revalidation and the drop are now ONE critical section rather than two, for the
same reason at one level down: a frame arriving between them was silently
discarded by a drop already decided on.

Also: phase 1 no longer starts the maintenance goroutines, and the watch bus's
phase is logged at startup — an operator cannot read an absence of idle_timeout
without knowing whether the detector was running, and the two flags are
independent.

DOCS still described the workspace model in the section that claims to cover
both buses: one heartbeat "per subscribed workspace", a phase table naming only
PAD_EVENTS_HEARTBEAT, and coverage described as a workspace's. Generalised.

Two more instrument gaps, both found by mutation: nothing asserted a straggler
cannot enter the replacement's BUFFER (only the stamp was covered), and the
phase-1 goroutine gate is untested by design — removing it changes no behaviour,
only goroutine count, and the only assertion is a census that would be flaky
here. Said out loud rather than left to look like coverage.

Claude-Session: https://claude.ai/code/session_01JVDBKbgn3Xt7ndW1YoYd8X

* test(watchevents): prove each generation fence on its own

Round 3's fix put a generation check in each of the four places a frame
from a replaced subscription can mutate shared state, rather than one
check at the top of the receive path — a check in one lock acquisition
and a write in another is a TOCTOU, which is what codex blocked.

Four checks means four mutations, and the matrix found the first pass of
tests could not tell them apart: removing the append's check, or the
coverage drop's, left every test green. Not because the guards were
redundant — because no test drove those paths with a stale generation.
The straggler tests all enter through fanOutFromRedis, whose own guard
returns early and hides the one below it, and nothing at all drove
dropCoverageForGen with a straggler.

So the fences are asserted one at a time, each through the entry point
that actually reaches it:

  epoch bookkeeping   fanOutFromRedis with a foreign epoch — the loudest
                      of the four, since an accepted straggler would
                      rewrite the id space and resync every client on the
                      instance
  buffer append       fanOutLocally directly, under the guard above it
  coverage drop       dropCoverageForGen, previously undriven
  liveness stamp      stampLastSeen, which would otherwise let a dead
                      socket's traffic hold detection open

Each fails against removal of the single check it names (M16/M17/M19 and
the existing stamp mutation), and the four together still pass the
end-to-end straggler tests unchanged.

Refs BUG-2769

* test(metrics): prove the two new watch signals reach the registry

Both were wired and neither was asserted at the metrics layer, which is
where docs/deployment.md's claims about them actually live. A reason or
a callback that never reaches the registry is a runbook pointing at a
series that does not exist, and nothing in internal/watchevents can
catch that — its observer is an interface, satisfied by a test double.

  pad_watchevents_heartbeat_publish_failures_total  incremented six
  times, a count no other assertion in that test uses, so a callback
  wired to the wrong counter cannot land on the right number by
  coincidence. Fails when the increment is pointed at a neighbour.

  sequence_resets_total{reason="idle_timeout"}  asserted with the
  literal label, alongside the four spellings already pinned there and
  for the same reason BUG-2739's rename left that test behind. Fails
  when the constant drifts.

Also corrects the shared "what happens if you run them out of order"
paragraph, which moved under a heading covering both buses while still
describing only one: it said the frame travels on "the workspace's event
channel" and that an un-upgraded instance resyncs "for every workspace",
neither of which is the watch bus, where there is one channel and one
buffer per instance. The blast radius differs in scale between the two
and the paragraph now says so.

Refs BUG-2769

* docs(watchevents): correct three counted claims that stopped being true

All three said "three" where the code now has four, and each was
accurate when written — the fourth fence (the epoch bookkeeping in
fanOutFromRedis) was identified after them, in the pass that found the
matrix could not tell the guards apart.

That is the whole failure mode: a count is a claim, and a claim written
before the last change is wrong afterwards with nothing to notice it.
Two of the three sat inside a comment ABOUT how carefully the guards
were enumerated, and one names them now instead of counting them, so
the next site added has to appear in the list or contradict it visibly.

Found by sweeping the branch diff for counted prose rather than by
rereading, which is what had already missed them twice.

Refs BUG-2769

* test(config): close the other half of the two-flag independence claim

The flag tests asserted PAD_WATCH_HEARTBEAT does not move
EventsHeartbeat and stopped there, while the comment above them and the
deployment doc both claim the two buses roll INDEPENDENTLY. That is a
biconditional and one leg does not establish it: a Load() that pointed
PAD_EVENTS_HEARTBEAT at both fields passed everything. Now both
directions are asserted, and the events leg checks its own premise
first, so a fixture that stopped setting the flag fails as a fixture
rather than as a pass.

Also pins env-over-file precedence for the watch flag, in the direction
that actually matters: PAD_WATCH_HEARTBEAT=false over
watch_heartbeat=true in config.toml. That is the rollback for a bad
phase-2 flip, and an operator reaching for it mid-incident cannot be
editing a file on every host.

Mutation matrix, each detected: the env var wired to the neighbouring
field, the env var never read at all, and the toml tag dropped.

Refs BUG-2769

* test(watchevents): fix five tests that passed for the wrong reason

Codex round 4 went at test honesty rather than correctness and found no
BLOCK, but it found five assertions that hold whether or not the thing
they name works. Each is now driven through the path it claims, and each
was mutation-checked against the specific defect it exists to catch.

  the malformed-frame contract  only ever called isWatchHeartbeat. The
  predicate can be perfect while the receive loop routes every "hb|…"
  payload to the ignore arm without asking it, which is the defect, and
  the test's name promises coverage ends — a claim about the loop. Now
  published on the real channel, with a well-formed frame as the control
  so the assertion cannot be satisfied by a loop that finds everything
  undecodable.

  the receive-loop wiring test  published, slept 300ms, and asserted
  nothing had changed. A loop that stalled or never started satisfies
  that perfectly. There is no natural signal to wait on instead, because
  a frame the fence refuses is by design invisible — hence a seam that
  fires after the loop handles a frame whichever arm it took. Bounded,
  so a stalled loop fails with a message rather than a package timeout,
  and followed by a control that the same loop still accepts a frame
  whose generation matches.

  the quiet-exit test  asserted only that no loud exit was reported,
  which a replaced goroutine that never exits at all also satisfies —
  a leak, and the worse outcome. Now joins the loop first via a
  process-wide live-loop count, then checks the counter, so it is a
  statement about a goroutine that has finished.

  the maintenance-loop wiring test  claimed both halves and observed a
  heartbeat, which a loop that started only the publisher passes. The
  idle half cannot be proved there at all: against a live miniredis this
  bus's own heartbeats come back and refresh liveness every cadence, so
  wedging it with the loop running is a race against the publisher —
  which is what my first fix for this turned out to be, flaky at 2 in 3.
  Renamed to what it proves, pointing at the blackhole end-to-end test,
  which drives the scanner for real and detects both mutations.

  the straggler test  never delivered a straggler. It incremented subGen
  by hand, called isCurrentGen, and compared an unchanged timestamp
  without touching a mutation path — green with every fence removed.
  Deleted rather than repaired: the four-way per-fence test added
  earlier covers it properly, and isCurrentGen went with it.

Plus two ordering changes in Close/resubscribe that ARE NOT fixes for an
observed race, and say so in the test. Making b.pubsub reassignable made
Close's unlocked read of it look wrong, and resubscribe's wg.Add outside
the lock look like it could land after Close reached Wait. Both windows
turn out to be shut already by resubscribe's b.closed check, which sits
under the same acquisition as the count — reverting either fix leaves
the new Close-during-cycle test green. Kept as defence because the
invariant they lean on is three functions away, and documented so
nobody later reads them as evidence of a bug that existed.

Also corrects the metric help and two comments that said an idle cycle
"replaced the connection" when it attempts a replacement that can fail;
the deployment doc already said attempted. And the deployment doc's
rollback, frame-validation, what-to-watch and startup-log paragraphs,
all of which moved under a heading covering both buses while still
describing only the activity one.

Refs BUG-2769

* refactor(watchevents): drop an always-empty return and the branch reading it

dropCoverageIfStillIdle returned (string, bool) where the string was
never anything but empty — the reset it reports goes out through the
pending/flush path inside the lock, so the caller's `if report != ""`
was unreachable. A second reporting path that exists in the signature
and never fires is a thing a later change wires up by accident.

Refs BUG-2769

* fix(watchevents): a failed re-dial retries without re-dropping coverage

Codex round 5, on behaviour across a full Redis outage. No BLOCK; this
was its one P2 and it is real.

The probe-failure suspension does not cover this case, and the reason is
worth stating because the suspension looks like it should. Suspension
asks "did our last probe get through", and that can be YES with the
route already gone: the last successful publish stamps lastProbeOK,
Redis dies before that frame comes back, and lastSeen stays behind it.
From there both timestamps are frozen — the probe fails so nothing
stamps lastProbeOK, nothing arrives so nothing stamps lastSeen — and the
cycle's precondition stays true for the whole outage. Every pass then
dropped coverage, announced to every subscriber, and re-dialled.

Only the re-dial is owed. The second drop empties an already-empty
buffer and re-announces a hole every subscriber has been told about,
and it moves pad_watchevents_sequence_resets_total{reason="idle_timeout"}
once per cadence — so a five-minute outage read as ten incidents on the
series operators are told to alert on.

cycleIfIdle now has a retry-only arm ahead of the decision, entered when
there is no subscription at all, and the teardown clears b.pubsub /
b.subCancel so that state is representable. Clearing them also stops
Close closing an already-closed PubSub a second time.

Two tests, discriminating in OPPOSITE directions, because the obvious
fix for the noise is to suspend the pass and that would trade a noisy
outage for one the instance never returns from — retrying the dial IS
the recovery path:

  three passes with Redis away        one reset, not three
  Redis returns after a failed pass   the subscription is re-established
                                      and the counter does not move again

Matrix: removing the retry arm, making it return without retrying, and
leaving the torn-down subscription in place are each detected, the
middle one only by the recovery test.

internal/events has no equivalent defect. Its teardown deletes the
workspace's subscription entry, so its next scan finds nothing live and
abandons; recovery there runs off the request path.

Refs BUG-2769

* fix(watchevents): only one caller may install a replacement subscription

Codex round 6, verifying round 5's fix. No BLOCK; this was its P2.

Both the cycle and its new retry arm dial with the lock RELEASED, which
is deliberate — a Redis round trip under the bus's hot mutex would stall
every fan-out on the instance — so two passes can each find no
subscription and each dial one. Installing both is wrong twice over: two
receive loops would run on the SAME generation, so both accept every
frame and each notification is processed twice, and the loser's PubSub
would be untracked, closed by nothing including Close.

The install is what needs serialising, not the dial, so the loser
discards its own connection under the lock rather than the two racing to
overwrite b.pubsub.

Only the idle scanner calls this today, so this guards an invariant
rather than fixing an observed fault. Written down because the invariant
lives in a different file from the code relying on it, and because the
failure is silent duplication rather than a crash.

The test races two resubscribes through the install seam. Two details it
needed, both found by running it rather than reading it:

  the loop count is incremented INSIDE the goroutine, so sampling it
  right after the constructor returns reads zero — the first version
  did, and measured every later count against that wrong baseline. It
  waits for the loop now.

  the seam release is deferred, because without it the guard's mutation
  parks both callers in the callback, Close waits on receive loops that
  cannot start, and the detection arrives as a package-wide hang with no
  message. That is how the mutation first appeared to pass.

Also completes the idle_timeout reason in three comment/help sites that
still enumerated four reasons and said "the last two" — the same stale
count corrected in the observer contract earlier on this branch, missed
in its neighbours because I fixed the one the reviewer named instead of
grepping for the claim.

Refs BUG-2769

* test(watchevents): count installs instead of waiting for one that never comes

Codex round 7 returned no BLOCK and no P2 on the production code, and
two NITs on what round 6 added. Both are real.

The concurrency test synchronised on a WaitGroup expecting BOTH callers
to reach the install seam. Only the winner does — that is the property
under test — so in the passing case the goroutine waiting on it blocks
forever. A leak inside a test written to prove a leak does not happen is
not a shape to leave standing. An atomic the abandoning caller never
touches carries the same information and blocks nobody, and it removes
the release channel and its deferred close along with it.

The final assertion also moved off liveReceiveLoops and onto that
count. A loop starts AFTER its install, so reading the loop count can
catch a second caller's goroutine before it has begun and see the
passing value on a failing run. Both callers have returned by the time
the install count is read, so it is final. Detection over ten runs with
the guard removed: 10/10, where the loop-count version was a race
against a goroutine's first instruction.

Also softens the retry arm's log line. It said the instance receives no
notifications until an attempt succeeds, which is true for today's
single scanner and stale the moment there are two: one caller's dial can
fail while another has already installed. It now claims only what the
failing call knows.

Refs BUG-2769

* test(watchevents): hold both callers at the window, and say what that misses

Codex round 8's P2, on the test the previous commit rewrote. Starting
two goroutines from a start gate makes overlap likely and guarantees
nothing: one can finish resubscribe before the other begins, so the
window the install guard closes need never have been open.

A seam at the dial/install boundary — connection dialled, lock not yet
taken — lets both callers announce their arrival and wait for each
other. Now the window is open by construction rather than by luck, and
the test fails as a fixture if only one caller ever reaches it, instead
of passing on evidence it never gathered.

AND IT STILL DOES NOT DETECT EVERYTHING, which the test now says in
place of leaving it implied. Measured:

  guard removed entirely                        10 runs, 10 detected
  guard checked in its own acquisition, then    10 runs,  0 detected
  the lock retaken to install

The second is the regression round 8 asked about, and catching it would
mean landing the second caller inside a check-to-install gap that exists
only in the mutant — there is nothing to yield on there, and no seam can
be placed in code that is not written. So this test covers "a guard
exists", not "the guard is in the right critical section". The latter is
held by the comment at the guard and by review, and a test comment
claiming otherwise would be worth less than the honest note.

Refs BUG-2769

* fix(watchevents): make the frame seam and the cycle log tell the truth

Codex round 9 was asked whether this should merge and said hold for a
cleanup pass. Five findings, no correctness blocker, and every one of
them a claim that had stopped matching the code.

  the frame seam did not fire for every arm, though its comment said so.
  The arms that decline to act — a heartbeat, an undecodable payload, an
  unsubscribe confirmation — were `continue` statements, which skipped
  everything after the switch. A test waiting on the seam for one of
  those frames would have HUNG rather than failed, which is the worst
  way to find this out. The switch is now its own method so every arm
  ends the frame by returning, and a test drives one frame per
  publisher-reachable arm and counts three. Detected against restoring
  the skip.

  the idle-cycle warning was emitted before the revalidation that can
  abandon the cycle, so it could announce coverage ending and resumes
  answering sync_required for a subscription that was then left alone —
  a log line with no counter behind it, and an on-call hunting a bug
  that is not there. internal/events learned this at its own round 6;
  the reason did not come across with the port. Moved after the decision
  is final, still saying "attempting" to replace because the resubscribe
  can fail.

  the quiet-exit test sampled liveReceiveLoops instead of waiting for
  it, so its "the replaced loop left" assertion could be satisfied by a
  loop that never ran. Same defect fixed in the sibling concurrency test
  a commit earlier and missed here, because I looked at the test the
  reviewer named rather than at the pattern. Latent rather than
  observed: sampling survives 10 runs, so this removes a possibility.

  the probe-failure log and metric help called an errored Publish a
  failure to publish. A returned error can also mean the reply was lost
  after Redis accepted the frame, so the honest claim is that the probe
  is UNCONFIRMED. It changes no behaviour — an unconfirmed probe is not
  evidence about the receive path either, so detection suspends the same
  way — but an operator reading the counter should not be told more than
  the instance knows.

  the deployment doc said the watch stream differs in "three things" and
  listed four, the fourth being the bullet I added last round. Third
  instance of that species on this branch; the count is gone rather than
  corrected.

Refs BUG-2769

* docs(watchevents): stop one unconfirmed probe standing in for a broken path

Codex round 10 confirmed four of round 9's five fixes and held the fifth
as partial. It was right on all three residual sites.

Renaming the condition to "could not confirm" did not fix the sentences
downstream of it. The log still said silence cannot be read as a finding
"when we could not ask" — but we may well have asked, and lost only the
answer. And both the metric help and the observer contract said an
instance in this state "is also failing to deliver its own notifications
to every other instance", which is a conclusion about the outbound path
drawn from a single call that did not come back.

The inference is sound at a SUSTAINED rate and worthless at one
increment, so both now say which is which. That distinction is the whole
value of the counter to an on-call: a blip is a lost reply, a rate is a
broken path, and the same wording for both makes the first look like the
second.

No behaviour change. An unconfirmed probe suspends detection exactly as
a definite failure does, because it is not evidence about the receive
path either way.

Refs BUG-2769

* docs: sweep the BUG-2738 prose this change makes false

BUG-2738 shipped documentation that describes the watch stream as still
carrying the half-open defect. Merging this makes those sentences wrong,
and I flagged the sweep as owed twice during the groundwork and then did
not do it — the lead caught that the package said nothing about it.

Five sites, each re-read after editing rather than grepped for, because
grepping for a phrasing I chose is how I have twice verified a sweep
that had not landed:

  the residual enumeration opened "One gap remains everywhere, and a
  second remains on the watch stream only", then described one gap and
  said it was open on both. The second WAS the half-open case. Now
  states one gap, on both streams, and says where the second went.

  the half-open paragraph already said "closed on both streams" — the
  one site I had fixed — but omitted that each half is behind its own
  phase-2 flag, so a reader takes it as closed on their deployment when
  it is closed only once they turn it on.

  "A third residual" counted the item it followed. With the second gone
  the ordinal was wrong; it does not need one.

  "these two gaps" in the closing sentence, same arithmetic.

  the pad_event_subscription_cycled_total row told an operator to read
  heartbeat_phase off the startup log. There are now two such fields on
  two lines under two flags, and only one bears on that counter. It
  names the line.

No code change; suite 28/28 and lint 0 re-run because the branch is
under review and a docs commit that skips them is a commit nobody
checked.

Refs BUG-2769
2026-08-25 11:04:18 -04:00
xarmian cc3cfeef2b fix(redis): honour the caller's context on TLS dials (BUG-2754) (#1198)
* fix(redis): honour the caller's context on TLS dials (BUG-2754)

go-redis's default dialer (v9.22.0, options.go NewDialer) honours the caller's
context on plaintext and NOT on TLS: the TLS branch returns
tls.DialWithDialer, which takes no context at all, so a cancelled caller could
not shorten the dial and it was bounded only by DialTimeout.

BUG-2749 put SSE subscription establishment on the request's context so a
client that disconnects stops holding its admission slots. On plaintext that
covered the dial. On TLS the dial was the one segment cancellation could not
reach, so the guarantee shrank from "released at once" to "released after up to
DialTimeout" — and a managed Redis is a rediss:// URL, which is the ordinary
production shape rather than an exotic one.

Fixed at CLIENT CONSTRUCTION rather than in any consumer, because the same dial
serves Publish, the Lua scripts, the presence registry and the watch bus's
reads. internal/redisdial is a small package so the thing can be tested
directly; cmd/pad/cmd_server.go installs it on the one client Pad builds.

THREE THINGS THAT FAIL QUIETLY IF THE REPLACEMENT GETS THEM WRONG, each with a
test that fails against getting it wrong:

ServerName. tls.DialWithDialer infers it from the dialled address when the
config leaves it empty; a hand-rolled tls.Client does not, and an empty
ServerName leaves certificate verification with no name to check. That would
turn a latency fix into a silent authentication regression. Replicated, on a
CLONE — mutating the caller's config would leak one host's name into every
later dial that shares it. Tested by dialling a certificate issued for another
name and requiring an x509.HostnameError, per the lead's correction: asserting
the field is set proves the code sets a field, not that the name is checked.
Verified in the pinned source rather than assumed — redis.ParseURL DOES set
ServerName for rediss:// (options.go:708), so Pad's path does not depend on the
fallback today; it is there because it is what the replaced code did.

The timeout must bound the HANDSHAKE, not just the connect. Otherwise a server
that accepts and then stalls hangs for as long as the context lives — trading a
bounded failure for an unbounded one, worse than the bug being fixed.

It must not EXTEND an earlier deadline. context.WithTimeout takes the sooner of
the two, matching go-redis's own promise about DialTimeout.

PROSE SWEPT, and the sweep found two sites my first pass missed because it only
grepped non-test files: five comments across internal/events said the TLS dial
could not be cancelled, including one carrying an explicit "See BUG-2754 for
the TLS half" forward reference. All five now say what is true, and the
confirmTimeout budget comment records that it has been amended twice.

Two instrument corrections: the certificate fixture put IP literals in DNSNames
where x509 will never match them, and two tests detected their mutations BY
HANGING — which is not a result anyone can act on, and which stranded the
mutation harness with its edit still applied. Both bound the dial in a
goroutine now, so a hang is a named failure.

Claude-Session: https://claude.ai/code/session_01JVDBKbgn3Xt7ndW1YoYd8X

* fix(redis): resolve DialTimeout, keep the keep-alive, share one budget (r1)

Codex round 1 found three, and the P1 was introduced BY the first draft of this
fix rather than inherited — the worst kind, since the diff was sold as closing
a hang.

DIALTIMEOUT READ AS ZERO. go-redis's Options.init() defaults it to 5s, but
NewClient CLONES the options first (redis.go:1924), so a caller reading
opt.DialTimeout in order to install a Dialer — the only time it can — reads
zero for the ordinary URL that sets none. And PubSubPool.NewConn calls the
dialer DIRECTLY with no timeout of its own (internal/pool/pubsub.go:45), so
nothing downstream supplies one either. Resolved in the package, with the
coupling named.

The mutation matrix then refused to confirm the failure mode the finding
described, which changed the test rather than the fix. An unresolved zero does
not hang HERE: this dialer wraps the dial in context.WithTimeout, and a zero
duration is an already-expired deadline, so every dial would fail INSTANTLY —
nothing connects at all. The original draft would have hung; this one refuses.
The assertion that separates them is a healthy server being reached, not a
stalled one giving up, and the comment says which draft did which.

KEEPALIVECONFIG DROPPED. go-redis's default dialer sets it (options.go:608) and
it governs how quickly a dead peer is noticed on every Redis connection this
process holds. Reverting to OS defaults would change that across the whole
client as an invisible side effect of a cancellation fix — invisible because
nothing fails.

My first test for it compared our copy against go-redis's published numbers,
which says nothing about whether the dialer USES it: deleting the field from
the dialer left that test green. Replaced with a Linux-tagged test that reads
SO_KEEPALIVE and TCP_KEEPIDLE off the accepted socket. Honest partial, stated
at the test: the property is platform-independent, the observation is not, and
the Smoke jobs on macOS and Windows skip the file. The value-comparison test is
kept as well — it catches the copy drifting from what it mirrors, which the
socket test cannot.

TWO SEPARATE BUDGETS. DialTimeout was applied to the TCP connect and then a
fresh one started for the handshake, allowing up to 2x on the pub/sub path,
which has no outer deadline to mask it. tls.DialWithDialer bounds both as one
interval; this must not be laxer than the code it replaces.

Claude-Session: https://claude.ai/code/session_01JVDBKbgn3Xt7ndW1YoYd8X

* fix(redis): honour an explicitly disabled dial timeout; finish the sweep (r2)

Codex round 2, four findings and one correction to a claim I had already made.

EXPLICIT dial_timeout=0 WAS BEING OVERRULED. ParseURL encodes an explicit zero
or negative as -1, which go-redis preserves as "no timeout at all". Treating
every non-positive value as unset collapsed that into the 5s default and
silently overruled an operator who had deliberately disabled the bound — the
same defect as the one round 1 found, in the opposite direction. `== 0` for the
unresolved case now, with a negative carried through as no bound, and a test
that goes through ParseURL rather than passing -1 by hand so it pins the real
path.

A DRIFT GUARD for the two copied constants, compared against go-redis's
RESOLVED options (NewClient runs init() on its clone and Options() returns the
result) rather than against a literal. A copy that silently diverges from what
it mirrors is what would make this package worse than none.

TWO STALE COMMENTS I HAD CLAIMED WERE FIXED. My sweep commit said "all five now
say what is true"; it was three. The first patch batch aborted on a failed
anchor and, because that helper writes only after every pair matches, none of
its edits landed — I re-applied some by hand and did not re-verify the rest. The
grep I ran afterwards searched for phrasings the surviving comments did not use.
Both now corrected: establishSubscription's two-bullet plaintext/TLS split and
the mutex comment that named TLS as the case cancellation could not reach.

TWO TESTS RELABELLED RATHER THAN LEFT LOOKING LIKE COVERAGE. The single-budget
test does not discriminate — the server accepts immediately, so the connect
consumes none of the budget and the two-budget implementation finishes in the
same time. Staging a slow connect against a local listener is not deterministic,
so what holds that property is structural (one context, created before the
connect, passed through the handshake) and the test says so. And the Linux-only
keepalive test now states what its build tag does and does not cost: the
behaviour is platform-independent and the full suite runs on Linux CI, so a
removal is caught; the macOS and Windows Smoke jobs are build-and-start checks
and were never the guard.

Claude-Session: https://claude.ai/code/session_01JVDBKbgn3Xt7ndW1YoYd8X
2026-08-25 00:27:27 -04:00
xarmian effd0199cd fix(events): detect a half-open Redis connection with a bus heartbeat (BUG-2738) (#1195)
* fix(events): detect a half-open Redis connection with a bus heartbeat (BUG-2738)

A Redis connection can stop carrying traffic without closing -- no FIN, no
RST, just a route that stopped working. The instance blocks on a read that
never returns, receives nothing, and its replay buffer goes on looking
complete, so every resume is answered "caught up" from a coverage window that
ended when the route did.

go-redis cannot see it: PubSub.Ping writes the command and never reads a
reply (v9.22.0), so its health check reports healthy for as long as the socket
accepts writes. Measured on day-52 against a proxy that silently stopped
forwarding: no reconnect in 24 seconds.

Each subscription now records when it last received ANYTHING -- event,
heartbeat, or subscription acknowledgement -- and a background pass ends the
coverage of any workspace whose stamp goes stale past 3T, then REPLACES the
connection. Drop alone would not recover: the resync it demands is served from
the same dead socket, so the detector fires again on the next pass.

Dave's day-49 ruling dissolves the threshold rather than tuning it. The bus
publishes its own frame every T=30s and fires at 3T=90s, which turns "is this
workspace quiet or is the route dead?" -- unanswerable, deployment-dependent --
into "did our heartbeat arrive?".

TWO PHASES, ORDER NOT OPTIONAL. The frame must travel on the workspace's event
channel, because that connection is what needs proving. A pre-phase-1 binary
cannot classify it: the frame reaches the event decoder, fails, and since
BUG-2739 that is a hole in coverage -- so an early flip makes every un-upgraded
instance drop its buffer and resync all its clients, every 30s, per workspace,
for the length of a mixed deployment. Phase 1 recognises and ignores;
PAD_EVENTS_HEARTBEAT is phase 2, a constructor parameter with no default so
every call site states its phase.

The idle detector is a THIRD actor in a region whose invariants were designed
around request goroutines plus Close. Four rules, each commented at
cycleIdleSubscriptions and each with a test:

  1. It refuses to cycle while pendingSubs holds a record, and MINTS the
     record itself before tearing anything down -- subscribeAndReplay checks
     pendingSubs before wsSubs, so a subscriber arriving mid-cycle joins the
     replacement instead of being admitted into the doomed subscription.
  2. lastSeen is stamped at INSTALL, not left at the zero value, which reads
     as 1970 and would cycle hardest on an unconfirmed admission -- the
     workspaces already having a bad time.
  3. wsCounts is re-read under the lock that performs the teardown.
  4. Re-establishment runs on b.ctx with a nil establisher; the bus has no
     subscriber registration of its own to unwind.

Two decisions beyond the plan:

A NEW COUNTER, not just the reset reason. dropWorkspaceCoverage reports a
reset only when a buffer existed to drop, and the incidents this detector
exists for skew hard toward having none -- a route that wedged early on a
quiet workspace. Reading cycles off the reset label alone would under-report
exactly the case it was built to find, so pad_event_subscription_cycled_total
is the dependable count and idle_timeout is corroboration. Both comments say
which is which.

THE CADENCE IS A LIVE TUNABLE -- a timer re-read under b.mu each pass plus a
buffered kick, not a ticker constructed once. A ticker captures the interval
at goroutine start, which makes the field write-once while its comment calls
it a tunable and makes any later write a data race; it also leaves no
deterministic way to test the WIRING other than a test-only constructor.

decodePayload's signature grew a payloadKind. The classification belongs to
the decoder, not the call site, so no future caller can reintroduce the
coverage drop; and the prefix (rather than an exact payload) means a later
frame version needs no third roll.

Also swept, per the team's prose convention: receiveMessages' doc comment and
deployment.md both said this gap was open and needed a decision. Both now say
what closes it -- and deployment.md says the watch stream still has the same
defect by the same mechanism, which is its own unit.

Trio kept together: ResetReasonIdleTimeout, the metric Help strings, and
docs/deployment.md's rollout order with the mixed-fleet failure named.

Claude-Session: https://claude.ai/code/session_01JVDBKbgn3Xt7ndW1YoYd8X

* test(events): rebuild the instruments the BUG-2738 matrix showed were blind

The mutation matrix found a defect in the fix itself and three tests that
could not have caught what they were named for.

THE DEFECT: the idle scan skipped a subscription whose lastSeen was the zero
value. That reads as belt-and-braces beside the install-time stamp and is the
opposite -- it makes a subscription that has NEVER received anything
permanently uncyclable, which is the BUG-2747 unconfirmed admission: the one
population the plan singles out as mattering most, and the one where a wedged
route would then be undetectable forever. It was also masking rule 2: with the
skip present, removing the install stamp survived every test. Skip removed;
that mutation is now caught. Re-adding it is undetectable by construction and
the comment says so, because a guard that only acts once a real one has broken
converts a caught defect into a silent one.

THREE INSTRUMENTS THAT WERE NOT MEASURING:

- "Drop only, never cycle" passed because establishSubscription overwrites
  wsSubs, so a generation check cannot see a replacement installed WITHOUT
  tearing the old connection down -- a leaked PubSub, connection and receive
  goroutine per cycle, forever, on exactly the wedged route where they never
  die on their own. Now asserted on the receive loop exiting.

- The Close test was vacuous. Close drains wsSubs, so a loop that ignored
  b.ctx entirely would find no workspaces and publish nothing: silence after
  Close was evidence of nothing. maintenanceStopped makes the goroutine's exit
  observable, which is the same reason Observer.ReceiveLoopExited exists.

- The joint test HUNG rather than failing under the drop-only mutation: the
  seam never fires, so the joiner goroutine was never spawned and an unbounded
  receive waited forever. The harness then aborted mid-run and LEFT THE
  MUTATION APPLIED to the working tree, which a grep caught and a green test
  run would not have. The wait is bounded and names the failure; the harness
  bounds each run, reports a hang as its own status, and restores in a finally.

Added: a direct test that a straggler frame from a replaced generation cannot
refresh its successor's liveness -- on a wedged route, the dead connection's
buffered tail would otherwise suppress the detector for the replacement.

RULE 3 IS AN OPTIMISATION, NOT A CORRECTNESS GUARD, and the matrix says so
rather than an argument: removing the whole second read -- liveness, generation
and count terms together -- survives every test, because
establishSubscription's abandon path already refuses to install for an emptied
workspace and retires the record in the same critical section (BUG-2749). The
first read is redundant more sharply still: reaching zero takes the
subscription down with it, so this loop never sees such a workspace. Both are
kept, because neither DEPENDS on that coupling, and both comments now carry the
per-term reading instead of describing tested defence in depth. The generation
term is unreachable while the establishment record is held, by rule 1's own
mechanism.

Matrix: 16/22 detected, plus 4 follow-ups. Every survivor is documented at its
line with why it survives.

Claude-Session: https://claude.ai/code/session_01JVDBKbgn3Xt7ndW1YoYd8X

* fix(events): gate idle detection on heartbeat phase 2 (BUG-2738, codex r1)

Codex round 1 found a defect the first draft had shipped WITH A COMMENT
JUSTIFYING IT, plus two coupling hazards.

P2-as-filed, P1 in effect: idle detection ran on every instance from phase 1,
on the reasoning that it could "detect off whatever traffic the deployment
already carries". That holds only for a BUSY workspace. A QUIET one on phase 1
has no events and no heartbeat, so a perfectly healthy subscription crossed
the 90s threshold on every pass and was cycled: replay coverage dropped, every
live subscriber told to resync, indefinitely -- on the DEFAULT configuration
every deployment lands in before it flips anything. A resync storm shipped as
the default, by the feature whose stated purpose is to avoid exactly that load
inversion.

Publishing and detecting are now one switch, which is what they always were:
an instance detects off its OWN frames -- it publishes to the channels it
subscribes to and receives them back -- so it never depended on peers having
flipped, and there was never a reason for the two to be separable. Phase 1 is
"recognise the frame so a phase-2 peer costs you nothing", and nothing else.
Regression test plus its counterfactual, so "no cycles" cannot be satisfied by
a detector that has simply stopped working.

P1: the maintenance loop published heartbeats and scanned for idleness on one
goroutine. publishHeartbeats makes N synchronous Redis publishes, and against
the failure this feature exists to detect those are precisely the calls that
block -- bounded by go-redis's own Dial/Read/WriteTimeout, not by any context
we can pass. A stalled publisher could therefore delay detection for as long
as those timeouts take, on the very instance whose connections had wedged, and
for longer the more workspaces it carried. Two goroutines with their own kick
channels; a stalled publisher now just produces silence, which is what the
detector reads.

P3: the cycle held the workspace's establishment record across a synchronous
observer report, so an Observer callback that subscribed to that workspace
would wait on a record only the reporting goroutine could retire. Moved the
SubscriptionCycled report past establishment. The narrower half is older than
this code -- confirmSubscription's late-acknowledgement path already reported
from inside that window -- so it is documented on the Observer interface as a
contract rather than silently worked around: a callback may publish, read and
unsubscribe; it may not subscribe.

Prose swept for what the gate falsified, per the team convention: the
constructor comment that argued for the defect, config.EventsHeartbeat's
rollback paragraph, the config test's inverted-rationale comment,
ResetReasonIdleTimeout, both metric Help strings, and deployment.md's phase
table and rollback section. All of them now say that phase 1 detects nothing
and that the cycled counter is STRUCTURALLY zero there -- a zero on phase 1
says nothing about whether a route has wedged, which is the reading an
operator would otherwise get wrong.

Claude-Session: https://claude.ai/code/session_01JVDBKbgn3Xt7ndW1YoYd8X

* test(events): prove a resuming joiner is told sync_required across a cycle

Codex round 2 raised that a subscriber arriving DURING an idle cycle gets no
gap signal, because dropWorkspaceCoverage only signals subscribers present
when it runs. True, and for a RESUMING caller the gap signal is not what
protects it: the registration mark is. It registers while the workspace has no
buffer, so its mark cannot match whatever buffer exists by the time it reads,
and eventsSinceMarkLocked answers nil -- sync_required rather than a false
"caught up".

A FRESH caller is deliberately not signalled and the finding is DECLINED for
that case, with reasons recorded at the test: it holds no prior position, so
there is no span it could be missing; it is admitted only after the
replacement subscription is acknowledged, because it waits on the cycle's
establishment record which finishPending closes after the confirmation; and on
the unconfirmed-admission path it IS told to reconcile when the acknowledgement
lands. Signalling it anyway would demand a resync of a client with nothing to
reconcile -- the load inversion this unit already had to fix once.

THE FIRST TWO VERSIONS OF THIS TEST DID NOT DISCRIMINATE, which is the part
worth keeping. Version one asserted the empty case: the cycle leaves no buffer,
so eventsSinceMarkLocked returned nil from its `!ok` term and removing the mark
check entirely still passed. Version two published inside
afterSubscriptionConfirmed so a FRESH buffer exists before the joiner reads --
and deleting the `mark.buffer == nil` term still survived, because the keep
arithmetic in that function already reduces to zero for a nil mark. Only
replacing eventsSinceMarkLocked with the unmarked eventsSinceLocked fails the
test, handing the joiner the post-cycle event as though it followed its cursor.
That is the mutation the test is built against, and the redundancy inside
eventsSinceMarkLocked is recorded rather than mistaken for coverage.

Claude-Session: https://claude.ai/code/session_01JVDBKbgn3Xt7ndW1YoYd8X

* fix(events): only count a cycle that actually replaced the connection (codex r3)

Three findings from a fresh-angle round on shutdown, wire format and doc
accuracy. The wire-format angle came back clean -- events:<workspace> cannot
collide with watchevents under validated namespaces, and no valid activity
payload can be mistaken for an hb| frame.

P3, and the one that stings: config.EventsHeartbeat still said phase 1
"already runs idle detection off whatever traffic exists". That is the exact
sentence the previous commit's sweep existed to remove, in a file that sweep
edited. A grep for the phrasing I remembered writing missed the paraphrase
sitting four lines above the paragraph I did fix.

P3: SubscriptionCycled was reported unconditionally after establishSubscription
returned, but establishment has two reasons to install nothing -- the bus
closed, or the workspace emptied while we dialled. The counter's documented
meaning is "torn down AND replaced", and counting an aborted establishment is
wrong in the direction that matters: an operator reading a non-zero rate
concludes connections are being blackholed, so a shutdown would manufacture
that signal. Now reported only when a replacement is installed, verified by
generation. Both Help strings and deployment.md say "counts replacements, not
teardowns"; the teardown stays visible through the idle_timeout reset reason.

P2: Close does not join the maintenance goroutines. Kept that way and
documented on Close, because the publish half makes synchronous Redis calls
bounded by go-redis's own timeouts -- the calls that stall on exactly the
wedged route this feature detects -- so joining would let a dead network hold
shutdown open. What has to hold instead is that a cycle already past its ctx
check leaves nothing behind, which is now pinned by a test that closes the bus
from inside the cycle's establishment: no subscription installed, no
establishment record stranded, no counter moved.

liveGen moved from the test file into the package -- production needs it now.

Claude-Session: https://claude.ai/code/session_01JVDBKbgn3Xt7ndW1YoYd8X

* test(events): restore the coverage the phase gate silently removed

The mutation matrix, re-run against the post-codex code, showed M3 -- removing
the install-time lastSeen stamp -- going from DETECTED back to SURVIVED. The
cause was my own round-1 fix: gating idle detection on heartbeat phase 2 means
a phase-1 bus never scans, and TestAnUnconfirmedAdmissionIsNotCycledAsIdle
built its own phase-1 bus. It was the only test that could observe a zero
lastSeen, because the plain fresh-subscription case is stamped twice over --
at install, and again by the acknowledgement. Flipped to phase 2 and
re-verified: removing the stamp fails it again.

Worth naming the shape rather than just the fix. A behaviour change that
narrows when code runs silently narrows what the tests reach, and nothing in a
green suite says so -- the tests still pass, they just stopped asking. Only
re-running the matrix after the change surfaced it.

Two harness bugs fixed alongside, both of which had been reporting
non-results as if they were readings:

- A mutation that INSERTS keeps its own anchor, so the "did the edit land?"
  check read every insertion as ANCHOR-ERROR. It compares the file now.
- The two rule-3 mutations left `sub`/`live` unused and came back BUILD-BREAK
  rather than answering the question; they carry the same discard the
  follow-up harness already used.

Claude-Session: https://claude.ai/code/session_01JVDBKbgn3Xt7ndW1YoYd8X

* test(events): close the wiring and barrier gaps codex round 4 found

Concurrency and lock discipline came back CLEAN -- the establishment record
and the generation checks cover two racing cycles, Unsubscribe, Publish and a
stale resubscription frame, with no lock-order deadlock. The four findings
were all about whether the tests measure what they claim.

P2, and it is the convention I had cited three commits earlier: the heartbeat
flip had no wiring test. internal/events proves a bus built with
publishHeartbeat=true emits frames and detects idleness, and every one of
those tests passes if newObservedEventBus hardcodes false -- the deployment
would simply never detect a wedged connection, which is indistinguishable from
a deployment that has none. Both directions asserted, because a helper that
ignored its config and hardcoded EITHER value passes a one-directional test.
Mutation-checked against exactly that edit.

P2: the metrics adapter test never touched SubscriptionCycled or the
idle_timeout reason, so an adapter that folded the counter into the reset
series -- destroying the very distinction those two are built to keep apart --
would have passed. Both added with counts that differ from their neighbours',
the pattern that file already uses so a label-dropping adapter cannot satisfy
the totals by coincidence.

P3: TestAHeartbeatConsumesNoEventID "waited" on a predicate that returned true
unconditionally. Not a slow wait -- no wait at all: the counter was read with
the publishes still in flight, so a heartbeat that DID consume an id could
land afterwards and the test would still pass. It now waits on the frames
arriving, and fails against a mutation that publishes an event alongside each
heartbeat.

P3: the maintenance goroutines started on phase 1, where both halves are
guaranteed no-ops -- two goroutines and two timers per process waking every
30s for the life of a deployment that asked for none of it, and phase 1 is the
DEFAULT. The flag is constructor-only so the decision is taken once. The
in-function gates stay: those are the correctness ones, and the tests reach
them directly without a loop.

Claude-Session: https://claude.ai/code/session_01JVDBKbgn3Xt7ndW1YoYd8X

* fix(events): validate the heartbeat frame and stop serialising recovery (r5)

Client-facing behaviour came back CLEAN: an idle cycle signals each local
subscriber, the SSE handler emits an in-band sync_required with an empty id
while holding the connection open, EventSource retires its cursor and the web
client runs the documented reconciliation. Two P2s on the other angles.

FRAME VALIDATION. Accepting any "hb|..." created a silently-ignored class on
the workspace event channel, where before this feature EVERY unreadable
payload ended coverage loudly and moved undecodable_message -- the counter
whose documented job is "suspect a namespace collision". A foreign or buggy
publisher whose bytes happened to start with the prefix slipped through that
signal without a trace. A frame is now hb|<version> plus optional short tokens
under a length cap; anything else wearing the prefix goes back to being a
coverage-ending decode failure, and the forward compatibility the prefix was
chosen for survives for a disciplined future frame.

What this deliberately does NOT try to fix, because it is not a hole: a forged
frame cannot fake liveness. Liveness means "this socket carried traffic", and a
frame that ARRIVES demonstrates exactly that whoever sent it -- which is why
stampLastSeen already fires for undecodable frames. There is no coverage claim
inside a heartbeat to forge.

CADENCE DRIFT, which was self-defeating rather than merely untidy. The timer
restarted after each pass, so the real period was T plus however long the pass
took. For the publisher that means an instance whose publishes are slow emits
heartbeats further apart, its own subscription sees them further apart, and it
can cross its own 3T threshold and cycle connections that were never wedged --
the slowness manufacturing the incident. Scheduling is deadline-based now, and
resets rather than bursting when a pass overruns badly.

SERIAL RECOVERY. One idle pass re-established every due workspace in sequence,
each re-dial bounded by go-redis's own timeouts, so recovery took N x that
timeout with the last workspaces reporting themselves uncovered throughout.
The failure that puts many workspaces on the due list at once is a Redis
failover, so the serial case was the common one. Bounded-parallel at 8 -- each
entry already owns its establishment record so they are independent by
construction, and an unbounded fan-out would answer a struggling Redis with one
dial per workspace at once. Test covers more workspaces than the cap, and
fails against a version that drops the overflow.

Claude-Session: https://claude.ai/code/session_01JVDBKbgn3Xt7ndW1YoYd8X

* docs(events): idle_timeout means coverage ended, not connection replaced (r6)

Codex round 6 came back clean on the non-Redis path (MemoryBus ignores the
Redis-only flag; EventBus and Close have not drifted), on the rollback
rehearsal (phase-2 to phase-1 and a mixed fleet are safe as documented,
including a bus mid-cycle -- Close cancels it, prevents installation and
retires its pending record), and on the operator surface
(PAD_EVENTS_HEARTBEAT is a server env/TOML setting; `pad configure` is client
connection config and needs no new surface).

The one finding is a contract drift I introduced two commits ago and then
wrote prose for in the same commit. Making SubscriptionCycled mean "replaced"
was right; what I missed is that the idle_timeout RESET REASON is emitted
earlier -- dropWorkspaceCoverage runs before the re-establishment -- so it can
fire when nothing is replaced, which is exactly the shutdown case the counter
was changed to exclude. Three doc sites and one log line said "replaced the
connection" anyway.

They now say what is true at the moment each fires: idle_timeout means
COVERAGE ENDED, only pad_event_subscription_cycled_total proves a replacement,
and the log says "attempting to replace" rather than "replacing". The log
wording matters on its own -- an operator correlating it with the counter
would otherwise find the log without the counter and go hunting a bug that
isn't there.

Third time this unit has produced prose the next change falsified, and each
time a different reviewer angle caught it rather than the sweep I ran at the
time. The pattern is that a behaviour change and the prose describing it land
in one commit, so there is no diff between them to notice.

Claude-Session: https://claude.ai/code/session_01JVDBKbgn3Xt7ndW1YoYd8X

* test(cmd): drive the heartbeat wiring test instead of sleeping at it (r7)

Codex round 7 found no leftovers across seven rounds of edits, and confirmed
the mass-cycle case does NOT produce a reconnect storm -- the SSE connections
stay open across a sync_required, so the admission limits are never consulted.

P3, and it is the failure I have been criticising in other people's tests: the
wiring test used a 300ms sleep as its ordering barrier. Under -race or on a
loaded CI box, a phase-1 bus that is correctly silent and a phase-2 goroutine
that merely has not been scheduled yet are indistinguishable, so the test could
pass or fail for reasons unrelated to the flip it exists to check. It now
drives one publish pass synchronously through a named test hook and uses an
ordinary event on the same channel as the barrier, which Redis delivers in
publish order. No timing left. Verified: still fails against the flag being
hardcoded false, and ten consecutive -race runs are green.

That replaces SetMaintenanceCadenceForTest with PublishHeartbeatsForTest rather
than adding to the exported test surface -- the loop's own wiring is covered
inside internal/events, where the unexported setter is available.

P2 is FILED, NOT FIXED, as BUG-2761: a mass coverage drop tells every connected
subscriber of every affected workspace to resync at once, and each browser tab
independently calls /changes with per-tab coalescing but no jitter and no
global budget. The fix is a web-client change plus possibly a wire-format hint,
which is independent of half-open detection and would materially expand this
diff. Worth filing rather than shrugging at because this unit makes the
simultaneous case MORE likely: it adds a third trigger of a class that already
existed (Redis failover, epoch change), and its natural cause is exactly a
network event that wedges many routes at once. deployment.md carries the
residual with the bug ref so an operator meets it before the incident does.

Claude-Session: https://claude.ai/code/session_01JVDBKbgn3Xt7ndW1YoYd8X

* test(events): make the tests prove what their comments claim (codex r8)

Round 8 was claim verification rather than bug hunting -- check the diff's
load-bearing assertions against the actual code -- and it was the highest-yield
round of the eight. The go-redis assertions (Ping writes without reading, the
channel path sets no read deadline, TLS dials ignore cancellation) and the four
claims about neighbouring functions all held. Seven other assertions did not.

TESTS THAT DID NOT PROVE THEIR OWN HEADLINE. This is the substance of the
round, and every one of these passed before and after:

- The JOINT TEST -- this unit's flagship -- claimed to discriminate the
  two-subscriptions failure and did not. Fan-out is per subscriber, so a joiner
  that opened its OWN second subscription still delivers the event to everyone
  exactly as the test expected. Nothing separates one subscription from two
  except counting them, which it now does at Redis, plus a duplicate-delivery
  check for the second receive loop. Fails against the pending record not being
  minted in the scan.
- The remedy test said "the old connection must also be gone" and waited for a
  receive-loop exit. stopRedisSubscription does two things and the loop exits on
  the first alone, so it passed against a version that cancelled the loop and
  left the PubSub and its health check open. Counted at Redis now; fails against
  exactly that mutation.
- The parallel-recovery test could not tell serial from parallel -- a serial
  pass cycles all thirteen workspaces too. It now uses a rendezvous, asserts the
  peak concurrency is above one AND within the cap, and fails against a serial
  implementation.
- The prefixed-garbage test only exercised the classifier. Whether
  receiveMessages ACTS on the error is a different claim, now driven through
  the real Redis path.
- The metrics adapter test's comment said "every reason this bus can emit"
  while subscription_unconfirmed was missing; its zero-assertion proved
  non-leakage, not mapping. Emitted now with a count distinct from its
  neighbour's, so a merging adapter cannot satisfy both.

PROSE THAT OUTLIVED THE CODE, again. The latency arithmetic still described the
single shared ticker that round 5 replaced with two independent loops; from
lastSeen [3T,4T) still holds, but from FAULT ONSET it is roughly [2T,4T)
because the publisher has its own phase. And a second "and replaces the
connection" in deployment.md that round 6's sweep missed.

Claude-Session: https://claude.ai/code/session_01JVDBKbgn3Xt7ndW1YoYd8X

* docs(events): correct three contract statements (codex r9)

Round 9 was cross-artifact conformance: every commitment the plan made was
checked against the code. All met -- wire classifier, lastSeen placement and
locking and install stamp and every-frame stamping, heartbeats bypassing
Publish and the shared counter, the drop-and-cycle remedy under the
single-establisher invariant, all four joint rules, the two-phase rollout with
its inverted-rationale test, and the reason/Help/deployment.md trio with the
rollout order. It also confirmed the three documented mutation survivors are
correctly dispositioned: both wsCounts checks are redundant-but-cheap under the
current invariant, and omitting the lastSeen.IsZero() skip is right because
adding it would mask a regression in the install stamp.

Three statements were wrong.

The env-var contract. My test comment said an unparseable PAD_EVENTS_HEARTBEAT
"must leave the flip off", which is true from a default config and false from a
config file that set it true -- there the value is left alone, as the
precedence test already asserts. The BEHAVIOUR is right and matches the epoch
flag: a typo must not move a migration in either direction, and silently
rolling an operator back to phase 1 would disable detection on a fleet that had
opted in with nothing saying so. Only the prose overclaimed, and it overclaimed
in the direction that invites someone to "fix" the ignore into a fail-closed
reset.

The constructor. NewRedisBusWithKeys documented publishEpoch and said nothing
about publishHeartbeat sitting next to it -- two adjacent booleans of the same
type belonging to two independent migrations, which is a shape that gets
swapped or dropped in a maintenance edit. Both now documented in order, with a
note that any combination is valid.

A stale count. EventSequenceResetsTotal's comment said "Five reasons" and there
are seven; it was already wrong by one before this unit added another. Replaced
with the count plus a pointer to the three artifacts that are authoritative and
move together, since the count itself is the part that goes stale first and is
read last.

Claude-Session: https://claude.ai/code/session_01JVDBKbgn3Xt7ndW1YoYd8X

* test(events): make the cadence arithmetic testable, and justify a guard pair

Matrix 5 (29 mutations, 21 detected) surfaced two things the previous run
could not, because both concern code the codex rounds added.

THE DRIFT FIX HAD NO TEST. Restoring the sleep-after-work form survived every
test in the package, and would have kept surviving: the only way to observe
drift through the loop is to time it, and a timing assertion is a flaky
assertion. Extracting nextTick makes the arithmetic checkable without a clock,
and the four cases now pin what the schedule is for -- a slow pass does not
push the next tick out, ten slow passes accumulate no drift, an overrun beyond
one interval resets instead of replaying the missed ticks, and an overrun
WITHIN one interval still catches up rather than re-phasing the schedule
permanently. Both directions mutation-checked.

The property is worth this much because breaking it is self-defeating rather
than merely untidy: an instance whose passes are slow emits heartbeats further
apart, its own subscription sees them further apart, and it crosses its own 3T
threshold and cycles connections that were never wedged.

A GUARD PAIR THAT ONLY DIES TOGETHER, which the team lesson says to treat as a
question rather than a clearance. The loop's ctx.Done select arm and its
post-wait ctx check each survive removal alone. Checked rather than assumed:
they cover disjoint moments and each is independently right -- the select arm
is the exit while WAITING, which is where the goroutine spends its life, and
the post-wait check stops a bus that closed DURING a pass from starting
another one against a cancelled context and a drained wsSubs. Removing BOTH is
detected. Reasoning recorded at the code, and the combined mutation added to
the matrix so the pair cannot quietly become a single point of failure.

Also fixed an ineffassign the lint gate caught in the new test.

Claude-Session: https://claude.ai/code/session_01JVDBKbgn3Xt7ndW1YoYd8X

* docs(events): state what the detector does not cover (codex r10)

Round 10 was adversarial: refute the unit's central claim rather than look for
defects in it. It partly succeeded, and the corrections are worth more than
most of the bug findings.

The claim was "a wedged connection is detected, coverage is ended, and the
connection is replaced so delivery resumes". Three parts of that were too
strong, and all three limits were checked against go-redis v9.22.0 rather than
argued:

IT IS A RECEIVE-SIDE DETECTOR, not a round-trip health check. It measures
whether frames ARRIVE. A subscription whose outbound direction is broken but
which still receives reads as healthy -- correctly, since nothing is lost, but
that is a narrower claim than "the connection is healthy".

IT CANNOT COVER THE PUBLISH PATH. PUBLISH travels on the client's connPool
while a subscription holds a connection from the separate pubSubPool
(redis.go:363, :1956) -- different sockets, different fates, and a reconnect of
one repairs nothing about the other. An instance whose publish path is wedged
loses its own events for every other instance and this feature will not say so.
That is a real gap in the family's coverage, now written down rather than
implied away.

REPLACEMENT IS ATTEMPTED, NOT GUARANTEED. If the path is still blackholed when
the cycle re-dials, the replacement cannot receive either. Coverage stays ended
so nothing is falsely claimed, but delivery resuming is a statement about the
network rather than about this code.

Filed BUG-2764 rather than folded in: establishSubscription's
`b.client.Subscribe(dialCtx, channel)` silently discards the SUBSCRIBE error,
because go-redis's own Client.Subscribe drops it (`_ = pubsub.Subscribe(...)`,
redis.go). A failed subscribe therefore installs a connection that looks live
and is subscribed to nothing. It is pre-existing, it lives in the establishment
path three bugs have already converged on, and changing how that function
issues its SUBSCRIBE does not belong in a diff about idle detection. Worth
knowing here because it is the one way the replacement can fail on a HEALTHY
network -- and because the detector now cycles it on the next pass, which is
why it self-heals on phase 2 and stays dead forever on phase 1.

Claude-Session: https://claude.ai/code/session_01JVDBKbgn3Xt7ndW1YoYd8X

* fix(events): do not cycle a workspace that recovered before its turn (r11 P1)

Codex round 11 attacked three claims. Phase-1 safety and rollback safety both
came back clean -- a phase-1 receiver stamps lastSeen and nothing else, touches
no buffer, metric, client, ID or epoch, and its maintenance loop is not started
at all, so that timestamp is inert; heartbeats leave no state in Redis or
across a process replacement, and a mid-cycle shutdown rechecks b.ctx before
installing. The third claim did not survive.

FALSE POSITIVES ON A HEALTHY SYSTEM, which is the property this design cares
about most: cycling a working subscription drops its coverage and resyncs every
one of its subscribers for nothing.

cycleIdleSubscriptions selects its victims under the lock and releases it; the
cycles run afterwards. Its re-checks asked about generation, subscriber count
and bus liveness -- and never re-asked the question the scan had asked. A
subscription that started receiving again in that window was cycled anyway.

The window is not theoretical, and this unit widened it itself: the 8-way
concurrency cap added in round 5 makes a workspace wait behind earlier batches
of slow replacement dials, and a GC or CPU pause leaves a backlog of heartbeats
undrained in the receive loop. Both are ordinary conditions on a loaded box.

cycleOne now validates, ends coverage and tears down WITHOUT RELEASING THE LOCK
in between, which needed dropWorkspaceCoverage split into a locked variant that
returns its reason for the caller to report after unlocking. That also removes
the ordering fragility the previous version documented rather than fixed: there
is no longer any window in which coverage is ended for a workspace this
function then decides to leave alone. The log moved after the decision for the
same reason -- it could previously describe a cycle that then abandoned.

The freshness term is load-bearing and says so, next to the three neighbouring
terms whose mutation survivals are recorded as redundant-but-cheap. Removing it
is detected, by a test that lands the recovery in the exact gap through a new
positional seam.

NTP steps were checked and are not a hazard: time.Time carries a monotonic
reading, so a wall-clock step cannot make a subscription look idle.

Claude-Session: https://claude.ai/code/session_01JVDBKbgn3Xt7ndW1YoYd8X

* perf(events): take logging and PubSub.Close off the global lock (codex r12)

Round 12 verified round 11's freshness fix: validation, coverage invalidation
and teardown are atomic under b.mu with no lock cycle,
dropWorkspaceCoverageLocked preserved the original semantics exactly including
the no-buffer branch that still signals subscribers, reset reporting happens
after unlocking, and the replacement metric still lands only when a new
generation does. Slow establishment stays outside b.mu, wg.Wait only delays the
next pass, and Close cancellation retires pending records.

Two P2s, both about what round 11 put UNDER that lock:

slog.Warn ran while b.mu was held. slog invokes the installed handler
synchronously, and b.mu is the lock every fan-out and every Subscribe on the
instance contends for -- a slow or custom handler stalls all of them, and one
that calls back into the bus deadlocks. Moved after the unlock; it still has to
come after the DECISION, for round 6's reason, so both constraints are now
stated together at the call.

PubSub.Close ran under b.mu too. It takes go-redis's own mutex, which the
health check can hold across reconnect work, so a network-bound wait sat inside
the instance's hottest lock. That was survivable when teardown only happened as
a workspace lost its last subscriber; the idle detector makes it happen on
every cycle, which is what turned a latent cost into a real one. Handed off to
a goroutine: nothing references the PubSub once the map entry is gone, and
cancel() -- which is what actually stops delivery -- still happens under the
lock.

Claude-Session: https://claude.ai/code/session_01JVDBKbgn3Xt7ndW1YoYd8X

* fix(events): do not read our own failed probe as a dead peer (codex r13)

Round 13 asked for a production-approval review. Four findings; the second is
the sharpest of the whole run because it is the mirror image of the failure
this feature exists to find.

A FAILED HEARTBEAT PUBLISH WAS READ AS A DEAD SUBSCRIPTION. The detector's
inference is "we published a frame and nothing came back, so the receive path
is dead" -- valid only if the publish actually happened. PUBLISH travels on the
client's connPool while the subscription holds a connection from the separate
pubSubPool, so a publish-side failure (pool exhaustion, a wedged outbound
route, Redis refusing writes) says nothing about whether that subscription can
receive. The detector was reading its own inability to probe as evidence about
the peer, and tearing down healthy connections on a schedule: a resync for
every subscriber of every workspace, every 90s, for as long as the outbound
path stayed broken. The third load inversion this unit has had to fix.

redisSub.lastProbeOK now records the last SUCCESSFUL publish, and detection is
suspended while it is stale -- checked in the scan and again in cycleOne, which
is a pair that only dies together and is therefore justified at the code:
the scan's keeps a workspace off the due list so no record is minted and no
joiner waits, cycleOne's covers the probe failing AFTER selection, a window the
concurrency cap makes real. Neither subsumes the other; removing both is
detected. New counter pad_event_heartbeat_publish_failures_total, documented as
DETECTION DEGRADED rather than as a peer being broken.

THE END-TO-END TEST THAT DID NOT EXIST. Every other test drives this through a
fake clock -- necessary, since the threshold is 90s by construction and
miniredis always answers, but it means they all ASSUME the wedge rather than
produce it. A TCP proxy that stops delivering server->client on the connections
already open, while writes keep succeeding and new connections stay healthy,
produces the real thing. The test asserts both halves of the claim: the wedge
is detected, and the replacement delivers. Both halves mutation-checked
(detector disabled; drop-only with no replacement).

The proxy's first version was vacuous -- a global flag consulted at read time
meant re-enabling delivery for future connections also revived the ones meant
to be dark. Per-connection now, and the comment says why.

Also: PubSub.Close taken off b.mu in Close() too (round 12 fixed only the cycle
path), and the replacement counter now takes an explicit installed result from
establishSubscription rather than inferring one from the live generation --
inference misattributed an unrelated caller's fresh subscription as this
cycle's replacement, and missed a real replacement that had lost its last
subscriber. Both mutation-checked.

Claude-Session: https://claude.ai/code/session_01JVDBKbgn3Xt7ndW1YoYd8X

* fix(events): bind the probe stamp to a generation; make the proxy test honest

Round 14 returned a BLOCK verdict on two P2s, both mine, both in the fix that
round 13 had just added.

lastProbeOK WAS NOT GENERATION-BOUND. publishHeartbeats snapshots the workspace
list, publishes off the lock -- for as long as go-redis's timeouts allow -- and
then stamped whatever subscription occupied that workspace by the time it
returned. A probe sent for generation A could credit generation B, which never
received one; if later probes then failed, B could be cycled while looking
recently probed. Exactly the hazard stampLastSeen already guards on the same
map, and I did not carry it across. The generation now travels with the
snapshot and is validated before stamping.

THE END-TO-END TEST COULD PASS WITHOUT EXERCISING WHAT IT CLAIMED. It darkened
the receive direction of every open connection, including the ordinary pooled
connection PUBLISH uses -- so the probe may have been failing too, and the run
would then have been exercising the cannot-probe path rather than a half-open
route, which is the very distinction round 13 added the premise check for. The
proxy now classifies connections as it forwards and darkens only one that has
carried a SUBSCRIBE, leaving the publish path healthy, and the test asserts
zero probe failures so a run that drifts back into the other case fails loudly
instead of passing quietly. Still fails against a disabled detector and against
drop-only.

Also covered the new counter's mapping in the metrics adapter test, with a
count distinct from both neighbours -- cycled, idle_timeout and
heartbeat-publish-failure say three different things and an operator acts on
the difference.

Verified by the same round: install-time stamping does not permanently suppress
detection, establishSubscription returns false only on abandon and true on all
three installed paths including the cancelled-establisher goroutine, and
Close's deferred PubSub.Close runs after the unlock.

Claude-Session: https://claude.ai/code/session_01JVDBKbgn3Xt7ndW1YoYd8X

* test(events): pin the probe-across-replacement race (closes r15's residual)

Round 15 returned CLEAN and approve-with-comments, naming one residual: the
generation binding on lastProbeOK had no deterministic test, only the argument
that it mirrors stampLastSeen. This closes it with a positional seam between
the publish and the stamp, which is the only place that interleave can be
forced.

TWO INSTRUMENT DEFECTS ON THE WAY, both caught by mutation rather than by
reading:

The first version compared the credited stamp against the PROBE's timestamp.
On a frozen clock the replacement's install stamp and a wrongly-credited probe
are the same value, so it could not tell them apart -- it failed on the install
stamp while claiming a credit had happened, and removing the generation binding
still passed. It now compares against what the replacement was INSTALLED with,
and the clock advances inside the seam so a buggy write lands strictly later.

The second version was FLAKY: 2 failures in 3 runs. The heartbeat that was just
published comes back through miniredis on another goroutine, and if it lands
between the forced-stale write and the scan it refreshes lastSeen, the
workspace is not due, and no replacement happens. Retried until the generation
actually moves. Now 5 of 5 green unmutated and 5 of 5 detected mutated -- which
is the bar, because a 2-in-3 detector reads as coverage while being noise.

Claude-Session: https://claude.ai/code/session_01JVDBKbgn3Xt7ndW1YoYd8X

* fix(events): on-call signals — log the cycle outcome, correct two claims (r16)

Round 16 read the diff as the person paged at 3am. Four findings.

THE CYCLE LOGGED ITS ATTEMPT AND NEVER ITS OUTCOME. The line says "attempting
to replace", which is correct and, on the one path where the replacement does
not happen, left an on-call with a warning, no counter movement, and no
explanation. Now there is a second line naming the reason.

pad_event_receive_loop_exits_total's documentation was falsified by this unit
and neither doc site said so: every idle cycle stops a receive loop while its
subscribers are still connected, and the comment still claimed exits happen
only at shutdown or when the last subscriber leaves. Both sites corrected, with
the expectation that it tracks the cycle counter during an incident.

A CLAIM I MADE AND THEN COULD NOT SUPPORT, recorded rather than quietly kept.
Round 16 argued the age-based premise check ("has a probe succeeded within the
threshold") failed to suspend detection where an ordering rule ("has a probe
succeeded since anything last arrived") would, and I rewrote the rule on that
argument and wrote a test named for the defect. The mutation matrix then
refused to confirm it: reverting to the age form leaves the test green, and so
does removing both copies of the check, and no case separates the two — on any
healthy path the two stamps advance together, because a probe whose frame
arrives sets both, and they diverge only on the wedge where both forms cycle.

The ordering rule is kept, because it states the intent exactly and is never
weaker. But the test and the comment now say what they actually establish —
that a probe which has started failing stops the detector concluding from
silence, which is the property both forms share and neither had before — rather
than claiming a fixed defect I cannot demonstrate.

The two remaining P2s are already-filed residuals: the cycled counter proves an
install rather than a working replacement (BUG-2764), and repeated cycling
amplifies /changes load with no jitter or global budget (BUG-2761). Both are
documented in deployment.md with their refs.

Claude-Session: https://claude.ai/code/session_01JVDBKbgn3Xt7ndW1YoYd8X

* docs(events): record what the final matrix actually says about four guards

Final matrix: 34 mutations, 22 detected, baseline restored green. Every
survivor is now documented at its line with why it survives, and two of them
turned out to be instrument defects rather than coverage gaps.

lastProbeOK's INSTALL STAMP IS REDUNDANT and the comment claimed otherwise. It
said a zero value "would permanently disqualify a subscription from ever being
cycled" -- true of the age-based premise it was written for, false under the
ordering rule that replaced it, because a zero value fails
`lastProbeOK.After(lastSeen)` exactly as an install stamp equal to lastSeen
does. Kept, for a reason it earns: it makes the field's invariant true by
construction, so a future rule reasoning about this value's AGE gets a real
timestamp rather than 1970 -- which is the trap the age-based rule fell into
one field over.

THE TWO cycleOne ABANDON GUARDS DIE ONLY TOGETHER AND ARE NOT REDUNDANT, which
took checking rather than assuming. They catch different shapes of the same
recovery: an arrival that has not been re-probed pushes lastSeen past
lastProbeOK so the premise case fires and the freshness case is unreachable --
that is the shape the test produces, and it is why removing either alone stays
green. But the publisher runs on its own goroutine at its own cadence and can
land a successful probe between the arrival and the decision, putting
lastProbeOK ahead again; there only the freshness case stops a healthy
subscription being torn down. Deleting it on the strength of the matrix would
remove the second shape's only guard.

Close's off-the-lock PubSub.Close is UNTESTED BY DESIGN, recorded rather than
papered over. It is a contention property, and the only assertion that
separates it is a timing one, which in this suite is a flaky one.

TWO HARNESS DEFECTS, both of which produced false survivors that would have
gone into the evidence package as findings. M11a inserted its mutation AFTER
the gate it was meant to disable -- unique anchor, wrong placement, so the
early return still fired and nothing changed; with a correct anchor it is
detected. M20 left variables unused and came back BUILD-BREAK rather than
answering; in compiling form it genuinely survives, consistent with
establishSubscription's abandon path already covering it.

The lesson worth keeping: when I rewrote all 34 anchors against current source
I verified each matched exactly ONCE, and uniqueness is not placement. An
anchor can be unique and still land somewhere that changes no behaviour.

Claude-Session: https://claude.ai/code/session_01JVDBKbgn3Xt7ndW1YoYd8X

* test(events): barrier the probe test on delivery — it was flaky, CI caught it

Go (PostgreSQL) failed on af7001ab, in a test I added two commits ago. Not a
timeout and not the race step: TestAFailedProbeAfterASuccessfulOneStillSuspends
Detection asserted no cycle and got one.

The test killed Redis immediately after a successful probe, without waiting for
that probe's frame to be delivered back. If the frame never lands, lastSeen
stays at the install stamp, the successful probe is then legitimately "after
the last arrival", the workspace is genuinely due — and the code cycles it FOR
THE RIGHT REASON under a test asserting it should not. The premise the test is
named for simply did not hold on a slower machine.

So this was not a false alarm in CI and not a defect in the code: it was my
test asserting an outcome whose precondition it never established. Waiting for
lastSeen to move makes the precondition real. Eight consecutive local runs
green, and removing both premise checks still fails it, so the barrier did not
neuter what it was measuring.

Worth naming because it is the third instrument defect in this unit found by
something other than reading it — after the harness restore that ate an edit
and the unique-but-misplaced mutation anchor. A test that depends on an
unsynchronised delivery is a test that passes on the machine that wrote it.

Claude-Session: https://claude.ai/code/session_01JVDBKbgn3Xt7ndW1YoYd8X
2026-08-24 16:57:52 -04:00
xarmian 5003718802 fix(push): apply delivery's visibility gate to delivered_sessions (BUG-2725) (#1187)
deliveredSessionCount applied three of watchNotificationVisible's four
gates, missing the first thing delivery checks: vis.allows(CollectionID,
ItemID). Broadcast over-reported. Targeted was worse — the publish-skip
reads this count, so the gate passed, the push went out, the stream
dropped it on visibility, and the response said delivered_sessions: 1.
An instruction lost behind a success.

Per Dave's day-49 ruling, visibility is RE-RESOLVED at push time rather
than snapshotted: membership and grants are revocable, so a value cached
at connect goes wrong exactly when revocation is what makes it matter.

The one input that cannot be re-resolved is the target connection's auth
transport — computeWatchAccessVisibility consults isBearerAuth exactly
once, inside the admin bypass, and the pushing request only knows its
own. So SessionOrigin.BearerAuth is recorded at Add(). That is NOT the
snapshot the ruling rejected: auth transport is a property of the
connection, fixed when it opened and not revocable while held, so it
cannot go stale. Armed is the precedent. SessionOrigin is kept separate
from SessionIdentity because that type documents itself as self-declared
and never verified; folding a server-derived security fact in there
would silently retract the warning for one field. Both comments state
the rule for future extenders: connection properties are admissible,
derived authorization state never is.

computeWatchAccessVisibility now takes a bool instead of an
*http.Request, which makes the per-connection input visible in the
signature and lets the count answer for a connection it is not serving.

COST: "re-resolve per counted session" reads like N access checks per
push. It is at most TWO, and sessionVisibility's memo makes that true by
construction rather than by careful calling — every other input is
per-user and identical across the sessions counted, so one varying
boolean bounds the answers at two. Pinned by a test with 50 sessions.

Codex round 1 (P1): the first version swallowed store errors into "not
visible", reintroducing BUG-2698 through this fix — a targeted push
reporting 0 SKIPS the publish, so a DB blip would drop the instruction
and answer 200, in a function whose own doc comment says why 0 is
load-bearing. Round 2 (P1): the same class one layer down —
computeWatchAccessVisibility collapsed FOUR store failures into a
denial, two discarded into underscores. Fixed as a class per CONVE-18.
Resolution and policy are now separate: stream-side callers discard the
error explicitly with reasons, only the counting caller propagates.
Round 3 CLEAN.

CONVE-23 sweep found three consumer-facing artifacts still describing
the old mechanism, none on a line this diff touched: the plugin skill
doc, the web push dialog, and pad push --help. All three corrected to
name what actually remains rather than deleting the caveat. Plugin
0.3.1 -> 0.3.2, since installed plugins are version-pinned at install.

NOT fixed, deliberately: the UNDER-count. A stream past
maxSessionsPerUser receives broadcasts while never entering the
registry. delivered_sessions remains an estimate with error in both
directions, and every consumer-facing description now says so.

Two coverage gaps recorded rather than rounded off: mutation M11
survives (the reporting test reaches only the first of four store calls,
because closing the DB fails it first), and no test drives the whole
chain store-fault-to-503 (the DB-close instrument kills the request
earlier, so such a test would have gone green against the wrong 500 —
deleted rather than relaxed).

Also lands the BUG-2752 refutation sentinel: that item claimed the OAuth
workspace allow-list went unenforced on /api/v1/events/stream. Refuted —
no allow-list-bearing credential can authenticate to /api/v1/* at all.
The test guards that format gate, so if it ever widens, the refutation's
premise fails loudly instead of silently reopening a leak.

Gates on the merged tip: make test 27 pkgs, make lint 0 issues, full
Postgres suite 27 pkgs, govulncheck, codex CLEAN, CI 7/7.

Claude-Session: https://claude.ai/code/session_01Cpr3teiHHsgcTmg2xhHA86
2026-08-24 09:45:48 -04:00
xarmian 72336aacb5 fix(events): release SSE admission slots when a client leaves mid-establishment (BUG-2749) (#1186)
`GET /api/v1/events` reserved its admission slot, then blocked in
`SubscribeAndReplaySince` while the workspace's Redis subscription was
dialled and — since BUG-2747 — acknowledged. Nothing propagated the
request's cancellation into that wait, so a client that disconnected
during establishment left a process-wide slot, a per-principal slot and a
per-workspace slot held for the whole of it. The connection was gone; the
capacity was not.

Cancellation is now DEREGISTRATION, and `wsCounts` — which already
answers "is anyone still here" — decides everything downstream. No
ownership hand-off and no reaper: the arbiter already existed. (One thing
IS handed off, and only one — the remainder of the confirmation wait; see
below.)

The two cancellation positions take different paths, and only one of them
owes the joiners anything:

- Before the install: the existing post-dial critical section already
  abandons and retires correctly when nobody is left. It needed one
  ordering rule — the departed establisher stops being counted IN THAT
  SAME SECTION, before the count is read. If joiners registered while we
  dialled, the count is still non-zero and they get the subscription;
  that is the hand-off the filing asked about, expressed as a count
  rather than a transfer of ownership.
- During the confirmation wait: the subscription is already installed
  with its receive loop running, so the connection is not at risk — but
  the WAIT is what releases the joiners, and dropping it would admit them
  into a subscription Redis has not acknowledged while telling them
  nothing. That is BUG-2747's defect re-created at the seam between the
  two designs. So the remainder of the wait moves to a goroutine that
  finishes exactly as the caller would have: same arms, same
  `markUnconfirmedAdmission` on the bound, same `finishPending`. Bounded
  by `confirmTimeout`; no reaper needed, because teardown stays
  count-driven.

A departure is not a refusal. `ok bool` is replaced by a
`SubscribeOutcome` enum across the three `EventBus` Subscribe methods, so
`SubscribeWorkspaceLimit` and `SubscribeCancelled` cannot be collapsed:
answering a departed client with 429 would have written a limit refusal
into the logs and counters that anyone would use to tune that limit. An
enum rather than a second bool or an error because the switch has to name
the case — by construction rather than by argument.

Caller population, with its search boundary: 3 production implementations
(events.MemoryBus, events.RedisBus, metrics.InstrumentedBus), 1 test
double (server.gapEventBus, which embeds the interface), 2 production call
sites (both in handlers_events.go). Searched this repo four ways — the
three method names, `.Subscribe(`, method declarations, and interface
embedding. collab.OpBus and watchevents.Bus are different interfaces and
are out of scope; no other repo links this package.

WHAT THIS DOES NOT FIX, verified in go-redis v9.22.0 rather than inferred
from its doc comment (which says Subscribe "does not wait on a response
from Redis" and so reads as though no dial happens on the request path —
it does; only the reply is unawaited). On plaintext, dialConn derives its
per-attempt deadline from the caller's context and the default dialer is
net.Dialer.DialContext, so cancellation aborts the dial. Under TLS the
same dialer calls tls.DialWithDialer, which takes no context, so the dial
stays bounded by DialTimeout alone. On a TLS deployment this shrinks the
held slot from (dial + confirm bound) to (dial), not to zero.

Review round 2 (codex) found a P1 in this unit's own first draft, of
exactly the shape the filing warned about. A cancellation check at the top
of the establish loop could return while the caller still OWNED an
unretired establishment record: section 1 had already named it the
establisher, so the record stayed in pendingSubs with nobody behind it,
its done channel never closed. The next subscriber for that workspace
would join it and wait forever — and its own registration keeps wsCounts
non-zero, so no later caller would establish either. A permanently dead
stream that looks alive, produced by a guard whose only purpose was to
save a dial. The guard is gone: a cancelled caller now goes THROUGH
establishSubscription, which is the only code that knows how to put the
record down. Regression test included, and reinstating the guard turns it
red.

Round 2 also found a P2 shutdown regression: routing the dial to the
caller's context alone took away Close()'s ability to interrupt a stalled
dial, which it had before. The dial now runs on a context ended by EITHER
the caller or the bus, and each half is pinned by its own test — dropping
either one is detected.

Review round 1 (codex): no P1. One nit fixed as a class — three comments
elsewhere in the file asserted the dial was "NOT bounded by the context we
pass", which this change falsified; the sweep found and corrected all
three (establishSubscription, defaultSubscribeConfirmTimeout, Subscribe).
The TLS half of its P2 is filed as BUG-2754: the fix belongs at client
construction, where it covers every Redis call rather than this one.

Class sweep filed separately as BUG-2751 (lead-ruled: one region, one
design per diff): internal/watchevents has no per-request establishment,
but its resume path blocks on a 250ms settle window bound to the bus's
context rather than the request's, while /api/v1/events/stream holds the
same admission slots across it.

Tests: five cancellation cases in internal/events (before install, during
the wait alone, during the wait with a joiner, a cancelled joiner, an
already-dead caller), a dial-binding assertion, and the handler-level
binding in internal/server asserting the admission slot itself is
released — the half of the bug that does not live in the bus.

Mutation matrix, 8 mutations: 7 detected, each by the test named for it.
The one survivor is the ctx term in the retry re-decide, and it survives
because it is an OPTIMISATION rather than a correctness guard — a departed
caller that mints a second record still establishes, deregisters and
retires correctly; the term only saves a pointless dial. The code says so
rather than implying the guard is load-bearing.

The earlier draft's entry guard and loop-top break formed a redundant pair
the matrix could only detect when both were removed. That redundancy was
the smell, and round 2 found the substance under it: one of the two was
not redundant, it was wrong. With it gone the entry guard is detected on
its own.
2026-08-24 08:24:34 -04:00
xarmian 2e9ace4194 docs: nine overclaims across code, metrics, docs and the CLI (BUG-2739, codex round 5)
A cross-artifact pass, which is the angle that keeps paying on this family.
Every item below was a statement of mine that was false or unsupported; the
code did not change.

WRONG FACTS:
- Watch epochs are opaque UUIDs, not numeric generations. I had copied
  internal/events' wording, where they ARE numeric — the distinction is the
  subject of internal/idspace's package comment.
- undecodable_message was described as proof a notification was missed. The
  instance knows only that something it could not read arrived on its
  channel; it cannot tell whether that was ours. It stops vouching BECAUSE it
  cannot tell, which is a different and weaker claim. Corrected in four
  places.
- The failover-cost paragraph said every SSE client on the instance
  reconciles. Wrong twice: a watch-bus resubscription ends the WATCH stream's
  coverage (activity coverage is per-workspace), and the one client that uses
  that stream today — pad watch --stream — answers sync_required by clearing
  its cursor and keeping the connection open, so it issues no request at all.
  Verified in cmd_watch.go rather than assumed.
- The midstream/reset ratio is not fan-out in aggregate: the announcement
  counter also carries gaps and slow-subscriber drops and coalesces per
  connection. Only a reset observed in isolation reads that way.
- 'The watch stream's only signal was a later non-contiguous notification'
  is true for a client HOLDING A STREAM OPEN. A reconnecting client was
  always covered, because a resume asks the shared counter instead of local
  state. Scoped in the doc and in the test header.
- The dropped-confirmation fallback said coverage still ends. Usually, not
  necessarily: with no traffic during the outage nothing was lost, and if
  the drops continue through whatever would expose the hole and the stream
  goes quiet, nothing ever does — BUG-2727's boundary. Named both.
- The new Observer Close warning was overbroad: reports run on the receive
  goroutine only on the RedisBus receive path, while a ResumeGap runs on the
  caller's and MemoryBus has no such goroutine. The rule stays
  unconditional, since a callback cannot tell which case it is in, but it
  now says why.

STALE AFTER THIS BRANCH:
- metrics.go's WatchSequenceResetsTotal comment listed two reset reasons.
- observer_test.go said 'both reset reasons'.
- The constructor comment named Channel() after the loop moved to
  ChannelWithSubscriptions, and did not say the Receive beneath it is
  load-bearing for that loop having no skip-the-first flag. It does now, and
  names the test that fails if it goes away.
- cmd_watch.go's sync_required cause list predated BUG-2739 (and did not
  mention the mid-stream delivery BUG-2730 added).

Claude-Session: https://claude.ai/code/session_01JVDBKbgn3Xt7ndW1YoYd8X
2026-08-23 13:33:58 +00:00
xarmian b7ae022b6f refactor(events): every way of subscribing hands back the gap signal (BUG-2730, codex round 5)
Subscribe allocated and raised a gap channel its callers could not read,
which round 5 called dead work. The read is right and the disposition is
the other one: an interface method whose subscribers CANNOT be told they
missed something is a silent under-delivery waiting for its first
production caller, and internal/watchevents' Subscribe already returns
the signal, so the asymmetry was the defect rather than the allocation.

Subscribe now returns it too, on all three implementations. No production
caller changes — the handlers use SubscribeIfAllowed and
SubscribeAndReplaySince — so this is a test-call-site sweep plus one
signature.
2026-08-23 01:01:59 +00:00
xarmian 6afe683389 fix(events): do not arm reset detection where interleave is ordinary traffic (BUG-2736)
Codex round 9, from the 3am-operator angle. Six findings; one of them was a
regression this diff would have shipped in the DEFAULT configuration, and the
review framed it as a log-volume problem.

THE REGRESSION. Phase 1 publishes with a two-call INCR-then-PUBLISH, so on any
multi-instance deployment two publishers interleave routinely and a lower ID
arrives after a higher one as ordinary traffic. main has no counter-backwards
detection at all; this diff added it. Armed unconditionally, it would have
fired on that ordinary interleave, dropped EVERY workspace's replay buffer,
and resynced every client -- in phase 1, which is where every deployment sits
until an operator flips phase 2.

The check is now armed only once an epoch has been adopted. What that costs is
stated rather than hidden: a genuine counter reset on a never-flipped
deployment goes undetected, which is exactly the behaviour before this change
and precisely the case phase 2 exists to fix.

The new test asserts the gate, and also asserts what is NOT claimed -- the
interleaved workspace's own buffer still holds ids out of order, so a cursor
at the higher one reads as foreign. That is pre-existing, unchanged here, and
strictly less harmful than a global drop; it is asserted rather than described
so a future change to since() surfaces there.

THE REST ARE THE OPERATOR'S SIGNALS, which were unreadable:

- The effective phase was invisible. pad_event_sequence_resets_total cannot be
  interpreted without it -- a counter_backward rate is expected on phase 1 and
  an anomaly on phase 2 -- and the setting can arrive from an env var, a TOML
  file, or neither. It is now on the startup line as id_space_phase.
- An unparseable PAD_EVENTS_PUBLISH_EPOCH was silently ignored, so an operator
  who typed "yes" believed they had flipped. Ignoring it stays the right
  behaviour; being silent about it does not.
- Both publish-failure logs said only "failed to publish". They now say what
  the operator needs, which differs by phase: phase 1 may or may not have
  reached subscribers, and phase 2's script is atomic so it did not
  half-execute, but a lost reply means it may have published anyway -- do not
  re-publish by hand.
- Adopting an epoch with empty buffers is the moment the documented residual
  becomes possible on that replica, and it happened silently. It now logs at
  INFO -- not a reset count, deliberately, since counting it would give the
  reset metric a per-deploy baseline.

Declined with reasons: a cause label on the resume-gap counter and a publish
failure counter are both pre-existing shapes rather than anything this diff
changed, and the straggler log is already bounded by the recovery window.

Claude-Session: https://claude.ai/code/session_01JVDBKbgn3Xt7ndW1YoYd8X
2026-08-22 20:05:48 +00:00
xarmian 736a8c48f7 docs: thirteen claims about code that had moved under them (BUG-2736)
Codex round 5, cross-artifact consistency. Every one was a claim in a comment,
help string, doc, or test name that the code no longer supported — and this
diff created most of them by moving the code.

The ones that would have misled an operator:

- pad_event_sequence_resets_total documented ONE reason in both the Go doc
  comment and the Prometheus help text, and the deployment table said the same.
  It has emitted three since this branch. An operator reading the help string
  to build an alert would have alerted on a third of the signal.
- deployment.md said every published message carries an epoch prefix. Phase 1
  publishes bare JSON — which is the entire point of having two phases.
- deployment.md said the first flipped message reaches each replica and every
  resuming client gets sync_required. A replica learns the epoch only from a
  message it RECEIVES, so only replicas subscribed to a workspace with traffic
  see it; and a replica with empty buffers adopts without dropping or
  counting, deliberately.
- deployment.md said a restart's IDs cannot collide. internal/idspace documents
  a bounded case — the earlier process publishing more than 2^20 events per
  millisecond of its life. Stated as the bound it is, with the
  backwards-clock direction named as the safe one.
- cmd_watch.go described sync_required as eviction-only. It has had four other
  causes since BUG-2731 and gained a fifth here.

The ones that would have misled the next person editing this code:

- bus.go said the Redis half was unwritten and a reset counter could still
  merge two ID spaces. It is written, three commits back on this branch.
- bus.go and watchevents.go said in-memory IDs restart from 1. They count from
  an incarnation base.
- redis_bus.go described this bus's epoch as an opaque uuid equivalent to the
  watch bus's, twice, after round 3 made it a Redis-minted generation. Only
  the watch bus still uses uuids.
- observer.go said counter_backward happens only during mixed-version rolls.
  Phase 1's two-call publish produces it in steady state too.
- redisns.go said the publish script spans four keys (it is five here now, plus
  a two-key assign script), and its hand-kept reserved-name inventory never
  gained event_epoch or event_epoch_gen — so a namespace equal to either would
  have nested one installation inside another's keyspace unrefused.
- A test comment referenced idIncarnationShift, which moved to
  internal/idspace.Shift when the package was extracted.
- Two tests called themselves process-restart tests while constructing
  successive buses in one process. They test bus incarnations; the comment now
  says so and says why that is the equivalent thing.

And one reasoning error rather than a stale fact: the counter-backwards branch
justified raising the floor by asserting the arriving ID is necessarily in the
SAME numeric space. It is not — a phase-1 counter reset publishes low IDs with
no epoch to explain them, which is a NEW space we cannot see. The behaviour is
unchanged and still correct (the lead's day-52 ruling: raise unconditionally,
prefer a loud bounded resync loop to a silent skip), but it now says what it
actually knows, which is nothing, and names the cost on a real phase-1 reset.

Claude-Session: https://claude.ai/code/session_01JVDBKbgn3Xt7ndW1YoYd8X
2026-08-22 19:28:17 +00:00
xarmian d393126d80 test(events): close the gaps a tests-as-production-code pass found (BUG-2736)
Codex round 4, on the tests themselves. Eight findings, all real; three of
them were behaviours in this diff with no test at all.

NO TEST AT ALL:

- The real receive path. Every reconciliation test drove fanOutFromRedis
  directly and the publish tests read the wire with a raw subscriber, so a
  regression that decoded the epoch correctly and then handed 0 to the fan-out
  would have passed all of them -- reconciliation silently never running in
  production. Now driven through Subscribe/Publish and back through Redis,
  with the mutation checked.
- The atomic script's ordering claim. Every phase-2 test published once or ran
  the script sequentially, so a two-call INCR-then-PUBLISH implementation
  passed them all -- and that ordering is load-bearing, because the receive
  path reads a descending id as a counter reset. 300 concurrent publishes now
  assert arrival order equals id order; verified to FAIL 5 of 5 against a
  two-call implementation and pass 3 of 3 against the script, so the
  instrument discriminates rather than merely being green.
- The TOML tag. The env-var test proved PAD_EVENTS_PUBLISH_EPOCH reaches the
  field and said nothing about the toml:"events_publish_epoch" tag -- the exact
  form the rollback procedure warns about, since a file value outlives an unset
  env var.

PASSING FOR THE WRONG REASON:

- The production config wiring was still unexecuted: passing an empty
  config.Config at both RunE call sites compiled and passed everything. The
  source-text guard that already counts those call sites now also requires
  them to pass the loaded config.
- The phase-2 wire assertions accepted a well-formed payload with an empty
  event body. They now assert the body survives.
- The Redis metrics subtest had no served-resume control, so a bus that
  refused every resume would have passed. It now round-trips a publish through
  Redis first.
- TestResumeGapIsReportedForBothWaysOfNotServing never proved ws-warm HAD a
  buffer, so its second half could silently duplicate its first.
- internal/server's cold-resume tests still sent a literal 4200, which the
  incarnation guard now answers before the handler's no-buffer path is
  reached. My own round-1 sweep of this class stopped at four packages and
  never looked at internal/server: reviewer-named instances are a sample
  (team CONVE-18), and so, evidently, are self-named ones.

Claude-Session: https://claude.ai/code/session_01JVDBKbgn3Xt7ndW1YoYd8X
2026-08-22 19:15:12 +00:00
xarmian a9544a57ba test(events): repair the resume tests the base guard made vacuous (BUG-2736)
Codex round 1 named three sites; the class was six, across four packages.

Every MemoryBus test that spelled out a cursor as a small literal now passes
through the incarnation guard before reaching the branch it is named for. The
worst were the two that exist precisely to distinguish branches: the
both-ways-of-not-serving observer test would have gone green with BOTH of its
branches deleted, and the watch bus's eviction test would have gone green with
eviction deleted.

Cursors are now base-relative or read back from what the bus issued. Where the
test has only the EventBus interface and no access to the base, the cursor is
derived from a published event's id instead.

Mutation-checked in the direction that matters: with the no-buffer branch, the
coverage check, and the eviction check each made inert in turn, the tests named
for them fail. Before this commit they did not.

Claude-Session: https://claude.ai/code/session_01JVDBKbgn3Xt7ndW1YoYd8X
2026-08-22 18:41:10 +00:00
xarmian 4a6a748c85 feat(events): identify the shared Redis ID space, behind a two-phase flip (BUG-2736)
The activity event counter lives in Redis and is shared by every instance, so
no instance can compute an identity for it the way MemoryBus computes its own
incarnation base. If that counter is ever reset -- evicted under maxmemory,
deleted by hand, a fresh Redis after a restore -- IDs start again from 1, and
a replica buffering the old sequence cannot tell the new 101 from the old 101.
It merges two ID spaces into one replay buffer and answers a resume across the
boundary as though nothing was missed.

Numeric detection alone cannot see it. By the time the new sequence passes the
replica's high-water mark it looks like ordinary progress -- which is the case
the epoch exists for, and the high-water check is what catches the OTHER case
(a publisher that never learned the epoch), so both are kept.

So the identity travels WITH each message, as an opaque token in a
"<epoch>|<id>|<json>" prefix. A prefix rather than an envelope field: an older
instance would unmarshal an envelope object SILENTLY -- no matching keys, no
error, a zero-valued Event delivered to its clients -- and fails loudly on the
prefix instead.

TWO PHASES, because the failure is asymmetric. Every instance ACCEPTS both
wire forms from this release; only emission is gated, on
PAD_EVENTS_PUBLISH_EPOCH. Phase 1 rolls the binary everywhere publishing the
historical bare JSON; phase 2 sets the flag and rolls again. Flipping before
every instance is upgraded is the one direction that LOSES events rather than
resyncing: a pre-phase-1 binary cannot parse the prefix at all. Rollback is
symmetric and safe. docs/deployment.md carries the procedure both ways, what
the reset counters should read during each roll, and what remains unfixed.

Phase 2 also moves ID assignment into one atomic script. The two-call
INCR-then-PUBLISH lets two instances interleave, so a receiving instance can
append 6 before 5 -- a window older than this change, and already wrong, but
load-bearing here because counter-backwards detection reads a descending ID as
a reset. The script carries a dedupe token for the same reason
internal/watchevents' does: go-redis retries a command whose REPLY was lost, so
a publish can happen AND return an error, and the retry would deliver a second
copy that looks perfectly valid.

THE COUNTER-BACKWARDS FLOOR STAYS, and the earlier hope that this unit would
delete it was wrong. Its trigger is mixed-VERSION ordering -- an older binary
assigning and publishing in two calls -- not mixed-FORMAT payloads, so
publish-old-until-flip removes the format window only. It lives for as long as
a deployment can run two publisher versions at once, which is every rolling
upgrade, and the code now says so where it fires.

THE ASYMMETRY WITH MemoryBus IS DECLARED IN BOTH BUSES, in both packages: an
opaque epoch where the counter is shared, a numeric base where one process
owns it. They are not two spellings of one idea and must not be symmetrized.
A numeric base for Redis would close more -- it would refuse cross-incarnation
cursors, which the epoch cannot -- and is deferred rather than rejected: at the
flip, IDs would jump to ~1.8e18 in one step and every un-flipped publisher's
message would read as a massive backwards jump, dropping every buffer across
the whole roll. It is a candidate follow-on once the flip has soaked.

What this does NOT fix is stated in the code and the docs rather than implied:
the client cursor is still a bare integer with no epoch, so an old and a new ID
of the same value remain indistinguishable TO A RESUME even though the buffers
can no longer mix them.

The flip is read inside newObservedEventBus, which now takes the whole Config.
As a hand-picked argument at the two RunE call sites it was untested wiring:
replacing it with `false` compiled, passed the entire tree, and left the
deployment silently on phase 1 -- indistinguishable from a correct phase-1
deployment, since phase 1 is the default. Mutation-checked in both directions,
because a helper that ignores its config and hardcodes either value would pass
a one-directional test.

Also: the epoch and dedupe keys join the namespace assertions (an epoch shared
between two installations is a cross-feed with teeth -- each would read the
other's ID-space changes as its own), and this package's four-key EVAL is now
recorded on BUG-2724's cluster deferral, which had one call site and now has
two.

Claude-Session: https://claude.ai/code/session_01JVDBKbgn3Xt7ndW1YoYd8X
2026-08-22 18:31:02 +00:00
xarmian c017ad359d fix(events): give each in-memory bus incarnation its own ID space (BUG-2736)
Both in-process buses assigned Last-Event-ID values from a counter that
restarted at 1 on every process start. A client holding cursor 2 from a
previous incarnation could reconnect to a restarted server, pass every
coverage check BUG-2731 added, and be replayed the NEW space's 3, 4, 5 as
though they followed the OLD space's 2 -- silently missing everything the
dead space held above 2.

Nothing local could tell the two 2s apart. The cursor carries no epoch, and
in internal/events per-workspace IDs are non-consecutive by construction, so
"did we issue this ID?" was numerically undecidable. The four adjacent levers
were checked rather than assumed: comparing in memory has nothing to compare
against; persisting the counter makes single-process Pad carry durable
event-bus state and still resets on data loss; refusing cursors we did not
issue is the undecidable one; and a nonce on a second channel is unavailable
because EventSource echoes Last-Event-ID and nothing else, and cannot rewrite
its URL on an automatic reconnect.

So the ID space's identity goes in the ID's VALUE while its FORMAT is
unchanged: still a bare int64, still ParseInt on the way back. internal/idspace
mints a base of processStartUnixMilli<<20 and each bus counts up from it. Two
incarnations can only collide if the earlier process published more than 2^20
events per millisecond of its own lifetime -- a deterministic bound, not the
probabilistic one BUG-2736's body rules out. A CAS makes bases strictly
increasing within a process too, which the clock alone does not do for two
buses constructed in the same millisecond.

A backwards clock step degrades in the SAFE direction: a lower base puts old
cursors ABOVE the new buffer's newest ID, so they are refused rather than
answered wrongly. The overflow bound is computed, not estimated: the last
start instant that fits is 2248-09-26T15:10:22Z.

Each bus then answers the resume question exactly instead of inferring it: a
non-zero cursor at or below this incarnation's base was issued by a dead
space. That is strictly stronger than the coverage check alone, which serves
the ADJACENT cursor on reasoning that only holds within one ID space.

In internal/watchevents the check lives in one helper both entry points call.
Written inline in EventsSince it was absent from SubscribeAndReplaySince --
the path the SSE handler actually uses -- so the component was fixed and its
wiring was not (team CONVE-19). A test now drives both.

web's ItemEvent no longer declares `id?: number`. Nothing read it, which is
the only reason it was harmless; a base of ~1.8e18 is past JavaScript's
MAX_SAFE_INTEGER, so the first reader would have silently got a rounded
number. Defused while still unread.

Tests that spelled out IDs now read back what the bus assigned -- a literal 1
is a cursor from a dead space, which turned two negative controls into their
own opposite. The two watchevents guards (cold buffer, dead incarnation) are
tested separately, because a single test covering both would keep passing
with either deleted.

The Redis half is not here. Its counter is shared across processes, so
identifying its ID space needs an epoch travelling with each message; that is
the next commit on this branch.

Claude-Session: https://claude.ai/code/session_01JVDBKbgn3Xt7ndW1YoYd8X
2026-08-22 18:14:45 +00:00
xarmian 9f88e94832 fix(events): a resume must not be answered from coverage we never had (BUG-2731)
internal/events answered a Last-Event-ID resume with an empty-but-non-nil
slice whenever the workspace's replay buffer could not speak to the span
being asked about. The SSE handler reads that as "caught up", so the client
sat on a live stream believing it was current while everything between its
cursor and now was silently gone.

COVERAGE. replayBuffer gains knownFrom: the lowest event ID from which this
instance's coverage of a workspace can be vouched for. A resume from below
it answers nil, which the handler already turns into sync_required. Covers
a buffer that does not exist (cold start, restart, scale-up, or simply the
first connection to a workspace on this instance), a buffer that exists but
starts above the cursor — NOT full and NOT empty, reachable on any
multi-instance deployment with no eviction and no restart — and a non-zero
cursor from a previous incarnation of a single process.

knownFrom here means RECEIVING-continuity, never ID-contiguity, and the
defining comment says so with the measurement attached.
internal/watchevents has a field of the same name that ALSO detects holes
by noticing a non-consecutive ID; porting that would have been a serious
regression, because this bus has a global counter and per-workspace
buffers, so a workspace's buffer holds non-consecutive IDs by construction
(four publishes alternating across two workspaces measure as W=[1 4],
X=[2 3]). An ID-contiguity check would fire on nearly every append and turn
every resume into sync_required — the false-positive inversion of this bug.

LIFECYCLE. Coverage now ends where it really ends:

  - a stopped workspace subscription drops its replay buffer. Keeping it
    "in case they come back" looks like a free win and is the bug: events
    published elsewhere never enter it while it goes on looking complete.
  - subscriptions are generation-numbered, so a straggler from an ended
    subscription cannot re-create a buffer and vouch for coverage that
    ended with it — including the case where the workspace has already been
    resubscribed under the stale goroutine.
  - a pub/sub reconnect ends that workspace's coverage. PubSub.Channel
    resubscribes transparently, so a Redis failover left a hole the buffer
    had no idea about; the loop reads pubsub.Receive instead. It must
    RECOVER rather than exit — returning on a transient error would leave
    an instance publishing fine and receiving nothing — and it drops ONE
    workspace's buffer, since a dropped subscription says nothing about any
    other channel.

Subscribers are indexed by workspace because the replay buffers moved under
the same mutex (necessary for the straggler race): scanning every local
subscriber under that lock would make one hot workspace the serialization
point for every other workspace's fan-out and every resume.

Also removes Publish's local-counter fallback on a failed INCR, which
minted an ID from a process-local space and published it — every receiving
instance reads that as the counter having been reset. It bought nothing:
this bus has no local fan-out path, so an event that does not reach Redis
reaches no subscriber here either.

SIBLING. internal/watchevents had the identical cold-resume defect on its
MemoryBus — its RedisBus guards it, MemoryBus reached the buffer directly —
so a single-process instance answered a post-restart resume as caught up.
Found by a cross-artifact review pass; the guard goes in `since` so both
implementations inherit it, and is tested through SubscribeAndReplaySince
as well as EventsSince because that is the path the handler uses.

Refs BUG-2731
2026-08-22 14:23:40 +00:00
xarmian bb003dd6bb fix: five claims the final comment-truth round found (BUG-2724, BUG-2726)
The bounded process the lead set: N rounds, an author prune pass, one
final comment-truth round. This is that round's output, and the loop
stops here.

Two were mechanisms I had wrong, and both are the kind a reader would
reuse without re-deriving:

- "Different Redis DB numbers do not help" was half true. Ordinary keys
  ARE DB-scoped, so two installations on different DBs keep separate
  presence registries; it is pub/sub that ignores DBs entirely, which is
  why the buses cross-feed regardless. Stating it as "does not help" made
  the namespace look like the only fix for a problem it only half is.
- A namespace cutover's client resync was attributed to the epoch check.
  That check needs an OLD epoch to compare against and a freshly
  namespaced bus has none — the resync comes from the cold replay-buffer
  coverage check instead (knownFrom is zero, so every resume falls below
  it). Same honest outcome, different mechanism, and the mechanism is
  what someone reasoning about a cutover would use.

Three were stale or over-general after earlier changes: the admission
comment still said the global limit is passed to the bus as 0 (that
parameter is gone), `pad watch --help` and the plugin monitor description
lumped a missing .pad.toml's hourly retry in with the 5s-to-5min backoff,
and CLAUDE.md said clients must back off without the browser exception
docs/deployment.md spells out.

Claude-Session: https://claude.ai/code/session_01JVDBKbgn3Xt7ndW1YoYd8X
2026-08-22 05:21:01 +00:00
xarmian 461c5a3e3d refactor: prune the claim surface, and turn one prose claim into a test
The review could not converge on this diff's comments because each round
of corrections re-expanded the surface it was reviewing — rounds 16 and
17 found errors inside 15 and 16's fixes. That is a production rate being
measured, not a backlog being drained, so the treatment is to write
fewer claims rather than review the same ones again.

PRUNED, ~135 comment lines: process narration. "An earlier version said
X", "found by mutation testing", "codex round N caught this", the
scoreboards. Every one of those is already in a commit message, which is
where the archaeology belongs; in the source they are claims a future
reader has to verify, about a past that no longer exists.

KEPT, because they earn it and a reader would otherwise re-derive them:
metric semantics, reachability boundaries, what a test does and does not
discriminate, why the obvious alternative was rejected, and the hazards
that cannot be enforced in code.

MOVED TO A TEST, per the rule this run earned the hard way: a comment
asserting countable behaviour belongs in the suite. Two test comments in
internal/watchevents relied on "this constructor waits for its SUBSCRIBE
to be confirmed" — prose, and the same assumption applied to the OTHER
bus (which subscribes asynchronously) is what made a namespace test
flake. It is now asserted with no polling and no sleep, and the mutation
that removes the wait fails it.

That rule generalises and is why round 17's find mattered: "counts every
unservable resume" was prose, so its falseness could hide a real metric
gap. Prose is for claims that cannot be asserted.

Claude-Session: https://claude.ai/code/session_01JVDBKbgn3Xt7ndW1YoYd8X
2026-08-22 05:09:29 +00:00
xarmian 35e564298b fix: seven more prose claims, one real metric gap, and a flaky test of my own (codex round 17)
The prose angle again, and it is still finding things — which is itself
the finding: this diff's comment density is generating wrong beliefs
faster than the review is removing them, in the one dimension where the
defect is a reader's understanding rather than the program's behaviour.
Everything below was a claim I wrote.

ONE WAS A REAL GAP, not just wording. pad_watchevents_resume_gaps_total
was documented as counting every unservable resume, and counted only the
half decided by the shared counter. The LOCAL half — a cursor below what
this instance can vouch for, from a hole or a cold start — returns nil
from replaySince, becomes sync_required for the client, and reported
nothing. Now counted, on the deferred path so it fires with the lock
released.

Its test needed a second pass to be an instrument: the first version
arranged a hole and asserted the counter moved, but the shared counter
disagreed too, so resumeOutrunsLocalView reported and the mutation
survived. It now sets the counter to AGREE with what the instance has
seen, which is the only arrangement that isolates the local path.

The prose corrections, swept by grep rather than by instance this time:

- MemoryBus's comment said a single-process deployment never wires an
  observer. cmd_server wires one, deliberately — that is what makes the
  drop counter meaningful there, which is a claim I had just added
  elsewhere.
- "Every write path works with Redis down" was too strong in three
  places. Push answers 503 for an unresolvable targeted push and 502
  push_unconfirmed on publish failure — the paths whose job IS
  cross-instance delivery.
- Presence-failure consequences were stated as certainties in four more
  places after round 16 fixed one. A failure means an error was
  REPORTED; Redis can fail a pipeline after applying it.
- The deployment metrics table still described pad_eventbus_publish_total
  as "Events published" after the Help string had been corrected to
  attempts.
- The reserved-namespace rationale called prefix nesting a "collision".
  It is nesting; an exact collision would need the namespace to match a
  workspace UUID. Refused anyway, and now for the reason that is true.
- A presence cutover was described as stranding one renewal interval of
  stale entries. It is the full 90s TTL — three intervals.

AND A FLAKE OF MY OWN, caught by the full suite rather than by the
targeted runs: the activity-bus namespace test asserted subscription
state immediately, but that bus subscribes ASYNCHRONOUSLY (the watch bus
waits for confirmation; the two differ). It now polls, and the asymmetry
is named in both tests so the next reader does not assume symmetry the
way I did.

Claude-Session: https://claude.ai/code/session_01JVDBKbgn3Xt7ndW1YoYd8X
2026-08-22 05:00:18 +00:00
xarmian 2fa1316853 fix: three claims round 15's corrections got wrong or missed (codex round 16)
Reviewing the corrections found three more, which is the honest shape of
this: the prose angle keeps paying because the errors are in prose.

- My round-15 correction said the Redis counters "stay at zero" on a
  single-process binary. That is wrong for one of them:
  pad_watchevents_notifications_dropped_total moves there, because
  MemoryBus has the same slow-subscriber drop and is wired to the same
  observer. So the comment was wrong before AND after, in opposite
  directions. It now says which counters are Redis-only by construction
  (everything sequence-related — MemoryBus assigns contiguous ids and has
  no subscription to lose) and which are not, and a test pins both halves.
- "The three keyspaces cannot drift" survived in cmd_server.go. Round 15
  fixed the copy in redisns.go and not this one — a two-member class,
  fixed one member, which is team CONVE-18 for the second time in this
  branch.
- The presence-failure consequences were stated as certainties. Redis can
  fail a pipeline or a script AFTER it applied, so a failure means the
  operation reported an error, not that it did not happen. Now phrased as
  what a failure risks.

Codex's reserved-namespace audit came back complete: the set covers every
current suffix root (watchevents:pub: is covered by watchevents), every
configuration path goes through Parse, all three production constructors
receive the parsed value, and the exact-match controls do not reject
names that merely contain a reserved word.

Claude-Session: https://claude.ai/code/session_01JVDBKbgn3Xt7ndW1YoYd8X
2026-08-22 04:43:14 +00:00
xarmian 9d54f24626 fix(server,cli): the half of round 12's fix I missed (BUG-2726)
Codex round 13, unanchored, found that my previous commit fixed one of
the two refusal paths on /api/v1/events. The admission check moved above
the SSE headers; the PER-WORKSPACE check stayed below them, so half the
429s on that endpoint still carried the JSON error envelope under
Content-Type: text/event-stream — the exact defect the commit said it
fixed.

Team CONVE-18 in its own shape: the reviewer named one instance, I fixed
that instance, and the class had two members. The enumeration I owed was
"how many ways can this handler refuse", and it takes ten seconds to
read. Every refusal is now above the header block, with a line saying
nothing below it refuses.

The contract test made the same omission and is the reason this reached
another round: it drove the admission bound on both endpoints and never
the per-workspace one, so it agreed with a handler that was half fixed.
It now enumerates all three refusal paths, and the mutation that
reintroduces the defect fails it by name.

Also from round 13: `pad project watch`'s 429 message named the two knobs
that cover both streams and omitted PAD_SSE_MAX_PER_WORKSPACE, which is
the one most likely to be the cause on a busy workspace — true as far as
it went, and pointing the reader away from the answer.

Claude-Session: https://claude.ai/code/session_01JVDBKbgn3Xt7ndW1YoYd8X
2026-08-22 04:18:18 +00:00
xarmian 3e3170e915 fix(server,cli,docs): the consumer contract, per codex round 12 (BUG-2726)
An angle no earlier round took: what does a CLIENT see. Two of the five
findings were about consumers I had never opened.

- `pad project watch` returned "event stream returned 429: {json}" and
  exited, which sends the reader looking for a bug rather than at a
  limit. It now says what happened and which knobs govern it, and names
  the fact that those knobs cover this stream and the agent watch stream
  together. It still exits rather than backing off — it is interactive,
  and a human can decide — unlike the unattended monitor, which already
  folds 429 into its ladder.

- Both endpoints now answer a refusal through one helper: same status,
  same code, same message, plus `Retry-After`. `/api/v1/events` was
  setting `Content-Type: text/event-stream` BEFORE the admission check,
  so its 429 carried the JSON error envelope under an SSE content type —
  a different contract from its sibling's for the same refusal. Admission
  moved above the headers, which is where it belonged anyway.

- The anonymous-caller rule was documented as if it applied to both
  endpoints. It applies to `/api/v1/events` only; the watch stream
  requires a resolved user and answers 401 without one.

- docs/architecture.md described one SSE endpoint and one bus. It now has
  the table: two streams, two buses, different scopes and consumers, one
  shared connection budget, one Redis namespace.

FILED, not fixed: the web UI's `EventSource` cannot see a 429 or a
`Retry-After` — the spec exposes neither to the page — so a refused
browser tab reconnects at a constant rate while the CLI backs off. That
asymmetry means reaching the limit sheds load from the population that
respects it and not from the one that grows fastest under it. No
server-side change closes it; the fix is a client-side reconnect wrapper.
BUG-2733, and docs/deployment.md warns operators to size the limit with
it in mind.

Claude-Session: https://claude.ai/code/session_01JVDBKbgn3Xt7ndW1YoYd8X
2026-08-22 04:10:06 +00:00