mirror of
https://github.com/PerpetualSoftware/pad.git
synced 2026-09-11 13:28:57 +00:00
83f00a2c66e604359134aa26520d979f3682c702
892 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
110045578c |
fix(build): make install proves what it installed and what it restarted (BUG-2897, TASK-2787) (#1272)
`make install` made three claims it did not check.
1. It brought the server back BY SIDE EFFECT -- `pad auth whoami` triggers
an auto-start, which does not know the killed process's argv. A server
running `--host 0.0.0.0` came back bound to the default host alone:
curl 127.0.0.1:7777 -> 000 while the LAN address -> 200, with the
process count and the version both reading correct (BUG-2897).
2. `cp` copied whatever was at the repo path, not what the invocation
built. Two sessions sharing the checkout interleave and the loser's
build is installed by the winner, every exit code green (TASK-2787).
3. "Server restarted." was printed after a command ending in `|| true`,
with no probe of any kind. Not "the wrong address went unverified" --
nothing was verified. Found reading the recipe; neither filing names it,
and it is what made the other two invisible.
The logic moves to scripts/install-refresh.sh for one reason above
readability: a script can be TESTED. internal/buildtools drives it against
a compiled stub `pad`, including the branches where a check must FAIL.
Recipe-inline logic is only exercisable by running `make install`, which
stops the developer's server -- a test nobody runs twice, which is how this
target accumulated three unverified claims.
The script checks OUTCOMES rather than steps: what got installed, and what
is answering afterwards. Two commit checks, deliberately, answering
different questions -- the ARTIFACT before the kill, so a wrong build costs
an error instead of an outage, and the DESTINATION after the copy, which is
the shared-path race TASK-2787 names. The restart uses the argv read from
/proc before the kill, and nothing is printed about a restart until the
server answers on BOTH 127.0.0.1 and the configured host (`--host 0.0.0.0`
resolving to loopback plus the primary LAN address, since 0.0.0.0 is a bind
spec, not something to curl).
Three defects were found in the fix itself, by its own tests and by the
first real run:
- The restart redirected to $HOME/.pad/server.log with nothing creating
that directory. On a fresh HOME the redirect fails, the server never
starts, and the probe reports "did not answer" -- true, and three steps
downstream. Invisible to every hand-run this script could have had,
because a developer's box has ~/.pad by luck of history.
- The post-copy check SURVIVED its mutant: both fixtures reported the same
version from source and destination, so the guard and its absence were
indistinguishable. Fixed by making the stub's version depend on its path.
- Comparing commits by equality rejected a healthy build. `git rev-parse
--short` returns the shortest UNAMBIGUOUS prefix, so its width grows with
the object database: the binary embedded `a3a1d58` and the Makefile
produced `a3a1d586` minutes later. Now a prefix comparison in either
direction, with a negative leg pinning that a genuinely different commit
of the same width is still refused.
Five mutants, each verified to compile, each detected by its own leg.
Verified end to end on the real box: captured `--host 0.0.0.0`, installed
|
||
|
|
a3a1d5862b | fix(store): preserve valid item references across workspace import (#1271) | ||
|
|
bb8ec04ef1 |
fix(server,store): both workspace mint doors enforce their preconditions from one place (BUG-2809) (#1268)
handleCreateWorkspace and handleImportWorkspace mint the same thing
through the same store.CreateWorkspace, and enforced preconditions in two
places. Two had already diverged and been fixed one at a time, each found
by a reviewer rather than by the door that lacked it: the OAuth consent
grant (IDEA-2756) and the user-scoped plan limit (BUG-2793). A third was
live.
The shared place is internal/server/workspace_mint.go, split by WHEN a
precondition can run, and the split is load-bearing rather than tidy:
beginWorkspaceMint — everything that does not need the body (consent,
plan limit, and the owner/source attributions). Runs before the body
read, so a refused caller never uploads a bundle and a refusal cannot
be probed by body shape; on the import route it sits above the
Content-Type dispatch, so one line covers both body shapes.
validateWorkspaceMintPayload — the payload-shaped rules. Returns an
error rather than writing one, because the JSON doors answer 400
bad_request and the bundle door answers 400 bad_bundle through
importStatusError. The rule is shared; the envelope stays each door's.
Callers: handleCreateWorkspace, handleImportWorkspace, and importBundle.
The mint context reaches the bundle path as an ARGUMENT rather than on the
Server, because it is per-request state and the two things it carries are
exactly what two concurrent requests would differ on.
THE LIVE DEFECT. Import accepted an empty workspace name. Measured before
the fix: it created a workspace with name="" and slug="", and a second
such import landed on slug "-2" -- the first had taken the empty slug,
globally, and a slug is a routing key. Both import doors now refuse it,
checking the EFFECTIVE name (the ?name= override when given, the bundle's
own otherwise) because that is what becomes the slug. A control leg covers
the override, or the rule would be indistinguishable from "reject any
bundle whose payload name is empty" and would break rename-on-import.
SETTINGS: the item's premise was wrong and this corrects it rather than
fixing it. Malformed settings never reached the store unnormalized --
createWorkspaceQ calls NormalizeWorkspaceSettings itself and refuses. What
diverged was the STATUS: create answers 400, import answered 500
import_failed because handleImportWorkspace maps every store error that
way. Validating in the shared payload step makes both 400. Context stays
create-only: an export carries none, so applying it on import would be
inventing input.
SOURCE: imported workspaces got no attribution at all (BUG-1557).
store.ImportWorkspace now takes a source parameter, derived by the caller
from the request's auth shape exactly as create derives it -- a parameter
rather than an export field, because a bundle says what the workspace WAS
and where this copy is minted from is a fact about this request. The
operator path (pad db migrate-to-pg) passes "": it is a copy, not a
creation surface, and inventing "cli" would relabel every migrated
workspace's origin.
Userless callers (the inventory's fourth item) are deliberately unchanged.
beginWorkspaceMint preserves the userID != "" guard exactly as both doors
had it rather than changing behaviour under cover of a refactor; the
measurement and the ruling are on BUG-2914.
Five mutants, each verified to COMPILE first and each detected by its own
leg: either import door skipping the payload check, the create door
skipping it, checking the payload name instead of the effective name, and
passing "" for source. Two of them initially did not compile, and go test
answers a build failure with FAIL <pkg> [build failed], which in a
filtered run reads exactly like detection -- a false DETECTED, the mirror
of the false SURVIVED. Re-run with the orphaned variable kept alive.
Claude-Session: https://claude.ai/code/session_01HeChkgZVYb3NTgTcckF5KR
|
||
|
|
cdc5b301e2 |
fix(server): minting or rotating an API token requires an interactive session (BUG-2890) (#1267)
A PAT that could reach POST /auth/tokens minted further tokens with
independent names and expiries. Those survive the revocation of the token
that created them, and nothing in the token list records which token
minted which -- so revoking a leaked credential did not end the access it
had been used to establish.
Ruled on the item's trail day 57: create and rotate require session auth
and refuse a PAT-authenticated call with 403 session_required; list and
revoke stay PAT-reachable.
Population is THREE doors, not the two the filing named. Enumerating every
route that reaches store.CreateAPIToken / RotateAPIToken turns up the
workspace-scoped mint at POST /workspaces/{ws}/tokens, which had only
requireMinRole("owner") in front of it -- and a user-owned PAT held by an
owner satisfies that. Measured, not read: with the fixture's owner also
written into workspace_members, that door returned 201 with a live token
on the unfixed tree. All three doors are gated; list and revoke on both
the user and workspace routes are deliberately untouched, because neither
extends access and revocation is the compromised-credential response.
The gate is on the CREDENTIAL, not on the door: isAPITokenAuth is false
for a session cookie AND for a padsess_ CLI bearer, a distinction that
predates this fix (see ctxValidatedSessionBearer's note in
middleware_auth.go) because a CLI session IS an interactive session. At
the workspace door the check runs before requireMinRole, so a PAT-borne
caller cannot learn its own membership status from the difference between
the two 403s.
Five mutants run, each detected by its own leg: dropping the gate at each
of the three doors fails that door's test; gating on Authorization != ""
instead of the credential kind fails the CLI-session leg AND NOTHING ELSE;
over-applying the gate to list and revoke fails both PAT-still-works
controls. The fourth is why the CLI leg exists -- a header-shaped gate
passes every PAT assertion and breaks every logged-in CLI.
Also documents the refusal on the CLAUDE.md route line: the routes were
listed without saying anything about credential kind, so an agent holding
a PAT would have met an unexplained 403.
Claude-Session: https://claude.ai/code/session_01HeChkgZVYb3NTgTcckF5KR
|
||
|
|
e32d4f468e |
fix(store): a relation value pointing at an orphaned item is carried, not rewritten into a minted id (BUG-2895) (#1265)
ImportWorkspace's second pass handed remapFieldIDs the unfiltered itemMap. That map holds an entry for EVERY item in the bundle including orphans -- items whose collection the bundle does not carry -- because the entry is written before the orphan skip and parent resolution inside the same loop reads it for items the loop has not reached yet. So a relation field pointing at an orphan was rewritten to the id that orphan WOULD have received: an id that names no row in the destination and has never identified anything in any workspace. That is worse than dangling. TASK-2878's carry rule imports an unresolvable relation value verbatim precisely so nothing is invented -- the source id is evidence a human or a repair tool can act on -- and ImportWorkspace runs no MigrateRelationReferents pass afterwards to re-examine what this wrote. The fix is the same correction BUG-2884 made for parent_id: filter the map to items that actually landed. It arrives separately because every site that unit fixed writes a FOREIGN KEY, so the database objected and the sweep found them. This one writes into a JSON blob, so nothing objected. Population, re-verified at this tip rather than carried from the recon: itemMap has eight consumers inside ImportWorkspace and seven already check insertedItems; this call was the eighth. collMap needs no equivalent -- the collections loop has no `continue`, so every collection either INSERTs or the import returns an error, and no collMap entry can name a row that does not exist. Repo-wide, remapFieldIDs has one caller. Search boundary: internal/store. Also drops remapFieldIDs' collMap parameter, which was never read. A relation's target collection travels in the SCHEMA as a slug, not as an id in the fields blob, so an unread parameter there reads as a claim that collection ids are remapped -- they are not. Test: TestImportWorkspace_DoesNotMintIDsForOrphanRelationReferents. The bundle is hand-built because ExportWorkspace cannot produce an orphan, which is the population the carry rule exists for. Three mutants run, all detected: restoring the defect fails the orphan leg with a fresh UUID; deleting the second-pass remap fails the CONTROL leg; filtering on the wrong key space (itemMap is old->new, insertedItems is keyed by NEW) fails the control too. Under the first mutant the pre-existing carry test stays GREEN -- its unresolvable ids are absent from the bundle, so they are absent from itemMap, which is why no existing test could see this. Claude-Session: https://claude.ai/code/session_01HeChkgZVYb3NTgTcckF5KR |
||
|
|
3c93472613 |
perf(store): run the store suite in parallel (TASK-2900 deliverable 3) (#1263)
The SQLite leg was CPU-bound and serial: 154.9s of CPU inside 185.6s of wall on an eight-core box, which is 0.84 cores. Unlike the Postgres leg there was no setup cost to reclaim — the unattributed gap between package wall-clock and the sum of its tests measured MINUS 1.4s, so per-binary overhead was already nil — and no fat test to find beyond one 48s outlier. The only lever left was the seven idle cores. t.Parallel() on 844 of 906 tests. Every test in this package builds its own isolated store (a SQLite file copy, or since |
||
|
|
7659ad3cd3 |
feat(server,web,cli): say when a relation's copy target is unusable, instead of offering a picker that cannot answer (IDEA-2899) (#1262)
* feat(server): the copy preflight says when a relation's target is not usable (IDEA-2899) TASK-2869 made a `needs_value` relation row collectable as soon as it names a target collection. Naming one is not having one: the slug can name a collection that has been DELETED, or one this caller cannot READ. The dialog then mounts a picker that can return nothing and, because the row is not blocked, Confirm stays disabled carrying only the generic required-field message — the user is told a value is missing and never told that no value is reachable. `collection_unavailable` on the needs_value row is the server saying so. THE CLIENT CANNOT COMPUTE THIS, which is why it belongs here. The dialog's destination collection list is filtered through `canEditCollection`, because it drives the copy-INTO picker; a relation TARGET needs only READ access, so a perfectly usable target routinely does not appear in that list. Testing against it would refuse rows the user could have filled in — over-blocking, which is the worse failure and invisible to whoever hits it. `visibleCollectionIDs` is the read-scoped view, and its NAV-LENIENT shape is right here rather than merely tolerable: it includes a collection reachable only through an item-level grant, and the question is "could a picker here return anything at all". One granted item is a picker with one row. DELETED and UNREADABLE are deliberately not distinguished. Same consequence, no client branch would differ — and separating them would tell a caller who cannot read a collection that it nonetheless exists. `omitempty` on a BOOL drops `false`, so the field is phrased NEGATIVELY. Present-and-true means the server checked and the target is unusable; ABSENT means available, or a server that does not report. A client must block only on an explicit true, so absence stays "no information" rather than becoming a value — the rule `access_epoch` follows on the item doors, and the one whose violation cost two review rounds on IDEA-2898 this morning. Costs nothing on the common path: a destination schema declaring no relation field runs no query at all. Claude-Session: https://claude.ai/code/session_01Xk9M5UVPdc84xL5E1mZkm8 * test(server): pin the type gate on collection_unavailable (IDEA-2899) Found by a surviving mutant rather than by inspection: dropping the `def.Type == "relation"` gate left every other test in the file green. Nothing stops a schema declaring `collection` on a field of another type — the validator does not police keys it has no use for — and such a field would then pick up a flag whose meaning is defined only for relations. The dialog would block a perfectly collectable `select` because some relation elsewhere in the same schema points at a collection that happens to be gone. The fixture is the discriminating one: ONE deleted collection, TWO required rows that name it, and only one of them means anything by it. Six mutants on this half, all killed: flag never set, flag always set, deleted target not flagged, unreadable target not flagged, type gate dropped, and the nil-visible-set case (an admin's "no filtering" read as "nothing visible", which would flag every target for the callers who can see everything). Claude-Session: https://claude.ai/code/session_01Xk9M5UVPdc84xL5E1mZkm8 * feat(web): block a relation whose target is unavailable, and stop advising a command that cannot work (IDEA-2899) The client half. `isCollectable` now refuses a relation row the server has flagged, so the row lands in `blockedFields`, Confirm is disabled with a reason, and no picker mounts that could only come back empty. `collection_unavailable !== true` is STRICT on purpose. The field is absent when the target is fine and absent from a server that predates it, so absence must read as "no information". (Over the domain the type admits — `boolean | undefined` — the truthiness spelling is EQUIVALENT and a mutant swapping it in survives; that is recorded in the source rather than papered over with an off-contract fixture. The strict form is kept because it states the contract where the next edit will read it, and the inverse spelling would block every row against an older server.) THE PART THAT IS NOT WIRING: the existing blocked-field notice said the field "is a required <type> field. This dialog can't collect a value for that type safely" and then printed `pad item copy … --field key=value`. Both halves are FALSE here. The type is perfectly collectable; the TARGET is gone. And the CLI runs as the same user against the same referent validation, so the command it prints is refused for exactly the reason the user is already stuck — advice that sends someone to do work that cannot succeed is worse than no advice. So the message branches on `uncollectableReason`, names the collection and the destination workspace, and the CLI line is now gated on `cliFillableField` — the first blocked row the CLI can ACTUALLY fill. `blockedFields[0]` was correct while every blocked row was type-shaped; with an unavailable relation sorted first it named the one field `--field` cannot set either. Eleven unit tests on `copyNeedsValue`, plus a source pin on the dialog whose own measured limit is in its docblock. Client mutants: 7 real, 6 killed, 1 recorded as equivalent with the domain argument that makes it equivalent. Claude-Session: https://claude.ai/code/session_01Xk9M5UVPdc84xL5E1mZkm8 * feat(cli): the copy preview marks an unavailable relation target and stops suggesting it (IDEA-2899) Caught by `TestItemCopyMirrorsMatchServerShapes`, not by me. The CLI keeps a mirror of the preflight response, and adding a field server-side without mirroring it fails that test by design — a mirror that silently lags is a mirror that lies. Working exactly as intended, and the reason this half exists at all. Mirroring the field turned out to be the smaller part. The CLI already prints `target collection: people` for a relation row, and it builds an `Add: --field owner_ref=<value>` suggestion from every unsupplied row. Both are wrong when the target is unavailable: the first sends a user looking for a ref in a collection they cannot read, and the second hands them a command the referent validation refuses for exactly the reason they are already stuck. So the target line is marked NOT AVAILABLE, and the row is excluded from the suggestion with a sentence saying why — modelled on the empty-key branch, which was written for the identical reason (a `--field =<value>` nobody can run) and is three lines away. That the same defect had to be fixed in two places is the shape worth naming: the dialog and the CLI independently built "here is how to supply it" from "here is a field needing a value", and neither had a notion of a field that CANNOT be supplied. The empty-key case was the first instance and was fixed locally; this is the second. Five mutants on this half, all killed: suppression removed, suppression applied to everything, the unavailable label dropped, the explanation dropped, and the mirror field ignored. The available-target control leg is a separate test so the omitempty contract is exercised on this surface too. Claude-Session: https://claude.ai/code/session_01Xk9M5UVPdc84xL5E1mZkm8 * fix(cli): route all three "how to supply it" sites through one predicate (IDEA-2899) Review found the fix applied at one door and not its siblings — my own recurring shape, arriving again. THREE places tell a CLI user how to resolve an unsatisfied field: the detailed `renderItemCopyNeedsValue`, the `--dry-run` summary, and the error the command returns. The first commit fixed the render. The other two went on printing `--field key=value` at someone for whom no value exists — and the ERROR is the line a script or a hurried reader actually sees, so it was the worst of the three to leave. `itemCopyUnfillable` is now the single definition all three consult. Not because three call sites are tidier than one, but because three sites independently answering "how do I supply this" is exactly how they diverged in the first place. The dry-run summary branches three ways rather than two, because the MIXED case is the one a boolean gets wrong: some fields can be supplied and some cannot, and collapsing that either suppresses advice the user needs or offers advice they cannot use. The error hint is suppressed only when NO field can be supplied — with one fillable field left, `--field key=value` is still true. Also pins the BOUNDARY the same review probed: a target collection that is live and readable but EMPTY is deliberately not flagged. The symptom looks identical — an empty picker — but the cases differ where it matters. An unavailable target is unfixable from inside the dialog, so blocking costs the user nothing they had; an empty collection is resolved by creating the item and retrying, and blocking would refuse a copy they were about to complete. It would also cost a live-visible-item count per relation target on a dry run the UI calls on every keystroke. The weaker case — an empty picker that says nothing about WHY — is filed as IDEA-2905 and belongs to the picker. Ten mutants across this round, all killed, including both directions on the error hint and both directions on the dry-run branch. Claude-Session: https://claude.ai/code/session_01Xk9M5UVPdc84xL5E1mZkm8 * fix: unfillable means EITHER reason, and a select never names a relation target (IDEA-2899) Review round 2, two findings, both real and both about a rule stated in one place and enforced in another. **"Unfillable" answered for one of two reasons.** An EMPTY KEY cannot be supplied either — `--field =value` is rejected by this command's own parser, and the detailed render has explained that since Codex round 6. Only that render knew: the --dry-run summary and the returned error went on advising `--field` for those rows, because the predicate I extracted last commit covered the relation reason alone. A predicate named "unfillable" that answers for half its name is a worse trap than no predicate — right at the site that defined it, wrong everywhere it was reused, which is precisely what extracting it was meant to prevent. Two functions now: `itemCopyUnfillable` (either reason — advice), and `itemCopyUnavailableTarget` (the relation half — the render's own sentence, since the two explanations are not interchangeable to a reader). `itemCopyUnavailableTarget` deliberately does NOT also exclude empty keys, though my first version did. A row can carry both faults, and a mutant removing that exclusion survived every test — correctly, because all it changes is printing two sentences that are both TRUE about such a row. The guard was tidiness dressed as a rule; a condition nothing can distinguish is one the next reader has to re-derive. **`Collection` was emitted for non-relation fields**, while its own doc said it is empty for every other type. That was a claim about the schemas people write, not a property of the code: a `select` carrying `"collection": "people"` is storable — field validation has no use for the key and does not police it — and the value was copied straight through, so the CLI printed "target collection: people" beneath a select. A relation fact asserted about a field that has none. `relationTargetSlug` makes the documented contract true at the only place that can make it true; my own type-gate test had created exactly that shape and asserted only the FLAG, not the slug. Three mutants on these fixes, all killed. Claude-Session: https://claude.ai/code/session_01Xk9M5UVPdc84xL5E1mZkm8 * fix(cli): the explanation now names the reason that actually applies (IDEA-2899) Review round 3, and the sharpest miss of this unit — my own, one commit old. Broadening what a predicate ACTS on silently broadened what a sentence SAYS. Once `itemCopyUnfillable` counted empty keys as well as unavailable relation targets, a set of empty-key rows selected the all-unfillable branch and was explained as "the relation target is not available to you" — a false statement about rows that contain no relation at all. Same in the returned error, which is the line a script sees. The tell was there to be read: a sentence that was TRUE while the predicate was narrower is a sentence to re-read the moment it widens. I broadened the predicate deliberately, wrote a commit message about how a half-answering predicate is a trap, and left the sentence describing the half. `itemCopyUnfillableWhy` names the reasons actually present — relation targets, empty keys, or both — and the two one-sentence sites consult it. The detailed render is unchanged: it explains each reason where the row is printed, which is why it uses the narrower count. Four mutants, all killed, including the two that matter: the explanation always saying "relation" (the defect) and never saying it (the same defect pointing the other way). The test carries a mixed-reason leg, because a sentence that picks one of two true reasons is the failure a single-reason fixture cannot see. Also corrected: three comments claiming `itemCopyUnfillable` is relation-only or that the detailed render consults it. Both stopped being true last commit. Claude-Session: https://claude.ai/code/session_01Xk9M5UVPdc84xL5E1mZkm8 * fix: one row can carry both faults, and four docs said this was simpler than it is (IDEA-2899) Review round 4. Four findings, no P1s, and the first is the one worth the round. **A `continue` between the two counts.** `itemCopyUnfillableWhy` counted a row as an unavailable relation target and then skipped the empty-key check, so ONE row carrying both faults reported only the first. My mixed-case test used TWO rows with one fault each — a different input, and the only one it exercised. Two rows with one fault each and one row with two are not the same fixture, and I built the weaker one while writing a commit message about fixtures that cannot discriminate. **The dialog could still print `--field =value`.** `cliFillableField` excluded unavailable relation targets and not empty keys, so a required `json` field the destination reported with no key was type-shaped, blocked, and still offered a command the CLI's own parser rejects. The CLI has refused those since Codex round 6; the web side had never learned it. Same defect, other surface — which is the third time this unit has fixed one door and not its sibling. **Cardinality.** "no --field can supply it" for several fields, and "reported them with an empty key" for one. Both sites now agree with their counts, and the empty-key phrase is neutral on number so it reads correctly after either. **Four documents claimed every needs_value row is resolvable with an override** — the CLI renderer's docblock, the server's `NeedsValue` field, the CLI mirror type, and the dialog's collectability comment. That was true when each was written and this unit falsified all four; a reader following any of them would conclude the CLI had simply forgotten to print a flag. Two mutants on the fixes, both killed: the `continue` restored, and the dialog's empty-key exclusion removed. Claude-Session: https://claude.ai/code/session_01Xk9M5UVPdc84xL5E1mZkm8 * refactor(cli): one tally, because the rounds said the branching was the problem (IDEA-2899) Four review rounds returned 2, 2, 2 and 4 findings. The counts looked like slow convergence; the DISTRIBUTION was the finding. Every defect after round 1 lived in this one layer — how the CLI and the dialog say "here is how to supply it" — while the server half that computes availability stayed clean throughout. The layer had accreted exactly the way IDEA-2898's cold path did: a count, then a second count for the other reason, then a phrase function, then a `continue` between two counters that made a dual-fault row report half of itself. Round 4 fixed something round 3 introduced to fix something round 2 introduced. That is not a run of bad luck, it is a shape. So this round removes branches instead of adding a seventh guard. `itemCopyTally` walks the rows once and returns what every caller needs; `AllUnfillable()` is the condition both one-sentence sites test, and `Why()` is the phrase both interpolate. Three helpers become one type. There is no second definition of "unfillable" to drift from the first, and no sentence describing a subset of what a predicate counts, because the sentence and the count come from the same walk. `Unfillable` is deliberately NOT `UnavailableTarget + EmptyKey`: one row can carry both, and double-counting makes `Unfillable == Total` false for a set that is entirely unfillable — the comparison every caller makes. A mutant does the addition and dies. Five mutants, all killed. The last needed a new test rather than a new fixture: `AllUnfillable`'s `Total > 0` guard is unreachable from both current callers, so a mutant removing it survived every command-level test. Keeping an unreachable guard and calling it defence is how a promise becomes a lie, so the tally is now unit-tested directly — an empty set is not "entirely unfillable", and a future caller outside the `len() > 0` gate would otherwise be told silently that nothing can be supplied. Claude-Session: https://claude.ai/code/session_01Xk9M5UVPdc84xL5E1mZkm8 |
||
|
|
00d650a861 |
feat(server,web): detect a revocation that writes no item, and evict the cache it left stale (IDEA-2898) (#1261)
* feat(server): fingerprint the caller's visible set on the item doors (IDEA-2898)
The delta stream can only express changes to ROWS. A revocation that writes
no item — the ordinary shape of revocation — therefore produces no signal at
all, and a client's warm local index goes on serving rows for a collection
the caller can no longer see. `ItemPicker` lists those titles.
`computeAccessEpoch` hashes the caller's EFFECTIVE visible set (collection
ids + item grants, canonicalised by sorting, separated by a byte no id can
contain) and `/items-index`, `/items-changes` and the 60s SSE tick all carry
it. The value is opaque, derived only from the caller's own access, and
costs no extra query — both lists are already resolved before the response
is written.
ONE DEFINITION, deliberately. The delta door computes its epoch from the
LIVE grant set rather than from its own include-deleted query set: a caller
holding a grant on a soft-deleted item would otherwise get a different epoch
from each door, forever, and the client would resync to its page cap on
every poll. `handlers_items_access_epoch_test.go` pins the two doors against
each other on exactly that fixture — a grant on a soft-deleted item is the
one input that discriminates, and an earlier version of this test agreed for
the wrong reason because its fixture had no grants at all.
The unrestricted caller gets a SENTINEL ("all") rather than a hash of the
empty set, because the empty set is a real and opposite state: a restricted
member with zero visible collections. Hashing both the same would make the
widest and narrowest access indistinguishable.
Claude-Session: https://claude.ai/code/session_01Xk9M5UVPdc84xL5E1mZkm8
* feat(web): compare the access fingerprint and evict the cache when it changes (IDEA-2898)
The client half of the signal. `WorkspaceState.accessEpoch` is the baseline
the cache was built under; it is persisted beside the cursor because the
revocation this closes can land while the tab is CLOSED, and a cache that
forgot its scope on reload could not detect that at all.
`ensureAccessScope` is the comparison, and it is ONE function with two call
sites — the page's `/items-changes` poll and bootstrap's own reconcile loop.
The first draft inlined it in the loop and had already diverged: the inline
copy skipped the null-baseline case, which is precisely the
offline-revocation cache the change exists for.
Three properties the tests pin, each of which was wrong at some point:
- ABSENCE IS NOT A VALUE. A server that does not send the field is a
mid-deploy older build, not a changed scope; and a snapshot carrying no
epoch must not erase a baseline we already know, which sent the reconcile
loop resyncing to its 50-page cap.
- THE RESYNC MUST TERMINATE. When the snapshot carries no epoch the
baseline would be unchanged and the next poll would ask again forever, so
the epoch we were TOLD goes in as the fallback — and again after the
await, for the case where the resync was JOINED rather than started and
the fallback was never seen.
- RAM AND DISK AGREE. `persistReplace` runs inside the resync, so the
baseline is set before it, not patched after: a durable meta row a
version behind the in-memory one makes the next warm boot resync for a
scope that never changed.
SCOPE, stated plainly because it is a real limitation and not a rounding
error: this closes the single-tab case and the offline case. A write from a
tab that has not yet learned the new scope can still reinsert a row into the
durable cache while the cache advertises the current epoch, and nothing
re-fires until the next access change. That is F2, and PLAN-2903 owns it
together with the cross-tab cache-coherence work. Strictly better than the
pre-change state, where no revocation without a row change was detected at
all.
`LOCAL_INDEX_SCHEMA_VERSION` 3 -> 4: a cache written before this has no
baseline, and adopting the incoming epoch for it would be exactly the silent
adopt the change exists to prevent. Two fixtures that hard-coded 3 now read
the constant, so the next bump does not turn them into stale-cache fixtures
by accident.
Claude-Session: https://claude.ai/code/session_01Xk9M5UVPdc84xL5E1mZkm8
* test(web): close the four seams a mutation run on the reduced tip found open (IDEA-2898)
The gates that ran on the full branch describe a different tree. A fresh
25-mutant run scoped to what this branch actually ships found four seams with
no coverage, three of them the same shape: the baseline being ADOPTED is
untested everywhere the adoption is silent.
That shape is why they survived a suite written for the eviction. Dropping an
adopt line does not stop a revocation being detected — the cache simply has no
baseline, and a null baseline over a populated cache resyncs, which LOOKS like
the change working. The discriminating case is the QUIET one: an unchanged
scope must cost nothing. All three new tests assert the absence of a resync.
- warm hydrate adopting the PERSISTED epoch (the offline-revocation case;
needs the mocked persistence module, since jsdom has no IndexedDB and the
warm branch is otherwise unreachable)
- the cold snapshot's epoch becoming the first-ever baseline
- `applyDelta` handing `persistDelta` the baseline its rows were applied
under, so the durable meta row is not stamped null after every delta
The fourth is the collection route's `ensureAccessScope` call, whose source
pin moved to PLAN-2903 with the pairing-guard argument it also covered. Half
of what it pinned still ships here, and the mutation run proved that half
uncovered — so the pin comes back NARROWED to the one surviving call site.
Its measured limit is in its docblock rather than assumed: deleting the call
kills it; short-circuiting the call (`if (false && await ...)`) SURVIVES,
because the text it matches is still on the line. A source pin cannot see
reachability. That survivor is reported, not hidden.
Instrument corrections made before any of this counted, both caught by the
runner's own controls rather than by inspection:
- `go vet` was in the Go build gate. It flags unreachable code, so the
positive control — an early `return` — scored BUILD-FAIL while compiling
perfectly. A vet failure is not a build failure.
- the web runs passed `--reporter=basic`, which this vitest does not have.
Every web mutant failed to START and the classifier read that as a
verdict. Fixed, and the classifier now requires the FULL baseline
population (50 tests across 7 files) to have run before it will call
anything killed — a mutant that stops a file loading also prints a
failing summary.
Claude-Session: https://claude.ai/code/session_01Xk9M5UVPdc84xL5E1mZkm8
* fix(web): an empty RAM state is not an empty cache, and a joined resync owes a durable epoch (IDEA-2898)
Two real defects from the review round on the reduced tip. Both are the
SILENT ADOPT class the change exists to prevent, arriving in the two places
the change itself created.
**A cache that has not answered yet is not an empty cache.** `bootstrap`
awaits `hydrate` before merging anything, but the SSE-driven `deltaSync` runs
on its own subscription rather than behind that await — so it can reach
`ensureAccessScope` with RAM empty and IDB holding rows from a scope nobody
has checked. The "nothing to evict, adopt silently" branch then stamps the
new epoch onto the durable cache through the delta that follows, and the
stale row hydrates under an epoch that agrees with the server FOREVER. Not
merely wrong once: permanently inert, which is worse than the defect this
change closes.
`cacheRead` makes the distinction the code was eliding. Declining to adopt
costs nothing and fails in the safe direction — the delta persists a null
epoch, and a null baseline over a populated cache resyncs on the next
reconcile.
**A joined resync updates RAM and leaves the disk behind.** Resyncs are
deduplicated per workspace. When `ensureAccessScope` JOINS one, that resync
already ran its own `persistReplace` under its own baseline, so assigning the
told epoch afterwards leaves the meta row recording the old one. The session
converges and every RELOAD hydrates the stale baseline and pays a full resync
for a scope that has not changed since. The code's own comment claimed the
epoch "lands in RAM and IDB together", which was true of the started path and
false of the joined one three lines below it.
`persistAccessEpoch` repairs just the epoch on an existing cache, on BOTH
branches — the first version of the fix had it on one, and no test in the
file could tell them apart until a mutant did.
Two tests changed because a PROPERTY changed, said out loud rather than
quietly rewritten: "adopts silently when there is no baseline and nothing
cached" held for a workspace that had never been bootstrapped, and no longer
does. It is now "…and the cache is known empty", and the unbootstrapped case
is its own test asserting the opposite. The `seedUnder` helper acquires its
baseline through a cold bootstrap, which is how a real session gets one
anyway.
Mutation matrix re-run whole on this tip: 31 real mutants, 30 killed. The
survivor is the page pin's short-circuit case, documented in its docblock.
Six of the mutants target these two fixes; one of them (`persistAccessEpoch`
minting a meta row) had to be rewritten after it survived for the wrong
reason — the naive version put a keyless row that IDB rejects, so the error
path compensated for the defect and the test never had to.
Claude-Session: https://claude.ai/code/session_01Xk9M5UVPdc84xL5E1mZkm8
* fix(web): a failed cache read is not an empty cache, and the epoch patch is a compare-and-set (IDEA-2898)
Round 2 on the reduced tip, and both findings are the previous round's fixes
being not quite finished.
**`hydrate` returns the same empty payload for a FAILURE as for an empty
cache** — deliberately, since a best-effort cache should not take the app
down. `cacheRead` was set unconditionally after the await, so a transient IDB
failure read as "there is nothing stored" and re-opened the exact silent
adopt round 1 closed: the durable cache may hold rows from a scope nobody
checked, and adopting stamps the new epoch onto them through the next delta.
`HydrateResult.durableRead` carries the difference the payload cannot. True
when the read succeeded, true when IndexedDB is unsupported (nothing durable
can contradict anything later), true when the read found an incompatible
cache and wiped it — false only when a database that might hold rows could
not be opened or read. Failing that way costs a resync and hides nothing.
**`persistAccessEpoch` was a blind read-modify-write on a row other writers
own.** IDB serializes transactions, but a `persistDelta` or `persistReplace`
carrying a NEWER epoch can commit between the resync this caller joined and
the patch — and the overwrite would then stamp the older epoch onto rows
fetched under the newer one, so the next comparison reports a change that
never happened and pays a full resync for it. Now a compare-and-set against
the epoch the caller believes it is repairing.
Both hydrate failure paths are covered, and they needed different fixtures:
a database at a HIGHER format version (open fails outright) and one whose
`items` store is missing (the open succeeds, the read throws). They are one
statement written twice, and a mutant on either is invisible to a test of the
other — which is how the second one was found, by a mutant surviving a test
written for the first.
The mutation matrix is 36 real mutants on this tip, 35 killed; the survivor
remains the page pin's short-circuit case, documented in that file. One
mutant is worth naming because it survived for the WRONG reason twice: the
naive removal of `persistAccessEpoch`'s "no meta row" guard cannot be
detected, because without it the code dereferences `undefined` and the catch
swallows the throw, so the outcome is identical. The faithful version — a
guard replaced by code that actually mints a valid row — is killed. A mutant
has to be the defect, not a crash that happens to look like it.
Also corrected: `accessEpoch`'s doc claimed null exists only before the first
response of a session. Two things falsify that now — a server that sends no
epoch, and the pre-hydration guard.
Claude-Session: https://claude.ai/code/session_01Xk9M5UVPdc84xL5E1mZkm8
* docs,test(web): the epoch patch's expected value is asserted, and null is not unreachable (IDEA-2898)
Round 3, two P3s and no behavioural findings.
The `HydrateResult.accessEpoch` doc said the version bump to 4 made a null
epoch unreachable for caches this build writes. It does not: a server that
predates `access_epoch` omits it, and `persistDelta`/`persistReplace` record
that absence honestly rather than inventing a value — a case the tests
already cover. What the bump actually rules out is a PRE-IDEA-2898 cache
being READ as though it had a baseline. Comment corrected to say the thing
that is true.
The two joined-resync tests asserted the epoch `persistAccessEpoch` is given
but not the value it expects to be REPLACING. The patch is a compare-and-set,
so a wrong `expectedPrevious` makes it a silent no-op that still looks right
in RAM, and the disagreement surfaces only as a resync on the next reload —
invisible to a single-session test. Both call sites now have their expected
value asserted, and two mutants (each call site given the wrong expectation)
are killed by them.
Claude-Session: https://claude.ai/code/session_01Xk9M5UVPdc84xL5E1mZkm8
|
||
|
|
6f8b6eb958 |
perf(store): clone Postgres test databases from a migrated template (TASK-2900) (#1260)
`storetest.NewPostgres` and its in-package twin `testStorePostgres` ran the
full 87-file migration chain for EVERY test that asked for a store. The schema
is identical for every one of them, so building it per test bought nothing.
Measured on this box, at
|
||
|
|
8910ad0679 |
fix(store): mint the imported workspace inside the import transaction (BUG-2892) (#1259)
ImportWorkspace called CreateWorkspace — its own committed write — before opening the transaction that carries every other row. Eight error returns sit between that INSERT and the commit (begin, de-duplicate declarations, import collection, item slug-after-truncation, import item, remap item, import comment, and the commit itself), and each returned an error while leaving the workspace row behind: named, slugged, owned by the caller, holding no collections and no items. The husk was not only clutter. uniqueWorkspaceSlug probes `WHERE slug = ? AND deleted_at IS NULL`, and a husk is not soft-deleted, so it kept the slug: an operator who fixed the bundle and retried landed on `name-2`. Measured on the unfixed build — the retry leg returns `retry-slug-2`. The attempt that stored nothing took the name from the one that worked. CreateWorkspace is now a thin wrapper over createWorkspaceQ, which takes the caller's executor, and ImportWorkspace opens its transaction first and mints on it. Both of createWorkspaceQ's reads — the slug probe and the read-back — take that executor too, which is load-bearing rather than tidy: a read routed through the pool while the caller's transaction holds its connection can wait for a free one, and under MaxOpenConns(1) there is none (BUG-2778 / BUG-2409). workspaces.slug has been globally UNIQUE since 001_initial, so the constraint still covers the race the in-transaction probe cannot. The filing said this needed a look rather than a one-liner because CreateWorkspace "does more than one INSERT (owner membership, seeding hooks)". That premise was wrong, and reading the function is what retired it: it does one INSERT. Owner membership and template seeding are both handler-level, and neither is on the import path inside the store — handleImportWorkspace calls AddWorkspaceMember after ImportWorkspace returns, and import never seeds because the bundle carries the collections. Tests: two legs, both red on the unfixed build for the reasons named above and green here. The negative fixture is a bundle whose two collections share a slug and differ in everything else, so it fails for exactly one reason, asserted on by message. CONVE-23 sweep: two comments said the workspace "survives as a husk" when the de-duplication rolls an import back (export.go's trait-dedupe note and TestImportDeduplicatesAConflictingArchive). Both were true when written and are now false in that clause only; corrected without disturbing the argument they were making, which is unchanged. CONVE-18 sweep: the class is a committed pool write followed by a transaction whose failure orphans it. One instance across the 64 non-test files of internal/store — this one. The one other hit is a false positive (CreateItemLink's early `return s.SetParentLink(...)`, a mutually exclusive path). Search boundary: internal/store non-test files, by syntax; the instrument recognises `s.db.Exec(` and s.<write-verb> calls, does not model control flow, and does not follow writes into helpers named otherwise. It was run against the pre-fix export.go as a positive control and fired on exactly the defect. Claude-Session: https://claude.ai/code/session_01HeChkgZVYb3NTgTcckF5KR |
||
|
|
07b2e439f2 |
TASK-2869 (U2b): the preflight names a relation's target collection, and the copy dialog scopes its picker to the destination (#1258)
* feat: the preflight's needs-value row names its relation target, and the copy dialog scopes its picker to the destination (TASK-2869)
U2b, per the day-55 ruling. Both blockers (U1 referent validation, U2 the
FieldEditor branch + picker) are in.
THE DEFECT. `ItemCopyPreflightNeedsValue` carried `type: "relation"` and no
target. FieldEditor gates its relation branch on `wsSlug` AND
`field.collection`, so a required relation in the destination reached the copy
dialog as a field it knew was a relation with no idea what to point at, and
rendered as FREE TEXT. Before U1 the copy stored whatever was typed.
SERVER. `collection` is added to the needs-value row, populated from the
DESTINATION schema's `def.Collection` — the only place it is known, since the
row is built from that schema. Additive and `omitempty`: a client that does not
read it is unaffected, and a row for a non-relation field is byte-identical to
before. Not a wire-version question, for the same reason
`models.ItemWriteWarnings` was not.
CLIENT. `toFieldDef` carries the collection through, and the FieldEditor call
passes `wsSlug={destWs}` — the DESTINATION, never the source. A relation
resolves at the destination, so the picker must list items the copy can
actually point at; that is same-workspace resolution AT the destination, not
the cross-workspace case PLAN-2857 rules out.
`relation` becomes collectable ONLY IF THE ROW NAMES ITS TARGET. Without a
collection, FieldEditor's gate renders the non-editable state, so offering the
row would produce a control that cannot be filled and a Confirm that cannot be
satisfied. Such a row now lands in the blocked list and the user is told which
field and why — the same disposition `multi_select` gets, for the same reason:
a control that silently cannot do its job is worse than an honest refusal.
TWO THINGS THIS UNIT TAUGHT ME THAT ARE NOT IN THE RULING.
1. THE WEB MUTANT SURVIVED, AND THAT IS WHY `isCollectable` MOVED. My first
version left the predicate inline in `CopyItemDialog.svelte`. Making
`relation` unconditionally collectable — the exact defect the negative leg
of the proving test is about — passed EVERY suite in the repo. That is
IDEA-2894's lesson arriving one unit later in the same file, so the
predicate now lives in `$lib/items/copyNeedsValue` with tests. Two mutants
die there: relation-always-collectable, and `multi_select` slipped into the
collectable set.
The Go half was pinned from the start (drop `Collection: def.Collection` ->
FAIL naming the empty value and the expected slug). Only the client half was
unpinned, and only because of where the code lived.
2. A U2-ERA TEST ASSERTED THE ABSENCE THIS UNIT CLOSES, and asserted it
CORRECTLY. `fieldEditorRelationCallers.test.ts` required that the dialog
pass no `wsSlug` and build its FieldDef from a shape with no `collection` —
which was the behaviour, and withholding `wsSlug` was what kept an unscoped
picker out. It also named this task by ref and told its successor to revisit
the gate WITH the change rather than let it drift. Inverted here: the block
now asserts the destination slug is passed, that the collection reaches the
FieldDef, and — the half that is easy to lose — that the dialog still
DELEGATES the collectability decision, so a future inlined predicate would
pass the unit tests and fail this.
Worth keeping: a test that pins a temporary absence should name what would
make it wrong. This one did, and that is the only reason its inversion was a
five-minute job instead of an argument about whether it was load-bearing.
Gates: `internal/server` ok 158.335s, `go vet` and `gofmt` clean,
`npm run check` 0 errors (6 pre-existing warnings), `make web-test` 127 files /
2123 tests. Postgres and CI are owed on this tip.
Claude-Session: https://claude.ai/code/session_01Xk9M5UVPdc84xL5E1mZkm8
* fix(cli): mirror the needs-value collection, and name the relation target in the CLI (TASK-2869)
TWO CONSUMERS I DID NOT SWEEP. The previous commit added `collection` to the
preflight's needs-value row and updated the server struct and the TypeScript
type — "both sides", as its own message put it. There are THREE sides.
`internal/cli` keeps a mirror of the preflight shape and
`TestItemCopyMirrorsMatchServerShapes` requires it to match the server field
for field. It failed in the Postgres gate, in a package the change did not
touch.
That is the third time this session a producer change broke a consumer I had
not enumerated, and the shape is always the same: I name the surfaces I edited
and call that the population. The instrument that would have caught it is not
"run more tests" but "grep for the type's name before claiming the sweep is
done" — `ItemCopyPreflightNeedsValue` appears in exactly three files and I
looked at two.
Mirrored, with a comment saying why it exists: a mirror that silently lags is a
mirror that lies, and the CLI renders these rows.
AND THE CLI NOW NAMES THE TARGET COLLECTION, which is the point of the unit on
the surface that has no picker at all. A row reading
owner_ref (Owner, relation) required — required, with no value…
tells a user a value is needed and nothing about what kind of value exists.
The dialog answers that with a scoped picker; the CLI had no answer. It now
prints the relation analogue of the `options:` line a select already gets:
owner_ref (Owner, relation) required — …
target collection: people
Test asserts the line appears for the relation row, appears EXACTLY ONCE with a
select row rendered alongside — so it cannot pass by printing unconditionally —
and that the select's own `options:` line still renders, so this did not
displace it.
Gates: `internal/cli` and `cmd/pad` green, build and gofmt clean. The full
Postgres run and CI are owed on this tip; the earlier PG run is the one that
caught the mirror and is superseded.
Claude-Session: https://claude.ai/code/session_01Xk9M5UVPdc84xL5E1mZkm8
* fix(web): the relation picker searches the preflight's canonical destination slug (TASK-2869)
Codex review, finding 3 of 3, and the only one of the three that belongs in
this unit.
`wsSlug={destWs}` handed `FieldEditor` a value that is NOT always a workspace
slug. An item can be opened through a workspace-UUID URL, the route parameter
is passed straight through as `sourceWsSlug`, and a same-workspace copy then
puts that UUID in `destWs`. `/search` resolves a workspace by SLUG only, so a
picker handed a UUID searches nothing and returns no results — a control that
looks usable, is not, and says nothing about why.
Now `pickerWsSlug`, which is the preflight response's own
`destination.workspace_slug`. The preflight IS the canonicalising round-trip:
the server resolved whatever it was given and answered with the real slug.
Falls back to `destWs` only before the first preflight returns, at which point
no needs-value row is rendered anyway.
The caller test asserts the prop AND the derivation, because asserting only the
prop would pass against a `pickerWsSlug` that was just `destWs` renamed.
THE OTHER TWO FINDINGS ARE REAL AND ARE FILED, NOT FIXED HERE.
IDEA-2898 — `ItemPicker` serves warm local-index results without re-authorising
them, so a collection whose access was revoked can still be listed. The cold
`/search` path is visibility-filtered and correct; the warm path is not. This
is PRE-EXISTING and applies to every caller of the picker, `ItemDetail`
included — last touched by TASK-2877, not by this unit. U2b widened the
exposure by adding a caller; it did not create the defect, and rewriting the
picker's cache-authorisation model inside a feature branch would be an
unrelated change riding along. The fix needs a decision about where the client
learns its access set from, which no current signal provides.
IDEA-2899 — a relation row can name a target collection that is DELETED or
UNREADABLE, so `isCollectable` says yes on a non-empty string and the user
meets a picker with nothing in it. I could not fix this correctly here, and the
reason is measured rather than assumed: the obvious test is to check the target
against `destCollections`, and that list is filtered by `canEditCollection` —
it is what the user may copy INTO. A relation TARGET needs only READ access, so
a perfectly usable target routinely is not in it. Using it would OVER-BLOCK,
refusing rows the user could have filled, which is a worse failure than the one
being fixed and invisible to whoever hits it. The right shape is probably the
server reporting the target's availability on the row it already builds — an
additive field on the same row this unit just changed, worth doing deliberately
rather than bolted on at the end of a branch.
Gates: `npm run check` 0 errors (6 pre-existing warnings), `make web-test` 127
files / 2123 tests. Postgres is running on this tip; CI is owed.
Claude-Session: https://claude.ai/code/session_01Xk9M5UVPdc84xL5E1mZkm8
|
||
|
|
47d2b15e10 |
feat(store,server): make collection trait uniqueness a database invariant — de-dup pass + partial unique indexes (TASK-2710) (#1257)
* feat(store): make one-collection-per-trait a database invariant (TASK-2710) Partial unique indexes on both drivers over the artifact_kind and invocation_field declarations, with the de-duplication pass that has to precede them. The de-dup is Go, not SQL, and runs BEFORE migrate(). The ruling requires every resolution to be REPORTED, because it silently changes which collection owns a kernel behaviour, and a SQL migration cannot log — Postgres RAISE NOTICE goes nowhere here and SQLite has no equivalent. Splitting decide-in-SQL from report-in-Go would give one rule two spellings to keep in step, which is the defect class IDEA-2883 closed. It runs before migrate() because the CREATE statements fail on exactly the databases needing repair. The rule, after three proposals each retired by a measurement: most user-written items wins; ties break on lowest (created_at, id), reported as ARBITRARY rather than as age, because created_at is second-resolution and newID() is a random uuid v4, so same-second rows carry no age at all. The loser keeps every item and loses only the declaration. Routing MAY change on affected deployments; there is no current behaviour to preserve, since with two declarations live the winner was measured flipping between runs on Postgres. SeedCollectionsFromTemplate now skips a definition whose artifact kind is already declared. Renaming re-slugs, so a workspace whose conventions became house-rules looked slug-empty while its kind was still claimed; seeding used to mint the duplicate and would now fail the whole seed instead. Claude-Session: https://claude.ai/code/session_01HeChkgZVYb3NTgTcckF5KR * test(server): keep the shadowing tests under the new trait invariant (TASK-2710) TestResolvePlaybookIgnoresInvisibleCollections and TestCollectionIDForKindIgnoresInvisibleCollections build two collections declaring one trait, which the partial unique indexes now forbid. They are not obsolete and I did not weaken them. They guard the round-2 shadowing fix: when two collections declare one kind and the first-sorting one is invisible to the caller, resolution must return the visible one instead of failing. TASK-2710 makes that state unrepresentable going forward and repairs it at startup on databases holding it, but the resolver is what stands between a legacy database and a wrong answer in the window before that repair, and on any deployment where an operator dropped the index. So the fixture now constructs the state the way it exists in the wild — with the constraint suspended — and the assertions are untouched. Same reasoning as IDEA-2883's disagreeing-reminder fixture. Claude-Session: https://claude.ai/code/session_01HeChkgZVYb3NTgTcckF5KR * test(store): SuspendTraitUniquenessForTesting beside its neighbour, with a restore (TASK-2710) Lead ruling: shape it like SetBcryptCostForTesting — same file, ForTesting suffix, returns a restore the caller defers, so a suspended constraint cannot outlive the test that suspended it. The restore recreates the indexes from the SHIPPED migration text rather than a hand-copied approximation, which would drift and then attest to an index the product does not have. Its failure is information, not noise: recreating a unique index while a duplicate is live is exactly what the migration would hit. The de-dup tests therefore assert the restore SUCCEEDS, which is the migration's precondition checked rather than assumed; the server shadowing tests leave the duplicate live for their whole duration and ignore the error explicitly rather than by omission. Claude-Session: https://claude.ai/code/session_01HeChkgZVYb3NTgTcckF5KR * fix(store): de-dup uses the index's own extraction, strips every declaration once, counts NULL source as user-written (TASK-2710) Three P1s from codex round 1, each verified before accepting. A collection can lose BOTH declarations — the playbooks definition declares artifact_kind and invocation_field — and resolving them in two passes, each re-parsing the row's ORIGINAL traits, made the second write restore what the first stripped. The duplicate survived and the migration would still have failed on it, surfacing as a broken upgrade rather than a test. Now one write per collection accumulating every strip, with a regression test. Detection asked Go what a declaration is while the indexes ask json_extract / ->>, so the two could disagree: a row the Go parser rejects still carries a value the index sees, and the de-dup would leave a pair the CREATE then refuses. It now asks the database the same question the index asks, which is the same one-rule-one-spelling reasoning that put the report in Go. source IS NULL now counts as user-written. The column is nullable and legacy rows predate it; 'source <> template' alone is NULL for those, which SQL treats as not-true, so a workspace whose only user content is old would have had it ignored when picking the winner. Claude-Session: https://claude.ai/code/session_01HeChkgZVYb3NTgTcckF5KR * fix(server,store): seed skips either held declaration; unique violations map to 409 on both drivers (TASK-2710) Two P2s from codex round 1. The seeder checked only artifact_kind, but the playbooks definition also declares invocation_field and TASK-2710 adds an index for each — so a workspace whose invocation-routing collection had been renamed would still have failed its seed. It now skips when EITHER declaration is held, and says which. Collection create recognised only SQLite's "UNIQUE constraint" text, so the identical race answered 409 on SQLite and 500 on Postgres; the update path recognised neither. Both now use one named isUniqueViolation covering both drivers, matching what every item handler already did. The conflict message also distinguishes the indexes: "a collection with this name already exists" is actively misleading for a trait conflict, where the name is fine and the declaration is taken — a user told to rename would rename forever. Claude-Session: https://claude.ai/code/session_01HeChkgZVYb3NTgTcckF5KR * chore(store): drop the unused order slice from the de-dup pass (TASK-2710) Leftover scaffolding from the one-write-per-collection rewrite; staticcheck caught it (SA4010). Mine to catch earlier — I ran lint at the tip BEFORE that rewrite and not after it, so the gate found what a re-run would have. Claude-Session: https://claude.ai/code/session_01HeChkgZVYb3NTgTcckF5KR * fix(store): de-duplicate a conflicting archive on import; guard json_extract against malformed traits (TASK-2710) Both from codex round 2, both real. Import warned about a duplicate declaration and inserted both, which was right while nothing forbade the pair. With the unique indexes the second INSERT is refused, the whole transaction rolls back, and the workspace minted beforehand survives as a husk (BUG-2892) — so an archive carrying a duplicate would become unimportable, and those archives are exactly the ones this release repairs. This is the task's item 4, which I had not done. The first declaring collection in bundle order keeps it and later ones are stripped and reported; the rule cannot use user-item counts here because items are inserted after collections, so bundle order IS the terminator and the log says so. SQLite's json_extract RAISES on malformed JSON rather than returning NULL, so the unguarded expressions in the index predicates and the de-dup scan would have failed STARTUP on any database holding one bad blob. json_valid now guards both, matching what every other reader does with malformed traits — treat the row as declaring nothing. Postgres needs no equivalent: traits is JSONB, so the column type makes malformed content unrepresentable at rest. Claude-Session: https://claude.ai/code/session_01HeChkgZVYb3NTgTcckF5KR * test(store): assert the malformed-traits asymmetry per driver (TASK-2710) My own Postgres gate caught this: the test planted a malformed traits blob and Postgres refused it — invalid input syntax for type json — because traits is JSONB there. That refusal IS the reason migration 064 carries no json_valid guard while 087 does, and it was prose in the migration until the gate turned it into an observation. The test now asserts it per driver: on Postgres the plant must be REFUSED, on SQLite it must succeed and the guard must keep both the de-dup pass and index creation working. It therefore also catches someone 'fixing' the asymmetry later — adding a guard Postgres does not need, or dropping the one SQLite does. Claude-Session: https://claude.ai/code/session_01HeChkgZVYb3NTgTcckF5KR * fix(store,server): de-duplicate on the bytes being written; stop claiming a name conflict for item-index violations (TASK-2710) Round 3, two findings. P1, a regression I introduced: the import de-duplication pre-computed its strips from the traits the BUNDLE carries, which is not what gets written — coercion, validation-discard and canonical inference all run afterwards. A pre-traits archive carrying conventions with an empty blob has its declaration INFERRED from the slug (BUG-2702), so a bundle pairing that with an explicit declarer showed the pre-pass one declaration and the database two, and the index aborted the whole import. The check now runs immediately before the INSERT, on the final bytes, which turns the question from 'what did the file say' into 'what am I about to write'. Reproduced first, then fixed. P2: a collection UPDATE can migrate item field values, and an item-level unique index can fail there — invocation_slug is the live example. My catch-all reported that as 'a collection with this name already exists', sending the caller to rename something that is not the problem. The name message now requires the error to name the collections table; otherwise it says what it knows. Claude-Session: https://claude.ai/code/session_01HeChkgZVYb3NTgTcckF5KR * test(store): fail when an archived collection takes the live one's declaration (TASK-2710) The test for the rebase onto BUG-2884, written before the fix and failing against the naive resolution: live collections declaring convention = [], want exactly [conventions] BUG-2884 made the bundle carry soft-deleted collections. This branch moved import's duplicate-declaration check out of a pre-pass and into the insert loop, so it operates on the bytes actually being written (round 3's P1) — but `dropDuplicateImportDeclarations` has no notion of liveness, which the pre-pass had gained on main. An archived collection travelling ahead of the live one that replaced it therefore CLAIMS the kind, and the live collection is stripped of it. Every resolver filters deleted_at IS NULL, so the workspace imports with no live convention routing at all. The second assertion is the other direction: the archived collection must KEEP its declaration. Nothing routes to it, both partial unique indexes exclude it, and stripping it would edit data the operator archived rather than deleted. Claude-Session: https://claude.ai/code/session_01HeChkgZVYb3NTgTcckF5KR * fix(store): an archived collection neither takes nor loses a trait declaration on import (TASK-2710) The rebase fix for the test in the previous commit. dropDuplicateImportDeclarations now returns an archived collection's traits untouched. A soft-deleted row sits outside both partial unique indexes (each carries `AND deleted_at IS NULL`) and outside every trait resolver, so it can neither create the conflict this function prevents nor be harmed by holding a stale declaration. Letting it take a claim was the real damage: the live collection later in the bundle lost the declaration and the workspace imported with no routing for that kind at all. BUG-2884's pre-pass had grown the same condition; this branch replaced that pre-pass with an in-loop check on the final bytes (round 3's P1) and the condition did not come with it. Keeping both is what the rebase owes. TestImportRoutingIgnoresSoftDeletedCollections builds its fixture in a new order — declare, archive, then seed — because the unique index refuses two LIVE collections declaring one kind. The order is not a workaround: it is the production path that mints this state (delete the conventions collection, seed again), every step legal under the invariant, and it needs no test-only suspension of the constraint. Its assertions are unchanged. Claude-Session: https://claude.ai/code/session_01HeChkgZVYb3NTgTcckF5KR * docs(server): checkTraitConflicts no longer claims the index it now has (TASK-2710) CONVE-23 sweep. The doc comment describing that gate was written when the invariant did not exist and this branch falsified three of its sentences: "workspace IMPORT bypasses it entirely by design" (import de-duplicates on the way in now), "the database-level version is deliberately NOT added in phase 0" (migration 087 adds it), and the closing paragraph handing duplicates back to the resolvers' order-dependent behaviour. Rewritten to say what the division of labour actually is — the pre-check survives for the MESSAGE, because a unique violation is a 409 about a name unless something tells the handler otherwise and "rename your collection" is useless when the name is fine and the declaration is taken; the index is what holds. It also states the two things the invariant genuinely does not cover: import (which de-duplicates rather than refusing a restore) and archived collections (outside both indexes and every resolver, so a soft-deleted row may hold a declaration a live one also holds). Claude-Session: https://claude.ai/code/session_01HeChkgZVYb3NTgTcckF5KR * fix(store): take the pre-migration snapshot before the trait repair writes (TASK-2710) Codex round 5, P2, verified. dedupeTraitDeclarations ran from the two constructors, ahead of migrate() — and snapshotBeforeMigrate() runs INSIDE migrate(). The repair changes data: it strips a declaration, moving which collection owns a kernel behavior. So the altered ownership was already committed when the snapshot was copied, and `<db>.pre-<version>` — the operator's rollback for a bad upgrade — contained it. Restoring after a failed migration handed back the old schema with the repair applied and unrecorded: the one thing the rollback could not undo was the only thing that had silently changed routing. The call moves into migrate(), immediately after the snapshot and before the migration loop, which satisfies both constraints at once — 087 / 064 still cannot run against a database holding duplicates, and the snapshot now precedes the write. The Postgres path takes the same position for symmetry; there is no snapshot there to sit after, so the ordering argument is one-sided on that dialect and the comment says so. TestTheSnapshotIsTakenBeforeTheRepairWrites pins it end to end: plant a duplicate, un-apply 087 so a migration is genuinely pending, reopen, then read the snapshot with the RAW driver — New() would migrate and repair the snapshot too, destroying the thing being measured — and assert it still holds both declarations while the live database holds one. Claude-Session: https://claude.ai/code/session_01HeChkgZVYb3NTgTcckF5KR |
||
|
|
d6771abeea |
IDEA-2886: state the drop-vs-refuse rule once, and route all three sites through it (#1256)
* refactor(store): state drop-vs-refuse once, and route all three sites through it (IDEA-2886)
The rule — a relation value the CALLER ASSERTED is refused if it does not
resolve; one asserted by NOBODY is dropped and reported — was stated three
times, each in its own control flow and its own prose:
- the origin split in MigrateRelationReferentsQ (supplied -> refusals, the
other two -> dropped);
- ResolveLateRelationDefaultsQ, whose return is named `dropped` because
everything it sees is a default;
- IssuesForCallerInput, which kept an issue if its key was present before
validation.
Three statements of one rule is the shape that invites a fourth door to state
it slightly differently, which is the drift TASK-2878 existed to remove. I
filed this out of that unit after codex round 10 stated the rule for the third
time.
The polarity now lives in ONE place, `RelationOrigin.Refuses()`, and the three
sites consult it:
- the supplied branch asks the rule instead of assuming its own identity.
Its `else` is unreachable while Refuses() answers as it does, and that IS
the point: if the rule changes, this site follows it rather than
contradicting it.
- the late-default pass opens with an assertion that its origin does not
refuse, so an edit that starts refusing there fails loudly instead of
quietly forking the rule. Its prose now says "this is
RelationOriginDestinationDefault.Refuses() being false", not "this is a
decision this function makes".
- IssuesForCallerInput CLASSIFIES to an origin and then asks, rather than
restating the rule as "keep it if it was there before".
BEHAVIOUR IS UNCHANGED AT EVERY DOOR, which is the spec and not a caveat.
WHAT THE MUTANTS CAN AND CANNOT SHOW, stated plainly because a behaviour-
preserving refactor is exactly where a mutation matrix lies:
- PER-SITE mutants are impossible BY CONSTRUCTION. Restoring any one site's
old form is behaviour-identical, so no test can distinguish it. Reporting
"3/3 survived" would be true and would mean nothing.
- The MASTER mutant is the real one. Flipping Refuses() to `!=` compiles and
kills 116 server subtests and 1 store subtest, so the rule is load-bearing
rather than decorative.
- PER-SITE REACHABILITY is the instrument that answers the question the
mutants cannot: does each site actually consult the rule, or keep a private
copy? Making Refuses() panic and running one test per site:
TestRelationDoors_Move -> panics
TestRelationDoors_Create -> panics
TestRelationDoors_MoveDropsInvisibleDestinationDefault -> panics
TestRelationDoors_LateDefaultWrongCollectionDoesNotDisclose -> panics
TestRelationDoors_RequiredInvisibleDefaultRefuses -> panics
My first run of that probe reported ZERO panics for the late-default site
and I nearly wrote it up as unreached. The filter I used
(`-run TestResolveLateRelationDefaults`) matched NO TESTS — the reading was
my instrument, not the code. I have the `-run`-narrower-than-the-population
trap in my own notes from a previous unit and walked into it again, so the
re-run above counts `=== RUN` lines alongside the panics: an instrument
that cannot say how many tests it ran cannot tell silence from absence.
Gates: `internal/store` ok 188.174s, `internal/server` ok 206.077s (SQLite),
`gofmt` and `go vet` clean. Postgres and CI are owed on this tip.
Claude-Session: https://claude.ai/code/session_01Xk9M5UVPdc84xL5E1mZkm8
* fix(store): make the supplied write-back survivor-guarded, as its sibling already is
Review finding on the previous commit, and the finding is exactly the one I
asked the round to look for: the new `else` in the supplied path is
UNREACHABLE, and it was WRONG.
It deleted each unresolvable key from `fieldMap` and then fell into a
write-back loop that copied every entry of `suppliedRelations` back in —
restoring the value it had just dropped. The result would report a drop while
retaining the dropped value, which is worse than either outcome alone. An
unreachable branch that is wrong is a trap for whoever makes it reachable, and
the whole reason that branch exists is that someone might.
FIXED BY CONSTRUCTION RATHER THAN BY BOOKKEEPING. My first fix deleted the key
from `suppliedRelations` as well, which works and leaves two loops that have to
agree with each other. The write-back is now guarded on survival —
if _, survived := fieldMap[k]; !survived { continue }
— which is character-for-character what the DESTINATION-DEFAULT branch twenty
lines below already does, and for the same reason its comment gives. The two
branches differ in their disposition; they no longer differ in their
write-back. When I notice I am writing a second loop to compensate for the
first, the compensation is usually the sign that the structure is wrong.
BEHAVIOUR IS UNCHANGED ON THE REACHABLE PATH, and here is why rather than an
assertion that it is: in the refusing path nothing is deleted from `fieldMap`,
`suppliedRelations` was built from `fieldMap`'s own keys, and
`ResolveRelationReferentsQ` canonicalises the copy it is handed without
touching `fieldMap` — so every supplied key survives the guard and the loop
does what it did before.
WHY THERE IS NO TEST FOR THE BRANCH, named rather than left looking covered:
exercising it requires inverting `Refuses()`, which inverts the rule for every
door at once and fails 116 subtests for unrelated reasons. A probe that cannot
isolate what it is probing proves nothing, so I did not keep the one I tried.
The reviewer's exhibit — schema `{Key:"r", Type:"relation", Collection:""}`,
`fieldMap={"r":"bad"}`, `supplied={"r":"bad"}` — is recorded here instead, and
the structural fix means the branch is now correct without needing to be run.
Gates on this tip: `internal/store` ok 183.002s, `internal/server` ok 160.804s,
`gofmt` clean. Postgres was green on the parent (29 packages, EXIT=0) and is
owed again here.
Claude-Session: https://claude.ai/code/session_01Xk9M5UVPdc84xL5E1mZkm8
|
||
|
|
47cc106ab8 |
IDEA-2893 + IDEA-2894: record the accepted carry disclosure, and make the drop-reason mapper testable (#1254)
* refactor(web): extract the copy dialog's drop-reason mapper so it can be tested (IDEA-2894)
The mapping from a server drop reason to the sentence a user reads lived inline
in `CopyItemDialog.svelte`, unexported, with no test file for the component at
all. Two separate review rounds found defects in it and NEITHER FIX WAS PINNED
BY ANYTHING:
- round 12: the UI asserted NON-EXISTENCE from `not_found`, which the server
also emits for a target the caller merely cannot see. Telling those apart
is the existence oracle the collapse exists to prevent.
- round 18: `referent_not_portable` read "it points at something in the
source workspace", claiming both existence and location for a reason
emitted WITHOUT resolving the target — and which `github_pr` reaches too,
where the referent is in no workspace at all.
A third defect of the same shape would have been found the same way, by a
reviewer happening to read it, or not at all.
Moved to `$lib/items/copyDropReasons` with the reason vocabulary as an explicit
exported list, which makes two tests possible that could not be written before:
1. EVERY reason the server can emit has a sentence. This is round 12's other
finding as a test: BUG-2674 added `referent_not_portable` server-side,
nothing here learned it, and it rendered through the fallback as a raw enum
string in front of a user. A reason that maps to itself IS that defect.
Mutant: delete `referent_not_portable`'s message -> FAIL.
2. THE TWO HAZARDOUS REASONS STAY NEUTRAL. `not_found` and
`referent_not_portable` must not claim a target exists, does not exist, or
say where it is. Asserted over those two rather than all ten:
`wrong_collection` legitimately says the target is outside the field's
collection, and it may, because the server only emits it to a caller who
can SEE the target.
Mutant: restore round 18's wording -> FAIL, naming the sentence and why.
The unknown-reason fallback returns the raw string, and a third test pins that
deliberately: a reason this build has never heard of means the server is ahead
of the client, and showing the enum is more honest than inventing a sentence or
hiding the row. It is also what keeps test 1 from being vacuous.
WHAT THIS DOES NOT FIX, stated because the list is the thing a future reader
will trust: the reason vocabulary is DUPLICATED from Go (five constants in
`handlers_items_copy_preflight.go`, five in `internal/store/relation_referents.go`)
rather than generated, so it can still go stale in the one direction that
matters — a reason added to Go and not added here. The test cannot see that.
What it can see is a reason listed here without a sentence, and the list is now
the single place to update. Generating it from the Go constants would close the
gap properly and is a bigger change than this one.
Gates: `npm run check` 0 errors (6 pre-existing warnings in unrelated files),
`make web-test` 126 files / 2118 tests passed. Frontend-only; no Go touched.
Claude-Session: https://claude.ai/code/session_01Xk9M5UVPdc84xL5E1mZkm8
* docs(store): record why a carried relation's survival is observable and accepted (IDEA-2893)
Comment only; no behaviour change.
A carried relation value naming a live item in a collection the mover cannot
see resolves and survives a same-workspace move, while one naming nothing is
dropped — so a mover can tell those apart, and on a stored REF they also learn
the target's canonical id. The lead's ruling is ACCEPT AND DOCUMENT, and this
is the documentation, placed at the branch that produces the behaviour rather
than in a doc nobody reading that code will open.
Four measurements decided it, and the comment carries the two that matter so a
future reviewer reaches the reasoning instead of re-deriving it:
- NOT ENUMERABLE. A caller cannot choose what to test: create, update and
fields_patch all refuse a caller-supplied ref naming an item they cannot
see, with the COLLAPSED `not_found` wording, so no door turns a chosen
value into a carried one. It can only confirm a value already sitting in
an item the caller can read and did not put there.
- THEY ALREADY HAVE THE VALUE. An ordinary GET returns the raw stored
relation value verbatim; reads apply no redaction. The increment is "it
currently resolves" plus the ref-to-id mapping — not the target's
existence in any general sense, not its title, and backlinks do not widen
it either.
The comment also records why each candidate fix costs more than the increment,
because "we considered it" is worth nothing to a successor without the reasons:
redacting the response closes nothing (the id is in the blob and comes back
from a plain GET); not canonicalising removes only the id half and makes a
relation value stop meaning one thing everywhere; dropping by the MOVER's
visibility silently destroys a valid relation because of who moved the item;
and canonicalising only for movers who can see the target would make the
STORED BYTES depend on who performed the move.
The last paragraph is the one I most want read. The ONE change that would close
this is carrying unresolvable values verbatim instead of dropping them — which
is precisely the drop-and-report rule three lines below, the rule #1246 exists
to enforce. So the comment doubles as a warning: changing that rule changes
this, in the other direction. A design decision that is load-bearing for a
second decision should say so at the place where someone would change it.
I also WITHDREW my own earlier proposal rather than leaving it standing:
"evaluate the carry against the ITEM's access rather than the mover's" was
confused, because an item has no access identity and resolving without a
requester IS the status quo. A plausible-sounding option left in a trail is a
successor's wasted afternoon.
Gates: `gofmt` clean, `go build ./internal/store/`, `go vet ./internal/store/`,
`go test ./internal/store -run 'TestRelation|TestMigrateRelation'` ok. Scoped
to the package because this worktree has no `web/build` for the root embed, and
the change is a comment in one file.
Claude-Session: https://claude.ai/code/session_01Xk9M5UVPdc84xL5E1mZkm8
* fix(test): repoint the drop-reason parity gate at the extracted module (IDEA-2894)
`f2c4a722` moved the drop-reason mapper out of `CopyItemDialog.svelte` and
broke `TestCopyPreflightDropReasonsAreRenderedByTheDialog`, which reads the
component for that function. CI caught it; my local gates did not, because this
worktree had no `web/build` for the root embed so I had scoped the Go run to
`internal/store` — and I stated that boundary in the commit message while it
was hiding a real failure. Naming a gate's scope is not the same as the scope
being adequate.
THE GATE FAILED THE RIGHT WAY, and that is worth recording. It does not search
the file for `case 'not_found':` and shrug when the file changes; it looks for
the declaration by name and calls `t.Fatalf` if it is gone, saying "this gate
is reading for a function that moved or was renamed, so its green means nothing
until it is repointed". A parity gate that cannot tell "no such reason" from
"no such function" is worse than none, because the second reads as the first
passing.
Repointed at `web/src/lib/items/copyDropReasons.ts` and STRENGTHENED, because
the extraction split the thing it was checking in two. It now requires each
server reason to appear in BOTH:
- `COPY_DROP_REASONS`, the exported list;
- the `MESSAGES` map.
They fail differently, and the first is the one that matters. The module's own
completeness test ITERATES that list, so a reason missing from the list is
invisible to that test as well — this gate is the only place it shows. A test
driven by a list cannot notice something absent from the list.
I ALSO HAVE A CORRECTION TO MAKE, to my own prose in `f2c4a722`. That commit
message and the PR body say the Go-to-TypeScript direction "can still go stale
in the one direction that matters — a reason added to Go and not added here.
The test cannot see that." That is FALSE, and I wrote it without checking: this
parity gate has enumerated the Go vocabulary and required a renderer for every
entry since BUG-2674, which is precisely that direction. My TS test cannot see
it; the repo already had a test that could, and I asserted its absence rather
than looking. Same failure as the fixture I designed around a hazard yesterday
instead of asking whether the product had it — a claim about what is NOT
covered owes a grep exactly as much as a claim about what is.
The PR body is corrected in the same push.
Mutants, all three arms, each restored after:
- remove a reason from the LIST -> FAIL, naming the list and why the module's
own test cannot see it;
- remove its entry from the MESSAGES map -> FAIL, naming the fallthrough;
- rename the `MESSAGES` declaration -> FATAL with the repoint message, so the
fail-safe itself is exercised rather than assumed.
Gates: `go test ./internal/server` ok **279.535s** — the full package this
time, with `web/build` populated so the root embed resolves. `gofmt` clean.
Claude-Session: https://claude.ai/code/session_01Xk9M5UVPdc84xL5E1mZkm8
* docs(store): point the carry comment at the trail rather than at a person (IDEA-2893)
Comment wording only.
The attribution now reads `(IDEA-2893, lead ruling day 58; the measurements it
rests on are on that idea's trail, which is where to check this reasoning
rather than take it)`.
The lead asked for `confirmed by Dave in chat` and I declined to write it: Dave
had said nothing to me about this disposition, so the only evidence was a relay
through a channel BUG-2542 proved cannot carry provenance, and the artifact is
a permanent comment asserting what a specific person decided. The lead withdrew
the line and agreed the hold was right. Dave then ruled the general case — no
code comment needs to name him — which is the wording above and is better than
either version, because a ref is CHECKABLE and a name is not. A reader who
doubts this comment can open the idea and read the measurement; a reader who
meets a name can only take it or leave it.
Recorded team-side as CONVE-32 so the successor does not relearn it: code
comments cite the trail, never a person by name. Its scope is source comments
only — commit messages, PR bodies and trail comments are where naming who
decided something is often the entire content, and those artifacts sit beside
their own evidence.
Two things the convention says out loud rather than gloss:
- The rule reached me as a RELAY, and I acted on it because it only ever
REMOVES a claim about a person. Acting on a relay to stop asserting
something is safe in a way that acting on a relay to start asserting it is
not — which is the same distinction that made the hold correct an hour
earlier. Read as general licence to act on relayed instruction it would be
a misreading; the direction is the whole point.
- About nineteen comments already in `internal/` name a person. They are NOT
rewritten. Churning merged history to apply a new rule retroactively costs
more than it returns and a sibling rebasing onto it pays the bill. Fix one
only while editing that comment for another reason.
Gates: `gofmt` clean, `go build ./internal/store/`, and the parity gate green
(`TestCopyPreflightDropReasonsAreRenderedByTheDialog` ok) since this touches
the same file the previous commit repointed it away from.
Claude-Session: https://claude.ai/code/session_01Xk9M5UVPdc84xL5E1mZkm8
|
||
|
|
5cb2297f67 |
fix(store): export soft-deleted collections so their live items survive a round trip (BUG-2884) (#1255)
* test(store): BUG-2884 regression tests — items under a soft-deleted collection
Measured against
|
||
|
|
5aa4bbe319 |
fix(build): give each worktree its own Postgres test port, and refuse to run when it is unreachable (TASK-2708) (#1253)
* fix(build): give each worktree its own Postgres test port, and refuse to run when it is unreachable (TASK-2708) docker-compose.test.yml bound the host port to 5445, so exactly one worktree could run make test-pg at a time. With concurrent worktrees the normal operating mode that produced three incidents in an afternoon: a port-already-allocated collision, a container dying mid-run under concurrent suites, and a stack orphaned by a removed worktree blocking the port for everyone. The worst of the three forged a gate leg: go test exited 2 having executed NO TESTS because the database was unreachable, and exit 2 with zero FAIL lines reads like a pass at a glance. Docker now assigns the host port and the Makefile reads it back with docker compose port. Before running anything the target probes the HOST path the tests will use, from a throwaway container, and refuses with a banner saying no tests executed rather than letting an unreachable database look like a result. If the suite fails and the database is gone afterwards, it says the failures are infrastructure. The compose project name was already per-directory, so teardown never could reach a sibling; the orphan recovery command is now documented where someone looking for it will be. Claude-Session: https://claude.ai/code/session_01HeChkgZVYb3NTgTcckF5KR * docs: record that make test-pg is now safe from concurrent worktrees (TASK-2708) The worktree section is where a reader learns what is safe to run alongside a sibling, so it is where this belongs — including that a privately-started container is no longer needed, and the recovery command for a stack orphaned by a deleted worktree. Claude-Session: https://claude.ai/code/session_01HeChkgZVYb3NTgTcckF5KR * docs(store): the mutation-harness recipe reads the port back instead of hardcoding it (TASK-2708) A paste-ready recipe in a comment is a consumed artifact: it said 5445, and after the ephemeral-port change pasting it would connect to whatever else is on that port, or to nothing. Found by re-running the prose sweep with a path-scoped exclusion — the first pass piped through 'grep -v node_modules', which filters by LINE CONTENT and had silently eaten the hits in files whose matching line mentions node_modules. Claude-Session: https://claude.ai/code/session_01HeChkgZVYb3NTgTcckF5KR * docs(build): say that the banner discriminates, not the exit code (TASK-2708) Measured while building the counterfactual matrix: make collapses every failed recipe to exit 2, so the infrastructure refusals and an ordinary test failure are indistinguishable by status. The banners are the only discriminator, and a reader who assumed otherwise would build automation on a difference that does not exist. Claude-Session: https://claude.ai/code/session_01HeChkgZVYb3NTgTcckF5KR * fix(build): unique compose project, guarded startup, loud teardown failure (TASK-2708) All four from codex round 1, each verified in the recipe before accepting. Compose defaults the project name to the directory BASENAME, so two checkouts sharing a basename share a stack and one down -v tears down the other's database mid-run — the cross-worktree teardown this task exists to prevent, reached through a second door. The project name is now explicit and keyed to the absolute path. My compose comment had claimed the default was already sufficient, in the place the next reader would believe it. up --wait now runs inside the guarded block: a health-check timeout used to abort the recipe before teardown, leaving the stack behind and creating exactly the orphan this task was filed about. A failed teardown is announced with the command to reap the stack instead of being swallowed. It does NOT fail the build: the tests genuinely ran and their status is honest; the leak is a separate fact and is now a loud one. make test-pg-project prints the name so an orphan can be reaped without re-deriving it. Claude-Session: https://claude.ai/code/session_01HeChkgZVYb3NTgTcckF5KR * fix(build): portable readiness probe, safe project derivation, honest recovery docs (TASK-2708) Five from codex round 2, each verified before accepting. The readiness probe used `docker run --network host`, which is Linux-only by default. On Docker Desktop a perfectly healthy database would have been reported unreachable and the target would have REFUSED TO RUN where it used to work — a guard against false greens turned into a false red. It now uses the host's pg_isready when present and falls back to an in-container check, which is weaker but never lies about the platform. The project name interpolated CURDIR into shell command text, so a checkout path containing a quote would have broken the quoting. The shell now reads its own working directory instead. Teardown failures on the three guard exits were silenced by >/dev/null, contradicting the loud-teardown promise those same guards make. Two docs were falsified by my own earlier commit in this branch: CLAUDE.md still told the reader to reap a stack by directory name, and the mutation recipe in the store test omitted -p entirely, which is exactly the same-basename collision the change exists to prevent. Claude-Session: https://claude.ai/code/session_01HeChkgZVYb3NTgTcckF5KR * fix(build): tear the stack down on interrupt; record why a post-run outage is not reported (TASK-2708) Round 3, one finding accepted and one refused. ACCEPTED: Ctrl-C during go test killed the recipe shell before down -v, leaving an orphaned stack — the exact failure this task was filed about. An INT/TERM trap set before up covers startup as well. REFUSED, with the premise checked rather than argued: the reviewer asked for the post-run banner's EXIT_CODE gate to be dropped so a database dying after a passing run is reported. storetest.NewPostgres skips only when the env var is EMPTY; a database that is gone produces t.Fatalf, not a skip. So exit 0 means every Postgres-backed test completed against a live database, and failing the leg because the container stopped afterwards would convert honest greens into reds. Written into the Makefile so the next reviewer does not re-raise it. Claude-Session: https://claude.ai/code/session_01HeChkgZVYb3NTgTcckF5KR * fix(build): probe the host-published port on every platform; interrupt reports teardown honestly (TASK-2708) Round 4, both findings real. The readiness fallback ran 'compose exec pg_isready', which answers whether the server is alive INSIDE the container — a broken host port mapping passes it and the guard is bypassed. Not a rarely-exercised path either: this box has no host pg_isready, so the fallback is the branch that has been running all along. It now reaches back through host.docker.internal from a throwaway container, which is native on Docker Desktop and resolves on Linux via --add-host=...:host-gateway. Verified against a live stack, with a negative control on a port nothing listens on. That is the third version of this probe. --network host was Linux-only and would have falsely refused on Desktop; compose exec was portable but asked a narrower question than the claim it carried. The interrupt trap announced 'stack torn down' unconditionally, so an interrupted run whose teardown failed reported successful cleanup. It now reports what happened and names the command to reap the stack. Claude-Session: https://claude.ai/code/session_01HeChkgZVYb3NTgTcckF5KR |
||
|
|
14cb97593f |
WIP feat(server,store): referent validation for relation values — 6 of 8 doors (TASK-2878) (#1246)
* feat(store): referent resolution for relation values (TASK-2878)
PLAN-2857 U1, first slice: the rule itself, with no door wired to it yet.
`ResolveRelationReferents` canonicalises every `relation` value in a field
map to the target item's ID and reports the ones that cannot be resolved —
same workspace, and the collection the field DECLARES.
WHERE IT LIVES was forced, not chosen. `internal/items` is DB-free by
construction and keeps the shape check only. `internal/server` cannot own
it either: six of the eight coercion doors live there, but the eighth is
`store.migrateFieldsForCopy`, and `store` does not import `server`. Putting
it here is what lets the cross-workspace copy door and the preflight door
reach the SAME function instead of two implementations of one rule — those
two already carry a comment saying they sit in different packages and that
is how they drift unnoticed.
VISIBILITY IS NOT HERE, deliberately. "Can this requester see that item" is
request-scoped and needs the user, role and auth mode; the server layer
adds it via `checkItemVisible`, which already exists as the context-free
predicate for exactly this reason.
NO SLUG FALLBACK, which is a deliberate divergence from `ResolveItem`
(UUID, then ref, then slug). Found by a test failing rather than by
reading: "red" resolved, because it is the slug of the live Red colour. A
relation field's contract is that it stores an item ID; a slug is neither
an ID nor stable, so the same stored value could point elsewhere tomorrow.
Worse, "red" is exactly the free-text value the pre-U2 editor wrote into
these fields, so accepting it makes the corruption this unit exists to
stop indistinguishable from a legitimate write. The client refuses the
same match for the same reason (TASK-2868). Exact-TITLE resolution is U6.
Issues are reported in SCHEMA order, not map order, because the copy
preflight is one of the callers and is specified to be safe to call
repeatedly and return identical results.
Unresolvable values are left EXACTLY as supplied: the caller quotes them
back, and a half-canonicalised map would make a drop report lie about what
the source held.
Verified rather than asserted: both lookups exclude soft-deleted rows
(`ResolveItem` by contrast with `ResolveItemIncludeDeleted`; `GetItem` via
`getItemScanQ`, which appends `AND i.deleted_at IS NULL`). That is what
keeps "target was deleted" distinguishable from "never resolved" — the
read half U2 shipped.
* feat(server): refuse unresolvable relation values at the four write doors (TASK-2878)
PLAN-2857 U1, second slice: the doors that take CALLER-SUPPLIED field
values now refuse a relation value that does not name a live item in the
declared target collection — create, update (full fields), update
(fields_patch), and bulk update.
The server half adds the one thing the store resolver deliberately does
not: visibility. It folds into the SAME `not_found` reason rather than
getting its own, because "that item exists but you may not see it" is an
existence oracle, and this codebase has a standing rule against handing
one out.
Ordering at every door is after the shape check and after coercion, so one
bad value produces one error rather than two describing it differently,
and so the value is in its final form when it is resolved.
`fields_patch` examines only the keys the patch carries — the resolver
skips absent keys — so an unresolvable value already stored on an item is
not re-litigated by an update that does not touch it. That mirrors the
undeclared-key rule immediately above it, and it is what stops this
turning every edit of a legacy item into a failure.
Refusals use the ORDINARY `validation_error` shape with no new details
key. The MCP stdio transport classifies errors by matching CLI stderr
prose, so a structured field it cannot see would help nobody there, and a
new error shape is a contract change for every client.
Existing suites unchanged: internal/server ok (224.5s), internal/store ok
(258.0s), internal/items ok. Nothing in the tree was writing a bogus
relation value through these doors, which is what made this slice safe to
land before the per-door pins.
* feat(store): one migrate decision for all four carrying doors (TASK-2878)
PLAN-2857 U1, third slice, on the lead's refined ruling: PROVENANCE
decides, not which door you came through.
* SUPPLIED (an explicit `--field` override on a move or copy) is a write
like any other, so an unresolvable value REFUSES.
* CARRIED (everything the source item already held) was asserted by
nobody. `internal/items` has accepted any string for a relation all
along, so most stored values are legacy — refusing them would make
those items unmovable and uncopyable. Dropped and REPORTED instead.
And carried values are not all alike, which is the refinement that keeps
this from being one rule wearing four coats:
* WITHIN a workspace (move, bulk move) the targets are still here, so a
valid relation SURVIVES the move and only an unresolvable one is
dropped, through the `dropped_fields` channel BUG-2674 established.
* ACROSS workspaces (copy, and its preflight) every carried relation is
dropped WITHOUT a lookup: the value names a source-workspace row and
v1 excludes cross-workspace targets, so no amount of resolving in the
destination changes what it means. Reported as `referent_not_portable`
— the same reason `github_pr` uses, because it is the same fact about
the same kind of value.
`MigrateRelationReferents` is one function because the four doors sharing
it is the point, not tidiness: the preflight lives in `internal/server`
and the copy in `internal/store`, and the code already carries a comment
saying those two sit in different packages and that is how they drift
unnoticed. A preflight that says "carried" while the copy drops is one
request answered two ways.
Tests drive both provenances against both modes, because the same bad
value must be a drop when carried and a refusal when supplied — a suite
that only drove carried values would pass against a build that never
refuses anything.
* feat(server): the two same-workspace migrate doors resolve and report (TASK-2878)
PLAN-2857 U1: `handleMoveItem` and `bulkMoveCollection` now take their
relation decision from `store.MigrateRelationReferents` — the same
function the two copy doors will call, which is the point of it existing.
Within a workspace the targets are still present, so a correctly-related
item KEEPS its relation across a move; only an unresolvable value is
dropped, and it joins the `dropped_fields` report BUG-2674 established
rather than failing the move. Refusing carried values here would make
every legacy item permanently unmovable, and `internal/items` has accepted
any string for a relation all along, so "legacy" is most of them.
The bulk path carries no per-field overrides — only `status` — so every
relation value reaching it is CARRIED and nothing there can refuse. It
passes nil for `supplied` to say so, and keeps the refusal branch: it is
unreachable today and stops being a silent no-op the day that path grows
overrides.
internal/server ok (270.8s), internal/store ok.
KNOWN GAP, recorded rather than half-built: the two CROSS-workspace doors
are not wired yet, and the reason is a real constraint rather than
running out of road. `migrateCopyFields` is called from
`copyItemAcrossWorkspacesTx` with a transaction already open
(`s.db.Begin()` at the top of that function), so resolving a SUPPLIED
override there would issue POOL reads while holding a tx — the deadlock
shape this repo keeps a deterministic test for. The carried half needs no
lookup at all and is safe; the supplied half needs
`GetCollectionBySlugQ` / `GetItemByRefQ` so the resolver can run on the
tx's connection, which is exactly the `...Q` convention the store already
uses (`GetItemQ`, `getCollectionInWorkspaceTx`, `uniqueSlugQ`). Adding
those two is the remaining work, and it is what makes one function
genuinely serve all four doors.
* refactor(store): thread a Queryer through referent resolution (TASK-2878)
Preparation for the two cross-workspace copy doors, landed on its own
because it is independently correct and the doors are not.
`migrateCopyFields` runs inside `copyItemAcrossWorkspacesTx`, which opens
a transaction as its second statement. A resolver reading from the POOL
there would issue pool reads while holding a tx — the deadlock this repo
keeps a deterministic test for. So `ResolveRelationReferentsQ` and
`MigrateRelationReferentsQ` take the executor, following the store's own
convention (`GetItemQ`, `uniqueSlugQ`, `getCollectionInWorkspaceTx`); the
pool-backed names stay as one-line shims for the six wired doors.
Two small read helpers come with it. `collectionIDBySlugQ` returns the ID
only — the referent check compares `item.CollectionID`, and the full model
would pull in per-collection counts nothing here uses. `itemByRefQ` keeps
`GetItemByRef`'s fallback to a bare item-number lookup, because a relation
written as COLO-3 must keep resolving after its target collection is
renamed, which is exactly what BUG-2873 made possible.
internal/store ok (341.7s), vet and gofmt clean.
WHY THE COPY DOORS ARE NOT IN THIS COMMIT. They were written and building,
and I reverted them. Team CONVE-29 and the lead's condition both say the
copy pair lands WITH its pin — one case driving BOTH doors, asserting
identical drop-and-report for a carried relation and refusal for a
supplied override — and I measured 58.9% context against a 65% ceiling,
which is not enough for that pin plus the 270s server suite plus the
commit. Landing the behaviour change unpinned would have been worse than
landing nothing: the preflight and the store copy are the pair the code
already warns will drift unnoticed, so they are the last place to accept
an untested agreement.
The design is complete and on the trail: derive the carry mode from the
existing `items.MigrateScope` rather than a second flag, pass `tx` on the
store side and the pool on the preflight side, refusals through the copy's
existing validation-error channel, drops appended to `migrated.Dropped`.
* feat(copy): the two cross-workspace doors resolve referents, with their pin (TASK-2878)
PLAN-2857 U1, doors seven and eight. `migrateCopyFields` and
`handleCopyItemPreflight` now take their relation decision from
`store.MigrateRelationReferents` — the same function the four write doors
and the two move doors already call, which is the entire reason it exists.
The defect this closes: MigrateFields matches on key and TYPE, so a
same-named `relation` field carried a SOURCE-workspace item id across the
boundary and the preflight reported it as a clean carry. What landed in
workspace B was a value naming a row in workspace A — unrenderable, and
indistinguishable on read from a legitimate reference.
Provenance decides, as at the move doors. A CARRIED value on a
cross-workspace copy is dropped without a lookup (no id from A can mean
anything in B) and reported through the `dropped_fields` channel BUG-2674
established; a SUPPLIED override is an ordinary write and an unresolvable
one is refused, 400 validation_error on both doors, rendered by the same
`store.RelationIssuesMessage` so one refusal cannot acquire two phrasings.
`internal/server`'s `relationIssuesMessage` now delegates to it: the eighth
door refuses from inside `store`, so the sentence had to be reachable there.
Two things are threaded rather than re-derived, and both are load-bearing:
- The TRANSACTION, not the pool. `migrateCopyFields` becomes a method
taking a Queryer, and `copyItemAcrossWorkspacesTx` passes its `tx`. That
function has held a transaction since its second statement, so a pool read
from inside it can wait for a free connection while every pooled
connection is blocked on this transaction's locks — the starvation shape
BUG-2409 fixed for the attachment planner and this repo keeps a
deterministic test for. This is what the day-70 handoff named as the
reason these two doors were not wired with the other six.
- The MODE comes from the `scope` MigrateFields was already given, not from
a second boundary test. Two independent answers to "is this crossing a
workspace" is how one request gets migrated one way and validated the
other, and this path also serves a copy whose target IS the source
workspace, where relations resolve and survive exactly as on a move.
The destination workspace id is the resolution scope: a supplied override
is a write into B and must name something that exists there.
THE PIN, and why it is not a per-door table. These two doors sit in
different PACKAGES and the code at both sites says so is how they drift
unnoticed. A table with a row per door can be fully green while the two
disagree about one request, which is the defect rather than a gap in
coverage of it. So every case sends ONE body to BOTH endpoints:
- carried relation — must drop on both, and the preflight must say
referent_not_portable rather than the generic no_target_field, which is
false here (the destination DOES declare the key, so that answer sends
the reader to fix a schema that is fine);
- supplied + unresolvable — both refuse, same status, same code, both name
the offending field and value, and nothing is written;
- supplied + resolvable — the positive control, supplied as a REF so
resolution is visible in the result. Without it the first two legs are
equally consistent with "relations always fail".
Negative controls run, all three mutants BUILD-CHECKED first (a
non-compiling mutant produces no `--- FAIL` lines and reads as survived):
both doors unwired = DETECTED; preflight unwired alone = DETECTED; store
unwired alone = DETECTED. Each single-door mutant failing is the pin's
whole claim — neither door can be wired without the other.
CONVE-23 sweep: the preflight's LIMITATION comment said this gap belonged
in MigrateFields "for both callers at once". That is now false in its
prescription as well as its premise — `internal/items` is DB-free by
construction and cannot ask whether a string names a live item — so the
comment records where the fix actually went and what of it remains open
(`computed`, `terminal_options`, `unique_scope`).
Gates: internal/server ok 170.0s · internal/store ok 296.1s · internal/items
ok · go vet clean · gofmt clean · make lint 0 issues.
* test(server): the per-door x per-provenance table, and the defect it found (TASK-2878)
PLAN-2857 U1. `internal/store` already tests the resolver exhaustively, but
those tests call it DIRECTLY: they vouch for the component and say nothing
about whether any door is bound to it. A door that never calls the resolver
passes every one of them. This table is the binding claim — one leg per
(door, provenance) pair, driven through the handler a client reaches.
PROVENANCE IS THE SECOND AXIS BECAUSE THE ANSWER DEPENDS ON IT, not for
symmetry. A SUPPLIED value is the caller's assertion and an unresolvable one
is refused; a CARRIED value was asserted by nobody, and refusing it makes
legacy items un-updatable, un-movable and un-copyable — the failure this
unit would otherwise CAUSE while fixing another. Which provenance a door
sees is a property OF THE DOOR, and getting it wrong is invisible until a
legacy item meets it.
THE TABLE FOUND ONE, ON ITS FIRST RUN. `bulkFieldUpdate` merges the item's
STORED fields blob with the caller's `changes` before validating, and the
resolver was pointed at the MERGED map — so a bulk status move or
set-priority re-litigated every stored relation value and REFUSED the item.
An item carrying a legacy relation value had its status and priority frozen
by a field the operation never mentioned. Fixed the way the fields_patch
door already handles it: resolve only the keys the operation CHANGES, read
out of the coerced map so the value is final, written back so a supplied ref
is still canonicalised. Verified by prediction before the run and by the leg
failing against the unfixed code.
THE DISPATCH MUTANT, which is what makes the table's coverage a measurement
rather than a hope. Wire ONLY the two `extractParentLink` doors (update
fields, update fields_patch) and neuter the other six by swapping their
field-map argument for an empty one — types unchanged, so the mutant
compiles and its verdict means something:
build OK · failing legs: create (both), move (both), bulk move,
bulk update supplied, copy, preflight. Update and update-patch pass.
Exactly the six unwired doors, and only those. Then each door alone, eight
runs: every one detected by its own legs and no other door's. That is the
claim the table exists to make — each leg reaches its OWN door rather than
being satisfied by a neighbour's check.
ONE DOOR NEEDS A WEAKER INSTRUMENT, AND THE TEST SAYS SO. No bulk op puts a
relation key into `changes` — `op` is a closed list and the only field
values any of them set are `status` and `priority` — so that door's SUPPLIED
branch is unreachable from outside. The first mutant run proved it: unwiring
door 4 alone left every black-box leg passing. It gets a direct-call leg,
labelled as vouching for the FUNCTION and not for a binding that does not
exist yet, and kept for the same reason `bulkMoveCollection`'s refusal
branch is kept: the day the bulk path grows per-field overrides, the branch
must already refuse rather than be a silent no-op nobody notices is missing.
Every refusing leg has a resolvable counterpart. Without them the table is
equally consistent with a build that refuses every relation value.
Gates: internal/server ok 156.1s · go vet clean · gofmt clean · make lint 0
issues.
* feat(mcp): ToolSurfaceVersion 0.29, and the drop-reason renderer it exposed (TASK-2878)
PLAN-2857 U1. The bump, its documentation sweep, and the consumer this
change turned from a rare wart into a routine one.
THE BUMP, at 0.29 rather than 0.28. Rebasing onto main found
|
||
|
|
a0e1bbf53c |
feat(store): composite FK making a reminder's workspace agree with its item's (IDEA-2883) (#1252)
* feat(store): composite FK making a reminder's workspace agree with its item's (IDEA-2883) item_reminders.workspace_id is denormalized from the item, and until now nothing but a read-side predicate repeated at five call sites stopped the two from disagreeing — a class IDEA-2641 closed one read at a time across two codex rounds. The composite foreign key makes the disagreeing row unrepresentable; the predicates stay as belt on top of braces. SQLite rebuilds the table (it cannot ADD a constraint), Postgres alters in place. Both add a unique index on items(id, workspace_id) as the referenceable parent key, and both repair before constraining — a pre-existing disagreeing row would otherwise fail a deployment's startup. Such a row is DROPPED rather than repaired: every read predicate already filters it, so it can never fire, and pointing it back at its item would resurrect it and deliver a notification for an instant long past. No ON UPDATE CASCADE: items.workspace_id is never updated anywhere, and a future feature that moves an item between workspaces should be refused here rather than silently rewriting its reminders. Claude-Session: https://claude.ai/code/session_01HeChkgZVYb3NTgTcckF5KR * test(store): build the disagreeing-reminder fixture with enforcement suspended (IDEA-2883) The IDEA-2641 test that proves a workspace-mismatched reminder is inert everywhere could no longer construct its row: the composite foreign key refuses it, on both drivers. Kept rather than deleted, for a specific reason rather than a general nod to defence in depth. The constraint protects rows written through a connection enforcing it, and SQLite's enforcement is a per-connection pragma that table-rebuild migrations legitimately turn off — so a row can still arrive from a restored pre-086 backup, from an operator who dropped the constraint, or through a future rebuild's window. Those five reads are what keep it inert, and this is the only test that says so. Its comment said the table has nothing that forbids such a row. That is now false and it says so. Claude-Session: https://claude.ai/code/session_01HeChkgZVYb3NTgTcckF5KR * docs(store): the table now forbids the disagreement the predicate compensated for (IDEA-2883) reminderFireable's note said the table has an FK to the item and no constraint tying the two columns. That was the premise for the predicate and it is now false. It also says why the predicate stays: enforcement binds to the connection doing the write, and SQLite's is a per-connection pragma that table-rebuild migrations legitimately turn off. Claude-Session: https://claude.ai/code/session_01HeChkgZVYb3NTgTcckF5KR * fix(store): drop EVERY legacy item_id foreign key, not just the first (IDEA-2883) Self-review: plpgsql's SELECT ... INTO takes the first row and does not error on several, so a table carrying two single-column FKs on item_id would keep one alongside the composite key while the migration reported success. A loop costs two lines. Nothing creates a duplicate today, which is exactly why the branch needed a test rather than a comment. Claude-Session: https://claude.ai/code/session_01HeChkgZVYb3NTgTcckF5KR |
||
|
|
5c4fa22999 |
fix(store): collection slug races — allocate under the row lock, serialize CreateCollection with rename (TASK-2885) (#1249)
* fix(store): allocate collection slugs under the row lock; serialize CreateCollection with rename (TASK-2885) IDEA-2874: UpdateCollection allocated its new slug on the pool before its transaction opened, so two renames deriving the same base both chose it and the loser failed on the UNIQUE index. Allocation now runs on the tx after the workspace lock and the FOR UPDATE re-read, matching documents.go and items.go. BUG-2875: CreateCollection took no lock, so a create could land between a rename's FOR UPDATE scan of existing siblings and its commit. It now opens a transaction, takes the workspace advisory lock first (the order item-create and rename share), and allocates its slug under it. Population: every collection rename and every collection create; the import door inserts into a workspace nobody else holds yet and is out by construction. * test(store): pool-I/O guard for the collection slug writers under the workspace lock (TASK-2885) |
||
|
|
02846a6785 |
fix(cli): the session's registered agent is the name its writes carry (BUG-2882) (#1248)
* fix(cli): the session's registered agent is the name its writes carry (BUG-2882) Two seats booted under one name; one re-registered under the right one with `pad session register --agent`; every write it made afterwards still carried the wrong name. The registry row and $PAD_AGENT were two self-declarations of "the same value" — the row could be rewritten, the environment could not, and nothing reconciled them or warned. Night 10 read both live seats inverted against each other. ResolveAgentName now consults the registry record for the session that owns this process FIRST: a stat-and-read of one file, no MkdirAll, no lock, ignored when malformed, legacy, or carrying a different process-start token (pid reuse). A non-empty registered name wins over .pad.toml and $PAD_AGENT; an anonymous row leaves an environment name in force. `pad session register` without --agent keeps the current name, as before, because its default is the resolver. Help text, the record's doc comment and README's precedence list say what is now true. Test: register as rook under PAD_AGENT=wren and a .pad.toml name → rook; default re-register → rook; anonymous → the .pad.toml name; a record for this pid with another process's start token → ignored. The registry-step mutant fails the first two assertions. Fixes BUG-2882 * fix(cli): a registry record names this session only when it is verifiably this session's, and the identity tests stop reading the real registry Codex round 1 on #1248. (1) registeredAgentForThisSession compared process-start tokens only when both sides had one, so a stale row with no token under a reused pid — a dead session's — would have named a live one. Fail closed: when this process can read a token, the record must carry the same one; and the record must pass the same OwnerLiveness verdict `pad session list` applies. The token-less case is now a test row, with a positive control after it. (2) The pre-existing resolver and header tests cleared PAD_AGENT/CLAUDECODE but not HOME or the session-pid variables, so run inside a registered seat they would have read that seat's row as the resolver's first answer. They now run from a scratch HOME with no session identity. Refs BUG-2882 * fix(cli): a registry row names this session only if its owner is this process or an ancestor, where that can be checked Codex round 2 on #1248. (1) TestPushItemSendsResolvedAgentHeader was the one identity test round 1's hermeticity fix missed; it now isolates HOME and the session env like its neighbours. (2) The registry step accepted a row whose owner pid was alive and token-matched but NOT this process or an ancestor — a misconfigured CLAUDE_PID pointing at a sibling session could borrow that session's name. Refused where the platform can walk ancestry. Not gated on PIDVerified: CaptureSessionOwner records "cannot check" and "checked and wrong" as the same false, and the flag alone would have disabled the step on every non-Linux platform. The check is re-run and only a checked-and-wrong answer refuses. Test uses a live non-ancestor child as the claimed owner; the refusal-dropped mutant lets it name us. Refs BUG-2882 * docs(cli): session list help — a row's name outranks PAD_AGENT; change it with --agent, not by re-registering (BUG-2882, codex round 3) |
||
|
|
b437cc582d |
feat: item reminders — the fire-at-an-instant primitive, and one overdue rule for all four surfaces (IDEA-2641, closes #1010) (#1244)
* feat(store): item reminders — the fire-at-an-instant primitive (IDEA-2641) Adds the storage, the scheduler tick, and the canonical event for one-shot item reminders (GitHub #1010). Nothing in Pad acted at a target time before this: a due_date makes an item show up as overdue once somebody asks the dashboard, so "revisit TASK-X on the 1st" had to live in an external cron. A TABLE, NOT A SCHEMA-FIELD ANNOTATION. The design sketch proposed marking schema date fields with a `reminds: true` key on models.FieldDef; recon overturned it. Such a key does not survive an ordinary collection edit, two independent ways: the web editor destructures each field into an EditableField and rebuilds a fresh definition key-by-key on save, so unknown keys are dropped (`pattern` and `unique_scope` survive only because two lines were hand-added for them), and models.CollectionSchema has fixed fields with no catch-all, so any Go unmarshal+marshal round-trip strips unknown properties — the hazard retargetRelationFieldsTx mutates raw JSON to avoid. Both failures are silent and both disarm a whole collection's reminders at once. It is the same defect class that moved traits out of the schema column in TASK-2657. The table also gives the lifecycle a home. A reminder is armed, then fired, then acknowledged, and a re-arm returns it to armed — per-reminder state a field definition has nowhere to keep. remind_at is an RFC3339 UTC instant, deliberately not a `date` schema value: those admit both YYYY-MM-DD and full RFC3339 and are compared against the SERVER'S LOCAL calendar day. A fire-at time cannot carry that ambiguity. The remaining timezone question for due_date is filed separately. Firing is one transaction per reminder carrying BOTH the fired_at write and the outbox insert. That pairing is the point: a fired_at committed without its event is a reminder that silently notifies nobody and can never be retried, because the row has left the armed set; an event without fired_at fires every tick forever. The UPDATE's own `fired_at IS NULL` predicate is the arbiter, so two instances ticking at once produce exactly one winner. item.reminder_due is admitted to the closed events/1 set as v1.2, with a new PayloadReminder family and no SSE name. The subject is the REMINDER, not the item: two reminders can be armed on one item, so an item-subject event could not say which fired, and the reminder id is what an acknowledgement addresses. A new payload family rather than reusing the item snapshot for the same reason — a snapshot would validate and still not answer the only question the event exists to answer. No SSE name in v1 because the poll surface is the contract; adding one later is additive, removing one is not. Ack is explicit and nothing else acks. An item reaching a terminal status deliberately does NOT ack: that would make every status write a reminder mutation, and it would silently consume a reminder set to fire after the work was done. * feat(server): reminder surfaces, and one shared overdue rule for all four Second half of IDEA-2641: the HTTP surface, the scheduler tick's wiring, and the fix for the finding that justified the unit — `ready` / `next` did no date handling at all. OVERDUE NOW HAS ONE IMPLEMENTATION. It used to live inline in the dashboard's attention loop, which meant `pad project stale` inherited it (it filters that very list) and the recommendation surface never saw it. So a deadline reached the two surfaces that REPORT on work and never the one an agent PULLS from. overdue.go is now the only place that decides, and all four call it. Two behaviour changes fall out, both deliberate: - An overdue item bypasses the orphan branch's high/critical priority gate. That gate was where a deadline quietly stopped: a low-priority item three weeks late was reported by `stale` and never suggested by `next`. - Overdue sorts above in-progress. The list is capped at three, so a rank below in-progress would not merely order the deadline lower — on any workspace with three things in flight it would keep an overdue item off the surface entirely, which is indistinguishable from not shipping this. The server-local-today comparison is UNCHANGED and known to be wrong for multi-timezone deployments; it is filed as its own item with the cloud case stated. Changing what "overdue" means on every existing instance inside a change about where the rule LIVES is the kind of behaviour change nobody reviews. Fired reminders reach `next` / `ready` two ways, from one filtered list: PendingReminders is the addressable form (it carries the id an ack needs), and a prepended suggestion is the rendered form. They are prepended AFTER the cap rather than entered as ranking candidates — a reminder is not a task competing on priority, and whether it appeared should not depend on how busy the workspace is. Terminal-item reminders are FILTERED from the surface, never acked. Acking on terminal status would couple every status write to reminder state and would consume a reminder armed to fire after the work was done. The row stays exactly as the user left it; the distinction is observable, and asserted. Three guard tests caught this change and each was answered rather than silenced: - The request-body reader guard was right: the handlers now go through decodeJSON, inheriting the NUL refusal and the size cap. - The canonical-events guard was right: item.reminder_due is admitted to the duplicated contract table as SPEC-3 v1.7, with the reminder subject kind and the new payload family. SPEC-3's own text owes the same amendment. - The NUL census asked for a decision on eight new columns. None carries caller text: ids and FKs are server-generated, four are the server clock, and remind_at is now re-parsed and re-formatted in the STORE as well as at the edge — so the stored value is always machine-produced from a parsed time and no caller bytes reach the column. The doc comment that used to say "the caller normalizes" protected nothing. Regenerating the baseline also found that GEN_NUL_BASELINE=1, which the test's own instructions name, was never implemented — the flag did nothing, so the documented path was hand-editing the file. Implemented, so the next reader gets the mechanism the instructions promise. * test(reminders): the lifecycle, the four surfaces, and 22 killed mutants Every test here was designed against a specific mutation and the mutation was RUN. A green suite proves nothing about a suite nobody tried to break, and three of the mutants I first wrote were not experiments at all. Store (10 mutants, all killed): candidate predicate <= flipped to >=; the event emission lifted out of the fire transaction; the fire UPDATE's `fired_at IS NULL` arbiter removed; the RowsAffected check ignored; re-arm clearing fired_at but not acked_at; ack losing `fired_at IS NOT NULL`; the poll surface losing `acked_at IS NULL`; normalizeRemindAt no longer refusing; it dropping .UTC(); GetReminder losing its workspace scope. Surfaces (12, all killed): the priority gate no longer bypassing on overdue; the sort no longer ranking overdue first; attention leaving the shared helper; the reason losing its OVERDUE prefix; the comparison flipped to >; terminal items no longer skipped; terminal reminders no longer filtered; the filter ACKING instead of hiding; reminders appended instead of prepended; the tick running on a far-future clock; ack answering 200 for an unfired reminder; parseRemindAt accepting a bare date. THREE MUTANTS DID NOT COUNT ON THE FIRST PASS and were rewritten. Two failed to compile (`if false` orphaned a variable; deleting a parse orphaned an import) and one had an anchor matching two call sites. A non-compiling mutant emits zero FAIL lines and reads exactly like a surviving one — it invents a hole that is not there — so the harness reports BUILD-FAIL and ANCHOR-BAD as outcomes distinct from SURVIVED. It also restores files from an in-memory copy rather than `git checkout`, which would delete uncommitted work in the tree. ONE MUTANT GENUINELY SURVIVED and the test was at fault, not the mutant: appending rather than prepending reminder suggestions was undetectable because the fixture had a single item, so the reminder sat at index 0 either way. The fixture now fills the three-item cap with in-progress work, where an appended reminder lands fourth and vanishes. Faithful mutant, weak test — checked in that order. The same lesson shapes the four-surface fixture: it is a LOW-priority open orphan, because that is the case the old code handled worst. A high-priority task would have made the ready/next leg pass against the unfixed tree, which is a green that measures nothing. Negative controls throughout: a future deadline is not overdue and does not reach the gate bypass; a tick with nothing due fires nothing; a completed item is neither overdue nor suggested. Without them a helper that reported every date, or a tick that fired everything, would satisfy every positive leg. The lead's pin is asserted in both directions: a fired reminder on a done item is ABSENT from the surface and PRESENT and still unacknowledged in the table. Asserting only the absence would pass against an implementation that consumed the row, which is the behaviour the pin exists to forbid. * feat(mcp): pad_item.remind + ack-reminder, ToolSurfaceVersion 0.28 An agent that can RECEIVE a reminder but not set one has half the primitive. The poll surface is pad_project.next / ready, both long exposed, so reminders already reached agents — what was missing is the other half: deferring a piece of work is exactly the moment an agent knows when it wants to be asked again, and it had no way to say so. Two additive actions, two optional params. Nothing existing moved, so a v0.27 consumer enumerating neither is unaffected — the v0.13 / v0.11 / v0.8 disposition, which likewise wired existing CLI verbs onto the catalog. remind_at REFUSES a bare date rather than reading it as midnight. Worth stating because the `date` schema type accepts YYYY-MM-DD and a caller will reasonably try it here: a bare date names a 24-hour span, and choosing an hour inside it would fire at a time nobody picked. Re-arm and disarm stay CLI-only. Both address a reminder by an id the agent would have to list first, and no listing action exists on this surface — a door with no handle. Adding them later is additive. Five guards had to be taught, and each was answered on its merits rather than excluded: the HTTP parity test (route mappers added, so the actions work on the remote transport rather than being advertised and unrouted), the read-only catalog's cmdhelp fixture and expected cmdPath map, the field- conflict classifier (remind_at / reminder_id are NOT field writers — a reminder is a row in its own table addressed by its own id, so listing them as classified sources would have pointed detectFieldConflicts at something that is not a field source), and the instructions.md / README action tables. That machinery is why the version bump is safe to make now, and it earned its keep on this change: every one of the five failed on the first build after the catalog entry landed. CONVE-23 sweep for prose this falsifies: - SPEC-3 (DOC-2653) amended to v1.7 in the room, recording item.reminder_due with its new subject kind and payload family — the first canonical event with no user mutation behind it, since a scheduler tick produces it. - CLAUDE.md gains the reminder routes, the CLI verbs, and the v0.28 entry. It was also stale at 0.26 with NO v0.27 entry at all: the 0.27 unit swept instructions.md and README.md and missed this file. Both added. - skills/pad/SKILL.md gains the verbs and a routing entry, including the two things an agent will get wrong — the time is an instant, so ask for a time of day rather than picking one, and finishing the item does not acknowledge the reminder. * fix(reminders): codex round 1 — four findings, all real, all with a pin Round 1 found four defects and refuted none of them. Each fix carries a test that fails against the code as it was, and each of those was mutation-checked. **P1 — pending reminders bypassed item-level visibility.** Every other dashboard section reads `allItems`, which the store already scoped to the caller's collections AND their granted item ids. The pending-reminder list is a direct workspace-wide query and inherited none of that, so a guest holding a grant on ONE item could read the refs and titles of every other item in the collection through its reminders — an item-level leak wearing a notification's clothes. Now filtered with the same `isItemVisibleToGuest` call the sibling sections use. The test's two items share a COLLECTION on purpose: a collection-level filter was already applied, so separate collections would have made it pass against the unfixed code. **P1 — soft-deleted items could starve the queue permanently.** Candidate selection ignored `deleted_at`, and `fireOneReminder` rolls back when it finds the item gone — which leaves the reminder ARMED and therefore a candidate again on the next pass. Candidates are ordered oldest-first and bounded by a limit, so enough archived reminders fill every batch and no live reminder ever fires. Silent, too: the tick reports zero fired and looks idle. Excluded in the candidate query rather than skipped downstream, so those rows never occupy a slot; the reminders themselves are kept, so restoring an item restores its reminder with it — asserted, because a fix that reaped them would pass the starvation test alone. **P2 — the pass stopped at the first failing reminder.** The per-reminder transaction exists precisely so one unfireable row cannot hold back the rest, and `return fired, err` made that comment false — with candidates oldest-first, one persistently broken old reminder blocks every newer one forever. Now continues and joins the errors, so a pass that fired seven and failed three reports both halves rather than reading as clean. The loop is split behind an injected seam because a real mid-transaction failure is not reachable from outside: the database refuses the corrupt rows that would cause one (verified — invalid JSON in items.fields is rejected by the schema). **P2 — suggestions dropped the reminder id.** The docs tell an agent to acknowledge what it sees in next/ready, and the payload carried no handle: a stateless poller could read the reminder and had no way to retire it, so it would be shown the same item forever. `DashboardSuggestion` now carries `reminder_id` (omitempty), `pad project next` prints the exact ack command, and the test acks with the id the surface handed out rather than merely checking the field is populated — a wrong-but-present id satisfies equality with itself. Four mutants, four killed; one was rewritten first because its anchor matched two call sites and was therefore not an experiment. * fix(reminders): codex round 2 — four findings, all real **`--rearm` was unusable.** `ExactArgs(1)` forced an item ref that the rearm branch then ignored, so the flag could not be reached without supplying a ref that was silently discarded. Now `MaximumNArgs(1)`, with each mode checked explicitly: a ref is required to arm, and a ref supplied ALONGSIDE `--rearm` is refused rather than ignored — it names an item the reminder may not even belong to, and quietly dropping it is how a user learns nothing about the reminder they just moved. **`unremind --format json` emitted plain text**, breaking the parseable-output contract every sibling command honours. **The MCP `ref` param did not list `remind`.** Agents read that flat description to decide what to send, so an action missing from it is an invalid call waiting to happen. It now also says what `ack-reminder` takes instead, and why: a reminder is addressed by its own id because an item can carry several. **Fractional seconds fired early.** `time.Parse` accepts `09:00:00.900Z` and `Format(RFC3339)` drops the fraction, so it was stored as `09:00:00Z` and fired 900ms BEFORE the moment the caller named — silently, having rewritten their value on the way in. Seconds are genuinely the stored resolution (the column is compared as a string against a whole-second clock, and the tick runs every 30s), so the only question was which way to resolve it, and truncation resolved it the wrong way. `NormalizeInstant` now rounds UP: at most a second of lateness, in exchange for a guarantee that can be stated — a reminder never fires before the instant it was set for. Late is a reminder; early is a wrong answer. Whole seconds round-trip exactly, which is asserted, because an implementation that added a second unconditionally would otherwise pass. Three mutants for this round, three killed (round-up→truncate, round-up→unconditional-add, MaximumNArgs→ExactArgs). Thirty across the unit. Two fixes carry no dedicated test and it is worth being explicit rather than implying coverage: the `--format json` branch on `unremind` is a one-line output change with no server-free way to drive it, and the MCP `ref` description is prose the drift tests do not read — they assert an action is DOCUMENTED, not that a param's sentence lists it. * docs(reminders): the ack id is on the surface an agent polls, not only on the arm response CONVE-23 follow-through on the round-1 fix. Both agent-facing docs told a caller to acknowledge a reminder with the id "returned when you armed it" — true, and useless to the caller that matters: a poller reading next/ready never armed anything. The suggestion now carries reminder_id and `pad project next` prints the exact ack command, so the docs say that instead. The prose was written before the fix existed, which is exactly the case CONVE-23 is about: a change that makes an instruction stale without touching the file the instruction lives in. * test(reminders): bind the tick LOOP to the work, not just the pass (CONVE-19) Every other test in this file calls runReminderTick directly. That vouches for the component and says nothing about whether anything ever calls it — a tick that is never started is indistinguishable, from those tests, from one that is. It is the convention's exact case, and the failure I recorded on my own identity doc three times in one unit: I test the component and not the binding. Driven through the injectable tick channel so the assertion pins a SPECIFIC pass instead of racing a 30-second ticker, and polled to a bounded deadline so a loop that never runs FAILS rather than hanging the suite. Mutant: drop `s.runReminderTick()` from the select and this goes red while every direct-call test stays green. Killed. The idempotence leg exists because a second Start spawning a second loop would leave one running after Stop, making the BUG-842 drain invariant false for this sweeper specifically — the one property a copied lifecycle is most likely to get right by accident and least likely to be checked. The cmd/pad call site (cmd_server.go, alongside StartTokenReaper) stays verified by inspection: a source-scanning guard for it would be an instrument asserting facts about source, which is code with an adversary and not worth it for one line that sits in the middle of five identical neighbours. * fix(reminders): codex round 3 — a deferred reminder fired anyway, and the poll surface was unbounded **A re-arm mid-pass did not stop the fire.** The candidate scan selects an id; before the UPDATE runs, a `--rearm` can move that reminder into the future. Re-arm clears `fired_at`, so a predicate checking only `fired_at IS NULL` still matched — the pass fired a reminder the user had just deferred and emitted its event. The re-arm cannot undo that: it can clear the mark, but the event is already on the outbox and at-least-once means a consumer has seen it. The fire UPDATE now revalidates `remind_at <= nowTS` against the SAME nowTS the candidate scan used. Same-value deliberately: the arbiter and the scan must agree about when this pass is, or a reminder could pass one and fail the other for no reason but clock drift inside a single pass. **The poll surface was unbounded.** Every fired-and-unacknowledged reminder was loaded and turned into a suggestion prepended to a list that is otherwise capped at three, so a workspace with five hundred unacknowledged reminders returned five hundred suggestions — in the dashboard response, the hottest read in the product, growing until somebody acknowledged them. Two bounds, because they are two different guarantees: the query takes a window (default 50, oldest-fired first, so it holds what has waited longest), and the prepended suggestions are capped at 5 so `suggested_next` stays a recommendation rather than a second inbox. The full set stays addressable in `pending_reminders`. Truncation is REPORTED as a boolean, not a count. A count would have to be post-visibility-filter to be true for the caller reading it, and the store cannot compute that — the filter runs per item, above. "There are more than you can see here" is the strongest claim the data supports, so it is the one made. Four mutants; two killed outright, two survived and were run down under CONVE-28: - **Uncapped suggestions survived because the fixture had ONE reminder** — capped and uncapped are the same list at n=1. That is the SECOND time a single-item fixture hid a count-or-order property in this file. Fixture now arms eight; it also asserts all eight remain in `pending_reminders`, so the cap is pinned to the recommendation and not to the data. - **Removing the SQL LIMIT survived, correctly, and the test comment now says so.** The Go slice cap bounds the PAYLOAD; the SQL LIMIT bounds the DATABASE'S work. Only the first is observable at this level — with the LIMIT gone the response is still bounded, while the query silently goes back to materialising every pending row before discarding most of them. That is a memory and I/O property with no assertion available here, so it is stated as a coverage boundary rather than papered over with a green that would not have measured it. * docs(reminders): the fire predicate arbitrates against two actors, not one CONVE-23 inside the file the round-3 fix touched. The comment described the UPDATE as an arbiter for concurrent TICKS, which is what it was written for and is why I did not re-read it when asked whether a user edit could race the pass. It now says what it actually defends against, and names the general shape: an arbiter is only an arbiter with respect to the writers it can see. * fix(reminders): codex round 4 — the round-3 bound recreated the round-1 starvation Round 3 bounded the poll surface. Round 4 caught what that bound did: the query took the first N rows and the dashboard then discarded the ones it could not show — hidden items, unauthorised items, completed items — so N such rows hide a visible reminder behind them indefinitely, with no continuation to reach it. That is the SAME defect I had removed from the fire path one round earlier, reintroduced in the read path within the hour. The general form is worth stating because I clearly did not hold it: **a bounded window is only safe when the discarding happens BEFORE the bound.** Filtering above a limit is a starvation every time, and it does not matter what the filter is for. Two halves, because the two filters are not the same kind of thing: **Visibility is now scoped IN SQL**, using the same collection-id / item-id sets every other dashboard section gets through `allItems` — the same three-way shape as ItemListParams, where holding both collection grants and item grants is an OR. Invisible rows no longer occupy the window at all, which is strictly better than filtering them out afterwards and is what the sibling sections have always done. **Terminality is paged**, because SQL cannot evaluate it — a collection's schema defines which statuses are terminal. The collector refills from the next page when a page comes back short, bounded by a max scan so a workspace full of completed items cannot turn a dashboard read into a table scan. The bound is 10x the window: the common shape fills on the first page, and the pathological shape terminates in a fixed number of indexed reads. Stopping at the scan bound reports truncation, which is honest — there may be more, and we did not look. The empty-scope case is a THIRD state that reads like the second: nil CollectionIDs means unrestricted, a non-nil EMPTY slice means this caller sees no collections. Without an explicit guard they collapse, because the switch matches none of its cases at length zero and adds no clause at all — so "nothing visible" would return the whole workspace. Three mutants, one survived: the empty-scope guard, because no dashboard-level test produces that state (callers that would are refused earlier by workspace access). Faithful mutant, missing test — it now has a direct one, with a sanity leg so a build returning nothing cannot pass it by accident. A guard for a state nothing exercises is exactly the one that rots. * fix(reminders): codex round 5 — the MCP action I shipped did not work over stdio **P1: local stdio MCP `remind` was unusable.** cmdhelp derives positionals by regex from a command's `Use` string, and `<instant>` inside `remind <ref> --remind-at <instant>` matched — it became a second REQUIRED positional, so dispatch failed with `missing required argument "instant"`. The action was advertised on a transport where it could not run. **The MCP catalog's own tests did not catch it, and the reason is the finding.** That suite builds its cmdhelp document BY HAND: I wrote `Args: mkArgs("ref")` in it, so the fixture agreed with what I meant rather than with what the CLI says. Five parity and drift tests passed against a document I authored to match my own intention — the "a test that agrees with whatever the table says is not a test of the table" shape, which the canonical-events test warns about in its own comment two packages away. The new test reads the REAL command tree via cmdhelp.Build, which is the only thing in this repo that can disagree with me about what the CLI declares. **P2: `pad project ready` withheld the ack handle** that `next` prints. Showing a fired reminder on the surface an agent polls while withholding the id it needs to retire it means the same entry comes back on every poll, forever. **P2: suggestions asserted a collection they did not have.** The orphan branch admits ANY collection — its own comment claimed it gated on tasks "mirroring the active-plan branch", and that comment was simply false — while the output hardcoded `Collection: "tasks"` and the reason said "Open task". Pre-existing for high-priority items since BUG-1082; my overdue bypass widened it to any overdue item, which is how it surfaced. Fixed by carrying the item's REAL collection rather than by narrowing the branch: narrowing would silently drop the non-task items this has surfaced for a year, and the defect is the mislabelling, not the inclusion. The false comment is replaced with what the code actually does. The first version of that test used an overdue IDEA and SKIPPED — ideas use `new`, and the branch requires `open` or an active status, so it never became a candidate. A test that cannot fire is a failed reconstruction, not a pass; the fixture is now a bug-like collection whose vocabulary contains `open`, which is the population the defect can actually reach. Three mutants, three killed. Forty-one across the unit. * fix(reminders): codex round 6 — reminders fired from soft-deleted workspaces **P1, and the only defect in this unit whose consequence leaves the process.** Workspace soft-delete deliberately keeps items for the 30-day restore window, so the candidate query's filter on the ITEM's deleted_at found nothing wrong — and the tick kept firing, emitting outbound webhook events for a workspace whose owner had deleted it, possibly while deleting their account. Both queries now join workspaces and require `w.deleted_at IS NULL`. Nothing is destroyed: a restored workspace resumes firing, which the test asserts, because "stops firing" and "is destroyed" are very different answers to someone who restores a workspace and only one of them is right. That test first failed for the WRONG REASON and the fixture was at fault: it counted every outbox row in the workspace, and item creation writes its own, so the assertion was satisfiable by the fixture itself and discriminated nothing. Scoped to the reminder event type. **`Use: "remind <ref>"` declared a requirement the command contradicts.** cmdhelp derives the machine-readable arg spec from that string, and `--rearm` takes no ref — so the published contract said "required" for something optional. The requirement is CONDITIONAL, which cmdhelp cannot express, so the honest declaration is `[ref]` plus the explicit check that names both call shapes. The round-5 test grew a `required` column, which is what makes this observable at all: asserting only the arg NAMES would have passed. **The pad_item tool description omitted both new actions.** The params were declared and the actions dispatched, but the prose an agent reads to decide what a tool can do did not mention them — discoverable only by someone who already knew to look. It now describes both, including the two things an agent gets wrong: remind_at is an instant, and nothing but an explicit ack retires a fired reminder. Three mutants, three killed. Forty-four across the unit. * fix(reminders): codex round 7 — one predicate for the scan and the arbiter Third instance of one class, so this fixes the SHAPE rather than the instance. The class: the candidate scan filters on something the fire transaction does not revalidate, so a change committed between them fires a reminder that no longer qualifies. Round 3 was a re-armed instant. Round 1's soft-deleted item was the same thing caught from the other side. Round 7 is a workspace deleted between the scan and the fire — the round-6 fix added the condition to the SCAN only, and the arbiter went on not knowing about it. Fixing those one at a time is what let the third happen. `reminderFireable` is now a single string that both sites reference: the scan asks it and the fire UPDATE re-asks it, so they cannot disagree, and a fourth condition is one edit in one place rather than two edits someone has to remember are paired. Written as a correlated EXISTS on item_reminders.item_id rather than a JOIN precisely so the identical text is valid in both a SELECT and an UPDATE, and the scan drops its table alias so the two uses are the same characters. What deliberately stays outside it: `fired_at IS NULL` and `remind_at <= ?` live on the reminder row itself, are already spelled identically at both sites, and folding them in would need a parameter order the shared form cannot express. Said in the comment so the omission reads as a decision. Both directions are now tested at the arbiter — a workspace deleted mid-pass and an item deleted mid-pass — because the item case previously relied on the item load coming back nil, and someone simplifying the EXISTS down to the workspace check alone would otherwise still see green. Three mutants, three killed: the arbiter dropping the shared predicate, and the predicate dropping each of its two halves. Forty-seven across the unit. * fix(reminders): codex round 8 — workspace export silently dropped every reminder WorkspaceExport is a hand-maintained field list, so a new table joins it only if someone remembers. Reminders did not: a backup/restore, or a SQLite→Postgres migration via `pad db migrate-to-pg`, dropped every pending reminder with nothing in the destination to show anything had gone. The line that list has always drawn is item-scoped workspace CONTENT (comments, links, versions — exported) versus per-user state (stars, watches — not). A reminder has no user column and hangs off an item, which puts it on the exported side. Stating the rule rather than just adding the field, because the next person adding a table needs to know which side they are on. LIFECYCLE MARKS ARE CARRIED, not reset. A fired-and-unacknowledged reminder is still owed to whoever armed it, so it arrives pending; an armed one whose instant has passed fires once on the destination's first tick, which is what would have happened had the workspace never moved. Re-arming everything on import would invent a schedule the user did not set. NULL rather than empty string for the unset marks — the lifecycle is defined by NULL-ness, and "" would make a never-fired reminder read as fired at "". TestMigratedTablesCoversTheExport caught the second half, which I would have missed: `pad db migrate-to-pg`'s NUL preflight decides what to REFUSE on from MigratedTables, so a table the migration copies and the preflight does not know about is a gap in exactly the guard that exists to prevent one. Added there too, with the reason it can never actually fire — every column is machine-produced, so it is listed for coverage rather than expectation — and the "six tables" prose it falsified is now seven. Two mutants, two killed: export dropping the block, and import discarding the marks. Forty-nine across the unit. * test(reminders): state the fire-path invariant and pin it from the invariant The lead's read on why rounds 4 and 7 were the same class: the fire path had no stated invariant, so each fix defended an instance. This states it, and derives the pin from the paragraph rather than from the bug history. THE INVARIANT: the candidate scan is a hint and may be assumed to prove nothing. Every condition that made a row a candidate is re-asserted inside the transaction that marks it fired, in the same statement that does the marking, so checking and writing are one atomic act. Worded as "the scan proves nothing" rather than as a list on purpose — a list invites the next person to add a condition to the scan and stop, which is exactly what happened four times here. TestFirePathInvariant is the pin: one table, one row per scan-side condition, each invalidating that condition in the window between the scan and the fire and asserting the same three things — nothing fires, no event leaves, the reminder is not consumed. The earlier per-defect tests are folded in as rows; they said the same thing one instance at a time, which is how four of these shipped. Adding a fifth condition to the scan without a row here should feel like an omission. It carries a positive control, because four cases that all assert nothing happens would pass against a build that never fires at all. The matrix immediately falsified a claim in the paragraph I had just written. I wrote that the item load inside the transaction is "for the payload, not for the check"; removing the item half of reminderFireable alone changes no observable behaviour, because the load then returns nil and the deferred rollback undoes the write. Item liveness is defended TWICE and a single-mutant experiment cannot say which guard is carrying it — removing both is what kills the test. Both are kept, the predicate is named as primary (the row never matches, so no write happens at all), and the asymmetry is stated: workspace liveness has no second line, which is why dropping ITS half does fail the pin. Six mutants: five singles plus the pair. Five killed alone; the item single survives by design and is documented as such rather than left as an unexplained green. Fifty-five across the unit. * fix(reminders): codex round 9 — one legacy row could hide every reminder **P1: items.item_number is NULLABLE and I scanned it into an int.** Migration 006 added the column to existing rows, so a pre-numbering item still carries NULL — and scanning NULL into an int fails the Scan, which fails the QUERY, which degrades the whole pending-reminder section. One old row, and the feature is dark for everyone in that workspace. ListWatchesForUser, which this query was modelled on, uses sql.NullInt64 for exactly this column. I copied its shape and dropped the part that handles the column's actual nullability — the same way of being wrong as the round-5 cmdhelp fixture: borrowing a form without borrowing what it knows. The legacy row now carries no ref rather than a fabricated "PREFIX-0", which would name a different item. **P1: export shipped reminders that import could only discard.** The items section filters on deleted_at IS NULL, so a soft-deleted item is not in the bundle and its reminder can never be reunited with it. My comment claimed the item_links rationale — round-trip the raw graph so a restore reunites them — which is true for links and false here, because links keep soft-deleted endpoints in the bundle and items do not. A link is a row ABOUT two items; a reminder whose item is absent is a dangling schedule. **P2: import wrote remind_at raw.** Import is a writer, and a bundle is not necessarily one this server produced — hand-edited, or from another instance. A local offset or a bare date would land in the one column every comparison downstream treats as a UTC instant, firing early, late, or never. It now normalizes like every other door. An unparseable value is SKIPPED with a warning rather than failing the restore, matching the lenient import-side precedent already in this file, and the raw value's LENGTH is logged rather than its content. Three mutants, three killed; two needed rewriting because the single-line form did not compile — reverting the nullable scan also requires reverting the render, and dropping the normalization orphans a variable. PROCESS FAULT, recorded because it makes this round's findings weaker than they look: I edited the tree while this review was reading it — committed the invariant work and ran five mutation experiments, which write and restore source, over the same files. A review binds to the tree it read and I moved it underneath. Every finding above was re-verified against the current tree before being acted on, and the next round runs with no concurrent edits. * fix(reminders): codex round 10 — one orphaned item aborted a whole restore An ORPHANED item — one whose collection is missing from the bundle — still gets an itemMap entry. It has to: the entry is written before the skip because parent resolution inside the same loop reads the map for items it has not reached yet. So `itemMap[x] != ""` is satisfied by an id that names no row, and inserting a foreign key to it fails (SQLite enforces FKs here via the DSN's `_pragma=foreign_keys(on)`; Postgres always does). The pre-existing mapping is the sharp edge. The aggravating half was mine: this loop treated a failed reminder insert as FATAL, where item_links and item_versions both skip, so one orphaned item carrying a reminder rolled back an entire 900-item workspace restore. A reminder is the least critical thing in a bundle and it had the strictest failure handling in the file. Both halves fixed: the loop gates on items that actually landed, and a failed insert warns and skips like its siblings. TWO GUARDS THAT ONLY DIE TOGETHER, and this is measured rather than assumed. Reverting either alone leaves the test green — with the map gate restored the skip survives the FK failure, and with the fatal return restored the gate means the insert never fails. Removing both is what fails it. They are kept as a pair because they defend the same failure at different depths (prevent the bad write / survive a bad write arriving some other way), and the pair is recorded in the code so a future reader does not delete one as dead after watching its mutant survive. Second time this shape appeared today; the first was item liveness on the fire path. The bundle in the test is hand-built, because ExportWorkspace cannot produce an orphan — which is the reason it needed a test. That shape only arrives from a hand-edited or foreign bundle, and surviving those is what import is for. Three mutants: two singles that survive by design, plus the pair that kills. Sixty-one across the unit. * fix(reminders): codex round 11 — four contract slips, one of them another unit's **suggested_next returned up to eight entries against a cap of three.** Round 3 prepended reminders PAST the list's own cap, reasoning they should not compete for slots. Every consumer — the web dashboard, `pad project next`, `pad project ready` — is written for three. Worse, it silently falsified a decision recorded elsewhere: BootstrapDashboard deliberately has no suggested_next_overflow_count BECAUSE this list is capped at three upstream, and its comment names raising that cap as the moment to add one. My change made another unit's reasoning wrong in a file I never opened. The combined list is now trimmed back to three, reminders still leading — a reminder can push a task suggestion out, which is the right way round, and the full set stays addressable in pending_reminders. My first version of that trim used `limit`, which is REASSIGNED above to len(candidates) — so on a workspace whose only entries are reminders it would have truncated to zero, killing precisely the case the surface exists for. Caught by reading the surrounding lines before running anything; it has its own test now. **pending_reminders was uncapped in the bootstrap projection.** BootstrapDashboard embeds *DashboardResponse, so every new field joins the boot payload automatically — here, a window of up to 50, which is the budget PLAN-1410 spent a unit trimming. Capped at 5 with an overflow count, under its own constant rather than borrowing bootstrapAttentionCap: they answer different questions and a future change to one must not silently move the other. **Truncation was reported from the wrong question.** The collector used the store's `more` flag, which answers "is there another PAGE", not "did I read all of THIS one" — so a window filling part way through the final page reported that the caller had seen everything while unread rows sat behind the fill point. The paging bounds are now injectable so the case is testable at all: building it with a window of 50 needs ~75 rows in a specific pattern, with a window of 3 it is four. **Import accepted acked-without-fired**, which is not one of the lifecycle's three states. Such a row fires, is excluded from the pending surface because it is already acked, and can never be acknowledged because AckReminder requires acked_at IS NULL — an event emitted into permanent invisibility. The acknowledgement is dropped and the schedule kept, since an ack of something that never fired means nothing. Five mutants, five killed (one rewritten — removing the flag orphans a variable). Sixty-six across the unit. * fix(reminders): codex round 12 — a read is not a hold; scope the arm; ack from the ack Four P2s from round 12 (two independent runs, both landing on the same line of the fire path), each closed at the layer where it lives: - fireOneReminder pins the item and workspace rows FOR NO KEY UPDATE on Postgres before the arbiter UPDATE. reminderFireable re-asserted liveness at the predicate's instant and nothing held it to the commit instant; under READ COMMITTED an archival could commit in between and the event left the process about a deleted resource. Same idiom and same lock strength as CreateAttachmentForLiveItem; SQLite is excluded by its BEGIN IMMEDIATE, not skipped for convenience. Two PG-only pins verify "blocked" in pg_stat_activity, not by elapsed time; the pin-removed mutant fails both. - CreateReminder asserts "live item of THIS workspace" in the INSERT's own SELECT and returns ErrReminderItemGone otherwise. The table had an FK and no same-workspace constraint; a mismatched pair fed another workspace's title to this one's dashboard and webhooks. Handler maps it to 404. - AckReminder matches every fired row (COALESCE keeps the first ack, updated_at moves only when acked_at does), so a no-match means exactly "not fired at the instant of the ack". The handler no longer decides 409-vs-200 from the row it read before the UPDATE. - The invariant paragraph gains its missing sentence: "at that instant" means the commit instant, and the pin is what makes the predicate's instant and the commit instant the same one. Round-12 caveat carried: both runs were static reads (sandbox blocked Go's build cache), so "four" is a floor, not a measurement. Refs IDEA-2641 * fix(reminders): codex round 13 — a reminder's workspace must agree with its item's, at every read Every reader scoped by r.workspace_id and then joined the item without asserting the two agree. No door writes a disagreeing row today (CreateReminder derives the pair from the item; import maps within the workspace), and the table has nothing that forbids one — so a hand-edited bundle, a future move door, or a direct write would carry one workspace's item into another's dashboard, export, and webhooks. The identity goes into reminderFireable (scan + arbiter), the Postgres row pin, ListPendingReminders and the export query. One test writes the row raw — the only way one can exist — and asserts it is inert at each site; the predicate-removed mutant scans and fires it. Refs IDEA-2641 * fix(reminders): codex round 14 — the by-id and by-item reads assert the same identity as every other read GetReminder scoped by the row's own workspace_id and ListRemindersForItem by item_id alone, so a row whose two columns disagree — the class rounds 12 and 13 closed at the scan, the arbiter, the pin, the pending surface and the export — was still readable through the two reads that reach a single row. reminderOwned is that identity on its own, without the liveness half those two reads must not have (a fired reminder on an archived item is history worth showing). The write paths reach a row only through GetReminder, so scoping it scopes them; a row no door can write needs no door to delete it. ListRemindersForItem now takes the workspace its caller already resolved the item in. The raw-row test asserts both reads refuse the row from both sides; the reminderOwned-removed mutant surfaces it through GetReminder. Refs IDEA-2641 * fix(reminders): codex round 16 — an archived item's reminders are readable, and its verbs say "archived" The doors resolved the item live. Listing an archived item's reminders answered 409 from a GET, and ack/re-arm/delete answered a bare 404 for a reminder that exists on an item that exists — while the store, since round 14, deliberately keeps that history readable. The API already has a posture for archived items: GET reads them, mutations answer 409 "archived … restore it before editing" (writeItemResolveError). The list now follows handleGetItem; the lifecycle verbs load the item include-deleted, run the visibility check first, and then answer the same 409 every other item mutation does. One test walks archive → list 200 / ack 409 / arm 409 → restore → ack 200 on the same rows. Refs IDEA-2641 * fix(reminders): codex round 17 — one suggestion per item, the archived 409 by slug, and the door courtesy named Three findings on the server pass. (1) An item that was both a fired reminder and an ordinary candidate appeared in suggested_next twice; the ordinary entry is dropped, the reminder entry (which carries the ack id) stays, and two reminders on one item remain two entries. (2) Round 16's 409 for an archived item's reminder was written by re-resolving item.Ref, which is derived and empty for a legacy item with no item_number — so the class most likely to be legacy fell through to a bare 404. The slug is handed over instead. (3) The archived check in resolveReminderForWrite is check-then-write, and an archive landing in between lets the verb through: accepted and documented — it is the posture of every item mutation here (UpdateItem's UPDATE has no liveness clause), the outcome is benign, and putting liveness in AckReminder's WHERE would re-create the no-match ambiguity round 12 removed. Refs IDEA-2641 |
||
|
|
e94e9afbea |
Merge pull request #1240 from PerpetualSoftware/fix/bug-2850-field-coercion
fix(server,mcp,cli): type field values server-side; carry the fields object natively (BUG-2850) |
||
|
|
80be76a3ce |
docs(mcp): the detectFieldConflicts header stated the pre-round-14 reach (BUG-2850)
Comment-only. The lead caught it in the package review. The "SCOPE:" paragraph still said the pass runs only when a `fields` object is present, and that a top-level-vs-`field:[]` collision without one is outside it. Round 14 falsified both halves — the pass runs from both action entry points on every call — and the body's own comment said so while the header contradicted it. On the one boundary this loop spent eighteen rounds on, the header is what a future reader trusts. CONVE-23 is exactly this and I missed it: the round-14 commit swept the version.go prose and the test comments, and left the header of the function it had just changed. The paragraph now states the actual reach: both entry points regardless of `fields`; alias collisions adjudicated always and refused even on equal values; same-name collisions adjudicated only when the `fields` object carries THAT key, which is a per-key question; the round-7 exemption as the sole carve-out, itself narrowed to keys the CLI can express (the compat IDs refuse, but only when a top-level compat value is present), with padded entries outside it. It closes with the instruction the loop earned: say a new condition's QUANTIFIER out loud before writing it. Rounds 15, 16, 17, 19, 20 and 21 were each that question answered by assumption. gofmt clean · go vet clean · go test ./internal/mcp/ ./cmd/pad/ green |
||
|
|
9ecc59af1e |
fix(store): refuse a schema with trailing content instead of truncating it (BUG-2873)
Codex round 3, one P2. `json.Decoder.Decode` stops at the end of the FIRST value and ignores whatever follows, where `json.Unmarshal` refuses it — so a stored schema with junk after the object would be silently truncated by the rewrite. It is now treated as unparseable and left alone, the same posture as any other schema this migration cannot faithfully reproduce. **The Postgres gate then failed the new test, and the failure is the finding.** It failed at the SEED, not the assertion: `ERROR: invalid input syntax for type json (SQLSTATE 22P02)`. `collections.schema` is TEXT on SQLite (`005_collections.sql:10`) and JSONB on Postgres (`pgmigrations/001_initial.sql:114`), so a value with trailing content cannot be STORED on Postgres at all. The state this guard defends against is reachable on one dialect and forbidden by the column type on the other. So the test skips on Postgres with that reason recorded. Asserting there would be asserting about a state that cannot exist — and reading WHICH LINE failed is what separated "my test is not portable" from "the product is broken on PG". **Second instance of the mutation harness reporting a false survivor**, same cause as the last: deleting the guard leaves `io` unused, the mutant fails to compile, and counting `--- FAIL` lines sees zero. With `_ = err` in place of the return it dies immediately. Twice in one unit makes it a harness defect, not bad luck: a runner that counts test failures must check the BUILD separately, or every non-compiling mutant reads as a hole in the tests. Mutation matrix 7 of 7 killed. Gates: `gofmt` clean, `go vet ./...` ok, full `go test ./...` green on SQLite, full `internal/store` green on Postgres (502.2s, private container at 127.0.0.1:5473, detached with a sentinel). |
||
|
|
a1b8e63a29 |
fix(mcp): a nil top-level value is absence, for every key (BUG-2850)
Codex round 21, one P2 and no P1 — a false refusal, and the finding
named a strict subset of it.
topLevelValueProvided returned true unconditionally for the compat IDs,
and fell through to true for everything else, so a nil counted as a
supplied value and refused against a `fields` entry for the same key.
Nothing writes a nil: the HTTP mapper's `.(string)` assertion drops it
and BuildCLIArgs has no flag value to emit, so both doors resolve to the
`fields` value.
The finding named `assigned_user_id` / `agent_role_id`. Probing the
population first — the habit this unit has been beating into me — showed
all five top-level keys behaving identically, because the non-compat
path fell through to `return true` as well. Fixing the named pair alone
would have left `status: null` refusing.
Mutation matrix, three directions:
remove the nil check -> all five key legs fail
fix ONLY the named compat pair -> status / priority / parent fail
treat the empty compat clear
as absence too -> both clear-semantics tests fail
The middle mutant is the population-versus-instance distinction made
executable: a fix that satisfies the reviewer's example and nothing else
is red, by name, in three legs.
Gates: gofmt clean · go vet clean · go test ./... green (29 packages) ·
contract-drift gate green
|
||
|
|
499387e99d |
fix(store): never regress a migrated token; keep large integers intact (BUG-2873)
Codex round 2: four findings, two fixed here and two filed as their own items.
**Migrated siblings' OCC tokens could REGRESS.** A sibling updated between this
rename's timestamp and the scan already holds a newer `updated_at`; stamping the
rename's value on it moved the token BACKWARDS — breaking the strictly-increasing
invariant the transaction above exists to maintain, and re-validating a token the
client should have lost. Each rewritten row now takes `max(current + 1ns,
renameToken)`, computed in Go from the value read under the row lock rather than
by comparing timestamp TEXT, which the existing comment warns is never safe.
**Large integers in unknown properties were corrupted.** Round 1 fixed the typed
round-trip dropping unknown keys, but decoding into `interface{}` turns every
JSON number into float64, so `9007199254740993` came back CHANGED. A rename would
silently damage a property it exists only to carry through. `UseNumber` keeps the
literal text.
## Filed, not absorbed — both because their dependents are not the relation feature
- **BUG-2875** — a collection CREATED during a rename escapes the scan's
`FOR UPDATE` and keeps a relation aimed at the old slug. Closing it means
`CreateCollection` takes the workspace lock, which changes the concurrency
behaviour of every collection creation on the instance. Same reasoning that
split IDEA-2874 out; this unit's reviewability rests on affecting zero live rows.
- **IDEA-2876** — migrated siblings emit no `collection_updated` event, so an open
page keeps the pre-rename schema until reload. Handler/event layer; the store
publishes nothing.
## The mutation harness was reporting a false survivor
Counting `--- FAIL` lines treats a mutant that FAILS TO COMPILE as one that
survived — zero failures either way. Removing the token guard leaves
`rowUpdatedAt` and `renameToken` unused, so that is exactly what happened, and it
read as "the guard is untested". With a compiling mutant (`_ = rowUpdatedAt`) it
dies immediately. Worth stating because the failure mode is silent and points the
wrong way: it invents doubt about code that is fine, and would equally hide a
real survivor behind an unrelated build break.
Mutation matrix 6 of 6 killed.
Gates: `gofmt` clean, `go vet ./...` ok, full `go test ./...` green on SQLite,
full `internal/store` green on Postgres — 460.9s, private container at
127.0.0.1:5473, run detached with a sentinel.
|
||
|
|
a9f2405903 |
fix(mcp): the compat exception turns on a top-level value, not on the key (BUG-2850)
Codex round 20, one P2 and no P1 — another false refusal from a reason
of mine applied past the source it was verified on.
Round 15's reason was specific: a TOP-LEVEL compat param has no CLI
flag, so BuildCLIArgs drops it while HTTP reads it, and the doors
receive different writes. I then keyed the exception on the KEY being a
compat one, which caught `field:["assigned_user_id=A",
"assigned_user_id=B"]` — two array entries, no top-level value, no
asymmetry: both doors keep the last and lift the same column. Refused a
call that resolves deterministically.
The gate now asks whether a top-level compat value is actually present,
which is the condition the reason describes.
Mutation matrix, both directions:
broaden it back to any compat-keyed contribution -> only the two-entry legs fail
drop the exception entirely -> only the top-level legs fail
(round 15's defect returns)
Round 15's own case is a leg of the new test deliberately: without it
this pin would pass on a build that dropped the compat exception
altogether, which is the defect round 15 existed to fix.
Gates: gofmt clean · go vet clean · go test ./... green (29 packages) ·
contract-drift gate green
|
||
|
|
6428e7db31 |
fix(store): lock, stamp and preserve on the relation retarget (BUG-2873)
Codex round 1: four findings, three P1, all real. **The migrated siblings' concurrency token was not advanced.** `collections.updated_at` doubles as the OCC token (BUG-2265), so rewriting a sibling's schema without touching it left a client holding the PRE-rename schema — and a token that still matched — able to write it straight back and undo the migration. Every rewritten row now takes the rename's own token, so the whole rename shares one instant. Pinned by asserting the stale token now 409s. **The scan did not lock the rows it rewrites.** A concurrent schema update to a sibling could commit between the SELECT and the UPDATE, and this transaction would then overwrite the newer schema with its stale copy. `FOR UPDATE` on Postgres, ordered by id so the multi-row acquisition is deterministic; SQLite is covered by its BEGIN IMMEDIATE write lock. **The old slug came from the pre-transaction snapshot.** Two tokenless concurrent renames of the same collection both read the ORIGINAL slug outside the lock; the loser would migrate `original -> its own new slug` while the relations already said the WINNER's, matching nothing and stranding them at a name no collection holds. The slug is now re-read alongside the token under the row lock. This is the READ — the ALLOCATION of the new slug is still outside the transaction and still IDEA-2874's, deliberately. **Re-marshaling through `models.CollectionSchema` dropped unknown properties.** That struct has fixed fields, so unmarshal+marshal silently erased anything it does not declare — a rename would quietly strip forward-compatible metadata from every relation-bearing schema in the workspace. It now edits the raw decoded JSON, touching only `fields[i].collection`. ## Two instruments that were not instruments Both found by mutation, not by reading: - **The deadlock test passed against its own mutant in 0.44s.** Two unsynchronised goroutines never collided. With a start barrier and 40 rounds it now fails in 1.4s with `ERROR: deadlock detected (SQLSTATE 40P01)` — so Rook's hazard was reproducible, not theoretical. - **The pre-tx-slug mutant survived the first matrix**, because nothing forced the interleaving. Rather than call it untestable, it is pinned by an end-state invariant that holds under ANY interleaving — whatever slug the collection ends up with, every relation aimed at it points there — over 40 concurrent rounds. It fails at round 1 under the mutant. Mutation matrix 4 of 4 killed. Gates: `gofmt` clean, `go vet ./...` ok, full `go test ./...` green on SQLite, and the full `internal/store` suite green on **Postgres** — 448.6s on a private container at 127.0.0.1:5473, never the shared 5445 a sibling seat may tear down. Run detached with a sentinel after the first attempt was killed at a turn boundary; the harness kills backgrounded tasks, it does not kill disowned ones. |
||
|
|
052850f613 |
fix(mcp): both gates ask the per-key question; compare like with like (BUG-2850)
Codex round 19, two P2 and no P1. Both were FALSE REFUSALS my own fixes
introduced — the first is round 17's mistake in the sibling gate.
[P2] THE SAME-NAME GATE WAS STILL PER-REQUEST. Round 17 made the padded
gate per-key and left this one asking whether the request has any
`fields` object. With `fields:{"other":"x"}` and
`field:["effort=l","effort=s"]`, `effort` is not in the object, nothing
arbitrates it but the doors themselves, and both keep the last entry —
so a call that resolves deterministically was refused. The predicate is
now `canonicalized`, the same per-key question the other gate asks.
`fieldsPresent` no longer exists anywhere in the pass, and its absence
is commented as the fix's shape: a future gate reaching for "does the
request have a fields object" is almost certainly this mistake a third
time.
[P2] TRIMMED AND UNTRIMMED VALUES WERE COMPARED. Entry values are
trimmed for comparison because ingestFieldKVP trims them; the `fields`
object's value was compared raw. `fields:{"note":" x "}` with
`field:["note= x "]` read as " x " vs "x" and refused, though both doors
write " x ". Only the COMPARISON key is trimmed now — `raw` and the
re-emitted wire value keep the caller's whitespace.
Mutation matrix:
same-name gate back to per-request -> only the unrelated-key leg fails
stop trimming for comparison -> only the whitespace-equal leg fails
trim the EMITTED value too -> only the re-emission leg fails
THE THIRD MUTANT SURVIVED AT FIRST, and it was unreachable rather than
unobserved: the whitespace test's entry is already canonical, so the
re-emission path never ran and a mutant trimming the emitted value
changed nothing it could see. Added a leg whose KEY is padded, which
forces the re-emission, and asserted the emitted value still carries the
caller's whitespace. Third time this loop that asking "is the mutant
faithful" before "is the test weak" found a real hole (CONVE-28).
Control legs, both directions: a `fields` object that DOES carry the key
still refuses differing values, and genuinely different values are still
refused however they are padded.
Gates: gofmt clean · go vet clean · go test ./... green (29 packages) ·
contract-drift gate green
|
||
|
|
4687c46f94 |
fix(store): migrate relation fields when their target collection is renamed (BUG-2873)
`models.FieldDef.Collection` holds the target's SLUG, and it is the ONLY pointer a relation field carries — there is no id beside it to fall back on. `UpdateCollection` re-slugifies on rename and nothing migrated the definitions aimed at the renamed collection, so every relation field pointing at it was stranded: the picker filters on a slug that resolves to nothing and the field silently stops being fillable. `retargetRelationFieldsTx` re-points them in the SAME transaction as the rename, for the reason the field-value migrations already run there: a failure must roll the rename back rather than commit collections pointing at a slug that no longer exists. **It parses instead of string-replacing.** A schema's JSON contains the old slug in places that must not move — a text field's `default`, a select's `options`, a label. Only `FieldDef.Collection` on a `relation` field is a reference. Export's `remapFieldIDs` gets away with a blind replace because it substitutes UUIDs, which cannot collide with prose; a slug is a word. A control test pins that. **The renamed collection is included deliberately** — a relation targeting ITSELF needs the same rewrite — and the rewrite lands after the caller's own `schema` write in the transaction, so a simultaneous schema edit composes rather than being reverted. Both have tests. ## The deadlock hazard, and why the existing comment does not cover it The lock-order comment above this transaction is a Codex P1 fix that orders the workspace lock against ONE collection row lock, because until now nothing took more than one. This change writes SIBLING collection rows, so two concurrent renames of mutually-referencing collections take those locks in opposite orders. **Reproduced, not theorised:** with the serialization removed, the test fails in 1.4s with `ERROR: deadlock detected (SQLSTATE 40P01)` on Postgres. Renames now take the workspace lock — previously acquired only when `len(input.Migrations) > 0` — BEFORE the row lock, which closes it without inventing a second ordering rule to keep in sync with the first. **The first version of that test was not an instrument.** Two unsynchronised goroutines passed against the same mutant in 0.44s, having simply never collided. It takes a start barrier and 40 rounds to be evidence. ## Scope The out-of-tx slug allocation (`uniqueSlugExcluding(s.db, …)` at :420, before `s.db.Begin()` at :503) is deliberately NOT touched — filed as IDEA-2874. Its dependents are every collection rename in every workspace, not the relation feature, so it does not belong in a change whose reviewability rests on affecting zero live rows. `UNIQUE(workspace_id, slug)` makes today's behaviour loud rather than lossy, so it can wait. A census found ZERO relation fields across all 11 accessible workspaces on this instance, and no shipped template declares one — this repairs the rename path before PLAN-2857 creates the population, which is why a migration is not needed. Gates: `gofmt` clean, `go vet ./...` ok, `go build ./...` ok, full `go test ./...` green on SQLite, and the full `internal/store` suite green on **Postgres** (private container on 127.0.0.1:5473, never the shared 5445 a sibling seat may tear down). Pin written and run BEFORE the fix per team CONVE-29: 2 propagation tests failed, the control passed. |
||
|
|
21a3057389 |
fix(mcp): keep per-entry multiplicity in the conflict pass (BUG-2850)
Codex round 18, one P2 and no P1 — and the fix is upstream of the rules rather than another rule. parseFieldArray indexes by NORMALIZED key, so two entries naming one key collapsed into a single index slot, and this pass walked that index. Its own input was lossy: `field:["effort=l", " effort=l"]` arrived as ONE contribution, fell under the len < 2 early exit, and passed unchecked — HTTP trims both to `effort` while stdio writes `effort` AND a junk `" effort"`. The pass claims to adjudicate one canonical key offered by multiple sources; two array entries ARE multiple sources, and it could not see them. It now walks the raw entries, so multiplicity survives and the existing rules apply unchanged — no new branch. The index is deliberately discarded here and the discard is commented, because reaching for it is the natural thing to do next. I NEARLY CHANGED THE CODE TO SATISFY A WRONG TEST. The third leg was first written asserting that two canonical entries with DIFFERING values are refused. The code disagreed, and the code was right: ingestFieldKVP (HTTP) and the --field loop in cmd_item.go (CLI) both do `map[key] = val` in entry order, so each door keeps the LAST entry and they agree. That is the round-7 boundary exactly — a visible duplicate with a resolution the caller can predict — and refusing it would have contradicted the boundary the lead confirmed, on two doors pinned to agree. Verified by reading both loops before touching anything. The leg is kept, inverted, because it is the one that stops a future "refuse every repeated key" simplification from looking correct. Mutation: restoring the collapse (walk one contribution per key) fails exactly the padded-twin leg, and leaves the two control legs green. Gates: gofmt clean · go vet clean · go test ./... green (29 packages) · contract-drift gate green (go test ./internal/mcp/ -run 'CoversEveryCatalogAction|VersionMatchesToolSurface') |
||
|
|
130c854052 |
fix(mcp): canonicalization is a per-KEY property, not a per-request one (BUG-2850)
Codex round 17, one P1, and it is the third consecutive round where my
own fix generalized a property verified on one subset to the whole.
Round 16 gated the padded-entry refusal on "no `fields` object",
reasoning that a `fields` object makes reshapeItemFields re-emit the
entry canonically. That holds for keys IN that object. With `fields:{}`,
or a `fields` carrying some OTHER key, nothing canonicalizes
`field:["status = done"]` and it reaches the doors padded exactly as it
does with no `fields` at all — HTTP writes `status`, the CLI writes a
junk `"status "` beside it.
The predicate is now the actual question: will anything canonicalize
THIS key.
The pattern is worth naming because it is now a habit rather than an
accident:
round 15 — a premise true of schema-declared params, applied to the
compat IDs, which are undeclared precisely so it cannot hold
round 16 — the check placed below the exemption, so it covered one key
class (caught by my own test's control leg)
round 17 — a per-key property read as per-request
Each time the fix was correct for the case in front of me and wrong for
its siblings, which is the same shape as the defects this unit started
with. The canonical restructure removed it from the CODE; it evidently
did not remove it from how I reason about the code.
Mutation matrix, both directions:
revert to the per-request gate -> only the two uncovered-key legs fail
refuse even when canonicalized -> only the canonicalized control fails
The third leg of the new test is the control that makes the distinction
real: with the key present in `fields` the entry IS canonicalized, so
the call must still succeed. A fix that refused whenever anything was
padded passes the first two legs and fails this one.
Gates: gofmt clean · go vet clean · go test ./... green (29 packages) ·
the contract-drift gate ran green
(go test ./internal/mcp/ -run 'CoversEveryCatalogAction|VersionMatchesToolSurface')
|
||
|
|
3696686377 |
fix(mcp): a padded entry colliding with a param is not an equal duplicate (BUG-2850)
Codex round 16, one P1, and its placement is the whole lesson. The conflict index is normalized — that is what lets a padded entry be recognized as a collision at all — so `field:["k = A"]` compared EQUAL to a top-level `k:"A"` and the pair was accepted while the entry stayed padded on the wire. HTTP trims it and writes `k`; the CLI does not, and writes a junk `"k "` key instead. The normalization that makes the collision VISIBLE is exactly what made accepting it wrong: equality on the normalized form licensed a collapse of the RAW forms, which are not equal at all. So equality only licenses a collapse when both doors receive the same write, and this check now runs BEFORE the same-name exemption. MY FIRST DRAFT PUT IT AFTER, which is round 15's mistake repeated one round later. Round 15 was a premise verified for declared params and generalized to two keys it did not hold for; placing this below the exemption meant only the compat IDs were checked, when padding breaks the exemption's own premise — both doors resolve the duplicate identically — for EVERY key class. The declared-param leg of the new test caught it, and the mutation matrix pins the placement rather than just the behaviour. Scope held: a padded entry standing ALONE, with no colliding param, is untouched. That is BUG-2870, ruled out of this PR, and fixing it changes what every CLI caller receives rather than only callers who supplied one key twice. There is an executable test for that boundary, so growing into BUG-2870's territory without a ruling goes red. Mutation matrix, three directions: remove the check -> both padded legs fail restrict it to compat IDs -> only the declared-param leg fails (the first draft) extend it to a lone entry -> the BUG-2870 scope-boundary test fails gofmt clean · go vet clean · go test ./... green (29 packages) context: 56.4% (session-shape) |
||
|
|
dae7bbf233 |
fix(mcp): the same-name exemption holds only where the doors agree (BUG-2850)
Codex round 15, one P1, and it lands on the boundary I defended at round 7 and the lead confirmed. The exemption's premise is "both doors resolve a same-name duplicate identically". That is TRUE for a schema-declared param: the CLI has a real flag, so stdio receives BOTH forms (`--status open --field status=done`) and its overlay order resolves them exactly as the HTTP mapper does. I verified that per door and pinned it. It is FALSE for the v0.16 compat IDs, and being undeclared is precisely why. BuildCLIArgs emits the CLI's real flags; there is none behind `assigned_user_id`, so the top-level value is DROPPED and stdio sees only the field entry while HTTP reads the param. `assigned_user_id:"A"` with `field:["assigned_user_id=B"]` assigns two different people depending on transport — with no `fields` object anywhere, which is why the canonical pass's `fields`-gated half never saw it. So I generalised a premise from the params I had verified to the two whose whole nature is being unverifiable that way. The exemption is now narrowed to keys the CLI can express; the compat pair refuses. Swept for the prose this falsifies (CONVE-23): the v0.27 changelog entry said the no-`fields` same-name case is "NOT refused, deliberately", full stop. It now states the narrower rule and why the compat IDs are outside it — the entry is the artifact a consumer reads to decide what this version does, and leaving it broader than the code would have been worse than never writing it. Mutation matrix, both directions: restore the blanket exemption -> only the two compat legs fail remove the exemption entirely -> only the declared-param control fails The over-narrow mutant did NOT compile on the first attempt (`declared and not used: fieldsPresent`), which proves nothing about the tests, so it was rewritten to compile before being counted. A compiler-killed mutant is a failed experiment, not a passing one. The declared-param control is in the same test deliberately: without it this pin would pass on a build that abandoned the round-7 boundary outright, which is the opposite defect and just as wrong. gofmt clean · go vet clean · go test ./... green (29 packages) |
||
|
|
4b6e321924 |
feat(mcp): bump ToolSurfaceVersion to 0.27 (BUG-2850)
The contract this branch changes is advertised in the handshake under
capabilities.experimental.padToolSurface, and agents branch on it. It
still said 0.26.
Caught by reading the artifact a CONSUMER reads rather than the code —
and the repo already had the instrument: tool_surface_drift_test.go
fails when instructions.md or README.md drift from the constant, so the
bump immediately named both documents. They are updated with what
actually changed, not just re-titled.
Bump grounds are v0.26's own, and v0.25's, v0.16's, v0.10's and v0.9's:
no tool name, action enum or parameter shape changed, and the behaviour
did. This branch REFUSES calls 0.26 accepted:
- two names for one target in a single call (parent/plan,
assign/assigned_user_id, role/agent_role_id), refused even when the
values match, because the names address one thing through
incomparable vocabularies and the doors resolved them differently;
- the same key through the `fields` object and another source with
differing values (equal ones collapse);
- a non-string `assign`/`role`, which one door dropped silently and
the other rejected;
- an empty hierarchy value inside `fields`, which promoted onto a
param both doors read as "not supplied" and so reported success
having detached nothing.
Every one of those replaced a call that SUCCEEDED while doing something
other than what it said, so the break is the fix in each case.
The additive halves are in the same entry because they are one contract
change: server-side coercion (a declared number/json field was
unwritable from the remote transport at all), the `fields` object
carrying native types, and `warnings.undeclared_fields` on write
responses.
Deliberately NOT refused, and stated in the entry so it reads as a
decision: a top-level param colliding with a `field:[]` entry under the
same name with no `fields` object. Both doors resolve that identically
and it is visibly a duplicate.
gofmt clean · go vet clean · go test ./... green (29 packages)
|
||
|
|
bb62cfdc95 |
test(mcp): derive the conflict property's population from the declared schema (BUG-2850)
The lead's finding after round 14, and it is a sharper statement of what went wrong than mine was. The property test enumerated its sources by hand, and that hand-written list came from the same head as detectFieldConflicts. A property whose input list mirrors the implementation cannot see a source the implementation forgot — which is exactly how the round-14 defect survived it: the property SKIPPED param-vs-array pairs as "out of scope", which was the implementation's assumption restated as a test assumption. So the population now comes from the DOOR'S DECLARED CONTRACT — the live pad_item ToolDef's parameter list, which is what agents read — and every declared param must be either classified by the conflict machinery or explicitly excluded with a reason. The two lists have genuinely different origins (the tool schema vs the four key sets), which is the whole point: a test that derives its expectations from the thing it checks cannot fail. It earned its place immediately by naming ten declared params I had not classified. Each is now excluded WITH its reason, because a bare list would let a future field-writing param be silenced by adding one word to it — the round-14 mistake in miniature. The interesting group is summary/details/decision/rationale. Those DO change item state, so excluding them is a real claim rather than a shrug: they write implementation_notes / decision_log through their own actions, and those exact keys are REFUSED through `field` and `fields` (BUG-2627 / BUG-2675), so they cannot reach one key by two routes — which is the only thing this pass adjudicates. The reverse direction is checked too: every key the machinery classifies must be reachable through the declared schema, or be a documented undeclared form (the v0.16 compat IDs, and `plan`, a fields_patch pseudo-key with no top-level param). That fails if a key set goes stale against the schema. gofmt clean · go test ./internal/mcp/ green |
||
|
|
eb37c0e53c |
fix(mcp): give the canonical pass full reach; equal structures collapse (BUG-2850)
Codex round 14, two P1 — both in the round-13 restructure, and the first
is the restructure repeating the mistake it was built to end.
[P1] THE CANONICAL PASS DID NOT REACH THE NO-`fields` CASE. It ran from
inside reshapeItemFields, which returns early without a `fields` object,
so an ALIAS pair arriving through the top level and the `field` array
alone slipped past: `assigned_user_id:"B"` with `field:["assign=dave"]`
applies the compat ID over HTTP while stdio drops it and sends only the
generic field. Two different people assigned, from one call.
The tell is that round 7 had ALREADY built an always-run alias guard —
for the hierarchy pair only. So the restructure meant to end guard
accretion had itself left two alias mechanisms with different reach, and
the pair the older one covered is exactly the pair that kept working.
detectFieldConflicts now parses its own inputs and runs from both action
entry points regardless of `fields`; checkHierarchyAliasAmbiguity is
deleted as subsumed. One mechanism.
The alias half is ungated; the SAME-NAME half stays gated on `fields`,
which is the round-7 boundary the lead confirmed and is unchanged.
Last-write-wins is defensible when both sources name one key and
indefensible when two names address one target through different
vocabularies — a slug and a UUID cannot be compared, so co-occurrence is
ambiguous however it arrives.
[P1] EQUAL STRUCTURES WERE REFUSED. `tags:["a"]` plus
`fields:{"tags":["a"]}` is one unambiguous value, and scalarEqual had
always collapsed it; the round-13 pass refused whenever either side was
structured. A regression I introduced, that nothing in the suite caught
because every existing tags test passes a structure on one side only.
Equal structures now collapse, differing ones still refuse.
The property test is widened rather than merely extended: it previously
SKIPPED param-vs-array pairs as out of scope, and that exclusion is
precisely where the round-14 defect lived. It now covers every source
pair — 72 combinations, up from 48.
Mutation matrix:
re-gate the alias half on `fields` -> property fails on the param×array legs
revert equal-structure collapse -> only EqualStructuredDuplicateCollapses fails
make the collapse unconditional -> only DifferingStructuredDuplicateRefused fails
The last two are the both-directions pair: under-apply and over-apply
each fail exactly one leg, so the pins bracket the behaviour instead of
agreeing with it from one side.
Control kept explicit: the round-7 same-name boundary still passes on
BOTH doors (internal/mcp + cmd/pad SameNameDuplicate tests), so widening
alias detection did not quietly swallow the case that resolves.
gofmt clean · go vet clean · go test ./... green (29 packages)
context: 45.1% (session-shape)
|
||
|
|
1de4fd48dc |
refactor(mcp): one canonical view, one conflict check (BUG-2850)
Codex round 13 found a fourth consecutive defect in a prior round's fix,
which fired the lead's restructure trigger. The finding and the ruling
are the same observation from two directions.
THE FINDING. `assign`/`assigned_user_id` and `role`/`agent_role_id` are
two names for one target, exactly like `parent`/`plan` — and none of the
five guards standing at round 12 compared them. The alias guard knew
only about hierarchy; the compat guard only about same-name collisions.
So `assigned_user_id:"B"` with `fields:{"assign":"A"}` was accepted and
the doors disagreed: resolveAssignName gives the explicit ID precedence
over HTTP, while BuildCLIArgs drops the compat ID and emits `--assign A`.
One call, two different people assigned.
THE RULING. Stop adding guards. Conflict handling had accreted one at
every site that noticed a problem — generic path, promoted block, alias
check, compat block, canonicalization predicate — and each covered only
the sources its author thought about. Round 13's finding is that shape's
signature, not a new one.
WHAT CHANGED. `detectFieldConflicts` resolves every source (the `fields`
object, the `field:[]` entries, the promoted params, the compat IDs) to
a canonical key via `fieldAliasGroups`, then refuses on the resulting
map. Alias collision refuses even when the values match — the names
address one target through different vocabularies (a slug vs a UUID), so
"equal" is not a question this layer can answer. Same name from two
sources keeps the old rule: equal collapses, differing refuses. The five
guards are gone; the per-key branches now do emission only, and they run
knowing the input is unambiguous.
A new alias pair is one line in a map rather than a sixth guard.
SCOPE, unchanged and stated: this runs only when a `fields` object is
present, because reshapeItemFields does. A top-level param colliding
with a `field:[]` entry and no `fields` object keeps its documented
last-write-wins resolution — the round-7 boundary the lead confirmed,
still pinned on both doors.
EVIDENCE. All 30-odd existing case tests pass unchanged against the new
structure; they are the regression net the ruling asked to keep. Added:
the round-13 case tests over both directions of both pairs, and a
PROPERTY test derived from `fieldAliasGroups` itself — for every alias
class, every ordered pair of member names, and every pair of distinct
sources, the call refuses. It covers 48 combinations and a future alias
pair extends it with no edit.
Mutation matrix:
drop assigned_user_id from the alias map -> property fails
unwire detectFieldConflicts entirely -> 54 tests fail
remove the alias-collision branch -> property fails (equal-value legs)
The third mutant SURVIVED at first, and the reason is the useful part:
my property used differing values, so the ordinary same-canonical-key
comparison refused anyway and a build with alias detection wholly
removed still passed. The equal-value leg is the one only alias
detection catches — exactly the semantics the parent/plan ruling
established — so the property now drives both value shapes. CONVE-28
again: on SURVIVED, the mutant was faithful and the test was weak.
gofmt clean · go vet clean · go test ./... green (29 packages)
context: 45.1% (session-shape)
|
||
|
|
af686c350d |
fix(mcp): a blank top-level param does not block the fields answer (BUG-2850)
Codex round 12, one P1 — and the fix is bigger than the finding, twice
over.
THE FINDING NAMED ONE KEY. `{status: "", fields: {status: "done"}}` was
refused, because the duplicate check treated a present-but-empty
top-level param as a competing value. `""` is "not supplied" everywhere
else on this surface — promotedParamValue treats it as absent, the CLI's
`status != ""` guards do, `assign: ""` is documented inert — so a client
that zero-fills its optional params was refused for asking one question.
Driving the whole class the key belongs to (CONVE-18: the reviewer names
an instance, the fix owes the population) turned `status` into all six
promoted keys, which behave identically. Round 10 fixed exactly this for
the hierarchy keys and I never asked whether the same reasoning covered
their siblings; it did. Third time in this unit that a guard was written
for the path in front of me, and this time the guard was mine.
THE SAME PROBE FOUND THE EXCEPTION, which matters more than the fix. For
`assigned_user_id` / `agent_role_id` an empty string is NOT absence — it
is a CLEAR to NULL, the deliberate v0.16 semantics
(dispatch_http_advanced.go forwards "" verbatim for exactly these two).
Applying the finding uniformly, as its wording invites, would have
discarded a clear in favour of the `fields` value: a spurious refusal
traded for a silent wrong write, which is the worse half. They stay
conflicts, with their own pin.
AND THEN A SURVIVING MUTANT CAUGHT A FALSE COMMENT OF MINE. Deleting the
compat carve-out failed nothing — because the carve-out lived in a
helper that only the PROMOTED block called, while the compat keys were
checked in a separate block that never consulted it. Dead code, and the
comment I had just written called it "the whole reason this is a
function". CONVE-28's rule is what caught it: on SURVIVED, ask whether
the mutant is faithful before blaming the test. The mutant was faithful;
the code was redundant and the prose was wrong. Both call sites now
consult the predicate, so the rule has one home and the carve-out is
live — re-running the same mutant against the corrected code fails
BlankCompatIDIsAClearAndStillConflicts, as it always should have.
Mutation matrix:
revert the blank-param exclusion -> only BlankTopLevelParamDoesNotBlockFields fails (6/6 subtests)
delete the compat carve-out -> only BlankCompatIDIsAClearAndStillConflicts fails (2/2)
[survived before the redundancy was removed — see above]
gofmt clean · go vet clean · go test ./... green (29 packages)
|
||
|
|
c2bb1bad34 |
fix(mcp,server): close three codex round-11 findings, one of them my own bad refutation (BUG-2850)
Two P1 and one P2. The P2 is the important one, because I had already
dismissed it in round 10 and was wrong.
[P2 — CORRECTION] Artifact imports DO drop undeclared-field warnings,
and the case is reachable. Round 10 refuted this on the grounds that
artifact.Decode populates Fields only from FieldKeysForKind, so no
undeclared key could arrive. That check was real, and it was the WRONG
SIDE of the comparison: UndeclaredFieldKeys compares the field map
against the DESTINATION COLLECTION'S SCHEMA, not against the artifact
format's key list. The destination schema is editable, so a canonical
artifact key can be undeclared THERE while being perfectly legal in the
artifact.
Verified before reinstating, not argued: narrow the conventions
collection's schema to declare only `status`, import an ordinary
convention carrying trigger/scope/priority — the blob stores all three
and UndeclaredFieldKeys names all three. The merge is back, and the
comment now records the correction rather than the refutation, so the
next reader inherits the right reason.
What I got wrong is worth naming exactly: I verified a true fact and
then drew a conclusion one step wider than it supported, because I never
asked what the OTHER operand of the comparison could be. "No key outside
the artifact's list arrives" does not imply "no undeclared key arrives"
unless the destination declares every key on that list — an assumption I
never stated and never checked. The test I wrote at the time could not
pass, and I read that as confirming the refutation instead of as the
setup being wrong.
[P1] An empty hierarchy value inside `fields` was a silent no-op.
`fields:{"parent":""}` promotes onto the top-level `parent`, where both
doors treat empty as NOT PROVIDED — so the call reported success and
detached nothing. Refused now, pointing at clear_parent and the raw
`field:["parent="]` form. Not silently promoted to a clear: that decides
what this door MEANS, and v0.19 already made clear_parent canonical so
the empty string would not have to carry it.
[P1] The v0.16 compat ID params were not conflict-checked against
`fields`. Never schema-declared by design, they are invisible to
padItemPromotedFieldKeys and took the generic path, where the check only
consults the `field` array. `assigned_user_id:"A"` with
`fields:{"assigned_user_id":"B"}` made the doors disagree outright — the
remote mapper reads A, stdio emits only `--field assigned_user_id=B`,
because the top-level form has no CLI flag behind it. One call, two
different people assigned. Conflicting values refuse; equal ones
collapse to one form.
Also corrected, per CONVE-23: the round-10 test asserting that
`fields:{"parent":""}` conflicts with the plan alias now refuses for a
DIFFERENT reason and its stated rationale had become false. The case
moved to the new test with the right reason, and the field-array clear —
which really is an effective directive — stays where it was, with a
control leg proving the new refusal did not swallow it.
Mutation matrix, each mutant from a file backup:
revert the empty-hierarchy refusal -> only EmptyHierarchyValueInFieldsRefused fails (both subtests)
revert the compat-ID check -> only CompatIDConflictRefused fails (both subtests)
revert the import warning merge -> only the narrowed-schema import test fails
gofmt clean · go vet clean · go test ./... green (29 packages)
|
||
|
|
0a71ad3aa9 |
fix(mcp): an empty parent param is not a hierarchy directive (BUG-2850)
Codex round 10: three P2, no P1. One fixed, one already filed, one
refuted and reverted.
[FIXED] An empty top-level `parent` was counted as an alias directive,
so `parent: ""` with `fields:{"plan":"X"}` refused a perfectly good
call. Every declared string param on this tool treats "" as NOT
PROVIDED — it is why promotedParamValue does, and why `assign: ""` is
deliberately inert — so a client that fills declared optional params
with their zero value rather than omitting them got a refusal for
asking one hierarchy question. My own round-9 snapshot carried this
forward from the out[]-based check it replaced.
Deliberately NOT applied to the other empty forms: `field:["parent="]`
and `fields:{"parent":""}` are the documented CLEAR signal
(BUG-2013 / BUG-2078), so they are semantically effective and still
conflict. One is a param left blank, the other is an instruction that
happens to look like one. Both directions are pinned, and the mutation
matrix drives both:
remove the empty-param exclusion -> only EmptyParentParamIsNotAnAliasConflict fails
extend it to the field-array clear -> only EmptyClearFormsStillConflict fails
[ALREADY FILED] The transport-dependent whitespace finding is BUG-2870,
ruled out of this PR's scope. Round 10 did add something the filing
missed and BUG-2870 now records it: the divergence covers VALUES too
(`--field "cost= 3"` stores the number 3 remotely and the string " 3"
over stdio), which is worse than the key half because both doors report
success and only the stored type differs.
[REFUTED, REVERTED] "Artifact imports discard createItemChecked's
undeclared-field warnings." True as a code reading — this handler builds
its own response shape and ignores item.Warnings — but the condition is
unreachable. artifact.Decode populates Fields exclusively from
FieldKeysForKind via a closed switch over a typed frontmatter struct, so
a key outside that per-kind list never enters the map. Verified both
ways before reverting, because the import door takes raw bytes and the
hand-written case is the one that mattered: Encode drops extra keys, and
Decode of a hand-written artifact carrying extra frontmatter keys drops
them too.
I had written the merge and a test for it before checking; the test
could not pass through the public door, which is what exposed the
finding rather than my fix. Reverted to a comment recording the
mechanism, so the next reader — or the next round — does not re-find it.
A branch nothing can enter is not defence in depth, it is a claim that
something is handled when it never happens.
gofmt clean · go vet clean · go test ./... green (29 packages)
|
||
|
|
56ee3a7e95 |
fix(mcp): require strings for fields.assign/role; fix the alias refusal's mechanism (BUG-2850)
Codex round 9: one P1 and one P2. The P1 was REFUTED on inspection and
the P2 confirmed; both produced a change, for different reasons.
[P2, real] `fields.assign` / `fields.role` accepted a number, and the two
doors then disagreed about it. The HTTP dispatcher's
`rawAssign.(string)` turns a float64 into "" and treats it as NOT
PROVIDED, silently dropping the write; stdio emits `--assign 123` and the
CLI fails loudly on the lookup. Same call, one door silent and one red.
Refused now at the door-independent layer, which is what stops them
drifting apart again rather than teaching each dispatcher separately.
Deliberately narrow: this does NOT walk back round 6's decision to accept
non-string promoted values in general. `priority` may legitimately be a
number in a custom schema and create has always passed such values
through — a control leg pins that. `assign` and `role` are references
that NAME something, where a number has no meaning at all.
[P1, refuted] `fields:{"parent":"A","plan":"B"}` was already refused. But
it was refused by ACCIDENT: keys process in sorted order, so `parent` was
promoted into out["parent"] and `plan` collided with it one iteration
later. Right answer, wrong mechanism — the refusal depended on `parent`
sorting before `plan` AND on `parent` being a promoted key, and it told
the caller their value conflicted with "the top-level parent param" when
no such param was passed. The fields-vs-fields case is now checked
against `obj` directly, and a snapshot of the original top-level params
keeps that message honest.
Reported as verified rather than as agreement: the finding's mechanism
was wrong, and shipping "fixed" against a refuted claim would have put a
false statement on the trail.
Mutation matrix, from file backups:
revert the identity-ref requirement -> only NonStringIdentityRefRefused fails (both subtests)
revert to the out[]-only alias check -> only BothAliasesInOneFieldsObject fails
Plus a PROBE that is not a mutant of the fix: removing `parent` from
padItemPromotedFieldKeys leaves the alias pair still refused. Under the
old code that mutation made the guard go silent, since nothing would
write out["parent"] — which is the latent coupling this change removes.
gofmt clean · go vet clean · go test ./... green (29 packages)
|
||
|
|
49e533d478 |
test(mcp,cli): pin same-name duplicate precedence on both doors (BUG-2850)
The lead's condition on the round-7 boundary. checkHierarchyAliasAmbiguity refuses parent+plan — two NAMES for one target, which a caller can collide without knowing — but deliberately does NOT refuse a same-name duplicate (`--status A --field status=B`), because those are visibly duplicates and both doors resolve them identically. "Both doors resolve them identically" is the load-bearing half of that argument and nothing enforced it. Two tests now do, one per door, asserting the SAME outcome: the `field` entry overlays the named param, because cmd_item.go and dispatch_http_advanced.go both apply named flags first and overlay --field after. Per-door mutation matrix, run this turn from file backups: make the named param win on the HTTP door -> only the mcp test fails make the named flag win on the CLI door -> only the cmd/pad test fails Neither mutant reddens the other door's test, which is the property worth having: the doors cannot drift apart again without exactly one of these going red and the boundary getting re-examined rather than silently becoming untrue. Also filed, per the lead's ruling: BUG-2870, the padded-`field`-key divergence with NO `fields` object (`--field " effort=l"` stores an undeclared " effort" key on the CLI door and writes `effort` on the remote one). Out of scope here — it predates this PR's claim rather than defending it — and its fix is a policy call on the CLI's input contract, so it wants a ruling, not a quick patch. gofmt clean · go vet clean · go test ./... green (29 packages) |
||
|
|
4937fd84f6 |
fix(mcp): canonicalize when ANY entry for the key is padded (BUG-2850)
Codex round 8, one P2 and no P1 — the first round of this unit that did
not turn up a correctness defect on the fields-vs-field seam.
Round 7's canonicalization asked whether a canonical entry was PRESENT
and left the array alone if one was. So `field:["effort=l", " effort=l"]`
with `fields:{"effort":"l"}` kept the padded twin, and the doors then
disagreed about it: HTTP trims and writes `effort`, the CLI does not and
writes an undeclared `" effort"`. Transport divergence out of a call both
doors accept — the shape this unit exists to remove, reintroduced one
round earlier by the fix for its sibling.
The predicate is now "any entry for this key is non-canonical", so the
key is re-emitted once and cleanly. Collapsing the duplicate pair is not
lossy: parseFieldArray already indexes both to a single value, so two
entries for one key were never two writes.
Mutation: restoring the round-7 predicate verbatim fails only
MixedCanonicalAndPaddedDuplicatesCollapse. Round 7's two pins still pass
under that mutant, which is correct — neither exercises the mixed case,
and that is exactly why the new one was owed.
NOT fixed here, and named so it is not mistaken for an oversight: a
padded entry with NO `fields` object at all (`field:[" effort=l"]` alone)
still reaches the CLI door untrimmed. That predates BUG-2850, is
unrelated to the fields merge, and normalizing every entry
unconditionally changes what the CLI receives for every caller — a
policy change, not a defect fix. Flagged to the lead on the trail.
gofmt clean · go vet clean · go test ./... green (29 packages)
|
||
|
|
13892fecf7 |
fix(mcp): close two codex round-7 findings on the same seam (BUG-2850)
Both are consequences of round 6's own fixes, which is the tell that the
seam — a guard written for one key shape, and the keys that do not take
that path — is still the thing to keep hitting.
[P1] The alias guard did not fire without a `fields` object.
Round 6 put it inside reshapeItemFields' per-key loop, and
reshapeItemFields returns early when `fields` is absent — so
`field:["parent=A","plan=B"]` walked straight past it and
extractParentLink's no-early-exit loop applied `plan` while the caller
had every reason to believe `parent` was what they set. A guard a caller
can step around by moving the same two values into a different param is
not a guard. checkHierarchyAliasAmbiguity now runs on create and update
regardless of `fields`, over the merged input.
SCOPE, stated rather than smuggled: the pure-`field` form was accepted
before BUG-2850 too, so this closes a pre-existing silent mis-write, not
a regression. Fixed here rather than filed because shipping round 6's
guard without it would advertise a refusal that any caller bypasses in
one edit. Deliberately NOT extended to same-name duplicates (a `parent`
param plus `field:["parent=B"]`) — those resolve last-write-wins
identically on both doors, which is documented behaviour, and widening
the refusal to cover them is a policy change rather than a defect fix.
[P2] A padded equal duplicate was retained raw. Round 6 normalized the
conflict INDEX so ` effort=l` matches `fields:{"effort":"l"}` — correct,
and it closed the padded-key bypass — but the raw entry stayed in
`field`, and the CLI door does not trim. Over stdio that stored an
undeclared `" effort"` key and left `effort` untouched: the
normalization that made the duplicate visible is what made the retained
entry wrong, so the fix belongs at the same place. The entry is now
re-emitted canonically, and only when the raw form actually differs, so
a well-formed array keeps its contents and its order.
Mutation matrix, run this turn, each mutant from a file backup:
unwire checkHierarchyAliasAmbiguity -> only the round-7 alias test fails (4/4 subtests)
revert the canonical re-emission -> only PaddedEqualDuplicateIsCanonicalized fails
Neither mutant touches round 6's in-loop alias test, which is right: that
one enters through the `fields` object and is a different path — the
distinction this finding exists about.
Control legs: a lone hierarchy key through the array still dispatches,
an already-canonical duplicate is left byte-identical and unreordered,
and the padded-entry case asserts exactly one --field is emitted.
gofmt clean · go vet clean · go test ./... green (29 packages) · 18 files
|
||
|
|
dfee13896d |
fix(mcp): close four codex round-6 findings in the fields-object merge (BUG-2850)
All four sit on the same seam this unit keeps failing at: a guard written
for one key shape, and a class of keys that does not take that path.
[P1] parent/plan alias conflicts bypassed every guard. extractParentLink
resolves the hierarchy link with `for _, key := range {"parent","plan"}`
and no early exit, so when both arrive the LATER key wins — but every
check in reshapeItemFields matched on the SAME key name. So
`fields:{"parent":"PLAN-12"}` with `field:["plan=PLAN-9"]` passed and
relinked the item to PLAN-9 while reporting PLAN-12. The same alias
bypass BUG-2078's round-1 review found on clear_parent, reached through
a different door. Refused now in both directions and against the
top-level param — and refused even when the two values are EQUAL, which
is what v0.19 already does for parent + clear_parent "including via the
plan alias".
[P1] The conflict index was not normalized the way the door normalizes.
ingestFieldKVP TrimSpaces both halves of a `key=value` entry;
parseFieldArray indexed the raw halves, so `field:[" status=cancelled"]`
sat under " status", missed the guard against `fields:{"status":…}` and
then silently overrode it. Trimming the value fixes the mirror-image
false refusal (`status= done` vs `done`). ONLY the index is normalized —
`entries` stay verbatim, because the CLI door does not trim and must
keep receiving exactly what the caller sent.
[P1] A non-string promoted value silently no-op'd on remote update.
reshapeItemFields promotes `fields:{"priority":3}` with its type intact,
but hasFieldChanges and the patch loop both read `.(string)` — so the
dispatcher skipped the fields_patch branch entirely and answered SUCCESS
having sent no PATCH. A silent no-op reintroduced by the fix for silent
no-ops, and asymmetric with create, which has always passed non-strings
through. promotedParamValue now accepts any scalar; empty string still
means "not supplied".
[P2] Equal promoted duplicates did not collapse. `fields:{"role":"x"}`
plus `field:["role=x"]` resolved the role to agent_role_id AND wrote a
literal `role` key into the fields blob that no schema declares — one
value, two writes, one of them an undeclared field with a warning
naming it. The array entry is now dropped so the value applies once
through its dedicated param.
Mutation matrix, run this turn, each mutant applied and reverted from a
file backup (never `git checkout` — the tests were uncommitted):
drop the alias block -> only HierarchyAliasConflictRefused fails (4/4 subtests)
revert index normalization -> only FieldArrayKeysNormalizedForConflicts fails (both legs)
revert the duplicate drop -> only the two EqualDuplicate tests fail
revert promotedParamValue -> only NonStringPromotedValueIsNotDropped fails
No cross-talk: each mutant is the defect at the site its test targets,
and each kills exactly that test. Control legs included on purpose — a
lone hierarchy key is still accepted, padding-only value differences are
not conflicts, an unrelated `field` entry survives the duplicate drop,
and an empty promoted string still produces no fields_patch.
gofmt clean · go vet clean · go test ./... green (29 packages)
|
||
|
|
7f25283a41 |
test(server): pin the last three coercion call sites (BUG-2850)
Five of the eight `CoerceFields` sites had a test that goes red if that site alone is dropped. Move, bulk move and bulk update did not, so the PR's "typed on every door" claim rested on reading the code — CONVE-19 and the shape rounds 2-5 of this unit kept finding. Why the move pins are faithful rather than green-for-free: migrateValue already permits text->number (migrate.go:190), but it returns `value`, the ORIGINAL, not a parsed float. So a text field holding "42" reaches a number-typed destination as the STRING "42", and only CoerceFields at the move site turns it into a number before validation. Assertions are on the STORED NATIVE TYPE, re-read from the item rather than taken from the mutation's own response, so a handler that answered 200 and stored the string is still red. Bulk update merges request STRINGS (status, priority), so it is observable only where the schema declares one of those keys as a non-string type; a collection declaring `priority` as a number is unusual but legal and is the honest way to reach that site. Bulk ops answer 200 with per-item failures in the envelope, so the pins read the envelope too — a status-code-only assertion would pass on a dropped coercion. Per-site mutation matrix, run this turn against these tests, each mutant applied and reverted from a file backup (never `git checkout`, which would have taken the uncommitted tests with it): drop coercion at handlers_items.go:2314 -> only TestItemFieldsCoercedOnMove fails drop coercion at handlers_items_bulk.go:683 -> only TestItemFieldsCoercedOnBulkMove fails drop coercion at handlers_items_bulk.go:499 -> only TestItemFieldsCoercedOnBulkFieldUpdate fails Each mutant is the defect at the site the test targets — not a call-site patch next to a still-correct function (CONVE-28) — and each kills exactly one test, which is the per-site discrimination the PR claims. Files restored and verified identical after the matrix; suite green. |
||
|
|
baaa236d83 |
fix(mcp): apply the field-array conflict guard to promoted keys too (BUG-2850)
Codex round 5, one P1. The `fields` object vs `field: ["k=v"]` conflict
guard lived on the generic path, which a promoted key never reaches:
`status`, `priority`, `category`, `parent`, `role`, `assign` and `tags`
all return from the promoted branch above it. So the one ambiguity this
function did not refuse was the one on the keys that matter most.
It did not fail closed either. The promoted branch writes the top-level
param (`out["status"]`) while the array entry stays in `out["field"]`,
and both the HTTP mappers and the CLI overlay `--field` entries AFTER
the named flags — so the array silently won. `fields:{"status":"done"}`
with `field:["status=cancelled"]` cancels the item. The same shape on
`parent` relinks or detaches it.
The guard now runs first in the promoted branch, before the existing
top-level-param check, with the generic path's semantics: differing
values refuse, an equal duplicate falls through and still promotes, and
a structure against a string entry (the `tags` case) refuses as
"one key cannot be both".
This is the fourth consecutive round where the defect was a guard
written on the generic path only, and round 4's test is why: it pins
`effort`, a key that takes the generic path, so it vouched for the path
the guard is on rather than for the class of keys that skips it
(CONVE-19). The new test drives all four promoted shapes plus an
equal-duplicate control leg.
Negative control: all four conflict cases fail on the unfixed tree
(run before the fix, not against a synthetic mutant); the
equal-duplicate leg passes both before and after, so it discriminates
refuse-on-ambiguity from refuse-on-agreement.
gofmt clean · go vet clean · go test ./internal/mcp/ ok
|
||
|
|
b23c0dbc6f |
fix(mcp): apply the null and hierarchy guards to promoted keys too (BUG-2850)
Codex round 4, and both findings are the same defect in my own round-3 fix:
ORDERING. The null guard and the hierarchy-key guard sat BELOW the branch
that promotes status/priority/category/parent/role/assign/tags onto dedicated
params, so every promoted key walked around both.
- `fields: {"tags": null}` reached the promoted branch and became a silent
no-op, instead of the refusal the null rule documents.
- `fields: {"parent": 42}` was accepted here and dropped later by the
handler — the same silent-drop shape this whole bug is about, reintroduced
by a guard I added to prevent a different instance of it.
Both guards now run before that branch, so they apply to every key. A guard a
whole class of keys bypasses is not a guard.
Pinned separately from the generic-path cases, because the generic-path tests
passed throughout: they never exercised a promoted key, which is exactly why
the hole survived round 3. Reverting the hoist fails the new test. Well-formed
promoted keys still promote — tested, so the hoist did not break promotion
while closing the bypass.
Gates: gofmt clean, go vet clean, go test ./... 29 packages ok.
Claude-Session: https://claude.ai/code/session_011Q4b1iHtJtSyMs7BA2ySxo
|
||
|
|
f15f60831e |
fix(mcp): guard hierarchy pseudo-keys and correct the fields description (BUG-2850)
Codex round 3.
1. [P1] A structured `plan` could silently DETACH an item. `plan` is not in
padItemPromotedFieldKeys, so it fell to the generic path — and once this
branch stopped refusing structures, fields:{"plan":{…}} reached the server
natively. There extractParentLink reads any PRESENT non-string plan/parent
as a hierarchy directive, drops the key, and on update clears the existing
parent link. Lifting the nested refusal quietly opened a path where a
malformed value detaches an item from its parent.
Guarded specifically: these keys take a string ref and nothing else. Both
`parent` and `plan` are listed, so the guard does not depend on which
other set a key happens to belong to. A string ref still works — tested,
so the fix did not re-refuse the normal case while closing the hole.
2. [P2] The MCP `fields` param description still told agents that
multi_select and json non-scalars are "refused, not written". This diff
makes that false, and a schema an agent reads is the artifact that decides
what it attempts — the reporter's agent rewrote seven playbooks after
believing exactly this kind of line. Rewritten to state what is true now,
including the two remaining refusals (null, structured parent/plan) and
that structured values need the remote transport.
Gates: gofmt clean, go vet clean, go test ./... 29 packages ok. Removing the
hierarchy guard fails its test.
Claude-Session: https://claude.ai/code/session_011Q4b1iHtJtSyMs7BA2ySxo
|