mirror of
https://github.com/PerpetualSoftware/pad.git
synced 2026-09-23 11:03:41 +00:00
25c7cd20f5
* feat(watchevents): Redis-backed bus so watch notifications cross instances (BUG-2651) internal/watchevents shipped MemoryBus only, so in a multi-instance deployment a notification published on instance A never reached a stream held open on instance B — watches appeared to work and silently dropped. Bus was an interface from day one for exactly this; adding RedisBus changed no producer and no consumer. NOT A MECHANICAL PORT of internal/events.RedisBus. Three deliberate divergences, each documented at the point someone diffing the two files would call it a mistake: - ONE channel and ONE replay buffer, because this package has exactly one logical stream by contract (DOC-2479 DR-2: all per-caller filtering happens in the consumer). Most of the template's bookkeeping — per- workspace counts, subscriptions, buffers — has nothing to key on here. - EAGER subscription for the bus's lifetime, not lazily on first local subscriber. The replay buffer fills from the RECEIVE path, so a lazily torn-down subscription stops filling it at precisely the moment before a Last-Event-ID resume — for one harness monitor holding one stream, that makes resume structurally useless. The template can afford lazy because per-workspace means N idle subscriptions; here it is one. - ONE mutex across subscriber membership and the replay buffer, held through the whole local fan-out. The template uses two and offers only separate Subscribe + EventsSince, which cannot provide SubscribeAndReplaySince's guarantee. Copying its locking would have handed back the double-delivery window this package's interface exists to close. Publish fails CLOSED when INCR fails, where the template falls back to a local counter. Two instances falling back at once mint ids from independent counters into a shared stream, and replayBuffer.since() reasons on monotonicity — so the damage is silent replay corruption, not a visible error. INCR and PUBLISH share a connection anyway, so the fallback mostly lets a doomed publish proceed carrying a poisoned id. Both load-bearing tests were VACUOUS as first written; the mutation matrix is the only reason I know: - the concurrency test's producer finished before the subscriber joined, so the channel leg was never exercised and a split-lock mutant survived 50 iterations. Now paced, with a both-legs-non-empty precondition that fails a run which never approached the boundary, plus a dedicated detector (600 attempts, 8/8 kills, 0.02s after switching the drain to non-blocking — exact, because the duplicate is already buffered when the call returns). - the fail-closed test asserted nothing was delivered, which is true of the fallback too: Publish never delivers locally, so with Redis down neither policy delivers. Rewritten around a go-redis ProcessHook that records attempted commands, which is where the policies actually differ (INCR-then-stop vs INCR-then-PUBLISH). Also corrects session_presence.go, which told the next person these two had to be fixed together. Delivery is now cross-instance; the registry's under-report is unchanged, so the remaining defect is a picker that under-reports rather than a push that lies. The PLAN-2558 S3 gate stays, for that reason instead of the old one. Claude-Session: https://claude.ai/code/session_017jD6t1zjxGSq47SQpZfp1V * fix(watchevents): make id assignment and publish atomic; close the bus on shutdown (Codex round 1) P1 — INCR and PUBLISH as two client calls are not order-preserving, and the failure is concrete: A gets id 1 and is descheduled, B gets id 2 and publishes, A publishes 1. Every subscriber receives 2 before 1, the replay buffer appends in ARRIVAL order, and replayBuffer.since() reasons on monotonicity — so a resume from 2 hits the sinceID > newestID branch and answers 'gap too large', turning a healthy reconnect into a spurious sync_required, while a resume from 1 silently skips the late arrival. Fixed at the source with a Lua script: Redis runs it atomically on its single thread, so INCR and PUBLISH for one instance both complete before another's script begins, and publish order equals id order globally with no coordination on our side. The id rides as a '<id>|<json>' prefix rather than being edited into the JSON from Lua; the id is digits and the FIRST '|' separates, so a '|' in the body is unambiguous. A pleasant consequence: there is no longer a window where an id exists but the publish has not happened, so the fail-closed decision and the publish decision became the same decision. P2 — Stop() never closed the watch bus. That was survivable for MemoryBus, whose Close only drops channels; RedisBus holds a receive goroutine and a Redis subscription from construction, so it leaked both for the process's life. Closed after bg.Wait(), so a background producer cannot publish into a bus already tearing down. nits, all real, all in artifacts someone reads: - 'exactly-once delivery' was simply wrong. Redis pub/sub is at-most-once and the local send is deliberately non-blocking. The property the round trip actually buys is NO DOUBLE DELIVERY to the publishing instance; the comment now says that and names the replay buffer as the bounded recovery mechanism for the rest. - the Bus interface comment still said only MemoryBus existed. - cmd_server.go's session-presence note still claimed the same caveat as 'the watch bus directly above', which had just stopped applying. - session_presence.go now says delivery is fixed WHEN PAD_REDIS_URL is set, rather than unconditionally. Tests: the fail-closed assertion moved from 'nothing was delivered' — still true under the two-call version — to 'no bare INCR or PUBLISH was issued', which is what distinguishes atomic from not. Mutation-verified by splitting the script back into two calls. Added a decode round-trip test covering the new wire format, a '|' inside the body, and four malformed payloads, since that decoder consumes bytes from a channel any holder of the Redis credentials can publish to. Claude-Session: https://claude.ai/code/session_017jD6t1zjxGSq47SQpZfp1V * fix(watchevents,server): correct the targeted-push claim; close the bus before HTTP shutdown (Codex round 2) P2 — I claimed cross-instance DELIVERY was fixed. Half true, and the false half was mine to catch: handlers_push.go gates a session-targeted push on the LOCAL presence registry and skips the publish entirely when the id is not there, so a POST landing on A for a session held on B still delivers nothing. The bus would carry it; the gate means it never reaches the bus. Broadcast pushes and every other notification kind ARE fixed. I asserted that behaviour from reading the bus and session_presence.go without reading the push handler — the exact thing I hold myself to not doing. Corrected in all three places the claim was made (the package doc, session_presence.go, and the KindPush comment), with the correction recorded rather than quietly overwritten. The gate's own justification is now stale too, and worth more than a tweak: 'a target this instance cannot see is a guaranteed no-op' was TRUE under MemoryBus and is FALSE under RedisBus, where another instance may hold that session. Left in place deliberately — publishing unconditionally would fix delivery and immediately make delivered_sessions=0 a lie in the other direction, which is a question about what that field promises. It belongs with the shared-state SessionPresence that PLAN-2558 S3 already gates on: fixing the registry makes the snapshot right, and then the skip is correct again for its original reason. Both open halves collapse into that one implementation. P2 — the watch bus was closed only in Server.Stop(), which runs AFTER http.Server.Shutdown. The event bus is closed before Shutdown precisely so its SSE handlers unblock; the watch stream is the same shape, so an open one would have held Shutdown to its full 30s deadline. Now closed alongside eventBus, with the Stop() close kept as the path for other callers — both implementations are idempotent. nit — MemoryBus and RedisBus disagreed after Close: RedisBus handed a late Subscribe an already-closed channel, MemoryBus registered one nobody would ever close, so a consumer racing shutdown blocked forever. MemoryBus now matches, and its Close is idempotent, which the CLI's double close relies on. Claude-Session: https://claude.ai/code/session_017jD6t1zjxGSq47SQpZfp1V * fix(watchevents): report a missed notification as a replay gap (Codex round 3) P2 — a divergence MemoryBus structurally cannot have. It assigns every id itself, so its replay buffer is contiguous and the only gap it can report is eviction. RedisBus receives ids over at-most-once pub/sub, so a blipped subscription can miss 101 and receive 102: the buffer holds a hole, is nowhere near full, and replayBuffer.since() answers a resume from 100 with just [102]. The consumer loses a nudge and is never told. RedisBus now tracks the id at which the sequence resumed after the most recent hole, and answers nil — the same signal eviction already gives, which the SSE handler already turns into sync_required — for a resume that would have to span it. Resumes that do not span it still replay normally, and sinceID=0 is treated as a fresh subscriber rather than a resume, so a hole nobody spanned is not turned into a spurious resync. The atomic publish script is what makes this readable: publish order is id order globally, so a non-consecutive id means MISSED, not reordered. Mutation-verified by disabling the check; the test fails on both the spanning resumes and would have failed the over-broad version too (it asserts the non-spanning resumes still work). Two residuals documented rather than fixed, both because the fix is the same shared-state SessionPresence that PLAN-2558 S3 gates on: - delivered_sessions is now wrong in BOTH directions for a broadcast push — the count is local while delivery is global, so a replica can report 1 while two sessions receive it, or 0 while a remote one does. No local arithmetic fixes that; it is asking one replica what all of them are doing. - the Redis channel and counter names are not deployment-scoped, so two installations sharing a Redis endpoint cross-feed (and picking different logical DBs does not help — pub/sub ignores them). Left flat to match internal/events rather than giving one of the two buses a prefix the other lacks; the rule is one Redis endpoint per installation, and relaxing it should cover both buses at once. Claude-Session: https://claude.ai/code/session_017jD6t1zjxGSq47SQpZfp1V * fix(watchevents): a cold-started replica must report a gap too (Codex round 4) P1 — the round-3 hole check only fired BETWEEN two received messages, so it never fired for the first one. A replica restarting while Redis is already at 101 has an empty buffer; its first received message is 102, nothing looks like a hole, and a client reconnecting to that replica with Last-Event-ID 100 was handed [102] — skipping 101 exactly as silently as the case round 3 fixed, by a different route. Replaced contiguousFrom with knownFrom: the lowest id from which this instance's buffer is contiguous. SET on the first append (before which this instance knows nothing) and RESET on every hole (before which it no longer knows anything usable). One variable, both failures. The boundary is pinned in both directions, which is what stops this being an over-broad 'always gap after a restart': a resume from exactly the id before our first (101 when we started at 102) IS contiguous with our view and replays normally. Mutation-verified by disabling the cold-start arm. Claude-Session: https://claude.ai/code/session_017jD6t1zjxGSq47SQpZfp1V * fix(watchevents): idempotent publish, confirmed subscription, and real Redis tests (Codex round 5) P2 — go-redis retries a command whose reply is lost to a network error, and the publish script was not idempotent: the same notification would be published twice under two different ids. Both copies look valid — ordered, distinct — so nothing downstream could tell them apart, and on the push path a duplicate is a duplicate DISPATCH into an agent harness. The script now takes a caller-generated token and SET NX's it, so a retry carrying the same arguments returns 0 without publishing. TWO THINGS THIS UNIT OWES ITS TESTS, both found within minutes of each other and both invisible to the hermetic ones: 1. The idempotency script shipped indexing ARGV[3] while Publish passed two arguments. Caught by re-reading, which is not a control worth relying on for the next Lua edit. 2. NewRedisBus returned before go-redis had established the subscription, so notifications published in that window were lost to this instance, silently. Surfaced as a test flake; the production shape is a rolling deploy, where a replica takes traffic before its subscription is live. The constructor now waits for the confirmation (bounded, and a failure is logged rather than fatal since Channel() re-subscribes on reconnect). So miniredis is now a test dependency, and the round-trip tests it enables cover what fanOutLocally-driven tests structurally cannot: the channel name, the KEYS/ARGV mapping, the id prefix wire format, the shared counter across two buses, cross-instance delivery (the actual bug), the dedupe token, and Close tearing down the SERVER-side subscription rather than just local channels. Verified by restoring the ARGV[3] bug: the round-trip test fails on it. The two findings I am NOT fixing here are unchanged and documented where the reasoning is met — the targeted-push gate and delivered_sessions are both consequences of the per-process presence registry, and both are closed by the shared-state SessionPresence that PLAN-2558 S3 gates on, not by anything in this package. make vuln: 0 vulnerabilities in imported packages. Claude-Session: https://claude.ai/code/session_017jD6t1zjxGSq47SQpZfp1V * fix(watchevents): survive a Redis counter reset without replaying stale ids (Codex round 6) P2 — pad:watchevents_seq has no TTL but can still vanish: evicted under maxmemory, dropped by a FLUSHDB, or restored from an older snapshot. Ids then restart at 1 while this instance's ring still holds the hundreds. Keeping both is what corrupts replay — the two id spaces are not comparable, so a resume from 2 in the NEW space would be handed the stale 99/100/101 as though they were newer. A backwards id now drops the replay buffer and re-anchors knownFrom. Every resume from the old space then exceeds the newest id held and gets nil — the resync signal that is the only honest answer once the ids stopped meaning what the client thinks they mean — while clients in the new space keep working immediately. The test asserts BOTH halves, which is what makes it a detector rather than a description: a build that logged the reset and kept the buffer passes 'the old resume reports a gap' and fails 'the new resume never returns a pre-reset entry'. Mutation-verified on exactly that. Hardened while I was here: the epoch-reset path REBUILDS the buffer at runtime, so a bus constructed with a non-positive replay size would have turned a counter reset into a panic (newReplayBuffer(0)'s first append indexes a zero-length slice) rather than a resync. The constructor now normalizes. MemoryBus has the same trap for a caller passing 0; left alone as pre-existing and off this path, but named in the comment rather than silently fixed or silently ignored. nit — this file's header still claimed there was no miniredis dependency and no round-trip coverage, which the previous commit made false. Claude-Session: https://claude.ai/code/session_017jD6t1zjxGSq47SQpZfp1V * docs(watchevents): actually correct the hermetic test header (Codex round 7) The previous commit's message claimed this fix. It did not contain it: the edit ran as one of two scripts in a single command, its assertion failed with a traceback, and the second script's success is what I read. The header kept saying there was no miniredis dependency and no round-trip coverage — both false since two commits ago, in the file a reader consults to find out what IS covered. That is the adjacent-success-signal failure exactly: a success line from the step next to the one I cared about. The tell was in the output and I walked past it, then asserted the change in a commit message. Recording it here rather than quietly fixing, because a commit that claims a change it does not make is worse than one that omits it. Verified this time by reading the file back and grepping for the stale phrases: zero. Round 7's other three findings are the documented residuals re-raised for the third time — the targeted-push gate, delivered_sessions, and the unnamespaced Redis keys. All three are dispositioned at the line a reader meets them, all three are consequences of the per-process SessionPresence registry or of matching internal/events' existing convention, and none is fixable inside this package. They stay open, on the record, and with the lead. Claude-Session: https://claude.ai/code/session_017jD6t1zjxGSq47SQpZfp1V * docs(watchevents,cli): correct pad push --help; document the reset-window residual (Codex round 8) nit, and the one that stings — cmd_push.go's Long help still said pushes go over the 'in-memory watch-events bus'. That is the text a user reads when they run pad push --help, and it has been false since this branch's first commit. I have a standing pre-push step to grep the artifacts a CONSUMER reads for exactly this, and I ran it as a code search (watchevents.New) rather than a prose search, so --help never came up. The help now distinguishes broadcast (reaches every instance) from session-targeted (still resolved against the handling server) and names the bug. P2 — the counter-reset handling fires when the first post-reset notification ARRIVES, so there is a window between Redis losing the counter and the next publish in which this instance still replays old ids to a reconnecting client. Documented as accepted rather than closed: nothing local can detect the reset earlier (the counter is in Redis and we learn of it by receiving something), and the two shapes that would — a GET per resume, or a background poller — put network I/O on a latency-sensitive path or spend a goroutine and a round trip per tick forever against a condition measured in years. The exposure is redelivery of notifications the client already has, bounded by the window and self-healing on the next publish. Claude-Session: https://claude.ai/code/session_017jD6t1zjxGSq47SQpZfp1V * fix(watchevents): a replica that has received nothing must not answer 'caught up' (Codex round 9) P1 — the coverage check was skipped entirely while knownFrom was still 0, so a bus that had received NOTHING answered any cursor with an empty-but-non-nil replay, which the SSE handler reads as caught-up. The scenario is a restart, not an exotic one: replica B comes up while Redis is at 100, id 101 is published before B's subscription is live, and a client reconnects to B with Last-Event-ID 100 before 102 arrives. B says caught-up, then delivers 102 live, and 101 is gone with nothing to tell anyone. The principle the code now follows: having received nothing is strictly LESS knowledge than 'contiguous from X', so it must produce at least as strong a signal. A non-zero cursor against an empty bus is a gap. Both sides pinned, because the over-broad version is a real risk here — answering every fresh connection with a resync would be its own bug. A sinceID of 0 is not a resume and still gets an empty replay rather than a gap. Mutation-verified on the new arm. Claude-Session: https://claude.ai/code/session_017jD6t1zjxGSq47SQpZfp1V * docs(watchevents,cli): name the trailing-gap and shutdown trades (Codex round 10) Two findings that are decisions rather than defects, so both are documented at the line where the reasoning is met and taken to the plan instead of being settled unilaterally after ten review rounds. P1 as reported — the TRAILING gap. Everything the coverage bookkeeping does reasons about what this instance HAS received; it cannot see a notification missed at the END of the sequence. Hold 100, miss 101 to a disconnect, and a client resuming from 100 before 102 arrives is told caught-up. The hole only becomes visible when 102 lands, which is too late for that connection. What would reveal it is a GET of the sequence key: a value above lastAppendedID means ids exist we never saw, and a value BELOW it reveals the counter reset documented last round — one mechanism, both open windows. It is not done here because it is product-visible in the other direction: INCR happens before the message propagates, so the counter legitimately runs ahead of every instance for microseconds after each publish, and a strict comparison turns ordinary in-flight traffic into spurious sync_required responses with no principled tolerance to pick. A resync is recoverable and a lost nudge is not, which is the argument for doing it — but that is a call about how chatty the resync path should be. P2 — closing the watch bus before Shutdown drains handlers means a push already in flight can publish into a closed bus and still return 200 with pushed:true. Closing after would instead hold every shutdown to its 30s deadline on any open stream. eventBus already makes the same trade the same way; naming it rather than inheriting it silently. The honest fix is Bus.Publish reporting the drop so the handler can, which is an interface change and a different unit. Claude-Session: https://claude.ai/code/session_017jD6t1zjxGSq47SQpZfp1V * feat(watchevents): close the trailing gap with a settle-window authority check (lead ruling) Lead's ruling on BUG-2651: a silently lost nudge is unbounded staleness, a spurious resync costs one redundant fetch, so the gap must not survive — and don't pick a magnitude tolerance, because the reason the counter legitimately runs ahead is in-flight propagation, which is TIME-bounded while a genuinely missed message never arrives. So the discriminator is time. On a resume (and only on a resume), read the shared counter: if it disagrees with this instance's high-water mark, wait out one settle window and read again. In-flight ids land during the beat and the resume proceeds normally; missed ones never do and the resume is answered with a gap. That converts an unprincipled 'how many ids behind is too many' threshold into a principled propagation bound. The same read also catches the counter having gone BACKWARDS, so the counter-reset window documented last round is closed by the same mechanism rather than needing its own — the arrival-time reset handling stays, because it is what repairs the instance's own state and what covers a bus with no reconnecting clients. Ordering matters and is documented at the call: the check runs WITHOUT the mutex (it sleeps and does network I/O, neither of which may happen inside the lock fan-out needs) and BEFORE subscribing rather than between subscribe and replay, which would reopen the double-delivery window SubscribeAndReplaySince exists to close. Nothing is lost by waiting first — fanOutLocally buffers regardless of subscribers. An unreadable counter falls back to local knowledge rather than failing closed: turning a Redis hiccup into a resync for every reconnecting client at once is a worse failure than the one being guarded against. EventsSince deliberately does NOT do this and says so — it is the local primitive the Bus interface already describes as being for tests and non-resuming callers, and making it sleep and hit the network would surprise every one of them. Five tests, each pinning a different half: the missed tail reports a gap; a current instance does NOT (the control that stops this being 'always resync'); an id arriving mid-settle is tolerated; an unreadable counter falls back; a fresh subscriber neither waits nor gets a gap. Mutation-verified twice — disabling the check, and removing the settle beat — each killed by the test that names it. Also filed at the lead's direction, so the two remaining cross-instance defects have tracked homes rather than only comments: BUG-2698 (targeted push resolved against local presence, plus the delivered_sessions inaccuracy — one shared-state SessionPresence closes both) and BUG-2699 (push returns 200 pushed:true for a dropped publish; Bus.Publish reports nothing, and fixing it is an interface change). Every disposition comment now cites its item. Claude-Session: https://claude.ai/code/session_017jD6t1zjxGSq47SQpZfp1V * fix(watchevents): compare two FRESH reads, not one stale snapshot (Codex round 11) P1 — the settle beat re-read only the local side, so the comparison was against a counter SNAPSHOT taken before the wait. Id 2 arrives during the beat while id 3 is published and missed: the stale remote is still 2, the check declares convergence, and 3 is silently lost — the exact failure this whole mechanism exists to prevent, reintroduced inside it. P2 — the same staleness in the other direction. A GET can land just before a publish completes and report a value BELOW what this instance already holds; that never matches, so a client who had missed nothing got a full resync. Both are one defect: agreement between the authority and this instance has to be evaluated on two FRESH reads or it is not agreement. Now re-reads both sides after the beat, and treats any remaining disagreement as a gap in either direction — still behind means ids never reached us, still ahead means the counter was reset under us and our buffer belongs to a dead id space. Two tests, one per direction, each mutation-verified against the re-read-locally-only version: the second counter advance must produce a gap, and the raced read must NOT produce a resync. Without the second test the fix could have been 'always report a gap', which passes the first. Documented the cost side of the lead's ruling while I was in here: the condition is agreement, so a resume during CONTINUOUS publishing across the whole settle window can disagree every time and resync. Bounded by this stream being low-volume by design and resumes only happening on reconnect; if a workload makes it chatty, the answer is a longer window, not a magnitude threshold. Claude-Session: https://claude.ai/code/session_017jD6t1zjxGSq47SQpZfp1V * fix(watchevents): an absent sequence key is zero, not unreadable (Codex round 12) P2 — the counter key can DISAPPEAR after this bus has seen ids (FLUSHDB, eviction). Reading redis.Nil as 'unreadable' meant falling back to local knowledge and cheerfully replaying an id space the authority no longer has — while the next publish starts again at 1 and collides with it. Absent is a VALUE. Returning zero-and-readable makes the case fall out of the ordinary comparison with no special branch: an instance holding 101 disagrees with an authority at 0, does not converge, and the resume is answered with a gap. A genuinely fresh deployment still agrees at zero and is not resynced — which is the control leg, and the reason 'absent means gap' would have been the wrong fix: it passes the first test while resyncing every first connection on a new install. P1 as reported — the equality fast path returning without settling — is not closed, and the comment now says why rather than leaving it to be re-found. A notification published AFTER that read and missed by this instance is invisible to any check made here, and settling anyway would not close it: the same race exists in the instant after the function returns. The check's honest scope is what was missed BEFORE the resume. A message missed after it is a property of at-most-once pub/sub with no per-connection ack, and the real answer is a durable stream (Redis Streams with consumer groups), not a longer wait. Mutation-verified: restoring redis.Nil to the unreadable branch fails the disappearing-counter test. Claude-Session: https://claude.ai/code/session_017jD6t1zjxGSq47SQpZfp1V * feat(watchevents): epoch marker, so a reset that caught up is still a reset (Codex round 13) P2 — numeric detection is blind to a reset that has already climbed past this instance's high-water mark. Hold 100, lose the connection, the counter resets and ids 1-101 are published, and the only one that reaches us is 101 — the perfect contiguous successor of 100. Every arithmetic check passes, the buffer quietly mixes two id spaces, and a client resuming from OLD 100 is handed NEW 101 having silently missed the new space's 1-100. No amount of comparing numbers fixes that, because the question is not 'is this bigger' but 'is this the same sequence'. The publish script now mints an epoch once per id space (SET NX, so every publisher can offer one and the first wins) and carries it on every message; a change drops the buffer and re-anchors. The subtle half, and the one the first attempt got wrong: after an epoch change the cold-start rule must NOT admit its usual contiguous-with-our-view cursor. Within an epoch, a client at n.ID-1 is genuinely adjacent to our first id. Across one it is ambiguous — id spaces overlap, so that cursor may be the OLD sequence's n.ID-1, a different notification entirely — and admitting it hands them the new epoch's id as though it followed theirs, which is exactly the failure the epoch exists to prevent. Letting it back in one line later would have been a poor joke. The test caught it; the control leg (a cursor genuinely inside the new epoch is still served) is what stops the fix becoming 'resync everyone forever after any reset'. Wire format changed to <epoch>|<id>|<json>. Free of compat cost, checked rather than assumed: redis_bus.go does not exist on origin/main, so no released build produces or consumes the old shape. The numeric backward check stays — it covers a counter reset where the epoch key survived (eviction picks keys individually), and it is what repairs an instance with no reconnecting clients at all. Mutation-verified: ignoring the epoch change fails the new test. Claude-Session: https://claude.ai/code/session_017jD6t1zjxGSq47SQpZfp1V * docs(watchevents): the wire format comments say <epoch>|<id>|<json> (Codex round 14) Three comments still described the pre-epoch format. Worth more than a tidy-up: a maintainer following them would conclude the epoch prefix is vestigial and remove it, which reintroduces exactly the cross-epoch replay corruption round 13 existed to fix. The publishScript comment now also says outright that the epoch is not decoration and points at redisWatchEpochKey before anyone considers it removable. Verified by grepping for the old shape rather than by trusting the edits — zero remaining, which is the check I owed after getting this wrong in round 7. Claude-Session: https://claude.ai/code/session_017jD6t1zjxGSq47SQpZfp1V * chore(nix): update vendorHash for the miniredis test dependency (BUG-2651) CI's Nix job failed on a fixed-output hash mismatch, and it is neither a flake nor a surprise once seen: nix/package.nix pins the vendored module set, and adding miniredis (plus gopher-lua, its Lua interpreter) to go.mod changed it. Regenerated per the procedure the file itself documents — build and read the 'got:' line. Run on CI rather than locally because this box has no nix; the hash is a content hash of the module set determined by go.mod/go.sum, so the same inputs produce it in either place. Worth naming as a gate lesson rather than just fixing: my pre-merge matrix had build, lint, test, test-pg, vuln and Codex, and none of them can see this. A dependency change has a SEVENTH consumer — the Nix packaging — and the only thing that checks it is the CI job that just did. Adding a dependency means checking the packaging, not only the security scan. Claude-Session: https://claude.ai/code/session_017jD6t1zjxGSq47SQpZfp1V
2629 lines
116 KiB
Go
2629 lines
116 KiB
Go
package server
|
||
|
||
import (
|
||
"bytes"
|
||
"context"
|
||
"crypto/subtle"
|
||
"encoding/base64"
|
||
"encoding/json"
|
||
"fmt"
|
||
"io/fs"
|
||
"log/slog"
|
||
"net"
|
||
"net/http"
|
||
"net/netip"
|
||
"net/url"
|
||
"os"
|
||
"runtime/debug"
|
||
"strconv"
|
||
"strings"
|
||
"sync"
|
||
"sync/atomic"
|
||
"time"
|
||
|
||
"github.com/go-chi/chi/v5"
|
||
chimiddleware "github.com/go-chi/chi/v5/middleware"
|
||
"github.com/go-chi/cors"
|
||
"github.com/prometheus/client_golang/prometheus/promhttp"
|
||
|
||
"github.com/PerpetualSoftware/pad/internal/attachments"
|
||
"github.com/PerpetualSoftware/pad/internal/billing"
|
||
"github.com/PerpetualSoftware/pad/internal/collab"
|
||
"github.com/PerpetualSoftware/pad/internal/email"
|
||
"github.com/PerpetualSoftware/pad/internal/events"
|
||
"github.com/PerpetualSoftware/pad/internal/metrics"
|
||
"github.com/PerpetualSoftware/pad/internal/models"
|
||
"github.com/PerpetualSoftware/pad/internal/oauth"
|
||
"github.com/PerpetualSoftware/pad/internal/store"
|
||
"github.com/PerpetualSoftware/pad/internal/watchevents"
|
||
"github.com/PerpetualSoftware/pad/internal/webhooks"
|
||
)
|
||
|
||
type Server struct {
|
||
store *store.Store
|
||
router *chi.Mux
|
||
routerOnce sync.Once // ensures setupRouter runs once, after all config
|
||
httpServer *http.Server // underlying HTTP server (set during ListenAndServe)
|
||
webFS fs.FS // embedded web UI static files (optional)
|
||
events events.EventBus // real-time event bus (optional)
|
||
watchEvents watchevents.Bus // watch/nudge notification bus (optional, TASK-2533)
|
||
sessionPresence SessionPresence // live event-stream connections per user (optional, PLAN-2558 S1)
|
||
collab *collab.RoomManager // Yjs collab room manager (PLAN-1248); optional
|
||
webhooks *webhooks.Dispatcher // webhook dispatcher (optional)
|
||
email *email.Sender // transactional email sender (optional)
|
||
emailAPIKey string // Maileroo API key (used for unsubscribe HMAC)
|
||
emailEnvConfigured bool // email was wired from env vars (SetEmailSender); reconfigureEmail must not tear this down when platform settings clear the key
|
||
rateLimiters *RateLimiters // per-endpoint rate limiters
|
||
baseURL string // public base URL for generating links (e.g. invite URLs)
|
||
corsOrigins string // comma-separated CORS origins (empty = localhost defaults)
|
||
secureCookies bool // set Secure flag on cookies (for TLS deployments)
|
||
metrics *metrics.Metrics // Prometheus metrics (optional)
|
||
metricsToken string // shared bearer token for /metrics scrapes ("" = loopback-only)
|
||
trustedProxyCIDRs []*net.IPNet // CIDRs allowed to set X-Forwarded-For (nil = proxy headers untrusted)
|
||
ipChangeEnforceStrict bool // when true, revoke+reject sessions whose client IP OR User-Agent hash differs from the one recorded at session creation
|
||
sseMaxConnections int // global SSE connection limit (0 = unlimited)
|
||
sseMaxPerWorkspace int // per-workspace SSE connection limit (0 = unlimited)
|
||
cloudMode bool // true when running as Pad Cloud (PAD_CLOUD=true or PAD_MODE=cloud)
|
||
cloudSecrets []string // shared secrets for sidecar ↔ pad communication (supports rotation)
|
||
cloudSidecar CloudSidecar // reverse pad → pad-cloud client (e.g. Stripe cancel on account delete); nil = not configured
|
||
billingAvailable bool // true when PAD_BILLING_AVAILABLE=true — gates Stripe Checkout CTAs in the web UI (TASK-800)
|
||
version string // release version (e.g. "dev", "1.2.3")
|
||
commit string // git commit hash
|
||
buildTime string // build timestamp
|
||
twoFAChallengeSecret []byte // HMAC key for 2FA challenge tokens
|
||
|
||
// Attachments storage. Wired via SetAttachments at startup; nil-checked
|
||
// by handlers so a server constructed for a test that doesn't need
|
||
// uploads still compiles and serves every other endpoint.
|
||
attachments *attachments.Registry
|
||
attachmentMaxBytes int64 // per-file upload cap; 0 = use defaultAttachmentMaxBytes
|
||
|
||
// Image processor used by the upload handler to derive thumbnail
|
||
// variants (TASK-878) and by the editor's rotate / crop tools
|
||
// (TASK-879/880). Wired via SetImageProcessor; nil-checked by
|
||
// callers so a server without image processing — e.g. a self-host
|
||
// build that doesn't want the dependency — still serves every
|
||
// other endpoint and stores originals untouched.
|
||
imageProcessor attachments.Processor
|
||
|
||
// MCP Streamable HTTP transport (PLAN-943 TASK-950). Wired via
|
||
// SetMCPTransport at startup when the deployment is in cloud mode.
|
||
// nil on self-hosted deployments and on any cloud build that hasn't
|
||
// constructed the MCP server yet — registerMCPRoutes nil-checks so
|
||
// the routes don't mount in either case. See handlers_mcp.go.
|
||
mcpTransport http.Handler
|
||
mcpPublicURL string // canonical public URL of the MCP vhost (e.g. https://mcp.getpad.dev)
|
||
mcpAuthServerURL string // canonical URL of the OAuth auth server (e.g. https://app.getpad.dev), TASK-951
|
||
|
||
// MCP tool-surface descriptor source (PLAN-1888 / TASK-1891). Wired
|
||
// via SetToolSurfaceHandler at startup from mcp.ToolSurfaceJSON. The
|
||
// injection mirrors SetMCPTransport and exists for the same reason:
|
||
// internal/mcp imports internal/server (dispatch_http.go), so this
|
||
// package CANNOT import internal/mcp to build the catalog JSON
|
||
// itself. cmd/pad/main.go imports both and injects the serializer.
|
||
// nil → GET /api/v1/mcp/tool-surface returns 404 (handler not wired).
|
||
toolSurfaceJSON func() ([]byte, error)
|
||
|
||
// OAuth 2.1 authorization server (PLAN-943 TASK-1024 sub-PR B,
|
||
// HTTP handlers in TASK-1025 sub-PR C). Wired via SetOAuthServer
|
||
// at startup when the deployment is in cloud mode + has the
|
||
// fosite-backed server constructed. nil disables the OAuth
|
||
// surface — registerOAuthRoutes nil-checks so the routes don't
|
||
// mount on self-hosted deployments. See handlers_oauth.go.
|
||
oauthServer *oauth.Server
|
||
|
||
// claimSecret is the HMAC key for stateless 6-digit claim codes
|
||
// (PLAN-1519 / TASK-1521 / IDEA-1517 §4). Wired by SetClaimSecret
|
||
// at startup — production reuses the deployment's 32-byte
|
||
// encryption key (cfg.EncryptionKey) since both are server-
|
||
// stable secrets with equivalent rotation cadence. nil/short →
|
||
// /api/v1/oauth/claim returns 412 "claim_disabled" on every
|
||
// request, surfacing a clear misconfiguration signal rather than
|
||
// silently accepting forgeable codes.
|
||
claimSecret []byte
|
||
|
||
// oauthMetricsWired records whether wireOAuthMetricsObserver has
|
||
// already attached the active-tokens callback collector. Re-
|
||
// registering would panic via prometheus.MustRegister, so the flag
|
||
// guards the one-shot registration. The TTL observer side is
|
||
// idempotent (just a function-pointer set) and runs unconditionally.
|
||
oauthMetricsWired bool
|
||
|
||
// MCP audit log async writer (PLAN-943 TASK-960). Spawned by
|
||
// startMCPAuditWriter at startup when MCP is wired; shut down
|
||
// from Server.Stop. nil-safe: every audit-emitting code path
|
||
// nil-checks so MCP-less builds + tests that don't start the
|
||
// writer still work. See middleware_mcp_audit.go.
|
||
mcpAudit *mcpAuditWriter
|
||
|
||
// MCP session tracker (PLAN-943 TASK-1120). Replaces the naive
|
||
// +1/-1 active-sessions accounting from TASK-961. Wired by
|
||
// startMCPSessionTracker (called from SetMCPTransport in cloud
|
||
// mode); shut down from Server.Stop alongside the audit writer.
|
||
// nil-safe: trackMCPSession + the gauge-update path both
|
||
// nil-check so non-cloud builds + tests run without the tracker.
|
||
// See middleware_mcp_session.go.
|
||
mcpSessions *mcpSessionTracker
|
||
mcpSessionTTL time.Duration // 0 → defaultMCPSessionTTL
|
||
mcpSessionSweepInterval time.Duration // 0 → defaultMCPSessionSweepInterval
|
||
|
||
// storageInfoCache memoizes per-workspace storage usage summaries
|
||
// behind a short TTL (storageInfoTTL). Reduces DB load on the
|
||
// Settings → Storage page and quota-aware UI surfaces. Initialized
|
||
// in newServer; never nil so handlers can call get/set without
|
||
// guarding.
|
||
storageInfoCache *storageInfoCache
|
||
|
||
// copyItemFn indirects Store.CopyItemAcrossWorkspaces for the
|
||
// cross-workspace copy endpoint (PLAN-2357 / TASK-2365).
|
||
//
|
||
// It exists because DR-13 forbids the endpoint from ever transparently
|
||
// retrying a mutating copy — there is no idempotency key in v1, so a
|
||
// retry after a post-commit failure duplicates the item — and the only
|
||
// falsifiable way to assert "called exactly once" is to count the calls.
|
||
// Several tests also use it to inject the store's typed errors and to
|
||
// land a concurrent mutation deterministically between the handler's
|
||
// authorization and the store call. See handleCopyItem and
|
||
// TestCopyEndpoint_DoesNotRetryOnAmbiguousError.
|
||
//
|
||
// It is nil on every production path — nothing outside package server can
|
||
// set it, no constructor or setter assigns it, and only _test.go files do
|
||
// (Codex round 6: it is compiled into the binary, so "test-only" describes
|
||
// the convention, not a compiler-enforced guarantee).
|
||
copyItemFn func(store.CrossWorkspaceCopyRequest) (*store.CrossWorkspaceCopyResult, error)
|
||
|
||
// importBundleMaxBytes caps a single workspace import bundle.
|
||
// 0 → defaultImportBundleMaxBytes (2 GiB). Set via
|
||
// SetImportBundleMaxBytes from cmd/pad/main.go using the
|
||
// PAD_IMPORT_BUNDLE_MAX_BYTES env var so operators with larger
|
||
// exports can opt in without recompiling.
|
||
importBundleMaxBytes int64
|
||
|
||
// importArtifactMaxBytes caps a single playbook/convention artifact
|
||
// import (POST /workspaces/{ws}/import-artifact). 0 →
|
||
// defaultImportArtifactMaxBytes (1 MiB). A single artifact is tiny;
|
||
// the cap is the first line of defense against an oversized body
|
||
// being materialized before the YAML-bomb guard runs. Set via
|
||
// SetImportArtifactMaxBytes from cmd/pad/main.go.
|
||
importArtifactMaxBytes int64
|
||
|
||
// orphanGC holds the periodic-sweep config + lifecycle for the
|
||
// attachment orphan garbage collector (TASK-886). Configured via
|
||
// SetOrphanGCConfig and started via StartOrphanGC. Stop() signals
|
||
// the loop to exit and waits for it via the bg WaitGroup.
|
||
orphanGC orphanGCConfig
|
||
|
||
// opLogGC holds the periodic-sweep config + lifecycle for the
|
||
// Yjs op-log prune sweeper (TASK-1309). Mirrors orphanGC's
|
||
// pattern. Configured via SetOpLogGCConfig + started via
|
||
// StartOpLogGC; Stop() signals the loop via stopOpLogGC.
|
||
opLogGC opLogGCConfig
|
||
|
||
// tokenReaper holds the periodic-sweep config + lifecycle for the
|
||
// short-lived-credential reaper (PLAN-1933 DR-5 / TASK-1936).
|
||
// Mirrors orphanGC/opLogGC. Configured via SetTokenReaperConfig +
|
||
// started via StartTokenReaper; Stop() signals the loop via
|
||
// stopTokenReaper.
|
||
tokenReaper tokenReaperConfig
|
||
|
||
// workspacePurge holds the periodic-sweep config + lifecycle for the
|
||
// soft-deleted-workspace hard-purge sweeper (TASK-1966 — the 30-day
|
||
// GDPR erasure SLA). Mirrors orphanGC. Configured via
|
||
// SetWorkspacePurgeConfig + started via StartWorkspacePurgeSweeper;
|
||
// Stop() signals the loop via stopWorkspacePurgeSweeper.
|
||
workspacePurge workspacePurgeConfig
|
||
|
||
// inFlightUploadHashes tracks content_hash values for uploads
|
||
// that have called AttachmentStore.Put but not yet inserted the
|
||
// attachments row. Without this, the orphan GC could delete a
|
||
// blob between Put and CreateAttachment, leaving a live row that
|
||
// references a missing blob (Codex P2 on PR #307 round 1).
|
||
//
|
||
// A plain map + mutex rather than sync.Map: counters need
|
||
// atomic-with-delete semantics (decrement-then-delete-if-zero
|
||
// must be one critical section, not two — sync.Map.CompareAndDelete
|
||
// addresses the entry but not the inc/dec interleaving). Codex
|
||
// P1 round 2 caught the prior sync.Map version racing on
|
||
// release-vs-reload of the same hash.
|
||
inFlightHashesMu sync.Mutex
|
||
inFlightHashes map[string]int64
|
||
|
||
// rowlessNoListerOnce gates the once-per-process notice that a
|
||
// registered attachment backend lacks the Lister capability, leaving
|
||
// the rowless-blob sweep (BUG-2406) inert for it. Logged rather than
|
||
// silently skipped so an operator can tell the leak class is
|
||
// unguarded on that backend; once, so a 24h-cadence sweep doesn't
|
||
// turn it into log spam.
|
||
rowlessNoListerOnce sync.Once
|
||
|
||
// rowlessPreDeleteHook, when non-nil, runs inside the rowless
|
||
// sweep's in-flight critical section immediately BEFORE the
|
||
// delete-time row re-check. Test seam only (injectedStageFailure
|
||
// precedent): it lets a test commit a row for the hash at exactly
|
||
// the point that distinguishes the batched subtraction from the
|
||
// re-check, making the TOCTOU leg deterministic.
|
||
rowlessPreDeleteHook func(hash string)
|
||
|
||
// bg tracks fire-and-forget goroutines spawned by request handlers
|
||
// (TouchUserActivity in middleware_auth, async email sends, etc.) so
|
||
// the server can drain them before shutdown / test cleanup. Without
|
||
// this, tests using t.TempDir() race the still-running goroutine's
|
||
// SQLite WAL write against TempDir RemoveAll, leaving "directory not
|
||
// empty" cleanup errors in CI. See BUG-842.
|
||
bg sync.WaitGroup
|
||
|
||
// First-run bootstrap token (TASK-1167 / PLAN-1166). When non-empty,
|
||
// handleBootstrap accepts the value via the X-Bootstrap-Token header
|
||
// from non-loopback peers (self-host mode only — cloud mode never
|
||
// loads or honors a token, D2/D10). Wired at startup via
|
||
// SetBootstrapToken; cleared by consumeBootstrapToken after the first
|
||
// admin is created.
|
||
//
|
||
// The mutex protects the token field AND the entire validate-token →
|
||
// check-UserCount → CreateUser → consume sequence in handleBootstrap.
|
||
// Two simultaneous valid-token requests with different emails would
|
||
// otherwise create two admins from one token (F5). Bootstrap happens
|
||
// once per install, so the contention window is irrelevant.
|
||
bootstrapMu sync.Mutex
|
||
bootstrapToken string
|
||
bootstrapTokenPath string
|
||
|
||
// bypassSetupToken, when true, allows the first-admin bootstrap POST to
|
||
// succeed from any IP without an X-Bootstrap-Token header — i.e. the
|
||
// /setup form on the web UI works directly, without the operator having
|
||
// to copy a token out of `docker logs`. Wired from PAD_BYPASS_SETUP_TOKEN
|
||
// at startup via SetBypassSetupToken (cmd/pad/main.go).
|
||
//
|
||
// Self-host only — cloud mode IGNORES this flag entirely (D2/D10 from
|
||
// the original logs-token design: cloud bootstrap stays loopback-only).
|
||
// The UserCount==0 gate in handleBootstrap is unchanged: once the first
|
||
// admin exists, the bootstrap endpoint returns 409 "already initialized"
|
||
// regardless of bypass. This matches the operator's mental model — the
|
||
// flag opens up the *first-run* surface, not registration in general.
|
||
//
|
||
// Operators on trusted networks (Unraid LAN, Tailscale-only deployments,
|
||
// homelabs behind a firewall) typically prefer this; operators with
|
||
// public exposure should leave it off and use the logs-token path.
|
||
bypassSetupToken bool
|
||
|
||
// restoreAckFault is a TEST SEAM (always nil in production). When non-nil,
|
||
// handleRestoreItemVersion's collab commit closure invokes it AFTER the restore
|
||
// transaction has durably committed; a non-nil return simulates a Postgres commit
|
||
// whose acknowledgement was lost at the connection boundary (the tx landed, but
|
||
// the driver surfaces an error), exercising BUG-2276 residual 1's commit-outcome
|
||
// reconciliation end-to-end through the real handler.
|
||
restoreAckFault func() error
|
||
|
||
// watchPredicatesLoadFault is a TEST SEAM (always nil in production,
|
||
// TASK-2533). When non-nil, loadWatchPredicates calls it before
|
||
// touching the store; a non-nil return short-circuits the real
|
||
// ListWatchesForUser call and is returned as the reload error —
|
||
// exercising GET /api/v1/events/stream's reval-tick error path
|
||
// (codex round 4: a watch-list reload failure must not also skip
|
||
// the identity/visibility refresh) deterministically, without
|
||
// needing to actually break the DB connection mid-test.
|
||
//
|
||
// atomic.Pointer, not a plain func field (codex round 5 finding 2):
|
||
// unlike restoreAckFault — set once, synchronously, before the single
|
||
// HTTP request that will read it, so goroutine-creation's own
|
||
// happens-before edge makes a plain field safe there — this seam is
|
||
// set by a test AFTER the SSE stream's background goroutine is
|
||
// already running and reading it on every reval tick. A plain field
|
||
// written from the test's goroutine while that goroutine reads it
|
||
// concurrently is a genuine, if timing-dependent, data race
|
||
// (verified: restoreAckFault's OWN usage doesn't share this flaw,
|
||
// since it's never touched after the goroutine that reads it starts,
|
||
// so it was intentionally left as a plain field rather than changed
|
||
// too).
|
||
watchPredicatesLoadFault atomic.Pointer[func() error]
|
||
|
||
// watchRevalTickOverride is a TEST SEAM (always nil in production,
|
||
// BUG-2570). When non-nil, GET /api/v1/events/stream selects reval
|
||
// ticks from this channel instead of the interval ticker, letting a
|
||
// test drive each revalidation tick explicitly. That is the only way
|
||
// to pin assertions to a SPECIFIC tick: with a free-running ticker,
|
||
// no test ordering can guarantee that an unwanted extra tick doesn't
|
||
// fire between two test steps — an extra SUCCESSFUL tick resets the
|
||
// visibility cache and reloads the watch list, masking exactly the
|
||
// reset-skipped-on-fault regression the reval-fault test guards, and
|
||
// enough extra FAULTING ticks clear the watch set (both observed as
|
||
// codex-round findings on BUG-2570's first fix attempt).
|
||
//
|
||
// Same atomic.Pointer rationale as watchPredicatesLoadFault above:
|
||
// read by the stream's background goroutine (once, at stream setup)
|
||
// while tests may write it. Tests must set it BEFORE connecting the
|
||
// stream they want to drive — a write after setup is not observed.
|
||
watchRevalTickOverride atomic.Pointer[chan time.Time]
|
||
}
|
||
|
||
// goAsync spawns fn in a goroutine that's tracked by s.bg, so Stop() can
|
||
// wait for in-flight background work to finish. Use this for any
|
||
// fire-and-forget work that touches the database, filesystem, or external
|
||
// services from inside a request handler — never bare `go func() {...}()`.
|
||
func (s *Server) goAsync(fn func()) {
|
||
s.bg.Add(1)
|
||
go func() {
|
||
defer s.bg.Done()
|
||
// Recover from panics in fn so a single bad background task
|
||
// (e.g. deriveThumbnails hitting a Go image-decoder panic on a
|
||
// crafted upload, or an email send) can't crash the whole
|
||
// single-binary server for every tenant. chi's Recoverer only
|
||
// covers request goroutines, not these detached ones. The
|
||
// deferred Done() above still fires because recover() keeps the
|
||
// goroutine from unwinding past this point.
|
||
defer func() {
|
||
if r := recover(); r != nil {
|
||
slog.Error("background task panicked",
|
||
"panic", r,
|
||
"stack", string(debug.Stack()))
|
||
}
|
||
}()
|
||
fn()
|
||
}()
|
||
}
|
||
|
||
// recoverSweeper is the panic firewall for the long-running background
|
||
// sweeper loops (orphan GC, op-log GC, token reaper, workspace purge).
|
||
// Each of those manages its own s.bg.Add/Done + stop-channel lifecycle,
|
||
// so — unlike fire-and-forget work — they can't just route through
|
||
// goAsync without breaking their shutdown handling or double-counting
|
||
// s.bg. Instead each spawns `defer s.recoverSweeper("<name>")` as a
|
||
// deferred call INSIDE its goroutine (BUG-2071): a panic in the loop
|
||
// body is logged with a stack (matching goAsync's style) and unwinds
|
||
// cleanly, the goroutine's own deferred s.bg.Done() still fires because
|
||
// recover() stops the unwind here, and Stop() still returns. Without it a
|
||
// panic in any sweeper takes down the whole single-binary server for
|
||
// every tenant. Must be `defer`-called directly in the goroutine for
|
||
// recover() to catch the panic.
|
||
func (s *Server) recoverSweeper(name string) {
|
||
if r := recover(); r != nil {
|
||
slog.Error("background sweeper panicked",
|
||
"sweeper", name,
|
||
"panic", r,
|
||
"stack", string(debug.Stack()))
|
||
}
|
||
}
|
||
|
||
// Stop waits for all background goroutines started via goAsync to finish
|
||
// AND drains the rate-limiter cleanup goroutines spawned at construction
|
||
// time (BUG-851). Safe to call multiple times. Should be called before
|
||
// Store.Close() so in-flight DB writes don't race a closed connection
|
||
// (or worse, the SQLite -wal/-shm file removal in t.TempDir cleanup).
|
||
func (s *Server) Stop() {
|
||
// Signal long-running background loops (orphan GC, etc.) to exit.
|
||
// Each loop registers itself on s.bg, so the Wait() below blocks
|
||
// until they actually finish and any in-flight goroutines drain.
|
||
s.stopOrphanGC()
|
||
// Yjs op-log prune sweeper (TASK-1309). Same lifecycle pattern;
|
||
// signals BEFORE Wait() so the goroutine sees the close and exits.
|
||
s.stopOpLogGC()
|
||
// Short-lived-credential reaper (PLAN-1933 DR-5 / TASK-1936). Same
|
||
// lifecycle pattern; signal BEFORE Wait() so the goroutine exits.
|
||
s.stopTokenReaper()
|
||
// Soft-deleted-workspace hard-purge sweeper (TASK-1966). Same
|
||
// lifecycle pattern; signal BEFORE Wait() so the goroutine exits.
|
||
s.stopWorkspacePurgeSweeper()
|
||
// MCP audit writer / sweeper run on s.bg too. Signal first so
|
||
// the workers see the close BEFORE Wait() blocks; without the
|
||
// signal Wait would hang forever on the writer's blocking
|
||
// queue receive.
|
||
s.stopMCPAuditWriter()
|
||
// MCP session tracker (TASK-1120) runs its sweeper on s.bg too.
|
||
// Order with the audit writer doesn't matter — both are
|
||
// independent goroutines; we just need the close BEFORE Wait().
|
||
s.stopMCPSessionTracker()
|
||
// Close the collab room manager BEFORE bg.Wait() so any in-flight
|
||
// op-log GC sweep (TASK-1309) blocked on a per-item lock behind
|
||
// an active Join can drain. collab.Close() tears down the Joins
|
||
// (their WS readLoops return, runConn unwinds, itemLocks
|
||
// release), which unblocks the GC's per-item PruneItemOpLogIfDormantBefore
|
||
// call. Without this ordering, Stop() can deadlock: GC waits on
|
||
// itemLock; Join holds itemLock until WS closes; WS only closes
|
||
// when collab.Close() runs; collab.Close() only runs after
|
||
// bg.Wait(); bg.Wait() never returns because GC is stuck.
|
||
// Per Codex review of TASK-1309 [P2]. nil-safe: collab is optional.
|
||
if s.collab != nil {
|
||
s.collab.Close()
|
||
}
|
||
s.bg.Wait()
|
||
// Watch/nudge bus (BUG-2651). Closed AFTER bg.Wait() so a background
|
||
// producer cannot publish into a bus that is already tearing down; both
|
||
// implementations are safe if one does anyway (MemoryBus finds no
|
||
// subscribers, RedisBus finds a cancelled context and fails closed).
|
||
//
|
||
// This matters more than it did for MemoryBus, whose Close only dropped
|
||
// channels: RedisBus holds a receive goroutine and a Redis subscription
|
||
// from construction, so skipping it leaks both for the process's life
|
||
// (Codex round 1 P2). Closing also closes every subscriber channel, which
|
||
// is how a still-open SSE stream learns to unwind.
|
||
if s.watchEvents != nil {
|
||
s.watchEvents.Close()
|
||
}
|
||
s.rateLimiters.Stop() // nil-safe via the RateLimiters receiver guard
|
||
}
|
||
|
||
func New(s *store.Store) *Server {
|
||
rl := NewRateLimiters()
|
||
// PAD_DISABLE_RATE_LIMITS turns off ALL HTTP rate limiting when set to a
|
||
// truthy value. It exists ONLY for the E2E harness (BUG-2089): every
|
||
// Playwright test shares one loopback IP (127.0.0.1), so the auth limiter
|
||
// (5 logins/min/IP) trips the moment a spec logs in a couple of browser
|
||
// clients — collab-persistence.spec.ts logs in two per test. A nil
|
||
// rateLimiters makes RateLimit() a pass-through (see middleware_ratelimit.go
|
||
// line 379), and Stop() + the MCP path are already nil-safe. Never set this
|
||
// in production or self-host; it's an explicit opt-in so it can't flip on
|
||
// by accident.
|
||
if disabled, _ := strconv.ParseBool(os.Getenv("PAD_DISABLE_RATE_LIMITS")); disabled {
|
||
rl = nil
|
||
}
|
||
return &Server{
|
||
store: s,
|
||
rateLimiters: rl,
|
||
storageInfoCache: newStorageInfoCache(storageInfoTTL),
|
||
}
|
||
}
|
||
|
||
// Init2FASecret loads the 2FA challenge signing key from platform_settings.
|
||
// If no key exists (first run), a new random key is generated and persisted.
|
||
// This must be called before the server handles requests so that challenge
|
||
// tokens survive process restarts and work across multiple instances.
|
||
func (s *Server) Init2FASecret() error {
|
||
const settingKey = "2fa_challenge_secret"
|
||
|
||
existing, err := s.store.GetPlatformSetting(settingKey)
|
||
if err != nil {
|
||
return fmt.Errorf("load 2FA secret: %w", err)
|
||
}
|
||
|
||
if existing != "" {
|
||
decoded, err := base64.StdEncoding.DecodeString(existing)
|
||
if err != nil {
|
||
return fmt.Errorf("decode 2FA secret: %w", err)
|
||
}
|
||
s.twoFAChallengeSecret = decoded
|
||
return nil
|
||
}
|
||
|
||
// First run — generate and persist a new secret.
|
||
// Multiple instances may race here on a fresh database; after persisting,
|
||
// re-read the winning value so all instances converge on the same key.
|
||
secret, err := generateTwoFASecret()
|
||
if err != nil {
|
||
return err
|
||
}
|
||
encoded := base64.StdEncoding.EncodeToString(secret)
|
||
if err := s.store.SetPlatformSetting(settingKey, encoded); err != nil {
|
||
return fmt.Errorf("persist 2FA secret: %w", err)
|
||
}
|
||
|
||
// Re-read to pick up whichever instance won the race (upsert may have
|
||
// been overwritten by a concurrent instance between our check and write).
|
||
final, err := s.store.GetPlatformSetting(settingKey)
|
||
if err != nil {
|
||
return fmt.Errorf("re-read 2FA secret: %w", err)
|
||
}
|
||
decoded, err := base64.StdEncoding.DecodeString(final)
|
||
if err != nil {
|
||
return fmt.Errorf("decode 2FA secret after re-read: %w", err)
|
||
}
|
||
s.twoFAChallengeSecret = decoded
|
||
slog.Info("initialized 2FA challenge signing key")
|
||
return nil
|
||
}
|
||
|
||
// SetCloudMode enables cloud mode with the shared sidecar secret(s).
|
||
// Accepts a comma-separated list of secrets for rotation support:
|
||
// "new-key,old-key" — both are accepted for INBOUND calls from pad-cloud.
|
||
// The OUTBOUND direction (pad → pad-cloud, see SetCloudSidecar) is
|
||
// configured separately via PAD_CLOUD_OUTBOUND_SECRET or derived from the
|
||
// last entry of this list — see cmd/pad/main.go for the resolution order.
|
||
func (s *Server) SetCloudMode(secret string) {
|
||
s.cloudMode = true
|
||
for _, k := range strings.Split(secret, ",") {
|
||
k = strings.TrimSpace(k)
|
||
if k != "" {
|
||
s.cloudSecrets = append(s.cloudSecrets, k)
|
||
}
|
||
}
|
||
// Propagate to the email sender so transactional emails carry the
|
||
// getpad.dev marketing footer (docs/brand.md §7) on Cloud installs.
|
||
// Self-hosted deployments leave cloudMode false on the sender, keeping
|
||
// outgoing mail neutral so operators can ship under their own brand.
|
||
if s.email != nil {
|
||
s.email.SetCloudMode(true)
|
||
}
|
||
}
|
||
|
||
// CloudSidecar is the reverse pad → pad-cloud client interface. Concrete
|
||
// implementation lives in internal/billing so server has no direct Stripe
|
||
// dependency. Kept as an interface so tests can inject fakes without
|
||
// spinning up a real HTTP server or touching Stripe.
|
||
type CloudSidecar interface {
|
||
// CancelCustomer asks pad-cloud to cancel every active Stripe subscription
|
||
// for customerID and then delete the Stripe customer object. Used by
|
||
// handleDeleteAccount to cascade account deletion through to Stripe billing
|
||
// (TASK-690).
|
||
//
|
||
// Failure contract: any non-nil error means the caller MUST abort the
|
||
// local delete. pad-cloud normalizes Stripe's "already gone" cases to a
|
||
// 200 on its side (see pad-cloud stripe.go isStripeAlreadyGone), so
|
||
// every error we see here is a real failure — transport, 4xx (ops
|
||
// misconfig), or 5xx (upstream breakage). Continuing after an error
|
||
// would wipe the user's StripeCustomerID while leaving the subscription
|
||
// billing, which is exactly the regression TASK-690 exists to prevent.
|
||
CancelCustomer(customerID string) error
|
||
|
||
// GetBillingMetrics fetches an aggregated Stripe-derived snapshot from
|
||
// pad-cloud's /admin/metrics/billing endpoint (active subs, MRR, ARR,
|
||
// churn, cancellations). Used by handleAdminBillingStats to power the
|
||
// admin Billing dashboard (TASK-827 / PLAN-825).
|
||
//
|
||
// Failure contract: returns an error on transport failure or non-200
|
||
// status. The admin handler treats any error as "degrade to local-only"
|
||
// and surfaces the distinction in its response via cloud_unreachable —
|
||
// it never propagates the upstream failure to the operator's browser.
|
||
GetBillingMetrics() (*billing.BillingMetricsResponse, error)
|
||
}
|
||
|
||
// SetCloudSidecar installs the reverse pad → pad-cloud client. Called from
|
||
// cmd/pad/main.go when PAD_CLOUD_SIDECAR_URL + PAD_CLOUD_SECRET are set.
|
||
// When unset, handleDeleteAccount skips the Stripe cancel step (self-hosted
|
||
// deploys that don't run a Stripe-backed sidecar have nothing to cascade).
|
||
func (s *Server) SetCloudSidecar(c CloudSidecar) {
|
||
s.cloudSidecar = c
|
||
}
|
||
|
||
// SetBillingAvailable marks this deployment as having Stripe Checkout wired
|
||
// up. Called from cmd/pad/main.go when PAD_BILLING_AVAILABLE=true is set.
|
||
// When false (the default), the session payload advertises billing_available=false
|
||
// so the web UI hides Stripe CTAs rather than dead-ending at a 503. TASK-800.
|
||
func (s *Server) SetBillingAvailable(v bool) {
|
||
s.billingAvailable = v
|
||
}
|
||
|
||
// IsCloud reports whether the server is running in cloud mode.
|
||
func (s *Server) IsCloud() bool {
|
||
return s.cloudMode
|
||
}
|
||
|
||
// SetVersion stores the build version info for the health endpoint.
|
||
func (s *Server) SetVersion(version, commit, buildTime string) {
|
||
s.version = version
|
||
s.commit = commit
|
||
s.buildTime = buildTime
|
||
}
|
||
|
||
// SetBaseURL sets the public base URL used for generating shareable links.
|
||
//
|
||
// If the supplied URL has an unspecified bind-all host ("0.0.0.0", "::",
|
||
// "[::]"), this logs a WARN: such a URL is the right thing to *bind* to
|
||
// but the wrong thing to *send* to a recipient (their browser cannot
|
||
// resolve 0.0.0.0 / :: as a connect target). Callers shipping email
|
||
// links from such a deployment should set PAD_URL or PUBLIC_URL to the
|
||
// real public hostname (e.g. https://app.getpad.dev). See BUG-899.
|
||
func (s *Server) SetBaseURL(rawURL string) {
|
||
s.baseURL = strings.TrimRight(rawURL, "/")
|
||
if s.baseURL == "" {
|
||
return
|
||
}
|
||
if u, err := url.Parse(s.baseURL); err == nil {
|
||
switch u.Hostname() {
|
||
case "", "0.0.0.0", "::":
|
||
slog.Warn("server base URL has an unspecified host; emailed links (password reset, invites, share links) will not be reachable. Set PAD_URL or PUBLIC_URL to the deployment's public URL (e.g. https://app.getpad.dev).", "base_url", s.baseURL)
|
||
}
|
||
}
|
||
}
|
||
|
||
// specialUseTLDs is the (finite) set of reserved top-level names from the IANA
|
||
// Special-Use Domain Names registry + related RFCs that never resolve to a
|
||
// public web host, so an emailed link using one is undeliverable. Matched on
|
||
// the final label so example.com (public) is allowed while foo.example /
|
||
// foo.test / bar.internal / pad.home.arpa / x.onion / y.alt (non-public) are
|
||
// rejected. Sources: RFC 6761 (localhost/invalid/test/example), RFC 6762
|
||
// (local), RFC 8375 (home.arpa) + the "arpa" infrastructure TLD, RFC 7686
|
||
// (onion), RFC 9476 (alt), and ICANN-reserved "internal".
|
||
var specialUseTLDs = map[string]bool{
|
||
"localhost": true, "local": true, "internal": true,
|
||
"invalid": true, "test": true, "example": true, "arpa": true,
|
||
"onion": true, "alt": true,
|
||
}
|
||
|
||
// hasUsableBaseURL reports whether s.baseURL is a URL a verification-email
|
||
// recipient on the public internet can actually reach. It supersets the
|
||
// unreachable-host warning in SetBaseURL (BUG-899): the bind-all hosts
|
||
// 0.0.0.0 / :: are the right thing to bind() to but the wrong thing to email,
|
||
// and so is every other host an external recipient can't resolve or route to.
|
||
//
|
||
// This gates the cloud self-serve signup path (PLAN-1933 DR-6): creating an
|
||
// UNVERIFIED user whose only way out of the write-lock is an emailed link we
|
||
// can't deliver would strand them permanently. So "usable" is conservative —
|
||
// anything not clearly a public web endpoint disqualifies self-serve signup:
|
||
//
|
||
// - scheme must be http/https (a browser can't follow ftp://, file://, …);
|
||
// - no query/fragment (the link is built by concatenation, so either would
|
||
// push the /verify-email/<token> route into the query/fragment);
|
||
// - a present port must be a valid TCP port (1–65535);
|
||
// - the host must be a valid public DNS FQDN, NOT a literal IP. A real cloud
|
||
// verification endpoint is a hostname (Pad Cloud is app.getpad.dev);
|
||
// bare-IP base URLs aren't used for emailed links, and exhaustively
|
||
// enumerating every non-public IP range (loopback / private / CGNAT /
|
||
// TEST-NET / 6to4 / benchmarking / reserved / … across IPv4 and IPv6) is a
|
||
// losing game — so we require a hostname and fail closed on any IP literal.
|
||
//
|
||
// No usable base URL → no self-serve signup (registration stays closed rather
|
||
// than minting a write-locked user). Self-host and admin/invitation signup are
|
||
// unaffected either way.
|
||
func (s *Server) hasUsableBaseURL() bool {
|
||
if s.baseURL == "" {
|
||
return false
|
||
}
|
||
u, err := url.Parse(s.baseURL)
|
||
if err != nil {
|
||
return false
|
||
}
|
||
if u.Scheme != "http" && u.Scheme != "https" {
|
||
return false
|
||
}
|
||
// The verification link is built by concatenation (baseURL +
|
||
// "/verify-email/" + token), so a base URL carrying a query or fragment
|
||
// would push the route into the query/fragment and break the link.
|
||
if u.RawQuery != "" || u.Fragment != "" {
|
||
return false
|
||
}
|
||
host := u.Hostname()
|
||
if host == "" {
|
||
return false
|
||
}
|
||
// A present port must be a valid TCP port (1–65535); url.Parse accepts
|
||
// out-of-range numeric ports that no client can actually connect to.
|
||
if p := u.Port(); p != "" {
|
||
n, perr := strconv.Atoi(p)
|
||
if perr != nil || n < 1 || n > 65535 {
|
||
return false
|
||
}
|
||
}
|
||
|
||
// Reject literal IPs outright — a usable public verification endpoint is a
|
||
// DNS hostname, and "is this IP publicly reachable" is not decidable from a
|
||
// finite denylist. Fail closed on any IP.
|
||
if _, aerr := netip.ParseAddr(host); aerr == nil {
|
||
return false
|
||
}
|
||
|
||
// The host must be a syntactically-valid, multi-label public FQDN (rejects
|
||
// malformed hosts like ".com", "foo..com", "-a.com" and special-use TLDs).
|
||
return isPublicDNSName(host)
|
||
}
|
||
|
||
// isPublicDNSName reports whether host is a syntactically-valid public FQDN
|
||
// (RFC 1123 labels) whose TLD is neither a special-use reserved name nor
|
||
// all-numeric. Empty labels, over-length labels, and invalid characters are
|
||
// rejected. Assumes host is not an IP literal (that's handled before this is
|
||
// called). Punycode/IDN TLDs (xn--…) are accepted since they aren't all-digit.
|
||
func isPublicDNSName(host string) bool {
|
||
host = strings.ToLower(strings.TrimSuffix(host, "."))
|
||
if host == "" || len(host) > 253 {
|
||
return false
|
||
}
|
||
labels := strings.Split(host, ".")
|
||
if len(labels) < 2 {
|
||
return false
|
||
}
|
||
for _, l := range labels {
|
||
if !isValidDNSLabel(l) {
|
||
return false
|
||
}
|
||
}
|
||
tld := labels[len(labels)-1]
|
||
if specialUseTLDs[tld] || isAllDigits(tld) {
|
||
return false
|
||
}
|
||
return true
|
||
}
|
||
|
||
// isValidDNSLabel reports whether l is a valid RFC 1123 hostname label:
|
||
// 1–63 chars of [a-z0-9-], not starting or ending with a hyphen.
|
||
func isValidDNSLabel(l string) bool {
|
||
if len(l) == 0 || len(l) > 63 {
|
||
return false
|
||
}
|
||
if l[0] == '-' || l[len(l)-1] == '-' {
|
||
return false
|
||
}
|
||
for i := 0; i < len(l); i++ {
|
||
c := l[i]
|
||
if !((c >= 'a' && c <= 'z') || (c >= '0' && c <= '9') || c == '-') {
|
||
return false
|
||
}
|
||
}
|
||
return true
|
||
}
|
||
|
||
// isAllDigits reports whether s is non-empty and entirely ASCII digits. A
|
||
// public TLD is never all-numeric (RFC 3696), so an all-digit final label
|
||
// signals a malformed host rather than a reachable name.
|
||
func isAllDigits(s string) bool {
|
||
if s == "" {
|
||
return false
|
||
}
|
||
for i := 0; i < len(s); i++ {
|
||
if s[i] < '0' || s[i] > '9' {
|
||
return false
|
||
}
|
||
}
|
||
return true
|
||
}
|
||
|
||
// emailConfigured reports whether this instance can actually SEND an emailed
|
||
// link — a sender is wired AND the public base URL is usable. This is the
|
||
// DR-6 gate for cloud email self-registration: the sender-only check (s.email
|
||
// != nil) is insufficient because link generation also needs a reachable
|
||
// public base URL (see handleRegister's verification email + BUG-899).
|
||
func (s *Server) emailConfigured() bool {
|
||
return s.email != nil && s.hasUsableBaseURL()
|
||
}
|
||
|
||
// SetEventBus attaches an event bus for real-time SSE streaming.
|
||
func (s *Server) SetEventBus(bus events.EventBus) {
|
||
s.events = bus
|
||
}
|
||
|
||
// SetWatchEventsBus attaches the watch/nudge notification bus consumed by
|
||
// GET /api/v1/events/stream (TASK-2533). Nil-checked by every producer and
|
||
// by the stream handler, so a server constructed without one (e.g. a test
|
||
// that doesn't exercise watches) still serves every other endpoint.
|
||
func (s *Server) SetWatchEventsBus(bus watchevents.Bus) {
|
||
s.watchEvents = bus
|
||
}
|
||
|
||
// SetSessionPresence attaches the live-session registry read by
|
||
// GET /api/v1/sessions and written by GET /api/v1/events/stream
|
||
// (PLAN-2558 S1). Nil-checked at both ends, so a server constructed
|
||
// without one still streams events — it just can't answer "who is
|
||
// listening?", and says so with a 503 rather than an empty list (see
|
||
// handleListSessions).
|
||
func (s *Server) SetSessionPresence(p SessionPresence) {
|
||
s.sessionPresence = p
|
||
}
|
||
|
||
// SetCollabRoomManager attaches a Yjs collab RoomManager (PLAN-1248).
|
||
// When set, the /api/v1/collab/{itemID} WebSocket endpoint hands new
|
||
// connections to the manager for op-log replay + fan-out. When nil,
|
||
// the endpoint exists but answers 503 — that's intentional so a
|
||
// self-host build that wants the editor without collab can leave
|
||
// this unwired without surfacing surprise behaviour.
|
||
func (s *Server) SetCollabRoomManager(rm *collab.RoomManager) {
|
||
s.collab = rm
|
||
}
|
||
|
||
// SetWebhookDispatcher attaches a webhook dispatcher for outgoing
|
||
// notifications. Delivery goroutines are routed through s.goAsync so they're
|
||
// tracked on s.bg — Server.Stop() waits for in-flight deliveries (closing the
|
||
// BUG-842 shutdown race where a detached delivery writes to a closed store)
|
||
// and inherits goAsync's panic recovery (BUG-2011).
|
||
func (s *Server) SetWebhookDispatcher(d *webhooks.Dispatcher) {
|
||
if d != nil {
|
||
d.SetSpawn(s.goAsync)
|
||
}
|
||
s.webhooks = d
|
||
}
|
||
|
||
// SetEmailSender attaches a transactional email sender.
|
||
// The apiKey is stored separately for deriving the unsubscribe HMAC secret.
|
||
//
|
||
// If the server is already in cloud mode when this is called (i.e.
|
||
// SetCloudMode ran before email config arrived from main.go), propagate
|
||
// the flag so the new sender adds the getpad.dev marketing footer to
|
||
// outgoing emails. Without this, the cloud-mode flag would silently
|
||
// fail to take effect when callers wired email and cloud mode in
|
||
// either order.
|
||
func (s *Server) SetEmailSender(e *email.Sender, apiKey ...string) {
|
||
s.email = e
|
||
// A sender wired here comes from an out-of-band source (env vars at startup).
|
||
// Mark it so reconfigureEmail leaves it in place when platform settings carry
|
||
// no key — env is the deployment baseline, not something the admin UI disables.
|
||
s.emailEnvConfigured = e != nil
|
||
if len(apiKey) > 0 {
|
||
s.emailAPIKey = apiKey[0]
|
||
}
|
||
if s.cloudMode && s.email != nil {
|
||
s.email.SetCloudMode(true)
|
||
}
|
||
}
|
||
|
||
// SetCORSOrigins configures allowed CORS origins (comma-separated).
|
||
func (s *Server) SetCORSOrigins(origins string) {
|
||
s.corsOrigins = origins
|
||
}
|
||
|
||
// SetAttachments wires the attachment storage Registry that the upload
|
||
// and download handlers use. Pass maxBytes = 0 to keep the
|
||
// defaultAttachmentMaxBytes ceiling (25 MiB).
|
||
func (s *Server) SetAttachments(reg *attachments.Registry, maxBytes int64) {
|
||
s.attachments = reg
|
||
s.attachmentMaxBytes = maxBytes
|
||
}
|
||
|
||
// SetImageProcessor wires the image processor that the upload handler
|
||
// uses to derive thumbnail variants (TASK-878). Optional — without it
|
||
// uploads still succeed but no thumbnails are generated; the
|
||
// download handler's variant fallback path returns the original blob.
|
||
// The capabilities endpoint reflects whichever processor is wired.
|
||
func (s *Server) SetImageProcessor(p attachments.Processor) {
|
||
s.imageProcessor = p
|
||
}
|
||
|
||
// markUploadInFlight increments the in-flight counter for a content
|
||
// hash. Returns a release func the caller MUST defer; the release
|
||
// decrements and removes the entry once it hits zero. Used by the
|
||
// upload handler to fence Put + CreateAttachment against orphan-GC
|
||
// blob deletions of the same hash.
|
||
//
|
||
// Increment + map-store + decrement + delete all run under one
|
||
// mutex so a concurrent uploadInFlight call can't observe a stale
|
||
// "0" between the last release-decrement and the next-upload
|
||
// increment. The earlier sync.Map version split increment from
|
||
// LoadOrStore-then-atomic-add and missed that window (Codex P1 on
|
||
// PR #307 round 2).
|
||
func (s *Server) markUploadInFlight(hash string) func() {
|
||
s.inFlightHashesMu.Lock()
|
||
if s.inFlightHashes == nil {
|
||
s.inFlightHashes = make(map[string]int64)
|
||
}
|
||
s.inFlightHashes[hash]++
|
||
s.inFlightHashesMu.Unlock()
|
||
return func() {
|
||
s.inFlightHashesMu.Lock()
|
||
defer s.inFlightHashesMu.Unlock()
|
||
s.inFlightHashes[hash]--
|
||
if s.inFlightHashes[hash] <= 0 {
|
||
delete(s.inFlightHashes, hash)
|
||
}
|
||
}
|
||
}
|
||
|
||
// uploadInFlight reports whether any upload is currently materializing
|
||
// a blob with the given hash. The orphan GC consults this before
|
||
// deleting a blob — if an upload just finished Put but hasn't
|
||
// inserted the row yet, GC must NOT reclaim the blob.
|
||
func (s *Server) uploadInFlight(hash string) bool {
|
||
s.inFlightHashesMu.Lock()
|
||
defer s.inFlightHashesMu.Unlock()
|
||
return s.inFlightHashes[hash] > 0
|
||
}
|
||
|
||
// SetImportBundleMaxBytes overrides the default 2 GiB cap on a
|
||
// single workspace import bundle. Set to 0 to fall back to the
|
||
// default. Wired from PAD_IMPORT_BUNDLE_MAX_BYTES in cmd/pad/main.go
|
||
// so operators with workspaces over 2 GiB can opt in without
|
||
// recompiling. Larger caps trade memory headroom (one blob in
|
||
// flight at a time, ≤25 MiB) for a longer import wall-clock.
|
||
func (s *Server) SetImportBundleMaxBytes(n int64) {
|
||
s.importBundleMaxBytes = n
|
||
}
|
||
|
||
// SetImportArtifactMaxBytes overrides the default 1 MiB cap on a single
|
||
// playbook/convention artifact import. Set to 0 to fall back to the
|
||
// default. Wired from PAD_IMPORT_ARTIFACT_MAX_BYTES in cmd/pad/main.go.
|
||
func (s *Server) SetImportArtifactMaxBytes(n int64) {
|
||
s.importArtifactMaxBytes = n
|
||
}
|
||
|
||
// SetSecureCookies enables the Secure flag on all cookies.
|
||
func (s *Server) SetSecureCookies(secure bool) {
|
||
s.secureCookies = secure
|
||
}
|
||
|
||
// SetMetrics attaches Prometheus metrics to the server.
|
||
// Must be called before the first request is served.
|
||
//
|
||
// Side effect (TASK-961): when both metrics AND the OAuth server are
|
||
// wired, this also attaches the OAuth-active-tokens callback collector
|
||
// and the revocation TTL observer. Order-independent — both
|
||
// SetMetrics and SetOAuthServer call wireOAuthMetricsObserver, which
|
||
// no-ops until both prerequisites are present.
|
||
func (s *Server) SetMetrics(m *metrics.Metrics) {
|
||
s.metrics = m
|
||
s.wireOAuthMetricsObserver()
|
||
}
|
||
|
||
// wireOAuthMetricsObserver attaches the OAuth metrics that need both
|
||
// the metrics registry AND the OAuth server: the active-tokens
|
||
// callback collector (reads via the store) and the per-revocation
|
||
// TTL observer (fires from internal/oauth/storage.go on every
|
||
// access-token family revocation).
|
||
//
|
||
// Idempotent — re-registering the same collector would panic via
|
||
// prometheus.MustRegister, so we guard with a flag. Setting the
|
||
// observer multiple times is harmless (just replaces the function
|
||
// pointer).
|
||
//
|
||
// Why this lives on Server rather than in cmd/pad: it composes two
|
||
// optional Server fields whose set-order isn't guaranteed by the
|
||
// boot sequence, and centralizing the wiring here keeps the cmd/pad
|
||
// startup path declarative ("set X, set Y") without an explicit
|
||
// "now wire the cross-cut" call.
|
||
func (s *Server) wireOAuthMetricsObserver() {
|
||
if s.metrics == nil || s.oauthServer == nil {
|
||
return
|
||
}
|
||
if !s.oauthMetricsWired {
|
||
s.metrics.RegisterOAuthActiveTokensCollector(s.store.CountActiveOAuthAccessTokens)
|
||
s.oauthMetricsWired = true
|
||
}
|
||
s.oauthServer.Storage().SetRevocationObserver(func(kind string, ttl time.Duration) {
|
||
s.metrics.OAuthTokenRevocationsTotal.WithLabelValues(kind).Inc()
|
||
s.metrics.OAuthTokenTTLSeconds.Observe(ttl.Seconds())
|
||
})
|
||
}
|
||
|
||
// SetMetricsToken configures the static bearer token required to scrape
|
||
// /metrics. When empty (the default), /metrics is exposed only to loopback
|
||
// callers so a self-hosted Prometheus on the same host keeps working
|
||
// without config — but LAN/internet scrapes are refused. A non-empty
|
||
// token requires "Authorization: Bearer <token>" regardless of source.
|
||
func (s *Server) SetMetricsToken(token string) {
|
||
s.metricsToken = strings.TrimSpace(token)
|
||
}
|
||
|
||
// metricsAuth gates the /metrics endpoint. See SetMetricsToken for the
|
||
// policy. Uses constant-time comparison to avoid leaking the configured
|
||
// token via response timing.
|
||
func (s *Server) metricsAuth(next http.Handler) http.Handler {
|
||
return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
|
||
if s.metricsToken == "" {
|
||
// No token configured → loopback-only access.
|
||
if !requestIsLoopback(r) {
|
||
writeError(w, http.StatusForbidden, "forbidden",
|
||
"/metrics is restricted to loopback when PAD_METRICS_TOKEN is unset")
|
||
return
|
||
}
|
||
next.ServeHTTP(w, r)
|
||
return
|
||
}
|
||
|
||
const prefix = "Bearer "
|
||
authHeader := r.Header.Get("Authorization")
|
||
if !strings.HasPrefix(authHeader, prefix) {
|
||
w.Header().Set("WWW-Authenticate", `Bearer realm="metrics"`)
|
||
writeError(w, http.StatusUnauthorized, "unauthorized",
|
||
"Missing Bearer token for /metrics")
|
||
return
|
||
}
|
||
given := strings.TrimSpace(strings.TrimPrefix(authHeader, prefix))
|
||
if subtle.ConstantTimeCompare([]byte(given), []byte(s.metricsToken)) != 1 {
|
||
w.Header().Set("WWW-Authenticate", `Bearer realm="metrics"`)
|
||
writeError(w, http.StatusUnauthorized, "unauthorized",
|
||
"Invalid Bearer token for /metrics")
|
||
return
|
||
}
|
||
next.ServeHTTP(w, r)
|
||
})
|
||
}
|
||
|
||
// SetSSELimits configures global and per-workspace SSE connection limits.
|
||
// A value of 0 means unlimited.
|
||
func (s *Server) SetSSELimits(global, perWorkspace int) {
|
||
s.sseMaxConnections = global
|
||
s.sseMaxPerWorkspace = perWorkspace
|
||
}
|
||
|
||
// SetTrustedProxies configures which direct TCP peers are allowed to set
|
||
// X-Real-IP / X-Forwarded-For on incoming requests. Accepts a comma-
|
||
// separated list of CIDRs or bare IPs (e.g. "10.0.0.0/8, 172.16.0.0/12").
|
||
// When empty (the default), proxy headers are ignored entirely — the
|
||
// actual TCP peer address is used for rate limiting, the bootstrap
|
||
// loopback check, and audit logging.
|
||
func (s *Server) SetTrustedProxies(spec string) {
|
||
s.trustedProxyCIDRs = ParseTrustedProxyCIDRs(spec)
|
||
}
|
||
|
||
// SetIPChangeEnforce controls how the auth middleware reacts when a
|
||
// session's binding (client IP OR User-Agent hash) changes mid-lifetime:
|
||
// - mode == "strict": revoke the session and reject the request (the token
|
||
// is treated as possibly stolen). Covers BOTH the IP and the UA signal —
|
||
// one flag arms the whole session-binding enforcement.
|
||
// - anything else (default): log to the audit log, update the stored IP,
|
||
// and let the request through. Strict mode breaks legitimate mobility
|
||
// (mobile roaming, VPN toggles for IP; browser/WebView updates for UA) so
|
||
// it is opt-in for high-sensitivity deployments via the
|
||
// PAD_IP_CHANGE_ENFORCE env var. See handleSessionIPChange /
|
||
// handleSessionUAChange for the per-signal semantics.
|
||
func (s *Server) SetIPChangeEnforce(mode string) {
|
||
s.ipChangeEnforceStrict = strings.EqualFold(strings.TrimSpace(mode), "strict")
|
||
}
|
||
|
||
// reconfigureEmail reads email settings from the platform_settings table
|
||
// and updates (or creates) the email sender. Called after admin settings change.
|
||
func (s *Server) reconfigureEmail() {
|
||
apiKey, _ := s.store.GetPlatformSetting(settingMailerooAPIKey)
|
||
fromAddr, _ := s.store.GetPlatformSetting(settingEmailFrom)
|
||
fromName, _ := s.store.GetPlatformSetting(settingEmailFromName)
|
||
|
||
if apiKey == "" {
|
||
// No platform-settings key. If email was wired from env vars, leave that
|
||
// sender in place — env is the deployment baseline. Otherwise the admin
|
||
// cleared the only email config (e.g. selecting provider "None"), so tear
|
||
// down the live sender: clearing the DB key alone left the running process
|
||
// sending mail until restart (BUG-1890).
|
||
if !s.emailEnvConfigured {
|
||
s.email = nil
|
||
s.emailAPIKey = ""
|
||
}
|
||
return
|
||
}
|
||
|
||
s.emailAPIKey = apiKey
|
||
if s.email == nil {
|
||
// Create a new sender from platform settings
|
||
s.email = email.NewSender(apiKey, fromAddr, fromName, s.baseURL)
|
||
} else {
|
||
// Update existing sender
|
||
s.email.Configure(apiKey, fromAddr, fromName, s.baseURL)
|
||
}
|
||
// Propagate cloud mode whichever way email was wired — see SetEmailSender
|
||
// for the matching note. Configure() preserves cloudMode on existing
|
||
// senders since SetCloudMode is independent; this branch covers the
|
||
// fresh-NewSender path.
|
||
if s.cloudMode {
|
||
s.email.SetCloudMode(true)
|
||
}
|
||
}
|
||
|
||
// InitEmailFromSettings loads email config from platform settings on startup,
|
||
// merging with any env-var-based sender that was already attached.
|
||
func (s *Server) InitEmailFromSettings() {
|
||
s.reconfigureEmail()
|
||
}
|
||
|
||
func (s *Server) setupRouter() {
|
||
r := chi.NewRouter()
|
||
|
||
// Infrastructure middleware (applies to all routes including /metrics)
|
||
// CapturePeerAddr MUST run before TrustedProxyRealIP so downstream code
|
||
// that needs to verify the real TCP peer (e.g. the bootstrap loopback
|
||
// check) can read the untampered value from request context even on
|
||
// deployments with a trusted reverse proxy in front.
|
||
r.Use(CapturePeerAddr)
|
||
// RealIP is gated on PAD_TRUSTED_PROXIES. When unset (the default), proxy
|
||
// headers are ignored and the real TCP peer address is used everywhere.
|
||
// This prevents X-Forwarded-For spoofing from bypassing rate limits, the
|
||
// bootstrap loopback check, or audit logs on direct-exposed deployments.
|
||
r.Use(TrustedProxyRealIP(s.trustedProxyCIDRs))
|
||
r.Use(chimiddleware.RequestID)
|
||
r.Use(StructuredLogger)
|
||
if s.metrics != nil {
|
||
r.Use(MetricsMiddleware(s.metrics))
|
||
}
|
||
r.Use(chimiddleware.Recoverer)
|
||
|
||
// Security headers (applies to all routes)
|
||
r.Use(SecurityHeaders)
|
||
if s.secureCookies {
|
||
r.Use(StrictTransportSecurity)
|
||
}
|
||
|
||
// MCP Streamable HTTP transport + OAuth discovery endpoints
|
||
// (PLAN-943 TASK-950). Mounted outside the standard /api/v1
|
||
// auth-required group because:
|
||
//
|
||
// - /mcp uses Bearer auth via its own MCPBearerAuth middleware,
|
||
// producing the spec-shape 401 + WWW-Authenticate that MCP
|
||
// clients expect (the API-stack 401 envelope is JSON-only and
|
||
// would fail Claude Desktop's discovery handshake).
|
||
// - /.well-known/oauth-protected-resource and
|
||
// /.well-known/oauth-authorization-server are public discovery
|
||
// documents (RFC 9728 / RFC 8414); routing them through
|
||
// TokenAuth+SessionAuth+RequireAuth would 401 unauth probes.
|
||
//
|
||
// No-op when SetMCPTransport hasn't been called or cloud mode is
|
||
// off — see registerMCPRoutes for the gating.
|
||
s.registerMCPRoutes(r)
|
||
|
||
// OAuth 2.1 authorization-server flow endpoints (PLAN-943
|
||
// TASK-1025 sub-PR C). /oauth/{register,authorize,token,
|
||
// authorize/decide} mounted alongside /mcp + /.well-known/*,
|
||
// outside /api/v1's auth-required group. CSRF middleware runs
|
||
// only on /api/* paths so /oauth/* is naturally exempt; the
|
||
// consent-decision endpoint adds its own form-token check
|
||
// using the existing __Host-pad_csrf cookie.
|
||
//
|
||
// SessionAuth runs in this group so /oauth/authorize can detect
|
||
// whether the user is logged in via the __Host-pad_session
|
||
// cookie. SessionAuth falls through gracefully when no cookie
|
||
// is present (handlers see currentUser(r)==nil and redirect to
|
||
// /login). RequireAuth is intentionally NOT used — /oauth/authorize
|
||
// must be reachable anonymously to trigger the login redirect.
|
||
//
|
||
// RateLimit gates /oauth/register specifically (per Codex review
|
||
// #372 round 2 — the DCR endpoint is open by RFC 7591 design,
|
||
// but unlimited writes to oauth_clients are an obvious DoS
|
||
// surface). The middleware short-circuits other /oauth/* paths
|
||
// because they're either session-bound or PKCE-bound; explicit
|
||
// per-endpoint limits arrive with TASK-959.
|
||
//
|
||
// No-op when SetOAuthServer hasn't been called or cloud mode is off.
|
||
r.Group(func(r chi.Router) {
|
||
r.Use(s.requireCloudMode)
|
||
r.Use(s.SessionAuth)
|
||
r.Use(s.RateLimit)
|
||
s.registerOAuthRoutes(r)
|
||
})
|
||
|
||
// Prometheus scrape endpoint — exempt from the standard auth/CSRF stack
|
||
// (Prometheus can't present a session cookie or pass a CSRF header), but
|
||
// gated by a dedicated static bearer token. Without the gate, any
|
||
// unauthenticated caller on the network can read workspace counts, API
|
||
// usage patterns, and — via label enumeration — user/workspace IDs.
|
||
//
|
||
// The gate runs in three layers:
|
||
// 1. No PAD_METRICS_TOKEN → endpoint is open ONLY to loopback. Safe
|
||
// default for self-hosters running Prometheus on the same box.
|
||
// 2. PAD_METRICS_TOKEN set → "Authorization: Bearer <token>" required.
|
||
// Compared in constant time; empty/missing header → 401.
|
||
// 3. In either case the SecurityHeaders / rate-limit / logging chain
|
||
// already wraps this group from the outer r.Use() calls above.
|
||
if s.metrics != nil {
|
||
r.Group(func(r chi.Router) {
|
||
r.Use(s.metricsAuth)
|
||
r.Handle("/metrics", promhttp.HandlerFor(s.metrics.Registry, promhttp.HandlerOpts{}))
|
||
})
|
||
}
|
||
|
||
// All other routes — full middleware stack
|
||
r.Group(func(r chi.Router) {
|
||
r.Use(cors.Handler(cors.Options{
|
||
AllowedOrigins: parseCORSOrigins(s.corsOrigins),
|
||
AllowedMethods: []string{"GET", "POST", "PATCH", "PUT", "DELETE", "OPTIONS"},
|
||
AllowedHeaders: []string{"Accept", "Authorization", "Content-Type", "X-CSRF-Token", "X-Share-Password", "X-Bootstrap-Token"},
|
||
// Credentials flag is gated on an operator explicitly listing
|
||
// PAD_CORS_ORIGINS. The CLI uses Bearer tokens so the default
|
||
// "no CORS_ORIGINS set" path doesn't need credential sharing;
|
||
// leaving it off by default prevents cross-origin fetches from
|
||
// a browser on a different site from piggy-backing cookies
|
||
// on the victim's session.
|
||
AllowCredentials: corsAllowCredentials(s.corsOrigins),
|
||
MaxAge: 300,
|
||
}))
|
||
r.Use(s.TokenAuth)
|
||
r.Use(s.SessionAuth)
|
||
r.Use(s.RateLimit)
|
||
r.Use(s.CSRFProtect)
|
||
r.Use(s.RequireAuth)
|
||
// PLAN-1933 DR-4: block content-mutating requests from an
|
||
// authenticated cloud user whose email is unverified. Mounted
|
||
// AFTER RequireAuth so currentUser is already resolved; a no-op
|
||
// on self-host and for verified / unauthenticated callers. The
|
||
// method gate here covers the /api/v1 surface (session + PAT);
|
||
// the collab GET-upgrade, the OAuth-provider flow, and the MCP
|
||
// write path are gated at their own out-of-band mounts.
|
||
r.Use(s.RequireVerifiedEmail)
|
||
r.Use(jsonContentType)
|
||
|
||
// SSE endpoint (outside jsonContentType middleware — but inherits auth)
|
||
r.Get("/api/v1/events", s.handleSSE)
|
||
|
||
// User-scoped watch/nudge event stream (TASK-2533, DOC-2479).
|
||
// Unlike /api/v1/events above, this is NOT workspace-scoped — a
|
||
// caller's watches and addressed pushes can span
|
||
// every workspace they belong to. Lives alongside the other SSE
|
||
// endpoint for the same "outside jsonContentType, inherits auth"
|
||
// reason.
|
||
r.Get("/api/v1/events/stream", s.handleWatchEventsStream)
|
||
|
||
// Live-session presence (PLAN-2558 S1) — the READ side of the
|
||
// registry the stream above writes. Mounted here beside the
|
||
// endpoint it reports on rather than in the /api/v1 Route block
|
||
// below: the two are one feature, and a reader asking "what
|
||
// fills this list?" should find the answer on the adjacent
|
||
// line. Self-scoped; see handleListSessions on why there is
|
||
// deliberately no admin view.
|
||
r.Get("/api/v1/sessions", s.handleListSessions)
|
||
|
||
// WebSocket endpoint for Yjs-based collaborative editing on a
|
||
// single item (PLAN-1248). Lives outside jsonContentType for
|
||
// the same reason as SSE: the response is a WS upgrade, not
|
||
// JSON. Inherits the auth middleware chain — handleCollab
|
||
// then re-checks workspace access keyed on the item's
|
||
// workspace ID (the URL only carries itemID).
|
||
r.Get("/api/v1/collab/{itemID}", s.handleCollab)
|
||
|
||
// API routes
|
||
r.Route("/api/v1", func(r chi.Router) {
|
||
r.Get("/health", s.handleHealth)
|
||
r.Get("/health/live", s.handleHealthLive)
|
||
r.Get("/health/ready", s.handleHealthReady)
|
||
r.Get("/plan-limits", s.handleGetPlanLimits) // Public: billing page reads plan limits
|
||
r.Get("/unsubscribe", s.handleUnsubscribe) // Public: email opt-out (HMAC-signed)
|
||
|
||
// Server capabilities — public so the editor can fetch it
|
||
// pre-login and gate per-format rotate / crop UI on the
|
||
// processor's reach (TASK-878). The response is static for
|
||
// the lifetime of the binary; clients can cache freely.
|
||
r.Get("/server/capabilities", s.handleServerCapabilities)
|
||
|
||
// Auth endpoints (exempt from auth middleware)
|
||
r.Route("/auth", func(r chi.Router) {
|
||
r.Get("/session", s.handleSessionCheck)
|
||
r.Post("/bootstrap", s.handleBootstrap)
|
||
r.Post("/register", s.handleRegister)
|
||
r.Get("/check-username", s.handleCheckUsername)
|
||
r.Post("/login", s.handleLogin)
|
||
r.Post("/logout", s.handleLogout)
|
||
r.Get("/me", s.handleGetCurrentUser)
|
||
r.Patch("/me", s.handleUpdateCurrentUser)
|
||
|
||
// Password reset
|
||
r.Post("/forgot-password", s.handleForgotPassword)
|
||
r.Post("/reset-password", s.handleResetPassword)
|
||
// Localhost-only recovery escape hatch (self-host, non-cloud).
|
||
r.Post("/local-reset", s.handleLocalReset)
|
||
|
||
// Email verification (PLAN-1933 Wave 3b). Both are
|
||
// enumeration-safe and rate-limited (middleware_ratelimit.go
|
||
// reuses the PasswordReset bucket) and are already in
|
||
// RequireVerifiedEmail's exempt list so an unverified user
|
||
// can reach them to clear their own unverified state.
|
||
r.Post("/verify-email", s.handleVerifyEmail)
|
||
r.Post("/resend-verification", s.handleResendVerification)
|
||
|
||
// Two-factor authentication
|
||
r.Post("/2fa/setup", s.handleTOTPSetup)
|
||
r.Post("/2fa/verify", s.handleTOTPVerify)
|
||
r.Post("/2fa/disable", s.handleTOTPDisable)
|
||
r.Post("/2fa/login-verify", s.handleTOTPLoginVerify)
|
||
|
||
// Account management (GDPR)
|
||
r.Post("/delete-account", s.handleDeleteAccount)
|
||
r.Get("/export", s.handleExportAccount)
|
||
|
||
// User-scoped API tokens
|
||
r.Get("/tokens", s.handleListUserTokens)
|
||
r.Post("/tokens", s.handleCreateUserToken)
|
||
r.Delete("/tokens/{tokenID}", s.handleDeleteUserToken)
|
||
r.Post("/tokens/{tokenID}/rotate", s.handleRotateUserToken)
|
||
|
||
// Cloud: OAuth login/linking (called by pad-cloud sidecar, protected by cloud secret)
|
||
r.Post("/oauth-login", s.handleOAuthLogin)
|
||
r.Post("/oauth-link", s.handleOAuthLink)
|
||
r.Post("/oauth-unlink", s.handleOAuthUnlink)
|
||
|
||
// CLI browser-based auth flow
|
||
r.Post("/cli/sessions", s.handleCreateCLIAuthSession)
|
||
r.Get("/cli/sessions/{code}", s.handlePollCLIAuthSession)
|
||
r.Post("/cli/sessions/{code}/approve", s.handleApproveCLIAuthSession)
|
||
})
|
||
|
||
// Admin endpoints (admin-only, handlers check role internally)
|
||
r.Route("/admin", func(r chi.Router) {
|
||
r.Get("/settings", s.handleGetPlatformSettings)
|
||
r.Patch("/settings", s.handleUpdatePlatformSettings)
|
||
r.Post("/test-email", s.handleTestEmail)
|
||
|
||
// Cloud sidecar endpoints — only exist in cloud mode. requireCloudMode
|
||
// returns 404 outside cloud mode so a self-hosted deployment doesn't
|
||
// expose "Cloud mode not configured" to unauthenticated probes.
|
||
r.Group(func(r chi.Router) {
|
||
r.Use(s.requireCloudMode)
|
||
r.Post("/plan", s.handleSetPlan) // Cloud: sidecar sets user plans; also accessible to admins
|
||
r.Post("/stripe-customer-id", s.handleSetStripeCustomerID) // Cloud: sidecar stores Stripe customer ID after checkout
|
||
r.Get("/user-by-customer", s.handleGetUserByCustomerID) // Cloud: sidecar looks up user by Stripe customer ID
|
||
r.Post("/stripe-event-processed", s.handleStripeEventProcessed) // Cloud: sidecar webhook idempotency (TASK-696)
|
||
r.Post("/stripe-event-unmark", s.handleStripeEventUnmark) // Cloud: sidecar handler-failure rollback (TASK-736)
|
||
r.Post("/payment-failed", s.handlePaymentFailed) // Cloud: sidecar forwards invoice.payment_failed to trigger email (TASK-712)
|
||
|
||
// Admin Billing dashboard data (TASK-827 / PLAN-825). Proxies
|
||
// pad-cloud's /admin/metrics/billing for Stripe-derived stats
|
||
// (active subs, MRR, ARR, churn) and merges with local
|
||
// users-table aggregates (customers_by_plan, new_signups_30d).
|
||
// Always returns 200; degraded states (sidecar unreachable,
|
||
// Stripe not configured) are surfaced as flags in the body.
|
||
r.Get("/billing-stats", s.handleAdminBillingStats)
|
||
})
|
||
|
||
// User management
|
||
r.Get("/users", s.handleAdminListUsers)
|
||
r.Get("/users/{userID}", s.handleAdminGetUser)
|
||
r.Patch("/users/{userID}", s.handleAdminUpdateUser)
|
||
r.Post("/users/{userID}/reset-password", s.handleAdminResetPassword)
|
||
r.Get("/users/{userID}/workspaces", s.handleAdminGetUserWorkspaces)
|
||
r.Get("/users/{userID}/detail", s.handleAdminGetUserDetail)
|
||
r.Get("/users/{userID}/activity", s.handleAdminGetUserActivity)
|
||
r.Get("/users/{userID}/metrics", s.handleAdminGetUserMetrics)
|
||
r.Post("/users/{userID}/disable", s.handleAdminDisableUser)
|
||
r.Post("/users/{userID}/enable", s.handleAdminEnableUser)
|
||
r.Post("/users/{userID}/verify-email", s.handleAdminVerifyEmail)
|
||
|
||
// Invitations
|
||
r.Get("/invitations", s.handleAdminListInvitations)
|
||
r.Post("/invitations/{invID}/resend", s.handleAdminResendInvitation)
|
||
r.Delete("/invitations/{invID}", s.handleAdminDeleteInvitation)
|
||
|
||
// Plan limits
|
||
r.Get("/limits", s.handleAdminGetLimits)
|
||
r.Patch("/limits", s.handleAdminUpdateLimits)
|
||
|
||
// Platform stats
|
||
r.Get("/stats", s.handleAdminStats)
|
||
|
||
// MCP audit log — admin-only full-table view (TASK-960).
|
||
// Powers /console/admin/mcp-audit. Per-connection
|
||
// drilldown that users see for their own connections
|
||
// lives at /api/v1/connected-apps/{id}/audit (registered
|
||
// outside the admin group so non-admin users can read
|
||
// their own).
|
||
r.Get("/mcp-audit", s.handleAdminMCPAudit)
|
||
})
|
||
|
||
// Audit log (admin-only)
|
||
r.Get("/audit-log", s.handleAuditLog)
|
||
|
||
// MCP per-connection audit (TASK-960). Owner-only via the
|
||
// store query (user_id is one of the WHERE clauses);
|
||
// returns the requesting user's own MCP activity for one
|
||
// connection. The handler runs inside the standard
|
||
// /api/v1 auth-required group, so unauthenticated callers
|
||
// 401 here just like every other API endpoint.
|
||
r.Get("/connected-apps/{id}/audit", s.handleMCPConnectionAudit)
|
||
|
||
// Connected-apps management (TASK-954). Lists every
|
||
// active OAuth grant chain the user has authorized
|
||
// (Claude Desktop, Cursor, …) and lets them revoke one.
|
||
// Cloud-mode-gated because OAuth is a cloud-only
|
||
// surface — self-hosted deployments would always see
|
||
// an empty list.
|
||
r.Group(func(r chi.Router) {
|
||
r.Use(s.requireCloudMode)
|
||
r.Get("/connected-apps", s.handleListConnectedApps)
|
||
r.Delete("/connected-apps/{id}", s.handleRevokeConnectedApp)
|
||
// PLAN-1519 / TASK-1524 / IDEA-1517 §3: mutation
|
||
// endpoints for the connections-page UI. Per-field
|
||
// patches rather than a general PATCH for cleaner
|
||
// error envelopes + audit shape.
|
||
r.Patch("/connected-apps/{id}/name", s.handleRenameConnectedApp)
|
||
r.Patch("/connected-apps/{id}/flags", s.handleUpdateConnectedAppFlags)
|
||
r.Post("/connected-apps/{id}/workspaces", s.handleAddConnectedAppWorkspace)
|
||
r.Delete("/connected-apps/{id}/workspaces/{slug}", s.handleRemoveConnectedAppWorkspace)
|
||
})
|
||
|
||
// Templates
|
||
r.Get("/templates", s.handleListTemplates)
|
||
|
||
// Convention Library
|
||
r.Get("/convention-library", s.handleConventionLibrary)
|
||
|
||
// Playbook Library
|
||
r.Get("/playbook-library", s.handlePlaybookLibrary)
|
||
|
||
// Single library entry by title (conventions first, then playbooks).
|
||
// TASK-1561 / PLAN-1560.
|
||
r.Get("/library/entry", s.handleLibraryEntry)
|
||
|
||
// URL import — fetch a remote page and return markdown.
|
||
// Side-effect-free; the client decides what to do with the
|
||
// markdown. See PLAN-1467 / TASK-1472 / internal/urlimport.
|
||
r.Post("/import/url", s.handleImportURL)
|
||
|
||
// Invitations (outside workspace scope)
|
||
r.Post("/invitations/{code}/accept", s.handleAcceptInvitation)
|
||
|
||
// Non-consuming invitation preview (BUG-1934). Public/pre-auth
|
||
// (exempted in isPublicAPIPath) so the logged-out /join page can
|
||
// prefill the invited email read-only and pick register-vs-login
|
||
// mode. Always HTTP 200 + rate limited (see middleware_ratelimit.go)
|
||
// so it can't be used to enumerate invite codes.
|
||
r.Get("/invitations/{code}/preview", s.handlePreviewInvitation)
|
||
|
||
// OAuth client public-info (PLAN-943 TASK-1027 sub-PR E).
|
||
// Read-only consent-screen support for OAuth clients
|
||
// registered via /oauth/register. Auth-required (inherits
|
||
// RequireAuth from the parent group); cloud-mode-gated so
|
||
// self-hosted deployments without an OAuth server don't
|
||
// expose a hollow endpoint. Returns four non-sensitive
|
||
// fields (client_id, client_name, logo_uri, redirect_uris)
|
||
// — see handlers_oauth_clients.go for the full leak-surface
|
||
// rationale.
|
||
r.Group(func(r chi.Router) {
|
||
r.Use(s.requireCloudMode)
|
||
r.Get("/oauth/clients/{id}/public-info", s.handleOAuthClientPublicInfo)
|
||
})
|
||
|
||
// Share link resolution (outside workspace scope, no auth required)
|
||
r.Get("/s/{token}", s.handleResolveShareLink)
|
||
// Share-link asset bytes (BUG-2389 2b / TASK-2637): rendered image
|
||
// VARIANTS for attachments embedded in the shared content. Same
|
||
// public/no-auth group; protected links gate on a short-lived
|
||
// signed ref minted by handleResolveShareLink. Originals and
|
||
// file downloads are out of scope by authorization.
|
||
r.Get("/s/{token}/attachments/{attachmentID}", s.handleGetShareLinkAttachment)
|
||
|
||
// Claim-code redemption (PLAN-1519 / TASK-1521 / IDEA-1517 §4).
|
||
// POST /api/v1/oauth/claim with body {workspace, code} grants
|
||
// the calling OAuth connection access to one workspace via a
|
||
// stateless 6-digit HMAC code the user generated in the web
|
||
// UI's "Connect project" modal. Auth: standard /api/v1 chain
|
||
// (TokenAuth + RequireAuth); the handler itself short-circuits
|
||
// the side effect when the caller isn't an OAuth grant (PAT /
|
||
// CLI session) and 412s when the claim secret isn't wired.
|
||
r.Post("/oauth/claim", s.handleOAuthClaim)
|
||
|
||
// Workspaces
|
||
r.Route("/workspaces", func(r chi.Router) {
|
||
r.Get("/", s.handleListWorkspaces)
|
||
r.Post("/", s.handleCreateWorkspace)
|
||
r.Post("/import", s.handleImportWorkspace)
|
||
r.Put("/reorder", s.handleReorderWorkspaces)
|
||
|
||
// Soft-delete recovery (PLAN-1969 / TASK-1970). Both live
|
||
// OUTSIDE the /{slug} RequireWorkspaceAccess subrouter
|
||
// because that middleware resolves only LIVE workspaces
|
||
// (deleted_at IS NULL) and would 404 a soft-deleted one
|
||
// before the handler ran. The static "/deleted" segment is
|
||
// registered before the /{slug} param route so chi matches
|
||
// it exactly (static beats param); it lists the caller's own
|
||
// deleted-but-restorable workspaces. "/{slug}/restore"
|
||
// resolves the soft-deleted row itself and enforces
|
||
// owner-only authz inside the handler.
|
||
r.Get("/deleted", s.handleListDeletedWorkspaces)
|
||
r.Post("/{slug}/restore", s.handleRestoreWorkspace)
|
||
|
||
r.Route("/{slug}", func(r chi.Router) {
|
||
r.Use(s.RequireWorkspaceAccess)
|
||
|
||
r.Get("/", s.handleGetWorkspace)
|
||
r.Patch("/", s.handleUpdateWorkspace)
|
||
r.Delete("/", s.handleDeleteWorkspace)
|
||
r.Get("/export", s.handleExportWorkspace)
|
||
// Import a single playbook/convention artifact (Markdown
|
||
// + YAML frontmatter) into this workspace. Editor+ gate
|
||
// is enforced inside the handler against the destination
|
||
// collection.
|
||
r.Post("/import-artifact", s.handleImportArtifact)
|
||
|
||
// Activity (workspace level)
|
||
r.Get("/activity", s.handleListWorkspaceActivity)
|
||
|
||
// Claim-code generation + smart suppression (PLAN-1519
|
||
// / TASK-1525 / IDEA-1517 §4). Inherits
|
||
// RequireWorkspaceAccess so any member can pull a code
|
||
// for any workspace they belong to — membership IS
|
||
// the consent. See handlers_claim_code.go.
|
||
r.Get("/claim-code", s.handleWorkspaceClaimCode)
|
||
|
||
// Documents (v1 — will be replaced by items in Phase 2)
|
||
r.Route("/documents", func(r chi.Router) {
|
||
r.Get("/", s.handleListDocuments)
|
||
r.Post("/", s.handleCreateDocument)
|
||
|
||
r.Route("/{docID}", func(r chi.Router) {
|
||
r.Get("/", s.handleGetDocument)
|
||
r.Patch("/", s.handleUpdateDocument)
|
||
r.Delete("/", s.handleDeleteDocument)
|
||
r.Post("/restore", s.handleRestoreDocument)
|
||
|
||
// Versions
|
||
r.Get("/versions", s.handleListVersions)
|
||
r.Get("/versions/{versionID}", s.handleGetVersion)
|
||
|
||
// Activity (document level)
|
||
r.Get("/activity", s.handleListDocumentActivity)
|
||
})
|
||
})
|
||
|
||
// Collections (v2)
|
||
r.Route("/collections", func(r chi.Router) {
|
||
r.Get("/", s.handleListCollections)
|
||
r.Post("/", s.handleCreateCollection)
|
||
r.Route("/{collSlug}", func(r chi.Router) {
|
||
r.Get("/", s.handleGetCollection)
|
||
r.Patch("/", s.handleUpdateCollection)
|
||
r.Delete("/", s.handleDeleteCollection)
|
||
// Items within collection
|
||
r.Get("/items", s.handleListCollectionItems)
|
||
r.Post("/items", s.handleCreateItem)
|
||
// Pairs with /items-index — server-side checkbox
|
||
// progress so the collection page can render
|
||
// list/board/table progress badges without
|
||
// fetching item content (TASK-1349).
|
||
r.Get("/checkbox-progress", s.handleCollectionCheckboxProgress)
|
||
// Child-item completion progress for any collection
|
||
// (BUG-1509). Same visibility/guest-grant semantics
|
||
// as /plans-progress but collection-generic.
|
||
r.Get("/child-progress", s.handleCollectionChildrenProgress)
|
||
// Collection grants
|
||
r.Get("/grants", s.handleListCollectionGrants)
|
||
r.Post("/grants", s.handleCreateCollectionGrant)
|
||
r.Delete("/grants/{grantID}", s.handleDeleteCollectionGrant)
|
||
r.Get("/share-links", s.handleListCollectionShareLinks)
|
||
r.Post("/share-links", s.handleCreateCollectionShareLink)
|
||
// Saved views within collection
|
||
r.Get("/views", s.handleListViews)
|
||
r.Post("/views", s.handleCreateView)
|
||
r.Route("/views/{viewID}", func(r chi.Router) {
|
||
r.Patch("/", s.handleUpdateView)
|
||
r.Delete("/", s.handleDeleteView)
|
||
})
|
||
})
|
||
})
|
||
|
||
// Plans progress
|
||
r.Get("/plans-progress", s.handlePlansProgress)
|
||
|
||
// Skinny-projection cross-collection items list for the
|
||
// local-first read model bootstrap (PLAN-1343 / TASK-1344).
|
||
// Lives at workspace level — sibling to /plans-progress
|
||
// and /starred — so the path can't ever collide with an
|
||
// item slug under /items/{itemSlug}.
|
||
r.Get("/items-index", s.handleListItemsIndex)
|
||
|
||
// Delta-fetch sibling of /items-index: returns rows
|
||
// where seq > since, including tombstones, so a
|
||
// local-first read-model client can resume without
|
||
// re-downloading the whole index (PLAN-1343 / TASK-1354).
|
||
r.Get("/items-changes", s.handleListItemsChanges)
|
||
|
||
// User grants (all grants for a specific user in this workspace)
|
||
r.Get("/users/{userID}/grants", s.handleListUserGrants)
|
||
|
||
// Starred items
|
||
r.Get("/starred", s.handleListStarredItems)
|
||
|
||
// Distinct tags across the workspace (with item counts)
|
||
r.Get("/tags", s.handleListTags)
|
||
|
||
// Items (cross-collection, v2)
|
||
r.Get("/items", s.handleListItems)
|
||
// Bulk mutation (TASK-1668). Static segment must be
|
||
// registered before the /items/{itemSlug} param route
|
||
// so "bulk" isn't captured as an item slug.
|
||
r.Post("/items/bulk", s.handleBulkItems)
|
||
r.Route("/items/{itemSlug}", func(r chi.Router) {
|
||
r.Get("/", s.handleGetItem)
|
||
r.Patch("/", s.handleUpdateItem)
|
||
r.Delete("/", s.handleDeleteItem)
|
||
r.Post("/restore", s.handleRestoreItem)
|
||
r.Post("/move", s.handleMoveItem)
|
||
// Cross-workspace copy PREFLIGHT (PLAN-2357 /
|
||
// TASK-2364). Reports what a copy into another
|
||
// workspace would carry, drop and need, and
|
||
// leaves no trace a copy would have left — see
|
||
// handlers_items_copy_preflight.go for the exact
|
||
// scope of that guarantee. POST because the
|
||
// request carries a body (destination + override
|
||
// map), not because it mutates. The mutating
|
||
// sibling lands at /copy in TASK-2365.
|
||
r.Post("/copy/preflight", s.handleCopyItemPreflight)
|
||
// Cross-workspace copy, the MUTATION (PLAN-2357 /
|
||
// TASK-2365). Same request shape as the preflight
|
||
// above; with archive_source it is the move. Post-
|
||
// commit fanout is asymmetric — see
|
||
// handlers_items_copy.go. Registered after the more
|
||
// specific /copy/preflight, though chi's trie makes
|
||
// the order immaterial.
|
||
r.Post("/copy", s.handleCopyItem)
|
||
// Export a single playbook/convention item as a
|
||
// portable artifact (Markdown + YAML frontmatter).
|
||
// Gated by per-item visibility, not the workspace-
|
||
// export owner gate — a viewer who can see the item
|
||
// may export it.
|
||
r.Get("/export", s.handleExportItemArtifact)
|
||
r.Get("/versions", s.handleListItemVersions)
|
||
r.Get("/versions/{versionID}", s.handleGetItemVersion)
|
||
r.Post("/versions/{versionID}/restore", s.handleRestoreItemVersion)
|
||
r.Get("/activity", s.handleListItemActivity)
|
||
r.Get("/links", s.handleGetItemLinks)
|
||
r.Post("/links", s.handleCreateItemLink)
|
||
r.Get("/comments", s.handleListComments)
|
||
r.Post("/comments", s.handleCreateComment)
|
||
r.Get("/timeline", s.handleListItemTimeline)
|
||
r.Get("/children", s.handleGetItemChildren)
|
||
r.Get("/progress", s.handleGetItemProgress)
|
||
r.Get("/backlinks", s.handleGetItemBacklinks)
|
||
r.Get("/tasks", s.handleGetItemChildren) // deprecated alias
|
||
r.Get("/grants", s.handleListItemGrants)
|
||
r.Post("/grants", s.handleCreateItemGrant)
|
||
r.Delete("/grants/{grantID}", s.handleDeleteItemGrant)
|
||
r.Get("/share-links", s.handleListItemShareLinks)
|
||
r.Post("/share-links", s.handleCreateItemShareLink)
|
||
// Stars
|
||
r.Get("/star", s.handleGetItemStarStatus)
|
||
r.Post("/star", s.handleStarItem)
|
||
r.Delete("/star", s.handleUnstarItem)
|
||
// Watches (TASK-2533): durable per-item subscriptions
|
||
// for the padd event-stream / plugin-monitor nudge
|
||
// pipeline. `pad watch <ref>` / `pad watch remove <ref>`.
|
||
r.Post("/watch", s.handleCreateWatch)
|
||
r.Delete("/watch", s.handleDeleteWatch)
|
||
// Push (IDEA-2544 Phase 1): transient, self-addressed
|
||
// human→harness dispatch over the SAME watch-events
|
||
// bus/stream — no durable row, see handlePushToItem's
|
||
// doc comment. `pad push <ref> -m "message"`.
|
||
r.Post("/push", s.handlePushToItem)
|
||
})
|
||
|
||
// Links (v2)
|
||
r.Delete("/links/{linkID}", s.handleDeleteItemLink)
|
||
|
||
// Share links (workspace-scoped management)
|
||
r.Delete("/share-links/{linkID}", s.handleDeleteShareLink)
|
||
r.Get("/share-links/{linkID}/views", s.handleShareLinkViews)
|
||
|
||
// Comments (v2)
|
||
r.Route("/comments/{commentID}", func(r chi.Router) {
|
||
r.Patch("/", s.handleUpdateComment)
|
||
r.Delete("/", s.handleDeleteComment)
|
||
r.Post("/replies", s.handleCreateReply)
|
||
r.Post("/reactions", s.handleAddReaction)
|
||
r.Delete("/reactions/{emoji}", s.handleRemoveReaction)
|
||
})
|
||
|
||
// Role Board (cross-collection role-based view)
|
||
r.Get("/roles/board", s.handleRoleBoard)
|
||
r.Put("/roles/board/reorder", s.handleRoleBoardReorder)
|
||
r.Put("/roles/board/lane-order", s.handleRoleBoardLaneReorder)
|
||
|
||
// Agent Roles
|
||
r.Route("/agent-roles", func(r chi.Router) {
|
||
r.Get("/", s.handleListAgentRoles)
|
||
r.Post("/", s.handleCreateAgentRole)
|
||
r.Route("/{roleID}", func(r chi.Router) {
|
||
r.Get("/", s.handleGetAgentRole)
|
||
r.Patch("/", s.handleUpdateAgentRole)
|
||
r.Delete("/", s.handleDeleteAgentRole)
|
||
})
|
||
})
|
||
|
||
// Attachments
|
||
// POST /attachments — upload (TASK-871)
|
||
// GET /attachments/{attachmentID} — serve blob (TASK-872, supports ?variant=)
|
||
// HEAD /attachments/{attachmentID} — metadata only (TASK-877 file-chip enrichment)
|
||
// POST /attachments/{attachmentID}/transform — server-side rotate/crop (TASK-879/880)
|
||
//
|
||
// chi does not auto-route HEAD to the GET handler, so the
|
||
// editor's HEAD probe for size + MIME has to be registered
|
||
// explicitly. The handler short-circuits the streaming
|
||
// path on HEAD; http.ServeContent already strips the body
|
||
// on the seekable path.
|
||
r.Post("/attachments", s.handleUploadAttachment)
|
||
r.Get("/attachments", s.handleListWorkspaceAttachments)
|
||
r.Get("/attachments/{attachmentID}", s.handleGetAttachment)
|
||
r.Head("/attachments/{attachmentID}", s.handleGetAttachment)
|
||
r.Post("/attachments/{attachmentID}/transform", s.handleTransformAttachment)
|
||
r.Delete("/attachments/{attachmentID}", s.handleDeleteWorkspaceAttachment)
|
||
|
||
// Storage usage summary for Settings → Storage and other
|
||
// quota-aware UI surfaces (TASK-881). Cached behind a
|
||
// short TTL — see handleGetWorkspaceStorageUsage.
|
||
r.Get("/storage/usage", s.handleGetWorkspaceStorageUsage)
|
||
|
||
// Webhooks
|
||
r.Route("/webhooks", func(r chi.Router) {
|
||
r.Get("/", s.handleListWebhooks)
|
||
r.Post("/", s.handleCreateWebhook)
|
||
r.Route("/{webhookID}", func(r chi.Router) {
|
||
r.Delete("/", s.handleDeleteWebhook)
|
||
r.Post("/test", s.handleTestWebhook)
|
||
})
|
||
})
|
||
|
||
// API Tokens
|
||
r.Route("/tokens", func(r chi.Router) {
|
||
r.Get("/", s.handleListTokens)
|
||
r.Post("/", s.handleCreateToken)
|
||
r.Delete("/{tokenID}", s.handleDeleteToken)
|
||
})
|
||
|
||
// Members
|
||
r.Route("/members", func(r chi.Router) {
|
||
r.Get("/", s.handleListMembers)
|
||
r.Post("/invite", s.handleInviteMember)
|
||
r.Delete("/invitations/{invID}", s.handleCancelInvitation)
|
||
r.Delete("/{userID}", s.handleRemoveMember)
|
||
r.Patch("/{userID}", s.handleUpdateMemberRole)
|
||
r.Get("/{userID}/collection-access", s.handleGetMemberCollectionAccess)
|
||
r.Put("/{userID}/collection-access", s.handleSetMemberCollectionAccess)
|
||
})
|
||
|
||
// Me — current user's effective workspace context (role,
|
||
// collection access, grants). Open to any principal admitted
|
||
// by RequireWorkspaceAccess (members + guests).
|
||
r.Get("/me", s.handleGetMe)
|
||
|
||
// Dashboard (v2)
|
||
r.Get("/dashboard", s.handleGetDashboard)
|
||
|
||
// Workspace graph — {nodes, edges} for the 3D
|
||
// graph view (PLAN-1730 / TASK-1731). Active
|
||
// items by default; ?include_terminal=true for
|
||
// the full history.
|
||
r.Get("/graph", s.handleGetWorkspaceGraph)
|
||
|
||
// Project report — windowed throughput/flow/status
|
||
// stats (PLAN-1628 / TASK-1630).
|
||
r.Get("/report", s.handleGetReport)
|
||
// Per-user Insights layout prefs (PLAN-1628 / TASK-1634).
|
||
r.Get("/report/layout", s.handleGetReportLayout)
|
||
r.Put("/report/layout", s.handleSaveReportLayout)
|
||
|
||
// Project intelligence reads — next/standup/changelog
|
||
// (PLAN-1888 / TASK-1894). Mirror `pad project
|
||
// next|standup|changelog` (cmd/pad/main.go) — KEEP IN
|
||
// SYNC, see handlers_project_intel.go's doc comments.
|
||
// The MCP HTTP transport's dispatchProjectNext/Standup/
|
||
// Changelog (internal/mcp/dispatch_http_project.go)
|
||
// proxy directly to these three handlers (TASK-1916),
|
||
// so they need no separate sync-keeping.
|
||
r.Get("/next", s.handleGetProjectNext)
|
||
r.Get("/standup", s.handleGetProjectStandup)
|
||
r.Get("/changelog", s.handleGetProjectChangelog)
|
||
|
||
// Agent bootstrap (PLAN-1377 / TASK-1379) — single
|
||
// round-trip that returns workspace + user +
|
||
// collections + always-on conventions + roles +
|
||
// playbook metadata + dashboard + recent activity.
|
||
// Replaces the four /pad context-loading calls the
|
||
// skill used to make. Same shape via the MCP
|
||
// surfaces in TASK-1380.
|
||
r.Get("/agent/bootstrap", s.handleGetBootstrap)
|
||
|
||
// Playbook surface (PLAN-1377 / TASK-1382) — list /
|
||
// show / run for first-class invokable procedures.
|
||
// run is side-effect-free: it parses args per the
|
||
// playbook's declared spec and returns the body +
|
||
// bound args. The agent (skill or MCP-driven)
|
||
// executes the body; the server does not.
|
||
r.Get("/playbooks", s.handleListPlaybooks)
|
||
r.Get("/playbooks/{ref}", s.handleShowPlaybook)
|
||
r.Post("/playbooks/{ref}/run", s.handleRunPlaybook)
|
||
|
||
// Incremental sync — returns items changed since a timestamp
|
||
r.Get("/changes", s.handleGetChanges)
|
||
})
|
||
})
|
||
|
||
// Search
|
||
r.Get("/search", s.handleSearch)
|
||
|
||
// My watches (TASK-2533), cross-workspace — mirrors
|
||
// /auth/tokens' shape for a user-scoped-not-workspace-scoped
|
||
// resource. `pad watch list`. Create/delete are per-item and
|
||
// live under /workspaces/{ws}/items/{itemSlug}/watch instead
|
||
// (they need the item's workspace context to resolve the
|
||
// ref/slug the CLI's positional arg names).
|
||
r.Get("/watches", s.handleListWatches)
|
||
|
||
// MCP tool-surface descriptor (PLAN-1888 / TASK-1891). Serves
|
||
// the catalog JSON (the nine env.Catalog tools + per-action
|
||
// read_only flags) for the browser-side WebMCP layer to build
|
||
// tool descriptors. Inside the authed group so it inherits
|
||
// TokenAuth/SessionAuth/CSRFProtect/RequireAuth — same-origin
|
||
// session/token only, NOT the bearer-gated /mcp infra path.
|
||
// The handler nil-checks toolSurfaceJSON: 404 when the
|
||
// serializer hasn't been injected (mirrors the SetMCPTransport
|
||
// gating). Exposes only catalog descriptors — no route table,
|
||
// handler internals, or other server state.
|
||
r.Get("/mcp/tool-surface", s.handleMCPToolSurface)
|
||
})
|
||
|
||
// Cross-workspace wiki-link resolver (IDEA-1492). Resolves
|
||
// `[[workspace::REF]]` links emitted by the markdown renderer to
|
||
// the canonical item URL via a 302 redirect. Lives outside /api/v1
|
||
// because rendered HTML hrefs target user-facing paths, not API
|
||
// endpoints. Registered at the outer group level so chi matches
|
||
// these URLs ahead of the catch-all SPA handler. ACL check matches
|
||
// existing workspace-access semantics — 404 (not 403) on no-access
|
||
// so we don't leak whether a workspace exists.
|
||
//
|
||
// URL shape: `/-/r/{workspace}/{ref}` — the leading `-/r/` prefix
|
||
// is structurally impossible to collide with any user-namespace
|
||
// URL because username slugs require a leading letter (slugify
|
||
// rule), so no existing or future page route under
|
||
// /{username}/... can shadow this resolver, and no collection
|
||
// slug under /{u}/{ws}/{coll}/... can intercept it
|
||
// (slug grammar also requires letter-led). This replaces the
|
||
// earlier `/{username}/{workspace}/ref/{ref}` shape that risked
|
||
// collision with collection slugs named "ref" on pre-existing
|
||
// data (Codex round-2 P1.4 — picked Option B over a migration
|
||
// because the feature is unshipped, the new shape is more
|
||
// defensive, and the only cost is a frontend emit-shape change).
|
||
r.Get("/-/r/{workspace}/{ref}", s.handleResolveCrossWorkspaceRef)
|
||
}) // end r.Group (full middleware stack)
|
||
|
||
s.router = r
|
||
}
|
||
|
||
// SetWebUI sets the embedded web UI filesystem for serving the SPA.
|
||
func (s *Server) SetWebUI(fsys fs.FS) {
|
||
s.webFS = fsys
|
||
s.ensureRouter()
|
||
s.router.Handle("/*", s.spaHandler())
|
||
}
|
||
|
||
func (s *Server) spaHandler() http.Handler {
|
||
fileServer := http.FileServer(http.FS(s.webFS))
|
||
indexHTML, err := fs.ReadFile(s.webFS, "index.html")
|
||
if err != nil {
|
||
// Embedded web UI is missing — fail fast instead of silently
|
||
// serving blank HTML to every request. This indicates a broken
|
||
// build, so the server should refuse to start.
|
||
panic(fmt.Sprintf("spaHandler: failed to read embedded index.html: %v", err))
|
||
}
|
||
|
||
return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
|
||
path := r.URL.Path
|
||
if strings.HasPrefix(path, "/api/") {
|
||
http.NotFound(w, r)
|
||
return
|
||
}
|
||
|
||
cleanPath := strings.TrimPrefix(path, "/")
|
||
if cleanPath != "" {
|
||
if _, err := fs.Stat(s.webFS, cleanPath); err == nil {
|
||
if strings.Contains(path, "/immutable/") {
|
||
w.Header().Set("Cache-Control", "public, max-age=31536000, immutable")
|
||
} else {
|
||
w.Header().Set("Cache-Control", "no-cache")
|
||
}
|
||
fileServer.ServeHTTP(w, r)
|
||
return
|
||
}
|
||
}
|
||
|
||
// Generate per-request nonce for inline script CSP
|
||
nonce := generateCSPNonce()
|
||
|
||
// Inject nonce into inline <script> tags (SvelteKit bootstrap)
|
||
html := bytes.Replace(indexHTML, []byte("<script>"), []byte(fmt.Sprintf(`<script nonce="%s">`, nonce)), -1)
|
||
|
||
// Set nonce-based CSP (overrides the strict default from SecurityHeaders).
|
||
// - 'nonce-<N>' authorizes the SvelteKit bootstrap <script> we inject below.
|
||
// - 'strict-dynamic' lets that trusted script dynamically import() the
|
||
// SvelteKit runtime chunks without listing every build-hashed path. In
|
||
// browsers that honor CSP L3, 'strict-dynamic' supersedes the 'self'
|
||
// host-list, so an XSS gap that injects <script src="//evil.com"> is
|
||
// rejected even though 'self' is present. 'self' stays as a fallback
|
||
// for older browsers that don't implement strict-dynamic.
|
||
// - script-src-attr 'none' blocks inline event handlers regardless of the
|
||
// script-src nonce — per CSP spec, event attributes bypass script-src.
|
||
w.Header().Set("Content-Security-Policy", fmt.Sprintf(
|
||
"default-src 'self'; script-src 'self' 'nonce-%s' 'strict-dynamic'; script-src-attr 'none'; style-src 'self' 'unsafe-inline'; img-src 'self' data:; font-src 'self'; connect-src 'self'; frame-ancestors 'none'",
|
||
nonce))
|
||
|
||
w.Header().Set("Content-Type", "text/html; charset=utf-8")
|
||
w.Header().Set("Cache-Control", "no-cache, no-store, must-revalidate")
|
||
w.WriteHeader(http.StatusOK)
|
||
w.Write(html)
|
||
})
|
||
}
|
||
|
||
// ensureRouter lazily initializes the router on first use, so all Set*
|
||
// configuration is applied before the middleware chain is built.
|
||
func (s *Server) ensureRouter() {
|
||
s.routerOnce.Do(func() {
|
||
s.setupRouter()
|
||
})
|
||
}
|
||
|
||
func (s *Server) ServeHTTP(w http.ResponseWriter, r *http.Request) {
|
||
s.ensureRouter()
|
||
s.router.ServeHTTP(w, r)
|
||
}
|
||
|
||
// httpIdleTimeout caps the keep-alive idle window on every HTTP
|
||
// connection. Tracked here (not in handlers_events.go) because it
|
||
// applies to ALL connections, not just SSE — but it has a hard
|
||
// invariant relationship with sseKeepaliveInterval: an idle SSE
|
||
// stream is kept alive by periodic comment writes, and those must
|
||
// land more frequently than IdleTimeout or the connection will be
|
||
// closed mid-stream by the http.Server. The guard in
|
||
// handlers_events.go's init() enforces 3 × sseKeepaliveInterval <
|
||
// httpIdleTimeout so we tolerate one or two missed/dropped writes
|
||
// (network blip, scheduler hiccup) before tripping the deadline.
|
||
const httpIdleTimeout = 120 * time.Second
|
||
|
||
func (s *Server) ListenAndServe(addr string) error {
|
||
s.ensureRouter()
|
||
|
||
s.httpServer = &http.Server{
|
||
Addr: addr,
|
||
Handler: s.router,
|
||
ReadTimeout: 15 * time.Second,
|
||
ReadHeaderTimeout: 5 * time.Second,
|
||
IdleTimeout: httpIdleTimeout,
|
||
// Cap total header bytes (default 1 MB) to 64 KB — well above any
|
||
// legitimate request (cookies, auth, content-type, a few CSRF/CORS
|
||
// headers) and tight enough to cheaply reject header-flood DoS.
|
||
MaxHeaderBytes: 64 * 1024,
|
||
// WriteTimeout left at 0 — SSE connections are long-lived.
|
||
// Non-SSE handlers should use per-request context deadlines.
|
||
}
|
||
|
||
slog.Info("Pad server listening", "addr", addr)
|
||
return s.httpServer.ListenAndServe()
|
||
}
|
||
|
||
// Shutdown gracefully drains in-flight requests and stops the HTTP server.
|
||
// The provided context controls how long to wait for active connections.
|
||
func (s *Server) Shutdown(ctx context.Context) error {
|
||
if s.httpServer == nil {
|
||
return nil
|
||
}
|
||
return s.httpServer.Shutdown(ctx)
|
||
}
|
||
|
||
// Handler returns the configured HTTP handler (router).
|
||
// Useful for testing with httptest.NewServer.
|
||
func (s *Server) Handler() http.Handler {
|
||
s.ensureRouter()
|
||
return s.router
|
||
}
|
||
|
||
// --- helpers ---
|
||
|
||
func jsonContentType(next http.Handler) http.Handler {
|
||
return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
|
||
if len(r.URL.Path) >= 7 && r.URL.Path[:7] == "/api/v1" {
|
||
w.Header().Set("Content-Type", "application/json")
|
||
}
|
||
next.ServeHTTP(w, r)
|
||
})
|
||
}
|
||
|
||
func writeJSON(w http.ResponseWriter, status int, v interface{}) {
|
||
w.WriteHeader(status)
|
||
if err := json.NewEncoder(w).Encode(v); err != nil {
|
||
slog.Error("failed to encode JSON response", "error", err)
|
||
}
|
||
}
|
||
|
||
func writeError(w http.ResponseWriter, status int, code, message string) {
|
||
writeJSON(w, status, map[string]interface{}{
|
||
"error": map[string]string{
|
||
"code": code,
|
||
"message": message,
|
||
},
|
||
})
|
||
}
|
||
|
||
// writeInternalError logs the real error server-side and sends a generic
|
||
// message to the client. This prevents leaking SQL errors, file paths,
|
||
// and other internal details.
|
||
func writeInternalError(w http.ResponseWriter, err error) {
|
||
slog.Error("internal server error", "error", err)
|
||
writeError(w, http.StatusInternalServerError, "internal_error", "An internal error occurred")
|
||
}
|
||
|
||
// defaultJSONBodyLimit is the default cap applied to JSON request bodies
|
||
// by decodeJSON. Every /api/* POST/PATCH is comfortably small in practice
|
||
// (items, collections, auth payloads — all well under 100 KB), so the
|
||
// 2 MB cap is several orders of magnitude above real traffic while still
|
||
// cheap to hold in memory per request. Callers who legitimately need
|
||
// more — bulk imports — should call decodeJSONWithLimit explicitly.
|
||
const defaultJSONBodyLimit = 2 << 20 // 2 MiB
|
||
|
||
// decodeJSON reads and unmarshals the JSON body into v. Wraps the body in
|
||
// http.MaxBytesReader so an attacker can't exhaust memory by POSTing a
|
||
// multi-GB JSON blob — without this, json.NewDecoder.Decode happily
|
||
// streams the whole body into a single allocation.
|
||
func decodeJSON(r *http.Request, v interface{}) error {
|
||
return decodeJSONWithLimit(r, v, defaultJSONBodyLimit)
|
||
}
|
||
|
||
// decodeJSONWithLimit is the size-configurable variant. Use this for
|
||
// endpoints that accept large payloads (e.g. bulk-import) where the
|
||
// default cap is too small — but always pass an explicit cap, never
|
||
// remove the wrapper.
|
||
func decodeJSONWithLimit(r *http.Request, v interface{}, maxBytes int64) error {
|
||
// http.MaxBytesReader.Close() is a no-op; the decoder leaves r.Body at
|
||
// EOF anyway. Setting this here also lets the server return a 413
|
||
// automatically via the error we wrap below.
|
||
if r.Body != nil {
|
||
r.Body = http.MaxBytesReader(nil, r.Body, maxBytes)
|
||
}
|
||
if err := json.NewDecoder(r.Body).Decode(v); err != nil {
|
||
return fmt.Errorf("invalid JSON: %w", err)
|
||
}
|
||
return nil
|
||
}
|
||
|
||
// getWorkspaceID resolves workspace slug/ID from the request.
|
||
// If RequireWorkspaceAccess already resolved the workspace, reads from context.
|
||
// Otherwise falls back to direct resolution (for unauthenticated paths).
|
||
func (s *Server) getWorkspaceID(w http.ResponseWriter, r *http.Request) (string, bool) {
|
||
// Fast path: already resolved by RequireWorkspaceAccess middleware
|
||
if wsID, ok := r.Context().Value(ctxResolvedWorkspaceID).(string); ok && wsID != "" {
|
||
return wsID, true
|
||
}
|
||
|
||
// Slow path: resolve directly (should rarely happen — only for routes
|
||
// that don't go through RequireWorkspaceAccess)
|
||
slugOrID := chi.URLParam(r, "slug")
|
||
ws, err := s.resolveWorkspace(slugOrID, currentUser(r))
|
||
if err != nil {
|
||
writeInternalError(w, err)
|
||
return "", false
|
||
}
|
||
if ws == nil {
|
||
writeError(w, http.StatusNotFound, "not_found", "Workspace not found")
|
||
return "", false
|
||
}
|
||
return ws.ID, true
|
||
}
|
||
|
||
// getWorkspace returns the full workspace object resolved by middleware.
|
||
// Falls back to direct resolution for routes without RequireWorkspaceAccess.
|
||
func (s *Server) getWorkspace(w http.ResponseWriter, r *http.Request) (*models.Workspace, bool) {
|
||
// Fast path: use middleware-resolved ID
|
||
if wsID, ok := r.Context().Value(ctxResolvedWorkspaceID).(string); ok && wsID != "" {
|
||
ws, err := s.store.GetWorkspaceByID(wsID)
|
||
if err != nil {
|
||
writeInternalError(w, err)
|
||
return nil, false
|
||
}
|
||
if ws != nil {
|
||
return ws, true
|
||
}
|
||
}
|
||
|
||
// Slow path: resolve from URL param
|
||
slugOrID := chi.URLParam(r, "slug")
|
||
ws, err := s.resolveWorkspace(slugOrID, currentUser(r))
|
||
if err != nil {
|
||
writeInternalError(w, err)
|
||
return nil, false
|
||
}
|
||
if ws == nil {
|
||
writeError(w, http.StatusNotFound, "not_found", "Workspace not found")
|
||
return nil, false
|
||
}
|
||
return ws, true
|
||
}
|
||
|
||
// visibleCollectionIDs returns the set of collection IDs the current user can
|
||
// see in the given workspace. Returns nil if the user has "all" access (no
|
||
// filtering needed), or a non-nil slice for "specific" access. Unauthenticated
|
||
// users (fresh install) always get nil (all access), as do platform admins —
|
||
// but ONLY over a cookie session. Bearer-authed admins (CLI / PAT / MCP —
|
||
// detected via isBearerAuth) fall through to the store lookup and are scoped
|
||
// to their actual membership, matching RequireWorkspaceAccess's admin-bypass
|
||
// suppression for bearer auth (BUG-1616/1617) and reportVisibleCollections'
|
||
// gate (handlers_reports.go). Fixes BUG-1917 — this was the last of the
|
||
// bearer-gate consumers (buildDashboardResponse, handleListItems, the graph
|
||
// handler, handleCreateItem's collection-visibility check, and every other
|
||
// direct caller) still granting a bearer admin an unrestricted view.
|
||
func (s *Server) visibleCollectionIDs(r *http.Request, workspaceID string) ([]string, error) {
|
||
user := currentUser(r)
|
||
if user == nil || (user.Role == "admin" && !isBearerAuth(r)) {
|
||
return nil, nil // No filtering for admins (cookie session) or unauthenticated
|
||
}
|
||
return s.store.VisibleCollectionIDs(workspaceID, user.ID)
|
||
}
|
||
|
||
// requireCollectionFullyVisible checks that the collection is visible to the
|
||
// requesting user under FULL-collection-access semantics (BUG-1920 —
|
||
// codex R2 follow-up). This is deliberately STRICTER than
|
||
// handleGetCollection's inline visibleCollectionIDs + isCollectionVisible
|
||
// check: VisibleCollectionIDs (workspace_members.go) intentionally folds in
|
||
// collections that are visible ONLY via an item-level grant, "so the
|
||
// collection appears in navigation" — item-level filtering is left to the
|
||
// handlers. That nav-lenient shape is correct for handleGetCollection
|
||
// (viewing collection metadata), but WRONG here: this helper's only callers
|
||
// are the four handlers that mint or list a collection-wide share link or
|
||
// grant, where passing would hand out (or reveal) access to the ENTIRE
|
||
// collection — an item grant on a single item inside it must NOT qualify.
|
||
//
|
||
// Mirrors reportVisibleCollections' fullCollIDs narrowing
|
||
// (handlers_reports.go): when the caller holds any item-level grants, the
|
||
// acceptable set narrows from the nav-lenient VisibleCollectionIDs set to
|
||
// the full-access-only set (guestResourceFilter's fullCollIDs — collection
|
||
// grants + member_collection_access + system collections, excluding
|
||
// item-grant-only collections).
|
||
//
|
||
// Writes a 404 and returns false if not visible; callers should invoke this
|
||
// immediately after resolving a collection by slug/ID.
|
||
func (s *Server) requireCollectionFullyVisible(w http.ResponseWriter, r *http.Request, workspaceID string, coll *models.Collection) bool {
|
||
visible, err := s.checkCollectionFullyVisible(r, workspaceID, coll.ID)
|
||
if err != nil {
|
||
writeInternalError(w, err)
|
||
return false
|
||
}
|
||
if !visible {
|
||
writeError(w, http.StatusNotFound, "not_found", "Collection not found")
|
||
return false
|
||
}
|
||
return true
|
||
}
|
||
|
||
// requireItemVisible checks that the item's collection is visible to the
|
||
// requesting user. For guests with item-level grants, also verifies that the
|
||
// specific item is granted (not just the collection). Writes a 404 and returns
|
||
// false if not. Callers should invoke this immediately after resolving an item
|
||
// by slug/ID.
|
||
//
|
||
// Thin shim over checkItemVisible — see that helper for the rules. This
|
||
// wrapper exists for the legacy call-sites that already hold a *http.Request
|
||
// pre-populated by RequireWorkspaceAccess; new callers without that middleware
|
||
// (e.g. handlers_ref_resolver.go) should use checkItemVisible directly with
|
||
// a manually-derived role.
|
||
func (s *Server) requireItemVisible(w http.ResponseWriter, r *http.Request, workspaceID string, item *models.Item) bool {
|
||
visible, err := s.checkItemVisible(workspaceID, item, currentUser(r), workspaceRole(r), isBearerAuth(r))
|
||
if err != nil {
|
||
writeInternalError(w, err)
|
||
return false
|
||
}
|
||
if !visible {
|
||
writeError(w, http.StatusNotFound, "not_found", "Item not found")
|
||
return false
|
||
}
|
||
return true
|
||
}
|
||
|
||
// checkItemVisible is the context-free visibility decision. Returns (true,
|
||
// nil) when the (user, role) pair can see `item` under the same rules
|
||
// `requireItemVisible` enforces. Centralizes the rule set so the
|
||
// resolver route (IDEA-1492) and the middleware-gated handlers can't drift.
|
||
//
|
||
// Inputs:
|
||
//
|
||
// - workspaceID — the resolved workspace's UUID.
|
||
// - item — already-loaded item (so the helper doesn't re-resolve and
|
||
// accidentally apply a different lookup path).
|
||
// - user — currentUser(r) at the call site; nil for unauthenticated.
|
||
// - role — workspaceRole(r) at the call site, or the role derived
|
||
// manually by callers operating outside RequireWorkspaceAccess.
|
||
// - isBearer — isBearerAuth(r) at the call site (BUG-1918). Narrows
|
||
// rule 3 below the same way BUG-1616/1617 narrowed the analogous
|
||
// bypasses in visibleCollectionIDs, resolverWorkspaceRole, and
|
||
// guestResourceFilterCore: a platform admin's global read access is
|
||
// a cookie-session / web-UI affordance only. A bearer-borne admin
|
||
// (CLI / PAT / MCP) who is a restricted workspace member must fall
|
||
// through to the same per-collection filter every other member
|
||
// faces — otherwise BUG-1917's list-level scoping is bypassable by
|
||
// guessing a ref and hitting the single-item endpoints directly.
|
||
//
|
||
// Rules (in order):
|
||
//
|
||
// 1. Tokenized-nil-user bypass: when currentUser == nil AND role is one
|
||
// of the synthesized-by-middleware roles ("owner" for fresh-install,
|
||
// "editor" for legacy workspace-scoped API tokens),
|
||
// RequireWorkspaceAccess has already authorized the request — there
|
||
// is no per-user filter to apply, and the user-nil rejection at
|
||
// rule 2 would false-404 these callers. The bypass is SCOPED TO
|
||
// user == nil — real authenticated members with role "owner" /
|
||
// "editor" must still fall through to the per-collection filter,
|
||
// otherwise a restricted editor member would bypass their own
|
||
// collection_access="specific" gate (Codex round-3 regression of
|
||
// the round-2 P1.1 fix).
|
||
// 2. nil user past rule 1 → not visible. Anonymous viewers without a
|
||
// tokenized role have no item-read access (share links own the
|
||
// public-read surface via /s/{token}).
|
||
// 3. Admin user via cookie session (user.Role == "admin" && !isBearer)
|
||
// → always visible. Bearer-borne admins fall through to rule 4.
|
||
// 4. Otherwise: replay the guestResourceFilterCore + member-collection-
|
||
// access logic that requireItemVisible used to inline, with a system-
|
||
// collections union added to the item-grants branch (Codex round-2
|
||
// P1.2 — restricted members with conventions/playbooks access plus
|
||
// an unrelated item grant previously 404'd on system-collection items
|
||
// because the item-grants branch only checked direct grants + the
|
||
// member's explicit collection-access list).
|
||
func (s *Server) checkItemVisible(workspaceID string, item *models.Item, user *models.User, role string, isBearer bool) (bool, error) {
|
||
return s.checkItemVisibleQ(s.store.Q(), workspaceID, item, user, role, isBearer)
|
||
}
|
||
|
||
// checkItemVisibleQ is checkItemVisible parameterized over its executor, so
|
||
// the cross-workspace copy's attachment authorizer can run the same
|
||
// visibility rule on the copy transaction's own connection instead of the
|
||
// pool while the transaction holds both workspace advisory locks
|
||
// (BUG-2409). Every other caller goes through the pool wrapper above; the
|
||
// decision logic is identical by construction — one body, two executors.
|
||
func (s *Server) checkItemVisibleQ(q store.Queryer, workspaceID string, item *models.Item, user *models.User, role string, isBearer bool) (bool, error) {
|
||
// Tokenized-nil-user bypass. RequireWorkspaceAccess synthesizes
|
||
// "owner" on fresh installs (UserCount == 0, currentUser == nil) and
|
||
// "editor" for legacy workspace-scoped API tokens (currentUser ==
|
||
// nil but tokenWorkspaceID matches). Both are authorized by the
|
||
// middleware already. Real authenticated users with these roles
|
||
// (workspace owners, member.Role=="editor", …) must NOT short-circuit
|
||
// here — they have to walk the per-collection filter so
|
||
// collection_access="specific" + member_collection_access actually
|
||
// gates them (Codex round-3 — the round-2 fix dropped the
|
||
// `user == nil` qualifier and accidentally disabled the gate for
|
||
// every real editor too).
|
||
if user == nil && (role == "owner" || role == "editor") {
|
||
return true, nil
|
||
}
|
||
if user == nil {
|
||
return false, nil
|
||
}
|
||
// Admin sees everything, but only for cookie-session auth (matches
|
||
// visibleCollectionIDs's nil-filter shape). Bearer-borne admins
|
||
// (BUG-1918) fall through to the same per-collection filter every
|
||
// other member faces.
|
||
if user.Role == "admin" && !isBearer {
|
||
return true, nil
|
||
}
|
||
|
||
// Visibility filter: nil = unrestricted; non-nil = restricted to the slice.
|
||
visibleIDs, err := s.store.VisibleCollectionIDsQ(q, workspaceID, user.ID)
|
||
if err != nil {
|
||
return false, err
|
||
}
|
||
if !isCollectionVisible(item.CollectionID, visibleIDs) {
|
||
return false, nil
|
||
}
|
||
|
||
// Replay guestResourceFilterCore's logic without the *http.Request
|
||
// dependency. Member-with-all-access short-circuits to "no item-level
|
||
// filter"; guests + restricted members get the grant filter.
|
||
if role != "guest" {
|
||
member, err := s.store.GetWorkspaceMemberQ(q, workspaceID, user.ID)
|
||
if err != nil {
|
||
return false, err
|
||
}
|
||
if member != nil && (member.CollectionAccess == "all" || member.CollectionAccess == "") {
|
||
// Full collection access — visibleIDs filter already passed.
|
||
return true, nil
|
||
}
|
||
}
|
||
|
||
grantCollIDs, grantedItemIDs, err := s.store.GuestVisibleResourcesQ(q, workspaceID, user.ID)
|
||
if err != nil {
|
||
return false, err
|
||
}
|
||
if len(grantedItemIDs) == 0 {
|
||
// No item-level grants in play. visibleCollectionIDs already
|
||
// determined the collection is reachable; visibility stands.
|
||
return true, nil
|
||
}
|
||
|
||
// Item-level grants are active. The item is visible when:
|
||
// a) the collection itself has a full grant (any item passes), OR
|
||
// b) for restricted members: the collection is in member_collection_access
|
||
// (the member's explicit collection-access list), OR
|
||
// c) the item's collection is a system collection — restricted
|
||
// members always retain access to system collections (conventions,
|
||
// playbooks, …); pre-round-2 this branch missed the system-
|
||
// collections union that guestResourceFilterCore performed, so a
|
||
// restricted member with an item grant in a non-system collection
|
||
// was 404'd on a system-collection item they were entitled to see.
|
||
// d) the specific item is in the granted-items list.
|
||
for _, id := range grantCollIDs {
|
||
if id == item.CollectionID {
|
||
return true, nil
|
||
}
|
||
}
|
||
if role != "guest" {
|
||
// member_collection_access path — restricted members see their
|
||
// explicit collection-access list as full grants alongside any
|
||
// item-level grants.
|
||
memberColls, err := s.store.GetMemberCollectionAccessQ(q, workspaceID, user.ID)
|
||
if err != nil {
|
||
return false, err
|
||
}
|
||
for _, id := range memberColls {
|
||
if id == item.CollectionID {
|
||
return true, nil
|
||
}
|
||
}
|
||
// System-collections union — mirror guestResourceFilterCore's
|
||
// pre-round-2 behavior. ListSystemCollectionIDs is a workspace-
|
||
// scoped lookup (no per-user filter), so the same call is correct
|
||
// for every restricted member in the workspace.
|
||
sysColls, err := s.store.ListSystemCollectionIDsQ(q, workspaceID)
|
||
if err != nil {
|
||
return false, err
|
||
}
|
||
for _, id := range sysColls {
|
||
if id == item.CollectionID {
|
||
return true, nil
|
||
}
|
||
}
|
||
}
|
||
for _, id := range grantedItemIDs {
|
||
if id == item.ID {
|
||
return true, nil
|
||
}
|
||
}
|
||
return false, nil
|
||
}
|
||
|
||
// isItemVisibleToGuest checks if an item is visible given grant-based access,
|
||
// considering both full-collection grants and individual item grants.
|
||
// When fullCollIDs and grantedItemIDs are both nil, always returns true (no grant filtering).
|
||
func (s *Server) isItemVisibleToGuest(r *http.Request, workspaceID string, item *models.Item, fullCollIDs, grantedItemIDs []string) bool {
|
||
if fullCollIDs == nil && grantedItemIDs == nil {
|
||
return true
|
||
}
|
||
// Full collection grant covers all items in the collection
|
||
for _, id := range fullCollIDs {
|
||
if id == item.CollectionID {
|
||
return true
|
||
}
|
||
}
|
||
// Otherwise, the specific item must be in the granted items list
|
||
for _, id := range grantedItemIDs {
|
||
if id == item.ID {
|
||
return true
|
||
}
|
||
}
|
||
return false
|
||
}
|
||
|
||
// guestResourceFilter returns the full-collection IDs and granted item IDs for
|
||
// the current user if they need item-level grant filtering. Returns nil/nil for:
|
||
// - unauthenticated users
|
||
// - admin users
|
||
// - members with "all" collection access (grants should merge, not replace)
|
||
// For guests: returns direct collection grants as fullCollIDs + item grants.
|
||
// For restricted members: returns member_collection_access + system collections
|
||
// + direct collection grants as fullCollIDs, plus item grants as grantedItemIDs.
|
||
// This ensures item grants are additive to the member's existing access.
|
||
func (s *Server) guestResourceFilter(r *http.Request, workspaceID string) (fullCollIDs, grantedItemIDs []string, err error) {
|
||
return s.guestResourceFilterCore(r, workspaceID, false)
|
||
}
|
||
|
||
// guestResourceFilterIncludeDeletedItems is the delta-sync variant
|
||
// of guestResourceFilter. It uses GuestVisibleResourcesIncludeDeleted
|
||
// under the hood so soft-deleted granted items still surface in the
|
||
// resulting ID set. Used by /items-changes (TASK-1354) so a guest /
|
||
// restricted member with an item-level grant still receives the
|
||
// `deleted:true` row when their granted item is soft-deleted —
|
||
// without this variant the grant ID vanishes before the delta
|
||
// query runs and the client keeps the stale entry forever (Codex
|
||
// review of TASK-1354 round 1 [P1]).
|
||
func (s *Server) guestResourceFilterIncludeDeletedItems(r *http.Request, workspaceID string) (fullCollIDs, grantedItemIDs []string, err error) {
|
||
return s.guestResourceFilterCore(r, workspaceID, true)
|
||
}
|
||
|
||
// guestResourceFilterCore is the request-scoped wrapper around
|
||
// Store.ResolveBacklinksVisibility. Delegates the role-determination
|
||
// + merge logic to the store helper so cross-workspace backlinks
|
||
// callers can reuse the same code path without a request context
|
||
// (PLAN-1593 / TASK-1597).
|
||
//
|
||
// The wrapper still exists for two reasons:
|
||
// - Admin bypass uses currentUser(r).Role rather than re-fetching
|
||
// the user (saves one DB roundtrip per request on the hot path).
|
||
// - The signature `(r *http.Request, workspaceID, includeDeleted) →
|
||
// (fullCollIDs, grantedItemIDs, err)` is established across many
|
||
// handlers — keeping it stable avoids a sprawling refactor.
|
||
//
|
||
// Admin bypass policy (BUG-1617 — companion to BUG-1616): the
|
||
// short-circuit only fires for cookie session auth. Bearer-borne
|
||
// admins (CLI / PAT / MCP — detected via isBearerAuth) fall through
|
||
// to the store helper which runs the regular member/grants pipeline
|
||
// against the platform-admin's actual workspace_members row. Without
|
||
// this, an admin's MCP token could pass RequireWorkspaceAccess's
|
||
// bearer gate (BUG-1616) on the URL's workspace but still see
|
||
// unrestricted visibility filters in any downstream backlinks /
|
||
// activity / delta-sync query that ran for the same workspace.
|
||
//
|
||
// The includeDeletedItems flag swaps the underlying grant query.
|
||
func (s *Server) guestResourceFilterCore(r *http.Request, workspaceID string, includeDeletedItems bool) (fullCollIDs, grantedItemIDs []string, err error) {
|
||
return s.guestResourceFilterCoreQ(s.store.Q(), r, workspaceID, includeDeletedItems)
|
||
}
|
||
|
||
// guestResourceFilterCoreQ is guestResourceFilterCore parameterized over its
|
||
// executor — the cross-workspace copy's attachment authorizer runs it on the
|
||
// copy transaction's connection (BUG-2409); everything else uses the pool
|
||
// wrapper above.
|
||
func (s *Server) guestResourceFilterCoreQ(q store.Queryer, r *http.Request, workspaceID string, includeDeletedItems bool) (fullCollIDs, grantedItemIDs []string, err error) {
|
||
user := currentUser(r)
|
||
if user == nil {
|
||
return nil, nil, nil
|
||
}
|
||
authIsBearer := isBearerAuth(r)
|
||
if user.Role == "admin" && !authIsBearer {
|
||
return nil, nil, nil
|
||
}
|
||
// Delegate to the request-independent helper. The store-side
|
||
// helper duplicates the admin check via GetUser, but for the
|
||
// request hot path we short-circuit above (cookie admin only)
|
||
// so the duplicate lookup never fires for the common case.
|
||
return s.store.ResolveBacklinksVisibilityQ(q, user.ID, workspaceID, includeDeletedItems, authIsBearer)
|
||
}
|
||
|
||
// isCollectionVisible checks if a collection ID is in the visible set.
|
||
// If visibleIDs is nil, all collections are visible.
|
||
func isCollectionVisible(collectionID string, visibleIDs []string) bool {
|
||
if visibleIDs == nil {
|
||
return true
|
||
}
|
||
for _, id := range visibleIDs {
|
||
if id == collectionID {
|
||
return true
|
||
}
|
||
}
|
||
return false
|
||
}
|
||
|
||
// filterUserGrantsForCaller narrows collGrants/itemGrants — the TARGET
|
||
// user's grants, already loaded by the caller — down to what the CALLER can
|
||
// see. Only meaningful when caller != target; handleListUserGrants skips
|
||
// calling this for self-queries (a user can always see their own grants).
|
||
//
|
||
// BUG-1928: handleListUserGrants returned the target's raw grants
|
||
// (including collection_id/item_id) to any workspace owner unconditionally.
|
||
// A restricted owner (collection_access="specific") could enumerate
|
||
// hidden-resource IDs this way — the disclosure half of the primitive
|
||
// BUG-1923's handlers fixed the action half of (know-the-ID → operate-on-it).
|
||
//
|
||
// Reuses the existing guestResourceFilter/isCollectionVisible/
|
||
// isItemVisibleToGuest helpers rather than a bespoke visibility pass:
|
||
// guestResourceFilter's fullCollIDs is already the STRICT full-access set
|
||
// (member_collection_access ∪ system collections ∪ direct collection
|
||
// grants, excluding item-grant-only collections) — the same strict set
|
||
// requireCollectionFullyVisible narrows to — so collection grants are
|
||
// filtered directly against it with no extra narrowing step.
|
||
//
|
||
// Filtering is pure ID-set membership: no parent collection/item lookup is
|
||
// needed to decide visibility, so a grant on a soft- (or even hard-)
|
||
// deleted parent is filtered the same as any other grant, matching #798's
|
||
// "still revocable/inspectable" precedent for grants on archived resources.
|
||
// The one exception is item grants, which only carry an item_id — those are
|
||
// resolved to their collection_id via a single bulk GetItemCollectionRefs
|
||
// call (state-agnostic; no deleted_at filter) rather than N per-grant
|
||
// lookups.
|
||
func (s *Server) filterUserGrantsForCaller(r *http.Request, workspaceID string, collGrants []models.CollectionGrant, itemGrants []models.ItemGrant) ([]models.CollectionGrant, []models.ItemGrant, error) {
|
||
fullCollIDs, grantedItemIDs, err := s.guestResourceFilter(r, workspaceID)
|
||
if err != nil {
|
||
return nil, nil, err
|
||
}
|
||
if fullCollIDs == nil && grantedItemIDs == nil {
|
||
// Unrestricted caller (admin/cookie session, or a member with
|
||
// full collection access) — no filtering, and no further store
|
||
// calls needed.
|
||
return collGrants, itemGrants, nil
|
||
}
|
||
|
||
filteredColl := make([]models.CollectionGrant, 0, len(collGrants))
|
||
for _, g := range collGrants {
|
||
if isCollectionVisible(g.CollectionID, fullCollIDs) {
|
||
filteredColl = append(filteredColl, g)
|
||
}
|
||
}
|
||
|
||
filteredItem := make([]models.ItemGrant, 0, len(itemGrants))
|
||
if len(itemGrants) > 0 {
|
||
itemIDs := make([]string, len(itemGrants))
|
||
for i, g := range itemGrants {
|
||
itemIDs[i] = g.ItemID
|
||
}
|
||
refs, err := s.store.GetItemCollectionRefs(workspaceID, itemIDs)
|
||
if err != nil {
|
||
return nil, nil, err
|
||
}
|
||
collByItem := make(map[string]string, len(refs))
|
||
for _, ref := range refs {
|
||
collByItem[ref.ID] = ref.CollectionID
|
||
}
|
||
for _, g := range itemGrants {
|
||
collID, ok := collByItem[g.ItemID]
|
||
if !ok {
|
||
// item_grants.item_id is ON DELETE CASCADE, so a grant
|
||
// row can't outlive its item — this should be
|
||
// unreachable. Exclude defensively rather than show a
|
||
// grant with no resolvable parent.
|
||
continue
|
||
}
|
||
item := &models.Item{ID: g.ItemID, CollectionID: collID}
|
||
if s.isItemVisibleToGuest(r, workspaceID, item, fullCollIDs, grantedItemIDs) {
|
||
filteredItem = append(filteredItem, g)
|
||
}
|
||
}
|
||
}
|
||
|
||
return filteredColl, filteredItem, nil
|
||
}
|
||
|
||
// requireEditPermission checks if the user has edit access to the given item.
|
||
// For regular members (editor/owner), this uses the standard role check.
|
||
// For members with insufficient roles (e.g., viewers), it falls back to
|
||
// grant-based permissions so grants can override the base role.
|
||
// For guests, it resolves the effective permission from grants directly.
|
||
// Returns true if the request should continue, false if it was rejected with a 403.
|
||
//
|
||
// NEVER call this with a workspace ID other than the one the current
|
||
// request's URL resolved to. The `workspaceID` parameter makes it look
|
||
// reusable for a second workspace; it is not. The editor/owner fast path
|
||
// below reads workspaceRole(r), which RequireWorkspaceAccess populates only
|
||
// for the URL's workspace, so passing workspace B's ID applies workspace A's
|
||
// role — privilege escalation. Use AuthorizeCrossWorkspaceEdit
|
||
// (authz_cross_workspace.go) for any other workspace; it also checks the
|
||
// OAuth/MCP consent allow-list, which this helper does not (DR-10 of
|
||
// PLAN-2357).
|
||
func (s *Server) requireEditPermission(w http.ResponseWriter, r *http.Request, workspaceID string, itemID, collectionID string) bool {
|
||
role := workspaceRole(r)
|
||
|
||
// Editors and owners always have edit access
|
||
if role != "guest" && requireRole(r, "editor") {
|
||
return true
|
||
}
|
||
|
||
// For guests and members with insufficient role (e.g., viewers),
|
||
// check grant-based permissions as an override.
|
||
user := currentUser(r)
|
||
if user == nil {
|
||
writeError(w, http.StatusForbidden, "forbidden", "Insufficient permissions")
|
||
return false
|
||
}
|
||
|
||
perm, err := s.store.ResolveUserPermission(workspaceID, user.ID, itemID, collectionID)
|
||
if err != nil {
|
||
writeInternalError(w, err)
|
||
return false
|
||
}
|
||
if permissionLevel(perm) < permissionLevel("edit") {
|
||
writeError(w, http.StatusForbidden, "forbidden", "Insufficient permissions")
|
||
return false
|
||
}
|
||
return true
|
||
}
|
||
|
||
// resolveWorkspace resolves a workspace by slug or UUID, scoped to the
|
||
// authenticated user's accessible workspaces when a user context is present.
|
||
// Returns nil (not an error) if no workspace is found.
|
||
func (s *Server) resolveWorkspace(slugOrID string, user *models.User) (*models.Workspace, error) {
|
||
// 1. Is it a UUID? Try resolving by ID first, then fall back to slug.
|
||
// A workspace slug could be UUID-shaped (e.g. imported data), so we
|
||
// can't short-circuit here.
|
||
if isUUID(slugOrID) {
|
||
ws, err := s.store.GetWorkspaceByID(slugOrID)
|
||
if ws != nil || err != nil {
|
||
return ws, err
|
||
}
|
||
// Not found by ID — fall through to slug-based resolution
|
||
}
|
||
|
||
// 2. No authenticated user — fall back to global slug lookup
|
||
// (fresh install, or pre-auth paths)
|
||
if user == nil {
|
||
return s.store.GetWorkspaceBySlug(slugOrID)
|
||
}
|
||
|
||
// 3. Admin users — global slug lookup (admins can see all workspaces)
|
||
if user.Role == "admin" {
|
||
return s.store.GetWorkspaceBySlug(slugOrID)
|
||
}
|
||
|
||
// 4. Auth-scoped slug resolution: find workspaces where user is owner or member
|
||
workspaces, err := s.store.GetWorkspacesBySlugForUser(slugOrID, user.ID)
|
||
if err != nil {
|
||
return nil, err
|
||
}
|
||
|
||
if len(workspaces) == 1 {
|
||
return &workspaces[0], nil
|
||
}
|
||
if len(workspaces) == 0 {
|
||
return nil, nil
|
||
}
|
||
|
||
// Ambiguous: multiple workspaces match — this should be rare.
|
||
// For now, return the first one. The 409 disambiguation is only needed
|
||
// when we actually have per-owner slug uniqueness (after the unique
|
||
// constraint is changed). Currently slugs are globally unique.
|
||
return &workspaces[0], nil
|
||
}
|
||
|
||
// isUUID is defined in handlers_items.go
|
||
|
||
// getWorkspaceDocument resolves workspace slug and document ID from URL params.
|
||
func (s *Server) getWorkspaceDocument(w http.ResponseWriter, r *http.Request) (string, *models.Document, bool) {
|
||
workspaceID, ok := s.getWorkspaceID(w, r)
|
||
if !ok {
|
||
return "", nil, false
|
||
}
|
||
|
||
docID := chi.URLParam(r, "docID")
|
||
doc, err := s.store.GetDocument(docID)
|
||
if err != nil {
|
||
writeInternalError(w, err)
|
||
return "", nil, false
|
||
}
|
||
if doc == nil || doc.WorkspaceID != workspaceID {
|
||
writeError(w, http.StatusNotFound, "not_found", "Document not found")
|
||
return "", nil, false
|
||
}
|
||
return workspaceID, doc, true
|
||
}
|