* feat(attachments): orphan GC sweep with periodic scheduler (TASK-886)
Background job that reclaims attachments past the grace period. Two
qualification criteria, both with a 30-day default grace:
- item_id IS NULL AND deleted_at IS NULL AND created_at < cutoff
(never-attached uploads — editor uploaded then tab-closed before
attaching to an item)
- deleted_at IS NOT NULL AND deleted_at < cutoff
(soft-deleted via the Settings → Storage delete button or the
DELETE /attachments/{id} endpoint)
Reclamation is dedupe-aware: content-addressed storage means the same
hash can be referenced by multiple rows, so the on-disk blob is only
removed when the GC'd row is the LAST live reference to its
content_hash. Otherwise the row drops and the blob stays for the
remaining references. CountLiveAttachmentsForHash is the predicate.
Per-row failures (resolve backend, blob delete, hard-delete) are
logged and skipped; the sweep keeps making progress. Catastrophic
errors (DB failure) return up to the loop, which logs and waits for
the next tick rather than crashing the server.
Lifecycle:
- SetOrphanGCConfig overrides the default 24h interval / 30-day
grace. cmd/pad reads PAD_ORPHAN_GC_INTERVAL / PAD_ORPHAN_GC_GRACE
(Go duration syntax — 1m, 24h, 720h) so operators can tune
without recompiling and tests can crank the interval down to 1ms
to see sweeps land in CI.
- StartOrphanGC kicks the loop. Idempotent — second call is a
no-op so a misconfigured caller can't double-spawn.
- Server.Stop() now signals the loop via stopOrphanGC() before
s.bg.Wait(), so process shutdown drains the goroutine cleanly
(BUG-842 invariant).
- Each tick wraps the sweep in a 30m context timeout so a slow
scan can't pin the goroutine across multiple intervals.
Tests:
- TestOrphanGC_ReclaimsSoftDeleted: upload → soft-delete → sweep
with future cutoff → DB row gone + blob gone from FSStore.
- TestOrphanGC_ReclaimsLongOrphans: upload → backdate created_at
31d → sweep with 30d grace cutoff → row reclaimed.
- TestOrphanGC_KeepsRecentRows: upload → soft-delete → sweep with
past cutoff → row stays. Catches a typo in the WHERE clause that
would silently destroy live attachments.
- TestOrphanGC_PreservesSharedBlob: two uploads with identical
bytes (same hash, same blob), soft-delete only one → sweep →
one row reclaimed BUT BlobsReclaimed=0 because the other row
still references the blob. Pin for content-addressed dedupe.
- TestOrphanGC_StartStop: loop spins up at 1ms interval, second
StartOrphanGC is a no-op, Stop drains via testServer's cleanup.
Parent: PLAN-866. Closes the phase 1 plan with full export →
import → orphan-cleanup round-trip.
* fix(attachments): protect referenced/in-flight blobs from orphan GC per Codex (round 1)
Two real correctness issues Codex caught on PR #307:
P1. The editor's normal upload flow leaves attachments.item_id NULL.
The canonical association lives in markdown content (the editor
PATCHes "pad-attachment:UUID" into the item) — but the GC's
"never-attached past 30d" predicate only checked item_id. So a
legitimate inline image could be hard-deleted 30 days after upload
even though item content still references it.
Added store.AttachmentReferencedInItems(workspaceID, attachmentID)
that scans items.content + items.fields for "pad-attachment:UUID".
The GC sweep now runs this check before reclaiming any
never-attached row; if any live item references the attachment,
the row is left alone (and re-checked next sweep).
P2. Race between concurrent upload and GC. Upload calls
AttachmentStore.Put (blob lands on disk) → THEN inserts the DB row.
Between those two steps an orphan-GC sweep could count zero live
refs for the hash, delete the blob, and the upload's row insert
would then point at a missing blob.
Added Server.inFlightUploadHashes (sync.Map of *atomic.Int64
counters) with markUploadInFlight / uploadInFlight helpers. Every
Put + CreateAttachment site fences itself via markUploadInFlight:
the upload handler, the transform handler, the thumbnail
derivation pipeline, and the bundle-import rehydrate path. The GC
sweep treats an in-flight hash as "another live ref" so it leaves
the blob alone.
Tests:
- TestOrphanGC_KeepsReferencedNeverAttachedRows: upload (item_id
NULL) → create item with pad-attachment: ref → backdate 31d →
sweep with 30d cutoff → row stays.
- TestOrphanGC_RespectsInFlightUploads: upload → soft-delete →
register an in-flight upload at the same hash → sweep → DB row
goes (it's tombstoned past grace) but blob stays so the
in-flight upload can complete cleanly.
The DB row still gets reclaimed in the in-flight case because the
soft-deleted row is independently past grace; only the blob delete
is fenced. That's correct: the blob remains usable for the
incoming upload and the new upload will register its own
attachments row.
* fix(attachments): mutex-protect in-flight tracker + portable JSONB scan per Codex (round 2)
Two fixes for the round-2 findings on PR #307:
P1. Same-hash race in the in-flight upload tracker. The sync.Map +
*atomic.Int64 design split increment from LoadOrStore-then-add and
release-decrement from delete, so a release could see "0" and start
deleting while another upload concurrently reloaded the same map
entry and incremented to "1" — the second upload's signal then
lived in a doomed map slot, invisible to subsequent uploadInFlight
calls.
Replaced with a plain map[string]int64 + sync.Mutex. Inc, dec,
delete-when-zero all run under one critical section, so any
inspection sees a consistent snapshot. Net cost is one mutex per
mark/release; uncontended this is ~10ns and the upload path is
already doing far more expensive work (Put + DB insert).
Stress test: 20 goroutines × 500 iterations of mark→check→release
on a shared hash. Every check must observe in-flight=true while
the calling goroutine holds the mark. Final state must be empty.
Runs cleanly under -race -count=3.
P2. Postgres JSONB compatibility. items.fields is TEXT on SQLite
but JSONB on PostgreSQL (per pgmigrations/001_initial.sql). LIKE
on JSONB fails with a type error, so the orphan GC's reference
scan would error on Postgres and skip every never-attached row —
breaking orphan reclamation for those rows entirely.
Cast fields::text in the Postgres dialect path:
fieldsExpr := "fields"
if s.dialect.Driver() == DriverPostgres {
fieldsExpr = "fields::text"
}
Same approach used elsewhere in the store for dialect-sensitive
text searches.
* fix(attachments): close GC/upload TOCTOU + protect in-grace peers per Codex (round 3)
P1 round 3: TOCTOU race between uploadInFlight check and store.Delete.
The mutex protected the in-flight counter but not the GC's
check-and-delete sequence. A new upload could call markUploadInFlight
between our check and our blob delete, then run Put after the blob
was gone — its CreateAttachment would insert a live row pointing at
the missing hash.
Fixed by holding inFlightHashesMu across the check + FS Delete:
s.inFlightHashesMu.Lock()
inFlight := s.inFlightHashes[hash] > 0
if !inFlight && others == 0 {
store.Delete(ctx, key)
}
s.inFlightHashesMu.Unlock()
A concurrent markUploadInFlight blocks until either we skip (because
we observed in-flight) or finish deleting. Lock window is ms-class
on FSStore; a per-hash lock can replace this server-wide mutex when
S3 lands in Phase 2.
P2 round 3: CountLiveAttachmentsForHash counted only live rows, so
GC could reclaim the blob from row A (soft-deleted 31d ago) even
when row B was also soft-deleted but only 1 day old — within
grace, so its blob must stay reachable until its own grace lapses.
Replaced with CountProtectingAttachmentsForHash which counts rows
where deleted_at IS NULL OR deleted_at >= graceCutoff. The blob is
preserved until every soft-deleted peer has aged past its own
grace window.
Tests:
- TestOrphanGC_RespectsSoftDeletedInGracePeer: two rows sharing a
hash, soft-delete both, backdate only one past 30d → sweep with
30d cutoff → older row reclaimed but blob stays for the still-in-
grace peer.
- existing TestOrphanGC_RespectsInFlightUploads still passes
(still uses the in-flight signal correctly).
* fix(attachments): dedupe blob-reclaim metric across same-hash peers per Codex (round 4)
Codex round 4 noted that when multiple soft-deleted peers share a
content_hash and all are past grace, the GC sweep would inflate
BlobsReclaimed and BytesReclaimed: AttachmentStore.Delete treats a
missing key as success, so the second peer's idempotent no-op
delete still bumped the counter.
Functional cleanup was correct (the blob really was gone after the
first peer); only the metric / log line was wrong, which makes
operator dashboards report fictitious bytes-reclaimed values.
Track per-sweep reclaimed hashes in a map and skip the Delete call
+ counter increment for repeats. The DB row still gets hard-deleted
on each peer.
Test: TestOrphanGC_DedupesBlobReclaimMetric uploads twice with
identical bytes (single shared blob), soft-deletes both, backdates
deleted_at past grace → sweep deletes 2 rows and reports
BlobsReclaimed=1 / BytesReclaimed=blobLen rather than 2 / 2*blobLen.
Pad
Collaborate with your AI agents.
One binary. Local-first. No accounts required. Pad gives you a CLI, a web UI, and an AI agent skill — all backed by SQLite, all running on your machine. Your project data never leaves your laptop.
Quick Start
brew install PerpetualSoftware/tap/pad
cd your-project
pad init # configure, auth, workspace, AI skill — all in one
pad server open # opens the web UI at localhost:7777
pad init is the smart entry point — it auto-detects what's needed, walks you through each step, and is safe to re-run anytime (it skips finished steps and prints a status summary).
Why Pad?
Tools like Linear, Jira, and Notion are built for teams on the cloud. Pad is built for developers on their machine — and for the AI agents working alongside them.
| Pad | Linear / Jira | Notion | |
|---|---|---|---|
| Setup | pad init |
Create account, invite team, configure | Create account, pick template |
| AI agents | Native /pad skill for 7+ tools |
Third-party integrations | Third-party integrations |
| Data | Local SQLite, you own it | Their cloud | Their cloud |
| Offline | Full functionality | Read-only cache at best | Limited |
| CLI | First-class | Afterthought | None |
| Price | Free, open source | Per-seat pricing | Per-seat pricing |
Features
For Developers
CLI that doesn't get in your way. Create tasks, search items, check status — without leaving the terminal.
pad item create task "Fix OAuth redirect" --priority high
pad item create idea "Real-time collaboration" --category infrastructure
pad item list tasks --status in-progress
pad item search "authentication"
pad project dashboard # Project dashboard
pad project next # What should I work on?
pad server info # How this client is connected to Pad
Web UI that stays out of your way. A clean, dark-themed interface at localhost:7777 with:
- Board, list, and table views — drag-and-drop between status columns
- Keyboard navigation —
j/kto move,Enterto open,Escto go back,Cmd+Kto search - Rich text editor — Tiptap-based with markdown, formatting toolbar, and auto-save
- Wiki-links — type
[[Title]]to link between items - Real-time updates — agent creates a task in the terminal, it appears in the browser instantly (via SSE)
- Dashboard — collection overview, active work, plan tracking, activity feed
For AI Agents
Your agent becomes a project partner. Install the /pad skill once, and your AI coding tool can read, create, and update project items through natural language.
pad agent install # Auto-detects your tools and installs the skill
Works with Claude Code, Cursor, Windsurf, Codex, GitHub Copilot, Amazon Q, and JetBrains Junie.
Then just talk to your project:
> /pad what should I work on next?
> /pad I finished the OAuth fix
> /pad create a task to add rate limiting
> /pad let's brainstorm about the API redesign
Conventions and playbooks teach agents how your project works:
- Conventions — trigger-based rules like "run tests before marking a task done" or "use conventional commits"
- Playbooks — multi-step workflows like "when implementing a feature: read the spec, create a branch, write tests first, then implement"
pad item create convention "Run tests before completing tasks" \
--field trigger=on-task-complete \
--field scope=all \
--field priority=must
Agents load relevant conventions automatically. All agent actions are attributed in the activity feed, so you always know what the AI changed.
Onboard agents to a new codebase:
pad workspace onboard # Analyzes project structure, saves workspace context, and suggests conventions
Collections & Custom Fields
Pad organizes work into collections — typed containers with structured fields.
Built-in collections:
| Collection | Purpose |
|---|---|
| Tasks | Work items with status, priority, assignee, effort, due date |
| Ideas | Feature ideas with impact and category |
| Plans | Project milestones with progress tracking |
| Docs | Documentation, decisions, reference material |
| Conventions | Project rules that guide agent behavior |
| Playbooks | Multi-step workflows for agents to follow |
Create your own with typed fields — select, text, date, number, url, relation, checkbox:
pad collection create "Bug Reports" \
--fields "severity:select:low,medium,high,critical; browser:text; reproducible:checkbox"
Items get reference numbers automatically (TASK-5, BUG-12) and can be moved between collections with field migration.
Installation
Homebrew (macOS and Linux)
brew install PerpetualSoftware/tap/pad
Build from Source
git clone https://github.com/PerpetualSoftware/pad
cd pad
make build
cp pad ~/.local/bin/ # or /usr/local/bin/
Requires Go 1.26+ and Node.js 22+.
The go install github.com/PerpetualSoftware/pad/cmd/pad@latest path is not supported for the full Pad binary, because the web UI must be built and embedded during the source build.
Docker
docker run -p 127.0.0.1:7777:7777 -v pad-data:/data ghcr.io/perpetualsoftware/pad
This publishes Pad to localhost:7777 on the host machine, which is the recommended default for local use.
Single user, more than one device? Publish to all interfaces so you can reach Pad from your phone, tablet, or another machine on the same LAN, Tailscale network, or home VPN:
docker run -p 7777:7777 -v pad-data:/data ghcr.io/perpetualsoftware/pad
For multi-instance deployments, Pad supports Postgres + Redis via docker-compose.yml — see docs/deployment.md for the full setup.
Binary Download
Pre-built binaries for macOS, Linux, and Windows are available on the releases page.
Getting Started
1. Set up Pad
cd ~/projects/myapp
pad init "My App"
pad init is the smart entry point that handles everything in one command:
- Configures this client's connection (local server, remote, or Docker)
- Auto-starts the local server
- Creates the first admin account on a fresh local install (Docker / remote hosts run
pad auth setupon the server instead) - Logs you in if needed
- Creates or links a workspace for the current directory (writes
.pad.toml) - Installs the
/padskill for any AI tools detected in the project
Run from your project root. Safe to re-run anytime — it skips finished steps and prints a status summary if nothing's needed.
Choose a template with --template, or omit it for an interactive picker grouped by category (Software / People / …):
pad workspace init --list-templates # See the full catalog grouped by category
pad init "My App" --template scrum # Scrum-style with sprints
pad init "My App" --template product # Product management focused
pad init "My Hiring" --template hiring # Company-side: requisitions, candidates, interview loops, feedback
pad init "Job Search" --template interviewing # Candidate-side: applications, interviews, companies, contacts
Pad ships templates for software (startup / scrum / product), people workflows (hiring, interviewing), and has reserved categories for research, content, operations, and personal use so the same project-management primitives fit well beyond code projects.
2. Start working
# From the CLI
pad item create task "Set up CI pipeline" --priority high
pad item create idea "Add WebSocket support" --category infrastructure
pad project dashboard
# From the web UI
pad server open # Opens localhost:7777 in your browser
# From your AI agent
# Just use /pad in Claude Code, Cursor, etc.
3. Teach your agents the rules
pad workspace onboard # Auto-analyze project, save workspace context, and suggest conventions
# Or browse the convention library
pad library list --type conventions # Pre-built conventions you can adopt
pad library list --type playbooks # Pre-built multi-step workflows
CLI Reference
pad auth configure Configure how this client connects to Pad
pad auth setup Initialize the first admin account
pad auth login Sign in
pad auth whoami Show current user
pad server start Start the Pad API server
pad server stop Stop the Pad server
pad server info Show client, connection, and local server status
pad server open Open web UI in browser
pad workspace init [name] Initialize workspace in current directory
pad workspace link <workspace> Link current directory to an existing workspace
pad workspace list List all workspaces
pad workspace switch <workspace> Switch active workspace
pad workspace context Show structured workspace context
pad workspace context set --file X Update structured workspace context from JSON
pad workspace onboard Analyze project, save workspace context, and suggest conventions
pad workspace members List workspace members
pad workspace invite <email> Invite a workspace member
pad workspace join <code> Accept an invitation
pad workspace export Export workspace data
pad workspace import <file> Import workspace data
pad project dashboard Project dashboard
pad project next Recommended next task
pad project ready Query actionable next items
pad project stale Query stalled or attention-worthy items
pad project standup [--days N] Daily standup report
pad project changelog [--days N] Release notes from completed items
pad project watch Real-time activity stream
pad project reconcile Reconcile item and PR state
pad item create <coll> "title" Create item (task, idea, plan, doc, ...)
pad item list [collection] List items (filters: --status, --priority, --all)
pad item show <ref> Show item detail
pad item update <ref> Update item fields
pad item delete <ref> Delete item
pad item move <ref> <collection> Move item between collections
pad item edit <ref> Open item in $EDITOR
pad item search "query" Full-text search across all items
pad item comment <ref> "text" Add comment to an item
pad item comments <ref> View item comments
pad item note <ref> "summary" Append an implementation note to an item
pad item decide <ref> "decision" Append a decision log entry to an item
pad item block <src> <target> Create dependency
pad item blocked-by <item> <blk> Mark item as blocked
pad item deps <ref> Show dependencies
pad item unblock <src> <target> Remove dependency
pad item related <ref> Show direct relationships for an item
pad item implemented-by <ref> Show incoming implementers for an item
pad item bulk-update --status X Batch update multiple items
pad collection list List collections with item counts
pad collection create <name> Create a custom collection
pad library list Browse convention and playbook library
pad library activate <title> Activate a convention or playbook
pad agent install [tool] Install /pad skill for AI coding tools
pad agent status Show supported tools and installation status
pad agent update Update installed tool integrations
pad github link [item-ref] Link current branch's PR to item
pad github status [item-ref] Show PR status for linked items
pad github unlink <item-ref> Remove PR link from item
pad webhook list List workspace webhooks
pad webhook create <url> Create webhook
All commands accept --format json for machine-readable output and --workspace to target a specific workspace.
Authentication
Pad runs without authentication by default for frictionless local use. For local installs, pad init creates the first admin account inline. The lower-level commands are useful when you're hosting a Pad server (Docker / remote) and need to set up auth on the server host directly:
pad auth setup # Initialize the first admin account (server host, non-local mode)
pad auth login # Sign in
pad auth whoami # Show current user
pad auth logout # Sign out
Once a user exists, all API requests and web UI access require authentication. Credentials are stored in ~/.pad/credentials.json. Multiple users can be invited to workspaces with role-based access control (owner, editor, viewer).
pad workspace members # List workspace members
pad workspace invite user@example.com
pad workspace join <code>
Architecture
┌──────────────────────────────────────────────┐
│ pad (single binary) │
│ │
│ ┌──────────┐ ┌──────────┐ ┌────────────┐ │
│ │ CLI │ │ REST │ │ Embedded │ │
│ │ (Cobra) │ │ API │ │ Web UI │ │
│ └────┬─────┘ └────┬─────┘ │ (SvelteKit)│ │
│ │ HTTP │ └────────────┘ │
│ └──────────────┤ │
│ ┌─────▼─────┐ │
│ │ SQLite │ │
│ │ + FTS5 │ │
│ └───────────┘ │
└───────────────────────────────────────────────┘
- Go backend — chi router, SQLite via modernc.org/sqlite (pure Go, no CGO), FTS5 full-text search, SSE for real-time updates
- SvelteKit frontend — Svelte 5, Tiptap editor, drag-and-drop, adapter-static, embedded via
go:embed - Single binary — serves the API and web UI, runs on macOS, Linux, and Windows
- Workspace-per-project — each project gets its own workspace linked by a
.pad.tomlfile
All data lives in ~/.pad/pad.db. Your data. Your machine. No telemetry, no cloud, no accounts required.
Contributing
See CONTRIBUTING.md for the development guide.
make build # Build web UI + Go binary
make test # Run Go tests
make dev-web # SvelteKit dev server with hot reload
make install # Build, install to ~/.local/bin, restart server
Security
See SECURITY.md for reporting vulnerabilities.

