* fix(stack-activity): per-stack history integrity, attribution, sanitization
Address the Stack Activity audit findings (PR 1 of 2):
- Per-stack history integrity: drop the per-insert 100-row prune in
addNotificationHistory that evicted quieter stacks' history whenever
another stack got chatty. Periodic cleanupOldNotifications now caps
per (node, stack) at 500 rows and per-node unattached system events
at 1000 rows, on top of the existing 30-day retention. Signature
takes an options bag and returns a per-stage summary so MonitorService
can log what actually ran each cycle.
- Actor attribution: thread req.user?.username through every
notifyActionFailure call site and add synthetic actors at service
emit sites (system:autoheal, system:scheduler, system:image-update,
system:docker-events, system:blueprint, system:monitor, system:policy).
The timeline renders system actors as "via <Label>" so an autoheal
redeploy is no longer indistinguishable from a user redeploy.
- Message sanitization: new sanitizeNotificationMessage at
NotificationService.dispatchAlert strips KEY=VALUE pairs whose key
ends in TOKEN/KEY/PASSWORD/SECRET/CREDENTIALS/AUTH, scrubs HTTP basic
auth in URLs and Bearer tokens, collapses COMPOSE_DIR paths, and
truncates to 1000 chars. Applied to the stored history and to every
downstream Discord/Slack/webhook channel. The ImageUpdateService
recovery-path direct DB write also runs through the sanitizer.
- Composite pagination cursor: getStackActivity now accepts a
(timestamp, id) cursor (?before=&beforeId=). The legacy timestamp-only
form silently dropped events when a single compose up emitted many
events sharing one millisecond. Route rejects beforeId without before.
- Frontend hardening: distinct error state with retry button (initial
fetch failure no longer renders as the genuine empty state), strict
positive-integer parsing on cursor params, overrequest-by-1 pagination
so the last page does not leave a dead "Load more" click, runtime
guard on liveEvents merge that validates the level union, per-minute
day-bucket recompute so an open panel does not stay on "Today" past
midnight.
No tier, role, or capability gate touched. Route permission gate
remains stack:read on the named stack.
* fix(stack-activity): sanitizer covers lowercase env vars and per-node compose dir
External review surfaced two leak paths in the message sanitizer:
- The sensitive-key regex was uppercase-only. Compose env names are
conventionally uppercase but lowercase forms (db_password, jwt_secret,
github_token) are valid and do leak through the same Docker and
compose-parse error paths. Make the regex case-insensitive and tighten
it to also catch bare TOKEN= / KEY= / PASSWORD= without a prefix word,
while still leaving BYPASS, COMPASS, and similar non-secret keys alone.
- The compose-dir path collapse only read process.env.COMPOSE_DIR, but
the real resolution chain is node.compose_dir (per-node DB override)
-> process.env.COMPOSE_DIR -> /app/compose. A node with a custom
compose_dir could still leak absolute paths into stored history and
downstream channels. Route both the dispatchAlert call and the
ImageUpdateService recovery-path direct write through
NodeRegistry.getInstance().getComposeDir(localNodeId) so the
collapse covers every resolution outcome.
Tests now assert lowercase keys are redacted and that BYPASS-style
non-secrets stay intact in both cases. notification-routing mock
extended to stub the new getComposeDir call.
* chore(stack-activity): a11y roles, visibility-aware tick, live-disconnect signal
Close three small follow-ups on the per-stack activity timeline:
- A11y: each day-group gets role="list" and each event row gets
role="listitem" so screen readers traverse the timeline as a list
instead of a wall of text. The day-group container also carries an
aria-label naming the bucket.
- Visibility-aware day-bucket tick: the 60s setInterval that re-derives
Today/Yesterday/Earlier now short-circuits when document.hidden, so a
backgrounded panel does not re-render every minute for no visible
effect.
- Live-disconnect signal: useNotifications dispatches a
sencho:notifications-connection custom event on WebSocket open and
close. The timeline listens and, when explicitly disconnected, shows
a one-line "Live updates offline; reconnecting…" hint above the list.
The sidebar ticker already surfaces fleet-wide connection state; this
adds an in-context cue for users who are focused on a single stack.
Stack-name case normalization was considered and rejected: stack names
are case-permissive per the isValidStackName validator, and lowercasing
on read or write would silently rename or hide a user's "MyApp" stack.
* ci(stack-activity): drop unnecessary escape in URL_BASIC_AUTH regex
ESLint no-useless-escape errored on \- inside the character class
[a-zA-Z0-9+.\-] at notificationMessage.ts:14. Move the dash to the
end of the class so it's an unambiguous literal and the escape is no
longer required. Behavior is identical; sanitizer tests still pass.
* revert(stack-activity): drop unvalidated E2E spec from this PR
The spec was committed without ever running against a real Docker
daemon, then failed in CI when it ran for the first time: deploy
returned 200 but no notification appeared on the activity endpoint
within the polling window, suggesting either a deploy-notification
race or a node-id resolution mismatch in the CI environment.
Backend unit tests (route + composite cursor + sanitizer) and
frontend component tests cover the same logic. The E2E spec will
land in a dedicated follow-up once it has been authored against a
working CI environment.
Flip Security (Trivy), Notifications (agents + history), and App Store from
global-and-hidden-on-remote to node-scoped so operators can manage them when a
remote node is selected in the node picker. The primary instance proxies the
calls to each remote, which resolves the correct per-instance binary state,
agent config, and template registry.
Backend: key `agents` and `notification_history` by `node_id` with idempotent
column-add migrations and a `(node_id, type)` unique index on agents, matching
the Labels pattern. Thread `req.nodeId` through the /api/agents and
/api/notifications routes. Internal NotificationService and ImageUpdateService
writes resolve the middleware default via `NodeRegistry.getDefaultNodeId()` so
monitor-emitted rows share a bucket with user-facing ones (avoids split-brain
where the UI sees test notifications but not internal alerts).
Frontend: split Security on remote to render only the scanner card and hide
scan policies and CVE suppressions (those remain control-plane-only). Drop the
misleading "Always Local" badge on Developer since retention windows govern
backend jobs, not UI state. Flip the App Store registry to node-scoped.
Docs: add a "What Settings apply per node" table to multi-node, clarify
remote alert setup in alerts-notifications, and note Trivy's per-host install
in vulnerability-scanning.
* docs: remove security-sensitive implementation details from public documentation
Generalize or remove internal architecture details that could aid targeted
attacks — CVE tables, database schema, rate limit thresholds, proxy internals,
encryption algorithm names, and WebSocket middleware bypass info.
* test(metrics): fix flaky minute-bucket aggregation test
Floor baseTime to the start of the current minute so baseTime + 5000
never crosses a minute boundary and produces 2 buckets instead of 1.