Remove the leftover paid-only blockingEnabled switch so enabled
block-on-deploy policies enforce on every tier, matching the
documented every-tier security surface. Existing Community policies
begin blocking immediately with no migration.
The auto-update, bulk-label, scheduler, and blueprint deploy block
messages hardcoded "image(s) exceed <max_severity>", which is wrong
under the risk-first policy model: a block can be driven by a
known-exploited (KEV) or fixable Critical/High input while the severity
threshold was never the trigger. In those cases the message named a
severity ceiling the policy did not enforce.
Route all four message paths through a shared summarizeBlockReasons
helper (the same reason text the deploy-gate 409 response and the block
dialog already use), so every surface names the inputs that actually
matched. Falls back to a generic phrase when no reason was recorded.
Scan-policy deploy gates can now block on a known-exploited CVE (CISA KEV)
and on a fixable Critical/High finding, in addition to an optional severity
threshold. New policies default risk-first (KEV and fixable on, severity off);
existing policies keep their severity-only behavior. CVSS stays captured for
context but is never the sole basis for a block, and a finding whose
exploitability cannot be confirmed is treated as risky rather than safe
(incomplete scan detail fails closed on KEV/fixable inputs).
The decision logic is shared between the pre-deploy gate and the informational
post-scan banner via a pure helper, so the two never disagree. Block messages
and the block dialog now name the conditions an image matched. Backend and
frontend gates move together, the new inputs replicate across the fleet, and a
blocking policy with no active input is rejected on both sides.
* fix(stack-activity): per-stack history integrity, attribution, sanitization
Address the Stack Activity audit findings (PR 1 of 2):
- Per-stack history integrity: drop the per-insert 100-row prune in
addNotificationHistory that evicted quieter stacks' history whenever
another stack got chatty. Periodic cleanupOldNotifications now caps
per (node, stack) at 500 rows and per-node unattached system events
at 1000 rows, on top of the existing 30-day retention. Signature
takes an options bag and returns a per-stage summary so MonitorService
can log what actually ran each cycle.
- Actor attribution: thread req.user?.username through every
notifyActionFailure call site and add synthetic actors at service
emit sites (system:autoheal, system:scheduler, system:image-update,
system:docker-events, system:blueprint, system:monitor, system:policy).
The timeline renders system actors as "via <Label>" so an autoheal
redeploy is no longer indistinguishable from a user redeploy.
- Message sanitization: new sanitizeNotificationMessage at
NotificationService.dispatchAlert strips KEY=VALUE pairs whose key
ends in TOKEN/KEY/PASSWORD/SECRET/CREDENTIALS/AUTH, scrubs HTTP basic
auth in URLs and Bearer tokens, collapses COMPOSE_DIR paths, and
truncates to 1000 chars. Applied to the stored history and to every
downstream Discord/Slack/webhook channel. The ImageUpdateService
recovery-path direct DB write also runs through the sanitizer.
- Composite pagination cursor: getStackActivity now accepts a
(timestamp, id) cursor (?before=&beforeId=). The legacy timestamp-only
form silently dropped events when a single compose up emitted many
events sharing one millisecond. Route rejects beforeId without before.
- Frontend hardening: distinct error state with retry button (initial
fetch failure no longer renders as the genuine empty state), strict
positive-integer parsing on cursor params, overrequest-by-1 pagination
so the last page does not leave a dead "Load more" click, runtime
guard on liveEvents merge that validates the level union, per-minute
day-bucket recompute so an open panel does not stay on "Today" past
midnight.
No tier, role, or capability gate touched. Route permission gate
remains stack:read on the named stack.
* fix(stack-activity): sanitizer covers lowercase env vars and per-node compose dir
External review surfaced two leak paths in the message sanitizer:
- The sensitive-key regex was uppercase-only. Compose env names are
conventionally uppercase but lowercase forms (db_password, jwt_secret,
github_token) are valid and do leak through the same Docker and
compose-parse error paths. Make the regex case-insensitive and tighten
it to also catch bare TOKEN= / KEY= / PASSWORD= without a prefix word,
while still leaving BYPASS, COMPASS, and similar non-secret keys alone.
- The compose-dir path collapse only read process.env.COMPOSE_DIR, but
the real resolution chain is node.compose_dir (per-node DB override)
-> process.env.COMPOSE_DIR -> /app/compose. A node with a custom
compose_dir could still leak absolute paths into stored history and
downstream channels. Route both the dispatchAlert call and the
ImageUpdateService recovery-path direct write through
NodeRegistry.getInstance().getComposeDir(localNodeId) so the
collapse covers every resolution outcome.
Tests now assert lowercase keys are redacted and that BYPASS-style
non-secrets stay intact in both cases. notification-routing mock
extended to stub the new getComposeDir call.
* chore(stack-activity): a11y roles, visibility-aware tick, live-disconnect signal
Close three small follow-ups on the per-stack activity timeline:
- A11y: each day-group gets role="list" and each event row gets
role="listitem" so screen readers traverse the timeline as a list
instead of a wall of text. The day-group container also carries an
aria-label naming the bucket.
- Visibility-aware day-bucket tick: the 60s setInterval that re-derives
Today/Yesterday/Earlier now short-circuits when document.hidden, so a
backgrounded panel does not re-render every minute for no visible
effect.
- Live-disconnect signal: useNotifications dispatches a
sencho:notifications-connection custom event on WebSocket open and
close. The timeline listens and, when explicitly disconnected, shows
a one-line "Live updates offline; reconnecting…" hint above the list.
The sidebar ticker already surfaces fleet-wide connection state; this
adds an in-context cue for users who are focused on a single stack.
Stack-name case normalization was considered and rejected: stack names
are case-permissive per the isValidStackName validator, and lowercasing
on read or write would silently rename or hide a user's "MyApp" stack.
* ci(stack-activity): drop unnecessary escape in URL_BASIC_AUTH regex
ESLint no-useless-escape errored on \- inside the character class
[a-zA-Z0-9+.\-] at notificationMessage.ts:14. Move the dash to the
end of the class so it's an unambiguous literal and the escape is no
longer required. Behavior is identical; sanitizer tests still pass.
* revert(stack-activity): drop unvalidated E2E spec from this PR
The spec was committed without ever running against a real Docker
daemon, then failed in CI when it ran for the first time: deploy
returned 200 but no notification appeared on the activity endpoint
within the polling window, suggesting either a deploy-notification
race or a node-id resolution mismatch in the CI environment.
Backend unit tests (route + composite cursor + sanitizer) and
frontend component tests cover the same logic. The E2E spec will
land in a dedicated follow-up once it has been authored against a
working CI environment.
triggerPostDeployScan was fire-and-forget. When Trivy was missing on a
node, when the registry refused the digest lookup, or when a single
image scan threw, the failure went to console.error and the user
never learned. Open the security tab later, see stale data, no
indicator that the scan even tried.
Backend:
- New stack_scan_attempts table (node_id, stack_name, status,
attempted_at, error_message). One row per stack; latest attempt
overwrites the previous one.
- DatabaseService gains recordStackScanAttempt /
getStackScanAttempt / clearStackScanAttempts. Status is one of
'ok' | 'partial' | 'failed' | 'skipped'.
- triggerPostDeployScan in helpers/policyGate.ts now records every
exit path: 'skipped' when Trivy is unavailable or no images to
scan; 'failed' when container enumeration or all images fail;
'partial' when some images scan and others fail; 'ok' on full
success.
- New GET /api/stacks/:name/scan-status returns { status,
attemptedAt, errorMessage } or { status: null } when never tried.
- DELETE /:stackName cleanup chain now clears the row alongside
the existing update-status / auto-update cleanups.
Frontend:
- StackAnatomyPanel fetches /scan-status on stackName change.
- Renders a small warning strip below the update banner when
status !== 'ok' (failed / partial / skipped). Hidden when status
is 'ok' or unknown (never attempted). Title attribute carries
the full error message for hover inspection.
Cross-feature note: the audit doc flagged this as M-6 with a
coordination note for the pending Security feature audit. The
schema kept intentionally narrow (one row per stack, simple
status enum) so the Security audit can extend it (richer history,
per-image-row breakdown, etc.) without a destructive migration.
Resolves M-6 from the stack-management audit.
* fix: harden deploy enforcement paths
* fix: update Docker toolchain to Go 1.26.3
* fix: repair Dockerfile tr argument split across lines
* fix: bump protobufjs to clear npm audit high-severity advisories
* fix(test): add execFile to child_process mock in compose-images test
* fix: resolve merge conflicts with main
* fix: resolve merge conflicts with main
* fix: resolve merge conflicts with main
* refactor(backend): sanitize user input before logging to close CRLF injection
Adds a small sanitizeForLog helper that strips CR, LF, tab, and ASCII
control characters (0x00-0x1F, 0x7F) from a value before it is embedded
in a console.log/warn/error/debug call. Wraps every call site where a
user-controlled value (req.params, req.body, req.query, or a value
derived from them) flows into a log message.
Closes the bulk of the open CodeQL alerts in this family:
- 96 js/log-injection
- 28 js/tainted-format-string
The helper is in backend/src/utils/safeLog.ts. Routes still pre-validate
input at the request boundary; this is the second line of defense and
gives static analyzers a sanitizer they can trace through. JSON
responses, Docker filter labels, and other non-log call sites are
intentionally left unwrapped.
* refactor(backend): printf-style format strings for tainted-log call sites
CodeQL's js/tainted-format-string rule flags template literals in the first
arg of console.X when any interpolated value is user-controlled, regardless
of whether each value is sanitized inline. The canonical mitigation is to
use a static format string and pass values as positional args.
Converts the 28 flagged template literals to printf-style ("%s") format
strings, with sanitizeForLog applied to each positional arg. Also fills in
the log-injection wraps on 9 sites where a user-controlled value was
missed in the first sweep (agents, fleet, gitSources, imageUpdates,
GitSourceService).
No behavior change at runtime. Node's util.format substitutes %s tokens
identically to template-literal interpolation.
* fix(backend): wrap nodeId/snapshotId in fleet restore debug log
CodeQL flagged the unwrapped numeric args even though they cannot
contain control chars in practice. Apply the sanitizer for taint-flow
recognition.
Introduce a NotificationCategory string-literal union (11 values) and
thread it through dispatchAlert as a required second argument. All
callers (DockerEventService, AutoHealService, ImageUpdateService,
MonitorService, PolicyEnforcement, policyGate, SchedulerService,
imageUpdates route) pass an explicit category at every call site,
giving TypeScript compile-time enforcement that no new emit site can
be added without choosing a category.
DatabaseService gains an idempotent migration that adds a nullable
category TEXT column to notification_history; existing rows keep
category=NULL (displayed as Uncategorized in the UI). The
getNotificationHistory method accepts an optional category filter
that is forwarded from the GET /api/notifications/history route via
a ?category= query param.
NotificationPanel gains a category Select dropdown so users can
filter history by category. The frontend types mirror the backend
union so API responses are type-safe end-to-end.
All 75 test files (1410 tests) updated to the new 4-arg dispatchAlert
signature and passing.
* refactor(backend): extract types, constants, and guards from index.ts (phase 0)
Additive, behavior-preserving first step of the modular backend refactor.
Moves purely static artifacts out of backend/src/index.ts so later phases can
extract routes and middleware without touching shared symbols.
New modules:
- types/express.ts: Express Request augmentation
- helpers/constants.ts: PORT, password policy, label colors, cookie names,
MFA TTLs, hot-path cache TTLs
- helpers/proxyExemptPaths.ts: PROXY_EXEMPT_PREFIXES + isProxyExemptPath
- helpers/cookies.ts: isSecureRequest, getCookieOptions
- helpers/policyGate.ts: buildPolicyGateOptions, runPolicyGate,
triggerPostDeployScan
- middleware/permissions.ts: ROLE_PERMISSIONS, checkPermission,
requirePermission
- middleware/tierGates.ts: requirePaid, requireAdmiral, requireAdmin,
requireNodeProxy, requireScheduledTaskTier + effectiveTier/Variant
index.ts shrinks by ~260 lines; no runtime behavior changes. All 64 vitest
files and 1,278 tests pass.
* refactor(backend): drop unused imports left after phase 0 extraction
LicenseTier, LicenseVariant, DIGEST_CACHE_TTL_MS, and isProxyExemptPath
were imported into index.ts but no longer referenced there after the
phase 0 move; CI lint flagged them as errors.
isProxyExemptPath will be re-imported in phase 1 when the JSON parser
bypass and nodeContext middleware get extracted. Silence the
no-namespace warning on the Express augmentation since the namespace
syntax is required for TypeScript module augmentation.