* fix(frontend): tighten bell notification panel toolbar
Restructure the popover so the segmented filter (All / Unread / Alerts)
and the action icons (filter toggle, mark all read, clear all) share a
single row inside the 360px panel. The previous layout wrapped the
segmented control onto a second row once the type filter was added.
Move the unread badge to the right of the title, collapse the node and
type dropdowns behind a filter-toggle icon (accent dot signals an active
filter), equalize both Select widths, and extract the duplicated
className strings into local constants.
* docs: refresh bell notification popover screenshot
Reflects the reworked toolbar (segmented filter + filter toggle +
mark-all-read + clear-all on a single row).
Flip Security (Trivy), Notifications (agents + history), and App Store from
global-and-hidden-on-remote to node-scoped so operators can manage them when a
remote node is selected in the node picker. The primary instance proxies the
calls to each remote, which resolves the correct per-instance binary state,
agent config, and template registry.
Backend: key `agents` and `notification_history` by `node_id` with idempotent
column-add migrations and a `(node_id, type)` unique index on agents, matching
the Labels pattern. Thread `req.nodeId` through the /api/agents and
/api/notifications routes. Internal NotificationService and ImageUpdateService
writes resolve the middleware default via `NodeRegistry.getDefaultNodeId()` so
monitor-emitted rows share a bucket with user-facing ones (avoids split-brain
where the UI sees test notifications but not internal alerts).
Frontend: split Security on remote to render only the scanner card and hide
scan policies and CVE suppressions (those remain control-plane-only). Drop the
misleading "Always Local" badge on Developer since retention windows govern
backend jobs, not UI state. Flip the App Store registry to node-scoped.
Docs: add a "What Settings apply per node" table to multi-node, clarify
remote alert setup in alerts-notifications, and note Trivy's per-host install
in vulnerability-scanning.
Enrich scheduled vulnerability scan completion notifications with
per-severity CVE counts so recipients can triage from the message body
alone. Expose the scan action in the schedule creation UI, require an
explicit node_id, and harden fire-and-forget alert dispatches so a
failing webhook cannot crash the scheduler.
Notification body now reports scanned/skipped/failed counts plus
critical/high/medium totals aggregated across fresh and cached scans,
reflecting the current node posture rather than only what was newly
scanned on this run.
Scheduled scan tasks now dispatch a completion alert through the
existing notification system: info when every image scanned cleanly,
warning when one or more images failed. The alert includes the task
name and the run's scanned/cached/failed summary so operators do not
need to open the task history.
* fix(notifications): stop Sencho version notifications from silently skipping
Three independent defects combined to make version-update notifications
silently fail even while Fleet overview correctly surfaced an update button:
- The in-memory 6-hour cooldown was advanced before the network fetch,
so a single transient failure at boot could lock the check for the
rest of the container lifetime. Moved the cooldown update inside the
success branch so failures retry on the next eval cycle.
- MonitorService called the raw version fetch directly, bypassing the
CacheService wrapper (TTL, inflight dedup, stale-on-error) that Fleet
uses, so the two paths could diverge. Unified both on a shared
getLatestVersion() helper in utils/version-check.ts.
- The dedup key could carry stale state from a previous build and never
self-clear. It now self-heals when the running version reaches the
previously-notified version, so future releases re-fire as expected.
Added diagnostic logs gated on debug mode for each skip branch, plus
three regression tests covering cooldown-on-failure, cooldown-on-success,
and dedup self-heal.
* docs(notifications): drop legacy-upgrade framing from alerts troubleshooting
Sencho has not shipped publicly, so troubleshooting entries written in
'this used to happen but now does Y' mode reference a past that does not
exist for any reader. Rewrote the version-notification, image-update,
and crash-alert troubleshooting entries to describe current behavior
positively without referring to prior builds, upgrade paths, or legacy
fixes.
* chore(security): accept CVE-2026-33810 in bundled Docker CLI 29.4.0
Trivy now flags CVE-2026-33810 (Go stdlib crypto/x509 DNS constraint
bypass, fixed in Go 1.26.2) in the Docker CLI static binary we ship.
Docker CLI 29.4.0 is the latest upstream release and still links Go
1.26.1; no newer static binary exists yet.
Same exposure profile as the already-accepted CVE-2026-32280: the
Docker CLI and compose plugin only validate certificates from
well-known registry CAs and the local Docker socket, not from
attacker-controlled CAs with crafted DNS name constraints. Revisit
on the next Docker CLI release that rebuilds against Go 1.26.2 or
later.
* fix(docker-events): harden crash detection against edge cases
- Isolate per-container failures in the reconcile path so one failed
container inspect cannot abort classification for the rest of the
batch after a reconnect.
- Fall back to inspecting State.OOMKilled when a container exits with
code 137 and no oom event preceded the die, so cgroup OOM kills that
lose the oom event are still classified correctly. The dedup check
runs before the fallback so crashloops do not hammer the daemon.
- Guard the intentional-kill window against wildly future-dated Docker
timestamps by switching to a signed age comparison with a bounded
negative-skew tolerance, so clock skew cannot flip a genuine crash
into an intentional stop.
- Align diagnostic logging with the codebase [Service:diag] convention
and add one informational lifecycle log on boot and shutdown so
operators running in developer mode can confirm the watcher started.
- New tests cover gap-inspect isolation, the OOM inspect fallback (both
success and failure paths), duplicate die events collapsing within
the grace window, and clock-skew bounds on the classifier.
* docs(alerts): add troubleshooting entries for crash detection toggle and rate limits
* fix(notifications): replace polling with Docker event stream for container lifecycle detection
Replaces the 30-second MonitorService crash-detection poll with a causal,
per-node Docker events stream. Eliminates false crash alerts on intentional
stops (docker stop, compose down, stack restart/update), detects OOM kills
as a distinct alert category, and surfaces real crashes in real time.
A new DockerEventManager spawns one DockerEventService per local node. Each
service consumes the filtered container event stream, classifies die events
against recent kill/oom state, and reconciles container state via snapshot
diffing on connect and reconnect. Rate limiting, exponential backoff with
jitter, and parse-error tolerance keep the stream resilient under load and
during daemon interruptions.
MonitorService retains host limits, janitor, version check, and stack metric
alerts; crash and healthcheck detection move out entirely.
* fix(tests): silence require-imports lint in hoisted mock factory
- Read running Sencho version from the packaged manifest via getSenchoVersion() instead of process.env.npm_package_version, which is undefined when launched via node dist/index.js (production Docker). Skip the update check entirely when the version cannot be resolved.
- Add a one-time backfill pass in ImageUpdateService so users who upgraded to a Sencho version with the notification pipeline receive a catch-up entry for stacks already flagged as having updates before the upgrade.
- Surface dispatch failures as error-level entries in the in-app notification bell via a direct notification_history write, so misconfigured webhooks are visible without tailing logs.
- Extend test coverage for both paths: mock getSenchoVersion (including the null-version case) and add dispatch-path tests for transition, no-re-fire, backfill, and error surfacing.
- Expand alerts-notifications docs with per-instance 6-hour check cadence, an example version message, a backfill note, and three new troubleshooting entries.
* fix(alerts): harden with security fixes, design compliance, and test coverage
Add authMiddleware to all alert endpoints, validate notification test
dispatch inputs, fix restart_count metric via Docker inspect, correct
network metric units, replace any types with DockerContainerStats
interface, add webhook timeouts and dispatch error tracking.
Frontend: migrate Select to Combobox, add ScrollArea and delete
confirmation AlertDialog, fix icon strokeWidth to 1.5.
Add update availability notifications for both Sencho version updates
(6-hour check in MonitorService) and stack image updates (state
transition detection in ImageUpdateService). Extract shared version
fetch logic into utils/version-check.ts.
Add diagnostic logging gated behind developer_mode for MonitorService
breach state machine and NotificationService dispatch routing.
Tests: 24 new alert API integration tests, restart_count and version
check unit tests (688 total passing). Docs updated with HTTPS
requirement, update notifications section, and troubleshooting guide.
* fix(alerts): remove unused TEST_USERNAME import in alerts-api tests
- Add 5 new Tier 1 doc pages: configuration, stack-management, editor,
multi-node, and alerts-notifications
- Update introduction, quickstart, and features/overview to reflect
current feature set and link to new pages
- Restructure mint.json with Getting Started / Features / Reference /
Operations navigation groups
- Add Playwright-captured screenshots for all major UI screens