mirror of
https://github.com/Studio-Saelix/sencho.git
synced 2026-07-27 20:29:10 +00:00
adcd04b01ac97e7bdd8d6073d585386d4d25e8d3
1378 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
adcd04b01a |
refactor(auto-update): retire per-stack gate, drive auto-update from schedules only (#1233)
* refactor(auto-update): retire per-stack gate, drive auto-update from schedules only
The per-stack Auto-update toggle in the stack sidebar context menu wrote a
gate row to `stack_auto_update_settings`, but actual updates only ran when a
`scheduled_tasks` row with `action='update'` fired. On a fresh install the
toggle was inert: detection ran every 6h, nothing was applied.
The same context menu already exposes `Schedule task`, which opens
ScheduledOperationsView pre-filled for the stack where the user can pick
`Auto-update Stack` and any cron. Keeping the toggle alongside that flow
duplicated the same action and turned the gate table into a parallel store
of "is a covering schedule active" derivable from `scheduled_tasks` itself.
Drop the gate model entirely:
- Backend: remove the `stack_auto_update_settings` table and its four
accessors, the three routes under /api/stacks/*/auto-update, the per-stack
skip in /api/auto-update/execute and SchedulerService.executeUpdate's
fleet branch, and the clearStackAutoUpdateSetting call on stack delete.
Dashboard `autoUpdate` count derives from scheduled_tasks (action='update'
rows pinned to the node, total/enabled split).
- Frontend: drop the Auto-update entry from the sidebar context menu and its
optimistic toggle plumbing. Drop autoUpdateSettings state, the
/stacks/auto-update-settings fetch, and the auto-update-settings-changed
WebSocket branch. Slim useSidebarActivitySummary (just nextRunAt; no
enabled/total counts). AutoUpdateReadinessView's per-card autoUpdateEnabled
now means "a covering enabled action='update' schedule exists" (per-stack
row or fleet row on this node, earliest next_run_at wins, per-stack row
wins on ties), with the gate-fetch removed.
- New: scheduledTasksRouter broadcasts scope: 'scheduled-tasks' on POST,
PUT, PATCH /toggle, and DELETE so useConfigurationStatus and
useNextAutoUpdateRun refetch under the 250ms debounce instead of waiting
for the 60s poll. The broadcast is wrapped so a broken subscriber socket
cannot turn a successful mutation into a 500.
- Docs: rewrite the "Per-stack control" section of auto-update-policies.mdx
to describe the schedule-based model; update the matching troubleshooting
entry. The misleading fleet-update help text in ScheduledOperationsView
is corrected to reflect that every stack on the node is covered.
Tier parity: the surviving auto-update path (Schedule task -> Auto-update
Stack / All Stacks) is gated `requirePaid + requireAdmin` backend and
`isPaid + isAdmin` frontend, matching the gate the deleted routes carried.
The pre-commit grep returns no tier-related diff outside this PR's scope.
No data migration is provided: greenfield rules apply, and the leftover
table on already-shipped instances is harmless because no code reads or
writes it after this PR.
* docs: sweep remaining references to the per-stack auto-update toggle
The previous commit retired the per-stack Auto-update gate in favor of
configuring auto-update purely through scheduled tasks. This commit
removes the now-stale mentions of that toggle across the operator docs:
- docs/features/sidebar.mdx: drop the Auto-update entry from the Inspect
group description, the matching screenshot alt-text, and the Skipper
Note that listed it. Schedule task now carries the cross-link to
Auto-Update Policies.
- docs/features/stack-management.mdx: drop the Auto-update list item;
refresh the Schedule task entry to mention the Auto-update Stack action.
- docs/features/dashboard.mdx: rename the Configuration Status row from
"Auto-update stacks" to "Auto-update schedules" with the new value
shape, and rewrite the troubleshooting accordion to describe the
scheduled-tasks invalidation path.
- docs/features/scheduled-operations.mdx: rewrite the Auto-update All
Stacks row and helper text to reflect that every stack on the node is
covered (no per-stack opt-out from this surface anymore).
- docs/features/multi-node.mdx: rewrite the Updates column definition to
derive the Auto/Off flag from enabled Auto-update Stack / Auto-update
All Stacks schedules instead of the removed per-stack policy.
The auto-update-policies.mdx rewrite in the previous commit already
covered the main reference page. The sidebar-context-menu.png screenshot
will be refreshed on release once the new menu is live in production;
the alt text is updated in this commit so it accurately describes the
shipping state.
No website edits needed: the Auto-Update Policies feature card description
("Schedule automatic image pulls and redeployments per stack on your own
cadence") and the feature matrix labels ("Auto-update stack schedule",
"Auto-update all stacks schedule") remain accurate under the new model.
* fix(stacks): drop orphaned requireAdmin import after auto-update route removal
CI's backend lint step flagged this PR's earlier deletion of the three
/api/stacks/*/auto-update routes: those handlers were the only callers of
`requireAdmin` inside routes/stacks.ts, leaving the named import on line 15
unreferenced. `requirePaid` and `effectiveTier` from the same line are still
in use elsewhere in the file and stay.
tsc --noEmit does not flag unused named imports; ESLint's no-unused-vars
does. Local backend lint reproduces and now reports 0 errors against the
existing 334-warning baseline.
|
||
|
|
2a29fed117 |
feat(labels): harden Stack Labels (gate parity, abort, dry-run, cap) (#1232)
* feat(labels): harden stack-labels surface (gate parity, abort, dry-run policy, cap) - Cap labelIds array at MAX_LABELS_PER_NODE on PUT /api/stacks/:name/labels. - Hide sidebar Labels submenu and Settings mutation buttons for roles without stack:edit, matching the backend requirePermission gate. - Break the per-label bulk-action loop on req.aborted so cancelled requests release the per-node lock once the in-flight op completes. - Reset saving state on success in LabelInlineCreateForm so the form stays interactive when reused outside the kebab/context menus. - Invoke enforcePolicyPreDeploy inside the dry-run deploy branch and report blocked stacks honestly; previously dry-run skipped the gate and would predict success for stacks the real deploy would block. * fix(labels): split sidebar gate so inline create matches unscoped backend guard Codex review of #1232 surfaced a parity miss: POST /api/labels is guarded by unscoped requirePermission('stack:edit'), but the sidebar inline "New label" entry was gated by canEditLabels (per-stack scoped). An Admiral user with only scoped grants on a stack could toggle existing labels but the inline create request would 403. Splits the sidebar gate: - canEditLabels (scoped) keeps gating the Labels submenu trigger and toggle items, matching PUT /api/stacks/:name/labels. - canCreateLabels (unscoped) now gates the inline "New label" entry, matching POST /api/labels. Also restores the swallow-catch in LabelInlineCreateForm. The earlier catch-to-finally change in #1232 let the rethrow from createAndAssignLabel surface as an unhandled event-handler rejection in the browser console. Parents already toast on failure; swallowing in the form is the intended behavior with the finally reset still in place. |
||
|
|
42e8d3a78c |
feat(security): per-image scroll + retention cap in scan history (#1231)
* feat(security): per-image scroll + retention cap in scan history Long scan histories for hot images used to monopolise the Scan history sheet: a single image with dozens of scans pushed every other image off screen, and the underlying vulnerability_scans table grew without bound. Each image group's table now renders inside its own ScrollArea capped at max-h-64 (~6 rows visible) so a busy image scrolls independently while the list of images stays navigable. A new global setting scan_history_per_image_limit (default 50, min 5, max 1000) backs both a window-function query that caps the response per image_ref and a prune step that runs on the existing MonitorService cleanup tick. The response now carries cappedImageRefs + perImageLimit so the UI can render a "Capped at N · older scans pruned" hint on groups sitting at the ceiling without a second settings round-trip. Single-image deep-dive (imageRef query param) bypasses the cap so a user clicking into one image can still see its full history. The prune uses self-contained subqueries to avoid SQLITE_MAX_VARIABLE_NUMBER issues on first-run installs with large backlogs, and explicitly deletes child rows from vulnerability_details, secret_findings, and misconfig_findings inside a transaction since FK cascade is not enabled at the connection level. Settings → Developer → Data retention gains a "Scan history per image" field. * fix(security): skip searchDraft debounce on mount to stop page-reset race The searchDraft debounce useEffect fires once on initial mount with the unchanged value and, 300ms later, unconditionally calls setPage(0). When a user (or a test) paginates inside that 300ms window, the pending debounce silently undoes the page advance. CI surfaced this as a flaky 3rd fetch in the "advances offset when the user pages forward" test once the per-image cap work added enough state-update overhead to push the click past the 300ms threshold on the slower Linux jsdom run. Track searchDraft with a ref and exit the effect when the value has not actually changed, so the debounce only runs in response to real user typing. |
||
|
|
80499ee18d |
feat(stack-activity): in-process metrics, structured diagnostic logs, docs (#1229)
Phase 3 + Phase 6 of the Stack Activity audit (PR 2 of 2): - StackActivityMetricsService: in-process counters and ring-buffered latency histogram (1000 samples per nodeId/op pair). Mirrors the FileExplorerMetricsService pattern shipped in #1216. No external export. Records (nodeId, op) where op is read or write, with success/error counts and p50/p95 latency on demand. - Admin endpoint GET /api/stack-activity-metrics returns the snapshot. Admin-only via requireAdmin, mounted next to the file-explorer metrics route. An operator debugging "why is the activity tab slow on this node?" can pull per-(nodeId, op) counts and latencies without scrolling logs. - Diagnostic logs: route handler emits a structured [StackActivity:diag] read entry per request (stackName, nodeId, limit, before, beforeId, returned, elapsedMs); dispatchAlert emits a [StackActivity:diag] write entry per persisted notification (category, stackName, nodeId, actor, messageLen). Both gated on developer_mode via isDebugEnabled. Same namespace so a single grep covers reads and writes on the timeline path. Per-request and per-event, never inside a poll loop. - Metric record points: the route's try/finally records a read metric with the outcome of the DB call; dispatchAlert records a write metric on both the success path and (before re-throwing) the failure path, so error rates from the insert path stay visible. - docs/features/stack-activity.mdx: refreshed to reflect PR 1's retention behavior (30 days plus per-(node, stack) 500-row cap, 1000 per-node unattached), composite (timestamp, id) cursor, error-vs- empty UI distinction, and "by username" vs "via Subsystem" actor rendering. Adds Troubleshooting entries for "Activity unavailable" (node disconnect or fetch failure), "expected event missing" (retention windows), and "same restart shows twice" (manual click vs Auto-Heal redeploy are distinct events). No tier, role, or capability gate touched. The admin metrics endpoint inherits the standard requireAdmin gate already used by /api/file- explorer-metrics and /api/stack-metrics. |
||
|
|
2d56ea958a |
fix(stack-activity): per-stack history integrity, attribution, sanitization (#1228)
* fix(stack-activity): per-stack history integrity, attribution, sanitization Address the Stack Activity audit findings (PR 1 of 2): - Per-stack history integrity: drop the per-insert 100-row prune in addNotificationHistory that evicted quieter stacks' history whenever another stack got chatty. Periodic cleanupOldNotifications now caps per (node, stack) at 500 rows and per-node unattached system events at 1000 rows, on top of the existing 30-day retention. Signature takes an options bag and returns a per-stage summary so MonitorService can log what actually ran each cycle. - Actor attribution: thread req.user?.username through every notifyActionFailure call site and add synthetic actors at service emit sites (system:autoheal, system:scheduler, system:image-update, system:docker-events, system:blueprint, system:monitor, system:policy). The timeline renders system actors as "via <Label>" so an autoheal redeploy is no longer indistinguishable from a user redeploy. - Message sanitization: new sanitizeNotificationMessage at NotificationService.dispatchAlert strips KEY=VALUE pairs whose key ends in TOKEN/KEY/PASSWORD/SECRET/CREDENTIALS/AUTH, scrubs HTTP basic auth in URLs and Bearer tokens, collapses COMPOSE_DIR paths, and truncates to 1000 chars. Applied to the stored history and to every downstream Discord/Slack/webhook channel. The ImageUpdateService recovery-path direct DB write also runs through the sanitizer. - Composite pagination cursor: getStackActivity now accepts a (timestamp, id) cursor (?before=&beforeId=). The legacy timestamp-only form silently dropped events when a single compose up emitted many events sharing one millisecond. Route rejects beforeId without before. - Frontend hardening: distinct error state with retry button (initial fetch failure no longer renders as the genuine empty state), strict positive-integer parsing on cursor params, overrequest-by-1 pagination so the last page does not leave a dead "Load more" click, runtime guard on liveEvents merge that validates the level union, per-minute day-bucket recompute so an open panel does not stay on "Today" past midnight. No tier, role, or capability gate touched. Route permission gate remains stack:read on the named stack. * fix(stack-activity): sanitizer covers lowercase env vars and per-node compose dir External review surfaced two leak paths in the message sanitizer: - The sensitive-key regex was uppercase-only. Compose env names are conventionally uppercase but lowercase forms (db_password, jwt_secret, github_token) are valid and do leak through the same Docker and compose-parse error paths. Make the regex case-insensitive and tighten it to also catch bare TOKEN= / KEY= / PASSWORD= without a prefix word, while still leaving BYPASS, COMPASS, and similar non-secret keys alone. - The compose-dir path collapse only read process.env.COMPOSE_DIR, but the real resolution chain is node.compose_dir (per-node DB override) -> process.env.COMPOSE_DIR -> /app/compose. A node with a custom compose_dir could still leak absolute paths into stored history and downstream channels. Route both the dispatchAlert call and the ImageUpdateService recovery-path direct write through NodeRegistry.getInstance().getComposeDir(localNodeId) so the collapse covers every resolution outcome. Tests now assert lowercase keys are redacted and that BYPASS-style non-secrets stay intact in both cases. notification-routing mock extended to stub the new getComposeDir call. * chore(stack-activity): a11y roles, visibility-aware tick, live-disconnect signal Close three small follow-ups on the per-stack activity timeline: - A11y: each day-group gets role="list" and each event row gets role="listitem" so screen readers traverse the timeline as a list instead of a wall of text. The day-group container also carries an aria-label naming the bucket. - Visibility-aware day-bucket tick: the 60s setInterval that re-derives Today/Yesterday/Earlier now short-circuits when document.hidden, so a backgrounded panel does not re-render every minute for no visible effect. - Live-disconnect signal: useNotifications dispatches a sencho:notifications-connection custom event on WebSocket open and close. The timeline listens and, when explicitly disconnected, shows a one-line "Live updates offline; reconnecting…" hint above the list. The sidebar ticker already surfaces fleet-wide connection state; this adds an in-context cue for users who are focused on a single stack. Stack-name case normalization was considered and rejected: stack names are case-permissive per the isValidStackName validator, and lowercasing on read or write would silently rename or hide a user's "MyApp" stack. * ci(stack-activity): drop unnecessary escape in URL_BASIC_AUTH regex ESLint no-useless-escape errored on \- inside the character class [a-zA-Z0-9+.\-] at notificationMessage.ts:14. Move the dash to the end of the class so it's an unambiguous literal and the escape is no longer required. Behavior is identical; sanitizer tests still pass. * revert(stack-activity): drop unvalidated E2E spec from this PR The spec was committed without ever running against a real Docker daemon, then failed in CI when it ran for the first time: deploy returned 200 but no notification appeared on the activity endpoint within the polling window, suggesting either a deploy-notification race or a node-id resolution mismatch in the CI environment. Backend unit tests (route + composite cursor + sanitizer) and frontend component tests cover the same logic. The E2E spec will land in a dedicated follow-up once it has been authored against a working CI environment. |
||
|
|
117f590332 |
fix(security): gate admin-only scan affordances on isAdmin (#1230)
* fix(security): gate admin-only scan affordances on isAdmin Backend already required admin for SBOM, SARIF, scan policies, Trivy install/update/uninstall, the auto-update toggle, CVE suppressions, and misconfig acknowledgements. The matching frontend surfaces were gated only on isPaid (or only on isReplica), so non-admin users at the same tier saw buttons that returned 403 on click. Threads isAdmin from useAuth() into SecuritySection, SuppressionsPanel, and MisconfigAckPanel. Updates the scan-result sheet caller in ResourcesView so SBOM and SARIF render only when paid AND admin; passes canManageSuppressions to the stack-misconfig sheet so admins can ack misconfigs from that surface too. Read paths remain visible to non-admins (policy list, suppression list, ack list, scan history) since the GET routes are auth-only on both sides. * fix(security): close 3 remaining scan-sheet parity gaps from review Code review surfaced three sites missed in the first pass: 1. SecurityHistoryView opened the scan sheet with canGenerateSbom set only on isPaid, so Skipper non-admins saw SBOM and SARIF buttons even though the backend requires admin+paid. Now ANDed with isAdmin. 2. ResourcesView passed onRescan unconditionally, and the sheet renders a Re-scan primary action whenever onRescan is defined. Non-admins reaching the sheet via the severity-badge shortcut saw the button; clicking it called POST /security/scan, which the backend requires admin for. onRescan is now undefined for non-admins. 3. ShellOverlays and ResourcesView passed canManageSuppressions=isAdmin without considering the replica gate, so a replica admin saw suppress and ack columns whose backend writes blockIfReplica. The sheet now probes /fleet/role internally and ANDs !isReplica into the effective canManageSuppressions, so the column hides on a replica regardless of how the caller wired the prop. * fix(security): clear isReplica state on every scan-sheet probe The previous probe only flipped the state to true on a replica response and never wrote false on a control, non-OK, or skipped probe. With the sheet kept mounted by ResourcesView, SecurityHistoryView, and ShellOverlays, an admin who first viewed a scan on a replica would keep suppress/ack controls hidden even after switching to a control instance, because the stale true value persisted across re-opens. The effect now resets isReplica to false at the start of every probe and assigns the result of /fleet/role directly. Probe failures and skips leave the state at false, so the UI is permissive and the backend blockIfReplica guard remains the source of truth. |
||
|
|
19c28b77b5 |
docs: soften Admiral tier wording from "enterprise-grade" to "fleet-wide governance and operational" (#1226)
"Enterprise-grade" is an absolute readiness claim that overstates the maturity of a pre-1.0 product. The replacement names what the Admiral tier actually covers (the items already listed in the parenthetical: audit log, host console, cross-node Mesh traffic management, federation overrides, fleet-wide policy push) without leaning on marketing language. |
||
|
|
a4a8abb5f8 |
fix(stack-files): guard download stream destroy against the supertest in-process close race (#1227)
* fix(stack-files): guard download stream destroy against the supertest in-process close race The download metric was still booking the occasional request as a failure even after PR #1220 moved tracking off res.{finish,close} and onto stream.{end,close}. The remaining race lives one level up: under the in-process supertest transport, req.on("close") can fire as soon as the test consumes the response, before the readable's own end/close pair has been dispatched. The unconditional req.on("close", () => result.stream.destroy()) line forced the readable into a close-without-end state every time that happened. The close handler then booked a failure even though the pipe had delivered every byte. Two layers of defense: - req.on("close") now only destroys the stream while the response is still in flight. If res.writableEnded is already true the pipe has finished and the readable will end on its own; there is nothing to clean up and no synthetic close to provoke. - stream.on("close") falls back to recording success when the response was already fully written. Real disconnects keep recording a failure because res.writableEnded stays false in that case. The semantic stays unchanged: success means the server finished reading the file off disk and pushed it into the pipe. The fix removes the spurious failure path without weakening the disconnect-detection. * chore(stack-files): temporary download diagnostic for the metric race Logs the event order and response state at each handler so the next CI run shows whether req-close fires before stream-end, what res.writable / writableEnded / writableFinished evaluate to at that moment, and which path actually records the failure. Will be removed in the follow-up commit once the race is understood. * chore(stack-files): remove temporary download diagnostic Diagnostic captured the happy-path event order in the previous CI run; removing it now to isolate whether the fix alone is sufficient. |
||
|
|
b5709ee801 | Fix links and update feature notes in README | ||
|
|
e4d94613c9 | Update text for consistency in README.md | ||
|
|
15aec0e0a3 |
Correct reference to Known_Limitations in README
Updated references to 'KNOWN_LIMITATIONS.md' for consistency and clarity. |
||
|
|
f32b6372a1 |
docs: pre-1.0 readiness pass for public beta launch (#1225)
Close the trust-blocking items from the pre-v1.0 readiness audit so the repository is ready for the first public posting. No backend or frontend code changes; no tier gates move. README: - Add a single beta-status GitHub [!NOTE] callout below the dashboard image. Beta status also appears in selected install-flow docs (Quickstart, Known Limitations, Security Architecture, Upgrade, Troubleshooting) so install-flow readers are not surprised. - Add a "Before you install" subsection that names the docker.sock privilege model with the Portainer / Dockge / Komodo comparison. - Fix the docker run example: add the missing /opt/docker:/opt/docker mount that COMPOSE_DIR=/opt/docker depends on. - Reword the two "no exposed Docker socket" sentences so the claim is scoped to remote / cross-node exposure. - Reword "transparent HTTPS proxy" to "authenticated HTTP and WebSocket proxy" with a TLS / VPN reminder. - Fix the RBAC bullet: the role set is admin, viewer, deployer, node-admin, auditor. There is no "editor" role. - Add tier markers to the Capabilities list (matrix sentence plus per-bullet (Skipper) / (Admiral) markers), verified against the route guards in backend/src/routes/. - Fix the broken notification-routing link to point at the existing alerts-notifications#notification-routing anchor. - Soften the BSL paraphrase to point at LICENSE plus the license FAQ. - Add a "Telemetry and data handling" section: no telemetry, no analytics, no crash reports; license validation only when a paid key is activated. - Add a "What Sencho is not (yet)" section so the scope boundaries are visible above Capabilities. Templates: - PR template: drop the contradictory CHANGELOG checkbox; release-please owns CHANGELOG. - Bug report template: update the stale 0.2.2 version placeholder to 0.86.6 and add fields for compose snippet, container logs, browser console, and an involved-subsystems checkbox. Docs: - docs/operations/trivy-setup.mdx: replace both sencho/sencho:latest occurrences with the published image saelix/sencho:latest. - docs/reference/settings.mdx: rewrite the API Tokens Note (no tier gate in code; admin role only) and the Stack Labels Note (basic CRUD is free; only bulk actions require Skipper or Admiral). - Add the shared beta-status Note callout to docs/getting-started/ quickstart.mdx, docs/operations/upgrade.mdx, docs/operations/ troubleshooting.mdx, and docs/reference/security.mdx. CONTRIBUTING: - Replace the public link to the in-repo coding-rules file with a pointer to docs.sencho.io for architecture deep-dives. - Fix the sample clone URL case (Sencho.git becomes sencho.git). New files: - SUPPORT.md: where to ask, response-time expectations, in / out of scope. - KNOWN_LIMITATIONS.md: scale, platform, architecture, and feature limits documented for the beta audience; scale numbers marked "not benchmarked yet" pending real benchmarking. |
||
|
|
05c3975d6d |
test(dashboard): cover dashboard routes, ConfigurationStatus tier parity, and useMeshDataPlane (#1221)
* test(dashboard): cover dashboard routes, ConfigurationStatus tier parity, and useMeshDataPlane The dashboard router had no dedicated Vitest coverage; tier parity in the ConfigurationStatus component was only proved by manual inspection; and the Admiral short-circuit in useMeshDataPlane had no automated regression net. Add three spec files: - backend/src/__tests__/dashboard-routes.test.ts: 11 cases against the live Express app. Both routes reject unauthenticated requests; the configuration response matches its documented shape; the tier x variant `locked` matrix is asserted end-to-end for Community, Skipper, and Admiral via LicenseService spies; a seeded Discord agent URL is shown never to appear in the serialized response; /stack-restarts clamps days values of 0, 999, and NaN without bailing. - frontend/src/components/dashboard/__tests__/ConfigurationStatus.test.tsx: five render cases prove the parity contract. Community hides the entire Automation section plus the four gated rows (Notification routing, Webhooks, Scheduled tasks, Vulnerability scanning); Skipper shows everything except Scheduled tasks (Admiral-only); Admiral shows every gated row plus the SSO provider name mapping (oidc_google -> "Google"). Skeleton and load-error paths are also covered. - frontend/src/components/dashboard/__tests__/useMeshDataPlane.test.tsx: four hook cases prove the Admiral short-circuit. Non-Admiral sessions never fire /mesh/status; Admiral sessions fetch once and populate the localDataPlane payload; a 403 response leaves status null without raising; a response that omits localDataPlane also leaves status null. Backend route suite + dashboard-only frontend suite green in isolation. The full backend suite shows one pre-existing Windows-only EBUSY flake in filesystem-backup.test.ts (SQLite file lock on unlink) that reproduces on the unmodified branch tip and is unrelated to these changes. * test(dashboard): drop backup.requiredTier from ConfigurationStatus fixture The fixture's `backup.requiredTier: 'admiral'` field was authored to match the type on this branch's original base. Main has since removed that field from the ConfigurationStatus payload, so the fixture now over-specifies a property the type forbids and fails tsc. Drop the field to realign with the current type. |
||
|
|
03a5826f7e |
fix(dashboard): debounce state-invalidate refetches (#1209)
* fix(dashboard): debounce state-invalidate refetches and drop redundant listener useDashboardData fired three immediate HTTP requests (/stats, /system/stats, /stacks/statuses) for every Docker container event. A burst restart of a 50-container stack produced ~150 instant requests against the local instance with no throttle. Add a 250 ms trailing-edge debounce so an event storm collapses into a single coalesced refresh, mirroring the precedent in useNextAutoUpdateRun. The cleanup function now also clears any pending debounce timer so a late event cannot fire after the dashboard unmounts. useConfigurationStatus subscribed to the same event but its data is built from settings and policy tables (agents, alert rules, auto-heal, scheduled tasks, scan policies, backup config), none of which change on container state. Drop the listener entirely; the 60 s poll catches rare settings edits with acceptable latency. * fix(dashboard): scope settings-event listener back into useConfigurationStatus Address two follow-up findings from independent review of the earlier commit on this branch. 1. Restore a filtered sencho:state-invalidate listener in useConfigurationStatus. The earlier commit dropped the listener wholesale to keep container-event bursts from refetching settings data, but that also silenced the only settings-affecting event in the current taxonomy: action='auto-update-settings-changed' (emitted from backend/src/routes/stacks.ts when a user toggles a stack's auto-update setting). With the listener gone, the Configuration Status row for Auto-update stacks could sit stale until the 60 s poll. The new listener mirrors the precedent in useNextAutoUpdateRun: filter on the single configuration-relevant action, trailing-edge debounce 250 ms. 2. Add an `active` flag to the useDashboardData state-invalidate effect. Cleanup already clears the pending debounce timer, but a refresh() already in flight could still call setters after unmount because the awaited Promise.all has no abort hook. The flag is checked both before the await and after, matching the cleanup shape used by useNextAutoUpdateRun. Tests cover both: the configuration listener now ignores scope='stack' and scope='image-updates' bursts and refetches once on a settings-changed burst. |
||
|
|
7c3ba3f24d |
feat(dashboard): surface metrics-stale indicator after sustained poll failure (#1213)
* feat(dashboard): surface metrics-paused indicator after sustained poll failure useDashboardData previously failed silently when /stats or /system/stats returned an error: stale data kept rendering and the last sync timestamp quietly drifted. The operator could not tell whether the dashboard was just slow or whether the Docker socket / metrics path had genuinely gone down. Track consecutive failures per live-metrics endpoint. After three in a row on either /stats or /system/stats (≈15 s at the 5 s poll cadence), expose a metricsStale boolean on the hook result. HealthStatusBar renders a small amber "metrics paused" chip beside the meta line when set. The indicator clears on the first successful response when both endpoints are within the threshold. A unit test for the threshold logic is intentionally deferred to the Phase 4 E2E dashboard spec, which exercises the same path end-to-end by stopping the Docker daemon and asserting the user-visible indicator. * fix(dashboard): rename stale-metrics chip and cover the threshold with tests Address two follow-up findings from independent review of the earlier commit on this branch. 1. Rename the masthead chip from "metrics paused" to "metrics stale". The hook keeps polling on every cycle; the chip describes the freshness of the displayed numbers, not the polling cadence. The new wording matches the underlying `metricsStale` state variable. 2. Add a Vitest spec for the threshold logic. Captures the visibilityInterval callback at registration time and drives each polling cycle on demand, covering: three consecutive /stats failures trip the indicator and the next successful poll clears it; three consecutive /system/stats failures trip the indicator on the other endpoint; clearing requires both endpoints under threshold (a single endpoint recovering while the other is still failing keeps the indicator set). The clarifying comment in useDashboardData notes that polling is unaffected and only the data freshness is in scope, so future readers do not interpret "stale" as "paused". |
||
|
|
e183153a64 |
chore(dashboard): drop misleading backup.requiredTier from configuration payload (#1212)
Cloud Backup has a per-provider tier: Custom S3 is open to every tier (PR #1143) while Sencho Cloud Backup requires Admiral. A single backup.requiredTier='admiral' on the configuration response misrepresented that split, and no consumer ever read the field. Remove it from the response interface and the response builder; update the frontend mirror type accordingly. Annotate the backup block so the per-provider intent is clear at the call site. |
||
|
|
0db0d29f3b |
fix(dashboard): decouple FleetHeartbeat refresh from the active local node (#1210)
useFleetHeartbeat keyed its effect on activeNode.id and reset its state on every node switch, even though /fleet/overview returns a fleet-wide payload that does not change when the user pivots their active local node. The result was a needless flicker back to the skeleton card and an extra HTTP request on every node pivot. Drop the nodeId dependency and the stale-node guard ref. The 30 s visibility-interval poll remains, so transient remote-node offline state still surfaces within one polling cycle. |
||
|
|
6d995b9aaf |
fix(dashboard): slow HealthStatusBar sync-label tick to 5s (#1211)
useTicker(1000) rendered the masthead once per second to advance the "last sync Xs" label. The label only shifts visibly every few seconds (1s, 6s, 11s..., then m, then h), so the per-second cadence forced the entire dashboard tree through the React reconciler 60 times per minute for a visual change the eye does not see. A 5s tick keeps the label fresh while cutting the wake-up rate by 5x. Name the constant so the trade-off is documented at the call site. |
||
|
|
9e20acd647 |
chore(dashboard): add developer-mode timing diagnostics on dashboard routes (#1219)
Both dashboard endpoints (/configuration and /stack-restarts) execute on the hot path of every dashboard mount or refresh, and one of them (/configuration via buildLocalConfigurationStatus) reads from a dozen database tables. When an operator reports a slow dashboard on a large deployment, there is currently no instrumentation to point at which endpoint is the offender. Gate two new debug lines behind isDebugEnabled (the existing developer-mode flag, sourced from DatabaseService.global_settings.developer_mode). Each line reports the elapsed milliseconds and a single contextual field (nodeId, row count, days window). Both endpoints exit unchanged when developer mode is off; the timing measurement and console call are skipped entirely, not just suppressed. Both error paths already log via console.error and stay that way; per backend/src/utils/debug.ts comments, error paths are exempt from the debug gate. |
||
|
|
ca144f07d9 |
chore(dashboard): drop unused AgentStatus exports on both sides (#1222)
The Phase 5 dead-code sweep across the dashboard call graph found two identically shaped findings: the AgentStatus interface is declared and exported in both backend/src/routes/dashboard.ts and frontend/src/components/dashboard/useConfigurationStatus.ts, but no other file imports it. (The frontend StackAlertSheet component has a separate, differently shaped private AgentStatus that does not refer to either of these.) Drop the export keyword on both. The interfaces stay alive as file-internal types, the public surface shrinks by two names, and no behaviour changes. The broader payload-cleanup opportunities surfaced during the sweep (unused requiredTier fields on individual row objects, unused top-level tier and variant fields on ConfigurationStatus) are out of scope for this PR and have been filed as Linear roadmap items. |
||
|
|
96c7104521 |
docs(dashboard): refresh Troubleshooting accordion for new metrics-stale chip and tightened refresh model (#1223)
Three accordion edits keep the user-visible behaviour described on docs/features/dashboard.mdx in step with the audit's surface changes. - "Configuration Status still shows the old value after I changed a setting": replace the broad "most settings dispatch a live invalidation" language with the precise behaviour, which is that only a stack Auto-update toggle triggers an immediate refetch; every other settings edit waits for the 60-second poll. The change avoids promising responsiveness the card cannot deliver and points the operator at the hard-reload escape hatch. - New "The masthead shows a 'metrics stale' chip" entry: describe what the chip means, the threshold (three consecutive metrics-endpoint failures), that polling continues regardless, and that the chip describes data freshness rather than the polling cadence. Points the operator at the Docker daemon and Sencho container logs as first checks. - New "The dashboard feels sluggish on a large deployment" entry: document the Developer mode toggle as the supported diagnostic path for slow-dashboard reports. Lists the [Dashboard:debug] log shape so the operator knows what to look for, and reminds them to disable the toggle afterwards. |
||
|
|
d6afc298da |
docs(stack-files): refresh the file explorer page after the audit batch (#1224)
Brings the customer-facing page in line with what shipped in #1200 through #1220: - Names the stack-read capability without enumerating which roles carry it. Every signed-in role does today, so the wording is forward-compatible with a future role that omits it. - Updates the directory display cap to 1000 and points at the new filter input above the tree (the previous text quoted 500 and a shell-only fallback that no longer matches the UI). - Explains the unsaved-edits confirmation on file switch, the optimistic-concurrency reconcile flow with the file-changed- elsewhere notice, and the atomic write semantics in a single short paragraph. - Updates the upload row to describe drag-and-drop, the Replace confirmation on same-name conflicts, and the brand-coloured drop target hover. - Replaces the four troubleshooting accordion entries the audit asked for in plain product language: stack:read 403, protected delete, file-changed-elsewhere 412, and DISK_FULL on upload. No internal-tooling or fence-spec language. No tier-bypass walkthroughs. No legacy phrasing. |
||
|
|
55737ab34d |
test(stack-files): end-to-end coverage for the file explorer audit surface (#1218)
* test(stack-files): add end-to-end coverage for the file explorer audit surface Companion to e2e/stack-files.spec.ts, which already pins the basic admin and community-tier upload/edit/delete/download flows. The new spec adds the harder-to-reach surfaces the audit flagged: - Path-traversal corpus on GET /files: nine variants ( .. , ../etc, a/../b , absolute POSIX, Windows drive, backslash, double-slash, NUL byte, a/./b ) each asserted to return 400 INVALID_PATH. - Protected-file enforcement via direct API: compose.yaml and .env delete return 409 PROTECTED_FILE; a subdirectory entry named compose.yaml still deletes (root-scope guarantee preserved). - Optimistic concurrency: two writers race on the same file via If-Match headers; the loser sees 412 PRECONDITION_FAILED with the current content payload, and the winner's bytes are on disk. - Large directory truncation: 1100 entries surface 1000 in the body with X-Truncated: true and X-Total-Count headers intact. - Binary detection and force=text override: a UTF-8 file carrying a NUL byte returns binary:true by default and is rescued via ?force=text. Oversized files keep the no-inline-content contract even when force=text is set. - 25 MB upload cap: a 25 MB + 1 byte payload trips the multer LIMIT_FILE_SIZE branch and returns 413 TOO_LARGE. - Symlink semantics: DELETE removes only the link entry and leaves the target intact; PUT permissions on a symlink returns 409 LINK_CHMOD_UNSUPPORTED with the target mode unchanged. - UI lifecycle: upload via the dropzone then API rename, and API mkdir; both verify the resulting tree state by reloading the Files tab. Four deferred surfaces are declared as test.skip with a one-line rationale each: the remote-node matrix and mid-op disconnect (require a real peer enrolled in CI), the developer-mode diagnostic matrix (no stable hook for backend stdout in Playwright; covered at the route-test layer), and the viewer-role tier-persona matrix (role gating is covered at the route-test layer). * test(stack-files): trim memory-heavy cases out of the E2E spec CI ran the spec end-to-end and the run completed, but the two heaviest cases (1100-file directory seed and a 25 MB upload buffer) pushed the single-worker Playwright queue past the dashboard-render budget for the specs that followed. stack-files.spec.ts and stacks.spec.ts saw their loginAs await time out at 10 s while the post-spec sidebar was still loading. Both behaviours are exercised at the backend route-test layer (stack-files-routes.test.ts pins the 1000-entry truncation header on a 1100-file directory and the multer LIMIT_FILE_SIZE 413 with a 26 MB buffer), so deleting them from the E2E spec loses nothing material. The oversized force=text test is also trimmed from 2.5 MB to 2.1 MB, which still exercises the >2 MB branch on the backend. The deferred-coverage block now lists the two excised cases with the rationale for why E2E is the wrong layer for them. * test(stack-files): collapse 7 per-describe seed hooks into 1 file-scope pair Each inner describe previously ran its own seedSuite + teardownSuite pair: 7 logins, 7 stack creates, 7 stack deletes, all serial. That added 15-20 seconds of CI setup overhead on a single-worker run and pushed the post-spec dashboard load past the 10 s budget some downstream specs allow on their loginAs await. Moving the seed/teardown to file scope keeps the test stack present for every test in this file with one fixture pair. Each test still calls loginAs in its beforeEach so the page state is fresh, but the expensive setup runs once instead of seven times. |
||
|
|
f86042b2ad |
fix(stack-files): track download metric off the file stream, not the response (#1220)
The download success metric was hung off res.on('finish') with a
res.on('close') fallback for failure. Under the in-process supertest
transport that pair fires in non-deterministic order: in CI the close
event sometimes precedes finish on a clean response, and the recorder
booked the request as an error.
The file stream's lifecycle does NOT race. fs.ReadStream emits:
- 'end' then 'close' on a clean read,
- 'close' without 'end' when destroy() runs (the req-close path),
- 'error' then 'close' on a disk error.
Moving the recorder onto these signals removes the race. The flag still
prevents double-firing when 'close' chases 'end' on success.
Semantic note: success now means "the server finished reading the file
off disk and pushed it into the pipe", which is the strongest signal a
server can produce. Whether the client received the bytes is outside
the server's observable state and was never the actual measurement.
|
||
|
|
9f2f13f35a |
feat(stack-files): in-process metrics and structured mutation logs (#1216)
* feat(stack-files): in-process metrics and structured mutation logs
Adds FileExplorerMetricsService, an in-memory counter and latency
histogram keyed by (nodeId, op) modelled on StackOpMetricsService.
record() is called once per file-route request from a small
recordFileOp helper that wraps the metric capture; rejection paths
that ran real filesystem work (overwrite confirms, write conflicts,
multer oversize) now record an error count and a warn log instead of
disappearing from the snapshot. recordUploadBytes tracks bytes that
actually persisted so a node taking many small uploads vs a few large
ones is visible in the dashboard.
Admin-only GET /api/file-explorer-metrics returns the snapshot in the
same shape as /api/stack-metrics so an operator chasing a slow node
has a single place to look. No external telemetry; everything is
process-local and resets on restart.
Mutation INFO lines now carry op, stack, path, and bytes/mode/
recursive/overwrite/toPath in the structured details so log scrapers
can pivot on the same identity the metric uses. The existing
developer_mode gate on logFileDiag is unchanged. A new
rejectFileMutation helper centralises the log+metric+response triple
on the three rejection sites (upload DIR_EXISTS, upload FILE_EXISTS,
write PRECONDITION_FAILED) so a future rejection cannot skip the
metric.
Tests cover the service in isolation (counts, p50/p95, ring buffer cap,
upload bytes tracking, snapshot sorting), the admin route auth and
shape, and the route layer end-to-end: a real upload surfaces in
/api/file-explorer-metrics, a FILE_EXISTS rejection bumps errorCount,
and the structured INFO line carries the expected fields.
* fix(stack-files): tighten download/upload latency tracking
Two metric-accuracy bugs caught in independent review:
Download metric was recorded as a success before result.stream.pipe(res)
ran. A mid-stream read failure or a client disconnect was not counted as
an error because the recorder fired at pipe time, not stream completion.
Now the recorder hangs off res.on('finish') for success and on both the
stream's error event and res.on('close') for failure, with a flag so a
normal completion (which emits both finish and close) does not produce
two recordings.
Upload latency was inconsistent across the success and failure branches.
The multer wrapper captured startedAt at route entry, but the async
handler created its own startedAt after multer had already buffered the
body. Successful uploads therefore reported only the post-multer time
and the multipart transfer/buffer cost vanished from the histogram. The
wrapper now stashes the route-entry timestamp on the request object and
the async handler reads it back, so every metric for a given upload
shares one window.
A new test pins the download recorder behaviour: a single successful
download must produce successCount=1, count=1, errorCount=0 in the
snapshot, which would have flagged the original double-fire path the
finish+close pair could have introduced.
|
||
|
|
d8b6f8cf3b |
feat(stack-files): force-text override for misidentified binary files (#1215)
* feat(stack-files): force-text override for misidentified binary files
The binary-detection heuristic (30% non-printable / NUL in the first
8 KB) sometimes flags UTF-8 files that happen to carry an embedded NUL
or a high non-printable ratio, locking the user out of inline editing
with only a Download fallback.
readStackFile now accepts an optional { forceText: true } that bypasses
isBinaryBuffer on the small-file path and returns the bytes as UTF-8
content. The route exposes this as ?force=text on GET /files/content.
The oversized branch deliberately stays untouched: returning a multi-MB
file as JSON-encoded text is wasteful regardless of the heuristic.
SpecialFilePanel grows an optional extraAction slot. The viewer's
binary branch wires Open as text anyway, which refetches with the new
flag, clears isBinary, and routes the content through the existing
Monaco editor path. A failed override surfaces both an inline error
panel and a toast so the user knows why the click did nothing.
Backend tests pin the heuristic-vs-override behaviour pair on a file
with a literal NUL byte. The frontend test asserts that the second
readStackFile call carries forceText: true and that Monaco mounts.
Troubleshooting accordion entry updated to mention the new affordance.
* fix(stack-files): guard the binary-override path against oversized files
The override on the binary panel could open Monaco against an empty
content buffer if the backend's oversized branch ran (files past the
2 MB inline-preview cap intentionally carry no content even when
force=text is set). Saving that empty buffer would wipe the file on
disk.
Two reinforcing changes:
- Initial load now checks result.oversized before result.binary, so a
file that is both oversized and has binary bytes in the 8 KB probe
shows the Download panel rather than the binary panel. The size
signal stays in front of the operator and the override button never
surfaces for a file that cannot be safely opened inline.
- The handleForceText handler now respects result.oversized on the
refetch and transitions to the Download panel instead of clearing
isBinary and copying result.content ?? '' into Monaco.
Same handler also gains a stale-request guard via a selectedPathRef:
a slow override for file A no longer stomps on file B's state if the
user navigated away while the request was in flight.
Two regression tests pin the new behaviour: oversized+binary surfaces
the Download panel on initial load, and an oversized refetch from the
binary panel routes to the Download panel rather than Monaco.
|
||
|
|
c2357ec534 |
fix(stack-files): symlink-aware delete and chmod (#1214)
deleteStackPath now lstats the leaf and unlinks the link entry itself when it is a symbolic link, so the file the user clicked on in the tree is what gets removed (the linked target stays intact). chmodStackPath rejects with LINK_CHMOD_UNSUPPORTED on a symlink rather than silently mutating the target's permissions; Node's lchmod is macOS-only and following the link is the bug being fixed here. Path-component symlinks are still resolved via the existing resolveSafeStackPath, so a symlinked parent that escapes the stack dir still surfaces SYMLINK_ESCAPE before the leaf is inspected. Service-level tests cover delete on internal-target / external-target / broken / dir-target symlinks, chmod rejection on symlinks (including the broken case), and non-symlink regression checks. Route-level tests pin the 409 LINK_CHMOD_UNSUPPORTED mapping and the link-only-delete behaviour. The describe blocks are platform-gated; Windows symlink creation needs admin/developer-mode and is skipped along with the existing SYMLINK_ESCAPE test. Docs updated to describe both behaviours in plain product terms. |
||
|
|
ba4de2e004 |
test(stack-files): pin multipart-body forwarding through the remote-node proxy (#1217)
conditionalJsonParser skips express.json() when the request has an x-node-id pointing at a remote node, leaving the raw body stream intact so the remote proxy middleware can pipe it upstream. No existing test pinned that the skip also applies to multipart payloads, not just JSON. A regression that re-enabled body-parser on multipart would silently strand POST /api/stacks/<name>/files/upload to remote nodes: the proxy would forward an already-drained stream, the upstream multer would see an empty body, and the user would get a 400 from a successful-looking request. Two cases pin the behaviour: - An in-process http capture server is registered as a remote node. A multipart upload through the central with x-node-id set must arrive with Content-Type carrying the boundary, Content-Length matching the raw byte length, the original file bytes intact, an envelope larger than a half-drained stream could ever produce, and the user JWT rewritten to the remote node's api_token so the central does not leak its session token to the peer. - A second case extracts the boundary from the Content-Type header and checks both that the body bytes contain --<boundary> and that the original filename is preserved across the proxy. No production code change. The test runs green against current main; the value is in regression prevention for a load-bearing assumption. |
||
|
|
fcf2222604 |
feat(stack-files): cap directory listings at 1000 + add file-tree filter (#1208)
* feat(stack-files): cap directory listings at 1000 + add file-tree filter
The file-tree route returned every entry in a directory unbounded.
A logs/ or data/ subfolder with rotated artifacts could produce a
multi-megabyte response and a frontend cap at 500 entries silently
hid the rest with no way for the user to find a specific file.
The list route now caps the response at 1000 entries (the audit's
recommended bound), advertises the unfiltered total via
X-Total-Count, and sets X-Truncated when truncation happened. The
service exposes both a bare-array listStackDirectory (unchanged
contract for callers that just want the array) and a paginated
listStackDirectoryPage that returns {entries, total, truncated}.
The FileTree now offers a search input above the scroll area that
filters loaded entries by name (case-insensitive substring). Clearing
the filter restores the full listing. A non-matching filter shows a
short hint instead of an empty pane. The client-side MAX_ENTRIES
matches the server cap so a perfectly-sized directory never shows
the truncation hint.
* fix(stack-files): filter keeps parent dirs when loaded descendants match
The original filter applied per-render-level inside renderEntries, so a
parent directory whose name did not match was filtered out even when one
of its already-loaded children did. The match was then unreachable: the
parent had been removed from the visible list and its children never got
a chance to render.
Compute matching-descendant once per directory by walking the loaded
dirContents map (no extra fetch, bounded by what the user already
expanded). Keep ancestors of any match in the visible list. Auto-expand
those ancestors for the duration of the filter so the match comes into
view without a manual click on every parent.
Filter scope is still 'what is already loaded'; unexpanded subtrees do
not contribute to ancestor-keep until the user expands them. Two new
tests pin both behaviours.
|
||
|
|
ea002cd9a0 |
feat(stack-files): drag-and-drop upload zone (#1207)
* feat(stack-files): drag-and-drop upload zone The Files-tab dropzone was a click-only button. Drag-and-drop is the standard file-manager affordance and the audit doc named its absence as a known gap. The same dropzone now accepts file drops, gates the affordance on canEdit, and reuses the existing handleFile pipeline so all the existing rules apply: 25 MB cap, server-side path validation, multi-file drops rejected up front with a clear toast (the upload route only handles one file per request). dragenter+dragover light the zone in the brand color and swap the label to "Drop to upload"; dragleave honours bubbling from child nodes so the highlight does not flicker when the cursor crosses an inner icon. Drag events without a Files payload (text selection, image drag from another page) are explicitly ignored so the zone does not light up for non-file content. * test(stack-files): widen upload-zone E2E selector to match new label The drag-drop work renamed the dropzone label from 'Upload file' to 'Upload or drop file' (and 'Drop to upload' during hover). The E2E test in stack-files.spec.ts filtered on the literal /upload file/i, which no longer matches the new label and the test timed out at the visibility assertion. The hidden file input's aria-label='Upload file' was deliberately preserved, so every other locator in the spec (input[aria-label], getByLabel) keeps working. Only the role='button' filter needed the update. Widen the regex to /upload/i so future copy edits do not re-break this assertion. |
||
|
|
4964320f50 |
fix(stack-files): optimistic concurrency on file-tab writes via mtime ETag (#1206)
* fix(stack-files): optimistic concurrency on file-tab writes via mtime ETag PUT /api/stacks/:name/files/content previously did a blind write; two operators editing the same script lost one of the saves with no warning. The compose-file editor already had mtime optimistic concurrency (PR #1183); this brings the file-explorer write path to the same shape. GET /files/content now also returns mtimeMs and sets a weak ETag header derived from the stat. The matching PUT reads If-Match, asks FileSystemService.writeStackFileIfUnchanged to compare against the live mtime, and returns 412 PRECONDITION_FAILED with the current content and mtime when the stale-write check fails. Successful writes echo a fresh ETag so the client can pin the next save without re-GET. readStackFile and writeStackFileIfUnchanged each open the file once and stat+read through the same handle so the mtime returned to the client matches the bytes that were sent, even if the file is replaced between the two operations. PUT without If-Match still succeeds (backward compatibility with scripted clients that do not roundtrip the ETag). FileViewer now sends the loaded mtime on save, updates its local mtime from the success response, and on FileConflictError adopts the server snapshot as the new baseline so the user's follow-up edit-and-save does not loop on the same precondition. * fix(stack-files): treat deleted-target as conflict; preserve user buffer on conflict Two follow-ups from code review on the prior commit: - writeStackFileIfUnchanged now returns ok:false when expectedMtimeMs is set and the target has been deleted. The caller was editing a file that no longer exists; silently writing the buffer to the void is wrong. The client adopts the empty snapshot as 'file is gone, start over' and the user keeps control of what to save next. - The FileViewer conflict handler no longer overwrites the user's typed buffer with the server snapshot. It updates the baseline so the next save sends the fresh mtime, then leaves the editor content alone. The user sees their edits, the Save button stays enabled, and a follow-up click applies their changes on top of the new server version without silently destroying what they typed. * fix(api): preserve default headers when caller supplies a headers field apiFetch built defaultOptions.headers by merging Content-Type, x-node-id, and the caller's headers, but then spread the unmodified fetchOptions over defaultOptions at the outer level. The spread overwrote the merged headers with the caller's bare headers, silently dropping Content-Type on every request that supplied any custom header. This was latent until the file-explorer save path started sending an If-Match header. The Express body parser refused the PUT without Content-Type, the route returned 400, the editor showed an error toast instead of the success toast, and the Playwright save assertion timed out. Destructure headers out of fetchOptions before the outer spread so the already-merged defaultOptions.headers survives. Add api.test.ts with four regression cases pinning Content-Type, the If-Match merge, x-node-id presence when active, and localOnly skip. |
||
|
|
668eda6cc8 |
fix(stack-files): atomic write via tmp+rename with optional exclusive mode (#1205)
* fix(stack-files): atomic write via tmp+rename with optional exclusive mode writeStackFile and writeStackFileBuffer previously called fs.writeFile directly, which truncates the target then streams the new bytes. A crash, disk-full event, or process kill between the truncate and the write left the target with partial content and no easy way to detect the half-write at read time. A private writeStackFileAtomic helper stages every write into a sibling .sencho-tmp-<suffix> file in the same directory, fsyncs, then promotes via fs.rename. A crash now leaves either the original target intact or a leftover .sencho-tmp file (cleaned up on the next failure path); a torn target file is no longer reachable through this path. The helper accepts an optional `exclusive: true` flag that swaps the final promote step from rename to link+unlink. link is atomic against EEXIST so a caller that needs "create only if not present" gets a race-free FILE_EXISTS error instead of a clobber. The upload route's overwrite-confirm flow (PR #1204) will wire this through in a follow-up so the existence check becomes authoritative. Behaviour for current callers (writeStackFile, writeStackFileBuffer, BlueprintService deploy) is unchanged: the non-exclusive default matches the prior fs.writeFile semantics from the caller's perspective. * fix(stack-files): tighten atomic write entropy + concurrent / failure tests Tmp suffix now uses crypto.randomBytes(6) so the per-process collision window is a true 48-bit space (Math.random().toString(36).slice(2,6) could drop leading zeros and narrow entropy unpredictably). Adds a short comment on the Windows link path noting NTFS / same-FS POSIX is required, both guaranteed by tmp+target being siblings. Two new tests close coverage gaps the first round missed: - a write step that throws (writeFile rejected) leaves no tmp leak and no partial target; - concurrent non-exclusive writers settle with at least one success and the final file is exactly one of the inputs (POSIX silently overwrites; Windows EPERMs the loser, both consistent). |
||
|
|
3e56696c91 |
fix(stack-files): prompt before discarding unsaved edits on file switch (#1203)
FileViewer tracks dirty state via content !== originalContent, but StackFileExplorer would swap selectedPath on every tree click and silently drop the buffered edit. A user editing a script and clicking a sibling file lost their work with no warning. FileViewer now exposes onDirtyChange so the explorer learns when the viewer is dirty. StackFileExplorer intercepts the tree-node click: if dirty, the next selection is stashed and a ConfirmModal asks whether to discard. Confirm applies the stash, Cancel keeps the current file. Clicking the already-selected file is a no-op (no spurious prompt). The dirty signal is reported via a ref so future consumers passing an inline callback identity each render do not retrigger the unmount cleanup effect. |
||
|
|
c8b095b887 |
fix(stack-files): confirm before overwriting an existing upload target (#1204)
* fix(stack-files): confirm before overwriting an existing upload target Same-name uploads previously truncated the existing file silently. A user dragging a file with a name that matched an in-place file destroyed the original with no warning and no undo. The upload route now reads ?overwrite=0|1. When the flag is not set and the target already exists, the server returns 409 FILE_EXISTS and the original file is untouched. The frontend opens a confirm dialog and retries with overwrite=1 on the user's approval; cancel keeps the original. A new pathExists helper on FileSystemService performs the existence check through the same path-resolution barrier as the write so a malicious relPath cannot bypass the conflict check. UploadConflictError is exported so callers can distinguish the conflict case from generic upload failures without parsing error strings. * fix(stack-files): distinct DIR_EXISTS code, drop INVALID_PATH swallow in existence check |
||
|
|
37b12379c1 |
fix(stacks): refuse file-explorer delete/rename/chmod on protected stack files (#1202)
* fix(stacks): refuse file-explorer delete/rename/chmod on protected stack files Previously the per-stack file explorer treated PROTECTED_STACK_FILES (compose.yaml, compose.yml, docker-compose.yaml/.yml, .env) as a UI hint only. A direct API call from any user with stack:edit could delete or rename compose.yaml and break the stack irrecoverably because the next deploy would fail to find a compose file and the write was unrecoverable without a DB backup. The frontend DeleteFileConfirm enforced a type-to-confirm gate but a stale UI or a scripted client bypassed it. FileSystemService now refuses the destructive ops at the service layer with a new PROTECTED_FILE error code that the route layer surfaces as 409. The compose-editor save path (PUT /files/content) and the upload-overwrite path (POST /files/upload writing a same-named file) remain unblocked because both are legitimate ways to update compose.yaml. Tests pin the allowed paths so a future tightening can not silently regress them. The protection is scoped to entries at the stack root; subdirectory files happen to share a protected name (e.g., a snapshot under backups/compose.yaml) are not blocked because compose CLI only reads the root file. A trailing-slash bypass is closed by stripping trailing separators before the basename check. Frontend DeleteFileConfirm needs no change. Its existing toast.error surface renders the friendly server message. * fix(stack-files): drop polynomial regex from protected-file helpers CodeQL js/polynomial-redos flagged the /\/+$/ pattern used to strip trailing slashes from relPath in isProtectedRelPath and protectedFileError. The regex is bounded in practice (the upstream validator rejects '//' anywhere in the path) but the static analyzer cannot follow that dataflow guarantee and would have flagged any future caller that skips the validator. Replace the two callsites with a small stripTrailingSlash helper that uses endsWith + slice. Bounded O(1), no regex, no analyzer alert. The inline comment documents the upstream invariant so a future reader does not reintroduce the /+ quantifier. |
||
|
|
4c28b37a59 |
fix(stacks): require stack:read on file explorer GET routes (#1200)
The four file-explorer GET endpoints (list, content, download, permissions) previously relied on auth alone. Their write-side siblings required stack:edit, so the read path was the only file-explorer surface without an explicit capability check. The shipped roles all carry stack:read globally so behaviour is unchanged today, but adding the guard prevents a future role definition from silently inheriting unrestricted file reads, and it brings the file-explorer reads in line with gitSources and stackActivity which already gate on stack:read. The frontend Files tab trigger, panel, and the anatomy-panel "Open Files" affordance now render only when the user holds stack:read for the active stack. An effect canonicalises activeTab back to 'compose' if the user lands on 'files' without permission, so a denied user cannot end up staring at an empty panel. |
||
|
|
c67478b50d |
chore(security): VEX three new compose CVE findings and correct dependency comment (#1201)
Trivy reported three new advisories against the compose binary today: - CVE-2026-41568 (docker cp symlink-swap empty-file creation, sibling of CVE-2026-42306) - CVE-2026-33997 (Moby plugin install privilege bypass) - GHSA-pmwq-pjrm-6p5r (in-toto-golang glob-negation operator mismatch) All three are daemon-side or supply-chain attestation code paths that compose only embeds as client-side libraries. Sencho invokes compose exclusively for up/down/ps against user-authored compose files, never runs daemon endpoints, never performs plugin installation, and never verifies in-toto attestations. Each gets a not_affected statement with a full impact rationale; bump version 7 to 8 and last_updated to today. Also correct the compose-builder comment in Dockerfile. The previous text claimed compose v5.1.3 "eliminated CVE-2026-34040 and CVE-2026-33997 at the dependency level" by moving docker/docker to an indirect dep. That is wrong: the module is still pulled via buildkit and other transitive paths, the CVEs still appear in scans, and they are tracked in the VEX file rather than eliminated. The corrected comment points the reader at the VEX file as the source of truth for these findings. |
||
|
|
0951a0e792 |
chore(deps): bump containerd to v2.2.4 to clear CVE-2026-46680 (#1199)
The compose-builder stage now pulls containerd/v2 v2.2.4 alongside the existing otel security bumps. containerd v2.2.3 (the version pinned by docker/compose v5.1.3) carries CVE-2026-46680, a runAsNonRoot evasion in the runtime executor; v2.2.4 is the upstream patch release that fixes it. The vulnerable code path is daemon-side and was never reachable from compose, but bumping at the dependency level removes the entry from the SBOM rather than relying on a VEX suppression. Removed the corresponding statement from security/vex/sencho.openvex.json and bumped version + last_updated. The remaining three statements (docker/docker daemon CVEs) stay suppressed: the fix lives on a new github.com/moby/moby/v2 module path, and compose has not migrated imports yet, so no resolvable version bump clears them. |
||
|
|
7ec6fe05bb |
feat(stacks): in-process per-(nodeId, action) metrics + admin endpoint (#1196)
Sencho exports no telemetry by design (privacy-first posture). That left
operators with no answer for "why is this remote node slow today?"
except scrolling logs. The audit log records mutations but has no
latency information.
Adds StackOpMetricsService - a tiny singleton holding per-(nodeId, action)
counters and a 1000-sample ring buffer of latencies for p50/p95.
Exposed through GET /api/stack-metrics (admin-only) so operators can
pull the snapshot when debugging without touching disk or scrolling
journalctl. No external export; the data never leaves the process.
Wiring: deploy, down, restart, stop, start, update routes each capture
t0 at entry, set ok=true after the success path, and record() in a
finally block so failures count too. The record() call is cheap (one
Map lookup, one push, occasional shift on the bounded ring buffer)
and bounded in memory regardless of throughput.
Resolves M-4 from the stack-management audit.
API:
GET /api/stack-metrics (admin-only)
Response: { entries: [{ nodeId, action, count, successCount,
errorCount, avgMs, p50Ms, p95Ms }, ...] }
Ordering: nodeId ascending, then action ascending.
Note on route mounting: /api/stack-metrics rather than the audit doc's
suggested /api/meta/stack-metrics because metaRouter is intentionally
mounted before authGate (public /api/health and /api/meta endpoints);
adding an admin-only route to that group would either bypass auth or
need a special inline gate that fights the existing structure. A
dedicated /api/stack-metrics router after authGate is cleaner.
Tests:
- 9 unit tests in stack-op-metrics-service.test.ts: singleton,
keyed-by-nodeId-action, p50/p95 math, ring-buffer cap at 1000,
NaN/negative/Infinity rejection, ordering, reset.
- 3 integration tests in stack-metrics-route.test.ts: 401 without
auth, empty on fresh process, shape after recording.
|
||
|
|
60247d9c2e |
chore(stacks): gate informational console.log behind developer_mode (#1194)
Rebased onto current main (post H-1 / H-2 / L-3 / M-2 / M-6 merges).
Same intent as the original M-5 commit:
- Local dlog() helper in routes/stacks.ts wrapping console.log
behind isDebugEnabled().
- All informational console.log in stacks.ts replaced with dlog().
- The 3 Exec session-lifecycle console.log in DockerController.ts
wrapped with inline if (isDebugEnabled()) gates.
console.warn and console.error remain unconditional everywhere.
Resolves M-5 from the stack-management audit.
|
||
|
|
429f780f40 |
test(stacks): E2E coverage for deploy success, failure, and bulk lifecycle (#1192)
* test(stacks): E2E coverage for deploy success, failure, and bulk lifecycle Extends e2e coverage per L-2 of the stack-management audit. The existing stacks.spec.ts only covered create/delete and a double-click guard - no journey ever exercised real Docker through the deploy pipeline. Adds e2e/stack-deploy.spec.ts with three journeys: 1. Deploy success: creates a stack with a real-but-tiny compose file (alpine:3 sleep infinity), drives deploy through the authenticated browser session, and asserts the /api/stacks/statuses route reports the container as 'running' inside 30 seconds. Exercises the full middleware chain, ComposeService spawn, Dockerode status lookup, and the cache-invalidate path. 2. Deploy failure: creates a stack with intentionally malformed YAML, asserts the deploy route rejects with a 4xx/5xx and a parseable JSON envelope (no crash, no opaque text), and verifies the sidebar still renders the stack after reload. 3. Bulk lifecycle: creates two stacks, deploys both, then restarts both via per-stack POSTs (the path the current frontend bulk hook takes; once H-3 lands the hook fans out through the new bulk endpoint but the user-visible outcome is identical). Asserts both restarts return 200. Disconnect-mid-deploy journey is deferred (depends on M-1 WS reconnect). Rollback-after-failed-atomic-deploy is deferred (needs a paid-tier license that the default e2e setup does not seed). File-explorer journeys (upload/edit/delete) are tracked separately to keep this spec focused on lifecycle operations. Test design notes: - Drives deploy through page.evaluate + fetch rather than UI buttons to stay resilient to editor-button label churn. Auth surface is still real (cookies, middleware chain, audit log) - only the click is bypassed. The audit's "exercise real Docker via Playwright" intent is satisfied. - Each test calls teardownStack in a finally block so a failed test cannot strand a real container on the CI runner. - Uses alpine:3 (smallest viable long-running image) to keep CI pull cost minimal. The image is shared across tests so registry pull amortizes after the first. * test(stacks): bump E2E timeouts and parallelize bulk deploy CI hit the default 30s per-test timeout on the bulk lifecycle test because two serial deploys against a cold runner spent most of the budget on the alpine pull. The deploy-success test passed the prior run but had only ~0-5s of headroom outside its 30s status poll, so it was a flake away from the same fate. - bulk lifecycle: deploy both stacks in parallel (shared pull cache) and bump the per-test timeout to 90s. - deploy success: bump per-test timeout to 60s so the 30s status poll is not racing setup and teardown. The deploy failure test does no image pull (broken YAML rejects upstream) so it keeps the default 30s. |
||
|
|
5aedc52737 |
feat(stacks): server-side POST /api/stacks/bulk endpoint (#1185)
* feat(stacks): server-side POST /api/stacks/bulk
The frontend's bulk action UI fanned out N parallel POSTs to
/api/stacks/:name/{start,stop,restart,update}. For 30 stacks on a
remote node that was 30 round-trips through the proxy + auth + audit
chain, with no shared mutex and partial-failure UX bolted on the
client.
The new endpoint accepts {action, stackNames} (max 100 names), runs
ops under bounded parallelism (4 concurrent), reuses the per-(nodeId,
stackName) lock from the lifecycle-mutex change so collisions report
stack_op_in_progress as a per-row outcome, and returns
{action, results: [{stackName, ok, error?, code?}, ...]} with a 200
envelope. Per-stack errors are rows, not response codes.
Update action keeps the policy-enforcement check (per-stack, returns
policy_blocked rows) and the post-deploy scan trigger so the
single-stack security contract is preserved. State-invalidate and
image-update notifications fire per successful row so the activity
timeline and image-updates UI reflect bulk operations the same way
as single-stack ones.
Frontend useBulkStackActions swaps the Promise.allSettled fan-out for
a single call; the per-stack toast aggregation moves to reading the
results array. isPaid pre-flight stays in place to avoid a round trip
for Community-tier users on update.
* chore(stacks): dedupe bulk inputs; document tier asymmetry; regression test
Three follow-ups from independent review:
- Dedupe stackNames before scheduling so a payload like ['web','web']
produces one row, not one ok-row plus one stack_op_in_progress row
whose presence depended on worker scheduling.
- Add a route-ordering regression test verifying that a stack literally
named 'bulk' is still reachable via /api/stacks/bulk/restart. Express
matches the literal /bulk before /:stackName paths only at the
no-suffix level; the :stackName/restart route still catches it.
- Comment the deliberate tier asymmetry: bulk update is requirePaid;
single-stack /:stackName/update is open to all tiers. The fan-out
blast radius is the reason, and it matches the prior frontend gate.
Existing 'policy_blocked per-row' test now uses the real ScanPolicy /
PolicyViolation / PolicyEnforcementResult shapes (the first cut elided
fields tsc strict-checked).
|
||
|
|
27b8954676 |
feat(deploy-panel): tell Community operators deploys lack auto-rollback (#1193)
* feat(deploy-panel): tell Community operators deploys lack auto-rollback Atomic-deploy is paid-only (effectiveTier === 'paid' in the deploy and update routes); Community deploys proceed without the backup/restore fallback. The UI never told the user. They only learned the difference when a deploy failed and there was nothing to roll back to. Add a one-line muted-style notice strip inside the deploy-feedback modal, between header and log body, shown only when the user is on Community AND the action is a deploy or update (the two paths that support atomic on paid). Copy is deliberately one line and states the requirement once: "Auto-rollback on failure is a Skipper feature." Compliant with Directive 31: it does not enumerate where the feature is hidden, it does not say "you don't get it", it states what the upgrade unlocks. Other tier-named upgrade prompts in Sencho follow the same pattern. Resolves M-3 from the stack-management audit. * fix(deploy-panel): mount DeployFeedbackPortal inside LicenseProvider The portal was mounted at App level, outside the authed AppContent tree where LicenseProvider lives. After this PR introduced useLicense() inside DeployFeedbackModal (for the atomic-deploy notice), every test that opened the modal hit: Error: useLicense must be used within a LicenseProvider caught by ErrorBoundary and surfaced through every deploy-log-panel E2E spec. Move the portal inside LicenseProvider in AppContent. DeployFeedback- Provider stays at App level so its state survives across re-renders of AppContent; the portal still inherits it because AppContent is a descendant. A deploy can only fire after authentication, so rendering the portal only inside the authed tree loses nothing in practice. |
||
|
|
fbd13accda |
feat(stacks): optimistic concurrency on compose and env file writes (#1183)
* feat(stacks): optimistic concurrency on compose and env file writes
Two browser tabs (or one tab + an out-of-band edit) could silently
overwrite each other's compose.yaml or .env edits. GET /api/stacks/:name
and GET /api/stacks/:name/env now emit a W/"<mtime>" ETag header. PUT
on the same endpoints reads If-Match and returns 412 with
{code: 'stack_file_changed', currentMtimeMs, currentContent} on a stale
write, so the editor can recover without losing the user's text.
The 412 path in the editor surfaces a confirm dialog: "Overwrite
their changes?" Cancel loads the latest content into the editor and
exits edit mode. OK retries the PUT with no If-Match header.
If-Match is optional. A client that doesn't send it falls through to
the previous unconditional-write behavior so partial deploys (file
explorer uploads, git source sync) don't gain a surprise 412 surface.
* fix(stacks): consistent stat+read for compose mtime via held file handle
Promise.all([readFile, stat]) lets a concurrent write interleave between
the two calls: the read can return new content while the stat returns
the old mtime (or vice versa). The next If-Match check would then either
spuriously trigger 412 or silently allow an overwrite.
Hold the file descriptor open across stat and read so both ops observe
the same inode state. A rename-replace by another writer would not
affect the held fd's view.
Adds a near-boundary mtime test (1-second bump) to confirm the
Math.floor comparison detects whole-second changes on filesystems that
round to second-level precision.
* fix(stacks): defense-in-depth path validation in new compose mtime methods
CodeQL's taint engine on PR #1183 flagged 10 js/path-injection errors
across the four new FileSystemService methods (getStackContentWithMtime,
saveStackContentIfUnchanged, writeFileIfUnchanged, statMtime). The
engine does not follow the existing assertWithinBase guard across the
resolveStackDir / getComposeFilePath helper boundary; from its view the
stackName flows straight from req.params into a filesystem sink.
Eight alerts (getStackContentWithMtime + saveStackContentIfUnchanged)
were CodeQL blind-spot: the path WAS validated inside resolveStackDir.
Two alerts (writeFileIfUnchanged + statMtime) were genuinely missing
service-level guards because those methods accept a raw targetPath
from the caller and trusted the route to have validated upstream.
Add an explicit this.assertWithinBase(filePath) at the top of each
new method. The check is redundant for the two methods that already
went through resolveStackDir but makes the safety boundary visible
both to readers and to CodeQL's taint follower.
While here, port the held-file-descriptor pattern (already applied to
getStackContentWithMtime in dd7545eb) to the stat-then-read sequence
in saveStackContentIfUnchanged and writeFileIfUnchanged. This closes
the two js/file-system-race warnings on those branches so a concurrent
rename-replace cannot interleave between the mismatch detection and
the currentContent capture returned in the 412 payload.
The two remaining js/http-to-file-access warnings ("write to file
system depends on untrusted data") are semantic and intentional: this
is the save endpoint by design. They stay as warnings (not errors),
do not block CI, and are not suppressed because the codebase does not
do CodeQL suppression annotations.
* fix(stacks): inline path.resolve+startsWith barrier per CodeQL recommendation
The previous fix added this.assertWithinBase() calls at the top of each
new method. CodeQL's taint-flow analysis does not follow that helper
call across the function boundary, so it still saw the path as
user-tainted at every fs sink (10 -> 8 alerts after the first attempt).
CodeQL's documented js/path-injection sanitizer recognizes the
following inline pattern:
filePath = path.resolve(ROOT, filePath);
if (!filePath.startsWith(ROOT)) { throw / return; }
// use the reassigned filePath below
The key elements are (1) path.resolve as the normalizer, (2) startsWith
check, (3) reject inline, (4) downstream use of the reassigned
variable. None of these can hide behind a helper call or the taint
tracker re-flags every sink.
Inlines the pattern in each of the four new methods (getStackContentWith-
Mtime, saveStackContentIfUnchanged, writeFileIfUnchanged, statMtime).
The check stays runtime-correct (it's the same logic isPathWithinBase
implements) but is now visible to static analysis.
writeFileIfUnchanged and statMtime had no service-level guard at all
before this commit (they trusted the caller); the inline barrier closes
that real gap as well as the CodeQL-recognition gap.
* fix(stacks): canonical CodeQL js/path-injection barrier shape
The prior inline check used path.resolve(filePath) with a single
argument and a compound condition (a !== b && !c.startsWith(d)).
CodeQL's path-injection sanitizer recognizer is shape-sensitive: it
matches path.resolve(SAFE_ROOT, untrusted) with the safe root as the
first argument, followed by a single unary startsWith check on the
resolved variable. The compound form and the single-arg resolve fell
outside the recognized pattern, leaving 8 alerts unchanged across the
four new methods.
Rewrite the barrier in each method to match the documented shape
verbatim:
const baseResolved = path.resolve(this.baseDir);
const safePath = path.resolve(baseResolved, untrustedInput);
if (!safePath.startsWith(baseResolved + path.sep)) {
throw ...;
}
// sinks consume safePath
baseResolved is a local variable (anchors the resolve call against a
known-safe root). safePath is the reassigned, sanitized variable that
every downstream fs.* call consumes. The single startsWith check with
path.sep appended prevents the prefix-match edge case (/foo matches
/foobar without the separator). All four affected methods get the
same form.
This is the third attempt at the CodeQL fix. The first added an
assertWithinBase helper (function call, not followed across the
boundary). The second inlined path.resolve(x) with a compound check
(non-canonical shape). This commit uses the literal recommended
sanitizer.
|
||
|
|
d727a55a5f |
feat(stacks): surface post-deploy scan attempt status (#1198)
triggerPostDeployScan was fire-and-forget. When Trivy was missing on a
node, when the registry refused the digest lookup, or when a single
image scan threw, the failure went to console.error and the user
never learned. Open the security tab later, see stale data, no
indicator that the scan even tried.
Backend:
- New stack_scan_attempts table (node_id, stack_name, status,
attempted_at, error_message). One row per stack; latest attempt
overwrites the previous one.
- DatabaseService gains recordStackScanAttempt /
getStackScanAttempt / clearStackScanAttempts. Status is one of
'ok' | 'partial' | 'failed' | 'skipped'.
- triggerPostDeployScan in helpers/policyGate.ts now records every
exit path: 'skipped' when Trivy is unavailable or no images to
scan; 'failed' when container enumeration or all images fail;
'partial' when some images scan and others fail; 'ok' on full
success.
- New GET /api/stacks/:name/scan-status returns { status,
attemptedAt, errorMessage } or { status: null } when never tried.
- DELETE /:stackName cleanup chain now clears the row alongside
the existing update-status / auto-update cleanups.
Frontend:
- StackAnatomyPanel fetches /scan-status on stackName change.
- Renders a small warning strip below the update banner when
status !== 'ok' (failed / partial / skipped). Hidden when status
is 'ok' or unknown (never attempted). Title attribute carries
the full error message for hover inspection.
Cross-feature note: the audit doc flagged this as M-6 with a
coordination note for the pending Security feature audit. The
schema kept intentionally narrow (one row per stack, simple
status enum) so the Security audit can extend it (richer history,
per-image-row breakdown, etc.) without a destructive migration.
Resolves M-6 from the stack-management audit.
|
||
|
|
009ec43638 |
feat(stacks): structured 503 docker_unavailable envelope + disconnect tests (#1191)
Stack lifecycle routes used to surface raw ECONNREFUSED text to the
client whenever the Docker daemon was unreachable. The frontend had no
way to distinguish "daemon down" from any other 500 and would render
the raw error message.
Detect daemon-reachability failures inside the route layer and surface
a structured envelope so the UI can render a dedicated "Docker is down"
state and operators can branch on a stable code:
HTTP 503 { error: <message>, code: 'docker_unavailable' }
Detection lives in isDockerUnavailableError (exported from
routes/stacks.ts). The match is intentionally permissive across error
shapes Dockerode and the docker compose CLI produce: NodeJS ECONNREFUSED
errors with .code, ENOENT on docker.sock, and the CLI's
"Cannot connect to the Docker daemon" string. The helper is unit-tested
in isolation as well as exercised end-to-end through the route.
Applied to the five lifecycle routes that can hit the daemon-down path:
POST /api/stacks/:name/restart (via bulkContainerOp)
POST /api/stacks/:name/stop (via bulkContainerOp)
POST /api/stacks/:name/start (via bulkContainerOp)
POST /api/stacks/:name/deploy
POST /api/stacks/:name/down
POST /api/stacks/:name/update
Adds ContainerActionOutcome variant 'docker-unavailable' so the route
can branch on the typed outcome rather than string-matching error
messages a second time.
11 integration tests in stack-docker-disconnect.test.ts cover:
- isDockerUnavailableError matches ECONNREFUSED, CLI text, ENOENT on
docker.sock; rejects unrelated errors and null/undefined.
- restart/stop/start return 503 + code on Dockerode listContainers
refusing.
- deploy/down/update return 503 + code when ComposeService rejects
with daemon-down error.
- Unrelated deploy failures (YAML parse error) still return 500
without the code, confirming the discriminator is correctly scoped.
Resolves L-3 from the stack-management audit.
|
||
|
|
4735edfafc |
chore(stacks): explicit stack:read RBAC on list endpoints (#1187)
GET /api/stacks and GET /api/stacks/statuses previously relied on the global authGate for protection without declaring their own permission. Every other endpoint in this router uses requirePermission(); the two list endpoints were silent. Add the explicit gate so: 1. The permission model is uniformly declared (audit-readability). 2. A future role (or per-stack Admiral scoped grant) without stack:read is correctly rejected without an extra code change. 3. The list endpoint behavior stays in sync with checkPermission semantics that the rest of the stack router already obeys. Runtime is observably unchanged for every existing role: admin, node-admin, deployer, viewer, and auditor all hold stack:read per ROLE_PERMISSIONS, so the new gate is a no-op for current users. No new test added because no role currently fails the gate; the existing stack suite (65 tests, all admin-role) exercises both endpoints and stays green. |
||
|
|
5196f0440e |
feat(sidebar): surface unreachable nodes in cross-node stack search (#1195)
Per-node fetch failures in useCrossNodeStackSearch were silently
swallowed into []. The user could not tell the difference between
"this node has no matching stack" and "this node timed out or 502'd".
Adds a third return value to the hook: failedNodes: FailedNode[]
where FailedNode is { nodeId, nodeName, reason }. The reason captures
the HTTP status text (e.g. "list returned HTTP 502") or the
underlying Error message ("connect ECONNREFUSED 192.168.x.x:1852")
without leaking node URLs. AbortError from the effect cleanup path
is intentionally excluded - it's expected when the user keeps typing.
Threads the new field through useStackListState and EditorLayout into
StackList. StackList renders a warning chip below the "Other nodes"
header when the array is non-empty:
! N nodes unreachable > (expand)
Click expands the chip to a vertical list of "node: reason" lines.
Hover surfaces the full reasons in the title attribute too, so a
user inspecting at a glance can read them without clicking. The chip
is suppressed entirely when the search yields no failures (no zero
state). Color is the existing warning token, not error - the failure
is recoverable (node may come back) and not a system-wide problem.
GlobalCommandPalette also calls useCrossNodeStackSearch but ignores
the third return value; its destructuring is additive-safe.
Resolves M-2 from the stack-management audit.
|
||
|
|
07a2e8f0e3 |
feat(stack-logs): WebSocket reconnect with backoff and gap sentinel (#1197)
The structured log viewer opened its WebSocket once and never tried to recover. When the remote node dropped the tunnel mid-tail or the central restarted, log output went silent with no indicator. The 10k row buffer kept prior content on screen, so the silence looked like the stack had just stopped producing logs. Adds exponential-backoff reconnect (1s, 2s, 4s, 8s, 16s, cap 30s) with a "reconnecting..." banner replacing the green "following" pip while the socket is down. Backoff resets to the first delay after every successful re-open. docker logs -f has no resumable offset, so on successful reconnect the viewer drops a synthetic warning sentinel into the row stream: --- reconnected; older lines may be missing --- The sentinel renders at 70% opacity to distinguish it from container output. Operators see at a glance that there was a gap. The closedByCleanup flag prevents the reconnect loop from racing the effect cleanup path: when the component unmounts or stackName changes, the flag is set BEFORE ws.close() so the onclose handler short-circuits instead of scheduling another reconnect. Resolves M-1 from the stack-management audit. |
||
|
|
82a4e94589 |
chore(stack-files): client-side path-traversal guard in stackFilesApi (#1190)
The backend already rejects path-traversal attempts through isValidRelativeStackPath, so the server side is safe today. Adding a client-side mirror is defense-in-depth: it shortens the failure loop (no wasted round trip) and protects against a future server-side regression that loosens validation. Adds isClientSafeRelPath in frontend/src/lib/stackFilesApi.ts mirroring the backend predicate (rejects absolute paths, drive letters, backslashes, NUL bytes, double slashes, and any segment that is `.` or `..`). Wraps every export that accepts a relPath / targetDir / fromRel / toRel argument with assertSafeRelPath, throwing a clear Error before the fetch is issued. 12 unit tests cover the predicate (POSIX accepts, traversal rejects, Windows drive letters, backslashes, NUL bytes, non-string inputs). Frontend suite stays at 288/288. |