Files
sencho/backend/src/services/fleetSyncConstants.ts
T
Anso 3b650523c1 Audit-hardening pass for secret and misconfiguration scanning (#977)
* fix(security): dedupe concurrent compose-stack scans

Track stack scans in scanningImages keyed stack:<nodeId>:<stackName>.
The /scan/stack route returns 409 when an in-flight scan exists, and
the service-side check is the real correctness barrier (the route
pre-check is a fast-path optimization that mirrors scanImage). The
dedup key release lives in a try/finally so failed scans free the
slot for retry.

Why: scanComposeStack had no equivalent of scanImage's scanningImages
guard, so two simultaneous calls for the same stack would both run
trivy config, both insert a vulnerability_scans row, and double-
process the result.

* feat(security): acknowledge misconfig findings

Adds a parallel acknowledgement system for Trivy misconfig findings
that mirrors cve_suppressions: a new misconfig_acknowledgements table,
read-time enrichment via the new misconfig-ack-filter utility, REST
CRUD endpoints, fleet-sync replication from control to replicas, a
Settings panel, and an Acknowledge button on the Misconfigs tab.

Schema and behavior parity with cve_suppressions:
  - UNIQUE(rule_id, COALESCE(stack_pattern, '')) so fleet-wide acks
    collide as expected
  - blockIfReplica on every write
  - Audit-log entries name the scope (rule_id, stack_pattern) but
    never the reason text
  - replicated_from_control flag controls UI delete affordance and
    drives clearReplicatedRows on demote/reanchor
  - Validators reused: validateStackPatternForRedos for glob safety,
    sanitizeForLog for log fragments

SARIF export emits an external/accepted suppression entry per
acknowledged misconfig, matching the CVE pattern.

Per-row Acknowledge dialog prefills stack_pattern with the scan's
stack_context so the default scope is "rule + this stack only" and an
operator must broaden explicitly.

Tests: misconfig-ack-filter (15) and misconfig-ack-routes (23)
including the duplicate-409 case for both pinned and fleet-wide acks.

* fix(security): reap orphaned trivy tmp dirs at startup

When the buildEnv path writes a per-scan DOCKER_CONFIG dir under
os.tmpdir() and the process crashes before the finally block runs,
the dir leaks. Mirrors GitSourceService.sweepStaleTempDirs:
exported sweepStaleTrivyTempDirs is fire-and-forget at boot,
removes prefix-matching dirs older than 1 hour, swallows
permission/race failures, logs a single line if any were reaped.

* perf(security): emit per-batch summary for scanAllNodeImages

Adds one diag() line at the end of scanAllNodeImages summarising
unique image count, scanned, skipped, failed, violation count, and
elapsed time. Per-image diag inside scanImage stays useful for
debugging individual scans; the summary gives operators a single
fleet-level checkpoint when developer_mode is on.

* perf(security): cap SARIF export at 5000 findings per type

Replace the unbounded fetchAllPages walk on /scans/:id/sarif with a
hard limit of 5000 findings per type. When any type trips the cap,
emit run-level properties.truncated=true plus row_limit and per-type
totals so downstream tooling can flag the export as partial.
Console-warns for ops visibility.

A scan with 50k vulns previously streamed every row into memory
before serialising; the cap bounds memory and serialisation time at
the cost of completeness on pathological scans.

* docs(env): document TRIVY_BIN host-binary override

The env var is honored by TrivyService.detectTrivy as a fallback when
no managed install is present, but it was undocumented in
.env.example. Adds the var with a comment explaining precedence
(managed > TRIVY_BIN > PATH).

* test(security): cover scanComposeStack failure modes

Two new cases drive the existing try/catch through real failure
paths:
  - Malformed Trivy stdout: row flips to status='failed' with the
    parser error preserved on `error`.
  - execFile rejection: row flips to status='failed' with a string
    error message.

Pairs with the existing dedup tests so the failure path now also
verifies the scan row state, not just the thrown exception.

* test(e2e): security scanner + misconfig acknowledgement flow

Seven Playwright tests covering the scanner UI and the new
acknowledgement system end-to-end:
  - Trivy availability gate (skips suite when binary absent so CI
    without Trivy can opt out via E2E_SKIP_TRIVY=1)
  - Stack config scan completes and records misconfig findings
  - Concurrent stack scan returns 409 from the dedup gate
  - Misconfig ack POST creates and lists on Settings
  - Duplicate (rule_id, stack_pattern) returns 409
  - Malformed rule_id (shell metacharacters) returns 400
  - Misconfigs tab renders against a real stack scan

Tests drive the API for behaviour assertions and the UI only for
shell-rendering checks; the visual snapshot suite owns screenshots.

* docs(features): add misconfig acknowledgement workflow and SARIF cap

Refreshes vulnerability-scanning.mdx with:
  - Misconfig acknowledgements section covering the per-row dialog,
    Settings panel, scope/matching rules, and SARIF emission
  - Tier table row for the new feature
  - SARIF section note on the 5000 row-per-type cap and the
    properties.truncated marker for partial exports
  - Troubleshooting entries: SARIF cap, hidden Acknowledge button,
    findings resurfacing after delete, Trivy DB phone-home, and
    409 on concurrent compose-stack scans

* fix(ci): clear backend lint and CodeQL alerts

- Remove the dead fetchAllPages helper in routes/security.ts. It lost
  its callers when the SARIF endpoint switched to direct paged reads
  for the truncation cap. ESLint flagged it as unused.
- Switch the trivy-tmp-cleanup test helper to fs.mkdtempSync. Building
  paths under os.tmpdir() with predictable names tripped CodeQL's
  js/insecure-temporary-file rule (high severity), which warns about
  symlink-pre-creation attacks even in test code. mkdtempSync appends
  a process-random suffix and creates the dir atomically; the
  sencho-trivy- prefix is preserved so the production sweep still
  matches the test fixtures.
2026-05-07 19:23:11 -04:00

83 lines
3.3 KiB
TypeScript

/**
* Shared constants for the Fleet Sync wire protocol.
*
* The control instance pushes scan_policies and cve_suppressions to remote
* nodes via POST /api/fleet/sync/:resource. The payload schema is defined
* here so both the sender (FleetSyncService) and the receiver (routes/fleet)
* agree on limits and field names.
*/
/**
* Maximum number of rows accepted per sync push.
*
* Enforced on both ends:
* - Sender: `FleetSyncService.loadResource` truncates at this cap and
* emits a warning notification when it triggers. Operators with
* >5000 policies on the control are exceedingly rare; truncating is
* safer than failing every push.
* - Receiver: `POST /api/fleet/sync/:resource` rejects payloads above
* this cap with HTTP 413.
*/
export const MAX_SYNC_ROWS = 5000;
/**
* Maximum body size for POST /api/fleet/sync/:resource. Sized to comfortably
* fit MAX_SYNC_ROWS rows of either resource at the per-field length caps
* enforced by the row validators (~1 KB per row worst-case = ~5 MB).
*
* The global JSON body limit stays at 100 KB; only this one route allows
* larger bodies. See middleware/jsonParser.ts for the dispatch logic.
*/
export const SYNC_BODY_LIMIT = '5mb';
/**
* Path prefix that the JSON body parser uses to dispatch to the larger
* limit. Kept as a constant so any path change updates both the parser
* and the route mount.
*
* CAUTION: prefix match, not exact route. Any new route under /api/fleet/sync/
* inherits the elevated body limit. If a future route under this prefix should
* keep the standard 100 KB cap, narrow the dispatch in middleware/jsonParser.ts
* to method+exact-path instead of prefix.
*/
export const SYNC_PATH_PREFIX = '/api/fleet/sync/';
/** How long to suppress repeat truncation alerts after one fires. */
export const TRUNCATION_ALERT_COOLDOWN_MS = 6 * 60 * 60 * 1000;
/** How far back the retry service looks for failed sync targets. */
export const RETRY_MAX_AGE_MS = 24 * 60 * 60 * 1000;
/**
* Failure-window threshold for the retry-service stale-target notification.
* A previously-working node whose `last_failure_at - last_success_at` exceeds
* this triggers one warning per cooldown.
*/
export const STALE_THRESHOLD_MS = 60 * 60 * 1000;
/**
* Resource enum kept here so the state-key helpers below can type-check
* their arguments without a cycle through FleetSyncService. The ordering
* below mirrors `FLEET_RESOURCES` in FleetSyncService.
*/
export type FleetResource = 'scan_policies' | 'cve_suppressions' | 'misconfig_acknowledgements';
/**
* `system_state` keys read or written by Fleet Sync. Centralized so a typo
* cannot silently bypass a stale-push check or a cooldown gate.
*/
export const SYNC_STATE_KEYS = {
fleetRole: 'fleet_role',
fleetSelfIdentity: 'fleet_self_identity',
fleetControlIdentity: 'fleet_control_identity',
receivedPushedAt: (resource: FleetResource): string => `received_pushed_at:${resource}`,
truncationAlertAt: (resource: FleetResource): string => `fleet_sync_truncation_alert_at:${resource}`,
} as const;
/** Structured error codes returned by the receive endpoint. */
export const SYNC_ERROR_CODES = {
staleSyncPush: 'STALE_SYNC_PUSH',
payloadTooLarge: 'SYNC_PAYLOAD_TOO_LARGE',
controlIdentityMismatch: 'CONTROL_IDENTITY_MISMATCH',
} as const;