Audit-hardening pass for secret and misconfiguration scanning (#977)

* fix(security): dedupe concurrent compose-stack scans

Track stack scans in scanningImages keyed stack:<nodeId>:<stackName>.
The /scan/stack route returns 409 when an in-flight scan exists, and
the service-side check is the real correctness barrier (the route
pre-check is a fast-path optimization that mirrors scanImage). The
dedup key release lives in a try/finally so failed scans free the
slot for retry.

Why: scanComposeStack had no equivalent of scanImage's scanningImages
guard, so two simultaneous calls for the same stack would both run
trivy config, both insert a vulnerability_scans row, and double-
process the result.

* feat(security): acknowledge misconfig findings

Adds a parallel acknowledgement system for Trivy misconfig findings
that mirrors cve_suppressions: a new misconfig_acknowledgements table,
read-time enrichment via the new misconfig-ack-filter utility, REST
CRUD endpoints, fleet-sync replication from control to replicas, a
Settings panel, and an Acknowledge button on the Misconfigs tab.

Schema and behavior parity with cve_suppressions:
  - UNIQUE(rule_id, COALESCE(stack_pattern, '')) so fleet-wide acks
    collide as expected
  - blockIfReplica on every write
  - Audit-log entries name the scope (rule_id, stack_pattern) but
    never the reason text
  - replicated_from_control flag controls UI delete affordance and
    drives clearReplicatedRows on demote/reanchor
  - Validators reused: validateStackPatternForRedos for glob safety,
    sanitizeForLog for log fragments

SARIF export emits an external/accepted suppression entry per
acknowledged misconfig, matching the CVE pattern.

Per-row Acknowledge dialog prefills stack_pattern with the scan's
stack_context so the default scope is "rule + this stack only" and an
operator must broaden explicitly.

Tests: misconfig-ack-filter (15) and misconfig-ack-routes (23)
including the duplicate-409 case for both pinned and fleet-wide acks.

* fix(security): reap orphaned trivy tmp dirs at startup

When the buildEnv path writes a per-scan DOCKER_CONFIG dir under
os.tmpdir() and the process crashes before the finally block runs,
the dir leaks. Mirrors GitSourceService.sweepStaleTempDirs:
exported sweepStaleTrivyTempDirs is fire-and-forget at boot,
removes prefix-matching dirs older than 1 hour, swallows
permission/race failures, logs a single line if any were reaped.

* perf(security): emit per-batch summary for scanAllNodeImages

Adds one diag() line at the end of scanAllNodeImages summarising
unique image count, scanned, skipped, failed, violation count, and
elapsed time. Per-image diag inside scanImage stays useful for
debugging individual scans; the summary gives operators a single
fleet-level checkpoint when developer_mode is on.

* perf(security): cap SARIF export at 5000 findings per type

Replace the unbounded fetchAllPages walk on /scans/:id/sarif with a
hard limit of 5000 findings per type. When any type trips the cap,
emit run-level properties.truncated=true plus row_limit and per-type
totals so downstream tooling can flag the export as partial.
Console-warns for ops visibility.

A scan with 50k vulns previously streamed every row into memory
before serialising; the cap bounds memory and serialisation time at
the cost of completeness on pathological scans.

* docs(env): document TRIVY_BIN host-binary override

The env var is honored by TrivyService.detectTrivy as a fallback when
no managed install is present, but it was undocumented in
.env.example. Adds the var with a comment explaining precedence
(managed > TRIVY_BIN > PATH).

* test(security): cover scanComposeStack failure modes

Two new cases drive the existing try/catch through real failure
paths:
  - Malformed Trivy stdout: row flips to status='failed' with the
    parser error preserved on `error`.
  - execFile rejection: row flips to status='failed' with a string
    error message.

Pairs with the existing dedup tests so the failure path now also
verifies the scan row state, not just the thrown exception.

* test(e2e): security scanner + misconfig acknowledgement flow

Seven Playwright tests covering the scanner UI and the new
acknowledgement system end-to-end:
  - Trivy availability gate (skips suite when binary absent so CI
    without Trivy can opt out via E2E_SKIP_TRIVY=1)
  - Stack config scan completes and records misconfig findings
  - Concurrent stack scan returns 409 from the dedup gate
  - Misconfig ack POST creates and lists on Settings
  - Duplicate (rule_id, stack_pattern) returns 409
  - Malformed rule_id (shell metacharacters) returns 400
  - Misconfigs tab renders against a real stack scan

Tests drive the API for behaviour assertions and the UI only for
shell-rendering checks; the visual snapshot suite owns screenshots.

* docs(features): add misconfig acknowledgement workflow and SARIF cap

Refreshes vulnerability-scanning.mdx with:
  - Misconfig acknowledgements section covering the per-row dialog,
    Settings panel, scope/matching rules, and SARIF emission
  - Tier table row for the new feature
  - SARIF section note on the 5000 row-per-type cap and the
    properties.truncated marker for partial exports
  - Troubleshooting entries: SARIF cap, hidden Acknowledge button,
    findings resurfacing after delete, Trivy DB phone-home, and
    409 on concurrent compose-stack scans

* fix(ci): clear backend lint and CodeQL alerts

- Remove the dead fetchAllPages helper in routes/security.ts. It lost
  its callers when the SARIF endpoint switched to direct paged reads
  for the truncation cap. ESLint flagged it as unused.
- Switch the trivy-tmp-cleanup test helper to fs.mkdtempSync. Building
  paths under os.tmpdir() with predictable names tripped CodeQL's
  js/insecure-temporary-file rule (high severity), which warns about
  symlink-pre-creation attacks even in test code. mkdtempSync appends
  a process-random suffix and creates the dir atomically; the
  sencho-trivy- prefix is preserved so the production sweep still
  matches the test fixtures.
This commit is contained in:
Anso
2026-05-07 19:23:11 -04:00
committed by GitHub
parent 4b1de35dda
commit 3b650523c1
20 changed files with 2376 additions and 162 deletions
+122
View File
@@ -0,0 +1,122 @@
/**
* Read-time misconfiguration acknowledgement filter.
*
* Acknowledgements never modify stored finding rows. They are applied at read
* time so deleting an ack resurfaces findings without rescanning.
*
* An acknowledgement matches a finding when:
* - rule_id equals the finding's rule_id, AND
* - stack_pattern is null OR matches the scan's stack_context (glob), AND
* - expires_at is null OR still in the future.
*
* Mirrors the design of `suppression-filter.ts`. The bucketing pass is shared
* spirit: pre-group by rule_id once so a multi-thousand-finding scan does not
* cross-multiply with the fleet ack list on every render.
*/
import type { MisconfigAcknowledgement } from '../services/DatabaseService';
export interface MisconfigAcknowledgementDecision {
acknowledged: boolean;
acknowledgement_id?: number;
acknowledgement_reason?: string;
}
export interface AcknowledgeableFinding {
rule_id: string;
}
function matchesStackPattern(pattern: string | null, stackContext: string | null): boolean {
// No pattern means fleet-wide; matches any stack including null contexts
// (e.g. image scans where stack_context is null).
if (!pattern) return true;
// Stack-scoped acks against an image scan (no stack_context) cannot match.
if (stackContext === null) return false;
const escaped = pattern.replace(/[.+?^${}()|[\]\\]/g, '\\$&').replace(/\*/g, '.*');
return new RegExp(`^${escaped}$`).test(stackContext);
}
function isActive(ack: MisconfigAcknowledgement, now: number): boolean {
return ack.expires_at === null || ack.expires_at > now;
}
function specificityScore(a: MisconfigAcknowledgement): number {
return a.stack_pattern ? 1 : 0;
}
/**
* Pick the highest-specificity active ack from a candidate bucket already
* filtered to a single rule_id. A stack-scoped ack beats a fleet-wide ack.
*/
function pickFromBucket(
bucket: MisconfigAcknowledgement[],
stackContext: string | null,
now: number,
): MisconfigAcknowledgement | null {
let best: MisconfigAcknowledgement | null = null;
let bestScore = -1;
for (const a of bucket) {
if (!isActive(a, now)) continue;
if (!matchesStackPattern(a.stack_pattern, stackContext)) continue;
const score = specificityScore(a);
if (score > bestScore) {
best = a;
bestScore = score;
}
}
return best;
}
/**
* Find the most specific active acknowledgement matching a single finding.
* For one-shot lookups; prefer applyMisconfigAcknowledgements when enriching
* a list because that path amortizes the bucketing.
*/
export function findMisconfigAcknowledgement(
finding: AcknowledgeableFinding,
stackContext: string | null,
acks: MisconfigAcknowledgement[],
now: number = Date.now(),
): MisconfigAcknowledgement | null {
const bucket: MisconfigAcknowledgement[] = [];
for (const a of acks) {
if (a.rule_id === finding.rule_id) bucket.push(a);
}
if (bucket.length === 0) return null;
return pickFromBucket(bucket, stackContext, now);
}
/**
* Enrich a list of misconfig findings with acknowledgement decisions. Does
* not mutate inputs.
*
* Acks are bucketed by rule_id once before the per-finding scan, so the
* per-finding work is O(matching-rule-acks) rather than O(acks).
*/
export function applyMisconfigAcknowledgements<T extends AcknowledgeableFinding>(
findings: T[],
stackContext: string | null,
acks: MisconfigAcknowledgement[],
now: number = Date.now(),
): Array<T & MisconfigAcknowledgementDecision> {
if (findings.length === 0) return [];
const buckets = new Map<string, MisconfigAcknowledgement[]>();
for (const a of acks) {
const existing = buckets.get(a.rule_id);
if (existing) {
existing.push(a);
} else {
buckets.set(a.rule_id, [a]);
}
}
return findings.map((f) => {
const bucket = buckets.get(f.rule_id);
const match = bucket ? pickFromBucket(bucket, stackContext, now) : null;
if (!match) return { ...f, acknowledged: false };
return {
...f,
acknowledged: true,
acknowledgement_id: match.id,
acknowledgement_reason: match.reason,
};
});
}