mirror of
https://github.com/Studio-Saelix/sencho.git
synced 2026-07-26 11:49:16 +00:00
5f1baa7522
* fix: harden deploy/update concurrency and node-targeting safety Release stabilization for deploy/update operational safety. Per-stack operation locking is now global. Background lifecycle paths (scheduler auto stop/down/start/backup/update, webhook execute, Git source auto-deploy, image auto-update, label bulk actions, fleet snapshot redeploy, and mesh redeploy) acquire the per-node, per-stack lock through a new StackOpLockService.runExclusive helper and skip rather than race a manual deploy/update/rollback/backup on the same stack and node. Skips surface honestly (a failed scheduled run, a recorded webhook failure, a per-stack batch result, or a thrown error) instead of a silent no-op. Update readiness and policy-bypass now run against the node captured when the dialog opened, not the live active node, so switching nodes while a dialog is open cannot retarget the update or the bypass retry. Rollback readiness no longer presents a moving-tag or unpinned image as a ready image revert. Restoring files does not revert a moving tag, so those stacks read as partial, and the rollback success message states that the compose and env files were restored. * fix: lock blueprint reconcile against manual ops and correct rollback wording Follow-up to the deploy/update safety hardening, closing two more gaps from a verification pass. BlueprintService.deployLocal and withdrawLocal called ComposeService directly, so blueprint reconciliation could race a manual deploy/update/rollback/backup on an owned stack. Both now run their compose lifecycle call through StackOpLockService.runExclusive and skip (recorded as a failed reconcile, retried on the next cycle) on conflict. The withdraw holds the lock across both the compose down and the directory delete so neither races a manual operation. The runtime rollback messages overstated recovery: a rollback restores the compose and env files and recreates containers, but does not revert an image behind a moving tag. The auto-rollback deploy-progress output, the recovery panel and chip, the failure toasts, and the manual rollback route message now state that the compose and env files were restored, with the matching OpenAPI example and atomic-deployments doc updated. * fix: acquire stack lock before blueprint deploy mutates compose and marker files Local blueprint deploy wrote the compose and marker files and ran the policy assert before acquiring the per-stack lock; the lock only wrapped the deploy itself. A reconcile could therefore rewrite an owned stack's files while a manual deploy/update/rollback/backup was running. The lock now wraps the whole critical section (create, write compose, write marker, policy assert, deploy), so on conflict nothing is written and the reconcile records a failed outcome. Adds a test asserting a deploy under a held lock records failed, writes no marker file, and leaves the manual lock untouched. * fix: make remote blueprint apply atomic under the receiving node's stack lock Remote blueprint deploy wrote the compose and marker files to the target node via separate HTTP calls and only locked on the final deploy, so the file writes could race a manual operation on that node. A node's operation lock is process-local and cannot be held by the hub across HTTP calls, so the locked create/write/deploy now runs on the receiving node. The locked critical section is extracted into BlueprintService.applyLocalUnderLock and exposed via POST /api/blueprints/apply-local. The hub posts the blueprint to that endpoint in one call; the receiving node runs create + write compose+marker + deploy under its own per-stack lock. Older nodes without the route answer 404 and fall back to the legacy multi-call flow. The endpoint is gated by paid tier and the same per-stack stack:edit and stack:deploy permissions as the PUT-compose + deploy it bundles, validates the stack name, compose size, and marker structure, and returns 409 on a lock conflict without writing anything. Adds tests for the atomic single-call path, the 404 legacy fallback, the 409 lock-conflict mapping, the route validation and permission paths, and the write-compose-then-marker-then-deploy ordering of the shared locked apply. * fix(deps): bump undici to 7.28.0 to clear high-severity advisory The frontend CI npm audit gate (--audit-level=high) failed on a transitive undici 7.25.0 (a dev-only dependency via jsdom): TLS certificate validation bypass (GHSA-vmh5-mc38-953g) and cross-user cache information disclosure (GHSA-pr7r-676h-xcf6). Bumping undici within jsdom's existing ^7.25.0 range to 7.28.0 clears the high-severity advisory and unblocks the frontend job. Lockfile only; no direct dependency or source change.
372 lines
17 KiB
TypeScript
372 lines
17 KiB
TypeScript
import crypto from 'crypto';
|
|
import { ComposeService } from './ComposeService';
|
|
import { StackOpLockService, stackOpSkipMessage, type StackOpAction } from './StackOpLockService';
|
|
import { DatabaseService, type Webhook } from './DatabaseService';
|
|
import { FileSystemService } from './FileSystemService';
|
|
import { GitSourceService } from './GitSourceService';
|
|
import { HealthGateService } from './HealthGateService';
|
|
import { LicenseService } from './LicenseService';
|
|
import { PROXY_TIER_HEADER } from './license-headers';
|
|
import { NodeRegistry } from './NodeRegistry';
|
|
import { getErrorMessage } from '../utils/errors';
|
|
import { redactSensitiveText } from '../utils/safeLog';
|
|
import { isValidStackName } from '../utils/validation';
|
|
import { assertPolicyGateAllows, buildSystemPolicyGateOptions } from '../helpers/policyGate';
|
|
|
|
type ExecutionResult = { success: boolean; error?: string; duration_ms: number };
|
|
type ExecutionStatus = 'success' | 'failure';
|
|
|
|
const REMOTE_WEBHOOK_REQUEST_TIMEOUT_MS = 30_000;
|
|
|
|
// Maps a webhook lifecycle action to the per-stack lock action. 'pull' updates,
|
|
// so it locks as 'update'; 'git-pull' is excluded (it locks inside GitSourceService).
|
|
const WEBHOOK_LOCK_ACTION: Record<string, StackOpAction | undefined> = {
|
|
deploy: 'deploy',
|
|
restart: 'restart',
|
|
stop: 'stop',
|
|
start: 'start',
|
|
pull: 'update',
|
|
};
|
|
|
|
export class WebhookService {
|
|
private static instance: WebhookService;
|
|
// Stable per-process decoy secret used to keep HMAC work non-skippable on
|
|
// reject paths (unknown webhook id, disabled, etc.). Never accepts a
|
|
// signature: the trigger handler decides the final 202 / 404 outcome from
|
|
// independent conditions and only consults the HMAC result when every
|
|
// other check has already passed.
|
|
private static decoySecret: string | null = null;
|
|
|
|
public static getInstance(): WebhookService {
|
|
if (!WebhookService.instance) {
|
|
WebhookService.instance = new WebhookService();
|
|
}
|
|
return WebhookService.instance;
|
|
}
|
|
|
|
public static getDecoySecret(): string {
|
|
if (!WebhookService.decoySecret) {
|
|
WebhookService.decoySecret = crypto.randomBytes(32).toString('hex');
|
|
}
|
|
return WebhookService.decoySecret;
|
|
}
|
|
|
|
public generateSecret(): string {
|
|
return crypto.randomBytes(32).toString('hex');
|
|
}
|
|
|
|
public validateSignature(payload: string, secret: string, signature: string): boolean {
|
|
// Always perform HMAC and timingSafeEqual against a fixed-length
|
|
// buffer so the wall-clock cost of this call is independent of the
|
|
// shape of the input signature. Without this, the early-return paths
|
|
// (missing header, wrong prefix, malformed hex) would skip the HMAC
|
|
// and let an attacker distinguish those cases from a real-shape
|
|
// wrong-secret case through repeated near-rate-limit probes with a
|
|
// large attacker-controlled body. Timing now depends only on the
|
|
// size of `payload`, which the attacker already controls and which
|
|
// does not reveal anything about the webhook id.
|
|
const expected = crypto.createHmac('sha256', secret).update(payload).digest();
|
|
const provided = Buffer.alloc(32);
|
|
let formatOk = false;
|
|
|
|
const parts = signature.split('=');
|
|
if (parts.length === 2 && parts[0] === 'sha256' && /^[0-9a-fA-F]{64}$/.test(parts[1])) {
|
|
provided.write(parts[1], 'hex');
|
|
formatOk = true;
|
|
}
|
|
|
|
const sigEq = crypto.timingSafeEqual(expected, provided);
|
|
return formatOk && sigEq;
|
|
}
|
|
|
|
public async gitSourceExists(stackName: string, nodeId: number): Promise<boolean> {
|
|
const node = NodeRegistry.getInstance().getNode(nodeId);
|
|
if (!node) return false;
|
|
if (node.type !== 'remote') return GitSourceService.getInstance().get(stackName) !== undefined;
|
|
|
|
const response = await this.remoteStackRequest(nodeId, stackName, 'git-source', 'GET');
|
|
return response.ok;
|
|
}
|
|
|
|
public async execute(
|
|
webhook: Webhook,
|
|
action: string,
|
|
triggerSource: string | null,
|
|
atomic?: boolean,
|
|
): Promise<ExecutionResult> {
|
|
if (webhook.id === undefined) {
|
|
throw new Error('Webhook must be loaded from the database before execution');
|
|
}
|
|
const webhookId = webhook.id;
|
|
|
|
const nodeId = webhook.node_id || NodeRegistry.getInstance().getDefaultNodeId();
|
|
const node = NodeRegistry.getInstance().getNode(nodeId);
|
|
if (!node) {
|
|
const error = `Node for webhook "${webhook.name}" was not found`;
|
|
this.recordExecution(webhookId, action, 'failure', triggerSource, 0, error);
|
|
return { success: false, error, duration_ms: 0 };
|
|
}
|
|
|
|
if (node.type === 'remote') {
|
|
return this.executeRemote(webhookId, nodeId, webhook.stack_name, action, triggerSource, atomic);
|
|
}
|
|
|
|
return this.executeLocal(webhookId, nodeId, webhook.stack_name, action, triggerSource, atomic);
|
|
}
|
|
|
|
public maskSecret(secret: string): string {
|
|
if (secret.length <= 8) return '********';
|
|
return '********' + secret.slice(-4);
|
|
}
|
|
|
|
private async executeLocal(
|
|
webhookId: number,
|
|
nodeId: number,
|
|
stackName: string,
|
|
action: string,
|
|
triggerSource: string | null,
|
|
atomic?: boolean,
|
|
): Promise<ExecutionResult> {
|
|
const stacks = await FileSystemService.getInstance(nodeId).getStacks();
|
|
if (!stacks.includes(stackName)) {
|
|
const error = `Stack "${stackName}" not found`;
|
|
this.recordExecution(webhookId, action, 'failure', triggerSource, 0, error);
|
|
return { success: false, error, duration_ms: 0 };
|
|
}
|
|
|
|
const startTime = Date.now();
|
|
try {
|
|
// git-pull pulls then deploys through GitSourceService, which holds
|
|
// the per-stack lock itself; locking here too would self-conflict.
|
|
if (action === 'git-pull') {
|
|
return this.executeLocalGitPull(webhookId, stackName, action, triggerSource, startTime);
|
|
}
|
|
const lockAction = WEBHOOK_LOCK_ACTION[action];
|
|
if (!lockAction) throw new Error(`Unknown action: ${action}`);
|
|
const compose = ComposeService.getInstance(nodeId);
|
|
// Run the lifecycle op under the per-stack lock so a webhook cannot
|
|
// race a manual deploy/update/rollback/backup on the same stack.
|
|
const lock = await StackOpLockService.getInstance().runExclusive(
|
|
nodeId, stackName, lockAction, 'system',
|
|
async () => {
|
|
switch (action) {
|
|
case 'deploy':
|
|
await assertPolicyGateAllows(
|
|
stackName,
|
|
nodeId,
|
|
buildSystemPolicyGateOptions('webhook', { auditPath: `/api/webhooks/${webhookId}/execute` }),
|
|
);
|
|
await compose.deployStack(stackName, undefined, atomic);
|
|
HealthGateService.getInstance().begin(nodeId, stackName, 'deploy', 'system:webhook');
|
|
break;
|
|
case 'restart':
|
|
await compose.runCommand(stackName, 'restart');
|
|
break;
|
|
case 'stop':
|
|
await compose.runCommand(stackName, 'stop');
|
|
break;
|
|
case 'start':
|
|
await compose.runCommand(stackName, 'start');
|
|
break;
|
|
case 'pull':
|
|
await assertPolicyGateAllows(
|
|
stackName,
|
|
nodeId,
|
|
buildSystemPolicyGateOptions('webhook', { auditPath: `/api/webhooks/${webhookId}/execute` }),
|
|
);
|
|
await compose.updateStack(stackName, undefined, atomic);
|
|
HealthGateService.getInstance().begin(nodeId, stackName, 'update', 'system:webhook');
|
|
break;
|
|
}
|
|
},
|
|
);
|
|
|
|
const durationMs = Date.now() - startTime;
|
|
if (!lock.ran) {
|
|
const error = stackOpSkipMessage(stackName, lock.existing.action);
|
|
this.recordExecution(webhookId, action, 'failure', triggerSource, durationMs, error);
|
|
return { success: false, error, duration_ms: durationMs };
|
|
}
|
|
this.recordExecution(webhookId, action, 'success', triggerSource, durationMs, null);
|
|
return { success: true, duration_ms: durationMs };
|
|
} catch (err) {
|
|
const durationMs = Date.now() - startTime;
|
|
const error = getErrorMessage(err, 'Unknown error');
|
|
this.recordExecution(webhookId, action, 'failure', triggerSource, durationMs, error);
|
|
return { success: false, error, duration_ms: durationMs };
|
|
}
|
|
}
|
|
|
|
private async executeLocalGitPull(
|
|
webhookId: number,
|
|
stackName: string,
|
|
action: string,
|
|
triggerSource: string | null,
|
|
startTime: number,
|
|
): Promise<ExecutionResult> {
|
|
const result = await GitSourceService.getInstance().handleWebhookPull(stackName);
|
|
const durationMs = Date.now() - startTime;
|
|
if (result.status === 'error') {
|
|
this.recordExecution(webhookId, action, 'failure', triggerSource, durationMs, result.message);
|
|
return { success: false, error: result.message, duration_ms: durationMs };
|
|
}
|
|
|
|
const skipped = result.status === 'skipped';
|
|
// A debounced pull is rate-limited, not failed (the route answers 202
|
|
// Accepted). Record it as a success carrying the debounce note so it
|
|
// does not pollute the webhook's failure history.
|
|
this.recordExecution(
|
|
webhookId,
|
|
action,
|
|
'success',
|
|
triggerSource,
|
|
durationMs,
|
|
skipped ? result.message : null,
|
|
);
|
|
return { success: true, error: undefined, duration_ms: durationMs };
|
|
}
|
|
|
|
private async executeRemote(
|
|
webhookId: number,
|
|
nodeId: number,
|
|
stackName: string,
|
|
action: string,
|
|
triggerSource: string | null,
|
|
atomic?: boolean,
|
|
): Promise<ExecutionResult> {
|
|
const startTime = Date.now();
|
|
try {
|
|
const endpoint = action === 'git-pull'
|
|
? 'git-source/webhook-pull'
|
|
: action === 'pull'
|
|
? 'update'
|
|
: action;
|
|
const body = atomic === undefined ? undefined : { atomic };
|
|
const response = await this.remoteStackRequest(nodeId, stackName, endpoint, 'POST', body);
|
|
const durationMs = Date.now() - startTime;
|
|
const payload = await response.json().catch(() => ({})) as { error?: string; message?: string; status?: string };
|
|
|
|
if (!response.ok || payload.status === 'error') {
|
|
const error = payload.error || payload.message || `Remote ${action} failed with status ${response.status}`;
|
|
this.recordExecution(webhookId, action, 'failure', triggerSource, durationMs, error);
|
|
return { success: false, error, duration_ms: durationMs };
|
|
}
|
|
|
|
// A debounced remote pull comes back 202 with status "skipped": it was
|
|
// accepted and rate-limited, not failed. Record it as a success with
|
|
// the debounce note rather than failure noise.
|
|
const skipped = payload.status === 'skipped';
|
|
this.recordExecution(webhookId, action, 'success', triggerSource, durationMs, skipped ? (payload.message ?? null) : null);
|
|
return { success: true, duration_ms: durationMs };
|
|
} catch (err) {
|
|
const durationMs = Date.now() - startTime;
|
|
const error = getErrorMessage(err, 'Remote node operation failed');
|
|
this.recordExecution(webhookId, action, 'failure', triggerSource, durationMs, error);
|
|
return { success: false, error, duration_ms: durationMs };
|
|
}
|
|
}
|
|
|
|
private async remoteStackRequest(
|
|
nodeId: number,
|
|
stackName: string,
|
|
endpoint: string,
|
|
method: 'GET' | 'POST',
|
|
body?: unknown,
|
|
): Promise<Response> {
|
|
const target = NodeRegistry.getInstance().getProxyTarget(nodeId);
|
|
if (!target) throw new Error('Remote node is unreachable or not configured');
|
|
|
|
const headers: Record<string, string> = { 'Content-Type': 'application/json' };
|
|
if (target.apiToken) headers.Authorization = `Bearer ${target.apiToken}`;
|
|
|
|
const licenseHeaders = LicenseService.getInstance().getProxyHeaders();
|
|
headers[PROXY_TIER_HEADER] = licenseHeaders.tier;
|
|
|
|
const controller = new AbortController();
|
|
const timer = setTimeout(() => controller.abort(), REMOTE_WEBHOOK_REQUEST_TIMEOUT_MS);
|
|
try {
|
|
// nodeId selects a server-controlled entry from the registry
|
|
// (CodeQL GOOD pattern: user input maps to known values, not concatenated into the URL).
|
|
const targetBase = new URL(target.apiUrl);
|
|
const protocol = targetBase.protocol;
|
|
const host = targetBase.host;
|
|
const hostname = targetBase.hostname;
|
|
|
|
// Verify the hostname is in the configured-node allow-list.
|
|
const allowedHosts = DatabaseService.getInstance().getNodes()
|
|
.filter(n => n.api_url)
|
|
.map(n => new URL(n.api_url!).hostname);
|
|
if (!allowedHosts.includes(hostname)) {
|
|
throw new Error('Remote node hostname is not a configured node');
|
|
}
|
|
|
|
// Restrict protocol to http/https (prevents file://, ftp://, etc.).
|
|
if (protocol !== 'http:' && protocol !== 'https:') {
|
|
throw new Error('Remote node URL must use http:// or https://');
|
|
}
|
|
|
|
// Validate path components to prevent traversal.
|
|
if (!isValidStackName(stackName)) {
|
|
throw new Error('Invalid stack name');
|
|
}
|
|
if (!/^[a-z][a-z0-9/-]*$/.test(endpoint) || endpoint.includes('..')) {
|
|
throw new Error('Invalid endpoint');
|
|
}
|
|
|
|
// Build URL from validated, server-controlled components.
|
|
const url = `${protocol}//${host}/api/stacks/${encodeURIComponent(stackName)}/${endpoint}`;
|
|
return await fetch(url, {
|
|
method,
|
|
headers,
|
|
body: method === 'GET' || body === undefined ? undefined : JSON.stringify(body),
|
|
signal: controller.signal,
|
|
});
|
|
} catch (err) {
|
|
if (controller.signal.aborted) {
|
|
throw new Error('Remote node request timed out', { cause: err });
|
|
}
|
|
throw err;
|
|
} finally {
|
|
clearTimeout(timer);
|
|
}
|
|
}
|
|
|
|
private recordExecution(
|
|
webhookId: number,
|
|
action: string,
|
|
status: ExecutionStatus,
|
|
triggerSource: string | null,
|
|
durationMs: number,
|
|
error: string | null,
|
|
): void {
|
|
// Execution history is readable in the UI; scrub bearer tokens,
|
|
// JWTs, URL credentials, and homedir paths before persisting so a
|
|
// compose / remote-node error surfacing on the dashboard cannot leak
|
|
// operator secrets or infrastructure details.
|
|
const safeError = error === null ? null : redactSensitiveText(error);
|
|
try {
|
|
DatabaseService.getInstance().addWebhookExecution({
|
|
webhook_id: webhookId,
|
|
action,
|
|
status,
|
|
trigger_source: triggerSource,
|
|
duration_ms: durationMs,
|
|
error: safeError,
|
|
executed_at: Date.now(),
|
|
});
|
|
} catch (err) {
|
|
// The webhook_executions table has ON DELETE CASCADE on webhook_id,
|
|
// so a delete that races an in-flight execution removes the parent
|
|
// row and any insert here fails the FK constraint. Swallow that
|
|
// race: the trigger already returned 202 and the action either
|
|
// ran or failed before reaching this point. Other write errors
|
|
// are still worth logging as warnings so a structural problem
|
|
// does not go silent.
|
|
console.warn(
|
|
`[Webhooks] Could not record execution for webhook ${webhookId} ` +
|
|
`(parent webhook may have been deleted mid-flight): ${getErrorMessage(err, 'Unknown error')}`,
|
|
);
|
|
}
|
|
}
|
|
}
|