Files
sencho/docs/features/scheduled-operations.mdx
T
Anso abee078741 feat(scheduler): add auto_backup, auto_stop, auto_down, auto_start and delete_after_run one-shot mode (#777)
* feat(scheduler): add auto_backup, auto_stop, auto_down, auto_start actions and delete_after_run one-shot mode

Extends the scheduler with four new stack-targeted actions:

- auto_backup: backs up stack compose files and .env using the existing
  FileSystemService.backupStackFiles primitive
- auto_stop: runs compose stop (containers preserved)
- auto_down: runs compose down (containers removed)
- auto_start: runs compose up -d via deployStack (universal start for both
  stopped and down stacks)

Adds delete_after_run boolean column to scheduled_tasks. When enabled,
the task self-deletes after its first successful execution; failures keep
the task so the user can debug and retry.

All four new actions gate at Admiral tier, consistent with restart/snapshot/prune.
Migration is idempotent (maybeAddCol).

* docs(scheduler): update scheduled-operations doc with new lifecycle actions and delete-after-run

Adds the four new actions (Backup Stack Files, Stop Stack, Take Stack Down,
Start Stack) to the action table. Documents the delete-after-run one-shot mode
with its success-only deletion semantics. Adds the Stack Lifecycle Scheduling
section explaining stop-vs-down semantics and the local-execution boundary.
Adds three troubleshooting entries: auto-start on a missing compose folder,
auto-backup single-slot overwrite by design, and one-shot task disappearing
after successful run.

Updates the timeline description from four to five lanes. Refreshes
screenshots to show the new dialog layout with the Lifecycle lane visible.
2026-04-25 16:11:13 -04:00

263 lines
17 KiB
Plaintext
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
---
title: Scheduled Operations
description: Automate recurring Docker operations like stack restarts, lifecycle management, fleet snapshots, and system prunes on a cron schedule.
---
<Note>
Scheduled Operations requires a Sencho **Admiral** license. Skipper users see only the **Auto-update Stack** action; Admiral users see every action.
</Note>
## Overview
Scheduled Operations lets you automate recurring maintenance tasks across your infrastructure. Define a cron schedule, choose an action, and Sencho handles the rest, including a full execution history log so you always know what ran and when.
The view opens on a **Timeline** that plots the next 24 hours of scheduled work across five lanes (Restart, Update, Scan, Prune, Lifecycle) so you can see, at a glance, what is about to fire and when. Toggle to **All tasks** for the full CRUD table.
<Frame>
<img src="/images/scheduled-operations/timeline.png" alt="Schedules timeline showing the next 24 hours across four lanes with a cyan now rail" />
</Frame>
## Timeline view
The timeline is the default view. It shows:
- A hero with the current 24-hour window as a date range, and the **next firing** on the right (time, task name, and relative countdown).
- Five color-coded lanes: **Restart** (cyan), **Update** (green), **Scan** (purple), **Prune** (amber), and **Lifecycle** (blue). Snapshot tasks share the Prune lane; stop, down, start, and backup tasks share the Lifecycle lane.
- One pill per firing within the window, positioned proportionally to the task's next run time. Click any pill to open that task's execution history.
- A vertical cyan **now rail** at the left edge, and six mono time ticks along the bottom axis.
Tasks that fire more than once in the window (e.g. an hourly cron) render a pill for each firing. Disabled tasks do not appear on the timeline.
Toggle to **All tasks** from the header to see every schedule in a table, regardless of whether it fires in the next 24 hours.
## Supported Actions
| Action | Target | Description |
|--------|--------|-------------|
| **Restart Stack** | A specific stack (or specific services within it) on a specific node | Restarts all or selected containers in the stack |
| **Auto-update Stack** | A specific stack on a specific node | Checks each image for updates and recreates the stack if any image has a newer version. See [Auto-Update Readiness](/features/auto-update-policies) for the companion board. Available on Skipper and Admiral. |
| **Fleet Snapshot** | All nodes | Creates a fleet-wide backup of all compose files and `.env` files |
| **System Prune** | The default node | Prunes selected resources, optionally filtered by Docker label |
| **Vulnerability Scan** | All images on a specific node | Runs Trivy against every image on the target node and records the results. Requires Trivy to be installed, see [Installing Trivy](/operations/trivy-setup). Available on Skipper and Admiral. |
| **Backup Stack Files** | A specific stack on a specific node | Backs up the stack's compose file and `.env` to `<DATA_DIR>/backups/<stackName>/`. The most recent backup per stack is kept; each run overwrites the previous one. |
| **Stop Stack** | A specific stack on a specific node | Runs `docker compose stop` — containers are stopped but preserved. Use this for off-hours power saving when you want a fast start later. |
| **Take Stack Down** | A specific stack on a specific node | Runs `docker compose down` — containers are removed. Use this to fully release resources when the stack is not needed for an extended period. |
| **Start Stack** | A specific stack on a specific node | Runs `docker compose up -d`. Works whether the stack was stopped or taken down: if containers exist they are started; if not, they are created and started from the compose file. |
## Creating a Scheduled Task
1. Navigate to the **Schedules** tab in the top navigation bar (visible to Admiral admins).
2. Click **New Schedule**.
3. Fill in the form:
- **Name**: A descriptive label (e.g. "Nightly staging restart").
- **Action**: Choose from the full action list. The form fields below change based on your selection.
- **Node**: (Restart Stack, Auto-update Stack, and Vulnerability Scan) Select the node to run against. For stack actions it determines where the target stack lives; for Vulnerability Scan it determines which node's images are scanned.
- **Stack**: (Restart Stack and Auto-update Stack) Select the stack to target. Becomes available after choosing a node.
- **Services**: (Restart Stack only) Optionally select specific services within the stack to restart. Leave empty to restart all services.
- **Prune Targets**: (System Prune only) Select which resources to prune: containers, images, networks, volumes. All are selected by default.
- **Label Filter**: (System Prune only) Optionally filter prune operations to resources matching a specific Docker label (e.g. `com.docker.compose.project=mystack`).
- **Cron Expression**: Standard 5-field cron format. A human-readable preview appears below the input.
- **Enabled**: Toggle the task on or off.
- **Delete after successful run**: When checked, the task automatically removes itself after its first successful execution. This turns the task into a one-time operation. Failures keep the task so you can inspect the error, adjust the target if needed, and trigger a retry with **Run Now** or by re-enabling the schedule.
4. Click **Create**.
<Frame>
<img src="/images/scheduled-operations/create-dialog.png" alt="Create scheduled task dialog with action, cron expression, and prune target options" />
</Frame>
## Task List
The task list is displayed as a table with the following columns:
| Column | Description |
|--------|-------------|
| **Name** | The task name |
| **Action** | Task type badge: Restart Stack, Fleet Snapshot, System Prune, or Vulnerability Scan |
| **Target** | Stack name (with selected services, if any), or the target type for non-stack actions |
| **Schedule** | Human-readable description with the raw cron expression below |
| **Status** | Last run result: **Success** (green), **Failed** (red), or "Never run" |
| **Next Run** | When the task will next execute |
| **Enabled** | Toggle switch to enable or disable the task |
| **Actions** | Action buttons (see [Managing Tasks](#managing-tasks)) |
## Granular Targeting
### Per-Service Restart
When creating a Restart Stack schedule, you can target individual services instead of restarting the entire stack. After selecting a stack, Sencho reads the compose file and displays checkboxes for each defined service. Select the services you want to restart, or leave all unchecked to restart every service in the stack.
<Frame>
<img src="/images/scheduled-operations/per-service-restart.png" alt="Service checkboxes displayed when creating a per-service restart schedule" />
</Frame>
### Scheduled Vulnerability Scans
A Vulnerability Scan task runs Trivy against every image on the selected node and persists the results. The scan uses the same digest-based 24-hour cache as manual scans, so unchanged images are not rescanned on every run. See [Vulnerability Scanning](/features/vulnerability-scanning) for how results are surfaced in the UI and [Installing Trivy](/operations/trivy-setup) for setup on each node.
When a scheduled scan finishes, Sencho dispatches a completion notification with a summary of what was scanned and a breakdown of findings by severity. The full message format is documented in [Alerts & Notifications → Scheduled scan completion](/features/alerts-notifications#scheduled-scan-completion).
### Prune Label Filter
When creating a System Prune schedule, you can scope the prune to resources matching a specific Docker label. This lets you target resources from a particular stack or project without affecting unrelated containers, images, or volumes.
Enter a label in `key=value` format (e.g. `com.docker.compose.project=mystack`). Leave the field empty to prune all unused resources of the selected types.
<Frame>
<img src="/images/scheduled-operations/prune-label-filter.png" alt="Label filter input for scoping prune operations to specific Docker labels" />
</Frame>
### Stack Lifecycle Scheduling
The Stop Stack, Take Stack Down, Start Stack, and Backup Stack Files actions let you schedule lifecycle events for individual stacks on a per-node basis.
A common pattern for development or staging stacks is:
- A **Stop Stack** or **Take Stack Down** task scheduled for the end of the working day (e.g. `0 19 * * 1-5` — 7 PM on weekdays).
- A **Start Stack** task scheduled for the start of the working day (e.g. `0 8 * * 1-5` — 8 AM on weekdays).
**Stop vs Down:** Use Stop Stack when you want containers ready to resume quickly; Docker keeps the container filesystem in place and `Start Stack` simply restarts the existing containers. Use Take Stack Down when you want to fully release resources; `Start Stack` recreates the containers from the compose file on the next run.
**Backup Stack Files** keeps only the most recent backup per stack under `<DATA_DIR>/backups/<stackName>/`. For a point-in-time archive across all nodes, use a [Fleet Snapshot](/features/fleet-backups) schedule instead.
All four actions run against the local Sencho instance only. To schedule lifecycle operations on a remote node, manage the schedule from within that node's own Sencho UI.
### Delete after Successful Run
When **Delete after successful run** is enabled, the task removes itself from the schedule after its first successful execution. The entire task record, including run history, is removed. Use this for one-time preparatory or clean-up tasks: pre-scaling a stack before a deployment, a one-off backup before a config change, or stopping a stack once a migration is complete.
If the run fails, the task stays in the schedule unchanged so you can inspect the error and retry. Disabling the option at any time before the successful run prevents auto-deletion.
<Note>
Manual runs via **Run Now** also trigger the delete-after-run logic. If you run the task manually and it succeeds, the task removes itself.
</Note>
## Cron Expression Reference
Sencho uses standard 5-field cron expressions:
```
┌───────────── minute (059)
│ ┌─────────── hour (0–23)
│ │ ┌───────── day of month (131)
│ │ │ ┌─────── month (112)
│ │ │ │ ┌───── day of week (07, 0 and 7 = Sunday)
│ │ │ │ │
* * * * *
```
### Common Examples
| Expression | Description |
|-----------|-------------|
| `0 3 * * *` | Every day at 3:00 AM |
| `0 */6 * * *` | Every 6 hours |
| `0 3 * * 0` | Every Sunday at 3:00 AM |
| `30 2 1 * *` | 1st of every month at 2:30 AM |
| `0 0 * * 1-5` | Midnight on weekdays |
## Filtering by Node
When managing a multi-node fleet, you can filter the schedule list to show only tasks targeting a specific node. There are two ways to access this:
- **From the Nodes table:** Click the **calendar icon** on any node row in **Settings → Nodes** to jump directly to the Schedules view filtered to that node.
- **From the Schedules view:** A filter bar appears at the top showing which node you're viewing, with a **Clear filter** button to return to the full list.
When creating a new task while a node filter is active, Sencho pre-selects that node in the create dialog.
## Managing Tasks
Each task row has four action buttons:
- **Run Now** (play icon): Immediately execute the task without waiting for the next scheduled run. Manual runs are labeled "Manual" in the execution history.
- **Execution History** (clock icon): Open the run history panel for this task.
- **Edit** (pencil icon): Modify the task name, action, target, or schedule.
- **Delete** (trash icon): Permanently remove the task and all its execution history after confirmation.
Use the **Enabled** toggle switch in the task list to pause or resume a schedule without deleting it.
## Failure Notifications
When a scheduled task fails, Sencho automatically dispatches an **error-level alert** through your configured notification channels (Discord, Slack, or custom webhooks). The alert includes the task name, action type, and error message so you can diagnose the issue immediately.
When a previously failing task succeeds again, Sencho sends an **info-level recovery notification** to confirm the issue is resolved. This recovery-only approach avoids notification noise from tasks that succeed on every run.
Vulnerability Scan tasks always send a completion notification, even on a clean run, because the message carries severity counts you may want to react to. See [Alerts & Notifications → Scheduled scan completion](/features/alerts-notifications#scheduled-scan-completion) for the full message format.
To configure notification channels, go to **Settings > Notifications**.
<Frame>
<img src="/images/scheduled-operations/failure-notification.png" alt="Failure notification alert shown in the notification bell popover" />
</Frame>
## Execution History
Click the clock icon on any task to open the run history panel. The history is displayed as a table with the following columns:
| Column | Description |
|--------|-------------|
| **Time** | When the run started |
| **Source** | Whether the run was triggered by the **Scheduler** or **Manual** (via Run Now) |
| **Status** | Success, Failed, or Running |
| **Duration** | How long the run took (in seconds) |
| **Details** | Output summary or error message |
Run history is paginated at 20 entries per page. You can export the full history as CSV using the download button in the panel header.
Execution history is retained for 30 days.
<Frame>
<img src="/images/scheduled-operations/run-history.png" alt="Execution history showing run source, status, duration, and details with export button" />
</Frame>
## How It Works
The Scheduler Service runs in the background and checks for due tasks every 60 seconds. When a task's next run time has passed:
1. The scheduler verifies your Admiral license is active.
2. It executes the configured action using the same internal services that power the UI buttons (restart, snapshot, prune).
3. Results are logged to the execution history.
4. On failure, an alert is dispatched via your configured notification channels.
5. The next run time is recalculated from the cron expression.
If a task is still running from a previous execution, the scheduler skips it to prevent overlap.
If the server restarts while a task is mid-execution, the orphaned run record is automatically marked as failed with a "Server restarted during execution" message on the next startup. No manual cleanup is needed.
## Troubleshooting
### Task was automatically disabled
If a task's cron expression becomes invalid after creation (for example, due to a corrupt database edit), the scheduler disables the task and records the reason in the last error column. To fix this:
1. Open the task in edit mode.
2. Re-enter a valid cron expression.
3. Re-enable the task using the toggle switch.
### Run shows "Server restarted during execution"
This means Sencho was restarted (or crashed) while this task was mid-execution. The run was marked as failed automatically on startup. The task itself is still enabled and will run at its next scheduled time. If you want to re-run it immediately, use the **Run Now** button.
### Scheduled scan completed but no notification arrived
The scan notification shares the same delivery path as every other Sencho alert. Check, in order:
1. At least one channel is enabled in **Settings > Notifications** and the **Test** button succeeds for it.
2. If the stack the scan is associated with has notification routes defined, make sure at least one matching route is enabled; the routing layer takes priority over the global channels.
3. Open the notification bell. The in-app bell receives every notification regardless of channel configuration, and a failed external delivery is logged there as an error entry you can inspect.
4. On a remote node, notification channels must be configured on the remote instance itself; channel settings are per-node.
### Scan task fails with "Trivy binary is not available"
Trivy must be installed on the node that runs the scan, not just on the primary instance. Follow [Installing Trivy](/operations/trivy-setup) on the target node, then trigger **Run Now** from the schedule to confirm the scan succeeds before waiting for the next cron tick.
### Start Stack fails with a compose file error
The Start Stack action runs `docker compose up -d` against the stack's compose directory. If the stack folder has been removed or its compose file is missing, the run fails. Re-create or restore the stack in the [Editor](/features/editor), then trigger **Run Now** to confirm the start succeeds before the next cron tick.
### Auto-backup always overwrites the previous backup
Backup Stack Files keeps a single backup slot per stack under `<DATA_DIR>/backups/<stackName>/`. Each run overwrites the previous backup. This is by design for simplicity. If you need a timestamped archive of compose files across all nodes, use a [Fleet Snapshot](/features/fleet-backups) schedule, which stores versioned snapshots.
### One-shot task disappeared after a successful run
A task with **Delete after successful run** enabled removes itself automatically after the first successful execution, including its run history. This is expected behavior. If you need a record of the run before deletion, export the execution history to CSV via the clock icon before the next successful run.