Audit cleanup of CHANGELOG.md accumulated over 43 releases:
- Delete 23 orphaned `[Unreleased]` sections. Release-please does not
consolidate manually-added Unreleased content into the next release, so
each previous contributor's notes were left stranded between release
blocks. The auto-generated release sections above each orphan already
captured the commits.
- Sanitize internal implementation details that should not be in a
public changelog: internal service/class/middleware names, library
names, container file paths, internal endpoint paths, exact rate-limit
thresholds, encryption primitives, and CVE/version-specific remediation
details.
- Normalize tier naming throughout: legacy "Team Pro" / "Personal Pro" /
"Sencho Pro" / "(Pro)" references rewritten to "Skipper" / "Admiral" /
"Skipper and Admiral" per current product tiering.
- Deduplicate the v0.39.0 bulk dump and other sections where
release-please swept the same commit subjects in multiple times.
- Rewrite the v0.1.0 section from a raw engineering log (with file
paths, attack payloads, and internal module names) into user-facing
release notes organised by Security / Added / Fixed / Changed /
Removed.
- Update CONTRIBUTING.md to tell contributors not to edit CHANGELOG.md
directly, matching the actual release-please flow. This is the root
cause that created the 23 orphan Unreleased sections in the first
place.
Verification:
- `## [Unreleased]` count: 0
- Legacy tier name count: 0
- Internal service/library name count: 0
- CVE IDs: 0
- File shrunk from 1,599 to 901 lines
Every filesystem operation against user compose folders (save, create,
deploy, update, rollback, template install, fleet snapshot restore)
previously failed with EACCES whenever a stack container had chowned
its own bind mount to another UID, which is extremely common with
linuxserver/* images and anything that runs as root by default.
Running Sencho as root eliminates the entire class of permission bugs
at the source and matches the default posture of Portainer, Dockge,
Komodo, and Yacht. Mounting /var/run/docker.sock is already equivalent
to root-on-host, so the previous non-root hardening provided essentially
no additional isolation while breaking real features.
Changes:
- docker-entrypoint.sh: default path stays root, no GID dance, no
privilege drop. Opt-out via SENCHO_USER=sencho restores the legacy
behavior bit-for-bit (chown data dir, match Docker socket GID,
su-exec to the user). Fails fast if SENCHO_USER names a nonexistent
account. Kubernetes / OpenShift forced-non-root compat preserved via
the existing id -u = 0 guard.
- FileSystemService: delete forceDeleteViaDocker (the ~40-line helper
that shelled out to an alpine container to work around EACCES during
deleteStack) and simplify deleteStack to a single fsPromises.rm call.
Tests updated accordingly.
- Dockerfile: keep the sencho user+group pre-created so the opt-out
path works out of the box; comments updated to document the new
default.
- Docs: new "Container user" section in configuration.mdx documenting
the root default and the SENCHO_USER opt-out; troubleshooting and
self-hosting updated to match.
The Skipper/Admiral atomic deploy/update path used to create
.sencho-backup/ inside the user's stack folder, which silently failed
with EACCES whenever a container had chowned the bind mount (swag,
tautulli, linuxserver/* images, etc). That broke auto-rollback and the
manual rollback endpoint for those stacks. Stack backups now live under
<DATA_DIR>/backups/<stackName>/ next to sencho.db, which is always
writable by the Sencho user.
While stress-testing the same scenario, MonitorService also flooded the
error log with "Error parsing stats for container ... 404 no such
container" because per-container stats polls (30s tick) raced with
docker compose recreating containers. The 404 case is now skipped
silently; non-404 stats failures still log at error level.
The helper container that runs `docker compose up --force-recreate` was
spawned with `docker run -d`, so the command returned immediately with
just the container ID. Any failure happening INSIDE the helper (bad
compose file, image mismatch, permission issue, socket problem) was
invisible: execFile's callback only fired for `docker run` command
errors, never for errors inside the detached helper. The UI fell back to
the generic 3-minute "Local update did not complete" heuristic with no
actionable information.
The helper now runs attached, so execFile's callback receives the
helper's exit code and stderr directly for any failure that happens
before the recreate kills this process. Additionally, the helper
persists exit code + stderr to `/app/data/.sencho-update-error` before
exiting, so the error survives the gateway's own death. On startup,
`SelfUpdateService.recoverPreviousError()` reads and deletes that file,
routing the real error through the existing `getLastError()` path so the
freshly booted gateway reports exactly why the previous attempt failed
instead of the generic timeout.
Every published Sencho Docker image is now signed with Sigstore cosign
using GitHub Actions keyless OIDC, and ships with an embedded SBOM plus
SLSA provenance attestation attached as OCI referrers on the image.
Users can verify the signature and inspect the attestations themselves
before pulling an image into production.
cosign verify saelix/sencho:<tag> \
--certificate-identity-regexp "https://github.com/AnsoCode/Sencho/.*" \
--certificate-oidc-issuer https://token.actions.githubusercontent.com
The release pipeline now also publishes a new moving minor tag
(saelix/sencho:X.Y, e.g. 0.42) alongside the existing latest and
immutable X.Y.Z tags, so operators who want the latest patch on a given
minor line can pin without chasing every release. Since every 0.x minor
is potentially breaking, the {{major}}-only tag is deferred until 1.0.
cosign is installed via the sigstore/cosign-installer action pinned to
a full commit SHA. The signing step batches every tag produced by
metadata-action into a single cosign invocation pinned to the built
digest, so all tags resolve to the same manifest and we save N-1
Rekor/Fulcio round trips. The job declares id-token: write so the
ambient OIDC token can be minted.
A new docs page (docs/reference/verifying-images.mdx) walks users
through installing cosign, verifying an image, and inspecting the SBOM
and provenance with docker buildx imagetools. A note calls out that
signatures start shipping with v0.43.0, so older images failing
verification is expected behavior and not tampering.
The "Updating Sencho..." overlay used to dismiss prematurely while the
image pull was still running, after which the local node card would get
stuck in "updating" and eventually surface a generic "Timed Out" error
while the container remained on the old version.
Three root causes are addressed:
1. The image pull was synchronous (`execFileSync`), which blocked the
Node event loop. The overlay's health probe saw the server come back
the moment the pull finished and reloaded the page, even though the
container had not restarted yet. The pull is now async via
`promisify(execFile)`, so /api/health and /api/fleet/update-status
keep serving throughout.
2. The overlay reloaded on the first 200 from /api/health regardless of
whether the underlying process had actually restarted. /api/health
now exposes the gateway boot timestamp, and the overlay captures it
pre-update and only reloads when it observes a different value. A
wasOffline-then-online fallback handles the case where the pre-update
fetch failed.
3. Helper container spawn errors from `docker run` were silently
discarded, so a failed compose recreate never surfaced anywhere.
Errors are now captured into `lastUpdateError` via the execFile
callback and surfaced through the existing /api/fleet/update-status
error path.
A 3-minute early-fail heuristic on the local node block surfaces a clear
failure message when the helper fails silently, instead of waiting the
full 5-minute timeout for an unknown failure.
Replaces five ad-hoc in-process caches (project name map, templates, latest
version, fleet update status, remote node meta) with a single internal
CacheService that provides TTL, inflight-promise deduplication to protect
against thundering herd, stale-on-error fallback, and per-namespace
hit/miss/stale/size counters for observability.
Wraps the hot-path dashboard endpoints in the cache with write-path
invalidation: /api/stats (2s), /api/system/stats (3s), and
/api/stacks/statuses (3s). Keys are namespaced by nodeId so switching nodes
never serves another node's data. Every route that mutates container or
stack state calls invalidateNodeCaches(nodeId), which also drops the global
project-name-map, so user actions stay instantly reflected in the UI.
For /api/system/stats the cheap per-request network rx/tx block is kept
outside the cache so live-updating charts stay smooth while the expensive
systeminformation.currentLoad() CPU sample (~200ms) is reused across the
TTL.
Adds admin-only GET /api/system/cache-stats returning per-namespace
counters for operators who want to observe cache effectiveness.
Enables the compression middleware site-wide for JSON responses. Large
payloads like /api/templates shrink roughly 5x on the wire. SSE endpoints
are explicitly excluded via a Content-Type filter so live log tails and
metric streams are not buffered.
Bumps vitest hookTimeout to match testTimeout (15s) so parallel fork
workers do not hit the default 10s hook limit under CPU contention.
Adds 35 new tests (26 unit for CacheService, 9 integration for cached
endpoints) covering TTL expiry, inflight dedup, stale-on-error,
namespace invalidation, entry-cap safety guard, and write-path
invalidation end-to-end through Express routes.
* fix(fleet): add Docker Hub fallback for version detection on private repos
The GitHub Releases API returns 404 for private repos, causing the
latest version fetch to silently fail and fall back to the gateway's
own version (defeating the update detection fix from PR #454).
Now tries GitHub first, then falls back to Docker Hub tags API which
is always public. Adds console.warn logging on fetch failures per
Directive 7.
* ci: trigger CI re-run
* fix(api): add tiered rate limiting to prevent polling lockouts
Replace the single global rate limiter (100 req/min/IP) with a tiered
system that separates high-frequency polling endpoints from standard
API traffic:
- Polling tier (300/min): /stats, /system/stats, /stacks/statuses,
/metrics/historical, /health, /meta, /auth/status, /auth/sso/providers,
/license. Exempt from the global limiter but governed by their own
safety net to prevent resource exhaustion.
- Standard tier (200/min): All other endpoints, raised from 100.
- Webhook tier (500/min): POST /webhooks/:id/trigger, dedicated limiter
for CI/CD platforms sharing datacenter IPs.
- Auth tier: Unchanged (5-10 attempts / 15 min).
Enterprise adaptations:
- Authenticated requests keyed by user session (JWT sub/username) instead
of IP, preventing shared NAT/VPN environments from pooling budgets.
- Internal node-to-node traffic (node_proxy tokens) bypasses all rate
limiters entirely.
Includes comprehensive stress tests (21 cases) validating tier
separation, node proxy bypass, and per-user keying.
* chore(deps): bump axios to 1.15.0 to fix SSRF vulnerability
Addresses GHSA-3p68-rc4w-qgx5 (NO_PROXY hostname normalization bypass).
* fix(fleet): detect updates via GitHub Releases instead of gateway self-comparison
The fleet update check compared each node's version against the gateway's
own version, so the local node could never appear outdated. Now fetches
the actual latest release from GitHub Releases API with a 30-minute
in-memory cache and thundering-herd protection. The Recheck button
invalidates this cache via ?recheck=true to force a fresh lookup.
* docs(fleet): update docs to reflect GitHub Releases version detection
Replace "Gateway version" references with "Latest version" to match the
new label. Document that version comparison uses the latest GitHub release
rather than the gateway's own version, and that Recheck refreshes the
cached latest version.