* feat(recovery): add safe-mode recovery surface and emergency CLI
Add a read-only Recovery tab under Settings (admin-only) backed by a new
GET /api/diagnostics endpoint reporting app version, database integrity,
encryption-key status, Docker reachability, account and SSO counts, and
non-secret configuration. The endpoint loads without Docker or live metrics
so it stays available when the dashboard does not, requires a genuine admin
session, and builds its config block from a non-secret allowlist so no
credentials are ever exposed.
Expand the emergency command-line toolkit beyond the two-factor reset with
seven host-level commands: reset-password, create-emergency-admin,
clear-sessions, disable-sso, diagnostics, validate-db, and backup-data. Each
prints its result, exits with a meaningful status code, and writes an audit
entry where it changes state.
Document the toolkit in a new operator guide and link it from the recovery
and two-factor pages.
* feat(recovery): download the emergency command reference as a text file
The recovery commands are needed exactly when the dashboard is unreachable,
so reading them only in-app is a chicken-and-egg problem. Add a Download
button to the command-line section that saves the full
`docker compose exec sencho ...` reference as a text file, letting operators
keep it on hand before they need it. Reuses a shared download helper with the
existing diagnostics export.
* fix(recovery): harden diagnostics, backup, and emergency-admin against edge cases
Address findings from an independent review of the recovery toolkit:
- DiagnosticsService now degrades instead of throwing when a queried table is
missing or corrupt: each read falls back and is folded into database.ok, so a
broken database reports "problem detected" rather than failing the whole
endpoint or showing a misleading healthy state with zeroed counts.
- backup-data refuses a destination that resolves to the live database, which
would otherwise report success while producing no separate copy.
- create-emergency-admin now applies the same username rule as the user-
management route, extracted to a shared helper so both stay in sync.
Adds tests for a missing read table, a malformed emergency-admin username, and
the backup same-target rejection.
Every mutating /api/* request runs an individual INSERT into audit_log
which serializes against other writers under burst load (SQLite's
single-writer model). Buffer the writes in DatabaseService and flush
them in a single transaction either every second or once the buffer
reaches 100 entries, whichever comes first.
Read paths (getAuditLogs, getAuditLogsInRange, cleanupOldAuditLogs)
drain the buffer first so callers always see a consistent view, which
keeps the existing test pattern of insert-then-read working.
Graceful shutdown flushes before db.close() so no entries are lost on
clean exit. The 1s flush timer is unref'd so the buffer cannot keep
the process alive on its own. The CLI resetMfa script flushes
explicitly before returning since it exits before the timer fires.
* feat(auth): add TOTP two-factor authentication with backup codes
Adds RFC 6238 time-based one-time password support to every tier,
integrated with the existing password and SSO login paths.
Backend:
- New MfaService wrapping otplib with a plus or minus 1 step tolerance,
base32 secret generation, and hashed single-use backup codes (bcrypt).
- user_mfa and mfa_used_tokens tables in DatabaseService. The second
table is a DB-backed replay blacklist, purged on a 60s interval.
- authMiddleware now recognizes an mfa_pending scope. A token carrying
that scope is rejected on every route except the MFA challenge and
logout, so no API surface is reachable before the second factor
clears.
- /api/auth/login issues only a short-lived mfa_pending cookie when the
user has MFA enrolled. /api/auth/login/mfa consumes that cookie,
verifies the code (or backup code), and swaps in a real session.
- /api/auth/mfa/* routes for status, enrol/start, enrol/confirm,
disable, backup-code regenerate, and SSO-bypass opt-in.
- Admin recovery path: POST /api/users/:id/mfa/reset clears the target's
MFA state, bumps token_version, and writes an audit log entry.
- CLI emergency fallback: backend/src/cli/resetMfa.ts is wired via
`npm run reset-mfa <username>` and also exported for tests.
- SSO flows (LDAP and OIDC) gate on user_mfa.sso_enforce_mfa before
issuing a session; default behaviour keeps the SSO path frictionless.
- Per-user lockout after 5 consecutive failed codes (15 min).
Frontend:
- AppStatus gains an mfa-challenge branch driven by /api/auth/status.
- New MfaChallenge screen, MfaEnrollDialog (QR plus manual secret plus
backup codes), MfaDisableDialog, MfaBackupCodesDialog.
- Account section shows a Two-factor authentication card with enrol,
regenerate, disable, and the SSO-enforce toggle (shown only when SSO
providers are configured).
- Users section gains a Reset 2FA action for admins.
Docs:
- New user guide at features/two-factor-authentication.mdx.
- New admin guide at operations/two-factor-admin.mdx.
- SSO page cross-links to the 2FA doc.
* fix(mfa): drop unused TEST_PASSWORD import and stale eslint disable
* fix(mfa): simplify e2e openAccountSettings helper to match working pattern
* fix(mfa): make e2e suite self-contained and always clean up
Test #2 called loginAs() before the MFA challenge step, which waited for
the dashboard indicator that never appears once the previous test enrolled
the user. That timeout skipped the rest of the serial block, including
the disable step, leaving MFA enabled and breaking every later spec.
Two fixes:
- Tests #2 and #3 now navigate directly to the login page instead of
piggybacking on loginAs, which only handles the password-only path.
- A new afterAll hook unconditionally disables MFA via the API using two
unused backup codes, so the DB is reset even if a test fails midway.
* fix(e2e): use backup code for mfa recovery to avoid totp replay race
The final recovery step in the backup-code replay test previously
generated a fresh TOTP to sign back in. When the timing landed inside
the same 30-second window that test #2 consumed, the server's replay
blacklist correctly rejected it, producing a ~50% flake rate. Backup
codes are single-use and sidestep the replay window, so the recovery
becomes deterministic.
* fix(e2e): drive mfa disable test through the challenge screen
Test #4 called loginAs after test #3 left MFA enabled, but loginAs
waits for the dashboard indicator and does not handle the challenge
screen, so it timed out. Drive the login manually, satisfy the
challenge with a backup code, and use a backup code for the disable
step too to avoid any TOTP replay-window race against earlier tests
in the serial block.