10 Commits

Author SHA1 Message Date
Anand fadc1f24e3 Wave 3 (3/3): scheduled fleet health digests + Terraform/Ansible export
- internal/digest: fleet summary builder + scheduler, delivered via the existing notify.Notifier (SMTP/Gotify), admin settings + send-now endpoint
- internal/export: Terraform (proxmox_vm_qemu/proxmox_lxc) and Ansible YAML inventory generators from live fleet inventory, downloadable via /connections/{id}/export/{terraform,ansible}
- Migration renumbered 00031 to avoid colliding with 00026-00030 already on this branch
2026-09-10 15:08:22 +05:30
Anand 542a9b2eb5 Wave 3 (2/3): config drift detection + guest lifecycle policies
- internal/api/drift.go: per-guest config/firewall baseline capture + diff, fleet-wide drift summary
- internal/poller/lifecycle.go + internal/api/lifecycle.go: snapshot retention sweep (tag or fleet default, dry-run until explicitly enforced) and orphaned-disk detection (report-only, surfaced as alerts)
- Migrations renumbered 00029/00030 to avoid colliding with 00026-00028 already on this branch
2026-09-10 15:04:30 +05:30
Anand 2c8e114aeb Wave 3 (1/3): capacity forecasting + fleet health score
- internal/api/forecast.go: OLS trend fit over node RRD history, per-node and fleet-wide capacity-warning endpoints
- internal/api/health.go: 0-100 health score (reachability, alerts, node/guest status, backup reliability) per-connection and fleet-wide
- web: CapacityForecastWidget, HealthScoreWidget dashboard widgets, registered in the widget catalog
2026-09-10 15:02:59 +05:30
Anand 4dfeb9848e Wave 2: real-time events/webhooks, cross-remote tags/search/bulk-ops, cert & session monitoring
- internal/events: in-process pub/sub bus; GET /api/v1/events SSE stream
- internal/notify/webhooks.go: signed outgoing webhook dispatcher w/ retry + delivery log
- internal/api/tags.go, search.go, bulk.go: cross-remote tag aggregation, global search, fleet-wide bulk guest actions
- internal/auth/sessions.go: session listing/revocation (self-service + admin), extends the existing stateful session table
- internal/poller/certificates.go: cert-expiry + connection-staleness alerting via the existing alert_rules pipeline
- web: useEventStream hook, GlobalSearch + BulkOperationsPage (not yet wired into nav), SessionsCard on ProfilePage
2026-09-10 11:49:11 +05:30
Anand ad43bec538 Merge PBS remotes + cross-cluster migration (backend) and UI modernization pass (frontend)
- internal/pbs: PBS client, datastores/namespaces/snapshots/prune/GC/sync/verify jobs
- connections: type column (pve|pbs), PBS-aware client resolution
- pve/guests: RemoteMigrateGuest + POST .../remote-migrate endpoint for true cross-cluster live migration
- web: shared ListSearch/Textarea primitives, DataTable search modernized
2026-09-10 11:31:57 +05:30
Anand 846cb8ce63 Bundle Needle 2 as a built-in, default AI assistant
Ships the official Needle 2 CLI binary (Apache-2.0) baked into the ferrum
binary itself for windows/amd64, linux/amd64, linux/arm64, and
darwin/arm64 via go:embed behind per-platform build tags. No download,
no FERRUM_NEEDLE_BIN, no manual provider setup on those platforms.

- internal/needle: resolveBinPath prefers an explicit FERRUM_NEEDLE_BIN,
  otherwise extracts the embedded binary to a cache file on first use.
- Fixed a port collision: Needle's --serve defaulted to :8080, the same
  default as ferrum's own server; it now runs on a dedicated port.
- Fixed the real 'assistant times out' bug: Needle's tool-retrieval does
  a one-time embedding pass on its first request once the tool catalog
  exceeds 5 tools (ferrum declares 10), which routinely took longer than
  the old readiness check's 500ms-per-attempt retry loop allowed. Each
  cancelled attempt kept occupying Needle's single-threaded request
  loop, so retries piled up and never let a real response through.
  Replaced with two phases: cheap/retryable raw TCP dials until the
  socket accepts a connection, then exactly one real request allowed to
  run for the full startup budget.
  Subprocess stdout/stderr are now captured so a future failure surfaces
  a real reason instead of a bare timeout.
- seedBuiltinNeedleProvider now runs on every startup (not just fresh
  installs), idempotently upserting the built-in provider and always
  reasserting its model as the global default AI assistant.
- .gitignore: carved out an exception for the bundled Windows binary,
  which the blanket *.exe rule would otherwise have silently excluded.
2026-09-06 20:26:00 +05:30
Anand 22c1dd382c Add API keys, MCP server, admin AI providers, and a built-in local LLM option
- User-scoped API keys (Profile > API Keys) for 3rd-party REST API access
  and MCP clients, each locked to one scope at creation, with expiry,
  revocation, and last-used tracking.
- A hand-rolled MCP (Model Context Protocol) server exposing the fleet
  (connections, nodes, guests, storage, pools, alerts, cluster status) as
  read tools plus one admin-gated power-action tool, so Claude Code/Desktop
  or any other MCP client can query and operate the fleet directly.
- Both the REST API and MCP are off by default and toggleable instance-wide
  from Settings > API & MCP, enforced live on every request.
- Admin-managed AI providers (any OpenAI-chat-completions-compatible
  endpoint) backing the AI Assistant's tool-calling loop, replacing the
  single hardcoded provider.
- A built-in, zero-config, no-API-key local provider backed by Needle 2
  (internal/needle) for fully offline tool-calling, wired in as a one-click
  preset. Requires the operator to separately download the Needle 2 binary
  and point FERRUM_NEEDLE_BIN at it -- Ferrum never fetches executable
  content from the network itself; see README "Built-in LLM (Needle 2)".
- System settings (CORS allow-list, instance-wide toggles) moved to the
  admin Settings UI; environment variables are now scoped to true
  bootstrap-level config only (listen address, TLS, DB connection, secret,
  optional Needle binary path).
- Fixed: node Journal tab 502'ing with "unexpected end of JSON input" on an
  empty response, and separately with a decode error on PVE versions that
  return a bare-string journal line instead of the documented {n,t} object.
- Fixed: bottom content padding disappearing on every page except the AI
  Assistant (an unconditional h-full on the content wrapper let overflowing
  content bleed through where the padding should render).
- Fixed: Profile page felt cramped despite a wide viewport (stray max-w-2xl
  cap not present on the equivalent Settings page).
- Test coverage added for the previously-untested MCP package and the new
  Needle adapter (20 new Go tests), plus a regression test for the journal
  decode fix.
2026-09-06 13:26:30 +05:30
Anand 7295a28ec9 Opt-in SSO single logout, nice chart scales, storage usage donut, topology export/snap, disk SMART + temperature
- OIDC: RP-Initiated Logout is now opt-in (default off) via a new
  'single_logout' setting, with the exact post-logout redirect URL shown
  in Settings for the admin to register at their provider. Fixes the
  regression from last time: enabling it unconditionally broke sign-out
  for anyone whose IdP hadn't been told to trust the redirect yet
  (Keycloak's invalid_redirect_uri, browser stuck on a stale page).
- formatBytes shows up to 2 decimals (was an adaptive 0-or-1 rule); every
  bytes/rate chart now computes a rounded 'nice' axis scale (0/5/10/15/20
  GB, the standard Heckbert algorithm) and locks every tick + the tooltip
  to one consistent unit derived from the axis's own max.
- Storage page: the capacity donut is now sized by used bytes + a free
  remainder instead of by total capacity share, which was always 100%
  the moment there was only one pool — completely disconnected from the
  "224 GB used" text next to it.
- Topology: Export SVG (fits the full diagram regardless of current pan/
  zoom) and a Snap-to-grid toggle.
- New Disks tab on the node detail page: every physical disk with model/
  serial/size/type, PASSED/FAILED health, SSD/NVMe wearout %, and a
  per-drive temperature read from SMART, plus a full SMART attribute
  table per disk. Backed by new /nodes/{node}/disks and
  /nodes/{node}/disks/smart endpoints.
- CPU/GPU temperature is not exposed by Proxmox's own API (no built-in
  lm-sensors/nvidia-smi integration) and isn't something this can add
  without a node-side agent Proxmox doesn't ship — disk temperature via
  SMART is the thermal data actually available.
2026-09-04 00:22:53 +05:30
Anand 61ee806af6 Add realtime polling defaults, Gotify/SMTP notifications, UI-configurable OIDC, and expanded admin settings
- Default all queries to a 20s poll + refetch-on-focus (main.tsx) instead of
  a per-page opt-in, so every page/widget stays live without manual tuning.
- New internal/notify package: Gotify and SMTP (stdlib net/smtp, STARTTLS
  and implicit-TLS-on-465) notifications, each independently optional. Fires
  from the alert evaluator on new alert triggers; admin-configurable from
  Settings with a send-test-notification action per channel.
- OIDC/SSO moved from config.yaml-only to a DB-backed, admin-editable
  Settings card — swaps the live client with no restart. config.yaml is
  used to seed the database once on first boot after upgrading.
- New Security settings: session TTL, login lockout policy, and a real
  "require 2FA for admins" enforcement (requireTOTPEnrolled middleware)
  that blocks non-enrolled admins from everything but /profile and logout.
- New org-wide default preferences (theme/accent/look/landing page) for
  brand-new accounts, plus a personal landing-page picker and an
  email-me-alerts opt-in on Profile.
- Storage page: separate Local vs Shared/External storage tables and
  capacity donuts, fixing shared-storage totals that were being summed once
  per node that mounts them (e.g. a 2TB NFS share on 4 nodes read as 8TB).
- RankedBarChart: stop the longest bar's value label wrapping onto two
  lines (recharts auto-wraps LabelList when space is tight).
2026-09-03 22:00:41 +05:30
Anand 9ab6b4340a Initial commit 2026-09-03 08:36:04 +05:30