Commit Graph

1103 Commits

Author SHA1 Message Date
rcourtman 2f8cf462bb feat: add temperature monitoring status badge to node cards
- Display green 'Temperature' badge when temperature monitoring is active
- Shows on both PVE and PBS nodes in Settings page
- Badge appears alongside VMs, Containers, Storage, Backups badges
- Only displays when node.temperature.available is true

This provides clear visual feedback that SSH temperature monitoring
was successfully configured during setup.
2025-10-01 08:18:18 +00:00
rcourtman d0f049d373 refactor: improve setup script output professionalism
- Remove excessive emojis and decorative elements
- Use clear, concise language throughout
- Simplify progress indicators to simple checkmarks
- Remove unnecessary tips and verbose explanations
- Improve error message clarity
- Use proper capitalization (not ALL CAPS for emphasis)
- Clean up temperature monitoring prompt to be more direct

The setup script now presents a more professional, enterprise-ready
appearance while maintaining all functionality.
2025-10-01 08:14:49 +00:00
rcourtman fdf0e0b958 feat: automate SSH key generation and embedding in setup scripts
- Add getOrGenerateSSHKey() function that automatically generates SSH keypair if needed
- Embed SSH public key directly in setup scripts (no manual copy/paste required)
- Simplify temperature monitoring setup - user just types 'y' and it's done
- Improves UX: removes manual steps for SSH key setup

Changes:
- internal/api/config_handlers.go: Add SSH key generation and auto-embedding
- frontend-modern/src/components/Settings/NodeModal.tsx: Remove dead setupCode modal code
- Setup script now includes embedded SSH_PUBLIC_KEY variable

User workflow before:
1. Run setup script
2. Prompted to run commands on Pulse server
3. Copy SSH public key manually
4. Paste into setup script
5. Done

User workflow now:
1. Run setup script
2. Type 'y' for temperature monitoring
3. Done (SSH key automatically installed)
2025-10-01 08:10:48 +00:00
rcourtman 524210468f fix: copy command to clipboard when generating PVE quick setup URL
The copy button in Quick Setup for PVE was generating the setup URL
but not actually copying the command to clipboard. Users had to click
the separate 'Copy Command' button after generation.

Fixed to automatically copy the curl command to clipboard immediately
after generating the setup URL, matching the PBS behavior.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
2025-10-01 07:41:25 +00:00
rcourtman ec1d8b3303 fix: ensure PULSE_DATA_DIR is exported in dev mode and improve sync validation
Additional safeguards to prevent dev/production config conflicts:

1. **hot-dev.sh**: Explicitly export PULSE_DATA_DIR before starting backend
   - Ensures backend always uses /opt/pulse/tmp/dev-config in dev mode
   - Prevents accidental fallback to /etc/pulse
   - Adds logging to show which config directory is being used

2. **sync-production-config.sh**: Smart encryption key handling
   - Never overwrites existing dev encryption key
   - Warns if production key is newer (unusual scenario)
   - Keeps dev key to avoid breaking encrypted configs
   - Adds detailed logging of sync decisions

These changes ensure that when Vite restarts:
- Backend always uses the correct dev-config directory
- Sync script never breaks working dev configuration
- All decisions are logged clearly for debugging

Related to previous commit fixing nodes.enc corruption.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
2025-10-01 07:33:20 +00:00
rcourtman 27373587c6 fix: prevent nodes.enc corruption and data loss with comprehensive safeguards
This commit addresses critical issues where nodes configuration was being
lost or corrupted, causing user frustration and data loss.

## Changes:

### 1. Sync Script Protection (sync-production-config.sh)
- Never overwrites newer dev config with older production files
- Validates timestamps before syncing
- Shows detailed logging of sync decisions
- Prevents accidental overwrites of working configuration

### 2. Timestamped Backups (persistence.go)
- Creates timestamped backup before EVERY save (e.g., nodes.enc.backup-20251001-073000)
- Maintains "latest" backup for quick recovery
- Auto-cleans old backups (keeps last 10)
- Ensures we can always recover from corruption

### 3. Empty Config Protection (persistence.go)
- BLOCKS attempts to save empty nodes config when existing nodes exist
- Prevents accidental data wipes
- Returns error with clear message about what was blocked

### 4. Enhanced Corruption Recovery (persistence.go)
- Detects "cipher: message authentication failed" errors
- Automatically attempts recovery from backup files
- Renames corrupted files with timestamps for forensics
- Logs detailed recovery process

### 5. Performance Logging (GuestRow.tsx)
- Added timing for individual metadata API calls
- Helps identify performance bottlenecks

## Why This Matters:
Previous behavior allowed:
- Corrupted files to overwrite working configs
- Empty configs to delete all nodes
- No way to recover from corruption
- Race conditions during rapid restarts

New behavior ensures:
- Multiple backup copies always exist
- Corruption auto-recovers from backups
- Empty saves are blocked
- Sync script validates before overwriting

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
2025-10-01 07:31:50 +00:00
rcourtman 6d517e46b2 fix: add disk object to VM/container API responses
addresses #481

The frontend expects a disk object with {total, used, free, usage} fields
but the backend was only sending flat diskUsed/diskTotal values. This
caused the DISK column to show disk I/O values instead of disk usage.

Added DiskObj field to VMFrontend and ContainerFrontend structs and
populated it in the converters, matching how Memory is already handled.
2025-10-01 07:01:23 +00:00
rcourtman a662bdb9a5 fix: preserve zero values for alert thresholds in UI
addresses #480

Changed from || to ?? operator when loading alert config to properly
handle zero values. Previously, setting a threshold to 0 would revert
to defaults when navigating away due to || treating 0 as falsy.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
2025-10-01 06:57:09 +00:00
rcourtman 178bb5aa34 fix: clarify alert delay label text
Changed "seconds before triggering" to "seconds above threshold before triggering" to make it clearer that the delay applies to how long a value must stay above the threshold.

addresses #470

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
2025-10-01 06:55:30 +00:00
rcourtman 9737735cae fix: stable node group sorting in dashboard for duplicate hostnames
addresses #479

when displaying grouped guests in the dashboard, node groups with the same
hostname were being sorted inconsistently, causing the groups to swap
positions on each data refresh.

fixed by adding instance ID as a secondary sort key when node names are
equal. this ensures stable, consistent ordering even when multiple nodes
share the same hostname.
2025-10-01 06:53:15 +00:00
rcourtman c3d9be23a7 fix: use instance ID for node selection filtering across all views
addresses #476

comprehensive fix for duplicate node name handling when users click to
filter by a specific node:

- NodeSummaryTable: pass node.id instead of node.name when user clicks a node
- Dashboard: filter guests by instance ID instead of hostname
- Storage: filter storage by instance ID instead of hostname
- DiskList: filter physical disks by instance ID instead of hostname
- UnifiedBackups: add instance field to UnifiedBackup type and filter by
  instance ID instead of hostname

this ensures that when users select a node with a duplicate hostname, they
only see resources from that specific node, not from all nodes sharing the
same hostname.
2025-09-30 22:21:20 +00:00
rcourtman 6b1a5f76cf fix: use instance ID for grouping/matching in GuestURLs and UnifiedBackups
addresses #476

- GuestURLs: group guests by instance ID instead of hostname
- UnifiedBackups: match snapshots to VMs/containers using instance ID
  instead of hostname when finding guest names

this prevents incorrect grouping and matching when multiple nodes share
the same hostname.
2025-09-30 22:08:48 +00:00
rcourtman 1cf61c3025 fix: hide temperature column when no SSH configured
Only show the temperature column in node tables when at least one node
has SSH configured and temperature monitoring available.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
2025-09-30 22:02:13 +00:00
rcourtman b47845ecb8 fix: add instance field to backup/snapshot structs for duplicate node names
addresses #476

added Instance field to StorageBackup and GuestSnapshot structs in both
backend and frontend to properly handle nodes with duplicate hostnames.
updated backup and snapshot counting logic to use instance ID instead of
hostname, consistent with the VM/container/storage count fixes.

this completes the fix for #476 - all counts and groupings now use unique
instance IDs instead of hostnames.
2025-09-30 21:57:05 +00:00
rcourtman 312f1b5862 fix: correct VM/container/storage counts for duplicate node names
addresses #476

the node summary table was still using node.name (hostname) to count
VMs, containers, and storage, which caused incorrect counts when multiple
nodes share the same hostname. changed the count logic to use the unique
instance ID (vm.instance === node.id) instead of hostname matching, making
it consistent with the grouping fix in 5180b84d4.
2025-09-30 21:50:12 +00:00
rcourtman 645c97850b fix temperature struct field names in mock generator
fixes build error from using wrong field names (ID/Temperature instead of Core/Temp)

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
2025-09-30 21:42:12 +00:00
rcourtman 286eba9985 add temperature data to mock nodes
addresses mock data not keeping up with new features added in v4.16.0. mock nodes now generate realistic CPU package, core, and NVMe temperatures to match the temperature monitoring feature.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
2025-09-30 21:36:43 +00:00
rcourtman 86d240e70e chore: bump version to v4.16.0 2025-09-30 21:10:14 +00:00
rcourtman 6bfaa8b79a fix: OIDC redirect URL now respects X-Forwarded-Proto header
Addresses #327 - Users behind reverse proxies (Traefik, nginx, etc) were
experiencing redirect loop issues because the redirect URL was being built
with http:// instead of https:// when X-Forwarded-Proto was set.

Changes:
- Build OIDC redirect URL dynamically from each request instead of at startup
- Respect X-Forwarded-Proto and X-Forwarded-Host headers from reverse proxies
- Update UI help text to clarify auto-detection behavior
- Add debug logging to show how redirect URL is constructed

When redirect URL is not explicitly configured, Pulse now builds it from
the incoming request headers, properly detecting HTTPS when behind a proxy.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
2025-09-30 21:06:20 +00:00
rcourtman 0778b8f002 fix: display backup-id for PBS host backups instead of 0
addresses #454

PBS host backups (created with standalone proxmox-backup-client) use
the hostname as backup-id (e.g. "delly", "krom-pc") instead of a
numeric VMID. Now the VMID column displays this hostname string for
host backups instead of incorrectly showing 0.

Changes:
- Allow vmid field to be string or number in UnifiedBackup type
- Keep backup-id as string for host backups, numeric for VMs/LXCs
- Update column header to show "VMID/Host" when host backups exist
- Fixed in both code paths: direct PBS API and PVE storage backups
- Sorting and filtering work correctly with both strings and numbers

Also includes security improvement:
- Show setup command to user before copying to clipboard
- Prevents blindly pasting commands without review

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
2025-09-30 21:02:18 +00:00
rcourtman 386bee1aa6 fix: improve OIDC redirect URL validation and help text
addresses #327

Fixed issues when PUBLIC_URL is not set:
- Better error message explaining how to fix missing redirect URL
- Help text now shows actionable guidance instead of incomplete message
- Hide IdP redirect URL hint when no default is available
2025-09-30 20:33:11 +00:00
rcourtman 6a2b053c7a add URL routing for better navigation
Implements @solidjs/router to provide proper URL-based navigation:
- Main routes: /, /storage, /backups, /alerts, /settings
- Settings sub-routes: /settings/pve, /settings/pbs, /settings/system, etc.
- Browser back/forward buttons now work
- URLs are bookmarkable and shareable
- Clearer indication of current page in URL bar

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
2025-09-30 20:31:10 +00:00
rcourtman cffd90ac39 fix: storage node grouping not matching Node.ID format
Related to #478

Storage was grouping by storage.instance but Node.ID uses the format
"instance-nodename". This caused node headers to not display in storage
view when grouped by node.

Applied the same fix as Dashboard - group by instance-nodename to match
the Node.ID format from the backend.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
2025-09-30 20:19:24 +00:00
rcourtman 9a6f6eefd6 fix: grouped/list toggle not working in dashboard
addresses #478

The grouped view toggle wasn't working because guests were grouped by
guest.instance (e.g., "delly.lan") but node IDs are formatted as
"instance-nodename" (e.g., "delly.lan-delly"). This mismatch prevented
the node lookup from working, so grouped mode showed no node headers.

Changed the grouping key to match the Node.ID format by combining
guest.instance and guest.node.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
2025-09-30 20:15:34 +00:00
rcourtman 2a67ccbd7d remove CI workflow - redundant with local pre-commit hooks 2025-09-30 20:02:17 +00:00
rcourtman edb8702e77 fix CI errors: remove unused imports and format Go code
addresses unused TypeScript variables and gofmt formatting issues
2025-09-30 19:59:55 +00:00
rcourtman fd52a7add1 improve oidc error logging and documentation
addresses #327

- added detailed logging when ID token verification fails
- added better error messages for common OIDC issues
- updated docs with Authentik-specific configuration
- added troubleshooting section for redirect loops and invalid_id_token errors

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
2025-09-30 19:52:55 +00:00
rcourtman 4f74d6d9b0 standardize alert tab headers and styling
unified section headers across all alert tabs to match settings styling - all main headers now use size="md", subsection headers use consistent typography, added proper headers to history tab

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
2025-09-30 19:49:47 +00:00
rcourtman 9a2634a239 clean up authentication panel layout
Removed the messy metadata grid and restructured the layout:
- Two buttons in a clean horizontal row
- Username shown on the right
- Removed redundant metadata (last updated, coverage, etc)
- Simplified button styles
- 136 lines removed

Result: Much cleaner, more focused interface.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
2025-09-30 19:41:13 +00:00
rcourtman 37bdc383ea reduce verbose text in security settings
The Security settings page had way too much explanatory text.
Simplified all sections to be more concise:

- Security posture: shorter status descriptions
- Auth warning: compact banner instead of essay
- Authentication panel: removed redundant explanation
- OIDC panel: 4-line setup steps instead of paragraphs
- API token: removed verbose preamble
- Misc: removed duplicate SSO explanation box

Result: 150 lines of fluff removed, much cleaner UI.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
2025-09-30 19:38:36 +00:00
rcourtman 0d79a56d18 fix OIDC auto-redirect and clarify configuration
addresses #327

Changes:
- Removed automatic redirect to OIDC when password auth is configured
- Users can now choose between password auth or SSO button
- Clarified that client_secret is optional for PKCE-supporting providers
- Improved setup instructions to mention PUBLIC_URL requirement
- Made it clear that redirect URL is auto-generated from PUBLIC_URL

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
2025-09-30 19:34:57 +00:00
rcourtman 50036fde58 add temperature display and alerting to UI
- add dedicated Temperature column in node summary table with color coding
- add temperature threshold configuration in alert settings
- default threshold: 80°C trigger / 75°C clear
- temperature alerts integrated with existing alert system
- configurable per-node or globally via alert overrides

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
2025-09-30 19:31:51 +00:00
rcourtman 745c2b4c6b rebalance temperature monitoring messaging - reassuring but honest
Changed from scary warnings to confident, reassuring tone:

Before:
- "⚠️ IMPORTANT: This grants SSH access..."
- Emphasized risks and compromise scenarios
- Made users feel unsafe enabling the feature

After:
- "Works just like Ansible, Saltstack, etc."
- Emphasizes this is industry-standard approach
- Compares to trusted automation tools
- Focuses on what it does, not what could go wrong
- Still transparent about security model
- Removes duplicate/contradictory sections

The feature is secure and follows best practices. The messaging should
reflect confidence in the design while still being transparent.

Users should feel good about enabling it, not scared.
2025-09-30 19:16:30 +00:00
rcourtman d5bd6c7676 improve SSH setup security messaging for temperature monitoring
- Make it clear SSH setup is OPTIONAL
- Explain security model upfront before user commits
- Detail exactly what access is being granted (root SSH, sensors only)
- Warn users to only proceed if they trust Pulse server
- Better differentiate public vs private keys
- Show exactly where the key is stored
- Explain how to revoke access
- Add comprehensive security documentation
- Include advanced option for command restrictions in authorized_keys
- Add risk assessment and best practices

This ensures users make informed decisions about SSH access to their
critical Proxmox infrastructure.
2025-09-30 19:13:23 +00:00
rcourtman f14259915c docs: add temperature monitoring documentation 2025-09-30 19:11:12 +00:00
rcourtman d78c388cd0 add SSH key setup to auto-setup script for temperature monitoring
- Prompts user to set up SSH access during auto-setup
- Guides user to paste their Pulse server's public key
- Adds key to /root/.ssh/authorized_keys
- Installs lm-sensors automatically
- Runs sensors-detect --auto for proper sensor detection
- Optional: user can skip and set up later manually
- Includes validation of SSH key format
- Shows clear instructions for manual setup if skipped

This ensures temperature monitoring works out-of-the-box for users
who run the auto-setup script.
2025-09-30 19:10:45 +00:00
rcourtman 92b3bcd33c add node temperature monitoring via SSH
addresses #101

- Implement SSH-based temperature collector using lm-sensors
- Add Temperature struct to node models (CPU package, cores, NVMe)
- Collect temps during node polling (5s timeout, non-blocking)
- Display temperature in node cards with color coding:
  - Green: <60°C
  - Yellow: 60-80°C
  - Red: >80°C
- Shows CPU temp or falls back to load average if unavailable
- Tooltip includes NVMe drive temps when present
- Uses root SSH access (no additional auth setup needed for now)
- Temperature data only collected for online nodes
2025-09-30 19:08:31 +00:00
rcourtman f7842a0892 improve: add comprehensive debug logging for OIDC troubleshooting
Added detailed debug-level logs throughout the OIDC flow:
- Provider initialization (issuer, endpoints, scopes)
- Login flow tracking (client ID, redirect URL)
- Token exchange success/failure details
- Claims extraction (username, email, groups)
- Access control checks (why restrictions passed/failed)

Enhanced error logs to include issuer URL and actual error details in
audit events instead of generic "failed" messages.

Updated docs with Debug Logging section showing example output and
troubleshooting guidance for common issues like group restrictions.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
2025-09-30 18:45:30 +00:00
rcourtman 10177597a3 improve: show actual curl command after generating setup URL
Previously the quick token setup showed placeholder text like "click copy
to generate". Now it displays the actual curl command with the one-time
URL after generation, giving users transparency about what they're running.

Changes:
- Display generated curl command with URL in code block
- Update prompt text for better clarity ("above" instead of generic)
- Add visual feedback with color states (gray/blue/green)
- Store URL in setupCode state for display
- Apply to both PVE and PBS setup flows

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
2025-09-30 18:26:57 +00:00
rcourtman 113c20ffe6 fix: prevent silent encryption key regeneration that orphans data
addresses data loss issue where encryption key regeneration silently
orphaned all encrypted configuration (nodes, email, webhooks).

Changes:
- Check for existing .enc files before generating new encryption key
- Refuse to start if encrypted data exists but key is missing/invalid
- Forces explicit user action (restore key backup or delete .enc files)
- Prevents silent data loss from key regeneration

This ensures encrypted data is never accidentally orphaned when the
encryption key is lost or corrupted.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
2025-09-30 18:23:24 +00:00
rcourtman 6d35c210be fix: expand timezone list in quiet hours configuration
addresses #477

Expanded the timezone dropdown from 11 options to 70+ common IANA
timezones covering all major regions (Africa, Americas, Asia,
Australia, Europe, Pacific).

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
2025-09-30 18:13:15 +00:00
rcourtman 88915dce4c chore: bump version to v4.15.1 2025-09-30 17:31:28 +00:00
rcourtman 3f9748dc1f fix: handle non-numeric wearout values for HDDs and RAID controllers
addresses #449

proxmox returns 'N/A' or empty string for the wearout field on disks that
don't support wear reporting (HDDs, hardware RAID controllers, etc). pulse
was expecting an integer, causing JSON unmarshal errors that prevented ALL
disks from being displayed on affected nodes.

added custom UnmarshalJSON method for the Disk type to gracefully handle:
- numeric values (SSDs with wear reporting)
- string values like 'N/A' (HDDs, RAID controllers) - converts to 0
- null values - converts to 0

this allows nodes with mixed disk types (SSDs, HDDs, RAID) to display all
their disks correctly. wearout value of 0 indicates no wear reporting
available, which is expected for HDDs.
2025-09-30 17:19:15 +00:00
rcourtman 22fb25ac00 fix: send notifications for critical alerts restored from disk after restart
addresses #471

when pulse restarts (service restart, container restart, etc), active alerts
are loaded from disk but notifications were never sent for these restored
alerts. this caused users to miss critical ongoing alerts that existed before
the restart.

the issue was particularly noticeable with memory alerts on VMs - if a VM's
memory was genuinely high and an alert was created, then pulse restarted, the
alert would show in 'Active Alerts' but no webhook notification would be sent.
however, manually creating a 'fake' alert by lowering thresholds would work
because those are new alerts.

fix: now sends notifications for restored critical alerts that started within
the last 2 hours. adds a 10-second delay after restart to allow the system
to stabilize before sending notifications. warning-level alerts are not
re-notified to avoid spam on restart.
2025-09-30 17:13:53 +00:00
rcourtman a90e5b90cc fix: correctly display PBS host backups as 'Host' type instead of 'LXC'
addresses #454

PBS supports backing up physical hosts using proxmox-backup-client with
backup-type='host'. these backups were incorrectly displayed as 'LXC' type
because the frontend only checked for 'vm'/'qemu' types and defaulted
everything else to 'LXC'.

now properly handles all three PBS backup types: vm, ct (lxc), and host
2025-09-30 16:56:03 +00:00
rcourtman cf25415a0a fix: resolve grouping issues when nodes have duplicate hostnames
addresses #476

when multiple nodes have the same hostname, the dashboard and storage views
were incorrectly grouping VMs/containers/storage by hostname instead of by
unique node instance ID. this caused:
- incorrect VM/container counts in node summary
- mixed display of resources from different nodes
- incorrect grouping in storage view

changed grouping logic to use guest.instance (unique node ID) instead of
guest.node (hostname). updated both Dashboard and Storage components to
properly map instance IDs to node objects for display while maintaining
correct data separation.
2025-09-30 16:50:51 +00:00
rcourtman 334dbc490d chore: bump version to v4.15.0 2025-09-30 16:21:43 +00:00
rcourtman f9b8037486 fix: resolve install script unbound variables and add update command
Addresses #450, #451, #406

- Initialize all variables at top of script to prevent "unbound variable" errors with set -u
  - BUILD_FROM_SOURCE, SKIP_DOWNLOAD, IN_CONTAINER, IN_DOCKER now set at line 20-27
  - ENABLE_AUTO_UPDATES, FORCE_VERSION, FORCE_CHANNEL, SOURCE_BRANCH also moved to top
  - Removed duplicate assignments from argument parsing section

- Restore /bin/update command creation for ProxmoxVE LXC installations
  - Creates update script that re-runs install.sh for easy updates
  - Allows backend to properly detect ProxmoxVE deployment type
  - Users can now run "update" in LXC console as documented

- Update deployment detection to recognize install.sh in update command
  - Previously only looked for legacy "pulse.sh" reference
  - Now checks for both pulse.sh and install.sh
2025-09-30 16:16:10 +00:00
rcourtman 413ef73953 improve webhook system security and robustness
addresses security vulnerabilities and improves webhook reliability

Changes:
- Add SSRF protection with redirect controls and strict URL validation
- Add response size limits (1MB cap) to prevent memory exhaustion
- Fix race condition in SendTestNotification
- Add per-webhook rate limiting (10 req/min)
- Add Retry-After header support for proper backoff
- Extract magic numbers to configurable constants
- Block localhost, link-local, and cloud metadata endpoints
- Add secure HTTP client with redirect validation
- Remove duplicate function definitions
- Clean up unused code

Security improvements:
- Prevents SSRF attacks via redirect chains
- Protects against DoS via large responses
- Rate limits prevent webhook flooding
- Thread-safe webhook operations

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
2025-09-30 15:57:28 +00:00
rcourtman 552173b262 fix: improve alert system robustness and security
Addresses multiple issues identified during comprehensive alert system audit:

1. Fix ZFS device loop lock issue
   - Moved lock acquisition outside loop in checkZFSPoolHealth
   - Changed clearAlert to clearAlertNoLock when lock already held
   - Prevents multiple lock acquisitions in same iteration

2. Add alert deduplication on restore
   - Prevents duplicate alerts after service restart
   - Tracks seen alert IDs during LoadActiveAlerts
   - Logs warnings for any duplicates found

3. Add API input validation
   - validateAlertID function prevents DOS attacks
   - Limit alert ID length to 500 characters
   - Whitelist allowed characters (alphanumeric, -, _, :, /, .)
   - Cap history limit parameter at 10,000 records
   - Applied validation to acknowledge, unacknowledge, and clear endpoints

4. Add panic recovery to goroutines
   - All SaveActiveAlerts goroutines now have defer/recover
   - Cleanup goroutines protected from panics
   - Contextual error logging for each goroutine type

5. Document lock ordering
   - Added comprehensive documentation for Manager mutexes
   - Explains m.mu and resolvedMutex relationship
   - Clarifies acquisition rules to prevent deadlocks
   - Inline comments for resolvedMutex field

These fixes improve stability, security, data integrity, and maintainability
of the alert system without breaking API compatibility.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
2025-09-30 15:35:39 +00:00