Files
PSProxmoxVE/tests/infrastructure/scripts/diagnose-cluster.sh
T
goodolclint-claude[bot] def8dc6b67 fix: verify checksums for downloaded ISOs/images, keep sshpass off argv (#166)
* fix: verify checksums for downloaded ISOs/images, keep sshpass off argv

ensure-base-iso.sh downloaded the PVE install ISO over plain HTTP with no
checksum, caching it on the persistent /opt/pve-integration mount and
booting it as the nested trust root the integration suite relies on.
ensure-cloud-images.sh fetched the Ubuntu cloud image and OVA over HTTPS
but never checked them either. prepare-test-environment.sh and
diagnose-cluster.sh passed the nested root password to sshpass via -p,
putting it in the process table. create-api-token.sh, unused anywhere in
the repo, minted a privsep=0 root token and echoed the secret unmasked.

- ensure-base-iso.sh now downloads from https://enterprise.proxmox.com/iso
  and verifies against its SHA256SUMS on every run, including a cache hit.
  download.proxmox.com's own TLS cert does not list download.proxmox.com in
  its SAN (confirmed with curl/openssl from this environment), so https to
  that name fails certificate validation; enterprise.proxmox.com serves the
  identical ISO tree over a valid cert. Verification happens before the
  downloaded file is moved to its canonical cache path.
- ensure-cloud-images.sh verifies the cloud image and OVA against Ubuntu's
  published SHA256SUMS the same way, matching by upstream filename since
  the cloud image is cached locally under a different extension (.img
  upstream, .qcow2 cached — the bytes are already qcow2-formatted).
- prepare-test-environment.sh and diagnose-cluster.sh now export SSHPASS
  and call sshpass -e, keeping the password out of argv/ps. This also fixes
  a latent bug: the old unquoted `sshpass -p ${ROOT_PASS}` word-split any
  password containing whitespace.
- create-api-token.sh deleted; grep across the repo found no caller.

Reviewers (codex:codex-rescue, correctness-reviewer, security-reviewer) all
independently found the same blocking bug in the first pass: when a cached
file failed verification and the subsequent redownload then failed,
ensure-cloud-images.sh fell through to a "keep the stale copy" branch and
returned that same known-bad file with exit 0 — verification could be
bypassed by inducing one failed redownload. Fixed by deleting the file
immediately on a failed verification, before the redownload is attempted,
so the later "is there a safe stale copy" check can no longer find it.
Added a test case (case 5) that reproduces this exact sequence and
mutation-tested it against the unfixed code. The three reviews also
flagged a real but separate bug already fixed in this same change: `trap
... RETURN` inside a function nested in another function is not scoped to
that function in bash — it re-fires on the OUTER function's return,
referencing an out-of-scope local. Both verify_checksum() helpers now
clean up their temp file explicitly instead of via trap.

Findings not acted on, judged out of scope for this fix:
- SHA256SUMS-fetch failures are treated the same as a checksum mismatch
  (delete + fail) rather than left untouched — a transient network blip
  destroys a good multi-GB cached ISO. This is the safer failure direction
  (never silently trust unverified bytes) and was a deliberate trade-off,
  not a defect.
- ensure-cloud-images.sh's 7-day cache window can span an upstream
  republish of noble/current, causing a legitimate re-verification churn
  (not a security issue, a cache-hit-rate one). Pre-existing cache design,
  unrelated to adding verification.
- wait-for-pve.sh (curl -d with the password on argv) and
  prepare-test-environment.sh's own positional password argument (from
  run-integration.sh) carry the same password-on-argv pattern this issue
  targeted in create-api-token.sh, sshpass -p and diagnose-cluster.sh, but
  neither script nor run-integration.sh was named in the issue. Left
  untouched per scope; worth a follow-up issue.
- GPG/detached-signature verification of the upstream SHA256SUMS was not
  added — the new checks defend against cache poisoning and transit
  corruption, not a compromised origin. Worth a follow-up issue.
- The two new self-checks (ensure-base-iso.test.sh,
  ensure-cloud-images.test.sh) are not wired into
  .github/workflows/unit-tests.yml's shell-selfchecks job. That file is
  code-owned and out of scope for this change; needs an operator follow-up.

Password rotation (the Testpass123! value from before it moved to a
secret) is unaddressed here per the contract — flagged for the operator.

Mutation-tested: broke the post-download checksum check in
ensure-base-iso.sh, confirmed the affected test cases failed, restored it.
Broke the sshpass -e change back to -p, confirmed the new assertions in
prepare-test-environment.test.sh failed, restored it. Broke the fail-open
fix in ensure-cloud-images.sh, confirmed case 5 failed, restored it.

Closes #149

* fix: also verify the stale-by-age fallback copy in ensure-cloud-images.sh

PR review on #166 (COMMENTED, non-blocking) found the sibling of the
fail-open bug already fixed in this branch: when the cached cloud image
is stale by *age* (>= 7 days) rather than failed verification, the
redownload-failure fallback could hand back that file with exit 0
without ever re-verifying it in this run. A file that failed the
earlier verification is already deleted by the time the fallback runs,
but a stale-by-age file skips verification entirely on the way in.

Fixed by verifying the stale-by-age file at the point of actual
fallback use — after the redownload has failed, not proactively before
it's attempted, so a copy the redownload was about to replace anyway
isn't deleted along a path that would have succeeded. Added two test
cases (6, 7): a still-verifying stale-by-age copy is used as a
fallback; one that no longer verifies is not. Mutation-tested by
reverting to the unfixed fallback and confirming case 7 fails, then
restored.

---------

Co-authored-by: goodolclint-claude[bot] <323206664+goodolclint-claude[bot]@users.noreply.github.com>
2026-09-02 17:20:36 +00:00

108 lines
3.8 KiB
Bash

#!/usr/bin/env bash
# Dump corosync state from both nested PVE nodes after a cluster test failure.
#
# Usage: diagnose-cluster.sh [8|9]
#
# The PVE API reports a joined-but-offline node as online=0 with no further
# detail; corosync's own view lives only on the nodes, which the cleanup job
# destroys minutes later. Best-effort: never fails the caller.
#
# Required env vars:
# PVE_PASSWORD Root password for the nested PVE instances
#
# Optional env vars:
# CONFIG_FILE Test config JSON (default: $CACHE_DIR/work/config.json)
# CACHE_DIR Shared cache mount (default: /opt/pve-integration)
VERSION="${1:-9}"
CACHE_DIR="${CACHE_DIR:-/opt/pve-integration}"
CONFIG_FILE="${CONFIG_FILE:-$CACHE_DIR/work/config.json}"
if [[ ! -f "$CONFIG_FILE" ]]; then
echo "diagnose-cluster: no config at $CONFIG_FILE — nothing to inspect"
exit 0
fi
if [[ -z "${PVE_PASSWORD:-}" ]]; then
echo "diagnose-cluster: PVE_PASSWORD unset — cannot reach the nodes"
exit 0
fi
SSH_OPTS=(-o StrictHostKeyChecking=no -o UserKnownHostsFile=/dev/null -o LogLevel=ERROR -o ConnectTimeout=10)
export SSHPASS="$PVE_PASSWORD"
dump_node() {
local label="$1" ip="$2"
echo
echo "══════════ $label ($ip) ══════════"
if [[ -z "$ip" || "$ip" == "null" ]]; then
echo " no address in $CONFIG_FILE"
return
fi
sshpass -e ssh "${SSH_OPTS[@]}" "root@${ip}" bash -s <<'REMOTE' 2>&1 || echo " ssh to $ip failed (rc=$?)"
set +e
echo "--- hostname / resolution ---"
hostname -f
echo "hostname -i: $(hostname -i 2>&1)"
grep -vE '^\s*#' /etc/hosts | grep -vE '^\s*$'
echo
echo "--- addresses ---"
ip -4 -o addr show scope global
echo
echo "--- pmxcfs mode: cluster or local? ---"
# /etc/pve/corosync.conf is database-backed; pmxcfs only creates it when it
# starts with no config.db and imports /etc/corosync/corosync.conf. A surviving
# standalone config.db means silent local mode with corosync otherwise healthy.
dpkg-query -W pve-cluster corosync 2>&1
tr '\0' ' ' < "/proc/$(systemctl show pve-cluster -p MainPID --value)/cmdline" 2>&1; echo
findmnt --target /etc/pve --output TARGET,SOURCE,FSTYPE 2>&1
echo ".members: $(cat /etc/pve/.members 2>&1 | tr -d '\n')"
ls -la /var/lib/pve-cluster/ 2>&1
ls -la /var/lib/pve-cluster/backup/ 2>&1
if command -v sqlite3 >/dev/null 2>&1; then
sqlite3 -readonly /var/lib/pve-cluster/config.db \
"PRAGMA quick_check; SELECT name,version,writer,mtime,length(data) FROM tree WHERE name='corosync.conf';" 2>&1
else
echo "sqlite3 absent; corosync.conf occurrences in config.db: $(strings /var/lib/pve-cluster/config.db 2>/dev/null | grep -c '^corosync\.conf$')"
fi
echo
echo "--- corosync-cpgtool (pmxcfs joins dcdb/status CPG groups when clustered) ---"
corosync-cpgtool 2>&1
echo
echo "--- corosync.conf ---"
cat /etc/pve/corosync.conf 2>&1 || cat /etc/corosync/corosync.conf 2>&1
echo
echo "--- corosync-cfgtool -s ---"
corosync-cfgtool -s 2>&1
echo
echo "--- pvecm status ---"
pvecm status 2>&1
echo
echo "--- corosync service ---"
systemctl is-active corosync pve-cluster 2>&1
echo
echo "--- journalctl -u corosync (last 60) ---"
journalctl -u corosync -n 60 --no-pager 2>&1
echo
echo "--- journalctl -u pve-cluster (last 30) ---"
journalctl -u pve-cluster -n 30 --no-pager 2>&1
echo
echo "--- cluster task logs ---"
# "Cluster join aborted!" is generic; the reason is only in the task log.
find /var/log/pve/tasks -type f \( -name '*clusterjoin*' -o -name '*clustercreate*' \) \
-exec echo "== {} ==" \; -exec cat {} \; 2>&1 | tail -80
REMOTE
}
echo "=== Cluster diagnostics for PVE $VERSION ==="
node_a="$(jq -r ".pve${VERSION}.nodes.a.host // empty" "$CONFIG_FILE")"
node_b="$(jq -r ".pve${VERSION}.nodes.b.host // empty" "$CONFIG_FILE")"
dump_node "node A" "$node_a"
dump_node "node B" "$node_b"
echo
echo "=== End cluster diagnostics ==="
exit 0