mirror of
https://github.com/GoodOlClint/PSProxmoxVE.git
synced 2026-09-03 18:55:33 +00:00
def8dc6b67
* fix: verify checksums for downloaded ISOs/images, keep sshpass off argv ensure-base-iso.sh downloaded the PVE install ISO over plain HTTP with no checksum, caching it on the persistent /opt/pve-integration mount and booting it as the nested trust root the integration suite relies on. ensure-cloud-images.sh fetched the Ubuntu cloud image and OVA over HTTPS but never checked them either. prepare-test-environment.sh and diagnose-cluster.sh passed the nested root password to sshpass via -p, putting it in the process table. create-api-token.sh, unused anywhere in the repo, minted a privsep=0 root token and echoed the secret unmasked. - ensure-base-iso.sh now downloads from https://enterprise.proxmox.com/iso and verifies against its SHA256SUMS on every run, including a cache hit. download.proxmox.com's own TLS cert does not list download.proxmox.com in its SAN (confirmed with curl/openssl from this environment), so https to that name fails certificate validation; enterprise.proxmox.com serves the identical ISO tree over a valid cert. Verification happens before the downloaded file is moved to its canonical cache path. - ensure-cloud-images.sh verifies the cloud image and OVA against Ubuntu's published SHA256SUMS the same way, matching by upstream filename since the cloud image is cached locally under a different extension (.img upstream, .qcow2 cached — the bytes are already qcow2-formatted). - prepare-test-environment.sh and diagnose-cluster.sh now export SSHPASS and call sshpass -e, keeping the password out of argv/ps. This also fixes a latent bug: the old unquoted `sshpass -p ${ROOT_PASS}` word-split any password containing whitespace. - create-api-token.sh deleted; grep across the repo found no caller. Reviewers (codex:codex-rescue, correctness-reviewer, security-reviewer) all independently found the same blocking bug in the first pass: when a cached file failed verification and the subsequent redownload then failed, ensure-cloud-images.sh fell through to a "keep the stale copy" branch and returned that same known-bad file with exit 0 — verification could be bypassed by inducing one failed redownload. Fixed by deleting the file immediately on a failed verification, before the redownload is attempted, so the later "is there a safe stale copy" check can no longer find it. Added a test case (case 5) that reproduces this exact sequence and mutation-tested it against the unfixed code. The three reviews also flagged a real but separate bug already fixed in this same change: `trap ... RETURN` inside a function nested in another function is not scoped to that function in bash — it re-fires on the OUTER function's return, referencing an out-of-scope local. Both verify_checksum() helpers now clean up their temp file explicitly instead of via trap. Findings not acted on, judged out of scope for this fix: - SHA256SUMS-fetch failures are treated the same as a checksum mismatch (delete + fail) rather than left untouched — a transient network blip destroys a good multi-GB cached ISO. This is the safer failure direction (never silently trust unverified bytes) and was a deliberate trade-off, not a defect. - ensure-cloud-images.sh's 7-day cache window can span an upstream republish of noble/current, causing a legitimate re-verification churn (not a security issue, a cache-hit-rate one). Pre-existing cache design, unrelated to adding verification. - wait-for-pve.sh (curl -d with the password on argv) and prepare-test-environment.sh's own positional password argument (from run-integration.sh) carry the same password-on-argv pattern this issue targeted in create-api-token.sh, sshpass -p and diagnose-cluster.sh, but neither script nor run-integration.sh was named in the issue. Left untouched per scope; worth a follow-up issue. - GPG/detached-signature verification of the upstream SHA256SUMS was not added — the new checks defend against cache poisoning and transit corruption, not a compromised origin. Worth a follow-up issue. - The two new self-checks (ensure-base-iso.test.sh, ensure-cloud-images.test.sh) are not wired into .github/workflows/unit-tests.yml's shell-selfchecks job. That file is code-owned and out of scope for this change; needs an operator follow-up. Password rotation (the Testpass123! value from before it moved to a secret) is unaddressed here per the contract — flagged for the operator. Mutation-tested: broke the post-download checksum check in ensure-base-iso.sh, confirmed the affected test cases failed, restored it. Broke the sshpass -e change back to -p, confirmed the new assertions in prepare-test-environment.test.sh failed, restored it. Broke the fail-open fix in ensure-cloud-images.sh, confirmed case 5 failed, restored it. Closes #149 * fix: also verify the stale-by-age fallback copy in ensure-cloud-images.sh PR review on #166 (COMMENTED, non-blocking) found the sibling of the fail-open bug already fixed in this branch: when the cached cloud image is stale by *age* (>= 7 days) rather than failed verification, the redownload-failure fallback could hand back that file with exit 0 without ever re-verifying it in this run. A file that failed the earlier verification is already deleted by the time the fallback runs, but a stale-by-age file skips verification entirely on the way in. Fixed by verifying the stale-by-age file at the point of actual fallback use — after the redownload has failed, not proactively before it's attempted, so a copy the redownload was about to replace anyway isn't deleted along a path that would have succeeded. Added two test cases (6, 7): a still-verifying stale-by-age copy is used as a fallback; one that no longer verifies is not. Mutation-tested by reverting to the unfixed fallback and confirming case 7 fails, then restored. --------- Co-authored-by: goodolclint-claude[bot] <323206664+goodolclint-claude[bot]@users.noreply.github.com>
113 lines
4.5 KiB
Bash
Executable File
113 lines
4.5 KiB
Bash
Executable File
#!/usr/bin/env bash
|
|
# Prepares the test environment on the nested PVE node.
|
|
# Only performs operations that have no PVE API equivalent.
|
|
#
|
|
# Usage: prepare-test-environment.sh <nested-pve-ip> <root-password> [dist-upgrade] [pkg-out]
|
|
#
|
|
# Operations:
|
|
# - Optionally dist-upgrade the node, record its package set, and reboot
|
|
# (currency lane only; off unless <dist-upgrade> is 1)
|
|
# - Enable snippets+import content types on local storage (pvesm set)
|
|
# - Upload cloud-init user-data snippet (SCP — no snippet upload API)
|
|
|
|
set -euo pipefail
|
|
|
|
NESTED_IP="${1:?Usage: prepare-test-environment.sh <ip> <password> [dist-upgrade] [pkg-out]}"
|
|
ROOT_PASS="$2"
|
|
DIST_UPGRADE="${3:-0}"
|
|
PKG_OUT="${4:-}"
|
|
|
|
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
|
|
|
SSH_OPTS="-o StrictHostKeyChecking=no -o UserKnownHostsFile=/dev/null -o LogLevel=ERROR"
|
|
export SSHPASS="${ROOT_PASS}"
|
|
SSH_CMD="sshpass -e ssh ${SSH_OPTS} root@${NESTED_IP}"
|
|
SCP_CMD="sshpass -e scp ${SSH_OPTS}"
|
|
|
|
echo "=== Preparing test environment on ${NESTED_IP} ==="
|
|
|
|
if [[ "${DIST_UPGRADE}" == "1" ]]; then
|
|
echo "Running dist-upgrade (currency lane)..."
|
|
${SSH_CMD} "DEBIAN_FRONTEND=noninteractive apt-get update -qq && \
|
|
DEBIAN_FRONTEND=noninteractive apt-get -y \
|
|
-o Dpkg::Options::=--force-confold \
|
|
-o Dpkg::Options::=--force-confdef \
|
|
dist-upgrade"
|
|
|
|
if [[ -n "${PKG_OUT}" ]]; then
|
|
echo "Recording package set to ${PKG_OUT}..."
|
|
# pipefail must be set in the REMOTE shell. ssh returns the remote
|
|
# pipeline's status, which is sort's — and sort succeeds on the empty
|
|
# input a failed dpkg-query produces, so a broken query would otherwise
|
|
# leave a zero-byte file and still exit 0.
|
|
${SSH_CMD} "set -o pipefail; dpkg-query -W -f='\${binary:Package}\t\${Version}\n' | sort" > "${PKG_OUT}"
|
|
if [[ ! -s "${PKG_OUT}" ]]; then
|
|
echo "ERROR: empty package set from ${NESTED_IP}" >&2
|
|
exit 1
|
|
fi
|
|
fi
|
|
|
|
# A PVE dist-upgrade pulls proxmox-kernel-*; without a reboot the node runs
|
|
# new userspace on the old kernel. Reboot unconditionally rather than
|
|
# testing /var/run/reboot-required — that file comes from
|
|
# update-notifier-common, which is not guaranteed on a PVE node.
|
|
#
|
|
# boot_id is the evidence that the reboot happened. Without it the `|| true`
|
|
# below swallows every ssh failure, the node stays up, and wait-for-api.sh
|
|
# matches the still-running pre-reboot pveproxy on its first poll.
|
|
boot_before="$(${SSH_CMD} "cat /proc/sys/kernel/random/boot_id")"
|
|
|
|
echo "Rebooting after dist-upgrade..."
|
|
${SSH_CMD} "systemctl reboot" || true
|
|
|
|
# Order matters: prove the reboot first (ssh returns before pveproxy does),
|
|
# then wait for the API, then for pmxcfs.
|
|
boot_after=""
|
|
for _ in $(seq 1 60); do
|
|
boot_after="$(${SSH_CMD} "cat /proc/sys/kernel/random/boot_id" 2>/dev/null || true)"
|
|
[[ -n "${boot_after}" && "${boot_after}" != "${boot_before}" ]] && break
|
|
sleep 5
|
|
done
|
|
if [[ -z "${boot_after}" || "${boot_after}" == "${boot_before}" ]]; then
|
|
echo "ERROR: ${NESTED_IP} did not reboot (boot_id unchanged)" >&2
|
|
exit 1
|
|
fi
|
|
|
|
bash "${SCRIPT_DIR}/wait-for-api.sh" "${NESTED_IP}" 8006 600
|
|
|
|
# wait-for-api.sh only proves pveproxy answers. `pvesm set` below writes
|
|
# /etc/pve/storage.cfg, which needs pmxcfs to have mounted /etc/pve — on a
|
|
# freshly rebooted node those are seconds apart.
|
|
for _ in $(seq 1 30); do
|
|
${SSH_CMD} "test -f /etc/pve/storage.cfg" 2>/dev/null && break
|
|
sleep 5
|
|
done
|
|
|
|
# printf, not echo: bash's builtin echo does not interpret \t without -e,
|
|
# which would make this the one row in the file without a real tab.
|
|
if [[ -n "${PKG_OUT}" ]]; then
|
|
${SSH_CMD} "printf '# running-kernel\t%s\n' \"\$(uname -r)\"" >> "${PKG_OUT}"
|
|
fi
|
|
fi
|
|
|
|
# Enable snippets and import content types on local storage
|
|
echo "Configuring local storage content types..."
|
|
${SSH_CMD} "mkdir -p /var/lib/vz/snippets && pvesm set local --content images,iso,vztmpl,snippets,import"
|
|
|
|
# Upload cloud-init user-data snippet (no API for snippet upload)
|
|
echo "Uploading cloud-init user-data snippet..."
|
|
USERDATA=$(mktemp)
|
|
cat > "${USERDATA}" <<'YAML'
|
|
#cloud-config
|
|
package_update: true
|
|
packages:
|
|
- qemu-guest-agent
|
|
runcmd:
|
|
- systemctl enable --now qemu-guest-agent
|
|
YAML
|
|
|
|
${SCP_CMD} "${USERDATA}" "root@${NESTED_IP}:/var/lib/vz/snippets/test-vm-userdata.yml"
|
|
rm -f "${USERDATA}"
|
|
|
|
echo "Environment preparation complete."
|