Files
k7/CHALLENGES.md
T
G 13fafe0997 Release 0.4.0
Per-key node pins, k7d-fc pause/resume/exec, and HA-soak fixes. Playbook
pins k7d 0.7.0. GitHub .deb, Launchpad PPA, and PyPI k7-sdk are 0.4.0.
2026-09-19 23:22:14 +02:00

38 KiB
Raw Blame History

Challenges & solutions log

Tracking non-obvious bugs in k7 (the sandbox management layer) and how they were solved. Same format as k7d-dev's CHALLENGES.md.

1. k7 install --backend extra-var silently clobbered inventory k7_backends (spec 18e)

Symptom: A 3-node HA install with an inventory declaring k7_backends=kfd,kql,k7d per host would have provisioned only the two Kata backends — the k7d backend would silently disappear from every node.

Root cause: k7 install always forwarded k7_backends as an Ansible extra-var (built from the --backend option's default value even when the user never passed --backend). Extra-vars have the highest precedence in Ansible, so the CLI default overrode the per-host inventory var.

Fix: install() now checks ctx.get_parameter_source("backend"); when the user supplied -i <inventory> without an explicit --backend, the k7_backends extra-var is dropped so the inventory wins. An explicit --backend alongside -i prints a warning that it overrides the inventory.

Reference: none (Ansible variable-precedence rules).

Time lost: caught in pre-flight review (spec 18e Phase 0), ~1 hour of code reading. Would have cost a full reset+reinstall cycle if it had shipped.


2. Longhorn StorageClass numberOfReplicas parameter makes default-replica-count a no-op (spec 18e)

Symptom: bench_docker_perf.py's r1/r2 legs patched Longhorn's default-replica-count setting — but sandbox volumes kept the replica count baked into the longhorn StorageClass (numberOfReplicas: "3" on the HA cluster). The r1/r2/r3 benches would all have silently measured the same replica count.

Root cause: Longhorn only consults the default-replica-count setting when the StorageClass has no numberOfReplicas parameter. k7's install playbook pins the parameter in the SC (topology-aware SC from the longhorn-storageclass ConfigMap), so the setting never applies to k7 volumes.

Fix: the bench now patches spec.numberOfReplicas on the sandbox's Longhorn Volume CRs after creation (_set_sandbox_volume_replicas), waits until exactly N running replicas exist, and records the replica → node placement in the log header as proof.

Reference: Longhorn docs (volume-level replica count update).

Time lost: ~1 hour. The old benches "passed" — the mismeasurement was invisible without checking actual replica CRs.


3. k7d VM operations are node-local; multi-node scheduler breaks pause/fork tests (spec 18e)

Symptom: On the 3-node cluster, test_k7d.py pause/fork tests failed with sandbox X runs on node k7-node-02, but this k7 process runs on k7-node-01; k7d VM operations must run on the pod's node.

Root cause: k7d pause/resume/fork go through the node-local /run/k7d/k7d.sock; cross-node VM ops are explicitly out of scope (k7d spec 9a M12) and core fails loudly on the mismatch. On a single-node cluster the tests never noticed; with 3 schedulable nodes the sandbox lands anywhere.

Fix: tests that exercise VM ops pin their sandboxes with SandboxConfig(node_name=os.uname().nodename) — the field that exists precisely for host-side-inspection tests. The centralized-API implication (k7-api pod can only pause/fork k7d sandboxes co-located on the first master) is recorded as a release-readiness limitation.

Reference: k7d spec 9a M12 (cross-node fork out of scope).

Time lost: ~30 min (the loud error message made it easy).


4. kql fork loses guest writes made just before the fork (crash-consistency)

Symptom: test_api.py::test_sdk_pause_resume_fork_round_trip flaked on the HA cluster: a file written via exec seconds before fork() did not exist in the fork (cat: can't open '/mnt/state/marker'). The near-identical test_qemu.py::test_fork_clones_data data check passed in the same run — pure timing luck.

Root cause: the kql fork path cuts a block-level Longhorn VolumeSnapshot of the source's root PVC. That is only crash-consistent: guest writes still sitting in the VM's page cache are not on the block device yet and are missing from the clone.

Fix: fork_sandbox now execs sync in the source sandbox before creating the snapshot (only when the deployment has ready replicas — a paused source has no writers), failing loudly if the flush fails.

Reference: none (standard crash-vs-application consistency).

Time lost: ~1 hour including re-runs. Note: named snapshots of live sandboxes (k7 snapshot) remain crash-consistent by design — documented behavior, unchanged.


5. NVMe enumeration swaps across reboots — hardcoded k7_devmapper_disk hit the OS disk

Symptom: The second full reset+reinstall loop of spec 18e failed on k7-node-03: Device '/dev/nvme1n1' has partitions; wipe it with utils/wipe-disk.sh or choose another disk. The identical inventory had just worked on the first loop.

Root cause: Linux NVMe controller enumeration (nvme0n1 vs nvme1n1) is not stable across reboots. After the second reset, node-03 booted with its OS on the disk now enumerated nvme1n1, and the raw spare as nvme0n1 — the inventory's hardcoded k7_devmapper_disk=/dev/nvme1n1 pointed at the OS disk. The playbook's safety checks caught it (fail-loud worked as designed).

Fix: omit k7_devmapper_disk on identical dual-NVMe boxes — the playbook's auto-detect ("first empty, non-removable, non-root whole disk") is enumeration-proof. inventory.ini.example now documents this.

Reference: none (kernel device-naming behavior).

Time lost: ~30 min (one wasted install attempt + one extra reset+reinstall loop of all three nodes).


6. Orphaned Firecracker microVMs leak on pod deletion and each burns a full CPU core

Symptom: During spec 18e Phase 3, the k7-ql-r2 bench leg started failing mid-run with Pod is not running (status: Pending) and the k7-ql-r3 leg failed entirely; longhorn-manager / cilium-envoy / coredns readiness probes were flapping cluster-wide. k7-node-03 had a load average of ~21.

Root cause: 14 orphaned /firecracker processes (2 on node-01, 2 on node-02, 10 on node-03) whose pods had been deleted hours earlier — zero live Kata pods existed cluster-wide. Each orphan spun at ~97% CPU (TIME ≈ ETIME in ps), starving Longhorn/Cilium/CoreDNS and the bench sandbox itself. The kata-fc shim intermittently fails to kill the microVM on pod deletion under parallel pod churn (~14 leaks over ~30 kfd pod deletions that day). Evidence: /tmp/leaked-firecracker-vms.txt (agent run artifact).

Fix (remediation): verified no live Kata pods, then pkill -9 firecracker on all three nodes; loads recovered and the r2/r3 bench legs were re-run green. Root-cause fix still open — tracked as a release blocker in spec 18f-release-blockers (investigate containerd-shim-kata-v2 / jailer cleanup path; add a leak-detection integration test that asserts zero firecracker processes after suite teardown).

Reference: none yet (kata-containers shim lifecycle).

Time lost: ~1.5 hours (failed bench legs + diagnosis + re-run).


7. Remote test loop tied to the SSH session died mid-run (Broken pipe)

Symptom: A multi-suite pytest loop launched over plain ssh host 'for f in ...; do pytest ...; done' died silently when the SSH connection dropped (client_loop: send disconnect: Broken pipe) — the remote shell got SIGHUP'd between suites.

Root cause: the remote loop was a child of the SSH session; NAT idle timeouts kill long-lived connections even with keepalives.

Fix: write the loop to a script on the node and launch it with setsid nohup ... < /dev/null &, then poll a progress file. (Same class of issue utils/run-integration-tests.sh already documents for its keepalive settings.)

Reference: none.

Time lost: ~20 min (one interrupted suite sequence, test_restore had finished right before the drop).


8. Firecracker leak root cause: jailer --daemonize makes the kata shim signal a dead PID (spec 18f)

Symptom: Follow-up to #6. Reproduced at will on the 18f run: create a naked kfd pod (sleep workload), delete it — the pod terminates cleanly but its /firecracker process survives with PPID 1 and climbs to ~100% CPU. Two out of two attempts leaked. Shim logs at 18e leak time showed Agent did not stop sandbox: Dead agent + failed to ping agent: CheckRequest timed out.

Root cause: kata 3.24.0 virtcontainers/fc.go. When jailed (spec 8a enabled the jailer), fcInit launches jailer --daemonize, which double-forks — firecracker reparents to init immediately, and fc.info.PID = cmd.Process.Pid records the jailer's PID, which is already dead. fcEnd() then calls WaitLocalProcess(pid, …, SIGTERM) on that stale PID: a no-op. The VMM normally exits because the in-guest agent shuts the VM down; whenever that graceful path fails (dead/hung agent under churn, wedged guest IO), nothing ever kills the firecracker process. The getting vm status failed … firecracker.socket: no such file or directory error seen at every kfd VM boot is a side effect of the same daemonize handling (the shim polls the jailed API socket path before it exists) and is harmless noise.

Fix: upstream fix belongs in kata (record the real VMM PID when jailed). In k7: (a) k7 install now deploys a per-node systemd timer k7-vmm-reaper.timer (1 min cadence) that SIGKILLs firecracker processes whose 32-hex --id matches no live containerd-shim-kata-v2 … -id (a live jailed firecracker always has PPID 1, so parentage cannot be used) and qemu processes reparented to init; (b) tests/integration/test_zz_leaks.py runs last in the suite and asserts every node's VMM process count equals its live Kata pod count via hostPID scan pods.

Reference: kata-containers src/runtime/virtcontainers/fc.go (fcInit/fcEnd), firecracker jailer docs (--daemonize).

Time lost: ~1.5 h (live repro + kata source dive), on top of the ~1.5 h in #6.


9. Cilium matchPattern * never crosses label boundaries — *.docker.com silently misses CDN blob hosts (spec 18f)

Symptom: docker pull inside a sandbox with --egress '*.docker.io' --egress '*.docker.com' --egress docker.io --egress '*.cloudfront.net' fetches the manifest fine but times out downloading blobs (dial tcp 108.156.22.x:443: i/o timeout), even though cilium fqdn cache list shows production.cloudfront.docker.com being learned. Hubble showed the SYNs Policy denied DROPPED with the CloudFront IPs still carrying identity world; cilium ip list had fqdn:*.docker.io entries (single-label subdomain registry-1) but nothing for the blob host.

Root cause: in Cilium's FQDN matchPattern grammar (pkg/fqdn/matchpattern), * expands to [-a-zA-Z0-9_]* — DNS characters within a single label. production.cloudfront.docker.com therefore does not match *.docker.com (and it is not under cloudfront.net at all, so that entry never helped). The multi-label subdomain wildcard is the non-obvious **. prefix form. Not a Cilium bug — a semantics trap between k7's documented "wildcards like *.huggingface.co" UX and Cilium's grammar.

Fix: K7Core._apply_cilium_egress_policy now translates a leading *. into **. (explicit **. and mid-label wildcards pass through). Verified live: the same pull that timed out for 2m40s completes in ~8s. Also set Cilium dnsProxy.minTtl=3600 at install: CDN DNS TTLs are 3060s while dockerd's blob downloader keeps dialing its cached IP for minutes, so with minTtl=0 the learned FQDN→identity mapping can expire mid-download. Integration coverage: test_docker_pull_through_fqdn_whitelist.

Reference: cilium pkg/fqdn/matchpattern/matchpattern.go (escapeRegexpCharacters), Cilium docs "DNS based" policies.

Time lost: ~2 h (repro, hubble/ipcache/fqdn-cache spelunking, a wrong first hypothesis on TTL expiry that the live test disproved).

10. kql-r3 dind IO wedge: single-threaded virtiofsd starves the kata-agent health ping (spec 18g)

Symptom: A kql (kata-qemu-longhorn) sandbox with the docker sidecar, running the spec-10b run_io workload (2k files + 512 MB dd conv=fsync) on an r=3 Longhorn volume, would intermittently (~50% per rep) "wedge": exec 500s, both containers restarted, Pod sandbox changed, it will be killed and re-created. First seen 2/2 in the spec-18e bench (run_read unmeasurable on r3).

Root cause: Not the guest, not dockerd, not Longhorn. A dmesg -c + /proc/meminfo stream running inside the guest right through the death showed a healthy VM (load 0.4, 1.5 GB free, zero dirty/writeback, no OOM, no hung tasks) — the last log lines were normal container veth setup. The containerd log on the node had the smoking gun: floods of ttrpc: received message on inactive stream, then failed to ping agent: CheckRequest timed outDead agentsandbox stopped unexpectedly — the kata shim killed a healthy VM. Kata's default virtiofsd runs --thread-pool-size=1, so ALL virtio-fs IO (container rootfs + the Longhorn-PVC /var/lib/docker) serializes through one thread. docker run on a vfs-driver dind copies the whole ~790 MB image rootfs and then fsyncs 512 MB through that single thread against an r=3 volume (6080 s saturated). Agent RPCs that touch virtio-fs queue behind the convoy; the shim's health ping starves and it declares the agent dead. r≥2 matters only because Longhorn write amplification makes the convoy long enough to exceed the ping deadline.

Misdiagnoses ruled out on the way: guest memory sizing (MemAvailable 1.8 GB throughout), vCPU count (4-vCPU guest wedged faster), Longhorn backpressure/faults (volume attached healthy, no rebuilds), kubelet exec-probe pressure (relaxing docker info/true probe timeouts from the kubelet default 1 s reduced cancelled-ttrpc noise but did NOT stop the kill).

Fix: playbook now sets virtio_fs_extra_args = ["--thread-pool-size=16", "--announce-submounts"] in configuration-qemu.toml (kata reads it per sandbox start, no restart needed). Verified on the live 3-node HA cluster: 6/6 run_io reps + run_read (~45 s) with zero VM restarts on the same r=3 volume. The probe relaxations in core.py were kept as well (less cancelled-exec churn on the shim↔agent ttrpc channel).

Reference: kata-containers virtiofsd integration (default --thread-pool-size=1), virtiofsd docs on request queueing.

Time lost: ~3 h (bench-faithful repro, guest-side dmesg/meminfo streaming, two disproven hypotheses, virtiofsd A/B).

11. Control-plane SSRF via sandbox image + missing API-key namespace authz (spec 10h)

Symptom: Authenticated callers of k7-api 0.2.0 could point the control plane at internal/loopback/metadata addresses by supplying a crafted container image (e.g. 169.254.169.254/... or 127.0.0.1:PORT/...). Separately, any valid API key could operate on any Kubernetes namespace — keys were authenticated but not authorized.

Root cause: _get_registry_image_config parsed the registry host straight from the user-controlled image reference and issued httpx GETs with no allowlist and no private/loopback/link-local rejection; the localhost case even downgraded to plaintext http. On the authz side, verify_api_key returned key metadata that no handler consulted, and namespace was a free query/body parameter on every route.

Fix:

  • _assert_registry_host_allowed — allowlist (default public registries + K7_REGISTRY_ALLOWLIST) plus resolve-and-deny for non-public addresses; called before any registry HTTP; follow_redirects=False; localhost→http downgrade removed. Also enforced early in create_sandbox.
  • Optional "namespaces": [...] on API key records; CLI generate-api-key -n; authorize_namespace applied on every namespace-bearing endpoint. Absent/empty scope remains unrestricted.

Reference: Responsible disclosure against k7-api 0.2.0 (SSRF ≈ CVSS 7.1; missing namespace authz ≈ CVSS 9.1 in multi-tenant). spec 10h-security-ssrf-and-namespace-authz.

Time lost: n/a (implemented from disclosure + spec).


12. k7 exec swallows --rm / -c (Show HN apt 0.2.1 smoke)

Symptom: k7 --core exec NAME docker run --rm hello-world returned in ~0.8 s on every backend with no Hello from Docker. Separately, k7 exec NAME sh -c 'echo x > /tmp/m' aborted with PyInstaller's tried to call itself with '-c'.

Root cause: Typer treats --rm as an option of k7 exec, so it never reaches docker. The packaged CLI is a PyInstaller binary; a guest command that includes -c trips its self-execution guard.

Fix: pass a separator and avoid -c in the guest command: k7 --core exec NAME -- docker run --rm hello-world and k7 --core exec NAME -- 'echo x > /tmp/m'. No product change this round.

Reference: none (Typer + PyInstaller).

Time lost: ~20 min (misread as sidecar/egress failure).


13. k7d --docker: BuildKit TLS vs dockerd pull; -p is guest-host netns (spec 37a-inc2)

Symptom: docker run --rm hello-world and docker run alpine:3.21 succeeded in a --docker k7d sandbox, but docker build -t myapp . from alpine:3.21 failed with x509: certificate signed by unknown authority on auth.docker.io, and wget http://127.0.0.1:8080 from k7 exec never saw an nginx published with -p 8080:80. K7Core.exec_command also always reports exit_code=0, so a failed docker build looked like a successful tag that docker run myapp then could not find.

Root cause: dockerd's pull path uses the payload CA bundle (guest/docker/payload/etc/ssl/certs/ca-certificates.crt). Default BuildKit metadata fetch (docker driver, buildx 0.20 / buildkit v0.18) does not, so docker build / compose build: fail TLS while docker pull works. docker run -p 8080:80 publishes in the guest host netns (k7d inc1 probes it via agent vsock wget). k7 exec is the CRI container netns, so localhost:8080 is the wrong place. exec_command never reads the kubectl-exec exit code.

Fix: tests use DOCKER_BUILDKIT=0 docker build (classic builder → dockerd pull + CA) and probe the inner nginx with docker exec web wget http://127.0.0.1:80. Guest exit codes are asserted via an __K7_EC:$? marker. Do not "fix" this by setting DOCKER_BUILDKIT=0 as a product default.

Reference: k7d crates/k7d/tests/test_docker_service.rs test_docker_warm_fork_nginx_survives (agent_sh wget 8080); k7d guest/docker/payload/etc/ssl/certs/ca-certificates.crt.

Time lost: ~40 min (first integration pass).


14. k3s containerd does not stamp the runtime-handler; k7d_version 0.5.0 would clobber the new shim (spec 38a-inc5)

Symptom: RuntimeClass k7-fc is not enough for the shim to see handler k7-fc — this node's k3s 1.36 / containerd 2.3.3 never sets io.kubernetes.cri.runtime-handler on the sandbox OCI spec (k7d CHALLENGES #228). Also, k7 install --backend k7d-fc with the playbook default k7d_version: 0.5.0 would download the public GitHub tarball and overwrite /usr/local/bin/containerd-shim-k7-v1 / k7d with a build that does not know ConfigPath.

Root cause: ConfigPath on runtimes.k7-fc is the working signal. Kata Firecracker lives at /opt/kata/bin (playbook pin ~v1.14); k7d-fc installs upstream v1.16.1 at /usr/local/bin — different paths, do not share binaries. grep runtimes.k7 matches k7-fc.

Fix: playbook writes /etc/k7d/shim-k7-fc.toml and a ConfigPath [options] table without BinaryName. Install Firecracker via vendored src/k7/deploy/k7d-fc/install-firecracker.sh (sha fatal). Smoke-test the playbook with --tags k7d-fc (facts tagged always) and --k7d-artifact pointing at a tarball built from the k7d tree that just passed make remote-check, never the 0.5.0 GitHub URL. k7_has_k7d is true for k7d or k7d-fc. The allowed-backend difference() list must include k7d-fc or a k7d-fc-only install fails validation before any task runs. A fifth trap: the playbook replaces the whole containerd template from the selected backends. A full k7 install --backend k7d,k7d-fc on a mixed k7d-dev node would drop kata/nvidia/wasm runtime blocks. --tags k7d-fc installs Firecracker + toml + RuntimeClass + labels and does not rewrite the template (runtimes.k7-fc is also self-patched by k7d's ensure_k7_fc_runtime_registered).

Reference: k7d CHALLENGES #228, docs/backends.md.

Time lost: named in the spec before implementation (~15 min confirming live sandbox inspect on the node).


15. CRI exec into a k7-fc guest hung; kube Ready never flipped (spec 38a-inc5 / fixed 38a-inc6)

Symptom: tests/integration/test_k7d_fc.py created a Running runtimeClassName: k7-fc pod (shim log backend=Firecracker) then hung forever on K7Core.exec_command. timeout 15 k3s kubectl exec … -- echo hello exited 124. The sandbox's exec readiness probe (/bin/sh -c true, 5s) also timed out, so the container never became Ready.

Root cause: Firecracker guests are reached over the jail vsock UDS (FcUds), not host AF_VSOCK. The daemon publishes guest_cid = 0 (NO_HOST_CID). A host that dialled CID 0 blocked in connect() with no timeout, so CRI Wait never saw an exit code. This was not ConfigPath / RuntimeClass selection (k7d #228 / #230) and not virtiofs (k7d-fc has none).

Fix (inc6): k7d refuses AF_VSOCK CIDs 02 and proves timeout 15 kubectl exec … -- echo hello plus kube Ready on k7-fc (test_k7_fc_exec_probe_ready_and_native_exec). Install that artifact with --k7d-artifact (not GitHub latest). k7-fc tests wait on Ready again; test_k7d_fc_exec_echo_hello and TestDockerK7dFc exec into the guest. Ready is a valid signal.

Reference: k7d CHALLENGES #228 / #230 (selection) and #232 (CID 0).

Time lost: ~45 min in inc5 (pytest sat on exec until the SSH session dropped; leftover ns k7-test-fe13c2d7 had to be deleted by hand).


16. Kata --docker vehicle: persist-bind hides CLI; two-PVC fork; shareProcessNamespace (spec 22a)

Symptom / traps while bringing --docker to kfd/kql:

  1. kql persist-bind overlays /usr. The spec mounts the docker CLI via emptyDir subPath onto /usr/local/bin/docker and /usr/local/lib/docker/cli-plugins/*. kql's persist-bind then mount --binds the PVC's /usr over that path, so compose/buildx vanish. Fix: also mount the CLI emptyDir at /run/k7/docker-cli (persist-bind does not overlay /run) and restage the canonical paths immediately after the /usr bind. Vehicle path-sharing fingerprints that staging file, not /usr.
  2. hostPID is the node, not the Kata guest. Path sharing needs the vehicle to see the sandbox container's rootfs. shareProcessNamespace: true shares the guest PID namespace between the two CRI containers. hostPID: true would be the Kubernetes node (and is the wrong trust domain).
  3. Two Longhorn VolumeSnapshots are not atomic. kql fork/restore snapshots the root PVC and the docker-graph Block PVC separately. Each is crash-consistent (sync in sandbox + vehicle first); containerd's boltdb recovers. Do not claim cross-volume atomicity.
  4. Never fall back to virtio-fs for the graph. CHALLENGES #10's vfs tax and virtiofsd wedge are why the graph is volumeDevices + ext4 + overlay2 or fail loud. Keep the relaxed exec-probe timings (3/15/12/4).
  5. Alpine dind VOLUME /var/lib/docker is virtio-fs on Kata. Unmount it, mkfs only when blkid is not already ext4 (busybox blkid ignores -o value), then mount -t ext4.
  6. mount --bind from /proc/PID/root is EINVAL on Kata. Reading through that path works. Path sharing symlinks /tmp /home /root /opt /workspace in the vehicle onto the sandbox rootfs so docker run -v /tmp/x:/x resolves.
  7. OpenEBS LVM GetCapacity reports VG VFree, not thin-pool free. kata-vg is almost entirely the thin-pool plus a few GiB leftover. With storageCapacity: true (chart default) the scheduler sees ~5Gi and never binds a 20Gi k7-docker-lvm PVC. Disable capacity tracking; thin LVs come from the pool (thinProvision: yes).
  8. Kata FC virtio-fs file subPath mounts are invisible to Docker plugin discovery. docker compose is not a docker command on kfd unless compose/buildx are copied into ~/.docker/cli-plugins from the directory-mounted emptyDir. kql persist-bind already restages them as regular files.

Reference: spec 22a-kata-docker-vehicle; CHALLENGES #10, #13.

Time lost: kql integration (vehicle CrashLoopBackOff on graph mount and path-share).


17. k7-fc --docker Ready wait saw OutOfcpu corpses; CRI exec after pause hung

Symptom (bench wait): test_bench_k7d_fc at 4 CPU Guaranteed left a Ready pod plus a pile of Failed/OutOfcpu siblings (requested: 4050, used: 9890, capacity: 12000 — the just-deleted k7d sandbox's 4 CPU still counted). _wait_all_containers_ready only looked at items[0], which was a Failed pod, so the wait would have timed out at 300s while one replica was already Ready.

Fix: wait until any Running pod has all containers Ready (same shape as bench_backend_lifecycle._pod_ready). Forked-child docker info asserts overlay2 on both k7d and k7d-fc.

Symptom (lifecycle resume): alpine k7d-fc pause returned in 0.18s and k7d logged resumed vm-… (guest_cid=0). _wait_exec after resume never returned: kubelet Readiness probe failed: "/bin/sh -c true" timed out after 5s. Create→Ready and exec before pause had worked (2.14s / 0.18s). This is #15 again on the post-pause path, not a docker-perf failure. Lifecycle was aborted; do not quote k7d-fc resume→exec from that run.

Fixed (lifecycle resume), spec 42a: not #15. Two bugs, one per layer. (1) Firecracker v1.16.0/v1.16.1's vsock device armed its TRANSPORT_RESET RX gate on every PATCH /vm Resumed with no reset event for the guest to ack, so after a bare pause → resume no host→guest packet — no CONNECT reply — was ever delivered; upstream fixed it as #6100 in v1.16.2, and k7d's guest/fc/pins.env plus the vendored src/k7/deploy/k7d-fc/pins.env now pin v1.16.2. k7d's resume_vm also dials the resumed FC guest and fails loud if it does not answer. (2) The shim's exec bridge never opened containerd's stdout/stderr FIFOs on its failure path, so a probe exec that could not reach the paused guest left kubelet's prober parked in ExecSync for its 2-minute gRPC deadline — the pod stayed Ready through a whole pause on k7d and k7d-fc. tests/integration/test_pause_resume.py::TestPauseResumeExecAnswers covers pause → resume → exec_command("echo hi") on both; the PERFORMANCE.md k7d vs k7d-fc lifecycle table has the resume→exec number. Root cause and timings: k7d CHALLENGES #240 (vsock gate) and #241 (FIFOs).

Reference: CHALLENGES #15; k7d #232, #240, #241; Firecracker #6100.

Time lost: ~15 min on the OutOfcpu wait; lifecycle resume hung until the pytest process was killed (~8 min).


18. k7 install --backend k7d-fc cannot find vendored Firecracker installer

Symptom: A 3-node HA install with k7_backends=kfd,kql,k7d,k7d-fc (public k7 0.3.0) failed at K7d-fc — stage install-firecracker.sh and pins on every node:

Could not find or access 'k7d-fc/install-firecracker.sh'
Searched in:
  /tmp/files/k7d-fc/install-firecracker.sh
  /tmp/k7d-fc/install-firecracker.sh
  ... on the Ansible Controller.

K3s HA, Cilium, Longhorn, kfd thin-pool, and k7d itself had already succeeded.

Root cause: k7 install writes the embedded playbook to a tempfile (/tmp/tmp….yaml) and runs ansible-playbook against that. Ansible copy without remote_src looks up src: k7d-fc/install-firecracker.sh next to the playbook file, i.e. /tmp/k7d-fc/…. The vendored files live at src/k7/deploy/k7d-fc/ in the source tree. The API-manifest copy already documents this trap and uses k7_repo_root; k7d-fc did not.

Fix: copy from {{ k7_repo_root }}/src/k7/deploy/k7d-fc/{install-firecracker.sh,pins.env} on the controller (the checkout the CLI already requires for k7-api:local).

Reference: playbook comment on "Copy K7 API manifests on first master".

Time lost: one full 3-node TWO_DISK reset + 6 min install (~45 min).


19. k7 create --docker via the API always said "this k7d has no docker service"

Symptom: After a successful 3-node HA install of k7d 0.6.0 (payload at /usr/local/share/k7d/docker/bin/dockerd, /etc/k7/k7d_version = 0.6.0), the docs path k7 create --docker --backend k7d --egress-open builder ubuntu:24.04 failed immediately with this k7d has no docker service; upgrade. The same create via k7 --core (and the integration suite, which uses --core) succeeded.

Root cause: k7d_supports_docker() treated a recorded version ≥ 0.6.0 as "not old" and then still required os.path.isfile of the host dockerd payload. k7-api hostPath-mounts /etc/k7 (so it can read the version file) but not /usr/local/share/k7d, so the payload check always failed inside the API pod. Default CLI routing is the API, so every laptop/docs user hit this; --core on the node never did.

Fix: if the playbook recorded a parseable version, that version is authoritative (>= 0.6.0 → supported). The payload path is only consulted when the version file is missing (CLI --core / incomplete install).

Reference: src/k7/deploy/manifests/k7-api/deployment.yaml (/etc/k7 hostPath); K7D_DOCKER_PAYLOAD_DOCKERD in src/k7/core/docker.py.

Time lost: ~30 min diagnosing why the live cluster had dockerd but the API refused --docker.


20. Laptop k7 api status crashed; docs k7.yaml 128Mi never went Ready

Symptom: After apt install k7 (PPA 0.3.1) on a laptop with k7 config set api.url / api.ca / api.key, k7 list and k7 nodes storage worked, but the Quickstart's next commands k7 api status and k7 api endpoint raised FileNotFoundError: kubectl. The same Quickstart's examples/k7.yaml (cpu: 100m, memory: 128Mi, before_script: apk add curl, default backend kfd) timed out with "Timed out waiting for sandbox container to start".

Root cause: k7 api status / endpoint always shelled out to kubectl (or k3s kubectl) without checking the binary exists. A laptop that only has the .deb has no kubeconfig and no kubectl. Separately, the docs k7.yaml set memory: 128Mi, which Kata stamps as io.katacontainers.config.hypervisor.default_memory: 128. The Firecracker shim refuses anything below 256Mi, so the pod stays ContainerCreating (FailedCreatePodSandBox) until create times out. before_script never runs.

Fix: missing kubectl falls back to GET /health on the configured API URL (status) / prints that URL (endpoint). Create rejects Kata memory below 256Mi immediately. Docs + examples/k7.yaml use backend: k7d and cpu: "1" / memory: "1Gi". PPA is 0.3.1, not 0.2.2.

Reference: src/k7/cli/k7.py (_kubectl_run, _api_status_via_https); examples/k7.yaml.

Time lost: ~20 min reproducing on a 3-node HA soak after a public 0.3.1 release; the install itself was one command and succeeded.


21. --expose-port NodePort is dead until Ready; k7 exec -- sh -c self-nests

Symptom: Docs k7 create --expose-port 8000 --before-script 'nohup python3 -m http.server 8000 &' printed a NodePort, but curling it from a laptop timed out. Cilium showed the Service in maintenance. k7 exec NAME -- sh -c 'echo hi > /tmp/x' failed with Nuitka/PyInstaller-style "tried to call itself with '-c'". k7 restore --latest and k7 delete-all -y do not exist.

Root cause: externalTrafficPolicy: Local plus Cilium keeps a NodePort in maintenance while endpoints are notReadyAddresses. k7d/kfd pid 1 is sleep 365d; a bare & in before_script is killed when the script exits, so the Ready probe never sees the http server (and a missing touch of the done file has the same effect). k7 exec already wraps the joined argv in sh -c, so a nested sh -c is the CLI binary eating -c. Restore takes two positionals; delete-all confirms interactively with no -y.

Fix: Docs: trailing sleep 1 after nohup so before_script can finish and Ready can fire; curl the pod's node, not an arbitrary master. Exec: one quoted string, no extra sh -c. Restore/delete-all examples match the CLI. Create success text points at k7 list, not k7 list --name.

Reference: src/k7/cli/k7.py exec/create; docs k7/guides/cli.mdx.

Time lost: ~40 min on the public 0.3.1 HA soak (expose looked like a CNI bug until Ready flipped).


22. k7d docker graph image is a teardown race, not a leak (spec 39a)

Symptom: TestDockerK7d / TestDockerK7dFc failed after k7 delete + 3s with leaked k7d docker volume images: {scratch-vm-…-docker.img}. Hours later, same k7d PID, the files were gone.

Root cause: k7 delete returns when Kubernetes objects are gone. The VM's ScratchDisk::drop unlinks the graph image asynchronously when the containerd shim Delete lands. On a busy 3-node HA box that is more than 3s.

Fix: bounded poll (~60s) in the test. delete_sandbox does not wait for the k7d VM — coupling API latency to shim teardown would stall every delete.

Reference: tests/integration/test_docker.py _wait_k7d_docker_disks_gone; K7Core.delete_sandbox.

Time lost: the soak already established this; the 3s sleep was the only defect.


23. Firecracker jailer test missed firecracker-v1. (spec 39a)

Symptom: test_jailer_active asserted "No firecracker processes found on the host" while a jailed kfd VM was running.

Root cause: Linux truncates comm to 15 characters. The pinned binary is firecracker-v1.16.1, so comm is firecracker-v1.. The helper compared == "firecracker". The chroot binary is also versioned (firecracker-v1.16.1, not /firecracker). The jail itself is correct: /etc/passwd, /etc/shadow, /usr, /boot absent; vmlinux + rootfs present. readlink /proc/<pid>/root is / because the jailer pivot-roots in a private mount ns.

Fix: match comm by prefix firecracker-, require fcConfig.json / --config-file on the cmdline, skip (deleted) orphans, require vmlinux+rootfs, glob firecracker* in the chroot. Host-FS-unreachable asserts stay.

Reference: tests/integration/test_firecracker.py _get_live_firecracker_pids.

Time lost: ~20 min of /proc on the soak node.


24. Expose tests timed out because the pod was on another node (spec 39a)

Symptom: TestSandboxExpose curled http://<other-node>:<nodeport> from the first master and timed out. Off-cluster, the pod's node answered HTTP 200 and the other nodes refused connect.

Root cause: externalTrafficPolicy: Local is required (without it, cidr: rules see SNAT). Cilium socket-LB intercepts in-cluster-node origin to a NodePort on a different node. The tests did not pin node_name.

Fix: pin expose sandboxes to os.uname().nodename. Do not switch the Service to Cluster.

Reference: tests/integration/test_sandbox_ingress.py TestSandboxExpose._exposed.

Time lost: ~15 min confirming Local vs Cilium vs a wrong-node curl.


25. kql live --docker fork never went Ready: overlay2 was crash-inconsistent (spec 39a)

Symptom: TestDockerKQL.test_fork_clones_both_pvcs — both VolumeSnapshots ready, both child PVCs Bound, child never Ready in 240s.

Root cause (live, spec 39a): the child stuck in Init:0/2 with FailedAttachVolume: volume is not ready for workloads. Kubernetes Bound is not Longhorn-ready-to-attach; HA r=3 clone hydration takes longer than 240s. The qemu fork test already waits 600s for this. sync is also not enough for a busy overlay2 once the guest does start.

Fix: wait 600s for the child (same bound as test_qemu.test_fork_clones_data). Plus fsfreeze the graph: alpine docker:27.5.1-dind has no fsfreeze, so stage it from the ubuntu sandbox via the shared /tmp emptyDir, copy into the vehicle rootfs, freeze around VolumeSnapshot create, thaw in finally.

Reference: K7Core._create_kata_snapshots_quiesced.

Time lost: soak diagnosis; freeze is the product answer rather than refusing live forks.


26. Partner-facing 0.3.1 traps: kfd fork 404, SDK snippet missing CA, k7 logs empty

Symptom: Following docs.katakate.org / the k7 README on a 3-node HA PPA 0.3.1 cluster:

  • k7 fork of a kfd sandbox printed Source root PVC <name>-root-lh not found; cannot fork storage instead of "kfd cannot fork".
  • k7 api status printed Client(endpoint=..., api_key=...) with no verify_ssl; that fails against the playbook-minted cluster CA.
  • k7 logs demo --tail 20 printed nothing (exit 0) after only k7 exec.
  • python3 -m venv on the node failed (ensurepip is not available).
  • k7 resume of a kql sandbox returned immediately while the pod was still Pending.

Root cause: fork_sandbox treated missing kfd PVCs as a generic storage 404. The status command's help snippet never grew verify_ssl when HTTPS-by-default landed. k7 logs is a CRI container snapshot; exec goes through the agent. Ubuntu 24.04 cloud images omit python3-venv. HA Longhorn attach is slower than resume() returning.

Fix: reject every kfd fork up front (KFD_FORK_REJECT). Print verify_ssl='./k7-ca.crt' in k7 api status. Docs: wait-until-Ready after kql resume, empty logs are success, SDK is a client install (apt install python3-venv on the node).

Reference: none.

Time lost: ~30 min walking the quickstart on the HA cluster.


27. k7 --core leaked kubernetes_asyncio/aiohttp sessions (Event loop is closed)

Symptom: TestSnapshotCrud.test_create_list_inspect_delete_round_trip failed with assert snap in cp.stdout and cp.stdout == '\n'. Interpreter also logged Unclosed client session / Event loop is closed after k7 --core snapshot create.

Root cause: Each CoreV1Api() / AppsV1Api() / CustomObjectsApi() constructed its own ApiClient (aiohttp session) bound to the asyncio.run loop. Typer handlers never called close(), so loop teardown raced the session destructor. Separately, _create_kata_snapshots_quiesced dropped the successful OperationResult.message from _create_volume_snapshot, so the CLI echoed a blank line even when the snapshot existed.

Fix: One shared ApiClient per K7Core, async def aclose(), CLI _core_run / API get_k7_core / snapshot-gc always close. Quiesced snapshot success now keeps the create message (Snapshot <name> created for PVC …).

Reference: kubernetes_asyncio ApiClient.closerest_client.close() (aiohttp).

Time lost: caught on the HA integration run after the partner walkthrough.