Per-key node pins, k7d-fc pause/resume/exec, and HA-soak fixes. Playbook pins k7d 0.7.0. GitHub .deb, Launchpad PPA, and PyPI k7-sdk are 0.4.0.
38 KiB
Challenges & solutions log
Tracking non-obvious bugs in k7 (the sandbox management layer) and how they
were solved. Same format as k7d-dev's CHALLENGES.md.
1. k7 install --backend extra-var silently clobbered inventory k7_backends (spec 18e)
Symptom: A 3-node HA install with an inventory declaring
k7_backends=kfd,kql,k7d per host would have provisioned only the two Kata
backends — the k7d backend would silently disappear from every node.
Root cause: k7 install always forwarded k7_backends as an Ansible
extra-var (built from the --backend option's default value even when
the user never passed --backend). Extra-vars have the highest precedence in
Ansible, so the CLI default overrode the per-host inventory var.
Fix: install() now checks ctx.get_parameter_source("backend"); when
the user supplied -i <inventory> without an explicit --backend, the
k7_backends extra-var is dropped so the inventory wins. An explicit
--backend alongside -i prints a warning that it overrides the inventory.
Reference: none (Ansible variable-precedence rules).
Time lost: caught in pre-flight review (spec 18e Phase 0), ~1 hour of code reading. Would have cost a full reset+reinstall cycle if it had shipped.
2. Longhorn StorageClass numberOfReplicas parameter makes default-replica-count a no-op (spec 18e)
Symptom: bench_docker_perf.py's r1/r2 legs patched Longhorn's
default-replica-count setting — but sandbox volumes kept the replica count
baked into the longhorn StorageClass (numberOfReplicas: "3" on the HA
cluster). The r1/r2/r3 benches would all have silently measured the same
replica count.
Root cause: Longhorn only consults the default-replica-count setting
when the StorageClass has no numberOfReplicas parameter. k7's install
playbook pins the parameter in the SC (topology-aware SC from the
longhorn-storageclass ConfigMap), so the setting never applies to k7
volumes.
Fix: the bench now patches spec.numberOfReplicas on the sandbox's
Longhorn Volume CRs after creation (_set_sandbox_volume_replicas),
waits until exactly N running replicas exist, and records the replica → node
placement in the log header as proof.
Reference: Longhorn docs (volume-level replica count update).
Time lost: ~1 hour. The old benches "passed" — the mismeasurement was invisible without checking actual replica CRs.
3. k7d VM operations are node-local; multi-node scheduler breaks pause/fork tests (spec 18e)
Symptom: On the 3-node cluster, test_k7d.py pause/fork tests failed
with sandbox X runs on node k7-node-02, but this k7 process runs on k7-node-01; k7d VM operations must run on the pod's node.
Root cause: k7d pause/resume/fork go through the node-local
/run/k7d/k7d.sock; cross-node VM ops are explicitly out of scope (k7d spec
9a M12) and core fails loudly on the mismatch. On a single-node cluster the
tests never noticed; with 3 schedulable nodes the sandbox lands anywhere.
Fix: tests that exercise VM ops pin their sandboxes with
SandboxConfig(node_name=os.uname().nodename) — the field that exists
precisely for host-side-inspection tests. The centralized-API implication
(k7-api pod can only pause/fork k7d sandboxes co-located on the first
master) is recorded as a release-readiness limitation.
Reference: k7d spec 9a M12 (cross-node fork out of scope).
Time lost: ~30 min (the loud error message made it easy).
4. kql fork loses guest writes made just before the fork (crash-consistency)
Symptom: test_api.py::test_sdk_pause_resume_fork_round_trip flaked on
the HA cluster: a file written via exec seconds before fork() did not
exist in the fork (cat: can't open '/mnt/state/marker'). The near-identical
test_qemu.py::test_fork_clones_data data check passed in the same run —
pure timing luck.
Root cause: the kql fork path cuts a block-level Longhorn VolumeSnapshot of the source's root PVC. That is only crash-consistent: guest writes still sitting in the VM's page cache are not on the block device yet and are missing from the clone.
Fix: fork_sandbox now execs sync in the source sandbox before
creating the snapshot (only when the deployment has ready replicas — a
paused source has no writers), failing loudly if the flush fails.
Reference: none (standard crash-vs-application consistency).
Time lost: ~1 hour including re-runs. Note: named snapshots of live
sandboxes (k7 snapshot) remain crash-consistent by design — documented
behavior, unchanged.
5. NVMe enumeration swaps across reboots — hardcoded k7_devmapper_disk hit the OS disk
Symptom: The second full reset+reinstall loop of spec 18e failed on
k7-node-03: Device '/dev/nvme1n1' has partitions; wipe it with utils/wipe-disk.sh or choose another disk. The identical inventory had just
worked on the first loop.
Root cause: Linux NVMe controller enumeration (nvme0n1 vs nvme1n1)
is not stable across reboots. After the second reset, node-03 booted with
its OS on the disk now enumerated nvme1n1, and the raw spare as
nvme0n1 — the inventory's hardcoded k7_devmapper_disk=/dev/nvme1n1
pointed at the OS disk. The playbook's safety checks caught it (fail-loud
worked as designed).
Fix: omit k7_devmapper_disk on identical dual-NVMe boxes — the
playbook's auto-detect ("first empty, non-removable, non-root whole disk")
is enumeration-proof. inventory.ini.example now documents this.
Reference: none (kernel device-naming behavior).
Time lost: ~30 min (one wasted install attempt + one extra reset+reinstall loop of all three nodes).
6. Orphaned Firecracker microVMs leak on pod deletion and each burns a full CPU core
Symptom: During spec 18e Phase 3, the k7-ql-r2 bench leg started
failing mid-run with Pod is not running (status: Pending) and the
k7-ql-r3 leg failed entirely; longhorn-manager / cilium-envoy /
coredns readiness probes were flapping cluster-wide. k7-node-03 had a
load average of ~21.
Root cause: 14 orphaned /firecracker processes (2 on node-01, 2 on
node-02, 10 on node-03) whose pods had been deleted hours earlier —
zero live Kata pods existed cluster-wide. Each orphan spun at ~97% CPU
(TIME ≈ ETIME in ps), starving Longhorn/Cilium/CoreDNS and the bench
sandbox itself. The kata-fc shim intermittently fails to kill the
microVM on pod deletion under parallel pod churn (~14 leaks over ~30
kfd pod deletions that day). Evidence:
/tmp/leaked-firecracker-vms.txt (agent run artifact).
Fix (remediation): verified no live Kata pods, then pkill -9 firecracker on all three nodes; loads recovered and the r2/r3 bench
legs were re-run green. Root-cause fix still open — tracked as a
release blocker in spec 18f-release-blockers (investigate
containerd-shim-kata-v2 / jailer cleanup path; add a leak-detection
integration test that asserts zero firecracker processes after suite
teardown).
Reference: none yet (kata-containers shim lifecycle).
Time lost: ~1.5 hours (failed bench legs + diagnosis + re-run).
7. Remote test loop tied to the SSH session died mid-run (Broken pipe)
Symptom: A multi-suite pytest loop launched over plain ssh host 'for f in ...; do pytest ...; done' died silently when the SSH connection
dropped (client_loop: send disconnect: Broken pipe) — the remote shell got
SIGHUP'd between suites.
Root cause: the remote loop was a child of the SSH session; NAT idle timeouts kill long-lived connections even with keepalives.
Fix: write the loop to a script on the node and launch it with
setsid nohup ... < /dev/null &, then poll a progress file. (Same class of
issue utils/run-integration-tests.sh already documents for its keepalive
settings.)
Reference: none.
Time lost: ~20 min (one interrupted suite sequence, test_restore had
finished right before the drop).
8. Firecracker leak root cause: jailer --daemonize makes the kata shim signal a dead PID (spec 18f)
Symptom: Follow-up to #6. Reproduced at will on the 18f run: create a
naked kfd pod (sleep workload), delete it — the pod terminates cleanly
but its /firecracker process survives with PPID 1 and climbs to ~100%
CPU. Two out of two attempts leaked. Shim logs at 18e leak time showed
Agent did not stop sandbox: Dead agent + failed to ping agent: CheckRequest timed out.
Root cause: kata 3.24.0 virtcontainers/fc.go. When jailed (spec 8a
enabled the jailer), fcInit launches jailer --daemonize, which
double-forks — firecracker reparents to init immediately, and
fc.info.PID = cmd.Process.Pid records the jailer's PID, which is
already dead. fcEnd() then calls WaitLocalProcess(pid, …, SIGTERM) on
that stale PID: a no-op. The VMM normally exits because the in-guest agent
shuts the VM down; whenever that graceful path fails (dead/hung agent under
churn, wedged guest IO), nothing ever kills the firecracker process. The
getting vm status failed … firecracker.socket: no such file or directory
error seen at every kfd VM boot is a side effect of the same daemonize
handling (the shim polls the jailed API socket path before it exists) and
is harmless noise.
Fix: upstream fix belongs in kata (record the real VMM PID when
jailed). In k7: (a) k7 install now deploys a per-node systemd timer
k7-vmm-reaper.timer (1 min cadence) that SIGKILLs firecracker processes
whose 32-hex --id matches no live containerd-shim-kata-v2 … -id
(a live jailed firecracker always has PPID 1, so parentage cannot be used)
and qemu processes reparented to init; (b)
tests/integration/test_zz_leaks.py runs last in the suite and asserts
every node's VMM process count equals its live Kata pod count via hostPID
scan pods.
Reference: kata-containers src/runtime/virtcontainers/fc.go
(fcInit/fcEnd), firecracker jailer docs (--daemonize).
Time lost: ~1.5 h (live repro + kata source dive), on top of the ~1.5 h in #6.
9. Cilium matchPattern * never crosses label boundaries — *.docker.com silently misses CDN blob hosts (spec 18f)
Symptom: docker pull inside a sandbox with
--egress '*.docker.io' --egress '*.docker.com' --egress docker.io --egress '*.cloudfront.net' fetches the manifest fine but times out
downloading blobs (dial tcp 108.156.22.x:443: i/o timeout), even though
cilium fqdn cache list shows production.cloudfront.docker.com being
learned. Hubble showed the SYNs Policy denied DROPPED with the CloudFront
IPs still carrying identity world; cilium ip list had fqdn:*.docker.io
entries (single-label subdomain registry-1) but nothing for the blob host.
Root cause: in Cilium's FQDN matchPattern grammar
(pkg/fqdn/matchpattern), * expands to [-a-zA-Z0-9_]* — DNS characters
within a single label. production.cloudfront.docker.com therefore
does not match *.docker.com (and it is not under cloudfront.net at
all, so that entry never helped). The multi-label subdomain wildcard is the
non-obvious **. prefix form. Not a Cilium bug — a semantics trap between
k7's documented "wildcards like *.huggingface.co" UX and Cilium's grammar.
Fix: K7Core._apply_cilium_egress_policy now translates a leading *.
into **. (explicit **. and mid-label wildcards pass through). Verified
live: the same pull that timed out for 2m40s completes in ~8s. Also set
Cilium dnsProxy.minTtl=3600 at install: CDN DNS TTLs are 30–60s while
dockerd's blob downloader keeps dialing its cached IP for minutes, so with
minTtl=0 the learned FQDN→identity mapping can expire mid-download.
Integration coverage: test_docker_pull_through_fqdn_whitelist.
Reference: cilium pkg/fqdn/matchpattern/matchpattern.go
(escapeRegexpCharacters), Cilium docs "DNS based" policies.
Time lost: ~2 h (repro, hubble/ipcache/fqdn-cache spelunking, a wrong first hypothesis on TTL expiry that the live test disproved).
10. kql-r3 dind IO wedge: single-threaded virtiofsd starves the kata-agent health ping (spec 18g)
Symptom: A kql (kata-qemu-longhorn) sandbox with the docker sidecar,
running the spec-10b run_io workload (2k files + 512 MB dd conv=fsync)
on an r=3 Longhorn volume, would intermittently (~50% per rep) "wedge":
exec 500s, both containers restarted, Pod sandbox changed, it will be killed and re-created. First seen 2/2 in the spec-18e bench (run_read
unmeasurable on r3).
Root cause: Not the guest, not dockerd, not Longhorn. A dmesg -c +
/proc/meminfo stream running inside the guest right through the death
showed a healthy VM (load 0.4, 1.5 GB free, zero dirty/writeback, no OOM,
no hung tasks) — the last log lines were normal container veth setup. The
containerd log on the node had the smoking gun: floods of
ttrpc: received message on inactive stream, then
failed to ping agent: CheckRequest timed out → Dead agent →
sandbox stopped unexpectedly — the kata shim killed a healthy VM.
Kata's default virtiofsd runs --thread-pool-size=1, so ALL virtio-fs IO
(container rootfs + the Longhorn-PVC /var/lib/docker) serializes through
one thread. docker run on a vfs-driver dind copies the whole ~790 MB
image rootfs and then fsyncs 512 MB through that single thread against an
r=3 volume (60–80 s saturated). Agent RPCs that touch virtio-fs queue
behind the convoy; the shim's health ping starves and it declares the
agent dead. r≥2 matters only because Longhorn write amplification makes
the convoy long enough to exceed the ping deadline.
Misdiagnoses ruled out on the way: guest memory sizing (MemAvailable
1.8 GB throughout), vCPU count (4-vCPU guest wedged faster), Longhorn
backpressure/faults (volume attached healthy, no rebuilds), kubelet
exec-probe pressure (relaxing docker info/true probe timeouts from the
kubelet default 1 s reduced cancelled-ttrpc noise but did NOT stop the
kill).
Fix: playbook now sets
virtio_fs_extra_args = ["--thread-pool-size=16", "--announce-submounts"]
in configuration-qemu.toml (kata reads it per sandbox start, no restart
needed). Verified on the live 3-node HA cluster: 6/6 run_io reps +
run_read (~45 s) with zero VM restarts on the same r=3 volume. The probe
relaxations in core.py were kept as well (less cancelled-exec churn on
the shim↔agent ttrpc channel).
Reference: kata-containers virtiofsd integration (default
--thread-pool-size=1), virtiofsd docs on request queueing.
Time lost: ~3 h (bench-faithful repro, guest-side dmesg/meminfo streaming, two disproven hypotheses, virtiofsd A/B).
11. Control-plane SSRF via sandbox image + missing API-key namespace authz (spec 10h)
Symptom: Authenticated callers of k7-api 0.2.0 could point the
control plane at internal/loopback/metadata addresses by supplying a
crafted container image (e.g. 169.254.169.254/... or
127.0.0.1:PORT/...). Separately, any valid API key could operate on any
Kubernetes namespace — keys were authenticated but not authorized.
Root cause: _get_registry_image_config parsed the registry host
straight from the user-controlled image reference and issued httpx
GETs with no allowlist and no private/loopback/link-local rejection; the
localhost case even downgraded to plaintext http. On the authz side,
verify_api_key returned key metadata that no handler consulted, and
namespace was a free query/body parameter on every route.
Fix:
_assert_registry_host_allowed— allowlist (default public registries +K7_REGISTRY_ALLOWLIST) plus resolve-and-deny for non-public addresses; called before any registry HTTP;follow_redirects=False; localhost→http downgrade removed. Also enforced early increate_sandbox.- Optional
"namespaces": [...]on API key records; CLIgenerate-api-key -n;authorize_namespaceapplied on every namespace-bearing endpoint. Absent/empty scope remains unrestricted.
Reference: Responsible disclosure against k7-api 0.2.0
(SSRF ≈ CVSS 7.1; missing namespace authz ≈ CVSS 9.1 in multi-tenant).
spec 10h-security-ssrf-and-namespace-authz.
Time lost: n/a (implemented from disclosure + spec).
12. k7 exec swallows --rm / -c (Show HN apt 0.2.1 smoke)
Symptom: k7 --core exec NAME docker run --rm hello-world returned in
~0.8 s on every backend with no Hello from Docker. Separately,
k7 exec NAME sh -c 'echo x > /tmp/m' aborted with PyInstaller's
tried to call itself with '-c'.
Root cause: Typer treats --rm as an option of k7 exec, so it never
reaches docker. The packaged CLI is a PyInstaller binary; a guest command
that includes -c trips its self-execution guard.
Fix: pass a separator and avoid -c in the guest command:
k7 --core exec NAME -- docker run --rm hello-world and
k7 --core exec NAME -- 'echo x > /tmp/m'. No product change this round.
Reference: none (Typer + PyInstaller).
Time lost: ~20 min (misread as sidecar/egress failure).
13. k7d --docker: BuildKit TLS vs dockerd pull; -p is guest-host netns (spec 37a-inc2)
Symptom: docker run --rm hello-world and docker run alpine:3.21
succeeded in a --docker k7d sandbox, but docker build -t myapp .
from alpine:3.21 failed with x509: certificate signed by unknown authority on auth.docker.io, and wget http://127.0.0.1:8080 from
k7 exec never saw an nginx published with -p 8080:80.
K7Core.exec_command also always reports exit_code=0, so a failed
docker build looked like a successful tag that docker run myapp
then could not find.
Root cause: dockerd's pull path uses the payload CA bundle
(guest/docker/payload/etc/ssl/certs/ca-certificates.crt). Default
BuildKit metadata fetch (docker driver, buildx 0.20 / buildkit v0.18)
does not, so docker build / compose build: fail TLS while docker pull works. docker run -p 8080:80 publishes in the guest host
netns (k7d inc1 probes it via agent vsock wget). k7 exec is the CRI
container netns, so localhost:8080 is the wrong place. exec_command
never reads the kubectl-exec exit code.
Fix: tests use DOCKER_BUILDKIT=0 docker build (classic builder →
dockerd pull + CA) and probe the inner nginx with docker exec web wget http://127.0.0.1:80. Guest exit codes are asserted via an __K7_EC:$?
marker. Do not "fix" this by setting DOCKER_BUILDKIT=0 as a product
default.
Reference: k7d crates/k7d/tests/test_docker_service.rs
test_docker_warm_fork_nginx_survives (agent_sh wget 8080);
k7d guest/docker/payload/etc/ssl/certs/ca-certificates.crt.
Time lost: ~40 min (first integration pass).
14. k3s containerd does not stamp the runtime-handler; k7d_version 0.5.0 would clobber the new shim (spec 38a-inc5)
Symptom: RuntimeClass k7-fc is not enough for the shim to see
handler k7-fc — this node's k3s 1.36 / containerd 2.3.3 never sets
io.kubernetes.cri.runtime-handler on the sandbox OCI spec (k7d
CHALLENGES #228). Also, k7 install --backend k7d-fc with the playbook
default k7d_version: 0.5.0 would download the public GitHub tarball
and overwrite /usr/local/bin/containerd-shim-k7-v1 / k7d with a
build that does not know ConfigPath.
Root cause: ConfigPath on runtimes.k7-fc is the working signal.
Kata Firecracker lives at /opt/kata/bin (playbook pin ~v1.14);
k7d-fc installs upstream v1.16.1 at /usr/local/bin — different
paths, do not share binaries. grep runtimes.k7 matches k7-fc.
Fix: playbook writes /etc/k7d/shim-k7-fc.toml and a ConfigPath
[options] table without BinaryName. Install Firecracker via
vendored src/k7/deploy/k7d-fc/install-firecracker.sh (sha fatal).
Smoke-test the playbook with --tags k7d-fc (facts tagged always)
and --k7d-artifact pointing at a tarball built from the k7d tree
that just passed make remote-check, never the 0.5.0 GitHub URL.
k7_has_k7d is true for k7d or k7d-fc. The allowed-backend
difference() list must include k7d-fc or a k7d-fc-only install
fails validation before any task runs.
A fifth trap: the playbook replaces the whole containerd template
from the selected backends. A full k7 install --backend k7d,k7d-fc
on a mixed k7d-dev node would drop kata/nvidia/wasm runtime blocks.
--tags k7d-fc installs Firecracker + toml + RuntimeClass + labels
and does not rewrite the template (runtimes.k7-fc is also
self-patched by k7d's ensure_k7_fc_runtime_registered).
Reference: k7d CHALLENGES #228, docs/backends.md.
Time lost: named in the spec before implementation (~15 min confirming live sandbox inspect on the node).
15. CRI exec into a k7-fc guest hung; kube Ready never flipped (spec 38a-inc5 / fixed 38a-inc6)
Symptom: tests/integration/test_k7d_fc.py created a Running
runtimeClassName: k7-fc pod (shim log backend=Firecracker) then
hung forever on K7Core.exec_command. timeout 15 k3s kubectl exec … -- echo hello exited 124. The sandbox's exec readiness probe
(/bin/sh -c true, 5s) also timed out, so the container never
became Ready.
Root cause: Firecracker guests are reached over the jail vsock
UDS (FcUds), not host AF_VSOCK. The daemon publishes
guest_cid = 0 (NO_HOST_CID). A host that dialled CID 0 blocked
in connect() with no timeout, so CRI Wait never saw an exit code.
This was not ConfigPath / RuntimeClass selection (k7d #228 / #230)
and not virtiofs (k7d-fc has none).
Fix (inc6): k7d refuses AF_VSOCK CIDs 0–2 and proves
timeout 15 kubectl exec … -- echo hello plus kube Ready on
k7-fc (test_k7_fc_exec_probe_ready_and_native_exec). Install that
artifact with --k7d-artifact (not GitHub latest). k7-fc tests
wait on Ready again; test_k7d_fc_exec_echo_hello and
TestDockerK7dFc exec into the guest. Ready is a valid signal.
Reference: k7d CHALLENGES #228 / #230 (selection) and #232 (CID 0).
Time lost: ~45 min in inc5 (pytest sat on exec until the SSH
session dropped; leftover ns k7-test-fe13c2d7 had to be deleted
by hand).
16. Kata --docker vehicle: persist-bind hides CLI; two-PVC fork; shareProcessNamespace (spec 22a)
Symptom / traps while bringing --docker to kfd/kql:
- kql persist-bind overlays
/usr. The spec mounts the docker CLI via emptyDirsubPathonto/usr/local/bin/dockerand/usr/local/lib/docker/cli-plugins/*. kql's persist-bind thenmount --binds the PVC's/usrover that path, so compose/buildx vanish. Fix: also mount the CLI emptyDir at/run/k7/docker-cli(persist-bind does not overlay/run) and restage the canonical paths immediately after the/usrbind. Vehicle path-sharing fingerprints that staging file, not/usr. hostPIDis the node, not the Kata guest. Path sharing needs the vehicle to see the sandbox container's rootfs.shareProcessNamespace: trueshares the guest PID namespace between the two CRI containers.hostPID: truewould be the Kubernetes node (and is the wrong trust domain).- Two Longhorn VolumeSnapshots are not atomic. kql fork/restore
snapshots the root PVC and the docker-graph Block PVC separately.
Each is crash-consistent (
syncin sandbox + vehicle first); containerd's boltdb recovers. Do not claim cross-volume atomicity. - Never fall back to virtio-fs for the graph. CHALLENGES #10's vfs
tax and virtiofsd wedge are why the graph is
volumeDevices+ ext4 + overlay2 or fail loud. Keep the relaxed exec-probe timings (3/15/12/4). - Alpine dind
VOLUME /var/lib/dockeris virtio-fs on Kata. Unmount it, mkfs only when blkid is not already ext4 (busybox blkid ignores-o value), thenmount -t ext4. mount --bindfrom/proc/PID/rootis EINVAL on Kata. Reading through that path works. Path sharing symlinks/tmp/home/root/opt/workspacein the vehicle onto the sandbox rootfs sodocker run -v /tmp/x:/xresolves.- OpenEBS LVM
GetCapacityreports VG VFree, not thin-pool free.kata-vgis almost entirely the thin-pool plus a few GiB leftover. WithstorageCapacity: true(chart default) the scheduler sees ~5Gi and never binds a 20Gik7-docker-lvmPVC. Disable capacity tracking; thin LVs come from the pool (thinProvision: yes). - Kata FC virtio-fs file
subPathmounts are invisible to Docker plugin discovery.docker composeis not a docker command on kfd unless compose/buildx are copied into~/.docker/cli-pluginsfrom the directory-mounted emptyDir. kql persist-bind already restages them as regular files.
Reference: spec 22a-kata-docker-vehicle; CHALLENGES #10, #13.
Time lost: kql integration (vehicle CrashLoopBackOff on graph mount and path-share).
17. k7-fc --docker Ready wait saw OutOfcpu corpses; CRI exec after pause hung
Symptom (bench wait): test_bench_k7d_fc at 4 CPU Guaranteed left a
Ready pod plus a pile of Failed/OutOfcpu siblings (requested: 4050, used: 9890, capacity: 12000 — the just-deleted k7d sandbox's 4 CPU
still counted). _wait_all_containers_ready only looked at items[0],
which was a Failed pod, so the wait would have timed out at 300s while
one replica was already Ready.
Fix: wait until any Running pod has all containers Ready (same
shape as bench_backend_lifecycle._pod_ready). Forked-child docker info
asserts overlay2 on both k7d and k7d-fc.
Symptom (lifecycle resume): alpine k7d-fc pause returned in 0.18s
and k7d logged resumed vm-… (guest_cid=0). _wait_exec after resume
never returned: kubelet Readiness probe failed: "/bin/sh -c true" timed out after 5s. Create→Ready and exec before pause had worked
(2.14s / 0.18s). This is #15 again on the post-pause path, not a
docker-perf failure. Lifecycle was aborted; do not quote k7d-fc
resume→exec from that run.
Fixed (lifecycle resume), spec 42a: not #15. Two bugs, one per
layer. (1) Firecracker v1.16.0/v1.16.1's vsock device armed its
TRANSPORT_RESET RX gate on every PATCH /vm Resumed with no reset
event for the guest to ack, so after a bare pause → resume no
host→guest packet — no CONNECT reply — was ever delivered; upstream
fixed it as #6100 in v1.16.2, and k7d's guest/fc/pins.env plus the
vendored src/k7/deploy/k7d-fc/pins.env now pin v1.16.2. k7d's
resume_vm also dials the resumed FC guest and fails loud if it does
not answer. (2) The shim's exec bridge never opened containerd's
stdout/stderr FIFOs on its failure path, so a probe exec that could
not reach the paused guest left kubelet's prober parked in ExecSync
for its 2-minute gRPC deadline — the pod stayed Ready through a whole
pause on k7d and k7d-fc. tests/integration/test_pause_resume.py::TestPauseResumeExecAnswers
covers pause → resume → exec_command("echo hi") on both; the
PERFORMANCE.md k7d vs k7d-fc lifecycle table has the resume→exec
number. Root cause and timings: k7d CHALLENGES #240 (vsock gate) and
#241 (FIFOs).
Reference: CHALLENGES #15; k7d #232, #240, #241; Firecracker #6100.
Time lost: ~15 min on the OutOfcpu wait; lifecycle resume hung until the pytest process was killed (~8 min).
18. k7 install --backend k7d-fc cannot find vendored Firecracker installer
Symptom: A 3-node HA install with k7_backends=kfd,kql,k7d,k7d-fc (public
k7 0.3.0) failed at K7d-fc — stage install-firecracker.sh and pins on every
node:
Could not find or access 'k7d-fc/install-firecracker.sh'
Searched in:
/tmp/files/k7d-fc/install-firecracker.sh
/tmp/k7d-fc/install-firecracker.sh
... on the Ansible Controller.
K3s HA, Cilium, Longhorn, kfd thin-pool, and k7d itself had already succeeded.
Root cause: k7 install writes the embedded playbook to a tempfile
(/tmp/tmp….yaml) and runs ansible-playbook against that. Ansible copy
without remote_src looks up src: k7d-fc/install-firecracker.sh next to the
playbook file, i.e. /tmp/k7d-fc/…. The vendored files live at
src/k7/deploy/k7d-fc/ in the source tree. The API-manifest copy already
documents this trap and uses k7_repo_root; k7d-fc did not.
Fix: copy from
{{ k7_repo_root }}/src/k7/deploy/k7d-fc/{install-firecracker.sh,pins.env}
on the controller (the checkout the CLI already requires for k7-api:local).
Reference: playbook comment on "Copy K7 API manifests on first master".
Time lost: one full 3-node TWO_DISK reset + 6 min install (~45 min).
19. k7 create --docker via the API always said "this k7d has no docker service"
Symptom: After a successful 3-node HA install of k7d 0.6.0 (payload at
/usr/local/share/k7d/docker/bin/dockerd, /etc/k7/k7d_version = 0.6.0),
the docs path k7 create --docker --backend k7d --egress-open builder ubuntu:24.04
failed immediately with this k7d has no docker service; upgrade. The same
create via k7 --core (and the integration suite, which uses --core)
succeeded.
Root cause: k7d_supports_docker() treated a recorded version ≥ 0.6.0 as
"not old" and then still required os.path.isfile of the host dockerd
payload. k7-api hostPath-mounts /etc/k7 (so it can read the version file)
but not /usr/local/share/k7d, so the payload check always failed inside the
API pod. Default CLI routing is the API, so every laptop/docs user hit this;
--core on the node never did.
Fix: if the playbook recorded a parseable version, that version is
authoritative (>= 0.6.0 → supported). The payload path is only consulted
when the version file is missing (CLI --core / incomplete install).
Reference: src/k7/deploy/manifests/k7-api/deployment.yaml (/etc/k7
hostPath); K7D_DOCKER_PAYLOAD_DOCKERD in src/k7/core/docker.py.
Time lost: ~30 min diagnosing why the live cluster had dockerd but the
API refused --docker.
20. Laptop k7 api status crashed; docs k7.yaml 128Mi never went Ready
Symptom: After apt install k7 (PPA 0.3.1) on a laptop with
k7 config set api.url / api.ca / api.key, k7 list and
k7 nodes storage worked, but the Quickstart's next commands
k7 api status and k7 api endpoint raised FileNotFoundError: kubectl.
The same Quickstart's examples/k7.yaml (cpu: 100m, memory: 128Mi,
before_script: apk add curl, default backend kfd) timed out with
"Timed out waiting for sandbox container to start".
Root cause: k7 api status / endpoint always shelled out to
kubectl (or k3s kubectl) without checking the binary exists. A laptop
that only has the .deb has no kubeconfig and no kubectl. Separately, the
docs k7.yaml set memory: 128Mi, which Kata stamps as
io.katacontainers.config.hypervisor.default_memory: 128. The Firecracker
shim refuses anything below 256Mi, so the pod stays ContainerCreating
(FailedCreatePodSandBox) until create times out. before_script never
runs.
Fix: missing kubectl falls back to GET /health on the configured API
URL (status) / prints that URL (endpoint). Create rejects Kata memory
below 256Mi immediately. Docs + examples/k7.yaml use backend: k7d
and cpu: "1" / memory: "1Gi". PPA is 0.3.1, not 0.2.2.
Reference: src/k7/cli/k7.py (_kubectl_run, _api_status_via_https);
examples/k7.yaml.
Time lost: ~20 min reproducing on a 3-node HA soak after a public 0.3.1 release; the install itself was one command and succeeded.
21. --expose-port NodePort is dead until Ready; k7 exec -- sh -c self-nests
Symptom: Docs k7 create --expose-port 8000 --before-script 'nohup python3 -m http.server 8000 &' printed a NodePort, but curling it from a laptop timed out. Cilium showed the Service in maintenance. k7 exec NAME -- sh -c 'echo hi > /tmp/x' failed with Nuitka/PyInstaller-style "tried to call itself with '-c'". k7 restore --latest and k7 delete-all -y do not exist.
Root cause: externalTrafficPolicy: Local plus Cilium keeps a NodePort in maintenance while endpoints are notReadyAddresses. k7d/kfd pid 1 is sleep 365d; a bare & in before_script is killed when the script exits, so the Ready probe never sees the http server (and a missing touch of the done file has the same effect). k7 exec already wraps the joined argv in sh -c, so a nested sh -c is the CLI binary eating -c. Restore takes two positionals; delete-all confirms interactively with no -y.
Fix: Docs: trailing sleep 1 after nohup so before_script can finish and Ready can fire; curl the pod's node, not an arbitrary master. Exec: one quoted string, no extra sh -c. Restore/delete-all examples match the CLI. Create success text points at k7 list, not k7 list --name.
Reference: src/k7/cli/k7.py exec/create; docs k7/guides/cli.mdx.
Time lost: ~40 min on the public 0.3.1 HA soak (expose looked like a CNI bug until Ready flipped).
22. k7d docker graph image is a teardown race, not a leak (spec 39a)
Symptom: TestDockerK7d / TestDockerK7dFc failed after k7 delete + 3s with leaked k7d docker volume images: {scratch-vm-…-docker.img}. Hours later, same k7d PID, the files were gone.
Root cause: k7 delete returns when Kubernetes objects are gone. The VM's ScratchDisk::drop unlinks the graph image asynchronously when the containerd shim Delete lands. On a busy 3-node HA box that is more than 3s.
Fix: bounded poll (~60s) in the test. delete_sandbox does not wait for the k7d VM — coupling API latency to shim teardown would stall every delete.
Reference: tests/integration/test_docker.py _wait_k7d_docker_disks_gone; K7Core.delete_sandbox.
Time lost: the soak already established this; the 3s sleep was the only defect.
23. Firecracker jailer test missed firecracker-v1. (spec 39a)
Symptom: test_jailer_active asserted "No firecracker processes found on the host" while a jailed kfd VM was running.
Root cause: Linux truncates comm to 15 characters. The pinned binary is firecracker-v1.16.1, so comm is firecracker-v1.. The helper compared == "firecracker". The chroot binary is also versioned (firecracker-v1.16.1, not /firecracker). The jail itself is correct: /etc/passwd, /etc/shadow, /usr, /boot absent; vmlinux + rootfs present. readlink /proc/<pid>/root is / because the jailer pivot-roots in a private mount ns.
Fix: match comm by prefix firecracker-, require fcConfig.json / --config-file on the cmdline, skip (deleted) orphans, require vmlinux+rootfs, glob firecracker* in the chroot. Host-FS-unreachable asserts stay.
Reference: tests/integration/test_firecracker.py _get_live_firecracker_pids.
Time lost: ~20 min of /proc on the soak node.
24. Expose tests timed out because the pod was on another node (spec 39a)
Symptom: TestSandboxExpose curled http://<other-node>:<nodeport> from the first master and timed out. Off-cluster, the pod's node answered HTTP 200 and the other nodes refused connect.
Root cause: externalTrafficPolicy: Local is required (without it, cidr: rules see SNAT). Cilium socket-LB intercepts in-cluster-node origin to a NodePort on a different node. The tests did not pin node_name.
Fix: pin expose sandboxes to os.uname().nodename. Do not switch the Service to Cluster.
Reference: tests/integration/test_sandbox_ingress.py TestSandboxExpose._exposed.
Time lost: ~15 min confirming Local vs Cilium vs a wrong-node curl.
25. kql live --docker fork never went Ready: overlay2 was crash-inconsistent (spec 39a)
Symptom: TestDockerKQL.test_fork_clones_both_pvcs — both VolumeSnapshots ready, both child PVCs Bound, child never Ready in 240s.
Root cause (live, spec 39a): the child stuck in Init:0/2 with FailedAttachVolume: volume is not ready for workloads. Kubernetes Bound is not Longhorn-ready-to-attach; HA r=3 clone hydration takes longer than 240s. The qemu fork test already waits 600s for this. sync is also not enough for a busy overlay2 once the guest does start.
Fix: wait 600s for the child (same bound as test_qemu.test_fork_clones_data). Plus fsfreeze the graph: alpine docker:27.5.1-dind has no fsfreeze, so stage it from the ubuntu sandbox via the shared /tmp emptyDir, copy into the vehicle rootfs, freeze around VolumeSnapshot create, thaw in finally.
Reference: K7Core._create_kata_snapshots_quiesced.
Time lost: soak diagnosis; freeze is the product answer rather than refusing live forks.
26. Partner-facing 0.3.1 traps: kfd fork 404, SDK snippet missing CA, k7 logs empty
Symptom: Following docs.katakate.org / the k7 README on a 3-node HA PPA 0.3.1 cluster:
k7 forkof a kfd sandbox printedSource root PVC <name>-root-lh not found; cannot fork storageinstead of "kfd cannot fork".k7 api statusprintedClient(endpoint=..., api_key=...)with noverify_ssl; that fails against the playbook-minted cluster CA.k7 logs demo --tail 20printed nothing (exit 0) after onlyk7 exec.python3 -m venvon the node failed (ensurepip is not available).k7 resumeof a kql sandbox returned immediately while the pod was stillPending.
Root cause: fork_sandbox treated missing kfd PVCs as a generic storage 404. The status command's help snippet never grew verify_ssl when HTTPS-by-default landed. k7 logs is a CRI container snapshot; exec goes through the agent. Ubuntu 24.04 cloud images omit python3-venv. HA Longhorn attach is slower than resume() returning.
Fix: reject every kfd fork up front (KFD_FORK_REJECT). Print verify_ssl='./k7-ca.crt' in k7 api status. Docs: wait-until-Ready after kql resume, empty logs are success, SDK is a client install (apt install python3-venv on the node).
Reference: none.
Time lost: ~30 min walking the quickstart on the HA cluster.
27. k7 --core leaked kubernetes_asyncio/aiohttp sessions (Event loop is closed)
Symptom: TestSnapshotCrud.test_create_list_inspect_delete_round_trip failed with assert snap in cp.stdout and cp.stdout == '\n'. Interpreter also logged Unclosed client session / Event loop is closed after k7 --core snapshot create.
Root cause: Each CoreV1Api() / AppsV1Api() / CustomObjectsApi() constructed its own ApiClient (aiohttp session) bound to the asyncio.run loop. Typer handlers never called close(), so loop teardown raced the session destructor. Separately, _create_kata_snapshots_quiesced dropped the successful OperationResult.message from _create_volume_snapshot, so the CLI echoed a blank line even when the snapshot existed.
Fix: One shared ApiClient per K7Core, async def aclose(), CLI _core_run / API get_k7_core / snapshot-gc always close. Quiesced snapshot success now keeps the create message (Snapshot <name> created for PVC …).
Reference: kubernetes_asyncio ApiClient.close → rest_client.close() (aiohttp).
Time lost: caught on the HA integration run after the partner walkthrough.