Files
rustfs/docs/superpowers/plans/profiling-numa-capability-inventory.md
T
Zhengchao An f63af3df63 chore: retire completed-migration scaffolding, wire orphaned boundary check (#4719)
The ecstore/global-state migrations are done (backlog#815, #939, #1052 all
closed). Review of every migration-era test/gate measure found three things
actually retirable or broken — everything else is a live anti-regression
guard and stays.

Remove:
- scripts/check_metrics_migration_refs.sh — guards a migration that
  finished: rustfs_metrics:: has zero hits, the metrics crate no longer
  exists, and the script was never wired into CI or make (only reference
  was one line in config-model-boundary-adr.md, also removed).
- crates/obs init_metrics_collectors — the "backward-compatible alias kept
  during migration" the removed script was guarding. Zero callers; pure
  delegate to init_metrics_runtime.

Archive (docs/superpowers/plans/, continuing the 2026-07 convention,
with the standard archived banner):
- startup-timeline.md, scheduler-baseline.md,
  profiling-numa-capability-inventory.md,
  kms-development-defaults-inventory.md — one-shot snapshots whose only
  consumer is the already-archived migration-progress ledger (their
  same-dir links there start resolving again after the move); zero script
  pins; fed the closed backlog#660/#665 architecture-review ledger.
  Fixed the one outbound link (startup-timeline -> readiness-matrix) that
  the move would have broken — check_doc_paths.sh deliberately does not
  scan plans/, so nothing else would have caught it.

Wire (found orphaned by the same review):
- scripts/check_extension_schema_boundaries.sh guards a live contract
  crate but was never invoked anywhere. Add lint-fmt.mak target, include
  in pre-commit/pre-pr/dev-check, add ci.yml Quick Checks step (job
  already installs ripgrep), sync the CONTRIBUTING.md enumerated list,
  and harden the script against a silently-passing rg probe when src/
  is missing.

Keep (verified live, documented so the next cleanup pass does not repeat
this analysis):
- scripts/check_architecture_migration_rules.sh — added a header stating
  it is a permanent boundary guard, not retirable migration scaffolding;
  'migration' in the name is historical.
- check_migration_gate_count.sh + floor, delete-marker e2e proof, all
  pinned docs, compat-cleanup-register sync, remaining inventories
  (referenced by live docs).

Verification: all 7 guard scripts pass, actionlint clean,
cargo check --workspace (excl e2e) clean, cargo fmt --check clean.
Adversarially reviewed by two independent skeptic passes; their 7
findings (alias left behind, broken outbound link, missing banners,
wrong backlog attribution, CONTRIBUTING drift, rg exit-2 hole, missing
header rationale) are all folded in.
2026-07-11 13:42:56 +08:00

5.3 KiB

Archived migration snapshot — moved from docs/architecture/ (2026-07) when the architecture-review ledger it fed closed out. Kept for history; not maintained.

Profiling And NUMA Capability Inventory

This inventory covers G-013 for rustfs/backlog#667. It records the current profiling, memory sampling, allocator, and NUMA baseline before optional runtime sidecars are designed.

Platform Support Matrix

Capability Current support Current owner Baseline recommendation
CPU pprof dump Linux and macOS builds use pprof; other targets return an unsupported-platform error. rustfs/src/profiling.rs Keep CPU profiling opt-in through existing env flags and cancellation token.
Continuous CPU profiling Linux and macOS builds can hold a continuous ProfilerGuard when enabled. rustfs/src/profiling.rs Preserve single-guard ownership and avoid starting multiple continuous guards.
Periodic CPU profiling Linux and macOS builds can spawn a periodic sampling loop. rustfs/src/profiling.rs Keep the loop cancellation-driven and non-fatal.
Jemalloc memory pprof Only linux + gnu + x86_64 exposes jemalloc pprof dumping. Other supported builds return an unsupported-target error. rustfs/src/profiling.rs Treat memory pprof as optional and target-gated.
Periodic memory pprof Only runs where jemalloc profiling control is available and active. rustfs/src/profiling.rs Keep inactive jemalloc as a skipped dump, not a startup failure.
Process/system memory sampling Uses rustfs_io_metrics::snapshot_process_resource_and_system plus sysinfo total memory. rustfs/src/memory_observability.rs Keep sampling portable and metric-gated.
cgroup memory sampling Reads Linux cgroup v2 or v1 memory files when present. Missing files produce no cgroup split. rustfs/src/memory_observability.rs Keep cgroup data opportunistic and absent-safe.
Allocator reclaim Uses jemalloc backend on linux + gnu + x86_64; otherwise mimalloc variants. rustfs/src/allocator_reclaim.rs Keep backend detection read-only and preserve effective-force behavior.
eBPF No runtime eBPF sidecar is currently wired into startup. N/A Treat eBPF as future optional Linux-only inventory, never as a required baseline.
NUMA No NUMA placement or topology controller is currently wired into startup. N/A Treat NUMA as future optional capability with no-op fallback.

Cross-Platform Baseline

The current safe baseline is:

  • Profiling is opt-in through env flags and must not make startup fatal.
  • Startup and shutdown call profiling through startup_profiling lifecycle hooks; profiling.rs remains the CPU/memory profiling implementation and admin dump API owner.
  • Unsupported profiling targets return structured unsupported errors or skip startup tasks.
  • Memory observability records process/system metrics and adds cgroup split only when cgroup files exist.
  • Allocator reclaim observes active HTTP, delete-tail, scanner, heal, erasure, and GET-buffer activity before reclaiming.
  • Runtime thread sizing remains owned by the Tokio runtime builder and sysinfo core detection, not NUMA topology.

Optional Sidecar Invariants

Future sidecars for profiling, eBPF, or NUMA must preserve these invariants:

  • Sidecars must be disabled by default or target-gated until explicitly enabled.
  • Unsupported targets must degrade to no-op status, not panic or fail startup.
  • Sidecars must use the runtime cancellation token or an equivalent explicit shutdown handle.
  • Sidecars must not mutate Tokio worker counts after runtime creation.
  • Profiling output directory fallback must stay local to profiling and must not affect object storage paths.
  • NUMA fallback must preserve current runtime thread defaults, storage set placement, and request admission behavior.

First Implementation Candidates

API-013:

  • Define a read-only capability contract for profiling, cgroup memory, eBPF, allocator backend, and NUMA availability.
  • Keep the contract in a low-dependency crate and report unsupported states explicitly.

R-016:

  • Wire storage runtime startup to consume capability snapshots read-only.
  • Do not start sidecars or mutate runtime worker ownership in the same PR.

X-012:

  • Define the ops.profiler.v1 extension schema for profiling capability reporting, backend status, redaction requirements, and provenance.
  • Keep the schema capability-only; it must not request profiler execution, start sidecars, or change profile export behavior.
  • Keep unsupported targets, disabled sidecars, and unknown future backends representable as no-op capability states.

X-013:

  • Add the extension capability snapshot contract for disabled, unsupported, and enabled profiler backends.
  • Verify optional profiler sidecar and Wasm runtimes stay disabled by default and cannot declare a startup fatal boundary.

R-021:

  • If a runtime service sidecar is added later, enter it through the optional runtime boundary with explicit shutdown ownership.
  • Preserve current service order, KMS/audit/notification fatal boundaries, and scanner/heal startup semantics.

R-022:

  • Keep optional runtime startup handoff in startup_optional_runtimes while leaving concrete protocol adapters in startup_protocols.
  • Preserve KMS-before-protocol startup ordering and disabled protocol no-op behavior.