mirror of
https://github.com/rustfs/rustfs.git
synced 2026-07-26 08:18:18 +00:00
f63af3df63
The ecstore/global-state migrations are done (backlog#815, #939, #1052 all closed). Review of every migration-era test/gate measure found three things actually retirable or broken — everything else is a live anti-regression guard and stays. Remove: - scripts/check_metrics_migration_refs.sh — guards a migration that finished: rustfs_metrics:: has zero hits, the metrics crate no longer exists, and the script was never wired into CI or make (only reference was one line in config-model-boundary-adr.md, also removed). - crates/obs init_metrics_collectors — the "backward-compatible alias kept during migration" the removed script was guarding. Zero callers; pure delegate to init_metrics_runtime. Archive (docs/superpowers/plans/, continuing the 2026-07 convention, with the standard archived banner): - startup-timeline.md, scheduler-baseline.md, profiling-numa-capability-inventory.md, kms-development-defaults-inventory.md — one-shot snapshots whose only consumer is the already-archived migration-progress ledger (their same-dir links there start resolving again after the move); zero script pins; fed the closed backlog#660/#665 architecture-review ledger. Fixed the one outbound link (startup-timeline -> readiness-matrix) that the move would have broken — check_doc_paths.sh deliberately does not scan plans/, so nothing else would have caught it. Wire (found orphaned by the same review): - scripts/check_extension_schema_boundaries.sh guards a live contract crate but was never invoked anywhere. Add lint-fmt.mak target, include in pre-commit/pre-pr/dev-check, add ci.yml Quick Checks step (job already installs ripgrep), sync the CONTRIBUTING.md enumerated list, and harden the script against a silently-passing rg probe when src/ is missing. Keep (verified live, documented so the next cleanup pass does not repeat this analysis): - scripts/check_architecture_migration_rules.sh — added a header stating it is a permanent boundary guard, not retirable migration scaffolding; 'migration' in the name is historical. - check_migration_gate_count.sh + floor, delete-marker e2e proof, all pinned docs, compat-cleanup-register sync, remaining inventories (referenced by live docs). Verification: all 7 guard scripts pass, actionlint clean, cargo check --workspace (excl e2e) clean, cargo fmt --check clean. Adversarially reviewed by two independent skeptic passes; their 7 findings (alias left behind, broken outbound link, missing banners, wrong backlog attribution, CONTRIBUTING drift, rg exit-2 hole, missing header rationale) are all folded in.
5.3 KiB
5.3 KiB
Archived migration snapshot — moved from
docs/architecture/(2026-07) when the architecture-review ledger it fed closed out. Kept for history; not maintained.
Profiling And NUMA Capability Inventory
This inventory covers G-013 for rustfs/backlog#667. It records the current
profiling, memory sampling, allocator, and NUMA baseline before optional runtime
sidecars are designed.
Platform Support Matrix
| Capability | Current support | Current owner | Baseline recommendation |
|---|---|---|---|
| CPU pprof dump | Linux and macOS builds use pprof; other targets return an unsupported-platform error. |
rustfs/src/profiling.rs |
Keep CPU profiling opt-in through existing env flags and cancellation token. |
| Continuous CPU profiling | Linux and macOS builds can hold a continuous ProfilerGuard when enabled. |
rustfs/src/profiling.rs |
Preserve single-guard ownership and avoid starting multiple continuous guards. |
| Periodic CPU profiling | Linux and macOS builds can spawn a periodic sampling loop. | rustfs/src/profiling.rs |
Keep the loop cancellation-driven and non-fatal. |
| Jemalloc memory pprof | Only linux + gnu + x86_64 exposes jemalloc pprof dumping. Other supported builds return an unsupported-target error. |
rustfs/src/profiling.rs |
Treat memory pprof as optional and target-gated. |
| Periodic memory pprof | Only runs where jemalloc profiling control is available and active. | rustfs/src/profiling.rs |
Keep inactive jemalloc as a skipped dump, not a startup failure. |
| Process/system memory sampling | Uses rustfs_io_metrics::snapshot_process_resource_and_system plus sysinfo total memory. |
rustfs/src/memory_observability.rs |
Keep sampling portable and metric-gated. |
| cgroup memory sampling | Reads Linux cgroup v2 or v1 memory files when present. Missing files produce no cgroup split. | rustfs/src/memory_observability.rs |
Keep cgroup data opportunistic and absent-safe. |
| Allocator reclaim | Uses jemalloc backend on linux + gnu + x86_64; otherwise mimalloc variants. |
rustfs/src/allocator_reclaim.rs |
Keep backend detection read-only and preserve effective-force behavior. |
| eBPF | No runtime eBPF sidecar is currently wired into startup. | N/A | Treat eBPF as future optional Linux-only inventory, never as a required baseline. |
| NUMA | No NUMA placement or topology controller is currently wired into startup. | N/A | Treat NUMA as future optional capability with no-op fallback. |
Cross-Platform Baseline
The current safe baseline is:
- Profiling is opt-in through env flags and must not make startup fatal.
- Startup and shutdown call profiling through
startup_profilinglifecycle hooks;profiling.rsremains the CPU/memory profiling implementation and admin dump API owner. - Unsupported profiling targets return structured unsupported errors or skip startup tasks.
- Memory observability records process/system metrics and adds cgroup split only when cgroup files exist.
- Allocator reclaim observes active HTTP, delete-tail, scanner, heal, erasure, and GET-buffer activity before reclaiming.
- Runtime thread sizing remains owned by the Tokio runtime builder and sysinfo core detection, not NUMA topology.
Optional Sidecar Invariants
Future sidecars for profiling, eBPF, or NUMA must preserve these invariants:
- Sidecars must be disabled by default or target-gated until explicitly enabled.
- Unsupported targets must degrade to no-op status, not panic or fail startup.
- Sidecars must use the runtime cancellation token or an equivalent explicit shutdown handle.
- Sidecars must not mutate Tokio worker counts after runtime creation.
- Profiling output directory fallback must stay local to profiling and must not affect object storage paths.
- NUMA fallback must preserve current runtime thread defaults, storage set placement, and request admission behavior.
First Implementation Candidates
API-013:
- Define a read-only capability contract for profiling, cgroup memory, eBPF, allocator backend, and NUMA availability.
- Keep the contract in a low-dependency crate and report unsupported states explicitly.
R-016:
- Wire storage runtime startup to consume capability snapshots read-only.
- Do not start sidecars or mutate runtime worker ownership in the same PR.
X-012:
- Define the
ops.profiler.v1extension schema for profiling capability reporting, backend status, redaction requirements, and provenance. - Keep the schema capability-only; it must not request profiler execution, start sidecars, or change profile export behavior.
- Keep unsupported targets, disabled sidecars, and unknown future backends representable as no-op capability states.
X-013:
- Add the extension capability snapshot contract for disabled, unsupported, and enabled profiler backends.
- Verify optional profiler sidecar and Wasm runtimes stay disabled by default and cannot declare a startup fatal boundary.
R-021:
- If a runtime service sidecar is added later, enter it through the optional runtime boundary with explicit shutdown ownership.
- Preserve current service order, KMS/audit/notification fatal boundaries, and scanner/heal startup semantics.
R-022:
- Keep optional runtime startup handoff in
startup_optional_runtimeswhile leaving concrete protocol adapters instartup_protocols. - Preserve KMS-before-protocol startup ordering and disabled protocol no-op behavior.