mirror of
https://github.com/rustfs/rustfs.git
synced 2026-07-26 08:18:18 +00:00
refactor(obs): make dial9 telemetry opt-in and actually record events (#4663)
* refactor(obs): make dial9 telemetry opt-in and actually record events
The dial9 Tokio-runtime profiler was disabled by default, yet every build
paid for it, and enabling it produced trace files with no events in them.
Recorded empty traces
---------------------
`build_traced_runtime` called `TracedRuntime::builder()...build(..)`, but dial9
only starts recording in `build_and_start*`. `build` still returns a live guard
whose `is_enabled()` reports true, and still creates and seals segment files —
they just contain a header and no events. It also skipped `with_trace_path`, so
the background worker driving the segment pipeline was never spawned.
Measured on the new smoke example: 310 bytes of bare segment header, against
5640 bytes for the same workload once recording actually starts.
Switch to `with_trace_path(..).build_and_start(..)`.
Cost was unconditional
----------------------
`--cfg tokio_unstable` was a global `[build] rustflags` entry and `rustfs-obs`
depended on `dial9-tokio-telemetry` unconditionally, so all builds depended on
Tokio's non-semver API. Worse, an environment `RUSTFLAGS` replaces (never
appends to) the config-file value, so any caller exporting their own RUSTFLAGS
silently dropped the flag — the long comment in build.yml was a scar from that.
dial9 is now an opt-in feature (`dial9`, plus `dial9-s3` and `dial9-taskdump`),
the global rustflag is gone, and `crates/obs/build.rs` fails the compile if the
feature is on without the flag. Telemetry builds go through `make build-profiling`.
Metrics that could not lie
--------------------------
`rustfs_dial9_{events_total,bytes_written_total,rotations_total,cpu_overhead_percent}`
were hard-coded to zero — a Counter pinned at 0 reads as "nothing happened".
Removed. `rustfs_dial9_enabled` was sourced from the environment, so it read 1
even when the traced runtime failed and the process fell back to a standard
runtime; it is replaced by `rustfs_dial9_supported` (compile-time),
`rustfs_dial9_configured` (intent) and `rustfs_dial9_active_sessions` (reality).
No `writer_healthy` gauge is exported: dial9's `RotatingWriter` can enter its
`Finished` state and stop writing, but exposes no way to observe that, so the
gauge could only ever be hard-coded to 1. Documented as a known gap instead.
Final events were lost
----------------------
The `TelemetryGuard` lived in a `static OnceLock`, which is never dropped, so
buffered events were never flushed at exit. `build_tokio_runtime` now returns
the guard and `run_process` drops it before any exit path.
Also
----
- `disk_usage_bytes` was a `read_dir` + per-file `stat` on the metrics
collection path. It is now sampled by a background task into an atomic.
- `SAMPLING_RATE`/`S3_BUCKET`/`S3_PREFIX` were parsed, warned about, and
discarded. S3 upload is now wired to dial9's `with_s3_uploader` behind
`dial9-s3`; `SAMPLING_RATE` has no upstream equivalent and is removed.
- Wire `with_task_dumps` (async backtraces of stalled tasks), configurable via
`RUSTFS_RUNTIME_DIAL9_TASK_DUMP_{ENABLED,IDLE_THRESHOLD_MS}`.
- Split `telemetry/dial9.rs` into `config`/`state`/`enabled`/`disabled`; the
stub keeps the public API identical so callers need no `#[cfg]`.
- Drop four print-only examples and the manual test bin that exercised the
removed `init_session` scaffolding.
Verified: cargo check/clippy/test across default, `dial9`, and `dial9-s3`;
build.rs correctly rejects `dial9` without `--cfg tokio_unstable`;
`make pre-commit` passes.
Co-Authored-By: heihutu <heihutu@gmail.com>
* docs(obs): document dial9 as an on-demand profiler
scripts/run.sh advertised a `SAMPLING_RATE` knob that was never passed to dial9,
and claimed "CPU overhead < 5% (with sampling rate 1.0)" and "lower values reduce
CPU overhead" on the strength of it. The knob is gone; the guidance built on it
had to go too.
Replace it with what is actually true: dial9 needs a `make build-profiling`
binary, its disk budget evicts oldest-first (so a high poll rate can overwrite
the incident you are chasing), and it cannot be toggled without a restart.
Add docs/operations/dial9-runtime-profiling.md covering the build variants, an
investigation walkthrough, the configuration table, how to read the three
supported/configured/active_sessions gauges against each other, and the upstream
gap that makes writer death only indirectly observable.
Co-Authored-By: heihutu <heihutu@gmail.com>
* test(obs): add a dial9 smoke example that proves events are recorded
The bug this guards against is invisible to every existing signal: with
`build` instead of `build_and_start`, dial9 creates the trace file, seals
segments, and reports `TelemetryGuard::is_enabled() == true` — it simply
records no events. Only the segment's byte count tells the two apart.
Measured on this workload: 5640 bytes when recording, 310 bytes (a bare
segment header) when not. The example asserts >= 2048 bytes, and was verified
to fail with the `build` call restored.
Also correct the comment on the `is_enabled` check in `finish_traced_runtime`.
It claimed to catch "recording silently off"; it does not. It only rejects the
inert guard a lenient config yields after a build failure. Recording is
guaranteed by `build_and_start`, not by that check.
Co-Authored-By: heihutu <heihutu@gmail.com>
* test(rustfs): accept Unsupported runtime telemetry capability
A binary built without the `dial9` feature now reports the runtime-telemetry
capability as `Unsupported` rather than `Disabled`. The distinction matters to
operators: `Disabled` implies the capability can be switched on by setting an
environment variable, which is not true here — telemetry needs a rebuild.
Widen the assertion and pin the new semantics: when `dial9::is_supported()` is
false, the state must be exactly `Unsupported`.
Co-Authored-By: heihutu <heihutu@gmail.com>
* fix(obs): drop the dial9-s3 feature, its TLS stack is vulnerable
CI's Dependency Review and `cargo deny` both reject the branch: dial9's
`worker-s3` feature depends on aws-sdk-s3-transfer-manager 0.1.3, which pins
aws-smithy-http-client onto hyper-rustls 0.24 and rustls-webpki 0.101.7. That
webpki carries RUSTSEC-2026-0098, -0099 and -0104.
0.1.3 is the latest release of the transfer manager, and 1.2.0 the latest of the
smithy client, so there is nothing to upgrade to. Cargo's feature unification can
add features but cannot drop a transitive dependency, so it cannot be worked
around from here either — the rest of the workspace already resolves to the safe
rustls-webpki 0.103 / hyper-rustls 0.27.
Remove the `dial9-s3` feature and the `with_s3_uploader` wiring. The two S3
environment variables stay parsed and warned about, now naming the real reason
rather than a missing build feature. Trace segments are collected from the output
directory instead. Tracked as D9-14 in rustfs/backlog#1157.
With this, Cargo.lock is byte-identical to main: the PR no longer touches the
dependency graph at all.
Also correct the `dial9-taskdump` documentation. It claimed the feature "compiles
to a no-op elsewhere"; in fact `tokio/taskdump` raises a `compile_error!` on any
target other than linux/{aarch64,x86,x86_64}. Verified by trying to build it on
macOS, which is how the claim was found to be wrong.
Co-Authored-By: heihutu <heihutu@gmail.com>
---------
Co-authored-by: heihutu <heihutu@gmail.com>
This commit is contained in:
@@ -94,7 +94,8 @@ checked_files=(
|
||||
"crates/protocols/src/swift/quota.rs"
|
||||
"crates/protocols/src/swift/symlink.rs"
|
||||
"crates/protocols/src/common/gateway.rs"
|
||||
"crates/obs/src/telemetry/dial9.rs"
|
||||
"crates/obs/src/telemetry/dial9/config.rs"
|
||||
"crates/obs/src/telemetry/dial9/enabled.rs"
|
||||
"crates/obs/src/telemetry/local.rs"
|
||||
"crates/obs/src/metrics/scheduler.rs"
|
||||
"crates/obs/src/cleaner/core.rs"
|
||||
|
||||
+29
-25
@@ -115,18 +115,24 @@ export RUSTFS_RUNTIME_GLOBAL_QUEUE_INTERVAL=31
|
||||
# ============================================================================
|
||||
# dial9 Tokio Runtime Telemetry Configuration
|
||||
# ============================================================================
|
||||
# dial9 provides low-overhead Tokio runtime-level telemetry for performance diagnostics.
|
||||
# It captures events like PollStart/End, WorkerPark/Unpark, QueueSample, TaskSpawn.
|
||||
# dial9 captures Tokio runtime-level events (poll start/end, worker park/unpark,
|
||||
# task spawn/terminate) into binary trace segments. It sees executor-level faults
|
||||
# — a long poll stalling a worker, a task that never yields — that request-level
|
||||
# metrics and spans cannot.
|
||||
#
|
||||
# Features:
|
||||
# - CPU overhead < 5% (with sampling rate 1.0)
|
||||
# - Automatic file rotation (configurable size and count)
|
||||
# - Graceful degradation if initialization fails
|
||||
# This is an on-demand profiler, not always-on telemetry:
|
||||
#
|
||||
# Note: Disabled by default. Enable only when needed for runtime diagnostics.
|
||||
# Note: Requires build flag --cfg tokio_unstable (set in .cargo/config.toml).
|
||||
# - The stock binary does NOT support it. Build with `make build-profiling`,
|
||||
# which enables the `dial9` feature and `--cfg tokio_unstable`. Setting the
|
||||
# variables below on a stock binary logs a warning and changes nothing.
|
||||
# - Trace segments are written continuously, and the oldest are deleted once
|
||||
# MAX_FILE_SIZE * ROTATION_COUNT bytes are retained. Under a high poll rate
|
||||
# that budget can wrap in minutes — size it against the window you need.
|
||||
# - Turn it off again when the investigation is over.
|
||||
#
|
||||
# See docs/operations/dial9-runtime-profiling.md.
|
||||
|
||||
# Enable dial9 telemetry (default: false)
|
||||
# Enable dial9 telemetry (default: false; requires a `make build-profiling` binary)
|
||||
#export RUSTFS_RUNTIME_DIAL9_ENABLED=true
|
||||
|
||||
# Output directory for trace files (default: /var/log/rustfs/telemetry)
|
||||
@@ -138,33 +144,31 @@ export RUSTFS_RUNTIME_GLOBAL_QUEUE_INTERVAL=31
|
||||
# Maximum trace file size in bytes (default: 104857600 = 100MB)
|
||||
#export RUSTFS_RUNTIME_DIAL9_MAX_FILE_SIZE=104857600
|
||||
|
||||
# Number of rotated files to keep (default: 10)
|
||||
# Number of rotated files to keep (default: 10). Total disk budget is
|
||||
# MAX_FILE_SIZE * ROTATION_COUNT; older segments are evicted, not kept.
|
||||
#export RUSTFS_RUNTIME_DIAL9_ROTATION_COUNT=10
|
||||
|
||||
# Sampling rate: 0.0 to 1.0 (default: 1.0 = 100% sampling)
|
||||
# Lower values reduce CPU overhead. Recommended: 0.1-0.5 for production.
|
||||
#export RUSTFS_RUNTIME_DIAL9_SAMPLING_RATE=1.0
|
||||
# Capture async backtraces of tasks that stall (default: false).
|
||||
# Needs a binary built with the `dial9-taskdump` feature (Linux only).
|
||||
#export RUSTFS_RUNTIME_DIAL9_TASK_DUMP_ENABLED=true
|
||||
# Mean idle duration for task-dump Poisson sampling, in ms (default: 10)
|
||||
#export RUSTFS_RUNTIME_DIAL9_TASK_DUMP_IDLE_THRESHOLD_MS=10
|
||||
|
||||
# S3 upload settings (not yet implemented; reserved for future use):
|
||||
# S3 upload is NOT available: dial9's uploader depends on a rustls-webpki with
|
||||
# known CVEs, so the feature is not built. These are parsed and warned about,
|
||||
# never honoured. Collect segments from OUTPUT_DIR instead.
|
||||
#export RUSTFS_RUNTIME_DIAL9_S3_BUCKET=my-trace-bucket
|
||||
#export RUSTFS_RUNTIME_DIAL9_S3_PREFIX=telemetry/
|
||||
|
||||
# --- Scenario 1: Development / Debugging ---
|
||||
# Full tracing with local storage, high sampling rate
|
||||
# --- Scenario 1: local development ---
|
||||
#export RUSTFS_RUNTIME_DIAL9_ENABLED=true
|
||||
#export RUSTFS_RUNTIME_DIAL9_OUTPUT_DIR="$current_dir/deploy/telemetry"
|
||||
#export RUSTFS_RUNTIME_DIAL9_SAMPLING_RATE=1.0
|
||||
|
||||
# --- Scenario 2: Production Diagnostics ---
|
||||
# Reduced sampling rate to minimize overhead
|
||||
#export RUSTFS_RUNTIME_DIAL9_ENABLED=true
|
||||
#export RUSTFS_RUNTIME_DIAL9_SAMPLING_RATE=0.1
|
||||
|
||||
# --- Scenario 3: Performance Investigation ---
|
||||
# Short-term tracing with high detail, manual cleanup
|
||||
# --- Scenario 2: investigating a worker stall ---
|
||||
# Task dumps show where stalled tasks are parked. Keep the run short.
|
||||
#export RUSTFS_RUNTIME_DIAL9_ENABLED=true
|
||||
#export RUSTFS_RUNTIME_DIAL9_OUTPUT_DIR=/tmp/rustfs-telemetry-investigation
|
||||
#export RUSTFS_RUNTIME_DIAL9_SAMPLING_RATE=1.0
|
||||
#export RUSTFS_RUNTIME_DIAL9_TASK_DUMP_ENABLED=true
|
||||
#export RUSTFS_RUNTIME_DIAL9_ROTATION_COUNT=3
|
||||
|
||||
export OTEL_INSTRUMENTATION_NAME="rustfs"
|
||||
|
||||
Reference in New Issue
Block a user