* feat(obs): add drive topology detail metrics Expose additive drive info, topology, state, and per-drive API metrics while preserving the existing drive metric label sets. Backlog: rustfs/backlog#1655 Co-Authored-By: heihutu <heihutu@gmail.com> * fix(obs): preserve suspect drive runtime state Keep suspect as a bounded drive runtime state and avoid all-zero runtime_state samples for that storage health state. Co-Authored-By: heihutu <heihutu@gmail.com> * fix(obs): skip unknown drive inode samples Avoid exporting zero inode gauges for missing or stale drive snapshots and ignore zero-count API latency buckets. Co-Authored-By: heihutu <heihutu@gmail.com> * feat(obs): add scanner source work detail metrics Expose additive scanner source and cycle work metrics with bounded server/source/state labels while leaving the existing aggregate scanner metrics unchanged. Co-Authored-By: heihutu <heihutu@gmail.com> * feat(obs): add ilm action detail metrics Expose additive ILM action/state task metrics with a server label while preserving the existing aggregate ILM series. Co-Authored-By: heihutu <heihutu@gmail.com> * feat(obs): add delivery target server metrics Expose additive audit and notification delivery target metrics with server labels and extend removed-target tombstones for the server-aware series. Co-Authored-By: heihutu <heihutu@gmail.com> * feat(obs): add replication target flow metrics Expose additive bucket replication target sent and failed-flow metrics while preserving existing bucket aggregates and target backlog series. Co-Authored-By: heihutu <heihutu@gmail.com> * feat(obs): add request server metrics Expose additive API request metrics with server labels while preserving the existing request and traffic metric label sets. Co-Authored-By: heihutu <heihutu@gmail.com> * style(obs): apply rustfmt to metrics changes Apply rustfmt output to the metrics dimension changes without altering behavior. Co-Authored-By: heihutu <heihutu@gmail.com> * style(obs): reuse audit target label constant Use the exported audit target_id label constant for legacy audit target metrics. Co-Authored-By: heihutu <heihutu@gmail.com> * feat(obs): populate drive disk metrics Co-Authored-By: heihutu <heihutu@gmail.com> * feat(obs): add scanner bucket drive result metrics Co-Authored-By: heihutu <heihutu@gmail.com> * feat(obs): add replication proxy server metrics Co-Authored-By: heihutu <heihutu@gmail.com> * fix(obs): address metric liveness review Use checked division for drive API latency aggregation and keep recovered drive, scanner current-cycle, replication flow, audit target, and notification target series from retaining stale values. Co-Authored-By: heihutu <heihutu@gmail.com> * fix(obs): address metric dimension review Co-Authored-By: heihutu <heihutu@gmail.com> * fix(obs): address additional metric review Co-Authored-By: heihutu <heihutu@gmail.com> * fix(obs): count drive calls at start Co-Authored-By: heihutu <heihutu@gmail.com> * fix(obs): address metrics dimension review Co-Authored-By: heihutu <heihutu@gmail.com> * fix(metrics): address dimension review gaps Co-Authored-By: heihutu <heihutu@gmail.com> * fix(metrics): address scanner review follow-ups Co-Authored-By: heihutu <heihutu@gmail.com> * fix(metrics): address runtime review follow-ups Co-Authored-By: heihutu <heihutu@gmail.com> * fix(metrics): reduce disk metric contention Co-Authored-By: heihutu <heihutu@gmail.com> * fix(metrics): address runtime review follow-ups Co-Authored-By: heihutu <heihutu@gmail.com> * fix(metrics): retire stale dimension series Co-Authored-By: heihutu <heihutu@gmail.com> --------- Co-authored-by: heihutu <heihutu@gmail.com>
RustFS Observability Stack
This directory contains the comprehensive observability stack for RustFS, designed to provide deep insights into application performance, logs, and traces.
Components
The stack is composed of the following best-in-class open-source components:
- Prometheus (v2.53.1): The industry standard for metric collection and alerting.
- Grafana (v11.1.0): The leading platform for observability visualization.
- Loki (v3.1.0): A horizontally-scalable, highly-available, multi-tenant log aggregation system.
- Tempo (v2.5.0): A high-volume, minimal dependency distributed tracing backend.
- Jaeger (v1.59.0): Distributed tracing system (configured as a secondary UI/storage).
- OpenTelemetry Collector (v0.104.0): A vendor-agnostic implementation for receiving, processing, and exporting telemetry data.
By default, this stack uses Tempo in single-binary mode and does not require Kafka/Redpanda.
If you want the Kafka-backed HA Tempo path, use docker-compose-example-for-rustfs.yml together with docker-compose-tempo-ha-override.yml.
Architecture
- Telemetry Collection: Applications send OTLP (OpenTelemetry Protocol) data (Metrics, Logs, Traces) to the OpenTelemetry Collector.
- Processing & Exporting: The Collector processes the data (batching, memory limiting) and exports it to the respective backends:
- Traces -> Tempo (Primary) & Jaeger (Secondary/Optional)
- Metrics -> Prometheus (via scraping the Collector's exporter)
- Logs -> Loki
- Visualization: Grafana connects to all backends (Prometheus, Tempo, Loki, Jaeger) to provide a unified dashboard experience.
Features
- Full Persistence: All data (Metrics, Logs, Traces) is persisted to Docker volumes, ensuring no data loss on restart.
- Correlation: Seamless navigation between Metrics, Logs, and Traces in Grafana.
- Jump from a Metric spike to relevant Traces.
- Jump from a Trace to relevant Logs.
- High Performance: Optimized configurations for batching, compression, and memory management.
- Standardized Protocols: Built entirely on OpenTelemetry standards.
GET Performance Optimization Dashboards
Three pre-built Grafana dashboards are included for monitoring RustFS GET performance optimization rollout:
Available Dashboards
| Dashboard | File | Description |
|---|---|---|
| GET Rollout Health | grafana-get-rollout-health.json |
Monitors optimization rollout: latency by reader path, early-stop hit rate, codec streaming usage, pipeline failures |
| GET Data Integrity | grafana-get-data-integrity.json |
Monitors data safety: bitrot verify failures, decode errors, short reads, shard read outcomes |
| GET Resource Impact | grafana-get-resource-impact.json |
Monitors resource usage: concurrent requests, IO queue utilization, disk permit wait, RSS trend |
| Object Data Cache | grafana-object-data-cache.json |
Monitors the GET body cache (rustfs_object_data_cache_*): hit ratio, lookup/plan/fill outcomes, fill duration quantiles, hit vs fill throughput, entries/weighted bytes, inflight fills, memory-pressure skips, invalidations, and size-class breakdowns |
Prometheus Alert Rules
The file prometheus-rules/rustfs-get-optimization-alerts.yaml contains pre-configured alerting rules:
| Alert | Severity | Condition |
|---|---|---|
GetP99Regression |
Critical | GET p99 latency > 2x baseline for 10m |
PipelineFailureSpike |
Critical | Pipeline failure rate > 5x baseline for 5m |
BitrotMismatchSpike |
Critical | Bitrot mismatch rate > 3x baseline for 5m |
EarlyStopInsufficientQuorum |
Warning | Early-stop insufficient quorum rate > 0.1/s for 5m |
CodecStreamingFallbackSpike |
Warning | Codec streaming fallback > 10x baseline for 10m |
IoQueueSaturation |
Warning | IO queue utilization > 90% for 5m |
The file prometheus-rules/rustfs-kms-alerts.yml contains alerting rules for the KMS backend operation metrics. Thresholds are conservative defaults pending staging baseline calibration; response procedures live in docs/operations/kms-observability-runbook.md, and the matching dashboard is deploy/observability/grafana/rustfs-kms-observability.json.
| Alert | Severity | Condition |
|---|---|---|
KmsBackendFatalErrors |
Critical | Fatal (non-retryable) attempt failures > 0 for 5m |
KmsBackendHighErrorRate |
Critical | Non-success operation ratio > 5% for 10m (with traffic guard) |
KmsBackendP99LatencyHigh |
Warning | Operation p99 duration (incl. retries) > 2s for 10m |
KmsBackendAttemptFailureSpike |
Warning | Attempt failure rate > 0.5/s for 10m |
KmsBackendRetryBudgetExhausted |
Warning | budget_exhausted / deadline_exceeded outcomes > 0.05/s for 10m |
Enabling Alert Rules
Add the alert rules file to your Prometheus configuration:
# prometheus.yml
rule_files:
- "/etc/prometheus/rules/*.yml"
# Or mount the file in docker-compose.yml:
# volumes:
# - ./prometheus-rules:/etc/prometheus/rules
Dashboard Usage
The dashboards are automatically provisioned when Grafana starts. They use the ${DS_PROMETHEUS} datasource variable, so you need a Prometheus datasource configured in Grafana.
Key panels to monitor during optimization rollout:
- GET Latency by Reader Path - Compare
codec_streamingvslegacy_duplexlatency - Early-Stop Hit Rate - Verify early-stop is triggering effectively
- Pipeline Failure Rate - Detect any new failure modes introduced by optimizations
- Bitrot Verify Failures - Ensure data integrity is maintained
Quick Start
Prerequisites
- Docker
- Docker Compose
Deploy
Run the following command to start the entire stack:
docker compose up -d
High Availability Tempo
The default docker-compose.yml is the single-node stack.
If you need the Kafka-backed HA Tempo configuration, start it with:
docker compose -f docker-compose-example-for-rustfs.yml -f docker-compose-tempo-ha-override.yml up -d
Access Dashboards
| Service | URL | Credentials | Description |
|---|---|---|---|
| Grafana | http://localhost:3000 | admin / admin |
Main visualization hub. |
| Prometheus | http://localhost:9090 | - | Metric queries and status. |
| Jaeger UI | http://localhost:16686 | - | Secondary trace visualization. |
| Tempo | http://localhost:3200 | - | Tempo status/metrics. |
Configuration
Data Persistence
Data is stored in the following Docker volumes:
prometheus-data: Prometheus metricstempo-data: Tempo traces (WAL and Blocks)loki-data: Loki logs (Chunks and Rules)jaeger-data: Jaeger traces (Badger DB)
To clear all data:
docker compose down -v
Customization
- Prometheus: Edit
prometheus.ymlto add scrape targets or alerting rules. - Grafana: Dashboards and datasources are provisioned from the
grafana/directory. - Collector: Edit
otel-collector-config.yamlto modify pipelines, processors, or exporters.
Verifying RustFS Traces
When RustFS points RUSTFS_OBS_ENDPOINT at this stack, treat the value as the
OTLP/HTTP base URL, for example:
export RUSTFS_OBS_ENDPOINT=http://host.docker.internal:4318
RustFS automatically expands that base URL to:
/v1/traces/v1/metrics/v1/logs
Important behavior notes:
- Logs and metrics usually appear during startup, so seeing those two signals first is expected.
- Visible trace data usually requires real HTTP/S3/gRPC request traffic after startup, because request-path spans are created on demand.
RUSTFS_OBS_LOGGER_LEVEL=infokeeps the top-level request span but filters many nesteddebugspans. If Tempo or Jaeger looks sparse, retry withRUSTFS_OBS_LOGGER_LEVEL=debugbefore suspecting collector or Tempo issues.
Minimal validation flow:
# 1. Start this observability stack.
docker compose up -d
# 2. Start RustFS with OTLP/HTTP export and richer span visibility.
export RUSTFS_OBS_ENDPOINT=http://host.docker.internal:4318
export RUSTFS_OBS_LOGGER_LEVEL=debug
# 3. Generate real request traffic.
curl -I http://127.0.0.1:9000/health
curl -I http://127.0.0.1:9000/health/ready
# 4. Inspect Grafana or Jaeger.
# Grafana: http://localhost:3000
# Jaeger: http://localhost:16686
If logs and metrics are present but traces are sparse, the most common cause is
"no real request traffic yet" or "info level filtered nested spans", not an
OTLP routing failure.
Troubleshooting
- Service Health: Check the health of services using
docker compose ps. - Logs: View logs for a specific service using
docker compose logs -f <service_name>. - Otel Collector: Check
http://localhost:13133for health status andhttp://localhost:1888/debug/pprof/for profiling.