Files
rustfs/crates/config
houseme 60ad15a7f9 fix(obs): remove the dial9 task-dump switch that could never work; correct measured claims (#4688)
* fix(obs): remove the dial9 task-dump switch that could never do anything

Measured on a bench host (Linux x86_64) against the code merged in #4663: with
`RUSTFS_RUNTIME_DIAL9_TASK_DUMP_ENABLED=true`, a `dial9-taskdump` build, and
`--cfg tokio_taskdump`, dial9 recorded **zero** TaskDump events.

dial9 captures a task dump only for futures it wrapped itself — those spawned
through `dial9_tokio_telemetry::spawn`, which is where `TaskDumped<F>` gets
applied. `tokio::spawn` gets no wrapper, and RustFS spawns with `tokio::spawn`
throughout. Same workload, same binary, only the spawner changed:

    tokio::spawn   ->      0 dumps
    dial9::spawn   ->  14709 dumps, all with callchains

Upstream documents this (README line 151) and tracks the doc gap at
dial9-rs/dial9#477. I did not read it before wiring `with_task_dumps` in #4663,
and so shipped exactly the kind of lying configuration knob that PR set out to
delete. Remove it: the two environment variables, the config fields, the
`with_task_dumps` call, and the `dial9-taskdump` feature — whose only effect was
to constrain the build to Linux while recording nothing.

Re-adding it only makes sense together with migrating the paths under
investigation to dial9's spawner. Tracked as D9-16 in rustfs/backlog#1157.

Also drop the `--cfg tokio_taskdump` requirement from the Makefile. Measured:
dumps are captured with and without it (14709 vs 14674, within noise), and
upstream never asked for it. That requirement was mine, invented and untested.

Cargo.lock loses tokio's `backtrace` dependency, which `tokio/taskdump` pulled in.

Co-Authored-By: heihutu <heihutu@gmail.com>

* docs(obs): replace guessed dial9 retention numbers with measured ones

Three corrections, all to claims I wrote in #4663 without measuring them.

"Under a high poll rate that budget can wrap in minutes" was a guess. Measured on
a single-node 4-drive cluster under warp mixed (66 MiB/s, 110 obj/s, 32 concurrent):
13023 events/s, 0.16 MiB/s, so the default 1 GiB budget wraps after roughly 108
minutes. Even at ten times the throughput that is ~11 minutes. State the measured
rate and how to scale it instead.

dial9 was described as the tool for drive stalls. It is not. RustFS does disk I/O
on the blocking pool and through io_uring, never on an async worker, so a slow
drive never lengthens a poll. Injecting 200 ms of latency on one of four drives
cut throughput by 64% and left the poll distribution unchanged (polls >= 5 ms:
49 -> 56; p999: 2.67 ms -> 2.75 ms). Enabling dial9's CPU and sched profilers
does not help: sched events are per-worker only, and the CPU profiler samples
on-CPU while a stalled drive is an off-CPU wait. Say so plainly, and point at
the `rustfs_io_*` metrics instead.

What dial9 *is* good for, on the same traces: single polls of 418-625 ms with no
fault injected at all — real worker stalls nothing else in the obs stack surfaces.
Lead with that.

Also link the two upstream issues filed for the gaps we documented:
dial9-rs/dial9#658 (writer death unobservable) and #659 (worker-s3 CVEs).

Measurements: rustfs/backlog#1157 (D9-11, D9-13, D9-18).

Co-Authored-By: heihutu <heihutu@gmail.com>

---------

Co-authored-by: heihutu <heihutu@gmail.com>
2026-07-10 17:34:59 +00:00
..

RustFS

RustFS Config - Configuration Management

Configuration management and validation module for RustFS distributed object storage

CI 📖 Documentation · 🐛 Bug Reports · 💬 Discussions


📖 Overview

RustFS Config provides configuration management and validation capabilities for the RustFS distributed object storage system. For the complete RustFS experience, please visit the main RustFS repository.

Features

  • Multi-format configuration support (TOML, YAML, JSON, ENV)
  • Environment variable integration and override
  • Configuration validation and type safety
  • Hot-reload capabilities for dynamic updates
  • Default value management and fallbacks
  • Secure credential handling and encryption

📚 Documentation

For comprehensive documentation, examples, and usage guides, please visit the main RustFS repository.

Environment Variable Naming Conventions

RustFS uses a flat naming style for top-level configuration: environment variables are RUSTFS_* without nested module segments.

Examples:

  • RUSTFS_REGION
  • RUSTFS_ADDRESS
  • RUSTFS_VOLUMES
  • RUSTFS_LICENSE
  • RUSTFS_LICENSE_PUBLIC_KEY

Current guidance:

  • Prefer module-specific names only when they are not top-level product configuration.
  • Renamed variables must keep backward-compatible aliases until before beta.
  • Alias usage must emit deprecation warnings and be treated as transitional only.
  • Deprecated example:
    • RUSTFS_ENABLE_SCANNER -> RUSTFS_SCANNER_ENABLED
    • RUSTFS_ENABLE_HEAL -> RUSTFS_HEAL_ENABLED
    • RUSTFS_DATA_SCANNER_START_DELAY_SECS -> RUSTFS_SCANNER_START_DELAY_SECS

License environment variables

  • RUSTFS_LICENSE contains the signed license token.
  • RUSTFS_LICENSE_PUBLIC_KEY contains the RSA public key used to verify signed license tokens.

CORS environment variables

  • RUSTFS_CORS_ALLOWED_ORIGINS defaults to empty, so the S3 endpoint emits no generic CORS headers unless configured. Set * for wildcard origins without credentials, or a comma-separated allow-list for credentialed explicit origins.
  • RUSTFS_CONSOLE_CORS_ALLOWED_ORIGINS defaults to * for the console service.

Browser redirect environment variables

  • RUSTFS_BROWSER_REDIRECT_URL sets the externally reachable browser origin used for OIDC callback, console success redirect, and logout fallback URLs. Configure it to the public scheme and authority without a path, for example https://console.example.com. In load-balancer deployments, keep OIDC authorize and callback requests on the same backend node because the in-flight OIDC state is local to the RustFS node.

Scanner environment aliases

  • RUSTFS_SCANNER_SPEED (canonical, also accepts MINIO_SCANNER_SPEED)
  • RUSTFS_SCANNER_DELAY (canonical)
  • RUSTFS_SCANNER_MAX_WAIT_SECS (canonical)
  • RUSTFS_SCANNER_CYCLE (canonical, also accepts MINIO_SCANNER_CYCLE)
  • RUSTFS_SCANNER_START_DELAY_SECS (canonical)
  • RUSTFS_DATA_SCANNER_START_DELAY_SECS (deprecated alias for compatibility)
  • RUSTFS_SCANNER_IDLE_MODE (canonical)
  • RUSTFS_SCANNER_CACHE_SAVE_TIMEOUT_SECS (canonical)
  • RUSTFS_SCANNER_CYCLE_MAX_DURATION_SECS (canonical)
  • RUSTFS_SCANNER_CYCLE_MAX_OBJECTS (canonical)
  • RUSTFS_SCANNER_CYCLE_MAX_DIRECTORIES (canonical)

Mmap read environment aliases

  • RUSTFS_OBJECT_MMAP_READ_ENABLE (canonical)
  • RUSTFS_OBJECT_ZERO_COPY_ENABLE (deprecated alias for compatibility)

Health compatibility switches

  • RUSTFS_HEALTH_ENDPOINT_ENABLE
    • controls canonical /health, /health/live, and /health/ready endpoint exposure.
  • RUSTFS_HEALTH_MINIMAL_RESPONSE_ENABLE
    • enables minimal payload mode for GET health responses (status, ready only).
  • RUSTFS_HEALTH_READINESS_CACHE_TTL_MS
    • TTL for readiness cache evaluation.
  • RUSTFS_HEALTH_COMPAT_BUSY_CHECK_ENABLE
    • enables busy protection behavior for health probes.
    • default is false.
  • RUSTFS_HEALTH_COMPAT_BUSY_MAX_ACTIVE_REQUESTS
    • max active HTTP requests; health probes return 429 when active requests reach or exceed this value.
    • 0 disables thresholding even if busy protection is enabled.
  • RUSTFS_HEALTH_COMPAT_KMS_READY_CHECK_ENABLE
    • enables KMS readiness enforcement for /health/ready.
    • default is false.

Drive timeout environment variables

  • RUSTFS_DRIVE_METADATA_TIMEOUT_SECS
  • RUSTFS_DRIVE_DISK_INFO_TIMEOUT_SECS
  • RUSTFS_DRIVE_LIST_DIR_TIMEOUT_SECS
  • RUSTFS_DRIVE_WALKDIR_TIMEOUT_SECS
  • RUSTFS_DRIVE_WALKDIR_STALL_TIMEOUT_SECS

Legacy compatibility fallback:

  • RUSTFS_DRIVE_MAX_TIMEOUT_DURATION This legacy variable is treated as a deprecated fallback for the operation-specific drive timeout variables above when a canonical variable is unset.

Drive timeout health-action policy:

  • RUSTFS_DRIVE_TIMEOUT_HEALTH_ACTION
    • mark_failure (default): timeout marks failure and may transition drive runtime state.
    • ignore_scanner: timeout does not mark failure for scanner-sensitive operations (walk_dir, read_metadata, list_dir, disk_info).

Drive timeout profile preset:

  • RUSTFS_DRIVE_TIMEOUT_PROFILE
    • default (default): keep current timeout defaults.
    • high_latency: use 60s default timeout for scanner-sensitive operations when no per-operation timeout override is set (read_metadata, disk_info, list_dir, walk_dir, walk_dir_stall).
  • Precedence:
    • Explicit per-operation timeout env (RUSTFS_DRIVE_*_TIMEOUT_SECS) takes highest precedence.
    • Then RUSTFS_DRIVE_MAX_TIMEOUT_DURATION legacy fallback.
    • Then the profile-derived default (default or high_latency).

Startup filesystem boundary policy

  • RUSTFS_UNSUPPORTED_FS_POLICY controls startup behavior when RustFS detects local endpoint filesystems that are outside the supported production boundary.
    • warn (default): log warning and continue startup.
    • fail: abort startup with an error.

RustFS production guidance remains direct-attached local POSIX filesystems. Network-mounted filesystems (for example nfs, cifs, smb2, and fuse.*) are treated as unsupported by this startup guard.

📄 License

This project is licensed under the Apache License 2.0 - see the LICENSE file for details.