mirror of
https://github.com/rustfs/rustfs.git
synced 2026-08-06 13:27:43 +00:00
264b2dd48076fa19af59aa2980e31a020b28c748
181 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
264b2dd480 |
perf(metrics): drop needless per-emission work on hot metric paths (#4743)
* perf(metrics): drop needless per-emission work on hot metric paths Audit of the metrics hot paths surfaced four low-risk wins where emission did work it did not need to: - `record_file_cache_reclaim_success/error` (disk/local.rs) called `.to_string()` on `kind` (already `&'static str`) and on the `"ok"`/`"err"` literals, heap- allocating up to four `String`s per page-cache reclaim window — which runs per read-stream reclaim. The `metrics` macros accept `&'static str` label values directly, so pass them as-is. - `record_read_repair_dedup` (set_disk/core/io_primitives.rs) likewise `.to_string()`-ed an already-`&'static str` `reason`. - `SetDisks::get_object_reader` (set_disk/ops/object.rs) captured `Instant::now()` and emitted the `rustfs.lock.acquire.*` counter and histogram unconditionally on every GET, right beside an already-gated stage timer. Gate them behind `get_stage_metrics_enabled()` too, so an inactive observability config pays no per-GET clock read or recorder lookups. - The per-response-body-chunk counter in server/http.rs re-ran the `counter!` registry lookup on every chunk (a streamed GET emits many). Resolve the label-less handle once into a `LazyLock<metrics::Counter>`; the global recorder is installed at startup before any response streams, so the cached handle binds to the final recorder. No metric names or label values change. The only behavior change is that the `rustfs.lock.acquire.*` GET-path metrics now follow the GET stage-metrics flag, consistent with the neighbouring stage timings. Co-Authored-By: heihutu <heihutu@gmail.com> * perf(metrics): gate page-cache reclaim metrics behind metrics_enabled() `record_file_cache_reclaim_success/error` run per read-stream reclaim window on large-object reads and emitted unconditionally. When general metrics are disabled the `counter!`/`histogram!` macros still construct three metric keys per call for nothing. Skip the emission behind `rustfs_io_metrics::metrics_enabled()`, matching how the io-metrics free functions self-gate. The serial reclaim-metrics test now enables the flag (save/restore) alongside the existing stage gate. Left ungated deliberately: `record_read_repair_dedup` (rare read-repair path, and its non-serial test would need a global-flag toggle), and the HTTP body-chunk counter (its cached handle already makes the disabled case a no-op increment). Co-Authored-By: heihutu <heihutu@gmail.com> --------- Co-authored-by: heihutu <heihutu@gmail.com> |
||
|
|
1ab66e124c |
test(ecstore): pin the remaining io_uring fd-cache invariants (#1180) (#4739)
Close the three test-completeness items in rustfs/backlog#1180 that the earlier hardening (rustfs/rustfs#4726, #4729) left unpinned; the other two of the five (sharded cancel routing, bailout-handle error) already landed in rustfs/uring. - rename_data end-to-end invalidation: drive the real `LocalDisk::rename_data` commit (non-inline part, so `invalidate_part_paths` is non-empty) and assert the destination part descriptor is dropped — not merely `rename_file`/`delete`. Both production `rename_data` call sites share this invalidation, so a "fix one copy, miss the other" regression is now caught. The test is non-vacuous: it seeds the cache, removes the on-disk data dir out of band (the cached fd keeps the old inode alive and clears the path for the directory rename `rename_data` performs), and asserts a read still returns the OLD bytes before the commit — which fails outright if the cache is off, so it cannot pass without a live cache. - FD_CACHE_TTL backstop: an injected short TTL proves the cache self-evicts a descriptor with no explicit invalidation; a static check pins the 5s value. - zero-length read bounds parity on the cache-HIT path: a `length == 0` read past EOF must be rejected identically to the miss path and StdBackend, pinning the #1173 fix against regression. Refactors `FdCache::new` to delegate to a private `with_ttl(ttl)` helper so the TTL backstop can be exercised with a short TTL instead of a multi-second wait. Verified: `cargo test -p rustfs-ecstore` on Linux with real io_uring (seccomp=unconfined, RLIMIT_NOFILE raised, RUSTFS_URING_TESTS_MUST_RUN=1 so a degraded skip fails rather than passing vacuously). Co-authored-by: heihutu <heihutu@gmail.com> |
||
|
|
793b2a06e2 | perf(ecstore): snapshot per-IO env config on local disk read/write paths (#4736) | ||
|
|
9d64d71bd7 | fix(ecstore): reduce GET reader setup shard fanout (#4735) | ||
|
|
c79366d42c |
refactor(ecstore): tidy io_uring read backend (docs + fd-cache handle) (#4732)
Two behavior-preserving cleanups to the io_uring local read backend that accumulated as later work layered onto it: - Reunite the `UringBackend` struct doc comment. The backlog#1145 fd-cache constants were inserted into the middle of the struct's doc block, so its opening sentences were orphaned onto `ENV_RUSTFS_IO_URING_FD_CACHE` and the struct itself was documented by a sentence fragment starting mid-clause with "through rustfs-uring's...". Move the opening back onto the struct and give the constant its own one-line doc. - Fold the fd-cache handle and its lookup key into a single `Option<(&FdCache, FdKey)>` in `pread_uring`, so the get and the insert_if_fresh sites stop re-deriving `self.fd_cache.as_ref()` and re-matching presence. Semantics are identical, including the backlog#1176 generation guard (`gen_at_open` snapshot before open, insert only when the generation is unchanged). No functional change. The borrow/move pattern was validated against a host-compilable reduction (the cfg(linux) path cannot be built from macOS); CI compiles the Linux backend. Co-authored-by: heihutu <heihutu@gmail.com> |
||
|
|
ffca98cdbf |
fix(ecstore): harden io_uring integration (#4726)
* fix(ecstore): close the fd-cache open-then-insert race with a generation guard (rustfs/backlog#1176) pread_uring's miss path opened a descriptor on the blocking pool and only then inserted it into the moka cache. moka's invalidations cover only entries present at call time, so a heal/delete commit that invalidated between the open and the insert could not stop the just-opened stale inode from being cached afterwards — serving the pre-heal/pre-delete inode for up to the TTL and defeating the heal. Add an invalidation generation to FdCache, bumped by invalidate_exact and invalidate_under before they touch moka. The read path snapshots the generation before opening and inserts via insert_if_fresh, which refuses the insert if the generation moved during the open and, with a post-insert re-check, removes the entry if an invalidation raced the insert itself. Reads that never miss are unaffected. Co-Authored-By: heihutu <heihutu@gmail.com> * fix(ecstore): invalidate the fd cache on the primary object-delete paths (rustfs/backlog#1175) The fd-cache invalidation contract was only wired into DiskAPI::delete, rename_file and rename_data, but object deletion almost never goes through LocalDisk::delete — DeleteObject(s) reach delete_version, delete_versions -> delete_versions_internal, and delete_paths, all of which remove a version's data dir (move_to_trash / rename_all staging) with no invalidation. A cached io_uring descriptor kept the deleted part.N inode readable for up to the TTL, so a GET in that window could still return deleted data. Invalidate every cached fd under the removed data dir at each site: in delete_version and delete_versions_internal the data_dir uuid and object path are in hand (invalidate_cached_fds_under(volume, "{path}/{uuid}")); delete_paths invalidates under each removed path. A later rollback that restores a data dir just causes the next read to re-open it. Co-Authored-By: heihutu <heihutu@gmail.com> * fix(ecstore): close remaining fd-cache invalidation gaps (rustfs/backlog#1177) Three residual paths could keep serving a stale descriptor: - delete_volume removed the whole bucket tree (remove_dir_all/remove_dir) with no invalidation, and the cache-hit read path skips the volume-access check, so a cached fd kept a removed object readable. Add invalidate_cached_fds_for_volume (a per-volume moka predicate) and call it after the bucket is removed. - A retired LocalDisk instance (renew_disk on reconnect builds a fresh one) kept its populated cache alive while still referenced by in-flight ops, so invalidations through the new instance never reached it. close() now clears the backend's cache via clear_cached_fds. - rename_data's post-commit rollback (a commit-metadata fsync failure under strict durability) restored the old data dir without dropping fds cached during the committed window; the streaming branch now invalidates the dst part fds on those rollback paths. The inline branch's rollback runs inside spawn_blocking and is left to the TTL backstop. Co-Authored-By: heihutu <heihutu@gmail.com> * fix(ecstore): narrow the io_uring latch classes to match StdBackend (rustfs/backlog#1171) The runtime degradation classification reused the probe-time restriction errnos, which the driver's C7 contract explicitly warns against, so a single per-file error could latch a whole disk off io_uring: - is_io_uring_unsupported no longer includes EACCES: at read time on an already-open fd it is per-file (an LSM hooks security_file_permission on every read) and StdBackend hits the same denial, so falling back masks nothing and a full-disk latch would be wrong. ENOSYS and EPERM (seccomp/LSM applied after startup) remain. EOPNOTSUPP is now classified per-path by the caller. - pread_uring_direct's read-error arm now mirrors StdBackend: an O_DIRECT-shape error (EINVAL/EOPNOTSUPP) latches only direct_uring.supported so eligible reads take StdBackend's aligned path, instead of over-latching the whole io_uring backend or never latching a read-side EINVAL at all. - try_new only negative-caches genuine restriction-class probe failures in URING_UNSUPPORTED_DISKS; an unexpected (possibly transient) probe failure now falls back without latching, so the next reconnect re-probes. Co-Authored-By: heihutu <heihutu@gmail.com> * feat(ecstore): log when a disk latches io_uring off at runtime (rustfs/backlog#1172) A probe-gated gray release was flying blind: the permanent per-disk `active` latch flipped with no log and no metric, so the only message operators ever saw was the startup "io_uring read backend enabled" line — which stayed true on dashboards even after the very first read latched the disk back to StdBackend forever. Add latch_active_off, which flips the latch with `swap` and logs the true->false transition exactly once at warn with a dedicated event constant, disk root, and errno. Both the buffered and O_DIRECT read paths use it. A fallback/latch metric counter and periodic export of the driver StatsSnapshot (cq_overflow, cancel_already) remain as follow-ups that need rustfs_io_metrics plumbing. Co-Authored-By: heihutu <heihutu@gmail.com> * chore(audit): correct the stale rustfs-uring license-allow rationale (rustfs/backlog#1181) The dependency-review allow said rustfs-uring is "pulled as a git dependency", but ecstore now pins it from crates.io. Update the rationale and scope the allow to the exact pinned version (pkg:cargo/rustfs-uring@0.1.0) so a future version bump forces a conscious re-review of the license/provenance claim instead of being waved through on an outdated justification. Co-Authored-By: heihutu <heihutu@gmail.com> * fix(ecstore): offload io_uring driver teardown off the tokio worker (rustfs/backlog#1170) UringBackend held Arc<UringDriver> and had no Drop, so when the last LocalDisk reference dropped in async context (disk reconnect via renew_disk, or shutdown), UringDriver's own Drop ran on that thread — sending Shutdown and joining each shard thread, which can block up to the bounded-drain timeout (5s) on a hung / D-state disk, stalling a tokio worker. Wrap the driver in ManuallyDrop (deref is transparent, so read call sites are unchanged) and add a Drop that takes the Arc and, when a runtime is present, drops it on a blocking thread so the potentially-blocking join never runs on a runtime worker. Off-runtime it drops inline. Co-Authored-By: heihutu <heihutu@gmail.com> * fix(ecstore): restore StdBackend read parity on the uring paths (rustfs/backlog#1173) Two byte-for-byte parity breaks against StdBackend on the io_uring read paths: - A zero-length read on an fd-cache hit returned Ok(empty) without any bounds check, while StdBackend and the uring miss path return FileCorrupt for an offset past EOF. Fstat the cached descriptor on the length==0 path and match. - reclaim_read_range fadvise(DONTNEED)'d the raw unaligned [offset, offset+len) range, but fadvise only drops fully-covered pages, so the head partial page stayed resident — whereas StdBackend's mmap path reclaims the page-aligned superset. Bitrot shards' 32-byte block headers keep offsets off page boundaries, so this diverged on the common case. Page-align the reclaim window to match the mmap path exactly. (The third parity item from the audit — a failed reclaim fadvise failing the read — is already parity: StdBackend's mmap path propagates the same fadvise error with `?`, so no change is needed.) Co-Authored-By: heihutu <heihutu@gmail.com> * fix(ecstore): bound worst-case in-flight memory by chunking huge uring reads (rustfs/backlog#1174) The driver's backpressure permits count operations, not bytes, and it zero-fills a full-size buffer per op, so a single unbounded read could pin ~length bytes per permit (128 permits x shards x up to ~2 GiB). ecstore passes a whole part's shard range as one pread_bytes with no upstream chunking. On the buffered path, split reads larger than URING_MAX_OP_LEN (128 MiB) into sequential chunks, awaited one at a time, so worst-case in-flight memory is bounded by permits x URING_MAX_OP_LEN per shard. The threshold is high enough that ordinary shard reads keep the single-op, zero-copy fast path unchanged. The O_DIRECT path (opt-in, alignment-constrained) is left for a follow-up. Co-Authored-By: heihutu <heihutu@gmail.com> * fix(ecstore): gate the io_uring fd cache on RLIMIT_NOFILE headroom (rustfs/backlog#1178) The fd cache holds up to FD_CACHE_CAPACITY (512) descriptors per disk, but try_new cannot know the disk count and nothing checked the process fd budget. On a bare-metal / non-systemd run with the common 1024 soft RLIMIT_NOFILE, two disks would already exhaust fds with EMFILE surfacing on reads and probes. Check the soft limit at try_new: enable the cache only with ample headroom (>= 16384), otherwise log a warning once and fall back to open-per-read. The packaged systemd unit sets 1,048,576, so tuned deployments are unaffected. Co-Authored-By: heihutu <heihutu@gmail.com> * test(ecstore): make io_uring test skips visible and gate non-vacuity (rustfs/backlog#1179) The ecstore io_uring tests degrade to a silent pass when io_uring is unavailable (bare `return`s or plain eprintlns), so a CI leg on a restricted runner never exercises the real UringBackend/FdCache/latch paths yet still goes green — an integration regression could merge unseen. Add uring_test_skip: it emits a grep-able `SKIP <name>` line and, when RUSTFS_URING_TESTS_MUST_RUN is set (a CI leg that guarantees io_uring, e.g. a seccomp=unconfined container), panics instead of skipping. Route the silent-skip sites through it. Wiring a dedicated CI leg that sets that env on a capable runner is tracked in the issue; this provides the enforcement mechanism. Co-Authored-By: heihutu <heihutu@gmail.com> * test(ecstore): cover delete_paths fd-cache invalidation (rustfs/backlog#1180) Add an end-to-end test that seeds the descriptor cache with a read, removes the part via disk.delete_paths (one of the primary object-delete entry points that does not go through LocalDisk::delete), and asserts the next read no longer returns the removed inode — pinning the invalidation added in #1175. The sharded cancel-routing half of #1180 is covered in the rustfs-uring PR. Co-Authored-By: heihutu <heihutu@gmail.com> * io_uring audit follow-ups: O_DIRECT chunking, inline invalidation, metrics, CI leg (backlog#1160) (#4729) * fix(ecstore): chunk large O_DIRECT reads too, bounding in-flight memory (rustfs/backlog#1174) The buffered read path already splits reads above URING_MAX_OP_LEN into sequential chunks; do the same for the O_DIRECT path, which was left for a follow-up. read_at_direct aligns each chunk's sub-range internally, and chunk sizes are a multiple of URING_MAX_OP_LEN so a boundary re-read is at most one block. Extract classify_direct_read_error so the single-op and chunked paths share one copy of the EINVAL/EOPNOTSUPP-vs-subsystem latch classification rather than duplicating it. Co-Authored-By: heihutu <heihutu@gmail.com> * fix(ecstore): invalidate cached fds on the inline rename_data rollback (rustfs/backlog#1177) The streaming rename_data branch invalidates cached part fds on its post-commit rollback paths, but the inline branch runs its commit and rollback inside a single spawn_blocking closure where the async invalidate cannot be called, so it was left to the TTL backstop. Capture the closure's result instead of `??`-propagating it: on error (a commit-metadata fsync failure under strict durability rolls the committed rename back), invalidate the dst part paths at the async level before returning. Inline objects keep their data in xl.meta rather than separate part inodes, so this is largely defensive, but it removes the caveat and keeps the two branches consistent. Co-Authored-By: heihutu <heihutu@gmail.com> * feat(ecstore): export io_uring latch/fallback and driver stats metrics (rustfs/backlog#1172) Complete the gray-release observability. Beyond the warn log added earlier, emit metrics so a dashboard can answer "how much traffic is on io_uring vs falling back, and is any disk degrading": - rustfs_io_uring_latch_off_total — a disk latching io_uring off at runtime. - rustfs_io_uring_read_fallback_total — each io_uring -> StdBackend read fallback (latched-off short-circuit, O_DIRECT error, buffered error). - a low-frequency per-disk exporter of the driver StatsSnapshot as gauges (in_flight, cq_overflow, cancel_already), spawned in try_new. It holds only a Weak reference so it never keeps the driver alive, and drops any temporary strong reference on the blocking pool so a last-reference UringDriver::Drop join never runs on an async worker (rustfs/backlog#1170). submit_errors is deliberately not exported yet: it is a field added in the unreleased rustfs-uring 0.2.0, and ecstore still pins 0.1.0. It lands once the dependency is bumped (rustfs/backlog#1181). Co-Authored-By: heihutu <heihutu@gmail.com> * ci: add a real-io_uring integration leg on ubuntu-latest (rustfs/backlog#1179) The existing self-hosted sm-standard runners cannot guarantee io_uring is available (a container seccomp filter can block io_uring_setup), so the ecstore uring tests degrade to a silent skip and never exercise the real UringBackend/FdCache/latch paths in CI. Add a job on GitHub-hosted ubuntu-latest, which runs a recent kernel with no container seccomp filter, running the uring-named ecstore tests with RUSTFS_IO_URING_READ_ENABLE=true and RUSTFS_URING_TESTS_MUST_RUN=1 — the non-vacuity gate makes the leg fail rather than skip if io_uring is unavailable, so an integration regression can no longer merge green behind a vacuous pass. Co-Authored-By: heihutu <heihutu@gmail.com> --------- Co-authored-by: heihutu <heihutu@gmail.com> --------- Co-authored-by: heihutu <heihutu@gmail.com> |
||
|
|
3f25426534 |
fix(ecstore): reject incomplete listing usage refreshes (#4698)
Co-authored-by: Henry Guo <marshawcoco@users.noreply.github.com> |
||
|
|
d0ca14d8df |
test(ecstore): cover post-commit hard-crash in rename_data crash harness (#4689)
The rename_data crash-consistency harness (backlog#935/#878) only armed pre-commit crash points (AfterDataRename, AfterBackupBeforeMetaCommit), which assert the *old* version survives. The other half of the "partial commit -> old or new, never mixed" invariant — a hard power loss *after* the xl.meta commit rename must leave the *new* version — had no crash-point coverage. The existing after-metadata-commit failpoint only exercises the graceful in-process rollback (restores old), a different path from a hard crash (no rollback runs, commit stays on disk). Add a RenameDataCrashPoint::AfterMetaCommit injection right after the commit rename (cfg(test), compiles to a const-false no-op in production, like the existing points) and overwrite/fresh x strict/relaxed tests asserting the object reads back as the new version. The fresh case pins that the point is genuinely post-commit: a pre-commit misplacement would leave no readable object and fail the assertion. cargo test -p rustfs-ecstore crash_consistency: 9 passed. clippy -D warnings clean; fmt and arch guards clean. |
||
|
|
6f05a740b3 |
refactor(ecstore): extract the shared io_uring read preamble (backlog#1145) (#4696)
`pread_uring` and `pread_uring_direct` accumulated across seven rounds (#4632/#4635/#4645/#4649/#4653/#4658/#4662) and each carried its own copy of the same open preamble: resolve the bucket path, run the volume access check, resolve the object path, and check the path length. The only real difference is what happens after — one opens buffered and errors as `DiskError`, the other opens `O_DIRECT` and errors as `DirectOpenError` so it can latch an O_DIRECT refusal. Extract that shared sequence into `resolve_uring_object_path`, a blocking helper called inside both `spawn_blocking` closures. Each caller keeps exactly what differs — its open flags, its error type, its metadata/bounds/align handling — and no longer restates the resolution. No behavior change: the same checks run in the same order and produce the same errors (the O_DIRECT path still maps them through `DirectOpenError::Disk`). Net -12 lines. Verified on a real Linux host (16-core, real io_uring): clippy -p rustfs-ecstore --all-targets -D warnings clean; disk::local tests 144 passed, 0 failed — including the io_uring, O_DIRECT, fd-cache, and page-cache-reclaim cases that exercise both read paths. Co-authored-by: heihutu <heihutu@gmail.com> |
||
|
|
80ddd8fa7e |
fix(ecstore): roll back delete on disks that staged then errored (#4676)
fix(ecstore): roll back delete on disks that staged then errored (backlog#1158) #4300 rolls back a failed delete only on disks that returned Ok, skipping any disk that staged its rollback backup, applied the delete, and then errored -- leaving that disk deleted while its peers are restored. On rollback, fan the undo out to every online disk instead; the disk-side restore_delete_rollback is already idempotent (Ok no-op when nothing was staged), so unstaged disks are unaffected. The err.is_some() skip now applies only to the success/cleanup path. Covers both single-object and batch delete. Refs backlog#1158. |
||
|
|
f83f9ada13 |
fix(ecstore): reclaim page cache after io_uring reads (#4662)
fix(ecstore): io_uring reads must reclaim the page cache like StdBackend (backlog#1145) `StdBackend::pread_bytes` calls `fadvise(DONTNEED)` over the range it just read whenever `should_reclaim_file_cache_after_read(length)` holds — `RUSTFS_OBJECT_FILE_CACHE_RECLAIM_READ_ENABLE` (on by default) above a 4 MiB threshold. Large object reads are usually cold, and leaving them resident evicts everything else, so this is a deliberate policy rather than a side effect of how StdBackend happens to read. `pread_uring` and `pread_uring_direct` never did it. Enabling io_uring on a disk therefore turned the policy off silently: shard reads at or above the threshold stayed resident and grew the page cache without bound. This is the same class of mistake as the O_DIRECT loss fixed earlier — copying the read itself while dropping the policy attached to it. It also badly distorts any benchmark. An end-to-end warp GET A/B on a 16-core host (4 MiB objects, conc 64) showed io_uring at 8782 MiB/s against StdBackend's 116 MiB/s. That is not a 76x speedup: with io_uring the run issued *zero* device reads and grew the page cache, while StdBackend read 2950 MiB from the device and shrank it. The two legs were not running the same policy. (Disabling the mmap read path changed nothing — 1.01x — so the mmap-vs-pread difference was not the cause either.) Both io_uring read paths now reclaim the range they read, with the same gate, the same range, and the same error handling as StdBackend. O_DIRECT should leave nothing resident anyway, but a filesystem that quietly buffered the read still has to honour the policy, so `pread_uring_direct` reclaims too. The test measures the policy, not the call: it reads an 8 MiB file through each backend and asks `mincore(2)` how much stayed resident. Crucially it also asserts the reclaim-disabled case leaves the range resident — without that gate, a backend that never reclaimed would pass silently. Verified: with the fix removed, the test fails with "uring: reclaim is on ... 2048/2048 pages remain"; with the fix it passes. Verified on a real Linux host (16-core, real io_uring): clippy --tests -D warnings clean; disk::local tests 143 passed, 0 failed. Co-authored-by: heihutu <heihutu@gmail.com> |
||
|
|
99eef032c6 |
fix(ecstore): restore Windows async read trait import (#4652)
* fix(ecstore): restore Windows async read trait import * fix(ecstore): gate async read import to non-Unix |
||
|
|
15b8b13698 |
feat(ecstore): cache part-file descriptors for io_uring reads (#4658)
* feat(ecstore): run one io_uring ring per shard on each disk (backlog#1145)
A buffered read that hits the page cache completes inline inside
`io_uring_enter`, so the thread driving a ring performs that read's
memcpy. One ring per disk therefore capped cache-hit reads at a single
core's memory bandwidth: measured on a 16-core host, one driver thread sat
pinned at 100% CPU while throughput stayed flat at ~5 GB/s regardless of
read size, against 50 GB/s for the blocking-pool baseline.
rustfs/uring#6 taught the driver to hold N independent rings, each with its
own thread, pending table, backpressure semaphore, and eventfd. Wire it up:
`UringBackend::try_new` now calls `probe_and_start_sharded`, and
`RUSTFS_IO_URING_SHARDS` selects the count per disk.
The default is a quarter of the available parallelism clamped to `1..=4`,
because the cost is `disks × shards` driver threads (each normally blocked
in `poll(2)`). Any override is clamped to `1..=16`, so a mistyped value can
neither disable the driver (0) nor spawn threads without bound; an
unparseable value falls back to the default.
Effect (warm page cache, 16-core, rustfs/uring's concurrent_pread_bench):
1 MiB, conc 8: 1 shard 4911 MB/s -> 8 shards 47361 MB/s (9.6x);
the blocking-pool baseline is 50662 MB/s
64 KiB, conc 32: StdBackend 153678 IOPS, p999 3030 us
8 shards 345402 IOPS, p999 897 us
64 KiB, conc 128: StdBackend 135155 IOPS, p999 10716 us
8 shards 389047 IOPS, p999 4092 us
Sharding removes the throughput deficit *and* keeps io_uring's tail-latency
advantage, rather than trading one for the other.
Unchanged: io_uring read stays gray-off by default
(`RUSTFS_IO_URING_READ_ENABLE`), reads are byte-for-byte identical to
StdBackend, the per-disk degradation latches and probe cache (backlog#1101)
and the O_DIRECT tiered fallback (backlog#1102) all still apply. Rings stay
per-disk, so a stalled disk cannot starve another disk's rings
(backlog#1055). Bumps the rustfs-uring pin to the merged #6 commit.
Verified on a real Linux host (16-core, real io_uring): cargo clippy
--tests -D warnings clean; disk::local tests 132 passed, 0 failed —
including the existing io_uring and O_DIRECT cases now running on the
sharded driver, plus a new test covering the shard-count default, override,
and clamping.
Co-Authored-By: heihutu <heihutu@gmail.com>
* feat(ecstore): cache part-file descriptors for io_uring reads (backlog#1145)
`pread_uring` opened the file on the blocking pool for every read, so each
read paid a `spawn_blocking` round trip — the very thread hop io_uring
exists to avoid. Sharding the driver (backlog#1145) removed the previous
ceiling and left this as the binding cost. Measured on a 16-core host with
a 4-shard driver, warm page cache:
64 KiB, conc 8: 143942 -> 263054 IOPS (+83%), p999 240 -> 65 us
64 KiB, conc 32: 150128 -> 204876 IOPS (+36%), p999 2508 -> 871 us
64 KiB, conc 128: 129172 -> 361287 IOPS (+180%), p999 15329 -> 3046 us
1 MiB, conc 32: 33875 -> 42301 IOPS (+25%)
At 64 KiB / conc 128 the open is what masked io_uring entirely: with it,
io_uring beat StdBackend by 3.5%; without it, by 189%.
Add a bounded per-disk descriptor cache used by the io_uring read path.
A hit takes no `open` and no `spawn_blocking`, so the read never leaves the
runtime worker.
Why caching a part-file descriptor is safe:
* only `<object>/<data_dir>/part.N` reaches this backend's `pread_bytes`;
`xl.meta` — the one path replaced in place — is read through `read_all`
/ `read_metadata` and never gets here;
* part files are never rewritten in place. A replacement is always
write-new-tmp then `rename`, which swaps the inode, so a cached
descriptor can never observe a torn shard.
Why invalidation is nevertheless REQUIRED: heal reuses the existing
version's `data_dir` and renames a rebuilt shard onto the SAME part path.
A cached descriptor would keep serving the pre-heal (corrupt) inode,
defeating the heal and eroding read quorum. `delete` likewise unlinks a
part that a cached descriptor would keep readable. So `rename_data`,
`rename_file`, and `delete` all call the new
`LocalIoBackend::invalidate_cached_fds` after they mutate, and a 5s TTL
bounds the blast radius should a future mutation path forget to.
Two preamble checks the miss path runs are not silently lost on a hit:
* bounds — the driver only short-reads at EOF (it resubmits otherwise),
so `bytes.len() != length` is exactly the old `meta.len() < end_offset`
check, and now yields the same `FileCorrupt`;
* volume access — skipped while an entry is live. An unreachable disk
keeps serving already-open descriptors for at most the TTL, after which
the re-open re-runs the check. Disk health is tracked independently of
this per-read probe.
Scope: buffered io_uring reads only. The O_DIRECT path keeps opening per
read (its reads are >= 4 MiB, so the open is a small fraction), and
StdBackend is untouched — it must take the blocking hop for the pread
regardless, so caching would buy it only 2-6% while carrying the same
invalidation risk. `RUSTFS_IO_URING_FD_CACHE=false` restores open-per-read.
Verified on a real Linux host (16-core, real io_uring): clippy --tests
-D warnings clean; disk::local tests 135 passed, 0 failed. The new
heal-staleness test first asserts a read still returns the PRE-heal bytes —
proving the cache is live and the hazard real — then that invalidation makes
the healed shard visible. A second test drives `rename_file` and `delete`
through `LocalDisk` to prove those paths actually invalidate, and a unit
test pins prefix invalidation to component boundaries (`a/b` must not drop
`a/bc`).
Co-Authored-By: heihutu <heihutu@gmail.com>
---------
Co-authored-by: heihutu <heihutu@gmail.com>
|
||
|
|
aaffd3ede8 | fix(ecstore): align Deep verifier shard geometry (#4656) | ||
|
|
a269f8df05 | fix(ecstore): require write quorum for metadata early stop (#4300) | ||
|
|
fe682e16e9 |
feat(ecstore): run one io_uring ring per shard on each disk (backlog#1145) (#4653)
A buffered read that hits the page cache completes inline inside
`io_uring_enter`, so the thread driving a ring performs that read's
memcpy. One ring per disk therefore capped cache-hit reads at a single
core's memory bandwidth: measured on a 16-core host, one driver thread sat
pinned at 100% CPU while throughput stayed flat at ~5 GB/s regardless of
read size, against 50 GB/s for the blocking-pool baseline.
rustfs/uring#6 taught the driver to hold N independent rings, each with its
own thread, pending table, backpressure semaphore, and eventfd. Wire it up:
`UringBackend::try_new` now calls `probe_and_start_sharded`, and
`RUSTFS_IO_URING_SHARDS` selects the count per disk.
The default is a quarter of the available parallelism clamped to `1..=4`,
because the cost is `disks × shards` driver threads (each normally blocked
in `poll(2)`). Any override is clamped to `1..=16`, so a mistyped value can
neither disable the driver (0) nor spawn threads without bound; an
unparseable value falls back to the default.
Effect (warm page cache, 16-core, rustfs/uring's concurrent_pread_bench):
1 MiB, conc 8: 1 shard 4911 MB/s -> 8 shards 47361 MB/s (9.6x);
the blocking-pool baseline is 50662 MB/s
64 KiB, conc 32: StdBackend 153678 IOPS, p999 3030 us
8 shards 345402 IOPS, p999 897 us
64 KiB, conc 128: StdBackend 135155 IOPS, p999 10716 us
8 shards 389047 IOPS, p999 4092 us
Sharding removes the throughput deficit *and* keeps io_uring's tail-latency
advantage, rather than trading one for the other.
Unchanged: io_uring read stays gray-off by default
(`RUSTFS_IO_URING_READ_ENABLE`), reads are byte-for-byte identical to
StdBackend, the per-disk degradation latches and probe cache (backlog#1101)
and the O_DIRECT tiered fallback (backlog#1102) all still apply. Rings stay
per-disk, so a stalled disk cannot starve another disk's rings
(backlog#1055). Bumps the rustfs-uring pin to the merged #6 commit.
Verified on a real Linux host (16-core, real io_uring): cargo clippy
--tests -D warnings clean; disk::local tests 132 passed, 0 failed —
including the existing io_uring and O_DIRECT cases now running on the
sharded driver, plus a new test covering the shard-count default, override,
and clamping.
Co-authored-by: heihutu <heihutu@gmail.com>
|
||
|
|
baf7e899b9 |
fix(ecstore): bound every walk filesystem call by the stall budget (#4650)
The listing path no longer has a wrapper-level total timeout, so any filesystem call the walk makes without a stall bound would have no liveness bound at all. The first pass covered list_dir and read_metadata but left four reads unbounded: the volume access() probe and fs::metadata() in walk_dir, the xl.meta read in its dir-object branch, and is_empty_dir() in scan_dir. Split the helper so reads that do not yield a Result go through the same budget, and route the four remaining calls through it. A hung drive now fails those reads with DiskError::Timeout instead of parking the walk until the merge consumer's own stall timeout aborts it. Co-authored-by: heihutu <heihutu@gmail.com> |
||
|
|
8b233cdd94 |
fix(ecstore): keep O_DIRECT for eligible reads via native io_uring (backlog#1102) (#4649)
fix(ecstore): keep O_DIRECT for eligible reads via native io_uring (rustfs/backlog#1102)
Wire ecstore's io_uring read backend to rustfs-uring's native aligned
O_DIRECT read (`read_at_direct`, merged in rustfs/uring#3), so an
O_DIRECT-eligible read keeps BOTH io_uring's async submission AND
O_DIRECT's page-cache bypass instead of trading one for the other.
`UringBackend::pread_bytes` now selects the read shape by a tiered,
per-disk-latched ladder — never a blanket downgrade:
1. io_uring latched off for this disk (backlog#1101) -> StdBackend.
2. O_DIRECT-eligible + native path usable -> pread_uring_direct: open
the file O_DIRECT, probe the device alignment once (cached), and
read via read_at_direct. Best path: async + no cache pollution.
3. O_DIRECT-eligible but the fs refused O_DIRECT earlier (latched)
-> StdBackend aligned path (which itself degrades to buffered).
4. non-O_DIRECT read -> buffered pread_uring.
Latching keeps a failure from being re-attempted per-read: an fs that
refuses O_DIRECT (EINVAL/EOPNOTSUPP on open) latches the native direct
path off for that disk; a restriction-class errno from the read latches
io_uring off wholesale (backlog#1101). Any other error falls back for
that one read without masking a genuine data problem as a permanent
downgrade. The alignment comes from the existing probe_direct_io_align
(statx STATX_DIOALIGN, default 4096) and is cached in a per-disk
OnceLock so the probe runs at most once.
This replaces the interim guard that routed every O_DIRECT-eligible read
to StdBackend (which disabled io_uring on the hottest shard reads); the
native path is now the default and StdBackend is only a fallback tier.
Bumps the rustfs-uring pin to the merged #3 commit.
Verified on a real Linux host (16-core, real io_uring + O_DIRECT):
cargo check --tests, clippy -D warnings, and disk::local tests all pass
(128 passed, 0 failed), including uring_preserves_o_direct_for_eligible_reads
across unaligned ranges.
Co-authored-by: heihutu <heihutu@gmail.com>
|
||
|
|
399461c33e |
fix(ecstore): bound listing walks by drive stall, not total duration (#4647)
Foreground S3 listings wrapped the entire walk_dir stream in a single wall-clock timeout (RUSTFS_DRIVE_WALKDIR_TIMEOUT_SECS, default 5s). That budget measured how much data the walk had to produce rather than whether the drive was still answering, so a healthy but large prefix on slow media returned 500 InternalError / "Io error: timeout" to the client. The timer also kept running while the walk was blocked writing to a slow consumer, charging merge-side backpressure to the producer. WalkDirOptions::stall_timeout_ms already carried the right semantics but was only honored by the remote-disk RPC walk; LocalDisk::walk_dir ignored it entirely. A single-drive deployment therefore had no way to distinguish a hung drive from a big directory. Teach LocalDisk::walk_dir to bound each individual drive read with the stall budget, defaulting it from RUSTFS_DRIVE_WALKDIR_STALL_TIMEOUT_SECS when the caller does not pin one, and let the foreground listing path skip the wrapper-level total timeout. A walk that keeps making progress now runs to completion; a drive that stops answering still fails with DiskError::Timeout. Time spent blocked on the consumer stays outside the budget. Heal and rebalance walks already skipped the total timeout and previously ran unbounded on local drives. Give them an explicit, generous 60s stall budget so this change does not tighten them from "no bound" to the 5s default. Fixes #4644 Refs #2999, #3001 Co-authored-by: heihutu <heihutu@gmail.com> |
||
|
|
0ad4a26056 |
fix(ecstore): keep O_DIRECT for eligible reads when io_uring is enabled (backlog#1102) (#4645)
fix(ecstore): keep O_DIRECT for eligible reads when io_uring is enabled (rustfs/backlog#1102) UringBackend::pread_uring opens the file buffered, so with io_uring enabled an O_DIRECT-eligible read (RUSTFS_OBJECT_DIRECT_IO_READ_ENABLE + length >= the threshold) silently lost O_DIRECT and polluted the page cache — a behavior change the moment io_uring was turned on for a disk that uses O_DIRECT. The io_uring backend is gray-off by default, so no deployment is affected, but the gap had to close before it can be enabled anywhere O_DIRECT is in use. Route O_DIRECT-eligible reads back through StdBackend's aligned path (which itself falls back to buffered when the filesystem rejects O_DIRECT). A native aligned O_DIRECT read in the driver remains part of rustfs/backlog#1102. Adds uring_preserves_o_direct_for_eligible_reads: with BOTH flags on and the threshold at 1, unaligned ranges still return exactly the requested bytes. Co-authored-by: heihutu <heihutu@gmail.com> |
||
|
|
e6e4aef45b |
test(ecstore): cover the io_uring per-disk probe cache hit (backlog#1101) (#4641)
test(ecstore): cover the io_uring per-disk probe cache hit (rustfs/backlog#1101) Follow-up to #4635. The per-disk negative probe cache was implemented but had no dedicated test. Add uring_probe_cache_skips_known_unsupported_disk: a root recorded in URING_UNSUPPORTED_DISKS makes try_new return None without a fresh probe. Completes the #1101 acceptance ("cache hit does not re-probe"). Co-authored-by: heihutu <heihutu@gmail.com> |
||
|
|
7dc03603da |
feat(ecstore): add io_uring degradation latch and per-disk probe cache (#4635)
feat(ecstore): io_uring runtime degradation latch + per-disk probe cache (rustfs/backlog#1101) Follow-up to #1104. Add the runtime half of the io_uring degradation contract: - Runtime latch: UringBackend gains an `active` flag (mirrors DirectIoReadState::supported). A read that returns a restriction-class errno (EPERM/EACCES/ENOSYS/EOPNOTSUPP) — io_uring became unusable on this disk — trips the latch, and every further read goes straight to StdBackend with no more per-read io_uring attempts. Data errors (EIO), missing files, and parameter errors do NOT latch, since StdBackend would hit them too. - Per-disk probe cache: a process-wide negative cache of disks whose probe failed, so try_new skips re-probing (creating a ring + driver thread) on reconnect. Tests: the errno classifier (restriction-only) and a latched-off read that still returns correct bytes via StdBackend; the #1104 differential test still passes under real io_uring. Co-authored-by: heihutu <heihutu@gmail.com> |
||
|
|
5e80980698 | test(ecstore): pin sorted scan_dir stream for dir-marker with children (backlog#1068) (#4626) | ||
|
|
da755eb66b |
feat(ecstore): runtime-probed io_uring read backend, gray-off by default (backlog#1104) (#4632)
feat(ecstore): add runtime-probed io_uring read backend, gray-off by default (rustfs/backlog#1104) Wire the standalone rustfs-uring crate (https://github.com/rustfs/uring) into LocalIoBackend as an opt-in read backend. When RUSTFS_IO_URING_READ_ENABLE is set AND the per-disk io_uring probe succeeds, positioned reads go through the cancel-safe UringDriver; otherwise — default, or on a restricted host, or on any per-read driver error — reads fall back to StdBackend byte-for-byte, so the default build is unchanged. - Cargo.toml: rustfs-uring as a Linux-only git dependency (pinned rev). The guard scripts/check_no_tokio_io_uring.sh allows an explicit io-uring integration; only the tokio "io-uring" runtime feature is banned. - UringBackend mirrors StdBackend::pread_bytes's resolution/access/bounds preamble and only swaps the raw byte read; the other three trait methods delegate to the inner StdBackend. - Backend selection at LocalDisk construction via build_local_io_backend. - Differential test uring_backend_reads_match_std: with the flag set, every positioned read returns byte-identical data (driver-served when io_uring is available, StdBackend fallback otherwise). Verified: cargo check -p rustfs-ecstore on Linux; the differential test passes under real io_uring (seccomp=unconfined). The runtime errno degradation latch, per-disk probe cache, and productionization are tracked in rustfs/backlog#1101/#1102. Co-authored-by: heihutu <heihutu@gmail.com> |
||
|
|
4f999bb6b8 | perf(ecstore): backfill rename_data old size and gate the PUT prelookup (#4598) | ||
|
|
579cdf52dc | fix(ecstore): harden object-prefix probe and parse markers before forward_past (#4600) | ||
|
|
e0eee89a3b |
fix(ecstore): emit CommonPrefix for object with same-named prefix in non-recursive listings (#4596)
fix(ecstore): emit CommonPrefix for object with same-named prefix in non-recursive listings (backlog#1042) On single-disk / consistent deployments a non-recursive scan_dir that finds a/xl.meta classifies `a` as an object and does not descend, so the source never emits the prefix `a/`, dropping the CommonPrefix from delimiter listings even when `a/b` exists. This is the single-disk follow-up to backlog#880 (the multi-disk merge path was fixed in #4563). Fix: in the scan_dir Ok branch, for a non-recursive listing of a plain object, probe whether `a/` holds a listing entry other than the object itself (new object_prefix_has_sibling_listing_entry, reusing directory_has_visible_listing_entry and skipping the object's own xl.meta and data dirs); if so, additionally schedule the prefix dir `a/`. The upstream fold in from_meta_cache_entries_sorted_infos already emits object `a` as Contents and dir `a/` as a CommonPrefix, so nothing changes there. Leaf objects (the common case) pay a single list_dir and return false immediately. Verification: new scan_dir unit test (positive: `a` + `a/b` coexist; negative: leaf `c` produces no spurious `c/`); new end-to-end e2e (object `a` + `a/b` with delimiter "/" -> Contents `a` + CommonPrefix `a/`); existing list_objects_duplicates e2e (#1797 / #2439 regressions) stays green; set_disk::read 114 passed. Note: lib-test compilation depends on #4573 (now merged), which added the allow_inplace_legacy_fallback argument to the codec streaming arity tests after #4560 / backlog#879. Co-authored-by: heihutu <heihutu@gmail.com> |
||
|
|
ce6bc30b26 | test(ecstore): harden validation gate and EC coverage tests (#4590) | ||
|
|
ed81d2f6b8 |
test(ecstore): complete EC validation coverage gate
* test(ecstore): complete EC validation coverage gate * test(ecstore): stabilize validation suite after rebase * test(ecstore): fix rio-v2 clippy lint |
||
|
|
608ab14d7d |
perf(ecstore): fold metadata read open+fstat+read into a single spawn_blocking (HP-12 item 1) (#4554)
Sub-change A only: the two local-disk metadata read paths (`read_metadata_with_dmtime`, `read_all_data_with_dmtime` in crates/ecstore/src/disk/local.rs) previously dispatched open, fstat and the xl.meta read as three separate async fs hops (three spawn_blocking round-trips under tokio::fs). Each is now folded into a single tokio::task::spawn_blocking closure using std::fs, cutting per-metadata-read dispatch from 3 to 1. This does NOT touch sub-change B (removing the entry-point access() call). Correctness first: this is the hottest metadata read path, so the refactor preserves byte-for-byte-equivalent Results. Every error mapping is kept identical: - open failure -> to_file_error - is_dir -> Error::FileNotFound (not to_file_error(EISDIR)) - metadata failure -> to_file_error - xl.meta parse -> propagated verbatim from the parser - try_reserve -> Error::other - read_to_end -> to_file_error For read_all_data_with_dmtime the async NotFound -> access(volume_dir) -> VolumeNotFound fallback (and its warn! event) is preserved on the async side: the closure returns the raw open error unmapped; only open() can yield ENOENT once the fd is valid, so gating the fallback on the open error is equivalent to the original open-arm-only fallback. To run the parser inside a blocking closure, filemeta gains a synchronous twin `read_xl_meta_no_data_sync` (+ `read_more_sync`) in crates/filemeta/src/filemeta/version.rs that line-for-line mirrors the async version, differing only in std vs tokio read_exact. A new equivalence test (`read_xl_meta_sync_equivalence_tests`) feeds identical buffers to both the async and sync readers and asserts equal Ok bytes / equal Err variants across: v1.0; v1.1/v1.2/v1.3; large meta triggering read_more; header truncation -> UnexpectedEof; CRC-trailer truncation -> FileCorrupt; unknown major/minor -> InvalidData; and want boundaries (exact fit, inline-data drop, read_more EOF). Co-authored-by: heihutu <heihutu@gmail.com> |
||
|
|
f7d2b25638 |
fix(ecstore): propagate disk delete/rename failures instead of swallowing them (#4546)
Two disk-layer correctness fixes in crates/ecstore/src/disk/local.rs. move_to_trash (#948, ECA-07) previously handled only DiskFull in its tail error block and let every other rename failure (EIO/EACCES/ENOTDIR/ cross-device) fall through to `return Ok(())`, so a delete that never landed on a faulty disk was reported as success and the drive fault was hidden from heal/offline logic. It now keeps the DiskFull in-place-remove fallback, treats a missing source (FileNotFound) as benign, and propagates every other error (already mapped by to_file_error, e.g. I/O -> FaultyDisk, permission -> FileAccessDenied), matching MinIO's deleteFile. rename_file and rename_part (#960, ECA-19) directory (trailing-slash) branch unconditionally removed the destination before rename and propagated the resulting NotFound, so renaming a directory to a new location always failed. Both now tolerate ErrorKind::NotFound on that pre-rename remove and continue, matching MinIO's RenameFile. Adds regression tests covering benign-missing-source, real-error propagation, the unchanged move happy path, and directory rename to a missing destination for both rename_file and rename_part. Co-authored-by: heihutu <heihutu@gmail.com> |
||
|
|
20d61c73bc |
fix(ecstore): stop reduce_errs leaking nil placeholder as dominant error (#4551)
reduce_errs seeds best_err with a synthetic Error::other("nil") placeholder
via unwrap_or. When the error slice is empty or every error is ignored,
err_counts is empty and nil_count is 0, so both nil tie-break conditions are
false and the else branch returned (0, Some(Error::other("nil"))). That
placeholder leaked to build_write_quorum_failure_summary, where dominant_error
became Some("nil"), skipping the nil_dominated early return and falling back to
the "other_error" metric label. Quorum decisions were unaffected (max_count 0
still fails any positive quorum), but write-quorum failure metric labels and
tracing fields were misclassified.
Add an explicit short-circuit: when best_count == 0 and nil_count == 0, return
the (0, None) sentinel so the placeholder never escapes. Add regression tests
for the all-ignored slice, empty slice, and the summary label path, which the
existing test_build_write_quorum_failure_summary (nil_count > 0) did not cover.
Fixes #957
Co-authored-by: heihutu <heihutu@gmail.com>
|
||
|
|
2df315baf0 |
fix(ecstore): fsync ancestor dirs for first object writes (#4493)
* fix(ecstore): fsync new object's ancestor dirs to close a power-loss gap reliable_mkdir_all creates an object's directory (and any missing prefix dirs) with plain mkdir and never fsyncs the parent chain. The commit-point fsync in rename_data persists the object dir's *contents* (its xl.meta and data dir), but not the object dir's own entry in the bucket/prefix directory. So on the first PUT of an object, a power loss after the write is acknowledged could drop the whole object directory even though its contents were durable — an acknowledged write silently lost (rustfs/backlog#922 step 4). For a new object (no prior xl.meta) under a durability tier that syncs commit metadata, fsync the ancestor chain from the object dir's parent up to and including the bucket after the commit rename, so the newly created directory entries survive power loss. A starts_with guard bounds the walk to the bucket subtree. Overwrites already have a durable object dir and are unaffected; relaxed/none accept the wider window like the existing commit fsync. Durability regressions are invisible to ordinary behavior tests, so the new tests assert directly (via the fsync_dir recorder) that a first PUT under a new prefix fsyncs both the prefix and bucket dirs, and that relaxed does not. Scope: the non-inline (erasure) commit path. The inline branch has the same gap and is a separate follow-up. Refs: rustfs/backlog#922 (HP-1 step 4), rustfs/backlog#936 Co-Authored-By: heihutu <heihutu@gmail.com> * fix(ecstore): fsync new inline object's ancestor dirs too Extend the backlog#922 step 4 mkdir-gap fix to the inline commit branch. Like the non-inline path, a first PUT of an inline object creates its directory (and any missing prefix dirs) whose entry in the parent chain reliable_mkdir_all never fsynced; the commit fsync persists the object dir's contents, not its own entry. For a new inline object under a commit-metadata-syncing tier, fsync the ancestor chain up to and including the bucket after the commit rename, using the same starts_with-bounded walk (via the synchronous os::fsync_dir_std inside the inline spawn_blocking closure). Adds a test asserting a new inline object under a new prefix fsyncs both the prefix and bucket dirs. rename_data now closes the ack'd-write power-loss gap on both the erasure and inline paths. Refs: rustfs/backlog#922 (HP-1 step 4), rustfs/backlog#936 Co-Authored-By: heihutu <heihutu@gmail.com> --------- Co-authored-by: heihutu <heihutu@gmail.com> |
||
|
|
651ccac130 |
perf(ecstore): parallelize tmp xl.meta write and shard fdatasync on commit (#4487)
The non-inline rename_data commit path made the tmp xl.meta durable and then fdatasynced the shard files in two sequential awaits, each its own blocking round-trip. The two operations touch disjoint paths (the tmp xl.meta under the tmp bucket vs the shard data dir) and have no ordering constraint between them: both only need to be durable before the commit renames that follow. Run them concurrently with tokio::join!, dropping a blocking round-trip from the PUT commit critical path (backlog#922 step 2). The commit ordering is unchanged — both futures complete before any rename, and a failure in either aborts before the rename exactly as the sequential version did (tmp-meta error is still surfaced first). Payload and metadata durability semantics are identical under strict and relaxed tiers. Validated by the rename_data crash-consistency harness (backlog#935): a crash at any pre-commit point still reopens as old-or-new, never mixed. Existing rename_data / durability-tier / disk::os tests are unchanged and pass. Refs: rustfs/backlog#922 (HP-1 step 2), rustfs/backlog#936 Co-authored-by: heihutu <heihutu@gmail.com> |
||
|
|
54872d52dc |
test(ecstore): deterministic rename_data crash-consistency harness (#4478)
The rename_data commit sequence is the highest-risk durability path in the
store, and the HP-1/HP-4/HP-5 work (fsync coalescing, group commit, relaxed
durability tiers) all need a standing gate that proves a power loss mid-commit
can never leave a mixed or corrupt object. There was one graceful-rollback
failpoint but no crash-consistency coverage.
Add a deterministic harness that models a hard power loss — the commit
sequence stops dead at the armed step with no in-process cleanup — and then
reopens the disk to assert the raw on-disk state is coherent:
- Two pre-commit injection points, RenameDataCrashPoint::{AfterDataRename,
AfterBackupBeforeMetaCommit}, constructed at the real commit-path call sites
but gated so the production build compiles them to a const-false no-op
(mirrors the existing should_fail_before_old_metadata_backup pattern).
- A parameterized scenario over {strict, relaxed} durability x {both crash
points} x {overwrite, fresh object}: seed, stage a replacement, inject,
reopen, and assert the object reads back as exactly the old version (or does
not exist when there was no old version) with its data dir intact, never the
half-committed new one. The un-injected run asserts the commit makes the new
version visible. Relaxed is held to the same old-or-new invariant as strict
(only the durability window widens), which is exactly the property the
durability-relaxation work must not break (rustfs/backlog#878 hard rule).
- Wire an explicit `ecstore-crash-consistency` gate into the destructive
profile of run_ecstore_validation_suite.sh so it runs in the standing suite
(coordinates with the #878 destructive profile rather than building a
parallel one).
Production behavior is unchanged: the injection guards are no-ops outside
tests, and the existing rename_data rollback tests still pass.
Refs: rustfs/backlog#935 (HP-14 crash-consistency harness), rustfs/backlog#896
(test plan), rustfs/backlog#878 (ECStore validation suite), rustfs/backlog#936
Co-authored-by: heihutu <heihutu@gmail.com>
|
||
|
|
506cd156bb |
fix(scanner): scope long walk timeouts (#4376)
* fix(scanner): scope long walk timeouts * fix(scanner): bound IAM config walks --------- Co-authored-by: Henry Guo <marshawcoco@users.noreply.github.com> |
||
|
|
92c8c6db75 |
perf(ecstore): Linux O_DIRECT shard-write path off the commit critical section (#4411)
`create_file` writes erasure shard / multipart part data through the page cache, so the commit-point `sync_dir_files` fdatasync in `rename_data` has to flush the whole shard's dirty pages (profiling: create_file avg 67.52ms vs 1.51ms nosync; BitrotWriter::write p95 40.86ms), and concurrent PUTs on one device stall each other's writeback inside the rename critical section. Add an opt-in Linux O_DIRECT streaming writer that reuses the direct-io read infrastructure (statx DIOALIGN probe, aligned bounce buffer, supported latch, one-time fallback log). Incoming bytes are staged into an aligned buffer; full aligned batches and the shutdown tail are flushed on the blocking pool (same offloading posture as the buffered writer it replaces). The sub-alignment tail is written buffered after clearing O_DIRECT via fcntl (MinIO's recipe). Shard bytes reach the device during the write phase, so the unchanged commit-point `sync_dir_files` fdatasync degrades to a cheap metadata/device FLUSH instead of flushing ~2MiB of dirty pages. Gated by `RUSTFS_OBJECT_DIRECT_IO_WRITE_ENABLE` (default false), Linux-only, and durability-neutral: `sync_dir_files` still fdatasyncs every shard at the commit point, so the strict/relaxed/none durability tiers keep their exact guarantees. Filesystems that reject O_DIRECT (tmpfs, overlayfs, 9p) latch the path off and fall back to buffered writes; an O_DIRECT open/write EINVAL is never surfaced as `InvalidInput` (which `to_file_error` maps to `FileNotFound` and would masquerade as a missing shard, triggering a spurious EC rebuild). Unit tests cover the env gate, the alignment/staging-capacity and tail-split math, an end-to-end `create_file` round trip across sizes crossing the alignment and multi-batch boundaries (buffered fallback on macOS / O_DIRECT on a block device), and the writer state machine over a plain file on Linux. Real O_DIRECT throughput on Linux NVMe (ext4/xfs) is deferred to the azure 4-node bench per the issue's GO/NO-GO gate; the development host is macOS and cannot exercise O_DIRECT. Refs: rustfs/backlog#927, rustfs/backlog#936 Co-authored-by: heihutu <heihutu@gmail.com> |
||
|
|
13e48d93aa |
perf(ecstore): support per-bucket durability tier overrides (#4407)
perf(ecstore): per-bucket durability tier overrides (HP-5 phase 2) Let a bucket override the process-wide RUSTFS_DURABILITY_MODE with its own strict/relaxed/none tier, stored as a durability.json extension entry in the bucket metadata file and resolved at commit points via effective_durability. System-critical buckets (.rustfs.sys, .minio.sys) can never carry an override and stay pinned to strict; the legacy full-off switch keeps its historical semantics and per-bucket overrides do not apply under it. Overrides are published and cleared through the existing bucket metadata cache-invalidation path, and an admin GET/PUT handler exposes the configuration. Default behavior is unchanged: with no override a bucket follows the global mode, which defaults to strict and stays byte-for-byte identical to before. Refs: https://github.com/rustfs/backlog/issues/938, https://github.com/rustfs/backlog/issues/936 Co-authored-by: heihutu <heihutu@gmail.com> |
||
|
|
f64f0ca91a | fix(ecstore): report bitrot verifier mismatches as corrupt (#4414) | ||
|
|
d25ddb0e1e |
fix(logging): reduce listing cancellation error noise (#4372)
* fix(logging): reduce listing cancellation error noise * fix(ecstore): preserve filemeta io error kind * fix(ecstore): avoid redundant clone lint in test --------- Co-authored-by: houseme <housemecn@gmail.com> Co-authored-by: overtrue <anzhengchao@gmail.com> |
||
|
|
eaff17cade | perf(ecstore): add strict/relaxed/none durability modes over the drive-sync switch (#4397) | ||
|
|
3f13d098b4 |
feat(observability): feature-gated hotpath instrumentation for the data path (#4394)
Merge the hotpath-rs wall-time instrumentation from the backlog#936 analysis worktree behind an opt-in 'hotpath' cargo feature, keeping the default build at zero overhead and zero dependency. - hotpath is an optional dependency everywhere (dep:hotpath feature syntax); the default dependency tree contains no hotpath crate at all - 40+ measurement points across S3 handlers, ECStore/SetDisks object and multipart ops, erasure encode/decode, bitrot, LocalDisk I/O, FileMeta codec, and HashReader - attribute sites use #[cfg_attr(feature = "hotpath", hotpath::measure)]; async_trait bodies use per-crate hp_guard! macros (ecstore + rustfs bin); rio gates measure_block! behind hp_measure_block! - feature chain: rustfs -> rustfs-ecstore -> rustfs-rio / rustfs-filemeta, each crate owning its own gate - hotpath-alloc is intentionally not wired up (hotpath 0.21.x TLS panic on cross-thread guard drop under tokio, see backlog#935); mimalloc stays the unconditional global allocator - docs/development/hotpath-profiling.md documents building, HOTPATH_* env vars, SIGTERM report flow, and how to reproduce the backlog#936 timing reports Refs: https://github.com/rustfs/backlog/issues/935 (HP-14, item 2), https://github.com/rustfs/backlog/issues/936 Co-authored-by: heihutu <heihutu@gmail.com> |
||
|
|
65953bfdb3 | fix(ecstore): reduce rename data version signatures (#4383) | ||
|
|
e7cc719c17 |
perf(ecstore): move speculative PUT-tail tmp cleanup off the hot path (#4389)
* perf(ecstore): move speculative tmp cleanup off the PUT hot path On a successful PUT, rename_data has already moved the data dir out of the tmp workspace, so the delete_all(RUSTFS_META_TMP_BUCKET) at the end of SetDisks::put_object is a speculative no-op safety net. It was awaited inline on the response path, where profiling (backlog#924 / HP-3) showed the same-disk queueing behind fsync-heavy load turns a ~49us no-op into ~9ms average (p99 77ms, macOS F_FULLFSYNC amplified) added to every PUT. Run that cleanup on a spawned task instead, keeping it as a real backstop (rename_data's remove_std only removes empty dirs and silently ignores failures). The failure path (quorum loss / rollback) keeps the cleanup inline so a failed PUT never returns with tmp shards still on disk. If the process dies before the spawned task runs, cleanup_stale_tmp_objects (24h expiry, 5-minute loop) reclaims the entry. Scope note: ops/multipart.rs delete_all on RUSTFS_META_MULTIPART_BUCKET is intentionally untouched; it removes real leftovers and deferring it would widen CompleteMultipartUpload/Abort races. Regression tests (hermetic SetDisks on formatted local disks, no global state): PUT success drains the tmp workspace (polling the spawned task), and PUT failure (missing bucket volume, rename_data quorum error after tmp shards were written) cleans the workspace inline before returning. Ref: https://github.com/rustfs/backlog/issues/924 Co-Authored-By: heihutu <heihutu@gmail.com> * fix(ecstore): do not retry NotFound in reliable_rename_inner reliable_rename_inner blindly retried the rename once on any error. A NotFound retry cannot succeed: nothing recreates the missing source or parent directory between attempts, so the second rename fails identically and speculative cleanup renames (e.g. move_to_trash on an already-removed tmp path) always paid for two syscalls. Extract the retry decision into should_retry_rename: NotFound returns immediately, any other error keeps the existing single retry. This helper is shared by the rename_data commit path via rename_all, so behavior there is covered by a new rename_all success regression test alongside the retry-predicate tests. Ref: https://github.com/rustfs/backlog/issues/924 Co-Authored-By: heihutu <heihutu@gmail.com> --------- Co-authored-by: heihutu <heihutu@gmail.com> |
||
|
|
062a68d151 |
perf(ecstore): skip tmp parent dir fsync for write-then-rename paths (#4387)
Step 1 of 4 for rustfs/backlog#922 (HP-1): write_all_internal previously coupled file-content durability (fdatasync) with directory-entry durability (parent dir fsync) behind a single bool. For tmp files that are immediately renamed away, the tmp parent dir fsync contributes nothing to crash consistency: the safe-rename recipe only needs file content fdatasync -> rename -> fsync of the destination parent, because the rename removes the tmp directory entry and a crash before the rename means the PUT was never acknowledged. Replace the bool with a SyncMode enum (None / FileAndDir / FileOnly) and use FileOnly at exactly the two write-then-rename tmp write points: the tmp xl.meta write in the non-inline rename_data path and the tmp write inside write_all_meta. write_all_public (format.json etc.) and the old-metadata rollback backup keep FileAndDir since those files stay where they are written. The rest of the commit sequence (shard sync_dir_files, rename, destination parent fsync) is untouched, and behavior with drive sync disabled is unchanged. This saves one directory fsync per disk per non-inline PUT (4 on a 4-drive set). Unit tests assert, via a test-only fsync-dir recorder, that tmp write points no longer fsync their parent while the public write point, the rollback backup, and the commit-rename destination parent still do. Co-authored-by: heihutu <heihutu@gmail.com> |
||
|
|
62a31e4ec4 | fix(storage): address pending metadata and health gaps (#4380) | ||
|
|
5533e080b0 |
feat(ecstore): LocalIoBackend trait and O_DIRECT read with fallback (#4365)
* refactor(ecstore): extract LocalIoBackend trait behind LocalDisk Model the LocalDisk per-file I/O hot path as a LocalIoBackend trait (pread_bytes / open_read_stream / open_full_read / open_write) and move the existing method bodies verbatim into a default StdBackend. LocalDisk holds Arc<dyn LocalIoBackend> and the DiskAPI methods forward to it. This is a behavior-preserving refactor: no logic, error-mapping, metrics, or cfg-branch changes. It creates the seam for an alternative runtime-probed io_uring backend (rustfs/backlog#894) without touching DiskAPI callers. Commit-point durability (fdatasync -> rename -> fsync-dir in rename_data) deliberately stays outside the trait. Add a differential test asserting all four read shapes return identical bytes across page-boundary and file-tail ranges. Tracking: rustfs/backlog#891 (parent rustfs/backlog#897) Co-Authored-By: heihutu <heihutu@gmail.com> * feat(ecstore): true O_DIRECT read path with per-disk graceful fallback (#4366) * feat(ecstore): true O_DIRECT read path with per-disk graceful fallback Implement a real O_DIRECT positioned read inside StdBackend::pread_bytes (Linux only) and wire up the previously dead RUSTFS_OBJECT_DIRECT_IO_READ_* knobs, which had zero call sites. Open with rustix OFlags::DIRECT, probe the DIO alignment once per disk via statx STATX_DIOALIGN (4096 fallback), read the aligned superset into an alignment-allocated bounce buffer with a short-read loop, and slice out the exact logical range so padding never reaches BitrotReader. Any O_DIRECT failure falls back to the buffered read methods; EINVAL or EOPNOTSUPP (tmpfs, overlayfs, 9p) latches the path off per disk with one warning. O_DIRECT errors never surface: EINVAL maps to FileNotFound in to_file_error and would trigger spurious EC rebuilds. Default behavior is unchanged (knob off). macOS keeps F_NOCACHE. Tracking: rustfs/backlog#892 (parent rustfs/backlog#897) Co-Authored-By: heihutu <heihutu@gmail.com> * test(ecstore): add P1.5 benchmark gate harness for O_DIRECT reads (#4369) * test(ecstore): add P1.5 benchmark gate harness for O_DIRECT reads Add an ignored, Linux-only release-mode test (direct_read_bench_gate) that measures DiskAPI::read_file_mmap_copy through a real LocalDisk with cold-cache enforcement (fadvise DONTNEED between rounds), reporting p50/p95/p99/mean latency, wall time, process CPU time, and throughput as one JSON line for A/B diffing between the buffered baseline and the O_DIRECT candidate selected by the production env knobs. Content is verified byte-for-byte before any timing starts. Tracking: rustfs/backlog#893 (parent rustfs/backlog#897) Co-Authored-By: heihutu <heihutu@gmail.com> * test(ecstore): add concurrency dimension to P1.5 bench harness Model the EC GET shape (FuturesUnordered over concurrent shard reads) via RUSTFS_BENCH_CONCURRENCY (default 1, sequential as before). Needed to evaluate blocking-pool pressure for the io_uring gate decision. Tracking: rustfs/backlog#893 Co-Authored-By: heihutu <heihutu@gmail.com> --------- Co-authored-by: heihutu <heihutu@gmail.com> --------- Co-authored-by: heihutu <heihutu@gmail.com> * ci: retrigger after cancelled required check Co-Authored-By: heihutu <heihutu@gmail.com> * fix(ecstore): use is_multiple_of in O_DIRECT bench cache-drop check clippy's manual_is_multiple_of (rust 1.96) fails -D warnings on the benchmark helper's `done % file_count == 0`. Verification: - cargo clippy -p rustfs-ecstore --all-targets Co-Authored-By: heihutu <heihutu@gmail.com> --------- Co-authored-by: heihutu <heihutu@gmail.com> |
||
|
|
0ce0388fc0 |
fix(ecstore): fsync inline rename_data rollback backup (#4346)
* fix(ecstore): fsync inline rename_data rollback backup The inline branch of LocalDisk::rename_data writes the previous version's xl.meta rollback backup (<old_data_dir>/xl.meta.bkp) with a bare std::fs::write, fsyncing neither the file nor its new directory entry, even when drive_sync_enabled() is true. The same closure already syncs the new xl.meta and the commit rename directory, and the non-inline branch persists its backup durably. That backup is the sole restore source for delete_version(undo_write=true) on a set-level write-quorum failure. With sync enabled, an inline overwrite plus power loss plus a quorum failure that triggers undo_write could restore a lost or truncated backup and fail to roll back to the prior committed object. Write the inline backup via OpenOptions + write_all and, when sync is enabled, sync_data() it and fsync_dir_std its parent, mirroring the non-inline branch. Extend the inline backup test to assert the .bkp contents equal the previous metadata bytes. Refs backlog#868 (disk-durability-01). * fix(ecstore): preserve to_file_error mapping on inline backup write Address review: the rewritten inline rollback-backup write must keep the same DiskError classification as the previous std::fs::write(...).map_err( to_file_error)? call. Without it, io errors (PermissionDenied/NotFound/ StorageFull/...) would surface as DiskError::Io(_) instead of the specific DiskError variants (FileAccessDenied/FileNotFound/DiskFull/...), since From<io::Error> for DiskError only recovers variants that were wrapped via to_file_error. Map open/write_all/sync_data/fsync_dir_std through to_file_error, matching write_all_internal. --------- Co-authored-by: houseme <housemecn@gmail.com> |
||
|
|
5594b18912 |
fix(ecstore): make delete_volume non-recursive by default to prevent bucket-heal wipe (backlog#799 B1) (#4339)
* fix(ecstore): make delete_volume non-recursive by default to prevent bucket-heal wipe (backlog#799 B1) `delete_volume` unconditionally `remove_dir_all`'d the whole bucket tree, and the bucket-heal "remove" branch called it fire-and-forget on every local disk. A mis-classified "dangling" bucket (or a non-force S3 DeleteBucket on a populated bucket) was therefore recursively wiped — a potential whole-bucket data loss. The `VolumeNotEmpty` -> recreate/`BucketNotEmpty` handling already present in both delete_bucket paths was dead code because the primitive never refused. Add an explicit `force_delete` flag to `DiskAPI::delete_volume` and default the non-force path to a non-recursive `remove_dir` (rmdir), which fails atomically with `VolumeNotEmpty` if the bucket still holds any object data. Only an explicit force delete (S3 force bucket delete) removes recursively. Mirrors MinIO's `xlStorage.DeleteVol` (`Remove` vs `RemoveAll`). - Trait + all impls (local behavior, dispatch, disk_store, remote RPC) take the flag; the gRPC `DeleteVolumeRequest` gains a `force` field (proto3 default false → old peers get the safe non-recursive behavior on rolling upgrade). - Heal remove branch passes `false` and no longer discards the result: a `VolumeNotEmpty` refusal is logged (the bucket is not dangling) instead of wiping data. - Both `delete_bucket` paths pass `opts.force`, activating the previously-dead `VolumeNotEmpty` -> `BucketNotEmpty`/recreate handling (correct S3 semantics). Adds a regression test: non-force delete of a non-empty bucket returns VolumeNotEmpty and preserves the data; force delete removes it. Design converged by two independent expert reviews (MinIO-fidelity + defense-in-depth) referencing MinIO xl-storage.go. Refs backlog#799 (B1), issue rustfs/backlog#850. The safety expert's deeper hardening (typed capability instead of a bool, trash-instead-of-in-place for force, quorum re-verification of dangling) is noted on #850 as follow-up. * fix(ecstore): reword 'mis-classified' -> 'misclassified' to satisfy typos (backlog#799 B1) * fix(rustfs): thread force_delete through StorageDiskRpcExt::delete_volume + test literal (backlog#799 B1) |
||
|
|
31c6859965 |
chore: converge stale TODOs and apply safe fills (backlog#646) (#4322)
Second TODO-convergence round over the current tree (backlog#646). All line numbers in the old inventory had gone stale after the set_disk / diagnostics / cluster refactors, so this re-scans and reduces the marker count from 144 to 99. STALE removals (comment describes already-implemented behavior, or dead commented-out blocks) across ecstore (set_disk ops/core, store, cluster/rpc, bucket/metadata_sys, services), iam, filemeta, s3select and rustfs auth/object_usecase. No behavior change. Safe fills, each verified: - filemeta: replication_info_equals now also compares replication_state_internal (function currently has no callers; adds a regression test). - bitrot: drop the confirmed-unused `_want` parameter from bitrot_verify and the now-unused `sum` on LocalDisk::bitrot_verify, removing a Bytes::copy_from_slice allocation. Streaming verify uses the file's embedded per-shard hash, never the passed sum. - signer: rename v4_ignored_headers -> V4_IGNORED_HEADERS and drop the non_upper_case_globals allow. - admin/heal: test_decode was #[ignore]d and used serde_urlencoded on a JSON body (would panic); rewire to serde_json::from_slice to match the production decode path, add assertions, un-ignore. Verified: cargo fmt; cargo check on touched crates; tests pass (filemeta, signer, bitrot, heal::test_decode); arch guardrail scripts pass. |