mirror of
https://github.com/rustfs/rustfs.git
synced 2026-08-26 05:56:50 +00:00
Address P2 follow-ups from the 2026-07-10..12 merged-PR review (backlog#1210-1220) (#4783)
* fix(obs): open cleaner compression source with O_NOFOLLOW The compressor opened the source log via File::open, which follows a symlink at the final path component. Between the scanner selecting a regular file and this open, an attacker with write access to the log directory could swap the entry for a symlink (TOCTOU) pointing at, say, /etc/shadow, whose contents would then be copied into an archive. Open the source with O_NOFOLLOW on Unix so such a swap fails with ELOOP; the temp/archive path already refused symlinks, this closes the source side. Refs rustfs/backlog#1210 Co-Authored-By: heihutu <heihutu@gmail.com> * fix(obs): recompress instead of trusting leftover cleaner archives archive_header_ok only checked the first 2-4 magic bytes before treating an existing .gz/.zst as a completed prior result and letting the caller delete the source log. A file with valid magic but a truncated or forged body passes that check, so an attacker with write access to the log directory (or a crashed prior run) could plant such a stub and make the cleaner delete the real log without ever producing a usable archive — silent audit-data loss. Chosen fix: stop trusting cross-process leftovers entirely and always recompress the source in this pass, rather than fully decoding every leftover to validate it. Full-decode validation would add real CPU cost and decode-bug surface for a rare crash-recovery case; the existing atomic create_new+rename already overwrites whatever sits at the archive path (a planted symlink is replaced, never followed) with a freshly written, fsync'd archive, so a partial/forged leftover can never gate source deletion. This is the lowest-regression option. Refs rustfs/backlog#1211 Co-Authored-By: heihutu <heihutu@gmail.com> * fix(object-data-cache): cap memory-gate reservation at cache growth headroom The memory gate subtracts `admitted_since_refresh` from the snapshot's available bytes so a burst arriving faster than the 5 s refresh cannot over-allocate. That counter is GROSS: it only rolls over on the refresh and never rolls back when a fill is later evicted, cancelled, or loses the invalidation race. Under sustained high-throughput churn (net footprint flat and far below `max_capacity`) the raw counter balloons past the memory the cache actually holds, so `effective_available` collapses and the gate reports false memory pressure — skipping the hottest fills with SkippedMemoryPressure until the next 5 s refresh. This only lowers hit rate; it never returns wrong data and self-heals each refresh. Fix direction 1 (minimal regression): cap the reservation deduction at the cache's own growth headroom (`max_capacity - weighted_size()`) instead of letting the unbounded gross counter shrink the system-available budget. The cache can never hold more than `max_capacity`, so a burst adds at most that headroom of real memory before moka evicts to stay bounded (net-zero churn beyond that point) — capping the deduction there keeps the reservation honest without treating gross churn as growth. Chosen over net-accounting (direction 2, releasing bytes on every failure/cancel/eviction path) because that only plugs the leak on failed fills and would not address the core defect: churn of *successful* insert/evict fills over the 5 s window. It also touches only the gate plus one call site rather than every failure path in moka_backend. The cap only ever raises `effective_available`, so real memory pressure (a low snapshot at refresh) still suppresses fills; when the cache is at capacity the headroom is 0 and the deduction vanishes, correctly reflecting net-zero churn. `MokaBackend` now stores `max_capacity` and passes the live headroom into `allows_fill`. Adds targeted gate tests: gross churn far above headroom no longer falsely suppresses, yet the reservation still bounds a burst while the cache can genuinely grow. Refs rustfs/backlog#1212 Co-Authored-By: heihutu <heihutu@gmail.com> * test(ecstore): assert native O_DIRECT path runs in uring read test uring_preserves_o_direct_for_eligible_reads only compared bytes through LocalDisk::read_file_mmap_copy. On a filesystem that rejects O_DIRECT the read silently degrades to the buffered StdBackend fallback and the byte check still passes, so the test could go green without the native read_at_direct path ever executing -- a vacuous pass. Add a per-disk native_direct_reads counter on UringBackend, incremented only when pread_uring_direct completes, and rebuild the test to drive a real UringBackend's pread_bytes and assert the counter is non-zero (every eligible read went through the native tier). When io_uring or O_DIRECT is unavailable on the host filesystem (restricted CI runners, tmpfs), the test skips loudly via eprintln instead of asserting a tautology, while still checking byte-correctness on whatever tier served the read. The counter also gives a gray release a positive signal that the O_DIRECT tier is serving reads, not just a fallback count. Refs rustfs/backlog#1213 Co-Authored-By: heihutu <heihutu@gmail.com> * fix(ecstore): warn + count read-time EINVAL on native O_DIRECT reads classify_direct_read_error is only reached from the read side: the O_DIRECT open in pread_uring_direct already succeeded (an open-time refusal is handled earlier as DirectOpenError::ODirectRefused). So an EINVAL/EOPNOTSUPP arriving here is a read-time error on an fd the kernel accepted for O_DIRECT -- far more likely an alignment bug in the aligned read path than an unsupported filesystem. The old code latched the disk's native path off with only a once-per-disk debug trace, hiding a potential correctness regression behind a silent buffered-read downgrade. Diagnostics only: the fallback behaviour is unchanged (the native path is still latched off and the caller still reads via StdBackend). This adds a rustfs_io_uring_direct_read_einval_total counter and promotes the once-per-disk trace from debug to warn so an operator can see an alignment regression instead of an unexplained latency/CPU shift. Refs rustfs/backlog#1214 Co-Authored-By: heihutu <heihutu@gmail.com> * docs(ecstore): document data-blocks-first default and its tail-latency cost DEFAULT_RUSTFS_GET_DATA_BLOCKS_FIRST_READER_SETUP is true and must stay true: deferred-parity is the deliberate, already-rolled-out full-object GET default from backlog#1159/#923. Flipping it back to false in code would silently revert that rollout for every deployment that has not set the env var, so this commit only documents -- no behaviour change. The added notes explain what data-blocks-first does (schedule data shards up front, engage parity lazily on a missing/corrupt data shard), the known trade-off (parity is engaged late, so a slow-but-not-dead data drive raises GET p99 because the faster parity shards are not raced against it until a data shard is declared missing), and the operational rollback switch (RUSTFS_GET_DATA_BLOCKS_FIRST_READER_SETUP=false), which is intentionally an env override rather than a code default change. No metric was added: the low-risk observability hook for "slow data drive engaged deferred parity" would live at the deferred-stripe engage point, which is out of this file's scope; this change stays documentation-only to avoid touching the hot GET path. Refs rustfs/backlog#1215 Co-Authored-By: heihutu <heihutu@gmail.com> * docs(ecstore): document wide-directory walk stall hazard and tuning list_dir enumerates a whole directory in one os::read_dir call (count = -1), and the walk caller bounds that entire enumeration with the per-read stall budget (default 5s) as if it were a single read. For a wide, flat prefix -- one directory holding millions of immediate children -- a single readdir can exceed the budget on a healthy disk, trip DiskError::Timeout, and surface as a ListObjects 500 quorum failure though the drive is fine (a #2999 sub-class). This is documented, not rewritten: turning the one-shot readdir into a streaming/batched enumeration that refreshes the stall deadline between chunks is an architecture-level change with high regression surface (ordering, the count contract, quorum merge) and belongs in a separate follow-up. The supported mitigation today is operational, so the comments point wide-directory deployments at RUSTFS_DRIVE_WALKDIR_STALL_TIMEOUT_SECS and the high-latency drive-timeout profile, which widen the budget with no code change. Notes were added at list_dir, the scan_dir call site, and get_drive_walkdir_stall_timeout. No behaviour change. Refs rustfs/backlog#1216 Co-Authored-By: heihutu <heihutu@gmail.com> * docs(ecstore): document consumer-peek vs producer-stall coupling In list_path_raw the consumer's peek_timeout is drawn from the same source and same value (walkdir_stall_timeout, default 5s) as the producer-side walk stall budget, but the two measure different things: the producer stall bounds a single drive read, while the consumer peek bounds the gap between two ADJACENT entries arriving from a reader. Because they share a value, the consumer cannot wait meaningfully longer for the next entry than the producer is allowed to spend producing one. Walking a region dense with non-listable internal items can make a HEALTHY drive miss the budget between visible entries; the consumer then declares it stalled and detaches it, dropping a good drive from the merge and capping the "large prefix succeeds" guarantee. Documented, not decoupled: giving the consumer peek an independent, strictly-larger budget would cut these false detaches but equally delays detaching a genuinely dead drive and shifts listing tail-latency semantics, so it wants soak data before changing the default. The comment records the invariant any such follow-up must keep -- consumer peek >= producer stall, never stricter -- so it can never fail a drive before the producer would. No behaviour change. Refs rustfs/backlog#1217 Co-Authored-By: heihutu <heihutu@gmail.com> * fix(io-metrics): add time-based trigger for low-IOPS latency percentiles Percentiles were recomputed only every 128 IOs and seeded to 0, so a low-traffic deployment exported p95/p99 = 0/stale for a long time after startup. Add a 10s wall-clock trigger alongside the count throttle so the first recompute can fire before 128 samples accrue. Hot-path per-op mean update is unchanged. Refs rustfs/backlog#1218 Co-Authored-By: heihutu <heihutu@gmail.com> * test(e2e): cover codec-streaming parity under fault injection and NoSuchKey The codec-streaming compat A/B previously ran only against a healthy 4-disk EC set with successful full GETs: the DiskFaultHarness was constructed but never faulted, the error path was untested, and the range assertion silently compared legacy-vs-legacy (ranges always fall back to the duplex path), overstating what it proved. Add two genuinely-failable scenarios reusing the existing harness and fixtures: - Parity reconstruction A/B: take one data disk offline and re-run the full object matrix on both phases while the EC 2+2 set rebuilds each large object from the surviving shards. Assert codec == legacy byte-for-byte (sha256) and header-for-header, and assert the codec phase served the reconstructed objects with zero duplex-pipe fallback (the reader gate is drive-health-independent, so the codec fast path is really exercised through reconstruction). - NoSuchKey negative path: compare the HTTP status + S3 error code of a missing-key GET across the legacy and codec phases and require them to be identical (404/NoSuchKey), guarding against the codec env perturbing the error path. Also clarify the range-phase comment so it is not misread as codec-range correctness coverage: both sides are served by the same legacy range path, so the assertion only proves ranges keep working and keep falling back to legacy with the gates open. Verified: cargo check/--no-run pass and the test passes locally (1 passed; dup_codec=0 confirms the codec path ran). Refs rustfs/backlog#1219 Co-Authored-By: heihutu <heihutu@gmail.com> * ci(ecstore): exercise native O_DIRECT read path on an ext4 loopback The uring-integration leg ran on the runner's default TMPDIR, which may sit on tmpfs/overlayfs where open(O_DIRECT) fails and the native read_at_direct path silently latches off to the aligned StdBackend fallback. Mount a dedicated ext4 loopback and point TMPDIR at it so the real io_uring dep (bumped git->0.1.0->0.2.0->0.2.1) and the native O_DIRECT read path are actually covered rather than validated only by signature diffing. Refs rustfs/backlog#1220 Co-Authored-By: heihutu <heihutu@gmail.com> --------- Co-authored-by: heihutu <heihutu@gmail.com>
This commit is contained in:
@@ -189,6 +189,18 @@ pub fn get_drive_walkdir_timeout() -> Duration {
|
||||
)
|
||||
}
|
||||
|
||||
/// Per-read stall budget for a directory walk: a walk read is failed only if
|
||||
/// the drive stops answering for this long, not for a walk simply taking a
|
||||
/// while (see `with_walk_stall_deadline` in `disk/local.rs`).
|
||||
///
|
||||
/// Wide-directory tuning (rustfs/backlog#1216): because a whole-directory
|
||||
/// enumeration (`list_dir` with `count = -1`) is bounded by this budget as one
|
||||
/// unit, a very wide flat prefix (millions of immediate children) can make a
|
||||
/// single `readdir` exceed the default on a healthy disk and fail ListObjects.
|
||||
/// Deployments with such directories should raise
|
||||
/// `RUSTFS_DRIVE_WALKDIR_STALL_TIMEOUT_SECS`, or select the high-latency
|
||||
/// drive-timeout profile (which raises this default automatically), to widen
|
||||
/// the budget without a code change.
|
||||
pub fn get_drive_walkdir_stall_timeout() -> Duration {
|
||||
get_drive_timeout_duration(
|
||||
rustfs_config::ENV_DRIVE_WALKDIR_STALL_TIMEOUT_SECS,
|
||||
|
||||
@@ -266,6 +266,13 @@ const METRIC_URING_IN_FLIGHT: &str = "rustfs_io_uring_in_flight";
|
||||
const METRIC_URING_CQ_OVERFLOW: &str = "rustfs_io_uring_cq_overflow";
|
||||
#[cfg(target_os = "linux")]
|
||||
const METRIC_URING_CANCEL_ALREADY: &str = "rustfs_io_uring_cancel_already";
|
||||
/// Read-side EINVAL/EOPNOTSUPP from a native O_DIRECT read that happened AFTER a
|
||||
/// successful O_DIRECT open (rustfs/backlog#1214). Unlike an open-time refusal
|
||||
/// (unsupported filesystem), this most likely means an alignment bug in the
|
||||
/// aligned read path, so it is surfaced with a counter + warn instead of a
|
||||
/// once-per-disk debug trace.
|
||||
#[cfg(target_os = "linux")]
|
||||
const METRIC_URING_DIRECT_READ_EINVAL_TOTAL: &str = "rustfs_io_uring_direct_read_einval_total";
|
||||
/// How often the per-disk driver StatsSnapshot is exported to metrics
|
||||
/// (rustfs/backlog#1172).
|
||||
#[cfg(target_os = "linux")]
|
||||
@@ -2961,6 +2968,14 @@ pub(crate) struct UringBackend {
|
||||
/// `active` gates io_uring as a whole, this gates only the native O_DIRECT
|
||||
/// read shape.
|
||||
direct_uring: DirectIoReadState,
|
||||
/// Count of reads that completed through the native io_uring + O_DIRECT path
|
||||
/// (`pread_uring_direct`) on this disk (rustfs/backlog#1213). Incremented only
|
||||
/// on success, so a value `> 0` is proof the native path actually executed
|
||||
/// rather than silently degrading to the StdBackend fallback. Tests assert on
|
||||
/// it to avoid a vacuous pass on filesystems that reject O_DIRECT; it also
|
||||
/// gives a gray release a positive signal that the O_DIRECT tier is serving
|
||||
/// reads instead of only ever counting fallbacks.
|
||||
native_direct_reads: std::sync::atomic::AtomicU64,
|
||||
/// Per-disk descriptor cache (backlog#1145). `None` when
|
||||
/// `RUSTFS_IO_URING_FD_CACHE` is off, which restores the open-per-read path.
|
||||
fd_cache: Option<FdCache>,
|
||||
@@ -3101,6 +3116,7 @@ impl UringBackend {
|
||||
active: std::sync::atomic::AtomicBool::new(true),
|
||||
fallback_logged: std::sync::atomic::AtomicBool::new(false),
|
||||
direct_uring: DirectIoReadState::new(),
|
||||
native_direct_reads: std::sync::atomic::AtomicU64::new(0),
|
||||
fd_cache,
|
||||
})
|
||||
}
|
||||
@@ -3209,6 +3225,36 @@ impl UringBackend {
|
||||
/// classification lives in one place (rustfs/backlog#1174).
|
||||
fn classify_direct_read_error(&self, io_err: &std::io::Error) {
|
||||
if is_direct_io_unsupported(io_err) {
|
||||
// This helper is only ever reached from the READ side: the O_DIRECT
|
||||
// `open` in `pread_uring_direct` already succeeded, and an open-time
|
||||
// refusal is handled separately as `DirectOpenError::ODirectRefused`
|
||||
// before any read is issued. So an EINVAL/EOPNOTSUPP arriving here is
|
||||
// a *read-time* error on an fd the kernel accepted for O_DIRECT. That
|
||||
// is far more likely an alignment bug in the aligned read path than a
|
||||
// filesystem that does not support O_DIRECT -- yet the old code
|
||||
// latched the whole disk's native path off with only a once-per-disk
|
||||
// debug trace, making a real correctness bug effectively invisible
|
||||
// (rustfs/backlog#1214).
|
||||
//
|
||||
// Diagnostics only: the fallback behaviour is unchanged. The native
|
||||
// O_DIRECT path is still latched off and the caller still falls back
|
||||
// to StdBackend for this and every future eligible read. We only make
|
||||
// the event observable -- a counter plus a once-per-disk `warn!`
|
||||
// instead of a silent `debug!` -- so an operator can see an alignment
|
||||
// regression rather than a mystery latency/CPU shift from buffered
|
||||
// reads.
|
||||
counter!(METRIC_URING_DIRECT_READ_EINVAL_TOTAL, "root" => self.root_label.clone()).increment(1);
|
||||
if !self.direct_uring.fallback_logged.swap(true, Ordering::Relaxed) {
|
||||
warn!(
|
||||
component = LOG_COMPONENT_ECSTORE,
|
||||
subsystem = LOG_SUBSYSTEM_DISK_LOCAL,
|
||||
root = %self.root.display(),
|
||||
error = ?io_err,
|
||||
"io_uring O_DIRECT read returned EINVAL/EOPNOTSUPP AFTER a successful O_DIRECT open; \
|
||||
this is more likely an alignment bug than an unsupported filesystem. Latching the \
|
||||
native O_DIRECT path off and reading via StdBackend (logged once per disk)"
|
||||
);
|
||||
}
|
||||
self.direct_uring.supported.store(false, Ordering::Relaxed);
|
||||
} else if is_io_uring_unsupported(io_err) {
|
||||
self.latch_active_off(io_err);
|
||||
@@ -3478,6 +3524,10 @@ impl UringBackend {
|
||||
if should_reclaim_file_cache_after_read(length) {
|
||||
reclaim_read_range(&file_for_reclaim, offset_u64, length)?;
|
||||
}
|
||||
// The native io_uring + O_DIRECT read completed (rustfs/backlog#1213):
|
||||
// record it so callers/tests can distinguish this path from the
|
||||
// StdBackend fallback, which never reaches here.
|
||||
self.native_direct_reads.fetch_add(1, Ordering::Relaxed);
|
||||
Ok(Bytes::from(bytes))
|
||||
}
|
||||
}
|
||||
@@ -5037,6 +5087,15 @@ impl LocalDisk {
|
||||
|
||||
let stall = opts.stall_timeout_duration();
|
||||
|
||||
// `count = -1` enumerates the whole directory in one `list_dir` call, and
|
||||
// the stall budget bounds that entire enumeration as a single unit. On a
|
||||
// WIDE, FLAT directory (millions of immediate children) that one readdir
|
||||
// can exceed `stall` on a healthy disk and fail the whole walk -- see the
|
||||
// wide-directory stall hazard documented on `list_dir`
|
||||
// (rustfs/backlog#1216, a #2999 sub-class). Mitigate operationally with a
|
||||
// larger `RUSTFS_DRIVE_WALKDIR_STALL_TIMEOUT_SECS` or the high-latency
|
||||
// drive-timeout profile; a streaming readdir rewrite is a separate,
|
||||
// higher-risk follow-up and is intentionally not done here.
|
||||
let mut entries = match with_walk_stall_timeout(stall, self.list_dir("", &opts.bucket, ¤t, -1)).await {
|
||||
Ok(res) => res,
|
||||
Err(e) => {
|
||||
@@ -6419,6 +6478,29 @@ impl DiskAPI for LocalDisk {
|
||||
self.io_backend.pread_bytes(volume, path, offset, length, metrics).await
|
||||
}
|
||||
|
||||
/// List a single directory. `count < 0` enumerates the *whole* directory in
|
||||
/// one `os::read_dir` call.
|
||||
///
|
||||
/// Wide-directory stall hazard (rustfs/backlog#1216, a #2999 sub-class):
|
||||
/// the walk caller wraps this whole call in the per-read stall budget
|
||||
/// (`with_walk_stall_timeout`, default 5s via
|
||||
/// `RUSTFS_DRIVE_WALKDIR_STALL_TIMEOUT_SECS`) as if the entire directory
|
||||
/// enumeration were a single read. For a *wide, flat* directory -- one
|
||||
/// bucket prefix holding millions of immediate children -- a single
|
||||
/// `readdir` of the whole directory can itself exceed the stall budget on a
|
||||
/// healthy disk. That trips `DiskError::Timeout`, which the listing path can
|
||||
/// escalate to a quorum failure and surface to the client as a ListObjects
|
||||
/// 500, even though nothing is actually wrong with the drive.
|
||||
///
|
||||
/// This is deliberately NOT fixed here by rewriting the one-shot
|
||||
/// `os::read_dir` into a streaming/batched readdir that would refresh the
|
||||
/// stall deadline between chunks: that is an architecture-level change with
|
||||
/// high regression surface (ordering, the `count` contract, quorum merge
|
||||
/// semantics) and is tracked as a separate follow-up. The supported
|
||||
/// mitigation for wide-directory deployments today is operational -- raise
|
||||
/// `RUSTFS_DRIVE_WALKDIR_STALL_TIMEOUT_SECS` or run with the high-latency
|
||||
/// drive-timeout profile (see `get_drive_walkdir_stall_timeout`), both of
|
||||
/// which widen the budget without any code change.
|
||||
#[tracing::instrument(level = "trace", skip_all)]
|
||||
async fn list_dir(&self, origvolume: &str, volume: &str, dir_path: &str, count: i32) -> Result<Vec<String>> {
|
||||
if !origvolume.is_empty() {
|
||||
@@ -6433,6 +6515,9 @@ impl DiskAPI for LocalDisk {
|
||||
let volume_dir = self.get_bucket_path(volume)?;
|
||||
let dir_path_abs = self.get_object_path(volume, dir_path.trim_start_matches(SLASH_SEPARATOR))?;
|
||||
|
||||
// Whole-directory enumeration in one syscall path (see the wide-directory
|
||||
// stall hazard on this fn): with `count < 0` this reads every entry, and
|
||||
// the caller's stall budget bounds the entire call as a unit.
|
||||
let entries = match os::read_dir(&dir_path_abs, count).await {
|
||||
Ok(res) => res,
|
||||
Err(e) => {
|
||||
@@ -14292,12 +14377,23 @@ mod test {
|
||||
/// an O_DIRECT-eligible read keeps O_DIRECT semantics via the native
|
||||
/// `read_at_direct` path (or, if that disk can't do io_uring+O_DIRECT, the
|
||||
/// StdBackend aligned fallback) and still returns exactly the requested
|
||||
/// bytes for unaligned ranges. This asserts byte-correctness regardless of
|
||||
/// which tier serves the read — the native path is exercised directly by
|
||||
/// rustfs-uring's own O_DIRECT test under real io_uring.
|
||||
/// bytes for unaligned ranges.
|
||||
///
|
||||
/// The old shape of this test only checked byte-equivalence through
|
||||
/// `LocalDisk::read_file_mmap_copy`. On a filesystem that rejects O_DIRECT
|
||||
/// the read silently degrades to the buffered fallback and the byte check
|
||||
/// still passes, so the test could go green without the native O_DIRECT path
|
||||
/// ever running — a vacuous pass (rustfs/backlog#1213). It now builds a real
|
||||
/// `UringBackend`, drives `pread_bytes` (which routes eligible reads into
|
||||
/// `pread_uring_direct`), and asserts the native path actually executed via
|
||||
/// the `native_direct_reads` counter. When the backing filesystem cannot do
|
||||
/// io_uring or O_DIRECT (restricted CI runners, tmpfs/overlayfs), the test
|
||||
/// skips loudly with `eprintln!` instead of asserting a tautology — but it
|
||||
/// still checks byte-correctness on whatever tier served the read.
|
||||
#[cfg(target_os = "linux")]
|
||||
#[tokio::test(flavor = "multi_thread")]
|
||||
async fn uring_preserves_o_direct_for_eligible_reads() {
|
||||
use std::sync::atomic::Ordering;
|
||||
use tempfile::tempdir;
|
||||
|
||||
// Unaligned on purpose: 3 blocks + 7 bytes.
|
||||
@@ -14319,27 +14415,43 @@ mod test {
|
||||
(FILE_LEN - 7, 7),
|
||||
];
|
||||
|
||||
// Threshold 1 makes every non-empty read O_DIRECT-eligible, so each read
|
||||
// below exercises the O_DIRECT tier. The env must be set before
|
||||
// LocalDisk::new so the io_uring backend is the one selected.
|
||||
let root_dir = tempdir().expect("operation should succeed");
|
||||
let root = root_dir.path().to_path_buf();
|
||||
|
||||
// Lay out the shard with LocalDisk, then read it back through a real
|
||||
// UringBackend so the test can inspect the O_DIRECT latch and the native
|
||||
// read counter directly.
|
||||
{
|
||||
let endpoint = Endpoint::try_from(root.to_string_lossy().as_ref()).expect("operation should succeed");
|
||||
let disk = LocalDisk::new(&endpoint, false).await.expect("operation should succeed");
|
||||
disk.make_volume("test-volume").await.expect("operation should succeed");
|
||||
disk.write_all("test-volume", "shard.bin", content.clone())
|
||||
.await
|
||||
.expect("operation should succeed");
|
||||
}
|
||||
|
||||
// Skip if io_uring is unavailable on this host (restricted env, e.g. the
|
||||
// Kubernetes CI runners): there is no native O_DIRECT path to exercise.
|
||||
let Some(backend) = UringBackend::try_new(root) else {
|
||||
uring_test_skip("uring_preserves_o_direct_for_eligible_reads");
|
||||
return;
|
||||
};
|
||||
|
||||
// Threshold 1 makes every non-empty read O_DIRECT-eligible, so each read
|
||||
// below drives `pread_bytes` into the native `pread_uring_direct` path.
|
||||
// These knobs are read per-read, so setting them around the reads is
|
||||
// enough (the backend was already constructed above).
|
||||
let got_ranges = temp_env::async_with_vars(
|
||||
[
|
||||
(ENV_RUSTFS_IO_URING_READ_ENABLE, Some("true")),
|
||||
(ENV_RUSTFS_OBJECT_DIRECT_IO_READ_ENABLE, Some("true")),
|
||||
(ENV_RUSTFS_OBJECT_DIRECT_IO_READ_THRESHOLD, Some("1")),
|
||||
],
|
||||
async {
|
||||
let endpoint = Endpoint::try_from(root_dir.path().to_string_lossy().as_ref()).expect("operation should succeed");
|
||||
let disk = LocalDisk::new(&endpoint, false).await.expect("operation should succeed");
|
||||
disk.make_volume("test-volume").await.expect("operation should succeed");
|
||||
disk.write_all("test-volume", "shard.bin", content.clone())
|
||||
.await
|
||||
.expect("operation should succeed");
|
||||
let mut out = Vec::new();
|
||||
for (offset, length) in ranges {
|
||||
out.push(
|
||||
disk.read_file_mmap_copy("test-volume", "shard.bin", offset, length)
|
||||
backend
|
||||
.pread_bytes("test-volume", "shard.bin", offset, length, None)
|
||||
.await
|
||||
.expect("O_DIRECT-eligible read must succeed under io_uring (direct or fallback)"),
|
||||
);
|
||||
@@ -14349,6 +14461,7 @@ mod test {
|
||||
)
|
||||
.await;
|
||||
|
||||
// Byte-correctness holds regardless of which tier served the read.
|
||||
for ((offset, length), got) in ranges.into_iter().zip(got_ranges) {
|
||||
assert_eq!(
|
||||
got,
|
||||
@@ -14356,6 +14469,26 @@ mod test {
|
||||
"O_DIRECT read mismatch at offset={offset} length={length}"
|
||||
);
|
||||
}
|
||||
|
||||
// The point of backlog#1213: prove the NATIVE O_DIRECT path executed
|
||||
// rather than silently passing on the StdBackend fallback. If the
|
||||
// filesystem refuses O_DIRECT, `direct_uring.supported` latches off and
|
||||
// no native read is counted — skip loudly instead of asserting nothing.
|
||||
let native_hits = backend.native_direct_reads.load(Ordering::Relaxed);
|
||||
let still_supported = backend.direct_uring.supported.load(Ordering::Relaxed);
|
||||
if still_supported && native_hits > 0 {
|
||||
assert_eq!(
|
||||
native_hits,
|
||||
ranges.len() as u64,
|
||||
"every eligible read should have gone through the native io_uring O_DIRECT path"
|
||||
);
|
||||
} else {
|
||||
eprintln!(
|
||||
"SKIP uring_preserves_o_direct_for_eligible_reads: native O_DIRECT path not \
|
||||
exercised on this filesystem (direct_uring.supported={still_supported}, \
|
||||
native_direct_reads={native_hits}); byte-correctness was still asserted"
|
||||
);
|
||||
}
|
||||
}
|
||||
|
||||
/// Once io_uring is latched off (backlog#1101), reads still return correct
|
||||
|
||||
Reference in New Issue
Block a user