Address P2 follow-ups from the 2026-07-10..12 merged-PR review (backlog#1210-1220) (#4783)

* fix(obs): open cleaner compression source with O_NOFOLLOW

The compressor opened the source log via File::open, which follows a
symlink at the final path component. Between the scanner selecting a
regular file and this open, an attacker with write access to the log
directory could swap the entry for a symlink (TOCTOU) pointing at, say,
/etc/shadow, whose contents would then be copied into an archive. Open
the source with O_NOFOLLOW on Unix so such a swap fails with ELOOP; the
temp/archive path already refused symlinks, this closes the source side.

Refs rustfs/backlog#1210
Co-Authored-By: heihutu <heihutu@gmail.com>

* fix(obs): recompress instead of trusting leftover cleaner archives

archive_header_ok only checked the first 2-4 magic bytes before treating
an existing .gz/.zst as a completed prior result and letting the caller
delete the source log. A file with valid magic but a truncated or forged
body passes that check, so an attacker with write access to the log
directory (or a crashed prior run) could plant such a stub and make the
cleaner delete the real log without ever producing a usable archive —
silent audit-data loss.

Chosen fix: stop trusting cross-process leftovers entirely and always
recompress the source in this pass, rather than fully decoding every
leftover to validate it. Full-decode validation would add real CPU cost
and decode-bug surface for a rare crash-recovery case; the existing
atomic create_new+rename already overwrites whatever sits at the archive
path (a planted symlink is replaced, never followed) with a freshly
written, fsync'd archive, so a partial/forged leftover can never gate
source deletion. This is the lowest-regression option.

Refs rustfs/backlog#1211
Co-Authored-By: heihutu <heihutu@gmail.com>

* fix(object-data-cache): cap memory-gate reservation at cache growth headroom

The memory gate subtracts `admitted_since_refresh` from the snapshot's
available bytes so a burst arriving faster than the 5 s refresh cannot
over-allocate. That counter is GROSS: it only rolls over on the refresh and
never rolls back when a fill is later evicted, cancelled, or loses the
invalidation race. Under sustained high-throughput churn (net footprint flat
and far below `max_capacity`) the raw counter balloons past the memory the
cache actually holds, so `effective_available` collapses and the gate reports
false memory pressure — skipping the hottest fills with SkippedMemoryPressure
until the next 5 s refresh. This only lowers hit rate; it never returns wrong
data and self-heals each refresh.

Fix direction 1 (minimal regression): cap the reservation deduction at the
cache's own growth headroom (`max_capacity - weighted_size()`) instead of
letting the unbounded gross counter shrink the system-available budget. The
cache can never hold more than `max_capacity`, so a burst adds at most that
headroom of real memory before moka evicts to stay bounded (net-zero churn
beyond that point) — capping the deduction there keeps the reservation honest
without treating gross churn as growth. Chosen over net-accounting (direction
2, releasing bytes on every failure/cancel/eviction path) because that only
plugs the leak on failed fills and would not address the core defect: churn of
*successful* insert/evict fills over the 5 s window. It also touches only the
gate plus one call site rather than every failure path in moka_backend.

The cap only ever raises `effective_available`, so real memory pressure (a low
snapshot at refresh) still suppresses fills; when the cache is at capacity the
headroom is 0 and the deduction vanishes, correctly reflecting net-zero churn.
`MokaBackend` now stores `max_capacity` and passes the live headroom into
`allows_fill`. Adds targeted gate tests: gross churn far above headroom no
longer falsely suppresses, yet the reservation still bounds a burst while the
cache can genuinely grow.

Refs rustfs/backlog#1212
Co-Authored-By: heihutu <heihutu@gmail.com>

* test(ecstore): assert native O_DIRECT path runs in uring read test

uring_preserves_o_direct_for_eligible_reads only compared bytes through
LocalDisk::read_file_mmap_copy. On a filesystem that rejects O_DIRECT the
read silently degrades to the buffered StdBackend fallback and the byte
check still passes, so the test could go green without the native
read_at_direct path ever executing -- a vacuous pass.

Add a per-disk native_direct_reads counter on UringBackend, incremented
only when pread_uring_direct completes, and rebuild the test to drive a
real UringBackend's pread_bytes and assert the counter is non-zero (every
eligible read went through the native tier). When io_uring or O_DIRECT is
unavailable on the host filesystem (restricted CI runners, tmpfs), the
test skips loudly via eprintln instead of asserting a tautology, while
still checking byte-correctness on whatever tier served the read.

The counter also gives a gray release a positive signal that the O_DIRECT
tier is serving reads, not just a fallback count.

Refs rustfs/backlog#1213
Co-Authored-By: heihutu <heihutu@gmail.com>

* fix(ecstore): warn + count read-time EINVAL on native O_DIRECT reads

classify_direct_read_error is only reached from the read side: the
O_DIRECT open in pread_uring_direct already succeeded (an open-time
refusal is handled earlier as DirectOpenError::ODirectRefused). So an
EINVAL/EOPNOTSUPP arriving here is a read-time error on an fd the kernel
accepted for O_DIRECT -- far more likely an alignment bug in the aligned
read path than an unsupported filesystem. The old code latched the disk's
native path off with only a once-per-disk debug trace, hiding a potential
correctness regression behind a silent buffered-read downgrade.

Diagnostics only: the fallback behaviour is unchanged (the native path is
still latched off and the caller still reads via StdBackend). This adds a
rustfs_io_uring_direct_read_einval_total counter and promotes the
once-per-disk trace from debug to warn so an operator can see an alignment
regression instead of an unexplained latency/CPU shift.

Refs rustfs/backlog#1214
Co-Authored-By: heihutu <heihutu@gmail.com>

* docs(ecstore): document data-blocks-first default and its tail-latency cost

DEFAULT_RUSTFS_GET_DATA_BLOCKS_FIRST_READER_SETUP is true and must stay
true: deferred-parity is the deliberate, already-rolled-out full-object
GET default from backlog#1159/#923. Flipping it back to false in code
would silently revert that rollout for every deployment that has not set
the env var, so this commit only documents -- no behaviour change.

The added notes explain what data-blocks-first does (schedule data shards
up front, engage parity lazily on a missing/corrupt data shard), the known
trade-off (parity is engaged late, so a slow-but-not-dead data drive
raises GET p99 because the faster parity shards are not raced against it
until a data shard is declared missing), and the operational rollback
switch (RUSTFS_GET_DATA_BLOCKS_FIRST_READER_SETUP=false), which is
intentionally an env override rather than a code default change.

No metric was added: the low-risk observability hook for "slow data drive
engaged deferred parity" would live at the deferred-stripe engage point,
which is out of this file's scope; this change stays documentation-only to
avoid touching the hot GET path.

Refs rustfs/backlog#1215
Co-Authored-By: heihutu <heihutu@gmail.com>

* docs(ecstore): document wide-directory walk stall hazard and tuning

list_dir enumerates a whole directory in one os::read_dir call (count =
-1), and the walk caller bounds that entire enumeration with the per-read
stall budget (default 5s) as if it were a single read. For a wide, flat
prefix -- one directory holding millions of immediate children -- a single
readdir can exceed the budget on a healthy disk, trip DiskError::Timeout,
and surface as a ListObjects 500 quorum failure though the drive is fine
(a #2999 sub-class).

This is documented, not rewritten: turning the one-shot readdir into a
streaming/batched enumeration that refreshes the stall deadline between
chunks is an architecture-level change with high regression surface
(ordering, the count contract, quorum merge) and belongs in a separate
follow-up. The supported mitigation today is operational, so the comments
point wide-directory deployments at RUSTFS_DRIVE_WALKDIR_STALL_TIMEOUT_SECS
and the high-latency drive-timeout profile, which widen the budget with no
code change. Notes were added at list_dir, the scan_dir call site, and
get_drive_walkdir_stall_timeout. No behaviour change.

Refs rustfs/backlog#1216
Co-Authored-By: heihutu <heihutu@gmail.com>

* docs(ecstore): document consumer-peek vs producer-stall coupling

In list_path_raw the consumer's peek_timeout is drawn from the same source
and same value (walkdir_stall_timeout, default 5s) as the producer-side
walk stall budget, but the two measure different things: the producer
stall bounds a single drive read, while the consumer peek bounds the gap
between two ADJACENT entries arriving from a reader. Because they share a
value, the consumer cannot wait meaningfully longer for the next entry
than the producer is allowed to spend producing one. Walking a region
dense with non-listable internal items can make a HEALTHY drive miss the
budget between visible entries; the consumer then declares it stalled and
detaches it, dropping a good drive from the merge and capping the "large
prefix succeeds" guarantee.

Documented, not decoupled: giving the consumer peek an independent,
strictly-larger budget would cut these false detaches but equally delays
detaching a genuinely dead drive and shifts listing tail-latency
semantics, so it wants soak data before changing the default. The comment
records the invariant any such follow-up must keep -- consumer peek >=
producer stall, never stricter -- so it can never fail a drive before the
producer would. No behaviour change.

Refs rustfs/backlog#1217
Co-Authored-By: heihutu <heihutu@gmail.com>

* fix(io-metrics): add time-based trigger for low-IOPS latency percentiles

Percentiles were recomputed only every 128 IOs and seeded to 0, so a
low-traffic deployment exported p95/p99 = 0/stale for a long time after
startup. Add a 10s wall-clock trigger alongside the count throttle so the
first recompute can fire before 128 samples accrue. Hot-path per-op mean
update is unchanged.

Refs rustfs/backlog#1218
Co-Authored-By: heihutu <heihutu@gmail.com>

* test(e2e): cover codec-streaming parity under fault injection and NoSuchKey

The codec-streaming compat A/B previously ran only against a healthy
4-disk EC set with successful full GETs: the DiskFaultHarness was
constructed but never faulted, the error path was untested, and the
range assertion silently compared legacy-vs-legacy (ranges always fall
back to the duplex path), overstating what it proved.

Add two genuinely-failable scenarios reusing the existing harness and
fixtures:

- Parity reconstruction A/B: take one data disk offline and re-run the
  full object matrix on both phases while the EC 2+2 set rebuilds each
  large object from the surviving shards. Assert codec == legacy
  byte-for-byte (sha256) and header-for-header, and assert the codec
  phase served the reconstructed objects with zero duplex-pipe fallback
  (the reader gate is drive-health-independent, so the codec fast path
  is really exercised through reconstruction).
- NoSuchKey negative path: compare the HTTP status + S3 error code of a
  missing-key GET across the legacy and codec phases and require them to
  be identical (404/NoSuchKey), guarding against the codec env
  perturbing the error path.

Also clarify the range-phase comment so it is not misread as
codec-range correctness coverage: both sides are served by the same
legacy range path, so the assertion only proves ranges keep working and
keep falling back to legacy with the gates open.

Verified: cargo check/--no-run pass and the test passes locally
(1 passed; dup_codec=0 confirms the codec path ran).

Refs rustfs/backlog#1219
Co-Authored-By: heihutu <heihutu@gmail.com>

* ci(ecstore): exercise native O_DIRECT read path on an ext4 loopback

The uring-integration leg ran on the runner's default TMPDIR, which may sit
on tmpfs/overlayfs where open(O_DIRECT) fails and the native read_at_direct
path silently latches off to the aligned StdBackend fallback. Mount a
dedicated ext4 loopback and point TMPDIR at it so the real io_uring dep
(bumped git->0.1.0->0.2.0->0.2.1) and the native O_DIRECT read path are
actually covered rather than validated only by signature diffing.

Refs rustfs/backlog#1220
Co-Authored-By: heihutu <heihutu@gmail.com>

---------

Co-authored-by: heihutu <heihutu@gmail.com>
This commit is contained in:
houseme
2026-07-13 00:03:28 +08:00
committed by GitHub
parent 63b4568f85
commit 1553dc3f62
10 changed files with 663 additions and 75 deletions
+82 -7
View File
@@ -262,7 +262,12 @@ impl ObjectDataCacheMemoryGate {
///
/// This is lock-free and does no blocking sysinfo read: it only reads the
/// atomic snapshot maintained by the periodic refresher.
pub fn allows_fill(&self, required_bytes: u64) -> bool {
///
/// `cache_growth_headroom` is how many more bytes the cache itself can hold
/// before it is at capacity (`max_capacity - weighted_size()`). It caps how
/// far the in-window reservation may shrink the budget: see the reservation
/// note below.
pub fn allows_fill(&self, required_bytes: u64, cache_growth_headroom: u64) -> bool {
// A zero floor opts out of the gate, so fill admission never depends on
// a live memory reading — which differs between a host and a container.
// This must short-circuit before any snapshot read (see 51a97a81c).
@@ -280,7 +285,20 @@ impl ObjectDataCacheMemoryGate {
// burst that arrives faster than the 5 s refresh: each admission shrinks
// the budget the next one sees, so cumulative admission cannot exceed the
// real headroom even though every fill reads the same (stale) snapshot.
let effective_available = snapshot.available_bytes.saturating_sub(self.snapshot.admitted());
//
// `admitted_since_refresh` counts GROSS admitted bytes and never rolls
// back on a fill that is later evicted, cancelled, or loses the
// invalidation race; it only resets on the 5 s refresh. Under sustained
// churn (net footprint flat, far below capacity) the raw counter would
// balloon past the real memory the cache consumes and falsely trip the
// gate, skipping the hottest fills until the next refresh. The cache can
// never hold more than `max_capacity`, so a burst adds at most
// `cache_growth_headroom` bytes of real memory before moka evicts to stay
// bounded (net-zero churn beyond that point). Capping the deduction there
// keeps the reservation honest without treating gross churn as growth
// (backlog#1212).
let reserved = self.snapshot.admitted().min(cache_growth_headroom);
let effective_available = snapshot.available_bytes.saturating_sub(reserved);
let min_free = u64::from(self.min_free_memory_percent);
let has_percent_budget = effective_available.saturating_mul(100) >= snapshot.total_bytes.saturating_mul(min_free);
@@ -383,7 +401,7 @@ mod tests {
available_bytes: 16 * 1024 * 1024,
}));
assert!(!gate.allows_fill(1024));
assert!(!gate.allows_fill(1024, u64::MAX));
assert_eq!(stats.snapshot().memory_pressure_events, 1);
}
@@ -396,7 +414,7 @@ mod tests {
available_bytes: 500,
}));
assert!(gate.allows_fill(100));
assert!(gate.allows_fill(100, u64::MAX));
}
#[test]
@@ -416,7 +434,7 @@ mod tests {
available_bytes: 1,
}));
assert!(gate.allows_fill(512));
assert!(gate.allows_fill(512, u64::MAX));
assert_eq!(stats.snapshot().memory_pressure_events, 0);
}
@@ -429,10 +447,67 @@ mod tests {
available_bytes: 100,
}));
assert!(!gate.allows_fill(128));
assert!(!gate.allows_fill(128, u64::MAX));
assert_eq!(stats.snapshot().memory_pressure_events, 1);
}
// backlog#1212: `admitted_since_refresh` counts GROSS admitted bytes and
// never rolls back on an evicted/cancelled/lost-race fill, so a churn window
// (net footprint flat, far below capacity) balloons the raw counter past the
// memory the cache actually holds. Deducting it wholesale falsely trips the
// gate; capping the deduction at the cache's growth headroom fixes it.
#[test]
fn reservation_deduction_capped_at_growth_headroom_avoids_false_pressure() {
let stats = Arc::new(ObjectDataCacheStats::default());
let gate = ObjectDataCacheMemoryGate::new(&ObjectDataCacheConfig::default(), Arc::clone(&stats));
gate.set_test_snapshot(Some(ObjectDataCacheMemorySnapshot {
total_bytes: 1_000_000,
available_bytes: 500_000, // 50% free, well above the 20% floor
}));
// A churn window admitted far more GROSS bytes than the cache can hold;
// repeated insert/evict never rolled the counter back, so it now dwarfs
// the real available memory even though the live footprint stays tiny.
gate.snapshot.reserve(10_000_000);
// Uncapped, the raw gross counter swamps the budget and falsely signals
// memory pressure even though the cache's net footprint is flat.
assert!(
!gate.allows_fill(1_000, u64::MAX),
"raw gross admitted-bytes deduction should falsely suppress (the bug being fixed)"
);
// Capping the deduction at the cache's growth headroom (net size flat,
// far below capacity) restores admission: gross churn is no longer
// mistaken for real memory growth.
assert!(
gate.allows_fill(1_000, 100_000),
"capping the reservation at cache growth headroom must not falsely suppress"
);
}
// backlog#1212: capping the deduction must not defeat the reservation under
// genuine cache growth. When the cache still has room to grow, the in-window
// reservation must still shrink the budget so a burst cannot over-admit.
#[test]
fn reservation_still_bounds_burst_within_growth_headroom() {
let stats = Arc::new(ObjectDataCacheStats::default());
let gate = ObjectDataCacheMemoryGate::new(&ObjectDataCacheConfig::default(), Arc::clone(&stats));
gate.set_test_snapshot(Some(ObjectDataCacheMemorySnapshot {
total_bytes: 1_000_000,
available_bytes: 300_000, // 30% free; floor is 20% = 200_000
}));
// The cache can still grow well past the reserved amount, so the cap does
// not bind and the reservation is deducted in full.
gate.snapshot.reserve(150_000);
// 300_000 available - 150_000 reserved = 150_000 effective, below the
// 200_000 floor: the reservation must still suppress the fill.
assert!(
!gate.allows_fill(1_000, u64::MAX),
"reservation must still bound a burst while the cache can genuinely grow"
);
}
// ODC-14: `allows_fill` must read the atomic snapshot without performing an
// inline (blocking) refresh. A synchronous test has no tokio runtime, so no
// refresher task exists; if `allows_fill` refreshed inline it would clobber
@@ -450,7 +525,7 @@ mod tests {
gate.store_raw_snapshot_for_test(sentinel);
// Exercise the gate on the atomic path (no test_override installed).
let _ = gate.allows_fill(1);
let _ = gate.allows_fill(1, u64::MAX);
let after = gate.raw_snapshot_for_test();
assert_eq!(
+15 -1
View File
@@ -32,6 +32,10 @@ pub struct MokaBackend {
index: Arc<ObjectDataCacheIdentityIndex>,
singleflight: ObjectDataCacheSingleflight,
memory_gate: ObjectDataCacheMemoryGate,
/// Cache capacity in weighted bytes. Used to derive how much the cache can
/// still grow (`max_capacity - weighted_size()`), which caps the memory
/// gate's in-window reservation deduction (backlog#1212).
max_capacity: u64,
/// Bounds the number of concurrent distinct-key fills. Singleflight only
/// dedups per key, so without this limiter distinct-key fills are unbounded.
fill_semaphore: Arc<Semaphore>,
@@ -144,6 +148,7 @@ impl MokaBackend {
index,
singleflight: ObjectDataCacheSingleflight::new(Arc::clone(&stats)),
memory_gate: ObjectDataCacheMemoryGate::new(config, stats),
max_capacity,
fill_semaphore: Arc::new(Semaphore::new(fill_permits)),
#[cfg(test)]
fill_barrier: std::sync::Mutex::new(None),
@@ -204,7 +209,16 @@ impl MokaBackend {
Err(_) => return leader.finish(ObjectDataCacheFillResult::SkippedFillConcurrency),
};
if !self.memory_gate.allows_fill(u64::try_from(bytes.len()).unwrap_or(u64::MAX)) {
// How much the cache can still grow before it is at capacity. This caps
// the gate's in-window reservation so sustained gross churn (net size
// flat, far below capacity) cannot be mistaken for real memory growth
// and falsely skip fills (backlog#1212). weighted_size() is moka's
// lazily-maintained approximation, which is all this bound needs.
let cache_growth_headroom = self.max_capacity.saturating_sub(self.cache.weighted_size());
if !self
.memory_gate
.allows_fill(u64::try_from(bytes.len()).unwrap_or(u64::MAX), cache_growth_headroom)
{
return leader.finish(ObjectDataCacheFillResult::SkippedMemoryPressure);
}