mirror of
https://github.com/rustfs/rustfs.git
synced 2026-07-28 00:58:59 +00:00
05f52beea3ea9c7e514baddfe111e48055feb305
14 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
1553dc3f62 |
Address P2 follow-ups from the 2026-07-10..12 merged-PR review (backlog#1210-1220) (#4783)
* fix(obs): open cleaner compression source with O_NOFOLLOW The compressor opened the source log via File::open, which follows a symlink at the final path component. Between the scanner selecting a regular file and this open, an attacker with write access to the log directory could swap the entry for a symlink (TOCTOU) pointing at, say, /etc/shadow, whose contents would then be copied into an archive. Open the source with O_NOFOLLOW on Unix so such a swap fails with ELOOP; the temp/archive path already refused symlinks, this closes the source side. Refs rustfs/backlog#1210 Co-Authored-By: heihutu <heihutu@gmail.com> * fix(obs): recompress instead of trusting leftover cleaner archives archive_header_ok only checked the first 2-4 magic bytes before treating an existing .gz/.zst as a completed prior result and letting the caller delete the source log. A file with valid magic but a truncated or forged body passes that check, so an attacker with write access to the log directory (or a crashed prior run) could plant such a stub and make the cleaner delete the real log without ever producing a usable archive — silent audit-data loss. Chosen fix: stop trusting cross-process leftovers entirely and always recompress the source in this pass, rather than fully decoding every leftover to validate it. Full-decode validation would add real CPU cost and decode-bug surface for a rare crash-recovery case; the existing atomic create_new+rename already overwrites whatever sits at the archive path (a planted symlink is replaced, never followed) with a freshly written, fsync'd archive, so a partial/forged leftover can never gate source deletion. This is the lowest-regression option. Refs rustfs/backlog#1211 Co-Authored-By: heihutu <heihutu@gmail.com> * fix(object-data-cache): cap memory-gate reservation at cache growth headroom The memory gate subtracts `admitted_since_refresh` from the snapshot's available bytes so a burst arriving faster than the 5 s refresh cannot over-allocate. That counter is GROSS: it only rolls over on the refresh and never rolls back when a fill is later evicted, cancelled, or loses the invalidation race. Under sustained high-throughput churn (net footprint flat and far below `max_capacity`) the raw counter balloons past the memory the cache actually holds, so `effective_available` collapses and the gate reports false memory pressure — skipping the hottest fills with SkippedMemoryPressure until the next 5 s refresh. This only lowers hit rate; it never returns wrong data and self-heals each refresh. Fix direction 1 (minimal regression): cap the reservation deduction at the cache's own growth headroom (`max_capacity - weighted_size()`) instead of letting the unbounded gross counter shrink the system-available budget. The cache can never hold more than `max_capacity`, so a burst adds at most that headroom of real memory before moka evicts to stay bounded (net-zero churn beyond that point) — capping the deduction there keeps the reservation honest without treating gross churn as growth. Chosen over net-accounting (direction 2, releasing bytes on every failure/cancel/eviction path) because that only plugs the leak on failed fills and would not address the core defect: churn of *successful* insert/evict fills over the 5 s window. It also touches only the gate plus one call site rather than every failure path in moka_backend. The cap only ever raises `effective_available`, so real memory pressure (a low snapshot at refresh) still suppresses fills; when the cache is at capacity the headroom is 0 and the deduction vanishes, correctly reflecting net-zero churn. `MokaBackend` now stores `max_capacity` and passes the live headroom into `allows_fill`. Adds targeted gate tests: gross churn far above headroom no longer falsely suppresses, yet the reservation still bounds a burst while the cache can genuinely grow. Refs rustfs/backlog#1212 Co-Authored-By: heihutu <heihutu@gmail.com> * test(ecstore): assert native O_DIRECT path runs in uring read test uring_preserves_o_direct_for_eligible_reads only compared bytes through LocalDisk::read_file_mmap_copy. On a filesystem that rejects O_DIRECT the read silently degrades to the buffered StdBackend fallback and the byte check still passes, so the test could go green without the native read_at_direct path ever executing -- a vacuous pass. Add a per-disk native_direct_reads counter on UringBackend, incremented only when pread_uring_direct completes, and rebuild the test to drive a real UringBackend's pread_bytes and assert the counter is non-zero (every eligible read went through the native tier). When io_uring or O_DIRECT is unavailable on the host filesystem (restricted CI runners, tmpfs), the test skips loudly via eprintln instead of asserting a tautology, while still checking byte-correctness on whatever tier served the read. The counter also gives a gray release a positive signal that the O_DIRECT tier is serving reads, not just a fallback count. Refs rustfs/backlog#1213 Co-Authored-By: heihutu <heihutu@gmail.com> * fix(ecstore): warn + count read-time EINVAL on native O_DIRECT reads classify_direct_read_error is only reached from the read side: the O_DIRECT open in pread_uring_direct already succeeded (an open-time refusal is handled earlier as DirectOpenError::ODirectRefused). So an EINVAL/EOPNOTSUPP arriving here is a read-time error on an fd the kernel accepted for O_DIRECT -- far more likely an alignment bug in the aligned read path than an unsupported filesystem. The old code latched the disk's native path off with only a once-per-disk debug trace, hiding a potential correctness regression behind a silent buffered-read downgrade. Diagnostics only: the fallback behaviour is unchanged (the native path is still latched off and the caller still reads via StdBackend). This adds a rustfs_io_uring_direct_read_einval_total counter and promotes the once-per-disk trace from debug to warn so an operator can see an alignment regression instead of an unexplained latency/CPU shift. Refs rustfs/backlog#1214 Co-Authored-By: heihutu <heihutu@gmail.com> * docs(ecstore): document data-blocks-first default and its tail-latency cost DEFAULT_RUSTFS_GET_DATA_BLOCKS_FIRST_READER_SETUP is true and must stay true: deferred-parity is the deliberate, already-rolled-out full-object GET default from backlog#1159/#923. Flipping it back to false in code would silently revert that rollout for every deployment that has not set the env var, so this commit only documents -- no behaviour change. The added notes explain what data-blocks-first does (schedule data shards up front, engage parity lazily on a missing/corrupt data shard), the known trade-off (parity is engaged late, so a slow-but-not-dead data drive raises GET p99 because the faster parity shards are not raced against it until a data shard is declared missing), and the operational rollback switch (RUSTFS_GET_DATA_BLOCKS_FIRST_READER_SETUP=false), which is intentionally an env override rather than a code default change. No metric was added: the low-risk observability hook for "slow data drive engaged deferred parity" would live at the deferred-stripe engage point, which is out of this file's scope; this change stays documentation-only to avoid touching the hot GET path. Refs rustfs/backlog#1215 Co-Authored-By: heihutu <heihutu@gmail.com> * docs(ecstore): document wide-directory walk stall hazard and tuning list_dir enumerates a whole directory in one os::read_dir call (count = -1), and the walk caller bounds that entire enumeration with the per-read stall budget (default 5s) as if it were a single read. For a wide, flat prefix -- one directory holding millions of immediate children -- a single readdir can exceed the budget on a healthy disk, trip DiskError::Timeout, and surface as a ListObjects 500 quorum failure though the drive is fine (a #2999 sub-class). This is documented, not rewritten: turning the one-shot readdir into a streaming/batched enumeration that refreshes the stall deadline between chunks is an architecture-level change with high regression surface (ordering, the count contract, quorum merge) and belongs in a separate follow-up. The supported mitigation today is operational, so the comments point wide-directory deployments at RUSTFS_DRIVE_WALKDIR_STALL_TIMEOUT_SECS and the high-latency drive-timeout profile, which widen the budget with no code change. Notes were added at list_dir, the scan_dir call site, and get_drive_walkdir_stall_timeout. No behaviour change. Refs rustfs/backlog#1216 Co-Authored-By: heihutu <heihutu@gmail.com> * docs(ecstore): document consumer-peek vs producer-stall coupling In list_path_raw the consumer's peek_timeout is drawn from the same source and same value (walkdir_stall_timeout, default 5s) as the producer-side walk stall budget, but the two measure different things: the producer stall bounds a single drive read, while the consumer peek bounds the gap between two ADJACENT entries arriving from a reader. Because they share a value, the consumer cannot wait meaningfully longer for the next entry than the producer is allowed to spend producing one. Walking a region dense with non-listable internal items can make a HEALTHY drive miss the budget between visible entries; the consumer then declares it stalled and detaches it, dropping a good drive from the merge and capping the "large prefix succeeds" guarantee. Documented, not decoupled: giving the consumer peek an independent, strictly-larger budget would cut these false detaches but equally delays detaching a genuinely dead drive and shifts listing tail-latency semantics, so it wants soak data before changing the default. The comment records the invariant any such follow-up must keep -- consumer peek >= producer stall, never stricter -- so it can never fail a drive before the producer would. No behaviour change. Refs rustfs/backlog#1217 Co-Authored-By: heihutu <heihutu@gmail.com> * fix(io-metrics): add time-based trigger for low-IOPS latency percentiles Percentiles were recomputed only every 128 IOs and seeded to 0, so a low-traffic deployment exported p95/p99 = 0/stale for a long time after startup. Add a 10s wall-clock trigger alongside the count throttle so the first recompute can fire before 128 samples accrue. Hot-path per-op mean update is unchanged. Refs rustfs/backlog#1218 Co-Authored-By: heihutu <heihutu@gmail.com> * test(e2e): cover codec-streaming parity under fault injection and NoSuchKey The codec-streaming compat A/B previously ran only against a healthy 4-disk EC set with successful full GETs: the DiskFaultHarness was constructed but never faulted, the error path was untested, and the range assertion silently compared legacy-vs-legacy (ranges always fall back to the duplex path), overstating what it proved. Add two genuinely-failable scenarios reusing the existing harness and fixtures: - Parity reconstruction A/B: take one data disk offline and re-run the full object matrix on both phases while the EC 2+2 set rebuilds each large object from the surviving shards. Assert codec == legacy byte-for-byte (sha256) and header-for-header, and assert the codec phase served the reconstructed objects with zero duplex-pipe fallback (the reader gate is drive-health-independent, so the codec fast path is really exercised through reconstruction). - NoSuchKey negative path: compare the HTTP status + S3 error code of a missing-key GET across the legacy and codec phases and require them to be identical (404/NoSuchKey), guarding against the codec env perturbing the error path. Also clarify the range-phase comment so it is not misread as codec-range correctness coverage: both sides are served by the same legacy range path, so the assertion only proves ranges keep working and keep falling back to legacy with the gates open. Verified: cargo check/--no-run pass and the test passes locally (1 passed; dup_codec=0 confirms the codec path ran). Refs rustfs/backlog#1219 Co-Authored-By: heihutu <heihutu@gmail.com> * ci(ecstore): exercise native O_DIRECT read path on an ext4 loopback The uring-integration leg ran on the runner's default TMPDIR, which may sit on tmpfs/overlayfs where open(O_DIRECT) fails and the native read_at_direct path silently latches off to the aligned StdBackend fallback. Mount a dedicated ext4 loopback and point TMPDIR at it so the real io_uring dep (bumped git->0.1.0->0.2.0->0.2.1) and the native O_DIRECT read path are actually covered rather than validated only by signature diffing. Refs rustfs/backlog#1220 Co-Authored-By: heihutu <heihutu@gmail.com> --------- Co-authored-by: heihutu <heihutu@gmail.com> |
||
|
|
3b139e5267 |
fix(obs/cleaner): harden log cleaner durability, symlink safety, and retention (audit OLC-01..14) (#4776)
* fix(obs/cleaner): fsync archive and log dir before deleting source logs OLC-01: the compression path flushed the BufWriter but never synced the archive data or the parent directory before renaming, and the source was then unlinked with no durability barrier. A crash after rename but before the page cache reached disk could leave a truncated/zero-length archive while the source was already gone — permanent log/audit data loss. Because the archive is always renamed to a brand-new name (guaranteed by the existing exists() guard), ext4 auto_da_alloc does not mask this. Hand the underlying File back from the writer closure, sync_all() it before rename, fsync the parent directory, and fsync the log directory after the unlinks so a delete cannot be reordered ahead of the archive it justified. Guard the temp file with an RAII cleanup so an early return or panic cannot leak a *.tmp orphan. Ref: rustfs/backlog#1194 (audit rustfs/backlog#1193) Co-Authored-By: heihutu <heihutu@gmail.com> * fix(obs/cleaner): create compression temp file with O_EXCL and O_NOFOLLOW OLC-03: the temp archive was opened with File::create at a predictable `<source>.gz.tmp` path with no O_EXCL/O_NOFOLLOW, so an actor with write access to the log directory could pre-plant that path as a symlink and have the compressor follow it — truncating and overwriting an arbitrary external file, then chmod-ing it to the source log's mode. This mirrors the symlink refusal already enforced on the deletion path (secure_delete). Route temp creation through create_tmp_archive(), which uses create_new (O_CREAT|O_EXCL) to refuse a pre-existing entry and, on Unix, O_NOFOLLOW to refuse a symlink at the final path component. Ref: rustfs/backlog#1196 (audit rustfs/backlog#1193) Co-Authored-By: heihutu <heihutu@gmail.com> * fix(obs/cleaner): create compression temp file with restrictive mode OLC-06: File::create left the temp archive world-readable (0644 & ~umask) for the entire duration of compressing a large log, exposing the full plaintext of a possibly-0600 audit log on shared hosts until the mode was copied only after the write completed. Pass the source mode into create_tmp_archive and open the temp file with it (default 0600) so it is restrictive from creation; the post-write chmod still tightens/matches the source mode exactly. Ref: rustfs/backlog#1199 (audit rustfs/backlog#1193) Co-Authored-By: heihutu <heihutu@gmail.com> * fix(obs/cleaner): validate existing archive before skipping recompression OLC-02: the idempotency guard used Path::exists() (which follows symlinks) and trusted whatever it found, then the caller deleted the source. A planted `<archive>.gz` symlink, or a zero-length/truncated archive left by a crashed run (OLC-01), would green-light deleting the source with no valid backup — data loss / log destruction. Replace exists() with symlink_metadata (no follow) and only treat the entry as a completed prior result when it is a regular, non-empty file whose header matches the codec magic (gzip 1f 8b / zstd 28 b5 2f fd). Anything else falls through to recompression, whose atomic create_new+rename replaces the bad entry (a symlink is replaced, never followed or deleted through). Ref: rustfs/backlog#1195 (audit rustfs/backlog#1193) Co-Authored-By: heihutu <heihutu@gmail.com> * fix(obs/cleaner): stop dry-run overstating reclaimed bytes for compression OLC-08: in dry-run, compress_with_writer returned output_bytes = 0, so projected_freed_bytes = input and delete_files reported the full input as freed. A real run keeps the archive on disk (freed = input - archive), so dry-run overstated reclaim by the whole archive footprint. Estimate the archive with a deliberately conservative ratio so the projection never exceeds what a real run reclaims. Ref: rustfs/backlog#1201 (audit rustfs/backlog#1193) Co-Authored-By: heihutu <heihutu@gmail.com> * fix(obs/cleaner): make freed-byte accounting resilient; document steal metric OLC-12: input/output byte sizes were read via metadata().unwrap_or(0), which silently reports 0 on failure and skews freed-byte metrics (input - 0 = full input, overstating reclaim). Use the copy() byte count as the authoritative input size and, when the archive metadata read fails, conservatively assume no savings instead of 0. Also document that the steal_success_rate counts only victim steals (batch = one success), so it reads as a relative rebalancing signal, not absolute task acquisition. Ref: rustfs/backlog#1205 (audit rustfs/backlog#1193) Co-Authored-By: heihutu <heihutu@gmail.com> * fix(obs/cleaner): preserve level-0 semantics, allow zstd 22, log effective levels OLC-09: build() and the codec calls clamped gzip/zstd levels to [1,9]/[1,21], silently rewriting gzip level 0 (store) and zstd level 0 (codec default) to 1 and blocking the legal zstd maximum of 22. Clamp to [0,9]/[0,22] so those meanings survive, and echo the effective (post-clamp) levels in the startup log via new effective_gzip_level()/effective_zstd_level() getters so the log matches what actually runs. Ref: rustfs/backlog#1202 (audit rustfs/backlog#1193) Co-Authored-By: heihutu <heihutu@gmail.com> * fix(obs/cleaner): bound compressed archives by byte cap; warn on retention=0 OLC-04: archive expiry was gated on compressed_file_retention_days > 0, so retention=0 disabled it entirely while compression kept producing archives, and max_total_size_bytes only ever bounded uncompressed logs — unbounded disk growth. Replace select_expired_compressed with select_archives_to_delete, which applies age expiry (when retention is on) and, regardless of retention, trims the oldest archives until the set fits under max_total_size_bytes. Also warn at startup when compression is on with retention=0 so the "keep forever" semantics are not mistaken for "delete immediately". Ref: rustfs/backlog#1197 (audit rustfs/backlog#1193) Co-Authored-By: heihutu <heihutu@gmail.com> * fix(obs/cleaner): warn on invalid exclude glob instead of dropping silently OLC-05: build() dropped unparseable exclude globs via filter_map(...ok()), so a typo (or a literal comma splitting a char-class in the config string) turned "protect this file" into "delete this file" with no signal. Log a warning per rejected pattern with the raw string and parse error so the misconfiguration is visible. Ref: rustfs/backlog#1198 (audit rustfs/backlog#1193) Co-Authored-By: heihutu <heihutu@gmail.com> * perf(obs/cleaner): backoff idle workers, cap worker count, lower small-host floor OLC-11: the work-stealing loop re-spun on Steal::Retry with no yield and used yield_now on the empty path, burning CPU during redistribution windows; worker_count had no upper bound so a mis-set parallel_workers over a directory of thousands of logs could spawn thousands of threads; and default_parallel_workers forced >=4 workers even on 1-2 vCPU hosts. Use crossbeam_utils::Backoff (spin->yield, reset on work) on the idle paths, clamp worker_count to MAX_PARALLEL_COMPRESS_WORKERS, and lower the default floor to 1. Ref: rustfs/backlog#1204 (audit rustfs/backlog#1193) Co-Authored-By: heihutu <heihutu@gmail.com> * fix(obs/cleaner): warn when active-file guard is disabled by empty filename OLC-13 (defense-in-depth): the scanner protects the live log purely by exact filename equality against active_filename. An empty active_filename silently disables that protection, so a non-empty file_pattern could make the live log a deletion candidate via the public builder. Warn in build() when that unsafe combination is configured. The audit's "never delete the newest match" structural guard is intentionally not implemented: it would conflict with the legitimate keep_files=0 semantics (purge all rotated logs). The naming contract is instead locked by regression tests (OLC-14). Ref: rustfs/backlog#1206 (audit rustfs/backlog#1193) Co-Authored-By: heihutu <heihutu@gmail.com> * fix(obs/cleaner): warn on unknown algorithm/match_mode, echo match_mode OLC-07: from_config_str silently fell back to defaults for unrecognized compression algorithm and match mode (any non-"prefix" value became Suffix), hiding operator typos like "prefixx" that could make the cleaner match no rotated logs. Warn on a non-empty unrecognized value in both parsers, and echo the resolved match_mode in the startup log alongside the algorithm. Ref: rustfs/backlog#1200 (audit rustfs/backlog#1193) Co-Authored-By: heihutu <heihutu@gmail.com> * fix(obs/cleaner): derive orphan .tmp suffixes and exempt them from min age OLC-10: orphan `*.gz.tmp`/`*.zst.tmp` cleanup was gated by min_file_age_seconds (default 3600), so crash-left orphans lingered up to an hour, and the tmp suffix list was hardcoded rather than derived from compressed_suffixes() — a new codec would leave `*.<ext>.tmp` orphans the scanner never recognizes. Derive the temp suffix from CompressionAlgorithm::compressed_suffixes(), and gate orphan removal on a small fixed grace window instead of min_file_age (orphans are never live-written after the rename that would promote them). Ref: rustfs/backlog#1203 (audit rustfs/backlog#1193) Co-Authored-By: heihutu <heihutu@gmail.com> * test(obs/cleaner): cover symlink, archive expiry, idempotency, and edge cases OLC-14: add regression tests for the previously-untested safety/correctness branches — symlink rejection (external target never deleted), archive age expiry vs fresh retention, archive byte-cap trim with retention disabled, gz/zst classification, max_single_file_size selection, min_age protecting a fresh non-empty log, active-file exclusion when the active name also matches the pattern, invalid exclude glob not aborting build, dry-run + compression creating no archive, gzip round-trip validity, and the idempotent-archive branch trusting a valid prior archive. Ref: rustfs/backlog#1207 (audit rustfs/backlog#1193) Co-Authored-By: heihutu <heihutu@gmail.com> * style(obs/cleaner): apply rustfmt and collapse nested if (clippy) Formatting-only cleanup over the audit fix series: rustfmt normalization of the multi-line expressions introduced in compress.rs/core.rs, plus collapsing the delete_files directory-fsync into a single let-chain to satisfy clippy::collapsible_if. No behavior change. Co-Authored-By: heihutu <heihutu@gmail.com> * refactor(obs/cleaner): collapse redundant source stats in compression path The per-issue fixes to compress_with_writer accumulated three metadata() syscalls on the source in the real compression path: an input_bytes read that was immediately shadowed by the copied byte count (a dead read), a source_mode read (OLC-06), and the pre-OLC-06 post-write chmod re-reading the same mode. Collapse to a single fd-based read — move the dry-run input_bytes read into the dry-run branch, read source_mode from the already-open fd (no path stat, no TOCTOU), and reuse it for the post-write chmod. Behavior is unchanged (same inode's mode, written for input size); 3 source stats -> 1. Co-Authored-By: heihutu <heihutu@gmail.com> * chore(obs/cleaner): fix typo flagged by CI (mis-set -> misconfigured) Co-Authored-By: heihutu <heihutu@gmail.com> --------- Co-authored-by: heihutu <heihutu@gmail.com> |
||
|
|
83bacd5b84 |
fix(obs): harden cleaner file matching (#4486)
Co-authored-by: heihutu <heihutu@gmail.com> |
||
|
|
87044b2378 |
fix(obs): tighten low-risk telemetry correctness (#4483)
Refs rustfs/backlog#1007 Refs rustfs/backlog#986 - respect explicit log-target directives when injecting noisy-crate suppressions - fix daily rotation day-boundary checks and add a short retry cooldown after failed rotations - preserve recorder metadata across label variants and serialize gauge updates before export - correct profiling target_env reporting, trace all=true parsing, cleaner accounting/docs, and legacy system interval rounding Co-authored-by: heihutu <heihutu@gmail.com> |
||
|
|
3ddade24f2 |
fix(obs): validate numeric env settings (#4474)
Co-authored-by: heihutu <heihutu@gmail.com> |
||
|
|
e0a46e940c |
fix(obs): drain cleaner results inside scope (#4429)
Co-authored-by: heihutu <heihutu@gmail.com> |
||
|
|
efa89a98ed |
refactor(logging): standardize protocol and observability events (#3419)
* refactor(logging): standardize object capacity events * refactor(logging): standardize protocol server events * refactor(logging): standardize swift protocol events * refactor(logging): standardize observability events * refactor(logging): move masking helper and extend guardrails |
||
|
|
bca8b08c2b |
fix: handle Windows paths in pre-commit tests (#2974)
Co-authored-by: houseme <housemecn@gmail.com> |
||
|
|
59f41eb86a | feat(obs): improve metrics coverage and dashboard performance (#2682) | ||
|
|
49366ee200 | chore(lint): clippy rules redundant_clone (#2554) | ||
|
|
94cdb89e29 |
feat(obs): add init_obs_with_config API and signature guard test (#2175)
Signed-off-by: houseme <housemecn@gmail.com> Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com> Co-authored-by: Copilot <198982749+Copilot@users.noreply.github.com> |
||
|
|
593a58c161 |
refactor(obs): optimize logging with custom RollingAppender and improved cleanup (#2151)
Signed-off-by: houseme <housemecn@gmail.com> Signed-off-by: heihutu <30542132+heihutu@users.noreply.github.com> Co-authored-by: Copilot <198982749+Copilot@users.noreply.github.com> Co-authored-by: houseme <4829346+houseme@users.noreply.github.com> Co-authored-by: heihutu <30542132+heihutu@users.noreply.github.com> Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com> |
||
|
|
2ac07c95a8 | refactor(obs): enhance log cleanup and rotation (#2040) | ||
|
|
c452f24487 |
Optimize log cleanup and rotation, update dependencies (#2032)
Co-authored-by: heihutu <heihutu@gmail.com> |