Zhengchao An
7db3882777
chore(obs): adjudicate 19 bare dead_code allows ( #6162 )
...
backlog#1823 step 10, batch 2. Eighteen of the nineteen suppress nothing and are deleted; one was real and keeps an allow that now says why.
Rotation::Never is constructed only by the rolling-appender tests at rolling.rs:456, 477 and 498, so the lib target reports it as never constructed. Its allow is restored with that reason.
Finding it corrected the method used for batch 1. Removing all nineteen and running cargo check -p rustfs-obs --tests reported zero warnings even after touching every source file, while clippy --lib --tests -D warnings caught Rotation::Never. cargo's warning output is not a reliable completeness check — it does not re-emit for cached compilations, and touching the sources did not cover the lib target here. Later batches should treat clippy -D warnings as the gate; batch 1's six crates were re-checked under clippy and are clean.
Taken with #6086 , which cleared this crate's 44 module-level blankets and left six real items, obs has now had 63 dead-code suppressions examined, of which seven were suppressing anything at all. The rest sat on items that are publicly reachable, where dead_code never applied — the same shape as the swift module and kms's dek.rs.
Verification: clippy --lib --tests -D warnings clean in the default, gpu and pyroscope lanes; cargo nextest run -p rustfs-obs 324 passed; make pre-commit exit 0.
Ref rustfs/backlog#1823 (step 10).
2026-08-17 11:32:23 +08:00
Zhengchao An
7c2b513613
chore(obs): drop 44 dead_code blankets from the metrics tree ( #6086 )
2026-08-14 08:10:49 +08:00
Henry Guo
4c5e73b2f2
fix(scanner): defer cycles during data movement ( #5970 )
...
* fix(scanner): defer cycles during data movement
* fix(scanner): distinguish deferred scan cycles
---------
Co-authored-by: Henry Guo <marshawcoco@users.noreply.github.com >
Co-authored-by: houseme <housemecn@gmail.com >
2026-08-12 20:46:06 +08:00
anthonymartin
706a8b6061
fix(scanner): publish bounded observational usage ( #5742 )
...
* fix(scanner): publish bounded observational usage
* test(ci): serialize embedded integration ports
* test(cache): isolate generation-change timeout
* fix(scanner): address observational usage review
Co-Authored-By: heihutu <heihutu@gmail.com >
---------
Co-authored-by: Anthony Martin <949506+anthonymartin@users.noreply.github.com >
Co-authored-by: houseme <housemecn@gmail.com >
Co-authored-by: heihutu <heihutu@gmail.com >
2026-08-07 08:52:34 +08:00
houseme
035ce5d784
feat(obs): add bounded metrics dimensions ( #5645 )
...
* feat(obs): add drive topology detail metrics
Expose additive drive info, topology, state, and per-drive API metrics while preserving the existing drive metric label sets.
Backlog: rustfs/backlog#1655
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(obs): preserve suspect drive runtime state
Keep suspect as a bounded drive runtime state and avoid all-zero runtime_state samples for that storage health state.
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(obs): skip unknown drive inode samples
Avoid exporting zero inode gauges for missing or stale drive snapshots and ignore zero-count API latency buckets.
Co-Authored-By: heihutu <heihutu@gmail.com >
* feat(obs): add scanner source work detail metrics
Expose additive scanner source and cycle work metrics with bounded server/source/state labels while leaving the existing aggregate scanner metrics unchanged.
Co-Authored-By: heihutu <heihutu@gmail.com >
* feat(obs): add ilm action detail metrics
Expose additive ILM action/state task metrics with a server label while preserving the existing aggregate ILM series.
Co-Authored-By: heihutu <heihutu@gmail.com >
* feat(obs): add delivery target server metrics
Expose additive audit and notification delivery target metrics with server labels and extend removed-target tombstones for the server-aware series.
Co-Authored-By: heihutu <heihutu@gmail.com >
* feat(obs): add replication target flow metrics
Expose additive bucket replication target sent and failed-flow metrics while preserving existing bucket aggregates and target backlog series.
Co-Authored-By: heihutu <heihutu@gmail.com >
* feat(obs): add request server metrics
Expose additive API request metrics with server labels while preserving the existing request and traffic metric label sets.
Co-Authored-By: heihutu <heihutu@gmail.com >
* style(obs): apply rustfmt to metrics changes
Apply rustfmt output to the metrics dimension changes without altering behavior.
Co-Authored-By: heihutu <heihutu@gmail.com >
* style(obs): reuse audit target label constant
Use the exported audit target_id label constant for legacy audit target metrics.
Co-Authored-By: heihutu <heihutu@gmail.com >
* feat(obs): populate drive disk metrics
Co-Authored-By: heihutu <heihutu@gmail.com >
* feat(obs): add scanner bucket drive result metrics
Co-Authored-By: heihutu <heihutu@gmail.com >
* feat(obs): add replication proxy server metrics
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(obs): address metric liveness review
Use checked division for drive API latency aggregation and keep recovered drive, scanner current-cycle, replication flow, audit target, and notification target series from retaining stale values.
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(obs): address metric dimension review
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(obs): address additional metric review
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(obs): count drive calls at start
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(obs): address metrics dimension review
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(metrics): address dimension review gaps
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(metrics): address scanner review follow-ups
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(metrics): address runtime review follow-ups
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(metrics): reduce disk metric contention
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(metrics): address runtime review follow-ups
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(metrics): retire stale dimension series
Co-Authored-By: heihutu <heihutu@gmail.com >
---------
Co-authored-by: heihutu <heihutu@gmail.com >
2026-08-03 09:03:34 +08:00
houseme
fbec33bd29
Expose target-scoped durable MRF backlog metrics ( #5584 )
...
* feat(replication): expose target durable mrf backlog
Add target ARN attribution to durable MRF entries and surface target-scoped durable backlog metrics without changing existing bucket-only metric labels.
Keep legacy MRF files bucket-only by defaulting missing targetARNs to an empty list, and expose target snapshots through an additive API so existing DurableMrfBacklogSummary callers remain source-compatible.
Co-Authored-By: heihutu <heihutu@gmail.com >
* feat(replication): expose runtime target backlog (#5586 )
Track runtime replication backlog by target ARN for regular, large, delete, and MRF admission paths while preserving the existing bucket-level backlog semantics.
Add target-scoped current backlog metrics and merge them with durable target backlog snapshots for observability.
Co-authored-by: heihutu <heihutu@gmail.com >
---------
Co-authored-by: heihutu <heihutu@gmail.com >
2026-08-01 17:31:50 +00:00
houseme
da389c0e21
fix(replication): harden backlog observability ( #5564 )
...
Add RAII guards for replication runtime backlog tickets so active worker and queue counters unwind on every terminal path.
Expose node-local MRF pending, dropped, missed, and flush-failure metrics through the bucket replication Prometheus collector while keeping the existing current backlog and durable MRF gauges additive.
Update durable MRF summary maintenance to aggregate incrementally during the persister loop, avoiding repeated full-entry scans on each successful flush.
Co-authored-by: heihutu <heihutu@gmail.com >
Co-authored-by: zhi22915 <qiuzgang@gmail.com >
2026-08-01 11:12:58 +00:00
houseme
62d44d10b8
Expose replication backlog gauges ( #5557 )
...
* fix(replication): count backlog at queue admission
Co-Authored-By: heihutu <heihutu@gmail.com >
* feat(obs): expose bucket replication backlog gauges
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(obs): report recent backlog from queued work
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(obs): preserve legacy backlog metric semantics
Co-Authored-By: heihutu <heihutu@gmail.com >
* feat(obs): expose durable MRF backlog gauges
Co-Authored-By: heihutu <heihutu@gmail.com >
* test(obs): cover replication backlog metric scope
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(obs): keep backlog metrics API-compatible
Co-Authored-By: heihutu <heihutu@gmail.com >
* refactor(obs): streamline replication backlog metrics
Keep MRF backlog accounting and OBS metric collection on a single, cheaper path.
Co-Authored-By: heihutu <heihutu@gmail.com >
* test(kms): update aws capability snapshot
Co-Authored-By: heihutu <heihutu@gmail.com >
---------
Co-authored-by: heihutu <heihutu@gmail.com >
2026-08-01 08:48:45 +00:00
houseme
30dc04c94b
fix(obs): label node-local metrics by server ( #5465 )
...
Add stable server labels to node-local Prometheus metrics and OTLP resource attributes so dashboards can distinguish per-node CPU, memory, host network, and internode traffic series.
Co-authored-by: heihutu <heihutu@gmail.com >
2026-07-30 04:57:22 +00:00
Henry Guo
a63b79004c
fix(scanner): make distributed usage convergence authoritative ( #5151 )
...
* fix(scanner): make distributed usage cycles authoritative
* fix(scanner): close distributed refresh races
* fix(config): align scanner reload integration
* fix(admin): scope config test helpers
* fix(scanner): harden distributed usage convergence
* fix(scanner): preserve rolling activity compatibility
* fix(admin): expose non-secret optional config values
* fix(scanner): acknowledge distributed dirty usage
* fix(ecstore): make bucket mutations cancellation safe
* fix(scanner): preserve pending dirty acknowledgements
* test(obs): account for superseded scanner metric
* fix(api): reject excess detached bucket mutations
* test: close scanner convergence coverage gaps
* fix(scanner): make path tracking cleanup one-shot
---------
Co-authored-by: Henry Guo <marshawcoco@users.noreply.github.com >
Co-authored-by: houseme <housemecn@gmail.com >
2026-07-25 18:45:16 +08:00
Henry Guo
bb7bba3237
fix(obs): clarify cluster bucket usage metrics ( #5081 )
...
Co-authored-by: Henry Guo <marshawcoco@users.noreply.github.com >
2026-07-21 17:23:49 +08:00
escapecode
a80699b6dd
feat: add an opt-in NATS JetStream publish path for the notify and audit targets ( #4634 )
...
feat(targets): add an opt-in NATS JetStream publish path for the notify and audit targets
The NATS notify and audit targets publish through NATS Core, which returns
before the server has durably accepted the message. A broker restart or a
connection drop between the publish and the flush loses the event, even though
the send queue has already cleared it, and no acknowledgement gates that clear.
An opt-in JetStream publish path clears a queued event only after the server
returns a durable PublishAck, so delivery is at-least-once across a broker
restart or a reconnect. It applies to both the notify and audit NATS targets, is
off by default, and is byte-identical to the NATS Core path when disabled.
The path includes durable store-and-forward, a stable dedup id sent as the
Nats-Msg-Id header so a replayed event is collapsed by the stream duplicate
window, pre-flight stream validation, and a bounded failed-events store for
terminally-failed and retry-exhausted events. Three configuration keys per
target select it: JETSTREAM_ENABLE, JETSTREAM_STREAM_NAME, and
JETSTREAM_ACK_TIMEOUT_SECS, under the RUSTFS_NOTIFY_NATS_ and RUSTFS_AUDIT_NATS_
prefixes.
The on-disk batch filename separator changes from colon to underscore so
batch names are valid on Windows filesystems, with transparent read-back
of files written under the previous separator. The migration affects the
shared queue store for every target type and lands with this feature
because the store gains its first Windows-exercised paths here.
Co-authored-by: houseme <housemecn@gmail.com >
2026-07-14 15:36:14 +08:00
houseme
f0bf8cfe03
fix(obs): hide unwired request metrics ( #4481 )
...
Refs rustfs/backlog#1006
- aggregate request traffic samples per type to avoid future counter collisions
- keep request schema and collector crate-internal until a production stats source exists
- preserve focused regression coverage for the internal request collector logic
Co-authored-by: heihutu <heihutu@gmail.com >
2026-07-08 13:08:57 +00:00
houseme
1eb393cab1
fix(obs): align metrics schema and collector contracts ( #4476 )
...
fix(obs): align schema and collector contracts
Refs rustfs/backlog#1005
- align obs schema descriptors with emitted labels and metric types
- fix bucket traffic help text and TTFB bucket descriptor semantics
- add regression tests for drive, network host, bucket, and node bucket contracts
Co-authored-by: heihutu <heihutu@gmail.com >
2026-07-08 11:23:23 +00:00
houseme
a30a9c0aba
fix(obs): align cluster capacity semantics ( #4457 )
...
Co-authored-by: heihutu <heihutu@gmail.com >
2026-07-08 09:11:01 +00:00
houseme
021c955c21
fix(obs): report live replication backlog ( #4448 )
...
Co-authored-by: heihutu <heihutu@gmail.com >
2026-07-08 16:39:06 +08:00
houseme
acb1b765db
fix(obs): export resettable metrics as gauges ( #4432 )
...
Co-authored-by: heihutu <heihutu@gmail.com >
2026-07-08 16:23:11 +08:00
houseme
f968129945
fix(obs): stop exporting fake cpu categories ( #4439 )
...
Co-authored-by: heihutu <heihutu@gmail.com >
Co-authored-by: Zhengchao An <anzhengchao@gmail.com >
2026-07-08 16:05:11 +08:00
Zhengchao An
9a7255540b
fix(replication): add resync metrics ( #4408 )
2026-07-08 15:01:28 +08:00
wood
7484a61fa3
feat(compression): show cluster-level compression stat in grafana ( #4112 )
2026-06-30 19:05:05 +08:00
Zhengchao An
e1272f2aba
revert: restore #![allow(dead_code)] - CI clippy -D warnings conflict ( #3979 )
...
revert: restore #![allow(dead_code)] - clippy -D warnings treats warn as error
The #742 PR changed #![allow(dead_code)] to #![warn(dead_code)], but
CI runs clippy with -D warnings which turns warnings into errors.
This caused CI failures across multiple PRs.
Reverting to #![allow(dead_code)] until the dead code is actually
cleaned up. The 189 warnings in ecstore should be fixed incrementally
by deleting dead code and adding item-level allows, not by changing
the crate-level policy.
2026-06-28 08:32:34 +08:00
Zhengchao An
113058af54
chore: replace blanket #![allow(dead_code)] with #![warn(dead_code)] ( #742 ) ( #3974 )
2026-06-28 07:50:51 +08:00
Henry Guo
ad1a489f75
feat(scanner): add scanner budget progress controls ( #3185 )
...
Co-authored-by: Henry Guo <marshawcoco@users.noreply.github.com >
2026-06-03 14:37:58 +00:00
Henry Guo
cc07946782
feat(scanner): add cycle budget observability ( #3166 )
...
* feat(scanner): add cycle budget observability
* fix(scanner): clear cycle state after budget stop
* test(scanner): stabilize timeout-based scanner tests
* fix(scanner): keep cycle ILM counts scanner-only
* test(scanner): avoid test-only pending import
---------
Co-authored-by: Henry Guo <marshawcoco@users.noreply.github.com >
Co-authored-by: houseme <housemecn@gmail.com >
Co-authored-by: 安正超 <anzhengchao@gmail.com >
2026-06-03 03:32:41 +00:00
Henry Guo
1d46047d6f
feat(scanner): expand scanner observability metrics ( #3159 )
...
* feat(scanner): expand scanner observability metrics
* chore(scanner): align bucket-drive metric wording
---------
Co-authored-by: Henry Guo <marshawcoco@users.noreply.github.com >
2026-06-02 01:29:53 +00:00
Henry Guo
f3bd838925
feat(scanner): expose cycle progress metrics ( #3152 )
...
Co-authored-by: Henry Guo <marshawcoco@users.noreply.github.com >
Co-authored-by: houseme <housemecn@gmail.com >
2026-06-01 07:54:39 +00:00
Henry Guo
76da2a48d0
feat(scanner): expose cycle observability controls ( #3147 )
...
Co-authored-by: Henry Guo <marshawcoco@users.noreply.github.com >
2026-05-31 21:55:46 +00:00
Henry Guo
a99ef64db2
feat(scanner): add scanner budgets and progress metrics ( #3145 )
...
* fix(scanner): preserve maintenance scan cadence
* feat(scanner): add scanner concurrency budget
* feat(scanner): expose scanner runtime progress
* fix(scanner): address scanner review feedback
---------
Co-authored-by: Henry Guo <marshawcoco@users.noreply.github.com >
Co-authored-by: houseme <housemecn@gmail.com >
2026-05-31 16:42:38 +00:00
安正超
c684438625
fix(obs): add proxied PUT replication metrics ( #3020 )
...
fix(obs): add proxied put replication metrics
2026-05-20 05:30:57 +00:00
安正超
c727589161
fix(obs): remove stale replication metric TODOs ( #3024 )
2026-05-20 03:59:16 +00:00
houseme
bd1e57293f
fix: harden lifecycle transition compensation and regression coverage ( #2995 )
...
* fix(ecstore): honor transition worker configuration
* fix(ecstore): add transition queue backpressure metrics
* fix(ecstore): schedule transition compensation on enqueue pressure
* fix(ecstore): log transition compensation scheduling
* test(rustfs): add transition compensation fault-injection coverage
* test(rustfs): cover delete after transition compensation
* test(scanner): cover cleanup after transition compensation
* test(rustfs): extend compensation transition coverage
* test(scanner): cover backfill idempotency after compensation
* test(scanner): cover noncurrent expiry after compensation
* test(rustfs): cover versioned delete after compensation
* test(rustfs): cover delete marker lifecycle after compensation
* test(scanner): extend versioned lifecycle compensation coverage
* test(scanner): model versioned delete after compensation
* test(scanner): clarify modeled versioned delete helper
* refactor(ecstore): optimize transition enqueue hot path
* refactor(ecstore): centralize transition runtime constants
* style(ecstore): apply rustfmt for transition timeout helper
* fix(ilm): align queue-full metric semantics
* refactor(ecstore): unify immediate enqueue failure handling
* refactor(ecstore): reuse transition worker env constant
* ci(actions): update setup action inputs
2026-05-18 12:50:43 +00:00
houseme
c90bfe2b23
fix(ecstore): harden runtime read-path quorum handling ( #2872 )
2026-05-08 09:56:39 +00:00
houseme
13b4500212
feat(obs): improve telemetry stack, replication metrics, and Grafana alignment ( #2672 )
...
Co-authored-by: Filipe Monteiro <a22407332@alunos.ulht.pt >
Co-authored-by: cxymds <Cxymds@qq.com >
Co-authored-by: weisd <im@weisd.in >
Co-authored-by: loverustfs <hello@rustfs.com >
Co-authored-by: 安正超 <anzhengchao@gmail.com >
2026-04-24 13:50:17 +00:00
houseme
116db4f5d9
refactor(metrics): unify process sampling and split network IO ( #2590 )
2026-04-18 15:30:44 +00:00
houseme
1cbf156559
refactor(obs): migrate metrics runtime/schema and tighten migration guards ( #2584 )
2026-04-18 07:51:15 +00:00