houseme
035ce5d784
feat(obs): add bounded metrics dimensions ( #5645 )
...
* feat(obs): add drive topology detail metrics
Expose additive drive info, topology, state, and per-drive API metrics while preserving the existing drive metric label sets.
Backlog: rustfs/backlog#1655
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(obs): preserve suspect drive runtime state
Keep suspect as a bounded drive runtime state and avoid all-zero runtime_state samples for that storage health state.
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(obs): skip unknown drive inode samples
Avoid exporting zero inode gauges for missing or stale drive snapshots and ignore zero-count API latency buckets.
Co-Authored-By: heihutu <heihutu@gmail.com >
* feat(obs): add scanner source work detail metrics
Expose additive scanner source and cycle work metrics with bounded server/source/state labels while leaving the existing aggregate scanner metrics unchanged.
Co-Authored-By: heihutu <heihutu@gmail.com >
* feat(obs): add ilm action detail metrics
Expose additive ILM action/state task metrics with a server label while preserving the existing aggregate ILM series.
Co-Authored-By: heihutu <heihutu@gmail.com >
* feat(obs): add delivery target server metrics
Expose additive audit and notification delivery target metrics with server labels and extend removed-target tombstones for the server-aware series.
Co-Authored-By: heihutu <heihutu@gmail.com >
* feat(obs): add replication target flow metrics
Expose additive bucket replication target sent and failed-flow metrics while preserving existing bucket aggregates and target backlog series.
Co-Authored-By: heihutu <heihutu@gmail.com >
* feat(obs): add request server metrics
Expose additive API request metrics with server labels while preserving the existing request and traffic metric label sets.
Co-Authored-By: heihutu <heihutu@gmail.com >
* style(obs): apply rustfmt to metrics changes
Apply rustfmt output to the metrics dimension changes without altering behavior.
Co-Authored-By: heihutu <heihutu@gmail.com >
* style(obs): reuse audit target label constant
Use the exported audit target_id label constant for legacy audit target metrics.
Co-Authored-By: heihutu <heihutu@gmail.com >
* feat(obs): populate drive disk metrics
Co-Authored-By: heihutu <heihutu@gmail.com >
* feat(obs): add scanner bucket drive result metrics
Co-Authored-By: heihutu <heihutu@gmail.com >
* feat(obs): add replication proxy server metrics
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(obs): address metric liveness review
Use checked division for drive API latency aggregation and keep recovered drive, scanner current-cycle, replication flow, audit target, and notification target series from retaining stale values.
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(obs): address metric dimension review
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(obs): address additional metric review
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(obs): count drive calls at start
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(obs): address metrics dimension review
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(metrics): address dimension review gaps
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(metrics): address scanner review follow-ups
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(metrics): address runtime review follow-ups
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(metrics): reduce disk metric contention
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(metrics): address runtime review follow-ups
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(metrics): retire stale dimension series
Co-Authored-By: heihutu <heihutu@gmail.com >
---------
Co-authored-by: heihutu <heihutu@gmail.com >
2026-08-03 09:03:34 +08:00
houseme
fbec33bd29
Expose target-scoped durable MRF backlog metrics ( #5584 )
...
* feat(replication): expose target durable mrf backlog
Add target ARN attribution to durable MRF entries and surface target-scoped durable backlog metrics without changing existing bucket-only metric labels.
Keep legacy MRF files bucket-only by defaulting missing targetARNs to an empty list, and expose target snapshots through an additive API so existing DurableMrfBacklogSummary callers remain source-compatible.
Co-Authored-By: heihutu <heihutu@gmail.com >
* feat(replication): expose runtime target backlog (#5586 )
Track runtime replication backlog by target ARN for regular, large, delete, and MRF admission paths while preserving the existing bucket-level backlog semantics.
Add target-scoped current backlog metrics and merge them with durable target backlog snapshots for observability.
Co-authored-by: heihutu <heihutu@gmail.com >
---------
Co-authored-by: heihutu <heihutu@gmail.com >
2026-08-01 17:31:50 +00:00
houseme
da389c0e21
fix(replication): harden backlog observability ( #5564 )
...
Add RAII guards for replication runtime backlog tickets so active worker and queue counters unwind on every terminal path.
Expose node-local MRF pending, dropped, missed, and flush-failure metrics through the bucket replication Prometheus collector while keeping the existing current backlog and durable MRF gauges additive.
Update durable MRF summary maintenance to aggregate incrementally during the persister loop, avoiding repeated full-entry scans on each successful flush.
Co-authored-by: heihutu <heihutu@gmail.com >
Co-authored-by: zhi22915 <qiuzgang@gmail.com >
2026-08-01 11:12:58 +00:00
houseme
62d44d10b8
Expose replication backlog gauges ( #5557 )
...
* fix(replication): count backlog at queue admission
Co-Authored-By: heihutu <heihutu@gmail.com >
* feat(obs): expose bucket replication backlog gauges
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(obs): report recent backlog from queued work
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(obs): preserve legacy backlog metric semantics
Co-Authored-By: heihutu <heihutu@gmail.com >
* feat(obs): expose durable MRF backlog gauges
Co-Authored-By: heihutu <heihutu@gmail.com >
* test(obs): cover replication backlog metric scope
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(obs): keep backlog metrics API-compatible
Co-Authored-By: heihutu <heihutu@gmail.com >
* refactor(obs): streamline replication backlog metrics
Keep MRF backlog accounting and OBS metric collection on a single, cheaper path.
Co-Authored-By: heihutu <heihutu@gmail.com >
* test(kms): update aws capability snapshot
Co-Authored-By: heihutu <heihutu@gmail.com >
---------
Co-authored-by: heihutu <heihutu@gmail.com >
2026-08-01 08:48:45 +00:00
Zhengchao An
9a7255540b
fix(replication): add resync metrics ( #4408 )
2026-07-08 15:01:28 +08:00
Zhengchao An
e1272f2aba
revert: restore #![allow(dead_code)] - CI clippy -D warnings conflict ( #3979 )
...
revert: restore #![allow(dead_code)] - clippy -D warnings treats warn as error
The #742 PR changed #![allow(dead_code)] to #![warn(dead_code)], but
CI runs clippy with -D warnings which turns warnings into errors.
This caused CI failures across multiple PRs.
Reverting to #![allow(dead_code)] until the dead code is actually
cleaned up. The 189 warnings in ecstore should be fixed incrementally
by deleting dead code and adding item-level allows, not by changing
the crate-level policy.
2026-06-28 08:32:34 +08:00
Zhengchao An
113058af54
chore: replace blanket #![allow(dead_code)] with #![warn(dead_code)] ( #742 ) ( #3974 )
2026-06-28 07:50:51 +08:00
安正超
c684438625
fix(obs): add proxied PUT replication metrics ( #3020 )
...
fix(obs): add proxied put replication metrics
2026-05-20 05:30:57 +00:00
安正超
c727589161
fix(obs): remove stale replication metric TODOs ( #3024 )
2026-05-20 03:59:16 +00:00
houseme
13b4500212
feat(obs): improve telemetry stack, replication metrics, and Grafana alignment ( #2672 )
...
Co-authored-by: Filipe Monteiro <a22407332@alunos.ulht.pt >
Co-authored-by: cxymds <Cxymds@qq.com >
Co-authored-by: weisd <im@weisd.in >
Co-authored-by: loverustfs <hello@rustfs.com >
Co-authored-by: 安正超 <anzhengchao@gmail.com >
2026-04-24 13:50:17 +00:00
houseme
1cbf156559
refactor(obs): migrate metrics runtime/schema and tighten migration guards ( #2584 )
2026-04-18 07:51:15 +00:00