anthonymartin
706a8b6061
fix(scanner): publish bounded observational usage ( #5742 )
...
* fix(scanner): publish bounded observational usage
* test(ci): serialize embedded integration ports
* test(cache): isolate generation-change timeout
* fix(scanner): address observational usage review
Co-Authored-By: heihutu <heihutu@gmail.com >
---------
Co-authored-by: Anthony Martin <949506+anthonymartin@users.noreply.github.com >
Co-authored-by: houseme <housemecn@gmail.com >
Co-authored-by: heihutu <heihutu@gmail.com >
2026-08-07 08:52:34 +08:00
houseme
035ce5d784
feat(obs): add bounded metrics dimensions ( #5645 )
...
* feat(obs): add drive topology detail metrics
Expose additive drive info, topology, state, and per-drive API metrics while preserving the existing drive metric label sets.
Backlog: rustfs/backlog#1655
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(obs): preserve suspect drive runtime state
Keep suspect as a bounded drive runtime state and avoid all-zero runtime_state samples for that storage health state.
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(obs): skip unknown drive inode samples
Avoid exporting zero inode gauges for missing or stale drive snapshots and ignore zero-count API latency buckets.
Co-Authored-By: heihutu <heihutu@gmail.com >
* feat(obs): add scanner source work detail metrics
Expose additive scanner source and cycle work metrics with bounded server/source/state labels while leaving the existing aggregate scanner metrics unchanged.
Co-Authored-By: heihutu <heihutu@gmail.com >
* feat(obs): add ilm action detail metrics
Expose additive ILM action/state task metrics with a server label while preserving the existing aggregate ILM series.
Co-Authored-By: heihutu <heihutu@gmail.com >
* feat(obs): add delivery target server metrics
Expose additive audit and notification delivery target metrics with server labels and extend removed-target tombstones for the server-aware series.
Co-Authored-By: heihutu <heihutu@gmail.com >
* feat(obs): add replication target flow metrics
Expose additive bucket replication target sent and failed-flow metrics while preserving existing bucket aggregates and target backlog series.
Co-Authored-By: heihutu <heihutu@gmail.com >
* feat(obs): add request server metrics
Expose additive API request metrics with server labels while preserving the existing request and traffic metric label sets.
Co-Authored-By: heihutu <heihutu@gmail.com >
* style(obs): apply rustfmt to metrics changes
Apply rustfmt output to the metrics dimension changes without altering behavior.
Co-Authored-By: heihutu <heihutu@gmail.com >
* style(obs): reuse audit target label constant
Use the exported audit target_id label constant for legacy audit target metrics.
Co-Authored-By: heihutu <heihutu@gmail.com >
* feat(obs): populate drive disk metrics
Co-Authored-By: heihutu <heihutu@gmail.com >
* feat(obs): add scanner bucket drive result metrics
Co-Authored-By: heihutu <heihutu@gmail.com >
* feat(obs): add replication proxy server metrics
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(obs): address metric liveness review
Use checked division for drive API latency aggregation and keep recovered drive, scanner current-cycle, replication flow, audit target, and notification target series from retaining stale values.
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(obs): address metric dimension review
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(obs): address additional metric review
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(obs): count drive calls at start
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(obs): address metrics dimension review
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(metrics): address dimension review gaps
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(metrics): address scanner review follow-ups
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(metrics): address runtime review follow-ups
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(metrics): reduce disk metric contention
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(metrics): address runtime review follow-ups
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(metrics): retire stale dimension series
Co-Authored-By: heihutu <heihutu@gmail.com >
---------
Co-authored-by: heihutu <heihutu@gmail.com >
2026-08-03 09:03:34 +08:00
houseme
fbec33bd29
Expose target-scoped durable MRF backlog metrics ( #5584 )
...
* feat(replication): expose target durable mrf backlog
Add target ARN attribution to durable MRF entries and surface target-scoped durable backlog metrics without changing existing bucket-only metric labels.
Keep legacy MRF files bucket-only by defaulting missing targetARNs to an empty list, and expose target snapshots through an additive API so existing DurableMrfBacklogSummary callers remain source-compatible.
Co-Authored-By: heihutu <heihutu@gmail.com >
* feat(replication): expose runtime target backlog (#5586 )
Track runtime replication backlog by target ARN for regular, large, delete, and MRF admission paths while preserving the existing bucket-level backlog semantics.
Add target-scoped current backlog metrics and merge them with durable target backlog snapshots for observability.
Co-authored-by: heihutu <heihutu@gmail.com >
---------
Co-authored-by: heihutu <heihutu@gmail.com >
2026-08-01 17:31:50 +00:00
houseme
da389c0e21
fix(replication): harden backlog observability ( #5564 )
...
Add RAII guards for replication runtime backlog tickets so active worker and queue counters unwind on every terminal path.
Expose node-local MRF pending, dropped, missed, and flush-failure metrics through the bucket replication Prometheus collector while keeping the existing current backlog and durable MRF gauges additive.
Update durable MRF summary maintenance to aggregate incrementally during the persister loop, avoiding repeated full-entry scans on each successful flush.
Co-authored-by: heihutu <heihutu@gmail.com >
Co-authored-by: zhi22915 <qiuzgang@gmail.com >
2026-08-01 11:12:58 +00:00
houseme
62d44d10b8
Expose replication backlog gauges ( #5557 )
...
* fix(replication): count backlog at queue admission
Co-Authored-By: heihutu <heihutu@gmail.com >
* feat(obs): expose bucket replication backlog gauges
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(obs): report recent backlog from queued work
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(obs): preserve legacy backlog metric semantics
Co-Authored-By: heihutu <heihutu@gmail.com >
* feat(obs): expose durable MRF backlog gauges
Co-Authored-By: heihutu <heihutu@gmail.com >
* test(obs): cover replication backlog metric scope
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(obs): keep backlog metrics API-compatible
Co-Authored-By: heihutu <heihutu@gmail.com >
* refactor(obs): streamline replication backlog metrics
Keep MRF backlog accounting and OBS metric collection on a single, cheaper path.
Co-Authored-By: heihutu <heihutu@gmail.com >
* test(kms): update aws capability snapshot
Co-Authored-By: heihutu <heihutu@gmail.com >
---------
Co-authored-by: heihutu <heihutu@gmail.com >
2026-08-01 08:48:45 +00:00
houseme
021c955c21
fix(obs): report live replication backlog ( #4448 )
...
Co-authored-by: heihutu <heihutu@gmail.com >
2026-07-08 16:39:06 +08:00
Zhengchao An
9a7255540b
fix(replication): add resync metrics ( #4408 )
2026-07-08 15:01:28 +08:00
Zhengchao An
644a472e52
refactor(obs): hide replication stats handle ( #4149 )
2026-07-01 23:37:23 +08:00
wood
7484a61fa3
feat(compression): show cluster-level compression stat in grafana ( #4112 )
2026-06-30 19:05:05 +08:00
Zhengchao An
27a491a730
refactor(runtime): guard replication stats globals ( #4058 )
2026-06-29 19:55:38 +08:00
Zhengchao An
ea88f5b67b
refactor(obs): use lifecycle runtime facades ( #4031 )
2026-06-29 09:29:50 +08:00
Zhengchao An
1b3dea012e
refactor: route ecstore runtime globals through facade ( #3941 )
2026-06-27 12:27:03 +08:00
Zhengchao An
72ae43cb90
refactor: segment external storage contract imports ( #3908 )
2026-06-26 16:44:59 +08:00
Zhengchao An
1f7e159388
refactor: segment external storage api boundaries ( #3903 )
2026-06-26 15:10:49 +08:00
Zhengchao An
e37e367390
refactor: route remaining external storage boundaries ( #3889 )
2026-06-26 06:22:51 +08:00