houseme
360bceafce
feat(heal): add progress and trace observability ( #6179 )
...
* feat(heal): track erasure set progress baseline
Record erasure-set heal byte progress from per-object results and seed progress totals from complete usage-cache snapshots when available.
Keep usage-cache failures observational so heal execution continues without a baseline.
Co-Authored-By: heihutu <heihutu@gmail.com >
* feat(heal): skip filtered erasure set versions
Skip erasure-set versions written after the durable heal start time, and queue lifecycle-expired versions for expiry before skipping them.
Track new-version and ILM-expired skips separately so progress can explain completed baseline work without treating these skips as retry-blocking failures.
Co-Authored-By: heihutu <heihutu@gmail.com >
* feat(heal): wire abandoned data-dir cleanup check
Connect check_abandoned_parts through ECStore, pool, and set layers so heal can invoke the existing orphan data-dir reclaim path instead of returning NotImplemented.
Add dry-run support to the reclaim scan and cover dry-run plus scoped set behavior with regression tests.
Co-Authored-By: heihutu <heihutu@gmail.com >
* feat(obs): add heal scanner trace bus
Introduce an in-process broadcast trace bus with typed heal and scanner events, lazy event construction, and bounded lagged-subscriber behavior.
Cover zero-subscriber publishing, subscription delivery, drop accounting, and lagged receivers with focused common-crate tests.
Co-Authored-By: heihutu <heihutu@gmail.com >
* feat(obs): stream heal trace events from admin API
Wire the admin trace endpoint to the common trace bus for heal/scanner events, including kind, regex, and threshold filtering.
Co-Authored-By: heihutu <heihutu@gmail.com >
* feat(obs): emit heal trace events
Publish heal task lifecycle and abandoned-parts cleanup events through the common trace bus so the admin trace stream has live heal diagnostics.
Co-Authored-By: heihutu <heihutu@gmail.com >
* feat(obs): emit scanner trace events
Publish scanner folder, lifecycle action, and heal-candidate events through the common trace bus for live admin scanner diagnostics.
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): route data usage loader through storage api
Keep ECStore data-usage facade access behind the heal storage_api boundary so architecture migration guards can validate the heal progress path.
Co-Authored-By: heihutu <heihutu@gmail.com >
* perf(heal): avoid lifecycle snapshots on ordinary heal pages
Only request lifecycle object snapshots when the heal pass has lifecycle expiry context. This keeps ordinary listing and disk-walk pages from cloning FileInfo/ObjectInfo payloads while preserving the skip path that queues expired versions.
Co-Authored-By: heihutu <heihutu@gmail.com >
* test(heal): update bug-fix mocks for lifecycle snapshots
Carry the lifecycle snapshot opt-in argument through the remaining heal bug-fix test mocks so all-targets clippy covers the updated storage trait.
Co-Authored-By: heihutu <heihutu@gmail.com >
* test(rustfs): sync heal storage mock signature
Update the rustfs storage RPC test mock for the lifecycle snapshot opt-in argument and cover it with rustfs all-targets clippy.
Co-Authored-By: heihutu <heihutu@gmail.com >
* test(e2e): allocate smoke ports across nextest processes
Serialize E2E port selection with a small /tmp allocator so nextest workers do not reuse the same just-released ephemeral port before RustFS binds it.
Co-Authored-By: heihutu <heihutu@gmail.com >
---------
Co-authored-by: heihutu <heihutu@gmail.com >
2026-08-18 08:29:29 +08:00
houseme
f17ea7f146
fix(heal): harden replacement rebuild tracking ( #5892 )
...
* fix(heal): gate auto replacement formatting
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): require replacement target outcomes
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): bind resumes to replacement targets
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): fence healing marker ownership
Co-Authored-By: heihutu <heihutu@gmail.com >
* test(heal): cover replacement target completion
Co-Authored-By: heihutu <heihutu@gmail.com >
* docs(heal): clarify replacement recovery status
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): canonicalize replacement target checks
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): satisfy marker test module lint
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): scope automatic replacement format
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): require a mounted replacement target
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(ecstore): avoid cloned ref slice in test
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): revalidate replacement before scanning
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): reset stale resume checkpoints
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): release scanner disk map before probing
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): persist replacement intent before format
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): fail closed on mountinfo read errors
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): fence replacement target identity
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): order replacement completion cleanup
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): atomically seal replacement completion
Co-Authored-By: heihutu <heihutu@gmail.com >
* test(heal): census replacement target shards
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): fence replacement recovery ownership
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): preserve replacement recovery anchors
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): satisfy replacement recovery lint gates
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(ecstore): bind replacement identity to mount lease
Co-Authored-By: heihutu <heihutu@gmail.com >
* test(heal): cover durable replacement recovery states
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): validate persisted resume task identifiers
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(ecstore): avoid blocking replacement marker CAS
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): report failed marker rollback
Co-Authored-By: heihutu <heihutu@gmail.com >
* test(heal): pin replacement resume schema compatibility
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): preserve durable recovery anchors
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(ecstore): preserve public disk path semantics
Co-Authored-By: heihutu <heihutu@gmail.com >
* test(heal): use canonical replacement task ids
Co-Authored-By: heihutu <heihutu@gmail.com >
* test(heal): cover automatic replacement in 3x4 cluster
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): verify replacement target commits
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): persist replacement completion proof
Co-Authored-By: heihutu <heihutu@gmail.com >
* feat(heal): expose durable replacement status
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): bound durable replacement discovery
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): remove replacement readiness bypass
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): retry terminal replacement cleanup
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): isolate replacement intents from legacy resume
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): migrate legacy replacement intents at startup
Co-Authored-By: heihutu <heihutu@gmail.com >
* style(heal): apply strict clippy fix
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): prioritize active replacement recovery state
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): bind readiness to the admitted mount lease
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): atomically publish replacement intents
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): isolate replacement recovery directory
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): tolerate an empty recovery directory
Co-Authored-By: heihutu <heihutu@gmail.com >
* style(heal): remove redundant disk bytes conversion
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): reconcile proof-first replacement recovery
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): fence torn intent recovery
Co-Authored-By: heihutu <heihutu@gmail.com >
* test(heal): cover replacement migration conflicts
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(ecstore): fence replacement lease mount identity
Co-Authored-By: heihutu <heihutu@gmail.com >
* test(heal): cover missing replacement path admission
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): reject conflicting legacy completion proof
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(ecstore): fall back to proc mount identity
Co-Authored-By: heihutu <heihutu@gmail.com >
* feat(admin): expose replacement recovery status
Surface the local durable replacement recovery snapshot in the background heal status response so operators can tell whether replacement cleanup is definitive or still pending.
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): keep replacement status compatible
Keep the existing background heal status response wire-compatible while retaining the Linux mount lease cleanup needed for the replacement recovery branch.
Co-Authored-By: heihutu <heihutu@gmail.com >
* style(ecstore): match linux mount lease formatting
Keep Linux rustfmt output stable for the replacement mount lease comparison.
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(ecstore): qualify mount lease test constant
Use the disk module path for the format config constant in the Linux mount lease regression test.
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(ecstore): keep procfd mount roots directory-safe
Use a procfd path with an explicit directory component so Unix directory guards can open the replacement mount lease root with O_NOFOLLOW while preserving handle-relative I/O semantics.
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(ecstore): delete empty leased buckets via dirfd
Use the held mount lease fd as the parent for non-force empty bucket deletion on Linux so procfd-rooted paths do not get rejected as BucketNotEmpty. Also make the download-part OpenOptions truncate behavior explicit and keep fsync test recording stable across procfd canonicalization.
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(ecstore): scan leased bucket paths for emptiness
Use the local disk I/O root for bucket emptiness probes before non-force bucket deletion and table-bucket metadata checks. This keeps validation on the same mount instance as the subsequent local disk delete path.
Co-Authored-By: heihutu <heihutu@gmail.com >
* test(ecstore): align lease path test probes
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): block unsafe replacement recovery restarts
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): defer blocked replacement candidates
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): retry transient replacement discovery
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): keep transient recovery errors retryable
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): block corrupt legacy replacement state
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): classify flat replacement intent corruption
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): keep transient resume loads retryable
Classify malformed legacy replacement state as blocking corruption while preserving disk and transient load failures for retry. This avoids permanently blocking replacement recovery on temporary storage errors.
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): avoid latching transient legacy publishes
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): retry blocked legacy migrations
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): defer blocked startup recoveries
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(ecstore): preserve disk sync limiter across lease roots
Co-Authored-By: heihutu <heihutu@gmail.com >
---------
Co-authored-by: heihutu <heihutu@gmail.com >
Co-authored-by: zhi22915 <qiuzgang@gmail.com >
2026-08-10 08:32:47 +08:00
cxymds
4290f390dd
fix(heal): aggregate status across cluster nodes ( #4990 )
...
* fix(rpc): bind internode auth to exact targets
* fix(heal): initialize the runtime atomically
* fix(heal): aggregate status across cluster nodes
---------
Co-authored-by: Zhengchao An <anzhengchao@gmail.com >
2026-07-19 15:25:52 +00:00
cxymds
1ac0841f6f
fix(heal): initialize the runtime atomically ( #4989 )
...
* fix(rpc): bind internode auth to exact targets
* fix(heal): initialize the runtime atomically
2026-07-19 14:20:37 +00:00
Zhengchao An
535d672b1f
fix(admin): report heal runtime state ( #4786 )
2026-07-13 11:02:13 +08:00
cxymds
ae641a33ac
feat(admin): expose global heal progress ( #3897 )
2026-06-26 15:19:48 +08:00
cxymds
90638bfc19
fix(heal): add backpressure to repair admission ( #3900 )
2026-06-26 15:19:37 +08:00
Henry Guo
b387689f26
feat(heal): expose scanner-aware operations status ( #3483 )
...
Co-authored-by: Henry Guo <marshawcoco@users.noreply.github.com >
2026-06-15 22:28:53 +08:00
houseme
7da10db852
refactor(logging): standardize heal and scanner events ( #3414 )
...
* refactor(logging): standardize heal and scanner events
* chore(git): untrack local logging governance note
* chore(git): ignore local logging governance note
2026-06-14 01:47:39 +08:00
houseme
50d03ef021
perf(memory): add reclaim signals and cache controls ( #2689 )
2026-04-26 16:42:35 +00:00
安正超
5625f04697
fix(common): remove panic paths in runtime helpers ( #2116 )
...
Co-authored-by: houseme <housemecn@gmail.com >
Co-authored-by: heihutu <30542132+heihutu@users.noreply.github.com >
2026-03-11 18:12:37 +08:00
weisd
dce117840c
refactor: NamespaceLock (nslock), AHM→Heal Crate, and Lock/Clippy Fixes ( #1664 )
...
Co-authored-by: Copilot <198982749+Copilot@users.noreply.github.com >
Co-authored-by: weisd <2057561+weisd@users.noreply.github.com >
Co-authored-by: houseme <housemecn@gmail.com >
2026-01-30 13:13:41 +08:00