houseme
06eeb886f8
feat(heal): add MRF queue, durable journal, and intent consumer (HS-01)
...
Consumer half of the mission repair feed: a bounded pending queue
(100k intents / 8 MiB dual ceiling, drop-newest on overflow), a durable
journal at buckets/.heal/mrf/journal.bin holding the unaccepted pending
snapshot, and a consumer task that batches intents off the global
channel, translates them into prioritized heal requests (decode
failure -> Urgent ECDecode, metadata corruption -> High Metadata,
partial write -> Normal object heal), and retries full admissions with
a 5s backoff and a 3-attempt ceiling.
Durability: every journal record carries its own CRC32 and a
format/version header, so a torn tail truncates cleanly at replay; the
journal is deleted after a successful replay and when the pending set
drains (mirroring MinIO's post-replay list.bin unlink). Losing the last
500 ms flush window is acceptable: replayed duplicates merge via the
manager dedup key and read-repair remains the safety net.
Metrics: rustfs_heal_mrf_queue_depth/_queue_bytes, _dropped_total
{reason}, _replayed_total, _journal_bytes, _journal_fsync_total.
The consumer is wired at heal runtime bootstrap right after manager
start, honoring RUSTFS_HEAL_MRF_ENABLE (default on, rollback = off).
Tests: unit tests for the dual ceiling, record roundtrip, torn-tail
truncation, and the priority mapping; integration tests against a real
4-disk ECStore proving channel intents reach the manager queue as
Urgent/mrf-attributed requests and journal replay arms intents, drops
torn tails, and removes the file.
Part of backlog#1865 (option a).
Co-Authored-By: heihutu <heihutu@gmail.com >
2026-08-18 09:03:11 +08:00
houseme
f17ea7f146
fix(heal): harden replacement rebuild tracking ( #5892 )
...
* fix(heal): gate auto replacement formatting
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): require replacement target outcomes
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): bind resumes to replacement targets
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): fence healing marker ownership
Co-Authored-By: heihutu <heihutu@gmail.com >
* test(heal): cover replacement target completion
Co-Authored-By: heihutu <heihutu@gmail.com >
* docs(heal): clarify replacement recovery status
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): canonicalize replacement target checks
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): satisfy marker test module lint
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): scope automatic replacement format
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): require a mounted replacement target
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(ecstore): avoid cloned ref slice in test
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): revalidate replacement before scanning
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): reset stale resume checkpoints
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): release scanner disk map before probing
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): persist replacement intent before format
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): fail closed on mountinfo read errors
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): fence replacement target identity
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): order replacement completion cleanup
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): atomically seal replacement completion
Co-Authored-By: heihutu <heihutu@gmail.com >
* test(heal): census replacement target shards
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): fence replacement recovery ownership
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): preserve replacement recovery anchors
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): satisfy replacement recovery lint gates
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(ecstore): bind replacement identity to mount lease
Co-Authored-By: heihutu <heihutu@gmail.com >
* test(heal): cover durable replacement recovery states
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): validate persisted resume task identifiers
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(ecstore): avoid blocking replacement marker CAS
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): report failed marker rollback
Co-Authored-By: heihutu <heihutu@gmail.com >
* test(heal): pin replacement resume schema compatibility
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): preserve durable recovery anchors
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(ecstore): preserve public disk path semantics
Co-Authored-By: heihutu <heihutu@gmail.com >
* test(heal): use canonical replacement task ids
Co-Authored-By: heihutu <heihutu@gmail.com >
* test(heal): cover automatic replacement in 3x4 cluster
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): verify replacement target commits
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): persist replacement completion proof
Co-Authored-By: heihutu <heihutu@gmail.com >
* feat(heal): expose durable replacement status
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): bound durable replacement discovery
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): remove replacement readiness bypass
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): retry terminal replacement cleanup
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): isolate replacement intents from legacy resume
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): migrate legacy replacement intents at startup
Co-Authored-By: heihutu <heihutu@gmail.com >
* style(heal): apply strict clippy fix
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): prioritize active replacement recovery state
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): bind readiness to the admitted mount lease
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): atomically publish replacement intents
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): isolate replacement recovery directory
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): tolerate an empty recovery directory
Co-Authored-By: heihutu <heihutu@gmail.com >
* style(heal): remove redundant disk bytes conversion
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): reconcile proof-first replacement recovery
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): fence torn intent recovery
Co-Authored-By: heihutu <heihutu@gmail.com >
* test(heal): cover replacement migration conflicts
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(ecstore): fence replacement lease mount identity
Co-Authored-By: heihutu <heihutu@gmail.com >
* test(heal): cover missing replacement path admission
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): reject conflicting legacy completion proof
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(ecstore): fall back to proc mount identity
Co-Authored-By: heihutu <heihutu@gmail.com >
* feat(admin): expose replacement recovery status
Surface the local durable replacement recovery snapshot in the background heal status response so operators can tell whether replacement cleanup is definitive or still pending.
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): keep replacement status compatible
Keep the existing background heal status response wire-compatible while retaining the Linux mount lease cleanup needed for the replacement recovery branch.
Co-Authored-By: heihutu <heihutu@gmail.com >
* style(ecstore): match linux mount lease formatting
Keep Linux rustfmt output stable for the replacement mount lease comparison.
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(ecstore): qualify mount lease test constant
Use the disk module path for the format config constant in the Linux mount lease regression test.
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(ecstore): keep procfd mount roots directory-safe
Use a procfd path with an explicit directory component so Unix directory guards can open the replacement mount lease root with O_NOFOLLOW while preserving handle-relative I/O semantics.
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(ecstore): delete empty leased buckets via dirfd
Use the held mount lease fd as the parent for non-force empty bucket deletion on Linux so procfd-rooted paths do not get rejected as BucketNotEmpty. Also make the download-part OpenOptions truncate behavior explicit and keep fsync test recording stable across procfd canonicalization.
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(ecstore): scan leased bucket paths for emptiness
Use the local disk I/O root for bucket emptiness probes before non-force bucket deletion and table-bucket metadata checks. This keeps validation on the same mount instance as the subsequent local disk delete path.
Co-Authored-By: heihutu <heihutu@gmail.com >
* test(ecstore): align lease path test probes
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): block unsafe replacement recovery restarts
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): defer blocked replacement candidates
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): retry transient replacement discovery
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): keep transient recovery errors retryable
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): block corrupt legacy replacement state
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): classify flat replacement intent corruption
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): keep transient resume loads retryable
Classify malformed legacy replacement state as blocking corruption while preserving disk and transient load failures for retry. This avoids permanently blocking replacement recovery on temporary storage errors.
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): avoid latching transient legacy publishes
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): retry blocked legacy migrations
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(heal): defer blocked startup recoveries
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(ecstore): preserve disk sync limiter across lease roots
Co-Authored-By: heihutu <heihutu@gmail.com >
---------
Co-authored-by: heihutu <heihutu@gmail.com >
Co-authored-by: zhi22915 <qiuzgang@gmail.com >
2026-08-10 08:32:47 +08:00
cxymds
4290f390dd
fix(heal): aggregate status across cluster nodes ( #4990 )
...
* fix(rpc): bind internode auth to exact targets
* fix(heal): initialize the runtime atomically
* fix(heal): aggregate status across cluster nodes
---------
Co-authored-by: Zhengchao An <anzhengchao@gmail.com >
2026-07-19 15:25:52 +00:00
cxymds
1ac0841f6f
fix(heal): initialize the runtime atomically ( #4989 )
...
* fix(rpc): bind internode auth to exact targets
* fix(heal): initialize the runtime atomically
2026-07-19 14:20:37 +00:00
Zhengchao An
535d672b1f
fix(admin): report heal runtime state ( #4786 )
2026-07-13 11:02:13 +08:00
cxymds
ae641a33ac
feat(admin): expose global heal progress ( #3897 )
2026-06-26 15:19:48 +08:00
cxymds
90638bfc19
fix(heal): add backpressure to repair admission ( #3900 )
2026-06-26 15:19:37 +08:00
Henry Guo
b387689f26
feat(heal): expose scanner-aware operations status ( #3483 )
...
Co-authored-by: Henry Guo <marshawcoco@users.noreply.github.com >
2026-06-15 22:28:53 +08:00
houseme
7da10db852
refactor(logging): standardize heal and scanner events ( #3414 )
...
* refactor(logging): standardize heal and scanner events
* chore(git): untrack local logging governance note
* chore(git): ignore local logging governance note
2026-06-14 01:47:39 +08:00
houseme
50d03ef021
perf(memory): add reclaim signals and cache controls ( #2689 )
2026-04-26 16:42:35 +00:00
安正超
5625f04697
fix(common): remove panic paths in runtime helpers ( #2116 )
...
Co-authored-by: houseme <housemecn@gmail.com >
Co-authored-by: heihutu <30542132+heihutu@users.noreply.github.com >
2026-03-11 18:12:37 +08:00
weisd
dce117840c
refactor: NamespaceLock (nslock), AHM→Heal Crate, and Lock/Clippy Fixes ( #1664 )
...
Co-authored-by: Copilot <198982749+Copilot@users.noreply.github.com >
Co-authored-by: weisd <2057561+weisd@users.noreply.github.com >
Co-authored-by: houseme <housemecn@gmail.com >
2026-01-30 13:13:41 +08:00