cxymds
6018dd372f
perf(ilm): reduce transition transaction mutations ( #7320 )
...
* perf(ilm): reduce transition transaction mutations
* test(ilm): rename transition kill points
---------
Co-authored-by: Zhengchao An <anzhengchao@gmail.com >
2026-09-07 02:40:57 +00:00
cxymds
a70c96d520
feat(ilm): execute legacy recovery dispositions ( #7304 )
...
* feat(ilm): execute legacy recovery dispositions
* feat(ilm): retry retained transition recovery (#7308 )
* fix(admin): use gateway errors for recovery retries
---------
Co-authored-by: houseme <housemecn@gmail.com >
Co-authored-by: Zhengchao An <anzhengchao@gmail.com >
2026-09-07 01:46:43 +00:00
Zhengchao An
22bff27aee
fix(storage): derive multipart identity from stored parts ( #7305 )
2026-09-06 22:04:15 +08:00
cxymds
002ac9544c
feat(ilm): add immutable legacy recovery exports ( #7283 )
...
* feat(ilm): add immutable legacy recovery exports
* feat(ilm): add legacy recovery disposition records (#7292 )
* fix(ilm): stabilize legacy recovery decode errors
* fix(admin): route recovery auth errors through gateway
* fix(ilm): resolve recovery disposition clippy errors
* fix(ilm): remove redundant recovery test clones
2026-09-06 21:20:16 +08:00
cxymds
f17f31a3df
feat(ilm): persist recovery controls for legacy tier journals ( #7252 )
2026-09-06 14:13:09 +08:00
GatewayJ
9fc9b5e69c
fix(ecstore): align platform and test helper compilation ( #7214 )
...
Co-authored-by: Zhengchao An <anzhengchao@gmail.com >
2026-09-06 10:50:48 +08:00
cxymds
941fae61a0
fix(ilm): persist bounded transition recovery controls ( #7200 )
...
* fix(ilm): persist bounded transition recovery controls
* test(ilm): keep expiry sentinel within control range
* fix(ilm): repair recovery control CI regressions
---------
Co-authored-by: overtrue <anzhengchao@gmail.com >
2026-09-06 10:10:24 +08:00
Zhengchao An
dd368f0f5b
fix(odm): fence source work against bucket recreation ( #7231 )
...
* fix(odm): fence backfill checkpoints by bucket incarnation
* fix(odm): bind source work to the bucket incarnation
* fix(odm): retain checkpoint fences through owned commit tails
* docs(odm): explain application service and incarnation boundaries
* test(odm): probe lifecycle fence after checkpoint waiter aborts
* fix(odm): defer source identity errors past local reads
* docs(metadata): clarify MinIO target recovery limits
* fix(odm): keep source-free reads independent of capture errors
* fix(odm): retain one source policy snapshot across lookup
* test(odm): name recorded metadata hook snapshots
2026-09-06 02:05:43 +08:00
Zhengchao An
6d8606412e
fix(odm): compile relocated instance-bound backfill service ( #7230 )
2026-09-06 01:43:14 +08:00
Zhengchao An
e2a921bc16
fix(storage): harden ODM and scanner publication ( #7187 )
...
* fix(storage): harden ODM and scanner publication
* fix(app): simplify absent SSE configuration matching
* test(heal): settle PUT rename tails before disk-wipe fixtures
* fix(ecstore): remove duplicate local rename implementation
Keep the canonical commit module after concurrent storage changes merged.
The control-write and rollback changes are already present there.
Co-Authored-By: heihutu <heihutu@gmail.com >
Co-Authored-By: zhi22915 <qiuzgang@gmail.com >
* fix(ci): satisfy new clippy lints
* style(scanner): order merged test imports
* fix(scanner): invalidate bucket work after namespace completion
* fix(scanner): fence cached snapshots by scan execution
---------
Co-authored-by: houseme <housemecn@gmail.com >
Co-authored-by: heihutu <heihutu@gmail.com >
Co-authored-by: zhi22915 <qiuzgang@gmail.com >
2026-09-05 13:47:12 +00:00
Zhengchao An
cbfd5b92f4
refactor(ecstore): isolate local object rename commit ( #7166 )
...
* fix(ecstore): drain durable control-plane write tails
* fix(ecstore): retain PUT staging after incomplete rollback
* fix(ecstore): drain backfill checkpoint before confirmation
* refactor(ecstore): isolate local object rename commit
* refactor(ecstore): remove moved quota fence import
* fix(ecstore): retain per-disk rename rollback outcomes
* fix(ecstore): retain indeterminate rename recovery evidence
* test(ecstore): mark rollback fixtures as inline data
* test(ecstore): match sealed context fixture map type
* test(ecstore): match sealed context fixture map type
* fix(ecstore): preserve known preflight rename rejections
* test(ecstore): cover observed rename outer failures
* test(ecstore): count decommission faults across retry restarts
2026-09-05 07:19:29 +00:00
cxymds
a3b8183be9
test(ecstore): narrow barrier re-export cfgs ( #7152 )
...
Co-authored-by: Zhengchao An <anzhengchao@gmail.com >
2026-09-05 06:12:45 +00:00
cxymds
4dbc58887a
fix(tier): probe legacy transition version state ( #7138 )
2026-09-05 06:00:14 +00:00
cxymds
6eb60f8e72
feat(tier): add durable probe intent protocol ( #7151 )
...
* feat(tier): add durable probe intent protocol
* test(tier): remove redundant intent clones
2026-09-05 09:06:05 +08:00
cxymds
65ed86f76e
test(ecstore): run checkpoint publication test on large stack ( #7134 )
2026-09-04 12:39:18 +00:00
cxymds
ff3c5a4989
fix(ilm): drain tier-delete recovery pages ( #7133 )
2026-09-04 10:46:19 +00:00
cxymds
4d226998e2
fix(ilm): chunk large tier-delete dispatches ( #7123 )
...
* fix(ilm): chunk large tier-delete dispatches
* fix(ilm): bound tier-delete chunk dispatch stack use
2026-09-04 10:30:32 +00:00
cxymds
3654c147e2
fix(ilm): defer aborted dispatch cleanup to recovery ( #7126 )
2026-09-04 08:49:14 +00:00
cxymds
80c88a9031
fix(ilm): delete historical null versions by exact identity ( #7109 )
2026-09-04 08:14:47 +08:00
cxymds
0181a583a6
fix(ilm): recover orphaned restore generations ( #7104 )
2026-09-03 20:03:51 +08:00
cxymds
a6cb34c7a4
fix: fence transition transaction recovery ( #7095 )
2026-09-03 10:38:30 +00:00
houseme
0e6ee3bf62
feat(scanner): coordinate usage and workload boundaries ( #7093 )
...
* test(scanner): wire usage and heal rebuild gates
* docs(scanner): define usage authority protocol
* docs(heal): clarify scanner and ecstore boundaries
* refactor(scanner): split metrics from contracts
* feat(scanner): use shared workload snapshots
* fix(ecstore): recheck capacity before decommission drain
2026-09-03 17:02:43 +08:00
cxymds
54a7e9f307
fix(ecstore): stabilize decommission config and retry tests ( #7091 )
2026-09-03 14:45:26 +08:00
cxymds
5e58b1d3a2
test(ecstore): stabilize subquorum free-version fixture ( #7090 )
2026-09-03 14:06:11 +08:00
cxymds
df30dff1a7
feat(storage): complete Snowball and decommission follow-ups ( #7039 )
...
* feat(storage): complete Snowball and capacity follow-ups
* fix(ecstore): clarify V3 capacity gate guidance
* fix(ecstore): keep target contention retryable
* fix(ecstore): preserve typed target lock errors
* fix(ecstore): harden decommission recovery
* fix(ecstore): close decommission recovery races
* fix(ecstore): fail closed on multipart cleanup gaps
* fix(ecstore): model capacity mutation parameters
* fix(ecstore): settle checkpoint capacity retries
2026-09-03 03:58:59 +00:00
Henry Guo
98f7e63396
fix(heal): recover replacement after transient disk errors ( #7059 )
...
* fix(heal): recover replacement after transient disk errors
* fix(heal): satisfy replacement status clippy lint
2026-09-03 07:03:01 +08:00
houseme
ba20af77bb
fix(ecstore): wait for multipart copy readiness ( #7065 )
2026-09-02 14:31:29 +00:00
cxymds
afc66b7182
fix(ilm): enqueue committed tier free versions ( #7041 )
...
* fix(ilm): enqueue committed tier free versions
* fix(ilm): stabilize causal cleanup CI coverage
* test(ilm): make expire GET race deterministic
* test(ilm): synchronize expiry with active GET
2026-09-02 11:08:28 +00:00
唐小鸭
32eb116cbc
fix(ecstore): report unreachable bucket-delete residue at error level ( #7048 )
...
DeleteBucket answers from a raw per-disk residue scan rather than from a
listing, so it can refuse for a reason no S3 request can observe: the
client drains every version the API will show, DeleteBucket still returns
BucketNotEmpty, and the client-visible message is the generic "The bucket
you tried to delete is not empty" for every blocker kind.
The server does know which residue blocked it, and where — that is what
`bucket_delete_blocked` carries. But it was emitted at `debug`, below
both the `error` DEFAULT_LOG_LEVEL and the `info` the CI s3-tests lane
runs at, so it was never actually written down. An intermittent
BucketNotEmpty in that lane leaves a server log with no trace of the
refusal at all, which is not a diagnosable state: confirmed against the
artifact log of a failing run, where the rejected bucket appears only in
span-close lines and the blocker event is absent entirely.
Split the blocker kinds by whether the client can still reach the
residue. A visible version or a tier free-version is an ordinary 409 —
the bucket really is not empty and the caller can list and delete what is
left — so that stays at `warn`. UnknownXlMeta, OrphanDirectory, and
DiagnosticBudgetExceeded are on-disk state no S3 request can remove; that
is a server-side integrity problem and is now reported at `error`, with
the blocker kind, the residue counts, and the sample path.
This does not change what DeleteBucket accepts or rejects, and does not
retry or suppress anything — it makes the existing diagnosis reachable.
Refs #7005 , #7010
2026-09-02 18:22:20 +08:00
cxymds
1bbfa71b11
fix(ecstore): preserve buckets after pool expansion ( #7040 )
...
* fix(ecstore): preserve buckets after pool expansion
* fix(ecstore): scope bucket operations by erasure set
* fix(ecstore): preserve bucket metadata load errors
2026-09-02 15:33:55 +08:00
cxymds
b422d1fea9
fix(ecstore): make publication part matching bijective ( #7037 )
...
* fix(ecstore): make publication part matching bijective
* test(ecstore): persist opaque retry etag
2026-09-02 03:00:11 +00:00
cxymds
87bc9d14ea
fix(ecstore): defer zero-evidence delete diagnostics ( #7036 )
2026-09-02 01:50:39 +00:00
cxymds
397dbcf102
fix(ecstore): reconcile pending capacity before exact delete ( #7016 )
...
fix(ecstore): reconcile capacity before exact delete
2026-09-02 00:25:55 +08:00
houseme
5720c5c748
fix(ecstore): bootstrap verified MinIO adoption metadata ( #7020 )
2026-09-01 23:15:41 +08:00
唐小鸭
43450df589
fix(ecstore): keep degraded objects listable when drives are offline ( #7010 )
2026-09-01 20:16:57 +08:00
cxymds
6e26769265
fix(ecstore): make transitioned cleanup crash-safe ( #6978 )
...
* fix(ecstore): fence transitioned object cleanup
* fix(ecstore): address ILM recovery review findings
* fix(ecstore): complete crash-safe tier cleanup recovery
* test(ecstore): avoid typo false positive
* fix(ecstore): stabilize decommission error buckets
* fix(ecstore): stabilize transition delete validation
* fix(ecstore): resume authorized tier delete dispatch
* fix(ecstore): satisfy feature clippy
2026-09-01 19:09:22 +08:00
cxymds
03aecc5c3e
fix(heal): avoid pool metadata lock recursion ( #6991 )
2026-09-01 18:31:21 +08:00
Zhengchao An
23ab078c56
fix(ecstore): reclaim stale object prefixes ( #6974 )
2026-09-01 10:09:00 +00:00
Zhengchao An
14a77f9d79
fix(ecstore): supplement split latest listings ( #6977 )
2026-09-01 07:08:25 +08:00
houseme
7541bb2c5d
fix(ecstore): stabilize decommission capacity retries ( #6959 )
...
* fix(heal): retry unavailable recreate targets
* fix(heal): refresh put-file epochs after target restart
* test(e2e): harden heal restart evidence
Co-Authored-By: heihutu <heihutu@gmail.com >
* test(e2e): cancel competing heal before restart
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(ecstore): complete decommission capacity recovery
* fix(ecstore): stabilize decommission capacity tests
Keep decommission test capacity snapshots deterministic across startup and mutation probes, serialize capacity-ledger entries during retries, and avoid reacquiring a multipart fence already covered by the outer migration fence.
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(ecstore): satisfy decommission test lint
Co-Authored-By: heihutu <heihutu@gmail.com >
* fix(ecstore): restore free-version decommission owner
Co-Authored-By: heihutu <heihutu@gmail.com >
---------
Co-authored-by: marshawcoco <marshawcoco@gmail.com >
Co-authored-by: heihutu <heihutu@gmail.com >
Co-authored-by: overtrue <anzhengchao@gmail.com >
2026-08-31 22:52:47 +08:00
Zhengchao An
9d4ccb7884
fix(ecstore): finalize decommission capacity recovery ( #6955 )
2026-08-31 18:09:09 +08:00
Zhengchao An
9a22cb85f3
fix(ecstore): complete decommission capacity recovery ( #6949 )
2026-08-31 16:53:14 +08:00
Zhengchao An
6c67086d0b
fix(ecstore): reserve decommission capacity safely ( #6917 )
2026-08-31 15:20:09 +08:00
houseme
c876df53f5
fix(ecstore): fence snapshot stream polls on lock loss ( #6930 )
...
Co-authored-by: heihutu <heihutu@gmail.com >
2026-08-31 02:26:24 +00:00
Zhengchao An
c4ac11d22e
fix(scanner): persist decommission catch-up debt ( #6922 )
2026-08-31 08:45:36 +08:00
houseme
602ed2cbcd
test(ecstore): add targeted refresh-loss harness ( #6924 )
...
Co-authored-by: heihutu <heihutu@gmail.com >
2026-08-31 08:45:04 +08:00
Zhengchao An
e6234d3714
test(ecstore): pin default bucket config bytes ( #6920 )
2026-08-31 00:03:07 +00:00
Zhengchao An
9945c67f7e
fix(ecstore): supervise decommission worker recovery ( #6908 )
2026-08-31 06:18:00 +08:00
houseme
489408c0b0
perf(ecstore): reuse prepared Select metadata ( #6911 )
...
Co-authored-by: heihutu <heihutu@gmail.com >
2026-08-30 20:41:24 +00:00
houseme
2f9c75d04f
perf(ecstore): reuse prepared metadata across pools ( #6889 )
...
Co-authored-by: heihutu <heihutu@gmail.com >
2026-08-30 16:15:10 +00:00