fix(ecstore): invalidate metadata cache after ILM transition persists (#4951)

A duplicate transition task admitted after the winner released its
in-flight claim (#4839) re-reads the version before uploading, but on
unversioned buckets that read could hit a stale pre-transition entry in
the 2s-TTL GET metadata cache: transition_object never invalidated the
cache after delete_object_version persisted transition_status=complete
and freed the local data. The stale hit defeated the TRANSITION_COMPLETE
early-return, so the duplicate streamed the already-deleted local data
to the remote tier (NotFound reader errors + rejected duplicate tier
PUT with UnexpectedContent).

Invalidate the cache right after the transitioned metadata is persisted,
matching the other metadata-mutating paths, and add a regression test
that runs a duplicate transition against an already-transitioned version
and asserts no second tier upload and unchanged remote object metadata.

Fixes #4827
This commit is contained in:
Zhengchao An
2026-07-18 15:03:27 +08:00
committed by GitHub
parent 2c113542f8
commit 04bfd48eb1
3 changed files with 104 additions and 2 deletions
@@ -2354,6 +2354,14 @@ impl crate::storage_api_contracts::object::ObjectOperations for SetDisks {
self.record_capacity_scope_if_needed(opts.capacity_scope_token, &disks);
}
// delete_object_version persisted transition_status=complete and freed the
// local data, but does not touch the GET metadata cache. Drop any cached
// pre-transition entry so a late duplicate transition task (or a plain GET)
// re-reads the fresh state; a stale hit here defeats the TRANSITION_COMPLETE
// early-return above and streams the already-deleted local data to the
// remote tier again (rustfs/rustfs#4827).
self.invalidate_get_object_metadata_cache(bucket, object).await;
for disk in disks.iter() {
if let Some(disk) = disk {
continue;