Files
rustfs/crates/object-data-cache/src/index.rs
T
houseme eebd16d8a4 feat(cache): add object data cache engine and app flow (#4187)
* feat(cache): add object data cache engine

* feat(cache): wire app-layer object cache flow

* refactor(cache): streamline app-layer cache flow

* refactor(cache): tighten cache flow internals

* refactor: address final clippy cleanup

* chore(deps): update quick-xml to 0.41.0

* feat(cache): wire object data cache env config

* fix(cache): gate materialize fill by cache plan

* chore(cache): add object data cache benchmark gate

* fix(cache): guard object cache fill size mismatches

* refactor(cache): streamline object cache body planning

* fix(cache): align object cache rollout config

* test(cache): cover buffered object cache benchmark

* test(cache): isolate object cache benchmark metrics

* test(cache): mark materialize rollout experimental

* test(cache): tighten object cache benchmark gate

* fix(cache): address review findings for object data cache

- singleflight: clean up leader entry on cancellation (Drop impl) so a dropped GET future can no longer wedge all subsequent fills for the same key; switch the fill map to a std Mutex and add a regression test

- adapter: honor RUSTFS_OBJECT_DATA_CACHE_ENABLE=true by defaulting to hit_only when no explicit mode is set (explicit mode still wins)

- planner: treat nil version UUIDs as "no value" per repo convention so unversioned objects key under the canonical "null" instead of fragmenting the key space

- multipart: invalidate the object cache on the quota-exceeded rollback delete after complete-multipart, closing a stale-cache window

- layering: move the disabled-cache fallback into app::context and drop the new infra->app layer-dependency baseline entry

* fix(cache): close invalidation races and drop full-cache scan on writes

- index: make identity-index insert/remove/prune atomic via starshard compute_if_present/compute_if_absent so concurrent fills can no longer drop each other's keys (lost keys made entries unreachable to invalidation until TTL); add a concurrency regression test

- fill: register the key in the identity index before the entry becomes visible in the cache and re-check the index afterwards, undoing the fill when an invalidation raced in between (new skipped_invalidation_race fill result)

- invalidate: with the index now authoritative, remove the full-cache iter() fallback that made every PUT/DELETE of a never-cached object O(total cache entries) (two scans per PUT, 2N per batch delete)

- materialize-fill: fail the GET instead of falling back to the partially consumed stream after a mid-read error (the fallback would send a body missing its prefix under a full-length Content-Length), and log the same size-mismatch warning as the sibling buffering paths

Co-Authored-By: heihutu <heihutu@gmail.com>

* test(storage): fix media-dependent buffer clamp expectation

test_concurrency_manager_multi_factor_strategy_buffer_clamp asserted media_cap.min(MI_B), but the implementation's final safety clamp is [32KiB, media_cap.max(MI_B)] — deliberately so a media cap above 1MiB (NVMe's 2MiB default) stays effective. The test only passed on machines detected as SSD/Unknown (cap == 1MiB) and failed on NVMe-backed CI runners with 2MiB != 1MiB. Assert the media cap itself, which is what the strategy actually guarantees on every environment.

Co-Authored-By: heihutu <heihutu@gmail.com>

* test(storage): format buffer clamp assertion

* chore(logging): update tier guardrail path

---------

Signed-off-by: houseme <housemecn@gmail.com>
Co-authored-by: cxymds <cxymds@gmail.com>
Co-authored-by: overtrue <anzhengchao@gmail.com>
Co-authored-by: heihutu <heihutu@gmail.com>
2026-07-03 18:11:14 +08:00

128 lines
4.0 KiB
Rust

// Copyright 2024 RustFS Team
//
// Licensed under the Apache License, Version 2.0 (the "License");
// you may not use this file except in compliance with the License.
// You may obtain a copy of the License at
//
// http://www.apache.org/licenses/LICENSE-2.0
//
// Unless required by applicable law or agreed to in writing, software
// distributed under the License is distributed on an "AS IS" BASIS,
// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
// See the License for the specific language governing permissions and
// limitations under the License.
use crate::key::ObjectDataCacheKey;
use crate::starshard_index::StarshardIdentityIndex;
/// Result of inserting a cache key into the identity index.
#[derive(Debug, Clone, PartialEq, Eq)]
pub enum ObjectDataCacheIndexInsertResult {
/// The key was inserted into the identity set.
Inserted,
/// The key was already tracked for the identity.
Duplicate,
/// The identity exceeded its configured key budget and was conservatively cleared.
Overflow {
/// Keys removed while clearing the identity.
cleared_keys: Vec<ObjectDataCacheKey>,
},
}
#[derive(Debug, Clone, Default, PartialEq, Eq)]
pub(crate) struct ObjectDataCacheKeySet {
keys: Vec<ObjectDataCacheKey>,
}
impl ObjectDataCacheKeySet {
pub(crate) fn insert(&mut self, key: ObjectDataCacheKey, max_keys: usize) -> ObjectDataCacheIndexInsertResult {
if self.keys.iter().any(|existing| existing == &key) {
return ObjectDataCacheIndexInsertResult::Duplicate;
}
if self.keys.len() >= max_keys {
let cleared_keys = self.drain();
return ObjectDataCacheIndexInsertResult::Overflow { cleared_keys };
}
self.keys.push(key);
ObjectDataCacheIndexInsertResult::Inserted
}
pub(crate) fn remove_key(&mut self, key: &ObjectDataCacheKey) -> bool {
let original_len = self.keys.len();
self.keys.retain(|existing| existing != key);
original_len != self.keys.len()
}
pub(crate) fn retain<F>(&mut self, mut keep: F)
where
F: FnMut(&ObjectDataCacheKey) -> bool,
{
self.keys.retain(|key| keep(key));
}
pub(crate) fn is_empty(&self) -> bool {
self.keys.is_empty()
}
pub(crate) fn contains(&self, key: &ObjectDataCacheKey) -> bool {
self.keys.iter().any(|existing| existing == key)
}
#[cfg(test)]
pub(crate) fn len(&self) -> usize {
self.keys.len()
}
pub(crate) fn cloned(&self) -> Vec<ObjectDataCacheKey> {
self.keys.clone()
}
pub(crate) fn drain(&mut self) -> Vec<ObjectDataCacheKey> {
std::mem::take(&mut self.keys)
}
}
/// Public identity-index façade used by the cache backend.
pub type ObjectDataCacheIdentityIndex = StarshardIdentityIndex;
#[cfg(test)]
mod tests {
use super::{ObjectDataCacheIndexInsertResult, ObjectDataCacheKeySet};
use crate::key::{ObjectDataCacheBodyVariant, ObjectDataCacheKey};
fn make_key(id: &str) -> ObjectDataCacheKey {
ObjectDataCacheKey::new("bucket", "object", Some(id), "etag", 1, ObjectDataCacheBodyVariant::FullObjectPlainV1)
}
#[test]
fn key_set_deduplicates_existing_key() {
let mut set = ObjectDataCacheKeySet::default();
let key = make_key("v1");
let first = set.insert(key.clone(), 4);
let second = set.insert(key, 4);
assert_eq!(first, ObjectDataCacheIndexInsertResult::Inserted);
assert_eq!(second, ObjectDataCacheIndexInsertResult::Duplicate);
assert_eq!(set.len(), 1);
}
#[test]
fn key_set_overflow_clears_existing_keys() {
let mut set = ObjectDataCacheKeySet::default();
let key_a = make_key("v1");
let key_b = make_key("v2");
let _ = set.insert(key_a.clone(), 1);
let result = set.insert(key_b, 1);
assert!(matches!(
result,
ObjectDataCacheIndexInsertResult::Overflow { cleared_keys } if cleared_keys == vec![key_a]
));
assert!(set.is_empty());
}
}