fix(tier): recover (#3182)

* fix(tier): stop sending nil/garbage versionId to warm backend S3

Three bugs caused NoSuchVersion errors when reading tiered objects:

1. warm_backend_s3sdk: GET and DELETE ignored rv/range opts entirely —
   fixed to forward version_id and byte-range to the SDK request.

2. version.rs (MetaObject + MetaDeleteMarker): transition_version_id was
   parsed with unwrap_or_default(), turning invalid/wrong-length bytes
   into Uuid::nil(). The nil UUID was then serialized and sent as
   ?versionId=00000000-... to the tier backend -> NoSuchVersion.
   Fixed: .and_then(.ok()).filter(!is_nil()) so only valid non-nil UUIDs
   are forwarded as versionId.

3. bucket_lifecycle_ops: add debug/error logs in
   get_transitioned_object_reader to record tier, tier_object, and
   tier_version_id before and on failure of the tier GET.

Also adds tier transition fields to dump_fileinfo example for offline
xl.meta inspection, and fixes Docker build (cargo path + entrypoint).
Adds CLAUDE.md with tier architecture and debugging notes.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* more fixes for versionId

* Potential fix for pull request finding

Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
Signed-off-by: Marcelo Bartsch <marcelo@bartsch.cl>

* remove branch

* Add tests and fix cargo path, add load to build-docker

* update documentation (CLAUDE.md)

* more fixes for recover

* More fixes to ILM recover

* final fix

* chore: add missing-shard first-scene diagnostics (#3213)

chore(ecstore): add missing-shard first-scene diagnostics

Log rename_data quorum context behind RUSTFS_ISSUE3031_DIAG_ENABLE so partial-disk success can be correlated with later missing shard reads.

Also log put_object commit success and tmp cleanup boundaries to capture when successful quorum writes are followed by tmp_dir cleanup.

* fix test anmd fmt

* fix cargo path
fix test

* fix(tier): format copy_object self-copy guard

---------

Signed-off-by: Marcelo Bartsch <marcelo@bartsch.cl>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: 安正超 <anzhengchao@gmail.com>
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
Co-authored-by: houseme <housemecn@gmail.com>
Co-authored-by: cxymds <Cxymds@qq.com>
Co-authored-by: loverustfs <hello@rustfs.com>
This commit is contained in:
Marcelo Bartsch
2026-06-07 11:38:28 +02:00
committed by GitHub
parent 069d1e5a75
commit f00898d070
10 changed files with 506 additions and 34 deletions
+66 -3
View File
@@ -1635,10 +1635,17 @@ impl ObjectOperations for SetDisks {
src_opts: &ObjectOptions,
dst_opts: &ObjectOptions,
) -> Result<ObjectInfo> {
// FIXME: TODO:
if !src_info.metadata_only {
return Err(StorageError::NotImplemented);
if path_join_buf(&[src_bucket, src_object]) != path_join_buf(&[dst_bucket, dst_object]) {
return Err(StorageError::NotImplemented);
}
// Self-copy with a data reader: write tier data back locally (de-tiering).
// Handles `mc cp --storage-class STANDARD obj obj` on a transitioned object.
if let Some(mut put_reader) = src_info.put_object_reader.take() {
return self.put_object(dst_bucket, dst_object, &mut put_reader, dst_opts).await;
}
// Same-key tiered copy without a pre-fetched reader: fall through to the metadata
// path so the caller gets a disk/quorum error rather than NotImplemented.
}
if path_join_buf(&[src_bucket, src_object]) != path_join_buf(&[dst_bucket, dst_object]) {
@@ -4727,6 +4734,7 @@ pub fn is_infrequent_access_class(storage_class: &str) -> bool {
#[cfg(test)]
mod tests {
use super::*;
use crate::bucket::lifecycle::bucket_lifecycle_ops::TransitionedObject;
use crate::disk::CHECK_PART_UNKNOWN;
use crate::disk::CHECK_PART_VOLUME_NOT_FOUND;
use crate::disk::RUSTFS_META_BUCKET;
@@ -6735,4 +6743,59 @@ mod tests {
assert!(!is_infrequent_access_class(storageclass::DEEP_ARCHIVE));
assert!(!is_infrequent_access_class(storageclass::EXPRESS_ONEZONE));
}
// Regression test: `mc cp --storage-class STANDARD` on a tiered object (self-copy) must not
// return NotImplemented. When the source object is tiered (transitioned_object.tier is
// non-empty) the usecase layer in object_usecase.rs intentionally leaves metadata_only=false
// so that the full copy path is taken. SetDisks::copy_object must therefore accept a
// same-bucket/same-key call even when metadata_only=false.
//
// Currently this test FAILS because the guard at set_disk.rs:1579 unconditionally rejects
// !metadata_only with StorageError::NotImplemented. Once the fix is applied the test will
// pass (or progress further through the copy path before failing on missing disk data).
#[tokio::test(flavor = "multi_thread")]
#[serial]
async fn copy_object_tiered_self_copy_does_not_return_not_implemented() {
let _setup_type_guard = SetupTypeGuard::switch_to(SetupType::Erasure).await;
let set_disks = make_test_set_disks(vec![Arc::new(LocalClient::with_manager(Arc::new(
rustfs_lock::GlobalLockManager::new(),
)))])
.await;
// Simulate a tiered object: metadata_only is false (set_disk must handle the full copy),
// and transitioned_object.tier is non-empty (the object lives on a remote tier).
let mut src_info = ObjectInfo {
metadata_only: false,
transitioned_object: TransitionedObject {
tier: "NEXTCLOUD".to_string(),
..Default::default()
},
..Default::default()
};
let result = set_disks
.copy_object(
"bucket",
"object",
"bucket",
"object",
&mut src_info,
&ObjectOptions::default(),
&ObjectOptions {
no_lock: true,
..Default::default()
},
)
.await;
// The copy must not be rejected with NotImplemented. Any other outcome (Ok or a
// different error such as missing-disk / quorum) is acceptable here.
if let Err(ref err) = result {
assert!(
!matches!(err, StorageError::NotImplemented),
"tiered self-copy returned NotImplemented — copy_object must handle \
metadata_only=false for same-key copies of tiered objects, got: {err}"
);
}
}
}