Compare commits

..

30 Commits

Author SHA1 Message Date
Henry Guo 9e6e02ea09 fix(table-catalog): assign fresh schema IDs on create (#6146)
* fix(table-catalog): assign fresh schema IDs on create

* fix(table-catalog): accept negative create schema IDs

---------

Co-authored-by: Henry Guo <marshawcoco@users.noreply.github.com>
2026-08-17 01:06:25 +08:00
houseme 39274fc37c feat(ecstore): default bounded metadata fanout (#6156)
Co-authored-by: heihutu <heihutu@gmail.com>
2026-08-17 00:56:20 +08:00
houseme 33eff4c3c4 test(ecstore): add metadata slow-tail fault hook (#6150)
Add a diagnostic metadata-only read_version delay hook for GET data-read fanout so bounded/default behavior can be compared under controlled slow-tail metadata responses.

Co-authored-by: heihutu <heihutu@gmail.com>
2026-08-16 23:36:46 +08:00
Zhengchao An a2f16aa066 test(tier): pin the compressed transitioned read against its stored bytes (#6151)
#6107 routed the transitioned read through the object's own ReadPlan so a tiered SSE object stops serving ciphertext. Compression rides that same plan and got fixed with it, but nothing pins it: revert the routing and a compressed object that ILM moved to a warm tier returns its stored (compressed) bytes under the compressed size, with every existing test still green.

The gap is easy to reopen because transition genuinely uploads the stored representation — the upload side is correct and the read side is the only place that can decode it. These tests state that contract at the boundary where it broke.

Four cases, all through SetDisks::get_object_reader against a mock warm tier:

- a full GET of a transitioned compressed object returns the plaintext and publishes the plaintext size (the test also asserts the remote copy holds the compressed bytes, so it fails loudly if the upload side ever changes instead);
- a ranged GET returns that plaintext slice, with the range deliberately starting past the compressed size so a range still measured in stored coordinates cannot produce it;
- a restore read still receives the stored bytes under the stored size — restore_request_active holds it on the Plain branch, and decompressing there would write plaintext under compressed metadata;
- a plain transitioned object still reads back byte-identical, full and ranged.

Verified as guards, not decoration: forcing the tiered read back onto the Plain branch turns the two compressed tests red and leaves the plain and restore tests green.

What these do not pin, so the gap stays recorded rather than implied covered: the fixture carries no compression index, so part.index stays None and the plan's storage offset is always 0 — the compressed-offset translation itself is still untested, as are multipart compressed objects, partNumber reads, and the encrypted tiered read that #6107 targeted.
2026-08-16 15:27:59 +00:00
Zhengchao An 4c8b9f87e1 chore(ecstore): annotate the two dead fns the blanket removals missed (#6152) 2026-08-16 22:42:51 +08:00
Zhengchao An 3272730c13 fix(ecstore): silence two dead_code warnings left on main (#6153) 2026-08-16 22:42:45 +08:00
Zhengchao An 1862112d0c chore: drop the remaining product-code dead_code blankets (#6149) 2026-08-16 14:18:00 +00:00
Zhengchao An cd0ac02879 test(interop): let the MinIO fixture lab build from registry mirrors (#6148) 2026-08-16 14:05:27 +00:00
Zhengchao An 6cf9cf7bb5 chore(ecstore): drop the bucket dead_code blanket (#6147)
* chore(ecstore): drop the bucket dead_code blanket

The last blanket of the backlog#1823 burn-down, and the largest: 71 items across lifecycle, replication, metadata, quota, object lock and bucket utils. Four are deleted.

Deleted, all trivial:

- check_valid_object_name and check_valid_object_name_prefix, a pair that only calls into each other with no external caller. Worth stating plainly so nobody reads this as a validation gap: object names are validated through check_object_name_for_length_and_slash, which is live; this pair is a second, unwired entry point.
- DEFAULT_HEALTH_CHECK_RELOAD_DURATION, a lone unused constant.
- The LifecycleReplicationConfig alias, which orphaned a re-export in replication/mod.rs that goes with it.

Everything else is kept, in four groups, because the blanket here was hiding structure rather than rot:

Windows platform gating. WINDOWS_RESERVED_NAMES, the two reason constants and object_name_has_windows_incompatible_segment are called from inside the #[cfg(target_os = "windows")] block in check_object_name_for_length_and_slash (utils.rs:228-255), so they only read as dead on non-Windows hosts. As with the Linux gating in the disk root, this cannot be adjudicated locally: cargo check for both x86_64-pc-windows-msvc and x86_64-unknown-linux-gnu fails in the aws-lc-sys build script for want of a cross C toolchain. CI covers both.

Declared boundary surface. The *_boundary.rs and *_bridge.rs files carry the replication split plan's contracts, which scripts/check_architecture_migration_rules.sh pins through the EcstoreReplicationBoundaryImports section of the split-plan doc. Their unused items are declarations, not leftovers.

test-util seams. ConfigWriteLockProbe with install/wait_until_attempted follows the same pattern as the barriers in the services and set_disk roots.

MinIO-parity tier/lifecycle entry points that this port never wired: apply_lifecycle_action, get_transitioned_object_reader, recover_tier_free_versions, delete_object_from_remote_tier, abort_tier_delete_journal_entry and the replication pool's worker-management surface. These are complete, substantial machinery with no caller — the same shape as data_usage's local_snapshot feature. Removing them is a product decision, so they are made explicit here rather than deleted.

Verification, four lanes warning-free: default, --tests, --features rio-v2 --tests, --features test-util --tests. cargo nextest run -p rustfs-ecstore 4096 passed; clippy --lib --tests -D warnings clean; make pre-commit exit 0. Note that clippy is what caught the orphaned re-export above: cargo check and pre-commit both treat unused_imports as a warning.

Ref rustfs/backlog#1823 (step 2, final root).

* chore(ecstore): correct inaccurate dead_code reasons in the bucket root

Six items were labelled 'asserted by this file's tests' or as MinIO-parity
entry points while having no caller at all - free get_bucket_acl_config and
created_at only reach their own live methods (production goes through
created_at_in), BucketVersioningSys::get_in, utils::serialize_content and
ServiceType have no reference anywhere, and with_transition_queue_env_async
is an unused test fixture, not a tier entry point. Name what each one is so
the next reader does not assume coverage that is not there.

Ref rustfs/backlog#1823.
2026-08-16 21:39:04 +08:00
Zhengchao An f1f86ee9d0 chore(ecstore): drop the set_disk dead_code blanket (#6141)
* chore(ecstore): drop the set_disk dead_code blanket

Removing the blanket exposes 39 items; exactly one is deleted. The low share is a finding, not caution: unlike the disk root, where platform gating made local adjudication impossible, here the items were checked and nearly all of them are live.

Deleted: HealEntryResult, the only item with no reference anywhere.

What the checks turned up, in the order the warnings suggest deleting them:

SetDisks::rename_data looked like the head of a dead chain feeding into_legacy_tuple and RenameDataLegacyTuple. It is not: production goes through rename_data_owned, and rename_data itself has test callers at mod.rs:5809 and 5880. The chain below it is therefore live through the tests, and inferring "this is dead, so its callee is dead" would have removed three working items.

create_bitrot_readers_until_quorum, read_multiple_files and map_cleanup_join_result all have callers inside their files' test modules, so they only look dead in the lib target.

TransitionCommitBarrier and TransitionUploadedSaveProbe, with their install/wait_until_paused/release surfaces, are installed by tests behind #[cfg(all(test, feature = "test-util"))].

ctx.rs's SetDisksCtx accessors are the split seam left by the SetDisks god-object break-up (backlog#815).

heal_object_dir's two apparent references are comments, and they document an index-alignment contract that live code maintains for it, so they stay as they are.

Worth a maintainer decision: the metadata early-stop switch has a complete percentage-rollout facet — ENV_RUSTFS_GET_METADATA_EARLY_STOP_ROLLOUT_PCT, get_metadata_early_stop_rollout_pct and should_use_metadata_early_stop — with no caller, no test and no documentation, while its sibling enable flag is live. It is kept with an allow that says so rather than removed, since a rollout knob is a product call.

One placement note for anyone adding allows near heal code: check_logging_guardrails.sh requires #[instrument(level = "trace")] to sit immediately before async fn heal_object_dir, so the allow goes above the instrument attribute. Putting it between the two drops the guard's match count and fails the check.

Verification, four lanes warning-free: default, --tests, --features rio-v2 --tests, --features test-util --tests. cargo nextest run -p rustfs-ecstore 4096 passed; clippy --lib --tests -D warnings clean; make pre-commit exit 0.

Ref rustfs/backlog#1823 (step 2).

* chore(ecstore): fix duplicated and inaccurate dead_code reasons in set_disk

format_lock_error carried the same #[allow] twice. Five items in the
locking/heal roots were labelled 'asserted by this file's tests' while
having no reference at all - heal_object_dir's only two references are
comments, as this branch's own notes point out. Say what each item
actually is instead, so the next reader does not assume test coverage
that is not there.

Ref rustfs/backlog#1823.

* chore(ecstore): correct the bounded_spare_disk_index dead_code reason

The mod.rs copy is an unused test fixture, not something this module's
tests assert; the namesake that is exercised lives in the io_primitives
test module.

Ref rustfs/backlog#1823.
2026-08-16 21:38:46 +08:00
Zhengchao An 1eef0de003 chore(ecstore): drop the disk dead_code blanket (#6139)
* chore(ecstore): drop the disk dead_code blanket

Removing the blanket exposes 36 items in the lowest storage layer: 7 deleted, 29 kept with reasoned item-level allows. That is the smallest deletion share of this burn-down, and the reason is a verification limit rather than a judgement call.

disk/local.rs carries 141 `#[cfg(target_os = "linux")]` sites — the densest platform gating in the tree, because O_DIRECT and io_uring only exist there. The direct-I/O cluster (six ENV_RUSTFS_OBJECT_DIRECT_IO_* constants plus is_direct_io_read_enabled, is_direct_io_write_enabled, get_direct_io_read_threshold, direct_write_staging_capacity, direct_write_tail_split and DIRECT_WRITE_STAGING_BYTES) reads as dead on macOS purely because its production callers at local.rs:1766, 3114 and 4605 sit inside Linux-gated blocks. direct_write_staging_capacity even documents itself as "Platform-independent (no O_DIRECT), so it is unit-tested on any host".

Deleting those would leave every local check green — 4096 tests pass, clippy is clean, make pre-commit exits 0 — and break the Linux build in CI, because all four local lanes compile for aarch64-apple-darwin. Cross-checking locally is not available either: cargo check --target x86_64-unknown-linux-gnu fails in the aws-lc-sys build script for want of a Linux C cross-compiler. Their allows name the platform reason so the next reader on a non-Linux host does not repeat the investigation.

Deleted, all in files with no target_os gating at all (os.rs, disk_store.rs):

- HealthDiskCtxKey and HealthDiskCtxValue with its private log_success. Note that DiskHealthTracker::log_success is a different method of the same name and is live from cluster/rpc/peer_s3_client.rs and remote_disk.rs — the two have to be told apart by type, not by name.
- LocalDiskWrapper::new_with_health and check_id.
- os.rs file_exists and lock_destination_directory_for_path_access.

Kept with allows: DiskHealthTracker's set_faulty, mark_offline, waiting_count and last_success have test callers in remote_disk.rs, so they only look dead in the lib target. to_disk_error, remove_all and sync_dir_files are asserted by their own files' tests. The reclaim, mmap and path-cache field groups are written but never read back.

Placement follows the same rule as the earlier roots: per-method allows inside impl DiskHealthTracker and impl LocalDisk, since both are mostly live and a block-level allow would be a smaller version of the blanket this issue removes. Struct-level allows are used only where the warning covers that struct's own fields. The three cached_read_env! functions take their allow inside the macro invocation, before the fn line, because the macro forwards $(#[$meta:meta])* onto the generated item.

Verification, four lanes warning-free: default, --tests, --features rio-v2 --tests, --features test-util --tests. cargo nextest run -p rustfs-ecstore 4096 passed; clippy --lib --tests -D warnings clean; make pre-commit exit 0. The Linux lane is not covered locally and is left to CI.

Ref rustfs/backlog#1823 (step 2).

* chore(ecstore): correct two dead_code reasons in the disk root

check_valid_path and reject_symlink_components have no caller at all -
not even a test - so 'asserted by this file's tests' misreads them as
covered. Both are method wrappers over live free functions; say that
instead.

Ref rustfs/backlog#1823.
2026-08-16 21:38:37 +08:00
houseme a118d7e4fd perf(ecstore): enable inline data read early-stop by default (#6140)
* perf(ecstore): enable inline data read early-stop by default

Co-Authored-By: heihutu <heihutu@gmail.com>

* test(scanner): box large ILM transition flow future

Co-Authored-By: heihutu <heihutu@gmail.com>

* test(ecstore): align internal meta early-stop miss

Co-Authored-By: heihutu <heihutu@gmail.com>

---------

Co-authored-by: heihutu <heihutu@gmail.com>
2026-08-16 06:38:51 +00:00
Zhengchao An ed1bedf1fb fix(storage): skip table guards for multipart parts (#6143) 2026-08-16 14:34:46 +08:00
houseme 4392f94e1a chore(deps): update flake.lock (#6142)
Flake lock file updates:

• Updated input 'nixpkgs':
    'github:NixOS/nixpkgs/70ce234' (2026-08-06)
  → 'github:NixOS/nixpkgs/8be7bd0' (2026-08-14)
• Updated input 'rust-overlay':
    'github:oxalica/rust-overlay/6df7076' (2026-08-09)
  → 'github:oxalica/rust-overlay/b211ead' (2026-08-16)
2026-08-16 13:35:41 +08:00
Zhengchao An 81d7b7d07a chore(ecstore): drop the client dead_code blanket (#6138) 2026-08-16 03:20:52 +00:00
唐小鸭 e26668e62c fix(ecstore): mint bucket-target ARNs in the madmin arn:minio partition (#6128)
* test(ecstore): pin madmin-compatible ARN partition contract

Red-light evidence for backlog#1675 P1-7: madmin-go's ParseARN
hard-rejects any ARN that does not start with 'arn:minio:', while RustFS
generates and only accepts 'arn:rustfs:'. mc/madmin tooling therefore
cannot decode RustFS remote-target listings, and MinIO-era replication
configs are rejected as StaleTarget when re-registered. The new tests
pin the target contract (generate arn:minio:, parse both partitions,
reject unknown partitions) and fail against the current single-partition
gate.

* fix(ecstore): mint bucket-target ARNs in the madmin arn:minio partition

madmin-go's ParseARN hard-rejects any partition other than 'arn:minio:',
so native mc/madmin tooling could not decode RustFS remote-target
listings, and re-registering a MinIO-era replication config failed its
StaleTarget check against freshly minted arn:rustfs: targets
(backlog#1675 P1-7, route A).

- ARN Display now emits 'arn:minio:'; FromStr accepts a {minio, rustfs}
  partition whitelist (the legacy partition stays readable forever for
  persisted bucket-targets.json / replication configs). The whitelist is
  the only structural gate — BucketTargetType::from_str never fails —
  so it deliberately rejects foreign partitions such as arn:aws:.
- No data migration: every runtime match between targets, rules and
  stats keys is full-string equality, so existing arn:rustfs: targets
  keep matching their persisted rules; site replication already
  preserves MinIO-era ARNs on reconcile (pinned by existing tests).
- Rolling upgrade note: upgrade all cluster nodes before creating new
  remote targets — a not-yet-upgraded node rejects remove-remote-target
  for a freshly minted arn:minio: ARN with BucketRemoteArnInvalid.
- Out of scope: notification/SQS ARNs (crates/targets) keep the
  arn:rustfs:sqs: partition; they have their own compatibility story.
2026-08-16 10:28:45 +08:00
Henry Guo 8d3511c1b3 fix(heal): preserve timeout budget across retries (#6101)
Co-authored-by: Henry Guo <marshawcoco@users.noreply.github.com>
2026-08-16 10:28:11 +08:00
Henry Guo d172d05e86 fix(ecstore): overlap metacache reader deadlines (#6098)
* fix(ecstore): overlap metacache reader deadlines

* fix(admin): avoid span guards across awaits

---------

Co-authored-by: Henry Guo <marshawcoco@users.noreply.github.com>
2026-08-16 10:27:59 +08:00
Zhengchao An 0d86c50760 fix(ecstore): classify remote inline early-stop misses (#6136) 2026-08-16 09:50:16 +08:00
houseme 526d6f667e perf(ecstore): defer pending inline data shards (#6137) 2026-08-16 09:44:00 +08:00
唐小鸭 dcf3e4b9e8 fix(replication): transport and persist LWW timestamps for tag, retention, and legal hold (#6129)
* test(replication): pin missing LWW timestamp header transport

Red-light tests for the replication timestamp three-header contract:

- put_object_headers_carry_replication_timestamp_headers pins that
  PutObjectOptions::header() must emit the
  x-{rustfs,minio}-source-replication-{tagging,retention,legalhold}-timestamp
  headers when the internal timestamps are set (currently missing).
- test_put_opts_from_headers_gates_replication_timestamp_persistence_on_authorization
  and test_complete_multipart_opts_persist_replication_timestamps_when_authorized
  pin that an authorized replication PUT / multipart complete must persist
  the inbound timestamps into the internal metadata keys while unauthorized
  requests must not (currently never persisted).
- fake_s3_target journals the three timestamp headers per request
  (ReplicationTimestampHeaders on RequestRecord) so sender-side e2e
  assertions can observe what a real target receives; self-test included.

* fix(replication): transport and persist LWW timestamps for tag, retention, and legal hold

Active-active conflict resolution for concurrent tag/retention/legal-hold
edits needs the source's per-category modification times on both sides of
the wire; the three AdvancedPutOptions timestamp fields were dead and the
headers were neither sent nor parsed.

- Emit x-{rustfs,minio}-source-replication-{tagging,retention,legalhold}-
  timestamp from PutObjectOptions::header(); names and RFC3339 values
  interoperate with MinIO (minio-go constants.go, object-api-options.go),
  pinned by a header_compat wire-name test.
- Default the three AdvancedPutOptions timestamps to UNIX_EPOCH and skip
  epoch values in header(), so "never modified" is not sent as a
  modification made now.
- Parse the headers only on authorized replication PUTs and multipart
  completes, expose them as Option<OffsetDateTime> on ObjectOptions, and
  persist them into the dual-prefix internal metadata keys so the
  outbound pass (replication_target_boundary) reads the source's
  timestamps instead of the mod_time fallback.
- Record the local tagging timestamp in the PutObjectTagging and
  DeleteObjectTagging eval metadata, mirroring the object-lock handlers;
  without it the sender only ever had the mod_time fallback to offer.

Receiver-side LWW comparison (keep newer stored category metadata over a
stale inbound copy) is left as a TODO at the parse site.

* fix(replication): load the stored tagging timestamp independently of remaining tags

Review: DeleteObjectTagging persists the tagging-timestamp internal key
but leaves the object tagless, and the outbound mapper only loaded the
key inside the user_tags-nonempty branch — the deletion's LWW timestamp
stayed at the epoch and the header was omitted, so the deletion could
never win conflict resolution on the replica. The stored key is now
loaded unconditionally; the mod_time fallback still applies only while
tags exist (MinIO parity), and a tagless object without the key keeps
the epoch default (no header). Deletion-path regression test added.

* fix(storage): reserve replication transport names at metadata ingest

Second review round: a client PUT of
x-amz-meta-x-rustfs-source-replication-tagging-timestamp materialized
the bare transport key as stored user metadata. The outbound
replication header builder forwards user metadata verbatim on a
server-authorized request, so the receiver would persist the
attacker-chosen value as trusted internal LWW state — and for a
tagless object nothing later overwrites it.

The ingest namespacing guard now reserves the whole
x-rustfs-source- / x-minio-source- families (the new timestamps and
their siblings: source-mtime/-etag/-version-id/-replication-request),
folding forged keys back under x-amz-meta-. Forged-ingress regression
covers both prefixes and a sibling.

* fix(replication): harden timestamp replay

* fix(app): route retention helper through facade

---------

Co-authored-by: overtrue <anzhengchao@gmail.com>
2026-08-16 05:56:04 +08:00
唐小鸭 04b9c8fd36 fix(admin): stream madmin ReplicationMRF documents from /v3/replication/mrf (#6126)
* test(admin): pin madmin ReplicationMRF stream contract for /v3/replication/mrf

Red-light evidence for backlog#1675 P1-13 (mrf half): madmin's
BucketReplicationMRF decodes the response one ReplicationMRF document at
a time, so the current aggregate envelope decodes as a single phantom
row with an empty object in 'mc replicate backlog'. The new contract
tests assert the desired bare-document stream (exact madmin json tags,
empty body for an empty backlog) and fail against the current
render_mrf_backlog extraction, which preserves the envelope-only
behavior:

- mrf_stream_renders_bare_madmin_documents: envelope keys leak, no
  per-entry documents
- mrf_stream_renders_empty_body_for_no_entries: empty backlog still
  renders the envelope (phantom row)
- mrf_aggregate_envelope_retains_counters: PerObjectEntriesAvailable
  never advertises the enumerable stream

* fix(admin): stream madmin ReplicationMRF documents from /v3/replication/mrf

The mrf endpoint returned a single aggregate envelope, which madmin's
json.Decoder loop decoded as one phantom row (empty object) in
'mc replicate backlog' (backlog#1675 P1-13, mrf half; the diff half was
fixed in #5799 and this mirrors its pattern).

- Default response is now a bare stream of ReplicationMRF documents
  (exact madmin json tags; Size/TargetARNs as ignored extension keys)
  built from the durable backlog ledger; an empty backlog renders an
  empty body, so mc shows zero rows instead of a phantom row.
- The aggregate counter envelope moves behind ?aggregate=true (RustFS
  extension) and now advertises PerObjectEntriesAvailable whenever the
  durable backlog is readable.
- An unreadable backlog is signalled out-of-band via
  x-rustfs-replication-mrf-backlog-unavailable (mirrors the diff
  truncation header) plus a warn event, since the bare stream cannot
  carry source health.
- The madmin node parameter is accepted but documented as a no-op: the
  durable ledger is cluster-shared with no per-node attribution.
- Delete-marker purge entries fall back to the marker version id so
  those rows keep a version identity.

* fix(admin): fail the mrf stream request when the durable ledger is unreadable

Review: madmin only decodes the body of a 200, so the out-of-band
unavailability header was invisible to it and an unreadable ledger read
as a clean zero-row backlog. Stream mode now returns 503; aggregate
mode keeps the availability fields.

* fix(admin): gate, bound, and null-map the mrf stream

Second review round:

- Authorization: the default stream enumerates object names and version
  ids, which a metrics-only principal must not see — it now requires
  admin:ReplicationDiff (MinIO parity, route policy updated);
  ?aggregate=true carries no object identities and keeps
  admin:GetReplicationMetrics.
- The nil UUID is RustFS's in-memory null-version sentinel and now
  leaves as the S3 wire token 'null' instead of a zero UUID (a
  pre-versioning object scanned after versioning + existing-object
  replication can persist it into the ledger).
- The durable ledger is not bounded by the in-memory pending cap and
  the body is buffered before send; the stream now stops at 10,000
  documents and signals truncation via
  x-rustfs-replication-mrf-truncated (mirroring the diff endpoint)
  plus a warn event, instead of staging an unbounded body.

* fix(admin): reject truncated MRF streams

---------

Co-authored-by: Zhengchao An <anzhengchao@gmail.com>
2026-08-16 05:55:15 +08:00
唐小鸭 c1f66969d7 fix(site-replication): merge incoming ILM expiry documents instead of overwriting (#6130)
* test(site-replication): pin ILM expiry merge contract for incoming lc-config

Red-light evidence for backlog#1675 P1-1: the lc-config receiver
overwrites the whole local lifecycle config with whatever the peer
sends (and deletes it wholesale on peer delete), so an expiry-only
document erases the receiver's local tier/transition rules, and peer
transition rules get installed across sites. The new tests pin the
MinIO mergeWithCurrentLCConfig semantics plus RustFS hardening:

- incoming expiry documents merge with (never replace) local rules
- local transition sides are authoritative for same-id rules
- incoming transition fields are discarded at the trust boundary
- dropped expiry rules strip the expiry side but keep transitions;
  pure-expiry rules are removed
- delete merges with the empty set instead of dropping the config
- disabled rules survive; abort-mpu-only rules stay site-local
- deterministic order (idempotent re-delivery) and expiry_updated_at
  stamping for the staleness axis

All fail against the current overwrite implementation (identity
extraction of merge_incoming_lifecycle_config).

* fix(site-replication): merge incoming ILM expiry documents instead of overwriting

The lc-config receiver replaced the whole local lifecycle config with
the peer's document (and deleted it wholesale on peer delete), so an
expiry-only update erased the receiver's local tier/transition rules,
and a peer's transition rules were installed across sites
(backlog#1675 P1-1).

Receiver (apply_bucket_meta_item):
- lc-config now merges via merge_incoming_lifecycle_config, mirroring
  MinIO's mergeWithCurrentLCConfig with a trust-boundary hardening:
  incoming transition fields are discarded outright; the local
  transition side of a same-id rule is authoritative. A peer delete
  merges with the empty set — pure-expiry rules go away, transition
  rules survive with their expiry side cleared, and only an empty
  result deletes the config file.
- Staleness moves to the expiry axis (config.expiry_updated_at):
  lifecycle_config_updated_at also moves on local transition-only
  edits, which shadowed newer peer expiry updates.
- Receiver-side replicateILMExpiry gate, symmetric with the sender
  hook (previously any peer could install expiry rules while the
  option was off).
- Rule order is deterministic (local order, incoming-new appended), so
  re-delivering the same document is byte-stable and does not rewrite
  bucket metadata per broadcast.

Sender:
- Both admin choke points — the bucket-meta hook and the SRInfo bucket
  entry feeding bootstrap/repair and consistency views — now emit only
  the expiry subset (transition fields stripped, non-expiry rules
  dropped). MinIO receivers install incoming rules verbatim, so
  transition rules must never leave the site. An unparseable local
  config is forwarded unfiltered rather than degraded to a delete.

Not covered here (follow-up): a two-site e2e with a real tier backend
to exercise transition-rule preservation end to end; receiver-side
validate_transition_tier for merged configs.

* fix(site-replication): close ILM merge review findings

Adversarial review of the lc-config merge surfaced four real defects,
all fixed here:

- Deletion tombstone regression: with the staleness axis moved to the
  in-config expiry_updated_at, a deleted lifecycle config fell back to
  UNIX_EPOCH and any delayed stale broadcast could resurrect deleted
  expiry rules. The axis now falls back to the whole-config write time
  (which survives deletion in bucket metadata as the deletion's lower
  bound), also covering legacy configs that predate the axis field.
- MinIO zero-rule documents: MinIO's delete tombstone / transition-only
  state marshals a lifecycle document with no <Rule>, which the strict
  s3s deserializer rejects — the receiver now recognizes it as the 'no
  expiry rules here' statement (delete semantics) instead of erroring
  on every MinIO heal pass.
- Inflated expiry axis at the sender: PutBucketLifecycle stamped
  expiry_updated_at unconditionally, so a transition-only edit advanced
  the axis and let this site's stale expiry subset shadow and roll back
  newer peer expiry edits fleet-wide. The stamp is now conditional
  (expiry subset present before or after the edit, MinIO parity), the
  hook item travels with the config's expiry axis (UNIX_EPOCH when the
  site has none), and the SRInfo bucket entry feeds bootstrap/repair
  the same axis instead of the whole-config write time.
- Del-marker parity: MinIO's CloneNonTransition never emits del-marker
  or abort-mpu fields, so treating del_marker_expiration as traveling
  expiry let a MinIO broadcast delete this site's del-marker-only
  rules. Both fields are now site-local on every edge: stripped from
  outbound subsets and inbound rules, restored from the local side on
  same-id merges, and never a deletion criterion.

Receiver-side validation of merged configs (object-lock / tier
constraints, MinIO runs finalLcCfg.Validate) remains a follow-up.

* fix(site-replication): close the second ILM review round

- Missed-delete repair: a deleted expiry state now travels through
  bootstrap/repair as an explicit timestamped lc-config delete item
  (lifecycle_expiry_statement distinguishes deletion — whole-config
  write time advanced past the created backfill — from never-configured
  buckets and from transition-only configs without an expiry axis,
  which say nothing). A peer that missed the live delete converges on
  repair; the receiver's staleness guard protects newer peer state.
- Strict tombstone recognition: only a well-delimited zero-rule
  <LifecycleConfiguration> document maps to delete semantics; truncated
  or foreign payloads that fail the strict deserializer are rejected
  instead of being treated as a delete that erases local expiry rules.
- Staleness fallback axis narrowed: the whole-config write time is used
  only for deleted or legacy-with-expiry state. A present
  transition-only config without an expiry axis compares at epoch — its
  whole-config time moves on transition edits and must not shadow or
  block independent peer expiry updates and same-timestamp repairs.

* fix(site-replication): validate tombstone children structurally

Second review round: a well-delimited root could still smuggle
malformed content — e.g. <LifecycleConfiguration><ExpiryUpdatedAt>
</LifecycleConfiguration> passed the no-<Rule check and was applied as
a delete. The tombstone body must now be a sequence of well-formed
simple children (matching open/close or self-closing, no nested markup,
no stray text, none named Rule); anything else surfaces InvalidRequest.
Malformed-child cases pinned in the recognition test.

* fix(site-replication): serialize lifecycle merges

---------

Co-authored-by: overtrue <anzhengchao@gmail.com>
2026-08-16 05:34:17 +08:00
唐小鸭 cfa9276fad fix(admin): serialize replication metrics in minio-go wire shapes (#6127)
* test(admin): pin minio-go Metrics/MetricsV2 wire contract for replication metrics

Red-light evidence for backlog#1675 P1-11: ?replication-metrics[=2]
serializes the internal snake_case BucketStats family straight onto the
wire, while minio-go's replication.Metrics/MetricsV2 expect camelCase
tags (currStats/queueStats/replicaCount/queued/...). Go's decoder is
case-insensitive but does not ignore underscores, so 'mc replicate
status' shows all zeros without any error. The rewritten snapshot tests
assert the minio-go tags (plus a synthesized queueStats node — the
aggregation path leaves queue_stats.nodes empty today) and fail against
the current pass-through serialization.

* fix(admin): serialize replication metrics in minio-go wire shapes

?replication-metrics[=2] and the admin replicationmetrics endpoint
serialized the internal snake_case BucketStats family straight onto the
wire, so 'mc replicate status' decoded all zeros without any error
(backlog#1675 P1-11). The internal structs cannot be renamed: they are
the intra-cluster peer-RPC wire format (rmp_serde to_vec_named in
node_service.rs), pinned by a new regression test.

- New admin/replication_metrics_wire.rs: Serialize-only projections onto
  minio-go replication.Metrics (v1 body, currStats) and MetricsV2
  (uptime/currStats/queueStats/downtimeInfo) with the exact json tags;
  per-target failed becomes the TimedErrStats envelope fed from the
  FailStats rolling window; the queue peak is dual-emitted as max
  (MinIO server tag) and peak (minio-go tag).
- queueStats synthesizes one node from the bucket queue snapshot — the
  aggregation path leaves queue_stats.nodes empty, and mc treats an
  empty node list as 'no data' — and carries transfer summaries
  (Large/Small/Total) derived from the per-target xfer rates.
- Both endpoints share the DTOs; source-health extension keys
  (provider_available/cluster_complete/...) ride along and are ignored
  by Go decoders.
- Widen the ecstore replication_stats_boundary re-exports
  (BucketReplicationStat/InQueueMetric/XferStats) so the admin facade
  chain can name the projected types.

* fix(replication): carry failure rolling windows through cluster aggregation

Review: both metrics endpoints aggregate first, and FailStats::merge
dropped the process-local samples (which also never cross the peer-RPC
wire — serde-skipped), so lastMinute/lastHour serialized as zero right
after a failure while totals was nonzero.

- FailStats gains serializable last_minute/last_hour window snapshots
  (serde default: old nodes read zeros, new fields are ignored by old
  decoders), recomputed on every add_size and re-stamped at the
  per-node collection point (get_latest_replication_stats), and summed
  by merge.
- The wire DTO takes the component-wise max of the live samples and the
  snapshot, so both the single-node and the aggregated path report the
  window.
- Regression test drives a stat through rmp round trip + merge before
  serialization, as requested.

Also restore the #[allow(dead_code)] attribute to route_policy — the
new module declaration had been inserted between the attribute and its
item, which broke the -D warnings CI lanes.

* fix(replication): bin transfer summaries at 128 MiB and keep window refresh off the hot path

Second review round:

- update_xfer_rate split at 1 MiB while the minio-go transferSummary
  labels (and RustFS's own worker-pool split) mean >= 128 MiB for
  Large, so a 2 MiB replication reported under Large with Small stuck
  at zero. The producer now bins on MIN_LARGE_OBJ_SIZE; a MetricsV2
  assertion covers 2 MiB / 127 MiB / exactly 128 MiB.
- add_size no longer recomputes the rolling windows: two full
  one-hour-deque scans per failure under the bucket-stats write lock
  made failure bursts quadratic (30k events ~2.1s). The windows are
  stamped only at the collection point (get_latest_replication_stats,
  which serves both the local leg and the peer RPC); the aggregation
  regression now drives that path explicitly before the RPC round trip
  and merge.

* fix(replication): average transfer summaries

---------

Co-authored-by: overtrue <anzhengchao@gmail.com>
2026-08-16 05:27:23 +08:00
Henry Guo db8f55cb97 feat(table-catalog): finalize Iceberg REST behavior (#6072)
* feat(table-catalog): finalize Iceberg REST behavior

* fix(table-catalog): address REST finalization regressions

* test(table-catalog): expect REST commit conflicts

* test(table-catalog): avoid serialized view test deadlocks

* fix(table-catalog): adapt shared test backend

* fix(table-catalog): enforce Iceberg metadata invariants

* fix(table-catalog): preserve manifest length in test

* test(table-catalog): use valid metadata fixtures

* test(table-catalog): seed manifests before manifest lists

* fix(table-catalog): restore validation gates

---------

Co-authored-by: Henry Guo <marshawcoco@users.noreply.github.com>
Co-authored-by: overtrue <anzhengchao@gmail.com>
2026-08-16 03:05:09 +08:00
houseme 7f23a1ba91 feat(ecstore): report inline early-stop miss reasons (#6134)
Co-authored-by: heihutu <heihutu@gmail.com>
2026-08-15 17:18:26 +00:00
Henry Guo 1619c4be60 fix(scanner): add context to corrupt metadata logs (#6099)
Co-authored-by: Henry Guo <marshawcoco@users.noreply.github.com>
2026-08-15 21:32:05 +08:00
Zhengchao An 72fd7339c9 test(utils): allow ephemeral port reuse (#6122)
* test(utils): allow ephemeral port reuse

* test(kms): allow any ciphertext prefix
2026-08-15 08:32:10 +08:00
Zhengchao An 71e83aeec4 fix(ci): pin Docker images to release source (#6121) 2026-08-15 07:13:37 +08:00
唐小鸭 9138c24571 fix(site-replication): lift a rejoined site's restarted edit counter over stale marks (#6119)
fix(site-replication): lift a rejoined site's restarted edit counter over stale fence marks

A site removed while unreachable (unilateral removal: the receiver never
dropped it from its peer map, so parse_site_replication_state's load-time
mark pruning never fired) that later rejoins recreates its state object
and restarts edit_generation at zero. The receiver's surviving high-water
mark then silently fences out every stamped delivery from that origin —
peer edits and the add finalize fan-out alike are acked without applying
— until the restarted counter catches up.

Allocate the generation as a hybrid logical clock instead:
max(wall clock in unix nanoseconds, previous + 1), still inside the state
transaction under the distributed state-object lock. Every value a
lifetime hands out is capped by the wall clock at its own allocation, so
a recreated lifetime's first allocation exceeds them all and clears the
stale mark, while a pre-removal delivery still in flight stays below the
new floor and remains correctly fenced. previous+1 keeps allocations
strictly increasing across same-tick allocations and mid-lifetime clock
regressions.

Nothing changes on the wire or in the persisted schema: editGeneration
stays the single fence param and edit_generation the single counter
field, so pre-hybrid receivers get the fix as soon as the sender
upgrades, old binaries preserve the field across rolling up/downgrades,
and marks recorded by plain-counter receivers (small values) are cleared
by any wall-clock allocation. A clock that regresses across a
delete/recreate degrades to a fence that self-heals once real time
passes the previous lifetime's last allocation, and introduces no
rollback window beyond what the plain counter already had.

An epoch-based design (editEpoch wire param + per-origin epoch marks)
was built first and rejected under adversarial review: old binaries
rewriting the state object drop the unknown epoch fields, which both
disarms the fix mid-rolling-upgrade and — because epoch adoption lowers
the generation mark — reopens the pre-restart rollback the fence exists
to prevent; a backwards clock also fences an origin permanently instead
of self-healing. The hybrid clock has none of these modes.
2026-08-15 01:50:35 +08:00
126 changed files with 14301 additions and 2019 deletions
+25 -1
View File
@@ -94,6 +94,7 @@ jobs:
short_sha: ${{ steps.check.outputs.short_sha }} short_sha: ${{ steps.check.outputs.short_sha }}
is_prerelease: ${{ steps.check.outputs.is_prerelease }} is_prerelease: ${{ steps.check.outputs.is_prerelease }}
create_latest: ${{ steps.check.outputs.create_latest }} create_latest: ${{ steps.check.outputs.create_latest }}
source_ref: ${{ steps.check.outputs.source_ref }}
steps: steps:
- name: Checkout repository - name: Checkout repository
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7 uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
@@ -118,6 +119,7 @@ jobs:
short_sha="" short_sha=""
is_prerelease=false is_prerelease=false
create_latest=false create_latest=false
source_ref="$GITHUB_SHA"
if [[ "${{ github.event_name }}" == "workflow_run" ]]; then if [[ "${{ github.event_name }}" == "workflow_run" ]]; then
# Triggered by build workflow completion # Triggered by build workflow completion
@@ -137,6 +139,7 @@ jobs:
# Extract version info from commit message or use commit SHA # Extract version info from commit message or use commit SHA
# Use Git to generate consistent short SHA (ensures uniqueness like build.yml) # Use Git to generate consistent short SHA (ensures uniqueness like build.yml)
short_sha=$(git rev-parse --short "$HEAD_SHA") short_sha=$(git rev-parse --short "$HEAD_SHA")
source_ref="$HEAD_SHA"
# Determine build type based on triggering workflow event and ref # Determine build type based on triggering workflow event and ref
triggering_event="$TRIGGERING_EVENT" triggering_event="$TRIGGERING_EVENT"
@@ -261,6 +264,23 @@ jobs:
echo "⚠️ Only release versions (latest, v1.0.0, 1.0.0) and prereleases (v1.0.0-alpha1, 1.0.0-beta2) are supported" echo "⚠️ Only release versions (latest, v1.0.0, 1.0.0) and prereleases (v1.0.0-alpha1, 1.0.0-beta2) are supported"
;; ;;
esac esac
if [[ "$should_build" == true && "$input_version" != "latest" ]]; then
tag_ref="refs/tags/$input_version"
if ! git ls-remote --exit-code origin "$tag_ref" >/dev/null 2>&1; then
if [[ "$input_version" == v* ]]; then
tag_ref="refs/tags/${input_version#v}"
else
tag_ref="refs/tags/v$input_version"
fi
fi
if ! git ls-remote --exit-code origin "$tag_ref" >/dev/null 2>&1; then
echo "❌ Release tag not found for Docker build: $input_version"
exit 1
fi
source_ref="$tag_ref"
fi
fi fi
{ {
@@ -271,6 +291,7 @@ jobs:
echo "short_sha=$short_sha" echo "short_sha=$short_sha"
echo "is_prerelease=$is_prerelease" echo "is_prerelease=$is_prerelease"
echo "create_latest=$create_latest" echo "create_latest=$create_latest"
echo "source_ref=$source_ref"
} >> "$GITHUB_OUTPUT" } >> "$GITHUB_OUTPUT"
echo "🐳 Docker Build Summary:" echo "🐳 Docker Build Summary:"
@@ -281,6 +302,7 @@ jobs:
echo " - Short SHA: $short_sha" echo " - Short SHA: $short_sha"
echo " - Is prerelease: $is_prerelease" echo " - Is prerelease: $is_prerelease"
echo " - Create latest: $create_latest" echo " - Create latest: $create_latest"
echo " - Source ref: $source_ref"
# Build multi-arch Docker images # Build multi-arch Docker images
# Strategy: Build images using pre-built binaries from dl.rustfs.com # Strategy: Build images using pre-built binaries from dl.rustfs.com
@@ -308,6 +330,7 @@ jobs:
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7 uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
with: with:
persist-credentials: false persist-credentials: false
ref: ${{ needs.build-check.outputs.source_ref }}
- name: Login to Docker Hub - name: Login to Docker Hub
uses: docker/login-action@c94ce9fb468520275223c153574b00df6fe4bcc9 # v3 uses: docker/login-action@c94ce9fb468520275223c153574b00df6fe4bcc9 # v3
@@ -397,7 +420,8 @@ jobs:
LABELS="org.opencontainers.image.title=RustFS" LABELS="org.opencontainers.image.title=RustFS"
LABELS="$LABELS,org.opencontainers.image.description=RustFS distributed object storage system" LABELS="$LABELS,org.opencontainers.image.description=RustFS distributed object storage system"
LABELS="$LABELS,org.opencontainers.image.version=$VERSION" LABELS="$LABELS,org.opencontainers.image.version=$VERSION"
LABELS="$LABELS,org.opencontainers.image.revision=${{ github.sha }}" SOURCE_REVISION="$(git rev-parse HEAD)"
LABELS="$LABELS,org.opencontainers.image.revision=$SOURCE_REVISION"
LABELS="$LABELS,org.opencontainers.image.source=${{ github.server_url }}/${{ github.repository }}" LABELS="$LABELS,org.opencontainers.image.source=${{ github.server_url }}/${{ github.repository }}"
LABELS="$LABELS,org.opencontainers.image.created=$(date -u +'%Y-%m-%dT%H:%M:%SZ')" LABELS="$LABELS,org.opencontainers.image.created=$(date -u +'%Y-%m-%dT%H:%M:%SZ')"
LABELS="$LABELS,org.opencontainers.image.build-type=$BUILD_TYPE" LABELS="$LABELS,org.opencontainers.image.build-type=$BUILD_TYPE"
Generated
+4
View File
@@ -278,6 +278,7 @@ checksum = "312c1ea69e5fe9966e0029fb95aca8790100b85aff4f0d3b00a9337c74069a9c"
dependencies = [ dependencies = [
"bigdecimal", "bigdecimal",
"bon", "bon",
"crc32fast",
"digest 0.11.3", "digest 0.11.3",
"log", "log",
"miniz_oxide 0.9.1", "miniz_oxide 0.9.1",
@@ -289,9 +290,11 @@ dependencies = [
"serde", "serde",
"serde_bytes", "serde_bytes",
"serde_json", "serde_json",
"snap",
"strum", "strum",
"thiserror 2.0.20", "thiserror 2.0.20",
"uuid", "uuid",
"zstd",
] ]
[[package]] [[package]]
@@ -9200,6 +9203,7 @@ dependencies = [
"serial_test", "serial_test",
"sha2 0.11.0", "sha2 0.11.0",
"shadow-rs", "shadow-rs",
"snap",
"socket2", "socket2",
"subtle", "subtle",
"sysinfo", "sysinfo",
+1 -1
View File
@@ -171,7 +171,7 @@ tower = { version = "0.5.3" }
tower-http = { version = "0.7.0" } tower-http = { version = "0.7.0" }
# Serialization and Data Formats # Serialization and Data Formats
apache-avro = "0.22.0" apache-avro = { version = "0.22.0", features = ["snappy", "zstandard"] }
bytes = { version = "1.12.1" } bytes = { version = "1.12.1" }
bytesize = "2.7.0" bytesize = "2.7.0"
byteorder = "1.5.0" byteorder = "1.5.0"
-1
View File
@@ -11,7 +11,6 @@
// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. // WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
// See the License for the specific language governing permissions and // See the License for the specific language governing permissions and
// limitations under the License. // limitations under the License.
#![allow(dead_code)]
use base64_simd::STANDARD; use base64_simd::STANDARD;
+91 -1
View File
@@ -76,6 +76,18 @@ const SOURCE_MTIME_HEADERS: [&str; 2] = ["x-rustfs-source-mtime", "x-minio-sourc
const SOURCE_REPLICATION_REQUEST_HEADERS: [&str; 2] = const SOURCE_REPLICATION_REQUEST_HEADERS: [&str; 2] =
["x-rustfs-source-replication-request", "x-minio-source-replication-request"]; ["x-rustfs-source-replication-request", "x-minio-source-replication-request"];
const SOURCE_ETAG_HEADERS: [&str; 2] = ["x-rustfs-source-etag", "x-minio-source-etag"]; const SOURCE_ETAG_HEADERS: [&str; 2] = ["x-rustfs-source-etag", "x-minio-source-etag"];
const SOURCE_TAGGING_TIMESTAMP_HEADERS: [&str; 2] = [
"x-rustfs-source-replication-tagging-timestamp",
"x-minio-source-replication-tagging-timestamp",
];
const SOURCE_RETENTION_TIMESTAMP_HEADERS: [&str; 2] = [
"x-rustfs-source-replication-retention-timestamp",
"x-minio-source-replication-retention-timestamp",
];
const SOURCE_LEGALHOLD_TIMESTAMP_HEADERS: [&str; 2] = [
"x-rustfs-source-replication-legalhold-timestamp",
"x-minio-source-replication-legalhold-timestamp",
];
const RESERVED_BUCKET_PREFIXES: [&str; 3] = ["xn--", "sthree-", "amzn-s3-demo-"]; const RESERVED_BUCKET_PREFIXES: [&str; 3] = ["xn--", "sthree-", "amzn-s3-demo-"];
const RESERVED_BUCKET_SUFFIXES: [&str; 6] = ["-s3alias", "--ol-s3", ".mrap", "--x-s3", "--table-s3", "-an"]; const RESERVED_BUCKET_SUFFIXES: [&str; 6] = ["-s3alias", "--ol-s3", ".mrap", "--x-s3", "--table-s3", "-an"];
@@ -118,6 +130,25 @@ pub enum FaultAction {
WrongEtag, WrongEtag,
} }
/// Replication LWW timestamp headers observed on a request, journaled so
/// sender-side tests can assert what a real target would receive.
#[derive(Debug, Clone, Default, PartialEq, Eq)]
pub struct ReplicationTimestampHeaders {
pub tagging: Option<String>,
pub retention: Option<String>,
pub legalhold: Option<String>,
}
impl ReplicationTimestampHeaders {
fn from_headers(headers: &HeaderMap) -> Self {
Self {
tagging: header_value(headers, &SOURCE_TAGGING_TIMESTAMP_HEADERS).map(bounded_journal_value),
retention: header_value(headers, &SOURCE_RETENTION_TIMESTAMP_HEADERS).map(bounded_journal_value),
legalhold: header_value(headers, &SOURCE_LEGALHOLD_TIMESTAMP_HEADERS).map(bounded_journal_value),
}
}
}
/// Credential-free request metadata retained for deterministic assertions. /// Credential-free request metadata retained for deterministic assertions.
#[derive(Debug, Clone, PartialEq, Eq)] #[derive(Debug, Clone, PartialEq, Eq)]
pub struct RequestRecord { pub struct RequestRecord {
@@ -131,6 +162,7 @@ pub struct RequestRecord {
pub part_number: Option<i32>, pub part_number: Option<i32>,
pub content_length: Option<u64>, pub content_length: Option<u64>,
pub consumed_bytes: Option<usize>, pub consumed_bytes: Option<usize>,
pub replication_timestamps: ReplicationTimestampHeaders,
pub fault: Option<FaultAction>, pub fault: Option<FaultAction>,
} }
@@ -536,7 +568,15 @@ impl S3Access for FaultAccess {
.get(CONTENT_LENGTH) .get(CONTENT_LENGTH)
.and_then(|value| value.to_str().ok()) .and_then(|value| value.to_str().ok())
.and_then(|value| value.parse().ok()); .and_then(|value| value.parse().ok());
let fault = record_request(&self.control, operation, context.method().clone(), parsed, content_length); let replication_timestamps = ReplicationTimestampHeaders::from_headers(context.headers());
let fault = record_request(
&self.control,
operation,
context.method().clone(),
parsed,
content_length,
replication_timestamps,
);
if let Some(RequestFault { if let Some(RequestFault {
action: FaultAction::Status(status), action: FaultAction::Status(status),
.. ..
@@ -589,6 +629,7 @@ fn record_request(
method: Method, method: Method,
parsed: ParsedRequest, parsed: ParsedRequest,
content_length: Option<u64>, content_length: Option<u64>,
replication_timestamps: ReplicationTimestampHeaders,
) -> Option<RequestFault> { ) -> Option<RequestFault> {
let mut state = lock(control); let mut state = lock(control);
let action = parsed let action = parsed
@@ -613,6 +654,7 @@ fn record_request(
part_number: parsed.part_number, part_number: parsed.part_number,
content_length, content_length,
consumed_bytes: None, consumed_bytes: None,
replication_timestamps,
fault: action.clone(), fault: action.clone(),
}); });
action.map(|action| RequestFault { sequence, action }) action.map(|action| RequestFault { sequence, action })
@@ -1699,6 +1741,52 @@ mod tests {
.await?) .await?)
} }
#[tokio::test]
async fn journals_replication_timestamp_headers() -> Result<(), BoxError> {
let target = FakeS3Target::start().await?;
target.create_bucket("target-bucket");
let client = client(&target);
client
.put_object()
.bucket("target-bucket")
.key("plain")
.body(ByteStream::from_static(b"plain"))
.send()
.await?;
client
.put_object()
.bucket("target-bucket")
.key("stamped")
.body(ByteStream::from_static(b"stamped"))
.customize()
.map_request(move |mut request| {
let headers = request.headers_mut();
headers.insert("x-rustfs-source-replication-tagging-timestamp", "2026-01-02T03:04:05Z");
headers.insert("x-minio-source-replication-retention-timestamp", "2026-01-02T03:04:06Z");
headers.insert("x-rustfs-source-replication-legalhold-timestamp", "2026-01-02T03:04:07Z");
Ok::<_, std::convert::Infallible>(request)
})
.send()
.await?;
let requests = target.requests();
let plain = requests
.iter()
.find(|record| record.operation == Operation::PutObject && record.key.as_deref() == Some("plain"))
.expect("plain PUT must be journaled");
assert_eq!(plain.replication_timestamps, ReplicationTimestampHeaders::default());
let stamped = requests
.iter()
.find(|record| record.operation == Operation::PutObject && record.key.as_deref() == Some("stamped"))
.expect("stamped PUT must be journaled");
assert_eq!(stamped.replication_timestamps.tagging.as_deref(), Some("2026-01-02T03:04:05Z"));
assert_eq!(stamped.replication_timestamps.retention.as_deref(), Some("2026-01-02T03:04:06Z"));
assert_eq!(stamped.replication_timestamps.legalhold.as_deref(), Some("2026-01-02T03:04:07Z"));
Ok(())
}
macro_rules! assert_sdk_error { macro_rules! assert_sdk_error {
($error:expr, $status:expr, $code:expr) => {{ ($error:expr, $status:expr, $code:expr) => {{
let error = &$error; let error = &$error;
@@ -2985,6 +3073,7 @@ mod tests {
part_number: None, part_number: None,
}, },
Some(0), Some(0),
ReplicationTimestampHeaders::default(),
); );
} }
let records = lock(&control).requests.clone(); let records = lock(&control).requests.clone();
@@ -3006,6 +3095,7 @@ mod tests {
part_number: None, part_number: None,
}, },
None, None,
ReplicationTimestampHeaders::default(),
); );
{ {
let bounded_records = lock(&bounded_control); let bounded_records = lock(&bounded_control);
+14 -12
View File
@@ -135,7 +135,8 @@ pub mod bucket {
pub use crate::bucket::metadata_sys::ConfigWriteLockProbe; pub use crate::bucket::metadata_sys::ConfigWriteLockProbe;
pub use crate::bucket::metadata_sys::{ pub use crate::bucket::metadata_sys::{
BucketMetadataMutationGuard, BucketMetadataSys, ObjectLockConfigState, acquire_bucket_metadata_transaction_lock, BucketMetadataMutationGuard, BucketMetadataSys, ObjectLockConfigState, acquire_bucket_metadata_transaction_lock,
capture_bucket_metadata_incarnation, delete, delete_if_incarnation, get, get_accelerate_config, get_bucket_policy, acquire_bucket_metadata_transaction_lock_for_incarnation, capture_bucket_metadata_incarnation, delete,
delete_if_incarnation, delete_under_transaction_lock, get, get_accelerate_config, get_bucket_policy,
get_bucket_policy_raw, get_bucket_targets_config, get_config_from_disk, get_cors_config, get_durability_config, get_bucket_policy_raw, get_bucket_targets_config, get_config_from_disk, get_cors_config, get_durability_config,
get_global_bucket_metadata_sys, get_lifecycle_config, get_logging_config, get_notification_config, get_global_bucket_metadata_sys, get_lifecycle_config, get_logging_config, get_notification_config,
get_object_lock_config, get_object_lock_config_state, get_public_access_block_config, get_quota_config, get_object_lock_config, get_object_lock_config_state, get_public_access_block_config, get_quota_config,
@@ -184,17 +185,18 @@ pub mod bucket {
mrf_backlog_observability_snapshot, mrf_backlog_observability_snapshot,
}; };
pub use crate::bucket::replication::{ pub use crate::bucket::replication::{
BucketReplicationResyncStatus, BucketReplicationStats, BucketStats, DeleteReplicationConfigSnapshot, BucketReplicationResyncStatus, BucketReplicationStat, BucketReplicationStats, BucketStats,
DeletedObjectReplicationInfo, DurableMrfBacklog, DynReplicationPool, MrfOpKind, MrfReplicateEntry, DeleteReplicationConfigSnapshot, DeletedObjectReplicationInfo, DurableMrfBacklog, DynReplicationPool, InQueueMetric,
MustReplicateOptions, ObjectOpts, REMOTE_TARGET_CAPABILITY_CONTRACT_VERSION, REMOTE_TARGET_UNSUPPORTED_FIELDS, MrfOpKind, MrfReplicateEntry, MustReplicateOptions, ObjectOpts, REMOTE_TARGET_CAPABILITY_CONTRACT_VERSION,
REMOTE_TARGET_WRITABLE_FIELDS, REPLICATE_INCOMING_DELETE, REPLICATION_CAPABILITY_CONTRACT_VERSION, REMOTE_TARGET_UNSUPPORTED_FIELDS, REMOTE_TARGET_WRITABLE_FIELDS, REPLICATE_INCOMING_DELETE,
REPLICATION_READ_ONLY_HISTORICAL_FIELDS, REPLICATION_WRITABLE_FIELDS, ReplicateDecision, ReplicateObjectInfo, REPLICATION_CAPABILITY_CONTRACT_VERSION, REPLICATION_READ_ONLY_HISTORICAL_FIELDS, REPLICATION_WRITABLE_FIELDS,
ReplicationBatchAdmission, ReplicationConfig, ReplicationConfigStructureError, ReplicationConfigurationExt, ReplicateDecision, ReplicateObjectInfo, ReplicationBatchAdmission, ReplicationConfig,
ReplicationDeleteScheduleInput, ReplicationDeleteStateSource, ReplicationHealQueueResult, ReplicationObjectBridge, ReplicationConfigStructureError, ReplicationConfigurationExt, ReplicationDeleteScheduleInput,
ReplicationObjectIO, ReplicationOperation, ReplicationPoolTrait, ReplicationPriority, ReplicationQueueAdmission, ReplicationDeleteStateSource, ReplicationHealQueueResult, ReplicationObjectBridge, ReplicationObjectIO,
ReplicationScannerBridge, ReplicationState, ReplicationStats, ReplicationStatusType, ReplicationStorage, ReplicationOperation, ReplicationPoolTrait, ReplicationPriority, ReplicationQueueAdmission, ReplicationScannerBridge,
ReplicationTargetValidationError, ReplicationType, ResyncOpts, ResyncStatusType, RuntimeReplicationTargetBacklog, ReplicationState, ReplicationStats, ReplicationStatusType, ReplicationStorage, ReplicationTargetValidationError,
TargetReplicationResyncStatus, VersionPurgeStatusType, commit_force_delete_intent, complete_force_delete_intent, ReplicationType, ResyncOpts, ResyncStatusType, RuntimeReplicationTargetBacklog, TargetReplicationResyncStatus,
VersionPurgeStatusType, XferStats, commit_force_delete_intent, complete_force_delete_intent,
delete_replication_state_from_config, delete_replication_version_id, get_global_replication_pool, delete_replication_state_from_config, delete_replication_version_id, get_global_replication_pool,
get_global_replication_stats, init_background_replication, invalid_replication_config_status_field, get_global_replication_stats, init_background_replication, invalid_replication_config_status_field,
persist_force_delete_intent, read_durable_mrf_backlog, replication_state_to_filemeta, replication_status_to_filemeta, persist_force_delete_intent, read_durable_mrf_backlog, replication_state_to_filemeta, replication_status_to_filemeta,
+70 -5
View File
@@ -58,7 +58,9 @@ use rustfs_utils::http::{
}; };
use rustfs_utils::http::{ use rustfs_utils::http::{
SUFFIX_FORCE_DELETE, SUFFIX_SOURCE_DELETEMARKER, SUFFIX_SOURCE_ETAG, SUFFIX_SOURCE_MTIME, SUFFIX_SOURCE_REPLICATION_CHECK, SUFFIX_FORCE_DELETE, SUFFIX_SOURCE_DELETEMARKER, SUFFIX_SOURCE_ETAG, SUFFIX_SOURCE_MTIME, SUFFIX_SOURCE_REPLICATION_CHECK,
SUFFIX_SOURCE_REPLICATION_REQUEST, SUFFIX_SOURCE_VERSION_ID, insert_header, SUFFIX_SOURCE_REPLICATION_LEGALHOLD_TIMESTAMP, SUFFIX_SOURCE_REPLICATION_REQUEST,
SUFFIX_SOURCE_REPLICATION_RETENTION_TIMESTAMP, SUFFIX_SOURCE_REPLICATION_TAGGING_TIMESTAMP, SUFFIX_SOURCE_VERSION_ID,
insert_header,
}; };
use rustls_pki_types::pem::PemObject; use rustls_pki_types::pem::PemObject;
use serde::{Deserialize, Serialize}; use serde::{Deserialize, Serialize};
@@ -80,7 +82,6 @@ use tracing::warn;
use url::Url; use url::Url;
use uuid::Uuid; use uuid::Uuid;
const DEFAULT_HEALTH_CHECK_RELOAD_DURATION: Duration = Duration::from_secs(30 * 60);
const MAX_CONCURRENT_TARGET_HEALTH_CHECKS: usize = 16; const MAX_CONCURRENT_TARGET_HEALTH_CHECKS: usize = 16;
const REDACTED_CREDENTIAL: &str = "<redacted>"; const REDACTED_CREDENTIAL: &str = "<redacted>";
@@ -1476,9 +1477,12 @@ impl Default for AdvancedPutOptions {
replication_status: ReplicationStatusType::Pending, replication_status: ReplicationStatusType::Pending,
source_mtime: OffsetDateTime::now_utc(), source_mtime: OffsetDateTime::now_utc(),
replication_request: false, replication_request: false,
retention_timestamp: OffsetDateTime::now_utc(), // UNIX_EPOCH means "never modified": header() must not emit a
tagging_timestamp: OffsetDateTime::now_utc(), // timestamp header for it, otherwise a receiver would treat an
legalhold_timestamp: OffsetDateTime::now_utc(), // unset category as a modification made right now.
retention_timestamp: OffsetDateTime::UNIX_EPOCH,
tagging_timestamp: OffsetDateTime::UNIX_EPOCH,
legalhold_timestamp: OffsetDateTime::UNIX_EPOCH,
replication_validity_check: false, replication_validity_check: false,
} }
} }
@@ -1675,6 +1679,16 @@ impl PutObjectOptions {
); );
} }
for (suffix, timestamp) in [
(SUFFIX_SOURCE_REPLICATION_TAGGING_TIMESTAMP, self.internal.tagging_timestamp),
(SUFFIX_SOURCE_REPLICATION_RETENTION_TIMESTAMP, self.internal.retention_timestamp),
(SUFFIX_SOURCE_REPLICATION_LEGALHOLD_TIMESTAMP, self.internal.legalhold_timestamp),
] {
if timestamp.unix_timestamp() != 0 {
insert_header(&mut header, suffix, timestamp.format(&Rfc3339).unwrap_or_default());
}
}
if self.internal.replication_request { if self.internal.replication_request {
insert_header(&mut header, SUFFIX_SOURCE_REPLICATION_REQUEST, "true"); insert_header(&mut header, SUFFIX_SOURCE_REPLICATION_REQUEST, "true");
} }
@@ -2842,6 +2856,57 @@ mod tests {
); );
} }
#[test]
fn put_object_headers_carry_replication_timestamp_headers() {
// MinIO receivers resolve concurrent tag/retention/legal-hold edits by
// last-writer-wins on these headers (object-api-options.go parses them
// as RFC3339); a replica without them loses every conflict resolution.
let mut opts = PutObjectOptions::default();
opts.internal.replication_request = true;
let tagging = OffsetDateTime::from_unix_timestamp(1_700_000_001).expect("valid timestamp");
let retention = OffsetDateTime::from_unix_timestamp(1_700_000_002).expect("valid timestamp");
let legalhold = OffsetDateTime::from_unix_timestamp(1_700_000_003).expect("valid timestamp");
opts.internal.tagging_timestamp = tagging;
opts.internal.retention_timestamp = retention;
opts.internal.legalhold_timestamp = legalhold;
let header = opts.header();
for (suffix, expected) in [
("source-replication-tagging-timestamp", tagging),
("source-replication-retention-timestamp", retention),
("source-replication-legalhold-timestamp", legalhold),
] {
assert_eq!(
rustfs_utils::http::get_header(&header, suffix).as_deref(),
Some(expected.format(&Rfc3339).expect("RFC3339 timestamp").as_str()),
"replication put requests must carry the {suffix} header"
);
}
}
#[test]
fn put_object_headers_omit_unset_replication_timestamps() {
// UNIX_EPOCH means "never modified on the source"; sending it would
// make the receiver treat an unset category as a fresh modification.
let mut opts = PutObjectOptions::default();
opts.internal.replication_request = true;
opts.internal.tagging_timestamp = OffsetDateTime::UNIX_EPOCH;
opts.internal.retention_timestamp = OffsetDateTime::UNIX_EPOCH;
opts.internal.legalhold_timestamp = OffsetDateTime::UNIX_EPOCH;
let header = opts.header();
for suffix in [
"source-replication-tagging-timestamp",
"source-replication-retention-timestamp",
"source-replication-legalhold-timestamp",
] {
assert!(
rustfs_utils::http::get_header(&header, suffix).is_none(),
"unset {suffix} must not be sent to replication targets"
);
}
}
#[tokio::test] #[tokio::test]
async fn get_remote_target_client_internal_rejects_loopback_endpoint() { async fn get_remote_target_client_internal_rejects_loopback_endpoint() {
let sys = BucketTargetSys::default(); let sys = BucketTargetSys::default();
@@ -126,11 +126,23 @@ const EVENT_LIFECYCLE_EXPIRED_DETECTED: &str = "lifecycle_expired_detected";
const EVENT_LIFECYCLE_NOT_ENQUEUED: &str = "lifecycle_not_enqueued"; const EVENT_LIFECYCLE_NOT_ENQUEUED: &str = "lifecycle_not_enqueued";
const EVENT_LIFECYCLE_DELETE_DISPATCHED: &str = "lifecycle_delete_dispatched"; const EVENT_LIFECYCLE_DELETE_DISPATCHED: &str = "lifecycle_delete_dispatched";
const EVENT_LIFECYCLE_DELETE_COMPLETED: &str = "lifecycle_delete_completed"; const EVENT_LIFECYCLE_DELETE_COMPLETED: &str = "lifecycle_delete_completed";
#[allow(
dead_code,
reason = "MinIO-parity tier/lifecycle entry point that this port never wired (backlog#1823)"
)]
const EVENT_LIFECYCLE_TIER_AUDIT: &str = "lifecycle_tier_audit"; const EVENT_LIFECYCLE_TIER_AUDIT: &str = "lifecycle_tier_audit";
const EVENT_LIFECYCLE_TIER_OPERATION_FAILED: &str = "lifecycle_tier_operation_failed"; const EVENT_LIFECYCLE_TIER_OPERATION_FAILED: &str = "lifecycle_tier_operation_failed";
const EVENT_LIFECYCLE_DELETE_FAILED: &str = "lifecycle_delete_failed"; const EVENT_LIFECYCLE_DELETE_FAILED: &str = "lifecycle_delete_failed";
#[allow(
dead_code,
reason = "MinIO-parity tier/lifecycle entry point that this port never wired (backlog#1823)"
)]
pub type TimeFn = Arc<dyn Fn() -> Pin<Box<dyn Future<Output = ()> + Send>> + Send + Sync + 'static>; pub type TimeFn = Arc<dyn Fn() -> Pin<Box<dyn Future<Output = ()> + Send>> + Send + Sync + 'static>;
#[allow(
dead_code,
reason = "MinIO-parity tier/lifecycle entry point that this port never wired (backlog#1823)"
)]
pub type TraceFn = pub type TraceFn =
Arc<dyn Fn(String, HashMap<String, String>) -> Pin<Box<dyn Future<Output = ()> + Send>> + Send + Sync + 'static>; Arc<dyn Fn(String, HashMap<String, String>) -> Pin<Box<dyn Future<Output = ()> + Send>> + Send + Sync + 'static>;
pub type ExpiryOpType = Box<dyn ExpiryOp + Send + Sync + 'static>; pub type ExpiryOpType = Box<dyn ExpiryOp + Send + Sync + 'static>;
@@ -140,9 +152,21 @@ static TIER_FREE_VERSION_RECOVERY_STARTED: OnceLock<()> = OnceLock::new();
static MANUAL_TRANSITION_JOB_RECOVERY_STARTED: OnceLock<()> = OnceLock::new(); static MANUAL_TRANSITION_JOB_RECOVERY_STARTED: OnceLock<()> = OnceLock::new();
pub const AMZ_OBJECT_TAGGING: &str = "X-Amz-Tagging"; pub const AMZ_OBJECT_TAGGING: &str = "X-Amz-Tagging";
#[allow(
dead_code,
reason = "MinIO-parity tier/lifecycle entry point that this port never wired (backlog#1823)"
)]
pub const AMZ_TAG_COUNT: &str = "x-amz-tagging-count"; pub const AMZ_TAG_COUNT: &str = "x-amz-tagging-count";
#[allow(
dead_code,
reason = "MinIO-parity tier/lifecycle entry point that this port never wired (backlog#1823)"
)]
pub const AMZ_TAG_DIRECTIVE: &str = "X-Amz-Tagging-Directive"; pub const AMZ_TAG_DIRECTIVE: &str = "X-Amz-Tagging-Directive";
pub const AMZ_ENCRYPTION_AES: &str = "AES256"; pub const AMZ_ENCRYPTION_AES: &str = "AES256";
#[allow(
dead_code,
reason = "MinIO-parity tier/lifecycle entry point that this port never wired (backlog#1823)"
)]
pub const AMZ_ENCRYPTION_KMS: &str = "aws:kms"; pub const AMZ_ENCRYPTION_KMS: &str = "aws:kms";
pub const ERR_INVALID_STORAGECLASS: &str = "invalid tier."; pub const ERR_INVALID_STORAGECLASS: &str = "invalid tier.";
@@ -280,6 +304,10 @@ impl LifecycleSys {
} }
} }
#[allow(
dead_code,
reason = "MinIO-parity tier/lifecycle entry point that this port never wired (backlog#1823)"
)]
pub fn trace(oi: &ObjectInfo) -> TraceFn { pub fn trace(oi: &ObjectInfo) -> TraceFn {
let bucket = oi.bucket.clone(); let bucket = oi.bucket.clone();
let name = oi.name.clone(); let name = oi.name.clone();
@@ -570,6 +598,10 @@ async fn delete_free_version_remote_object(
Ok(()) Ok(())
} }
#[allow(
dead_code,
reason = "MinIO-parity tier/lifecycle entry point that this port never wired (backlog#1823)"
)]
async fn delete_free_version_remote_object_then<T, F, Fut>( async fn delete_free_version_remote_object_then<T, F, Fut>(
oi: &ObjectInfo, oi: &ObjectInfo,
tier_config_mgr: &Arc<RwLock<TierConfigMgr>>, tier_config_mgr: &Arc<RwLock<TierConfigMgr>>,
@@ -2868,6 +2900,10 @@ fn stale_upload_default_due(initiated: OffsetDateTime, default_expiry: StdDurati
initiated + time::Duration::seconds(default_expiry.as_secs() as i64) initiated + time::Duration::seconds(default_expiry.as_secs() as i64)
} }
#[allow(
dead_code,
reason = "MinIO-parity tier/lifecycle entry point that this port never wired (backlog#1823)"
)]
async fn stale_upload_current_size(set: &Arc<SetDisks>, metadata: &HashMap<String, String>, upload_dir: &str) -> Option<usize> { async fn stale_upload_current_size(set: &Arc<SetDisks>, metadata: &HashMap<String, String>, upload_dir: &str) -> Option<usize> {
stale_upload_current_size_with_opts(set, metadata, upload_dir, false).await stale_upload_current_size_with_opts(set, metadata, upload_dir, false).await
} }
@@ -3352,6 +3388,10 @@ pub async fn validate_transition_tier(lc: &BucketLifecycleConfiguration) -> Resu
Ok(()) Ok(())
} }
#[allow(
dead_code,
reason = "MinIO-parity tier/lifecycle entry point that this port never wired (backlog#1823)"
)]
fn mark_delete_opts_skip_decommissioned_on_remote_success(opts: &mut ObjectOptions, remote_delete_succeeded: bool) { fn mark_delete_opts_skip_decommissioned_on_remote_success(opts: &mut ObjectOptions, remote_delete_succeeded: bool) {
if remote_delete_succeeded { if remote_delete_succeeded {
opts.skip_decommissioned = true; opts.skip_decommissioned = true;
@@ -4339,6 +4379,10 @@ pub async fn expire_transitioned_object(
Ok(dobj) Ok(dobj)
} }
#[allow(
dead_code,
reason = "MinIO-parity tier/lifecycle entry point that this port never wired (backlog#1823)"
)]
pub fn gen_transition_objname(bucket: &str) -> Result<String, Error> { pub fn gen_transition_objname(bucket: &str) -> Result<String, Error> {
let us = Uuid::new_v4().to_string(); let us = Uuid::new_v4().to_string();
let mut hasher = Sha256::new(); let mut hasher = Sha256::new();
@@ -4373,6 +4417,10 @@ pub async fn transition_object(api: Arc<ECStore>, oi: &ObjectInfo, lae: LcAuditE
result result
} }
#[allow(
dead_code,
reason = "MinIO-parity tier/lifecycle entry point that this port never wired (backlog#1823)"
)]
pub fn audit_tier_actions(_tier: &str, bytes: i64) -> TimeFn { pub fn audit_tier_actions(_tier: &str, bytes: i64) -> TimeFn {
let tier = _tier.to_string(); let tier = _tier.to_string();
Arc::new(move || { Arc::new(move || {
@@ -4391,6 +4439,10 @@ pub fn audit_tier_actions(_tier: &str, bytes: i64) -> TimeFn {
}) })
} }
#[allow(
dead_code,
reason = "MinIO-parity tier/lifecycle entry point that this port never wired (backlog#1823)"
)]
pub async fn get_transitioned_object_reader( pub async fn get_transitioned_object_reader(
bucket: &str, bucket: &str,
object: &str, object: &str,
@@ -5145,6 +5197,10 @@ async fn lifecycle_delete_config_snapshot(api: &ECStore, oi: &ObjectInfo) -> Res
ReplicationObjectBridge::delete_request_config(api, &oi.bucket).await ReplicationObjectBridge::delete_request_config(api, &oi.bucket).await
} }
#[allow(
dead_code,
reason = "MinIO-parity tier/lifecycle entry point that this port never wired (backlog#1823)"
)]
pub async fn apply_lifecycle_action(event: &lifecycle::Event, src: &LcEventSrc, oi: &ObjectInfo) -> bool { pub async fn apply_lifecycle_action(event: &lifecycle::Event, src: &LcEventSrc, oi: &ObjectInfo) -> bool {
let mut success = false; let mut success = false;
match event.action { match event.action {
@@ -7422,6 +7478,10 @@ mod tests {
// process environment while `env::set_var`/`env::remove_var` is active. // process environment while `env::set_var`/`env::remove_var` is active.
// SAFETY: keep this note adjacent to the allowance for the repository guard. // SAFETY: keep this note adjacent to the allowance for the repository guard.
#[allow(unsafe_code)] #[allow(unsafe_code)]
#[allow(
dead_code,
reason = "transition-queue env fixture kept for tests that scope those vars; no test uses it today (backlog#1823)"
)]
async fn with_transition_queue_env_async<F, Fut>(capacity: Option<&str>, timeout_ms: Option<&str>, test_fn: F) async fn with_transition_queue_env_async<F, Fut>(capacity: Option<&str>, timeout_ms: Option<&str>, test_fn: F)
where where
F: FnOnce() -> Fut, F: FnOnce() -> Fut,
@@ -759,6 +759,10 @@ pub struct ManualTransitionWorkerResultRecord {
} }
impl ManualTransitionWorkerResultRecord { impl ManualTransitionWorkerResultRecord {
#[allow(
dead_code,
reason = "MinIO-parity tier/lifecycle entry point that this port never wired (backlog#1823)"
)]
pub fn new(job_id: Uuid, task_key: impl Into<String>, result: ManualTransitionWorkerResult) -> Self { pub fn new(job_id: Uuid, task_key: impl Into<String>, result: ManualTransitionWorkerResult) -> Self {
Self::new_with_reason(job_id, task_key, result, None) Self::new_with_reason(job_id, task_key, result, None)
} }
@@ -1257,6 +1261,10 @@ pub(crate) async fn save_manual_transition_task_if_absent(
} }
} }
#[allow(
dead_code,
reason = "MinIO-parity tier/lifecycle entry point that this port never wired (backlog#1823)"
)]
pub async fn load_manual_transition_task_record( pub async fn load_manual_transition_task_record(
api: Arc<ECStore>, api: Arc<ECStore>,
job_id: Uuid, job_id: Uuid,
@@ -1320,6 +1328,10 @@ async fn scan_manual_transition_task_journal(api: Arc<ECStore>, job_id: Uuid) ->
} }
} }
#[allow(
dead_code,
reason = "MinIO-parity tier/lifecycle entry point that this port never wired (backlog#1823)"
)]
pub async fn load_manual_transition_worker_result_stats( pub async fn load_manual_transition_worker_result_stats(
api: Arc<ECStore>, api: Arc<ECStore>,
job_id: Uuid, job_id: Uuid,
@@ -1455,6 +1467,10 @@ async fn scan_manual_transition_worker_result_journal(
} }
} }
#[allow(
dead_code,
reason = "MinIO-parity tier/lifecycle entry point that this port never wired (backlog#1823)"
)]
pub async fn reconcile_manual_transition_worker_results( pub async fn reconcile_manual_transition_worker_results(
api: Arc<ECStore>, api: Arc<ECStore>,
job_id: Uuid, job_id: Uuid,
@@ -15,25 +15,35 @@
use rustfs_common::metrics::IlmAction; use rustfs_common::metrics::IlmAction;
use crate::bucket::lifecycle::lifecycle::ObjectOpts; use crate::bucket::lifecycle::lifecycle::ObjectOpts;
use crate::bucket::replication::ReplicationLifecycleBridge;
pub(crate) use crate::bucket::replication::ReplicationStatusType; pub(crate) use crate::bucket::replication::ReplicationStatusType;
#[cfg(test)] #[cfg(test)]
pub(crate) use crate::bucket::replication::VersionPurgeStatusType; pub(crate) use crate::bucket::replication::VersionPurgeStatusType;
pub(crate) use crate::bucket::replication::{ pub(crate) use crate::bucket::replication::{
DeleteReplicationConfigSnapshot, ReplicationObjectBridge, replication_state_to_filemeta, DeleteReplicationConfigSnapshot, ReplicationObjectBridge, replication_state_to_filemeta,
}; };
use crate::bucket::replication::{ReplicationLifecycleBridge, ReplicationLifecycleConfig};
use crate::storage_api_contracts::object::DeletedObject; use crate::storage_api_contracts::object::DeletedObject;
pub(crate) type LifecycleReplicationConfig = ReplicationLifecycleConfig; #[allow(
dead_code,
reason = "declared boundary surface for the ECStore replication split plan; no caller in this port (backlog#1823)"
)]
pub(crate) fn has_pending_version_purge(obj: &ObjectOpts) -> bool { pub(crate) fn has_pending_version_purge(obj: &ObjectOpts) -> bool {
obj.version_purge_status.is_pending() obj.version_purge_status.is_pending()
} }
#[allow(
dead_code,
reason = "declared boundary surface for the ECStore replication split plan; no caller in this port (backlog#1823)"
)]
pub(crate) fn has_pending_object_replication(obj: &ObjectOpts) -> bool { pub(crate) fn has_pending_object_replication(obj: &ObjectOpts) -> bool {
replication_status_blocks_lifecycle(&obj.replication_status) replication_status_blocks_lifecycle(&obj.replication_status)
} }
#[allow(
dead_code,
reason = "declared boundary surface for the ECStore replication split plan; no caller in this port (backlog#1823)"
)]
pub(crate) fn has_pending_lifecycle_replication(obj: &ObjectOpts) -> bool { pub(crate) fn has_pending_lifecycle_replication(obj: &ObjectOpts) -> bool {
has_pending_object_replication(obj) || has_pending_version_purge(obj) has_pending_object_replication(obj) || has_pending_version_purge(obj)
} }
@@ -14,6 +14,10 @@
use std::collections::HashMap; use std::collections::HashMap;
#[allow(
dead_code,
reason = "declared boundary surface for the ECStore replication split plan; no caller in this port (backlog#1823)"
)]
pub(crate) fn decode_tags_to_map(tags: &str) -> HashMap<String, String> { pub(crate) fn decode_tags_to_map(tags: &str) -> HashMap<String, String> {
crate::bucket::tagging::decode_tags_to_map(tags) crate::bucket::tagging::decode_tags_to_map(tags)
} }
@@ -331,6 +331,10 @@ where
persist_tier_delete_journal_entry(api, &committed).await persist_tier_delete_journal_entry(api, &committed).await
} }
#[allow(
dead_code,
reason = "MinIO-parity tier/lifecycle entry point that this port never wired (backlog#1823)"
)]
pub async fn abort_tier_delete_journal_entry<S>(api: Arc<S>, je: &Jentry) -> std::io::Result<()> pub async fn abort_tier_delete_journal_entry<S>(api: Arc<S>, je: &Jentry) -> std::io::Result<()>
where where
S: ObjectOperations< S: ObjectOperations<
@@ -148,6 +148,10 @@ struct RecoveryCursor {
object: String, object: String,
} }
#[allow(
dead_code,
reason = "MinIO-parity tier/lifecycle entry point that this port never wired (backlog#1823)"
)]
pub async fn recover_tier_free_versions( pub async fn recover_tier_free_versions(
api: Arc<ECStore>, api: Arc<ECStore>,
limit: usize, limit: usize,
@@ -385,6 +385,10 @@ impl ExpiryOp for Jentry {
} }
} }
#[allow(
dead_code,
reason = "MinIO-parity tier/lifecycle entry point that this port never wired (backlog#1823)"
)]
pub async fn delete_object_from_remote_tier(obj_name: &str, rv_id: &str, tier_name: &str) -> Result<(), std::io::Error> { pub async fn delete_object_from_remote_tier(obj_name: &str, rv_id: &str, tier_name: &str) -> Result<(), std::io::Error> {
let result = delete_object_from_remote_tier_raw(obj_name, rv_id, tier_name).await; let result = delete_object_from_remote_tier_raw(obj_name, rv_id, tier_name).await;
if let Err(err) = &result if let Err(err) = &result
@@ -395,6 +399,10 @@ pub async fn delete_object_from_remote_tier(obj_name: &str, rv_id: &str, tier_na
result result
} }
#[allow(
dead_code,
reason = "MinIO-parity tier/lifecycle entry point that this port never wired (backlog#1823)"
)]
async fn delete_object_from_remote_tier_raw(obj_name: &str, rv_id: &str, tier_name: &str) -> Result<(), std::io::Error> { async fn delete_object_from_remote_tier_raw(obj_name: &str, rv_id: &str, tier_name: &str) -> Result<(), std::io::Error> {
#[cfg(test)] #[cfg(test)]
if let Some(result) = run_remote_tier_delete_test_hook(obj_name, rv_id, tier_name) { if let Some(result) = run_remote_tier_delete_test_hook(obj_name, rv_id, tier_name) {
@@ -405,6 +413,10 @@ async fn delete_object_from_remote_tier_raw(obj_name: &str, rv_id: &str, tier_na
delete_object_from_remote_tier_raw_with_manager(obj_name, rv_id, tier_name, &tier_config_mgr).await delete_object_from_remote_tier_raw_with_manager(obj_name, rv_id, tier_name, &tier_config_mgr).await
} }
#[allow(
dead_code,
reason = "MinIO-parity tier/lifecycle entry point that this port never wired (backlog#1823)"
)]
async fn delete_object_from_remote_tier_raw_with_manager( async fn delete_object_from_remote_tier_raw_with_manager(
obj_name: &str, obj_name: &str,
rv_id: &str, rv_id: &str,
@@ -485,6 +497,10 @@ pub enum RemoteTierDeleteOutcome {
AlreadyRemoved, AlreadyRemoved,
} }
#[allow(
dead_code,
reason = "MinIO-parity tier/lifecycle entry point that this port never wired (backlog#1823)"
)]
pub async fn delete_object_from_remote_tier_idempotent( pub async fn delete_object_from_remote_tier_idempotent(
obj_name: &str, obj_name: &str,
rv_id: &str, rv_id: &str,
@@ -50,8 +50,16 @@ pub type Result<T> = std::result::Result<T, TransitionTransactionError>;
#[derive(Debug, thiserror::Error)] #[derive(Debug, thiserror::Error)]
pub enum TransitionTransactionError { pub enum TransitionTransactionError {
#[error("transition transaction already exists")] #[error("transition transaction already exists")]
#[allow(
dead_code,
reason = "MinIO-parity tier/lifecycle entry point that this port never wired (backlog#1823)"
)]
AlreadyExists, AlreadyExists,
#[error("transition transaction is not found")] #[error("transition transaction is not found")]
#[allow(
dead_code,
reason = "MinIO-parity tier/lifecycle entry point that this port never wired (backlog#1823)"
)]
NotFound, NotFound,
#[error("transition transaction is corrupt: {0}")] #[error("transition transaction is corrupt: {0}")]
Corrupt(&'static str), Corrupt(&'static str),
+31
View File
@@ -60,12 +60,14 @@ struct ConfigWriteLockProbeState {
static CONFIG_WRITE_LOCK_PROBES: std::sync::OnceLock<StdMutex<Vec<Arc<ConfigWriteLockProbeState>>>> = std::sync::OnceLock::new(); static CONFIG_WRITE_LOCK_PROBES: std::sync::OnceLock<StdMutex<Vec<Arc<ConfigWriteLockProbeState>>>> = std::sync::OnceLock::new();
#[cfg(any(test, feature = "test-util"))] #[cfg(any(test, feature = "test-util"))]
#[allow(dead_code, reason = "installed by tests behind `--features test-util` (backlog#1823)")]
pub struct ConfigWriteLockProbe { pub struct ConfigWriteLockProbe {
state: Arc<ConfigWriteLockProbeState>, state: Arc<ConfigWriteLockProbeState>,
} }
#[cfg(any(test, feature = "test-util"))] #[cfg(any(test, feature = "test-util"))]
impl ConfigWriteLockProbe { impl ConfigWriteLockProbe {
#[allow(dead_code, reason = "installed by tests behind `--features test-util` (backlog#1823)")]
pub fn install(bucket: &str) -> Self { pub fn install(bucket: &str) -> Self {
let state = Arc::new(ConfigWriteLockProbeState { let state = Arc::new(ConfigWriteLockProbeState {
bucket: bucket.to_string(), bucket: bucket.to_string(),
@@ -84,6 +86,7 @@ impl ConfigWriteLockProbe {
Self { state } Self { state }
} }
#[allow(dead_code, reason = "installed by tests behind `--features test-util` (backlog#1823)")]
pub async fn wait_until_attempted(&self) { pub async fn wait_until_attempted(&self) {
tokio::time::timeout(Duration::from_secs(30), self.state.arrived.notified()) tokio::time::timeout(Duration::from_secs(30), self.state.arrived.notified())
.await .await
@@ -656,6 +659,16 @@ pub async fn update_under_transaction_lock(
update_under_config_write_guard(get_bucket_metadata_sys()?, guard, config_file, data).await update_under_config_write_guard(get_bucket_metadata_sys()?, guard, config_file, data).await
} }
/// Clear one config file while the caller holds this bucket's transaction lock.
pub async fn delete_under_transaction_lock(
guard: &BucketMetadataMutationGuard,
bucket: &str,
config_file: &str,
) -> Result<OffsetDateTime> {
guard.ensure_valid(bucket)?;
delete_under_config_write_guard(get_bucket_metadata_sys()?, guard, config_file).await
}
pub async fn update_quota_if_incarnation( pub async fn update_quota_if_incarnation(
bucket: &str, bucket: &str,
data: Vec<u8>, data: Vec<u8>,
@@ -795,6 +808,14 @@ pub async fn acquire_bucket_metadata_transaction_lock(bucket: &str) -> Result<Bu
acquire_config_write_guard(get_bucket_metadata_sys()?, bucket).await acquire_config_write_guard(get_bucket_metadata_sys()?, bucket).await
} }
/// Acquire the bucket transaction lock only if its incarnation still matches.
pub async fn acquire_bucket_metadata_transaction_lock_for_incarnation(
bucket: &str,
expected_incarnation_id: Uuid,
) -> Result<BucketMetadataMutationGuard> {
acquire_config_write_guard_for_incarnation(get_bucket_metadata_sys()?, bucket, Some(expected_incarnation_id)).await
}
pub(crate) async fn acquire_bucket_metadata_transaction_lock_in( pub(crate) async fn acquire_bucket_metadata_transaction_lock_in(
ctx: &crate::runtime::instance::InstanceContext, ctx: &crate::runtime::instance::InstanceContext,
bucket: &str, bucket: &str,
@@ -872,6 +893,10 @@ pub async fn get_bucket_policy_raw(bucket: &str) -> Result<(String, OffsetDateTi
bucket_meta_sys.get_bucket_policy_raw(bucket).await bucket_meta_sys.get_bucket_policy_raw(bucket).await
} }
#[allow(
dead_code,
reason = "free-function facade over the live BucketMetadataSys::get_bucket_acl_config; no caller in this port (backlog#1823)"
)]
pub async fn get_bucket_acl_config(bucket: &str) -> Result<(String, OffsetDateTime)> { pub async fn get_bucket_acl_config(bucket: &str) -> Result<(String, OffsetDateTime)> {
let bucket_meta_sys_lock = get_bucket_metadata_sys()?; let bucket_meta_sys_lock = get_bucket_metadata_sys()?;
let bucket_meta_sys = bucket_meta_sys_lock.read().await; let bucket_meta_sys = bucket_meta_sys_lock.read().await;
@@ -1086,6 +1111,10 @@ pub async fn get_config_from_disk(bucket: &str) -> Result<BucketMetadata> {
bucket_meta_sys.get_config_from_disk(bucket).await bucket_meta_sys.get_config_from_disk(bucket).await
} }
#[allow(
dead_code,
reason = "ambient-facade variant of the live created_at_in; no caller in this port (backlog#1823)"
)]
pub async fn created_at(bucket: &str) -> Result<OffsetDateTime> { pub async fn created_at(bucket: &str) -> Result<OffsetDateTime> {
let bucket_meta_sys_lock = get_bucket_metadata_sys()?; let bucket_meta_sys_lock = get_bucket_metadata_sys()?;
let bucket_meta_sys = bucket_meta_sys_lock.read().await; let bucket_meta_sys = bucket_meta_sys_lock.read().await;
@@ -1599,6 +1628,7 @@ impl BucketMetadataSys {
/// [`Self::update`], with the payload computed from the loaded metadata /// [`Self::update`], with the payload computed from the loaded metadata
/// instead of supplied up front. Loads through this system's own store so /// instead of supplied up front. Loads through this system's own store so
/// the read and the persisted write target the same instance. /// the read and the persisted write target the same instance.
#[allow(dead_code, reason = "asserted by this file's tests (backlog#1823)")]
async fn update_config_with<F>(&self, bucket: &str, config_file: &str, mutate: F) -> Result<OffsetDateTime> async fn update_config_with<F>(&self, bucket: &str, config_file: &str, mutate: F) -> Result<OffsetDateTime>
where where
F: FnOnce(&BucketMetadata) -> Result<Vec<u8>> + Send, F: FnOnce(&BucketMetadata) -> Result<Vec<u8>> + Send,
@@ -1703,6 +1733,7 @@ impl BucketMetadataSys {
/// A miss is never published as an authoritative default, and a snapshot /// A miss is never published as an authoritative default, and a snapshot
/// read before delete plus same-name recreation cannot replace the new /// read before delete plus same-name recreation cannot replace the new
/// generation. /// generation.
#[allow(dead_code, reason = "asserted by this file's tests (backlog#1823)")]
pub(crate) async fn reload_from_store(&self, bucket: &str) -> Result<()> { pub(crate) async fn reload_from_store(&self, bucket: &str) -> Result<()> {
if is_meta_bucketname(bucket) { if is_meta_bucketname(bucket) {
return Err(Error::other("errInvalidArgument")); return Err(Error::other("errInvalidArgument"));
-1
View File
@@ -13,7 +13,6 @@
// limitations under the License. // limitations under the License.
// #730: bucket subsystems still contain staged ECStore migration code. // #730: bucket subsystems still contain staged ECStore migration code.
#![allow(dead_code)]
pub mod bandwidth; pub mod bandwidth;
pub mod bucket_target_sys; pub mod bucket_target_sys;
@@ -136,6 +136,7 @@ pub fn add_years(dt: OffsetDateTime, years: i32) -> OffsetDateTime {
/// Check if an object has legal hold enabled. /// Check if an object has legal hold enabled.
/// Returns true if legal hold is ON. /// Returns true if legal hold is ON.
#[allow(dead_code, reason = "asserted by this file's tests (backlog#1823)")]
fn has_legal_hold(user_defined: &std::collections::HashMap<String, String>) -> bool { fn has_legal_hold(user_defined: &std::collections::HashMap<String, String>) -> bool {
let lhold = objectlock::get_object_legalhold_meta(user_defined); let lhold = objectlock::get_object_legalhold_meta(user_defined);
matches!(lhold.status, Some(ref st) if st.as_str() == ObjectLockLegalHoldStatus::ON) matches!(lhold.status, Some(ref st) if st.as_str() == ObjectLockLegalHoldStatus::ON)
@@ -151,6 +152,7 @@ fn has_legal_hold(user_defined: &std::collections::HashMap<String, String>) -> b
/// # Returns /// # Returns
/// * `true` if the object is locked (cannot be deleted/modified) /// * `true` if the object is locked (cannot be deleted/modified)
/// * `false` if the object is not locked /// * `false` if the object is not locked
#[allow(dead_code, reason = "asserted by this file's tests (backlog#1823)")]
pub fn is_object_locked_by_metadata(user_defined: &std::collections::HashMap<String, String>, is_delete_marker: bool) -> bool { pub fn is_object_locked_by_metadata(user_defined: &std::collections::HashMap<String, String>, is_delete_marker: bool) -> bool {
// Delete markers are never locked // Delete markers are never locked
if is_delete_marker { if is_delete_marker {
+2
View File
@@ -193,6 +193,7 @@ pub enum QuotaError {
} }
#[derive(Debug, Serialize)] #[derive(Debug, Serialize)]
#[allow(dead_code, reason = "asserted by this file's tests (backlog#1823)")]
pub struct QuotaErrorResponse { pub struct QuotaErrorResponse {
#[serde(rename = "Code")] #[serde(rename = "Code")]
pub code: String, pub code: String,
@@ -208,6 +209,7 @@ pub struct QuotaErrorResponse {
} }
impl QuotaErrorResponse { impl QuotaErrorResponse {
#[allow(dead_code, reason = "asserted by this file's tests (backlog#1823)")]
pub fn new(quota_error: &QuotaError, request_id: &str, host_id: &str) -> Self { pub fn new(quota_error: &QuotaError, request_id: &str, host_id: &str) -> Self {
match quota_error { match quota_error {
QuotaError::QuotaExceeded { .. } => Self { QuotaError::QuotaExceeded { .. } => Self {
@@ -899,6 +899,7 @@ async fn save_ledger_locked(
} }
#[cfg(any(test, feature = "test-util"))] #[cfg(any(test, feature = "test-util"))]
#[allow(dead_code, reason = "asserted by this file's tests (backlog#1823)")]
pub fn fail_next_quota_ledger_save_for_test() { pub fn fail_next_quota_ledger_save_for_test() {
FAIL_NEXT_LEDGER_SAVE.store(true, std::sync::atomic::Ordering::SeqCst); FAIL_NEXT_LEDGER_SAVE.store(true, std::sync::atomic::Ordering::SeqCst);
} }
+2 -2
View File
@@ -60,7 +60,7 @@ pub use replication_filemeta_boundary::{
pub(crate) use replication_filemeta_boundary::{ pub(crate) use replication_filemeta_boundary::{
replication_state_from_filemeta, replication_status_from_filemeta, version_purge_status_from_filemeta, replication_state_from_filemeta, replication_status_from_filemeta, version_purge_status_from_filemeta,
}; };
pub(crate) use replication_lifecycle_bridge::{ReplicationLifecycleBridge, ReplicationLifecycleConfig}; pub(crate) use replication_lifecycle_bridge::ReplicationLifecycleBridge;
pub(crate) use replication_migration_bridge::ReplicationMigrationBridge; pub(crate) use replication_migration_bridge::ReplicationMigrationBridge;
pub use replication_object_bridge::ReplicationObjectBridge; pub use replication_object_bridge::ReplicationObjectBridge;
pub use replication_object_config::{DeleteReplicationConfigSnapshot, ReplicationConfig}; pub use replication_object_config::{DeleteReplicationConfigSnapshot, ReplicationConfig};
@@ -81,6 +81,6 @@ pub use replication_queue_boundary::{
pub use replication_resync_boundary::{BucketReplicationResyncStatus, ResyncOpts, TargetReplicationResyncStatus}; pub use replication_resync_boundary::{BucketReplicationResyncStatus, ResyncOpts, TargetReplicationResyncStatus};
pub use replication_scanner_bridge::ReplicationScannerBridge; pub use replication_scanner_bridge::ReplicationScannerBridge;
pub use replication_state::{ReplicationStats, RuntimeReplicationTargetBacklog}; pub use replication_state::{ReplicationStats, RuntimeReplicationTargetBacklog};
pub use replication_stats_boundary::{BucketReplicationStats, BucketStats}; pub use replication_stats_boundary::{BucketReplicationStat, BucketReplicationStats, BucketStats, InQueueMetric, XferStats};
pub use replication_storage_boundary::{ReplicationObjectIO, ReplicationStorage}; pub use replication_storage_boundary::{ReplicationObjectIO, ReplicationStorage};
pub(crate) use replication_target_config_bridge::ReplicationTargetConfigBridge; pub(crate) use replication_target_config_bridge::ReplicationTargetConfigBridge;
@@ -37,6 +37,10 @@ impl ReplicationConfigStore {
com::read_config_limited(api, file, max_bytes).await com::read_config_limited(api, file, max_bytes).await
} }
#[allow(
dead_code,
reason = "MinIO-parity replication surface with no caller in this port (backlog#1823)"
)]
pub(crate) async fn read_no_lock<S>(api: Arc<S>, file: &str) -> Result<Vec<u8>> pub(crate) async fn read_no_lock<S>(api: Arc<S>, file: &str) -> Result<Vec<u8>>
where where
S: ReplicationObjectIO, S: ReplicationObjectIO,
@@ -24,15 +24,27 @@ use super::replication_storage_boundary::{
DeletedObject, ObjectInfo, ObjectOptions, ObjectToDelete, deleted_object_for_replication, DeletedObject, ObjectInfo, ObjectOptions, ObjectToDelete, deleted_object_for_replication,
}; };
#[allow(
dead_code,
reason = "declared boundary surface for the ECStore replication split plan; no caller in this port (backlog#1823)"
)]
pub(crate) type ReplicationLifecycleConfig = ReplicationConfig; pub(crate) type ReplicationLifecycleConfig = ReplicationConfig;
pub(crate) struct ReplicationLifecycleBridge; pub(crate) struct ReplicationLifecycleBridge;
impl ReplicationLifecycleBridge { impl ReplicationLifecycleBridge {
#[allow(
dead_code,
reason = "declared boundary surface for the ECStore replication split plan; no caller in this port (backlog#1823)"
)]
pub(crate) fn new_config(config: ReplicationConfiguration) -> ReplicationLifecycleConfig { pub(crate) fn new_config(config: ReplicationConfiguration) -> ReplicationLifecycleConfig {
ReplicationConfig::new(Some(config), None) ReplicationConfig::new(Some(config), None)
} }
#[allow(
dead_code,
reason = "declared boundary surface for the ECStore replication split plan; no caller in this port (backlog#1823)"
)]
pub(crate) fn has_pending_version_purge( pub(crate) fn has_pending_version_purge(
config: &ReplicationLifecycleConfig, config: &ReplicationLifecycleConfig,
object_name: &str, object_name: &str,
@@ -45,6 +57,10 @@ impl ReplicationLifecycleBridge {
.is_some_and(|config| config.has_active_rules(object_name, true)) .is_some_and(|config| config.has_active_rules(object_name, true))
} }
#[allow(
dead_code,
reason = "declared boundary surface for the ECStore replication split plan; no caller in this port (backlog#1823)"
)]
pub(crate) async fn check_delete_replication( pub(crate) async fn check_delete_replication(
bucket: &str, bucket: &str,
object: &ObjectToDelete, object: &ObjectToDelete,
@@ -54,6 +70,10 @@ impl ReplicationLifecycleBridge {
check_replicate_delete(bucket, object, source, opts, None).await check_replicate_delete(bucket, object, source, opts, None).await
} }
#[allow(
dead_code,
reason = "declared boundary surface for the ECStore replication split plan; no caller in this port (backlog#1823)"
)]
pub(crate) fn version_delete_replication_state(decision: &ReplicateDecision) -> ReplicationState { pub(crate) fn version_delete_replication_state(decision: &ReplicateDecision) -> ReplicationState {
let pending_status = decision.pending_status(); let pending_status = decision.pending_status();
ReplicationState { ReplicationState {
@@ -19,17 +19,33 @@ use time::OffsetDateTime;
use super::replication_error_boundary::Result; use super::replication_error_boundary::Result;
use crate::bucket::msgp_decode; use crate::bucket::msgp_decode;
#[allow(
dead_code,
reason = "declared boundary surface for the ECStore replication split plan; no caller in this port (backlog#1823)"
)]
pub(crate) struct ReplicationMsgpCodec; pub(crate) struct ReplicationMsgpCodec;
impl ReplicationMsgpCodec { impl ReplicationMsgpCodec {
#[allow(
dead_code,
reason = "declared boundary surface for the ECStore replication split plan; no caller in this port (backlog#1823)"
)]
pub(crate) fn read_ext8_time<R: Read>(rd: &mut R) -> Result<OffsetDateTime> { pub(crate) fn read_ext8_time<R: Read>(rd: &mut R) -> Result<OffsetDateTime> {
msgp_decode::read_msgp_ext8_time(rd) msgp_decode::read_msgp_ext8_time(rd)
} }
#[allow(
dead_code,
reason = "declared boundary surface for the ECStore replication split plan; no caller in this port (backlog#1823)"
)]
pub(crate) fn skip_value<R: Read>(rd: &mut R) -> Result<()> { pub(crate) fn skip_value<R: Read>(rd: &mut R) -> Result<()> {
msgp_decode::skip_msgp_value(rd) msgp_decode::skip_msgp_value(rd)
} }
#[allow(
dead_code,
reason = "declared boundary surface for the ECStore replication split plan; no caller in this port (backlog#1823)"
)]
pub(crate) fn write_time<W: Write>(wr: &mut W, time: OffsetDateTime) -> Result<()> { pub(crate) fn write_time<W: Write>(wr: &mut W, time: OffsetDateTime) -> Result<()> {
msgp_decode::write_msgp_time(wr, time) msgp_decode::write_msgp_time(wr, time)
} }
@@ -77,6 +77,10 @@ impl ReplicationObjectBridge {
load_delete_request_config_in(ctx, bucket).await load_delete_request_config_in(ctx, bucket).await
} }
#[allow(
dead_code,
reason = "declared boundary surface for the ECStore replication split plan; no caller in this port (backlog#1823)"
)]
pub(crate) async fn delete_config_snapshot_in( pub(crate) async fn delete_config_snapshot_in(
ctx: &ReplicationInstanceContext, ctx: &ReplicationInstanceContext,
bucket: &str, bucket: &str,
@@ -231,6 +231,10 @@ pub(crate) async fn load_delete_replication_config(
delete_snapshot_from_metadata(ReplicationMetadataStore::delete_metadata(bucket).await?) delete_snapshot_from_metadata(ReplicationMetadataStore::delete_metadata(bucket).await?)
} }
#[allow(
dead_code,
reason = "MinIO-parity replication surface with no caller in this port (backlog#1823)"
)]
pub(crate) async fn load_delete_replication_config_in( pub(crate) async fn load_delete_replication_config_in(
ctx: &ReplicationInstanceContext, ctx: &ReplicationInstanceContext,
bucket: &str, bucket: &str,
@@ -217,6 +217,10 @@ impl DurableMrfBacklogTracker {
} }
} }
#[allow(
dead_code,
reason = "MinIO-parity replication surface with no caller in this port (backlog#1823)"
)]
fn durable_mrf_backlog_tracker_from_entries(entries: &[MrfReplicateEntry]) -> DurableMrfBacklogTracker { fn durable_mrf_backlog_tracker_from_entries(entries: &[MrfReplicateEntry]) -> DurableMrfBacklogTracker {
let mut tracker = DurableMrfBacklogTracker { let mut tracker = DurableMrfBacklogTracker {
available: true, available: true,
@@ -712,6 +716,10 @@ pub struct ReplicationPool<S: ReplicationStorage> {
// MRF worker lifecycle // MRF worker lifecycle
mrf_worker_cancellations: Mutex<Vec<CancellationToken>>, mrf_worker_cancellations: Mutex<Vec<CancellationToken>>,
#[allow(
dead_code,
reason = "MinIO-parity replication surface with no caller in this port (backlog#1823)"
)]
mrf_stop_tx: Sender<()>, mrf_stop_tx: Sender<()>,
// Worker size tracking // Worker size tracking
@@ -940,6 +948,10 @@ impl<S: ReplicationStorage> ReplicationPool<S> {
} }
/// Resizes worker priority and counts /// Resizes worker priority and counts
#[allow(
dead_code,
reason = "MinIO-parity replication surface with no caller in this port (backlog#1823)"
)]
pub async fn resize_worker_priority( pub async fn resize_worker_priority(
&self, &self,
pri: ReplicationPriority, pri: ReplicationPriority,
@@ -1180,6 +1192,10 @@ impl<S: ReplicationStorage> ReplicationPool<S> {
} }
/// Queues an MRF save operation /// Queues an MRF save operation
#[allow(
dead_code,
reason = "MinIO-parity replication surface with no caller in this port (backlog#1823)"
)]
async fn queue_mrf_save(&self, entry: MrfReplicateEntry) { async fn queue_mrf_save(&self, entry: MrfReplicateEntry) {
let _ = self.queue_mrf_save_admission(entry, "mrf_worker").await; let _ = self.queue_mrf_save_admission(entry, "mrf_worker").await;
} }
@@ -1651,6 +1667,10 @@ impl<S: ReplicationStorage> ReplicationPool<S> {
} }
/// Worker function for handling regular replication operations /// Worker function for handling regular replication operations
#[allow(
dead_code,
reason = "MinIO-parity replication surface with no caller in this port (backlog#1823)"
)]
async fn add_worker( async fn add_worker(
&self, &self,
mut rx: Receiver<ReplicationOperation>, mut rx: Receiver<ReplicationOperation>,
@@ -1664,6 +1684,10 @@ impl<S: ReplicationStorage> ReplicationPool<S> {
} }
/// Worker function for handling large object replication operations /// Worker function for handling large object replication operations
#[allow(
dead_code,
reason = "MinIO-parity replication surface with no caller in this port (backlog#1823)"
)]
async fn add_large_worker( async fn add_large_worker(
&self, &self,
mut rx: Receiver<ReplicationOperation>, mut rx: Receiver<ReplicationOperation>,
@@ -1678,6 +1702,10 @@ impl<S: ReplicationStorage> ReplicationPool<S> {
} }
/// Worker function for handling MRF (Most Recent Failures) operations /// Worker function for handling MRF (Most Recent Failures) operations
#[allow(
dead_code,
reason = "MinIO-parity replication surface with no caller in this port (backlog#1823)"
)]
async fn add_mrf_worker( async fn add_mrf_worker(
&self, &self,
mut rx: Receiver<ReplicationOperation>, mut rx: Receiver<ReplicationOperation>,
@@ -1691,6 +1719,10 @@ impl<S: ReplicationStorage> ReplicationPool<S> {
} }
/// Delete resync metadata from replication resync state in memory /// Delete resync metadata from replication resync state in memory
#[allow(
dead_code,
reason = "MinIO-parity replication surface with no caller in this port (backlog#1823)"
)]
pub async fn delete_resync_metadata(&self, bucket: &str) { pub async fn delete_resync_metadata(&self, bucket: &str) {
let mut status_map = self.resyncer.status_map.write().await; let mut status_map = self.resyncer.status_map.write().await;
status_map.remove(bucket); status_map.remove(bucket);
@@ -21,11 +21,31 @@ pub(crate) use rustfs_replication::{
should_count_head_proxy_failure, should_count_head_proxy_failure,
}; };
#[allow(
dead_code,
reason = "declared boundary surface for the ECStore replication split plan; no caller in this port (backlog#1823)"
)]
pub(crate) const RESYNC_META_FORMAT: u16 = rustfs_replication::resync::RESYNC_META_FORMAT; pub(crate) const RESYNC_META_FORMAT: u16 = rustfs_replication::resync::RESYNC_META_FORMAT;
#[allow(
dead_code,
reason = "declared boundary surface for the ECStore replication split plan; no caller in this port (backlog#1823)"
)]
pub(crate) const RESYNC_META_VERSION: u16 = rustfs_replication::resync::RESYNC_META_VERSION; pub(crate) const RESYNC_META_VERSION: u16 = rustfs_replication::resync::RESYNC_META_VERSION;
pub(crate) const RESYNC_FILE_MAX_BYTES: usize = rustfs_replication::RESYNC_FILE_MAX_BYTES; pub(crate) const RESYNC_FILE_MAX_BYTES: usize = rustfs_replication::RESYNC_FILE_MAX_BYTES;
#[allow(
dead_code,
reason = "declared boundary surface for the ECStore replication split plan; no caller in this port (backlog#1823)"
)]
pub(crate) const WIRE_ZERO_TIME_UNIX: i64 = rustfs_replication::resync::WIRE_ZERO_TIME_UNIX; pub(crate) const WIRE_ZERO_TIME_UNIX: i64 = rustfs_replication::resync::WIRE_ZERO_TIME_UNIX;
#[allow(
dead_code,
reason = "declared boundary surface for the ECStore replication split plan; no caller in this port (backlog#1823)"
)]
pub(crate) const MRF_META_FORMAT: u16 = rustfs_replication::mrf::MRF_META_FORMAT; pub(crate) const MRF_META_FORMAT: u16 = rustfs_replication::mrf::MRF_META_FORMAT;
#[allow(
dead_code,
reason = "declared boundary surface for the ECStore replication split plan; no caller in this port (backlog#1823)"
)]
pub(crate) const MRF_META_VERSION: u16 = rustfs_replication::mrf::MRF_META_VERSION; pub(crate) const MRF_META_VERSION: u16 = rustfs_replication::mrf::MRF_META_VERSION;
fn map_replication_error(err: rustfs_replication::Error) -> Error { fn map_replication_error(err: rustfs_replication::Error) -> Error {
@@ -122,6 +122,10 @@ const REPLICATION_TARGET_OFFLINE_ERROR_MARKERS: &[&str] = &[
"tcp connect error", "tcp connect error",
]; ];
#[allow(
dead_code,
reason = "MinIO-parity replication surface with no caller in this port (backlog#1823)"
)]
const RESYNC_TIME_INTERVAL: TokioDuration = TokioDuration::from_secs(60); const RESYNC_TIME_INTERVAL: TokioDuration = TokioDuration::from_secs(60);
static WARNED_MONITOR_UNINIT: std::sync::Once = std::sync::Once::new(); static WARNED_MONITOR_UNINIT: std::sync::Once = std::sync::Once::new();
@@ -328,6 +332,10 @@ fn bounded_resync_max_jobs(value: usize) -> usize {
#[derive(Debug)] #[derive(Debug)]
pub struct ReplicationResyncer { pub struct ReplicationResyncer {
pub status_map: Arc<RwLock<HashMap<String, BucketReplicationResyncStatus>>>, pub status_map: Arc<RwLock<HashMap<String, BucketReplicationResyncStatus>>>,
#[allow(
dead_code,
reason = "MinIO-parity replication surface with no caller in this port (backlog#1823)"
)]
pub worker_size: usize, pub worker_size: usize,
pub(crate) cancel_tokens: Arc<RwLock<HashMap<ResyncCancelKey, CancellationToken>>>, pub(crate) cancel_tokens: Arc<RwLock<HashMap<ResyncCancelKey, CancellationToken>>>,
resync_admission: Arc<Semaphore>, resync_admission: Arc<Semaphore>,
@@ -544,6 +552,10 @@ impl ReplicationResyncer {
.is_some_and(|status| status.failed_count > 0) .is_some_and(|status| status.failed_count > 0)
} }
#[allow(
dead_code,
reason = "MinIO-parity replication surface with no caller in this port (backlog#1823)"
)]
pub async fn persist_to_disk<S>(&self, cancel_token: CancellationToken, api: Arc<S>) pub async fn persist_to_disk<S>(&self, cancel_token: CancellationToken, api: Arc<S>)
where where
S: ReplicationObjectIO, S: ReplicationObjectIO,
@@ -340,6 +340,10 @@ impl ReplicationStats {
} }
/// Site replication update replica statistics /// Site replication update replica statistics
#[allow(
dead_code,
reason = "MinIO-parity replication surface with no caller in this port (backlog#1823)"
)]
fn sr_update_replica_stat(&self, size: i64) { fn sr_update_replica_stat(&self, size: i64) {
self.sr_stats.replica_size.fetch_add(size, Ordering::Relaxed); self.sr_stats.replica_size.fetch_add(size, Ordering::Relaxed);
self.sr_stats.replica_count.fetch_add(1, Ordering::Relaxed); self.sr_stats.replica_count.fetch_add(1, Ordering::Relaxed);
@@ -704,6 +708,12 @@ impl ReplicationStats {
} else { } else {
BucketReplicationStats::new() BucketReplicationStats::new()
}; };
// Stamp the serializable failure windows from the live samples: the
// samples themselves do not cross the peer-RPC wire, so this snapshot
// is what cluster aggregation and the metrics endpoints see.
for stat in replication_stats.stats.values_mut() {
stat.fail_stats.refresh_windows();
}
let uptime = if cache.contains_key(bucket) { let uptime = if cache.contains_key(bucket) {
SystemTime::now() SystemTime::now()
.duration_since(SystemTime::UNIX_EPOCH) .duration_since(SystemTime::UNIX_EPOCH)
@@ -15,7 +15,9 @@
#[cfg(test)] #[cfg(test)]
pub(crate) use rustfs_replication::FailStats; pub(crate) use rustfs_replication::FailStats;
pub(crate) use rustfs_replication::{ pub(crate) use rustfs_replication::{
ActiveWorkerStat, BucketReplicationStat, InQueueMetric, ProxyMetric, ProxyStatsCache, QueueCache, ReplicationMetricScope, ActiveWorkerStat, ProxyMetric, ProxyStatsCache, QueueCache, ReplicationMetricScope, SRMetricsSummary,
SRMetricsSummary, XferStats,
}; };
pub use rustfs_replication::{BucketReplicationStats, BucketStats}; // Public so the admin wire DTOs (rustfs/src/admin/replication_metrics_wire.rs)
// can project the internal stats onto the minio-go response shapes through
// the storage_api facade chain.
pub use rustfs_replication::{BucketReplicationStat, BucketReplicationStats, BucketStats, InQueueMetric, XferStats};
@@ -27,8 +27,10 @@ use rustfs_utils::http::{
AMZ_OBJECT_TAGGING, AMZ_SERVER_SIDE_ENCRYPTION, AMZ_SERVER_SIDE_ENCRYPTION_KMS_CONTEXT, AMZ_SERVER_SIDE_ENCRYPTION_KMS_ID, AMZ_OBJECT_TAGGING, AMZ_SERVER_SIDE_ENCRYPTION, AMZ_SERVER_SIDE_ENCRYPTION_KMS_CONTEXT, AMZ_SERVER_SIDE_ENCRYPTION_KMS_ID,
AMZ_STORAGE_CLASS, AMZ_TAG_COUNT, CACHE_CONTROL, CONTENT_DISPOSITION, CONTENT_ENCODING, CONTENT_LANGUAGE, CONTENT_TYPE, AMZ_STORAGE_CLASS, AMZ_TAG_COUNT, CACHE_CONTROL, CONTENT_DISPOSITION, CONTENT_ENCODING, CONTENT_LANGUAGE, CONTENT_TYPE,
HeaderExt as _, SUFFIX_OBJECTLOCK_LEGALHOLD_TIMESTAMP, SUFFIX_OBJECTLOCK_RETENTION_TIMESTAMP, HeaderExt as _, SUFFIX_OBJECTLOCK_LEGALHOLD_TIMESTAMP, SUFFIX_OBJECTLOCK_RETENTION_TIMESTAMP,
SUFFIX_REPLICATION_ACTUAL_OBJECT_SIZE, SUFFIX_REPLICATION_SSEC_CRC, SUFFIX_TAGGING_TIMESTAMP, get_str, insert_header_map, SUFFIX_REPLICATION_ACTUAL_OBJECT_SIZE, SUFFIX_REPLICATION_SSEC_CRC, SUFFIX_SOURCE_REPLICATION_LEGALHOLD_TIMESTAMP,
is_internal_key, is_object_encryption_marker, is_replication_stripped_encryption_key, ssec_replication_transport_header, SUFFIX_SOURCE_REPLICATION_RETENTION_TIMESTAMP, SUFFIX_SOURCE_REPLICATION_TAGGING_TIMESTAMP, SUFFIX_TAGGING_TIMESTAMP,
get_str, insert_header_map, is_internal_key, is_object_encryption_marker, is_replication_stripped_encryption_key,
ssec_replication_transport_header,
}; };
use time::OffsetDateTime; use time::OffsetDateTime;
use time::format_description::well_known::Rfc3339; use time::format_description::well_known::Rfc3339;
@@ -119,6 +121,27 @@ fn classify_replication_source_encryption(metadata: &HashMap<String, String>) ->
} }
} }
fn is_legacy_source_replication_timestamp_key(key: &str) -> bool {
fn has_prefix_and_suffix(key: &str, prefix: &str, suffix: &str) -> bool {
let key = key.as_bytes();
key.len() == prefix.len() + suffix.len()
&& key[..prefix.len()].eq_ignore_ascii_case(prefix.as_bytes())
&& key[prefix.len()..].eq_ignore_ascii_case(suffix.as_bytes())
}
[
SUFFIX_SOURCE_REPLICATION_TAGGING_TIMESTAMP,
SUFFIX_SOURCE_REPLICATION_RETENTION_TIMESTAMP,
SUFFIX_SOURCE_REPLICATION_LEGALHOLD_TIMESTAMP,
]
.iter()
.any(|suffix| {
["x-rustfs-", "x-minio-"]
.iter()
.any(|prefix| has_prefix_and_suffix(key, prefix, suffix))
})
}
pub(crate) fn replication_object_is_ssec_encrypted(user_defined: &HashMap<String, String>) -> bool { pub(crate) fn replication_object_is_ssec_encrypted(user_defined: &HashMap<String, String>) -> bool {
rustfs_replication::is_ssec_encrypted(user_defined) rustfs_replication::is_ssec_encrypted(user_defined)
} }
@@ -176,6 +199,11 @@ pub(crate) fn replication_put_object_options(sc: &str, object_info: &ObjectInfo)
continue; continue;
} }
if is_legacy_source_replication_timestamp_key(key) {
meta.insert(format!("x-amz-meta-{key}"), value.to_string());
continue;
}
if is_internal_key(key) || is_standard_header(key) { if is_internal_key(key) || is_standard_header(key) {
continue; continue;
} }
@@ -259,15 +287,23 @@ pub(crate) fn replication_put_object_options(sc: &str, object_info: &ObjectInfo)
if !tags.is_empty() { if !tags.is_empty() {
put_options.user_tags = tags; put_options.user_tags = tags;
put_options.internal.tagging_timestamp =
if let Some(timestamp) = get_str(&object_info.user_defined, SUFFIX_TAGGING_TIMESTAMP) {
OffsetDateTime::parse(&timestamp, &Rfc3339)
.map_err(|err| Error::other(format!("Failed to parse tagging timestamp: {err}")))?
} else {
object_info.mod_time.unwrap_or(OffsetDateTime::UNIX_EPOCH)
};
} }
} }
// Load the stored tagging timestamp independently of whether any tags
// remain: DeleteObjectTagging leaves the object tagless but stamps this
// key, and the deletion's LWW timestamp must still reach the replica.
// With no stored key, fall back to mod_time only while tags exist
// (MinIO parity); a tagless object without the key was never tagged and
// keeps the epoch default (no header).
put_options.internal.tagging_timestamp = if let Some(timestamp) = get_str(&object_info.user_defined, SUFFIX_TAGGING_TIMESTAMP)
{
OffsetDateTime::parse(&timestamp, &Rfc3339)
.map_err(|err| Error::other(format!("Failed to parse tagging timestamp: {err}")))?
} else if !put_options.user_tags.is_empty() {
object_info.mod_time.unwrap_or(OffsetDateTime::UNIX_EPOCH)
} else {
OffsetDateTime::UNIX_EPOCH
};
let metadata = &*object_info.user_defined; let metadata = &*object_info.user_defined;
@@ -283,13 +319,15 @@ pub(crate) fn replication_put_object_options(sc: &str, object_info: &ObjectInfo)
put_options.cache_control = cache_control.to_string(); put_options.cache_control = cache_control.to_string();
} }
if let Some(mode) = metadata.lookup(AMZ_OBJECT_LOCK_MODE) { if let Some(mode) = metadata.lookup(AMZ_OBJECT_LOCK_MODE).filter(|mode| !mode.is_empty()) {
put_options.mode = Some(ObjectLockRetentionMode::from(mode.to_uppercase().as_str())); put_options.mode = Some(ObjectLockRetentionMode::from(mode.to_uppercase().as_str()));
} }
if let Some(retain_until_date) = metadata.lookup(AMZ_OBJECT_LOCK_RETAIN_UNTIL_DATE) { if let Some(retain_until_date) = metadata.lookup(AMZ_OBJECT_LOCK_RETAIN_UNTIL_DATE) {
put_options.retain_until_date = OffsetDateTime::parse(retain_until_date, &Rfc3339) if !retain_until_date.is_empty() {
.map_err(|err| Error::other(format!("Failed to parse retain until date: {err}")))?; put_options.retain_until_date = OffsetDateTime::parse(retain_until_date, &Rfc3339)
.map_err(|err| Error::other(format!("Failed to parse retain until date: {err}")))?;
}
put_options.internal.retention_timestamp = put_options.internal.retention_timestamp =
if let Some(timestamp) = get_str(&object_info.user_defined, SUFFIX_OBJECTLOCK_RETENTION_TIMESTAMP) { if let Some(timestamp) = get_str(&object_info.user_defined, SUFFIX_OBJECTLOCK_RETENTION_TIMESTAMP) {
OffsetDateTime::parse(&timestamp, &Rfc3339).unwrap_or(OffsetDateTime::UNIX_EPOCH) OffsetDateTime::parse(&timestamp, &Rfc3339).unwrap_or(OffsetDateTime::UNIX_EPOCH)
@@ -694,6 +732,110 @@ mod tests {
assert!(options.internal.replication_request); assert!(options.internal.replication_request);
} }
/// DeleteObjectTagging leaves the object tagless but stamps the
/// tagging-timestamp internal key; the deletion's LWW timestamp must
/// still be loaded (and therefore sent) so the replica can order the
/// deletion against concurrent tag edits.
#[test]
fn replication_put_options_carry_tagging_timestamp_after_tag_deletion() {
let mut metadata = std::collections::HashMap::new();
rustfs_utils::http::insert_str(&mut metadata, SUFFIX_TAGGING_TIMESTAMP, "2026-01-02T03:04:05Z".to_string());
let object_info = ObjectInfo {
user_defined: Arc::new(metadata),
user_tags: Arc::new(String::new()),
mod_time: Some(OffsetDateTime::UNIX_EPOCH),
version_id: Some(Uuid::nil()),
..Default::default()
};
let (options, _) = replication_put_object_options("", &object_info).expect("build put options");
assert!(options.user_tags.is_empty());
assert_eq!(
options.internal.tagging_timestamp,
OffsetDateTime::parse("2026-01-02T03:04:05Z", &Rfc3339).expect("valid timestamp"),
"the stored tagging timestamp must load independently of remaining tags"
);
// A tagless object without the stored key was never tagged: the epoch
// default keeps the header unsent.
let untagged = ObjectInfo {
user_tags: Arc::new(String::new()),
mod_time: Some(OffsetDateTime::from_unix_timestamp(1_700_000_000).expect("timestamp")),
version_id: Some(Uuid::nil()),
..Default::default()
};
let (options, _) = replication_put_object_options("", &untagged).expect("build put options");
assert_eq!(options.internal.tagging_timestamp, OffsetDateTime::UNIX_EPOCH);
}
#[test]
fn replication_put_options_do_not_promote_legacy_user_timestamp_metadata() {
let legacy_keys = [
"x-rustfs-source-replication-tagging-timestamp",
"x-rustfs-source-replication-retention-timestamp",
"x-rustfs-source-replication-legalhold-timestamp",
"x-minio-source-replication-tagging-timestamp",
"x-minio-source-replication-retention-timestamp",
"x-minio-source-replication-legalhold-timestamp",
];
let object_info = ObjectInfo {
user_defined: Arc::new(
legacy_keys
.iter()
.map(|key| (key.to_string(), "2099-01-02T03:04:05Z".to_string()))
.collect(),
),
..Default::default()
};
let (options, _) = replication_put_object_options("", &object_info).expect("build put options");
for legacy_key in legacy_keys {
assert!(!options.user_metadata.contains_key(legacy_key));
assert_eq!(
options
.user_metadata
.get(&format!("x-amz-meta-{legacy_key}"))
.map(String::as_str),
Some("2099-01-02T03:04:05Z")
);
}
assert_eq!(options.internal.tagging_timestamp, OffsetDateTime::UNIX_EPOCH);
assert_eq!(options.internal.retention_timestamp, OffsetDateTime::UNIX_EPOCH);
assert_eq!(options.internal.legalhold_timestamp, OffsetDateTime::UNIX_EPOCH);
}
#[test]
fn replication_put_options_carry_retention_timestamp_after_clear() {
let mut metadata = HashMap::from([
(AMZ_OBJECT_LOCK_MODE.to_string(), String::new()),
(AMZ_OBJECT_LOCK_RETAIN_UNTIL_DATE.to_string(), String::new()),
]);
rustfs_utils::http::insert_str(&mut metadata, SUFFIX_OBJECTLOCK_RETENTION_TIMESTAMP, "2026-01-02T03:04:05Z".to_string());
let object_info = ObjectInfo {
user_defined: Arc::new(metadata),
..Default::default()
};
let (options, _) = replication_put_object_options("", &object_info).expect("retention clear must replicate");
assert!(options.mode.is_none());
assert_eq!(options.retain_until_date, OffsetDateTime::UNIX_EPOCH);
assert_eq!(
options.internal.retention_timestamp,
OffsetDateTime::parse("2026-01-02T03:04:05Z", &Rfc3339).expect("valid timestamp")
);
let headers = options.header();
assert!(!headers.contains_key(AMZ_OBJECT_LOCK_MODE));
assert!(!headers.contains_key(AMZ_OBJECT_LOCK_RETAIN_UNTIL_DATE));
assert_eq!(
rustfs_utils::http::get_header(&headers, SUFFIX_SOURCE_REPLICATION_RETENTION_TIMESTAMP).as_deref(),
Some("2026-01-02T03:04:05Z")
);
}
#[test] #[test]
fn replication_put_options_strip_encryption_metadata_from_plaintext_objects() { fn replication_put_options_strip_encryption_metadata_from_plaintext_objects() {
use rustfs_utils::http::object_encryption_keys::{INTERNAL_ENCRYPTION_ORIGINAL_SIZE_HEADER, SSEC_ORIGINAL_SIZE_HEADER}; use rustfs_utils::http::object_encryption_keys::{INTERNAL_ENCRYPTION_ORIGINAL_SIZE_HEADER, SSEC_ORIGINAL_SIZE_HEADER};
+52 -4
View File
@@ -40,7 +40,14 @@ impl ARN {
impl Display for ARN { impl Display for ARN {
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result { fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
write!(f, "arn:rustfs:{}:{}:{}:{}", self.arn_type, self.region, self.id, self.bucket) // The `minio` partition is deliberate: madmin-go's ParseARN
// hard-rejects any other partition, so native mc/madmin tooling can
// only decode remote-target ARNs minted in this form (backlog#1675
// P1-7). Legacy `arn:rustfs:` ARNs persisted by older releases stay
// readable via the FromStr whitelist below; runtime matching between
// targets and replication rules is by full-string equality, so mixed
// partitions coexist safely.
write!(f, "arn:minio:{}:{}:{}:{}", self.arn_type, self.region, self.id, self.bucket)
} }
} }
@@ -48,7 +55,12 @@ impl FromStr for ARN {
type Err = std::io::Error; type Err = std::io::Error;
fn from_str(s: &str) -> Result<Self, Self::Err> { fn from_str(s: &str) -> Result<Self, Self::Err> {
if !s.starts_with("arn:rustfs:") { // Partition whitelist, not just an `arn:` check: `BucketTargetType::
// from_str(...).unwrap_or_default()` below never fails, so this is
// the only structural gate rejecting foreign ARNs. `arn:rustfs:` is
// the legacy partition and must stay accepted forever (persisted
// bucket-targets.json / replication configs from older releases).
if !s.starts_with("arn:minio:") && !s.starts_with("arn:rustfs:") {
return Err(std::io::Error::new(std::io::ErrorKind::InvalidInput, "Invalid ARN format")); return Err(std::io::Error::new(std::io::ErrorKind::InvalidInput, "Invalid ARN format"));
} }
@@ -101,14 +113,50 @@ mod tests {
} }
/// RustFS commonly generates ARNs with an empty region: /// RustFS commonly generates ARNs with an empty region:
/// `arn:rustfs:replication::<deployment_id>:<bucket>`. /// `arn:minio:replication::<deployment_id>:<bucket>`.
#[test] #[test]
fn from_str_handles_empty_region_segment() { fn from_str_handles_empty_region_segment() {
let parsed = ARN::from_str("arn:rustfs:replication::depl-123:bucket-a").expect("valid ARN must parse"); let parsed = ARN::from_str("arn:minio:replication::depl-123:bucket-a").expect("valid ARN must parse");
assert_eq!(parsed.arn_type, BucketTargetType::ReplicationService); assert_eq!(parsed.arn_type, BucketTargetType::ReplicationService);
assert_eq!(parsed.region, "", "region segment is empty in this form"); assert_eq!(parsed.region, "", "region segment is empty in this form");
assert_eq!(parsed.id, "depl-123"); assert_eq!(parsed.id, "depl-123");
assert_eq!(parsed.bucket, "bucket-a"); assert_eq!(parsed.bucket, "bucket-a");
} }
/// madmin-go's `ParseARN` hard-rejects anything that does not start with
/// `arn:minio:`, so generated ARNs must use the `minio` partition or the
/// native mc/madmin tooling cannot decode remote-target listings.
#[test]
fn display_emits_minio_partition() {
let arn = ARN::new(
BucketTargetType::ReplicationService,
"depl-123".to_string(),
String::new(),
"bucket-a".to_string(),
);
assert_eq!(arn.to_string(), "arn:minio:replication::depl-123:bucket-a");
}
/// Persisted bucket-targets.json files from older RustFS releases carry
/// `arn:rustfs:` ARNs; the legacy partition must stay parseable forever.
#[test]
fn from_str_accepts_legacy_rustfs_partition() {
let parsed = ARN::from_str("arn:rustfs:replication:us-east-1:depl-123:bucket-a").expect("legacy ARN must parse");
assert_eq!(parsed.arn_type, BucketTargetType::ReplicationService);
assert_eq!(parsed.region, "us-east-1");
assert_eq!(parsed.id, "depl-123");
assert_eq!(parsed.bucket, "bucket-a");
}
/// The partition whitelist is the only structural gate: `BucketTargetType::
/// from_str(...).unwrap_or_default()` never fails, so any 6-segment string
/// would otherwise parse as `type=None`.
#[test]
fn from_str_rejects_unknown_partition() {
assert!(ARN::from_str("arn:aws:replication::depl-123:bucket-a").is_err());
assert!(ARN::from_str("not-an-arn").is_err());
}
} }
@@ -59,6 +59,10 @@ impl fmt::Debug for Credentials {
} }
#[derive(Debug, Deserialize, Serialize, Default, Clone)] #[derive(Debug, Deserialize, Serialize, Default, Clone)]
#[allow(
dead_code,
reason = "MinIO-parity bucket-target service discriminator with no caller in this port (backlog#1823)"
)]
pub enum ServiceType { pub enum ServiceType {
#[default] #[default]
Replication, Replication,
+20 -17
View File
@@ -73,23 +73,6 @@ pub fn check_valid_bucket_name_strict(bucket_name: &str) -> Result<()> {
check_bucket_name_common(bucket_name, true) check_bucket_name_common(bucket_name, true)
} }
pub fn check_valid_object_name_prefix(object_name: &str) -> Result<()> {
if object_name.len() > 1024 {
return Err(Error::other("Object name cannot be longer than 1024 characters"));
}
if !object_name.is_ascii() {
return Err(Error::other("Object name with non-UTF-8 strings are not supported"));
}
Ok(())
}
pub fn check_valid_object_name(object_name: &str) -> Result<()> {
if object_name.trim().is_empty() {
return Err(Error::other("Object name cannot be empty"));
}
check_valid_object_name_prefix(object_name)
}
pub fn deserialize<T>(input: &[u8]) -> xml::DeResult<T> pub fn deserialize<T>(input: &[u8]) -> xml::DeResult<T>
where where
T: for<'xml> xml::Deserialize<'xml>, T: for<'xml> xml::Deserialize<'xml>,
@@ -100,6 +83,10 @@ where
Ok(ans) Ok(ans)
} }
#[allow(
dead_code,
reason = "xml serialize helper with no caller in this port; the live sibling is deserialize (backlog#1823)"
)]
pub fn serialize_content<T: xml::SerializeContent>(val: &T) -> xml::SerResult<String> { pub fn serialize_content<T: xml::SerializeContent>(val: &T) -> xml::SerResult<String> {
let mut buf = Vec::with_capacity(256); let mut buf = Vec::with_capacity(256);
{ {
@@ -186,15 +173,27 @@ pub fn is_valid_object_name(object: &str) -> bool {
/// Client-facing reason attached to rejections of object keys that Win32/NTFS /// Client-facing reason attached to rejections of object keys that Win32/NTFS
/// cannot represent as file paths (issue #3299). Deployments on Linux/macOS /// cannot represent as file paths (issue #3299). Deployments on Linux/macOS
/// accept the full S3 key character set. /// accept the full S3 key character set.
#[allow(
dead_code,
reason = "live on Windows: callers sit inside the #[cfg(target_os = \"windows\")] block in check_object_name_for_length_and_slash (backlog#1823)"
)]
pub const WINDOWS_RESERVED_CHARACTERS_REASON: &str = pub const WINDOWS_RESERVED_CHARACTERS_REASON: &str =
"object key contains characters unsupported on Windows hosts (one of ':', '*', '?', '\"', '|', '<', '>')"; "object key contains characters unsupported on Windows hosts (one of ':', '*', '?', '\"', '|', '<', '>')";
/// Client-facing reason for path segments Windows can store but not address /// Client-facing reason for path segments Windows can store but not address
/// afterwards (issue #3449): trailing dot/space or reserved DOS device names. /// afterwards (issue #3449): trailing dot/space or reserved DOS device names.
#[allow(
dead_code,
reason = "live on Windows: callers sit inside the #[cfg(target_os = \"windows\")] block in check_object_name_for_length_and_slash (backlog#1823)"
)]
pub const WINDOWS_RESERVED_SEGMENT_REASON: &str = "object key contains a path segment unsupported on Windows hosts (trailing dot or space, or a reserved device name such as NUL/CON/COM1)"; pub const WINDOWS_RESERVED_SEGMENT_REASON: &str = "object key contains a path segment unsupported on Windows hosts (trailing dot or space, or a reserved device name such as NUL/CON/COM1)";
/// Reserved DOS device names that shadow regular files on Windows, even when /// Reserved DOS device names that shadow regular files on Windows, even when
/// an extension is appended (e.g. `NUL.txt` resolves to the `NUL` device). /// an extension is appended (e.g. `NUL.txt` resolves to the `NUL` device).
#[allow(
dead_code,
reason = "live on Windows: callers sit inside the #[cfg(target_os = \"windows\")] block in check_object_name_for_length_and_slash (backlog#1823)"
)]
const WINDOWS_RESERVED_NAMES: &[&str] = &[ const WINDOWS_RESERVED_NAMES: &[&str] = &[
"CON", "PRN", "AUX", "NUL", "COM1", "COM2", "COM3", "COM4", "COM5", "COM6", "COM7", "COM8", "COM9", "LPT1", "LPT2", "LPT3", "CON", "PRN", "AUX", "NUL", "COM1", "COM2", "COM3", "COM4", "COM5", "COM6", "COM7", "COM8", "COM9", "LPT1", "LPT2", "LPT3",
"LPT4", "LPT5", "LPT6", "LPT7", "LPT8", "LPT9", "LPT4", "LPT5", "LPT6", "LPT7", "LPT8", "LPT9",
@@ -204,6 +203,10 @@ const WINDOWS_RESERVED_NAMES: &[&str] = &[
/// the Win32 API cannot address afterwards (issue #3449): segments ending in a /// the Win32 API cannot address afterwards (issue #3449): segments ending in a
/// dot or a space, and reserved DOS device names — bare or with an extension /// dot or a space, and reserved DOS device names — bare or with an extension
/// (`NUL.txt`), matching classic Win32 path resolution semantics. /// (`NUL.txt`), matching classic Win32 path resolution semantics.
#[allow(
dead_code,
reason = "live on Windows: callers sit inside the #[cfg(target_os = \"windows\")] block in check_object_name_for_length_and_slash (backlog#1823)"
)]
pub fn object_name_has_windows_incompatible_segment(object: &str) -> bool { pub fn object_name_has_windows_incompatible_segment(object: &str) -> bool {
object.split(['/', '\\']).any(|segment| { object.split(['/', '\\']).any(|segment| {
if segment.ends_with('.') || segment.ends_with(' ') { if segment.ends_with('.') || segment.ends_with(' ') {
@@ -90,6 +90,10 @@ impl BucketVersioningSys {
/// caller's own instance context so a second in-process store never /// caller's own instance context so a second in-process store never
/// answers with the first instance's versioning state; falls back to the /// answers with the first instance's versioning state; falls back to the
/// ambient system when the instance cell is not initialized. /// ambient system when the instance cell is not initialized.
#[allow(
dead_code,
reason = "instance-scoped seam (backlog#1052) with no caller in this port (backlog#1823)"
)]
pub(crate) async fn get_in(ctx: &crate::runtime::instance::InstanceContext, bucket: &str) -> Result<VersioningConfiguration> { pub(crate) async fn get_in(ctx: &crate::runtime::instance::InstanceContext, bucket: &str) -> Result<VersioningConfiguration> {
if bucket == RUSTFS_META_BUCKET || bucket.starts_with(RUSTFS_META_BUCKET) { if bucket == RUSTFS_META_BUCKET || bucket.starts_with(RUSTFS_META_BUCKET) {
return Ok(VersioningConfiguration::default()); return Ok(VersioningConfiguration::default());
@@ -15,6 +15,7 @@
use crate::disk::disk_store::{get_drive_walkdir_peek_timeout, get_drive_walkdir_stall_timeout}; use crate::disk::disk_store::{get_drive_walkdir_peek_timeout, get_drive_walkdir_stall_timeout};
use crate::disk::error::DiskError; use crate::disk::error::DiskError;
use crate::disk::{self, DiskAPI, DiskStore, WalkDirOptions}; use crate::disk::{self, DiskAPI, DiskStore, WalkDirOptions};
use futures::future::join_all;
use metrics::counter; use metrics::counter;
use rustfs_filemeta::{MetaCacheEntries, MetaCacheEntry, MetacacheReader, is_io_eof}; use rustfs_filemeta::{MetaCacheEntries, MetaCacheEntry, MetacacheReader, is_io_eof};
use std::{ use std::{
@@ -655,6 +656,7 @@ async fn list_path_raw_inner(
errs.push(None); errs.push(None);
} }
let mut pending_entries: Vec<Option<MetaCacheEntry>> = vec![None; readers.len()]; let mut pending_entries: Vec<Option<MetaCacheEntry>> = vec![None; readers.len()];
let mut peek_outcomes: Vec<Option<PeekOutcome>> = std::iter::repeat_with(|| None).take(readers.len()).collect();
loop { loop {
let mut current = MetaCacheEntry::default(); let mut current = MetaCacheEntry::default();
@@ -676,6 +678,21 @@ async fn list_path_raw_inner(
let mut has_err = 0; let mut has_err = 0;
let mut agree = 0; let mut agree = 0;
// Start every missing head read in the same round so one stalled
// disk cannot multiply the wait budget by the erasure-set width.
// Outcomes are still consumed below in stable disk-index order.
let concurrent_peeks = readers.iter_mut().enumerate().filter_map(|(i, reader)| {
if errs[i].is_some() || pending_entries[i].is_some() {
return None;
}
let cancel = &revjob_rx;
Some(async move { (i, peek_with_timeout(cancel, reader, peek_timeout).await) })
});
for (i, outcome) in join_all(concurrent_peeks).await {
peek_outcomes[i] = Some(outcome);
}
for (i, r) in readers.iter_mut().enumerate() { for (i, r) in readers.iter_mut().enumerate() {
if errs[i].is_some() { if errs[i].is_some() {
has_err += 1; has_err += 1;
@@ -685,7 +702,10 @@ async fn list_path_raw_inner(
let entry = if let Some(entry) = pending_entries[i].take() { let entry = if let Some(entry) = pending_entries[i].take() {
entry entry
} else { } else {
match peek_with_timeout(&revjob_rx, r, peek_timeout).await { let Some(outcome) = peek_outcomes[i].take() else {
return Err(DiskError::Unexpected);
};
match outcome {
PeekOutcome::Ready(res) => { PeekOutcome::Ready(res) => {
if let Some(entry) = res { if let Some(entry) = res {
// info!("read entry disk: {}, name: {}", i, entry.name); // info!("read entry disk: {}, name: {}", i, entry.name);
@@ -1295,6 +1315,36 @@ mod tests {
assert_eq!(err, DiskError::Timeout); assert_eq!(err, DiskError::Timeout);
} }
#[tokio::test(start_paused = true)]
async fn list_path_raw_bounds_multiple_stalled_readers_by_one_peek_deadline() {
let peek_timeout = Duration::from_millis(20);
let started = tokio::time::Instant::now();
let err = list_path_raw(
CancellationToken::new(),
ListPathRawOptions {
disks: vec![None, None, None, None],
min_disks: 1,
test_reader_behaviors: vec![
TestReaderBehavior::Stall,
TestReaderBehavior::Stall,
TestReaderBehavior::Stall,
TestReaderBehavior::Stall,
],
peek_timeout: Some(peek_timeout),
..Default::default()
},
)
.await
.expect_err("all stalled readers should fail the listing");
assert_eq!(err, DiskError::Timeout);
assert_eq!(
started.elapsed(),
peek_timeout,
"reader deadlines must overlap instead of accumulating once per disk"
);
}
#[tokio::test] #[tokio::test]
async fn list_path_raw_waits_past_producer_stall_for_slow_progressing_reader() { async fn list_path_raw_waits_past_producer_stall_for_slow_progressing_reader() {
let entry = MetaCacheEntry { let entry = MetaCacheEntry {
@@ -229,17 +229,6 @@ pub fn http_resp_to_error_response(
err_resp err_resp
} }
pub fn err_transfer_acceleration_bucket(bucket_name: &str) -> ErrorResponse {
ErrorResponse {
status_code: StatusCode::BAD_REQUEST,
code: S3ErrorCode::InvalidArgument,
message: "The name of the bucket used for Transfer Acceleration must be DNS-compliant and must not contain periods .."
.to_string(),
bucket_name: bucket_name.to_string(),
..Default::default()
}
}
pub fn err_entity_too_large(total_size: i64, max_object_size: i64, bucket_name: &str, object_name: &str) -> ErrorResponse { pub fn err_entity_too_large(total_size: i64, max_object_size: i64, bucket_name: &str, object_name: &str) -> ErrorResponse {
let msg = format!( let msg = format!(
"Your proposed upload size {} exceeds the maximum allowed object size {} for single PUT operation.", "Your proposed upload size {} exceeds the maximum allowed object size {} for single PUT operation.",
@@ -295,16 +284,6 @@ pub fn err_invalid_argument(message: &str) -> ErrorResponse {
} }
} }
pub fn err_api_not_supported(message: &str) -> ErrorResponse {
ErrorResponse {
status_code: StatusCode::NOT_IMPLEMENTED,
code: S3ErrorCode::Custom("APINotSupported".into()),
message: message.to_string(),
request_id: "rustfs".to_string(),
..Default::default()
}
}
#[cfg(test)] #[cfg(test)]
mod tests { mod tests {
use super::*; use super::*;
@@ -135,6 +135,10 @@ impl Object {
Self { ..Default::default() } Self { ..Default::default() }
} }
#[allow(
dead_code,
reason = "MinIO-parity reader surface with no caller in this port (backlog#1823)"
)]
fn do_get_request(&self, request: &GetRequest) -> Result<GetResponse, std::io::Error> { fn do_get_request(&self, request: &GetRequest) -> Result<GetResponse, std::io::Error> {
let _ = request.did_offset_change; let _ = request.did_offset_change;
let _ = request.offset; let _ = request.offset;
@@ -150,12 +154,20 @@ impl Object {
)) ))
} }
#[allow(
dead_code,
reason = "MinIO-parity Object reader method with no caller in this port (backlog#1823)"
)]
fn set_offset(&mut self, bytes_read: i64) -> Result<(), std::io::Error> { fn set_offset(&mut self, bytes_read: i64) -> Result<(), std::io::Error> {
self.curr_offset += bytes_read; self.curr_offset += bytes_read;
Ok(()) Ok(())
} }
#[allow(
dead_code,
reason = "MinIO-parity Object reader method with no caller in this port (backlog#1823)"
)]
fn read(&mut self, b: &[u8]) -> Result<i64, std::io::Error> { fn read(&mut self, b: &[u8]) -> Result<i64, std::io::Error> {
let mut read_req = GetRequest { let mut read_req = GetRequest {
is_read_op: true, is_read_op: true,
@@ -180,6 +192,10 @@ impl Object {
Ok(response.size) Ok(response.size)
} }
#[allow(
dead_code,
reason = "MinIO-parity Object reader method with no caller in this port (backlog#1823)"
)]
fn stat(&self) -> Result<ObjectInfo, std::io::Error> { fn stat(&self) -> Result<ObjectInfo, std::io::Error> {
if !self.is_started || !self.object_info_set { if !self.is_started || !self.object_info_set {
let _ = self.do_get_request(&GetRequest { let _ = self.do_get_request(&GetRequest {
@@ -192,6 +208,10 @@ impl Object {
Ok(self.object_info.clone()) Ok(self.object_info.clone())
} }
#[allow(
dead_code,
reason = "MinIO-parity Object reader method with no caller in this port (backlog#1823)"
)]
fn read_at(&mut self, b: &[u8], offset: i64) -> Result<i64, std::io::Error> { fn read_at(&mut self, b: &[u8], offset: i64) -> Result<i64, std::io::Error> {
self.curr_offset = offset; self.curr_offset = offset;
@@ -219,6 +239,10 @@ impl Object {
Ok(response.size) Ok(response.size)
} }
#[allow(
dead_code,
reason = "MinIO-parity Object reader method with no caller in this port (backlog#1823)"
)]
fn seek(&mut self, offset: i64, whence: i64) -> Result<i64, std::io::Error> { fn seek(&mut self, offset: i64, whence: i64) -> Result<i64, std::io::Error> {
if !self.is_started || !self.object_info_set { if !self.is_started || !self.object_info_set {
let seek_req = GetRequest { let seek_req = GetRequest {
@@ -253,6 +277,10 @@ impl Object {
Ok(self.curr_offset) Ok(self.curr_offset)
} }
#[allow(
dead_code,
reason = "MinIO-parity Object reader method with no caller in this port (backlog#1823)"
)]
fn close(&mut self) -> Result<(), std::io::Error> { fn close(&mut self) -> Result<(), std::io::Error> {
self.is_closed = true; self.is_closed = true;
Ok(()) Ok(())
+1 -1
View File
@@ -37,7 +37,7 @@ use crate::client::{
api_put_object_common::optimal_part_info, api_put_object_common::optimal_part_info,
api_put_object_multipart::UploadPartParams, api_put_object_multipart::UploadPartParams,
api_s3_datatypes::{CompleteMultipartUpload, CompletePart, ObjectPart}, api_s3_datatypes::{CompleteMultipartUpload, CompletePart, ObjectPart},
constants::{ISO8601_DATEFORMAT, MAX_MULTIPART_PUT_OBJECT_SIZE, MIN_PART_SIZE, TOTAL_WORKERS}, constants::{ISO8601_DATEFORMAT, MAX_MULTIPART_PUT_OBJECT_SIZE, MIN_PART_SIZE},
credentials::SignatureType, credentials::SignatureType,
transition_api::{ReaderImpl, TransitionClient, UploadInfo}, transition_api::{ReaderImpl, TransitionClient, UploadInfo},
utils::{is_amz_header, is_minio_header, is_rustfs_header, is_standard_header, is_storageclass_header}, utils::{is_amz_header, is_minio_header, is_rustfs_header, is_standard_header, is_storageclass_header},
@@ -30,10 +30,6 @@ pub fn is_object(reader: &ReaderImpl) -> bool {
matches!(reader, ReaderImpl::ObjectBody(_)) matches!(reader, ReaderImpl::ObjectBody(_))
} }
pub fn is_read_at(reader: ReaderImpl) -> bool {
matches!(reader, ReaderImpl::ObjectBody(_))
}
pub fn optimal_part_info(object_size: i64, configured_part_size: u64) -> Result<(i64, i64, i64), std::io::Error> { pub fn optimal_part_info(object_size: i64, configured_part_size: u64) -> Result<(i64, i64, i64), std::io::Error> {
let unknown_size; let unknown_size;
let mut object_size = object_size; let mut object_size = object_size;
@@ -81,18 +81,6 @@ async fn read_multipart_part(reader: &mut ReaderImpl, want: usize) -> Result<Vec
} }
} }
pub struct UploadedPartRes {
pub error: std::io::Error,
pub part_num: i64,
pub size: i64,
pub part: ObjectPart,
}
pub struct UploadPartReq {
pub part_num: i64,
pub part: ObjectPart,
}
impl TransitionClient { impl TransitionClient {
pub async fn put_object_multipart_stream( pub async fn put_object_multipart_stream(
self: Arc<Self>, self: Arc<Self>,
+18 -38
View File
@@ -29,10 +29,6 @@ use crate::client::utils::base64_decode;
use super::transition_api; use super::transition_api;
pub struct ListAllMyBucketsResult {
pub owner: Owner,
}
#[derive(Debug, Default, Serialize, Deserialize)] #[derive(Debug, Default, Serialize, Deserialize)]
pub struct CommonPrefix { pub struct CommonPrefix {
pub prefix: String, pub prefix: String,
@@ -89,6 +85,10 @@ pub struct ListVersionsResult {
pub next_version_id_marker: String, pub next_version_id_marker: String,
} }
#[allow(
dead_code,
reason = "fields of a MinIO-parity list result that this port builds but never reads back (backlog#1823)"
)]
pub struct ListBucketResult { pub struct ListBucketResult {
common_prefixes: Vec<CommonPrefix>, common_prefixes: Vec<CommonPrefix>,
contents: Vec<transition_api::ObjectInfo>, contents: Vec<transition_api::ObjectInfo>,
@@ -102,6 +102,10 @@ pub struct ListBucketResult {
prefix: String, prefix: String,
} }
#[allow(
dead_code,
reason = "fields of a MinIO-parity list result that this port builds but never reads back (backlog#1823)"
)]
pub struct ListMultipartUploadsResult { pub struct ListMultipartUploadsResult {
bucket: String, bucket: String,
key_marker: String, key_marker: String,
@@ -117,16 +121,15 @@ pub struct ListMultipartUploadsResult {
common_prefixes: Vec<CommonPrefix>, common_prefixes: Vec<CommonPrefix>,
} }
#[allow(
dead_code,
reason = "fields of a MinIO-parity list result that this port builds but never reads back (backlog#1823)"
)]
pub struct Initiator { pub struct Initiator {
id: String, id: String,
display_name: String, display_name: String,
} }
pub struct CopyObjectResult {
pub etag: String,
pub last_modified: OffsetDateTime,
}
#[derive(Debug, Clone)] #[derive(Debug, Clone)]
pub struct ObjectPart { pub struct ObjectPart {
pub etag: String, pub etag: String,
@@ -260,6 +263,7 @@ pub struct CompletePart {
} }
impl CompletePart { impl CompletePart {
#[allow(dead_code, reason = "MinIO-parity accessor with no caller in this port (backlog#1823)")]
fn checksum(&self, t: &ChecksumMode) -> String { fn checksum(&self, t: &ChecksumMode) -> String {
match t { match t {
ChecksumMode::ChecksumCRC32C => { ChecksumMode::ChecksumCRC32C => {
@@ -284,11 +288,6 @@ impl CompletePart {
} }
} }
pub struct CopyObjectPartResult {
pub etag: String,
pub last_modified: OffsetDateTime,
}
#[derive(Debug, Default, serde::Serialize)] #[derive(Debug, Default, serde::Serialize)]
#[serde(rename = "CompleteMultipartUpload")] #[serde(rename = "CompleteMultipartUpload")]
pub struct CompleteMultipartUpload { pub struct CompleteMultipartUpload {
@@ -357,10 +356,10 @@ impl CompleteMultipartUpload {
} }
} }
pub struct CreateBucketConfiguration { #[allow(
pub location: String, dead_code,
} reason = "live via quick_xml::de::from_str in bucket_cache.rs; serde deserialization is not a construction (backlog#1823)"
)]
#[derive(serde::Serialize)] #[derive(serde::Serialize)]
pub struct DeleteObject { pub struct DeleteObject {
//api has //api has
@@ -368,21 +367,6 @@ pub struct DeleteObject {
pub version_id: String, pub version_id: String,
} }
pub struct DeletedObject {
//s3s has
pub key: String,
pub version_id: String,
pub deletemarker: bool,
pub deletemarker_version_id: String,
}
pub struct NonDeletedObject {
pub key: String,
pub code: String,
pub message: String,
pub version_id: String,
}
#[derive(serde::Serialize)] #[derive(serde::Serialize)]
pub struct DeleteMultiObjects { pub struct DeleteMultiObjects {
pub quiet: bool, pub quiet: bool,
@@ -402,6 +386,7 @@ impl DeleteMultiObjects {
Ok(buf) Ok(buf)
} }
#[allow(dead_code, reason = "MinIO-parity XML helper with no caller in this port (backlog#1823)")]
pub fn unmarshal(buf: &[u8]) -> Result<Self, std::io::Error> { pub fn unmarshal(buf: &[u8]) -> Result<Self, std::io::Error> {
#[derive(Debug, Deserialize)] #[derive(Debug, Deserialize)]
struct WireDeleteObject { struct WireDeleteObject {
@@ -436,8 +421,3 @@ impl DeleteMultiObjects {
}) })
} }
} }
pub struct DeleteMultiObjectsResult {
pub deleted_objects: Vec<DeletedObject>,
pub undeleted_objects: Vec<NonDeletedObject>,
}
+4
View File
@@ -365,6 +365,10 @@ mod tests {
pub struct Checksum { pub struct Checksum {
checksum_type: ChecksumMode, checksum_type: ChecksumMode,
r: Vec<u8>, r: Vec<u8>,
#[allow(
dead_code,
reason = "checksum bookkeeping field kept beside the value it guards (backlog#1823)"
)]
computed: bool, computed: bool,
} }
-3
View File
@@ -32,8 +32,5 @@ pub const MAX_MULTIPART_PUT_OBJECT_SIZE: i64 = 1024 * 1024 * 1024 * 1024 * 5;
pub const UNSIGNED_PAYLOAD: &str = "UNSIGNED-PAYLOAD"; pub const UNSIGNED_PAYLOAD: &str = "UNSIGNED-PAYLOAD";
pub const UNSIGNED_PAYLOAD_TRAILER: &str = "STREAMING-UNSIGNED-PAYLOAD-TRAILER"; pub const UNSIGNED_PAYLOAD_TRAILER: &str = "STREAMING-UNSIGNED-PAYLOAD-TRAILER";
pub const TOTAL_WORKERS: i64 = 4;
pub const SIGN_V4_ALGORITHM: &str = "AWS4-HMAC-SHA256";
pub const ISO8601_DATEFORMAT: &[FormatItem<'_>] = pub const ISO8601_DATEFORMAT: &[FormatItem<'_>] =
format_description!("[year]-[month]-[day]T[hour]:[minute]:[second].[subsecond]Z"); format_description!("[year]-[month]-[day]T[hour]:[minute]:[second].[subsecond]Z");
+12 -19
View File
@@ -67,6 +67,10 @@ impl<P: Provider + Default> Credentials<P> {
Ok(self.creds.clone()) Ok(self.creds.clone())
} }
#[allow(
dead_code,
reason = "MinIO-parity credential surface with no caller in this port (backlog#1823)"
)]
fn expire(&mut self) { fn expire(&mut self) {
self.force_refresh = true; self.force_refresh = true;
} }
@@ -133,6 +137,10 @@ impl Provider for Static {
#[derive(Debug, Clone, Default)] #[derive(Debug, Clone, Default)]
pub struct STSError { pub struct STSError {
#[allow(
dead_code,
reason = "MinIO-parity STS error detail that this port never reads back (backlog#1823)"
)]
pub r#type: String, pub r#type: String,
pub code: String, pub code: String,
pub message: String, pub message: String,
@@ -141,6 +149,10 @@ pub struct STSError {
#[derive(Debug, Clone, thiserror::Error)] #[derive(Debug, Clone, thiserror::Error)]
pub struct ErrorResponse { pub struct ErrorResponse {
pub sts_error: STSError, pub sts_error: STSError,
#[allow(
dead_code,
reason = "MinIO-parity STS error detail that this port never reads back (backlog#1823)"
)]
pub request_id: String, pub request_id: String,
} }
@@ -158,22 +170,3 @@ impl ErrorResponse {
return self.sts_error.message.clone(); return self.sts_error.message.clone();
} }
} }
pub fn xml_decoder<T>(body: &[u8]) -> Result<T, Error>
where
for<'de> T: Deserialize<'de>,
{
match std::str::from_utf8(body) {
Ok(xml_body) => quick_xml::de::from_str::<T>(xml_body).map_err(|err| Error::new(ErrorKind::InvalidData, err.to_string())),
Err(err) => Err(Error::new(ErrorKind::InvalidData, err.to_string())),
}
}
pub fn xml_decode_and_body<T>(body_reader: &[u8]) -> Result<(Vec<u8>, T), std::io::Error>
where
for<'de> T: Deserialize<'de>,
{
let body = body_reader.to_vec();
let parsed = xml_decoder(&body)?;
Ok((body, parsed))
}
-1
View File
@@ -13,7 +13,6 @@
// limitations under the License. // limitations under the License.
// #730: S3 client compatibility models are kept while ECStore callers move to narrower facades. // #730: S3 client compatibility models are kept while ECStore callers move to narrower facades.
#![allow(dead_code)]
pub mod admin_handler_utils; pub mod admin_handler_utils;
pub mod api_error_response; pub mod api_error_response;
@@ -77,39 +77,6 @@ fn part_number_to_rangespec(oi: ObjectInfo, part_number: usize) -> Option<HTTPRa
}) })
} }
fn get_compressed_offsets(oi: ObjectInfo, offset: i64) -> (i64, i64, i64, i64, u64) {
let mut skip_length: i64 = 0;
let mut cumulative_actual_size: i64 = 0;
let mut first_part_idx: i64 = 0;
let mut compressed_offset: i64 = 0;
let mut part_skip: i64 = 0;
let mut decrypt_skip: i64 = 0;
let mut seq_num: u64 = 0;
for (i, part) in oi.parts.iter().enumerate() {
cumulative_actual_size += part.actual_size as i64;
if cumulative_actual_size <= offset {
compressed_offset += part.size as i64;
} else {
first_part_idx = i as i64;
skip_length = cumulative_actual_size - part.actual_size as i64;
break;
}
}
skip_length = offset - skip_length;
let parts: &[ObjectPartInfo] = &oi.parts;
if skip_length > 0
&& parts.len() > first_part_idx as usize
&& parts[first_part_idx as usize].index.as_ref().is_some_and(|idx| idx.len() > 0)
{
let _ = part_skip;
let _ = decrypt_skip;
let _ = seq_num;
}
(compressed_offset, part_skip, first_part_idx, decrypt_skip, seq_num)
}
pub fn new_getobjectreader<'a>( pub fn new_getobjectreader<'a>(
rs: &Option<HTTPRangeSpec>, rs: &Option<HTTPRangeSpec>,
oi: &'a ObjectInfo, oi: &'a ObjectInfo,
@@ -23,6 +23,7 @@ const X_OBS_VERSION_ID: &str = "x-obs-version-id";
const MAX_REMOTE_VERSION_ID_LEN: usize = 1024; const MAX_REMOTE_VERSION_ID_LEN: usize = 1024;
#[derive(Clone, Copy, Debug, Eq, PartialEq)] #[derive(Clone, Copy, Debug, Eq, PartialEq)]
#[allow(dead_code, reason = "bucket versioning states kept as a complete vocabulary (backlog#1823)")]
pub(crate) enum BucketVersioningState { pub(crate) enum BucketVersioningState {
Unknown, Unknown,
Disabled, Disabled,
@@ -47,6 +48,7 @@ impl RemoteVersion {
} }
} }
#[allow(dead_code, reason = "MinIO-parity accessor with no caller in this port (backlog#1823)")]
pub(crate) fn exact_request_id(&self) -> Result<Option<&str>, Error> { pub(crate) fn exact_request_id(&self) -> Result<Option<&str>, Error> {
match self { match self {
Self::Unknown => Err(Error::new( Self::Unknown => Err(Error::new(
@@ -101,6 +101,10 @@ where
const C_UNKNOWN: i32 = -1; const C_UNKNOWN: i32 = -1;
const C_OFFLINE: i32 = 0; const C_OFFLINE: i32 = 0;
#[allow(
dead_code,
reason = "reachable only from the unused transition client methods below (backlog#1823)"
)]
const C_ONLINE: i32 = 1; const C_ONLINE: i32 = 1;
fn invalid_utf8_header_error(scope: &str, header_name: &str) -> std::io::Error { fn invalid_utf8_header_error(scope: &str, header_name: &str) -> std::io::Error {
@@ -320,6 +324,10 @@ impl TransitionClient {
Ok(client) Ok(client)
} }
#[allow(
dead_code,
reason = "MinIO-parity transition client surface with no caller in this port (backlog#1823)"
)]
fn endpoint_url(&self) -> Url { fn endpoint_url(&self) -> Url {
self.endpoint_url.clone() self.endpoint_url.clone()
} }
@@ -348,12 +356,20 @@ impl TransitionClient {
.to_string()) .to_string())
} }
#[allow(
dead_code,
reason = "MinIO-parity transition client method with no caller in this port (backlog#1823)"
)]
fn trace_errors_only_off(&self) { fn trace_errors_only_off(&self) {
if let Ok(mut trace_errors_only) = self.trace_errors_only.lock() { if let Ok(mut trace_errors_only) = self.trace_errors_only.lock() {
*trace_errors_only = false; *trace_errors_only = false;
} }
} }
#[allow(
dead_code,
reason = "MinIO-parity transition client method with no caller in this port (backlog#1823)"
)]
fn trace_off(&self) { fn trace_off(&self) {
if let Ok(mut is_trace_enabled) = self.is_trace_enabled.lock() { if let Ok(mut is_trace_enabled) = self.is_trace_enabled.lock() {
*is_trace_enabled = false; *is_trace_enabled = false;
@@ -363,12 +379,20 @@ impl TransitionClient {
} }
} }
#[allow(
dead_code,
reason = "MinIO-parity transition client method with no caller in this port (backlog#1823)"
)]
fn set_s3_transfer_accelerate(&self, accelerate_endpoint: &str) { fn set_s3_transfer_accelerate(&self, accelerate_endpoint: &str) {
if let Ok(mut endpoint) = self.s3_accelerate_endpoint.lock() { if let Ok(mut endpoint) = self.s3_accelerate_endpoint.lock() {
*endpoint = accelerate_endpoint.to_string(); *endpoint = accelerate_endpoint.to_string();
} }
} }
#[allow(
dead_code,
reason = "MinIO-parity transition client method with no caller in this port (backlog#1823)"
)]
fn set_s3_enable_dual_stack(&self, enabled: bool) { fn set_s3_enable_dual_stack(&self, enabled: bool) {
if let Ok(mut dual_stack) = self.s3_dual_stack_enabled.lock() { if let Ok(mut dual_stack) = self.s3_dual_stack_enabled.lock() {
*dual_stack = enabled; *dual_stack = enabled;
@@ -398,10 +422,18 @@ impl TransitionClient {
(hash_algos, hash_sums) (hash_algos, hash_sums)
} }
#[allow(
dead_code,
reason = "MinIO-parity transition client method with no caller in this port (backlog#1823)"
)]
fn is_online(&self) -> bool { fn is_online(&self) -> bool {
!self.is_offline() !self.is_offline()
} }
#[allow(
dead_code,
reason = "MinIO-parity transition client method with no caller in this port (backlog#1823)"
)]
fn mark_offline(&self) { fn mark_offline(&self) {
self.health_status self.health_status
.compare_exchange(C_ONLINE, C_OFFLINE, Ordering::SeqCst, Ordering::SeqCst); .compare_exchange(C_ONLINE, C_OFFLINE, Ordering::SeqCst, Ordering::SeqCst);
@@ -411,10 +443,18 @@ impl TransitionClient {
self.health_status.load(Ordering::SeqCst) == C_OFFLINE self.health_status.load(Ordering::SeqCst) == C_OFFLINE
} }
#[allow(
dead_code,
reason = "MinIO-parity transition client method with no caller in this port (backlog#1823)"
)]
fn health_check(hc_duration: Duration) { fn health_check(hc_duration: Duration) {
let _ = hc_duration; let _ = hc_duration;
} }
#[allow(
dead_code,
reason = "MinIO-parity transition client method with no caller in this port (backlog#1823)"
)]
fn dump_http(&self, req: &Request<s3s::Body>, resp: &Response<Incoming>) -> Result<(), std::io::Error> { fn dump_http(&self, req: &Request<s3s::Body>, resp: &Response<Incoming>) -> Result<(), std::io::Error> {
let mut resp_trace: Vec<u8>; let mut resp_trace: Vec<u8>;
@@ -1102,6 +1142,7 @@ impl Default for ObjectInfo {
} }
impl ObjectInfo { impl ObjectInfo {
#[allow(dead_code, reason = "MinIO-parity accessor with no caller in this port (backlog#1823)")]
pub(crate) fn remote_version( pub(crate) fn remote_version(
&self, &self,
capabilities: ProviderVersionCapabilities, capabilities: ProviderVersionCapabilities,
-4
View File
@@ -48,10 +48,6 @@ lazy_static! {
}; };
} }
pub fn is_standard_query_value(qs_key: &str) -> bool {
SUPPORTED_QUERY_VALUES[qs_key]
}
pub fn is_storageclass_header(header_key: &str) -> bool { pub fn is_storageclass_header(header_key: &str) -> bool {
header_key.to_lowercase() == X_AMZ_STORAGE_CLASS.as_str().to_lowercase() header_key.to_lowercase() == X_AMZ_STORAGE_CLASS.as_str().to_lowercase()
} }
+37
View File
@@ -190,6 +190,17 @@ pub(crate) const GET_METADATA_CACHE_REASON_VERSION_SUSPENDED: &str = "version_su
pub(crate) const GET_METADATA_CACHE_REASON_VERSIONED: &str = "versioned"; pub(crate) const GET_METADATA_CACHE_REASON_VERSIONED: &str = "versioned";
pub(crate) const GET_METADATA_EARLY_STOP_REASON_CONFLICTING_METADATA: &str = "conflicting_metadata"; pub(crate) const GET_METADATA_EARLY_STOP_REASON_CONFLICTING_METADATA: &str = "conflicting_metadata";
pub(crate) const GET_METADATA_EARLY_STOP_REASON_DELETE_MARKER: &str = "delete_marker"; pub(crate) const GET_METADATA_EARLY_STOP_REASON_DELETE_MARKER: &str = "delete_marker";
pub(crate) const GET_METADATA_EARLY_STOP_REASON_DATA_READ_INLINE_BODY_VERIFY: &str = "data_read_inline_body_verify";
pub(crate) const GET_METADATA_EARLY_STOP_REASON_DATA_READ_INLINE_DELETED: &str = "data_read_inline_deleted";
pub(crate) const GET_METADATA_EARLY_STOP_REASON_DATA_READ_INLINE_GEOMETRY: &str = "data_read_inline_geometry";
pub(crate) const GET_METADATA_EARLY_STOP_REASON_DATA_READ_INLINE_IDENTITY_MISMATCH: &str = "data_read_inline_identity_mismatch";
pub(crate) const GET_METADATA_EARLY_STOP_REASON_DATA_READ_INLINE_MISSING_PAYLOAD: &str = "data_read_inline_missing_payload";
pub(crate) const GET_METADATA_EARLY_STOP_REASON_DATA_READ_INLINE_MISSING_SHARD: &str = "data_read_inline_missing_shard";
pub(crate) const GET_METADATA_EARLY_STOP_REASON_DATA_READ_INLINE_NOT_INLINE: &str = "data_read_inline_not_inline";
pub(crate) const GET_METADATA_EARLY_STOP_REASON_DATA_READ_INLINE_PART_SHAPE: &str = "data_read_inline_part_shape";
pub(crate) const GET_METADATA_EARLY_STOP_REASON_DATA_READ_INLINE_REMOTE: &str = "data_read_inline_remote";
pub(crate) const GET_METADATA_EARLY_STOP_REASON_DATA_READ_INLINE_SIZE: &str = "data_read_inline_size";
pub(crate) const GET_METADATA_EARLY_STOP_REASON_DATA_READ_INLINE_TRANSFORMED: &str = "data_read_inline_transformed";
pub(crate) const GET_METADATA_EARLY_STOP_REASON_ERROR: &str = "error"; pub(crate) const GET_METADATA_EARLY_STOP_REASON_ERROR: &str = "error";
pub(crate) const GET_METADATA_EARLY_STOP_REASON_INSUFFICIENT_QUORUM: &str = "insufficient_quorum"; pub(crate) const GET_METADATA_EARLY_STOP_REASON_INSUFFICIENT_QUORUM: &str = "insufficient_quorum";
pub(crate) const GET_METADATA_EARLY_STOP_REASON_NOT_FOUND: &str = "not_found"; pub(crate) const GET_METADATA_EARLY_STOP_REASON_NOT_FOUND: &str = "not_found";
@@ -551,6 +562,32 @@ mod tests {
assert_eq!(GET_METADATA_CACHE_REASON_VERSIONED, "versioned"); assert_eq!(GET_METADATA_CACHE_REASON_VERSIONED, "versioned");
assert_eq!(GET_METADATA_EARLY_STOP_REASON_CONFLICTING_METADATA, "conflicting_metadata"); assert_eq!(GET_METADATA_EARLY_STOP_REASON_CONFLICTING_METADATA, "conflicting_metadata");
assert_eq!(GET_METADATA_EARLY_STOP_REASON_DELETE_MARKER, "delete_marker"); assert_eq!(GET_METADATA_EARLY_STOP_REASON_DELETE_MARKER, "delete_marker");
assert_eq!(
GET_METADATA_EARLY_STOP_REASON_DATA_READ_INLINE_BODY_VERIFY,
"data_read_inline_body_verify"
);
assert_eq!(GET_METADATA_EARLY_STOP_REASON_DATA_READ_INLINE_DELETED, "data_read_inline_deleted");
assert_eq!(GET_METADATA_EARLY_STOP_REASON_DATA_READ_INLINE_GEOMETRY, "data_read_inline_geometry");
assert_eq!(
GET_METADATA_EARLY_STOP_REASON_DATA_READ_INLINE_IDENTITY_MISMATCH,
"data_read_inline_identity_mismatch"
);
assert_eq!(
GET_METADATA_EARLY_STOP_REASON_DATA_READ_INLINE_MISSING_PAYLOAD,
"data_read_inline_missing_payload"
);
assert_eq!(
GET_METADATA_EARLY_STOP_REASON_DATA_READ_INLINE_MISSING_SHARD,
"data_read_inline_missing_shard"
);
assert_eq!(GET_METADATA_EARLY_STOP_REASON_DATA_READ_INLINE_NOT_INLINE, "data_read_inline_not_inline");
assert_eq!(GET_METADATA_EARLY_STOP_REASON_DATA_READ_INLINE_PART_SHAPE, "data_read_inline_part_shape");
assert_eq!(GET_METADATA_EARLY_STOP_REASON_DATA_READ_INLINE_REMOTE, "data_read_inline_remote");
assert_eq!(GET_METADATA_EARLY_STOP_REASON_DATA_READ_INLINE_SIZE, "data_read_inline_size");
assert_eq!(
GET_METADATA_EARLY_STOP_REASON_DATA_READ_INLINE_TRANSFORMED,
"data_read_inline_transformed"
);
assert_eq!(GET_METADATA_EARLY_STOP_REASON_ERROR, "error"); assert_eq!(GET_METADATA_EARLY_STOP_REASON_ERROR, "error");
assert_eq!(GET_METADATA_EARLY_STOP_REASON_INSUFFICIENT_QUORUM, "insufficient_quorum"); assert_eq!(GET_METADATA_EARLY_STOP_REASON_INSUFFICIENT_QUORUM, "insufficient_quorum");
assert_eq!(GET_METADATA_EARLY_STOP_REASON_NOT_FOUND, "not_found"); assert_eq!(GET_METADATA_EARLY_STOP_REASON_NOT_FOUND, "not_found");
+13 -33
View File
@@ -637,14 +637,23 @@ impl Default for DiskOperationMetrics {
} }
impl DiskOperationMetrics { impl DiskOperationMetrics {
#[allow(
dead_code,
reason = "internal metrics recorder reached only from record() below (backlog#1823)"
)]
fn record_call(&mut self) { fn record_call(&mut self) {
self.lifetime_calls.fetch_add(1, Ordering::Relaxed); self.lifetime_calls.fetch_add(1, Ordering::Relaxed);
} }
#[allow(
dead_code,
reason = "internal metrics recorder reached only from record() below (backlog#1823)"
)]
fn record_latency(&mut self, now_sec: u64, elapsed: Duration) { fn record_latency(&mut self, now_sec: u64, elapsed: Duration) {
self.record_latency_atomic(now_sec, elapsed); self.record_latency_atomic(now_sec, elapsed);
} }
#[allow(dead_code, reason = "metrics roll-up with no caller in this port (backlog#1823)")]
fn record(&mut self, now_sec: u64, elapsed: Duration) { fn record(&mut self, now_sec: u64, elapsed: Duration) {
self.record_call(); self.record_call();
self.record_latency(now_sec, elapsed); self.record_latency(now_sec, elapsed);
@@ -770,6 +779,7 @@ impl DiskHealthTracker {
} }
/// Set disk as faulty /// Set disk as faulty
#[allow(dead_code, reason = "asserted by this file's tests (backlog#1823)")]
pub fn set_faulty(&self) { pub fn set_faulty(&self) {
self.status.store(DISK_HEALTH_FAULTY, Ordering::Release); self.status.store(DISK_HEALTH_FAULTY, Ordering::Release);
} }
@@ -850,6 +860,7 @@ impl DiskHealthTracker {
became_offline became_offline
} }
#[allow(dead_code, reason = "asserted by this file's tests (backlog#1823)")]
pub fn mark_offline(&self, endpoint: &Endpoint, reason: &'static str) -> bool { pub fn mark_offline(&self, endpoint: &Endpoint, reason: &'static str) -> bool {
let current = self.runtime_state(); let current = self.runtime_state();
if current == RuntimeDriveHealthState::Offline { if current == RuntimeDriveHealthState::Offline {
@@ -980,11 +991,13 @@ impl DiskHealthTracker {
} }
/// Get waiting operations count /// Get waiting operations count
#[allow(dead_code, reason = "asserted by this file's tests (backlog#1823)")]
pub fn waiting_count(&self) -> u32 { pub fn waiting_count(&self) -> u32 {
self.waiting.load(Ordering::Relaxed) self.waiting.load(Ordering::Relaxed)
} }
/// Get last success timestamp /// Get last success timestamp
#[allow(dead_code, reason = "asserted by this file's tests (backlog#1823)")]
pub fn last_success(&self) -> i64 { pub fn last_success(&self) -> i64 {
self.last_success.load(Ordering::Acquire) self.last_success.load(Ordering::Acquire)
} }
@@ -1026,21 +1039,6 @@ impl Default for DiskHealthTracker {
} }
} }
/// Health check context key for tracking disk operations
#[derive(Debug, Clone)]
struct HealthDiskCtxKey;
#[derive(Debug)]
struct HealthDiskCtxValue {
last_success: Arc<AtomicI64>,
}
impl HealthDiskCtxValue {
fn log_success(&self) {
self.last_success.store(current_unix_nanos(), Ordering::Relaxed);
}
}
/// LocalDiskWrapper wraps a DiskStore with health tracking capabilities. /// LocalDiskWrapper wraps a DiskStore with health tracking capabilities.
/// This is similar to Go's xlStorageDiskIDCheck. /// This is similar to Go's xlStorageDiskIDCheck.
#[derive(Debug, Clone)] #[derive(Debug, Clone)]
@@ -1072,10 +1070,6 @@ impl LocalDiskWrapper {
) )
} }
pub(crate) fn new_with_health(disk: Arc<LocalDisk>, health_check: bool, health: Arc<DiskHealthTracker>) -> Self {
Self::new_with_health_and_metrics(disk, health_check, health, Arc::new(DiskHealthMetricEpoch::default()))
}
pub(crate) fn new_with_reconnect_state( pub(crate) fn new_with_reconnect_state(
disk: Arc<LocalDisk>, disk: Arc<LocalDisk>,
health_check: bool, health_check: bool,
@@ -1438,20 +1432,6 @@ impl LocalDiskWrapper {
} }
} }
async fn check_id(&self, want_id: Option<Uuid>) -> Result<()> {
if want_id.is_none() {
return Ok(());
}
let stored_disk_id = self.disk.get_disk_id().await?;
if stored_disk_id != want_id {
return Err(Error::other(format!("Disk ID mismatch wanted {want_id:?}, got {stored_disk_id:?}")));
}
Ok(())
}
/// Check if disk ID is stale /// Check if disk ID is stale
async fn check_disk_stale(&self) -> Result<()> { async fn check_disk_stale(&self) -> Result<()> {
let Some(current_disk_id) = *self.disk_id.read().await else { let Some(current_disk_id) = *self.disk_id.read().await else {
+1
View File
@@ -48,6 +48,7 @@ pub fn to_volume_error(io_err: std::io::Error) -> std::io::Error {
} }
} }
#[allow(dead_code, reason = "asserted by this file's tests (backlog#1823)")]
pub fn to_disk_error(io_err: std::io::Error) -> std::io::Error { pub fn to_disk_error(io_err: std::io::Error) -> std::io::Error {
match io_err.kind() { match io_err.kind() {
std::io::ErrorKind::NotFound => DiskError::DiskNotFound.into(), std::io::ErrorKind::NotFound => DiskError::DiskNotFound.into(),
+1
View File
@@ -178,6 +178,7 @@ pub async fn remove(path: impl AsRef<Path>) -> io::Result<()> {
} }
} }
#[allow(dead_code, reason = "asserted by this file's tests (backlog#1823)")]
pub async fn remove_all(path: impl AsRef<Path>) -> io::Result<()> { pub async fn remove_all(path: impl AsRef<Path>) -> io::Result<()> {
// Try remove_file first; fall back to remove_dir_all if it's a directory // Try remove_file first; fall back to remove_dir_all if it's a directory
match fs::remove_file(path.as_ref()).await { match fs::remove_file(path.as_ref()).await {
+69
View File
@@ -665,6 +665,7 @@ async fn remove_empty_directory_tree_under_mount_lease(
} }
#[cfg(unix)] #[cfg(unix)]
#[allow(dead_code, reason = "asserted by this file's tests (backlog#1823)")]
async fn remove_empty_directory_tree_with( async fn remove_empty_directory_tree_with(
root: &Path, root: &Path,
before_descend: impl FnMut(&Path) -> std::io::Result<()>, before_descend: impl FnMut(&Path) -> std::io::Result<()>,
@@ -1016,13 +1017,29 @@ fn record_direct_read_page_fault_delta(path: &'static str, stage: &'static str,
/// When enabled, shard reads bypass the page cache using O_DIRECT flag. /// When enabled, shard reads bypass the page cache using O_DIRECT flag.
/// Requires aligned buffers (typically 512 bytes or 4096 bytes). /// Requires aligned buffers (typically 512 bytes or 4096 bytes).
/// Default: false (uses page cache via mmap/pread). /// Default: false (uses page cache via mmap/pread).
#[allow(
dead_code,
reason = "platform-conditional: production callers are inside #[cfg(target_os = \"linux\")] blocks, so this reads as dead on non-Linux hosts (backlog#1823)"
)]
const ENV_RUSTFS_OBJECT_DIRECT_IO_READ_ENABLE: &str = "RUSTFS_OBJECT_DIRECT_IO_READ_ENABLE"; const ENV_RUSTFS_OBJECT_DIRECT_IO_READ_ENABLE: &str = "RUSTFS_OBJECT_DIRECT_IO_READ_ENABLE";
#[allow(
dead_code,
reason = "platform-conditional: production callers are inside #[cfg(target_os = \"linux\")] blocks, so this reads as dead on non-Linux hosts (backlog#1823)"
)]
const DEFAULT_RUSTFS_OBJECT_DIRECT_IO_READ_ENABLE: bool = false; const DEFAULT_RUSTFS_OBJECT_DIRECT_IO_READ_ENABLE: bool = false;
/// Minimum shard size threshold for O_DIRECT reads. /// Minimum shard size threshold for O_DIRECT reads.
/// Only shards larger than this threshold will use O_DIRECT. /// Only shards larger than this threshold will use O_DIRECT.
/// Default: 4MB. /// Default: 4MB.
#[allow(
dead_code,
reason = "platform-conditional: production callers are inside #[cfg(target_os = \"linux\")] blocks, so this reads as dead on non-Linux hosts (backlog#1823)"
)]
const ENV_RUSTFS_OBJECT_DIRECT_IO_READ_THRESHOLD: &str = "RUSTFS_OBJECT_DIRECT_IO_READ_THRESHOLD"; const ENV_RUSTFS_OBJECT_DIRECT_IO_READ_THRESHOLD: &str = "RUSTFS_OBJECT_DIRECT_IO_READ_THRESHOLD";
#[allow(
dead_code,
reason = "platform-conditional: production callers are inside #[cfg(target_os = \"linux\")] blocks, so this reads as dead on non-Linux hosts (backlog#1823)"
)]
const DEFAULT_RUSTFS_OBJECT_DIRECT_IO_READ_THRESHOLD: usize = 4 * 1024 * 1024; const DEFAULT_RUSTFS_OBJECT_DIRECT_IO_READ_THRESHOLD: usize = 4 * 1024 * 1024;
/// Enable O_DIRECT for erasure shard / multipart part data writes (Linux only). /// Enable O_DIRECT for erasure shard / multipart part data writes (Linux only).
@@ -1036,7 +1053,15 @@ const DEFAULT_RUSTFS_OBJECT_DIRECT_IO_READ_THRESHOLD: usize = 4 * 1024 * 1024;
/// EINVAL/EOPNOTSUPP (tmpfs, overlayfs, 9p, ...) latch the path off and fall /// EINVAL/EOPNOTSUPP (tmpfs, overlayfs, 9p, ...) latch the path off and fall
/// back to buffered writes for the whole disk. Non-Linux always falls back. /// back to buffered writes for the whole disk. Non-Linux always falls back.
/// Default: false (buffered writes via the page cache, as before). /// Default: false (buffered writes via the page cache, as before).
#[allow(
dead_code,
reason = "platform-conditional: production callers are inside #[cfg(target_os = \"linux\")] blocks, so this reads as dead on non-Linux hosts (backlog#1823)"
)]
const ENV_RUSTFS_OBJECT_DIRECT_IO_WRITE_ENABLE: &str = "RUSTFS_OBJECT_DIRECT_IO_WRITE_ENABLE"; const ENV_RUSTFS_OBJECT_DIRECT_IO_WRITE_ENABLE: &str = "RUSTFS_OBJECT_DIRECT_IO_WRITE_ENABLE";
#[allow(
dead_code,
reason = "platform-conditional: production callers are inside #[cfg(target_os = \"linux\")] blocks, so this reads as dead on non-Linux hosts (backlog#1823)"
)]
const DEFAULT_RUSTFS_OBJECT_DIRECT_IO_WRITE_ENABLE: bool = false; const DEFAULT_RUSTFS_OBJECT_DIRECT_IO_WRITE_ENABLE: bool = false;
const ENV_RUSTFS_OBJECT_MMAP_POPULATE_ENABLE: &str = "RUSTFS_OBJECT_MMAP_POPULATE_ENABLE"; const ENV_RUSTFS_OBJECT_MMAP_POPULATE_ENABLE: &str = "RUSTFS_OBJECT_MMAP_POPULATE_ENABLE";
const DEFAULT_RUSTFS_OBJECT_MMAP_POPULATE_ENABLE: bool = false; const DEFAULT_RUSTFS_OBJECT_MMAP_POPULATE_ENABLE: bool = false;
@@ -1095,12 +1120,14 @@ macro_rules! cached_read_env {
cached_read_env! { cached_read_env! {
/// Check if O_DIRECT reads are enabled. /// Check if O_DIRECT reads are enabled.
#[allow(dead_code, reason = "platform-conditional: production callers are inside #[cfg(target_os = \"linux\")] blocks, so this reads as dead on non-Linux hosts (backlog#1823)")]
fn is_direct_io_read_enabled() -> bool = fn is_direct_io_read_enabled() -> bool =
rustfs_utils::get_env_bool(ENV_RUSTFS_OBJECT_DIRECT_IO_READ_ENABLE, DEFAULT_RUSTFS_OBJECT_DIRECT_IO_READ_ENABLE); rustfs_utils::get_env_bool(ENV_RUSTFS_OBJECT_DIRECT_IO_READ_ENABLE, DEFAULT_RUSTFS_OBJECT_DIRECT_IO_READ_ENABLE);
} }
cached_read_env! { cached_read_env! {
/// Check if O_DIRECT shard/part data writes are enabled. /// Check if O_DIRECT shard/part data writes are enabled.
#[allow(dead_code, reason = "platform-conditional: production callers are inside #[cfg(target_os = \"linux\")] blocks, so this reads as dead on non-Linux hosts (backlog#1823)")]
fn is_direct_io_write_enabled() -> bool = fn is_direct_io_write_enabled() -> bool =
rustfs_utils::get_env_bool(ENV_RUSTFS_OBJECT_DIRECT_IO_WRITE_ENABLE, DEFAULT_RUSTFS_OBJECT_DIRECT_IO_WRITE_ENABLE); rustfs_utils::get_env_bool(ENV_RUSTFS_OBJECT_DIRECT_IO_WRITE_ENABLE, DEFAULT_RUSTFS_OBJECT_DIRECT_IO_WRITE_ENABLE);
} }
@@ -1456,6 +1483,7 @@ pub(crate) fn effective_durability(volume: &str) -> DurabilityMode {
cached_read_env! { cached_read_env! {
/// Get the O_DIRECT read threshold size. /// Get the O_DIRECT read threshold size.
#[allow(dead_code, reason = "platform-conditional: production callers are inside #[cfg(target_os = \"linux\")] blocks, so this reads as dead on non-Linux hosts (backlog#1823)")]
fn get_direct_io_read_threshold() -> usize = fn get_direct_io_read_threshold() -> usize =
rustfs_utils::get_env_usize(ENV_RUSTFS_OBJECT_DIRECT_IO_READ_THRESHOLD, DEFAULT_RUSTFS_OBJECT_DIRECT_IO_READ_THRESHOLD); rustfs_utils::get_env_usize(ENV_RUSTFS_OBJECT_DIRECT_IO_READ_THRESHOLD, DEFAULT_RUSTFS_OBJECT_DIRECT_IO_READ_THRESHOLD);
} }
@@ -1673,12 +1701,20 @@ impl DirectIoWriteState {
/// Target staging size for O_DIRECT writes, rounded up to the DIO alignment. /// Target staging size for O_DIRECT writes, rounded up to the DIO alignment.
/// Bounds the per-writer aligned bounce buffer and batches many shard blocks /// Bounds the per-writer aligned bounce buffer and batches many shard blocks
/// into one positioned write to keep the syscall count low. /// into one positioned write to keep the syscall count low.
#[allow(
dead_code,
reason = "platform-conditional: production callers are inside #[cfg(target_os = \"linux\")] blocks, so this reads as dead on non-Linux hosts (backlog#1823)"
)]
const DIRECT_WRITE_STAGING_BYTES: usize = 1024 * 1024; const DIRECT_WRITE_STAGING_BYTES: usize = 1024 * 1024;
/// Aligned bounce-buffer capacity for a given DIO alignment: the target staging /// Aligned bounce-buffer capacity for a given DIO alignment: the target staging
/// size rounded up to a whole multiple of `align` so the buffer address, every /// size rounded up to a whole multiple of `align` so the buffer address, every
/// flushed batch length, and every write offset stay alignment-correct. /// flushed batch length, and every write offset stay alignment-correct.
/// Platform-independent (no O_DIRECT), so it is unit-tested on any host. /// Platform-independent (no O_DIRECT), so it is unit-tested on any host.
#[allow(
dead_code,
reason = "platform-conditional: production callers are inside #[cfg(target_os = \"linux\")] blocks, so this reads as dead on non-Linux hosts (backlog#1823)"
)]
fn direct_write_staging_capacity(align: usize) -> usize { fn direct_write_staging_capacity(align: usize) -> usize {
debug_assert!(align.is_power_of_two() && align >= 512); debug_assert!(align.is_power_of_two() && align >= 512);
DIRECT_WRITE_STAGING_BYTES.div_ceil(align) * align DIRECT_WRITE_STAGING_BYTES.div_ceil(align) * align
@@ -1687,6 +1723,10 @@ fn direct_write_staging_capacity(align: usize) -> usize {
/// Split `filled` staged bytes into the alignment-sized prefix written with /// Split `filled` staged bytes into the alignment-sized prefix written with
/// O_DIRECT and the sub-alignment tail written buffered. Platform-independent, /// O_DIRECT and the sub-alignment tail written buffered. Platform-independent,
/// so the tail-boundary math is unit-tested on any host. /// so the tail-boundary math is unit-tested on any host.
#[allow(
dead_code,
reason = "platform-conditional: production callers are inside #[cfg(target_os = \"linux\")] blocks, so this reads as dead on non-Linux hosts (backlog#1823)"
)]
fn direct_write_tail_split(filled: usize, align: usize) -> (usize, usize) { fn direct_write_tail_split(filled: usize, align: usize) -> (usize, usize) {
let aligned = filled - (filled % align); let aligned = filled - (filled % align);
(aligned, filled - aligned) (aligned, filled - aligned)
@@ -2142,6 +2182,7 @@ fn set_delete_version_fail_after_data_staged(path: &str) {
} }
#[cfg(test)] #[cfg(test)]
#[allow(dead_code, reason = "asserted by this file's tests (backlog#1823)")]
pub(crate) fn set_delete_version_fail_after_commit(root: &Path, path: &str) { pub(crate) fn set_delete_version_fail_after_commit(root: &Path, path: &str) {
DELETE_VERSION_FAIL_AFTER_COMMIT DELETE_VERSION_FAIL_AFTER_COMMIT
.lock() .lock()
@@ -2447,6 +2488,10 @@ enum SyncMode {
FileOnly, FileOnly,
} }
#[allow(
dead_code,
reason = "reclaim bookkeeping fields written by Drop but never read back (backlog#1823)"
)]
struct FileCacheReclaimWriter { struct FileCacheReclaimWriter {
inner: File, inner: File,
reclaim_len: usize, reclaim_len: usize,
@@ -2454,6 +2499,10 @@ struct FileCacheReclaimWriter {
reclaimed: bool, reclaimed: bool,
} }
#[allow(
dead_code,
reason = "reclaim bookkeeping fields written by Drop but never read back (backlog#1823)"
)]
struct FileCacheReclaimReader { struct FileCacheReclaimReader {
inner: File, inner: File,
reclaim_offset: u64, reclaim_offset: u64,
@@ -2519,6 +2568,10 @@ impl<R: AsyncRead + Unpin> AsyncRead for StallTimeoutReader<R> {
} }
} }
#[allow(
dead_code,
reason = "reclaim metrics emitter reached only from the Linux-gated reclaim paths (backlog#1823)"
)]
fn record_file_cache_reclaim_success(kind: &'static str, reclaim_len: usize, started: std::time::Instant) { fn record_file_cache_reclaim_success(kind: &'static str, reclaim_len: usize, started: std::time::Instant) {
// Runs per read-stream page-cache reclaim window; skip the whole emission // Runs per read-stream page-cache reclaim window; skip the whole emission
// (three metric-key constructions) when general metrics are disabled. // (three metric-key constructions) when general metrics are disabled.
@@ -3071,6 +3124,7 @@ impl LocalIoBackend for StdBackend {
use memmap2::MmapOptions; use memmap2::MmapOptions;
use std::time::{Duration as StdDuration, Instant as StdInstant}; use std::time::{Duration as StdDuration, Instant as StdInstant};
#[allow(dead_code, reason = "mmap copy result slot kept beside the mapping it owns (backlog#1823)")]
struct MmapCopyReadResult { struct MmapCopyReadResult {
bytes: Bytes, bytes: Bytes,
access_check_duration: StdDuration, access_check_duration: StdDuration,
@@ -4704,6 +4758,10 @@ fn build_local_io_backend(root: PathBuf) -> Arc<dyn LocalIoBackend> {
Arc::new(StdBackend::new(root)) Arc::new(StdBackend::new(root))
} }
#[allow(
dead_code,
reason = "path cache and cwd slots retained beside the disk root they derive from (backlog#1823)"
)]
pub struct LocalDisk { pub struct LocalDisk {
pub root: PathBuf, pub root: PathBuf,
publication_root: os::PublicationRoot, publication_root: os::PublicationRoot,
@@ -5490,6 +5548,7 @@ impl LocalDisk {
Ok(Self::resolve_abs_path_from(&self.root, path.as_ref())) Ok(Self::resolve_abs_path_from(&self.root, path.as_ref()))
} }
#[allow(dead_code, reason = "asserted by this file's tests (backlog#1823)")]
fn io_resolve_abs_path(&self, path: impl AsRef<Path>) -> PathBuf { fn io_resolve_abs_path(&self, path: impl AsRef<Path>) -> PathBuf {
let path_ref = path.as_ref(); let path_ref = path.as_ref();
let path_str = path_ref.to_string_lossy(); let path_str = path_ref.to_string_lossy();
@@ -5567,15 +5626,24 @@ impl LocalDisk {
} }
// Check if a path is valid // Check if a path is valid
#[allow(
dead_code,
reason = "method wrapper over the live free function check_local_disk_valid_path; no caller in this port (backlog#1823)"
)]
fn check_valid_path<P: AsRef<Path>>(&self, path: P) -> Result<()> { fn check_valid_path<P: AsRef<Path>>(&self, path: P) -> Result<()> {
check_local_disk_valid_path(self.io_root(), path) check_local_disk_valid_path(self.io_root(), path)
} }
#[allow(
dead_code,
reason = "method wrapper over the live free function reject_local_disk_symlink_components; no caller in this port (backlog#1823)"
)]
fn reject_symlink_components(&self, path: &Path) -> Result<()> { fn reject_symlink_components(&self, path: &Path) -> Result<()> {
reject_local_disk_symlink_components(self.io_root(), path) reject_local_disk_symlink_components(self.io_root(), path)
} }
// Batch path generation with single lock acquisition // Batch path generation with single lock acquisition
#[allow(dead_code, reason = "asserted by this file's tests (backlog#1823)")]
fn get_object_paths_batch(&self, requests: &[(String, String)]) -> Result<Vec<PathBuf>> { fn get_object_paths_batch(&self, requests: &[(String, String)]) -> Result<Vec<PathBuf>> {
let mut results = Vec::with_capacity(requests.len()); let mut results = Vec::with_capacity(requests.len());
let mut cache_misses = Vec::new(); let mut cache_misses = Vec::new();
@@ -6488,6 +6556,7 @@ impl LocalDisk {
Ok(f) Ok(f)
} }
#[allow(dead_code, reason = "asserted by this file's tests (backlog#1823)")]
async fn open_file_read_only(&self, path: impl AsRef<Path>) -> Result<File> { async fn open_file_read_only(&self, path: impl AsRef<Path>) -> Result<File> {
let f = super::fs::open_file(path.as_ref(), O_RDONLY).await.map_err(to_file_error)?; let f = super::fs::open_file(path.as_ref(), O_RDONLY).await.map_err(to_file_error)?;
Ok(f) Ok(f)
+5 -1
View File
@@ -13,7 +13,6 @@
// limitations under the License. // limitations under the License.
// #730: disk abstractions still carry staged health and direct-I/O migration paths. // #730: disk abstractions still carry staged health and direct-I/O migration paths.
#![allow(dead_code)]
pub mod disk_store; pub mod disk_store;
pub mod endpoint; pub mod endpoint;
@@ -1114,6 +1113,10 @@ pub struct DiskInfo {
} }
#[derive(Clone, Debug, Default)] #[derive(Clone, Debug, Default)]
#[allow(
dead_code,
reason = "MinIO-parity disk info shape with no constructor in this port (backlog#1823)"
)]
pub struct Info { pub struct Info {
pub total: u64, pub total: u64,
pub free: u64, pub free: u64,
@@ -1372,6 +1375,7 @@ pub fn conv_part_err_to_int(err: &Option<Error>) -> usize {
} }
} }
#[allow(dead_code, reason = "asserted by this file's tests (backlog#1823)")]
pub fn has_part_err(part_errs: &[usize]) -> bool { pub fn has_part_err(part_errs: &[usize]) -> bool {
part_errs.iter().any(|err| *err != CHECK_PART_SUCCESS) part_errs.iter().any(|err| *err != CHECK_PART_SUCCESS)
} }
+5 -11
View File
@@ -571,6 +571,10 @@ fn regular_files(dir: &Path) -> io::Result<Vec<PathBuf>> {
/// Fdatasync every regular file directly inside `dir`, then fsync the directory /// Fdatasync every regular file directly inside `dir`, then fsync the directory
/// itself. /// itself.
#[allow(
dead_code,
reason = "reached only through sync_dir_files, whose callers are tests (backlog#1823)"
)]
pub fn sync_dir_files_std(dir: impl AsRef<Path>) -> io::Result<()> { pub fn sync_dir_files_std(dir: impl AsRef<Path>) -> io::Result<()> {
for entry in std::fs::read_dir(dir.as_ref())? { for entry in std::fs::read_dir(dir.as_ref())? {
let entry = entry?; let entry = entry?;
@@ -583,6 +587,7 @@ pub fn sync_dir_files_std(dir: impl AsRef<Path>) -> io::Result<()> {
/// Async wrapper around [`sync_dir_files_std`]. Large directories flush files /// Async wrapper around [`sync_dir_files_std`]. Large directories flush files
/// concurrently, bounded both per directory and process-wide. /// concurrently, bounded both per directory and process-wide.
#[allow(dead_code, reason = "asserted by this file's tests (backlog#1823)")]
pub async fn sync_dir_files(dir: impl AsRef<Path>) -> io::Result<()> { pub async fn sync_dir_files(dir: impl AsRef<Path>) -> io::Result<()> {
sync_dir_files_with_limiter(dir, Arc::new(Semaphore::new(MAX_PARALLEL_FILE_SYNCS))).await sync_dir_files_with_limiter(dir, Arc::new(Semaphore::new(MAX_PARALLEL_FILE_SYNCS))).await
} }
@@ -1809,10 +1814,6 @@ impl RenameCommitGuard {
}) })
} }
pub(crate) fn lock_destination_directory_for_path_access(&self, directory: &Path) -> io::Result<RenameDestinationPathGuard> {
self.destination_directory_guard(directory, false)
}
pub(crate) fn create_destination_directory_for_path_access( pub(crate) fn create_destination_directory_for_path_access(
&self, &self,
directory: &Path, directory: &Path,
@@ -2858,13 +2859,6 @@ pub async fn os_mkdir_all(dir_path: impl AsRef<Path>, base_dir: impl AsRef<Path>
Ok(()) Ok(())
} }
/// Check if a file exists.
/// Returns true if the file exists, false otherwise.
#[tracing::instrument(level = "debug", skip_all)]
pub fn file_exists(path: impl AsRef<Path>) -> bool {
std::fs::metadata(path.as_ref()).map(|_| true).unwrap_or(false)
}
/// Whether an [`io::Error`] means "the directory is not empty". /// Whether an [`io::Error`] means "the directory is not empty".
/// ///
/// POSIX lets `rmdir`/`rename` report a non-empty directory as either /// POSIX lets `rmdir`/`rename` report a non-empty directory as either
+1
View File
@@ -704,6 +704,7 @@ pub(crate) async fn create_bitrot_reader_from_bytes_with_stage_metrics(
} }
#[allow(clippy::too_many_arguments)] #[allow(clippy::too_many_arguments)]
#[allow(dead_code, reason = "asserted by this file's tests (backlog#1823)")]
pub fn create_deferred_bitrot_reader( pub fn create_deferred_bitrot_reader(
inline_data: Option<Bytes>, inline_data: Option<Bytes>,
disk: Option<DiskStore>, disk: Option<DiskStore>,
+6
View File
@@ -277,6 +277,12 @@ pub struct ObjectOptions {
/// fence avoids recursively acquiring the read lock behind a queued writer. /// fence avoids recursively acquiring the read lock behind a queued writer.
pub bucket_lifecycle_lock_fence: Option<NamespaceLockFence>, pub bucket_lifecycle_lock_fence: Option<NamespaceLockFence>,
pub replication_request: bool, pub replication_request: bool,
/// Source-cluster LWW timestamps carried by an authorized replication
/// request; None when the source never modified the category. Only the
/// replication-authorized options builders may set these.
pub replication_tagging_timestamp: Option<OffsetDateTime>,
pub replication_retention_timestamp: Option<OffsetDateTime>,
pub replication_legalhold_timestamp: Option<OffsetDateTime>,
/// Authorized SSE-C replication passthrough: the body is already /// Authorized SSE-C replication passthrough: the body is already
/// ciphertext, so the write path must not encrypt or compress it and /// ciphertext, so the write path must not encrypt or compress it and
/// stores the restored encryption metadata verbatim. Only the /// stores the restored encryption metadata verbatim. Only the
+629 -68
View File
@@ -32,15 +32,22 @@ use crate::diagnostics::get::{
GET_METADATA_CACHE_REASON_NOT_READ_DATA, GET_METADATA_CACHE_REASON_PART_NUMBER, GET_METADATA_CACHE_REASON_NOT_READ_DATA, GET_METADATA_CACHE_REASON_PART_NUMBER,
GET_METADATA_CACHE_REASON_RAW_DATA_MOVEMENT_READ, GET_METADATA_CACHE_REASON_USABLE, GET_METADATA_CACHE_REASON_VERSION_ID, GET_METADATA_CACHE_REASON_RAW_DATA_MOVEMENT_READ, GET_METADATA_CACHE_REASON_USABLE, GET_METADATA_CACHE_REASON_VERSION_ID,
GET_METADATA_CACHE_REASON_VERSION_SUSPENDED, GET_METADATA_CACHE_REASON_VERSIONED, GET_METADATA_CACHE_REASON_VERSION_SUSPENDED, GET_METADATA_CACHE_REASON_VERSIONED,
GET_METADATA_EARLY_STOP_REASON_CONFLICTING_METADATA, GET_METADATA_EARLY_STOP_REASON_DELETE_MARKER, GET_METADATA_EARLY_STOP_REASON_CONFLICTING_METADATA, GET_METADATA_EARLY_STOP_REASON_DATA_READ_INLINE_BODY_VERIFY,
GET_METADATA_EARLY_STOP_REASON_ERROR, GET_METADATA_EARLY_STOP_REASON_INSUFFICIENT_QUORUM, GET_METADATA_EARLY_STOP_REASON_DATA_READ_INLINE_DELETED, GET_METADATA_EARLY_STOP_REASON_DATA_READ_INLINE_GEOMETRY,
GET_METADATA_EARLY_STOP_REASON_NOT_FOUND, GET_METADATA_EARLY_STOP_REASON_UNSAFE_REQUEST, GET_METADATA_EARLY_STOP_REASON_DATA_READ_INLINE_IDENTITY_MISMATCH,
GET_METADATA_EARLY_STOP_REASON_VALID_QUORUM, GET_METADATA_EARLY_STOP_REASON_VERSION_MATCH_QUORUM, GET_METADATA_EARLY_STOP_REASON_DATA_READ_INLINE_MISSING_PAYLOAD,
GET_METADATA_EARLY_STOP_REASON_VERSION_NOT_FOUND, GET_METADATA_RESPONSE_CORRUPT, GET_METADATA_RESPONSE_DISK_NOT_FOUND, GET_METADATA_EARLY_STOP_REASON_DATA_READ_INLINE_MISSING_SHARD, GET_METADATA_EARLY_STOP_REASON_DATA_READ_INLINE_NOT_INLINE,
GET_METADATA_RESPONSE_ERROR, GET_METADATA_RESPONSE_IGNORED, GET_METADATA_RESPONSE_NOT_FOUND, GET_METADATA_RESPONSE_TIMEOUT, GET_METADATA_EARLY_STOP_REASON_DATA_READ_INLINE_PART_SHAPE, GET_METADATA_EARLY_STOP_REASON_DATA_READ_INLINE_REMOTE,
GET_METADATA_RESPONSE_VALID, GET_METADATA_RESPONSE_VERSION_NOT_FOUND, GET_OBJECT_PATH_CODEC_STREAMING, GET_METADATA_EARLY_STOP_REASON_DATA_READ_INLINE_SIZE, GET_METADATA_EARLY_STOP_REASON_DATA_READ_INLINE_TRANSFORMED,
GET_OBJECT_PATH_DIRECT_MEMORY, GET_OBJECT_PATH_INTERNAL_META, GET_OBJECT_PATH_LEGACY_DUPLEX, GET_OBJECT_PATH_SET_DISK, GET_METADATA_EARLY_STOP_REASON_DELETE_MARKER, GET_METADATA_EARLY_STOP_REASON_ERROR,
GET_STAGE_DECODE, GET_STAGE_METADATA_CACHE_LOOKUP, GET_STAGE_METADATA_RESOLVE, GET_STAGE_RANGE, GET_STAGE_READER_SETUP, GET_METADATA_EARLY_STOP_REASON_INSUFFICIENT_QUORUM, GET_METADATA_EARLY_STOP_REASON_NOT_FOUND,
GET_METADATA_EARLY_STOP_REASON_UNSAFE_REQUEST, GET_METADATA_EARLY_STOP_REASON_VALID_QUORUM,
GET_METADATA_EARLY_STOP_REASON_VERSION_MATCH_QUORUM, GET_METADATA_EARLY_STOP_REASON_VERSION_NOT_FOUND,
GET_METADATA_RESPONSE_CORRUPT, GET_METADATA_RESPONSE_DISK_NOT_FOUND, GET_METADATA_RESPONSE_ERROR,
GET_METADATA_RESPONSE_IGNORED, GET_METADATA_RESPONSE_NOT_FOUND, GET_METADATA_RESPONSE_TIMEOUT, GET_METADATA_RESPONSE_VALID,
GET_METADATA_RESPONSE_VERSION_NOT_FOUND, GET_OBJECT_PATH_CODEC_STREAMING, GET_OBJECT_PATH_DIRECT_MEMORY,
GET_OBJECT_PATH_INTERNAL_META, GET_OBJECT_PATH_LEGACY_DUPLEX, GET_OBJECT_PATH_SET_DISK, GET_STAGE_DECODE,
GET_STAGE_METADATA_CACHE_LOOKUP, GET_STAGE_METADATA_RESOLVE, GET_STAGE_RANGE, GET_STAGE_READER_SETUP,
GET_STAGE_READER_SETUP_DROP_PENDING, GET_STAGE_READER_SETUP_SCHEDULE, GET_STAGE_READER_SETUP_WAIT_QUORUM, GET_STAGE_READER_SETUP_DROP_PENDING, GET_STAGE_READER_SETUP_SCHEDULE, GET_STAGE_READER_SETUP_WAIT_QUORUM,
GET_STAGE_READER_TASK_BITROT_READER_INIT, GET_STAGE_READER_TASK_FILE_OPEN, GET_STAGE_READER_TASK_READER_CONSTRUCTION, GET_STAGE_READER_TASK_BITROT_READER_INIT, GET_STAGE_READER_TASK_FILE_OPEN, GET_STAGE_READER_TASK_READER_CONSTRUCTION,
GetObjectFailureReason, classify_disk_error, get_stage_timer_if_enabled, record_get_object_pipeline_failure, GetObjectFailureReason, classify_disk_error, get_stage_timer_if_enabled, record_get_object_pipeline_failure,
@@ -173,11 +180,13 @@ pub(in crate::set_disk) enum GetCodecStreamingReaderBuildOutcome {
Fallback(GetCodecStreamingFallbackReason), Fallback(GetCodecStreamingFallbackReason),
} }
#[allow(dead_code, reason = "asserted by this file's tests (backlog#1823)")]
pub(in crate::set_disk) struct MultipartCodecStreamingReader { pub(in crate::set_disk) struct MultipartCodecStreamingReader {
pub(in crate::set_disk) readers: VecDeque<Box<dyn AsyncRead + Unpin + Send + Sync>>, pub(in crate::set_disk) readers: VecDeque<Box<dyn AsyncRead + Unpin + Send + Sync>>,
} }
impl MultipartCodecStreamingReader { impl MultipartCodecStreamingReader {
#[allow(dead_code, reason = "asserted by this file's tests (backlog#1823)")]
pub(in crate::set_disk) fn new(readers: Vec<Box<dyn AsyncRead + Unpin + Send + Sync>>) -> Self { pub(in crate::set_disk) fn new(readers: Vec<Box<dyn AsyncRead + Unpin + Send + Sync>>) -> Self {
Self { Self {
readers: VecDeque::from(readers), readers: VecDeque::from(readers),
@@ -652,36 +661,15 @@ pub(in crate::set_disk) fn metadata_early_stop_candidate_matches(left: &FileInfo
&& left.erasure.distribution == right.erasure.distribution && left.erasure.distribution == right.erasure.distribution
} }
pub(in crate::set_disk) async fn data_read_early_stop_inline_body_verified( pub(in crate::set_disk) async fn data_read_early_stop_inline_body_miss_reason(
bucket: &str, bucket: &str,
object: &str, object: &str,
candidate: &FileInfo, candidate: &FileInfo,
parts_metadata: &[FileInfo], parts_metadata: &[FileInfo],
disks: &[Option<DiskStore>], disks: &[Option<DiskStore>],
) -> bool { ) -> Option<&'static str> {
if !candidate.inline_data() if let Some(reason) = data_read_early_stop_inline_candidate_miss_reason(candidate) {
|| candidate.is_compressed() return Some(reason);
|| candidate
.metadata
.keys()
.any(|key| rustfs_utils::http::is_object_encryption_marker(key))
|| candidate.is_remote()
|| candidate.deleted
|| candidate.size <= 0
|| candidate.parts.len() != 1
|| !candidate.has_valid_erasure_geometry()
{
return false;
}
let Ok(object_size) = usize::try_from(candidate.size) else {
return false;
};
if candidate.parts.first().is_none_or(|part| part.size != object_size) {
return false;
}
if !can_try_inline_data_shards_direct(object_size, candidate.erasure.block_size) {
return false;
} }
let Ok(erasure) = coding::Erasure::try_new_with_options( let Ok(erasure) = coding::Erasure::try_new_with_options(
@@ -690,18 +678,21 @@ pub(in crate::set_disk) async fn data_read_early_stop_inline_body_verified(
candidate.erasure.block_size, candidate.erasure.block_size,
candidate.uses_legacy_checksum, candidate.uses_legacy_checksum,
) else { ) else {
return false; return Some(GET_METADATA_EARLY_STOP_REASON_DATA_READ_INLINE_GEOMETRY);
}; };
let Some(data_files) = let data_files =
collect_inline_data_shard_fileinfos_by_index(parts_metadata, candidate, erasure.data_shards, |index| { match collect_inline_data_shard_fileinfos_by_index_or_reason(parts_metadata, candidate, erasure.data_shards, |index| {
disks.get(index).is_some_and(Option::is_some) disks.get(index).is_some_and(Option::is_some)
}) }) {
else { Ok(data_files) => data_files,
return false; Err(reason) => return Some(reason),
}; };
let Some(part) = candidate.parts.first() else { let Some(part) = candidate.parts.first() else {
return false; return Some(GET_METADATA_EARLY_STOP_REASON_DATA_READ_INLINE_PART_SHAPE);
};
let Ok(object_size) = usize::try_from(candidate.size) else {
return Some(GET_METADATA_EARLY_STOP_REASON_DATA_READ_INLINE_SIZE);
}; };
let checksum_info = candidate.erasure.get_checksum_info(part.number); let checksum_info = candidate.erasure.get_checksum_info(part.number);
let checksum_algo = if candidate.uses_legacy_checksum && checksum_info.algorithm == HashAlgorithm::HighwayHash256S { let checksum_algo = if candidate.uses_legacy_checksum && checksum_info.algorithm == HashAlgorithm::HighwayHash256S {
@@ -721,12 +712,111 @@ pub(in crate::set_disk) async fn data_read_early_stop_inline_body_verified(
let Ok(mut readers) = let Ok(mut readers) =
build_inline_bitrot_readers_from_refs(&data_files, bucket, object, read_length, shard_size, &checksum_algo, false).await build_inline_bitrot_readers_from_refs(&data_files, bucket, object, read_length, shard_size, &checksum_algo, false).await
else { else {
return false; return Some(GET_METADATA_EARLY_STOP_REASON_DATA_READ_INLINE_BODY_VERIFY);
}; };
try_read_inline_data_shards_direct(&mut readers, erasure.data_shards, read_length, object_size) match try_read_inline_data_shards_direct(&mut readers, erasure.data_shards, read_length, object_size).await {
.await Some(body) if body.len() == object_size => None,
.is_some_and(|body| body.len() == object_size) _ => Some(GET_METADATA_EARLY_STOP_REASON_DATA_READ_INLINE_BODY_VERIFY),
}
}
fn data_read_early_stop_inline_candidate_miss_reason(candidate: &FileInfo) -> Option<&'static str> {
// `inline_data` excludes remote objects; this diagnostic reports them separately.
if !rustfs_utils::http::contains_key_str(&candidate.metadata, rustfs_utils::http::SUFFIX_INLINE_DATA) {
return Some(GET_METADATA_EARLY_STOP_REASON_DATA_READ_INLINE_NOT_INLINE);
}
if candidate.is_compressed()
|| candidate
.metadata
.keys()
.any(|key| rustfs_utils::http::is_object_encryption_marker(key))
{
return Some(GET_METADATA_EARLY_STOP_REASON_DATA_READ_INLINE_TRANSFORMED);
}
if candidate.is_remote() {
return Some(GET_METADATA_EARLY_STOP_REASON_DATA_READ_INLINE_REMOTE);
}
if candidate.deleted {
return Some(GET_METADATA_EARLY_STOP_REASON_DATA_READ_INLINE_DELETED);
}
if candidate.size <= 0 {
return Some(GET_METADATA_EARLY_STOP_REASON_DATA_READ_INLINE_SIZE);
}
if candidate.parts.len() != 1 {
return Some(GET_METADATA_EARLY_STOP_REASON_DATA_READ_INLINE_PART_SHAPE);
}
if !candidate.has_valid_erasure_geometry() {
return Some(GET_METADATA_EARLY_STOP_REASON_DATA_READ_INLINE_GEOMETRY);
}
let Ok(object_size) = usize::try_from(candidate.size) else {
return Some(GET_METADATA_EARLY_STOP_REASON_DATA_READ_INLINE_SIZE);
};
if candidate.parts.first().is_none_or(|part| part.size != object_size) {
return Some(GET_METADATA_EARLY_STOP_REASON_DATA_READ_INLINE_PART_SHAPE);
}
if !can_try_inline_data_shards_direct(object_size, candidate.erasure.block_size) {
return Some(GET_METADATA_EARLY_STOP_REASON_DATA_READ_INLINE_SIZE);
}
None
}
fn data_read_inline_missing_shards_are_pending(
candidate: &FileInfo,
parts_metadata: &[FileInfo],
errors: &[Option<DiskError>],
disks: &[Option<DiskStore>],
fanout_order: &[usize],
scheduled_fanout_len: usize,
) -> bool {
let Ok(erasure) = coding::Erasure::try_new_with_options(
candidate.erasure.data_blocks,
candidate.erasure.parity_blocks,
candidate.erasure.block_size,
candidate.uses_legacy_checksum,
) else {
return false;
};
let distribution = &candidate.erasure.distribution;
let mut data_shards_seen_or_pending = vec![false; erasure.data_shards];
let mut missing_pending_data_shards = 0usize;
for (disk_index, file_info) in parts_metadata.iter().enumerate() {
let Some(&block_index) = distribution.get(disk_index) else {
return false;
};
if block_index == 0 || block_index > erasure.data_shards {
continue;
}
if !disks.get(disk_index).is_some_and(Option::is_some) {
return false;
}
let data_slot = block_index - 1;
if file_info.name.is_empty() {
let scheduled_and_not_failed = fanout_order
.get(..scheduled_fanout_len)
.is_some_and(|scheduled_disks| scheduled_disks.contains(&disk_index))
&& errors.get(disk_index).is_some_and(Option::is_none);
if scheduled_and_not_failed {
data_shards_seen_or_pending[data_slot] = true;
missing_pending_data_shards = missing_pending_data_shards.saturating_add(1);
continue;
}
return false;
}
if file_info.erasure.index != block_index
|| !file_info.has_valid_erasure_geometry()
|| !metadata_early_stop_candidate_matches(file_info, candidate)
|| file_info.data.as_ref().is_none_or(|data| data.is_empty())
{
return false;
}
data_shards_seen_or_pending[data_slot] = true;
}
missing_pending_data_shards > 0 && data_shards_seen_or_pending.into_iter().all(|seen_or_pending| seen_or_pending)
} }
pub(in crate::set_disk) fn classify_metadata_response_error(err: &DiskError) -> &'static str { pub(in crate::set_disk) fn classify_metadata_response_error(err: &DiskError) -> &'static str {
@@ -1758,6 +1848,7 @@ pub(in crate::set_disk) async fn create_bitrot_readers_until_quorum_all_shards(
} }
#[allow(clippy::too_many_arguments)] #[allow(clippy::too_many_arguments)]
#[allow(dead_code, reason = "asserted by this file's tests (backlog#1823)")]
pub(in crate::set_disk) async fn create_bitrot_readers_until_quorum( pub(in crate::set_disk) async fn create_bitrot_readers_until_quorum(
files: &[FileInfo], files: &[FileInfo],
disks: &[Option<DiskStore>], disks: &[Option<DiskStore>],
@@ -2048,6 +2139,7 @@ pub(in crate::set_disk) async fn create_data_block_bitrot_readers(
setup setup
} }
#[allow(dead_code, reason = "asserted by this file's tests (backlog#1823)")]
pub(in crate::set_disk) async fn collect_read_multiple_results<F>( pub(in crate::set_disk) async fn collect_read_multiple_results<F>(
tasks: Vec<F>, tasks: Vec<F>,
read_quorum: usize, read_quorum: usize,
@@ -2364,6 +2456,7 @@ impl SetDisks {
let bucket: Arc<str> = Arc::from(bucket); let bucket: Arc<str> = Arc::from(bucket);
let object: Arc<str> = Arc::from(object); let object: Arc<str> = Arc::from(object);
let version_id: Arc<str> = Arc::from(version_id); let version_id: Arc<str> = Arc::from(version_id);
let slowtail_fault = get_metadata_slowtail_fault_request(bucket.as_ref(), object.as_ref(), read_data);
let futures = disks.iter().enumerate().map(|(disk_index, disk)| { let futures = disks.iter().enumerate().map(|(disk_index, disk)| {
let disk = disk.clone(); let disk = disk.clone();
let task_opts = opts; let task_opts = opts;
@@ -2371,10 +2464,14 @@ impl SetDisks {
let bucket = bucket.clone(); let bucket = bucket.clone();
let object = object.clone(); let object = object.clone();
let version_id = version_id.clone(); let version_id = version_id.clone();
let slowtail_fault = slowtail_fault.clone();
tokio::spawn(async move { tokio::spawn(async move {
let response_start = observe.then(Instant::now); let response_start = observe.then(Instant::now);
let result = if let Some(disk) = disk { let result = if let Some(disk) = disk {
Self::record_read_version_call(&object, disk_index); Self::record_read_version_call(&object, disk_index);
if let Some(delay) = slowtail_fault.as_ref().and_then(|fault| fault.delay_for_disk(disk_index)) {
tokio::time::sleep(delay).await;
}
disk.read_version(&org_bucket, &bucket, &object, &version_id, &task_opts) disk.read_version(&org_bucket, &bucket, &object, &version_id, &task_opts)
.await .await
} else { } else {
@@ -2469,6 +2566,8 @@ impl SetDisks {
let mut next_fanout_index = 0usize; let mut next_fanout_index = 0usize;
let mut scheduled_count = 0usize; let mut scheduled_count = 0usize;
let mut force_full_wait = false; let mut force_full_wait = false;
let mut final_miss_reason_override = None;
let slowtail_fault = get_metadata_slowtail_fault_request(bucket.as_ref(), object.as_ref(), read_data);
let spawn_read_version = let spawn_read_version =
|join_set: &mut JoinSet<(usize, disk::error::Result<FileInfo>, Duration)>, index: usize, disk: Option<DiskStore>| { |join_set: &mut JoinSet<(usize, disk::error::Result<FileInfo>, Duration)>, index: usize, disk: Option<DiskStore>| {
let task_opts = opts; let task_opts = opts;
@@ -2476,6 +2575,7 @@ impl SetDisks {
let bucket = bucket.clone(); let bucket = bucket.clone();
let object = object.clone(); let object = object.clone();
let version_id = version_id.clone(); let version_id = version_id.clone();
let slowtail_fault = slowtail_fault.clone();
join_set.spawn(async move { join_set.spawn(async move {
let response_start = Instant::now(); let response_start = Instant::now();
let result = if let Some(disk) = disk { let result = if let Some(disk) = disk {
@@ -2484,6 +2584,9 @@ impl SetDisks {
Self::record_read_version_call(&object, index); Self::record_read_version_call(&object, index);
#[cfg(test)] #[cfg(test)]
Self::read_version_fanout_barrier(&object, index).await; Self::read_version_fanout_barrier(&object, index).await;
if let Some(delay) = slowtail_fault.as_ref().and_then(|fault| fault.delay_for_disk(index)) {
tokio::time::sleep(delay).await;
}
disk.read_version(&org_bucket, &bucket, &object, &version_id, &task_opts) disk.read_version(&org_bucket, &bucket, &object, &version_id, &task_opts)
.await .await
} else { } else {
@@ -2511,11 +2614,20 @@ impl SetDisks {
} }
while let Some(result) = join_set.join_next().await { while let Some(result) = join_set.join_next().await {
let mut defer_pending_inline_data_shard = false;
match result { match result {
Ok((index, res, elapsed)) => match res { Ok((index, res, elapsed)) => match res {
Ok(file_info) => { Ok(file_info) => {
observations.push(MetadataFanoutObservation::from_file_info(&file_info, elapsed)); observations.push(MetadataFanoutObservation::from_file_info(&file_info, elapsed));
accumulator.observe_file_info(&file_info); accumulator.observe_file_info(&file_info);
if bounded_fanout
&& read_data
&& !force_full_wait
&& let Some(reason) = data_read_early_stop_inline_candidate_miss_reason(&file_info)
{
force_full_wait = true;
final_miss_reason_override.get_or_insert(reason);
}
if let Some(slot) = ress.get_mut(index) { if let Some(slot) = ress.get_mut(index) {
*slot = file_info; *slot = file_info;
} }
@@ -2541,17 +2653,43 @@ impl SetDisks {
.or_else(|| accumulator.version_early_stop_decision()) .or_else(|| accumulator.version_early_stop_decision())
{ {
let should_return_early = if read_data { let should_return_early = if read_data {
let allow_data_read_early_stop = match accumulator.candidate.as_ref() { match accumulator.candidate.as_ref() {
Some(candidate) => { Some(candidate) => match data_read_early_stop_inline_body_miss_reason(
data_read_early_stop_inline_body_verified(bucket.as_ref(), object.as_ref(), candidate, &ress, disks) bucket.as_ref(),
.await object.as_ref(),
candidate,
&ress,
disks,
)
.await
{
None => true,
Some(reason) => {
final_miss_reason_override = Some(reason);
if bounded_fanout
&& reason == GET_METADATA_EARLY_STOP_REASON_DATA_READ_INLINE_MISSING_SHARD
&& data_read_inline_missing_shards_are_pending(
candidate,
&ress,
&errors,
disks,
&fanout_order,
next_fanout_index,
)
{
defer_pending_inline_data_shard = true;
} else {
force_full_wait = true;
}
false
}
},
None => {
force_full_wait = true;
final_miss_reason_override = Some(GET_METADATA_EARLY_STOP_REASON_INSUFFICIENT_QUORUM);
false
} }
None => false,
};
if !allow_data_read_early_stop {
force_full_wait = true;
} }
allow_data_read_early_stop
} else { } else {
true true
}; };
@@ -2588,6 +2726,7 @@ impl SetDisks {
let pending_responses = join_set.len(); let pending_responses = join_set.len();
let should_hedge_single_pending_data_read = read_data let should_hedge_single_pending_data_read = read_data
&& !force_full_wait && !force_full_wait
&& !defer_pending_inline_data_shard
&& pending_responses == 1 && pending_responses == 1
&& accumulator.can_still_reach_early_stop_with_pending(pending_responses); && accumulator.can_still_reach_early_stop_with_pending(pending_responses);
if bounded_fanout && force_full_wait { if bounded_fanout && force_full_wait {
@@ -2600,6 +2739,7 @@ impl SetDisks {
next_fanout_index = next_fanout_index.saturating_add(1); next_fanout_index = next_fanout_index.saturating_add(1);
} }
} else if bounded_fanout } else if bounded_fanout
&& !defer_pending_inline_data_shard
&& next_fanout_index < disks.len() && next_fanout_index < disks.len()
&& (!accumulator.can_still_reach_early_stop_with_pending(pending_responses) && (!accumulator.can_still_reach_early_stop_with_pending(pending_responses)
|| should_hedge_single_pending_data_read) || should_hedge_single_pending_data_read)
@@ -2613,7 +2753,12 @@ impl SetDisks {
} }
} }
rustfs_io_metrics::record_get_object_metadata_early_stop_miss(metrics_path, accumulator.final_miss_reason()); let accumulator_miss_reason = accumulator.final_miss_reason();
let final_miss_reason = match (final_miss_reason_override, accumulator_miss_reason) {
(Some(reason), GET_METADATA_EARLY_STOP_REASON_INSUFFICIENT_QUORUM) => reason,
_ => accumulator_miss_reason,
};
rustfs_io_metrics::record_get_object_metadata_early_stop_miss(metrics_path, final_miss_reason);
rustfs_io_metrics::record_get_object_metadata_early_stop_saved_responses(metrics_path, 0); rustfs_io_metrics::record_get_object_metadata_early_stop_saved_responses(metrics_path, 0);
rustfs_io_metrics::record_get_object_metadata_fanout_lifecycle(metrics_path, scheduled_count, scheduled_count, 0); rustfs_io_metrics::record_get_object_metadata_fanout_lifecycle(metrics_path, scheduled_count, scheduled_count, 0);
let diagnostics = MetadataFanoutDiagnostics::new(fanout_start.elapsed(), observations); let diagnostics = MetadataFanoutDiagnostics::new(fanout_start.elapsed(), observations);
@@ -2842,6 +2987,7 @@ impl SetDisks {
(meta_file_infos, errs) (meta_file_infos, errs)
} }
#[allow(dead_code, reason = "asserted by this file's tests (backlog#1823)")]
pub(in crate::set_disk) async fn read_multiple_files( pub(in crate::set_disk) async fn read_multiple_files(
disks: &[Option<DiskStore>], disks: &[Option<DiskStore>],
req: ReadMultipleReq, req: ReadMultipleReq,
@@ -3021,6 +3167,7 @@ pub(in crate::set_disk) struct RenameDataCommit {
pub(in crate::set_disk) committed_file_info: FileInfo, pub(in crate::set_disk) committed_file_info: FileInfo,
} }
#[allow(dead_code, reason = "asserted by this file's tests (backlog#1823)")]
type RenameDataLegacyTuple = ( type RenameDataLegacyTuple = (
Vec<Option<DiskStore>>, Vec<Option<DiskStore>>,
RenameConvergence, RenameConvergence,
@@ -3030,6 +3177,7 @@ type RenameDataLegacyTuple = (
); );
impl RenameDataCommit { impl RenameDataCommit {
#[allow(dead_code, reason = "asserted by this file's tests (backlog#1823)")]
fn into_legacy_tuple(self) -> RenameDataLegacyTuple { fn into_legacy_tuple(self) -> RenameDataLegacyTuple {
( (
self.online_disks, self.online_disks,
@@ -3148,6 +3296,7 @@ impl SetDisks {
#[tracing::instrument(level = "debug", skip(disks, file_infos))] #[tracing::instrument(level = "debug", skip(disks, file_infos))]
#[allow(clippy::type_complexity)] #[allow(clippy::type_complexity)]
#[allow(dead_code, reason = "asserted by this file's tests (backlog#1823)")]
pub(in crate::set_disk) async fn rename_data( pub(in crate::set_disk) async fn rename_data(
disks: &[Option<DiskStore>], disks: &[Option<DiskStore>],
src_bucket: &str, src_bucket: &str,
@@ -4960,6 +5109,7 @@ fn is_cleanup_not_found(e: &DiskError) -> bool {
/// normalized to `DiskNotFound`: a panic is not a "disk absent" condition and /// normalized to `DiskNotFound`: a panic is not a "disk absent" condition and
/// must not be silently swallowed as an ignorable error (fixes the historical /// must not be silently swallowed as an ignorable error (fixes the historical
/// `Unexpected`/`DiskNotFound` misclassification). /// `Unexpected`/`DiskNotFound` misclassification).
#[allow(dead_code, reason = "asserted by this file's tests (backlog#1823)")]
fn map_cleanup_join_result(joined: std::result::Result<Option<DiskError>, tokio::task::JoinError>) -> Option<DiskError> { fn map_cleanup_join_result(joined: std::result::Result<Option<DiskError>, tokio::task::JoinError>) -> Option<DiskError> {
match joined { match joined {
Ok(res) => res, Ok(res) => res,
@@ -5184,6 +5334,7 @@ pub(in crate::set_disk) mod rename_fanout_barrier_phase {
/// The per-disk old-data-dir cleanup phase of the commit fan-out. /// The per-disk old-data-dir cleanup phase of the commit fan-out.
pub const CLEANUP: &str = "cleanup"; pub const CLEANUP: &str = "cleanup";
/// The per-disk `read_version` phase of metadata read fan-out. /// The per-disk `read_version` phase of metadata read fan-out.
#[allow(dead_code, reason = "asserted by this file's tests (backlog#1823)")]
pub const READ_VERSION: &str = "read_version"; pub const READ_VERSION: &str = "read_version";
} }
@@ -5621,6 +5772,130 @@ mod tests {
(dirs, disks) (dirs, disks)
} }
#[test]
fn metadata_slowtail_fault_delay_parses_and_filters_request() {
temp_env::with_vars(
[
(ENV_RUSTFS_GET_METADATA_SLOWTAIL_FAULT_DELAY_MS, Some("25")),
(ENV_RUSTFS_GET_METADATA_SLOWTAIL_FAULT_DISKS, Some("1,3")),
(ENV_RUSTFS_GET_METADATA_SLOWTAIL_FAULT_BUCKET, Some("bench-bucket")),
(ENV_RUSTFS_GET_METADATA_SLOWTAIL_FAULT_OBJECT_PREFIX, Some("objects/")),
],
|| {
assert_eq!(
get_metadata_slowtail_fault_delay("bench-bucket", "objects/000001", 3, true),
Some(Duration::from_millis(25))
);
assert!(get_metadata_slowtail_fault_delay("bench-bucket", "objects/000001", 2, true).is_none());
assert!(get_metadata_slowtail_fault_delay("other-bucket", "objects/000001", 3, true).is_none());
assert!(get_metadata_slowtail_fault_delay("bench-bucket", "other/000001", 3, true).is_none());
assert!(get_metadata_slowtail_fault_delay("bench-bucket", "objects/000001", 3, false).is_none());
},
);
}
#[test]
fn metadata_slowtail_fault_delay_disables_invalid_disk_list() {
temp_env::with_vars(
[
(ENV_RUSTFS_GET_METADATA_SLOWTAIL_FAULT_DELAY_MS, Some("25")),
(ENV_RUSTFS_GET_METADATA_SLOWTAIL_FAULT_DISKS, Some("1,nope")),
],
|| {
assert!(get_metadata_slowtail_fault_delay("bucket", "object", 1, true).is_none());
},
);
}
#[tokio::test(flavor = "multi_thread", worker_threads = 2)]
async fn metadata_slowtail_fault_delays_only_data_read_metadata_task() {
const DISKS: usize = 4;
let bucket = "metadata-slowtail-fault-bucket";
let object = "objects/metadata-slowtail-fault-object";
let (dirs, disks) = call_counter_local_disks(bucket, DISKS).await;
install_metadata_fanout_fileinfo(&disks, bucket, object, None).await;
temp_env::async_with_vars(
[
(ENV_RUSTFS_GET_METADATA_EARLY_STOP_ENABLE, Some("false")),
(ENV_RUSTFS_GET_METADATA_SLOWTAIL_FAULT_DELAY_MS, Some("150")),
(ENV_RUSTFS_GET_METADATA_SLOWTAIL_FAULT_DISKS, Some("3")),
(ENV_RUSTFS_GET_METADATA_SLOWTAIL_FAULT_BUCKET, Some(bucket)),
(ENV_RUSTFS_GET_METADATA_SLOWTAIL_FAULT_OBJECT_PREFIX, Some("objects/")),
],
async {
let read_without_data =
SetDisks::read_all_fileinfo_observed(&disks, bucket, bucket, object, "", false, false, false, true, 2);
tokio::time::timeout(Duration::from_millis(100), read_without_data)
.await
.expect("non-data metadata fanout must not be delayed by the data-read slowtail hook")
.expect("metadata fanout without read_data should resolve");
let mut read_with_data = Box::pin(SetDisks::read_all_fileinfo_observed(
&disks, bucket, bucket, object, "", true, false, false, true, 2,
));
assert!(
tokio::time::timeout(Duration::from_millis(40), &mut read_with_data)
.await
.is_err(),
"data-read metadata fanout must wait for the injected slow read_version response"
);
let (parts_metadata, errs, diagnostics) = tokio::time::timeout(Duration::from_secs(2), read_with_data)
.await
.expect("injected slowtail should eventually complete")
.expect("data-read metadata fanout should resolve");
assert_eq!(parts_metadata.iter().filter(|fi| fi.name == object).count(), DISKS);
assert!(errs.iter().all(Option::is_none));
assert_eq!(diagnostics.total_responses(), DISKS);
},
)
.await;
drop(dirs);
}
#[tokio::test(flavor = "multi_thread", worker_threads = 2)]
async fn metadata_slowtail_fault_delays_early_stop_metadata_task() {
const DISKS: usize = 4;
let bucket = "metadata-slowtail-early-stop-bucket";
let object = "objects/metadata-slowtail-early-stop-object";
let (dirs, disks) = call_counter_local_disks(bucket, DISKS).await;
install_metadata_fanout_fileinfo(&disks, bucket, object, None).await;
temp_env::async_with_vars(
[
(ENV_RUSTFS_GET_METADATA_EARLY_STOP_ENABLE, Some("true")),
(ENV_RUSTFS_GET_METADATA_DATA_READ_EARLY_STOP_ENABLE, Some("true")),
(ENV_RUSTFS_GET_METADATA_EARLY_STOP_BOUNDED_FANOUT, Some("false")),
(ENV_RUSTFS_GET_METADATA_SLOWTAIL_FAULT_DELAY_MS, Some("150")),
(ENV_RUSTFS_GET_METADATA_SLOWTAIL_FAULT_DISKS, Some("3")),
(ENV_RUSTFS_GET_METADATA_SLOWTAIL_FAULT_BUCKET, Some(bucket)),
(ENV_RUSTFS_GET_METADATA_SLOWTAIL_FAULT_OBJECT_PREFIX, Some("objects/")),
],
async {
let mut read_with_data = Box::pin(SetDisks::read_all_fileinfo_observed(
&disks, bucket, bucket, object, "", true, false, false, true, 2,
));
assert!(
tokio::time::timeout(Duration::from_millis(40), &mut read_with_data)
.await
.is_err(),
"early-stop metadata fanout must still wait for the injected slow response after fallback to full wait"
);
let (parts_metadata, errs, diagnostics) = tokio::time::timeout(Duration::from_secs(2), read_with_data)
.await
.expect("injected early-stop slowtail should eventually complete")
.expect("early-stop metadata fanout should resolve");
assert_eq!(parts_metadata.iter().filter(|fi| fi.name == object).count(), DISKS);
assert!(errs.iter().all(Option::is_none));
assert_eq!(diagnostics.total_responses(), DISKS);
},
)
.await;
drop(dirs);
}
/// Demo / regression guard for the backlog#1325 per-disk call counters. /// Demo / regression guard for the backlog#1325 per-disk call counters.
/// ///
/// The metadata fan-out issues each `read_version` inside its own /// The metadata fan-out issues each `read_version` inside its own
@@ -5752,9 +6027,20 @@ mod tests {
object: &str, object: &str,
payload: &[u8], payload: &[u8],
uses_legacy_checksum: bool, uses_legacy_checksum: bool,
) -> Vec<FileInfo> {
inline_metadata_fanout_fileinfos_with_geometry(bucket, object, payload, uses_legacy_checksum, 2, 2).await
}
async fn inline_metadata_fanout_fileinfos_with_geometry(
bucket: &str,
object: &str,
payload: &[u8],
uses_legacy_checksum: bool,
data_shards: usize,
parity_shards: usize,
) -> Vec<FileInfo> { ) -> Vec<FileInfo> {
let distribution_key = metadata_distribution_key(bucket, object); let distribution_key = metadata_distribution_key(bucket, object);
let mut base = FileInfo::new(&distribution_key, 2, 2); let mut base = FileInfo::new(&distribution_key, data_shards, parity_shards);
base.volume = bucket.to_string(); base.volume = bucket.to_string();
base.name = object.to_string(); base.name = object.to_string();
base.size = i64::try_from(payload.len()).expect("test payload should fit i64"); base.size = i64::try_from(payload.len()).expect("test payload should fit i64");
@@ -5817,6 +6103,21 @@ mod tests {
install_inline_metadata_fanout_files(disks, bucket, object, files).await; install_inline_metadata_fanout_files(disks, bucket, object, files).await;
} }
async fn install_inline_metadata_fanout_fileinfo_with_geometry(
disks: &[Option<DiskStore>],
bucket: &str,
object: &str,
payload: &[u8],
data_shards: usize,
parity_shards: usize,
mutate: impl FnOnce(&mut [FileInfo]),
) {
let mut files =
inline_metadata_fanout_fileinfos_with_geometry(bucket, object, payload, false, data_shards, parity_shards).await;
mutate(&mut files);
install_inline_metadata_fanout_files(disks, bucket, object, files).await;
}
async fn install_inline_metadata_fanout_files(disks: &[Option<DiskStore>], bucket: &str, object: &str, files: Vec<FileInfo>) { async fn install_inline_metadata_fanout_files(disks: &[Option<DiskStore>], bucket: &str, object: &str, files: Vec<FileInfo>) {
let distribution = files let distribution = files
.first() .first()
@@ -6037,6 +6338,118 @@ mod tests {
drop(dirs); drop(dirs);
} }
#[tokio::test(flavor = "multi_thread", worker_threads = 2)]
async fn bounded_metadata_early_stop_waits_for_pending_inline_data_shard() {
const DISKS: usize = 6;
const DATA_SHARDS: usize = 4;
const PARITY_SHARDS: usize = 2;
let bucket = "bounded-inline-data-get-pending-shard-bucket";
let object =
object_with_initial_data_shards(bucket, "bounded-inline-data-get-pending-shard-object", DATA_SHARDS, DATA_SHARDS);
let (dirs, disks) = call_counter_local_disks(bucket, DISKS).await;
install_inline_metadata_fanout_fileinfo_with_geometry(
&disks,
bucket,
&object,
b"verified inline payload",
DATA_SHARDS,
PARITY_SHARDS,
|_| {},
)
.await;
temp_env::async_with_vars(
[
("RUSTFS_GET_METADATA_EARLY_STOP_ENABLE", Some("true")),
("RUSTFS_GET_METADATA_DATA_READ_EARLY_STOP_ENABLE", Some("true")),
("RUSTFS_GET_METADATA_EARLY_STOP_BOUNDED_FANOUT", Some("true")),
],
async {
let fanout_order = bounded_metadata_fanout_order(bucket, &object, DISKS, PARITY_SHARDS);
let distribution_key = metadata_distribution_key(bucket, &object);
let distribution = FileInfo::new(&distribution_key, DATA_SHARDS, PARITY_SHARDS)
.erasure
.distribution;
let paused_data_disk = *fanout_order
.iter()
.take(DATA_SHARDS)
.find(|disk_index| {
distribution
.get(**disk_index)
.is_some_and(|block_index| (1..=DATA_SHARDS).contains(block_index))
})
.expect("initial fanout should include a data shard to pause");
let hedged_parity_disk = fanout_order[DATA_SHARDS];
let unscheduled_parity_disk = fanout_order[DATA_SHARDS + 1];
let barrier = rename_fanout_barrier::arm(&object, paused_data_disk, rename_fanout_barrier::PHASE_READ_VERSION);
let tracker = rename_fanout_barrier::observe_tasks(&object);
let calls = disk_call_counters::observe(&object);
let disks_for_read = disks.clone();
let object_for_read = object.clone();
let mut read = tokio::spawn(async move {
SetDisks::read_all_fileinfo_observed(
&disks_for_read,
bucket,
bucket,
&object_for_read,
"",
true,
false,
false,
true,
PARITY_SHARDS,
)
.await
});
tokio::time::timeout(BARRIER_PAUSE_GUARD, barrier.wait_until_paused())
.await
.expect("initial data shard should pause before returning");
tokio::time::timeout(BARRIER_PAUSE_GUARD, async {
while calls.for_disk(disk_call_counters::KIND_READ_VERSION, hedged_parity_disk) == 0 {
tokio::task::yield_now().await;
}
})
.await
.expect("bounded fanout should hedge one parity disk while the data shard is pending");
assert!(
tokio::time::timeout(BARRIER_PAUSE_GUARD, &mut read).await.is_err(),
"inline data-read early-stop must wait for a scheduled missing data shard instead of forcing full wait"
);
barrier.release();
let (parts_metadata, errs, diagnostics) = read
.await
.expect("metadata read task should not panic")
.expect("pending data shard should let the inline verifier finish");
assert_eq!(
calls.total(disk_call_counters::KIND_READ_VERSION),
5,
"pending data-shard defer should not schedule the final parity disk"
);
assert_eq!(
calls.for_disk(disk_call_counters::KIND_READ_VERSION, unscheduled_parity_disk),
0,
"the remaining parity disk must stay unissued when pending data verification succeeds"
);
assert_eq!(
tracker.running(),
0,
"early-stop should drain spawned read_version tasks before returning"
);
assert_eq!(diagnostics.total_responses(), 5);
assert_eq!(parts_metadata.iter().filter(|fi| fi.name == object).count(), 5);
assert!(errs.iter().all(Option::is_none));
},
)
.await;
drop(dirs);
}
#[tokio::test] #[tokio::test]
async fn data_read_early_stop_verifies_legacy_inline_checksum_payload() { async fn data_read_early_stop_verifies_legacy_inline_checksum_payload() {
let bucket = "legacy-inline-data-get-fanout-bucket"; let bucket = "legacy-inline-data-get-fanout-bucket";
@@ -6067,11 +6480,133 @@ mod tests {
.clone(); .clone();
assert!( assert!(
data_read_early_stop_inline_body_verified(bucket, object, &candidate, &parts_metadata, &disks).await, data_read_early_stop_inline_body_miss_reason(bucket, object, &candidate, &parts_metadata, &disks)
.await
.is_none(),
"legacy inline metadata must use the legacy bitrot shard sizing and checksum algorithm" "legacy inline metadata must use the legacy bitrot shard sizing and checksum algorithm"
); );
} }
#[tokio::test]
async fn data_read_early_stop_reports_inline_miss_reasons() {
let bucket = "inline-data-get-miss-reason-bucket";
let object = "inline-data-get-miss-reason-object";
let payload = b"verified inline payload";
let (_dirs, disks) = call_counter_local_disks(bucket, 4).await;
let files = inline_metadata_fanout_fileinfos_with_mode(bucket, object, payload, false).await;
let distribution = files
.first()
.map(|file| file.erasure.distribution.clone())
.expect("fixture should include metadata");
let order = bounded_metadata_fanout_order(bucket, object, 4, 2);
let mut parts_metadata = vec![FileInfo::default(); 4];
for disk_index in order.into_iter().take(3) {
let block_index = distribution
.get(disk_index)
.copied()
.expect("fixture distribution should cover every disk");
parts_metadata[disk_index] = files
.get(block_index.checked_sub(1).expect("erasure block indexes are one-based"))
.expect("fixture should include every distributed shard")
.clone();
}
let candidate = parts_metadata
.iter()
.find(|file| file.name == object)
.expect("fixture should include observed metadata")
.clone();
let data_disk = distribution
.iter()
.position(|block_index| *block_index == 1)
.expect("fixture distribution should include first data shard");
assert_eq!(
data_read_early_stop_inline_body_miss_reason(bucket, object, &candidate, &parts_metadata, &disks).await,
None
);
let mut not_inline = candidate.clone();
rustfs_utils::http::remove_str(&mut not_inline.metadata, rustfs_utils::http::SUFFIX_INLINE_DATA);
assert_eq!(
data_read_early_stop_inline_body_miss_reason(bucket, object, &not_inline, &parts_metadata, &disks).await,
Some(GET_METADATA_EARLY_STOP_REASON_DATA_READ_INLINE_NOT_INLINE)
);
let mut remote = candidate.clone();
remote.transition_status = TRANSITION_COMPLETE.to_string();
assert_eq!(
data_read_early_stop_inline_body_miss_reason(bucket, object, &remote, &parts_metadata, &disks).await,
Some(GET_METADATA_EARLY_STOP_REASON_DATA_READ_INLINE_REMOTE)
);
let mut transformed = candidate.clone();
rustfs_utils::http::insert_str(&mut transformed.metadata, rustfs_utils::http::SUFFIX_COMPRESSION, "zstd".to_string());
assert_eq!(
data_read_early_stop_inline_body_miss_reason(bucket, object, &transformed, &parts_metadata, &disks).await,
Some(GET_METADATA_EARLY_STOP_REASON_DATA_READ_INLINE_TRANSFORMED)
);
let mut deleted = candidate.clone();
deleted.deleted = true;
assert_eq!(
data_read_early_stop_inline_body_miss_reason(bucket, object, &deleted, &parts_metadata, &disks).await,
Some(GET_METADATA_EARLY_STOP_REASON_DATA_READ_INLINE_DELETED)
);
let mut zero_size = candidate.clone();
zero_size.size = 0;
assert_eq!(
data_read_early_stop_inline_body_miss_reason(bucket, object, &zero_size, &parts_metadata, &disks).await,
Some(GET_METADATA_EARLY_STOP_REASON_DATA_READ_INLINE_SIZE)
);
let mut multipart = candidate.clone();
multipart.parts.push(multipart.parts[0].clone());
assert_eq!(
data_read_early_stop_inline_body_miss_reason(bucket, object, &multipart, &parts_metadata, &disks).await,
Some(GET_METADATA_EARLY_STOP_REASON_DATA_READ_INLINE_PART_SHAPE)
);
let mut invalid_geometry = candidate.clone();
invalid_geometry.erasure.data_blocks = 0;
assert_eq!(
data_read_early_stop_inline_body_miss_reason(bucket, object, &invalid_geometry, &parts_metadata, &disks).await,
Some(GET_METADATA_EARLY_STOP_REASON_DATA_READ_INLINE_GEOMETRY)
);
let mut missing_shard = parts_metadata.clone();
missing_shard[data_disk] = FileInfo::default();
assert_eq!(
data_read_early_stop_inline_body_miss_reason(bucket, object, &candidate, &missing_shard, &disks).await,
Some(GET_METADATA_EARLY_STOP_REASON_DATA_READ_INLINE_MISSING_SHARD)
);
let mut missing_payload = parts_metadata.clone();
missing_payload[data_disk].data = None;
assert_eq!(
data_read_early_stop_inline_body_miss_reason(bucket, object, &candidate, &missing_payload, &disks).await,
Some(GET_METADATA_EARLY_STOP_REASON_DATA_READ_INLINE_MISSING_PAYLOAD)
);
let mut identity_mismatch = parts_metadata.clone();
identity_mismatch[data_disk].version_id = Some(Uuid::new_v4());
assert_eq!(
data_read_early_stop_inline_body_miss_reason(bucket, object, &candidate, &identity_mismatch, &disks).await,
Some(GET_METADATA_EARLY_STOP_REASON_DATA_READ_INLINE_IDENTITY_MISMATCH)
);
let mut corrupt = parts_metadata.clone();
if let Some(data) = corrupt[data_disk].data.as_mut() {
let mut corrupt_data = data.to_vec();
corrupt_data[0] ^= 0x01;
*data = Bytes::from(corrupt_data);
}
assert_eq!(
data_read_early_stop_inline_body_miss_reason(bucket, object, &candidate, &corrupt, &disks).await,
Some(GET_METADATA_EARLY_STOP_REASON_DATA_READ_INLINE_BODY_VERIFY)
);
}
#[test] #[test]
#[serial_test::serial] #[serial_test::serial]
fn metadata_fanout_lifecycle_records_real_early_stop_abort() { fn metadata_fanout_lifecycle_records_real_early_stop_abort() {
@@ -6161,7 +6696,7 @@ mod tests {
&[ &[
("path", GET_OBJECT_PATH_INTERNAL_META), ("path", GET_OBJECT_PATH_INTERNAL_META),
("decision", "miss"), ("decision", "miss"),
("reason", GET_METADATA_EARLY_STOP_REASON_INSUFFICIENT_QUORUM), ("reason", GET_METADATA_EARLY_STOP_REASON_DATA_READ_INLINE_BODY_VERIFY),
], ],
), ),
1, 1,
@@ -6173,7 +6708,7 @@ mod tests {
&[ &[
("path", GET_OBJECT_PATH_LEGACY_DUPLEX), ("path", GET_OBJECT_PATH_LEGACY_DUPLEX),
("decision", "miss"), ("decision", "miss"),
("reason", GET_METADATA_EARLY_STOP_REASON_INSUFFICIENT_QUORUM), ("reason", GET_METADATA_EARLY_STOP_REASON_DATA_READ_INLINE_BODY_VERIFY),
], ],
), ),
0, 0,
@@ -6708,7 +7243,7 @@ mod tests {
} }
#[tokio::test(flavor = "multi_thread", worker_threads = 2)] #[tokio::test(flavor = "multi_thread", worker_threads = 2)]
async fn bounded_non_inline_data_get_hedges_then_waits_for_full_fanout() { async fn bounded_non_inline_data_get_immediately_forces_full_fanout() {
const DISKS: usize = 4; const DISKS: usize = 4;
let bucket = "bounded-data-get-hedge-bucket"; let bucket = "bounded-data-get-hedge-bucket";
let object = "bounded-data-get-hedge-object"; let object = "bounded-data-get-hedge-object";
@@ -6739,7 +7274,7 @@ mod tests {
} }
}) })
.await .await
.expect("bounded data-read fanout should hedge by starting the spare disk"); .expect("bounded non-inline data-read fanout should immediately schedule the spare disk");
let pending = tokio::time::timeout(BARRIER_PAUSE_GUARD, &mut read).await; let pending = tokio::time::timeout(BARRIER_PAUSE_GUARD, &mut read).await;
assert!( assert!(
@@ -6755,7 +7290,7 @@ mod tests {
assert_eq!( assert_eq!(
calls.total(disk_call_counters::KIND_READ_VERSION), calls.total(disk_call_counters::KIND_READ_VERSION),
DISKS as u64, DISKS as u64,
"bounded data-read fanout should issue the paused disk plus one spare hedge" "bounded non-inline data-read fanout should issue the paused disk plus the remaining spare"
); );
assert_eq!(diagnostics.total_responses(), DISKS); assert_eq!(diagnostics.total_responses(), DISKS);
assert_eq!(parts_metadata.iter().filter(|fi| fi.name == object).count(), DISKS); assert_eq!(parts_metadata.iter().filter(|fi| fi.name == object).count(), DISKS);
@@ -6768,7 +7303,7 @@ mod tests {
} }
#[tokio::test] #[tokio::test]
async fn bounded_metadata_early_stop_defaults_keep_data_get_full_fanout() { async fn bounded_metadata_early_stop_defaults_keep_non_inline_data_get_full_fanout() {
const DISKS: usize = 4; const DISKS: usize = 4;
let bucket = "bounded-data-get-default-bucket"; let bucket = "bounded-data-get-default-bucket";
let object = "bounded-data-get-default-object"; let object = "bounded-data-get-default-object";
@@ -6782,16 +7317,42 @@ mod tests {
("RUSTFS_GET_METADATA_EARLY_STOP_BOUNDED_FANOUT", None::<&str>), ("RUSTFS_GET_METADATA_EARLY_STOP_BOUNDED_FANOUT", None::<&str>),
], ],
async { async {
let barrier = rename_fanout_barrier::arm(object, 2, rename_fanout_barrier::PHASE_READ_VERSION);
let calls = disk_call_counters::observe(object); let calls = disk_call_counters::observe(object);
let (parts_metadata, errs, diagnostics) = let disks_for_read = disks.clone();
SetDisks::read_all_fileinfo_observed(&disks, bucket, bucket, object, "", true, false, false, true, 2) let mut read = tokio::spawn(async move {
SetDisks::read_all_fileinfo_observed(&disks_for_read, bucket, bucket, object, "", true, false, false, true, 2)
.await .await
.expect("default data-read metadata should resolve"); });
tokio::time::timeout(BARRIER_PAUSE_GUARD, barrier.wait_until_paused())
.await
.expect("default bounded non-inline read should schedule the paused metadata task");
tokio::time::timeout(BARRIER_PAUSE_GUARD, async {
while calls.for_disk(disk_call_counters::KIND_READ_VERSION, 3) == 0 {
tokio::task::yield_now().await;
}
})
.await
.expect(
"default bounded non-inline read should immediately force full fanout after the first non-inline response",
);
let pending = tokio::time::timeout(BARRIER_PAUSE_GUARD, &mut read).await;
assert!(
pending.is_err(),
"default non-inline data reads must not return before the paused metadata response"
);
barrier.release();
let (parts_metadata, errs, diagnostics) = read
.await
.expect("metadata read task should not panic")
.expect("default data-read metadata should resolve");
assert_eq!( assert_eq!(
calls.total(disk_call_counters::KIND_READ_VERSION), calls.total(disk_call_counters::KIND_READ_VERSION),
DISKS as u64, DISKS as u64,
"default GET data-read metadata must keep full fanout for read-failure tolerance" "default non-inline GET data-read metadata must keep full fanout without waiting for a quorum miss first"
); );
assert_eq!(diagnostics.total_responses(), DISKS); assert_eq!(diagnostics.total_responses(), DISKS);
assert_eq!(parts_metadata.iter().filter(|fi| fi.name == object).count(), DISKS); assert_eq!(parts_metadata.iter().filter(|fi| fi.name == object).count(), DISKS);
+28
View File
@@ -42,12 +42,20 @@ impl<'a> SetDisksCtx<'a> {
} }
/// The borrowed core, for state not yet fronted by a typed accessor. /// The borrowed core, for state not yet fronted by a typed accessor.
#[allow(
dead_code,
reason = "SetDisks split seam (backlog#815) with no caller in this port (backlog#1823)"
)]
pub(crate) fn core(&self) -> &'a SetDisks { pub(crate) fn core(&self) -> &'a SetDisks {
self.core self.core
} }
// --- Immutable topology / config (fixed after construction) --- // --- Immutable topology / config (fixed after construction) ---
#[allow(
dead_code,
reason = "SetDisks split seam (backlog#815) with no caller in this port (backlog#1823)"
)]
pub(crate) fn set_index(&self) -> usize { pub(crate) fn set_index(&self) -> usize {
self.core.set_index self.core.set_index
} }
@@ -56,14 +64,26 @@ impl<'a> SetDisksCtx<'a> {
self.core.pool_index self.core.pool_index
} }
#[allow(
dead_code,
reason = "SetDisks split seam (backlog#815) with no caller in this port (backlog#1823)"
)]
pub(crate) fn set_drive_count(&self) -> usize { pub(crate) fn set_drive_count(&self) -> usize {
self.core.set_drive_count self.core.set_drive_count
} }
#[allow(
dead_code,
reason = "SetDisks split seam (backlog#815) with no caller in this port (backlog#1823)"
)]
pub(crate) fn default_parity_count(&self) -> usize { pub(crate) fn default_parity_count(&self) -> usize {
self.core.default_parity_count self.core.default_parity_count
} }
#[allow(
dead_code,
reason = "SetDisks split seam (backlog#815) with no caller in this port (backlog#1823)"
)]
pub(crate) fn set_endpoints(&self) -> &'a [Endpoint] { pub(crate) fn set_endpoints(&self) -> &'a [Endpoint] {
&self.core.set_endpoints &self.core.set_endpoints
} }
@@ -72,6 +92,10 @@ impl<'a> SetDisksCtx<'a> {
&self.core.format &self.core.format
} }
#[allow(
dead_code,
reason = "SetDisks split seam (backlog#815) with no caller in this port (backlog#1823)"
)]
pub(crate) fn locker_owner(&self) -> &'a str { pub(crate) fn locker_owner(&self) -> &'a str {
&self.core.locker_owner &self.core.locker_owner
} }
@@ -84,6 +108,10 @@ impl<'a> SetDisksCtx<'a> {
// --- Locker trio --- // --- Locker trio ---
#[allow(
dead_code,
reason = "SetDisks split seam (backlog#815) with no caller in this port (backlog#1823)"
)]
pub(crate) fn lockers(&self) -> &'a [Arc<dyn LockClient>] { pub(crate) fn lockers(&self) -> &'a [Arc<dyn LockClient>] {
&self.core.lockers &self.core.lockers
} }
+177 -48
View File
@@ -39,7 +39,6 @@
//! - `metadata.rs`, `replication.rs`, `shard_source.rs` — supporting helpers. //! - `metadata.rs`, `replication.rs`, `shard_source.rs` — supporting helpers.
// #730: SetDisks still hosts staged read/heal/write migration helpers. // #730: SetDisks still hosts staged read/heal/write migration helpers.
#![allow(dead_code)]
#![allow(unused_imports)] #![allow(unused_imports)]
#![allow(unused_variables)] #![allow(unused_variables)]
@@ -59,7 +58,10 @@ use crate::client::{object_api_utils::get_raw_etag, transition_api::ReaderImpl};
use crate::cluster::rpc::heal_bucket_local_on_disks; use crate::cluster::rpc::heal_bucket_local_on_disks;
use crate::data_usage::record_compression_total_memory; use crate::data_usage::record_compression_total_memory;
use crate::diagnostics::get::{ use crate::diagnostics::get::{
GET_CODEC_STREAMING_OBJECT_CLASS_PLAIN_SINGLE_PART, GET_OBJECT_PATH_BODY_CACHE, GET_OBJECT_PATH_CODEC_STREAMING, GET_CODEC_STREAMING_OBJECT_CLASS_PLAIN_SINGLE_PART, GET_METADATA_EARLY_STOP_REASON_DATA_READ_INLINE_GEOMETRY,
GET_METADATA_EARLY_STOP_REASON_DATA_READ_INLINE_IDENTITY_MISMATCH,
GET_METADATA_EARLY_STOP_REASON_DATA_READ_INLINE_MISSING_PAYLOAD,
GET_METADATA_EARLY_STOP_REASON_DATA_READ_INLINE_MISSING_SHARD, GET_OBJECT_PATH_BODY_CACHE, GET_OBJECT_PATH_CODEC_STREAMING,
GET_OBJECT_PATH_CODEC_STREAMING_LEGACY_ENGINE, GET_OBJECT_PATH_CODEC_STREAMING_RUSTFS_ENGINE, GET_OBJECT_PATH_DIRECT_MEMORY, GET_OBJECT_PATH_CODEC_STREAMING_LEGACY_ENGINE, GET_OBJECT_PATH_CODEC_STREAMING_RUSTFS_ENGINE, GET_OBJECT_PATH_DIRECT_MEMORY,
GET_OBJECT_PATH_EMPTY, GET_OBJECT_PATH_INLINE_DIRECT, GET_OBJECT_PATH_INTERNAL_META, GET_OBJECT_PATH_LEGACY_DUPLEX, GET_OBJECT_PATH_EMPTY, GET_OBJECT_PATH_INLINE_DIRECT, GET_OBJECT_PATH_INTERNAL_META, GET_OBJECT_PATH_LEGACY_DUPLEX,
GET_OBJECT_PATH_REMOTE_TRANSITION, GET_OBJECT_PATH_SET_DISK, GET_STAGE_DECODE, GET_STAGE_EMIT, GET_STAGE_INLINE_PREPARE, GET_OBJECT_PATH_REMOTE_TRANSITION, GET_OBJECT_PATH_SET_DISK, GET_STAGE_DECODE, GET_STAGE_EMIT, GET_STAGE_INLINE_PREPARE,
@@ -100,9 +102,7 @@ use crate::storage_api_contracts::{
}; };
use crate::store::utils::is_reserved_or_invalid_bucket; use crate::store::utils::is_reserved_or_invalid_bucket;
use crate::{ use crate::{
bucket::lifecycle::bucket_lifecycle_ops::{ bucket::lifecycle::bucket_lifecycle_ops::{LifecycleOps, get_transitioned_object_reader_with_tier_manager, put_restore_opts},
LifecycleOps, gen_transition_objname, get_transitioned_object_reader_with_tier_manager, put_restore_opts,
},
cache_value::metacache_set::{ListPathRawOptions, list_path_raw}, cache_value::metacache_set::{ListPathRawOptions, list_path_raw},
config::storageclass, config::storageclass,
disk::{ disk::{
@@ -174,15 +174,14 @@ use std::future::Future;
use std::hash::{BuildHasher, Hash, Hasher}; use std::hash::{BuildHasher, Hash, Hasher};
use std::mem::{self}; use std::mem::{self};
use std::pin::Pin; use std::pin::Pin;
use std::sync::OnceLock;
use std::sync::atomic::{AtomicBool, AtomicU64, Ordering}; use std::sync::atomic::{AtomicBool, AtomicU64, Ordering};
use std::sync::{Arc, OnceLock};
use std::task::{Context, Poll}; use std::task::{Context, Poll};
use std::time::{Instant, SystemTime, UNIX_EPOCH}; use std::time::{Instant, SystemTime, UNIX_EPOCH};
use std::{ use std::{
collections::{HashMap, HashSet}, collections::{HashMap, HashSet},
io::{Cursor, Write}, io::{Cursor, Write},
path::Path, path::Path,
sync::Arc,
time::Duration, time::Duration,
}; };
use time::OffsetDateTime; use time::OffsetDateTime;
@@ -621,7 +620,9 @@ fn adaptive_duplex_buffer_size(object_size: i64) -> usize {
// Each flag has a corresponding `*_ROLLOUT_PCT` for percentage-based gradual rollout. // Each flag has a corresponding `*_ROLLOUT_PCT` for percentage-based gradual rollout.
// ============================================================================ // ============================================================================
#[allow(dead_code, reason = "asserted by this file's tests (backlog#1823)")]
const DISK_ONLINE_TIMEOUT: Duration = Duration::from_secs(1); const DISK_ONLINE_TIMEOUT: Duration = Duration::from_secs(1);
#[allow(dead_code, reason = "asserted by this file's tests (backlog#1823)")]
const DISK_HEALTH_CACHE_TTL: Duration = Duration::from_millis(750); const DISK_HEALTH_CACHE_TTL: Duration = Duration::from_millis(750);
const GET_OBJECT_METADATA_CACHE_TTL: Duration = Duration::from_secs(2); // Increased from 250ms to 2s const GET_OBJECT_METADATA_CACHE_TTL: Duration = Duration::from_secs(2); // Increased from 250ms to 2s
const DEFAULT_GET_OBJECT_METADATA_CACHE_MAX_ENTRIES: usize = 4096; // Increased from 1024 to 4096 const DEFAULT_GET_OBJECT_METADATA_CACHE_MAX_ENTRIES: usize = 4096; // Increased from 1024 to 4096
@@ -689,22 +690,36 @@ const DEFAULT_RUSTFS_GET_SMALL_OBJECT_DIRECT_MEMORY_THRESHOLD: usize = 128 * 102
const ENV_RUSTFS_GET_METADATA_EARLY_STOP_ENABLE: &str = "RUSTFS_GET_METADATA_EARLY_STOP_ENABLE"; const ENV_RUSTFS_GET_METADATA_EARLY_STOP_ENABLE: &str = "RUSTFS_GET_METADATA_EARLY_STOP_ENABLE";
// Enabled by default (backlog#872): the early-stop path only engages for // Enabled by default (backlog#872): the early-stop path only engages for
// requests `should_allow_metadata_early_stop` classifies as safe (latest-version // requests `should_allow_metadata_early_stop` classifies as safe (latest-version
// metadata-only reads by default, without version_id / healing / free-version // reads by default, without version_id / healing / free-version needs) and still
// needs) and still requires a full read-quorum agreement before stopping. Set // requires a full read-quorum agreement before stopping. Data-read requests add
// a separate inline-shard verifier before cancelling the remaining fanout. Set
// the env var to `false` to fall back to full-wait metadata fanout. // the env var to `false` to fall back to full-wait metadata fanout.
const DEFAULT_RUSTFS_GET_METADATA_EARLY_STOP_ENABLE: bool = true; const DEFAULT_RUSTFS_GET_METADATA_EARLY_STOP_ENABLE: bool = true;
#[allow(
dead_code,
reason = "percentage-rollout facet of the metadata early-stop switch; its predicate has no caller while the sibling enable flag is live (backlog#1823)"
)]
const ENV_RUSTFS_GET_METADATA_EARLY_STOP_ROLLOUT_PCT: &str = "RUSTFS_GET_METADATA_EARLY_STOP_ROLLOUT_PCT"; const ENV_RUSTFS_GET_METADATA_EARLY_STOP_ROLLOUT_PCT: &str = "RUSTFS_GET_METADATA_EARLY_STOP_ROLLOUT_PCT";
#[allow(
dead_code,
reason = "percentage-rollout facet of the metadata early-stop switch; its predicate has no caller while the sibling enable flag is live (backlog#1823)"
)]
const DEFAULT_RUSTFS_GET_METADATA_EARLY_STOP_ROLLOUT_PCT: u32 = 100; const DEFAULT_RUSTFS_GET_METADATA_EARLY_STOP_ROLLOUT_PCT: u32 = 100;
const ENV_RUSTFS_GET_METADATA_VERSION_EARLY_STOP_ENABLE: &str = "RUSTFS_GET_METADATA_VERSION_EARLY_STOP_ENABLE"; const ENV_RUSTFS_GET_METADATA_VERSION_EARLY_STOP_ENABLE: &str = "RUSTFS_GET_METADATA_VERSION_EARLY_STOP_ENABLE";
const DEFAULT_RUSTFS_GET_METADATA_VERSION_EARLY_STOP_ENABLE: bool = false; const DEFAULT_RUSTFS_GET_METADATA_VERSION_EARLY_STOP_ENABLE: bool = false;
const ENV_RUSTFS_GET_METADATA_DATA_READ_EARLY_STOP_ENABLE: &str = "RUSTFS_GET_METADATA_DATA_READ_EARLY_STOP_ENABLE"; const ENV_RUSTFS_GET_METADATA_DATA_READ_EARLY_STOP_ENABLE: &str = "RUSTFS_GET_METADATA_DATA_READ_EARLY_STOP_ENABLE";
const DEFAULT_RUSTFS_GET_METADATA_DATA_READ_EARLY_STOP_ENABLE: bool = false; const DEFAULT_RUSTFS_GET_METADATA_DATA_READ_EARLY_STOP_ENABLE: bool = true;
const ENV_RUSTFS_GET_METADATA_EARLY_STOP_BOUNDED_FANOUT: &str = "RUSTFS_GET_METADATA_EARLY_STOP_BOUNDED_FANOUT"; const ENV_RUSTFS_GET_METADATA_EARLY_STOP_BOUNDED_FANOUT: &str = "RUSTFS_GET_METADATA_EARLY_STOP_BOUNDED_FANOUT";
const DEFAULT_RUSTFS_GET_METADATA_EARLY_STOP_BOUNDED_FANOUT: bool = false; const DEFAULT_RUSTFS_GET_METADATA_EARLY_STOP_BOUNDED_FANOUT: bool = true;
const ENV_RUSTFS_GET_METADATA_SLOWTAIL_FAULT_DELAY_MS: &str = "RUSTFS_GET_METADATA_SLOWTAIL_FAULT_DELAY_MS";
const ENV_RUSTFS_GET_METADATA_SLOWTAIL_FAULT_DISKS: &str = "RUSTFS_GET_METADATA_SLOWTAIL_FAULT_DISKS";
const ENV_RUSTFS_GET_METADATA_SLOWTAIL_FAULT_BUCKET: &str = "RUSTFS_GET_METADATA_SLOWTAIL_FAULT_BUCKET";
const ENV_RUSTFS_GET_METADATA_SLOWTAIL_FAULT_OBJECT_PREFIX: &str = "RUSTFS_GET_METADATA_SLOWTAIL_FAULT_OBJECT_PREFIX";
// --- Multipart Reader-Setup Prefetch Configuration (backlog#870) --- // --- Multipart Reader-Setup Prefetch Configuration (backlog#870) ---
@@ -906,18 +921,16 @@ mod prepared_get_object_metadata_tests {
.expect("test should find an object whose initial fanout covers both data shards") .expect("test should find an object whose initial fanout covers both data shards")
} }
#[allow(
dead_code,
reason = "test fixture no assertion in this module uses today; the live namesake lives in io_primitives tests (backlog#1823)"
)]
fn bounded_spare_disk_index(bucket: &str, object: &str) -> usize { fn bounded_spare_disk_index(bucket: &str, object: &str) -> usize {
*bounded_metadata_fanout_order(bucket, object, 4, 2) *bounded_metadata_fanout_order(bucket, object, 4, 2)
.get(3) .get(3)
.expect("4-disk test geometry should leave one bounded spare disk") .expect("4-disk test geometry should leave one bounded spare disk")
} }
fn bounded_slow_initial_disk_index(bucket: &str, object: &str) -> usize {
*bounded_metadata_fanout_order(bucket, object, 4, 2)
.get(2)
.expect("4-disk test geometry should include a third initial metadata disk")
}
#[tokio::test] #[tokio::test]
async fn prepared_metadata_is_consumed_exactly_once() { async fn prepared_metadata_is_consumed_exactly_once() {
let snapshot = GetObjectFileInfo::owned(FileInfo::default(), Vec::new(), Vec::new()); let snapshot = GetObjectFileInfo::owned(FileInfo::default(), Vec::new(), Vec::new());
@@ -1036,7 +1049,7 @@ mod prepared_get_object_metadata_tests {
#[test] #[test]
#[serial_test::serial(body_cache_hook)] #[serial_test::serial(body_cache_hook)]
fn inline_data_read_early_stop_reader_returns_exact_body() { fn inline_data_read_early_stop_defaults_return_exact_body() {
let runtime = tokio::runtime::Builder::new_current_thread() let runtime = tokio::runtime::Builder::new_current_thread()
.enable_all() .enable_all()
.build() .build()
@@ -1068,14 +1081,14 @@ mod prepared_get_object_metadata_tests {
temp_env::async_with_vars( temp_env::async_with_vars(
[ [
("RUSTFS_GET_METADATA_EARLY_STOP_ENABLE", Some("true")), ("RUSTFS_GET_METADATA_EARLY_STOP_ENABLE", None::<&str>),
("RUSTFS_GET_METADATA_DATA_READ_EARLY_STOP_ENABLE", Some("true")), ("RUSTFS_GET_METADATA_DATA_READ_EARLY_STOP_ENABLE", None::<&str>),
("RUSTFS_GET_METADATA_EARLY_STOP_BOUNDED_FANOUT", Some("true")), ("RUSTFS_GET_METADATA_EARLY_STOP_BOUNDED_FANOUT", None::<&str>),
], ],
async { async {
let slow_initial_disk = bounded_slow_initial_disk_index(bucket, &object); let slow_parity_disk = bounded_spare_disk_index(bucket, &object);
let barrier = let barrier =
rename_fanout_barrier::arm(&object, slow_initial_disk, rename_fanout_barrier::PHASE_READ_VERSION); rename_fanout_barrier::arm(&object, slow_parity_disk, rename_fanout_barrier::PHASE_READ_VERSION);
let calls = disk_call_counters::observe(&object); let calls = disk_call_counters::observe(&object);
let set_disks_for_read = Arc::clone(&set_disks); let set_disks_for_read = Arc::clone(&set_disks);
let opts_for_read = opts.clone(); let opts_for_read = opts.clone();
@@ -1088,10 +1101,10 @@ mod prepared_get_object_metadata_tests {
tokio::time::timeout(READ_VERSION_BARRIER_GUARD, barrier.wait_until_paused()) tokio::time::timeout(READ_VERSION_BARRIER_GUARD, barrier.wait_until_paused())
.await .await
.expect("bounded inline GET should pause a slow initial metadata read"); .expect("default inline GET should pause a slow parity metadata read");
let mut reader = tokio::time::timeout(READ_VERSION_BARRIER_GUARD, &mut open_reader) let mut reader = tokio::time::timeout(READ_VERSION_BARRIER_GUARD, &mut open_reader)
.await .await
.expect("production inline GET should return before the paused metadata response") .expect("default production inline GET should return before the paused parity metadata response")
.expect("inline GET reader task should not panic") .expect("inline GET reader task should not panic")
.expect("inline GET reader should open"); .expect("inline GET reader should open");
let object_size = reader.object_info.size; let object_size = reader.object_info.size;
@@ -1112,14 +1125,17 @@ mod prepared_get_object_metadata_tests {
assert_eq!(object_size, payload.len() as i64); assert_eq!(object_size, payload.len() as i64);
assert_eq!(restored, payload); assert_eq!(restored, payload);
assert_eq!(calls_total, 4, "bounded production GET should schedule the initial quorum plus one spare"); assert_eq!(
calls_total, 4,
"default production inline GET should schedule the initial bounded quorum plus one hedge"
);
assert_eq!( assert_eq!(
recorder.histogram_values( recorder.histogram_values(
"rustfs_io_get_object_metadata_fanout_scheduled", "rustfs_io_get_object_metadata_fanout_scheduled",
&[("path", GET_OBJECT_PATH_LEGACY_DUPLEX)] &[("path", GET_OBJECT_PATH_LEGACY_DUPLEX)]
), ),
vec![4.0], vec![4.0],
"bounded production GET should record all scheduled metadata tasks" "default production GET should record all scheduled metadata tasks"
); );
assert_eq!( assert_eq!(
recorder.histogram_values( recorder.histogram_values(
@@ -1127,7 +1143,7 @@ mod prepared_get_object_metadata_tests {
&[("path", GET_OBJECT_PATH_LEGACY_DUPLEX)] &[("path", GET_OBJECT_PATH_LEGACY_DUPLEX)]
), ),
vec![3.0], vec![3.0],
"bounded production GET should record only observed metadata responses as completed" "default production GET should record only observed metadata responses as completed"
); );
assert_eq!( assert_eq!(
recorder.histogram_values( recorder.histogram_values(
@@ -1135,7 +1151,7 @@ mod prepared_get_object_metadata_tests {
&[("path", GET_OBJECT_PATH_LEGACY_DUPLEX)] &[("path", GET_OBJECT_PATH_LEGACY_DUPLEX)]
), ),
vec![1.0], vec![1.0],
"bounded production GET should record the aborted slow metadata task" "default production GET should record the aborted slow parity metadata task"
); );
} }
@@ -1282,9 +1298,9 @@ mod prepared_get_object_metadata_tests {
temp_env::async_with_vars( temp_env::async_with_vars(
[ [
("RUSTFS_GET_METADATA_EARLY_STOP_ENABLE", Some("true")), ("RUSTFS_GET_METADATA_EARLY_STOP_ENABLE", None::<&str>),
("RUSTFS_GET_METADATA_DATA_READ_EARLY_STOP_ENABLE", Some("true")), ("RUSTFS_GET_METADATA_DATA_READ_EARLY_STOP_ENABLE", None::<&str>),
("RUSTFS_GET_METADATA_EARLY_STOP_BOUNDED_FANOUT", Some("true")), ("RUSTFS_GET_METADATA_EARLY_STOP_BOUNDED_FANOUT", None::<&str>),
], ],
async { async {
let calls = disk_call_counters::observe(&object); let calls = disk_call_counters::observe(&object);
@@ -1686,6 +1702,95 @@ fn is_get_metadata_early_stop_bounded_fanout_enabled() -> bool {
} }
} }
#[derive(Debug)]
struct GetMetadataSlowtailFaultConfig {
delay: Duration,
disks: Arc<[usize]>,
bucket: Option<String>,
object_prefix: Option<String>,
}
#[derive(Clone, Debug)]
struct GetMetadataSlowtailFaultRequest {
delay: Duration,
disks: Arc<[usize]>,
}
impl GetMetadataSlowtailFaultRequest {
fn delay_for_disk(&self, disk_index: usize) -> Option<Duration> {
self.disks.contains(&disk_index).then_some(self.delay)
}
}
fn parse_get_metadata_slowtail_fault_disks(raw: &str) -> Option<Vec<usize>> {
let mut disks = Vec::new();
for item in raw.split(',').map(str::trim).filter(|item| !item.is_empty()) {
let Ok(index) = item.parse::<usize>() else {
return None;
};
if !disks.contains(&index) {
disks.push(index);
}
}
(!disks.is_empty()).then_some(disks)
}
fn load_get_metadata_slowtail_fault_config() -> Option<GetMetadataSlowtailFaultConfig> {
let delay_ms = rustfs_utils::get_env_u64(ENV_RUSTFS_GET_METADATA_SLOWTAIL_FAULT_DELAY_MS, 0);
if delay_ms == 0 {
return None;
}
let disks = parse_get_metadata_slowtail_fault_disks(&std::env::var(ENV_RUSTFS_GET_METADATA_SLOWTAIL_FAULT_DISKS).ok()?)?;
let bucket = std::env::var(ENV_RUSTFS_GET_METADATA_SLOWTAIL_FAULT_BUCKET)
.ok()
.filter(|value| !value.is_empty());
let object_prefix = std::env::var(ENV_RUSTFS_GET_METADATA_SLOWTAIL_FAULT_OBJECT_PREFIX)
.ok()
.filter(|value| !value.is_empty());
Some(GetMetadataSlowtailFaultConfig {
delay: Duration::from_millis(delay_ms),
disks: Arc::from(disks.into_boxed_slice()),
bucket,
object_prefix,
})
}
fn get_metadata_slowtail_fault_request(bucket: &str, object: &str, read_data: bool) -> Option<GetMetadataSlowtailFaultRequest> {
if !read_data {
return None;
}
#[cfg(test)]
let config = load_get_metadata_slowtail_fault_config();
#[cfg(test)]
let config = config.as_ref()?;
#[cfg(not(test))]
let config = ({
static CACHED: OnceLock<Option<GetMetadataSlowtailFaultConfig>> = OnceLock::new();
CACHED.get_or_init(load_get_metadata_slowtail_fault_config).as_ref()
})?;
if let Some(expected_bucket) = &config.bucket
&& expected_bucket != bucket
{
return None;
}
if let Some(expected_prefix) = &config.object_prefix
&& !object.starts_with(expected_prefix)
{
return None;
}
Some(GetMetadataSlowtailFaultRequest {
delay: config.delay,
disks: config.disks.clone(),
})
}
#[cfg(test)]
fn get_metadata_slowtail_fault_delay(bucket: &str, object: &str, disk_index: usize, read_data: bool) -> Option<Duration> {
get_metadata_slowtail_fault_request(bucket, object, read_data)?.delay_for_disk(disk_index)
}
/// Check if multipart reads prefetch the next part's bitrot reader setup /// Check if multipart reads prefetch the next part's bitrot reader setup
/// while the current part decodes (backlog#870). /// while the current part decodes (backlog#870).
/// ///
@@ -1711,6 +1816,10 @@ fn is_multipart_reader_setup_prefetch_enabled() -> bool {
} }
} }
#[allow(
dead_code,
reason = "percentage-rollout facet of the metadata early-stop switch; its predicate has no caller while the sibling enable flag is live (backlog#1823)"
)]
fn get_metadata_early_stop_rollout_pct() -> u32 { fn get_metadata_early_stop_rollout_pct() -> u32 {
static CACHED: OnceLock<u32> = OnceLock::new(); static CACHED: OnceLock<u32> = OnceLock::new();
*CACHED.get_or_init(|| { *CACHED.get_or_init(|| {
@@ -1750,6 +1859,10 @@ fn should_use_codec_streaming(config: GetCodecStreamingConfig, bucket: &str, obj
} }
/// Should this specific request use metadata early-stop? /// Should this specific request use metadata early-stop?
#[allow(
dead_code,
reason = "percentage-rollout facet of the metadata early-stop switch; its predicate has no caller while the sibling enable flag is live (backlog#1823)"
)]
pub fn should_use_metadata_early_stop(bucket: &str, object: &str) -> bool { pub fn should_use_metadata_early_stop(bucket: &str, object: &str) -> bool {
let base = is_get_metadata_early_stop_enabled(); let base = is_get_metadata_early_stop_enabled();
let pct = get_metadata_early_stop_rollout_pct(); let pct = get_metadata_early_stop_rollout_pct();
@@ -2183,6 +2296,7 @@ fn classify_get_codec_streaming_object_class(
GetCodecStreamingObjectClass::PlainSinglePart GetCodecStreamingObjectClass::PlainSinglePart
} }
#[allow(dead_code, reason = "asserted by this file's tests (backlog#1823)")]
fn is_get_small_object_direct_memory_eligible_with_threshold( fn is_get_small_object_direct_memory_eligible_with_threshold(
range: &Option<HTTPRangeSpec>, range: &Option<HTTPRangeSpec>,
object_info: &ObjectInfo, object_info: &ObjectInfo,
@@ -2788,6 +2902,7 @@ pub struct SetDisks {
/// Stable namespace shared by every object lock created for this set. /// Stable namespace shared by every object lock created for this set.
set_lock_namespace: Arc<str>, set_lock_namespace: Arc<str>,
pub format: FormatV3, pub format: FormatV3,
#[allow(dead_code, reason = "asserted by this file's tests (backlog#1823)")]
disk_health_cache: Arc<RwLock<Vec<Option<DiskHealthEntry>>>>, disk_health_cache: Arc<RwLock<Vec<Option<DiskHealthEntry>>>>,
get_object_metadata_cache: moka::future::Cache<GetObjectMetadataCacheKey, Arc<GetObjectMetadataCacheEntry>>, get_object_metadata_cache: moka::future::Cache<GetObjectMetadataCacheKey, Arc<GetObjectMetadataCacheEntry>>,
get_object_metadata_cache_hash_builder: std::collections::hash_map::RandomState, get_object_metadata_cache_hash_builder: std::collections::hash_map::RandomState,
@@ -3063,11 +3178,13 @@ struct GetObjectMetadataCacheEntry {
#[derive(Clone, Debug)] #[derive(Clone, Debug)]
struct DiskHealthEntry { struct DiskHealthEntry {
#[allow(dead_code, reason = "asserted by this file's tests (backlog#1823)")]
last_check: Instant, last_check: Instant,
online: bool, online: bool,
} }
impl DiskHealthEntry { impl DiskHealthEntry {
#[allow(dead_code, reason = "asserted by this file's tests (backlog#1823)")]
fn cached_value(&self) -> Option<bool> { fn cached_value(&self) -> Option<bool> {
if self.last_check.elapsed() <= DISK_HEALTH_CACHE_TTL { if self.last_check.elapsed() <= DISK_HEALTH_CACHE_TTL {
Some(self.online) Some(self.online)
@@ -3661,6 +3778,7 @@ fn multipart_put_large_batch_min_size_bytes() -> usize {
}) })
} }
#[allow(dead_code, reason = "asserted by this file's tests (backlog#1823)")]
fn classify_small_write_path(is_inline_buffer: bool, object_size: i64, block_size: usize) -> SmallWritePath { fn classify_small_write_path(is_inline_buffer: bool, object_size: i64, block_size: usize) -> SmallWritePath {
if should_use_inline_small_fast_path(is_inline_buffer, object_size, block_size) { if should_use_inline_small_fast_path(is_inline_buffer, object_size, block_size) {
SmallWritePath::Inline SmallWritePath::Inline
@@ -3866,8 +3984,17 @@ fn collect_inline_data_shard_fileinfos_by_index<'a>(
parts_metadata: &'a [FileInfo], parts_metadata: &'a [FileInfo],
fi: &FileInfo, fi: &FileInfo,
data_shards: usize, data_shards: usize,
mut disk_is_online: impl FnMut(usize) -> bool, disk_is_online: impl FnMut(usize) -> bool,
) -> Option<Vec<&'a FileInfo>> { ) -> Option<Vec<&'a FileInfo>> {
collect_inline_data_shard_fileinfos_by_index_or_reason(parts_metadata, fi, data_shards, disk_is_online).ok()
}
fn collect_inline_data_shard_fileinfos_by_index_or_reason<'a>(
parts_metadata: &'a [FileInfo],
fi: &FileInfo,
data_shards: usize,
mut disk_is_online: impl FnMut(usize) -> bool,
) -> std::result::Result<Vec<&'a FileInfo>, &'static str> {
let distribution = &fi.erasure.distribution; let distribution = &fi.erasure.distribution;
let mut data_files = vec![None; data_shards]; let mut data_files = vec![None; data_shards];
@@ -3875,27 +4002,35 @@ fn collect_inline_data_shard_fileinfos_by_index<'a>(
if !disk_is_online(disk_index) { if !disk_is_online(disk_index) {
continue; continue;
} }
let block_index = *distribution.get(disk_index)?; let Some(&block_index) = distribution.get(disk_index) else {
return Err(GET_METADATA_EARLY_STOP_REASON_DATA_READ_INLINE_GEOMETRY);
};
if block_index == 0 || block_index > data_shards { if block_index == 0 || block_index > data_shards {
continue; continue;
} }
if file_info.name.is_empty() {
return Err(GET_METADATA_EARLY_STOP_REASON_DATA_READ_INLINE_MISSING_SHARD);
}
if file_info.erasure.index != block_index { if file_info.erasure.index != block_index {
continue; return Err(GET_METADATA_EARLY_STOP_REASON_DATA_READ_INLINE_IDENTITY_MISMATCH);
} }
if !file_info.has_valid_erasure_geometry() { if !file_info.has_valid_erasure_geometry() {
continue; return Err(GET_METADATA_EARLY_STOP_REASON_DATA_READ_INLINE_GEOMETRY);
} }
if !core::io_primitives::metadata_early_stop_candidate_matches(file_info, fi) { if !core::io_primitives::metadata_early_stop_candidate_matches(file_info, fi) {
continue; return Err(GET_METADATA_EARLY_STOP_REASON_DATA_READ_INLINE_IDENTITY_MISMATCH);
} }
if file_info.data.as_ref().is_none_or(|data| data.is_empty()) { if file_info.data.as_ref().is_none_or(|data| data.is_empty()) {
continue; return Err(GET_METADATA_EARLY_STOP_REASON_DATA_READ_INLINE_MISSING_PAYLOAD);
} }
data_files[block_index - 1] = Some(file_info); data_files[block_index - 1] = Some(file_info);
} }
data_files.into_iter().collect() data_files
.into_iter()
.collect::<Option<Vec<_>>>()
.ok_or(GET_METADATA_EARLY_STOP_REASON_DATA_READ_INLINE_MISSING_SHARD)
} }
impl SetDisks { impl SetDisks {
@@ -4222,6 +4357,7 @@ fn check_object_lock_retention_update(bucket: &str, object: &str, obj_info: &Obj
/// ///
/// Fail closed: when bucket metadata cannot be resolved the check stays on, so /// Fail closed: when bucket metadata cannot be resolved the check stays on, so
/// object-lock protection is never skipped because of a metadata lookup miss. /// object-lock protection is never skipped because of a metadata lookup miss.
#[allow(dead_code, reason = "asserted by this file's tests (backlog#1823)")]
pub(crate) fn object_lock_delete_check_required(bucket_meta: Option<&crate::bucket::metadata::BucketMetadata>) -> bool { pub(crate) fn object_lock_delete_check_required(bucket_meta: Option<&crate::bucket::metadata::BucketMetadata>) -> bool {
bucket_meta.is_none_or(|meta| meta.object_locking()) bucket_meta.is_none_or(|meta| meta.object_locking())
} }
@@ -4497,15 +4633,6 @@ impl Hash for ObjProps {
} }
} }
#[derive(Default, Clone, Debug)]
pub struct HealEntryResult {
pub bytes: usize,
pub success: bool,
pub skipped: bool,
pub entry_done: bool,
pub name: String,
}
fn is_object_dangling( fn is_object_dangling(
meta_arr: &[FileInfo], meta_arr: &[FileInfo],
errs: &[Option<DiskError>], errs: &[Option<DiskError>],
@@ -5282,6 +5409,7 @@ pub fn is_valid_storage_class(storage_class: &str) -> bool {
} }
/// Returns true if the storage class is a cold storage tier that requires special handling /// Returns true if the storage class is a cold storage tier that requires special handling
#[allow(dead_code, reason = "asserted by this file's tests (backlog#1823)")]
pub fn is_cold_storage_class(storage_class: &str) -> bool { pub fn is_cold_storage_class(storage_class: &str) -> bool {
matches!( matches!(
storage_class, storage_class,
@@ -5290,6 +5418,7 @@ pub fn is_cold_storage_class(storage_class: &str) -> bool {
} }
/// Returns true if the storage class is an infrequent access tier /// Returns true if the storage class is an infrequent access tier
#[allow(dead_code, reason = "asserted by this file's tests (backlog#1823)")]
pub fn is_infrequent_access_class(storage_class: &str) -> bool { pub fn is_infrequent_access_class(storage_class: &str) -> bool {
matches!( matches!(
storage_class, storage_class,
+4
View File
@@ -1716,6 +1716,10 @@ impl SetDisks {
Ok((result, None)) Ok((result, None))
} }
#[allow(
dead_code,
reason = "lock-taking wrapper over the live heal_object_dir_locked; only comments reference it (backlog#1823)"
)]
#[tracing::instrument(level = "trace", skip(self), fields(bucket = %bucket, object = %object))] #[tracing::instrument(level = "trace", skip(self), fields(bucket = %bucket, object = %object))]
pub(in crate::set_disk) async fn heal_object_dir( pub(in crate::set_disk) async fn heal_object_dir(
&self, &self,
@@ -66,6 +66,7 @@ impl crate::storage_api_contracts::namespace::NamespaceLocking for SetDisks {
} }
impl SetDisks { impl SetDisks {
#[allow(dead_code, reason = "lock diagnostics formatter with no caller in this port (backlog#1823)")]
pub(in crate::set_disk) fn format_lock_error(&self, bucket: &str, object: &str, mode: &str, err: &LockResult) -> String { pub(in crate::set_disk) fn format_lock_error(&self, bucket: &str, object: &str, mode: &str, err: &LockResult) -> String {
match err { match err {
LockResult::Timeout => { LockResult::Timeout => {
@@ -79,6 +80,7 @@ impl SetDisks {
} }
} }
#[allow(dead_code, reason = "lock diagnostics formatter with no caller in this port (backlog#1823)")]
pub(in crate::set_disk) fn format_lock_error_from_error( pub(in crate::set_disk) fn format_lock_error_from_error(
&self, &self,
bucket: &str, bucket: &str,
@@ -143,6 +145,7 @@ impl SetDisks {
disks disks
} }
#[allow(dead_code, reason = "asserted by this file's tests (backlog#1823)")]
pub(in crate::set_disk) async fn get_online_disks(&self) -> Vec<Option<DiskStore>> { pub(in crate::set_disk) async fn get_online_disks(&self) -> Vec<Option<DiskStore>> {
let snapshot = self.drive_membership_snapshot().await; let snapshot = self.drive_membership_snapshot().await;
let mut disks = snapshot.strict_online_candidates().into_iter().map(Some).collect::<Vec<_>>(); let mut disks = snapshot.strict_online_candidates().into_iter().map(Some).collect::<Vec<_>>();
@@ -153,6 +156,10 @@ impl SetDisks {
disks disks
} }
#[allow(
dead_code,
reason = "local-only sibling of the test-covered get_online_disks; no caller in this port (backlog#1823)"
)]
pub(in crate::set_disk) async fn get_online_local_disks(&self) -> Vec<Option<DiskStore>> { pub(in crate::set_disk) async fn get_online_local_disks(&self) -> Vec<Option<DiskStore>> {
let snapshot = self.drive_membership_snapshot().await; let snapshot = self.drive_membership_snapshot().await;
let mut disks = snapshot let mut disks = snapshot
@@ -432,6 +439,10 @@ impl SetDisks {
Ok((disk, fm)) Ok((disk, fm))
} }
#[allow(
dead_code,
reason = "MinIO-parity healing-disk accessor with no caller in this port (backlog#1823)"
)]
pub(in crate::set_disk) async fn get_online_disk_with_healing( pub(in crate::set_disk) async fn get_online_disk_with_healing(
&self, &self,
incl_healing: bool, incl_healing: bool,
@@ -440,6 +451,10 @@ impl SetDisks {
Ok((new_disks, healing > 0)) Ok((new_disks, healing > 0))
} }
#[allow(
dead_code,
reason = "reached only from get_online_disk_with_healing, itself uncalled in this port (backlog#1823)"
)]
pub(in crate::set_disk) async fn get_online_disk_with_healing_and_info( pub(in crate::set_disk) async fn get_online_disk_with_healing_and_info(
&self, &self,
incl_healing: bool, incl_healing: bool,
@@ -415,6 +415,7 @@ fn reduce_quorum_part_numbers(object_parts: Vec<Vec<String>>, read_quorum: usize
/// never returned, but flips `is_truncated` to `true` and yields a /// never returned, but flips `is_truncated` to `true` and yields a
/// `next_upload_id_marker` pointing at the last returned upload so the caller can /// `next_upload_id_marker` pointing at the last returned upload so the caller can
/// resume paging. /// resume paging.
#[allow(dead_code, reason = "asserted by this file's tests (backlog#1823)")]
fn paginate_upload_page(remaining: &[MultipartInfo], max_uploads: usize) -> (Vec<MultipartInfo>, bool, Option<String>) { fn paginate_upload_page(remaining: &[MultipartInfo], max_uploads: usize) -> (Vec<MultipartInfo>, bool, Option<String>) {
let is_truncated = remaining.len() > max_uploads; let is_truncated = remaining.len() > max_uploads;
let page: Vec<MultipartInfo> = remaining.iter().take(max_uploads).cloned().collect(); let page: Vec<MultipartInfo> = remaining.iter().take(max_uploads).cloned().collect();
@@ -557,6 +558,7 @@ impl SetDisks {
} }
#[tracing::instrument(level = "debug", skip(self))] #[tracing::instrument(level = "debug", skip(self))]
#[allow(dead_code, reason = "asserted by this file's tests (backlog#1823)")]
pub(super) async fn check_upload_id_exists( pub(super) async fn check_upload_id_exists(
&self, &self,
bucket: &str, bucket: &str,
+329 -11
View File
@@ -3507,6 +3507,10 @@ struct TransitionUploadedSaveProbeState {
} }
#[cfg(test)] #[cfg(test)]
#[allow(
dead_code,
reason = "installed by set_disk tests behind `--features test-util` (backlog#1823)"
)]
struct TransitionUploadedSaveProbe { struct TransitionUploadedSaveProbe {
state: Arc<TransitionUploadedSaveProbeState>, state: Arc<TransitionUploadedSaveProbeState>,
} }
@@ -3517,6 +3521,10 @@ static TRANSITION_UPLOADED_SAVE_PROBE: std::sync::OnceLock<std::sync::Mutex<Opti
#[cfg(test)] #[cfg(test)]
impl TransitionUploadedSaveProbe { impl TransitionUploadedSaveProbe {
#[allow(
dead_code,
reason = "installed by set_disk tests behind `--features test-util` (backlog#1823)"
)]
fn install(bucket: &str, object: &str) -> Self { fn install(bucket: &str, object: &str) -> Self {
let state = Arc::new(TransitionUploadedSaveProbeState { let state = Arc::new(TransitionUploadedSaveProbeState {
bucket: bucket.to_string(), bucket: bucket.to_string(),
@@ -3533,6 +3541,10 @@ impl TransitionUploadedSaveProbe {
Self { state } Self { state }
} }
#[allow(
dead_code,
reason = "installed by set_disk tests behind `--features test-util` (backlog#1823)"
)]
fn attempts(&self) -> usize { fn attempts(&self) -> usize {
self.state.attempts.load(std::sync::atomic::Ordering::Acquire) self.state.attempts.load(std::sync::atomic::Ordering::Acquire)
} }
@@ -3738,6 +3750,10 @@ struct TransitionCommitBarrierState {
} }
#[cfg(test)] #[cfg(test)]
#[allow(
dead_code,
reason = "installed by set_disk tests behind `--features test-util` (backlog#1823)"
)]
struct TransitionCommitBarrier { struct TransitionCommitBarrier {
state: Arc<TransitionCommitBarrierState>, state: Arc<TransitionCommitBarrierState>,
} }
@@ -3748,14 +3764,26 @@ static TRANSITION_COMMIT_BARRIER: std::sync::OnceLock<std::sync::Mutex<Option<Ar
#[cfg(test)] #[cfg(test)]
impl TransitionCommitBarrier { impl TransitionCommitBarrier {
#[allow(
dead_code,
reason = "installed by set_disk tests behind `--features test-util` (backlog#1823)"
)]
fn install_before_lock_lost_check(bucket: &str, object: &str) -> Self { fn install_before_lock_lost_check(bucket: &str, object: &str) -> Self {
Self::install_at(bucket, object, TransitionCommitPause::BeforeLockLost) Self::install_at(bucket, object, TransitionCommitPause::BeforeLockLost)
} }
#[allow(
dead_code,
reason = "installed by set_disk tests behind `--features test-util` (backlog#1823)"
)]
fn install(bucket: &str, object: &str) -> Self { fn install(bucket: &str, object: &str) -> Self {
Self::install_at(bucket, object, TransitionCommitPause::BeforeLeaseValidation) Self::install_at(bucket, object, TransitionCommitPause::BeforeLeaseValidation)
} }
#[allow(
dead_code,
reason = "installed by set_disk tests behind `--features test-util` (backlog#1823)"
)]
fn install_after_lease_check(bucket: &str, object: &str) -> Self { fn install_after_lease_check(bucket: &str, object: &str) -> Self {
Self::install_at(bucket, object, TransitionCommitPause::AfterLeaseValidation) Self::install_at(bucket, object, TransitionCommitPause::AfterLeaseValidation)
} }
@@ -3778,12 +3806,20 @@ impl TransitionCommitBarrier {
Self { state } Self { state }
} }
#[allow(
dead_code,
reason = "installed by set_disk tests behind `--features test-util` (backlog#1823)"
)]
async fn wait_until_paused(&self) { async fn wait_until_paused(&self) {
tokio::time::timeout(Duration::from_secs(30), self.state.arrived.notified()) tokio::time::timeout(Duration::from_secs(30), self.state.arrived.notified())
.await .await
.expect("transition should reach the deterministic commit barrier"); .expect("transition should reach the deterministic commit barrier");
} }
#[allow(
dead_code,
reason = "installed by set_disk tests behind `--features test-util` (backlog#1823)"
)]
fn release(&self) { fn release(&self) {
self.state.release.notify_one(); self.state.release.notify_one();
} }
@@ -7505,7 +7541,7 @@ mod get_object_downstream_close_accounting_tests {
use super::hermetic_set_disks_support::hermetic_set_disks; use super::hermetic_set_disks_support::hermetic_set_disks;
use super::*; use super::*;
use crate::diagnostics::get::{ use crate::diagnostics::get::{
GET_METADATA_EARLY_STOP_REASON_UNSAFE_REQUEST, GET_OBJECT_PATH_INTERNAL_META, GET_STAGE_DECODE, GET_STAGE_EMIT, GET_METADATA_EARLY_STOP_REASON_NOT_FOUND, GET_OBJECT_PATH_INTERNAL_META, GET_STAGE_DECODE, GET_STAGE_EMIT,
GetObjectFailureReason, GetObjectFailureReason,
}; };
use crate::disk::RUSTFS_META_BUCKET; use crate::disk::RUSTFS_META_BUCKET;
@@ -7637,8 +7673,8 @@ mod get_object_downstream_close_accounting_tests {
legacy_completed, legacy_completed,
internal_cancelled, internal_cancelled,
legacy_cancelled, legacy_cancelled,
internal_unsafe_miss, internal_not_found_miss,
legacy_unsafe_miss, legacy_not_found_miss,
internal_saved, internal_saved,
legacy_saved, legacy_saved,
) = metrics::with_local_recorder(&recorder, || { ) = metrics::with_local_recorder(&recorder, || {
@@ -7714,7 +7750,7 @@ mod get_object_downstream_close_accounting_tests {
&[ &[
("path", GET_OBJECT_PATH_INTERNAL_META), ("path", GET_OBJECT_PATH_INTERNAL_META),
("decision", "miss"), ("decision", "miss"),
("reason", GET_METADATA_EARLY_STOP_REASON_UNSAFE_REQUEST), ("reason", GET_METADATA_EARLY_STOP_REASON_NOT_FOUND),
], ],
), ),
recorder.counter_value( recorder.counter_value(
@@ -7722,7 +7758,7 @@ mod get_object_downstream_close_accounting_tests {
&[ &[
("path", GET_OBJECT_PATH_LEGACY_DUPLEX), ("path", GET_OBJECT_PATH_LEGACY_DUPLEX),
("decision", "miss"), ("decision", "miss"),
("reason", GET_METADATA_EARLY_STOP_REASON_UNSAFE_REQUEST), ("reason", GET_METADATA_EARLY_STOP_REASON_NOT_FOUND),
], ],
), ),
recorder.histogram_values( recorder.histogram_values(
@@ -7773,21 +7809,21 @@ mod get_object_downstream_close_accounting_tests {
"internal metadata lifecycle cancelled count must not leak into legacy_duplex" "internal metadata lifecycle cancelled count must not leak into legacy_duplex"
); );
assert_eq!( assert_eq!(
internal_unsafe_miss, 1, internal_not_found_miss, 1,
"internal metadata unsafe early-stop miss must retain its path label" "internal metadata not-found early-stop miss must retain its path label"
); );
assert_eq!( assert_eq!(
legacy_unsafe_miss, 0, legacy_not_found_miss, 0,
"internal metadata unsafe early-stop miss must not leak into legacy_duplex" "internal metadata not-found early-stop miss must not leak into legacy_duplex"
); );
assert_eq!( assert_eq!(
internal_saved, internal_saved,
vec![0.0], vec![0.0],
"internal metadata unsafe miss must record zero saved responses on internal_meta" "internal metadata not-found miss must record zero saved responses on internal_meta"
); );
assert!( assert!(
legacy_saved.is_empty(), legacy_saved.is_empty(),
"internal metadata unsafe miss saved responses must not leak into legacy_duplex" "internal metadata not-found miss saved responses must not leak into legacy_duplex"
); );
} }
} }
@@ -10170,6 +10206,288 @@ mod transition_upload_integrity_tests {
assert!(backend.contains(remote_object).await, "committed remote object should remain available"); assert!(backend.contains(remote_object).await, "committed remote object should remain available");
} }
/// Compresses `plaintext` with the codec the PUT path uses, so the stored
/// bytes round-trip through the read path's decompressor.
async fn compress_for_storage(plaintext: &[u8]) -> Vec<u8> {
let mut reader = crate::io_support::rio::compression_reader(
Cursor::new(plaintext.to_vec()),
rustfs_utils::CompressionAlgorithm::default(),
false,
);
let mut compressed = Vec::new();
reader.read_to_end(&mut compressed).await.expect("plaintext should compress");
assert!(compressed.len() < plaintext.len(), "test payload must actually compress");
compressed
}
/// Writes a genuinely compressed object: stored data is `compressed`, and the
/// metadata marks it compressed with the plaintext length as its actual size,
/// exactly as the app-layer compress path records it.
async fn write_compressed_source(
set_disks: &Arc<SetDisks>,
disk_stores: &[DiskStore],
bucket: &str,
object: &str,
plaintext: &[u8],
compressed: &[u8],
) -> ObjectInfo {
for disk in disk_stores {
disk.make_volume(bucket).await.expect("bucket volume should be created");
}
let mut user_defined = HashMap::new();
rustfs_utils::http::insert_str(
&mut user_defined,
rustfs_utils::http::SUFFIX_COMPRESSION,
crate::io_support::rio::compression_metadata_value(rustfs_utils::CompressionAlgorithm::default()),
);
rustfs_utils::http::insert_str(&mut user_defined, rustfs_utils::http::SUFFIX_ACTUAL_SIZE, plaintext.len().to_string());
let stream = crate::io_support::rio::HashReader::from_stream(
Cursor::new(compressed.to_vec()),
compressed.len() as i64,
plaintext.len() as i64,
None,
None,
false,
)
.expect("hash reader over compressed bytes");
let mut reader = PutObjReader::new(stream);
set_disks
.put_object(
bucket,
object,
&mut reader,
&ObjectOptions {
no_lock: true,
user_defined,
..Default::default()
},
)
.await
.expect("compressed object should be written")
}
async fn read_transitioned(
set_disks: &Arc<SetDisks>,
bucket: &str,
object: &str,
range: Option<HTTPRangeSpec>,
opts: &ObjectOptions,
) -> (Vec<u8>, i64) {
let mut reader = set_disks
.get_object_reader(bucket, object, range, HeaderMap::new(), opts)
.await
.expect("transitioned object reader should open");
let published_size = reader.object_info.size;
let mut body = Vec::new();
reader
.stream
.read_to_end(&mut body)
.await
.expect("transitioned body should drain");
(body, published_size)
}
/// Transition uploads the object's STORED bytes, so a tiered read has to
/// apply the same transform an erasure read would. #6107 routed this path
/// through `ReadPlan` to stop serving an encrypted object's ciphertext;
/// compression rides the same plan, and nothing pinned it (backlog#1851).
/// Without the transform this GET returns the compressed bytes under the
/// compressed size — silent corruption for every client of a compressed
/// object that ILM has moved to a warm tier.
#[tokio::test]
#[serial_test::serial]
async fn transitioned_compressed_object_get_returns_plaintext() {
let (_temp_dirs, disk_stores, set_disks) = hermetic_set_disks(4).await;
let bucket = "transitioned-compressed-get-bucket";
let object = "object.txt";
let plaintext = b"transitioned compressed objects must decompress on read ".repeat(20_000);
let compressed = compress_for_storage(&plaintext).await;
let original = write_compressed_source(&set_disks, &disk_stores, bucket, object, &plaintext, &compressed).await;
let opts = ObjectOptions {
no_lock: true,
..Default::default()
};
let (local_body, local_size) = read_transitioned(&set_disks, bucket, object, None, &opts).await;
assert_eq!(local_body, plaintext, "control: the pre-transition read must decompress");
assert_eq!(
local_size,
plaintext.len() as i64,
"control: the pre-transition read publishes the plaintext size"
);
let tier_name = format!("COLDTIER{}", &Uuid::new_v4().simple().to_string()[..8]).to_uppercase();
let backend = register_mock_tier(&runtime_sources::global_tier_config_mgr(), &tier_name).await;
set_disks
.transition_object(bucket, object, &transition_options(&original, tier_name))
.await
.expect("transition should commit");
let put_versions = backend.put_versions().await;
assert_eq!(put_versions.len(), 1, "transition should upload one remote candidate");
let remote_bytes = backend
.bytes(&put_versions[0].0)
.await
.expect("remote candidate should be stored");
assert_eq!(
remote_bytes, compressed,
"transition uploads the stored representation; the read side is what has to decode it"
);
let (body, published_size) = read_transitioned(&set_disks, bucket, object, None, &opts).await;
assert_eq!(body, plaintext, "a tiered read must return the object's content, not its stored bytes");
assert_eq!(
published_size,
plaintext.len() as i64,
"a tiered read must publish the plaintext size, not the compressed one"
);
}
/// A ranged tiered read is expressed in plaintext coordinates, so the plan
/// has to translate it into the remote copy's compressed extent and skip
/// into the decompressed stream — the same translation the erasure path does.
#[tokio::test]
#[serial_test::serial]
async fn transitioned_compressed_object_range_get_returns_plaintext_slice() {
let (_temp_dirs, disk_stores, set_disks) = hermetic_set_disks(4).await;
let bucket = "transitioned-compressed-range-bucket";
let object = "object.txt";
let plaintext = b"ranged reads of transitioned compressed objects must land in plaintext ".repeat(20_000);
let compressed = compress_for_storage(&plaintext).await;
let original = write_compressed_source(&set_disks, &disk_stores, bucket, object, &plaintext, &compressed).await;
let tier_name = format!("COLDTIER{}", &Uuid::new_v4().simple().to_string()[..8]).to_uppercase();
register_mock_tier(&runtime_sources::global_tier_config_mgr(), &tier_name).await;
set_disks
.transition_object(bucket, object, &transition_options(&original, tier_name))
.await
.expect("transition should commit");
let opts = ObjectOptions {
no_lock: true,
..Default::default()
};
// Deliberately past the compressed size, so a range still measured in
// stored coordinates could not produce this slice.
let start = compressed.len() as i64 + 4096;
let end = start + 511;
let range = HTTPRangeSpec {
is_suffix_length: false,
start,
end,
};
let (body, published_size) = read_transitioned(&set_disks, bucket, object, Some(range), &opts).await;
let expected = &plaintext[start as usize..=end as usize];
assert_eq!(body, expected, "a ranged tiered read must return that plaintext slice");
assert_eq!(published_size, expected.len() as i64, "a ranged tiered read publishes the slice length");
}
/// The restore copy-back re-writes the object under its original metadata,
/// which still says "compressed". It therefore has to keep receiving the
/// STORED bytes: `restore_request_active` holds it on the plan's `Plain`
/// branch, and decompressing there would write plaintext under compressed
/// metadata.
#[tokio::test]
#[serial_test::serial]
async fn restore_read_of_transitioned_compressed_object_keeps_stored_bytes() {
let (_temp_dirs, disk_stores, set_disks) = hermetic_set_disks(4).await;
let bucket = "transitioned-compressed-restore-bucket";
let object = "object.txt";
let plaintext = b"restore copy-back must keep the stored representation intact ".repeat(20_000);
let compressed = compress_for_storage(&plaintext).await;
let original = write_compressed_source(&set_disks, &disk_stores, bucket, object, &plaintext, &compressed).await;
let tier_name = format!("COLDTIER{}", &Uuid::new_v4().simple().to_string()[..8]).to_uppercase();
register_mock_tier(&runtime_sources::global_tier_config_mgr(), &tier_name).await;
set_disks
.transition_object(bucket, object, &transition_options(&original, tier_name))
.await
.expect("transition should commit");
let oi = set_disks
.get_object_info(
bucket,
object,
&ObjectOptions {
no_lock: true,
..Default::default()
},
)
.await
.expect("transitioned metadata should resolve");
let restore_opts = ObjectOptions {
no_lock: true,
part_number: Some(1),
transition: TransitionOptions {
restore_request: s3s::dto::RestoreRequest {
days: Some(1),
..Default::default()
},
..Default::default()
},
..Default::default()
};
let mut reader = get_transitioned_object_reader_with_tier_manager(
bucket,
object,
&None,
&HeaderMap::new(),
&oi,
&restore_opts,
&set_disks.ctx.tier_config_mgr(),
set_disks.ctx.object_encryption_resolver(),
)
.await
.expect("restore read of the tiered copy should open");
let published_size = reader.object_info.size;
let mut body = Vec::new();
reader.stream.read_to_end(&mut body).await.expect("restore body should drain");
assert_eq!(body, compressed, "a restore read must copy the stored bytes back verbatim");
assert_eq!(
published_size,
compressed.len() as i64,
"a restore read must keep publishing the stored size"
);
}
/// Plain objects must keep streaming the remote bytes through untouched:
/// their plan is `Plain`, so the tiered read stays byte-identical.
#[tokio::test]
#[serial_test::serial]
async fn transitioned_plain_object_get_is_unchanged() {
let (_temp_dirs, disk_stores, set_disks) = hermetic_set_disks(4).await;
let bucket = "transitioned-plain-get-bucket";
let object = "object.bin";
let payload = b"plain transitioned objects must keep reading back byte-identical ".repeat(1024);
let original = write_source(&set_disks, &disk_stores, bucket, object, &payload).await;
let tier_name = format!("COLDTIER{}", &Uuid::new_v4().simple().to_string()[..8]).to_uppercase();
register_mock_tier(&runtime_sources::global_tier_config_mgr(), &tier_name).await;
set_disks
.transition_object(bucket, object, &transition_options(&original, tier_name))
.await
.expect("transition should commit");
let opts = ObjectOptions {
no_lock: true,
..Default::default()
};
let (body, published_size) = read_transitioned(&set_disks, bucket, object, None, &opts).await;
assert_eq!(body, payload);
assert_eq!(published_size, payload.len() as i64);
let range = HTTPRangeSpec {
is_suffix_length: false,
start: 100,
end: 611,
};
let (ranged_body, ranged_size) = read_transitioned(&set_disks, bucket, object, Some(range), &opts).await;
assert_eq!(ranged_body, &payload[100..=611]);
assert_eq!(ranged_size, payload.len() as i64, "a plain ranged read keeps publishing the object size");
}
async fn corrupt_beyond_read_quorum( async fn corrupt_beyond_read_quorum(
temp_dirs: &[tempfile::TempDir], temp_dirs: &[tempfile::TempDir],
bucket: &str, bucket: &str,
+35 -3
View File
@@ -116,6 +116,7 @@ impl SetDisks {
.then_some(GET_METADATA_CACHE_REASON_DIST_ERASURE) .then_some(GET_METADATA_CACHE_REASON_DIST_ERASURE)
} }
#[allow(dead_code, reason = "asserted by this file's tests (backlog#1823)")]
async fn cached_get_object_fileinfo(&self, bucket: &str, object: &str) -> Option<Arc<GetObjectMetadataCacheEntry>> { async fn cached_get_object_fileinfo(&self, bucket: &str, object: &str) -> Option<Arc<GetObjectMetadataCacheEntry>> {
match self.lookup_cached_get_object_fileinfo(bucket, object).await { match self.lookup_cached_get_object_fileinfo(bucket, object).await {
MetadataCacheLookup::Hit(entry) => Some(entry), MetadataCacheLookup::Hit(entry) => Some(entry),
@@ -1826,6 +1827,7 @@ fn get_object_metadata_cache_request_bypass_reason(bucket: &str, opts: &ObjectOp
.then_some(GET_METADATA_CACHE_REASON_META_BUCKET) .then_some(GET_METADATA_CACHE_REASON_META_BUCKET)
} }
#[allow(dead_code, reason = "asserted by this file's tests (backlog#1823)")]
fn is_get_object_metadata_cache_request_eligible(bucket: &str, opts: &ObjectOptions, read_data: bool) -> bool { fn is_get_object_metadata_cache_request_eligible(bucket: &str, opts: &ObjectOptions, read_data: bool) -> bool {
get_object_metadata_cache_request_bypass_reason(bucket, opts, read_data).is_none() get_object_metadata_cache_request_bypass_reason(bucket, opts, read_data).is_none()
} }
@@ -3886,13 +3888,15 @@ mod tests {
assert!(metadata_early_stop_permitted(true, true, false, "", false, false)); assert!(metadata_early_stop_permitted(true, true, false, "", false, false));
// observe=false (non-observed fanout) also disables early-stop. // observe=false (non-observed fanout) also disables early-stop.
assert!(!metadata_early_stop_permitted(true, false, false, "", false, false)); assert!(!metadata_early_stop_permitted(true, false, false, "", false, false));
assert!(!metadata_early_stop_permitted(true, true, true, "", false, false)); // Whole/latest data-read metadata is now allowed by default;
// the inline verifier still decides whether it can stop early.
assert!(metadata_early_stop_permitted(true, true, true, "", false, false));
}, },
); );
} }
#[test] #[test]
fn metadata_early_stop_keeps_data_reads_opt_in_by_default() { fn metadata_early_stop_allows_safe_data_reads_by_default() {
temp_env::with_vars( temp_env::with_vars(
[ [
(ENV_RUSTFS_GET_METADATA_EARLY_STOP_ENABLE, Some("true")), (ENV_RUSTFS_GET_METADATA_EARLY_STOP_ENABLE, Some("true")),
@@ -3900,7 +3904,7 @@ mod tests {
(ENV_RUSTFS_GET_METADATA_DATA_READ_EARLY_STOP_ENABLE, None), (ENV_RUSTFS_GET_METADATA_DATA_READ_EARLY_STOP_ENABLE, None),
], ],
|| { || {
assert!(!should_allow_metadata_early_stop(true, "", false, false)); assert!(should_allow_metadata_early_stop(true, "", false, false));
assert!(!should_allow_metadata_early_stop(true, "version-id", false, false)); assert!(!should_allow_metadata_early_stop(true, "version-id", false, false));
assert!(should_allow_metadata_early_stop(false, "", false, false)); assert!(should_allow_metadata_early_stop(false, "", false, false));
assert!(!should_allow_metadata_early_stop(false, "version-id", false, false)); assert!(!should_allow_metadata_early_stop(false, "version-id", false, false));
@@ -3932,6 +3936,34 @@ mod tests {
); );
} }
#[test]
fn metadata_early_stop_bounded_fanout_defaults_to_enabled() {
temp_env::with_vars(
[
(ENV_RUSTFS_GET_METADATA_EARLY_STOP_ENABLE, Some("true")),
(ENV_RUSTFS_GET_METADATA_DATA_READ_EARLY_STOP_ENABLE, None),
(ENV_RUSTFS_GET_METADATA_EARLY_STOP_BOUNDED_FANOUT, None),
],
|| {
assert!(is_get_metadata_data_read_early_stop_enabled());
assert!(is_get_metadata_early_stop_bounded_fanout_enabled());
},
);
temp_env::with_vars([(ENV_RUSTFS_GET_METADATA_EARLY_STOP_BOUNDED_FANOUT, Some("false"))], || {
assert!(!is_get_metadata_early_stop_bounded_fanout_enabled());
});
temp_env::with_vars(
[
(ENV_RUSTFS_GET_METADATA_DATA_READ_EARLY_STOP_ENABLE, Some("false")),
(ENV_RUSTFS_GET_METADATA_EARLY_STOP_BOUNDED_FANOUT, Some("true")),
],
|| {
assert!(!is_get_metadata_data_read_early_stop_enabled());
assert!(is_get_metadata_early_stop_bounded_fanout_enabled());
},
);
}
#[test] #[test]
fn metadata_early_stop_rejects_healing_and_free_version_requests() { fn metadata_early_stop_rejects_healing_and_free_version_requests() {
temp_env::with_vars( temp_env::with_vars(
+1 -4
View File
@@ -135,10 +135,7 @@ impl FileMeta {
let i = buf.len() as u64; let i = buf.len() as u64;
// check version, buf = buf[8..] // check version, buf = buf[8..]
let (buf, _, _) = Self::check_xl2_v1(buf).map_err(|e| { let (buf, _, _) = Self::check_xl2_v1(buf)?;
error!("failed to check XL2 v1 format: {}", e);
e
})?;
if buf.len() < 5 { if buf.len() < 5 {
error!( error!(
+7 -2
View File
@@ -82,8 +82,8 @@ impl Error {
/// Whether a heal operation can be retried without changing its inputs. /// Whether a heal operation can be retried without changing its inputs.
pub(crate) fn is_recoverable_heal(&self) -> bool { pub(crate) fn is_recoverable_heal(&self) -> bool {
match self { match self {
Error::TaskCancelled => false, Error::TaskCancelled | Error::TaskTimeout => false,
Error::TaskTimeout | Error::TransientSkip { .. } => true, Error::TransientSkip { .. } => true,
Error::Storage(err) => { Error::Storage(err) => {
err.is_quorum_error() err.is_quorum_error()
|| matches!( || matches!(
@@ -165,4 +165,9 @@ mod tests {
assert!(Error::Storage(EcstoreError::DiskNotFound).is_recoverable_heal()); assert!(Error::Storage(EcstoreError::DiskNotFound).is_recoverable_heal());
assert!(Error::Storage(EcstoreError::VolumeNotFound).is_recoverable_heal()); assert!(Error::Storage(EcstoreError::VolumeNotFound).is_recoverable_heal());
} }
#[test]
fn task_timeout_is_terminal() {
assert!(!Error::TaskTimeout.is_recoverable_heal());
}
} }
+30 -20
View File
@@ -673,6 +673,12 @@ fn retry_request_for_result(task: &HealTask, result: &Result<()>) -> Option<(Hea
Some((request, delay, error)) Some((request, delay, error))
} }
async fn retry_request_for_result_with_budget(task: &HealTask, result: &Result<()>) -> Option<(HealRequest, Duration, String)> {
let (_, delay, error) = retry_request_for_result(task, result)?;
let request = task.retry_request_with_remaining_timeout().await.ok()?;
Some((request, delay, error))
}
fn recoverable_heal_retry_delay(retry_attempt: u32) -> Duration { fn recoverable_heal_retry_delay(retry_attempt: u32) -> Duration {
let retry_attempt = retry_attempt.clamp(1, 5); let retry_attempt = retry_attempt.clamp(1, 5);
let delay = Duration::from_secs(2_u64.saturating_pow(retry_attempt)); let delay = Duration::from_secs(2_u64.saturating_pow(retry_attempt));
@@ -690,7 +696,7 @@ pub struct HealConfig {
pub max_concurrent_heals: usize, pub max_concurrent_heals: usize,
/// Maximum concurrent heal tasks allowed for a single erasure set /// Maximum concurrent heal tasks allowed for a single erasure set
pub max_concurrent_per_set: usize, pub max_concurrent_per_set: usize,
/// Task timeout /// Aggregate task execution timeout across recoverable retries
pub task_timeout: Duration, pub task_timeout: Duration,
/// Queue size /// Queue size
pub queue_size: usize, pub queue_size: usize,
@@ -3106,7 +3112,7 @@ impl HealManager {
"Heal scheduler task started" "Heal scheduler task started"
); );
let result = task.execute().await; let result = task.execute().await;
let retry_request = retry_request_for_result(task.as_ref(), &result); let retry_request = retry_request_for_result_with_budget(task.as_ref(), &result).await;
match &result { match &result {
Ok(_) => { Ok(_) => {
debug!( debug!(
@@ -4539,6 +4545,25 @@ mod tests {
assert!(retry_error.contains("Lock acquisition timeout")); assert!(retry_error.contains("Lock acquisition timeout"));
} }
#[tokio::test]
async fn retry_request_for_result_preserves_remaining_timeout_budget() {
let storage: Arc<dyn HealStorageAPI> = Arc::new(MockStorage);
let mut request = HealRequest::object("retry-transition".to_string(), "object".to_string(), None);
request.options.timeout = Some(Duration::from_secs(60));
let task = HealTask::from_request(request, storage);
let result = task.execute().await;
let (retry_request, _, _) = retry_request_for_result_with_budget(&task, &result)
.await
.expect("read quorum failure should retain the unused timeout budget");
let remaining = retry_request
.options
.timeout
.expect("configured timeout should remain present");
assert!(remaining < Duration::from_secs(60));
assert!(remaining > Duration::from_secs(59));
}
#[test] #[test]
fn test_retry_request_for_incomplete_heal_rename() { fn test_retry_request_for_incomplete_heal_rename() {
let storage: Arc<dyn HealStorageAPI> = Arc::new(MockStorage); let storage: Arc<dyn HealStorageAPI> = Arc::new(MockStorage);
@@ -6054,7 +6079,7 @@ mod tests {
process_manager_queue_once(&manager).await; process_manager_queue_once(&manager).await;
let defaulted_status = tokio::time::timeout(Duration::from_secs(1), async { let defaulted_status = tokio::time::timeout(Duration::from_secs(1), async {
loop { loop {
if let Ok(status @ HealTaskStatus::Retrying { .. }) = manager.get_task_status(&defaulted_id).await { if let Ok(status @ HealTaskStatus::Timeout) = manager.get_task_status(&defaulted_id).await {
break status; break status;
} }
tokio::task::yield_now().await; tokio::task::yield_now().await;
@@ -6062,23 +6087,8 @@ mod tests {
}) })
.await .await
.expect("configured timeout should finish the task"); .expect("configured timeout should finish the task");
assert!(matches!(defaulted_status, HealTaskStatus::Retrying { .. })); assert_eq!(defaulted_status, HealTaskStatus::Timeout);
assert_eq!( assert!(manager.retrying_heals.lock().await.get(&defaulted_id).is_none());
manager
.retrying_heals
.lock()
.await
.get(&defaulted_id)
.expect("timed out task should retain its retry request")
.request
.options
.timeout,
Some(Duration::ZERO)
);
manager
.cancel_task(&defaulted_id)
.await
.expect("retrying timeout task should be cancelled");
let mut explicit = bucket_request("explicit-timeout", HealPriority::Normal, HealRequestSource::Admin); let mut explicit = bucket_request("explicit-timeout", HealPriority::Normal, HealRequestSource::Admin);
explicit.options.timeout = Some(Duration::from_secs(60)); explicit.options.timeout = Some(Duration::from_secs(60));
+39 -1
View File
@@ -196,7 +196,7 @@ pub struct HealOptions {
/// Whether to skip namespace locking /// Whether to skip namespace locking
#[serde(default)] #[serde(default)]
pub no_lock: bool, pub no_lock: bool,
/// Timeout /// Aggregate execution timeout across recoverable manager retries
pub timeout: Option<Duration>, pub timeout: Option<Duration>,
/// pool index /// pool index
pub pool_index: Option<usize>, pub pool_index: Option<usize>,
@@ -442,6 +442,14 @@ impl HealTask {
} }
} }
pub(crate) async fn retry_request_with_remaining_timeout(&self) -> Result<HealRequest> {
let mut request = self.retry_request();
if self.options.timeout.is_some() {
request.options.timeout = self.remaining_timeout().await?;
}
Ok(request)
}
pub(crate) fn from_replacement_recovery_request( pub(crate) fn from_replacement_recovery_request(
request: HealRequest, request: HealRequest,
storage: Arc<dyn HealStorageAPI>, storage: Arc<dyn HealStorageAPI>,
@@ -2657,6 +2665,36 @@ mod tests {
use super::super::storage_api::status::BucketInfo; use super::super::storage_api::status::BucketInfo;
#[tokio::test]
async fn retry_request_carries_remaining_timeout_budget() {
let storage: Arc<dyn HealStorageAPI> = Arc::new(MockStorage::default());
let mut request = HealRequest::bucket("bucket".to_string());
request.options.timeout = Some(Duration::from_secs(100));
let task = HealTask::from_request(request, storage.clone());
*task.task_start_instant.write().await = Some(Instant::now() - Duration::from_secs(40));
let retry = task
.retry_request_with_remaining_timeout()
.await
.expect("first retry should retain the unused timeout budget");
let first_remaining = retry.options.timeout.expect("configured timeout should remain present");
assert!(first_remaining <= Duration::from_secs(60));
assert!(first_remaining > Duration::from_secs(59));
let retry_task = HealTask::from_request(retry, storage);
*retry_task.task_start_instant.write().await = Some(Instant::now() - Duration::from_secs(20));
let second_retry = retry_task
.retry_request_with_remaining_timeout()
.await
.expect("second retry should retain only the unused aggregate budget");
let second_remaining = second_retry
.options
.timeout
.expect("configured timeout should remain present");
assert!(second_remaining <= Duration::from_secs(40));
assert!(second_remaining > Duration::from_secs(39));
}
#[test] #[test]
fn format_result_requires_every_requested_target_to_be_ok() { fn format_result_requires_every_requested_target_to_be_ok() {
let result = HealResultItem { let result = HealResultItem {
-2
View File
@@ -18,8 +18,6 @@
//! data encryption keys using master keys. It abstracts the encryption //! data encryption keys using master keys. It abstracts the encryption
//! operations so that different backends can share the same encryption logic. //! operations so that different backends can share the same encryption logic.
#![allow(dead_code)] // Trait methods may be used by implementations
use crate::error::{KmsError, Result}; use crate::error::{KmsError, Result};
use crate::persisted_observability::{BoundedUnknownFieldName, UnknownFieldSummary}; use crate::persisted_observability::{BoundedUnknownFieldName, UnknownFieldSummary};
use async_trait::async_trait; use async_trait::async_trait;
-2
View File
@@ -225,8 +225,6 @@ async fn nothing_readable_leaves_the_bundle_unwrapped() {
"artifact {} carries the raw on-disk record", "artifact {} carries the raw on-disk record",
artifact.path artifact.path
); );
// A cheap structural check too: an encrypted payload is not JSON.
assert_ne!(payload.first(), Some(&b'{'), "artifact {} looks like plaintext JSON", artifact.path);
} }
// The manifest itself is not encrypted, so assert directly that it carries // The manifest itself is not encrypted, so assert directly that it carries
+32 -1
View File
@@ -520,6 +520,14 @@ struct FailureSample {
pub struct FailStats { pub struct FailStats {
pub count: i64, pub count: i64,
pub size: i64, pub size: i64,
/// Rolling-window snapshots refreshed at collection time
/// ([`Self::refresh_windows`]). The raw samples (`recent`) are process
/// local (serde-skipped), so these fields are what survives the peer-RPC
/// wire and [`Self::merge`]-based cluster aggregation.
#[serde(default)]
pub last_minute: FailedMetric,
#[serde(default)]
pub last_hour: FailedMetric,
#[serde(skip)] #[serde(skip)]
recent: VecDeque<FailureSample>, recent: VecDeque<FailureSample>,
} }
@@ -537,6 +545,17 @@ impl FailStats {
self.prune(observed_at); self.prune(observed_at);
} }
/// Recompute the serializable rolling-window snapshots from the local
/// samples. Called at the collection point (per-node stats snapshot),
/// never on the failure hot path — the two deque scans are O(window) and
/// `add_size` runs under the bucket-stats write lock. Only meaningful on
/// the live per-node struct: a deserialized or merged struct has no
/// samples, and refreshing it would wipe the aggregated windows.
pub fn refresh_windows(&mut self) {
self.last_minute = self.recent_since(Duration::from_secs(60));
self.last_hour = self.recent_since(Duration::from_secs(3600));
}
fn prune(&mut self, observed_at: Instant) { fn prune(&mut self, observed_at: Instant) {
while self while self
.recent .recent
@@ -565,6 +584,16 @@ impl FailStats {
Self { Self {
count: self.count.saturating_add(other.count), count: self.count.saturating_add(other.count),
size: self.size.saturating_add(other.size), size: self.size.saturating_add(other.size),
// The window snapshots sum across nodes; the raw samples do not
// travel and stay empty on aggregated structs.
last_minute: FailedMetric {
count: self.last_minute.count.saturating_add(other.last_minute.count),
size: self.last_minute.size.saturating_add(other.last_minute.size),
},
last_hour: FailedMetric {
count: self.last_hour.count.saturating_add(other.last_hour.count),
size: self.last_hour.size.saturating_add(other.last_hour.size),
},
recent: VecDeque::new(), recent: VecDeque::new(),
} }
} }
@@ -636,7 +665,9 @@ impl BucketReplicationStat {
} }
pub fn update_xfer_rate(&mut self, size: i64, duration: Duration) { pub fn update_xfer_rate(&mut self, size: i64, duration: Duration) {
if size > 1024 * 1024 { // Same boundary as the worker-pool split and minio-go's
// Large/Small transfer-summary labels: >= 128 MiB is "large".
if size >= crate::runtime::MIN_LARGE_OBJ_SIZE {
self.xfer_rate_lrg.add_size(size, duration); self.xfer_rate_lrg.add_size(size, duration);
} else { } else {
self.xfer_rate_sml.add_size(size, duration); self.xfer_rate_sml.add_size(size, duration);
@@ -6,9 +6,18 @@
# #
# The MinIO release is pinned so the captured fixture format is reproducible; # The MinIO release is pinned so the captured fixture format is reproducible;
# this is the release the interop tests were validated against. # this is the release the interop tests were validated against.
FROM minio/minio:RELEASE.2025-09-07T16-13-09Z AS minio #
# Both base images are build args so a network that cannot reach Docker Hub can
# point them at a mirror carrying the same content — quay.io publishes the MinIO
# releases, and public.ecr.aws mirrors the official Python images. CI keeps the
# Docker Hub defaults. Override with:
# --build-arg MINIO_IMAGE=quay.io/minio/minio:RELEASE.2025-09-07T16-13-09Z \
# --build-arg PYTHON_IMAGE=public.ecr.aws/docker/library/python:3.12-slim
ARG MINIO_IMAGE=minio/minio:RELEASE.2025-09-07T16-13-09Z
ARG PYTHON_IMAGE=python:3.12-slim
FROM ${MINIO_IMAGE} AS minio
FROM python:3.12-slim FROM ${PYTHON_IMAGE}
RUN apt-get update \ RUN apt-get update \
&& apt-get install -y --no-install-recommends openssl ca-certificates \ && apt-get install -y --no-install-recommends openssl ca-certificates \
&& rm -rf /var/lib/apt/lists/* && rm -rf /var/lib/apt/lists/*
@@ -22,6 +22,18 @@ Use the automated path when you want the lab to:
- upload a predefined SSE fixture case - upload a predefined SSE fixture case
- export the generated backend tree into the lab layout - export the generated backend tree into the lab layout
## Networks without Docker Hub access
`capture_via_docker.sh` pulls its two base images from Docker Hub by default. Where that registry is unreachable, point the build at mirrors carrying the same content — quay.io publishes the MinIO releases and public.ecr.aws mirrors the official Python images:
```bash
MINIO_LAB_MINIO_IMAGE=quay.io/minio/minio:RELEASE.2025-09-07T16-13-09Z \
MINIO_LAB_PYTHON_IMAGE=public.ecr.aws/docker/library/python:3.12-slim \
./capture_via_docker.sh
```
Pin the MinIO tag to the same release the Dockerfile names; an unpinned `:latest` captures whatever format that day's build writes, which is not what the interop tests were validated against.
## Layout ## Layout
The default root is `artifacts/minio-fixture-lab`, which is already ignored by the repository. The default root is `artifacts/minio-fixture-lab`, which is already ignored by the repository.
@@ -34,8 +34,19 @@ if [ "${cases[0]}" != "all" ]; then
done done
fi fi
# Base images are overridable so a network without Docker Hub access can point
# them at a mirror (see the Dockerfile header). Unset by default, which keeps the
# Dockerfile's Docker Hub defaults for CI.
build_args=()
if [ -n "${MINIO_LAB_MINIO_IMAGE:-}" ]; then
build_args+=(--build-arg "MINIO_IMAGE=${MINIO_LAB_MINIO_IMAGE}")
fi
if [ -n "${MINIO_LAB_PYTHON_IMAGE:-}" ]; then
build_args+=(--build-arg "PYTHON_IMAGE=${MINIO_LAB_PYTHON_IMAGE}")
fi
echo ">> building ${IMAGE}" echo ">> building ${IMAGE}"
docker build -f "${SCRIPT_DIR}/Dockerfile" -t "${IMAGE}" "${SCRIPT_DIR}" docker build -f "${SCRIPT_DIR}/Dockerfile" -t "${IMAGE}" "${build_args[@]}" "${SCRIPT_DIR}"
echo ">> capturing fixtures into ${FIXTURE_REL}" echo ">> capturing fixtures into ${FIXTURE_REL}"
docker run --rm -v "${REPO_ROOT}:/repo" "${IMAGE}" \ docker run --rm -v "${REPO_ROOT}:/repo" "${IMAGE}" \
+1 -1
View File
@@ -102,7 +102,7 @@ bytes.workspace = true
hex-simd.workspace = true hex-simd.workspace = true
[dev-dependencies] [dev-dependencies]
tracing-subscriber = { workspace = true, features = ["env-filter", "time"] } tracing-subscriber = { workspace = true, features = ["json", "env-filter", "time"] }
serial_test = { workspace = true } serial_test = { workspace = true }
temp-env = { workspace = true } temp-env = { workspace = true }
tempfile = { workspace = true } tempfile = { workspace = true }
+111 -12
View File
@@ -65,6 +65,7 @@ const LOG_SUBSYSTEM_FOLDER: &str = "folder";
const LOG_SUBSYSTEM_LIFECYCLE: &str = "lifecycle"; const LOG_SUBSYSTEM_LIFECYCLE: &str = "lifecycle";
const LOG_SUBSYSTEM_HEAL: &str = "heal"; const LOG_SUBSYSTEM_HEAL: &str = "heal";
const EVENT_SCANNER_FOLDER_STATE: &str = "scanner_folder_state"; const EVENT_SCANNER_FOLDER_STATE: &str = "scanner_folder_state";
const EVENT_SCANNER_METADATA_CORRUPT: &str = "scanner_metadata_corrupt";
const EVENT_SCANNER_LIFECYCLE_ACTION: &str = "scanner_lifecycle_action"; const EVENT_SCANNER_LIFECYCLE_ACTION: &str = "scanner_lifecycle_action";
const EVENT_SCANNER_HEAL_ADMISSION: &str = "scanner_heal_admission"; const EVENT_SCANNER_HEAL_ADMISSION: &str = "scanner_heal_admission";
const EVENT_SCANNER_ALERT_STATE: &str = "scanner_alert_state"; const EVENT_SCANNER_ALERT_STATE: &str = "scanner_alert_state";
@@ -2154,17 +2155,34 @@ impl FolderScanner {
self.record_failed(&item.path); self.record_failed(&item.path);
if should_log_failed_object(into.failed_objects) { if should_log_failed_object(into.failed_objects) {
warn!( if let GetSizeFailureAction::HealMetadata { object } = &failure_action {
target: "rustfs::scanner::folder", error!(
event = EVENT_SCANNER_FOLDER_STATE, target: "rustfs::scanner::folder",
component = LOG_COMPONENT_SCANNER, event = EVENT_SCANNER_METADATA_CORRUPT,
subsystem = LOG_SUBSYSTEM_FOLDER, component = LOG_COMPONENT_SCANNER,
path = %item.path, subsystem = LOG_SUBSYSTEM_FOLDER,
failed_objects = into.failed_objects, drive = %self.local_disk.path().display(),
state = "get_size_failed", bucket = %item.bucket,
error = %e, object = %object,
"Scanner folder failed to get object size" metadata_path = %item.path,
); failed_objects = into.failed_objects,
state = "metadata_corrupt",
error = %e,
"Scanner detected corrupt object metadata"
);
} else {
warn!(
target: "rustfs::scanner::folder",
event = EVENT_SCANNER_FOLDER_STATE,
component = LOG_COMPONENT_SCANNER,
subsystem = LOG_SUBSYSTEM_FOLDER,
path = %item.path,
failed_objects = into.failed_objects,
state = "get_size_failed",
error = %e,
"Scanner folder failed to get object size"
);
}
} }
} }
@@ -3054,12 +3072,59 @@ mod tests {
use crate::{DiskOption, Endpoint, STORAGE_FORMAT_FILE, TierStats, new_disk, storageclass}; use crate::{DiskOption, Endpoint, STORAGE_FORMAT_FILE, TierStats, new_disk, storageclass};
use rustfs_filemeta::{FileInfo, FileMeta}; use rustfs_filemeta::{FileInfo, FileMeta};
use serial_test::serial; use serial_test::serial;
use std::io::Write;
#[cfg(unix)] #[cfg(unix)]
use std::os::unix::fs::{PermissionsExt, symlink}; use std::os::unix::fs::{PermissionsExt, symlink};
use std::sync::Mutex;
use std::sync::atomic::{AtomicBool, AtomicUsize, Ordering}; use std::sync::atomic::{AtomicBool, AtomicUsize, Ordering};
use temp_env::{with_var, with_var_unset}; use temp_env::{with_var, with_var_unset};
use tracing_subscriber::fmt::MakeWriter;
use uuid::Uuid; use uuid::Uuid;
#[derive(Clone, Default)]
struct CapturedLogs {
buffer: Arc<Mutex<Vec<u8>>>,
}
struct CapturedLogWriter {
buffer: Arc<Mutex<Vec<u8>>>,
}
impl CapturedLogs {
fn contents(&self) -> String {
let buffer = self
.buffer
.lock()
.expect("captured logs mutex should not be poisoned")
.clone();
String::from_utf8(buffer).expect("captured logs should be valid UTF-8")
}
}
impl Write for CapturedLogWriter {
fn write(&mut self, buf: &[u8]) -> std::io::Result<usize> {
self.buffer
.lock()
.expect("captured logs mutex should not be poisoned")
.extend_from_slice(buf);
Ok(buf.len())
}
fn flush(&mut self) -> std::io::Result<()> {
Ok(())
}
}
impl<'a> MakeWriter<'a> for CapturedLogs {
type Writer = CapturedLogWriter;
fn make_writer(&'a self) -> Self::Writer {
CapturedLogWriter {
buffer: Arc::clone(&self.buffer),
}
}
}
#[test] #[test]
fn scanner_size_summary_application_saturates_usage_counters() { fn scanner_size_summary_application_saturates_usage_counters() {
let target = "arn:minio:replication::target".to_string(); let target = "arn:minio:replication::target".to_string();
@@ -4542,9 +4607,19 @@ mod tests {
assert!(budget.entries_visited() >= 1); assert!(budget.entries_visited() >= 1);
} }
#[tokio::test] #[tokio::test(flavor = "current_thread")]
#[serial] #[serial]
async fn test_scan_folder_corrupt_xl_meta_stops_erasure_data_dir_descent() { async fn test_scan_folder_corrupt_xl_meta_stops_erasure_data_dir_descent() {
let logs = CapturedLogs::default();
let subscriber = tracing_subscriber::fmt()
.json()
.with_max_level(tracing::Level::ERROR)
.with_writer(logs.clone())
.with_ansi(false)
.without_time()
.finish();
let _subscriber_guard = tracing::subscriber::set_default(subscriber);
let (mut scanner, temp_dir) = build_test_scanner().await; let (mut scanner, temp_dir) = build_test_scanner().await;
let _guard = TestGuard::new(60, 100, &mut scanner, temp_dir.clone()); let _guard = TestGuard::new(60, 100, &mut scanner, temp_dir.clone());
@@ -4596,6 +4671,30 @@ mod tests {
assert!(!budget.budget_elapsed()); assert!(!budget.budget_elapsed());
assert_eq!(budget.reason(), None); assert_eq!(budget.reason(), None);
let captured = logs.contents();
assert!(
!captured.contains("failed to check XL2 v1 format"),
"the context-free filemeta parser error must not be emitted"
);
let events = captured
.lines()
.map(|line| serde_json::from_str::<serde_json::Value>(line).expect("captured scanner log should be valid JSON"))
.filter(|line| line["fields"]["event"] == EVENT_SCANNER_METADATA_CORRUPT)
.collect::<Vec<_>>();
assert_eq!(
events.len(),
1,
"one corrupt metadata observation must emit one scanner-owned diagnostic event"
);
let fields = &events[0]["fields"];
assert_eq!(fields["component"], LOG_COMPONENT_SCANNER);
assert_eq!(fields["subsystem"], LOG_SUBSYSTEM_FOLDER);
assert_eq!(fields["drive"], temp_dir.to_string_lossy().as_ref());
assert_eq!(fields["bucket"], "bucket");
assert_eq!(fields["object"], "object");
assert_eq!(fields["metadata_path"], metadata_path.to_string_lossy().as_ref());
assert_eq!(fields["state"], "metadata_corrupt");
let retry_budget = ScannerCycleBudget::new_with_progress_tracking( let retry_budget = ScannerCycleBudget::new_with_progress_tracking(
&parent, &parent,
crate::scanner_budget::ScannerCycleBudgetConfig { crate::scanner_budget::ScannerCycleBudgetConfig {
-11
View File
@@ -3849,17 +3849,6 @@ impl ScannerIODisk for Disk {
let fivs = match meta.get_file_info_versions(item.bucket.as_str(), item.object_path().as_str(), false) { let fivs = match meta.get_file_info_versions(item.bucket.as_str(), item.object_path().as_str(), false) {
Ok(versions) => versions, Ok(versions) => versions,
Err(e) => { Err(e) => {
error!(
target: "rustfs::scanner::io",
event = EVENT_SCANNER_DISK_BUCKET_STATE,
component = LOG_COMPONENT_SCANNER,
subsystem = LOG_SUBSYSTEM_IO,
bucket = %item.bucket,
object = %item.object_path(),
state = "file_info_versions_failed",
error = %e,
"Scanner disk bucket failed to resolve file info versions"
);
return Err(scanner_metadata_corrupt_error( return Err(scanner_metadata_corrupt_error(
format!("failed to resolve file info versions: {e}"), format!("failed to resolve file info versions: {e}"),
&item.bucket, &item.bucket,
+284 -279
View File
@@ -1069,309 +1069,314 @@ mod serial_tests {
#[serial] #[serial]
#[ignore = "global-state ILM integration test: runs serialized in the CI ILM Integration (serial) lane, see ci.yml test-ilm-integration-serial and rustfs/backlog#1148 (ilm-1)"] #[ignore = "global-state ILM integration test: runs serialized in the CI ILM Integration (serial) lane, see ci.yml test-ilm-integration-serial and rustfs/backlog#1148 (ilm-1)"]
async fn test_transition_and_restore_flows() { async fn test_transition_and_restore_flows() {
let (disk_paths, ecstore) = setup_test_env().await; async move {
let (disk_paths, ecstore) = setup_test_env().await;
let tier_name = format!("COLDTIER{}", &Uuid::new_v4().simple().to_string()[..8]).to_uppercase(); let tier_name = format!("COLDTIER{}", &Uuid::new_v4().simple().to_string()[..8]).to_uppercase();
let backend = register_mock_tier(&tier_name).await; let backend = register_mock_tier(&tier_name).await;
let put_bucket = format!("test-immediate-put-{}", &Uuid::new_v4().simple().to_string()[..8]); let put_bucket = format!("test-immediate-put-{}", &Uuid::new_v4().simple().to_string()[..8]);
let put_object = "test/object.txt"; let put_object = "test/object.txt";
let put_payload = b"Hello, immediate transition!"; let put_payload = b"Hello, immediate transition!";
create_test_bucket(&ecstore, put_bucket.as_str()).await; create_test_bucket(&ecstore, put_bucket.as_str()).await;
set_bucket_lifecycle_transition_with_tier(put_bucket.as_str(), &tier_name) set_bucket_lifecycle_transition_with_tier(put_bucket.as_str(), &tier_name)
.await
.expect("Failed to set lifecycle configuration");
let mut reader = PutObjReader::from_vec(put_payload.to_vec());
let mut metadata = HashMap::new();
metadata.insert("content-type".to_string(), "text/plain".to_string());
ecstore
.put_object(
put_bucket.as_str(),
put_object,
&mut reader,
&ObjectOptions {
user_defined: metadata,
..Default::default()
},
)
.await
.expect("Failed to upload transition metadata test object");
enqueue_transition_for_existing_objects(ecstore.clone(), put_bucket.as_str())
.await
.expect("Failed to enqueue transitioned put object");
let put_info = wait_for_transition(&ecstore, put_bucket.as_str(), put_object, TRANSITION_WAIT_TIMEOUT)
.await
.expect("object should transition after enqueueing existing objects");
assert_eq!(put_info.transitioned_object.status, "complete");
assert_eq!(put_info.transitioned_object.tier, tier_name);
assert!(backend.contains(&put_info.transitioned_object.name).await);
{
let transitioned = backend
.stored(&put_info.transitioned_object.name)
.await .await
.expect("transitioned object should be present in mock backend"); .expect("Failed to set lifecycle configuration");
assert_eq!(transitioned.metadata.get("content-type"), Some(&"text/plain".to_string()));
assert!(
!transitioned.metadata.contains_key("x-amz-replication-status"),
"transitioned objects must not inherit replication status defaults"
);
assert!(
!transitioned.metadata.contains_key("x-amz-object-lock-legal-hold"),
"transitioned objects must not invent object lock headers"
);
}
// Cross-shard xl.meta transition assertion helper (rustfs/backlog#1148 ilm-6): let mut reader = PutObjReader::from_vec(put_payload.to_vec());
// every disk must agree on the transition tuple for the object. let mut metadata = HashMap::new();
let put_meta = assert_transition_meta_consistent(&disk_paths, put_bucket.as_str(), put_object).await; metadata.insert("content-type".to_string(), "text/plain".to_string());
assert_eq!(put_meta.status, "complete"); ecstore
assert_eq!(put_meta.tier, tier_name); .put_object(
put_bucket.as_str(),
put_object,
&mut reader,
&ObjectOptions {
user_defined: metadata,
..Default::default()
},
)
.await
.expect("Failed to upload transition metadata test object");
let multipart_bucket = format!("test-immediate-mpu-{}", &Uuid::new_v4().simple().to_string()[..8]); enqueue_transition_for_existing_objects(ecstore.clone(), put_bucket.as_str())
let multipart_object = "test/multipart.txt"; .await
.expect("Failed to enqueue transitioned put object");
create_test_bucket(&ecstore, multipart_bucket.as_str()).await; let put_info = wait_for_transition(&ecstore, put_bucket.as_str(), put_object, TRANSITION_WAIT_TIMEOUT)
set_bucket_lifecycle_transition_with_tier(multipart_bucket.as_str(), &tier_name) .await
.await .expect("object should transition after enqueueing existing objects");
.expect("Failed to set lifecycle configuration");
let upload = ecstore assert_eq!(put_info.transitioned_object.status, "complete");
.new_multipart_upload(multipart_bucket.as_str(), multipart_object, &ObjectOptions::default()) assert_eq!(put_info.transitioned_object.tier, tier_name);
.await assert!(backend.contains(&put_info.transitioned_object.name).await);
.expect("Failed to create multipart upload"); {
let transitioned = backend
.stored(&put_info.transitioned_object.name)
.await
.expect("transitioned object should be present in mock backend");
assert_eq!(transitioned.metadata.get("content-type"), Some(&"text/plain".to_string()));
assert!(
!transitioned.metadata.contains_key("x-amz-replication-status"),
"transitioned objects must not inherit replication status defaults"
);
assert!(
!transitioned.metadata.contains_key("x-amz-object-lock-legal-hold"),
"transitioned objects must not invent object lock headers"
);
}
let part_data = b"multipart immediate transition"; // Cross-shard xl.meta transition assertion helper (rustfs/backlog#1148 ilm-6):
let mut reader = PutObjReader::from_vec(part_data.to_vec()); // every disk must agree on the transition tuple for the object.
let part = ecstore let put_meta = assert_transition_meta_consistent(&disk_paths, put_bucket.as_str(), put_object).await;
.put_object_part( assert_eq!(put_meta.status, "complete");
multipart_bucket.as_str(), assert_eq!(put_meta.tier, tier_name);
multipart_object,
&upload.upload_id,
1,
&mut reader,
&ObjectOptions::default(),
)
.await
.expect("Failed to upload multipart part");
ecstore let multipart_bucket = format!("test-immediate-mpu-{}", &Uuid::new_v4().simple().to_string()[..8]);
.clone() let multipart_object = "test/multipart.txt";
.complete_multipart_upload(
multipart_bucket.as_str(),
multipart_object,
&upload.upload_id,
vec![CompletePart {
part_num: 1,
etag: part.etag.clone(),
..Default::default()
}],
&ObjectOptions::default(),
)
.await
.expect("Failed to complete multipart upload");
enqueue_transition_for_existing_objects(ecstore.clone(), multipart_bucket.as_str()) create_test_bucket(&ecstore, multipart_bucket.as_str()).await;
.await set_bucket_lifecycle_transition_with_tier(multipart_bucket.as_str(), &tier_name)
.expect("Failed to enqueue transitioned multipart object"); .await
.expect("Failed to set lifecycle configuration");
let multipart_info = wait_for_transition(&ecstore, multipart_bucket.as_str(), multipart_object, TRANSITION_WAIT_TIMEOUT) let upload = ecstore
.await .new_multipart_upload(multipart_bucket.as_str(), multipart_object, &ObjectOptions::default())
.expect("object should transition after enqueueing existing objects"); .await
.expect("Failed to create multipart upload");
assert_eq!(multipart_info.transitioned_object.status, "complete"); let part_data = b"multipart immediate transition";
assert_eq!(multipart_info.transitioned_object.tier, tier_name); let mut reader = PutObjReader::from_vec(part_data.to_vec());
assert!(backend.contains(&multipart_info.transitioned_object.name).await); let part = ecstore
.put_object_part(
multipart_bucket.as_str(),
multipart_object,
&upload.upload_id,
1,
&mut reader,
&ObjectOptions::default(),
)
.await
.expect("Failed to upload multipart part");
let src_bucket = format!("test-immediate-copy-src-{}", &Uuid::new_v4().simple().to_string()[..8]); ecstore
let dst_bucket = format!("test-immediate-copy-dst-{}", &Uuid::new_v4().simple().to_string()[..8]); .clone()
let src_object = "test/source.txt"; .complete_multipart_upload(
let dst_object = "test/copied.txt"; multipart_bucket.as_str(),
let payload = b"copy object immediate transition"; multipart_object,
&upload.upload_id,
create_test_bucket(&ecstore, src_bucket.as_str()).await; vec![CompletePart {
create_test_bucket(&ecstore, dst_bucket.as_str()).await;
set_bucket_lifecycle_transition_with_tier(dst_bucket.as_str(), &tier_name)
.await
.expect("Failed to set destination lifecycle configuration");
upload_test_object(&ecstore, src_bucket.as_str(), src_object, payload).await;
let mut src_info = ecstore
.get_object_info(src_bucket.as_str(), src_object, &ObjectOptions::default())
.await
.expect("Failed to load source object info");
src_info.put_object_reader = Some(PutObjReader::from_vec(payload.to_vec()));
ecstore
.copy_object(
src_bucket.as_str(),
src_object,
dst_bucket.as_str(),
dst_object,
&mut src_info,
&ObjectOptions::default(),
&ObjectOptions::default(),
)
.await
.expect("Failed to copy object");
enqueue_transition_for_existing_objects(ecstore.clone(), dst_bucket.as_str())
.await
.expect("Failed to enqueue transitioned copied object");
let copy_info = wait_for_transition(&ecstore, dst_bucket.as_str(), dst_object, TRANSITION_WAIT_TIMEOUT)
.await
.expect("copied object should transition after enqueueing existing objects");
assert_eq!(copy_info.transitioned_object.status, "complete");
assert_eq!(copy_info.transitioned_object.tier, tier_name);
assert!(backend.contains(&copy_info.transitioned_object.name).await);
let bucket_name = format!("test-lifecycle-update-{}", &Uuid::new_v4().simple().to_string()[..8]);
let object_name = "test/existing.txt";
let payload = b"existing object before lifecycle";
create_test_bucket(&ecstore, bucket_name.as_str()).await;
upload_test_object(&ecstore, bucket_name.as_str(), object_name, payload).await;
set_bucket_lifecycle_transition_with_tier(bucket_name.as_str(), &tier_name)
.await
.expect("Failed to set lifecycle configuration");
enqueue_transition_for_existing_objects(ecstore.clone(), bucket_name.as_str())
.await
.expect("Failed to enqueue transition for existing objects");
let info = wait_for_transition(&ecstore, bucket_name.as_str(), object_name, TRANSITION_WAIT_TIMEOUT)
.await
.expect("existing object should transition after lifecycle update");
assert_eq!(info.transitioned_object.status, "complete");
assert_eq!(info.transitioned_object.tier, tier_name);
assert!(backend.contains(&info.transitioned_object.name).await);
let bucket_name = format!("test-restore-mpu-{}", &Uuid::new_v4().simple().to_string()[..8]);
let object_name = "test/restore.txt";
let part1 = vec![b'a'; 5 * 1024 * 1024];
let part2 = b"restored-tail".to_vec();
let expected = [part1.clone(), part2.clone()].concat();
create_test_bucket(&ecstore, bucket_name.as_str()).await;
set_bucket_lifecycle_transition_with_tier(bucket_name.as_str(), &tier_name)
.await
.expect("Failed to set lifecycle configuration");
let upload = ecstore
.new_multipart_upload(bucket_name.as_str(), object_name, &ObjectOptions::default())
.await
.expect("Failed to create multipart upload");
let mut part1_reader = PutObjReader::from_vec(part1);
let uploaded_part1 = ecstore
.put_object_part(
bucket_name.as_str(),
object_name,
&upload.upload_id,
1,
&mut part1_reader,
&ObjectOptions::default(),
)
.await
.expect("Failed to upload first multipart part");
let mut part2_reader = PutObjReader::from_vec(part2);
let uploaded_part2 = ecstore
.put_object_part(
bucket_name.as_str(),
object_name,
&upload.upload_id,
2,
&mut part2_reader,
&ObjectOptions::default(),
)
.await
.expect("Failed to upload second multipart part");
ecstore
.clone()
.complete_multipart_upload(
bucket_name.as_str(),
object_name,
&upload.upload_id,
vec![
CompletePart {
part_num: 1, part_num: 1,
etag: uploaded_part1.etag.clone(), etag: part.etag.clone(),
..Default::default() ..Default::default()
}, }],
CompletePart { &ObjectOptions::default(),
part_num: 2, )
etag: uploaded_part2.etag.clone(), .await
..Default::default() .expect("Failed to complete multipart upload");
},
],
&ObjectOptions::default(),
)
.await
.expect("Failed to complete multipart upload");
enqueue_transition_for_existing_objects(ecstore.clone(), bucket_name.as_str()) enqueue_transition_for_existing_objects(ecstore.clone(), multipart_bucket.as_str())
.await .await
.expect("Failed to enqueue transitioned restore object"); .expect("Failed to enqueue transitioned multipart object");
let transitioned = wait_for_transition(&ecstore, bucket_name.as_str(), object_name, TRANSITION_WAIT_TIMEOUT) let multipart_info =
.await wait_for_transition(&ecstore, multipart_bucket.as_str(), multipart_object, TRANSITION_WAIT_TIMEOUT)
.expect("multipart object should transition after enqueueing existing objects"); .await
assert_eq!(transitioned.parts.len(), 2); .expect("object should transition after enqueueing existing objects");
ecstore assert_eq!(multipart_info.transitioned_object.status, "complete");
.clone() assert_eq!(multipart_info.transitioned_object.tier, tier_name);
.restore_transitioned_object( assert!(backend.contains(&multipart_info.transitioned_object.name).await);
bucket_name.as_str(),
object_name, let src_bucket = format!("test-immediate-copy-src-{}", &Uuid::new_v4().simple().to_string()[..8]);
&ObjectOptions { let dst_bucket = format!("test-immediate-copy-dst-{}", &Uuid::new_v4().simple().to_string()[..8]);
transition: TransitionOptions { let src_object = "test/source.txt";
restore_request: RestoreRequest { let dst_object = "test/copied.txt";
days: Some(1), let payload = b"copy object immediate transition";
description: None,
glacier_job_parameters: None, create_test_bucket(&ecstore, src_bucket.as_str()).await;
output_location: None, create_test_bucket(&ecstore, dst_bucket.as_str()).await;
select_parameters: None, set_bucket_lifecycle_transition_with_tier(dst_bucket.as_str(), &tier_name)
tier: None, .await
type_: None, .expect("Failed to set destination lifecycle configuration");
upload_test_object(&ecstore, src_bucket.as_str(), src_object, payload).await;
let mut src_info = ecstore
.get_object_info(src_bucket.as_str(), src_object, &ObjectOptions::default())
.await
.expect("Failed to load source object info");
src_info.put_object_reader = Some(PutObjReader::from_vec(payload.to_vec()));
ecstore
.copy_object(
src_bucket.as_str(),
src_object,
dst_bucket.as_str(),
dst_object,
&mut src_info,
&ObjectOptions::default(),
&ObjectOptions::default(),
)
.await
.expect("Failed to copy object");
enqueue_transition_for_existing_objects(ecstore.clone(), dst_bucket.as_str())
.await
.expect("Failed to enqueue transitioned copied object");
let copy_info = wait_for_transition(&ecstore, dst_bucket.as_str(), dst_object, TRANSITION_WAIT_TIMEOUT)
.await
.expect("copied object should transition after enqueueing existing objects");
assert_eq!(copy_info.transitioned_object.status, "complete");
assert_eq!(copy_info.transitioned_object.tier, tier_name);
assert!(backend.contains(&copy_info.transitioned_object.name).await);
let bucket_name = format!("test-lifecycle-update-{}", &Uuid::new_v4().simple().to_string()[..8]);
let object_name = "test/existing.txt";
let payload = b"existing object before lifecycle";
create_test_bucket(&ecstore, bucket_name.as_str()).await;
upload_test_object(&ecstore, bucket_name.as_str(), object_name, payload).await;
set_bucket_lifecycle_transition_with_tier(bucket_name.as_str(), &tier_name)
.await
.expect("Failed to set lifecycle configuration");
enqueue_transition_for_existing_objects(ecstore.clone(), bucket_name.as_str())
.await
.expect("Failed to enqueue transition for existing objects");
let info = wait_for_transition(&ecstore, bucket_name.as_str(), object_name, TRANSITION_WAIT_TIMEOUT)
.await
.expect("existing object should transition after lifecycle update");
assert_eq!(info.transitioned_object.status, "complete");
assert_eq!(info.transitioned_object.tier, tier_name);
assert!(backend.contains(&info.transitioned_object.name).await);
let bucket_name = format!("test-restore-mpu-{}", &Uuid::new_v4().simple().to_string()[..8]);
let object_name = "test/restore.txt";
let part1 = vec![b'a'; 5 * 1024 * 1024];
let part2 = b"restored-tail".to_vec();
let expected = [part1.clone(), part2.clone()].concat();
create_test_bucket(&ecstore, bucket_name.as_str()).await;
set_bucket_lifecycle_transition_with_tier(bucket_name.as_str(), &tier_name)
.await
.expect("Failed to set lifecycle configuration");
let upload = ecstore
.new_multipart_upload(bucket_name.as_str(), object_name, &ObjectOptions::default())
.await
.expect("Failed to create multipart upload");
let mut part1_reader = PutObjReader::from_vec(part1);
let uploaded_part1 = ecstore
.put_object_part(
bucket_name.as_str(),
object_name,
&upload.upload_id,
1,
&mut part1_reader,
&ObjectOptions::default(),
)
.await
.expect("Failed to upload first multipart part");
let mut part2_reader = PutObjReader::from_vec(part2);
let uploaded_part2 = ecstore
.put_object_part(
bucket_name.as_str(),
object_name,
&upload.upload_id,
2,
&mut part2_reader,
&ObjectOptions::default(),
)
.await
.expect("Failed to upload second multipart part");
ecstore
.clone()
.complete_multipart_upload(
bucket_name.as_str(),
object_name,
&upload.upload_id,
vec![
CompletePart {
part_num: 1,
etag: uploaded_part1.etag.clone(),
..Default::default()
},
CompletePart {
part_num: 2,
etag: uploaded_part2.etag.clone(),
..Default::default()
},
],
&ObjectOptions::default(),
)
.await
.expect("Failed to complete multipart upload");
enqueue_transition_for_existing_objects(ecstore.clone(), bucket_name.as_str())
.await
.expect("Failed to enqueue transitioned restore object");
let transitioned = wait_for_transition(&ecstore, bucket_name.as_str(), object_name, TRANSITION_WAIT_TIMEOUT)
.await
.expect("multipart object should transition after enqueueing existing objects");
assert_eq!(transitioned.parts.len(), 2);
ecstore
.clone()
.restore_transitioned_object(
bucket_name.as_str(),
object_name,
&ObjectOptions {
transition: TransitionOptions {
restore_request: RestoreRequest {
days: Some(1),
description: None,
glacier_job_parameters: None,
output_location: None,
select_parameters: None,
tier: None,
type_: None,
},
..Default::default()
}, },
..Default::default() ..Default::default()
}, },
..Default::default() )
}, .await
) .expect("Failed to restore transitioned multipart object");
.await
.expect("Failed to restore transitioned multipart object");
let restored = ecstore let restored = ecstore
.get_object_info(bucket_name.as_str(), object_name, &ObjectOptions::default()) .get_object_info(bucket_name.as_str(), object_name, &ObjectOptions::default())
.await .await
.expect("Failed to load restored object info"); .expect("Failed to load restored object info");
assert_eq!(restored.parts.len(), 2); assert_eq!(restored.parts.len(), 2);
assert!(restored.restore_expires.is_some()); assert!(restored.restore_expires.is_some());
assert!(!restored.restore_ongoing); assert!(!restored.restore_ongoing);
let mut reader = ecstore let mut reader = ecstore
.get_object_reader(bucket_name.as_str(), object_name, None, http::HeaderMap::new(), &ObjectOptions::default()) .get_object_reader(bucket_name.as_str(), object_name, None, http::HeaderMap::new(), &ObjectOptions::default())
.await .await
.expect("Failed to read restored object"); .expect("Failed to read restored object");
let mut data = Vec::new(); let mut data = Vec::new();
reader reader
.stream .stream
.read_to_end(&mut data) .read_to_end(&mut data)
.await .await
.expect("Failed to consume restored object stream"); .expect("Failed to consume restored object stream");
assert_eq!(data, expected); assert_eq!(data, expected);
}
.boxed_local()
.await;
} }
#[tokio::test(flavor = "multi_thread", worker_threads = 1)] #[tokio::test(flavor = "multi_thread", worker_threads = 1)]
+29
View File
@@ -50,6 +50,14 @@ pub const SUFFIX_SOURCE_DELETEMARKER: &str = "source-deletemarker";
pub const SUFFIX_SOURCE_PROXY_REQUEST: &str = "source-proxy-request"; pub const SUFFIX_SOURCE_PROXY_REQUEST: &str = "source-proxy-request";
pub const SUFFIX_SOURCE_REPLICATION_REQUEST: &str = "source-replication-request"; pub const SUFFIX_SOURCE_REPLICATION_REQUEST: &str = "source-replication-request";
pub const SUFFIX_SOURCE_REPLICATION_CHECK: &str = "source-replication-check"; pub const SUFFIX_SOURCE_REPLICATION_CHECK: &str = "source-replication-check";
// LWW timestamps for replicated tag/retention/legal-hold modifications. MinIO
// declares these with mixed case (internal/http/headers.go:
// X-Minio-Source-Replication-Tagging-Timestamp / -Retention-Timestamp /
// -LegalHold-Timestamp); HTTP header names compare case-insensitively, so the
// lowercase suffix forms interoperate. Values are RFC3339 on the wire.
pub const SUFFIX_SOURCE_REPLICATION_TAGGING_TIMESTAMP: &str = "source-replication-tagging-timestamp";
pub const SUFFIX_SOURCE_REPLICATION_RETENTION_TIMESTAMP: &str = "source-replication-retention-timestamp";
pub const SUFFIX_SOURCE_REPLICATION_LEGALHOLD_TIMESTAMP: &str = "source-replication-legalhold-timestamp";
pub const SUFFIX_REPLICATION_SSEC_CRC: &str = "replication-ssec-crc"; pub const SUFFIX_REPLICATION_SSEC_CRC: &str = "replication-ssec-crc";
/// Returns true if the key is object-encryption metadata understood by RustFS or MinIO. /// Returns true if the key is object-encryption metadata understood by RustFS or MinIO.
@@ -196,6 +204,27 @@ mod tests {
assert_eq!(get_object_encryption_original_size(&metadata).expect("valid size"), Some(42)); assert_eq!(get_object_encryption_original_size(&metadata).expect("valid size"), Some(42));
} }
#[test]
fn replication_timestamp_headers_match_minio_wire_names() {
let mut headers = HeaderMap::new();
insert_header(&mut headers, SUFFIX_SOURCE_REPLICATION_TAGGING_TIMESTAMP, "2026-01-02T03:04:05Z");
insert_header(&mut headers, SUFFIX_SOURCE_REPLICATION_RETENTION_TIMESTAMP, "2026-01-02T03:04:06Z");
insert_header(&mut headers, SUFFIX_SOURCE_REPLICATION_LEGALHOLD_TIMESTAMP, "2026-01-02T03:04:07Z");
// The exact names MinIO's object-api-options.go reads (its Get()
// canonicalizes case, so a case-insensitive match is wire-equivalent).
for name in [
"X-Minio-Source-Replication-Tagging-Timestamp",
"X-Minio-Source-Replication-Retention-Timestamp",
"X-Minio-Source-Replication-LegalHold-Timestamp",
"x-rustfs-source-replication-tagging-timestamp",
"x-rustfs-source-replication-retention-timestamp",
"x-rustfs-source-replication-legalhold-timestamp",
] {
assert!(headers.contains_key(name), "replication timestamp header {name} must be written");
}
}
#[test] #[test]
fn test_get_header() { fn test_get_header() {
let mut headers = HeaderMap::new(); let mut headers = HeaderMap::new();
-3
View File
@@ -659,9 +659,6 @@ mod test {
// Port should be in valid range (u16 max is always <= 65535) // Port should be in valid range (u16 max is always <= 65535)
assert!(port1 > 0); assert!(port1 > 0);
assert!(port2 > 0); assert!(port2 > 0);
// Different calls should typically return different ports
assert_ne!(port1, port2);
} }
#[test] #[test]
@@ -61,17 +61,17 @@ catalog extension.
| Area | Status | Covered behavior | | Area | Status | Covered behavior |
|---|---|---| |---|---|---|
| Catalog config | Supported | `GET /v1/config` advertises RustFS catalog defaults and route capabilities. | | Catalog config | Supported | `GET /v1/config` advertises RustFS catalog defaults and only the supported OpenAPI REST paths in `endpoints`. RustFS administration, maintenance, migration, diagnostics, refs, and metadata-location extensions remain available but are not presented as standard Iceberg REST endpoints. |
| Table bucket discovery | Supported | `PUT` and `GET /v1/buckets/{warehouse}` enable and inspect table bucket state. | | Table bucket discovery | Supported | `PUT` and `GET /v1/buckets/{warehouse}` enable and inspect table bucket state. |
| Namespaces | Supported | Create, list, load, existence check, and drop namespace routes are registered on both catalog prefixes. List responses support Iceberg REST `pageSize`/`pageToken` pagination with context-bound tokens and bounded catalog-store reads. Namespace identifiers are limited to 512 ASCII characters so persisted paths and stateless continuation tokens remain bounded. | | Namespaces | Supported | Create, list, load, existence check, and drop namespace routes are registered on both catalog prefixes. List responses support Iceberg REST `pageSize`/`pageToken` pagination with context-bound tokens and bounded catalog-store reads. Namespace identifiers are limited to 512 ASCII characters so persisted paths and stateless continuation tokens remain bounded. |
| Tables | Supported | Create, register, list, load, existence check, commit, metadata-location get/update, and drop table routes are registered on both catalog prefixes. Table and view listings support Iceberg REST `pageSize`/`pageToken` pagination with context-bound tokens and bounded catalog-store reads. | | Tables | Supported | Create, register, list, load, existence check, commit, metadata-location get/update, and drop table routes are registered on both catalog prefixes. Table and view listings support Iceberg REST `pageSize`/`pageToken` pagination with context-bound tokens and bounded catalog-store reads. Commit identifiers must match the URL resource; unknown requirements, updates, and snapshot operations fail as bad requests; staged create, register overwrite, purge-on-drop, and v3-only encryption-key updates return an explicit unsupported-operation response. Standard statistics, partition statistics, and schema/spec cleanup updates are accepted. |
| Commit CAS | Supported | Single-table commits validate base metadata, expected version token, referenced object existence, warehouse scope, and Iceberg commit requirements before advancing the current metadata pointer. Standard commits preserve the normal commit-token file name and use an immutable-table-scoped fallback when rename followed by source-name reuse would otherwise collide at the same generation and commit ID. | | Commit CAS | Supported | Single-table commits validate base metadata, expected version token, referenced object existence, warehouse scope, and Iceberg commit requirements before advancing the current metadata pointer. Externally supplied metadata transitions preserve monotonic column, partition, and sequence assignment watermarks and immutable definitions for retained schemas, partition specs, sort orders, and snapshots. Standard commits preserve the normal commit-token file name and use an immutable-table-scoped fallback when rename followed by source-name reuse would otherwise collide at the same generation and commit ID. The catalog does not advertise `idempotency-key-lifetime`; clients must treat standard mutation-wide `Idempotency-Key` semantics as unsupported. |
| Commit recovery | Supported | Commit log, idempotency lookup, diagnostics, and recovery routes expose staged/finalization gaps and repair safe idempotency gaps without moving the table pointer. | | Commit recovery | Supported | Commit log, idempotency lookup, diagnostics, and recovery routes expose staged/finalization gaps and repair safe idempotency gaps without moving the table pointer. |
| Snapshot refs | Supported | Refs can be listed, created or replaced, and deleted through catalog commits. `main` is protected and refs with explicit retention require forced delete. | | Snapshot refs | Supported | Refs can be listed, created or replaced, and deleted through catalog commits. `main` is protected and refs with explicit retention require forced delete. |
| Iceberg views | Supported | Basic create, list, load, replace, existence check, and drop routes persist view metadata with view-scoped authorization. | | Iceberg views | Supported | Basic create, list, load, replace, existence check, and drop routes persist view metadata with view-scoped authorization. Replace identifiers must match the URL resource, `schema-id: -1` resolves to the last added schema, one commit timestamp is used consistently, and only Iceberg view format version 1 is accepted. |
| Table credentials endpoint | Supported | Returns an empty `storage-credentials` list by default. Returns table-scoped temporary credentials only when credential vending is enabled. | | Table credentials endpoint | Supported | Returns an empty `storage-credentials` list by default. Returns table-scoped temporary credentials only when credential vending is enabled. Credential responses set `Cache-Control: no-store, private`, `Pragma: no-cache`, and `Expires: 0`. |
| Catalog diagnostics and export | Supported | Exposes recovery state, consistency state, backing manifest, recoverable commit-log WAL state, strong backing migration target, single-active-writer policy, and scale validation matrix. | | Catalog diagnostics and export | Supported | Exposes recovery state, consistency state, backing manifest, recoverable commit-log WAL state, strong backing migration target, single-active-writer policy, and scale validation matrix. |
| Catalog import and rollback | Supported | Import/register and rollback use catalog validation and commit paths rather than direct pointer mutation. | | Catalog import and rollback | Supported | Import/register and online rollback use catalog validation and commit paths rather than direct pointer mutation. Online rollback accepts only a forward-safe metadata target that preserves assignment watermarks and retained definitions. Restoring an older target that lowers those watermarks is an offline disaster-recovery operation and requires every writer to be stopped. |
| External catalog bridge | Supported operator path | Operator-supplied metadata pointer sync/import is supported for external catalog identity boundaries. Online vendor SDK polling and policy mirroring are not claimed. | | External catalog bridge | Supported operator path | Operator-supplied metadata pointer sync/import is supported for external catalog identity boundaries. Online vendor SDK polling and policy mirroring are not claimed. |
| Multi-table transactions | Not claimed | RustFS currently claims single-table commit atomicity only. | | Multi-table transactions | Not claimed | RustFS currently claims single-table commit atomicity only. |
Generated
+6 -6
View File
@@ -2,11 +2,11 @@
"nodes": { "nodes": {
"nixpkgs": { "nixpkgs": {
"locked": { "locked": {
"lastModified": 1785975029, "lastModified": 1786719841,
"narHash": "sha256-X44cn5rzytELc3NNoQsh0aLkjWA/QzPfc6HPQmsG3sU=", "narHash": "sha256-QcpQOT0NQEFkI77t+YXPZqDJc35iIodG7zinieOwFUg=",
"owner": "NixOS", "owner": "NixOS",
"repo": "nixpkgs", "repo": "nixpkgs",
"rev": "70ce234312134a463ba7728e94da2486a1d237ac", "rev": "8be7bd0c83f12e2e3bbba07c9044d6fed9e66f7f",
"type": "github" "type": "github"
}, },
"original": { "original": {
@@ -29,11 +29,11 @@
] ]
}, },
"locked": { "locked": {
"lastModified": 1786247265, "lastModified": 1786849542,
"narHash": "sha256-cVTcTqAhzSNMOVddtBhFpuv+Vx6/SgrhgZWJiI61H24=", "narHash": "sha256-MS8cOa/ii1+QA8R3JWVzqEIoXimu9n50GwQDiBVTFOM=",
"owner": "oxalica", "owner": "oxalica",
"repo": "rust-overlay", "repo": "rust-overlay",
"rev": "6df7076ea0f5697e3242719a4c71211f00411cf5", "rev": "b211eadeba8b180da9453ec3413a8a3535c85b3f",
"type": "github" "type": "github"
}, },
"original": { "original": {
+2 -1
View File
@@ -278,6 +278,8 @@ rustfs-signer.workspace = true
serde = { workspace = true, features = ["derive"] } serde = { workspace = true, features = ["derive"] }
serde_json = { workspace = true, features = ["raw_value"] } serde_json = { workspace = true, features = ["raw_value"] }
serde_urlencoded = { workspace = true } serde_urlencoded = { workspace = true }
snap.workspace = true
zstd.workspace = true
# Cryptography and Security # Cryptography and Security
rustls = { workspace = true, default-features = false, features = ["aws-lc-rs", "logging", "tls12", "prefer-post-quantum", "std"] } rustls = { workspace = true, default-features = false, features = ["aws-lc-rs", "logging", "tls12", "prefer-post-quantum", "std"] }
@@ -355,7 +357,6 @@ rcgen = { workspace = true }
rustfs-test-utils.workspace = true rustfs-test-utils.workspace = true
# diagnose_e2e fixtures (archives are generated in-test, never checked in) # diagnose_e2e fixtures (archives are generated in-test, never checked in)
zip = { workspace = true } zip = { workspace = true }
zstd = { workspace = true }
# Enables the shared MockWarmBackend / xl.meta assertion helpers exposed via # Enables the shared MockWarmBackend / xl.meta assertion helpers exposed via
# the ecstore `api::tier::test_util` facade module (rustfs/backlog#1148 ilm-6). # the ecstore `api::tier::test_util` facade module (rustfs/backlog#1148 ilm-6).
rustfs-ecstore = { workspace = true, features = ["test-util"] } rustfs-ecstore = { workspace = true, features = ["test-util"] }
+8 -7
View File
@@ -40,7 +40,7 @@ use s3s::{Body, S3Request, S3Response, S3Result, s3_error};
use serde::{Deserialize, Serialize}; use serde::{Deserialize, Serialize};
use std::collections::HashMap; use std::collections::HashMap;
use std::sync::LazyLock; use std::sync::LazyLock;
use tracing::{Span, error, info, warn}; use tracing::{error, info, warn};
const LOG_COMPONENT_ADMIN_API: &str = "admin_api"; const LOG_COMPONENT_ADMIN_API: &str = "admin_api";
const LOG_SUBSYSTEM_AUDIT_TARGET: &str = "audit_target"; const LOG_SUBSYSTEM_AUDIT_TARGET: &str = "audit_target";
@@ -278,8 +278,6 @@ pub struct AuditTargetConfig {}
#[async_trait::async_trait] #[async_trait::async_trait]
impl Operation for AuditTargetConfig { impl Operation for AuditTargetConfig {
async fn call(&self, req: S3Request<Body>, params: Params<'_, '_>) -> S3Result<S3Response<(StatusCode, Body)>> { async fn call(&self, req: S3Request<Body>, params: Params<'_, '_>) -> S3Result<S3Response<(StatusCode, Body)>> {
let span = Span::current();
let _enter = span.enter();
let (target_type, target_name) = extract_target_params(&params)?; let (target_type, target_name) = extract_target_params(&params)?;
authorize_audit_admin_request(&req, AdminAction::SetBucketTargetAction).await?; authorize_audit_admin_request(&req, AdminAction::SetBucketTargetAction).await?;
@@ -345,8 +343,6 @@ pub struct ListAuditTargets {}
#[async_trait::async_trait] #[async_trait::async_trait]
impl Operation for ListAuditTargets { impl Operation for ListAuditTargets {
async fn call(&self, req: S3Request<Body>, _params: Params<'_, '_>) -> S3Result<S3Response<(StatusCode, Body)>> { async fn call(&self, req: S3Request<Body>, _params: Params<'_, '_>) -> S3Result<S3Response<(StatusCode, Body)>> {
let span = Span::current();
let _enter = span.enter();
authorize_audit_admin_request(&req, AdminAction::GetBucketTargetAction).await?; authorize_audit_admin_request(&req, AdminAction::GetBucketTargetAction).await?;
let mut runtime_statuses = HashMap::new(); let mut runtime_statuses = HashMap::new();
@@ -370,8 +366,6 @@ pub struct RemoveAuditTarget {}
#[async_trait::async_trait] #[async_trait::async_trait]
impl Operation for RemoveAuditTarget { impl Operation for RemoveAuditTarget {
async fn call(&self, req: S3Request<Body>, params: Params<'_, '_>) -> S3Result<S3Response<(StatusCode, Body)>> { async fn call(&self, req: S3Request<Body>, params: Params<'_, '_>) -> S3Result<S3Response<(StatusCode, Body)>> {
let span = Span::current();
let _enter = span.enter();
let (target_type, target_name) = extract_target_params(&params)?; let (target_type, target_name) = extract_target_params(&params)?;
authorize_audit_admin_request(&req, AdminAction::SetBucketTargetAction).await?; authorize_audit_admin_request(&req, AdminAction::SetBucketTargetAction).await?;
@@ -838,6 +832,13 @@ mod tests {
extract_block_between_markers(src, "impl Operation for ListAuditTargets", "pub struct RemoveAuditTarget"); extract_block_between_markers(src, "impl Operation for ListAuditTargets", "pub struct RemoveAuditTarget");
let delete_block = extract_block_between_markers(src, "impl Operation for RemoveAuditTarget", "#[cfg(test)]"); let delete_block = extract_block_between_markers(src, "impl Operation for RemoveAuditTarget", "#[cfg(test)]");
for block in [put_block, list_block, delete_block] {
assert!(
!block.contains(".enter()"),
"async audit handlers must rely on request-future instrumentation instead of holding span guards across awaits"
);
}
assert!( assert!(
put_block.contains("authorize_audit_admin_request(&req, AdminAction::SetBucketTargetAction).await?;"), put_block.contains("authorize_audit_admin_request(&req, AdminAction::SetBucketTargetAction).await?;"),
"audit target writes should require SetBucketTargetAction" "audit target writes should require SetBucketTargetAction"
+8 -9
View File
@@ -43,7 +43,7 @@ use s3s::{Body, S3Request, S3Response, S3Result, s3_error};
use serde::{Deserialize, Serialize}; use serde::{Deserialize, Serialize};
use std::collections::HashMap; use std::collections::HashMap;
use std::sync::LazyLock; use std::sync::LazyLock;
use tracing::{Span, error, info, warn}; use tracing::{error, info, warn};
const LOG_COMPONENT_ADMIN_API: &str = "admin_api"; const LOG_COMPONENT_ADMIN_API: &str = "admin_api";
@@ -333,8 +333,6 @@ pub struct NotificationTarget {}
#[async_trait::async_trait] #[async_trait::async_trait]
impl Operation for NotificationTarget { impl Operation for NotificationTarget {
async fn call(&self, req: S3Request<Body>, params: Params<'_, '_>) -> S3Result<S3Response<(StatusCode, Body)>> { async fn call(&self, req: S3Request<Body>, params: Params<'_, '_>) -> S3Result<S3Response<(StatusCode, Body)>> {
let span = Span::current();
let _enter = span.enter();
let (target_type, target_name) = extract_target_params(&params)?; let (target_type, target_name) = extract_target_params(&params)?;
let context = app_context_from_req(&req); let context = app_context_from_req(&req);
@@ -401,8 +399,6 @@ pub struct ListNotificationTargets {}
#[async_trait::async_trait] #[async_trait::async_trait]
impl Operation for ListNotificationTargets { impl Operation for ListNotificationTargets {
async fn call(&self, req: S3Request<Body>, _params: Params<'_, '_>) -> S3Result<S3Response<(StatusCode, Body)>> { async fn call(&self, req: S3Request<Body>, _params: Params<'_, '_>) -> S3Result<S3Response<(StatusCode, Body)>> {
let span = Span::current();
let _enter = span.enter();
authorize_notification_admin_request(&req, AdminAction::GetBucketTargetAction).await?; authorize_notification_admin_request(&req, AdminAction::GetBucketTargetAction).await?;
refresh_persisted_module_switches_from_store().await.map_err(|err| { refresh_persisted_module_switches_from_store().await.map_err(|err| {
warn!( warn!(
@@ -439,8 +435,6 @@ pub struct ListTargetsArns {}
#[async_trait::async_trait] #[async_trait::async_trait]
impl Operation for ListTargetsArns { impl Operation for ListTargetsArns {
async fn call(&self, req: S3Request<Body>, _params: Params<'_, '_>) -> S3Result<S3Response<(StatusCode, Body)>> { async fn call(&self, req: S3Request<Body>, _params: Params<'_, '_>) -> S3Result<S3Response<(StatusCode, Body)>> {
let span = Span::current();
let _enter = span.enter();
authorize_notification_admin_request(&req, AdminAction::GetBucketTargetAction).await?; authorize_notification_admin_request(&req, AdminAction::GetBucketTargetAction).await?;
if let Some(reason) = notification_target_operation_block_reason( if let Some(reason) = notification_target_operation_block_reason(
"querying notification target ARNs for bucket associations from the console", "querying notification target ARNs for bucket associations from the console",
@@ -485,8 +479,6 @@ pub struct RemoveNotificationTarget {}
#[async_trait::async_trait] #[async_trait::async_trait]
impl Operation for RemoveNotificationTarget { impl Operation for RemoveNotificationTarget {
async fn call(&self, req: S3Request<Body>, params: Params<'_, '_>) -> S3Result<S3Response<(StatusCode, Body)>> { async fn call(&self, req: S3Request<Body>, params: Params<'_, '_>) -> S3Result<S3Response<(StatusCode, Body)>> {
let span = Span::current();
let _enter = span.enter();
let (target_type, target_name) = extract_target_params(&params)?; let (target_type, target_name) = extract_target_params(&params)?;
let context = app_context_from_req(&req); let context = app_context_from_req(&req);
@@ -1006,6 +998,13 @@ mod tests {
extract_block_between_markers(src, "impl Operation for ListTargetsArns", "pub struct RemoveNotificationTarget"); extract_block_between_markers(src, "impl Operation for ListTargetsArns", "pub struct RemoveNotificationTarget");
let delete_block = extract_block_between_markers(src, "impl Operation for RemoveNotificationTarget", "fn extract_param"); let delete_block = extract_block_between_markers(src, "impl Operation for RemoveNotificationTarget", "fn extract_param");
for block in [put_block, list_block, arns_block, delete_block] {
assert!(
!block.contains(".enter()"),
"async notification handlers must rely on request-future instrumentation instead of holding span guards across awaits"
);
}
assert!( assert!(
put_block.contains("authorize_notification_admin_request(&req, AdminAction::SetBucketTargetAction).await?;"), put_block.contains("authorize_notification_admin_request(&req, AdminAction::SetBucketTargetAction).await?;"),
"notification target writes should require SetBucketTargetAction" "notification target writes should require SetBucketTargetAction"
+353 -19
View File
@@ -535,8 +535,13 @@ impl Operation for GetReplicationMetricsHandler {
let bucket_stats = cluster_replication_stats(bucket, app_context_from_req(&req)).await; let bucket_stats = cluster_replication_stats(bucket, app_context_from_req(&req)).await;
let data = serde_json::to_vec(&bucket_stats.replication_stats) // Same minio-go `replication.Metrics` wire shape as
.map_err(|_| S3Error::with_message(S3ErrorCode::InternalError, "serialize failed"))?; // `?replication-metrics` — the internal snake_case stats are the peer
// RPC wire format and must not leak here.
let data = serde_json::to_vec(&crate::admin::replication_metrics_wire::MetricsWire::from(
&bucket_stats.replication_stats,
))
.map_err(|_| S3Error::with_message(S3ErrorCode::InternalError, "serialize failed"))?;
let mut headers = HeaderMap::new(); let mut headers = HeaderMap::new();
headers.insert(CONTENT_TYPE, HeaderValue::from_static("application/json")); headers.insert(CONTENT_TYPE, HeaderValue::from_static("application/json"));
Ok(S3Response::with_headers((StatusCode::OK, Body::from(data)), headers)) Ok(S3Response::with_headers((StatusCode::OK, Body::from(data)), headers))
@@ -1153,7 +1158,7 @@ struct MrfResponse {
fn build_mrf_response( fn build_mrf_response(
bucket: String, bucket: String,
bucket_stats: &BucketStats, bucket_stats: &BucketStats,
durable: crate::admin::storage_api::replication::DurableMrfBacklog, durable: &crate::admin::storage_api::replication::DurableMrfBacklog,
) -> MrfResponse { ) -> MrfResponse {
let observation_scope = if bucket_stats.replication_stats.cluster_complete { let observation_scope = if bucket_stats.replication_stats.cluster_complete {
"cluster_aggregated" "cluster_aggregated"
@@ -1223,7 +1228,10 @@ fn build_mrf_response(
total_failed_size, total_failed_size,
queued_count: queued.count, queued_count: queued.count,
queued_size: queued.bytes, queued_size: queued.bytes,
per_object_entries_available: false, // The default (non-aggregate) response mode streams the durable
// backlog per object, so the enumerable API exists whenever the
// backlog is readable.
per_object_entries_available: durable.available,
runtime_stats_available: bucket_stats.replication_stats.provider_available, runtime_stats_available: bucket_stats.replication_stats.provider_available,
cluster_complete: bucket_stats.replication_stats.cluster_complete, cluster_complete: bucket_stats.replication_stats.cluster_complete,
observed_node_count: bucket_stats.replication_stats.observed_node_count, observed_node_count: bucket_stats.replication_stats.observed_node_count,
@@ -1235,23 +1243,165 @@ fn build_mrf_response(
} }
} }
/// One durable MRF backlog entry rendered for the default (madmin-compatible)
/// stream. Field names are the exact json tags of madmin-go `ReplicationMRF`
/// (replication-api.go), which `mc replicate backlog` decodes one JSON
/// document at a time. `Size` and `TargetARNs` are RustFS extension keys with
/// no madmin counterpart; Go decoders ignore unknown keys.
#[derive(Debug, Serialize)]
struct MrfEntryDocument {
/// The durable backlog is a cluster-shared ledger with no per-node
/// attribution, so the madmin `nodeName` tag is always empty.
#[serde(rename = "nodeName")]
node_name: String,
#[serde(rename = "bucket")]
bucket: String,
#[serde(rename = "object")]
object: String,
#[serde(rename = "versionId")]
version_id: String,
#[serde(rename = "retryCount")]
retry_count: i32,
#[serde(rename = "Size")]
size: i64,
#[serde(rename = "TargetARNs", skip_serializing_if = "Vec::is_empty")]
target_arns: Vec<String>,
}
/// Upper bound on the number of documents one stream response emits. The
/// durable ledger is not bounded by the in-memory pending cap (recovery can
/// persist far larger generations), and the body is buffered before send, so
/// an unbounded read could stage hundreds of MB per request. The handler
/// rejects a response beyond this bound instead of returning a partial 200.
const REPLICATION_MRF_MAX_STREAM_ENTRIES: usize = 10_000;
/// Project the durable backlog into madmin `ReplicationMRF` documents,
/// scoped to `bucket` when it is non-empty (madmin allows an empty bucket to
/// mean "across all buckets"), bounded by
/// [`REPLICATION_MRF_MAX_STREAM_ENTRIES`]. Returns the documents and whether
/// the backlog was truncated.
fn mrf_entry_documents(
bucket: &str,
durable: &crate::admin::storage_api::replication::DurableMrfBacklog,
) -> (Vec<MrfEntryDocument>, bool) {
let mut documents = Vec::new();
let mut truncated = false;
for entry in durable
.entries
.iter()
.filter(|entry| bucket.is_empty() || entry.bucket == bucket)
{
if documents.len() >= REPLICATION_MRF_MAX_STREAM_ENTRIES {
truncated = true;
break;
}
documents.push(MrfEntryDocument {
node_name: String::new(),
bucket: entry.bucket.clone(),
object: entry.object.clone(),
// Delete-marker purge entries track the marker version separately;
// fall back to it so those rows still carry a version identity.
// The nil UUID is RustFS's in-memory null-version sentinel and
// must leave as the S3 wire token, not a zero UUID.
version_id: entry
.version_id
.or(entry.delete_marker_version_id)
.map(|v| {
if v.is_nil() {
rustfs_filemeta::NULL_VERSION_ID.to_string()
} else {
v.to_string()
}
})
.unwrap_or_default(),
retry_count: entry.retry_count,
size: entry.size,
target_arns: entry.target_arns.clone(),
});
}
(documents, truncated)
}
/// Render the MRF backlog as a response body.
///
/// Default (madmin-compatible) mode emits one `ReplicationMRF` JSON document
/// per line with no envelope — madmin's `BucketReplicationMRF` reads the body
/// with a `json.Decoder` loop, so an envelope object would decode as a single
/// entry whose `"Bucket"` key case-insensitively matches
/// `ReplicationMRF.Bucket` (a phantom row in `mc replicate backlog`), and an
/// empty backlog must render an empty body so the loop ends on io.EOF with
/// zero rows.
///
/// `aggregate=true` (RustFS extension) keeps the enveloped counter shape;
/// backlog-source health (`RuntimeStatsAvailable`/`DurableBacklogAvailable`)
/// is only representable there — an unreadable ledger fails the stream
/// request outright in the handler (madmin only decodes the body of a 200,
/// so an empty stream would read as a healthy zero-row backlog).
fn render_mrf_backlog(
response: &MrfResponse,
durable: &crate::admin::storage_api::replication::DurableMrfBacklog,
aggregate: bool,
) -> Result<(Vec<u8>, bool), serde_json::Error> {
if aggregate {
return Ok((serde_json::to_vec(response)?, false));
}
let (documents, truncated) = mrf_entry_documents(&response.bucket, durable);
let mut data = Vec::new();
for entry in documents {
serde_json::to_writer(&mut data, &entry)?;
data.push(b'\n');
}
Ok((data, truncated))
}
fn ensure_complete_mrf_stream(truncated: bool) -> S3Result<()> {
if truncated {
return Err(S3Error::with_message(
S3ErrorCode::ServiceUnavailable,
"durable MRF backlog exceeds the stream limit; narrow the bucket scope or drain the backlog".to_string(),
));
}
Ok(())
}
/// `GET /v3/replication/mrf` /// `GET /v3/replication/mrf`
/// ///
/// Reports the failed-replication backlog (MinIO's MRF concept) for a bucket. /// Reports the failed-replication backlog (MinIO's MRF concept) for a bucket.
/// ///
/// Compatibility note: MinIO returns a stream of individual MRF entries. RustFS /// The default response is a madmin-compatible stream of `ReplicationMRF`
/// deliberately returns aggregate runtime and durable counters instead. /// documents built from the durable backlog ledger (in-memory failures that
/// `PerObjectEntriesAvailable` remains false until an enumerable API exists. /// have not been flushed yet — the persister runs every few seconds — are not
/// `PerTargetDurableEntriesAvailable` is false when the durable backlog includes /// visible). `?aggregate=true` (RustFS extension) returns the enveloped
/// older entries that cannot be attributed to a target. /// runtime + durable counter shape instead; `PerTargetDurableEntriesAvailable`
/// is false there when the durable backlog includes older entries that cannot
/// be attributed to a target.
///
/// The madmin `node` parameter is accepted but has no filtering effect: the
/// durable ledger is cluster-shared with no per-node attribution, so every
/// node serves the same (complete) backlog.
///
/// Authorization: the stream requires `admin:ReplicationDiff` (it enumerates
/// object names and version ids, MinIO parity); `?aggregate=true` carries no
/// object identities and requires only `admin:GetReplicationMetrics`.
pub struct ReplicationMrfHandler {} pub struct ReplicationMrfHandler {}
#[async_trait::async_trait] #[async_trait::async_trait]
impl Operation for ReplicationMrfHandler { impl Operation for ReplicationMrfHandler {
async fn call(&self, req: S3Request<Body>, _params: Params<'_, '_>) -> S3Result<S3Response<(StatusCode, Body)>> { async fn call(&self, req: S3Request<Body>, _params: Params<'_, '_>) -> S3Result<S3Response<(StatusCode, Body)>> {
validate_replication_admin_request(&req, AdminAction::GetReplicationMetricsAction).await?;
let queries = extract_query_params(&req.uri); let queries = extract_query_params(&req.uri);
let aggregate = queries.get("aggregate").map(String::as_str) == Some("true");
// The default stream enumerates object names and version ids, which
// a metrics-only principal must not see; gate it on the same action
// MinIO uses for this endpoint. The aggregate counters carry no
// object identities and keep the metrics action.
let action = if aggregate {
AdminAction::GetReplicationMetricsAction
} else {
AdminAction::ReplicationDiff
};
validate_replication_admin_request(&req, action).await?;
let Some(bucket) = queries.get("bucket").filter(|b| !b.is_empty()).cloned() else { let Some(bucket) = queries.get("bucket").filter(|b| !b.is_empty()).cloned() else {
return Err(s3_error!(InvalidRequest, "bucket is required")); return Err(s3_error!(InvalidRequest, "bucket is required"));
}; };
@@ -1275,14 +1425,48 @@ impl Operation for ReplicationMrfHandler {
return Err(ApiError::from(err).into()); return Err(ApiError::from(err).into());
} }
if let Some(node) = queries.get("node").filter(|node| !node.is_empty() && node.as_str() != "all") {
// The durable backlog ledger is cluster-shared with no per-node
// attribution, so a node-scoped request still sees the complete
// (superset) backlog.
debug!(node = %node, "replication mrf node filter has no effect on the cluster-shared backlog");
}
let durable = crate::admin::storage_api::replication::read_durable_mrf_backlog(store).await; let durable = crate::admin::storage_api::replication::read_durable_mrf_backlog(store).await;
let bucket_stats = cluster_replication_stats(&bucket, app_context_from_req(&req)).await; let bucket_stats = cluster_replication_stats(&bucket, app_context_from_req(&req)).await;
let response = build_mrf_response(bucket, &bucket_stats, durable); let response = build_mrf_response(bucket, &bucket_stats, &durable);
let data = serde_json::to_vec(&response) if !durable.available && !aggregate {
// The madmin stream has no envelope to carry source health, and
// madmin only decodes the body of a 200 — an empty stream would
// read as a clean, healthy zero-row backlog. Fail loudly instead;
// aggregate mode still reports the availability fields.
tracing::warn!(
bucket = %response.bucket,
"durable MRF backlog is unreadable; failing the stream request — use aggregate=true to see source health"
);
return Err(S3Error::with_message(
S3ErrorCode::ServiceUnavailable,
"durable MRF backlog is unreadable; retry, or use aggregate=true for source health".to_string(),
));
}
let (data, truncated) = render_mrf_backlog(&response, &durable, aggregate)
.map_err(|e| S3Error::with_message(S3ErrorCode::InternalError, format!("serialize failed: {e}")))?; .map_err(|e| S3Error::with_message(S3ErrorCode::InternalError, format!("serialize failed: {e}")))?;
let mut headers = HeaderMap::new(); let mut headers = HeaderMap::new();
headers.insert(CONTENT_TYPE, HeaderValue::from_static("application/json")); headers.insert(CONTENT_TYPE, HeaderValue::from_static("application/json"));
if truncated {
tracing::warn!(
event = "replication_mrf_stream_rejected",
component = "admin",
subsystem = "replication",
result = "rejected",
bucket = %response.bucket,
max_entries = REPLICATION_MRF_MAX_STREAM_ENTRIES,
"replication mrf stream exceeds the response limit"
);
}
ensure_complete_mrf_stream(truncated)?;
Ok(S3Response::with_headers((StatusCode::OK, Body::from(data)), headers)) Ok(S3Response::with_headers((StatusCode::OK, Body::from(data)), headers))
} }
} }
@@ -1292,7 +1476,8 @@ mod tests {
use super::{ use super::{
REMOTE_TARGET_UNSUPPORTED_FIELDS, REMOTE_TARGET_WRITABLE_FIELDS, RemoteTargetCredentialsRequest, RemoteTargetRequest, REMOTE_TARGET_UNSUPPORTED_FIELDS, REMOTE_TARGET_WRITABLE_FIELDS, RemoteTargetCredentialsRequest, RemoteTargetRequest,
ReplicationDiffEntry, SUPPORTED_REMOTE_TARGET_API, TargetUpdateOp, build_mrf_response, extract_query_params, ReplicationDiffEntry, SUPPORTED_REMOTE_TARGET_API, TargetUpdateOp, build_mrf_response, extract_query_params,
parse_remote_target_update_ops, render_replication_diff, unique_replication_peers, validate_remote_target_tls_settings, parse_remote_target_update_ops, render_mrf_backlog, render_replication_diff, unique_replication_peers,
validate_remote_target_tls_settings,
}; };
use crate::admin::storage_api::bucket::target::{BucketTarget, LatencyStat}; use crate::admin::storage_api::bucket::target::{BucketTarget, LatencyStat};
use crate::admin::storage_api::replication::{BucketStats, DurableMrfBacklog, MrfOpKind, MrfReplicateEntry}; use crate::admin::storage_api::replication::{BucketStats, DurableMrfBacklog, MrfOpKind, MrfReplicateEntry};
@@ -1510,7 +1695,7 @@ mod tests {
], ],
}; };
let response = build_mrf_response("bucket-a".to_string(), &stats, durable); let response = build_mrf_response("bucket-a".to_string(), &stats, &durable);
let json = serde_json::to_value(response).expect("MRF response should serialize"); let json = serde_json::to_value(response).expect("MRF response should serialize");
assert_eq!(json["TotalFailedCount"], 3); assert_eq!(json["TotalFailedCount"], 3);
@@ -1522,7 +1707,9 @@ mod tests {
assert_eq!(json["RuntimeStatsAvailable"], true); assert_eq!(json["RuntimeStatsAvailable"], true);
assert_eq!(json["ClusterComplete"], false); assert_eq!(json["ClusterComplete"], false);
assert_eq!(json["Targets"][0]["ObservationScope"], "partial_cluster"); assert_eq!(json["Targets"][0]["ObservationScope"], "partial_cluster");
assert_eq!(json["PerObjectEntriesAvailable"], false); // The bare stream enumerates the durable backlog per object, so a
// readable backlog advertises the enumerable API.
assert_eq!(json["PerObjectEntriesAvailable"], true);
assert_eq!(json["PerTargetDurableEntriesAvailable"], true); assert_eq!(json["PerTargetDurableEntriesAvailable"], true);
let targets = json["Targets"].as_array().expect("targets should serialize as an array"); let targets = json["Targets"].as_array().expect("targets should serialize as an array");
@@ -1569,7 +1756,7 @@ mod tests {
}], }],
}; };
let response = build_mrf_response("bucket-a".to_string(), &stats, durable); let response = build_mrf_response("bucket-a".to_string(), &stats, &durable);
let json = serde_json::to_value(response).expect("MRF response should serialize"); let json = serde_json::to_value(response).expect("MRF response should serialize");
assert_eq!(json["DurableBacklogAvailable"], true); assert_eq!(json["DurableBacklogAvailable"], true);
@@ -1587,7 +1774,7 @@ mod tests {
#[test] #[test]
fn mrf_response_distinguishes_unavailable_sources_from_valid_zero() { fn mrf_response_distinguishes_unavailable_sources_from_valid_zero() {
let unavailable = build_mrf_response("bucket-a".to_string(), &BucketStats::default(), DurableMrfBacklog::default()); let unavailable = build_mrf_response("bucket-a".to_string(), &BucketStats::default(), &DurableMrfBacklog::default());
let unavailable_json = serde_json::to_value(unavailable).expect("unavailable response should serialize"); let unavailable_json = serde_json::to_value(unavailable).expect("unavailable response should serialize");
assert_eq!(unavailable_json["RuntimeStatsAvailable"], false); assert_eq!(unavailable_json["RuntimeStatsAvailable"], false);
assert_eq!(unavailable_json["DurableBacklogAvailable"], false); assert_eq!(unavailable_json["DurableBacklogAvailable"], false);
@@ -1600,7 +1787,7 @@ mod tests {
let valid_empty = build_mrf_response( let valid_empty = build_mrf_response(
"bucket-a".to_string(), "bucket-a".to_string(),
&valid_empty_stats, &valid_empty_stats,
DurableMrfBacklog { &DurableMrfBacklog {
available: true, available: true,
entries: Vec::new(), entries: Vec::new(),
}, },
@@ -1613,6 +1800,153 @@ mod tests {
assert_eq!(valid_empty_json["PerTargetDurableEntriesAvailable"], true); assert_eq!(valid_empty_json["PerTargetDurableEntriesAvailable"], true);
} }
fn sample_durable_backlog() -> DurableMrfBacklog {
DurableMrfBacklog {
available: true,
entries: vec![
MrfReplicateEntry {
bucket: "bucket-a".to_string(),
object: "object-a".to_string(),
version_id: Some(uuid::Uuid::from_u128(7)),
retry_count: 2,
size: 250,
op: MrfOpKind::Object,
target_arns: vec!["arn-a".to_string()],
..Default::default()
},
MrfReplicateEntry {
bucket: "other-bucket".to_string(),
object: "object-b".to_string(),
version_id: None,
retry_count: 0,
size: 999,
op: MrfOpKind::Object,
target_arns: Vec::new(),
..Default::default()
},
],
}
}
/// madmin's `BucketReplicationMRF` decodes the body one `ReplicationMRF`
/// JSON document at a time; the default response must therefore be a bare
/// document stream with madmin's exact json tags, not an envelope.
#[test]
fn mrf_stream_renders_bare_madmin_documents() {
let durable = sample_durable_backlog();
let response = build_mrf_response("bucket-a".to_string(), &BucketStats::default(), &durable);
let (body, _) = render_mrf_backlog(&response, &durable, false).expect("stream body should serialize");
let text = String::from_utf8(body).expect("body should be utf-8");
let lines: Vec<&str> = text.lines().filter(|line| !line.trim().is_empty()).collect();
// Only the entry matching the requested bucket is streamed.
assert_eq!(lines.len(), 1, "expected one MRF document, got: {text}");
let doc: serde_json::Value = serde_json::from_str(lines[0]).expect("each line should be a JSON document");
assert_eq!(doc["bucket"], "bucket-a");
assert_eq!(doc["object"], "object-a");
assert_eq!(doc["versionId"], uuid::Uuid::from_u128(7).to_string());
assert_eq!(doc["retryCount"], 2);
// madmin `ReplicationMRF` has a `nodeName` tag; the durable backlog is
// cluster-shared, so RustFS reports an empty node name.
assert_eq!(doc["nodeName"], "");
// The envelope keys must not leak into the stream: a `"Bucket"` key
// would case-insensitively populate `ReplicationMRF.Bucket` and render
// a phantom row in `mc replicate backlog`.
assert!(doc.get("Bucket").is_none());
assert!(doc.get("Targets").is_none());
}
/// An empty backlog must produce an empty body: madmin's decoder loop then
/// terminates on io.EOF with zero rows instead of one phantom row.
#[test]
fn mrf_stream_renders_empty_body_for_no_entries() {
let durable = DurableMrfBacklog {
available: true,
entries: Vec::new(),
};
let response = build_mrf_response("bucket-a".to_string(), &BucketStats::default(), &durable);
let (body, _) = render_mrf_backlog(&response, &durable, false).expect("stream body should serialize");
assert!(
body.is_empty(),
"empty backlog must serialize to an empty body, got: {}",
String::from_utf8_lossy(&body)
);
}
/// `?aggregate=true` (RustFS extension) keeps the enveloped counter shape.
#[test]
fn mrf_aggregate_envelope_retains_counters() {
let durable = sample_durable_backlog();
let response = build_mrf_response("bucket-a".to_string(), &BucketStats::default(), &durable);
let (body, _) = render_mrf_backlog(&response, &durable, true).expect("aggregate body should serialize");
let json: serde_json::Value = serde_json::from_slice(&body).expect("aggregate body should be one JSON object");
assert_eq!(json["Bucket"], "bucket-a");
assert_eq!(json["DurableCount"], 1);
assert_eq!(json["DurableBacklogAvailable"], true);
// The bare stream is an enumerable per-object API, so the aggregate
// shell now truthfully advertises it whenever the backlog is readable.
assert_eq!(json["PerObjectEntriesAvailable"], true);
}
/// The nil UUID is RustFS's in-memory null-version sentinel; the wire
/// token is `null`, never the zero UUID (second review round).
#[test]
fn mrf_stream_maps_nil_version_to_null_token() {
let durable = DurableMrfBacklog {
available: true,
entries: vec![MrfReplicateEntry {
bucket: "bucket-a".to_string(),
object: "null-version-object".to_string(),
version_id: Some(uuid::Uuid::nil()),
retry_count: 1,
size: 10,
op: MrfOpKind::Object,
..Default::default()
}],
};
let response = build_mrf_response("bucket-a".to_string(), &BucketStats::default(), &durable);
let (body, truncated) = render_mrf_backlog(&response, &durable, false).expect("stream body should serialize");
assert!(!truncated);
let doc: serde_json::Value =
serde_json::from_str(String::from_utf8(body).expect("utf-8").lines().next().expect("one line"))
.expect("line should be a JSON document");
assert_eq!(doc["versionId"], "null");
}
/// The durable ledger is not bounded by the in-memory pending cap; the
/// stream must stop at the documented bound and signal truncation
/// (second review round).
#[test]
fn mrf_stream_truncates_at_the_documented_bound() {
let entries = (0..super::REPLICATION_MRF_MAX_STREAM_ENTRIES + 1)
.map(|index| MrfReplicateEntry {
bucket: "bucket-a".to_string(),
object: format!("object-{index}"),
retry_count: 1,
op: MrfOpKind::Object,
..Default::default()
})
.collect();
let durable = DurableMrfBacklog {
available: true,
entries,
};
let response = build_mrf_response("bucket-a".to_string(), &BucketStats::default(), &durable);
let (body, truncated) = render_mrf_backlog(&response, &durable, false).expect("stream body should serialize");
assert!(truncated, "one entry past the bound must signal truncation");
assert_eq!(
String::from_utf8(body).expect("utf-8").lines().count(),
super::REPLICATION_MRF_MAX_STREAM_ENTRIES
);
let error = super::ensure_complete_mrf_stream(truncated).expect_err("partial streams must not return 200");
assert_eq!(error.code(), &s3s::S3ErrorCode::ServiceUnavailable);
}
#[test] #[test]
fn test_extract_query_params_decodes_percent_encoded_values() { fn test_extract_query_params_decodes_percent_encoded_values() {
let uri: Uri = "/rustfs/admin/v3/list-remote-targets?bucket=foo%2Fbar&flag=a+b" let uri: Uri = "/rustfs/admin/v3/list-remote-targets?bucket=foo%2Fbar&flag=a+b"
File diff suppressed because it is too large Load Diff
@@ -29,6 +29,6 @@ impl Operation for RestLoadCredentialsHandler {
let issuer = IamTableCredentialIssuer::from_request(&req)?; let issuer = IamTableCredentialIssuer::from_request(&req)?;
let response = let response =
load_credentials_response(&store, &warehouse, &namespace, &table, &issuer, Some(&principal.credentials)).await?; load_credentials_response(&store, &warehouse, &namespace, &table, &issuer, Some(&principal.credentials)).await?;
build_json_response(StatusCode::OK, &response) build_sensitive_json_response(StatusCode::OK, &response)
} }
} }
File diff suppressed because it is too large Load Diff
@@ -125,7 +125,9 @@ impl Operation for RestLoadTableHandler {
ensure_table_bucket_enabled_from_extensions(&req.extensions, &warehouse).await?; ensure_table_bucket_enabled_from_extensions(&req.extensions, &warehouse).await?;
let metadata_backend = table_catalog_backend_from_extensions(&req.extensions)?; let metadata_backend = table_catalog_backend_from_extensions(&req.extensions)?;
let store = table_catalog_store_from_backend(metadata_backend.clone())?; let store = table_catalog_store_from_backend(metadata_backend.clone())?;
let response = load_table_response(&store, &metadata_backend, &warehouse, &namespace, &table).await?; let snapshot_selection = rest_table_snapshot_selection_from_query(&req.uri)?;
let mut response = load_table_response(&store, &metadata_backend, &warehouse, &namespace, &table).await?;
apply_rest_table_snapshot_selection(&mut response.metadata, snapshot_selection);
build_json_response(StatusCode::OK, &response) build_json_response(StatusCode::OK, &response)
} }
} }
@@ -158,7 +160,7 @@ impl Operation for RestCommitTableHandler {
let principal = authorize_table_catalog_resource_request(&req, &resource, AdminAction::CommitTableAction).await?; let principal = authorize_table_catalog_resource_request(&req, &resource, AdminAction::CommitTableAction).await?;
install_table_catalog_s3_request_info(&mut req, &principal)?; install_table_catalog_s3_request_info(&mut req, &principal)?;
ensure_table_bucket_enabled_from_extensions(&req.extensions, &warehouse).await?; ensure_table_bucket_enabled_from_extensions(&req.extensions, &warehouse).await?;
let request = read_json_body::<RestCommitTableRequest>(std::mem::take(&mut req.input)).await?; let request = read_rest_commit_table_request(std::mem::take(&mut req.input)).await?;
let metadata_backend = table_catalog_backend_from_extensions(&req.extensions)?; let metadata_backend = table_catalog_backend_from_extensions(&req.extensions)?;
let store = table_catalog_store_from_backend(metadata_backend.clone())?; let store = table_catalog_store_from_backend(metadata_backend.clone())?;
let commit_backend = TableCommitObjectBackend::for_request(metadata_backend, req); let commit_backend = TableCommitObjectBackend::for_request(metadata_backend, req);
@@ -178,6 +180,14 @@ impl Operation for RestDropTableHandler {
let table = table_name_from_params(&params)?; let table = table_name_from_params(&params)?;
let resource = TableCatalogResource::table(&warehouse, &namespace, &table); let resource = TableCatalogResource::table(&warehouse, &namespace, &table);
authorize_table_catalog_resource_request(&req, &resource, AdminAction::DeleteTableAction).await?; authorize_table_catalog_resource_request(&req, &resource, AdminAction::DeleteTableAction).await?;
let purge_requested = rest_purge_requested_from_query(&req.uri)?;
if purge_requested {
return Err(iceberg_rest_error(
ICEBERG_ERROR_UNSUPPORTED_OPERATION,
StatusCode::NOT_ACCEPTABLE,
"purgeRequested=true is not supported",
));
}
ensure_table_bucket_enabled_from_extensions(&req.extensions, &warehouse).await?; ensure_table_bucket_enabled_from_extensions(&req.extensions, &warehouse).await?;
let store = table_catalog_store_from_extensions(&req.extensions)?; let store = table_catalog_store_from_extensions(&req.extensions)?;
drop_table_in_store(&store, &warehouse, &namespace, &table).await?; drop_table_in_store(&store, &warehouse, &namespace, &table).await?;
File diff suppressed because it is too large Load Diff
@@ -43,8 +43,9 @@ impl Operation for RestCreateViewHandler {
let metadata_backend = table_catalog_backend_from_extensions(&req.extensions)?; let metadata_backend = table_catalog_backend_from_extensions(&req.extensions)?;
let store = table_catalog_store_from_backend(metadata_backend.clone())?; let store = table_catalog_store_from_backend(metadata_backend.clone())?;
let table_bucket_enabled = table_bucket_enabled_from_extensions(&req.extensions, &warehouse).await?; let table_bucket_enabled = table_bucket_enabled_from_extensions(&req.extensions, &warehouse).await?;
let publication_backend = TableCommitObjectBackend::preauthorized(metadata_backend);
let response = let response =
create_view_response(&store, &metadata_backend, &warehouse, &namespace, request, table_bucket_enabled).await?; create_view_response(&store, &publication_backend, &warehouse, &namespace, request, table_bucket_enabled).await?;
build_json_response(StatusCode::OK, &response) build_json_response(StatusCode::OK, &response)
} }
} }
@@ -87,17 +88,20 @@ pub struct RestReplaceViewHandler {}
#[async_trait::async_trait] #[async_trait::async_trait]
impl Operation for RestReplaceViewHandler { impl Operation for RestReplaceViewHandler {
async fn call(&self, req: S3Request<Body>, params: Params<'_, '_>) -> S3Result<S3Response<(StatusCode, Body)>> { async fn call(&self, mut req: S3Request<Body>, params: Params<'_, '_>) -> S3Result<S3Response<(StatusCode, Body)>> {
let warehouse = warehouse_from_params(&params)?; let warehouse = warehouse_from_params(&params)?;
let namespace = namespace_from_params(&params)?; let namespace = namespace_from_params(&params)?;
let view = view_name_from_params(&params)?; let view = view_name_from_params(&params)?;
let resource = TableCatalogResource::view(&warehouse, &namespace, &view); let resource = TableCatalogResource::view(&warehouse, &namespace, &view);
authorize_table_catalog_resource_request(&req, &resource, AdminAction::CommitTableAction).await?; let principal = authorize_table_catalog_resource_request(&req, &resource, AdminAction::CommitTableAction).await?;
install_table_catalog_s3_request_info(&mut req, &principal)?;
ensure_table_bucket_enabled_from_extensions(&req.extensions, &warehouse).await?; ensure_table_bucket_enabled_from_extensions(&req.extensions, &warehouse).await?;
let request = read_json_body::<RestCommitViewRequest>(req.input).await?; let request = read_rest_commit_view_request(std::mem::take(&mut req.input)).await?;
let metadata_backend = table_catalog_backend_from_extensions(&req.extensions)?; let metadata_backend = table_catalog_backend_from_extensions(&req.extensions)?;
let store = table_catalog_store_from_backend(metadata_backend.clone())?; let store = table_catalog_store_from_backend(metadata_backend.clone())?;
let response = replace_view_response(&store, &metadata_backend, &warehouse, &namespace, &view, request).await?; let commit_backend = TableCommitObjectBackend::for_request(metadata_backend, req);
let result = replace_view_response(&store, &commit_backend, &warehouse, &namespace, &view, request).await;
let response = commit_backend.finish(result).await?;
build_json_response(StatusCode::OK, &response) build_json_response(StatusCode::OK, &response)
} }
} }
+1
View File
@@ -17,6 +17,7 @@ mod auth;
pub mod console; pub mod console;
pub mod handlers; pub mod handlers;
mod plugin_contract; mod plugin_contract;
pub(crate) mod replication_metrics_wire;
// Contract inventory is validated by tests before later runtime integration. // Contract inventory is validated by tests before later runtime integration.
#[allow(dead_code)] #[allow(dead_code)]
pub(crate) mod route_policy; pub(crate) mod route_policy;
@@ -0,0 +1,673 @@
// Copyright 2024 RustFS Team
//
// Licensed under the Apache License, Version 2.0 (the "License");
// you may not use this file except in compliance with the License.
// You may obtain a copy of the License at
//
// http://www.apache.org/licenses/LICENSE-2.0
//
// Unless required by applicable law or agreed to in writing, software
// distributed under the License is distributed on an "AS IS" BASIS,
// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
// See the License for the specific language governing permissions and
// limitations under the License.
//! Serialize-only wire projections of the internal replication statistics
//! onto the minio-go `replication.Metrics` / `replication.MetricsV2` json
//! shapes consumed by `mc replicate status` (`?replication-metrics[=2]` and
//! the admin `replicationmetrics` endpoint).
//!
//! Red line: the internal `BucketStats` family in
//! `crates/replication/src/stats.rs` is ALSO the intra-cluster peer-RPC wire
//! format — `node_service.rs` encodes it with `rmp_serde::to_vec_named`, so
//! its Rust field names travel between nodes as msgpack map keys. Renaming
//! those serde names would break mixed-version clusters mid rolling upgrade.
//! All madmin/minio-go interop therefore happens in these DTOs; never add
//! `#[serde(rename)]` to the internal structs instead.
//!
//! Field names below are the exact json tags of minio-go
//! `pkg/replication/replication.go` (v7.0.91). Keys minio-go does not know
//! are RustFS extensions; Go decoders ignore unknown keys. `max`/`peak` are
//! both emitted for the queue peak because the MinIO server writes `max`
//! while minio-go reads `peak` (an upstream drift); emitting both keeps every
//! decoder working.
use serde::Serialize;
use std::collections::HashMap;
use std::time::Duration;
use crate::admin::storage_api::replication::{
BucketReplicationStat as InternalReplicationStat, BucketReplicationStats as InternalReplicationStats, BucketStats,
InQueueMetric as InternalInQueueMetric, XferStats as InternalXferStats,
};
/// minio-go `replication.RStat`.
#[derive(Debug, Default, Clone, Copy, Serialize)]
pub(crate) struct RStatWire {
#[serde(rename = "count")]
pub count: f64,
#[serde(rename = "bytes")]
pub bytes: i64,
}
/// minio-go `replication.TimedErrStats`.
#[derive(Debug, Default, Clone, Copy, Serialize)]
pub(crate) struct TimedErrStatsWire {
#[serde(rename = "lastMinute")]
pub last_minute: RStatWire,
#[serde(rename = "lastHour")]
pub last_hour: RStatWire,
#[serde(rename = "totals")]
pub totals: RStatWire,
}
impl TimedErrStatsWire {
fn add(self, other: TimedErrStatsWire) -> TimedErrStatsWire {
fn add(a: RStatWire, b: RStatWire) -> RStatWire {
RStatWire {
count: a.count + b.count,
bytes: a.bytes.saturating_add(b.bytes),
}
}
TimedErrStatsWire {
last_minute: add(self.last_minute, other.last_minute),
last_hour: add(self.last_hour, other.last_hour),
totals: add(self.totals, other.totals),
}
}
}
/// minio-go `replication.QStat`.
#[derive(Debug, Default, Clone, Copy, Serialize)]
pub(crate) struct QStatWire {
#[serde(rename = "count")]
pub count: f64,
#[serde(rename = "bytes")]
pub bytes: f64,
}
/// minio-go `replication.InQueueMetric`, with the queue peak emitted under
/// both `peak` (minio-go tag) and `max` (MinIO server tag).
#[derive(Debug, Default, Clone, Copy, Serialize)]
pub(crate) struct InQueueMetricWire {
#[serde(rename = "curr")]
pub curr: QStatWire,
#[serde(rename = "avg")]
pub avg: QStatWire,
#[serde(rename = "max")]
pub max: QStatWire,
#[serde(rename = "peak")]
pub peak: QStatWire,
}
impl From<&InternalInQueueMetric> for InQueueMetricWire {
fn from(metric: &InternalInQueueMetric) -> Self {
fn qstat(bytes: i64, count: i64) -> QStatWire {
QStatWire {
count: count as f64,
bytes: bytes as f64,
}
}
let peak = qstat(metric.max.bytes, metric.max.count);
InQueueMetricWire {
curr: qstat(metric.curr.bytes, metric.curr.count),
avg: qstat(metric.avg.bytes, metric.avg.count),
max: peak,
peak,
}
}
}
/// minio-go `replication.XferStats`.
#[derive(Debug, Default, Clone, Copy, Serialize)]
pub(crate) struct XferStatsWire {
#[serde(rename = "avgRate")]
pub avg_rate: f64,
#[serde(rename = "peakRate")]
pub peak_rate: f64,
#[serde(rename = "currRate")]
pub curr_rate: f64,
}
#[derive(Default)]
struct XferStatsAverage {
sum: XferStatsWire,
active: u32,
}
impl XferStatsAverage {
fn add_active(&mut self, stats: XferStatsWire) {
if stats.peak_rate <= 0.0 {
return;
}
self.add_raw(stats);
self.active += 1;
}
fn add_raw(&mut self, stats: XferStatsWire) {
self.sum.avg_rate += stats.avg_rate;
self.sum.curr_rate += stats.curr_rate;
self.sum.peak_rate = self.sum.peak_rate.max(stats.peak_rate);
}
fn finish(self) -> XferStatsWire {
let active = self.active;
self.finish_with_divisor(active)
}
fn finish_with_divisor(self, divisor: u32) -> XferStatsWire {
if divisor == 0 {
return self.sum;
}
XferStatsWire {
avg_rate: self.sum.avg_rate / f64::from(divisor),
peak_rate: self.sum.peak_rate,
curr_rate: self.sum.curr_rate / f64::from(divisor),
}
}
}
impl From<&InternalXferStats> for XferStatsWire {
fn from(stats: &InternalXferStats) -> Self {
XferStatsWire {
avg_rate: stats.avg,
peak_rate: stats.peak,
curr_rate: stats.curr,
}
}
}
/// minio-go `replication.WorkerStat`. RustFS does not track per-bucket worker
/// occupancy yet, so this always reports zeros.
#[derive(Debug, Default, Clone, Copy, Serialize)]
pub(crate) struct WorkerStatWire {
#[serde(rename = "curr")]
pub curr: i32,
#[serde(rename = "avg")]
pub avg: f32,
#[serde(rename = "max")]
pub max: i32,
}
/// minio-go `replication.ReplMRFStats`. RustFS does not track the 5-minute /
/// dropped MRF windows, so this always reports zeros; the durable backlog is
/// enumerable via `/v3/replication/mrf` instead.
#[derive(Debug, Default, Clone, Copy, Serialize)]
pub(crate) struct ReplMrfStatsWire {
#[serde(rename = "failedCount_last5min")]
pub last_failed_count: u64,
#[serde(rename = "droppedCount_since_uptime")]
pub total_dropped_count: u64,
#[serde(rename = "droppedBytes_since_uptime")]
pub total_dropped_bytes: u64,
}
/// minio-go `replication.CounterSummary`.
#[derive(Debug, Default, Clone, Copy, Serialize)]
pub(crate) struct CounterSummaryWire {
#[serde(rename = "last1hr")]
pub last1hr: u64,
#[serde(rename = "last1m")]
pub last1m: u64,
#[serde(rename = "total")]
pub total: u64,
}
/// minio-go `replication.TargetMetrics` (one remote target / ARN).
#[derive(Debug, Default, Serialize)]
pub(crate) struct TargetMetricsWire {
#[serde(rename = "replicationCount")]
pub replicated_count: i64,
#[serde(rename = "completedReplicationSize")]
pub replicated_size: i64,
/// Bandwidth limit for this target. The tag says "bits" but both MinIO
/// and minio-go treat the value as bytes/sec; keep bytes/sec.
#[serde(rename = "limitInBits")]
pub bandwidth_limit_bytes_per_sec: i64,
#[serde(rename = "currentBandwidth")]
pub current_bandwidth_bytes_per_sec: f64,
#[serde(rename = "failed")]
pub failed: TimedErrStatsWire,
#[serde(rename = "failedReplicationSize")]
pub failed_size: i64,
#[serde(rename = "failedReplicationCount")]
pub failed_count: i64,
}
fn target_timed_err_stats(stat: &InternalReplicationStat) -> TimedErrStatsWire {
// Cluster aggregation merges FailStats without the process-local samples,
// so the serializable window snapshots (refreshed at each node's
// collection point, summed by merge) are authoritative here; the live
// samples only ever agree with or lag them, so take the larger.
let sampled_minute = stat.fail_stats.recent_since(Duration::from_secs(60));
let sampled_hour = stat.fail_stats.recent_since(Duration::from_secs(3600));
let window = |sampled_count: i64, sampled_size: i64, snapshot_count: i64, snapshot_size: i64| RStatWire {
count: sampled_count.max(snapshot_count) as f64,
bytes: sampled_size.max(snapshot_size),
};
TimedErrStatsWire {
last_minute: window(
sampled_minute.count,
sampled_minute.size,
stat.fail_stats.last_minute.count,
stat.fail_stats.last_minute.size,
),
last_hour: window(
sampled_hour.count,
sampled_hour.size,
stat.fail_stats.last_hour.count,
stat.fail_stats.last_hour.size,
),
totals: RStatWire {
count: stat.failed.count as f64,
bytes: stat.failed.size,
},
}
}
impl From<&InternalReplicationStat> for TargetMetricsWire {
fn from(stat: &InternalReplicationStat) -> Self {
TargetMetricsWire {
replicated_count: stat.replicated_count,
replicated_size: stat.replicated_size,
bandwidth_limit_bytes_per_sec: stat.bandwidth_limit_bytes_per_sec,
current_bandwidth_bytes_per_sec: stat.current_bandwidth_bytes_per_sec,
failed: target_timed_err_stats(stat),
failed_size: stat.failed.size,
failed_count: stat.failed.count,
}
}
}
/// minio-go `replication.Metrics` — the `currStats` member of `MetricsV2` and
/// the whole v1 response body. The trailing snake_case fields are RustFS
/// source-health extension keys (ignored by Go decoders) carried over from
/// the previous response shape.
#[derive(Debug, Default, Serialize)]
pub(crate) struct MetricsWire {
#[serde(rename = "Stats")]
pub stats: HashMap<String, TargetMetricsWire>,
#[serde(rename = "completedReplicationSize")]
pub replicated_size: i64,
#[serde(rename = "replicaSize")]
pub replica_size: i64,
#[serde(rename = "replicaCount")]
pub replica_count: i64,
#[serde(rename = "replicationCount")]
pub replicated_count: i64,
#[serde(rename = "failed")]
pub failed: TimedErrStatsWire,
#[serde(rename = "queued")]
pub queued: InQueueMetricWire,
// RustFS extension keys (source health of the aggregation).
pub provider_available: bool,
pub cluster_complete: bool,
pub observed_node_count: u32,
pub expected_node_count: u32,
}
impl From<&InternalReplicationStats> for MetricsWire {
fn from(stats: &InternalReplicationStats) -> Self {
let mut failed = TimedErrStatsWire::default();
let mut targets = HashMap::with_capacity(stats.stats.len());
for (arn, stat) in &stats.stats {
let target = TargetMetricsWire::from(stat);
failed = failed.add(target.failed);
targets.insert(arn.clone(), target);
}
MetricsWire {
stats: targets,
replicated_size: stats.replicated_size,
replica_size: stats.replica_size,
replica_count: stats.replica_count,
replicated_count: stats.replicated_count,
failed,
queued: InQueueMetricWire::from(&stats.q_stat),
provider_available: stats.provider_available,
cluster_complete: stats.cluster_complete,
observed_node_count: stats.observed_node_count,
expected_node_count: stats.expected_node_count,
}
}
}
/// minio-go `replication.ReplQNodeStats`.
#[derive(Debug, Default, Serialize)]
pub(crate) struct ReplQNodeStatsWire {
#[serde(rename = "nodeName")]
pub node_name: String,
#[serde(rename = "uptime")]
pub uptime: i64,
#[serde(rename = "activeWorkers")]
pub workers: WorkerStatWire,
#[serde(rename = "transferSummary")]
pub xfer_stats: XferSummaryWire,
#[serde(rename = "tgtTransferStats")]
pub tgt_xfer_stats: TargetXferSummaryWire,
#[serde(rename = "queueStats")]
pub q_stats: InQueueMetricWire,
#[serde(rename = "mrfStats")]
pub mrf_stats: ReplMrfStatsWire,
#[serde(rename = "retries")]
pub retries: CounterSummaryWire,
#[serde(rename = "errors")]
pub errors: CounterSummaryWire,
}
/// minio-go `replication.ReplQueueStats`.
#[derive(Debug, Default, Serialize)]
pub(crate) struct ReplQueueStatsWire {
#[serde(rename = "nodes")]
pub nodes: Vec<ReplQNodeStatsWire>,
}
/// minio-go `replication.MetricsV2` — the `?replication-metrics=2` body.
#[derive(Debug, Default, Serialize)]
pub(crate) struct MetricsV2Wire {
#[serde(rename = "uptime")]
pub uptime: i64,
#[serde(rename = "currStats")]
pub current_stats: MetricsWire,
#[serde(rename = "queueStats")]
pub queue_stats: ReplQueueStatsWire,
#[serde(rename = "downtimeInfo")]
pub downtime_info: HashMap<String, serde_json::Value>,
}
/// `transferSummary` map keyed by minio-go `MetricName` (Large/Small/Total).
type XferSummaryWire = HashMap<&'static str, XferStatsWire>;
/// `tgtTransferStats` map keyed by target ARN.
type TargetXferSummaryWire = HashMap<String, XferSummaryWire>;
fn transfer_summaries(stats: &InternalReplicationStats) -> (XferSummaryWire, TargetXferSummaryWire) {
let mut per_target: TargetXferSummaryWire = HashMap::new();
let mut large_summary = XferStatsAverage::default();
let mut small_summary = XferStatsAverage::default();
let mut total_summary = XferStatsAverage::default();
let mut active_targets = 0;
for (arn, stat) in &stats.stats {
let large = XferStatsWire::from(&stat.xfer_rate_lrg);
let small = XferStatsWire::from(&stat.xfer_rate_sml);
let mut target_total = XferStatsAverage::default();
target_total.add_active(large);
target_total.add_active(small);
let total = target_total.finish();
per_target.insert(arn.clone(), HashMap::from([("Large", large), ("Small", small), ("Total", total)]));
if large.peak_rate > 0.0 || small.peak_rate > 0.0 {
active_targets += 1;
large_summary.add_raw(large);
small_summary.add_raw(small);
total_summary.add_raw(large);
total_summary.add_raw(small);
}
}
let summary = HashMap::from([
("Large", large_summary.finish_with_divisor(active_targets)),
("Small", small_summary.finish_with_divisor(active_targets)),
("Total", total_summary.finish_with_divisor(active_targets)),
]);
(summary, per_target)
}
impl MetricsV2Wire {
/// Project the aggregated internal stats onto the `MetricsV2` shape.
///
/// The aggregation path leaves `queue_stats.nodes` empty today, so a
/// single node entry is synthesized from the bucket queue snapshot —
/// `mc replicate status` derives its queue/worker panels from
/// `queueStats.nodes` and treats an empty list as "no data".
pub(crate) fn from_stats(bucket_stats: &BucketStats, node_name: &str) -> Self {
let (xfer_stats, tgt_xfer_stats) = transfer_summaries(&bucket_stats.replication_stats);
let mut nodes: Vec<ReplQNodeStatsWire> = bucket_stats
.queue_stats
.nodes
.iter()
.map(|node| ReplQNodeStatsWire {
node_name: node_name.to_string(),
uptime: bucket_stats.uptime,
q_stats: InQueueMetricWire::from(&node.q_stats),
..Default::default()
})
.collect();
if nodes.is_empty() {
nodes.push(ReplQNodeStatsWire {
node_name: node_name.to_string(),
uptime: bucket_stats.uptime,
q_stats: InQueueMetricWire::from(&bucket_stats.replication_stats.q_stat),
xfer_stats: xfer_stats.clone(),
tgt_xfer_stats: tgt_xfer_stats.clone(),
..Default::default()
});
} else {
// Attach the transfer summaries to the first node; the internal
// snapshot does not attribute transfer rates per node.
if let Some(first) = nodes.first_mut() {
first.xfer_stats = xfer_stats.clone();
first.tgt_xfer_stats = tgt_xfer_stats.clone();
}
}
MetricsV2Wire {
uptime: bucket_stats.uptime,
current_stats: MetricsWire::from(&bucket_stats.replication_stats),
queue_stats: ReplQueueStatsWire { nodes },
downtime_info: HashMap::new(),
}
}
}
#[cfg(test)]
mod tests {
use super::*;
fn sample_bucket_stats() -> BucketStats {
let mut stats = BucketStats {
uptime: 42,
..Default::default()
};
stats.replication_stats.replica_count = 2;
stats.replication_stats.replica_size = 128;
stats.replication_stats.replicated_count = 9;
stats.replication_stats.replicated_size = 4096;
let target = stats
.replication_stats
.stats
.entry("arn:minio:replication::t:b".to_string())
.or_default();
target.replicated_count = 9;
target.replicated_size = 4096;
target.failed.count = 3;
target.failed.size = 900;
target.bandwidth_limit_bytes_per_sec = 1024;
target.current_bandwidth_bytes_per_sec = 512.5;
stats
.replication_stats
.q_stat
.curr
.now_count
.store(4, std::sync::atomic::Ordering::Relaxed);
stats
.replication_stats
.q_stat
.curr
.now_bytes
.store(1200, std::sync::atomic::Ordering::Relaxed);
stats.replication_stats.q_stat = stats.replication_stats.q_stat.snapshot();
stats
}
#[test]
fn metrics_wire_matches_minio_go_tags() {
let stats = sample_bucket_stats();
let json = serde_json::to_value(MetricsWire::from(&stats.replication_stats)).expect("v1 wire should serialize");
assert_eq!(json["replicaCount"], 2);
assert_eq!(json["replicaSize"], 128);
assert_eq!(json["replicationCount"], 9);
assert_eq!(json["completedReplicationSize"], 4096);
assert_eq!(json["queued"]["curr"]["count"], 4.0);
assert_eq!(json["queued"]["curr"]["bytes"], 1200.0);
let target = &json["Stats"]["arn:minio:replication::t:b"];
assert_eq!(target["replicationCount"], 9);
assert_eq!(target["completedReplicationSize"], 4096);
assert_eq!(target["limitInBits"], 1024);
assert_eq!(target["currentBandwidth"], 512.5);
// failed is the madmin TimedErrStats envelope, not the internal
// {count,size} pair.
assert_eq!(target["failed"]["totals"]["count"], 3.0);
assert_eq!(target["failed"]["totals"]["bytes"], 900);
assert!(target["failed"].get("count").is_none());
// Aggregate failed mirrors the per-target totals.
assert_eq!(json["failed"]["totals"]["count"], 3.0);
}
#[test]
fn metrics_v2_wire_synthesizes_queue_node() {
let stats = sample_bucket_stats();
let json = serde_json::to_value(MetricsV2Wire::from_stats(&stats, "node-1:9000")).expect("v2 wire should serialize");
assert_eq!(json["uptime"], 42);
assert_eq!(json["currStats"]["replicaCount"], 2);
let node = &json["queueStats"]["nodes"][0];
assert_eq!(node["nodeName"], "node-1:9000");
assert_eq!(node["uptime"], 42);
assert_eq!(node["queueStats"]["curr"]["count"], 4.0);
// The queue peak is emitted under both the minio-go tag (`peak`) and
// the MinIO server tag (`max`).
assert_eq!(node["queueStats"]["peak"], node["queueStats"]["max"]);
assert!(node["activeWorkers"].get("curr").is_some());
assert!(node["transferSummary"].get("Total").is_some());
assert_eq!(json["downtimeInfo"], serde_json::json!({}));
}
/// minio-go's transferSummary labels mean >= 128 MiB for Large; the
/// producer must bin on the same boundary (MIN_LARGE_OBJ_SIZE, shared
/// with the worker-pool split), or a 2 MiB replication shows under Large
/// while Small stays zero.
#[test]
fn transfer_summary_bins_on_the_128_mib_boundary() {
const MIB: i64 = 1024 * 1024;
let mut stats = BucketStats::default();
let stat = stats
.replication_stats
.stats
.entry("arn:minio:replication::t:b".to_string())
.or_default();
stat.update_xfer_rate(2 * MIB, std::time::Duration::from_secs(1));
stat.update_xfer_rate(127 * MIB, std::time::Duration::from_secs(1));
stat.update_xfer_rate(128 * MIB, std::time::Duration::from_secs(1));
let json = serde_json::to_value(MetricsV2Wire::from_stats(&stats, "node-1")).expect("v2 wire should serialize");
let summary = &json["queueStats"]["nodes"][0]["tgtTransferStats"]["arn:minio:replication::t:b"];
let small_peak = summary["Small"]["peakRate"].as_f64().expect("Small peakRate");
let large_peak = summary["Large"]["peakRate"].as_f64().expect("Large peakRate");
assert!(
(small_peak - (127 * MIB) as f64).abs() < 1.0,
"2 MiB and 127 MiB transfers must bin as Small (peak {small_peak})"
);
assert!(
(large_peak - (128 * MIB) as f64).abs() < 1.0,
"exactly 128 MiB must bin as Large (peak {large_peak})"
);
}
#[test]
fn transfer_summaries_average_active_bins_and_targets() {
let mut stats = BucketStats::default();
let first = stats.replication_stats.stats.entry("target-a".to_string()).or_default();
first.xfer_rate_sml.avg = 50.0;
first.xfer_rate_sml.curr = 40.0;
first.xfer_rate_sml.peak = 60.0;
first.xfer_rate_lrg.avg = 100.0;
first.xfer_rate_lrg.curr = 80.0;
first.xfer_rate_lrg.peak = 120.0;
let second = stats.replication_stats.stats.entry("target-b".to_string()).or_default();
second.xfer_rate_sml.avg = 30.0;
second.xfer_rate_sml.curr = 20.0;
second.xfer_rate_sml.peak = 40.0;
let json = serde_json::to_value(MetricsV2Wire::from_stats(&stats, "node-1")).expect("v2 wire should serialize");
let node = &json["queueStats"]["nodes"][0];
let target_a = &node["tgtTransferStats"]["target-a"]["Total"];
assert_eq!(target_a["avgRate"], 75.0);
assert_eq!(target_a["currRate"], 60.0);
assert_eq!(target_a["peakRate"], 120.0);
let summary = &node["transferSummary"];
assert_eq!(summary["Small"]["avgRate"], 40.0);
assert_eq!(summary["Small"]["currRate"], 30.0);
assert_eq!(summary["Large"]["avgRate"], 50.0);
assert_eq!(summary["Total"]["avgRate"], 90.0);
assert_eq!(summary["Total"]["currRate"], 70.0);
assert_eq!(summary["Total"]["peakRate"], 120.0);
}
/// Review regression: both metrics endpoints aggregate first, and the
/// FailStats merge drops the process-local samples — the rolling windows
/// must survive a peer-RPC round trip plus aggregation and still reach
/// the wire body.
#[test]
fn failure_windows_survive_aggregation_before_serialization() {
// Node A: live failure; the windows are stamped at the collection
// point (get_latest_replication_stats calls refresh_windows before
// the stats cross the wire), never on the failure hot path.
let mut node_a = crate::admin::storage_api::replication::BucketReplicationStat::default();
node_a.fail_stats.add_size(512, None::<&std::io::Error>);
node_a.fail_stats.refresh_windows();
node_a.failed = node_a.fail_stats.to_metric();
// Node A's stats cross the peer RPC wire: the samples are dropped,
// the window snapshots travel.
let encoded = rmp_serde::to_vec_named(&node_a).expect("stat should encode");
let remote: crate::admin::storage_api::replication::BucketReplicationStat =
rmp_serde::from_slice(&encoded).expect("stat should decode");
// Aggregation merges the remote stat with an empty local one.
let merged_fail = remote.fail_stats.merge(&Default::default());
let aggregated = crate::admin::storage_api::replication::BucketReplicationStat {
failed: merged_fail.to_metric(),
fail_stats: merged_fail,
..Default::default()
};
let mut stats = BucketStats::default();
stats
.replication_stats
.stats
.insert("arn:minio:replication::t:b".to_string(), aggregated);
let json = serde_json::to_value(MetricsWire::from(&stats.replication_stats)).expect("wire should serialize");
let failed = &json["Stats"]["arn:minio:replication::t:b"]["failed"];
assert_eq!(failed["totals"]["count"], 1.0);
assert_eq!(
failed["lastMinute"]["count"], 1.0,
"the rolling minute window must survive RPC + aggregation"
);
assert_eq!(failed["lastMinute"]["bytes"], 512);
assert_eq!(failed["lastHour"]["count"], 1.0);
}
/// Pin the intra-cluster peer-RPC wire format of the internal stats: it
/// is msgpack with the Rust field names as map keys
/// (`rmp_serde::to_vec_named` in node_service.rs). If someone "fixes"
/// the interop bug by renaming the internal serde fields instead of using
/// these DTOs, this test fails and points them here.
#[test]
fn internal_bucket_stats_rpc_wire_stays_snake_case() {
let stats = sample_bucket_stats();
let encoded = rmp_serde::to_vec_named(&stats).expect("internal stats should encode");
let value: serde_json::Value = rmp_serde::from_slice(&encoded).expect("named msgpack should decode generically");
assert!(
value.get("replication_stats").is_some(),
"peer RPC key replication_stats must not be renamed"
);
assert!(value["replication_stats"].get("q_stat").is_some());
assert!(value.get("queue_stats").is_some());
assert!(value.get("proxy_stats").is_some());
let decoded: BucketStats = rmp_serde::from_slice(&encoded).expect("round-trip through the peer RPC wire");
assert_eq!(decoded.replication_stats.replica_count, 2);
}
}
+4 -1
View File
@@ -1459,10 +1459,13 @@ pub const ADMIN_ROUTE_POLICY_SPECS: &[AdminRouteSpec] = &[
REPLICATION_DIFF, REPLICATION_DIFF,
RouteRiskLevel::Sensitive, RouteRiskLevel::Sensitive,
), ),
// The default stream enumerates object names/version ids and requires
// ReplicationDiff (MinIO parity); only ?aggregate=true relaxes to
// GetReplicationMetrics in the handler.
admin( admin(
HttpMethod::Get, HttpMethod::Get,
"/rustfs/admin/v3/replication/mrf", "/rustfs/admin/v3/replication/mrf",
GET_REPLICATION_METRICS, REPLICATION_DIFF,
RouteRiskLevel::Sensitive, RouteRiskLevel::Sensitive,
), ),
]; ];
+67 -13
View File
@@ -1548,7 +1548,8 @@ async fn build_replication_metrics_response(
let bucket_stats = apply_replication_metrics_bandwidth_report(bucket_stats, collect_replication_metrics_bandwidth(bucket)); let bucket_stats = apply_replication_metrics_bandwidth_report(bucket_stats, collect_replication_metrics_bandwidth(bucket));
let bucket_stats = apply_replication_metrics_runtime_fields(bucket_stats, route, replication_metrics_uptime_seconds()); let bucket_stats = apply_replication_metrics_runtime_fields(bucket_stats, route, replication_metrics_uptime_seconds());
let body = serialize_replication_metrics_body(&bucket_stats, route)?; let node_name = crate::runtime_sources::current_local_node_name().await.unwrap_or_default();
let body = serialize_replication_metrics_body(&bucket_stats, route, &node_name)?;
let mut resp = S3Response::with_status(Body::from(body), StatusCode::OK); let mut resp = S3Response::with_status(Body::from(body), StatusCode::OK);
resp.headers resp.headers
@@ -1608,12 +1609,24 @@ fn apply_replication_metrics_runtime_fields(
bucket_stats bucket_stats
} }
fn serialize_replication_metrics_body(bucket_stats: &BucketStats, route: ReplicationExtRoute) -> S3Result<Vec<u8>> { /// Serialize the metrics body in the minio-go wire shapes
/// (`replication.Metrics` for v1, `replication.MetricsV2` for v2). The
/// internal `BucketStats` serde names are the intra-cluster peer-RPC wire
/// format and must never appear here — see
/// `crate::admin::replication_metrics_wire`.
fn serialize_replication_metrics_body(
bucket_stats: &BucketStats,
route: ReplicationExtRoute,
node_name: &str,
) -> S3Result<Vec<u8>> {
use crate::admin::replication_metrics_wire::{MetricsV2Wire, MetricsWire};
match route { match route {
ReplicationExtRoute::MetricsV1 => { ReplicationExtRoute::MetricsV1 => {
serde_json::to_vec(&bucket_stats.replication_stats).map_err(|e| s3_error!(InternalError, "{e}")) serde_json::to_vec(&MetricsWire::from(&bucket_stats.replication_stats)).map_err(|e| s3_error!(InternalError, "{e}"))
}
ReplicationExtRoute::MetricsV2 => {
serde_json::to_vec(&MetricsV2Wire::from_stats(bucket_stats, node_name)).map_err(|e| s3_error!(InternalError, "{e}"))
} }
ReplicationExtRoute::MetricsV2 => serde_json::to_vec(bucket_stats).map_err(|e| s3_error!(InternalError, "{e}")),
ReplicationExtRoute::Check | ReplicationExtRoute::ResetStart | ReplicationExtRoute::ResetStatus => { ReplicationExtRoute::Check | ReplicationExtRoute::ResetStart | ReplicationExtRoute::ResetStatus => {
Err(s3_error!(InternalError, "invalid route for metrics response")) Err(s3_error!(InternalError, "invalid route for metrics response"))
} }
@@ -4147,22 +4160,37 @@ mod tests {
assert!(err.message().unwrap_or_default().contains("rule-stale")); assert!(err.message().unwrap_or_default().contains("rule-stale"));
} }
/// The v1 body must decode into minio-go `replication.Metrics` (exact
/// json tags); Go's decoder matches case-insensitively but does not
/// ignore underscores, so the internal snake_case names read as all-zero.
#[test] #[test]
fn serialize_replication_metrics_body_v1_returns_replication_stats_only() { fn serialize_replication_metrics_body_v1_returns_minio_go_metrics_shape() {
let mut stats = BucketStats { let mut stats = BucketStats {
uptime: 99, uptime: 99,
..Default::default() ..Default::default()
}; };
stats.replication_stats.replica_count = 7; stats.replication_stats.replica_count = 7;
stats.replication_stats.replicated_size = 2048;
stats
.replication_stats
.stats
.entry("arn:minio:replication::t:b".to_string())
.or_default()
.replicated_count = 5;
stats.proxy_stats.put_total = 3; stats.proxy_stats.put_total = 3;
let body = let body = serialize_replication_metrics_body(&stats, ReplicationExtRoute::MetricsV1, "node-1:9000")
serialize_replication_metrics_body(&stats, ReplicationExtRoute::MetricsV1).expect("metrics v1 body should serialize"); .expect("metrics v1 body should serialize");
let payload: serde_json::Value = serde_json::from_slice(&body).expect("body should be json"); let payload: serde_json::Value = serde_json::from_slice(&body).expect("body should be json");
assert_eq!(payload["replica_count"], 7); assert_eq!(payload["replicaCount"], 7);
assert_eq!(payload["completedReplicationSize"], 2048);
assert_eq!(payload["Stats"]["arn:minio:replication::t:b"]["replicationCount"], 5);
assert!(payload.get("uptime").is_none()); assert!(payload.get("uptime").is_none());
assert!(payload.get("proxy_stats").is_none()); assert!(payload.get("proxy_stats").is_none());
// The internal snake_case names must not leak into the wire body.
assert!(payload.get("replica_count").is_none());
assert!(payload.get("q_stat").is_none());
} }
#[test] #[test]
@@ -4248,22 +4276,48 @@ mod tests {
assert_eq!(target.current_bandwidth_bytes_per_sec, 3000.0); assert_eq!(target.current_bandwidth_bytes_per_sec, 3000.0);
} }
/// The v2 body must decode into minio-go `replication.MetricsV2`
/// (`uptime`/`currStats`/`queueStats`); `mc replicate status` reads
/// `currStats` and `queueStats.nodes` and silently shows zeros when the
/// keys do not match.
#[test] #[test]
fn serialize_replication_metrics_body_v2_returns_full_bucket_stats() { fn serialize_replication_metrics_body_v2_returns_minio_go_metrics_v2_shape() {
let mut stats = BucketStats { let mut stats = BucketStats {
uptime: 99, uptime: 99,
..Default::default() ..Default::default()
}; };
stats.replication_stats.replica_count = 7; stats.replication_stats.replica_count = 7;
stats
.replication_stats
.q_stat
.curr
.now_count
.store(4, std::sync::atomic::Ordering::Relaxed);
stats
.replication_stats
.q_stat
.curr
.now_bytes
.store(1200, std::sync::atomic::Ordering::Relaxed);
stats.replication_stats.q_stat = stats.replication_stats.q_stat.snapshot();
stats.proxy_stats.put_total = 3; stats.proxy_stats.put_total = 3;
let body = let body = serialize_replication_metrics_body(&stats, ReplicationExtRoute::MetricsV2, "node-1:9000")
serialize_replication_metrics_body(&stats, ReplicationExtRoute::MetricsV2).expect("metrics v2 body should serialize"); .expect("metrics v2 body should serialize");
let payload: serde_json::Value = serde_json::from_slice(&body).expect("body should be json"); let payload: serde_json::Value = serde_json::from_slice(&body).expect("body should be json");
assert_eq!(payload["uptime"], 99); assert_eq!(payload["uptime"], 99);
assert_eq!(payload["replication_stats"]["replica_count"], 7); assert_eq!(payload["currStats"]["replicaCount"], 7);
assert_eq!(payload["proxy_stats"]["put_total"], 3); assert_eq!(payload["currStats"]["queued"]["curr"]["count"], 4.0);
// The queue snapshot must surface at least one node: mc derives the
// worker/queue panels from queueStats.nodes and treats an empty list
// as "no data".
assert_eq!(payload["queueStats"]["nodes"][0]["queueStats"]["curr"]["count"], 4.0);
assert_eq!(payload["queueStats"]["nodes"][0]["uptime"], 99);
// The internal snake_case names must not leak into the wire body.
assert!(payload.get("replication_stats").is_none());
assert!(payload.get("queue_stats").is_none());
assert!(payload.get("proxy_stats").is_none());
} }
#[test] #[test]

Some files were not shown because too many files have changed in this diff Show More