Compare commits

..

28 Commits

Author SHA1 Message Date
overtrue abca7ddfd7 test(kms): move the Vault KV2 Transit-wrapping doc guard into check_fips_wording.sh
`test_vault_kv2_sources_do_not_claim_transit_wrapping` asserted that four
`include_str!`-pinned files never describe the Vault KV2 backend as wrapping key
material through Vault's Transit engine. The invariant is a documentation-claim
invariant with no behavioral twin by construction, and the test form was weak in
both directions: it saw only four files (the same prose in a fifth file passed
silently) and it stopped compiling — rather than reporting a violation — as soon
as one of them was renamed.

Move the four literals verbatim into `scripts/check_fips_wording.sh`, which
already guards the adjacent cryptographic over-claim class (unsupported FIPS
validation wording) and is anchored to the same policy document. The guard now
greps every file under `crates/kms` for the same four case-sensitive literals and
separately reports a moved pinned source instead of failing to build.

`check_fips_wording.sh` previously ran only in `make pre-commit` / `pre-pr`, so
wire it into the Quick Checks job of both CI workflows to keep the invariant's
failure visibility at least as strong as the deleted test's.
2026-08-18 17:30:11 +08:00
Zhengchao An 68547ed7ea test(e2e): drop 187 no-op serial markers from the three densest e2e suites (#6209)
test(e2e): drop no-op serial markers from the three densest e2e suites

serial_test's #[serial] is an in-process mutex. cargo-nextest, this repo's
authoritative runner, executes every test in its own process, so the mutex is
never contended and the attribute is a documented no-op -- see the "Serial
execution & nextest profiles" section of docs/testing/README.md and the header
of .config/nextest.toml. Cross-process serialization is provided only by a
[test-groups] entry with max-threads = 1.

Remove 187 such markers (plus 3 now-unused imports) from the three
marker-densest modules of crates/e2e_test:

  multipart_auth_test.rs             85
  replication_extension_test.rs      68
  object_lock/object_lock_test.rs    34

None of these modules is covered by any [test-groups] entry, so the markers
were carrying no isolation for any lane. Every test in all three files builds
its own server via RustFSTestEnvironment::new(), which gives a UUID temp dir
and a uniquely allocated port -- the .config/nextest.toml comment on
replication_extension_test already states this explicitly ("parallel-safe by
construction"). No test mutates process env, binds a fixed port, or touches
process-global state, so nothing here needed temp_env or a test-group instead.

Pure deletion: 190 lines removed, 0 added, no test renamed, no behaviour
changed. serial_test stays in Cargo.toml -- 338 markers across 90 other files
in the crate still use it.

Refs: backlog#1846 (T1).
2026-08-18 17:05:30 +08:00
Zhengchao An f06a9c9cba fix(deps): record heal's bytes and crc-fast in the lockfile (#6211) 2026-08-18 17:04:55 +08:00
Zhengchao An bd296eff9e chore(io-core): drop eight zero-consumer modules (#6201) 2026-08-18 08:14:05 +00:00
houseme a5800033bd feat(heal): incremental status cursors and typed overlap policy (HS-06) (#6206)
* feat(heal): incremental heal status cursors and typed overlap policy (HS-06)

Incremental results: every retained result item now carries a monotonic
sequence number. The status query accepts a client cursor (sinceSeq on
the admin wire, Option<u64> internally) and returns only newer items,
plus nextSeq (the next cursor) and minSeq (the oldest retained
sequence). A cursor that fell behind the 1024-item retention window is
flagged through the existing truncated signal together with minSeq so
the client can restart from it. Sequencing survives task completion:
the completion archive stores the seq-stamped window. None keeps the
exact legacy full-snapshot behavior, so existing clients see no change.

Typed overlap handling for admin starts: RUSTFS_HEAL_OVERLAP_POLICY
(merge default | minio_error). Under minio_error, an admin start whose
path overlaps an active or queued task rejects with typed
already-running / overlapping-paths admission reasons (surfaced through
reason_label in the admin error body, sharing the existing
OperationAborted site because the s3s footprint ratchet forbids new
s3_error! sites); an exact duplicate start rejects with
already-running instead of silently merging. Scanner/autoheal/
read-repair sources never take the rejection path.

forceStart semantics now match MinIO for admin requests: an admin
forceStart first cancels the overlapping active admin task, then
admits the replacement.

Wire: the heal-control Query command grows an optional sinceSeq
(defaulted and skipped when absent, so older peers stay compatible);
the admin handler accepts the sinceSeq query parameter; the local
channel query gains the same cursor.

Tests: seq monotonicity and incremental slicing, window slide moving
minSeq with lagging-cursor flags, overlap matrix (same/containing/
contained/disjoint x policy x source), forceStart cancel-then-admit,
and the completion-archive window handoff.

Co-Authored-By: heihutu <heihutu@gmail.com>

* style: fmt after main merge

---------

Co-authored-by: heihutu <heihutu@gmail.com>
Co-authored-by: zhi22915 <qiuzgang@gmail.com>
2026-08-18 16:09:30 +08:00
cxymds a4ea36b298 fix(ecstore): single-flight remote disk recovery (#6096)
* fix(ecstore): single-flight remote disk recovery

* test(ecstore): cover remote recovery review cases

* test(ecstore): exercise recovery through disk slot

* test(ecstore): match format reads exactly

* test(ecstore): cover recovery teardown races

* style(ecstore): format recovery race tests

* fix(ecstore): remove unused health snapshot helper

* fix(ecstore): group recovery monitor test state

* fix(protos): preserve production source in compatibility checks
2026-08-18 14:49:53 +08:00
cxymds 4f68f117ba fix(lock): reap expired local lease guards (#6094)
* fix(lock): reap expired local lease guards

* fix(lock): reject refresh after guard expiry

* test(lock): isolate expired refresh regression

* style(lock): format expired refresh assertion
2026-08-18 14:49:39 +08:00
cxymds 0f30a75fdb fix(ecstore): resume remote shard reads once (#6091)
* fix(ecstore): preserve CopyObject producer errors

* fix(ecstore): resume remote shard reads once

* fix(app): resume preserved relocation I/O errors

* fix(ecstore): reserve remote read recovery budget
2026-08-18 14:49:19 +08:00
Zhengchao An 355c8d2e22 fix(admin): classify missing kms config by error variant (#6196) 2026-08-18 05:40:10 +00:00
Zhengchao An 60eb139db9 refactor: import x-amz-checksum header names from the shared constants (#6193)
Co-authored-by: houseme <housemecn@gmail.com>
2026-08-18 05:27:38 +00:00
hector 9ef059c908 ci(package): auto-trigger DEB/RPM packaging on releases and upload to GitHub release assets (#6202) 2026-08-18 05:18:13 +00:00
Zhengchao An 51497cb533 fix(ecstore): give peer REST failures op and bucket context (#6200) 2026-08-18 04:57:44 +00:00
Zhengchao An c7a29ec0a7 chore(storage): drop dead backpressure and lock-optimizer wrappers (#6195)
Neither rustfs/src/storage/backpressure.rs nor rustfs/src/storage/lock_optimizer.rs had a production caller: their only non-self references were the pub mod lines in storage/mod.rs and a cfg(test) module, so the object transfer path never applied this backpressure and never took these lock shortcuts.

The six removed tests in concurrent_fix_test.rs duplicated tests that lived inside the deleted files; the shared primitives they shadowed keep their own coverage in rustfs-io-core.
2026-08-18 12:51:44 +08:00
Zhengchao An deb0edb7cc chore: adjudicate 26 bare dead_code allows across five crates (#6187)
Remove every bare `#[allow(dead_code)]` in io-core, object-capacity, targets, rio, and scanner. Each allow was stripped first and clippy was then asked which ones the compiler actually missed, so the verdicts rest on the diagnostic rather than on inspection.

23 were inert: they sat on `pub fn`s inside `pub mod`s, where `dead_code` does not apply, or on scanner integration-test helpers that the tests in the same file do call.

The remaining 3 are in rio's private `compress_index` module and the code behind them is deleted rather than annotated. `remove_index_headers` is dead and also wrong — after skipping the 4-byte chunk header it matches against `S2_INDEX_TRAILER` where `S2_INDEX_HEADER` sits, so it returns `None` for every well-formed index; rio-v2 carries the correct equivalent that is actually in use. `restore_index_headers` is its unreachable counterpart, likewise duplicated live in rio-v2. `Index::reset` is a private method with no caller.

Refs backlog#1823
2026-08-18 12:45:42 +08:00
houseme a08de9229b feat(heal): wire MRF intents with durable repair journal (HS-01) (#6189)
* feat(common): add MRF intent channel and Mrf request source (HS-01)

Introduce the producer-facing half of the mission repair feed: a global
bounded (8192) channel carrying lightweight MrfIntent values from IO
error paths, plus the RUSTFS_HEAL_MRF_ENABLE delivery kill-switch and
config constants for queue/journal sizing. Delivery is strictly
non-blocking (try_send, drop-on-full) so it can sit on decode-failure
and partial-write paths without adding latency. HealRequestSource grows
a 'mrf' variant so admission accounting can attribute replayed intents.

Part of backlog#1865 (option a: wire HealEvent-style intents with a
durable retry ledger).

Co-Authored-By: heihutu <heihutu@gmail.com>

* feat(heal): add MRF queue, durable journal, and intent consumer (HS-01)

Consumer half of the mission repair feed: a bounded pending queue
(100k intents / 8 MiB dual ceiling, drop-newest on overflow), a durable
journal at buckets/.heal/mrf/journal.bin holding the unaccepted pending
snapshot, and a consumer task that batches intents off the global
channel, translates them into prioritized heal requests (decode
failure -> Urgent ECDecode, metadata corruption -> High Metadata,
partial write -> Normal object heal), and retries full admissions with
a 5s backoff and a 3-attempt ceiling.

Durability: every journal record carries its own CRC32 and a
format/version header, so a torn tail truncates cleanly at replay; the
journal is deleted after a successful replay and when the pending set
drains (mirroring MinIO's post-replay list.bin unlink). Losing the last
500 ms flush window is acceptable: replayed duplicates merge via the
manager dedup key and read-repair remains the safety net.

Metrics: rustfs_heal_mrf_queue_depth/_queue_bytes, _dropped_total
{reason}, _replayed_total, _journal_bytes, _journal_fsync_total.
The consumer is wired at heal runtime bootstrap right after manager
start, honoring RUSTFS_HEAL_MRF_ENABLE (default on, rollback = off).

Tests: unit tests for the dual ceiling, record roundtrip, torn-tail
truncation, and the priority mapping; integration tests against a real
4-disk ECStore proving channel intents reach the manager queue as
Urgent/mrf-attributed requests and journal replay arms intents, drops
torn tails, and removes the file.

Part of backlog#1865 (option a).

Co-Authored-By: heihutu <heihutu@gmail.com>

* feat(ecstore,scanner): deliver MRF intents from error paths (HS-01)

Wire the three production delivery points, each a single non-blocking
try_send next to the existing in-memory heal paths, which stay as the
fast path:

- read.rs decode-error branch: DecodeFailure intent beside the existing
  read-repair submit, so an Urgent ECDecode request survives restarts
  even when the Low-priority read-repair request was dropped or lost.
- add_partial: PartialWrite intent, giving partial-write recovery a
  durable Normal-priority object heal across restarts.
- scanner_folder metadata-corruption classification: MetadataCorruption
  intent beside the existing High-priority scanner heal request.

All three are on error paths only: zero cost on healthy IO.

Part of backlog#1865 (option a).

Co-Authored-By: heihutu <heihutu@gmail.com>

* fix: include mrf heal source counts

Co-Authored-By: heihutu <heihutu@gmail.com>

* fix: keep node heal status wire compatibility

Co-Authored-By: heihutu <heihutu@gmail.com>

---------

Co-authored-by: heihutu <heihutu@gmail.com>
2026-08-18 12:43:27 +08:00
Zhengchao An abffa5cf1b chore(storage): drop dead io-schedule metrics and helpers (#6199) 2026-08-18 04:35:36 +00:00
Zhengchao An b825c54850 refactor(admin): route kms management auth through shared gate (#6194) 2026-08-18 12:21:49 +08:00
houseme de9145e87a feat(storage): add default-off PUT admission gate (#6197)
Add an experimental fixed-count foreground PutObject admission gate for #1882 Phase 0 validation. The gate is default-off, returns SlowDown before body ingest when saturated, and keeps the admission permit with the spawned store commit owner until store PUT returns.

Co-authored-by: heihutu <heihutu@gmail.com>
2026-08-18 04:15:19 +00:00
houseme 84bd76a3ce chore(deps): refresh cargo dependencies (#6198)
Update workspace Cargo dependency requirements and lockfile after cargo update/upgrade, including rumqttc-next 0.34.0 and MQTT API compatibility adjustments.

Verification:

- cargo update --verbose

- cargo upgrade --verbose

- cargo update -p rumqttc-next --precise 0.34.0 --verbose

- cargo tree --invert rumqttc-next --locked

- cargo metadata --locked --no-deps --format-version 1

- cargo fmt --all --check

- cargo check -p rustfs-targets --all-targets --locked

- cargo test -p rustfs-targets mqtt --locked

- make pre-pr

Co-authored-by: heihutu <heihutu@gmail.com>
2026-08-18 04:11:53 +00:00
houseme 9a2d06b370 test(heal): lock heal vs delete/overwrite race invariants (HS-12) (#6183)
* test(heal): add concurrency invariants for heal vs delete/overwrite races (HS-12)

Audit conclusion for backlog#1874: RustFS does not need a persistent
object-level healing marker (MinIO x-minio-healing) because every path
that can touch the same (bucket, object) commit surface serializes on
the same namespace write lock, and the heal lock guard spans the whole
rename commit including the HEAL_RENAME_INCOMPLETE partial path.

Lock the conclusion in with two race regression tests:

- heal_racing_version_delete_never_resurrects_the_deleted_version:
  shard damage is injected on the doomed version so a Deep heal has real
  reconstruction work while a versioned DELETE runs concurrently; the
  deleted version must stay deleted and the survivor intact.
- heal_racing_unversioned_overwrites_preserves_the_last_commit:
  unversioned overwrites (activating the post-commit tail that deletes
  the replaced data dir without the ns lock) race a Deep heal in a
  loop; the final current version must be exactly the last commit.

Also adds docs/operations/heal-concurrency-safety-notes-zh.md with the
full intersection matrix (17 intersections), lock-coverage argument,
and the residual-window classification (commit tail races are
fail-into-retry safe; bare prefix delete has zero production callers;
admin no_lock is an explicit operator opt-in).

Co-Authored-By: heihutu <heihutu@gmail.com>

* test: remove redundant heal etag clone

Co-Authored-By: heihutu <heihutu@gmail.com>

---------

Co-authored-by: heihutu <heihutu@gmail.com>
2026-08-18 02:01:04 +00:00
Zhengchao An 00de43528c fix(ecstore): describe peer bucket RPC failures with no details (#6190)
heal_bucket, list_bucket, get_bucket_info and delete_bucket returned Error::other("") when a peer answered success=false without an error payload, so operators saw a bare "io error " after quorum reduction. Route all five bucket RPCs through peer_failure_without_details, which names the operation and bucket while staying identical across the peers of one operation so reduce_errs keeps grouping them into a single dominant error.
2026-08-18 01:55:29 +00:00
houseme 35a30cd614 feat(scanner): emit excess alerts as S3 notification events (HS-04) (#6176)
* feat(scanner): emit excess alerts as S3 notification events

The excess-versions / excess-version-size / excess-folders alerts were
metrics-and-logs only; consoles and external auditors had no way to hear
them (rustfs/backlog#1868, HS-04). MinIO emits s3:ObjectManyVersions /
s3:ObjectLargeVersions / s3:PrefixManyFolders for the same conditions —
RustFS carries those as EventName::Scanner* with s3:Scanner:* wire names
that already existed unpublished.

The three alert sites now also dispatch through the standard event
pipeline (send_event via the storage_api owner facade), carrying the
actual values and thresholds in req_params and UserAgent "Scanner".
Without a cooldown a single over-threshold object would re-emit on every
~60s scan cycle, so emissions are edge-held per (kind, bucket, object)
for 24h (RUSTFS_SCANNER_ALERT_COOLDOWN_SECS, 0 = every cycle), backed by
a process-global map with a 4096-key hard cap that clears rather than
grows. Metrics and structured logs stay level-triggered every cycle;
only the notification events are held back. A restart resets the
cooldown deliberately: one re-emission per still-hot key buys back
visibility after the restarts that accompany incident response.

Tests pin the edge-hold semantics (first fires, immediate re-check held,
independent keys, cooldown expiry re-fires, zero cooldown always emits,
hard bound) in one sequential test for the process-global map, and pin
the emitted wire names against EventName's canonical string forms so a
subscribed bucket notification can never silently stop matching.
docs/operations/scanner-excess-alerts.md documents the three events,
the metric-vs-event cadence difference, and the HS-15 threshold deltas
(alert_excess_folders 65538 vs MinIO 50000 is deliberate: Proxmox
Backup Server chunk layout compatibility).

Closes rustfs/backlog#1868.

Co-Authored-By: heihutu <heihutu@gmail.com>

* docs(operations): split scanner excess alerts into English and Chinese pages

The page shipped Chinese-only; keep it as scanner-excess-alerts_zh.md and
add a faithful English translation at the original path, cross-linked at
the top of both.

Co-Authored-By: heihutu <heihutu@gmail.com>

---------

Co-authored-by: heihutu <heihutu@gmail.com>
2026-08-18 08:46:32 +08:00
houseme 360bceafce feat(heal): add progress and trace observability (#6179)
* feat(heal): track erasure set progress baseline

Record erasure-set heal byte progress from per-object results and seed progress totals from complete usage-cache snapshots when available.

Keep usage-cache failures observational so heal execution continues without a baseline.

Co-Authored-By: heihutu <heihutu@gmail.com>

* feat(heal): skip filtered erasure set versions

Skip erasure-set versions written after the durable heal start time, and queue lifecycle-expired versions for expiry before skipping them.

Track new-version and ILM-expired skips separately so progress can explain completed baseline work without treating these skips as retry-blocking failures.

Co-Authored-By: heihutu <heihutu@gmail.com>

* feat(heal): wire abandoned data-dir cleanup check

Connect check_abandoned_parts through ECStore, pool, and set layers so heal can invoke the existing orphan data-dir reclaim path instead of returning NotImplemented.

Add dry-run support to the reclaim scan and cover dry-run plus scoped set behavior with regression tests.

Co-Authored-By: heihutu <heihutu@gmail.com>

* feat(obs): add heal scanner trace bus

Introduce an in-process broadcast trace bus with typed heal and scanner events, lazy event construction, and bounded lagged-subscriber behavior.

Cover zero-subscriber publishing, subscription delivery, drop accounting, and lagged receivers with focused common-crate tests.

Co-Authored-By: heihutu <heihutu@gmail.com>

* feat(obs): stream heal trace events from admin API

Wire the admin trace endpoint to the common trace bus for heal/scanner events, including kind, regex, and threshold filtering.

Co-Authored-By: heihutu <heihutu@gmail.com>

* feat(obs): emit heal trace events

Publish heal task lifecycle and abandoned-parts cleanup events through the common trace bus so the admin trace stream has live heal diagnostics.

Co-Authored-By: heihutu <heihutu@gmail.com>

* feat(obs): emit scanner trace events

Publish scanner folder, lifecycle action, and heal-candidate events through the common trace bus for live admin scanner diagnostics.

Co-Authored-By: heihutu <heihutu@gmail.com>

* fix(heal): route data usage loader through storage api

Keep ECStore data-usage facade access behind the heal storage_api boundary so architecture migration guards can validate the heal progress path.

Co-Authored-By: heihutu <heihutu@gmail.com>

* perf(heal): avoid lifecycle snapshots on ordinary heal pages

Only request lifecycle object snapshots when the heal pass has lifecycle expiry context. This keeps ordinary listing and disk-walk pages from cloning FileInfo/ObjectInfo payloads while preserving the skip path that queues expired versions.

Co-Authored-By: heihutu <heihutu@gmail.com>

* test(heal): update bug-fix mocks for lifecycle snapshots

Carry the lifecycle snapshot opt-in argument through the remaining heal bug-fix test mocks so all-targets clippy covers the updated storage trait.

Co-Authored-By: heihutu <heihutu@gmail.com>

* test(rustfs): sync heal storage mock signature

Update the rustfs storage RPC test mock for the lifecycle snapshot opt-in argument and cover it with rustfs all-targets clippy.

Co-Authored-By: heihutu <heihutu@gmail.com>

* test(e2e): allocate smoke ports across nextest processes

Serialize E2E port selection with a small /tmp allocator so nextest workers do not reuse the same just-released ephemeral port before RustFS binds it.

Co-Authored-By: heihutu <heihutu@gmail.com>

---------

Co-authored-by: heihutu <heihutu@gmail.com>
2026-08-18 08:29:29 +08:00
Zhengchao An 7cb91a0190 chore(ecstore): adjudicate 32 bare dead_code allows (#6173)
Replace every bare `#[allow(dead_code)]` in ecstore with either a deletion or a per-item allow carrying a `reason`. Blanket allows at module, struct, and impl level silence the lint for future members too, so each is narrowed to the members that are actually dead.

Delete the dead cluster in `config/heal.rs` (`Config`, its three methods, `RUSTFS_BITROT_CYCLE_IN_MONTHS`, `parse_bitrot_config`) rather than annotate it: it has no callers and is unreachable outside the crate, and `parse_bitrot_config` would panic on its disabled path via `Duration::from_secs_f64(-1.0)`. `DEFAULT_KVS` stays, since the config registry uses it.

Correct two `reason` strings on `Checksum::new` and `PutObjReader::md5_current_hex_string`, which are methods but carried a field-only rationale.

Refs backlog#1823

Co-authored-by: houseme <housemecn@gmail.com>
2026-08-17 23:51:56 +00:00
hector beb6e1383e feat(helm): add TLSRoute passthrough support for gateway api (#6169)
Add an optional TLS passthrough listener to the Gateway API support. When gatewayApi.listeners.tls.enabled is true, the Gateway gets a TLS listener with tls.mode: Passthrough and a TLSRoute is rendered to the RustFS service so TLS terminates at the backend (end-to-end encryption).

Refs rustfs/rustfs#3862.
2026-08-18 01:21:15 +08:00
houseme 59b7d13095 feat(scanner): expose prefix-level bucket usage via admin API (HS-08) (#6171)
feat(scanner): expose prefix-level bucket usage via admin API

The scanner's per-bucket, per-set usage caches already hold a path-keyed
prefix tree, but dui() flattened it only to bucket names — consoles and
operators had no way to ask "what does this prefix hold" without an S3
listing sweep (rustfs/backlog#1872, MinIO loadPrefixUsageFromBackend
parity).

Add:

- data-usage: prefix_usage_in_cache — a shared aggregation over the
  entry map (arbitrary prefix, full counters, one-level sub-prefix
  breakdown with names recovered from the literal-path cache keys),
  hardened like the scanner's checked flatten: cycles, dangling child
  links, over-deep trees, and overflowing counters yield None rather
  than unbounded recursion or wrapped totals.
- ecstore: ECStore::all_set_disks — iterate every erasure set so a
  query can read each set's own cache copy; the hash-routed store path
  would always land on one set.
- scanner: bucket_prefix_usage — per-set loads (5s budget each, a slow
  set degrades to not-reporting instead of stalling the caller),
  merged across sets with partial/compacted/truncated flags, served
  from a bounded 30s cache (128 entries, hard-capped) that bucket
  writes invalidate through the dirty-usage hook.
- admin: GET /rustfs/admin/v3/usage/{bucket}?prefix=&max-entries=
  behind the same any-of gate as datausageinfo (DataUsageInfoAdminAction
  OR ListBucketAction), rejecting unknown query parameters and
  clamping max-entries to 1..=10000. Route registered in the policy
  table (deferred MultipleActions, matching datausageinfo) and the
  route matrix test.

Closes rustfs/backlog#1872.

Co-authored-by: heihutu <heihutu@gmail.com>
2026-08-17 11:40:56 +00:00
唐小鸭 e0b87b0e7e fix(site-replication): admit only verifiable peer-edit fences (#6123) 2026-08-17 09:47:36 +00:00
houseme 984c705713 docs(ecstore): fix bitrot comment typo (#6168)
Co-authored-by: heihutu <heihutu@gmail.com>
2026-08-17 08:24:54 +00:00
147 changed files with 9585 additions and 6454 deletions
+2 -2
View File
@@ -66,8 +66,8 @@ s3s-footprint-check: ## Check the s3s dependency footprint ratchet stays frozen
./scripts/check_s3s_footprint.sh ./scripts/check_s3s_footprint.sh
.PHONY: fips-wording-check .PHONY: fips-wording-check
fips-wording-check: ## Check outward docs do not make unsupported FIPS claims fips-wording-check: ## Check docs and crates/kms do not over-claim crypto capabilities
@echo "📣 Checking FIPS wording guard..." @echo "📣 Checking cryptographic capability wording guard..."
./scripts/check_fips_wording.sh ./scripts/check_fips_wording.sh
.PHONY: log-analyzer-rules-check .PHONY: log-analyzer-rules-check
+3
View File
@@ -117,6 +117,9 @@ jobs:
- name: Check s3s footprint ratchet - name: Check s3s footprint ratchet
run: ./scripts/check_s3s_footprint.sh run: ./scripts/check_s3s_footprint.sh
- name: Check cryptographic capability wording
run: ./scripts/check_fips_wording.sh
- name: Check no planning docs committed - name: Check no planning docs committed
run: ./scripts/check_no_planning_docs.sh run: ./scripts/check_no_planning_docs.sh
+3
View File
@@ -152,6 +152,9 @@ jobs:
- name: Check s3s footprint ratchet - name: Check s3s footprint ratchet
run: ./scripts/check_s3s_footprint.sh run: ./scripts/check_s3s_footprint.sh
- name: Check cryptographic capability wording
run: ./scripts/check_fips_wording.sh
- name: Check no planning docs committed - name: Check no planning docs committed
run: ./scripts/check_no_planning_docs.sh run: ./scripts/check_no_planning_docs.sh
+84 -11
View File
@@ -15,28 +15,35 @@
# Package Workflow - Build DEB/RPM packages # Package Workflow - Build DEB/RPM packages
# #
# This workflow builds DEB and RPM packages from pre-built Linux binaries # This workflow builds DEB and RPM packages from pre-built Linux binaries
# and uploads them to Cloudflare R2. # and uploads them to Cloudflare R2 and the GitHub release.
# #
# Trigger: # Trigger:
# - release published: automatically package when a GitHub release is published # - workflow_run: automatically package after "Build and Release" completes
# - workflow_dispatch: manual trigger with optional tag/run_id # for a release tag (the mac/windows/linux binaries are already uploaded
# to the GitHub release before packaging starts)
# - workflow_dispatch: manual fallback (backfill / re-run) with optional tag/run_id
# #
# Flow: # Flow:
# 1. Find the Build workflow run for the release tag # 1. Resolve the triggering Build workflow run for the release tag
# 2. Download Linux binaries (x86_64-gnu, aarch64-gnu) from build artifacts # 2. Download Linux binaries (x86_64-gnu, aarch64-gnu) from build artifacts
# 3. Build DEB packages for amd64 and arm64 # 3. Build DEB packages for amd64 and arm64
# 4. Build RPM packages for x86_64 and aarch64 # 4. Build RPM packages for x86_64 and aarch64
# 5. Upload all packages to Cloudflare R2 # 5. Upload all packages to Cloudflare R2 and the GitHub release
name: Package DEB/RPM name: Package DEB/RPM
permissions: permissions:
contents: read # contents: write is required to upload packages to the GitHub release
contents: write
actions: read actions: read
on: on:
release: # Follows the same pattern as docker.yml: run after the release build
types: [ published ] # workflow completes, so packaging is triggered only by release tags
# (e.g. 1.0.0-rc.2, 1.0.0-rc.3), never by development builds.
workflow_run:
workflows: [ "Build and Release" ]
types: [ completed ]
workflow_dispatch: workflow_dispatch:
inputs: inputs:
tag: tag:
@@ -49,13 +56,26 @@ on:
type: string type: string
concurrency: concurrency:
group: ${{ github.workflow }}-${{ github.event.release.tag_name || github.event.inputs.tag || github.run_id }} group: ${{ github.workflow }}-${{ github.event.workflow_run.head_branch || github.event.inputs.tag || github.run_id }}
cancel-in-progress: true cancel-in-progress: true
env:
HEAD_BRANCH: ${{ github.event.workflow_run.head_branch }}
WORKFLOW_RUN_ID: ${{ github.event.workflow_run.id }}
jobs: jobs:
# Resolve which build run to use and extract version info # Resolve which build run to use and extract version info
resolve: resolve:
name: Resolve Build name: Resolve Build
# Auto-trigger only from successful tag builds of "Build and Release".
# Tag pushes arrive as event == push with head_branch != main (a
# non-main push head_branch is the release tag name). Manual dispatch
# stays available as a fallback for backfills and re-runs.
if: >-
github.event_name == 'workflow_dispatch' ||
(github.event.workflow_run.conclusion == 'success' &&
github.event.workflow_run.event == 'push' &&
github.event.workflow_run.head_branch != 'main')
runs-on: ubuntu-latest runs-on: ubuntu-latest
timeout-minutes: 10 timeout-minutes: 10
outputs: outputs:
@@ -75,8 +95,8 @@ jobs:
set -euo pipefail set -euo pipefail
# Determine tag # Determine tag
if [[ "${{ github.event_name }}" == "release" ]]; then if [[ "${{ github.event_name }}" == "workflow_run" ]]; then
TAG="${{ github.event.release.tag_name }}" TAG="${HEAD_BRANCH}"
elif [[ -n "$INPUT_TAG" ]]; then elif [[ -n "$INPUT_TAG" ]]; then
TAG="$INPUT_TAG" TAG="$INPUT_TAG"
else else
@@ -93,6 +113,11 @@ jobs:
BUILD_RUN_ID="$INPUT_RUN_ID" BUILD_RUN_ID="$INPUT_RUN_ID"
echo "Using explicit build run ID: $BUILD_RUN_ID" echo "Using explicit build run ID: $BUILD_RUN_ID"
elif [[ "${{ github.event_name }}" == "workflow_run" ]]; then
# Use the Build and Release run that triggered this workflow
BUILD_RUN_ID="${WORKFLOW_RUN_ID}"
echo "Using triggering workflow run: $BUILD_RUN_ID"
elif [[ -n "$TAG" ]]; then elif [[ -n "$TAG" ]]; then
# Find the build run that produced this tag # Find the build run that produced this tag
echo "Looking for build run for tag: $TAG" echo "Looking for build run for tag: $TAG"
@@ -456,6 +481,54 @@ jobs:
echo "✅ Latest packages updated" echo "✅ Latest packages updated"
fi fi
- name: Upload packages to GitHub Release
if: needs.resolve.outputs.tag != ''
env:
GH_TOKEN: ${{ github.token }}
shell: bash
run: |
set -euo pipefail
TAG="${{ needs.resolve.outputs.tag }}"
DEB_FILE="${{ steps.deb.outputs.deb_file }}"
RPM_FILE="${{ steps.rpm.outputs.rpm_file }}"
# Upload the packages, then refresh the release checksums so the new
# assets are covered, matching the binary release flow.
for f in "$DEB_FILE" "$RPM_FILE"; do
if [[ -n "$f" && -f "$f" ]]; then
echo "📤 Uploading $(basename "$f") to GitHub release ${TAG}..."
gh release upload "$TAG" "$f" --clobber
fi
done
CHECKSUM_DIR="$(mktemp -d)"
gh release download "$TAG" -p 'SHA256SUMS' -p 'SHA512SUMS' \
-D "$CHECKSUM_DIR" --clobber 2>/dev/null || true
for spec in "SHA256SUMS:sha256sum" "SHA512SUMS:sha512sum"; do
asset="${spec%%:*}"
checksum_cmd="${spec##*:}"
checksum_file="${CHECKSUM_DIR}/${asset}"
touch "$checksum_file"
for f in "$DEB_FILE" "$RPM_FILE"; do
if [[ -n "$f" && -f "$f" ]]; then
base="$(basename "$f")"
# Remove any stale entry, then append the fresh digest
grep -Fv -- "$base" "$checksum_file" > "${checksum_file}.tmp" || true
mv "${checksum_file}.tmp" "$checksum_file"
(cd "$(dirname "$f")" && "$checksum_cmd" -- "$base") >> "$checksum_file"
fi
done
echo "📤 Updating ${asset} for release ${TAG}..."
gh release upload "$TAG" "$checksum_file" --clobber
done
echo "✅ GitHub release assets updated"
# Summary # Summary
summary: summary:
name: Summary name: Summary
+4 -4
View File
@@ -31,7 +31,7 @@ HTTP request
→ storage/ecfs (erasure coding, encryption, checksums) → storage/ecfs (erasure coding, encryption, checksums)
→ ecstore (disk pool selection, data distribution) → ecstore (disk pool selection, data distribution)
→ rio (reader pipeline: encrypt → compress → hash → write) → rio (reader pipeline: encrypt → compress → hash → write)
→ io-core (zero-copy I/O, buffer pool, direct I/O) → io-core (buffer pool, storage profiling, admission control)
→ local disk / remote disk via RPC → local disk / remote disk via RPC
``` ```
@@ -55,7 +55,7 @@ rustfs/ # Workspace root (virtual manifest)
├── crates/ # library crates (authoritative list: Cargo.toml [workspace].members) ├── crates/ # library crates (authoritative list: Cargo.toml [workspace].members)
│ ├── ecstore/ # Erasure-coded storage engine │ ├── ecstore/ # Erasure-coded storage engine
│ ├── rio/ # Reader I/O pipeline (encrypt, compress, hash) │ ├── rio/ # Reader I/O pipeline (encrypt, compress, hash)
│ ├── io-core/ # Zero-copy I/O, scheduling, buffer pool │ ├── io-core/ # Buffer pool, storage profiling, admission control
│ ├── io-metrics/ # I/O metrics collection │ ├── io-metrics/ # I/O metrics collection
│ ├── common/ # Shared runtime state, globals, data usage types │ ├── common/ # Shared runtime state, globals, data usage types
│ ├── config/ # Configuration types and parsing │ ├── config/ # Configuration types and parsing
@@ -302,7 +302,7 @@ The binary (`main.rs`) boots in this order:
│ │ │ │ │ │
┌─────▼──────┐ ┌──────▼──────┐ ┌──────▼──────┐ ┌─────▼──────┐ ┌──────▼──────┐ ┌──────▼──────┐
│ ecstore │ │ rio │ │ io-core │ │ ecstore │ │ rio │ │ io-core │
│ (core) │ │ (readers) │ │ (zero-copy) │ (core) │ │ (readers) │ │ (buffers)
└─────┬──────┘ └─────────────┘ └─────────────┘ └─────┬──────┘ └─────────────┘ └─────────────┘
┌─────┬──┼──┬─────┬──────┐ ┌─────┬──┼──┬─────┬──────┐
@@ -314,7 +314,7 @@ The binary (`main.rs`) boots in this order:
- **"Where does S3 PutObject go?"** - **"Where does S3 PutObject go?"**
`server/` routes → `app/object_usecase` validates → `storage/ecfs` encodes → `server/` routes → `app/object_usecase` validates → `storage/ecfs` encodes →
`ecstore` distributes → `rio` encrypts/compresses → `io-core` writes `ecstore` distributes → `rio` encrypts/compresses → `io-core` supplies buffers
- **"Where are bucket policies enforced?"** - **"Where are bucket policies enforced?"**
`app/bucket_usecase` calls into `crates/policy/` `app/bucket_usecase` calls into `crates/policy/`
Generated
+109 -104
View File
@@ -964,9 +964,9 @@ dependencies = [
[[package]] [[package]]
name = "aws-sdk-kms" name = "aws-sdk-kms"
version = "1.114.0" version = "1.115.0"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "c0b7d906608ee41e7ddea9983577ba82200435644d567d63dc34e822e088b453" checksum = "d5b034f8b7ceadb873d0bc607c30bb4b0be68e09a84c837174e7c2c6878ff882"
dependencies = [ dependencies = [
"arc-swap", "arc-swap",
"aws-credential-types", "aws-credential-types",
@@ -990,9 +990,9 @@ dependencies = [
[[package]] [[package]]
name = "aws-sdk-s3" name = "aws-sdk-s3"
version = "1.141.0" version = "1.142.0"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "d9f9420d3a2467eed22ed3635ca653653162c386a0b0f65c78189f9bd3c1379e" checksum = "f9e15a5c55e05f4b0b7e483160b3c85cccdf77cff02c95504f3e71d460855cd2"
dependencies = [ dependencies = [
"arc-swap", "arc-swap",
"aws-credential-types", "aws-credential-types",
@@ -1027,9 +1027,9 @@ dependencies = [
[[package]] [[package]]
name = "aws-sdk-sso" name = "aws-sdk-sso"
version = "1.105.0" version = "1.106.0"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "6ffd0fbe7873cb548a7aa60f9573c268fff94155397fd4f14dc9f1ecaaab8516" checksum = "2d0efcee834347b6705eca3eea2defd88242f43774f55d7326604222e3c86260"
dependencies = [ dependencies = [
"arc-swap", "arc-swap",
"aws-credential-types", "aws-credential-types",
@@ -1053,9 +1053,9 @@ dependencies = [
[[package]] [[package]]
name = "aws-sdk-ssooidc" name = "aws-sdk-ssooidc"
version = "1.107.0" version = "1.108.0"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "175763eb222a46377df7aa257a3bca980ab3e96703fefc8f4d0b8da6ad2e254c" checksum = "a59312a04cf19c962cfee32b64ecfee758f8786407ff6da5b30fff46ae96f201"
dependencies = [ dependencies = [
"arc-swap", "arc-swap",
"aws-credential-types", "aws-credential-types",
@@ -1079,9 +1079,9 @@ dependencies = [
[[package]] [[package]]
name = "aws-sdk-sts" name = "aws-sdk-sts"
version = "1.110.0" version = "1.111.0"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "dd8b14781dfbff48984017d57167b6ea0b6471c6920ec52b44a2677c7feb3c13" checksum = "120e7eb63457a9e547f9986fe3b273f77c43679da4d04f46359fa881c5e19b6e"
dependencies = [ dependencies = [
"arc-swap", "arc-swap",
"aws-credential-types", "aws-credential-types",
@@ -1598,7 +1598,7 @@ version = "0.10.4"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "3078c7629b62d3f0439517fa394996acacc5cbc91c5a20d8c658e77abd503a71" checksum = "3078c7629b62d3f0439517fa394996acacc5cbc91c5a20d8c658e77abd503a71"
dependencies = [ dependencies = [
"generic-array 0.14.9", "generic-array 0.14.7",
] ]
[[package]] [[package]]
@@ -1617,7 +1617,7 @@ version = "0.3.3"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "a8894febbff9f758034a5b8e12d87918f56dfc64a8e1fe757d65e29041538d93" checksum = "a8894febbff9f758034a5b8e12d87918f56dfc64a8e1fe757d65e29041538d93"
dependencies = [ dependencies = [
"generic-array 0.14.9", "generic-array 0.14.7",
] ]
[[package]] [[package]]
@@ -1858,9 +1858,9 @@ dependencies = [
[[package]] [[package]]
name = "cc" name = "cc"
version = "1.4.2" version = "1.4.3"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "5d262e149917187838d5b42777c8253bcb64500067342904e7d429499a6f277e" checksum = "509591b7bcd67f4ef775afad7662703b4935daaa6ec0e5605cfb1090b32a2b6d"
dependencies = [ dependencies = [
"find-msvc-tools", "find-msvc-tools",
"jobserver", "jobserver",
@@ -1968,7 +1968,7 @@ version = "0.4.4"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "773f3b9af64447d2ce9850330c473515014aa235e6a783b02db81ff39e4a3dad" checksum = "773f3b9af64447d2ce9850330c473515014aa235e6a783b02db81ff39e4a3dad"
dependencies = [ dependencies = [
"crypto-common 0.1.6", "crypto-common 0.1.7",
"inout 0.1.4", "inout 0.1.4",
] ]
@@ -2428,7 +2428,7 @@ version = "0.5.5"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "0dc92fb57ca44df6db8059111ab3af99a63d5d0f8375d9972e319a379c6bab76" checksum = "0dc92fb57ca44df6db8059111ab3af99a63d5d0f8375d9972e319a379c6bab76"
dependencies = [ dependencies = [
"generic-array 0.14.9", "generic-array 0.14.7",
"rand_core 0.6.4", "rand_core 0.6.4",
"subtle", "subtle",
"zeroize", "zeroize",
@@ -2453,11 +2453,11 @@ dependencies = [
[[package]] [[package]]
name = "crypto-common" name = "crypto-common"
version = "0.1.6" version = "0.1.7"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "1bfb12502f3fc46cca1bb51ac28df9d618d813cdc3d2f25b9fe775a34af26bb3" checksum = "78c8292055d1c1df0cce5d180393dc8cce0abec0a7102adb6c7b1eef6016d60a"
dependencies = [ dependencies = [
"generic-array 0.14.9", "generic-array 0.14.7",
"typenum", "typenum",
] ]
@@ -3664,7 +3664,7 @@ checksum = "9ed9a281f7bc9b7576e61468ba615a66a5c8cfdff42420a70aa82701a3b1e292"
dependencies = [ dependencies = [
"block-buffer 0.10.4", "block-buffer 0.10.4",
"const-oid 0.9.6", "const-oid 0.9.6",
"crypto-common 0.1.6", "crypto-common 0.1.7",
"subtle", "subtle",
] ]
@@ -3924,7 +3924,7 @@ dependencies = [
"crypto-bigint 0.5.5", "crypto-bigint 0.5.5",
"digest 0.10.7", "digest 0.10.7",
"ff 0.13.1", "ff 0.13.1",
"generic-array 0.14.9", "generic-array 0.14.7",
"group 0.13.0", "group 0.13.0",
"hkdf 0.12.4", "hkdf 0.12.4",
"pem-rfc7468 0.7.0", "pem-rfc7468 0.7.0",
@@ -4148,9 +4148,9 @@ dependencies = [
[[package]] [[package]]
name = "find-msvc-tools" name = "find-msvc-tools"
version = "0.1.10" version = "0.1.11"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "26b73573e6edcd2af0cdf47bd6cb58f0b3839491263c314eaad1ccf24430e1de" checksum = "d45db016d36b838f563236e9193d0ee6ce38f3f68b6c94e914b4929c96bbb890"
[[package]] [[package]]
name = "findshlibs" name = "findshlibs"
@@ -4369,9 +4369,9 @@ dependencies = [
[[package]] [[package]]
name = "generic-array" name = "generic-array"
version = "0.14.9" version = "0.14.7"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "4bb6743198531e02858aeaea5398fcc883e71851fcbcb5a2f773e2fb6cb1edf2" checksum = "85649ca51fd72272d7821adaf274ad91c288277713d9c18820d8499a7ff69e9a"
dependencies = [ dependencies = [
"typenum", "typenum",
"version_check", "version_check",
@@ -4380,11 +4380,11 @@ dependencies = [
[[package]] [[package]]
name = "generic-array" name = "generic-array"
version = "1.4.4" version = "1.4.5"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "ab4e5aa225bc56696909483320f0ff9b600f1a971b52e07a17d70f3d9b43254b" checksum = "337d46834ee672ab3e48caca2cb0c78cc174fb12b3a68d0d88f99a0519a5e36e"
dependencies = [ dependencies = [
"generic-array 0.14.9", "generic-array 0.14.7",
"rustversion", "rustversion",
"typenum", "typenum",
] ]
@@ -4726,9 +4726,9 @@ dependencies = [
[[package]] [[package]]
name = "h2" name = "h2"
version = "0.4.15" version = "0.4.16"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "6cb093c84e8bd9b188d4c4a8cb6579fc016968d14c99882163cd3ff402a4f155" checksum = "a9f37a958b41b3b19ee2707c06439c0e9e547e847223eb791ecb0cb821c65e27"
dependencies = [ dependencies = [
"atomic-waker", "atomic-waker",
"bytes", "bytes",
@@ -5028,9 +5028,9 @@ dependencies = [
[[package]] [[package]]
name = "hotpath" name = "hotpath"
version = "0.23.2" version = "0.23.3"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "62e810bedda5a467ef5c9b5c8a20763fefebc89b63ef36f7ee44a143085204a2" checksum = "dce755d457a63bdd0c95e4c91511daad1b58b33209543b7f38027b676f387e5e"
dependencies = [ dependencies = [
"arc-swap", "arc-swap",
"async-channel", "async-channel",
@@ -5062,9 +5062,9 @@ dependencies = [
[[package]] [[package]]
name = "hotpath-macros" name = "hotpath-macros"
version = "0.23.2" version = "0.23.3"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "01bdc59bfc1a9984bee2ff5da63b2f6fccbaa57cd9a4119d709524632bddf341" checksum = "a903af89a8429cb07790c3818bc15270b394f80af1bc254e5ccf9c7de2961770"
dependencies = [ dependencies = [
"proc-macro2", "proc-macro2",
"quote", "quote",
@@ -5073,15 +5073,15 @@ dependencies = [
[[package]] [[package]]
name = "hotpath-macros-meta" name = "hotpath-macros-meta"
version = "0.23.2" version = "0.23.3"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "d9216e8a01abe1e1671c376dc8736fb1bf772d7a889538d25f9e1200120ced38" checksum = "bcc0ab94ffbb2ee77f4a897df02b5a137a10cf24d69bda936e59aff4dd456e61"
[[package]] [[package]]
name = "hotpath-meta" name = "hotpath-meta"
version = "0.23.2" version = "0.23.3"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "f22a9d20435fb79511b19dae37b3607224cd98f342a410702d84657cc38fc72f" checksum = "053481f6cec8f775a3276c7f6e2f21123111d28261e4edc15ea7421c445964bb"
dependencies = [ dependencies = [
"hotpath-macros-meta", "hotpath-macros-meta",
] ]
@@ -5280,9 +5280,9 @@ dependencies = [
[[package]] [[package]]
name = "icu_collections" name = "icu_collections"
version = "2.2.0" version = "2.3.0"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "2984d1cd16c883d7935b9e07e44071dca8d917fd52ecc02c04d5fa0b5a3f191c" checksum = "fa68d21081c4a05d5a901a1c62add574c77048b6a1c67be3b50ce0b60d4ca513"
dependencies = [ dependencies = [
"displaydoc", "displaydoc",
"potential_utf", "potential_utf",
@@ -5294,9 +5294,9 @@ dependencies = [
[[package]] [[package]]
name = "icu_locale_core" name = "icu_locale_core"
version = "2.2.0" version = "2.3.0"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "92219b62b3e2b4d88ac5119f8904c10f8f61bf7e95b640d25ba3075e6cac2c29" checksum = "d56e28588da92eee5c3201a6eff33fabdd49b62269c8938d4ff050ce4d900deb"
dependencies = [ dependencies = [
"displaydoc", "displaydoc",
"litemap", "litemap",
@@ -5307,9 +5307,9 @@ dependencies = [
[[package]] [[package]]
name = "icu_normalizer" name = "icu_normalizer"
version = "2.2.0" version = "2.3.0"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "c56e5ee99d6e3d33bd91c5d85458b6005a22140021cc324cea84dd0e72cff3b4" checksum = "12f9cf5f235641ed274641dd81c3f28d870e276763d0797aeeab72317b1c646f"
dependencies = [ dependencies = [
"icu_collections", "icu_collections",
"icu_normalizer_data", "icu_normalizer_data",
@@ -5321,16 +5321,17 @@ dependencies = [
[[package]] [[package]]
name = "icu_normalizer_data" name = "icu_normalizer_data"
version = "2.2.0" version = "2.3.0"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "da3be0ae77ea334f4da67c12f149704f19f81d1adf7c51cf482943e84a2bad38" checksum = "1563da1ed3e0b3bf3d74c9b85917ac9c56464d2f57242270c09c9e752f8021a0"
[[package]] [[package]]
name = "icu_properties" name = "icu_properties"
version = "2.2.0" version = "2.3.0"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "bee3b67d0ea5c2cca5003417989af8996f8604e34fb9ddf96208a033901e70de" checksum = "7e7ca276ad3145661a65914e6daf131ca5120cd3dcee8f8f3214b8875184a148"
dependencies = [ dependencies = [
"displaydoc",
"icu_collections", "icu_collections",
"icu_locale_core", "icu_locale_core",
"icu_properties_data", "icu_properties_data",
@@ -5341,15 +5342,15 @@ dependencies = [
[[package]] [[package]]
name = "icu_properties_data" name = "icu_properties_data"
version = "2.2.0" version = "2.3.0"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "8e2bbb201e0c04f7b4b3e14382af113e17ba4f63e2c9d2ee626b720cbce54a14" checksum = "e590f038c1464a96894fd6d10127e90a8be4509f56ff7ecef851b15cee0b7caa"
[[package]] [[package]]
name = "icu_provider" name = "icu_provider"
version = "2.2.0" version = "2.3.0"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "139c4cf31c8b5f33d7e199446eff9c1e02decfc2f0eec2c8d71f65befa45b421" checksum = "92a7ed671a6aad807a8651a2e1782a6598fda9ce5185dd8158549e95a91c6428"
dependencies = [ dependencies = [
"displaydoc", "displaydoc",
"icu_locale_core", "icu_locale_core",
@@ -5417,7 +5418,7 @@ source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "879f10e63c20629ecabbb64a8010319738c66a5cd0c29b02d63d272b03751d01" checksum = "879f10e63c20629ecabbb64a8010319738c66a5cd0c29b02d63d272b03751d01"
dependencies = [ dependencies = [
"block-padding 0.3.3", "block-padding 0.3.3",
"generic-array 0.14.9", "generic-array 0.14.7",
] ]
[[package]] [[package]]
@@ -5968,9 +5969,9 @@ dependencies = [
[[package]] [[package]]
name = "libredox" name = "libredox"
version = "0.1.19" version = "0.1.20"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "2026a5056764a10b2bf5d56488cba40da507f5493a6a429340e2004d9ed085fa" checksum = "28d0a00925a9f930d679b6789b721e3a7f9ed110f41b86d2497caa780c3a070a"
dependencies = [ dependencies = [
"libc", "libc",
] ]
@@ -6033,9 +6034,9 @@ checksum = "32a66949e030da00e8c7d4434b251670a91556f4144941d37452769c25d58a53"
[[package]] [[package]]
name = "litemap" name = "litemap"
version = "0.8.2" version = "0.8.3"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "92daf443525c4cce67b150400bc2316076100ce0b3686209eb8cf3c31612e6f0" checksum = "47d9d19d1d6efa0109d2f65ff4c85cddd50bd572e5a00127ab10987290bcefae"
[[package]] [[package]]
name = "local-ip-address" name = "local-ip-address"
@@ -6482,9 +6483,9 @@ dependencies = [
[[package]] [[package]]
name = "mqttbytes-core-next" name = "mqttbytes-core-next"
version = "0.33.3" version = "0.34.0"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "3ff7ae19c74aba9e0ed6e4071cd52aa364e020076fa3cc6ef17e43662f756f3c" checksum = "366b6ba2b4209ca4bc5ac731ccddf570d09831981eed07e5fbd63564cf0cf1aa"
dependencies = [ dependencies = [
"bytes", "bytes",
"thiserror 2.0.20", "thiserror 2.0.20",
@@ -6885,7 +6886,7 @@ version = "5.0.0"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "51e219e79014df21a225b1860a479e2dcd7cbd9130f4defd4bd0e191ea31d67d" checksum = "51e219e79014df21a225b1860a479e2dcd7cbd9130f4defd4bd0e191ea31d67d"
dependencies = [ dependencies = [
"base64 0.21.7", "base64 0.22.1",
"chrono", "chrono",
"getrandom 0.2.17", "getrandom 0.2.17",
"http 1.5.0", "http 1.5.0",
@@ -7344,9 +7345,9 @@ dependencies = [
[[package]] [[package]]
name = "pageant" name = "pageant"
version = "0.2.1" version = "0.2.2"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "4f3a5ae18f65a85c67a77d18d42d3606c07948e3c17c1e5f74852b26589e88a5" checksum = "3adadc44070da6f464b0918655a12f5792c156e088d8c4082d13e27d94c3e791"
dependencies = [ dependencies = [
"base16ct 1.0.0", "base16ct 1.0.0",
"byteorder", "byteorder",
@@ -7728,9 +7729,9 @@ dependencies = [
[[package]] [[package]]
name = "pkg-config" name = "pkg-config"
version = "0.3.33" version = "0.3.34"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "19f132c84eca552bf34cab8ec81f1c1dcc229b811638f9d283dceabe58c5569e" checksum = "f6b464fbc74e149a392436b17d523f769e057cb6877f6a5c4618bc6f11800548"
[[package]] [[package]]
name = "plotters" name = "plotters"
@@ -7836,9 +7837,9 @@ dependencies = [
[[package]] [[package]]
name = "potential_utf" name = "potential_utf"
version = "0.1.5" version = "0.1.6"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "0103b1cef7ec0cf76490e969665504990193874ea05c85ff9bab8b911d0a0564" checksum = "d83eb9bc6d8e5cf568e7a1101d60ee05e81ed50ea106026f3d18deeb046d7661"
dependencies = [ dependencies = [
"zerovec", "zerovec",
] ]
@@ -8046,7 +8047,7 @@ source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "be769465445e8c1474e9c5dac2018218498557af32d9ed057325ec9a41ae81bf" checksum = "be769465445e8c1474e9c5dac2018218498557af32d9ed057325ec9a41ae81bf"
dependencies = [ dependencies = [
"heck 0.5.0", "heck 0.5.0",
"itertools 0.10.5", "itertools 0.14.0",
"log", "log",
"multimap", "multimap",
"once_cell", "once_cell",
@@ -8066,7 +8067,7 @@ source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "03da047801ff44bb6a4d407d4860c05fd70bb81714e6b2f3812603d5b145b042" checksum = "03da047801ff44bb6a4d407d4860c05fd70bb81714e6b2f3812603d5b145b042"
dependencies = [ dependencies = [
"heck 0.5.0", "heck 0.5.0",
"itertools 0.10.5", "itertools 0.14.0",
"log", "log",
"multimap", "multimap",
"petgraph 0.8.3", "petgraph 0.8.3",
@@ -8087,7 +8088,7 @@ source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "8a56d757972c98b346a9b766e3f02746cde6dd1cd1d1d563472929fdd74bec4d" checksum = "8a56d757972c98b346a9b766e3f02746cde6dd1cd1d1d563472929fdd74bec4d"
dependencies = [ dependencies = [
"anyhow", "anyhow",
"itertools 0.10.5", "itertools 0.14.0",
"proc-macro2", "proc-macro2",
"quote", "quote",
"syn 2.0.119", "syn 2.0.119",
@@ -8100,7 +8101,7 @@ source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "b570b25f7617e43d59005d0990ccb79e950a423952cea19671b7a876da390adf" checksum = "b570b25f7617e43d59005d0990ccb79e950a423952cea19671b7a876da390adf"
dependencies = [ dependencies = [
"anyhow", "anyhow",
"itertools 0.10.5", "itertools 0.14.0",
"proc-macro2", "proc-macro2",
"quote", "quote",
"syn 2.0.119", "syn 2.0.119",
@@ -8216,7 +8217,7 @@ dependencies = [
"reqwest", "reqwest",
"serde_json", "serde_json",
"smallvec", "smallvec",
"spin 0.12.2", "spin 0.12.3",
"symbolic-demangle", "symbolic-demangle",
"tempfile", "tempfile",
"thiserror 2.0.20", "thiserror 2.0.20",
@@ -8285,9 +8286,9 @@ dependencies = [
[[package]] [[package]]
name = "quinn-proto" name = "quinn-proto"
version = "0.11.16" version = "0.11.17"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "2f4bfc015262b9df63c8845072ce59068853ff5872180c2ce2f13038b970e560" checksum = "04759210543be93709136e28212294a659ef5001836ff4eab4d663e4529bba83"
dependencies = [ dependencies = [
"aws-lc-rs", "aws-lc-rs",
"bytes", "bytes",
@@ -8553,9 +8554,9 @@ dependencies = [
[[package]] [[package]]
name = "redis" name = "redis"
version = "1.5.0" version = "1.6.0"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "3257df217f7eab0044627a268c9cc6cdb60c0c421c88f83ac41c4e31520b6b84" checksum = "e37a4ca5c6ca42aa3e6df2fd32b987a65d32a4c2159a6f3fe0fd1df306a2658f"
dependencies = [ dependencies = [
"arc-swap", "arc-swap",
"arcstr", "arcstr",
@@ -8567,7 +8568,7 @@ dependencies = [
"futures-channel", "futures-channel",
"futures-util", "futures-util",
"itoa", "itoa",
"num-bigint 0.4.8", "num-bigint 0.5.1",
"percent-encoding", "percent-encoding",
"pin-project-lite", "pin-project-lite",
"rustls", "rustls",
@@ -8868,9 +8869,9 @@ dependencies = [
[[package]] [[package]]
name = "rumqttc-core-next" name = "rumqttc-core-next"
version = "0.33.3" version = "0.34.0"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "7d7d9205738dd41a2546e82d27a634d07d8b303dcf7558565ff70caf3ceb0f9c" checksum = "249896ab27ed630590971738264baa8f722f18965d2e387c706c40a3c2a572cc"
dependencies = [ dependencies = [
"async-tungstenite", "async-tungstenite",
"futures-io", "futures-io",
@@ -8886,18 +8887,18 @@ dependencies = [
[[package]] [[package]]
name = "rumqttc-next" name = "rumqttc-next"
version = "0.33.3" version = "0.34.0"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "ed1bad2180ff539da671da9a996152a921bc5316eb6d8a9cc3bd441653138b08" checksum = "477c9bbfba8f3aecc7aad31c6de2eacb75822efaa18e7aeecb8d3d8e534fbf07"
dependencies = [ dependencies = [
"rumqttc-v5-next", "rumqttc-v5-next",
] ]
[[package]] [[package]]
name = "rumqttc-v5-next" name = "rumqttc-v5-next"
version = "0.33.3" version = "0.34.0"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "229576cbedfa9089f90c17c9454e9429ac1e89cdd223bac5cb39d837593f79bc" checksum = "3dfa6ddcc7a7dd5688f9bf78d8f81cb94f367bce56c055d8d94cf81ecb0518bf"
dependencies = [ dependencies = [
"async-tungstenite", "async-tungstenite",
"bytes", "bytes",
@@ -8920,9 +8921,9 @@ dependencies = [
[[package]] [[package]]
name = "russh" name = "russh"
version = "0.62.6" version = "0.62.7"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "b41043523e0edcbd4e31d00903e26f12994f63b21bae9904f7405c1ed92752a5" checksum = "9decb68e4e44e1079700e54f17c8f23806ec53d7e0db73ab1c71d9dabc666812"
dependencies = [ dependencies = [
"aes 0.9.2", "aes 0.9.2",
"aws-lc-rs", "aws-lc-rs",
@@ -8945,7 +8946,7 @@ dependencies = [
"enum_dispatch", "enum_dispatch",
"flate2", "flate2",
"futures", "futures",
"generic-array 1.4.4", "generic-array 1.4.5",
"getrandom 0.4.3", "getrandom 0.4.3",
"ghash", "ghash",
"hex-literal", "hex-literal",
@@ -9280,6 +9281,7 @@ dependencies = [
"s3s", "s3s",
"serde", "serde",
"serde_json", "serde_json",
"smallvec",
"tokio", "tokio",
"tonic", "tonic",
"tracing", "tracing",
@@ -9463,6 +9465,7 @@ dependencies = [
"tokio-stream", "tokio-stream",
"tokio-util", "tokio-util",
"tonic", "tonic",
"tonic-prost",
"tower", "tower",
"tracing", "tracing",
"tracing-core", "tracing-core",
@@ -9490,7 +9493,7 @@ dependencies = [
"parking_lot", "parking_lot",
"rayon", "rayon",
"smallvec", "smallvec",
"spin 0.12.2", "spin 0.12.3",
] ]
[[package]] [[package]]
@@ -9536,6 +9539,8 @@ version = "1.0.0-rc.2"
dependencies = [ dependencies = [
"async-trait", "async-trait",
"base64 0.23.1", "base64 0.23.1",
"bytes",
"crc-fast",
"futures", "futures",
"hotpath", "hotpath",
"http 1.5.0", "http 1.5.0",
@@ -9608,7 +9613,6 @@ version = "1.0.0-rc.2"
dependencies = [ dependencies = [
"bytes", "bytes",
"hotpath", "hotpath",
"memmap2",
"rustfs-io-metrics", "rustfs-io-metrics",
"thiserror 2.0.20", "thiserror 2.0.20",
"tokio", "tokio",
@@ -10251,6 +10255,7 @@ dependencies = [
"rustfs-ecstore", "rustfs-ecstore",
"rustfs-filemeta", "rustfs-filemeta",
"rustfs-lock", "rustfs-lock",
"rustfs-s3-types",
"rustfs-storage-api", "rustfs-storage-api",
"rustfs-utils", "rustfs-utils",
"s3s", "s3s",
@@ -10831,7 +10836,7 @@ checksum = "d3e97a565f76233a6003f9f5c54be1d9c5bdfa3eccfb189469f11ec4901c47dc"
dependencies = [ dependencies = [
"base16ct 0.2.0", "base16ct 0.2.0",
"der 0.7.10", "der 0.7.10",
"generic-array 0.14.9", "generic-array 0.14.7",
"pkcs8 0.10.2", "pkcs8 0.10.2",
"subtle", "subtle",
"zeroize", "zeroize",
@@ -11395,9 +11400,9 @@ checksum = "023a211cb3138dbc438680b32560ad89f699977624c9f8dbb95a47d5b4c07dd3"
[[package]] [[package]]
name = "spin" name = "spin"
version = "0.12.2" version = "0.12.3"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "8abadc99fd9c7bbb7d0ca2b31d72a067d0c0dcd7aad25ab8cac71ba91417694b" checksum = "0134f9043ed38b087ac4f7d4af44c79e2c9e5094421fe3164f435ce585953b10"
dependencies = [ dependencies = [
"lock_api", "lock_api",
] ]
@@ -11809,7 +11814,7 @@ source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "32497e9a4c7b38532efcdebeef879707aa9f794296a4f0244f6f69e9bc8574bd" checksum = "32497e9a4c7b38532efcdebeef879707aa9f794296a4f0244f6f69e9bc8574bd"
dependencies = [ dependencies = [
"fastrand", "fastrand",
"getrandom 0.3.4", "getrandom 0.4.3",
"once_cell", "once_cell",
"rustix", "rustix",
"windows-sys 0.61.2", "windows-sys 0.61.2",
@@ -11973,9 +11978,9 @@ dependencies = [
[[package]] [[package]]
name = "tinystr" name = "tinystr"
version = "0.8.3" version = "0.8.4"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "c8323304221c2a851516f22236c5722a72eaa19749016521d6dff0824447d96d" checksum = "b1e27c91459209c2986af3dcf603a5a74a4368754ce37414f59acc971167f643"
dependencies = [ dependencies = [
"displaydoc", "displaydoc",
"zerovec", "zerovec",
@@ -12649,9 +12654,9 @@ checksum = "06abde3611657adf66d383f00b093d7faecc7fa57071cce2578660c9f1010821"
[[package]] [[package]]
name = "uuid" name = "uuid"
version = "1.24.0" version = "1.24.1"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "bf3923a6f5c4c6382e0b653c4117f48d631ea17f38ed86e2a828e6f7412f5239" checksum = "2cefc03fd367c0c6d4305de1b312cf00248c4114f4a0418ce6a6af769e3b0bd9"
dependencies = [ dependencies = [
"getrandom 0.4.3", "getrandom 0.4.3",
"js-sys", "js-sys",
@@ -13166,9 +13171,9 @@ dependencies = [
[[package]] [[package]]
name = "writeable" name = "writeable"
version = "0.6.3" version = "0.6.4"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "1ffae5123b2d3fc086436f8834ae3ab053a283cfac8fe0a0b8eaae044768a4c4" checksum = "3ad82d2a33cdc9674dc7465672f271e096168fcdbe0f799d9e6db8c5892679dc"
[[package]] [[package]]
name = "x509-cert" name = "x509-cert"
@@ -13348,9 +13353,9 @@ dependencies = [
[[package]] [[package]]
name = "zerotrie" name = "zerotrie"
version = "0.2.4" version = "0.2.5"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "0f9152d31db0792fa83f70fb2f83148effb5c1f5b8c7686c3459e361d9bc20bf" checksum = "4ea269c3bd32f0a32c321907a2ae912ba6f4649bb0fc764a15627e99a7095a3f"
dependencies = [ dependencies = [
"displaydoc", "displaydoc",
"yoke", "yoke",
@@ -13359,9 +13364,9 @@ dependencies = [
[[package]] [[package]]
name = "zerovec" name = "zerovec"
version = "0.11.6" version = "0.11.7"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "90f911cbc359ab6af17377d242225f4d75119aec87ea711a880987b18cd7b239" checksum = "94b5c6b5976d66c1d703c4fd17d3f5e43c8cedaacf604961b171adc7130896d8"
dependencies = [ dependencies = [
"yoke", "yoke",
"zerofrom", "zerofrom",
@@ -13370,13 +13375,13 @@ dependencies = [
[[package]] [[package]]
name = "zerovec-derive" name = "zerovec-derive"
version = "0.11.3" version = "0.11.5"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "625dc425cab0dca6dc3c3319506e6593dcb08a9f387ea3b284dbd52a92c40555" checksum = "9f212a141d820099d57ffafb9569be9617a6f27d3dc881fbee8fb56642f917a9"
dependencies = [ dependencies = [
"proc-macro2", "proc-macro2",
"quote", "quote",
"syn 2.0.119", "syn 3.0.3",
] ]
[[package]] [[package]]
+8 -8
View File
@@ -228,9 +228,9 @@ atoi = "3.1.0"
atomic_enum = "0.3.0" atomic_enum = "0.3.0"
aws-config = { version = "1.10.1" } aws-config = { version = "1.10.1" }
aws-credential-types = { version = "1.3.0" } aws-credential-types = { version = "1.3.0" }
aws-sdk-kms = { default-features = false, version = "1.114.0" } aws-sdk-kms = { default-features = false, version = "1.115.0" }
aws-sdk-s3 = { default-features = false, version = "1.141.0" } aws-sdk-s3 = { default-features = false, version = "1.142.0" }
aws-sdk-sts = { default-features = false, version = "1.110.0" } aws-sdk-sts = { default-features = false, version = "1.111.0" }
aws-smithy-http-client = { default-features = false, version = "1.3.0" } aws-smithy-http-client = { default-features = false, version = "1.3.0" }
aws-smithy-runtime-api = { version = "1.14.0" } aws-smithy-runtime-api = { version = "1.14.0" }
aws-smithy-types = { version = "1.6.2" } aws-smithy-types = { version = "1.6.2" }
@@ -284,8 +284,8 @@ rayon = "1.12.0"
reed-solomon-erasure = { package = "rustfs-erasure-codec", version = "8.0.2" } reed-solomon-erasure = { package = "rustfs-erasure-codec", version = "8.0.2" }
reed-solomon-simd = "3.1.0" reed-solomon-simd = "3.1.0"
regex = { version = "1.13.1" } regex = { version = "1.13.1" }
rumqttc = { package = "rumqttc-next", version = "0.33.3" } rumqttc = { package = "rumqttc-next", version = "0.34.0" }
redis = { version = "1.5.0" } redis = { version = "1.6.0" }
rustify = { version = "0.7", default-features = false } rustify = { version = "0.7", default-features = false }
rustix = { version = "1.1.4" } rustix = { version = "1.1.4" }
rust-embed = { version = "8.12.0" } rust-embed = { version = "8.12.0" }
@@ -313,7 +313,7 @@ tracing-subscriber = { version = "0.3.23" }
transform-stream = "0.3.1" transform-stream = "0.3.1"
url = "2.5.8" url = "2.5.8"
urlencoding = "2.1.3" urlencoding = "2.1.3"
uuid = { version = "1.24.0" } uuid = { version = "1.24.1" }
vaultrs = { version = "0.8.0" } vaultrs = { version = "0.8.0" }
tar = "0.4.46" tar = "0.4.46"
walkdir = "2.5.0" walkdir = "2.5.0"
@@ -341,7 +341,7 @@ libunftp = { version = "0.23.0" }
unftp-core = "0.1.0" unftp-core = "0.1.0"
suppaftp = { version = "10.0.1" } suppaftp = { version = "10.0.1" }
rcgen = { version = "0.14.9", default-features = false, features = ["aws_lc_rs", "crypto", "pem"] } rcgen = { version = "0.14.9", default-features = false, features = ["aws_lc_rs", "crypto", "pem"] }
russh = { version = "0.62.6" } russh = { version = "0.62.7" }
russh-sftp = "2.4.0" russh-sftp = "2.4.0"
# WebDAV # WebDAV
@@ -350,7 +350,7 @@ dav-server = "0.11.0"
# Performance Analysis and Memory Profiling # Performance Analysis and Memory Profiling
mimalloc = { version = "0.1.52", git = "https://github.com/xonatius/mimalloc_rust.git", rev = "6d4c41bb10c6d9da1d1b6f07b38c4cc051667f11" } mimalloc = { version = "0.1.52", git = "https://github.com/xonatius/mimalloc_rust.git", rev = "6d4c41bb10c6d9da1d1b6f07b38c4cc051667f11" }
libmimalloc-sys = { version = "0.1.49", git = "https://github.com/xonatius/mimalloc_rust.git", rev = "6d4c41bb10c6d9da1d1b6f07b38c4cc051667f11", features = ["extended"] } libmimalloc-sys = { version = "0.1.49", git = "https://github.com/xonatius/mimalloc_rust.git", rev = "6d4c41bb10c6d9da1d1b6f07b38c4cc051667f11", features = ["extended"] }
hotpath = { version = "0.23.2", default-features = false } hotpath = { version = "0.23.3", default-features = false }
# Snapshot testing for output format regression detection # Snapshot testing for output format regression detection
insta = { version = "1.48" } insta = { version = "1.48" }
+1
View File
@@ -40,6 +40,7 @@ mak = "mak"
gae = "gae" gae = "gae"
GAE = "GAE" GAE = "GAE"
thr = "thr" thr = "thr"
mis = "mis"
# s3-tests original test names (cannot be changed) # s3-tests original test names (cannot be changed)
nonexisted = "nonexisted" nonexisted = "nonexisted"
consts = "consts" consts = "consts"
+1
View File
@@ -42,6 +42,7 @@ chrono = { workspace = true, features = ["serde"] }
jiff = { workspace = true, features = ["serde"] } jiff = { workspace = true, features = ["serde"] }
metrics = { workspace = true } metrics = { workspace = true }
serde = { workspace = true, features = ["derive"] } serde = { workspace = true, features = ["derive"] }
smallvec = { workspace = true }
rmp-serde = { workspace = true } rmp-serde = { workspace = true }
s3s = { workspace = true, features = ["minio"] } s3s = { workspace = true, features = ["minio"] }
tracing = { workspace = true } tracing = { workspace = true }
+27
View File
@@ -224,6 +224,13 @@ pub struct HealOpts {
pub enum HealAdmissionDropReason { pub enum HealAdmissionDropReason {
QueueFull, QueueFull,
PolicyDropped, PolicyDropped,
/// HS-06: an admin heal start overlaps (same bucket with mutually
/// containing prefixes, or the same erasure set) an already running or
/// queued task. Only produced when RUSTFS_HEAL_OVERLAP_POLICY=minio_error.
AlreadyRunning,
/// HS-06: same as [`Self::AlreadyRunning`] but for paths that merely
/// contain (or are contained by) the active task's path.
OverlappingPaths,
} }
impl HealAdmissionDropReason { impl HealAdmissionDropReason {
@@ -231,6 +238,8 @@ impl HealAdmissionDropReason {
match self { match self {
Self::QueueFull => "queue_full", Self::QueueFull => "queue_full",
Self::PolicyDropped => "policy_dropped", Self::PolicyDropped => "policy_dropped",
Self::AlreadyRunning => "already_running",
Self::OverlappingPaths => "overlapping_paths",
} }
} }
} }
@@ -287,6 +296,9 @@ pub enum HealRequestSource {
Scanner, Scanner,
AutoHeal, AutoHeal,
ReadRepair, ReadRepair,
/// Mission Repair Feed: intents delivered by error paths and replayed
/// from the durable MRF journal.
Mrf,
} }
impl HealRequestSource { impl HealRequestSource {
@@ -297,6 +309,7 @@ impl HealRequestSource {
Self::Scanner => "scanner", Self::Scanner => "scanner",
Self::AutoHeal => "auto_heal", Self::AutoHeal => "auto_heal",
Self::ReadRepair => "read_repair", Self::ReadRepair => "read_repair",
Self::Mrf => "mrf",
} }
} }
} }
@@ -313,6 +326,9 @@ pub enum HealChannelCommand {
Query { Query {
heal_path: String, heal_path: String,
client_token: String, client_token: String,
/// Incremental result cursor (HS-06): only items with a sequence
/// greater than this are returned; `None` keeps the full snapshot.
since_seq: Option<u64>,
response_tx: oneshot::Sender<Result<HealChannelResponse, String>>, response_tx: oneshot::Sender<Result<HealChannelResponse, String>>,
}, },
/// Cancel heal task /// Cancel heal task
@@ -518,10 +534,21 @@ async fn receive_heal_channel_response(
/// Send heal query request /// Send heal query request
pub async fn query_heal_status(heal_path: String, client_token: String) -> Result<HealChannelResponse, String> { pub async fn query_heal_status(heal_path: String, client_token: String) -> Result<HealChannelResponse, String> {
query_heal_status_since(heal_path, client_token, None).await
}
/// Incremental heal query (HS-06): pass the client's last seen sequence
/// number to receive only newer result items.
pub async fn query_heal_status_since(
heal_path: String,
client_token: String,
since_seq: Option<u64>,
) -> Result<HealChannelResponse, String> {
let (response_tx, response_rx) = oneshot::channel(); let (response_tx, response_rx) = oneshot::channel();
send_heal_command(HealChannelCommand::Query { send_heal_command(HealChannelCommand::Query {
heal_path, heal_path,
client_token, client_token,
since_seq,
response_tx, response_tx,
}) })
.await?; .await?;
+2
View File
@@ -17,8 +17,10 @@ pub mod globals;
pub mod heal_channel; pub mod heal_channel;
pub mod last_minute; pub mod last_minute;
pub mod metrics; pub mod metrics;
pub mod mrf_channel;
mod readiness; mod readiness;
pub mod table_catalog; pub mod table_catalog;
pub mod trace_bus;
pub use globals::*; pub use globals::*;
pub use readiness::{GlobalReadiness, SystemStage}; pub use readiness::{GlobalReadiness, SystemStage};
+203
View File
@@ -0,0 +1,203 @@
// Copyright 2024 RustFS Team
//
// Licensed under the Apache License, Version 2.0 (the "License");
// you may not use this file except in compliance with the License.
// You may obtain a copy of the License at
//
// http://www.apache.org/licenses/LICENSE-2.0
//
// Unless required by applicable law or agreed to in writing, software
// distributed under the License is distributed on an "AS IS" BASIS,
// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
// See the License for the specific language governing permissions and
// limitations under the License.
//! Mission Repair Feed (MRF) intent channel.
//!
//! Producers on error paths (read decode failure, scanner metadata
//! corruption, partial-write recovery) hand a lightweight [`MrfIntent`] to the
//! heal crate through a global bounded channel. Delivery is strictly
//! non-blocking: `try_send_mrf_intent` never awaits and drops the intent
//! (counting it) when the channel is full or uninitialized — losing one heal
//! hint is always preferred over stalling an IO path. Durable replay of
//! unconsumed intents is the consumer's job (see `rustfs-heal`
//! `heal::mrf_queue`), mirroring MinIO's `.heal/mrf/list.bin`.
use std::sync::{
Arc, OnceLock,
atomic::{AtomicBool, Ordering},
};
use tokio::sync::mpsc;
use uuid::Uuid;
/// Bounded capacity of the global MRF channel. Backpressure is resolved by
/// dropping (and counting) intents, never by blocking the producer.
const MRF_CHANNEL_CAPACITY: usize = 8192;
/// Why an intent was produced. Drives the heal priority mapping on the
/// consumer side (DecodeFailure -> Urgent, MetadataCorruption -> High,
/// PartialWrite -> Normal).
#[derive(Clone, Copy, Debug, PartialEq, Eq)]
pub enum MrfKind {
/// Erasure decode failed while serving a read (read path).
DecodeFailure,
/// Scanner classified object metadata as corrupt.
MetadataCorruption,
/// A write left the object with fewer committed shards than the set size.
PartialWrite,
}
impl MrfKind {
pub const fn as_str(self) -> &'static str {
match self {
MrfKind::DecodeFailure => "decode-failure",
MrfKind::MetadataCorruption => "metadata-corruption",
MrfKind::PartialWrite => "partial-write",
}
}
}
/// One repair intent. Kept deliberately small so the in-memory queue and the
/// journal stay bounded; `bucket`/`object` are `Arc<str>` so re-arming an
/// intent never re-allocates the strings.
#[derive(Clone, Debug)]
pub struct MrfIntent {
pub bucket: Arc<str>,
pub object: Arc<str>,
/// Version the intent targets, as raw UUID bytes.
pub version_id: Option<[u8; 16]>,
pub kind: MrfKind,
pub enqueued_at_ms: u64,
/// Times this intent has already been offered to the heal manager.
/// Dropped by the consumer once it reaches `MRF_MAX_ATTEMPTS`.
pub attempts: u8,
}
/// Consumer-side retry ceiling before an intent is given up on.
pub const MRF_MAX_ATTEMPTS: u8 = 3;
impl MrfIntent {
/// Rough in-memory footprint used by the queue's byte budget.
pub fn estimated_bytes(&self) -> usize {
// Struct + strings + version bytes; buckets and objects are usually
// far below this bound, so rounding up keeps the budget conservative.
64 + self.bucket.len() + self.object.len()
}
}
static GLOBAL_MRF_SENDER: OnceLock<mpsc::Sender<MrfIntent>> = OnceLock::new();
/// Delivery kill-switch, set from `RUSTFS_HEAL_MRF_ENABLE`. Producers check
/// this before touching the channel so the disabled path stays allocation- and
/// sync-free.
static MRF_DELIVERY_ENABLED: AtomicBool = AtomicBool::new(true);
/// Override delivery (used at heal-runtime startup from configuration).
pub fn set_mrf_delivery_enabled(enabled: bool) {
MRF_DELIVERY_ENABLED.store(enabled, Ordering::Relaxed);
}
/// Whether producers currently deliver intents.
pub fn mrf_delivery_enabled() -> bool {
MRF_DELIVERY_ENABLED.load(Ordering::Relaxed)
}
/// Create the global MRF channel and return the consumer half. Fails if the
/// channel is already initialized (the heal runtime is a singleton).
pub fn init_mrf_channel() -> Result<mpsc::Receiver<MrfIntent>, &'static str> {
let (sender, receiver) = mpsc::channel(MRF_CHANNEL_CAPACITY);
GLOBAL_MRF_SENDER
.set(sender)
.map_err(|_| "MRF channel sender already initialized")?;
Ok(receiver)
}
/// Best-effort, non-blocking intent delivery from an error path.
///
/// Returns `true` when the intent was accepted into the channel. `false`
/// means the intent was dropped (feature disabled, channel not yet
/// initialized, or channel full) — callers must not retry or await; the
/// existing read-repair / scanner heal paths remain the safety net.
///
/// This runs on IO error paths, so it stays synchronous and cheap: one
/// bounded allocation for the two `Arc<str>` handles plus the channel slot.
pub fn try_send_mrf_intent(kind: MrfKind, bucket: &str, object: &str, version_id: Option<Uuid>) -> bool {
if !mrf_delivery_enabled() {
return false;
}
let Some(sender) = GLOBAL_MRF_SENDER.get() else {
return false;
};
let intent = MrfIntent {
bucket: Arc::from(bucket),
object: Arc::from(object),
version_id: version_id.map(|vid| *vid.as_bytes()),
kind,
enqueued_at_ms: unix_now_ms(),
attempts: 0,
};
sender.try_send(intent).is_ok()
}
fn unix_now_ms() -> u64 {
// Kept trivial: the timestamp is diagnostic metadata only; wall-clock
// failure would be a bug rather than something to handle here.
std::time::SystemTime::now()
.duration_since(std::time::UNIX_EPOCH)
.map(|d| d.as_millis() as u64)
.unwrap_or(0)
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn intents_estimate_is_conservative() {
let intent = MrfIntent {
bucket: Arc::from("bucket"),
object: Arc::from("object"),
version_id: Some([0u8; 16]),
kind: MrfKind::DecodeFailure,
enqueued_at_ms: 0,
attempts: 0,
};
assert!(intent.estimated_bytes() >= intent.bucket.len() + intent.object.len());
}
#[tokio::test]
async fn try_send_delivers_and_respects_capacity() {
let mut receiver = init_mrf_channel().expect("first initialization should succeed");
assert!(init_mrf_channel().is_err(), "double initialization must fail");
assert!(try_send_mrf_intent(MrfKind::DecodeFailure, "b", "o", Some(Uuid::nil())));
let intent = receiver.recv().await.expect("intent should arrive");
assert_eq!(intent.kind, MrfKind::DecodeFailure);
assert_eq!(intent.bucket.as_ref(), "b");
// Disable delivery: producers become no-ops.
set_mrf_delivery_enabled(false);
assert!(!try_send_mrf_intent(MrfKind::PartialWrite, "b", "o", None));
set_mrf_delivery_enabled(true);
// Fill the bounded channel past capacity: excess intents are dropped,
// never blocking.
let mut accepted = 0;
for _ in 0..(MRF_CHANNEL_CAPACITY + 64) {
if try_send_mrf_intent(MrfKind::PartialWrite, "b", "o", None) {
accepted += 1;
}
}
assert_eq!(accepted, MRF_CHANNEL_CAPACITY);
}
#[test]
fn try_send_without_channel_is_false() {
// This test may run after the tokio test above in the same process;
// the singleton semantics make a clean "uninitialized" case hard, so
// assert the flag-off behavior only.
set_mrf_delivery_enabled(false);
assert!(!try_send_mrf_intent(MrfKind::MetadataCorruption, "b", "o", None));
set_mrf_delivery_enabled(true);
}
}
+333
View File
@@ -0,0 +1,333 @@
// Copyright 2024 RustFS Team
//
// Licensed under the Apache License, Version 2.0 (the "License");
// you may not use this file except in compliance with the License.
// You may obtain a copy of the License at
//
// http://www.apache.org/licenses/LICENSE-2.0
//
// Unless required by applicable law or agreed to in writing, software
// distributed under the License is distributed on an "AS IS" BASIS,
// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
// See the License for the specific language governing permissions and
// limitations under the License.
use smallvec::SmallVec;
use std::{
sync::{
Arc, OnceLock,
atomic::{AtomicUsize, Ordering},
},
time::{Duration, SystemTime},
};
use tokio::sync::broadcast;
const DEFAULT_TRACE_BUS_CAPACITY: usize = 1024;
const TRACE_ATTR_INLINE_CAPACITY: usize = 8;
static GLOBAL_TRACE_BUS: OnceLock<TraceBus> = OnceLock::new();
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub enum TraceKind {
Heal,
Scanner,
}
impl TraceKind {
pub const fn as_str(self) -> &'static str {
match self {
Self::Heal => "heal",
Self::Scanner => "scanner",
}
}
}
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub enum TraceFunc {
HealTask,
HealBucket,
HealObject,
HealCheckAbandonedParts,
HealErasureSetPage,
ScannerFolder,
ScannerIlmAction,
ScannerHealCandidate,
Dropped,
}
impl TraceFunc {
pub const fn as_str(self) -> &'static str {
match self {
Self::HealTask => "heal.Task",
Self::HealBucket => "heal.Bucket",
Self::HealObject => "heal.Object",
Self::HealCheckAbandonedParts => "heal.CheckAbandonedParts",
Self::HealErasureSetPage => "heal.ErasureSetPage",
Self::ScannerFolder => "scanner.Folder",
Self::ScannerIlmAction => "scanner.IlmAction",
Self::ScannerHealCandidate => "scanner.HealCandidate",
Self::Dropped => "trace.Dropped",
}
}
}
#[derive(Debug, Clone, PartialEq, Eq)]
pub enum TraceVal {
Bool(bool),
U64(u64),
I64(i64),
Str(Arc<str>),
}
impl From<bool> for TraceVal {
fn from(value: bool) -> Self {
Self::Bool(value)
}
}
impl From<u64> for TraceVal {
fn from(value: u64) -> Self {
Self::U64(value)
}
}
impl From<i64> for TraceVal {
fn from(value: i64) -> Self {
Self::I64(value)
}
}
impl From<&str> for TraceVal {
fn from(value: &str) -> Self {
Self::Str(Arc::from(value))
}
}
impl From<String> for TraceVal {
fn from(value: String) -> Self {
Self::Str(Arc::from(value))
}
}
#[derive(Debug, Clone, PartialEq, Eq)]
pub struct TraceAttr {
pub key: &'static str,
pub value: TraceVal,
}
#[derive(Debug, Clone, PartialEq, Eq)]
pub struct TraceEvent {
pub kind: TraceKind,
pub func: TraceFunc,
pub time: SystemTime,
pub bucket: Option<Arc<str>>,
pub object: Option<Arc<str>>,
pub duration: Duration,
pub bytes: u64,
pub attrs: SmallVec<[TraceAttr; TRACE_ATTR_INLINE_CAPACITY]>,
}
impl TraceEvent {
pub fn new(kind: TraceKind, func: TraceFunc) -> Self {
Self {
kind,
func,
time: SystemTime::now(),
bucket: None,
object: None,
duration: Duration::ZERO,
bytes: 0,
attrs: SmallVec::new(),
}
}
pub fn with_bucket(mut self, bucket: impl Into<Arc<str>>) -> Self {
self.bucket = Some(bucket.into());
self
}
pub fn with_object(mut self, object: impl Into<Arc<str>>) -> Self {
self.object = Some(object.into());
self
}
pub fn with_duration(mut self, duration: Duration) -> Self {
self.duration = duration;
self
}
pub fn with_bytes(mut self, bytes: u64) -> Self {
self.bytes = bytes;
self
}
pub fn with_attr(mut self, key: &'static str, value: impl Into<TraceVal>) -> Self {
self.attrs.push(TraceAttr {
key,
value: value.into(),
});
self
}
}
#[derive(Debug)]
pub struct TraceBus {
sender: broadcast::Sender<Arc<TraceEvent>>,
subscriber_count: Arc<AtomicUsize>,
}
impl TraceBus {
pub fn new(capacity: usize) -> Self {
let capacity = capacity.max(1);
let (sender, _receiver) = broadcast::channel(capacity);
Self {
sender,
subscriber_count: Arc::new(AtomicUsize::new(0)),
}
}
pub fn subscriber_count(&self) -> usize {
self.subscriber_count.load(Ordering::Acquire)
}
pub fn subscribe(&self) -> TraceSubscription {
let receiver = self.sender.subscribe();
self.subscriber_count.fetch_add(1, Ordering::AcqRel);
TraceSubscription {
receiver,
subscriber_count: Arc::clone(&self.subscriber_count),
}
}
pub fn emit(&self, build: impl FnOnce() -> TraceEvent) -> bool {
if self.subscriber_count() == 0 {
return false;
}
self.sender.send(Arc::new(build())).is_ok()
}
}
impl Default for TraceBus {
fn default() -> Self {
Self::new(DEFAULT_TRACE_BUS_CAPACITY)
}
}
#[derive(Debug)]
pub struct TraceSubscription {
receiver: broadcast::Receiver<Arc<TraceEvent>>,
subscriber_count: Arc<AtomicUsize>,
}
impl TraceSubscription {
pub async fn recv(&mut self) -> Result<Arc<TraceEvent>, broadcast::error::RecvError> {
self.receiver.recv().await
}
pub fn try_recv(&mut self) -> Result<Arc<TraceEvent>, broadcast::error::TryRecvError> {
self.receiver.try_recv()
}
}
impl Drop for TraceSubscription {
fn drop(&mut self) {
self.subscriber_count.fetch_sub(1, Ordering::AcqRel);
}
}
pub fn global_trace_bus() -> &'static TraceBus {
GLOBAL_TRACE_BUS.get_or_init(TraceBus::default)
}
pub fn subscribe_trace_events() -> TraceSubscription {
global_trace_bus().subscribe()
}
pub fn trace_emit(build: impl FnOnce() -> TraceEvent) -> bool {
global_trace_bus().emit(build)
}
pub fn trace_subscriber_count() -> usize {
global_trace_bus().subscriber_count()
}
#[cfg(test)]
mod tests {
use super::*;
use std::sync::atomic::AtomicUsize;
#[test]
fn trace_emit_skips_builder_without_subscribers() {
let bus = TraceBus::new(4);
let built = AtomicUsize::new(0);
let sent = bus.emit(|| {
built.fetch_add(1, Ordering::Relaxed);
TraceEvent::new(TraceKind::Heal, TraceFunc::HealTask)
});
assert!(!sent);
assert_eq!(built.load(Ordering::Relaxed), 0);
}
#[tokio::test]
async fn trace_subscriber_receives_event() {
let bus = TraceBus::new(4);
let mut subscription = bus.subscribe();
assert!(bus.emit(|| {
TraceEvent::new(TraceKind::Heal, TraceFunc::HealObject)
.with_bucket("bucket")
.with_object("object")
.with_duration(Duration::from_millis(7))
.with_bytes(11)
.with_attr("dry", true)
}));
let event = subscription
.recv()
.await
.expect("subscriber should receive emitted trace event");
assert_eq!(event.kind, TraceKind::Heal);
assert_eq!(event.func, TraceFunc::HealObject);
assert_eq!(event.bucket.as_deref(), Some("bucket"));
assert_eq!(event.object.as_deref(), Some("object"));
assert_eq!(event.duration, Duration::from_millis(7));
assert_eq!(event.bytes, 11);
assert_eq!(
event.attrs.as_slice(),
&[TraceAttr {
key: "dry",
value: TraceVal::Bool(true)
}]
);
}
#[test]
fn trace_subscription_drop_decrements_count() {
let bus = TraceBus::new(4);
let subscription = bus.subscribe();
assert_eq!(bus.subscriber_count(), 1);
drop(subscription);
assert_eq!(bus.subscriber_count(), 0);
}
#[tokio::test]
async fn lagged_subscriber_drops_events_without_blocking_publishers() {
let bus = TraceBus::new(2);
let mut subscription = bus.subscribe();
for index in 0_u64..4 {
assert!(bus.emit(|| { TraceEvent::new(TraceKind::Scanner, TraceFunc::ScannerFolder).with_attr("index", index) }));
}
let err = subscription
.recv()
.await
.expect_err("receiver should observe lag instead of blocking publishers");
assert!(matches!(err, broadcast::error::RecvError::Lagged(_)));
}
}
+2 -3
View File
@@ -14,9 +14,8 @@
//! Shared backpressure policy type. //! Shared backpressure policy type.
//! //!
//! The runtime backpressure implementation (byte-watermark pipes and //! This module only carries the watermark policy; the admission primitive it
//! monitors) lives in `rustfs/src/storage/backpressure.rs`; this module only //! projects into lives in `rustfs-io-core`.
//! carries the watermark policy type that implementation shares.
use rustfs_io_core::BackpressureConfig as CoreBackpressureConfig; use rustfs_io_core::BackpressureConfig as CoreBackpressureConfig;
+37
View File
@@ -177,3 +177,40 @@ pub const DEFAULT_HEAL_MAINLINE_WRITE_UTILIZATION_HIGH_PERCENT: usize = 80;
/// Default foreground pressure recheck delay for heal scheduler, in milliseconds. /// Default foreground pressure recheck delay for heal scheduler, in milliseconds.
pub const DEFAULT_HEAL_MAINLINE_MAX_SLEEP_MS: u64 = 250; pub const DEFAULT_HEAL_MAINLINE_MAX_SLEEP_MS: u64 = 250;
/// Environment variable that toggles the MRF (mission repair feed) intent
/// pipeline: error paths deliver repair intents to the heal runtime, and
/// unconsumed intents are replayed from the durable journal after a restart.
pub const ENV_HEAL_MRF_ENABLE: &str = "RUSTFS_HEAL_MRF_ENABLE";
/// Environment variable for the MRF in-memory queue capacity (intent count).
pub const ENV_HEAL_MRF_QUEUE_SIZE: &str = "RUSTFS_HEAL_MRF_QUEUE_SIZE";
/// Environment variable for the MRF journal byte budget. The journal is
/// compacted once its on-disk size crosses this bound.
pub const ENV_HEAL_MRF_JOURNAL_MAX_BYTES: &str = "RUSTFS_HEAL_MRF_JOURNAL_MAX_BYTES";
/// Environment variable for the MRF journal replay batch size (intents per
/// replay push round).
pub const ENV_HEAL_MRF_REPLAY_BATCH: &str = "RUSTFS_HEAL_MRF_REPLAY_BATCH";
/// Default behavior keeps the MRF intent pipeline enabled.
pub const DEFAULT_HEAL_MRF_ENABLE: bool = true;
/// Default MRF queue capacity (matches MinIO's 100k MRF list ceiling).
pub const DEFAULT_HEAL_MRF_QUEUE_SIZE: usize = 100_000;
/// Default MRF journal byte budget (8 MiB), mirroring the channel payload cap.
pub const DEFAULT_HEAL_MRF_JOURNAL_MAX_BYTES: usize = 8 * 1024 * 1024;
/// Default MRF replay batch size.
pub const DEFAULT_HEAL_MRF_REPLAY_BATCH: usize = 256;
/// Environment variable selecting how admin heal starts behave when the
/// requested path overlaps an already running or queued heal: `merge`
/// (default, keep today's dedup/merge semantics) or `minio_error` (return a
/// typed already-running / overlapping-paths rejection like madmin).
pub const ENV_HEAL_OVERLAP_POLICY: &str = "RUSTFS_HEAL_OVERLAP_POLICY";
/// Default overlap policy: merge duplicate/overlapping requests.
pub const DEFAULT_HEAL_OVERLAP_POLICY: &str = "merge";
+25
View File
@@ -234,6 +234,31 @@ pub const ENV_OBJECT_DISK_WRITE_ABSOLUTE_CAP: &str = "RUSTFS_OBJECT_DISK_WRITE_A
/// Default absolute per-object erasure write cap in seconds (`0` = disabled). /// Default absolute per-object erasure write cap in seconds (`0` = disabled).
pub const DEFAULT_OBJECT_DISK_WRITE_ABSOLUTE_CAP: u64 = 0; pub const DEFAULT_OBJECT_DISK_WRITE_ABSOLUTE_CAP: u64 = 0;
/// Enable foreground PutObject request admission.
///
/// This is an experimental, default-off foreground write backpressure gate for
/// strict commit tail investigations. When disabled, PUTs follow the legacy
/// path and only the existing request counters are updated.
pub const ENV_PUT_FOREGROUND_ADMISSION_ENABLE: &str = "RUSTFS_PUT_FOREGROUND_ADMISSION_ENABLE";
pub const DEFAULT_PUT_FOREGROUND_ADMISSION_ENABLE: bool = false;
/// Maximum foreground PutObject requests admitted concurrently per process.
///
/// The limit is used only when [`ENV_PUT_FOREGROUND_ADMISSION_ENABLE`] is true.
/// A value of `0` disables the gate even when the enable flag is present, so a
/// partially configured rollout cannot reject every PUT.
pub const ENV_PUT_FOREGROUND_ADMISSION_LIMIT: &str = "RUSTFS_PUT_FOREGROUND_ADMISSION_LIMIT";
pub const DEFAULT_PUT_FOREGROUND_ADMISSION_LIMIT: usize = 0;
/// Time in milliseconds a foreground PutObject waits for an admission permit.
///
/// Once this timeout expires the request fails before body ingest/storage
/// mutation with S3 `SlowDown`/503. `0` means fail fast when the limit is full.
pub const ENV_PUT_FOREGROUND_ADMISSION_WAIT_TIMEOUT_MS: &str = "RUSTFS_PUT_FOREGROUND_ADMISSION_WAIT_TIMEOUT_MS";
pub const DEFAULT_PUT_FOREGROUND_ADMISSION_WAIT_TIMEOUT_MS: u64 = 0;
const _: () = assert!(!DEFAULT_PUT_FOREGROUND_ADMISSION_ENABLE);
/// Environment variable for minimum GetObject timeout in seconds. /// Environment variable for minimum GetObject timeout in seconds.
/// ///
/// When dynamic timeout calculation is enabled, this is the minimum timeout /// When dynamic timeout calculation is enabled, this is the minimum timeout
+286
View File
@@ -870,6 +870,157 @@ pub struct DataUsageCacheInfo {
pub snapshot_complete: bool, pub snapshot_complete: bool,
} }
/// Prefix-level usage over a raw entry map — the shared core behind
/// [`DataUsageCache::prefix_usage`], usable by any cache-shaped reader (the
/// scanner's writer-side cache has the same map type).
///
/// Cache keys are cleaned literal paths (`bucket/pre/fix`), so sub-prefix
/// names come straight off the child keys — no reverse mapping exists or is
/// needed. A compacted prefix carries its aggregate but no children, which
/// the `compacted` flag reports so callers can say why the breakdown is
/// empty. `truncated` is set when the breakdown exceeded `max_entries` and
/// was cut (largest first).
pub fn prefix_usage_in_cache(
cache: &HashMap<String, DataUsageEntry>,
bucket: &str,
prefix: &str,
max_entries: usize,
) -> Option<PrefixUsageQuery> {
let prefix = prefix.trim_matches('/');
let root = if prefix.is_empty() {
bucket.to_string()
} else {
format!("{bucket}/{prefix}")
};
let entry = cache.get(&hash_path(&root).key())?.clone();
let usage = PrefixUsageSummary::from_entry(&flatten_entry(cache, &entry, 0)?);
let child_prefix = format!("{root}/");
let mut sub_prefixes: Vec<PrefixUsageEntry> = entry
.children
.iter()
.filter_map(|child_key| {
let child = cache.get(child_key)?;
let child_flat = flatten_entry(cache, child, 1)?;
// Child keys are literal `bucket/pre/name` paths; a trailing
// slash marks a directory object and is display-only here.
let name = child_key
.strip_prefix(child_prefix.as_str())
.unwrap_or(child_key.as_str())
.trim_end_matches('/')
.to_string();
Some(PrefixUsageEntry {
prefix: name,
usage: PrefixUsageSummary::from_entry(&child_flat),
})
})
.collect();
sub_prefixes.sort_by(|left, right| {
right
.usage
.size
.cmp(&left.usage.size)
.then_with(|| left.prefix.cmp(&right.prefix))
});
let truncated = sub_prefixes.len() > max_entries;
sub_prefixes.truncate(max_entries);
Some(PrefixUsageQuery {
usage,
compacted: entry.compacted,
truncated,
sub_prefixes,
})
}
/// Maximum subtree depth [`flatten_entry`] will walk before declaring the
/// cache corrupt — the same bound the scanner's checked flatten uses.
const PREFIX_USAGE_MAX_DEPTH: usize = 1024;
/// Flatten one entry's subtree into an aggregate: the free-function twin of
/// [`DataUsageCache::flatten`], carrying the scanner checked-flatten
/// hardening so a corrupt cache (cycles, over-deep trees, overflowing
/// counters) yields `None` instead of unbounded recursion or wrapped totals.
fn flatten_entry(cache: &HashMap<String, DataUsageEntry>, root: &DataUsageEntry, depth: usize) -> Option<DataUsageEntry> {
if depth > PREFIX_USAGE_MAX_DEPTH {
return None;
}
let mut flattened = DataUsageEntry::default();
if !flattened.checked_merge(root) {
return None;
}
flattened.compacted = root.compacted;
// The root itself is not pre-seeded: it is merged above, and a corrupt
// child edge pointing back at the root's own key is still terminated by
// the visited set on first encounter.
let mut visited: HashSet<&str> = HashSet::new();
let mut pending: Vec<(&String, usize)> = root.children.iter().map(|child| (child, depth + 1)).collect();
while let Some((key, child_depth)) = pending.pop() {
if child_depth > PREFIX_USAGE_MAX_DEPTH || !visited.insert(key.as_str()) {
return None;
}
let entry = cache.get(key)?;
if !flattened.checked_merge(entry) {
return None;
}
pending.extend(entry.children.iter().map(|child| (child, child_depth + 1)));
}
flattened.children.clear();
Some(flattened)
}
/// Flattened counters of one prefix subtree, as returned by
/// [`DataUsageCache::prefix_usage`].
#[derive(Clone, Copy, Debug, Default, PartialEq, Eq, serde::Serialize)]
#[serde(rename_all = "camelCase")]
pub struct PrefixUsageSummary {
pub size: u64,
pub objects: u64,
pub versions: u64,
pub delete_markers: u64,
}
impl PrefixUsageSummary {
fn from_entry(entry: &DataUsageEntry) -> Self {
Self {
size: entry.size as u64,
objects: entry.objects as u64,
versions: entry.versions as u64,
delete_markers: entry.delete_markers as u64,
}
}
/// Add another set's counters into this one (entries are partitioned by
/// set, so per-set results sum).
pub fn merge(&mut self, other: &Self) {
self.size = self.size.saturating_add(other.size);
self.objects = self.objects.saturating_add(other.objects);
self.versions = self.versions.saturating_add(other.versions);
self.delete_markers = self.delete_markers.saturating_add(other.delete_markers);
}
}
/// One first-level sub-prefix row of a [`PrefixUsageQuery`].
#[derive(Clone, Debug, PartialEq, Eq, serde::Serialize)]
pub struct PrefixUsageEntry {
pub prefix: String,
pub usage: PrefixUsageSummary,
}
/// Result of [`DataUsageCache::prefix_usage`].
#[derive(Clone, Debug, Default, PartialEq, Eq, serde::Serialize)]
#[serde(rename_all = "camelCase")]
pub struct PrefixUsageQuery {
pub usage: PrefixUsageSummary,
/// The prefix entry was compacted by the scanner: its aggregate is valid
/// but no sub-prefix breakdown exists on disk.
pub compacted: bool,
/// The breakdown had more entries than `max_entries`; the largest remain.
pub truncated: bool,
pub sub_prefixes: Vec<PrefixUsageEntry>,
}
/// Read-only projection of a scanner-written `.usage-cache.bin` file. /// Read-only projection of a scanner-written `.usage-cache.bin` file.
/// ///
/// The scanner-side `DataUsageCache` (`crates/scanner/src/data_usage_define.rs`) /// The scanner-side `DataUsageCache` (`crates/scanner/src/data_usage_define.rs`)
@@ -997,6 +1148,21 @@ impl DataUsageCache {
} }
} }
/// Prefix-level usage for one bucket subtree, plus the one-level
/// breakdown below it (rustfs/backlog#1872, MinIO
/// `loadPrefixUsageFromBackend` parity and beyond: arbitrary prefixes and
/// full counters instead of first-level sizes only).
///
/// Cache keys are cleaned literal paths (`bucket/pre/fix`), so sub-prefix
/// names come straight off the child keys — no reverse mapping exists or
/// is needed. A compacted prefix carries its aggregate but no children,
/// which the `compacted` flag reports so callers can say why the
/// breakdown is empty. `truncated` is set when the breakdown exceeded
/// `max_entries` and was cut (largest first).
pub fn prefix_usage(&self, bucket: &str, prefix: &str, max_entries: usize) -> Option<PrefixUsageQuery> {
prefix_usage_in_cache(&self.cache, bucket, prefix, max_entries)
}
pub fn force_compact(&mut self, limit: usize) { pub fn force_compact(&mut self, limit: usize) {
if self.cache.len() < limit { if self.cache.len() < limit {
return; return;
@@ -1898,6 +2064,126 @@ mod tests {
); );
} }
/// Build a cache shaped like `bucket/{a,b/{c,d}},bucket/loose` with
/// distinct counters so aggregation is observable.
fn prefix_usage_fixture_cache() -> DataUsageCache {
let mut cache = DataUsageCache::default();
let mut insert = |path: &str, parent: &str, size: usize, objects: usize, versions: usize, delete_markers: usize| {
cache.replace(
path,
parent,
DataUsageEntry {
size,
objects,
versions,
delete_markers,
..Default::default()
},
);
};
insert("bucket", "", 0, 0, 0, 0);
insert("bucket/a", "bucket", 100, 1, 1, 0);
insert("bucket/b", "bucket", 0, 0, 0, 0);
insert("bucket/b/c", "bucket/b", 200, 2, 2, 1);
insert("bucket/b/d", "bucket/b", 40, 1, 3, 0);
insert("bucket/loose", "bucket", 10, 1, 1, 1);
cache
}
#[test]
fn prefix_usage_aggregates_bucket_root_and_one_level_below() {
let cache = prefix_usage_fixture_cache();
let root = cache
.prefix_usage("bucket", "", 100)
.expect("root query must find the bucket entry");
assert_eq!(root.usage.size, 350, "root aggregate flattens the whole subtree");
assert_eq!(root.usage.objects, 5);
assert_eq!(root.usage.versions, 7);
assert_eq!(root.usage.delete_markers, 2);
assert!(!root.compacted);
assert!(!root.truncated);
// Breakdown is one level: b (240) before a (100) before loose (10),
// each flattened to its own subtree total.
let names: Vec<(&str, u64)> = root
.sub_prefixes
.iter()
.map(|entry| (entry.prefix.as_str(), entry.usage.size))
.collect();
assert_eq!(names, vec![("b", 240), ("a", 100), ("loose", 10)]);
}
#[test]
fn prefix_usage_drills_into_arbitrary_prefixes() {
let cache = prefix_usage_fixture_cache();
let b = cache.prefix_usage("bucket", "b", 100).expect("nested prefix must resolve");
assert_eq!(b.usage.size, 240);
assert_eq!(b.usage.versions, 5);
let names: Vec<&str> = b.sub_prefixes.iter().map(|entry| entry.prefix.as_str()).collect();
assert_eq!(names, vec!["c", "d"]);
// Prefix slashes are normalized away.
let slashed = cache.prefix_usage("bucket", "/b/", 100).expect("slash-insensitive lookup");
assert_eq!(slashed.usage.size, 240);
assert!(cache.prefix_usage("bucket", "absent", 100).is_none(), "unknown prefix must be a miss");
assert!(cache.prefix_usage("other", "", 100).is_none(), "unknown bucket must be a miss");
}
#[test]
fn prefix_usage_reports_and_respects_truncation() {
let cache = prefix_usage_fixture_cache();
let capped = cache.prefix_usage("bucket", "", 2).expect("root query");
assert!(capped.truncated, "three children capped to two must flag truncation");
let names: Vec<&str> = capped.sub_prefixes.iter().map(|entry| entry.prefix.as_str()).collect();
assert_eq!(names, vec!["b", "a"], "largest prefixes survive the cut");
}
#[test]
fn prefix_usage_marks_compacted_entries() {
let mut cache = DataUsageCache::default();
cache.replace(
"bucket",
"",
DataUsageEntry {
size: 999,
objects: 9,
compacted: true,
..Default::default()
},
);
let compacted = cache.prefix_usage("bucket", "", 100).expect("compacted root resolves");
assert!(compacted.compacted, "compaction must be visible to callers");
assert_eq!(compacted.usage.size, 999);
assert!(compacted.sub_prefixes.is_empty(), "a compacted entry carries no children");
}
#[test]
fn prefix_usage_rejects_cyclic_and_dangling_caches() {
// A self-referencing child (corrupt cache) must yield a miss for the
// whole query, not unbounded recursion.
let mut cache = prefix_usage_fixture_cache();
if let Some(entry) = cache.cache.get_mut("bucket/b") {
entry.children.insert("bucket/b".to_string());
}
assert!(cache.prefix_usage("bucket", "b", 100).is_none(), "a cyclic subtree must be rejected");
// The unaffected sibling still answers.
assert!(cache.prefix_usage("bucket", "a", 100).is_some());
// A child key with no entry (dangling link) is rejected rather than
// silently dropped: half a tree would under-report usage.
let mut dangling = prefix_usage_fixture_cache();
if let Some(entry) = dangling.cache.get_mut("bucket/b") {
entry.children.insert("bucket/b/ghost".to_string());
}
assert!(
dangling.prefix_usage("bucket", "b", 100).is_none(),
"a dangling child link must be rejected"
);
}
#[test] #[test]
fn hash_path_uses_portable_slash_semantics() { fn hash_path_uses_portable_slash_semantics() {
for (input, expected) in [ for (input, expected) in [
+79 -4
View File
@@ -32,6 +32,7 @@ use rustfs_signer::sign_v4;
use s3s::Body; use s3s::Body;
use std::ffi::OsStr; use std::ffi::OsStr;
use std::fs as stdfs; use std::fs as stdfs;
use std::io::ErrorKind;
use std::path::{Path, PathBuf}; use std::path::{Path, PathBuf};
use std::process::{Child, Command, Stdio}; use std::process::{Child, Command, Stdio};
use std::sync::Once; use std::sync::Once;
@@ -51,6 +52,11 @@ pub(crate) const FAST_DATA_USAGE_SCANNER_ENV: &[(&str, &str)] =
&[("RUSTFS_SCANNER_CYCLE", "1"), ("RUSTFS_SCANNER_START_DELAY_SECS", "0")]; &[("RUSTFS_SCANNER_CYCLE", "1"), ("RUSTFS_SCANNER_START_DELAY_SECS", "0")];
pub const TEST_BUCKET: &str = "e2e-test-bucket"; pub const TEST_BUCKET: &str = "e2e-test-bucket";
const RUSTFS_FULL_FEATURE: &str = "full"; const RUSTFS_FULL_FEATURE: &str = "full";
const TEST_PORT_MIN: u16 = 20_000;
const TEST_PORT_RANGE: u16 = 40_000;
const TEST_PORT_COUNTER_PATH: &str = "/tmp/rustfs_e2e_next_port";
const TEST_PORT_LOCK_DIR: &str = "/tmp/rustfs_e2e_port_allocator.lock";
const TEST_PORT_LOCK_STALE_AFTER: Duration = Duration::from_secs(30);
fn capture_log_path(log_dir: &Path, temp_dir: &str) -> Option<PathBuf> { fn capture_log_path(log_dir: &Path, temp_dir: &str) -> Option<PathBuf> {
let temp_name = Path::new(temp_dir).file_name()?.to_string_lossy(); let temp_name = Path::new(temp_dir).file_name()?.to_string_lossy();
@@ -67,6 +73,64 @@ fn configured_capture_log_path(temp_dir: &str) -> Option<String> {
capture_log_path(Path::new(&log_dir), temp_dir).map(|path| path.to_string_lossy().into_owned()) capture_log_path(Path::new(&log_dir), temp_dir).map(|path| path.to_string_lossy().into_owned())
} }
struct PortAllocatorGuard;
impl PortAllocatorGuard {
async fn acquire() -> Result<Self, Box<dyn std::error::Error + Send + Sync>> {
loop {
match stdfs::create_dir(TEST_PORT_LOCK_DIR) {
Ok(()) => return Ok(Self),
Err(err) if err.kind() == ErrorKind::AlreadyExists => {
remove_stale_port_allocator_lock();
sleep(Duration::from_millis(10)).await;
}
Err(err) => return Err(err.into()),
}
}
}
}
impl Drop for PortAllocatorGuard {
fn drop(&mut self) {
let _ = stdfs::remove_dir(TEST_PORT_LOCK_DIR);
}
}
fn advance_test_port(port: u16) -> u16 {
let offset = (port - TEST_PORT_MIN + 1) % TEST_PORT_RANGE;
TEST_PORT_MIN + offset
}
fn seeded_test_port() -> u16 {
let offset = (Uuid::new_v4().as_u128() % u128::from(TEST_PORT_RANGE)) as u16;
TEST_PORT_MIN + offset
}
fn read_next_test_port() -> u16 {
stdfs::read_to_string(TEST_PORT_COUNTER_PATH)
.ok()
.and_then(|value| value.trim().parse::<u16>().ok())
.filter(|port| (TEST_PORT_MIN..TEST_PORT_MIN + TEST_PORT_RANGE).contains(port))
.unwrap_or_else(seeded_test_port)
}
fn remove_stale_port_allocator_lock() {
let Ok(metadata) = stdfs::metadata(TEST_PORT_LOCK_DIR) else {
return;
};
let Ok(modified) = metadata.modified() else {
return;
};
if modified.elapsed().is_ok_and(|elapsed| elapsed > TEST_PORT_LOCK_STALE_AFTER) {
let _ = stdfs::remove_dir(TEST_PORT_LOCK_DIR);
}
}
fn write_next_test_port(port: u16) -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
stdfs::write(TEST_PORT_COUNTER_PATH, port.to_string())?;
Ok(())
}
pub(crate) fn capture_command_logs( pub(crate) fn capture_command_logs(
command: &mut Command, command: &mut Command,
log_path: Option<&str>, log_path: Option<&str>,
@@ -508,10 +572,21 @@ impl RustFSTestEnvironment {
/// Find an available port for the test /// Find an available port for the test
pub async fn find_available_port() -> Result<u16, Box<dyn std::error::Error + Send + Sync>> { pub async fn find_available_port() -> Result<u16, Box<dyn std::error::Error + Send + Sync>> {
use std::net::TcpListener; use std::net::TcpListener;
let listener = TcpListener::bind("127.0.0.1:0")?; let _guard = PortAllocatorGuard::acquire().await?;
let port = listener.local_addr()?.port(); let mut next_port = read_next_test_port();
drop(listener);
Ok(port) for _ in 0..TEST_PORT_RANGE {
let port = next_port;
next_port = advance_test_port(next_port);
write_next_test_port(next_port)?;
if let Ok(listener) = TcpListener::bind(("127.0.0.1", port)) {
drop(listener);
return Ok(port);
}
}
Err("no available E2E test port found".into())
} }
/// Kill any existing RustFS processes /// Kill any existing RustFS processes
@@ -30,7 +30,6 @@ use md5::{Digest as Md5Digest, Md5};
use rustfs_signer::constants::UNSIGNED_PAYLOAD; use rustfs_signer::constants::UNSIGNED_PAYLOAD;
use rustfs_signer::sign_v4; use rustfs_signer::sign_v4;
use s3s::Body; use s3s::Body;
use serial_test::serial;
use std::collections::HashMap; use std::collections::HashMap;
use std::error::Error; use std::error::Error;
use std::io::Cursor; use std::io::Cursor;
@@ -356,7 +355,6 @@ async fn run_post_object_policy_case(
/// smuggles one extra field the policy never declared, and the upload must be /// smuggles one extra field the policy never declared, and the upload must be
/// rejected with 403 AccessDenied naming the offending field. /// rejected with 403 AccessDenied naming the offending field.
#[tokio::test] #[tokio::test]
#[serial]
async fn test_anonymous_post_object_rejects_fields_missing_from_policy_conditions() async fn test_anonymous_post_object_rejects_fields_missing_from_policy_conditions()
-> Result<(), Box<dyn std::error::Error + Send + Sync>> { -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -484,7 +482,6 @@ async fn test_anonymous_post_object_rejects_fields_missing_from_policy_condition
/// sends a different one, and the upload must be rejected with 400 /// sends a different one, and the upload must be rejected with 400
/// InvalidPolicyDocument naming the field. /// InvalidPolicyDocument naming the field.
#[tokio::test] #[tokio::test]
#[serial]
async fn test_anonymous_post_object_rejects_exact_condition_policy_mismatches() async fn test_anonymous_post_object_rejects_exact_condition_policy_mismatches()
-> Result<(), Box<dyn std::error::Error + Send + Sync>> { -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -689,7 +686,6 @@ async fn test_anonymous_post_object_rejects_exact_condition_policy_mismatches()
/// one of them with a different value, and the upload must be rejected with /// one of them with a different value, and the upload must be rejected with
/// 400 InvalidPolicyDocument naming the mismatched field. /// 400 InvalidPolicyDocument naming the mismatched field.
#[tokio::test] #[tokio::test]
#[serial]
async fn test_anonymous_post_object_rejects_object_lock_policy_mismatches() -> Result<(), Box<dyn std::error::Error + Send + Sync>> async fn test_anonymous_post_object_rejects_object_lock_policy_mismatches() -> Result<(), Box<dyn std::error::Error + Send + Sync>>
{ {
init_logging(); init_logging();
@@ -757,7 +753,6 @@ async fn test_anonymous_post_object_rejects_object_lock_policy_mismatches() -> R
/// exact values, the form sends a different parameter value, and the upload /// exact values, the form sends a different parameter value, and the upload
/// must be rejected with 400 InvalidPolicyDocument naming the parameter. /// must be rejected with 400 InvalidPolicyDocument naming the parameter.
#[tokio::test] #[tokio::test]
#[serial]
async fn test_anonymous_post_object_rejects_sse_kms_policy_mismatches() -> Result<(), Box<dyn std::error::Error + Send + Sync>> { async fn test_anonymous_post_object_rejects_sse_kms_policy_mismatches() -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -839,7 +834,6 @@ async fn test_anonymous_post_object_rejects_sse_kms_policy_mismatches() -> Resul
/// NotImplemented (SSE-KMS POST uploads are not implemented), not with a /// NotImplemented (SSE-KMS POST uploads are not implemented), not with a
/// policy error. /// policy error.
#[tokio::test] #[tokio::test]
#[serial]
async fn test_anonymous_post_object_rejects_sse_kms_params_outside_policy_conditions() async fn test_anonymous_post_object_rejects_sse_kms_params_outside_policy_conditions()
-> Result<(), Box<dyn std::error::Error + Send + Sync>> { -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -894,7 +888,6 @@ async fn test_anonymous_post_object_rejects_sse_kms_params_outside_policy_condit
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_anonymous_multipart_control_apis_require_auth() -> Result<(), Box<dyn std::error::Error + Send + Sync>> { async fn test_anonymous_multipart_control_apis_require_auth() -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -968,7 +961,6 @@ async fn test_anonymous_multipart_control_apis_require_auth() -> Result<(), Box<
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_anonymous_post_object_requires_auth() -> Result<(), Box<dyn std::error::Error + Send + Sync>> { async fn test_anonymous_post_object_requires_auth() -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -1002,7 +994,6 @@ async fn test_anonymous_post_object_requires_auth() -> Result<(), Box<dyn std::e
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_anonymous_post_object_honors_success_action_status() -> Result<(), Box<dyn std::error::Error + Send + Sync>> { async fn test_anonymous_post_object_honors_success_action_status() -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -1066,7 +1057,6 @@ async fn test_anonymous_post_object_honors_success_action_status() -> Result<(),
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_anonymous_post_object_honors_success_action_redirect() -> Result<(), Box<dyn std::error::Error + Send + Sync>> { async fn test_anonymous_post_object_honors_success_action_redirect() -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -1139,7 +1129,6 @@ async fn test_anonymous_post_object_honors_success_action_redirect() -> Result<(
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_anonymous_post_object_defaults_to_no_content() -> Result<(), Box<dyn std::error::Error + Send + Sync>> { async fn test_anonymous_post_object_defaults_to_no_content() -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -1185,7 +1174,6 @@ async fn test_anonymous_post_object_defaults_to_no_content() -> Result<(), Box<d
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_anonymous_post_object_rejects_sse_kms() -> Result<(), Box<dyn std::error::Error + Send + Sync>> { async fn test_anonymous_post_object_rejects_sse_kms() -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -1232,7 +1220,6 @@ async fn test_anonymous_post_object_rejects_sse_kms() -> Result<(), Box<dyn std:
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_anonymous_post_object_accepts_sse_s3() -> Result<(), Box<dyn std::error::Error + Send + Sync>> { async fn test_anonymous_post_object_accepts_sse_s3() -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -1290,7 +1277,6 @@ async fn test_anonymous_post_object_accepts_sse_s3() -> Result<(), Box<dyn std::
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_anonymous_post_object_uses_bucket_default_sse_s3() -> Result<(), Box<dyn std::error::Error + Send + Sync>> { async fn test_anonymous_post_object_uses_bucket_default_sse_s3() -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -1363,7 +1349,6 @@ async fn test_anonymous_post_object_uses_bucket_default_sse_s3() -> Result<(), B
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_anonymous_post_object_uses_bucket_default_sse_kms() -> Result<(), Box<dyn std::error::Error + Send + Sync>> { async fn test_anonymous_post_object_uses_bucket_default_sse_kms() -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -1437,7 +1422,6 @@ async fn test_anonymous_post_object_uses_bucket_default_sse_kms() -> Result<(),
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_anonymous_post_object_rejects_sse_s3_policy_mismatch() -> Result<(), Box<dyn std::error::Error + Send + Sync>> { async fn test_anonymous_post_object_rejects_sse_s3_policy_mismatch() -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -1488,7 +1472,6 @@ async fn test_anonymous_post_object_rejects_sse_s3_policy_mismatch() -> Result<(
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_anonymous_post_object_accepts_sse_s3_missing_from_policy_conditions() async fn test_anonymous_post_object_accepts_sse_s3_missing_from_policy_conditions()
-> Result<(), Box<dyn std::error::Error + Send + Sync>> { -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -1552,7 +1535,6 @@ async fn test_anonymous_post_object_accepts_sse_s3_missing_from_policy_condition
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_anonymous_post_object_accepts_storage_class_exact_policy_match() async fn test_anonymous_post_object_accepts_storage_class_exact_policy_match()
-> Result<(), Box<dyn std::error::Error + Send + Sync>> { -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -1606,7 +1588,6 @@ async fn test_anonymous_post_object_accepts_storage_class_exact_policy_match()
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_anonymous_post_object_rejects_storage_class_missing_from_policy_conditions() async fn test_anonymous_post_object_rejects_storage_class_missing_from_policy_conditions()
-> Result<(), Box<dyn std::error::Error + Send + Sync>> { -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -1657,7 +1638,6 @@ async fn test_anonymous_post_object_rejects_storage_class_missing_from_policy_co
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_anonymous_post_object_rejects_invalid_storage_class_value() -> Result<(), Box<dyn std::error::Error + Send + Sync>> async fn test_anonymous_post_object_rejects_invalid_storage_class_value() -> Result<(), Box<dyn std::error::Error + Send + Sync>>
{ {
init_logging(); init_logging();
@@ -1709,7 +1689,6 @@ async fn test_anonymous_post_object_rejects_invalid_storage_class_value() -> Res
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_anonymous_post_object_rejects_checksum_algorithm_missing_from_policy_conditions() async fn test_anonymous_post_object_rejects_checksum_algorithm_missing_from_policy_conditions()
-> Result<(), Box<dyn std::error::Error + Send + Sync>> { -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -1765,7 +1744,6 @@ async fn test_anonymous_post_object_rejects_checksum_algorithm_missing_from_poli
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_anonymous_post_object_rejects_checksum_algorithm_policy_mismatch() async fn test_anonymous_post_object_rejects_checksum_algorithm_policy_mismatch()
-> Result<(), Box<dyn std::error::Error + Send + Sync>> { -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -1822,7 +1800,6 @@ async fn test_anonymous_post_object_rejects_checksum_algorithm_policy_mismatch()
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_anonymous_post_object_rejects_checksum_auxiliary_fields_missing_from_policy_conditions() async fn test_anonymous_post_object_rejects_checksum_auxiliary_fields_missing_from_policy_conditions()
-> Result<(), Box<dyn std::error::Error + Send + Sync>> { -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -1886,7 +1863,6 @@ async fn test_anonymous_post_object_rejects_checksum_auxiliary_fields_missing_fr
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_anonymous_post_object_allows_sse_c_fields_outside_policy_conditions() async fn test_anonymous_post_object_allows_sse_c_fields_outside_policy_conditions()
-> Result<(), Box<dyn std::error::Error + Send + Sync>> { -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -1963,7 +1939,6 @@ async fn test_anonymous_post_object_allows_sse_c_fields_outside_policy_condition
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_anonymous_post_object_rejects_sse_c_exact_policy_mismatch() -> Result<(), Box<dyn std::error::Error + Send + Sync>> async fn test_anonymous_post_object_rejects_sse_c_exact_policy_mismatch() -> Result<(), Box<dyn std::error::Error + Send + Sync>>
{ {
init_logging(); init_logging();
@@ -2022,7 +1997,6 @@ async fn test_anonymous_post_object_rejects_sse_c_exact_policy_mismatch() -> Res
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_anonymous_post_object_rejects_duplicate_key_form_values() -> Result<(), Box<dyn std::error::Error + Send + Sync>> { async fn test_anonymous_post_object_rejects_duplicate_key_form_values() -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -2072,7 +2046,6 @@ async fn test_anonymous_post_object_rejects_duplicate_key_form_values() -> Resul
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_anonymous_post_object_rejects_invalid_success_action_status() -> Result<(), Box<dyn std::error::Error + Send + Sync>> async fn test_anonymous_post_object_rejects_invalid_success_action_status() -> Result<(), Box<dyn std::error::Error + Send + Sync>>
{ {
init_logging(); init_logging();
@@ -2120,7 +2093,6 @@ async fn test_anonymous_post_object_rejects_invalid_success_action_status() -> R
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_anonymous_post_object_rejects_invalid_success_action_redirect() async fn test_anonymous_post_object_rejects_invalid_success_action_redirect()
-> Result<(), Box<dyn std::error::Error + Send + Sync>> { -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -2168,7 +2140,6 @@ async fn test_anonymous_post_object_rejects_invalid_success_action_redirect()
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_anonymous_post_object_rejects_form_fields_missing_from_policy_conditions() async fn test_anonymous_post_object_rejects_form_fields_missing_from_policy_conditions()
-> Result<(), Box<dyn std::error::Error + Send + Sync>> { -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -2223,7 +2194,6 @@ async fn test_anonymous_post_object_rejects_form_fields_missing_from_policy_cond
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_anonymous_post_object_accepts_form_fields_covered_by_policy_conditions() async fn test_anonymous_post_object_accepts_form_fields_covered_by_policy_conditions()
-> Result<(), Box<dyn std::error::Error + Send + Sync>> { -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -2280,7 +2250,6 @@ async fn test_anonymous_post_object_accepts_form_fields_covered_by_policy_condit
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_anonymous_post_object_rejects_starts_with_policy_mismatch() -> Result<(), Box<dyn std::error::Error + Send + Sync>> async fn test_anonymous_post_object_rejects_starts_with_policy_mismatch() -> Result<(), Box<dyn std::error::Error + Send + Sync>>
{ {
init_logging(); init_logging();
@@ -2335,7 +2304,6 @@ async fn test_anonymous_post_object_rejects_starts_with_policy_mismatch() -> Res
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_anonymous_post_object_rejects_content_length_range_violation() async fn test_anonymous_post_object_rejects_content_length_range_violation()
-> Result<(), Box<dyn std::error::Error + Send + Sync>> { -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -2388,7 +2356,6 @@ async fn test_anonymous_post_object_rejects_content_length_range_violation()
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_anonymous_post_object_accepts_success_action_status_exact_policy_match() async fn test_anonymous_post_object_accepts_success_action_status_exact_policy_match()
-> Result<(), Box<dyn std::error::Error + Send + Sync>> { -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -2445,7 +2412,6 @@ async fn test_anonymous_post_object_accepts_success_action_status_exact_policy_m
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_anonymous_post_object_rejects_success_action_redirect_policy_mismatch() async fn test_anonymous_post_object_rejects_success_action_redirect_policy_mismatch()
-> Result<(), Box<dyn std::error::Error + Send + Sync>> { -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -2502,7 +2468,6 @@ async fn test_anonymous_post_object_rejects_success_action_redirect_policy_misma
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_anonymous_post_object_accepts_success_action_redirect_exact_policy_match() async fn test_anonymous_post_object_accepts_success_action_redirect_exact_policy_match()
-> Result<(), Box<dyn std::error::Error + Send + Sync>> { -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -2568,7 +2533,6 @@ async fn test_anonymous_post_object_accepts_success_action_redirect_exact_policy
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_anonymous_post_object_rejects_success_action_redirect_missing_from_policy_conditions() async fn test_anonymous_post_object_rejects_success_action_redirect_missing_from_policy_conditions()
-> Result<(), Box<dyn std::error::Error + Send + Sync>> { -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -2621,7 +2585,6 @@ async fn test_anonymous_post_object_rejects_success_action_redirect_missing_from
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_anonymous_post_object_accepts_metadata_field_covered_by_starts_with() async fn test_anonymous_post_object_accepts_metadata_field_covered_by_starts_with()
-> Result<(), Box<dyn std::error::Error + Send + Sync>> { -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -2676,7 +2639,6 @@ async fn test_anonymous_post_object_accepts_metadata_field_covered_by_starts_wit
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_anonymous_post_object_accepts_content_type_field_exact_policy_match() async fn test_anonymous_post_object_accepts_content_type_field_exact_policy_match()
-> Result<(), Box<dyn std::error::Error + Send + Sync>> { -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -2734,7 +2696,6 @@ async fn test_anonymous_post_object_accepts_content_type_field_exact_policy_matc
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_anonymous_post_object_accepts_content_type_field_covered_by_starts_with() async fn test_anonymous_post_object_accepts_content_type_field_covered_by_starts_with()
-> Result<(), Box<dyn std::error::Error + Send + Sync>> { -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -2792,7 +2753,6 @@ async fn test_anonymous_post_object_accepts_content_type_field_covered_by_starts
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_anonymous_post_object_accepts_content_disposition_field_exact_policy_match() async fn test_anonymous_post_object_accepts_content_disposition_field_exact_policy_match()
-> Result<(), Box<dyn std::error::Error + Send + Sync>> { -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -2850,7 +2810,6 @@ async fn test_anonymous_post_object_accepts_content_disposition_field_exact_poli
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_anonymous_post_object_accepts_cache_control_field_exact_policy_match() async fn test_anonymous_post_object_accepts_cache_control_field_exact_policy_match()
-> Result<(), Box<dyn std::error::Error + Send + Sync>> { -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -2908,7 +2867,6 @@ async fn test_anonymous_post_object_accepts_cache_control_field_exact_policy_mat
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_anonymous_post_object_accepts_content_language_field_exact_policy_match() async fn test_anonymous_post_object_accepts_content_language_field_exact_policy_match()
-> Result<(), Box<dyn std::error::Error + Send + Sync>> { -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -2966,7 +2924,6 @@ async fn test_anonymous_post_object_accepts_content_language_field_exact_policy_
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_anonymous_post_object_accepts_content_encoding_field_exact_policy_match() async fn test_anonymous_post_object_accepts_content_encoding_field_exact_policy_match()
-> Result<(), Box<dyn std::error::Error + Send + Sync>> { -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -3024,7 +2981,6 @@ async fn test_anonymous_post_object_accepts_content_encoding_field_exact_policy_
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_anonymous_post_object_accepts_website_redirect_location_exact_policy_match() async fn test_anonymous_post_object_accepts_website_redirect_location_exact_policy_match()
-> Result<(), Box<dyn std::error::Error + Send + Sync>> { -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -3082,7 +3038,6 @@ async fn test_anonymous_post_object_accepts_website_redirect_location_exact_poli
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_anonymous_post_object_accepts_expires_field_exact_policy_match() async fn test_anonymous_post_object_accepts_expires_field_exact_policy_match()
-> Result<(), Box<dyn std::error::Error + Send + Sync>> { -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -3140,7 +3095,6 @@ async fn test_anonymous_post_object_accepts_expires_field_exact_policy_match()
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_anonymous_post_object_rejects_object_lock_retention_without_permission() async fn test_anonymous_post_object_rejects_object_lock_retention_without_permission()
-> Result<(), Box<dyn std::error::Error + Send + Sync>> { -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -3196,7 +3150,6 @@ async fn test_anonymous_post_object_rejects_object_lock_retention_without_permis
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_anonymous_post_object_rejects_object_lock_retention_missing_from_policy_conditions() async fn test_anonymous_post_object_rejects_object_lock_retention_missing_from_policy_conditions()
-> Result<(), Box<dyn std::error::Error + Send + Sync>> { -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -3256,7 +3209,6 @@ async fn test_anonymous_post_object_rejects_object_lock_retention_missing_from_p
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_anonymous_post_object_rejects_object_lock_legal_hold_without_permission() async fn test_anonymous_post_object_rejects_object_lock_legal_hold_without_permission()
-> Result<(), Box<dyn std::error::Error + Send + Sync>> { -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -3309,7 +3261,6 @@ async fn test_anonymous_post_object_rejects_object_lock_legal_hold_without_permi
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_anonymous_post_object_rejects_object_lock_legal_hold_policy_mismatch() async fn test_anonymous_post_object_rejects_object_lock_legal_hold_policy_mismatch()
-> Result<(), Box<dyn std::error::Error + Send + Sync>> { -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -3368,7 +3319,6 @@ async fn test_anonymous_post_object_rejects_object_lock_legal_hold_policy_mismat
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_anonymous_post_object_rejects_object_lock_legal_hold_missing_from_policy_conditions() async fn test_anonymous_post_object_rejects_object_lock_legal_hold_missing_from_policy_conditions()
-> Result<(), Box<dyn std::error::Error + Send + Sync>> { -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -3426,7 +3376,6 @@ async fn test_anonymous_post_object_rejects_object_lock_legal_hold_missing_from_
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_anonymous_post_object_accepts_tagging_field_exact_policy_match() async fn test_anonymous_post_object_accepts_tagging_field_exact_policy_match()
-> Result<(), Box<dyn std::error::Error + Send + Sync>> { -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -3492,7 +3441,6 @@ async fn test_anonymous_post_object_accepts_tagging_field_exact_policy_match()
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_anonymous_post_object_accepts_metadata_field_exact_policy_match() async fn test_anonymous_post_object_accepts_metadata_field_exact_policy_match()
-> Result<(), Box<dyn std::error::Error + Send + Sync>> { -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -3551,7 +3499,6 @@ async fn test_anonymous_post_object_accepts_metadata_field_exact_policy_match()
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_anonymous_post_object_allows_x_ignore_fields_outside_policy_conditions() async fn test_anonymous_post_object_allows_x_ignore_fields_outside_policy_conditions()
-> Result<(), Box<dyn std::error::Error + Send + Sync>> { -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -3604,7 +3551,6 @@ async fn test_anonymous_post_object_allows_x_ignore_fields_outside_policy_condit
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_anonymous_post_object_rejects_sigv4_date_policy_mismatch() -> Result<(), Box<dyn std::error::Error + Send + Sync>> { async fn test_anonymous_post_object_rejects_sigv4_date_policy_mismatch() -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -3657,7 +3603,6 @@ async fn test_anonymous_post_object_rejects_sigv4_date_policy_mismatch() -> Resu
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_anonymous_post_object_rejects_mismatched_bucket_form_field() -> Result<(), Box<dyn std::error::Error + Send + Sync>> async fn test_anonymous_post_object_rejects_mismatched_bucket_form_field() -> Result<(), Box<dyn std::error::Error + Send + Sync>>
{ {
init_logging(); init_logging();
@@ -3712,7 +3657,6 @@ async fn test_anonymous_post_object_rejects_mismatched_bucket_form_field() -> Re
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_anonymous_post_object_rejects_multiple_bucket_values() -> Result<(), Box<dyn std::error::Error + Send + Sync>> { async fn test_anonymous_post_object_rejects_multiple_bucket_values() -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -3764,7 +3708,6 @@ async fn test_anonymous_post_object_rejects_multiple_bucket_values() -> Result<(
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_anonymous_post_object_rejects_extra_content_disposition_field() async fn test_anonymous_post_object_rejects_extra_content_disposition_field()
-> Result<(), Box<dyn std::error::Error + Send + Sync>> { -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -3820,7 +3763,6 @@ async fn test_anonymous_post_object_rejects_extra_content_disposition_field()
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_signed_put_object_extract_expands_tar_entries_with_prefix_headers() async fn test_signed_put_object_extract_expands_tar_entries_with_prefix_headers()
-> Result<(), Box<dyn std::error::Error + Send + Sync>> { -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -3891,7 +3833,6 @@ async fn test_signed_put_object_extract_expands_tar_entries_with_prefix_headers(
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_signed_put_object_extract_preserves_request_metadata_on_extracted_objects() async fn test_signed_put_object_extract_preserves_request_metadata_on_extracted_objects()
-> Result<(), Box<dyn std::error::Error + Send + Sync>> { -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -3956,7 +3897,6 @@ async fn test_signed_put_object_extract_preserves_request_metadata_on_extracted_
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_signed_put_object_extract_preserves_sse_s3_and_redirect() -> Result<(), Box<dyn std::error::Error + Send + Sync>> { async fn test_signed_put_object_extract_preserves_sse_s3_and_redirect() -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -4004,7 +3944,6 @@ async fn test_signed_put_object_extract_preserves_sse_s3_and_redirect() -> Resul
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_signed_put_object_extract_preserves_storage_class() -> Result<(), Box<dyn std::error::Error + Send + Sync>> { async fn test_signed_put_object_extract_preserves_storage_class() -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -4047,7 +3986,6 @@ async fn test_signed_put_object_extract_preserves_storage_class() -> Result<(),
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_signed_put_object_extract_rejects_invalid_storage_class() -> Result<(), Box<dyn std::error::Error + Send + Sync>> { async fn test_signed_put_object_extract_rejects_invalid_storage_class() -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -4083,7 +4021,6 @@ async fn test_signed_put_object_extract_rejects_invalid_storage_class() -> Resul
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_signed_put_object_rejects_write_offset_bytes_header() -> Result<(), Box<dyn std::error::Error + Send + Sync>> { async fn test_signed_put_object_rejects_write_offset_bytes_header() -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -4137,7 +4074,6 @@ async fn test_signed_put_object_rejects_write_offset_bytes_header() -> Result<()
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_raw_signed_put_object_write_offset_bytes_returns_minio_compatible_error_body() async fn test_raw_signed_put_object_write_offset_bytes_returns_minio_compatible_error_body()
-> Result<(), Box<dyn std::error::Error + Send + Sync>> { -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -4176,7 +4112,6 @@ async fn test_raw_signed_put_object_write_offset_bytes_returns_minio_compatible_
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_anonymous_put_object_write_offset_bytes_returns_minio_compatible_error_body() async fn test_anonymous_put_object_write_offset_bytes_returns_minio_compatible_error_body()
-> Result<(), Box<dyn std::error::Error + Send + Sync>> { -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -4235,7 +4170,6 @@ async fn test_anonymous_put_object_write_offset_bytes_returns_minio_compatible_e
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_signed_put_object_extract_uses_bucket_default_sse_s3() -> Result<(), Box<dyn std::error::Error + Send + Sync>> { async fn test_signed_put_object_extract_uses_bucket_default_sse_s3() -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -4300,7 +4234,6 @@ async fn test_signed_put_object_extract_uses_bucket_default_sse_s3() -> Result<(
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_signed_put_object_extract_rejects_bucket_default_sse_kms() -> Result<(), Box<dyn std::error::Error + Send + Sync>> { async fn test_signed_put_object_extract_rejects_bucket_default_sse_kms() -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -4356,7 +4289,6 @@ async fn test_signed_put_object_extract_rejects_bucket_default_sse_kms() -> Resu
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_signed_put_object_extract_preserves_sse_c() -> Result<(), Box<dyn std::error::Error + Send + Sync>> { async fn test_signed_put_object_extract_preserves_sse_c() -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -4421,7 +4353,6 @@ async fn test_signed_put_object_extract_preserves_sse_c() -> Result<(), Box<dyn
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_signed_put_object_extract_preserves_object_lock_legal_hold() -> Result<(), Box<dyn std::error::Error + Send + Sync>> async fn test_signed_put_object_extract_preserves_object_lock_legal_hold() -> Result<(), Box<dyn std::error::Error + Send + Sync>>
{ {
init_logging(); init_logging();
@@ -4476,7 +4407,6 @@ async fn test_signed_put_object_extract_preserves_object_lock_legal_hold() -> Re
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_signed_put_object_extract_preserves_object_lock_retention() -> Result<(), Box<dyn std::error::Error + Send + Sync>> async fn test_signed_put_object_extract_preserves_object_lock_retention() -> Result<(), Box<dyn std::error::Error + Send + Sync>>
{ {
init_logging(); init_logging();
@@ -4536,7 +4466,6 @@ async fn test_signed_put_object_extract_preserves_object_lock_retention() -> Res
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_signed_put_object_extract_pax_retention_overrides_request_retention() async fn test_signed_put_object_extract_pax_retention_overrides_request_retention()
-> Result<(), Box<dyn std::error::Error + Send + Sync>> { -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -4600,7 +4529,6 @@ async fn test_signed_put_object_extract_pax_retention_overrides_request_retentio
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_signed_put_object_extract_returns_archive_etag() -> Result<(), Box<dyn std::error::Error + Send + Sync>> { async fn test_signed_put_object_extract_returns_archive_etag() -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -4634,7 +4562,6 @@ async fn test_signed_put_object_extract_returns_archive_etag() -> Result<(), Box
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_signed_put_object_extract_preserves_entry_mtime() -> Result<(), Box<dyn std::error::Error + Send + Sync>> { async fn test_signed_put_object_extract_preserves_entry_mtime() -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -4670,7 +4597,6 @@ async fn test_signed_put_object_extract_preserves_entry_mtime() -> Result<(), Bo
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_signed_put_object_extract_preserves_pax_metadata_and_version_id() async fn test_signed_put_object_extract_preserves_pax_metadata_and_version_id()
-> Result<(), Box<dyn std::error::Error + Send + Sync>> { -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -4724,7 +4650,6 @@ async fn test_signed_put_object_extract_preserves_pax_metadata_and_version_id()
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_signed_put_object_extract_authorizes_each_pax_privilege_and_retention_conditions() async fn test_signed_put_object_extract_authorizes_each_pax_privilege_and_retention_conditions()
-> Result<(), Box<dyn std::error::Error + Send + Sync>> { -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -5034,7 +4959,6 @@ async fn test_signed_put_object_extract_authorizes_each_pax_privilege_and_retent
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_signed_put_object_extract_accepts_compat_header() -> Result<(), Box<dyn std::error::Error + Send + Sync>> { async fn test_signed_put_object_extract_accepts_compat_header() -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -5076,7 +5000,6 @@ async fn test_signed_put_object_extract_accepts_compat_header() -> Result<(), Bo
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_signed_put_object_extract_preserves_directory_markers_by_default() async fn test_signed_put_object_extract_preserves_directory_markers_by_default()
-> Result<(), Box<dyn std::error::Error + Send + Sync>> { -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -5137,7 +5060,6 @@ async fn test_signed_put_object_extract_preserves_directory_markers_by_default()
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_signed_put_object_extract_expands_tar_gz_archive() -> Result<(), Box<dyn std::error::Error + Send + Sync>> { async fn test_signed_put_object_extract_expands_tar_gz_archive() -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -5189,7 +5111,6 @@ async fn test_signed_put_object_extract_expands_tar_gz_archive() -> Result<(), B
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_signed_put_object_extract_expands_tgz_archive() -> Result<(), Box<dyn std::error::Error + Send + Sync>> { async fn test_signed_put_object_extract_expands_tgz_archive() -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -5241,7 +5162,6 @@ async fn test_signed_put_object_extract_expands_tgz_archive() -> Result<(), Box<
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_signed_put_object_extract_expands_tbz2_archive() -> Result<(), Box<dyn std::error::Error + Send + Sync>> { async fn test_signed_put_object_extract_expands_tbz2_archive() -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -5293,7 +5213,6 @@ async fn test_signed_put_object_extract_expands_tbz2_archive() -> Result<(), Box
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_signed_put_object_extract_expands_txz_archive() -> Result<(), Box<dyn std::error::Error + Send + Sync>> { async fn test_signed_put_object_extract_expands_txz_archive() -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -5345,7 +5264,6 @@ async fn test_signed_put_object_extract_expands_txz_archive() -> Result<(), Box<
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_signed_put_object_extract_skips_invalid_entry_when_ignore_errors_enabled() async fn test_signed_put_object_extract_skips_invalid_entry_when_ignore_errors_enabled()
-> Result<(), Box<dyn std::error::Error + Send + Sync>> { -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -5419,7 +5337,6 @@ async fn test_signed_put_object_extract_skips_invalid_entry_when_ignore_errors_e
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_signed_put_object_extract_normalizes_prefix_header_value() -> Result<(), Box<dyn std::error::Error + Send + Sync>> { async fn test_signed_put_object_extract_normalizes_prefix_header_value() -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -5462,7 +5379,6 @@ async fn test_signed_put_object_extract_normalizes_prefix_header_value() -> Resu
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_signed_put_object_extract_expands_tzst_archive() -> Result<(), Box<dyn std::error::Error + Send + Sync>> { async fn test_signed_put_object_extract_expands_tzst_archive() -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -5514,7 +5430,6 @@ async fn test_signed_put_object_extract_expands_tzst_archive() -> Result<(), Box
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_signed_put_object_extract_rejects_missing_archive_extension() -> Result<(), Box<dyn std::error::Error + Send + Sync>> async fn test_signed_put_object_extract_rejects_missing_archive_extension() -> Result<(), Box<dyn std::error::Error + Send + Sync>>
{ {
init_logging(); init_logging();
@@ -5548,7 +5463,6 @@ async fn test_signed_put_object_extract_rejects_missing_archive_extension() -> R
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_signed_put_object_extract_rejects_invalid_tar_gz_payload() -> Result<(), Box<dyn std::error::Error + Send + Sync>> { async fn test_signed_put_object_extract_rejects_invalid_tar_gz_payload() -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
init_logging(); init_logging();
@@ -33,7 +33,6 @@ use aws_sdk_s3::types::{
ObjectLockMode, ObjectLockRetentionMode, ObjectLockMode, ObjectLockRetentionMode,
}; };
use chrono::{DateTime, Duration, Utc}; use chrono::{DateTime, Duration, Utc};
use serial_test::serial;
use tracing::info; use tracing::info;
/// Initialize test logging /// Initialize test logging
@@ -107,7 +106,6 @@ fn parse_s3_datetime(value: &aws_sdk_s3::primitives::DateTime) -> DateTime<Utc>
// ============================================================================ // ============================================================================
#[tokio::test] #[tokio::test]
#[serial]
async fn test_delete_object_blocked_by_compliance_retention() { async fn test_delete_object_blocked_by_compliance_retention() {
init_logging(); init_logging();
info!("🧪 Test: DeleteObject blocked by COMPLIANCE retention"); info!("🧪 Test: DeleteObject blocked by COMPLIANCE retention");
@@ -145,7 +143,6 @@ async fn test_delete_object_blocked_by_compliance_retention() {
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_delete_object_blocked_by_governance_without_bypass() { async fn test_delete_object_blocked_by_governance_without_bypass() {
init_logging(); init_logging();
info!("🧪 Test: DeleteObject blocked by GOVERNANCE retention without bypass"); info!("🧪 Test: DeleteObject blocked by GOVERNANCE retention without bypass");
@@ -175,7 +172,6 @@ async fn test_delete_object_blocked_by_governance_without_bypass() {
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_delete_object_allowed_by_governance_with_bypass() { async fn test_delete_object_allowed_by_governance_with_bypass() {
init_logging(); init_logging();
info!("🧪 Test: DeleteObject allowed by GOVERNANCE retention with bypass"); info!("🧪 Test: DeleteObject allowed by GOVERNANCE retention with bypass");
@@ -215,7 +211,6 @@ async fn test_delete_object_allowed_by_governance_with_bypass() {
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_delete_object_creates_delete_marker_for_retained_current_version() { async fn test_delete_object_creates_delete_marker_for_retained_current_version() {
init_logging(); init_logging();
info!("🧪 Test: DeleteObject creates delete marker for retained current version"); info!("🧪 Test: DeleteObject creates delete marker for retained current version");
@@ -266,7 +261,6 @@ async fn test_delete_object_creates_delete_marker_for_retained_current_version()
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_delete_object_blocked_by_legal_hold() { async fn test_delete_object_blocked_by_legal_hold() {
init_logging(); init_logging();
info!("🧪 Test: DeleteObject blocked by Legal Hold"); info!("🧪 Test: DeleteObject blocked by Legal Hold");
@@ -299,7 +293,6 @@ async fn test_delete_object_blocked_by_legal_hold() {
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_delete_object_allowed_with_legal_hold_off() { async fn test_delete_object_allowed_with_legal_hold_off() {
init_logging(); init_logging();
info!("🧪 Test: DeleteObject allowed with Legal Hold OFF"); info!("🧪 Test: DeleteObject allowed with Legal Hold OFF");
@@ -335,7 +328,6 @@ async fn test_delete_object_allowed_with_legal_hold_off() {
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_delete_object_after_legal_hold_removed() { async fn test_delete_object_after_legal_hold_removed() {
init_logging(); init_logging();
info!("🧪 Test: DeleteObject succeeds after Legal Hold is removed"); info!("🧪 Test: DeleteObject succeeds after Legal Hold is removed");
@@ -369,7 +361,6 @@ async fn test_delete_object_after_legal_hold_removed() {
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_get_object_legal_hold_returns_updated_status() { async fn test_get_object_legal_hold_returns_updated_status() {
init_logging(); init_logging();
info!("🧪 Test: GetObjectLegalHold returns updated status"); info!("🧪 Test: GetObjectLegalHold returns updated status");
@@ -425,7 +416,6 @@ async fn test_get_object_legal_hold_returns_updated_status() {
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_get_object_retention_returns_configured_values() { async fn test_get_object_retention_returns_configured_values() {
init_logging(); init_logging();
info!("🧪 Test: GetObjectRetention returns configured values"); info!("🧪 Test: GetObjectRetention returns configured values");
@@ -476,7 +466,6 @@ async fn test_get_object_retention_returns_configured_values() {
// creating a new current version. The lock protects the existing version // creating a new current version. The lock protects the existing version
// from deletion; it never blocks new versions. // from deletion; it never blocks new versions.
#[tokio::test] #[tokio::test]
#[serial]
async fn test_put_object_overwrite_creates_new_version_under_legal_hold() { async fn test_put_object_overwrite_creates_new_version_under_legal_hold() {
init_logging(); init_logging();
info!("🧪 Test: PutObject overwrite of a legal-hold version creates a new version"); info!("🧪 Test: PutObject overwrite of a legal-hold version creates a new version");
@@ -561,7 +550,6 @@ async fn test_put_object_overwrite_creates_new_version_under_legal_hold() {
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_copy_object_applies_requested_legal_hold() { async fn test_copy_object_applies_requested_legal_hold() {
init_logging(); init_logging();
info!("🧪 Test: CopyObject applies requested Legal Hold"); info!("🧪 Test: CopyObject applies requested Legal Hold");
@@ -613,7 +601,6 @@ async fn test_copy_object_applies_requested_legal_hold() {
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_copy_object_does_not_inherit_source_legal_hold() { async fn test_copy_object_does_not_inherit_source_legal_hold() {
init_logging(); init_logging();
info!("🧪 Test: CopyObject does not inherit source Legal Hold"); info!("🧪 Test: CopyObject does not inherit source Legal Hold");
@@ -707,7 +694,6 @@ async fn test_copy_object_does_not_inherit_source_legal_hold() {
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_copy_object_overwrite_creates_new_version_under_legal_hold() { async fn test_copy_object_overwrite_creates_new_version_under_legal_hold() {
init_logging(); init_logging();
info!("🧪 Test: CopyObject overwrite of a legal-hold destination creates a new version"); info!("🧪 Test: CopyObject overwrite of a legal-hold destination creates a new version");
@@ -787,7 +773,6 @@ async fn test_copy_object_overwrite_creates_new_version_under_legal_hold() {
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_create_multipart_upload_applies_requested_legal_hold() { async fn test_create_multipart_upload_applies_requested_legal_hold() {
init_logging(); init_logging();
info!("🧪 Test: CreateMultipartUpload applies requested Legal Hold"); info!("🧪 Test: CreateMultipartUpload applies requested Legal Hold");
@@ -853,7 +838,6 @@ async fn test_create_multipart_upload_applies_requested_legal_hold() {
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_create_multipart_upload_creates_new_version_under_compliance_retention() { async fn test_create_multipart_upload_creates_new_version_under_compliance_retention() {
init_logging(); init_logging();
info!("🧪 Test: CreateMultipartUpload over a COMPLIANCE-retained key creates a new version"); info!("🧪 Test: CreateMultipartUpload over a COMPLIANCE-retained key creates a new version");
@@ -933,7 +917,6 @@ async fn test_create_multipart_upload_creates_new_version_under_compliance_reten
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_delete_completed_multipart_object_blocked_by_legal_hold() { async fn test_delete_completed_multipart_object_blocked_by_legal_hold() {
init_logging(); init_logging();
info!("🧪 Test: Delete completed multipart object blocked by Legal Hold"); info!("🧪 Test: Delete completed multipart object blocked by Legal Hold");
@@ -993,7 +976,6 @@ async fn test_delete_completed_multipart_object_blocked_by_legal_hold() {
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_delete_completed_multipart_object_blocked_by_retention() { async fn test_delete_completed_multipart_object_blocked_by_retention() {
init_logging(); init_logging();
info!("🧪 Test: Delete completed multipart object blocked by retention"); info!("🧪 Test: Delete completed multipart object blocked by retention");
@@ -1055,7 +1037,6 @@ async fn test_delete_completed_multipart_object_blocked_by_retention() {
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_complete_multipart_upload_creates_new_version_under_legal_hold() { async fn test_complete_multipart_upload_creates_new_version_under_legal_hold() {
init_logging(); init_logging();
info!("🧪 Test: CompleteMultipartUpload creates a new version when the current version is under Legal Hold"); info!("🧪 Test: CompleteMultipartUpload creates a new version when the current version is under Legal Hold");
@@ -1135,7 +1116,6 @@ async fn test_complete_multipart_upload_creates_new_version_under_legal_hold() {
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_complete_multipart_upload_creates_new_version_under_compliance_retention() { async fn test_complete_multipart_upload_creates_new_version_under_compliance_retention() {
init_logging(); init_logging();
info!("🧪 Test: CompleteMultipartUpload creates a new version when the current version is under COMPLIANCE retention"); info!("🧪 Test: CompleteMultipartUpload creates a new version when the current version is under COMPLIANCE retention");
@@ -1209,7 +1189,6 @@ async fn test_complete_multipart_upload_creates_new_version_under_compliance_ret
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_write_paths_require_put_object_legal_hold_permission() { async fn test_write_paths_require_put_object_legal_hold_permission() {
init_logging(); init_logging();
info!("🧪 Test: write paths require PutObjectLegalHold permission"); info!("🧪 Test: write paths require PutObjectLegalHold permission");
@@ -1273,7 +1252,6 @@ async fn test_write_paths_require_put_object_legal_hold_permission() {
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_write_paths_require_put_object_retention_permission() { async fn test_write_paths_require_put_object_retention_permission() {
init_logging(); init_logging();
info!("🧪 Test: write paths require PutObjectRetention permission"); info!("🧪 Test: write paths require PutObjectRetention permission");
@@ -1345,7 +1323,6 @@ async fn test_write_paths_require_put_object_retention_permission() {
// ============================================================================ // ============================================================================
#[tokio::test] #[tokio::test]
#[serial]
async fn test_delete_objects_mixed_locked_unlocked() { async fn test_delete_objects_mixed_locked_unlocked() {
init_logging(); init_logging();
info!("🧪 Test: DeleteObjects with mixed locked and unlocked objects"); info!("🧪 Test: DeleteObjects with mixed locked and unlocked objects");
@@ -1427,7 +1404,6 @@ async fn test_delete_objects_mixed_locked_unlocked() {
// ============================================================================ // ============================================================================
#[tokio::test] #[tokio::test]
#[serial]
async fn test_put_retention_compliance_cannot_shorten() { async fn test_put_retention_compliance_cannot_shorten() {
init_logging(); init_logging();
info!("🧪 Test: PutObjectRetention cannot shorten COMPLIANCE retention"); info!("🧪 Test: PutObjectRetention cannot shorten COMPLIANCE retention");
@@ -1468,7 +1444,6 @@ async fn test_put_retention_compliance_cannot_shorten() {
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_put_retention_compliance_can_extend() { async fn test_put_retention_compliance_can_extend() {
init_logging(); init_logging();
info!("🧪 Test: PutObjectRetention can extend COMPLIANCE retention"); info!("🧪 Test: PutObjectRetention can extend COMPLIANCE retention");
@@ -1509,7 +1484,6 @@ async fn test_put_retention_compliance_can_extend() {
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_put_retention_governance_extend_without_bypass() { async fn test_put_retention_governance_extend_without_bypass() {
init_logging(); init_logging();
info!("🧪 Test: PutObjectRetention on GOVERNANCE can extend without bypass"); info!("🧪 Test: PutObjectRetention on GOVERNANCE can extend without bypass");
@@ -1553,7 +1527,6 @@ async fn test_put_retention_governance_extend_without_bypass() {
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_put_retention_governance_shorten_requires_bypass() { async fn test_put_retention_governance_shorten_requires_bypass() {
init_logging(); init_logging();
info!("🧪 Test: PutObjectRetention on GOVERNANCE requires bypass to shorten"); info!("🧪 Test: PutObjectRetention on GOVERNANCE requires bypass to shorten");
@@ -1615,7 +1588,6 @@ async fn test_put_retention_governance_shorten_requires_bypass() {
// ============================================================================ // ============================================================================
#[tokio::test] #[tokio::test]
#[serial]
async fn test_default_retention_applied_to_new_objects() { async fn test_default_retention_applied_to_new_objects() {
init_logging(); init_logging();
info!("🧪 Test: Default retention is applied to new objects"); info!("🧪 Test: Default retention is applied to new objects");
@@ -1685,7 +1657,6 @@ async fn test_default_retention_applied_to_new_objects() {
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_delete_object_creates_delete_marker_for_default_retained_current_version() { async fn test_delete_object_creates_delete_marker_for_default_retained_current_version() {
init_logging(); init_logging();
info!("🧪 Test: DeleteObject creates delete marker for default-retained current version"); info!("🧪 Test: DeleteObject creates delete marker for default-retained current version");
@@ -1770,7 +1741,6 @@ async fn test_delete_object_creates_delete_marker_for_default_retained_current_v
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_put_copy_and_multipart_reject_incomplete_retention_headers() { async fn test_put_copy_and_multipart_reject_incomplete_retention_headers() {
init_logging(); init_logging();
info!("🧪 Test: write paths reject incomplete Object Lock retention headers"); info!("🧪 Test: write paths reject incomplete Object Lock retention headers");
@@ -1869,7 +1839,6 @@ async fn test_put_copy_and_multipart_reject_incomplete_retention_headers() {
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_copy_object_retention_uses_destination_policy() { async fn test_copy_object_retention_uses_destination_policy() {
init_logging(); init_logging();
info!("🧪 Test: CopyObject retention follows destination policy"); info!("🧪 Test: CopyObject retention follows destination policy");
@@ -2051,7 +2020,6 @@ async fn test_copy_object_retention_uses_destination_policy() {
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_multipart_default_retention_fixed_at_create() { async fn test_multipart_default_retention_fixed_at_create() {
init_logging(); init_logging();
info!("🧪 Test: multipart default retention is fixed at CreateMultipartUpload"); info!("🧪 Test: multipart default retention is fixed at CreateMultipartUpload");
@@ -2122,7 +2090,6 @@ async fn test_multipart_default_retention_fixed_at_create() {
// ============================================================================ // ============================================================================
#[tokio::test] #[tokio::test]
#[serial]
async fn test_unretained_object_lock_object_delete_and_bucket_cleanup() { async fn test_unretained_object_lock_object_delete_and_bucket_cleanup() {
init_logging(); init_logging();
info!("🧪 Test: Unretained Object Lock object delete and bucket cleanup (Issue #5339)"); info!("🧪 Test: Unretained Object Lock object delete and bucket cleanup (Issue #5339)");
@@ -2243,7 +2210,6 @@ async fn test_unretained_object_lock_object_delete_and_bucket_cleanup() {
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_versioning_auto_enabled_with_object_lock() { async fn test_versioning_auto_enabled_with_object_lock() {
init_logging(); init_logging();
info!("🧪 Test: Versioning is auto-enabled when Object Lock is configured"); info!("🧪 Test: Versioning is auto-enabled when Object Lock is configured");
@@ -2302,7 +2268,6 @@ async fn test_versioning_auto_enabled_with_object_lock() {
// ============================================================================ // ============================================================================
#[tokio::test] #[tokio::test]
#[serial]
async fn test_error_message_distinguishes_legal_hold_from_retention() { async fn test_error_message_distinguishes_legal_hold_from_retention() {
init_logging(); init_logging();
info!("🧪 Test: Error messages distinguish Legal Hold from Retention"); info!("🧪 Test: Error messages distinguish Legal Hold from Retention");
@@ -60,7 +60,6 @@ use rustfs_signer::constants::UNSIGNED_PAYLOAD;
use rustfs_signer::sign_v4; use rustfs_signer::sign_v4;
use s3s::Body; use s3s::Body;
use s3s::header::X_AMZ_REPLICATION_STATUS; use s3s::header::X_AMZ_REPLICATION_STATUS;
use serial_test::serial;
use sha2::{Digest, Sha256}; use sha2::{Digest, Sha256};
use std::collections::BTreeMap; use std::collections::BTreeMap;
use std::convert::Infallible; use std::convert::Infallible;
@@ -2506,7 +2505,6 @@ async fn build_replication_pair(
/// metadata was inherited wholesale from the source, so the scanner heal pass /// metadata was inherited wholesale from the source, so the scanner heal pass
/// skipped it too — no PENDING/FAILED marker meant nothing to re-drive). /// skipped it too — no PENDING/FAILED marker meant nothing to re-drive).
#[tokio::test] #[tokio::test]
#[serial]
async fn test_copy_object_replicates_to_target() -> TestResult { async fn test_copy_object_replicates_to_target() -> TestResult {
init_logging(); init_logging();
@@ -2555,7 +2553,6 @@ async fn test_copy_object_replicates_to_target() -> TestResult {
/// independent object; every member must replicate to the remote target like a /// independent object; every member must replicate to the remote target like a
/// regular PUT (MinIO PutObjectExtract parity). /// regular PUT (MinIO PutObjectExtract parity).
#[tokio::test] #[tokio::test]
#[serial]
async fn test_snowball_extract_replicates_members_to_target() -> TestResult { async fn test_snowball_extract_replicates_members_to_target() -> TestResult {
init_logging(); init_logging();
@@ -2601,7 +2598,6 @@ async fn test_snowball_extract_replicates_members_to_target() -> TestResult {
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_replication_check_succeeds_with_remote_target() -> Result<(), Box<dyn Error + Send + Sync>> { async fn test_replication_check_succeeds_with_remote_target() -> Result<(), Box<dyn Error + Send + Sync>> {
init_logging(); init_logging();
@@ -2638,7 +2634,6 @@ async fn test_replication_check_succeeds_with_remote_target() -> Result<(), Box<
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_replication_check_rejects_target_without_object_lock() -> Result<(), Box<dyn Error + Send + Sync>> { async fn test_replication_check_rejects_target_without_object_lock() -> Result<(), Box<dyn Error + Send + Sync>> {
init_logging(); init_logging();
@@ -2692,7 +2687,6 @@ async fn test_replication_check_rejects_target_without_object_lock() -> Result<(
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_set_remote_target_rejects_unversioned_source_bucket() -> Result<(), Box<dyn Error + Send + Sync>> { async fn test_set_remote_target_rejects_unversioned_source_bucket() -> Result<(), Box<dyn Error + Send + Sync>> {
init_logging(); init_logging();
@@ -2731,7 +2725,6 @@ async fn test_set_remote_target_rejects_unversioned_source_bucket() -> Result<()
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_replication_check_rejects_unversioned_source_bucket() -> Result<(), Box<dyn Error + Send + Sync>> { async fn test_replication_check_rejects_unversioned_source_bucket() -> Result<(), Box<dyn Error + Send + Sync>> {
init_logging(); init_logging();
@@ -2755,7 +2748,6 @@ async fn test_replication_check_rejects_unversioned_source_bucket() -> Result<()
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_replication_check_rejects_missing_replication_config() -> Result<(), Box<dyn Error + Send + Sync>> { async fn test_replication_check_rejects_missing_replication_config() -> Result<(), Box<dyn Error + Send + Sync>> {
init_logging(); init_logging();
@@ -2779,7 +2771,6 @@ async fn test_replication_check_rejects_missing_replication_config() -> Result<(
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_replication_check_rejects_invalid_bucket() -> Result<(), Box<dyn Error + Send + Sync>> { async fn test_replication_check_rejects_invalid_bucket() -> Result<(), Box<dyn Error + Send + Sync>> {
init_logging(); init_logging();
@@ -2798,7 +2789,6 @@ async fn test_replication_check_rejects_invalid_bucket() -> Result<(), Box<dyn E
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_set_remote_target_rejects_same_bucket_on_same_deployment() -> Result<(), Box<dyn Error + Send + Sync>> { async fn test_set_remote_target_rejects_same_bucket_on_same_deployment() -> Result<(), Box<dyn Error + Send + Sync>> {
init_logging(); init_logging();
@@ -2842,7 +2832,6 @@ async fn test_set_remote_target_rejects_same_bucket_on_same_deployment() -> Resu
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_set_remote_target_rejects_unversioned_target_bucket() -> Result<(), Box<dyn Error + Send + Sync>> { async fn test_set_remote_target_rejects_unversioned_target_bucket() -> Result<(), Box<dyn Error + Send + Sync>> {
init_logging(); init_logging();
@@ -2876,7 +2865,6 @@ async fn test_set_remote_target_rejects_unversioned_target_bucket() -> Result<()
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_set_remote_target_update_requires_arn() -> Result<(), Box<dyn Error + Send + Sync>> { async fn test_set_remote_target_update_requires_arn() -> Result<(), Box<dyn Error + Send + Sync>> {
init_logging(); init_logging();
@@ -2928,7 +2916,6 @@ async fn test_set_remote_target_update_requires_arn() -> Result<(), Box<dyn Erro
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_set_remote_target_update_rejects_missing_target() -> Result<(), Box<dyn Error + Send + Sync>> { async fn test_set_remote_target_update_rejects_missing_target() -> Result<(), Box<dyn Error + Send + Sync>> {
init_logging(); init_logging();
@@ -3017,7 +3004,6 @@ async fn fetch_single_target(
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_set_remote_target_partial_update_preserves_credentials() -> Result<(), Box<dyn Error + Send + Sync>> { async fn test_set_remote_target_partial_update_preserves_credentials() -> Result<(), Box<dyn Error + Send + Sync>> {
init_logging(); init_logging();
@@ -3103,7 +3089,6 @@ async fn test_set_remote_target_partial_update_preserves_credentials() -> Result
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_set_remote_target_rejects_invalid_target_url() -> Result<(), Box<dyn Error + Send + Sync>> { async fn test_set_remote_target_rejects_invalid_target_url() -> Result<(), Box<dyn Error + Send + Sync>> {
init_logging(); init_logging();
@@ -3145,7 +3130,6 @@ async fn test_set_remote_target_rejects_invalid_target_url() -> Result<(), Box<d
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_set_remote_target_rejects_self_signed_https_target_without_skip_tls_verify() async fn test_set_remote_target_rejects_self_signed_https_target_without_skip_tls_verify()
-> Result<(), Box<dyn Error + Send + Sync>> { -> Result<(), Box<dyn Error + Send + Sync>> {
init_logging(); init_logging();
@@ -3230,7 +3214,6 @@ async fn test_set_remote_target_rejects_self_signed_https_target_without_skip_tl
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_set_remote_target_allows_self_signed_https_target_with_skip_tls_verify() -> Result<(), Box<dyn Error + Send + Sync>> async fn test_set_remote_target_allows_self_signed_https_target_with_skip_tls_verify() -> Result<(), Box<dyn Error + Send + Sync>>
{ {
init_logging(); init_logging();
@@ -3342,7 +3325,6 @@ async fn test_set_remote_target_allows_self_signed_https_target_with_skip_tls_ve
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_set_remote_target_rejects_private_ca_https_target_without_ca_cert_pem() -> Result<(), Box<dyn Error + Send + Sync>> async fn test_set_remote_target_rejects_private_ca_https_target_without_ca_cert_pem() -> Result<(), Box<dyn Error + Send + Sync>>
{ {
init_logging(); init_logging();
@@ -3427,7 +3409,6 @@ async fn test_set_remote_target_rejects_private_ca_https_target_without_ca_cert_
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_set_remote_target_allows_private_ca_https_target_with_ca_cert_pem() -> Result<(), Box<dyn Error + Send + Sync>> { async fn test_set_remote_target_allows_private_ca_https_target_with_ca_cert_pem() -> Result<(), Box<dyn Error + Send + Sync>> {
init_logging(); init_logging();
@@ -3518,7 +3499,6 @@ async fn test_set_remote_target_allows_private_ca_https_target_with_ca_cert_pem(
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_list_remote_targets_rejects_empty_bucket() -> Result<(), Box<dyn Error + Send + Sync>> { async fn test_list_remote_targets_rejects_empty_bucket() -> Result<(), Box<dyn Error + Send + Sync>> {
init_logging(); init_logging();
@@ -3538,7 +3518,6 @@ async fn test_list_remote_targets_rejects_empty_bucket() -> Result<(), Box<dyn E
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_list_remote_targets_rejects_invalid_bucket() -> Result<(), Box<dyn Error + Send + Sync>> { async fn test_list_remote_targets_rejects_invalid_bucket() -> Result<(), Box<dyn Error + Send + Sync>> {
init_logging(); init_logging();
@@ -3557,7 +3536,6 @@ async fn test_list_remote_targets_rejects_invalid_bucket() -> Result<(), Box<dyn
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_remove_remote_target_rejects_missing_target() -> Result<(), Box<dyn Error + Send + Sync>> { async fn test_remove_remote_target_rejects_missing_target() -> Result<(), Box<dyn Error + Send + Sync>> {
init_logging(); init_logging();
@@ -3598,7 +3576,6 @@ async fn test_remove_remote_target_rejects_missing_target() -> Result<(), Box<dy
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_remove_remote_target_rejects_missing_arn() -> Result<(), Box<dyn Error + Send + Sync>> { async fn test_remove_remote_target_rejects_missing_arn() -> Result<(), Box<dyn Error + Send + Sync>> {
init_logging(); init_logging();
@@ -3623,7 +3600,6 @@ async fn test_remove_remote_target_rejects_missing_arn() -> Result<(), Box<dyn E
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_remove_remote_target_rejects_invalid_bucket() -> Result<(), Box<dyn Error + Send + Sync>> { async fn test_remove_remote_target_rejects_invalid_bucket() -> Result<(), Box<dyn Error + Send + Sync>> {
init_logging(); init_logging();
@@ -3647,7 +3623,6 @@ async fn test_remove_remote_target_rejects_invalid_bucket() -> Result<(), Box<dy
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_remove_remote_target_rejects_target_used_by_replication() -> Result<(), Box<dyn Error + Send + Sync>> { async fn test_remove_remote_target_rejects_target_used_by_replication() -> Result<(), Box<dyn Error + Send + Sync>> {
init_logging(); init_logging();
@@ -3687,7 +3662,6 @@ async fn test_remove_remote_target_rejects_target_used_by_replication() -> Resul
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_delete_bucket_replication_removes_remote_target() -> Result<(), Box<dyn Error + Send + Sync>> { async fn test_delete_bucket_replication_removes_remote_target() -> Result<(), Box<dyn Error + Send + Sync>> {
init_logging(); init_logging();
@@ -3737,7 +3711,6 @@ async fn test_delete_bucket_replication_removes_remote_target() -> Result<(), Bo
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_bucket_replication_replicates_put_object_issue_2539() -> Result<(), Box<dyn Error + Send + Sync>> { async fn test_bucket_replication_replicates_put_object_issue_2539() -> Result<(), Box<dyn Error + Send + Sync>> {
init_logging(); init_logging();
@@ -3779,7 +3752,6 @@ async fn test_bucket_replication_replicates_put_object_issue_2539() -> Result<()
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_bucket_replication_converges_delete_marker_and_version_purge() -> TestResult { async fn test_bucket_replication_converges_delete_marker_and_version_purge() -> TestResult {
init_logging(); init_logging();
@@ -3878,7 +3850,6 @@ async fn test_bucket_replication_converges_delete_marker_and_version_purge() ->
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_bucket_replication_disabled_delete_marker_does_not_propagate() -> TestResult { async fn test_bucket_replication_disabled_delete_marker_does_not_propagate() -> TestResult {
init_logging(); init_logging();
@@ -3965,7 +3936,6 @@ async fn test_bucket_replication_disabled_delete_marker_does_not_propagate() ->
/// interoperability profile for a runner that provisions MinIO credentials /// interoperability profile for a runner that provisions MinIO credentials
/// and a reachable endpoint. /// and a reachable endpoint.
#[tokio::test] #[tokio::test]
#[serial]
async fn test_bucket_replication_acceptance_matrix_local_dual_targets() -> TestResult { async fn test_bucket_replication_acceptance_matrix_local_dual_targets() -> TestResult {
init_logging(); init_logging();
@@ -4293,7 +4263,6 @@ async fn test_bucket_replication_acceptance_matrix_local_dual_targets() -> TestR
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_single_bucket_multipart_replication_fans_out_to_multiple_targets() -> Result<(), Box<dyn Error + Send + Sync>> { async fn test_single_bucket_multipart_replication_fans_out_to_multiple_targets() -> Result<(), Box<dyn Error + Send + Sync>> {
init_logging(); init_logging();
@@ -4462,7 +4431,6 @@ async fn test_repl17_failure_observation_helpers() -> TestResult {
/// the replica is decryptable only with the original customer key. The /// the replica is decryptable only with the original customer key. The
/// backlog#1291 property still holds: never a silent plaintext replica. /// backlog#1291 property still holds: never a silent plaintext replica.
#[tokio::test] #[tokio::test]
#[serial]
async fn test_bucket_replication_sse_c_contract() -> TestResult { async fn test_bucket_replication_sse_c_contract() -> TestResult {
init_logging(); init_logging();
@@ -4540,7 +4508,6 @@ async fn test_bucket_replication_sse_c_contract() -> TestResult {
/// part — part boundaries and the encrypted-multipart marker survive so the /// part — part boundaries and the encrypted-multipart marker survive so the
/// replica decrypts each part with its part-derived nonce. /// replica decrypts each part with its part-derived nonce.
#[tokio::test] #[tokio::test]
#[serial]
async fn test_bucket_replication_sse_c_multipart_passthrough() -> TestResult { async fn test_bucket_replication_sse_c_multipart_passthrough() -> TestResult {
init_logging(); init_logging();
@@ -4657,7 +4624,6 @@ async fn test_bucket_replication_sse_c_multipart_passthrough() -> TestResult {
/// (independent KMS, so success proves target-owned envelopes), preserved /// (independent KMS, so success proves target-owned envelopes), preserved
/// source ETag, and a version that stays stable across scanner cycles. /// source ETag, and a version that stays stable across scanner cycles.
#[tokio::test] #[tokio::test]
#[serial]
async fn test_bucket_replication_sse_s3_contract() -> TestResult { async fn test_bucket_replication_sse_s3_contract() -> TestResult {
init_logging(); init_logging();
assert_managed_sse_replicates_and_reencrypts("sse-s3", false).await assert_managed_sse_replicates_and_reencrypts("sse-s3", false).await
@@ -4667,7 +4633,6 @@ async fn test_bucket_replication_sse_s3_contract() -> TestResult {
/// fail closed — replication FAILED, and no plaintext (or any) replica ever /// fail closed — replication FAILED, and no plaintext (or any) replica ever
/// materializes on the target. /// materializes on the target.
#[tokio::test] #[tokio::test]
#[serial]
async fn test_bucket_replication_sse_s3_fails_closed_without_target_kms() -> TestResult { async fn test_bucket_replication_sse_s3_fails_closed_without_target_kms() -> TestResult {
init_logging(); init_logging();
@@ -4711,7 +4676,6 @@ async fn test_bucket_replication_sse_s3_fails_closed_without_target_kms() -> Tes
/// the ETag comparison sees the preserved source ETag on the replica and does /// the ETag comparison sees the preserved source ETag on the replica and does
/// not rewrite it, so the replica's version stays stable through the resync. /// not rewrite it, so the replica's version stays stable through the resync.
#[tokio::test] #[tokio::test]
#[serial]
async fn test_bucket_replication_sse_s3_resync_converges() -> TestResult { async fn test_bucket_replication_sse_s3_resync_converges() -> TestResult {
init_logging(); init_logging();
@@ -4768,7 +4732,6 @@ async fn test_bucket_replication_sse_s3_resync_converges() -> TestResult {
/// re-encrypts under its own default key. The independent-KMS pair proves the /// re-encrypts under its own default key. The independent-KMS pair proves the
/// replica's envelope is target-owned. /// replica's envelope is target-owned.
#[tokio::test] #[tokio::test]
#[serial]
async fn test_bucket_replication_sse_kms_contract() -> TestResult { async fn test_bucket_replication_sse_kms_contract() -> TestResult {
init_logging(); init_logging();
assert_managed_sse_replicates_and_reencrypts("sse-kms", true).await assert_managed_sse_replicates_and_reencrypts("sse-kms", true).await
@@ -4779,7 +4742,6 @@ async fn test_bucket_replication_sse_kms_contract() -> TestResult {
/// carries the full header set (SSE intent, content-type, user metadata) and /// carries the full header set (SSE intent, content-type, user metadata) and
/// the completed replica preserves the source's multipart ETag. /// the completed replica preserves the source's multipart ETag.
#[tokio::test] #[tokio::test]
#[serial]
async fn test_bucket_replication_sse_s3_multipart_reencrypts() -> TestResult { async fn test_bucket_replication_sse_s3_multipart_reencrypts() -> TestResult {
init_logging(); init_logging();
@@ -4873,7 +4835,6 @@ async fn test_bucket_replication_sse_s3_multipart_reencrypts() -> TestResult {
/// still-running source's data scanner (short cycle via [`FAST_SCANNER_ENV`]) /// still-running source's data scanner (short cycle via [`FAST_SCANNER_ENV`])
/// re-drives the failed objects once the target is reachable again. /// re-drives the failed objects once the target is reachable again.
#[tokio::test] #[tokio::test]
#[serial]
async fn test_bucket_replication_recovers_after_target_outage() -> TestResult { async fn test_bucket_replication_recovers_after_target_outage() -> TestResult {
init_logging(); init_logging();
@@ -4953,7 +4914,6 @@ async fn test_bucket_replication_recovers_after_target_outage() -> TestResult {
/// must settle back to zero even though the historical failed counter remains /// must settle back to zero even though the historical failed counter remains
/// non-zero. /// non-zero.
#[tokio::test] #[tokio::test]
#[serial]
async fn test_bucket_replication_backlog_metrics_observe_outage_and_recovery() -> TestResult { async fn test_bucket_replication_backlog_metrics_observe_outage_and_recovery() -> TestResult {
init_logging(); init_logging();
@@ -5087,7 +5047,6 @@ async fn test_bucket_replication_backlog_metrics_observe_outage_and_recovery() -
/// must converge every persisted failure, including the replayed delete marker /// must converge every persisted failure, including the replayed delete marker
/// (whose replication decision is re-derived from the live config). /// (whose replication decision is re-derived from the live config).
#[tokio::test] #[tokio::test]
#[serial]
async fn test_bucket_replication_replays_failed_entries_after_source_restart() -> TestResult { async fn test_bucket_replication_replays_failed_entries_after_source_restart() -> TestResult {
init_logging(); init_logging();
@@ -5179,7 +5138,6 @@ async fn test_bucket_replication_replays_failed_entries_after_source_restart() -
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_bucket_replication_replayed_delete_marker_preserves_source_mtime_without_source_restart() -> TestResult { async fn test_bucket_replication_replayed_delete_marker_preserves_source_mtime_without_source_restart() -> TestResult {
init_logging(); init_logging();
@@ -5249,7 +5207,6 @@ async fn test_bucket_replication_replayed_delete_marker_preserves_source_mtime_w
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_sequential_bucket_replication_succeeds_for_multiple_buckets() -> Result<(), Box<dyn Error + Send + Sync>> { async fn test_sequential_bucket_replication_succeeds_for_multiple_buckets() -> Result<(), Box<dyn Error + Send + Sync>> {
init_logging(); init_logging();
@@ -5293,7 +5250,6 @@ async fn test_sequential_bucket_replication_succeeds_for_multiple_buckets() -> R
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_replication_recovers_after_runtime_target_cache_is_cleared() -> Result<(), Box<dyn Error + Send + Sync>> { async fn test_replication_recovers_after_runtime_target_cache_is_cleared() -> Result<(), Box<dyn Error + Send + Sync>> {
init_logging(); init_logging();
@@ -5337,7 +5293,6 @@ async fn test_replication_recovers_after_runtime_target_cache_is_cleared() -> Re
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_site_replication_allows_self_signed_https_with_skip_tls_verify_real_dual_node() -> TestResult { async fn test_site_replication_allows_self_signed_https_with_skip_tls_verify_real_dual_node() -> TestResult {
init_logging(); init_logging();
@@ -5416,7 +5371,6 @@ async fn test_site_replication_allows_self_signed_https_with_skip_tls_verify_rea
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_site_replication_allows_private_ca_https_with_ca_cert_pem_real_dual_node() -> TestResult { async fn test_site_replication_allows_private_ca_https_with_ca_cert_pem_real_dual_node() -> TestResult {
init_logging(); init_logging();
@@ -5495,7 +5449,6 @@ async fn test_site_replication_allows_private_ca_https_with_ca_cert_pem_real_dua
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_site_replication_resync_lifecycle_survives_real_server_restart() -> Result<(), Box<dyn Error + Send + Sync>> { async fn test_site_replication_resync_lifecycle_survives_real_server_restart() -> Result<(), Box<dyn Error + Send + Sync>> {
init_logging(); init_logging();
let resync_process_env = [ let resync_process_env = [
@@ -5715,7 +5668,6 @@ async fn test_site_replication_resync_lifecycle_survives_real_server_restart() -
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_site_replication_edit_and_status_peer_state_real_three_node() -> Result<(), Box<dyn Error + Send + Sync>> { async fn test_site_replication_edit_and_status_peer_state_real_three_node() -> Result<(), Box<dyn Error + Send + Sync>> {
init_logging(); init_logging();
@@ -5964,7 +5916,6 @@ async fn test_site_replication_edit_and_status_peer_state_real_three_node() -> R
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_site_replication_remove_all_real_dual_node() -> Result<(), Box<dyn Error + Send + Sync>> { async fn test_site_replication_remove_all_real_dual_node() -> Result<(), Box<dyn Error + Send + Sync>> {
init_logging(); init_logging();
@@ -6084,7 +6035,6 @@ async fn test_site_replication_remove_all_real_dual_node() -> Result<(), Box<dyn
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_site_replication_state_edit_fresh_and_stale_real_dual_node() -> Result<(), Box<dyn Error + Send + Sync>> { async fn test_site_replication_state_edit_fresh_and_stale_real_dual_node() -> Result<(), Box<dyn Error + Send + Sync>> {
init_logging(); init_logging();
@@ -6193,7 +6143,6 @@ async fn test_site_replication_state_edit_fresh_and_stale_real_dual_node() -> Re
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_site_replication_replicates_object_with_bucket_versioning_real_dual_node() -> TestResult { async fn test_site_replication_replicates_object_with_bucket_versioning_real_dual_node() -> TestResult {
init_logging(); init_logging();
@@ -6284,7 +6233,6 @@ async fn test_site_replication_replicates_object_with_bucket_versioning_real_dua
/// receiver was dropped with only a debug line, while `replicate status` still reported /// receiver was dropped with only a debug line, while `replicate status` still reported
/// "1/1 Buckets in sync" because both configs were byte-identical. /// "1/1 Buckets in sync" because both configs were byte-identical.
#[tokio::test] #[tokio::test]
#[serial]
async fn test_site_replication_config_broadcast_keeps_reverse_direction_real_dual_node() -> TestResult { async fn test_site_replication_config_broadcast_keeps_reverse_direction_real_dual_node() -> TestResult {
init_logging(); init_logging();
@@ -6423,7 +6371,6 @@ async fn wait_for_site_replication_rule(
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_site_replication_active_active_converges_without_loops_real_dual_node() -> TestResult { async fn test_site_replication_active_active_converges_without_loops_real_dual_node() -> TestResult {
init_logging(); init_logging();
@@ -6741,7 +6688,6 @@ async fn test_site_replication_active_active_converges_without_loops_real_dual_n
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_site_replication_replicates_policy_backed_user_access_real_dual_node() -> Result<(), Box<dyn Error + Send + Sync>> { async fn test_site_replication_replicates_policy_backed_user_access_real_dual_node() -> Result<(), Box<dyn Error + Send + Sync>> {
init_logging(); init_logging();
@@ -6829,7 +6775,6 @@ async fn test_site_replication_replicates_policy_backed_user_access_real_dual_no
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_site_replication_replicates_group_policy_backed_access_real_dual_node() -> Result<(), Box<dyn Error + Send + Sync>> async fn test_site_replication_replicates_group_policy_backed_access_real_dual_node() -> Result<(), Box<dyn Error + Send + Sync>>
{ {
init_logging(); init_logging();
@@ -6920,7 +6865,6 @@ async fn test_site_replication_replicates_group_policy_backed_access_real_dual_n
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_service_account_policy_from_accountinfo_round_trips_real_single_node() -> TestResult { async fn test_service_account_policy_from_accountinfo_round_trips_real_single_node() -> TestResult {
init_logging(); init_logging();
@@ -6972,7 +6916,6 @@ async fn test_service_account_policy_from_accountinfo_round_trips_real_single_no
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_site_replication_replicates_multiple_service_accounts_real_dual_node() -> Result<(), Box<dyn Error + Send + Sync>> { async fn test_site_replication_replicates_multiple_service_accounts_real_dual_node() -> Result<(), Box<dyn Error + Send + Sync>> {
init_logging(); init_logging();
@@ -7073,7 +7016,6 @@ async fn test_site_replication_replicates_multiple_service_accounts_real_dual_no
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_site_replication_replicates_service_accounts_created_from_sts_session_real_dual_node() -> TestResult { async fn test_site_replication_replicates_service_accounts_created_from_sts_session_real_dual_node() -> TestResult {
init_logging(); init_logging();
@@ -7214,7 +7156,6 @@ async fn wait_for_target_request_version_id(
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_bucket_resync_restart_revisits_objects_before_out_of_order_checkpoint() -> TestResult { async fn test_bucket_resync_restart_revisits_objects_before_out_of_order_checkpoint() -> TestResult {
init_logging(); init_logging();
@@ -7333,7 +7274,6 @@ async fn test_bucket_resync_restart_revisits_objects_before_out_of_order_checkpo
/// CreateMultipartUpload (the version is decided at initiate time) must both /// CreateMultipartUpload (the version is decided at initiate time) must both
/// carry the source version as `?versionId=`. /// carry the source version as `?versionId=`.
#[tokio::test] #[tokio::test]
#[serial]
async fn test_replication_put_and_create_multipart_carry_source_version_id_query() -> TestResult { async fn test_replication_put_and_create_multipart_carry_source_version_id_query() -> TestResult {
init_logging(); init_logging();
@@ -7448,7 +7388,6 @@ async fn test_replication_put_and_create_multipart_carry_source_version_id_query
/// flow to the onward bucket, proving B's outbound replication and scanner /// flow to the onward bucket, proving B's outbound replication and scanner
/// are live. /// are live.
#[tokio::test] #[tokio::test]
#[serial]
async fn test_scanner_never_cascades_inbound_replicas() -> TestResult { async fn test_scanner_never_cascades_inbound_replicas() -> TestResult {
init_logging(); init_logging();
@@ -7529,7 +7468,6 @@ async fn test_scanner_never_cascades_inbound_replicas() -> TestResult {
/// version ids and still mint its own there — the check must not report OK /// version ids and still mint its own there — the check must not report OK
/// while multipart deletes and heals would silently miss. /// while multipart deletes and heals would silently miss.
#[tokio::test] #[tokio::test]
#[serial]
async fn test_replication_check_flags_multipart_only_version_minting_target() -> TestResult { async fn test_replication_check_flags_multipart_only_version_minting_target() -> TestResult {
init_logging(); init_logging();
@@ -7602,7 +7540,6 @@ async fn test_replication_check_flags_multipart_only_version_minting_target() ->
} }
#[tokio::test] #[tokio::test]
#[serial]
async fn test_replication_check_aborts_failed_multipart_probes() -> TestResult { async fn test_replication_check_aborts_failed_multipart_probes() -> TestResult {
init_logging(); init_logging();
@@ -7765,7 +7702,6 @@ async fn test_replication_check_aborts_failed_multipart_probes() -> TestResult {
/// BucketRemoteTargetVersionMismatch — while still cleaning up the probe /// BucketRemoteTargetVersionMismatch — while still cleaning up the probe
/// object via the version id the target actually assigned. /// object via the version id the target actually assigned.
#[tokio::test] #[tokio::test]
#[serial]
async fn test_replication_check_flags_version_minting_target() -> TestResult { async fn test_replication_check_flags_version_minting_target() -> TestResult {
init_logging(); init_logging();
@@ -8006,7 +7942,6 @@ async fn wait_for_target_marker_purged(
/// the target forever. Contract under test: a failed purge attempt is retried /// the target forever. Contract under test: a failed purge attempt is retried
/// within the watch window and converges once the fault clears. /// within the watch window and converges once the fault clears.
#[tokio::test] #[tokio::test]
#[serial]
async fn test_delayed_delete_marker_purge_retries_after_transient_target_failure() -> TestResult { async fn test_delayed_delete_marker_purge_retries_after_transient_target_failure() -> TestResult {
init_logging(); init_logging();
let source_bucket = "delayed-purge-retry-src"; let source_bucket = "delayed-purge-retry-src";
@@ -8066,7 +8001,6 @@ async fn test_delayed_delete_marker_purge_retries_after_transient_target_failure
/// with an idempotent 204, which used to look like success and strand the /// with an idempotent 204, which used to look like success and strand the
/// real marker on the target forever. /// real marker on the target forever.
#[tokio::test] #[tokio::test]
#[serial]
async fn test_delayed_delete_marker_purge_uses_target_assigned_version() -> TestResult { async fn test_delayed_delete_marker_purge_uses_target_assigned_version() -> TestResult {
init_logging(); init_logging();
let source_bucket = "delayed-purge-mint-src"; let source_bucket = "delayed-purge-mint-src";
@@ -8099,7 +8033,6 @@ async fn test_delayed_delete_marker_purge_uses_target_assigned_version() -> Test
/// replayed purge succeeds, the entry must be acknowledged instead of being /// replayed purge succeeds, the entry must be acknowledged instead of being
/// retained as Missed forever. /// retained as Missed forever.
#[tokio::test] #[tokio::test]
#[serial]
async fn test_delayed_delete_marker_purge_exhaustion_persists_to_mrf_and_replays_on_restart() -> TestResult { async fn test_delayed_delete_marker_purge_exhaustion_persists_to_mrf_and_replays_on_restart() -> TestResult {
init_logging(); init_logging();
let source_bucket = "delayed-purge-mrf-src"; let source_bucket = "delayed-purge-mrf-src";
@@ -8236,7 +8169,6 @@ async fn build_scanner_compensation_pair(
/// nil-version objects entirely (`scanner_folder.rs` heal_replication), so it /// nil-version objects entirely (`scanner_folder.rs` heal_replication), so it
/// must NEVER be compensated. /// must NEVER be compensated.
#[tokio::test] #[tokio::test]
#[serial]
async fn test_scanner_compensates_existing_objects_across_write_paths() -> TestResult { async fn test_scanner_compensates_existing_objects_across_write_paths() -> TestResult {
init_logging(); init_logging();
let source_bucket = "scanner-comp-src"; let source_bucket = "scanner-comp-src";
@@ -8352,7 +8284,6 @@ async fn test_scanner_compensates_existing_objects_across_write_paths() -> TestR
/// written after the rule replicate normally (the setting only gates the /// written after the rule replicate normally (the setting only gates the
/// existing-object resync path). /// existing-object resync path).
#[tokio::test] #[tokio::test]
#[serial]
async fn test_scanner_never_compensates_when_existing_object_replication_disabled() -> TestResult { async fn test_scanner_never_compensates_when_existing_object_replication_disabled() -> TestResult {
init_logging(); init_logging();
let source_bucket = "scanner-disabled-src"; let source_bucket = "scanner-disabled-src";
+1
View File
@@ -273,6 +273,7 @@ proptest = "1"
rcgen.workspace = true rcgen.workspace = true
insta = { workspace = true, features = ["yaml", "json"] } insta = { workspace = true, features = ["yaml", "json"] }
rustfs-crypto = { workspace = true } rustfs-crypto = { workspace = true }
tonic-prost = { workspace = true }
[build-dependencies] [build-dependencies]
shadow-rs = { workspace = true, default-features = false, features = ["build", "metadata"] } shadow-rs = { workspace = true, default-features = false, features = ["build", "metadata"] }
+2 -1
View File
@@ -380,7 +380,7 @@ pub mod erasure {
pub mod event { pub mod event {
pub use crate::event::name::EventName; pub use crate::event::name::EventName;
pub use crate::services::event_notification::{EventArgs, register_event_dispatch_hook}; pub use crate::services::event_notification::{EventArgs, register_event_dispatch_hook, send_event};
} }
pub mod global { pub mod global {
@@ -483,6 +483,7 @@ pub mod store_list {
} }
pub mod storage { pub mod storage {
pub use crate::core::pools::HealLifecycleExpiryContext;
pub use crate::store::HealWalkVersion; pub use crate::store::HealWalkVersion;
pub use crate::store::{ pub use crate::store::{
ECStore, all_local_disk, all_local_disk_path, find_local_disk_by_ref, init_local_disks, ECStore, all_local_disk, all_local_disk_path, find_local_disk_by_ref, init_local_disks,
@@ -1549,8 +1549,8 @@ impl Default for PutObjectOptions {
} }
} }
#[allow(dead_code)]
impl PutObjectOptions { impl PutObjectOptions {
#[allow(dead_code, reason = "MinIO-parity surface with no caller in this port (backlog#1823)")]
fn set_match_etag(&mut self, etag: &str) { fn set_match_etag(&mut self, etag: &str) {
if etag == "*" { if etag == "*" {
self.custom_header self.custom_header
@@ -1561,6 +1561,7 @@ impl PutObjectOptions {
} }
} }
#[allow(dead_code, reason = "MinIO-parity surface with no caller in this port (backlog#1823)")]
fn set_match_etag_except(&mut self, etag: &str) { fn set_match_etag_except(&mut self, etag: &str) {
if etag == "*" { if etag == "*" {
self.custom_header self.custom_header
@@ -1696,6 +1697,7 @@ impl PutObjectOptions {
header header
} }
#[allow(dead_code, reason = "MinIO-parity surface with no caller in this port (backlog#1823)")]
fn validate(&self, _c: Arc<TargetClient>) -> Result<(), std::io::Error> { fn validate(&self, _c: Arc<TargetClient>) -> Result<(), std::io::Error> {
//if self.checksum.is_set() { //if self.checksum.is_set() {
/*if !self.trailing_header_support { /*if !self.trailing_header_support {
@@ -456,16 +456,23 @@ impl<'a> LifecycleExpiryTrace<'a> {
} }
} }
#[allow(dead_code)]
impl ExpiryStats { impl ExpiryStats {
pub fn missed_tasks(&self) -> i64 { pub fn missed_tasks(&self) -> i64 {
self.missed_expiry_tasks.load(Ordering::SeqCst) self.missed_expiry_tasks.load(Ordering::SeqCst)
} }
#[allow(
dead_code,
reason = "asserted by this file's tests; the lib target cannot see test-only consumers (backlog#1823)"
)]
fn missed_free_vers_tasks(&self) -> i64 { fn missed_free_vers_tasks(&self) -> i64 {
self.missed_freevers_tasks.load(Ordering::SeqCst) self.missed_freevers_tasks.load(Ordering::SeqCst)
} }
#[allow(
dead_code,
reason = "asserted by this file's tests; the lib target cannot see test-only consumers (backlog#1823)"
)]
fn missed_tier_journal_tasks(&self) -> i64 { fn missed_tier_journal_tasks(&self) -> i64 {
self.missed_tier_journal_tasks.load(Ordering::SeqCst) self.missed_tier_journal_tasks.load(Ordering::SeqCst)
} }
@@ -1776,7 +1783,7 @@ impl TransitionState {
.await; .await;
} }
global_metrics().record_scanner_transition_failed(1); global_metrics().record_scanner_transition_failed(1);
if !is_err_version_not_found(&err) && !is_err_object_not_found(&err) && !is_network_or_host_down(&err.to_string(), false) && !err.to_string().contains("use of closed network connection") { if !is_err_version_not_found(&err) && !is_err_object_not_found(&err) && !is_network_or_host_down(&err.to_string(), false) {
error!( error!(
event = EVENT_LIFECYCLE_TIER_OPERATION_FAILED, event = EVENT_LIFECYCLE_TIER_OPERATION_FAILED,
component = LOG_COMPONENT_ECSTORE, component = LOG_COMPONENT_ECSTORE,
+1 -1
View File
@@ -19,7 +19,7 @@ pub mod core;
pub mod evaluator; pub mod evaluator;
pub mod manual_transition_job; pub mod manual_transition_job;
mod metadata_boundary; mod metadata_boundary;
pub(crate) use metadata_boundary::get_expiry_configs; pub(crate) use metadata_boundary::{LifecycleExpiryConfigs, get_expiry_configs};
mod object_lock_boundary; mod object_lock_boundary;
pub use self::core as lifecycle; pub use self::core as lifecycle;
mod replication_sink; mod replication_sink;
@@ -80,7 +80,10 @@ impl LastDayTierStats {
} }
} }
#[allow(dead_code)] #[allow(
dead_code,
reason = "asserted by this file's tests; the lib target cannot see test-only consumers (backlog#1823)"
)]
fn merge(&self, m: LastDayTierStats) -> LastDayTierStats { fn merge(&self, m: LastDayTierStats) -> LastDayTierStats {
let mut cl = self.clone(); let mut cl = self.clone();
let mut cm = m; let mut cm = m;
@@ -177,9 +177,10 @@ fn should_record_remote_delete_failure(err: &std::io::Error) -> bool {
} }
#[derive(Default)] #[derive(Default)]
#[allow(dead_code)]
struct ObjSweeper { struct ObjSweeper {
#[allow(dead_code, reason = "written but never read back (backlog#1823)")]
object: String, object: String,
#[allow(dead_code, reason = "written but never read back (backlog#1823)")]
bucket: String, bucket: String,
version_id: Option<Uuid>, version_id: Option<Uuid>,
versioned: bool, versioned: bool,
@@ -191,9 +192,9 @@ struct ObjSweeper {
remote_object: String, remote_object: String,
} }
#[allow(dead_code)]
impl ObjSweeper { impl ObjSweeper {
#[allow(clippy::new_ret_no_self)] #[allow(clippy::new_ret_no_self)]
#[allow(dead_code, reason = "MinIO-parity surface with no caller in this port (backlog#1823)")]
pub async fn new(bucket: &str, object: &str) -> Result<Self, std::io::Error> { pub async fn new(bucket: &str, object: &str) -> Result<Self, std::io::Error> {
Ok(Self { Ok(Self {
object: object.into(), object: object.into(),
@@ -202,17 +203,20 @@ impl ObjSweeper {
}) })
} }
#[allow(dead_code, reason = "MinIO-parity surface with no caller in this port (backlog#1823)")]
pub fn with_version(&mut self, vid: Option<Uuid>) -> &Self { pub fn with_version(&mut self, vid: Option<Uuid>) -> &Self {
self.version_id = vid.clone(); self.version_id = vid.clone();
self self
} }
#[allow(dead_code, reason = "MinIO-parity surface with no caller in this port (backlog#1823)")]
pub fn with_versioning(&mut self, versioned: bool, suspended: bool) -> &Self { pub fn with_versioning(&mut self, versioned: bool, suspended: bool) -> &Self {
self.versioned = versioned; self.versioned = versioned;
self.suspended = suspended; self.suspended = suspended;
self self
} }
#[allow(dead_code, reason = "MinIO-parity surface with no caller in this port (backlog#1823)")]
pub fn get_opts(&self) -> lifecycle::ObjectOpts { pub fn get_opts(&self) -> lifecycle::ObjectOpts {
let mut opts = ObjectOpts { let mut opts = ObjectOpts {
version_id: self.version_id.clone(), version_id: self.version_id.clone(),
@@ -226,6 +230,7 @@ impl ObjSweeper {
opts opts
} }
#[allow(dead_code, reason = "MinIO-parity surface with no caller in this port (backlog#1823)")]
pub fn set_transition_state(&mut self, info: TransitionedObject) { pub fn set_transition_state(&mut self, info: TransitionedObject) {
self.transition_tier = info.tier; self.transition_tier = info.tier;
self.transition_status = info.status; self.transition_status = info.status;
@@ -266,6 +271,7 @@ impl ObjSweeper {
None None
} }
#[allow(dead_code, reason = "MinIO-parity surface with no caller in this port (backlog#1823)")]
pub async fn sweep(&self, api: Arc<ECStore>) { pub async fn sweep(&self, api: Arc<ECStore>) {
let Some(je) = self.should_remove_remote_object() else { let Some(je) = self.should_remove_remote_object() else {
return; return;
-2
View File
@@ -312,9 +312,7 @@ mod tests {
} }
#[derive(Deserialize)] #[derive(Deserialize)]
struct LegacyBucketQuota { struct LegacyBucketQuota {
#[allow(dead_code)]
quota: Option<u64>, quota: Option<u64>,
#[allow(dead_code)]
quota_type: LegacyQuotaType, quota_type: LegacyQuotaType,
} }
let legacy = serde_json::from_slice::<LegacyBucketQuota>(&json) let legacy = serde_json::from_slice::<LegacyBucketQuota>(&json)
+2 -2
View File
@@ -95,7 +95,6 @@ impl TransitionClient {
} }
#[derive(Default)] #[derive(Default)]
#[allow(dead_code)]
pub struct GetRequest { pub struct GetRequest {
pub buffer: Vec<u8>, pub buffer: Vec<u8>,
pub offset: i64, pub offset: i64,
@@ -107,11 +106,12 @@ pub struct GetRequest {
pub setting_object_info: bool, pub setting_object_info: bool,
} }
#[allow(dead_code)]
pub struct GetResponse { pub struct GetResponse {
pub size: i64, pub size: i64,
//pub error: error, //pub error: error,
#[allow(dead_code, reason = "written but never read back (backlog#1823)")]
pub did_read: bool, pub did_read: bool,
#[allow(dead_code, reason = "written but never read back (backlog#1823)")]
pub object_info: ObjectInfo, pub object_info: ObjectInfo,
} }
+2 -2
View File
@@ -20,6 +20,7 @@
#![allow(clippy::all)] #![allow(clippy::all)]
use http::{HeaderMap, HeaderName, HeaderValue}; use http::{HeaderMap, HeaderName, HeaderValue};
use rustfs_utils::http::headers::AMZ_CHECKSUM_MODE;
use std::collections::HashMap; use std::collections::HashMap;
use time::OffsetDateTime; use time::OffsetDateTime;
use tracing::warn; use tracing::warn;
@@ -27,7 +28,6 @@ use tracing::warn;
use crate::client::api_error_response::err_invalid_argument; use crate::client::api_error_response::err_invalid_argument;
#[derive(Default)] #[derive(Default)]
#[allow(dead_code)]
pub struct AdvancedGetOptions { pub struct AdvancedGetOptions {
pub replication_delete_marker: bool, pub replication_delete_marker: bool,
pub is_replication_ready_for_delete_marker: bool, pub is_replication_ready_for_delete_marker: bool,
@@ -77,7 +77,7 @@ impl GetObjectOptions {
} }
} }
if self.checksum { if self.checksum {
headers.insert(HeaderName::from_static("x-amz-checksum-mode"), HeaderValue::from_static("ENABLED")); headers.insert(HeaderName::from_static(AMZ_CHECKSUM_MODE), HeaderValue::from_static("ENABLED"));
} }
headers headers
} }
-1
View File
@@ -360,7 +360,6 @@ impl TransitionClient {
} }
#[derive(Default)] #[derive(Default)]
#[allow(dead_code)]
pub struct ListObjectsOptions { pub struct ListObjectsOptions {
reverse_versions: bool, reverse_versions: bool,
with_versions: bool, with_versions: bool,
+3 -1
View File
@@ -137,8 +137,8 @@ impl Default for PutObjectOptions {
} }
} }
#[allow(dead_code)]
impl PutObjectOptions { impl PutObjectOptions {
#[allow(dead_code, reason = "MinIO-parity surface with no caller in this port (backlog#1823)")]
fn set_match_etag(&mut self, etag: &str) { fn set_match_etag(&mut self, etag: &str) {
if etag == "*" { if etag == "*" {
self.custom_header.insert("If-Match", HeaderValue::from_static("*")); self.custom_header.insert("If-Match", HeaderValue::from_static("*"));
@@ -149,6 +149,7 @@ impl PutObjectOptions {
} }
} }
#[allow(dead_code, reason = "MinIO-parity surface with no caller in this port (backlog#1823)")]
fn set_match_etag_except(&mut self, etag: &str) { fn set_match_etag_except(&mut self, etag: &str) {
if etag == "*" { if etag == "*" {
self.custom_header.insert("If-None-Match", HeaderValue::from_static("*")); self.custom_header.insert("If-None-Match", HeaderValue::from_static("*"));
@@ -259,6 +260,7 @@ impl PutObjectOptions {
header header
} }
#[allow(dead_code, reason = "MinIO-parity surface with no caller in this port (backlog#1823)")]
fn validate(&self, c: TransitionClient) -> Result<(), std::io::Error> { fn validate(&self, c: TransitionClient) -> Result<(), std::io::Error> {
//if self.checksum.is_set() { //if self.checksum.is_set() {
/*if !self.trailing_header_support { /*if !self.trailing_header_support {
+2 -3
View File
@@ -55,7 +55,6 @@ pub struct RemoveBucketOptions {
const DELETE_RESPONSE_PREVIEW_LEN: usize = 1024; const DELETE_RESPONSE_PREVIEW_LEN: usize = 1024;
#[derive(Debug)] #[derive(Debug)]
#[allow(dead_code)]
pub struct AdvancedRemoveOptions { pub struct AdvancedRemoveOptions {
pub replication_delete_marker: bool, pub replication_delete_marker: bool,
pub replication_status: ReplicationStatus, pub replication_status: ReplicationStatus,
@@ -465,10 +464,10 @@ impl TransitionClient {
} }
#[derive(Debug, Default)] #[derive(Debug, Default)]
#[allow(dead_code)]
pub struct RemoveObjectError { pub struct RemoveObjectError {
#[allow(dead_code, reason = "written but never read back (backlog#1823)")]
object_name: String, object_name: String,
#[allow(dead_code)] #[allow(dead_code, reason = "written but never read back (backlog#1823)")]
version_id: String, version_id: String,
err: Option<std::io::Error>, err: Option<std::io::Error>,
} }
+3 -3
View File
@@ -372,8 +372,8 @@ pub struct Checksum {
computed: bool, computed: bool,
} }
#[allow(dead_code)]
impl Checksum { impl Checksum {
#[allow(dead_code, reason = "MinIO-parity surface with no caller in this port (backlog#1823)")]
fn new(t: ChecksumMode, b: &[u8]) -> Checksum { fn new(t: ChecksumMode, b: &[u8]) -> Checksum {
if t.is_set() && b.len() == t.raw_byte_len() { if t.is_set() && b.len() == t.raw_byte_len() {
return Checksum { return Checksum {
@@ -385,7 +385,7 @@ impl Checksum {
Checksum::default() Checksum::default()
} }
#[allow(dead_code)] #[allow(dead_code, reason = "MinIO-parity surface with no caller in this port (backlog#1823)")]
fn new_checksum_string(t: ChecksumMode, s: &str) -> Result<Checksum, std::io::Error> { fn new_checksum_string(t: ChecksumMode, s: &str) -> Result<Checksum, std::io::Error> {
let b = match base64_decode(s.as_bytes()) { let b = match base64_decode(s.as_bytes()) {
Ok(b) => b, Ok(b) => b,
@@ -412,7 +412,7 @@ impl Checksum {
base64_encode(&self.r) base64_encode(&self.r)
} }
#[allow(dead_code)] #[allow(dead_code, reason = "MinIO-parity surface with no caller in this port (backlog#1823)")]
fn raw(&self) -> Option<Vec<u8>> { fn raw(&self) -> Option<Vec<u8>> {
if !self.is_set() { if !self.is_set() {
return None; return None;
@@ -37,16 +37,17 @@ pub struct PutObjReader {
//pub sealMD5Fn: SealMD5CurrFn, //pub sealMD5Fn: SealMD5CurrFn,
} }
#[allow(dead_code)]
impl PutObjReader { impl PutObjReader {
pub fn new(reader: HashReader) -> Self { pub fn new(reader: HashReader) -> Self {
Self { reader } Self { reader }
} }
#[allow(dead_code, reason = "MinIO-parity surface with no caller in this port (backlog#1823)")]
fn md5_current_hex_string(&self) -> String { fn md5_current_hex_string(&self) -> String {
self.reader.checksum().map(|v| v.encoded).unwrap_or_default() self.reader.checksum().map(|v| v.encoded).unwrap_or_default()
} }
#[allow(dead_code, reason = "MinIO-parity surface with no caller in this port (backlog#1823)")]
fn with_encryption(&mut self, enc_reader: HashReader) -> Result<(), std::io::Error> { fn with_encryption(&mut self, enc_reader: HashReader) -> Result<(), std::io::Error> {
self.reader = enc_reader; self.reader = enc_reader;
+10 -6
View File
@@ -54,6 +54,10 @@ use rustfs_config::MAX_S3_CLIENT_RESPONSE_SIZE;
use rustfs_rio::HashReader; use rustfs_rio::HashReader;
use rustfs_utils::HashAlgorithm; use rustfs_utils::HashAlgorithm;
use rustfs_utils::{ use rustfs_utils::{
http::headers::{
AMZ_CHECKSUM_CRC32, AMZ_CHECKSUM_CRC32C, AMZ_CHECKSUM_CRC64NVME, AMZ_CHECKSUM_MODE, AMZ_CHECKSUM_SHA1,
AMZ_CHECKSUM_SHA256,
},
net::get_endpoint_url, net::get_endpoint_url,
retry::{DEFAULT_RETRY_CAP, DEFAULT_RETRY_UNIT, MAX_JITTER, MAX_RETRY, RetryTimer}, retry::{DEFAULT_RETRY_CAP, DEFAULT_RETRY_UNIT, MAX_JITTER, MAX_RETRY, RetryTimer},
}; };
@@ -1383,12 +1387,12 @@ pub(crate) fn to_object_info_for_provider(
}; };
// Extract checksums // Extract checksums
let checksum_crc32 = get_header("x-amz-checksum-crc32"); let checksum_crc32 = get_header(AMZ_CHECKSUM_CRC32);
let checksum_crc32c = get_header("x-amz-checksum-crc32c"); let checksum_crc32c = get_header(AMZ_CHECKSUM_CRC32C);
let checksum_sha1 = get_header("x-amz-checksum-sha1"); let checksum_sha1 = get_header(AMZ_CHECKSUM_SHA1);
let checksum_sha256 = get_header("x-amz-checksum-sha256"); let checksum_sha256 = get_header(AMZ_CHECKSUM_SHA256);
let checksum_crc64nvme = get_header("x-amz-checksum-crc64nvme"); let checksum_crc64nvme = get_header(AMZ_CHECKSUM_CRC64NVME);
let checksum_mode = get_header("x-amz-checksum-mode"); let checksum_mode = get_header(AMZ_CHECKSUM_MODE);
// Build and return the ObjectInfo struct // Build and return the ObjectInfo struct
Ok(ObjectInfo { Ok(ObjectInfo {
@@ -233,11 +233,17 @@ pub struct NsScannerCapabilityRequest {
#[async_trait] #[async_trait]
pub trait InternodeDataTransport: Send + Sync + std::fmt::Debug { pub trait InternodeDataTransport: Send + Sync + std::fmt::Debug {
async fn open_read(&self, request: ReadStreamRequest) -> Result<FileReader>; async fn open_read(&self, request: ReadStreamRequest) -> Result<FileReader>;
async fn open_read_fresh(&self, request: ReadStreamRequest) -> Result<FileReader> {
self.open_read(request).await
}
/// Opens an owned-chunk stream when this transport can retain receive-buffer /// Opens an owned-chunk stream when this transport can retain receive-buffer
/// ownership. `None` preserves the established `open_read` fallback. /// ownership. `None` preserves the established `open_read` fallback.
async fn open_read_chunks(&self, _request: ReadStreamRequest) -> Result<Option<ChunkReaderBox>> { async fn open_read_chunks(&self, _request: ReadStreamRequest) -> Result<Option<ChunkReaderBox>> {
Ok(None) Ok(None)
} }
async fn open_read_chunks_fresh(&self, request: ReadStreamRequest) -> Result<Option<ChunkReaderBox>> {
self.open_read_chunks(request).await
}
async fn open_write(&self, request: WriteStreamRequest) -> Result<FileWriter>; async fn open_write(&self, request: WriteStreamRequest) -> Result<FileWriter>;
async fn open_walk_dir(&self, request: WalkDirStreamRequest) -> Result<FileReader>; async fn open_walk_dir(&self, request: WalkDirStreamRequest) -> Result<FileReader>;
async fn open_ns_scanner(&self, _request: NsScannerStreamRequest) -> Result<FileReader> { async fn open_ns_scanner(&self, _request: NsScannerStreamRequest) -> Result<FileReader> {
@@ -269,6 +275,15 @@ impl InternodeDataTransport for TcpHttpInternodeDataTransport {
)) ))
} }
async fn open_read_fresh(&self, request: ReadStreamRequest) -> Result<FileReader> {
let url = build_read_file_stream_url(&request);
let mut headers = json_headers();
build_auth_headers(&url, &Method::GET, &mut headers)?;
Ok(Box::new(
HttpReader::new_fresh_connection_with_stall_timeout(url, Method::GET, headers, None, request.stall_timeout).await?,
))
}
async fn open_read_chunks(&self, request: ReadStreamRequest) -> Result<Option<ChunkReaderBox>> { async fn open_read_chunks(&self, request: ReadStreamRequest) -> Result<Option<ChunkReaderBox>> {
let url = build_read_file_stream_url(&request); let url = build_read_file_stream_url(&request);
let mut headers = json_headers(); let mut headers = json_headers();
@@ -278,6 +293,16 @@ impl InternodeDataTransport for TcpHttpInternodeDataTransport {
))) )))
} }
async fn open_read_chunks_fresh(&self, request: ReadStreamRequest) -> Result<Option<ChunkReaderBox>> {
let url = build_read_file_stream_url(&request);
let mut headers = json_headers();
build_auth_headers(&url, &Method::GET, &mut headers)?;
Ok(Some(Box::new(
HttpChunkReader::new_fresh_connection_with_stall_timeout(url, Method::GET, headers, None, request.stall_timeout)
.await?,
)))
}
async fn open_write(&self, request: WriteStreamRequest) -> Result<FileWriter> { async fn open_write(&self, request: WriteStreamRequest) -> Result<FileWriter> {
let server_epoch = self.put_file_auth_capability(&request.endpoint).await?; let server_epoch = self.put_file_auth_capability(&request.endpoint).await?;
let nonce = server_epoch.map(|_| Uuid::new_v4()); let nonce = server_epoch.map(|_| Uuid::new_v4());
@@ -86,6 +86,25 @@ const PEER_REST_RECOVERY_MAX_BACKOFF: Duration = Duration::from_secs(30);
const SCANNER_ACTIVITY_MAX_MESSAGE_SIZE: usize = 1024; const SCANNER_ACTIVITY_MAX_MESSAGE_SIZE: usize = 1024;
const REPLICATION_STATS_MAX_MESSAGE_SIZE: usize = 8 * 1024 * 1024; const REPLICATION_STATS_MAX_MESSAGE_SIZE: usize = 8 * 1024 * 1024;
/// Error for a peer that reported `success = false` without an `error_info` payload.
///
/// Same shape as `peer_s3_client::peer_failure_without_details`, over `StorageError`
/// instead of `DiskError`. The message names the operation (and the bucket, where the
/// operation has one) and nothing else, for two reasons:
///
/// - `finalize_result` classifies failures by message substring, so any text matching
/// `message_has_network_needle` would take an answering peer offline and evict its
/// connection over a plain application-level rejection.
/// - Quorum aggregation (`reduce_errs`) buckets `Io` errors by kind plus rendered
/// message, so a per-peer detail such as the peer address would split one shared
/// failure into single-count buckets and downgrade the dominant error.
fn peer_failure_without_details(op: &str, bucket: Option<&str>) -> Error {
match bucket {
Some(bucket) => Error::other(format!("{op}({bucket}): peer returned failure without error details")),
None => Error::other(format!("{op}: peer returned failure without error details")),
}
}
fn decode_bucket_stats_response(response: GetBucketStatsDataResponse) -> Result<BucketStats> { fn decode_bucket_stats_response(response: GetBucketStatsDataResponse) -> Result<BucketStats> {
if !response.success { if !response.success {
return Err(Error::other( return Err(Error::other(
@@ -696,7 +715,7 @@ impl PeerRestClient {
if let Some(msg) = response.error_info { if let Some(msg) = response.error_info {
return Err(Error::other(msg)); return Err(Error::other(msg));
} }
return Err(Error::other("")); return Err(peer_failure_without_details("local_storage_info", None));
} }
let data = response.storage_info; let data = response.storage_info;
@@ -719,7 +738,7 @@ impl PeerRestClient {
if let Some(msg) = response.error_info { if let Some(msg) = response.error_info {
return Err(Error::other(msg)); return Err(Error::other(msg));
} }
return Err(Error::other("")); return Err(peer_failure_without_details("server_info", None));
} }
let data = response.server_properties; let data = response.server_properties;
@@ -742,7 +761,7 @@ impl PeerRestClient {
if let Some(msg) = response.error_info { if let Some(msg) = response.error_info {
return Err(Error::other(msg)); return Err(Error::other(msg));
} }
return Err(Error::other("")); return Err(peer_failure_without_details("get_cpus", None));
} }
let data = response.cpus; let data = response.cpus;
@@ -765,7 +784,7 @@ impl PeerRestClient {
if let Some(msg) = response.error_info { if let Some(msg) = response.error_info {
return Err(Error::other(msg)); return Err(Error::other(msg));
} }
return Err(Error::other("")); return Err(peer_failure_without_details("get_net_info", None));
} }
let data = response.net_info; let data = response.net_info;
@@ -788,7 +807,7 @@ impl PeerRestClient {
if let Some(msg) = response.error_info { if let Some(msg) = response.error_info {
return Err(Error::other(msg)); return Err(Error::other(msg));
} }
return Err(Error::other("")); return Err(peer_failure_without_details("get_partitions", None));
} }
let data = response.partitions; let data = response.partitions;
@@ -811,7 +830,7 @@ impl PeerRestClient {
if let Some(msg) = response.error_info { if let Some(msg) = response.error_info {
return Err(Error::other(msg)); return Err(Error::other(msg));
} }
return Err(Error::other("")); return Err(peer_failure_without_details("get_os_info", None));
} }
let data = response.os_info; let data = response.os_info;
@@ -832,7 +851,7 @@ impl PeerRestClient {
if let Some(msg) = response.error_info { if let Some(msg) = response.error_info {
return Err(Error::other(msg)); return Err(Error::other(msg));
} }
return Err(Error::other("")); return Err(peer_failure_without_details("get_se_linux_info", None));
} }
let data = response.sys_services; let data = response.sys_services;
@@ -857,7 +876,7 @@ impl PeerRestClient {
if let Some(msg) = response.error_info { if let Some(msg) = response.error_info {
return Err(Error::other(msg)); return Err(Error::other(msg));
} }
return Err(Error::other("")); return Err(peer_failure_without_details("get_sys_config", None));
} }
let data = response.sys_config; let data = response.sys_config;
@@ -882,7 +901,7 @@ impl PeerRestClient {
if let Some(msg) = response.error_info { if let Some(msg) = response.error_info {
return Err(Error::other(msg)); return Err(Error::other(msg));
} }
return Err(Error::other("")); return Err(peer_failure_without_details("get_sys_errors", None));
} }
let data = response.sys_errors; let data = response.sys_errors;
@@ -907,7 +926,7 @@ impl PeerRestClient {
if let Some(msg) = response.error_info { if let Some(msg) = response.error_info {
return Err(Error::other(msg)); return Err(Error::other(msg));
} }
return Err(Error::other("")); return Err(peer_failure_without_details("get_mem_info", None));
} }
let data = response.mem_info; let data = response.mem_info;
@@ -939,7 +958,7 @@ impl PeerRestClient {
if let Some(msg) = response.error_info { if let Some(msg) = response.error_info {
return Err(Error::other(msg)); return Err(Error::other(msg));
} }
return Err(Error::other("")); return Err(peer_failure_without_details("get_metrics", None));
} }
let data = response.realtime_metrics; let data = response.realtime_metrics;
@@ -964,7 +983,7 @@ impl PeerRestClient {
if let Some(msg) = response.error_info { if let Some(msg) = response.error_info {
return Err(Error::other(msg)); return Err(Error::other(msg));
} }
return Err(Error::other("")); return Err(peer_failure_without_details("get_live_events", None));
} }
Ok(PeerLiveEventsBatch { Ok(PeerLiveEventsBatch {
@@ -989,7 +1008,7 @@ impl PeerRestClient {
if let Some(msg) = response.error_info { if let Some(msg) = response.error_info {
return Err(Error::other(msg)); return Err(Error::other(msg));
} }
return Err(Error::other("")); return Err(peer_failure_without_details("get_proc_info", None));
} }
let data = response.proc_info; let data = response.proc_info;
@@ -1016,7 +1035,7 @@ impl PeerRestClient {
if let Some(msg) = response.error_info { if let Some(msg) = response.error_info {
return Err(Error::other(msg)); return Err(Error::other(msg));
} }
return Err(Error::other("")); return Err(peer_failure_without_details("start_profiling", None));
} }
Ok(()) Ok(())
} }
@@ -1323,7 +1342,7 @@ impl PeerRestClient {
if let Some(msg) = response.error_info { if let Some(msg) = response.error_info {
return Err(Error::other(msg)); return Err(Error::other(msg));
} }
return Err(Error::other("")); return Err(peer_failure_without_details("load_bucket_metadata", Some(bucket)));
} }
Ok(()) Ok(())
} }
@@ -1346,7 +1365,7 @@ impl PeerRestClient {
if let Some(msg) = response.error_info { if let Some(msg) = response.error_info {
return Err(Error::other(msg)); return Err(Error::other(msg));
} }
return Err(Error::other("")); return Err(peer_failure_without_details("delete_bucket_metadata", Some(bucket)));
} }
Ok(()) Ok(())
} }
@@ -1369,7 +1388,7 @@ impl PeerRestClient {
if let Some(msg) = response.error_info { if let Some(msg) = response.error_info {
return Err(Error::other(msg)); return Err(Error::other(msg));
} }
return Err(Error::other("")); return Err(peer_failure_without_details("delete_policy", None));
} }
Ok(()) Ok(())
} }
@@ -1392,7 +1411,7 @@ impl PeerRestClient {
if let Some(msg) = response.error_info { if let Some(msg) = response.error_info {
return Err(Error::other(msg)); return Err(Error::other(msg));
} }
return Err(Error::other("")); return Err(peer_failure_without_details("load_policy", None));
} }
Ok(()) Ok(())
} }
@@ -1417,7 +1436,7 @@ impl PeerRestClient {
if let Some(msg) = response.error_info { if let Some(msg) = response.error_info {
return Err(Error::other(msg)); return Err(Error::other(msg));
} }
return Err(Error::other("")); return Err(peer_failure_without_details("load_policy_mapping", None));
} }
Ok(()) Ok(())
} }
@@ -1440,7 +1459,7 @@ impl PeerRestClient {
if let Some(msg) = response.error_info { if let Some(msg) = response.error_info {
return Err(Error::other(msg)); return Err(Error::other(msg));
} }
return Err(Error::other("")); return Err(peer_failure_without_details("delete_user", None));
} }
Ok(()) Ok(())
} }
@@ -1463,7 +1482,7 @@ impl PeerRestClient {
if let Some(msg) = response.error_info { if let Some(msg) = response.error_info {
return Err(Error::other(msg)); return Err(Error::other(msg));
} }
return Err(Error::other("")); return Err(peer_failure_without_details("delete_service_account", None));
} }
Ok(()) Ok(())
} }
@@ -1487,7 +1506,7 @@ impl PeerRestClient {
if let Some(msg) = response.error_info { if let Some(msg) = response.error_info {
return Err(Error::other(msg)); return Err(Error::other(msg));
} }
return Err(Error::other("")); return Err(peer_failure_without_details("load_user", None));
} }
Ok(()) Ok(())
} }
@@ -1510,7 +1529,7 @@ impl PeerRestClient {
if let Some(msg) = response.error_info { if let Some(msg) = response.error_info {
return Err(Error::other(msg)); return Err(Error::other(msg));
} }
return Err(Error::other("")); return Err(peer_failure_without_details("load_service_account", None));
} }
Ok(()) Ok(())
} }
@@ -1533,7 +1552,7 @@ impl PeerRestClient {
if let Some(msg) = response.error_info { if let Some(msg) = response.error_info {
return Err(Error::other(msg)); return Err(Error::other(msg));
} }
return Err(Error::other("")); return Err(peer_failure_without_details("load_group", None));
} }
Ok(()) Ok(())
} }
@@ -1554,7 +1573,7 @@ impl PeerRestClient {
if let Some(msg) = response.error_info { if let Some(msg) = response.error_info {
return Err(Error::other(msg)); return Err(Error::other(msg));
} }
return Err(Error::other("")); return Err(peer_failure_without_details("reload_site_replication_config", None));
} }
Ok(()) Ok(())
} }
@@ -1597,7 +1616,7 @@ impl PeerRestClient {
if let Some(msg) = response.error_info { if let Some(msg) = response.error_info {
return Err(Error::other(msg)); return Err(Error::other(msg));
} }
return Err(Error::other("")); return Err(peer_failure_without_details("signal_service", None));
} }
validate_signal_service_protocol(sig, sub_sys, response.protocol_version)?; validate_signal_service_protocol(sig, sub_sys, response.protocol_version)?;
Ok(response) Ok(response)
@@ -1667,7 +1686,7 @@ impl PeerRestClient {
if let Some(msg) = response.error_info { if let Some(msg) = response.error_info {
return Err(Error::other(msg)); return Err(Error::other(msg));
} }
return Err(Error::other("")); return Err(peer_failure_without_details("reload_pool_meta", None));
} }
Ok(()) Ok(())
@@ -1691,7 +1710,7 @@ impl PeerRestClient {
if let Some(msg) = response.error_info { if let Some(msg) = response.error_info {
return Err(Error::other(msg)); return Err(Error::other(msg));
} }
return Err(Error::other("")); return Err(peer_failure_without_details("stop_rebalance", None));
} }
Ok(()) Ok(())
@@ -1725,7 +1744,7 @@ impl PeerRestClient {
if let Some(msg) = response.error_info { if let Some(msg) = response.error_info {
return Err(Error::other(msg)); return Err(Error::other(msg));
} }
return Err(Error::other("")); return Err(peer_failure_without_details("load_rebalance_meta", None));
} }
Ok(()) Ok(())
@@ -1753,7 +1772,7 @@ impl PeerRestClient {
if let Some(msg) = response.error_info { if let Some(msg) = response.error_info {
return Err(Error::other(msg)); return Err(Error::other(msg));
} }
return Err(Error::other("")); return Err(peer_failure_without_details("start_decommission", None));
} }
Ok(()) Ok(())
@@ -1777,7 +1796,7 @@ impl PeerRestClient {
if let Some(msg) = response.error_info { if let Some(msg) = response.error_info {
return Err(Error::other(msg)); return Err(Error::other(msg));
} }
return Err(Error::other("")); return Err(peer_failure_without_details("decommission_cancel", None));
} }
Ok(()) Ok(())
@@ -1801,7 +1820,7 @@ impl PeerRestClient {
if let Some(msg) = response.error_info { if let Some(msg) = response.error_info {
return Err(Error::other(msg)); return Err(Error::other(msg));
} }
return Err(Error::other("")); return Err(peer_failure_without_details("clear_decommission", None));
} }
Ok(()) Ok(())
@@ -1947,6 +1966,8 @@ fn tier_config_reload_status_outcome(status: tonic::Status) -> TierConfigReloadO
mod tests { mod tests {
use super::*; use super::*;
use crate::config::com::STORAGE_CLASS_SUB_SYS; use crate::config::com::STORAGE_CLASS_SUB_SYS;
use crate::disk::error::DiskError;
use crate::disk::error_reduce::reduce_errs;
use crate::layout::{disks_layout::DisksLayout, endpoints::SetupType}; use crate::layout::{disks_layout::DisksLayout, endpoints::SetupType};
use rustfs_config::{ENV_KUBERNETES_SERVICE_HOST, ENV_LOCAL_ENDPOINT_HOST, ENV_STARTUP_TOPOLOGY_WAIT_MODE}; use rustfs_config::{ENV_KUBERNETES_SERVICE_HOST, ENV_LOCAL_ENDPOINT_HOST, ENV_STARTUP_TOPOLOGY_WAIT_MODE};
use serde_json::Value; use serde_json::Value;
@@ -3098,4 +3119,115 @@ mod tests {
&& span.get("request_id").and_then(Value::as_str) == Some("req-peer-rest") && span.get("request_id").and_then(Value::as_str) == Some("req-peer-rest")
})); }));
} }
/// Every operation name passed to `peer_failure_without_details` in this file.
const PEER_FAILURE_OPS: &[&str] = &[
"local_storage_info",
"server_info",
"get_cpus",
"get_net_info",
"get_partitions",
"get_os_info",
"get_se_linux_info",
"get_sys_config",
"get_sys_errors",
"get_mem_info",
"get_metrics",
"get_live_events",
"get_proc_info",
"start_profiling",
"load_bucket_metadata",
"delete_bucket_metadata",
"delete_policy",
"load_policy",
"load_policy_mapping",
"delete_user",
"delete_service_account",
"load_user",
"load_service_account",
"load_group",
"reload_site_replication_config",
"signal_service",
"reload_pool_meta",
"stop_rebalance",
"load_rebalance_meta",
"start_decommission",
"decommission_cancel",
"clear_decommission",
];
#[test]
fn peer_failure_without_details_names_operation_and_bucket() {
for op in PEER_FAILURE_OPS {
let message = peer_failure_without_details(op, None).to_string();
assert!(message.contains(op), "{op} message must name the operation: {message}");
}
for op in ["load_bucket_metadata", "delete_bucket_metadata"] {
let message = peer_failure_without_details(op, Some("ops-bucket")).to_string();
assert!(message.contains(op), "{op} message must name the operation: {message}");
assert!(message.contains("ops-bucket"), "{op} message must name the bucket: {message}");
}
}
#[test]
fn peer_failure_without_details_keeps_one_reduce_errs_bucket_per_operation() {
// reduce_errs groups Io errors by kind plus rendered message: peers failing the
// same operation must stay a single dominant error instead of one bucket per peer.
let per_peer_errs = (0..4)
.map(|_| Some(DiskError::from(peer_failure_without_details("load_bucket_metadata", Some("shared")))))
.collect::<Vec<_>>();
let (count, dominant) = reduce_errs(&per_peer_errs, &[]);
assert_eq!(count, 4, "one shared failure must not split into per-peer buckets");
assert_eq!(
dominant,
Some(DiskError::from(peer_failure_without_details("load_bucket_metadata", Some("shared"))))
);
assert_ne!(
peer_failure_without_details("load_bucket_metadata", Some("shared")).to_string(),
peer_failure_without_details("delete_bucket_metadata", Some("shared")).to_string()
);
assert_ne!(
peer_failure_without_details("load_bucket_metadata", Some("bucket-a")).to_string(),
peer_failure_without_details("load_bucket_metadata", Some("bucket-b")).to_string()
);
}
#[test]
fn peer_failure_without_details_never_reads_as_a_network_failure() {
// `finalize_result` marks the peer offline and evicts its connection whenever the
// message matches a network needle. A peer that answered `success = false` is alive,
// so no operation or bucket name may push this text over that classifier.
for op in PEER_FAILURE_OPS {
let err = peer_failure_without_details(op, None);
assert!(
!PeerRestClient::is_network_like_error(&err),
"{op} must not read as a transport failure: {err}"
);
let scoped = peer_failure_without_details(op, Some("bucket-name"));
assert!(
!PeerRestClient::is_network_like_error(&scoped),
"{op} must not read as a transport failure: {scoped}"
);
}
// The bucket name is caller-supplied. Every needle carries a space, which S3 bucket
// names cannot, and the name is closed by `)` before the literal text resumes, so no
// needle can straddle the boundary either.
for bucket in [
"timed-out",
"connection-reset",
"transport-error",
"broken-pipe",
"unavailable-logs",
] {
let err = peer_failure_without_details("load_bucket_metadata", Some(bucket));
assert!(
!PeerRestClient::is_network_like_error(&err),
"bucket {bucket} must not push the message over the network classifier: {err}"
);
}
}
} }
@@ -214,6 +214,21 @@ fn pool_write_quorum(participant_count: usize) -> usize {
(participant_count / 2) + 1 (participant_count / 2) + 1
} }
/// Error for a peer that reported `success = false` without an error payload.
///
/// The message must stay identical across the peers of one operation: `reduce_errs`
/// buckets `Error::Io` by kind plus rendered message, so any per-peer detail (address,
/// timing) would split one shared failure into single-count buckets and downgrade a real
/// dominant error into `ErasureWriteQuorum`.
///
/// `peer_rest_client` carries the same helper over `StorageError` for the same response shape.
fn peer_failure_without_details(op: &str, bucket: Option<&str>) -> Error {
match bucket {
Some(bucket) => Error::other(format!("{op}({bucket}): peer returned failure without error details")),
None => Error::other(format!("{op}: peer returned failure without error details")),
}
}
fn reduce_pool_write_quorum_errs(per_pool_errs: &[Option<Error>]) -> Option<Error> { fn reduce_pool_write_quorum_errs(per_pool_errs: &[Option<Error>]) -> Option<Error> {
if per_pool_errs.is_empty() { if per_pool_errs.is_empty() {
return Some(Error::ErasureWriteQuorum); return Some(Error::ErasureWriteQuorum);
@@ -1078,7 +1093,7 @@ impl PeerS3Client for RemotePeerS3Client {
return if let Some(err) = response.error { return if let Some(err) = response.error {
Err(err.into()) Err(err.into())
} else { } else {
Err(Error::other("")) Err(peer_failure_without_details("heal_bucket", Some(bucket)))
}; };
} }
@@ -1105,7 +1120,7 @@ impl PeerS3Client for RemotePeerS3Client {
return if let Some(err) = response.error { return if let Some(err) = response.error {
Err(err.into()) Err(err.into())
} else { } else {
Err(Error::other("")) Err(peer_failure_without_details("list_bucket", None))
}; };
} }
let bucket_infos = response let bucket_infos = response
@@ -1136,9 +1151,7 @@ impl PeerS3Client for RemotePeerS3Client {
return if let Some(err) = response.error { return if let Some(err) = response.error {
Err(err.into()) Err(err.into())
} else { } else {
Err(Error::other(format!( Err(peer_failure_without_details("make_bucket", Some(bucket)))
"make_bucket({bucket}): peer returned failure without error details"
)))
}; };
} }
@@ -1162,7 +1175,7 @@ impl PeerS3Client for RemotePeerS3Client {
return if let Some(err) = response.error { return if let Some(err) = response.error {
Err(err.into()) Err(err.into())
} else { } else {
Err(Error::other("")) Err(peer_failure_without_details("get_bucket_info", Some(bucket)))
}; };
} }
let bucket_info = serde_json::from_str::<BucketInfo>(&response.bucket_info)?; let bucket_info = serde_json::from_str::<BucketInfo>(&response.bucket_info)?;
@@ -1190,7 +1203,7 @@ impl PeerS3Client for RemotePeerS3Client {
return if let Some(err) = response.error { return if let Some(err) = response.error {
Err(err.into()) Err(err.into())
} else { } else {
Err(Error::other("")) Err(peer_failure_without_details("delete_bucket", Some(bucket)))
}; };
} }
@@ -2314,4 +2327,37 @@ mod tests {
.collect::<Vec<_>>(); .collect::<Vec<_>>();
assert_eq!(calls, vec![1, 1, 0, 0, 0, 0, 0, 0]); assert_eq!(calls, vec![1, 1, 0, 0, 0, 0, 0, 0]);
} }
#[test]
fn peer_failure_without_details_names_operation_and_bucket() {
for op in ["heal_bucket", "make_bucket", "get_bucket_info", "delete_bucket"] {
let message = peer_failure_without_details(op, Some("ops-bucket")).to_string();
assert!(message.contains(op), "{op} message must name the operation: {message}");
assert!(message.contains("ops-bucket"), "{op} message must name the bucket: {message}");
}
let message = peer_failure_without_details("list_bucket", None).to_string();
assert!(message.contains("list_bucket"), "cluster-wide message must name the operation");
assert!(!message.trim().is_empty());
}
#[test]
fn peer_failure_without_details_keeps_one_reduce_errs_bucket_per_operation() {
// reduce_errs groups Io errors by kind plus rendered message: peers failing the
// same operation on the same bucket must still reach quorum as one dominant error.
let per_pool_errs = vec![
Some(peer_failure_without_details("delete_bucket", Some("shared"))),
Some(peer_failure_without_details("delete_bucket", Some("shared"))),
Some(peer_failure_without_details("delete_bucket", Some("shared"))),
];
assert_eq!(
reduce_pool_write_quorum_errs(&per_pool_errs),
Some(peer_failure_without_details("delete_bucket", Some("shared")))
);
assert_ne!(
peer_failure_without_details("delete_bucket", Some("shared")),
peer_failure_without_details("get_bucket_info", Some("shared"))
);
}
} }
File diff suppressed because it is too large Load Diff
-3
View File
@@ -39,7 +39,6 @@ use rustfs_config::{
}; };
use std::sync::LazyLock; use std::sync::LazyLock;
#[allow(dead_code)]
#[allow(clippy::declare_interior_mutable_const)] #[allow(clippy::declare_interior_mutable_const)]
/// Default KVS for audit webhook settings. /// Default KVS for audit webhook settings.
pub static DEFAULT_AUDIT_WEBHOOK_KVS: LazyLock<KVS> = LazyLock::new(|| { pub static DEFAULT_AUDIT_WEBHOOK_KVS: LazyLock<KVS> = LazyLock::new(|| {
@@ -117,7 +116,6 @@ pub static DEFAULT_AUDIT_WEBHOOK_KVS: LazyLock<KVS> = LazyLock::new(|| {
]) ])
}); });
#[allow(dead_code)]
#[allow(clippy::declare_interior_mutable_const)] #[allow(clippy::declare_interior_mutable_const)]
/// Default KVS for audit MQTT settings. /// Default KVS for audit MQTT settings.
pub static DEFAULT_AUDIT_MQTT_KVS: LazyLock<KVS> = LazyLock::new(|| { pub static DEFAULT_AUDIT_MQTT_KVS: LazyLock<KVS> = LazyLock::new(|| {
@@ -375,7 +373,6 @@ pub static DEFAULT_AUDIT_NATS_KVS: LazyLock<KVS> = LazyLock::new(|| {
]) ])
}); });
#[allow(dead_code)]
pub static DEFAULT_AUDIT_PULSAR_KVS: LazyLock<KVS> = LazyLock::new(|| { pub static DEFAULT_AUDIT_PULSAR_KVS: LazyLock<KVS> = LazyLock::new(|| {
KVS(vec![ KVS(vec![
KV { KV {
-59
View File
@@ -12,12 +12,9 @@
// See the License for the specific language governing permissions and // See the License for the specific language governing permissions and
// limitations under the License. // limitations under the License.
use crate::error::{Error, Result};
use rustfs_config::server_config::{KV, KVS}; use rustfs_config::server_config::{KV, KVS};
use rustfs_config::{DEFAULT_HEAL_BITROT_CYCLE_SECS, HEAL_BITROT_CYCLE}; use rustfs_config::{DEFAULT_HEAL_BITROT_CYCLE_SECS, HEAL_BITROT_CYCLE};
use rustfs_utils::string::parse_bool;
use std::sync::LazyLock; use std::sync::LazyLock;
use std::time::Duration;
pub static DEFAULT_KVS: LazyLock<KVS> = LazyLock::new(|| { pub static DEFAULT_KVS: LazyLock<KVS> = LazyLock::new(|| {
KVS(vec![KV { KVS(vec![KV {
@@ -26,59 +23,3 @@ pub static DEFAULT_KVS: LazyLock<KVS> = LazyLock::new(|| {
hidden_if_empty: false, hidden_if_empty: false,
}]) }])
}); });
#[derive(Debug, Default)]
pub struct Config {
pub bitrot: String,
pub sleep: Duration,
pub io_count: usize,
pub drive_workers: usize,
pub cache: Duration,
}
impl Config {
pub fn bitrot_scan_cycle(&self) -> Duration {
self.cache
}
pub fn get_workers(&self) -> usize {
self.drive_workers
}
pub fn update(&mut self, nopts: &Config) {
self.bitrot = nopts.bitrot.clone();
self.io_count = nopts.io_count;
self.sleep = nopts.sleep;
self.drive_workers = nopts.drive_workers;
}
}
const RUSTFS_BITROT_CYCLE_IN_MONTHS: u64 = 1;
fn parse_bitrot_config(s: &str) -> Result<Duration> {
match parse_bool(s) {
Ok(enabled) => {
if enabled {
Ok(Duration::from_secs_f64(0.0))
} else {
Ok(Duration::from_secs_f64(-1.0))
}
}
Err(_) => {
if !s.ends_with("m") {
return Err(Error::other("unknown format"));
}
match s.trim_end_matches('m').parse::<u64>() {
Ok(months) => {
if months < RUSTFS_BITROT_CYCLE_IN_MONTHS {
return Err(Error::other(format!("minimum bitrot cycle is {RUSTFS_BITROT_CYCLE_IN_MONTHS} month(s)")));
}
Ok(Duration::from_secs(months * 30 * 24 * 60))
}
Err(err) => Err(Error::other(err)),
}
}
}
}
-1
View File
@@ -16,7 +16,6 @@
mod audit; mod audit;
pub mod com; pub mod com;
#[allow(dead_code)]
pub mod heal; pub mod heal;
mod notify; mod notify;
mod oidc; mod oidc;
+100 -3
View File
@@ -16,6 +16,7 @@ use crate::bucket::replication::replication_state_from_filemeta;
use crate::bucket::versioning_sys::BucketVersioningSys; use crate::bucket::versioning_sys::BucketVersioningSys;
use crate::bucket::{ use crate::bucket::{
lifecycle::{ lifecycle::{
LifecycleExpiryConfigs,
bucket_lifecycle_audit::LcEventSrc, bucket_lifecycle_audit::LcEventSrc,
bucket_lifecycle_ops::{ bucket_lifecycle_ops::{
LifecycleOps, apply_expiry_on_transitioned_object, apply_expiry_rule_in, eval_action_from_lifecycle, LifecycleOps, apply_expiry_on_transitioned_object, apply_expiry_rule_in, eval_action_from_lifecycle,
@@ -1996,11 +1997,11 @@ impl PoolMeta {
Ok(false) Ok(false)
} }
#[allow(dead_code)]
pub fn validate(&self, pools: Vec<Arc<Sets>>) -> Result<bool> { pub fn validate(&self, pools: Vec<Arc<Sets>>) -> Result<bool> {
struct PoolInfo { struct PoolInfo {
position: usize, position: usize,
completed: bool, completed: bool,
#[allow(dead_code, reason = "written but never read back (backlog#1823)")]
decom_started: bool, decom_started: bool,
} }
@@ -2335,6 +2336,10 @@ fn lifecycle_action_removes_data_movement_version(action: IlmAction) -> bool {
) )
} }
fn lifecycle_action_skips_heal_version(action: IlmAction) -> bool {
action.delete()
}
fn resolve_data_movement_lifecycle_expiry_result(action: IlmAction, apply_actions: bool, applied: bool) -> Result<bool> { fn resolve_data_movement_lifecycle_expiry_result(action: IlmAction, apply_actions: bool, applied: bool) -> Result<bool> {
if !apply_actions || applied { if !apply_actions || applied {
return Ok(true); return Ok(true);
@@ -2385,7 +2390,80 @@ pub(crate) async fn should_skip_lifecycle_for_data_movement(
} }
} }
pub struct HealLifecycleExpiryContext {
configs: LifecycleExpiryConfigs,
}
impl ECStore { impl ECStore {
pub async fn load_heal_lifecycle_expiry_context(&self, bucket: &str) -> Result<Option<HealLifecycleExpiryContext>> {
if bucket == RUSTFS_META_BUCKET {
return Ok(None);
}
let configs = get_expiry_configs(self, bucket).await?;
if configs.lifecycle.is_none() {
return Ok(None);
}
Ok(Some(HealLifecycleExpiryContext { configs }))
}
pub async fn enqueue_heal_lifecycle_expiry(
self: &Arc<Self>,
context: &HealLifecycleExpiryContext,
bucket: &str,
object: &str,
version_id: Option<&str>,
object_info: Option<&crate::object_api::ObjectInfo>,
) -> Result<bool> {
let Some(lifecycle_config) = context.configs.lifecycle.as_ref() else {
return Ok(false);
};
let object_info = if let Some(object_info) = object_info {
if object_info.bucket != bucket || object_info.name != object {
return Ok(false);
}
let snapshot_version_id = object_info
.version_id
.filter(|version_id| !version_id.is_nil())
.map(|version_id| version_id.to_string());
if snapshot_version_id.as_deref() != version_id {
return Ok(false);
}
object_info.clone()
} else {
match self
.get_object_info(
bucket,
object,
&ObjectOptions {
version_id: version_id.map(str::to_string),
versioned: version_id.is_some(),
expected_bucket_incarnation_id: Some(context.configs.bucket_incarnation_id),
..Default::default()
},
)
.await
{
Ok(object_info) => object_info,
Err(err) if is_err_object_not_found(&err) || is_err_version_not_found(&err) => return Ok(false),
Err(err) => return Err(err),
}
};
let event = eval_action_from_lifecycle(lifecycle_config, context.configs.object_lock.as_deref(), &object_info).await;
if !lifecycle_action_skips_heal_version(event.action) {
return Ok(false);
}
if lifecycle_delete_all_versions_blocked_by_replication(self.clone(), bucket, &object_info.name, event.action).await? {
return Ok(false);
}
Ok(apply_expiry_rule_in(self.clone(), &event, &LcEventSrc::Scanner, &object_info).await)
}
async fn save_current_pool_meta(&self) -> Result<()> { async fn save_current_pool_meta(&self) -> Result<()> {
let _save_guard = self.pool_meta_save_gate.lock().await; let _save_guard = self.pool_meta_save_gate.lock().await;
let snapshot = { let snapshot = {
@@ -4287,6 +4365,19 @@ mod tests {
)); ));
} }
#[test]
fn lifecycle_action_skips_heal_version_for_every_delete_action() {
assert!(lifecycle_action_skips_heal_version(IlmAction::DeleteAction));
assert!(lifecycle_action_skips_heal_version(IlmAction::DeleteVersionAction));
assert!(lifecycle_action_skips_heal_version(IlmAction::DeleteRestoredAction));
assert!(lifecycle_action_skips_heal_version(IlmAction::DeleteRestoredVersionAction));
assert!(lifecycle_action_skips_heal_version(IlmAction::DeleteAllVersionsAction));
assert!(lifecycle_action_skips_heal_version(IlmAction::DelMarkerDeleteAllVersionsAction));
assert!(!lifecycle_action_skips_heal_version(IlmAction::TransitionAction));
assert!(!lifecycle_action_skips_heal_version(IlmAction::TransitionVersionAction));
assert!(!lifecycle_action_skips_heal_version(IlmAction::NoneAction));
}
#[test] #[test]
fn resolve_data_movement_lifecycle_expiry_result_allows_dry_run_skip() { fn resolve_data_movement_lifecycle_expiry_result_allows_dry_run_skip() {
let skip = resolve_data_movement_lifecycle_expiry_result(IlmAction::DeleteVersionAction, false, false) let skip = resolve_data_movement_lifecycle_expiry_result(IlmAction::DeleteVersionAction, false, false)
@@ -4958,13 +5049,19 @@ fn is_disk_online_state(state: &str) -> bool {
} }
#[deprecated(since = "0.1.0", note = "Use fallback_total_capacity_dedup instead")] #[deprecated(since = "0.1.0", note = "Use fallback_total_capacity_dedup instead")]
#[allow(dead_code)] #[allow(
dead_code,
reason = "superseded by the replacement named in the comment at pools.rs:5071 (backlog#1823)"
)]
fn fallback_total_capacity(disks: &[rustfs_madmin::Disk]) -> usize { fn fallback_total_capacity(disks: &[rustfs_madmin::Disk]) -> usize {
fallback_total_capacity_dedup(disks) fallback_total_capacity_dedup(disks)
} }
#[deprecated(since = "0.1.0", note = "Use fallback_free_capacity_dedup instead")] #[deprecated(since = "0.1.0", note = "Use fallback_free_capacity_dedup instead")]
#[allow(dead_code)] #[allow(
dead_code,
reason = "superseded by the replacement named in the comment at pools.rs:5071 (backlog#1823)"
)]
fn fallback_free_capacity(disks: &[rustfs_madmin::Disk]) -> usize { fn fallback_free_capacity(disks: &[rustfs_madmin::Disk]) -> usize {
fallback_free_capacity_dedup(disks) fallback_free_capacity_dedup(disks)
} }
+20 -9
View File
@@ -1140,11 +1140,11 @@ impl crate::storage_api_contracts::heal::HealOperations for Sets {
Err(Error::DiskNotFound) Err(Error::DiskNotFound)
} }
#[tracing::instrument(skip(self))] #[tracing::instrument(level = "debug", skip(self, opts), fields(bucket = %bucket, object = %object, dry_run = opts.dry_run))]
async fn check_abandoned_parts(&self, _bucket: &str, _object: &str, _opts: &HealOpts) -> Result<()> { async fn check_abandoned_parts(&self, bucket: &str, object: &str, opts: &HealOpts) -> Result<()> {
// Multipart orphan reconciliation is intentionally retained above the pool/set layers self.get_disks_for_heal_object(object, opts)?
// until there is a concrete caller and a stable lower-level contract to implement. .check_abandoned_parts(bucket, object, opts)
Err(StorageError::NotImplemented) .await
} }
} }
@@ -1996,7 +1996,7 @@ mod tests {
} }
#[tokio::test] #[tokio::test]
async fn sets_check_abandoned_parts_returns_typed_not_implemented_error() { async fn sets_check_abandoned_parts_rejects_invalid_set_scope() {
let format = FormatV3::new(1, 1); let format = FormatV3::new(1, 1);
let sets = Sets { let sets = Sets {
id: format.id, id: format.id,
@@ -2021,10 +2021,21 @@ mod tests {
}; };
let err = sets let err = sets
.check_abandoned_parts("bucket", "object", &HealOpts::default()) .check_abandoned_parts(
"bucket",
"object",
&HealOpts {
set: Some(1),
..Default::default()
},
)
.await .await
.expect_err("abandoned-parts ownership should stay above the pool/set storage layers"); .expect_err("out-of-range abandoned-parts set scope must fail closed");
assert!(matches!(err, StorageError::NotImplemented)); assert!(
matches!(err, StorageError::InvalidArgument(_, ref field, ref reason)
if field == "set" && reason.contains("invalid heal set index 1")),
"unexpected invalid set error: {err:?}"
);
} }
// Builds a single-set `Sets` over `SET_DRIVE_COUNT` local temp-dir disks, // Builds a single-set `Sets` over `SET_DRIVE_COUNT` local temp-dir disks,
+104 -26
View File
@@ -418,6 +418,17 @@ pub struct DiskHealthTracker {
pub last_capacity_free: AtomicU64, pub last_capacity_free: AtomicU64,
/// Last successful capacity probe timestamp /// Last successful capacity probe timestamp
pub last_capacity_probe_unix_secs: AtomicI64, pub last_capacity_probe_unix_secs: AtomicI64,
/// Authoritative atomically published runtime/status pair.
state_snapshot: AtomicU64,
transition_lock: std::sync::Mutex<()>,
}
fn pack_health_state(runtime_state: RuntimeDriveHealthState, status: u32) -> u64 {
(u64::from(runtime_state as u32) << 32) | u64::from(status)
}
fn unpack_health_state(snapshot: u64) -> (RuntimeDriveHealthState, u32) {
(RuntimeDriveHealthState::from_u32((snapshot >> 32) as u32), snapshot as u32)
} }
#[derive(Debug)] #[derive(Debug)]
@@ -739,6 +750,8 @@ impl DiskHealthTracker {
last_capacity_used: AtomicU64::new(0), last_capacity_used: AtomicU64::new(0),
last_capacity_free: AtomicU64::new(0), last_capacity_free: AtomicU64::new(0),
last_capacity_probe_unix_secs: AtomicI64::new(0), last_capacity_probe_unix_secs: AtomicI64::new(0),
state_snapshot: AtomicU64::new(pack_health_state(RuntimeDriveHealthState::Online, DISK_HEALTH_OK)),
transition_lock: std::sync::Mutex::new(()),
} }
} }
@@ -775,39 +788,52 @@ impl DiskHealthTracker {
/// Check if disk is faulty /// Check if disk is faulty
pub fn is_faulty(&self) -> bool { pub fn is_faulty(&self) -> bool {
self.status.load(Ordering::Acquire) == DISK_HEALTH_FAULTY unpack_health_state(self.state_snapshot.load(Ordering::Acquire)).1 == DISK_HEALTH_FAULTY
}
fn publish_state(&self, runtime_state: RuntimeDriveHealthState, status: u32) {
self.state_snapshot
.store(pack_health_state(runtime_state, status), Ordering::Release);
self.runtime_state.store(runtime_state as u32, Ordering::Release);
self.status.store(status, Ordering::Release);
} }
/// Set disk as faulty /// Set disk as faulty
#[allow(dead_code, reason = "asserted by this file's tests (backlog#1823)")] #[allow(dead_code, reason = "asserted by this file's tests (backlog#1823)")]
pub fn set_faulty(&self) { pub fn set_faulty(&self) {
self.status.store(DISK_HEALTH_FAULTY, Ordering::Release); let _guard = self.transition_lock.lock().unwrap_or_else(|poisoned| poisoned.into_inner());
self.publish_state(RuntimeDriveHealthState::Offline, DISK_HEALTH_FAULTY);
} }
/// Set disk as OK /// Set disk as OK
pub fn set_ok(&self) { pub fn set_ok(&self) {
self.status.store(DISK_HEALTH_OK, Ordering::Release); let _guard = self.transition_lock.lock().unwrap_or_else(|poisoned| poisoned.into_inner());
self.publish_state(RuntimeDriveHealthState::Online, DISK_HEALTH_OK);
} }
#[cfg(test)] #[cfg(test)]
pub fn force_runtime_state_for_test(&self, state: RuntimeDriveHealthState) { pub fn force_runtime_state_for_test(&self, state: RuntimeDriveHealthState) {
self.runtime_state.store(state as u32, Ordering::Release); let _guard = self.transition_lock.lock().unwrap_or_else(|poisoned| poisoned.into_inner());
match state { let status = if state == RuntimeDriveHealthState::Offline {
RuntimeDriveHealthState::Offline => self.set_faulty(), DISK_HEALTH_FAULTY
RuntimeDriveHealthState::Online | RuntimeDriveHealthState::Suspect | RuntimeDriveHealthState::Returning => { } else {
self.set_ok(); DISK_HEALTH_OK
} };
} self.publish_state(state, status);
} }
pub fn swap_ok_to_faulty(&self) -> bool { pub fn swap_ok_to_faulty(&self) -> bool {
self.status let _guard = self.transition_lock.lock().unwrap_or_else(|poisoned| poisoned.into_inner());
.compare_exchange(DISK_HEALTH_OK, DISK_HEALTH_FAULTY, Ordering::AcqRel, Ordering::Relaxed) let (_, status) = unpack_health_state(self.state_snapshot.load(Ordering::Acquire));
.is_ok() if status != DISK_HEALTH_OK {
return false;
}
self.publish_state(RuntimeDriveHealthState::Offline, DISK_HEALTH_FAULTY);
true
} }
pub fn runtime_state(&self) -> RuntimeDriveHealthState { pub fn runtime_state(&self) -> RuntimeDriveHealthState {
RuntimeDriveHealthState::from_u32(self.runtime_state.load(Ordering::Acquire)) unpack_health_state(self.state_snapshot.load(Ordering::Acquire)).0
} }
pub fn offline_duration(&self) -> Option<Duration> { pub fn offline_duration(&self) -> Option<Duration> {
@@ -823,6 +849,7 @@ impl DiskHealthTracker {
} }
pub fn mark_failure(&self, endpoint: &Endpoint, reason: &'static str) -> bool { pub fn mark_failure(&self, endpoint: &Endpoint, reason: &'static str) -> bool {
let _guard = self.transition_lock.lock().unwrap_or_else(|poisoned| poisoned.into_inner());
let current = self.runtime_state(); let current = self.runtime_state();
let now = current_unix_secs(); let now = current_unix_secs();
let next = match current { let next = match current {
@@ -851,24 +878,19 @@ impl DiskHealthTracker {
}; };
let became_offline = next == RuntimeDriveHealthState::Offline && current != RuntimeDriveHealthState::Offline; let became_offline = next == RuntimeDriveHealthState::Offline && current != RuntimeDriveHealthState::Offline;
if next == RuntimeDriveHealthState::Offline {
self.status.store(DISK_HEALTH_FAULTY, Ordering::Release);
} else {
self.status.store(DISK_HEALTH_OK, Ordering::Release);
}
self.transition_state(endpoint, current, next, reason); self.transition_state(endpoint, current, next, reason);
became_offline became_offline
} }
#[allow(dead_code, reason = "asserted by this file's tests (backlog#1823)")] #[allow(dead_code, reason = "asserted by this file's tests (backlog#1823)")]
pub fn mark_offline(&self, endpoint: &Endpoint, reason: &'static str) -> bool { pub fn mark_offline(&self, endpoint: &Endpoint, reason: &'static str) -> bool {
let _guard = self.transition_lock.lock().unwrap_or_else(|poisoned| poisoned.into_inner());
let current = self.runtime_state(); let current = self.runtime_state();
if current == RuntimeDriveHealthState::Offline { if current == RuntimeDriveHealthState::Offline {
return false; return false;
} }
self.consecutive_successes.store(0, Ordering::Release); self.consecutive_successes.store(0, Ordering::Release);
self.status.store(DISK_HEALTH_FAULTY, Ordering::Release);
self.transition_state(endpoint, current, RuntimeDriveHealthState::Offline, reason); self.transition_state(endpoint, current, RuntimeDriveHealthState::Offline, reason);
true true
} }
@@ -882,11 +904,10 @@ impl DiskHealthTracker {
} }
fn reset_for_store_init_retry_at(&self, endpoint: &Endpoint, now: Duration) { fn reset_for_store_init_retry_at(&self, endpoint: &Endpoint, now: Duration) {
let _guard = self.transition_lock.lock().unwrap_or_else(|poisoned| poisoned.into_inner());
let now_nanos = unix_nanos(now); let now_nanos = unix_nanos(now);
let now_secs = unix_secs_i64(now); let now_secs = unix_secs_i64(now);
self.status.store(DISK_HEALTH_OK, Ordering::Release); self.publish_state(RuntimeDriveHealthState::Online, DISK_HEALTH_OK);
self.runtime_state
.store(RuntimeDriveHealthState::Online as u32, Ordering::Release);
self.consecutive_failures.store(0, Ordering::Release); self.consecutive_failures.store(0, Ordering::Release);
self.consecutive_successes.store(0, Ordering::Release); self.consecutive_successes.store(0, Ordering::Release);
self.offline_since_unix_secs.store(0, Ordering::Release); self.offline_since_unix_secs.store(0, Ordering::Release);
@@ -898,6 +919,7 @@ impl DiskHealthTracker {
} }
pub fn mark_recovery_success(&self, endpoint: &Endpoint, reason: &'static str) -> bool { pub fn mark_recovery_success(&self, endpoint: &Endpoint, reason: &'static str) -> bool {
let _guard = self.transition_lock.lock().unwrap_or_else(|poisoned| poisoned.into_inner());
let current = self.runtime_state(); let current = self.runtime_state();
let next = match current { let next = match current {
RuntimeDriveHealthState::Online => RuntimeDriveHealthState::Online, RuntimeDriveHealthState::Online => RuntimeDriveHealthState::Online,
@@ -918,7 +940,6 @@ impl DiskHealthTracker {
let became_online = next == RuntimeDriveHealthState::Online; let became_online = next == RuntimeDriveHealthState::Online;
if became_online { if became_online {
self.status.store(DISK_HEALTH_OK, Ordering::Release);
self.consecutive_failures.store(0, Ordering::Release); self.consecutive_failures.store(0, Ordering::Release);
self.consecutive_successes.store(0, Ordering::Release); self.consecutive_successes.store(0, Ordering::Release);
} }
@@ -948,7 +969,13 @@ impl DiskHealthTracker {
return; return;
} }
self.runtime_state.store(next as u32, Ordering::Release); let current_status = unpack_health_state(self.state_snapshot.load(Ordering::Acquire)).1;
let status = match next {
RuntimeDriveHealthState::Offline => DISK_HEALTH_FAULTY,
RuntimeDriveHealthState::Returning => current_status,
RuntimeDriveHealthState::Online | RuntimeDriveHealthState::Suspect => DISK_HEALTH_OK,
};
self.publish_state(next, status);
self.last_transition_unix_secs self.last_transition_unix_secs
.store(current_unix_secs() as i64, Ordering::Release); .store(current_unix_secs() as i64, Ordering::Release);
@@ -1217,7 +1244,7 @@ impl LocalDiskWrapper {
return; return;
} }
if health.status.load(Ordering::Relaxed) != DISK_HEALTH_OK { if health.is_faulty() {
continue; continue;
} }
@@ -2909,6 +2936,57 @@ mod tests {
}); });
} }
#[test]
#[serial_test::serial]
fn concurrent_failure_and_recovery_publish_one_health_snapshot() {
temp_env::with_var(rustfs_config::ENV_DRIVE_SUSPECT_FAILURE_THRESHOLD, Some("2"), || {
let endpoint = Endpoint::try_from("/tmp/concurrent-health-snapshot").expect("endpoint should parse");
let health = Arc::new(DiskHealthTracker::new());
let transition_guard = health
.transition_lock
.lock()
.expect("health transition lock should not be poisoned");
let start = Arc::new(std::sync::Barrier::new(3));
let (completed_tx, completed_rx) = std::sync::mpsc::channel();
let workers = (0..2)
.map(|_| {
let health = Arc::clone(&health);
let endpoint = endpoint.clone();
let start = Arc::clone(&start);
let completed_tx = completed_tx.clone();
std::thread::spawn(move || {
start.wait();
health.mark_failure(&endpoint, "concurrent_test");
completed_tx.send(()).expect("completion receiver should remain available");
})
})
.collect::<Vec<_>>();
start.wait();
assert!(
matches!(
completed_rx.recv_timeout(Duration::from_millis(250)),
Err(std::sync::mpsc::RecvTimeoutError::Timeout)
),
"concurrent transitions must wait for the serialization lock"
);
drop(transition_guard);
completed_rx
.recv_timeout(Duration::from_secs(1))
.expect("first failure transition should complete after lock release");
completed_rx
.recv_timeout(Duration::from_secs(1))
.expect("second failure transition should complete after lock release");
for worker in workers {
worker.join().expect("health transition worker should not panic");
}
assert_eq!(health.runtime_state(), RuntimeDriveHealthState::Offline);
assert!(health.is_faulty());
assert_eq!(health.consecutive_failures.load(Ordering::Acquire), 2);
});
}
#[test] #[test]
fn operation_success_recovers_suspect_drive_without_faulting() { fn operation_success_recovers_suspect_drive_without_faulting() {
let endpoint = Endpoint::try_from("/tmp/runtime-state-suspect-success").expect("endpoint should parse"); let endpoint = Endpoint::try_from("/tmp/runtime-state-suspect-success").expect("endpoint should parse");
+1 -1
View File
@@ -6562,7 +6562,7 @@ impl LocalDisk {
Ok(f) Ok(f)
} }
#[allow(dead_code)] #[allow(dead_code, reason = "MinIO-parity surface with no caller in this port (backlog#1823)")]
fn get_metrics(&self) -> DiskMetrics { fn get_metrics(&self) -> DiskMetrics {
DiskMetrics::default() DiskMetrics::default()
} }
+1 -1
View File
@@ -835,7 +835,7 @@ pub const BITROT_SELF_TEST_PAYLOAD_LEN: usize = 4096;
/// Known-answer digest of [`bitrot_self_test_payload`] under `HighwayHash256S` /// Known-answer digest of [`bitrot_self_test_payload`] under `HighwayHash256S`
/// (the production default). Pinned so any platform or build where the /// (the production default). Pinned so any platform or build where the
/// implementation drifts fails startup instead of mis-hashing shards. /// implementation drifts fails startup instead of miss-hashing shards.
const BITROT_SELF_TEST_KAT_HIGHWAY_HASH256S: [u8; 32] = [ const BITROT_SELF_TEST_KAT_HIGHWAY_HASH256S: [u8; 32] = [
0xb9, 0x32, 0xa2, 0xaa, 0x4a, 0xb7, 0x33, 0x6a, 0xa3, 0xca, 0x7e, 0x61, 0x9d, 0x86, 0x52, 0x14, 0x6e, 0x7f, 0xd8, 0x9e, 0xea, 0xb9, 0x32, 0xa2, 0xaa, 0x4a, 0xb7, 0x33, 0x6a, 0xa3, 0xca, 0x7e, 0x61, 0x9d, 0x86, 0x52, 0x14, 0x6e, 0x7f, 0xd8, 0x9e, 0xea,
0x08, 0xd9, 0x8c, 0x33, 0x85, 0x87, 0x19, 0x30, 0xd6, 0xed, 0x06, 0x08, 0xd9, 0x8c, 0x33, 0x85, 0x87, 0x19, 0x30, 0xd6, 0xed, 0x06,
+8
View File
@@ -1116,6 +1116,14 @@ mod tests {
assert!(encoder_source.is::<reed_solomon_erasure::Error>()); assert!(encoder_source.is::<reed_solomon_erasure::Error>());
} }
// The lifecycle transition worker relies on this arm alone to suppress the
// closed-connection noise (`bucket_lifecycle_ops.rs`); dropping it here would
// silently turn shutdown races back into `error!` log spam.
#[test]
fn is_network_or_host_down_covers_closed_network_connection() {
assert!(is_network_or_host_down("transition failed: use of closed network connection", false));
}
// Regression for #952 (ECA-11): an all-`DiskNotFound` slice (every drive in // Regression for #952 (ECA-11): an all-`DiskNotFound` slice (every drive in
// every set unreachable) must NOT be classified as "all not found", // every set unreachable) must NOT be classified as "all not found",
// otherwise ListObjects silently returns an empty listing and masks a full // otherwise ListObjects silently returns an empty listing and masks a full
@@ -132,7 +132,6 @@ impl RebalanceStopPropagationRecord {
} }
} }
#[allow(dead_code)]
#[derive(Debug, Clone, Default)] #[derive(Debug, Clone, Default)]
pub struct DiskStat { pub struct DiskStat {
pub total_space: u64, pub total_space: u64,
@@ -16,8 +16,16 @@ use serde::{Deserialize, Serialize};
use std::{fmt::Display, io}; use std::{fmt::Display, io};
use tracing::info; use tracing::info;
#[allow(
dead_code,
reason = "tier config wire version stamped by the parity constructors below (backlog#1823)"
)]
const C_TIER_CONFIG_VER: &str = "v1"; const C_TIER_CONFIG_VER: &str = "v1";
#[allow(
dead_code,
reason = "tier-name validation message reached only from the parity constructors below (backlog#1823)"
)]
const ERR_TIER_NAME_EMPTY: &str = "remote tier name empty"; const ERR_TIER_NAME_EMPTY: &str = "remote tier name empty";
const WASABI_US_EAST_ENDPOINT: &str = "https://s3.wasabisys.com"; const WASABI_US_EAST_ENDPOINT: &str = "https://s3.wasabisys.com";
const WASABI_ALTERNATIVE_ENDPOINTS: &[(&str, &str)] = &[ const WASABI_ALTERNATIVE_ENDPOINTS: &[(&str, &str)] = &[
@@ -264,7 +272,6 @@ impl Clone for TierConfig {
} }
} }
#[allow(dead_code)]
impl TierConfig { impl TierConfig {
pub(crate) fn clone_with_credentials(&self) -> Self { pub(crate) fn clone_with_credentials(&self) -> Self {
Self { Self {
@@ -284,6 +291,7 @@ impl TierConfig {
} }
} }
#[allow(dead_code, reason = "MinIO-parity surface with no caller in this port (backlog#1823)")]
fn endpoint(&self) -> String { fn endpoint(&self) -> String {
match self.tier_type { match self.tier_type {
TierType::S3 => self.s3.as_ref().map(|s| s.endpoint.clone()).unwrap_or_default(), TierType::S3 => self.s3.as_ref().map(|s| s.endpoint.clone()).unwrap_or_default(),
@@ -303,6 +311,7 @@ impl TierConfig {
} }
} }
#[allow(dead_code, reason = "MinIO-parity surface with no caller in this port (backlog#1823)")]
fn bucket(&self) -> String { fn bucket(&self) -> String {
match self.tier_type { match self.tier_type {
TierType::S3 => self.s3.as_ref().map(|s| s.bucket.clone()).unwrap_or_default(), TierType::S3 => self.s3.as_ref().map(|s| s.bucket.clone()).unwrap_or_default(),
@@ -322,6 +331,7 @@ impl TierConfig {
} }
} }
#[allow(dead_code, reason = "MinIO-parity surface with no caller in this port (backlog#1823)")]
fn prefix(&self) -> String { fn prefix(&self) -> String {
match self.tier_type { match self.tier_type {
TierType::S3 => self.s3.as_ref().map(|s| s.prefix.clone()).unwrap_or_default(), TierType::S3 => self.s3.as_ref().map(|s| s.prefix.clone()).unwrap_or_default(),
@@ -341,6 +351,7 @@ impl TierConfig {
} }
} }
#[allow(dead_code, reason = "MinIO-parity surface with no caller in this port (backlog#1823)")]
fn region(&self) -> String { fn region(&self) -> String {
match self.tier_type { match self.tier_type {
TierType::S3 => self.s3.as_ref().map(|s| s.region.clone()).unwrap_or_default(), TierType::S3 => self.s3.as_ref().map(|s| s.region.clone()).unwrap_or_default(),
@@ -457,7 +468,7 @@ impl TierWasabi {
} }
impl TierS3 { impl TierS3 {
#[allow(dead_code)] #[allow(dead_code, reason = "MinIO-parity surface with no caller in this port (backlog#1823)")]
fn create<F>( fn create<F>(
name: &str, name: &str,
access_key: &str, access_key: &str,
@@ -528,7 +539,7 @@ pub struct TierMinIO {
} }
impl TierMinIO { impl TierMinIO {
#[allow(dead_code)] #[allow(dead_code, reason = "MinIO-parity surface with no caller in this port (backlog#1823)")]
fn create<F>( fn create<F>(
name: &str, name: &str,
endpoint: &str, endpoint: &str,
@@ -14,7 +14,6 @@
use crate::services::tier::tier::TierConfigMgr; use crate::services::tier::tier::TierConfigMgr;
#[allow(dead_code)]
impl TierConfigMgr { impl TierConfigMgr {
pub fn msg_size(&self) -> usize { pub fn msg_size(&self) -> usize {
100 100
@@ -4860,6 +4860,14 @@ impl SetDisks {
/// is best-effort maintenance: individual delete failures are logged and /// is best-effort maintenance: individual delete failures are logged and
/// skipped rather than propagated. /// skipped rather than propagated.
pub(crate) async fn reclaim_orphan_data_dirs(&self, bucket: &str, object: &str) -> disk::error::Result<usize> { pub(crate) async fn reclaim_orphan_data_dirs(&self, bucket: &str, object: &str) -> disk::error::Result<usize> {
self.reclaim_orphan_data_dirs_inner(bucket, object, false).await
}
pub(crate) async fn dry_run_reclaim_orphan_data_dirs(&self, bucket: &str, object: &str) -> disk::error::Result<usize> {
self.reclaim_orphan_data_dirs_inner(bucket, object, true).await
}
async fn reclaim_orphan_data_dirs_inner(&self, bucket: &str, object: &str, dry_run: bool) -> disk::error::Result<usize> {
let disks = self.get_disks_internal().await; let disks = self.get_disks_internal().await;
// Phase 1 (read-only): build the referenced-data-dir union and record the // Phase 1 (read-only): build the referenced-data-dir union and record the
@@ -4967,6 +4975,20 @@ impl SetDisks {
continue; continue;
} }
let stray = format!("{object}/{dir}"); let stray = format!("{object}/{dir}");
if dry_run {
removed += 1;
debug!(
target: "rustfs_ecstore::set_disk",
event = "heal_abandoned_parts",
component = "ecstore",
subsystem = "heal",
state = "dry_run_matched",
result = "matched",
bucket, object, data_dir = %dir,
"Heal abandoned parts dry-run matched orphaned data directory"
);
continue;
}
match disk match disk
.delete( .delete(
bucket, bucket,
+105 -4
View File
@@ -6998,6 +6998,100 @@ mod tests {
assert!(object_dir.join(STORAGE_FORMAT_FILE).exists(), "metadata must be preserved"); assert!(object_dir.join(STORAGE_FORMAT_FILE).exists(), "metadata must be preserved");
} }
async fn recv_abandoned_parts_trace(
trace: &mut rustfs_common::trace_bus::TraceSubscription,
bucket: &str,
object: &str,
state: &str,
) -> rustfs_common::trace_bus::TraceEvent {
for _ in 0..32 {
let event = tokio::time::timeout(std::time::Duration::from_secs(1), trace.recv())
.await
.expect("abandoned-parts trace event should arrive")
.expect("trace bus should stay open");
if event.kind == rustfs_common::trace_bus::TraceKind::Heal
&& event.func == rustfs_common::trace_bus::TraceFunc::HealCheckAbandonedParts
&& event.bucket.as_deref() == Some(bucket)
&& event.object.as_deref() == Some(object)
&& trace_attr_string(&event, "state").as_deref() == Some(state)
{
return (*event).clone();
}
}
panic!("expected abandoned-parts trace state {state} for {bucket}/{object}");
}
fn trace_attr_string(event: &rustfs_common::trace_bus::TraceEvent, key: &str) -> Option<String> {
event.attrs.iter().find_map(|attr| {
if attr.key != key {
return None;
}
Some(match &attr.value {
rustfs_common::trace_bus::TraceVal::Bool(value) => value.to_string(),
rustfs_common::trace_bus::TraceVal::U64(value) => value.to_string(),
rustfs_common::trace_bus::TraceVal::I64(value) => value.to_string(),
rustfs_common::trace_bus::TraceVal::Str(value) => value.to_string(),
})
})
}
#[tokio::test]
async fn check_abandoned_parts_dry_run_counts_without_deleting() {
let mut trace = rustfs_common::trace_bus::subscribe_trace_events();
let (dir, disk) = make_single_local_disk().await;
let live = Uuid::new_v4();
let orphan = Uuid::new_v4();
let object_dir = dir.path().join("bucket").join("obj");
write_object_meta_with_data_dirs(&object_dir, "bucket", "obj", &[live]).await;
fs::create_dir_all(object_dir.join(live.to_string()))
.await
.expect("live data dir should be created");
fs::create_dir_all(object_dir.join(orphan.to_string()))
.await
.expect("orphan data dir should be created");
let set = make_set_disks_with(vec![Some(disk)]).await;
set.check_abandoned_parts(
"bucket",
"obj",
&HealOpts {
dry_run: true,
no_lock: true,
..Default::default()
},
)
.await
.expect("dry-run abandoned-parts check should succeed");
let dry_run_trace = recv_abandoned_parts_trace(&mut trace, "bucket", "obj", "dry_run_matched").await;
assert_eq!(trace_attr_string(&dry_run_trace, "dry_run").as_deref(), Some("true"));
assert_eq!(trace_attr_string(&dry_run_trace, "data_dirs").as_deref(), Some("1"));
assert!(object_dir.join(live.to_string()).exists(), "referenced data dir must be preserved");
assert!(object_dir.join(orphan.to_string()).exists(), "dry-run must not remove orphaned data dir");
set.check_abandoned_parts(
"bucket",
"obj",
&HealOpts {
no_lock: true,
..Default::default()
},
)
.await
.expect("abandoned-parts check should reclaim stale data dir");
let reclaim_trace = recv_abandoned_parts_trace(&mut trace, "bucket", "obj", "reclaimed").await;
assert_eq!(trace_attr_string(&reclaim_trace, "dry_run").as_deref(), Some("false"));
assert_eq!(trace_attr_string(&reclaim_trace, "data_dirs").as_deref(), Some("1"));
assert!(
object_dir.join(live.to_string()).exists(),
"referenced data dir must remain after reclaim"
);
assert!(!object_dir.join(orphan.to_string()).exists(), "orphaned data dir must be removed");
}
#[tokio::test] #[tokio::test]
async fn reclaim_orphan_data_dirs_recovers_deferred_cleanup_after_restart() { async fn reclaim_orphan_data_dirs_recovers_deferred_cleanup_after_restart() {
let (dir, disk) = make_single_local_disk().await; let (dir, disk) = make_single_local_disk().await;
@@ -12233,11 +12327,18 @@ mod tests {
.expect_err("unsupported copy_object_part should return a typed error"); .expect_err("unsupported copy_object_part should return a typed error");
assert!(matches!(copy_part_err, StorageError::NotImplemented)); assert!(matches!(copy_part_err, StorageError::NotImplemented));
let abandoned_err = set_disks set_disks
.check_abandoned_parts("bucket", "object", &HealOpts::default()) .check_abandoned_parts(
"bucket",
"object",
&HealOpts {
dry_run: true,
no_lock: true,
..Default::default()
},
)
.await .await
.expect_err("abandoned-parts check should stay in the upper reconciliation layer"); .expect("abandoned-parts check should be callable on empty disk sets");
assert!(matches!(abandoned_err, StorageError::NotImplemented));
} }
#[tokio::test] #[tokio::test]
+275 -5
View File
@@ -16,6 +16,7 @@ use super::super::*;
use crate::disk::disk_store::DiskStoreRenameDataExt; use crate::disk::disk_store::DiskStoreRenameDataExt;
use crate::io_support::bitrot::object_mmap_read_enabled; use crate::io_support::bitrot::object_mmap_read_enabled;
use crate::storage_api_contracts::namespace::NamespaceLocking as _; use crate::storage_api_contracts::namespace::NamespaceLocking as _;
use rustfs_common::trace_bus::{TraceEvent, TraceFunc, TraceKind, trace_emit};
use tracing::trace; use tracing::trace;
const LOG_COMPONENT_ECSTORE: &str = "ecstore"; const LOG_COMPONENT_ECSTORE: &str = "ecstore";
@@ -2057,11 +2058,61 @@ impl crate::storage_api_contracts::heal::HealOperations for SetDisks {
Err(Error::DiskNotFound) Err(Error::DiskNotFound)
} }
#[tracing::instrument(skip(self))] #[tracing::instrument(level = "debug", skip(self, opts), fields(bucket = %bucket, object = %object, dry_run = opts.dry_run))]
async fn check_abandoned_parts(&self, _bucket: &str, _object: &str, _opts: &HealOpts) -> Result<()> { async fn check_abandoned_parts(&self, bucket: &str, object: &str, opts: &HealOpts) -> Result<()> {
// Multipart orphan reconciliation is intentionally retained above the set layer let started_at = std::time::Instant::now();
// until there is a concrete caller and a stable lower-level contract to implement. let _write_lock_guard = if !opts.no_lock {
Err(StorageError::NotImplemented) let ns_lock = self.new_ns_lock(bucket, object).await?;
Some(
ns_lock
.get_write_lock(get_lock_acquire_timeout())
.await
.map_err(|e| self.map_namespace_lock_error(bucket, object, "write", e))?,
)
} else {
None
};
let removed = if opts.dry_run {
self.dry_run_reclaim_orphan_data_dirs(bucket, object).await?
} else {
self.reclaim_orphan_data_dirs(bucket, object).await?
};
let state = if opts.dry_run && removed > 0 {
"dry_run_matched"
} else if removed > 0 {
"reclaimed"
} else {
"checked"
};
let data_dirs = u64::try_from(removed).unwrap_or(u64::MAX);
trace_emit(|| {
TraceEvent::new(TraceKind::Heal, TraceFunc::HealCheckAbandonedParts)
.with_bucket(bucket)
.with_object(object)
.with_duration(started_at.elapsed())
.with_attr("state", state)
.with_attr("dry_run", opts.dry_run)
.with_attr("data_dirs", data_dirs)
});
if removed > 0 {
trace!(
event = "heal_abandoned_parts",
component = LOG_COMPONENT_ECSTORE,
subsystem = LOG_SUBSYSTEM_HEAL,
state = if opts.dry_run { "dry_run_matched" } else { "reclaimed" },
result = "ok",
bucket,
object,
dry_run = opts.dry_run,
data_dirs = removed,
"Heal abandoned parts checked object data directories"
);
}
Ok(())
} }
} }
@@ -3246,4 +3297,223 @@ mod heal_result_report_tests {
assert!(result.detail.contains("part 1")); assert!(result.detail.contains("part 1"));
assert!(result.detail.contains("bitrot_failure=true")); assert!(result.detail.contains("bitrot_failure=true"));
} }
// HS-12 (backlog#1874): a versioned DELETE racing an object heal must never
// resurrect the deleted version. The heal has real reconstruction work (a
// shard of the doomed version is removed), so both sides touch the same
// (bucket, object, data_dir); whichever order the ns write lock serializes
// them in, the committed delete must win.
#[tokio::test]
#[serial_test::serial]
async fn heal_racing_version_delete_never_resurrects_the_deleted_version() {
let (temp_dirs, disks, set) = hermetic_set_disks_isolated(4).await;
let bucket = "heal-race-delete-no-resurrect";
let object = "object.bin";
set.make_bucket(
bucket,
&MakeBucketOptions {
versioning_enabled: true,
..Default::default()
},
)
.await
.expect("versioned bucket should be created");
let mut first_reader = PutObjReader::from_vec(vec![0x11; 1024 * 1024]);
let first_info = set
.put_object(
bucket,
object,
&mut first_reader,
&ObjectOptions {
versioned: true,
..Default::default()
},
)
.await
.expect("first version should be written");
let first_version = first_info
.version_id
.expect("versioned put should return the first version id")
.to_string();
let mut second_reader = PutObjReader::from_vec(vec![0x22; 1024 * 1024]);
let second_info = set
.put_object(
bucket,
object,
&mut second_reader,
&ObjectOptions {
versioned: true,
..Default::default()
},
)
.await
.expect("second version should be written");
let second_version = second_info
.version_id
.expect("versioned put should return the second version id")
.to_string();
// Damage one shard of the doomed version so the racing heal performs an
// actual reconstruction over its data dir instead of an early exit.
let doomed_source = disks[0]
.read_version("", bucket, object, &first_version, &ReadOptions::default())
.await
.expect("doomed version metadata should be readable");
let doomed_data_dir = doomed_source
.data_dir
.expect("non-inline version should have a data directory");
tokio::fs::remove_file(
temp_dirs[1]
.path()
.join(bucket)
.join(object)
.join(doomed_data_dir.to_string())
.join("part.1"),
)
.await
.expect("shard damage should be injected before the race");
let delete_set = set.clone();
let (delete_res, heal_res) = tokio::join!(
async {
delete_set
.delete_object(
bucket,
object,
ObjectOptions {
versioned: true,
version_id: Some(first_version.clone()),
object_lock_config_snapshot: Some(Arc::new(crate::set_disk::ObjectLockConfigSnapshot::new(
crate::bucket::metadata_sys::ObjectLockConfigState::ConfirmedAbsent,
))),
..Default::default()
},
)
.await
},
async {
set.heal_object(
bucket,
object,
"",
&HealOpts {
scan_mode: HealScanMode::Deep,
..Default::default()
},
)
.await
},
);
delete_res.expect("version delete must succeed under lock serialization");
// The heal may legitimately report a transient failure when the version
// it was rebuilding disappears mid-flight; only the end state matters.
drop(heal_res);
let resurrected = set
.get_object_info(
bucket,
object,
&ObjectOptions {
versioned: true,
version_id: Some(first_version.clone()),
..Default::default()
},
)
.await;
assert!(
matches!(&resurrected, Err(Error::FileVersionNotFound) | Err(Error::ObjectNotFound(..))),
"a racing heal must not resurrect the deleted version: {resurrected:?}"
);
let survivor = set
.get_object_info(
bucket,
object,
&ObjectOptions {
versioned: true,
version_id: Some(second_version.clone()),
..Default::default()
},
)
.await
.expect("surviving version must remain readable after the race");
assert_eq!(survivor.size, 1024 * 1024, "survivor size must be intact");
}
// HS-12 (backlog#1874): unversioned overwrite commits race a Deep heal on
// the same object. The overwrite's post-commit tail deletes the replaced
// data dir without the ns lock (object.rs commit tail), which is exactly
// the intersection the audit flagged: the heal must tolerate the tail race
// (retryable outcome) and every committed overwrite must survive — the
// final current version is exactly the last payload written.
#[tokio::test]
#[serial_test::serial]
async fn heal_racing_unversioned_overwrites_preserves_the_last_commit() {
let (temp_dirs, disks, set) = hermetic_set_disks_isolated(4).await;
let bucket = "heal-race-put-overwrite";
let object = "object.bin";
set.make_bucket(bucket, &MakeBucketOptions::default())
.await
.expect("bucket should be created");
const ROUNDS: usize = 8;
const PAYLOAD_SIZE: usize = 256 * 1024;
let mut last_etag = String::new();
for round in 0..ROUNDS {
// Give the heal something to rebuild on alternating rounds: remove a
// shard of the current data dir right before the race.
if round % 2 == 1 {
let current = disks[2]
.read_version("", bucket, object, "", &ReadOptions::default())
.await
.expect("current metadata should be readable");
if let Some(data_dir) = current.data_dir {
let shard = temp_dirs[3]
.path()
.join(bucket)
.join(object)
.join(data_dir.to_string())
.join("part.1");
if shard.exists() {
tokio::fs::remove_file(&shard)
.await
.expect("shard damage should be injectable mid-race");
}
}
}
let payload = vec![round as u8; PAYLOAD_SIZE];
let mut put_reader = PutObjReader::from_vec(payload);
let put_opts = ObjectOptions::default();
let heal_opts = HealOpts {
scan_mode: HealScanMode::Deep,
..Default::default()
};
let (put_res, heal_res) = tokio::join!(
set.put_object(bucket, object, &mut put_reader, &put_opts),
set.heal_object(bucket, object, "", &heal_opts),
);
let put_info = put_res.expect("overwrite must succeed under lock serialization");
last_etag = put_info.etag.clone().unwrap_or_default();
// Heal outcome is unconstrained (may hit the tail race and report a
// retryable error); the invariant is checked on the end state.
drop(heal_res);
}
let final_info = set
.get_object_info(bucket, object, &ObjectOptions::default())
.await
.expect("object must remain readable after the race loop");
assert_eq!(
final_info.size, PAYLOAD_SIZE as i64,
"final current version must be the last committed overwrite"
);
assert_eq!(
final_info.etag.unwrap_or_default(),
last_etag,
"the racing heal loop must never leave a stale or resurrected current version"
);
}
} }
+46 -4
View File
@@ -23,6 +23,7 @@
//! per-version `SetDisks::heal_object`. //! per-version `SetDisks::heal_object`.
use super::super::*; use super::super::*;
use crate::object_api::ObjectInfo;
use std::collections::HashSet; use std::collections::HashSet;
use std::sync::Mutex; use std::sync::Mutex;
use std::sync::atomic::{AtomicBool, AtomicUsize, Ordering}; use std::sync::atomic::{AtomicBool, AtomicUsize, Ordering};
@@ -39,12 +40,16 @@ const BACKGROUND_WALKDIR_STALL_TIMEOUT: Duration = Duration::from_secs(60);
/// it must not gate healing logic — the delete-marker vs data path is chosen /// it must not gate healing logic — the delete-marker vs data path is chosen
/// inside `ops/heal.rs` from the resolved latest metadata. `version_id` is /// inside `ops/heal.rs` from the resolved latest metadata. `version_id` is
/// normalized (nil/absent UUID => `None`). /// normalized (nil/absent UUID => `None`).
#[derive(Debug, Clone, PartialEq, Eq)] #[derive(Debug, Clone)]
pub struct HealWalkVersion { pub struct HealWalkVersion {
/// object key /// object key
pub name: String, pub name: String,
/// normalized version id (`None` when the version is nil/absent) /// normalized version id (`None` when the version is nil/absent)
pub version_id: Option<String>, pub version_id: Option<String>,
/// version modification time as Unix nanoseconds
pub mod_time_unix_nanos: Option<i128>,
/// object snapshot for lifecycle evaluation
pub lifecycle_object_info: Option<ObjectInfo>,
/// whether this version is a delete marker (observability only) /// whether this version is a delete marker (observability only)
pub is_delete_marker: bool, pub is_delete_marker: bool,
} }
@@ -63,6 +68,7 @@ struct HealWalkCollector {
bucket: String, bucket: String,
batch_objects: usize, batch_objects: usize,
version_budget: usize, version_budget: usize,
include_lifecycle_object_info: bool,
objects: Mutex<Vec<HealWalkObject>>, objects: Mutex<Vec<HealWalkObject>>,
decode_error: Mutex<Option<DiskError>>, decode_error: Mutex<Option<DiskError>>,
version_total: AtomicUsize, version_total: AtomicUsize,
@@ -116,10 +122,25 @@ impl HealWalkCollector {
let mut versions = Vec::with_capacity(fiv.versions.len() + fiv.free_versions.len()); let mut versions = Vec::with_capacity(fiv.versions.len() + fiv.free_versions.len());
for fi in fiv.versions.iter().chain(fiv.free_versions.iter()) { for fi in fiv.versions.iter().chain(fiv.free_versions.iter()) {
let version_uuid = fi.version_id.filter(|version_id| !version_id.is_nil());
let lifecycle_object_info = if self.include_lifecycle_object_info {
let mut lifecycle_fi = fi.clone();
lifecycle_fi.version_id = version_uuid;
Some(ObjectInfo::from_file_info(
&lifecycle_fi,
&self.bucket,
&entry.name,
version_uuid.is_some(),
))
} else {
None
};
versions.push(HealWalkVersion { versions.push(HealWalkVersion {
name: entry.name.clone(), name: entry.name.clone(),
// Normalize: nil/absent version id => None. // Normalize: nil/absent version id => None.
version_id: fi.version_id.filter(|u| !u.is_nil()).map(|u| u.to_string()), version_id: version_uuid.map(|u| u.to_string()),
mod_time_unix_nanos: fi.mod_time.map(|mod_time| mod_time.unix_timestamp_nanos()),
lifecycle_object_info,
is_delete_marker: fi.deleted, is_delete_marker: fi.deleted,
}); });
} }
@@ -173,11 +194,26 @@ impl HealWalkCollector {
} }
}; };
for fi in fiv.versions.iter().chain(fiv.free_versions.iter()) { for fi in fiv.versions.iter().chain(fiv.free_versions.iter()) {
let vid = fi.version_id.filter(|u| !u.is_nil()).map(|u| u.to_string()); let version_uuid = fi.version_id.filter(|version_id| !version_id.is_nil());
let vid = version_uuid.map(|u| u.to_string());
if seen.insert(vid.clone()) { if seen.insert(vid.clone()) {
let lifecycle_object_info = if self.include_lifecycle_object_info {
let mut lifecycle_fi = fi.clone();
lifecycle_fi.version_id = version_uuid;
Some(ObjectInfo::from_file_info(
&lifecycle_fi,
&self.bucket,
&entry.name,
version_uuid.is_some(),
))
} else {
None
};
versions.push(HealWalkVersion { versions.push(HealWalkVersion {
name: entry.name.clone(), name: entry.name.clone(),
version_id: vid, version_id: vid,
mod_time_unix_nanos: fi.mod_time.map(|mod_time| mod_time.unix_timestamp_nanos()),
lifecycle_object_info,
is_delete_marker: fi.deleted, is_delete_marker: fi.deleted,
}); });
} }
@@ -255,6 +291,7 @@ impl SetDisks {
forward_to: Option<&str>, forward_to: Option<&str>,
batch_objects: usize, batch_objects: usize,
version_budget: usize, version_budget: usize,
include_lifecycle_object_info: bool,
) -> disk::error::Result<(Vec<HealWalkVersion>, Option<String>, bool)> { ) -> disk::error::Result<(Vec<HealWalkVersion>, Option<String>, bool)> {
assert!(batch_objects >= 2, "heal_walk_versions_page requires batch_objects >= 2"); assert!(batch_objects >= 2, "heal_walk_versions_page requires batch_objects >= 2");
@@ -264,6 +301,7 @@ impl SetDisks {
bucket: bucket.to_string(), bucket: bucket.to_string(),
batch_objects, batch_objects,
version_budget: version_budget.max(1), version_budget: version_budget.max(1),
include_lifecycle_object_info,
objects: Mutex::new(Vec::new()), objects: Mutex::new(Vec::new()),
decode_error: Mutex::new(None), decode_error: Mutex::new(None),
version_total: AtomicUsize::new(0), version_total: AtomicUsize::new(0),
@@ -347,6 +385,7 @@ mod tests {
bucket: "bucket".to_string(), bucket: "bucket".to_string(),
batch_objects: 2, batch_objects: 2,
version_budget: 2, version_budget: 2,
include_lifecycle_object_info: false,
objects: Mutex::new(Vec::new()), objects: Mutex::new(Vec::new()),
decode_error: Mutex::new(None), decode_error: Mutex::new(None),
version_total: AtomicUsize::new(0), version_total: AtomicUsize::new(0),
@@ -388,6 +427,8 @@ mod tests {
HealWalkVersion { HealWalkVersion {
name: name.to_string(), name: name.to_string(),
version_id: Some(id.to_string()), version_id: Some(id.to_string()),
mod_time_unix_nanos: None,
lifecycle_object_info: None,
is_delete_marker: dm, is_delete_marker: dm,
} }
} }
@@ -491,6 +532,7 @@ mod tests {
bucket: "bucket".to_string(), bucket: "bucket".to_string(),
batch_objects: 1000, batch_objects: 1000,
version_budget: 10_000, version_budget: 10_000,
include_lifecycle_object_info: false,
objects: Mutex::new(Vec::new()), objects: Mutex::new(Vec::new()),
version_total: AtomicUsize::new(0), version_total: AtomicUsize::new(0),
decode_error: Mutex::new(None), decode_error: Mutex::new(None),
@@ -567,7 +609,7 @@ mod tests {
.expect("corrupt test metadata should be written"); .expect("corrupt test metadata should be written");
let error = set_disks let error = set_disks
.heal_walk_versions_page(bucket, "", None, 2, 2) .heal_walk_versions_page(bucket, "", None, 2, 2, false)
.await .await
.expect_err("semantic metadata corruption must fail the heal disk walk"); .expect_err("semantic metadata corruption must fail the heal disk walk");
@@ -5845,6 +5845,14 @@ impl crate::storage_api_contracts::object::ObjectOperations for SetDisks {
#[tracing::instrument(skip(self))] #[tracing::instrument(skip(self))]
async fn add_partial(&self, bucket: &str, object: &str, version_id: &str) -> Result<()> { async fn add_partial(&self, bucket: &str, object: &str, version_id: &str) -> Result<()> {
// MRF journal intent: partial-write recovery must survive a restart
// (HS-01); the heal request below remains the in-memory fast path.
rustfs_common::mrf_channel::try_send_mrf_intent(
rustfs_common::mrf_channel::MrfKind::PartialWrite,
bucket,
object,
uuid::Uuid::try_parse(version_id).ok(),
);
let mut request = rustfs_common::heal_channel::create_heal_request_with_options( let mut request = rustfs_common::heal_channel::create_heal_request_with_options(
bucket.to_string(), bucket.to_string(),
Some(object.to_string()), Some(object.to_string()),
+9
View File
@@ -1077,6 +1077,15 @@ impl SetDisks {
"Recoverable decode error triggered read repair" "Recoverable decode error triggered read repair"
); );
let version_id = fi.version_id.as_ref().map(ToString::to_string); let version_id = fi.version_id.as_ref().map(ToString::to_string);
// MRF journal intent: keeps a durable Urgent ECDecode
// request alive across restarts even when the in-memory
// read-repair request is dropped or lost (HS-01).
rustfs_common::mrf_channel::try_send_mrf_intent(
rustfs_common::mrf_channel::MrfKind::DecodeFailure,
bucket,
object,
fi.version_id,
);
submit_read_repair_heal( submit_read_repair_heal(
bucket, bucket,
object, object,
+35 -7
View File
@@ -18,6 +18,7 @@ use tracing::trace;
const LOG_COMPONENT_ECSTORE: &str = "ecstore"; const LOG_COMPONENT_ECSTORE: &str = "ecstore";
const LOG_SUBSYSTEM_HEAL: &str = "heal"; const LOG_SUBSYSTEM_HEAL: &str = "heal";
const EVENT_HEAL_ABANDONED_PARTS: &str = "heal_abandoned_parts";
const EVENT_HEAL_FORMAT_COMPLETED: &str = "heal_format_completed"; const EVENT_HEAL_FORMAT_COMPLETED: &str = "heal_format_completed";
const EVENT_HEAL_OBJECT_STARTED: &str = "heal_object_started"; const EVENT_HEAL_OBJECT_STARTED: &str = "heal_object_started";
@@ -256,13 +257,40 @@ impl ECStore {
#[instrument(skip(self))] #[instrument(skip(self))]
pub(super) async fn handle_check_abandoned_parts(&self, bucket: &str, object: &str, opts: &HealOpts) -> Result<()> { pub(super) async fn handle_check_abandoned_parts(&self, bucket: &str, object: &str, opts: &HealOpts) -> Result<()> {
let _ = (bucket, object, opts); let object = encode_dir_object(object);
// Stale multipart reconciliation is already owned by the lifecycle-driven let pools = self.get_pools_for_heal_object(opts)?;
// background cleanup path in `bucket_lifecycle_ops.rs`. There is currently
// no stable object-heal contract that should fan this request out through let mut futures = Vec::with_capacity(pools.len());
// pool/set storage layers, so keep the placeholder explicit at the ECStore for pool in pools.iter() {
// boundary instead of dispatching into lower layers. futures.push(pool.check_abandoned_parts(bucket, &object, opts));
Err(StorageError::NotImplemented) }
let mut first_error = None;
for result in join_all(futures).await {
if let Err(err) = result
&& first_error.is_none()
{
first_error = Some(err);
}
}
if let Some(err) = first_error {
return Err(err);
}
trace!(
event = EVENT_HEAL_ABANDONED_PARTS,
component = LOG_COMPONENT_ECSTORE,
subsystem = LOG_SUBSYSTEM_HEAL,
state = "completed",
result = "ok",
bucket,
object,
dry_run = opts.dry_run,
"Heal abandoned parts completed"
);
Ok(())
} }
} }
+2 -1
View File
@@ -34,6 +34,7 @@ impl ECStore {
forward_to: Option<&str>, forward_to: Option<&str>,
batch_objects: usize, batch_objects: usize,
version_budget: usize, version_budget: usize,
include_lifecycle_object_info: bool,
) -> Result<(Vec<HealWalkVersion>, Option<String>, bool)> { ) -> Result<(Vec<HealWalkVersion>, Option<String>, bool)> {
if pool_idx >= self.pools.len() || set_idx >= self.pools[pool_idx].disk_set.len() { if pool_idx >= self.pools.len() || set_idx >= self.pools[pool_idx].disk_set.len() {
return Err(Error::other(format!( return Err(Error::other(format!(
@@ -43,7 +44,7 @@ impl ECStore {
} }
self.pools[pool_idx].disk_set[set_idx] self.pools[pool_idx].disk_set[set_idx]
.heal_walk_versions_page(bucket, prefix, forward_to, batch_objects, version_budget) .heal_walk_versions_page(bucket, prefix, forward_to, batch_objects, version_budget, include_lifecycle_object_info)
.await .await
.map_err(Error::from) .map_err(Error::from)
} }
+10
View File
@@ -216,6 +216,16 @@ impl std::fmt::Debug for ECStore {
/// These delegate to the process-global statics. No local state — the globals /// These delegate to the process-global statics. No local state — the globals
/// remain the single source of truth until the migration is complete. /// remain the single source of truth until the migration is complete.
impl ECStore { impl ECStore {
/// Every erasure set across all pools, pool-major order.
///
/// Read-only queries that must consult each set's own copy of a
/// per-bucket object (e.g. the scanner's `.usage-cache.bin`) iterate
/// this instead of the hash-routed store path, which would always land
/// on one set (rustfs/backlog#1872).
pub fn all_set_disks(&self) -> Vec<Arc<crate::set_disk::SetDisks>> {
self.pools.iter().flat_map(|pool| pool.disk_set.iter().cloned()).collect()
}
/// Get server configuration (delegates to global) /// Get server configuration (delegates to global)
pub fn get_server_config(&self) -> Option<Config> { pub fn get_server_config(&self) -> Option<Config> {
runtime_sources::server_config() runtime_sources::server_config()
+2
View File
@@ -89,6 +89,8 @@ async-trait = { workspace = true }
futures = { workspace = true } futures = { workspace = true }
metrics = { workspace = true } metrics = { workspace = true }
base64 = { workspace = true } base64 = { workspace = true }
bytes = { workspace = true }
crc-fast = { workspace = true }
[dev-dependencies] [dev-dependencies]
serde_json = { workspace = true, features = ["raw_value"] } serde_json = { workspace = true, features = ["raw_value"] }
+100 -17
View File
@@ -66,21 +66,37 @@ struct HealTaskStatusPayload<'a> {
summary: &'a str, summary: &'a str,
items: &'a [HealResultItem], items: &'a [HealResultItem],
truncated: bool, truncated: bool,
/// Cursor for incremental consumption (HS-06): sequence of the next item
/// to be produced. Absent on responses without sequencing (0).
#[serde(skip_serializing_if = "u64_is_zero")]
next_seq: u64,
/// Oldest sequence still retained; with `truncated`, tells a lagging
/// client where to restart its cursor.
#[serde(skip_serializing_if = "u64_is_zero")]
min_seq: u64,
#[serde(skip_serializing_if = "Option::is_none")] #[serde(skip_serializing_if = "Option::is_none")]
progress: Option<&'a HealProgress>, progress: Option<&'a HealProgress>,
} }
fn u64_is_zero(value: &u64) -> bool {
*value == 0
}
fn encode_heal_task_status_payload( fn encode_heal_task_status_payload(
summary: &str, summary: &str,
mut items: Vec<HealResultItem>, mut items: Vec<HealResultItem>,
progress: Option<&HealProgress>, progress: Option<&HealProgress>,
mut truncated: bool, mut truncated: bool,
next_seq: u64,
min_seq: u64,
) -> Result<(Vec<u8>, bool)> { ) -> Result<(Vec<u8>, bool)> {
loop { loop {
let data = serde_json::to_vec(&HealTaskStatusPayload { let data = serde_json::to_vec(&HealTaskStatusPayload {
summary, summary,
items: &items, items: &items,
truncated, truncated,
next_seq,
min_seq,
progress, progress,
}) })
.map_err(|e| Error::Serialization(format!("failed to serialize heal task status: {e}")))?; .map_err(|e| Error::Serialization(format!("failed to serialize heal task status: {e}")))?;
@@ -109,8 +125,10 @@ fn encode_heal_status_response(
progress: Option<&HealProgress>, progress: Option<&HealProgress>,
detail: Option<String>, detail: Option<String>,
truncated: bool, truncated: bool,
next_seq: u64,
min_seq: u64,
) -> Result<(Vec<u8>, Option<String>)> { ) -> Result<(Vec<u8>, Option<String>)> {
let (data, truncated) = encode_heal_task_status_payload(summary, items, progress, truncated)?; let (data, truncated) = encode_heal_task_status_payload(summary, items, progress, truncated, next_seq, min_seq)?;
Ok((data, heal_status_detail(detail, truncated))) Ok((data, heal_status_detail(detail, truncated)))
} }
@@ -138,8 +156,19 @@ impl HealChannelProcessor {
/// Execute a token query directly against the manager. /// Execute a token query directly against the manager.
pub async fn execute_query_request(&self, heal_path: String, client_token: String) -> Result<HealChannelResponse> { pub async fn execute_query_request(&self, heal_path: String, client_token: String) -> Result<HealChannelResponse> {
self.execute_query_request_since(heal_path, client_token, None).await
}
/// Incremental variant of [`Self::execute_query_request`] (HS-06).
pub async fn execute_query_request_since(
&self,
heal_path: String,
client_token: String,
since_seq: Option<u64>,
) -> Result<HealChannelResponse> {
let (response_tx, response_rx) = oneshot::channel(); let (response_tx, response_rx) = oneshot::channel();
self.process_query_request(heal_path, client_token, response_tx).await?; self.process_query_request(heal_path, client_token, since_seq, response_tx)
.await?;
response_rx response_rx
.await .await
.map_err(|err| Error::other(format!("heal query channel closed: {err}")))? .map_err(|err| Error::other(format!("heal query channel closed: {err}")))?
@@ -262,8 +291,12 @@ impl HealChannelProcessor {
HealChannelCommand::Query { HealChannelCommand::Query {
heal_path, heal_path,
client_token, client_token,
since_seq,
response_tx, response_tx,
} => self.process_query_request(heal_path, client_token, response_tx).await, } => {
self.process_query_request(heal_path, client_token, since_seq, response_tx)
.await
}
HealChannelCommand::Cancel { HealChannelCommand::Cancel {
heal_path, heal_path,
client_token, client_token,
@@ -384,6 +417,7 @@ impl HealChannelProcessor {
&self, &self,
heal_path: String, heal_path: String,
client_token: String, client_token: String,
since_seq: Option<u64>,
response_tx: oneshot::Sender<std::result::Result<HealChannelResponse, String>>, response_tx: oneshot::Sender<std::result::Result<HealChannelResponse, String>>,
) -> Result<()> { ) -> Result<()> {
debug!( debug!(
@@ -398,72 +432,118 @@ impl HealChannelProcessor {
); );
let report = if heal_path.trim_matches('/').is_empty() { let report = if heal_path.trim_matches('/').is_empty() {
self.heal_manager.get_task_report(&client_token).await self.heal_manager.get_task_report_since(&client_token, since_seq).await
} else { } else {
self.heal_manager.get_task_report_for_path(&heal_path, &client_token).await self.heal_manager
.get_task_report_for_path_since(&heal_path, &client_token, since_seq)
.await
}; };
let (summary, detail, items, truncated, progress) = match report { let (summary, detail, items, truncated, progress, next_seq, min_seq) = match report {
Ok(HealTaskReport { Ok(HealTaskReport {
status: HealTaskStatus::Pending | HealTaskStatus::Running, status: HealTaskStatus::Pending | HealTaskStatus::Running,
result_items, result_items,
result_items_truncated, result_items_truncated,
progress, progress,
}) => ("running".to_string(), None, result_items, result_items_truncated, progress), next_seq,
min_seq,
}) => (
"running".to_string(),
None,
result_items,
result_items_truncated,
progress,
next_seq,
min_seq,
),
Ok(HealTaskReport { Ok(HealTaskReport {
status: HealTaskStatus::Retrying { error, retry_attempt }, status: HealTaskStatus::Retrying { error, retry_attempt },
result_items, result_items,
result_items_truncated, result_items_truncated,
progress, progress,
next_seq,
min_seq,
}) => ( }) => (
"running".to_string(), "running".to_string(),
Some(format!("heal task retrying after recoverable failure, attempt {retry_attempt}: {error}")), Some(format!("heal task retrying after recoverable failure, attempt {retry_attempt}: {error}")),
result_items, result_items,
result_items_truncated, result_items_truncated,
progress, progress,
next_seq,
min_seq,
), ),
Ok(HealTaskReport { Ok(HealTaskReport {
status: HealTaskStatus::Completed, status: HealTaskStatus::Completed,
result_items, result_items,
result_items_truncated, result_items_truncated,
progress, progress,
}) => ("finished".to_string(), None, result_items, result_items_truncated, progress), next_seq,
min_seq,
}) => (
"finished".to_string(),
None,
result_items,
result_items_truncated,
progress,
next_seq,
min_seq,
),
Ok(HealTaskReport { Ok(HealTaskReport {
status: HealTaskStatus::Cancelled, status: HealTaskStatus::Cancelled,
result_items, result_items,
result_items_truncated, result_items_truncated,
progress, progress,
next_seq,
min_seq,
}) => ( }) => (
"stopped".to_string(), "stopped".to_string(),
Some("heal task cancelled".to_string()), Some("heal task cancelled".to_string()),
result_items, result_items,
result_items_truncated, result_items_truncated,
progress, progress,
next_seq,
min_seq,
), ),
Ok(HealTaskReport { Ok(HealTaskReport {
status: HealTaskStatus::Timeout, status: HealTaskStatus::Timeout,
result_items, result_items,
result_items_truncated, result_items_truncated,
progress, progress,
next_seq,
min_seq,
}) => ( }) => (
"stopped".to_string(), "stopped".to_string(),
Some("heal task timed out".to_string()), Some("heal task timed out".to_string()),
result_items, result_items,
result_items_truncated, result_items_truncated,
progress, progress,
next_seq,
min_seq,
), ),
Ok(HealTaskReport { Ok(HealTaskReport {
status: HealTaskStatus::Failed { error }, status: HealTaskStatus::Failed { error },
result_items, result_items,
result_items_truncated, result_items_truncated,
progress, progress,
}) => ("stopped".to_string(), Some(error), result_items, result_items_truncated, progress), next_seq,
min_seq,
}) => (
"stopped".to_string(),
Some(error),
result_items,
result_items_truncated,
progress,
next_seq,
min_seq,
),
Err(crate::Error::TaskNotFound { .. }) => ( Err(crate::Error::TaskNotFound { .. }) => (
"notFound".to_string(), "notFound".to_string(),
Some("heal task not found or expired".to_string()), Some("heal task not found or expired".to_string()),
Vec::new(), Vec::new(),
false, false,
None, None,
0,
0,
), ),
Err(crate::Error::InvalidClientToken) => { Err(crate::Error::InvalidClientToken) => {
let response = HealChannelResponse { let response = HealChannelResponse {
@@ -490,7 +570,8 @@ impl HealChannelProcessor {
} }
}; };
let (data, detail) = encode_heal_status_response(&summary, items, progress.as_ref(), detail, truncated)?; let (data, detail) =
encode_heal_status_response(&summary, items, progress.as_ref(), detail, truncated, next_seq, min_seq)?;
let response = HealChannelResponse { let response = HealChannelResponse {
request_id: client_token, request_id: client_token,
@@ -612,7 +693,8 @@ impl HealChannelProcessor {
HealRequestSource::Admin HealRequestSource::Admin
| HealRequestSource::AutoHeal | HealRequestSource::AutoHeal
| HealRequestSource::Internal | HealRequestSource::Internal
| HealRequestSource::ReadRepair => true, | HealRequestSource::ReadRepair
| HealRequestSource::Mrf => true,
}); });
// Build HealOptions with all available fields // Build HealOptions with all available fields
@@ -767,6 +849,7 @@ mod tests {
_bucket: &str, _bucket: &str,
_prefix: &str, _prefix: &str,
_continuation_token: Option<&str>, _continuation_token: Option<&str>,
_include_lifecycle_object_info: bool,
) -> crate::Result<(Vec<crate::heal::storage::HealListItem>, Option<String>, bool)> { ) -> crate::Result<(Vec<crate::heal::storage::HealListItem>, Option<String>, bool)> {
Ok((vec![], None, false)) Ok((vec![], None, false))
} }
@@ -803,7 +886,7 @@ mod tests {
..Default::default() ..Default::default()
}]; }];
let (data, detail) = encode_heal_status_response("running", items, None, None, false).unwrap(); let (data, detail) = encode_heal_status_response("running", items, None, None, false, 0, 0).unwrap();
assert!(data.len() <= MAX_HEAL_STATUS_PAYLOAD_SIZE); assert!(data.len() <= MAX_HEAL_STATUS_PAYLOAD_SIZE);
let payload: serde_json::Value = serde_json::from_slice(&data).unwrap(); let payload: serde_json::Value = serde_json::from_slice(&data).unwrap();
@@ -1573,7 +1656,7 @@ mod tests {
let (tx, rx) = oneshot::channel(); let (tx, rx) = oneshot::channel();
processor processor
.process_query_request("bucket".to_string(), "completed-token".to_string(), tx) .process_query_request("bucket".to_string(), "completed-token".to_string(), None, tx)
.await .await
.expect("query should process"); .expect("query should process");
@@ -1608,7 +1691,7 @@ mod tests {
let (tx, rx) = oneshot::channel(); let (tx, rx) = oneshot::channel();
processor processor
.process_query_request("bucket".to_string(), task_id.clone(), tx) .process_query_request("bucket".to_string(), task_id.clone(), None, tx)
.await .await
.expect("query should process"); .expect("query should process");
@@ -1641,7 +1724,7 @@ mod tests {
let (tx, rx) = oneshot::channel(); let (tx, rx) = oneshot::channel();
processor processor
.process_query_request("bucket".to_string(), "wrong-token".to_string(), tx) .process_query_request("bucket".to_string(), "wrong-token".to_string(), None, tx)
.await .await
.expect("query should process"); .expect("query should process");
@@ -1666,7 +1749,7 @@ mod tests {
let (tx, rx) = oneshot::channel(); let (tx, rx) = oneshot::channel();
processor processor
.process_query_request(String::new(), "wrong-token".to_string(), tx) .process_query_request(String::new(), "wrong-token".to_string(), None, tx)
.await .await
.expect("query should process"); .expect("query should process");
@@ -1703,7 +1786,7 @@ mod tests {
let (tx, rx) = oneshot::channel(); let (tx, rx) = oneshot::channel();
processor processor
.process_query_request(String::new(), task_id.clone(), tx) .process_query_request(String::new(), task_id.clone(), None, tx)
.await .await
.expect("query should process"); .expect("query should process");
+317 -27
View File
@@ -23,13 +23,14 @@ use crate::heal::{
}; };
use crate::{Error, Result}; use crate::{Error, Result};
use futures::{StreamExt, stream::FuturesUnordered}; use futures::{StreamExt, stream::FuturesUnordered};
use metrics::gauge; use metrics::{counter, gauge};
use rustfs_common::heal_channel::{HealOpts, HealRequestSource, HealScanMode}; use rustfs_common::heal_channel::{HealOpts, HealRequestSource, HealScanMode};
use rustfs_madmin::heal_commands::HealResultItem; use rustfs_madmin::heal_commands::HealResultItem;
use std::sync::{ use std::sync::{
Arc, Arc,
atomic::{AtomicUsize, Ordering}, atomic::{AtomicUsize, Ordering},
}; };
use std::time::{Duration, UNIX_EPOCH};
use tokio::sync::{RwLock, Semaphore}; use tokio::sync::{RwLock, Semaphore};
use tracing::{debug, error, warn}; use tracing::{debug, error, warn};
@@ -47,6 +48,21 @@ enum HealObjectOutcome {
Failed, Failed,
} }
fn result_object_size_u64(result: &HealResultItem) -> u64 {
u64::try_from(result.object_size).unwrap_or(u64::MAX)
}
const NEW_VERSION_SKIP_GRACE_SECS: u64 = 60;
const NANOS_PER_SECOND: i128 = 1_000_000_000;
fn should_skip_new_version(mod_time_unix_nanos: Option<i128>, started_at_secs: u64) -> bool {
let Some(mod_time_unix_nanos) = mod_time_unix_nanos else {
return false;
};
let cutoff_secs = started_at_secs.saturating_add(NEW_VERSION_SKIP_GRACE_SECS);
mod_time_unix_nanos > i128::from(cutoff_secs).saturating_mul(NANOS_PER_SECOND)
}
struct PageConcurrencyGuard { struct PageConcurrencyGuard {
in_flight: Arc<AtomicUsize>, in_flight: Arc<AtomicUsize>,
set_label: String, set_label: String,
@@ -492,6 +508,7 @@ impl ErasureSetHealer {
&mut skipped_objects, &mut skipped_objects,
resume_manager, resume_manager,
checkpoint_manager, checkpoint_manager,
state.start_time,
) )
.await; .await;
@@ -658,6 +675,7 @@ impl ErasureSetHealer {
skipped_objects: &mut u64, skipped_objects: &mut u64,
resume_manager: &ResumeManager, resume_manager: &ResumeManager,
checkpoint_manager: &CheckpointManager, checkpoint_manager: &CheckpointManager,
started_at_secs: u64,
) -> Result<()> { ) -> Result<()> {
debug!( debug!(
target: "rustfs::heal::erasure_healer", target: "rustfs::heal::erasure_healer",
@@ -710,6 +728,7 @@ impl ErasureSetHealer {
// The end-of-pass summary reports the full failed/skipped counts. // The end-of-pass summary reports the full failed/skipped counts.
let mut transient_skip_samples_logged = 0_u64; let mut transient_skip_samples_logged = 0_u64;
let mut failure_samples_logged = 0_u64; let mut failure_samples_logged = 0_u64;
let mut bytes_processed = self.progress.read().await.bytes_processed;
// backlog#920: select the per-erasure-set DISK-WALK union enumerator when // backlog#920: select the per-erasure-set DISK-WALK union enumerator when
// the scan is Deep OR the request came from AutoHeal — these are the paths // the scan is Deep OR the request came from AutoHeal — these are the paths
@@ -718,17 +737,25 @@ impl ErasureSetHealer {
// which stays the default. // which stays the default.
let use_disk_walk = let use_disk_walk =
matches!(self.heal_opts.scan_mode, HealScanMode::Deep) || matches!(self.source, HealRequestSource::AutoHeal); matches!(self.heal_opts.scan_mode, HealScanMode::Deep) || matches!(self.source, HealRequestSource::AutoHeal);
let lifecycle_expiry_context = self.storage.load_heal_lifecycle_expiry_context(bucket).await?;
let include_lifecycle_object_info = lifecycle_expiry_context.is_some();
loop { loop {
self.verify_replacement_identity_fence("page scan").await?; self.verify_replacement_identity_fence("page scan").await?;
// Get one page of object versions // Get one page of object versions
let (objects, next_token, is_truncated) = if use_disk_walk { let (objects, next_token, is_truncated) = if use_disk_walk {
self.storage self.storage
.list_versions_for_heal_page_disk_walk(set_disk_id, bucket, "", continuation_token.as_deref()) .list_versions_for_heal_page_disk_walk(
set_disk_id,
bucket,
"",
continuation_token.as_deref(),
include_lifecycle_object_info,
)
.await? .await?
} else { } else {
self.storage self.storage
.list_objects_for_heal_page(bucket, "", continuation_token.as_deref()) .list_objects_for_heal_page(bucket, "", continuation_token.as_deref(), include_lifecycle_object_info)
.await? .await?
}; };
let page_is_empty = objects.is_empty(); let page_is_empty = objects.is_empty();
@@ -736,6 +763,7 @@ impl ErasureSetHealer {
let page_resume_index = *current_object_index; let page_resume_index = *current_object_index;
let semaphore = Arc::new(Semaphore::new(page_concurrency_limit)); let semaphore = Arc::new(Semaphore::new(page_concurrency_limit));
let mut page_tasks = FuturesUnordered::new(); let mut page_tasks = FuturesUnordered::new();
let mut completed_in_page = 0usize;
// Capture the last version identity of this page for the anti-loop guard. // Capture the last version identity of this page for the anti-loop guard.
let page_last = objects.last().map(|item| (item.name.clone(), item.version_id.clone())); let page_last = objects.last().map(|item| (item.name.clone(), item.version_id.clone()));
@@ -751,6 +779,75 @@ impl ErasureSetHealer {
continue; continue;
} }
if should_skip_new_version(item.mod_time_unix_nanos, started_at_secs) {
checkpoint_manager.add_processed_object(key).await?;
*processed_objects = processed_objects.saturating_add(1);
completed_in_page = completed_in_page.saturating_add(1);
counter!("rustfs_heal_skipped_new_versions_total").increment(1);
{
let mut progress = self.progress.write().await;
progress.record_skipped_new_version();
progress.set_current_object(Some(format!("skipped_new: {bucket}/{}", item.name)));
progress.update_progress(*processed_objects, *successful_objects, *failed_objects, bytes_processed);
}
debug!(
target: "rustfs::heal::erasure_healer",
event = EVENT_HEAL_ERASURE_OBJECT_STATE,
component = LOG_COMPONENT_HEAL,
subsystem = LOG_SUBSYSTEM_ERASURE_HEALER,
set_disk_id,
bucket,
object = %item.name,
version_id = ?item.version_id,
state = "skipped_new_version",
"Erasure set object version skipped because it was written after heal started"
);
if completed_in_page.is_multiple_of(100) {
checkpoint_manager.update_position(bucket_index, page_resume_index).await?;
}
continue;
}
if let Some(context) = lifecycle_expiry_context.as_ref()
&& self
.storage
.enqueue_heal_lifecycle_expiry(
context,
bucket,
&item.name,
item.version_id.as_deref(),
item.lifecycle_object_info.as_ref(),
)
.await?
{
checkpoint_manager.add_processed_object(key).await?;
*processed_objects = processed_objects.saturating_add(1);
completed_in_page = completed_in_page.saturating_add(1);
counter!("rustfs_heal_skipped_ilm_expired_total").increment(1);
{
let mut progress = self.progress.write().await;
progress.record_skipped_ilm_expired();
progress.set_current_object(Some(format!("skipped_ilm: {bucket}/{}", item.name)));
progress.update_progress(*processed_objects, *successful_objects, *failed_objects, bytes_processed);
}
debug!(
target: "rustfs::heal::erasure_healer",
event = EVENT_HEAL_ERASURE_OBJECT_STATE,
component = LOG_COMPONENT_HEAL,
subsystem = LOG_SUBSYSTEM_ERASURE_HEALER,
set_disk_id,
bucket,
object = %item.name,
version_id = ?item.version_id,
state = "skipped_ilm_expired",
"Erasure set object version skipped because lifecycle expiry was queued"
);
if completed_in_page.is_multiple_of(100) {
checkpoint_manager.update_position(bucket_index, page_resume_index).await?;
}
continue;
}
resume_manager resume_manager
.set_current_item(Some(bucket.to_string()), Some(item.name.clone())) .set_current_item(Some(bucket.to_string()), Some(item.name.clone()))
.await?; .await?;
@@ -777,7 +874,7 @@ impl ErasureSetHealer {
let _permit = match permit { let _permit = match permit {
Ok(permit) => permit, Ok(permit) => permit,
Err(err) => return (dedup_key, object_name, version_id, Err(err)), Err(err) => return (dedup_key, object_name, version_id, (0, Err(err))),
}; };
let _in_flight_guard = PageConcurrencyGuard::new(in_flight, set_label); let _in_flight_guard = PageConcurrencyGuard::new(in_flight, set_label);
@@ -788,7 +885,7 @@ impl ErasureSetHealer {
// recorded as skipped-ok rather than failed. The delete-marker // recorded as skipped-ok rather than failed. The delete-marker
// vs data path is chosen internally in ops/heal.rs. // vs data path is chosen internally in ops/heal.rs.
let result = if cancel_token.is_cancelled() { let result = if cancel_token.is_cancelled() {
Err(Error::TaskCancelled) (0, Err(Error::TaskCancelled))
} else { } else {
match storage match storage
.heal_object(&bucket_name, &object_name, version_id.as_deref(), &heal_opts) .heal_object(&bucket_name, &object_name, version_id.as_deref(), &heal_opts)
@@ -797,8 +894,9 @@ impl ErasureSetHealer {
Ok((result, None)) Ok((result, None))
if target_outcomes_complete(&result, &target_endpoints) => if target_outcomes_complete(&result, &target_endpoints) =>
{ {
let object_size = result_object_size_u64(&result);
if !replacement_commit_evidence_required { if !replacement_commit_evidence_required {
Ok(true) (object_size, Ok(true))
} else { } else {
match storage match storage
.replacement_targets_have_version( .replacement_targets_have_version(
@@ -810,27 +908,42 @@ impl ErasureSetHealer {
) )
.await .await
{ {
Ok(true) => Ok(true), Ok(true) => (object_size, Ok(true)),
Ok(false) => Err(Error::transient_skip(format!( Ok(false) => (object_size, Err(Error::transient_skip(format!(
"Skipped heal for {bucket_name}/{object_name} because replacement target readback did not confirm the committed version" "Skipped heal for {bucket_name}/{object_name} because replacement target readback did not confirm the committed version"
))), )))),
Err(err) => Err(Error::transient_skip(format!( Err(err) => (object_size, Err(Error::transient_skip(format!(
"Skipped heal for {bucket_name}/{object_name} because replacement target readback failed: {err}" "Skipped heal for {bucket_name}/{object_name} because replacement target readback failed: {err}"
))), )))),
} }
} }
} },
Ok((_result, None)) if !target_endpoints.is_empty() => Err(Error::transient_skip(format!( Ok((result, None)) if !target_endpoints.is_empty() => (
"Skipped heal for {bucket_name}/{object_name} because a replacement target was not committed" result_object_size_u64(&result),
))), Err(Error::transient_skip(format!(
Ok((_result, None)) => Ok(true), "Skipped heal for {bucket_name}/{object_name} because a replacement target was not committed"
Ok((_, Some(err))) if is_missing_object_dir_heal_result(&object_name, &err) => Ok(false),
Ok((_, Some(err))) | Err(err) => match Self::classify_heal_object_error(&err) {
HealObjectOutcome::Absent => Ok(false),
HealObjectOutcome::Transient => Err(Error::transient_skip(format!(
"Skipped heal for {bucket_name}/{object_name} due to transient error: {err}"
))), ))),
HealObjectOutcome::Failed => Err(err), ),
Ok((result, None)) => (result_object_size_u64(&result), Ok(true)),
Ok((result, Some(err))) if is_missing_object_dir_heal_result(&object_name, &err) => {
(result_object_size_u64(&result), Ok(false))
}
Ok((result, Some(err))) => {
let object_size = result_object_size_u64(&result);
match Self::classify_heal_object_error(&err) {
HealObjectOutcome::Absent => (object_size, Ok(false)),
HealObjectOutcome::Transient => (object_size, Err(Error::transient_skip(format!(
"Skipped heal for {bucket_name}/{object_name} due to transient error: {err}"
)))),
HealObjectOutcome::Failed => (object_size, Err(err)),
}
}
Err(err) => match Self::classify_heal_object_error(&err) {
HealObjectOutcome::Absent => (0, Ok(false)),
HealObjectOutcome::Transient => (0, Err(Error::transient_skip(format!(
"Skipped heal for {bucket_name}/{object_name} due to transient error: {err}"
)))),
HealObjectOutcome::Failed => (0, Err(err)),
}, },
} }
}; };
@@ -839,11 +952,12 @@ impl ErasureSetHealer {
}); });
} }
let mut completed_in_page = 0usize;
while let Some((key, object, version_id, result)) = page_tasks.next().await { while let Some((key, object, version_id, result)) = page_tasks.next().await {
let (object_size, result) = result;
match result { match result {
Ok(true) => { Ok(true) => {
*successful_objects += 1; *successful_objects += 1;
bytes_processed = bytes_processed.saturating_add(object_size);
checkpoint_manager.add_processed_object(key).await?; checkpoint_manager.add_processed_object(key).await?;
debug!( debug!(
target: "rustfs::heal::erasure_healer", target: "rustfs::heal::erasure_healer",
@@ -861,6 +975,7 @@ impl ErasureSetHealer {
Ok(false) => { Ok(false) => {
checkpoint_manager.add_processed_object(key).await?; checkpoint_manager.add_processed_object(key).await?;
*successful_objects += 1; *successful_objects += 1;
bytes_processed = bytes_processed.saturating_add(object_size);
debug!( debug!(
target: "rustfs::heal::erasure_healer", target: "rustfs::heal::erasure_healer",
event = EVENT_HEAL_ERASURE_OBJECT_STATE, event = EVENT_HEAL_ERASURE_OBJECT_STATE,
@@ -877,6 +992,7 @@ impl ErasureSetHealer {
Err(err @ Error::TaskCancelled) | Err(err @ Error::TaskTimeout) => return Err(err), Err(err @ Error::TaskCancelled) | Err(err @ Error::TaskTimeout) => return Err(err),
Err(Error::TransientSkip { message }) => { Err(Error::TransientSkip { message }) => {
*skipped_objects += 1; *skipped_objects += 1;
bytes_processed = bytes_processed.saturating_add(object_size);
checkpoint_manager.add_skipped_object(key).await?; checkpoint_manager.add_skipped_object(key).await?;
demote_to_debug_when!(!take_failure_log_sample(&mut transient_skip_samples_logged), warn, target: "rustfs::heal::erasure_healer", { demote_to_debug_when!(!take_failure_log_sample(&mut transient_skip_samples_logged), warn, target: "rustfs::heal::erasure_healer", {
event = EVENT_HEAL_ERASURE_OBJECT_STATE, event = EVENT_HEAL_ERASURE_OBJECT_STATE,
@@ -893,6 +1009,7 @@ impl ErasureSetHealer {
} }
Err(err) => { Err(err) => {
*failed_objects += 1; *failed_objects += 1;
bytes_processed = bytes_processed.saturating_add(object_size);
checkpoint_manager.add_failed_object(key).await?; checkpoint_manager.add_failed_object(key).await?;
demote_to_debug_when!(!take_failure_log_sample(&mut failure_samples_logged), warn, target: "rustfs::heal::erasure_healer", { demote_to_debug_when!(!take_failure_log_sample(&mut failure_samples_logged), warn, target: "rustfs::heal::erasure_healer", {
event = EVENT_HEAL_ERASURE_OBJECT_STATE, event = EVENT_HEAL_ERASURE_OBJECT_STATE,
@@ -911,6 +1028,11 @@ impl ErasureSetHealer {
*processed_objects += 1; *processed_objects += 1;
completed_in_page += 1; completed_in_page += 1;
{
let mut progress = self.progress.write().await;
progress.set_current_object(Some(format!("{bucket}/{object}")));
progress.update_progress(*processed_objects, *successful_objects, *failed_objects, bytes_processed);
}
if completed_in_page.is_multiple_of(100) { if completed_in_page.is_multiple_of(100) {
checkpoint_manager.update_position(bucket_index, page_resume_index).await?; checkpoint_manager.update_position(bucket_index, page_resume_index).await?;
@@ -964,7 +1086,9 @@ impl ErasureSetHealer {
progress.objects_scanned = state.total_objects; progress.objects_scanned = state.total_objects;
progress.objects_healed = state.successful_objects; progress.objects_healed = state.successful_objects;
progress.objects_failed = state.failed_objects; progress.objects_failed = state.failed_objects;
progress.bytes_processed = 0; // set to 0 for now, can be extended later progress.bytes_processed = 0; // Resume state tracks object counts, not byte counters.
progress.start_time = UNIX_EPOCH.checked_add(Duration::from_secs(state.start_time));
progress.last_update_time = UNIX_EPOCH.checked_add(Duration::from_secs(state.last_update));
progress.set_current_object(state.current_object.clone()); progress.set_current_object(state.current_object.clone());
} }
} }
@@ -1135,13 +1259,15 @@ mod resume_loop_tests {
//! that emits programmable multi-version pages. These exercise the real loop //! that emits programmable multi-version pages. These exercise the real loop
//! logic (cursor seeding, per-version dedup, anti-loop guard, absence //! logic (cursor seeding, per-version dedup, anti-loop guard, absence
//! handling) — not merely a mock's own output. //! handling) — not merely a mock's own output.
use super::{ErasureSetHealer, target_outcomes_complete}; use super::{
ErasureSetHealer, NANOS_PER_SECOND, NEW_VERSION_SKIP_GRACE_SECS, should_skip_new_version, target_outcomes_complete,
};
use crate::heal::progress::HealProgress; use crate::heal::progress::HealProgress;
use crate::heal::resume::{ use crate::heal::resume::{
CheckpointManager, RESUME_CHECKPOINT_FILE, ReplacementTargetIdentity, ResumeDeleteFailure, ResumeManager, ResumeUtils, CheckpointManager, RESUME_CHECKPOINT_FILE, ReplacementTargetIdentity, ResumeDeleteFailure, ResumeManager, ResumeUtils,
compose_key, compose_key,
}; };
use crate::heal::storage::{DiskStatus, HealListItem, HealObjectInfo, HealStorageAPI}; use crate::heal::storage::{DiskStatus, HealLifecycleExpiryContext, HealListItem, HealObjectInfo, HealStorageAPI};
use crate::heal::storage_api::status::BucketInfo; use crate::heal::storage_api::status::BucketInfo;
use crate::heal::{ use crate::heal::{
BUCKET_META_PREFIX, DiskOption, DiskStore, EcstoreError, Endpoint, HealDiskExt as _, RUSTFS_META_BUCKET, new_disk, BUCKET_META_PREFIX, DiskOption, DiskStore, EcstoreError, Endpoint, HealDiskExt as _, RUSTFS_META_BUCKET, new_disk,
@@ -1149,7 +1275,7 @@ mod resume_loop_tests {
use crate::{Error, Result}; use crate::{Error, Result};
use rustfs_common::heal_channel::{HealOpts, HealRequestSource}; use rustfs_common::heal_channel::{HealOpts, HealRequestSource};
use rustfs_madmin::heal_commands::{HealDriveInfo, HealResultItem, Infos}; use rustfs_madmin::heal_commands::{HealDriveInfo, HealResultItem, Infos};
use std::collections::{HashMap, VecDeque}; use std::collections::{HashMap, HashSet, VecDeque};
use std::sync::atomic::{AtomicBool, Ordering}; use std::sync::atomic::{AtomicBool, Ordering};
use std::sync::{Arc, Mutex}; use std::sync::{Arc, Mutex};
use tempfile::TempDir; use tempfile::TempDir;
@@ -1160,10 +1286,37 @@ mod resume_loop_tests {
HealListItem { HealListItem {
name: name.to_string(), name: name.to_string(),
version_id: version.map(str::to_string), version_id: version.map(str::to_string),
mod_time_unix_nanos: None,
lifecycle_object_info: None,
is_delete_marker: delete_marker, is_delete_marker: delete_marker,
} }
} }
fn item_with_mod_time(name: &str, version: Option<&str>, mod_time_secs: u64) -> HealListItem {
HealListItem {
name: name.to_string(),
version_id: version.map(str::to_string),
mod_time_unix_nanos: Some(i128::from(mod_time_secs).saturating_mul(NANOS_PER_SECOND)),
lifecycle_object_info: None,
is_delete_marker: false,
}
}
#[test]
fn new_version_filter_respects_grace_boundary() {
let started_at = 1_700_000_000;
assert!(!should_skip_new_version(None, started_at));
assert!(!should_skip_new_version(
Some(i128::from(started_at + NEW_VERSION_SKIP_GRACE_SECS).saturating_mul(NANOS_PER_SECOND)),
started_at,
));
assert!(should_skip_new_version(
Some(i128::from(started_at + NEW_VERSION_SKIP_GRACE_SECS + 1).saturating_mul(NANOS_PER_SECOND)),
started_at,
));
}
#[test] #[test]
fn target_outcomes_require_each_requested_endpoint_once_and_ok() { fn target_outcomes_require_each_requested_endpoint_once_and_ok() {
let result = HealResultItem { let result = HealResultItem {
@@ -1246,8 +1399,10 @@ mod resume_loop_tests {
/// Target-specific physical readback evidence per `compose_key`; the /// Target-specific physical readback evidence per `compose_key`; the
/// fake models a healthy backend unless a test explicitly revokes it. /// fake models a healthy backend unless a test explicitly revokes it.
replacement_commit_evidence: Mutex<HashMap<String, ReplacementCommitEvidence>>, replacement_commit_evidence: Mutex<HashMap<String, ReplacementCommitEvidence>>,
lifecycle_expired: Mutex<HashSet<String>>,
/// every heal_object call recorded as (name, version_id) /// every heal_object call recorded as (name, version_id)
heal_calls: Mutex<Vec<(String, Option<String>)>>, heal_calls: Mutex<Vec<(String, Option<String>)>>,
list_include_lifecycle_object_info: Mutex<Vec<bool>>,
replacement_target_identity_sequences: Mutex<VecDeque<Vec<ReplacementTargetIdentity>>>, replacement_target_identity_sequences: Mutex<VecDeque<Vec<ReplacementTargetIdentity>>>,
fail_listing: AtomicBool, fail_listing: AtomicBool,
} }
@@ -1274,9 +1429,15 @@ mod resume_loop_tests {
.unwrap() .unwrap()
.insert(compose_key(name, version), ReplacementCommitEvidence::Error(message.to_string())); .insert(compose_key(name, version), ReplacementCommitEvidence::Error(message.to_string()));
} }
fn set_lifecycle_expired(&self, name: &str, version: Option<&str>) {
self.lifecycle_expired.lock().unwrap().insert(compose_key(name, version));
}
fn calls(&self) -> Vec<(String, Option<String>)> { fn calls(&self) -> Vec<(String, Option<String>)> {
self.heal_calls.lock().unwrap().clone() self.heal_calls.lock().unwrap().clone()
} }
fn list_include_lifecycle_object_info_calls(&self) -> Vec<bool> {
self.list_include_lifecycle_object_info.lock().unwrap().clone()
}
fn fail_listing(&self) { fn fail_listing(&self) {
self.fail_listing.store(true, Ordering::SeqCst); self.fail_listing.store(true, Ordering::SeqCst);
} }
@@ -1330,6 +1491,23 @@ mod resume_loop_tests {
async fn get_object_checksum(&self, _b: &str, _o: &str) -> Result<Option<String>> { async fn get_object_checksum(&self, _b: &str, _o: &str) -> Result<Option<String>> {
Ok(None) Ok(None)
} }
async fn load_heal_lifecycle_expiry_context(&self, _bucket: &str) -> Result<Option<HealLifecycleExpiryContext>> {
Ok((!self.lifecycle_expired.lock().unwrap().is_empty()).then(HealLifecycleExpiryContext::test))
}
async fn enqueue_heal_lifecycle_expiry(
&self,
_context: &HealLifecycleExpiryContext,
_bucket: &str,
object: &str,
version_id: Option<&str>,
_object_info: Option<&HealObjectInfo>,
) -> Result<bool> {
Ok(self
.lifecycle_expired
.lock()
.unwrap()
.contains(&compose_key(object, version_id)))
}
async fn heal_object( async fn heal_object(
&self, &self,
_bucket: &str, _bucket: &str,
@@ -1386,7 +1564,12 @@ mod resume_loop_tests {
_bucket: &str, _bucket: &str,
_prefix: &str, _prefix: &str,
continuation_token: Option<&str>, continuation_token: Option<&str>,
include_lifecycle_object_info: bool,
) -> Result<(Vec<HealListItem>, Option<String>, bool)> { ) -> Result<(Vec<HealListItem>, Option<String>, bool)> {
self.list_include_lifecycle_object_info
.lock()
.unwrap()
.push(include_lifecycle_object_info);
if self.fail_listing.load(Ordering::SeqCst) { if self.fail_listing.load(Ordering::SeqCst) {
return Err(Error::other("injected listing failure")); return Err(Error::other("injected listing failure"));
} }
@@ -1476,6 +1659,7 @@ mod resume_loop_tests {
/// Drive one bucket heal pass; returns (processed, successful, failed, skipped, result). /// Drive one bucket heal pass; returns (processed, successful, failed, skipped, result).
async fn run(env: &Env) -> (u64, u64, u64, u64, Result<()>) { async fn run(env: &Env) -> (u64, u64, u64, u64, Result<()>) {
let state = env.resume.get_state().await;
let mut current_object_index = 0usize; let mut current_object_index = 0usize;
let mut processed = 0u64; let mut processed = 0u64;
let mut successful = 0u64; let mut successful = 0u64;
@@ -1494,6 +1678,7 @@ mod resume_loop_tests {
&mut skipped, &mut skipped,
&env.resume, &env.resume,
&env.checkpoint, &env.checkpoint,
state.start_time,
) )
.await; .await;
(processed, successful, failed, skipped, result) (processed, successful, failed, skipped, result)
@@ -1559,6 +1744,7 @@ mod resume_loop_tests {
let mut successful = 0; let mut successful = 0;
let mut failed = 0; let mut failed = 0;
let mut skipped = 0; let mut skipped = 0;
let started_at = env.resume.get_state().await.start_time;
let error = healer let error = healer
.heal_bucket_with_resume( .heal_bucket_with_resume(
@@ -1572,6 +1758,7 @@ mod resume_loop_tests {
&mut skipped, &mut skipped,
&env.resume, &env.resume,
&env.checkpoint, &env.checkpoint,
started_at,
) )
.await .await
.expect_err("a remounted target must not begin a new page scan"); .expect_err("a remounted target must not begin a new page scan");
@@ -1641,6 +1828,109 @@ mod resume_loop_tests {
assert_eq!(skipped, 0); assert_eq!(skipped, 0);
} }
#[tokio::test]
async fn erasure_set_progress_accumulates_healed_object_bytes() {
let env = make_env().await;
env.storage.set_page(
None,
Page {
items: vec![item("first", Some("v1"), false), item("second", Some("v2"), false)],
next: None,
truncated: false,
},
);
env.storage.set_result(
"first",
Some("v1"),
HealResultItem {
object_size: 1024,
..Default::default()
},
);
env.storage.set_result(
"second",
Some("v2"),
HealResultItem {
object_size: 2048,
..Default::default()
},
);
let (processed, successful, failed, skipped, result) = run(&env).await;
result.expect("page heal should succeed");
assert_eq!(processed, 2);
assert_eq!(successful, 2);
assert_eq!(failed, 0);
assert_eq!(skipped, 0);
let progress = env.healer.progress.read().await;
assert_eq!(progress.objects_scanned, 2);
assert_eq!(progress.objects_healed, 2);
assert_eq!(progress.objects_failed, 0);
assert_eq!(progress.bytes_processed, 3072);
assert!(matches!(progress.current_object.as_deref(), Some("b/first" | "b/second")));
}
#[tokio::test]
async fn erasure_set_skips_versions_written_after_heal_started() {
let env = make_env().await;
let started_at = env.resume.get_state().await.start_time;
env.storage.set_page(
None,
Page {
items: vec![
item_with_mod_time("old", Some("v1"), started_at + NEW_VERSION_SKIP_GRACE_SECS),
item_with_mod_time("new", Some("v2"), started_at + NEW_VERSION_SKIP_GRACE_SECS + 1),
],
next: None,
truncated: false,
},
);
let (processed, successful, failed, skipped, result) = run(&env).await;
result.expect("page heal should succeed");
assert_eq!(processed, 2);
assert_eq!(successful, 1);
assert_eq!(failed, 0);
assert_eq!(skipped, 0);
assert_eq!(env.storage.calls(), vec![("old".to_string(), Some("v1".to_string()))]);
let progress = env.healer.progress.read().await;
assert_eq!(progress.skipped_new_versions, 1);
assert_eq!(progress.objects_scanned, 2);
assert_eq!(progress.objects_healed, 1);
assert_eq!(progress.objects_failed, 0);
}
#[tokio::test]
async fn erasure_set_skips_versions_queued_for_lifecycle_expiry() {
let env = make_env().await;
env.storage.set_page(
None,
Page {
items: vec![item("expired", Some("v1"), false), item("kept", Some("v2"), false)],
next: None,
truncated: false,
},
);
env.storage.set_lifecycle_expired("expired", Some("v1"));
let (processed, successful, failed, skipped, result) = run(&env).await;
result.expect("page heal should succeed");
assert_eq!(processed, 2);
assert_eq!(successful, 1);
assert_eq!(failed, 0);
assert_eq!(skipped, 0);
assert_eq!(env.storage.calls(), vec![("kept".to_string(), Some("v2".to_string()))]);
assert_eq!(env.storage.list_include_lifecycle_object_info_calls(), vec![true]);
let progress = env.healer.progress.read().await;
assert_eq!(progress.skipped_ilm_expired, 1);
assert_eq!(progress.objects_scanned, 2);
assert_eq!(progress.objects_healed, 1);
assert_eq!(progress.objects_failed, 0);
}
#[tokio::test] #[tokio::test]
async fn bucket_listing_failure_does_not_mark_set_completed() { async fn bucket_listing_failure_does_not_mark_set_completed() {
let env = make_env().await; let env = make_env().await;
+468 -62
View File
@@ -220,6 +220,11 @@ struct CompletedHealStatus {
result_items: Vec<HealResultItem>, result_items: Vec<HealResultItem>,
result_items_truncated: bool, result_items_truncated: bool,
completed_at: SystemTime, completed_at: SystemTime,
/// Sequence-stamped retained window, archived with the completion so
/// incremental consumers keep their cursor across the transition (HS-06).
seqed_items: Vec<(u64, HealResultItem)>,
next_seq: u64,
min_seq: u64,
} }
#[derive(Debug, Clone)] #[derive(Debug, Clone)]
@@ -240,6 +245,65 @@ pub struct HealTaskReport {
pub result_items: Vec<HealResultItem>, pub result_items: Vec<HealResultItem>,
pub result_items_truncated: bool, pub result_items_truncated: bool,
pub progress: Option<HealProgress>, pub progress: Option<HealProgress>,
/// Cursor for incremental consumption: sequence number of the next item
/// to be produced. `0` on reports from sources without sequencing.
pub next_seq: u64,
/// Oldest sequence still retained (`0` together with `next_seq` when
/// sequencing is unavailable).
pub min_seq: u64,
}
/// Report from a live task, honoring the client's incremental cursor.
async fn active_task_report(task: &HealTask, since: Option<u64>) -> HealTaskReport {
let window = task.get_result_items_since(since).await;
HealTaskReport {
status: task.get_status().await,
result_items: window.items,
// The legacy flag stays set once anything was evicted; a lagging
// incremental cursor additionally marks this response truncated so
// the client knows to restart from `min_seq`.
result_items_truncated: task.result_items_truncated() || window.lagged,
progress: Some(task.get_progress().await),
next_seq: window.next_seq,
min_seq: window.min_seq,
}
}
fn empty_task_report(status: HealTaskStatus) -> HealTaskReport {
HealTaskReport {
status,
result_items: Vec::new(),
result_items_truncated: false,
progress: None,
next_seq: 0,
min_seq: 0,
}
}
fn completed_task_report(completed: &CompletedHealStatus, since: Option<u64>) -> HealTaskReport {
let mut lagged = false;
let result_items = match since {
None => completed.result_items.clone(),
Some(cursor) => {
if cursor + 1 < completed.min_seq {
lagged = true;
}
completed
.seqed_items
.iter()
.filter(|(seq, _)| *seq > cursor)
.map(|(_, item)| item.clone())
.collect()
}
};
HealTaskReport {
status: completed.status.clone(),
result_items,
result_items_truncated: completed.result_items_truncated || lagged,
progress: None,
next_seq: completed.next_seq,
min_seq: completed.min_seq,
}
} }
#[derive(Debug, Clone, Copy, Default, PartialEq, Eq, serde::Deserialize, serde::Serialize)] #[derive(Debug, Clone, Copy, Default, PartialEq, Eq, serde::Deserialize, serde::Serialize)]
@@ -270,6 +334,8 @@ pub struct HealSourceCounts {
pub auto_heal: u64, pub auto_heal: u64,
pub internal: u64, pub internal: u64,
pub read_repair: u64, pub read_repair: u64,
#[serde(default)]
pub mrf: u64,
} }
impl HealSourceCounts { impl HealSourceCounts {
@@ -280,6 +346,7 @@ impl HealSourceCounts {
HealRequestSource::AutoHeal => self.auto_heal += 1, HealRequestSource::AutoHeal => self.auto_heal += 1,
HealRequestSource::Internal => self.internal += 1, HealRequestSource::Internal => self.internal += 1,
HealRequestSource::ReadRepair => self.read_repair += 1, HealRequestSource::ReadRepair => self.read_repair += 1,
HealRequestSource::Mrf => self.mrf += 1,
} }
} }
} }
@@ -528,6 +595,11 @@ impl PriorityHealQueue {
self.dedup_keys.contains_key(&key) self.dedup_keys.contains_key(&key)
} }
/// Iterate queued requests (used by the admin overlap check).
fn requests(&self) -> impl Iterator<Item = &HealRequest> {
self.heap.iter().map(|item| &item.request)
}
fn contains_request_id(&self, request_id: &str) -> bool { fn contains_request_id(&self, request_id: &str) -> bool {
self.heap.iter().any(|item| item.request.id == request_id) self.heap.iter().any(|item| item.request.id == request_id)
} }
@@ -686,6 +758,80 @@ fn recoverable_heal_retry_delay(retry_attempt: u32) -> Duration {
} }
/// Heal config /// Heal config
/// HS-06 admin overlap policy.
#[derive(Debug, Clone, Copy, PartialEq, Eq, Default)]
pub enum HealOverlapPolicy {
/// Default: overlapping admin starts merge into the existing task
/// (today's dedup semantics).
#[default]
Merge,
/// Return a typed already-running / overlapping-paths rejection like
/// madmin's ErrHealAlreadyRunning / ErrHealOverlappingPaths.
MinioError,
}
/// Path view of a heal type for overlap comparison: a bucket plus a
/// prefix/object path inside it (`None` bucket = cluster-wide, overlaps
/// everything).
fn heal_type_path_view(heal_type: &HealType) -> (Option<&str>, &str) {
match heal_type {
HealType::Cluster => (None, ""),
HealType::Bucket { bucket } => (Some(bucket), ""),
HealType::Prefix { bucket, prefix } => (Some(bucket), prefix),
HealType::Object { bucket, object, .. }
| HealType::Metadata { bucket, object }
| HealType::ECDecode { bucket, object, .. } => (Some(bucket), object),
// MRF/MetaPath heal keys on a meta path; treat the whole set of
// buckets as one namespace so it only overlaps itself exactly.
HealType::MRF { meta_path } => (Some("\u{0}mrf"), meta_path),
// Erasure-set heal: the set id is the overlap dimension.
HealType::ErasureSet { set_disk_id, .. } => (Some("\u{0}set"), set_disk_id),
}
}
/// How two heal paths relate for the admin overlap check (HS-06).
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
enum OverlapVerdict {
/// Distinct targets: no conflict.
Disjoint,
/// Same target: an identical heal is already in flight.
SameTarget,
/// One target contains the other.
Overlapping,
}
fn prefix_paths_overlap(a: &str, b: &str) -> OverlapVerdict {
if a == b {
return OverlapVerdict::SameTarget;
}
if a.is_empty() || b.is_empty() || a.starts_with(b) || b.starts_with(a) {
return OverlapVerdict::Overlapping;
}
OverlapVerdict::Disjoint
}
fn heal_types_overlap(left: &HealType, right: &HealType) -> OverlapVerdict {
let (left_bucket, left_path) = heal_type_path_view(left);
let (right_bucket, right_path) = heal_type_path_view(right);
match (left_bucket, right_bucket) {
// Cluster-wide overlaps everything (but an exact cluster match is
// SameTarget).
(None, _) | (_, None) => {
if matches!(left, HealType::Cluster) && matches!(right, HealType::Cluster) {
OverlapVerdict::SameTarget
} else {
OverlapVerdict::Overlapping
}
}
(Some(lb), Some(rb)) => {
if lb != rb {
return OverlapVerdict::Disjoint;
}
prefix_paths_overlap(left_path, right_path)
}
}
}
#[derive(Debug, Clone)] #[derive(Debug, Clone)]
pub struct HealConfig { pub struct HealConfig {
/// Whether to enable auto heal /// Whether to enable auto heal
@@ -706,6 +852,9 @@ pub struct HealConfig {
pub low_priority_drop_when_full: bool, pub low_priority_drop_when_full: bool,
/// Whether notify-driven scheduler wakeups are enabled. /// Whether notify-driven scheduler wakeups are enabled.
pub event_driven_scheduler_enable: bool, pub event_driven_scheduler_enable: bool,
/// How admin heal starts behave on path overlap (HS-06): merge into the
/// existing task (default) or return a typed already-running rejection.
pub overlap_policy: HealOverlapPolicy,
/// Whether per-set bulkhead scheduling is enabled. /// Whether per-set bulkhead scheduling is enabled.
pub set_bulkhead_enable: bool, pub set_bulkhead_enable: bool,
/// Whether erasure-set page parallelism is enabled. /// Whether erasure-set page parallelism is enabled.
@@ -754,6 +903,14 @@ impl Default for HealConfig {
rustfs_config::ENV_HEAL_EVENT_DRIVEN_SCHEDULER_ENABLE, rustfs_config::ENV_HEAL_EVENT_DRIVEN_SCHEDULER_ENABLE,
rustfs_config::DEFAULT_HEAL_EVENT_DRIVEN_SCHEDULER_ENABLE, rustfs_config::DEFAULT_HEAL_EVENT_DRIVEN_SCHEDULER_ENABLE,
); );
let overlap_policy =
match rustfs_utils::get_env_str(rustfs_config::ENV_HEAL_OVERLAP_POLICY, rustfs_config::DEFAULT_HEAL_OVERLAP_POLICY)
.to_lowercase()
.as_str()
{
"minio_error" => HealOverlapPolicy::MinioError,
_ => HealOverlapPolicy::Merge,
};
let set_bulkhead_enable = rustfs_utils::get_env_bool( let set_bulkhead_enable = rustfs_utils::get_env_bool(
rustfs_config::ENV_HEAL_SET_BULKHEAD_ENABLE, rustfs_config::ENV_HEAL_SET_BULKHEAD_ENABLE,
rustfs_config::DEFAULT_HEAL_SET_BULKHEAD_ENABLE, rustfs_config::DEFAULT_HEAL_SET_BULKHEAD_ENABLE,
@@ -790,6 +947,7 @@ impl Default for HealConfig {
low_priority_merge_enable, low_priority_merge_enable,
low_priority_drop_when_full, low_priority_drop_when_full,
event_driven_scheduler_enable, event_driven_scheduler_enable,
overlap_policy,
set_bulkhead_enable, set_bulkhead_enable,
page_parallel_enable, page_parallel_enable,
mainline_throttle_enable, mainline_throttle_enable,
@@ -1756,6 +1914,50 @@ impl HealManager {
request: HealRequest, request: HealRequest,
preserve_alias: bool, preserve_alias: bool,
) -> Result<HealAdmissionReceipt> { ) -> Result<HealAdmissionReceipt> {
// HS-06 forceStart semantics (admin only): MinIO stops the old task
// first and then starts the new one. Cancel any active admin task
// overlapping this request's path before entering admission, so the
// fresh task is never merged into the one being replaced.
if request.source == HealRequestSource::Admin && request.force_start {
let overlapping: Vec<String> = {
let active_heals = self.active_heals.lock().await;
active_heals
.iter()
.filter(|(task_id, task)| {
task.source == HealRequestSource::Admin
&& heal_types_overlap(&request.heal_type, &task.heal_type) != OverlapVerdict::Disjoint
&& *task_id != &request.id
})
.map(|(task_id, _)| task_id.clone())
.collect()
};
for task_id in overlapping {
match self.cancel_task(&task_id).await {
Ok(_) => info!(
target: "rustfs::heal::manager",
event = EVENT_HEAL_QUEUE_ADMISSION,
component = LOG_COMPONENT_HEAL,
subsystem = LOG_SUBSYSTEM_MANAGER,
request_id = %request.id,
cancelled_task_id = %task_id,
result = "force_start_cancelled_overlap",
"Admin forceStart cancelled an overlapping heal task"
),
Err(err) => warn!(
target: "rustfs::heal::manager",
event = EVENT_HEAL_QUEUE_ADMISSION,
component = LOG_COMPONENT_HEAL,
subsystem = LOG_SUBSYSTEM_MANAGER,
request_id = %request.id,
cancelled_task_id = %task_id,
error = %err,
result = "force_start_cancel_failed",
"Admin forceStart failed to cancel an overlapping heal task"
),
}
}
}
let config = self.config.read().await; let config = self.config.read().await;
let dedup_key = PriorityHealQueue::make_dedup_key(&request); let dedup_key = PriorityHealQueue::make_dedup_key(&request);
@@ -1778,7 +1980,15 @@ impl HealManager {
.or_else(|| retrying_heal_for_dedup_key(&retrying_heals, &dedup_key).map(|(task_id, _)| (task_id, "retrying"))) .or_else(|| retrying_heal_for_dedup_key(&retrying_heals, &dedup_key).map(|(task_id, _)| (task_id, "retrying")))
}); });
if let Some((merged_task_id, duplicate_state)) = duplicate.flatten() { if let Some((merged_task_id, duplicate_state)) = duplicate.flatten() {
let admission = Self::duplicate_admission_for_request(&request, &config); // HS-06: under the minio_error overlap policy an exact duplicate
// admin start reports the typed AlreadyRunning rejection instead
// of the silent merge (MinIO's ErrHealAlreadyRunning).
let admission =
if request.source == HealRequestSource::Admin && config.overlap_policy == HealOverlapPolicy::MinioError {
HealAdmissionResult::Dropped(HealAdmissionDropReason::AlreadyRunning)
} else {
Self::duplicate_admission_for_request(&request, &config)
};
drop(retrying_heals); drop(retrying_heals);
drop(queue); drop(queue);
drop(active_heals); drop(active_heals);
@@ -1824,6 +2034,62 @@ impl HealManager {
}); });
} }
// HS-06 typed overlap rejection (admin only, minio_error policy):
// paths containing or contained by an active/queued task reject with
// AlreadyRunning / OverlappingPaths instead of merging. Exact
// duplicates already merged above; scanner/autoheal/read-repair
// sources never take this path.
if request.source == HealRequestSource::Admin && config.overlap_policy == HealOverlapPolicy::MinioError {
let mut rejection = None;
for (task_id, task) in active_heals.iter() {
match heal_types_overlap(&request.heal_type, &task.heal_type) {
OverlapVerdict::SameTarget => {
rejection = Some((HealAdmissionDropReason::AlreadyRunning, task_id.clone()));
break;
}
OverlapVerdict::Overlapping => {
rejection = Some((HealAdmissionDropReason::OverlappingPaths, task_id.clone()));
}
OverlapVerdict::Disjoint => {}
}
}
if rejection.is_none() {
for queued in queue.requests() {
match heal_types_overlap(&request.heal_type, &queued.heal_type) {
OverlapVerdict::SameTarget => {
rejection = Some((HealAdmissionDropReason::AlreadyRunning, queued.id.clone()));
break;
}
OverlapVerdict::Overlapping => {
rejection = Some((HealAdmissionDropReason::OverlappingPaths, queued.id.clone()));
}
OverlapVerdict::Disjoint => {}
}
}
}
if let Some((reason, overlap_task_id)) = rejection {
drop(retrying_heals);
drop(queue);
drop(active_heals);
Self::record_admission_metric(request.source, HealAdmissionResult::Dropped(reason), "overlap_rejected");
warn!(
target: "rustfs::heal::manager",
event = EVENT_HEAL_QUEUE_ADMISSION,
component = LOG_COMPONENT_HEAL,
subsystem = LOG_SUBSYSTEM_MANAGER,
request_id = %request.id,
overlap_task_id = %overlap_task_id,
reason = reason.as_str(),
result = "overlap_rejected",
"Admin heal start rejected by overlap policy"
);
return Ok(HealAdmissionReceipt {
result: HealAdmissionResult::Dropped(reason),
task_id: overlap_task_id,
});
}
}
let mut task_id = request.id.clone(); let mut task_id = request.id.clone();
let admission = Self::admit_request_to_queue(&mut queue, request, &config, "submit"); let admission = Self::admit_request_to_queue(&mut queue, request, &config, "submit");
if admission == HealAdmissionResult::Merged if admission == HealAdmissionResult::Merged
@@ -1896,28 +2162,25 @@ impl HealManager {
} }
pub async fn get_task_report(&self, task_id: &str) -> Result<HealTaskReport> { pub async fn get_task_report(&self, task_id: &str) -> Result<HealTaskReport> {
self.get_task_report_since(task_id, None).await
}
/// Incremental variant of [`Self::get_task_report`] (HS-06): `since` is
/// the client's last seen sequence number; `None` keeps the legacy
/// full-snapshot semantics.
pub async fn get_task_report_since(&self, task_id: &str, since: Option<u64>) -> Result<HealTaskReport> {
let canonical_task_id = self.canonical_task_id(task_id).await; let canonical_task_id = self.canonical_task_id(task_id).await;
{ {
let active_heals = self.active_heals.lock().await; let active_heals = self.active_heals.lock().await;
if let Some(task) = active_heals.get(&canonical_task_id) { if let Some(task) = active_heals.get(&canonical_task_id) {
return Ok(HealTaskReport { return Ok(active_task_report(task, since).await);
status: task.get_status().await,
result_items: task.get_result_items().await,
result_items_truncated: task.result_items_truncated(),
progress: Some(task.get_progress().await),
});
} }
} }
{ {
let retrying_heals = self.retrying_heals.lock().await; let retrying_heals = self.retrying_heals.lock().await;
if let Some(retrying) = retrying_heals.get(&canonical_task_id) { if let Some(retrying) = retrying_heals.get(&canonical_task_id) {
return Ok(HealTaskReport { return Ok(empty_task_report(retrying.status()));
status: retrying.status(),
result_items: Vec::new(),
result_items_truncated: false,
progress: None,
});
} }
} }
@@ -1927,36 +2190,21 @@ impl HealManager {
if let Some(completed) = completed_heals.get(&canonical_task_id) if let Some(completed) = completed_heals.get(&canonical_task_id)
&& completed_status_is_retrying(&completed.status) && completed_status_is_retrying(&completed.status)
{ {
return Ok(HealTaskReport { return Ok(completed_task_report(completed, since));
status: completed.status.clone(),
result_items: completed.result_items.clone(),
result_items_truncated: completed.result_items_truncated,
progress: None,
});
} }
} }
{ {
let queue = self.heal_queue.lock().await; let queue = self.heal_queue.lock().await;
if queue.contains_request_id(&canonical_task_id) { if queue.contains_request_id(&canonical_task_id) {
return Ok(HealTaskReport { return Ok(empty_task_report(HealTaskStatus::Pending));
status: HealTaskStatus::Pending,
result_items: Vec::new(),
result_items_truncated: false,
progress: None,
});
} }
} }
let mut completed_heals = self.completed_heals.lock().await; let mut completed_heals = self.completed_heals.lock().await;
prune_completed_heal_statuses(&mut completed_heals); prune_completed_heal_statuses(&mut completed_heals);
if let Some(completed) = completed_heals.get(&canonical_task_id) { if let Some(completed) = completed_heals.get(&canonical_task_id) {
return Ok(HealTaskReport { return Ok(completed_task_report(completed, since));
status: completed.status.clone(),
result_items: completed.result_items.clone(),
result_items_truncated: completed.result_items_truncated,
progress: None,
});
} }
Err(Error::TaskNotFound { Err(Error::TaskNotFound {
@@ -1965,18 +2213,23 @@ impl HealManager {
} }
pub async fn get_task_report_for_path(&self, heal_path: &str, task_id: &str) -> Result<HealTaskReport> { pub async fn get_task_report_for_path(&self, heal_path: &str, task_id: &str) -> Result<HealTaskReport> {
self.get_task_report_for_path_since(heal_path, task_id, None).await
}
/// Incremental variant of [`Self::get_task_report_for_path`] (HS-06).
pub async fn get_task_report_for_path_since(
&self,
heal_path: &str,
task_id: &str,
since: Option<u64>,
) -> Result<HealTaskReport> {
let canonical_task_id = self.canonical_task_id(task_id).await; let canonical_task_id = self.canonical_task_id(task_id).await;
{ {
let active_heals = self.active_heals.lock().await; let active_heals = self.active_heals.lock().await;
if let Some(task) = active_heals.get(&canonical_task_id) if let Some(task) = active_heals.get(&canonical_task_id)
&& heal_type_matches_path(&task.heal_type, heal_path) && heal_type_matches_path(&task.heal_type, heal_path)
{ {
return Ok(HealTaskReport { return Ok(active_task_report(task, since).await);
status: task.get_status().await,
result_items: task.get_result_items().await,
result_items_truncated: task.result_items_truncated(),
progress: Some(task.get_progress().await),
});
} }
} }
@@ -1985,12 +2238,7 @@ impl HealManager {
if let Some(retrying) = retrying_heals.get(&canonical_task_id) if let Some(retrying) = retrying_heals.get(&canonical_task_id)
&& heal_type_matches_path(&retrying.request.heal_type, heal_path) && heal_type_matches_path(&retrying.request.heal_type, heal_path)
{ {
return Ok(HealTaskReport { return Ok(empty_task_report(retrying.status()));
status: retrying.status(),
result_items: Vec::new(),
result_items_truncated: false,
progress: None,
});
} }
} }
@@ -2001,24 +2249,14 @@ impl HealManager {
&& heal_type_matches_path(&completed.heal_type, heal_path) && heal_type_matches_path(&completed.heal_type, heal_path)
&& completed_status_is_retrying(&completed.status) && completed_status_is_retrying(&completed.status)
{ {
return Ok(HealTaskReport { return Ok(completed_task_report(completed, since));
status: completed.status.clone(),
result_items: completed.result_items.clone(),
result_items_truncated: completed.result_items_truncated,
progress: None,
});
} }
} }
{ {
let queue = self.heal_queue.lock().await; let queue = self.heal_queue.lock().await;
if queue.contains_request_id_matching_path(&canonical_task_id, heal_path) { if queue.contains_request_id_matching_path(&canonical_task_id, heal_path) {
return Ok(HealTaskReport { return Ok(empty_task_report(HealTaskStatus::Pending));
status: HealTaskStatus::Pending,
result_items: Vec::new(),
result_items_truncated: false,
progress: None,
});
} }
} }
@@ -2028,12 +2266,7 @@ impl HealManager {
if let Some(completed) = completed_heals.get(&canonical_task_id) if let Some(completed) = completed_heals.get(&canonical_task_id)
&& heal_type_matches_path(&completed.heal_type, heal_path) && heal_type_matches_path(&completed.heal_type, heal_path)
{ {
return Ok(HealTaskReport { return Ok(completed_task_report(completed, since));
status: completed.status.clone(),
result_items: completed.result_items.clone(),
result_items_truncated: completed.result_items_truncated,
progress: None,
});
} }
} }
@@ -2385,8 +2618,27 @@ impl HealManager {
snapshot.objects_scanned = snapshot.objects_scanned.saturating_add(progress.objects_scanned); snapshot.objects_scanned = snapshot.objects_scanned.saturating_add(progress.objects_scanned);
snapshot.objects_healed = snapshot.objects_healed.saturating_add(progress.objects_healed); snapshot.objects_healed = snapshot.objects_healed.saturating_add(progress.objects_healed);
snapshot.objects_failed = snapshot.objects_failed.saturating_add(progress.objects_failed); snapshot.objects_failed = snapshot.objects_failed.saturating_add(progress.objects_failed);
snapshot.skipped_new_versions = snapshot.skipped_new_versions.saturating_add(progress.skipped_new_versions);
snapshot.skipped_ilm_expired = snapshot.skipped_ilm_expired.saturating_add(progress.skipped_ilm_expired);
snapshot.objects_total_count = snapshot.objects_total_count.saturating_add(progress.objects_total_count);
snapshot.objects_total_size = snapshot.objects_total_size.saturating_add(progress.objects_total_size);
snapshot.bytes_processed = snapshot.bytes_processed.saturating_add(progress.bytes_processed); snapshot.bytes_processed = snapshot.bytes_processed.saturating_add(progress.bytes_processed);
snapshot.start_time = match (snapshot.start_time, progress.start_time) {
(Some(current), Some(next)) => Some(current.min(next)),
(None, next) => next,
(current, None) => current,
};
snapshot.last_update_time = match (snapshot.last_update_time, progress.last_update_time) {
(Some(current), Some(next)) => Some(current.max(next)),
(None, next) => next,
(current, None) => current,
};
if progress.current_object.is_some() {
snapshot.current_object = progress.current_object;
}
} }
snapshot.refresh_progress_percentage();
snapshot.refresh_estimated_completion_time();
Some(snapshot) Some(snapshot)
} }
@@ -3208,12 +3460,17 @@ impl HealManager {
} else { } else {
completed_task.get_status().await completed_task.get_status().await
}; };
let completed_progress = completed_task.get_progress().await;
let final_window = completed_task.get_result_items_since(None).await;
let completed_status_entry = CompletedHealStatus { let completed_status_entry = CompletedHealStatus {
heal_type: completed_task.heal_type.clone(), heal_type: completed_task.heal_type.clone(),
status: completed_status.clone(), status: completed_status.clone(),
result_items: completed_task.get_result_items().await, result_items: final_window.items.clone(),
result_items_truncated: completed_task.result_items_truncated(), result_items_truncated: completed_task.result_items_truncated(),
completed_at: SystemTime::now(), completed_at: SystemTime::now(),
seqed_items: completed_task.get_seqed_result_items().await,
next_seq: final_window.next_seq,
min_seq: final_window.min_seq,
}; };
let mut completed_heals_guard = completed_heals_clone.lock().await; let mut completed_heals_guard = completed_heals_clone.lock().await;
prune_completed_heal_statuses(&mut completed_heals_guard); prune_completed_heal_statuses(&mut completed_heals_guard);
@@ -3223,6 +3480,7 @@ impl HealManager {
match completed_status { match completed_status {
HealTaskStatus::Completed => { HealTaskStatus::Completed => {
stats.update_task_completion(true); stats.update_task_completion(true);
stats.add_healed_objects(completed_progress.objects_healed, completed_progress.bytes_processed);
} }
HealTaskStatus::Retrying { .. } => {} HealTaskStatus::Retrying { .. } => {}
_ => { _ => {
@@ -3749,6 +4007,7 @@ mod tests {
_bucket: &str, _bucket: &str,
_prefix: &str, _prefix: &str,
_continuation_token: Option<&str>, _continuation_token: Option<&str>,
_include_lifecycle_object_info: bool,
) -> Result<(Vec<crate::heal::storage::HealListItem>, Option<String>, bool)> { ) -> Result<(Vec<crate::heal::storage::HealListItem>, Option<String>, bool)> {
Ok((Vec::new(), None, false)) Ok((Vec::new(), None, false))
} }
@@ -4983,6 +5242,9 @@ mod tests {
}, },
result_items: Vec::new(), result_items: Vec::new(),
result_items_truncated: false, result_items_truncated: false,
seqed_items: Vec::new(),
next_seq: 0,
min_seq: 0,
completed_at: SystemTime::now(), completed_at: SystemTime::now(),
}, },
); );
@@ -5264,6 +5526,136 @@ mod tests {
assert_eq!(snapshot.queued_by_source.internal, 0); assert_eq!(snapshot.queued_by_source.internal, 0);
} }
// HS-06 (backlog#1870): overlap policy + forceStart semantics.
fn manager_with_policy(policy: HealOverlapPolicy) -> HealManager {
let storage: Arc<dyn HealStorageAPI> = Arc::new(MockStorage);
HealManager::new(
storage,
Some(HealConfig {
overlap_policy: policy,
..Default::default()
}),
)
}
fn admin_prefix_request(bucket: &str, prefix: &str) -> HealRequest {
let mut request = HealRequest::new(
HealType::Prefix {
bucket: bucket.to_string(),
prefix: prefix.to_string(),
},
HealOptions::default(),
HealPriority::Normal,
);
request.source = HealRequestSource::Admin;
request
}
async fn insert_active_task(manager: &HealManager, request: HealRequest) -> String {
let task = Arc::new(HealTask::from_request(request, manager.storage.clone()));
let task_id = task.id.clone();
manager.active_heals.lock().await.insert(task_id.clone(), task);
task_id
}
#[tokio::test]
async fn overlap_policy_minio_error_rejects_same_and_containing_paths() {
let manager = manager_with_policy(HealOverlapPolicy::MinioError);
insert_active_task(&manager, admin_prefix_request("bucket-a", "logs/")).await;
// Same target: typed AlreadyRunning.
let same = manager
.submit_heal_request(admin_prefix_request("bucket-a", "logs/"))
.await
.expect("admission must decide");
assert_eq!(
same,
HealAdmissionResult::Dropped(HealAdmissionDropReason::AlreadyRunning),
"an identical target must reject with already-running"
);
// Contained path: typed OverlappingPaths.
let nested = manager
.submit_heal_request(admin_prefix_request("bucket-a", "logs/app/"))
.await
.expect("admission must decide");
assert_eq!(
nested,
HealAdmissionResult::Dropped(HealAdmissionDropReason::OverlappingPaths),
"a path inside the active task's path must reject with overlapping-paths"
);
// Containing path (bucket-wide vs nested active): also overlapping.
let wide = manager
.submit_heal_request(admin_prefix_request("bucket-a", ""))
.await
.expect("admission must decide");
assert_eq!(
wide,
HealAdmissionResult::Dropped(HealAdmissionDropReason::OverlappingPaths),
"a bucket-wide start overlapping a nested active heal must reject"
);
// Disjoint bucket: unaffected.
let disjoint = manager
.submit_heal_request(admin_prefix_request("bucket-b", "logs/"))
.await
.expect("admission must decide");
assert_eq!(disjoint, HealAdmissionResult::Accepted);
}
#[tokio::test]
async fn overlap_policy_default_merge_keeps_today_semantics() {
let manager = manager_with_policy(HealOverlapPolicy::Merge);
insert_active_task(&manager, admin_prefix_request("bucket-a", "logs/")).await;
// Different-dedup-key overlap still merges under the default policy:
// the nested path dedups to its own key but nothing rejects it.
let nested = manager
.submit_heal_request(admin_prefix_request("bucket-a", "logs/app/"))
.await
.expect("admission must decide");
assert_eq!(nested, HealAdmissionResult::Accepted, "default policy must not reject overlaps");
// Non-admin sources never get overlap rejections even under minio_error.
let manager = manager_with_policy(HealOverlapPolicy::MinioError);
insert_active_task(&manager, admin_prefix_request("bucket-a", "logs/")).await;
let mut scanner_request = admin_prefix_request("bucket-a", "logs/app/");
scanner_request.source = HealRequestSource::Scanner;
let admitted = manager
.submit_heal_request(scanner_request)
.await
.expect("admission must decide");
assert_eq!(admitted, HealAdmissionResult::Accepted, "scanner sources must never be overlap-rejected");
}
#[tokio::test]
async fn admin_force_start_cancels_overlapping_active_task_first() {
let manager = manager_with_policy(HealOverlapPolicy::Merge);
let old_id = insert_active_task(&manager, admin_prefix_request("bucket-a", "logs/")).await;
let mut replacement = admin_prefix_request("bucket-a", "logs/");
replacement.force_start = true;
let receipt = manager
.submit_heal_request_with_receipt(replacement)
.await
.expect("force-start submission must decide");
assert!(receipt.result.is_admitted(), "the new task must be admitted (Accepted or Merged)");
let old_task_gone = {
let active_heals = manager.active_heals.lock().await;
!active_heals.contains_key(&old_id)
};
assert!(
old_task_gone,
"the overlapping admin task must be cancelled (removed from the active table) before the new one starts"
);
assert!(
matches!(manager.get_task_status(&old_id).await, Err(Error::TaskNotFound { .. })),
"a cancelled task must no longer resolve as an active heal"
);
}
#[tokio::test] #[tokio::test]
async fn test_operations_snapshot_counts_active_by_source_and_priority() { async fn test_operations_snapshot_counts_active_by_source_and_priority() {
let storage: Arc<dyn HealStorageAPI> = Arc::new(MockStorage); let storage: Arc<dyn HealStorageAPI> = Arc::new(MockStorage);
@@ -5396,6 +5788,8 @@ mod tests {
)); ));
{ {
let mut progress = first.progress.write().await; let mut progress = first.progress.write().await;
progress.start_time = Some(SystemTime::now() - Duration::from_secs(20));
progress.set_total_baseline(12, 8192);
progress.update_progress(7, 3, 1, 4096); progress.update_progress(7, 3, 1, 4096);
} }
@@ -5405,6 +5799,8 @@ mod tests {
)); ));
{ {
let mut progress = second.progress.write().await; let mut progress = second.progress.write().await;
progress.start_time = Some(SystemTime::now() - Duration::from_secs(10));
progress.set_total_baseline(8, 4096);
progress.update_progress(11, 5, 2, 2048); progress.update_progress(11, 5, 2, 2048);
} }
@@ -5419,7 +5815,11 @@ mod tests {
assert_eq!(progress.objects_scanned, 18); assert_eq!(progress.objects_scanned, 18);
assert_eq!(progress.objects_healed, 8); assert_eq!(progress.objects_healed, 8);
assert_eq!(progress.objects_failed, 3); assert_eq!(progress.objects_failed, 3);
assert_eq!(progress.objects_total_count, 20);
assert_eq!(progress.objects_total_size, 12288);
assert_eq!(progress.bytes_processed, 6144); assert_eq!(progress.bytes_processed, 6144);
assert!((progress.progress_percentage - 50.0).abs() < 0.001);
assert!(progress.estimated_completion_time.is_some());
} }
#[tokio::test] #[tokio::test]
@@ -5558,6 +5958,9 @@ mod tests {
status: HealTaskStatus::Completed, status: HealTaskStatus::Completed,
result_items: Vec::new(), result_items: Vec::new(),
result_items_truncated: false, result_items_truncated: false,
seqed_items: Vec::new(),
next_seq: 0,
min_seq: 0,
completed_at: SystemTime::now(), completed_at: SystemTime::now(),
}, },
); );
@@ -5592,6 +5995,9 @@ mod tests {
..Default::default() ..Default::default()
}], }],
result_items_truncated: true, result_items_truncated: true,
seqed_items: Vec::new(),
next_seq: 0,
min_seq: 0,
completed_at: SystemTime::now(), completed_at: SystemTime::now(),
}, },
); );
+1
View File
@@ -16,6 +16,7 @@ pub mod channel;
pub mod erasure_healer; pub mod erasure_healer;
pub mod event; pub mod event;
pub mod manager; pub mod manager;
pub mod mrf_queue;
pub mod progress; pub mod progress;
pub(crate) mod replacement_readiness; pub(crate) mod replacement_readiness;
pub mod resume; pub mod resume;
+682
View File
@@ -0,0 +1,682 @@
// Copyright 2024 RustFS Team
//
// Licensed under the Apache License, Version 2.0 (the "License");
// you may not use this file except in compliance with the License.
// You may obtain a copy of the License at
//
// http://www.apache.org/licenses/LICENSE-2.0
//
// Unless required by applicable law or agreed to in writing, software
// distributed under the License is distributed on an "AS IS" BASIS,
// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
// See the License for the specific language governing permissions and
// limitations under the License.
//! Mission Repair Feed (MRF) queue, journal, and consumer.
//!
//! Intents arriving on the global channel (see `rustfs_common::mrf_channel`)
//! are buffered in a bounded in-memory queue, translated into prioritized
//! heal requests, and — while they are not yet accepted by the heal manager —
//! mirrored into a durable journal so a crash or restart can replay them.
//! This is the RustFS counterpart of MinIO's `.heal/mrf/list.bin` replay,
//! layered on top of (not replacing) read-repair and scanner heal.
//!
//! Durability model: the journal is a snapshot of the *unaccepted* pending
//! set, rewritten on a group-commit cadence (every flush interval or flush
//! threshold new intents). A rewrite is atomic at the record level only — a
//! torn tail simply truncates during replay because every record carries its
//! own CRC32. Losing the last flush window (≤500 ms) is acceptable: replayed
//! duplicates are merged by the manager's dedup key, and read-repair remains
//! the safety net.
use super::{DiskStore, HealDiskExt as _, local_disk_map_read};
use crate::heal::manager::HealManager;
use metrics::{counter, gauge};
use rustfs_common::heal_channel::{HealAdmissionDropReason, HealAdmissionResult};
use rustfs_common::mrf_channel::{MRF_MAX_ATTEMPTS, MrfIntent};
use std::collections::VecDeque;
use std::sync::Arc;
use std::time::Duration;
use tokio::sync::mpsc;
use uuid::Uuid;
use crate::heal::task::{HealOptions, HealPriority, HealRequest, HealType};
/// Journal location inside the metadata bucket, following the resume-state
/// layout.
pub(crate) const MRF_JOURNAL_PATH: &str = "buckets/.heal/mrf/journal.bin";
/// Record format tag.
const MRF_JOURNAL_FORMAT: u8 = 1;
/// Record layout version.
const MRF_JOURNAL_VERSION: u8 = 1;
/// Fixed header size: format, version, kind, attempts, enqueued_at_ms,
/// has_version flag.
const MRF_RECORD_FIXED_HEAD: usize = 1 + 1 + 1 + 1 + 8 + 1;
#[derive(Debug, Clone)]
pub(crate) struct MrfConsumerConfig {
/// In-memory queue capacity in intents.
pub queue_capacity: usize,
/// Journal byte budget; a pending snapshot above this bound is rejected
/// oldest-first so the journal can never grow unbounded.
pub journal_max_bytes: usize,
/// How many journal intents to re-arm per replay round.
pub replay_batch: usize,
/// Group-commit cadence for the journal snapshot.
pub flush_interval: Duration,
/// New intents between flushes that force an early snapshot.
pub flush_threshold: usize,
/// Backoff after the heal manager reports a full admission.
pub admission_backoff: Duration,
}
impl Default for MrfConsumerConfig {
fn default() -> Self {
Self {
queue_capacity: rustfs_utils::get_env_usize(
rustfs_config::ENV_HEAL_MRF_QUEUE_SIZE,
rustfs_config::DEFAULT_HEAL_MRF_QUEUE_SIZE,
),
journal_max_bytes: rustfs_utils::get_env_usize(
rustfs_config::ENV_HEAL_MRF_JOURNAL_MAX_BYTES,
rustfs_config::DEFAULT_HEAL_MRF_JOURNAL_MAX_BYTES,
),
replay_batch: rustfs_utils::get_env_usize(
rustfs_config::ENV_HEAL_MRF_REPLAY_BATCH,
rustfs_config::DEFAULT_HEAL_MRF_REPLAY_BATCH,
),
flush_interval: Duration::from_millis(500),
flush_threshold: 1000,
admission_backoff: Duration::from_secs(5),
}
}
}
/// Bounded pending set with count and byte ceilings. Overflow drops the
/// incoming intent (never a resident one) and counts the loss.
pub(crate) struct MrfQueue {
pending: VecDeque<MrfIntent>,
bytes: usize,
capacity: usize,
byte_budget: usize,
}
impl MrfQueue {
pub(crate) fn new(capacity: usize, byte_budget: usize) -> Self {
Self {
pending: VecDeque::new(),
bytes: 0,
capacity,
byte_budget,
}
}
/// Returns `false` (after counting) when either ceiling would be crossed.
pub(crate) fn try_push(&mut self, intent: MrfIntent) -> bool {
let cost = intent.estimated_bytes();
if self.pending.len() >= self.capacity || self.bytes + cost > self.byte_budget {
counter!("rustfs_heal_mrf_dropped_total", "reason" => "queue_overflow").increment(1);
return false;
}
self.bytes += cost;
self.pending.push_back(intent);
true
}
pub(crate) fn pop_front(&mut self) -> Option<MrfIntent> {
let intent = self.pending.pop_front()?;
self.bytes = self.bytes.saturating_sub(intent.estimated_bytes());
Some(intent)
}
pub(crate) fn push_back(&mut self, intent: MrfIntent) {
self.bytes += intent.estimated_bytes();
self.pending.push_back(intent);
}
pub(crate) fn depth(&self) -> usize {
self.pending.len()
}
pub(crate) fn bytes(&self) -> usize {
self.bytes
}
pub(crate) fn intents(&self) -> impl Iterator<Item = &MrfIntent> {
self.pending.iter()
}
}
// ---------------------------------------------------------------------------
// Journal record codec
// ---------------------------------------------------------------------------
/// Append one encoded record to `out`.
pub(crate) fn encode_intent(intent: &MrfIntent, out: &mut Vec<u8>) {
let start = out.len();
out.push(MRF_JOURNAL_FORMAT);
out.push(MRF_JOURNAL_VERSION);
out.push(match intent.kind {
rustfs_common::mrf_channel::MrfKind::DecodeFailure => 1,
rustfs_common::mrf_channel::MrfKind::MetadataCorruption => 2,
rustfs_common::mrf_channel::MrfKind::PartialWrite => 3,
});
out.push(intent.attempts);
out.extend_from_slice(&intent.enqueued_at_ms.to_le_bytes());
match intent.version_id {
Some(bytes) => {
out.push(1);
out.extend_from_slice(&bytes);
}
None => out.push(0),
}
out.extend_from_slice(&(intent.bucket.len() as u32).to_le_bytes());
out.extend_from_slice(&(intent.object.len() as u32).to_le_bytes());
out.extend_from_slice(intent.bucket.as_bytes());
out.extend_from_slice(intent.object.as_bytes());
let mut hasher = crc_fast::Digest::new(crc_fast::CrcAlgorithm::Crc32IsoHdlc);
hasher.update(&out[start..]);
out.extend_from_slice(&(hasher.finalize() as u32).to_le_bytes());
}
fn decode_one(data: &[u8]) -> Option<(MrfIntent, usize)> {
if data.len() < MRF_RECORD_FIXED_HEAD + 8 {
return None;
}
if data[0] != MRF_JOURNAL_FORMAT || data[1] != MRF_JOURNAL_VERSION {
return None;
}
let kind = match data[2] {
1 => rustfs_common::mrf_channel::MrfKind::DecodeFailure,
2 => rustfs_common::mrf_channel::MrfKind::MetadataCorruption,
3 => rustfs_common::mrf_channel::MrfKind::PartialWrite,
_ => return None,
};
let attempts = data[3];
let enqueued_at_ms = u64::from_le_bytes(data[4..12].try_into().expect("slice length checked"));
let has_version = data[12] != 0;
let mut cursor = MRF_RECORD_FIXED_HEAD;
let version_id = if has_version {
if data.len() < cursor + 16 {
return None;
}
let bytes: [u8; 16] = data[cursor..cursor + 16].try_into().expect("slice length checked");
cursor += 16;
Some(bytes)
} else {
None
};
if data.len() < cursor + 8 {
return None;
}
let bucket_len = u32::from_le_bytes(data[cursor..cursor + 4].try_into().expect("slice length checked")) as usize;
let object_len = u32::from_le_bytes(data[cursor + 4..cursor + 8].try_into().expect("slice length checked")) as usize;
cursor += 8;
let body_end = cursor.checked_add(bucket_len)?.checked_add(object_len)?;
let record_end = body_end.checked_add(4)?;
if data.len() < record_end {
return None;
}
let mut hasher = crc_fast::Digest::new(crc_fast::CrcAlgorithm::Crc32IsoHdlc);
hasher.update(&data[..body_end]);
if (hasher.finalize() as u32) != u32::from_le_bytes(data[body_end..record_end].try_into().expect("slice length checked")) {
return None;
}
let bucket = std::sync::Arc::from(std::str::from_utf8(&data[cursor..cursor + bucket_len]).ok()?);
let object = std::sync::Arc::from(std::str::from_utf8(&data[cursor + bucket_len..body_end]).ok()?);
Some((
MrfIntent {
bucket,
object,
version_id,
kind,
enqueued_at_ms,
attempts,
},
record_end,
))
}
/// Decode a whole journal, stopping at the first torn or corrupt record.
/// Returns the decoded intents and the number of trailing bytes discarded.
pub(crate) fn decode_journal(data: &[u8]) -> (Vec<MrfIntent>, usize) {
let mut intents = Vec::new();
let mut cursor = 0usize;
while cursor < data.len() {
match decode_one(&data[cursor..]) {
Some((intent, consumed)) => {
intents.push(intent);
cursor += consumed;
}
None => break,
}
}
let truncated = data.len() - cursor;
(intents, truncated)
}
// ---------------------------------------------------------------------------
// Journal disk IO (all local disks, first successful read wins)
// ---------------------------------------------------------------------------
async fn journal_disks() -> Vec<DiskStore> {
let map = local_disk_map_read().await;
map.values().flatten().cloned().collect()
}
async fn read_journal() -> Option<Vec<u8>> {
for disk in journal_disks().await {
match disk.read_all(super::RUSTFS_META_BUCKET, MRF_JOURNAL_PATH).await {
Ok(bytes) => return Some(bytes.to_vec()),
Err(_) => continue,
}
}
None
}
async fn write_journal(data: &[u8]) {
let payload = bytes::Bytes::copy_from_slice(data);
for disk in journal_disks().await {
if let Err(err) = disk
.write_all(super::RUSTFS_META_BUCKET, MRF_JOURNAL_PATH, payload.clone())
.await
{
warn_mrf_journal_write(&err);
}
}
if !data.is_empty() {
counter!("rustfs_heal_mrf_journal_fsync_total").increment(1);
}
gauge!("rustfs_heal_mrf_journal_bytes").set(data.len() as f64);
}
async fn delete_journal() {
for disk in journal_disks().await {
let _ = disk
.delete(
super::RUSTFS_META_BUCKET,
MRF_JOURNAL_PATH,
crate::heal::storage_api::owner::EcstoreDeleteOptions::default(),
)
.await;
}
}
fn warn_mrf_journal_write(err: &super::DiskError) {
tracing::warn!(
target: "rustfs::heal::mrf",
error = %err,
"MRF journal write failed; unconsumed intents may be lost on restart"
);
}
// ---------------------------------------------------------------------------
// Consumer
// ---------------------------------------------------------------------------
/// Translate an intent into the prioritized heal request the issue specifies:
/// decode failures go Urgent ECDecode, metadata corruption goes High
/// Metadata, partial writes go Normal object heal.
pub(crate) fn build_heal_request(intent: &MrfIntent) -> HealRequest {
let bucket = intent.bucket.to_string();
let object = intent.object.to_string();
let version_id = intent.version_id.map(|bytes| Uuid::from_bytes(bytes).to_string());
let (heal_type, priority) = match intent.kind {
rustfs_common::mrf_channel::MrfKind::DecodeFailure => (
HealType::ECDecode {
bucket,
object,
version_id,
},
HealPriority::Urgent,
),
rustfs_common::mrf_channel::MrfKind::MetadataCorruption => (HealType::Metadata { bucket, object }, HealPriority::High),
rustfs_common::mrf_channel::MrfKind::PartialWrite => (
HealType::Object {
bucket,
object,
version_id,
},
HealPriority::Normal,
),
};
let mut request = HealRequest::new(heal_type, HealOptions::default(), priority);
request.source = rustfs_common::heal_channel::HealRequestSource::Mrf;
request
}
struct MrfRuntime {
queue: MrfQueue,
config: MrfConsumerConfig,
new_since_flush: usize,
/// True while a journal snapshot exists on disk that no longer reflects
/// an all-consumed pending set; the next idle tick removes it (MinIO
/// deletes its `list.bin` after replay for the same reason).
journal_on_disk: bool,
/// Earliest instant a full-admission retry may proceed.
backoff_until: Option<tokio::time::Instant>,
}
impl MrfRuntime {
fn record_accept(&mut self) {
// Accepted intents leave the pending set; the next flush persists the
// smaller snapshot, which is the journal's compaction.
}
fn snapshot(&self) -> Vec<u8> {
let mut buf = Vec::new();
for intent in self.queue.intents() {
encode_intent(intent, &mut buf);
}
buf
}
async fn flush(&mut self) {
write_journal(&self.snapshot()).await;
self.new_since_flush = 0;
self.journal_on_disk = true;
}
/// Drain pending intents into the heal manager until it is full, the
/// queue empties, or attempts are exhausted.
async fn dispatch(&mut self, manager: &HealManager) {
if let Some(until) = self.backoff_until {
if tokio::time::Instant::now() < until {
return;
}
self.backoff_until = None;
}
while let Some(mut intent) = self.queue.pop_front() {
let request = build_heal_request(&intent);
match manager.submit_heal_request(request).await {
Ok(HealAdmissionResult::Accepted) | Ok(HealAdmissionResult::Merged) => self.record_accept(),
Ok(HealAdmissionResult::Full) | Ok(HealAdmissionResult::Dropped(HealAdmissionDropReason::QueueFull)) => {
intent.attempts = intent.attempts.saturating_add(1);
if intent.attempts >= MRF_MAX_ATTEMPTS {
counter!("rustfs_heal_mrf_dropped_total", "reason" => "attempts_exhausted").increment(1);
continue;
}
self.queue.push_back(intent);
self.backoff_until = Some(tokio::time::Instant::now() + self.config.admission_backoff);
break;
}
Ok(HealAdmissionResult::Dropped(_)) => {
counter!("rustfs_heal_mrf_dropped_total", "reason" => "admission_policy").increment(1);
}
Err(_) => {
intent.attempts = intent.attempts.saturating_add(1);
if intent.attempts >= MRF_MAX_ATTEMPTS {
counter!("rustfs_heal_mrf_dropped_total", "reason" => "attempts_exhausted").increment(1);
continue;
}
self.queue.push_back(intent);
self.backoff_until = Some(tokio::time::Instant::now() + self.config.admission_backoff);
break;
}
}
}
gauge!("rustfs_heal_mrf_queue_depth").set(self.queue.depth() as f64);
gauge!("rustfs_heal_mrf_queue_bytes").set(self.queue.bytes() as f64);
}
}
/// Initialize the global MRF channel (honoring `RUSTFS_HEAL_MRF_ENABLE`) and
/// spawn the consumer task. Called once from the heal runtime bootstrap right
/// after the manager started; a disabled feature or a double call is a no-op.
/// Public for integration tests that drive the real consumer loop.
pub fn spawn_mrf_consumer(manager: Arc<HealManager>) {
let enabled = rustfs_utils::get_env_bool(rustfs_config::ENV_HEAL_MRF_ENABLE, rustfs_config::DEFAULT_HEAL_MRF_ENABLE);
rustfs_common::mrf_channel::set_mrf_delivery_enabled(enabled);
if !enabled {
tracing::info!(
target: "rustfs::heal::mrf",
"MRF intent pipeline disabled by configuration; producers will not deliver"
);
return;
}
let receiver = match rustfs_common::mrf_channel::init_mrf_channel() {
Ok(receiver) => receiver,
Err(err) => {
tracing::warn!(
target: "rustfs::heal::mrf",
error = err,
"MRF channel initialization failed; intents will be dropped at producers"
);
return;
}
};
tokio::spawn(async move {
run_mrf_consumer(manager, receiver).await;
});
tracing::info!(target: "rustfs::heal::mrf", "MRF intent consumer started");
}
/// Replay the durable journal into a fresh pending queue and submit whatever
/// it armed. Returns the number of intact intents replayed. Duplicates are
/// merged by the manager's dedup key; the journal file is removed once read
/// (torn tails truncate via the per-record CRC). Public for integration tests;
/// the live consumer invokes this through [`replay_into`] at startup.
pub async fn replay_journal_once(manager: &Arc<HealManager>) -> usize {
let config = MrfConsumerConfig::default();
let mut queue = MrfQueue::new(config.queue_capacity, config.journal_max_bytes);
let mut backoff_until: Option<tokio::time::Instant> = None;
replay_into(manager, &mut queue, &mut backoff_until).await
}
/// Shared replay core: read + decode + re-arm + delete, then drain what fits.
async fn replay_into(
manager: &Arc<HealManager>,
queue: &mut MrfQueue,
backoff_until: &mut Option<tokio::time::Instant>,
) -> usize {
let Some(data) = read_journal().await else {
return 0;
};
let (intents, truncated) = decode_journal(&data);
if truncated > 0 {
tracing::warn!(
target: "rustfs::heal::mrf",
truncated_bytes = truncated,
"MRF journal had a torn tail; truncated records were discarded"
);
}
counter!("rustfs_heal_mrf_replayed_total").increment(intents.len() as u64);
let replayed = intents.len();
for intent in intents {
queue.try_push(intent);
}
delete_journal().await;
// Drain the replayed intents immediately; whatever the manager refuses
// stays armed in `queue` for the consumer's retry loop.
if backoff_until.is_none() {
while let Some(mut intent) = queue.pop_front() {
let request = build_heal_request(&intent);
match manager.submit_heal_request(request).await {
Ok(HealAdmissionResult::Accepted) | Ok(HealAdmissionResult::Merged) => {}
Ok(HealAdmissionResult::Full) | Ok(HealAdmissionResult::Dropped(HealAdmissionDropReason::QueueFull)) => {
intent.attempts = intent.attempts.saturating_add(1);
if intent.attempts < MRF_MAX_ATTEMPTS {
queue.push_back(intent);
*backoff_until = Some(tokio::time::Instant::now());
}
break;
}
Ok(HealAdmissionResult::Dropped(_)) | Err(_) => {}
}
}
}
replayed
}
/// Replay the journal, then keep draining the channel into the heal manager
/// while persisting the pending snapshot.
async fn run_mrf_consumer(manager: Arc<HealManager>, mut receiver: mpsc::Receiver<MrfIntent>) {
let config = MrfConsumerConfig::default();
let mut runtime = MrfRuntime {
queue: MrfQueue::new(config.queue_capacity, config.journal_max_bytes),
config: config.clone(),
new_since_flush: 0,
journal_on_disk: false,
backoff_until: None,
};
// Replay: read the journal, re-arm intents (duplicates are merged by the
// manager's dedup key), then drop the file so the next flush starts clean.
replay_into(&manager, &mut runtime.queue, &mut runtime.backoff_until).await;
let mut flush_tick = tokio::time::interval(runtime.config.flush_interval);
flush_tick.set_missed_tick_behavior(tokio::time::MissedTickBehavior::Delay);
let mut batch: Vec<MrfIntent> = Vec::with_capacity(runtime.config.replay_batch);
loop {
tokio::select! {
received = receiver.recv_many(&mut batch, runtime.config.replay_batch) => {
if received == 0 {
// Channel closed: flush once more and stop.
runtime.flush().await;
tracing::info!(
target: "rustfs::heal::mrf",
"MRF channel closed; consumer stopped after final flush"
);
return;
}
for intent in batch.drain(..) {
runtime.queue.try_push(intent);
runtime.new_since_flush += 1;
}
runtime.dispatch(manager.as_ref()).await;
if runtime.new_since_flush >= runtime.config.flush_threshold {
runtime.flush().await;
}
}
_ = flush_tick.tick() => {
if runtime.new_since_flush > 0 || runtime.queue.depth() > 0 {
runtime.flush().await;
runtime.dispatch(manager.as_ref()).await;
} else if runtime.journal_on_disk {
// All intents consumed: remove the journal so a restart
// replays nothing (mirrors MinIO's post-replay unlink).
delete_journal().await;
runtime.journal_on_disk = false;
gauge!("rustfs_heal_mrf_journal_bytes").set(0.0);
}
gauge!("rustfs_heal_mrf_queue_depth").set(runtime.queue.depth() as f64);
}
}
}
}
#[cfg(test)]
mod tests {
use super::*;
use rustfs_common::mrf_channel::{MrfIntent, MrfKind};
use std::sync::Arc as StdArc;
fn intent(bucket: &str, object: &str, attempts: u8) -> MrfIntent {
MrfIntent {
bucket: StdArc::from(bucket),
object: StdArc::from(object),
version_id: Some([7u8; 16]),
kind: MrfKind::DecodeFailure,
enqueued_at_ms: 1_700_000_000_000,
attempts,
}
}
#[test]
fn queue_enforces_count_and_byte_ceilings() {
let mut queue = MrfQueue::new(2, usize::MAX);
assert!(queue.try_push(intent("b", "o", 0)));
assert!(queue.try_push(intent("b", "o", 0)));
assert!(!queue.try_push(intent("b", "o", 0)), "count ceiling must drop");
let mut tiny = MrfQueue::new(usize::MAX, intent("bucket", "object", 0).estimated_bytes());
assert!(tiny.try_push(intent("bucket", "object", 0)));
assert!(
!tiny.try_push(intent("bucket", "object", 0)),
"byte budget must drop before the second intent fits"
);
}
#[test]
fn journal_roundtrip_preserves_intents() {
let intents = vec![
intent("bucket-a", "object/a", 0),
intent("bucket-b", "object/b", 2),
MrfIntent {
bucket: StdArc::from("bucket-c"),
object: StdArc::from("object/c"),
version_id: None,
kind: MrfKind::MetadataCorruption,
enqueued_at_ms: 5,
attempts: 1,
},
];
let mut buf = Vec::new();
for intent in &intents {
encode_intent(intent, &mut buf);
}
let (decoded, truncated) = decode_journal(&buf);
assert_eq!(truncated, 0);
assert_eq!(decoded.len(), intents.len());
for (left, right) in decoded.iter().zip(intents.iter()) {
assert_eq!(left.bucket, right.bucket);
assert_eq!(left.object, right.object);
assert_eq!(left.version_id, right.version_id);
assert_eq!(left.kind, right.kind);
assert_eq!(left.attempts, right.attempts);
}
}
#[test]
fn journal_torn_tail_is_truncated() {
let mut buf = Vec::new();
encode_intent(&intent("b", "o", 0), &mut buf);
let mut torn = buf.clone();
torn.extend_from_slice(&buf[..buf.len() / 2]);
let (decoded, truncated) = decode_journal(&torn);
assert_eq!(decoded.len(), 1, "the intact record must survive");
assert!(truncated > 0, "the partial tail must be discarded");
// A corrupted body (CRC mismatch) also truncates from that record on.
let mut corrupt = buf.clone();
let mid = MRF_RECORD_FIXED_HEAD + 4;
corrupt[mid] ^= 0xff;
let (decoded, truncated) = decode_journal(&corrupt);
assert!(decoded.is_empty());
assert_eq!(truncated, corrupt.len());
}
#[test]
fn heal_request_mapping_follows_priority_matrix() {
let decode = build_heal_request(&intent("b", "o", 0));
assert!(matches!(decode.heal_type, HealType::ECDecode { .. }));
assert_eq!(decode.priority, HealPriority::Urgent);
let metadata = build_heal_request(&MrfIntent {
bucket: StdArc::from("b"),
object: StdArc::from("o"),
version_id: None,
kind: MrfKind::MetadataCorruption,
enqueued_at_ms: 0,
attempts: 0,
});
assert!(matches!(metadata.heal_type, HealType::Metadata { .. }));
assert_eq!(metadata.priority, HealPriority::High);
let partial = build_heal_request(&MrfIntent {
bucket: StdArc::from("b"),
object: StdArc::from("o"),
version_id: None,
kind: MrfKind::PartialWrite,
enqueued_at_ms: 0,
attempts: 0,
});
assert!(matches!(partial.heal_type, HealType::Object { .. }));
assert_eq!(partial.priority, HealPriority::Normal);
}
}
+160 -6
View File
@@ -13,7 +13,7 @@
// limitations under the License. // limitations under the License.
use serde::{Deserialize, Serialize}; use serde::{Deserialize, Serialize};
use std::time::SystemTime; use std::time::{Duration, SystemTime};
#[derive(Debug, Default, Clone, Serialize, Deserialize)] #[derive(Debug, Default, Clone, Serialize, Deserialize)]
#[serde(rename_all = "camelCase")] #[serde(rename_all = "camelCase")]
@@ -24,6 +24,14 @@ pub struct HealProgress {
pub objects_healed: u64, pub objects_healed: u64,
/// Objects failed /// Objects failed
pub objects_failed: u64, pub objects_failed: u64,
/// Versions skipped because they were written after this heal started
pub skipped_new_versions: u64,
/// Versions skipped because lifecycle already selected them for expiry
pub skipped_ilm_expired: u64,
/// Baseline object count from the latest complete usage snapshot
pub objects_total_count: u64,
/// Baseline object bytes from the latest complete usage snapshot
pub objects_total_size: u64,
/// Bytes processed /// Bytes processed
pub bytes_processed: u64, pub bytes_processed: u64,
/// Current object /// Current object
@@ -54,10 +62,56 @@ impl HealProgress {
self.bytes_processed = bytes; self.bytes_processed = bytes;
self.last_update_time = Some(SystemTime::now()); self.last_update_time = Some(SystemTime::now());
// calculate progress percentage self.refresh_progress_percentage();
let total = scanned + healed + failed; self.refresh_estimated_completion_time();
}
pub fn set_total_baseline(&mut self, objects_total_count: u64, objects_total_size: u64) {
self.objects_total_count = objects_total_count;
self.objects_total_size = objects_total_size;
self.last_update_time = Some(SystemTime::now());
self.refresh_progress_percentage();
self.refresh_estimated_completion_time();
}
pub fn record_skipped_new_version(&mut self) {
self.skipped_new_versions = self.skipped_new_versions.saturating_add(1);
self.last_update_time = Some(SystemTime::now());
self.refresh_progress_percentage();
self.refresh_estimated_completion_time();
}
pub fn record_skipped_ilm_expired(&mut self) {
self.skipped_ilm_expired = self.skipped_ilm_expired.saturating_add(1);
self.last_update_time = Some(SystemTime::now());
self.refresh_progress_percentage();
self.refresh_estimated_completion_time();
}
fn completed_for_baseline(&self) -> u64 {
self.objects_healed
.saturating_add(self.objects_failed)
.saturating_add(self.skipped_new_versions)
.saturating_add(self.skipped_ilm_expired)
}
pub(crate) fn refresh_progress_percentage(&mut self) {
if self.objects_total_size > 0 {
self.progress_percentage = ((self.bytes_processed as f64 / self.objects_total_size as f64) * 100.0).min(100.0);
return;
}
if self.objects_total_count > 0 {
let completed = self.completed_for_baseline();
self.progress_percentage = ((completed as f64 / self.objects_total_count as f64) * 100.0).min(100.0);
return;
}
let total = self
.objects_scanned
.saturating_add(self.objects_healed)
.saturating_add(self.objects_failed);
if total > 0 { if total > 0 {
self.progress_percentage = (healed as f64 / total as f64) * 100.0; self.progress_percentage = (self.objects_healed as f64 / total as f64) * 100.0;
} }
} }
@@ -66,9 +120,36 @@ impl HealProgress {
self.last_update_time = Some(SystemTime::now()); self.last_update_time = Some(SystemTime::now());
} }
pub fn refresh_estimated_completion_time(&mut self) {
let Some(start_time) = self.start_time else {
self.estimated_completion_time = None;
return;
};
if self.is_completed() || !(0.0..100.0).contains(&self.progress_percentage) || self.bytes_processed == 0 {
self.estimated_completion_time = None;
return;
}
let elapsed = match SystemTime::now().duration_since(start_time) {
Ok(elapsed) if !elapsed.is_zero() => elapsed,
_ => {
self.estimated_completion_time = None;
return;
}
};
let estimated_total_secs = elapsed.as_secs_f64() * 100.0 / self.progress_percentage;
self.estimated_completion_time = start_time.checked_add(Duration::from_secs_f64(estimated_total_secs));
}
pub fn is_completed(&self) -> bool { pub fn is_completed(&self) -> bool {
self.progress_percentage >= 100.0 if self.progress_percentage >= 100.0 {
|| self.objects_scanned > 0 && self.objects_healed + self.objects_failed >= self.objects_scanned return true;
}
if self.objects_total_count > 0 || self.objects_total_size > 0 {
return false;
}
self.objects_scanned > 0 && self.objects_healed.saturating_add(self.objects_failed) >= self.objects_scanned
} }
pub fn get_success_rate(&self) -> f64 { pub fn get_success_rate(&self) -> f64 {
@@ -158,6 +239,10 @@ mod tests {
assert_eq!(progress.objects_scanned, 0); assert_eq!(progress.objects_scanned, 0);
assert_eq!(progress.objects_healed, 0); assert_eq!(progress.objects_healed, 0);
assert_eq!(progress.objects_failed, 0); assert_eq!(progress.objects_failed, 0);
assert_eq!(progress.skipped_new_versions, 0);
assert_eq!(progress.skipped_ilm_expired, 0);
assert_eq!(progress.objects_total_count, 0);
assert_eq!(progress.objects_total_size, 0);
assert_eq!(progress.bytes_processed, 0); assert_eq!(progress.bytes_processed, 0);
assert_eq!(progress.progress_percentage, 0.0); assert_eq!(progress.progress_percentage, 0.0);
assert!(progress.start_time.is_some()); assert!(progress.start_time.is_some());
@@ -181,6 +266,73 @@ mod tests {
assert!(progress.last_update_time.is_some()); assert!(progress.last_update_time.is_some());
} }
#[test]
fn test_heal_progress_estimates_completion_time_from_progress() {
let mut progress = HealProgress::new();
progress.start_time = Some(SystemTime::now() - Duration::from_secs(10));
progress.update_progress(100, 25, 0, 4096);
let eta = progress
.estimated_completion_time
.expect("partial byte progress should estimate completion");
assert!(eta > SystemTime::now());
}
#[test]
fn test_heal_progress_uses_byte_baseline_for_percentage() {
let mut progress = HealProgress::new();
progress.set_total_baseline(10, 8192);
progress.update_progress(100, 25, 0, 4096);
assert!((progress.progress_percentage - 50.0).abs() < 0.001);
}
#[test]
fn test_heal_progress_uses_object_baseline_when_bytes_unknown() {
let mut progress = HealProgress::new();
progress.set_total_baseline(10, 0);
progress.update_progress(100, 3, 2, 0);
assert!((progress.progress_percentage - 50.0).abs() < 0.001);
}
#[test]
fn test_heal_progress_counts_skipped_versions_for_object_baseline() {
let mut progress = HealProgress::new();
progress.set_total_baseline(10, 0);
progress.update_progress(100, 3, 2, 0);
progress.record_skipped_new_version();
assert_eq!(progress.skipped_new_versions, 1);
assert!((progress.progress_percentage - 60.0).abs() < 0.001);
}
#[test]
fn test_heal_progress_does_not_estimate_completion_without_bytes() {
let mut progress = HealProgress::new();
progress.start_time = Some(SystemTime::now() - Duration::from_secs(10));
progress.update_progress(100, 25, 0, 0);
assert!(progress.estimated_completion_time.is_none());
}
#[test]
fn test_heal_progress_with_baseline_is_not_completed_by_processed_count() {
let mut progress = HealProgress::new();
progress.start_time = Some(SystemTime::now() - Duration::from_secs(10));
progress.set_total_baseline(10, 8192);
progress.update_progress(1, 1, 0, 1024);
assert!(!progress.is_completed());
assert!(progress.estimated_completion_time.is_some());
}
#[test] #[test]
fn test_heal_progress_update_progress_zero_total() { fn test_heal_progress_update_progress_zero_total() {
let mut progress = HealProgress::new(); let mut progress = HealProgress::new();
@@ -251,6 +403,8 @@ mod tests {
assert_eq!(json["objectsScanned"], 10); assert_eq!(json["objectsScanned"], 10);
assert_eq!(json["objectsHealed"], 8); assert_eq!(json["objectsHealed"], 8);
assert_eq!(json["objectsFailed"], 2); assert_eq!(json["objectsFailed"], 2);
assert_eq!(json["skippedNewVersions"], 0);
assert_eq!(json["skippedIlmExpired"], 0);
assert_eq!(json["bytesProcessed"], 1024); assert_eq!(json["bytesProcessed"], 1024);
assert_eq!(json["currentObject"], "test-bucket/test-object"); assert_eq!(json["currentObject"], "test-bucket/test-object");
assert!(json["progressPercentage"].is_number()); assert!(json["progressPercentage"].is_number());
+169 -7
View File
@@ -22,6 +22,7 @@ use serde::{Deserialize, Serialize};
use std::sync::Arc; use std::sync::Arc;
use tracing::{debug, error, warn}; use tracing::{debug, error, warn};
use super::storage_api::owner::{EcstoreHealLifecycleExpiryContext, ecstore_load_admin_data_usage_from_backend_cached};
use super::storage_api::storage::{ use super::storage_api::storage::{
BucketInfo, BucketOperations, DiskSetSelector, HealOperations as _, ListOperations as _, ObjectIO as _, BucketInfo, BucketOperations, DiskSetSelector, HealOperations as _, ListOperations as _, ObjectIO as _,
ObjectOperations as _, StorageAdminApi, ObjectOperations as _, StorageAdminApi,
@@ -29,6 +30,37 @@ use super::storage_api::storage::{
use super::{DiskStore, ECStore, Endpoint, HealDiskExt as _, StorageError, resume::ReplacementTargetIdentity}; use super::{DiskStore, ECStore, Endpoint, HealDiskExt as _, StorageError, resume::ReplacementTargetIdentity};
pub use super::{HealObjectInfo, HealObjectOptions, HealPutObjReader}; pub use super::{HealObjectInfo, HealObjectOptions, HealPutObjReader};
#[derive(Debug, Clone, Copy, Default, PartialEq, Eq)]
pub struct HealBucketUsageBaseline {
pub objects_count: u64,
pub bytes: u64,
}
pub struct HealLifecycleExpiryContext {
inner: HealLifecycleExpiryContextInner,
}
enum HealLifecycleExpiryContextInner {
Ecstore(EcstoreHealLifecycleExpiryContext),
#[allow(dead_code)]
Test,
}
impl HealLifecycleExpiryContext {
fn ecstore(inner: EcstoreHealLifecycleExpiryContext) -> Self {
Self {
inner: HealLifecycleExpiryContextInner::Ecstore(inner),
}
}
#[cfg(test)]
pub(crate) fn test() -> Self {
Self {
inner: HealLifecycleExpiryContextInner::Test,
}
}
}
const LOG_COMPONENT_HEAL: &str = "heal"; const LOG_COMPONENT_HEAL: &str = "heal";
const LOG_SUBSYSTEM_STORAGE: &str = "storage"; const LOG_SUBSYSTEM_STORAGE: &str = "storage";
const EVENT_HEAL_STORAGE_OBJECT_IO: &str = "heal_storage_object_io"; const EVENT_HEAL_STORAGE_OBJECT_IO: &str = "heal_storage_object_io";
@@ -272,6 +304,10 @@ pub struct HealListItem {
pub name: String, pub name: String,
/// normalized version id (`None` when the version is nil/absent) /// normalized version id (`None` when the version is nil/absent)
pub version_id: Option<String>, pub version_id: Option<String>,
/// version modification time as Unix nanoseconds
pub mod_time_unix_nanos: Option<i128>,
/// object snapshot for lifecycle evaluation
pub lifecycle_object_info: Option<HealObjectInfo>,
/// whether this version is a delete marker (observability only) /// whether this version is a delete marker (observability only)
pub is_delete_marker: bool, pub is_delete_marker: bool,
} }
@@ -329,6 +365,28 @@ pub trait HealStorageAPI: Send + Sync {
/// Get bucket info /// Get bucket info
async fn get_bucket_info(&self, bucket: &str) -> Result<Option<BucketInfo>>; async fn get_bucket_info(&self, bucket: &str) -> Result<Option<BucketInfo>>;
/// Aggregate usage-cache baselines for the requested buckets.
async fn erasure_set_usage_baseline(&self, _buckets: &[String]) -> Result<Option<HealBucketUsageBaseline>> {
Ok(None)
}
/// Load per-bucket lifecycle expiry context for heal skips.
async fn load_heal_lifecycle_expiry_context(&self, _bucket: &str) -> Result<Option<HealLifecycleExpiryContext>> {
Ok(None)
}
/// Queue lifecycle expiry for a version that heal can skip.
async fn enqueue_heal_lifecycle_expiry(
&self,
_context: &HealLifecycleExpiryContext,
_bucket: &str,
_object: &str,
_version_id: Option<&str>,
_object_info: Option<&HealObjectInfo>,
) -> Result<bool> {
Ok(false)
}
/// Fix bucket metadata /// Fix bucket metadata
async fn heal_bucket_metadata(&self, bucket: &str) -> Result<()>; async fn heal_bucket_metadata(&self, bucket: &str) -> Result<()>;
@@ -409,6 +467,7 @@ pub trait HealStorageAPI: Send + Sync {
bucket: &str, bucket: &str,
prefix: &str, prefix: &str,
continuation_token: Option<&str>, continuation_token: Option<&str>,
include_lifecycle_object_info: bool,
) -> Result<(Vec<HealListItem>, Option<String>, bool)>; ) -> Result<(Vec<HealListItem>, Option<String>, bool)>;
/// List versions for healing via a per-erasure-set DISK-WALK union enumerator /// List versions for healing via a per-erasure-set DISK-WALK union enumerator
@@ -427,8 +486,10 @@ pub trait HealStorageAPI: Send + Sync {
bucket: &str, bucket: &str,
prefix: &str, prefix: &str,
continuation_token: Option<&str>, continuation_token: Option<&str>,
include_lifecycle_object_info: bool,
) -> Result<(Vec<HealListItem>, Option<String>, bool)> { ) -> Result<(Vec<HealListItem>, Option<String>, bool)> {
self.list_objects_for_heal_page(bucket, prefix, continuation_token).await self.list_objects_for_heal_page(bucket, prefix, continuation_token, include_lifecycle_object_info)
.await
} }
/// Get disk for resume functionality. /// Get disk for resume functionality.
@@ -1021,6 +1082,85 @@ impl HealStorageAPI for ECStoreHealStorage {
} }
} }
async fn erasure_set_usage_baseline(&self, buckets: &[String]) -> Result<Option<HealBucketUsageBaseline>> {
if buckets.is_empty() {
return Ok(None);
}
let info = match ecstore_load_admin_data_usage_from_backend_cached(self.ecstore.clone()).await {
Ok(info) if info.is_complete_bucket_usage_snapshot() => info,
Ok(_) | Err(_) => return Ok(None),
};
let mut baseline = HealBucketUsageBaseline::default();
for bucket in buckets {
if let Some(usage) = info.buckets_usage.get(bucket) {
baseline.objects_count = baseline.objects_count.saturating_add(usage.objects_count);
baseline.bytes = baseline.bytes.saturating_add(usage.size);
}
}
Ok(Some(baseline))
}
async fn load_heal_lifecycle_expiry_context(&self, bucket: &str) -> Result<Option<HealLifecycleExpiryContext>> {
match self.ecstore.load_heal_lifecycle_expiry_context(bucket).await {
Ok(Some(context)) => Ok(Some(HealLifecycleExpiryContext::ecstore(context))),
Ok(None) => Ok(None),
Err(err) => {
debug!(
target: "rustfs::heal::storage",
event = EVENT_HEAL_STORAGE_ADMIN_OP,
component = LOG_COMPONENT_HEAL,
subsystem = LOG_SUBSYSTEM_STORAGE,
operation = "load_heal_lifecycle_expiry_context",
bucket,
result = "failed",
error = %err,
"Heal storage lifecycle expiry context load failed"
);
Ok(None)
}
}
}
async fn enqueue_heal_lifecycle_expiry(
&self,
context: &HealLifecycleExpiryContext,
bucket: &str,
object: &str,
version_id: Option<&str>,
object_info: Option<&HealObjectInfo>,
) -> Result<bool> {
let context = match &context.inner {
HealLifecycleExpiryContextInner::Ecstore(context) => context,
HealLifecycleExpiryContextInner::Test => return Ok(false),
};
match self
.ecstore
.enqueue_heal_lifecycle_expiry(context, bucket, object, version_id, object_info)
.await
{
Ok(queued) => Ok(queued),
Err(err) => {
debug!(
target: "rustfs::heal::storage",
event = EVENT_HEAL_STORAGE_ADMIN_OP,
component = LOG_COMPONENT_HEAL,
subsystem = LOG_SUBSYSTEM_STORAGE,
operation = "enqueue_heal_lifecycle_expiry",
bucket,
object,
version_id = ?version_id,
result = "failed",
error = %err,
"Heal storage lifecycle expiry check failed"
);
Ok(false)
}
}
}
async fn heal_bucket_metadata(&self, bucket: &str) -> Result<()> { async fn heal_bucket_metadata(&self, bucket: &str) -> Result<()> {
debug!( debug!(
target: "rustfs::heal::storage", target: "rustfs::heal::storage",
@@ -1436,7 +1576,7 @@ impl HealStorageAPI for ECStoreHealStorage {
loop { loop {
let (page_objects, next_token, is_truncated) = self let (page_objects, next_token, is_truncated) = self
.list_objects_for_heal_page(bucket, prefix, continuation_token.as_deref()) .list_objects_for_heal_page(bucket, prefix, continuation_token.as_deref(), false)
.await?; .await?;
all_objects.extend(page_objects); all_objects.extend(page_objects);
@@ -1471,6 +1611,7 @@ impl HealStorageAPI for ECStoreHealStorage {
bucket: &str, bucket: &str,
prefix: &str, prefix: &str,
continuation_token: Option<&str>, continuation_token: Option<&str>,
include_lifecycle_object_info: bool,
) -> Result<(Vec<HealListItem>, Option<String>, bool)> { ) -> Result<(Vec<HealListItem>, Option<String>, bool)> {
debug!( debug!(
target: "rustfs::heal::storage", target: "rustfs::heal::storage",
@@ -1522,10 +1663,19 @@ impl HealStorageAPI for ECStoreHealStorage {
let page_objects: Vec<HealListItem> = list_info let page_objects: Vec<HealListItem> = list_info
.objects .objects
.into_iter() .into_iter()
.map(|obj| HealListItem { .map(|mut obj| {
name: obj.name, obj.version_id = obj.version_id.filter(|u| !u.is_nil());
version_id: obj.version_id.filter(|u| !u.is_nil()).map(|u| u.to_string()), let version_id = obj.version_id.map(|u| u.to_string());
is_delete_marker: obj.delete_marker, let mod_time_unix_nanos = obj.mod_time.map(|mod_time| mod_time.unix_timestamp_nanos());
let is_delete_marker = obj.delete_marker;
let lifecycle_object_info = include_lifecycle_object_info.then(|| obj.clone());
HealListItem {
name: obj.name,
version_id,
mod_time_unix_nanos,
lifecycle_object_info,
is_delete_marker,
}
}) })
.collect(); .collect();
let page_count = page_objects.len(); let page_count = page_objects.len();
@@ -1562,6 +1712,7 @@ impl HealStorageAPI for ECStoreHealStorage {
bucket: &str, bucket: &str,
prefix: &str, prefix: &str,
continuation_token: Option<&str>, continuation_token: Option<&str>,
include_lifecycle_object_info: bool,
) -> Result<(Vec<HealListItem>, Option<String>, bool)> { ) -> Result<(Vec<HealListItem>, Option<String>, bool)> {
// Per-page bounds for the disk-walk union enumerator. Objects are atomic // Per-page bounds for the disk-walk union enumerator. Objects are atomic
// (never split across pages), so version_budget only bounds how many // (never split across pages), so version_budget only bounds how many
@@ -1590,7 +1741,16 @@ impl HealStorageAPI for ECStoreHealStorage {
let (versions, next_forward, is_truncated) = self let (versions, next_forward, is_truncated) = self
.ecstore .ecstore
.heal_walk_versions_page(pool_idx, set_idx, bucket, prefix, forward_to.as_deref(), BATCH_OBJECTS, VERSION_BUDGET) .heal_walk_versions_page(
pool_idx,
set_idx,
bucket,
prefix,
forward_to.as_deref(),
BATCH_OBJECTS,
VERSION_BUDGET,
include_lifecycle_object_info,
)
.await .await
.map_err(|e| { .map_err(|e| {
error!( error!(
@@ -1614,6 +1774,8 @@ impl HealStorageAPI for ECStoreHealStorage {
.map(|v| HealListItem { .map(|v| HealListItem {
name: v.name, name: v.name,
version_id: v.version_id, version_id: v.version_id,
mod_time_unix_nanos: v.mod_time_unix_nanos,
lifecycle_object_info: v.lifecycle_object_info,
is_delete_marker: v.is_delete_marker, is_delete_marker: v.is_delete_marker,
}) })
.collect(); .collect();
+9 -4
View File
@@ -12,7 +12,10 @@
// See the License for the specific language governing permissions and // See the License for the specific language governing permissions and
// limitations under the License. // limitations under the License.
pub(crate) use rustfs_ecstore::api::data_usage::DATA_USAGE_CACHE_NAME as ECSTORE_DATA_USAGE_CACHE_NAME; pub(crate) use rustfs_ecstore::api::data_usage::{
DATA_USAGE_CACHE_NAME as ECSTORE_DATA_USAGE_CACHE_NAME,
load_admin_data_usage_from_backend_cached as ecstore_load_admin_data_usage_from_backend_cached,
};
pub(crate) use rustfs_ecstore::api::disk::endpoint::Endpoint as EcstoreEndpoint; pub(crate) use rustfs_ecstore::api::disk::endpoint::Endpoint as EcstoreEndpoint;
pub(crate) use rustfs_ecstore::api::disk::error::{DiskError as EcstoreDiskError, Result as EcstoreDiskResult}; pub(crate) use rustfs_ecstore::api::disk::error::{DiskError as EcstoreDiskError, Result as EcstoreDiskResult};
pub(crate) use rustfs_ecstore::api::disk::{ pub(crate) use rustfs_ecstore::api::disk::{
@@ -25,7 +28,9 @@ pub(crate) use rustfs_ecstore::api::disk::{
pub(crate) use rustfs_ecstore::api::disk::{DiskOption as EcstoreDiskOption, new_disk as ecstore_new_disk}; pub(crate) use rustfs_ecstore::api::disk::{DiskOption as EcstoreDiskOption, new_disk as ecstore_new_disk};
pub(crate) use rustfs_ecstore::api::error::{Error as EcstoreErrorType, StorageError as EcstoreStorageError}; pub(crate) use rustfs_ecstore::api::error::{Error as EcstoreErrorType, StorageError as EcstoreStorageError};
pub(crate) use rustfs_ecstore::api::runtime::local_disk_map_read as ecstore_local_disk_map_read; pub(crate) use rustfs_ecstore::api::runtime::local_disk_map_read as ecstore_local_disk_map_read;
pub(crate) use rustfs_ecstore::api::storage::ECStore as EcstoreStore; pub(crate) use rustfs_ecstore::api::storage::{
ECStore as EcstoreStore, HealLifecycleExpiryContext as EcstoreHealLifecycleExpiryContext,
};
use rustfs_storage_api as storage_contracts; use rustfs_storage_api as storage_contracts;
pub(crate) mod owner { pub(crate) mod owner {
@@ -34,8 +39,8 @@ pub(crate) mod owner {
pub(crate) use super::{ pub(crate) use super::{
ECSTORE_BUCKET_META_PREFIX, ECSTORE_DATA_USAGE_CACHE_NAME, ECSTORE_HEALING_MARKER_PATH, ECSTORE_RUSTFS_META_BUCKET, ECSTORE_BUCKET_META_PREFIX, ECSTORE_DATA_USAGE_CACHE_NAME, ECSTORE_HEALING_MARKER_PATH, ECSTORE_RUSTFS_META_BUCKET,
EcstoreConditionalFileUpdate, EcstoreDeleteOptions, EcstoreDiskAPI, EcstoreDiskBytes, EcstoreDiskError, EcstoreConditionalFileUpdate, EcstoreDeleteOptions, EcstoreDiskAPI, EcstoreDiskBytes, EcstoreDiskError,
EcstoreDiskResult, EcstoreDiskStore, EcstoreEndpoint, EcstoreErrorType, EcstoreStorageError, EcstoreStore, EcstoreDiskResult, EcstoreDiskStore, EcstoreEndpoint, EcstoreErrorType, EcstoreHealLifecycleExpiryContext,
ecstore_local_disk_map_read, EcstoreStorageError, EcstoreStore, ecstore_load_admin_data_usage_from_backend_cached, ecstore_local_disk_map_read,
}; };
#[cfg(test)] #[cfg(test)]
+371 -7
View File
@@ -19,11 +19,12 @@ use crate::heal::{
resume::{ resume::{
CheckpointManager, ReplacementPhase, ReplacementTargetIdentity, ResumeManager, replacement_target_identities_match, CheckpointManager, ReplacementPhase, ReplacementTargetIdentity, ResumeManager, replacement_target_identities_match,
}, },
storage::{HealStorageAPI, next_heal_listing_token}, storage::{HealBucketUsageBaseline, HealStorageAPI, next_heal_listing_token},
}; };
use crate::{Error, Result}; use crate::{Error, Result};
use metrics::{counter, histogram}; use metrics::{counter, histogram};
use rustfs_common::heal_channel::{HealOpts, HealRequestSource, HealScanMode}; use rustfs_common::heal_channel::{HealOpts, HealRequestSource, HealScanMode};
use rustfs_common::trace_bus::{TraceEvent, TraceFunc, TraceKind, trace_emit};
use rustfs_madmin::heal_commands::HealResultItem; use rustfs_madmin::heal_commands::HealResultItem;
use rustfs_utils::path::SLASH_SEPARATOR; use rustfs_utils::path::SLASH_SEPARATOR;
use serde::{Deserialize, Serialize}; use serde::{Deserialize, Serialize};
@@ -31,7 +32,7 @@ use std::{
future::Future, future::Future,
sync::{ sync::{
Arc, Arc,
atomic::{AtomicBool, Ordering}, atomic::{AtomicBool, AtomicU64, Ordering},
}, },
time::{Duration, Instant, SystemTime}, time::{Duration, Instant, SystemTime},
}; };
@@ -178,6 +179,17 @@ pub enum HealPriority {
Urgent = 3, Urgent = 3,
} }
impl HealPriority {
fn as_str(self) -> &'static str {
match self {
Self::Low => "low",
Self::Normal => "normal",
Self::High => "high",
Self::Urgent => "urgent",
}
}
}
/// Heal options /// Heal options
#[derive(Debug, Clone, Serialize, Deserialize)] #[derive(Debug, Clone, Serialize, Deserialize)]
pub struct HealOptions { pub struct HealOptions {
@@ -339,6 +351,20 @@ impl HealRequest {
} }
/// Heal task /// Heal task
/// Incremental view over a task's retained result items (HS-06).
///
/// `next_seq` is the cursor a client should pass on its next poll; `min_seq`
/// is the oldest sequence still retained; `lagged` means the client's cursor
/// fell behind `min_seq` and items were skipped — the client should restart
/// from `min_seq`.
#[derive(Debug, Clone)]
pub struct HealResultWindow {
pub items: Vec<HealResultItem>,
pub next_seq: u64,
pub min_seq: u64,
pub lagged: bool,
}
pub struct HealTask { pub struct HealTask {
/// Task ID /// Task ID
pub id: String, pub id: String,
@@ -361,8 +387,16 @@ pub struct HealTask {
pub status: Arc<RwLock<HealTaskStatus>>, pub status: Arc<RwLock<HealTaskStatus>>,
/// Progress tracking /// Progress tracking
pub progress: Arc<RwLock<HealProgress>>, pub progress: Arc<RwLock<HealProgress>>,
/// Result items collected from storage heal calls. /// Result items collected from storage heal calls, each stamped with a
pub result_items: Arc<RwLock<Vec<HealResultItem>>>, /// monotonically increasing sequence number for incremental consumption
/// (the client passes the last seen seq back and receives only newer
/// items; see `get_result_items_since`).
pub result_items: Arc<RwLock<Vec<(u64, HealResultItem)>>>,
/// Next sequence number to assign; starts at 1.
next_item_seq: Arc<AtomicU64>,
/// Sequence number of the oldest item still inside the retention window;
/// equals `next_item_seq` while the window is empty.
min_available_seq: Arc<AtomicU64>,
result_items_truncated: Arc<AtomicBool>, result_items_truncated: Arc<AtomicBool>,
batch_failure: Arc<RwLock<Option<BatchHealFailure>>>, batch_failure: Arc<RwLock<Option<BatchHealFailure>>>,
batch_failure_recorded: Arc<AtomicBool>, batch_failure_recorded: Arc<AtomicBool>,
@@ -414,6 +448,8 @@ impl HealTask {
status: Arc::new(RwLock::new(HealTaskStatus::Pending)), status: Arc::new(RwLock::new(HealTaskStatus::Pending)),
progress: Arc::new(RwLock::new(HealProgress::new())), progress: Arc::new(RwLock::new(HealProgress::new())),
result_items: Arc::new(RwLock::new(Vec::new())), result_items: Arc::new(RwLock::new(Vec::new())),
next_item_seq: Arc::new(AtomicU64::new(1)),
min_available_seq: Arc::new(AtomicU64::new(1)),
result_items_truncated: Arc::new(AtomicBool::new(false)), result_items_truncated: Arc::new(AtomicBool::new(false)),
batch_failure: Arc::new(RwLock::new(None)), batch_failure: Arc::new(RwLock::new(None)),
batch_failure_recorded: Arc::new(AtomicBool::new(false)), batch_failure_recorded: Arc::new(AtomicBool::new(false)),
@@ -498,6 +534,61 @@ impl HealTask {
} }
} }
fn emit_trace_task_state(&self, state: &'static str, duration: Duration, error: Option<&Error>) {
trace_emit(|| {
let mut event = TraceEvent::new(TraceKind::Heal, TraceFunc::HealTask)
.with_duration(duration)
.with_attr("task_id", self.id.as_str())
.with_attr("heal_type", self.heal_type.log_kind())
.with_attr("state", state)
.with_attr("source", self.source.as_str())
.with_attr("priority", self.priority.as_str())
.with_attr("retry_attempts", u64::from(self.retry_attempts))
.with_attr("dry_run", self.options.dry_run);
event = match &self.heal_type {
HealType::Cluster => event,
HealType::Object {
bucket,
object,
version_id,
} => {
let event = event.with_bucket(bucket.as_str()).with_object(object.as_str());
match version_id {
Some(version_id) => event.with_attr("version_id", version_id.as_str()),
None => event,
}
}
HealType::Bucket { bucket } => event.with_bucket(bucket.as_str()),
HealType::Prefix { bucket, prefix } => event.with_bucket(bucket.as_str()).with_object(prefix.as_str()),
HealType::ErasureSet { buckets, set_disk_id } => {
let bucket_count = u64::try_from(buckets.len()).unwrap_or(u64::MAX);
event
.with_attr("set_disk_id", set_disk_id.as_str())
.with_attr("bucket_count", bucket_count)
}
HealType::Metadata { bucket, object } => event.with_bucket(bucket.as_str()).with_object(object.as_str()),
HealType::ECDecode {
bucket,
object,
version_id,
} => {
let event = event.with_bucket(bucket.as_str()).with_object(object.as_str());
match version_id {
Some(version_id) => event.with_attr("version_id", version_id.as_str()),
None => event,
}
}
HealType::MRF { meta_path } => event.with_object(meta_path.as_str()),
};
match error {
Some(error) => event.with_attr("error", error.to_string()),
None => event,
}
});
}
async fn remaining_timeout(&self) -> Result<Option<Duration>> { async fn remaining_timeout(&self) -> Result<Option<Duration>> {
if let Some(total) = self.options.timeout { if let Some(total) = self.options.timeout {
let start_instant = { *self.task_start_instant.read().await }; let start_instant = { *self.task_start_instant.read().await };
@@ -717,6 +808,7 @@ impl HealTask {
queue_delay = ?queue_delay, queue_delay = ?queue_delay,
"Heal task started" "Heal task started"
}); });
self.emit_trace_task_state("started", Duration::ZERO, None);
let result = match &self.heal_type { let result = match &self.heal_type {
HealType::Cluster => self.heal_cluster().await, HealType::Cluster => self.heal_cluster().await,
@@ -805,6 +897,14 @@ impl HealTask {
} }
} }
let terminal_state = match &result {
Ok(_) => "completed",
Err(Error::TaskCancelled) => "cancelled",
Err(Error::TaskTimeout) => "timed_out",
Err(_) => "failed",
};
self.emit_trace_task_state(terminal_state, start_instant.elapsed(), result.as_ref().err());
result result
} }
@@ -835,18 +935,63 @@ impl HealTask {
} }
pub async fn get_result_items(&self) -> Vec<HealResultItem> { pub async fn get_result_items(&self) -> Vec<HealResultItem> {
self.result_items.read().await.iter().map(|(_, item)| item.clone()).collect()
}
/// Sequence-stamped retained window, used when archiving a completed
/// task so incremental cursors survive the transition (HS-06).
pub async fn get_seqed_result_items(&self) -> Vec<(u64, HealResultItem)> {
self.result_items.read().await.clone() self.result_items.read().await.clone()
} }
/// Incremental result window (HS-06): `since = None` returns the full
/// retained window (legacy snapshot semantics); `since = Some(seq)`
/// returns only items stamped with a sequence greater than `seq`.
/// `lagged` warns that the caller's cursor fell behind the window start
/// and items were skipped (the response carries `min_seq` as the catch-up
/// cursor).
pub async fn get_result_items_since(&self, since: Option<u64>) -> HealResultWindow {
let result_items = self.result_items.read().await;
let next_seq = self.next_item_seq.load(Ordering::Relaxed);
let min_seq = self.min_available_seq.load(Ordering::Relaxed);
let mut lagged = false;
let items = match since {
None => result_items.iter().map(|(_, item)| item.clone()).collect::<Vec<_>>(),
Some(cursor) => {
if cursor + 1 < min_seq {
lagged = true;
}
result_items
.iter()
.filter(|(seq, _)| *seq > cursor)
.map(|(_, item)| item.clone())
.collect::<Vec<_>>()
}
};
HealResultWindow {
items,
next_seq,
min_seq,
lagged,
}
}
pub fn result_items_truncated(&self) -> bool { pub fn result_items_truncated(&self) -> bool {
self.result_items_truncated.load(Ordering::Relaxed) self.result_items_truncated.load(Ordering::Relaxed)
} }
async fn record_result_item(&self, result: HealResultItem) { async fn record_result_item(&self, result: HealResultItem) {
let seq = self.next_item_seq.fetch_add(1, Ordering::Relaxed);
let mut result_items = self.result_items.write().await; let mut result_items = self.result_items.write().await;
if result_items.len() < MAX_RETAINED_HEAL_RESULT_ITEMS { if result_items.len() < MAX_RETAINED_HEAL_RESULT_ITEMS {
result_items.push(result); result_items.push((seq, result));
} else { } else {
// Slide the window: the oldest item leaves and the cursor for the
// oldest still-available item moves forward with it.
result_items.remove(0);
self.min_available_seq
.store(result_items.first().map_or(seq, |(oldest, _)| *oldest), Ordering::Relaxed);
result_items.push((seq, result));
self.result_items_truncated.store(true, Ordering::Relaxed); self.result_items_truncated.store(true, Ordering::Relaxed);
} }
} }
@@ -1535,7 +1680,7 @@ impl HealTask {
let (objects, next_token, is_truncated) = self let (objects, next_token, is_truncated) = self
.await_with_control( .await_with_control(
self.storage self.storage
.list_objects_for_heal_page(bucket, prefix, continuation_token.as_deref()), .list_objects_for_heal_page(bucket, prefix, continuation_token.as_deref(), false),
) )
.await?; .await?;
@@ -1697,6 +1842,23 @@ impl HealTask {
Ok(()) Ok(())
} }
async fn apply_erasure_set_usage_baseline(&self, buckets: &[String]) -> Result<()> {
let baseline = match self
.await_with_control(self.storage.erasure_set_usage_baseline(buckets))
.await
{
Ok(Some(baseline)) => baseline,
Ok(None) => return Ok(()),
Err(err @ Error::TaskCancelled) | Err(err @ Error::TaskTimeout) => return Err(err),
Err(_) => return Ok(()),
};
let HealBucketUsageBaseline { objects_count, bytes } = baseline;
let mut progress = self.progress.write().await;
progress.set_total_baseline(objects_count, bytes);
Ok(())
}
async fn heal_metadata(&self, bucket: &str, object: &str) -> Result<()> { async fn heal_metadata(&self, bucket: &str, object: &str) -> Result<()> {
debug!( debug!(
target: "rustfs::heal::task", target: "rustfs::heal::task",
@@ -2298,6 +2460,8 @@ impl HealTask {
None None
}; };
self.apply_erasure_set_usage_baseline(&buckets).await?;
let healing_marker = format!("{set_disk_id}:{}", self.id); let healing_marker = format!("{set_disk_id}:{}", self.id);
if let Some((disk, resume_manager, _)) = replacement_resume.as_ref() { if let Some((disk, resume_manager, _)) = replacement_resume.as_ref() {
let state = resume_manager.get_state().await; let state = resume_manager.get_state().await;
@@ -2602,7 +2766,8 @@ impl HealTask {
{ {
let mut progress = self.progress.write().await; let mut progress = self.progress.write().await;
progress.update_progress(4, 4, 0, 0); let bytes_processed = progress.bytes_processed;
progress.update_progress(4, 4, 0, bytes_processed);
} }
match result { match result {
@@ -2658,6 +2823,7 @@ mod tests {
use super::super::{DiskOption, DiskStore, Endpoint, HealDiskExt as _, new_disk}; use super::super::{DiskOption, DiskStore, Endpoint, HealDiskExt as _, new_disk};
use super::*; use super::*;
use crate::heal::storage::{DiskStatus, HealListItem, HealObjectInfo}; use crate::heal::storage::{DiskStatus, HealListItem, HealObjectInfo};
use rustfs_common::trace_bus::{TraceEvent, TraceFunc, TraceKind, TraceSubscription, TraceVal, subscribe_trace_events};
use rustfs_madmin::heal_commands::{HealDriveInfo, HealResultItem, Infos}; use rustfs_madmin::heal_commands::{HealDriveInfo, HealResultItem, Infos};
use std::collections::{HashMap, VecDeque}; use std::collections::{HashMap, VecDeque};
use std::sync::Mutex; use std::sync::Mutex;
@@ -3203,6 +3369,8 @@ mod tests {
block_heal_object: Mutex<bool>, block_heal_object: Mutex<bool>,
resume_disk: Mutex<Option<DiskStore>>, resume_disk: Mutex<Option<DiskStore>>,
replacement_resume_disk: Mutex<Option<DiskStore>>, replacement_resume_disk: Mutex<Option<DiskStore>>,
usage_baseline: Mutex<Option<HealBucketUsageBaseline>>,
usage_baseline_error: Mutex<bool>,
} }
#[test] #[test]
@@ -3265,11 +3433,69 @@ mod tests {
assert_eq!(samples_logged, MAX_BUCKET_FAILURE_LOG_SAMPLES); assert_eq!(samples_logged, MAX_BUCKET_FAILURE_LOG_SAMPLES);
} }
#[tokio::test]
async fn execute_emits_heal_trace_task_state() {
let mut trace = subscribe_trace_events();
let storage = Arc::new(MockStorage::default());
let task = HealTask::from_request(
HealRequest::object("bucket-a".to_string(), "object-a".to_string(), Some("version-a".to_string())),
storage,
);
task.execute().await.expect("mock object heal should complete");
let started = recv_trace_task_state(&mut trace, &task.id, "started").await;
assert_eq!(started.kind, TraceKind::Heal);
assert_eq!(started.func, TraceFunc::HealTask);
assert_eq!(started.bucket.as_deref(), Some("bucket-a"));
assert_eq!(started.object.as_deref(), Some("object-a"));
assert_eq!(trace_attr_string(&started, "heal_type").as_deref(), Some("object"));
assert_eq!(trace_attr_string(&started, "source").as_deref(), Some("internal"));
assert_eq!(trace_attr_string(&started, "version_id").as_deref(), Some("version-a"));
let completed = recv_trace_task_state(&mut trace, &task.id, "completed").await;
assert_eq!(completed.kind, TraceKind::Heal);
assert_eq!(completed.func, TraceFunc::HealTask);
assert_eq!(trace_attr_string(&completed, "state").as_deref(), Some("completed"));
}
async fn recv_trace_task_state(trace: &mut TraceSubscription, task_id: &str, state: &str) -> TraceEvent {
for _ in 0..32 {
let event = tokio::time::timeout(Duration::from_secs(1), trace.recv())
.await
.expect("trace event should arrive")
.expect("trace bus should stay open");
if trace_attr_string(&event, "task_id").as_deref() == Some(task_id)
&& trace_attr_string(&event, "state").as_deref() == Some(state)
{
return (*event).clone();
}
}
panic!("expected trace state {state} for task {task_id}");
}
fn trace_attr_string(event: &TraceEvent, key: &str) -> Option<String> {
event.attrs.iter().find_map(|attr| {
if attr.key != key {
return None;
}
Some(match &attr.value {
TraceVal::Bool(value) => value.to_string(),
TraceVal::U64(value) => value.to_string(),
TraceVal::I64(value) => value.to_string(),
TraceVal::Str(value) => value.to_string(),
})
})
}
/// Build a latest, non-delete-marker heal list item with no version id. /// Build a latest, non-delete-marker heal list item with no version id.
fn heal_item(name: &str) -> HealListItem { fn heal_item(name: &str) -> HealListItem {
HealListItem { HealListItem {
name: name.to_string(), name: name.to_string(),
version_id: None, version_id: None,
mod_time_unix_nanos: None,
lifecycle_object_info: None,
is_delete_marker: false, is_delete_marker: false,
} }
} }
@@ -3357,6 +3583,13 @@ mod tests {
})) }))
} }
async fn erasure_set_usage_baseline(&self, _buckets: &[String]) -> Result<Option<HealBucketUsageBaseline>> {
if *self.usage_baseline_error.lock().unwrap() {
return Err(Error::Other("usage baseline unavailable".to_string()));
}
Ok(*self.usage_baseline.lock().unwrap())
}
async fn heal_bucket_metadata(&self, _bucket: &str) -> Result<()> { async fn heal_bucket_metadata(&self, _bucket: &str) -> Result<()> {
Ok(()) Ok(())
} }
@@ -3540,6 +3773,7 @@ mod tests {
bucket: &str, bucket: &str,
prefix: &str, prefix: &str,
continuation_token: Option<&str>, continuation_token: Option<&str>,
_include_lifecycle_object_info: bool,
) -> Result<(Vec<HealListItem>, Option<String>, bool)> { ) -> Result<(Vec<HealListItem>, Option<String>, bool)> {
self.listed_prefixes.lock().unwrap().push(prefix.to_string()); self.listed_prefixes.lock().unwrap().push(prefix.to_string());
if *self.truncate_without_token.lock().unwrap() { if *self.truncate_without_token.lock().unwrap() {
@@ -3715,6 +3949,69 @@ mod tests {
assert!(task.result_items_truncated()); assert!(task.result_items_truncated());
} }
// HS-06 (backlog#1870): incremental result windows.
#[tokio::test]
async fn result_items_seq_is_monotonic_and_incremental_slices_work() {
let storage = Arc::new(MockStorage::default());
let task = HealTask::from_request(HealRequest::bucket("bucket-a".to_string()), storage);
for round in 0..5u64 {
let item = HealResultItem {
object_size: round as usize,
..Default::default()
};
task.record_result_item(item).await;
}
let full = task.get_result_items_since(None).await;
assert_eq!(full.items.len(), 5, "None keeps the full-snapshot semantics");
assert_eq!(full.next_seq, 6, "next_seq is one past the last assigned");
assert_eq!(full.min_seq, 1, "nothing was evicted yet");
assert!(!full.lagged);
// Incremental: only items newer than the cursor.
let incremental = task.get_result_items_since(Some(3)).await;
assert_eq!(
incremental.items.iter().map(|item| item.object_size).collect::<Vec<_>>(),
vec![3, 4],
"only sequences greater than the cursor are returned"
);
assert_eq!(incremental.next_seq, 6);
// A cursor at the head is not lagging.
assert!(!task.get_result_items_since(Some(0)).await.lagged);
}
#[tokio::test]
async fn result_items_window_slide_moves_min_seq_and_flags_lagging_cursors() {
let storage = Arc::new(MockStorage::default());
let task = HealTask::from_request(HealRequest::bucket("bucket-a".to_string()), storage);
// Fill the window completely, then push two more items: seq 1 and 2
// are evicted by the slide.
for _ in 0..(MAX_RETAINED_HEAL_RESULT_ITEMS + 2) {
task.record_result_item(HealResultItem::default()).await;
}
let full = task.get_result_items_since(None).await;
assert_eq!(full.items.len(), MAX_RETAINED_HEAL_RESULT_ITEMS);
assert_eq!(full.min_seq, 3, "each evicted head item moved the oldest-available cursor");
assert!(task.result_items_truncated());
// A client still polling from before the eviction is lagging.
let lagging = task.get_result_items_since(Some(0)).await;
assert!(lagging.lagged, "a cursor behind min_seq must be flagged");
assert_eq!(lagging.min_seq, 3, "the response tells the client where to restart");
// A cursor inside the window is fine.
assert!(!task.get_result_items_since(Some(3)).await.lagged);
// The lagging client restarts from min_seq and gets the full window.
let catch_up = task.get_result_items_since(Some(3)).await;
assert_eq!(catch_up.items.len(), MAX_RETAINED_HEAL_RESULT_ITEMS - 1);
assert!(!catch_up.lagged);
}
#[tokio::test] #[tokio::test]
async fn test_recursive_bucket_heal_skips_object_dir_candidates() { async fn test_recursive_bucket_heal_skips_object_dir_candidates() {
let storage = Arc::new(MockStorage { let storage = Arc::new(MockStorage {
@@ -4654,6 +4951,73 @@ mod tests {
assert!(storage.object_heal_opts.lock().unwrap().is_empty()); assert!(storage.object_heal_opts.lock().unwrap().is_empty());
} }
#[tokio::test]
async fn erasure_set_heal_applies_usage_baseline_to_progress() {
let temp = TempDir::new().expect("temporary directory should be created");
let disk = make_resume_disk(&temp).await;
let storage = Arc::new(MockStorage {
resume_disk: Mutex::new(Some(disk)),
usage_baseline: Mutex::new(Some(HealBucketUsageBaseline {
objects_count: 10,
bytes: 8,
})),
..Default::default()
});
let request = HealRequest::new(
HealType::ErasureSet {
buckets: vec!["bucket-a".to_string()],
set_disk_id: "pool_0_set_0".to_string(),
},
HealOptions {
timeout: None,
..Default::default()
},
HealPriority::Normal,
);
let task = HealTask::from_request(request, storage);
task.heal_erasure_set(vec!["bucket-a".to_string()], "pool_0_set_0".to_string())
.await
.expect("erasure set heal should complete");
let progress = task.get_progress().await;
assert_eq!(progress.objects_total_count, 10);
assert_eq!(progress.objects_total_size, 8);
assert_eq!(progress.bytes_processed, 2);
assert!((progress.progress_percentage - 25.0).abs() < 0.001);
}
#[tokio::test]
async fn erasure_set_heal_ignores_usage_baseline_errors() {
let temp = TempDir::new().expect("temporary directory should be created");
let disk = make_resume_disk(&temp).await;
let storage = Arc::new(MockStorage {
resume_disk: Mutex::new(Some(disk)),
usage_baseline_error: Mutex::new(true),
..Default::default()
});
let request = HealRequest::new(
HealType::ErasureSet {
buckets: vec!["bucket-a".to_string()],
set_disk_id: "pool_0_set_0".to_string(),
},
HealOptions {
timeout: None,
..Default::default()
},
HealPriority::Normal,
);
let task = HealTask::from_request(request, storage);
task.heal_erasure_set(vec!["bucket-a".to_string()], "pool_0_set_0".to_string())
.await
.expect("usage baseline failures should not fail erasure set heal");
let progress = task.get_progress().await;
assert_eq!(progress.objects_total_count, 0);
assert_eq!(progress.objects_total_size, 0);
}
#[tokio::test] #[tokio::test]
async fn resumable_erasure_set_execution_is_cancelled_while_object_heal_is_pending() { async fn resumable_erasure_set_execution_is_cancelled_while_object_heal_is_pending() {
let temp = TempDir::new().expect("temporary directory should be created"); let temp = TempDir::new().expect("temporary directory should be created");
+5
View File
@@ -158,6 +158,10 @@ pub async fn init_heal_manager_with_workload_provider(
return Err(err); return Err(err);
} }
// Start the MRF intent consumer (error-path repair intents + durable
// journal replay) now that the manager can accept submissions.
heal::mrf_queue::spawn_mrf_consumer(heal_manager.clone());
#[cfg(test)] #[cfg(test)]
test_hook_after_manager_start().await; test_hook_after_manager_start().await;
@@ -445,6 +449,7 @@ mod tests {
_bucket: &str, _bucket: &str,
_prefix: &str, _prefix: &str,
_continuation_token: Option<&str>, _continuation_token: Option<&str>,
_include_lifecycle_object_info: bool,
) -> Result<(Vec<HealListItem>, Option<String>, bool), Error> { ) -> Result<(Vec<HealListItem>, Option<String>, bool), Error> {
Ok((Vec::new(), None, false)) Ok((Vec::new(), None, false))
} }
@@ -176,7 +176,7 @@ async fn enumerate_all_versions(heal_storage: &Arc<ECStoreHealStorage>, bucket:
let mut token: Option<String> = None; let mut token: Option<String> = None;
loop { loop {
let (page, next, truncated) = heal_storage let (page, next, truncated) = heal_storage
.list_objects_for_heal_page(bucket, "", token.as_deref()) .list_objects_for_heal_page(bucket, "", token.as_deref(), false)
.await .await
.expect("list_objects_for_heal_page failed"); .expect("list_objects_for_heal_page failed");
items.extend(page); items.extend(page);
@@ -166,7 +166,7 @@ async fn enumerate_b5(heal_storage: &Arc<ECStoreHealStorage>, bucket: &str) -> V
let mut token: Option<String> = None; let mut token: Option<String> = None;
loop { loop {
let (page, next, truncated) = heal_storage let (page, next, truncated) = heal_storage
.list_objects_for_heal_page(bucket, "", token.as_deref()) .list_objects_for_heal_page(bucket, "", token.as_deref(), false)
.await .await
.expect("b5 list page failed"); .expect("b5 list page failed");
items.extend(page); items.extend(page);
@@ -187,7 +187,7 @@ async fn enumerate_disk_walk(heal_storage: &Arc<ECStoreHealStorage>, bucket: &st
let mut token: Option<String> = None; let mut token: Option<String> = None;
loop { loop {
let (page, next, truncated) = heal_storage let (page, next, truncated) = heal_storage
.list_versions_for_heal_page_disk_walk(SET_DISK_ID, bucket, "", token.as_deref()) .list_versions_for_heal_page_disk_walk(SET_DISK_ID, bucket, "", token.as_deref(), false)
.await .await
.expect("disk-walk list page failed"); .expect("disk-walk list page failed");
items.extend(page); items.extend(page);
@@ -418,7 +418,7 @@ mod serial_tests {
let mut pages = 0usize; let mut pages = 0usize;
loop { loop {
let (versions, next_forward, truncated) = ecstore let (versions, next_forward, truncated) = ecstore
.heal_walk_versions_page(0, 0, bucket, "", forward.as_deref(), 2, 100_000) .heal_walk_versions_page(0, 0, bucket, "", forward.as_deref(), 2, 100_000, false)
.await .await
.expect("heal_walk_versions_page failed"); .expect("heal_walk_versions_page failed");
pages += 1; pages += 1;
+2
View File
@@ -242,6 +242,7 @@ fn test_heal_task_status_atomic_update() {
_bucket: &str, _bucket: &str,
_prefix: &str, _prefix: &str,
_continuation_token: Option<&str>, _continuation_token: Option<&str>,
_include_lifecycle_object_info: bool,
) -> rustfs_heal::Result<(Vec<HealListItem>, Option<String>, bool)> { ) -> rustfs_heal::Result<(Vec<HealListItem>, Option<String>, bool)> {
Ok((vec![], None, false)) Ok((vec![], None, false))
} }
@@ -385,6 +386,7 @@ async fn test_heal_task_transient_object_exists_skip_avoids_recreate() {
_bucket: &str, _bucket: &str,
_prefix: &str, _prefix: &str,
_continuation_token: Option<&str>, _continuation_token: Option<&str>,
_include_lifecycle_object_info: bool,
) -> rustfs_heal::Result<(Vec<HealListItem>, Option<String>, bool)> { ) -> rustfs_heal::Result<(Vec<HealListItem>, Option<String>, bool)> {
Ok((Vec::new(), None, false)) Ok((Vec::new(), None, false))
} }
+189
View File
@@ -0,0 +1,189 @@
// Copyright 2024 RustFS Team
//
// Licensed under the Apache License, Version 2.0 (the "License");
// you may not use this file except in compliance with the License.
// You may obtain a copy of the License at
//
// http://www.apache.org/licenses/LICENSE-2.0
//
// Unless required by applicable law or agreed to in writing, software
// distributed under the License is distributed on an "AS IS" BASIS,
// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
// See the License for the specific language governing permissions and
// limitations under the License.
//! HS-01 (rustfs/backlog#1865): MRF intent pipeline integration tests.
//!
//! Drives the real consumer loop (`spawn_mrf_consumer`) against a real
//! 4-disk `ECStore` heal storage and a `HealManager` that has not started its
//! scheduler, so submitted intents stay observable in the admission queue.
//! Under `cargo nextest` each test runs in its own process, which keeps the
//! process-global MRF channel singleton safe.
use rustfs_common::mrf_channel::{self, MrfKind};
use rustfs_heal::heal::{
manager::{HealConfig, HealManager},
mrf_queue,
storage::{ECStoreHealStorage, HealStorageAPI},
};
use serial_test::serial;
use std::{path::Path, sync::Arc, time::Duration};
mod storage_api;
use storage_api::endpoint_index::{Endpoint, EndpointServerPools, Endpoints, PoolEndpoints, init_local_disks};
const META_BUCKET: &str = ".rustfs.sys";
const JOURNAL_REL: &str = "buckets/.heal/mrf/journal.bin";
async fn heal_env() -> (Vec<std::path::PathBuf>, Arc<dyn HealStorageAPI>) {
let env = rustfs_test_utils::TestECStoreEnv::builder()
.prefix("rustfs_heal_mrf_test")
.build()
.await;
let heal_storage: Arc<dyn HealStorageAPI> = Arc::new(ECStoreHealStorage::new(env.ecstore.clone()));
(env.disk_paths, heal_storage)
}
fn make_manager(storage: Arc<dyn HealStorageAPI>) -> Arc<HealManager> {
Arc::new(HealManager::new(
storage,
Some(HealConfig {
// Keep the scheduler from draining the queue before assertions.
heal_interval: Duration::from_secs(3600),
enable_auto_heal: false,
..Default::default()
}),
))
}
/// Encode one journal record independently of the implementation, so a format
/// drift between writer and this fixture fails loudly here.
fn journal_record(kind: u8, bucket: &str, object: &str, version: Option<[u8; 16]>, attempts: u8) -> Vec<u8> {
let mut body = vec![1u8, 1, kind, attempts];
body.extend_from_slice(&1_700_000_000_000u64.to_le_bytes());
match version {
Some(bytes) => {
body.push(1);
body.extend_from_slice(&bytes);
}
None => body.push(0),
}
body.extend_from_slice(&(bucket.len() as u32).to_le_bytes());
body.extend_from_slice(&(object.len() as u32).to_le_bytes());
body.extend_from_slice(bucket.as_bytes());
body.extend_from_slice(object.as_bytes());
let mut hasher = crc_fast::Digest::new(crc_fast::CrcAlgorithm::Crc32IsoHdlc);
hasher.update(&body);
body.extend_from_slice(&(hasher.finalize() as u32).to_le_bytes());
body
}
fn write_journal_to_disks(disk_paths: &[std::path::PathBuf], data: &[u8]) {
for path in disk_paths {
let journal = path.join(META_BUCKET).join(JOURNAL_REL);
std::fs::create_dir_all(journal.parent().expect("journal parent")).expect("create journal dir");
std::fs::write(&journal, data).expect("write journal fixture");
}
}
async fn wait_until<F, Fut>(deadline: Duration, mut probe: F) -> bool
where
F: FnMut() -> Fut,
Fut: std::future::Future<Output = bool>,
{
let start = std::time::Instant::now();
while start.elapsed() < deadline {
if probe().await {
return true;
}
tokio::time::sleep(Duration::from_millis(50)).await;
}
false
}
/// A decode-failure intent delivered on the global channel must surface in the
/// heal manager as an Urgent request attributed to the MRF source.
#[tokio::test]
#[serial]
async fn decode_failure_intent_maps_to_urgent_mrf_heal_request() {
let (_disk_paths, storage) = heal_env().await;
let manager = make_manager(storage);
mrf_queue::spawn_mrf_consumer(manager.clone());
assert!(
mrf_channel::try_send_mrf_intent(MrfKind::DecodeFailure, "mrf-bucket", "mrf-object", None),
"intent should be accepted while the consumer holds the channel"
);
let appeared = wait_until(Duration::from_secs(10), || async {
let snapshot = manager.operations_snapshot().await;
snapshot.queued_by_source.mrf >= 1 && snapshot.queued_by_priority.urgent >= 1
})
.await;
assert!(
appeared,
"MRF intent must reach the manager queue as an Urgent request (snapshot: {:?})",
manager.operations_snapshot().await
);
}
/// A journal left behind by a previous process must be replayed into the
/// manager queue and then removed, and a torn tail must not block replay of
/// the intact records.
#[tokio::test(flavor = "multi_thread", worker_threads = 4)]
#[serial]
async fn journal_replay_arms_intents_and_deletes_the_file() {
let (disk_paths, storage) = heal_env().await;
// The journal reader resolves disks through the process-local disk map;
// register the environment's disks the same way server startup does.
let mut endpoints: Vec<Endpoint> = disk_paths
.iter()
.map(|p| Endpoint::try_from(p.to_string_lossy().as_ref()).expect("endpoint from disk path"))
.collect();
for (i, endpoint) in endpoints.iter_mut().enumerate() {
endpoint.set_pool_index(0);
endpoint.set_set_index(0);
endpoint.set_disk_index(i);
}
let pool = PoolEndpoints {
legacy: false,
set_count: 1,
drives_per_set: endpoints.len(),
endpoints: Endpoints::from(endpoints),
cmd_line: "mrf-test".to_string(),
platform: String::new(),
};
init_local_disks(EndpointServerPools::from(vec![pool]))
.await
.expect("local disks should register");
let mut journal = journal_record(1, "replay-bucket", "replay-object", Some([9u8; 16]), 0);
journal.extend(journal_record(3, "replay-bucket", "partial-object", None, 1));
// Torn tail: a third record truncated mid-way must not block the two
// intact records above.
journal.extend_from_slice(&journal_record(2, "replay-bucket", "metadata-object", None, 0)[..8]);
write_journal_to_disks(&disk_paths, &journal);
let manager = make_manager(storage);
// Replay directly (not via the process-global channel consumer, which the
// sibling test already claimed in this process under plain `cargo test`).
let replayed = mrf_queue::replay_journal_once(&manager).await;
assert_eq!(replayed, 2, "the two intact records must be replayed");
let snapshot = manager.operations_snapshot().await;
assert_eq!(snapshot.queued_by_source.mrf, 2, "replayed intents must be attributed to the MRF source");
assert!(
disk_paths
.iter()
.all(|path| !Path::new(path).join(META_BUCKET).join(JOURNAL_REL).exists()),
"the journal file must be removed after a successful replay"
);
let snapshot = manager.operations_snapshot().await;
assert_eq!(snapshot.queued_by_priority.urgent, 1, "the decode-failure record must replay as Urgent");
assert!(snapshot.queued_by_priority.normal >= 1, "the partial-write record must replay as Normal");
}
+3
View File
@@ -9,6 +9,9 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
### Removed ### Removed
#### rustfs-io-core
- **Zero-consumer modules** (added in 0.0.5): `reader`, `writer`, `bufreader_optimizer`, `shared_memory`, `direct_io`, `timeout_wrapper`, `io_priority_queue`, and `scheduler` had no caller in the workspace and were removed (rustfs/backlog#1824). The scheduling algorithm and the request timeout wrapper that RustFS actually runs live in `rustfs/src/storage/`; this crate keeps the config shapes they project into. `OperationProgress` moved to the new `progress` module and is still exported as `rustfs_io_core::OperationProgress`.
#### rustfs-io-metrics #### rustfs-io-metrics
- **Unified configuration** (added in 0.0.5): the zero-consumer `IoConfig`, `CacheSettings`, `IoSchedulerSettings`, `BackpressureSettings`, `TimeoutSettings`, `DeadlockDetectionSettings` types and their `DEFAULT_*` constants were removed (rustfs/rustfs#6008); rustfs-io-core's `IoSchedulerConfig`/`BackpressureConfig` remain the canonical configuration types. - **Unified configuration** (added in 0.0.5): the zero-consumer `IoConfig`, `CacheSettings`, `IoSchedulerSettings`, `BackpressureSettings`, `TimeoutSettings`, `DeadlockDetectionSettings` types and their `DEFAULT_*` constants were removed (rustfs/rustfs#6008); rustfs-io-core's `IoSchedulerConfig`/`BackpressureConfig` remain the canonical configuration types.
+2 -3
View File
@@ -20,8 +20,8 @@ license.workspace = true
repository.workspace = true repository.workspace = true
rust-version.workspace = true rust-version.workspace = true
homepage.workspace = true homepage.workspace = true
description = "Buffered I/O reader and writer implementations for RustFS (mmap-then-copy, aligned pread)" description = "Shared I/O primitives for RustFS (buffer pool, storage profiling, backpressure, deadlock detection)"
keywords = ["io", "reader", "writer", "rustfs", "mmap"] keywords = ["io", "buffer", "pool", "rustfs", "backpressure"]
categories = ["development-tools", "filesystem"] categories = ["development-tools", "filesystem"]
[lints] [lints]
@@ -38,7 +38,6 @@ hotpath.workspace = true
bytes = { workspace = true, features = ["serde"] } bytes = { workspace = true, features = ["serde"] }
thiserror = { workspace = true } thiserror = { workspace = true }
tokio = { workspace = true, features = ["io-util", "fs", "sync", "rt-multi-thread"] } tokio = { workspace = true, features = ["io-util", "fs", "sync", "rt-multi-thread"] }
memmap2 = { workspace = true }
rustfs-io-metrics = { workspace = true } rustfs-io-metrics = { workspace = true }
tracing = { workspace = true } tracing = { workspace = true }
+18 -120
View File
@@ -23,67 +23,20 @@
## Overview ## Overview
**rustfs-io-core** is the core I/O scheduling module for [RustFS](https://rustfs.com), a distributed object storage system. It provides: **rustfs-io-core** holds the shared I/O primitives for [RustFS](https://rustfs.com), a distributed object storage system. It provides:
- **I/O Scheduler**: Adaptive buffer size calculation and load management - **Buffer Pool**: Tiered `BytesPool` for buffer reuse
- **Priority Queue**: Request priority scheduling with starvation prevention - **Storage Profiling**: Storage-media and access-pattern model (`io_profile`)
- **Scheduler Configuration**: The `IoSchedulerConfig` / `IoPriorityQueueConfig` shapes the storage layer projects into
- **Backpressure Control**: System overload protection with graceful degradation - **Backpressure Control**: System overload protection with graceful degradation
- **Deadlock Detection**: Wait-for graph based deadlock detection algorithm - **Deadlock Detection**: Wait-for graph based deadlock detection algorithm
- **Lock Optimizer**: Adaptive spin lock optimization - **Lock Optimizer**: Adaptive spin lock optimization
- **Timeout Wrapper**: Dynamic timeout calculation and operation progress tracking - **Progress Tracking**: Byte progress and staleness for long-running operations
The scheduling algorithm itself lives in `rustfs/src/storage/concurrency/io_schedule.rs`; this crate carries the configuration shapes it projects into, not a second implementation.
## Features ## Features
### I/O Scheduler
Adaptive I/O scheduling with dynamic buffer size calculation based on file size, access pattern, and system load:
```rust
use rustfs_io_core::{IoScheduler, IoSchedulerConfig, IoLoadLevel};
use rustfs_io_core::io_profile::{StorageMedia, AccessPattern};
// Create scheduler
let config = IoSchedulerConfig {
max_concurrent_reads: 64,
base_buffer_size: 64 * 1024, // 64 KB
max_buffer_size: 1024 * 1024, // 1 MB
..Default::default()
};
let scheduler = IoScheduler::new(config);
// Calculate optimal buffer size
let buffer_size = calculate_optimal_buffer_size(
10 * 1024 * 1024, // 10 MB file
64 * 1024, // base buffer
true, // sequential access
4, // concurrent requests
StorageMedia::Ssd,
IoLoadLevel::Low,
);
```
### Priority Queue
Priority queue with starvation prevention:
```rust
use rustfs_io_core::{IoPriorityQueue, IoPriority, IoQueueStatus};
let queue = IoPriorityQueue::<()>::new(100);
// Enqueue request
let request_id = queue.enqueue(IoPriority::High, (), 1024);
// Dequeue request
if let Some((priority, data)) = queue.dequeue() {
println!("Processing priority {:?} request", priority);
}
// Check queue status
let status = queue.status();
println!("High priority waiting: {}", status.high_priority_waiting);
```
### Backpressure Control ### Backpressure Control
System overload protection: System overload protection:
@@ -148,71 +101,23 @@ let stats = optimizer.stats();
println!("Locks acquired: {}", stats.total_acquired()); println!("Locks acquired: {}", stats.total_acquired());
``` ```
### Timeout Wrapper ### Progress Tracking
Dynamic timeout calculation: Byte progress and staleness for long-running operations:
```rust ```rust
use rustfs_io_core::{RequestTimeoutWrapper, TimeoutConfig}; use rustfs_io_core::OperationProgress;
use std::time::Duration; use std::time::Duration;
let config = TimeoutConfig { let progress = OperationProgress::new(Some(1000), Duration::from_secs(5));
base_timeout: Duration::from_secs(5),
timeout_per_mb: Duration::from_millis(100),
max_timeout: Duration::from_secs(300),
..Default::default()
};
let wrapper = RequestTimeoutWrapper::new(config);
// Calculate operation timeout progress.update(500);
let timeout = wrapper.calculate_timeout(10 * 1024 * 1024); // 10 MB assert_eq!(progress.progress_percent(), Some(50.0));
``` assert!(!progress.is_stale());
## Buffer Size Calculation
Multiple buffer size calculation functions are provided:
```rust
use rustfs_io_core::{
get_concurrency_aware_buffer_size,
get_advanced_buffer_size,
get_buffer_size_for_media,
calculate_optimal_buffer_size,
KI_B, MI_B,
};
use rustfs_io_core::io_profile::StorageMedia;
// Basic calculation
let size1 = get_concurrency_aware_buffer_size(1024 * 1024, 64 * 1024);
// Advanced calculation (considering access pattern)
let size2 = get_advanced_buffer_size(10 * 1024 * 1024, 64 * 1024, true);
// Media type optimization
let size3 = get_buffer_size_for_media(64 * 1024, StorageMedia::Ssd);
// Comprehensive calculation
let size4 = calculate_optimal_buffer_size(
100 * 1024 * 1024, // 100 MB file
64 * 1024, // base buffer
true, // sequential access
4, // concurrent requests
StorageMedia::Nvme,
IoLoadLevel::Low,
);
``` ```
## Configuration ## Configuration
### Environment Variables
| Variable | Description | Default |
|----------|-------------|---------|
| `RUSTFS_MAX_CONCURRENT_READS` | Max concurrent reads | 64 |
| `RUSTFS_BASE_BUFFER_SIZE` | Base buffer size | 65536 |
| `RUSTFS_MAX_BUFFER_SIZE` | Max buffer size | 1048576 |
| `RUSTFS_IO_TIMEOUT_SECS` | I/O timeout seconds | 30 |
### Code Configuration ### Code Configuration
```rust ```rust
@@ -240,12 +145,11 @@ rustfs-io-core/
├── src/ ├── src/
│ ├── lib.rs # Module entry │ ├── lib.rs # Module entry
│ ├── config.rs # Configuration types │ ├── config.rs # Configuration types
│ ├── scheduler.rs # I/O scheduler │ ├── pool.rs # Tiered buffer pool
│ ├── io_priority_queue.rs # Priority queue
│ ├── backpressure.rs # Backpressure control │ ├── backpressure.rs # Backpressure control
│ ├── deadlock_detector.rs # Deadlock detection │ ├── deadlock_detector.rs # Deadlock detection
│ ├── lock_optimizer.rs # Lock optimization │ ├── lock_optimizer.rs # Lock optimization
│ ├── timeout_wrapper.rs # Timeout wrapper │ ├── progress.rs # Operation progress tracking
│ └── io_profile.rs # I/O profile │ └── io_profile.rs # I/O profile
└── Cargo.toml └── Cargo.toml
``` ```
@@ -254,21 +158,15 @@ rustfs-io-core/
```bash ```bash
# Run all tests # Run all tests
cargo test --package rustfs-io-core cargo nextest run --package rustfs-io-core
# Run specific tests # Run specific tests
cargo test --package rustfs-io-core --lib scheduler cargo nextest run --package rustfs-io-core -E 'test(backpressure)'
# Run benchmarks
cargo bench --package rustfs-io-core
``` ```
## Documentation ## Documentation
- [API Documentation](https://docs.rs/rustfs-io-core) - [API Documentation](https://docs.rs/rustfs-io-core)
- [I/O Scheduler Design](./docs/scheduler-design.md)
- [Backpressure Control Design](./docs/backpressure-design.md)
- [Deadlock Detection Algorithm](./docs/deadlock-detection.md)
## Related Modules ## Related Modules
+18 -131
View File
@@ -23,71 +23,20 @@
## 📖 概述 ## 📖 概述
**rustfs-io-core** 是 [RustFS](https://rustfs.com) 分布式对象存储系统的核心 I/O 调度模块。它提供了: **rustfs-io-core** 是 [RustFS](https://rustfs.com) 分布式对象存储系统的共享 I/O 基础组件。它提供了:
- **I/O 调度器**:自适应缓冲区大小计算和负载管理 - **缓冲池**:分级复用的 `BytesPool`
- **优先级队列**支持饥饿预防的请求优先级调度 - **存储画像**存储介质与访问模式模型(`io_profile`
- **调度配置**:存储层投影使用的 `IoSchedulerConfig` / `IoPriorityQueueConfig`
- **背压控制**:系统过载保护和优雅降级 - **背压控制**:系统过载保护和优雅降级
- **死锁检测**:基于等待图的死锁检测算法 - **死锁检测**:基于等待图的死锁检测算法
- **锁优化**:自适应自旋锁优化 - **锁优化**:自适应自旋锁优化
- **超时包装器**动态超时计算和操作进度追踪 - **进度追踪**长耗时操作的字节进度与停滞判定
调度算法本身位于 `rustfs/src/storage/concurrency/io_schedule.rs`;本 crate 只承载它投影使用的配置形状,不是第二套实现。
## ✨ 核心功能 ## ✨ 核心功能
### I/O 调度器 (IoScheduler)
自适应 I/O 调度,根据文件大小、访问模式和系统负载动态调整缓冲区大小:
```rust
use rustfs_io_core::{IoScheduler, IoSchedulerConfig, IoLoadLevel};
use rustfs_io_core::io_profile::{StorageMedia, AccessPattern};
// 创建调度器
let config = IoSchedulerConfig {
max_concurrent_reads: 64,
base_buffer_size: 64 * 1024, // 64 KB
max_buffer_size: 1024 * 1024, // 1 MB
..Default::default()
};
let scheduler = IoScheduler::new(config);
// 计算最优缓冲区大小
let buffer_size = scheduler.calculate_buffer_size(
10 * 1024 * 1024, // 10 MB 文件
true, // 顺序访问
StorageMedia::Ssd,
IoLoadLevel::Low,
);
println!("缓冲区大小: {} bytes", buffer_size);
```
### 优先级队列 (IoPriorityQueue)
支持饥饿预防的优先级队列:
```rust
use rustfs_io_core::{IoPriorityQueue, IoPriority, IoQueueStatus};
let queue = IoPriorityQueue::<()>::new(100);
// 入队请求
let request_id = queue.enqueue(
IoPriority::High,
(), // 请求数据
1024, // 请求大小
);
// 出队请求
if let Some((priority, data)) = queue.dequeue() {
println!("处理优先级 {:?} 的请求", priority);
}
// 检查队列状态
let status = queue.status();
println!("高优先级等待: {}", status.high_priority_waiting);
println!("低优先级等待: {}", status.low_priority_waiting);
```
### 背压控制 (BackpressureMonitor) ### 背压控制 (BackpressureMonitor)
系统过载保护: 系统过载保护:
@@ -165,78 +114,23 @@ let stats = optimizer.stats();
println!("获取锁次数: {}", stats.locks_acquired.load(std::sync::atomic::Ordering::Relaxed)); println!("获取锁次数: {}", stats.locks_acquired.load(std::sync::atomic::Ordering::Relaxed));
``` ```
### 超时包装器 (RequestTimeoutWrapper) ### 进度追踪 (OperationProgress)
动态超时计算 长耗时操作的字节进度与停滞判定
```rust ```rust
use rustfs_io_core::{RequestTimeoutWrapper, TimeoutConfig}; use rustfs_io_core::OperationProgress;
use std::time::Duration; use std::time::Duration;
let config = TimeoutConfig { let progress = OperationProgress::new(Some(1000), Duration::from_secs(5));
base_timeout: Duration::from_secs(5),
timeout_per_mb: Duration::from_millis(100),
max_timeout: Duration::from_secs(300),
..Default::default()
};
let wrapper = RequestTimeoutWrapper::new(config);
// 计算操作超时 progress.update(500);
let timeout = wrapper.calculate_timeout(10 * 1024 * 1024); // 10 MB assert_eq!(progress.progress_percent(), Some(50.0));
println!("超时时间: {:?}", timeout); assert!(!progress.is_stale());
// 执行带超时的操作
let result = wrapper.execute_with_timeout(async {
// 异步操作
Ok::<_, std::io::Error>(())
}, timeout).await;
```
## 📊 缓冲区大小计算
模块提供了多种缓冲区大小计算函数:
```rust
use rustfs_io_core::{
get_concurrency_aware_buffer_size,
get_advanced_buffer_size,
get_buffer_size_for_media,
calculate_optimal_buffer_size,
KI_B, MI_B,
};
use rustfs_io_core::io_profile::StorageMedia;
// 基础计算
let size1 = get_concurrency_aware_buffer_size(1024 * 1024, 64 * 1024);
// 高级计算(考虑访问模式)
let size2 = get_advanced_buffer_size(10 * 1024 * 1024, 64 * 1024, true);
// 媒体类型优化
let size3 = get_buffer_size_for_media(64 * 1024, StorageMedia::Ssd);
// 综合计算
let size4 = calculate_optimal_buffer_size(
100 * 1024 * 1024, // 100 MB 文件
64 * 1024, // 基础缓冲区
true, // 顺序访问
4, // 并发请求数
StorageMedia::Nvme,
IoLoadLevel::Low,
);
``` ```
## 🔧 配置 ## 🔧 配置
### 环境变量
| 变量名 | 描述 | 默认值 |
|--------|------|--------|
| `RUSTFS_MAX_CONCURRENT_READS` | 最大并发读数 | 64 |
| `RUSTFS_BASE_BUFFER_SIZE` | 基础缓冲区大小 | 65536 |
| `RUSTFS_MAX_BUFFER_SIZE` | 最大缓冲区大小 | 1048576 |
| `RUSTFS_IO_TIMEOUT_SECS` | I/O 超时秒数 | 30 |
### 代码配置 ### 代码配置
```rust ```rust
@@ -264,12 +158,11 @@ rustfs-io-core/
├── src/ ├── src/
│ ├── lib.rs # 模块入口 │ ├── lib.rs # 模块入口
│ ├── config.rs # 配置类型 │ ├── config.rs # 配置类型
│ ├── scheduler.rs # I/O 调度器 │ ├── pool.rs # 分级缓冲池
│ ├── io_priority_queue.rs # 优先级队列
│ ├── backpressure.rs # 背压控制 │ ├── backpressure.rs # 背压控制
│ ├── deadlock_detector.rs # 死锁检测 │ ├── deadlock_detector.rs # 死锁检测
│ ├── lock_optimizer.rs # 锁优化 │ ├── lock_optimizer.rs # 锁优化
│ ├── timeout_wrapper.rs # 超时包装器 │ ├── progress.rs # 操作进度追踪
│ └── io_profile.rs # I/O 配置文件 │ └── io_profile.rs # I/O 配置文件
└── Cargo.toml └── Cargo.toml
``` ```
@@ -278,21 +171,15 @@ rustfs-io-core/
```bash ```bash
# 运行所有测试 # 运行所有测试
cargo test --package rustfs-io-core cargo nextest run --package rustfs-io-core
# 运行特定测试 # 运行特定测试
cargo test --package rustfs-io-core --lib scheduler cargo nextest run --package rustfs-io-core -E 'test(backpressure)'
# 运行基准测试
cargo bench --package rustfs-io-core
``` ```
## 📚 文档 ## 📚 文档
- [API 文档](https://docs.rs/rustfs-io-core) - [API 文档](https://docs.rs/rustfs-io-core)
- [I/O 调度器设计](./docs/scheduler-design.md)
- [背压控制原理](./docs/backpressure-design.md)
- [死锁检测算法](./docs/deadlock-detection.md)
## 🔗 相关模块 ## 🔗 相关模块
@@ -1,190 +0,0 @@
// Copyright 2024 RustFS Team
//
// Licensed under the Apache License, Version 2.0 (the "License");
// you may not use this file except in compliance with the License.
// You may obtain a copy of the License at
//
// http://www.apache.org/licenses/LICENSE-2.0
//
// Unless required by applicable law or agreed to in writing, software
// distributed under the License is distributed on an "AS IS" BASIS,
// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
// See the License for the specific language governing permissions and
// limitations under the License.
//! Example demonstrating I/O scheduler usage.
use rustfs_io_core::io_profile::StorageMedia;
use rustfs_io_core::{
BackpressureMonitor, BackpressureState, DeadlockDetector, IoLoadLevel, IoScheduler, IoSchedulerConfig, KI_B, LockOptimizer,
LockType, MI_B, calculate_optimal_buffer_size, get_buffer_size_for_media,
};
use std::time::Duration;
fn main() {
println!("=== rustfs-io-core Example ===\n");
// 1. I/O scheduler example
io_scheduler_example();
// 2. Buffer size calculation example
buffer_size_example();
// 3. Backpressure control example
backpressure_example();
// 4. Deadlock detection example
deadlock_detection_example();
// 5. Lock optimizer example
lock_optimizer_example();
}
fn io_scheduler_example() {
println!("--- I/O Scheduler ---");
// Create scheduler with configuration
let config = IoSchedulerConfig {
max_concurrent_reads: 64,
base_buffer_size: 64 * KI_B,
max_buffer_size: MI_B,
..Default::default()
};
let scheduler = IoScheduler::new(config);
println!(" Max concurrent reads: {}", scheduler.config().max_concurrent_reads);
println!(" Base buffer size: {} KB", scheduler.config().base_buffer_size / KI_B);
println!(" Max buffer size: {} KB", scheduler.config().max_buffer_size / KI_B);
// Calculate buffer sizes for different scenarios
let scenarios = [
("Small file", 10 * KI_B as i64, true, StorageMedia::Ssd),
("Medium file", MI_B as i64, true, StorageMedia::Ssd),
("Large sequential", 100 * MI_B as i64, true, StorageMedia::Ssd),
("Large random", 100 * MI_B as i64, false, StorageMedia::Ssd),
("NVMe large", 100 * MI_B as i64, true, StorageMedia::Nvme),
("HDD large", 100 * MI_B as i64, true, StorageMedia::Hdd),
];
for (name, size, sequential, media) in scenarios {
let buffer = calculate_optimal_buffer_size(size, 64 * KI_B, sequential, 4, media, IoLoadLevel::Low);
println!(" {}: {} bytes ({} KB)", name, buffer, buffer / KI_B);
}
println!();
}
fn buffer_size_example() {
println!("--- Buffer Size Calculation ---");
// Comprehensive calculation
let size1 = calculate_optimal_buffer_size(10 * MI_B as i64, 64 * KI_B, true, 4, StorageMedia::Ssd, IoLoadLevel::Low);
println!(" Comprehensive (10MB, sequential, SSD): {} KB", size1 / KI_B);
// Media type optimization
let media_types = [
StorageMedia::Nvme,
StorageMedia::Ssd,
StorageMedia::Hdd,
StorageMedia::Unknown,
];
for media in media_types {
let size = get_buffer_size_for_media(64 * KI_B, media);
println!(" {} optimized: {} KB", media.as_str(), size / KI_B);
}
println!();
}
fn backpressure_example() {
println!("--- Backpressure Control ---");
let monitor = BackpressureMonitor::with_defaults();
// Check initial state
let state = monitor.state();
let state_str = match state {
BackpressureState::Normal => "Normal",
BackpressureState::Warning => "Warning",
BackpressureState::Critical => "Critical",
};
println!(" Initial state: {}", state_str);
// Check if active
let is_active = monitor.is_active();
println!(" Backpressure active: {}", is_active);
// Try to acquire permit
if monitor.try_acquire() {
println!(" Successfully acquired permit");
monitor.release();
println!(" Released permit");
}
// View statistics
println!(" Total processed: {}", monitor.total_processed());
println!(" Total rejected: {}", monitor.total_rejected());
println!();
}
fn deadlock_detection_example() {
println!("--- Deadlock Detection ---");
let detector = DeadlockDetector::with_defaults();
// Register locks
let mutex1 = detector.register_lock(LockType::Mutex);
let mutex2 = detector.register_lock(LockType::Mutex);
println!(" Registered locks: mutex1={}, mutex2={}", mutex1, mutex2);
// Simulate normal operation
detector.record_acquire(mutex1, 1); // Thread 1 acquires mutex1
detector.record_acquire(mutex2, 2); // Thread 2 acquires mutex2
println!(" Normal operation: no deadlock");
// Detect deadlock
if detector.detect_deadlock().is_none() {
println!(" Detection result: no deadlock");
}
// Simulate deadlock scenario
detector.record_wait(mutex2, 1); // Thread 1 waits for mutex2
detector.record_wait(mutex1, 2); // Thread 2 waits for mutex1
// Detect deadlock
if let Some(deadlock) = detector.detect_deadlock() {
println!(" Detection result: deadlock found {:?}", deadlock);
}
// Cleanup
detector.unregister_lock(mutex1);
detector.unregister_lock(mutex2);
println!();
}
fn lock_optimizer_example() {
println!("--- Lock Optimizer ---");
let optimizer = LockOptimizer::with_defaults();
// Simulate lock operations
for _i in 0..5 {
optimizer.on_acquire();
// Simulate work
std::thread::sleep(Duration::from_millis(10));
optimizer.on_release(Duration::from_millis(10));
}
// View statistics
let stats = optimizer.stats();
let acquired = stats.total_acquired();
let avg_hold = stats.avg_hold_time();
let contention = stats.contention_rate();
println!(" Locks acquired: {}", acquired);
println!(" Average hold time: {:?}", avg_hold);
println!(" Contention rate: {:.2}%", contention * 100.0);
println!();
}
-227
View File
@@ -1,227 +0,0 @@
// Copyright 2024 RustFS Team
//
// Licensed under the Apache License, Version 2.0 (the "License");
// you may not use this file except in compliance with the License.
// You may obtain a copy of the License at
//
// http://www.apache.org/licenses/LICENSE-2.0
//
// Unless required by applicable law or agreed to in writing, software
// distributed under the License is distributed on an "AS IS" BASIS,
// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
// See the License for the specific language governing permissions and
// limitations under the License.
//! BufReader layer optimizer for minimizing redundant buffering layers.
//!
//! This module provides optimization for BufReader usage in data paths,
//! including layer count limiting and dynamic buffer size adjustment.
use std::sync::atomic::{AtomicU64, Ordering};
/// BufReader optimization configuration.
#[derive(Debug, Clone)]
pub struct BufReaderConfig {
/// Maximum number of nested BufReader layers (default: 2)
pub max_layers: u32,
/// Buffer size for small files (default: 8KB)
pub small_file_buffer: usize,
/// Buffer size for large files (default: 64KB)
pub large_file_buffer: usize,
/// Threshold for large file classification (default: 1MB)
pub large_file_threshold: usize,
}
impl Default for BufReaderConfig {
fn default() -> Self {
Self {
max_layers: 2,
small_file_buffer: 8 * 1024, // 8KB
large_file_buffer: 64 * 1024, // 64KB
large_file_threshold: 1024 * 1024, // 1MB
}
}
}
/// BufReader optimization statistics.
#[derive(Debug, Default)]
pub struct BufReaderStats {
/// Total number of readers created
pub total_readers: AtomicU64,
/// Number of redundant layers eliminated
pub eliminated_layers: AtomicU64,
/// Number of buffer size adjustments
pub buffer_size_adjustments: AtomicU64,
}
/// BufReader layer optimizer.
///
/// Analyzes and optimizes BufReader nesting in data paths,
/// dynamically adjusting buffer sizes based on data characteristics.
pub struct BufReaderOptimizer {
config: BufReaderConfig,
stats: BufReaderStats,
}
impl BufReaderOptimizer {
/// Create a new BufReader optimizer with the given configuration.
pub fn new(config: BufReaderConfig) -> Self {
Self {
config,
stats: BufReaderStats::default(),
}
}
/// Create a new BufReader optimizer with default configuration.
pub fn with_defaults() -> Self {
Self::new(BufReaderConfig::default())
}
/// Calculate the optimal buffer size based on data size.
///
/// Returns the appropriate buffer size based on whether the data
/// is classified as a small or large file.
pub fn optimal_buffer_size(&self, data_size: Option<usize>) -> usize {
match data_size {
Some(size) if size >= self.config.large_file_threshold => self.config.large_file_buffer,
Some(_) => self.config.small_file_buffer,
None => self.config.small_file_buffer,
}
}
/// Optimize a reader by wrapping it with an appropriately sized BufReader.
///
/// This method applies the optimal buffer size based on the expected
/// data size and tracks statistics.
pub fn optimize<R: tokio::io::AsyncRead + Unpin>(&self, reader: R, data_size: Option<usize>) -> tokio::io::BufReader<R> {
let buffer_size = self.optimal_buffer_size(data_size);
self.stats.total_readers.fetch_add(1, Ordering::Relaxed);
tokio::io::BufReader::with_capacity(buffer_size, reader)
}
/// Get the statistics for this optimizer.
pub fn stats(&self) -> &BufReaderStats {
&self.stats
}
/// Get the configuration for this optimizer.
pub fn config(&self) -> &BufReaderConfig {
&self.config
}
}
/// Marker trait for buffered sources.
///
/// Types implementing this trait are considered already buffered
/// and should not be wrapped with additional BufReader layers.
pub trait BufferedSource: tokio::io::AsyncRead {}
impl BufReaderOptimizer {
/// Check if a reader is already a buffered source.
///
/// Returns true if the reader implements `BufferedSource`,
/// indicating it should not be wrapped with BufReader.
pub fn is_buffered_source<R: BufferedSource + ?Sized>(&self, _reader: &R) -> bool {
true
}
/// Eliminate redundant BufReader layers if possible.
///
/// This method attempts to reduce the nesting depth of BufReader
/// layers to improve performance.
pub fn eliminate_redundant_layers<R: tokio::io::AsyncRead + Unpin>(&self, reader: R) -> R {
// For now, just return the reader as-is
// Future implementation could detect and unwrap nested BufReaders
self.stats.eliminated_layers.fetch_add(0, Ordering::Relaxed);
reader
}
}
#[cfg(test)]
mod tests {
use super::*;
use tokio::io::AsyncReadExt;
#[test]
fn test_default_config() {
let config = BufReaderConfig::default();
assert_eq!(config.max_layers, 2);
assert_eq!(config.small_file_buffer, 8 * 1024);
assert_eq!(config.large_file_buffer, 64 * 1024);
assert_eq!(config.large_file_threshold, 1024 * 1024);
}
#[test]
fn test_optimal_buffer_size_small_file() {
let optimizer = BufReaderOptimizer::with_defaults();
// Small file (< 1MB)
assert_eq!(optimizer.optimal_buffer_size(Some(100)), 8 * 1024);
assert_eq!(optimizer.optimal_buffer_size(Some(1024)), 8 * 1024);
assert_eq!(optimizer.optimal_buffer_size(Some(512 * 1024)), 8 * 1024);
}
#[test]
fn test_optimal_buffer_size_large_file() {
let optimizer = BufReaderOptimizer::with_defaults();
// Large file (>= 1MB)
assert_eq!(optimizer.optimal_buffer_size(Some(1024 * 1024)), 64 * 1024);
assert_eq!(optimizer.optimal_buffer_size(Some(10 * 1024 * 1024)), 64 * 1024);
}
#[test]
fn test_optimal_buffer_size_unknown() {
let optimizer = BufReaderOptimizer::with_defaults();
// Unknown size
assert_eq!(optimizer.optimal_buffer_size(None), 8 * 1024);
}
#[tokio::test]
async fn test_optimize_creates_bufreader() {
let optimizer = BufReaderOptimizer::with_defaults();
let data = vec![1u8, 2, 3, 4, 5];
let cursor = std::io::Cursor::new(data.clone());
let mut reader = optimizer.optimize(cursor, Some(5));
let mut buf = vec![0u8; 5];
let n = reader.read(&mut buf).await.unwrap();
assert_eq!(n, 5);
assert_eq!(buf, data);
}
#[test]
fn test_stats_tracking() {
let optimizer = BufReaderOptimizer::with_defaults();
assert_eq!(optimizer.stats().total_readers.load(Ordering::Relaxed), 0);
let cursor = std::io::Cursor::new(vec![1u8, 2, 3]);
let _reader = optimizer.optimize(cursor, Some(3));
assert_eq!(optimizer.stats().total_readers.load(Ordering::Relaxed), 1);
}
#[test]
fn test_custom_config() {
let config = BufReaderConfig {
max_layers: 3,
small_file_buffer: 4 * 1024,
large_file_buffer: 128 * 1024,
large_file_threshold: 2 * 1024 * 1024,
};
let optimizer = BufReaderOptimizer::new(config);
assert_eq!(optimizer.optimal_buffer_size(Some(1024 * 1024)), 4 * 1024);
assert_eq!(optimizer.optimal_buffer_size(Some(3 * 1024 * 1024)), 128 * 1024);
}
}
-332
View File
@@ -1,332 +0,0 @@
// Copyright 2024 RustFS Team
//
// Licensed under the Apache License, Version 2.0 (the "License");
// you may not use this file except in compliance with the License.
// You may obtain a copy of the License at
//
// http://www.apache.org/licenses/LICENSE-2.0
//
// Unless required by applicable law or agreed to in writing, software
// distributed under the License is distributed on an "AS IS" BASIS,
// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
// See the License for the specific language governing permissions and
// limitations under the License.
//! Aligned pread-based file reader.
//!
//! This module provides an aligned, position-based file reader that uses
//! `pread`/`FileExt::read_at` for I/O operations. It performs reads at
//! 512-byte-aligned offsets and sizes, making it suitable as a foundation
//! for workloads where alignment matters.
//!
//! Note: This reader does **not** set the `O_DIRECT` flag and therefore does
//! not bypass the OS page cache. It is an aligned `pread`-based reader, not
//! true Direct I/O. To implement true O_DIRECT on Linux, the file must be
//! opened with `O_DIRECT` via `libc::open`.
//!
//! # Platform Support
//!
//! The `read_at` implementation is only available on Unix-like platforms.
//! On other platforms, this reader will return an error.
use std::io::{self};
use std::pin::Pin;
use std::task::{Context, Poll};
use tokio::io::{AsyncRead, ReadBuf};
/// Errors that can occur during aligned pread operations.
#[derive(Debug, Clone)]
pub enum AlignedPreadError {
/// Platform doesn't support `read_at`-based I/O
UnsupportedPlatform,
/// File descriptor doesn't support this reader
UnsupportedFile,
/// I/O error occurred
Io(String),
/// Invalid alignment (reads require 512-byte-aligned offset and size)
AlignmentError { offset: u64, size: usize },
}
impl std::fmt::Display for AlignedPreadError {
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
match self {
Self::UnsupportedPlatform => write!(f, "Aligned pread not supported on this platform"),
Self::UnsupportedFile => write!(f, "File doesn't support this reader"),
Self::Io(msg) => write!(f, "I/O error: {}", msg),
Self::AlignmentError { offset, size } => {
write!(f, "Alignment error: offset={}, size={}", offset, size)
}
}
}
}
impl std::error::Error for AlignedPreadError {}
impl From<io::Error> for AlignedPreadError {
fn from(err: io::Error) -> Self {
Self::Io(err.to_string())
}
}
/// Aligned pread-based file reader for Unix platforms.
///
/// This reader performs I/O using `pread`/`FileExt::read_at` at
/// 512-byte-aligned offsets and sizes, without modifying the file's
/// current position.
///
/// **Note:** This reader does **not** set the `O_DIRECT` flag and therefore
/// does **not** bypass the OS page cache. It is an aligned `pread`-based
/// reader. To implement true O_DIRECT, the file must be opened with
/// `O_DIRECT` via `libc::open`.
///
/// # Platform Support
///
/// Only available on Linux (uses `FileExt::read_at`). On other platforms,
/// use `BytesBufferedReader` instead.
///
/// # Alignment Requirements
///
/// Reads have strict alignment requirements:
/// - File offset must be aligned to 512 bytes
/// - Buffer size must be a multiple of 512 bytes
/// - Buffer address must be aligned (handled internally)
///
/// # Example
///
/// ```ignore
/// use rustfs_io_core::AlignedPreadReader;
///
/// // Linux only
/// #[cfg(target_os = "linux")]
/// let reader = AlignedPreadReader::new(file, offset, size)?;
/// ```
#[cfg(target_os = "linux")]
pub struct AlignedPreadReader {
/// Underlying file handle used for aligned pread I/O
file: std::fs::File,
/// Current read position
pos: u64,
/// Remaining bytes to read
remaining: usize,
/// Buffer for aligned reads
buffer: Vec<u8>,
/// Current position in the buffer
buffer_pos: usize,
/// Amount of data in the buffer
buffer_len: usize,
}
#[cfg(target_os = "linux")]
impl AlignedPreadReader {
/// Alignment requirement for reads (512 bytes for most systems)
pub const ALIGNMENT: usize = 512;
/// Create a new aligned pread-based reader.
///
/// # Arguments
///
/// * `file` - File to read from
/// * `offset` - Starting offset in the file (must be 512-byte aligned)
/// * `size` - Number of bytes to read (must be 512-byte aligned)
///
/// # Returns
///
/// An `AlignedPreadReader` that reads the file at the given offset.
///
/// # Errors
///
/// Returns an error if offset or size are not 512-byte aligned.
pub fn new(file: std::fs::File, offset: u64, size: usize) -> Result<Self, AlignedPreadError> {
// Check alignment
if !offset.is_multiple_of(Self::ALIGNMENT as u64) {
return Err(AlignedPreadError::AlignmentError { offset, size });
}
if !size.is_multiple_of(Self::ALIGNMENT) {
return Err(AlignedPreadError::AlignmentError { offset, size });
}
Ok(Self {
file,
pos: offset,
remaining: size,
buffer: Vec::new(),
buffer_pos: 0,
buffer_len: 0,
})
}
/// Read a chunk of data using aligned pread.
///
/// This method performs aligned reads and handles the buffering required
/// by this aligned pread implementation. It does not use `O_DIRECT`.
fn read_chunk(&mut self, buf: &mut [u8]) -> io::Result<usize> {
// If buffer is exhausted, read more data
if self.buffer_pos >= self.buffer_len {
if self.remaining == 0 {
return Ok(0);
}
// Allocate aligned buffer
let chunk_size = (self.remaining).min(64 * 1024); // 64KB chunks
let aligned_size = chunk_size.div_ceil(Self::ALIGNMENT) * Self::ALIGNMENT;
self.buffer = vec![0u8; aligned_size];
// Use pread for atomic read at position (no file offset modification)
use std::os::unix::fs::FileExt;
let n = self.file.read_at(&mut self.buffer, self.pos)?;
self.buffer_pos = 0;
self.buffer_len = n;
self.pos += n as u64;
self.remaining -= n;
if n == 0 {
return Ok(0);
}
}
// Copy from buffer to user buffer
let available = self.buffer_len - self.buffer_pos;
let to_copy = buf.len().min(available);
buf[..to_copy].copy_from_slice(&self.buffer[self.buffer_pos..self.buffer_pos + to_copy]);
self.buffer_pos += to_copy;
Ok(to_copy)
}
}
#[cfg(target_os = "linux")]
impl AsyncRead for AlignedPreadReader {
fn poll_read(mut self: Pin<&mut Self>, _cx: &mut Context<'_>, buf: &mut ReadBuf<'_>) -> Poll<io::Result<()>> {
let filled = buf.filled().len();
let mut remaining = buf.initialize_unfilled();
while !remaining.is_empty() {
match self.read_chunk(remaining) {
Ok(0) => break,
Ok(n) => {
remaining = &mut remaining[n..];
}
Err(e) => return Poll::Ready(Err(e)),
}
}
let _n_read = buf.filled().len() - filled;
Poll::Ready(Ok(()))
}
}
/// Aligned pread reader stub for non-Linux platforms.
///
/// On non-Linux platforms, `read_at`-based I/O is not available through this
/// type. This stub exists to provide a consistent API across platforms.
#[cfg(not(target_os = "linux"))]
pub struct AlignedPreadReader {
_priv: (),
}
#[cfg(not(target_os = "linux"))]
impl AlignedPreadReader {
/// Create a new aligned pread reader (not supported on this platform).
///
/// Always returns an error on non-Linux platforms.
pub fn new(_file: std::fs::File, _offset: u64, _size: usize) -> Result<Self, AlignedPreadError> {
Err(AlignedPreadError::UnsupportedPlatform)
}
}
#[cfg(not(target_os = "linux"))]
impl AsyncRead for AlignedPreadReader {
fn poll_read(self: Pin<&mut Self>, _cx: &mut Context<'_>, _buf: &mut ReadBuf<'_>) -> Poll<io::Result<()>> {
Poll::Ready(Err(io::Error::new(
io::ErrorKind::Unsupported,
"Aligned pread-based I/O not supported on this platform",
)))
}
}
impl std::fmt::Debug for AlignedPreadReader {
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
#[cfg(target_os = "linux")]
{
f.debug_struct("AlignedPreadReader")
.field("pos", &self.pos)
.field("remaining", &self.remaining)
.field("buffer_len", &self.buffer_len)
.finish()
}
#[cfg(not(target_os = "linux"))]
{
f.debug_struct("AlignedPreadReader")
.field("platform", &"unsupported")
.finish()
}
}
}
/// Historical name for aligned pread errors.
#[deprecated(since = "1.0.0-beta.8", note = "use AlignedPreadError; this reader does not set O_DIRECT")]
pub type DirectIoError = AlignedPreadError;
/// Historical name for the aligned pread-based reader.
#[deprecated(since = "1.0.0-beta.8", note = "use AlignedPreadReader; this reader does not set O_DIRECT")]
pub type DirectIoReader = AlignedPreadReader;
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn test_alignment_check() {
#[cfg(target_os = "linux")]
{
// Valid alignment
let file = std::fs::File::open("/dev/zero").unwrap();
assert!(
AlignedPreadReader::new(file, 0, 512).is_ok(),
"Should succeed with aligned offset and size"
);
let file = std::fs::File::open("/dev/zero").expect("open /dev/zero for alias");
assert!(
AlignedPreadReader::new(file, 0, 512).is_ok(),
"Should succeed through aligned pread alias"
);
// Invalid offset
let file = std::fs::File::open("/dev/zero").unwrap();
assert!(AlignedPreadReader::new(file, 1, 512).is_err(), "Should fail with unaligned offset");
// Invalid size
let file = std::fs::File::open("/dev/zero").unwrap();
assert!(AlignedPreadReader::new(file, 0, 511).is_err(), "Should fail with unaligned size");
}
#[cfg(not(target_os = "linux"))]
{
// Non-Linux should return UnsupportedPlatform
let file = std::fs::File::open(std::env::current_exe().unwrap()).unwrap();
assert!(matches!(
AlignedPreadReader::new(file, 0, 512),
Err(AlignedPreadError::UnsupportedPlatform)
));
}
}
#[test]
#[allow(deprecated)]
fn test_legacy_direct_io_alias() {
#[cfg(target_os = "linux")]
{
let file = std::fs::File::open("/dev/zero").unwrap();
assert!(DirectIoReader::new(file, 0, 512).is_ok());
}
#[cfg(not(target_os = "linux"))]
{
let file = std::fs::File::open(std::env::current_exe().unwrap()).unwrap();
assert!(matches!(DirectIoReader::new(file, 0, 512), Err(AlignedPreadError::UnsupportedPlatform)));
}
}
}
-381
View File
@@ -1,381 +0,0 @@
// Copyright 2024 RustFS Team
//
// Licensed under the Apache License, Version 2.0 (the "License");
// you may not use this file except in compliance with the License.
// You may obtain a copy of the License at
//
// http://www.apache.org/licenses/LICENSE-2.0
//
// Unless required by applicable law or agreed to in writing, software
// distributed under the License is distributed on an "AS IS" BASIS,
// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
// See the License for the specific language governing permissions and
// limitations under the License.
//! I/O priority queue for scheduling I/O operations.
//!
//! This module provides a priority queue implementation for I/O operations
//! with support for starvation prevention and fair scheduling.
use crate::config::IoPriorityQueueConfig;
use crate::scheduler::IoPriority;
use std::collections::VecDeque;
use std::time::{Duration, Instant};
/// A queued I/O request.
#[derive(Debug, Clone)]
pub struct IoRequest {
/// Request ID.
pub id: u64,
/// Request priority.
pub priority: IoPriority,
/// Request size in bytes.
pub size: usize,
/// Queue time.
pub queued_at: Instant,
/// Whether this is a sequential read.
pub is_sequential: bool,
}
impl IoRequest {
/// Create a new I/O request.
pub fn new(id: u64, priority: IoPriority, size: usize, is_sequential: bool) -> Self {
Self {
id,
priority,
size,
queued_at: Instant::now(),
is_sequential,
}
}
/// Get the wait time in the queue.
pub fn wait_time(&self) -> Duration {
self.queued_at.elapsed()
}
}
/// Queue status for a priority level.
#[derive(Debug, Clone, Default)]
pub struct IoQueueStatus {
/// Number of requests in the queue.
pub count: usize,
/// Total size of all requests.
pub total_size: usize,
/// Oldest request wait time.
pub oldest_wait: Option<Duration>,
/// Number of requests processed.
pub processed: u64,
}
impl IoQueueStatus {
/// Create new queue status.
pub fn new() -> Self {
Self::default()
}
}
/// I/O priority queue.
pub struct IoPriorityQueue {
/// Queue configuration.
config: IoPriorityQueueConfig,
/// High priority queue.
high: VecDeque<IoRequest>,
/// Normal priority queue.
normal: VecDeque<IoRequest>,
/// Low priority queue.
low: VecDeque<IoRequest>,
/// Next request ID.
next_id: u64,
/// Last dequeue time for each priority (for starvation prevention).
last_dequeue: [Option<Instant>; 3],
/// Statistics for each queue.
stats: [IoQueueStatus; 3],
}
impl IoPriorityQueue {
/// Create a new priority queue with the given configuration.
pub fn new(config: IoPriorityQueueConfig) -> Self {
Self {
config,
high: VecDeque::with_capacity(100),
normal: VecDeque::with_capacity(500),
low: VecDeque::with_capacity(200),
next_id: 0,
last_dequeue: [None, None, None],
stats: [IoQueueStatus::new(), IoQueueStatus::new(), IoQueueStatus::new()],
}
}
/// Create with default configuration.
pub fn with_defaults() -> Self {
Self::new(IoPriorityQueueConfig::default())
}
/// Get the configuration.
pub fn config(&self) -> &IoPriorityQueueConfig {
&self.config
}
/// Enqueue a request.
pub fn enqueue(&mut self, priority: IoPriority, size: usize, is_sequential: bool) -> u64 {
let id = self.next_id;
self.next_id += 1;
let request = IoRequest::new(id, priority, size, is_sequential);
match priority {
IoPriority::High => {
if self.high.len() < self.config.high_capacity {
self.high.push_back(request);
}
}
IoPriority::Normal => {
if self.normal.len() < self.config.normal_capacity {
self.normal.push_back(request);
}
}
IoPriority::Low => {
if self.low.len() < self.config.low_capacity {
self.low.push_back(request);
}
}
}
id
}
/// Dequeue the next request.
///
/// Uses weighted fair queuing with starvation prevention.
pub fn dequeue(&mut self) -> Option<IoRequest> {
let now = Instant::now();
// Check for starvation: if a lower priority queue hasn't been served in a while,
// give it priority
let normal_starved = self.is_starved(IoPriority::Normal, now);
let low_starved = self.is_starved(IoPriority::Low, now);
// Priority order with starvation consideration
// Check conditions first, then dequeue
let dequeue_high = !self.high.is_empty() && !low_starved && !normal_starved;
let dequeue_normal = !self.normal.is_empty() && !low_starved;
let dequeue_low = !self.low.is_empty();
let dequeue_high_fallback = !self.high.is_empty();
let dequeue_normal_fallback = !self.normal.is_empty();
if dequeue_high {
let request = self.high.pop_front();
if request.is_some() {
self.last_dequeue[0] = Some(Instant::now());
self.stats[0].processed += 1;
}
request
} else if dequeue_normal {
let request = self.normal.pop_front();
if request.is_some() {
self.last_dequeue[1] = Some(Instant::now());
self.stats[1].processed += 1;
}
request
} else if dequeue_low {
let request = self.low.pop_front();
if request.is_some() {
self.last_dequeue[2] = Some(Instant::now());
self.stats[2].processed += 1;
}
request
} else if dequeue_high_fallback {
let request = self.high.pop_front();
if request.is_some() {
self.last_dequeue[0] = Some(Instant::now());
self.stats[0].processed += 1;
}
request
} else if dequeue_normal_fallback {
let request = self.normal.pop_front();
if request.is_some() {
self.last_dequeue[1] = Some(Instant::now());
self.stats[1].processed += 1;
}
request
} else {
None
}
}
/// Check if a priority level is starved.
fn is_starved(&self, priority: IoPriority, now: Instant) -> bool {
let idx = match priority {
IoPriority::High => 0,
IoPriority::Normal => 1,
IoPriority::Low => 2,
};
if let Some(last) = self.last_dequeue[idx] {
now.duration_since(last) > self.config.starvation_threshold
} else {
false
}
}
/// Get the total number of queued requests.
pub fn len(&self) -> usize {
self.high.len() + self.normal.len() + self.low.len()
}
/// Check if the queue is empty.
pub fn is_empty(&self) -> bool {
self.high.is_empty() && self.normal.is_empty() && self.low.is_empty()
}
/// Get queue status for a priority level.
pub fn status(&self, priority: IoPriority) -> IoQueueStatus {
let (queue, idx) = match priority {
IoPriority::High => (&self.high, 0),
IoPriority::Normal => (&self.normal, 1),
IoPriority::Low => (&self.low, 2),
};
let mut status = self.stats[idx].clone();
status.count = queue.len();
status.total_size = queue.iter().map(|r| r.size).sum();
status.oldest_wait = queue.front().map(|r| r.wait_time());
status
}
/// Get the total queue status.
pub fn total_status(&self) -> IoQueueStatus {
let mut total = IoQueueStatus::new();
total.count = self.len();
total.total_size = self
.high
.iter()
.chain(self.normal.iter())
.chain(self.low.iter())
.map(|r| r.size)
.sum();
total.processed = self.stats.iter().map(|s| s.processed).sum();
total.oldest_wait = self
.high
.front()
.map(|r| r.wait_time())
.or_else(|| self.normal.front().map(|r| r.wait_time()))
.or_else(|| self.low.front().map(|r| r.wait_time()));
total
}
/// Clear all queues.
pub fn clear(&mut self) {
self.high.clear();
self.normal.clear();
self.low.clear();
}
/// Peek at the next request without removing it.
pub fn peek(&self) -> Option<&IoRequest> {
if !self.high.is_empty() {
self.high.front()
} else if !self.normal.is_empty() {
self.normal.front()
} else {
self.low.front()
}
}
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn test_enqueue_dequeue() {
let mut queue = IoPriorityQueue::with_defaults();
let id1 = queue.enqueue(IoPriority::High, 1024, true);
let id2 = queue.enqueue(IoPriority::Normal, 2048, false);
let id3 = queue.enqueue(IoPriority::Low, 4096, true);
assert_eq!(queue.len(), 3);
// High priority should be dequeued first
let req1 = queue.dequeue().unwrap();
assert_eq!(req1.id, id1);
assert_eq!(req1.priority, IoPriority::High);
let req2 = queue.dequeue().unwrap();
assert_eq!(req2.id, id2);
assert_eq!(req2.priority, IoPriority::Normal);
let req3 = queue.dequeue().unwrap();
assert_eq!(req3.id, id3);
assert_eq!(req3.priority, IoPriority::Low);
assert!(queue.is_empty());
}
#[test]
fn test_queue_status() {
let mut queue = IoPriorityQueue::with_defaults();
queue.enqueue(IoPriority::High, 1024, true);
queue.enqueue(IoPriority::High, 2048, true);
queue.enqueue(IoPriority::Normal, 4096, false);
let high_status = queue.status(IoPriority::High);
assert_eq!(high_status.count, 2);
assert_eq!(high_status.total_size, 3072);
let normal_status = queue.status(IoPriority::Normal);
assert_eq!(normal_status.count, 1);
assert_eq!(normal_status.total_size, 4096);
let total = queue.total_status();
assert_eq!(total.count, 3);
assert_eq!(total.total_size, 7168);
}
#[test]
fn test_queue_capacity() {
let config = IoPriorityQueueConfig {
high_capacity: 2,
normal_capacity: 2,
low_capacity: 2,
..Default::default()
};
let mut queue = IoPriorityQueue::new(config);
queue.enqueue(IoPriority::High, 1024, true);
queue.enqueue(IoPriority::High, 1024, true);
queue.enqueue(IoPriority::High, 1024, true); // Should be dropped
assert_eq!(queue.status(IoPriority::High).count, 2);
}
#[test]
fn test_clear() {
let mut queue = IoPriorityQueue::with_defaults();
queue.enqueue(IoPriority::High, 1024, true);
queue.enqueue(IoPriority::Normal, 2048, false);
queue.enqueue(IoPriority::Low, 4096, true);
assert_eq!(queue.len(), 3);
queue.clear();
assert!(queue.is_empty());
}
#[test]
fn test_peek() {
let mut queue = IoPriorityQueue::with_defaults();
queue.enqueue(IoPriority::Normal, 2048, false);
queue.enqueue(IoPriority::High, 1024, true);
let peeked = queue.peek().unwrap();
assert_eq!(peeked.priority, IoPriority::High);
// Peek shouldn't remove the item
assert_eq!(queue.len(), 2);
}
}
-6
View File
@@ -26,7 +26,6 @@ pub enum StorageMedia {
} }
impl StorageMedia { impl StorageMedia {
#[allow(dead_code)]
pub fn as_str(&self) -> &'static str { pub fn as_str(&self) -> &'static str {
match self { match self {
Self::Nvme => "nvme", Self::Nvme => "nvme",
@@ -60,7 +59,6 @@ pub enum AccessPattern {
} }
impl AccessPattern { impl AccessPattern {
#[allow(dead_code)]
pub fn as_str(&self) -> &'static str { pub fn as_str(&self) -> &'static str {
match self { match self {
Self::Sequential => "sequential", Self::Sequential => "sequential",
@@ -71,25 +69,21 @@ impl AccessPattern {
} }
/// Check if this is a sequential access pattern. /// Check if this is a sequential access pattern.
#[allow(dead_code)]
pub fn is_sequential(&self) -> bool { pub fn is_sequential(&self) -> bool {
matches!(self, Self::Sequential) matches!(self, Self::Sequential)
} }
/// Check if this is a random access pattern. /// Check if this is a random access pattern.
#[allow(dead_code)]
pub fn is_random(&self) -> bool { pub fn is_random(&self) -> bool {
matches!(self, Self::Random) matches!(self, Self::Random)
} }
/// Check if this is a mixed access pattern. /// Check if this is a mixed access pattern.
#[allow(dead_code)]
pub fn is_mixed(&self) -> bool { pub fn is_mixed(&self) -> bool {
matches!(self, Self::Mixed) matches!(self, Self::Mixed)
} }
/// Check if this pattern is unknown. /// Check if this pattern is unknown.
#[allow(dead_code)]
pub fn is_unknown(&self) -> bool { pub fn is_unknown(&self) -> bool {
matches!(self, Self::Unknown) matches!(self, Self::Unknown)
} }
+12 -61
View File
@@ -12,85 +12,39 @@
// See the License for the specific language governing permissions and // See the License for the specific language governing permissions and
// limitations under the License. // limitations under the License.
//! Buffered I/O reader and writer implementations for RustFS. //! Shared I/O primitives for RustFS.
//! //!
//! This crate provides buffered readers and writers for I/O operations. //! This crate holds the buffer pool and the concurrency-control primitives
//! Prefer `BytesBufferedReader`, `BytesMutWriter`, and `AlignedPreadReader` //! that the storage layer builds on:
//! for new code. Historical `ZeroCopy*` and `DirectIo*` names remain exported
//! for backward compatibility.
//! //!
//! # Features //! - Tiered `BytesPool` for buffer management
//! //! - Storage-media and access-pattern profiling (`io_profile`)
//! - Memory-mapped file reading (mmap-then-copy) on Unix platforms //! - Scheduler and priority-queue configuration shapes
//! - Bytes-based buffered wrapping //! - Backpressure admission, deadlock detection, lock optimization
//! - AsyncRead trait implementations //! - Progress tracking for long-running operations
//! - Tiered BytesPool for buffer management
//! - Aligned pread-based reader (NOT true Direct I/O / O_DIRECT)
//! //!
//! # Example //! # Example
//! //!
//! ```ignore //! ```ignore
//! use rustfs_io_core::{BytesBufferedReader, BytesPool}; //! use rustfs_io_core::BytesPool;
//! use bytes::Bytes;
//! //!
//! // Create from existing bytes (zero-copy)
//! let data = Bytes::from("hello world");
//! let reader = BytesBufferedReader::from_bytes(data);
//!
//! // Create from file using buffered reads
//! let reader = BytesBufferedReader::from_file_read(&file, 0, 1024).await?;
//!
//! // Use BytesPool
//! let pool = BytesPool::new_tiered(); //! let pool = BytesPool::new_tiered();
//! let mut buffer = pool.acquire_buffer(8192).await; //! let mut buffer = pool.acquire_buffer(8192).await;
//! ``` //! ```
pub mod backpressure; pub mod backpressure;
pub mod bufreader_optimizer;
pub mod config; pub mod config;
pub mod deadlock_detector; pub mod deadlock_detector;
pub mod direct_io;
pub mod io_priority_queue;
pub mod io_profile; pub mod io_profile;
pub mod lock_optimizer; pub mod lock_optimizer;
pub mod pool; pub mod pool;
pub mod reader; pub mod progress;
pub mod scheduler;
pub mod shared_memory;
pub mod timeout_wrapper;
pub mod writer;
#[cfg(target_os = "linux")]
pub use direct_io::{AlignedPreadError, AlignedPreadReader};
#[cfg(target_os = "linux")]
#[allow(deprecated)]
pub use direct_io::{DirectIoError, DirectIoReader};
pub use pool::{BytesPool, BytesPoolConfig, BytesPoolMetrics, PooledBuffer}; pub use pool::{BytesPool, BytesPoolConfig, BytesPoolMetrics, PooledBuffer};
#[allow(deprecated)]
pub use reader::ZeroCopyObjectReader;
pub use reader::{BytesBufferedReader, ZeroCopyReadError};
#[allow(deprecated)]
pub use writer::ZeroCopyObjectWriter;
pub use writer::{BytesMutWriter, ZeroCopyWriteError};
// BufReader optimizer exports
pub use bufreader_optimizer::{BufReaderConfig, BufReaderOptimizer, BufReaderStats, BufferedSource};
// Shared memory exports
pub use shared_memory::{ArcData, ArcMetadata, SharedMemoryConfig, SharedMemoryPool, SharedMemoryStats};
// Config exports // Config exports
pub use config::{ConfigError, IoPriorityQueueConfig, IoSchedulerConfig}; pub use config::{ConfigError, IoPriorityQueueConfig, IoSchedulerConfig};
// Scheduler exports
pub use scheduler::{
BandwidthTier, IoLoadLevel, IoLoadMetrics, IoPriority, IoScheduler, IoSchedulingContext, IoStrategy, KI_B, MI_B,
calculate_optimal_buffer_size, get_advanced_buffer_size, get_buffer_size_for_media, get_concurrency_aware_buffer_size,
};
// Priority queue exports
pub use io_priority_queue::{IoPriorityQueue, IoQueueStatus, IoRequest};
// Backpressure exports // Backpressure exports
pub use backpressure::{BackpressureConfig, BackpressureError, BackpressureMonitor, BackpressureState}; pub use backpressure::{BackpressureConfig, BackpressureError, BackpressureMonitor, BackpressureState};
@@ -100,8 +54,5 @@ pub use deadlock_detector::{DeadlockDetector, DeadlockDetectorConfig, LockInfo,
// Lock optimizer exports // Lock optimizer exports
pub use lock_optimizer::{LockGuard, LockOptimizeConfig, LockOptimizer, LockStats}; pub use lock_optimizer::{LockGuard, LockOptimizeConfig, LockOptimizer, LockStats};
// Timeout wrapper exports // Progress tracking exports
pub use timeout_wrapper::{ pub use progress::OperationProgress;
OperationProgress, RequestTimeoutWrapper, TimeoutConfig, TimeoutError, TimeoutStats, calculate_adaptive_timeout,
estimate_bytes_per_second,
};
+138
View File
@@ -0,0 +1,138 @@
// Copyright 2024 RustFS Team
//
// Licensed under the Apache License, Version 2.0 (the "License");
// you may not use this file except in compliance with the License.
// You may obtain a copy of the License at
//
// http://www.apache.org/licenses/LICENSE-2.0
//
// Unless required by applicable law or agreed to in writing, software
// distributed under the License is distributed on an "AS IS" BASIS,
// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
// See the License for the specific language governing permissions and
// limitations under the License.
//! Progress tracking for long-running I/O operations.
//!
//! Re-exported as `rustfs_concurrency::OperationProgress` for the storage
//! timeout implementation, which uses `is_stale` to tell a slow transfer
//! apart from a stalled one.
use std::sync::atomic::{AtomicU64, Ordering};
use std::time::{Duration, Instant};
/// Operation progress tracker.
#[derive(Debug)]
pub struct OperationProgress {
/// Total size (if known).
pub total_size: Option<u64>,
/// Bytes processed.
bytes_processed: AtomicU64,
/// Last update time.
last_update: std::sync::Mutex<Instant>,
/// Stale timeout.
stale_timeout: Duration,
/// Start time for transfer rate calculation.
start_time: Instant,
}
impl OperationProgress {
/// Create new operation progress.
pub fn new(total_size: Option<u64>, stale_timeout: Duration) -> Self {
Self {
total_size,
bytes_processed: AtomicU64::new(0),
last_update: std::sync::Mutex::new(Instant::now()),
stale_timeout,
start_time: Instant::now(),
}
}
/// Update progress.
pub fn update(&self, bytes: u64) {
self.bytes_processed.store(bytes, Ordering::Relaxed);
if let Ok(mut last) = self.last_update.lock() {
*last = Instant::now();
}
}
/// Add to progress.
pub fn add(&self, bytes: u64) {
self.bytes_processed.fetch_add(bytes, Ordering::Relaxed);
if let Ok(mut last) = self.last_update.lock() {
*last = Instant::now();
}
}
/// Get current progress.
pub fn current(&self) -> u64 {
self.bytes_processed.load(Ordering::Relaxed)
}
/// Check if progress is stale.
pub fn is_stale(&self) -> bool {
if let Ok(last) = self.last_update.lock() {
last.elapsed() > self.stale_timeout
} else {
false
}
}
/// Get progress percentage.
pub fn progress_percent(&self) -> Option<f64> {
self.total_size.map(|total| {
if total == 0 {
100.0
} else {
let processed = self.bytes_processed.load(Ordering::Relaxed);
(processed as f64 / total as f64 * 100.0).min(100.0)
}
})
}
/// Get remaining bytes.
pub fn remaining(&self) -> Option<u64> {
self.total_size.map(|total| {
let processed = self.bytes_processed.load(Ordering::Relaxed);
total.saturating_sub(processed)
})
}
/// Calculate transfer rate in bytes per second.
///
/// Returns 0 if no time has elapsed or no data transferred.
pub fn transfer_rate(&self) -> u64 {
let processed = self.bytes_processed.load(Ordering::Relaxed);
if processed == 0 {
return 0;
}
let elapsed = self.start_time.elapsed().as_secs_f64();
if elapsed > 0.0 {
(processed as f64 / elapsed) as u64
} else {
0
}
}
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn test_operation_progress() {
let progress = OperationProgress::new(Some(1000), Duration::from_secs(5));
assert_eq!(progress.current(), 0);
assert_eq!(progress.progress_percent(), Some(0.0));
progress.update(500);
assert_eq!(progress.current(), 500);
assert_eq!(progress.progress_percent(), Some(50.0));
progress.add(300);
assert_eq!(progress.current(), 800);
assert_eq!(progress.remaining(), Some(200));
}
}
-412
View File
@@ -1,412 +0,0 @@
// Copyright 2024 RustFS Team
//
// Licensed under the Apache License, Version 2.0 (the "License");
// you may not use this file except in compliance with the License.
// You may obtain a copy of the License at
//
// http://www.apache.org/licenses/LICENSE-2.0
//
// Unless required by applicable law or agreed to in writing, software
// distributed under the License is distributed on an "AS IS" BASIS,
// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
// See the License for the specific language governing permissions and
// limitations under the License.
//! Bytes-backed object reader implementation.
use bytes::Bytes;
use std::io;
use std::pin::Pin;
use std::task::{Context, Poll};
use tokio::io::{AsyncRead, ReadBuf};
/// Errors that can occur during Bytes-backed read operations.
#[derive(Debug, Clone)]
pub enum ZeroCopyReadError {
/// I/O error occurred.
Io(String),
/// Memory mapping error.
Mmap(String),
/// Invalid offset or size.
InvalidRange,
}
impl std::fmt::Display for ZeroCopyReadError {
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
match self {
Self::Io(msg) => write!(f, "I/O error: {}", msg),
Self::Mmap(msg) => write!(f, "Mmap error: {}", msg),
Self::InvalidRange => write!(f, "Invalid offset or size"),
}
}
}
impl std::error::Error for ZeroCopyReadError {}
impl From<io::Error> for ZeroCopyReadError {
fn from(err: io::Error) -> Self {
Self::Io(err.to_string())
}
}
/// Bytes-backed object reader.
///
/// `from_bytes` wraps existing `Bytes` without copying, but file constructors
/// copy file data into owned `Bytes` after mmap or normal reads.
///
/// # Example
///
/// ```ignore
/// use bytes::Bytes;
/// use rustfs_io_core::BytesBufferedReader;
///
/// // Create from bytes without copying the `Bytes` buffer
/// let data = Bytes::from("hello world");
/// let reader = BytesBufferedReader::from_bytes(data);
///
/// // Read using AsyncRead trait
/// let mut buf = vec![0u8; 1024];
/// let n = reader.read(&mut buf[..]).await?;
/// ```
pub struct BytesBufferedReader {
/// Internal data source (could be mmap or owned bytes)
data: Bytes,
/// Current read position
pos: usize,
}
/// Historical name for the bytes-backed object reader.
#[deprecated(
since = "1.0.0-beta.8",
note = "use BytesBufferedReader; file constructors copy into owned Bytes"
)]
pub type ZeroCopyObjectReader = BytesBufferedReader;
impl BytesBufferedReader {
/// Create a reader from existing bytes.
///
/// This is a true zero-copy operation - the Bytes are wrapped
/// without any allocation or copying.
///
/// # Arguments
///
/// * `data` - Bytes to wrap
///
/// # Example
///
/// ```ignore
/// let data = Bytes::from("hello world");
/// let reader = BytesBufferedReader::from_bytes(data);
/// ```
pub fn from_bytes(data: Bytes) -> Self {
Self { data, pos: 0 }
}
/// Create a Bytes-backed reader from a file using mmap-then-copy.
///
/// This maps the requested file range and copies it into owned `Bytes`
/// before returning. It does not expose the mmap as a zero-copy buffer.
///
/// # Arguments
///
/// * `path` - Path to the file to memory map
/// * `offset` - Offset within the file to start reading
/// * `size` - Number of bytes to read
///
/// # Returns
///
/// A reader backed by copied file data.
///
/// # Errors
///
/// Returns an error if the file cannot be memory mapped.
///
/// # Example
///
/// ```ignore
/// let reader = BytesBufferedReader::from_file_mmap_path("large_file.bin", 0, 1024).await?;
/// ```
#[cfg(unix)]
// SAFETY: The mmap is created from a read-only file handle for the
// caller-provided range, then copied into owned `Bytes` before the file and
// mapping are dropped.
#[allow(unsafe_code)]
pub async fn from_file_mmap_path(path: &std::path::Path, offset: u64, size: usize) -> Result<Self, ZeroCopyReadError> {
use memmap2::MmapOptions;
let path = path.to_path_buf();
let (offset, size) = (offset, size);
tokio::task::spawn_blocking(move || {
// Open the file in sync context
let std_file = std::fs::File::open(&path).map_err(|e| ZeroCopyReadError::Io(e.to_string()))?;
// SAFETY: `std_file` remains open while the mapping is created and
// copied, and the mapped bytes are not exposed beyond this closure.
let mmap = unsafe { MmapOptions::new().offset(offset).len(size).map(&std_file) }
.map_err(|e| ZeroCopyReadError::Mmap(e.to_string()))?;
// Convert to Bytes (this is a copy, but only done once)
Ok(Self {
data: Bytes::copy_from_slice(&mmap),
pos: 0,
})
})
.await
.map_err(|e| ZeroCopyReadError::Io(e.to_string()))?
}
/// Create a Bytes-backed reader from a file using normal reads.
///
/// This path reads the requested range into an owned buffer and wraps it in
/// `Bytes`. It does not perform mmap or zero-copy file I/O.
///
/// # Arguments
///
/// * `file` - File to read from
/// * `offset` - Offset within the file to start reading
/// * `size` - Number of bytes to map
///
/// # Returns
///
/// A reader backed by copied file data.
///
/// # Errors
///
/// Returns an error if the file cannot be read.
///
/// # Example
///
/// ```ignore
/// let file = tokio::fs::File::open("large_file.bin").await?;
/// let reader = BytesBufferedReader::from_file_read(&file, 0, 1024).await?;
/// ```
#[cfg(unix)]
pub async fn from_file_read(file: &tokio::fs::File, offset: u64, size: usize) -> Result<Self, ZeroCopyReadError> {
use tokio::io::{AsyncReadExt, AsyncSeekExt, SeekFrom};
let mut cloned = file.try_clone().await?;
cloned.seek(SeekFrom::Start(offset)).await?;
let mut buffer = vec![0u8; size];
cloned.read_exact(&mut buffer).await?;
Ok(Self {
data: Bytes::from(buffer),
pos: 0,
})
}
/// Create a Bytes-backed reader from a file (non-Unix fallback).
///
/// On platforms that don't support mmap, this falls back to regular file I/O.
#[cfg(not(unix))]
pub async fn from_file_read(file: &tokio::fs::File, offset: u64, size: usize) -> Result<Self, ZeroCopyReadError> {
use tokio::io::{AsyncReadExt, AsyncSeekExt, SeekFrom};
let mut cloned = file.try_clone().await?;
cloned.seek(SeekFrom::Start(offset)).await?;
let mut buffer = vec![0u8; size];
cloned.read_exact(&mut buffer).await?;
Ok(Self {
data: Bytes::from(buffer),
pos: 0,
})
}
/// Historical name for `from_file_read`.
#[deprecated(
since = "1.0.0-beta.8",
note = "use from_file_read; this method performs normal reads into owned Bytes"
)]
pub async fn from_file_mmap(file: &tokio::fs::File, offset: u64, size: usize) -> Result<Self, ZeroCopyReadError> {
Self::from_file_read(file, offset, size).await
}
/// Get the remaining data as Bytes (zero-copy).
///
/// This returns a slice of the remaining data without copying.
/// The returned Bytes shares the underlying memory with this reader.
///
/// # Example
///
/// ```ignore
/// let remaining = reader.remaining_bytes();
/// println!("Remaining: {} bytes", remaining.len());
/// ```
pub fn remaining_bytes(&self) -> Bytes {
self.data.slice(self.pos..)
}
/// Get the total length of the data.
pub fn len(&self) -> usize {
self.data.len()
}
/// Check if the reader has reached the end.
pub fn is_empty(&self) -> bool {
self.pos >= self.data.len()
}
/// Get the current read position.
pub fn position(&self) -> usize {
self.pos
}
}
impl AsyncRead for BytesBufferedReader {
fn poll_read(mut self: Pin<&mut Self>, _cx: &mut Context<'_>, buf: &mut ReadBuf<'_>) -> Poll<io::Result<()>> {
let remaining = self.data.len() - self.pos;
if remaining == 0 {
return Poll::Ready(Ok(()));
}
let to_read = std::cmp::min(remaining, buf.remaining());
let slice = &self.data[self.pos..self.pos + to_read];
buf.put_slice(slice);
self.pos += to_read;
Poll::Ready(Ok(()))
}
}
impl std::fmt::Debug for BytesBufferedReader {
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
f.debug_struct("BytesBufferedReader")
.field("data_len", &self.data.len())
.field("pos", &self.pos)
.field("remaining", &(self.data.len() - self.pos))
.finish()
}
}
#[cfg(test)]
mod tests {
use super::*;
use std::path::PathBuf;
use tokio::io::AsyncReadExt;
fn temp_file_path(test_name: &str) -> PathBuf {
let nonce = std::time::SystemTime::now()
.duration_since(std::time::UNIX_EPOCH)
.expect("system time should be after unix epoch")
.as_nanos();
std::env::temp_dir().join(format!("rustfs-io-core-{test_name}-{}-{nonce}", std::process::id()))
}
#[tokio::test]
async fn test_from_bytes() {
let data = Bytes::from("hello world");
let mut reader = BytesBufferedReader::from_bytes(data.clone());
let mut buf = [0u8; 11];
let n = reader.read(&mut buf[..]).await.unwrap();
assert_eq!(n, 11);
assert_eq!(&buf[..n], b"hello world");
}
#[tokio::test]
async fn test_preferred_reader_alias() {
let data = Bytes::from("hello world");
let mut reader = BytesBufferedReader::from_bytes(data);
let mut buf = [0u8; 5];
let n = reader.read(&mut buf[..]).await.expect("read bytes from alias");
assert_eq!(n, 5);
assert_eq!(&buf[..n], b"hello");
}
#[tokio::test]
async fn test_from_file_read_reads_requested_range() {
let path = temp_file_path("from-file-read");
tokio::fs::write(&path, b"hello world")
.await
.expect("write temp file for reader test");
let file = tokio::fs::File::open(&path).await.expect("open temp file for reader test");
let mut reader = BytesBufferedReader::from_file_read(&file, 6, 5)
.await
.expect("read requested range into Bytes");
let mut output = Vec::new();
reader.read_to_end(&mut output).await.expect("drain reader output");
assert_eq!(output, b"world");
let _ = tokio::fs::remove_file(path).await;
}
#[tokio::test]
#[allow(deprecated)]
async fn test_from_file_mmap_legacy_alias_reads_requested_range() {
let path = temp_file_path("from-file-mmap");
tokio::fs::write(&path, b"hello world")
.await
.expect("write temp file for legacy reader test");
let file = tokio::fs::File::open(&path)
.await
.expect("open temp file for legacy reader test");
let mut reader = BytesBufferedReader::from_file_mmap(&file, 0, 5)
.await
.expect("read requested range through legacy alias");
let mut output = Vec::new();
reader.read_to_end(&mut output).await.expect("drain legacy reader output");
assert_eq!(output, b"hello");
let _ = tokio::fs::remove_file(path).await;
}
#[tokio::test]
async fn test_remaining_bytes() {
let data = Bytes::from("hello world");
let reader = BytesBufferedReader::from_bytes(data);
let remaining = reader.remaining_bytes();
assert_eq!(remaining.len(), 11);
assert_eq!(&remaining[..], b"hello world");
}
#[tokio::test]
async fn test_position() {
let data = Bytes::from("hello world");
let mut reader = BytesBufferedReader::from_bytes(data);
assert_eq!(reader.position(), 0);
let mut buf = [0u8; 5];
reader.read_exact(&mut buf[..]).await.unwrap();
assert_eq!(reader.position(), 5);
}
#[tokio::test]
async fn test_is_empty() {
let data = Bytes::from("");
let reader = BytesBufferedReader::from_bytes(data);
assert!(reader.is_empty());
let data = Bytes::from("hello");
let reader = BytesBufferedReader::from_bytes(data);
assert!(!reader.is_empty());
}
#[tokio::test]
#[allow(deprecated)]
async fn test_legacy_reader_alias() {
let data = Bytes::from("hello world");
let mut reader = ZeroCopyObjectReader::from_bytes(data);
let mut buf = [0u8; 5];
let n = reader.read(&mut buf[..]).await.expect("read bytes through legacy alias");
assert_eq!(n, 5);
assert_eq!(&buf[..n], b"hello");
}
}
-882
View File
@@ -1,882 +0,0 @@
// Copyright 2024 RustFS Team
//
// Licensed under the Apache License, Version 2.0 (the "License");
// you may not use this file except in compliance with the License.
// You may obtain a copy of the License at
//
// http://www.apache.org/licenses/LICENSE-2.0
//
// Unless required by applicable law or agreed to in writing, software
// distributed under the License is distributed on an "AS IS" BASIS,
// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
// See the License for the specific language governing permissions and
// limitations under the License.
//! I/O scheduler for adaptive buffer sizing and load management.
//!
//! This module provides the core I/O scheduling logic that determines
//! optimal buffer sizes, I/O strategies, and load management decisions.
use crate::config::IoSchedulerConfig;
use crate::io_profile::{AccessPattern, StorageMedia, StorageProfile};
use std::sync::atomic::{AtomicUsize, Ordering};
use std::time::Duration;
/// I/O priority levels.
#[derive(Debug, Clone, Copy, PartialEq, Eq, Hash, Default)]
pub enum IoPriority {
/// High priority for small, latency-sensitive operations.
High,
/// Normal priority for standard operations.
#[default]
Normal,
/// Low priority for large, throughput-oriented operations.
Low,
}
impl IoPriority {
/// Determine priority based on request size.
///
/// A negative `size` means the size is unknown (-1 by convention) and maps
/// to `Normal`; casting it to `usize` would wrap to a huge value and
/// misclassify the request as `Low`.
pub fn from_size(size: i64, high_threshold: usize, low_threshold: usize) -> Self {
if size < 0 {
return IoPriority::Normal;
}
let size = size as usize;
if size < high_threshold {
IoPriority::High
} else if size > low_threshold {
IoPriority::Low
} else {
IoPriority::Normal
}
}
/// Get the priority as a string for metrics labels.
pub fn as_str(&self) -> &'static str {
match self {
IoPriority::High => "high",
IoPriority::Normal => "normal",
IoPriority::Low => "low",
}
}
/// Check if this is high priority.
pub fn is_high(&self) -> bool {
matches!(self, IoPriority::High)
}
/// Check if this is normal priority.
pub fn is_normal(&self) -> bool {
matches!(self, IoPriority::Normal)
}
/// Check if this is low priority.
pub fn is_low(&self) -> bool {
matches!(self, IoPriority::Low)
}
}
impl std::fmt::Display for IoPriority {
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
write!(f, "{}", self.as_str())
}
}
/// I/O load level.
#[derive(Debug, Clone, Copy, PartialEq, Eq, Hash, PartialOrd, Default)]
pub enum IoLoadLevel {
/// Low load - system is underutilized.
Low,
/// Medium load - system is moderately utilized.
#[default]
Medium,
/// High load - system is heavily utilized.
High,
/// Critical load - system is overloaded.
Critical,
}
impl IoLoadLevel {
/// Get the load level as a string for metrics labels.
pub fn as_str(&self) -> &'static str {
match self {
IoLoadLevel::Low => "low",
IoLoadLevel::Medium => "medium",
IoLoadLevel::High => "high",
IoLoadLevel::Critical => "critical",
}
}
/// Determine load level from wait time.
pub fn from_wait_time(wait_time: Duration, low_threshold: Duration, high_threshold: Duration) -> Self {
if wait_time <= low_threshold {
IoLoadLevel::Low
} else if wait_time <= high_threshold {
IoLoadLevel::Medium
} else if wait_time <= high_threshold * 2 {
IoLoadLevel::High
} else {
IoLoadLevel::Critical
}
}
}
impl std::fmt::Display for IoLoadLevel {
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
write!(f, "{}", self.as_str())
}
}
/// Bandwidth tier classification.
#[derive(Debug, Clone, Copy, PartialEq, Eq, Hash, Default)]
pub enum BandwidthTier {
/// Low bandwidth (< 100 MB/s).
Low,
/// Medium bandwidth (100-500 MB/s).
#[default]
Medium,
/// High bandwidth (> 500 MB/s).
High,
/// Unknown bandwidth.
Unknown,
}
impl BandwidthTier {
/// Determine bandwidth tier from bytes per second.
pub fn from_bps(bps: u64) -> Self {
const MB: u64 = 1024 * 1024;
if bps < 100 * MB {
BandwidthTier::Low
} else if bps < 500 * MB {
BandwidthTier::Medium
} else {
BandwidthTier::High
}
}
/// Get the tier as a string for metrics labels.
pub fn as_str(&self) -> &'static str {
match self {
BandwidthTier::Low => "low",
BandwidthTier::Medium => "medium",
BandwidthTier::High => "high",
BandwidthTier::Unknown => "unknown",
}
}
}
/// I/O strategy decision.
#[derive(Debug, Clone)]
pub struct IoStrategy {
/// Buffer size to use for I/O operations.
pub buffer_size: usize,
/// Buffer multiplier based on storage media.
pub buffer_multiplier: f64,
/// Whether to enable readahead.
pub enable_readahead: bool,
/// Whether to use buffered I/O.
pub use_buffered_io: bool,
// Performance state
/// Current number of concurrent requests.
pub concurrent_requests: usize,
/// Observed bandwidth in bytes per second.
pub observed_bandwidth_bps: Option<u64>,
/// Bandwidth tier classification.
pub bandwidth_tier: BandwidthTier,
/// Current load level.
pub load_level: IoLoadLevel,
// Priority
/// I/O priority for this operation.
pub priority: IoPriority,
// Decision flags
/// Whether to throttle random I/O.
pub should_throttle_random_io: bool,
/// Whether to expand buffer for sequential access.
pub should_expand_for_sequential: bool,
/// Whether to reduce buffer due to concurrency.
pub should_reduce_for_concurrency: bool,
/// Whether to reduce buffer due to low bandwidth.
pub should_reduce_for_bandwidth: bool,
}
impl Default for IoStrategy {
fn default() -> Self {
Self {
buffer_size: 128 * 1024,
buffer_multiplier: 1.0,
enable_readahead: true,
use_buffered_io: true,
concurrent_requests: 0,
observed_bandwidth_bps: None,
bandwidth_tier: BandwidthTier::Medium,
load_level: IoLoadLevel::Low,
priority: IoPriority::Normal,
should_throttle_random_io: false,
should_expand_for_sequential: false,
should_reduce_for_concurrency: false,
should_reduce_for_bandwidth: false,
}
}
}
impl IoStrategy {
/// Create a new strategy with default values.
pub fn new() -> Self {
Self::default()
}
/// Create a strategy for sequential access.
pub fn sequential(buffer_size: usize) -> Self {
Self {
buffer_size,
enable_readahead: true,
should_expand_for_sequential: true,
..Self::default()
}
}
/// Create a strategy for random access.
pub fn random(buffer_size: usize) -> Self {
Self {
buffer_size,
enable_readahead: false,
should_throttle_random_io: true,
..Self::default()
}
}
}
/// I/O load metrics.
#[derive(Debug, Clone, Default)]
pub struct IoLoadMetrics {
/// Number of samples in the current window.
pub sample_count: usize,
/// Total wait time in the window.
pub total_wait_time: Duration,
/// Maximum wait time in the window.
pub max_wait_time: Duration,
/// Average wait time.
pub avg_wait_time: Duration,
/// Current load level.
pub load_level: IoLoadLevel,
}
impl IoLoadMetrics {
/// Create new load metrics.
pub fn new() -> Self {
Self::default()
}
/// Add a wait time sample.
pub fn add_sample(&mut self, wait_time: Duration) {
self.sample_count += 1;
self.total_wait_time += wait_time;
if wait_time > self.max_wait_time {
self.max_wait_time = wait_time;
}
self.avg_wait_time = if self.sample_count > 0 {
self.total_wait_time / self.sample_count as u32
} else {
Duration::ZERO
};
}
/// Update load level based on thresholds.
pub fn update_load_level(&mut self, low_threshold: Duration, high_threshold: Duration) {
self.load_level = IoLoadLevel::from_wait_time(self.avg_wait_time, low_threshold, high_threshold);
}
/// Reset the metrics.
pub fn reset(&mut self) {
*self = Self::default();
}
}
/// I/O scheduler.
pub struct IoScheduler {
/// Scheduler configuration.
config: IoSchedulerConfig,
/// Active request counter.
active_requests: AtomicUsize,
/// Load metrics.
load_metrics: std::sync::Mutex<IoLoadMetrics>,
}
impl IoScheduler {
/// Create a new I/O scheduler with the given configuration.
pub fn new(config: IoSchedulerConfig) -> Self {
Self {
config,
active_requests: AtomicUsize::new(0),
load_metrics: std::sync::Mutex::new(IoLoadMetrics::new()),
}
}
/// Create a new I/O scheduler with default configuration.
pub fn with_defaults() -> Self {
Self::new(IoSchedulerConfig::default())
}
/// Get the scheduler configuration.
pub fn config(&self) -> &IoSchedulerConfig {
&self.config
}
/// Get the current number of active requests.
pub fn active_requests(&self) -> usize {
self.active_requests.load(Ordering::Relaxed)
}
/// Increment the active request count.
pub fn increment_requests(&self) {
self.active_requests.fetch_add(1, Ordering::Relaxed);
}
/// Decrement the active request count.
pub fn decrement_requests(&self) {
self.active_requests.fetch_sub(1, Ordering::Relaxed);
}
/// Calculate I/O strategy for a request.
pub fn calculate_strategy(&self, file_size: i64, permit_wait_time: Duration, is_sequential: bool) -> IoStrategy {
let concurrent_requests = self.active_requests.load(Ordering::Relaxed);
// Determine priority based on file size
let priority = IoPriority::from_size(
file_size,
self.config.high_priority_size_threshold,
self.config.low_priority_size_threshold,
);
// Determine load level
let load_level =
IoLoadLevel::from_wait_time(permit_wait_time, self.config.load_low_threshold(), self.config.load_high_threshold());
// Calculate base buffer size
let base_buffer = self.config.base_buffer_size;
// Adjust for concurrency
let concurrency_factor = match concurrent_requests {
0..=2 => 1.0,
3..=4 => 0.75,
5..=8 => 0.5,
_ => 0.4,
};
// Adjust for load level
let load_factor = match load_level {
IoLoadLevel::Low => 1.2,
IoLoadLevel::Medium => 1.0,
IoLoadLevel::High => 0.7,
IoLoadLevel::Critical => 0.5,
};
// Adjust for access pattern
let sequential_factor = if is_sequential { 1.5 } else { 1.0 };
// Calculate final buffer size
let buffer_size = (base_buffer as f64 * concurrency_factor * load_factor * sequential_factor) as usize;
let buffer_size = buffer_size.clamp(self.config.min_buffer_size, self.config.max_buffer_size);
IoStrategy {
buffer_size,
buffer_multiplier: concurrency_factor * load_factor * sequential_factor,
enable_readahead: is_sequential && load_level != IoLoadLevel::Critical,
use_buffered_io: true,
concurrent_requests,
observed_bandwidth_bps: None,
bandwidth_tier: BandwidthTier::Unknown,
load_level,
priority,
should_throttle_random_io: !is_sequential && load_level >= IoLoadLevel::High,
should_expand_for_sequential: is_sequential && load_level <= IoLoadLevel::Medium,
should_reduce_for_concurrency: concurrent_requests > 4,
should_reduce_for_bandwidth: false,
}
}
/// Calculate multi-factor I/O strategy.
pub fn calculate_multi_factor_strategy(
&self,
file_size: i64,
permit_wait_time: Duration,
is_sequential: bool,
storage_profile: Option<&StorageProfile>,
) -> IoStrategy {
let mut strategy = self.calculate_strategy(file_size, permit_wait_time, is_sequential);
// Apply storage profile adjustments
if let Some(profile) = storage_profile {
// Adjust buffer size based on storage media
let media_factor = match profile.media {
StorageMedia::Nvme => 1.5,
StorageMedia::Ssd => 1.2,
StorageMedia::Hdd => 0.8,
StorageMedia::Unknown => 1.0,
};
strategy.buffer_size = (strategy.buffer_size as f64 * media_factor).min(self.config.max_buffer_size as f64) as usize;
// Apply sequential boost if applicable
if is_sequential {
strategy.buffer_size = (strategy.buffer_size as f64 * profile.sequential_boost_multiplier)
.min(self.config.max_buffer_size as f64) as usize;
}
// Apply random penalty if applicable
if !is_sequential {
strategy.buffer_size = (strategy.buffer_size as f64 * profile.random_penalty_multiplier)
.max(self.config.min_buffer_size as f64) as usize;
}
// Update readahead preference
strategy.enable_readahead = strategy.enable_readahead && profile.prefers_readahead;
}
strategy
}
/// Record a wait time sample for load tracking.
pub fn record_wait_time(&self, wait_time: Duration) {
if let Ok(mut metrics) = self.load_metrics.lock() {
metrics.add_sample(wait_time);
metrics.update_load_level(self.config.load_low_threshold(), self.config.load_high_threshold());
}
}
/// Get current load metrics.
pub fn load_metrics(&self) -> IoLoadMetrics {
if let Ok(metrics) = self.load_metrics.lock() {
metrics.clone()
} else {
IoLoadMetrics::default()
}
}
}
impl Default for IoScheduler {
fn default() -> Self {
Self::with_defaults()
}
}
// ============================================================================
// Buffer Size Calculation Functions
// ============================================================================
/// Constants for buffer size calculations.
pub const KI_B: usize = 1024;
pub const MI_B: usize = 1024 * 1024;
/// Get concurrency-aware buffer size.
///
/// Adjusts buffer size based on the current level of concurrent requests.
/// Higher concurrency leads to smaller buffers to reduce memory pressure.
///
/// # Arguments
///
/// * `file_size` - Size of the file being read (-1 if unknown)
/// * `base_buffer_size` - Base buffer size from workload profile
///
/// # Returns
///
/// Adjusted buffer size in bytes
pub fn get_concurrency_aware_buffer_size(file_size: i64, base_buffer_size: usize) -> usize {
// Get current concurrency level from global counter
let concurrent_requests = 1; // Default to 1 if no global counter available
// Define concurrency thresholds
let medium_threshold = 4;
let high_threshold = 8;
// Calculate adaptive multiplier based on concurrency
let adaptive_multiplier = if concurrent_requests <= 2 {
// Low concurrency (1-2): use full buffer size
1.0
} else if concurrent_requests <= medium_threshold {
// Medium concurrency (3-4): slightly reduce buffer size (75% of base)
0.75
} else if concurrent_requests <= high_threshold {
// Higher concurrency (5-8): more aggressive reduction (50% of base)
0.5
} else {
// Very high concurrency (>8): minimize memory per request (40% of base)
0.4
};
// Calculate the adjusted buffer size
let adjusted_size = (base_buffer_size as f64 * adaptive_multiplier) as usize;
// Ensure we stay within reasonable bounds
let min_buffer = if file_size > 0 && file_size < 100 * KI_B as i64 {
32 * KI_B // For very small files, use minimum buffer
} else {
64 * KI_B // Standard minimum buffer size
};
let max_buffer = if concurrent_requests > high_threshold {
256 * KI_B // Cap at 256KB for high concurrency
} else {
MI_B // Cap at 1MB for lower concurrency
};
adjusted_size.clamp(min_buffer, max_buffer)
}
/// Advanced concurrency-aware buffer sizing with file size optimization.
///
/// This enhanced version considers both concurrency level and file size patterns
/// to provide even better performance characteristics.
///
/// # Arguments
///
/// * `file_size` - Size of the file being read (-1 if unknown)
/// * `base_buffer_size` - Baseline buffer size from workload profile
/// * `is_sequential` - Whether this is a sequential read (hint for optimization)
/// * `concurrent_requests` - Current number of concurrent requests
///
/// # Returns
///
/// Optimized buffer size in bytes
pub fn get_advanced_buffer_size(
file_size: i64,
base_buffer_size: usize,
is_sequential: bool,
concurrent_requests: usize,
) -> usize {
// For very small files, use smaller buffers regardless of concurrency
if file_size > 0 && file_size < 256 * KI_B as i64 {
return (file_size as usize / 4).clamp(16 * KI_B, 64 * KI_B);
}
// Base calculation from standard function
let standard_size = get_concurrency_aware_buffer_size(file_size, base_buffer_size);
let medium_threshold = 4;
let high_threshold = 8;
// For sequential reads, we can be more aggressive with buffer sizes
if is_sequential && concurrent_requests <= medium_threshold {
// Boost buffer size for sequential reads under low concurrency
let boosted = (standard_size as f64 * 1.5) as usize;
return boosted.min(MI_B);
}
// For random reads under high concurrency, reduce buffer size
if !is_sequential && concurrent_requests > high_threshold {
let reduced = (standard_size as f64 * 0.7) as usize;
return reduced.max(32 * KI_B);
}
standard_size
}
/// Get buffer size with storage media optimization.
///
/// Adjusts buffer size based on storage media characteristics.
///
/// # Arguments
///
/// * `base_size` - Base buffer size
/// * `media` - Storage media type
///
/// # Returns
///
/// Optimized buffer size for the storage media
pub fn get_buffer_size_for_media(base_size: usize, media: StorageMedia) -> usize {
let multiplier = match media {
StorageMedia::Nvme => 1.5, // NVMe can handle larger buffers
StorageMedia::Ssd => 1.2, // SSD benefits from moderate buffers
StorageMedia::Hdd => 0.8, // HDD prefers smaller buffers to reduce seek overhead
StorageMedia::Unknown => 1.0,
};
(base_size as f64 * multiplier).min(MI_B as f64) as usize
}
/// Calculate optimal buffer size using multi-factor analysis.
///
/// This is the main entry point for buffer size calculation, considering
/// all factors: concurrency, storage media, access pattern, and load.
///
/// # Arguments
///
/// * `file_size` - Size of the file being read
/// * `base_buffer_size` - Base buffer size
/// * `is_sequential` - Whether access is sequential
/// * `concurrent_requests` - Current concurrency level
/// * `media` - Storage media type
/// * `load_level` - Current I/O load level
///
/// # Returns
///
/// Optimally calculated buffer size
pub fn calculate_optimal_buffer_size(
file_size: i64,
base_buffer_size: usize,
is_sequential: bool,
concurrent_requests: usize,
media: StorageMedia,
load_level: IoLoadLevel,
) -> usize {
// Start with advanced buffer size calculation
let mut buffer_size = get_advanced_buffer_size(file_size, base_buffer_size, is_sequential, concurrent_requests);
// Apply storage media optimization
buffer_size = get_buffer_size_for_media(buffer_size, media);
// Apply load-based adjustment
let load_multiplier = match load_level {
IoLoadLevel::Low => 1.2,
IoLoadLevel::Medium => 1.0,
IoLoadLevel::High => 0.7,
IoLoadLevel::Critical => 0.5,
};
buffer_size = (buffer_size as f64 * load_multiplier) as usize;
// Final bounds check
buffer_size.clamp(32 * KI_B, MI_B)
}
/// I/O scheduling context for multi-factor strategy calculation.
#[derive(Debug, Clone)]
pub struct IoSchedulingContext {
/// File size in bytes (-1 if unknown).
pub file_size: i64,
/// Base buffer size from configuration.
pub base_buffer_size: usize,
/// Time spent waiting for permit.
pub permit_wait_duration: Duration,
/// Whether access is sequential.
pub is_sequential_hint: bool,
/// Detected access pattern.
pub access_pattern: AccessPattern,
/// Detected storage media.
pub storage_media: StorageMedia,
/// Observed bandwidth in bytes per second.
pub observed_bandwidth_bps: Option<u64>,
/// Current concurrent request count.
pub concurrent_requests: usize,
}
impl Default for IoSchedulingContext {
fn default() -> Self {
Self {
file_size: -1,
base_buffer_size: 128 * KI_B,
permit_wait_duration: Duration::ZERO,
is_sequential_hint: true,
access_pattern: AccessPattern::Unknown,
storage_media: StorageMedia::Unknown,
observed_bandwidth_bps: None,
concurrent_requests: 1,
}
}
}
impl IoSchedulingContext {
/// Create a new scheduling context.
pub fn new(file_size: i64, base_buffer_size: usize) -> Self {
Self {
file_size,
base_buffer_size,
..Self::default()
}
}
/// Builder pattern: set sequential hint.
pub fn with_sequential(mut self, is_sequential: bool) -> Self {
self.is_sequential_hint = is_sequential;
self.access_pattern = if is_sequential {
AccessPattern::Sequential
} else {
AccessPattern::Random
};
self
}
/// Builder pattern: set storage media.
pub fn with_media(mut self, media: StorageMedia) -> Self {
self.storage_media = media;
self
}
/// Builder pattern: set bandwidth.
pub fn with_bandwidth(mut self, bps: u64) -> Self {
self.observed_bandwidth_bps = Some(bps);
self
}
/// Builder pattern: set concurrency.
pub fn with_concurrency(mut self, count: usize) -> Self {
self.concurrent_requests = count;
self
}
/// Builder pattern: set wait duration.
pub fn with_wait_duration(mut self, duration: Duration) -> Self {
self.permit_wait_duration = duration;
self
}
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn test_io_priority() {
assert_eq!(IoPriority::from_size(1024, 64 * 1024, 4 * 1024 * 1024), IoPriority::High);
assert_eq!(IoPriority::from_size(1024 * 1024, 64 * 1024, 4 * 1024 * 1024), IoPriority::Normal);
assert_eq!(IoPriority::from_size(10 * 1024 * 1024, 64 * 1024, 4 * 1024 * 1024), IoPriority::Low);
}
#[test]
fn test_io_priority_unknown_size_is_normal() {
// -1 means "size unknown" and must not wrap to usize::MAX (=> Low).
assert_eq!(IoPriority::from_size(-1, 64 * 1024, 4 * 1024 * 1024), IoPriority::Normal);
assert_eq!(IoPriority::from_size(i64::MIN, 64 * 1024, 4 * 1024 * 1024), IoPriority::Normal);
}
#[test]
fn test_io_load_level() {
let low = Duration::from_millis(5);
let high = Duration::from_millis(50);
assert_eq!(IoLoadLevel::from_wait_time(Duration::from_millis(1), low, high), IoLoadLevel::Low);
assert_eq!(IoLoadLevel::from_wait_time(Duration::from_millis(20), low, high), IoLoadLevel::Medium);
assert_eq!(IoLoadLevel::from_wait_time(Duration::from_millis(60), low, high), IoLoadLevel::High);
assert_eq!(IoLoadLevel::from_wait_time(Duration::from_millis(150), low, high), IoLoadLevel::Critical);
}
#[test]
fn test_bandwidth_tier() {
assert_eq!(BandwidthTier::from_bps(50 * 1024 * 1024), BandwidthTier::Low);
assert_eq!(BandwidthTier::from_bps(200 * 1024 * 1024), BandwidthTier::Medium);
assert_eq!(BandwidthTier::from_bps(600 * 1024 * 1024), BandwidthTier::High);
}
#[test]
fn test_io_strategy_default() {
let strategy = IoStrategy::default();
assert!(strategy.buffer_size > 0);
assert!(strategy.enable_readahead);
}
#[test]
fn test_io_scheduler() {
let scheduler = IoScheduler::with_defaults();
let strategy = scheduler.calculate_strategy(1024 * 1024, Duration::from_millis(5), true);
assert!(strategy.buffer_size > 0);
assert!(strategy.enable_readahead);
assert_eq!(strategy.load_level, IoLoadLevel::Low);
}
#[test]
fn test_io_scheduler_with_concurrency() {
let scheduler = IoScheduler::with_defaults();
// Simulate concurrent requests
scheduler.increment_requests();
scheduler.increment_requests();
scheduler.increment_requests();
let strategy = scheduler.calculate_strategy(1024 * 1024, Duration::from_millis(5), true);
assert_eq!(strategy.concurrent_requests, 3);
}
#[test]
fn test_load_metrics() {
let mut metrics = IoLoadMetrics::new();
metrics.add_sample(Duration::from_millis(10));
metrics.add_sample(Duration::from_millis(20));
metrics.add_sample(Duration::from_millis(30));
assert_eq!(metrics.sample_count, 3);
assert_eq!(metrics.avg_wait_time, Duration::from_millis(20));
assert_eq!(metrics.max_wait_time, Duration::from_millis(30));
}
#[test]
fn test_get_concurrency_aware_buffer_size() {
// Test with default concurrency (1)
let size = get_concurrency_aware_buffer_size(1024 * 1024, 128 * KI_B);
assert!(size >= 64 * KI_B);
assert!(size <= MI_B);
// Test with small file
let size = get_concurrency_aware_buffer_size(50 * KI_B as i64, 128 * KI_B);
assert!(size >= 32 * KI_B);
}
#[test]
fn test_get_advanced_buffer_size() {
// Sequential read with low concurrency
let size = get_advanced_buffer_size(10 * MI_B as i64, 128 * KI_B, true, 2);
assert!(size >= 128 * KI_B);
// Random read with high concurrency
let size = get_advanced_buffer_size(10 * MI_B as i64, 128 * KI_B, false, 10);
assert!(size >= 32 * KI_B);
// Very small file
let size = get_advanced_buffer_size(100 * KI_B as i64, 128 * KI_B, true, 1);
assert!(size <= 64 * KI_B);
}
#[test]
fn test_get_buffer_size_for_media() {
let base = 128 * KI_B;
// NVMe should get larger buffers
let nvme_size = get_buffer_size_for_media(base, StorageMedia::Nvme);
assert!(nvme_size > base);
// SSD should get slightly larger buffers
let ssd_size = get_buffer_size_for_media(base, StorageMedia::Ssd);
assert!(ssd_size > base);
// HDD should get smaller buffers
let hdd_size = get_buffer_size_for_media(base, StorageMedia::Hdd);
assert!(hdd_size < base);
}
#[test]
fn test_calculate_optimal_buffer_size() {
// Low load, sequential, NVMe
let size = calculate_optimal_buffer_size(10 * MI_B as i64, 128 * KI_B, true, 2, StorageMedia::Nvme, IoLoadLevel::Low);
assert!(size >= 32 * KI_B);
assert!(size <= MI_B);
// Critical load, random, HDD
let size =
calculate_optimal_buffer_size(10 * MI_B as i64, 128 * KI_B, false, 10, StorageMedia::Hdd, IoLoadLevel::Critical);
assert!(size >= 32 * KI_B);
assert!(size <= MI_B);
}
#[test]
fn test_io_scheduling_context() {
let ctx = IoSchedulingContext::new(10 * MI_B as i64, 256 * KI_B)
.with_sequential(true)
.with_media(StorageMedia::Nvme)
.with_bandwidth(500 * MI_B as u64)
.with_concurrency(4);
assert_eq!(ctx.file_size, 10 * MI_B as i64);
assert_eq!(ctx.base_buffer_size, 256 * KI_B);
assert!(ctx.is_sequential_hint);
assert_eq!(ctx.storage_media, StorageMedia::Nvme);
assert_eq!(ctx.observed_bandwidth_bps, Some(500 * MI_B as u64));
assert_eq!(ctx.concurrent_requests, 4);
}
}
-320
View File
@@ -1,320 +0,0 @@
// Copyright 2024 RustFS Team
//
// Licensed under the Apache License, Version 2.0 (the "License");
// you may not use this file except in compliance with the License.
// You may obtain a copy of the License at
//
// http://www.apache.org/licenses/LICENSE-2.0
//
// Unless required by applicable law or agreed to in writing, software
// distributed under the License is distributed on an "AS IS" BASIS,
// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
// See the License for the specific language governing permissions and
// limitations under the License.
//! Shared memory pool for zero-copy data sharing.
//!
//! This module provides Arc-based shared memory management for
//! efficient cross-task data passing without serialization.
use std::convert::AsRef;
use std::ops::Deref;
use std::sync::Arc;
use std::sync::atomic::{AtomicU64, Ordering};
use std::time::Instant;
/// Shared memory pool configuration.
#[derive(Debug, Clone)]
pub struct SharedMemoryConfig {
/// Whether shared memory is enabled
pub enabled: bool,
/// Maximum pool size in bytes
pub max_pool_size: usize,
/// Maximum object size in bytes
pub max_object_size: usize,
}
impl Default for SharedMemoryConfig {
fn default() -> Self {
Self {
enabled: true,
max_pool_size: 100 * 1024 * 1024, // 100MB
max_object_size: 10 * 1024 * 1024, // 10MB
}
}
}
/// Shared memory pool statistics.
#[derive(Debug, Default)]
pub struct SharedMemoryStats {
/// Total number of objects created
pub total_objects: AtomicU64,
/// Total number of shared references
pub total_shared_refs: AtomicU64,
/// Current memory usage in bytes
pub current_memory: AtomicU64,
/// Peak memory usage in bytes
pub peak_memory: AtomicU64,
}
/// Arc data metadata.
#[derive(Clone, Debug)]
pub struct ArcMetadata {
/// Size of the data (if measurable)
pub size: Option<usize>,
/// Creation timestamp
pub created_at: Instant,
}
/// Arc-based data wrapper for zero-copy sharing.
///
/// This wrapper uses Arc to enable shared ownership of data
/// across multiple tasks without copying.
pub struct ArcData<T> {
/// The wrapped data
inner: Arc<T>,
/// Metadata about the data
metadata: ArcMetadata,
}
impl<T> Clone for ArcData<T> {
fn clone(&self) -> Self {
Self {
inner: Arc::clone(&self.inner),
metadata: self.metadata.clone(),
}
}
}
impl<T> ArcData<T> {
/// Create a new ArcData wrapper.
pub fn new(data: T) -> Self {
ArcData {
inner: Arc::new(data),
metadata: ArcMetadata {
size: None,
created_at: Instant::now(),
},
}
}
/// Create a new ArcData wrapper with known size.
pub fn with_size(data: T, size: usize) -> Self {
ArcData {
inner: Arc::new(data),
metadata: ArcMetadata {
size: Some(size),
created_at: Instant::now(),
},
}
}
/// Get the reference count.
pub fn ref_count(&self) -> usize {
Arc::strong_count(&self.inner)
}
/// Convert into the underlying Arc.
pub fn into_arc(self) -> Arc<T> {
self.inner
}
/// Get the metadata.
pub fn metadata(&self) -> &ArcMetadata {
&self.metadata
}
/// Get the size if known.
pub fn size(&self) -> Option<usize> {
self.metadata.size
}
}
impl<T> AsRef<T> for ArcData<T> {
fn as_ref(&self) -> &T {
&self.inner
}
}
impl<T> Deref for ArcData<T> {
type Target = T;
fn deref(&self) -> &Self::Target {
&self.inner
}
}
impl<T> std::fmt::Debug for ArcData<T> {
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
f.debug_struct("ArcData")
.field("ref_count", &self.ref_count())
.field("metadata", &self.metadata)
.finish()
}
}
/// Shared memory pool for managing Arc-based shared data.
pub struct SharedMemoryPool {
config: SharedMemoryConfig,
stats: SharedMemoryStats,
}
impl SharedMemoryPool {
/// Create a new shared memory pool with the given configuration.
pub fn new(config: SharedMemoryConfig) -> Self {
Self {
config,
stats: SharedMemoryStats::default(),
}
}
/// Create a new shared memory pool with default configuration.
pub fn with_defaults() -> Self {
Self::new(SharedMemoryConfig::default())
}
/// Create shared data.
///
/// This method wraps the data in an ArcData for zero-copy sharing.
pub fn create<T>(&self, data: T) -> ArcData<T> {
self.stats.total_objects.fetch_add(1, Ordering::Relaxed);
ArcData::new(data)
}
/// Create shared data with known size.
///
/// This method tracks memory usage for statistics.
pub fn create_with_size<T>(&self, data: T, size: usize) -> ArcData<T> {
self.stats.total_objects.fetch_add(1, Ordering::Relaxed);
// Update memory statistics
self.stats.current_memory.fetch_add(size as u64, Ordering::Relaxed);
// Update peak memory
let current = self.stats.current_memory.load(Ordering::Relaxed);
let mut peak = self.stats.peak_memory.load(Ordering::Relaxed);
if current > peak {
peak = current;
self.stats.peak_memory.store(peak, Ordering::Relaxed);
}
ArcData::with_size(data, size)
}
/// Share data by increasing reference count.
///
/// This method creates a new ArcData that shares the underlying data
/// without copying.
pub fn share<T>(&self, data: &ArcData<T>) -> ArcData<T> {
self.stats.total_shared_refs.fetch_add(1, Ordering::Relaxed);
data.clone()
}
/// Get the statistics for this pool.
pub fn stats(&self) -> &SharedMemoryStats {
&self.stats
}
/// Get the configuration for this pool.
pub fn config(&self) -> &SharedMemoryConfig {
&self.config
}
/// Check if the pool is enabled.
pub fn is_enabled(&self) -> bool {
self.config.enabled
}
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn test_arc_data_new() {
let data = vec![1u8, 2, 3, 4, 5];
let arc_data = ArcData::new(data.clone());
assert_eq!(arc_data.as_ref(), &data);
assert_eq!(arc_data.ref_count(), 1);
}
#[test]
fn test_arc_data_clone() {
let data = vec![1u8, 2, 3, 4, 5];
let arc_data = ArcData::new(data);
assert_eq!(arc_data.ref_count(), 1);
let arc_data2 = arc_data.clone();
assert_eq!(arc_data.ref_count(), 2);
assert_eq!(arc_data2.ref_count(), 2);
let arc_data3 = arc_data.clone();
assert_eq!(arc_data.ref_count(), 3);
assert_eq!(arc_data2.ref_count(), 3);
assert_eq!(arc_data3.ref_count(), 3);
}
#[test]
fn test_arc_data_deref() {
let data = vec![1u8, 2, 3, 4, 5];
let arc_data = ArcData::new(data);
// Test Deref trait
assert_eq!(arc_data.len(), 5);
assert_eq!(arc_data[0], 1);
}
#[test]
fn test_shared_memory_pool_create() {
let pool = SharedMemoryPool::with_defaults();
let data = vec![1u8, 2, 3, 4, 5];
let arc_data = pool.create(data.clone());
assert_eq!(arc_data.as_ref(), &data);
assert_eq!(pool.stats().total_objects.load(Ordering::Relaxed), 1);
}
#[test]
fn test_shared_memory_pool_share() {
let pool = SharedMemoryPool::with_defaults();
let data = vec![1u8, 2, 3, 4, 5];
let arc_data = pool.create(data);
assert_eq!(arc_data.ref_count(), 1);
let shared = pool.share(&arc_data);
assert_eq!(arc_data.ref_count(), 2);
assert_eq!(shared.ref_count(), 2);
assert_eq!(pool.stats().total_shared_refs.load(Ordering::Relaxed), 1);
}
#[test]
fn test_shared_memory_pool_with_size() {
let pool = SharedMemoryPool::with_defaults();
let data = vec![1u8; 1024];
let arc_data = pool.create_with_size(data, 1024);
assert_eq!(arc_data.size(), Some(1024));
assert_eq!(pool.stats().current_memory.load(Ordering::Relaxed), 1024);
}
#[test]
fn test_default_config() {
let config = SharedMemoryConfig::default();
assert!(config.enabled);
assert_eq!(config.max_pool_size, 100 * 1024 * 1024);
assert_eq!(config.max_object_size, 10 * 1024 * 1024);
}
}
-501
View File
@@ -1,501 +0,0 @@
// Copyright 2024 RustFS Team
//
// Licensed under the Apache License, Version 2.0 (the "License");
// you may not use this file except in compliance with the License.
// You may obtain a copy of the License at
//
// http://www.apache.org/licenses/LICENSE-2.0
//
// Unless required by applicable law or agreed to in writing, software
// distributed under the License is distributed on an "AS IS" BASIS,
// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
// See the License for the specific language governing permissions and
// limitations under the License.
//! Timeout wrapper for I/O operations.
//!
//! This module provides timeout management for I/O operations with
//! dynamic timeout calculation based on operation size.
use std::sync::atomic::{AtomicU64, Ordering};
use std::time::{Duration, Instant};
/// Timeout configuration.
#[derive(Debug, Clone)]
pub struct TimeoutConfig {
/// Base timeout for small operations.
pub base_timeout: Duration,
/// Timeout per MB of data.
pub timeout_per_mb: Duration,
/// Maximum timeout.
pub max_timeout: Duration,
/// Minimum timeout.
pub min_timeout: Duration,
/// GetObject operation timeout.
pub get_object_timeout: Duration,
/// PutObject operation timeout.
pub put_object_timeout: Duration,
/// ListObjects operation timeout.
pub list_objects_timeout: Duration,
/// Whether dynamic timeout is enabled.
pub enable_dynamic_timeout: bool,
}
impl Default for TimeoutConfig {
fn default() -> Self {
Self {
base_timeout: Duration::from_secs(5),
timeout_per_mb: Duration::from_millis(100),
max_timeout: Duration::from_secs(300),
min_timeout: Duration::from_secs(1),
get_object_timeout: Duration::from_secs(30),
put_object_timeout: Duration::from_secs(60),
list_objects_timeout: Duration::from_secs(10),
enable_dynamic_timeout: true,
}
}
}
impl TimeoutConfig {
/// Create new timeout configuration.
pub fn new() -> Self {
Self::default()
}
/// Calculate dynamic timeout based on size.
pub fn calculate_timeout(&self, size_bytes: u64) -> Duration {
if !self.enable_dynamic_timeout {
return self.base_timeout;
}
let mb = size_bytes as f64 / (1024.0 * 1024.0);
let timeout = self.base_timeout + self.timeout_per_mb.mul_f64(mb);
timeout.clamp(self.min_timeout, self.max_timeout)
}
/// Validate the configuration.
pub fn validate(&self) -> Result<(), TimeoutError> {
if self.min_timeout > self.max_timeout {
return Err(TimeoutError::InvalidConfig("min_timeout must be <= max_timeout".to_string()));
}
if self.base_timeout < self.min_timeout || self.base_timeout > self.max_timeout {
return Err(TimeoutError::InvalidConfig(
"base_timeout must be between min_timeout and max_timeout".to_string(),
));
}
Ok(())
}
}
/// Timeout error.
#[derive(Debug, Clone, thiserror::Error)]
pub enum TimeoutError {
/// Operation timed out.
#[error("Operation timed out after {0:?}")]
TimedOut(Duration),
/// Invalid configuration.
#[error("Invalid timeout config: {0}")]
InvalidConfig(String),
}
/// Operation progress tracker.
#[derive(Debug)]
pub struct OperationProgress {
/// Total size (if known).
pub total_size: Option<u64>,
/// Bytes processed.
bytes_processed: AtomicU64,
/// Last update time.
last_update: std::sync::Mutex<Instant>,
/// Stale timeout.
stale_timeout: Duration,
/// Start time for transfer rate calculation.
start_time: Instant,
}
impl OperationProgress {
/// Create new operation progress.
pub fn new(total_size: Option<u64>, stale_timeout: Duration) -> Self {
Self {
total_size,
bytes_processed: AtomicU64::new(0),
last_update: std::sync::Mutex::new(Instant::now()),
stale_timeout,
start_time: Instant::now(),
}
}
/// Update progress.
pub fn update(&self, bytes: u64) {
self.bytes_processed.store(bytes, Ordering::Relaxed);
if let Ok(mut last) = self.last_update.lock() {
*last = Instant::now();
}
}
/// Add to progress.
pub fn add(&self, bytes: u64) {
self.bytes_processed.fetch_add(bytes, Ordering::Relaxed);
if let Ok(mut last) = self.last_update.lock() {
*last = Instant::now();
}
}
/// Get current progress.
pub fn current(&self) -> u64 {
self.bytes_processed.load(Ordering::Relaxed)
}
/// Check if progress is stale.
pub fn is_stale(&self) -> bool {
if let Ok(last) = self.last_update.lock() {
last.elapsed() > self.stale_timeout
} else {
false
}
}
/// Get progress percentage.
pub fn progress_percent(&self) -> Option<f64> {
self.total_size.map(|total| {
if total == 0 {
100.0
} else {
let processed = self.bytes_processed.load(Ordering::Relaxed);
(processed as f64 / total as f64 * 100.0).min(100.0)
}
})
}
/// Get remaining bytes.
pub fn remaining(&self) -> Option<u64> {
self.total_size.map(|total| {
let processed = self.bytes_processed.load(Ordering::Relaxed);
total.saturating_sub(processed)
})
}
/// Calculate transfer rate in bytes per second.
///
/// Returns 0 if no time has elapsed or no data transferred.
pub fn transfer_rate(&self) -> u64 {
let processed = self.bytes_processed.load(Ordering::Relaxed);
if processed == 0 {
return 0;
}
let elapsed = self.start_time.elapsed().as_secs_f64();
if elapsed > 0.0 {
(processed as f64 / elapsed) as u64
} else {
0
}
}
}
/// Request timeout wrapper.
pub struct RequestTimeoutWrapper {
/// Configuration.
config: TimeoutConfig,
/// Start time.
start_time: Instant,
/// Operation progress.
progress: Option<OperationProgress>,
}
impl RequestTimeoutWrapper {
/// Create a new timeout wrapper.
pub fn new(config: TimeoutConfig) -> Self {
Self {
config,
start_time: Instant::now(),
progress: None,
}
}
/// Create with progress tracking.
pub fn with_progress(config: TimeoutConfig, total_size: Option<u64>, stale_timeout: Duration) -> Self {
Self {
config,
start_time: Instant::now(),
progress: Some(OperationProgress::new(total_size, stale_timeout)),
}
}
/// Get the configuration.
pub fn config(&self) -> &TimeoutConfig {
&self.config
}
/// Get elapsed time.
pub fn elapsed(&self) -> Duration {
self.start_time.elapsed()
}
/// Get remaining time.
pub fn remaining(&self, timeout: Duration) -> Option<Duration> {
let elapsed = self.elapsed();
if elapsed >= timeout { None } else { Some(timeout - elapsed) }
}
/// Check if timed out.
pub fn is_timed_out(&self, size: Option<u64>) -> bool {
let timeout = self.get_timeout(size);
self.elapsed() > timeout
}
/// Get the timeout for a given size.
pub fn get_timeout(&self, size: Option<u64>) -> Duration {
if self.config.enable_dynamic_timeout {
if let Some(s) = size {
self.config.calculate_timeout(s)
} else {
self.config.base_timeout
}
} else {
self.config.base_timeout
}
}
/// Check if timed out and return error if so.
pub fn check_timeout(&self, size: Option<u64>) -> Result<(), TimeoutError> {
if self.is_timed_out(size) {
Err(TimeoutError::TimedOut(self.get_timeout(size)))
} else {
Ok(())
}
}
/// Get progress.
pub fn progress(&self) -> Option<&OperationProgress> {
self.progress.as_ref()
}
/// Update progress.
pub fn update_progress(&self, bytes: u64) {
if let Some(ref progress) = self.progress {
progress.update(bytes);
}
}
/// Check if operation is stalled (no progress for a while).
pub fn is_stalled(&self) -> bool {
self.progress.as_ref().is_some_and(|p| p.is_stale())
}
/// Get progress percentage.
pub fn progress_percent(&self) -> Option<f64> {
self.progress.as_ref().and_then(|p| p.progress_percent())
}
}
/// Timeout statistics.
#[derive(Debug, Default)]
pub struct TimeoutStats {
/// Total operations.
pub total_operations: AtomicU64,
/// Timed out operations.
pub timed_out: AtomicU64,
/// Total wait time in nanoseconds.
pub total_wait_time_ns: AtomicU64,
/// Maximum wait time in nanoseconds.
pub max_wait_time_ns: AtomicU64,
}
impl TimeoutStats {
/// Create new timeout statistics.
pub fn new() -> Self {
Self::default()
}
/// Record an operation.
pub fn record_operation(&self, wait_time: Duration) {
self.total_operations.fetch_add(1, Ordering::Relaxed);
let ns = wait_time.as_nanos() as u64;
self.total_wait_time_ns.fetch_add(ns, Ordering::Relaxed);
let mut current = self.max_wait_time_ns.load(Ordering::Relaxed);
while ns > current {
match self
.max_wait_time_ns
.compare_exchange_weak(current, ns, Ordering::Relaxed, Ordering::Relaxed)
{
Ok(_) => break,
Err(actual) => current = actual,
}
}
}
/// Record a timeout.
pub fn record_timeout(&self) {
self.timed_out.fetch_add(1, Ordering::Relaxed);
}
/// Get timeout rate.
pub fn timeout_rate(&self) -> f64 {
let total = self.total_operations.load(Ordering::Relaxed);
let timed_out = self.timed_out.load(Ordering::Relaxed);
if total == 0 { 0.0 } else { timed_out as f64 / total as f64 }
}
/// Get average wait time.
pub fn avg_wait_time(&self) -> Duration {
let total = self.total_wait_time_ns.load(Ordering::Relaxed);
let count = self.total_operations.load(Ordering::Relaxed);
total.checked_div(count).map(Duration::from_nanos).unwrap_or(Duration::ZERO)
}
/// Reset statistics.
pub fn reset(&self) {
self.total_operations.store(0, Ordering::Relaxed);
self.timed_out.store(0, Ordering::Relaxed);
self.total_wait_time_ns.store(0, Ordering::Relaxed);
self.max_wait_time_ns.store(0, Ordering::Relaxed);
}
}
/// Calculate adaptive timeout based on historical data and current conditions.
///
/// This function adjusts the timeout based on:
/// - Historical transfer rate
/// - Recent timeout count
/// - Object size
pub fn calculate_adaptive_timeout(
base_timeout: Duration,
historical_rate_bps: Option<u64>,
recent_timeout_count: u32,
object_size: u64,
) -> Duration {
// If we have recent timeouts, increase timeout
let timeout_multiplier = if recent_timeout_count > 3 {
2.0 // Double timeout if many recent timeouts
} else if recent_timeout_count > 1 {
1.5 // 50% increase if some timeouts
} else {
1.0 // No adjustment
};
// Adaptive timeout bounds: 5 seconds minimum, 10 minutes maximum.
const MIN_SECS: f64 = 5.0;
const MAX_SECS: f64 = 600.0;
// If we have historical rate data, use it for estimation
let estimated_secs = match historical_rate_bps {
Some(rate) if rate > 0 => (object_size as f64 / rate as f64) * 1.2, // 20% buffer
_ => base_timeout.as_secs_f64(),
};
// Clamp BEFORE constructing the Duration: `from_secs_f64` panics when the
// estimate overflows Duration (huge object_size with a tiny historical rate).
Duration::from_secs_f64((estimated_secs * timeout_multiplier).clamp(MIN_SECS, MAX_SECS))
}
/// Estimate bytes per second transfer rate.
///
/// This is used for adaptive timeout calculation.
pub fn estimate_bytes_per_second(object_size: u64, expected_duration: Duration) -> u64 {
let secs = expected_duration.as_secs_f64();
if secs > 0.0 {
(object_size as f64 / secs) as u64
} else {
// Return a reasonable default (1 MB/s)
1024 * 1024
}
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn test_timeout_config() {
let config = TimeoutConfig::default();
assert!(config.validate().is_ok());
// Small file
let timeout = config.calculate_timeout(1024);
assert!(timeout >= config.min_timeout);
// Large file
let timeout = config.calculate_timeout(100 * 1024 * 1024);
assert!(timeout <= config.max_timeout);
}
#[test]
fn test_timeout_config_validation() {
let config = TimeoutConfig {
min_timeout: Duration::from_secs(10),
max_timeout: Duration::from_secs(5),
..Default::default()
};
assert!(config.validate().is_err());
}
#[test]
fn test_adaptive_timeout_extreme_estimate_does_not_panic() {
// A huge object with a tiny historical rate used to overflow
// Duration::from_secs_f64 and panic; it must clamp to the upper bound.
let timeout = calculate_adaptive_timeout(Duration::from_secs(30), Some(1), 0, u64::MAX);
assert_eq!(timeout, Duration::from_secs(600));
// Tiny estimates clamp to the lower bound.
let timeout = calculate_adaptive_timeout(Duration::from_secs(30), Some(u64::MAX), 0, 1);
assert_eq!(timeout, Duration::from_secs(5));
}
#[test]
fn test_operation_progress() {
let progress = OperationProgress::new(Some(1000), Duration::from_secs(5));
assert_eq!(progress.current(), 0);
assert_eq!(progress.progress_percent(), Some(0.0));
progress.update(500);
assert_eq!(progress.current(), 500);
assert_eq!(progress.progress_percent(), Some(50.0));
progress.add(300);
assert_eq!(progress.current(), 800);
assert_eq!(progress.remaining(), Some(200));
}
#[test]
fn test_request_timeout_wrapper() {
let config = TimeoutConfig {
base_timeout: Duration::from_millis(100),
enable_dynamic_timeout: false,
..Default::default()
};
let wrapper = RequestTimeoutWrapper::new(config);
assert!(!wrapper.is_timed_out(None));
std::thread::sleep(Duration::from_millis(150));
assert!(wrapper.is_timed_out(None));
assert!(wrapper.check_timeout(None).is_err());
}
#[test]
fn test_timeout_stats() {
let stats = TimeoutStats::new();
stats.record_operation(Duration::from_millis(10));
stats.record_operation(Duration::from_millis(20));
stats.record_timeout();
assert_eq!(stats.total_operations.load(Ordering::Relaxed), 2);
assert_eq!(stats.timed_out.load(Ordering::Relaxed), 1);
assert!((stats.timeout_rate() - 0.5).abs() < 0.01);
}
#[test]
fn test_progress_tracking() {
let config = TimeoutConfig::default();
let wrapper = RequestTimeoutWrapper::with_progress(config, Some(1000), Duration::from_secs(1));
wrapper.update_progress(500);
assert_eq!(wrapper.progress_percent(), Some(50.0));
assert!(!wrapper.is_stalled());
}
}
-443
View File
@@ -1,443 +0,0 @@
// Copyright 2024 RustFS Team
//
// Licensed under the Apache License, Version 2.0 (the "License");
// you may not use this file except in compliance with the License.
// You may obtain a copy of the License at
//
// http://www.apache.org/licenses/LICENSE-2.0
//
// Unless required by applicable law or agreed to in writing, software
// distributed under the License is distributed on an "AS IS" BASIS,
// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
// See the License for the specific language governing permissions and
// limitations under the License.
//! BytesMut-backed object writer for optimized write operations.
//!
//! It uses `BytesMut` for efficient buffering; writes into that buffer may
//! still copy input bytes. The historical `ZeroCopyObjectWriter` name remains
//! available as a deprecated compatibility alias.
use bytes::{BufMut, Bytes, BytesMut};
use std::pin::Pin;
use std::task::{Context, Poll};
use tokio::io::AsyncWrite;
/// BytesMut-backed object writer for optimized write operations.
///
/// This writer minimizes memory allocations by:
/// - Using BytesMut for efficient buffer growth
/// - Accepting `Bytes` inputs for efficient buffer handling
/// - Optional integration with BytesPool for buffer reuse
///
/// # Example
///
/// ```ignore
/// use rustfs_io_core::BytesMutWriter;
/// use bytes::Bytes;
///
/// #[tokio::main]
/// async fn main() -> Result<(), Box<dyn std::error::Error>> {
/// let mut writer = BytesMutWriter::new();
///
/// // Write into the internal BytesMut buffer
/// let data = Bytes::from("hello world");
/// writer.write_buffered(data).await?;
///
/// // Get the result as Bytes (zero-copy conversion)
/// let result = writer.into_bytes();
///
/// Ok(())
/// }
/// ```
pub struct BytesMutWriter {
/// Internal buffer using BytesMut for efficient growth
buffer: BytesMut,
/// Total bytes written
bytes_written: usize,
/// Whether the writer has been finalized
finalized: bool,
}
/// Historical name for the BytesMut-backed object writer.
#[deprecated(since = "1.0.0-beta.8", note = "use BytesMutWriter; writes append into a BytesMut buffer")]
pub type ZeroCopyObjectWriter = BytesMutWriter;
impl BytesMutWriter {
/// Create a new bytes-backed object writer with default capacity (8KB).
///
/// # Example
///
/// ```ignore
/// let writer = BytesMutWriter::new();
/// ```
pub fn new() -> Self {
Self::with_capacity(8 * 1024)
}
/// Create a new bytes-backed object writer with specified capacity.
///
/// # Arguments
///
/// * `capacity` - Initial buffer capacity in bytes
///
/// # Example
///
/// ```ignore
/// let writer = BytesMutWriter::with_capacity(64 * 1024);
/// ```
pub fn with_capacity(capacity: usize) -> Self {
Self {
buffer: BytesMut::with_capacity(capacity),
bytes_written: 0,
finalized: false,
}
}
/// Write data into the internal buffer.
///
/// This method accepts `Bytes` for API compatibility, then appends the
/// bytes into the internal `BytesMut` buffer.
///
/// # Arguments
///
/// * `data` - Data to append to the internal buffer
///
/// # Returns
///
/// * `Ok(usize)` - Number of bytes written
/// * `Err(ZeroCopyWriteError)` - Write error
///
/// # Example
///
/// ```ignore
/// let data = Bytes::from("hello world");
/// let written = writer.write_buffered(data).await?;
/// ```
pub async fn write_buffered(&mut self, data: Bytes) -> Result<usize, ZeroCopyWriteError> {
if self.finalized {
return Err(ZeroCopyWriteError::Finalized("Cannot write to finalized writer".to_string()));
}
let len = data.len();
self.buffer.put(data);
self.bytes_written += len;
Ok(len)
}
/// Historical name for `write_buffered`.
#[deprecated(
since = "1.0.0-beta.8",
note = "use write_buffered; this method appends bytes into an internal buffer"
)]
pub async fn write_zero_copy(&mut self, data: Bytes) -> Result<usize, ZeroCopyWriteError> {
self.write_buffered(data).await
}
/// Write a slice of data.
///
/// # Arguments
///
/// * `data` - Data slice to write
///
/// # Returns
///
/// * `Ok(usize)` - Number of bytes written
/// * `Err(ZeroCopyWriteError)` - Write error
pub async fn write_slice(&mut self, data: &[u8]) -> Result<usize, ZeroCopyWriteError> {
if self.finalized {
return Err(ZeroCopyWriteError::Finalized("Cannot write to finalized writer".to_string()));
}
let len = data.len();
self.buffer.put_slice(data);
self.bytes_written += len;
Ok(len)
}
/// Finalize the writer and consume it, returning the written data as Bytes.
///
/// This converts the internal BytesMut to Bytes, which is a zero-copy
/// operation that freezes the buffer.
///
/// # Returns
///
/// The written data as Bytes
///
/// # Example
///
/// ```ignore
/// let result = writer.into_bytes();
/// ```
pub fn into_bytes(mut self) -> Bytes {
self.finalized = true;
self.buffer.freeze()
}
/// Get the current buffer as a slice (without consuming).
///
/// # Returns
///
/// Slice of the current buffer content
pub fn as_slice(&self) -> &[u8] {
&self.buffer[..]
}
/// Get the total number of bytes written.
///
/// # Returns
///
/// Number of bytes written
pub fn bytes_written(&self) -> usize {
self.bytes_written
}
/// Get the current buffer capacity.
///
/// # Returns
///
/// Current buffer capacity in bytes
pub fn capacity(&self) -> usize {
self.buffer.capacity()
}
/// Get the current buffer length.
///
/// # Returns
///
/// Current buffer length in bytes
pub fn len(&self) -> usize {
self.buffer.len()
}
/// Check if the buffer is empty.
///
/// # Returns
///
/// `true` if buffer is empty, `false` otherwise
pub fn is_empty(&self) -> bool {
self.buffer.is_empty()
}
/// Clear the buffer, resetting it to empty.
///
/// This does not change the capacity, just resets the length to 0.
pub fn clear(&mut self) {
self.buffer.clear();
self.bytes_written = 0;
self.finalized = false;
}
/// Reserve additional capacity in the buffer.
///
/// # Arguments
///
/// * `additional` - Additional capacity to reserve
pub fn reserve(&mut self, additional: usize) {
self.buffer.reserve(additional);
}
}
impl Default for BytesMutWriter {
fn default() -> Self {
Self::new()
}
}
impl std::fmt::Debug for BytesMutWriter {
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
f.debug_struct("BytesMutWriter")
.field("buffer_len", &self.buffer.len())
.field("buffer_capacity", &self.buffer.capacity())
.field("bytes_written", &self.bytes_written)
.field("finalized", &self.finalized)
.finish()
}
}
/// AsyncWrite implementation for BytesMutWriter.
///
/// This allows the writer to be used with tokio's async I/O utilities.
impl AsyncWrite for BytesMutWriter {
fn poll_write(mut self: Pin<&mut Self>, _cx: &mut Context<'_>, buf: &[u8]) -> Poll<Result<usize, tokio::io::Error>> {
if self.finalized {
return Poll::Ready(Err(tokio::io::Error::new(
tokio::io::ErrorKind::WriteZero,
"Cannot write to finalized writer",
)));
}
let len = buf.len();
self.buffer.put_slice(buf);
self.bytes_written += len;
Poll::Ready(Ok(len))
}
fn poll_flush(self: Pin<&mut Self>, _cx: &mut Context<'_>) -> Poll<Result<(), tokio::io::Error>> {
// Nothing to flush for in-memory buffer
Poll::Ready(Ok(()))
}
fn poll_shutdown(mut self: Pin<&mut Self>, _cx: &mut Context<'_>) -> Poll<Result<(), tokio::io::Error>> {
self.finalized = true;
Poll::Ready(Ok(()))
}
}
/// Zero-copy write error types.
#[derive(Debug, thiserror::Error)]
pub enum ZeroCopyWriteError {
/// I/O error occurred
#[error("I/O error: {0}")]
Io(#[from] tokio::io::Error),
/// Writer has been finalized and cannot accept more writes
#[error("Writer finalized: {0}")]
Finalized(String),
/// Invalid input provided
#[error("Invalid input: {0}")]
InvalidInput(String),
}
#[cfg(test)]
mod tests {
use super::*;
#[tokio::test]
async fn test_new_writer() {
let writer = BytesMutWriter::new();
assert!(writer.is_empty());
assert_eq!(writer.bytes_written(), 0);
assert!(writer.capacity() >= 8 * 1024);
}
#[tokio::test]
async fn test_write_buffered() {
let mut writer = BytesMutWriter::new();
let data = Bytes::from("hello world");
let written = writer.write_buffered(data).await.unwrap();
assert_eq!(written, 11);
assert_eq!(writer.bytes_written(), 11);
assert_eq!(writer.as_slice(), b"hello world");
}
#[tokio::test]
async fn test_preferred_writer_alias() {
let mut writer = BytesMutWriter::new();
let written = writer
.write_buffered(Bytes::from("hello world"))
.await
.expect("write bytes through alias");
assert_eq!(written, 11);
assert_eq!(writer.as_slice(), b"hello world");
}
#[tokio::test]
async fn test_write_slice() {
let mut writer = BytesMutWriter::new();
let data = b"hello world";
let written = writer.write_slice(data).await.unwrap();
assert_eq!(written, 11);
assert_eq!(writer.bytes_written(), 11);
assert_eq!(writer.as_slice(), b"hello world");
}
#[tokio::test]
async fn test_into_bytes() {
let mut writer = BytesMutWriter::new();
let data = Bytes::from("hello world");
writer.write_buffered(data).await.unwrap();
let result = writer.into_bytes();
assert_eq!(result.as_ref(), b"hello world");
}
#[tokio::test]
async fn test_write_after_finalize() {
let mut writer = BytesMutWriter::new();
let data = Bytes::from("hello");
writer.write_buffered(data).await.unwrap();
let _result = writer.into_bytes();
// Create new writer and try to write after finalize
let mut writer2 = BytesMutWriter::new();
writer2.write_buffered(Bytes::from("test")).await.unwrap();
let _ = writer2.into_bytes();
// Writing to a consumed writer should work via new writer
let mut writer3 = BytesMutWriter::new();
let result = writer3.write_buffered(Bytes::from("final")).await;
assert!(result.is_ok());
}
#[tokio::test]
async fn test_clear() {
let mut writer = BytesMutWriter::new();
writer.write_slice(b"hello").await.unwrap();
writer.clear();
assert!(writer.is_empty());
assert_eq!(writer.bytes_written(), 0);
// Capacity should remain
assert!(writer.capacity() > 0);
}
#[tokio::test]
async fn test_reserve() {
let mut writer = BytesMutWriter::with_capacity(10);
let initial_capacity = writer.capacity();
writer.reserve(1000);
// Reserve ensures at least the additional capacity can be added
// but may allocate more than requested
assert!(writer.capacity() >= initial_capacity);
}
#[tokio::test]
async fn test_multiple_writes() {
let mut writer = BytesMutWriter::new();
writer.write_buffered(Bytes::from("hello ")).await.unwrap();
writer.write_slice(b"world").await.unwrap();
assert_eq!(writer.as_slice(), b"hello world");
assert_eq!(writer.bytes_written(), 11);
}
#[tokio::test]
async fn test_async_write() {
use tokio::io::AsyncWriteExt;
let mut writer = BytesMutWriter::new();
let data = b"hello world";
let written = writer.write(data).await.unwrap();
assert_eq!(written, 11);
assert_eq!(writer.as_slice(), b"hello world");
}
#[tokio::test]
async fn test_debug() {
let writer = BytesMutWriter::new();
let debug_str = format!("{:?}", writer);
assert!(debug_str.contains("BytesMutWriter"));
assert!(debug_str.contains("buffer_len"));
}
#[tokio::test]
#[allow(deprecated)]
async fn test_legacy_writer_alias() {
let mut writer = ZeroCopyObjectWriter::new();
let written = writer.write_zero_copy(Bytes::from("hello")).await.unwrap();
assert_eq!(written, 5);
assert_eq!(writer.as_slice(), b"hello");
}
}
+4 -24
View File
@@ -1729,30 +1729,10 @@ mod tests {
assert!(config.validate().is_ok(), "deprecated mount_path must not be required"); assert!(config.validate().is_ok(), "deprecated mount_path must not be required");
} }
#[test] // The "VaultKv2 must not claim Transit wrapping" documentation-claim
fn test_vault_kv2_sources_do_not_claim_transit_wrapping() { // invariant is enforced by scripts/check_fips_wording.sh, which scans every
let sources = [ // file in crates/kms rather than a fixed include_str! list
("config.rs", include_str!("config.rs")), // (rustfs/backlog#1884).
("api_types.rs", include_str!("api_types.rs")),
("backends/vault.rs", include_str!("backends/vault.rs")),
("lib.rs", include_str!("lib.rs")),
];
// Assemble the needles at runtime so this guard does not match its own source.
let needles = [
format!("wrapping via {}", "Transit"),
format!("KV v2 + {}", "Transit"),
format!("KV2+{}", "Transit"),
format!("you would use Vault's {} engine", "transit"),
];
for (name, source) in sources {
for needle in &needles {
assert!(
!source.contains(needle.as_str()),
"{name} still describes the Vault KV2 backend with `{needle}`"
);
}
}
}
#[test] #[test]
fn test_legacy_persisted_vault_transit_config_uses_metadata_defaults() { fn test_legacy_persisted_vault_transit_config_uses_metadata_defaults() {
+282 -69
View File
@@ -15,8 +15,10 @@
use std::collections::HashMap; use std::collections::HashMap;
use std::hash::{Hash, Hasher}; use std::hash::{Hash, Hasher};
use std::sync::Arc; use std::sync::Arc;
use std::sync::atomic::{AtomicBool, Ordering};
use std::time::{Duration, SystemTime}; use std::time::{Duration, SystemTime};
use tokio::sync::RwLock; use tokio::sync::RwLock;
use tokio::time::Instant;
use crate::{ use crate::{
FastLockGuard, GlobalLockManager, LockClient, LockId, LockInfo, LockManager, LockMetadata, LockPriority, LockRequest, FastLockGuard, GlobalLockManager, LockClient, LockId, LockInfo, LockManager, LockMetadata, LockPriority, LockRequest,
@@ -26,43 +28,51 @@ use crate::{
/// Default shard count for guard storage (must be power of 2) /// Default shard count for guard storage (must be power of 2)
const DEFAULT_GUARD_SHARD_COUNT: usize = 64; const DEFAULT_GUARD_SHARD_COUNT: usize = 64;
type GuardShard = Arc<RwLock<HashMap<LockId, LocalGuardEntry>>>;
type GuardStorage = Arc<Vec<GuardShard>>;
/// Local lock client using FastLock with sharded guard storage for better concurrency /// Local lock client using FastLock with sharded guard storage for better concurrency
#[derive(Debug)] #[derive(Debug)]
pub struct LocalClient { pub struct LocalClient {
/// Sharded guard storage to reduce lock contention /// Sharded guard storage to reduce lock contention
guard_storage: Vec<Arc<RwLock<HashMap<LockId, LocalGuardEntry>>>>, guard_storage: GuardStorage,
/// Mask for fast shard index calculation (shard_count - 1) /// Mask for fast shard index calculation (shard_count - 1)
shard_mask: usize, shard_mask: usize,
/// Optional lock manager (if None, uses global singleton) /// Optional lock manager (if None, uses global singleton)
manager: Option<Arc<GlobalLockManager>>, manager: Option<Arc<GlobalLockManager>>,
reaper_started: AtomicBool,
reaper_interval: Duration,
} }
#[derive(Debug)] #[derive(Debug)]
struct LocalGuardEntry { struct LocalGuardEntry {
guard: FastLockGuard, guard: FastLockGuard,
expires_at: SystemTime, expires_at: SystemTime,
deadline: Instant,
ttl: Duration, ttl: Duration,
/// Owner recorded at acquire time; used only for reclaim diagnostics (#899).
owner: String,
} }
impl LocalGuardEntry { impl LocalGuardEntry {
fn new(guard: FastLockGuard, ttl: Duration, owner: String) -> Self { fn new(guard: FastLockGuard, ttl: Duration) -> Self {
let now = SystemTime::now(); let now = SystemTime::now();
let monotonic_now = Instant::now();
Self { Self {
guard, guard,
expires_at: now + ttl, expires_at: now.checked_add(ttl).unwrap_or(now),
deadline: monotonic_now.checked_add(ttl).unwrap_or(monotonic_now),
ttl, ttl,
owner,
} }
} }
fn is_expired(&self) -> bool { fn is_expired(&self) -> bool {
self.expires_at <= SystemTime::now() self.deadline <= Instant::now()
} }
fn refresh(&mut self) { fn refresh(&mut self) {
self.expires_at = SystemTime::now() + self.ttl; let now = SystemTime::now();
let monotonic_now = Instant::now();
self.expires_at = now.checked_add(self.ttl).unwrap_or(now);
self.deadline = monotonic_now.checked_add(self.ttl).unwrap_or(monotonic_now);
} }
} }
@@ -77,26 +87,38 @@ impl LocalClient {
pub fn with_shard_count(shard_count: usize) -> Self { pub fn with_shard_count(shard_count: usize) -> Self {
assert!(shard_count.is_power_of_two(), "Shard count must be power of 2"); assert!(shard_count.is_power_of_two(), "Shard count must be power of 2");
let guard_storage: Vec<Arc<RwLock<HashMap<LockId, LocalGuardEntry>>>> = let guard_storage: Vec<GuardShard> = (0..shard_count).map(|_| Arc::new(RwLock::new(HashMap::new()))).collect();
(0..shard_count).map(|_| Arc::new(RwLock::new(HashMap::new()))).collect();
Self::with_storage(Arc::new(guard_storage), None, crate::fast_lock::CLEANUP_INTERVAL)
}
fn with_storage(guard_storage: GuardStorage, manager: Option<Arc<GlobalLockManager>>, reaper_interval: Duration) -> Self {
let shard_count = guard_storage.len();
debug_assert!(shard_count.is_power_of_two());
Self { Self {
guard_storage, guard_storage,
shard_mask: shard_count - 1, shard_mask: shard_count - 1,
manager: None, manager,
reaper_started: AtomicBool::new(false),
reaper_interval,
} }
} }
/// Create new local client with a specific lock manager /// Create new local client with a specific lock manager
/// This allows simulating multi-node environments where each node has its own lock backend /// This allows simulating multi-node environments where each node has its own lock backend
pub fn with_manager(manager: Arc<GlobalLockManager>) -> Self { pub fn with_manager(manager: Arc<GlobalLockManager>) -> Self {
Self { let guard_storage = (0..DEFAULT_GUARD_SHARD_COUNT)
guard_storage: (0..DEFAULT_GUARD_SHARD_COUNT) .map(|_| Arc::new(RwLock::new(HashMap::new())))
.map(|_| Arc::new(RwLock::new(HashMap::new()))) .collect();
.collect(), Self::with_storage(Arc::new(guard_storage), Some(manager), crate::fast_lock::CLEANUP_INTERVAL)
shard_mask: DEFAULT_GUARD_SHARD_COUNT - 1, }
manager: Some(manager),
} #[cfg(test)]
pub(crate) fn with_manager_and_reaper_interval(manager: Arc<GlobalLockManager>, reaper_interval: Duration) -> Self {
let guard_storage = (0..DEFAULT_GUARD_SHARD_COUNT)
.map(|_| Arc::new(RwLock::new(HashMap::new())))
.collect();
Self::with_storage(Arc::new(guard_storage), Some(manager), reaper_interval)
} }
/// Get the lock manager (injected manager if available, otherwise global singleton) /// Get the lock manager (injected manager if available, otherwise global singleton)
@@ -118,52 +140,63 @@ impl LocalClient {
} }
async fn reclaim_expired_guards_for_resource(&self, resource: &crate::ObjectKey) -> usize { async fn reclaim_expired_guards_for_resource(&self, resource: &crate::ObjectKey) -> usize {
let mut reclaimed = 0usize; let expired_entries = Self::extract_expired_guards(&self.guard_storage, Some(resource)).await;
Self::release_reclaimed_guards(expired_entries, Some(resource))
}
for shard in &self.guard_storage { async fn extract_expired_guards(storage: &GuardStorage, resource: Option<&crate::ObjectKey>) -> Vec<LocalGuardEntry> {
let expired_entries = { let mut expired_entries = Vec::new();
let mut guards = shard.write().await; for shard in storage.iter() {
let mut retained = HashMap::with_capacity(guards.len()); let mut guards = shard.write().await;
let mut expired_entries = Vec::new(); expired_entries.extend(
guards
.extract_if(|lock_id, entry| {
resource.is_none_or(|resource| &lock_id.resource == resource) && entry.is_expired()
})
.map(|(_, entry)| entry),
);
}
expired_entries
}
for (lock_id, entry) in std::mem::take(&mut *guards) { fn release_reclaimed_guards(
if &lock_id.resource == resource && entry.is_expired() { entries: impl IntoIterator<Item = LocalGuardEntry>,
expired_entries.push(entry); resource: Option<&crate::ObjectKey>,
} else { ) -> usize {
retained.insert(lock_id, entry); let mut reclaimed = 0;
} for mut entry in entries {
} let _ = entry.guard.release();
rustfs_io_metrics::record_lock_reclaimed();
*guards = retained; reclaimed += 1;
expired_entries }
}; if reclaimed > 0 {
if let Some(resource) = resource {
for mut entry in expired_entries { tracing::debug!(event = "lock_guard_reclaimed", resource = %resource, count = reclaimed, "expired lock guards reclaimed");
// An expired entry whose owner never refreshed it (a dead coordinator, #698) is } else {
// reclaimed so a live contender can re-form quorum. With guard heartbeats in place tracing::debug!(event = "lock_guard_reaper_sweep", count = reclaimed, "expired lock guards reclaimed");
// (#899) a live owner keeps its entry from expiring, so reaching here means the
// lease genuinely lapsed. Surface it for observability; the reclaim decision itself
// is unchanged.
let since_last_refresh = entry
.expires_at
.checked_sub(entry.ttl)
.and_then(|last_refresh| SystemTime::now().duration_since(last_refresh).ok())
.unwrap_or(entry.ttl);
tracing::warn!(
owner = %entry.owner,
resource = %resource,
ttl_ms = entry.ttl.as_millis() as u64,
since_last_refresh_ms = since_last_refresh.as_millis() as u64,
"reclaiming expired lock guard whose lease was not refreshed"
);
rustfs_io_metrics::record_lock_reclaimed();
let _ = entry.guard.release();
reclaimed = reclaimed.saturating_add(1);
} }
} }
reclaimed reclaimed
} }
fn ensure_reaper(&self) {
if self.reaper_started.swap(true, Ordering::AcqRel) {
return;
}
let storage = Arc::downgrade(&self.guard_storage);
let interval = self.reaper_interval;
tokio::spawn(async move {
let mut ticker = tokio::time::interval(interval);
loop {
ticker.tick().await;
let Some(storage) = storage.upgrade() else {
break;
};
let expired_entries = Self::extract_expired_guards(&storage, None).await;
Self::release_reclaimed_guards(expired_entries, None);
}
});
}
} }
impl Default for LocalClient { impl Default for LocalClient {
@@ -175,28 +208,36 @@ impl Default for LocalClient {
#[async_trait::async_trait] #[async_trait::async_trait]
impl LockClient for LocalClient { impl LockClient for LocalClient {
async fn acquire_lock(&self, request: &LockRequest) -> Result<LockResponse> { async fn acquire_lock(&self, request: &LockRequest) -> Result<LockResponse> {
self.ensure_reaper();
let lock_manager = self.get_lock_manager(); let lock_manager = self.get_lock_manager();
let reclaimed_before_acquire = self.reclaim_expired_guards_for_resource(&request.resource).await; let reclaimed_before_acquire = self.reclaim_expired_guards_for_resource(&request.resource).await;
let acquire_deadline = Instant::now()
.checked_add(request.acquire_timeout)
.unwrap_or_else(Instant::now);
let build_lock_request = || match request.lock_type { let build_lock_request = |acquire_timeout| match request.lock_type {
LockType::Exclusive => crate::ObjectLockRequest::new_write(request.resource.clone(), request.owner.clone()) LockType::Exclusive => crate::ObjectLockRequest::new_write(request.resource.clone(), request.owner.clone())
.with_acquire_timeout(request.acquire_timeout), .with_acquire_timeout(acquire_timeout),
LockType::Shared => crate::ObjectLockRequest::new_read(request.resource.clone(), request.owner.clone()) LockType::Shared => crate::ObjectLockRequest::new_read(request.resource.clone(), request.owner.clone())
.with_acquire_timeout(request.acquire_timeout), .with_acquire_timeout(acquire_timeout),
}; };
let mut retried_after_reclaim = reclaimed_before_acquire > 0; let mut retried_after_reclaim = reclaimed_before_acquire > 0;
loop { loop {
match lock_manager.acquire_lock(build_lock_request()).await { let remaining = acquire_deadline.saturating_duration_since(Instant::now());
if remaining.is_zero() {
return Ok(LockResponse::failure("Lock acquisition timeout", request.acquire_timeout));
}
match lock_manager.acquire_lock(build_lock_request(remaining)).await {
Ok(guard) => { Ok(guard) => {
let lock_id = request.lock_id.clone(); let lock_id = request.lock_id.clone();
let acquired_at = SystemTime::now(); let acquired_at = SystemTime::now();
let expires_at = acquired_at + request.ttl; let expires_at = acquired_at.checked_add(request.ttl).unwrap_or(acquired_at);
{ {
let shard = self.get_shard(&lock_id); let shard = self.get_shard(&lock_id);
let mut guards = shard.write().await; let mut guards = shard.write().await;
guards.insert(lock_id.clone(), LocalGuardEntry::new(guard, request.ttl, request.owner.clone())); guards.insert(lock_id.clone(), LocalGuardEntry::new(guard, request.ttl));
} }
let lock_info = LockInfo { let lock_info = LockInfo {
@@ -256,12 +297,24 @@ impl LockClient for LocalClient {
async fn refresh(&self, lock_id: &LockId) -> Result<bool> { async fn refresh(&self, lock_id: &LockId) -> Result<bool> {
let shard = self.get_shard(lock_id); let shard = self.get_shard(lock_id);
let mut guards = shard.write().await; let expired_entry = {
if let Some(entry) = guards.get_mut(lock_id) { let mut guards = shard.write().await;
entry.refresh(); let Some(entry) = guards.get_mut(lock_id) else {
Ok(true) return Ok(false);
} else { };
if entry.is_expired() {
guards.remove(lock_id)
} else {
entry.refresh();
None
}
};
if let Some(entry) = expired_entry {
Self::release_reclaimed_guards([entry], Some(&lock_id.resource));
Ok(false) Ok(false)
} else {
Ok(true)
} }
} }
@@ -317,3 +370,163 @@ impl LockClient for LocalClient {
true true
} }
} }
#[cfg(test)]
mod tests {
use super::*;
use crate::{GlobalLockManager, LockClient, LockRequest, LockType};
fn request(resource: crate::ObjectKey, owner: &str, ttl: Duration) -> LockRequest {
LockRequest::new(resource, LockType::Exclusive, owner)
.with_ttl(ttl)
.with_acquire_timeout(Duration::from_millis(80))
}
async fn wait_until_reaped(client: &LocalClient, lock_id: &LockId) {
for _ in 0..80 {
if client.check_status(lock_id).await.unwrap().is_none() {
return;
}
tokio::time::sleep(Duration::from_millis(5)).await;
}
panic!("lock guard was not reaped before test deadline");
}
#[tokio::test(flavor = "current_thread")]
async fn expired_guard_is_reaped_without_resource_reacquire() {
let manager = Arc::new(GlobalLockManager::new());
let client = LocalClient::with_manager_and_reaper_interval(manager.clone(), Duration::from_millis(5));
let request = request(crate::ObjectKey::new("bucket", "unique-chunk"), "owner-a", Duration::from_millis(10));
let lock_id = request.lock_id.clone();
assert!(client.acquire_lock(&request).await.unwrap().success);
assert!(client.check_status(&lock_id).await.unwrap().is_some());
tokio::time::sleep(Duration::from_millis(15)).await;
wait_until_reaped(&client, &lock_id).await;
let direct = manager
.acquire_lock(crate::ObjectLockRequest::new_write(request.resource.clone(), "owner-b"))
.await;
assert!(direct.is_ok());
}
#[tokio::test(flavor = "current_thread")]
async fn sibling_client_cannot_reclaim_but_owner_reaper_releases_shared_lock() {
let manager = Arc::new(GlobalLockManager::new());
let owner = LocalClient::with_manager_and_reaper_interval(manager.clone(), Duration::from_millis(5));
let contender = LocalClient::with_manager_and_reaper_interval(manager, Duration::from_millis(5));
let request_a = request(crate::ObjectKey::new("bucket", "shared-resource"), "owner-a", Duration::from_millis(10));
assert!(owner.acquire_lock(&request_a).await.unwrap().success);
let request_b = request(request_a.resource.clone(), "owner-b", Duration::from_millis(20))
.with_acquire_timeout(Duration::from_millis(5));
assert!(!contender.acquire_lock(&request_b).await.unwrap().success);
tokio::time::sleep(Duration::from_millis(25)).await;
assert!(owner.check_status(&request_a.lock_id).await.unwrap().is_none());
assert!(contender.acquire_lock(&request_b).await.unwrap().success);
}
#[tokio::test(flavor = "current_thread")]
async fn refresh_wins_before_deadline_and_reaper_wins_after_deadline() {
let manager = Arc::new(GlobalLockManager::new());
let client = LocalClient::with_manager_and_reaper_interval(manager, Duration::from_millis(5));
let request = request(crate::ObjectKey::new("bucket", "refresh-race"), "owner-a", Duration::from_millis(25));
let lock_id = request.lock_id.clone();
assert!(client.acquire_lock(&request).await.unwrap().success);
tokio::time::sleep(Duration::from_millis(10)).await;
assert!(client.refresh(&lock_id).await.unwrap());
tokio::time::sleep(Duration::from_millis(15)).await;
assert!(client.check_status(&lock_id).await.unwrap().is_some());
wait_until_reaped(&client, &lock_id).await;
}
#[tokio::test(start_paused = true)]
async fn refresh_after_expiry_releases_guard_without_reviving_it() {
let manager = Arc::new(GlobalLockManager::new());
let client = LocalClient::with_manager_and_reaper_interval(manager, Duration::from_secs(60));
client.reaper_started.store(true, Ordering::Release);
let lock_request = request(
crate::ObjectKey::new("bucket", "refresh-after-expiry"),
"owner-a",
Duration::from_secs(10),
);
let lock_id = lock_request.lock_id.clone();
assert!(
client
.acquire_lock(&lock_request)
.await
.expect("initial owner should acquire the lock")
.success
);
tokio::time::advance(Duration::from_secs(11)).await;
assert!(
!client
.refresh(&lock_id)
.await
.expect("expired refresh should return a result"),
"an expired guard must not be refreshed"
);
assert!(
client
.check_status(&lock_id)
.await
.expect("expired guard status should be readable")
.is_none(),
"expired guard should be removed after refresh"
);
let contender = request(
crate::ObjectKey::new("bucket", "refresh-after-expiry"),
"owner-b",
Duration::from_secs(10),
);
assert!(
client
.acquire_lock(&contender)
.await
.expect("contender should receive an acquisition result")
.success,
"released guard must be acquirable by a new owner"
);
}
#[tokio::test(flavor = "current_thread")]
async fn zero_ttl_is_reaped_and_oversized_ttl_does_not_panic() {
let manager = Arc::new(GlobalLockManager::new());
let client = LocalClient::with_manager_and_reaper_interval(manager, Duration::from_millis(5));
let zero = request(crate::ObjectKey::new("bucket", "zero-ttl"), "owner-zero", Duration::ZERO);
let zero_id = zero.lock_id.clone();
assert!(client.acquire_lock(&zero).await.unwrap().success);
wait_until_reaped(&client, &zero_id).await;
let huge = request(crate::ObjectKey::new("bucket", "huge-ttl"), "owner-huge", Duration::MAX);
let huge_id = huge.lock_id.clone();
assert!(client.acquire_lock(&huge).await.unwrap().success);
wait_until_reaped(&client, &huge_id).await;
}
#[tokio::test(flavor = "current_thread")]
async fn acquire_retry_preserves_total_deadline() {
let manager = Arc::new(GlobalLockManager::new());
let client = LocalClient::with_manager_and_reaper_interval(manager, Duration::from_secs(60));
let first = request(crate::ObjectKey::new("bucket", "deadline-budget"), "owner-a", Duration::from_millis(10));
assert!(client.acquire_lock(&first).await.unwrap().success);
let second =
request(first.resource.clone(), "owner-b", Duration::from_millis(30)).with_acquire_timeout(Duration::from_millis(60));
let started = Instant::now();
let response = client.acquire_lock(&second).await.unwrap();
assert!(!response.success, "the first attempt consumed the caller's acquire budget");
assert!(
started.elapsed() < Duration::from_millis(100),
"reclaim retry must not double the acquire budget"
);
let recovered = client.acquire_lock(&second).await.unwrap();
assert!(recovered.success, "the reclaimed guard must be available to the next request");
}
}
+111
View File
@@ -840,6 +840,117 @@ async fn test_namespace_lock_distributed_reclaims_expired_same_resource_after_fa
); );
} }
#[tokio::test]
async fn four_node_failed_release_converges_without_replica_repair() {
let managers = (0..4).map(|_| Arc::new(GlobalLockManager::new())).collect::<Vec<_>>();
let flaky_clients = managers
.iter()
.map(|manager| {
Arc::new(FlakyReleaseClient {
inner: LocalClient::with_manager_and_reaper_interval(manager.clone(), Duration::from_millis(5)),
failed_releases_remaining: AtomicUsize::new(usize::MAX),
release_attempts: AtomicUsize::new(0),
})
})
.collect::<Vec<_>>();
let clients = flaky_clients
.iter()
.map(|client| client.clone() as Arc<dyn LockClient>)
.collect::<Vec<_>>();
let lock = NamespaceLock::Distributed(DistributedLock::new("four-node-expired-lease".to_string(), clients, 3));
let resource = create_test_object_key("bucket", "object-four-node-expired");
let request = LockRequest::new(resource.clone(), LockType::Exclusive, "owner-a")
.with_acquire_timeout(Duration::from_millis(300))
.with_ttl(Duration::from_millis(40));
let mut guard = lock
.acquire_guard(&request)
.await
.expect("initial acquire should not error")
.expect("initial acquire should reach quorum");
assert!(guard.release(), "release should be acknowledged while RPC cleanup is pending");
for _ in 0..40 {
if flaky_clients.iter().all(|client| client.release_attempts() >= 3) {
break;
}
tokio::time::sleep(Duration::from_millis(5)).await;
}
let deadline = tokio::time::Instant::now() + Duration::from_secs(2);
loop {
let all_reaped =
futures::future::join_all(flaky_clients.iter().map(|client| client.inner.check_status(&request.lock_id)))
.await
.into_iter()
.all(|status| status.expect("status should not error").is_none());
if all_reaped {
break;
}
assert!(tokio::time::Instant::now() < deadline, "all four local lease entries must converge");
tokio::time::sleep(Duration::from_millis(10)).await;
}
for suffix in ["chunk-0", "chunk-1", ".rustfs.sys/multipart/upload-0"] {
for client in &flaky_clients {
let orphan = LockRequest::new(create_test_object_key("bucket", suffix), LockType::Exclusive, "orphan")
.with_ttl(Duration::from_millis(25));
assert!(client.inner.acquire_lock(&orphan).await.expect("orphan acquire").success);
}
}
tokio::time::sleep(Duration::from_millis(80)).await;
let recovered = lock
.acquire_guard(
&LockRequest::new(resource, LockType::Exclusive, "owner-b")
.with_acquire_timeout(Duration::from_millis(300))
.with_ttl(Duration::from_millis(40)),
)
.await
.expect("recovery acquire should not error")
.expect("four-node quorum should recover after local reapers run");
drop(recovered);
}
#[tokio::test]
async fn four_node_stale_quorum_contention_respects_acquire_deadline() {
let managers = (0..4).map(|_| Arc::new(GlobalLockManager::new())).collect::<Vec<_>>();
let node_clients = managers
.iter()
.map(|manager| Arc::new(LocalClient::with_manager_and_reaper_interval(manager.clone(), Duration::from_millis(5))))
.collect::<Vec<_>>();
let resource = create_test_object_key("bucket", "stale-quorum");
let stale = LockRequest::new(resource.clone(), LockType::Exclusive, "stale-owner").with_ttl(Duration::from_millis(180));
for client in &node_clients {
assert!(client.acquire_lock(&stale).await.expect("stale acquire").success);
}
let clients = node_clients
.iter()
.map(|client| client.clone() as Arc<dyn LockClient>)
.collect::<Vec<_>>();
let lock = NamespaceLock::Distributed(DistributedLock::new("stale-quorum-deadline".to_string(), clients, 3));
let contender = LockRequest::new(resource.clone(), LockType::Exclusive, "new-owner")
.with_acquire_timeout(Duration::from_millis(150))
.with_ttl(Duration::from_millis(100));
let started = tokio::time::Instant::now();
let response = lock.acquire_guard(&contender).await.expect("contention should not error");
assert!(response.is_none(), "unexpired leases must not be force-reclaimed");
assert!(started.elapsed() < Duration::from_millis(350), "acquire must respect its deadline");
tokio::time::sleep(Duration::from_millis(80)).await;
let recovered = lock
.acquire_guard(
&LockRequest::new(resource, LockType::Exclusive, "new-owner")
.with_acquire_timeout(Duration::from_millis(300))
.with_ttl(Duration::from_millis(100)),
)
.await
.expect("post-expiry acquire should not error")
.expect("quorum should recover after local reapers clear stale leases");
drop(recovered);
}
#[tokio::test] #[tokio::test]
async fn test_namespace_lock_distributed_retries_transient_acquire_timeout() { async fn test_namespace_lock_distributed_retries_transient_acquire_timeout() {
let managers = (0..3).map(|_| Arc::new(GlobalLockManager::new())).collect::<Vec<_>>(); let managers = (0..3).map(|_| Arc::new(GlobalLockManager::new())).collect::<Vec<_>>();
+9 -1
View File
@@ -43,7 +43,7 @@ pub struct ServiceTraceOpts {
#[allow(dead_code)] #[allow(dead_code)]
impl ServiceTraceOpts { impl ServiceTraceOpts {
fn trace_types(&self) -> TraceType { pub fn trace_types(&self) -> TraceType {
let mut tt = TraceType::default(); let mut tt = TraceType::default();
tt.set_if(self.s3, &TraceType::S3); tt.set_if(self.s3, &TraceType::S3);
tt.set_if(self.internal, &TraceType::INTERNAL); tt.set_if(self.internal, &TraceType::INTERNAL);
@@ -72,6 +72,14 @@ impl ServiceTraceOpts {
tt tt
} }
pub fn only_errors(&self) -> bool {
self.only_errors
}
pub fn threshold(&self) -> Duration {
self.threshold
}
pub fn parse_params(&mut self, uri: &Uri) -> Result<(), String> { pub fn parse_params(&mut self, uri: &Uri) -> Result<(), String> {
let query_pairs: HashMap<_, _> = uri let query_pairs: HashMap<_, _> = uri
.query() .query()
@@ -427,7 +427,6 @@ pub enum DataSource {
/// Write triggered /// Write triggered
WriteTriggered, WriteTriggered,
/// Fallback value /// Fallback value
#[allow(dead_code)]
Fallback, Fallback,
} }
@@ -603,7 +602,6 @@ impl WriteRecord {
/// Hybrid strategy configuration /// Hybrid strategy configuration
#[derive(Debug, Clone)] #[derive(Debug, Clone)]
#[allow(dead_code)]
pub struct HybridStrategyConfig { pub struct HybridStrategyConfig {
/// Scheduled update interval /// Scheduled update interval
pub scheduled_update_interval: Duration, pub scheduled_update_interval: Duration,
@@ -998,14 +996,12 @@ impl HybridCapacityManager {
} }
/// Get cache age /// Get cache age
#[allow(dead_code)]
pub async fn get_cache_age(&self) -> Option<Duration> { pub async fn get_cache_age(&self) -> Option<Duration> {
let cache = self.cache.read().await; let cache = self.cache.read().await;
cache.as_ref().map(|c| c.last_update.elapsed()) cache.as_ref().map(|c| c.last_update.elapsed())
} }
/// Get write frequency (writes/minute) /// Get write frequency (writes/minute)
#[allow(dead_code)]
pub async fn get_write_frequency(&self) -> usize { pub async fn get_write_frequency(&self) -> usize {
let record = &self.write_record; let record = &self.write_record;
record.recent_write_count(record.monotonic_second()) record.recent_write_count(record.monotonic_second())
@@ -1300,7 +1296,6 @@ pub fn get_capacity_manager() -> Arc<HybridCapacityManager> {
/// .update_capacity(CapacityUpdate::exact(1000, 0), DataSource::RealTime) /// .update_capacity(CapacityUpdate::exact(1000, 0), DataSource::RealTime)
/// .await; /// .await;
/// ``` /// ```
#[allow(dead_code)]
pub fn create_isolated_manager(config: HybridStrategyConfig) -> Arc<HybridCapacityManager> { pub fn create_isolated_manager(config: HybridStrategyConfig) -> Arc<HybridCapacityManager> {
Arc::new(HybridCapacityManager::new(config)) Arc::new(HybridCapacityManager::new(config))
} }
+66 -10
View File
@@ -177,16 +177,38 @@ impl StartCommand {
#[derive(Debug, Clone, PartialEq, Eq, Serialize, Deserialize)] #[derive(Debug, Clone, PartialEq, Eq, Serialize, Deserialize)]
#[serde(tag = "action", rename_all = "snake_case", deny_unknown_fields)] #[serde(tag = "action", rename_all = "snake_case", deny_unknown_fields)]
pub enum Command { pub enum Command {
Start { request: StartCommand }, Start {
Query { heal_path: String, client_token: String }, request: StartCommand,
Cancel { heal_path: String, client_token: String }, },
Query {
heal_path: String,
client_token: String,
/// Incremental result cursor (HS-06): only items with a sequence
/// greater than this are returned. Absent = legacy full snapshot.
/// Optional + defaulted so older peers stay wire-compatible.
#[serde(default, skip_serializing_if = "Option::is_none")]
since_seq: Option<u64>,
},
Cancel {
heal_path: String,
client_token: String,
},
} }
#[derive(Debug)] #[derive(Debug)]
pub enum ExecutableCommand { pub enum ExecutableCommand {
Start { request: HealChannelRequest }, Start {
Query { heal_path: String, client_token: String }, request: HealChannelRequest,
Cancel { heal_path: String, client_token: String }, },
Query {
heal_path: String,
client_token: String,
since_seq: Option<u64>,
},
Cancel {
heal_path: String,
client_token: String,
},
} }
#[derive(Debug, Clone, PartialEq, Eq, Serialize, Deserialize)] #[derive(Debug, Clone, PartialEq, Eq, Serialize, Deserialize)]
@@ -227,8 +249,22 @@ impl Envelope {
) )
} }
pub fn query(request_id: String, metadata: RequestMetadata, heal_path: String, client_token: String) -> Result<Self, String> { pub fn query(
Self::new(request_id, metadata, Command::Query { heal_path, client_token }) request_id: String,
metadata: RequestMetadata,
heal_path: String,
client_token: String,
since_seq: Option<u64>,
) -> Result<Self, String> {
Self::new(
request_id,
metadata,
Command::Query {
heal_path,
client_token,
since_seq,
},
)
} }
pub fn cancel( pub fn cancel(
@@ -286,7 +322,15 @@ impl Envelope {
Command::Start { request } => ExecutableCommand::Start { Command::Start { request } => ExecutableCommand::Start {
request: request.into_channel_request(self.request_id.clone())?, request: request.into_channel_request(self.request_id.clone())?,
}, },
Command::Query { heal_path, client_token } => ExecutableCommand::Query { heal_path, client_token }, Command::Query {
heal_path,
client_token,
since_seq,
} => ExecutableCommand::Query {
heal_path,
client_token,
since_seq,
},
Command::Cancel { heal_path, client_token } => ExecutableCommand::Cancel { heal_path, client_token }, Command::Cancel { heal_path, client_token } => ExecutableCommand::Cancel { heal_path, client_token },
}; };
Ok((self.request_id, self.coordinator_epoch, command)) Ok((self.request_id, self.coordinator_epoch, command))
@@ -305,6 +349,11 @@ pub enum Admission {
Full, Full,
DroppedQueueFull, DroppedQueueFull,
DroppedPolicy, DroppedPolicy,
/// HS-06: admin start rejected because the same target is already being
/// healed (RUSTFS_HEAL_OVERLAP_POLICY=minio_error only).
DroppedAlreadyRunning,
/// HS-06: admin start rejected because its path overlaps an active heal.
DroppedOverlappingPaths,
} }
impl From<HealAdmissionResult> for Admission { impl From<HealAdmissionResult> for Admission {
@@ -315,6 +364,8 @@ impl From<HealAdmissionResult> for Admission {
HealAdmissionResult::Full => Self::Full, HealAdmissionResult::Full => Self::Full,
HealAdmissionResult::Dropped(HealAdmissionDropReason::QueueFull) => Self::DroppedQueueFull, HealAdmissionResult::Dropped(HealAdmissionDropReason::QueueFull) => Self::DroppedQueueFull,
HealAdmissionResult::Dropped(HealAdmissionDropReason::PolicyDropped) => Self::DroppedPolicy, HealAdmissionResult::Dropped(HealAdmissionDropReason::PolicyDropped) => Self::DroppedPolicy,
HealAdmissionResult::Dropped(HealAdmissionDropReason::AlreadyRunning) => Self::DroppedAlreadyRunning,
HealAdmissionResult::Dropped(HealAdmissionDropReason::OverlappingPaths) => Self::DroppedOverlappingPaths,
} }
} }
} }
@@ -331,6 +382,8 @@ impl Admission {
Self::Full => HealAdmissionResult::Full, Self::Full => HealAdmissionResult::Full,
Self::DroppedQueueFull => HealAdmissionResult::Dropped(HealAdmissionDropReason::QueueFull), Self::DroppedQueueFull => HealAdmissionResult::Dropped(HealAdmissionDropReason::QueueFull),
Self::DroppedPolicy => HealAdmissionResult::Dropped(HealAdmissionDropReason::PolicyDropped), Self::DroppedPolicy => HealAdmissionResult::Dropped(HealAdmissionDropReason::PolicyDropped),
Self::DroppedAlreadyRunning => HealAdmissionResult::Dropped(HealAdmissionDropReason::AlreadyRunning),
Self::DroppedOverlappingPaths => HealAdmissionResult::Dropped(HealAdmissionDropReason::OverlappingPaths),
} }
} }
} }
@@ -592,6 +645,7 @@ mod tests {
metadata(2, 7), metadata(2, 7),
"bucket/prefix".to_string(), "bucket/prefix".to_string(),
"token".to_string(), "token".to_string(),
None,
) )
.unwrap(); .unwrap();
let cancel = Envelope::cancel( let cancel = Envelope::cancel(
@@ -667,6 +721,7 @@ mod tests {
RequestMetadata::new([0x11; 16], 1_700_000_000_000, 1_700_000_030_000, 9), RequestMetadata::new([0x11; 16], 1_700_000_000_000, 1_700_000_030_000, 9),
"bucket/prefix".to_string(), "bucket/prefix".to_string(),
"client-token".to_string(), "client-token".to_string(),
None,
) )
.unwrap(); .unwrap();
let cancel = Envelope::cancel( let cancel = Envelope::cancel(
@@ -749,7 +804,7 @@ mod tests {
assert!(Envelope::start(test_request(request_id.clone()), metadata(0, 7)).is_err()); assert!(Envelope::start(test_request(request_id.clone()), metadata(0, 7)).is_err());
assert!(Envelope::start(test_request(request_id.clone()), metadata(1, 0)).is_err()); assert!(Envelope::start(test_request(request_id.clone()), metadata(1, 0)).is_err());
assert!(Envelope::start(test_request(request_id.clone()), RequestMetadata::new([1; 16], 1_000, 31_001, 7),).is_err()); assert!(Envelope::start(test_request(request_id.clone()), RequestMetadata::new([1; 16], 1_000, 31_001, 7),).is_err());
assert!(Envelope::query(request_id.clone(), metadata(1, 7), String::new(), String::new()).is_err()); assert!(Envelope::query(request_id.clone(), metadata(1, 7), String::new(), String::new(), None).is_err());
assert!(Envelope::cancel(request_id.clone(), metadata(1, 7), String::new(), String::new()).is_ok()); assert!(Envelope::cancel(request_id.clone(), metadata(1, 7), String::new(), String::new()).is_ok());
let mut noncanonical_request = test_request(request_id.to_uppercase()); let mut noncanonical_request = test_request(request_id.to_uppercase());
@@ -782,6 +837,7 @@ mod tests {
metadata(1, 7), metadata(1, 7),
"x".repeat(ENVELOPE_MAX_SIZE), "x".repeat(ENVELOPE_MAX_SIZE),
"token".to_string(), "token".to_string(),
None,
) )
.unwrap(); .unwrap();
let error = super::encode_envelope(&oversized).unwrap_err(); let error = super::encode_envelope(&oversized).unwrap_err();
+1 -1
View File
@@ -2557,7 +2557,7 @@ mod tests {
fn production_source(source: &'static str, file_name: &str) -> &'static str { fn production_source(source: &'static str, file_name: &str) -> &'static str {
source source
.split("\n#[cfg(test)]") .split("\n#[cfg(test)]\nmod tests")
.next() .next()
.unwrap_or_else(|| panic!("{file_name} should contain production source before tests")) .unwrap_or_else(|| panic!("{file_name} should contain production source before tests"))
} }

Some files were not shown because too many files have changed in this diff Show More