Compare commits

...

75 Commits

Author SHA1 Message Date
houseme fbec33bd29 Expose target-scoped durable MRF backlog metrics (#5584)
* feat(replication): expose target durable mrf backlog

Add target ARN attribution to durable MRF entries and surface target-scoped durable backlog metrics without changing existing bucket-only metric labels.

Keep legacy MRF files bucket-only by defaulting missing targetARNs to an empty list, and expose target snapshots through an additive API so existing DurableMrfBacklogSummary callers remain source-compatible.

Co-Authored-By: heihutu <heihutu@gmail.com>

* feat(replication): expose runtime target backlog (#5586)

Track runtime replication backlog by target ARN for regular, large, delete, and MRF admission paths while preserving the existing bucket-level backlog semantics.

Add target-scoped current backlog metrics and merge them with durable target backlog snapshots for observability.

Co-authored-by: heihutu <heihutu@gmail.com>

---------

Co-authored-by: heihutu <heihutu@gmail.com>
2026-08-01 17:31:50 +00:00
GatewayJ 4c83bff81c fix(s3select): support two-byte CSV record delimiters (#5565) 2026-08-01 17:08:31 +00:00
GatewayJ 51f9d1b74f fix(s3select): guarantee terminal event delivery (#5563) 2026-08-02 00:53:04 +08:00
houseme ed02899d31 test(perf): cover ABBA evidence binding matrix (#5585)
Co-authored-by: heihutu <heihutu@gmail.com>
2026-08-01 15:58:32 +00:00
Zhengchao An 81fc61db41 feat(admin): expose KMS key description and tag endpoints (#5575)
* feat(kms): add admin endpoints for key description and tag updates

Wire the KMS key metadata updates landed at the service layer to the admin
API: POST /v3/kms/keys/{update-description,tag,untag}. Each endpoint gates on
a dedicated KMS action scoped to the key its body names, records a handler-owned
audit entry for both the authorization denial and the outcome, and maps a
backend that cannot update key metadata to 501 rather than 404.

Refs rustfs/backlog#1586 (part of rustfs/backlog#1562)

* fix(admin): audit metadata attempts refused by an unavailable KMS

* test(admin): register the new KMS metadata routes in the matrix
2026-08-01 15:57:26 +00:00
houseme a08fb56607 test(perf): strengthen HotPath evidence contracts (#5581)
* test(perf): mark nonformal ABBA artifacts

Co-Authored-By: heihutu <heihutu@gmail.com>

* test(io-metrics): lock disabled EC stage accounting

Co-Authored-By: heihutu <heihutu@gmail.com>

---------

Co-authored-by: heihutu <heihutu@gmail.com>
2026-08-01 15:07:17 +00:00
唐小鸭 c8016cbcdb fix(replication): enforce bucket replication switches (#5449)
* fix(replication): enforce bucket replication switches

* fix(replication): satisfy delete admission clippy lint

* fix(replication): restore MinIO tag filter behavior

---------

Co-authored-by: Zhengchao An <anzhengchao@gmail.com>
Co-authored-by: cxymds <cxymds@gmail.com>
2026-08-01 15:03:57 +00:00
Zhengchao An 56a7e3b707 chore(s3-types): guard the EventName mask bit budget (#5582)
EventName::mask() turns a leaf variant's discriminant `v` into
`1 << (v - 1)`, so the enum can hold at most 64 variants in total —
compound "All" variants included, since they consume discriminants
even though they own no bit and push every later leaf further up the
range. The last variant is KmsServiceStopped at 61, leaving 3 slots.

Past the budget, `1u64 << (v - 1)` shifts by 64 or more: debug builds
panic with a shift overflow, release builds silently mask the shift
amount down and hand back a bit that already belongs to another event.
Either failure surfaces far from the line that added the variant.

The existing test_mask_bit_budget_is_not_exhausted only bounds the
events listed in the test-local ALL_EVENT_NAMES, so a variant nobody
remembered to list went unchecked. Add three layers instead:

- A const assertion on LAST_EVENT_NAME_VALUE, anchored on the last
  variant, that fails `cargo build` once the discriminant passes 64.
- An exhaustive-match tripwire next to ALL_EVENT_NAMES, so adding a
  variant breaks compilation with a non-exhaustive-patterns error that
  points at the budget notes.
- Tests pinning that discriminants stay dense and fully listed, that
  every leaf mask is non-zero and owns a unique bit (catching the
  release-mode wrapped-shift collision), and that the last variant
  still lands inside the u64.

Document the budget and the three guards on mask() so the next person
adding an event sees the remaining headroom.

Refs: rustfs/backlog#1572
2026-08-01 14:49:16 +00:00
Zhengchao An f2d09d1426 feat(admin): expose KMS backup and restore behind explicit guards (#5579)
* feat(kms): add backup and restore admin API

Wires the merged KMS backup contract, Local export and Local restore into
the admin API: export a sealed bundle, run a zero-write restore preflight,
execute a confirmed restore, roll an interrupted restore back, and report
subsystem readiness.

- Dedicated kms:Backup / kms:Restore actions, recorded in the admin route
  matrix. Neither is reachable through any other KMS action.
- Restore requires two independent confirmations: an echo of the bundle
  manifest's backup id, and an explicitly named conflict policy (the
  default never writes).
- The backup KEK comes from the environment and is refused when it reuses
  a secret of the configured backend, compared both as the literal value
  and as raw key bytes.
- No endpoint accepts a path: bundles are addressed by a validated name
  under a configured root, and the restore target is always the server's
  own configured key directory.
- Bundles now carry a sanitized configuration artifact built as an
  allowlist projection, so a future backend credential field cannot leak
  into a bundle by default. Restore verifies it and never applies it.
- Audit entries go through the existing KMS admin wiring and carry
  identifiers only.

* test(kms): pin the backup admin API gates

Fixes the test KEK to a real 32-byte value and drives the export refusal
from the configured backend rather than from the handle that happens to
be available, so a Local handle cannot export on behalf of a backend
whose material RustFS does not own.
2026-08-01 14:31:25 +00:00
Zhengchao An 02aa383598 fix(kms): refuse to rotate a Vault key whose baseline version was erased (#5578)
`VaultKeyData::baseline_version` pins the master key version that every
pre-versioning DEK envelope (one with no `master_key_version`) resolves to.
Builds released before versioned rotation do not know the field, and every KV2
lifecycle write — enable, disable, schedule/cancel deletion, tag, untag —
rewrites the whole record, so a single lifecycle call from an older node during
a rolling upgrade silently drops the baseline. Serde attributes cannot prevent
this: the code doing the dropping already shipped.

The loss is only latent until the next rotation. With the baseline gone,
rotation takes the first-rotation path again and freezes a *new* baseline at the
current version, so every legacy envelope permanently resolves to material that
never wrapped it. That is the point of no return, and it is the one this commit
blocks: rotation now lists the key's immutable version records first and refuses
when records exist while the record carries no baseline. Those two are created
by the same commit, so the combination can only mean the baseline was erased
afterwards. The refusal is decided from reads alone, before any write, and names
the version to restore — the oldest recorded version *is* the lost baseline,
since version records start at the baseline the first rotation froze.

Reads are diagnosed rather than blocked. A version-less envelope on a key with
no baseline resolves to the current version, which AES-256-GCM refuses to
unwrap when it is the wrong one, so no wrong plaintext can be returned. Only
after that failure does decrypt list the version records and re-report the
failure as the lost baseline. Refusing up front on the same evidence would break
reads that work today: an older node writes version-less envelopes wrapped with
whatever material is current, and those still decrypt. This also keeps the extra
listing off every read of pre-versioning data.

Self-healing (writing back `baseline_version = min(recorded)`) is deliberately
not done: during the mixed-version window that caused the loss, an older node
can erase it again on the next lifecycle call, so healing would mask an
unfinished upgrade instead of surfacing it. Moving the baseline to a KV path
older builds cannot rewrite is the real fix and is scheduled for GA.

Refs rustfs/backlog#1581 (part of rustfs/backlog#1562)
2026-08-01 14:02:52 +00:00
Zhengchao An 860ad68afd feat(sse): attach KMS attribution to S3 audit entries (#5577)
feat(kms): attach the SSE data-plane KMS summary to S3 audit entries
2026-08-01 13:58:30 +00:00
Zhengchao An 95e96e1bd5 docs(agents): require worktree and disk hygiene (#5580)
* docs(agents): require worktree and disk hygiene

* docs(agents): add PR lifecycle monitoring
2026-08-01 21:52:05 +08:00
GatewayJ 816849a8ee fix(s3select): cancel queries after client disconnect (#5560) 2026-08-01 13:38:30 +00:00
Zhengchao An 5ef5eb8ea9 ci: add nightly GNU build (#5576) 2026-08-01 21:30:21 +08:00
Zhengchao An ad721bff42 fix(kms): honour the configured metadata cache TTL and metrics switch (#5569)
* fix(kms): honour the configured metadata cache TTL and metrics switch

KmsManager::new built the KmsCache from cache_config.max_keys alone, so
cache_config.ttl was dead configuration: every deployment ran the
hardcoded 300s window whatever the admin configure API was given, while
CacheSummary and the KMS config endpoint reported the configured value
back. cache_config.enable_metrics was never read anywhere.

Build the cache from the whole CacheConfig. The documented default is
reconciled down to the 300s the cache has always used rather than up to
the advertised 3600s, and now lives in one place (DEFAULT_CACHE_TTL)
instead of being duplicated across the four configure-request
converters, so the default path behaves exactly as before.

Behaviour change: a deployment configured through the admin API already
has ttl 3600 persisted, because the old converters wrote that default
into the stored config, so its describe_key staleness window widens from
an effective 300s to the 3600s it asked for. No cryptographic or
authorization path widens - encrypt, decrypt and generate_data_key go
straight to the backend and never read this cache. The Vault Transit
backend's own metadata cache, which does gate crypto through
ensure_key_state_allows, stays fixed at 300s and is now documented as
deliberately not operator-tunable.

The configured duration now reaches moka's builder, which panics above
1000 years, so CacheConfig::effective_ttl clamps to a 24h maximum the
way effective_timeout already clamps its own, and validate rejects a
zero TTL beside the existing max_keys check. Both config summaries and
the KMS config endpoint report the effective value, so the admin API
cannot advertise a lifetime the cache does not honour.

enable_metrics gates publication of the rustfs_kms_metadata_cache_*
families only; the counters behind the admin status API keep running
either way. No configure-request field sets it yet.

Refs rustfs/backlog#1584

* docs(kms): state why the Transit metadata TTL is not bound to the default

The comment claimed the constant matches config::DEFAULT_CACHE_TTL, which
reads as an invariant the code does not enforce. Say plainly that the
equality is a coincidence rather than a contract, and why binding the two
would be wrong: this cache gates crypto through ensure_key_state_allows,
so a later change to the operator-facing describe-cache default must not
be able to widen its staleness window.
2026-08-01 13:27:33 +00:00
houseme 29dedcc7fd fix(obs): retire replication backlog metric series (#5574)
Emit zero tombstones for removed bucket-level replication backlog series and retire the corresponding metric descriptors after the configured tombstone window, matching the existing replication bandwidth lifecycle.

Co-authored-by: heihutu <heihutu@gmail.com>
2026-08-01 13:21:41 +00:00
houseme 749cb2c1a1 perf: harden HotPath closure evidence (#5573)
* test(ecstore): cover raw shard write errors

Co-Authored-By: heihutu <heihutu@gmail.com>

* perf(ecstore): track encode payload stage peaks

Co-Authored-By: heihutu <heihutu@gmail.com>

* test(perf): harden formal warp ABBA evidence

Co-Authored-By: heihutu <heihutu@gmail.com>

---------

Co-authored-by: heihutu <heihutu@gmail.com>
Co-authored-by: zhi22915 <qiuzgang@gmail.com>
2026-08-01 21:16:31 +08:00
houseme eb4f2d61ad test(replication): distinguish backlog from failed counters (#5570)
test(replication): distinguish backlog from failures

Extend the replication backlog e2e to assert that historical failed counters can remain non-zero after recovery while current backlog and MRF pending gauges settle to zero.

Co-authored-by: heihutu <heihutu@gmail.com>
Co-authored-by: zhi22915 <qiuzgang@gmail.com>
2026-08-01 12:36:14 +00:00
Zhengchao An 3c00ad6048 fix(kms): make the deletion waiting window non-bypassable (#5535) 2026-08-01 12:23:28 +00:00
Zhengchao An 85bc0d3ce2 ci: let the CARGO_BUILD_JOBS experiment collect samples on demand (#5572)
Gate 2 in rustfs/backlog#1601 needs ten terminal Test and Lint samples at 3.
Push alone cannot supply them: this workflow cancels superseded runs on main,
and only 4 of the last 20 push-triggered Test and Lint jobs reached a terminal
state — 15 were cancelled. At that rate ten samples take roughly fifty merges,
so the experiment would sit open for days while the branch it measures keeps
moving.

The concurrency group is scoped by event_name, so a dispatched run has its own
group and is not cancelled by merge traffic. Including workflow_dispatch in the
raised value makes the sample collectable deliberately rather than by waiting
for pushes to survive.

PRs still get 2. The merge path is untouched.

Refs: rustfs/backlog#1598, rustfs/backlog#1601
2026-08-01 20:19:38 +08:00
Zhengchao An b8d2f1c84c feat(kms): orchestrate and verify Vault-backed restores (#5549) 2026-08-01 12:18:11 +00:00
Zhengchao An 59f69fbfb5 ci: raise CARGO_BUILD_JOBS to 3 on main pushes as a measured experiment (#5571)
The #5394 mitigation set it to 2 on the belief that three concurrent workspace
test links saturate the runner's overlay I/O. The sampler's cgroup v2 readings
show the pod has 14 CPUs and 28GB with a 2.1GB peak, so 2 throttles compilation
to a seventh of what is available and memory was never the constraint. The label
name sm-standard-4 had led everyone, including that mitigation and every
description written during this series, to assume four cores.

Raised to 3 for push events only. PRs keep 2, so the merge path is untouched
while the experiment runs.

Baseline over 17 non-cancelled samples at 2: median nextest/clippy step ratio
1.95, spread 1.85-2.06. Judging by that ratio rather than absolute wall clock is
what makes the signal usable — clippy runs on the same node at the same time and
is check-only for workspace members, so it never links the ~100 test binaries
this limit throttles, making it a control arm rather than a second measurement.

The criterion is the ratio dropping at least 10%, below about 1.76, with no 75m
timeout and no run showing three consecutive samples of rustc, collect2 or
rust-lld in D state. If it does not drop, the conclusion is that this limit is
not the bottleneck: fix it back at 2 and record the experiment. That is a
result, not a failure.

Kept at step level deliberately. rust-cache hashes CARGO/CC/CFLAGS/CXX/CMAKE/RUST
prefixed variables from process.env into its key, so promoting this to job level
would rotate every cache key on this lane.

Refs: rustfs/backlog#1598, rustfs/backlog#1601
2026-08-01 20:00:31 +08:00
Zhengchao An be7ba1bba9 ci: make the cache size report readable from the API (#5568)
The figures went only to $GITHUB_STEP_SUMMARY, which renders in the UI but
never appears in the job log — and the REST API exposes the log, not the
summary. So the one number the cache-all-crates decision rests on could not be
fetched by the script that needed it. Print to stdout as well, between markers,
and keep the summary rendering.
2026-08-01 19:23:42 +08:00
GatewayJ 6363263f09 fix(s3select): parse JSON source paths from SQL AST (#5559) 2026-08-01 11:20:33 +00:00
GatewayJ 533896d045 fix(s3select): honor custom CSV record delimiters (#5558) 2026-08-01 11:18:04 +00:00
Zhengchao An 0bf077f918 ci: stop main pushes from starving dispatched Cache Warm runs (#5567)
Cache Warm used one concurrency group for every trigger. GitHub keeps one
running plus one pending run per group, so a manually dispatched run sat as
pending until the next merge displaced and cancelled it — observed three times
in a row. That made the --timings gate in rustfs/backlog#1601 impossible to
trigger while main was busy, which is precisely when someone wants to measure.

Scoping the group by event lets the push and dispatch paths queue independently.
Neither gains cancel-in-progress, so a burst of merges still collapses into
"current run finishes, newest queued run follows".

The two paths can now overlap and race to save the same key. That is benign:
the loser finds the key already present and skips, and both builds produce the
same artifacts from the same commit.
2026-08-01 19:16:32 +08:00
GatewayJ 284faec03f fix(s3select): preserve ScanRange across file partitions (#5562) 2026-08-01 19:15:51 +08:00
Zhengchao An bfec547d36 feat(health): drive KMS readiness from the probe status (#5552)
* feat(health): drive KMS readiness from the probe status

Readiness reported the KMS ready on the service status bit alone, which says
the manager started, not that the backend can still serve a request. Consume
the background probe snapshot instead: a running service is withdrawn only
when a fresh snapshot shows failures at the threshold.

The readiness path stays free of backend calls — it reads the lock-free
snapshot — so a struggling KMS cannot be amplified into probe traffic of its
own. Everything the probe cannot speak to (no worker, no completed round,
a snapshot older than three probe intervals) leaves the previous status-bit
verdict standing, and the check remains off by default.

Refs rustfs/backlog#1584 (part of rustfs/backlog#1562)

* test(health): pin the readiness default at compile time
2026-08-01 19:15:12 +08:00
Zhengchao An 0bd53becb5 feat(admin): emit KMS management audit events (#5554)
* feat(s3-types): add KMS service-control audit events

Configuration changes and service start/stop are management-plane actions
with no event name of their own, so they could not reach the audit
pipeline at all. Append three variants for them, following the existing
rule that KMS events are audit-only and live outside the `s3:` namespace,
so no bucket notification selector can expand to them.

* feat(admin): audit KMS management operations

Every KMS admin endpoint now builds an OperationContext from the
authenticated caller and hands it to the KMS layer, so the record the
manager already produces carries the principal, source address and
canonical request id instead of the internal placeholder.

A new adapter maps those records onto the server's existing AuditEntry
format and installs itself as the KMS audit sink at service assembly, so
KMS activity reaches the targets a deployment already operates. The
handlers emit directly for what the KMS layer cannot see: a request the
authorization gate rejects, and the endpoints with no context-aware KMS
entry point (data-key derivation and service control).

Only the failure class is recorded, never the error message, and the new
module joins the logging guardrail's checked files alongside the handlers
it serves.
2026-08-01 19:14:31 +08:00
houseme da389c0e21 fix(replication): harden backlog observability (#5564)
Add RAII guards for replication runtime backlog tickets so active worker and queue counters unwind on every terminal path.

Expose node-local MRF pending, dropped, missed, and flush-failure metrics through the bucket replication Prometheus collector while keeping the existing current backlog and durable MRF gauges additive.

Update durable MRF summary maintenance to aggregate incrementally during the persister loop, avoiding repeated full-entry scans on each successful flush.

Co-authored-by: heihutu <heihutu@gmail.com>
Co-authored-by: zhi22915 <qiuzgang@gmail.com>
2026-08-01 11:12:58 +00:00
Zhengchao An b7805caa58 ci: stop archiving every dependency's unpacked source in each cache (#5566)
The four consolidated ci keys plus the build keys still do not fit the
repository's fixed 10GB Actions cache quota, so LRU keeps evicting them:
measured demand is ci-dev 2331MB + ci-feat-proto 2310MB + ci-feat-rio 1951MB +
ci-uring 1317MB + two build legs at ~1429MB + cargo-deny 844MB, and main pushes
add two more build legs. That is roughly 14.4GB against 10.24GB. The symptom is
misleading: Cache Warm reports every job successful and the restores log "full
match: true", yet ci-feat-rio and ci-uring disappear from the cache list between
runs.

cache-all-crates was the wrong default for this repository. With it set to true,
rust-cache's cleanup returns before pruning ~/.cargo/registry/src and its config
archives the whole registry, so every cache carried the unpacked source tree of
every dependency — not, as the name suggests, just a few extra crates.

Setting it to false is rust-cache's own default and loses no coverage: the
package set comes from `cargo metadata --all-features`, a strict superset of any
single lane's feature closure; -sys crates are explicitly exempt from pruning,
since their source timestamps would otherwise trigger rebuilds; and everything
pruned is re-unpacked from the .crate files still in registry/cache, whose
mtimes crates.io normalises, so cargo fingerprints stay valid.

Applied to the setup composite and to audit.yml's own rust-cache. Cache Warm
now also reports the sizes of registry/src, registry/cache, registry/index,
~/.cargo/git and target/ to the step summary, immediately before rust-cache's
post step archives them, so the size of the effect is measured rather than
assumed.

The new sizes only appear once the cache key next rotates, since rust-cache
skips the save entirely on an exact key hit. Deliberately not forcing that by
bumping prefix-key: it would invalidate every family at once and produce a
repository-wide cold build.

Refs: rustfs/backlog#1598, rustfs/backlog#1600
2026-08-01 18:57:23 +08:00
Zhengchao An b0bb0bbd3a test(e2e): classify failed lock nodes as quorum loss (#5513) 2026-08-01 18:56:57 +08:00
GatewayJ bc41e567a5 fix(oidc): warn on request-header redirect fallback (#5561)
* fix(oidc): warn on request-header redirect fallback

* test(oidc): cover startup warning publication
2026-08-01 18:10:37 +08:00
houseme d5c6ba99d5 fix(ecstore): harden HotPath profiling boundaries (#5555)
* test(hotpath): gate mimalloc heap test by platform

Co-Authored-By: heihutu <heihutu@gmail.com>

* fix(ecstore): attribute HotPath CPU measurements to impls

Co-Authored-By: heihutu <heihutu@gmail.com>

* feat(ecstore): trace raw shard I/O with HotPath

Co-Authored-By: heihutu <heihutu@gmail.com>

* fix(obs): redact profiler and OPA endpoint diagnostics

Co-Authored-By: heihutu <heihutu@gmail.com>

* fix(ecstore): settle encoded queue accounting

Co-Authored-By: heihutu <heihutu@gmail.com>

* fix(ecstore): abort encoder producer on cancellation

Co-Authored-By: heihutu <heihutu@gmail.com>

---------

Co-authored-by: heihutu <heihutu@gmail.com>
2026-08-01 08:49:48 +00:00
houseme 62d44d10b8 Expose replication backlog gauges (#5557)
* fix(replication): count backlog at queue admission

Co-Authored-By: heihutu <heihutu@gmail.com>

* feat(obs): expose bucket replication backlog gauges

Co-Authored-By: heihutu <heihutu@gmail.com>

* fix(obs): report recent backlog from queued work

Co-Authored-By: heihutu <heihutu@gmail.com>

* fix(obs): preserve legacy backlog metric semantics

Co-Authored-By: heihutu <heihutu@gmail.com>

* feat(obs): expose durable MRF backlog gauges

Co-Authored-By: heihutu <heihutu@gmail.com>

* test(obs): cover replication backlog metric scope

Co-Authored-By: heihutu <heihutu@gmail.com>

* fix(obs): keep backlog metrics API-compatible

Co-Authored-By: heihutu <heihutu@gmail.com>

* refactor(obs): streamline replication backlog metrics

Keep MRF backlog accounting and OBS metric collection on a single, cheaper path.

Co-Authored-By: heihutu <heihutu@gmail.com>

* test(kms): update aws capability snapshot

Co-Authored-By: heihutu <heihutu@gmail.com>

---------

Co-authored-by: heihutu <heihutu@gmail.com>
2026-08-01 08:48:45 +00:00
houseme 8218248000 fix(hotpath): pin mimalloc allocator backend (#5550)
* fix(hotpath): pin mimalloc allocator backend

* test(hotpath): verify mimalloc allocator backend

Co-Authored-By: heihutu <heihutu@gmail.com>

* chore(hotpath): document unsafe allocator tests

Co-Authored-By: heihutu <heihutu@gmail.com>

* feat(kms): record real cache hit, miss and eviction metrics (#5531)

* feat(kms): record real cache hit, miss and eviction metrics

The metadata cache reported (entry_count, 0) because moka exposes no hit
or miss counts, so the miss half of every cache report was a constant.

Track lookups and removals in the cache itself: hit/miss counters on the
lookup path, a moka eviction listener classifying removals by cause, and
an entry gauge refreshed whenever the entry set changes. The counters are
exported through the metrics facade under the rustfs_kms_ prefix with
static label values only, matching the operation-policy metrics, and are
also returned as a KmsCacheStats snapshot in place of the old tuple.

Cache semantics are unchanged: capacity, TTL and invalidation points are
the same, and remove now flushes pending maintenance so the gauge and the
removal notification describe the cache the caller sees.

Refs rustfs/backlog#1584

* fix(kms): report real cache counters through the admin status API

KmsStatusResponse.cache_stats mapped the old (entry_count, 0) tuple onto
hit_count and miss_count, so operators polling KMS status read the entry
count as a hit count and a miss count that was always zero.

Map the fields to the counters they claim to be, and add entry_count and
eviction_count as additive, defaulted fields so the entry number that
hit_count used to carry is still available.

Refs rustfs/backlog#1584

* fix(kms): refresh the cache entry gauge on lookup misses

The entry gauge was published only from the write paths, so an entry
dropped by TTL expiry left `rustfs_kms_metadata_cache_entries` reporting
a population that no longer existed until the next put, remove or clear.
A cache that goes quiet — entries ageing out with no further writes —
kept over-reporting indefinitely.

Republish the gauge from the lookup path when the lookup misses. A miss
is where expiry surfaces, and moka reaps expired entries in the
maintenance it runs during that same lookup, so the count read
afterwards reflects the reaping. Hits stay free of the extra work.

* docs(kms): correct the entry gauge convergence claim on the miss path

The comment on the miss-path gauge refresh said moka reaps expired
entries in the maintenance it runs on that same lookup. It does not:
`should_apply_reads` is gated on a full read log or an elapsed
housekeeping interval, so the removal that decrements `entry_count` and
reaches the eviction listener may land on a later lookup.

The behaviour and the test are unchanged — the gauge still converges,
and the test drives `run_pending_tasks` explicitly rather than riding on
that interval. Only the stated guarantee was wrong, so say interval
instead of same-lookup and record why forcing maintenance on the read
path was not the trade taken.

* chore(deps): refresh cargo dependencies

Co-Authored-By: heihutu <heihutu@gmail.com>

---------

Co-authored-by: heihutu <heihutu@gmail.com>
Co-authored-by: Zhengchao An <anzhengchao@gmail.com>
2026-08-01 07:04:54 +00:00
Zhengchao An fd36bdfb1a feat(kms): allow key description and tag updates (#5546) 2026-08-01 06:32:47 +00:00
Zhengchao An 322ce21b9a feat(kms): add an AWS KMS backend (#5553) 2026-08-01 06:07:32 +00:00
Zhengchao An 35e4415ed9 feat(kms): expose key lifecycle and Vault credential gauges (#5542)
* feat(kms): observe key lifecycle from the deletion sweep

The sweep already pages through the whole key set, so the lifecycle
gauges come out of the pages it has in hand: no extra backend call is
made for them. It publishes the number of keys awaiting their deletion
deadline, the number of tombstones an interrupted removal left behind,
and how long ago the least recently rotated usable key was rotated
(counting from creation for keys that were never rotated), plus a
per-outcome counter of what the sweep acted on.

Every gauge is a label-less aggregate: a per-key label would carry key
identifiers into the metric stream and grow the series count with the
key set, so "this key is overdue for rotation" stays a threshold for an
alerting rule to apply to the aggregate. Keys the sweep destroys drop
out of the census, and a sweep that could not finish listing leaves the
gauges at their last complete values rather than understating them.

* feat(kms): expose Vault token TTL and fail-closed state as gauges

The renewal loop already tracks token expiry, so it now publishes the
seconds left on the active Vault token and whether the provider is
refusing to serve it. The fail-closed gauge re-evaluates the very gate
`VaultCredentialProvider::current` applies, so what operators see and
what the request path does cannot drift apart.

Both waits in the loop republish on a bounded cadence, so a scrape
landing between refresh cycles never reads a TTL frozen at the last
refresh or a fail-closed state that flipped after it. That costs a timer
and no Vault traffic, and the request path stays free of metric work.
Renewal successes and failures already land in the auth operation
counters, so nothing is double-counted here. Neither gauge carries a
label: the address, mount, auth path and token are all off limits as
label values, and there is one generation to describe.
2026-08-01 13:52:43 +08:00
Zhengchao An 4e34f97dd7 feat(kms): converge KMS configuration across nodes after a runtime change (#5551) 2026-08-01 05:41:02 +00:00
Zhengchao An 782c78e0ef feat(sse): enforce per-key KMS authorization on the SSE-KMS data path (#5538) 2026-08-01 05:25:18 +00:00
Zhengchao An 8387528c9b feat(kms): record real cache hit, miss and eviction metrics (#5531)
* feat(kms): record real cache hit, miss and eviction metrics

The metadata cache reported (entry_count, 0) because moka exposes no hit
or miss counts, so the miss half of every cache report was a constant.

Track lookups and removals in the cache itself: hit/miss counters on the
lookup path, a moka eviction listener classifying removals by cause, and
an entry gauge refreshed whenever the entry set changes. The counters are
exported through the metrics facade under the rustfs_kms_ prefix with
static label values only, matching the operation-policy metrics, and are
also returned as a KmsCacheStats snapshot in place of the old tuple.

Cache semantics are unchanged: capacity, TTL and invalidation points are
the same, and remove now flushes pending maintenance so the gauge and the
removal notification describe the cache the caller sees.

Refs rustfs/backlog#1584

* fix(kms): report real cache counters through the admin status API

KmsStatusResponse.cache_stats mapped the old (entry_count, 0) tuple onto
hit_count and miss_count, so operators polling KMS status read the entry
count as a hit count and a miss count that was always zero.

Map the fields to the counters they claim to be, and add entry_count and
eviction_count as additive, defaulted fields so the entry number that
hit_count used to carry is still available.

Refs rustfs/backlog#1584

* fix(kms): refresh the cache entry gauge on lookup misses

The entry gauge was published only from the write paths, so an entry
dropped by TTL expiry left `rustfs_kms_metadata_cache_entries` reporting
a population that no longer existed until the next put, remove or clear.
A cache that goes quiet — entries ageing out with no further writes —
kept over-reporting indefinitely.

Republish the gauge from the lookup path when the lookup misses. A miss
is where expiry surfaces, and moka reaps expired entries in the
maintenance it runs during that same lookup, so the count read
afterwards reflects the reaping. Hits stay free of the extra work.

* docs(kms): correct the entry gauge convergence claim on the miss path

The comment on the miss-path gauge refresh said moka reaps expired
entries in the maintenance it runs on that same lookup. It does not:
`should_apply_reads` is gated on a full read log or an elapsed
housekeeping interval, so the removal that decrements `entry_count` and
reaches the eviction listener may land on a later lookup.

The behaviour and the test are unchanged — the gauge still converges,
and the test drives `run_pending_tasks` explicitly rather than riding on
that interval. Only the stated guarantee was wrong, so say interval
instead of same-lookup and record why forcing maintenance on the read
path was not the trade taken.
2026-08-01 05:19:25 +00:00
Zhengchao An 4ce0e280f2 feat(kms): add a synthetic encrypt-decrypt probe worker (#5543) 2026-08-01 12:30:35 +08:00
Zhengchao An 793c193a6b feat(admin): scope KMS admin authorization to the target key (#5533) 2026-08-01 04:18:52 +00:00
Zhengchao An 5a6e850c67 feat(kms): wire OperationContext into an audit event contract (#5534) 2026-08-01 04:12:00 +00:00
Zhengchao An 2d8ad5caee ci: drop cla.yml to least privilege (#5548) 2026-08-01 12:00:46 +08:00
GatewayJ 3d4f4bb86d fix(sts): align AssumeRole authorization (#5281)
* fix(sts): align AssumeRole authorization with MinIO

* test(sts): cover AssumeRole OPA contract

* fix(iam): fail closed on unresolved policies

* fix(iam): fail closed while OPA initializes

---------

Co-authored-by: cxymds <cxymds@gmail.com>
Co-authored-by: Zhengchao An <anzhengchao@gmail.com>
2026-08-01 11:48:58 +08:00
Zhengchao An 2b31bda6d1 ci: clear the checkout token in every job that does not push (#5547)
actions/checkout writes its token into .git/config as an http extraheader,
where it stays for the rest of the job. That matters more here than usual:
pull_request jobs run on self-hosted runners and execute the PR's own build.rs,
proc-macros and tests, any of which can read that file. Test and Lint holds
actions: write on top of that, so its token can cancel runs and delete the
Actions caches the whole pipeline now depends on — it was given
persist-credentials: false when that permission was added, and this extends the
same treatment to the other 45 checkouts.

Two are exempt because the token IS the credential the job needs.
helm-package's publish job pushes to rustfs/helm with it; clearing it would
break chart publishing. nix-flake-update is exempt pending verification: it
passes FLAKE_UPDATE_TOKEN to create-pull-request directly rather than reusing
.git/config, so it very likely does not need persistence, but that is unproven
and a broken weekly bot is not worth the guess. Both carry a
persist-credentials-exempt comment saying which.

scripts/security/check_persist_credentials.sh requires every checkout to either
clear its credentials or carry that comment, so the decision stays visible in
review rather than being an omission nobody notices.

Refs: rustfs/backlog#1598, rustfs/backlog#1602
2026-08-01 11:48:47 +08:00
Zhengchao An 68e344bf03 fix(scanner): remove total timeout from heal walks (#5502)
Co-authored-by: cxymds <cxymds@gmail.com>
2026-08-01 03:47:42 +00:00
Zhengchao An 48c2fcb62b ci: add an opt-in cargo --timings run to decide the sccache question (#5545)
sccache can only cache compilation units whose --emit includes link. That
covers workspace rlibs and nothing else: clippy is metadata-only, and the ~100
test binaries, the rustfs bin and every build script invoke the system linker.
So the headline claim that started this — s3select-query costing 17m13s inside
a 19m32s nextest build — does not by itself justify sccache, because cargo
prints "Compiling" when a crate starts and never when it finishes: that number
is the whole compilation tail, not one crate.

Adding --timings behind a workflow_dispatch input on Cache Warm answers it with
data instead. Read two shares off the report: workspace lib codegen against the
whole build, and s3select-query's own rlib. The plan in rustfs/backlog#1601
adopts sccache only above 50% and 25%; if linking dominates instead, the answer
is mold/lld plus split-debuginfo, which is precisely the region sccache cannot
reach.

Off by default: it doubles the ci-dev build, and this only needs running once
per question. Nothing about the normal warm path changes.

Refs: rustfs/backlog#1598, rustfs/backlog#1601
2026-08-01 11:36:48 +08:00
Zhengchao An 428fde069d ci: drop the docker layer cache and stop installing unused tooling (#5544)
Two cleanups, neither changing what is built or published.

docker.yml loses its type=gha layer cache. The image build compiles nothing —
it downloads a release zip and runs apk/apt — so the cache could only save the
minute or two those take, against a real correctness problem: with
RELEASE=latest the binary URL is resolved by curl inside a RUN layer, and the
layer key does not include what it resolved to. A rebuild at the same RELEASE
value (a dispatch with version=latest, or a re-run of the same version) would
hit the old layer and ship the previous release's binary. mode=max also drew on
the same repo-wide 10GB Actions cache quota the Rust lanes are contending for.

Only RELEASE is passed as a build-arg now; it is the sole one the Dockerfiles
declare besides TARGETARCH. BUILDTIME, VERSION, BUILD_TYPE, REVISION and CHANNEL
were read by no stage, and BUILDTIME's $(date ...) was a literal string in the
YAML block rather than a substitution. The DOCKER_CHANNEL computation that fed
CHANNEL evaluated to "release" down every branch and is gone with it. BUILD_DATE
and VCS_REF stay unset even though the Dockerfiles declare them: supplying them
would change the published image labels.

The setup composite stops installing tools its callers do not use. protobuf-
compiler comes out of the apt list entirely — setup-protoc installs 34.1 into
the tool cache and prepends it to PATH, so the apt build was shadowed on every
run and never used; the same duplicate install is removed from the io_uring lane.
musl-tools, zip and unzip move behind install-build-packaging-tools, off for the
CI and coverage lanes and left on for build.yml, whose native musl leg needs
musl-gcc and whose packaging steps need zip. cargo-nextest and the
rustfmt/clippy components move behind install-test-tools, off only for build.yml,
which runs no tests and no lints; coverage and the nightly replication lane keep
them because both invoke nextest.

The action's github-token input is deleted along with its 20 call sites. Its
runs block never referenced it — setup-protoc uses github.token directly — so it
was 20 places handing a token to something that ignored it.

Refs: rustfs/backlog#1598, rustfs/backlog#1600, rustfs/backlog#1603
2026-08-01 11:35:34 +08:00
Zhengchao An 364168c0ba docs(kms): record the compliance position and mixed-version constraints (#5541)
* docs(kms): record the cryptographic compliance position

RustFS links no FIPS-validated cryptographic module: the rustls provider is
the ordinary aws-lc-rs build, and every data-path AEAD is RustCrypto. The
crypto crate's default-on `fips` feature only selects PBKDF2+AES-GCM over
Argon2id, with the same RustCrypto implementations behind both branches, so
it cannot support a validation claim either.

Document that status, the terminology rules for external material, the real
semantics of the `fips` feature with a rename direction, the cost of the
three routes to a stronger position, and the sequencing rules for retiring an
algorithm.

Refs rustfs/backlog#1587 (part of rustfs/backlog#1562)

* docs(kms): document the mixed-version cluster constraints

Collect the cross-version constraints that landed with versioned rotation and
the check-and-set lifecycle work: which persisted formats decode both ways,
which guarantees only hold once every node is upgraded, how long nodes can
disagree on lifecycle state, and that reconfigure is persisted cluster-wide
but applied only on the node that handled it.

Adds the recommended rolling-upgrade sequence and the list of operations to
avoid while two builds are running.

Refs rustfs/backlog#1581 (part of rustfs/backlog#1562)
2026-08-01 11:34:04 +08:00
Zhengchao An 3f4f31129e ci: narrow the io_uring lane and make the sampler attributable (#5540)
Two independent changes to the Test and Lint area.

The io_uring lane compiled 7 integration binaries to run none of their tests.
Job log 91309868055 shows the lib target reporting "18 passed; 3453 filtered
out" while every binary under crates/ecstore/tests/ reported "running 0 tests".
Adding --lib drops them from the build without changing the selected set.

The `uring_` filter itself must not be touched. libtest matches on substring, so
it also selects names containing `during_` — 6 of the 18 selected tests are such
incidental matches, and narrowing the filter to `io_uring` would silently drop
them. scripts/check_uring_lane_lib_only.sh asserts the precondition --lib
depends on: no test function whose name contains `uring_` may live under
crates/ecstore/tests/. A count floor would not do, because the dangerous case —
someone adding a matching test there — leaves the lib count unchanged and CI
green.

The resource sampler was measuring the wrong machine. The runners are ARC pods,
so /proc/loadavg, /proc/pressure/*, `free` and `df` are node-level and include
every other runner pod on the same Kubernetes node: one sample reported loadavg
6.67 with 832 threads node-wide while `ps` inside the pod showed about 10
processes. Judging a CARGO_BUILD_JOBS change on those numbers cannot work.

The sampler moves to scripts/ci/resource_sampler.sh and now records
/sys/fs/cgroup cpu.max, memory.max, memory.peak and the cpu/io/memory pressure
files, which are this pod's own. The node-level readings stay — co-tenancy is a
real cause of stalls, and a 9m57s plain `git checkout` was traced to it — but
are labelled NODE-LEVEL so nobody reads them as this job's load. Phase markers
are written on start, and clippy is now sampled too: it is the natural control
arm for a CARGO_BUILD_JOBS experiment, since --all-targets is check-only for
workspace members and never links the ~100 test binaries the limit throttles.

No numbers are changed in this commit. CARGO_BUILD_JOBS stays at 2 until there
is attributable data to change it on.

Refs: rustfs/backlog#1598, rustfs/backlog#1601
2026-08-01 11:32:00 +08:00
Zhengchao An 6c99d4fe22 ci: stop the weekly flake.lock bot from running the whole pipeline (#5539)
nix-flake-update opens a PR every Sunday that changes flake.lock and nothing
else. flake.lock is consumed only by Nix packaging — cargo never reads it — yet
it was in no paths filter, so each of those PRs ran the full Continuous
Integration pipeline and the merge then ran Build and Release too. Added to all
four lists that have to agree: ci.yml's push and pull_request paths-ignore,
build.yml's push paths-ignore, and ci-docs-only.yml's paths.

The cron stays. flake.lock still has consumers — anyone running `nix build` or
`nix develop` — so stopping the updates would remove the workflow's output, not
just its CI cost.

Those four lists drifting is a silent failure, so scripts/check_ci_paths_sync.sh
now asserts the pair that matters: ci.yml's pull_request paths-ignore must equal
ci-docs-only.yml's paths. An entry present only in the first means a PR touching
those files triggers neither workflow, nobody reports "Test and Lint" or "Quick
Checks", and the PR waits on a required check forever. It also asserts both
companion job names still exist, since renaming one produces the same hang. The
push list is not compared: no required check is reported for push events.

Also in this cleanup:

- nix-flake-update's GITHUB_TOKEN drops to contents: read. The branch push and
  the pull request are both made by update-flake-lock with the
  FLAKE_UPDATE_TOKEN PAT, so the write scopes were an unused repo-write
  credential on an unattended weekly job.
- build.yml's build_docker dispatch input is documented as advisory. docker.yml
  triggers on workflow_run and requires the triggering event to be a tag push,
  so a manual dispatch never reaches it whatever this input says.
- The nine disabled workflow files get a banner saying so. Their
  disabled_manually state lives in GitHub's UI and is invisible when reading the
  file, which has already misled one audit into treating dead workflows as live.

Refs: rustfs/backlog#1598, rustfs/backlog#1603
2026-08-01 11:28:14 +08:00
Zhengchao An f5348d5cc4 fix(ecstore): settle lease-deferred data-dir deletes before bucket removal (#5516)
A streaming GET holds snapshot leases on the object's data directories,
and DeleteObjects defers their physical cleanup until the leases are
released. A DeleteBucket issued inside that window passes the xl.meta
emptiness check but fails closed in the non-force delete_volume tree
removal on the leftover part files, returning BucketNotEmpty for a
logically empty bucket.

This is what intermittently failed the s3tests
test_encryption_sse_c_multipart_bad_download teardown in CI: the test
never reads its 30MiB GET body, so the server-side stream (and its
leases) stays alive until the connection drops, racing the teardown's
DeleteObjects + DeleteBucket sequence. With the body held open the
failure reproduces 5/5 locally; after this change it passes 20/20, and
the real s3-tests case passes 20 consecutive runs.

delete_volume now executes the registry-tracked pending deferred
deletions for the volume before removing the directory tree. Only data
dirs whose logical delete already committed are touched; unknown files
still fail closed with VolumeNotEmpty.
2026-08-01 03:26:58 +00:00
Zhengchao An e2b2bdcc34 ci: bound every job's runtime and stop pasting inputs into shell (#5537)
Three hardening changes with no effect on what any workflow produces.

Declare timeout-minutes on the 25 jobs that lacked it. GitHub's default is 360
minutes, and this repository has a history of runners stalling intermittently
(#5394) plus a measured 9m57s plain `git checkout` under node-level I/O
contention, so one wedged job could hold a runner for six hours out of a pool of
roughly 15-21. Budgets follow what the jobs actually do: 10 minutes for
echo-only and guard-script jobs, 30 for anything calling the GitHub API,
uploading release assets or pushing over the network.
scripts/security/check_job_timeouts.sh keeps it that way, checking only jobs
that declare runs-on so reusable-workflow callers are not flagged.

Pass workflow inputs and workflow_run fields through env instead of `${{ }}`
interpolation in run blocks. A git ref name may contain `$(...)` — any string
without a space is a legal tag — and interpolation pastes it into the script
where bash evaluates it. The worst instance was helm-package's final commit
message: it is built from the triggering tag name inside the job that holds the
cross-repository push token with rustfs/helm already checked out. Also converted
in build.yml, docker.yml and performance-ab.yml; the last is currently disabled,
but a disabled workflow can be re-enabled. Not touched: helm-package's
`contains(head_branch, '.')` tag test, since GitHub expressions have no regex
and this repository's tags carry no `v` prefix, so rewriting the condition would
change which builds publish a chart.

Give audit.yml a scheduled-failure alert and run it daily. A scheduled
cargo-deny failure usually means the dependency tree just matched a newly
published RustSec advisory — the most important signal this workflow produces,
and until now it was visible only to whoever happened to open the Actions tab.
coverage.yml and e2e-replication-nightly.yml already use this ci-8 mechanism.
The cron moves from weekly to daily so a new advisory against an unchanged tree
surfaces within a day instead of seven; the check list is untouched, since
splitting it into a light daily run and a weekly full run would create runs
where sources, bans and licenses go unverified.

Refs: rustfs/backlog#1598, rustfs/backlog#1602
2026-08-01 11:25:26 +08:00
houseme b965bd6eef fix(storage): harden scanner and recovery edge cases (#5521)
* fix(ecstore): handle benign listing and GET disconnects

Co-Authored-By: heihutu <heihutu@gmail.com>

* fix(scanner): scope cache locks by set

Co-Authored-By: heihutu <heihutu@gmail.com>

* test(kms): stabilize Vault transport retry coverage

Co-Authored-By: heihutu <heihutu@gmail.com>

* fix(scanner): fence scoped cache locks by protocol

Co-Authored-By: heihutu <heihutu@gmail.com>

* test(ecstore): stabilize topology DNS fallback coverage

Co-Authored-By: heihutu <heihutu@gmail.com>

* fix(kms): remove stale local export test import

Co-Authored-By: heihutu <heihutu@gmail.com>

---------

Co-authored-by: heihutu <heihutu@gmail.com>
2026-08-01 03:25:13 +00:00
Zhengchao An 1524ed891f ci: write the Rust caches from a workflow that is not cancelled (#5536)
#5532 consolidated ci.yml's nine cache sites into four keys with one writer
each, but the writers never run to completion: ci.yml cancels superseded runs
on main, and merges land far faster than its 70-minute pipeline. Measured over
15 consecutive main pushes — 12 cancelled, 2 failed, 0 succeeded. A cancelled
run never reaches rust-cache's post step (cache-on-failure does not cover
cancellation), so nothing was being saved and every PR paid a cold restore:
11.8-20.9 minutes of "Setup Rust environment" against 0.7-3.4 warm.

Move cache writing into its own workflow whose concurrency group does not
cancel in progress, and make every lane in ci.yml a pure reader. ci.yml keeps
cancelling superseded runs, which is correct — nobody needs test results for a
commit that is already three merges behind — while the caches still get
written.

Not simply disabling cancel-in-progress for main pushes in ci.yml: that would
run several full pipelines concurrently on a 15-21 runner self-hosted pool that
is already the bottleneck, which is the opposite of the goal. GitHub keeps at
most one running plus one pending run per concurrency group, so the new
workflow collapses a burst of merges into "current finishes, newest queued
follows" and occupies one runner at a time.

The warm builds are supersets of what the reading lanes compile, because a
reader restores only what the writer saved: --all-targets for the test binaries
nextest builds (including e2e_test, which test-and-lint excludes), plus the
e2e-test-hooks and rio-v2 feature combinations, whose different feature
resolution yields a different -Cmetadata. The ci-dev build carries the same
CARGO_BUILD_JOBS limit ci.yml puts on nextest, since it links the same ~100
test binaries and three concurrent links can wedge Cargo (#5394). ci-uring is
warmed on ubuntu-latest to match its reader: rust-cache's key covers runner.os
and arch but not the runner label or image.

Refs: rustfs/backlog#1598, rustfs/backlog#1600
2026-08-01 11:20:02 +08:00
Zhengchao An 7354a5663d ci: fit the Rust caches back inside the 10GB quota (#5532)
The repository has 18 rust-cache families of 1.2-3.1GB each against GitHub's
fixed 10GB per-repo quota. Measured usage sat at 9.72GB, 9.96GB and 11.63GB on
three samples, so LRU eviction is continuous and the main lanes lose: ci-test
and ci-e2e were repeatedly absent from the surviving entries. That is what
makes "Setup Rust environment" bimodal — 0.7-3.4 minutes warm against
11.8-20.9 minutes cold.

Four changes, all cache-only. No job builds or tests anything different.

Collapse ci.yml's nine cache sites into four keys, each with exactly one
writer: ci-dev (test-and-lint writes; ILM, debug-binary, e2e-tests and e2e-full
read), ci-feat-rio (rio-v2 lint writes, its debug-binary reads), ci-feat-proto
(the swift leg writes for both protocol legs), and ci-uring, which stays alone.
e2e-full is easy to miss here: it already shared ci-e2e with e2e-tests and both
saved, so without an explicit 'false' it would have become a second unnamed
writer of ci-dev. rio/swift/sftp are deliberately NOT merged — measured at
2365/2307/1200MB they are not near-identical, and none is a superset of the
others.

Give each writer a superset warm-up on main. A writer's own steps are not
automatically a superset of its readers': clippy emits metadata only, the
nextest pass excludes e2e_test, and no lint lane enables e2e-test-hooks, whose
different feature resolution yields a different -Cmetadata. Without these
builds the readers would restore a cache missing precisely what they need.
Guarded to main, so the PR critical path is unaffected.

Stop writing tag-scoped caches in build.yml. A cache saved on refs/tags/X can
only be restored by a re-run of that same tag, so each release cycle wrote up
to 12 unreadable 1-2GB entries that evicted the hot lanes. Tag builds still
restore the main-scoped cache. The one real cost is that re-running a failed
leg of the same tag now falls back to main's cache.

Flip the composite action's cache-save-if default to 'false' and make the input
mandatory in practice. audit.yml was relying on the old "true" default: every
PR touching Cargo.toml or Cargo.lock saved a second, PR-scoped copy (~843MB
measured, job 91048468127) that pushed main-scoped lanes out of the quota. The
fail-safe direction is a cold cache, not a stolen quota slice.
scripts/security/check_cache_save_if.sh now asserts every call site states it,
wired into audit.yml next to the existing pin check.

While there, cargo-deny stops pulling the full setup composite. It compiles
nothing, so apt, protoc, flatc and nextest were pure overhead — but it does run
cargo metadata, and Cargo.toml pins datafusion and s3s as git dependencies that
must be materialised into ~/.cargo/git, so the cache itself stays.

Refs: rustfs/backlog#1598, rustfs/backlog#1600
2026-08-01 11:06:31 +08:00
Zhengchao An 790bdc0e63 ci: gate the seven expensive jobs on Quick Checks (#5529)
* ci: stop running expensive jobs that cannot inform the result

Three independent fixes that all avoid burning self-hosted runners on work
whose outcome is already determined. None of them changes what is tested.

- ci-docs-only: add a "Quick Checks" companion job. It is a prerequisite for
  gating ci.yml's expensive jobs behind quick-checks (rustfs/backlog#1599):
  once "Quick Checks" is a required check, a docs-only PR would otherwise wait
  on it forever. The steps are a byte-identical copy of ci.yml's quick-checks
  rather than an `echo`, so that on a mixed PR the two same-named check runs
  execute the same commands against the same merge ref and cannot disagree —
  GitHub has no written contract for how it picks between same-named required
  check runs, and the real job only takes 45-51s, leaving no timing margin to
  rely on.

- ci: guard uring-integration with the same `closed` check every other job
  already has. The pull_request trigger includes `closed` only so the
  concurrency group cancels in-flight runs; this job had no guard and no
  `needs`, so every closed or merged PR ran the full io_uring suite (4m17s,
  7m19s and 7m31s on runs 30678272341, 30678117601 and 30662728539).

- ci: gate s3-lifecycle-behavior-tests on e2e-tests, matching
  s3-implemented-tests. Both lanes only download the prebuilt debug binary, and
  s3-implemented-tests already finishes later, so a green PR's wall clock is
  unchanged; a red one stops holding a sm-standard-4 for up to 30 minutes.

Refs: rustfs/backlog#1598, rustfs/backlog#1599

* ci: gate the seven expensive jobs on Quick Checks

Every expensive job started in parallel with quick-checks, so a formatting or
architecture-guard failure still paid for the full pipeline. On run
30673292690 Quick Checks failed after 0.8 minutes and the run went on to burn
424.5 runner-minutes — 99.8% of it after the gate had already failed. The
self-hosted pool is 15-21 ARC runners and one full PR run needs about seven
sm-standard-4 concurrently, so those minutes come straight out of other PRs'
queue time (six runs measured 72-488 minutes queued).

quick-checks itself is compile-free and takes 45-51s, so a passing PR pays
about a minute of extra critical path.

REQUIRES the branch ruleset to list "Quick Checks" as a required check BEFORE
this merges. Adding `needs` gives these jobs a `skipped` conclusion for the
first time, and GitHub treats a skipped required check as satisfied — with
required_approving_review_count=0, a failing quick-checks would otherwise let
a broken PR merge. Ordering is tracked in rustfs/backlog#1599.

Refs: rustfs/backlog#1598, rustfs/backlog#1599

* ci: stop a PR run once Test and Lint has failed (#5530)

On run 30674613104 the e2e, ILM and sftp lanes had all failed while Test and
Lint and the rio-v2 variant kept running past 70 minutes. The run's verdict was
settled; the remaining lanes were spending sm-standard-4 time on a result
nobody could act on, and with one full PR run needing about seven of those
runners, that time comes out of other PRs' queue time.

Two mechanisms, both scoped to pull_request so main pushes, the merge queue and
the weekly schedule keep the full failure signal:

- test-and-lint-protocols: fail-fast on PRs, so one failing protocol leg stops
  its sibling. This is the only part that also covers fork PRs, since it needs
  no token.
- test-and-lint: on failure, cancel the run through the REST API.

Only test-and-lint may cancel. The lanes that are not required checks
(protocols, ILM, e2e, s3-tests) must never hold that power: a flake in one of
them would turn the required "Test and Lint" into `cancelled`, which blocks the
merge. A maintainer can merge today with sftp red, and that has to stay true.

The cancel step uses curl, not `gh`: every existing `gh` call in this repo runs
on ubuntu-latest, and the sm-standard-* images are custom and trimmed, so `gh`
is not known to exist there. Fork PRs are excluded by an explicit condition
rather than left to fail, since their GITHUB_TOKEN is forced read-only and
job-level permissions cannot raise it.

Job-level permissions must list contents: read alongside actions: write —
job-level permissions replace the workflow block instead of merging with it,
and dropping contents would break this job's checkout and the repo-token the
setup action passes to setup-protoc. Because that token can now cancel runs and
delete Actions caches, the checkout also sets persist-credentials: false so a
PR's own build.rs or proc-macro cannot read it back out of .git/config.

Refs: rustfs/backlog#1598, rustfs/backlog#1599

* ci: correct the companion-workflow comments

Addresses review feedback on #5528, which merged before these fixes were
pushed.

The ci-docs-only header claimed the ruleset already requires "Quick Checks".
It does not — that ruleset change is a separate step, and this file's whole
purpose is to land first so that change does not strand docs-only PRs. Say
what is true today.

Also move the byte-identical requirement onto ci.yml's quick-checks job, which
is the more likely edit site, instead of pointing at a comment that was not
there.
2026-08-01 10:57:02 +08:00
Zhengchao An 707d062174 ci: stop running expensive jobs that cannot inform the result (#5528)
Three independent fixes that all avoid burning self-hosted runners on work
whose outcome is already determined. None of them changes what is tested.

- ci-docs-only: add a "Quick Checks" companion job. It is a prerequisite for
  gating ci.yml's expensive jobs behind quick-checks (rustfs/backlog#1599):
  once "Quick Checks" is a required check, a docs-only PR would otherwise wait
  on it forever. The steps are a byte-identical copy of ci.yml's quick-checks
  rather than an `echo`, so that on a mixed PR the two same-named check runs
  execute the same commands against the same merge ref and cannot disagree —
  GitHub has no written contract for how it picks between same-named required
  check runs, and the real job only takes 45-51s, leaving no timing margin to
  rely on.

- ci: guard uring-integration with the same `closed` check every other job
  already has. The pull_request trigger includes `closed` only so the
  concurrency group cancels in-flight runs; this job had no guard and no
  `needs`, so every closed or merged PR ran the full io_uring suite (4m17s,
  7m19s and 7m31s on runs 30678272341, 30678117601 and 30662728539).

- ci: gate s3-lifecycle-behavior-tests on e2e-tests, matching
  s3-implemented-tests. Both lanes only download the prebuilt debug binary, and
  s3-implemented-tests already finishes later, so a green PR's wall clock is
  unchanged; a red one stops holding a sm-standard-4 for up to 30 minutes.

Refs: rustfs/backlog#1598, rustfs/backlog#1599
2026-08-01 10:49:19 +08:00
houseme 4b6b6f14bd feat(hotpath): use mimalloc in inner counting allocator (#5523)
* feat(hotpath): use mimalloc in counting allocator

* fix(hotpath): adapt mimalloc for allocation counting

---------

Co-authored-by: overtrue <anzhengchao@gmail.com>
2026-08-01 10:43:52 +08:00
Zhengchao An 9080ea8ea0 test(object-lock): cover unretained version cleanup (#5504)
Co-authored-by: cxymds <cxymds@gmail.com>
2026-08-01 10:14:32 +08:00
Zhengchao An 739efaaea1 feat(kms): restore local backend key material from sealed backup bundles (#5522)
* feat(kms): add link_durably primitive and restore-marker startup guard

The local backend gains the two pieces the bundle restore path builds on:
a no-clobber hard-link publish primitive whose AlreadyExists case is
idempotent only for byte-identical content, and a fail-closed startup
guard that refuses to open a key directory holding a restore cutover
marker. The durable commit protocol and the key-id containment check
become pub(crate) so the restore module reuses them instead of copies.

* feat(kms): record a master-key verifier and pre-seal decrypt probe in export

Fill the manifest's master_key_verifier slot with an opaque one-way
value (scheme-prefixed, bound to the backup id and the KDF salt) so a
restore can detect a wrong operator-supplied master key before touching
any target state, and probe-decrypt every artifact as stored under the
backup KEK before the manifest may seal — digest equality alone only
proves the ciphertext landed intact. The payload decryption tail is
factored out and shared with the restore side so producer and consumer
cannot drift on the framing.

Also drops the stale KmsClient test import orphaned by the backend
refactor (#5501); the test suite did not compile without this.

* feat(kms): restore local backend key material from sealed backup bundles

The consumer side of the Local bundle export, as a four-phase protocol:

- Dry-run: full in-memory bundle decode (digest and AEAD verification of
  every artifact), KDF-drift detection against the compiled-in
  derivation, deployment and injected-generation checks (strictly lower
  is rejected, equal stays allowed for repeated drills), master-key
  verifier check, and target conflict enumeration - with zero writes.
- Staging: artifacts are committed durably into the .restore-staging/
  subdirectory (invisible to the backend's key scan and orphan-temp
  matcher) and every record is decryption-probed with the derived
  master key both in memory before staging and again from the staged
  bytes.
- Commit marker + cutover: the durably published .restore-commit.json
  marker is the single commit point; cutover publishes staged files via
  link_durably (salt first, keys after), then durably removes the
  marker and drops staging.
- Crash re-entry: before the marker the target top level is untouched
  and a re-run starts over; with the marker published, backend startup
  fails closed and a re-run with the same bundle rolls forward while
  abort_local_restore rolls back. Every interruption converges to the
  complete old or complete new state.

Restore never goes through LocalKmsClient::new (which would mint a
fresh salt); the only write mode is the explicit
restore-into-empty-target policy, where an orphan salt or a foreign
marker already counts as non-empty. The bundle source stays strictly
read-only.

Refs rustfs/backlog#1572
2026-08-01 10:05:47 +08:00
Zhengchao An b6d4689c75 feat(policy): parse and match KMS key resources with a kms:Decrypt action (#5515)
Bring the dormant Resource::Kms variant to life: arn:aws:kms:::key/<key_id>
patterns (empty-account form, wildcards allowed in the id, alias/<name>
reserved as parse-only) now parse, validate, serialize back, and are matched
by KMS statements against the requested key id carried in Args::object.
Statements without KMS resources keep the legacy match-every-key behaviour,
as do call sites that pass no key resource, so nothing changes until the
admin/SSE authorization paths start passing key ids.

Statement validation rejects KMS resources on non-KMS statements, and bucket
policy validation rejects KMS actions and resources outright while stored
policies keep deserializing; evaluation skips pure-KMS bucket policy
statements with a warning. A kms:Decrypt action is added for the upcoming
SSE-KMS read path.

Refs rustfs/backlog#1582 (part of rustfs/backlog#1562)
2026-08-01 10:05:40 +08:00
Zhengchao An 8763cd0c67 fix(kms): CAS Transit metadata writes and bound the metadata cache (#5520)
* fix(kms): drop the KmsClient trait import removed by #5501

The local backup export tests (#5499) merged after #5501 folded the
KmsClient trait into KmsBackend, leaving a dead trait import that breaks
'cargo test -p rustfs-kms' compilation on main. create_key is an inherent
method on LocalKmsClient since #5501, so the import is unnecessary.

* fix(kms): CAS Transit metadata writes and bound the metadata cache

Transit KV metadata writes were whole-record overwrites with no
precondition, so two nodes mutating the same key could silently clobber
each other's lifecycle state, and the process-local metadata cache had
neither a TTL nor a capacity bound, so a key disabled or scheduled for
deletion on one node stayed usable on every other node until restart.

- Replace write_metadata_to_kv with a versioned read
  (read_metadata_from_kv_versioned) plus a check-and-set write
  (cas_write_metadata_to_kv); mutate_key_metadata re-reads the
  authoritative record and re-runs the state gate on every attempt, and a
  lost CAS race retries with a fresh snapshot (bounded budget) instead of
  replaying the stale one.
- Migrate every read-modify-write caller: enable, disable, schedule and
  cancel deletion, rotate version bump, the expired-key tombstone, and
  both create paths (create-only CAS that read-confirms the winner on a
  lost race).
- Bound the metadata cache with moka (300s TTL, 1024 entries) and drop a
  key's entry when a transit data call reports it gone server-side.
- Fail closed when the synthesized-metadata fallback cannot be read or
  persisted: the fabricated Enabled record is only served after a durable
  create-only CAS write, closing the gate weakening documented as a KNOWN
  RISK; the persistence fallback for pre-metadata keys is kept.

Refs rustfs/backlog#1581 (part of rustfs/backlog#1562)
2026-08-01 10:05:33 +08:00
Zhengchao An fdac60b0e2 fix(kms): make Vault KV2 lifecycle writes check-and-set (#5518)
* fix(kms): drop stale KmsClient trait import in local_export tests

The KmsClient trait was folded into KmsBackend (#5501), but the backup
export tests merged afterwards (#5499) still imported it, breaking the
crate's test build; create_key is an inherent LocalKmsClient method, so
the import is simply unused.

* fix(kms): make Vault KV2 lifecycle writes check-and-set

Every KV2 lifecycle write used to be a blind whole-record overwrite, so
two nodes racing on the same key could lose updates: a disable racing a
rotation wrote the pre-rotation record back (rolling back the version
and material of a committed rotation), concurrent same-name creates let
the later material win (orphaning DEKs wrapped under the earlier one),
and a cancellation racing the deletion sweep could be overwritten by
the tombstone (or resurrect an already tombstoned key).

All lifecycle mutations now go through a bounded check-and-set
read-modify-write loop: each attempt re-reads the record pinned to its
KV2 secret version, re-runs the state gate against the fresh snapshot,
and writes back check-and-set against exactly that version; after
LIFECYCLE_CAS_ATTEMPTS lost races the typed conflict error surfaces.
The loop composes with the operation policy's single-attempt rule for
non-idempotent writes: each write is still sent at most once, only the
whole read-gate-write cycle repeats. create_key becomes a create-only
write (cas=0) so exactly one of two concurrent creates commits and the
loser reports KeyAlreadyExists. The blind store_key_data primitive is
now test-only.

Reads and rotation additionally fail closed when the version history is
inconsistent: resolving material through a version record above the
current pointer is refused (that state only arises when a lost update
rolled back a committed rotation), and rotation refuses to extend a
history whose records reach more than one step past the current pointer
(one step ahead is the footprint of an interrupted rotation and still
recovers through the adopt path).

Refs rustfs/backlog#1581
2026-08-01 09:36:12 +08:00
Zhengchao An 98b20b4231 test(s3): cover delete marker list visibility (#5503) 2026-08-01 09:31:21 +08:00
Zhengchao An ffc9de72cb docs: make validation proportional to change risk (#5527) 2026-08-01 09:31:03 +08:00
cxymds c09d11ff3b fix(config): fence persisted config updates and reloads (#5512)
* fix(config): fence persisted config updates and reloads

* fix(ci): unblock config and e2e checks

Co-Authored-By: heihutu <heihutu@gmail.com>

---------

Co-authored-by: houseme <housemecn@gmail.com>
Co-authored-by: heihutu <heihutu@gmail.com>
2026-08-01 00:33:46 +00:00
Zhengchao An 792f2ef204 docs(kms): add observability dashboard, alert rules and runbook (#5517)
Add a Grafana dashboard for the four KMS backend operation metrics
emitted at the policy choke point, Prometheus alert rules with
conservative default thresholds (pending staging baseline calibration),
and an operations runbook documenting the metric contract and the
response procedure for each alert. Register the new alert rules file in
the observability READMEs.

Refs rustfs/backlog#1584 (part of rustfs/backlog#1562)
2026-07-31 21:28:10 +00:00
Zhengchao An 74019845c4 fix(kms): drop stale KmsClient trait import in local_export tests (#5519)
The KmsClient trait was folded into KmsBackend (#5501), but the backup
export tests merged afterwards (#5499) still imported it, breaking the
crate's test build; create_key is an inherent LocalKmsClient method, so
the import is simply unused.
2026-07-31 20:24:04 +00:00
houseme 04c5921850 fix(s3): allow CopyObject COPY request metadata (#5514)
* fix(s3): allow CopyObject COPY request metadata

Co-Authored-By: heihutu <heihutu@gmail.com>

* fix(kms): remove stale KmsClient import

Co-Authored-By: heihutu <heihutu@gmail.com>

---------

Co-authored-by: heihutu <heihutu@gmail.com>
2026-07-31 17:52:00 +00:00
Zhengchao An db8039dece feat(kms): export local backend key material as sealed backup bundles (#5499)
* feat(kms): export local backend key material as sealed backup bundles

Adds the producer side of the Local backup series on top of the #5483
contract: a directory-wide export fence gives the snapshot a single
consistent generation, every artifact is AEAD-wrapped under a
caller-supplied backup KEK that is separate from the business trust
hierarchy, and the sealed manifest with completeness marker is written
last so an interrupted export can never be mistaken for a restorable
bundle. Restore and the admin API land in follow-up changes.

* fix(kms): verify manifest digest against raw bytes, not re-serialized fields

The decode path recomputed the digest by re-serializing the parsed
manifest, which silently assumes every field's stored spelling survives
a parse-and-reprint round trip. Timestamps do not guarantee that: the
time zone annotation jiff emits depends on the host (IANA name, POSIX
TZ string, Etc/Unknown), and the legacy-compat parser rewrites
bracket-less spellings to +00:00[UTC]. On CI this made freshly written
bundles fail digest verification while passing locally.

Digest verification now operates on the raw stored bytes, normalized
only through the JSON value layer with the digest slot emptied in
place; parsed typed fields are never re-serialized on the decode path.
Sealing uses the same value-layer canonical form, and the export
additionally pins created_at to UTC so bundles are host-independent. A
regression test seals a manifest whose created_at spelling cannot
round-trip and proves decoding still verifies. One behavior sharpens:
inserting an explicit null reserved slot after sealing is now rejected
as a digest mismatch instead of being tolerated.

* fix(kms): make manifest digest canonicalization independent of map ordering

The canonical digest form serialized serde_json values directly, which
inherits the key order of serde_json's map type: sorted by default, but
insertion-ordered when any crate in the unified build graph enables the
preserve_order feature. The workspace-wide CI build unified that
feature while a per-crate local build did not, so the frozen fixture
digest matched in one environment and not the other — and a bundle
sealed by one build flavor would fail verification in the other.

Canonicalization now rebuilds every JSON object with bytewise-sorted
keys (array order preserved) before hashing, so the digest bytes are
identical regardless of feature unification. Reproduced by enabling
preserve_order in dev-dependencies (fixture test red), then verified
green with the fix under both map flavors.
2026-07-31 23:57:39 +08:00
houseme 40eee6177a test: add formal hotpath ABBA runner (#5507)
Co-authored-by: heihutu <heihutu@gmail.com>
2026-07-31 11:33:10 +00:00
218 changed files with 37054 additions and 2997 deletions
+2
View File
@@ -28,6 +28,8 @@ script-tests: ## Run shell script tests
./scripts/test_entrypoint_credentials.sh ./scripts/test_entrypoint_credentials.sh
./scripts/test_internode_grpc_ab_bench.sh ./scripts/test_internode_grpc_ab_bench.sh
./scripts/test_object_batch_bench_enhanced.sh ./scripts/test_object_batch_bench_enhanced.sh
./scripts/test_hotpath_warp_ab_gate.sh
./scripts/test_hotpath_warp_abba.sh
./scripts/test_exact_1mib_handoff_abba.sh ./scripts/test_exact_1mib_handoff_abba.sh
./scripts/test_pinned_paired_abba_bench.sh ./scripts/test_pinned_paired_abba_bench.sh
./scripts/test_manual_transition_runbooks.sh ./scripts/test_manual_transition_runbooks.sh
-4
View File
@@ -300,9 +300,6 @@ path = "junit.xml"
# negative-path siblings of each family stay in as regression guards. # negative-path siblings of each family stay in as regression guards.
# * rustfs#4843 — over-limit archive entry paths hard-reject the whole # * rustfs#4843 — over-limit archive entry paths hard-reject the whole
# archive even under ignore-errors semantics. # archive even under ignore-errors semantics.
# * rustfs#4846 — distributed-lock quorum tests misclassify as timeout
# under parallel load (multi-node in-process clusters; natural home is
# ci-7's nightly cluster lane).
[profile.e2e-full] [profile.e2e-full]
default-filter = """ default-filter = """
package(e2e_test) package(e2e_test)
@@ -311,7 +308,6 @@ default-filter = """
& !test(/^replication_extension_test::/) & !test(/^replication_extension_test::/)
& !test(/^multipart_auth_test::test_signed_put_object_extract_skips_invalid_entry_when_ignore_errors_enabled$/) & !test(/^multipart_auth_test::test_signed_put_object_extract_skips_invalid_entry_when_ignore_errors_enabled$/)
& !test(/^snowball_auto_extract_test::tests::snowball_auto_extract_(ignores_invalid_entries_when_requested|supports_standard_headers_with_combined_extract_options)$/) & !test(/^snowball_auto_extract_test::tests::snowball_auto_extract_(ignores_invalid_entries_when_requested|supports_standard_headers_with_combined_extract_options)$/)
& !test(/^reliant::lock::test_distributed_lock_(2_nodes_grpc_read_survives_failed_node|4_nodes_grpc_read_write_quorum_split_with_two_failed_nodes)$/)
""" """
fail-fast = false fail-fast = false
+10
View File
@@ -60,6 +60,16 @@ The file `prometheus-rules/rustfs-get-optimization-alerts.yaml` contains pre-con
| `CodecStreamingFallbackSpike` | Warning | Codec streaming fallback > 10x baseline for 10m | | `CodecStreamingFallbackSpike` | Warning | Codec streaming fallback > 10x baseline for 10m |
| `IoQueueSaturation` | Warning | IO queue utilization > 90% for 5m | | `IoQueueSaturation` | Warning | IO queue utilization > 90% for 5m |
The file `prometheus-rules/rustfs-kms-alerts.yml` contains alerting rules for the KMS backend operation metrics. Thresholds are conservative defaults pending staging baseline calibration; response procedures live in `docs/operations/kms-observability-runbook.md`, and the matching dashboard is `deploy/observability/grafana/rustfs-kms-observability.json`.
| Alert | Severity | Condition |
|-------|----------|-----------|
| `KmsBackendFatalErrors` | Critical | Fatal (non-retryable) attempt failures > 0 for 5m |
| `KmsBackendHighErrorRate` | Critical | Non-success operation ratio > 5% for 10m (with traffic guard) |
| `KmsBackendP99LatencyHigh` | Warning | Operation p99 duration (incl. retries) > 2s for 10m |
| `KmsBackendAttemptFailureSpike` | Warning | Attempt failure rate > 0.5/s for 10m |
| `KmsBackendRetryBudgetExhausted` | Warning | budget_exhausted / deadline_exceeded outcomes > 0.05/s for 10m |
### Enabling Alert Rules ### Enabling Alert Rules
Add the alert rules file to your Prometheus configuration: Add the alert rules file to your Prometheus configuration:
+10
View File
@@ -60,6 +60,16 @@
| `CodecStreamingFallbackSpike` | 警告 | Codec streaming 回退 > 10x 基线,持续 10 分钟 | | `CodecStreamingFallbackSpike` | 警告 | Codec streaming 回退 > 10x 基线,持续 10 分钟 |
| `IoQueueSaturation` | 警告 | IO 队列利用率 > 90%,持续 5 分钟 | | `IoQueueSaturation` | 警告 | IO 队列利用率 > 90%,持续 5 分钟 |
文件 `prometheus-rules/rustfs-kms-alerts.yml` 包含 KMS 后端操作指标的告警规则。阈值为保守默认值,待 staging 基线校准;响应流程见 `docs/operations/kms-observability-runbook.md`,配套仪表盘为 `deploy/observability/grafana/rustfs-kms-observability.json`
| 告警 | 级别 | 条件 |
|------|------|------|
| `KmsBackendFatalErrors` | 严重 | fatal(不可重试)尝试失败 > 0,持续 5 分钟 |
| `KmsBackendHighErrorRate` | 严重 | 非 success 操作占比 > 5%,持续 10 分钟(含流量下限保护) |
| `KmsBackendP99LatencyHigh` | 警告 | 操作 p99 耗时(含重试)> 2s,持续 10 分钟 |
| `KmsBackendAttemptFailureSpike` | 警告 | 尝试失败率 > 0.5/s,持续 10 分钟 |
| `KmsBackendRetryBudgetExhausted` | 警告 | budget_exhausted / deadline_exceeded 结果 > 0.05/s,持续 10 分钟 |
### 启用告警规则 ### 启用告警规则
在 Prometheus 配置中添加告警规则文件: 在 Prometheus 配置中添加告警规则文件:
@@ -0,0 +1,188 @@
# Copyright 2024 RustFS Team
#
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
# =============================================================================
# RustFS KMS backend — Prometheus alerting rules
# =============================================================================
#
# Metric source: the KMS operation-policy choke point in
# crates/kms/src/policy.rs. All label values are static enum strings
# (operation, op_class, outcome, error_class); key identifiers, key material,
# and tokens never appear in labels.
#
# Response procedures: docs/operations/kms-observability-runbook.md
#
# IMPORTANT — threshold status: every numeric threshold below is a
# conservative default chosen without a production baseline. Calibrate against
# a staging baseline before relying on these alerts for paging, and prefer
# loosening over tightening until the baseline exists. Formal SLO targets are
# deliberately not encoded here (see rustfs/backlog#1584).
#
# NOTE: prometheus.yml loads /etc/prometheus/rules/*.yml — keep the .yml
# extension or the file is silently ignored by the docker-compose stack.
#
# Validate: promtool check rules rustfs-kms-alerts.yml
# =============================================================================
groups:
# ==========================================================================
# Critical alerts — immediate action required
# ==========================================================================
- name: rustfs-kms-critical
interval: 30s
rules:
# ------------------------------------------------------------------
# 1. KmsBackendFatalErrors
# Any attempt failure classified as fatal (non-retryable): auth
# or permission errors, malformed requests, missing keys. The
# policy never retries these, so even a low rate means real
# operations are failing right now.
# ------------------------------------------------------------------
- alert: KmsBackendFatalErrors
expr: |
sum by (operation) (rate(rustfs_kms_backend_attempt_failures_total{error_class="fatal"}[5m])) > 0
for: 5m
labels:
severity: critical
component: kms
annotations:
summary: "KMS backend fatal errors on operation {{ $labels.operation }}"
description: >-
Attempt failures classified as fatal are occurring at
{{ $value | printf "%.3f" }}/s on operation
{{ $labels.operation }}. Fatal failures are not retried:
each one is a KMS backend call that failed permanently
(authentication, permissions, malformed request, or a
missing key/version).
runbook_url: "https://github.com/rustfs/rustfs/blob/main/docs/operations/kms-observability-runbook.md#kmsbackendfatalerrors"
# ------------------------------------------------------------------
# 2. KmsBackendHighErrorRate
# Sustained share of operations terminating without success
# (fatal, budget_exhausted, deadline_exceeded). The cancelled
# outcome is excluded because shutdowns legitimately produce it.
# The traffic guard keeps a single failure on a near-idle
# cluster from firing the alert.
# Threshold: 5% for 10m — conservative default, calibrate
# against a staging baseline.
# ------------------------------------------------------------------
- alert: KmsBackendHighErrorRate
expr: |
(
sum(rate(rustfs_kms_backend_operations_total{outcome!~"success|cancelled"}[5m]))
/
clamp_min(sum(rate(rustfs_kms_backend_operations_total[5m])), 1e-9)
) > 0.05
and
sum(rate(rustfs_kms_backend_operations_total[5m])) > 0.02
for: 10m
labels:
severity: critical
component: kms
annotations:
summary: "KMS backend non-success ratio above 5% for 10m"
description: >-
{{ $value | humanizePercentage }} of KMS backend operations
are terminating in fatal, budget_exhausted, or
deadline_exceeded. Object encryption and decryption paths
depending on the KMS are degraded or failing.
runbook_url: "https://github.com/rustfs/rustfs/blob/main/docs/operations/kms-observability-runbook.md#kmsbackendhigherrorrate"
# ==========================================================================
# Warning alerts — investigation needed
# ==========================================================================
- name: rustfs-kms-warning
interval: 30s
rules:
# ------------------------------------------------------------------
# 3. KmsBackendP99LatencyHigh
# p99 wall-clock duration of whole operations (attempts plus
# backoff) is sustained above 2 seconds. Because the histogram
# includes retries, a high p99 usually means the retry policy
# is absorbing backend failures, not that every call is slow.
# Threshold: 2s for 10m — conservative default, calibrate
# against a staging baseline.
# ------------------------------------------------------------------
- alert: KmsBackendP99LatencyHigh
expr: |
histogram_quantile(0.99,
sum by (le) (rate(rustfs_kms_backend_operation_duration_seconds_bucket[5m]))
) > 2
for: 10m
labels:
severity: warning
component: kms
annotations:
summary: "KMS backend operation p99 latency above 2s for 10m"
description: >-
The 99th-percentile KMS backend operation duration is
{{ $value | humanizeDuration }}, including retries and
backoff. Encryption and decryption latency is leaking into
S3 request latency.
runbook_url: "https://github.com/rustfs/rustfs/blob/main/docs/operations/kms-observability-runbook.md#kmsbackendp99latencyhigh"
# ------------------------------------------------------------------
# 4. KmsBackendAttemptFailureSpike
# Aggregate attempt-failure rate (all error classes) sustained
# above an absolute floor. An absolute threshold is used instead
# of an offset-1d baseline ratio because fresh deployments have
# no baseline and an empty offset vector would keep a ratio
# alert from ever firing; switch to a baseline-relative form
# (see rustfs-get-optimization-alerts.yaml for the pattern)
# once a stable staging baseline exists.
# Threshold: 0.5/s for 10m — conservative default, calibrate
# against a staging baseline.
# ------------------------------------------------------------------
- alert: KmsBackendAttemptFailureSpike
expr: |
sum(rate(rustfs_kms_backend_attempt_failures_total[5m])) > 0.5
for: 10m
labels:
severity: warning
component: kms
annotations:
summary: "KMS backend attempt failures above 0.5/s for 10m"
description: >-
KMS backend attempts are failing at
{{ $value | printf "%.2f" }}/s across all error classes.
The retry policy may still be masking these from callers —
check the error-class breakdown before it stops absorbing
them.
runbook_url: "https://github.com/rustfs/rustfs/blob/main/docs/operations/kms-observability-runbook.md#kmsbackendattemptfailurespike"
# ------------------------------------------------------------------
# 5. KmsBackendRetryBudgetExhausted
# Operations are running out of retry budget (budget_exhausted)
# or operation deadline (deadline_exceeded). These surface to
# callers as failed KMS operations even though every individual
# failure was retryable — the backend is unhealthy for longer
# than the policy can bridge.
# Threshold: 0.05/s for 10m — conservative default, calibrate
# against a staging baseline.
# ------------------------------------------------------------------
- alert: KmsBackendRetryBudgetExhausted
expr: |
sum by (outcome) (rate(rustfs_kms_backend_operations_total{outcome=~"budget_exhausted|deadline_exceeded"}[5m])) > 0.05
for: 10m
labels:
severity: warning
component: kms
annotations:
summary: "KMS backend operations exhausting retry budget ({{ $labels.outcome }})"
description: >-
KMS backend operations are terminating as
{{ $labels.outcome }} at {{ $value | printf "%.3f" }}/s.
Retryable failures are outlasting the retry budget, so
callers are seeing hard failures.
runbook_url: "https://github.com/rustfs/rustfs/blob/main/docs/operations/kms-observability-runbook.md#kmsbackendretrybudgetexhausted"
+43 -12
View File
@@ -25,9 +25,13 @@ inputs:
required: false required: false
default: "rustfs-deps" default: "rustfs-deps"
cache-save-if: cache-save-if:
description: "Condition for saving cache" description: >-
Whether to save the cache. The fail-safe default is 'false': a caller that
wants to populate a cache must opt in explicitly, so a forgotten input
costs a cold cache (minutes) rather than silently consuming the
repository-wide 10GB Actions cache quota and evicting other lanes.
required: false required: false
default: "true" default: "false"
install-cross-tools: install-cross-tools:
description: "Install cross-compilation tools" description: "Install cross-compilation tools"
required: false required: false
@@ -36,28 +40,43 @@ inputs:
description: "Target architecture to add" description: "Target architecture to add"
required: false required: false
default: "" default: ""
github-token: install-build-packaging-tools:
description: "GitHub token for API access" description: >-
Install musl-tools/zip/unzip, needed for musl linking and release
packaging. Off for CI test lanes, which use none of them.
required: false required: false
default: "" default: "true"
install-test-tools:
description: >-
Install cargo-nextest and the rustfmt/clippy components. Off for release
and audit lanes, which run no tests and no lints.
required: false
default: "true"
runs: runs:
using: "composite" using: "composite"
steps: steps:
# protobuf-compiler is deliberately absent: the setup-protoc step below
# installs 34.1 into the tool cache and prepends it to PATH, so the apt
# build (older, and never version-matched) was shadowed on every run and
# simply never used.
- name: Install system dependencies (Ubuntu) - name: Install system dependencies (Ubuntu)
if: runner.os == 'Linux' if: runner.os == 'Linux'
shell: bash shell: bash
run: | run: |
sudo apt-get update sudo apt-get update
sudo apt-get install -y \ sudo apt-get install -y \
musl-tools \
build-essential \ build-essential \
pkg-config \ pkg-config \
libssl-dev \ libssl-dev \
ripgrep \ ripgrep
unzip \
zip \ # musl-gcc is needed by the native musl release leg, and zip/unzip by the
protobuf-compiler # release packaging steps. No CI test lane touches any of them.
- name: Install packaging and cross-linking dependencies (Ubuntu)
if: runner.os == 'Linux' && inputs.install-build-packaging-tools == 'true'
shell: bash
run: sudo apt-get install -y musl-tools zip unzip
- name: Install protoc - name: Install protoc
uses: rustfs/setup-protoc@a3705324d8f9bf5b6c3573fb6cf8ae421db55dd6 # v3.0.1 uses: rustfs/setup-protoc@a3705324d8f9bf5b6c3573fb6cf8ae421db55dd6 # v3.0.1
@@ -75,7 +94,7 @@ runs:
with: with:
toolchain: ${{ inputs.rust-version }} toolchain: ${{ inputs.rust-version }}
targets: ${{ inputs.target }} targets: ${{ inputs.target }}
components: rustfmt, clippy components: ${{ inputs.install-test-tools == 'true' && 'rustfmt, clippy' || '' }}
- name: Install Zig - name: Install Zig
if: inputs.install-cross-tools == 'true' if: inputs.install-cross-tools == 'true'
@@ -86,12 +105,24 @@ runs:
uses: taiki-e/install-action@a21ae4029b089b9ddc45704028756f51ab8abe48 # cargo-zigbuild uses: taiki-e/install-action@a21ae4029b089b9ddc45704028756f51ab8abe48 # cargo-zigbuild
- name: Install cargo-nextest - name: Install cargo-nextest
if: inputs.install-test-tools == 'true'
uses: taiki-e/install-action@96c7780c1d8a2b8723e12031def873a434d39d8d # nextest uses: taiki-e/install-action@96c7780c1d8a2b8723e12031def873a434d39d8d # nextest
- name: Setup Rust cache - name: Setup Rust cache
uses: Swatinem/rust-cache@e18b497796c12c097a38f9edb9d0641fb99eee32 # v2 uses: Swatinem/rust-cache@e18b497796c12c097a38f9edb9d0641fb99eee32 # v2
with: with:
cache-all-crates: true # false is rust-cache's own default. With true, cleanup.ts returns
# *before* pruning ~/.cargo/registry/src, and config.ts archives the
# whole registry — so every cache carried the unpacked source tree of
# every dependency, not just "a few extra crates".
#
# No coverage is lost: getPackages runs `cargo metadata --all-features`,
# a strict superset of any single lane's feature closure, and -sys crates
# are explicitly exempted from pruning (their src timestamps would
# otherwise trigger rebuilds). Anything pruned is re-unpacked from the
# .crate files still in registry/cache, whose mtimes crates.io
# normalises, so cargo fingerprints stay valid.
cache-all-crates: false
cache-on-failure: true cache-on-failure: true
shared-key: ${{ inputs.cache-shared-key }} shared-key: ${{ inputs.cache-shared-key }}
save-if: ${{ inputs.cache-save-if }} save-if: ${{ inputs.cache-save-if }}
@@ -37,6 +37,7 @@ jobs:
name: Cancel Closed PR Runs name: Cancel Closed PR Runs
if: github.event_name == 'pull_request' && github.event.action == 'closed' if: github.event_name == 'pull_request' && github.event.action == 'closed'
runs-on: ubuntu-latest runs-on: ubuntu-latest
timeout-minutes: 10
steps: steps:
- name: Explain cancellation run - name: Explain cancellation run
run: echo "PR closed; this run only cancels older runs in the same concurrency group." run: echo "PR closed; this run only cancels older runs in the same concurrency group."
@@ -45,8 +46,11 @@ jobs:
name: Architecture Migration Rules name: Architecture Migration Rules
if: github.event_name != 'pull_request' || github.event.action != 'closed' if: github.event_name != 'pull_request' || github.event.action != 'closed'
runs-on: ubuntu-latest runs-on: ubuntu-latest
timeout-minutes: 10
steps: steps:
- uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7 - uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
with:
persist-credentials: false
- name: Install ripgrep - name: Install ripgrep
run: | run: |
+71 -5
View File
@@ -39,7 +39,12 @@ on:
- 'scripts/security/check_preview_release_workflow.sh' - 'scripts/security/check_preview_release_workflow.sh'
- 'scripts/security/check_workflow_pins.sh' - 'scripts/security/check_workflow_pins.sh'
schedule: schedule:
- cron: '0 3 * * 0' # Weekly on Sunday 03:00 UTC (staggered after the midnight ci/build crons) # Daily, not weekly. This schedule exists to catch RustSec advisories
# published against an unchanged dependency tree; at weekly cadence a new
# advisory could sit unnoticed for seven days. The check list is unchanged —
# splitting it into a light daily advisories-only run and a weekly full run
# would create runs where sources/bans/licenses go unverified.
- cron: '0 3 * * *' # Daily 03:00 UTC (staggered after the midnight ci/build crons)
workflow_dispatch: workflow_dispatch:
permissions: permissions:
@@ -59,6 +64,7 @@ jobs:
name: Cancel Closed PR Runs name: Cancel Closed PR Runs
if: github.event_name == 'pull_request' && github.event.action == 'closed' if: github.event_name == 'pull_request' && github.event.action == 'closed'
runs-on: ubuntu-latest runs-on: ubuntu-latest
timeout-minutes: 10
steps: steps:
- name: Explain cancellation run - name: Explain cancellation run
run: echo "PR closed; this run only cancels older runs in the same concurrency group." run: echo "PR closed; this run only cancels older runs in the same concurrency group."
@@ -74,11 +80,32 @@ jobs:
steps: steps:
- name: Checkout repository - name: Checkout repository
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7 uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
- name: Setup Rust environment
uses: ./.github/actions/setup
with: with:
cache-shared-key: rustfs-cargo-deny persist-credentials: false
# cargo-deny compiles nothing, so the full setup composite (apt packages,
# protoc, flatc, nextest, rustfmt/clippy) was pure overhead here. It does
# still need a real cargo: `cargo deny check` runs `cargo metadata`, and
# Cargo.toml pins datafusion and s3s as git dependencies, which must be
# materialised into ~/.cargo/git — a cold clone is hundreds of MB, so the
# cache stays.
- name: Install Rust toolchain
uses: dtolnay/rust-toolchain@29eef336d9b2848a0b548edc03f92a220660cdb8 # stable
# Was relying on the composite's default, which used to be "true": every
# PR touching Cargo.toml/Cargo.lock saved a second, PR-scoped copy of this
# cache and pushed the main-scoped lanes out of the 10GB quota. The
# default is now "false", but state it explicitly — see
# scripts/security/check_cache_save_if.sh.
- name: Setup Rust cache
uses: Swatinem/rust-cache@e18b497796c12c097a38f9edb9d0641fb99eee32 # v2
with:
# Same reasoning as the setup composite: true archives every
# dependency's unpacked source tree.
cache-all-crates: false
cache-on-failure: true
shared-key: rustfs-cargo-deny
save-if: ${{ github.ref == 'refs/heads/main' }}
- name: Install cargo-deny - name: Install cargo-deny
uses: taiki-e/install-action@bffeee26d4db9be238a4ea78d8826604ebcb594d # v2 uses: taiki-e/install-action@bffeee26d4db9be238a4ea78d8826604ebcb594d # v2
@@ -96,16 +123,28 @@ jobs:
steps: steps:
- name: Checkout repository - name: Checkout repository
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7 uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
with:
persist-credentials: false
- name: Report unpinned GitHub Actions - name: Report unpinned GitHub Actions
run: ./scripts/security/check_workflow_pins.sh --enforce run: ./scripts/security/check_workflow_pins.sh --enforce
- name: Check setup cache-save-if is explicit
run: ./scripts/security/check_cache_save_if.sh
- name: Check every job declares a timeout
run: ./scripts/security/check_job_timeouts.sh
- name: Check checkouts clear their credentials
run: ./scripts/security/check_persist_credentials.sh
- name: Check preview release workflow policy - name: Check preview release workflow policy
run: ./scripts/security/check_preview_release_workflow.sh run: ./scripts/security/check_preview_release_workflow.sh
dependency-review: dependency-review:
name: Dependency Review name: Dependency Review
runs-on: ubuntu-latest runs-on: ubuntu-latest
timeout-minutes: 30
if: github.event_name == 'pull_request' && github.event.action != 'closed' if: github.event_name == 'pull_request' && github.event.action != 'closed'
permissions: permissions:
contents: read contents: read
@@ -113,6 +152,8 @@ jobs:
steps: steps:
- name: Checkout repository - name: Checkout repository
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7 uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
with:
persist-credentials: false
- name: Dependency Review - name: Dependency Review
uses: actions/dependency-review-action@a1d282b36b6f3519aa1f3fc636f609c47dddb294 # v5 uses: actions/dependency-review-action@a1d282b36b6f3519aa1f3fc636f609c47dddb294 # v5
@@ -125,3 +166,28 @@ jobs:
# conscious re-review of the license/provenance claim (backlog#1181). # conscious re-review of the license/provenance claim (backlog#1181).
allow-dependencies-licenses: pkg:cargo/rustfs-uring@0.1.0 allow-dependencies-licenses: pkg:cargo/rustfs-uring@0.1.0
comment-summary-in-pr: always comment-summary-in-pr: always
alert-on-failure:
name: Alert on scheduled failure
# dependency-review is deliberately excluded: it only runs on pull_request,
# so it can never contribute a failure to a scheduled run.
needs: [cargo-deny, workflow-pin-report]
# A scheduled cargo-deny failure usually means the dependency tree just
# matched a newly published advisory — the single most important signal this
# workflow produces, and until now it was only visible to whoever happened to
# open the Actions tab. Same ci-8 mechanism coverage.yml and
# e2e-replication-nightly.yml already use.
if: always() && github.event_name == 'schedule' && contains(needs.*.result, 'failure')
runs-on: ubuntu-latest
timeout-minutes: 10
permissions:
contents: read
issues: write
steps:
- uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
with:
persist-credentials: false
- name: Open or update failure-tracking issue
uses: ./.github/actions/schedule-failure-issue
with:
github-token: ${{ secrets.GITHUB_TOKEN }}
+40 -5
View File
@@ -50,12 +50,18 @@ on:
- "**/*.svg" - "**/*.svg"
- ".gitignore" - ".gitignore"
- ".dockerignore" - ".dockerignore"
- "flake.lock"
schedule: schedule:
- cron: "0 1 * * 0" # Weekly on Sunday 01:00 UTC (staggered after the ci.yml midnight cron) - cron: "0 1 * * 0" # Weekly on Sunday 01:00 UTC (staggered after the ci.yml midnight cron)
workflow_dispatch: workflow_dispatch:
inputs: inputs:
build_docker: build_docker:
description: "Build and push Docker images after binary build" # Advisory only. docker.yml triggers on workflow_run and its job-level
# condition requires the triggering event to be a tag push, so a manual
# dispatch of this workflow never produces images regardless of this
# value. Kept because the summary step reports it; wiring it up would
# mean teaching docker.yml's version parser a second event shape.
description: "Build and push Docker images after binary build (ignored: dispatch runs never reach docker.yml)"
required: false required: false
default: true default: true
type: boolean type: boolean
@@ -83,6 +89,7 @@ jobs:
build-check: build-check:
name: Build Strategy Check name: Build Strategy Check
runs-on: ubuntu-latest runs-on: ubuntu-latest
timeout-minutes: 30
outputs: outputs:
should_build: ${{ steps.check.outputs.should_build }} should_build: ${{ steps.check.outputs.should_build }}
build_type: ${{ steps.check.outputs.build_type }} build_type: ${{ steps.check.outputs.build_type }}
@@ -92,6 +99,8 @@ jobs:
steps: steps:
- name: Checkout repository - name: Checkout repository
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7 uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
with:
persist-credentials: false
- name: Determine build strategy - name: Determine build strategy
id: check id: check
@@ -164,6 +173,7 @@ jobs:
name: Prepare Platform Matrix name: Prepare Platform Matrix
needs: build-check needs: build-check
runs-on: ubuntu-latest runs-on: ubuntu-latest
timeout-minutes: 30
outputs: outputs:
matrix: ${{ steps.select.outputs.matrix }} matrix: ${{ steps.select.outputs.matrix }}
selected: ${{ steps.select.outputs.selected }} selected: ${{ steps.select.outputs.selected }}
@@ -171,10 +181,14 @@ jobs:
- name: Select target platforms - name: Select target platforms
id: select id: select
shell: bash shell: bash
env:
# via env, not interpolation: a dispatch input is free-form text and
# would otherwise be pasted into the script for bash to evaluate.
RAW_PLATFORMS: ${{ github.event_name == 'workflow_dispatch' && github.event.inputs.platforms || 'all' }}
run: | run: |
set -euo pipefail set -euo pipefail
selected="${{ github.event_name == 'workflow_dispatch' && github.event.inputs.platforms || 'all' }}" selected="$RAW_PLATFORMS"
selected="$(echo "${selected}" | tr -d '[:space:]')" selected="$(echo "${selected}" | tr -d '[:space:]')"
if [[ -z "${selected}" ]]; then if [[ -z "${selected}" ]]; then
selected="all" selected="all"
@@ -245,6 +259,7 @@ jobs:
- name: Checkout repository - name: Checkout repository
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7 uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
with: with:
persist-credentials: false
fetch-depth: 0 fetch-depth: 0
- name: Setup Rust environment - name: Setup Rust environment
@@ -253,9 +268,17 @@ jobs:
rust-version: stable rust-version: stable
target: ${{ matrix.target }} target: ${{ matrix.target }}
cache-shared-key: build-${{ matrix.target }} cache-shared-key: build-${{ matrix.target }}
github-token: ${{ secrets.GITHUB_TOKEN }} # main only. A cache saved on refs/tags/X is scoped to that tag: no
cache-save-if: ${{ github.ref == 'refs/heads/main' || startsWith(github.ref, 'refs/tags/') }} # other tag, no main run and no PR can restore it, so every release
# cycle wrote up to 12 entries of 1-2GB (preview tag plus final tag,
# six legs each) that nobody could read, evicting the hot lanes from
# the repo-wide 10GB quota. Tag builds still restore the main-scoped
# cache, since default-branch caches are readable from every ref.
# The one real cost: re-running a failed leg of the same tag no longer
# finds that tag's own warm cache and falls back to main's.
cache-save-if: ${{ github.ref == 'refs/heads/main' }}
install-cross-tools: ${{ matrix.cross }} install-cross-tools: ${{ matrix.cross }}
install-test-tools: 'false'
- name: Download static console assets - name: Download static console assets
shell: bash shell: bash
@@ -702,9 +725,14 @@ jobs:
needs: [ build-check, build-rustfs ] needs: [ build-check, build-rustfs ]
if: always() && needs.build-check.outputs.should_build == 'true' if: always() && needs.build-check.outputs.should_build == 'true'
runs-on: ubuntu-latest runs-on: ubuntu-latest
timeout-minutes: 10
steps: steps:
- name: Build completion summary - name: Build completion summary
shell: bash shell: bash
env:
# dispatch input via env: free-form text must not be pasted into the
# script for bash to evaluate.
INPUT_BUILD_DOCKER: ${{ github.event.inputs.build_docker }}
run: | run: |
BUILD_TYPE="${{ needs.build-check.outputs.build_type }}" BUILD_TYPE="${{ needs.build-check.outputs.build_type }}"
VERSION="${{ needs.build-check.outputs.version }}" VERSION="${{ needs.build-check.outputs.version }}"
@@ -746,7 +774,7 @@ jobs:
echo "🐳 Docker Images:" echo "🐳 Docker Images:"
if [[ "$BUILD_TYPE" == "preview" ]]; then if [[ "$BUILD_TYPE" == "preview" ]]; then
echo "⏭️ Preview tags do not publish Docker images" echo "⏭️ Preview tags do not publish Docker images"
elif [[ "${{ github.event.inputs.build_docker }}" == "false" ]]; then elif [[ "$INPUT_BUILD_DOCKER" == "false" ]]; then
echo "⏭️ Docker image build was skipped (binary only build)" echo "⏭️ Docker image build was skipped (binary only build)"
elif [[ "$BUILD_STATUS" == "success" ]]; then elif [[ "$BUILD_STATUS" == "success" ]]; then
echo "🔄 Docker images will be built and pushed automatically via workflow_run event" echo "🔄 Docker images will be built and pushed automatically via workflow_run event"
@@ -760,6 +788,7 @@ jobs:
needs: [ build-check, build-rustfs ] needs: [ build-check, build-rustfs ]
if: startsWith(github.ref, 'refs/tags/') && (needs.build-check.outputs.build_type == 'preview' || needs.build-check.outputs.build_type == 'release' || needs.build-check.outputs.build_type == 'prerelease') if: startsWith(github.ref, 'refs/tags/') && (needs.build-check.outputs.build_type == 'preview' || needs.build-check.outputs.build_type == 'release' || needs.build-check.outputs.build_type == 'prerelease')
runs-on: ubuntu-latest runs-on: ubuntu-latest
timeout-minutes: 30
permissions: permissions:
contents: write contents: write
outputs: outputs:
@@ -769,6 +798,7 @@ jobs:
- name: Checkout repository - name: Checkout repository
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7 uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
with: with:
persist-credentials: false
fetch-depth: 0 fetch-depth: 0
- name: Create GitHub Release - name: Create GitHub Release
@@ -819,12 +849,15 @@ jobs:
needs: [ build-check, build-rustfs, create-release ] needs: [ build-check, build-rustfs, create-release ]
if: startsWith(github.ref, 'refs/tags/') && (needs.build-check.outputs.build_type == 'preview' || needs.build-check.outputs.build_type == 'release' || needs.build-check.outputs.build_type == 'prerelease') if: startsWith(github.ref, 'refs/tags/') && (needs.build-check.outputs.build_type == 'preview' || needs.build-check.outputs.build_type == 'release' || needs.build-check.outputs.build_type == 'prerelease')
runs-on: ubuntu-latest runs-on: ubuntu-latest
timeout-minutes: 30
permissions: permissions:
contents: write contents: write
actions: read actions: read
steps: steps:
- name: Checkout repository - name: Checkout repository
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7 uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
with:
persist-credentials: false
- name: Download all build artifacts - name: Download all build artifacts
uses: actions/download-artifact@37930b1c2abaa49bbe596cd826c3c89aef350131 # v7 uses: actions/download-artifact@37930b1c2abaa49bbe596cd826c3c89aef350131 # v7
@@ -910,6 +943,7 @@ jobs:
needs: [ build-check, publish-release ] needs: [ build-check, publish-release ]
if: startsWith(github.ref, 'refs/tags/') && (needs.build-check.outputs.build_type == 'release' || needs.build-check.outputs.build_type == 'prerelease') if: startsWith(github.ref, 'refs/tags/') && (needs.build-check.outputs.build_type == 'release' || needs.build-check.outputs.build_type == 'prerelease')
runs-on: ubuntu-latest runs-on: ubuntu-latest
timeout-minutes: 30
steps: steps:
- name: Update latest.json - name: Update latest.json
env: env:
@@ -969,6 +1003,7 @@ jobs:
needs: [ build-check, create-release, upload-release-assets ] needs: [ build-check, create-release, upload-release-assets ]
if: startsWith(github.ref, 'refs/tags/') && (needs.build-check.outputs.build_type == 'preview' || needs.build-check.outputs.build_type == 'release' || needs.build-check.outputs.build_type == 'prerelease') if: startsWith(github.ref, 'refs/tags/') && (needs.build-check.outputs.build_type == 'preview' || needs.build-check.outputs.build_type == 'release' || needs.build-check.outputs.build_type == 'prerelease')
runs-on: ubuntu-latest runs-on: ubuntu-latest
timeout-minutes: 30
permissions: permissions:
contents: write contents: write
steps: steps:
+265
View File
@@ -0,0 +1,265 @@
# Copyright 2026 RustFS Team
#
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
# Sole writer of the Rust dependency caches that ci.yml restores.
#
# Why this is a separate workflow rather than steps inside ci.yml: ci.yml's
# concurrency group cancels in-progress runs on main pushes, and merges land far
# faster than its 70-minute pipeline. Measured over 15 consecutive main pushes:
# 12 cancelled, 2 failed, 0 succeeded. A cancelled run never reaches
# Swatinem/rust-cache's post step (cache-on-failure does not cover cancellation),
# so the writer lanes were saving nothing and every PR paid a cold restore —
# 11.8-20.9 minutes of "Setup Rust environment" against 0.7-3.4 warm.
#
# Splitting cache writing out of the test pipeline lets ci.yml keep cancelling
# superseded runs (which is correct — nobody needs test results for a commit
# that is already three merges behind) while the caches still get written.
#
# The group below deliberately does NOT cancel in progress; see the comment on
# it for how that bounds concurrency and why it is scoped by event.
#
# Each job below owns exactly one shared-key and is the only place that sets
# cache-save-if to anything but 'false' for it; every lane in ci.yml reads.
# scripts/security/check_cache_save_if.sh keeps the declarations explicit.
#
# The builds are supersets of what the reading lanes compile, because a reader
# restores only what the writer saved. Feature resolution matters here: a lane
# built with e2e-test-hooks resolves dependency features differently, which
# changes -Cmetadata, so the plain build does not cover it. See
# rustfs/backlog#1600.
name: Cache Warm
on:
push:
branches: [ main ]
# Mirrors ci.yml's push paths-ignore: if a commit cannot change what ci.yml
# compiles, it cannot change what ci.yml needs restored either.
paths-ignore:
- "**.md"
- "docs/**"
- "deploy/**"
- "scripts/dev_*.sh"
- "scripts/probe.sh"
- "LICENSE*"
- ".gitignore"
- ".dockerignore"
- "README*"
- "**/*.png"
- "**/*.jpg"
- "**/*.svg"
- ".github/workflows/build.yml"
- ".github/workflows/docker.yml"
- ".github/workflows/audit.yml"
workflow_dispatch:
inputs:
emit_timings:
description: >-
Also emit cargo --timings for the ci-dev build and upload it. Used to
decide whether sccache is worth adopting (rustfs/backlog#1601 gate).
required: false
default: false
type: boolean
permissions:
contents: read
# Scoped by event. A push run and a dispatch run do not compete: GitHub keeps
# one running plus one pending per group, so with a single shared group a
# manually dispatched run was displaced as pending by the next merge and
# cancelled — observed three times in a row, which made the --timings gate in
# rustfs/backlog#1601 effectively impossible to trigger while main was busy.
#
# Still no cancel-in-progress: a burst of merges collapses into "current run
# finishes, newest queued run follows" rather than a pile-up, which is what
# bounds this workflow to one self-hosted runner per event type.
#
# The two paths can now overlap and race to save the same key. That is benign:
# the loser finds the key already present and skips, and both builds produce the
# same artifacts from the same commit.
concurrency:
group: cache-warm-${{ github.event_name }}
cancel-in-progress: false
env:
CARGO_TERM_COLOR: always
jobs:
# Readers: test-and-lint, test-ilm-integration-serial, build-rustfs-debug-binary,
# e2e-tests, e2e-full.
warm-ci-dev:
name: Warm ci-dev
runs-on: sm-standard-4
timeout-minutes: 90
env:
FORCE_JAVASCRIPT_ACTIONS_TO_NODE24: "true"
steps:
- name: Checkout repository
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
with:
persist-credentials: false
- name: Setup Rust environment
uses: ./.github/actions/setup
with:
rust-version: stable
cache-shared-key: ci-dev
cache-save-if: 'true'
install-build-packaging-tools: 'false'
# rustfs/backlog#1601 gate. sccache can only cache compilation units whose
# --emit includes link, so it covers workspace rlibs and nothing else:
# clippy is metadata-only, and the ~100 test binaries, the rustfs bin and
# every build script invoke the system linker. Before spending a bucket,
# credentials and a supply-chain boundary on it, measure how much of the
# build is actually rlib codegen.
#
# Read from the report: workspace lib codegen as a share of the build, and
# s3select-query's own rlib as a share. The plan adopts sccache only above
# 50% and 25% respectively; if linking dominates instead, the answer is
# mold/lld plus split-debuginfo, which is exactly the part sccache cannot
# touch. Off by default — this doubles the ci-dev build.
- name: Build ci-dev superset (with --timings)
if: inputs.emit_timings
env:
CARGO_BUILD_JOBS: "2"
run: cargo build --workspace --all-targets --timings
- name: Upload cargo timings report
if: inputs.emit_timings
uses: actions/upload-artifact@b7c566a772e6b6bfb58ed0dc250532a479d7789f # v6
with:
name: cargo-timings-ci-dev
path: target/cargo-timings/
retention-days: 30
if-no-files-found: error
# --all-targets covers the test binaries nextest builds, including
# e2e_test, which test-and-lint's own run excludes. The second build adds
# the e2e-test-hooks feature resolution that build-rustfs-debug-binary uses
# and that no lint lane enables.
- name: Build ci-dev superset
env:
# Same limit ci.yml puts on its nextest step: this builds the same
# ~100 workspace test binaries, and three concurrent links saturate the
# self-hosted runner's overlay I/O and can wedge Cargo (#5394).
CARGO_BUILD_JOBS: "2"
run: |
cargo build --workspace --all-targets
cargo build -p rustfs --bins --features e2e-test-hooks
# Runs before rust-cache's post step, so these are the sizes it is about
# to archive. Reported so the cache-all-crates decision stays evidence-led:
# registry/src is what that flag prunes, registry/cache is what the pruned
# sources are re-unpacked from. See rustfs/backlog#1600.
- name: Report cache input sizes
if: always()
run: |
# tee, not a plain redirect: sent only to $GITHUB_STEP_SUMMARY these
# numbers are readable in the UI but absent from the job log, and the
# REST API exposes the log, not the summary — which made the figures
# unreachable for exactly the scripted comparison they exist for.
sizes="$(du -sh ~/.cargo/registry/src ~/.cargo/registry/cache \
~/.cargo/registry/index ~/.cargo/git target 2>/dev/null || true)"
echo "cache-input-sizes-begin"
printf '%s\n' "$sizes"
echo "cache-input-sizes-end"
{
echo "### Cache input sizes (ci-dev)"
echo '```'
printf '%s\n' "$sizes"
echo '```'
} >> "$GITHUB_STEP_SUMMARY"
# Readers: test-and-lint-rio-v2, build-rustfs-debug-binary-rio-v2.
warm-ci-feat-rio:
name: Warm ci-feat-rio
runs-on: sm-standard-4
timeout-minutes: 90
env:
FORCE_JAVASCRIPT_ACTIONS_TO_NODE24: "true"
steps:
- name: Checkout repository
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
with:
persist-credentials: false
- name: Setup Rust environment
uses: ./.github/actions/setup
with:
rust-version: stable
cache-shared-key: ci-feat-rio
cache-save-if: 'true'
install-build-packaging-tools: 'false'
- name: Build ci-feat-rio superset
run: |
cargo build -p rustfs -p rustfs-ecstore --all-targets --features rio-v2
cargo build -p rustfs --bins --features rio-v2,e2e-test-hooks
# Readers: the swift and sftp legs of test-and-lint-protocols. Built in
# sequence rather than as `--features swift,sftp`, which is a combination no
# lane actually compiles; running both leaves the union in target/.
warm-ci-feat-proto:
name: Warm ci-feat-proto
runs-on: sm-standard-4
timeout-minutes: 90
env:
FORCE_JAVASCRIPT_ACTIONS_TO_NODE24: "true"
steps:
- name: Checkout repository
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
with:
persist-credentials: false
- name: Setup Rust environment
uses: ./.github/actions/setup
with:
rust-version: stable
cache-shared-key: ci-feat-proto
cache-save-if: 'true'
install-build-packaging-tools: 'false'
- name: Build ci-feat-proto superset
run: |
cargo build -p rustfs -p rustfs-protocols --all-targets --features swift
cargo build -p rustfs -p rustfs-protocols --all-targets --features sftp
# Reader: uring-integration. Runs on ubuntu-latest to match it: rust-cache's
# key covers runner.os and arch but not the runner label or image, so a cache
# written on sm-standard-4 would be restored by the hosted runner as if it
# belonged to it.
warm-ci-uring:
name: Warm ci-uring
runs-on: ubuntu-latest
timeout-minutes: 60
steps:
- name: Checkout repository
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
with:
persist-credentials: false
- name: Setup Rust environment
uses: ./.github/actions/setup
with:
rust-version: stable
cache-shared-key: ci-uring
cache-save-if: 'true'
install-build-packaging-tools: 'false'
- name: Install build dependencies
run: sudo apt-get update && sudo apt-get install -y protobuf-compiler
- name: Build ci-uring superset
run: cargo build -p rustfs-ecstore --all-targets
+81 -10
View File
@@ -12,18 +12,24 @@
# See the License for the specific language governing permissions and # See the License for the specific language governing permissions and
# limitations under the License. # limitations under the License.
# Companion to ci.yml for the required "Test and Lint" status check. # Companion to ci.yml for required status checks.
# #
# ci.yml skips docs-only pull requests via paths-ignore, but the branch # ci.yml skips docs-only pull requests via paths-ignore, but the branch ruleset
# ruleset requires a check named "Test and Lint" — without this workflow a # requires a check named "Test and Lint" — without this workflow a docs-only PR
# docs-only PR would wait on that check forever. This workflow triggers on # would wait on it forever. This workflow triggers on exactly the paths ci.yml
# exactly the paths ci.yml ignores and reports an instant success under the # ignores and reports success under the same job name. Mixed PRs trigger both
# same job name. Mixed PRs trigger both workflows and the real check still # workflows and the real check still gates: a required check with any failing
# gates: a required check with any failing run blocks the merge. # run blocks the merge.
# https://docs.github.com/en/repositories/configuring-branches-and-merges-in-your-repository/defining-the-mergeability-of-pull-requests/troubleshooting-required-status-checks#handling-skipped-but-required-checks # https://docs.github.com/en/repositories/configuring-branches-and-merges-in-your-repository/defining-the-mergeability-of-pull-requests/troubleshooting-required-status-checks#handling-skipped-but-required-checks
# #
# "Quick Checks" is mirrored here ahead of the ruleset change that will make it
# required too (rustfs/backlog#1599). Until that change lands this job is
# inert; mirroring it first is what lets the ruleset change happen without
# stranding docs-only PRs on a check nobody reports.
#
# Keep the paths list below in sync with the pull_request paths-ignore list # Keep the paths list below in sync with the pull_request paths-ignore list
# in ci.yml. # in ci.yml, and keep the quick-checks steps below byte-identical to the
# quick-checks job in ci.yml.
name: Continuous Integration (docs only) name: Continuous Integration (docs only)
@@ -47,17 +53,82 @@ on:
- ".github/workflows/build.yml" - ".github/workflows/build.yml"
- ".github/workflows/docker.yml" - ".github/workflows/docker.yml"
- ".github/workflows/audit.yml" - ".github/workflows/audit.yml"
- "flake.lock"
permissions: permissions:
contents: read contents: read
jobs: jobs:
test-and-lint: # Deliberately NOT a bare `echo`. Once "Quick Checks" becomes a required
name: Test and Lint # check, ci.yml gates every expensive job behind it, so a mixed PR reports
# two check runs with this name: the real one (45-51s) and this companion.
# GitHub has no written contract for how it picks between same-named
# required check runs ("latest wins" vs "any failure blocks"), so instead of
# relying on ordering we make both runs execute the same commands against
# the same merge ref — their conclusions are then necessarily identical and
# the choice does not matter. Keep these steps byte-identical to the
# quick-checks job in ci.yml (a guard script that asserts this, and the paths
# sync below, is tracked in rustfs/backlog#1603).
#
# For a genuinely docs-only PR this adds no strictness (no code changed, so
# fmt and the guards always pass) and costs ~50s of ubuntu-latest.
quick-checks:
name: Quick Checks
runs-on: ubuntu-latest runs-on: ubuntu-latest
timeout-minutes: 10
steps: steps:
- name: Checkout repository - name: Checkout repository
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7 uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
with:
persist-credentials: false
- name: Install ripgrep
run: sudo apt-get update && sudo apt-get install -y ripgrep
- name: Install Rust toolchain
uses: dtolnay/rust-toolchain@29eef336d9b2848a0b548edc03f92a220660cdb8 # stable
with:
components: rustfmt
- name: Check code formatting
run: cargo fmt --all --check
- name: Check unsafe code allowances
run: ./scripts/check_unsafe_code_allowances.sh
- name: Check layered dependencies
run: ./scripts/check_layer_dependencies.sh
- name: Check architecture migration rules
run: ./scripts/check_architecture_migration_rules.sh
- name: Check tokio io-uring feature guard
run: ./scripts/check_no_tokio_io_uring.sh
- name: Check extension schema boundaries
run: ./scripts/check_extension_schema_boundaries.sh
- name: Check body-cache whitelist guard
run: ./scripts/check_body_cache_whitelist.sh
- name: Check no planning docs committed
run: ./scripts/check_no_planning_docs.sh
- name: Check CI paths stay in sync
run: ./scripts/check_ci_paths_sync.sh
- name: Check io_uring lane --lib precondition
run: ./scripts/check_uring_lane_lib_only.sh
test-and-lint:
name: Test and Lint
runs-on: ubuntu-latest
timeout-minutes: 10
steps:
- name: Checkout repository
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
with:
persist-credentials: false
# Docs-only PRs skip the full code CI, but they are exactly where a # Docs-only PRs skip the full code CI, but they are exactly where a
# planning-type document could be slipped in (git add -f bypasses # planning-type document could be slipped in (git add -f bypasses
+214 -63
View File
@@ -33,6 +33,7 @@ on:
- ".github/workflows/build.yml" - ".github/workflows/build.yml"
- ".github/workflows/docker.yml" - ".github/workflows/docker.yml"
- ".github/workflows/audit.yml" - ".github/workflows/audit.yml"
- "flake.lock"
pull_request: pull_request:
types: [ opened, synchronize, reopened, closed ] types: [ opened, synchronize, reopened, closed ]
branches: [ main ] branches: [ main ]
@@ -54,6 +55,7 @@ on:
- ".github/workflows/build.yml" - ".github/workflows/build.yml"
- ".github/workflows/docker.yml" - ".github/workflows/docker.yml"
- ".github/workflows/audit.yml" - ".github/workflows/audit.yml"
- "flake.lock"
merge_group: merge_group:
types: [ checks_requested ] types: [ checks_requested ]
schedule: schedule:
@@ -81,6 +83,7 @@ jobs:
name: Cancel Closed PR Runs name: Cancel Closed PR Runs
if: github.event_name == 'pull_request' && github.event.action == 'closed' if: github.event_name == 'pull_request' && github.event.action == 'closed'
runs-on: ubuntu-latest runs-on: ubuntu-latest
timeout-minutes: 10
steps: steps:
- name: Explain cancellation run - name: Explain cancellation run
run: echo "PR closed; this run only cancels older runs in the same concurrency group." run: echo "PR closed; this run only cancels older runs in the same concurrency group."
@@ -89,13 +92,20 @@ jobs:
name: Typos name: Typos
if: github.event_name != 'pull_request' || github.event.action != 'closed' if: github.event_name != 'pull_request' || github.event.action != 'closed'
runs-on: ubuntu-latest runs-on: ubuntu-latest
timeout-minutes: 10
steps: steps:
- uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7 - uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
with:
persist-credentials: false
- name: Typos check with custom config file - name: Typos check with custom config file
uses: crate-ci/typos@37bb98842b0d8c4ffebdb75301a13db0267cef89 # master uses: crate-ci/typos@37bb98842b0d8c4ffebdb75301a13db0267cef89 # master
# Fast, compile-free checks that fail early so contributors get feedback in # Fast, compile-free checks that fail early so contributors get feedback in
# ~1 minute instead of waiting for the full test job. # ~1 minute instead of waiting for the full test job.
#
# These steps are mirrored byte-for-byte in ci-docs-only.yml so that a mixed
# PR, which reports two check runs named "Quick Checks", cannot get one red
# and one green. Edit both jobs together.
quick-checks: quick-checks:
name: Quick Checks name: Quick Checks
if: github.event_name != 'pull_request' || github.event.action != 'closed' if: github.event_name != 'pull_request' || github.event.action != 'closed'
@@ -104,6 +114,8 @@ jobs:
steps: steps:
- name: Checkout repository - name: Checkout repository
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7 uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
with:
persist-credentials: false
- name: Install ripgrep - name: Install ripgrep
run: sudo apt-get update && sudo apt-get install -y ripgrep run: sudo apt-get update && sudo apt-get install -y ripgrep
@@ -137,24 +149,48 @@ jobs:
- name: Check no planning docs committed - name: Check no planning docs committed
run: ./scripts/check_no_planning_docs.sh run: ./scripts/check_no_planning_docs.sh
- name: Check CI paths stay in sync
run: ./scripts/check_ci_paths_sync.sh
- name: Check io_uring lane --lib precondition
run: ./scripts/check_uring_lane_lib_only.sh
test-and-lint: test-and-lint:
name: Test and Lint name: Test and Lint
if: github.event_name != 'pull_request' || github.event.action != 'closed' if: github.event_name != 'pull_request' || github.event.action != 'closed'
needs: [ quick-checks ]
runs-on: sm-standard-4 runs-on: sm-standard-4
timeout-minutes: 90 timeout-minutes: 90
# Both lines are required. Job-level `permissions` replaces the workflow
# block rather than merging with it, so declaring only `actions: write`
# would drop `contents: read` and break this job's checkout and the
# repo-token the setup action hands to setup-protoc.
permissions:
contents: read
actions: write
env: env:
FORCE_JAVASCRIPT_ACTIONS_TO_NODE24: "true" FORCE_JAVASCRIPT_ACTIONS_TO_NODE24: "true"
steps: steps:
- name: Checkout repository - name: Checkout repository
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7 uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
with:
# This job's token can cancel runs and delete Actions caches. Checkout
# otherwise writes it into .git/config, where a PR's own build.rs or
# proc-macro could read it back out.
persist-credentials: false
- name: Setup Rust environment - name: Setup Rust environment
uses: ./.github/actions/setup uses: ./.github/actions/setup
with: with:
rust-version: stable rust-version: stable
cache-shared-key: ci-test # Every lane in this workflow reads its cache and none writes it.
github-token: ${{ secrets.GITHUB_TOKEN }} # cache-warm.yml is the sole writer for all four keys: this workflow
cache-save-if: ${{ github.ref == 'refs/heads/main' }} # cancels superseded runs on main, and a cancelled run never reaches
# rust-cache's post step, so writing from here saved nothing (12 of 15
# consecutive main-push runs were cancelled). See rustfs/backlog#1600.
cache-shared-key: ci-dev
cache-save-if: 'false'
install-build-packaging-tools: 'false'
- name: Prepare test evidence - name: Prepare test evidence
run: | run: |
@@ -169,43 +205,54 @@ jobs:
# Clippy runs before the test pass: lint failures are the most common # Clippy runs before the test pass: lint failures are the most common
# CI-only breakage and should surface in minutes, not after 20+ minutes # CI-only breakage and should surface in minutes, not after 20+ minutes
# of tests. # of tests.
# Sampled too: clippy is the natural control arm for any CARGO_BUILD_JOBS
# experiment, since --all-targets is check-only for workspace members and
# never links the ~100 test binaries the limit exists to throttle.
- name: Run clippy lints - name: Run clippy lints
run: cargo clippy --all-targets -- -D warnings run: |
./scripts/ci/resource_sampler.sh start clippy
trap './scripts/ci/resource_sampler.sh stop' EXIT
cargo clippy --all-targets -- -D warnings
- name: Run nextest tests - name: Run nextest tests
env: env:
# Three concurrent workspace test links saturate the self-hosted # #5394 mitigation, now under a measured experiment (backlog#1601).
# runner's overlay I/O and can wedge Cargo until the 75m timeout. #
CARGO_BUILD_JOBS: "2" # 2 was chosen when three concurrent workspace test links were believed
# to saturate the runner's overlay I/O and wedge Cargo until the 75m
# timeout. cgroup v2 readings from the sampler show the pod actually
# has 14 CPUs and 28GB (peak use 2.1GB), so 2 throttles compilation to
# a seventh of what is available and memory was never the constraint —
# the label name "sm-standard-4" had led everyone, including the
# original mitigation, to assume 4 cores.
#
# Raised to 3 on main pushes and manual dispatches; PRs keep 2 so the
# merge path is untouched while the experiment runs.
#
# Dispatch is included because push alone cannot supply the samples:
# this workflow cancels superseded runs on main, and only 4 of the last
# 20 push-triggered Test and Lint jobs reached a terminal state — at
# that rate ten samples would take roughly fifty merges. The
# concurrency group is scoped by event_name, so a dispatched run has
# its own group and is not cancelled by merge traffic, which makes the
# sample collectable on demand rather than by waiting.
#
# Baseline over 17 samples at 2:
# median nextest/clippy step ratio 1.95, spread 1.85-2.06. The gate-2
# criterion is that ratio dropping at least 10% (below ~1.76) with no
# 75m timeout and no run showing three consecutive samples of
# rustc/collect2/rust-lld in D state. If it does not, the conclusion is
# "this limit is not the bottleneck" — fix it back at 2 and record the
# experiment, which is a result, not a failure.
#
# Must stay step-level: rust-cache hashes CARGO/CC/CFLAGS/CXX/CMAKE/RUST
# prefixed variables from process.env into the cache key, so promoting
# this to job level would rotate every key on this lane.
CARGO_BUILD_JOBS: ${{ (github.event_name == 'push' || github.event_name == 'workflow_dispatch') && '3' || '2' }}
run: | run: |
mkdir -p artifacts/test-and-lint mkdir -p artifacts/test-and-lint
# Evidence sampler for issue #5394: the post-mortem pgrep below runs ./scripts/ci/resource_sampler.sh start nextest
# only after `timeout` has already TERM'd the whole cargo process trap './scripts/ci/resource_sampler.sh stop' EXIT
# group, so it cannot name a wedged process. Sample system and
# process state every 60s instead; the last samples before the
# timeout show what was stuck (rustc, linker, build script, memory
# pressure, ...). The log rides along in the existing artifact.
(
while true; do
{
echo "=== $(date --utc --iso-8601=seconds)"
echo "--- load"; cat /proc/loadavg
echo "--- psi"; grep -H . /proc/pressure/* 2>/dev/null || true
echo "--- mem"; free -m
echo "--- disk"; df -h / /home/runner 2>/dev/null || df -h /
echo "--- top-rss"
ps -eo pid,ppid,stat,etime,rss,pcpu,args --sort=-rss | head -15
echo "--- build/test processes"
ps -eo pid,ppid,stat,etime,rss,pcpu,args | grep -E '[c]argo|[r]ustc|[n]extest|[c]ollect2|rust-ll[d]|[b]uild-script|deps[/]' || true
echo "--- d-state (uninterruptible IO)"
ps -eo pid,stat,etime,args | awk 'NR > 1 && $2 ~ /D/' || true
echo
} >> artifacts/test-and-lint/sampler.log 2>&1 || true
sleep 60
done
) &
sampler_pid=$!
trap 'kill "${sampler_pid}" 2>/dev/null || true' EXIT
set +e set +e
NEXTEST_HIDE_PROGRESS_BAR=1 timeout --verbose --signal=TERM --kill-after=30s 75m \ NEXTEST_HIDE_PROGRESS_BAR=1 timeout --verbose --signal=TERM --kill-after=30s 75m \
cargo nextest run --profile ci --all --exclude e2e_test \ cargo nextest run --profile ci --all --exclude e2e_test \
@@ -277,6 +324,48 @@ jobs:
- name: Run rebalance/decommission migration proofs - name: Run rebalance/decommission migration proofs
run: ./scripts/check_migration_gate_count.sh run: ./scripts/check_migration_gate_count.sh
# Early stop. Once this job has failed the PR cannot merge, so the sibling
# lanes are burning runners on a result nobody can act on: on run
# 30674613104 three lanes had already failed while Test and Lint and the
# rio-v2 variant kept going past 70 minutes.
#
# Only this job may cancel. The lanes that are NOT required checks
# (protocols, ILM, e2e, s3-tests) must never hold that power: a flake in
# one of them would turn the required "Test and Lint" into `cancelled`,
# which blocks the merge. Today a maintainer can merge with sftp red, and
# that has to stay true.
#
# These steps run last so the `if: always()` artifact upload above still
# captures logs and diagnostics before the run goes away.
- name: Annotate early-stop reason
if: failure() && github.event_name == 'pull_request'
run: |
echo "## CI early-stop" >> "$GITHUB_STEP_SUMMARY"
echo "Job \`${GITHUB_JOB}\` (Test and Lint) failed; cancelling run ${GITHUB_RUN_ID} to free runners." >> "$GITHUB_STEP_SUMMARY"
echo "Sibling jobs showing **cancelled** were stopped by this job, not by their own failure." >> "$GITHUB_STEP_SUMMARY"
# curl rather than `gh`: every existing `gh` call in this repo runs on
# ubuntu-latest, and the sm-standard-* images are custom and trimmed (they
# ship no C toolchain, see the e2e job below), so `gh` is not known to
# exist here.
#
# Fork PRs are excluded explicitly instead of relying on the error path:
# their GITHUB_TOKEN is forced read-only and job-level permissions cannot
# raise it, so the call would always 403. Skipping keeps their logs clean.
- name: Cancel run on failure (same-repo PR only)
if: >-
failure() && github.event_name == 'pull_request'
&& github.event.pull_request.head.repo.full_name == github.repository
continue-on-error: true
env:
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
run: |
curl -fsS -X POST \
-H "Authorization: Bearer ${GH_TOKEN}" \
-H "Accept: application/vnd.github+json" \
-H "X-GitHub-Api-Version: 2022-11-28" \
"${GITHUB_API_URL}/repos/${GITHUB_REPOSITORY}/actions/runs/${GITHUB_RUN_ID}/cancel" || true
# Dedicated serial lane for the ILM / lifecycle integration tests. These tests # Dedicated serial lane for the ILM / lifecycle integration tests. These tests
# drive the object layer through process-global singletons (the GLOBAL_ENV # drive the object layer through process-global singletons (the GLOBAL_ENV
# ECStore, the global tier-config manager, background-expiry workers) and bind # ECStore, the global tier-config manager, background-expiry workers) and bind
@@ -290,6 +379,7 @@ jobs:
test-ilm-integration-serial: test-ilm-integration-serial:
name: ILM Integration (serial) name: ILM Integration (serial)
if: github.event_name != 'pull_request' || github.event.action != 'closed' if: github.event_name != 'pull_request' || github.event.action != 'closed'
needs: [ quick-checks ]
runs-on: sm-standard-4 runs-on: sm-standard-4
timeout-minutes: 45 timeout-minutes: 45
env: env:
@@ -297,14 +387,16 @@ jobs:
steps: steps:
- name: Checkout repository - name: Checkout repository
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7 uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
with:
persist-credentials: false
- name: Setup Rust environment - name: Setup Rust environment
uses: ./.github/actions/setup uses: ./.github/actions/setup
with: with:
rust-version: stable rust-version: stable
cache-shared-key: ci-ilm-serial cache-shared-key: ci-dev
github-token: ${{ secrets.GITHUB_TOKEN }} cache-save-if: 'false'
cache-save-if: ${{ github.ref == 'refs/heads/main' }} install-build-packaging-tools: 'false'
# test_transition_and_restore_flows was re-enabled by rustfs/backlog#1303: # test_transition_and_restore_flows was re-enabled by rustfs/backlog#1303:
# its "missing xl.meta on disk2" was a test-util bug (open_disk hardcoded # its "missing xl.meta on disk2" was a test-util bug (open_disk hardcoded
@@ -327,6 +419,7 @@ jobs:
test-and-lint-rio-v2: test-and-lint-rio-v2:
name: Test and Lint (rio-v2) name: Test and Lint (rio-v2)
if: github.event_name != 'pull_request' || github.event.action != 'closed' if: github.event_name != 'pull_request' || github.event.action != 'closed'
needs: [ quick-checks ]
runs-on: sm-standard-4 runs-on: sm-standard-4
timeout-minutes: 60 timeout-minutes: 60
env: env:
@@ -334,14 +427,16 @@ jobs:
steps: steps:
- name: Checkout repository - name: Checkout repository
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7 uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
with:
persist-credentials: false
- name: Setup Rust environment - name: Setup Rust environment
uses: ./.github/actions/setup uses: ./.github/actions/setup
with: with:
rust-version: stable rust-version: stable
cache-shared-key: ci-test-rio-v2 cache-shared-key: ci-feat-rio
github-token: ${{ secrets.GITHUB_TOKEN }} cache-save-if: 'false'
cache-save-if: ${{ github.ref == 'refs/heads/main' }} install-build-packaging-tools: 'false'
- name: Run rio-v2 clippy lints - name: Run rio-v2 clippy lints
run: cargo clippy -p rustfs -p rustfs-ecstore --all-targets --features rio-v2 -- -D warnings run: cargo clippy -p rustfs -p rustfs-ecstore --all-targets --features rio-v2 -- -D warnings
@@ -354,10 +449,17 @@ jobs:
test-and-lint-protocols: test-and-lint-protocols:
name: "Test and Lint (${{ matrix.features.name }})" name: "Test and Lint (${{ matrix.features.name }})"
if: github.event_name != 'pull_request' || github.event.action != 'closed' if: github.event_name != 'pull_request' || github.event.action != 'closed'
needs: [ quick-checks ]
runs-on: sm-standard-4 runs-on: sm-standard-4
timeout-minutes: 60 timeout-minutes: 60
strategy: strategy:
fail-fast: false # On a PR, one failing protocol leg is enough to know the PR is not ready,
# so stop the sibling leg instead of paying another ~40 minutes for it.
# Everywhere else (main pushes, the merge queue, the weekly schedule) keep
# the full signal: there we want to know whether swift AND sftp are broken,
# not just whichever failed first. This is the only part of the early-stop
# work that also covers fork PRs, since it needs no token.
fail-fast: ${{ github.event_name == 'pull_request' }}
matrix: matrix:
features: features:
- name: swift - name: swift
@@ -369,14 +471,16 @@ jobs:
steps: steps:
- name: Checkout repository - name: Checkout repository
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7 uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
with:
persist-credentials: false
- name: Setup Rust environment - name: Setup Rust environment
uses: ./.github/actions/setup uses: ./.github/actions/setup
with: with:
rust-version: stable rust-version: stable
cache-shared-key: ci-test-${{ matrix.features.name }} cache-shared-key: ci-feat-proto
github-token: ${{ secrets.GITHUB_TOKEN }} cache-save-if: 'false'
cache-save-if: ${{ github.ref == 'refs/heads/main' }} install-build-packaging-tools: 'false'
- name: Run clippy with ${{ matrix.features.name }} - name: Run clippy with ${{ matrix.features.name }}
run: | run: |
@@ -389,6 +493,7 @@ jobs:
build-rustfs-debug-binary: build-rustfs-debug-binary:
name: Build RustFS Debug Binary name: Build RustFS Debug Binary
if: github.event_name != 'pull_request' || github.event.action != 'closed' if: github.event_name != 'pull_request' || github.event.action != 'closed'
needs: [ quick-checks ]
runs-on: sm-standard-4 runs-on: sm-standard-4
timeout-minutes: 30 timeout-minutes: 30
env: env:
@@ -396,14 +501,16 @@ jobs:
steps: steps:
- name: Checkout repository - name: Checkout repository
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7 uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
with:
persist-credentials: false
- name: Setup Rust environment - name: Setup Rust environment
uses: ./.github/actions/setup uses: ./.github/actions/setup
with: with:
rust-version: stable rust-version: stable
cache-shared-key: ci-rustfs-debug-binary cache-shared-key: ci-dev
cache-save-if: ${{ github.ref == 'refs/heads/main' }} cache-save-if: 'false'
github-token: ${{ secrets.GITHUB_TOKEN }} install-build-packaging-tools: 'false'
- name: Build debug binary - name: Build debug binary
run: cargo build -p rustfs --bins --features e2e-test-hooks run: cargo build -p rustfs --bins --features e2e-test-hooks
@@ -419,6 +526,7 @@ jobs:
build-rustfs-debug-binary-rio-v2: build-rustfs-debug-binary-rio-v2:
name: Build RustFS Debug Binary (rio-v2) name: Build RustFS Debug Binary (rio-v2)
if: github.event_name != 'pull_request' || github.event.action != 'closed' if: github.event_name != 'pull_request' || github.event.action != 'closed'
needs: [ quick-checks ]
runs-on: sm-standard-4 runs-on: sm-standard-4
timeout-minutes: 30 timeout-minutes: 30
env: env:
@@ -426,14 +534,16 @@ jobs:
steps: steps:
- name: Checkout repository - name: Checkout repository
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7 uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
with:
persist-credentials: false
- name: Setup Rust environment - name: Setup Rust environment
uses: ./.github/actions/setup uses: ./.github/actions/setup
with: with:
rust-version: stable rust-version: stable
cache-shared-key: ci-rustfs-debug-binary-rio-v2 cache-shared-key: ci-feat-rio
cache-save-if: ${{ github.ref == 'refs/heads/main' }} cache-save-if: 'false'
github-token: ${{ secrets.GITHUB_TOKEN }} install-build-packaging-tools: 'false'
- name: Build debug binary with rio-v2 - name: Build debug binary with rio-v2
run: cargo build -p rustfs --bins --features rio-v2,e2e-test-hooks run: cargo build -p rustfs --bins --features rio-v2,e2e-test-hooks
@@ -448,6 +558,14 @@ jobs:
uring-integration: uring-integration:
name: io_uring Integration (real) name: io_uring Integration (real)
# The pull_request trigger includes `closed` purely so the concurrency
# group cancels in-flight runs of a closed PR; every other job opts out of
# that run with this guard (or is skipped through its `needs` chain). This
# job had neither, so each closed/merged PR really ran the whole io_uring
# suite (measured 4m17s / 7m19s / 7m31s on runs 30678272341 / 30678117601 /
# 30662728539) and kept the cancellation run in progress for minutes.
if: github.event_name != 'pull_request' || github.event.action != 'closed'
needs: [ quick-checks ]
# GitHub-hosted ubuntu-latest runs a recent kernel with io_uring and, unlike # GitHub-hosted ubuntu-latest runs a recent kernel with io_uring and, unlike
# a container, applies no seccomp filter that would block io_uring_setup — so # a container, applies no seccomp filter that would block io_uring_setup — so
# the probe succeeds and the tests exercise the real UringBackend/FdCache/ # the probe succeeds and the tests exercise the real UringBackend/FdCache/
@@ -458,17 +576,24 @@ jobs:
steps: steps:
- name: Checkout repository - name: Checkout repository
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7 uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
with:
persist-credentials: false
- name: Setup Rust environment - name: Setup Rust environment
uses: ./.github/actions/setup uses: ./.github/actions/setup
with: with:
rust-version: stable rust-version: stable
# Keeps its own key rather than joining ci-dev. rust-cache's key is
# built from runner.os/arch plus rustc and lockfile fingerprints — it
# does NOT include the runner label or image. ubuntu-latest and
# sm-standard-4 are therefore indistinguishable to it, so sharing a key
# would let two different system images overwrite each other's
# artifacts, and would make a 2-core hosted runner unpack ci-dev's ~3GB
# instead of this lane's ~1.3GB. cache-warm.yml warms this key on
# ubuntu-latest for the same reason.
cache-shared-key: ci-uring cache-shared-key: ci-uring
github-token: ${{ secrets.GITHUB_TOKEN }} cache-save-if: 'false'
cache-save-if: ${{ github.ref == 'refs/heads/main' }} install-build-packaging-tools: 'false'
- name: Install build dependencies
run: sudo apt-get update && sudo apt-get install -y protobuf-compiler
# ext4 supports O_DIRECT; the runner's default TMPDIR may sit on tmpfs or # ext4 supports O_DIRECT; the runner's default TMPDIR may sit on tmpfs or
# overlayfs, where open(O_DIRECT) returns EINVAL/EOPNOTSUPP and the native # overlayfs, where open(O_DIRECT) returns EINVAL/EOPNOTSUPP and the native
@@ -495,7 +620,17 @@ jobs:
RUSTFS_IO_URING_READ_ENABLE: "true" RUSTFS_IO_URING_READ_ENABLE: "true"
RUSTFS_URING_TESTS_MUST_RUN: "1" RUSTFS_URING_TESTS_MUST_RUN: "1"
TMPDIR: /mnt/rustfs-odirect TMPDIR: /mnt/rustfs-odirect
run: cargo test -p rustfs-ecstore uring_ -- --test-threads=1 --nocapture # --lib narrows what gets compiled, not what gets run: every selected
# test lives in the lib target. The 7 integration binaries under
# crates/ecstore/tests/ each reported "running 0 tests" here, so they
# were compiled and linked for nothing.
#
# The `uring_` filter must stay exactly as it is. libtest matches on
# substring, so it also selects names containing `during_` — 6 of the 18
# selected tests are such incidental matches. Narrowing the filter to
# `io_uring` would silently drop them, which is a coverage change.
# scripts/check_uring_lane_lib_only.sh guards the --lib precondition.
run: cargo test -p rustfs-ecstore --lib uring_ -- --test-threads=1 --nocapture
e2e-tests: e2e-tests:
name: End-to-End Tests name: End-to-End Tests
@@ -505,6 +640,8 @@ jobs:
steps: steps:
- name: Checkout repository - name: Checkout repository
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7 uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
with:
persist-credentials: false
# Full setup with dependency caching: the smoke-suite step below # Full setup with dependency caching: the smoke-suite step below
# compiles the e2e_test crate, which pulls in most of the workspace. # compiles the e2e_test crate, which pulls in most of the workspace.
@@ -513,9 +650,9 @@ jobs:
uses: ./.github/actions/setup uses: ./.github/actions/setup
with: with:
rust-version: stable rust-version: stable
cache-shared-key: ci-e2e cache-shared-key: ci-dev
github-token: ${{ secrets.GITHUB_TOKEN }} cache-save-if: 'false'
cache-save-if: ${{ github.ref == 'refs/heads/main' }} install-build-packaging-tools: 'false'
# Download after the cache restore so the freshly built binary from the # Download after the cache restore so the freshly built binary from the
# build job always wins over anything restored into target/debug. # build job always wins over anything restored into target/debug.
@@ -589,14 +726,16 @@ jobs:
steps: steps:
- name: Checkout repository - name: Checkout repository
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7 uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
with:
persist-credentials: false
- name: Setup Rust environment - name: Setup Rust environment
uses: ./.github/actions/setup uses: ./.github/actions/setup
with: with:
rust-version: stable rust-version: stable
cache-shared-key: ci-e2e cache-shared-key: ci-dev
github-token: ${{ secrets.GITHUB_TOKEN }} cache-save-if: 'false'
cache-save-if: ${{ github.ref == 'refs/heads/main' }} install-build-packaging-tools: 'false'
# Download after the cache restore so the freshly built binary from the # Download after the cache restore so the freshly built binary from the
# build job always wins over anything restored into target/debug. # build job always wins over anything restored into target/debug.
@@ -632,6 +771,8 @@ jobs:
steps: steps:
- name: Checkout repository - name: Checkout repository
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7 uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
with:
persist-credentials: false
- name: Clean up previous test run - name: Clean up previous test run
run: | run: |
@@ -686,6 +827,8 @@ jobs:
steps: steps:
- name: Checkout repository - name: Checkout repository
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7 uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
with:
persist-credentials: false
- name: Download debug binary - name: Download debug binary
uses: actions/download-artifact@37930b1c2abaa49bbe596cd826c3c89aef350131 # v7 uses: actions/download-artifact@37930b1c2abaa49bbe596cd826c3c89aef350131 # v7
@@ -746,12 +889,20 @@ jobs:
# evaluates ILM within ~2s of the due time, well inside the poll window. # evaluates ILM within ~2s of the due time, well inside the poll window.
s3-lifecycle-behavior-tests: s3-lifecycle-behavior-tests:
name: S3 Lifecycle Behavior Tests name: S3 Lifecycle Behavior Tests
needs: [ build-rustfs-debug-binary ] # Also gated on e2e-tests, matching s3-implemented-tests: when the e2e smoke
# suite is already red this lane cannot tell us anything new, and it holds a
# sm-standard-4 for up to 30 minutes doing so. Both lanes only download the
# prebuilt debug binary (no cargo build), and s3-implemented-tests — which
# already waits on e2e-tests — finishes later anyway, so a green PR's total
# wall clock is unchanged.
needs: [ build-rustfs-debug-binary, e2e-tests ]
runs-on: sm-standard-4 runs-on: sm-standard-4
timeout-minutes: 30 timeout-minutes: 30
steps: steps:
- name: Checkout repository - name: Checkout repository
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7 uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
with:
persist-credentials: false
- name: Download debug binary - name: Download debug binary
uses: actions/download-artifact@37930b1c2abaa49bbe596cd826c3c89aef350131 # v7 uses: actions/download-artifact@37930b1c2abaa49bbe596cd826c3c89aef350131 # v7
+23 -4
View File
@@ -22,11 +22,18 @@ on:
issue_comment: issue_comment:
types: [created, edited] types: [created, edited]
# Least privilege at the top, widened per job below. This workflow runs on
# pull_request_target and issue_comment, so it holds full secrets on every fork
# PR and on any comment anyone writes — the one place in this repository where a
# compromised action would be handed a repo-write token. It does not check out
# or execute PR code, so there is no pwn-request path today, but the blast
# radius should not depend on that staying true.
#
# contents: write in particular was never used: the signature records are
# written to rustfs/cla through the scoped app token created below, and nothing
# here writes to this repository's contents.
permissions: permissions:
contents: write contents: read
pull-requests: write
issues: write
checks: write
concurrency: concurrency:
group: ${{ github.workflow }}-${{ github.event.pull_request.number || github.event.issue.number || github.ref }} group: ${{ github.workflow }}-${{ github.event.pull_request.number || github.event.issue.number || github.ref }}
@@ -36,14 +43,26 @@ jobs:
cancel-closed-pr-runs: cancel-closed-pr-runs:
name: Cancel Closed PR Runs name: Cancel Closed PR Runs
if: github.event_name == 'pull_request_target' && github.event.action == 'closed' if: github.event_name == 'pull_request_target' && github.event.action == 'closed'
# Echoes one line; the run exists only so the concurrency group cancels the
# in-flight run of a closed PR.
permissions: {}
runs-on: ubuntu-latest runs-on: ubuntu-latest
timeout-minutes: 10
steps: steps:
- name: Explain cancellation run - name: Explain cancellation run
run: echo "PR closed; this run only cancels older runs in the same concurrency group." run: echo "PR closed; this run only cancels older runs in the same concurrency group."
cla: cla:
if: ${{ (github.event_name != 'issue_comment' || github.event.issue.pull_request) && (github.event_name != 'pull_request_target' || github.event.action != 'closed') }} if: ${{ (github.event_name != 'issue_comment' || github.event.issue.pull_request) && (github.event_name != 'pull_request_target' || github.event.action != 'closed') }}
# checks: write reports the merge-queue check run; pull-requests and issues
# let cla-bot comment and label. contents stays read — see the note above.
permissions:
contents: read
checks: write
issues: write
pull-requests: write
runs-on: ubuntu-latest runs-on: ubuntu-latest
timeout-minutes: 30
steps: steps:
- name: Report CLA result for merge queue - name: Report CLA result for merge queue
if: github.event_name == 'merge_group' if: github.event_name == 'merge_group'
+5 -1
View File
@@ -62,14 +62,16 @@ jobs:
steps: steps:
- name: Checkout repository - name: Checkout repository
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7 uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
with:
persist-credentials: false
- name: Setup Rust environment - name: Setup Rust environment
uses: ./.github/actions/setup uses: ./.github/actions/setup
with: with:
rust-version: stable rust-version: stable
cache-shared-key: ci-coverage cache-shared-key: ci-coverage
github-token: ${{ secrets.GITHUB_TOKEN }}
cache-save-if: ${{ github.ref == 'refs/heads/main' }} cache-save-if: ${{ github.ref == 'refs/heads/main' }}
install-build-packaging-tools: 'false'
- name: Install cargo-llvm-cov - name: Install cargo-llvm-cov
uses: taiki-e/install-action@bffeee26d4db9be238a4ea78d8826604ebcb594d # v2 uses: taiki-e/install-action@bffeee26d4db9be238a4ea78d8826604ebcb594d # v2
@@ -116,6 +118,8 @@ jobs:
issues: write issues: write
steps: steps:
- uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7 - uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
with:
persist-credentials: false
- name: Open or update failure-tracking issue - name: Open or update failure-tracking issue
uses: ./.github/actions/schedule-failure-issue uses: ./.github/actions/schedule-failure-issue
with: with:
+34 -20
View File
@@ -85,6 +85,7 @@ jobs:
github.event.workflow_run.head_branch != 'main' && github.event.workflow_run.head_branch != 'main' &&
!contains(github.event.workflow_run.head_branch, '-preview')) !contains(github.event.workflow_run.head_branch, '-preview'))
runs-on: ubuntu-latest runs-on: ubuntu-latest
timeout-minutes: 30
outputs: outputs:
should_build: ${{ steps.check.outputs.should_build }} should_build: ${{ steps.check.outputs.should_build }}
should_push: ${{ steps.check.outputs.should_push }} should_push: ${{ steps.check.outputs.should_push }}
@@ -97,11 +98,18 @@ jobs:
- name: Checkout repository - name: Checkout repository
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7 uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
with: with:
persist-credentials: false
# For workflow_run events, checkout the specific commit that triggered the workflow # For workflow_run events, checkout the specific commit that triggered the workflow
ref: ${{ github.event.workflow_run.head_sha || github.sha }} ref: ${{ github.event.workflow_run.head_sha || github.sha }}
- name: Check build conditions - name: Check build conditions
id: check id: check
env:
# dispatch inputs via env, not `${{ }}` interpolation: they are
# free-form strings and would otherwise be evaluated by bash.
INPUT_VERSION: ${{ github.event.inputs.version }}
INPUT_PUSH_IMAGES: ${{ github.event.inputs.push_images }}
INPUT_FORCE_REBUILD: ${{ github.event.inputs.force_rebuild }}
run: | run: |
should_build=false should_build=false
should_push=false should_push=false
@@ -202,9 +210,9 @@ jobs:
elif [[ "${{ github.event_name }}" == "workflow_dispatch" ]]; then elif [[ "${{ github.event_name }}" == "workflow_dispatch" ]]; then
# Manual trigger # Manual trigger
input_version="${{ github.event.inputs.version }}" input_version="$INPUT_VERSION"
version="${input_version}" version="${input_version}"
should_push="${{ github.event.inputs.push_images }}" should_push="$INPUT_PUSH_IMAGES"
should_build=true should_build=true
# Get short SHA # Get short SHA
@@ -212,7 +220,7 @@ jobs:
echo "🎯 Manual Docker build triggered:" echo "🎯 Manual Docker build triggered:"
echo " 📋 Requested version: $input_version" echo " 📋 Requested version: $input_version"
echo " 🔧 Force rebuild: ${{ github.event.inputs.force_rebuild }}" echo " 🔧 Force rebuild: $INPUT_FORCE_REBUILD"
echo " 🚀 Push images: $should_push" echo " 🚀 Push images: $should_push"
case "$input_version" in case "$input_version" in
@@ -298,6 +306,8 @@ jobs:
steps: steps:
- name: Checkout repository - name: Checkout repository
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7 uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
with:
persist-credentials: false
- name: Login to Docker Hub - name: Login to Docker Hub
uses: docker/login-action@c94ce9fb468520275223c153574b00df6fe4bcc9 # v3 uses: docker/login-action@c94ce9fb468520275223c153574b00df6fe4bcc9 # v3
@@ -333,32 +343,28 @@ jobs:
CREATE_LATEST="${{ needs.build-check.outputs.create_latest }}" CREATE_LATEST="${{ needs.build-check.outputs.create_latest }}"
VARIANT_SUFFIX="${{ matrix.suffix }}" VARIANT_SUFFIX="${{ matrix.suffix }}"
# Convert version format for Dockerfile compatibility # Convert version format for Dockerfile compatibility. The former
# DOCKER_CHANNEL was "release" down every branch and was passed as a
# build-arg no Dockerfile declares, so it is gone.
case "$VERSION" in case "$VERSION" in
"latest") "latest")
# For stable latest, use RELEASE=latest + release CHANNEL
DOCKER_RELEASE="latest" DOCKER_RELEASE="latest"
DOCKER_CHANNEL="release"
;; ;;
v*) v*)
# For versioned releases (v1.0.0), remove 'v' prefix for Dockerfile # For versioned releases (v1.0.0), remove 'v' prefix for Dockerfile
DOCKER_RELEASE="${VERSION#v}" DOCKER_RELEASE="${VERSION#v}"
DOCKER_CHANNEL="release"
;; ;;
*) *)
# For other versions, pass as-is # For other versions, pass as-is
DOCKER_RELEASE="${VERSION}" DOCKER_RELEASE="${VERSION}"
DOCKER_CHANNEL="release"
;; ;;
esac esac
echo "docker_release=$DOCKER_RELEASE" >> "$GITHUB_OUTPUT" echo "docker_release=$DOCKER_RELEASE" >> "$GITHUB_OUTPUT"
echo "docker_channel=$DOCKER_CHANNEL" >> "$GITHUB_OUTPUT"
echo "🐳 Docker build parameters:" echo "🐳 Docker build parameters:"
echo " - Original version: $VERSION" echo " - Original version: $VERSION"
echo " - Docker RELEASE: $DOCKER_RELEASE" echo " - Docker RELEASE: $DOCKER_RELEASE"
echo " - Docker CHANNEL: $DOCKER_CHANNEL"
# Generate tags based on build type # Generate tags based on build type
# Only support release and prerelease builds (no development builds) # Only support release and prerelease builds (no development builds)
@@ -412,18 +418,24 @@ jobs:
push: ${{ needs.build-check.outputs.should_push == 'true' }} push: ${{ needs.build-check.outputs.should_push == 'true' }}
tags: ${{ steps.meta.outputs.tags }} tags: ${{ steps.meta.outputs.tags }}
labels: ${{ steps.meta.outputs.labels }} labels: ${{ steps.meta.outputs.labels }}
cache-from: | # No layer cache. This build compiles nothing — it downloads a
type=gha,scope=docker-${{ matrix.variant }} # release zip and runs apk/apt — so the cache could only save the
cache-to: | # minute or two those take, while creating a correctness problem: with
type=gha,mode=max,scope=docker-${{ matrix.variant }} # RELEASE=latest the binary URL is resolved by curl *inside* a RUN
# layer, and the layer key does not include what that resolved to. A
# rebuild at the same RELEASE value (dispatch with version=latest, or
# a re-run of the same version) would hit the old layer and ship the
# previous release's binary. mode=max also consumed the same 10GB
# Actions cache quota the Rust lanes are fighting over.
#
# Only RELEASE is passed: it is the sole build-arg the Dockerfiles
# declare besides TARGETARCH. BUILDTIME, VERSION, BUILD_TYPE, REVISION
# and CHANNEL were never read by any stage (and BUILDTIME's $(date ...)
# was a literal here, not a shell substitution). BUILD_DATE and VCS_REF
# are declared by the Dockerfiles but deliberately left unset —
# supplying them would change the published image labels.
build-args: | build-args: |
BUILDTIME=$(date -u +'%Y-%m-%dT%H:%M:%SZ')
VERSION=${{ needs.build-check.outputs.version }}
BUILD_TYPE=${{ needs.build-check.outputs.build_type }}
REVISION=${{ github.sha }}
RELEASE=${{ steps.meta.outputs.docker_release }} RELEASE=${{ steps.meta.outputs.docker_release }}
CHANNEL=${{ steps.meta.outputs.docker_channel }}
BUILDKIT_INLINE_CACHE=1
provenance: true provenance: true
sbom: true sbom: true
# Add retry mechanism by splitting the build process # Add retry mechanism by splitting the build process
@@ -439,6 +451,7 @@ jobs:
needs: [ build-check, build-docker ] needs: [ build-check, build-docker ]
if: needs.build-check.outputs.should_build == 'true' && needs.build-check.outputs.should_push == 'true' if: needs.build-check.outputs.should_build == 'true' && needs.build-check.outputs.should_push == 'true'
runs-on: ubuntu-latest runs-on: ubuntu-latest
timeout-minutes: 30
permissions: permissions:
contents: read contents: read
security-events: write security-events: write
@@ -493,6 +506,7 @@ jobs:
needs: [ build-check, build-docker ] needs: [ build-check, build-docker ]
if: always() && needs.build-check.outputs.should_build == 'true' if: always() && needs.build-check.outputs.should_build == 'true'
runs-on: ubuntu-latest runs-on: ubuntu-latest
timeout-minutes: 10
steps: steps:
- name: Docker build completion summary - name: Docker build completion summary
run: | run: |
@@ -64,14 +64,16 @@ jobs:
steps: steps:
- name: Checkout repository - name: Checkout repository
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7 uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
with:
persist-credentials: false
- name: Setup Rust environment - name: Setup Rust environment
uses: ./.github/actions/setup uses: ./.github/actions/setup
with: with:
rust-version: stable rust-version: stable
cache-shared-key: ci-e2e-repl cache-shared-key: ci-e2e-repl
github-token: ${{ secrets.GITHUB_TOKEN }}
cache-save-if: ${{ github.ref == 'refs/heads/main' }} cache-save-if: ${{ github.ref == 'refs/heads/main' }}
install-build-packaging-tools: 'false'
# awscurl lets the STS dual-node test actually exercise its path. Without # awscurl lets the STS dual-node test actually exercise its path. Without
# it the test skips gracefully with a visible log line # it the test skips gracefully with a visible log line
@@ -124,6 +126,8 @@ jobs:
issues: write issues: write
steps: steps:
- uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7 - uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
with:
persist-credentials: false
- name: Open or update failure-tracking issue - name: Open or update failure-tracking issue
uses: ./.github/actions/schedule-failure-issue uses: ./.github/actions/schedule-failure-issue
with: with:
+11
View File
@@ -45,6 +45,13 @@
# The PR gate (ci.yml s3-implemented-tests) is unaffected: it avoids Docker # The PR gate (ci.yml s3-implemented-tests) is unaffected: it avoids Docker
# via DEPLOY_MODE=binary and defers all pip setup to run.sh's self-bootstrap. # via DEPLOY_MODE=binary and defers all pip setup to run.sh's self-bootstrap.
# DISABLED. This workflow is switched off in the repository's Actions settings
# (state: disabled_manually) and does not run on any trigger, including its cron
# and workflow_dispatch. That state lives in GitHub's UI and is invisible when
# reading this file, which has already misled at least one audit — hence this
# banner. Re-enabling is a UI action; anyone doing so should first check that the
# workflow still matches the current CI layout. See rustfs/backlog#1603.
#
name: e2e-s3tests name: e2e-s3tests
on: on:
@@ -135,6 +142,8 @@ jobs:
TEST_MODE: ${{ matrix.test-mode }} TEST_MODE: ${{ matrix.test-mode }}
steps: steps:
- uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7 - uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
with:
persist-credentials: false
# Provision Python explicitly rather than trusting the runner image to # Provision Python explicitly rather than trusting the runner image to
# ship a working pip (ci-1: a bare python3 without pip is what broke the # ship a working pip (ci-1: a bare python3 without pip is what broke the
@@ -354,6 +363,8 @@ jobs:
issues: write issues: write
steps: steps:
- uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7 - uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
with:
persist-credentials: false
- name: Open or update failure-tracking issue - name: Open or update failure-tracking issue
uses: ./.github/actions/schedule-failure-issue uses: ./.github/actions/schedule-failure-issue
with: with:
+16 -1
View File
@@ -12,6 +12,13 @@
# See the License for the specific language governing permissions and # See the License for the specific language governing permissions and
# limitations under the License. # limitations under the License.
# DISABLED. This workflow is switched off in the repository's Actions settings
# (state: disabled_manually) and does not run on any trigger, including its cron
# and workflow_dispatch. That state lives in GitHub's UI and is invisible when
# reading this file, which has already misled at least one audit — hence this
# banner. Re-enabling is a UI action; anyone doing so should first check that the
# workflow still matches the current CI layout. See rustfs/backlog#1603.
#
name: Fuzz name: Fuzz
on: on:
@@ -59,6 +66,7 @@ jobs:
name: Cancel Closed PR Runs name: Cancel Closed PR Runs
if: github.event_name == 'pull_request' && github.event.action == 'closed' if: github.event_name == 'pull_request' && github.event.action == 'closed'
runs-on: ubuntu-latest runs-on: ubuntu-latest
timeout-minutes: 10
steps: steps:
- name: Explain cancellation run - name: Explain cancellation run
run: echo "PR closed; this run only cancels older runs in the same concurrency group." run: echo "PR closed; this run only cancels older runs in the same concurrency group."
@@ -79,13 +87,14 @@ jobs:
steps: steps:
- name: Checkout repository - name: Checkout repository
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7 uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
with:
persist-credentials: false
- name: Setup Rust environment - name: Setup Rust environment
uses: ./.github/actions/setup uses: ./.github/actions/setup
with: with:
rust-version: nightly rust-version: nightly
cache-shared-key: fuzz-${{ hashFiles('fuzz/Cargo.lock') }} cache-shared-key: fuzz-${{ hashFiles('fuzz/Cargo.lock') }}
github-token: ${{ secrets.GITHUB_TOKEN }}
cache-save-if: ${{ github.ref == 'refs/heads/main' || github.event_name == 'schedule' }} cache-save-if: ${{ github.ref == 'refs/heads/main' || github.event_name == 'schedule' }}
- name: Install cargo-fuzz - name: Install cargo-fuzz
@@ -145,6 +154,8 @@ jobs:
steps: steps:
- name: Checkout repository - name: Checkout repository
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7 uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
with:
persist-credentials: false
- name: Download prebuilt fuzz binaries - name: Download prebuilt fuzz binaries
uses: actions/download-artifact@37930b1c2abaa49bbe596cd826c3c89aef350131 # v7 uses: actions/download-artifact@37930b1c2abaa49bbe596cd826c3c89aef350131 # v7
@@ -200,6 +211,8 @@ jobs:
steps: steps:
- name: Checkout repository - name: Checkout repository
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7 uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
with:
persist-credentials: false
- name: Download prebuilt fuzz binaries - name: Download prebuilt fuzz binaries
uses: actions/download-artifact@37930b1c2abaa49bbe596cd826c3c89aef350131 # v7 uses: actions/download-artifact@37930b1c2abaa49bbe596cd826c3c89aef350131 # v7
@@ -247,6 +260,8 @@ jobs:
issues: write issues: write
steps: steps:
- uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7 - uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
with:
persist-credentials: false
- name: Open or update failure-tracking issue - name: Open or update failure-tracking issue
uses: ./.github/actions/schedule-failure-issue uses: ./.github/actions/schedule-failure-issue
with: with:
+32 -7
View File
@@ -32,6 +32,7 @@ permissions:
jobs: jobs:
build-helm-package: build-helm-package:
runs-on: ubuntu-latest runs-on: ubuntu-latest
timeout-minutes: 30
if: | if: |
(github.event_name == 'workflow_dispatch' && !contains(github.event.inputs.version, '-preview')) || (github.event_name == 'workflow_dispatch' && !contains(github.event.inputs.version, '-preview')) ||
( (
@@ -49,16 +50,26 @@ jobs:
steps: steps:
- name: Checkout helm chart repo - name: Checkout helm chart repo
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7 uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
with:
persist-credentials: false
# Both inputs reach the shell through env rather than `${{ }}`
# interpolation. A git ref name may contain `$(...)` — anything without a
# space is a legal tag — and interpolation pastes it into the script
# verbatim, where bash would run it. Reading "$RAW_INPUT" instead makes it
# data.
- name: Normalize release version - name: Normalize release version
id: version id: version
env:
RAW_INPUT: ${{ github.event.inputs.version }}
RAW_BRANCH: ${{ github.event.workflow_run.head_branch }}
run: | run: |
set -eux set -eux
if [ "${{ github.event_name }}" = "workflow_dispatch" ]; then if [ "${{ github.event_name }}" = "workflow_dispatch" ]; then
RAW="${{ github.event.inputs.version }}" RAW="$RAW_INPUT"
else else
RAW="${{ github.event.workflow_run.head_branch }}" RAW="$RAW_BRANCH"
fi fi
case "$RAW" in case "$RAW" in
@@ -73,10 +84,13 @@ jobs:
./scripts/helm_chart_version.sh "$RAW_TAG" ./scripts/helm_chart_version.sh "$RAW_TAG"
- name: Replace chart version and app version - name: Replace chart version and app version
env:
CHART_VERSION: ${{ steps.version.outputs.chart_version }}
APP_VERSION: ${{ steps.version.outputs.app_version }}
run: | run: |
set -eux set -eux
sed -i -E 's/^version:.*/version: "${{ steps.version.outputs.chart_version }}"/' helm/rustfs/Chart.yaml sed -i -E "s/^version:.*/version: \"${CHART_VERSION}\"/" helm/rustfs/Chart.yaml
sed -i -E 's/^appVersion:.*/appVersion: "${{ steps.version.outputs.app_version }}"/' helm/rustfs/Chart.yaml sed -i -E "s/^appVersion:.*/appVersion: \"${APP_VERSION}\"/" helm/rustfs/Chart.yaml
- name: Set up Helm - name: Set up Helm
uses: azure/setup-helm@b9e51907a09c216f16ebe8536097933489208112 # v4.3.0 uses: azure/setup-helm@b9e51907a09c216f16ebe8536097933489208112 # v4.3.0
@@ -101,6 +115,7 @@ jobs:
publish-helm-package: publish-helm-package:
runs-on: ubuntu-latest runs-on: ubuntu-latest
timeout-minutes: 30
needs: [ build-helm-package ] needs: [ build-helm-package ]
if: needs.build-helm-package.result == 'success' if: needs.build-helm-package.result == 'success'
@@ -108,6 +123,8 @@ jobs:
- name: Checkout helm package repo - name: Checkout helm package repo
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7 uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
with: with:
# persist-credentials-exempt: this checkout's token IS the push credential —
# the job git-pushes to rustfs/helm below. Clearing it breaks chart publishing.
repository: rustfs/helm repository: rustfs/helm
token: ${{ secrets.RUSTFS_HELM_PACKAGE }} token: ${{ secrets.RUSTFS_HELM_PACKAGE }}
@@ -123,11 +140,19 @@ jobs:
- name: Generate index - name: Generate index
run: helm repo index . --url https://charts.rustfs.com run: helm repo index . --url https://charts.rustfs.com
# app_version is derived from the triggering tag name, and this job holds
# the cross-repository push token with rustfs/helm already checked out —
# the worst place in the repo to paste an attacker-influenced string into
# a shell line. Passed through env so bash treats it as data.
- name: Push helm package and index file - name: Push helm package and index file
env:
GIT_USERNAME: ${{ secrets.USERNAME }}
GIT_EMAIL: ${{ secrets.EMAIL_ADDRESS }}
APP_VERSION: ${{ needs.build-helm-package.outputs.app_version }}
run: | run: |
set -eux set -eux
git config --global user.name "${{ secrets.USERNAME }}" git config --global user.name "${GIT_USERNAME}"
git config --global user.email "${{ secrets.EMAIL_ADDRESS }}" git config --global user.email "${GIT_EMAIL}"
git add . git add .
git commit -m "Update rustfs helm package with ${{ needs.build-helm-package.outputs.app_version }}." || echo "No changes to commit" git commit -m "Update rustfs helm package with ${APP_VERSION}." || echo "No changes to commit"
git push origin main git push origin main
+8
View File
@@ -12,6 +12,13 @@
# See the License for the specific language governing permissions and # See the License for the specific language governing permissions and
# limitations under the License. # limitations under the License.
# DISABLED. This workflow is switched off in the repository's Actions settings
# (state: disabled_manually) and does not run on any trigger, including its cron
# and workflow_dispatch. That state lives in GitHub's UI and is invisible when
# reading this file, which has already misled at least one audit — hence this
# banner. Re-enabling is a UI action; anyone doing so should first check that the
# workflow still matches the current CI layout. See rustfs/backlog#1603.
#
name: "issue-translator" name: "issue-translator"
on: on:
issue_comment: issue_comment:
@@ -26,6 +33,7 @@ permissions:
jobs: jobs:
build: build:
runs-on: ubuntu-latest runs-on: ubuntu-latest
timeout-minutes: 10
steps: steps:
- uses: usthe/issues-translate-action@b41f55ddc81d7d54bd542a4f289fe28ec081898e # v2.7 - uses: usthe/issues-translate-action@b41f55ddc81d7d54bd542a4f289fe28ec081898e # v2.7
with: with:
+9 -1
View File
@@ -23,6 +23,13 @@
# Runner: GitHub-hosted `ubuntu-latest`. It reliably ships Docker + Python, # Runner: GitHub-hosted `ubuntu-latest`. It reliably ships Docker + Python,
# unlike the self-hosted fleet, whose pods drift in Docker/pip availability # unlike the self-hosted fleet, whose pods drift in Docker/pip availability
# (see the infra note in e2e-s3tests.yml). Nightly + manual only. # (see the infra note in e2e-s3tests.yml). Nightly + manual only.
# DISABLED. This workflow is switched off in the repository's Actions settings
# (state: disabled_manually) and does not run on any trigger, including its cron
# and workflow_dispatch. That state lives in GitHub's UI and is invisible when
# reading this file, which has already misled at least one audit — hence this
# banner. Re-enabling is a UI action; anyone doing so should first check that the
# workflow still matches the current CI layout. See rustfs/backlog#1603.
#
name: minio-interop name: minio-interop
on: on:
@@ -51,13 +58,14 @@ jobs:
steps: steps:
- name: Checkout repository - name: Checkout repository
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7 uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
with:
persist-credentials: false
- name: Setup Rust environment - name: Setup Rust environment
uses: ./.github/actions/setup uses: ./.github/actions/setup
with: with:
rust-version: stable rust-version: stable
cache-shared-key: ci-minio-interop cache-shared-key: ci-minio-interop
github-token: ${{ secrets.GITHUB_TOKEN }}
cache-save-if: ${{ github.ref == 'refs/heads/main' }} cache-save-if: ${{ github.ref == 'refs/heads/main' }}
- name: Generate real MinIO fixtures via Docker - name: Generate real MinIO fixtures via Docker
+11
View File
@@ -45,6 +45,13 @@
# docker-capable self-hosted `dind-sm-standard-2` label was the alternative but # docker-capable self-hosted `dind-sm-standard-2` label was the alternative but
# has fewer cores and reintroduces fleet-state risk for no reliability gain. # has fewer cores and reintroduces fleet-state risk for no reliability gain.
# DISABLED. This workflow is switched off in the repository's Actions settings
# (state: disabled_manually) and does not run on any trigger, including its cron
# and workflow_dispatch. That state lives in GitHub's UI and is invisible when
# reading this file, which has already misled at least one audit — hence this
# banner. Re-enabling is a UI action; anyone doing so should first check that the
# workflow still matches the current CI layout. See rustfs/backlog#1603.
#
name: mint name: mint
on: on:
@@ -118,6 +125,8 @@ jobs:
timeout-minutes: 120 timeout-minutes: 120
steps: steps:
- uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7 - uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
with:
persist-credentials: false
- name: Enable buildx - name: Enable buildx
uses: docker/setup-buildx-action@8d2750c68a42422c14e847fe6c8ac0403b4cbd6f # v3 uses: docker/setup-buildx-action@8d2750c68a42422c14e847fe6c8ac0403b4cbd6f # v3
@@ -263,6 +272,8 @@ jobs:
issues: write issues: write
steps: steps:
- uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7 - uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
with:
persist-credentials: false
- name: Open or update failure-tracking issue - name: Open or update failure-tracking issue
uses: ./.github/actions/schedule-failure-issue uses: ./.github/actions/schedule-failure-issue
with: with:
+57
View File
@@ -0,0 +1,57 @@
# Copyright 2024 RustFS Team
#
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
name: Nightly GNU Build
on:
schedule:
- cron: "0 0 * * *"
timezone: "Asia/Shanghai"
workflow_dispatch:
permissions:
contents: read
concurrency:
group: nightly-gnu-build-main-${{ github.event_name }}
cancel-in-progress: ${{ github.event_name == 'workflow_dispatch' }}
env:
CARGO_TERM_COLOR: always
RUST_BACKTRACE: 1
jobs:
build:
name: Build x86_64 GNU
runs-on: sm-standard-2
timeout-minutes: 150
env:
FORCE_JAVASCRIPT_ACTIONS_TO_NODE24: "true"
steps:
- name: Checkout main branch
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
with:
persist-credentials: false
ref: main
- name: Setup Rust environment
uses: ./.github/actions/setup
with:
cache-shared-key: build-x86_64-unknown-linux-gnu
cache-save-if: 'false'
install-build-packaging-tools: 'false'
install-test-tools: 'false'
- name: Build RustFS
run: cargo build --release --locked --target x86_64-unknown-linux-gnu -p rustfs --bins
+9 -2
View File
@@ -19,9 +19,12 @@ on:
schedule: schedule:
- cron: '0 5 * * 0' # Weekly on Sunday 05:00 UTC (staggered after the midnight ci/build crons) - cron: '0 5 * * 0' # Weekly on Sunday 05:00 UTC (staggered after the midnight ci/build crons)
# GITHUB_TOKEN only needs to read the repository here: the branch push and the
# pull request are both created by update-flake-lock using the
# FLAKE_UPDATE_TOKEN PAT below, not by this token. Leaving write on it hands a
# repo-write credential to an unattended weekly job that does not use it.
permissions: permissions:
contents: write contents: read
pull-requests: write
concurrency: concurrency:
group: ${{ github.workflow }}-${{ github.ref }} group: ${{ github.workflow }}-${{ github.ref }}
@@ -37,6 +40,10 @@ jobs:
steps: steps:
- name: Checkout repository - name: Checkout repository
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7 uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
# persist-credentials-exempt: update-flake-lock pushes the branch and opens
# the PR. It passes FLAKE_UPDATE_TOKEN to create-pull-request itself rather
# than reusing .git/config, but that is unverified — exempt until a
# workflow_dispatch run confirms it (rustfs/backlog#1602).
- name: Install Nix - name: Install Nix
uses: DeterminateSystems/determinate-nix-action@629b284231c2a82554b724e357e47fc6020833c8 # v3 uses: DeterminateSystems/determinate-nix-action@629b284231c2a82554b724e357e47fc6020833c8 # v3
+10
View File
@@ -12,6 +12,13 @@
# See the License for the specific language governing permissions and # See the License for the specific language governing permissions and
# limitations under the License. # limitations under the License.
# DISABLED. This workflow is switched off in the repository's Actions settings
# (state: disabled_manually) and does not run on any trigger, including its cron
# and workflow_dispatch. That state lives in GitHub's UI and is invisible when
# reading this file, which has already misled at least one audit — hence this
# banner. Re-enabling is a UI action; anyone doing so should first check that the
# workflow still matches the current CI layout. See rustfs/backlog#1603.
#
name: Nix CI name: Nix CI
on: on:
@@ -46,6 +53,7 @@ jobs:
name: Cancel Closed PR Runs name: Cancel Closed PR Runs
if: github.event_name == 'pull_request' && github.event.action == 'closed' if: github.event_name == 'pull_request' && github.event.action == 'closed'
runs-on: ubuntu-latest runs-on: ubuntu-latest
timeout-minutes: 10
steps: steps:
- name: Explain cancellation run - name: Explain cancellation run
run: echo "PR closed; this run only cancels older runs in the same concurrency group." run: echo "PR closed; this run only cancels older runs in the same concurrency group."
@@ -63,6 +71,8 @@ jobs:
steps: steps:
- name: Checkout repository - name: Checkout repository
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7 uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
with:
persist-credentials: false
- name: Install Nix - name: Install Nix
uses: DeterminateSystems/determinate-nix-action@4eea0b33e3d1f02ecfe37cf16e7204c424009606 # v3.21.0 uses: DeterminateSystems/determinate-nix-action@4eea0b33e3d1f02ecfe37cf16e7204c424009606 # v3.21.0
+71 -29
View File
@@ -22,6 +22,13 @@
# a deliberate correctness cost (e.g. the #4221 fsync durability fix) is # a deliberate correctness cost (e.g. the #4221 fsync durability fix) is
# recorded but does not block (rustfs/backlog#935 correction 1). # recorded but does not block (rustfs/backlog#935 correction 1).
# DISABLED. This workflow is switched off in the repository's Actions settings
# (state: disabled_manually) and does not run on any trigger, including its cron
# and workflow_dispatch. That state lives in GitHub's UI and is invisible when
# reading this file, which has already misled at least one audit — hence this
# banner. Re-enabling is a UI action; anyone doing so should first check that the
# workflow still matches the current CI layout. See rustfs/backlog#1603.
#
name: Performance A/B name: Performance A/B
on: on:
@@ -92,6 +99,8 @@ jobs:
steps: steps:
- name: Checkout repository - name: Checkout repository
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7 uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
with:
persist-credentials: false
- name: Setup Rust environment - name: Setup Rust environment
uses: ./.github/actions/setup uses: ./.github/actions/setup
@@ -99,7 +108,6 @@ jobs:
rust-version: stable rust-version: stable
cache-shared-key: warp-ab-${{ hashFiles('**/Cargo.lock') }} cache-shared-key: warp-ab-${{ hashFiles('**/Cargo.lock') }}
cache-save-if: ${{ github.ref == 'refs/heads/main' }} cache-save-if: ${{ github.ref == 'refs/heads/main' }}
github-token: ${{ secrets.GITHUB_TOKEN }}
- name: Build release rustfs - name: Build release rustfs
run: cargo build --release --bin rustfs run: cargo build --release --bin rustfs
@@ -142,6 +150,7 @@ jobs:
- name: Checkout repository - name: Checkout repository
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7 uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
with: with:
persist-credentials: false
fetch-depth: 0 # baseline is built from origin/main fetch-depth: 0 # baseline is built from origin/main
- name: Setup Rust environment - name: Setup Rust environment
@@ -150,7 +159,6 @@ jobs:
rust-version: stable rust-version: stable
cache-shared-key: warp-ab-${{ hashFiles('**/Cargo.lock') }} cache-shared-key: warp-ab-${{ hashFiles('**/Cargo.lock') }}
cache-save-if: ${{ github.ref == 'refs/heads/main' }} cache-save-if: ${{ github.ref == 'refs/heads/main' }}
github-token: ${{ secrets.GITHUB_TOKEN }}
- name: Install warp - name: Install warp
run: | run: |
@@ -162,13 +170,15 @@ jobs:
- name: Decide exemption - name: Decide exemption
id: exempt id: exempt
env:
INPUT_ALLOW_REGRESSION: ${{ github.event.inputs.allow_regression }}
run: | run: |
allow="false" allow="false"
if [[ "${{ github.event_name }}" == "pull_request" ]] \ if [[ "${{ github.event_name }}" == "pull_request" ]] \
&& ${{ contains(github.event.pull_request.labels.*.name, 'perf-deliberate-tradeoff') }}; then && ${{ contains(github.event.pull_request.labels.*.name, 'perf-deliberate-tradeoff') }}; then
allow="true" allow="true"
fi fi
if [[ "${{ github.event.inputs.allow_regression }}" == "true" ]]; then if [[ "$INPUT_ALLOW_REGRESSION" == "true" ]]; then
allow="true" allow="true"
fi fi
echo "allow_regression=$allow" >> "$GITHUB_OUTPUT" echo "allow_regression=$allow" >> "$GITHUB_OUTPUT"
@@ -215,8 +225,34 @@ jobs:
cp target/release/rustfs baseline-bin/rustfs cp target/release/rustfs baseline-bin/rustfs
echo "built=true" >> "$GITHUB_OUTPUT" echo "built=true" >> "$GITHUB_OUTPUT"
- name: Build baseline on cache miss (different candidate)
id: baseline_build
if: >-
steps.baseline_cache.outputs.cache-hit != 'true' &&
steps.commits.outputs.baseline_sha != steps.commits.outputs.candidate_sha
run: |
set -euo pipefail
baseline_root="$RUNNER_TEMP/rustfs-baseline-${{ github.run_id }}"
baseline_target="$RUNNER_TEMP/rustfs-baseline-target-${{ github.run_id }}"
git worktree add --detach "$baseline_root" "${{ steps.commits.outputs.baseline_sha }}"
cargo build --release --manifest-path "$baseline_root/Cargo.toml" --bin rustfs --target-dir "$baseline_target"
mkdir -p baseline-bin
cp "$baseline_target/release/rustfs" baseline-bin/rustfs
git worktree remove --force "$baseline_root"
echo "built=true" >> "$GITHUB_OUTPUT"
- name: Build candidate binary
id: candidate_build
if: steps.commits.outputs.baseline_sha != steps.commits.outputs.candidate_sha
run: |
set -euo pipefail
cargo build --release --bin rustfs
mkdir -p candidate-bin
cp target/release/rustfs candidate-bin/rustfs
echo "built=true" >> "$GITHUB_OUTPUT"
- name: Save self-healed baseline to cache - name: Save self-healed baseline to cache
if: steps.selfheal.outputs.built == 'true' if: steps.selfheal.outputs.built == 'true' || steps.baseline_build.outputs.built == 'true'
uses: actions/cache/save@55cc8345863c7cc4c66a329aec7e433d2d1c52a9 # v6 uses: actions/cache/save@55cc8345863c7cc4c66a329aec7e433d2d1c52a9 # v6
with: with:
path: baseline-bin/rustfs path: baseline-bin/rustfs
@@ -224,54 +260,58 @@ jobs:
- name: Run warp A/B and gate - name: Run warp A/B and gate
id: ab id: ab
env:
INPUT_DURATION: ${{ github.event.inputs.duration }}
run: | run: |
set -euo pipefail set -euo pipefail
# Budget note: with perf-3's cached baseline the nightly does no source # The formal runner executes A1 baseline -> B1 candidate -> B2 candidate
# build on a cache hit, so the wall-clock is dominated by the short warp # -> A2 baseline for each workload and drive-sync cell. It requires three
# matrix — duration/rounds/cooldown are kept small to fit all 24 cells # rounds per leg to emit tail latency and error-rate evidence.
# (6 workloads x 2 phases x 2 drive-sync) rather than dropping cells.
# --health-timeout 180 outlasts the server's own 120s startup-readiness # --health-timeout 180 outlasts the server's own 120s startup-readiness
# budget, which the rig's previous 60s health poll undershot (the first # budget, which the rig's previous 60s health poll undershot (the first
# two nightly failures). perf-6 recalibrates these once the noise study # two nightly failures). perf-6 recalibrates these once the noise study
# lands. # lands.
duration="${{ github.event.inputs.duration || '12s' }}" duration="${INPUT_DURATION:-12s}"
baseline_sha="${{ steps.commits.outputs.baseline_sha }}" baseline_sha="${{ steps.commits.outputs.baseline_sha }}"
candidate_sha="${{ steps.commits.outputs.candidate_sha }}" candidate_sha="${{ steps.commits.outputs.candidate_sha }}"
baseline_hit="${{ steps.baseline_cache.outputs.cache-hit }}" baseline_hit="${{ steps.baseline_cache.outputs.cache-hit }}"
selfheal_built="${{ steps.selfheal.outputs.built }}" selfheal_built="${{ steps.selfheal.outputs.built }}"
baseline_built="${{ steps.baseline_build.outputs.built }}"
candidate_built="${{ steps.candidate_build.outputs.built }}"
args=(--duration "$duration" --rounds 2 --cooldown 5 --health-timeout 180) args=(--duration "$duration" --rounds 3 --cooldown 5 --health-timeout 180 --baseline-revision "$baseline_sha" --candidate-revision "$candidate_sha")
if [[ "$baseline_hit" == "true" || "$selfheal_built" == "true" ]]; then if [[ "$baseline_hit" == "true" || "$selfheal_built" == "true" || "$baseline_built" == "true" ]]; then
chmod +x baseline-bin/rustfs chmod +x baseline-bin/rustfs
base_bin="$PWD/baseline-bin/rustfs" base_bin="$PWD/baseline-bin/rustfs"
args+=(--baseline-bin "$base_bin") args+=(--baseline-bin "$base_bin")
if [[ "$baseline_hit" == "true" ]]; then if [[ "$baseline_hit" == "true" ]]; then
base_src="actions-cache (rustfs-baseline-$baseline_sha)" base_src="actions-cache (rustfs-baseline-$baseline_sha)"
else elif [[ "$selfheal_built" == "true" ]]; then
base_src="source build (cache self-heal, saved as rustfs-baseline-$baseline_sha)" base_src="source build (cache self-heal, saved as rustfs-baseline-$baseline_sha)"
else
base_src="isolated origin/main source build (saved as rustfs-baseline-$baseline_sha)"
fi fi
if [[ "$candidate_sha" == "$baseline_sha" ]]; then if [[ "$candidate_sha" == "$baseline_sha" ]]; then
# Nightly on main: the candidate is the same commit as the baseline, # Nightly on main: the candidate is the same commit as the baseline,
# so reuse the one binary for both phases and skip all builds. # so reuse the one binary for both phases and skip all builds.
args+=(--candidate-bin "$base_bin" --skip-build) args+=(--candidate-bin "$base_bin")
cand_src="same binary as baseline (same commit)" cand_src="same binary as baseline (same commit)"
else elif [[ "$candidate_built" == "true" ]]; then
chmod +x candidate-bin/rustfs
args+=(--candidate-bin "$PWD/candidate-bin/rustfs")
cand_src="source build of the checked-out ref" cand_src="source build of the checked-out ref"
else
echo "::error::candidate binary was not built" >&2
exit 2
fi fi
else else
# Cache miss with candidate != baseline (opt-in PR gate only): fall echo "::error::baseline binary was not restored or built" >&2
# back to the source double-build. With the post-#4806 LTO profile exit 2
# this will overrun the job budget and alert; rerun once the push
# cache build for origin/main has completed, or wait for perf-7's
# merge-base caching.
args+=(--baseline-ref origin/main)
base_src="source build of origin/main (cache miss)"
cand_src="source build of the checked-out ref"
fi fi
args+=(--provenance-note "baseline commit: $baseline_sha - $base_src") echo "baseline binary: $base_src"
args+=(--provenance-note "candidate commit: $candidate_sha - $cand_src") echo "candidate binary: $cand_src"
if [[ "${{ steps.exempt.outputs.allow_regression }}" == "true" ]]; then if [[ "${{ steps.exempt.outputs.allow_regression }}" == "true" ]]; then
args+=(--allow-regression --exemption-reason "labeled perf-deliberate-tradeoff / dispatch override") args+=(--allow-regression --exemption-reason "labeled perf-deliberate-tradeoff / dispatch override")
@@ -279,7 +319,7 @@ jobs:
# Do not let a gate FAIL abort the job here; capture status and surface # Do not let a gate FAIL abort the job here; capture status and surface
# it after the PR comment is posted. # it after the PR comment is posted.
set +e set +e
bash scripts/run_hotpath_warp_ab.sh "${args[@]}" bash scripts/run_hotpath_warp_abba.sh "${args[@]}"
echo "status=$?" >> "$GITHUB_OUTPUT" echo "status=$?" >> "$GITHUB_OUTPUT"
set -e set -e
# Locate the newest run dir + gate.md for the summary/comment/artifact # Locate the newest run dir + gate.md for the summary/comment/artifact
@@ -287,10 +327,10 @@ jobs:
# holds server-logs/ for diagnosis. # holds server-logs/ for diagnosis.
# Run dirs are UTC-timestamp names (no special chars); ls is safe here. # Run dirs are UTC-timestamp names (no special chars); ls is safe here.
# shellcheck disable=SC2012 # shellcheck disable=SC2012
run_dir="$(ls -td target/hotpath-ab/*/ 2>/dev/null | head -n1 || true)" run_dir="$(ls -td target/hotpath-abba/*/ 2>/dev/null | head -n1 || true)"
echo "run_dir=${run_dir%/}" >> "$GITHUB_OUTPUT" echo "run_dir=${run_dir%/}" >> "$GITHUB_OUTPUT"
# shellcheck disable=SC2012 # shellcheck disable=SC2012
gate_md="$(ls -t target/hotpath-ab/*/gate.md 2>/dev/null | head -n1 || true)" gate_md="$(ls -t target/hotpath-abba/*/candidate_gate.md 2>/dev/null | head -n1 || true)"
echo "gate_md=$gate_md" >> "$GITHUB_OUTPUT" echo "gate_md=$gate_md" >> "$GITHUB_OUTPUT"
- name: Upload A/B results - name: Upload A/B results
@@ -298,10 +338,10 @@ jobs:
uses: actions/upload-artifact@b7c566a772e6b6bfb58ed0dc250532a479d7789f # v6 uses: actions/upload-artifact@b7c566a772e6b6bfb58ed0dc250532a479d7789f # v6
with: with:
name: hotpath-warp-ab-${{ github.run_number }} name: hotpath-warp-ab-${{ github.run_number }}
# Includes per-cell median_summary.csv / baseline_compare.csv, gate.md, # Includes per-cell median_summary.csv / baseline_compare.csv, both gates,
# and server-logs/ (rustfs.log + startup env per phase) so a failed run # and server-logs/ (rustfs.log + startup env per phase) so a failed run
# is diagnosable. Short retention: this is churny nightly debug data. # is diagnosable. Short retention: this is churny nightly debug data.
path: target/hotpath-ab/ path: target/hotpath-abba/
if-no-files-found: warn if-no-files-found: warn
retention-days: 14 retention-days: 14
@@ -385,6 +425,8 @@ jobs:
issues: write issues: write
steps: steps:
- uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7 - uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
with:
persist-credentials: false
- name: Open or update failure-tracking issue - name: Open or update failure-tracking issue
uses: ./.github/actions/schedule-failure-issue uses: ./.github/actions/schedule-failure-issue
with: with:
+81
View File
@@ -0,0 +1,81 @@
# Copyright 2026 RustFS Team
#
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
# Asserts that the self-hosted runners are still ephemeral — one job per pod.
#
# This repository is public and its pull_request jobs run on those runners,
# executing the PR's own build.rs, proc-macros and tests. The only thing keeping
# that code from reaching a later job is that each ARC pod handles exactly one
# job and is then destroyed. That guarantee lives in the ARC scale-set
# configuration, outside this repository, where it can be changed without any PR
# — so it is asserted here from the outside, against real run data, instead of
# being assumed.
#
# Monthly rather than per-PR: the property changes only when someone
# reconfigures the scale set, and the check costs a few dozen API calls.
# See docs/ci/runners.md and rustfs/backlog#1602.
name: Runner Hygiene
on:
schedule:
- cron: "0 6 1 * *" # Monthly, 1st at 06:00 UTC (after the daily audit cron)
workflow_dispatch:
permissions:
contents: read
concurrency:
group: runner-hygiene
cancel-in-progress: false
jobs:
check-ephemerality:
name: Check runner ephemerality
runs-on: ubuntu-latest
timeout-minutes: 15
steps:
- name: Checkout repository
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
with:
persist-credentials: false
# Exit 2 (inconclusive / broken) is deliberately not a pass: a window
# where every sm-* job was still queued would otherwise look identical to
# a clean bill of health.
- name: Assert one job per self-hosted runner
env:
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
run: ./scripts/ci/check_runner_ephemerality.sh 40
alert-on-failure:
name: Alert on scheduled failure
needs: [check-ephemerality]
# Same ci-8 mechanism as coverage.yml, audit.yml and the nightly lanes:
# scheduled runs file a tracking issue, manual dispatch stays quiet so
# debugging never produces a spurious alert.
if: always() && github.event_name == 'schedule' && contains(needs.*.result, 'failure')
runs-on: ubuntu-latest
timeout-minutes: 10
permissions:
contents: read
issues: write
steps:
- uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
with:
persist-credentials: false
- name: Open or update failure-tracking issue
uses: ./.github/actions/schedule-failure-issue
with:
github-token: ${{ secrets.GITHUB_TOKEN }}
@@ -24,6 +24,13 @@
# The run itself is expected to end red (the forced failure); only the # The run itself is expected to end red (the forced failure); only the
# alert-on-failure job result matters. # alert-on-failure job result matters.
# DISABLED. This workflow is switched off in the repository's Actions settings
# (state: disabled_manually) and does not run on any trigger, including its cron
# and workflow_dispatch. That state lives in GitHub's UI and is invisible when
# reading this file, which has already misled at least one audit — hence this
# banner. Re-enabling is a UI action; anyone doing so should first check that the
# workflow still matches the current CI layout. See rustfs/backlog#1603.
#
name: Schedule Failure Alert Drill name: Schedule Failure Alert Drill
on: on:
@@ -56,6 +63,8 @@ jobs:
issues: write issues: write
steps: steps:
- uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7 - uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
with:
persist-credentials: false
- name: Open or update failure-tracking issue - name: Open or update failure-tracking issue
uses: ./.github/actions/schedule-failure-issue uses: ./.github/actions/schedule-failure-issue
with: with:
+8
View File
@@ -12,6 +12,13 @@
# See the License for the specific language governing permissions and # See the License for the specific language governing permissions and
# limitations under the License. # limitations under the License.
# DISABLED. This workflow is switched off in the repository's Actions settings
# (state: disabled_manually) and does not run on any trigger, including its cron
# and workflow_dispatch. That state lives in GitHub's UI and is invisible when
# reading this file, which has already misled at least one audit — hence this
# banner. Re-enabling is a UI action; anyone doing so should first check that the
# workflow still matches the current CI layout. See rustfs/backlog#1603.
#
name: "Mark stale issues" name: "Mark stale issues"
on: on:
schedule: schedule:
@@ -20,6 +27,7 @@ on:
jobs: jobs:
stale: stale:
runs-on: ubuntu-latest runs-on: ubuntu-latest
timeout-minutes: 30
steps: steps:
- uses: actions/stale@5bef64f19d7facfb25b37b414482c7164d639639 # v9 - uses: actions/stale@5bef64f19d7facfb25b37b414482c7164d639639 # v9
with: with:
+1
View File
@@ -15,6 +15,7 @@ concurrency:
jobs: jobs:
update: update:
runs-on: ubuntu-latest runs-on: ubuntu-latest
timeout-minutes: 10
steps: steps:
- uses: overtrue/repo-visuals-action@72f34d24769ff5d341956da2f23952594ef2f1e2 # v1.3.0 - uses: overtrue/repo-visuals-action@72f34d24769ff5d341956da2f23952594ef2f1e2 # v1.3.0
with: with:
+80 -19
View File
@@ -21,6 +21,22 @@ If repo-level instructions conflict, follow the nearest file and keep behavior a
- Avoid redundant file reads, repeated commands, and unnecessary exploratory work once enough context is available. - Avoid redundant file reads, repeated commands, and unnecessary exploratory work once enough context is available.
- A good result is a minimal diff with clear assumptions, no over-engineering, and independent verification that survives Adversarial Validation (below). - A good result is a minimal diff with clear assumptions, no over-engineering, and independent verification that survives Adversarial Validation (below).
## Worktree and Disk Hygiene
- Unless the requester explicitly says otherwise, treat every new implementation task as isolated work: fetch the latest `origin/main`, confirm the requested change is not already present there, and create a dedicated feature branch and worktree from that exact upstream commit before editing. Do not implement new work directly in the primary checkout or reuse a worktree from another task.
- Check available disk space before creating the worktree or starting dependency downloads, builds, tests, coverage, or other artifact-heavy commands. For long-running or artifact-heavy work, re-check disk usage at natural phase boundaries and before broad validation; if remaining space may not safely accommodate the next command, stop and reclaim task-owned artifacts before continuing.
- Keep cleanup scoped and safe: remove generated build/test/coverage artifacts and temporary files created by the task when they are no longer needed, and never delete another task's worktree or uncommitted files. Prefer shared dependency caches where supported instead of duplicating large artifacts across worktrees.
- At handoff, report the disk-space checks, cleanup performed, and any retained worktree or artifacts with the reason they are still needed.
## PR Lifecycle Monitoring
- Creating or updating a PR is not the terminal state. Unless the requester explicitly limits the task to PR creation, monitor the PR through its terminal state: merged, closed, or explicitly handed off because progress requires user or maintainer action.
- While the task is active, monitor CI/check runs, review decisions and unresolved threads, mergeability and conflicts, and unexpected head/base changes. Prefer event-driven or bounded waits provided by the current environment over frequent polling; report only state changes, actionable failures, or meaningful prolonged delays.
- Investigate every failing check and review comment before changing code. Fix failures attributable to the task, run the verification required for the new diff, push the update, respond to or resolve the corresponding review threads, and resume monitoring. Do not weaken checks, dismiss valid feedback, or retry flaky failures merely to obtain a green result.
- Treat opening, green CI, approval, and mergeability as intermediate states. Never merge without the required reviewer approval or explicit authority. If progress depends on credentials, infrastructure, a maintainer decision, or another external action, report the exact blocker and the evidence already collected.
- If the current execution environment cannot remain active until the next PR event, use a supported automation, monitor, or thread wakeup when available and within scope. Otherwise leave an explicit handoff containing the PR, current state, next event to observe, and pending cleanup; do not imply that background monitoring exists when none is scheduled.
- After observing a merge, verify the commits are preserved on the upstream base, ensure the worktree is clean, remove the dedicated worktree, prune stale worktree metadata, and delete the local task branch when it is no longer in use. For a closed or abandoned PR, preserve any unmerged work unless deletion was explicitly authorized. Do not delete remote branches unless explicitly requested or repository automation owns that cleanup.
## Autonomy and Approval Boundaries ## Autonomy and Approval Boundaries
- Inquiry tasks (answer, explain, review, diagnose, plan): report findings; do not change files unless a fix is explicitly requested. - Inquiry tasks (answer, explain, review, diagnose, plan): report findings; do not change files unless a fix is explicitly requested.
@@ -105,33 +121,78 @@ CI) fails the build if anything is committed under `docs/superpowers/`, even via
## Verification Before PR ## Verification Before PR
Convert changes into independently verifiable outcomes. Prefer focused tests for behavior changes and run the relevant checks before declaring completion. Convert changes into independently verifiable outcomes. This section controls
Non-exempt changes must also pass Adversarial Validation (next section) before the checks below count as completion. agent-run local validation; preparing a commit or PR does not by itself require
the broadest gate. Inspect only the final task-owned diff, classify it by
behavioral impact rather than line count or path alone, and run the smallest
set of checks that provides meaningful coverage. Do not let unrelated
worktree changes or a generic contributor checklist expand the scope.
Non-exempt changes must also pass Adversarial Validation (next section) before
the checks below count as completion.
For code changes, run and pass the following before opening a PR: ### Validation floor
```bash - Every change that is not documentation-only must finish with
make pre-pr `cargo fmt --all --check` passing. An umbrella gate that runs this exact
``` check satisfies the requirement; do not run it twice. Use `cargo fmt --all`
only when formatting needs to be fixed. Run the configured formatter or
validator for other changed languages when one exists.
- Documentation-only or instruction-only means all task-owned changes are
prose or documentation assets and cannot affect runtime, builds, CI,
dependencies, generated code, or tests. Run `git diff --check` and any
relevant documentation guard, but skip Cargo formatting, compilation,
Clippy, tests, `make pre-commit`, and `make pre-pr`.
- Behavior changes require relevant existing or new tests. Prefer the most
focused test or affected package. A passing targeted test can also provide
sufficient compilation coverage when it builds every changed target and
feature involved; do not add a redundant `cargo check` in that case.
- `cargo check` supplements compilation coverage; it never substitutes for a
behavioral test. If a relevant test cannot reasonably be added or run, use
the narrowest compilation check and report the reason and remaining risk.
Before committing code changes, prefer focused verification for the touched ### Validation tiers
surface and use the faster local gate when a broad smoke check is needed:
```bash 1. **Documentation/instruction-only:** Apply the exemption above. Run a guard
make pre-commit such as `make doc-paths-check` only when it is relevant to the edited text.
``` 2. **Non-behavioral source change:** For comments, formatting, or another
demonstrably non-executable change, run the formatting floor. Compilation,
Clippy, and tests may be skipped only when the edit cannot affect
compilation or runtime behavior; run targeted doctests if executable
documentation examples changed.
3. **Localized or bounded behavior change:** Run the formatting floor and the
narrowest relevant tests. Add package-scoped `cargo check` or Clippy only
for changed targets, features, APIs, error handling, async behavior, or
control flow not already covered. When several crates are affected but the
dependency set is identifiable, validate those packages and known
dependents instead of the whole workspace. Use `make pre-commit` only when
a repository-wide fast gate adds useful confidence beyond those checks.
4. **Broad or high-risk change:** Run `make pre-pr` only when targeted coverage
cannot bound the impact, including:
- dependency, feature, build-script, procedural-macro, code-generation,
toolchain, or CI changes that alter compilation or the test matrix;
- cross-crate public APIs, shared foundational code, or broad refactors with
an unbounded dependent set;
- locking, storage durability or formats, erasure coding, replication,
RPC/protocol compatibility, IAM/KMS/auth, cryptography, or other
security-sensitive behavior;
- a targeted check that reveals wider impact, an explicit user request, or
a release policy that requires the full gate.
For migration batches, do not run the full `make pre-pr` gate before every Documentation-only and non-behavioral classifications take precedence over
intermediate commit. Use focused tests and `make pre-commit` during path-based triggers. A small diff can still be high-risk, while a CI comment,
development, then reserve `make pre-pr` for the final PR-ready branch. manifest comment, or release-note edit does not require full validation.
Before pushing code changes, make sure formatting is clean: `make pre-pr` includes `make pre-commit` coverage. Never run both for the same
unchanged diff, and do not repeat equivalent checks during PR preparation or
because a local hook already ran them. Rerun only checks whose scope is affected
by later edits. Full workspace checks do not replace a relevant integration or
E2E test for changed behavior; run that focused test when required and
available, or report why it was not run and the remaining risk.
- Run `cargo fmt --all`. If `make` is unavailable, run the equivalent checks defined under
- Run `cargo fmt --all --check` and ensure no files are modified unexpectedly. `.config/make/`. At handoff, list the checks actually run, checks intentionally
skipped, and the reason for the selected tier.
If `make` is unavailable, run the equivalent checks defined under `.config/make/`.
Documentation-only or instruction-only changes are exempt from the verification commands above (including the `.config/make/` equivalents), though any locally installed git pre-commit hooks may still run on commit unless explicitly skipped.
After build-based verification completes, clean generated build artifacts before wrapping up to avoid unnecessary disk usage. After build-based verification completes, clean generated build artifacts before wrapping up to avoid unnecessary disk usage.
Do not open a PR with code changes when the required checks fail. Do not open a PR with code changes when the required checks fail.
Make a failing check pass by fixing the cause, never by weakening the gate: Make a failing check pass by fixing the cause, never by weakening the gate:
Generated
+152 -36
View File
@@ -290,11 +290,11 @@ dependencies = [
[[package]] [[package]]
name = "ar_archive_writer" name = "ar_archive_writer"
version = "0.5.2" version = "0.5.3"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "4087686b4b0a3427190bae57a1d9a478dbb2d40c5dc1bd6e2b6d797913bdd348" checksum = "73cd58deff2140a0a8eae87e417bd01db68a33e148aa93d1e8cd837e55e312b6"
dependencies = [ dependencies = [
"object 0.37.3", "object 0.39.1",
] ]
[[package]] [[package]]
@@ -599,6 +599,16 @@ dependencies = [
"syn 2.0.119", "syn 2.0.119",
] ]
[[package]]
name = "assert-json-diff"
version = "2.0.2"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "47e4f2b81832e72834d7518d8487a0396a28cc408186a2e8854c0f98011faf12"
dependencies = [
"serde",
"serde_json",
]
[[package]] [[package]]
name = "astral-tokio-tar" name = "astral-tokio-tar"
version = "0.6.4" version = "0.6.4"
@@ -943,6 +953,32 @@ dependencies = [
"uuid", "uuid",
] ]
[[package]]
name = "aws-sdk-kms"
version = "1.114.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "c0b7d906608ee41e7ddea9983577ba82200435644d567d63dc34e822e088b453"
dependencies = [
"arc-swap",
"aws-credential-types",
"aws-runtime",
"aws-smithy-async",
"aws-smithy-http",
"aws-smithy-json",
"aws-smithy-observability",
"aws-smithy-runtime",
"aws-smithy-runtime-api",
"aws-smithy-schema",
"aws-smithy-types",
"aws-types",
"bytes",
"fastrand",
"http 0.2.12",
"http 1.5.0",
"regex-lite",
"tracing",
]
[[package]] [[package]]
name = "aws-sdk-s3" name = "aws-sdk-s3"
version = "1.140.0" version = "1.140.0"
@@ -1158,17 +1194,23 @@ source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "635d23afda0a6ab48d666c4d447c4873e8d1e83518a2be2093122397e50b838e" checksum = "635d23afda0a6ab48d666c4d447c4873e8d1e83518a2be2093122397e50b838e"
dependencies = [ dependencies = [
"aws-smithy-async", "aws-smithy-async",
"aws-smithy-protocol-test",
"aws-smithy-runtime-api", "aws-smithy-runtime-api",
"aws-smithy-types", "aws-smithy-types",
"bytes",
"h2", "h2",
"http 1.5.0", "http 1.5.0",
"http-body 1.1.0",
"hyper", "hyper",
"hyper-rustls", "hyper-rustls",
"hyper-util", "hyper-util",
"indexmap 2.14.0",
"pin-project-lite", "pin-project-lite",
"rustls", "rustls",
"rustls-native-certs", "rustls-native-certs",
"rustls-pki-types", "rustls-pki-types",
"serde",
"serde_json",
"tokio", "tokio",
"tokio-rustls", "tokio-rustls",
"tower", "tower",
@@ -1195,6 +1237,25 @@ dependencies = [
"aws-smithy-runtime-api", "aws-smithy-runtime-api",
] ]
[[package]]
name = "aws-smithy-protocol-test"
version = "0.64.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "f76511a0e223ce78deb6a78b8afebda99cb737cfbc8a58d96dcb190f012dd40a"
dependencies = [
"assert-json-diff",
"aws-smithy-runtime-api",
"base64-simd",
"cbor-diag",
"ciborium",
"http 0.2.12",
"pretty_assertions",
"regex-lite",
"roxmltree",
"serde_json",
"thiserror 2.0.19",
]
[[package]] [[package]]
name = "aws-smithy-query" name = "aws-smithy-query"
version = "0.62.0" version = "0.62.0"
@@ -1528,7 +1589,7 @@ version = "0.10.4"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "3078c7629b62d3f0439517fa394996acacc5cbc91c5a20d8c658e77abd503a71" checksum = "3078c7629b62d3f0439517fa394996acacc5cbc91c5a20d8c658e77abd503a71"
dependencies = [ dependencies = [
"generic-array 0.14.7", "generic-array 0.14.9",
] ]
[[package]] [[package]]
@@ -1547,7 +1608,7 @@ version = "0.3.3"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "a8894febbff9f758034a5b8e12d87918f56dfc64a8e1fe757d65e29041538d93" checksum = "a8894febbff9f758034a5b8e12d87918f56dfc64a8e1fe757d65e29041538d93"
dependencies = [ dependencies = [
"generic-array 0.14.7", "generic-array 0.14.9",
] ]
[[package]] [[package]]
@@ -1598,7 +1659,7 @@ version = "3.9.3"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "6dee98b0db6a962de883bf5d20362dee4d7ca0d12fe39a7c6c73c844e1cd7c1f" checksum = "6dee98b0db6a962de883bf5d20362dee4d7ca0d12fe39a7c6c73c844e1cd7c1f"
dependencies = [ dependencies = [
"darling 0.20.11", "darling 0.23.0",
"ident_case", "ident_case",
"prettyplease", "prettyplease",
"proc-macro2", "proc-macro2",
@@ -1679,9 +1740,9 @@ dependencies = [
[[package]] [[package]]
name = "bytesize" name = "bytesize"
version = "2.4.2" version = "2.6.0"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "3d7c8918969267b2932ffd5655509bbbea0833823058c378876953217f5fc50e" checksum = "351a3e803ee3c6eaeee6b00076b767514b37c32a73d326c3ec7abddb7d6c3493"
[[package]] [[package]]
name = "bytestring" name = "bytestring"
@@ -1758,6 +1819,25 @@ dependencies = [
"cipher 0.5.2", "cipher 0.5.2",
] ]
[[package]]
name = "cbor-diag"
version = "0.1.12"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "dc245b6ecd09b23901a4fbad1ad975701fd5061ceaef6afa93a2d70605a64429"
dependencies = [
"bs58",
"chrono",
"data-encoding",
"half",
"nom 7.1.3",
"num-bigint",
"num-rational",
"num-traits",
"separator",
"url",
"uuid",
]
[[package]] [[package]]
name = "cc" name = "cc"
version = "1.4.0" version = "1.4.0"
@@ -1870,7 +1950,7 @@ version = "0.4.4"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "773f3b9af64447d2ce9850330c473515014aa235e6a783b02db81ff39e4a3dad" checksum = "773f3b9af64447d2ce9850330c473515014aa235e6a783b02db81ff39e4a3dad"
dependencies = [ dependencies = [
"crypto-common 0.1.7", "crypto-common 0.1.6",
"inout 0.1.4", "inout 0.1.4",
] ]
@@ -1888,9 +1968,9 @@ dependencies = [
[[package]] [[package]]
name = "clap" name = "clap"
version = "4.6.4" version = "4.6.5"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "d91e0c145792ef73a6ad36d27c75ac09f1832222a3c209689d90f534685ee5b7" checksum = "301b56658598e48f3648647ac6fc887be7e7108eddfa4e9b63fcf3ec58c0cadf"
dependencies = [ dependencies = [
"clap_builder", "clap_builder",
"clap_derive", "clap_derive",
@@ -1898,9 +1978,9 @@ dependencies = [
[[package]] [[package]]
name = "clap_builder" name = "clap_builder"
version = "4.6.2" version = "4.6.5"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "f09628afdcc538b57f3c6341e9c8e9970f18e4a481690a64974d7023bd33548b" checksum = "94a65403d1a1bd28f7dc68eb8506e8874808ee5eecb59298de588e2e1407a078"
dependencies = [ dependencies = [
"anstream", "anstream",
"anstyle", "anstyle",
@@ -2328,7 +2408,7 @@ version = "0.5.5"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "0dc92fb57ca44df6db8059111ab3af99a63d5d0f8375d9972e319a379c6bab76" checksum = "0dc92fb57ca44df6db8059111ab3af99a63d5d0f8375d9972e319a379c6bab76"
dependencies = [ dependencies = [
"generic-array 0.14.7", "generic-array 0.14.9",
"rand_core 0.6.4", "rand_core 0.6.4",
"subtle", "subtle",
"zeroize", "zeroize",
@@ -2353,11 +2433,11 @@ dependencies = [
[[package]] [[package]]
name = "crypto-common" name = "crypto-common"
version = "0.1.7" version = "0.1.6"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "78c8292055d1c1df0cce5d180393dc8cce0abec0a7102adb6c7b1eef6016d60a" checksum = "1bfb12502f3fc46cca1bb51ac28df9d618d813cdc3d2f25b9fe775a34af26bb3"
dependencies = [ dependencies = [
"generic-array 0.14.7", "generic-array 0.14.9",
"typenum", "typenum",
] ]
@@ -3554,7 +3634,7 @@ checksum = "9ed9a281f7bc9b7576e61468ba615a66a5c8cfdff42420a70aa82701a3b1e292"
dependencies = [ dependencies = [
"block-buffer 0.10.4", "block-buffer 0.10.4",
"const-oid 0.9.6", "const-oid 0.9.6",
"crypto-common 0.1.7", "crypto-common 0.1.6",
"subtle", "subtle",
] ]
@@ -3814,7 +3894,7 @@ dependencies = [
"crypto-bigint 0.5.5", "crypto-bigint 0.5.5",
"digest 0.10.7", "digest 0.10.7",
"ff 0.13.1", "ff 0.13.1",
"generic-array 0.14.7", "generic-array 0.14.9",
"group 0.13.0", "group 0.13.0",
"hkdf 0.12.4", "hkdf 0.12.4",
"pem-rfc7468 0.7.0", "pem-rfc7468 0.7.0",
@@ -4259,9 +4339,9 @@ dependencies = [
[[package]] [[package]]
name = "generic-array" name = "generic-array"
version = "0.14.7" version = "0.14.9"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "85649ca51fd72272d7821adaf274ad91c288277713d9c18820d8499a7ff69e9a" checksum = "4bb6743198531e02858aeaea5398fcc883e71851fcbcb5a2f773e2fb6cb1edf2"
dependencies = [ dependencies = [
"typenum", "typenum",
"version_check", "version_check",
@@ -4274,7 +4354,7 @@ version = "1.4.4"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "ab4e5aa225bc56696909483320f0ff9b600f1a971b52e07a17d70f3d9b43254b" checksum = "ab4e5aa225bc56696909483320f0ff9b600f1a971b52e07a17d70f3d9b43254b"
dependencies = [ dependencies = [
"generic-array 0.14.7", "generic-array 0.14.9",
"rustversion", "rustversion",
"typenum", "typenum",
] ]
@@ -5301,7 +5381,7 @@ source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "879f10e63c20629ecabbb64a8010319738c66a5cd0c29b02d63d272b03751d01" checksum = "879f10e63c20629ecabbb64a8010319738c66a5cd0c29b02d63d272b03751d01"
dependencies = [ dependencies = [
"block-padding 0.3.3", "block-padding 0.3.3",
"generic-array 0.14.7", "generic-array 0.14.9",
] ]
[[package]] [[package]]
@@ -5835,8 +5915,7 @@ checksum = "b6d2cec3eae94f9f509c767b45932f1ada8350c4bdb85af2fcab4a3c14807981"
[[package]] [[package]]
name = "libmimalloc-sys" name = "libmimalloc-sys"
version = "0.1.49" version = "0.1.49"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "git+https://github.com/xonatius/mimalloc_rust.git?rev=1cdadea43e9c5a0f054b65be21200ce580e4eb13#1cdadea43e9c5a0f054b65be21200ce580e4eb13"
checksum = "6a45a52f43e1c16f667ccfe4dd8c85b7f7c204fd5e3bf46c5b0db9a5c3c0b8e9"
dependencies = [ dependencies = [
"cc", "cc",
"cty", "cty",
@@ -6245,8 +6324,7 @@ dependencies = [
[[package]] [[package]]
name = "mimalloc" name = "mimalloc"
version = "0.1.52" version = "0.1.52"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "git+https://github.com/xonatius/mimalloc_rust.git?rev=1cdadea43e9c5a0f054b65be21200ce580e4eb13#1cdadea43e9c5a0f054b65be21200ce580e4eb13"
checksum = "2d4139bb28d14ad1facf21d5eb8825051b326e172d216b39f6d31df53cc97862"
dependencies = [ dependencies = [
"libmimalloc-sys", "libmimalloc-sys",
] ]
@@ -6668,6 +6746,17 @@ dependencies = [
"num-traits", "num-traits",
] ]
[[package]]
name = "num-rational"
version = "0.4.2"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "f83d14da390562dca69fc84082e73e548e1ad308d24accdedd2720017cb37824"
dependencies = [
"num-bigint",
"num-integer",
"num-traits",
]
[[package]] [[package]]
name = "num-traits" name = "num-traits"
version = "0.2.19" version = "0.2.19"
@@ -6732,7 +6821,7 @@ version = "5.0.0"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "51e219e79014df21a225b1860a479e2dcd7cbd9130f4defd4bd0e191ea31d67d" checksum = "51e219e79014df21a225b1860a479e2dcd7cbd9130f4defd4bd0e191ea31d67d"
dependencies = [ dependencies = [
"base64 0.21.7", "base64 0.22.1",
"chrono", "chrono",
"getrandom 0.2.17", "getrandom 0.2.17",
"http 1.5.0", "http 1.5.0",
@@ -7857,7 +7946,7 @@ source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "be769465445e8c1474e9c5dac2018218498557af32d9ed057325ec9a41ae81bf" checksum = "be769465445e8c1474e9c5dac2018218498557af32d9ed057325ec9a41ae81bf"
dependencies = [ dependencies = [
"heck", "heck",
"itertools 0.13.0", "itertools 0.14.0",
"log", "log",
"multimap", "multimap",
"once_cell", "once_cell",
@@ -7877,7 +7966,7 @@ source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "03da047801ff44bb6a4d407d4860c05fd70bb81714e6b2f3812603d5b145b042" checksum = "03da047801ff44bb6a4d407d4860c05fd70bb81714e6b2f3812603d5b145b042"
dependencies = [ dependencies = [
"heck", "heck",
"itertools 0.13.0", "itertools 0.14.0",
"log", "log",
"multimap", "multimap",
"petgraph 0.8.3", "petgraph 0.8.3",
@@ -7898,7 +7987,7 @@ source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "8a56d757972c98b346a9b766e3f02746cde6dd1cd1d1d563472929fdd74bec4d" checksum = "8a56d757972c98b346a9b766e3f02746cde6dd1cd1d1d563472929fdd74bec4d"
dependencies = [ dependencies = [
"anyhow", "anyhow",
"itertools 0.13.0", "itertools 0.14.0",
"proc-macro2", "proc-macro2",
"quote", "quote",
"syn 2.0.119", "syn 2.0.119",
@@ -7911,7 +8000,7 @@ source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "b570b25f7617e43d59005d0990ccb79e950a423952cea19671b7a876da390adf" checksum = "b570b25f7617e43d59005d0990ccb79e950a423952cea19671b7a876da390adf"
dependencies = [ dependencies = [
"anyhow", "anyhow",
"itertools 0.13.0", "itertools 0.14.0",
"proc-macro2", "proc-macro2",
"quote", "quote",
"syn 2.0.119", "syn 2.0.119",
@@ -8631,6 +8720,15 @@ dependencies = [
"serde", "serde",
] ]
[[package]]
name = "roxmltree"
version = "0.14.1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "921904a62e410e37e215c40381b7117f830d9d89ba60ab5236170541dd25646b"
dependencies = [
"xmlparser",
]
[[package]] [[package]]
name = "rsa" name = "rsa"
version = "0.9.10" version = "0.9.10"
@@ -8724,9 +8822,9 @@ dependencies = [
[[package]] [[package]]
name = "russh" name = "russh"
version = "0.62.4" version = "0.62.5"
source = "registry+https://github.com/rust-lang/crates.io-index" source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "b8b67b5a0d8068c89dcbe9d95df986af7a851d1f3c604525274c37468e60464f" checksum = "da7c230e0ed9cbeb92fbad6c8848985d6df2a1464c0dc247a021abd666e9005e"
dependencies = [ dependencies = [
"aes 0.9.2", "aes 0.9.2",
"aws-lc-rs", "aws-lc-rs",
@@ -9027,6 +9125,7 @@ dependencies = [
"url", "url",
"urlencoding", "urlencoding",
"uuid", "uuid",
"zeroize",
"zip", "zip",
"zstd", "zstd",
] ]
@@ -9511,10 +9610,16 @@ dependencies = [
"arc-swap", "arc-swap",
"argon2", "argon2",
"async-trait", "async-trait",
"aws-config",
"aws-sdk-kms",
"aws-smithy-http-client",
"aws-smithy-runtime-api",
"aws-smithy-types",
"base64 0.23.0", "base64 0.23.0",
"chacha20poly1305", "chacha20poly1305",
"hex", "hex",
"hotpath", "hotpath",
"http 1.5.0",
"insta", "insta",
"jiff", "jiff",
"md-5 0.11.0", "md-5 0.11.0",
@@ -9523,6 +9628,7 @@ dependencies = [
"moka", "moka",
"rand 0.10.2", "rand 0.10.2",
"reqwest", "reqwest",
"rustfs-s3-types",
"rustfs-security-governance", "rustfs-security-governance",
"rustfs-utils", "rustfs-utils",
"rustify", "rustify",
@@ -9708,6 +9814,7 @@ dependencies = [
"hotpath", "hotpath",
"jiff", "jiff",
"libc", "libc",
"log",
"metrics", "metrics",
"num_cpus", "num_cpus",
"nvml-wrapper", "nvml-wrapper",
@@ -9743,6 +9850,7 @@ dependencies = [
"tracing-error", "tracing-error",
"tracing-opentelemetry", "tracing-opentelemetry",
"tracing-subscriber", "tracing-subscriber",
"url",
"zstd", "zstd",
] ]
@@ -9774,6 +9882,7 @@ dependencies = [
"time", "time",
"tokio", "tokio",
"tracing", "tracing",
"tracing-subscriber",
] ]
[[package]] [[package]]
@@ -10027,6 +10136,7 @@ dependencies = [
"rustfs-data-usage", "rustfs-data-usage",
"rustfs-ecstore", "rustfs-ecstore",
"rustfs-filemeta", "rustfs-filemeta",
"rustfs-lock",
"rustfs-storage-api", "rustfs-storage-api",
"rustfs-utils", "rustfs-utils",
"s3s", "s3s",
@@ -10607,7 +10717,7 @@ checksum = "d3e97a565f76233a6003f9f5c54be1d9c5bdfa3eccfb189469f11ec4901c47dc"
dependencies = [ dependencies = [
"base16ct 0.2.0", "base16ct 0.2.0",
"der 0.7.10", "der 0.7.10",
"generic-array 0.14.7", "generic-array 0.14.9",
"pkcs8 0.10.2", "pkcs8 0.10.2",
"subtle", "subtle",
"zeroize", "zeroize",
@@ -10660,6 +10770,12 @@ dependencies = [
"serde_core", "serde_core",
] ]
[[package]]
name = "separator"
version = "0.4.1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "f97841a747eef040fcd2e7b3b9a220a7205926e60488e673d9e4926d27772ce5"
[[package]] [[package]]
name = "seq-macro" name = "seq-macro"
version = "0.3.6" version = "0.3.6"
@@ -11597,7 +11713,7 @@ source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "32497e9a4c7b38532efcdebeef879707aa9f794296a4f0244f6f69e9bc8574bd" checksum = "32497e9a4c7b38532efcdebeef879707aa9f794296a4f0244f6f69e9bc8574bd"
dependencies = [ dependencies = [
"fastrand", "fastrand",
"getrandom 0.4.3", "getrandom 0.3.4",
"once_cell", "once_cell",
"rustix", "rustix",
"windows-sys 0.61.2", "windows-sys 0.61.2",
+6 -4
View File
@@ -173,7 +173,7 @@ tower-http = { version = "0.7.0" }
# Serialization and Data Formats # Serialization and Data Formats
apache-avro = "0.21.0" apache-avro = "0.21.0"
bytes = { version = "1.12.1" } bytes = { version = "1.12.1" }
bytesize = "2.4.2" bytesize = "2.6.0"
byteorder = "1.5.0" byteorder = "1.5.0"
flatbuffers = "25.12.19" flatbuffers = "25.12.19"
form_urlencoded = "1.2.2" form_urlencoded = "1.2.2"
@@ -227,6 +227,7 @@ atoi = "3.1.0"
atomic_enum = "0.3.0" atomic_enum = "0.3.0"
aws-config = { version = "1.10.1" } aws-config = { version = "1.10.1" }
aws-credential-types = { version = "1.3.0" } aws-credential-types = { version = "1.3.0" }
aws-sdk-kms = { default-features = false, version = "1.114.0" }
aws-sdk-s3 = { default-features = false, version = "1.140.0" } aws-sdk-s3 = { default-features = false, version = "1.140.0" }
aws-sdk-sts = { default-features = false, version = "1.110.0" } aws-sdk-sts = { default-features = false, version = "1.110.0" }
aws-smithy-http-client = { default-features = false, version = "1.2.0" } aws-smithy-http-client = { default-features = false, version = "1.2.0" }
@@ -235,7 +236,7 @@ aws-smithy-types = { version = "1.6.1" }
base64 = "0.23.0" base64 = "0.23.0"
base64-simd = "0.8.0" base64-simd = "0.8.0"
brotli = "8.0.4" brotli = "8.0.4"
clap = { version = "4.6.4" } clap = { version = "4.6.5" }
const-str = { version = "1.1.0" } const-str = { version = "1.1.0" }
convert_case = "0.11.0" convert_case = "0.11.0"
criterion = { version = "0.8" } criterion = { version = "0.8" }
@@ -340,14 +341,15 @@ libunftp = { version = "0.23.0" }
unftp-core = "0.1.0" unftp-core = "0.1.0"
suppaftp = { version = "10.0.1" } suppaftp = { version = "10.0.1" }
rcgen = { version = "0.14.8", default-features = false, features = ["aws_lc_rs", "crypto", "pem"] } rcgen = { version = "0.14.8", default-features = false, features = ["aws_lc_rs", "crypto", "pem"] }
russh = { version = "0.62.4" } russh = { version = "0.62.5" }
russh-sftp = "2.3.0" russh-sftp = "2.3.0"
# WebDAV # WebDAV
dav-server = "0.11.0" dav-server = "0.11.0"
# Performance Analysis and Memory Profiling # Performance Analysis and Memory Profiling
mimalloc = "0.1.52" mimalloc = { version = "0.1.52", git = "https://github.com/xonatius/mimalloc_rust.git", rev = "1cdadea43e9c5a0f054b65be21200ce580e4eb13" }
libmimalloc-sys = { version = "0.1.49", git = "https://github.com/xonatius/mimalloc_rust.git", rev = "1cdadea43e9c5a0f054b65be21200ce580e4eb13", features = ["extended"] }
hotpath = { version = "0.22.0", default-features = false } hotpath = { version = "0.22.0", default-features = false }
# Snapshot testing for output format regression detection # Snapshot testing for output format regression detection
insta = { version = "1.48" } insta = { version = "1.48" }
+13
View File
@@ -230,6 +230,19 @@ pub const ENV_RUSTFS_KMS_ENABLE: &str = "RUSTFS_KMS_ENABLE";
/// Default value: false /// Default value: false
pub const DEFAULT_KMS_ENABLE: bool = false; pub const DEFAULT_KMS_ENABLE: bool = false;
/// Environment variable enabling per-key KMS authorization on the SSE-KMS data path.
///
/// When enabled, an SSE-KMS write additionally requires `kms:GenerateDataKey` and an
/// SSE-KMS read additionally requires `kms:Decrypt` on the resolved key, evaluated as
/// the requesting identity. SSE-S3 and SSE-C are unaffected.
pub const ENV_RUSTFS_KMS_ENFORCE_SSE_KEY_POLICY: &str = "RUSTFS_KMS_ENFORCE_SSE_KEY_POLICY";
/// Default per-key KMS authorization mode for the SSE-KMS data path.
///
/// Off for now so deployments whose identity policies only grant s3 actions keep
/// working; the default flips to on in a later release.
pub const DEFAULT_KMS_ENFORCE_SSE_KEY_POLICY: bool = false;
/// Environment variable for server KMS backend. /// Environment variable for server KMS backend.
pub const ENV_RUSTFS_KMS_BACKEND: &str = "RUSTFS_KMS_BACKEND"; pub const ENV_RUSTFS_KMS_BACKEND: &str = "RUSTFS_KMS_BACKEND";
+1 -60
View File
@@ -26,75 +26,16 @@
//! Later batches tracked on backlog#1154: config get/set, info, pools status, //! Later batches tracked on backlog#1154: config get/set, info, pools status,
//! group lifecycle, import/export IAM. //! group lifecycle, import/export IAM.
use crate::common::{RustFSTestEnvironment, init_logging, local_http_client}; use crate::common::{RustFSTestEnvironment, admin_ok, admin_request, init_logging};
use aws_sdk_s3::config::{Credentials, Region}; use aws_sdk_s3::config::{Credentials, Region};
use aws_sdk_s3::primitives::ByteStream; use aws_sdk_s3::primitives::ByteStream;
use aws_sdk_s3::{Client, Config}; use aws_sdk_s3::{Client, Config};
use http::header::{CONTENT_TYPE, HOST};
use reqwest::StatusCode; use reqwest::StatusCode;
use rustfs_signer::constants::UNSIGNED_PAYLOAD;
use rustfs_signer::sign_v4;
use s3s::Body;
use serial_test::serial; use serial_test::serial;
use std::error::Error; use std::error::Error;
use tokio::time::{Duration, sleep}; use tokio::time::{Duration, sleep};
type TestResult = Result<(), Box<dyn Error + Send + Sync>>; type TestResult = Result<(), Box<dyn Error + Send + Sync>>;
type BoxError = Box<dyn Error + Send + Sync>;
/// Signs and sends an admin HTTP request with the given credential, returning
/// status and body. Native `/rustfs/admin/v3` requests and responses are plain
/// JSON (the MinIO-compat encryption applies only to `/minio/admin/v3` paths).
async fn admin_request(
base_url: &str,
method: http::Method,
path_and_query: &str,
body: Option<String>,
access_key: &str,
secret_key: &str,
) -> Result<(StatusCode, String), BoxError> {
let url = format!("{base_url}{path_and_query}");
let uri = url.parse::<http::Uri>()?;
let authority = uri.authority().ok_or("admin URL missing authority")?.to_string();
let mut builder = http::Request::builder()
.method(method.clone())
.uri(uri)
.header(HOST, authority)
.header("x-amz-content-sha256", UNSIGNED_PAYLOAD);
if body.is_some() {
builder = builder.header(CONTENT_TYPE, "application/json");
}
let content_len = body.as_ref().map(|b| b.len() as i64).unwrap_or_default();
let signed = sign_v4(builder.body(Body::empty())?, content_len, access_key, secret_key, "", "us-east-1");
let reqwest_method = reqwest::Method::from_bytes(method.as_str().as_bytes())?;
let mut request = local_http_client().request(reqwest_method, &url);
for (name, value) in signed.headers() {
request = request.header(name, value);
}
if let Some(body) = body {
request = request.body(body);
}
let response = request.send().await?;
let status = response.status();
let text = response.text().await.unwrap_or_default();
Ok((status, text))
}
/// Root-credential admin request that must succeed; returns the response body.
async fn admin_ok(
env: &RustFSTestEnvironment,
method: http::Method,
path_and_query: &str,
body: Option<String>,
) -> Result<String, BoxError> {
let (status, text) = admin_request(&env.url, method.clone(), path_and_query, body, &env.access_key, &env.secret_key).await?;
if !status.is_success() {
return Err(format!("{method} {path_and_query} failed: {status} {text}").into());
}
Ok(text)
}
fn build_s3_client(url: &str, access_key: &str, secret_key: &str) -> Client { fn build_s3_client(url: &str, access_key: &str, secret_key: &str) -> Client {
let config = Config::builder() let config = Config::builder()
+57
View File
@@ -24,7 +24,12 @@
use aws_sdk_s3::config::{Credentials, Region}; use aws_sdk_s3::config::{Credentials, Region};
use aws_sdk_s3::{Client, Config}; use aws_sdk_s3::{Client, Config};
use aws_smithy_http_client::Builder as SmithyHttpClientBuilder; use aws_smithy_http_client::Builder as SmithyHttpClientBuilder;
use http::header::{CONTENT_TYPE, HOST};
use reqwest::Client as HttpClient; use reqwest::Client as HttpClient;
use reqwest::StatusCode;
use rustfs_signer::constants::UNSIGNED_PAYLOAD;
use rustfs_signer::sign_v4;
use s3s::Body;
use std::ffi::OsStr; use std::ffi::OsStr;
use std::fs as stdfs; use std::fs as stdfs;
use std::path::{Path, PathBuf}; use std::path::{Path, PathBuf};
@@ -75,6 +80,58 @@ pub fn local_http_client() -> HttpClient {
.expect("failed to build local reqwest client") .expect("failed to build local reqwest client")
} }
/// Signs and sends an admin HTTP request with the given credentials.
pub(crate) async fn admin_request(
base_url: &str,
method: http::Method,
path_and_query: &str,
body: Option<String>,
access_key: &str,
secret_key: &str,
) -> Result<(StatusCode, String), Box<dyn std::error::Error + Send + Sync>> {
let url = format!("{base_url}{path_and_query}");
let uri = url.parse::<http::Uri>()?;
let authority = uri.authority().ok_or("admin URL missing authority")?.to_string();
let mut request = http::Request::builder()
.method(method.clone())
.uri(uri)
.header(HOST, authority)
.header("x-amz-content-sha256", UNSIGNED_PAYLOAD);
if body.is_some() {
request = request.header(CONTENT_TYPE, "application/json");
}
let content_length = i64::try_from(body.as_ref().map_or(0, String::len)).map_err(|_| "admin request body is too large")?;
let signed = sign_v4(request.body(Body::empty())?, content_length, access_key, secret_key, "", "us-east-1");
let mut request = local_http_client().request(method, &url);
for (name, value) in signed.headers() {
request = request.header(name, value);
}
if let Some(body) = body {
request = request.body(body);
}
let response = request.send().await?;
let status = response.status();
let body = response.text().await?;
Ok((status, body))
}
/// Sends a root-credential admin request and returns its successful response body.
pub(crate) async fn admin_ok(
env: &RustFSTestEnvironment,
method: http::Method,
path_and_query: &str,
body: Option<String>,
) -> Result<String, Box<dyn std::error::Error + Send + Sync>> {
let (status, response_body) =
admin_request(&env.url, method.clone(), path_and_query, body, &env.access_key, &env.secret_key).await?;
if !status.is_success() {
return Err(format!("{method} {path_and_query} failed: {status} {response_body}").into());
}
Ok(response_body)
}
/// Resolve the RustFS binary relative to the workspace. /// Resolve the RustFS binary relative to the workspace.
pub fn rustfs_binary_path() -> PathBuf { pub fn rustfs_binary_path() -> PathBuf {
rustfs_binary_path_with_features(requested_rustfs_build_features().as_deref()) rustfs_binary_path_with_features(requested_rustfs_build_features().as_deref())
@@ -117,9 +117,14 @@ mod tests {
.key("assets/explicit-copy.js") .key("assets/explicit-copy.js")
.copy_source(format!("{bucket}/{key}")) .copy_source(format!("{bucket}/{key}"))
.metadata_directive(MetadataDirective::Copy) .metadata_directive(MetadataDirective::Copy)
.customize()
.mutate_request(|request| {
request.headers_mut().insert("content-type", "application/octet-stream");
request.headers_mut().insert("x-amz-meta-request-only", "ignored");
})
.send() .send()
.await .await
.expect("explicit COPY directive failed"); .expect("explicit COPY directive with request metadata failed");
let explicit_copy_head = client let explicit_copy_head = client
.head_object() .head_object()
.bucket(bucket) .bucket(bucket)
@@ -128,6 +133,18 @@ mod tests {
.await .await
.expect("HEAD failed after explicit COPY"); .expect("HEAD failed after explicit COPY");
assert_eq!(explicit_copy_head.cache_control(), Some("max-age=60")); assert_eq!(explicit_copy_head.cache_control(), Some("max-age=60"));
assert_eq!(explicit_copy_head.content_type(), Some("text/javascript; charset=utf-8"));
assert_eq!(
explicit_copy_head.metadata().and_then(|metadata| metadata.get("mtime")),
Some(&"1777992333".to_string())
);
assert_eq!(
explicit_copy_head
.metadata()
.and_then(|metadata| metadata.get("request-only")),
None,
"COPY must ignore request metadata"
);
assert_eq!( assert_eq!(
explicit_copy_head.website_redirect_location(), explicit_copy_head.website_redirect_location(),
None, None,
@@ -571,20 +588,6 @@ mod tests {
Some("InvalidArgument") Some("InvalidArgument")
); );
let ignored_replacement = client
.copy_object()
.bucket(bucket)
.key(key)
.copy_source(format!("{bucket}/{key}"))
.content_type("application/ignored")
.send()
.await
.expect_err("Replacement fields without REPLACE should be rejected");
assert_eq!(
ignored_replacement.as_service_error().and_then(|error| error.code()),
Some("InvalidRequest")
);
let unchanged = client let unchanged = client
.get_object() .get_object()
.bucket(bucket) .bucket(bucket)
@@ -56,6 +56,21 @@ mod tests {
); );
} }
async fn assert_current_list_hides_delete_marker(client: &Client, bucket: &str, key: &str) {
let listed = client
.list_objects_v2()
.bucket(bucket)
.prefix(key)
.send()
.await
.expect("list current objects after delete marker");
assert!(
listed.contents().iter().all(|object| object.key() != Some(key)),
"ListObjectsV2 must hide an object whose latest version is a delete marker"
);
}
#[tokio::test] #[tokio::test]
#[serial] #[serial]
async fn test_versioning_only_delete_marker_has_minio_compatible_visibility_for_migration_proof() { async fn test_versioning_only_delete_marker_has_minio_compatible_visibility_for_migration_proof() {
@@ -94,6 +109,7 @@ mod tests {
assert_eq!(markers[0].version_id(), Some(delete_marker_version_id)); assert_eq!(markers[0].version_id(), Some(delete_marker_version_id));
assert_eq!(markers[0].is_latest(), Some(true)); assert_eq!(markers[0].is_latest(), Some(true));
assert_current_get_is_delete_marker_not_found(&client, bucket, key).await; assert_current_get_is_delete_marker_not_found(&client, bucket, key).await;
assert_current_list_hides_delete_marker(&client, bucket, key).await;
} }
#[tokio::test] #[tokio::test]
@@ -118,6 +134,17 @@ mod tests {
.await .await
.expect("put historical version"); .expect("put historical version");
let data_version_id = put.version_id().expect("put should return data version id"); let data_version_id = put.version_id().expect("put should return data version id");
let listed_before_delete = client
.list_objects_v2()
.bucket(bucket)
.prefix(key)
.send()
.await
.expect("list current object before creating delete marker");
assert!(
listed_before_delete.contents().iter().any(|object| object.key() == Some(key)),
"ListObjectsV2 must include the current object before it is deleted"
);
let delete_marker = client let delete_marker = client
.delete_object() .delete_object()
@@ -145,6 +172,7 @@ mod tests {
assert_eq!(markers[0].version_id(), Some(delete_marker_version_id)); assert_eq!(markers[0].version_id(), Some(delete_marker_version_id));
assert_eq!(markers[0].is_latest(), Some(true)); assert_eq!(markers[0].is_latest(), Some(true));
assert_current_get_is_delete_marker_not_found(&client, bucket, key).await; assert_current_get_is_delete_marker_not_found(&client, bucket, key).await;
assert_current_list_hides_delete_marker(&client, bucket, key).await;
let historical = client let historical = client
.get_object() .get_object()
@@ -126,7 +126,10 @@ async fn assert_key_deletion_lifecycle(base_url: &str, access_key: &str, secret_
assert_eq!(cancelled["success"], true); assert_eq!(cancelled["success"], true);
assert_eq!(cancelled["key_metadata"]["key_state"], "Enabled"); assert_eq!(cancelled["key_metadata"]["key_state"], "Enabled");
let removed = kms_admin_request( // A default server refuses to skip the waiting window (rustfs/backlog#1585):
// immediate deletion is unrecoverable and takes every object encrypted under
// the key with it, so the endpoint must reject it rather than honour it.
let refused = kms_admin_request(
base_url, base_url,
http::Method::DELETE, http::Method::DELETE,
"/rustfs/admin/v3/kms/keys/delete", "/rustfs/admin/v3/kms/keys/delete",
@@ -140,27 +143,39 @@ async fn assert_key_deletion_lifecycle(base_url: &str, access_key: &str, secret_
access_key, access_key,
secret_key, secret_key,
) )
.await?; .await
let removed: serde_json::Value = serde_json::from_str(&removed)?; .err()
assert_eq!(removed["success"], true); .ok_or("immediate KMS key deletion must be refused on a default server")?;
assert!(
refused.to_string().contains("400 Bad Request"),
"refused immediate deletion must report a client error: {refused}"
);
let listed = // The refused request left the key alone, so the window-bounded path still
kms_admin_request(base_url, http::Method::GET, "/rustfs/admin/v3/kms/keys", None, access_key, secret_key).await?; // has something to schedule.
let listed: serde_json::Value = serde_json::from_str(&listed)?; let described = kms_admin_request(
assert_eq!(listed["success"], true); base_url,
let keys = listed["keys"] http::Method::GET,
.as_array() &format!("/rustfs/admin/v3/kms/keys/{key_id}"),
.ok_or("list KMS keys response omitted keys after deletion")?; None,
if let Some(key) = keys.iter().find(|key| key["key_id"] == key_id) { access_key,
assert_eq!(key["status"], "PendingDeletion", "a retained force-deleted key must be pending deletion"); secret_key,
let removed = kms_admin_request( )
.await?;
let described: serde_json::Value = serde_json::from_str(&described)?;
assert_eq!(
described["key_metadata"]["key_state"], "Enabled",
"a refused immediate deletion must leave the key usable"
);
let rescheduled = kms_admin_request(
base_url, base_url,
http::Method::DELETE, http::Method::DELETE,
"/rustfs/admin/v3/kms/keys/delete", "/rustfs/admin/v3/kms/keys/delete",
Some( Some(
&serde_json::json!({ &serde_json::json!({
"key_id": key_id, "key_id": key_id,
"force_immediate": true "pending_window_in_days": 7
}) })
.to_string(), .to_string(),
), ),
@@ -168,9 +183,9 @@ async fn assert_key_deletion_lifecycle(base_url: &str, access_key: &str, secret_
secret_key, secret_key,
) )
.await?; .await?;
let removed: serde_json::Value = serde_json::from_str(&removed)?; let rescheduled: serde_json::Value = serde_json::from_str(&rescheduled)?;
assert_eq!(removed["success"], true); assert_eq!(rescheduled["success"], true);
} assert!(rescheduled["deletion_date"].is_string());
let listed = let listed =
kms_admin_request(base_url, http::Method::GET, "/rustfs/admin/v3/kms/keys", None, access_key, secret_key).await?; kms_admin_request(base_url, http::Method::GET, "/rustfs/admin/v3/kms/keys", None, access_key, secret_key).await?;
@@ -179,10 +194,11 @@ async fn assert_key_deletion_lifecycle(base_url: &str, access_key: &str, secret_
let keys = listed["keys"] let keys = listed["keys"]
.as_array() .as_array()
.ok_or("final list KMS keys response omitted keys after deletion")?; .ok_or("final list KMS keys response omitted keys after deletion")?;
assert!( let key = keys
keys.iter().all(|key| key["key_id"] != key_id), .iter()
"force-deleted KMS key must no longer appear in list" .find(|key| key["key_id"] == key_id)
); .ok_or("a key awaiting its deletion window must still be listed")?;
assert_eq!(key["status"], "PendingDeletion", "a scheduled key must be pending deletion");
Ok(()) Ok(())
} }
+17 -15
View File
@@ -449,29 +449,31 @@ async fn test_vault_kms_key_crud(
info!("✅ Delete verification: Key state correctly changed to: {}", key_state); info!("✅ Delete verification: Key state correctly changed to: {}", key_state);
// Force Delete - Force immediate deletion for PendingDeletion key // Force Delete - a default server refuses to skip the waiting window
let force_delete_response = crate::common::execute_awscurl( // (rustfs/backlog#1585): destroying the key material immediately would take
// every object encrypted under the key with it.
let force_delete_error = crate::common::execute_awscurl(
&format!("{base_url}/rustfs/admin/v3/kms/keys/delete?keyId={key_id}&force_immediate=true"), &format!("{base_url}/rustfs/admin/v3/kms/keys/delete?keyId={key_id}&force_immediate=true"),
"DELETE", "DELETE",
None, None,
access_key, access_key,
secret_key, secret_key,
) )
.await?; .await
.expect_err("Immediate KMS key deletion must be refused on a default server");
info!("✅ Force Delete: correctly refused for key {}: {}", key_id, force_delete_error);
// Parse and validate the force delete response // The refused request must leave the key exactly as it was: still present,
let force_delete_result: serde_json::Value = serde_json::from_str(&force_delete_response)?; // still pending deletion, still recoverable through cancel-deletion.
assert_eq!(force_delete_result["success"], true, "Force delete operation must return success=true"); let describe_after_refusal =
info!("✅ Force Delete: Successfully force deleted key: {}", key_id); crate::common::awscurl_get(&format!("{base_url}/rustfs/admin/v3/kms/keys/{key_id}"), access_key, secret_key).await?;
let describe_after_refusal: serde_json::Value = serde_json::from_str(&describe_after_refusal)?;
assert_eq!(
describe_after_refusal["key_metadata"]["key_state"], "PendingDeletion",
"A refused immediate deletion must leave the key pending deletion"
);
// Verify key no longer exists after force deletion (should return error) info!("✅ Force Delete verification: Key survived the refused immediate deletion");
let describe_force_deleted_result =
crate::common::awscurl_get(&format!("{base_url}/rustfs/admin/v3/kms/keys/{key_id}"), access_key, secret_key).await;
// After force deletion, key should not be found (GET should fail)
assert!(describe_force_deleted_result.is_err(), "Force deleted key should not be found");
info!("✅ Force Delete verification: Key was permanently deleted and is no longer accessible");
info!("Vault KMS key CRUD operations completed successfully"); info!("Vault KMS key CRUD operations completed successfully");
Ok(()) Ok(())
@@ -26,6 +26,7 @@
use super::common::*; use super::common::*;
use aws_sdk_s3::Client; use aws_sdk_s3::Client;
use aws_sdk_s3::error::ProvideErrorMetadata;
use aws_sdk_s3::primitives::{ByteStream, DateTimeFormat}; use aws_sdk_s3::primitives::{ByteStream, DateTimeFormat};
use aws_sdk_s3::types::{ use aws_sdk_s3::types::{
CompletedMultipartUpload, CompletedPart, Delete, MetadataDirective, ObjectIdentifier, ObjectLockLegalHoldStatus, CompletedMultipartUpload, CompletedPart, Delete, MetadataDirective, ObjectIdentifier, ObjectLockLegalHoldStatus,
@@ -2120,6 +2121,127 @@ async fn test_multipart_default_retention_fixed_at_create() {
// Versioning Auto-Enable Tests // Versioning Auto-Enable Tests
// ============================================================================ // ============================================================================
#[tokio::test]
#[serial]
async fn test_unretained_object_lock_object_delete_and_bucket_cleanup() {
init_logging();
info!("🧪 Test: Unretained Object Lock object delete and bucket cleanup (Issue #5339)");
let mut env = ObjectLockTestEnvironment::new()
.await
.expect("failed to create Object Lock test environment");
env.start_rustfs().await.expect("failed to start RustFS");
let bucket = "test-object-lock-delete-cleanup";
let key = "unretained-object";
env.create_object_lock_bucket(bucket)
.await
.expect("failed to create Object Lock bucket");
let client = env.s3_client();
let put_response = client
.put_object()
.bucket(bucket)
.key(key)
.body(ByteStream::from_static(b"unretained data"))
.send()
.await
.expect("failed to upload unretained object");
let object_version_id = put_response
.version_id()
.expect("Object Lock buckets must create versioned objects")
.to_string();
let delete_response = client
.delete_object()
.bucket(bucket)
.key(key)
.send()
.await
.expect("failed to create delete marker");
assert_eq!(delete_response.delete_marker(), Some(true));
let delete_marker_version_id = delete_response
.version_id()
.expect("Deleting without a version ID must create a delete marker")
.to_string();
let get_error = client
.get_object()
.bucket(bucket)
.key(key)
.send()
.await
.expect_err("GET must not return an object hidden by a delete marker");
assert_eq!(get_error.raw_response().map(|response| response.status().as_u16()), Some(404));
assert_eq!(get_error.as_service_error().and_then(|error| error.code()), Some("NoSuchKey"));
let listed_objects = client
.list_objects_v2()
.bucket(bucket)
.send()
.await
.expect("failed to list current objects");
assert!(
listed_objects.contents().iter().all(|object| object.key() != Some(key)),
"ListObjectsV2 must hide objects whose latest version is a delete marker"
);
let listed_versions = client
.list_object_versions()
.bucket(bucket)
.send()
.await
.expect("failed to list object versions");
assert!(
listed_versions
.versions()
.iter()
.any(|version| version.key() == Some(key) && version.version_id() == Some(object_version_id.as_str())),
"The data version must remain until it is explicitly deleted"
);
assert!(
listed_versions
.delete_markers()
.iter()
.any(|marker| marker.key() == Some(key) && marker.version_id() == Some(delete_marker_version_id.as_str())),
"ListObjectVersions must expose the delete marker"
);
client
.delete_object()
.bucket(bucket)
.key(key)
.version_id(object_version_id)
.send()
.await
.expect("failed to delete the data version");
client
.delete_object()
.bucket(bucket)
.key(key)
.version_id(delete_marker_version_id)
.send()
.await
.expect("failed to delete the delete marker");
let remaining_versions = client
.list_object_versions()
.bucket(bucket)
.send()
.await
.expect("failed to list versions after cleanup");
assert!(remaining_versions.versions().is_empty());
assert!(remaining_versions.delete_markers().is_empty());
client
.delete_bucket()
.bucket(bucket)
.send()
.await
.expect("Deleting every version must remove xl.meta so the bucket can be deleted normally");
}
#[tokio::test] #[tokio::test]
#[serial] #[serial]
async fn test_versioning_auto_enabled_with_object_lock() { async fn test_versioning_auto_enabled_with_object_lock() {
+6 -4
View File
@@ -15,9 +15,7 @@
use super::{grpc_lock_client::GrpcLockClient, grpc_lock_server::spawn_lock_server}; use super::{grpc_lock_client::GrpcLockClient, grpc_lock_server::spawn_lock_server};
use rustfs_lock::client::{LockClient, local::LocalClient}; use rustfs_lock::client::{LockClient, local::LocalClient};
use rustfs_lock::{ use rustfs_lock::{GlobalLockManager, LockInfo, LockRequest, LockResponse, LockStats, LockType, NamespaceLock, ObjectKey};
GlobalLockManager, LockError, LockInfo, LockRequest, LockResponse, LockStats, LockType, NamespaceLock, ObjectKey,
};
use std::sync::Arc; use std::sync::Arc;
use std::time::Duration; use std::time::Duration;
@@ -35,7 +33,11 @@ struct FailingClient;
#[async_trait::async_trait] #[async_trait::async_trait]
impl rustfs_lock::LockClient for FailingClient { impl rustfs_lock::LockClient for FailingClient {
async fn acquire_lock(&self, _request: &rustfs_lock::LockRequest) -> rustfs_lock::Result<LockResponse> { async fn acquire_lock(&self, _request: &rustfs_lock::LockRequest) -> rustfs_lock::Result<LockResponse> {
Err(LockError::internal("simulated gRPC node failure")) // Match RemoteClient's transport-failure response so the coordinator can count this node toward quorum loss.
Ok(LockResponse::failure(
"Remote lock RPC failed: simulated gRPC node failure",
Duration::ZERO,
))
} }
async fn release(&self, _lock_id: &rustfs_lock::LockId) -> rustfs_lock::Result<bool> { async fn release(&self, _lock_id: &rustfs_lock::LockId) -> rustfs_lock::Result<bool> {
+17 -45
View File
@@ -382,7 +382,6 @@ struct ManualTransitionRunReport {
skipped_delete_marker: u64, skipped_delete_marker: u64,
skipped_directory: u64, skipped_directory: u64,
skipped_replication: u64, skipped_replication: u64,
skipped_already_transitioned: u64,
skipped_already_in_flight: u64, skipped_already_in_flight: u64,
skipped_queue_full: u64, skipped_queue_full: u64,
skipped_queue_closed: u64, skipped_queue_closed: u64,
@@ -408,41 +407,6 @@ fn assert_completed_or_in_flight_partial(state: &str, report: &ManualTransitionR
} }
} }
fn assert_conflict_winner_report(state: &str, report: &ManualTransitionRunReport, expected_objects: u64, context: &str) {
assert_completed_or_in_flight_partial(state, report, context);
if report.skipped_already_in_flight > 0 {
assert!(
report.scanned <= expected_objects,
"{context}: scanned more objects than the conflict scope contains: {report:#?}"
);
assert!(
report.eligible <= expected_objects,
"{context}: marked more objects eligible than the conflict scope contains: {report:#?}"
);
assert_eq!(
report.enqueued + report.skipped_already_in_flight,
report.eligible,
"{context}: partial in-flight accounting must cover every eligible object: {report:#?}"
);
} else {
assert_eq!(report.scanned, expected_objects, "{context}: {report:#?}");
assert_eq!(
report.eligible + report.skipped_already_transitioned,
expected_objects,
"{context}: {report:#?}"
);
assert_eq!(
report.enqueued + report.skipped_already_in_flight,
expected_objects,
"{context}: {report:#?}"
);
}
assert_eq!(
report.transition_completed, report.enqueued,
"{context}: winner must wait for all queued transitions: {report:#?}"
);
}
#[derive(Debug, Deserialize)] #[derive(Debug, Deserialize)]
struct ManualTransitionQueueSnapshot { struct ManualTransitionQueueSnapshot {
queue_capacity: u64, queue_capacity: u64,
@@ -1276,7 +1240,14 @@ async fn test_manual_transition_async_scope_conflicts_report_active_job() -> Tes
cold_client.create_bucket().bucket(TIER_BUCKET).send().await?; cold_client.create_bucket().bucket(TIER_BUCKET).send().await?;
let mut hot = RustFSTestEnvironment::new().await?; let mut hot = RustFSTestEnvironment::new().await?;
hot.start_rustfs_server_with_env(vec![], &[("RUSTFS_SCANNER_ENABLED", "false"), ("RUSTFS_SCANNER_CYCLE", "3600")]) hot.start_rustfs_server_with_env(
vec![],
&[
("RUSTFS_SCANNER_ENABLED", "false"),
("RUSTFS_SCANNER_CYCLE", "3600"),
(MANUAL_TRANSITION_CANCEL_BARRIER_ENV, "1"),
],
)
.await?; .await?;
let hot_client = hot.create_s3_client(); let hot_client = hot.create_s3_client();
add_rustfs_tier(&hot, &cold).await?; add_rustfs_tier(&hot, &cold).await?;
@@ -1347,20 +1318,21 @@ async fn test_manual_transition_async_scope_conflicts_report_active_job() -> Tes
assert_eq!(conflict.cancel_endpoint, status_endpoint); assert_eq!(conflict.cancel_endpoint, status_endpoint);
assert!(!conflict.scope_key.is_empty()); assert!(!conflict.scope_key.is_empty());
manual_transition_job_cancel(&hot, cancel_endpoint).await?;
let terminal = wait_for_manual_transition_job_terminal(&hot, status_endpoint, MANUAL_ASYNC_CONFLICT_TERMINAL_TIMEOUT).await?; let terminal = wait_for_manual_transition_job_terminal(&hot, status_endpoint, MANUAL_ASYNC_CONFLICT_TERMINAL_TIMEOUT).await?;
assert_eq!(terminal.job_id, job_id); assert_eq!(terminal.job_id, job_id);
assert_eq!(terminal.status, "cancelled", "terminal conflict winner response: {terminal:#?}");
assert!(!terminal.report.dry_run); assert!(!terminal.report.dry_run);
assert_eq!(terminal.report.bucket, MANUAL_ASYNC_CONFLICT_BUCKET); assert_eq!(terminal.report.bucket, MANUAL_ASYNC_CONFLICT_BUCKET);
assert_eq!(terminal.report.prefix, accepted.report.prefix); assert_eq!(terminal.report.prefix, accepted.report.prefix);
assert_conflict_winner_report( assert!(terminal.report.cancelled, "terminal conflict winner response: {terminal:#?}");
&terminal.status, assert_eq!(terminal.report.scanned, 0, "terminal conflict winner response: {terminal:#?}");
&terminal.report, assert_eq!(terminal.report.enqueued, 0, "terminal conflict winner response: {terminal:#?}");
MANUAL_ASYNC_CONFLICT_OBJECTS as u64, assert_eq!(
"terminal conflict winner response", terminal.report.transition_completed, 0,
"terminal conflict winner response: {terminal:#?}"
); );
assert_eq!(terminal.report.dry_run_eligible, 0, "terminal conflict winner response: {terminal:#?}");
assert_eq!(terminal.report.transition_failed, 0, "terminal conflict winner response: {terminal:#?}");
assert_eq!(terminal.report.tier_failure, 0, "terminal conflict winner response: {terminal:#?}");
let after_remote_count = cold_tier_object_count(&cold_client).await?; let after_remote_count = cold_tier_object_count(&cold_client).await?;
assert!(after_remote_count >= before_remote_count); assert!(after_remote_count >= before_remote_count);
assert!(after_remote_count <= before_remote_count + MANUAL_ASYNC_CONFLICT_OBJECTS); assert!(after_remote_count <= before_remote_count + MANUAL_ASYNC_CONFLICT_OBJECTS);
@@ -29,8 +29,9 @@ use aws_sdk_s3::types::{
use aws_sdk_s3::{Client, Config}; use aws_sdk_s3::{Client, Config};
use base64::{Engine, engine::general_purpose::STANDARD as BASE64_STANDARD}; use base64::{Engine, engine::general_purpose::STANDARD as BASE64_STANDARD};
use bytes::Bytes; use bytes::Bytes;
use flate2::read::GzDecoder;
use futures::{Stream, StreamExt}; use futures::{Stream, StreamExt};
use http::header::{CONTENT_TYPE, HOST}; use http::header::{CONTENT_ENCODING, CONTENT_TYPE, HOST};
use http_body_util::{BodyExt, Full}; use http_body_util::{BodyExt, Full};
use hyper::body::Incoming; use hyper::body::Incoming;
use hyper::server::conn::http1; use hyper::server::conn::http1;
@@ -38,6 +39,10 @@ use hyper::service::service_fn;
use hyper::{Request, Response}; use hyper::{Request, Response};
use hyper_util::rt::TokioIo; use hyper_util::rt::TokioIo;
use local_ip_address::local_ip; use local_ip_address::local_ip;
use opentelemetry_proto::tonic::collector::metrics::v1::ExportMetricsServiceRequest;
use opentelemetry_proto::tonic::common::v1::{KeyValue, any_value::Value as AnyValue};
use opentelemetry_proto::tonic::metrics::v1::{Metric, metric, number_data_point};
use prost::Message;
use rcgen::{ use rcgen::{
BasicConstraints, CertificateParams, CertifiedIssuer, DnType, ExtendedKeyUsagePurpose, IsCa, KeyPair, KeyUsagePurpose, BasicConstraints, CertificateParams, CertifiedIssuer, DnType, ExtendedKeyUsagePurpose, IsCa, KeyPair, KeyUsagePurpose,
SanType, generate_simple_self_signed, SanType, generate_simple_self_signed,
@@ -56,6 +61,7 @@ use sha2::{Digest, Sha256};
use std::collections::BTreeMap; use std::collections::BTreeMap;
use std::convert::Infallible; use std::convert::Infallible;
use std::error::Error; use std::error::Error;
use std::io::Read;
use std::net::IpAddr; use std::net::IpAddr;
use std::path::Path; use std::path::Path;
use std::process::Command; use std::process::Command;
@@ -64,11 +70,13 @@ use std::sync::atomic::{AtomicU64, Ordering};
use time::{Duration as TimeDuration, OffsetDateTime}; use time::{Duration as TimeDuration, OffsetDateTime};
use tokio::fs; use tokio::fs;
use tokio::net::TcpListener; use tokio::net::TcpListener;
use tokio::sync::watch; use tokio::sync::{Mutex, watch};
use tokio::task::JoinHandle;
use tokio::task::JoinSet; use tokio::task::JoinSet;
use tokio::time::{Duration, sleep, timeout}; use tokio::time::{Duration, sleep, timeout};
type TestResult = Result<(), Box<dyn Error + Send + Sync>>; type TestResult = Result<(), Box<dyn Error + Send + Sync>>;
type BacklogMetricPoints = Arc<Mutex<BTreeMap<String, BTreeMap<String, (u64, f64)>>>>;
/// A replication source server validates the remote target endpoint, and the e2e /// A replication source server validates the remote target endpoint, and the e2e
/// target runs on loopback (127.0.0.1), which RustFS's SSRF egress guard rejects by /// target runs on loopback (127.0.0.1), which RustFS's SSRF egress guard rejects by
@@ -107,6 +115,252 @@ const REPL17_KMS_KEY_ID: &str = "repl17-local-key";
const REPL17_SSEC_KEY: &str = "01234567890123456789012345678901"; const REPL17_SSEC_KEY: &str = "01234567890123456789012345678901";
const REPLICATION_FAILED_EVENT: &str = "s3:Replication:OperationFailedReplication"; const REPLICATION_FAILED_EVENT: &str = "s3:Replication:OperationFailedReplication";
const REPLICATION_EVENT_MAX_BUFFER_BYTES: usize = 1024 * 1024; const REPLICATION_EVENT_MAX_BUFFER_BYTES: usize = 1024 * 1024;
const OTLP_METRICS_BODY_LIMIT: u64 = 4 * 1024 * 1024;
const BUCKET_LABEL: &str = "bucket";
const TOTAL_FAILED_COUNT_METRIC: &str = "rustfs_bucket_replication_total_failed_count";
const CURRENT_BACKLOG_COUNT_METRIC: &str = "rustfs_bucket_replication_current_backlog_count";
const CURRENT_BACKLOG_BYTES_METRIC: &str = "rustfs_bucket_replication_current_backlog_bytes";
const MRF_PENDING_COUNT_METRIC: &str = "rustfs_bucket_replication_mrf_pending_count";
const MRF_PENDING_BYTES_METRIC: &str = "rustfs_bucket_replication_mrf_pending_bytes";
struct ReplicationBacklogMetricCollector {
endpoint: String,
values: BacklogMetricPoints,
task: JoinHandle<()>,
}
impl ReplicationBacklogMetricCollector {
async fn start() -> Result<Self, Box<dyn Error + Send + Sync>> {
let listener = TcpListener::bind("127.0.0.1:0").await?;
let endpoint = format!("http://{}/v1/metrics", listener.local_addr()?);
let values = Arc::new(Mutex::new(BTreeMap::new()));
let task_values = values.clone();
let task = tokio::spawn(async move {
loop {
let Ok((stream, _)) = listener.accept().await else {
break;
};
let values = task_values.clone();
tokio::spawn(async move {
let _ = http1::Builder::new()
.serve_connection(
TokioIo::new(stream),
service_fn(move |request| handle_backlog_metric_export(request, values.clone())),
)
.await;
});
}
});
Ok(Self { endpoint, values, task })
}
fn root_endpoint(&self) -> &str {
self.endpoint.trim_end_matches("/v1/metrics")
}
async fn bucket_metric_value(&self, metric: &str, bucket: &str) -> f64 {
self.values
.lock()
.await
.get(metric)
.and_then(|buckets| buckets.get(bucket))
.map(|(_, value)| *value)
.unwrap_or_default()
}
async fn wait_for_bucket_metric(
&self,
metric: &str,
bucket: &str,
expected: impl Fn(f64) -> bool,
description: &str,
) -> Result<f64, Box<dyn Error + Send + Sync>> {
let deadline = tokio::time::Instant::now() + Duration::from_secs(30);
loop {
let value = self.bucket_metric_value(metric, bucket).await;
if expected(value) {
return Ok(value);
}
if tokio::time::Instant::now() >= deadline {
let snapshot = self.values.lock().await.clone();
return Err(format!("timed out waiting for {metric} on bucket {bucket} to satisfy {description}; last={value}, snapshot={snapshot:?}").into());
}
sleep(Duration::from_millis(200)).await;
}
}
}
impl Drop for ReplicationBacklogMetricCollector {
fn drop(&mut self) {
self.task.abort();
}
}
async fn handle_backlog_metric_export(
request: Request<Incoming>,
values: BacklogMetricPoints,
) -> Result<Response<Full<Bytes>>, Infallible> {
if request.uri().path() != "/v1/metrics" {
return Ok(empty_http_response(StatusCode::NOT_FOUND));
}
let gzip = request
.headers()
.get(CONTENT_ENCODING)
.and_then(|value| value.to_str().ok())
.is_some_and(|value| value.eq_ignore_ascii_case("gzip"));
let Ok(collected) = request.into_body().collect().await else {
return Ok(empty_http_response(StatusCode::BAD_REQUEST));
};
let body = collected.to_bytes();
if body.len() as u64 > OTLP_METRICS_BODY_LIMIT {
return Ok(empty_http_response(StatusCode::PAYLOAD_TOO_LARGE));
}
let payload = if gzip {
let mut decoder = GzDecoder::new(body.as_ref());
let mut decoded = Vec::new();
if decoder
.by_ref()
.take(OTLP_METRICS_BODY_LIMIT + 1)
.read_to_end(&mut decoded)
.is_err()
|| decoded.len() as u64 > OTLP_METRICS_BODY_LIMIT
{
return Ok(empty_http_response(StatusCode::BAD_REQUEST));
}
decoded
} else {
body.to_vec()
};
match ExportMetricsServiceRequest::decode(payload.as_slice()) {
Ok(export) => {
let mut values = values.lock().await;
record_backlog_metrics(&export, &mut values);
Ok(empty_http_response(StatusCode::OK))
}
Err(_) => Ok(empty_http_response(StatusCode::BAD_REQUEST)),
}
}
fn empty_http_response(status: StatusCode) -> Response<Full<Bytes>> {
Response::builder()
.status(status)
.body(Full::new(Bytes::new()))
.expect("static HTTP response is valid")
}
fn record_backlog_metrics(export: &ExportMetricsServiceRequest, values: &mut BTreeMap<String, BTreeMap<String, (u64, f64)>>) {
for resource_metrics in &export.resource_metrics {
for scope_metrics in &resource_metrics.scope_metrics {
for metric in &scope_metrics.metrics {
record_backlog_metric(metric, values);
}
}
}
}
fn record_backlog_metric(metric: &Metric, values: &mut BTreeMap<String, BTreeMap<String, (u64, f64)>>) {
if ![
TOTAL_FAILED_COUNT_METRIC,
CURRENT_BACKLOG_COUNT_METRIC,
CURRENT_BACKLOG_BYTES_METRIC,
MRF_PENDING_COUNT_METRIC,
MRF_PENDING_BYTES_METRIC,
]
.contains(&metric.name.as_str())
{
return;
}
let points = match &metric.data {
Some(metric::Data::Gauge(gauge)) => gauge.data_points.as_slice(),
Some(metric::Data::Sum(sum)) => sum.data_points.as_slice(),
_ => return,
};
for point in points {
let Some(bucket) = attribute_string(&point.attributes, BUCKET_LABEL) else {
continue;
};
let Some(value) = number_point_value(point.value.as_ref()) else {
continue;
};
values
.entry(metric.name.clone())
.or_default()
.entry(bucket.to_string())
.and_modify(|current| {
if point.time_unix_nano >= current.0 {
*current = (point.time_unix_nano, value);
}
})
.or_insert((point.time_unix_nano, value));
}
}
fn number_point_value(value: Option<&number_data_point::Value>) -> Option<f64> {
match value? {
number_data_point::Value::AsDouble(value) => Some(*value),
number_data_point::Value::AsInt(value) => Some(*value as f64),
}
}
fn attribute_string<'a>(attributes: &'a [KeyValue], wanted_key: &str) -> Option<&'a str> {
attributes.iter().find_map(|attribute| {
if attribute.key != wanted_key {
return None;
}
match attribute.value.as_ref()?.value.as_ref()? {
AnyValue::StringValue(value) => Some(value.as_str()),
_ => None,
}
})
}
struct SlowReplicationTargetGuard {
task: Option<JoinHandle<()>>,
}
impl SlowReplicationTargetGuard {
async fn bind(address: &str, response_delay: Duration) -> Result<Self, Box<dyn Error + Send + Sync>> {
let listener = TcpListener::bind(address).await?;
let task = tokio::spawn(async move {
loop {
let Ok((stream, _)) = listener.accept().await else {
break;
};
tokio::spawn(async move {
let _ = http1::Builder::new()
.serve_connection(
TokioIo::new(stream),
service_fn(move |_request| async move {
sleep(response_delay).await;
Ok::<_, Infallible>(empty_http_response(StatusCode::SERVICE_UNAVAILABLE))
}),
)
.await;
});
}
});
Ok(Self { task: Some(task) })
}
async fn stop(mut self) {
if let Some(task) = self.task.take() {
task.abort();
let _ = task.await;
}
}
}
impl Drop for SlowReplicationTargetGuard {
fn drop(&mut self) {
if let Some(task) = self.task.take() {
task.abort();
}
}
}
#[derive(Debug, Clone, serde::Deserialize)] #[derive(Debug, Clone, serde::Deserialize)]
struct ReplicationResetStatusResponse { struct ReplicationResetStatusResponse {
@@ -3659,6 +3913,139 @@ async fn test_bucket_replication_recovers_after_target_outage() -> TestResult {
Ok(()) Ok(())
} }
/// backlog#1610 - black-box bucket replication backlog observability.
///
/// The source exports metrics through the same OTLP path production uses. A slow
/// loopback target keeps replication workers occupied long enough for the metrics
/// runtime to publish non-zero bucket backlog gauges; after the real target
/// returns, replication must converge and the exported current/MRF pending gauges
/// must settle back to zero even though the historical failed counter remains
/// non-zero.
#[tokio::test]
#[serial]
async fn test_bucket_replication_backlog_metrics_observe_outage_and_recovery() -> TestResult {
init_logging();
let collector = ReplicationBacklogMetricCollector::start().await?;
let mut source_env = RustFSTestEnvironment::new().await?;
let metric_root = collector.root_endpoint().to_string();
let metric_endpoint = collector.endpoint.clone();
let mut source_env_vars: Vec<(&str, &str)> = replication_fast_env().into_iter().collect();
source_env_vars.extend_from_slice(LOOPBACK_REPLICATION_TARGET_ENV);
source_env_vars.extend_from_slice(FAST_SCANNER_ENV);
source_env_vars.extend_from_slice(&[
("RUSTFS_OBS_ENDPOINT", metric_root.as_str()),
("RUSTFS_OBS_METRIC_ENDPOINT", metric_endpoint.as_str()),
("RUSTFS_OBS_METRICS_EXPORT_ENABLED", "true"),
("RUSTFS_OBS_TRACES_EXPORT_ENABLED", "false"),
("RUSTFS_OBS_LOGS_EXPORT_ENABLED", "false"),
("RUSTFS_OBS_METER_INTERVAL", "1"),
("RUSTFS_OBS_USE_STDOUT", "false"),
("RUSTFS_METRICS_BUCKET_REPLICATION_BANDWIDTH_INTERVAL_SEC", "1"),
]);
source_env.start_rustfs_server_with_env(vec![], &source_env_vars).await?;
let mut target_env = RustFSTestEnvironment::new().await?;
target_env.start_rustfs_server_without_cleanup(vec![]).await?;
let source_bucket = "repl-backlog-metrics-src";
let target_bucket = "repl-backlog-metrics-dst";
let source_client = source_env.create_s3_client();
let target_client = target_env.create_s3_client();
source_client.create_bucket().bucket(source_bucket).send().await?;
target_client.create_bucket().bucket(target_bucket).send().await?;
enable_bucket_versioning(&source_env, source_bucket).await?;
enable_bucket_versioning(&target_env, target_bucket).await?;
let target_arn = set_replication_target(&source_env, source_bucket, &target_env, target_bucket).await?;
put_bucket_replication(&source_env, source_bucket, &target_arn).await?;
source_client
.put_object()
.bucket(source_bucket)
.key("before-outage.txt")
.body(ByteStream::from_static(b"baseline written before backlog metrics outage"))
.send()
.await?;
assert_replication_converged(&source_client, source_bucket, &target_client, target_bucket).await?;
target_env.stop_server();
let slow_target = SlowReplicationTargetGuard::bind(&target_env.address, Duration::from_secs(5)).await?;
let outage_keys = ["metrics-outage-1.txt", "metrics-outage-2.txt", "metrics-outage-3.txt"];
for key in outage_keys {
source_client
.put_object()
.bucket(source_bucket)
.key(key)
.body(ByteStream::from(
format!("written while backlog metrics target was slow: {key}").into_bytes(),
))
.send()
.await?;
}
let current_backlog = collector
.wait_for_bucket_metric(
CURRENT_BACKLOG_COUNT_METRIC,
source_bucket,
|value| value >= 1.0,
"be at least 1 during outage",
)
.await?;
let current_bytes = collector
.wait_for_bucket_metric(
CURRENT_BACKLOG_BYTES_METRIC,
source_bucket,
|value| value > 0.0,
"report bytes during outage",
)
.await?;
assert!(
current_bytes >= current_backlog,
"backlog bytes should be at least the object count while queued; count={current_backlog}, bytes={current_bytes}"
);
let failed_count = collector
.wait_for_bucket_metric(
TOTAL_FAILED_COUNT_METRIC,
source_bucket,
|value| value >= 1.0,
"record at least one failed replication attempt during outage",
)
.await?;
slow_target.stop().await;
target_env.restart_server_preserving_data(vec![], &[]).await?;
assert_replication_converged(&source_client, source_bucket, &target_client, target_bucket).await?;
for metric in [
CURRENT_BACKLOG_COUNT_METRIC,
CURRENT_BACKLOG_BYTES_METRIC,
MRF_PENDING_COUNT_METRIC,
MRF_PENDING_BYTES_METRIC,
] {
collector
.wait_for_bucket_metric(metric, source_bucket, |value| value == 0.0, "settle back to zero after recovery")
.await?;
}
let final_failed_count = collector.bucket_metric_value(TOTAL_FAILED_COUNT_METRIC, source_bucket).await;
assert!(
final_failed_count >= failed_count,
"historical failed counter should remain non-zero after recovery while current backlog is zero; before={failed_count}, after={final_failed_count}"
);
let target_state = list_replication_state(&target_client, target_bucket).await?;
for key in ["before-outage.txt"].into_iter().chain(outage_keys) {
assert!(
target_state.iter().any(|entry| entry.key == key && !entry.delete_marker),
"target missing object {key} after backlog metrics recovery; state={target_state:?}"
);
}
Ok(())
}
/// backlog#1147 repl-5, scenario (b) — failure state survives a source restart /// backlog#1147 repl-5, scenario (b) — failure state survives a source restart
/// (mirrors backlog#858 delete-decision re-derivation and #859 no-drop). /// (mirrors backlog#858 delete-decision re-derivation and #859 no-drop).
/// ///
+421 -37
View File
@@ -12,22 +12,35 @@
// See the License for the specific language governing permissions and // See the License for the specific language governing permissions and
// limitations under the License. // limitations under the License.
use crate::common::{RustFSTestEnvironment, init_logging, local_http_client}; use crate::common::{RustFSTestEnvironment, admin_ok, init_logging};
use aws_sdk_sts::config::retry::RetryConfig; use aws_sdk_sts::config::retry::RetryConfig;
use aws_sdk_sts::config::{Credentials, Region}; use aws_sdk_sts::config::{Credentials, Region};
use aws_sdk_sts::error::ProvideErrorMetadata; use aws_sdk_sts::error::ProvideErrorMetadata;
use aws_sdk_sts::operation::RequestId; use aws_sdk_sts::operation::RequestId;
use aws_sdk_sts::{Client, Config}; use aws_sdk_sts::{Client, Config};
use aws_smithy_http_client::Builder as SmithyHttpClientBuilder; use aws_smithy_http_client::Builder as SmithyHttpClientBuilder;
use http::header::{CONTENT_TYPE, HOST}; use bytes::Bytes;
use rustfs_signer::constants::UNSIGNED_PAYLOAD; use http::header::{AUTHORIZATION, CONTENT_TYPE};
use rustfs_signer::sign_v4; use http::{Request, Response};
use s3s::Body; use http_body_util::{BodyExt, Full};
use hyper::body::Incoming;
use hyper::server::conn::http1;
use hyper::service::service_fn;
use hyper_util::rt::TokioIo;
use serde_json::Value;
use serial_test::serial; use serial_test::serial;
use std::collections::BTreeSet;
use std::convert::Infallible;
use std::error::Error; use std::error::Error;
use std::sync::Arc;
use tokio::net::TcpListener;
use tokio::sync::{Notify, mpsc};
use tokio::task::{JoinHandle, JoinSet};
use tokio::time::{Duration, timeout};
type BoxError = Box<dyn Error + Send + Sync>; type BoxError = Box<dyn Error + Send + Sync>;
type TestResult = Result<(), BoxError>; type TestResult = Result<(), BoxError>;
const OPA_AUTH_TOKEN: &str = "sts-opa-token";
fn sts_client(url: &str, access_key: &str, secret_key: &str, session_token: Option<&str>) -> Client { fn sts_client(url: &str, access_key: &str, secret_key: &str, session_token: Option<&str>) -> Client {
let mut config = Config::builder() let mut config = Config::builder()
@@ -49,32 +62,14 @@ fn sts_client(url: &str, access_key: &str, secret_key: &str, session_token: Opti
} }
async fn create_root_service_account(env: &RustFSTestEnvironment) -> Result<(String, String), BoxError> { async fn create_root_service_account(env: &RustFSTestEnvironment) -> Result<(String, String), BoxError> {
let path = "/rustfs/admin/v3/add-service-accounts"; let body = admin_ok(
let url = format!("{}{path}", env.url); env,
let uri = url.parse::<http::Uri>()?; http::Method::PUT,
let authority = uri.authority().ok_or("admin URL missing authority")?.to_string(); "/rustfs/admin/v3/add-service-accounts",
let body = serde_json::json!({ "targetUser": env.access_key.clone() }).to_string(); Some(serde_json::json!({ "targetUser": env.access_key.clone() }).to_string()),
let request = http::Request::builder() )
.method(http::Method::PUT) .await?;
.uri(uri) let response: Value = serde_json::from_str(&body)?;
.header(HOST, authority)
.header(CONTENT_TYPE, "application/json")
.header("x-amz-content-sha256", UNSIGNED_PAYLOAD)
.body(Body::empty())?;
let content_length = i64::try_from(body.len()).map_err(|_| "service account request body is too large")?;
let signed = sign_v4(request, content_length, &env.access_key, &env.secret_key, "", "us-east-1");
let mut request = local_http_client().put(&url);
for (name, value) in signed.headers() {
request = request.header(name, value);
}
let response = request.body(body).send().await?;
let status = response.status();
let body = response.text().await?;
if !status.is_success() {
return Err(format!("create service account failed: {status} {body}").into());
}
let response: serde_json::Value = serde_json::from_str(&body)?;
let access_key = response["credentials"]["accessKey"] let access_key = response["credentials"]["accessKey"]
.as_str() .as_str()
.ok_or("service account response should contain credentials.accessKey")? .ok_or("service account response should contain credentials.accessKey")?
@@ -86,28 +81,237 @@ async fn create_root_service_account(env: &RustFSTestEnvironment) -> Result<(Str
Ok((access_key, secret_key)) Ok((access_key, secret_key))
} }
async fn assert_chaining_denied(client: &Client, credential_kind: &str) -> TestResult { async fn create_user_with_policy(
env: &RustFSTestEnvironment,
user: &str,
secret: &str,
policy_name: &str,
statements: Value,
) -> TestResult {
create_user(env, user, secret).await?;
admin_ok(
env,
http::Method::PUT,
&format!("/rustfs/admin/v3/add-canned-policy?name={policy_name}"),
Some(
serde_json::json!({
"Version": "2012-10-17",
"Statement": statements,
})
.to_string(),
),
)
.await?;
admin_ok(
env,
http::Method::POST,
"/rustfs/admin/v3/idp/builtin/policy/attach",
Some(serde_json::json!({ "policies": [policy_name], "user": user }).to_string()),
)
.await?;
Ok(())
}
async fn create_user(env: &RustFSTestEnvironment, user: &str, secret: &str) -> TestResult {
admin_ok(
env,
http::Method::PUT,
&format!("/rustfs/admin/v3/add-user?accessKey={user}"),
Some(serde_json::json!({ "secretKey": secret, "status": "enabled" }).to_string()),
)
.await?;
Ok(())
}
async fn assert_access_denied(client: &Client, context: &str) -> TestResult {
let error = client let error = client
.assume_role() .assume_role()
.role_arn("arn:aws:iam::123456789012:role/test") .role_arn("arn:aws:iam::123456789012:role/test")
.role_session_name("sts-query-compat-e2e") .role_session_name("sts-query-compat-e2e")
.send() .send()
.await .await
.expect_err("credential chaining must be denied"); .expect_err("AssumeRole must be denied");
let service_error = error let service_error = error
.as_service_error() .as_service_error()
.ok_or_else(|| format!("{credential_kind} denial should deserialize as an STS service error: {error:?}"))?; .ok_or_else(|| format!("{context} should deserialize as an STS service error: {error:?}"))?;
assert_eq!(error.raw_response().map(|response| response.status().as_u16()), Some(403)); assert_eq!(error.raw_response().map(|response| response.status().as_u16()), Some(403));
assert_eq!(service_error.code(), Some("AccessDenied")); assert_eq!(service_error.code(), Some("AccessDenied"));
assert_eq!(service_error.message(), Some("Access Denied")); assert_eq!(service_error.message(), Some("Access Denied"));
assert!( assert!(
error.request_id().is_some_and(|request_id| !request_id.is_empty()), error.request_id().is_some_and(|request_id| !request_id.is_empty()),
"{credential_kind} denial should include a request ID" "{context} should include a request ID"
); );
Ok(()) Ok(())
} }
async fn handle_opa_request(
request: Request<Incoming>,
requests: mpsc::UnboundedSender<Value>,
validation_started: mpsc::UnboundedSender<()>,
validation_mode: OpaValidationMode,
expected_authorization: Option<String>,
) -> Result<Response<Full<Bytes>>, Infallible> {
if let Some(expected_authorization) = expected_authorization
&& request.headers().get(AUTHORIZATION).and_then(|value| value.to_str().ok()) != Some(expected_authorization.as_str())
{
return Ok(Response::builder()
.status(401)
.body(Full::new(Bytes::new()))
.expect("static OPA unauthorized response must be valid"));
}
let body = match request.into_body().collect().await {
Ok(body) => body.to_bytes(),
Err(error) => {
return Ok(Response::builder()
.status(400)
.body(Full::new(Bytes::from(error.to_string())))
.expect("static OPA error response must be valid"));
}
};
let payload = if body.is_empty() {
None
} else {
match serde_json::from_slice::<Value>(&body) {
Ok(payload) => Some(payload),
Err(error) => {
return Ok(Response::builder()
.status(400)
.body(Full::new(Bytes::from(error.to_string())))
.expect("static OPA error response must be valid"));
}
}
};
if payload.is_none() {
let _ = validation_started.send(());
if let OpaValidationMode::DelayedUnavailable(release) = validation_mode {
release.notified().await;
return Ok(Response::builder()
.status(503)
.body(Full::new(Bytes::new()))
.expect("static OPA unavailable response must be valid"));
}
}
let allow = match payload.as_ref().and_then(|value| value.pointer("/input/identity/account")) {
Some(Value::String(account)) if account == "opaallow" => payload
.as_ref()
.and_then(|value| value.pointer("/input/context/deny_only"))
.and_then(Value::as_bool)
.unwrap_or(false),
Some(Value::String(account)) if account == "opadeny" => false,
None => true,
_ => false,
};
if let Some(payload) = payload {
let _ = requests.send(payload);
}
let body =
serde_json::to_vec(&serde_json::json!({ "result": { "allow": allow } })).expect("static OPA response must serialize");
Ok(Response::builder()
.header(CONTENT_TYPE, "application/json")
.body(Full::new(Bytes::from(body)))
.expect("static OPA response must be valid"))
}
#[derive(Clone)]
enum OpaValidationMode {
Ready,
DelayedUnavailable(Arc<Notify>),
}
struct OpaMock {
url: String,
requests: mpsc::UnboundedReceiver<Value>,
validation_started: mpsc::UnboundedReceiver<()>,
validation_release: Option<Arc<Notify>>,
task: JoinHandle<()>,
}
impl OpaMock {
async fn start() -> Result<Self, BoxError> {
Self::start_with_mode(OpaValidationMode::Ready, Some(OPA_AUTH_TOKEN)).await
}
async fn start_delayed_unavailable() -> Result<Self, BoxError> {
let release = Arc::new(Notify::new());
Self::start_with_mode(OpaValidationMode::DelayedUnavailable(release), None).await
}
async fn start_with_mode(validation_mode: OpaValidationMode, auth_token: Option<&str>) -> Result<Self, BoxError> {
let listener = TcpListener::bind("127.0.0.1:0").await?;
let url = format!("http://{}/v1/data/rustfs/authz/allow", listener.local_addr()?);
let (requests_tx, requests) = mpsc::unbounded_channel();
let (validation_started_tx, validation_started) = mpsc::unbounded_channel();
let expected_authorization = auth_token.map(|token| format!("Bearer {token}"));
let validation_release = match &validation_mode {
OpaValidationMode::Ready => None,
OpaValidationMode::DelayedUnavailable(release) => Some(Arc::clone(release)),
};
let task = tokio::spawn(async move {
let mut connections = JoinSet::new();
loop {
tokio::select! {
accepted = listener.accept() => {
let Ok((stream, _)) = accepted else { break };
let requests = requests_tx.clone();
let validation_started = validation_started_tx.clone();
let validation_mode = validation_mode.clone();
let expected_authorization = expected_authorization.clone();
connections.spawn(async move {
let handler = service_fn(move |request| {
handle_opa_request(
request,
requests.clone(),
validation_started.clone(),
validation_mode.clone(),
expected_authorization.clone(),
)
});
let _ = http1::Builder::new()
.serve_connection(TokioIo::new(stream), handler)
.await;
});
}
_ = connections.join_next(), if !connections.is_empty() => {}
}
}
});
Ok(Self {
url,
requests,
validation_started,
validation_release,
task,
})
}
async fn next_request(&mut self) -> Result<Value, BoxError> {
timeout(Duration::from_secs(5), self.requests.recv())
.await?
.ok_or_else(|| "OPA request channel closed".into())
}
async fn wait_for_validation(&mut self) -> TestResult {
timeout(Duration::from_secs(5), self.validation_started.recv())
.await?
.ok_or_else(|| "OPA validation channel closed".into())
}
fn release_validation(&self) {
if let Some(release) = &self.validation_release {
release.notify_one();
}
}
}
impl Drop for OpaMock {
fn drop(&mut self) {
self.task.abort();
}
}
#[tokio::test] #[tokio::test]
#[serial] #[serial]
async fn test_sts_query_responses_are_aws_sdk_compatible() -> TestResult { async fn test_sts_query_responses_are_aws_sdk_compatible() -> TestResult {
@@ -155,7 +359,7 @@ async fn test_sts_query_responses_are_aws_sdk_compatible() -> TestResult {
"signature rejection should include a request ID" "signature rejection should include a request ID"
); );
assert_chaining_denied( assert_access_denied(
&sts_client( &sts_client(
&env.url, &env.url,
temporary.access_key_id(), temporary.access_key_id(),
@@ -167,7 +371,187 @@ async fn test_sts_query_responses_are_aws_sdk_compatible() -> TestResult {
.await?; .await?;
let (service_access_key, service_secret_key) = create_root_service_account(&env).await?; let (service_access_key, service_secret_key) = create_root_service_account(&env).await?;
assert_chaining_denied(&sts_client(&env.url, &service_access_key, &service_secret_key, None), "service account").await?; assert_access_denied(
&sts_client(&env.url, &service_access_key, &service_secret_key, None),
"service-account denial",
)
.await?;
let implicit_user = "stsimplicit";
let explicit_allow_user = "stsallow";
let explicit_deny_user = "stsdeny";
let policyless_user = "stspolicyless";
let secret = "stsAuthzSecret123";
create_user(&env, policyless_user, secret).await?;
assert_access_denied(&sts_client(&env.url, policyless_user, secret, None), "policyless user").await?;
create_user_with_policy(
&env,
implicit_user,
secret,
"sts-implicit-policy",
serde_json::json!([{
"Effect": "Allow",
"Action": ["s3:ListAllMyBuckets"],
"Resource": ["arn:aws:s3:::*"],
}]),
)
.await?;
create_user_with_policy(
&env,
explicit_allow_user,
secret,
"sts-allow-policy",
serde_json::json!([{
"Effect": "Allow",
"Action": ["sts:AssumeRole"],
"Resource": ["arn:aws:s3:::*"],
}]),
)
.await?;
create_user_with_policy(
&env,
explicit_deny_user,
secret,
"sts-deny-policy",
serde_json::json!([
{
"Effect": "Allow",
"Action": ["sts:AssumeRole"],
"Resource": ["arn:aws:s3:::*"],
},
{
"Effect": "Deny",
"Action": ["sts:AssumeRole"],
"Resource": ["arn:aws:s3:::*"],
}
]),
)
.await?;
for user in [implicit_user, explicit_allow_user] {
let output = sts_client(&env.url, user, secret, None)
.assume_role()
.role_arn("arn:aws:iam::123456789012:role/test")
.role_session_name("sts-authz-e2e")
.send()
.await
.map_err(|error| format!("{user} should be allowed to call AssumeRole: {error:?}"))?;
let credentials = output
.credentials()
.ok_or_else(|| format!("{user} AssumeRole response should contain credentials"))?;
assert!(!credentials.access_key_id().is_empty());
assert!(!credentials.secret_access_key().is_empty());
assert!(!credentials.session_token().is_empty());
}
assert_access_denied(&sts_client(&env.url, explicit_deny_user, secret, None), "explicit sts:AssumeRole Deny").await?;
env.stop_server();
Ok(())
}
#[tokio::test]
#[serial]
async fn test_sts_assume_role_opa_contract() -> TestResult {
init_logging();
let mut opa = OpaMock::start().await?;
let mut env = RustFSTestEnvironment::new().await?;
env.start_rustfs_server_with_env(
vec![],
&[
("RUSTFS_POLICY_PLUGIN_URL", opa.url.as_str()),
("RUSTFS_POLICY_PLUGIN_AUTH_TOKEN", OPA_AUTH_TOKEN),
],
)
.await?;
let secret = "stsOpaSecret123";
create_user_with_policy(
&env,
"opaallow",
secret,
"sts-opa-local-deny-policy",
serde_json::json!([{
"Effect": "Deny",
"Action": ["sts:AssumeRole"],
"Resource": ["arn:aws:s3:::*"],
}]),
)
.await?;
create_user_with_policy(
&env,
"opadeny",
secret,
"sts-opa-local-allow-policy",
serde_json::json!([{
"Effect": "Allow",
"Action": ["sts:AssumeRole"],
"Resource": ["arn:aws:s3:::*"],
}]),
)
.await?;
sts_client(&env.url, "opaallow", secret, None)
.assume_role()
.role_arn("arn:aws:iam::123456789012:role/test")
.role_session_name("sts-opa-contract")
.send()
.await
.map_err(|error| format!("OPA allow should override the local explicit Deny: {error:?}"))?;
assert_access_denied(
&sts_client(&env.url, "opadeny", secret, None),
"OPA denial despite local sts:AssumeRole Allow",
)
.await?;
let mut accounts = BTreeSet::new();
for _ in 0..2 {
let request = opa.next_request().await?;
assert_eq!(request.pointer("/input/action").and_then(Value::as_str), Some("sts:AssumeRole"));
assert_eq!(request.pointer("/input/context/deny_only").and_then(Value::as_bool), Some(true));
let account = request
.pointer("/input/identity/account")
.and_then(Value::as_str)
.ok_or("OPA input should include identity.account")?;
accounts.insert(account.to_owned());
}
assert_eq!(accounts, BTreeSet::from(["opaallow".to_owned(), "opadeny".to_owned()]));
env.stop_server();
Ok(())
}
#[tokio::test]
#[serial]
async fn test_sts_assume_role_fails_closed_while_opa_is_unavailable() -> TestResult {
init_logging();
let mut opa = OpaMock::start_delayed_unavailable().await?;
let mut env = RustFSTestEnvironment::new().await?;
env.start_rustfs_server_with_env(vec![], &[("RUSTFS_POLICY_PLUGIN_URL", opa.url.as_str())])
.await?;
opa.wait_for_validation().await?;
let user = "opaunavailable";
let secret = "stsOpaUnavailableSecret123";
create_user_with_policy(
&env,
user,
secret,
"sts-opa-unavailable-local-policy",
serde_json::json!([{
"Effect": "Allow",
"Action": ["s3:ListAllMyBuckets"],
"Resource": ["arn:aws:s3:::*"],
}]),
)
.await?;
assert_access_denied(&sts_client(&env.url, user, secret, None), "configured OPA initialization").await?;
opa.release_validation();
tokio::time::sleep(Duration::from_millis(200)).await;
assert_access_denied(&sts_client(&env.url, user, secret, None), "configured OPA validation failure").await?;
env.stop_server(); env.stop_server();
Ok(()) Ok(())
+28 -22
View File
@@ -172,6 +172,11 @@ pub mod bucket {
} }
pub mod replication { pub mod replication {
pub use crate::bucket::replication::replication_pool::{
DurableMrfBacklogSummary, DurableMrfBucketBacklog, DurableMrfTargetBacklog, MrfBacklogObservabilitySummary,
MrfBucketBacklogObservability, durable_mrf_backlog_summary_snapshot, durable_mrf_target_backlog_snapshot,
mrf_backlog_observability_snapshot,
};
pub use crate::bucket::replication::{ pub use crate::bucket::replication::{
BucketReplicationResyncStatus, BucketStats, DeletedObjectReplicationInfo, DurableMrfBacklog, DynReplicationPool, BucketReplicationResyncStatus, BucketStats, DeletedObjectReplicationInfo, DurableMrfBacklog, DynReplicationPool,
MrfOpKind, MrfReplicateEntry, MustReplicateOptions, ObjectOpts, REPLICATE_INCOMING_DELETE, ReplicateDecision, MrfOpKind, MrfReplicateEntry, MustReplicateOptions, ObjectOpts, REPLICATE_INCOMING_DELETE, ReplicateDecision,
@@ -179,13 +184,13 @@ pub mod bucket {
ReplicationDeleteStateSource, ReplicationHealQueueResult, ReplicationObjectBridge, ReplicationObjectIO, ReplicationDeleteStateSource, ReplicationHealQueueResult, ReplicationObjectBridge, ReplicationObjectIO,
ReplicationOperation, ReplicationPoolTrait, ReplicationPriority, ReplicationQueueAdmission, ReplicationScannerBridge, ReplicationOperation, ReplicationPoolTrait, ReplicationPriority, ReplicationQueueAdmission, ReplicationScannerBridge,
ReplicationState, ReplicationStats, ReplicationStatusType, ReplicationStorage, ReplicationTargetValidationError, ReplicationState, ReplicationStats, ReplicationStatusType, ReplicationStorage, ReplicationTargetValidationError,
ReplicationType, ResyncOpts, ResyncStatusType, TargetReplicationResyncStatus, VersionPurgeStatusType, ReplicationType, ResyncOpts, ResyncStatusType, RuntimeReplicationTargetBacklog, TargetReplicationResyncStatus,
delete_replication_state_from_config, delete_replication_version_id, get_global_replication_pool, VersionPurgeStatusType, delete_replication_state_from_config, delete_replication_version_id,
get_global_replication_stats, init_background_replication, read_durable_mrf_backlog, replication_state_to_filemeta, get_global_replication_pool, get_global_replication_stats, init_background_replication, read_durable_mrf_backlog,
replication_status_to_filemeta, replication_statuses_map, replication_target_arns, resync_start_conflict_id, replication_state_to_filemeta, replication_status_to_filemeta, replication_statuses_map, replication_target_arns,
should_remove_replication_target, should_schedule_delete_replication, should_use_existing_delete_replication_info, resync_start_conflict_id, should_remove_replication_target, should_schedule_delete_replication,
should_use_existing_delete_replication_source, validate_replication_config_target_arns, should_use_existing_delete_replication_info, should_use_existing_delete_replication_source,
version_purge_status_to_filemeta, unsupported_replication_config_field, validate_replication_config_target_arns, version_purge_status_to_filemeta,
}; };
} }
@@ -267,12 +272,13 @@ pub mod config {
pub mod com { pub mod com {
pub use crate::config::com::{ pub use crate::config::com::{
COMMA_SEPARATED_LISTS, CONFIG_PREFIX, ENV_CONFIG_RECOVER_ON_CORRUPTION, STORAGE_CLASS_SUB_SYS, COMMA_SEPARATED_LISTS, CONFIG_PREFIX, ENV_CONFIG_RECOVER_ON_CORRUPTION, STORAGE_CLASS_SUB_SYS,
ServerConfigCorruptError, ServerConfigSnapshot, delete_config, is_server_config_corrupt_error, lookup_configs, ServerConfigCorruptError, ServerConfigSaveResult, ServerConfigSnapshot, delete_config,
read_config, read_config_no_lock, read_config_with_metadata, read_config_without_migrate, is_server_config_corrupt_error, lookup_configs, read_config, read_config_no_lock, read_config_with_metadata,
read_config_without_migrate_no_lock, read_existing_server_config_no_lock, read_server_config_snapshot, save_config, read_config_without_migrate, read_config_without_migrate_no_lock, read_existing_server_config_no_lock,
save_config_no_lock, save_config_with_opts, save_server_config, save_server_config_no_lock, read_server_config_snapshot, save_config, save_config_no_lock, save_config_with_opts, save_server_config,
save_server_config_snapshot, server_config_path, try_migrate_server_config, with_config_object_read_lock, save_server_config_no_lock, save_server_config_snapshot, save_server_config_snapshot_with_generation,
with_config_object_write_lock, with_server_config_read_lock, with_server_config_write_lock, server_config_path, try_migrate_server_config, with_config_object_read_lock, with_config_object_write_lock,
with_server_config_read_lock, with_server_config_write_lock,
}; };
} }
@@ -418,15 +424,15 @@ pub mod rio {
pub mod rpc { pub mod rpc {
pub use crate::cluster::rpc::{ pub use crate::cluster::rpc::{
AuthenticatedChannel, LocalPeerS3Client, PEER_RESTDRY_RUN, PEER_RESTSIGNAL, PEER_RESTSUB_SYS, PeerRestClient, AuthenticatedChannel, KMS_SIGNAL_SUBSYSTEM, LocalPeerS3Client, PEER_RESTDRY_RUN, PEER_RESTSIGNAL, PEER_RESTSUB_SYS,
PeerS3Client, S3PeerSys, SERVICE_SIGNAL_REFRESH_CONFIG, SERVICE_SIGNAL_RELOAD_DYNAMIC, ScannerBucketListing, PeerRestClient, PeerS3Client, S3PeerSys, SERVICE_SIGNAL_REFRESH_CONFIG, SERVICE_SIGNAL_RELOAD_DYNAMIC,
ScannerPeerActivity, TONIC_RPC_PREFIX, TonicInterceptor, gen_signature_headers, gen_tonic_replay_scope_headers, ScannerBucketListing, ScannerPeerActivity, TONIC_RPC_PREFIX, TonicInterceptor, gen_signature_headers,
gen_tonic_signature_headers, gen_tonic_signature_interceptor, node_service_time_out_client, gen_tonic_replay_scope_headers, gen_tonic_signature_headers, gen_tonic_signature_interceptor,
node_service_time_out_client_no_auth, normalize_tonic_rpc_audience, set_tonic_canonical_body_digest, node_service_time_out_client, node_service_time_out_client_no_auth, normalize_tonic_rpc_audience,
sign_ns_scanner_capability, sign_tonic_rpc_response_proof, tonic_boot_epoch_challenge, tonic_boot_epoch_response_headers, set_tonic_canonical_body_digest, sign_ns_scanner_capability, sign_tonic_rpc_response_proof, tonic_boot_epoch_challenge,
verify_rpc_signature, verify_tonic_boot_epoch_response, verify_tonic_canonical_body_digest, tonic_boot_epoch_response_headers, verify_rpc_signature, verify_tonic_boot_epoch_response,
verify_tonic_mutation_body_digest, verify_tonic_rpc_response_proof, verify_tonic_rpc_signature, verify_tonic_canonical_body_digest, verify_tonic_mutation_body_digest, verify_tonic_rpc_response_proof,
verify_tonic_rpc_signature_with_bootstrap, verify_tonic_rpc_signature, verify_tonic_rpc_signature_with_bootstrap,
}; };
} }
+2 -2
View File
@@ -46,7 +46,7 @@ mod runtime_boundary;
pub use datatypes::ResyncStatusType; pub use datatypes::ResyncStatusType;
pub use replication_config_boundary::{ pub use replication_config_boundary::{
ObjectOpts, ReplicationConfigurationExt, ReplicationTargetValidationError, replication_target_arns, ObjectOpts, ReplicationConfigurationExt, ReplicationTargetValidationError, replication_target_arns,
should_remove_replication_target, validate_replication_config_target_arns, should_remove_replication_target, unsupported_replication_config_field, validate_replication_config_target_arns,
}; };
#[cfg(test)] #[cfg(test)]
pub(crate) use replication_filemeta_boundary::ReplicateTargetDecision; pub(crate) use replication_filemeta_boundary::ReplicateTargetDecision;
@@ -78,7 +78,7 @@ pub use replication_queue_boundary::{
}; };
pub use replication_resync_boundary::{BucketReplicationResyncStatus, ResyncOpts, TargetReplicationResyncStatus}; pub use replication_resync_boundary::{BucketReplicationResyncStatus, ResyncOpts, TargetReplicationResyncStatus};
pub use replication_scanner_bridge::ReplicationScannerBridge; pub use replication_scanner_bridge::ReplicationScannerBridge;
pub use replication_state::ReplicationStats; pub use replication_state::{ReplicationStats, RuntimeReplicationTargetBacklog};
pub use replication_stats_boundary::BucketStats; pub use replication_stats_boundary::BucketStats;
pub use replication_storage_boundary::{ReplicationObjectIO, ReplicationStorage}; pub use replication_storage_boundary::{ReplicationObjectIO, ReplicationStorage};
pub(crate) use replication_target_config_bridge::ReplicationTargetConfigBridge; pub(crate) use replication_target_config_bridge::ReplicationTargetConfigBridge;
@@ -14,5 +14,5 @@
pub use rustfs_replication::{ pub use rustfs_replication::{
ObjectOpts, ReplicationConfigurationExt, ReplicationTargetValidationError, replication_target_arns, ObjectOpts, ReplicationConfigurationExt, ReplicationTargetValidationError, replication_target_arns,
should_remove_replication_target, validate_replication_config_target_arns, should_remove_replication_target, unsupported_replication_config_field, validate_replication_config_target_arns,
}; };
@@ -14,8 +14,11 @@
use std::{collections::HashMap, sync::Arc}; use std::{collections::HashMap, sync::Arc};
use super::replication_error_boundary::Result;
use super::replication_filemeta_boundary::{ReplicateDecision, ReplicationStatusType, ReplicationType}; use super::replication_filemeta_boundary::{ReplicateDecision, ReplicationStatusType, ReplicationType};
use super::replication_object_config::{check_replicate_delete, get_must_replicate_options, must_replicate}; use super::replication_object_config::{
check_replicate_delete, check_replicate_delete_strict, get_must_replicate_options, must_replicate,
};
use super::replication_object_decision_boundary::MustReplicateOptions; use super::replication_object_decision_boundary::MustReplicateOptions;
use super::replication_pool::{schedule_replication, schedule_replication_delete}; use super::replication_pool::{schedule_replication, schedule_replication_delete};
use super::replication_queue_boundary::DeletedObjectReplicationInfo; use super::replication_queue_boundary::DeletedObjectReplicationInfo;
@@ -50,6 +53,16 @@ impl ReplicationObjectBridge {
check_replicate_delete(bucket, object, source, opts, get_error).await check_replicate_delete(bucket, object, source, opts, get_error).await
} }
pub async fn check_delete_strict(
bucket: &str,
object: &ObjectToDelete,
source: &ObjectInfo,
opts: &ObjectOptions,
get_error: Option<String>,
) -> Result<ReplicateDecision> {
check_replicate_delete_strict(bucket, object, source, opts, get_error).await
}
pub async fn schedule_object<S: ReplicationStorage>( pub async fn schedule_object<S: ReplicationStorage>(
object: ObjectInfo, object: ObjectInfo,
storage: Arc<S>, storage: Arc<S>,
@@ -169,9 +169,8 @@ pub(crate) async fn check_replicate_delete(
del_opts: &ObjectOptions, del_opts: &ObjectOptions,
gerr: Option<String>, gerr: Option<String>,
) -> ReplicateDecision { ) -> ReplicateDecision {
let rcfg = match get_replication_config(bucket).await { match check_replicate_delete_strict(bucket, dobj, oi, del_opts, gerr).await {
Ok(Some(config)) => config, Ok(decision) => decision,
Ok(None) => return ReplicateDecision::default(),
Err(err) => { Err(err) => {
error!( error!(
event = EVENT_RESYNC_CONFIG_LOOKUP_SKIPPED, event = EVENT_RESYNC_CONFIG_LOOKUP_SKIPPED,
@@ -182,16 +181,30 @@ pub(crate) async fn check_replicate_delete(
error = %err, error = %err,
"Failed to look up replication config for delete replication" "Failed to look up replication config for delete replication"
); );
return ReplicateDecision::default(); ReplicateDecision::default()
} }
}
}
pub(crate) async fn check_replicate_delete_strict(
bucket: &str,
dobj: &ObjectToDelete,
oi: &ObjectInfo,
del_opts: &ObjectOptions,
gerr: Option<String>,
) -> Result<ReplicateDecision> {
let rcfg = match get_replication_config(bucket).await {
Ok(Some(config)) => config,
Ok(None) => return Ok(ReplicateDecision::default()),
Err(err) => return Err(err),
}; };
if del_opts.replication_request { if del_opts.replication_request {
return ReplicateDecision::default(); return Ok(ReplicateDecision::default());
} }
if !del_opts.versioned { if !del_opts.versioned && !del_opts.version_suspended {
return ReplicateDecision::default(); return Ok(ReplicateDecision::default());
} }
let replication_delete = object_to_delete_for_replication(dobj); let replication_delete = object_to_delete_for_replication(dobj);
@@ -209,7 +222,7 @@ pub(crate) async fn check_replicate_delete(
let mut dsc = ReplicateDecision::new(); let mut dsc = ReplicateDecision::new();
if tgt_arns.is_empty() { if tgt_arns.is_empty() {
return dsc; return Ok(dsc);
} }
for tgt_arn in tgt_arns { for tgt_arn in tgt_arns {
@@ -239,7 +252,7 @@ pub(crate) async fn check_replicate_delete(
dsc.set(tgt_dsc); dsc.set(tgt_dsc);
} }
dsc Ok(dsc)
} }
pub(crate) async fn must_replicate(bucket: &str, object: &str, mopts: MustReplicateOptions) -> ReplicateDecision { pub(crate) async fn must_replicate(bucket: &str, object: &str, mopts: MustReplicateOptions) -> ReplicateDecision {
File diff suppressed because it is too large Load Diff
@@ -1151,60 +1151,6 @@ pub async fn replicate_delete<S: ReplicationStorage>(dobj: DeletedObjectReplicat
dobj.delete_object.version_id dobj.delete_object.version_id
}; };
let _rcfg = match get_replication_config(&bucket).await {
Ok(Some(config)) => config,
Ok(None) => {
debug!(
event = EVENT_REPLICATION_DELETE_SKIPPED,
component = LOG_COMPONENT_ECSTORE,
subsystem = LOG_SUBSYSTEM_REPLICATION_RESYNC,
bucket = %bucket,
reason = "replication_config_missing",
"Skipping replication delete because replication config is missing"
);
send_local_event(EventArgs {
event_name: EventName::ObjectReplicationNotTracked.to_string(),
bucket_name: bucket.clone(),
object: ObjectInfo {
bucket: bucket.clone(),
name: dobj.delete_object.object_name.clone(),
version_id,
delete_marker: dobj.delete_object.delete_marker,
..Default::default()
},
user_agent: "Internal: [Replication]".to_string(),
..Default::default()
});
return;
}
Err(err) => {
debug!(
event = EVENT_REPLICATION_DELETE_SKIPPED,
component = LOG_COMPONENT_ECSTORE,
subsystem = LOG_SUBSYSTEM_REPLICATION_RESYNC,
bucket = %bucket,
error = %err,
reason = "replication_config_lookup_failed",
"Skipping replication delete because replication config lookup failed"
);
send_local_event(EventArgs {
event_name: EventName::ObjectReplicationNotTracked.to_string(),
bucket_name: bucket.clone(),
object: ObjectInfo {
bucket: bucket.clone(),
name: dobj.delete_object.object_name.clone(),
version_id,
delete_marker: dobj.delete_object.delete_marker,
..Default::default()
},
user_agent: "Internal: [Replication]".to_string(),
..Default::default()
});
return;
}
};
if dobj.delete_object.delete_marker if dobj.delete_object.delete_marker
&& let Some(delete_marker_version_id) = dobj.delete_object.delete_marker_version_id && let Some(delete_marker_version_id) = dobj.delete_object.delete_marker_version_id
{ {
@@ -22,9 +22,9 @@ use super::replication_stats_boundary::{
QueueCache, ReplicationMetricScope, SRMetricsSummary, XferStats, QueueCache, ReplicationMetricScope, SRMetricsSummary, XferStats,
}; };
use super::runtime_boundary as runtime_sources; use super::runtime_boundary as runtime_sources;
use std::collections::HashMap; use std::collections::{HashMap, hash_map::Entry};
use std::sync::Arc;
use std::sync::atomic::{AtomicI64, Ordering}; use std::sync::atomic::{AtomicI64, Ordering};
use std::sync::{Arc, LazyLock, Mutex as StdMutex, Weak};
use std::time::{Duration, SystemTime}; use std::time::{Duration, SystemTime};
use tokio::sync::{Mutex, RwLock}; use tokio::sync::{Mutex, RwLock};
use tokio::time::interval; use tokio::time::interval;
@@ -143,7 +143,7 @@ pub struct ReplicationStats {
// Active worker statistics // Active worker statistics
pub workers: Arc<Mutex<ActiveWorkerStat>>, pub workers: Arc<Mutex<ActiveWorkerStat>>,
// Queue statistics cache // Queue statistics cache
pub q_cache: Arc<Mutex<QueueCache>>, pub q_cache: Arc<StdMutex<QueueCache>>,
// Proxy statistics cache // Proxy statistics cache
pub p_cache: Arc<Mutex<ProxyStatsCache>>, pub p_cache: Arc<Mutex<ProxyStatsCache>>,
// MRF backlog statistics (simplified) // MRF backlog statistics (simplified)
@@ -153,12 +153,89 @@ pub struct ReplicationStats {
pub most_recent_stats: Arc<Mutex<HashMap<String, BucketStats>>>, pub most_recent_stats: Arc<Mutex<HashMap<String, BucketStats>>>,
} }
#[derive(Debug, Clone, Default, PartialEq, Eq)]
pub struct RuntimeReplicationTargetBacklog {
pub bucket: String,
pub target_arn: String,
pub count: u64,
pub bytes: u64,
}
type TargetQueueKey = (String, String);
type TargetQueueCache = HashMap<TargetQueueKey, InQueueMetric>;
struct TargetQueueCacheSlot {
owner: Weak<StdMutex<QueueCache>>,
metrics: TargetQueueCache,
}
impl TargetQueueCacheSlot {
fn new(owner: &Arc<StdMutex<QueueCache>>) -> Self {
Self {
owner: Arc::downgrade(owner),
metrics: TargetQueueCache::default(),
}
}
fn belongs_to(&self, owner: &Arc<StdMutex<QueueCache>>) -> bool {
self.owner.upgrade().is_some_and(|current| Arc::ptr_eq(&current, owner))
}
}
// Keep runtime target counters outside ReplicationStats to preserve its public struct shape.
static TARGET_QUEUE_CACHES: LazyLock<StdMutex<Vec<TargetQueueCacheSlot>>> = LazyLock::new(|| StdMutex::new(Vec::new()));
fn i64_to_u64_floor_zero(value: i64) -> u64 {
u64::try_from(value.max(0)).unwrap_or(0)
}
fn normalized_target_arns(target_arns: &[String]) -> Vec<&str> {
let mut target_arns = target_arns
.iter()
.map(String::as_str)
.filter(|target_arn| !target_arn.is_empty())
.collect::<Vec<_>>();
target_arns.sort_unstable();
target_arns.dedup();
target_arns
}
fn target_queue_cache_snapshot(cache: &TargetQueueCache) -> Vec<RuntimeReplicationTargetBacklog> {
cache
.iter()
.filter_map(|((bucket, target_arn), metric)| {
let count = i64_to_u64_floor_zero(metric.curr.get_current_count());
let bytes = i64_to_u64_floor_zero(metric.curr.get_current_bytes());
(count > 0 || bytes > 0).then(|| RuntimeReplicationTargetBacklog {
bucket: bucket.clone(),
target_arn: target_arn.clone(),
count,
bytes,
})
})
.collect()
}
fn prune_stale_target_queue_caches(caches: &mut Vec<TargetQueueCacheSlot>) {
caches.retain(|slot| slot.owner.strong_count() > 0);
}
fn with_target_queue_caches<T>(f: impl FnOnce(&mut Vec<TargetQueueCacheSlot>) -> T) -> T {
match TARGET_QUEUE_CACHES.lock() {
Ok(mut caches) => f(&mut caches),
Err(poisoned) => {
let mut caches = poisoned.into_inner();
f(&mut caches)
}
}
}
impl ReplicationStats { impl ReplicationStats {
pub fn new() -> Self { pub fn new() -> Self {
Self { Self {
sr_stats: Arc::new(SRStats::new()), sr_stats: Arc::new(SRStats::new()),
workers: Arc::new(Mutex::new(ActiveWorkerStat::new())), workers: Arc::new(Mutex::new(ActiveWorkerStat::new())),
q_cache: Arc::new(Mutex::new(QueueCache::new())), q_cache: Arc::new(StdMutex::new(QueueCache::new())),
p_cache: Arc::new(Mutex::new(ProxyStatsCache::new())), p_cache: Arc::new(Mutex::new(ProxyStatsCache::new())),
mrf_stats: HashMap::new(), mrf_stats: HashMap::new(),
cache: Arc::new(RwLock::new(HashMap::new())), cache: Arc::new(RwLock::new(HashMap::new())),
@@ -198,9 +275,10 @@ impl ReplicationStats {
let mut interval = interval(Duration::from_secs(2)); let mut interval = interval(Duration::from_secs(2));
loop { loop {
interval.tick().await; interval.tick().await;
let mut cache = q_cache_clone.lock().await; if let Ok(mut cache) = q_cache_clone.lock() {
cache.update(); cache.update();
} }
}
}); });
} }
@@ -391,7 +469,7 @@ impl ReplicationStats {
drop(cache); drop(cache);
{ {
let q_cache = self.q_cache.lock().await; if let Ok(q_cache) = self.q_cache.lock() {
for (bucket, queue_stats) in &q_cache.bucket_stats { for (bucket, queue_stats) in &q_cache.bucket_stats {
let bucket_stats = result.entry(bucket.clone()).or_insert_with(BucketReplicationStats::new); let bucket_stats = result.entry(bucket.clone()).or_insert_with(BucketReplicationStats::new);
bucket_stats.q_stat = queue_stats.snapshot(); bucket_stats.q_stat = queue_stats.snapshot();
@@ -399,6 +477,7 @@ impl ReplicationStats {
bucket_stats.queue_scope = ReplicationMetricScope::NodeLocal; bucket_stats.queue_scope = ReplicationMetricScope::NodeLocal;
} }
} }
}
{ {
let p_cache = self.p_cache.lock().await; let p_cache = self.p_cache.lock().await;
@@ -429,8 +508,11 @@ impl ReplicationStats {
let boot_time = SystemTime::UNIX_EPOCH; // simplified implementation let boot_time = SystemTime::UNIX_EPOCH; // simplified implementation
let uptime = SystemTime::now().duration_since(boot_time).unwrap_or_default().as_secs() as i64; let uptime = SystemTime::now().duration_since(boot_time).unwrap_or_default().as_secs() as i64;
let q_cache = self.q_cache.lock().await; let queued = self
let queued = q_cache.get_site_stats(); .q_cache
.lock()
.map(|q_cache| q_cache.get_site_stats())
.unwrap_or_default();
let p_cache = self.p_cache.lock().await; let p_cache = self.p_cache.lock().await;
let proxied = p_cache.get_site_stats(); let proxied = p_cache.get_site_stats();
@@ -633,8 +715,9 @@ impl ReplicationStats {
drop(cache); drop(cache);
{ {
let q_cache = self.q_cache.lock().await; if let Ok(q_cache) = self.q_cache.lock()
if let Some(queue_stats) = q_cache.bucket_stats.get(bucket) { && let Some(queue_stats) = q_cache.bucket_stats.get(bucket)
{
replication_stats.q_stat = queue_stats.snapshot(); replication_stats.q_stat = queue_stats.snapshot();
} }
} }
@@ -665,31 +748,79 @@ impl ReplicationStats {
} }
/// Increase queue statistics /// Increase queue statistics
pub async fn inc_q(&self, bucket: &str, size: i64, _is_delete_repl: bool, _op_type: ReplicationType) { pub fn inc_q(&self, bucket: &str, size: i64, _is_delete_repl: bool, _op_type: ReplicationType) {
let mut q_cache = self.q_cache.lock().await; if let Ok(mut q_cache) = self.q_cache.lock() {
let stats = q_cache q_cache.inc(bucket, size);
.bucket_stats }
.entry(bucket.to_string())
.or_insert_with(InQueueMetric::default);
stats.curr.now_bytes.fetch_add(size, Ordering::Relaxed);
stats.curr.now_count.fetch_add(1, Ordering::Relaxed);
q_cache.sr_queue_stats.curr.now_bytes.fetch_add(size, Ordering::Relaxed);
q_cache.sr_queue_stats.curr.now_count.fetch_add(1, Ordering::Relaxed);
} }
/// Decrease queue statistics /// Decrease queue statistics
pub async fn dec_q(&self, bucket: &str, size: i64, _is_del_marker: bool, _op_type: ReplicationType) { pub fn dec_q(&self, bucket: &str, size: i64, _is_del_marker: bool, _op_type: ReplicationType) {
let mut q_cache = self.q_cache.lock().await; if let Ok(mut q_cache) = self.q_cache.lock() {
let stats = q_cache q_cache.dec(bucket, size);
.bucket_stats }
.entry(bucket.to_string()) }
.or_insert_with(InQueueMetric::default);
stats.curr.now_bytes.fetch_sub(size, Ordering::Relaxed);
stats.curr.now_count.fetch_sub(1, Ordering::Relaxed);
q_cache.sr_queue_stats.curr.now_bytes.fetch_sub(size, Ordering::Relaxed); pub(crate) fn inc_target_q(&self, bucket: &str, target_arns: &[String], size: i64) {
q_cache.sr_queue_stats.curr.now_count.fetch_sub(1, Ordering::Relaxed); let target_arns = normalized_target_arns(target_arns);
if target_arns.is_empty() {
return;
}
with_target_queue_caches(|caches| {
prune_stale_target_queue_caches(caches);
let slot_index = match caches.iter().position(|slot| slot.belongs_to(&self.q_cache)) {
Some(index) => index,
None => {
caches.push(TargetQueueCacheSlot::new(&self.q_cache));
caches.len() - 1
}
};
let slot = &mut caches[slot_index];
let bucket = bucket.to_string();
for target_arn in target_arns {
let metric = match slot.metrics.entry((bucket.clone(), target_arn.to_string())) {
Entry::Occupied(entry) => entry.into_mut(),
Entry::Vacant(entry) => entry.insert(InQueueMetric::default()),
};
metric.curr.add_current(size, 1);
}
});
}
pub(crate) fn dec_target_q(&self, bucket: &str, target_arns: &[String], size: i64) {
let target_arns = normalized_target_arns(target_arns);
if target_arns.is_empty() {
return;
}
with_target_queue_caches(|caches| {
prune_stale_target_queue_caches(caches);
if let Some(slot) = caches.iter_mut().find(|slot| slot.belongs_to(&self.q_cache)) {
let bucket = bucket.to_string();
for target_arn in target_arns {
if let Some(metric) = slot.metrics.get_mut(&(bucket.clone(), target_arn.to_string())) {
metric.curr.subtract_current(size, 1);
}
}
slot.metrics
.retain(|_, metric| metric.curr.get_current_count() > 0 || metric.curr.get_current_bytes() > 0);
}
});
}
pub fn runtime_target_backlog_snapshot(&self) -> Vec<RuntimeReplicationTargetBacklog> {
let mut snapshot = with_target_queue_caches(|caches| {
caches
.iter()
.find(|slot| slot.belongs_to(&self.q_cache))
.map(|slot| target_queue_cache_snapshot(&slot.metrics))
.unwrap_or_default()
});
snapshot.sort_by(|left, right| {
left.bucket
.cmp(&right.bucket)
.then_with(|| left.target_arn.cmp(&right.target_arn))
});
snapshot
} }
/// Increase proxy metrics /// Increase proxy metrics
@@ -715,6 +846,94 @@ impl Default for ReplicationStats {
mod tests { mod tests {
use super::*; use super::*;
#[test]
fn runtime_target_backlog_snapshot_tracks_targets() {
let stats = ReplicationStats::new();
stats.inc_target_q(
"photos",
&[
"arn:rustfs:replication:target-b".to_string(),
"arn:rustfs:replication:target-a".to_string(),
"arn:rustfs:replication:target-a".to_string(),
],
1024,
);
let snapshot = stats.runtime_target_backlog_snapshot();
assert_eq!(
snapshot,
vec![
RuntimeReplicationTargetBacklog {
bucket: "photos".to_string(),
target_arn: "arn:rustfs:replication:target-a".to_string(),
count: 1,
bytes: 1024,
},
RuntimeReplicationTargetBacklog {
bucket: "photos".to_string(),
target_arn: "arn:rustfs:replication:target-b".to_string(),
count: 1,
bytes: 1024,
},
]
);
}
#[test]
fn runtime_target_backlog_ignores_empty_targets() {
let stats = ReplicationStats::new();
stats.inc_target_q("photos", &["".to_string()], 1024);
assert!(stats.runtime_target_backlog_snapshot().is_empty());
}
#[test]
fn runtime_target_backlog_is_scoped_to_stats_instance() {
let first = ReplicationStats::new();
let second = ReplicationStats::new();
first.inc_target_q("photos", &["arn:rustfs:replication:target-a".to_string()], 1024);
assert!(second.runtime_target_backlog_snapshot().is_empty());
assert_eq!(first.runtime_target_backlog_snapshot()[0].count, 1);
}
#[test]
fn runtime_target_backlog_decrements_with_saturation() {
let stats = ReplicationStats::new();
let target_arns = ["arn:rustfs:replication:target-a".to_string()];
stats.inc_target_q("photos", &target_arns, 1024);
stats.dec_target_q("photos", &target_arns, 2048);
assert!(stats.runtime_target_backlog_snapshot().is_empty());
}
#[test]
fn runtime_target_backlog_prunes_stale_sidecar_on_next_access() {
{
let stats = ReplicationStats::new();
stats.inc_target_q("photos", &["arn:rustfs:replication:target-a".to_string()], 1024);
assert!(
TARGET_QUEUE_CACHES
.lock()
.expect("target queue cache mutex")
.iter()
.any(|slot| slot.owner.strong_count() > 0)
);
}
let stats = ReplicationStats::new();
stats.inc_target_q("photos", &["arn:rustfs:replication:target-b".to_string()], 1024);
assert!(
TARGET_QUEUE_CACHES
.lock()
.expect("target queue cache mutex")
.iter()
.all(|slot| slot.owner.strong_count() > 0)
);
}
#[tokio::test] #[tokio::test]
async fn test_replication_stats_new() { async fn test_replication_stats_new() {
let stats = ReplicationStats::new(); let stats = ReplicationStats::new();
@@ -801,14 +1020,14 @@ mod tests {
async fn latest_stats_include_queue_until_drained() { async fn latest_stats_include_queue_until_drained() {
let stats = ReplicationStats::new(); let stats = ReplicationStats::new();
stats.inc_q("queued-bucket", 4096, false, ReplicationType::Object).await; stats.inc_q("queued-bucket", 4096, false, ReplicationType::Object);
let queued = stats.get_latest_replication_stats("queued-bucket").await; let queued = stats.get_latest_replication_stats("queued-bucket").await;
assert!(queued.replication_stats.provider_available); assert!(queued.replication_stats.provider_available);
assert_eq!(queued.replication_stats.q_stat.curr.count, 1); assert_eq!(queued.replication_stats.q_stat.curr.count, 1);
assert_eq!(queued.replication_stats.q_stat.curr.bytes, 4096); assert_eq!(queued.replication_stats.q_stat.curr.bytes, 4096);
assert_eq!(queued.replication_stats.queue_scope, ReplicationMetricScope::NodeLocal); assert_eq!(queued.replication_stats.queue_scope, ReplicationMetricScope::NodeLocal);
stats.dec_q("queued-bucket", 4096, false, ReplicationType::Object).await; stats.dec_q("queued-bucket", 4096, false, ReplicationType::Object);
let drained = stats.get_latest_replication_stats("queued-bucket").await; let drained = stats.get_latest_replication_stats("queued-bucket").await;
assert_eq!(drained.replication_stats.q_stat.curr.count, 0); assert_eq!(drained.replication_stats.q_stat.curr.count, 0);
assert_eq!(drained.replication_stats.q_stat.curr.bytes, 0); assert_eq!(drained.replication_stats.q_stat.curr.bytes, 0);
@@ -911,7 +1130,7 @@ mod tests {
for _ in 0..32 { for _ in 0..32 {
let stats = Arc::clone(&stats); let stats = Arc::clone(&stats);
tasks.push(tokio::spawn(async move { tasks.push(tokio::spawn(async move {
stats.inc_q("concurrent-bucket", 7, false, ReplicationType::Object).await; stats.inc_q("concurrent-bucket", 7, false, ReplicationType::Object);
})); }));
} }
for task in tasks { for task in tasks {
+1 -1
View File
@@ -44,7 +44,7 @@ pub(crate) use internode_data_transport::TcpHttpInternodeDataTransport;
pub use internode_data_transport::build_internode_data_transport_from_env; pub use internode_data_transport::build_internode_data_transport_from_env;
pub(crate) use peer_rest_client::TierConfigReloadOutcome; pub(crate) use peer_rest_client::TierConfigReloadOutcome;
pub use peer_rest_client::{ pub use peer_rest_client::{
PEER_RESTDRY_RUN, PEER_RESTSIGNAL, PEER_RESTSUB_SYS, PeerRestClient, SERVICE_SIGNAL_REFRESH_CONFIG, KMS_SIGNAL_SUBSYSTEM, PEER_RESTDRY_RUN, PEER_RESTSIGNAL, PEER_RESTSUB_SYS, PeerRestClient, SERVICE_SIGNAL_REFRESH_CONFIG,
SERVICE_SIGNAL_RELOAD_DYNAMIC, ScannerPeerActivity, SERVICE_SIGNAL_RELOAD_DYNAMIC, ScannerPeerActivity,
}; };
pub(crate) use peer_s3_client::heal_bucket_local_on_disks; pub(crate) use peer_s3_client::heal_bucket_local_on_disks;
@@ -45,9 +45,9 @@ use rustfs_protos::proto_gen::node_service::{
HealControlRequest, LoadBucketMetadataRequest, LoadGroupRequest, LoadPolicyMappingRequest, LoadPolicyRequest, HealControlRequest, LoadBucketMetadataRequest, LoadGroupRequest, LoadPolicyMappingRequest, LoadPolicyRequest,
LoadRebalanceMetaRequest, LoadServiceAccountRequest, LoadTransitionTierConfigRequest, LoadUserRequest, LoadRebalanceMetaRequest, LoadServiceAccountRequest, LoadTransitionTierConfigRequest, LoadUserRequest,
LocalStorageInfoRequest, Mss, ReloadPoolMetaRequest, ReloadSiteReplicationConfigRequest, ScannerActivityRequest, LocalStorageInfoRequest, Mss, ReloadPoolMetaRequest, ReloadSiteReplicationConfigRequest, ScannerActivityRequest,
ScannerActivityResponse, ServerInfoRequest, SignalServiceRequest, StartDecommissionRequest, StartProfilingRequest, ScannerActivityResponse, ServerInfoRequest, SignalServiceRequest, SignalServiceResponse, StartDecommissionRequest,
StopRebalanceRequest, TierMutationAbortRequest, TierMutationCommitRequest, TierMutationControlResponse, StartProfilingRequest, StopRebalanceRequest, TierMutationAbortRequest, TierMutationCommitRequest,
TierMutationPeerState, TierMutationPrepareRequest, node_service_client::NodeServiceClient, TierMutationControlResponse, TierMutationPeerState, TierMutationPrepareRequest, node_service_client::NodeServiceClient,
tier_mutation_control_service_client::TierMutationControlServiceClient, tier_mutation_control_service_client::TierMutationControlServiceClient,
}; };
pub use rustfs_protos::{PEER_RESTDRY_RUN, PEER_RESTSIGNAL, PEER_RESTSUB_SYS}; pub use rustfs_protos::{PEER_RESTDRY_RUN, PEER_RESTSIGNAL, PEER_RESTSUB_SYS};
@@ -71,6 +71,12 @@ use uuid::Uuid;
pub const SERVICE_SIGNAL_REFRESH_CONFIG: u64 = 1; pub const SERVICE_SIGNAL_REFRESH_CONFIG: u64 = 1;
pub const SERVICE_SIGNAL_RELOAD_DYNAMIC: u64 = 2; pub const SERVICE_SIGNAL_RELOAD_DYNAMIC: u64 = 2;
/// Dynamic config subsystem for the cluster-persisted KMS configuration.
///
/// KMS configuration lives in its own cluster object rather than in the server
/// config document, so it is not a `ServerConfig` subsystem; it only shares the
/// reload signal transport.
pub const KMS_SIGNAL_SUBSYSTEM: &str = "kms";
const BACKGROUND_HEAL_STATUS_MAX_MESSAGE_SIZE: usize = 64 * 1024; const BACKGROUND_HEAL_STATUS_MAX_MESSAGE_SIZE: usize = 64 * 1024;
const HEAL_CONTROL_FINGERPRINT_MAX_SIZE: usize = 256; const HEAL_CONTROL_FINGERPRINT_MAX_SIZE: usize = 256;
const HEAL_CONTROL_PAYLOAD_MAX_SIZE: usize = 64 * 1024; const HEAL_CONTROL_PAYLOAD_MAX_SIZE: usize = 64 * 1024;
@@ -99,8 +105,13 @@ fn decode_bucket_stats_response(response: GetBucketStatsDataResponse) -> Result<
} }
fn validate_signal_service_protocol(sig: u64, sub_sys: &str, protocol_version: u32) -> Result<()> { fn validate_signal_service_protocol(sig: u64, sub_sys: &str, protocol_version: u32) -> Result<()> {
// The version stays pinned to DYNAMIC_CONFIG_PROTOCOL_VERSION rather than
// being bumped per subsystem: the comparison is shared, so raising it would
// retire peers that already converge scanner and heal config correctly.
// Subsystems added after a peer was built are rejected by that peer's own
// subsystem allow-list, which surfaces as an explicit failed signal.
if sig == SERVICE_SIGNAL_RELOAD_DYNAMIC if sig == SERVICE_SIGNAL_RELOAD_DYNAMIC
&& matches!(sub_sys, SCANNER_SUB_SYS | HEAL_SUB_SYS) && matches!(sub_sys, SCANNER_SUB_SYS | HEAL_SUB_SYS | KMS_SIGNAL_SUBSYSTEM)
&& protocol_version < rustfs_protos::DYNAMIC_CONFIG_PROTOCOL_VERSION && protocol_version < rustfs_protos::DYNAMIC_CONFIG_PROTOCOL_VERSION
{ {
return Err(Error::other(format!("peer does not support dynamic {sub_sys} config convergence"))); return Err(Error::other(format!("peer does not support dynamic {sub_sys} config convergence")));
@@ -1500,6 +1511,22 @@ impl PeerRestClient {
} }
pub async fn signal_service(&self, sig: u64, sub_sys: &str, dry_run: bool, _exec_at: SystemTime) -> Result<()> { pub async fn signal_service(&self, sig: u64, sub_sys: &str, dry_run: bool, _exec_at: SystemTime) -> Result<()> {
self.signal_service_checked(sig, sub_sys, dry_run).await.map(|_| ())
}
/// Report the KMS configuration fingerprint the peer is currently running.
///
/// Sent as a dry-run reload signal so the peer answers without swapping its
/// own configuration. `None` means the peer has no KMS configuration. The
/// fingerprint is advisory and feeds cluster status reporting only, so the
/// response is not proof-signed.
pub async fn kms_config_fingerprint(&self) -> Result<Option<String>> {
self.signal_service_checked(SERVICE_SIGNAL_RELOAD_DYNAMIC, KMS_SIGNAL_SUBSYSTEM, true)
.await
.map(|response| response.config_fingerprint)
}
async fn signal_service_checked(&self, sig: u64, sub_sys: &str, dry_run: bool) -> Result<SignalServiceResponse> {
self.finalize_result( self.finalize_result(
async { async {
let mut client = self.get_client().await?; let mut client = self.get_client().await?;
@@ -1520,7 +1547,7 @@ impl PeerRestClient {
return Err(Error::other("")); return Err(Error::other(""));
} }
validate_signal_service_protocol(sig, sub_sys, response.protocol_version)?; validate_signal_service_protocol(sig, sub_sys, response.protocol_version)?;
Ok(()) Ok(response)
} }
.await, .await,
) )
@@ -2347,6 +2374,19 @@ mod tests {
.expect("full refresh compatibility is guarded by its scanner preflight"); .expect("full refresh compatibility is guarded by its scanner preflight");
} }
#[test]
fn dynamic_kms_config_requires_versioned_peer_acknowledgement() {
let err = validate_signal_service_protocol(SERVICE_SIGNAL_RELOAD_DYNAMIC, KMS_SIGNAL_SUBSYSTEM, 0)
.expect_err("an unversioned peer must not claim KMS config convergence");
assert!(err.to_string().contains("does not support dynamic"));
validate_signal_service_protocol(
SERVICE_SIGNAL_RELOAD_DYNAMIC,
KMS_SIGNAL_SUBSYSTEM,
rustfs_protos::DYNAMIC_CONFIG_PROTOCOL_VERSION,
)
.expect("a current peer should support dynamic KMS config");
}
#[test] #[test]
fn peer_rest_client_marks_network_like_errors() { fn peer_rest_client_marks_network_like_errors() {
assert!(PeerRestClient::is_network_like_error(&Error::other("transport error"))); assert!(PeerRestClient::is_network_like_error(&Error::other("transport error")));
+643 -54
View File
@@ -53,12 +53,14 @@ use std::sync::LazyLock;
use std::sync::{Arc, RwLock}; use std::sync::{Arc, RwLock};
use tokio::sync::{OwnedRwLockWriteGuard, RwLock as AsyncRwLock}; use tokio::sync::{OwnedRwLockWriteGuard, RwLock as AsyncRwLock};
use tracing::{debug, error, info, instrument, warn}; use tracing::{debug, error, info, instrument, warn};
use uuid::Uuid;
pub const CONFIG_PREFIX: &str = "config"; pub const CONFIG_PREFIX: &str = "config";
const SERVER_CONFIG_OBJECT: &str = "config/config.json"; const SERVER_CONFIG_OBJECT: &str = "config/config.json";
const CONFIG_TRANSACTION_LOCK_SUFFIX: &str = ".transaction.lock";
// Server-config lock order: SERVER_CONFIG_LOCK -> distributed namespace lock // Server-config lock order: SERVER_CONFIG_LOCK -> transaction lock ->
// for SERVER_CONFIG_OBJECT. Readers and writers must never reverse this order. // SERVER_CONFIG_OBJECT. Readers and writers must never reverse this order.
static SERVER_CONFIG_LOCK: LazyLock<Arc<AsyncRwLock<()>>> = LazyLock::new(|| Arc::new(AsyncRwLock::new(()))); static SERVER_CONFIG_LOCK: LazyLock<Arc<AsyncRwLock<()>>> = LazyLock::new(|| Arc::new(AsyncRwLock::new(())));
fn config_task_join_error(operation: &'static str, error: tokio::task::JoinError) -> Error { fn config_task_join_error(operation: &'static str, error: tokio::task::JoinError) -> Error {
@@ -76,8 +78,11 @@ where
T: Send + 'static, T: Send + 'static,
{ {
tokio::spawn(async move { tokio::spawn(async move {
// Lock order: SERVER_CONFIG_LOCK -> namespace write lock. // Lock order: SERVER_CONFIG_LOCK -> transaction lock -> object lock.
let _local_guard = SERVER_CONFIG_LOCK.write().await; let _local_guard = SERVER_CONFIG_LOCK.write().await;
let transaction_lock = server_config_transaction_lock_path();
let transaction_lock = store.new_ns_lock(RUSTFS_META_BUCKET, &transaction_lock).await?;
let _transaction_guard = transaction_lock.get_write_lock(get_lock_acquire_timeout()).await?;
let namespace_lock = store.new_ns_lock(RUSTFS_META_BUCKET, SERVER_CONFIG_OBJECT).await?; let namespace_lock = store.new_ns_lock(RUSTFS_META_BUCKET, SERVER_CONFIG_OBJECT).await?;
let _write_guard = namespace_lock.get_write_lock(get_lock_acquire_timeout()).await?; let _write_guard = namespace_lock.get_write_lock(get_lock_acquire_timeout()).await?;
Ok(operation().await) Ok(operation().await)
@@ -96,8 +101,11 @@ where
T: Send + 'static, T: Send + 'static,
{ {
tokio::spawn(async move { tokio::spawn(async move {
// Lock order: SERVER_CONFIG_LOCK -> namespace read lock. // Lock order: SERVER_CONFIG_LOCK -> transaction lock -> object lock.
let _local_guard = SERVER_CONFIG_LOCK.read().await; let _local_guard = SERVER_CONFIG_LOCK.read().await;
let transaction_lock = server_config_transaction_lock_path();
let transaction_lock = store.new_ns_lock(RUSTFS_META_BUCKET, &transaction_lock).await?;
let _transaction_guard = transaction_lock.get_read_lock(get_lock_acquire_timeout()).await?;
let namespace_lock = store.new_ns_lock(RUSTFS_META_BUCKET, SERVER_CONFIG_OBJECT).await?; let namespace_lock = store.new_ns_lock(RUSTFS_META_BUCKET, SERVER_CONFIG_OBJECT).await?;
let _read_guard = namespace_lock.get_read_lock(get_lock_acquire_timeout()).await?; let _read_guard = namespace_lock.get_read_lock(get_lock_acquire_timeout()).await?;
Ok(operation().await) Ok(operation().await)
@@ -578,12 +586,29 @@ where
PutObjectReader = PutObjReader, PutObjectReader = PutObjReader,
>, >,
{ {
let mut put_data = PutObjReader::from_vec(data); save_config_with_opts_and_metadata(api, file, data, opts).await.map(|_| ())
if let Err(err) = api.put_object(RUSTFS_META_BUCKET, file, &mut put_data, opts).await { }
error!("save_config_with_opts: err: {:?}, file: {}", err, file);
return Err(err); async fn save_config_with_opts_and_metadata<S>(api: Arc<S>, file: &str, data: Vec<u8>, opts: &ObjectOptions) -> Result<ObjectInfo>
where
S: ObjectIO<
Error = Error,
RangeSpec = HTTPRangeSpec,
HeaderMap = HeaderMap,
ObjectOptions = ObjectOptions,
ObjectInfo = ObjectInfo,
GetObjectReader = GetObjectReader,
PutObjectReader = PutObjReader,
>,
{
let mut put_data = PutObjReader::from_vec(data);
match api.put_object(RUSTFS_META_BUCKET, file, &mut put_data, opts).await {
Ok(object_info) => Ok(object_info),
Err(err) => {
error!("save_config_with_opts: err: {:?}, file: {}", err, file);
Err(err)
}
} }
Ok(())
} }
fn new_server_config() -> Config { fn new_server_config() -> Config {
@@ -594,8 +619,12 @@ async fn new_and_save_server_config<S>(api: Arc<S>) -> Result<Config>
where where
S: EcstoreObjectIO + StorageAdminApi + NamespaceLocking<Error = Error, NamespaceLock = rustfs_lock::NamespaceLockWrapper>, S: EcstoreObjectIO + StorageAdminApi + NamespaceLocking<Error = Error, NamespaceLock = rustfs_lock::NamespaceLockWrapper>,
{ {
let snapshot = read_server_config_snapshot(api.clone()).await?;
if snapshot.object_exists() {
return Ok(snapshot.config.clone());
}
let cfg = new_server_config(); let cfg = new_server_config();
save_server_config(api, &cfg).await?; save_server_config_snapshot(api, &cfg, &snapshot).await?;
Ok(cfg) Ok(cfg)
} }
@@ -617,6 +646,10 @@ pub fn server_config_path() -> String {
SERVER_CONFIG_OBJECT.to_string() SERVER_CONFIG_OBJECT.to_string()
} }
fn server_config_transaction_lock_path() -> String {
format!("{}{CONFIG_TRANSACTION_LOCK_SUFFIX}", server_config_path())
}
fn storage_class_kvs_mut(cfg: &mut Config) -> &mut KVS { fn storage_class_kvs_mut(cfg: &mut Config) -> &mut KVS {
let sub_cfg = cfg.0.entry(STORAGE_CLASS_SUB_SYS.to_string()).or_insert_with(|| { let sub_cfg = cfg.0.entry(STORAGE_CLASS_SUB_SYS.to_string()).or_insert_with(|| {
let mut section = HashMap::new(); let mut section = HashMap::new();
@@ -819,6 +852,9 @@ fn apply_external_scalar_config_map(
let Some(config_value) = root.get(descriptor.subsystem_key) else { let Some(config_value) = root.get(descriptor.subsystem_key) else {
return Ok(false); return Ok(false);
}; };
if descriptor.subsystem_key == HEAL_SUB_SYS && config_value.is_null() {
return Ok(false);
}
let overrides = decode_scalar_config_value(config_value, descriptor)?; let overrides = decode_scalar_config_value(config_value, descriptor)?;
if overrides.is_empty() { if overrides.is_empty() {
@@ -1463,21 +1499,142 @@ fn build_audit_object(cfg: &Config) -> Map<String, Value> {
build_target_object(cfg, &audit_target_descriptors()) build_target_object(cfg, &audit_target_descriptors())
} }
fn sync_rendered_target_instance(existing: Value, rendered: Option<&Value>, valid_keys: &[&str]) -> Option<Value> {
match existing {
Value::Object(mut instance) => {
for key in valid_keys {
instance.remove(*key);
}
if let Some(Value::Object(rendered)) = rendered {
instance.extend(rendered.clone());
}
(!instance.is_empty()).then_some(Value::Object(instance))
}
Value::Array(entries) => {
let mut pending = rendered
.and_then(Value::as_object)
.map(|rendered| {
rendered
.iter()
.filter_map(|(key, value)| parse_target_scalar_value(key, value).map(|value| (key.clone(), value)))
.collect::<HashMap<_, _>>()
})
.unwrap_or_default();
let mut updated = Vec::with_capacity(entries.len().saturating_add(pending.len()));
for entry in entries {
let Some(entry_obj) = entry.as_object() else {
updated.push(entry);
continue;
};
let Some(key) = entry_obj.get("key").and_then(Value::as_str) else {
updated.push(entry);
continue;
};
if !valid_keys.contains(&key) {
updated.push(entry);
continue;
}
let Some(value) = pending.remove(key) else {
continue;
};
let mut entry_obj = entry_obj.clone();
entry_obj.insert("value".to_string(), Value::String(value));
updated.push(Value::Object(entry_obj));
}
updated.extend(rendered_scalar_config_kvs_entries(&pending));
(!updated.is_empty()).then_some(Value::Array(updated))
}
value if rendered.is_none() => Some(value),
_ => rendered.cloned(),
}
}
fn sync_rendered_target_object( fn sync_rendered_target_object(
target_obj: &mut Map<String, Value>, target_obj: &mut Map<String, Value>,
rendered_target: &Map<String, Value>, rendered_target: &Map<String, Value>,
descriptors: &[TargetConfigDescriptor], descriptors: &[TargetConfigDescriptor],
) { ) {
for descriptor in descriptors { for descriptor in descriptors {
match rendered_target.get(descriptor.external_key) { let existing = target_obj.remove(descriptor.external_key);
Some(Value::Object(v)) => { let alias = target_obj.remove(descriptor.subsystem_key);
target_obj.insert(descriptor.external_key.to_string(), Value::Object(v.clone())); let mut section = existing
target_obj.remove(descriptor.subsystem_key); .or(alias)
.and_then(|value| value.as_object().cloned())
.unwrap_or_default();
let rendered = rendered_target.get(descriptor.external_key).and_then(Value::as_object);
if is_target_instance_shorthand(&section, descriptor.valid_keys) {
let has_named_instances = rendered.is_some_and(|instances| instances.keys().any(|name| name != "default"));
if !has_named_instances {
if let Some(section) = sync_rendered_target_instance(
Value::Object(section),
rendered.and_then(|instances| instances.get("default")),
descriptor.valid_keys,
) {
target_obj.insert(descriptor.external_key.to_string(), section);
} }
_ => { continue;
target_obj.remove(descriptor.external_key);
target_obj.remove(descriptor.subsystem_key);
} }
let mut nested = Map::new();
if let Some(default) = sync_rendered_target_instance(
Value::Object(section),
rendered.and_then(|instances| instances.get("default")),
descriptor.valid_keys,
) {
nested.insert("default".to_string(), default);
}
if let Some(rendered) = rendered {
for (instance_name, instance) in rendered {
if instance_name != "default" {
nested.insert(instance_name.clone(), instance.clone());
}
}
}
if !nested.is_empty() {
target_obj.insert(descriptor.external_key.to_string(), Value::Object(nested));
}
continue;
}
if let Some(default_alias) = section.remove(DEFAULT_DELIMITER) {
if let Some(default) = section.get_mut("default") {
if let Some(alias) = sync_rendered_target_instance(default_alias, None, descriptor.valid_keys) {
match (default, alias) {
(Value::Object(default), Value::Object(alias)) => {
for (key, value) in alias {
default.entry(key).or_insert(value);
}
}
(Value::Array(default), Value::Array(alias)) => default.extend(alias),
_ => {}
}
}
} else {
section.insert("default".to_string(), default_alias);
}
}
let mut merged = Map::new();
for (instance_name, instance) in section {
if let Some(instance) = sync_rendered_target_instance(
instance,
rendered.and_then(|instances| instances.get(&instance_name)),
descriptor.valid_keys,
) {
merged.insert(instance_name, instance);
}
}
if let Some(rendered) = rendered {
for (instance_name, instance) in rendered {
if !merged.contains_key(instance_name) {
merged.insert(instance_name.clone(), instance.clone());
}
}
}
if !merged.is_empty() {
target_obj.insert(descriptor.external_key.to_string(), Value::Object(merged));
} }
} }
} }
@@ -1496,6 +1653,14 @@ fn encode_server_config_blob(cfg: &Config, seed: Option<&[u8]>) -> Result<Vec<u8
Some(Value::Object(v)) => v, Some(Value::Object(v)) => v,
_ => Map::new(), _ => Map::new(),
}; };
for key in [
storageclass::CLASS_STANDARD,
storageclass::CLASS_RRS,
storageclass::OPTIMIZE,
storageclass::INLINE_BLOCK,
] {
sc_obj.remove(key);
}
for (k, v) in build_storageclass_object(cfg) { for (k, v) in build_storageclass_object(cfg) {
sc_obj.insert(k, v); sc_obj.insert(k, v);
} }
@@ -1503,7 +1668,10 @@ fn encode_server_config_blob(cfg: &Config, seed: Option<&[u8]>) -> Result<Vec<u8
root.remove("storage_class"); root.remove("storage_class");
for descriptor in [scanner_config_descriptor(), heal_config_descriptor()] { for descriptor in [scanner_config_descriptor(), heal_config_descriptor()] {
let existing = root.remove(descriptor.subsystem_key); let mut existing = root.remove(descriptor.subsystem_key);
if descriptor.subsystem_key == HEAL_SUB_SYS && existing.as_ref().is_some_and(Value::is_null) {
existing = None;
}
let rendered = build_scalar_config_object(cfg, descriptor); let rendered = build_scalar_config_object(cfg, descriptor);
if let Some(config_value) = sync_rendered_scalar_config_value(existing, &rendered, descriptor)? { if let Some(config_value) = sync_rendered_scalar_config_value(existing, &rendered, descriptor)? {
root.insert(descriptor.subsystem_key.to_string(), config_value); root.insert(descriptor.subsystem_key.to_string(), config_value);
@@ -1560,6 +1728,7 @@ fn is_standard_object_server_config(data: &[u8]) -> bool {
matches!(root.get("version"), Some(Value::String(v)) if !v.trim().is_empty()) matches!(root.get("version"), Some(Value::String(v)) if !v.trim().is_empty())
&& matches!(root.get("storageclass"), Some(Value::Object(_))) && matches!(root.get("storageclass"), Some(Value::Object(_)))
&& !root.contains_key("storage_class") && !root.contains_key("storage_class")
&& !matches!(root.get(HEAL_SUB_SYS), Some(Value::Null))
} }
fn configs_semantically_equal(lhs: &Config, rhs: &Config) -> bool { fn configs_semantically_equal(lhs: &Config, rhs: &Config) -> bool {
@@ -1593,7 +1762,7 @@ where
FileInfo = FileInfo, FileInfo = FileInfo,
ObjectToDelete = ObjectToDelete, ObjectToDelete = ObjectToDelete,
DeletedObject = DeletedObject, DeletedObject = DeletedObject,
>, > + NamespaceLocking<Error = Error, NamespaceLock = rustfs_lock::NamespaceLockWrapper>,
{ {
if let Some(decrypt) = &decrypt_fn { if let Some(decrypt) = &decrypt_fn {
register_server_config_decrypt_fn(decrypt.clone()); register_server_config_decrypt_fn(decrypt.clone());
@@ -1601,14 +1770,7 @@ where
let config_file = server_config_path(); let config_file = server_config_path();
match api match api
.get_object_info( .get_object_info(RUSTFS_META_BUCKET, &config_file, &ObjectOptions::default())
RUSTFS_META_BUCKET,
&config_file,
&ObjectOptions {
no_lock: true,
..Default::default()
},
)
.await .await
{ {
Ok(_) => { Ok(_) => {
@@ -1624,7 +1786,6 @@ where
let opts = ObjectOptions { let opts = ObjectOptions {
max_parity: true, max_parity: true,
no_lock: true,
..Default::default() ..Default::default()
}; };
@@ -1677,7 +1838,33 @@ where
} }
}; };
match save_config(api, &config_file, normalized).await { let snapshot = match read_server_config_snapshot(api.clone()).await {
Ok(snapshot) => snapshot,
Err(err) => {
warn!("recheck target server config failed, skip migration: {:?}", err);
return;
}
};
if snapshot.object_exists() {
debug!("server config was created while legacy migration was preparing, skip migration");
return;
}
match save_config_with_opts(
api,
&config_file,
normalized,
&ObjectOptions {
max_parity: true,
http_preconditions: Some(HTTPPreconditions {
if_none_match: Some("*".to_string()),
..Default::default()
}),
..Default::default()
},
)
.await
{
Ok(()) => { Ok(()) => {
info!("Migrated compatible server config from legacy metadata bucket"); info!("Migrated compatible server config from legacy metadata bucket");
} }
@@ -1769,8 +1956,13 @@ where
{ {
let config_file = server_config_path(); let config_file = server_config_path();
// Try to read the configuration file // Try to read the configuration file.
match read_config_no_lock(api.clone(), &config_file).await { let data = if namespace_lock_held {
read_config_no_lock(api.clone(), &config_file).await
} else {
read_config(api.clone(), &config_file).await
};
match data {
Ok(data) => read_server_config(api, &data, namespace_lock_held).await, Ok(data) => read_server_config(api, &data, namespace_lock_held).await,
Err(Error::ConfigNotFound) => handle_missing_config(api, "Read the main configuration", namespace_lock_held).await, Err(Error::ConfigNotFound) => handle_missing_config(api, "Read the main configuration", namespace_lock_held).await,
Err(err) => handle_config_read_error(err, &config_file), Err(err) => handle_config_read_error(err, &config_file),
@@ -1787,7 +1979,12 @@ where
warn!("Received empty configuration data, try to reread from '{}'", config_file); warn!("Received empty configuration data, try to reread from '{}'", config_file);
// Try to read the configuration again // Try to read the configuration again
match read_config_no_lock(api.clone(), &config_file).await { let data = if namespace_lock_held {
read_config_no_lock(api.clone(), &config_file).await
} else {
read_config(api.clone(), &config_file).await
};
match data {
Ok(cfg_data) => { Ok(cfg_data) => {
let cfg = decode_persisted_server_config(&cfg_data)?; let cfg = decode_persisted_server_config(&cfg_data)?;
return Ok(cfg.merge()); return Ok(cfg.merge());
@@ -2036,11 +2233,16 @@ pub struct ServerConfigSnapshot {
raw: Option<Vec<u8>>, raw: Option<Vec<u8>>,
seed: Option<Vec<u8>>, seed: Option<Vec<u8>>,
etag: Option<String>, etag: Option<String>,
generation: Option<Uuid>,
_local_guard: OwnedRwLockWriteGuard<()>, _local_guard: OwnedRwLockWriteGuard<()>,
_guard: rustfs_lock::NamespaceLockGuard, _guard: rustfs_lock::NamespaceLockGuard,
} }
impl ServerConfigSnapshot { impl ServerConfigSnapshot {
pub fn object_exists(&self) -> bool {
self.raw.is_some()
}
pub fn ensure_lock_held(&self) -> Result<()> { pub fn ensure_lock_held(&self) -> Result<()> {
if self._guard.is_lock_lost() { if self._guard.is_lock_lost() {
return Err(Error::other("server config transaction lock was lost")); return Err(Error::other("server config transaction lock was lost"));
@@ -2051,12 +2253,34 @@ impl ServerConfigSnapshot {
pub fn is_lock_lost(&self) -> bool { pub fn is_lock_lost(&self) -> bool {
self._guard.is_lock_lost() self._guard.is_lock_lost()
} }
pub fn generation(&self) -> Option<Uuid> {
self.generation
}
} }
/// Read a server config transaction snapshot while holding the same local and #[derive(Debug, Clone, PartialEq, Eq)]
/// distributed write locks used by every other server-config writer. Internal pub struct ServerConfigSaveResult {
/// reads and the later conditional write use no-lock object I/O; the guards persisted: bool,
/// remain live until the snapshot is dropped. generation: Option<Uuid>,
}
impl ServerConfigSaveResult {
pub fn persisted(&self) -> bool {
self.persisted
}
pub fn generation(&self) -> Option<Uuid> {
self.generation
}
}
/// Read a server config transaction snapshot while holding a dedicated
/// transaction lock. The config object's normal namespace lock remains
/// available to fence reads and the conditional write at commit time.
/// The transaction guard remains live until the snapshot is dropped,
/// serializing persistence and history ordering across admin nodes. Runtime
/// state is reloaded from the durable object after this guard is released.
pub async fn read_server_config_snapshot<S>(api: Arc<S>) -> Result<ServerConfigSnapshot> pub async fn read_server_config_snapshot<S>(api: Arc<S>) -> Result<ServerConfigSnapshot>
where where
S: ObjectIO< S: ObjectIO<
@@ -2071,12 +2295,10 @@ where
{ {
let config_file = server_config_path(); let config_file = server_config_path();
let local_guard = SERVER_CONFIG_LOCK.clone().write_owned().await; let local_guard = SERVER_CONFIG_LOCK.clone().write_owned().await;
let lock = api.new_ns_lock(RUSTFS_META_BUCKET, &config_file).await?; let transaction_lock = server_config_transaction_lock_path();
let lock = api.new_ns_lock(RUSTFS_META_BUCKET, &transaction_lock).await?;
let guard = lock.get_write_lock(get_lock_acquire_timeout()).await?; let guard = lock.get_write_lock(get_lock_acquire_timeout()).await?;
let read_options = ObjectOptions { let read_options = ObjectOptions::default();
no_lock: true,
..Default::default()
};
match read_config_with_metadata_inner(api, &config_file, &read_options, true).await { match read_config_with_metadata_inner(api, &config_file, &read_options, true).await {
Ok((raw, object_info)) => { Ok((raw, object_info)) => {
let (config, seed) = decode_persisted_server_config_with_seed(&raw)?; let (config, seed) = decode_persisted_server_config_with_seed(&raw)?;
@@ -2085,6 +2307,7 @@ where
raw: Some(raw), raw: Some(raw),
seed: Some(seed), seed: Some(seed),
etag: object_info.etag, etag: object_info.etag,
generation: object_info.data_dir.filter(|generation| !generation.is_nil()),
_local_guard: local_guard, _local_guard: local_guard,
_guard: guard, _guard: guard,
}) })
@@ -2094,6 +2317,7 @@ where
raw: None, raw: None,
seed: None, seed: None,
etag: None, etag: None,
generation: None,
_local_guard: local_guard, _local_guard: local_guard,
_guard: guard, _guard: guard,
}), }),
@@ -2108,6 +2332,27 @@ where
/// lock, so a concurrent update or transaction lease loss cannot commit an /// lock, so a concurrent update or transaction lease loss cannot commit an
/// unfenced overwrite. /// unfenced overwrite.
pub async fn save_server_config_snapshot<S>(api: Arc<S>, cfg: &Config, snapshot: &ServerConfigSnapshot) -> Result<bool> pub async fn save_server_config_snapshot<S>(api: Arc<S>, cfg: &Config, snapshot: &ServerConfigSnapshot) -> Result<bool>
where
S: ObjectIO<
Error = Error,
RangeSpec = HTTPRangeSpec,
HeaderMap = HeaderMap,
ObjectOptions = ObjectOptions,
ObjectInfo = ObjectInfo,
GetObjectReader = GetObjectReader,
PutObjectReader = PutObjReader,
> + NamespaceLocking<Error = Error, NamespaceLock = rustfs_lock::NamespaceLockWrapper>,
{
save_server_config_snapshot_with_generation(api, cfg, snapshot)
.await
.map(|result| result.persisted())
}
pub async fn save_server_config_snapshot_with_generation<S>(
api: Arc<S>,
cfg: &Config,
snapshot: &ServerConfigSnapshot,
) -> Result<ServerConfigSaveResult>
where where
S: ObjectIO< S: ObjectIO<
Error = Error, Error = Error,
@@ -2126,13 +2371,19 @@ where
&& configs_semantically_equal(&snapshot.config, cfg) && configs_semantically_equal(&snapshot.config, cfg)
{ {
debug!("server config unchanged and already in standard object shape, skip write"); debug!("server config unchanged and already in standard object shape, skip write");
return Ok(false); return Ok(ServerConfigSaveResult {
persisted: false,
generation: snapshot.generation(),
});
} }
let data = encode_server_config_blob(cfg, snapshot.seed.as_deref())?; let data = encode_server_config_blob(cfg, snapshot.seed.as_deref())?;
if snapshot.raw.as_deref().is_some_and(|current| current == data.as_slice()) { if snapshot.raw.as_deref().is_some_and(|current| current == data.as_slice()) {
debug!("server config bytes unchanged after encode, skip write"); debug!("server config bytes unchanged after encode, skip write");
return Ok(false); return Ok(ServerConfigSaveResult {
persisted: false,
generation: snapshot.generation(),
});
} }
let http_preconditions = if snapshot.raw.is_some() { let http_preconditions = if snapshot.raw.is_some() {
@@ -2152,19 +2403,22 @@ where
} }
}; };
save_config_with_opts( snapshot.ensure_lock_held()?;
let object_info = save_config_with_opts_and_metadata(
api, api,
&config_file, &config_file,
data, data,
&ObjectOptions { &ObjectOptions {
max_parity: true, max_parity: true,
no_lock: true,
http_preconditions: Some(http_preconditions), http_preconditions: Some(http_preconditions),
..Default::default() ..Default::default()
}, },
) )
.await?; .await?;
Ok(true) Ok(ServerConfigSaveResult {
persisted: true,
generation: object_info.data_dir.filter(|generation| !generation.is_nil()),
})
} }
/// Saves the server config while an upper layer holds the namespace write /// Saves the server config while an upper layer holds the namespace write
@@ -2301,8 +2555,9 @@ mod tests {
use super::{ use super::{
SERVER_CONFIG_LOCK, ServerConfigSnapshot, apply_dynamic_config_for_sub_sys_with, config_task_join_error, SERVER_CONFIG_LOCK, ServerConfigSnapshot, apply_dynamic_config_for_sub_sys_with, config_task_join_error,
configs_semantically_equal, decode_server_config_blob, encode_server_config_blob, is_standard_object_server_config, configs_semantically_equal, decode_server_config_blob, encode_server_config_blob, is_standard_object_server_config,
lookup_configs, read_config, read_config_preserve_empty, read_config_with_metadata, read_config_without_migrate, lookup_configs, new_and_save_server_config, read_config, read_config_preserve_empty, read_config_with_metadata,
read_server_config_snapshot, save_server_config, save_server_config_snapshot, server_config_path, storage_class_kvs_mut, read_config_without_migrate, read_server_config_snapshot, save_server_config, save_server_config_snapshot,
save_server_config_snapshot_with_generation, server_config_transaction_lock_path, storage_class_kvs_mut,
}; };
use crate::config::{audit, heal, notify, oidc, scanner}; use crate::config::{audit, heal, notify, oidc, scanner};
use crate::disk::endpoint::Endpoint; use crate::disk::endpoint::Endpoint;
@@ -2311,7 +2566,9 @@ mod tests {
use crate::object_api::{GetObjectReader, ObjectInfo, ObjectOptions, PutObjReader}; use crate::object_api::{GetObjectReader, ObjectInfo, ObjectOptions, PutObjReader};
use crate::runtime::sources as runtime_sources; use crate::runtime::sources as runtime_sources;
use crate::set_disk::SetDisks; use crate::set_disk::SetDisks;
use crate::storage_api_contracts::{admin::StorageAdminApi, namespace::NamespaceLocking as _, range::HTTPRangeSpec}; use crate::storage_api_contracts::{
admin::StorageAdminApi, namespace::NamespaceLocking as _, object::HTTPPreconditions, range::HTTPRangeSpec,
};
use http::HeaderMap; use http::HeaderMap;
use rustfs_config::audit::{AUDIT_AMQP_SUB_SYS, AUDIT_KAFKA_SUB_SYS, AUDIT_MQTT_SUB_SYS, AUDIT_WEBHOOK_SUB_SYS}; use rustfs_config::audit::{AUDIT_AMQP_SUB_SYS, AUDIT_KAFKA_SUB_SYS, AUDIT_MQTT_SUB_SYS, AUDIT_WEBHOOK_SUB_SYS};
use rustfs_config::notify::{ use rustfs_config::notify::{
@@ -3104,6 +3361,85 @@ mod tests {
); );
} }
#[test]
fn root_heal_null_decodes_as_no_override_and_is_canonicalized_on_save() {
let seed = br#"{
"version":"33",
"storageclass":{"standard":"","rrs":""},
"heal":null,
"future_root":{"mode":"keep"},
"openid":{"default":{
"config_url":"https://issuer.example/.well-known/openid-configuration",
"client_id":"console",
"client_secret":"oidc-secret",
"future_provider_control":"keep"
}},
"notify":{"webhook":{"primary":{
"enable":true,
"endpoint":"https://notify.example/hook",
"auth_token":"notify-secret",
"future_notify_control":"keep"
}}},
"logger":{"webhook":{"primary":{
"enable":true,
"endpoint":"https://audit.example/hook",
"auth_token":"audit-secret",
"future_audit_control":"keep"
}}}
}"#;
let cfg = decode_server_config_blob(seed).expect("root heal null should mean no persisted override");
assert!(cfg.get_value(HEAL_SUB_SYS, DEFAULT_DELIMITER).is_none());
assert!(!is_standard_object_server_config(seed));
let encoded = encode_server_config_blob(&cfg, Some(seed)).expect("legacy seed should canonicalize on an authorized save");
let value: Value = serde_json::from_slice(&encoded).expect("canonical config should be valid JSON");
assert!(value.get(HEAL_SUB_SYS).is_none());
assert_eq!(value["future_root"]["mode"].as_str(), Some("keep"));
assert_eq!(value["openid"]["default"]["client_secret"].as_str(), Some("oidc-secret"));
assert_eq!(value["openid"]["default"]["future_provider_control"].as_str(), Some("keep"));
assert_eq!(value["notify"]["webhook"]["primary"]["auth_token"].as_str(), Some("notify-secret"));
assert_eq!(value["notify"]["webhook"]["primary"]["future_notify_control"].as_str(), Some("keep"));
assert_eq!(value["logger"]["webhook"]["primary"]["auth_token"].as_str(), Some("audit-secret"));
assert_eq!(value["logger"]["webhook"]["primary"]["future_audit_control"].as_str(), Some("keep"));
assert!(is_standard_object_server_config(&encoded));
}
#[test]
fn invalid_scalar_and_nested_null_config_shapes_remain_rejected() {
let invalid_sections = [
r#""scanner":null"#,
r#""heal":"""#,
r#""heal":false"#,
r#""heal":0"#,
r#""heal":{"default":null}"#,
r#""heal":{"_":null}"#,
r#""heal":{"bitrot_cycle":null}"#,
r#""heal":[{"key":"bitrot_cycle","value":null}]"#,
];
for section in invalid_sections {
let input = format!(r#"{{"version":"33","storageclass":{{"standard":"","rrs":""}},{section}}}"#);
let err = decode_server_config_blob(input.as_bytes()).expect_err("invalid scalar shape must remain rejected");
assert!(
err.to_string().contains("expected"),
"invalid section {section} returned an unrelated error: {err}"
);
}
}
#[test]
fn valid_heal_object_and_kvs_array_shapes_remain_accepted() {
let empty_object = br#"{"version":"33","storageclass":{"standard":"","rrs":""},"heal":{}}"#;
let cfg = decode_server_config_blob(empty_object).expect("empty heal object should decode as no override");
assert!(cfg.get_value(HEAL_SUB_SYS, DEFAULT_DELIMITER).is_none());
let kvs_array =
br#"{"version":"33","storageclass":{"standard":"","rrs":""},"heal":[{"key":"bitrot_cycle","value":"off"}]}"#;
let cfg = decode_server_config_blob(kvs_array).expect("heal KVS array should decode");
assert!(cfg.get_value(HEAL_SUB_SYS, DEFAULT_DELIMITER).is_some());
}
#[test] #[test]
fn scanner_update_preserves_unknown_root_and_oidc_provider_fields() { fn scanner_update_preserves_unknown_root_and_oidc_provider_fields() {
let seed = br#"{ let seed = br#"{
@@ -3131,6 +3467,171 @@ mod tests {
assert_eq!(value["openid"]["default"]["client_id"].as_str(), Some("console")); assert_eq!(value["openid"]["default"]["client_id"].as_str(), Some("console"));
} }
#[test]
fn storageclass_reset_removes_stale_inline_block_from_seed() {
let seed = br#"{
"version":"33",
"storageclass":{
"standard":"EC:2",
"rrs":"EC:1",
"optimize":"availability",
"inline_block":"64KiB",
"future_storage_control":"keep"
}
}"#;
let encoded = encode_server_config_blob(&Config::new(), Some(seed)).expect("storageclass reset should encode");
let value: Value = serde_json::from_slice(&encoded).expect("encoded config should be valid json");
let storageclass = value["storageclass"].as_object().expect("storageclass object");
assert!(storageclass.get(crate::config::storageclass::INLINE_BLOCK).is_none());
assert_eq!(storageclass["future_storage_control"].as_str(), Some("keep"));
}
#[test]
fn target_update_preserves_unknown_fields_without_restoring_removed_instances() {
let seed = br#"{
"version":"33",
"storageclass":{"standard":"","rrs":""},
"notify":{"webhook":{
"primary":{
"enable":true,
"endpoint":"https://notify.example/old",
"auth_token":"notify-secret",
"future_control":"keep"
},
"removed":{"enable":true,"endpoint":"https://notify.example/removed"},
"retained":{"enable":true,"endpoint":"https://notify.example/retained","future_control":"keep"},
"enable":{"enable":true,"endpoint":"https://notify.example/named-enable","future_control":"keep"}
}}
}"#;
let mut cfg = decode_server_config_blob(seed).expect("target seed should decode");
let webhook = cfg
.0
.get_mut(NOTIFY_WEBHOOK_SUB_SYS)
.expect("notify webhook subsystem should exist");
webhook
.get_mut("primary")
.expect("primary target should exist")
.insert(rustfs_config::WEBHOOK_ENDPOINT.to_string(), "https://notify.example/new".to_string());
webhook.remove("removed");
webhook.remove("retained");
let encoded = encode_server_config_blob(&cfg, Some(seed)).expect("target update should encode");
let value: Value = serde_json::from_slice(&encoded).expect("encoded config should be valid json");
let webhook = value["notify"]["webhook"].as_object().expect("webhook section");
assert_eq!(webhook["primary"]["endpoint"].as_str(), Some("https://notify.example/new"));
assert_eq!(webhook["primary"]["auth_token"].as_str(), Some("notify-secret"));
assert_eq!(webhook["primary"]["future_control"].as_str(), Some("keep"));
assert!(webhook.get("removed").is_none());
assert_eq!(webhook["retained"]["future_control"].as_str(), Some("keep"));
assert!(webhook["retained"].get("enable").is_none());
assert!(webhook["retained"].get("endpoint").is_none());
assert_eq!(webhook["enable"]["endpoint"].as_str(), Some("https://notify.example/named-enable"));
assert_eq!(webhook["enable"]["future_control"].as_str(), Some("keep"));
}
#[test]
fn shorthand_target_update_preserves_shape_and_unknown_nested_fields() {
let seed = br#"{
"version":"33",
"storageclass":{"standard":"","rrs":""},
"notify":{"webhook":{
"enable":true,
"endpoint":"https://notify.example/old",
"future_control":{"endpoint":"leave-untouched","mode":"keep"}
}}
}"#;
let mut cfg = decode_server_config_blob(seed).expect("shorthand target should decode");
cfg.0
.get_mut(NOTIFY_WEBHOOK_SUB_SYS)
.and_then(|targets| targets.get_mut(DEFAULT_DELIMITER))
.expect("default webhook target should exist")
.insert(rustfs_config::WEBHOOK_ENDPOINT.to_string(), "https://notify.example/new".to_string());
let encoded = encode_server_config_blob(&cfg, Some(seed)).expect("shorthand target update should encode");
let value: Value = serde_json::from_slice(&encoded).expect("encoded config should be valid json");
let webhook = value["notify"]["webhook"].as_object().expect("webhook shorthand object");
assert_eq!(webhook["endpoint"].as_str(), Some("https://notify.example/new"));
assert!(webhook.get("default").is_none());
assert_eq!(webhook["future_control"]["endpoint"].as_str(), Some("leave-untouched"));
assert_eq!(webhook["future_control"]["mode"].as_str(), Some("keep"));
let decoded = decode_server_config_blob(&encoded).expect("updated shorthand target should remain decodable");
assert_eq!(
decoded
.get_value(NOTIFY_WEBHOOK_SUB_SYS, DEFAULT_DELIMITER)
.expect("updated default webhook target")
.get(rustfs_config::WEBHOOK_ENDPOINT),
"https://notify.example/new"
);
}
#[test]
fn target_kvs_update_preserves_unknown_entries_and_attributes() {
let seed = br#"{
"version":"33",
"storageclass":{"standard":"","rrs":""},
"notify":{"webhook":{"primary":[
{"key":"enable","value":"on","hidden_if_empty":false},
{"key":"endpoint","value":"https://notify.example/old","future_attribute":"keep-endpoint"},
{"key":"future_control","value":"keep","future_attribute":"keep-control"}
]}}
}"#;
let mut cfg = decode_server_config_blob(seed).expect("target KVS seed should decode");
cfg.0
.get_mut(NOTIFY_WEBHOOK_SUB_SYS)
.and_then(|targets| targets.get_mut("primary"))
.expect("primary webhook target should exist")
.insert(rustfs_config::WEBHOOK_ENDPOINT.to_string(), "https://notify.example/new".to_string());
let encoded = encode_server_config_blob(&cfg, Some(seed)).expect("target KVS update should encode");
let value: Value = serde_json::from_slice(&encoded).expect("encoded config should be valid json");
let entries = value["notify"]["webhook"]["primary"]
.as_array()
.expect("target KVS shape should be preserved");
let endpoint = entries
.iter()
.find(|entry| entry["key"].as_str() == Some(rustfs_config::WEBHOOK_ENDPOINT))
.expect("endpoint entry should remain");
let future = entries
.iter()
.find(|entry| entry["key"].as_str() == Some("future_control"))
.expect("unknown target entry should remain");
assert_eq!(endpoint["value"].as_str(), Some("https://notify.example/new"));
assert_eq!(endpoint["future_attribute"].as_str(), Some("keep-endpoint"));
assert_eq!(future["value"].as_str(), Some("keep"));
assert_eq!(future["future_attribute"].as_str(), Some("keep-control"));
}
#[test]
fn target_default_alias_is_canonicalized_without_losing_unknown_fields() {
let seed = br#"{
"version":"33",
"storageclass":{"standard":"","rrs":""},
"notify":{"webhook":{
"_":{"enable":false,"endpoint":"https://notify.example/alias","future_alias":"keep"},
"default":{"enable":true,"endpoint":"https://notify.example/default","future_default":"keep"}
}}
}"#;
let cfg = decode_server_config_blob(seed).expect("dual default aliases should decode");
let expected_endpoint = cfg
.get_value(NOTIFY_WEBHOOK_SUB_SYS, DEFAULT_DELIMITER)
.expect("default webhook target should exist")
.get(rustfs_config::WEBHOOK_ENDPOINT);
let encoded = encode_server_config_blob(&cfg, Some(seed)).expect("default alias should canonicalize");
let value: Value = serde_json::from_slice(&encoded).expect("encoded config should be valid json");
let webhook = value["notify"]["webhook"].as_object().expect("webhook section");
assert!(webhook.get(DEFAULT_DELIMITER).is_none());
assert_eq!(webhook["default"]["endpoint"].as_str(), Some(expected_endpoint.as_str()));
assert_eq!(webhook["default"]["future_alias"].as_str(), Some("keep"));
assert_eq!(webhook["default"]["future_default"].as_str(), Some("keep"));
}
#[test] #[test]
fn test_scanner_config_changes_are_semantically_significant() { fn test_scanner_config_changes_are_semantically_significant() {
let baseline = Config::new(); let baseline = Config::new();
@@ -4124,6 +4625,7 @@ mod tests {
/// What reads of the config object currently return. /// What reads of the config object currently return.
enum RecoveryReadState { enum RecoveryReadState {
Missing,
Blob(Vec<u8>), Blob(Vec<u8>),
QuorumError, QuorumError,
} }
@@ -4136,6 +4638,7 @@ mod tests {
heal_calls: AtomicUsize, heal_calls: AtomicUsize,
write_calls: AtomicUsize, write_calls: AtomicUsize,
last_put_no_lock: AtomicBool, last_put_no_lock: AtomicBool,
last_put_preconditions: Mutex<Option<HTTPPreconditions>>,
revision: AtomicUsize, revision: AtomicUsize,
drive_counts: Vec<usize>, drive_counts: Vec<usize>,
lock_manager: Arc<rustfs_lock::GlobalLockManager>, lock_manager: Arc<rustfs_lock::GlobalLockManager>,
@@ -4150,6 +4653,7 @@ mod tests {
heal_calls: AtomicUsize::new(0), heal_calls: AtomicUsize::new(0),
write_calls: AtomicUsize::new(0), write_calls: AtomicUsize::new(0),
last_put_no_lock: AtomicBool::new(false), last_put_no_lock: AtomicBool::new(false),
last_put_preconditions: Mutex::new(None),
revision: AtomicUsize::new(1), revision: AtomicUsize::new(1),
drive_counts: vec![2], drive_counts: vec![2],
lock_manager: Arc::new(rustfs_lock::GlobalLockManager::new()), lock_manager: Arc::new(rustfs_lock::GlobalLockManager::new()),
@@ -4206,6 +4710,7 @@ mod tests {
_opts: &ObjectOptions, _opts: &ObjectOptions,
) -> Result<GetObjectReader> { ) -> Result<GetObjectReader> {
let data = match &*self.state.lock().expect("state lock poisoned") { let data = match &*self.state.lock().expect("state lock poisoned") {
RecoveryReadState::Missing => return Err(Error::ConfigNotFound),
RecoveryReadState::Blob(data) => data.clone(), RecoveryReadState::Blob(data) => data.clone(),
RecoveryReadState::QuorumError => return Err(Error::ErasureReadQuorum), RecoveryReadState::QuorumError => return Err(Error::ErasureReadQuorum),
}; };
@@ -4213,6 +4718,9 @@ mod tests {
size: data.len() as i64, size: data.len() as i64,
actual_size: data.len() as i64, actual_size: data.len() as i64,
etag: Some(format!("config-{}", self.revision.load(Ordering::SeqCst))), etag: Some(format!("config-{}", self.revision.load(Ordering::SeqCst))),
data_dir: Some(uuid::Uuid::from_u128(
u128::try_from(self.revision.load(Ordering::SeqCst)).expect("test revision should fit in u128"),
)),
..Default::default() ..Default::default()
}; };
Ok(GetObjectReader { Ok(GetObjectReader {
@@ -4231,15 +4739,19 @@ mod tests {
opts: &ObjectOptions, opts: &ObjectOptions,
) -> Result<ObjectInfo> { ) -> Result<ObjectInfo> {
let current_etag = format!("config-{}", self.revision.load(Ordering::SeqCst)); let current_etag = format!("config-{}", self.revision.load(Ordering::SeqCst));
let object_exists = matches!(&*self.state.lock().expect("state lock poisoned"), RecoveryReadState::Blob(_));
if let Some(preconditions) = &opts.http_preconditions if let Some(preconditions) = &opts.http_preconditions
&& (preconditions.if_match_value().is_some_and(|etag| etag != current_etag) && (preconditions
|| preconditions.if_none_match_value() == Some("*")) .if_match_value()
.is_some_and(|etag| !object_exists || etag != current_etag)
|| (object_exists && preconditions.if_none_match_value() == Some("*")))
{ {
return Err(Error::PreconditionFailed); return Err(Error::PreconditionFailed);
} }
let mut body = Vec::new(); let mut body = Vec::new();
data.stream.read_to_end(&mut body).await?; data.stream.read_to_end(&mut body).await?;
self.last_put_no_lock.store(opts.no_lock, Ordering::SeqCst); self.last_put_no_lock.store(opts.no_lock, Ordering::SeqCst);
*self.last_put_preconditions.lock().expect("preconditions lock poisoned") = opts.http_preconditions.clone();
self.write_calls.fetch_add(1, Ordering::SeqCst); self.write_calls.fetch_add(1, Ordering::SeqCst);
*self.state.lock().expect("state lock poisoned") = RecoveryReadState::Blob(body.clone()); *self.state.lock().expect("state lock poisoned") = RecoveryReadState::Blob(body.clone());
let revision = self.revision.fetch_add(1, Ordering::SeqCst) + 1; let revision = self.revision.fetch_add(1, Ordering::SeqCst) + 1;
@@ -4247,6 +4759,7 @@ mod tests {
size: i64::try_from(body.len()).expect("test config should fit in i64"), size: i64::try_from(body.len()).expect("test config should fit in i64"),
actual_size: i64::try_from(body.len()).expect("test config should fit in i64"), actual_size: i64::try_from(body.len()).expect("test config should fit in i64"),
etag: Some(format!("config-{revision}")), etag: Some(format!("config-{revision}")),
data_dir: Some(uuid::Uuid::from_u128(u128::try_from(revision).expect("test revision should fit in u128"))),
..Default::default() ..Default::default()
}) })
} }
@@ -4284,10 +4797,18 @@ mod tests {
.expect("scanner-only config change should be persisted"); .expect("scanner-only config change should be persisted");
assert_eq!(store.write_calls.load(Ordering::SeqCst), 1); assert_eq!(store.write_calls.load(Ordering::SeqCst), 1);
assert!(store.last_put_no_lock.load(Ordering::SeqCst)); assert!(!store.last_put_no_lock.load(Ordering::SeqCst));
let preconditions = store
.last_put_preconditions
.lock()
.expect("preconditions lock poisoned")
.clone()
.expect("existing config update must be conditional");
assert_eq!(preconditions.if_match_value(), Some("config-1"));
assert_eq!(preconditions.if_none_match_value(), None);
assert_eq!( assert_eq!(
store.lock_resources.lock().expect("lock resources mutex poisoned").as_slice(), store.lock_resources.lock().expect("lock resources mutex poisoned").as_slice(),
&[server_config_path()] &[server_config_transaction_lock_path()]
); );
let decoded = read_config_without_migrate(store) let decoded = read_config_without_migrate(store)
.await .await
@@ -4301,6 +4822,73 @@ mod tests {
); );
} }
#[tokio::test]
async fn server_config_snapshot_save_returns_committed_generation() {
let baseline = encode_server_config_blob(&Config::new(), None).expect("baseline config should encode");
let store = Arc::new(RecoveryMockStore::new(RecoveryReadState::Blob(baseline), None));
let snapshot = read_server_config_snapshot(store.clone())
.await
.expect("server config snapshot");
let result = save_server_config_snapshot_with_generation(store, &config_with_scanner_cycle("61"), &snapshot)
.await
.expect("conditional config save");
assert!(result.persisted());
assert_eq!(result.generation(), Some(uuid::Uuid::from_u128(2)));
}
#[tokio::test]
async fn missing_server_config_is_created_with_if_none_match() {
let store = Arc::new(RecoveryMockStore::new(RecoveryReadState::Missing, None));
let cfg = config_with_scanner_cycle("61");
save_server_config(store.clone(), &cfg)
.await
.expect("missing config should be created conditionally");
assert_eq!(store.write_calls.load(Ordering::SeqCst), 1);
assert!(!store.last_put_no_lock.load(Ordering::SeqCst));
let preconditions = store
.last_put_preconditions
.lock()
.expect("preconditions lock poisoned")
.clone()
.expect("missing config create must be conditional");
assert_eq!(preconditions.if_match_value(), None);
assert_eq!(preconditions.if_none_match_value(), Some("*"));
let persisted = read_config_without_migrate(store)
.await
.expect("created config should reload");
assert_eq!(
persisted
.get_value(SCANNER_SUB_SYS, DEFAULT_DELIMITER)
.expect("persisted scanner config")
.get(SCANNER_CYCLE),
"61"
);
}
#[tokio::test]
async fn missing_config_initialization_recheck_preserves_concurrent_config() {
let existing = config_with_scanner_cycle("73");
let baseline = encode_server_config_blob(&existing, None).expect("existing config should encode");
let store = Arc::new(RecoveryMockStore::new(RecoveryReadState::Blob(baseline), None));
let observed = new_and_save_server_config(store.clone())
.await
.expect("initialization recheck should return the config created by another writer");
assert_eq!(store.write_calls.load(Ordering::SeqCst), 0);
assert_eq!(
observed
.get_value(SCANNER_SUB_SYS, DEFAULT_DELIMITER)
.expect("concurrent scanner config")
.get(SCANNER_CYCLE),
"73"
);
}
#[tokio::test] #[tokio::test]
async fn stale_server_config_snapshot_cannot_overwrite_newer_update() { async fn stale_server_config_snapshot_cannot_overwrite_newer_update() {
let baseline = encode_server_config_blob(&Config::new(), None).expect("baseline config should encode"); let baseline = encode_server_config_blob(&Config::new(), None).expect("baseline config should encode");
@@ -4357,7 +4945,7 @@ mod tests {
let lock = rustfs_lock::NamespaceLock::new("server-config-lease-loss".to_string(), client.clone()); let lock = rustfs_lock::NamespaceLock::new("server-config-lease-loss".to_string(), client.clone());
let guard = lock let guard = lock
.lock_guard( .lock_guard(
rustfs_lock::ObjectKey::new(crate::disk::RUSTFS_META_BUCKET, server_config_path()), rustfs_lock::ObjectKey::new(crate::disk::RUSTFS_META_BUCKET, server_config_transaction_lock_path()),
"server-config-lease-loss", "server-config-lease-loss",
std::time::Duration::from_secs(1), std::time::Duration::from_secs(1),
std::time::Duration::from_millis(120), std::time::Duration::from_millis(120),
@@ -4372,6 +4960,7 @@ mod tests {
raw: Some(baseline.clone()), raw: Some(baseline.clone()),
seed: None, seed: None,
etag: Some("config-0".to_string()), etag: Some("config-0".to_string()),
generation: Some(uuid::Uuid::from_u128(1)),
_local_guard: local_guard, _local_guard: local_guard,
_guard: guard, _guard: guard,
}; };
+121 -2
View File
@@ -5078,7 +5078,7 @@ impl LocalDisk {
Ok((buf, mtime)) Ok((buf, mtime))
} }
#[hotpath::measure] #[hotpath::measure(impl_type = "LocalDisk")]
async fn read_metadata_with_dmtime(&self, file_path: impl AsRef<Path>) -> Result<(Vec<u8>, Option<OffsetDateTime>)> { async fn read_metadata_with_dmtime(&self, file_path: impl AsRef<Path>) -> Result<(Vec<u8>, Option<OffsetDateTime>)> {
check_path_length(file_path.as_ref().to_string_lossy().as_ref())?; check_path_length(file_path.as_ref().to_string_lossy().as_ref())?;
@@ -5121,7 +5121,7 @@ impl LocalDisk {
Ok((data, modtime)) Ok((data, modtime))
} }
#[hotpath::measure] #[hotpath::measure(impl_type = "LocalDisk")]
async fn read_all_data(&self, volume: &str, volume_dir: impl AsRef<Path>, file_path: impl AsRef<Path>) -> Result<Vec<u8>> { async fn read_all_data(&self, volume: &str, volume_dir: impl AsRef<Path>, file_path: impl AsRef<Path>) -> Result<Vec<u8>> {
// TODO: timeout support // TODO: timeout support
let (data, _) = self.read_all_data_with_dmtime(volume, volume_dir, file_path).await?; let (data, _) = self.read_all_data_with_dmtime(volume, volume_dir, file_path).await?;
@@ -6641,6 +6641,49 @@ impl LocalDisk {
let xl_path = object_dir.join(STORAGE_FORMAT_FILE); let xl_path = object_dir.join(STORAGE_FORMAT_FILE);
restore_delete_rollback_after_error(object_dir, &xl_path, Some(rollback_dir), volume, object, stage, err).await restore_delete_rollback_after_error(object_dir, &xl_path, Some(rollback_dir), volume, object, stage, err).await
} }
/// Execute every deferred data-dir deletion pending on `volume` right now,
/// even while snapshot leases are still held. Bucket deletion requires it:
/// a streaming reader defers the physical cleanup of an already-deleted
/// version, and a non-force `delete_volume` would otherwise fail closed
/// with `VolumeNotEmpty` on those remnants even though the bucket is
/// logically empty. The still-active readers keep their open descriptors;
/// only path-based reopens observe the removal.
async fn settle_pending_snapshot_deletes(&self, volume: &str) {
let pending: Vec<(SnapshotLeaseKey, DeleteOptions)> = {
let mut registry = self.snapshot_leases.lock().await;
registry
.entries
.iter_mut()
.filter(|(key, entry)| key.volume == volume && !entry.deleting && entry.pending_delete.is_some())
.map(|(key, entry)| {
entry.deleting = true;
(key.clone(), entry.pending_delete.clone().expect("filtered on Some"))
})
.collect()
};
for (key, opts) in pending {
let result = self.delete_unleased(&key.volume, &key.path, &opts).await;
let mut registry = self.snapshot_leases.lock().await;
match result {
Ok(()) => {
registry.entries.remove(&key);
}
Err(err) => {
if let Some(entry) = registry.entries.get_mut(&key) {
entry.deleting = false;
}
warn!(
volume = %key.volume,
path = %key.path,
error = %err,
"failed to settle deferred data-dir deletion before volume removal"
);
}
}
}
}
} }
#[async_trait::async_trait] #[async_trait::async_trait]
@@ -9122,6 +9165,14 @@ impl DiskAPI for LocalDisk {
let p = self.get_bucket_path(volume)?; let p = self.get_bucket_path(volume)?;
let _volume_mutation_guard = os::disk_volume_mutation_lock(&self.root, volume).write_owned().await; let _volume_mutation_guard = os::disk_volume_mutation_lock(&self.root, volume).write_owned().await;
// A streaming reader's snapshot lease defers the physical cleanup of
// data dirs whose version delete already committed. Those remnants are
// logically deleted, so run the parked cleanups now instead of letting
// the non-force removal below fail closed on them (the s3-tests SSE-C
// teardown races exactly this way: DeleteObjects, then DeleteBucket
// while an abandoned GET body still pins the lease).
self.settle_pending_snapshot_deletes(volume).await;
// Non-force removes empty directory remnants children-first with // Non-force removes empty directory remnants children-first with
// non-recursive rmdir calls. A file that exists during the scan, or // non-recursive rmdir calls. A file that exists during the scan, or
// appears before its parent is removed, fails closed with // appears before its parent is removed, fails closed with
@@ -15981,6 +16032,74 @@ mod test {
assert!(matches!(disk.read_all(volume, &first_part).await, Err(DiskError::FileNotFound))); assert!(matches!(disk.read_all(volume, &first_part).await, Err(DiskError::FileNotFound)));
} }
#[tokio::test]
async fn delete_volume_settles_lease_deferred_cleanup() {
use tempfile::tempdir;
let root_dir = tempdir().expect("temp dir should be created");
let endpoint = Endpoint::try_from(root_dir.path().to_string_lossy().as_ref()).expect("endpoint should parse");
let disk = LocalDisk::new(&endpoint, false).await.expect("local disk should be created");
let volume = "snapshot-lease-bucket-delete";
let object = "multipart_enc";
let version_id = Uuid::new_v4();
let data_dir = Uuid::new_v4();
let rollback_dir = Uuid::new_v4();
let data_path = path_join_buf(&[object, &data_dir.to_string()]);
let part = path_join_buf(&[&data_path, "part.1"]);
ensure_test_volume(&disk, volume).await;
disk.write_all(volume, &part, Bytes::from_static(b"payload"))
.await
.expect("shard should be written");
let fi = test_file_info(object, version_id, Some(data_dir), None);
disk.write_all(volume, &path_join_buf(&[object, STORAGE_FORMAT_FILE]), test_meta(fi.clone()).into())
.await
.expect("metadata should be written");
// An abandoned streaming GET pins the data dir with a snapshot lease.
let snapshot = disk
.acquire_snapshot_lease(volume, &data_path)
.await
.expect("snapshot lease should be acquired");
disk.delete_version(
volume,
object,
fi.clone(),
false,
DeleteOptions {
old_data_dir: Some(rollback_dir),
..Default::default()
},
)
.await
.expect("version delete should commit metadata");
disk.delete(
volume,
&format!("{object}/{rollback_dir}"),
DeleteOptions {
recursive: true,
immediate: true,
..Default::default()
},
)
.await
.expect("version delete should schedule physical cleanup");
// The bucket is logically empty; a non-force volume delete must settle
// the deferred data-dir cleanup instead of failing with VolumeNotEmpty.
disk.delete_volume(volume, false)
.await
.expect("bucket delete must not observe lease-deferred remnants");
assert!(matches!(
disk.read_all(volume, &part).await,
Err(DiskError::FileNotFound | DiskError::VolumeNotFound)
));
// The late lease release finds nothing pending and stays idempotent.
disk.release_snapshot_lease(volume, &data_path, snapshot)
.await
.expect("releasing the lease after bucket deletion should be a no-op");
}
#[tokio::test] #[tokio::test]
async fn version_delete_cleanup_intent_survives_local_disk_restart() { async fn version_delete_cleanup_intent_survives_local_disk_restart() {
use tempfile::tempdir; use tempfile::tempdir;
+2 -2
View File
@@ -103,7 +103,7 @@ where
/// or `out` is larger than one shard. On error `out`'s contents are /// or `out` is larger than one shard. On error `out`'s contents are
/// unspecified but never contain bytes that failed the hash check — the copy /// unspecified but never contain bytes that failed the hash check — the copy
/// happens only after verification. /// happens only after verification.
#[hotpath::measure] #[hotpath::measure(impl_type = "BitrotReader")]
pub async fn read(&mut self, out: &mut [u8]) -> std::io::Result<usize> { pub async fn read(&mut self, out: &mut [u8]) -> std::io::Result<usize> {
let want = out.len(); let want = out.len();
self.begin_read(want)?; self.begin_read(want)?;
@@ -303,7 +303,7 @@ where
/// Write a (hash+data) block. Returns the number of data bytes written. /// Write a (hash+data) block. Returns the number of data bytes written.
/// Returns an error if called after a short write or if data exceeds shard_size. /// Returns an error if called after a short write or if data exceeds shard_size.
#[hotpath::measure(label = "BitrotWriter::write")] #[hotpath::measure(label = "BitrotWriter::write", impl_type = "BitrotWriter")]
pub async fn write(&mut self, buf: &[u8]) -> std::io::Result<usize> { pub async fn write(&mut self, buf: &[u8]) -> std::io::Result<usize> {
if buf.is_empty() { if buf.is_empty() {
return Ok(0); return Ok(0);
+111 -4
View File
@@ -691,7 +691,7 @@ impl<R> ParallelReader<R>
where where
R: crate::erasure::coding::ShardSource, R: crate::erasure::coding::ShardSource,
{ {
#[hotpath::measure] #[hotpath::measure(impl_type = "ParallelReader")]
pub async fn read(&mut self) -> (Vec<Option<Vec<u8>>>, Vec<Option<Error>>) { pub async fn read(&mut self) -> (Vec<Option<Vec<u8>>>, Vec<Option<Error>>) {
// On the reconstruction-verifying GET path, read every live shard reader // On the reconstruction-verifying GET path, read every live shard reader
// in lockstep so all readers advance one block per stripe and stay // in lockstep so all readers advance one block per stripe and stay
@@ -1505,7 +1505,7 @@ where
} }
impl Erasure { impl Erasure {
#[hotpath::measure] #[hotpath::measure(impl_type = "Erasure")]
pub async fn decode<W, R>( pub async fn decode<W, R>(
&self, &self,
writer: &mut W, writer: &mut W,
@@ -1645,15 +1645,28 @@ impl Erasure {
} }
Err(e) => { Err(e) => {
record_get_stage_duration_if_enabled(GET_OBJECT_PATH_LEGACY_DUPLEX, GET_STAGE_EMIT, emit_stage_start); record_get_stage_duration_if_enabled(GET_OBJECT_PATH_LEGACY_DUPLEX, GET_STAGE_EMIT, emit_stage_start);
let reason = classify_io_error(&e);
if reason == GetObjectFailureReason::DownstreamClosed {
debug!(
block_offset,
block_length,
bytes_written = *written,
stage = GET_STAGE_EMIT,
reason = reason.as_str(),
error = ?e,
"Erasure decode stopped after downstream closed"
);
} else {
error!( error!(
block_offset, block_offset,
block_length, block_length,
bytes_written = *written, bytes_written = *written,
stage = GET_STAGE_EMIT, stage = GET_STAGE_EMIT,
reason = classify_io_error(&e).as_str(), reason = reason.as_str(),
error = ?e, error = ?e,
"Erasure decode failed to emit reconstructed data" "Erasure decode failed to emit reconstructed data"
); );
}
*ret_err = Some(e); *ret_err = Some(e);
return StripeFlow::Stop; return StripeFlow::Stop;
} }
@@ -1945,7 +1958,7 @@ mod tests {
use std::io::Cursor; use std::io::Cursor;
use std::pin::Pin; use std::pin::Pin;
use std::sync::{ use std::sync::{
Arc, Arc, Mutex,
atomic::{AtomicUsize, Ordering}, atomic::{AtomicUsize, Ordering},
}; };
use std::task::{Context, Poll}; use std::task::{Context, Poll};
@@ -2120,6 +2133,59 @@ mod tests {
} }
} }
struct DownstreamClosedWriter;
impl AsyncWrite for DownstreamClosedWriter {
fn poll_write(self: Pin<&mut Self>, _cx: &mut Context<'_>, _buf: &[u8]) -> Poll<io::Result<usize>> {
Poll::Ready(Err(crate::diagnostics::get::mark_get_object_downstream_closed(io::Error::new(
ErrorKind::BrokenPipe,
"injected downstream close",
))))
}
fn poll_flush(self: Pin<&mut Self>, _cx: &mut Context<'_>) -> Poll<io::Result<()>> {
Poll::Ready(Ok(()))
}
fn poll_shutdown(self: Pin<&mut Self>, _cx: &mut Context<'_>) -> Poll<io::Result<()>> {
Poll::Ready(Ok(()))
}
}
#[derive(Clone, Default)]
struct CapturedLogs(Arc<Mutex<Vec<u8>>>);
struct CapturedLogWriter(Arc<Mutex<Vec<u8>>>);
impl CapturedLogs {
fn contents(&self) -> String {
String::from_utf8(self.0.lock().expect("captured logs mutex should not be poisoned").clone())
.expect("captured logs should be valid UTF-8")
}
}
impl std::io::Write for CapturedLogWriter {
fn write(&mut self, buf: &[u8]) -> std::io::Result<usize> {
self.0
.lock()
.expect("captured logs mutex should not be poisoned")
.extend_from_slice(buf);
Ok(buf.len())
}
fn flush(&mut self) -> std::io::Result<()> {
Ok(())
}
}
impl<'a> tracing_subscriber::fmt::MakeWriter<'a> for CapturedLogs {
type Writer = CapturedLogWriter;
fn make_writer(&'a self) -> Self::Writer {
CapturedLogWriter(Arc::clone(&self.0))
}
}
#[test] #[test]
fn parallel_reader_constructor_variants_preserve_read_cost_and_verification_flags() { fn parallel_reader_constructor_variants_preserve_read_cost_and_verification_flags() {
let erasure = Erasure::new(2, 1, 64); let erasure = Erasure::new(2, 1, 64);
@@ -2215,6 +2281,47 @@ mod tests {
assert_eq!(err.to_string(), "injected emit failure"); assert_eq!(err.to_string(), "injected emit failure");
} }
#[tokio::test(flavor = "current_thread")]
async fn erasure_decode_logs_reconstructed_downstream_close_at_debug() {
let logs = CapturedLogs::default();
let subscriber = tracing_subscriber::fmt()
.with_max_level(tracing::Level::DEBUG)
.with_writer(logs.clone())
.with_ansi(false)
.without_time()
.finish();
let _guard = tracing::subscriber::set_default(subscriber);
let erasure = Erasure::new(2, 1, 64);
let data: Vec<u8> = (0..64).collect();
let shard_size = erasure.shard_size();
let encoded = erasure.encode_data(&data).expect("test data should encode");
let readers = vec![
None,
Some(BitrotReader::new(
Cursor::new(encoded[1].to_vec()),
shard_size,
HashAlgorithm::None,
false,
)),
Some(BitrotReader::new(
Cursor::new(encoded[2].to_vec()),
shard_size,
HashAlgorithm::None,
false,
)),
];
let mut writer = DownstreamClosedWriter;
let (written, err) = erasure.decode(&mut writer, readers, 0, data.len(), data.len()).await;
assert_eq!(written, 0);
assert_eq!(err.expect("downstream close must still terminate the GET").kind(), ErrorKind::BrokenPipe);
let captured = logs.contents();
assert!(captured.contains("Erasure decode stopped after downstream closed"));
assert!(!captured.contains("Erasure decode failed to emit reconstructed data"));
}
#[tokio::test] #[tokio::test]
async fn test_erasure_decode_rejects_reader_count_and_range_overflow() { async fn test_erasure_decode_rejects_reader_count_and_range_overflow() {
let erasure = Erasure::new(2, 1, 64); let erasure = Erasure::new(2, 1, 64);
+658 -56
View File
@@ -28,8 +28,11 @@ use std::vec;
use tokio::io::AsyncRead; use tokio::io::AsyncRead;
use tokio::runtime::RuntimeFlavor; use tokio::runtime::RuntimeFlavor;
use tokio::sync::mpsc; use tokio::sync::mpsc;
use tokio::task::{JoinError, JoinHandle};
use tracing::error; use tracing::error;
/// Queue-capacity input for encoded blocks awaiting shard writers; it is not a
/// per-PUT or process-RSS memory limit.
const ENV_RUSTFS_ERASURE_ENCODE_MAX_INFLIGHT_BYTES: &str = "RUSTFS_ERASURE_ENCODE_MAX_INFLIGHT_BYTES"; const ENV_RUSTFS_ERASURE_ENCODE_MAX_INFLIGHT_BYTES: &str = "RUSTFS_ERASURE_ENCODE_MAX_INFLIGHT_BYTES";
const ENV_RUSTFS_ERASURE_ENCODE_BATCH_BLOCKS: &str = "RUSTFS_ERASURE_ENCODE_BATCH_BLOCKS"; const ENV_RUSTFS_ERASURE_ENCODE_BATCH_BLOCKS: &str = "RUSTFS_ERASURE_ENCODE_BATCH_BLOCKS";
const ENV_RUSTFS_ERASURE_ENCODE_BYTESMUT_INGEST: &str = "RUSTFS_ERASURE_ENCODE_BYTESMUT_INGEST"; const ENV_RUSTFS_ERASURE_ENCODE_BYTESMUT_INGEST: &str = "RUSTFS_ERASURE_ENCODE_BYTESMUT_INGEST";
@@ -87,6 +90,33 @@ fn use_bytesmut_ingest() -> bool {
rustfs_utils::get_env_bool(ENV_RUSTFS_ERASURE_ENCODE_BYTESMUT_INGEST, DEFAULT_RUSTFS_ERASURE_ENCODE_BYTESMUT_INGEST) rustfs_utils::get_env_bool(ENV_RUSTFS_ERASURE_ENCODE_BYTESMUT_INGEST, DEFAULT_RUSTFS_ERASURE_ENCODE_BYTESMUT_INGEST)
}) })
} }
/// Keeps the encoder producer scoped to its parent future. Tokio detaches a
/// task when its `JoinHandle` is dropped, so the producer must be aborted when
/// an upload is cancelled before the encode pipeline finishes.
struct AbortOnDropTask<T>(JoinHandle<T>);
impl<T> AbortOnDropTask<T> {
fn new(task: JoinHandle<T>) -> Self {
Self(task)
}
async fn abort_and_wait(&mut self) {
self.0.abort();
let _ = (&mut self.0).await;
}
async fn join(&mut self) -> Result<T, JoinError> {
(&mut self.0).await
}
}
impl<T> Drop for AbortOnDropTask<T> {
fn drop(&mut self) {
self.0.abort();
}
}
/// Read up to `limit` bytes into `buf`'s uninitialized spare capacity, appending after its /// Read up to `limit` bytes into `buf`'s uninitialized spare capacity, appending after its
/// current length, and distinguish a clean EOF from a short read. /// current length, and distinguish a clean EOF from a short read.
/// ///
@@ -133,22 +163,65 @@ fn queued_block_bytes(block: &[Bytes]) -> usize {
block.iter().map(Bytes::len).sum() block.iter().map(Bytes::len).sum()
} }
async fn drain_queued_inflight_bytes(rx: &mut mpsc::Receiver<Vec<Bytes>>) { /// Owns an encoded queue entry's gauge contribution until its consumer takes it.
while let Some(block) = rx.recv().await { struct QueuedInflightBytes {
rustfs_io_metrics::remove_ec_encode_inflight_bytes(queued_block_bytes(&block)); bytes: usize,
} }
impl QueuedInflightBytes {
fn new(bytes: usize) -> Self {
rustfs_io_metrics::add_ec_encode_inflight_bytes(bytes);
Self { bytes }
}
fn settle(&mut self) {
let bytes = std::mem::take(&mut self.bytes);
if bytes != 0 {
rustfs_io_metrics::remove_ec_encode_inflight_bytes(bytes);
}
}
}
impl Drop for QueuedInflightBytes {
fn drop(&mut self) {
self.settle();
}
}
/// Couples an encoded block with its queue gauge contribution. Keeping the
/// guard in the queue entry also covers Tokio sends that complete through a
/// permit after the receiver has closed.
struct InflightEntry<T> {
entry: T,
accounting: QueuedInflightBytes,
}
impl<T> InflightEntry<T> {
fn new(entry: T, bytes: usize) -> Self {
Self {
entry,
accounting: QueuedInflightBytes::new(bytes),
}
}
fn into_inner(mut self) -> T {
self.accounting.settle();
self.entry
}
}
async fn send_queued<T>(
sender: &mpsc::Sender<InflightEntry<T>>,
entry: T,
bytes: usize,
) -> Result<(), mpsc::error::SendError<InflightEntry<T>>> {
sender.send(InflightEntry::new(entry, bytes)).await
} }
fn queued_batch_bytes(batch: &[Vec<Bytes>]) -> usize { fn queued_batch_bytes(batch: &[Vec<Bytes>]) -> usize {
batch.iter().map(|block| queued_block_bytes(block)).sum() batch.iter().map(|block| queued_block_bytes(block)).sum()
} }
async fn drain_queued_batched_inflight_bytes(rx: &mut mpsc::Receiver<Vec<Vec<Bytes>>>) {
while let Some(batch) = rx.recv().await {
rustfs_io_metrics::remove_ec_encode_inflight_bytes(queued_batch_bytes(&batch));
}
}
fn dominant_error_summary_label(summary: &WriteQuorumFailureSummary) -> &'static str { fn dominant_error_summary_label(summary: &WriteQuorumFailureSummary) -> &'static str {
summary.dominant_error_label summary.dominant_error_label
} }
@@ -504,7 +577,7 @@ impl Erasure {
Ok((reader, total)) Ok((reader, total))
} }
#[hotpath::measure] #[hotpath::measure(impl_type = "Erasure")]
pub async fn encode<R>( pub async fn encode<R>(
self: Arc<Self>, self: Arc<Self>,
reader: R, reader: R,
@@ -540,13 +613,14 @@ impl Erasure {
)); ));
} }
// Bound queued encoded blocks by memory budget to avoid per-request spikes. // Bound queued encoded blocks by a queue budget; this does not bound
// reader, encoder, writer, allocator, or process-RSS memory.
let expanded_block_bytes = self.shard_size().saturating_mul(self.total_shard_count()); let expanded_block_bytes = self.shard_size().saturating_mul(self.total_shard_count());
let max_inflight_bytes = erasure_encode_max_inflight_bytes(); let max_inflight_bytes = erasure_encode_max_inflight_bytes();
let inflight_blocks = encode_channel_capacity(expanded_block_bytes, max_inflight_bytes); let inflight_blocks = encode_channel_capacity(expanded_block_bytes, max_inflight_bytes);
let (tx, mut rx) = mpsc::channel::<Vec<Bytes>>(inflight_blocks); let (tx, mut rx) = mpsc::channel::<InflightEntry<Vec<Bytes>>>(inflight_blocks);
let task = tokio::spawn(async move { let mut task = AbortOnDropTask::new(tokio::spawn(async move {
let block_size = self.block_size; let block_size = self.block_size;
let mut total = 0; let mut total = 0;
if use_bytesmut_ingest { if use_bytesmut_ingest {
@@ -567,10 +641,9 @@ impl Erasure {
let res = self.clone().encode_block_bytes_mut(encode_buf, n).await?; let res = self.clone().encode_block_bytes_mut(encode_buf, n).await?;
buf = BytesMut::with_capacity(ingest_capacity); buf = BytesMut::with_capacity(ingest_capacity);
let queued_bytes = queued_block_bytes(&res); let queued_bytes = queued_block_bytes(&res);
rustfs_io_metrics::add_ec_encode_inflight_bytes(queued_bytes); let _producer_stage = rustfs_io_metrics::track_ec_encode_producer_bytes(queued_bytes);
let send_wait_stage_start = stage_timer_if_enabled(); let send_wait_stage_start = stage_timer_if_enabled();
if let Err(err) = tx.send(res).await { if let Err(err) = send_queued(&tx, res, queued_bytes).await {
rustfs_io_metrics::remove_ec_encode_inflight_bytes(queued_bytes);
return Err(std::io::Error::other(format!("Failed to send encoded data : {err}"))); return Err(std::io::Error::other(format!("Failed to send encoded data : {err}")));
} }
record_internal_stage_if_enabled("erasure_encode_send_wait", send_wait_stage_start); record_internal_stage_if_enabled("erasure_encode_send_wait", send_wait_stage_start);
@@ -598,10 +671,9 @@ impl Erasure {
let (res, returned_buf) = self.clone().encode_block(encode_buf, n).await?; let (res, returned_buf) = self.clone().encode_block(encode_buf, n).await?;
buf = returned_buf; buf = returned_buf;
let queued_bytes = queued_block_bytes(&res); let queued_bytes = queued_block_bytes(&res);
rustfs_io_metrics::add_ec_encode_inflight_bytes(queued_bytes); let _producer_stage = rustfs_io_metrics::track_ec_encode_producer_bytes(queued_bytes);
let send_wait_stage_start = stage_timer_if_enabled(); let send_wait_stage_start = stage_timer_if_enabled();
if let Err(err) = tx.send(res).await { if let Err(err) = send_queued(&tx, res, queued_bytes).await {
rustfs_io_metrics::remove_ec_encode_inflight_bytes(queued_bytes);
return Err(std::io::Error::other(format!("Failed to send encoded data : {err}"))); return Err(std::io::Error::other(format!("Failed to send encoded data : {err}")));
} }
record_internal_stage_if_enabled("erasure_encode_send_wait", send_wait_stage_start); record_internal_stage_if_enabled("erasure_encode_send_wait", send_wait_stage_start);
@@ -626,7 +698,7 @@ impl Erasure {
} }
Ok((reader, total)) Ok((reader, total))
}); }));
let mut writers = MultiWriter::new(writers, quorum); let mut writers = MultiWriter::new(writers, quorum);
@@ -638,11 +710,11 @@ impl Erasure {
break; break;
}; };
record_internal_stage_if_enabled("erasure_encode_recv_wait", recv_wait_stage_start); record_internal_stage_if_enabled("erasure_encode_recv_wait", recv_wait_stage_start);
let block = block.into_inner();
if block.is_empty() { if block.is_empty() {
break; break;
} }
let queued_bytes = queued_block_bytes(&block); let _writer_stage = rustfs_io_metrics::track_ec_encode_writer_bytes(queued_block_bytes(&block));
rustfs_io_metrics::remove_ec_encode_inflight_bytes(queued_bytes);
let write_stage_start = stage_timer_if_enabled(); let write_stage_start = stage_timer_if_enabled();
if let Err(err) = writers.write(block).await { if let Err(err) = writers.write(block).await {
write_err = Some(err); write_err = Some(err);
@@ -652,9 +724,8 @@ impl Erasure {
} }
if let Some(err) = write_err { if let Some(err) = write_err {
task.abort(); task.abort_and_wait().await;
let _ = task.await; drop(rx);
drain_queued_inflight_bytes(&mut rx).await;
let shutdown_stage_start = stage_timer_if_enabled(); let shutdown_stage_start = stage_timer_if_enabled();
if let Err(shutdown_err) = writers.shutdown().await { if let Err(shutdown_err) = writers.shutdown().await {
error!("failed to shutdown erasure writers after write error: {:?}", shutdown_err); error!("failed to shutdown erasure writers after write error: {:?}", shutdown_err);
@@ -663,14 +734,14 @@ impl Erasure {
return Err(err); return Err(err);
} }
let (reader, total) = task.await??; let (reader, total) = task.join().await??;
let shutdown_stage_start = stage_timer_if_enabled(); let shutdown_stage_start = stage_timer_if_enabled();
writers.shutdown().await?; writers.shutdown().await?;
record_internal_stage_if_enabled("erasure_encode_shutdown", shutdown_stage_start); record_internal_stage_if_enabled("erasure_encode_shutdown", shutdown_stage_start);
Ok((reader, total)) Ok((reader, total))
} }
#[hotpath::measure] #[hotpath::measure(impl_type = "Erasure")]
pub async fn encode_batched<R>( pub async fn encode_batched<R>(
self: Arc<Self>, self: Arc<Self>,
mut reader: R, mut reader: R,
@@ -692,14 +763,15 @@ impl Erasure {
let inflight_blocks = encode_channel_capacity(expanded_block_bytes, max_inflight_bytes); let inflight_blocks = encode_channel_capacity(expanded_block_bytes, max_inflight_bytes);
let batch_blocks = encode_batch_block_count().min(inflight_blocks); let batch_blocks = encode_batch_block_count().min(inflight_blocks);
let channel_capacity = inflight_blocks.div_ceil(batch_blocks).max(1); let channel_capacity = inflight_blocks.div_ceil(batch_blocks).max(1);
let (tx, mut rx) = mpsc::channel::<Vec<Vec<Bytes>>>(channel_capacity); let (tx, mut rx) = mpsc::channel::<InflightEntry<Vec<Vec<Bytes>>>>(channel_capacity);
let task = tokio::spawn(async move { let mut task = AbortOnDropTask::new(tokio::spawn(async move {
let block_size = self.block_size; let block_size = self.block_size;
let mut total = 0; let mut total = 0;
let mut buf = vec![0u8; block_size]; let mut buf = vec![0u8; block_size];
let mut pending_batch = Vec::with_capacity(batch_blocks); let mut pending_batch = Vec::with_capacity(batch_blocks);
let mut pending_batch_bytes = 0usize; let mut pending_batch_bytes = 0usize;
let mut pending_batch_stage = None;
loop { loop {
match rustfs_utils::read_full_or_eof(&mut reader, &mut buf).await { match rustfs_utils::read_full_or_eof(&mut reader, &mut buf).await {
Ok(Some(n)) => { Ok(Some(n)) => {
@@ -711,15 +783,16 @@ impl Erasure {
let queued_bytes = queued_block_bytes(&res); let queued_bytes = queued_block_bytes(&res);
pending_batch_bytes = pending_batch_bytes.saturating_add(queued_bytes); pending_batch_bytes = pending_batch_bytes.saturating_add(queued_bytes);
pending_batch.push(res); pending_batch.push(res);
drop(pending_batch_stage.take());
pending_batch_stage = Some(rustfs_io_metrics::track_ec_encode_producer_bytes(pending_batch_bytes));
if pending_batch.len() >= batch_blocks { if pending_batch.len() >= batch_blocks {
rustfs_io_metrics::add_ec_encode_inflight_bytes(pending_batch_bytes);
let send_wait_stage_start = stage_timer_if_enabled(); let send_wait_stage_start = stage_timer_if_enabled();
if let Err(err) = tx.send(pending_batch).await { if let Err(err) = send_queued(&tx, pending_batch, pending_batch_bytes).await {
rustfs_io_metrics::remove_ec_encode_inflight_bytes(pending_batch_bytes);
return Err(std::io::Error::other(format!("Failed to send encoded data : {err}"))); return Err(std::io::Error::other(format!("Failed to send encoded data : {err}")));
} }
record_internal_stage_if_enabled("erasure_encode_batched_send_wait", send_wait_stage_start); record_internal_stage_if_enabled("erasure_encode_batched_send_wait", send_wait_stage_start);
drop(pending_batch_stage.take());
pending_batch = Vec::with_capacity(batch_blocks); pending_batch = Vec::with_capacity(batch_blocks);
pending_batch_bytes = 0; pending_batch_bytes = 0;
} }
@@ -742,17 +815,16 @@ impl Erasure {
} }
if !pending_batch.is_empty() { if !pending_batch.is_empty() {
rustfs_io_metrics::add_ec_encode_inflight_bytes(pending_batch_bytes);
let send_wait_stage_start = stage_timer_if_enabled(); let send_wait_stage_start = stage_timer_if_enabled();
if let Err(err) = tx.send(pending_batch).await { if let Err(err) = send_queued(&tx, pending_batch, pending_batch_bytes).await {
rustfs_io_metrics::remove_ec_encode_inflight_bytes(pending_batch_bytes);
return Err(std::io::Error::other(format!("Failed to send encoded data : {err}"))); return Err(std::io::Error::other(format!("Failed to send encoded data : {err}")));
} }
record_internal_stage_if_enabled("erasure_encode_batched_send_wait", send_wait_stage_start); record_internal_stage_if_enabled("erasure_encode_batched_send_wait", send_wait_stage_start);
drop(pending_batch_stage);
} }
Ok((reader, total)) Ok((reader, total))
}); }));
let mut writers = MultiWriter::new(writers, quorum); let mut writers = MultiWriter::new(writers, quorum);
let mut write_err = None; let mut write_err = None;
@@ -763,7 +835,8 @@ impl Erasure {
break; break;
}; };
record_internal_stage_if_enabled("erasure_encode_batched_recv_wait", recv_wait_stage_start); record_internal_stage_if_enabled("erasure_encode_batched_recv_wait", recv_wait_stage_start);
rustfs_io_metrics::remove_ec_encode_inflight_bytes(queued_batch_bytes(&batch)); let batch = batch.into_inner();
let _writer_stage = rustfs_io_metrics::track_ec_encode_writer_bytes(queued_batch_bytes(&batch));
let write_stage_start = stage_timer_if_enabled(); let write_stage_start = stage_timer_if_enabled();
for block in batch { for block in batch {
if let Err(err) = writers.write(block).await { if let Err(err) = writers.write(block).await {
@@ -778,9 +851,8 @@ impl Erasure {
} }
if let Some(err) = write_err { if let Some(err) = write_err {
task.abort(); task.abort_and_wait().await;
let _ = task.await; drop(rx);
drain_queued_batched_inflight_bytes(&mut rx).await;
let shutdown_stage_start = stage_timer_if_enabled(); let shutdown_stage_start = stage_timer_if_enabled();
if let Err(shutdown_err) = writers.shutdown().await { if let Err(shutdown_err) = writers.shutdown().await {
error!("failed to shutdown erasure writers after write error: {:?}", shutdown_err); error!("failed to shutdown erasure writers after write error: {:?}", shutdown_err);
@@ -789,7 +861,7 @@ impl Erasure {
return Err(err); return Err(err);
} }
let (reader, total) = task.await??; let (reader, total) = task.join().await??;
let shutdown_stage_start = stage_timer_if_enabled(); let shutdown_stage_start = stage_timer_if_enabled();
writers.shutdown().await?; writers.shutdown().await?;
record_internal_stage_if_enabled("erasure_encode_batched_shutdown", shutdown_stage_start); record_internal_stage_if_enabled("erasure_encode_batched_shutdown", shutdown_stage_start);
@@ -798,7 +870,7 @@ impl Erasure {
/// Fast path for small inline objects: skip tokio::spawn + mpsc channel. /// Fast path for small inline objects: skip tokio::spawn + mpsc channel.
/// Reads all data, encodes directly, writes shards sequentially. /// Reads all data, encodes directly, writes shards sequentially.
#[hotpath::measure] #[hotpath::measure(impl_type = "Erasure")]
pub async fn encode_inline_small<R>( pub async fn encode_inline_small<R>(
self: Arc<Self>, self: Arc<Self>,
reader: R, reader: R,
@@ -813,7 +885,7 @@ impl Erasure {
/// Fast path for single-block non-inline objects: avoids the producer/consumer /// Fast path for single-block non-inline objects: avoids the producer/consumer
/// pipeline in `encode()` while keeping the same writer/quorum/shutdown semantics. /// pipeline in `encode()` while keeping the same writer/quorum/shutdown semantics.
#[hotpath::measure] #[hotpath::measure(impl_type = "Erasure")]
pub async fn encode_single_block_non_inline<R>( pub async fn encode_single_block_non_inline<R>(
self: Arc<Self>, self: Arc<Self>,
reader: R, reader: R,
@@ -839,7 +911,104 @@ mod tests {
use std::sync::{Arc, Mutex}; use std::sync::{Arc, Mutex};
use std::task::{Context, Poll}; use std::task::{Context, Poll};
use std::time::Duration; use std::time::Duration;
use tokio::io::{AsyncWrite, AsyncWriteExt}; use tokio::io::{AsyncWrite, AsyncWriteExt, ReadBuf};
use tokio::sync::oneshot;
struct PendingReader {
entered: Option<oneshot::Sender<()>>,
dropped: Option<oneshot::Sender<()>>,
}
impl PendingReader {
fn new() -> (Self, oneshot::Receiver<()>, oneshot::Receiver<()>) {
let (entered_tx, entered_rx) = oneshot::channel();
let (dropped_tx, dropped_rx) = oneshot::channel();
(
Self {
entered: Some(entered_tx),
dropped: Some(dropped_tx),
},
entered_rx,
dropped_rx,
)
}
}
impl AsyncRead for PendingReader {
fn poll_read(mut self: Pin<&mut Self>, _cx: &mut Context<'_>, _buf: &mut ReadBuf<'_>) -> Poll<std::io::Result<()>> {
if let Some(entered) = self.entered.take() {
let _ = entered.send(());
}
Poll::Pending
}
}
impl Drop for PendingReader {
fn drop(&mut self) {
if let Some(dropped) = self.dropped.take() {
let _ = dropped.send(());
}
}
}
struct BlocksThenPendingReader {
blocks_remaining: usize,
block: Vec<u8>,
blocked: Option<oneshot::Sender<()>>,
dropped: Option<oneshot::Sender<()>>,
final_block: Option<oneshot::Sender<()>>,
}
impl BlocksThenPendingReader {
fn new(
blocks_remaining: usize,
block_size: usize,
) -> (Self, oneshot::Receiver<()>, oneshot::Receiver<()>, oneshot::Receiver<()>) {
let (blocked_tx, blocked_rx) = oneshot::channel();
let (dropped_tx, dropped_rx) = oneshot::channel();
let (final_block_tx, final_block_rx) = oneshot::channel();
(
Self {
blocks_remaining,
block: vec![0x5a; block_size],
blocked: Some(blocked_tx),
dropped: Some(dropped_tx),
final_block: Some(final_block_tx),
},
blocked_rx,
dropped_rx,
final_block_rx,
)
}
}
impl AsyncRead for BlocksThenPendingReader {
fn poll_read(mut self: Pin<&mut Self>, _cx: &mut Context<'_>, buf: &mut ReadBuf<'_>) -> Poll<std::io::Result<()>> {
if self.blocks_remaining == 0 {
if let Some(blocked) = self.blocked.take() {
let _ = blocked.send(());
}
return Poll::Pending;
}
if self.blocks_remaining == 1
&& let Some(final_block) = self.final_block.take()
{
let _ = final_block.send(());
}
self.blocks_remaining -= 1;
buf.put_slice(&self.block);
Poll::Ready(Ok(()))
}
}
impl Drop for BlocksThenPendingReader {
fn drop(&mut self) {
if let Some(dropped) = self.dropped.take() {
let _ = dropped.send(());
}
}
}
fn erasure_with_zero_block_size() -> Erasure { fn erasure_with_zero_block_size() -> Erasure {
let mut erasure = Erasure::default(); let mut erasure = Erasure::default();
@@ -897,6 +1066,54 @@ mod tests {
} }
} }
struct FailAfterReaderBlocksWriter {
reader_blocked: oneshot::Receiver<()>,
writes: Arc<std::sync::atomic::AtomicUsize>,
}
impl AsyncWrite for FailAfterReaderBlocksWriter {
fn poll_write(mut self: Pin<&mut Self>, cx: &mut Context<'_>, _buf: &[u8]) -> Poll<std::io::Result<usize>> {
match Pin::new(&mut self.reader_blocked).poll(cx) {
Poll::Pending => Poll::Pending,
Poll::Ready(_) => {
self.writes.fetch_add(1, std::sync::atomic::Ordering::SeqCst);
Poll::Ready(Err(std::io::Error::other("injected write failure after producer blocks")))
}
}
}
fn poll_flush(self: Pin<&mut Self>, _cx: &mut Context<'_>) -> Poll<std::io::Result<()>> {
Poll::Ready(Ok(()))
}
fn poll_shutdown(self: Pin<&mut Self>, _cx: &mut Context<'_>) -> Poll<std::io::Result<()>> {
Poll::Ready(Ok(()))
}
}
struct StallOnWriteWithSignal {
entered: Option<oneshot::Sender<()>>,
writes: Arc<std::sync::atomic::AtomicUsize>,
}
impl AsyncWrite for StallOnWriteWithSignal {
fn poll_write(mut self: Pin<&mut Self>, _cx: &mut Context<'_>, _buf: &[u8]) -> Poll<std::io::Result<usize>> {
self.writes.fetch_add(1, std::sync::atomic::Ordering::SeqCst);
if let Some(entered) = self.entered.take() {
let _ = entered.send(());
}
Poll::Pending
}
fn poll_flush(self: Pin<&mut Self>, _cx: &mut Context<'_>) -> Poll<std::io::Result<()>> {
Poll::Ready(Ok(()))
}
fn poll_shutdown(self: Pin<&mut Self>, _cx: &mut Context<'_>) -> Poll<std::io::Result<()>> {
Poll::Ready(Ok(()))
}
}
#[derive(Clone, Default)] #[derive(Clone, Default)]
struct ShortWriteWriter; struct ShortWriteWriter;
@@ -1046,6 +1263,262 @@ mod tests {
BitrotWriterWrapper::new(CustomWriter::new_tokio_writer(writer), shard_size, HashAlgorithm::None) BitrotWriterWrapper::new(CustomWriter::new_tokio_writer(writer), shard_size, HashAlgorithm::None)
} }
#[derive(Clone, Copy)]
enum EncodePipeline {
Vec,
BytesMut,
Batched,
}
async fn aborting_encode_drops_blocked_producer(pipeline: EncodePipeline) {
const BLOCK_SIZE: usize = 16;
let gauge_baseline = rustfs_io_metrics::current_ec_encode_inflight_bytes();
let committed = Arc::new(Mutex::new(Vec::new()));
let mut writers = vec![Some(bitrot_writer(DeferredCommitWriter::new(committed.clone()), BLOCK_SIZE))];
let (reader, entered, dropped) = PendingReader::new();
let erasure = Arc::new(Erasure::new(1, 0, BLOCK_SIZE));
let encode = match pipeline {
EncodePipeline::Vec => {
tokio::spawn(async move { erasure.encode_with_ingest_mode(reader, &mut writers, 1, false).await })
}
EncodePipeline::BytesMut => {
tokio::spawn(async move { erasure.encode_with_ingest_mode(reader, &mut writers, 1, true).await })
}
EncodePipeline::Batched => tokio::spawn(async move { erasure.encode_batched(reader, &mut writers, 1).await }),
};
tokio::time::timeout(Duration::from_secs(1), entered)
.await
.expect("producer should enter the blocked reader before cancellation")
.expect("blocked reader should signal entry");
encode.abort();
assert!(matches!(encode.await, Err(err) if err.is_cancelled()), "encode task should be cancelled");
tokio::time::timeout(Duration::from_secs(1), dropped)
.await
.expect("cancelling encode should drop the producer reader")
.expect("blocked reader should signal producer drop");
assert!(
committed.lock().expect("committed buffer should be lockable").is_empty(),
"cancelling before the first encoded block must not make data visible"
);
assert_eq!(
rustfs_io_metrics::current_ec_encode_inflight_bytes(),
gauge_baseline,
"cancelling the encode pipeline must preserve the inflight queue gauge"
);
}
async fn writer_error_aborts_blocked_producer(pipeline: EncodePipeline) {
let gauge_baseline = rustfs_io_metrics::current_ec_encode_inflight_bytes();
let producer_baseline = rustfs_io_metrics::current_ec_encode_producer_bytes();
let writer_baseline = rustfs_io_metrics::current_ec_encode_writer_bytes();
let producer_peak_before = rustfs_io_metrics::current_ec_encode_producer_bytes_peak();
let writer_peak_before = rustfs_io_metrics::current_ec_encode_writer_bytes_peak();
let block_size = match pipeline {
EncodePipeline::Batched => usize::try_from(producer_peak_before.max(writer_peak_before).saturating_add(1))
.expect("stage peak fits the test address space"),
EncodePipeline::Vec | EncodePipeline::BytesMut => 16,
};
rustfs_io_metrics::set_put_stage_metrics_enabled(true);
let erasure = Arc::new(Erasure::new(1, 0, block_size));
let batch_blocks = encode_batch_block_count().min(encode_channel_capacity(
erasure.shard_size().saturating_mul(erasure.total_shard_count()),
erasure_encode_max_inflight_bytes(),
));
let blocks_before_pending = match pipeline {
EncodePipeline::Batched => batch_blocks,
EncodePipeline::Vec | EncodePipeline::BytesMut => 1,
};
let (reader, reader_blocked, reader_dropped, _final_block) =
BlocksThenPendingReader::new(blocks_before_pending, block_size);
let writes = Arc::new(std::sync::atomic::AtomicUsize::new(0));
let mut writers = vec![Some(bitrot_writer(
FailAfterReaderBlocksWriter {
reader_blocked,
writes: writes.clone(),
},
block_size,
))];
let result = match pipeline {
EncodePipeline::Vec => erasure.encode_with_ingest_mode(reader, &mut writers, 1, false).await,
EncodePipeline::BytesMut => erasure.encode_with_ingest_mode(reader, &mut writers, 1, true).await,
EncodePipeline::Batched => erasure.encode_batched(reader, &mut writers, 1).await,
};
let err = match result {
Ok(_) => panic!("writer quorum failure should fail the encode pipeline"),
Err(err) => err,
};
assert!(err.to_string().contains("Failed to write data"));
tokio::time::timeout(Duration::from_secs(1), reader_dropped)
.await
.expect("writer failure should abort the blocked producer")
.expect("blocked producer should signal reader drop");
assert_eq!(
writes.load(std::sync::atomic::Ordering::SeqCst),
1,
"writer failure must stop the pipeline before any additional shard write"
);
assert_eq!(
rustfs_io_metrics::current_ec_encode_inflight_bytes(),
gauge_baseline,
"writer failure must settle all queued and pending encoded bytes"
);
assert_eq!(
rustfs_io_metrics::current_ec_encode_producer_bytes(),
producer_baseline,
"writer failure must settle producer stage bytes for every ingest pipeline"
);
assert_eq!(
rustfs_io_metrics::current_ec_encode_writer_bytes(),
writer_baseline,
"writer failure must settle writer stage bytes for every ingest pipeline"
);
if matches!(pipeline, EncodePipeline::Batched) {
let expected_batch_bytes = u64::try_from(block_size)
.expect("block size fits the stage gauge")
.checked_mul(u64::try_from(batch_blocks).expect("batch block count fits the stage gauge"))
.expect("test batch payload fits the stage gauge");
assert!(
rustfs_io_metrics::current_ec_encode_producer_bytes_peak() >= producer_peak_before.max(expected_batch_bytes),
"batched producer must expose its full pending batch before writer failure"
);
assert!(
rustfs_io_metrics::current_ec_encode_writer_bytes_peak() >= writer_peak_before.max(expected_batch_bytes),
"batched writer must expose its full batch before writer failure"
);
}
rustfs_io_metrics::set_put_stage_metrics_enabled(false);
}
async fn aborting_full_queue_settles_pending_send() {
const BLOCK_SIZE: usize = 16;
let gauge_baseline = rustfs_io_metrics::current_ec_encode_inflight_bytes();
let producer_baseline = rustfs_io_metrics::current_ec_encode_producer_bytes();
let writer_baseline = rustfs_io_metrics::current_ec_encode_writer_bytes();
rustfs_io_metrics::set_put_stage_metrics_enabled(true);
let erasure = Arc::new(Erasure::new(1, 0, BLOCK_SIZE));
let inflight_blocks = encode_channel_capacity(
erasure.shard_size().saturating_mul(erasure.total_shard_count()),
erasure_encode_max_inflight_bytes(),
);
let (reader, _reader_blocked, reader_dropped, final_block) =
BlocksThenPendingReader::new(inflight_blocks + 2, BLOCK_SIZE);
let (writer_entered_tx, writer_entered) = oneshot::channel();
let writes = Arc::new(std::sync::atomic::AtomicUsize::new(0));
let mut writers = vec![Some(bitrot_writer(
StallOnWriteWithSignal {
entered: Some(writer_entered_tx),
writes: writes.clone(),
},
BLOCK_SIZE,
))];
let erasure_for_task = erasure.clone();
let encode = tokio::spawn(async move { erasure_for_task.encode_with_ingest_mode(reader, &mut writers, 1, false).await });
tokio::time::timeout(Duration::from_secs(1), writer_entered)
.await
.expect("consumer should start the first writer call")
.expect("stalling writer should signal entry");
tokio::time::timeout(Duration::from_secs(1), final_block)
.await
.expect("producer should supply the block whose send fills the queue")
.expect("reader should signal final block");
let expected_queued_bytes =
u64::try_from((inflight_blocks + 1) * BLOCK_SIZE).expect("queued byte count should fit the gauge");
tokio::time::timeout(Duration::from_secs(1), async {
while rustfs_io_metrics::current_ec_encode_inflight_bytes() < gauge_baseline + expected_queued_bytes {
tokio::task::yield_now().await;
}
})
.await
.expect("producer should account for the pending send after the queue fills");
tokio::time::timeout(Duration::from_secs(1), async {
while rustfs_io_metrics::current_ec_encode_producer_bytes() == producer_baseline
|| rustfs_io_metrics::current_ec_encode_writer_bytes() == writer_baseline
{
tokio::task::yield_now().await;
}
})
.await
.expect("full queue must retain producer and writer stage ownership before cancellation");
encode.abort();
assert!(matches!(encode.await, Err(err) if err.is_cancelled()), "encode task should be cancelled");
tokio::time::timeout(Duration::from_secs(1), reader_dropped)
.await
.expect("cancelling a full queue should abort its producer")
.expect("full-queue producer should signal reader drop");
assert_eq!(
writes.load(std::sync::atomic::Ordering::SeqCst),
1,
"cancellation must not resume the stalled writer"
);
assert_eq!(
rustfs_io_metrics::current_ec_encode_inflight_bytes(),
gauge_baseline,
"cancelling a full queue must settle queued and pending bytes"
);
assert_eq!(
rustfs_io_metrics::current_ec_encode_producer_bytes(),
producer_baseline,
"cancelling a full queue must settle the pending producer stage"
);
assert_eq!(
rustfs_io_metrics::current_ec_encode_writer_bytes(),
writer_baseline,
"cancelling a full queue must settle the stalled writer stage"
);
rustfs_io_metrics::set_put_stage_metrics_enabled(false);
}
#[tokio::test]
#[serial_test::serial]
async fn cancelling_vec_encode_drops_blocked_producer() {
aborting_encode_drops_blocked_producer(EncodePipeline::Vec).await;
}
#[tokio::test]
#[serial_test::serial]
async fn cancelling_bytesmut_encode_drops_blocked_producer() {
aborting_encode_drops_blocked_producer(EncodePipeline::BytesMut).await;
}
#[tokio::test]
#[serial_test::serial]
async fn cancelling_batched_encode_drops_blocked_producer() {
aborting_encode_drops_blocked_producer(EncodePipeline::Batched).await;
}
#[tokio::test]
#[serial_test::serial]
async fn vec_writer_error_aborts_blocked_producer() {
writer_error_aborts_blocked_producer(EncodePipeline::Vec).await;
}
#[tokio::test]
#[serial_test::serial]
async fn bytesmut_writer_error_aborts_blocked_producer() {
writer_error_aborts_blocked_producer(EncodePipeline::BytesMut).await;
}
#[tokio::test]
#[serial_test::serial]
async fn batched_writer_error_aborts_blocked_producer() {
writer_error_aborts_blocked_producer(EncodePipeline::Batched).await;
}
#[tokio::test]
#[serial_test::serial]
async fn cancelling_full_queue_settles_pending_send() {
aborting_full_queue_settles_pending_send().await;
}
#[tokio::test] #[tokio::test]
async fn helper_writers_cover_flush_and_shutdown_paths() { async fn helper_writers_cover_flush_and_shutdown_paths() {
let mut failing_write = FailingWriteWriter; let mut failing_write = FailingWriteWriter;
@@ -1329,25 +1802,99 @@ mod tests {
} }
#[tokio::test] #[tokio::test]
async fn drain_queued_inflight_bytes_consumes_pending_blocks() { #[serial_test::serial]
let (tx, mut rx) = mpsc::channel(2); async fn queued_inflight_bytes_are_settled_on_all_queue_exit_paths() {
tx.send(vec![Bytes::from_static(b"queued")]).await.unwrap(); let baseline = rustfs_io_metrics::current_ec_encode_inflight_bytes();
drop(tx); let (tx, rx) = mpsc::channel(1);
let queued = vec![Bytes::from_static(b"queued")];
let queued_bytes = queued_block_bytes(&queued);
drain_queued_inflight_bytes(&mut rx).await; send_queued(&tx, queued, queued_bytes)
.await
.expect("first queue entry should fit");
assert!(rx.recv().await.is_none()); let blocked = vec![Bytes::from_static(b"blocked")];
let blocked_bytes = queued_block_bytes(&blocked);
{
let pending_send = send_queued(&tx, blocked, blocked_bytes);
tokio::pin!(pending_send);
assert!(
futures::poll!(pending_send.as_mut()).is_pending(),
"full queue must suspend producer send"
);
}
assert_eq!(
rustfs_io_metrics::current_ec_encode_inflight_bytes(),
baseline + u64::try_from(queued_bytes).expect("queue bytes fit the gauge"),
"dropping a pending producer send must compensate its bytes"
);
drop(rx);
assert_eq!(
rustfs_io_metrics::current_ec_encode_inflight_bytes(),
baseline,
"dropping the receiver must settle every buffered entry"
);
let rejected = vec![Bytes::from_static(b"rejected")];
let rejected_bytes = queued_block_bytes(&rejected);
assert!(
send_queued(&tx, rejected, rejected_bytes).await.is_err(),
"closed receiver must reject a new send"
);
assert_eq!(
rustfs_io_metrics::current_ec_encode_inflight_bytes(),
baseline,
"failed sends must compensate their bytes"
);
} }
#[tokio::test] #[tokio::test]
async fn drain_queued_batched_inflight_bytes_consumes_pending_batches() { #[serial_test::serial]
let (tx, mut rx) = mpsc::channel(2); async fn queued_batch_entry_settles_bytes_before_handoff() {
tx.send(vec![vec![Bytes::from_static(b"queued")]]).await.unwrap(); let baseline = rustfs_io_metrics::current_ec_encode_inflight_bytes();
drop(tx); let (tx, rx) = mpsc::channel(2);
let mut rx = rx;
let batch = vec![vec![Bytes::from_static(b"queued")], vec![Bytes::from_static(b"batch")]];
let batch_bytes = queued_batch_bytes(&batch);
drain_queued_batched_inflight_bytes(&mut rx).await; send_queued(&tx, batch, batch_bytes).await.expect("batch should be queued");
let batch = rx.recv().await.expect("queued batch should be received").into_inner();
assert_eq!(batch_bytes, queued_batch_bytes(&batch));
assert!(rx.recv().await.is_none()); assert_eq!(
rustfs_io_metrics::current_ec_encode_inflight_bytes(),
baseline,
"receiving a batch must settle all contained block bytes before shard writes"
);
}
#[tokio::test]
#[serial_test::serial]
async fn queued_entry_settles_when_a_closed_receiver_accepts_an_outstanding_permit() {
let baseline = rustfs_io_metrics::current_ec_encode_inflight_bytes();
let (tx, mut rx) = mpsc::channel(1);
let permit = tx
.clone()
.reserve_owned()
.await
.expect("open receiver should reserve queue capacity");
rx.close();
assert!(
matches!(rx.try_recv(), Err(mpsc::error::TryRecvError::Empty)),
"an outstanding permit must leave the closed queue observably empty"
);
let block = vec![Bytes::from_static(b"late-permit")];
let block_bytes = queued_block_bytes(&block);
permit.send(InflightEntry::new(block, block_bytes));
drop(rx);
assert_eq!(
rustfs_io_metrics::current_ec_encode_inflight_bytes(),
baseline,
"a queue entry sent through an outstanding permit must settle when Tokio drops it"
);
} }
#[tokio::test] #[tokio::test]
@@ -1394,6 +1941,61 @@ mod tests {
} }
} }
#[tokio::test]
#[serial_test::serial]
async fn bytesmut_streaming_encode_observes_all_payload_stage_peaks() {
let queue_baseline = rustfs_io_metrics::current_ec_encode_inflight_bytes();
let producer_peak_before = rustfs_io_metrics::current_ec_encode_producer_bytes_peak();
let queue_peak_before = rustfs_io_metrics::current_ec_encode_queue_bytes_peak();
let writer_peak_before = rustfs_io_metrics::current_ec_encode_writer_bytes_peak();
let prior_peak = producer_peak_before.max(queue_peak_before).max(writer_peak_before);
let block_size = usize::try_from(prior_peak.saturating_add(1)).expect("stage peak fits the test address space");
let erasure = Arc::new(Erasure::new(1, 0, block_size));
let encoded_block_bytes = erasure.shard_size() * erasure.total_shard_count();
let committed = Arc::new(Mutex::new(Vec::new()));
let mut writers = vec![Some(bitrot_writer(DeferredCommitWriter::new(committed.clone()), block_size))];
let reader = tokio::io::BufReader::new(Cursor::new(vec![0x5a; block_size * 2]));
rustfs_io_metrics::set_put_stage_metrics_enabled(true);
let result = erasure.encode_with_ingest_mode(reader, &mut writers, 1, true).await;
rustfs_io_metrics::set_put_stage_metrics_enabled(false);
let (_reader, written) = result.expect("bytesmut streaming encode should complete");
assert_eq!(written, block_size * 2);
assert_eq!(
rustfs_io_metrics::current_ec_encode_inflight_bytes(),
queue_baseline,
"completed streaming encode must not retain queue bytes"
);
assert_eq!(
rustfs_io_metrics::current_ec_encode_producer_bytes(),
0,
"completed streaming encode must not retain producer bytes"
);
assert_eq!(
rustfs_io_metrics::current_ec_encode_writer_bytes(),
0,
"completed streaming encode must not retain writer bytes"
);
let expected_peak = u64::try_from(encoded_block_bytes).expect("encoded block bytes fit the gauge");
assert!(
rustfs_io_metrics::current_ec_encode_producer_bytes_peak() >= producer_peak_before.max(expected_peak),
"producer peak must observe encoded bytes before queue hand-off"
);
assert!(
rustfs_io_metrics::current_ec_encode_queue_bytes_peak() >= queue_peak_before.max(expected_peak),
"queue peak must observe encoded bytes pending shard writers"
);
assert!(
rustfs_io_metrics::current_ec_encode_writer_bytes_peak() >= writer_peak_before.max(expected_peak),
"writer peak must observe encoded bytes after queue hand-off"
);
assert!(
!committed.lock().expect("committed buffer should be lockable").is_empty(),
"stage peak observation must not change writer commit behavior"
);
}
#[tokio::test] #[tokio::test]
async fn encode_streaming_write_quorum_failure_aborts_and_reports_error() { async fn encode_streaming_write_quorum_failure_aborts_and_reports_error() {
const DATA_SHARDS: usize = 2; const DATA_SHARDS: usize = 2;
+5 -5
View File
@@ -640,7 +640,7 @@ impl Erasure {
/// # Returns /// # Returns
/// A vector of encoded shards as `Bytes`. /// A vector of encoded shards as `Bytes`.
#[tracing::instrument(level = "debug", skip_all, fields(data_len=data.len()))] #[tracing::instrument(level = "debug", skip_all, fields(data_len=data.len()))]
#[hotpath::measure] #[hotpath::measure(impl_type = "Erasure")]
pub fn encode_data(&self, data: &[u8]) -> io::Result<Vec<Bytes>> { pub fn encode_data(&self, data: &[u8]) -> io::Result<Vec<Bytes>> {
let shard_size_fn = if self.uses_legacy { let shard_size_fn = if self.uses_legacy {
calc_shard_size_legacy calc_shard_size_legacy
@@ -688,7 +688,7 @@ impl Erasure {
/// Encode owned data, avoiding a copy when the caller already has a heap buffer. /// Encode owned data, avoiding a copy when the caller already has a heap buffer.
/// Falls back to copying into a new buffer if zero-copy conversion fails. /// Falls back to copying into a new buffer if zero-copy conversion fails.
#[hotpath::measure] #[hotpath::measure(impl_type = "Erasure")]
pub fn encode_data_owned(&self, data: Vec<u8>) -> io::Result<Vec<Bytes>> { pub fn encode_data_owned(&self, data: Vec<u8>) -> io::Result<Vec<Bytes>> {
let shard_size_fn = if self.uses_legacy { let shard_size_fn = if self.uses_legacy {
calc_shard_size_legacy calc_shard_size_legacy
@@ -752,7 +752,7 @@ impl Erasure {
/// block), the `resize(need_total_size)` below stays within capacity for every /// block), the `resize(need_total_size)` below stays within capacity for every
/// `data_len <= block_size` — both shard-size formulas are monotone in /// `data_len <= block_size` — both shard-size formulas are monotone in
/// `data_len` — so this function never reallocates the buffer. /// `data_len` — so this function never reallocates the buffer.
#[hotpath::measure] #[hotpath::measure(impl_type = "Erasure")]
pub fn encode_data_bytes_mut(&self, mut data_buffer: BytesMut, data_len: usize) -> io::Result<Vec<Bytes>> { pub fn encode_data_bytes_mut(&self, mut data_buffer: BytesMut, data_len: usize) -> io::Result<Vec<Bytes>> {
let shard_size_fn = if self.uses_legacy { let shard_size_fn = if self.uses_legacy {
calc_shard_size_legacy calc_shard_size_legacy
@@ -805,7 +805,7 @@ impl Erasure {
/// ///
/// # Returns /// # Returns
/// Ok if reconstruction succeeds, error otherwise. /// Ok if reconstruction succeeds, error otherwise.
#[hotpath::measure] #[hotpath::measure(impl_type = "Erasure")]
pub fn decode_data(&self, shards: &mut [Option<Vec<u8>>]) -> io::Result<()> { pub fn decode_data(&self, shards: &mut [Option<Vec<u8>>]) -> io::Result<()> {
if self.parity_shards > 0 { if self.parity_shards > 0 {
if self.uses_legacy { if self.uses_legacy {
@@ -825,7 +825,7 @@ impl Erasure {
} }
/// Decode and reconstruct missing data shards, then regenerate parity shards. /// Decode and reconstruct missing data shards, then regenerate parity shards.
#[hotpath::measure] #[hotpath::measure(impl_type = "Erasure")]
pub fn decode_data_and_parity(&self, shards: &mut [Option<Vec<u8>>]) -> io::Result<()> { pub fn decode_data_and_parity(&self, shards: &mut [Option<Vec<u8>>]) -> io::Result<()> {
if self.parity_shards > 0 { if self.parity_shards > 0 {
if self.uses_legacy { if self.uses_legacy {
+468 -20
View File
@@ -19,6 +19,8 @@ use crate::diagnostics::get::{
GET_STAGE_READER_MMAP_PATH_RESOLVE, GET_STAGE_READER_OPEN_MMAP_COPY_FALLBACK, GET_STAGE_READER_OPEN_MMAP_COPY_SUCCESS, GET_STAGE_READER_MMAP_PATH_RESOLVE, GET_STAGE_READER_OPEN_MMAP_COPY_FALLBACK, GET_STAGE_READER_OPEN_MMAP_COPY_SUCCESS,
GET_STAGE_READER_OPEN_STREAM, GET_STAGE_READER_STREAM_FIRST_READ, record_get_stage_duration_if_enabled, GET_STAGE_READER_OPEN_STREAM, GET_STAGE_READER_STREAM_FIRST_READ, record_get_stage_duration_if_enabled,
}; };
#[cfg(feature = "hotpath")]
use crate::disk::FileWriter;
use crate::disk::{self, DiskAPI as _, DiskStore, FileReader, MmapCopyStageMetrics, error::DiskError}; use crate::disk::{self, DiskAPI as _, DiskStore, FileReader, MmapCopyStageMetrics, error::DiskError};
use crate::erasure::coding::{BitrotReader, BitrotWriterWrapper, CustomWriter}; use crate::erasure::coding::{BitrotReader, BitrotWriterWrapper, CustomWriter};
use bytes::Bytes; use bytes::Bytes;
@@ -36,6 +38,11 @@ use std::time::Instant;
use tokio::io::{AsyncRead, ReadBuf}; use tokio::io::{AsyncRead, ReadBuf};
use tracing::debug; use tracing::debug;
#[cfg(all(test, feature = "hotpath"))]
tokio::task_local! {
static FORCE_MMAP_COPY_FAILURE_FOR_TEST: ();
}
/// A shard source for the bitrot reader. /// A shard source for the bitrot reader.
/// ///
/// `InMemory` keeps the `Bytes` concrete instead of erasing it behind /// `InMemory` keeps the `Bytes` concrete instead of erasing it behind
@@ -360,10 +367,21 @@ async fn open_disk_reader(
mmap_copy_stage: GET_STAGE_READER_MMAP_COPY_BUFFER, mmap_copy_stage: GET_STAGE_READER_MMAP_COPY_BUFFER,
direct_read_copy_stage: GET_STAGE_READER_MMAP_DIRECT_READ_COPY, direct_read_copy_stage: GET_STAGE_READER_MMAP_DIRECT_READ_COPY,
}); });
match disk let mmap_result = {
.read_file_mmap_copy_with_metrics(bucket, path, offset, length, mmap_metrics) #[cfg(all(test, feature = "hotpath"))]
if FORCE_MMAP_COPY_FAILURE_FOR_TEST.try_with(|_| ()).is_ok() {
Err(DiskError::other("forced mmap-copy failure for test"))
} else {
disk.read_file_mmap_copy_with_metrics(bucket, path, offset, length, mmap_metrics)
.await .await
}
#[cfg(not(all(test, feature = "hotpath")))]
{ {
disk.read_file_mmap_copy_with_metrics(bucket, path, offset, length, mmap_metrics)
.await
}
};
match mmap_result {
Ok(bytes) => { Ok(bytes) => {
let duration_ms = zero_copy_start.elapsed().as_secs_f64() * 1000.0; let duration_ms = zero_copy_start.elapsed().as_secs_f64() * 1000.0;
@@ -398,7 +416,11 @@ async fn open_disk_reader(
} }
return match stream_result { return match stream_result {
Ok(reader) => Ok(wrap_first_read_metrics(reader, metrics_path)), Ok(reader) => {
#[cfg(feature = "hotpath")]
let reader = instrument_raw_shard_reader(reader, disk.is_local());
Ok(wrap_first_read_metrics(reader, metrics_path))
}
Err(_) => Err(err), Err(_) => Err(err),
}; };
} }
@@ -407,6 +429,8 @@ async fn open_disk_reader(
let stream_start = stage_metrics_enabled.then(Instant::now); let stream_start = stage_metrics_enabled.then(Instant::now);
let reader = disk.read_file_stream(bucket, path, offset, length).await?; let reader = disk.read_file_stream(bucket, path, offset, length).await?;
#[cfg(feature = "hotpath")]
let reader = instrument_raw_shard_reader(reader, disk.is_local());
if let Some(metrics_path) = metrics_path { if let Some(metrics_path) = metrics_path {
record_get_stage_duration_if_enabled(metrics_path, GET_STAGE_READER_OPEN_STREAM, stream_start); record_get_stage_duration_if_enabled(metrics_path, GET_STAGE_READER_OPEN_STREAM, stream_start);
} }
@@ -427,6 +451,37 @@ fn wrap_first_read_metrics(reader: FileReader, metrics_path: Option<&'static str
ShardReader::Stream(reader) ShardReader::Stream(reader)
} }
// The labels are deliberately fixed: object keys, disk paths, and remote hosts
// are all high-cardinality or sensitive and belong nowhere in a profiling report.
#[cfg(feature = "hotpath")]
const RAW_SHARD_READ_LOCAL_LABEL: &str = "EC raw shard read local";
#[cfg(feature = "hotpath")]
const RAW_SHARD_READ_REMOTE_LABEL: &str = "EC raw shard read remote";
#[cfg(feature = "hotpath")]
const RAW_SHARD_WRITE_LOCAL_LABEL: &str = "EC raw shard write local";
#[cfg(feature = "hotpath")]
const RAW_SHARD_WRITE_REMOTE_LABEL: &str = "EC raw shard write remote";
#[cfg(feature = "hotpath")]
fn instrument_raw_shard_reader(reader: FileReader, is_local: bool) -> FileReader {
// `io!` aggregates by call site, so the local and remote branches must remain distinct.
if is_local {
Box::new(hotpath::io!(reader, label = RAW_SHARD_READ_LOCAL_LABEL))
} else {
Box::new(hotpath::io!(reader, label = RAW_SHARD_READ_REMOTE_LABEL))
}
}
#[cfg(feature = "hotpath")]
fn instrument_raw_shard_writer(writer: FileWriter, is_local: bool) -> FileWriter {
// `io!` aggregates by call site, so the local and remote branches must remain distinct.
if is_local {
Box::new(hotpath::io!(writer, label = RAW_SHARD_WRITE_LOCAL_LABEL))
} else {
Box::new(hotpath::io!(writer, label = RAW_SHARD_WRITE_REMOTE_LABEL))
}
}
fn bitrot_encoded_range(offset: usize, length: usize, shard_size: usize, checksum_algo: HashAlgorithm) -> (usize, usize) { fn bitrot_encoded_range(offset: usize, length: usize, shard_size: usize, checksum_algo: HashAlgorithm) -> (usize, usize) {
( (
offset.div_ceil(shard_size) * checksum_algo.size() + offset, offset.div_ceil(shard_size) * checksum_algo.size() + offset,
@@ -686,6 +741,8 @@ pub async fn create_bitrot_writer(
}; };
let file = disk.create_file("", volume, path, length).await?; let file = disk.create_file("", volume, path, length).await?;
#[cfg(feature = "hotpath")]
let file = instrument_raw_shard_writer(file, disk.is_local());
CustomWriter::new_tokio_writer(file) CustomWriter::new_tokio_writer(file)
} else { } else {
return Err(DiskError::DiskNotFound); return Err(DiskError::DiskNotFound);
@@ -698,6 +755,194 @@ pub async fn create_bitrot_writer(
mod tests { mod tests {
use super::*; use super::*;
#[cfg(feature = "hotpath")]
use crate::cluster::rpc::RemoteDisk;
#[cfg(feature = "hotpath")]
use crate::cluster::rpc::internode_data_transport::{
InternodeDataTransport, InternodeDataTransportCapabilities, ReadStreamRequest, WalkDirStreamRequest, WriteStreamRequest,
};
#[cfg(feature = "hotpath")]
use crate::disk::{Disk, DiskOption, error::Result};
#[cfg(feature = "hotpath")]
#[derive(Debug, Clone, Default)]
struct TestRemoteDataTransport {
bytes: Arc<Mutex<Vec<u8>>>,
write_error: Option<io::ErrorKind>,
}
#[cfg(feature = "hotpath")]
impl TestRemoteDataTransport {
fn with_write_error(kind: io::ErrorKind) -> Self {
Self {
bytes: Arc::default(),
write_error: Some(kind),
}
}
fn bytes(&self) -> Vec<u8> {
self.bytes
.lock()
.expect("test remote transport bytes lock should not be poisoned")
.clone()
}
}
#[cfg(feature = "hotpath")]
#[derive(Debug)]
struct TestRemoteWriter {
bytes: Arc<Mutex<Vec<u8>>>,
write_error: Option<io::ErrorKind>,
}
#[cfg(feature = "hotpath")]
impl tokio::io::AsyncWrite for TestRemoteWriter {
fn poll_write(self: Pin<&mut Self>, _cx: &mut Context<'_>, buf: &[u8]) -> Poll<io::Result<usize>> {
if let Some(kind) = self.write_error {
return Poll::Ready(Err(io::Error::from(kind)));
}
self.bytes
.lock()
.expect("test remote transport bytes lock should not be poisoned")
.extend_from_slice(buf);
Poll::Ready(Ok(buf.len()))
}
fn poll_flush(self: Pin<&mut Self>, _cx: &mut Context<'_>) -> Poll<io::Result<()>> {
Poll::Ready(Ok(()))
}
fn poll_shutdown(self: Pin<&mut Self>, _cx: &mut Context<'_>) -> Poll<io::Result<()>> {
Poll::Ready(Ok(()))
}
}
#[cfg(feature = "hotpath")]
#[async_trait::async_trait]
impl InternodeDataTransport for TestRemoteDataTransport {
async fn open_read(&self, _request: ReadStreamRequest) -> Result<FileReader> {
Ok(Box::new(Cursor::new(self.bytes())))
}
async fn open_write(&self, _request: WriteStreamRequest) -> Result<FileWriter> {
Ok(Box::new(TestRemoteWriter {
bytes: Arc::clone(&self.bytes),
write_error: self.write_error,
}))
}
async fn open_walk_dir(&self, _request: WalkDirStreamRequest) -> Result<FileReader> {
panic!("open_walk_dir must not be used by the raw shard I/O test")
}
fn name(&self) -> &'static str {
"bitrot-test-remote"
}
fn capabilities(&self) -> InternodeDataTransportCapabilities {
InternodeDataTransportCapabilities::tcp_http()
}
}
async fn local_test_disk() -> (DiskStore, tempfile::TempDir) {
use crate::disk::endpoint::Endpoint;
use crate::disk::{DiskOption, new_disk};
let dir = tempfile::tempdir().expect("tempdir should be created");
let mut endpoint =
Endpoint::try_from(dir.path().to_str().expect("tempdir path should be utf8")).expect("endpoint should parse");
endpoint.set_pool_index(0);
endpoint.set_set_index(0);
endpoint.set_disk_index(0);
let disk = new_disk(
&endpoint,
&DiskOption {
cleanup: false,
health_check: false,
},
)
.await
.expect("local disk should be created");
(disk, dir)
}
#[cfg(feature = "hotpath")]
async fn remote_test_disk(transport: TestRemoteDataTransport) -> (DiskStore, TestRemoteDataTransport) {
use crate::disk::endpoint::Endpoint;
let endpoint = Endpoint {
url: url::Url::parse("http://remote-node:9000/data/rustfs0").expect("test remote endpoint should parse"),
is_local: false,
pool_idx: 0,
set_idx: 0,
disk_idx: 0,
};
let remote = RemoteDisk::new(
&endpoint,
&DiskOption {
cleanup: false,
health_check: false,
},
Arc::new(transport.clone()),
)
.await
.expect("test remote disk should be created");
(Arc::new(Disk::Remote(Box::new(remote))), transport)
}
async fn round_trip_disk_bitrot(disk: &DiskStore, bucket: &str, path: &str, payload: &[u8], shard_size: usize) -> Vec<u8> {
disk.make_volume(bucket).await.expect("volume should be created");
let mut writer = create_bitrot_writer(
false,
Some(disk),
bucket,
path,
i64::try_from(payload.len()).expect("test payload length should fit i64"),
shard_size,
HashAlgorithm::None,
)
.await
.expect("disk bitrot writer should open the raw shard file");
for chunk in payload.chunks(shard_size) {
writer
.write(chunk)
.await
.expect("disk bitrot writer should preserve each shard block");
}
writer.shutdown().await.expect("disk bitrot writer should close cleanly");
let mut reader = create_bitrot_reader(
None,
Some(disk),
bucket,
path,
0,
payload.len(),
shard_size,
HashAlgorithm::None,
false,
false,
)
.await
.expect("disk bitrot reader should open the raw shard file")
.expect("disk bitrot reader should exist");
let mut actual = Vec::with_capacity(payload.len());
while actual.len() < payload.len() {
let remaining = payload.len() - actual.len();
let mut chunk = vec![0; remaining.min(shard_size)];
let read = reader
.read(&mut chunk)
.await
.expect("disk bitrot reader should return the complete shard body");
assert!(read > 0, "disk bitrot reader must not end before the expected shard body is complete");
actual.extend_from_slice(&chunk[..read]);
}
actual
}
#[test] #[test]
fn object_mmap_read_enabled_accepts_legacy_zero_copy_alias() { fn object_mmap_read_enabled_accepts_legacy_zero_copy_alias() {
temp_env::with_vars( temp_env::with_vars(
@@ -724,6 +969,214 @@ mod tests {
); );
} }
#[cfg(feature = "hotpath")]
#[test]
fn raw_shard_io_wrappers_report_fixed_labels_and_preserve_bytes() {
const CHILD_ENV: &str = "RUSTFS_HOTPATH_RAW_SHARD_IO_TEST_CHILD";
if std::env::var_os(CHILD_ENV).is_none() {
let status = std::process::Command::new(std::env::current_exe().expect("test executable path should be available"))
.arg("--exact")
.arg("io_support::bitrot::tests::raw_shard_io_wrappers_report_fixed_labels_and_preserve_bytes")
.arg("--nocapture")
.env(CHILD_ENV, "1")
.status()
.expect("isolated HotPath I/O test process should start");
assert!(status.success(), "isolated HotPath I/O test process should pass");
return;
}
tokio::runtime::Builder::new_current_thread()
.enable_all()
.build()
.expect("test runtime should be created")
.block_on(raw_shard_io_wrappers_report_fixed_labels_and_preserve_bytes_in_isolated_process());
}
#[cfg(feature = "hotpath")]
async fn raw_shard_io_wrappers_report_fixed_labels_and_preserve_bytes_in_isolated_process() {
use hotpath::{Format, HotpathGuardBuilder, Section};
use tokio::io::AsyncReadExt;
let report_dir = tempfile::tempdir().expect("report tempdir should be created");
let report_path = report_dir.path().join("hotpath-io.json");
let guard = HotpathGuardBuilder::new("raw_shard_io_test")
.format(Format::Json)
.output_path(&report_path)
.sections(vec![Section::Io])
.build();
let (disk, _dir) = local_test_disk().await;
let bucket = "test-bucket";
let path = "obj/hotpath-part.1";
let payload = b"local shard bytes";
let shard_size = 4;
let local_read = round_trip_disk_bitrot(&disk, bucket, path, payload, shard_size).await;
assert_eq!(local_read, payload, "local raw shard I/O instrumentation must not alter stored bytes");
let fallback_path = "obj/hotpath-mmap-fallback-part.1";
let fallback_payload = b"mmap fallback shard bytes";
disk.write_all("test-bucket", fallback_path, Bytes::from_static(fallback_payload))
.await
.expect("fallback shard file should be written");
let mut fallback_reader = FORCE_MMAP_COPY_FAILURE_FOR_TEST
.scope((), open_disk_reader(&disk, bucket, fallback_path, 0, fallback_payload.len(), true, None))
.await
.expect("mmap-copy failure should fall back to a raw shard stream");
assert!(
matches!(fallback_reader, ShardReader::Stream(_)),
"forced mmap-copy failure must use the streaming fallback"
);
let mut fallback_read = Vec::new();
fallback_reader
.read_to_end(&mut fallback_read)
.await
.expect("mmap-copy fallback stream should preserve bytes");
assert_eq!(fallback_read, fallback_payload);
let (remote_disk, remote_transport) = remote_test_disk(TestRemoteDataTransport::default()).await;
let remote_payload = b"remote shard bytes";
let mut remote_writer = create_bitrot_writer(
false,
Some(&remote_disk),
bucket,
"obj/hotpath-remote-part.1",
i64::try_from(remote_payload.len()).expect("remote payload length should fit i64"),
shard_size,
HashAlgorithm::None,
)
.await
.expect("remote bitrot writer should use the production raw writer path");
for chunk in remote_payload.chunks(shard_size) {
remote_writer
.write(chunk)
.await
.expect("remote bitrot writer should preserve bytes");
}
remote_writer
.shutdown()
.await
.expect("remote bitrot writer should close cleanly");
assert_eq!(remote_transport.bytes(), remote_payload);
let mut reader = create_bitrot_reader(
None,
Some(&remote_disk),
bucket,
"obj/hotpath-remote-part.1",
0,
remote_payload.len(),
shard_size,
HashAlgorithm::None,
false,
true,
)
.await
.expect("remote bitrot reader should use the production raw reader path")
.expect("remote bitrot reader should exist");
let mut remote_read = Vec::with_capacity(remote_payload.len());
while remote_read.len() < remote_payload.len() {
let remaining = remote_payload.len() - remote_read.len();
let mut chunk = vec![0; remaining.min(shard_size)];
let read = reader
.read(&mut chunk)
.await
.expect("remote bitrot reader should preserve bytes");
assert!(read > 0, "remote bitrot reader must not end before the expected shard body is complete");
remote_read.extend_from_slice(&chunk[..read]);
}
assert_eq!(remote_read, remote_payload);
let (failing_remote_disk, _) = remote_test_disk(TestRemoteDataTransport::with_write_error(io::ErrorKind::Other)).await;
let mut failing_writer = create_bitrot_writer(
false,
Some(&failing_remote_disk),
bucket,
"obj/hotpath-remote-write-error-part.1",
4,
shard_size,
HashAlgorithm::None,
)
.await
.expect("failing remote bitrot writer should open before its first write");
let write_error = failing_writer
.write(b"fail")
.await
.expect_err("raw shard writer failures must remain visible through the bitrot writer");
assert_eq!(write_error.kind(), io::ErrorKind::Other);
let (would_block_remote_disk, _) =
remote_test_disk(TestRemoteDataTransport::with_write_error(io::ErrorKind::WouldBlock)).await;
let mut would_block_writer = create_bitrot_writer(
false,
Some(&would_block_remote_disk),
bucket,
"obj/hotpath-remote-write-would-block-part.1",
4,
shard_size,
HashAlgorithm::None,
)
.await
.expect("would-block remote bitrot writer should open before its first write");
let would_block = would_block_writer
.write(b"wait")
.await
.expect_err("would-block must remain visible to the caller");
assert_eq!(would_block.kind(), io::ErrorKind::WouldBlock);
drop(guard);
let report = std::fs::read_to_string(&report_path).expect("HotPath I/O report should be written");
let report: serde_json::Value = serde_json::from_str(&report).expect("HotPath I/O report should be valid JSON");
let entries = report["io"]["data"]
.as_array()
.expect("HotPath I/O report should include data rows");
let io_bytes = |label: &str, direction: &str| {
let byte_count: u64 = entries
.iter()
.filter(|entry| entry["label"].as_str().is_some_and(|entry_label| entry_label == label))
.filter_map(|entry| entry[direction]["bytes"].as_u64())
.sum();
assert!(byte_count > 0, "report must include fixed label {label}");
byte_count
};
let io_errors = |label: &str, direction: &str| {
entries
.iter()
.filter(|entry| entry["label"].as_str().is_some_and(|entry_label| entry_label == label))
.filter_map(|entry| entry[direction]["errors"].as_u64())
.sum::<u64>()
};
let payload_len = u64::try_from(payload.len()).expect("test payload length should fit u64");
assert_eq!(
io_bytes(RAW_SHARD_READ_LOCAL_LABEL, "read"),
payload_len + u64::try_from(fallback_payload.len()).expect("fallback payload length should fit u64")
);
assert_eq!(io_bytes(RAW_SHARD_WRITE_LOCAL_LABEL, "write"), payload_len);
assert_eq!(
io_bytes(RAW_SHARD_READ_REMOTE_LABEL, "read"),
u64::try_from(remote_payload.len()).expect("remote payload length should fit u64")
);
assert_eq!(
io_bytes(RAW_SHARD_WRITE_REMOTE_LABEL, "write"),
u64::try_from(remote_payload.len()).expect("remote payload length should fit u64")
);
assert_eq!(
io_errors(RAW_SHARD_WRITE_REMOTE_LABEL, "write"),
1,
"the wrapper must record the real remote writer failure without changing it, while excluding WouldBlock"
);
for label in [
RAW_SHARD_READ_LOCAL_LABEL,
RAW_SHARD_READ_REMOTE_LABEL,
RAW_SHARD_WRITE_LOCAL_LABEL,
RAW_SHARD_WRITE_REMOTE_LABEL,
] {
assert!(
!label.contains(['/', ':', '?', '@']),
"raw shard I/O labels must not carry a path, host, query, or credential delimiter: {label}"
);
}
}
#[test] #[test]
fn object_mmap_read_max_length_defaults_and_env_override() { fn object_mmap_read_max_length_defaults_and_env_override() {
temp_env::with_var(ENV_OBJECT_MMAP_READ_MAX_LENGTH, None::<&str>, || { temp_env::with_var(ENV_OBJECT_MMAP_READ_MAX_LENGTH, None::<&str>, || {
@@ -741,25 +1194,9 @@ mod tests {
// be materialized in memory by the mmap-copy path; over-cap reads stream. // be materialized in memory by the mmap-copy path; over-cap reads stream.
#[tokio::test] #[tokio::test]
async fn open_disk_reader_streams_when_length_exceeds_mmap_cap() { async fn open_disk_reader_streams_when_length_exceeds_mmap_cap() {
use crate::disk::endpoint::Endpoint;
use crate::disk::{DiskOption, new_disk};
use tokio::io::AsyncReadExt; use tokio::io::AsyncReadExt;
let dir = tempfile::tempdir().expect("tempdir should be created"); let (disk, _dir) = local_test_disk().await;
let mut endpoint =
Endpoint::try_from(dir.path().to_str().expect("tempdir path should be utf8")).expect("endpoint should parse");
endpoint.set_pool_index(0);
endpoint.set_set_index(0);
endpoint.set_disk_index(0);
let disk = new_disk(
&endpoint,
&DiskOption {
cleanup: false,
health_check: false,
},
)
.await
.expect("local disk should be created");
let payload = vec![7u8; 4096]; let payload = vec![7u8; 4096];
disk.make_volume("test-bucket").await.expect("volume should be created"); disk.make_volume("test-bucket").await.expect("volume should be created");
@@ -814,6 +1251,17 @@ mod tests {
.await; .await;
} }
#[tokio::test]
async fn disk_bitrot_reader_and_writer_preserve_full_shard_body() {
let (disk, _dir) = local_test_disk().await;
let bucket = "test-bucket";
let path = "obj/wrapped-part.1";
let payload = b"wrapped shard body";
let actual = round_trip_disk_bitrot(&disk, bucket, path, payload, 4).await;
assert_eq!(actual, payload, "raw shard I/O instrumentation must not alter stored bytes");
}
#[tokio::test] #[tokio::test]
async fn test_create_bitrot_reader_with_inline_data() { async fn test_create_bitrot_reader_with_inline_data() {
let test_data = b"hello world test data"; let test_data = b"hello world test data";
+58 -1
View File
@@ -22,6 +22,8 @@ use rustfs_config::{
ENV_STARTUP_TOPOLOGY_WAIT_MODE, ENV_STARTUP_TOPOLOGY_WAIT_TIMEOUT, ENV_UNSAFE_BYPASS_DISK_CHECK, ENV_STARTUP_TOPOLOGY_WAIT_MODE, ENV_STARTUP_TOPOLOGY_WAIT_TIMEOUT, ENV_UNSAFE_BYPASS_DISK_CHECK,
}; };
use rustfs_utils::{XHost, check_local_server_addr, get_env_opt_str, get_host_ip, is_local_host}; use rustfs_utils::{XHost, check_local_server_addr, get_env_opt_str, get_host_ip, is_local_host};
#[cfg(test)]
use std::sync::{LazyLock, Mutex};
use std::{ use std::{
collections::{BTreeMap, BTreeSet, HashMap, HashSet, hash_map::Entry}, collections::{BTreeMap, BTreeSet, HashMap, HashSet, hash_map::Entry},
future::Future, future::Future,
@@ -598,6 +600,58 @@ const DNS_RETRY_JITTER_PERCENT: u64 = 20;
/// wait does not flood the log with one line per backoff tick. /// wait does not flood the log with one line per backoff tick.
const TOPOLOGY_WARN_THROTTLE: Duration = Duration::from_secs(30); const TOPOLOGY_WARN_THROTTLE: Duration = Duration::from_secs(30);
#[cfg(test)]
static FORCED_LOCAL_HOST_RESOLUTION_TIMEOUTS: LazyLock<Mutex<HashSet<String>>> = LazyLock::new(|| Mutex::new(HashSet::new()));
#[cfg(test)]
struct LocalHostResolutionTimeoutGuard {
hosts: Vec<String>,
}
#[cfg(test)]
impl Drop for LocalHostResolutionTimeoutGuard {
fn drop(&mut self) {
let mut forced_hosts = FORCED_LOCAL_HOST_RESOLUTION_TIMEOUTS
.lock()
.expect("local-host test resolver mutex poisoned");
for host in &self.hosts {
forced_hosts.remove(host);
}
}
}
#[cfg(test)]
fn force_local_host_resolution_timeout_for_test(hosts: &[&str]) -> LocalHostResolutionTimeoutGuard {
let hosts = hosts.iter().map(|host| (*host).to_string()).collect::<Vec<_>>();
let mut forced_hosts = FORCED_LOCAL_HOST_RESOLUTION_TIMEOUTS
.lock()
.expect("local-host test resolver mutex poisoned");
forced_hosts.extend(hosts.iter().cloned());
LocalHostResolutionTimeoutGuard { hosts }
}
#[cfg(test)]
fn local_host_resolution_timeout_forced(host: &Host<&str>) -> bool {
let host = match host {
Host::Domain(domain) => (*domain).to_string(),
Host::Ipv4(ip) => ip.to_string(),
Host::Ipv6(ip) => ip.to_string(),
};
FORCED_LOCAL_HOST_RESOLUTION_TIMEOUTS
.lock()
.expect("local-host test resolver mutex poisoned")
.contains(&host)
}
fn endpoint_is_local_host(host: Host<&str>, port: u16, local_port: u16) -> Result<bool> {
#[cfg(test)]
if local_host_resolution_timeout_forced(&host) {
return Err(Error::new(ErrorKind::TimedOut, "resolver timeout"));
}
is_local_host(host, port, local_port)
}
struct DnsRetryDeadline { struct DnsRetryDeadline {
started: Instant, started: Instant,
timeout: Duration, timeout: Duration,
@@ -694,7 +748,7 @@ async fn resolve_local_host_with_retry(
retry_dns_operation( retry_dns_operation(
|| { || {
let host = host.clone(); let host = host.clone();
async move { is_local_host(host, port, local_port) } async move { endpoint_is_local_host(host, port, local_port) }
}, },
async_sleep, async_sleep,
dns_retry_deadline, dns_retry_deadline,
@@ -2231,6 +2285,9 @@ mod test {
#[serial] #[serial]
#[tokio::test] #[tokio::test]
async fn create_server_endpoints_bounds_kubernetes_alias_dns_fallback() { async fn create_server_endpoints_bounds_kubernetes_alias_dns_fallback() {
let _resolution_timeout =
force_local_host_resolution_timeout_for_test(&["unrelated-0.example.invalid", "unrelated-1.example.invalid"]);
async_with_vars( async_with_vars(
[ [
(ENV_LOCAL_ENDPOINT_HOST, None), (ENV_LOCAL_ENDPOINT_HOST, None),
@@ -337,6 +337,15 @@ pub struct NotificationPeerErr {
pub err: Option<Error>, pub err: Option<Error>,
} }
/// One peer's answer to a KMS configuration fingerprint probe.
pub struct PeerKmsConfigFingerprint {
pub host: String,
/// `None` when the peer has no KMS configuration of its own, or could not
/// be asked at all, in which case `err` carries the reason.
pub fingerprint: Option<String>,
pub err: Option<Error>,
}
fn notification_peer_result<T>(host: String, result: Result<T>) -> NotificationPeerErr { fn notification_peer_result<T>(host: String, result: Result<T>) -> NotificationPeerErr {
NotificationPeerErr { host, err: result.err() } NotificationPeerErr { host, err: result.err() }
} }
@@ -518,6 +527,47 @@ impl NotificationSys {
self.signal_dynamic_config(sub_sys, false).await self.signal_dynamic_config(sub_sys, false).await
} }
/// Ask every peer to re-read the cluster-persisted KMS configuration.
///
/// Best-effort by contract: the caller has already switched locally, so a
/// peer that fails is reported rather than rolled back. Peers built before
/// the KMS subsystem existed reject the signal with an explicit error.
pub async fn reload_kms_config(&self) -> Vec<NotificationPeerErr> {
self.reload_dynamic_config(crate::cluster::rpc::KMS_SIGNAL_SUBSYSTEM).await
}
/// Collect the KMS configuration fingerprint each peer is running.
///
/// A peer whose build predates the KMS subsystem rejects the probe, so it
/// is reported as an error rather than silently agreeing with this node.
pub async fn kms_config_fingerprints(&self) -> Vec<PeerKmsConfigFingerprint> {
let mut futures = Vec::with_capacity(self.peer_clients.len());
for client in self.peer_clients.iter() {
futures.push(async move {
let Some(client) = client else {
return PeerKmsConfigFingerprint {
host: String::new(),
fingerprint: None,
err: Some(Error::other("peer is not reachable")),
};
};
match client.kms_config_fingerprint().await {
Ok(fingerprint) => PeerKmsConfigFingerprint {
host: client.host.to_string(),
fingerprint,
err: None,
},
Err(e) => PeerKmsConfigFingerprint {
host: client.host.to_string(),
fingerprint: None,
err: Some(e),
},
}
});
}
join_all(futures).await
}
pub async fn refresh_config_snapshot(&self) -> Vec<NotificationPeerErr> { pub async fn refresh_config_snapshot(&self) -> Vec<NotificationPeerErr> {
let mut futures = Vec::with_capacity(self.peer_clients.len()); let mut futures = Vec::with_capacity(self.peer_clients.len());
for client in self.peer_clients.iter() { for client in self.peer_clients.iter() {
+224 -1
View File
@@ -4714,6 +4714,7 @@ mod tests {
use crate::layout::endpoints::SetupType; use crate::layout::endpoints::SetupType;
use crate::object_api::BLOCK_SIZE_V2; use crate::object_api::BLOCK_SIZE_V2;
use crate::object_api::ObjectInfo; use crate::object_api::ObjectInfo;
use crate::set_disk::core::io_primitives::rename_fanout_barrier;
use crate::storage_api_contracts::{ use crate::storage_api_contracts::{
heal::HealOperations as _, lifecycle::TransitionedObject, list::ListOperations as _, multipart::CompletePart, heal::HealOperations as _, lifecycle::TransitionedObject, list::ListOperations as _, multipart::CompletePart,
namespace::NamespaceLocking as _, object::ObjectIO as _, object::ObjectOperations as _, namespace::NamespaceLocking as _, object::ObjectIO as _, object::ObjectOperations as _,
@@ -9565,6 +9566,15 @@ mod tests {
make_local_bucket_test_set_disks_with_drive_count(2).await make_local_bucket_test_set_disks_with_drive_count(2).await
} }
fn assert_exclusive_object_lock_held(set_disks: &SetDisks, bucket: &str, object: &str) {
let lock = set_disks
.local_lock_manager_for_test()
.get_lock_info(&ObjectKey::new(bucket, object))
.expect("object lock should be visible while rename is paused");
assert!(matches!(lock.mode, rustfs_lock::LockMode::Exclusive));
assert_eq!(lock.owner.as_ref(), set_disks.locker_owner.as_str());
}
async fn make_local_bucket_test_set_disks_with_drive_count(drive_count: usize) -> Arc<SetDisks> { async fn make_local_bucket_test_set_disks_with_drive_count(drive_count: usize) -> Arc<SetDisks> {
let format = FormatV3::new(1, drive_count); let format = FormatV3::new(1, drive_count);
let mut endpoints = Vec::new(); let mut endpoints = Vec::new();
@@ -9599,7 +9609,9 @@ mod tests {
disks.push(Some(disk)); disks.push(Some(disk));
} }
let set_disks = SetDisks::new( let instance_ctx = Arc::new(InstanceContext::new());
instance_ctx.update_erasure_type(SetupType::Erasure).await;
let set_disks = SetDisks::new_with_instance_ctx(
"test-owner".to_string(), "test-owner".to_string(),
Arc::new(RwLock::new(disks)), Arc::new(RwLock::new(disks)),
drive_count, drive_count,
@@ -9609,6 +9621,7 @@ mod tests {
endpoints, endpoints,
format, format,
Vec::new(), Vec::new(),
instance_ctx,
) )
.await; .await;
set_disks.set_test_storage_class_config( set_disks.set_test_storage_class_config(
@@ -10262,6 +10275,216 @@ mod tests {
)); ));
} }
#[tokio::test]
async fn conditional_replace_holds_object_lock_through_rename() {
let set_disks = make_local_bucket_test_set_disks().await;
let bucket = "bucket-conditional-replace-fence";
let object = "config/conditional-replace.json";
set_disks
.make_bucket(bucket, &MakeBucketOptions::default())
.await
.expect("bucket should be created");
let mut initial_reader = PutObjReader::from_vec(b"initial config".to_vec());
let initial = set_disks
.put_object(
bucket,
object,
&mut initial_reader,
&ObjectOptions {
no_lock: true,
..Default::default()
},
)
.await
.expect("initial config should be written");
let initial_etag = initial.etag.expect("initial config should have an ETag");
let barrier = rename_fanout_barrier::arm(object, 0, rename_fanout_barrier::PHASE_RENAME);
let writer_store = set_disks.clone();
let expected_etag = initial_etag.clone();
let writer = tokio::spawn(async move {
let mut reader = PutObjReader::from_vec(b"replacement config".to_vec());
writer_store
.put_object(
bucket,
object,
&mut reader,
&ObjectOptions {
preserve_etag: Some("replacement-etag".to_string()),
http_preconditions: Some(HTTPPreconditions {
if_match: Some(expected_etag),
..Default::default()
}),
..Default::default()
},
)
.await
});
tokio::time::timeout(std::time::Duration::from_secs(30), barrier.wait_until_paused())
.await
.expect("conditional replace should reach the rename barrier");
assert_exclusive_object_lock_held(&set_disks, bucket, object);
barrier.release();
writer
.await
.expect("conditional writer task should finish")
.expect("matching conditional replace should commit");
assert!(
set_disks
.local_lock_manager_for_test()
.get_lock_info(&ObjectKey::new(bucket, object))
.is_none(),
"conditional replace should release the object lock after commit"
);
let contender = set_disks
.new_ns_lock(bucket, object)
.await
.expect("contender namespace lock should be created");
let contender_guard = contender
.get_write_lock(std::time::Duration::from_secs(30))
.await
.expect("contender should acquire after conditional replace commits");
drop(contender_guard);
let mut stale_reader = PutObjReader::from_vec(b"stale config".to_vec());
let err = set_disks
.put_object(
bucket,
object,
&mut stale_reader,
&ObjectOptions {
http_preconditions: Some(HTTPPreconditions {
if_match: Some(initial_etag),
..Default::default()
}),
..Default::default()
},
)
.await
.expect_err("the old ETag must fail after the fenced replacement commits");
assert_eq!(err, StorageError::PreconditionFailed);
}
#[tokio::test]
async fn repeated_body_write_keeps_etag_but_changes_data_dir_generation() {
let set_disks = make_local_bucket_test_set_disks().await;
let bucket = "bucket-write-generation";
let object = "config/write-generation.json";
let body = b"identical config body".to_vec();
set_disks
.make_bucket(bucket, &MakeBucketOptions::default())
.await
.expect("bucket should be created");
let mut first_reader = PutObjReader::from_vec(body.clone());
let first = set_disks
.put_object(
bucket,
object,
&mut first_reader,
&ObjectOptions {
no_lock: true,
..Default::default()
},
)
.await
.expect("first config body should be written");
let mut second_reader = PutObjReader::from_vec(body);
let second = set_disks
.put_object(
bucket,
object,
&mut second_reader,
&ObjectOptions {
no_lock: true,
..Default::default()
},
)
.await
.expect("identical config body should be rewritten");
assert_eq!(first.etag, second.etag, "content ETag should expose the ABA collision");
assert_ne!(first.data_dir, second.data_dir, "each committed body write needs a unique generation");
assert!(first.data_dir.is_some() && second.data_dir.is_some());
}
#[tokio::test]
async fn conditional_create_holds_object_lock_through_rename() {
let set_disks = make_local_bucket_test_set_disks().await;
let bucket = "bucket-conditional-create-fence";
let object = "config/conditional-create.json";
set_disks
.make_bucket(bucket, &MakeBucketOptions::default())
.await
.expect("bucket should be created");
let barrier = rename_fanout_barrier::arm(object, 0, rename_fanout_barrier::PHASE_RENAME);
let writer_store = set_disks.clone();
let writer = tokio::spawn(async move {
let mut reader = PutObjReader::from_vec(b"created config".to_vec());
writer_store
.put_object(
bucket,
object,
&mut reader,
&ObjectOptions {
http_preconditions: Some(HTTPPreconditions {
if_none_match: Some("*".to_string()),
..Default::default()
}),
..Default::default()
},
)
.await
});
tokio::time::timeout(std::time::Duration::from_secs(30), barrier.wait_until_paused())
.await
.expect("conditional create should reach the rename barrier");
assert_exclusive_object_lock_held(&set_disks, bucket, object);
barrier.release();
writer
.await
.expect("conditional writer task should finish")
.expect("first conditional create should commit");
assert!(
set_disks
.local_lock_manager_for_test()
.get_lock_info(&ObjectKey::new(bucket, object))
.is_none(),
"conditional create should release the object lock after commit"
);
let contender = set_disks
.new_ns_lock(bucket, object)
.await
.expect("contender namespace lock should be created");
let contender_guard = contender
.get_write_lock(std::time::Duration::from_secs(30))
.await
.expect("contender should acquire after conditional create commits");
drop(contender_guard);
let mut duplicate_reader = PutObjReader::from_vec(b"duplicate config".to_vec());
let err = set_disks
.put_object(
bucket,
object,
&mut duplicate_reader,
&ObjectOptions {
http_preconditions: Some(HTTPPreconditions {
if_none_match: Some("*".to_string()),
..Default::default()
}),
..Default::default()
},
)
.await
.expect_err("a second create-only write must not replace the committed config");
assert_eq!(err, StorageError::PreconditionFailed);
}
#[tokio::test] #[tokio::test]
async fn set_level_if_none_match_fails_closed_without_read_quorum() { async fn set_level_if_none_match_fails_closed_without_read_quorum() {
let set_disks = make_local_bucket_test_set_disks_with_drive_count(4).await; let set_disks = make_local_bucket_test_set_disks_with_drive_count(4).await;
+6 -6
View File
@@ -199,7 +199,7 @@ impl SetDisks {
); );
} }
#[hotpath::measure] #[hotpath::measure(impl_type = "SetDisks")]
pub async fn read_version_optimized( pub async fn read_version_optimized(
&self, &self,
bucket: &str, bucket: &str,
@@ -238,7 +238,7 @@ impl SetDisks {
} }
#[tracing::instrument(level = "debug", skip(self))] #[tracing::instrument(level = "debug", skip(self))]
#[hotpath::measure] #[hotpath::measure(impl_type = "SetDisks")]
pub(super) async fn get_object_fileinfo( pub(super) async fn get_object_fileinfo(
&self, &self,
bucket: &str, bucket: &str,
@@ -410,7 +410,7 @@ impl SetDisks {
Ok((fi, parts_metadata, op_online_disks)) Ok((fi, parts_metadata, op_online_disks))
} }
#[hotpath::measure] #[hotpath::measure(impl_type = "SetDisks")]
pub(super) async fn get_object_info_and_quorum( pub(super) async fn get_object_info_and_quorum(
&self, &self,
bucket: &str, bucket: &str,
@@ -605,7 +605,7 @@ impl SetDisks {
} }
#[allow(clippy::too_many_arguments)] #[allow(clippy::too_many_arguments)]
#[hotpath::measure] #[hotpath::measure(impl_type = "SetDisks")]
pub(super) async fn get_object_with_fileinfo<W>( pub(super) async fn get_object_with_fileinfo<W>(
// &self, // &self,
bucket: &str, bucket: &str,
@@ -1140,7 +1140,7 @@ impl SetDisks {
} }
#[allow(clippy::too_many_arguments)] #[allow(clippy::too_many_arguments)]
#[hotpath::measure] #[hotpath::measure(impl_type = "SetDisks")]
pub(super) async fn get_object_decode_reader_with_fileinfo( pub(super) async fn get_object_decode_reader_with_fileinfo(
bucket: &str, bucket: &str,
object: &str, object: &str,
@@ -1296,7 +1296,7 @@ impl SetDisks {
} }
#[allow(clippy::too_many_arguments)] #[allow(clippy::too_many_arguments)]
#[hotpath::measure] #[hotpath::measure(impl_type = "SetDisks")]
async fn build_codec_streaming_part_reader( async fn build_codec_streaming_part_reader(
bucket: &str, bucket: &str,
object: &str, object: &str,
+23 -4
View File
@@ -4861,10 +4861,7 @@ async fn poll_merge_head(rx: &CancellationToken, in_channels: &mut [Receiver<Met
async fn send_or_cancel(rx: &CancellationToken, out_channel: &Sender<MetaCacheEntry>, entry: MetaCacheEntry) -> Result<bool> { async fn send_or_cancel(rx: &CancellationToken, out_channel: &Sender<MetaCacheEntry>, entry: MetaCacheEntry) -> Result<bool> {
tokio::select! { tokio::select! {
result = out_channel.send(entry) => { result = out_channel.send(entry) => Ok(result.is_ok()),
result.map_err(Error::other)?;
Ok(true)
}
_ = rx.cancelled() => Ok(false), _ = rx.cancelled() => Ok(false),
} }
} }
@@ -9858,6 +9855,28 @@ mod test {
assert_eq!(results, vec!["obj-a", "obj-b"]); assert_eq!(results, vec!["obj-a", "obj-b"]);
} }
#[tokio::test]
async fn merge_entry_channels_treats_dropped_output_receiver_as_completion() {
let (tx_a, rx_a) = mpsc::channel(4);
let (tx_b, rx_b) = mpsc::channel(4);
let (out_tx, out_rx) = mpsc::channel(1);
tx_a.send(test_meta_entry("obj-a")).await.unwrap();
tx_b.send(test_meta_entry("obj-b")).await.unwrap();
drop(tx_a);
drop(tx_b);
drop(out_rx);
let result = timeout(
Duration::from_secs(1),
merge_entry_channels(CancellationToken::new(), vec![rx_a, rx_b], out_tx, 1),
)
.await
.expect("merge should stop promptly when its consumer disconnects");
assert!(result.is_ok(), "consumer disconnect must not surface as a merge worker error");
}
#[test] #[test]
fn walk_ascending_versions_contract_reverses_newest_first_metadata() { fn walk_ascending_versions_contract_reverses_newest_first_metadata() {
// Documents the invariant the walk `versions_sort` handling relies on: // Documents the invariant the walk `versions_sort` handling relies on:
+1 -1
View File
@@ -238,7 +238,7 @@ impl ECStore {
} }
#[instrument(skip(self, data))] #[instrument(skip(self, data))]
#[hotpath::measure] #[hotpath::measure(impl_type = "ECStore")]
pub(super) async fn handle_put_object_part( pub(super) async fn handle_put_object_part(
&self, &self,
bucket: &str, bucket: &str,
+2 -2
View File
@@ -815,7 +815,7 @@ impl ECStore {
} }
#[instrument(level = "debug", skip(self))] #[instrument(level = "debug", skip(self))]
#[hotpath::measure] #[hotpath::measure(impl_type = "ECStore")]
pub(super) async fn handle_get_object_reader( pub(super) async fn handle_get_object_reader(
&self, &self,
bucket: &str, bucket: &str,
@@ -849,7 +849,7 @@ impl ECStore {
} }
#[instrument(level = "debug", skip(self, data))] #[instrument(level = "debug", skip(self, data))]
#[hotpath::measure] #[hotpath::measure(impl_type = "ECStore")]
pub(super) async fn handle_put_object( pub(super) async fn handle_put_object(
&self, &self,
bucket: &str, bucket: &str,
@@ -0,0 +1,95 @@
// Copyright 2024 RustFS Team
//
// Licensed under the Apache License, Version 2.0 (the "License");
// you may not use this file except in compliance with the License.
// You may obtain a copy of the License at
//
// http://www.apache.org/licenses/LICENSE-2.0
//
// Unless required by applicable law or agreed to in writing, software
// distributed under the License is distributed on an "AS IS" BASIS,
// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
// See the License for the specific language governing permissions and
// limitations under the License.
const STORE_MULTIPART: &str = include_str!("../src/store/multipart.rs");
const STORE_OBJECT: &str = include_str!("../src/store/object.rs");
const ERASURE: &str = include_str!("../src/erasure/coding/erasure.rs");
const DECODE: &str = include_str!("../src/erasure/coding/decode.rs");
const ENCODE: &str = include_str!("../src/erasure/coding/encode.rs");
const LOCAL_DISK: &str = include_str!("../src/disk/local.rs");
const BITROT: &str = include_str!("../src/erasure/coding/bitrot.rs");
const SET_DISK_READ: &str = include_str!("../src/set_disk/read.rs");
fn assert_measured_as(source: &str, impl_type: &str, function: &str) {
let attribute = format!("#[hotpath::measure(impl_type = \"{impl_type}\")]\n");
let function_start = format!("fn {function}");
let mut remaining = source;
while let Some(attribute_offset) = remaining.find(&attribute) {
let measured_source = &remaining[attribute_offset + attribute.len()..];
if measured_source.find(&function_start).is_some_and(|offset| offset < 256) {
return;
}
remaining = measured_source;
}
panic!("missing HotPath impl_type={impl_type} for {function}");
}
#[test]
fn inherent_hotpath_measurements_have_cpu_attribution_types() {
assert_measured_as(STORE_MULTIPART, "ECStore", "handle_put_object_part");
for function in ["handle_get_object_reader", "handle_put_object"] {
assert_measured_as(STORE_OBJECT, "ECStore", function);
}
for function in [
"encode_data",
"encode_data_owned",
"encode_data_bytes_mut",
"decode_data",
"decode_data_and_parity",
] {
assert_measured_as(ERASURE, "Erasure", function);
}
assert_measured_as(DECODE, "ParallelReader", "read");
assert_measured_as(DECODE, "Erasure", "decode");
for function in [
"encode",
"encode_batched",
"encode_inline_small",
"encode_single_block_non_inline",
] {
assert_measured_as(ENCODE, "Erasure", function);
}
for function in ["read_metadata_with_dmtime", "read_all_data"] {
assert_measured_as(LOCAL_DISK, "LocalDisk", function);
}
assert_measured_as(BITROT, "BitrotReader", "read");
assert!(
BITROT.contains("#[hotpath::measure(label = \"BitrotWriter::write\", impl_type = \"BitrotWriter\")]"),
"BitrotWriter::write must retain its stable label and CPU impl_type",
);
for function in [
"read_version_optimized",
"get_object_fileinfo",
"get_object_info_and_quorum",
"get_object_with_fileinfo",
"get_object_decode_reader_with_fileinfo",
"build_codec_streaming_part_reader",
] {
assert_measured_as(SET_DISK_READ, "SetDisks", function);
}
}
#[test]
fn module_hotpath_measurements_remain_untyped() {
assert!(
BITROT.contains("#[hotpath::measure]\npub async fn bitrot_verify"),
"bitrot_verify is a module function and must not claim an impl type",
);
assert!(
SET_DISK_READ.contains("#[hotpath::measure]\nasync fn setup_multipart_part_readers"),
"setup_multipart_part_readers is a module function and must not claim an impl type",
);
}
+235 -26
View File
@@ -41,7 +41,7 @@ use rustfs_policy::policy::Args;
use rustfs_policy::policy::opa; use rustfs_policy::policy::opa;
use rustfs_policy::policy::{Policy, PolicyDoc, iam_policy_claim_name_sa, policy_needs_existing_object_tag_for_args}; use rustfs_policy::policy::{Policy, PolicyDoc, iam_policy_claim_name_sa, policy_needs_existing_object_tag_for_args};
use serde_json::Value; use serde_json::Value;
use std::collections::HashMap; use std::collections::{HashMap, HashSet};
use std::sync::Arc; use std::sync::Arc;
use std::sync::OnceLock; use std::sync::OnceLock;
use time::OffsetDateTime; use time::OffsetDateTime;
@@ -61,10 +61,51 @@ const VERIFIED_FEDERATED_POLICY_CLAIM: &str = "x-rustfs-internal-federated-polic
const FEDERATED_POLICY_BOUNDARY_CLAIM: &str = "x-rustfs-internal-federated-boundary"; const FEDERATED_POLICY_BOUNDARY_CLAIM: &str = "x-rustfs-internal-federated-boundary";
pub(crate) const SITE_REPLICATOR_CLAIM: &str = "site-replicator"; pub(crate) const SITE_REPLICATOR_CLAIM: &str = "site-replicator";
static POLICY_PLUGIN_CLIENT: OnceLock<Arc<RwLock<Option<rustfs_policy::policy::opa::AuthZPlugin>>>> = OnceLock::new(); #[derive(Clone)]
enum PolicyPluginState {
Disabled,
Initializing,
Ready(opa::AuthZPlugin),
Failed,
}
fn get_policy_plugin_client() -> Arc<RwLock<Option<rustfs_policy::policy::opa::AuthZPlugin>>> { static POLICY_PLUGIN_STATE: OnceLock<Arc<RwLock<PolicyPluginState>>> = OnceLock::new();
POLICY_PLUGIN_CLIENT.get_or_init(|| Arc::new(RwLock::new(None))).clone()
fn get_policy_plugin_state() -> Arc<RwLock<PolicyPluginState>> {
POLICY_PLUGIN_STATE
.get_or_init(|| {
let configured = opa::is_configured();
let state = Arc::new(RwLock::new(if configured {
PolicyPluginState::Initializing
} else {
PolicyPluginState::Disabled
}));
if configured {
let state = Arc::clone(&state);
tokio::spawn(async move {
let next_state = match opa::lookup_config().await {
Ok(conf) if conf.enable() => {
info!("OPA plugin enabled");
PolicyPluginState::Ready(opa::AuthZPlugin::new(conf))
}
Ok(_) => PolicyPluginState::Failed,
Err(e) => {
error!(
component = "iam",
subsystem = "policy_plugin",
result = "configuration_load_failed",
error_kind = e.kind(),
"OPA plugin configuration load failed"
);
PolicyPluginState::Failed
}
};
*state.write().await = next_state;
});
}
state
})
.clone()
} }
pub struct IamSys<T> { pub struct IamSys<T> {
@@ -173,19 +214,7 @@ impl<T: Store> IamSys<T> {
/// # Returns /// # Returns
/// A new instance of IamSys /// A new instance of IamSys
pub fn new(store: Arc<IamCache<T>>) -> Self { pub fn new(store: Arc<IamCache<T>>) -> Self {
tokio::spawn(async move { get_policy_plugin_state();
match opa::lookup_config().await {
Ok(conf) => {
if conf.enable() {
Self::set_policy_plugin_client(opa::AuthZPlugin::new(conf)).await;
info!("OPA plugin enabled");
}
}
Err(e) => {
error!("Error loading OPA configuration err:{}", e);
}
};
});
Self { Self {
store, store,
@@ -206,15 +235,22 @@ impl<T: Store> IamSys<T> {
} }
pub async fn set_policy_plugin_client(client: rustfs_policy::policy::opa::AuthZPlugin) { pub async fn set_policy_plugin_client(client: rustfs_policy::policy::opa::AuthZPlugin) {
let policy_plugin_client = get_policy_plugin_client(); let policy_plugin_state = get_policy_plugin_state();
let mut guard = policy_plugin_client.write().await; let mut guard = policy_plugin_state.write().await;
*guard = Some(client); *guard = PolicyPluginState::Ready(client);
} }
pub async fn get_policy_plugin_client() -> Option<rustfs_policy::policy::opa::AuthZPlugin> { pub async fn get_policy_plugin_client() -> Option<rustfs_policy::policy::opa::AuthZPlugin> {
let policy_plugin_client = get_policy_plugin_client(); let policy_plugin_state = get_policy_plugin_state();
let guard = policy_plugin_client.read().await; let guard = policy_plugin_state.read().await;
guard.clone() match &*guard {
PolicyPluginState::Ready(client) => Some(client.clone()),
PolicyPluginState::Disabled | PolicyPluginState::Initializing | PolicyPluginState::Failed => None,
}
}
async fn policy_plugin_state() -> PolicyPluginState {
get_policy_plugin_state().read().await.clone()
} }
pub async fn load_group(&self, name: &str) -> Result<()> { pub async fn load_group(&self, name: &str) -> Result<()> {
@@ -1060,12 +1096,21 @@ impl<T: Store> IamSys<T> {
}; };
} }
if Self::get_policy_plugin_client().await.is_some() { match Self::policy_plugin_state().await {
PolicyPluginState::Ready(_) => {
return PreparedIamAuth { return PreparedIamAuth {
needs_existing_object_tag: false, needs_existing_object_tag: false,
mode: PreparedIamMode::Opa, mode: PreparedIamMode::Opa,
}; };
} }
PolicyPluginState::Initializing | PolicyPluginState::Failed => {
return PreparedIamAuth {
needs_existing_object_tag: false,
mode: PreparedIamMode::Deny,
};
}
PolicyPluginState::Disabled => {}
}
let Ok((is_svc, parent_user)) = self.is_service_account(args.account).await else { let Ok((is_svc, parent_user)) = self.is_service_account(args.account).await else {
return PreparedIamAuth { return PreparedIamAuth {
@@ -1163,7 +1208,22 @@ impl<T: Store> IamSys<T> {
}; };
} }
let combined_policy = self.get_combined_policy(&policies).await; let (resolved_policies, combined_policy) = self.store.merge_policies(&policies.join(",")).await;
if args.deny_only {
let resolved_policy_names: HashSet<&str> = resolved_policies
.split(',')
.filter(|policy_name| !policy_name.trim().is_empty())
.collect();
if policies
.iter()
.any(|policy_name| !resolved_policy_names.contains(policy_name.as_str()))
{
return PreparedIamAuth {
needs_existing_object_tag: false,
mode: PreparedIamMode::Deny,
};
}
}
PreparedIamAuth { PreparedIamAuth {
needs_existing_object_tag: policy_needs_existing_object_tag_for_args(&combined_policy, args).await, needs_existing_object_tag: policy_needs_existing_object_tag_for_args(&combined_policy, args).await,
mode: PreparedIamMode::Regular { combined_policy }, mode: PreparedIamMode::Regular { combined_policy },
@@ -1703,7 +1763,7 @@ mod tests {
fn deprecated_list_polices_api_is_available() { fn deprecated_list_polices_api_is_available() {
let _ = IamSys::<StsTestMockStore>::list_polices; let _ = IamSys::<StsTestMockStore>::list_polices;
} }
use rustfs_policy::policy::action::{Action, AdminAction, S3Action}; use rustfs_policy::policy::action::{Action, AdminAction, S3Action, StsAction};
use rustfs_policy::policy::policy_uses_existing_object_tag_conditions; use rustfs_policy::policy::policy_uses_existing_object_tag_conditions;
use serde_json::Value; use serde_json::Value;
use std::{ use std::{
@@ -3309,6 +3369,155 @@ mod tests {
assert!(!iam_sys.eval_prepared(&prepared, &args).await); assert!(!iam_sys.eval_prepared(&prepared, &args).await);
} }
#[tokio::test]
async fn regular_deny_only_rejects_unresolved_mapped_policy() {
let iam_sys = test_iam_sys().await;
let user = "regular-unresolved-policy";
let identity = UserIdentity::from(Credentials {
access_key: user.to_string(),
secret_key: "longenoughsecret".to_string(),
status: ACCOUNT_ON.to_string(),
..Default::default()
});
iam_sys.store.cache.with_write_lock(|cache| {
let now = OffsetDateTime::now_utc();
cache.add_or_update_user(user, &identity, now);
cache.add_or_update_user_policy(user, &MappedPolicy::new("readwrite,missing-policy"), now);
});
let groups = None;
let claims = HashMap::new();
let conditions = HashMap::new();
let args = Args {
account: user,
groups: &groups,
action: Action::StsAction(StsAction::AssumeRoleAction),
bucket: "",
conditions: &conditions,
is_owner: false,
object: "",
claims: &claims,
deny_only: true,
};
let prepared = iam_sys.prepare_regular_auth(&args).await;
assert!(matches!(prepared.mode, PreparedIamMode::Deny));
assert!(
!iam_sys.eval_prepared(&prepared, &args).await,
"deny-only evaluation must fail closed when a mapped policy document is unresolved"
);
}
#[tokio::test]
async fn regular_full_evaluation_keeps_resolved_part_of_partial_mapping() {
let iam_sys = test_iam_sys().await;
let user = "regular-partial-policy";
let identity = UserIdentity::from(Credentials {
access_key: user.to_string(),
secret_key: "longenoughsecret".to_string(),
status: ACCOUNT_ON.to_string(),
..Default::default()
});
iam_sys.store.cache.with_write_lock(|cache| {
let now = OffsetDateTime::now_utc();
cache.add_or_update_user(user, &identity, now);
cache.add_or_update_user_policy(user, &MappedPolicy::new("readwrite,missing-policy"), now);
});
let groups = None;
let claims = HashMap::new();
let conditions = HashMap::new();
let args = Args {
account: user,
groups: &groups,
action: Action::S3Action(S3Action::GetObjectAction),
bucket: "bucket",
conditions: &conditions,
is_owner: false,
object: "object",
claims: &claims,
deny_only: false,
};
let prepared = iam_sys.prepare_regular_auth(&args).await;
assert!(matches!(prepared.mode, PreparedIamMode::Regular { .. }));
assert!(
iam_sys.eval_prepared(&prepared, &args).await,
"full evaluation must preserve the historical partial-resolution behavior"
);
}
#[tokio::test]
async fn regular_deny_only_honors_group_derived_explicit_deny() {
let iam_sys = test_iam_sys().await;
let deny_policy =
Policy::parse_config(br#"{"Version":"2012-10-17","Statement":[{"Effect":"Deny","Action":["sts:AssumeRole"]}]}"#)
.expect("group deny policy should parse");
let now = OffsetDateTime::now_utc();
iam_sys
.store
.cache
.add_or_update_policy_doc("deny-assume-role", &PolicyDoc::new(deny_policy), now);
iam_sys
.store
.cache
.add_or_update_group_policy("testgroup", &MappedPolicy::new("deny-assume-role"), now);
let groups = Some(vec!["testgroup".to_string()]);
let claims = HashMap::new();
let conditions = HashMap::new();
let args = Args {
account: "sts-fallback-test-parent",
groups: &groups,
action: Action::StsAction(StsAction::AssumeRoleAction),
bucket: "",
conditions: &conditions,
is_owner: false,
object: "",
claims: &claims,
deny_only: true,
};
let prepared = iam_sys.prepare_regular_auth(&args).await;
assert!(matches!(prepared.mode, PreparedIamMode::Regular { .. }));
assert!(
!iam_sys.eval_prepared(&prepared, &args).await,
"group-derived explicit Deny must remain authoritative"
);
}
#[tokio::test]
async fn regular_deny_only_rejects_unresolved_group_mapping() {
let iam_sys = test_iam_sys().await;
iam_sys.store.cache.add_or_update_group_policy(
"testgroup",
&MappedPolicy::new("readwrite,missing-policy"),
OffsetDateTime::now_utc(),
);
let groups = Some(vec!["testgroup".to_string()]);
let claims = HashMap::new();
let conditions = HashMap::new();
let args = Args {
account: "sts-fallback-test-parent",
groups: &groups,
action: Action::StsAction(StsAction::AssumeRoleAction),
bucket: "",
conditions: &conditions,
is_owner: false,
object: "",
claims: &claims,
deny_only: true,
};
let prepared = iam_sys.prepare_regular_auth(&args).await;
assert!(matches!(prepared.mode, PreparedIamMode::Deny));
assert!(
!iam_sys.eval_prepared(&prepared, &args).await,
"deny-only evaluation must fail closed for an unresolved group-derived mapping"
);
}
#[tokio::test] #[tokio::test]
async fn test_sts_claim_policy_ignores_unsafe_and_missing_policy_names() { async fn test_sts_claim_policy_ignores_unsafe_and_missing_policy_names() {
let store = StsTestMockStore::new(true); let store = StsTestMockStore::new(true);
+354 -5
View File
@@ -49,7 +49,10 @@
#[macro_use] #[macro_use]
extern crate metrics; extern crate metrics;
use std::sync::atomic::{AtomicBool, AtomicU64, Ordering}; use std::sync::{
Mutex,
atomic::{AtomicBool, AtomicU64, Ordering},
};
/// Global switch for detailed per-stage PUT metrics (path label, stage durations). /// Global switch for detailed per-stage PUT metrics (path label, stage durations).
/// When `false`, `record_put_object_path` and `record_put_object_stage_duration` /// When `false`, `record_put_object_path` and `record_put_object_stage_duration`
@@ -269,6 +272,12 @@ pub use collector::MetricsCollector;
pub use performance::PerformanceMetrics; pub use performance::PerformanceMetrics;
static EC_ENCODE_INFLIGHT_BYTES: AtomicU64 = AtomicU64::new(0); static EC_ENCODE_INFLIGHT_BYTES: AtomicU64 = AtomicU64::new(0);
static EC_ENCODE_PRODUCER_BYTES_CURRENT: AtomicU64 = AtomicU64::new(0);
static EC_ENCODE_PRODUCER_BYTES_PEAK: AtomicU64 = AtomicU64::new(0);
static EC_ENCODE_QUEUE_BYTES_PEAK: AtomicU64 = AtomicU64::new(0);
static EC_ENCODE_WRITER_BYTES_CURRENT: AtomicU64 = AtomicU64::new(0);
static EC_ENCODE_WRITER_BYTES_PEAK: AtomicU64 = AtomicU64::new(0);
static EC_ENCODE_PEAK_PUBLISH_LOCK: Mutex<()> = Mutex::new(());
static GET_OBJECT_BUFFERED_BYTES: AtomicU64 = AtomicU64::new(0); static GET_OBJECT_BUFFERED_BYTES: AtomicU64 = AtomicU64::new(0);
const SHARD_READ_COST_LOCAL: &str = "local"; const SHARD_READ_COST_LOCAL: &str = "local";
const SHARD_READ_COST_REMOTE: &str = "remote"; const SHARD_READ_COST_REMOTE: &str = "remote";
@@ -313,6 +322,49 @@ fn saturating_sub_atomic(counter: &AtomicU64, bytes: u64) -> u64 {
} }
} }
#[inline(always)]
fn update_peak_atomic(counter: &AtomicU64, value: u64) -> Option<u64> {
let mut peak = counter.load(Ordering::Relaxed);
loop {
if value <= peak {
return None;
}
match counter.compare_exchange_weak(peak, value, Ordering::Relaxed, Ordering::Relaxed) {
Ok(_) => return Some(value),
Err(actual) => peak = actual,
}
}
}
#[derive(Clone, Copy)]
enum EcEncodePeakMetric {
Producer,
Queue,
Writer,
}
fn publish_ec_encode_peak_with(counter: &AtomicU64, value: u64, before_publish_lock: impl FnOnce(), set_gauge: impl FnOnce(f64)) {
if update_peak_atomic(counter, value).is_none() {
return;
}
before_publish_lock();
let _guard = EC_ENCODE_PEAK_PUBLISH_LOCK.lock().unwrap_or_else(|error| error.into_inner());
set_gauge(counter.load(Ordering::Relaxed) as f64);
}
fn publish_ec_encode_peak(counter: &AtomicU64, metric: EcEncodePeakMetric, value: u64) {
publish_ec_encode_peak_with(
counter,
value,
|| {},
|peak| match metric {
EcEncodePeakMetric::Producer => gauge_set_cached!("rustfs_ec_encode_producer_bytes_peak", peak),
EcEncodePeakMetric::Queue => gauge_set_cached!("rustfs_ec_encode_queue_bytes_peak", peak),
EcEncodePeakMetric::Writer => gauge_set_cached!("rustfs_ec_encode_writer_bytes_peak", peak),
},
);
}
#[inline(always)] #[inline(always)]
fn usize_to_f64(value: usize) -> f64 { fn usize_to_f64(value: usize) -> f64 {
value as f64 value as f64
@@ -2195,18 +2247,25 @@ pub fn record_allocator_memory_observation(backend: &'static str, observation: A
} }
} }
/// Track encoded bytes currently queued between erasure encode and disk writers. /// Track encoded bytes in the queue hand-off between erasure encode and disk writers.
///
/// This is queue occupancy, not a per-request or process-RSS memory limit. The
/// erasure encoder settles it on failed/cancelled sends, receiver drop, and
/// consumer hand-off before shard writes begin.
#[inline(always)] #[inline(always)]
pub fn add_ec_encode_inflight_bytes(bytes: usize) { pub fn add_ec_encode_inflight_bytes(bytes: usize) {
let next = EC_ENCODE_INFLIGHT_BYTES.fetch_add(bytes as u64, Ordering::Relaxed) + bytes as u64; let next = EC_ENCODE_INFLIGHT_BYTES.fetch_add(bytes as u64, Ordering::Relaxed) + bytes as u64;
gauge!("rustfs_ec_encode_inflight_bytes_current").set(next as f64); gauge_set_cached!("rustfs_ec_encode_inflight_bytes_current", next as f64);
if put_stage_metrics_enabled() {
publish_ec_encode_peak(&EC_ENCODE_QUEUE_BYTES_PEAK, EcEncodePeakMetric::Queue, next);
}
} }
/// Remove encoded bytes from the tracked erasure encode in-flight gauge. /// Remove encoded bytes from the tracked erasure encode in-flight gauge.
#[inline(always)] #[inline(always)]
pub fn remove_ec_encode_inflight_bytes(bytes: usize) { pub fn remove_ec_encode_inflight_bytes(bytes: usize) {
let next = saturating_sub_atomic(&EC_ENCODE_INFLIGHT_BYTES, bytes as u64); let next = saturating_sub_atomic(&EC_ENCODE_INFLIGHT_BYTES, bytes as u64);
gauge!("rustfs_ec_encode_inflight_bytes_current").set(next as f64); gauge_set_cached!("rustfs_ec_encode_inflight_bytes_current", next as f64);
} }
/// Return the current tracked EC encode in-flight bytes. /// Return the current tracked EC encode in-flight bytes.
@@ -2215,6 +2274,109 @@ pub fn current_ec_encode_inflight_bytes() -> u64 {
EC_ENCODE_INFLIGHT_BYTES.load(Ordering::Relaxed) EC_ENCODE_INFLIGHT_BYTES.load(Ordering::Relaxed)
} }
/// Tracks encoded payload bytes held before queue hand-off or during shard writes.
///
/// Each guard contributes to a process-wide stage total until it is dropped. The
/// reported peak therefore includes concurrent PUTs, but excludes reader,
/// allocator, and transport buffers; it is not a per-PUT or process-RSS limit.
pub struct EcEncodePayloadStageGuard {
counter: &'static AtomicU64,
bytes: u64,
enabled: bool,
current_metric: EcEncodePeakMetric,
}
impl Drop for EcEncodePayloadStageGuard {
fn drop(&mut self) {
if !self.enabled {
return;
}
let next = saturating_sub_atomic(self.counter, self.bytes);
match self.current_metric {
EcEncodePeakMetric::Producer => gauge_set_cached!("rustfs_ec_encode_producer_bytes_current", next as f64),
EcEncodePeakMetric::Queue => unreachable!("queue bytes use their own ownership guard"),
EcEncodePeakMetric::Writer => gauge_set_cached!("rustfs_ec_encode_writer_bytes_current", next as f64),
}
}
}
fn track_ec_encode_payload_stage(
bytes: usize,
counter: &'static AtomicU64,
peak: &'static AtomicU64,
metric: EcEncodePeakMetric,
) -> EcEncodePayloadStageGuard {
let enabled = put_stage_metrics_enabled();
let bytes = bytes as u64;
if enabled {
let next = counter.fetch_add(bytes, Ordering::Relaxed) + bytes;
match metric {
EcEncodePeakMetric::Producer => gauge_set_cached!("rustfs_ec_encode_producer_bytes_current", next as f64),
EcEncodePeakMetric::Queue => unreachable!("queue bytes use their own ownership guard"),
EcEncodePeakMetric::Writer => gauge_set_cached!("rustfs_ec_encode_writer_bytes_current", next as f64),
}
publish_ec_encode_peak(peak, metric, next);
}
EcEncodePayloadStageGuard {
counter,
bytes,
enabled,
current_metric: metric,
}
}
/// Track encoded producer payload bytes until queue hand-off completes.
#[inline(always)]
pub fn track_ec_encode_producer_bytes(bytes: usize) -> EcEncodePayloadStageGuard {
track_ec_encode_payload_stage(
bytes,
&EC_ENCODE_PRODUCER_BYTES_CURRENT,
&EC_ENCODE_PRODUCER_BYTES_PEAK,
EcEncodePeakMetric::Producer,
)
}
/// Track encoded payload bytes while shard writers own the batch.
#[inline(always)]
pub fn track_ec_encode_writer_bytes(bytes: usize) -> EcEncodePayloadStageGuard {
track_ec_encode_payload_stage(
bytes,
&EC_ENCODE_WRITER_BYTES_CURRENT,
&EC_ENCODE_WRITER_BYTES_PEAK,
EcEncodePeakMetric::Writer,
)
}
/// Return the process-lifetime high-water mark of encoded producer payload bytes.
#[inline(always)]
pub fn current_ec_encode_producer_bytes_peak() -> u64 {
EC_ENCODE_PRODUCER_BYTES_PEAK.load(Ordering::Relaxed)
}
/// Return the current process-wide encoded producer payload bytes.
#[inline(always)]
pub fn current_ec_encode_producer_bytes() -> u64 {
EC_ENCODE_PRODUCER_BYTES_CURRENT.load(Ordering::Relaxed)
}
/// Return the process-lifetime high-water mark of encoded queue payload bytes.
#[inline(always)]
pub fn current_ec_encode_queue_bytes_peak() -> u64 {
EC_ENCODE_QUEUE_BYTES_PEAK.load(Ordering::Relaxed)
}
/// Return the process-lifetime high-water mark of encoded writer payload bytes.
#[inline(always)]
pub fn current_ec_encode_writer_bytes_peak() -> u64 {
EC_ENCODE_WRITER_BYTES_PEAK.load(Ordering::Relaxed)
}
/// Return the current process-wide encoded writer payload bytes.
#[inline(always)]
pub fn current_ec_encode_writer_bytes() -> u64 {
EC_ENCODE_WRITER_BYTES_CURRENT.load(Ordering::Relaxed)
}
/// Track whole-object buffering on the GET path. /// Track whole-object buffering on the GET path.
#[inline(always)] #[inline(always)]
pub fn track_get_object_buffered_bytes(bytes: usize) -> Option<MemoryGaugeGuard> { pub fn track_get_object_buffered_bytes(bytes: usize) -> Option<MemoryGaugeGuard> {
@@ -2381,7 +2543,9 @@ pub fn record_io_latency_p99(latency_ms: f64) {
#[cfg(test)] #[cfg(test)]
mod tests { mod tests {
use super::*; use super::*;
use std::sync::Mutex; use metrics_util::MetricKind;
use metrics_util::debugging::{DebugValue, DebuggingRecorder};
use std::sync::{Arc, Barrier, Mutex};
// Serialize tests that mutate the process-global PUT_STAGE_METRICS_ENABLED flag. // Serialize tests that mutate the process-global PUT_STAGE_METRICS_ENABLED flag.
static METRICS_FLAG_LOCK: Mutex<()> = Mutex::new(()); static METRICS_FLAG_LOCK: Mutex<()> = Mutex::new(());
@@ -2834,6 +2998,9 @@ mod tests {
#[test] #[test]
fn test_ec_encode_inflight_bytes_tracking() { fn test_ec_encode_inflight_bytes_tracking() {
let _guard = METRICS_FLAG_LOCK.lock().unwrap_or_else(|e| e.into_inner());
set_put_stage_metrics_enabled(true);
let queue_peak_before = current_ec_encode_queue_bytes_peak();
EC_ENCODE_INFLIGHT_BYTES.store(0, Ordering::Relaxed); EC_ENCODE_INFLIGHT_BYTES.store(0, Ordering::Relaxed);
add_ec_encode_inflight_bytes(1024); add_ec_encode_inflight_bytes(1024);
add_ec_encode_inflight_bytes(2048); add_ec_encode_inflight_bytes(2048);
@@ -2841,6 +3008,188 @@ mod tests {
remove_ec_encode_inflight_bytes(2048); remove_ec_encode_inflight_bytes(2048);
remove_ec_encode_inflight_bytes(4096); remove_ec_encode_inflight_bytes(4096);
assert_eq!(current_ec_encode_inflight_bytes(), 0); assert_eq!(current_ec_encode_inflight_bytes(), 0);
assert!(
current_ec_encode_queue_bytes_peak() >= queue_peak_before.max(3072),
"queue peak must retain the largest observed queue occupancy"
);
set_put_stage_metrics_enabled(false);
}
#[test]
fn test_ec_encode_producer_and_writer_stage_guards_aggregate_and_settle() {
let _guard = METRICS_FLAG_LOCK.lock().unwrap_or_else(|e| e.into_inner());
set_put_stage_metrics_enabled(true);
let producer_bytes = 1024;
let producer_first = track_ec_encode_producer_bytes(producer_bytes);
let producer_second = track_ec_encode_producer_bytes(producer_bytes);
assert_eq!(current_ec_encode_producer_bytes(), 2048);
assert!(
current_ec_encode_producer_bytes_peak() >= 2048,
"producer peak must include simultaneous stage ownership"
);
drop((producer_first, producer_second));
assert_eq!(current_ec_encode_producer_bytes(), 0);
let writer_bytes = 2048;
let writer_first = track_ec_encode_writer_bytes(writer_bytes);
let writer_second = track_ec_encode_writer_bytes(writer_bytes);
assert_eq!(current_ec_encode_writer_bytes(), 4096);
assert!(
current_ec_encode_writer_bytes_peak() >= 4096,
"writer peak must include simultaneous stage ownership"
);
drop((writer_first, writer_second));
assert_eq!(current_ec_encode_writer_bytes(), 0);
set_put_stage_metrics_enabled(false);
}
#[test]
fn test_ec_encode_payload_stage_guards_are_noop_when_disabled() {
let _guard = METRICS_FLAG_LOCK.lock().unwrap_or_else(|e| e.into_inner());
set_put_stage_metrics_enabled(false);
let producer_current_before = current_ec_encode_producer_bytes();
let producer_peak_before = current_ec_encode_producer_bytes_peak();
let writer_current_before = current_ec_encode_writer_bytes();
let writer_peak_before = current_ec_encode_writer_bytes_peak();
let producer = track_ec_encode_producer_bytes(1024);
let writer = track_ec_encode_writer_bytes(2048);
assert_eq!(current_ec_encode_producer_bytes(), producer_current_before);
assert_eq!(current_ec_encode_producer_bytes_peak(), producer_peak_before);
assert_eq!(current_ec_encode_writer_bytes(), writer_current_before);
assert_eq!(current_ec_encode_writer_bytes_peak(), writer_peak_before);
drop((producer, writer));
assert_eq!(current_ec_encode_producer_bytes(), producer_current_before);
assert_eq!(current_ec_encode_producer_bytes_peak(), producer_peak_before);
assert_eq!(current_ec_encode_writer_bytes(), writer_current_before);
assert_eq!(current_ec_encode_writer_bytes_peak(), writer_peak_before);
}
fn assert_concurrent_ec_encode_stage_guards_aggregate_and_settle(
track: fn(usize) -> EcEncodePayloadStageGuard,
current: fn() -> u64,
peak: fn() -> u64,
) {
const WORKERS: usize = 4;
const STAGE_BYTES: usize = 1536;
let current_before = current();
let peak_before = peak();
let entered = Arc::new(Barrier::new(WORKERS + 1));
let release = Arc::new(Barrier::new(WORKERS + 1));
let mut workers = Vec::with_capacity(WORKERS);
for _ in 0..WORKERS {
let entered = Arc::clone(&entered);
let release = Arc::clone(&release);
workers.push(std::thread::spawn(move || {
let stage = track(STAGE_BYTES);
entered.wait();
release.wait();
drop(stage);
}));
}
entered.wait();
let expected_delta = u64::try_from(WORKERS).expect("worker count should fit in u64")
* u64::try_from(STAGE_BYTES).expect("stage bytes should fit in u64");
assert_eq!(current(), current_before + expected_delta);
assert!(
peak() >= peak_before.max(current_before + expected_delta),
"stage peak must expose process-wide concurrent ownership"
);
release.wait();
for worker in workers {
worker.join().expect("stage worker should not panic");
}
assert_eq!(current(), current_before);
}
#[test]
fn test_ec_encode_payload_stage_guards_track_concurrent_ownership() {
let _guard = METRICS_FLAG_LOCK.lock().unwrap_or_else(|e| e.into_inner());
set_put_stage_metrics_enabled(true);
assert_concurrent_ec_encode_stage_guards_aggregate_and_settle(
track_ec_encode_producer_bytes,
current_ec_encode_producer_bytes,
current_ec_encode_producer_bytes_peak,
);
assert_concurrent_ec_encode_stage_guards_aggregate_and_settle(
track_ec_encode_writer_bytes,
current_ec_encode_writer_bytes,
current_ec_encode_writer_bytes_peak,
);
set_put_stage_metrics_enabled(false);
}
#[test]
fn test_ec_encode_producer_peak_exports_the_high_water_mark() {
let _guard = METRICS_FLAG_LOCK.lock().unwrap_or_else(|e| e.into_inner());
assert_eq!(current_ec_encode_producer_bytes(), 0, "test must start without producer stage ownership");
let previous_peak = EC_ENCODE_PRODUCER_BYTES_PEAK.swap(0, Ordering::Relaxed);
let recorder = DebuggingRecorder::new();
let snapshotter = recorder.snapshotter();
metrics::with_local_recorder(&recorder, || {
set_put_stage_metrics_enabled(true);
let first = track_ec_encode_producer_bytes(1024);
let second = track_ec_encode_producer_bytes(2048);
drop((first, second));
set_put_stage_metrics_enabled(false);
});
let exported_peak = snapshotter
.snapshot()
.into_vec()
.into_iter()
.find_map(|(composite, _, _, value)| {
(composite.kind() == MetricKind::Gauge && composite.key().name() == "rustfs_ec_encode_producer_bytes_peak")
.then_some(value)
});
assert!(
matches!(exported_peak, Some(DebugValue::Gauge(value)) if value.0 == 3072.0),
"exported producer peak must retain the aggregate high-water mark"
);
EC_ENCODE_PRODUCER_BYTES_PEAK.fetch_max(previous_peak, Ordering::Relaxed);
}
#[test]
fn test_ec_encode_peak_publish_does_not_regress_after_out_of_order_cas() {
let _guard = METRICS_FLAG_LOCK.lock().unwrap_or_else(|e| e.into_inner());
let peak = Arc::new(AtomicU64::new(0));
let exported = Arc::new(AtomicU64::new(0));
let first_ready = Arc::new(Barrier::new(2));
let release_first = Arc::new(Barrier::new(2));
let first_peak = Arc::clone(&peak);
let first_exported = Arc::clone(&exported);
let first_ready_for_thread = Arc::clone(&first_ready);
let release_first_for_thread = Arc::clone(&release_first);
let first = std::thread::spawn(move || {
publish_ec_encode_peak_with(
&first_peak,
10,
|| {
first_ready_for_thread.wait();
release_first_for_thread.wait();
},
|value| first_exported.store(value as u64, Ordering::Relaxed),
);
});
first_ready.wait();
publish_ec_encode_peak_with(&peak, 20, || {}, |value| exported.store(value as u64, Ordering::Relaxed));
release_first.wait();
first.join().expect("first publisher should not panic");
assert_eq!(peak.load(Ordering::Relaxed), 20);
assert_eq!(exported.load(Ordering::Relaxed), 20, "late publisher must reload the high-water mark");
} }
#[test] #[test]
+17
View File
@@ -64,6 +64,10 @@ md-5 = { workspace = true }
arc-swap = { workspace = true } arc-swap = { workspace = true }
rustfs-utils = { workspace = true } rustfs-utils = { workspace = true }
rustfs-security-governance = { workspace = true } rustfs-security-governance = { workspace = true }
# `EventName` for KMS audit records. A leaf crate with no rustfs dependencies,
# so the audit sink can live outside this crate without a second, drifting
# copy of the event vocabulary.
rustfs-s3-types = { workspace = true }
# HTTP client for Vault # HTTP client for Vault
reqwest = { workspace = true } reqwest = { workspace = true }
@@ -73,6 +77,16 @@ vaultrs = { workspace = true }
rustify = { workspace = true } rustify = { workspace = true }
tokio-util = { workspace = true } tokio-util = { workspace = true }
# AWS KMS backend. Credentials come from the standard aws-config provider chain
# (environment, shared profile, IMDS/container roles); this crate never handles
# AWS credential material itself.
aws-config = { workspace = true }
aws-sdk-kms = { workspace = true, default-features = false, features = ["default-https-client", "rt-tokio"] }
# SdkError variants and raw HTTP status are needed to classify AWS failures for
# the operation policy's retry decisions.
aws-smithy-runtime-api = { workspace = true, features = ["http-1x"] }
aws-smithy-types = { workspace = true }
[dev-dependencies] [dev-dependencies]
anyhow = { workspace = true } anyhow = { workspace = true }
# Debugging recorder for asserting emitted metrics in tests. # Debugging recorder for asserting emitted metrics in tests.
@@ -82,6 +96,9 @@ tempfile = { workspace = true }
temp-env = { workspace = true } temp-env = { workspace = true }
# "net" backs the scripted loopback Vault used by the policy wiring tests. # "net" backs the scripted loopback Vault used by the policy wiring tests.
tokio = { workspace = true, features = ["net", "test-util"] } tokio = { workspace = true, features = ["net", "test-util"] }
# Replays canned AWS KMS HTTP exchanges so the AWS backend tests stay offline.
aws-smithy-http-client = { workspace = true, default-features = false, features = ["test-util"] }
http = { workspace = true }
[features] [features]
default = [] default = []
+5 -3
View File
@@ -227,10 +227,12 @@ async fn main() -> Result<(), Box<dyn std::error::Error>> {
// Step 12: Check cache statistics // Step 12: Check cache statistics
println!("12. Checking cache statistics..."); println!("12. Checking cache statistics...");
if let Some((hits, misses)) = encryption_service.cache_stats().await { if let Some(stats) = encryption_service.cache_stats().await {
println!(" ✓ Cache statistics:"); println!(" ✓ Cache statistics:");
println!(" - Cache hits: {}", hits); println!(" - Cached entries: {}", stats.entries);
println!(" - Cache misses: {}\n", misses); println!(" - Cache hits: {}", stats.hits);
println!(" - Cache misses: {}", stats.misses);
println!(" - Entries evicted: {}\n", stats.evictions);
} else { } else {
println!(" - Cache is disabled\n"); println!(" - Cache is disabled\n");
} }
+5 -3
View File
@@ -263,10 +263,12 @@ async fn main() -> Result<(), Box<dyn std::error::Error>> {
// Step 12: Check cache statistics // Step 12: Check cache statistics
println!("12. Checking cache statistics..."); println!("12. Checking cache statistics...");
if let Some((hits, misses)) = encryption_service.cache_stats().await { if let Some(stats) = encryption_service.cache_stats().await {
println!(" ✓ Cache statistics:"); println!(" ✓ Cache statistics:");
println!(" - Cache hits: {}", hits); println!(" - Cached entries: {}", stats.entries);
println!(" - Cache misses: {}\n", misses); println!(" - Cache hits: {}", stats.hits);
println!(" - Cache misses: {}", stats.misses);
println!(" - Entries evicted: {}\n", stats.evictions);
} else { } else {
println!(" - Cache is disabled\n"); println!(" - Cache is disabled\n");
} }
+52 -19
View File
@@ -15,9 +15,9 @@
//! API types for KMS dynamic configuration //! API types for KMS dynamic configuration
use crate::config::{ use crate::config::{
BackendConfig, CacheConfig, DEFAULT_VAULT_TRANSIT_METADATA_KEY_PREFIX, DEFAULT_VAULT_TRANSIT_METADATA_KV_MOUNT, KmsBackend, BackendConfig, CacheConfig, DEFAULT_CACHE_TTL, DEFAULT_MAX_CACHED_KEYS, DEFAULT_VAULT_TRANSIT_METADATA_KEY_PREFIX,
KmsConfig, LocalConfig, StaticConfig, TlsConfig, VaultAuthMethod, VaultConfig, VaultTransitConfig, redacted_secret, DEFAULT_VAULT_TRANSIT_METADATA_KV_MOUNT, KmsBackend, KmsConfig, LocalConfig, StaticConfig, TlsConfig, VaultAuthMethod,
redacted_secret_option, VaultConfig, VaultTransitConfig, allow_immediate_deletion_from_env, redacted_secret, redacted_secret_option,
}; };
use crate::service_manager::KmsServiceStatus; use crate::service_manager::KmsServiceStatus;
use crate::types::{KeyMetadata, KeyUsage}; use crate::types::{KeyMetadata, KeyUsage};
@@ -418,14 +418,26 @@ pub enum BackendSummary {
/// Configured key identifier /// Configured key identifier
key_id: String, key_id: String,
}, },
/// AWS KMS backend summary
Aws {
/// Configured region, when pinned instead of resolved by the AWS chain
region: Option<String>,
/// Endpoint override, when set for an emulator or private endpoint
endpoint_url: Option<String>,
},
} }
impl From<&KmsConfig> for KmsConfigSummary { impl From<&KmsConfig> for KmsConfigSummary {
fn from(config: &KmsConfig) -> Self { fn from(config: &KmsConfig) -> Self {
// Report the lifetime the cache was built with, not the raw configured
// value: an oversized `ttl` is clamped rather than rejected, and this
// response is what operators check the cache against.
let cache_ttl_seconds = config.cache_config.effective_ttl().as_secs();
let cache_summary = if config.enable_cache { let cache_summary = if config.enable_cache {
Some(CacheSummary { Some(CacheSummary {
max_keys: config.cache_config.max_keys, max_keys: config.cache_config.max_keys,
ttl_seconds: config.cache_config.ttl.as_secs(), ttl_seconds: cache_ttl_seconds,
enable_metrics: config.cache_config.enable_metrics, enable_metrics: config.cache_config.enable_metrics,
}) })
} else { } else {
@@ -467,6 +479,10 @@ impl From<&KmsConfig> for KmsConfigSummary {
BackendConfig::Static(static_config) => BackendSummary::Static { BackendConfig::Static(static_config) => BackendSummary::Static {
key_id: static_config.key_id.clone(), key_id: static_config.key_id.clone(),
}, },
BackendConfig::Aws(aws_config) => BackendSummary::Aws {
region: aws_config.region.clone(),
endpoint_url: aws_config.endpoint_url.clone(),
},
}; };
Self { Self {
@@ -476,7 +492,7 @@ impl From<&KmsConfig> for KmsConfigSummary {
retry_attempts: config.retry_attempts, retry_attempts: config.retry_attempts,
enable_cache: config.enable_cache, enable_cache: config.enable_cache,
max_cached_keys: config.cache_config.max_keys, max_cached_keys: config.cache_config.max_keys,
cache_ttl_seconds: config.cache_config.ttl.as_secs(), cache_ttl_seconds,
cache_summary, cache_summary,
backend_summary, backend_summary,
} }
@@ -495,13 +511,17 @@ impl ConfigureLocalKmsRequest {
file_permissions: self.file_permissions, file_permissions: self.file_permissions,
}), }),
allow_insecure_dev_defaults: self.allow_insecure_dev_defaults.unwrap_or(false), allow_insecure_dev_defaults: self.allow_insecure_dev_defaults.unwrap_or(false),
// Read from server configuration, never from the request body: the
// gate must mean the same thing whether KMS was configured at
// startup or through this endpoint.
allow_immediate_deletion: allow_immediate_deletion_from_env(),
timeout: Duration::from_secs(self.timeout_seconds.unwrap_or(30)), timeout: Duration::from_secs(self.timeout_seconds.unwrap_or(30)),
retry_attempts: self.retry_attempts.unwrap_or(3), retry_attempts: self.retry_attempts.unwrap_or(3),
enable_cache: self.enable_cache.unwrap_or(true), enable_cache: self.enable_cache.unwrap_or(true),
cache_config: CacheConfig { cache_config: CacheConfig {
max_keys: self.max_cached_keys.unwrap_or(1000), max_keys: self.max_cached_keys.unwrap_or(DEFAULT_MAX_CACHED_KEYS),
ttl: Duration::from_secs(self.cache_ttl_seconds.unwrap_or(3600)), ttl: self.cache_ttl_seconds.map_or(DEFAULT_CACHE_TTL, Duration::from_secs),
enable_metrics: true, ..CacheConfig::default()
}, },
} }
} }
@@ -532,13 +552,17 @@ impl ConfigureVaultKmsRequest {
}, },
})), })),
allow_insecure_dev_defaults: self.allow_insecure_dev_defaults.unwrap_or(false), allow_insecure_dev_defaults: self.allow_insecure_dev_defaults.unwrap_or(false),
// Read from server configuration, never from the request body: the
// gate must mean the same thing whether KMS was configured at
// startup or through this endpoint.
allow_immediate_deletion: allow_immediate_deletion_from_env(),
timeout: Duration::from_secs(self.timeout_seconds.unwrap_or(30)), timeout: Duration::from_secs(self.timeout_seconds.unwrap_or(30)),
retry_attempts: self.retry_attempts.unwrap_or(3), retry_attempts: self.retry_attempts.unwrap_or(3),
enable_cache: self.enable_cache.unwrap_or(true), enable_cache: self.enable_cache.unwrap_or(true),
cache_config: CacheConfig { cache_config: CacheConfig {
max_keys: self.max_cached_keys.unwrap_or(1000), max_keys: self.max_cached_keys.unwrap_or(DEFAULT_MAX_CACHED_KEYS),
ttl: Duration::from_secs(self.cache_ttl_seconds.unwrap_or(3600)), ttl: self.cache_ttl_seconds.map_or(DEFAULT_CACHE_TTL, Duration::from_secs),
enable_metrics: true, ..CacheConfig::default()
}, },
} }
} }
@@ -569,13 +593,17 @@ impl ConfigureVaultTransitKmsRequest {
}, },
})), })),
allow_insecure_dev_defaults: self.allow_insecure_dev_defaults.unwrap_or(false), allow_insecure_dev_defaults: self.allow_insecure_dev_defaults.unwrap_or(false),
// Read from server configuration, never from the request body: the
// gate must mean the same thing whether KMS was configured at
// startup or through this endpoint.
allow_immediate_deletion: allow_immediate_deletion_from_env(),
timeout: Duration::from_secs(self.timeout_seconds.unwrap_or(30)), timeout: Duration::from_secs(self.timeout_seconds.unwrap_or(30)),
retry_attempts: self.retry_attempts.unwrap_or(3), retry_attempts: self.retry_attempts.unwrap_or(3),
enable_cache: self.enable_cache.unwrap_or(true), enable_cache: self.enable_cache.unwrap_or(true),
cache_config: CacheConfig { cache_config: CacheConfig {
max_keys: self.max_cached_keys.unwrap_or(1000), max_keys: self.max_cached_keys.unwrap_or(DEFAULT_MAX_CACHED_KEYS),
ttl: Duration::from_secs(self.cache_ttl_seconds.unwrap_or(3600)), ttl: self.cache_ttl_seconds.map_or(DEFAULT_CACHE_TTL, Duration::from_secs),
enable_metrics: true, ..CacheConfig::default()
}, },
} }
} }
@@ -592,13 +620,17 @@ impl ConfigureStaticKmsRequest {
secret_key: self.secret_key.clone(), secret_key: self.secret_key.clone(),
}), }),
allow_insecure_dev_defaults: self.allow_insecure_dev_defaults.unwrap_or(false), allow_insecure_dev_defaults: self.allow_insecure_dev_defaults.unwrap_or(false),
// Read from server configuration, never from the request body: the
// gate must mean the same thing whether KMS was configured at
// startup or through this endpoint.
allow_immediate_deletion: allow_immediate_deletion_from_env(),
timeout: Duration::from_secs(self.timeout_seconds.unwrap_or(30)), timeout: Duration::from_secs(self.timeout_seconds.unwrap_or(30)),
retry_attempts: self.retry_attempts.unwrap_or(3), retry_attempts: self.retry_attempts.unwrap_or(3),
enable_cache: self.enable_cache.unwrap_or(true), enable_cache: self.enable_cache.unwrap_or(true),
cache_config: CacheConfig { cache_config: CacheConfig {
max_keys: self.max_cached_keys.unwrap_or(1000), max_keys: self.max_cached_keys.unwrap_or(DEFAULT_MAX_CACHED_KEYS),
ttl: Duration::from_secs(self.cache_ttl_seconds.unwrap_or(3600)), ttl: self.cache_ttl_seconds.map_or(DEFAULT_CACHE_TTL, Duration::from_secs),
enable_metrics: true, ..CacheConfig::default()
}, },
} }
} }
@@ -847,6 +879,7 @@ mod tests {
tls: None, tls: None,
})), })),
allow_insecure_dev_defaults: true, allow_insecure_dev_defaults: true,
allow_immediate_deletion: false,
timeout: Duration::from_secs(30), timeout: Duration::from_secs(30),
retry_attempts: 3, retry_attempts: 3,
enable_cache: true, enable_cache: true,
@@ -858,8 +891,8 @@ mod tests {
assert_eq!(summary.backend_type, KmsBackend::VaultTransit); assert_eq!(summary.backend_type, KmsBackend::VaultTransit);
assert_eq!(summary.timeout_seconds, 30); assert_eq!(summary.timeout_seconds, 30);
assert_eq!(summary.retry_attempts, 3); assert_eq!(summary.retry_attempts, 3);
assert_eq!(summary.max_cached_keys, 1000); assert_eq!(summary.max_cached_keys, DEFAULT_MAX_CACHED_KEYS);
assert_eq!(summary.cache_ttl_seconds, 3600); assert_eq!(summary.cache_ttl_seconds, DEFAULT_CACHE_TTL.as_secs());
match summary.backend_summary { match summary.backend_summary {
BackendSummary::VaultTransit { BackendSummary::VaultTransit {
+424
View File
@@ -0,0 +1,424 @@
// Copyright 2024 RustFS Team
//
// Licensed under the Apache License, Version 2.0 (the "License");
// you may not use this file except in compliance with the License.
// You may obtain a copy of the License at
//
// http://www.apache.org/licenses/LICENSE-2.0
//
// Unless required by applicable law or agreed to in writing, software
// distributed under the License is distributed on an "AS IS" BASIS,
// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
// See the License for the specific language governing permissions and
// limitations under the License.
//! Audit contract for KMS management operations.
//!
//! [`KmsManager`](crate::manager::KmsManager) builds a [`KmsAuditRecord`] for
//! every management operation it serves and hands it to the installed
//! [`KmsAuditSink`]. No sink is installed by default, so a deployment that
//! does not consume KMS audit records pays nothing beyond the `Option` check.
//!
//! The record is deliberately a KMS-side type rather than an audit-crate
//! entry: keeping the dependency edge out of this crate lets the server map
//! records onto its existing audit pipeline (and its delivery semantics)
//! without KMS having to know that pipeline exists.
//!
//! # Redaction
//!
//! A record has no field that can hold key material, and every value that
//! originates from a caller passes through [`redact_encryption_context`]
//! before it is stored. Adding a field here means re-checking that invariant.
use crate::error::KmsError;
use crate::types::OperationContext;
use rustfs_s3_types::EventName;
use sha2::{Digest, Sha256};
use std::collections::{BTreeMap, HashMap};
use std::time::Duration;
use uuid::Uuid;
/// Encryption-context keys that describe *where* an object lives rather than
/// anything secret about it. Their values are recorded verbatim; every other
/// key is reduced to a digest.
const ENCRYPTION_CONTEXT_ALLOWLIST: [&str; 5] = ["bucket", "object", "object_key", "algorithm", "sse_type"];
/// Prefix marking a value that was replaced by a digest of itself.
const DIGEST_PREFIX: &str = "sha256:";
/// Hex characters of the digest that are kept. Enough to correlate repeated
/// values across records without being reversible in practice.
const DIGEST_LEN: usize = 16;
/// A KMS management operation that produces an audit record.
#[derive(Debug, Clone, Copy, PartialEq, Eq, Hash)]
pub enum KmsAuditOperation {
/// Creation of a new master key.
CreateKey,
/// Metadata lookup for a single key.
DescribeKey,
/// Enumeration of keys.
ListKeys,
/// Scheduling a key for deletion after the pending window.
ScheduleKeyDeletion,
/// Cancelling a previously scheduled deletion.
CancelKeyDeletion,
/// Returning a disabled key to service.
EnableKey,
/// Taking a key out of service without destroying it.
DisableKey,
/// Rotating a key to a new version.
RotateKey,
/// Irreversible removal of key material once the pending window expired.
DeleteKey,
}
impl KmsAuditOperation {
/// Stable operation name for audit consumers.
pub fn as_str(&self) -> &'static str {
match self {
KmsAuditOperation::CreateKey => "CreateKey",
KmsAuditOperation::DescribeKey => "DescribeKey",
KmsAuditOperation::ListKeys => "ListKeys",
KmsAuditOperation::ScheduleKeyDeletion => "ScheduleKeyDeletion",
KmsAuditOperation::CancelKeyDeletion => "CancelKeyDeletion",
KmsAuditOperation::EnableKey => "EnableKey",
KmsAuditOperation::DisableKey => "DisableKey",
KmsAuditOperation::RotateKey => "RotateKey",
KmsAuditOperation::DeleteKey => "DeleteKey",
}
}
/// Notification/audit event name carried by records for this operation.
pub fn event_name(&self) -> EventName {
match self {
KmsAuditOperation::CreateKey => EventName::KmsKeyCreated,
KmsAuditOperation::DescribeKey | KmsAuditOperation::ListKeys => EventName::KmsKeyAccessed,
KmsAuditOperation::ScheduleKeyDeletion => EventName::KmsKeyDeletionScheduled,
KmsAuditOperation::CancelKeyDeletion => EventName::KmsKeyDeletionCancelled,
KmsAuditOperation::EnableKey => EventName::KmsKeyEnabled,
KmsAuditOperation::DisableKey => EventName::KmsKeyDisabled,
KmsAuditOperation::RotateKey => EventName::KmsKeyRotated,
KmsAuditOperation::DeleteKey => EventName::KmsKeyDeleted,
}
}
}
/// Whether the audited operation completed.
#[derive(Debug, Clone, Copy, PartialEq, Eq, Hash)]
pub enum KmsAuditOutcome {
/// The operation completed and its effect is durable.
Success,
/// The operation failed; `error_class` carries the reason category.
Failure,
}
impl KmsAuditOutcome {
/// Stable outcome name for audit consumers.
pub fn as_str(&self) -> &'static str {
match self {
KmsAuditOutcome::Success => "success",
KmsAuditOutcome::Failure => "failure",
}
}
}
/// Coarse, stable classification of a failed KMS operation.
///
/// Audit consumers alert on these, so the strings are a wire contract: add
/// new classes rather than renaming existing ones.
pub fn error_class(error: &KmsError) -> &'static str {
match error {
KmsError::ConfigurationError { .. } => "configuration",
KmsError::KeyNotFound { .. } => "key_not_found",
KmsError::InvalidKey { .. } => "invalid_key",
KmsError::CryptographicError { .. } => "cryptographic",
KmsError::BackendError { .. } => "backend",
KmsError::AccessDenied { .. } => "access_denied",
KmsError::KeyAlreadyExists { .. } => "key_already_exists",
KmsError::InvalidOperation { .. } => "invalid_operation",
KmsError::InternalError { .. } => "internal",
KmsError::SerializationError { .. } => "serialization",
KmsError::IoError { .. } => "io",
KmsError::CacheError { .. } => "cache",
KmsError::ValidationError { .. } => "validation",
KmsError::UnsupportedAlgorithm { .. } => "unsupported_algorithm",
KmsError::InvalidKeySize { .. } => "invalid_key_size",
KmsError::ContextMismatch { .. } => "context_mismatch",
KmsError::OperationTimedOut { .. } => "timeout",
KmsError::OperationCancelled { .. } => "cancelled",
KmsError::MaterialMissing { .. } => "material_missing",
KmsError::MaterialCorrupt { .. } => "material_corrupt",
KmsError::MaterialAuthenticationFailed { .. } => "material_authentication_failed",
KmsError::UnsupportedFormatVersion { .. } => "unsupported_format_version",
KmsError::KeyVersionNotFound { .. } => "key_version_not_found",
KmsError::Backup(_) => "backup",
KmsError::UnsupportedCapability { .. } => "unsupported_capability",
KmsError::CredentialsUnavailable { .. } => "credentials_unavailable",
KmsError::BaselineVersionLost { .. } => "baseline_version_lost",
}
}
/// One audited KMS management operation.
///
/// Everything an audit consumer needs to answer "who did what to which key,
/// against which backend, and did it work" — and nothing that could carry key
/// material. See the module docs for the redaction rules.
///
/// Tenant attribution is intentionally absent: multi-tenancy is not modelled
/// yet, and an invented tenant value is worse than a missing one. It will
/// arrive as an entry in [`Self::context`] once tenancy lands.
#[derive(Debug, Clone)]
pub struct KmsAuditRecord {
/// Correlates every record emitted for one logical request.
pub operation_id: Uuid,
/// The audited operation.
pub operation: KmsAuditOperation,
/// Event name published to audit consumers.
pub event: EventName,
/// Authenticated identity that requested the operation, or
/// [`OperationContext::INTERNAL_PRINCIPAL`] for server-initiated work.
pub principal: String,
/// Client address, when the caller supplied one.
pub source_ip: Option<String>,
/// Client user agent, when the caller supplied one.
pub user_agent: Option<String>,
/// Key the operation acted on. `None` for operations that span keys, such
/// as listing.
pub key_id: Option<String>,
/// Key version the operation resolved to, when the result identifies one.
/// Operations on a key as a whole leave this unset.
pub key_version: Option<u32>,
/// Whether the operation succeeded.
pub outcome: KmsAuditOutcome,
/// Failure category; `None` on success. See [`error_class`].
pub error_class: Option<&'static str>,
/// Backend that served the operation, e.g. `local` or `vault-transit`.
pub backend: &'static str,
/// Wall-clock time spent in the audited operation.
pub latency: Duration,
/// Retries performed above the audit point. `None` means the number is
/// not observable here: backend-internal retries are accounted for by the
/// `rustfs_kms_backend_operation_attempts` metric instead of being
/// duplicated (and possibly contradicted) in the audit trail.
pub retry_count: Option<u32>,
/// Server-supplied correlation values carried by the operation context.
/// Also the reserved slot for tenant attribution.
pub context: BTreeMap<String, String>,
/// Caller-supplied encryption context, after [`redact_encryption_context`].
pub encryption_context: BTreeMap<String, String>,
}
impl KmsAuditRecord {
/// Start a record for `operation` performed under `context`.
///
/// The outcome defaults to failure so that a record which somehow escapes
/// without [`Self::with_result`] under-reports success rather than
/// inventing it.
pub fn new(operation: KmsAuditOperation, context: &OperationContext, backend: &'static str) -> Self {
Self {
operation_id: context.operation_id,
operation,
event: operation.event_name(),
principal: context.principal.clone(),
source_ip: context.source_ip.clone(),
user_agent: context.user_agent.clone(),
key_id: None,
key_version: None,
outcome: KmsAuditOutcome::Failure,
error_class: None,
backend,
latency: Duration::ZERO,
retry_count: None,
context: context
.additional_context
.iter()
.map(|(k, v)| (k.clone(), v.clone()))
.collect(),
encryption_context: BTreeMap::new(),
}
}
/// Attach the key the operation acted on.
pub fn with_key_id(mut self, key_id: Option<impl Into<String>>) -> Self {
self.key_id = key_id.map(Into::into);
self
}
/// Attach the key version the operation resolved to.
pub fn with_key_version(mut self, key_version: Option<u32>) -> Self {
self.key_version = key_version;
self
}
/// Attach the measured duration of the operation.
pub fn with_latency(mut self, latency: Duration) -> Self {
self.latency = latency;
self
}
/// Derive outcome and error class from the operation result.
pub fn with_result<T>(mut self, result: &crate::error::Result<T>) -> Self {
match result {
Ok(_) => {
self.outcome = KmsAuditOutcome::Success;
self.error_class = None;
}
Err(error) => {
self.outcome = KmsAuditOutcome::Failure;
self.error_class = Some(error_class(error));
}
}
self
}
/// Attach the caller-supplied encryption context, redacted.
///
/// Redaction happens here rather than at the call site so no caller can
/// place a raw context value into a record by accident.
pub fn with_encryption_context(mut self, encryption_context: &HashMap<String, String>) -> Self {
self.encryption_context = redact_encryption_context(encryption_context);
self
}
}
/// Reduce a caller-supplied encryption context to something safe to persist.
///
/// Keys naming the object's location are kept verbatim because they are the
/// reason the context is audited at all. Every other value is replaced by a
/// truncated digest, which still correlates repeated values across records
/// but does not reproduce whatever the caller put there.
pub fn redact_encryption_context(encryption_context: &HashMap<String, String>) -> BTreeMap<String, String> {
encryption_context
.iter()
.map(|(key, value)| {
let recorded = if ENCRYPTION_CONTEXT_ALLOWLIST.contains(&key.as_str()) {
value.clone()
} else {
digest_value(value)
};
(key.clone(), recorded)
})
.collect()
}
fn digest_value(value: &str) -> String {
let digest = hex::encode(Sha256::digest(value.as_bytes()));
format!("{DIGEST_PREFIX}{}", &digest[..DIGEST_LEN])
}
/// Receives KMS audit records.
///
/// Implementations must not block: they are called on the task that served
/// the KMS operation. Delivery failures are the sink's problem — the KMS
/// operation has already completed by the time a record is emitted, and its
/// result is never changed by what the sink does.
pub trait KmsAuditSink: Send + Sync {
/// Handle one audit record.
fn emit(&self, record: KmsAuditRecord);
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn every_operation_maps_to_a_kms_event() {
let operations = [
KmsAuditOperation::CreateKey,
KmsAuditOperation::DescribeKey,
KmsAuditOperation::ListKeys,
KmsAuditOperation::ScheduleKeyDeletion,
KmsAuditOperation::CancelKeyDeletion,
KmsAuditOperation::EnableKey,
KmsAuditOperation::DisableKey,
KmsAuditOperation::RotateKey,
KmsAuditOperation::DeleteKey,
];
for operation in operations {
let event = operation.event_name();
assert!(event.is_kms(), "{} must map to a KMS event, got {event}", operation.as_str());
}
}
#[test]
fn record_defaults_to_failure_until_a_result_is_attached() {
let context = OperationContext::new("tester".to_string());
let record = KmsAuditRecord::new(KmsAuditOperation::CreateKey, &context, "local");
assert_eq!(record.outcome, KmsAuditOutcome::Failure);
assert!(record.error_class.is_none());
let ok: crate::error::Result<()> = Ok(());
assert_eq!(record.clone().with_result(&ok).outcome, KmsAuditOutcome::Success);
let denied: crate::error::Result<()> = Err(KmsError::access_denied("nope"));
let failed = record.with_result(&denied);
assert_eq!(failed.outcome, KmsAuditOutcome::Failure);
assert_eq!(failed.error_class, Some("access_denied"));
}
#[test]
fn record_carries_the_full_operation_context() {
let context = OperationContext::new("arn:aws:iam::user/alice".to_string())
.with_source_ip("10.0.0.7".to_string())
.with_user_agent("aws-cli/2".to_string())
.with_context("requestID".to_string(), "req-1".to_string());
let record = KmsAuditRecord::new(KmsAuditOperation::RotateKey, &context, "vault-transit")
.with_key_id(Some("key-1"))
.with_key_version(Some(3))
.with_latency(Duration::from_millis(12));
assert_eq!(record.operation_id, context.operation_id);
assert_eq!(record.principal, "arn:aws:iam::user/alice");
assert_eq!(record.source_ip.as_deref(), Some("10.0.0.7"));
assert_eq!(record.user_agent.as_deref(), Some("aws-cli/2"));
assert_eq!(record.key_id.as_deref(), Some("key-1"));
assert_eq!(record.key_version, Some(3));
assert_eq!(record.backend, "vault-transit");
assert_eq!(record.latency, Duration::from_millis(12));
assert_eq!(record.event, EventName::KmsKeyRotated);
assert_eq!(record.context.get("requestID").map(String::as_str), Some("req-1"));
}
#[test]
fn encryption_context_keeps_only_allowlisted_values_verbatim() {
let secret = "eyJhbGciOiJIUzI1NiJ9.super-secret-grant-token";
let context = HashMap::from([
("bucket".to_string(), "photos".to_string()),
("object_key".to_string(), "2026/cat.png".to_string()),
("algorithm".to_string(), "AES256".to_string()),
("grant_token".to_string(), secret.to_string()),
("customer_reference".to_string(), "acct-4711".to_string()),
]);
let redacted = redact_encryption_context(&context);
assert_eq!(redacted.get("bucket").map(String::as_str), Some("photos"));
assert_eq!(redacted.get("object_key").map(String::as_str), Some("2026/cat.png"));
assert_eq!(redacted.get("algorithm").map(String::as_str), Some("AES256"));
for key in ["grant_token", "customer_reference"] {
let value = redacted.get(key).expect("non-allowlisted key should be kept as a digest");
assert!(value.starts_with(DIGEST_PREFIX), "{key} should be digested, got {value}");
}
let rendered = format!("{redacted:?}");
assert!(!rendered.contains(secret), "redacted context must not reproduce the raw value");
assert!(!rendered.contains("acct-4711"), "redacted context must not reproduce the raw value");
}
#[test]
fn identical_values_digest_identically_across_records() {
// Correlation is the reason a digest is preferable to a fixed
// placeholder; assert it actually holds.
let first = redact_encryption_context(&HashMap::from([("tenant".to_string(), "alpha".to_string())]));
let second = redact_encryption_context(&HashMap::from([("tenant".to_string(), "alpha".to_string())]));
let other = redact_encryption_context(&HashMap::from([("tenant".to_string(), "beta".to_string())]));
assert_eq!(first.get("tenant"), second.get("tenant"));
assert_ne!(first.get("tenant"), other.get("tenant"));
}
}
File diff suppressed because it is too large Load Diff
@@ -108,6 +108,7 @@ fn schedule_request(key_id: &str) -> DeleteKeyRequest {
key_id: key_id.to_string(), key_id: key_id.to_string(),
pending_window_in_days: Some(7), pending_window_in_days: Some(7),
force_immediate: None, force_immediate: None,
confirm_key_id: None,
} }
} }
+302 -20
View File
@@ -14,7 +14,10 @@
//! Local file-based KMS backend implementation //! Local file-based KMS backend implementation
use crate::backends::{BackendCapabilities, ExpiredKeyRemoval, KmsBackend, StateGatedOperation, ensure_key_status_permits}; use crate::backends::{
BackendCapabilities, ExpiredKeyRemoval, KmsBackend, StateGatedOperation, ensure_key_status_permits,
ensure_tag_keys_are_mutable,
};
use crate::config::KmsConfig; use crate::config::KmsConfig;
use crate::config::LocalConfig; use crate::config::LocalConfig;
use crate::encryption::{AesDekCrypto, DataKeyEnvelope, DekCrypto, generate_key_material}; use crate::encryption::{AesDekCrypto, DataKeyEnvelope, DekCrypto, generate_key_material};
@@ -46,7 +49,10 @@ use zeroize::Zeroizing;
/// `key_dir` is accepted, so identifiers already in use by existing deployments keep /// `key_dir` is accepted, so identifiers already in use by existing deployments keep
/// resolving. Only separators, traversal and the degenerate cases are refused, which is /// resolving. Only separators, traversal and the degenerate cases are refused, which is
/// what stops `key_dir.join(...)` from escaping. /// what stops `key_dir.join(...)` from escaping.
fn validate_key_id(key_id: &str) -> Result<()> { ///
/// pub(crate) because the backup restore path applies the same containment
/// rule to key identifiers recovered from bundle artifacts.
pub(crate) fn validate_key_id(key_id: &str) -> Result<()> {
if key_id.is_empty() { if key_id.is_empty() {
return Err(KmsError::invalid_key("key identifier must not be empty")); return Err(KmsError::invalid_key("key identifier must not be empty"));
} }
@@ -68,7 +74,15 @@ fn validate_key_id(key_id: &str) -> Result<()> {
} }
} }
const LOCAL_KMS_MASTER_KEY_SALT_FILE: &str = ".master-key.salt"; // The salt and restore-marker file names are pub(crate) so the backup/restore
// modules (`crate::backup`) address the exact on-disk names instead of copies
// that could drift.
pub(crate) const LOCAL_KMS_MASTER_KEY_SALT_FILE: &str = ".master-key.salt";
/// Commit marker of an in-progress Local restore cutover (see
/// `crate::backup::local_restore`). Its presence means the key directory is
/// mid-cutover: startup must fail closed until the restore is rolled forward
/// or explicitly aborted.
pub(crate) const LOCAL_RESTORE_COMMIT_MARKER_FILE: &str = ".restore-commit.json";
// The KDF parameters are pub(crate) so the backup manifest contract // The KDF parameters are pub(crate) so the backup manifest contract
// (`crate::backup`) records the exact compiled-in derivation instead of a // (`crate::backup`) records the exact compiled-in derivation instead of a
// copy that could drift. // copy that could drift.
@@ -87,7 +101,11 @@ pub(crate) const LOCAL_KMS_ARGON2_P_COST: u32 = 1;
/// `foo.tmp-<uuid>` is stored as `foo.tmp-<uuid>.key` — so the `.key` guard /// `foo.tmp-<uuid>` is stored as `foo.tmp-<uuid>.key` — so the `.key` guard
/// plus the exact hyphenated-UUID check makes it impossible to match an /// plus the exact hyphenated-UUID check makes it impossible to match an
/// authoritative file. /// authoritative file.
fn is_orphan_commit_temp_name(file_name: &str) -> bool { ///
/// pub(crate) because the backup restore path applies the same classification
/// when it re-enters an interrupted run: a leftover commit temp is never
/// authoritative state, so it does not make a target non-empty.
pub(crate) fn is_orphan_commit_temp_name(file_name: &str) -> bool {
if file_name.ends_with(".key") { if file_name.ends_with(".key") {
return false; return false;
} }
@@ -110,12 +128,15 @@ fn is_orphan_commit_temp_name(file_name: &str) -> bool {
/// ///
/// This intentionally mirrors ecstore's fsync helpers without depending on the /// This intentionally mirrors ecstore's fsync helpers without depending on the
/// ecstore crate: the KMS backend stays decoupled from storage internals. /// ecstore crate: the KMS backend stays decoupled from storage internals.
mod durable_file { ///
/// pub(crate) because the backup restore path (`crate::backup::local_restore`)
/// commits staged files and its cutover marker through the same protocol.
pub(crate) mod durable_file {
use std::io::{self, Write}; use std::io::{self, Write};
use std::path::{Path, PathBuf}; use std::path::{Path, PathBuf};
/// How the fully written temp file becomes visible under its final name. /// How the fully written temp file becomes visible under its final name.
pub(super) enum Publish { pub(crate) enum Publish {
/// Atomically replace whatever is at the destination via `rename`. /// Atomically replace whatever is at the destination via `rename`.
Replace, Replace,
/// Publish via `hard_link`, failing with [`CommitError::AlreadyExists`] /// Publish via `hard_link`, failing with [`CommitError::AlreadyExists`]
@@ -124,7 +145,7 @@ mod durable_file {
} }
#[derive(Debug)] #[derive(Debug)]
pub(super) enum CommitError { pub(crate) enum CommitError {
AlreadyExists, AlreadyExists,
Io(io::Error), Io(io::Error),
/// Test-only simulated crash: the protocol stops after the given step /// Test-only simulated crash: the protocol stops after the given step
@@ -158,14 +179,14 @@ mod durable_file {
/// to prove that every interrupted prefix recovers to either the complete /// to prove that every interrupted prefix recovers to either the complete
/// old state or the complete new state. /// old state or the complete new state.
#[derive(Debug, Clone, Copy, PartialEq, Eq)] #[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub(super) enum CommitStep { pub(crate) enum CommitStep {
TempWritten, TempWritten,
FileSynced, FileSynced,
Published, Published,
DirSynced, DirSynced,
} }
pub(super) async fn commit( pub(crate) async fn commit(
temp_path: PathBuf, temp_path: PathBuf,
final_path: PathBuf, final_path: PathBuf,
content: Vec<u8>, content: Vec<u8>,
@@ -179,7 +200,7 @@ mod durable_file {
/// Remove a published file durably: without the parent directory fsync a /// Remove a published file durably: without the parent directory fsync a
/// deleted key could resurface after power loss. /// deleted key could resurface after power loss.
pub(super) async fn remove_durably(path: PathBuf) -> io::Result<()> { pub(crate) async fn remove_durably(path: PathBuf) -> io::Result<()> {
tokio::task::spawn_blocking(move || { tokio::task::spawn_blocking(move || {
std::fs::remove_file(&path)?; std::fs::remove_file(&path)?;
let parent = path let parent = path
@@ -191,6 +212,41 @@ mod durable_file {
.map_err(io::Error::other)? .map_err(io::Error::other)?
} }
/// Publish an already-durable file under a second name via `hard_link`,
/// then fsync the destination's parent directory.
///
/// This is the restore cutover primitive: the source (a staged file that
/// went through [`commit`]) is already durable, so linking plus a parent
/// fsync is a complete publish. `AlreadyExists` is idempotent success only
/// when the destination content is byte-identical to the source — that is
/// exactly the re-entry case of a cutover interrupted after this link —
/// and a hard failure otherwise, so the primitive can never clobber or
/// silently accept foreign state.
pub(crate) async fn link_durably(source: PathBuf, dest: PathBuf) -> io::Result<()> {
tokio::task::spawn_blocking(move || {
match std::fs::hard_link(&source, &dest) {
Ok(()) => {}
Err(error) if error.kind() == io::ErrorKind::AlreadyExists => {
let existing = std::fs::read(&dest)?;
let staged = std::fs::read(&source)?;
if existing != staged {
return Err(io::Error::new(
io::ErrorKind::AlreadyExists,
format!("destination {} already exists with different content", dest.display()),
));
}
}
Err(error) => return Err(error),
}
let parent = dest
.parent()
.ok_or_else(|| io::Error::other("destination has no parent directory"))?;
fsync_dir(parent)
})
.await
.map_err(io::Error::other)?
}
fn commit_blocking( fn commit_blocking(
temp_path: &Path, temp_path: &Path,
final_path: &Path, final_path: &Path,
@@ -355,7 +411,7 @@ mod durable_file {
/// Test-only failpoints simulating a crash after a given commit step. /// Test-only failpoints simulating a crash after a given commit step.
/// Armed per directory so parallel tests never affect each other. /// Armed per directory so parallel tests never affect each other.
#[cfg(test)] #[cfg(test)]
pub(super) mod failpoint { pub(crate) mod failpoint {
use super::CommitStep; use super::CommitStep;
use std::path::{Path, PathBuf}; use std::path::{Path, PathBuf};
use std::sync::Mutex; use std::sync::Mutex;
@@ -397,6 +453,20 @@ pub struct LocalKmsClient {
/// Per-key write locks serializing read-modify-write updates within this /// Per-key write locks serializing read-modify-write updates within this
/// process (see [`Self::lock_key_for_write`]). /// process (see [`Self::lock_key_for_write`]).
key_write_locks: Mutex<HashMap<String, Arc<tokio::sync::Mutex<()>>>>, key_write_locks: Mutex<HashMap<String, Arc<tokio::sync::Mutex<()>>>>,
/// Directory-wide writer fence for backup export (see
/// [`Self::acquire_export_fence`]). Writers hold the read side; an export
/// snapshot holds the write side so it observes a single-generation view.
export_fence: Arc<tokio::sync::RwLock<()>>,
}
/// Guard pairing the export-fence read lock with a per-key write mutex.
///
/// Dropping it releases both, so every existing `lock_key_for_write` call
/// site participates in the export fence without changes.
#[must_use]
struct KeyWriteGuard {
_fence: tokio::sync::OwnedRwLockReadGuard<()>,
_key: tokio::sync::OwnedMutexGuard<()>,
} }
// pub(crate) so the backup contract tests can anchor the manifest's // pub(crate) so the backup contract tests can anchor the manifest's
@@ -446,6 +516,11 @@ impl LocalKmsClient {
debug!(path = ?config.key_dir, "KMS key directory created"); debug!(path = ?config.key_dir, "KMS key directory created");
} }
// The restore-marker guard must run before anything else touches the
// directory (in particular before salt load/creation): a directory
// mid-cutover holds an arbitrary mix of old and new state.
Self::ensure_no_restore_marker(&config).await?;
// Initialize master cipher if master key is provided // Initialize master cipher if master key is provided
let (master_cipher, legacy_master_cipher) = if let Some(ref master_key) = config.master_key { let (master_cipher, legacy_master_cipher) = if let Some(ref master_key) = config.master_key {
let salt = Self::load_or_create_master_key_salt(&config).await?; let salt = Self::load_or_create_master_key_salt(&config).await?;
@@ -463,6 +538,7 @@ impl LocalKmsClient {
legacy_master_cipher, legacy_master_cipher,
dek_crypto: AesDekCrypto::new(), dek_crypto: AesDekCrypto::new(),
key_write_locks: Mutex::new(HashMap::new()), key_write_locks: Mutex::new(HashMap::new()),
export_fence: Arc::new(tokio::sync::RwLock::new(())),
}; };
client.validate_existing_keys().await?; client.validate_existing_keys().await?;
Ok(client) Ok(client)
@@ -476,6 +552,7 @@ impl LocalKmsClient {
if !fs::try_exists(&config.key_dir).await? { if !fs::try_exists(&config.key_dir).await? {
return Err(KmsError::configuration_error("Local KMS key directory does not exist")); return Err(KmsError::configuration_error("Local KMS key directory does not exist"));
} }
Self::ensure_no_restore_marker(&config).await?;
let (master_cipher, legacy_master_cipher) = if let Some(ref master_key) = config.master_key { let (master_cipher, legacy_master_cipher) = if let Some(ref master_key) = config.master_key {
let legacy_key = Self::derive_legacy_master_key(master_key)?; let legacy_key = Self::derive_legacy_master_key(master_key)?;
@@ -505,6 +582,7 @@ impl LocalKmsClient {
legacy_master_cipher, legacy_master_cipher,
dek_crypto: AesDekCrypto::new(), dek_crypto: AesDekCrypto::new(),
key_write_locks: Mutex::new(HashMap::new()), key_write_locks: Mutex::new(HashMap::new()),
export_fence: Arc::new(tokio::sync::RwLock::new(())),
}) })
} }
@@ -515,16 +593,72 @@ impl LocalKmsClient {
/// delete with a rewrite. Cross-process writers sharing a key directory /// delete with a rewrite. Cross-process writers sharing a key directory
/// remain unsupported. Entries live for the client's lifetime; the table /// remain unsupported. Entries live for the client's lifetime; the table
/// is bounded by the number of distinct key ids this process touches. /// is bounded by the number of distinct key ids this process touches.
async fn lock_key_for_write(&self, key_id: &str) -> tokio::sync::OwnedMutexGuard<()> { async fn lock_key_for_write(&self, key_id: &str) -> KeyWriteGuard {
// Fence first, per-key mutex second: the ordering is uniform across
// all writers, so an export waiting on the write side can never
// deadlock with a writer holding a key mutex.
let fence = Arc::clone(&self.export_fence).read_owned().await;
let lock = { let lock = {
let mut locks = self.key_write_locks.lock().expect("Local KMS key write lock table poisoned"); let mut locks = self.key_write_locks.lock().expect("Local KMS key write lock table poisoned");
Arc::clone(locks.entry(key_id.to_string()).or_default()) Arc::clone(locks.entry(key_id.to_string()).or_default())
}; };
lock.lock_owned().await KeyWriteGuard {
_fence: fence,
_key: lock.lock_owned().await,
}
}
/// Block every key-directory writer while a backup export collects its
/// snapshot, so all records belong to one generation.
///
/// Mutating operations hold the read side (via [`Self::lock_key_for_write`]
/// or [`Self::save_new_master_key`]); the export holds the write side only
/// for the collection phase, never while encrypting or writing the bundle.
pub(crate) async fn acquire_export_fence(&self) -> tokio::sync::OwnedRwLockWriteGuard<()> {
Arc::clone(&self.export_fence).write_owned().await
}
/// Key directory root, exposed for the backup export module.
pub(crate) fn key_directory(&self) -> &Path {
&self.config.key_dir
}
/// Absolute path of the master-key KDF salt file, exposed for the backup
/// export module.
pub(crate) fn master_key_salt_file(&self) -> PathBuf {
Self::master_key_salt_path(&self.config)
}
/// Operator-configured master key string, exposed for the backup export
/// module so it can record a one-way verifier in the bundle manifest.
/// Never log or persist this value.
pub(crate) fn configured_master_key(&self) -> Option<&str> {
self.config.master_key.as_deref()
}
/// Fail closed while a restore cutover marker is present: the directory
/// then holds an arbitrary mix of pre-restore and restored state, and the
/// only valid next steps are re-running the restore with the same bundle
/// (roll forward) or explicitly aborting it. This mirrors the missing-salt
/// guard: startup must never paper over a half-applied restore.
async fn ensure_no_restore_marker(config: &LocalConfig) -> Result<()> {
let marker = config.key_dir.join(LOCAL_RESTORE_COMMIT_MARKER_FILE);
if fs::try_exists(&marker).await? {
return Err(KmsError::configuration_error(format!(
"Local KMS key directory has an unfinished restore (marker {} present); \
re-run the restore with the same bundle to roll it forward or abort it explicitly",
marker.display()
)));
}
Ok(())
} }
/// Derive a 256-bit key from the master key string using a persistent Argon2id salt. /// Derive a 256-bit key from the master key string using a persistent Argon2id salt.
fn derive_master_key(master_key: &str, salt: &[u8]) -> Result<Key<Aes256Gcm>> { ///
/// pub(crate) because the backup restore path derives the same key from
/// the operator-supplied master key and the bundled salt for its verifier
/// check and staged decryption probe.
pub(crate) fn derive_master_key(master_key: &str, salt: &[u8]) -> Result<Key<Aes256Gcm>> {
let params = Params::new( let params = Params::new(
LOCAL_KMS_ARGON2_M_COST_KIB, LOCAL_KMS_ARGON2_M_COST_KIB,
LOCAL_KMS_ARGON2_T_COST, LOCAL_KMS_ARGON2_T_COST,
@@ -541,7 +675,7 @@ impl LocalKmsClient {
Ok(key) Ok(key)
} }
fn derive_legacy_master_key(master_key: &str) -> Result<Key<Aes256Gcm>> { pub(crate) fn derive_legacy_master_key(master_key: &str) -> Result<Key<Aes256Gcm>> {
let mut hasher = Sha256::new(); let mut hasher = Sha256::new();
hasher.update(master_key.as_bytes()); hasher.update(master_key.as_bytes());
hasher.update(b"rustfs-kms-local"); hasher.update(b"rustfs-kms-local");
@@ -797,6 +931,11 @@ impl LocalKmsClient {
} }
async fn save_new_master_key(&self, master_key: &MasterKeyInfo, key_material: &[u8]) -> Result<()> { async fn save_new_master_key(&self, master_key: &MasterKeyInfo, key_material: &[u8]) -> Result<()> {
// Creates never take the per-key write lock (`NoClobber` publishing
// already linearizes them), so they join the export fence here. This
// must stay the only fence acquisition on the create path: the fence
// read lock is not reentrant while an export waits for the write side.
let _fence = Arc::clone(&self.export_fence).read_owned().await;
let key_path = self.master_key_path(&master_key.key_id)?; let key_path = self.master_key_path(&master_key.key_id)?;
let content = self.encode_master_key(master_key, key_material)?; let content = self.encode_master_key(master_key, key_material)?;
let temp_path = key_path.with_extension(format!("tmp-{}", uuid::Uuid::new_v4())); let temp_path = key_path.with_extension(format!("tmp-{}", uuid::Uuid::new_v4()));
@@ -1158,6 +1297,32 @@ impl LocalKmsClient {
Ok(()) Ok(())
} }
/// Read-modify-write of a key's mutable metadata under the per-key write
/// lock, carrying the existing key material over untouched.
///
/// `mutate` reports whether it changed anything; an unchanged record is
/// never rewritten, so a repeated no-op update neither rewrites the key
/// file nor risks a failed commit on an unrelated call.
async fn update_key_metadata<F>(&self, key_id: &str, mutate: F) -> Result<()>
where
F: FnOnce(&mut MasterKeyInfo) -> Result<bool>,
{
let _write_guard = self.lock_key_for_write(key_id).await;
let mut master_key = self.load_master_key(key_id).await?;
if !mutate(&mut master_key)? {
return Ok(());
}
// Preserve the existing key material (see enable_key): a metadata edit
// must never regenerate the master key, or every DEK wrapped by it
// becomes undecryptable.
let key_material = self.get_key_material(key_id).await?;
self.save_master_key(&master_key, &key_material).await?;
debug!(key_id, "Local KMS key metadata updated");
Ok(())
}
/// Test-only lifecycle driver: the product path goes through [`KmsBackend`]. /// Test-only lifecycle driver: the product path goes through [`KmsBackend`].
#[cfg(test)] #[cfg(test)]
pub(crate) async fn schedule_key_deletion( pub(crate) async fn schedule_key_deletion(
@@ -1250,7 +1415,8 @@ impl LocalKmsBackend {
crate::config::BackendConfig::Local(local_config) => local_config.clone(), crate::config::BackendConfig::Local(local_config) => local_config.clone(),
crate::config::BackendConfig::VaultKv2(_) crate::config::BackendConfig::VaultKv2(_)
| crate::config::BackendConfig::VaultTransit(_) | crate::config::BackendConfig::VaultTransit(_)
| crate::config::BackendConfig::Static(_) => { | crate::config::BackendConfig::Static(_)
| crate::config::BackendConfig::Aws(_) => {
return Err(KmsError::configuration_error("Expected Local backend configuration")); return Err(KmsError::configuration_error("Expected Local backend configuration"));
} }
}; };
@@ -1275,12 +1441,15 @@ impl KmsBackend for LocalKmsBackend {
// Generate key material // Generate key material
let key_material = generate_key_material(algorithm)?; let key_material = generate_key_material(algorithm)?;
let master_key = MasterKeyInfo::new_with_description( let mut master_key = MasterKeyInfo::new_with_description(
key_id.clone(), key_id.clone(),
algorithm.to_string(), algorithm.to_string(),
Some("local-kms".to_string()), Some("local-kms".to_string()),
request.description.clone(), request.description.clone(),
); );
// Persist the caller's tags: the response below reports them, and
// describe_key reads them back out of this record.
master_key.metadata = request.tags.clone();
// Save to disk // Save to disk
self.client.save_new_master_key(&master_key, &key_material).await?; self.client.save_new_master_key(&master_key, &key_material).await?;
@@ -1448,9 +1617,15 @@ impl KmsBackend for LocalKmsBackend {
// Schedule for deletion (default 30 days) // Schedule for deletion (default 30 days)
ensure_key_status_permits(key_id, &master_key.status, StateGatedOperation::ScheduleDeletion)?; ensure_key_status_permits(key_id, &master_key.status, StateGatedOperation::ScheduleDeletion)?;
let days = request.pending_window_in_days.unwrap_or(30); // Defensive: KmsManager::delete_key is the enforcement point for the
if !(7..=30).contains(&days) { // waiting window and rejects out-of-range requests before any
return Err(KmsError::invalid_parameter("pending_window_in_days must be between 7 and 30".to_string())); // backend runs. This repeats the bound for callers holding a backend
// handle directly (tests, maintenance tasks).
let days = request.pending_window_in_days.unwrap_or(DEFAULT_PENDING_DELETION_WINDOW_DAYS);
if !(MIN_PENDING_DELETION_WINDOW_DAYS..=MAX_PENDING_DELETION_WINDOW_DAYS).contains(&days) {
return Err(KmsError::invalid_parameter(format!(
"pending_window_in_days must be between {MIN_PENDING_DELETION_WINDOW_DAYS} and {MAX_PENDING_DELETION_WINDOW_DAYS}"
)));
} }
let deletion_date = Zoned::now() + Duration::from_secs(days as u64 * 86400); let deletion_date = Zoned::now() + Duration::from_secs(days as u64 * 86400);
@@ -1548,10 +1723,52 @@ impl KmsBackend for LocalKmsBackend {
self.client.disable_key(key_id, None).await self.client.disable_key(key_id, None).await
} }
async fn update_key_description(&self, key_id: &str, description: Option<&str>) -> Result<()> {
self.client
.update_key_metadata(key_id, |master_key| {
if master_key.description.as_deref() == description {
return Ok(false);
}
master_key.description = description.map(str::to_string);
Ok(true)
})
.await
}
async fn tag_key(&self, key_id: &str, tags: &HashMap<String, String>) -> Result<()> {
ensure_tag_keys_are_mutable(tags.keys().map(String::as_str))?;
self.client
.update_key_metadata(key_id, |master_key| {
let mut changed = false;
for (tag_key, value) in tags {
changed |= master_key.metadata.insert(tag_key.clone(), value.clone()).as_ref() != Some(value);
}
Ok(changed)
})
.await
}
async fn untag_key(&self, key_id: &str, tag_keys: &[String]) -> Result<()> {
ensure_tag_keys_are_mutable(tag_keys.iter().map(String::as_str))?;
self.client
.update_key_metadata(key_id, |master_key| {
let mut changed = false;
for tag_key in tag_keys {
changed |= master_key.metadata.remove(tag_key).is_some();
}
Ok(changed)
})
.await
}
async fn health_check(&self) -> Result<bool> { async fn health_check(&self) -> Result<bool> {
self.client.health_check().await.map(|_| true) self.client.health_check().await.map(|_| true)
} }
fn local_backup_client(&self) -> Option<&LocalKmsClient> {
Some(&self.client)
}
fn capabilities(&self) -> BackendCapabilities { fn capabilities(&self) -> BackendCapabilities {
// Rotation stays unadvertised until historical key versions can be // Rotation stays unadvertised until historical key versions can be
// retained (see LocalKmsClient::rotate_key); without version history // retained (see LocalKmsClient::rotate_key); without version history
@@ -1560,6 +1777,7 @@ impl KmsBackend for LocalKmsBackend {
.with_enable_disable(true) .with_enable_disable(true)
.with_schedule_deletion(true) .with_schedule_deletion(true)
.with_physical_delete(true) .with_physical_delete(true)
.with_update_key_metadata(true)
} }
async fn remove_expired_key(&self, key_id: &str, now: &Zoned) -> Result<ExpiredKeyRemoval> { async fn remove_expired_key(&self, key_id: &str, now: &Zoned) -> Result<ExpiredKeyRemoval> {
@@ -2404,6 +2622,34 @@ mod tests {
assert!(!is_orphan_commit_temp_name("mykey.tmp-")); assert!(!is_orphan_commit_temp_name("mykey.tmp-"));
} }
#[tokio::test]
async fn link_durably_publishes_no_clobber_and_is_content_idempotent() {
let temp_dir = TempDir::new().expect("temp dir");
let source = temp_dir.path().join("staged");
let dest = temp_dir.path().join("published");
fs::write(&source, b"staged-content").await.expect("write source");
durable_file::link_durably(source.clone(), dest.clone())
.await
.expect("first link must succeed");
assert_eq!(fs::read(&dest).await.expect("read dest"), b"staged-content");
// Re-entry with identical content is idempotent success — exactly the
// resumed-cutover case.
durable_file::link_durably(source.clone(), dest.clone())
.await
.expect("re-linking identical content must be idempotent");
// Existing content that differs is a hard failure, never a clobber.
let foreign = temp_dir.path().join("foreign");
fs::write(&foreign, b"different-content").await.expect("write foreign");
let error = durable_file::link_durably(foreign, dest.clone())
.await
.expect_err("differing content must not be clobbered");
assert_eq!(error.kind(), std::io::ErrorKind::AlreadyExists);
assert_eq!(fs::read(&dest).await.expect("dest unchanged"), b"staged-content");
}
#[tokio::test] #[tokio::test]
async fn durable_commit_fsyncs_every_write_path() { async fn durable_commit_fsyncs_every_write_path() {
use durable_file::fsync_recorder; use durable_file::fsync_recorder;
@@ -2442,6 +2688,7 @@ mod tests {
key_id: "durable-key".to_string(), key_id: "durable-key".to_string(),
pending_window_in_days: None, pending_window_in_days: None,
force_immediate: Some(true), force_immediate: Some(true),
confirm_key_id: None,
}) })
.await .await
.expect("delete key"); .expect("delete key");
@@ -2449,6 +2696,41 @@ mod tests {
assert!(fsync_recorder::dir_sync_count(dir) > dirs_before, "delete must fsync the key directory"); assert!(fsync_recorder::dir_sync_count(dir) > dirs_before, "delete must fsync the key directory");
} }
/// KmsManager::delete_key is the enforcement point for the waiting window;
/// this pins the backend's defensive copy of the same bound, which is all
/// that stands between a direct backend caller and a one-day window.
#[tokio::test]
async fn delete_key_refuses_a_window_outside_the_supported_range() {
let (client, _temp_dir) = create_test_client().await;
let key_id = "window-bounds-key";
client.create_key(key_id, "AES_256", None).await.expect("create key");
let backend = LocalKmsBackend { client };
for days in [MIN_PENDING_DELETION_WINDOW_DAYS - 1, MAX_PENDING_DELETION_WINDOW_DAYS + 1] {
let result = backend
.delete_key(DeleteKeyRequest {
key_id: key_id.to_string(),
pending_window_in_days: Some(days),
..Default::default()
})
.await;
assert!(
matches!(result, Err(KmsError::InvalidOperation { .. })),
"a {days}-day window must be refused, got {result:?}"
);
}
let state = backend
.describe_key(DescribeKeyRequest {
key_id: key_id.to_string(),
})
.await
.expect("describe should succeed")
.key_metadata
.key_state;
assert_eq!(state, KeyState::Enabled, "a refused window must not schedule the key");
}
#[tokio::test] #[tokio::test]
async fn interrupted_update_commit_recovers_to_complete_old_or_new_state() { async fn interrupted_update_commit_recovers_to_complete_old_or_new_state() {
use durable_file::{CommitStep, failpoint}; use durable_file::{CommitStep, failpoint};
+103
View File
@@ -19,7 +19,9 @@ use crate::types::*;
use async_trait::async_trait; use async_trait::async_trait;
use jiff::Zoned; use jiff::Zoned;
use serde::{Deserialize, Serialize}; use serde::{Deserialize, Serialize};
use std::collections::HashMap;
pub mod aws;
#[cfg(test)] #[cfg(test)]
mod contract_tests; mod contract_tests;
pub mod local; pub mod local;
@@ -98,6 +100,32 @@ pub(crate) fn ensure_key_status_permits(key_id: &str, status: &KeyStatus, operat
ensure_key_state_permits(key_id, &state, operation) ensure_key_state_permits(key_id, &state, operation)
} }
/// Tag key that carries the key's identity rather than user metadata.
///
/// The key-creation path lifts `name` out of the caller's tag map and uses it
/// as the key id, so a key's `name` tag and its id are the same string.
/// Rewriting or dropping it after creation would leave the key addressable
/// under an id its own metadata no longer states. Metadata updates therefore
/// reject it; only creation may set it.
pub const RESERVED_KEY_NAME_TAG: &str = "name";
/// Reject a metadata update that would rewrite or remove
/// [`RESERVED_KEY_NAME_TAG`].
///
/// Enforced by every backend that implements tag updates, so no call path —
/// including a direct [`KmsBackend`] user that bypasses the manager — can
/// detach a key from its identity.
pub(crate) fn ensure_tag_keys_are_mutable<'a>(tag_keys: impl IntoIterator<Item = &'a str>) -> Result<()> {
for tag_key in tag_keys {
if tag_key == RESERVED_KEY_NAME_TAG {
return Err(KmsError::invalid_parameter(format!(
"Tag '{RESERVED_KEY_NAME_TAG}' identifies the key and cannot be updated or removed"
)));
}
}
Ok(())
}
/// Simplified KMS backend interface for manager /// Simplified KMS backend interface for manager
#[async_trait] #[async_trait]
pub trait KmsBackend: Send + Sync { pub trait KmsBackend: Send + Sync {
@@ -153,6 +181,38 @@ pub trait KmsBackend: Send + Sync {
Err(KmsError::unsupported_capability("backend without rotation support", "rotate_key")) Err(KmsError::unsupported_capability("backend without rotation support", "rotate_key"))
} }
/// Replace a key's free-form description; `None` clears it.
///
/// Backends that advertise [`BackendCapabilities::update_key_metadata`]
/// must override this method; the default rejects the operation.
async fn update_key_description(&self, _key_id: &str, _description: Option<&str>) -> Result<()> {
Err(KmsError::unsupported_capability(
"backend without key metadata updates",
"update_key_description",
))
}
/// Add or overwrite the given tags, leaving every other tag untouched.
///
/// Implementations must run [`ensure_tag_keys_are_mutable`] before
/// persisting anything. Backends that advertise
/// [`BackendCapabilities::update_key_metadata`] must override this method;
/// the default rejects the operation.
async fn tag_key(&self, _key_id: &str, _tags: &HashMap<String, String>) -> Result<()> {
Err(KmsError::unsupported_capability("backend without key metadata updates", "tag_key"))
}
/// Remove the given tags.
///
/// Tags that are not set are ignored, so repeating the call is a no-op
/// rather than an error. Implementations must run
/// [`ensure_tag_keys_are_mutable`] before persisting anything. Backends
/// that advertise [`BackendCapabilities::update_key_metadata`] must
/// override this method; the default rejects the operation.
async fn untag_key(&self, _key_id: &str, _tag_keys: &[String]) -> Result<()> {
Err(KmsError::unsupported_capability("backend without key metadata updates", "untag_key"))
}
/// Health check /// Health check
async fn health_check(&self) -> Result<bool>; async fn health_check(&self) -> Result<bool>;
@@ -181,6 +241,19 @@ pub trait KmsBackend: Send + Sync {
async fn remove_expired_key(&self, _key_id: &str, _now: &Zoned) -> Result<ExpiredKeyRemoval> { async fn remove_expired_key(&self, _key_id: &str, _now: &Zoned) -> Result<ExpiredKeyRemoval> {
Err(KmsError::unsupported_capability("backend without deletion support", "remove_expired_key")) Err(KmsError::unsupported_capability("backend without deletion support", "remove_expired_key"))
} }
/// The running client to export a full-material backup bundle from.
///
/// Only the Local backend owns key material RustFS is allowed to export in
/// full (see [`crate::backup::BackupResponsibility`]); every other backend
/// keeps its cryptographic root outside RustFS and returns `None` here.
///
/// The export must run against the *running* client so that its fence
/// actually blocks concurrent create/delete work — a second client opened
/// on the same key directory would fence nothing.
fn local_backup_client(&self) -> Option<&local::LocalKmsClient> {
None
}
} }
/// Outcome of [`KmsBackend::remove_expired_key`]. /// Outcome of [`KmsBackend::remove_expired_key`].
@@ -222,6 +295,8 @@ pub struct BackendCapabilities {
pub versioning: bool, pub versioning: bool,
/// Irreversible physical deletion of key material /// Irreversible physical deletion of key material
pub physical_delete: bool, pub physical_delete: bool,
/// Updating a key's description and tags after creation
pub update_key_metadata: bool,
} }
impl BackendCapabilities { impl BackendCapabilities {
@@ -238,6 +313,7 @@ impl BackendCapabilities {
schedule_deletion: false, schedule_deletion: false,
versioning: false, versioning: false,
physical_delete: false, physical_delete: false,
update_key_metadata: false,
} }
} }
@@ -288,6 +364,12 @@ impl BackendCapabilities {
self.physical_delete = physical_delete; self.physical_delete = physical_delete;
self self
} }
/// Set whether description and tag updates are supported
pub const fn with_update_key_metadata(mut self, update_key_metadata: bool) -> Self {
self.update_key_metadata = update_key_metadata;
self
}
} }
impl Default for BackendCapabilities { impl Default for BackendCapabilities {
@@ -366,6 +448,7 @@ mod tests {
assert!(!capabilities.schedule_deletion); assert!(!capabilities.schedule_deletion);
assert!(!capabilities.versioning); assert!(!capabilities.versioning);
assert!(!capabilities.physical_delete); assert!(!capabilities.physical_delete);
assert!(!capabilities.update_key_metadata);
} }
#[tokio::test] #[tokio::test]
@@ -374,6 +457,12 @@ mod tests {
("enable_key", MinimalBackend.enable_key("any-key").await), ("enable_key", MinimalBackend.enable_key("any-key").await),
("disable_key", MinimalBackend.disable_key("any-key").await), ("disable_key", MinimalBackend.disable_key("any-key").await),
("rotate_key", MinimalBackend.rotate_key("any-key").await), ("rotate_key", MinimalBackend.rotate_key("any-key").await),
(
"update_key_description",
MinimalBackend.update_key_description("any-key", Some("new")).await,
),
("tag_key", MinimalBackend.tag_key("any-key", &HashMap::new()).await),
("untag_key", MinimalBackend.untag_key("any-key", &[]).await),
] { ] {
let error = result.expect_err("backends must opt in to lifecycle operations by overriding them"); let error = result.expect_err("backends must opt in to lifecycle operations by overriding them");
assert!( assert!(
@@ -395,6 +484,20 @@ mod tests {
); );
} }
#[test]
fn identity_tag_is_rejected_by_metadata_updates() {
let error = ensure_tag_keys_are_mutable([RESERVED_KEY_NAME_TAG])
.expect_err("the identity tag must not be writable through a metadata update");
assert!(
matches!(&error, KmsError::InvalidOperation { message } if message.contains(RESERVED_KEY_NAME_TAG)),
"expected a typed rejection naming the tag, got {error:?}"
);
// Ordinary tags — including ones that merely contain the reserved name
// — stay writable.
ensure_tag_keys_are_mutable(["team", "nickname", "Name"]).expect("ordinary tags must remain writable");
}
#[tokio::test] #[tokio::test]
async fn local_backend_capabilities_golden() { async fn local_backend_capabilities_golden() {
let temp_dir = tempfile::tempdir().expect("temp dir should be created"); let temp_dir = tempfile::tempdir().expect("temp dir should be created");
+45 -22
View File
@@ -26,16 +26,16 @@ use std::sync::{Arc, Mutex};
use tokio::io::{AsyncReadExt, AsyncWriteExt}; use tokio::io::{AsyncReadExt, AsyncWriteExt};
use tokio::net::{TcpListener, TcpStream}; use tokio::net::{TcpListener, TcpStream};
/// One canned HTTP response. /// One scripted connection outcome.
pub(crate) struct ScriptedResponse { pub(crate) enum ScriptedResponse {
status: u16, Http { status: u16, body: String },
body: String, Close,
} }
impl ScriptedResponse { impl ScriptedResponse {
/// A 200 response carrying `data` inside the standard Vault envelope. /// A 200 response carrying `data` inside the standard Vault envelope.
pub(crate) fn ok(data: serde_json::Value) -> Self { pub(crate) fn ok(data: serde_json::Value) -> Self {
Self { Self::Http {
status: 200, status: 200,
body: serde_json::json!({ body: serde_json::json!({
"request_id": "scripted", "request_id": "scripted",
@@ -50,18 +50,23 @@ impl ScriptedResponse {
/// An error response in Vault's `{"errors": [...]}` format. /// An error response in Vault's `{"errors": [...]}` format.
pub(crate) fn error(status: u16, message: &str) -> Self { pub(crate) fn error(status: u16, message: &str) -> Self {
Self { Self::Http {
status, status,
body: serde_json::json!({ "errors": [message] }).to_string(), body: serde_json::json!({ "errors": [message] }).to_string(),
} }
} }
/// Close the connection after consuming a request without sending an HTTP response.
pub(crate) fn close() -> Self {
Self::Close
}
} }
/// A scripted stand-in Vault listening on a loopback port. /// A scripted stand-in Vault listening on a loopback port.
pub(crate) struct ScriptedVault { pub(crate) struct ScriptedVault {
/// Base address (`http://127.0.0.1:port`) to point a Vault client at. /// Base address (`http://127.0.0.1:port`) to point a Vault client at.
pub(crate) address: String, pub(crate) address: String,
requests: Arc<Mutex<Vec<String>>>, requests: Arc<Mutex<Vec<(String, String)>>>,
} }
impl ScriptedVault { impl ScriptedVault {
@@ -81,25 +86,22 @@ impl ScriptedVault {
let Ok((mut stream, _)) = listener.accept().await else { let Ok((mut stream, _)) = listener.accept().await else {
return; return;
}; };
let Some(request_line) = read_request(&mut stream).await else { let Some(request) = read_request(&mut stream).await else {
continue; continue;
}; };
recorded recorded.lock().expect("scripted vault request log poisoned").push(request);
.lock()
.expect("scripted vault request log poisoned")
.push(request_line);
let response = responses let response = responses
.next() .next()
.unwrap_or_else(|| ScriptedResponse::error(599, "scripted vault: script exhausted")); .unwrap_or_else(|| ScriptedResponse::error(599, "scripted vault: script exhausted"));
if let ScriptedResponse::Http { status, body } = response {
let payload = format!( let payload = format!(
"HTTP/1.1 {} Scripted\r\ncontent-type: application/json\r\ncontent-length: {}\r\nconnection: close\r\n\r\n{}", "HTTP/1.1 {status} Scripted\r\ncontent-type: application/json\r\ncontent-length: {}\r\nconnection: close\r\n\r\n{body}",
response.status, body.len(),
response.body.len(),
response.body
); );
let _ = stream.write_all(payload.as_bytes()).await; let _ = stream.write_all(payload.as_bytes()).await;
let _ = stream.shutdown().await; let _ = stream.shutdown().await;
} }
}
}); });
Self { address, requests } Self { address, requests }
@@ -107,14 +109,32 @@ impl ScriptedVault {
/// The `METHOD /path` lines of every request served so far, in order. /// The `METHOD /path` lines of every request served so far, in order.
pub(crate) fn requests(&self) -> Vec<String> { pub(crate) fn requests(&self) -> Vec<String> {
self.requests.lock().expect("scripted vault request log poisoned").clone() self.requests
.lock()
.expect("scripted vault request log poisoned")
.iter()
.map(|(line, _)| line.clone())
.collect()
}
/// The request bodies, in the same order as [`Self::requests`]; empty for
/// bodyless requests. Lets tests assert what a write actually persisted
/// (record contents, check-and-set options), not just that a write happened.
pub(crate) fn request_bodies(&self) -> Vec<String> {
self.requests
.lock()
.expect("scripted vault request log poisoned")
.iter()
.map(|(_, body)| body.clone())
.collect()
} }
} }
/// Read one HTTP/1.1 request (head plus content-length body) and return its /// Read one HTTP/1.1 request (head plus content-length body) and return its
/// `METHOD /path` line. Draining the body before responding keeps the client /// `METHOD /path` line together with the body. Draining the body before
/// from seeing a connection reset while it is still writing. /// responding keeps the client from seeing a connection reset while it is
async fn read_request(stream: &mut TcpStream) -> Option<String> { /// still writing.
async fn read_request(stream: &mut TcpStream) -> Option<(String, String)> {
let mut buffer = Vec::new(); let mut buffer = Vec::new();
let mut chunk = [0u8; 4096]; let mut chunk = [0u8; 4096];
let head_end = loop { let head_end = loop {
@@ -146,14 +166,17 @@ async fn read_request(stream: &mut TcpStream) -> Option<String> {
}) })
.next() .next()
.unwrap_or(0); .unwrap_or(0);
let mut remaining = content_length.saturating_sub(buffer.len() - head_end); let mut body = buffer[head_end..].to_vec();
let mut remaining = content_length.saturating_sub(body.len());
while remaining > 0 { while remaining > 0 {
let read = stream.read(&mut chunk).await.ok()?; let read = stream.read(&mut chunk).await.ok()?;
if read == 0 { if read == 0 {
break; break;
} }
body.extend_from_slice(&chunk[..read]);
remaining = remaining.saturating_sub(read); remaining = remaining.saturating_sub(read);
} }
body.truncate(content_length);
Some(format!("{method} {path}")) Some((format!("{method} {path}"), String::from_utf8_lossy(&body).into_owned()))
} }
@@ -0,0 +1,15 @@
---
source: crates/kms/src/backends/aws.rs
expression: capabilities_snapshot(backend.capabilities())
---
{
"decrypt": true,
"enable_disable": true,
"encrypt": true,
"generate_data_key": true,
"physical_delete": false,
"rotate": true,
"schedule_deletion": true,
"update_key_metadata": false,
"versioning": true
}
@@ -10,5 +10,6 @@ expression: capabilities_snapshot(backend.capabilities())
"physical_delete": true, "physical_delete": true,
"rotate": false, "rotate": false,
"schedule_deletion": true, "schedule_deletion": true,
"update_key_metadata": true,
"versioning": false "versioning": false
} }
@@ -10,5 +10,6 @@ expression: capabilities_snapshot(backend.capabilities())
"physical_delete": false, "physical_delete": false,
"rotate": false, "rotate": false,
"schedule_deletion": false, "schedule_deletion": false,
"update_key_metadata": false,
"versioning": false "versioning": false
} }
@@ -10,5 +10,6 @@ expression: capabilities_snapshot(backend.capabilities())
"physical_delete": true, "physical_delete": true,
"rotate": true, "rotate": true,
"schedule_deletion": true, "schedule_deletion": true,
"update_key_metadata": true,
"versioning": true "versioning": true
} }
@@ -10,5 +10,6 @@ expression: capabilities_snapshot(backend.capabilities())
"physical_delete": true, "physical_delete": true,
"rotate": true, "rotate": true,
"schedule_deletion": true, "schedule_deletion": true,
"update_key_metadata": true,
"versioning": true "versioning": true
} }
+27
View File
@@ -431,6 +431,32 @@ mod tests {
(backend, key_id, key) (backend, key_id, key)
} }
#[tokio::test]
async fn metadata_updates_report_the_capability_gap() {
let (backend, key_id, _key) = create_test_backend().await;
assert!(
!backend.capabilities().update_key_metadata,
"Static KMS owns no mutable key record, so it must not advertise metadata updates"
);
for (operation, result) in [
("update_key_description", backend.update_key_description(&key_id, Some("new")).await),
(
"tag_key",
backend
.tag_key(&key_id, &HashMap::from([("team".to_string(), "storage".to_string())]))
.await,
),
("untag_key", backend.untag_key(&key_id, &["team".to_string()]).await),
] {
let error = result.expect_err("Static KMS is read-only: metadata updates must be rejected");
assert!(
matches!(error, KmsError::UnsupportedCapability { .. }),
"expected UnsupportedCapability for {operation}, got {error:?}"
);
}
}
#[tokio::test] #[tokio::test]
async fn test_generate_and_decrypt_data_key() { async fn test_generate_and_decrypt_data_key() {
let (backend, key_id, _key) = create_test_backend().await; let (backend, key_id, _key) = create_test_backend().await;
@@ -657,6 +683,7 @@ mod tests {
key_id: key_id.clone(), key_id: key_id.clone(),
pending_window_in_days: Some(7), pending_window_in_days: Some(7),
force_immediate: None, force_immediate: None,
confirm_key_id: None,
}, },
) )
.await; .await;
File diff suppressed because it is too large Load Diff
+208 -8
View File
@@ -57,6 +57,42 @@ const DEFAULT_REFRESH_RETRY_INTERVAL: Duration = Duration::from_secs(5);
/// Default seconds between token file re-reads for [`TokenFileSource`]. /// Default seconds between token file re-reads for [`TokenFileSource`].
const DEFAULT_TOKEN_FILE_POLL_INTERVAL_SECS: u64 = 30; const DEFAULT_TOKEN_FILE_POLL_INTERVAL_SECS: u64 = 30;
// ---------------------------------------------------------------------------
// Metrics
//
// Both gauges describe the one credential generation currently installed, so
// they carry no labels: the Vault address, mount, auth path and token are all
// off limits as label values, and there is exactly one generation to describe.
// The renewal loop republishes them on a bounded cadence while it waits, so a
// scrape landing between refresh cycles never reads a TTL frozen at the last
// refresh, or a fail-closed state that flipped after it.
// ---------------------------------------------------------------------------
/// Gauge: seconds left before the active Vault token expires; `0` once it has.
const METRIC_TOKEN_TTL_SECONDS: &str = "rustfs_kms_vault_token_ttl_seconds";
/// Gauge: `1` while [`VaultCredentialProvider::current`] refuses to hand out
/// the token because it is inside the fail-closed safety window, `0` otherwise.
const METRIC_CREDENTIALS_FAIL_CLOSED: &str = "rustfs_kms_vault_credentials_fail_closed";
/// How often the renewal loop republishes the credential gauges while waiting.
/// Bounds how stale a scrape can be, without any additional Vault traffic.
const CREDENTIAL_GAUGE_INTERVAL: Duration = Duration::from_secs(10);
/// Register metric descriptions once per process.
fn describe_credential_metrics() {
static DESCRIBE: std::sync::Once = std::sync::Once::new();
DESCRIBE.call_once(|| {
metrics::describe_gauge!(
METRIC_TOKEN_TTL_SECONDS,
"Seconds remaining before the Vault token backing the KMS backend expires"
);
metrics::describe_gauge!(
METRIC_CREDENTIALS_FAIL_CLOSED,
"1 while the Vault credential provider refuses to serve its token because the token is inside the fail-closed safety window"
);
});
}
/// A crate-owned secret value, zeroized on drop and redacted in Debug output. /// A crate-owned secret value, zeroized on drop and redacted in Debug output.
#[derive(Clone, Zeroize, ZeroizeOnDrop)] #[derive(Clone, Zeroize, ZeroizeOnDrop)]
pub(crate) struct SecretString(String); pub(crate) struct SecretString(String);
@@ -627,6 +663,45 @@ impl VaultCredentialProvider {
Ok(handle) Ok(handle)
} }
/// Publish the credential gauges for the generation currently installed.
///
/// The fail-closed gauge re-evaluates the very gate
/// [`VaultCredentialProvider::current`] applies, so what operators see and
/// what the request path does cannot drift apart.
fn record_credential_gauges(&self) {
let handle = self.current.load();
let now = Instant::now();
let fail_closed = match handle.expires_at() {
Some(expires_at) => {
metrics::gauge!(METRIC_TOKEN_TTL_SECONDS).set(expires_at.saturating_duration_since(now).as_secs_f64());
now + self.policy.safety_window >= expires_at
}
// A generation without an expiry has no remaining TTL to report
// and can never lapse, so it can never fail closed either.
None => false,
};
metrics::gauge!(METRIC_CREDENTIALS_FAIL_CLOSED).set(if fail_closed { 1.0 } else { 0.0 });
}
/// Wait until `deadline`, republishing the credential gauges on the
/// observation cadence. Reports `false` when cancellation cut the wait
/// short.
async fn wait_publishing_gauges(&self, deadline: Instant, cancel: &CancellationToken) -> bool {
loop {
self.record_credential_gauges();
let now = Instant::now();
if now >= deadline {
return true;
}
let slice = (deadline - now).min(CREDENTIAL_GAUGE_INTERVAL);
tokio::select! {
biased;
_ = cancel.cancelled() => return false,
_ = tokio::time::sleep(slice) => {}
}
}
}
/// Refresh the credentials if generation `observed` is still current. /// Refresh the credentials if generation `observed` is still current.
/// ///
/// Single-flight: concurrent callers serialize on the refresh lock, and a /// Single-flight: concurrent callers serialize on the refresh lock, and a
@@ -711,17 +786,22 @@ impl fmt::Debug for VaultCredentialProvider {
/// again immediately (the renewal point is already in the past), so the /// again immediately (the renewal point is already in the past), so the
/// provider keeps trying to recover even after the fail-closed window has /// provider keeps trying to recover even after the fail-closed window has
/// been reached. /// been reached.
///
/// Both waits run through [`VaultCredentialProvider::wait_publishing_gauges`],
/// which is the only place the credential gauges are published: the loop is
/// already the component that tracks token expiry, and doing it here keeps the
/// request path free of any metric work.
async fn renewal_loop(provider: Arc<VaultCredentialProvider>, cancel: CancellationToken) { async fn renewal_loop(provider: Arc<VaultCredentialProvider>, cancel: CancellationToken) {
describe_credential_metrics();
loop { loop {
let handle = provider.snapshot(); let handle = provider.snapshot();
let Some(renew_at) = handle.renew_at() else { let Some(renew_at) = handle.renew_at() else {
// The current generation never expires; nothing left to schedule. // The current generation never expires; nothing left to schedule.
provider.record_credential_gauges();
return; return;
}; };
tokio::select! { if !provider.wait_publishing_gauges(renew_at, &cancel).await {
biased; return;
_ = cancel.cancelled() => return,
_ = tokio::time::sleep_until(renew_at) => {}
} }
match provider.refresh(handle.generation, &cancel).await { match provider.refresh(handle.generation, &cancel).await {
@@ -733,10 +813,9 @@ async fn renewal_loop(provider: Arc<VaultCredentialProvider>, cancel: Cancellati
error = %error, error = %error,
"Vault credential refresh failed; retrying until the credentials recover" "Vault credential refresh failed; retrying until the credentials recover"
); );
tokio::select! { let retry_at = Instant::now() + provider.policy.retry_interval;
biased; if !provider.wait_publishing_gauges(retry_at, &cancel).await {
_ = cancel.cancelled() => return, return;
_ = tokio::time::sleep(provider.policy.retry_interval) => {}
} }
} }
} }
@@ -1354,4 +1433,125 @@ mod tests {
assert!(!rendered.contains(TEST_TOKEN), "debug output must not leak the vault token: {rendered}"); assert!(!rendered.contains(TEST_TOKEN), "debug output must not leak the vault token: {rendered}");
assert!(rendered.contains("vault-token"), "the file path is not a secret"); assert!(rendered.contains("vault-token"), "the file path is not a secret");
} }
// -- Metric emission ----------------------------------------------------
//
// Each test installs a thread-local debugging recorder and drives a
// paused-clock current-thread runtime inside it, so the renewal loop's
// gauge publications land on exact virtual timestamps.
use metrics_util::MetricKind;
use metrics_util::debugging::{DebugValue, DebuggingRecorder};
type MetricEntry = (
metrics_util::CompositeKey,
Option<metrics::Unit>,
Option<metrics::SharedString>,
DebugValue,
);
/// Run `test` on a paused current-thread runtime under a debugging
/// recorder and return one snapshot of everything it emitted.
///
/// A single snapshot per test on purpose: `Snapshotter::snapshot` drains
/// the recorded state, so taking it per assertion would only show the
/// first assertion any data.
fn record_metrics<Out>(
test: impl FnOnce() -> std::pin::Pin<Box<dyn std::future::Future<Output = Out>>>,
) -> (Vec<MetricEntry>, Out) {
let recorder = DebuggingRecorder::new();
let snapshotter = recorder.snapshotter();
let out = metrics::with_local_recorder(&recorder, || {
let runtime = tokio::runtime::Builder::new_current_thread()
.enable_time()
.start_paused(true)
.build()
.expect("current-thread runtime must build");
runtime.block_on(test())
});
(snapshotter.snapshot().into_vec(), out)
}
/// Last value of a gauge, plus the labels it was published with.
fn gauge(snapshot: &[MetricEntry], name: &str) -> Option<(f64, Vec<String>)> {
snapshot.iter().find_map(|(composite, _unit, _description, value)| {
let matches = composite.kind() == MetricKind::Gauge && composite.key().name() == name;
match (matches, value) {
(true, DebugValue::Gauge(value)) => Some((
value.into_inner(),
composite.key().labels().map(|label| label.key().to_string()).collect(),
)),
_ => None,
}
})
}
#[test]
fn renewal_loop_republishes_token_ttl_while_it_waits() {
let (snapshot, ()) = record_metrics(|| {
Box::pin(async {
let (provider, _state) = scripted_provider(
Duration::from_secs(60),
true,
test_policy(Duration::from_secs(10), Duration::from_secs(5)),
)
.await;
let task = provider.spawn_renewal_task().expect("lease-bound tokens need renewal");
// Well inside the first half of the lease: nothing has been
// renewed yet, so only the observation cadence can have moved
// the gauge off its initial 60s.
tokio::time::sleep(Duration::from_secs(25)).await;
task.shutdown().await;
})
});
let (ttl, labels) = gauge(&snapshot, METRIC_TOKEN_TTL_SECONDS).expect("token TTL gauge must be published");
assert!(
(ttl - 40.0).abs() < 1.0,
"expected ~40s left of a 60s lease at the last observation, got {ttl}"
);
assert!(labels.is_empty(), "credential gauges must carry no labels, got {labels:?}");
assert_eq!(
gauge(&snapshot, METRIC_CREDENTIALS_FAIL_CLOSED).map(|(value, _)| value),
Some(0.0),
"a token outside its safety window must not report fail-closed"
);
}
#[test]
fn fail_closed_gauge_tracks_the_gate_the_request_path_applies() {
let (snapshot, refused) = record_metrics(|| {
Box::pin(async {
let (provider, state) = scripted_provider(
Duration::from_secs(60),
true,
test_policy(Duration::from_secs(10), Duration::from_secs(5)),
)
.await;
state.fail_renew.store(true, Ordering::SeqCst);
state.fail_login.store(true, Ordering::SeqCst);
let task = provider.spawn_renewal_task().expect("renewal task");
// 60s lease minus the 10s safety window: by 51s the provider
// is refusing the token, and the gauge must already say so.
tokio::time::sleep(Duration::from_secs(51)).await;
let refused = provider.current().is_err();
task.shutdown().await;
refused
})
});
assert!(refused, "a token inside the safety window must be refused");
assert_eq!(
gauge(&snapshot, METRIC_CREDENTIALS_FAIL_CLOSED).map(|(value, _)| value),
Some(1.0),
"the gauge must report the same fail-closed state the request path enforces"
);
let (ttl, _labels) = gauge(&snapshot, METRIC_TOKEN_TTL_SECONDS).expect("token TTL gauge must be published");
assert!(
(ttl - 10.0).abs() < 1.0,
"expected the TTL gauge to track the lease down into its safety window, got {ttl}"
);
}
} }
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
+448 -25
View File
@@ -306,6 +306,141 @@ impl LocalKdfDescriptor {
} }
} }
/// Transit engine reference recorded in a Vault-backed bundle.
///
/// Transit keys are non-exportable, so a bundle can only ever name the key and
/// the version window its content depends on. Restore compares these values
/// against the Transit key the operator's native Vault restore produced.
#[derive(Debug, Clone, PartialEq, Eq, Serialize, Deserialize)]
#[serde(deny_unknown_fields)]
pub struct VaultTransitReference {
/// Transit engine mount path.
pub mount_path: String,
/// Transit key name.
pub key_name: String,
/// Lowest Transit key version the bundle's content still needs. A target
/// whose `min_decryption_version` sits above this can no longer decrypt
/// the oldest state in the bundle.
pub required_min_version: u32,
/// Latest Transit key version at snapshot time. A target whose newest
/// version is below this was restored to a point before the bundle.
pub current_version: u32,
/// Operator-recorded immutable reference to the native Vault/HSM snapshot
/// protecting this key. RustFS neither produces nor consumes that
/// snapshot; the reference exists so restore evidence can name it.
#[serde(default, skip_serializing_if = "Option::is_none")]
pub native_snapshot_reference: Option<String>,
}
/// KV generation of one key's record at snapshot time.
#[derive(Debug, Clone, PartialEq, Eq, Serialize, Deserialize)]
#[serde(deny_unknown_fields)]
pub struct VaultKvRecordReference {
/// Stable key id the record belongs to.
pub key_id: String,
/// KV2 secret version holding the record when the snapshot was taken.
/// Restore compares the target's current version against it: a lower
/// value means the target's Vault was restored to a point before the
/// bundle, a higher value means the target moved ahead of it.
pub kv_version: u64,
}
/// Immutable references to the external Vault state a bundle depends on.
///
/// A Vault-backed bundle never carries the cryptographic root: Transit keys
/// cannot be exported and KV2 records live in Vault's own storage. What the
/// bundle can own is a precise description of *which* external state it was
/// captured against, so a restore refuses to proceed when the operator's
/// native Vault restore landed somewhere else. Credentials are structurally
/// absent — no token, AppRole secret id, or TLS material is representable in
/// this type.
#[derive(Debug, Clone, PartialEq, Eq, Serialize, Deserialize)]
#[serde(deny_unknown_fields)]
pub struct VaultExternalReferences {
/// Opaque identity of the Vault cluster the snapshot was taken from.
pub cluster_id: String,
/// Vault Enterprise namespace, when the deployment uses one.
#[serde(default, skip_serializing_if = "Option::is_none")]
pub namespace: Option<String>,
/// KV v2 mount holding the RustFS records.
pub kv_mount: String,
/// Path prefix under `kv_mount` holding the RustFS records.
pub kv_path_prefix: String,
/// Transit engine reference; present exactly when the bundle's material
/// is protected by a Transit key.
#[serde(default, skip_serializing_if = "Option::is_none")]
pub transit: Option<VaultTransitReference>,
/// Per-key KV generation observed at snapshot time, one entry per key the
/// bundle carries a record for.
pub kv_records: Vec<VaultKvRecordReference>,
}
impl VaultExternalReferences {
fn validate(&self, backend: BackupBackendKind, protection: AtRestProtection) -> Result<(), BackupError> {
require_non_empty("external_references.cluster_id", &self.cluster_id)?;
require_non_empty("external_references.kv_mount", &self.kv_mount)?;
require_non_empty("external_references.kv_path_prefix", &self.kv_path_prefix)?;
if self.namespace.as_deref() == Some("") {
return Err(BackupError::corrupted("external Vault namespace must not be empty when present"));
}
// The Transit reference is required exactly where the cryptographic
// root lives in Transit; a storage-only KV2 bundle that claims one
// would misdescribe its own trust root.
let transit_required = matches!(
(backend, protection),
(BackupBackendKind::VaultTransit, _) | (BackupBackendKind::VaultKv2, AtRestProtection::TransitWrapped)
);
match (&self.transit, transit_required) {
(None, true) => {
return Err(BackupError::corrupted(format!(
"({backend:?}, {protection:?}) bundles must reference the Transit key protecting them"
)));
}
(Some(_), false) => {
return Err(BackupError::corrupted(format!(
"({backend:?}, {protection:?}) bundles have no Transit trust root to reference"
)));
}
(Some(transit), true) => {
require_non_empty("external_references.transit.mount_path", &transit.mount_path)?;
require_non_empty("external_references.transit.key_name", &transit.key_name)?;
if transit.required_min_version == 0 {
return Err(BackupError::corrupted("Transit reference must require at least key version 1"));
}
if transit.current_version < transit.required_min_version {
return Err(BackupError::corrupted(format!(
"Transit reference declares current version {} below the required minimum {}",
transit.current_version, transit.required_min_version
)));
}
if transit.native_snapshot_reference.as_deref() == Some("") {
return Err(BackupError::corrupted("Transit native snapshot reference must not be empty when present"));
}
}
(None, false) => {}
}
let mut seen = BTreeSet::new();
for record in &self.kv_records {
require_non_empty("external_references.kv_records[].key_id", &record.key_id)?;
if record.kv_version == 0 {
return Err(BackupError::corrupted(format!(
"KV generation for key '{}' must be at least 1",
record.key_id
)));
}
if !seen.insert(record.key_id.as_str()) {
return Err(BackupError::corrupted(format!(
"KV reference for key '{}' appears more than once",
record.key_id
)));
}
}
Ok(())
}
}
/// Completeness marker of a bundle. /// Completeness marker of a bundle.
/// ///
/// A producer writes `in-progress` state (or no manifest at all) until the /// A producer writes `in-progress` state (or no manifest at all) until the
@@ -355,8 +490,10 @@ struct ManifestProbe {
/// ///
/// The manifest is the authoritative description of one backup bundle: what /// The manifest is the authoritative description of one backup bundle: what
/// was captured, under which snapshot generation, protected by which backup /// was captured, under which snapshot generation, protected by which backup
/// KEK, and which restore responsibility applies. Field order is part of the /// KEK, and which restore responsibility applies. The digest's canonical
/// canonical digest form and is frozen for this format version. /// form is the manifest's JSON value with the digest hex emptied (see
/// [`Self::compute_digest`]); decoders verify it against the raw stored
/// bytes and never re-serialize parsed fields.
#[derive(Debug, Clone, PartialEq, Eq, Serialize, Deserialize)] #[derive(Debug, Clone, PartialEq, Eq, Serialize, Deserialize)]
#[serde(deny_unknown_fields)] #[serde(deny_unknown_fields)]
pub struct BackupManifest { pub struct BackupManifest {
@@ -392,6 +529,10 @@ pub struct BackupManifest {
/// Local backend section; required exactly when `backend` is `Local`. /// Local backend section; required exactly when `backend` is `Local`.
#[serde(default, skip_serializing_if = "Option::is_none")] #[serde(default, skip_serializing_if = "Option::is_none")]
pub local_kdf: Option<LocalKdfDescriptor>, pub local_kdf: Option<LocalKdfDescriptor>,
/// External Vault references; required exactly when `backend` is a Vault
/// backend. Never carries credentials — see [`VaultExternalReferences`].
#[serde(default, skip_serializing_if = "Option::is_none")]
pub external_references: Option<VaultExternalReferences>,
/// Reserved for the per-key version/envelope inventory defined by the /// Reserved for the per-key version/envelope inventory defined by the
/// backlog#1565 contract. May not carry data in format version 1. /// backlog#1565 contract. May not carry data in format version 1.
#[serde(default, skip_serializing_if = "Option::is_none")] #[serde(default, skip_serializing_if = "Option::is_none")]
@@ -415,8 +556,14 @@ impl BackupManifest {
/// ///
/// Fail-closed order: truncated or malformed input, then unknown format /// Fail-closed order: truncated or malformed input, then unknown format
/// version, then a missing completeness marker, then schema decoding /// version, then a missing completeness marker, then schema decoding
/// (unknown fields, duplicate fields, missing fields), then semantic /// (unknown fields, duplicate fields, missing fields), then digest
/// validation including digest verification. /// verification against the raw input bytes, then semantic validation.
///
/// The digest is verified against the bytes as stored — parsed typed
/// fields are never re-serialized for verification, so a field whose
/// string form does not round-trip byte-identically through its parsed
/// representation (timestamps in environment-dependent time zone
/// spellings, for example) cannot produce a spurious mismatch.
pub fn decode(bytes: &[u8]) -> Result<Self, BackupError> { pub fn decode(bytes: &[u8]) -> Result<Self, BackupError> {
let probe: ManifestProbe = serde_json::from_slice(bytes).map_err(map_serde_error)?; let probe: ManifestProbe = serde_json::from_slice(bytes).map_err(map_serde_error)?;
if probe.format_version != Self::FORMAT_VERSION { if probe.format_version != Self::FORMAT_VERSION {
@@ -429,7 +576,9 @@ impl BackupManifest {
return Err(BackupError::incomplete_bundle("manifest has no completeness marker")); return Err(BackupError::incomplete_bundle("manifest has no completeness marker"));
} }
let manifest: Self = serde_json::from_slice(bytes).map_err(map_serde_error)?; let manifest: Self = serde_json::from_slice(bytes).map_err(map_serde_error)?;
manifest.validate()?; manifest.validate_pre_digest()?;
Self::verify_digest_in_bytes(bytes, &manifest.manifest_digest)?;
manifest.validate_content()?;
Ok(manifest) Ok(manifest)
} }
@@ -449,23 +598,30 @@ impl BackupManifest {
Ok(self) Ok(self)
} }
/// Compute the digest over the canonical manifest bytes. /// Compute the digest over the canonical manifest form.
/// ///
/// Canonical form: compact JSON serialization of this manifest with the /// Canonical form: the JSON *value* of the manifest with the digest hex
/// digest hex emptied. Field order is struct declaration order and is /// emptied, object keys rebuilt in bytewise-sorted order at every level
/// frozen for format version 1, so the same manifest content always /// (see [`canonicalize_value`]), then serialized compactly. The value
/// hashes to the same value. /// layer is what makes sealing and decoding agree byte-for-byte — a
/// decoder recovers the identical value from the raw stored bytes
/// without round-tripping any typed field through parse-and-reprint —
/// and the explicit key sort makes the bytes independent of
/// `serde_json`'s map implementation (`preserve_order` on or off).
pub fn compute_digest(&self) -> Result<ContentDigest, BackupError> { pub fn compute_digest(&self) -> Result<ContentDigest, BackupError> {
let mut unsealed = self.clone(); let mut unsealed = self.clone();
unsealed.manifest_digest = ContentDigest::placeholder(self.manifest_digest.algorithm); unsealed.manifest_digest = ContentDigest::placeholder(self.manifest_digest.algorithm);
let canonical = serde_json::to_vec(&unsealed) let value = serde_json::to_value(&unsealed)
.map_err(|error| BackupError::corrupted(format!("manifest canonicalization failed: {error}")))?; .map_err(|error| BackupError::corrupted(format!("manifest canonicalization failed: {error}")))?;
match self.manifest_digest.algorithm { Self::digest_of_canonical_value(value, self.manifest_digest.algorithm)
DigestAlgorithm::Sha256 => Ok(ContentDigest::sha256_of(&canonical)),
}
} }
/// Verify the sealed digest against the current manifest content. /// Verify the sealed digest against the current in-memory content.
///
/// This is the producer-side check (sealing and [`Self::encode`]).
/// Decoders must use the raw stored bytes instead (see [`Self::decode`]):
/// re-serializing parsed fields is not guaranteed to reproduce the
/// stored spelling byte-for-byte.
pub fn verify_digest(&self) -> Result<(), BackupError> { pub fn verify_digest(&self) -> Result<(), BackupError> {
if !self.manifest_digest.is_well_formed() { if !self.manifest_digest.is_well_formed() {
return Err(BackupError::corrupted("manifest digest is not a well-formed digest value")); return Err(BackupError::corrupted("manifest digest is not a well-formed digest value"));
@@ -478,6 +634,33 @@ impl BackupManifest {
Ok(()) Ok(())
} }
/// Verify a declared digest against raw manifest bytes, normalizing only
/// through the JSON value layer and emptying the digest slot in place.
fn verify_digest_in_bytes(bytes: &[u8], declared: &ContentDigest) -> Result<(), BackupError> {
if !declared.is_well_formed() {
return Err(BackupError::corrupted("manifest digest is not a well-formed digest value"));
}
let mut value: serde_json::Value = serde_json::from_slice(bytes).map_err(map_serde_error)?;
let Some(slot) = value.get_mut("manifest_digest").and_then(|digest| digest.get_mut("hex")) else {
return Err(BackupError::corrupted("manifest has no digest slot"));
};
*slot = serde_json::Value::String(String::new());
if Self::digest_of_canonical_value(value, declared.algorithm)? != *declared {
return Err(BackupError::corrupted(
"manifest digest mismatch: content does not match the sealed digest",
));
}
Ok(())
}
fn digest_of_canonical_value(value: serde_json::Value, algorithm: DigestAlgorithm) -> Result<ContentDigest, BackupError> {
let canonical = serde_json::to_vec(&canonicalize_value(value))
.map_err(|error| BackupError::corrupted(format!("manifest canonicalization failed: {error}")))?;
match algorithm {
DigestAlgorithm::Sha256 => Ok(ContentDigest::sha256_of(&canonical)),
}
}
/// Look up a required artifact by kind, failing closed when absent. /// Look up a required artifact by kind, failing closed when absent.
pub fn require_artifact(&self, kind: ArtifactKind) -> Result<&ArtifactDescriptor, BackupError> { pub fn require_artifact(&self, kind: ArtifactKind) -> Result<&ArtifactDescriptor, BackupError> {
self.artifacts self.artifacts
@@ -486,11 +669,21 @@ impl BackupManifest {
.ok_or_else(|| BackupError::missing_artifact(artifact_kind_name(kind))) .ok_or_else(|| BackupError::missing_artifact(artifact_kind_name(kind)))
} }
/// Validate the full manifest contract. /// Validate the full manifest contract against the in-memory content.
/// ///
/// This is decode-side validation and also guards [`Self::encode`], so a /// This guards [`Self::encode`], so a producer cannot publish a manifest
/// producer cannot publish a manifest a decoder would reject. /// a decoder would reject. [`Self::decode`] runs the same checks but
/// verifies the digest against the raw input bytes instead.
pub fn validate(&self) -> Result<(), BackupError> { pub fn validate(&self) -> Result<(), BackupError> {
self.validate_pre_digest()?;
self.verify_digest()?;
self.validate_content()
}
/// Checks that must run before any digest verification: an unknown
/// version or an unsealed bundle is reported as its own typed error, not
/// as a digest mismatch.
fn validate_pre_digest(&self) -> Result<(), BackupError> {
if self.format_version != Self::FORMAT_VERSION { if self.format_version != Self::FORMAT_VERSION {
return Err(BackupError::UnknownVersion { return Err(BackupError::UnknownVersion {
found: self.format_version, found: self.format_version,
@@ -500,7 +693,12 @@ impl BackupManifest {
if self.completeness != CompletenessState::Complete { if self.completeness != CompletenessState::Complete {
return Err(BackupError::incomplete_bundle("completeness marker records an in-progress bundle")); return Err(BackupError::incomplete_bundle("completeness marker records an in-progress bundle"));
} }
self.verify_digest()?; Ok(())
}
/// Semantic validation of everything except version, completeness, and
/// digest integrity.
fn validate_content(&self) -> Result<(), BackupError> {
require_non_empty("backup_id", &self.backup_id)?; require_non_empty("backup_id", &self.backup_id)?;
require_non_empty("rustfs_version", &self.rustfs_version)?; require_non_empty("rustfs_version", &self.rustfs_version)?;
require_non_empty("deployment_identity", &self.deployment_identity)?; require_non_empty("deployment_identity", &self.deployment_identity)?;
@@ -513,6 +711,7 @@ impl BackupManifest {
} }
self.validate_responsibility()?; self.validate_responsibility()?;
self.validate_local_kdf()?; self.validate_local_kdf()?;
self.validate_external_references()?;
self.validate_artifacts()?; self.validate_artifacts()?;
Ok(()) Ok(())
} }
@@ -542,6 +741,19 @@ impl BackupManifest {
} }
} }
fn validate_external_references(&self) -> Result<(), BackupError> {
let vault_backend = matches!(self.backend, BackupBackendKind::VaultKv2 | BackupBackendKind::VaultTransit);
match (&self.external_references, vault_backend) {
(None, true) => Err(BackupError::corrupted("Vault backend manifest must carry external Vault references")),
(Some(_), false) => Err(BackupError::corrupted(format!(
"external Vault references are not valid for backend {:?}",
self.backend
))),
(Some(references), true) => references.validate(self.backend, self.at_rest_protection),
(None, false) => Ok(()),
}
}
fn validate_artifacts(&self) -> Result<(), BackupError> { fn validate_artifacts(&self) -> Result<(), BackupError> {
let mut paths = BTreeSet::new(); let mut paths = BTreeSet::new();
for artifact in &self.artifacts { for artifact in &self.artifacts {
@@ -645,6 +857,29 @@ fn map_serde_error(error: serde_json::Error) -> BackupError {
} }
} }
/// Rebuild a JSON value with object keys in bytewise-sorted order at every
/// nesting level (array element order is preserved).
///
/// `serde_json`'s map keeps keys sorted by default but preserves insertion
/// order when the `preserve_order` feature is unified into the build by any
/// other crate. Digest bytes must not depend on that, so the ordering is
/// imposed explicitly here instead of being inherited from the map type.
fn canonicalize_value(value: serde_json::Value) -> serde_json::Value {
match value {
serde_json::Value::Object(map) => {
let mut entries: Vec<(String, serde_json::Value)> = map.into_iter().collect();
entries.sort_by(|a, b| a.0.cmp(&b.0));
let mut sorted = serde_json::Map::with_capacity(entries.len());
for (key, entry) in entries {
sorted.insert(key, canonicalize_value(entry));
}
serde_json::Value::Object(sorted)
}
serde_json::Value::Array(items) => serde_json::Value::Array(items.into_iter().map(canonicalize_value).collect()),
other => other,
}
}
#[cfg(test)] #[cfg(test)]
mod tests { mod tests {
use super::*; use super::*;
@@ -736,7 +971,7 @@ mod tests {
/// canonical form and frozen. If serialization layout or field order /// canonical form and frozen. If serialization layout or field order
/// changes, this value changes and the fixture test fails — which is the /// changes, this value changes and the fixture test fails — which is the
/// point: that is a format-version bump, not a patch. /// point: that is a format-version bump, not a patch.
const FIXTURE_DIGEST_HEX: &str = "01accb3e2bc51e12d17d1efc52ad6ab9441c50bec52c3fb29e0a9f92b725cdaa"; const FIXTURE_DIGEST_HEX: &str = "a8b104d61c358cd75cc9a7691d2f7d2d96f73c53653bef8940a49e96dbdf7775";
fn fixture() -> String { fn fixture() -> String {
FIXTURE.replace("SEALED_DIGEST_HEX", FIXTURE_DIGEST_HEX) FIXTURE.replace("SEALED_DIGEST_HEX", FIXTURE_DIGEST_HEX)
@@ -783,6 +1018,7 @@ mod tests {
vec![AtRestProtection::EncryptedMasterKey], vec![AtRestProtection::EncryptedMasterKey],
Some("verifier-opaque-1".to_string()), Some("verifier-opaque-1".to_string()),
)), )),
external_references: None,
key_versions: None, key_versions: None,
capability_discovery: None, capability_discovery: None,
completeness: CompletenessState::InProgress, completeness: CompletenessState::InProgress,
@@ -1010,12 +1246,39 @@ mod tests {
let with_discovery = fixture().replace("\"completeness\"", "\"capability_discovery\": [1], \"completeness\""); let with_discovery = fixture().replace("\"completeness\"", "\"capability_discovery\": [1], \"completeness\"");
expect_corrupted(BackupManifest::decode(with_discovery.as_bytes()), "reserved"); expect_corrupted(BackupManifest::decode(with_discovery.as_bytes()), "reserved");
// Explicit null carries no data and is tolerated as absence; digest // An explicit null slot decodes as absence at the schema layer, but
// verification still passes because null slots are skipped on // sealed bundles never contain the key (`skip_serializing_if`), so
// serialization. // inserting one after sealing is a byte-level modification and the
// raw-bytes digest check rejects it.
let with_null = fixture().replace("\"completeness\"", "\"key_versions\": null, \"completeness\""); let with_null = fixture().replace("\"completeness\"", "\"key_versions\": null, \"completeness\"");
let decoded = BackupManifest::decode(with_null.as_bytes()).expect("null reserved slot should decode"); expect_corrupted(BackupManifest::decode(with_null.as_bytes()), "digest mismatch");
assert_eq!(decoded.key_versions, None); }
#[test]
fn digest_verification_survives_non_round_tripping_timestamp_spellings() {
// A legacy `created_at` spelling (no time zone annotation) parses via
// the compat fallback and re-serializes differently ("+00:00[UTC]"),
// and host-dependent zone spellings can do the same. Digest
// verification therefore operates on the raw stored bytes and must
// never re-serialize parsed fields.
let sealed = seal(local_manifest_unsealed());
let mut value = serde_json::to_value(&sealed).expect("manifest should convert to a value");
value["created_at"] = serde_json::Value::String("2026-07-30T00:00:00+00:00".to_string());
value["manifest_digest"]["hex"] = serde_json::Value::String(String::new());
let digest = BackupManifest::digest_of_canonical_value(value.clone(), DigestAlgorithm::Sha256)
.expect("canonical digest should compute");
value["manifest_digest"]["hex"] = serde_json::Value::String(digest.hex);
let bytes = serde_json::to_vec(&value).expect("manifest bytes");
let decoded = BackupManifest::decode(&bytes).expect("a non-round-tripping timestamp spelling must not break decoding");
// Precondition: the spelling really does not survive a typed
// round-trip — otherwise this test is vacuous.
let reserialized = serde_json::to_value(&decoded).expect("decoded manifest should convert to a value");
assert_ne!(reserialized["created_at"], value["created_at"]);
// Which is exactly why the producer-side (in-memory) digest check
// cannot be used on decoded manifests.
assert!(decoded.verify_digest().is_err());
} }
#[test] #[test]
@@ -1162,4 +1425,164 @@ mod tests {
assert_eq!(decoded, sealed); assert_eq!(decoded, sealed);
assert_eq!(decoded.responsibility, BackupResponsibility::ReferenceOnly); assert_eq!(decoded.responsibility, BackupResponsibility::ReferenceOnly);
} }
fn transit_references() -> VaultExternalReferences {
VaultExternalReferences {
cluster_id: "vault-cluster-a".to_string(),
namespace: Some("tenant-a".to_string()),
kv_mount: "secret".to_string(),
kv_path_prefix: "rustfs/kms/transit-metadata".to_string(),
transit: Some(VaultTransitReference {
mount_path: "transit".to_string(),
key_name: "rustfs-master".to_string(),
required_min_version: 2,
current_version: 5,
native_snapshot_reference: Some("s3://dr/vault/2026-08-01.snap".to_string()),
}),
kv_records: vec![VaultKvRecordReference {
key_id: "object-key".to_string(),
kv_version: 4,
}],
}
}
fn transit_manifest_unsealed() -> BackupManifest {
BackupManifest {
backend: BackupBackendKind::VaultTransit,
at_rest_protection: AtRestProtection::ExternalNonExportable,
responsibility: BackupResponsibility::MetadataPlusExternalRoot,
artifacts: vec![artifact(ArtifactKind::KeyMetadata, "vault/records/object-key.json.enc")],
local_kdf: None,
external_references: Some(transit_references()),
..local_manifest_unsealed()
}
}
#[test]
fn vault_bundle_round_trips_with_external_references() {
let sealed = seal(transit_manifest_unsealed());
let bytes = sealed.encode().expect("encoding should succeed");
let decoded = BackupManifest::decode(&bytes).expect("decoding should succeed");
assert_eq!(decoded, sealed);
}
#[test]
fn vault_backend_requires_external_references() {
let manifest = BackupManifest {
external_references: None,
..transit_manifest_unsealed()
};
expect_corrupted(
BackupManifest::decode(&serde_json::to_vec(&seal(manifest)).expect("serialize")),
"must carry external Vault references",
);
}
#[test]
fn non_vault_backend_rejects_external_references() {
let manifest = BackupManifest {
external_references: Some(transit_references()),
..local_manifest_unsealed()
};
expect_corrupted(
BackupManifest::decode(&serde_json::to_vec(&seal(manifest)).expect("serialize")),
"not valid for backend",
);
}
/// One way to contradict an otherwise valid set of Vault references.
type ReferenceContradiction = Box<dyn Fn(&mut VaultExternalReferences)>;
#[test]
fn external_reference_contradictions_fail_closed() {
let cases: [(&str, ReferenceContradiction); 7] = [
("cluster_id", Box::new(|refs| refs.cluster_id.clear())),
("kv_mount", Box::new(|refs| refs.kv_mount.clear())),
("namespace must not be empty", Box::new(|refs| refs.namespace = Some(String::new()))),
(
"must reference the Transit key",
Box::new(|refs: &mut VaultExternalReferences| refs.transit = None),
),
(
"at least key version 1",
Box::new(|refs: &mut VaultExternalReferences| {
if let Some(transit) = refs.transit.as_mut() {
transit.required_min_version = 0;
}
}),
),
(
"below the required minimum",
Box::new(|refs: &mut VaultExternalReferences| {
if let Some(transit) = refs.transit.as_mut() {
transit.current_version = 1;
}
}),
),
(
"appears more than once",
Box::new(|refs: &mut VaultExternalReferences| {
let first = refs.kv_records[0].clone();
refs.kv_records.push(first);
}),
),
];
for (needle, mutate) in cases {
let mut references = transit_references();
mutate(&mut references);
let manifest = BackupManifest {
external_references: Some(references),
..transit_manifest_unsealed()
};
expect_corrupted(BackupManifest::decode(&serde_json::to_vec(&seal(manifest)).expect("serialize")), needle);
}
}
/// The reference schema is the only place a Vault coordinate reaches a
/// bundle, so its field set is pinned: a credential-carrying field could
/// only appear by editing this assertion.
#[test]
fn external_reference_field_set_is_frozen() {
let value = serde_json::to_value(transit_references()).expect("serialize");
let mut top: Vec<&str> = value.as_object().expect("object").keys().map(String::as_str).collect();
top.sort_unstable();
assert_eq!(
top,
[
"cluster_id",
"kv_mount",
"kv_path_prefix",
"kv_records",
"namespace",
"transit"
]
);
let mut transit: Vec<&str> = value["transit"]
.as_object()
.expect("transit object")
.keys()
.map(String::as_str)
.collect();
transit.sort_unstable();
assert_eq!(
transit,
[
"current_version",
"key_name",
"mount_path",
"native_snapshot_reference",
"required_min_version"
]
);
let mut record: Vec<&str> = value["kv_records"][0]
.as_object()
.expect("record object")
.keys()
.map(String::as_str)
.collect();
record.sort_unstable();
assert_eq!(record, ["key_id", "kv_version"]);
}
} }
+27 -7
View File
@@ -12,13 +12,16 @@
// See the License for the specific language governing permissions and // See the License for the specific language governing permissions and
// limitations under the License. // limitations under the License.
//! Backup/restore contract types for KMS state. //! Backup/restore contracts and backup production for KMS state.
//! //!
//! This module is contract-only: it defines the versioned backup manifest, //! The contract side defines the versioned backup manifest, the per-backend
//! the per-backend responsibility matrix, typed failure modes, and the //! responsibility matrix, typed failure modes, and the restore dry-run
//! restore dry-run report. Nothing here is wired into handlers or backends; //! report. [`local_export`] implements the producer side and
//! backup export, restore orchestration, and the admin API build on these //! [`local_restore`] the consumer side for the Local backend;
//! types in follow-up changes. //! [`vault_restore`] orchestrates the consumer side for the Vault backends,
//! whose cryptographic root is restored by Vault's own disaster-recovery
//! flow. All are crate-internal APIs; the admin API builds on these pieces in
//! follow-up changes.
//! //!
//! # Bundle model //! # Bundle model
//! //!
@@ -50,14 +53,31 @@
mod capability; mod capability;
mod dry_run; mod dry_run;
mod error; mod error;
pub mod local_export;
pub mod local_restore;
mod manifest; mod manifest;
pub mod vault_restore;
pub use capability::{AtRestProtection, BackupBackendKind, BackupResponsibility}; pub use capability::{AtRestProtection, BackupBackendKind, BackupResponsibility};
pub use dry_run::{ pub use dry_run::{
ExternalDependencyMismatch, RestoreBlocker, RestoreBlockerCode, RestoreConflict, RestoreConflictKind, RestoreDryRunReport, ExternalDependencyMismatch, RestoreBlocker, RestoreBlockerCode, RestoreConflict, RestoreConflictKind, RestoreDryRunReport,
}; };
pub use error::BackupError; pub use error::BackupError;
pub use local_export::{
BackupKek, LOCAL_BUNDLE_MANIFEST_FILE, LocalBackupExportRequest, decrypt_bundle_artifact, export_local_backup,
read_bundle_manifest, read_local_bundle_manifest,
};
pub use local_restore::{
LocalRestoreReport, LocalRestoreRequest, RestoreConflictPolicy, abort_local_restore, dry_run_local_restore,
restore_local_backup,
};
pub use manifest::{ pub use manifest::{
AeadAlgorithm, ArtifactDescriptor, ArtifactKind, BackupKekDescriptor, BackupManifest, CompletenessState, ContentDigest, AeadAlgorithm, ArtifactDescriptor, ArtifactKind, BackupKekDescriptor, BackupManifest, CompletenessState, ContentDigest,
DigestAlgorithm, LocalKdfDescriptor, LocalKeyDerivation, ReservedSlot, DigestAlgorithm, LocalKdfDescriptor, LocalKeyDerivation, ReservedSlot, VaultExternalReferences, VaultKvRecordReference,
VaultTransitReference,
};
pub use vault_restore::{
VaultRestoreAbortReport, VaultRestoreAbortSkip, VaultRestoreAbortSkipReason, VaultRestoreClient, VaultRestoreMismatch,
VaultRestoreReport, VaultRestoreRequest, VaultRestoreSequence, VaultRestoreStage, VaultRestoreTarget, abort_vault_restore,
dry_run_vault_restore, restore_vault_backup,
}; };
File diff suppressed because it is too large Load Diff
+369 -48
View File
@@ -14,30 +14,125 @@
//! Caching layer for KMS operations to improve performance //! Caching layer for KMS operations to improve performance
use crate::config::CacheConfig;
use crate::types::KeyMetadata; use crate::types::KeyMetadata;
use moka::future::Cache; use moka::future::Cache;
use std::time::Duration; use moka::notification::RemovalCause;
use std::sync::Arc;
use std::sync::atomic::{AtomicU64, Ordering};
// ---------------------------------------------------------------------------
// Metrics
//
// Lookup outcomes and removals are counted here because moka exposes neither
// hit nor miss counts; the numbers below are the cache's own, not derived.
// Label values are exclusively static strings (lookup result, removal cause) —
// key identifiers and key metadata must never reach a metric label.
//
// Publication is gated by `CacheConfig::enable_metrics`; the atomics backing
// `KmsCache::stats` are maintained regardless, so the admin status API keeps
// reporting real numbers with the metrics switch off.
// ---------------------------------------------------------------------------
/// Counter: key metadata lookups, by `result` (`hit` or `miss`).
const METRIC_CACHE_LOOKUPS_TOTAL: &str = "rustfs_kms_metadata_cache_lookups_total";
/// Counter: entries dropped from the cache, by `cause` (`expired`, `size`,
/// `explicit`, `replaced`); only `expired` and `size` are true evictions, the
/// rest are invalidations driven by key lifecycle operations.
const METRIC_CACHE_EVICTIONS_TOTAL: &str = "rustfs_kms_metadata_cache_evictions_total";
/// Gauge: entries currently held by the cache. Republished whenever the entry
/// set can have changed — every write path, plus the lookups that miss, since
/// TTL expiry drops entries without any write taking place.
const METRIC_CACHE_ENTRIES: &str = "rustfs_kms_metadata_cache_entries";
/// Register metric descriptions once per process.
fn describe_metrics() {
static DESCRIBE: std::sync::Once = std::sync::Once::new();
DESCRIBE.call_once(|| {
metrics::describe_counter!(METRIC_CACHE_LOOKUPS_TOTAL, "Total KMS key metadata cache lookups, by hit or miss");
metrics::describe_counter!(
METRIC_CACHE_EVICTIONS_TOTAL,
"Total entries dropped from the KMS key metadata cache, by removal cause"
);
metrics::describe_gauge!(METRIC_CACHE_ENTRIES, "Entries currently held by the KMS key metadata cache");
});
}
fn removal_cause_label(cause: RemovalCause) -> &'static str {
match cause {
RemovalCause::Expired => "expired",
RemovalCause::Explicit => "explicit",
RemovalCause::Replaced => "replaced",
RemovalCause::Size => "size",
}
}
/// Cumulative cache counters. Shared with moka's eviction listener, which runs
/// outside any `&self` borrow, hence the `Arc` and the atomics.
#[derive(Debug, Default)]
struct CacheCounters {
hits: AtomicU64,
misses: AtomicU64,
evictions: AtomicU64,
}
/// Snapshot of the KMS key metadata cache counters.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub struct KmsCacheStats {
/// Entries currently held. Eventually consistent: moka applies pending
/// maintenance lazily.
pub entries: u64,
/// Lookups served from the cache since process start.
pub hits: u64,
/// Lookups that fell through to the backend since process start.
pub misses: u64,
/// Entries dropped since process start, whatever the cause (expiry,
/// capacity, invalidation, replacement).
pub evictions: u64,
}
/// KMS cache for storing frequently accessed keys and metadata /// KMS cache for storing frequently accessed keys and metadata
pub struct KmsCache { pub struct KmsCache {
key_metadata_cache: Cache<String, KeyMetadata>, key_metadata_cache: Cache<String, KeyMetadata>,
counters: Arc<CacheCounters>,
metrics_enabled: bool,
} }
impl KmsCache { impl KmsCache {
/// Create a new KMS cache with the specified capacity /// Create a new KMS cache from the operator-supplied cache configuration
///
/// The entry lifetime is [`CacheConfig::effective_ttl`] rather than the raw
/// configured value: this value arrives straight from the admin configure
/// API, and moka panics when built with a time-to-live beyond 1000 years.
/// ///
/// # Arguments /// # Arguments
/// * `capacity` - Maximum number of entries in the cache /// * `config` - Capacity, metadata lifetime and metrics switch to build with
/// ///
/// # Returns /// # Returns
/// A new instance of `KmsCache` /// A new instance of `KmsCache`
/// ///
pub fn new(capacity: u64) -> Self { pub fn new(config: &CacheConfig) -> Self {
let metrics_enabled = config.enable_metrics;
if metrics_enabled {
describe_metrics();
}
let counters = Arc::new(CacheCounters::default());
let eviction_counters = Arc::clone(&counters);
Self { Self {
key_metadata_cache: Cache::builder() key_metadata_cache: Cache::builder()
.max_capacity(capacity) .max_capacity(config.max_keys as u64)
.time_to_live(Duration::from_secs(300)) // 5 minutes default TTL .time_to_live(config.effective_ttl())
.eviction_listener(move |_key: Arc<String>, _metadata: KeyMetadata, cause: RemovalCause| {
eviction_counters.evictions.fetch_add(1, Ordering::Relaxed);
if metrics_enabled {
metrics::counter!(METRIC_CACHE_EVICTIONS_TOTAL, "cause" => removal_cause_label(cause)).increment(1);
}
})
.build(), .build(),
counters,
metrics_enabled,
} }
} }
@@ -50,7 +145,26 @@ impl KmsCache {
/// An `Option` containing the `KeyMetadata` if found, or `None` if not found /// An `Option` containing the `KeyMetadata` if found, or `None` if not found
/// ///
pub async fn get_key_metadata(&self, key_id: &str) -> Option<KeyMetadata> { pub async fn get_key_metadata(&self, key_id: &str) -> Option<KeyMetadata> {
self.key_metadata_cache.get(key_id).await let cached = self.key_metadata_cache.get(key_id).await;
self.record_lookup(cached.is_some());
if cached.is_none() {
// TTL expiry is the one way the entry set shrinks without a write,
// and a miss is where it surfaces. moka expires lazily: the miss is
// reported at once, but the removal that decrements `entry_count`
// and reaches the eviction listener lands in the maintenance pass
// moka runs opportunistically on reads, on its own interval. So
// this converges the gauge within a housekeeping interval rather
// than on the first miss — enough to stop a cache that only ever
// expires, with no put, remove or clear to follow, from reporting a
// population that is gone. Forcing `run_pending_tasks` would
// tighten that at the cost of taking moka's maintenance lock on
// every miss, for a gauge `entry_count` only approximates anyway.
// Hits, the hot path, are left alone.
self.record_entry_count();
}
cached
} }
/// Put key metadata into cache /// Put key metadata into cache
@@ -62,6 +176,7 @@ impl KmsCache {
pub async fn put_key_metadata(&mut self, key_id: &str, metadata: &KeyMetadata) { pub async fn put_key_metadata(&mut self, key_id: &str, metadata: &KeyMetadata) {
self.key_metadata_cache.insert(key_id.to_string(), metadata.clone()).await; self.key_metadata_cache.insert(key_id.to_string(), metadata.clone()).await;
self.key_metadata_cache.run_pending_tasks().await; self.key_metadata_cache.run_pending_tasks().await;
self.record_entry_count();
} }
/// Remove key metadata from cache /// Remove key metadata from cache
@@ -71,6 +186,11 @@ impl KmsCache {
/// ///
pub async fn remove_key_metadata(&mut self, key_id: &str) { pub async fn remove_key_metadata(&mut self, key_id: &str) {
self.key_metadata_cache.remove(key_id).await; self.key_metadata_cache.remove(key_id).await;
// Flush maintenance so the removal notification and the entry gauge
// describe the cache as it is once this call returns.
self.key_metadata_cache.run_pending_tasks().await;
self.record_entry_count();
} }
/// Clear all cached entries /// Clear all cached entries
@@ -79,18 +199,39 @@ impl KmsCache {
// Wait for invalidation to complete // Wait for invalidation to complete
self.key_metadata_cache.run_pending_tasks().await; self.key_metadata_cache.run_pending_tasks().await;
self.record_entry_count();
} }
/// Get cache statistics (hit count, miss count) /// Get cache statistics
/// ///
/// # Returns /// # Returns
/// A tuple containing total entries and total misses /// A [`KmsCacheStats`] snapshot of entry count, hits, misses and evictions
/// ///
pub fn stats(&self) -> (u64, u64) { pub fn stats(&self) -> KmsCacheStats {
( KmsCacheStats {
self.key_metadata_cache.entry_count(), entries: self.key_metadata_cache.entry_count(),
0u64, // moka doesn't provide miss count directly hits: self.counters.hits.load(Ordering::Relaxed),
) misses: self.counters.misses.load(Ordering::Relaxed),
evictions: self.counters.evictions.load(Ordering::Relaxed),
}
}
fn record_lookup(&self, hit: bool) {
let (counter, result) = if hit {
(&self.counters.hits, "hit")
} else {
(&self.counters.misses, "miss")
};
counter.fetch_add(1, Ordering::Relaxed);
if self.metrics_enabled {
metrics::counter!(METRIC_CACHE_LOOKUPS_TOTAL, "result" => result).increment(1);
}
}
fn record_entry_count(&self) {
if self.metrics_enabled {
metrics::gauge!(METRIC_CACHE_ENTRIES).set(self.key_metadata_cache.entry_count() as f64);
}
} }
} }
@@ -99,40 +240,101 @@ mod tests {
use super::*; use super::*;
use crate::types::{KeyState, KeyUsage}; use crate::types::{KeyState, KeyUsage};
use jiff::Zoned; use jiff::Zoned;
use metrics_util::MetricKind;
use metrics_util::debugging::{DebugValue, DebuggingRecorder};
use std::time::Duration; use std::time::Duration;
#[derive(Debug, Clone)]
struct CacheInfo {
key_metadata_count: u64,
}
impl CacheInfo {
fn total_entries(&self) -> u64 {
self.key_metadata_count
}
}
impl KmsCache { impl KmsCache {
fn with_ttl_for_tests(capacity: u64, metadata_ttl: Duration) -> Self {
Self {
key_metadata_cache: Cache::builder().max_capacity(capacity).time_to_live(metadata_ttl).build(),
}
}
fn info_for_tests(&self) -> CacheInfo {
CacheInfo {
key_metadata_count: self.key_metadata_cache.entry_count(),
}
}
fn contains_key_metadata_for_tests(&self, key_id: &str) -> bool { fn contains_key_metadata_for_tests(&self, key_id: &str) -> bool {
self.key_metadata_cache.contains_key(key_id) self.key_metadata_cache.contains_key(key_id)
} }
} }
/// A cache holding `max_keys` entries for the default lifetime.
fn sized(max_keys: usize) -> KmsCache {
KmsCache::new(&CacheConfig {
max_keys,
..Default::default()
})
}
fn test_metadata(key_id: &str) -> KeyMetadata {
KeyMetadata {
key_id: key_id.to_string(),
key_state: KeyState::Enabled,
key_usage: KeyUsage::EncryptDecrypt,
description: None,
creation_date: Zoned::now(),
deletion_date: None,
origin: "KMS".to_string(),
key_manager: "CUSTOMER".to_string(),
tags: std::collections::HashMap::new(),
}
}
type MetricEntry = (
metrics_util::CompositeKey,
Option<metrics::Unit>,
Option<metrics::SharedString>,
DebugValue,
);
/// Drive `test` on a current-thread runtime under a thread-local debugging
/// recorder and return its output plus one snapshot of everything emitted.
///
/// A single snapshot per test on purpose: `Snapshotter::snapshot` drains
/// the recorded state, so taking it per assertion would leave every
/// assertion after the first with nothing to look at.
fn record_metrics<F, Fut, T>(test: F) -> (T, Vec<MetricEntry>)
where
F: FnOnce() -> Fut,
Fut: std::future::Future<Output = T>,
{
let recorder = DebuggingRecorder::new();
let snapshotter = recorder.snapshotter();
let output = metrics::with_local_recorder(&recorder, || {
let runtime = tokio::runtime::Builder::new_current_thread()
.enable_all()
.build()
.expect("current-thread runtime must build");
runtime.block_on(test())
});
(output, snapshotter.snapshot().into_vec())
}
/// Sum of counters with `name` whose labels include every `(key, value)` pair.
fn counter_value(snapshot: &[MetricEntry], name: &str, labels: &[(&str, &str)]) -> u64 {
snapshot
.iter()
.filter_map(|(composite, _unit, _description, value)| {
let key = composite.key();
let matches = composite.kind() == MetricKind::Counter
&& key.name() == name
&& labels
.iter()
.all(|(label, expected)| key.labels().any(|l| l.key() == *label && l.value() == *expected));
match (matches, value) {
(true, DebugValue::Counter(count)) => Some(*count),
_ => None,
}
})
.sum()
}
/// Last value recorded for the unlabelled gauge `name`.
fn gauge_value(snapshot: &[MetricEntry], name: &str) -> Option<f64> {
snapshot.iter().find_map(|(composite, _unit, _description, value)| {
let matches = composite.kind() == MetricKind::Gauge && composite.key().name() == name;
match (matches, value) {
(true, DebugValue::Gauge(gauge)) => Some(**gauge),
_ => None,
}
})
}
#[tokio::test] #[tokio::test]
async fn test_cache_operations() { async fn test_cache_operations() {
let mut cache = KmsCache::new(100); let mut cache = sized(100);
// Test key metadata caching // Test key metadata caching
let metadata = KeyMetadata { let metadata = KeyMetadata {
@@ -154,22 +356,20 @@ mod tests {
assert_eq!(retrieved.expect("metadata should be cached").key_id, "test-key-1"); assert_eq!(retrieved.expect("metadata should be cached").key_id, "test-key-1");
// Test cache info // Test cache info
let info = cache.info_for_tests(); assert_eq!(cache.stats().entries, 1);
assert_eq!(info.key_metadata_count, 1);
assert_eq!(info.total_entries(), 1);
// Test cache clearing // Test cache clearing
cache.clear().await; cache.clear().await;
let info_after_clear = cache.info_for_tests(); assert_eq!(cache.stats().entries, 0);
assert_eq!(info_after_clear.total_entries(), 0);
} }
#[tokio::test] #[tokio::test]
async fn test_cache_with_custom_ttl() { async fn test_cache_with_custom_ttl() {
let mut cache = KmsCache::with_ttl_for_tests( let mut cache = KmsCache::new(&CacheConfig {
100, max_keys: 100,
Duration::from_millis(100), // Short TTL for testing ttl: Duration::from_millis(100), // Short TTL for testing
); ..Default::default()
});
let metadata = KeyMetadata { let metadata = KeyMetadata {
key_id: "ttl-test-key".to_string(), key_id: "ttl-test-key".to_string(),
@@ -197,7 +397,7 @@ mod tests {
#[tokio::test] #[tokio::test]
async fn test_cache_contains_methods() { async fn test_cache_contains_methods() {
let mut cache = KmsCache::new(100); let mut cache = sized(100);
assert!(!cache.contains_key_metadata_for_tests("nonexistent")); assert!(!cache.contains_key_metadata_for_tests("nonexistent"));
@@ -217,4 +417,125 @@ mod tests {
assert!(cache.contains_key_metadata_for_tests("contains-test")); assert!(cache.contains_key_metadata_for_tests("contains-test"));
} }
#[test]
fn lookups_report_real_hit_and_miss_counts() {
let (stats, snapshot) = record_metrics(|| async {
let mut cache = sized(100);
assert!(cache.get_key_metadata("absent").await.is_none());
cache.put_key_metadata("present", &test_metadata("present")).await;
assert!(cache.get_key_metadata("present").await.is_some());
assert!(cache.get_key_metadata("absent").await.is_none());
cache.stats()
});
assert_eq!(stats.hits, 1);
assert_eq!(stats.misses, 2);
assert_eq!(counter_value(&snapshot, METRIC_CACHE_LOOKUPS_TOTAL, &[("result", "hit")]), 1);
assert_eq!(counter_value(&snapshot, METRIC_CACHE_LOOKUPS_TOTAL, &[("result", "miss")]), 2);
}
#[test]
fn disabling_metrics_stops_publication_without_blinding_the_status_api() {
let (stats, snapshot) = record_metrics(|| async {
let mut cache = KmsCache::new(&CacheConfig {
max_keys: 1,
enable_metrics: false,
..Default::default()
});
assert!(cache.get_key_metadata("absent").await.is_none());
cache.put_key_metadata("first", &test_metadata("first")).await;
assert!(cache.get_key_metadata("first").await.is_some());
// Capacity is one, so this insert evicts "first".
cache.put_key_metadata("second", &test_metadata("second")).await;
cache.stats()
});
// The counters behind the admin status API keep moving...
assert_eq!(stats.hits, 1);
assert_eq!(stats.misses, 1);
assert_eq!(stats.evictions, 1);
// ...while nothing reaches the metrics recorder.
assert_eq!(counter_value(&snapshot, METRIC_CACHE_LOOKUPS_TOTAL, &[]), 0);
assert_eq!(counter_value(&snapshot, METRIC_CACHE_EVICTIONS_TOTAL, &[]), 0);
assert_eq!(gauge_value(&snapshot, METRIC_CACHE_ENTRIES), None);
}
#[test]
fn removals_report_their_cause_and_the_resulting_entry_count() {
let (stats, snapshot) = record_metrics(|| async {
let mut cache = sized(100);
cache.put_key_metadata("key", &test_metadata("key")).await;
cache.put_key_metadata("key", &test_metadata("key")).await;
cache.remove_key_metadata("key").await;
cache.stats()
});
assert_eq!(counter_value(&snapshot, METRIC_CACHE_EVICTIONS_TOTAL, &[("cause", "replaced")]), 1);
assert_eq!(counter_value(&snapshot, METRIC_CACHE_EVICTIONS_TOTAL, &[("cause", "explicit")]), 1);
assert_eq!(stats.evictions, 2);
assert_eq!(stats.entries, 0);
assert_eq!(gauge_value(&snapshot, METRIC_CACHE_ENTRIES), Some(0.0));
}
#[test]
fn capacity_pressure_reports_size_evictions() {
let (stats, snapshot) = record_metrics(|| async {
let mut cache = sized(1);
cache.put_key_metadata("first", &test_metadata("first")).await;
cache.put_key_metadata("second", &test_metadata("second")).await;
cache.stats()
});
assert_eq!(counter_value(&snapshot, METRIC_CACHE_EVICTIONS_TOTAL, &[("cause", "size")]), 1);
assert_eq!(stats.evictions, 1);
assert_eq!(stats.entries, 1);
assert_eq!(gauge_value(&snapshot, METRIC_CACHE_ENTRIES), Some(1.0));
}
/// Entries that expire are dropped outside every write path, so the gauge
/// has to be republished from the lookup that observes the expiry — nothing
/// else runs afterwards to correct it.
///
/// moka's clock is internal and cannot be driven by tokio's paused time,
/// hence the one short real sleep. The explicit `run_pending_tasks` stands
/// in for the maintenance moka schedules by itself on reads, so the test
/// does not depend on moka's internal 300ms housekeeping interval.
#[test]
fn expiry_without_further_writes_converges_the_entry_gauge() {
let ttl = Duration::from_millis(100);
let (stats, snapshot) = record_metrics(move || async move {
let mut cache = KmsCache::new(&CacheConfig {
max_keys: 100,
ttl,
..Default::default()
});
cache.put_key_metadata("expiring", &test_metadata("expiring")).await;
assert_eq!(cache.stats().entries, 1);
tokio::time::sleep(Duration::from_millis(150)).await;
cache.key_metadata_cache.run_pending_tasks().await;
// The lookup that observes the expiry is the last thing to touch
// the cache: no put, remove or clear follows it.
assert!(cache.get_key_metadata("expiring").await.is_none());
cache.stats()
});
assert_eq!(counter_value(&snapshot, METRIC_CACHE_EVICTIONS_TOTAL, &[("cause", "expired")]), 1);
assert_eq!(stats.entries, 0);
assert_eq!(gauge_value(&snapshot, METRIC_CACHE_ENTRIES), Some(0.0));
}
} }
+300 -5
View File
@@ -24,6 +24,7 @@ use std::time::Duration;
use url::Url; use url::Url;
pub const ENV_KMS_ALLOW_INSECURE_DEV_DEFAULTS: &str = "RUSTFS_KMS_ALLOW_INSECURE_DEV_DEFAULTS"; pub const ENV_KMS_ALLOW_INSECURE_DEV_DEFAULTS: &str = "RUSTFS_KMS_ALLOW_INSECURE_DEV_DEFAULTS";
pub const ENV_KMS_ALLOW_IMMEDIATE_DELETION: &str = "RUSTFS_KMS_ALLOW_IMMEDIATE_DELETION";
pub const ENV_KMS_VAULT_SKIP_TLS_VERIFY: &str = "RUSTFS_KMS_VAULT_SKIP_TLS_VERIFY"; pub const ENV_KMS_VAULT_SKIP_TLS_VERIFY: &str = "RUSTFS_KMS_VAULT_SKIP_TLS_VERIFY";
pub const ENV_KMS_VAULT_TRANSIT_METADATA_KV_MOUNT: &str = "RUSTFS_KMS_VAULT_TRANSIT_METADATA_KV_MOUNT"; pub const ENV_KMS_VAULT_TRANSIT_METADATA_KV_MOUNT: &str = "RUSTFS_KMS_VAULT_TRANSIT_METADATA_KV_MOUNT";
pub const ENV_KMS_VAULT_TRANSIT_METADATA_PREFIX: &str = "RUSTFS_KMS_VAULT_TRANSIT_METADATA_PREFIX"; pub const ENV_KMS_VAULT_TRANSIT_METADATA_PREFIX: &str = "RUSTFS_KMS_VAULT_TRANSIT_METADATA_PREFIX";
@@ -34,6 +35,8 @@ pub const ENV_KMS_VAULT_APPROLE_SECRET_ID: &str = "RUSTFS_KMS_VAULT_APPROLE_SECR
pub const ENV_KMS_VAULT_APPROLE_SECRET_ID_FILE: &str = "RUSTFS_KMS_VAULT_APPROLE_SECRET_ID_FILE"; pub const ENV_KMS_VAULT_APPROLE_SECRET_ID_FILE: &str = "RUSTFS_KMS_VAULT_APPROLE_SECRET_ID_FILE";
pub const ENV_KMS_VAULT_APPROLE_MOUNT: &str = "RUSTFS_KMS_VAULT_APPROLE_MOUNT"; pub const ENV_KMS_VAULT_APPROLE_MOUNT: &str = "RUSTFS_KMS_VAULT_APPROLE_MOUNT";
pub const ENV_KMS_VAULT_TOKEN_FILE: &str = "RUSTFS_KMS_VAULT_TOKEN_FILE"; pub const ENV_KMS_VAULT_TOKEN_FILE: &str = "RUSTFS_KMS_VAULT_TOKEN_FILE";
pub const ENV_KMS_AWS_REGION: &str = "RUSTFS_KMS_AWS_REGION";
pub const ENV_KMS_AWS_ENDPOINT_URL: &str = "RUSTFS_KMS_AWS_ENDPOINT_URL";
pub const DEFAULT_VAULT_TRANSIT_METADATA_KV_MOUNT: &str = "secret"; pub const DEFAULT_VAULT_TRANSIT_METADATA_KV_MOUNT: &str = "secret";
pub const DEFAULT_VAULT_TRANSIT_METADATA_KEY_PREFIX: &str = "rustfs/kms/transit-metadata"; pub const DEFAULT_VAULT_TRANSIT_METADATA_KEY_PREFIX: &str = "rustfs/kms/transit-metadata";
pub const DEFAULT_VAULT_APPROLE_MOUNT: &str = "approle"; pub const DEFAULT_VAULT_APPROLE_MOUNT: &str = "approle";
@@ -47,6 +50,19 @@ pub(crate) const MAX_OPERATION_TIMEOUT: Duration = Duration::from_secs(300);
/// Upper bound applied to `KmsConfig::retry_attempts` when deriving backend behavior. /// Upper bound applied to `KmsConfig::retry_attempts` when deriving backend behavior.
pub(crate) const MAX_RETRY_ATTEMPTS: u32 = 10; pub(crate) const MAX_RETRY_ATTEMPTS: u32 = 10;
/// Default number of key metadata entries the cache holds.
pub const DEFAULT_MAX_CACHED_KEYS: usize = 1000;
/// Default lifetime of a cached key metadata entry.
pub const DEFAULT_CACHE_TTL: Duration = Duration::from_secs(300);
/// Upper bound applied to `CacheConfig::ttl` when building the metadata cache.
///
/// Out-of-range values are clamped at use rather than rejected, matching
/// `MAX_OPERATION_TIMEOUT`, so existing deployments with an oversized setting
/// keep starting after an upgrade.
pub(crate) const MAX_CACHE_TTL: Duration = Duration::from_secs(24 * 60 * 60);
fn default_vault_transit_metadata_kv_mount() -> String { fn default_vault_transit_metadata_kv_mount() -> String {
DEFAULT_VAULT_TRANSIT_METADATA_KV_MOUNT.to_string() DEFAULT_VAULT_TRANSIT_METADATA_KV_MOUNT.to_string()
} }
@@ -128,6 +144,26 @@ pub enum KmsBackend {
/// Static single-key backend that derives DEKs from a pre-configured key /// Static single-key backend that derives DEKs from a pre-configured key
#[serde(rename = "Static")] #[serde(rename = "Static")]
Static, Static,
/// AWS KMS backend: AWS is the cryptographic source of truth and owns key
/// state, versioning, and the deletion window.
#[serde(rename = "AWS", alias = "AwsKms")]
Aws,
}
impl KmsBackend {
/// Stable identifier for logs, metrics, and audit records.
///
/// External consumers key off these values, so treat them as a wire
/// contract rather than a rendering detail.
pub fn as_str(&self) -> &'static str {
match self {
KmsBackend::VaultKv2 => "vault-kv2",
KmsBackend::VaultTransit => "vault-transit",
KmsBackend::Local => "local",
KmsBackend::Static => "static",
KmsBackend::Aws => "aws",
}
}
} }
/// Main KMS configuration /// Main KMS configuration
@@ -142,6 +178,24 @@ pub struct KmsConfig {
/// Allow development-only insecure defaults such as plaintext local keys or HTTP Vault. /// Allow development-only insecure defaults such as plaintext local keys or HTTP Vault.
#[serde(default)] #[serde(default)]
pub allow_insecure_dev_defaults: bool, pub allow_insecure_dev_defaults: bool,
/// Allow `DeleteKey` requests to skip the pending-deletion waiting window and
/// destroy key material right away.
///
/// Off by default: an immediate deletion is unrecoverable and takes every
/// object encrypted under the key with it, so the waiting window (plus
/// `CancelKeyDeletion`) is the only recovery path. Operators who genuinely
/// need immediate deletion — throwaway test clusters, key material that was
/// never used — must turn it on through server configuration
/// ([`ENV_KMS_ALLOW_IMMEDIATE_DELETION`]); the request must still echo the
/// key id back for confirmation.
///
/// Not part of the serialized configuration, and not settable through the
/// admin configure API. It is per-server operator state that has to be
/// re-stated to survive a restart: persisting it would carry one operator's
/// one-time enablement into the cluster-wide config that every node reloads,
/// long after the deletion it was turned on for.
#[serde(skip)]
pub allow_immediate_deletion: bool,
/// Timeout for a single backend attempt. /// Timeout for a single backend attempt.
/// ///
/// This bounds one outbound request, not the whole operation: the operation /// This bounds one outbound request, not the whole operation: the operation
@@ -165,6 +219,7 @@ impl Default for KmsConfig {
default_key_id: None, default_key_id: None,
backend_config: BackendConfig::default(), backend_config: BackendConfig::default(),
allow_insecure_dev_defaults: false, allow_insecure_dev_defaults: false,
allow_immediate_deletion: false,
timeout: Duration::from_secs(30), timeout: Duration::from_secs(30),
retry_attempts: 3, retry_attempts: 3,
enable_cache: true, enable_cache: true,
@@ -185,6 +240,9 @@ pub enum BackendConfig {
VaultTransit(Box<VaultTransitConfig>), VaultTransit(Box<VaultTransitConfig>),
/// Static single-key backend configuration /// Static single-key backend configuration
Static(StaticConfig), Static(StaticConfig),
/// AWS KMS backend configuration
#[serde(rename = "AWS", alias = "AwsKms")]
Aws(Box<AwsKmsConfig>),
} }
impl Default for BackendConfig { impl Default for BackendConfig {
@@ -200,6 +258,7 @@ impl fmt::Debug for BackendConfig {
Self::VaultKv2(config) => f.debug_tuple("VaultKv2").field(config).finish(), Self::VaultKv2(config) => f.debug_tuple("VaultKv2").field(config).finish(),
Self::VaultTransit(config) => f.debug_tuple("VaultTransit").field(config).finish(), Self::VaultTransit(config) => f.debug_tuple("VaultTransit").field(config).finish(),
Self::Static(config) => f.debug_tuple("Static").field(config).finish(), Self::Static(config) => f.debug_tuple("Static").field(config).finish(),
Self::Aws(config) => f.debug_tuple("Aws").field(config).finish(),
} }
} }
} }
@@ -509,22 +568,59 @@ pub struct TlsConfig {
pub struct CacheConfig { pub struct CacheConfig {
/// Maximum number of keys to cache /// Maximum number of keys to cache
pub max_keys: usize, pub max_keys: usize,
/// TTL for cached keys /// Lifetime of a cached key metadata entry.
///
/// This bounds how long a describe can answer from metadata that another
/// node has since changed (disable, schedule-deletion); encrypt, decrypt
/// and data key generation never read the cache. Values above 24 hours are
/// clamped at use (see [`CacheConfig::effective_ttl`]).
pub ttl: Duration, pub ttl: Duration,
/// Enable cache metrics /// Publish the `rustfs_kms_metadata_cache_*` metrics.
///
/// Only metrics-recorder output is gated: the counters behind the admin
/// status API are maintained either way.
pub enable_metrics: bool, pub enable_metrics: bool,
} }
impl Default for CacheConfig { impl Default for CacheConfig {
fn default() -> Self { fn default() -> Self {
Self { Self {
max_keys: 1000, max_keys: DEFAULT_MAX_CACHED_KEYS,
ttl: Duration::from_secs(3600), // 1 hour ttl: DEFAULT_CACHE_TTL,
enable_metrics: true, enable_metrics: true,
} }
} }
} }
impl CacheConfig {
/// Metadata lifetime with the configured value clamped to the supported maximum.
///
/// This is the value the cache is built with and the value reported back to
/// operators, so what the admin API advertises is what the cache does.
pub fn effective_ttl(&self) -> Duration {
self.ttl.min(MAX_CACHE_TTL)
}
}
/// AWS KMS backend configuration.
///
/// Deliberately holds no credential material: the backend resolves credentials
/// through the standard `aws-config` provider chain (environment, shared
/// profile, container/IMDS role), so RustFS never stores, persists, or redacts
/// AWS secrets of its own.
#[derive(Clone, Debug, Default, Serialize, Deserialize)]
pub struct AwsKmsConfig {
/// AWS region hosting the KMS keys. When unset, the region is resolved by
/// the standard chain (`AWS_REGION`, profile, IMDS).
#[serde(default)]
pub region: Option<String>,
/// Override for the KMS endpoint, for local emulators and private
/// endpoints. Unset in production, where the SDK derives the regional
/// endpoint.
#[serde(default)]
pub endpoint_url: Option<String>,
}
impl KmsConfig { impl KmsConfig {
/// Create a new KMS configuration for local backend (for development and testing only) /// Create a new KMS configuration for local backend (for development and testing only)
pub fn local(key_dir: PathBuf) -> Self { pub fn local(key_dir: PathBuf) -> Self {
@@ -634,6 +730,29 @@ impl KmsConfig {
} }
} }
/// Create a new KMS configuration for the AWS KMS backend.
///
/// Credentials are resolved by the standard `aws-config` provider chain;
/// only the region is configured here.
pub fn aws(region: Option<String>) -> Self {
Self {
backend: KmsBackend::Aws,
backend_config: BackendConfig::Aws(Box::new(AwsKmsConfig {
region,
endpoint_url: None,
})),
..Default::default()
}
}
/// Get the AWS configuration if backend is AWS KMS
pub fn aws_kms_config(&self) -> Option<&AwsKmsConfig> {
match &self.backend_config {
BackendConfig::Aws(config) => Some(config),
_ => None,
}
}
/// Set default key ID /// Set default key ID
pub fn with_default_key(mut self, key_id: String) -> Self { pub fn with_default_key(mut self, key_id: String) -> Self {
self.default_key_id = Some(key_id); self.default_key_id = Some(key_id);
@@ -646,6 +765,12 @@ impl KmsConfig {
self self
} }
/// Explicitly allow deletions that bypass the pending-deletion waiting window.
pub fn with_immediate_deletion_allowed(mut self) -> Self {
self.allow_immediate_deletion = true;
self
}
/// Set operation timeout /// Set operation timeout
pub fn with_timeout(mut self, timeout: Duration) -> Self { pub fn with_timeout(mut self, timeout: Duration) -> Self {
self.timeout = timeout; self.timeout = timeout;
@@ -790,13 +915,41 @@ impl KmsConfig {
// Validate that the key can be decoded (right length, valid base64) // Validate that the key can be decoded (right length, valid base64)
config.decode_key()?; config.decode_key()?;
} }
BackendConfig::Aws(config) => {
if let Some(region) = &config.region
&& region.is_empty()
{
return Err(KmsError::configuration_error("AWS KMS region cannot be empty when set"));
}
if let Some(endpoint) = &config.endpoint_url {
if !endpoint.starts_with("http://") && !endpoint.starts_with("https://") {
return Err(KmsError::configuration_error("AWS KMS endpoint URL must use http or https scheme"));
}
// A plaintext endpoint override exposes every KMS request,
// including plaintext data keys, so it stays gated on the
// explicit development opt-in.
if endpoint.starts_with("http://") && !self.allow_insecure_dev_defaults {
return Err(development_default_error("AWS KMS endpoint URL must use https"));
}
}
}
} }
// Validate cache configuration // Validate cache configuration
if self.enable_cache && self.cache_config.max_keys == 0 { if self.enable_cache {
if self.cache_config.max_keys == 0 {
return Err(KmsError::configuration_error("Cache max_keys must be greater than 0")); return Err(KmsError::configuration_error("Cache max_keys must be greater than 0"));
} }
// A zero TTL expires every entry on insert, which is a cache that
// cannot serve anything; reject it instead of silently disabling
// the cache the operator asked for.
if self.cache_config.ttl.is_zero() {
return Err(KmsError::configuration_error("Cache ttl must be greater than 0"));
}
}
Ok(()) Ok(())
} }
@@ -811,6 +964,7 @@ impl KmsConfig {
"vault" | "vault-kv2" | "vault_kv2" => KmsBackend::VaultKv2, "vault" | "vault-kv2" | "vault_kv2" => KmsBackend::VaultKv2,
"vault-transit" | "vault_transit" => KmsBackend::VaultTransit, "vault-transit" | "vault_transit" => KmsBackend::VaultTransit,
"static" => KmsBackend::Static, "static" => KmsBackend::Static,
"aws" | "aws-kms" | "aws_kms" => KmsBackend::Aws,
_ => return Err(KmsError::configuration_error(format!("Unknown KMS backend: {backend_type}"))), _ => return Err(KmsError::configuration_error(format!("Unknown KMS backend: {backend_type}"))),
}; };
} }
@@ -839,6 +993,7 @@ impl KmsConfig {
config.enable_cache = get_env_bool("RUSTFS_KMS_ENABLE_CACHE", config.enable_cache); config.enable_cache = get_env_bool("RUSTFS_KMS_ENABLE_CACHE", config.enable_cache);
config.allow_insecure_dev_defaults = config.allow_insecure_dev_defaults =
get_env_bool(ENV_KMS_ALLOW_INSECURE_DEV_DEFAULTS, config.allow_insecure_dev_defaults); get_env_bool(ENV_KMS_ALLOW_INSECURE_DEV_DEFAULTS, config.allow_insecure_dev_defaults);
config.allow_immediate_deletion = get_env_bool(ENV_KMS_ALLOW_IMMEDIATE_DELETION, config.allow_immediate_deletion);
// Backend-specific configuration // Backend-specific configuration
match config.backend { match config.backend {
@@ -939,6 +1094,14 @@ impl KmsConfig {
}); });
config.default_key_id = Some(key_id); config.default_key_id = Some(key_id);
} }
KmsBackend::Aws => {
// Only non-credential settings are read here; access keys,
// profiles, and role assumption stay with the aws-config chain.
config.backend_config = BackendConfig::Aws(Box::new(AwsKmsConfig {
region: get_env_opt_str(ENV_KMS_AWS_REGION),
endpoint_url: get_env_opt_str(ENV_KMS_AWS_ENDPOINT_URL),
}));
}
} }
config.validate()?; config.validate()?;
@@ -946,6 +1109,15 @@ impl KmsConfig {
} }
} }
/// Read the immediate-deletion gate from the environment.
///
/// Callers that assemble a [`KmsConfig`] field by field instead of going
/// through [`KmsConfig::from_env`] use this, so the gate keeps one name, one
/// default, and one place to look it up.
pub fn allow_immediate_deletion_from_env() -> bool {
get_env_bool(ENV_KMS_ALLOW_IMMEDIATE_DELETION, false)
}
fn vault_tls_config(skip_tls_verify: bool) -> Option<TlsConfig> { fn vault_tls_config(skip_tls_verify: bool) -> Option<TlsConfig> {
skip_tls_verify.then_some(TlsConfig { skip_tls_verify.then_some(TlsConfig {
ca_cert_path: None, ca_cert_path: None,
@@ -1106,6 +1278,60 @@ mod tests {
assert_eq!(local_config.key_dir, temp_dir.path()); assert_eq!(local_config.key_dir, temp_dir.path());
} }
#[test]
fn oversized_cache_ttl_is_clamped_not_rejected() {
let temp_dir = TempDir::new().expect("Failed to create temp dir");
let config = KmsConfig::local(temp_dir.path().to_path_buf()).with_insecure_development_defaults();
assert_eq!(config.cache_config.ttl, DEFAULT_CACHE_TTL);
assert_eq!(config.cache_config.effective_ttl(), DEFAULT_CACHE_TTL);
// An oversized lifetime must not keep the service from starting, and
// must not reach moka, whose builder panics beyond 1000 years.
let config = KmsConfig {
cache_config: CacheConfig {
ttl: Duration::from_secs(u64::MAX),
..Default::default()
},
..config
};
assert!(config.validate().is_ok());
assert_eq!(config.cache_config.effective_ttl(), MAX_CACHE_TTL);
// In-range values pass through unchanged.
let config = KmsConfig {
cache_config: CacheConfig {
ttl: Duration::from_secs(600),
..Default::default()
},
..config
};
assert!(config.validate().is_ok());
assert_eq!(config.cache_config.effective_ttl(), Duration::from_secs(600));
}
#[test]
fn zero_cache_ttl_is_rejected_only_while_caching_is_enabled() {
let temp_dir = TempDir::new().expect("Failed to create temp dir");
let config = KmsConfig {
cache_config: CacheConfig {
ttl: Duration::ZERO,
..Default::default()
},
..KmsConfig::local(temp_dir.path().to_path_buf()).with_insecure_development_defaults()
};
// A cache that expires every entry on insert is a misconfiguration, not
// a way to turn caching off.
assert!(config.validate().is_err());
let config = KmsConfig {
enable_cache: false,
..config
};
assert!(config.validate().is_ok());
}
#[test] #[test]
fn test_oversized_timeout_and_retries_clamped_not_rejected() { fn test_oversized_timeout_and_retries_clamped_not_rejected() {
let temp_dir = TempDir::new().expect("Failed to create temp dir"); let temp_dir = TempDir::new().expect("Failed to create temp dir");
@@ -1540,6 +1766,43 @@ mod tests {
); );
} }
/// The AWS backend reads only non-credential settings from the
/// environment; access keys, profiles, and role assumption stay with the
/// aws-config provider chain.
#[test]
fn test_from_env_selects_aws_backend_without_credentials() {
with_vars(
vec![
("RUSTFS_KMS_BACKEND", Some("aws")),
(ENV_KMS_AWS_REGION, Some("eu-central-1")),
(ENV_KMS_AWS_ENDPOINT_URL, None::<&str>),
],
|| {
let config = KmsConfig::from_env().expect("kms config should load from env");
assert_eq!(config.backend, KmsBackend::Aws);
let aws = config.aws_kms_config().expect("aws backend config");
assert_eq!(aws.region.as_deref(), Some("eu-central-1"));
assert_eq!(aws.endpoint_url, None);
},
);
}
/// A plaintext endpoint override exposes every KMS request, including the
/// plaintext data keys, so it stays behind the explicit development opt-in.
#[test]
fn test_from_env_rejects_plaintext_aws_endpoint() {
with_vars(
vec![
("RUSTFS_KMS_BACKEND", Some("aws")),
(ENV_KMS_AWS_REGION, Some("us-east-1")),
(ENV_KMS_AWS_ENDPOINT_URL, Some("http://localhost:4566")),
],
|| {
KmsConfig::from_env().expect_err("a plaintext AWS endpoint must be rejected by default");
},
);
}
#[test] #[test]
fn test_from_env_approle_requires_secret_id_or_file() { fn test_from_env_approle_requires_secret_id_or_file() {
with_vars( with_vars(
@@ -1669,6 +1932,38 @@ mod tests {
assert_eq!(refresh_safety_window_secs, None); assert_eq!(refresh_safety_window_secs, None);
} }
/// The gate lives in server configuration only: it never rides along in a
/// serialized config, and a stored config that claims it must not be
/// believed. Otherwise one operator's one-time enablement would reach every
/// node that later reloads that config.
#[test]
fn immediate_deletion_gate_is_server_local_and_never_persisted() {
let persisted =
serde_json::to_value(KmsConfig::default().with_immediate_deletion_allowed()).expect("kms config should serialize");
assert!(
persisted.get("allow_immediate_deletion").is_none(),
"the gate must not be written into a persisted config: {persisted}"
);
let mut forged = persisted;
forged
.as_object_mut()
.expect("a persisted config must be a JSON object")
.insert("allow_immediate_deletion".to_string(), serde_json::json!(true));
let restored: KmsConfig = serde_json::from_value(forged).expect("an unknown gate field must not break loading");
assert!(!restored.allow_immediate_deletion, "a stored gate must fail closed");
with_vars(vec![(ENV_KMS_ALLOW_IMMEDIATE_DELETION, Some("true"))], || {
assert!(
allow_immediate_deletion_from_env(),
"the gate must be reachable from server configuration"
);
});
with_vars(vec![(ENV_KMS_ALLOW_IMMEDIATE_DELETION, None::<&str>)], || {
assert!(!allow_immediate_deletion_from_env());
});
}
#[test] #[test]
fn test_validate_rejects_incomplete_approle() { fn test_validate_rejects_incomplete_approle() {
let mut config = KmsConfig::vault_approle( let mut config = KmsConfig::vault_approle(
+380 -18
View File
@@ -22,18 +22,123 @@
//! every node of a deployment concurrently — a key is only ever removed while //! every node of a deployment concurrently — a key is only ever removed while
//! its (re-read) record is an expired pending deletion or a tombstone. //! its (re-read) record is an expired pending deletion or a tombstone.
use crate::audit::{KmsAuditOperation, KmsAuditRecord, KmsAuditSink};
use crate::backends::{ExpiredKeyRemoval, KmsBackend}; use crate::backends::{ExpiredKeyRemoval, KmsBackend};
use crate::types::{KeyStatus, ListKeysRequest}; use crate::error::Result;
use crate::types::{KeyInfo, KeyStatus, ListKeysRequest, OperationContext};
use async_trait::async_trait; use async_trait::async_trait;
use jiff::Zoned; use jiff::Zoned;
use std::sync::Arc; use std::sync::Arc;
use std::time::Duration; use std::time::{Duration, Instant};
use tokio_util::sync::CancellationToken; use tokio_util::sync::CancellationToken;
use tracing::{debug, info, warn}; use tracing::{debug, info, warn};
/// How often the worker looks for expired pending deletions. /// How often the worker looks for expired pending deletions.
pub const DEFAULT_SWEEP_INTERVAL: Duration = Duration::from_secs(60); pub const DEFAULT_SWEEP_INTERVAL: Duration = Duration::from_secs(60);
// ---------------------------------------------------------------------------
// Metrics
//
// The sweep already pages through the whole key set, so everything below is
// derived from the pages it has in hand: observing the key lifecycle costs no
// extra backend call. Each metric is an aggregate — the gauges carry no labels
// at all and the counter only a static outcome — because a per-key label would
// carry key identifiers into the metric stream and grow the series count with
// the key set. "This key is overdue for rotation" is a threshold on the
// aggregate age and belongs in an alerting rule, not in a label.
// ---------------------------------------------------------------------------
/// Gauge: keys scheduled for deletion whose deadline has not passed, as of the
/// end of the last sweep that saw the whole key set.
const METRIC_PENDING_DELETION_KEYS: &str = "rustfs_kms_pending_deletion_keys";
/// Gauge: keys left tombstoned by an interrupted removal and still awaiting
/// the sweep, as of the end of the last sweep that saw the whole key set.
const METRIC_TOMBSTONE_KEYS: &str = "rustfs_kms_deletion_tombstone_keys";
/// Gauge: seconds since the least recently rotated usable key was rotated
/// (its creation time when it was never rotated); `0` when there are none.
const METRIC_OLDEST_ROTATION_AGE_SECONDS: &str = "rustfs_kms_oldest_key_rotation_age_seconds";
/// Counter: keys the sweep acted on, by `outcome` (`removed`, `blocked`,
/// `skipped`, `failed`).
const METRIC_SWEEP_KEYS_TOTAL: &str = "rustfs_kms_deletion_sweep_keys_total";
/// Register metric descriptions once per process.
fn describe_metrics() {
static DESCRIBE: std::sync::Once = std::sync::Once::new();
DESCRIBE.call_once(|| {
metrics::describe_gauge!(
METRIC_PENDING_DELETION_KEYS,
"KMS keys scheduled for deletion whose deadline has not passed yet"
);
metrics::describe_gauge!(
METRIC_TOMBSTONE_KEYS,
"KMS keys left tombstoned by an interrupted removal, still awaiting the deletion sweep"
);
metrics::describe_gauge!(
METRIC_OLDEST_ROTATION_AGE_SECONDS,
"Seconds since the least recently rotated usable KMS key was last rotated, counting from creation for keys that were never rotated"
);
metrics::describe_counter!(METRIC_SWEEP_KEYS_TOTAL, "Total keys acted on by the KMS deletion sweep, by outcome");
});
}
/// Aggregate view of the key set, accumulated while the sweep pages through it.
///
/// Counts and one maximum age only: these feed label-less gauges, so no key
/// identifier can reach the metric stream through them.
#[derive(Debug, Default, Clone, Copy, PartialEq)]
struct KeyCensus {
pending_deletion: usize,
tombstones: usize,
/// Longest time since last rotation across keys that are still usable,
/// counting from creation for keys that were never rotated. Keys on their
/// way out are excluded: they will never be rotated again, and would
/// otherwise pin the gauge high until the sweep finishes removing them.
oldest_rotation_age_seconds: f64,
}
impl KeyCensus {
fn observe(&mut self, key: &KeyInfo, now: &Zoned) {
match key.status {
KeyStatus::PendingDeletion => self.pending_deletion += 1,
KeyStatus::Deleted => self.tombstones += 1,
KeyStatus::Active | KeyStatus::Disabled => {
let rotated_at = key.rotated_at.as_ref().unwrap_or(&key.created_at);
self.oldest_rotation_age_seconds = self.oldest_rotation_age_seconds.max(seconds_between(rotated_at, now));
}
}
}
}
/// Seconds from `earlier` to `now`, clamped at zero so clock skew (or a
/// timestamp persisted by a node running ahead) cannot produce a negative age.
fn seconds_between(earlier: &Zoned, now: &Zoned) -> f64 {
(now.timestamp().as_second() - earlier.timestamp().as_second()).max(0) as f64
}
/// Publish one sweep's counters, plus the lifecycle gauges when `census` covers
/// the whole key set.
fn record_sweep(report: &SweepReport, census: Option<KeyCensus>) {
describe_metrics();
for (outcome, count) in [
("removed", report.removed.len()),
("blocked", report.blocked.len()),
("skipped", report.skipped),
("failed", report.failed),
] {
// Emitted even at zero so every outcome series exists from the first
// sweep on and a rate over it is defined.
metrics::counter!(METRIC_SWEEP_KEYS_TOTAL, "outcome" => outcome).increment(count as u64);
}
// A sweep that could not finish listing saw only part of the key set;
// publishing its counts would understate every gauge, so the previous
// (complete) values are left standing instead.
let Some(census) = census else { return };
metrics::gauge!(METRIC_PENDING_DELETION_KEYS).set(census.pending_deletion as f64);
metrics::gauge!(METRIC_TOMBSTONE_KEYS).set(census.tombstones as f64);
metrics::gauge!(METRIC_OLDEST_ROTATION_AGE_SECONDS).set(census.oldest_rotation_age_seconds);
}
/// Reports configuration that still references a KMS key. /// Reports configuration that still references a KMS key.
/// ///
/// Consulted before any material is destroyed; a non-empty result blocks the /// Consulted before any material is destroyed; a non-empty result blocks the
@@ -68,6 +173,8 @@ pub(crate) struct DeletionWorker {
default_key_id: Option<String>, default_key_id: Option<String>,
reference_checker: Option<Arc<dyn DeletionReferenceChecker>>, reference_checker: Option<Arc<dyn DeletionReferenceChecker>>,
interval: Duration, interval: Duration,
backend_kind: &'static str,
audit_sink: Option<Arc<dyn KmsAuditSink>>,
} }
impl DeletionWorker { impl DeletionWorker {
@@ -75,15 +182,24 @@ impl DeletionWorker {
backend: Arc<dyn KmsBackend>, backend: Arc<dyn KmsBackend>,
default_key_id: Option<String>, default_key_id: Option<String>,
reference_checker: Option<Arc<dyn DeletionReferenceChecker>>, reference_checker: Option<Arc<dyn DeletionReferenceChecker>>,
backend_kind: &'static str,
) -> Self { ) -> Self {
Self { Self {
backend, backend,
default_key_id, default_key_id,
reference_checker, reference_checker,
backend_kind,
audit_sink: None,
interval: DEFAULT_SWEEP_INTERVAL, interval: DEFAULT_SWEEP_INTERVAL,
} }
} }
/// Audit each irreversible key removal to `sink`.
pub(crate) fn with_audit_sink(mut self, sink: Option<Arc<dyn KmsAuditSink>>) -> Self {
self.audit_sink = sink;
self
}
pub(crate) fn spawn(self, cancel: CancellationToken) -> tokio::task::JoinHandle<()> { pub(crate) fn spawn(self, cancel: CancellationToken) -> tokio::task::JoinHandle<()> {
tokio::spawn(async move { self.run(cancel).await }) tokio::spawn(async move { self.run(cancel).await })
} }
@@ -114,10 +230,15 @@ impl DeletionWorker {
/// Run one sweep at the given time. Exposed separately so tests can drive /// Run one sweep at the given time. Exposed separately so tests can drive
/// the expiry logic deterministically. /// the expiry logic deterministically.
///
/// Also feeds the lifecycle gauges: every key it lists is counted into a
/// [`KeyCensus`] on the way past, so the observability comes out of the
/// pages the sweep already had to fetch.
pub(crate) async fn sweep(&self, now: &Zoned) -> SweepReport { pub(crate) async fn sweep(&self, now: &Zoned) -> SweepReport {
let mut report = SweepReport::default(); let mut report = SweepReport::default();
let mut census = KeyCensus::default();
let mut marker: Option<String> = None; let mut marker: Option<String> = None;
loop { let listed_everything = loop {
let request = ListKeysRequest { let request = ListKeysRequest {
limit: Some(100), limit: Some(100),
marker: marker.clone(), marker: marker.clone(),
@@ -129,54 +250,99 @@ impl DeletionWorker {
Err(error) => { Err(error) => {
warn!(%error, "KMS deletion sweep could not list keys"); warn!(%error, "KMS deletion sweep could not list keys");
report.failed += 1; report.failed += 1;
return report; break false;
} }
}; };
for key in &response.keys { for key in &response.keys {
if matches!(key.status, KeyStatus::PendingDeletion | KeyStatus::Deleted) { // Keys this sweep destroys are left out of the census: the
self.process_key(&key.key_id, now, &mut report).await; // gauges describe the key set as it stands once the sweep is
// done, not as it was when the page was listed.
let removed = if matches!(key.status, KeyStatus::PendingDeletion | KeyStatus::Deleted) {
self.process_key(&key.key_id, now, &mut report).await
} else {
false
};
if !removed {
census.observe(key, now);
} }
} }
if !response.truncated { if !response.truncated {
break; break true;
} }
match response.next_marker { match response.next_marker {
Some(next_marker) => marker = Some(next_marker), Some(next_marker) => marker = Some(next_marker),
None => break, // Truncated without a marker: the rest of the key set is out
} // of reach, so the census is incomplete.
None => break false,
} }
};
record_sweep(&report, listed_everything.then_some(census));
report report
} }
async fn process_key(&self, key_id: &str, now: &Zoned, report: &mut SweepReport) { /// Handle one key that is on its way out; reports whether it was removed.
async fn process_key(&self, key_id: &str, now: &Zoned, report: &mut SweepReport) -> bool {
// Never remove a key that live configuration still points at. The // Never remove a key that live configuration still points at. The
// default key check is built in; broader references (bucket // default key check is built in; broader references (bucket
// encryption settings, ...) come from the injected checker. // encryption settings, ...) come from the injected checker.
if self.default_key_id.as_deref() == Some(key_id) { if self.default_key_id.as_deref() == Some(key_id) {
warn!(key_id, "expired KMS key is still the default key; refusing removal"); warn!(key_id, "expired KMS key is still the default key; refusing removal");
report.blocked.push(key_id.to_string()); report.blocked.push(key_id.to_string());
return; return false;
} }
if let Some(checker) = &self.reference_checker { if let Some(checker) = &self.reference_checker {
let references = checker.references(key_id).await; let references = checker.references(key_id).await;
if !references.is_empty() { if !references.is_empty() {
warn!(key_id, ?references, "expired KMS key is still referenced; refusing removal"); warn!(key_id, ?references, "expired KMS key is still referenced; refusing removal");
report.blocked.push(key_id.to_string()); report.blocked.push(key_id.to_string());
return; return false;
} }
} }
// The backend re-checks state and deadline under its own write // The backend re-checks state and deadline under its own write
// synchronization, so a cancellation racing this sweep wins there. // synchronization, so a cancellation racing this sweep wins there.
match self.backend.remove_expired_key(key_id, now).await { let started = Instant::now();
Ok(ExpiredKeyRemoval::Removed) => report.removed.push(key_id.to_string()), let outcome = self.backend.remove_expired_key(key_id, now).await;
Ok(ExpiredKeyRemoval::StateChanged | ExpiredKeyRemoval::NotExpired) => report.skipped += 1, match &outcome {
Ok(ExpiredKeyRemoval::Removed) => {
report.removed.push(key_id.to_string());
self.audit_removal(key_id, started, &outcome);
true
}
Ok(ExpiredKeyRemoval::StateChanged | ExpiredKeyRemoval::NotExpired) => {
report.skipped += 1;
false
}
Err(error) => { Err(error) => {
warn!(key_id, %error, "failed to remove expired KMS key; will retry next sweep"); warn!(key_id, %error, "failed to remove expired KMS key; will retry next sweep");
report.failed += 1; report.failed += 1;
self.audit_removal(key_id, started, &outcome);
false
} }
} }
} }
/// Record an attempted destruction of key material.
///
/// Only attempts that reached the backend are audited: a key the sweep
/// declined to touch (not yet due, still referenced, state changed under
/// us) was never at risk, and recording it as a deletion event would
/// drown the real ones.
fn audit_removal(&self, key_id: &str, started: Instant, outcome: &Result<ExpiredKeyRemoval>) {
let Some(sink) = self.audit_sink.as_ref() else {
return;
};
// Removal runs on a background sweep, so there is no request
// principal to attribute it to.
let context = OperationContext::internal();
sink.emit(
KmsAuditRecord::new(KmsAuditOperation::DeleteKey, &context, self.backend_kind)
.with_key_id(Some(key_id))
.with_latency(started.elapsed())
.with_result(outcome),
);
}
} }
#[cfg(test)] #[cfg(test)]
@@ -210,13 +376,14 @@ mod tests {
key_id: key_id.to_string(), key_id: key_id.to_string(),
pending_window_in_days: Some(7), pending_window_in_days: Some(7),
force_immediate: None, force_immediate: None,
confirm_key_id: None,
}) })
.await .await
.expect("deletion should be scheduled"); .expect("deletion should be scheduled");
} }
fn worker(backend: Arc<LocalKmsBackend>) -> DeletionWorker { fn worker(backend: Arc<LocalKmsBackend>) -> DeletionWorker {
DeletionWorker::new(backend, None, None) DeletionWorker::new(backend, None, None, "local")
} }
fn after_window() -> Zoned { fn after_window() -> Zoned {
@@ -307,7 +474,7 @@ mod tests {
schedule(&backend, &key_id).await; schedule(&backend, &key_id).await;
// Blocked while it is the configured default key. // Blocked while it is the configured default key.
let as_default = DeletionWorker::new(backend.clone(), Some(key_id.clone()), None); let as_default = DeletionWorker::new(backend.clone(), Some(key_id.clone()), None, "local");
let report = as_default.sweep(&after_window()).await; let report = as_default.sweep(&after_window()).await;
assert_eq!(report.blocked, vec![key_id.clone()]); assert_eq!(report.blocked, vec![key_id.clone()]);
assert!(report.removed.is_empty()); assert!(report.removed.is_empty());
@@ -317,6 +484,7 @@ mod tests {
backend.clone(), backend.clone(),
None, None,
Some(Arc::new(StaticReferences(vec!["bucket:sse-bucket".to_string()]))), Some(Arc::new(StaticReferences(vec!["bucket:sse-bucket".to_string()]))),
"local",
); );
let report = with_references.sweep(&after_window()).await; let report = with_references.sweep(&after_window()).await;
assert_eq!(report.blocked, vec![key_id.clone()]); assert_eq!(report.blocked, vec![key_id.clone()]);
@@ -327,11 +495,53 @@ mod tests {
.expect("blocked key must still exist"); .expect("blocked key must still exist");
// Removed once nothing references it anymore. // Removed once nothing references it anymore.
let unreferenced = DeletionWorker::new(backend.clone(), None, Some(Arc::new(StaticReferences(Vec::new())))); let unreferenced = DeletionWorker::new(backend.clone(), None, Some(Arc::new(StaticReferences(Vec::new()))), "local");
let report = unreferenced.sweep(&after_window()).await; let report = unreferenced.sweep(&after_window()).await;
assert_eq!(report.removed, vec![key_id.clone()]); assert_eq!(report.removed, vec![key_id.clone()]);
} }
#[tokio::test]
async fn only_attempted_removals_are_audited() {
#[derive(Default)]
struct CapturingSink {
records: std::sync::Mutex<Vec<KmsAuditRecord>>,
}
impl KmsAuditSink for CapturingSink {
fn emit(&self, record: KmsAuditRecord) {
self.records.lock().expect("sink lock should not be poisoned").push(record);
}
}
let temp_dir = tempfile::tempdir().expect("temp dir");
let backend = local_backend(&temp_dir).await;
let key_id = create_key(&backend, "audited-removal").await;
schedule(&backend, &key_id).await;
let sink = Arc::new(CapturingSink::default());
let worker = DeletionWorker::new(backend.clone(), None, None, "local").with_audit_sink(Some(sink.clone()));
// A key that is not yet due was never at risk, so it produces no record.
worker.sweep(&Zoned::now()).await;
assert!(
sink.records.lock().expect("sink lock").is_empty(),
"a key the sweep declined to touch must not be audited as a deletion"
);
worker.sweep(&after_window()).await;
let records = sink.records.lock().expect("sink lock");
assert_eq!(records.len(), 1, "the removal should be audited exactly once");
let record = &records[0];
assert_eq!(record.operation, crate::audit::KmsAuditOperation::DeleteKey);
assert_eq!(record.event, rustfs_s3_types::EventName::KmsKeyDeleted);
assert_eq!(record.outcome, crate::audit::KmsAuditOutcome::Success);
assert_eq!(record.key_id.as_deref(), Some(key_id.as_str()));
assert_eq!(record.backend, "local");
// Background work has no request principal to attribute.
assert_eq!(record.principal, OperationContext::INTERNAL_PRINCIPAL);
}
#[tokio::test] #[tokio::test]
async fn deadline_survives_backend_restart_and_sweep_completes_it() { async fn deadline_survives_backend_restart_and_sweep_completes_it() {
let temp_dir = tempfile::tempdir().expect("temp dir"); let temp_dir = tempfile::tempdir().expect("temp dir");
@@ -395,4 +605,156 @@ mod tests {
cancel.cancel(); cancel.cancel();
task.await.expect("worker task must stop after cancellation"); task.await.expect("worker task must stop after cancellation");
} }
// -- Metric emission ----------------------------------------------------
use metrics_util::MetricKind;
use metrics_util::debugging::{DebugValue, DebuggingRecorder};
type MetricEntry = (
metrics_util::CompositeKey,
Option<metrics::Unit>,
Option<metrics::SharedString>,
DebugValue,
);
/// Run `test` on a current-thread runtime under a debugging recorder and
/// return one snapshot of everything it emitted.
///
/// A single snapshot per test on purpose: `Snapshotter::snapshot` drains
/// the recorded state, so taking it per assertion would only show the
/// first assertion any data.
fn record_metrics<Out>(
test: impl FnOnce() -> std::pin::Pin<Box<dyn std::future::Future<Output = Out>>>,
) -> (Vec<MetricEntry>, Out) {
let recorder = DebuggingRecorder::new();
let snapshotter = recorder.snapshotter();
let out = metrics::with_local_recorder(&recorder, || {
let runtime = tokio::runtime::Builder::new_current_thread()
.enable_all()
.build()
.expect("current-thread runtime must build");
runtime.block_on(test())
});
(snapshotter.snapshot().into_vec(), out)
}
fn gauge_value(snapshot: &[MetricEntry], name: &str) -> Option<f64> {
snapshot.iter().find_map(|(composite, _unit, _description, value)| {
let matches = composite.kind() == MetricKind::Gauge && composite.key().name() == name;
match (matches, value) {
(true, DebugValue::Gauge(value)) => Some(value.into_inner()),
_ => None,
}
})
}
fn counter_value(snapshot: &[MetricEntry], name: &str, outcome: &str) -> u64 {
snapshot
.iter()
.filter_map(|(composite, _unit, _description, value)| {
let matches = composite.kind() == MetricKind::Counter
&& composite.key().name() == name
&& composite
.key()
.labels()
.any(|label| label.key() == "outcome" && label.value() == outcome);
match (matches, value) {
(true, DebugValue::Counter(count)) => Some(*count),
_ => None,
}
})
.sum()
}
fn key_info(key_id: &str, status: KeyStatus, created_at: Zoned, rotated_at: Option<Zoned>) -> KeyInfo {
KeyInfo {
key_id: key_id.to_string(),
description: None,
algorithm: "AES-256".to_string(),
usage: KeyUsage::EncryptDecrypt,
status,
version: 1,
metadata: std::collections::HashMap::new(),
tags: std::collections::HashMap::new(),
created_at,
rotated_at,
created_by: None,
}
}
#[test]
fn census_ages_from_rotation_and_ignores_departing_keys() {
let now = Zoned::now();
let day = Duration::from_secs(86400);
let mut census = KeyCensus::default();
// Rotation beats creation as the age baseline...
census.observe(
&key_info("rotated", KeyStatus::Active, now.clone() - 30 * day, Some(now.clone() - 2 * day)),
&now,
);
// ...and a key that was never rotated ages from its creation.
census.observe(&key_info("never-rotated", KeyStatus::Disabled, now.clone() - 5 * day, None), &now);
// Keys on their way out only ever move the counts: they will not be
// rotated again, so their age must not drive the rotation gauge.
census.observe(&key_info("pending", KeyStatus::PendingDeletion, now.clone() - 400 * day, None), &now);
census.observe(&key_info("tombstone", KeyStatus::Deleted, now.clone() - 400 * day, None), &now);
assert_eq!(census.pending_deletion, 1);
assert_eq!(census.tombstones, 1);
let expected = 5.0 * 86400.0;
assert!(
(census.oldest_rotation_age_seconds - expected).abs() < 1.0,
"expected the never-rotated key to set the age, got {}",
census.oldest_rotation_age_seconds
);
}
#[test]
fn sweep_publishes_lifecycle_gauges_without_key_labels() {
let (snapshot, key_ids) = record_metrics(|| {
Box::pin(async {
let temp_dir = tempfile::tempdir().expect("temp dir");
let backend = local_backend(&temp_dir).await;
let live = create_key(&backend, "live-key").await;
let doomed = create_key(&backend, "doomed-key").await;
schedule(&backend, &doomed).await;
// Three days in, still inside the seven-day window: the sweep
// observes the scheduled key instead of removing it.
let report = worker(backend.clone())
.sweep(&(Zoned::now() + Duration::from_secs(3 * 86400)))
.await;
assert_eq!(report.skipped, 1);
assert!(report.removed.is_empty());
vec![live, doomed]
})
});
assert_eq!(gauge_value(&snapshot, METRIC_PENDING_DELETION_KEYS), Some(1.0));
assert_eq!(gauge_value(&snapshot, METRIC_TOMBSTONE_KEYS), Some(0.0));
let age = gauge_value(&snapshot, METRIC_OLDEST_ROTATION_AGE_SECONDS).expect("rotation age gauge must be published");
// Both keys were just created, and only the usable one counts, so the
// reported age is the sweep's own offset into the future.
assert!(
(age - 3.0 * 86400.0).abs() < 60.0,
"expected the sweep offset as the rotation age, got {age}"
);
assert_eq!(counter_value(&snapshot, METRIC_SWEEP_KEYS_TOTAL, "skipped"), 1);
assert_eq!(counter_value(&snapshot, METRIC_SWEEP_KEYS_TOTAL, "removed"), 0);
for (composite, ..) in &snapshot {
for label in composite.key().labels() {
for key_id in &key_ids {
assert!(
!label.value().contains(key_id.as_str()),
"metric {} leaked a key identifier through label {}",
composite.key().name(),
label.key()
);
}
}
}
}
} }
+19
View File
@@ -137,6 +137,13 @@ pub enum KmsError {
/// fail closed instead of being sent with credentials that may lapse mid-flight /// fail closed instead of being sent with credentials that may lapse mid-flight
#[error("KMS credentials unavailable: {message}")] #[error("KMS credentials unavailable: {message}")]
CredentialsUnavailable { message: String }, CredentialsUnavailable { message: String },
/// Key has master key version records but no baseline version, so envelopes
/// written before versioned rotation can no longer be resolved
#[error(
"Baseline version lost for key {key_id}: master key version records exist (oldest {oldest_version}) but the key record carries no baseline version, so data keys written before versioned rotation can no longer be resolved to the master key version that wrapped them. A node older than versioned rotation rewrote the key record and dropped the field. Finish upgrading every node, restore baseline_version to {oldest_version} on the key record, then retry"
)]
BaselineVersionLost { key_id: String, oldest_version: u32 },
} }
impl KmsError { impl KmsError {
@@ -291,6 +298,18 @@ impl KmsError {
pub fn credentials_unavailable<S: Into<String>>(message: S) -> Self { pub fn credentials_unavailable<S: Into<String>>(message: S) -> Self {
Self::CredentialsUnavailable { message: message.into() } Self::CredentialsUnavailable { message: message.into() }
} }
/// Create a baseline version lost error
///
/// `oldest_version` is the lowest master key version that still has a
/// material record; it is exactly the baseline that was dropped, because
/// version records start at the baseline the first rotation froze.
pub fn baseline_version_lost<S: Into<String>>(key_id: S, oldest_version: u32) -> Self {
Self::BaselineVersionLost {
key_id: key_id.into(),
oldest_version,
}
}
} }
/// Convert from standard library errors /// Convert from standard library errors

Some files were not shown because too many files have changed in this diff Show More