mirror of
https://github.com/rustfs/rustfs.git
synced 2026-08-01 11:02:14 +00:00
Compare commits
27 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| 0d30b623a7 | |||
| 6900e67d32 | |||
| 428fde069d | |||
| 364168c0ba | |||
| 3f4f31129e | |||
| 6c99d4fe22 | |||
| f5348d5cc4 | |||
| e2b2bdcc34 | |||
| b965bd6eef | |||
| 1524ed891f | |||
| 7354a5663d | |||
| 790bdc0e63 | |||
| 707d062174 | |||
| 4b6b6f14bd | |||
| 9080ea8ea0 | |||
| 739efaaea1 | |||
| b6d4689c75 | |||
| 8763cd0c67 | |||
| fdac60b0e2 | |||
| 98b20b4231 | |||
| ffc9de72cb | |||
| c09d11ff3b | |||
| 792f2ef204 | |||
| 74019845c4 | |||
| 04c5921850 | |||
| db8039dece | |||
| 40eee6177a |
@@ -60,6 +60,16 @@ The file `prometheus-rules/rustfs-get-optimization-alerts.yaml` contains pre-con
|
||||
| `CodecStreamingFallbackSpike` | Warning | Codec streaming fallback > 10x baseline for 10m |
|
||||
| `IoQueueSaturation` | Warning | IO queue utilization > 90% for 5m |
|
||||
|
||||
The file `prometheus-rules/rustfs-kms-alerts.yml` contains alerting rules for the KMS backend operation metrics. Thresholds are conservative defaults pending staging baseline calibration; response procedures live in `docs/operations/kms-observability-runbook.md`, and the matching dashboard is `deploy/observability/grafana/rustfs-kms-observability.json`.
|
||||
|
||||
| Alert | Severity | Condition |
|
||||
|-------|----------|-----------|
|
||||
| `KmsBackendFatalErrors` | Critical | Fatal (non-retryable) attempt failures > 0 for 5m |
|
||||
| `KmsBackendHighErrorRate` | Critical | Non-success operation ratio > 5% for 10m (with traffic guard) |
|
||||
| `KmsBackendP99LatencyHigh` | Warning | Operation p99 duration (incl. retries) > 2s for 10m |
|
||||
| `KmsBackendAttemptFailureSpike` | Warning | Attempt failure rate > 0.5/s for 10m |
|
||||
| `KmsBackendRetryBudgetExhausted` | Warning | budget_exhausted / deadline_exceeded outcomes > 0.05/s for 10m |
|
||||
|
||||
### Enabling Alert Rules
|
||||
|
||||
Add the alert rules file to your Prometheus configuration:
|
||||
|
||||
@@ -60,6 +60,16 @@
|
||||
| `CodecStreamingFallbackSpike` | 警告 | Codec streaming 回退 > 10x 基线,持续 10 分钟 |
|
||||
| `IoQueueSaturation` | 警告 | IO 队列利用率 > 90%,持续 5 分钟 |
|
||||
|
||||
文件 `prometheus-rules/rustfs-kms-alerts.yml` 包含 KMS 后端操作指标的告警规则。阈值为保守默认值,待 staging 基线校准;响应流程见 `docs/operations/kms-observability-runbook.md`,配套仪表盘为 `deploy/observability/grafana/rustfs-kms-observability.json`。
|
||||
|
||||
| 告警 | 级别 | 条件 |
|
||||
|------|------|------|
|
||||
| `KmsBackendFatalErrors` | 严重 | fatal(不可重试)尝试失败 > 0,持续 5 分钟 |
|
||||
| `KmsBackendHighErrorRate` | 严重 | 非 success 操作占比 > 5%,持续 10 分钟(含流量下限保护) |
|
||||
| `KmsBackendP99LatencyHigh` | 警告 | 操作 p99 耗时(含重试)> 2s,持续 10 分钟 |
|
||||
| `KmsBackendAttemptFailureSpike` | 警告 | 尝试失败率 > 0.5/s,持续 10 分钟 |
|
||||
| `KmsBackendRetryBudgetExhausted` | 警告 | budget_exhausted / deadline_exceeded 结果 > 0.05/s,持续 10 分钟 |
|
||||
|
||||
### 启用告警规则
|
||||
|
||||
在 Prometheus 配置中添加告警规则文件:
|
||||
|
||||
@@ -0,0 +1,188 @@
|
||||
# Copyright 2024 RustFS Team
|
||||
#
|
||||
# Licensed under the Apache License, Version 2.0 (the "License");
|
||||
# you may not use this file except in compliance with the License.
|
||||
# You may obtain a copy of the License at
|
||||
#
|
||||
# http://www.apache.org/licenses/LICENSE-2.0
|
||||
#
|
||||
# Unless required by applicable law or agreed to in writing, software
|
||||
# distributed under the License is distributed on an "AS IS" BASIS,
|
||||
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
||||
# See the License for the specific language governing permissions and
|
||||
# limitations under the License.
|
||||
|
||||
# =============================================================================
|
||||
# RustFS KMS backend — Prometheus alerting rules
|
||||
# =============================================================================
|
||||
#
|
||||
# Metric source: the KMS operation-policy choke point in
|
||||
# crates/kms/src/policy.rs. All label values are static enum strings
|
||||
# (operation, op_class, outcome, error_class); key identifiers, key material,
|
||||
# and tokens never appear in labels.
|
||||
#
|
||||
# Response procedures: docs/operations/kms-observability-runbook.md
|
||||
#
|
||||
# IMPORTANT — threshold status: every numeric threshold below is a
|
||||
# conservative default chosen without a production baseline. Calibrate against
|
||||
# a staging baseline before relying on these alerts for paging, and prefer
|
||||
# loosening over tightening until the baseline exists. Formal SLO targets are
|
||||
# deliberately not encoded here (see rustfs/backlog#1584).
|
||||
#
|
||||
# NOTE: prometheus.yml loads /etc/prometheus/rules/*.yml — keep the .yml
|
||||
# extension or the file is silently ignored by the docker-compose stack.
|
||||
#
|
||||
# Validate: promtool check rules rustfs-kms-alerts.yml
|
||||
# =============================================================================
|
||||
|
||||
groups:
|
||||
# ==========================================================================
|
||||
# Critical alerts — immediate action required
|
||||
# ==========================================================================
|
||||
- name: rustfs-kms-critical
|
||||
interval: 30s
|
||||
rules:
|
||||
# ------------------------------------------------------------------
|
||||
# 1. KmsBackendFatalErrors
|
||||
# Any attempt failure classified as fatal (non-retryable): auth
|
||||
# or permission errors, malformed requests, missing keys. The
|
||||
# policy never retries these, so even a low rate means real
|
||||
# operations are failing right now.
|
||||
# ------------------------------------------------------------------
|
||||
- alert: KmsBackendFatalErrors
|
||||
expr: |
|
||||
sum by (operation) (rate(rustfs_kms_backend_attempt_failures_total{error_class="fatal"}[5m])) > 0
|
||||
for: 5m
|
||||
labels:
|
||||
severity: critical
|
||||
component: kms
|
||||
annotations:
|
||||
summary: "KMS backend fatal errors on operation {{ $labels.operation }}"
|
||||
description: >-
|
||||
Attempt failures classified as fatal are occurring at
|
||||
{{ $value | printf "%.3f" }}/s on operation
|
||||
{{ $labels.operation }}. Fatal failures are not retried:
|
||||
each one is a KMS backend call that failed permanently
|
||||
(authentication, permissions, malformed request, or a
|
||||
missing key/version).
|
||||
runbook_url: "https://github.com/rustfs/rustfs/blob/main/docs/operations/kms-observability-runbook.md#kmsbackendfatalerrors"
|
||||
|
||||
# ------------------------------------------------------------------
|
||||
# 2. KmsBackendHighErrorRate
|
||||
# Sustained share of operations terminating without success
|
||||
# (fatal, budget_exhausted, deadline_exceeded). The cancelled
|
||||
# outcome is excluded because shutdowns legitimately produce it.
|
||||
# The traffic guard keeps a single failure on a near-idle
|
||||
# cluster from firing the alert.
|
||||
# Threshold: 5% for 10m — conservative default, calibrate
|
||||
# against a staging baseline.
|
||||
# ------------------------------------------------------------------
|
||||
- alert: KmsBackendHighErrorRate
|
||||
expr: |
|
||||
(
|
||||
sum(rate(rustfs_kms_backend_operations_total{outcome!~"success|cancelled"}[5m]))
|
||||
/
|
||||
clamp_min(sum(rate(rustfs_kms_backend_operations_total[5m])), 1e-9)
|
||||
) > 0.05
|
||||
and
|
||||
sum(rate(rustfs_kms_backend_operations_total[5m])) > 0.02
|
||||
for: 10m
|
||||
labels:
|
||||
severity: critical
|
||||
component: kms
|
||||
annotations:
|
||||
summary: "KMS backend non-success ratio above 5% for 10m"
|
||||
description: >-
|
||||
{{ $value | humanizePercentage }} of KMS backend operations
|
||||
are terminating in fatal, budget_exhausted, or
|
||||
deadline_exceeded. Object encryption and decryption paths
|
||||
depending on the KMS are degraded or failing.
|
||||
runbook_url: "https://github.com/rustfs/rustfs/blob/main/docs/operations/kms-observability-runbook.md#kmsbackendhigherrorrate"
|
||||
|
||||
# ==========================================================================
|
||||
# Warning alerts — investigation needed
|
||||
# ==========================================================================
|
||||
- name: rustfs-kms-warning
|
||||
interval: 30s
|
||||
rules:
|
||||
# ------------------------------------------------------------------
|
||||
# 3. KmsBackendP99LatencyHigh
|
||||
# p99 wall-clock duration of whole operations (attempts plus
|
||||
# backoff) is sustained above 2 seconds. Because the histogram
|
||||
# includes retries, a high p99 usually means the retry policy
|
||||
# is absorbing backend failures, not that every call is slow.
|
||||
# Threshold: 2s for 10m — conservative default, calibrate
|
||||
# against a staging baseline.
|
||||
# ------------------------------------------------------------------
|
||||
- alert: KmsBackendP99LatencyHigh
|
||||
expr: |
|
||||
histogram_quantile(0.99,
|
||||
sum by (le) (rate(rustfs_kms_backend_operation_duration_seconds_bucket[5m]))
|
||||
) > 2
|
||||
for: 10m
|
||||
labels:
|
||||
severity: warning
|
||||
component: kms
|
||||
annotations:
|
||||
summary: "KMS backend operation p99 latency above 2s for 10m"
|
||||
description: >-
|
||||
The 99th-percentile KMS backend operation duration is
|
||||
{{ $value | humanizeDuration }}, including retries and
|
||||
backoff. Encryption and decryption latency is leaking into
|
||||
S3 request latency.
|
||||
runbook_url: "https://github.com/rustfs/rustfs/blob/main/docs/operations/kms-observability-runbook.md#kmsbackendp99latencyhigh"
|
||||
|
||||
# ------------------------------------------------------------------
|
||||
# 4. KmsBackendAttemptFailureSpike
|
||||
# Aggregate attempt-failure rate (all error classes) sustained
|
||||
# above an absolute floor. An absolute threshold is used instead
|
||||
# of an offset-1d baseline ratio because fresh deployments have
|
||||
# no baseline and an empty offset vector would keep a ratio
|
||||
# alert from ever firing; switch to a baseline-relative form
|
||||
# (see rustfs-get-optimization-alerts.yaml for the pattern)
|
||||
# once a stable staging baseline exists.
|
||||
# Threshold: 0.5/s for 10m — conservative default, calibrate
|
||||
# against a staging baseline.
|
||||
# ------------------------------------------------------------------
|
||||
- alert: KmsBackendAttemptFailureSpike
|
||||
expr: |
|
||||
sum(rate(rustfs_kms_backend_attempt_failures_total[5m])) > 0.5
|
||||
for: 10m
|
||||
labels:
|
||||
severity: warning
|
||||
component: kms
|
||||
annotations:
|
||||
summary: "KMS backend attempt failures above 0.5/s for 10m"
|
||||
description: >-
|
||||
KMS backend attempts are failing at
|
||||
{{ $value | printf "%.2f" }}/s across all error classes.
|
||||
The retry policy may still be masking these from callers —
|
||||
check the error-class breakdown before it stops absorbing
|
||||
them.
|
||||
runbook_url: "https://github.com/rustfs/rustfs/blob/main/docs/operations/kms-observability-runbook.md#kmsbackendattemptfailurespike"
|
||||
|
||||
# ------------------------------------------------------------------
|
||||
# 5. KmsBackendRetryBudgetExhausted
|
||||
# Operations are running out of retry budget (budget_exhausted)
|
||||
# or operation deadline (deadline_exceeded). These surface to
|
||||
# callers as failed KMS operations even though every individual
|
||||
# failure was retryable — the backend is unhealthy for longer
|
||||
# than the policy can bridge.
|
||||
# Threshold: 0.05/s for 10m — conservative default, calibrate
|
||||
# against a staging baseline.
|
||||
# ------------------------------------------------------------------
|
||||
- alert: KmsBackendRetryBudgetExhausted
|
||||
expr: |
|
||||
sum by (outcome) (rate(rustfs_kms_backend_operations_total{outcome=~"budget_exhausted|deadline_exceeded"}[5m])) > 0.05
|
||||
for: 10m
|
||||
labels:
|
||||
severity: warning
|
||||
component: kms
|
||||
annotations:
|
||||
summary: "KMS backend operations exhausting retry budget ({{ $labels.outcome }})"
|
||||
description: >-
|
||||
KMS backend operations are terminating as
|
||||
{{ $labels.outcome }} at {{ $value | printf "%.3f" }}/s.
|
||||
Retryable failures are outlasting the retry budget, so
|
||||
callers are seeing hard failures.
|
||||
runbook_url: "https://github.com/rustfs/rustfs/blob/main/docs/operations/kms-observability-runbook.md#kmsbackendretrybudgetexhausted"
|
||||
@@ -25,9 +25,13 @@ inputs:
|
||||
required: false
|
||||
default: "rustfs-deps"
|
||||
cache-save-if:
|
||||
description: "Condition for saving cache"
|
||||
description: >-
|
||||
Whether to save the cache. The fail-safe default is 'false': a caller that
|
||||
wants to populate a cache must opt in explicitly, so a forgotten input
|
||||
costs a cold cache (minutes) rather than silently consuming the
|
||||
repository-wide 10GB Actions cache quota and evicting other lanes.
|
||||
required: false
|
||||
default: "true"
|
||||
default: "false"
|
||||
install-cross-tools:
|
||||
description: "Install cross-compilation tools"
|
||||
required: false
|
||||
@@ -36,28 +40,43 @@ inputs:
|
||||
description: "Target architecture to add"
|
||||
required: false
|
||||
default: ""
|
||||
github-token:
|
||||
description: "GitHub token for API access"
|
||||
install-build-packaging-tools:
|
||||
description: >-
|
||||
Install musl-tools/zip/unzip, needed for musl linking and release
|
||||
packaging. Off for CI test lanes, which use none of them.
|
||||
required: false
|
||||
default: ""
|
||||
default: "true"
|
||||
install-test-tools:
|
||||
description: >-
|
||||
Install cargo-nextest and the rustfmt/clippy components. Off for release
|
||||
and audit lanes, which run no tests and no lints.
|
||||
required: false
|
||||
default: "true"
|
||||
|
||||
runs:
|
||||
using: "composite"
|
||||
steps:
|
||||
# protobuf-compiler is deliberately absent: the setup-protoc step below
|
||||
# installs 34.1 into the tool cache and prepends it to PATH, so the apt
|
||||
# build (older, and never version-matched) was shadowed on every run and
|
||||
# simply never used.
|
||||
- name: Install system dependencies (Ubuntu)
|
||||
if: runner.os == 'Linux'
|
||||
shell: bash
|
||||
run: |
|
||||
sudo apt-get update
|
||||
sudo apt-get install -y \
|
||||
musl-tools \
|
||||
build-essential \
|
||||
pkg-config \
|
||||
libssl-dev \
|
||||
ripgrep \
|
||||
unzip \
|
||||
zip \
|
||||
protobuf-compiler
|
||||
ripgrep
|
||||
|
||||
# musl-gcc is needed by the native musl release leg, and zip/unzip by the
|
||||
# release packaging steps. No CI test lane touches any of them.
|
||||
- name: Install packaging and cross-linking dependencies (Ubuntu)
|
||||
if: runner.os == 'Linux' && inputs.install-build-packaging-tools == 'true'
|
||||
shell: bash
|
||||
run: sudo apt-get install -y musl-tools zip unzip
|
||||
|
||||
- name: Install protoc
|
||||
uses: rustfs/setup-protoc@a3705324d8f9bf5b6c3573fb6cf8ae421db55dd6 # v3.0.1
|
||||
@@ -75,7 +94,7 @@ runs:
|
||||
with:
|
||||
toolchain: ${{ inputs.rust-version }}
|
||||
targets: ${{ inputs.target }}
|
||||
components: rustfmt, clippy
|
||||
components: ${{ inputs.install-test-tools == 'true' && 'rustfmt, clippy' || '' }}
|
||||
|
||||
- name: Install Zig
|
||||
if: inputs.install-cross-tools == 'true'
|
||||
@@ -86,6 +105,7 @@ runs:
|
||||
uses: taiki-e/install-action@a21ae4029b089b9ddc45704028756f51ab8abe48 # cargo-zigbuild
|
||||
|
||||
- name: Install cargo-nextest
|
||||
if: inputs.install-test-tools == 'true'
|
||||
uses: taiki-e/install-action@96c7780c1d8a2b8723e12031def873a434d39d8d # nextest
|
||||
|
||||
- name: Setup Rust cache
|
||||
|
||||
@@ -37,6 +37,7 @@ jobs:
|
||||
name: Cancel Closed PR Runs
|
||||
if: github.event_name == 'pull_request' && github.event.action == 'closed'
|
||||
runs-on: ubuntu-latest
|
||||
timeout-minutes: 10
|
||||
steps:
|
||||
- name: Explain cancellation run
|
||||
run: echo "PR closed; this run only cancels older runs in the same concurrency group."
|
||||
@@ -45,6 +46,7 @@ jobs:
|
||||
name: Architecture Migration Rules
|
||||
if: github.event_name != 'pull_request' || github.event.action != 'closed'
|
||||
runs-on: ubuntu-latest
|
||||
timeout-minutes: 10
|
||||
steps:
|
||||
- uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
|
||||
|
||||
|
||||
@@ -39,7 +39,12 @@ on:
|
||||
- 'scripts/security/check_preview_release_workflow.sh'
|
||||
- 'scripts/security/check_workflow_pins.sh'
|
||||
schedule:
|
||||
- cron: '0 3 * * 0' # Weekly on Sunday 03:00 UTC (staggered after the midnight ci/build crons)
|
||||
# Daily, not weekly. This schedule exists to catch RustSec advisories
|
||||
# published against an unchanged dependency tree; at weekly cadence a new
|
||||
# advisory could sit unnoticed for seven days. The check list is unchanged —
|
||||
# splitting it into a light daily advisories-only run and a weekly full run
|
||||
# would create runs where sources/bans/licenses go unverified.
|
||||
- cron: '0 3 * * *' # Daily 03:00 UTC (staggered after the midnight ci/build crons)
|
||||
workflow_dispatch:
|
||||
|
||||
permissions:
|
||||
@@ -59,6 +64,7 @@ jobs:
|
||||
name: Cancel Closed PR Runs
|
||||
if: github.event_name == 'pull_request' && github.event.action == 'closed'
|
||||
runs-on: ubuntu-latest
|
||||
timeout-minutes: 10
|
||||
steps:
|
||||
- name: Explain cancellation run
|
||||
run: echo "PR closed; this run only cancels older runs in the same concurrency group."
|
||||
@@ -75,10 +81,27 @@ jobs:
|
||||
- name: Checkout repository
|
||||
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
|
||||
|
||||
- name: Setup Rust environment
|
||||
uses: ./.github/actions/setup
|
||||
# cargo-deny compiles nothing, so the full setup composite (apt packages,
|
||||
# protoc, flatc, nextest, rustfmt/clippy) was pure overhead here. It does
|
||||
# still need a real cargo: `cargo deny check` runs `cargo metadata`, and
|
||||
# Cargo.toml pins datafusion and s3s as git dependencies, which must be
|
||||
# materialised into ~/.cargo/git — a cold clone is hundreds of MB, so the
|
||||
# cache stays.
|
||||
- name: Install Rust toolchain
|
||||
uses: dtolnay/rust-toolchain@29eef336d9b2848a0b548edc03f92a220660cdb8 # stable
|
||||
|
||||
# Was relying on the composite's default, which used to be "true": every
|
||||
# PR touching Cargo.toml/Cargo.lock saved a second, PR-scoped copy of this
|
||||
# cache and pushed the main-scoped lanes out of the 10GB quota. The
|
||||
# default is now "false", but state it explicitly — see
|
||||
# scripts/security/check_cache_save_if.sh.
|
||||
- name: Setup Rust cache
|
||||
uses: Swatinem/rust-cache@e18b497796c12c097a38f9edb9d0641fb99eee32 # v2
|
||||
with:
|
||||
cache-shared-key: rustfs-cargo-deny
|
||||
cache-all-crates: true
|
||||
cache-on-failure: true
|
||||
shared-key: rustfs-cargo-deny
|
||||
save-if: ${{ github.ref == 'refs/heads/main' }}
|
||||
|
||||
- name: Install cargo-deny
|
||||
uses: taiki-e/install-action@bffeee26d4db9be238a4ea78d8826604ebcb594d # v2
|
||||
@@ -100,12 +123,19 @@ jobs:
|
||||
- name: Report unpinned GitHub Actions
|
||||
run: ./scripts/security/check_workflow_pins.sh --enforce
|
||||
|
||||
- name: Check setup cache-save-if is explicit
|
||||
run: ./scripts/security/check_cache_save_if.sh
|
||||
|
||||
- name: Check every job declares a timeout
|
||||
run: ./scripts/security/check_job_timeouts.sh
|
||||
|
||||
- name: Check preview release workflow policy
|
||||
run: ./scripts/security/check_preview_release_workflow.sh
|
||||
|
||||
dependency-review:
|
||||
name: Dependency Review
|
||||
runs-on: ubuntu-latest
|
||||
timeout-minutes: 30
|
||||
if: github.event_name == 'pull_request' && github.event.action != 'closed'
|
||||
permissions:
|
||||
contents: read
|
||||
@@ -125,3 +155,26 @@ jobs:
|
||||
# conscious re-review of the license/provenance claim (backlog#1181).
|
||||
allow-dependencies-licenses: pkg:cargo/rustfs-uring@0.1.0
|
||||
comment-summary-in-pr: always
|
||||
|
||||
alert-on-failure:
|
||||
name: Alert on scheduled failure
|
||||
# dependency-review is deliberately excluded: it only runs on pull_request,
|
||||
# so it can never contribute a failure to a scheduled run.
|
||||
needs: [cargo-deny, workflow-pin-report]
|
||||
# A scheduled cargo-deny failure usually means the dependency tree just
|
||||
# matched a newly published advisory — the single most important signal this
|
||||
# workflow produces, and until now it was only visible to whoever happened to
|
||||
# open the Actions tab. Same ci-8 mechanism coverage.yml and
|
||||
# e2e-replication-nightly.yml already use.
|
||||
if: always() && github.event_name == 'schedule' && contains(needs.*.result, 'failure')
|
||||
runs-on: ubuntu-latest
|
||||
timeout-minutes: 10
|
||||
permissions:
|
||||
contents: read
|
||||
issues: write
|
||||
steps:
|
||||
- uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
|
||||
- name: Open or update failure-tracking issue
|
||||
uses: ./.github/actions/schedule-failure-issue
|
||||
with:
|
||||
github-token: ${{ secrets.GITHUB_TOKEN }}
|
||||
|
||||
@@ -50,12 +50,18 @@ on:
|
||||
- "**/*.svg"
|
||||
- ".gitignore"
|
||||
- ".dockerignore"
|
||||
- "flake.lock"
|
||||
schedule:
|
||||
- cron: "0 1 * * 0" # Weekly on Sunday 01:00 UTC (staggered after the ci.yml midnight cron)
|
||||
workflow_dispatch:
|
||||
inputs:
|
||||
build_docker:
|
||||
description: "Build and push Docker images after binary build"
|
||||
# Advisory only. docker.yml triggers on workflow_run and its job-level
|
||||
# condition requires the triggering event to be a tag push, so a manual
|
||||
# dispatch of this workflow never produces images regardless of this
|
||||
# value. Kept because the summary step reports it; wiring it up would
|
||||
# mean teaching docker.yml's version parser a second event shape.
|
||||
description: "Build and push Docker images after binary build (ignored: dispatch runs never reach docker.yml)"
|
||||
required: false
|
||||
default: true
|
||||
type: boolean
|
||||
@@ -83,6 +89,7 @@ jobs:
|
||||
build-check:
|
||||
name: Build Strategy Check
|
||||
runs-on: ubuntu-latest
|
||||
timeout-minutes: 30
|
||||
outputs:
|
||||
should_build: ${{ steps.check.outputs.should_build }}
|
||||
build_type: ${{ steps.check.outputs.build_type }}
|
||||
@@ -164,6 +171,7 @@ jobs:
|
||||
name: Prepare Platform Matrix
|
||||
needs: build-check
|
||||
runs-on: ubuntu-latest
|
||||
timeout-minutes: 30
|
||||
outputs:
|
||||
matrix: ${{ steps.select.outputs.matrix }}
|
||||
selected: ${{ steps.select.outputs.selected }}
|
||||
@@ -171,10 +179,14 @@ jobs:
|
||||
- name: Select target platforms
|
||||
id: select
|
||||
shell: bash
|
||||
env:
|
||||
# via env, not interpolation: a dispatch input is free-form text and
|
||||
# would otherwise be pasted into the script for bash to evaluate.
|
||||
RAW_PLATFORMS: ${{ github.event_name == 'workflow_dispatch' && github.event.inputs.platforms || 'all' }}
|
||||
run: |
|
||||
set -euo pipefail
|
||||
|
||||
selected="${{ github.event_name == 'workflow_dispatch' && github.event.inputs.platforms || 'all' }}"
|
||||
selected="$RAW_PLATFORMS"
|
||||
selected="$(echo "${selected}" | tr -d '[:space:]')"
|
||||
if [[ -z "${selected}" ]]; then
|
||||
selected="all"
|
||||
@@ -253,9 +265,17 @@ jobs:
|
||||
rust-version: stable
|
||||
target: ${{ matrix.target }}
|
||||
cache-shared-key: build-${{ matrix.target }}
|
||||
github-token: ${{ secrets.GITHUB_TOKEN }}
|
||||
cache-save-if: ${{ github.ref == 'refs/heads/main' || startsWith(github.ref, 'refs/tags/') }}
|
||||
# main only. A cache saved on refs/tags/X is scoped to that tag: no
|
||||
# other tag, no main run and no PR can restore it, so every release
|
||||
# cycle wrote up to 12 entries of 1-2GB (preview tag plus final tag,
|
||||
# six legs each) that nobody could read, evicting the hot lanes from
|
||||
# the repo-wide 10GB quota. Tag builds still restore the main-scoped
|
||||
# cache, since default-branch caches are readable from every ref.
|
||||
# The one real cost: re-running a failed leg of the same tag no longer
|
||||
# finds that tag's own warm cache and falls back to main's.
|
||||
cache-save-if: ${{ github.ref == 'refs/heads/main' }}
|
||||
install-cross-tools: ${{ matrix.cross }}
|
||||
install-test-tools: 'false'
|
||||
|
||||
- name: Download static console assets
|
||||
shell: bash
|
||||
@@ -702,9 +722,14 @@ jobs:
|
||||
needs: [ build-check, build-rustfs ]
|
||||
if: always() && needs.build-check.outputs.should_build == 'true'
|
||||
runs-on: ubuntu-latest
|
||||
timeout-minutes: 10
|
||||
steps:
|
||||
- name: Build completion summary
|
||||
shell: bash
|
||||
env:
|
||||
# dispatch input via env: free-form text must not be pasted into the
|
||||
# script for bash to evaluate.
|
||||
INPUT_BUILD_DOCKER: ${{ github.event.inputs.build_docker }}
|
||||
run: |
|
||||
BUILD_TYPE="${{ needs.build-check.outputs.build_type }}"
|
||||
VERSION="${{ needs.build-check.outputs.version }}"
|
||||
@@ -746,7 +771,7 @@ jobs:
|
||||
echo "🐳 Docker Images:"
|
||||
if [[ "$BUILD_TYPE" == "preview" ]]; then
|
||||
echo "⏭️ Preview tags do not publish Docker images"
|
||||
elif [[ "${{ github.event.inputs.build_docker }}" == "false" ]]; then
|
||||
elif [[ "$INPUT_BUILD_DOCKER" == "false" ]]; then
|
||||
echo "⏭️ Docker image build was skipped (binary only build)"
|
||||
elif [[ "$BUILD_STATUS" == "success" ]]; then
|
||||
echo "🔄 Docker images will be built and pushed automatically via workflow_run event"
|
||||
@@ -760,6 +785,7 @@ jobs:
|
||||
needs: [ build-check, build-rustfs ]
|
||||
if: startsWith(github.ref, 'refs/tags/') && (needs.build-check.outputs.build_type == 'preview' || needs.build-check.outputs.build_type == 'release' || needs.build-check.outputs.build_type == 'prerelease')
|
||||
runs-on: ubuntu-latest
|
||||
timeout-minutes: 30
|
||||
permissions:
|
||||
contents: write
|
||||
outputs:
|
||||
@@ -819,6 +845,7 @@ jobs:
|
||||
needs: [ build-check, build-rustfs, create-release ]
|
||||
if: startsWith(github.ref, 'refs/tags/') && (needs.build-check.outputs.build_type == 'preview' || needs.build-check.outputs.build_type == 'release' || needs.build-check.outputs.build_type == 'prerelease')
|
||||
runs-on: ubuntu-latest
|
||||
timeout-minutes: 30
|
||||
permissions:
|
||||
contents: write
|
||||
actions: read
|
||||
@@ -910,6 +937,7 @@ jobs:
|
||||
needs: [ build-check, publish-release ]
|
||||
if: startsWith(github.ref, 'refs/tags/') && (needs.build-check.outputs.build_type == 'release' || needs.build-check.outputs.build_type == 'prerelease')
|
||||
runs-on: ubuntu-latest
|
||||
timeout-minutes: 30
|
||||
steps:
|
||||
- name: Update latest.json
|
||||
env:
|
||||
@@ -969,6 +997,7 @@ jobs:
|
||||
needs: [ build-check, create-release, upload-release-assets ]
|
||||
if: startsWith(github.ref, 'refs/tags/') && (needs.build-check.outputs.build_type == 'preview' || needs.build-check.outputs.build_type == 'release' || needs.build-check.outputs.build_type == 'prerelease')
|
||||
runs-on: ubuntu-latest
|
||||
timeout-minutes: 30
|
||||
permissions:
|
||||
contents: write
|
||||
steps:
|
||||
|
||||
@@ -0,0 +1,196 @@
|
||||
# Copyright 2026 RustFS Team
|
||||
#
|
||||
# Licensed under the Apache License, Version 2.0 (the "License");
|
||||
# you may not use this file except in compliance with the License.
|
||||
# You may obtain a copy of the License at
|
||||
#
|
||||
# http://www.apache.org/licenses/LICENSE-2.0
|
||||
#
|
||||
# Unless required by applicable law or agreed to in writing, software
|
||||
# distributed under the License is distributed on an "AS IS" BASIS,
|
||||
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
||||
# See the License for the specific language governing permissions and
|
||||
# limitations under the License.
|
||||
|
||||
# Sole writer of the Rust dependency caches that ci.yml restores.
|
||||
#
|
||||
# Why this is a separate workflow rather than steps inside ci.yml: ci.yml's
|
||||
# concurrency group cancels in-progress runs on main pushes, and merges land far
|
||||
# faster than its 70-minute pipeline. Measured over 15 consecutive main pushes:
|
||||
# 12 cancelled, 2 failed, 0 succeeded. A cancelled run never reaches
|
||||
# Swatinem/rust-cache's post step (cache-on-failure does not cover cancellation),
|
||||
# so the writer lanes were saving nothing and every PR paid a cold restore —
|
||||
# 11.8-20.9 minutes of "Setup Rust environment" against 0.7-3.4 warm.
|
||||
#
|
||||
# Splitting cache writing out of the test pipeline lets ci.yml keep cancelling
|
||||
# superseded runs (which is correct — nobody needs test results for a commit
|
||||
# that is already three merges behind) while the caches still get written.
|
||||
#
|
||||
# The group below deliberately does NOT cancel in progress. GitHub keeps at most
|
||||
# one running plus one pending run per group, so a burst of merges collapses
|
||||
# into "current run finishes, newest queued run follows" rather than a pile-up.
|
||||
# That also bounds this workflow to one self-hosted runner at a time.
|
||||
#
|
||||
# Each job below owns exactly one shared-key and is the only place that sets
|
||||
# cache-save-if to anything but 'false' for it; every lane in ci.yml reads.
|
||||
# scripts/security/check_cache_save_if.sh keeps the declarations explicit.
|
||||
#
|
||||
# The builds are supersets of what the reading lanes compile, because a reader
|
||||
# restores only what the writer saved. Feature resolution matters here: a lane
|
||||
# built with e2e-test-hooks resolves dependency features differently, which
|
||||
# changes -Cmetadata, so the plain build does not cover it. See
|
||||
# rustfs/backlog#1600.
|
||||
|
||||
name: Cache Warm
|
||||
|
||||
on:
|
||||
push:
|
||||
branches: [ main ]
|
||||
# Mirrors ci.yml's push paths-ignore: if a commit cannot change what ci.yml
|
||||
# compiles, it cannot change what ci.yml needs restored either.
|
||||
paths-ignore:
|
||||
- "**.md"
|
||||
- "docs/**"
|
||||
- "deploy/**"
|
||||
- "scripts/dev_*.sh"
|
||||
- "scripts/probe.sh"
|
||||
- "LICENSE*"
|
||||
- ".gitignore"
|
||||
- ".dockerignore"
|
||||
- "README*"
|
||||
- "**/*.png"
|
||||
- "**/*.jpg"
|
||||
- "**/*.svg"
|
||||
- ".github/workflows/build.yml"
|
||||
- ".github/workflows/docker.yml"
|
||||
- ".github/workflows/audit.yml"
|
||||
workflow_dispatch:
|
||||
|
||||
permissions:
|
||||
contents: read
|
||||
|
||||
concurrency:
|
||||
group: cache-warm
|
||||
cancel-in-progress: false
|
||||
|
||||
env:
|
||||
CARGO_TERM_COLOR: always
|
||||
|
||||
jobs:
|
||||
# Readers: test-and-lint, test-ilm-integration-serial, build-rustfs-debug-binary,
|
||||
# e2e-tests, e2e-full.
|
||||
warm-ci-dev:
|
||||
name: Warm ci-dev
|
||||
runs-on: sm-standard-4
|
||||
timeout-minutes: 90
|
||||
env:
|
||||
FORCE_JAVASCRIPT_ACTIONS_TO_NODE24: "true"
|
||||
steps:
|
||||
- name: Checkout repository
|
||||
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
|
||||
with:
|
||||
persist-credentials: false
|
||||
|
||||
- name: Setup Rust environment
|
||||
uses: ./.github/actions/setup
|
||||
with:
|
||||
rust-version: stable
|
||||
cache-shared-key: ci-dev
|
||||
cache-save-if: 'true'
|
||||
install-build-packaging-tools: 'false'
|
||||
|
||||
# --all-targets covers the test binaries nextest builds, including
|
||||
# e2e_test, which test-and-lint's own run excludes. The second build adds
|
||||
# the e2e-test-hooks feature resolution that build-rustfs-debug-binary uses
|
||||
# and that no lint lane enables.
|
||||
- name: Build ci-dev superset
|
||||
env:
|
||||
# Same limit ci.yml puts on its nextest step: this builds the same
|
||||
# ~100 workspace test binaries, and three concurrent links saturate the
|
||||
# self-hosted runner's overlay I/O and can wedge Cargo (#5394).
|
||||
CARGO_BUILD_JOBS: "2"
|
||||
run: |
|
||||
cargo build --workspace --all-targets
|
||||
cargo build -p rustfs --bins --features e2e-test-hooks
|
||||
|
||||
# Readers: test-and-lint-rio-v2, build-rustfs-debug-binary-rio-v2.
|
||||
warm-ci-feat-rio:
|
||||
name: Warm ci-feat-rio
|
||||
runs-on: sm-standard-4
|
||||
timeout-minutes: 90
|
||||
env:
|
||||
FORCE_JAVASCRIPT_ACTIONS_TO_NODE24: "true"
|
||||
steps:
|
||||
- name: Checkout repository
|
||||
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
|
||||
with:
|
||||
persist-credentials: false
|
||||
|
||||
- name: Setup Rust environment
|
||||
uses: ./.github/actions/setup
|
||||
with:
|
||||
rust-version: stable
|
||||
cache-shared-key: ci-feat-rio
|
||||
cache-save-if: 'true'
|
||||
install-build-packaging-tools: 'false'
|
||||
|
||||
- name: Build ci-feat-rio superset
|
||||
run: |
|
||||
cargo build -p rustfs -p rustfs-ecstore --all-targets --features rio-v2
|
||||
cargo build -p rustfs --bins --features rio-v2,e2e-test-hooks
|
||||
|
||||
# Readers: the swift and sftp legs of test-and-lint-protocols. Built in
|
||||
# sequence rather than as `--features swift,sftp`, which is a combination no
|
||||
# lane actually compiles; running both leaves the union in target/.
|
||||
warm-ci-feat-proto:
|
||||
name: Warm ci-feat-proto
|
||||
runs-on: sm-standard-4
|
||||
timeout-minutes: 90
|
||||
env:
|
||||
FORCE_JAVASCRIPT_ACTIONS_TO_NODE24: "true"
|
||||
steps:
|
||||
- name: Checkout repository
|
||||
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
|
||||
with:
|
||||
persist-credentials: false
|
||||
|
||||
- name: Setup Rust environment
|
||||
uses: ./.github/actions/setup
|
||||
with:
|
||||
rust-version: stable
|
||||
cache-shared-key: ci-feat-proto
|
||||
cache-save-if: 'true'
|
||||
install-build-packaging-tools: 'false'
|
||||
|
||||
- name: Build ci-feat-proto superset
|
||||
run: |
|
||||
cargo build -p rustfs -p rustfs-protocols --all-targets --features swift
|
||||
cargo build -p rustfs -p rustfs-protocols --all-targets --features sftp
|
||||
|
||||
# Reader: uring-integration. Runs on ubuntu-latest to match it: rust-cache's
|
||||
# key covers runner.os and arch but not the runner label or image, so a cache
|
||||
# written on sm-standard-4 would be restored by the hosted runner as if it
|
||||
# belonged to it.
|
||||
warm-ci-uring:
|
||||
name: Warm ci-uring
|
||||
runs-on: ubuntu-latest
|
||||
timeout-minutes: 60
|
||||
steps:
|
||||
- name: Checkout repository
|
||||
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
|
||||
with:
|
||||
persist-credentials: false
|
||||
|
||||
- name: Setup Rust environment
|
||||
uses: ./.github/actions/setup
|
||||
with:
|
||||
rust-version: stable
|
||||
cache-shared-key: ci-uring
|
||||
cache-save-if: 'true'
|
||||
install-build-packaging-tools: 'false'
|
||||
|
||||
- name: Install build dependencies
|
||||
run: sudo apt-get update && sudo apt-get install -y protobuf-compiler
|
||||
|
||||
- name: Build ci-uring superset
|
||||
run: cargo build -p rustfs-ecstore --all-targets
|
||||
@@ -12,18 +12,24 @@
|
||||
# See the License for the specific language governing permissions and
|
||||
# limitations under the License.
|
||||
|
||||
# Companion to ci.yml for the required "Test and Lint" status check.
|
||||
# Companion to ci.yml for required status checks.
|
||||
#
|
||||
# ci.yml skips docs-only pull requests via paths-ignore, but the branch
|
||||
# ruleset requires a check named "Test and Lint" — without this workflow a
|
||||
# docs-only PR would wait on that check forever. This workflow triggers on
|
||||
# exactly the paths ci.yml ignores and reports an instant success under the
|
||||
# same job name. Mixed PRs trigger both workflows and the real check still
|
||||
# gates: a required check with any failing run blocks the merge.
|
||||
# ci.yml skips docs-only pull requests via paths-ignore, but the branch ruleset
|
||||
# requires a check named "Test and Lint" — without this workflow a docs-only PR
|
||||
# would wait on it forever. This workflow triggers on exactly the paths ci.yml
|
||||
# ignores and reports success under the same job name. Mixed PRs trigger both
|
||||
# workflows and the real check still gates: a required check with any failing
|
||||
# run blocks the merge.
|
||||
# https://docs.github.com/en/repositories/configuring-branches-and-merges-in-your-repository/defining-the-mergeability-of-pull-requests/troubleshooting-required-status-checks#handling-skipped-but-required-checks
|
||||
#
|
||||
# "Quick Checks" is mirrored here ahead of the ruleset change that will make it
|
||||
# required too (rustfs/backlog#1599). Until that change lands this job is
|
||||
# inert; mirroring it first is what lets the ruleset change happen without
|
||||
# stranding docs-only PRs on a check nobody reports.
|
||||
#
|
||||
# Keep the paths list below in sync with the pull_request paths-ignore list
|
||||
# in ci.yml.
|
||||
# in ci.yml, and keep the quick-checks steps below byte-identical to the
|
||||
# quick-checks job in ci.yml.
|
||||
|
||||
name: Continuous Integration (docs only)
|
||||
|
||||
@@ -47,14 +53,75 @@ on:
|
||||
- ".github/workflows/build.yml"
|
||||
- ".github/workflows/docker.yml"
|
||||
- ".github/workflows/audit.yml"
|
||||
- "flake.lock"
|
||||
|
||||
permissions:
|
||||
contents: read
|
||||
|
||||
jobs:
|
||||
# Deliberately NOT a bare `echo`. Once "Quick Checks" becomes a required
|
||||
# check, ci.yml gates every expensive job behind it, so a mixed PR reports
|
||||
# two check runs with this name: the real one (45-51s) and this companion.
|
||||
# GitHub has no written contract for how it picks between same-named
|
||||
# required check runs ("latest wins" vs "any failure blocks"), so instead of
|
||||
# relying on ordering we make both runs execute the same commands against
|
||||
# the same merge ref — their conclusions are then necessarily identical and
|
||||
# the choice does not matter. Keep these steps byte-identical to the
|
||||
# quick-checks job in ci.yml (a guard script that asserts this, and the paths
|
||||
# sync below, is tracked in rustfs/backlog#1603).
|
||||
#
|
||||
# For a genuinely docs-only PR this adds no strictness (no code changed, so
|
||||
# fmt and the guards always pass) and costs ~50s of ubuntu-latest.
|
||||
quick-checks:
|
||||
name: Quick Checks
|
||||
runs-on: ubuntu-latest
|
||||
timeout-minutes: 10
|
||||
steps:
|
||||
- name: Checkout repository
|
||||
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
|
||||
|
||||
- name: Install ripgrep
|
||||
run: sudo apt-get update && sudo apt-get install -y ripgrep
|
||||
|
||||
- name: Install Rust toolchain
|
||||
uses: dtolnay/rust-toolchain@29eef336d9b2848a0b548edc03f92a220660cdb8 # stable
|
||||
with:
|
||||
components: rustfmt
|
||||
|
||||
- name: Check code formatting
|
||||
run: cargo fmt --all --check
|
||||
|
||||
- name: Check unsafe code allowances
|
||||
run: ./scripts/check_unsafe_code_allowances.sh
|
||||
|
||||
- name: Check layered dependencies
|
||||
run: ./scripts/check_layer_dependencies.sh
|
||||
|
||||
- name: Check architecture migration rules
|
||||
run: ./scripts/check_architecture_migration_rules.sh
|
||||
|
||||
- name: Check tokio io-uring feature guard
|
||||
run: ./scripts/check_no_tokio_io_uring.sh
|
||||
|
||||
- name: Check extension schema boundaries
|
||||
run: ./scripts/check_extension_schema_boundaries.sh
|
||||
|
||||
- name: Check body-cache whitelist guard
|
||||
run: ./scripts/check_body_cache_whitelist.sh
|
||||
|
||||
- name: Check no planning docs committed
|
||||
run: ./scripts/check_no_planning_docs.sh
|
||||
|
||||
- name: Check CI paths stay in sync
|
||||
run: ./scripts/check_ci_paths_sync.sh
|
||||
|
||||
- name: Check io_uring lane --lib precondition
|
||||
run: ./scripts/check_uring_lane_lib_only.sh
|
||||
|
||||
test-and-lint:
|
||||
name: Test and Lint
|
||||
runs-on: ubuntu-latest
|
||||
timeout-minutes: 10
|
||||
steps:
|
||||
- name: Checkout repository
|
||||
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
|
||||
|
||||
+155
-60
@@ -33,6 +33,7 @@ on:
|
||||
- ".github/workflows/build.yml"
|
||||
- ".github/workflows/docker.yml"
|
||||
- ".github/workflows/audit.yml"
|
||||
- "flake.lock"
|
||||
pull_request:
|
||||
types: [ opened, synchronize, reopened, closed ]
|
||||
branches: [ main ]
|
||||
@@ -54,6 +55,7 @@ on:
|
||||
- ".github/workflows/build.yml"
|
||||
- ".github/workflows/docker.yml"
|
||||
- ".github/workflows/audit.yml"
|
||||
- "flake.lock"
|
||||
merge_group:
|
||||
types: [ checks_requested ]
|
||||
schedule:
|
||||
@@ -81,6 +83,7 @@ jobs:
|
||||
name: Cancel Closed PR Runs
|
||||
if: github.event_name == 'pull_request' && github.event.action == 'closed'
|
||||
runs-on: ubuntu-latest
|
||||
timeout-minutes: 10
|
||||
steps:
|
||||
- name: Explain cancellation run
|
||||
run: echo "PR closed; this run only cancels older runs in the same concurrency group."
|
||||
@@ -89,6 +92,7 @@ jobs:
|
||||
name: Typos
|
||||
if: github.event_name != 'pull_request' || github.event.action != 'closed'
|
||||
runs-on: ubuntu-latest
|
||||
timeout-minutes: 10
|
||||
steps:
|
||||
- uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
|
||||
- name: Typos check with custom config file
|
||||
@@ -96,6 +100,10 @@ jobs:
|
||||
|
||||
# Fast, compile-free checks that fail early so contributors get feedback in
|
||||
# ~1 minute instead of waiting for the full test job.
|
||||
#
|
||||
# These steps are mirrored byte-for-byte in ci-docs-only.yml so that a mixed
|
||||
# PR, which reports two check runs named "Quick Checks", cannot get one red
|
||||
# and one green. Edit both jobs together.
|
||||
quick-checks:
|
||||
name: Quick Checks
|
||||
if: github.event_name != 'pull_request' || github.event.action != 'closed'
|
||||
@@ -137,24 +145,48 @@ jobs:
|
||||
- name: Check no planning docs committed
|
||||
run: ./scripts/check_no_planning_docs.sh
|
||||
|
||||
- name: Check CI paths stay in sync
|
||||
run: ./scripts/check_ci_paths_sync.sh
|
||||
|
||||
- name: Check io_uring lane --lib precondition
|
||||
run: ./scripts/check_uring_lane_lib_only.sh
|
||||
|
||||
test-and-lint:
|
||||
name: Test and Lint
|
||||
if: github.event_name != 'pull_request' || github.event.action != 'closed'
|
||||
needs: [ quick-checks ]
|
||||
runs-on: sm-standard-4
|
||||
timeout-minutes: 90
|
||||
# Both lines are required. Job-level `permissions` replaces the workflow
|
||||
# block rather than merging with it, so declaring only `actions: write`
|
||||
# would drop `contents: read` and break this job's checkout and the
|
||||
# repo-token the setup action hands to setup-protoc.
|
||||
permissions:
|
||||
contents: read
|
||||
actions: write
|
||||
env:
|
||||
FORCE_JAVASCRIPT_ACTIONS_TO_NODE24: "true"
|
||||
steps:
|
||||
- name: Checkout repository
|
||||
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
|
||||
with:
|
||||
# This job's token can cancel runs and delete Actions caches. Checkout
|
||||
# otherwise writes it into .git/config, where a PR's own build.rs or
|
||||
# proc-macro could read it back out.
|
||||
persist-credentials: false
|
||||
|
||||
- name: Setup Rust environment
|
||||
uses: ./.github/actions/setup
|
||||
with:
|
||||
rust-version: stable
|
||||
cache-shared-key: ci-test
|
||||
github-token: ${{ secrets.GITHUB_TOKEN }}
|
||||
cache-save-if: ${{ github.ref == 'refs/heads/main' }}
|
||||
# Every lane in this workflow reads its cache and none writes it.
|
||||
# cache-warm.yml is the sole writer for all four keys: this workflow
|
||||
# cancels superseded runs on main, and a cancelled run never reaches
|
||||
# rust-cache's post step, so writing from here saved nothing (12 of 15
|
||||
# consecutive main-push runs were cancelled). See rustfs/backlog#1600.
|
||||
cache-shared-key: ci-dev
|
||||
cache-save-if: 'false'
|
||||
install-build-packaging-tools: 'false'
|
||||
|
||||
- name: Prepare test evidence
|
||||
run: |
|
||||
@@ -169,8 +201,14 @@ jobs:
|
||||
# Clippy runs before the test pass: lint failures are the most common
|
||||
# CI-only breakage and should surface in minutes, not after 20+ minutes
|
||||
# of tests.
|
||||
# Sampled too: clippy is the natural control arm for any CARGO_BUILD_JOBS
|
||||
# experiment, since --all-targets is check-only for workspace members and
|
||||
# never links the ~100 test binaries the limit exists to throttle.
|
||||
- name: Run clippy lints
|
||||
run: cargo clippy --all-targets -- -D warnings
|
||||
run: |
|
||||
./scripts/ci/resource_sampler.sh start clippy
|
||||
trap './scripts/ci/resource_sampler.sh stop' EXIT
|
||||
cargo clippy --all-targets -- -D warnings
|
||||
|
||||
- name: Run nextest tests
|
||||
env:
|
||||
@@ -179,33 +217,8 @@ jobs:
|
||||
CARGO_BUILD_JOBS: "2"
|
||||
run: |
|
||||
mkdir -p artifacts/test-and-lint
|
||||
# Evidence sampler for issue #5394: the post-mortem pgrep below runs
|
||||
# only after `timeout` has already TERM'd the whole cargo process
|
||||
# group, so it cannot name a wedged process. Sample system and
|
||||
# process state every 60s instead; the last samples before the
|
||||
# timeout show what was stuck (rustc, linker, build script, memory
|
||||
# pressure, ...). The log rides along in the existing artifact.
|
||||
(
|
||||
while true; do
|
||||
{
|
||||
echo "=== $(date --utc --iso-8601=seconds)"
|
||||
echo "--- load"; cat /proc/loadavg
|
||||
echo "--- psi"; grep -H . /proc/pressure/* 2>/dev/null || true
|
||||
echo "--- mem"; free -m
|
||||
echo "--- disk"; df -h / /home/runner 2>/dev/null || df -h /
|
||||
echo "--- top-rss"
|
||||
ps -eo pid,ppid,stat,etime,rss,pcpu,args --sort=-rss | head -15
|
||||
echo "--- build/test processes"
|
||||
ps -eo pid,ppid,stat,etime,rss,pcpu,args | grep -E '[c]argo|[r]ustc|[n]extest|[c]ollect2|rust-ll[d]|[b]uild-script|deps[/]' || true
|
||||
echo "--- d-state (uninterruptible IO)"
|
||||
ps -eo pid,stat,etime,args | awk 'NR > 1 && $2 ~ /D/' || true
|
||||
echo
|
||||
} >> artifacts/test-and-lint/sampler.log 2>&1 || true
|
||||
sleep 60
|
||||
done
|
||||
) &
|
||||
sampler_pid=$!
|
||||
trap 'kill "${sampler_pid}" 2>/dev/null || true' EXIT
|
||||
./scripts/ci/resource_sampler.sh start nextest
|
||||
trap './scripts/ci/resource_sampler.sh stop' EXIT
|
||||
set +e
|
||||
NEXTEST_HIDE_PROGRESS_BAR=1 timeout --verbose --signal=TERM --kill-after=30s 75m \
|
||||
cargo nextest run --profile ci --all --exclude e2e_test \
|
||||
@@ -277,6 +290,48 @@ jobs:
|
||||
- name: Run rebalance/decommission migration proofs
|
||||
run: ./scripts/check_migration_gate_count.sh
|
||||
|
||||
# Early stop. Once this job has failed the PR cannot merge, so the sibling
|
||||
# lanes are burning runners on a result nobody can act on: on run
|
||||
# 30674613104 three lanes had already failed while Test and Lint and the
|
||||
# rio-v2 variant kept going past 70 minutes.
|
||||
#
|
||||
# Only this job may cancel. The lanes that are NOT required checks
|
||||
# (protocols, ILM, e2e, s3-tests) must never hold that power: a flake in
|
||||
# one of them would turn the required "Test and Lint" into `cancelled`,
|
||||
# which blocks the merge. Today a maintainer can merge with sftp red, and
|
||||
# that has to stay true.
|
||||
#
|
||||
# These steps run last so the `if: always()` artifact upload above still
|
||||
# captures logs and diagnostics before the run goes away.
|
||||
- name: Annotate early-stop reason
|
||||
if: failure() && github.event_name == 'pull_request'
|
||||
run: |
|
||||
echo "## CI early-stop" >> "$GITHUB_STEP_SUMMARY"
|
||||
echo "Job \`${GITHUB_JOB}\` (Test and Lint) failed; cancelling run ${GITHUB_RUN_ID} to free runners." >> "$GITHUB_STEP_SUMMARY"
|
||||
echo "Sibling jobs showing **cancelled** were stopped by this job, not by their own failure." >> "$GITHUB_STEP_SUMMARY"
|
||||
|
||||
# curl rather than `gh`: every existing `gh` call in this repo runs on
|
||||
# ubuntu-latest, and the sm-standard-* images are custom and trimmed (they
|
||||
# ship no C toolchain, see the e2e job below), so `gh` is not known to
|
||||
# exist here.
|
||||
#
|
||||
# Fork PRs are excluded explicitly instead of relying on the error path:
|
||||
# their GITHUB_TOKEN is forced read-only and job-level permissions cannot
|
||||
# raise it, so the call would always 403. Skipping keeps their logs clean.
|
||||
- name: Cancel run on failure (same-repo PR only)
|
||||
if: >-
|
||||
failure() && github.event_name == 'pull_request'
|
||||
&& github.event.pull_request.head.repo.full_name == github.repository
|
||||
continue-on-error: true
|
||||
env:
|
||||
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
|
||||
run: |
|
||||
curl -fsS -X POST \
|
||||
-H "Authorization: Bearer ${GH_TOKEN}" \
|
||||
-H "Accept: application/vnd.github+json" \
|
||||
-H "X-GitHub-Api-Version: 2022-11-28" \
|
||||
"${GITHUB_API_URL}/repos/${GITHUB_REPOSITORY}/actions/runs/${GITHUB_RUN_ID}/cancel" || true
|
||||
|
||||
# Dedicated serial lane for the ILM / lifecycle integration tests. These tests
|
||||
# drive the object layer through process-global singletons (the GLOBAL_ENV
|
||||
# ECStore, the global tier-config manager, background-expiry workers) and bind
|
||||
@@ -290,6 +345,7 @@ jobs:
|
||||
test-ilm-integration-serial:
|
||||
name: ILM Integration (serial)
|
||||
if: github.event_name != 'pull_request' || github.event.action != 'closed'
|
||||
needs: [ quick-checks ]
|
||||
runs-on: sm-standard-4
|
||||
timeout-minutes: 45
|
||||
env:
|
||||
@@ -302,9 +358,9 @@ jobs:
|
||||
uses: ./.github/actions/setup
|
||||
with:
|
||||
rust-version: stable
|
||||
cache-shared-key: ci-ilm-serial
|
||||
github-token: ${{ secrets.GITHUB_TOKEN }}
|
||||
cache-save-if: ${{ github.ref == 'refs/heads/main' }}
|
||||
cache-shared-key: ci-dev
|
||||
cache-save-if: 'false'
|
||||
install-build-packaging-tools: 'false'
|
||||
|
||||
# test_transition_and_restore_flows was re-enabled by rustfs/backlog#1303:
|
||||
# its "missing xl.meta on disk2" was a test-util bug (open_disk hardcoded
|
||||
@@ -327,6 +383,7 @@ jobs:
|
||||
test-and-lint-rio-v2:
|
||||
name: Test and Lint (rio-v2)
|
||||
if: github.event_name != 'pull_request' || github.event.action != 'closed'
|
||||
needs: [ quick-checks ]
|
||||
runs-on: sm-standard-4
|
||||
timeout-minutes: 60
|
||||
env:
|
||||
@@ -339,9 +396,9 @@ jobs:
|
||||
uses: ./.github/actions/setup
|
||||
with:
|
||||
rust-version: stable
|
||||
cache-shared-key: ci-test-rio-v2
|
||||
github-token: ${{ secrets.GITHUB_TOKEN }}
|
||||
cache-save-if: ${{ github.ref == 'refs/heads/main' }}
|
||||
cache-shared-key: ci-feat-rio
|
||||
cache-save-if: 'false'
|
||||
install-build-packaging-tools: 'false'
|
||||
|
||||
- name: Run rio-v2 clippy lints
|
||||
run: cargo clippy -p rustfs -p rustfs-ecstore --all-targets --features rio-v2 -- -D warnings
|
||||
@@ -354,10 +411,17 @@ jobs:
|
||||
test-and-lint-protocols:
|
||||
name: "Test and Lint (${{ matrix.features.name }})"
|
||||
if: github.event_name != 'pull_request' || github.event.action != 'closed'
|
||||
needs: [ quick-checks ]
|
||||
runs-on: sm-standard-4
|
||||
timeout-minutes: 60
|
||||
strategy:
|
||||
fail-fast: false
|
||||
# On a PR, one failing protocol leg is enough to know the PR is not ready,
|
||||
# so stop the sibling leg instead of paying another ~40 minutes for it.
|
||||
# Everywhere else (main pushes, the merge queue, the weekly schedule) keep
|
||||
# the full signal: there we want to know whether swift AND sftp are broken,
|
||||
# not just whichever failed first. This is the only part of the early-stop
|
||||
# work that also covers fork PRs, since it needs no token.
|
||||
fail-fast: ${{ github.event_name == 'pull_request' }}
|
||||
matrix:
|
||||
features:
|
||||
- name: swift
|
||||
@@ -374,9 +438,9 @@ jobs:
|
||||
uses: ./.github/actions/setup
|
||||
with:
|
||||
rust-version: stable
|
||||
cache-shared-key: ci-test-${{ matrix.features.name }}
|
||||
github-token: ${{ secrets.GITHUB_TOKEN }}
|
||||
cache-save-if: ${{ github.ref == 'refs/heads/main' }}
|
||||
cache-shared-key: ci-feat-proto
|
||||
cache-save-if: 'false'
|
||||
install-build-packaging-tools: 'false'
|
||||
|
||||
- name: Run clippy with ${{ matrix.features.name }}
|
||||
run: |
|
||||
@@ -389,6 +453,7 @@ jobs:
|
||||
build-rustfs-debug-binary:
|
||||
name: Build RustFS Debug Binary
|
||||
if: github.event_name != 'pull_request' || github.event.action != 'closed'
|
||||
needs: [ quick-checks ]
|
||||
runs-on: sm-standard-4
|
||||
timeout-minutes: 30
|
||||
env:
|
||||
@@ -401,9 +466,9 @@ jobs:
|
||||
uses: ./.github/actions/setup
|
||||
with:
|
||||
rust-version: stable
|
||||
cache-shared-key: ci-rustfs-debug-binary
|
||||
cache-save-if: ${{ github.ref == 'refs/heads/main' }}
|
||||
github-token: ${{ secrets.GITHUB_TOKEN }}
|
||||
cache-shared-key: ci-dev
|
||||
cache-save-if: 'false'
|
||||
install-build-packaging-tools: 'false'
|
||||
|
||||
- name: Build debug binary
|
||||
run: cargo build -p rustfs --bins --features e2e-test-hooks
|
||||
@@ -419,6 +484,7 @@ jobs:
|
||||
build-rustfs-debug-binary-rio-v2:
|
||||
name: Build RustFS Debug Binary (rio-v2)
|
||||
if: github.event_name != 'pull_request' || github.event.action != 'closed'
|
||||
needs: [ quick-checks ]
|
||||
runs-on: sm-standard-4
|
||||
timeout-minutes: 30
|
||||
env:
|
||||
@@ -431,9 +497,9 @@ jobs:
|
||||
uses: ./.github/actions/setup
|
||||
with:
|
||||
rust-version: stable
|
||||
cache-shared-key: ci-rustfs-debug-binary-rio-v2
|
||||
cache-save-if: ${{ github.ref == 'refs/heads/main' }}
|
||||
github-token: ${{ secrets.GITHUB_TOKEN }}
|
||||
cache-shared-key: ci-feat-rio
|
||||
cache-save-if: 'false'
|
||||
install-build-packaging-tools: 'false'
|
||||
|
||||
- name: Build debug binary with rio-v2
|
||||
run: cargo build -p rustfs --bins --features rio-v2,e2e-test-hooks
|
||||
@@ -448,6 +514,14 @@ jobs:
|
||||
|
||||
uring-integration:
|
||||
name: io_uring Integration (real)
|
||||
# The pull_request trigger includes `closed` purely so the concurrency
|
||||
# group cancels in-flight runs of a closed PR; every other job opts out of
|
||||
# that run with this guard (or is skipped through its `needs` chain). This
|
||||
# job had neither, so each closed/merged PR really ran the whole io_uring
|
||||
# suite (measured 4m17s / 7m19s / 7m31s on runs 30678272341 / 30678117601 /
|
||||
# 30662728539) and kept the cancellation run in progress for minutes.
|
||||
if: github.event_name != 'pull_request' || github.event.action != 'closed'
|
||||
needs: [ quick-checks ]
|
||||
# GitHub-hosted ubuntu-latest runs a recent kernel with io_uring and, unlike
|
||||
# a container, applies no seccomp filter that would block io_uring_setup — so
|
||||
# the probe succeeds and the tests exercise the real UringBackend/FdCache/
|
||||
@@ -463,12 +537,17 @@ jobs:
|
||||
uses: ./.github/actions/setup
|
||||
with:
|
||||
rust-version: stable
|
||||
# Keeps its own key rather than joining ci-dev. rust-cache's key is
|
||||
# built from runner.os/arch plus rustc and lockfile fingerprints — it
|
||||
# does NOT include the runner label or image. ubuntu-latest and
|
||||
# sm-standard-4 are therefore indistinguishable to it, so sharing a key
|
||||
# would let two different system images overwrite each other's
|
||||
# artifacts, and would make a 2-core hosted runner unpack ci-dev's ~3GB
|
||||
# instead of this lane's ~1.3GB. cache-warm.yml warms this key on
|
||||
# ubuntu-latest for the same reason.
|
||||
cache-shared-key: ci-uring
|
||||
github-token: ${{ secrets.GITHUB_TOKEN }}
|
||||
cache-save-if: ${{ github.ref == 'refs/heads/main' }}
|
||||
|
||||
- name: Install build dependencies
|
||||
run: sudo apt-get update && sudo apt-get install -y protobuf-compiler
|
||||
cache-save-if: 'false'
|
||||
install-build-packaging-tools: 'false'
|
||||
|
||||
# ext4 supports O_DIRECT; the runner's default TMPDIR may sit on tmpfs or
|
||||
# overlayfs, where open(O_DIRECT) returns EINVAL/EOPNOTSUPP and the native
|
||||
@@ -495,7 +574,17 @@ jobs:
|
||||
RUSTFS_IO_URING_READ_ENABLE: "true"
|
||||
RUSTFS_URING_TESTS_MUST_RUN: "1"
|
||||
TMPDIR: /mnt/rustfs-odirect
|
||||
run: cargo test -p rustfs-ecstore uring_ -- --test-threads=1 --nocapture
|
||||
# --lib narrows what gets compiled, not what gets run: every selected
|
||||
# test lives in the lib target. The 7 integration binaries under
|
||||
# crates/ecstore/tests/ each reported "running 0 tests" here, so they
|
||||
# were compiled and linked for nothing.
|
||||
#
|
||||
# The `uring_` filter must stay exactly as it is. libtest matches on
|
||||
# substring, so it also selects names containing `during_` — 6 of the 18
|
||||
# selected tests are such incidental matches. Narrowing the filter to
|
||||
# `io_uring` would silently drop them, which is a coverage change.
|
||||
# scripts/check_uring_lane_lib_only.sh guards the --lib precondition.
|
||||
run: cargo test -p rustfs-ecstore --lib uring_ -- --test-threads=1 --nocapture
|
||||
|
||||
e2e-tests:
|
||||
name: End-to-End Tests
|
||||
@@ -513,9 +602,9 @@ jobs:
|
||||
uses: ./.github/actions/setup
|
||||
with:
|
||||
rust-version: stable
|
||||
cache-shared-key: ci-e2e
|
||||
github-token: ${{ secrets.GITHUB_TOKEN }}
|
||||
cache-save-if: ${{ github.ref == 'refs/heads/main' }}
|
||||
cache-shared-key: ci-dev
|
||||
cache-save-if: 'false'
|
||||
install-build-packaging-tools: 'false'
|
||||
|
||||
# Download after the cache restore so the freshly built binary from the
|
||||
# build job always wins over anything restored into target/debug.
|
||||
@@ -594,9 +683,9 @@ jobs:
|
||||
uses: ./.github/actions/setup
|
||||
with:
|
||||
rust-version: stable
|
||||
cache-shared-key: ci-e2e
|
||||
github-token: ${{ secrets.GITHUB_TOKEN }}
|
||||
cache-save-if: ${{ github.ref == 'refs/heads/main' }}
|
||||
cache-shared-key: ci-dev
|
||||
cache-save-if: 'false'
|
||||
install-build-packaging-tools: 'false'
|
||||
|
||||
# Download after the cache restore so the freshly built binary from the
|
||||
# build job always wins over anything restored into target/debug.
|
||||
@@ -746,7 +835,13 @@ jobs:
|
||||
# evaluates ILM within ~2s of the due time, well inside the poll window.
|
||||
s3-lifecycle-behavior-tests:
|
||||
name: S3 Lifecycle Behavior Tests
|
||||
needs: [ build-rustfs-debug-binary ]
|
||||
# Also gated on e2e-tests, matching s3-implemented-tests: when the e2e smoke
|
||||
# suite is already red this lane cannot tell us anything new, and it holds a
|
||||
# sm-standard-4 for up to 30 minutes doing so. Both lanes only download the
|
||||
# prebuilt debug binary (no cargo build), and s3-implemented-tests — which
|
||||
# already waits on e2e-tests — finishes later anyway, so a green PR's total
|
||||
# wall clock is unchanged.
|
||||
needs: [ build-rustfs-debug-binary, e2e-tests ]
|
||||
runs-on: sm-standard-4
|
||||
timeout-minutes: 30
|
||||
steps:
|
||||
|
||||
@@ -37,6 +37,7 @@ jobs:
|
||||
name: Cancel Closed PR Runs
|
||||
if: github.event_name == 'pull_request_target' && github.event.action == 'closed'
|
||||
runs-on: ubuntu-latest
|
||||
timeout-minutes: 10
|
||||
steps:
|
||||
- name: Explain cancellation run
|
||||
run: echo "PR closed; this run only cancels older runs in the same concurrency group."
|
||||
@@ -44,6 +45,7 @@ jobs:
|
||||
cla:
|
||||
if: ${{ (github.event_name != 'issue_comment' || github.event.issue.pull_request) && (github.event_name != 'pull_request_target' || github.event.action != 'closed') }}
|
||||
runs-on: ubuntu-latest
|
||||
timeout-minutes: 30
|
||||
steps:
|
||||
- name: Report CLA result for merge queue
|
||||
if: github.event_name == 'merge_group'
|
||||
|
||||
@@ -68,8 +68,8 @@ jobs:
|
||||
with:
|
||||
rust-version: stable
|
||||
cache-shared-key: ci-coverage
|
||||
github-token: ${{ secrets.GITHUB_TOKEN }}
|
||||
cache-save-if: ${{ github.ref == 'refs/heads/main' }}
|
||||
install-build-packaging-tools: 'false'
|
||||
|
||||
- name: Install cargo-llvm-cov
|
||||
uses: taiki-e/install-action@bffeee26d4db9be238a4ea78d8826604ebcb594d # v2
|
||||
|
||||
@@ -85,6 +85,7 @@ jobs:
|
||||
github.event.workflow_run.head_branch != 'main' &&
|
||||
!contains(github.event.workflow_run.head_branch, '-preview'))
|
||||
runs-on: ubuntu-latest
|
||||
timeout-minutes: 30
|
||||
outputs:
|
||||
should_build: ${{ steps.check.outputs.should_build }}
|
||||
should_push: ${{ steps.check.outputs.should_push }}
|
||||
@@ -102,6 +103,12 @@ jobs:
|
||||
|
||||
- name: Check build conditions
|
||||
id: check
|
||||
env:
|
||||
# dispatch inputs via env, not `${{ }}` interpolation: they are
|
||||
# free-form strings and would otherwise be evaluated by bash.
|
||||
INPUT_VERSION: ${{ github.event.inputs.version }}
|
||||
INPUT_PUSH_IMAGES: ${{ github.event.inputs.push_images }}
|
||||
INPUT_FORCE_REBUILD: ${{ github.event.inputs.force_rebuild }}
|
||||
run: |
|
||||
should_build=false
|
||||
should_push=false
|
||||
@@ -202,9 +209,9 @@ jobs:
|
||||
|
||||
elif [[ "${{ github.event_name }}" == "workflow_dispatch" ]]; then
|
||||
# Manual trigger
|
||||
input_version="${{ github.event.inputs.version }}"
|
||||
input_version="$INPUT_VERSION"
|
||||
version="${input_version}"
|
||||
should_push="${{ github.event.inputs.push_images }}"
|
||||
should_push="$INPUT_PUSH_IMAGES"
|
||||
should_build=true
|
||||
|
||||
# Get short SHA
|
||||
@@ -212,7 +219,7 @@ jobs:
|
||||
|
||||
echo "🎯 Manual Docker build triggered:"
|
||||
echo " 📋 Requested version: $input_version"
|
||||
echo " 🔧 Force rebuild: ${{ github.event.inputs.force_rebuild }}"
|
||||
echo " 🔧 Force rebuild: $INPUT_FORCE_REBUILD"
|
||||
echo " 🚀 Push images: $should_push"
|
||||
|
||||
case "$input_version" in
|
||||
@@ -333,32 +340,28 @@ jobs:
|
||||
CREATE_LATEST="${{ needs.build-check.outputs.create_latest }}"
|
||||
VARIANT_SUFFIX="${{ matrix.suffix }}"
|
||||
|
||||
# Convert version format for Dockerfile compatibility
|
||||
# Convert version format for Dockerfile compatibility. The former
|
||||
# DOCKER_CHANNEL was "release" down every branch and was passed as a
|
||||
# build-arg no Dockerfile declares, so it is gone.
|
||||
case "$VERSION" in
|
||||
"latest")
|
||||
# For stable latest, use RELEASE=latest + release CHANNEL
|
||||
DOCKER_RELEASE="latest"
|
||||
DOCKER_CHANNEL="release"
|
||||
;;
|
||||
v*)
|
||||
# For versioned releases (v1.0.0), remove 'v' prefix for Dockerfile
|
||||
DOCKER_RELEASE="${VERSION#v}"
|
||||
DOCKER_CHANNEL="release"
|
||||
;;
|
||||
*)
|
||||
# For other versions, pass as-is
|
||||
DOCKER_RELEASE="${VERSION}"
|
||||
DOCKER_CHANNEL="release"
|
||||
;;
|
||||
esac
|
||||
|
||||
echo "docker_release=$DOCKER_RELEASE" >> "$GITHUB_OUTPUT"
|
||||
echo "docker_channel=$DOCKER_CHANNEL" >> "$GITHUB_OUTPUT"
|
||||
|
||||
echo "🐳 Docker build parameters:"
|
||||
echo " - Original version: $VERSION"
|
||||
echo " - Docker RELEASE: $DOCKER_RELEASE"
|
||||
echo " - Docker CHANNEL: $DOCKER_CHANNEL"
|
||||
|
||||
# Generate tags based on build type
|
||||
# Only support release and prerelease builds (no development builds)
|
||||
@@ -412,18 +415,24 @@ jobs:
|
||||
push: ${{ needs.build-check.outputs.should_push == 'true' }}
|
||||
tags: ${{ steps.meta.outputs.tags }}
|
||||
labels: ${{ steps.meta.outputs.labels }}
|
||||
cache-from: |
|
||||
type=gha,scope=docker-${{ matrix.variant }}
|
||||
cache-to: |
|
||||
type=gha,mode=max,scope=docker-${{ matrix.variant }}
|
||||
# No layer cache. This build compiles nothing — it downloads a
|
||||
# release zip and runs apk/apt — so the cache could only save the
|
||||
# minute or two those take, while creating a correctness problem: with
|
||||
# RELEASE=latest the binary URL is resolved by curl *inside* a RUN
|
||||
# layer, and the layer key does not include what that resolved to. A
|
||||
# rebuild at the same RELEASE value (dispatch with version=latest, or
|
||||
# a re-run of the same version) would hit the old layer and ship the
|
||||
# previous release's binary. mode=max also consumed the same 10GB
|
||||
# Actions cache quota the Rust lanes are fighting over.
|
||||
#
|
||||
# Only RELEASE is passed: it is the sole build-arg the Dockerfiles
|
||||
# declare besides TARGETARCH. BUILDTIME, VERSION, BUILD_TYPE, REVISION
|
||||
# and CHANNEL were never read by any stage (and BUILDTIME's $(date ...)
|
||||
# was a literal here, not a shell substitution). BUILD_DATE and VCS_REF
|
||||
# are declared by the Dockerfiles but deliberately left unset —
|
||||
# supplying them would change the published image labels.
|
||||
build-args: |
|
||||
BUILDTIME=$(date -u +'%Y-%m-%dT%H:%M:%SZ')
|
||||
VERSION=${{ needs.build-check.outputs.version }}
|
||||
BUILD_TYPE=${{ needs.build-check.outputs.build_type }}
|
||||
REVISION=${{ github.sha }}
|
||||
RELEASE=${{ steps.meta.outputs.docker_release }}
|
||||
CHANNEL=${{ steps.meta.outputs.docker_channel }}
|
||||
BUILDKIT_INLINE_CACHE=1
|
||||
provenance: true
|
||||
sbom: true
|
||||
# Add retry mechanism by splitting the build process
|
||||
@@ -439,6 +448,7 @@ jobs:
|
||||
needs: [ build-check, build-docker ]
|
||||
if: needs.build-check.outputs.should_build == 'true' && needs.build-check.outputs.should_push == 'true'
|
||||
runs-on: ubuntu-latest
|
||||
timeout-minutes: 30
|
||||
permissions:
|
||||
contents: read
|
||||
security-events: write
|
||||
@@ -493,6 +503,7 @@ jobs:
|
||||
needs: [ build-check, build-docker ]
|
||||
if: always() && needs.build-check.outputs.should_build == 'true'
|
||||
runs-on: ubuntu-latest
|
||||
timeout-minutes: 10
|
||||
steps:
|
||||
- name: Docker build completion summary
|
||||
run: |
|
||||
|
||||
@@ -70,8 +70,8 @@ jobs:
|
||||
with:
|
||||
rust-version: stable
|
||||
cache-shared-key: ci-e2e-repl
|
||||
github-token: ${{ secrets.GITHUB_TOKEN }}
|
||||
cache-save-if: ${{ github.ref == 'refs/heads/main' }}
|
||||
install-build-packaging-tools: 'false'
|
||||
|
||||
# awscurl lets the STS dual-node test actually exercise its path. Without
|
||||
# it the test skips gracefully with a visible log line
|
||||
|
||||
@@ -45,6 +45,13 @@
|
||||
# The PR gate (ci.yml s3-implemented-tests) is unaffected: it avoids Docker
|
||||
# via DEPLOY_MODE=binary and defers all pip setup to run.sh's self-bootstrap.
|
||||
|
||||
# DISABLED. This workflow is switched off in the repository's Actions settings
|
||||
# (state: disabled_manually) and does not run on any trigger, including its cron
|
||||
# and workflow_dispatch. That state lives in GitHub's UI and is invisible when
|
||||
# reading this file, which has already misled at least one audit — hence this
|
||||
# banner. Re-enabling is a UI action; anyone doing so should first check that the
|
||||
# workflow still matches the current CI layout. See rustfs/backlog#1603.
|
||||
#
|
||||
name: e2e-s3tests
|
||||
|
||||
on:
|
||||
|
||||
@@ -12,6 +12,13 @@
|
||||
# See the License for the specific language governing permissions and
|
||||
# limitations under the License.
|
||||
|
||||
# DISABLED. This workflow is switched off in the repository's Actions settings
|
||||
# (state: disabled_manually) and does not run on any trigger, including its cron
|
||||
# and workflow_dispatch. That state lives in GitHub's UI and is invisible when
|
||||
# reading this file, which has already misled at least one audit — hence this
|
||||
# banner. Re-enabling is a UI action; anyone doing so should first check that the
|
||||
# workflow still matches the current CI layout. See rustfs/backlog#1603.
|
||||
#
|
||||
name: Fuzz
|
||||
|
||||
on:
|
||||
@@ -59,6 +66,7 @@ jobs:
|
||||
name: Cancel Closed PR Runs
|
||||
if: github.event_name == 'pull_request' && github.event.action == 'closed'
|
||||
runs-on: ubuntu-latest
|
||||
timeout-minutes: 10
|
||||
steps:
|
||||
- name: Explain cancellation run
|
||||
run: echo "PR closed; this run only cancels older runs in the same concurrency group."
|
||||
@@ -85,7 +93,6 @@ jobs:
|
||||
with:
|
||||
rust-version: nightly
|
||||
cache-shared-key: fuzz-${{ hashFiles('fuzz/Cargo.lock') }}
|
||||
github-token: ${{ secrets.GITHUB_TOKEN }}
|
||||
cache-save-if: ${{ github.ref == 'refs/heads/main' || github.event_name == 'schedule' }}
|
||||
|
||||
- name: Install cargo-fuzz
|
||||
|
||||
@@ -32,6 +32,7 @@ permissions:
|
||||
jobs:
|
||||
build-helm-package:
|
||||
runs-on: ubuntu-latest
|
||||
timeout-minutes: 30
|
||||
if: |
|
||||
(github.event_name == 'workflow_dispatch' && !contains(github.event.inputs.version, '-preview')) ||
|
||||
(
|
||||
@@ -50,15 +51,23 @@ jobs:
|
||||
- name: Checkout helm chart repo
|
||||
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7
|
||||
|
||||
# Both inputs reach the shell through env rather than `${{ }}`
|
||||
# interpolation. A git ref name may contain `$(...)` — anything without a
|
||||
# space is a legal tag — and interpolation pastes it into the script
|
||||
# verbatim, where bash would run it. Reading "$RAW_INPUT" instead makes it
|
||||
# data.
|
||||
- name: Normalize release version
|
||||
id: version
|
||||
env:
|
||||
RAW_INPUT: ${{ github.event.inputs.version }}
|
||||
RAW_BRANCH: ${{ github.event.workflow_run.head_branch }}
|
||||
run: |
|
||||
set -eux
|
||||
|
||||
if [ "${{ github.event_name }}" = "workflow_dispatch" ]; then
|
||||
RAW="${{ github.event.inputs.version }}"
|
||||
RAW="$RAW_INPUT"
|
||||
else
|
||||
RAW="${{ github.event.workflow_run.head_branch }}"
|
||||
RAW="$RAW_BRANCH"
|
||||
fi
|
||||
|
||||
case "$RAW" in
|
||||
@@ -73,10 +82,13 @@ jobs:
|
||||
./scripts/helm_chart_version.sh "$RAW_TAG"
|
||||
|
||||
- name: Replace chart version and app version
|
||||
env:
|
||||
CHART_VERSION: ${{ steps.version.outputs.chart_version }}
|
||||
APP_VERSION: ${{ steps.version.outputs.app_version }}
|
||||
run: |
|
||||
set -eux
|
||||
sed -i -E 's/^version:.*/version: "${{ steps.version.outputs.chart_version }}"/' helm/rustfs/Chart.yaml
|
||||
sed -i -E 's/^appVersion:.*/appVersion: "${{ steps.version.outputs.app_version }}"/' helm/rustfs/Chart.yaml
|
||||
sed -i -E "s/^version:.*/version: \"${CHART_VERSION}\"/" helm/rustfs/Chart.yaml
|
||||
sed -i -E "s/^appVersion:.*/appVersion: \"${APP_VERSION}\"/" helm/rustfs/Chart.yaml
|
||||
|
||||
- name: Set up Helm
|
||||
uses: azure/setup-helm@b9e51907a09c216f16ebe8536097933489208112 # v4.3.0
|
||||
@@ -101,6 +113,7 @@ jobs:
|
||||
|
||||
publish-helm-package:
|
||||
runs-on: ubuntu-latest
|
||||
timeout-minutes: 30
|
||||
needs: [ build-helm-package ]
|
||||
if: needs.build-helm-package.result == 'success'
|
||||
|
||||
@@ -123,11 +136,19 @@ jobs:
|
||||
- name: Generate index
|
||||
run: helm repo index . --url https://charts.rustfs.com
|
||||
|
||||
# app_version is derived from the triggering tag name, and this job holds
|
||||
# the cross-repository push token with rustfs/helm already checked out —
|
||||
# the worst place in the repo to paste an attacker-influenced string into
|
||||
# a shell line. Passed through env so bash treats it as data.
|
||||
- name: Push helm package and index file
|
||||
env:
|
||||
GIT_USERNAME: ${{ secrets.USERNAME }}
|
||||
GIT_EMAIL: ${{ secrets.EMAIL_ADDRESS }}
|
||||
APP_VERSION: ${{ needs.build-helm-package.outputs.app_version }}
|
||||
run: |
|
||||
set -eux
|
||||
git config --global user.name "${{ secrets.USERNAME }}"
|
||||
git config --global user.email "${{ secrets.EMAIL_ADDRESS }}"
|
||||
git config --global user.name "${GIT_USERNAME}"
|
||||
git config --global user.email "${GIT_EMAIL}"
|
||||
git add .
|
||||
git commit -m "Update rustfs helm package with ${{ needs.build-helm-package.outputs.app_version }}." || echo "No changes to commit"
|
||||
git commit -m "Update rustfs helm package with ${APP_VERSION}." || echo "No changes to commit"
|
||||
git push origin main
|
||||
|
||||
@@ -12,6 +12,13 @@
|
||||
# See the License for the specific language governing permissions and
|
||||
# limitations under the License.
|
||||
|
||||
# DISABLED. This workflow is switched off in the repository's Actions settings
|
||||
# (state: disabled_manually) and does not run on any trigger, including its cron
|
||||
# and workflow_dispatch. That state lives in GitHub's UI and is invisible when
|
||||
# reading this file, which has already misled at least one audit — hence this
|
||||
# banner. Re-enabling is a UI action; anyone doing so should first check that the
|
||||
# workflow still matches the current CI layout. See rustfs/backlog#1603.
|
||||
#
|
||||
name: "issue-translator"
|
||||
on:
|
||||
issue_comment:
|
||||
@@ -26,6 +33,7 @@ permissions:
|
||||
jobs:
|
||||
build:
|
||||
runs-on: ubuntu-latest
|
||||
timeout-minutes: 10
|
||||
steps:
|
||||
- uses: usthe/issues-translate-action@b41f55ddc81d7d54bd542a4f289fe28ec081898e # v2.7
|
||||
with:
|
||||
|
||||
@@ -23,6 +23,13 @@
|
||||
# Runner: GitHub-hosted `ubuntu-latest`. It reliably ships Docker + Python,
|
||||
# unlike the self-hosted fleet, whose pods drift in Docker/pip availability
|
||||
# (see the infra note in e2e-s3tests.yml). Nightly + manual only.
|
||||
# DISABLED. This workflow is switched off in the repository's Actions settings
|
||||
# (state: disabled_manually) and does not run on any trigger, including its cron
|
||||
# and workflow_dispatch. That state lives in GitHub's UI and is invisible when
|
||||
# reading this file, which has already misled at least one audit — hence this
|
||||
# banner. Re-enabling is a UI action; anyone doing so should first check that the
|
||||
# workflow still matches the current CI layout. See rustfs/backlog#1603.
|
||||
#
|
||||
name: minio-interop
|
||||
|
||||
on:
|
||||
@@ -57,7 +64,6 @@ jobs:
|
||||
with:
|
||||
rust-version: stable
|
||||
cache-shared-key: ci-minio-interop
|
||||
github-token: ${{ secrets.GITHUB_TOKEN }}
|
||||
cache-save-if: ${{ github.ref == 'refs/heads/main' }}
|
||||
|
||||
- name: Generate real MinIO fixtures via Docker
|
||||
|
||||
@@ -45,6 +45,13 @@
|
||||
# docker-capable self-hosted `dind-sm-standard-2` label was the alternative but
|
||||
# has fewer cores and reintroduces fleet-state risk for no reliability gain.
|
||||
|
||||
# DISABLED. This workflow is switched off in the repository's Actions settings
|
||||
# (state: disabled_manually) and does not run on any trigger, including its cron
|
||||
# and workflow_dispatch. That state lives in GitHub's UI and is invisible when
|
||||
# reading this file, which has already misled at least one audit — hence this
|
||||
# banner. Re-enabling is a UI action; anyone doing so should first check that the
|
||||
# workflow still matches the current CI layout. See rustfs/backlog#1603.
|
||||
#
|
||||
name: mint
|
||||
|
||||
on:
|
||||
|
||||
@@ -19,9 +19,12 @@ on:
|
||||
schedule:
|
||||
- cron: '0 5 * * 0' # Weekly on Sunday 05:00 UTC (staggered after the midnight ci/build crons)
|
||||
|
||||
# GITHUB_TOKEN only needs to read the repository here: the branch push and the
|
||||
# pull request are both created by update-flake-lock using the
|
||||
# FLAKE_UPDATE_TOKEN PAT below, not by this token. Leaving write on it hands a
|
||||
# repo-write credential to an unattended weekly job that does not use it.
|
||||
permissions:
|
||||
contents: write
|
||||
pull-requests: write
|
||||
contents: read
|
||||
|
||||
concurrency:
|
||||
group: ${{ github.workflow }}-${{ github.ref }}
|
||||
|
||||
@@ -12,6 +12,13 @@
|
||||
# See the License for the specific language governing permissions and
|
||||
# limitations under the License.
|
||||
|
||||
# DISABLED. This workflow is switched off in the repository's Actions settings
|
||||
# (state: disabled_manually) and does not run on any trigger, including its cron
|
||||
# and workflow_dispatch. That state lives in GitHub's UI and is invisible when
|
||||
# reading this file, which has already misled at least one audit — hence this
|
||||
# banner. Re-enabling is a UI action; anyone doing so should first check that the
|
||||
# workflow still matches the current CI layout. See rustfs/backlog#1603.
|
||||
#
|
||||
name: Nix CI
|
||||
|
||||
on:
|
||||
@@ -46,6 +53,7 @@ jobs:
|
||||
name: Cancel Closed PR Runs
|
||||
if: github.event_name == 'pull_request' && github.event.action == 'closed'
|
||||
runs-on: ubuntu-latest
|
||||
timeout-minutes: 10
|
||||
steps:
|
||||
- name: Explain cancellation run
|
||||
run: echo "PR closed; this run only cancels older runs in the same concurrency group."
|
||||
|
||||
@@ -22,6 +22,13 @@
|
||||
# a deliberate correctness cost (e.g. the #4221 fsync durability fix) is
|
||||
# recorded but does not block (rustfs/backlog#935 correction 1).
|
||||
|
||||
# DISABLED. This workflow is switched off in the repository's Actions settings
|
||||
# (state: disabled_manually) and does not run on any trigger, including its cron
|
||||
# and workflow_dispatch. That state lives in GitHub's UI and is invisible when
|
||||
# reading this file, which has already misled at least one audit — hence this
|
||||
# banner. Re-enabling is a UI action; anyone doing so should first check that the
|
||||
# workflow still matches the current CI layout. See rustfs/backlog#1603.
|
||||
#
|
||||
name: Performance A/B
|
||||
|
||||
on:
|
||||
@@ -99,7 +106,6 @@ jobs:
|
||||
rust-version: stable
|
||||
cache-shared-key: warp-ab-${{ hashFiles('**/Cargo.lock') }}
|
||||
cache-save-if: ${{ github.ref == 'refs/heads/main' }}
|
||||
github-token: ${{ secrets.GITHUB_TOKEN }}
|
||||
|
||||
- name: Build release rustfs
|
||||
run: cargo build --release --bin rustfs
|
||||
@@ -150,7 +156,6 @@ jobs:
|
||||
rust-version: stable
|
||||
cache-shared-key: warp-ab-${{ hashFiles('**/Cargo.lock') }}
|
||||
cache-save-if: ${{ github.ref == 'refs/heads/main' }}
|
||||
github-token: ${{ secrets.GITHUB_TOKEN }}
|
||||
|
||||
- name: Install warp
|
||||
run: |
|
||||
@@ -162,13 +167,15 @@ jobs:
|
||||
|
||||
- name: Decide exemption
|
||||
id: exempt
|
||||
env:
|
||||
INPUT_ALLOW_REGRESSION: ${{ github.event.inputs.allow_regression }}
|
||||
run: |
|
||||
allow="false"
|
||||
if [[ "${{ github.event_name }}" == "pull_request" ]] \
|
||||
&& ${{ contains(github.event.pull_request.labels.*.name, 'perf-deliberate-tradeoff') }}; then
|
||||
allow="true"
|
||||
fi
|
||||
if [[ "${{ github.event.inputs.allow_regression }}" == "true" ]]; then
|
||||
if [[ "$INPUT_ALLOW_REGRESSION" == "true" ]]; then
|
||||
allow="true"
|
||||
fi
|
||||
echo "allow_regression=$allow" >> "$GITHUB_OUTPUT"
|
||||
@@ -224,6 +231,8 @@ jobs:
|
||||
|
||||
- name: Run warp A/B and gate
|
||||
id: ab
|
||||
env:
|
||||
INPUT_DURATION: ${{ github.event.inputs.duration }}
|
||||
run: |
|
||||
set -euo pipefail
|
||||
# Budget note: with perf-3's cached baseline the nightly does no source
|
||||
@@ -234,7 +243,7 @@ jobs:
|
||||
# budget, which the rig's previous 60s health poll undershot (the first
|
||||
# two nightly failures). perf-6 recalibrates these once the noise study
|
||||
# lands.
|
||||
duration="${{ github.event.inputs.duration || '12s' }}"
|
||||
duration="${INPUT_DURATION:-12s}"
|
||||
baseline_sha="${{ steps.commits.outputs.baseline_sha }}"
|
||||
candidate_sha="${{ steps.commits.outputs.candidate_sha }}"
|
||||
baseline_hit="${{ steps.baseline_cache.outputs.cache-hit }}"
|
||||
|
||||
@@ -24,6 +24,13 @@
|
||||
# The run itself is expected to end red (the forced failure); only the
|
||||
# alert-on-failure job result matters.
|
||||
|
||||
# DISABLED. This workflow is switched off in the repository's Actions settings
|
||||
# (state: disabled_manually) and does not run on any trigger, including its cron
|
||||
# and workflow_dispatch. That state lives in GitHub's UI and is invisible when
|
||||
# reading this file, which has already misled at least one audit — hence this
|
||||
# banner. Re-enabling is a UI action; anyone doing so should first check that the
|
||||
# workflow still matches the current CI layout. See rustfs/backlog#1603.
|
||||
#
|
||||
name: Schedule Failure Alert Drill
|
||||
|
||||
on:
|
||||
|
||||
@@ -12,6 +12,13 @@
|
||||
# See the License for the specific language governing permissions and
|
||||
# limitations under the License.
|
||||
|
||||
# DISABLED. This workflow is switched off in the repository's Actions settings
|
||||
# (state: disabled_manually) and does not run on any trigger, including its cron
|
||||
# and workflow_dispatch. That state lives in GitHub's UI and is invisible when
|
||||
# reading this file, which has already misled at least one audit — hence this
|
||||
# banner. Re-enabling is a UI action; anyone doing so should first check that the
|
||||
# workflow still matches the current CI layout. See rustfs/backlog#1603.
|
||||
#
|
||||
name: "Mark stale issues"
|
||||
on:
|
||||
schedule:
|
||||
@@ -20,6 +27,7 @@ on:
|
||||
jobs:
|
||||
stale:
|
||||
runs-on: ubuntu-latest
|
||||
timeout-minutes: 30
|
||||
steps:
|
||||
- uses: actions/stale@5bef64f19d7facfb25b37b414482c7164d639639 # v9
|
||||
with:
|
||||
|
||||
@@ -15,6 +15,7 @@ concurrency:
|
||||
jobs:
|
||||
update:
|
||||
runs-on: ubuntu-latest
|
||||
timeout-minutes: 10
|
||||
steps:
|
||||
- uses: overtrue/repo-visuals-action@72f34d24769ff5d341956da2f23952594ef2f1e2 # v1.3.0
|
||||
with:
|
||||
|
||||
@@ -105,33 +105,78 @@ CI) fails the build if anything is committed under `docs/superpowers/`, even via
|
||||
|
||||
## Verification Before PR
|
||||
|
||||
Convert changes into independently verifiable outcomes. Prefer focused tests for behavior changes and run the relevant checks before declaring completion.
|
||||
Non-exempt changes must also pass Adversarial Validation (next section) before the checks below count as completion.
|
||||
Convert changes into independently verifiable outcomes. This section controls
|
||||
agent-run local validation; preparing a commit or PR does not by itself require
|
||||
the broadest gate. Inspect only the final task-owned diff, classify it by
|
||||
behavioral impact rather than line count or path alone, and run the smallest
|
||||
set of checks that provides meaningful coverage. Do not let unrelated
|
||||
worktree changes or a generic contributor checklist expand the scope.
|
||||
Non-exempt changes must also pass Adversarial Validation (next section) before
|
||||
the checks below count as completion.
|
||||
|
||||
For code changes, run and pass the following before opening a PR:
|
||||
### Validation floor
|
||||
|
||||
```bash
|
||||
make pre-pr
|
||||
```
|
||||
- Every change that is not documentation-only must finish with
|
||||
`cargo fmt --all --check` passing. An umbrella gate that runs this exact
|
||||
check satisfies the requirement; do not run it twice. Use `cargo fmt --all`
|
||||
only when formatting needs to be fixed. Run the configured formatter or
|
||||
validator for other changed languages when one exists.
|
||||
- Documentation-only or instruction-only means all task-owned changes are
|
||||
prose or documentation assets and cannot affect runtime, builds, CI,
|
||||
dependencies, generated code, or tests. Run `git diff --check` and any
|
||||
relevant documentation guard, but skip Cargo formatting, compilation,
|
||||
Clippy, tests, `make pre-commit`, and `make pre-pr`.
|
||||
- Behavior changes require relevant existing or new tests. Prefer the most
|
||||
focused test or affected package. A passing targeted test can also provide
|
||||
sufficient compilation coverage when it builds every changed target and
|
||||
feature involved; do not add a redundant `cargo check` in that case.
|
||||
- `cargo check` supplements compilation coverage; it never substitutes for a
|
||||
behavioral test. If a relevant test cannot reasonably be added or run, use
|
||||
the narrowest compilation check and report the reason and remaining risk.
|
||||
|
||||
Before committing code changes, prefer focused verification for the touched
|
||||
surface and use the faster local gate when a broad smoke check is needed:
|
||||
### Validation tiers
|
||||
|
||||
```bash
|
||||
make pre-commit
|
||||
```
|
||||
1. **Documentation/instruction-only:** Apply the exemption above. Run a guard
|
||||
such as `make doc-paths-check` only when it is relevant to the edited text.
|
||||
2. **Non-behavioral source change:** For comments, formatting, or another
|
||||
demonstrably non-executable change, run the formatting floor. Compilation,
|
||||
Clippy, and tests may be skipped only when the edit cannot affect
|
||||
compilation or runtime behavior; run targeted doctests if executable
|
||||
documentation examples changed.
|
||||
3. **Localized or bounded behavior change:** Run the formatting floor and the
|
||||
narrowest relevant tests. Add package-scoped `cargo check` or Clippy only
|
||||
for changed targets, features, APIs, error handling, async behavior, or
|
||||
control flow not already covered. When several crates are affected but the
|
||||
dependency set is identifiable, validate those packages and known
|
||||
dependents instead of the whole workspace. Use `make pre-commit` only when
|
||||
a repository-wide fast gate adds useful confidence beyond those checks.
|
||||
4. **Broad or high-risk change:** Run `make pre-pr` only when targeted coverage
|
||||
cannot bound the impact, including:
|
||||
- dependency, feature, build-script, procedural-macro, code-generation,
|
||||
toolchain, or CI changes that alter compilation or the test matrix;
|
||||
- cross-crate public APIs, shared foundational code, or broad refactors with
|
||||
an unbounded dependent set;
|
||||
- locking, storage durability or formats, erasure coding, replication,
|
||||
RPC/protocol compatibility, IAM/KMS/auth, cryptography, or other
|
||||
security-sensitive behavior;
|
||||
- a targeted check that reveals wider impact, an explicit user request, or
|
||||
a release policy that requires the full gate.
|
||||
|
||||
For migration batches, do not run the full `make pre-pr` gate before every
|
||||
intermediate commit. Use focused tests and `make pre-commit` during
|
||||
development, then reserve `make pre-pr` for the final PR-ready branch.
|
||||
Documentation-only and non-behavioral classifications take precedence over
|
||||
path-based triggers. A small diff can still be high-risk, while a CI comment,
|
||||
manifest comment, or release-note edit does not require full validation.
|
||||
|
||||
Before pushing code changes, make sure formatting is clean:
|
||||
`make pre-pr` includes `make pre-commit` coverage. Never run both for the same
|
||||
unchanged diff, and do not repeat equivalent checks during PR preparation or
|
||||
because a local hook already ran them. Rerun only checks whose scope is affected
|
||||
by later edits. Full workspace checks do not replace a relevant integration or
|
||||
E2E test for changed behavior; run that focused test when required and
|
||||
available, or report why it was not run and the remaining risk.
|
||||
|
||||
- Run `cargo fmt --all`.
|
||||
- Run `cargo fmt --all --check` and ensure no files are modified unexpectedly.
|
||||
If `make` is unavailable, run the equivalent checks defined under
|
||||
`.config/make/`. At handoff, list the checks actually run, checks intentionally
|
||||
skipped, and the reason for the selected tier.
|
||||
|
||||
If `make` is unavailable, run the equivalent checks defined under `.config/make/`.
|
||||
Documentation-only or instruction-only changes are exempt from the verification commands above (including the `.config/make/` equivalents), though any locally installed git pre-commit hooks may still run on commit unless explicitly skipped.
|
||||
After build-based verification completes, clean generated build artifacts before wrapping up to avoid unnecessary disk usage.
|
||||
Do not open a PR with code changes when the required checks fail.
|
||||
Make a failing check pass by fixing the cause, never by weakening the gate:
|
||||
|
||||
Generated
+1
@@ -10027,6 +10027,7 @@ dependencies = [
|
||||
"rustfs-data-usage",
|
||||
"rustfs-ecstore",
|
||||
"rustfs-filemeta",
|
||||
"rustfs-lock",
|
||||
"rustfs-storage-api",
|
||||
"rustfs-utils",
|
||||
"s3s",
|
||||
|
||||
@@ -117,9 +117,14 @@ mod tests {
|
||||
.key("assets/explicit-copy.js")
|
||||
.copy_source(format!("{bucket}/{key}"))
|
||||
.metadata_directive(MetadataDirective::Copy)
|
||||
.customize()
|
||||
.mutate_request(|request| {
|
||||
request.headers_mut().insert("content-type", "application/octet-stream");
|
||||
request.headers_mut().insert("x-amz-meta-request-only", "ignored");
|
||||
})
|
||||
.send()
|
||||
.await
|
||||
.expect("explicit COPY directive failed");
|
||||
.expect("explicit COPY directive with request metadata failed");
|
||||
let explicit_copy_head = client
|
||||
.head_object()
|
||||
.bucket(bucket)
|
||||
@@ -128,6 +133,18 @@ mod tests {
|
||||
.await
|
||||
.expect("HEAD failed after explicit COPY");
|
||||
assert_eq!(explicit_copy_head.cache_control(), Some("max-age=60"));
|
||||
assert_eq!(explicit_copy_head.content_type(), Some("text/javascript; charset=utf-8"));
|
||||
assert_eq!(
|
||||
explicit_copy_head.metadata().and_then(|metadata| metadata.get("mtime")),
|
||||
Some(&"1777992333".to_string())
|
||||
);
|
||||
assert_eq!(
|
||||
explicit_copy_head
|
||||
.metadata()
|
||||
.and_then(|metadata| metadata.get("request-only")),
|
||||
None,
|
||||
"COPY must ignore request metadata"
|
||||
);
|
||||
assert_eq!(
|
||||
explicit_copy_head.website_redirect_location(),
|
||||
None,
|
||||
@@ -571,20 +588,6 @@ mod tests {
|
||||
Some("InvalidArgument")
|
||||
);
|
||||
|
||||
let ignored_replacement = client
|
||||
.copy_object()
|
||||
.bucket(bucket)
|
||||
.key(key)
|
||||
.copy_source(format!("{bucket}/{key}"))
|
||||
.content_type("application/ignored")
|
||||
.send()
|
||||
.await
|
||||
.expect_err("Replacement fields without REPLACE should be rejected");
|
||||
assert_eq!(
|
||||
ignored_replacement.as_service_error().and_then(|error| error.code()),
|
||||
Some("InvalidRequest")
|
||||
);
|
||||
|
||||
let unchanged = client
|
||||
.get_object()
|
||||
.bucket(bucket)
|
||||
|
||||
@@ -56,6 +56,21 @@ mod tests {
|
||||
);
|
||||
}
|
||||
|
||||
async fn assert_current_list_hides_delete_marker(client: &Client, bucket: &str, key: &str) {
|
||||
let listed = client
|
||||
.list_objects_v2()
|
||||
.bucket(bucket)
|
||||
.prefix(key)
|
||||
.send()
|
||||
.await
|
||||
.expect("list current objects after delete marker");
|
||||
|
||||
assert!(
|
||||
listed.contents().iter().all(|object| object.key() != Some(key)),
|
||||
"ListObjectsV2 must hide an object whose latest version is a delete marker"
|
||||
);
|
||||
}
|
||||
|
||||
#[tokio::test]
|
||||
#[serial]
|
||||
async fn test_versioning_only_delete_marker_has_minio_compatible_visibility_for_migration_proof() {
|
||||
@@ -94,6 +109,7 @@ mod tests {
|
||||
assert_eq!(markers[0].version_id(), Some(delete_marker_version_id));
|
||||
assert_eq!(markers[0].is_latest(), Some(true));
|
||||
assert_current_get_is_delete_marker_not_found(&client, bucket, key).await;
|
||||
assert_current_list_hides_delete_marker(&client, bucket, key).await;
|
||||
}
|
||||
|
||||
#[tokio::test]
|
||||
@@ -118,6 +134,17 @@ mod tests {
|
||||
.await
|
||||
.expect("put historical version");
|
||||
let data_version_id = put.version_id().expect("put should return data version id");
|
||||
let listed_before_delete = client
|
||||
.list_objects_v2()
|
||||
.bucket(bucket)
|
||||
.prefix(key)
|
||||
.send()
|
||||
.await
|
||||
.expect("list current object before creating delete marker");
|
||||
assert!(
|
||||
listed_before_delete.contents().iter().any(|object| object.key() == Some(key)),
|
||||
"ListObjectsV2 must include the current object before it is deleted"
|
||||
);
|
||||
|
||||
let delete_marker = client
|
||||
.delete_object()
|
||||
@@ -145,6 +172,7 @@ mod tests {
|
||||
assert_eq!(markers[0].version_id(), Some(delete_marker_version_id));
|
||||
assert_eq!(markers[0].is_latest(), Some(true));
|
||||
assert_current_get_is_delete_marker_not_found(&client, bucket, key).await;
|
||||
assert_current_list_hides_delete_marker(&client, bucket, key).await;
|
||||
|
||||
let historical = client
|
||||
.get_object()
|
||||
|
||||
@@ -26,6 +26,7 @@
|
||||
|
||||
use super::common::*;
|
||||
use aws_sdk_s3::Client;
|
||||
use aws_sdk_s3::error::ProvideErrorMetadata;
|
||||
use aws_sdk_s3::primitives::{ByteStream, DateTimeFormat};
|
||||
use aws_sdk_s3::types::{
|
||||
CompletedMultipartUpload, CompletedPart, Delete, MetadataDirective, ObjectIdentifier, ObjectLockLegalHoldStatus,
|
||||
@@ -2120,6 +2121,127 @@ async fn test_multipart_default_retention_fixed_at_create() {
|
||||
// Versioning Auto-Enable Tests
|
||||
// ============================================================================
|
||||
|
||||
#[tokio::test]
|
||||
#[serial]
|
||||
async fn test_unretained_object_lock_object_delete_and_bucket_cleanup() {
|
||||
init_logging();
|
||||
info!("🧪 Test: Unretained Object Lock object delete and bucket cleanup (Issue #5339)");
|
||||
|
||||
let mut env = ObjectLockTestEnvironment::new()
|
||||
.await
|
||||
.expect("failed to create Object Lock test environment");
|
||||
env.start_rustfs().await.expect("failed to start RustFS");
|
||||
|
||||
let bucket = "test-object-lock-delete-cleanup";
|
||||
let key = "unretained-object";
|
||||
|
||||
env.create_object_lock_bucket(bucket)
|
||||
.await
|
||||
.expect("failed to create Object Lock bucket");
|
||||
let client = env.s3_client();
|
||||
|
||||
let put_response = client
|
||||
.put_object()
|
||||
.bucket(bucket)
|
||||
.key(key)
|
||||
.body(ByteStream::from_static(b"unretained data"))
|
||||
.send()
|
||||
.await
|
||||
.expect("failed to upload unretained object");
|
||||
let object_version_id = put_response
|
||||
.version_id()
|
||||
.expect("Object Lock buckets must create versioned objects")
|
||||
.to_string();
|
||||
|
||||
let delete_response = client
|
||||
.delete_object()
|
||||
.bucket(bucket)
|
||||
.key(key)
|
||||
.send()
|
||||
.await
|
||||
.expect("failed to create delete marker");
|
||||
assert_eq!(delete_response.delete_marker(), Some(true));
|
||||
let delete_marker_version_id = delete_response
|
||||
.version_id()
|
||||
.expect("Deleting without a version ID must create a delete marker")
|
||||
.to_string();
|
||||
|
||||
let get_error = client
|
||||
.get_object()
|
||||
.bucket(bucket)
|
||||
.key(key)
|
||||
.send()
|
||||
.await
|
||||
.expect_err("GET must not return an object hidden by a delete marker");
|
||||
assert_eq!(get_error.raw_response().map(|response| response.status().as_u16()), Some(404));
|
||||
assert_eq!(get_error.as_service_error().and_then(|error| error.code()), Some("NoSuchKey"));
|
||||
|
||||
let listed_objects = client
|
||||
.list_objects_v2()
|
||||
.bucket(bucket)
|
||||
.send()
|
||||
.await
|
||||
.expect("failed to list current objects");
|
||||
assert!(
|
||||
listed_objects.contents().iter().all(|object| object.key() != Some(key)),
|
||||
"ListObjectsV2 must hide objects whose latest version is a delete marker"
|
||||
);
|
||||
|
||||
let listed_versions = client
|
||||
.list_object_versions()
|
||||
.bucket(bucket)
|
||||
.send()
|
||||
.await
|
||||
.expect("failed to list object versions");
|
||||
assert!(
|
||||
listed_versions
|
||||
.versions()
|
||||
.iter()
|
||||
.any(|version| version.key() == Some(key) && version.version_id() == Some(object_version_id.as_str())),
|
||||
"The data version must remain until it is explicitly deleted"
|
||||
);
|
||||
assert!(
|
||||
listed_versions
|
||||
.delete_markers()
|
||||
.iter()
|
||||
.any(|marker| marker.key() == Some(key) && marker.version_id() == Some(delete_marker_version_id.as_str())),
|
||||
"ListObjectVersions must expose the delete marker"
|
||||
);
|
||||
|
||||
client
|
||||
.delete_object()
|
||||
.bucket(bucket)
|
||||
.key(key)
|
||||
.version_id(object_version_id)
|
||||
.send()
|
||||
.await
|
||||
.expect("failed to delete the data version");
|
||||
client
|
||||
.delete_object()
|
||||
.bucket(bucket)
|
||||
.key(key)
|
||||
.version_id(delete_marker_version_id)
|
||||
.send()
|
||||
.await
|
||||
.expect("failed to delete the delete marker");
|
||||
|
||||
let remaining_versions = client
|
||||
.list_object_versions()
|
||||
.bucket(bucket)
|
||||
.send()
|
||||
.await
|
||||
.expect("failed to list versions after cleanup");
|
||||
assert!(remaining_versions.versions().is_empty());
|
||||
assert!(remaining_versions.delete_markers().is_empty());
|
||||
|
||||
client
|
||||
.delete_bucket()
|
||||
.bucket(bucket)
|
||||
.send()
|
||||
.await
|
||||
.expect("Deleting every version must remove xl.meta so the bucket can be deleted normally");
|
||||
}
|
||||
|
||||
#[tokio::test]
|
||||
#[serial]
|
||||
async fn test_versioning_auto_enabled_with_object_lock() {
|
||||
|
||||
@@ -267,12 +267,13 @@ pub mod config {
|
||||
pub mod com {
|
||||
pub use crate::config::com::{
|
||||
COMMA_SEPARATED_LISTS, CONFIG_PREFIX, ENV_CONFIG_RECOVER_ON_CORRUPTION, STORAGE_CLASS_SUB_SYS,
|
||||
ServerConfigCorruptError, ServerConfigSnapshot, delete_config, is_server_config_corrupt_error, lookup_configs,
|
||||
read_config, read_config_no_lock, read_config_with_metadata, read_config_without_migrate,
|
||||
read_config_without_migrate_no_lock, read_existing_server_config_no_lock, read_server_config_snapshot, save_config,
|
||||
save_config_no_lock, save_config_with_opts, save_server_config, save_server_config_no_lock,
|
||||
save_server_config_snapshot, server_config_path, try_migrate_server_config, with_config_object_read_lock,
|
||||
with_config_object_write_lock, with_server_config_read_lock, with_server_config_write_lock,
|
||||
ServerConfigCorruptError, ServerConfigSaveResult, ServerConfigSnapshot, delete_config,
|
||||
is_server_config_corrupt_error, lookup_configs, read_config, read_config_no_lock, read_config_with_metadata,
|
||||
read_config_without_migrate, read_config_without_migrate_no_lock, read_existing_server_config_no_lock,
|
||||
read_server_config_snapshot, save_config, save_config_no_lock, save_config_with_opts, save_server_config,
|
||||
save_server_config_no_lock, save_server_config_snapshot, save_server_config_snapshot_with_generation,
|
||||
server_config_path, try_migrate_server_config, with_config_object_read_lock, with_config_object_write_lock,
|
||||
with_server_config_read_lock, with_server_config_write_lock,
|
||||
};
|
||||
}
|
||||
|
||||
|
||||
@@ -53,12 +53,14 @@ use std::sync::LazyLock;
|
||||
use std::sync::{Arc, RwLock};
|
||||
use tokio::sync::{OwnedRwLockWriteGuard, RwLock as AsyncRwLock};
|
||||
use tracing::{debug, error, info, instrument, warn};
|
||||
use uuid::Uuid;
|
||||
|
||||
pub const CONFIG_PREFIX: &str = "config";
|
||||
const SERVER_CONFIG_OBJECT: &str = "config/config.json";
|
||||
const CONFIG_TRANSACTION_LOCK_SUFFIX: &str = ".transaction.lock";
|
||||
|
||||
// Server-config lock order: SERVER_CONFIG_LOCK -> distributed namespace lock
|
||||
// for SERVER_CONFIG_OBJECT. Readers and writers must never reverse this order.
|
||||
// Server-config lock order: SERVER_CONFIG_LOCK -> transaction lock ->
|
||||
// SERVER_CONFIG_OBJECT. Readers and writers must never reverse this order.
|
||||
static SERVER_CONFIG_LOCK: LazyLock<Arc<AsyncRwLock<()>>> = LazyLock::new(|| Arc::new(AsyncRwLock::new(())));
|
||||
|
||||
fn config_task_join_error(operation: &'static str, error: tokio::task::JoinError) -> Error {
|
||||
@@ -76,8 +78,11 @@ where
|
||||
T: Send + 'static,
|
||||
{
|
||||
tokio::spawn(async move {
|
||||
// Lock order: SERVER_CONFIG_LOCK -> namespace write lock.
|
||||
// Lock order: SERVER_CONFIG_LOCK -> transaction lock -> object lock.
|
||||
let _local_guard = SERVER_CONFIG_LOCK.write().await;
|
||||
let transaction_lock = server_config_transaction_lock_path();
|
||||
let transaction_lock = store.new_ns_lock(RUSTFS_META_BUCKET, &transaction_lock).await?;
|
||||
let _transaction_guard = transaction_lock.get_write_lock(get_lock_acquire_timeout()).await?;
|
||||
let namespace_lock = store.new_ns_lock(RUSTFS_META_BUCKET, SERVER_CONFIG_OBJECT).await?;
|
||||
let _write_guard = namespace_lock.get_write_lock(get_lock_acquire_timeout()).await?;
|
||||
Ok(operation().await)
|
||||
@@ -96,8 +101,11 @@ where
|
||||
T: Send + 'static,
|
||||
{
|
||||
tokio::spawn(async move {
|
||||
// Lock order: SERVER_CONFIG_LOCK -> namespace read lock.
|
||||
// Lock order: SERVER_CONFIG_LOCK -> transaction lock -> object lock.
|
||||
let _local_guard = SERVER_CONFIG_LOCK.read().await;
|
||||
let transaction_lock = server_config_transaction_lock_path();
|
||||
let transaction_lock = store.new_ns_lock(RUSTFS_META_BUCKET, &transaction_lock).await?;
|
||||
let _transaction_guard = transaction_lock.get_read_lock(get_lock_acquire_timeout()).await?;
|
||||
let namespace_lock = store.new_ns_lock(RUSTFS_META_BUCKET, SERVER_CONFIG_OBJECT).await?;
|
||||
let _read_guard = namespace_lock.get_read_lock(get_lock_acquire_timeout()).await?;
|
||||
Ok(operation().await)
|
||||
@@ -567,6 +575,21 @@ where
|
||||
}
|
||||
|
||||
pub async fn save_config_with_opts<S>(api: Arc<S>, file: &str, data: Vec<u8>, opts: &ObjectOptions) -> Result<()>
|
||||
where
|
||||
S: ObjectIO<
|
||||
Error = Error,
|
||||
RangeSpec = HTTPRangeSpec,
|
||||
HeaderMap = HeaderMap,
|
||||
ObjectOptions = ObjectOptions,
|
||||
ObjectInfo = ObjectInfo,
|
||||
GetObjectReader = GetObjectReader,
|
||||
PutObjectReader = PutObjReader,
|
||||
>,
|
||||
{
|
||||
save_config_with_opts_and_metadata(api, file, data, opts).await.map(|_| ())
|
||||
}
|
||||
|
||||
async fn save_config_with_opts_and_metadata<S>(api: Arc<S>, file: &str, data: Vec<u8>, opts: &ObjectOptions) -> Result<ObjectInfo>
|
||||
where
|
||||
S: ObjectIO<
|
||||
Error = Error,
|
||||
@@ -579,11 +602,13 @@ where
|
||||
>,
|
||||
{
|
||||
let mut put_data = PutObjReader::from_vec(data);
|
||||
if let Err(err) = api.put_object(RUSTFS_META_BUCKET, file, &mut put_data, opts).await {
|
||||
error!("save_config_with_opts: err: {:?}, file: {}", err, file);
|
||||
return Err(err);
|
||||
match api.put_object(RUSTFS_META_BUCKET, file, &mut put_data, opts).await {
|
||||
Ok(object_info) => Ok(object_info),
|
||||
Err(err) => {
|
||||
error!("save_config_with_opts: err: {:?}, file: {}", err, file);
|
||||
Err(err)
|
||||
}
|
||||
}
|
||||
Ok(())
|
||||
}
|
||||
|
||||
fn new_server_config() -> Config {
|
||||
@@ -594,8 +619,12 @@ async fn new_and_save_server_config<S>(api: Arc<S>) -> Result<Config>
|
||||
where
|
||||
S: EcstoreObjectIO + StorageAdminApi + NamespaceLocking<Error = Error, NamespaceLock = rustfs_lock::NamespaceLockWrapper>,
|
||||
{
|
||||
let snapshot = read_server_config_snapshot(api.clone()).await?;
|
||||
if snapshot.object_exists() {
|
||||
return Ok(snapshot.config.clone());
|
||||
}
|
||||
let cfg = new_server_config();
|
||||
save_server_config(api, &cfg).await?;
|
||||
save_server_config_snapshot(api, &cfg, &snapshot).await?;
|
||||
|
||||
Ok(cfg)
|
||||
}
|
||||
@@ -617,6 +646,10 @@ pub fn server_config_path() -> String {
|
||||
SERVER_CONFIG_OBJECT.to_string()
|
||||
}
|
||||
|
||||
fn server_config_transaction_lock_path() -> String {
|
||||
format!("{}{CONFIG_TRANSACTION_LOCK_SUFFIX}", server_config_path())
|
||||
}
|
||||
|
||||
fn storage_class_kvs_mut(cfg: &mut Config) -> &mut KVS {
|
||||
let sub_cfg = cfg.0.entry(STORAGE_CLASS_SUB_SYS.to_string()).or_insert_with(|| {
|
||||
let mut section = HashMap::new();
|
||||
@@ -819,6 +852,9 @@ fn apply_external_scalar_config_map(
|
||||
let Some(config_value) = root.get(descriptor.subsystem_key) else {
|
||||
return Ok(false);
|
||||
};
|
||||
if descriptor.subsystem_key == HEAL_SUB_SYS && config_value.is_null() {
|
||||
return Ok(false);
|
||||
}
|
||||
let overrides = decode_scalar_config_value(config_value, descriptor)?;
|
||||
|
||||
if overrides.is_empty() {
|
||||
@@ -1463,21 +1499,142 @@ fn build_audit_object(cfg: &Config) -> Map<String, Value> {
|
||||
build_target_object(cfg, &audit_target_descriptors())
|
||||
}
|
||||
|
||||
fn sync_rendered_target_instance(existing: Value, rendered: Option<&Value>, valid_keys: &[&str]) -> Option<Value> {
|
||||
match existing {
|
||||
Value::Object(mut instance) => {
|
||||
for key in valid_keys {
|
||||
instance.remove(*key);
|
||||
}
|
||||
if let Some(Value::Object(rendered)) = rendered {
|
||||
instance.extend(rendered.clone());
|
||||
}
|
||||
(!instance.is_empty()).then_some(Value::Object(instance))
|
||||
}
|
||||
Value::Array(entries) => {
|
||||
let mut pending = rendered
|
||||
.and_then(Value::as_object)
|
||||
.map(|rendered| {
|
||||
rendered
|
||||
.iter()
|
||||
.filter_map(|(key, value)| parse_target_scalar_value(key, value).map(|value| (key.clone(), value)))
|
||||
.collect::<HashMap<_, _>>()
|
||||
})
|
||||
.unwrap_or_default();
|
||||
let mut updated = Vec::with_capacity(entries.len().saturating_add(pending.len()));
|
||||
for entry in entries {
|
||||
let Some(entry_obj) = entry.as_object() else {
|
||||
updated.push(entry);
|
||||
continue;
|
||||
};
|
||||
let Some(key) = entry_obj.get("key").and_then(Value::as_str) else {
|
||||
updated.push(entry);
|
||||
continue;
|
||||
};
|
||||
if !valid_keys.contains(&key) {
|
||||
updated.push(entry);
|
||||
continue;
|
||||
}
|
||||
let Some(value) = pending.remove(key) else {
|
||||
continue;
|
||||
};
|
||||
let mut entry_obj = entry_obj.clone();
|
||||
entry_obj.insert("value".to_string(), Value::String(value));
|
||||
updated.push(Value::Object(entry_obj));
|
||||
}
|
||||
updated.extend(rendered_scalar_config_kvs_entries(&pending));
|
||||
(!updated.is_empty()).then_some(Value::Array(updated))
|
||||
}
|
||||
value if rendered.is_none() => Some(value),
|
||||
_ => rendered.cloned(),
|
||||
}
|
||||
}
|
||||
|
||||
fn sync_rendered_target_object(
|
||||
target_obj: &mut Map<String, Value>,
|
||||
rendered_target: &Map<String, Value>,
|
||||
descriptors: &[TargetConfigDescriptor],
|
||||
) {
|
||||
for descriptor in descriptors {
|
||||
match rendered_target.get(descriptor.external_key) {
|
||||
Some(Value::Object(v)) => {
|
||||
target_obj.insert(descriptor.external_key.to_string(), Value::Object(v.clone()));
|
||||
target_obj.remove(descriptor.subsystem_key);
|
||||
let existing = target_obj.remove(descriptor.external_key);
|
||||
let alias = target_obj.remove(descriptor.subsystem_key);
|
||||
let mut section = existing
|
||||
.or(alias)
|
||||
.and_then(|value| value.as_object().cloned())
|
||||
.unwrap_or_default();
|
||||
let rendered = rendered_target.get(descriptor.external_key).and_then(Value::as_object);
|
||||
|
||||
if is_target_instance_shorthand(§ion, descriptor.valid_keys) {
|
||||
let has_named_instances = rendered.is_some_and(|instances| instances.keys().any(|name| name != "default"));
|
||||
if !has_named_instances {
|
||||
if let Some(section) = sync_rendered_target_instance(
|
||||
Value::Object(section),
|
||||
rendered.and_then(|instances| instances.get("default")),
|
||||
descriptor.valid_keys,
|
||||
) {
|
||||
target_obj.insert(descriptor.external_key.to_string(), section);
|
||||
}
|
||||
continue;
|
||||
}
|
||||
_ => {
|
||||
target_obj.remove(descriptor.external_key);
|
||||
target_obj.remove(descriptor.subsystem_key);
|
||||
|
||||
let mut nested = Map::new();
|
||||
if let Some(default) = sync_rendered_target_instance(
|
||||
Value::Object(section),
|
||||
rendered.and_then(|instances| instances.get("default")),
|
||||
descriptor.valid_keys,
|
||||
) {
|
||||
nested.insert("default".to_string(), default);
|
||||
}
|
||||
if let Some(rendered) = rendered {
|
||||
for (instance_name, instance) in rendered {
|
||||
if instance_name != "default" {
|
||||
nested.insert(instance_name.clone(), instance.clone());
|
||||
}
|
||||
}
|
||||
}
|
||||
if !nested.is_empty() {
|
||||
target_obj.insert(descriptor.external_key.to_string(), Value::Object(nested));
|
||||
}
|
||||
continue;
|
||||
}
|
||||
|
||||
if let Some(default_alias) = section.remove(DEFAULT_DELIMITER) {
|
||||
if let Some(default) = section.get_mut("default") {
|
||||
if let Some(alias) = sync_rendered_target_instance(default_alias, None, descriptor.valid_keys) {
|
||||
match (default, alias) {
|
||||
(Value::Object(default), Value::Object(alias)) => {
|
||||
for (key, value) in alias {
|
||||
default.entry(key).or_insert(value);
|
||||
}
|
||||
}
|
||||
(Value::Array(default), Value::Array(alias)) => default.extend(alias),
|
||||
_ => {}
|
||||
}
|
||||
}
|
||||
} else {
|
||||
section.insert("default".to_string(), default_alias);
|
||||
}
|
||||
}
|
||||
|
||||
let mut merged = Map::new();
|
||||
for (instance_name, instance) in section {
|
||||
if let Some(instance) = sync_rendered_target_instance(
|
||||
instance,
|
||||
rendered.and_then(|instances| instances.get(&instance_name)),
|
||||
descriptor.valid_keys,
|
||||
) {
|
||||
merged.insert(instance_name, instance);
|
||||
}
|
||||
}
|
||||
if let Some(rendered) = rendered {
|
||||
for (instance_name, instance) in rendered {
|
||||
if !merged.contains_key(instance_name) {
|
||||
merged.insert(instance_name.clone(), instance.clone());
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
if !merged.is_empty() {
|
||||
target_obj.insert(descriptor.external_key.to_string(), Value::Object(merged));
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -1496,6 +1653,14 @@ fn encode_server_config_blob(cfg: &Config, seed: Option<&[u8]>) -> Result<Vec<u8
|
||||
Some(Value::Object(v)) => v,
|
||||
_ => Map::new(),
|
||||
};
|
||||
for key in [
|
||||
storageclass::CLASS_STANDARD,
|
||||
storageclass::CLASS_RRS,
|
||||
storageclass::OPTIMIZE,
|
||||
storageclass::INLINE_BLOCK,
|
||||
] {
|
||||
sc_obj.remove(key);
|
||||
}
|
||||
for (k, v) in build_storageclass_object(cfg) {
|
||||
sc_obj.insert(k, v);
|
||||
}
|
||||
@@ -1503,7 +1668,10 @@ fn encode_server_config_blob(cfg: &Config, seed: Option<&[u8]>) -> Result<Vec<u8
|
||||
root.remove("storage_class");
|
||||
|
||||
for descriptor in [scanner_config_descriptor(), heal_config_descriptor()] {
|
||||
let existing = root.remove(descriptor.subsystem_key);
|
||||
let mut existing = root.remove(descriptor.subsystem_key);
|
||||
if descriptor.subsystem_key == HEAL_SUB_SYS && existing.as_ref().is_some_and(Value::is_null) {
|
||||
existing = None;
|
||||
}
|
||||
let rendered = build_scalar_config_object(cfg, descriptor);
|
||||
if let Some(config_value) = sync_rendered_scalar_config_value(existing, &rendered, descriptor)? {
|
||||
root.insert(descriptor.subsystem_key.to_string(), config_value);
|
||||
@@ -1560,6 +1728,7 @@ fn is_standard_object_server_config(data: &[u8]) -> bool {
|
||||
matches!(root.get("version"), Some(Value::String(v)) if !v.trim().is_empty())
|
||||
&& matches!(root.get("storageclass"), Some(Value::Object(_)))
|
||||
&& !root.contains_key("storage_class")
|
||||
&& !matches!(root.get(HEAL_SUB_SYS), Some(Value::Null))
|
||||
}
|
||||
|
||||
fn configs_semantically_equal(lhs: &Config, rhs: &Config) -> bool {
|
||||
@@ -1593,7 +1762,7 @@ where
|
||||
FileInfo = FileInfo,
|
||||
ObjectToDelete = ObjectToDelete,
|
||||
DeletedObject = DeletedObject,
|
||||
>,
|
||||
> + NamespaceLocking<Error = Error, NamespaceLock = rustfs_lock::NamespaceLockWrapper>,
|
||||
{
|
||||
if let Some(decrypt) = &decrypt_fn {
|
||||
register_server_config_decrypt_fn(decrypt.clone());
|
||||
@@ -1601,14 +1770,7 @@ where
|
||||
|
||||
let config_file = server_config_path();
|
||||
match api
|
||||
.get_object_info(
|
||||
RUSTFS_META_BUCKET,
|
||||
&config_file,
|
||||
&ObjectOptions {
|
||||
no_lock: true,
|
||||
..Default::default()
|
||||
},
|
||||
)
|
||||
.get_object_info(RUSTFS_META_BUCKET, &config_file, &ObjectOptions::default())
|
||||
.await
|
||||
{
|
||||
Ok(_) => {
|
||||
@@ -1624,7 +1786,6 @@ where
|
||||
|
||||
let opts = ObjectOptions {
|
||||
max_parity: true,
|
||||
no_lock: true,
|
||||
..Default::default()
|
||||
};
|
||||
|
||||
@@ -1677,7 +1838,33 @@ where
|
||||
}
|
||||
};
|
||||
|
||||
match save_config(api, &config_file, normalized).await {
|
||||
let snapshot = match read_server_config_snapshot(api.clone()).await {
|
||||
Ok(snapshot) => snapshot,
|
||||
Err(err) => {
|
||||
warn!("recheck target server config failed, skip migration: {:?}", err);
|
||||
return;
|
||||
}
|
||||
};
|
||||
if snapshot.object_exists() {
|
||||
debug!("server config was created while legacy migration was preparing, skip migration");
|
||||
return;
|
||||
}
|
||||
|
||||
match save_config_with_opts(
|
||||
api,
|
||||
&config_file,
|
||||
normalized,
|
||||
&ObjectOptions {
|
||||
max_parity: true,
|
||||
http_preconditions: Some(HTTPPreconditions {
|
||||
if_none_match: Some("*".to_string()),
|
||||
..Default::default()
|
||||
}),
|
||||
..Default::default()
|
||||
},
|
||||
)
|
||||
.await
|
||||
{
|
||||
Ok(()) => {
|
||||
info!("Migrated compatible server config from legacy metadata bucket");
|
||||
}
|
||||
@@ -1769,8 +1956,13 @@ where
|
||||
{
|
||||
let config_file = server_config_path();
|
||||
|
||||
// Try to read the configuration file
|
||||
match read_config_no_lock(api.clone(), &config_file).await {
|
||||
// Try to read the configuration file.
|
||||
let data = if namespace_lock_held {
|
||||
read_config_no_lock(api.clone(), &config_file).await
|
||||
} else {
|
||||
read_config(api.clone(), &config_file).await
|
||||
};
|
||||
match data {
|
||||
Ok(data) => read_server_config(api, &data, namespace_lock_held).await,
|
||||
Err(Error::ConfigNotFound) => handle_missing_config(api, "Read the main configuration", namespace_lock_held).await,
|
||||
Err(err) => handle_config_read_error(err, &config_file),
|
||||
@@ -1787,7 +1979,12 @@ where
|
||||
warn!("Received empty configuration data, try to reread from '{}'", config_file);
|
||||
|
||||
// Try to read the configuration again
|
||||
match read_config_no_lock(api.clone(), &config_file).await {
|
||||
let data = if namespace_lock_held {
|
||||
read_config_no_lock(api.clone(), &config_file).await
|
||||
} else {
|
||||
read_config(api.clone(), &config_file).await
|
||||
};
|
||||
match data {
|
||||
Ok(cfg_data) => {
|
||||
let cfg = decode_persisted_server_config(&cfg_data)?;
|
||||
return Ok(cfg.merge());
|
||||
@@ -2036,11 +2233,16 @@ pub struct ServerConfigSnapshot {
|
||||
raw: Option<Vec<u8>>,
|
||||
seed: Option<Vec<u8>>,
|
||||
etag: Option<String>,
|
||||
generation: Option<Uuid>,
|
||||
_local_guard: OwnedRwLockWriteGuard<()>,
|
||||
_guard: rustfs_lock::NamespaceLockGuard,
|
||||
}
|
||||
|
||||
impl ServerConfigSnapshot {
|
||||
pub fn object_exists(&self) -> bool {
|
||||
self.raw.is_some()
|
||||
}
|
||||
|
||||
pub fn ensure_lock_held(&self) -> Result<()> {
|
||||
if self._guard.is_lock_lost() {
|
||||
return Err(Error::other("server config transaction lock was lost"));
|
||||
@@ -2051,12 +2253,34 @@ impl ServerConfigSnapshot {
|
||||
pub fn is_lock_lost(&self) -> bool {
|
||||
self._guard.is_lock_lost()
|
||||
}
|
||||
|
||||
pub fn generation(&self) -> Option<Uuid> {
|
||||
self.generation
|
||||
}
|
||||
}
|
||||
|
||||
/// Read a server config transaction snapshot while holding the same local and
|
||||
/// distributed write locks used by every other server-config writer. Internal
|
||||
/// reads and the later conditional write use no-lock object I/O; the guards
|
||||
/// remain live until the snapshot is dropped.
|
||||
#[derive(Debug, Clone, PartialEq, Eq)]
|
||||
pub struct ServerConfigSaveResult {
|
||||
persisted: bool,
|
||||
generation: Option<Uuid>,
|
||||
}
|
||||
|
||||
impl ServerConfigSaveResult {
|
||||
pub fn persisted(&self) -> bool {
|
||||
self.persisted
|
||||
}
|
||||
|
||||
pub fn generation(&self) -> Option<Uuid> {
|
||||
self.generation
|
||||
}
|
||||
}
|
||||
|
||||
/// Read a server config transaction snapshot while holding a dedicated
|
||||
/// transaction lock. The config object's normal namespace lock remains
|
||||
/// available to fence reads and the conditional write at commit time.
|
||||
/// The transaction guard remains live until the snapshot is dropped,
|
||||
/// serializing persistence and history ordering across admin nodes. Runtime
|
||||
/// state is reloaded from the durable object after this guard is released.
|
||||
pub async fn read_server_config_snapshot<S>(api: Arc<S>) -> Result<ServerConfigSnapshot>
|
||||
where
|
||||
S: ObjectIO<
|
||||
@@ -2071,12 +2295,10 @@ where
|
||||
{
|
||||
let config_file = server_config_path();
|
||||
let local_guard = SERVER_CONFIG_LOCK.clone().write_owned().await;
|
||||
let lock = api.new_ns_lock(RUSTFS_META_BUCKET, &config_file).await?;
|
||||
let transaction_lock = server_config_transaction_lock_path();
|
||||
let lock = api.new_ns_lock(RUSTFS_META_BUCKET, &transaction_lock).await?;
|
||||
let guard = lock.get_write_lock(get_lock_acquire_timeout()).await?;
|
||||
let read_options = ObjectOptions {
|
||||
no_lock: true,
|
||||
..Default::default()
|
||||
};
|
||||
let read_options = ObjectOptions::default();
|
||||
match read_config_with_metadata_inner(api, &config_file, &read_options, true).await {
|
||||
Ok((raw, object_info)) => {
|
||||
let (config, seed) = decode_persisted_server_config_with_seed(&raw)?;
|
||||
@@ -2085,6 +2307,7 @@ where
|
||||
raw: Some(raw),
|
||||
seed: Some(seed),
|
||||
etag: object_info.etag,
|
||||
generation: object_info.data_dir.filter(|generation| !generation.is_nil()),
|
||||
_local_guard: local_guard,
|
||||
_guard: guard,
|
||||
})
|
||||
@@ -2094,6 +2317,7 @@ where
|
||||
raw: None,
|
||||
seed: None,
|
||||
etag: None,
|
||||
generation: None,
|
||||
_local_guard: local_guard,
|
||||
_guard: guard,
|
||||
}),
|
||||
@@ -2108,6 +2332,27 @@ where
|
||||
/// lock, so a concurrent update or transaction lease loss cannot commit an
|
||||
/// unfenced overwrite.
|
||||
pub async fn save_server_config_snapshot<S>(api: Arc<S>, cfg: &Config, snapshot: &ServerConfigSnapshot) -> Result<bool>
|
||||
where
|
||||
S: ObjectIO<
|
||||
Error = Error,
|
||||
RangeSpec = HTTPRangeSpec,
|
||||
HeaderMap = HeaderMap,
|
||||
ObjectOptions = ObjectOptions,
|
||||
ObjectInfo = ObjectInfo,
|
||||
GetObjectReader = GetObjectReader,
|
||||
PutObjectReader = PutObjReader,
|
||||
> + NamespaceLocking<Error = Error, NamespaceLock = rustfs_lock::NamespaceLockWrapper>,
|
||||
{
|
||||
save_server_config_snapshot_with_generation(api, cfg, snapshot)
|
||||
.await
|
||||
.map(|result| result.persisted())
|
||||
}
|
||||
|
||||
pub async fn save_server_config_snapshot_with_generation<S>(
|
||||
api: Arc<S>,
|
||||
cfg: &Config,
|
||||
snapshot: &ServerConfigSnapshot,
|
||||
) -> Result<ServerConfigSaveResult>
|
||||
where
|
||||
S: ObjectIO<
|
||||
Error = Error,
|
||||
@@ -2126,13 +2371,19 @@ where
|
||||
&& configs_semantically_equal(&snapshot.config, cfg)
|
||||
{
|
||||
debug!("server config unchanged and already in standard object shape, skip write");
|
||||
return Ok(false);
|
||||
return Ok(ServerConfigSaveResult {
|
||||
persisted: false,
|
||||
generation: snapshot.generation(),
|
||||
});
|
||||
}
|
||||
|
||||
let data = encode_server_config_blob(cfg, snapshot.seed.as_deref())?;
|
||||
if snapshot.raw.as_deref().is_some_and(|current| current == data.as_slice()) {
|
||||
debug!("server config bytes unchanged after encode, skip write");
|
||||
return Ok(false);
|
||||
return Ok(ServerConfigSaveResult {
|
||||
persisted: false,
|
||||
generation: snapshot.generation(),
|
||||
});
|
||||
}
|
||||
|
||||
let http_preconditions = if snapshot.raw.is_some() {
|
||||
@@ -2152,19 +2403,22 @@ where
|
||||
}
|
||||
};
|
||||
|
||||
save_config_with_opts(
|
||||
snapshot.ensure_lock_held()?;
|
||||
let object_info = save_config_with_opts_and_metadata(
|
||||
api,
|
||||
&config_file,
|
||||
data,
|
||||
&ObjectOptions {
|
||||
max_parity: true,
|
||||
no_lock: true,
|
||||
http_preconditions: Some(http_preconditions),
|
||||
..Default::default()
|
||||
},
|
||||
)
|
||||
.await?;
|
||||
Ok(true)
|
||||
Ok(ServerConfigSaveResult {
|
||||
persisted: true,
|
||||
generation: object_info.data_dir.filter(|generation| !generation.is_nil()),
|
||||
})
|
||||
}
|
||||
|
||||
/// Saves the server config while an upper layer holds the namespace write
|
||||
@@ -2301,8 +2555,9 @@ mod tests {
|
||||
use super::{
|
||||
SERVER_CONFIG_LOCK, ServerConfigSnapshot, apply_dynamic_config_for_sub_sys_with, config_task_join_error,
|
||||
configs_semantically_equal, decode_server_config_blob, encode_server_config_blob, is_standard_object_server_config,
|
||||
lookup_configs, read_config, read_config_preserve_empty, read_config_with_metadata, read_config_without_migrate,
|
||||
read_server_config_snapshot, save_server_config, save_server_config_snapshot, server_config_path, storage_class_kvs_mut,
|
||||
lookup_configs, new_and_save_server_config, read_config, read_config_preserve_empty, read_config_with_metadata,
|
||||
read_config_without_migrate, read_server_config_snapshot, save_server_config, save_server_config_snapshot,
|
||||
save_server_config_snapshot_with_generation, server_config_transaction_lock_path, storage_class_kvs_mut,
|
||||
};
|
||||
use crate::config::{audit, heal, notify, oidc, scanner};
|
||||
use crate::disk::endpoint::Endpoint;
|
||||
@@ -2311,7 +2566,9 @@ mod tests {
|
||||
use crate::object_api::{GetObjectReader, ObjectInfo, ObjectOptions, PutObjReader};
|
||||
use crate::runtime::sources as runtime_sources;
|
||||
use crate::set_disk::SetDisks;
|
||||
use crate::storage_api_contracts::{admin::StorageAdminApi, namespace::NamespaceLocking as _, range::HTTPRangeSpec};
|
||||
use crate::storage_api_contracts::{
|
||||
admin::StorageAdminApi, namespace::NamespaceLocking as _, object::HTTPPreconditions, range::HTTPRangeSpec,
|
||||
};
|
||||
use http::HeaderMap;
|
||||
use rustfs_config::audit::{AUDIT_AMQP_SUB_SYS, AUDIT_KAFKA_SUB_SYS, AUDIT_MQTT_SUB_SYS, AUDIT_WEBHOOK_SUB_SYS};
|
||||
use rustfs_config::notify::{
|
||||
@@ -3104,6 +3361,85 @@ mod tests {
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn root_heal_null_decodes_as_no_override_and_is_canonicalized_on_save() {
|
||||
let seed = br#"{
|
||||
"version":"33",
|
||||
"storageclass":{"standard":"","rrs":""},
|
||||
"heal":null,
|
||||
"future_root":{"mode":"keep"},
|
||||
"openid":{"default":{
|
||||
"config_url":"https://issuer.example/.well-known/openid-configuration",
|
||||
"client_id":"console",
|
||||
"client_secret":"oidc-secret",
|
||||
"future_provider_control":"keep"
|
||||
}},
|
||||
"notify":{"webhook":{"primary":{
|
||||
"enable":true,
|
||||
"endpoint":"https://notify.example/hook",
|
||||
"auth_token":"notify-secret",
|
||||
"future_notify_control":"keep"
|
||||
}}},
|
||||
"logger":{"webhook":{"primary":{
|
||||
"enable":true,
|
||||
"endpoint":"https://audit.example/hook",
|
||||
"auth_token":"audit-secret",
|
||||
"future_audit_control":"keep"
|
||||
}}}
|
||||
}"#;
|
||||
|
||||
let cfg = decode_server_config_blob(seed).expect("root heal null should mean no persisted override");
|
||||
assert!(cfg.get_value(HEAL_SUB_SYS, DEFAULT_DELIMITER).is_none());
|
||||
assert!(!is_standard_object_server_config(seed));
|
||||
|
||||
let encoded = encode_server_config_blob(&cfg, Some(seed)).expect("legacy seed should canonicalize on an authorized save");
|
||||
let value: Value = serde_json::from_slice(&encoded).expect("canonical config should be valid JSON");
|
||||
assert!(value.get(HEAL_SUB_SYS).is_none());
|
||||
assert_eq!(value["future_root"]["mode"].as_str(), Some("keep"));
|
||||
assert_eq!(value["openid"]["default"]["client_secret"].as_str(), Some("oidc-secret"));
|
||||
assert_eq!(value["openid"]["default"]["future_provider_control"].as_str(), Some("keep"));
|
||||
assert_eq!(value["notify"]["webhook"]["primary"]["auth_token"].as_str(), Some("notify-secret"));
|
||||
assert_eq!(value["notify"]["webhook"]["primary"]["future_notify_control"].as_str(), Some("keep"));
|
||||
assert_eq!(value["logger"]["webhook"]["primary"]["auth_token"].as_str(), Some("audit-secret"));
|
||||
assert_eq!(value["logger"]["webhook"]["primary"]["future_audit_control"].as_str(), Some("keep"));
|
||||
assert!(is_standard_object_server_config(&encoded));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn invalid_scalar_and_nested_null_config_shapes_remain_rejected() {
|
||||
let invalid_sections = [
|
||||
r#""scanner":null"#,
|
||||
r#""heal":"""#,
|
||||
r#""heal":false"#,
|
||||
r#""heal":0"#,
|
||||
r#""heal":{"default":null}"#,
|
||||
r#""heal":{"_":null}"#,
|
||||
r#""heal":{"bitrot_cycle":null}"#,
|
||||
r#""heal":[{"key":"bitrot_cycle","value":null}]"#,
|
||||
];
|
||||
|
||||
for section in invalid_sections {
|
||||
let input = format!(r#"{{"version":"33","storageclass":{{"standard":"","rrs":""}},{section}}}"#);
|
||||
let err = decode_server_config_blob(input.as_bytes()).expect_err("invalid scalar shape must remain rejected");
|
||||
assert!(
|
||||
err.to_string().contains("expected"),
|
||||
"invalid section {section} returned an unrelated error: {err}"
|
||||
);
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn valid_heal_object_and_kvs_array_shapes_remain_accepted() {
|
||||
let empty_object = br#"{"version":"33","storageclass":{"standard":"","rrs":""},"heal":{}}"#;
|
||||
let cfg = decode_server_config_blob(empty_object).expect("empty heal object should decode as no override");
|
||||
assert!(cfg.get_value(HEAL_SUB_SYS, DEFAULT_DELIMITER).is_none());
|
||||
|
||||
let kvs_array =
|
||||
br#"{"version":"33","storageclass":{"standard":"","rrs":""},"heal":[{"key":"bitrot_cycle","value":"off"}]}"#;
|
||||
let cfg = decode_server_config_blob(kvs_array).expect("heal KVS array should decode");
|
||||
assert!(cfg.get_value(HEAL_SUB_SYS, DEFAULT_DELIMITER).is_some());
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn scanner_update_preserves_unknown_root_and_oidc_provider_fields() {
|
||||
let seed = br#"{
|
||||
@@ -3131,6 +3467,171 @@ mod tests {
|
||||
assert_eq!(value["openid"]["default"]["client_id"].as_str(), Some("console"));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn storageclass_reset_removes_stale_inline_block_from_seed() {
|
||||
let seed = br#"{
|
||||
"version":"33",
|
||||
"storageclass":{
|
||||
"standard":"EC:2",
|
||||
"rrs":"EC:1",
|
||||
"optimize":"availability",
|
||||
"inline_block":"64KiB",
|
||||
"future_storage_control":"keep"
|
||||
}
|
||||
}"#;
|
||||
|
||||
let encoded = encode_server_config_blob(&Config::new(), Some(seed)).expect("storageclass reset should encode");
|
||||
let value: Value = serde_json::from_slice(&encoded).expect("encoded config should be valid json");
|
||||
let storageclass = value["storageclass"].as_object().expect("storageclass object");
|
||||
|
||||
assert!(storageclass.get(crate::config::storageclass::INLINE_BLOCK).is_none());
|
||||
assert_eq!(storageclass["future_storage_control"].as_str(), Some("keep"));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn target_update_preserves_unknown_fields_without_restoring_removed_instances() {
|
||||
let seed = br#"{
|
||||
"version":"33",
|
||||
"storageclass":{"standard":"","rrs":""},
|
||||
"notify":{"webhook":{
|
||||
"primary":{
|
||||
"enable":true,
|
||||
"endpoint":"https://notify.example/old",
|
||||
"auth_token":"notify-secret",
|
||||
"future_control":"keep"
|
||||
},
|
||||
"removed":{"enable":true,"endpoint":"https://notify.example/removed"},
|
||||
"retained":{"enable":true,"endpoint":"https://notify.example/retained","future_control":"keep"},
|
||||
"enable":{"enable":true,"endpoint":"https://notify.example/named-enable","future_control":"keep"}
|
||||
}}
|
||||
}"#;
|
||||
let mut cfg = decode_server_config_blob(seed).expect("target seed should decode");
|
||||
let webhook = cfg
|
||||
.0
|
||||
.get_mut(NOTIFY_WEBHOOK_SUB_SYS)
|
||||
.expect("notify webhook subsystem should exist");
|
||||
webhook
|
||||
.get_mut("primary")
|
||||
.expect("primary target should exist")
|
||||
.insert(rustfs_config::WEBHOOK_ENDPOINT.to_string(), "https://notify.example/new".to_string());
|
||||
webhook.remove("removed");
|
||||
webhook.remove("retained");
|
||||
|
||||
let encoded = encode_server_config_blob(&cfg, Some(seed)).expect("target update should encode");
|
||||
let value: Value = serde_json::from_slice(&encoded).expect("encoded config should be valid json");
|
||||
let webhook = value["notify"]["webhook"].as_object().expect("webhook section");
|
||||
|
||||
assert_eq!(webhook["primary"]["endpoint"].as_str(), Some("https://notify.example/new"));
|
||||
assert_eq!(webhook["primary"]["auth_token"].as_str(), Some("notify-secret"));
|
||||
assert_eq!(webhook["primary"]["future_control"].as_str(), Some("keep"));
|
||||
assert!(webhook.get("removed").is_none());
|
||||
assert_eq!(webhook["retained"]["future_control"].as_str(), Some("keep"));
|
||||
assert!(webhook["retained"].get("enable").is_none());
|
||||
assert!(webhook["retained"].get("endpoint").is_none());
|
||||
assert_eq!(webhook["enable"]["endpoint"].as_str(), Some("https://notify.example/named-enable"));
|
||||
assert_eq!(webhook["enable"]["future_control"].as_str(), Some("keep"));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn shorthand_target_update_preserves_shape_and_unknown_nested_fields() {
|
||||
let seed = br#"{
|
||||
"version":"33",
|
||||
"storageclass":{"standard":"","rrs":""},
|
||||
"notify":{"webhook":{
|
||||
"enable":true,
|
||||
"endpoint":"https://notify.example/old",
|
||||
"future_control":{"endpoint":"leave-untouched","mode":"keep"}
|
||||
}}
|
||||
}"#;
|
||||
let mut cfg = decode_server_config_blob(seed).expect("shorthand target should decode");
|
||||
cfg.0
|
||||
.get_mut(NOTIFY_WEBHOOK_SUB_SYS)
|
||||
.and_then(|targets| targets.get_mut(DEFAULT_DELIMITER))
|
||||
.expect("default webhook target should exist")
|
||||
.insert(rustfs_config::WEBHOOK_ENDPOINT.to_string(), "https://notify.example/new".to_string());
|
||||
|
||||
let encoded = encode_server_config_blob(&cfg, Some(seed)).expect("shorthand target update should encode");
|
||||
let value: Value = serde_json::from_slice(&encoded).expect("encoded config should be valid json");
|
||||
let webhook = value["notify"]["webhook"].as_object().expect("webhook shorthand object");
|
||||
|
||||
assert_eq!(webhook["endpoint"].as_str(), Some("https://notify.example/new"));
|
||||
assert!(webhook.get("default").is_none());
|
||||
assert_eq!(webhook["future_control"]["endpoint"].as_str(), Some("leave-untouched"));
|
||||
assert_eq!(webhook["future_control"]["mode"].as_str(), Some("keep"));
|
||||
let decoded = decode_server_config_blob(&encoded).expect("updated shorthand target should remain decodable");
|
||||
assert_eq!(
|
||||
decoded
|
||||
.get_value(NOTIFY_WEBHOOK_SUB_SYS, DEFAULT_DELIMITER)
|
||||
.expect("updated default webhook target")
|
||||
.get(rustfs_config::WEBHOOK_ENDPOINT),
|
||||
"https://notify.example/new"
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn target_kvs_update_preserves_unknown_entries_and_attributes() {
|
||||
let seed = br#"{
|
||||
"version":"33",
|
||||
"storageclass":{"standard":"","rrs":""},
|
||||
"notify":{"webhook":{"primary":[
|
||||
{"key":"enable","value":"on","hidden_if_empty":false},
|
||||
{"key":"endpoint","value":"https://notify.example/old","future_attribute":"keep-endpoint"},
|
||||
{"key":"future_control","value":"keep","future_attribute":"keep-control"}
|
||||
]}}
|
||||
}"#;
|
||||
let mut cfg = decode_server_config_blob(seed).expect("target KVS seed should decode");
|
||||
cfg.0
|
||||
.get_mut(NOTIFY_WEBHOOK_SUB_SYS)
|
||||
.and_then(|targets| targets.get_mut("primary"))
|
||||
.expect("primary webhook target should exist")
|
||||
.insert(rustfs_config::WEBHOOK_ENDPOINT.to_string(), "https://notify.example/new".to_string());
|
||||
|
||||
let encoded = encode_server_config_blob(&cfg, Some(seed)).expect("target KVS update should encode");
|
||||
let value: Value = serde_json::from_slice(&encoded).expect("encoded config should be valid json");
|
||||
let entries = value["notify"]["webhook"]["primary"]
|
||||
.as_array()
|
||||
.expect("target KVS shape should be preserved");
|
||||
let endpoint = entries
|
||||
.iter()
|
||||
.find(|entry| entry["key"].as_str() == Some(rustfs_config::WEBHOOK_ENDPOINT))
|
||||
.expect("endpoint entry should remain");
|
||||
let future = entries
|
||||
.iter()
|
||||
.find(|entry| entry["key"].as_str() == Some("future_control"))
|
||||
.expect("unknown target entry should remain");
|
||||
|
||||
assert_eq!(endpoint["value"].as_str(), Some("https://notify.example/new"));
|
||||
assert_eq!(endpoint["future_attribute"].as_str(), Some("keep-endpoint"));
|
||||
assert_eq!(future["value"].as_str(), Some("keep"));
|
||||
assert_eq!(future["future_attribute"].as_str(), Some("keep-control"));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn target_default_alias_is_canonicalized_without_losing_unknown_fields() {
|
||||
let seed = br#"{
|
||||
"version":"33",
|
||||
"storageclass":{"standard":"","rrs":""},
|
||||
"notify":{"webhook":{
|
||||
"_":{"enable":false,"endpoint":"https://notify.example/alias","future_alias":"keep"},
|
||||
"default":{"enable":true,"endpoint":"https://notify.example/default","future_default":"keep"}
|
||||
}}
|
||||
}"#;
|
||||
let cfg = decode_server_config_blob(seed).expect("dual default aliases should decode");
|
||||
let expected_endpoint = cfg
|
||||
.get_value(NOTIFY_WEBHOOK_SUB_SYS, DEFAULT_DELIMITER)
|
||||
.expect("default webhook target should exist")
|
||||
.get(rustfs_config::WEBHOOK_ENDPOINT);
|
||||
|
||||
let encoded = encode_server_config_blob(&cfg, Some(seed)).expect("default alias should canonicalize");
|
||||
let value: Value = serde_json::from_slice(&encoded).expect("encoded config should be valid json");
|
||||
let webhook = value["notify"]["webhook"].as_object().expect("webhook section");
|
||||
|
||||
assert!(webhook.get(DEFAULT_DELIMITER).is_none());
|
||||
assert_eq!(webhook["default"]["endpoint"].as_str(), Some(expected_endpoint.as_str()));
|
||||
assert_eq!(webhook["default"]["future_alias"].as_str(), Some("keep"));
|
||||
assert_eq!(webhook["default"]["future_default"].as_str(), Some("keep"));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_scanner_config_changes_are_semantically_significant() {
|
||||
let baseline = Config::new();
|
||||
@@ -4124,6 +4625,7 @@ mod tests {
|
||||
|
||||
/// What reads of the config object currently return.
|
||||
enum RecoveryReadState {
|
||||
Missing,
|
||||
Blob(Vec<u8>),
|
||||
QuorumError,
|
||||
}
|
||||
@@ -4136,6 +4638,7 @@ mod tests {
|
||||
heal_calls: AtomicUsize,
|
||||
write_calls: AtomicUsize,
|
||||
last_put_no_lock: AtomicBool,
|
||||
last_put_preconditions: Mutex<Option<HTTPPreconditions>>,
|
||||
revision: AtomicUsize,
|
||||
drive_counts: Vec<usize>,
|
||||
lock_manager: Arc<rustfs_lock::GlobalLockManager>,
|
||||
@@ -4150,6 +4653,7 @@ mod tests {
|
||||
heal_calls: AtomicUsize::new(0),
|
||||
write_calls: AtomicUsize::new(0),
|
||||
last_put_no_lock: AtomicBool::new(false),
|
||||
last_put_preconditions: Mutex::new(None),
|
||||
revision: AtomicUsize::new(1),
|
||||
drive_counts: vec![2],
|
||||
lock_manager: Arc::new(rustfs_lock::GlobalLockManager::new()),
|
||||
@@ -4206,6 +4710,7 @@ mod tests {
|
||||
_opts: &ObjectOptions,
|
||||
) -> Result<GetObjectReader> {
|
||||
let data = match &*self.state.lock().expect("state lock poisoned") {
|
||||
RecoveryReadState::Missing => return Err(Error::ConfigNotFound),
|
||||
RecoveryReadState::Blob(data) => data.clone(),
|
||||
RecoveryReadState::QuorumError => return Err(Error::ErasureReadQuorum),
|
||||
};
|
||||
@@ -4213,6 +4718,9 @@ mod tests {
|
||||
size: data.len() as i64,
|
||||
actual_size: data.len() as i64,
|
||||
etag: Some(format!("config-{}", self.revision.load(Ordering::SeqCst))),
|
||||
data_dir: Some(uuid::Uuid::from_u128(
|
||||
u128::try_from(self.revision.load(Ordering::SeqCst)).expect("test revision should fit in u128"),
|
||||
)),
|
||||
..Default::default()
|
||||
};
|
||||
Ok(GetObjectReader {
|
||||
@@ -4231,15 +4739,19 @@ mod tests {
|
||||
opts: &ObjectOptions,
|
||||
) -> Result<ObjectInfo> {
|
||||
let current_etag = format!("config-{}", self.revision.load(Ordering::SeqCst));
|
||||
let object_exists = matches!(&*self.state.lock().expect("state lock poisoned"), RecoveryReadState::Blob(_));
|
||||
if let Some(preconditions) = &opts.http_preconditions
|
||||
&& (preconditions.if_match_value().is_some_and(|etag| etag != current_etag)
|
||||
|| preconditions.if_none_match_value() == Some("*"))
|
||||
&& (preconditions
|
||||
.if_match_value()
|
||||
.is_some_and(|etag| !object_exists || etag != current_etag)
|
||||
|| (object_exists && preconditions.if_none_match_value() == Some("*")))
|
||||
{
|
||||
return Err(Error::PreconditionFailed);
|
||||
}
|
||||
let mut body = Vec::new();
|
||||
data.stream.read_to_end(&mut body).await?;
|
||||
self.last_put_no_lock.store(opts.no_lock, Ordering::SeqCst);
|
||||
*self.last_put_preconditions.lock().expect("preconditions lock poisoned") = opts.http_preconditions.clone();
|
||||
self.write_calls.fetch_add(1, Ordering::SeqCst);
|
||||
*self.state.lock().expect("state lock poisoned") = RecoveryReadState::Blob(body.clone());
|
||||
let revision = self.revision.fetch_add(1, Ordering::SeqCst) + 1;
|
||||
@@ -4247,6 +4759,7 @@ mod tests {
|
||||
size: i64::try_from(body.len()).expect("test config should fit in i64"),
|
||||
actual_size: i64::try_from(body.len()).expect("test config should fit in i64"),
|
||||
etag: Some(format!("config-{revision}")),
|
||||
data_dir: Some(uuid::Uuid::from_u128(u128::try_from(revision).expect("test revision should fit in u128"))),
|
||||
..Default::default()
|
||||
})
|
||||
}
|
||||
@@ -4284,10 +4797,18 @@ mod tests {
|
||||
.expect("scanner-only config change should be persisted");
|
||||
|
||||
assert_eq!(store.write_calls.load(Ordering::SeqCst), 1);
|
||||
assert!(store.last_put_no_lock.load(Ordering::SeqCst));
|
||||
assert!(!store.last_put_no_lock.load(Ordering::SeqCst));
|
||||
let preconditions = store
|
||||
.last_put_preconditions
|
||||
.lock()
|
||||
.expect("preconditions lock poisoned")
|
||||
.clone()
|
||||
.expect("existing config update must be conditional");
|
||||
assert_eq!(preconditions.if_match_value(), Some("config-1"));
|
||||
assert_eq!(preconditions.if_none_match_value(), None);
|
||||
assert_eq!(
|
||||
store.lock_resources.lock().expect("lock resources mutex poisoned").as_slice(),
|
||||
&[server_config_path()]
|
||||
&[server_config_transaction_lock_path()]
|
||||
);
|
||||
let decoded = read_config_without_migrate(store)
|
||||
.await
|
||||
@@ -4301,6 +4822,73 @@ mod tests {
|
||||
);
|
||||
}
|
||||
|
||||
#[tokio::test]
|
||||
async fn server_config_snapshot_save_returns_committed_generation() {
|
||||
let baseline = encode_server_config_blob(&Config::new(), None).expect("baseline config should encode");
|
||||
let store = Arc::new(RecoveryMockStore::new(RecoveryReadState::Blob(baseline), None));
|
||||
let snapshot = read_server_config_snapshot(store.clone())
|
||||
.await
|
||||
.expect("server config snapshot");
|
||||
|
||||
let result = save_server_config_snapshot_with_generation(store, &config_with_scanner_cycle("61"), &snapshot)
|
||||
.await
|
||||
.expect("conditional config save");
|
||||
|
||||
assert!(result.persisted());
|
||||
assert_eq!(result.generation(), Some(uuid::Uuid::from_u128(2)));
|
||||
}
|
||||
|
||||
#[tokio::test]
|
||||
async fn missing_server_config_is_created_with_if_none_match() {
|
||||
let store = Arc::new(RecoveryMockStore::new(RecoveryReadState::Missing, None));
|
||||
let cfg = config_with_scanner_cycle("61");
|
||||
|
||||
save_server_config(store.clone(), &cfg)
|
||||
.await
|
||||
.expect("missing config should be created conditionally");
|
||||
|
||||
assert_eq!(store.write_calls.load(Ordering::SeqCst), 1);
|
||||
assert!(!store.last_put_no_lock.load(Ordering::SeqCst));
|
||||
let preconditions = store
|
||||
.last_put_preconditions
|
||||
.lock()
|
||||
.expect("preconditions lock poisoned")
|
||||
.clone()
|
||||
.expect("missing config create must be conditional");
|
||||
assert_eq!(preconditions.if_match_value(), None);
|
||||
assert_eq!(preconditions.if_none_match_value(), Some("*"));
|
||||
let persisted = read_config_without_migrate(store)
|
||||
.await
|
||||
.expect("created config should reload");
|
||||
assert_eq!(
|
||||
persisted
|
||||
.get_value(SCANNER_SUB_SYS, DEFAULT_DELIMITER)
|
||||
.expect("persisted scanner config")
|
||||
.get(SCANNER_CYCLE),
|
||||
"61"
|
||||
);
|
||||
}
|
||||
|
||||
#[tokio::test]
|
||||
async fn missing_config_initialization_recheck_preserves_concurrent_config() {
|
||||
let existing = config_with_scanner_cycle("73");
|
||||
let baseline = encode_server_config_blob(&existing, None).expect("existing config should encode");
|
||||
let store = Arc::new(RecoveryMockStore::new(RecoveryReadState::Blob(baseline), None));
|
||||
|
||||
let observed = new_and_save_server_config(store.clone())
|
||||
.await
|
||||
.expect("initialization recheck should return the config created by another writer");
|
||||
|
||||
assert_eq!(store.write_calls.load(Ordering::SeqCst), 0);
|
||||
assert_eq!(
|
||||
observed
|
||||
.get_value(SCANNER_SUB_SYS, DEFAULT_DELIMITER)
|
||||
.expect("concurrent scanner config")
|
||||
.get(SCANNER_CYCLE),
|
||||
"73"
|
||||
);
|
||||
}
|
||||
|
||||
#[tokio::test]
|
||||
async fn stale_server_config_snapshot_cannot_overwrite_newer_update() {
|
||||
let baseline = encode_server_config_blob(&Config::new(), None).expect("baseline config should encode");
|
||||
@@ -4357,7 +4945,7 @@ mod tests {
|
||||
let lock = rustfs_lock::NamespaceLock::new("server-config-lease-loss".to_string(), client.clone());
|
||||
let guard = lock
|
||||
.lock_guard(
|
||||
rustfs_lock::ObjectKey::new(crate::disk::RUSTFS_META_BUCKET, server_config_path()),
|
||||
rustfs_lock::ObjectKey::new(crate::disk::RUSTFS_META_BUCKET, server_config_transaction_lock_path()),
|
||||
"server-config-lease-loss",
|
||||
std::time::Duration::from_secs(1),
|
||||
std::time::Duration::from_millis(120),
|
||||
@@ -4372,6 +4960,7 @@ mod tests {
|
||||
raw: Some(baseline.clone()),
|
||||
seed: None,
|
||||
etag: Some("config-0".to_string()),
|
||||
generation: Some(uuid::Uuid::from_u128(1)),
|
||||
_local_guard: local_guard,
|
||||
_guard: guard,
|
||||
};
|
||||
|
||||
@@ -6641,6 +6641,49 @@ impl LocalDisk {
|
||||
let xl_path = object_dir.join(STORAGE_FORMAT_FILE);
|
||||
restore_delete_rollback_after_error(object_dir, &xl_path, Some(rollback_dir), volume, object, stage, err).await
|
||||
}
|
||||
|
||||
/// Execute every deferred data-dir deletion pending on `volume` right now,
|
||||
/// even while snapshot leases are still held. Bucket deletion requires it:
|
||||
/// a streaming reader defers the physical cleanup of an already-deleted
|
||||
/// version, and a non-force `delete_volume` would otherwise fail closed
|
||||
/// with `VolumeNotEmpty` on those remnants even though the bucket is
|
||||
/// logically empty. The still-active readers keep their open descriptors;
|
||||
/// only path-based reopens observe the removal.
|
||||
async fn settle_pending_snapshot_deletes(&self, volume: &str) {
|
||||
let pending: Vec<(SnapshotLeaseKey, DeleteOptions)> = {
|
||||
let mut registry = self.snapshot_leases.lock().await;
|
||||
registry
|
||||
.entries
|
||||
.iter_mut()
|
||||
.filter(|(key, entry)| key.volume == volume && !entry.deleting && entry.pending_delete.is_some())
|
||||
.map(|(key, entry)| {
|
||||
entry.deleting = true;
|
||||
(key.clone(), entry.pending_delete.clone().expect("filtered on Some"))
|
||||
})
|
||||
.collect()
|
||||
};
|
||||
|
||||
for (key, opts) in pending {
|
||||
let result = self.delete_unleased(&key.volume, &key.path, &opts).await;
|
||||
let mut registry = self.snapshot_leases.lock().await;
|
||||
match result {
|
||||
Ok(()) => {
|
||||
registry.entries.remove(&key);
|
||||
}
|
||||
Err(err) => {
|
||||
if let Some(entry) = registry.entries.get_mut(&key) {
|
||||
entry.deleting = false;
|
||||
}
|
||||
warn!(
|
||||
volume = %key.volume,
|
||||
path = %key.path,
|
||||
error = %err,
|
||||
"failed to settle deferred data-dir deletion before volume removal"
|
||||
);
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
#[async_trait::async_trait]
|
||||
@@ -9122,6 +9165,14 @@ impl DiskAPI for LocalDisk {
|
||||
let p = self.get_bucket_path(volume)?;
|
||||
let _volume_mutation_guard = os::disk_volume_mutation_lock(&self.root, volume).write_owned().await;
|
||||
|
||||
// A streaming reader's snapshot lease defers the physical cleanup of
|
||||
// data dirs whose version delete already committed. Those remnants are
|
||||
// logically deleted, so run the parked cleanups now instead of letting
|
||||
// the non-force removal below fail closed on them (the s3-tests SSE-C
|
||||
// teardown races exactly this way: DeleteObjects, then DeleteBucket
|
||||
// while an abandoned GET body still pins the lease).
|
||||
self.settle_pending_snapshot_deletes(volume).await;
|
||||
|
||||
// Non-force removes empty directory remnants children-first with
|
||||
// non-recursive rmdir calls. A file that exists during the scan, or
|
||||
// appears before its parent is removed, fails closed with
|
||||
@@ -15981,6 +16032,74 @@ mod test {
|
||||
assert!(matches!(disk.read_all(volume, &first_part).await, Err(DiskError::FileNotFound)));
|
||||
}
|
||||
|
||||
#[tokio::test]
|
||||
async fn delete_volume_settles_lease_deferred_cleanup() {
|
||||
use tempfile::tempdir;
|
||||
|
||||
let root_dir = tempdir().expect("temp dir should be created");
|
||||
let endpoint = Endpoint::try_from(root_dir.path().to_string_lossy().as_ref()).expect("endpoint should parse");
|
||||
let disk = LocalDisk::new(&endpoint, false).await.expect("local disk should be created");
|
||||
let volume = "snapshot-lease-bucket-delete";
|
||||
let object = "multipart_enc";
|
||||
let version_id = Uuid::new_v4();
|
||||
let data_dir = Uuid::new_v4();
|
||||
let rollback_dir = Uuid::new_v4();
|
||||
let data_path = path_join_buf(&[object, &data_dir.to_string()]);
|
||||
let part = path_join_buf(&[&data_path, "part.1"]);
|
||||
ensure_test_volume(&disk, volume).await;
|
||||
disk.write_all(volume, &part, Bytes::from_static(b"payload"))
|
||||
.await
|
||||
.expect("shard should be written");
|
||||
let fi = test_file_info(object, version_id, Some(data_dir), None);
|
||||
disk.write_all(volume, &path_join_buf(&[object, STORAGE_FORMAT_FILE]), test_meta(fi.clone()).into())
|
||||
.await
|
||||
.expect("metadata should be written");
|
||||
|
||||
// An abandoned streaming GET pins the data dir with a snapshot lease.
|
||||
let snapshot = disk
|
||||
.acquire_snapshot_lease(volume, &data_path)
|
||||
.await
|
||||
.expect("snapshot lease should be acquired");
|
||||
disk.delete_version(
|
||||
volume,
|
||||
object,
|
||||
fi.clone(),
|
||||
false,
|
||||
DeleteOptions {
|
||||
old_data_dir: Some(rollback_dir),
|
||||
..Default::default()
|
||||
},
|
||||
)
|
||||
.await
|
||||
.expect("version delete should commit metadata");
|
||||
disk.delete(
|
||||
volume,
|
||||
&format!("{object}/{rollback_dir}"),
|
||||
DeleteOptions {
|
||||
recursive: true,
|
||||
immediate: true,
|
||||
..Default::default()
|
||||
},
|
||||
)
|
||||
.await
|
||||
.expect("version delete should schedule physical cleanup");
|
||||
|
||||
// The bucket is logically empty; a non-force volume delete must settle
|
||||
// the deferred data-dir cleanup instead of failing with VolumeNotEmpty.
|
||||
disk.delete_volume(volume, false)
|
||||
.await
|
||||
.expect("bucket delete must not observe lease-deferred remnants");
|
||||
assert!(matches!(
|
||||
disk.read_all(volume, &part).await,
|
||||
Err(DiskError::FileNotFound | DiskError::VolumeNotFound)
|
||||
));
|
||||
|
||||
// The late lease release finds nothing pending and stays idempotent.
|
||||
disk.release_snapshot_lease(volume, &data_path, snapshot)
|
||||
.await
|
||||
.expect("releasing the lease after bucket deletion should be a no-op");
|
||||
}
|
||||
|
||||
#[tokio::test]
|
||||
async fn version_delete_cleanup_intent_survives_local_disk_restart() {
|
||||
use tempfile::tempdir;
|
||||
|
||||
@@ -1645,15 +1645,28 @@ impl Erasure {
|
||||
}
|
||||
Err(e) => {
|
||||
record_get_stage_duration_if_enabled(GET_OBJECT_PATH_LEGACY_DUPLEX, GET_STAGE_EMIT, emit_stage_start);
|
||||
error!(
|
||||
block_offset,
|
||||
block_length,
|
||||
bytes_written = *written,
|
||||
stage = GET_STAGE_EMIT,
|
||||
reason = classify_io_error(&e).as_str(),
|
||||
error = ?e,
|
||||
"Erasure decode failed to emit reconstructed data"
|
||||
);
|
||||
let reason = classify_io_error(&e);
|
||||
if reason == GetObjectFailureReason::DownstreamClosed {
|
||||
debug!(
|
||||
block_offset,
|
||||
block_length,
|
||||
bytes_written = *written,
|
||||
stage = GET_STAGE_EMIT,
|
||||
reason = reason.as_str(),
|
||||
error = ?e,
|
||||
"Erasure decode stopped after downstream closed"
|
||||
);
|
||||
} else {
|
||||
error!(
|
||||
block_offset,
|
||||
block_length,
|
||||
bytes_written = *written,
|
||||
stage = GET_STAGE_EMIT,
|
||||
reason = reason.as_str(),
|
||||
error = ?e,
|
||||
"Erasure decode failed to emit reconstructed data"
|
||||
);
|
||||
}
|
||||
*ret_err = Some(e);
|
||||
return StripeFlow::Stop;
|
||||
}
|
||||
@@ -1945,7 +1958,7 @@ mod tests {
|
||||
use std::io::Cursor;
|
||||
use std::pin::Pin;
|
||||
use std::sync::{
|
||||
Arc,
|
||||
Arc, Mutex,
|
||||
atomic::{AtomicUsize, Ordering},
|
||||
};
|
||||
use std::task::{Context, Poll};
|
||||
@@ -2120,6 +2133,59 @@ mod tests {
|
||||
}
|
||||
}
|
||||
|
||||
struct DownstreamClosedWriter;
|
||||
|
||||
impl AsyncWrite for DownstreamClosedWriter {
|
||||
fn poll_write(self: Pin<&mut Self>, _cx: &mut Context<'_>, _buf: &[u8]) -> Poll<io::Result<usize>> {
|
||||
Poll::Ready(Err(crate::diagnostics::get::mark_get_object_downstream_closed(io::Error::new(
|
||||
ErrorKind::BrokenPipe,
|
||||
"injected downstream close",
|
||||
))))
|
||||
}
|
||||
|
||||
fn poll_flush(self: Pin<&mut Self>, _cx: &mut Context<'_>) -> Poll<io::Result<()>> {
|
||||
Poll::Ready(Ok(()))
|
||||
}
|
||||
|
||||
fn poll_shutdown(self: Pin<&mut Self>, _cx: &mut Context<'_>) -> Poll<io::Result<()>> {
|
||||
Poll::Ready(Ok(()))
|
||||
}
|
||||
}
|
||||
|
||||
#[derive(Clone, Default)]
|
||||
struct CapturedLogs(Arc<Mutex<Vec<u8>>>);
|
||||
|
||||
struct CapturedLogWriter(Arc<Mutex<Vec<u8>>>);
|
||||
|
||||
impl CapturedLogs {
|
||||
fn contents(&self) -> String {
|
||||
String::from_utf8(self.0.lock().expect("captured logs mutex should not be poisoned").clone())
|
||||
.expect("captured logs should be valid UTF-8")
|
||||
}
|
||||
}
|
||||
|
||||
impl std::io::Write for CapturedLogWriter {
|
||||
fn write(&mut self, buf: &[u8]) -> std::io::Result<usize> {
|
||||
self.0
|
||||
.lock()
|
||||
.expect("captured logs mutex should not be poisoned")
|
||||
.extend_from_slice(buf);
|
||||
Ok(buf.len())
|
||||
}
|
||||
|
||||
fn flush(&mut self) -> std::io::Result<()> {
|
||||
Ok(())
|
||||
}
|
||||
}
|
||||
|
||||
impl<'a> tracing_subscriber::fmt::MakeWriter<'a> for CapturedLogs {
|
||||
type Writer = CapturedLogWriter;
|
||||
|
||||
fn make_writer(&'a self) -> Self::Writer {
|
||||
CapturedLogWriter(Arc::clone(&self.0))
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn parallel_reader_constructor_variants_preserve_read_cost_and_verification_flags() {
|
||||
let erasure = Erasure::new(2, 1, 64);
|
||||
@@ -2215,6 +2281,47 @@ mod tests {
|
||||
assert_eq!(err.to_string(), "injected emit failure");
|
||||
}
|
||||
|
||||
#[tokio::test(flavor = "current_thread")]
|
||||
async fn erasure_decode_logs_reconstructed_downstream_close_at_debug() {
|
||||
let logs = CapturedLogs::default();
|
||||
let subscriber = tracing_subscriber::fmt()
|
||||
.with_max_level(tracing::Level::DEBUG)
|
||||
.with_writer(logs.clone())
|
||||
.with_ansi(false)
|
||||
.without_time()
|
||||
.finish();
|
||||
let _guard = tracing::subscriber::set_default(subscriber);
|
||||
|
||||
let erasure = Erasure::new(2, 1, 64);
|
||||
let data: Vec<u8> = (0..64).collect();
|
||||
let shard_size = erasure.shard_size();
|
||||
let encoded = erasure.encode_data(&data).expect("test data should encode");
|
||||
let readers = vec![
|
||||
None,
|
||||
Some(BitrotReader::new(
|
||||
Cursor::new(encoded[1].to_vec()),
|
||||
shard_size,
|
||||
HashAlgorithm::None,
|
||||
false,
|
||||
)),
|
||||
Some(BitrotReader::new(
|
||||
Cursor::new(encoded[2].to_vec()),
|
||||
shard_size,
|
||||
HashAlgorithm::None,
|
||||
false,
|
||||
)),
|
||||
];
|
||||
|
||||
let mut writer = DownstreamClosedWriter;
|
||||
let (written, err) = erasure.decode(&mut writer, readers, 0, data.len(), data.len()).await;
|
||||
|
||||
assert_eq!(written, 0);
|
||||
assert_eq!(err.expect("downstream close must still terminate the GET").kind(), ErrorKind::BrokenPipe);
|
||||
let captured = logs.contents();
|
||||
assert!(captured.contains("Erasure decode stopped after downstream closed"));
|
||||
assert!(!captured.contains("Erasure decode failed to emit reconstructed data"));
|
||||
}
|
||||
|
||||
#[tokio::test]
|
||||
async fn test_erasure_decode_rejects_reader_count_and_range_overflow() {
|
||||
let erasure = Erasure::new(2, 1, 64);
|
||||
|
||||
@@ -22,6 +22,8 @@ use rustfs_config::{
|
||||
ENV_STARTUP_TOPOLOGY_WAIT_MODE, ENV_STARTUP_TOPOLOGY_WAIT_TIMEOUT, ENV_UNSAFE_BYPASS_DISK_CHECK,
|
||||
};
|
||||
use rustfs_utils::{XHost, check_local_server_addr, get_env_opt_str, get_host_ip, is_local_host};
|
||||
#[cfg(test)]
|
||||
use std::sync::{LazyLock, Mutex};
|
||||
use std::{
|
||||
collections::{BTreeMap, BTreeSet, HashMap, HashSet, hash_map::Entry},
|
||||
future::Future,
|
||||
@@ -598,6 +600,58 @@ const DNS_RETRY_JITTER_PERCENT: u64 = 20;
|
||||
/// wait does not flood the log with one line per backoff tick.
|
||||
const TOPOLOGY_WARN_THROTTLE: Duration = Duration::from_secs(30);
|
||||
|
||||
#[cfg(test)]
|
||||
static FORCED_LOCAL_HOST_RESOLUTION_TIMEOUTS: LazyLock<Mutex<HashSet<String>>> = LazyLock::new(|| Mutex::new(HashSet::new()));
|
||||
|
||||
#[cfg(test)]
|
||||
struct LocalHostResolutionTimeoutGuard {
|
||||
hosts: Vec<String>,
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
impl Drop for LocalHostResolutionTimeoutGuard {
|
||||
fn drop(&mut self) {
|
||||
let mut forced_hosts = FORCED_LOCAL_HOST_RESOLUTION_TIMEOUTS
|
||||
.lock()
|
||||
.expect("local-host test resolver mutex poisoned");
|
||||
for host in &self.hosts {
|
||||
forced_hosts.remove(host);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
fn force_local_host_resolution_timeout_for_test(hosts: &[&str]) -> LocalHostResolutionTimeoutGuard {
|
||||
let hosts = hosts.iter().map(|host| (*host).to_string()).collect::<Vec<_>>();
|
||||
let mut forced_hosts = FORCED_LOCAL_HOST_RESOLUTION_TIMEOUTS
|
||||
.lock()
|
||||
.expect("local-host test resolver mutex poisoned");
|
||||
forced_hosts.extend(hosts.iter().cloned());
|
||||
LocalHostResolutionTimeoutGuard { hosts }
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
fn local_host_resolution_timeout_forced(host: &Host<&str>) -> bool {
|
||||
let host = match host {
|
||||
Host::Domain(domain) => (*domain).to_string(),
|
||||
Host::Ipv4(ip) => ip.to_string(),
|
||||
Host::Ipv6(ip) => ip.to_string(),
|
||||
};
|
||||
FORCED_LOCAL_HOST_RESOLUTION_TIMEOUTS
|
||||
.lock()
|
||||
.expect("local-host test resolver mutex poisoned")
|
||||
.contains(&host)
|
||||
}
|
||||
|
||||
fn endpoint_is_local_host(host: Host<&str>, port: u16, local_port: u16) -> Result<bool> {
|
||||
#[cfg(test)]
|
||||
if local_host_resolution_timeout_forced(&host) {
|
||||
return Err(Error::new(ErrorKind::TimedOut, "resolver timeout"));
|
||||
}
|
||||
|
||||
is_local_host(host, port, local_port)
|
||||
}
|
||||
|
||||
struct DnsRetryDeadline {
|
||||
started: Instant,
|
||||
timeout: Duration,
|
||||
@@ -694,7 +748,7 @@ async fn resolve_local_host_with_retry(
|
||||
retry_dns_operation(
|
||||
|| {
|
||||
let host = host.clone();
|
||||
async move { is_local_host(host, port, local_port) }
|
||||
async move { endpoint_is_local_host(host, port, local_port) }
|
||||
},
|
||||
async_sleep,
|
||||
dns_retry_deadline,
|
||||
@@ -2231,6 +2285,9 @@ mod test {
|
||||
#[serial]
|
||||
#[tokio::test]
|
||||
async fn create_server_endpoints_bounds_kubernetes_alias_dns_fallback() {
|
||||
let _resolution_timeout =
|
||||
force_local_host_resolution_timeout_for_test(&["unrelated-0.example.invalid", "unrelated-1.example.invalid"]);
|
||||
|
||||
async_with_vars(
|
||||
[
|
||||
(ENV_LOCAL_ENDPOINT_HOST, None),
|
||||
|
||||
@@ -4714,6 +4714,7 @@ mod tests {
|
||||
use crate::layout::endpoints::SetupType;
|
||||
use crate::object_api::BLOCK_SIZE_V2;
|
||||
use crate::object_api::ObjectInfo;
|
||||
use crate::set_disk::core::io_primitives::rename_fanout_barrier;
|
||||
use crate::storage_api_contracts::{
|
||||
heal::HealOperations as _, lifecycle::TransitionedObject, list::ListOperations as _, multipart::CompletePart,
|
||||
namespace::NamespaceLocking as _, object::ObjectIO as _, object::ObjectOperations as _,
|
||||
@@ -9565,6 +9566,15 @@ mod tests {
|
||||
make_local_bucket_test_set_disks_with_drive_count(2).await
|
||||
}
|
||||
|
||||
fn assert_exclusive_object_lock_held(set_disks: &SetDisks, bucket: &str, object: &str) {
|
||||
let lock = set_disks
|
||||
.local_lock_manager_for_test()
|
||||
.get_lock_info(&ObjectKey::new(bucket, object))
|
||||
.expect("object lock should be visible while rename is paused");
|
||||
assert!(matches!(lock.mode, rustfs_lock::LockMode::Exclusive));
|
||||
assert_eq!(lock.owner.as_ref(), set_disks.locker_owner.as_str());
|
||||
}
|
||||
|
||||
async fn make_local_bucket_test_set_disks_with_drive_count(drive_count: usize) -> Arc<SetDisks> {
|
||||
let format = FormatV3::new(1, drive_count);
|
||||
let mut endpoints = Vec::new();
|
||||
@@ -9599,7 +9609,9 @@ mod tests {
|
||||
disks.push(Some(disk));
|
||||
}
|
||||
|
||||
let set_disks = SetDisks::new(
|
||||
let instance_ctx = Arc::new(InstanceContext::new());
|
||||
instance_ctx.update_erasure_type(SetupType::Erasure).await;
|
||||
let set_disks = SetDisks::new_with_instance_ctx(
|
||||
"test-owner".to_string(),
|
||||
Arc::new(RwLock::new(disks)),
|
||||
drive_count,
|
||||
@@ -9609,6 +9621,7 @@ mod tests {
|
||||
endpoints,
|
||||
format,
|
||||
Vec::new(),
|
||||
instance_ctx,
|
||||
)
|
||||
.await;
|
||||
set_disks.set_test_storage_class_config(
|
||||
@@ -10262,6 +10275,216 @@ mod tests {
|
||||
));
|
||||
}
|
||||
|
||||
#[tokio::test]
|
||||
async fn conditional_replace_holds_object_lock_through_rename() {
|
||||
let set_disks = make_local_bucket_test_set_disks().await;
|
||||
let bucket = "bucket-conditional-replace-fence";
|
||||
let object = "config/conditional-replace.json";
|
||||
set_disks
|
||||
.make_bucket(bucket, &MakeBucketOptions::default())
|
||||
.await
|
||||
.expect("bucket should be created");
|
||||
|
||||
let mut initial_reader = PutObjReader::from_vec(b"initial config".to_vec());
|
||||
let initial = set_disks
|
||||
.put_object(
|
||||
bucket,
|
||||
object,
|
||||
&mut initial_reader,
|
||||
&ObjectOptions {
|
||||
no_lock: true,
|
||||
..Default::default()
|
||||
},
|
||||
)
|
||||
.await
|
||||
.expect("initial config should be written");
|
||||
let initial_etag = initial.etag.expect("initial config should have an ETag");
|
||||
|
||||
let barrier = rename_fanout_barrier::arm(object, 0, rename_fanout_barrier::PHASE_RENAME);
|
||||
let writer_store = set_disks.clone();
|
||||
let expected_etag = initial_etag.clone();
|
||||
let writer = tokio::spawn(async move {
|
||||
let mut reader = PutObjReader::from_vec(b"replacement config".to_vec());
|
||||
writer_store
|
||||
.put_object(
|
||||
bucket,
|
||||
object,
|
||||
&mut reader,
|
||||
&ObjectOptions {
|
||||
preserve_etag: Some("replacement-etag".to_string()),
|
||||
http_preconditions: Some(HTTPPreconditions {
|
||||
if_match: Some(expected_etag),
|
||||
..Default::default()
|
||||
}),
|
||||
..Default::default()
|
||||
},
|
||||
)
|
||||
.await
|
||||
});
|
||||
|
||||
tokio::time::timeout(std::time::Duration::from_secs(30), barrier.wait_until_paused())
|
||||
.await
|
||||
.expect("conditional replace should reach the rename barrier");
|
||||
assert_exclusive_object_lock_held(&set_disks, bucket, object);
|
||||
barrier.release();
|
||||
writer
|
||||
.await
|
||||
.expect("conditional writer task should finish")
|
||||
.expect("matching conditional replace should commit");
|
||||
assert!(
|
||||
set_disks
|
||||
.local_lock_manager_for_test()
|
||||
.get_lock_info(&ObjectKey::new(bucket, object))
|
||||
.is_none(),
|
||||
"conditional replace should release the object lock after commit"
|
||||
);
|
||||
let contender = set_disks
|
||||
.new_ns_lock(bucket, object)
|
||||
.await
|
||||
.expect("contender namespace lock should be created");
|
||||
let contender_guard = contender
|
||||
.get_write_lock(std::time::Duration::from_secs(30))
|
||||
.await
|
||||
.expect("contender should acquire after conditional replace commits");
|
||||
drop(contender_guard);
|
||||
|
||||
let mut stale_reader = PutObjReader::from_vec(b"stale config".to_vec());
|
||||
let err = set_disks
|
||||
.put_object(
|
||||
bucket,
|
||||
object,
|
||||
&mut stale_reader,
|
||||
&ObjectOptions {
|
||||
http_preconditions: Some(HTTPPreconditions {
|
||||
if_match: Some(initial_etag),
|
||||
..Default::default()
|
||||
}),
|
||||
..Default::default()
|
||||
},
|
||||
)
|
||||
.await
|
||||
.expect_err("the old ETag must fail after the fenced replacement commits");
|
||||
assert_eq!(err, StorageError::PreconditionFailed);
|
||||
}
|
||||
|
||||
#[tokio::test]
|
||||
async fn repeated_body_write_keeps_etag_but_changes_data_dir_generation() {
|
||||
let set_disks = make_local_bucket_test_set_disks().await;
|
||||
let bucket = "bucket-write-generation";
|
||||
let object = "config/write-generation.json";
|
||||
let body = b"identical config body".to_vec();
|
||||
set_disks
|
||||
.make_bucket(bucket, &MakeBucketOptions::default())
|
||||
.await
|
||||
.expect("bucket should be created");
|
||||
|
||||
let mut first_reader = PutObjReader::from_vec(body.clone());
|
||||
let first = set_disks
|
||||
.put_object(
|
||||
bucket,
|
||||
object,
|
||||
&mut first_reader,
|
||||
&ObjectOptions {
|
||||
no_lock: true,
|
||||
..Default::default()
|
||||
},
|
||||
)
|
||||
.await
|
||||
.expect("first config body should be written");
|
||||
let mut second_reader = PutObjReader::from_vec(body);
|
||||
let second = set_disks
|
||||
.put_object(
|
||||
bucket,
|
||||
object,
|
||||
&mut second_reader,
|
||||
&ObjectOptions {
|
||||
no_lock: true,
|
||||
..Default::default()
|
||||
},
|
||||
)
|
||||
.await
|
||||
.expect("identical config body should be rewritten");
|
||||
|
||||
assert_eq!(first.etag, second.etag, "content ETag should expose the ABA collision");
|
||||
assert_ne!(first.data_dir, second.data_dir, "each committed body write needs a unique generation");
|
||||
assert!(first.data_dir.is_some() && second.data_dir.is_some());
|
||||
}
|
||||
|
||||
#[tokio::test]
|
||||
async fn conditional_create_holds_object_lock_through_rename() {
|
||||
let set_disks = make_local_bucket_test_set_disks().await;
|
||||
let bucket = "bucket-conditional-create-fence";
|
||||
let object = "config/conditional-create.json";
|
||||
set_disks
|
||||
.make_bucket(bucket, &MakeBucketOptions::default())
|
||||
.await
|
||||
.expect("bucket should be created");
|
||||
|
||||
let barrier = rename_fanout_barrier::arm(object, 0, rename_fanout_barrier::PHASE_RENAME);
|
||||
let writer_store = set_disks.clone();
|
||||
let writer = tokio::spawn(async move {
|
||||
let mut reader = PutObjReader::from_vec(b"created config".to_vec());
|
||||
writer_store
|
||||
.put_object(
|
||||
bucket,
|
||||
object,
|
||||
&mut reader,
|
||||
&ObjectOptions {
|
||||
http_preconditions: Some(HTTPPreconditions {
|
||||
if_none_match: Some("*".to_string()),
|
||||
..Default::default()
|
||||
}),
|
||||
..Default::default()
|
||||
},
|
||||
)
|
||||
.await
|
||||
});
|
||||
|
||||
tokio::time::timeout(std::time::Duration::from_secs(30), barrier.wait_until_paused())
|
||||
.await
|
||||
.expect("conditional create should reach the rename barrier");
|
||||
assert_exclusive_object_lock_held(&set_disks, bucket, object);
|
||||
barrier.release();
|
||||
writer
|
||||
.await
|
||||
.expect("conditional writer task should finish")
|
||||
.expect("first conditional create should commit");
|
||||
assert!(
|
||||
set_disks
|
||||
.local_lock_manager_for_test()
|
||||
.get_lock_info(&ObjectKey::new(bucket, object))
|
||||
.is_none(),
|
||||
"conditional create should release the object lock after commit"
|
||||
);
|
||||
let contender = set_disks
|
||||
.new_ns_lock(bucket, object)
|
||||
.await
|
||||
.expect("contender namespace lock should be created");
|
||||
let contender_guard = contender
|
||||
.get_write_lock(std::time::Duration::from_secs(30))
|
||||
.await
|
||||
.expect("contender should acquire after conditional create commits");
|
||||
drop(contender_guard);
|
||||
|
||||
let mut duplicate_reader = PutObjReader::from_vec(b"duplicate config".to_vec());
|
||||
let err = set_disks
|
||||
.put_object(
|
||||
bucket,
|
||||
object,
|
||||
&mut duplicate_reader,
|
||||
&ObjectOptions {
|
||||
http_preconditions: Some(HTTPPreconditions {
|
||||
if_none_match: Some("*".to_string()),
|
||||
..Default::default()
|
||||
}),
|
||||
..Default::default()
|
||||
},
|
||||
)
|
||||
.await
|
||||
.expect_err("a second create-only write must not replace the committed config");
|
||||
assert_eq!(err, StorageError::PreconditionFailed);
|
||||
}
|
||||
|
||||
#[tokio::test]
|
||||
async fn set_level_if_none_match_fails_closed_without_read_quorum() {
|
||||
let set_disks = make_local_bucket_test_set_disks_with_drive_count(4).await;
|
||||
|
||||
@@ -4861,10 +4861,7 @@ async fn poll_merge_head(rx: &CancellationToken, in_channels: &mut [Receiver<Met
|
||||
|
||||
async fn send_or_cancel(rx: &CancellationToken, out_channel: &Sender<MetaCacheEntry>, entry: MetaCacheEntry) -> Result<bool> {
|
||||
tokio::select! {
|
||||
result = out_channel.send(entry) => {
|
||||
result.map_err(Error::other)?;
|
||||
Ok(true)
|
||||
}
|
||||
result = out_channel.send(entry) => Ok(result.is_ok()),
|
||||
_ = rx.cancelled() => Ok(false),
|
||||
}
|
||||
}
|
||||
@@ -9858,6 +9855,28 @@ mod test {
|
||||
assert_eq!(results, vec!["obj-a", "obj-b"]);
|
||||
}
|
||||
|
||||
#[tokio::test]
|
||||
async fn merge_entry_channels_treats_dropped_output_receiver_as_completion() {
|
||||
let (tx_a, rx_a) = mpsc::channel(4);
|
||||
let (tx_b, rx_b) = mpsc::channel(4);
|
||||
let (out_tx, out_rx) = mpsc::channel(1);
|
||||
|
||||
tx_a.send(test_meta_entry("obj-a")).await.unwrap();
|
||||
tx_b.send(test_meta_entry("obj-b")).await.unwrap();
|
||||
drop(tx_a);
|
||||
drop(tx_b);
|
||||
drop(out_rx);
|
||||
|
||||
let result = timeout(
|
||||
Duration::from_secs(1),
|
||||
merge_entry_channels(CancellationToken::new(), vec![rx_a, rx_b], out_tx, 1),
|
||||
)
|
||||
.await
|
||||
.expect("merge should stop promptly when its consumer disconnects");
|
||||
|
||||
assert!(result.is_ok(), "consumer disconnect must not surface as a merge worker error");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn walk_ascending_versions_contract_reverses_newest_first_metadata() {
|
||||
// Documents the invariant the walk `versions_sort` handling relies on:
|
||||
|
||||
@@ -46,7 +46,10 @@ use zeroize::Zeroizing;
|
||||
/// `key_dir` is accepted, so identifiers already in use by existing deployments keep
|
||||
/// resolving. Only separators, traversal and the degenerate cases are refused, which is
|
||||
/// what stops `key_dir.join(...)` from escaping.
|
||||
fn validate_key_id(key_id: &str) -> Result<()> {
|
||||
///
|
||||
/// pub(crate) because the backup restore path applies the same containment
|
||||
/// rule to key identifiers recovered from bundle artifacts.
|
||||
pub(crate) fn validate_key_id(key_id: &str) -> Result<()> {
|
||||
if key_id.is_empty() {
|
||||
return Err(KmsError::invalid_key("key identifier must not be empty"));
|
||||
}
|
||||
@@ -68,7 +71,15 @@ fn validate_key_id(key_id: &str) -> Result<()> {
|
||||
}
|
||||
}
|
||||
|
||||
const LOCAL_KMS_MASTER_KEY_SALT_FILE: &str = ".master-key.salt";
|
||||
// The salt and restore-marker file names are pub(crate) so the backup/restore
|
||||
// modules (`crate::backup`) address the exact on-disk names instead of copies
|
||||
// that could drift.
|
||||
pub(crate) const LOCAL_KMS_MASTER_KEY_SALT_FILE: &str = ".master-key.salt";
|
||||
/// Commit marker of an in-progress Local restore cutover (see
|
||||
/// `crate::backup::local_restore`). Its presence means the key directory is
|
||||
/// mid-cutover: startup must fail closed until the restore is rolled forward
|
||||
/// or explicitly aborted.
|
||||
pub(crate) const LOCAL_RESTORE_COMMIT_MARKER_FILE: &str = ".restore-commit.json";
|
||||
// The KDF parameters are pub(crate) so the backup manifest contract
|
||||
// (`crate::backup`) records the exact compiled-in derivation instead of a
|
||||
// copy that could drift.
|
||||
@@ -87,7 +98,11 @@ pub(crate) const LOCAL_KMS_ARGON2_P_COST: u32 = 1;
|
||||
/// `foo.tmp-<uuid>` is stored as `foo.tmp-<uuid>.key` — so the `.key` guard
|
||||
/// plus the exact hyphenated-UUID check makes it impossible to match an
|
||||
/// authoritative file.
|
||||
fn is_orphan_commit_temp_name(file_name: &str) -> bool {
|
||||
///
|
||||
/// pub(crate) because the backup restore path applies the same classification
|
||||
/// when it re-enters an interrupted run: a leftover commit temp is never
|
||||
/// authoritative state, so it does not make a target non-empty.
|
||||
pub(crate) fn is_orphan_commit_temp_name(file_name: &str) -> bool {
|
||||
if file_name.ends_with(".key") {
|
||||
return false;
|
||||
}
|
||||
@@ -110,12 +125,15 @@ fn is_orphan_commit_temp_name(file_name: &str) -> bool {
|
||||
///
|
||||
/// This intentionally mirrors ecstore's fsync helpers without depending on the
|
||||
/// ecstore crate: the KMS backend stays decoupled from storage internals.
|
||||
mod durable_file {
|
||||
///
|
||||
/// pub(crate) because the backup restore path (`crate::backup::local_restore`)
|
||||
/// commits staged files and its cutover marker through the same protocol.
|
||||
pub(crate) mod durable_file {
|
||||
use std::io::{self, Write};
|
||||
use std::path::{Path, PathBuf};
|
||||
|
||||
/// How the fully written temp file becomes visible under its final name.
|
||||
pub(super) enum Publish {
|
||||
pub(crate) enum Publish {
|
||||
/// Atomically replace whatever is at the destination via `rename`.
|
||||
Replace,
|
||||
/// Publish via `hard_link`, failing with [`CommitError::AlreadyExists`]
|
||||
@@ -124,7 +142,7 @@ mod durable_file {
|
||||
}
|
||||
|
||||
#[derive(Debug)]
|
||||
pub(super) enum CommitError {
|
||||
pub(crate) enum CommitError {
|
||||
AlreadyExists,
|
||||
Io(io::Error),
|
||||
/// Test-only simulated crash: the protocol stops after the given step
|
||||
@@ -158,14 +176,14 @@ mod durable_file {
|
||||
/// to prove that every interrupted prefix recovers to either the complete
|
||||
/// old state or the complete new state.
|
||||
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
|
||||
pub(super) enum CommitStep {
|
||||
pub(crate) enum CommitStep {
|
||||
TempWritten,
|
||||
FileSynced,
|
||||
Published,
|
||||
DirSynced,
|
||||
}
|
||||
|
||||
pub(super) async fn commit(
|
||||
pub(crate) async fn commit(
|
||||
temp_path: PathBuf,
|
||||
final_path: PathBuf,
|
||||
content: Vec<u8>,
|
||||
@@ -179,7 +197,7 @@ mod durable_file {
|
||||
|
||||
/// Remove a published file durably: without the parent directory fsync a
|
||||
/// deleted key could resurface after power loss.
|
||||
pub(super) async fn remove_durably(path: PathBuf) -> io::Result<()> {
|
||||
pub(crate) async fn remove_durably(path: PathBuf) -> io::Result<()> {
|
||||
tokio::task::spawn_blocking(move || {
|
||||
std::fs::remove_file(&path)?;
|
||||
let parent = path
|
||||
@@ -191,6 +209,41 @@ mod durable_file {
|
||||
.map_err(io::Error::other)?
|
||||
}
|
||||
|
||||
/// Publish an already-durable file under a second name via `hard_link`,
|
||||
/// then fsync the destination's parent directory.
|
||||
///
|
||||
/// This is the restore cutover primitive: the source (a staged file that
|
||||
/// went through [`commit`]) is already durable, so linking plus a parent
|
||||
/// fsync is a complete publish. `AlreadyExists` is idempotent success only
|
||||
/// when the destination content is byte-identical to the source — that is
|
||||
/// exactly the re-entry case of a cutover interrupted after this link —
|
||||
/// and a hard failure otherwise, so the primitive can never clobber or
|
||||
/// silently accept foreign state.
|
||||
pub(crate) async fn link_durably(source: PathBuf, dest: PathBuf) -> io::Result<()> {
|
||||
tokio::task::spawn_blocking(move || {
|
||||
match std::fs::hard_link(&source, &dest) {
|
||||
Ok(()) => {}
|
||||
Err(error) if error.kind() == io::ErrorKind::AlreadyExists => {
|
||||
let existing = std::fs::read(&dest)?;
|
||||
let staged = std::fs::read(&source)?;
|
||||
if existing != staged {
|
||||
return Err(io::Error::new(
|
||||
io::ErrorKind::AlreadyExists,
|
||||
format!("destination {} already exists with different content", dest.display()),
|
||||
));
|
||||
}
|
||||
}
|
||||
Err(error) => return Err(error),
|
||||
}
|
||||
let parent = dest
|
||||
.parent()
|
||||
.ok_or_else(|| io::Error::other("destination has no parent directory"))?;
|
||||
fsync_dir(parent)
|
||||
})
|
||||
.await
|
||||
.map_err(io::Error::other)?
|
||||
}
|
||||
|
||||
fn commit_blocking(
|
||||
temp_path: &Path,
|
||||
final_path: &Path,
|
||||
@@ -355,7 +408,7 @@ mod durable_file {
|
||||
/// Test-only failpoints simulating a crash after a given commit step.
|
||||
/// Armed per directory so parallel tests never affect each other.
|
||||
#[cfg(test)]
|
||||
pub(super) mod failpoint {
|
||||
pub(crate) mod failpoint {
|
||||
use super::CommitStep;
|
||||
use std::path::{Path, PathBuf};
|
||||
use std::sync::Mutex;
|
||||
@@ -397,6 +450,20 @@ pub struct LocalKmsClient {
|
||||
/// Per-key write locks serializing read-modify-write updates within this
|
||||
/// process (see [`Self::lock_key_for_write`]).
|
||||
key_write_locks: Mutex<HashMap<String, Arc<tokio::sync::Mutex<()>>>>,
|
||||
/// Directory-wide writer fence for backup export (see
|
||||
/// [`Self::acquire_export_fence`]). Writers hold the read side; an export
|
||||
/// snapshot holds the write side so it observes a single-generation view.
|
||||
export_fence: Arc<tokio::sync::RwLock<()>>,
|
||||
}
|
||||
|
||||
/// Guard pairing the export-fence read lock with a per-key write mutex.
|
||||
///
|
||||
/// Dropping it releases both, so every existing `lock_key_for_write` call
|
||||
/// site participates in the export fence without changes.
|
||||
#[must_use]
|
||||
struct KeyWriteGuard {
|
||||
_fence: tokio::sync::OwnedRwLockReadGuard<()>,
|
||||
_key: tokio::sync::OwnedMutexGuard<()>,
|
||||
}
|
||||
|
||||
// pub(crate) so the backup contract tests can anchor the manifest's
|
||||
@@ -446,6 +513,11 @@ impl LocalKmsClient {
|
||||
debug!(path = ?config.key_dir, "KMS key directory created");
|
||||
}
|
||||
|
||||
// The restore-marker guard must run before anything else touches the
|
||||
// directory (in particular before salt load/creation): a directory
|
||||
// mid-cutover holds an arbitrary mix of old and new state.
|
||||
Self::ensure_no_restore_marker(&config).await?;
|
||||
|
||||
// Initialize master cipher if master key is provided
|
||||
let (master_cipher, legacy_master_cipher) = if let Some(ref master_key) = config.master_key {
|
||||
let salt = Self::load_or_create_master_key_salt(&config).await?;
|
||||
@@ -463,6 +535,7 @@ impl LocalKmsClient {
|
||||
legacy_master_cipher,
|
||||
dek_crypto: AesDekCrypto::new(),
|
||||
key_write_locks: Mutex::new(HashMap::new()),
|
||||
export_fence: Arc::new(tokio::sync::RwLock::new(())),
|
||||
};
|
||||
client.validate_existing_keys().await?;
|
||||
Ok(client)
|
||||
@@ -476,6 +549,7 @@ impl LocalKmsClient {
|
||||
if !fs::try_exists(&config.key_dir).await? {
|
||||
return Err(KmsError::configuration_error("Local KMS key directory does not exist"));
|
||||
}
|
||||
Self::ensure_no_restore_marker(&config).await?;
|
||||
|
||||
let (master_cipher, legacy_master_cipher) = if let Some(ref master_key) = config.master_key {
|
||||
let legacy_key = Self::derive_legacy_master_key(master_key)?;
|
||||
@@ -505,6 +579,7 @@ impl LocalKmsClient {
|
||||
legacy_master_cipher,
|
||||
dek_crypto: AesDekCrypto::new(),
|
||||
key_write_locks: Mutex::new(HashMap::new()),
|
||||
export_fence: Arc::new(tokio::sync::RwLock::new(())),
|
||||
})
|
||||
}
|
||||
|
||||
@@ -515,16 +590,72 @@ impl LocalKmsClient {
|
||||
/// delete with a rewrite. Cross-process writers sharing a key directory
|
||||
/// remain unsupported. Entries live for the client's lifetime; the table
|
||||
/// is bounded by the number of distinct key ids this process touches.
|
||||
async fn lock_key_for_write(&self, key_id: &str) -> tokio::sync::OwnedMutexGuard<()> {
|
||||
async fn lock_key_for_write(&self, key_id: &str) -> KeyWriteGuard {
|
||||
// Fence first, per-key mutex second: the ordering is uniform across
|
||||
// all writers, so an export waiting on the write side can never
|
||||
// deadlock with a writer holding a key mutex.
|
||||
let fence = Arc::clone(&self.export_fence).read_owned().await;
|
||||
let lock = {
|
||||
let mut locks = self.key_write_locks.lock().expect("Local KMS key write lock table poisoned");
|
||||
Arc::clone(locks.entry(key_id.to_string()).or_default())
|
||||
};
|
||||
lock.lock_owned().await
|
||||
KeyWriteGuard {
|
||||
_fence: fence,
|
||||
_key: lock.lock_owned().await,
|
||||
}
|
||||
}
|
||||
|
||||
/// Block every key-directory writer while a backup export collects its
|
||||
/// snapshot, so all records belong to one generation.
|
||||
///
|
||||
/// Mutating operations hold the read side (via [`Self::lock_key_for_write`]
|
||||
/// or [`Self::save_new_master_key`]); the export holds the write side only
|
||||
/// for the collection phase, never while encrypting or writing the bundle.
|
||||
pub(crate) async fn acquire_export_fence(&self) -> tokio::sync::OwnedRwLockWriteGuard<()> {
|
||||
Arc::clone(&self.export_fence).write_owned().await
|
||||
}
|
||||
|
||||
/// Key directory root, exposed for the backup export module.
|
||||
pub(crate) fn key_directory(&self) -> &Path {
|
||||
&self.config.key_dir
|
||||
}
|
||||
|
||||
/// Absolute path of the master-key KDF salt file, exposed for the backup
|
||||
/// export module.
|
||||
pub(crate) fn master_key_salt_file(&self) -> PathBuf {
|
||||
Self::master_key_salt_path(&self.config)
|
||||
}
|
||||
|
||||
/// Operator-configured master key string, exposed for the backup export
|
||||
/// module so it can record a one-way verifier in the bundle manifest.
|
||||
/// Never log or persist this value.
|
||||
pub(crate) fn configured_master_key(&self) -> Option<&str> {
|
||||
self.config.master_key.as_deref()
|
||||
}
|
||||
|
||||
/// Fail closed while a restore cutover marker is present: the directory
|
||||
/// then holds an arbitrary mix of pre-restore and restored state, and the
|
||||
/// only valid next steps are re-running the restore with the same bundle
|
||||
/// (roll forward) or explicitly aborting it. This mirrors the missing-salt
|
||||
/// guard: startup must never paper over a half-applied restore.
|
||||
async fn ensure_no_restore_marker(config: &LocalConfig) -> Result<()> {
|
||||
let marker = config.key_dir.join(LOCAL_RESTORE_COMMIT_MARKER_FILE);
|
||||
if fs::try_exists(&marker).await? {
|
||||
return Err(KmsError::configuration_error(format!(
|
||||
"Local KMS key directory has an unfinished restore (marker {} present); \
|
||||
re-run the restore with the same bundle to roll it forward or abort it explicitly",
|
||||
marker.display()
|
||||
)));
|
||||
}
|
||||
Ok(())
|
||||
}
|
||||
|
||||
/// Derive a 256-bit key from the master key string using a persistent Argon2id salt.
|
||||
fn derive_master_key(master_key: &str, salt: &[u8]) -> Result<Key<Aes256Gcm>> {
|
||||
///
|
||||
/// pub(crate) because the backup restore path derives the same key from
|
||||
/// the operator-supplied master key and the bundled salt for its verifier
|
||||
/// check and staged decryption probe.
|
||||
pub(crate) fn derive_master_key(master_key: &str, salt: &[u8]) -> Result<Key<Aes256Gcm>> {
|
||||
let params = Params::new(
|
||||
LOCAL_KMS_ARGON2_M_COST_KIB,
|
||||
LOCAL_KMS_ARGON2_T_COST,
|
||||
@@ -541,7 +672,7 @@ impl LocalKmsClient {
|
||||
Ok(key)
|
||||
}
|
||||
|
||||
fn derive_legacy_master_key(master_key: &str) -> Result<Key<Aes256Gcm>> {
|
||||
pub(crate) fn derive_legacy_master_key(master_key: &str) -> Result<Key<Aes256Gcm>> {
|
||||
let mut hasher = Sha256::new();
|
||||
hasher.update(master_key.as_bytes());
|
||||
hasher.update(b"rustfs-kms-local");
|
||||
@@ -797,6 +928,11 @@ impl LocalKmsClient {
|
||||
}
|
||||
|
||||
async fn save_new_master_key(&self, master_key: &MasterKeyInfo, key_material: &[u8]) -> Result<()> {
|
||||
// Creates never take the per-key write lock (`NoClobber` publishing
|
||||
// already linearizes them), so they join the export fence here. This
|
||||
// must stay the only fence acquisition on the create path: the fence
|
||||
// read lock is not reentrant while an export waits for the write side.
|
||||
let _fence = Arc::clone(&self.export_fence).read_owned().await;
|
||||
let key_path = self.master_key_path(&master_key.key_id)?;
|
||||
let content = self.encode_master_key(master_key, key_material)?;
|
||||
let temp_path = key_path.with_extension(format!("tmp-{}", uuid::Uuid::new_v4()));
|
||||
@@ -2404,6 +2540,34 @@ mod tests {
|
||||
assert!(!is_orphan_commit_temp_name("mykey.tmp-"));
|
||||
}
|
||||
|
||||
#[tokio::test]
|
||||
async fn link_durably_publishes_no_clobber_and_is_content_idempotent() {
|
||||
let temp_dir = TempDir::new().expect("temp dir");
|
||||
let source = temp_dir.path().join("staged");
|
||||
let dest = temp_dir.path().join("published");
|
||||
fs::write(&source, b"staged-content").await.expect("write source");
|
||||
|
||||
durable_file::link_durably(source.clone(), dest.clone())
|
||||
.await
|
||||
.expect("first link must succeed");
|
||||
assert_eq!(fs::read(&dest).await.expect("read dest"), b"staged-content");
|
||||
|
||||
// Re-entry with identical content is idempotent success — exactly the
|
||||
// resumed-cutover case.
|
||||
durable_file::link_durably(source.clone(), dest.clone())
|
||||
.await
|
||||
.expect("re-linking identical content must be idempotent");
|
||||
|
||||
// Existing content that differs is a hard failure, never a clobber.
|
||||
let foreign = temp_dir.path().join("foreign");
|
||||
fs::write(&foreign, b"different-content").await.expect("write foreign");
|
||||
let error = durable_file::link_durably(foreign, dest.clone())
|
||||
.await
|
||||
.expect_err("differing content must not be clobbered");
|
||||
assert_eq!(error.kind(), std::io::ErrorKind::AlreadyExists);
|
||||
assert_eq!(fs::read(&dest).await.expect("dest unchanged"), b"staged-content");
|
||||
}
|
||||
|
||||
#[tokio::test]
|
||||
async fn durable_commit_fsyncs_every_write_path() {
|
||||
use durable_file::fsync_recorder;
|
||||
|
||||
@@ -26,16 +26,16 @@ use std::sync::{Arc, Mutex};
|
||||
use tokio::io::{AsyncReadExt, AsyncWriteExt};
|
||||
use tokio::net::{TcpListener, TcpStream};
|
||||
|
||||
/// One canned HTTP response.
|
||||
pub(crate) struct ScriptedResponse {
|
||||
status: u16,
|
||||
body: String,
|
||||
/// One scripted connection outcome.
|
||||
pub(crate) enum ScriptedResponse {
|
||||
Http { status: u16, body: String },
|
||||
Close,
|
||||
}
|
||||
|
||||
impl ScriptedResponse {
|
||||
/// A 200 response carrying `data` inside the standard Vault envelope.
|
||||
pub(crate) fn ok(data: serde_json::Value) -> Self {
|
||||
Self {
|
||||
Self::Http {
|
||||
status: 200,
|
||||
body: serde_json::json!({
|
||||
"request_id": "scripted",
|
||||
@@ -50,18 +50,23 @@ impl ScriptedResponse {
|
||||
|
||||
/// An error response in Vault's `{"errors": [...]}` format.
|
||||
pub(crate) fn error(status: u16, message: &str) -> Self {
|
||||
Self {
|
||||
Self::Http {
|
||||
status,
|
||||
body: serde_json::json!({ "errors": [message] }).to_string(),
|
||||
}
|
||||
}
|
||||
|
||||
/// Close the connection after consuming a request without sending an HTTP response.
|
||||
pub(crate) fn close() -> Self {
|
||||
Self::Close
|
||||
}
|
||||
}
|
||||
|
||||
/// A scripted stand-in Vault listening on a loopback port.
|
||||
pub(crate) struct ScriptedVault {
|
||||
/// Base address (`http://127.0.0.1:port`) to point a Vault client at.
|
||||
pub(crate) address: String,
|
||||
requests: Arc<Mutex<Vec<String>>>,
|
||||
requests: Arc<Mutex<Vec<(String, String)>>>,
|
||||
}
|
||||
|
||||
impl ScriptedVault {
|
||||
@@ -81,24 +86,21 @@ impl ScriptedVault {
|
||||
let Ok((mut stream, _)) = listener.accept().await else {
|
||||
return;
|
||||
};
|
||||
let Some(request_line) = read_request(&mut stream).await else {
|
||||
let Some(request) = read_request(&mut stream).await else {
|
||||
continue;
|
||||
};
|
||||
recorded
|
||||
.lock()
|
||||
.expect("scripted vault request log poisoned")
|
||||
.push(request_line);
|
||||
recorded.lock().expect("scripted vault request log poisoned").push(request);
|
||||
let response = responses
|
||||
.next()
|
||||
.unwrap_or_else(|| ScriptedResponse::error(599, "scripted vault: script exhausted"));
|
||||
let payload = format!(
|
||||
"HTTP/1.1 {} Scripted\r\ncontent-type: application/json\r\ncontent-length: {}\r\nconnection: close\r\n\r\n{}",
|
||||
response.status,
|
||||
response.body.len(),
|
||||
response.body
|
||||
);
|
||||
let _ = stream.write_all(payload.as_bytes()).await;
|
||||
let _ = stream.shutdown().await;
|
||||
if let ScriptedResponse::Http { status, body } = response {
|
||||
let payload = format!(
|
||||
"HTTP/1.1 {status} Scripted\r\ncontent-type: application/json\r\ncontent-length: {}\r\nconnection: close\r\n\r\n{body}",
|
||||
body.len(),
|
||||
);
|
||||
let _ = stream.write_all(payload.as_bytes()).await;
|
||||
let _ = stream.shutdown().await;
|
||||
}
|
||||
}
|
||||
});
|
||||
|
||||
@@ -107,14 +109,32 @@ impl ScriptedVault {
|
||||
|
||||
/// The `METHOD /path` lines of every request served so far, in order.
|
||||
pub(crate) fn requests(&self) -> Vec<String> {
|
||||
self.requests.lock().expect("scripted vault request log poisoned").clone()
|
||||
self.requests
|
||||
.lock()
|
||||
.expect("scripted vault request log poisoned")
|
||||
.iter()
|
||||
.map(|(line, _)| line.clone())
|
||||
.collect()
|
||||
}
|
||||
|
||||
/// The request bodies, in the same order as [`Self::requests`]; empty for
|
||||
/// bodyless requests. Lets tests assert what a write actually persisted
|
||||
/// (record contents, check-and-set options), not just that a write happened.
|
||||
pub(crate) fn request_bodies(&self) -> Vec<String> {
|
||||
self.requests
|
||||
.lock()
|
||||
.expect("scripted vault request log poisoned")
|
||||
.iter()
|
||||
.map(|(_, body)| body.clone())
|
||||
.collect()
|
||||
}
|
||||
}
|
||||
|
||||
/// Read one HTTP/1.1 request (head plus content-length body) and return its
|
||||
/// `METHOD /path` line. Draining the body before responding keeps the client
|
||||
/// from seeing a connection reset while it is still writing.
|
||||
async fn read_request(stream: &mut TcpStream) -> Option<String> {
|
||||
/// `METHOD /path` line together with the body. Draining the body before
|
||||
/// responding keeps the client from seeing a connection reset while it is
|
||||
/// still writing.
|
||||
async fn read_request(stream: &mut TcpStream) -> Option<(String, String)> {
|
||||
let mut buffer = Vec::new();
|
||||
let mut chunk = [0u8; 4096];
|
||||
let head_end = loop {
|
||||
@@ -146,14 +166,17 @@ async fn read_request(stream: &mut TcpStream) -> Option<String> {
|
||||
})
|
||||
.next()
|
||||
.unwrap_or(0);
|
||||
let mut remaining = content_length.saturating_sub(buffer.len() - head_end);
|
||||
let mut body = buffer[head_end..].to_vec();
|
||||
let mut remaining = content_length.saturating_sub(body.len());
|
||||
while remaining > 0 {
|
||||
let read = stream.read(&mut chunk).await.ok()?;
|
||||
if read == 0 {
|
||||
break;
|
||||
}
|
||||
body.extend_from_slice(&chunk[..read]);
|
||||
remaining = remaining.saturating_sub(read);
|
||||
}
|
||||
body.truncate(content_length);
|
||||
|
||||
Some(format!("{method} {path}"))
|
||||
Some((format!("{method} {path}"), String::from_utf8_lossy(&body).into_owned()))
|
||||
}
|
||||
|
||||
+936
-154
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
@@ -306,6 +306,141 @@ impl LocalKdfDescriptor {
|
||||
}
|
||||
}
|
||||
|
||||
/// Transit engine reference recorded in a Vault-backed bundle.
|
||||
///
|
||||
/// Transit keys are non-exportable, so a bundle can only ever name the key and
|
||||
/// the version window its content depends on. Restore compares these values
|
||||
/// against the Transit key the operator's native Vault restore produced.
|
||||
#[derive(Debug, Clone, PartialEq, Eq, Serialize, Deserialize)]
|
||||
#[serde(deny_unknown_fields)]
|
||||
pub struct VaultTransitReference {
|
||||
/// Transit engine mount path.
|
||||
pub mount_path: String,
|
||||
/// Transit key name.
|
||||
pub key_name: String,
|
||||
/// Lowest Transit key version the bundle's content still needs. A target
|
||||
/// whose `min_decryption_version` sits above this can no longer decrypt
|
||||
/// the oldest state in the bundle.
|
||||
pub required_min_version: u32,
|
||||
/// Latest Transit key version at snapshot time. A target whose newest
|
||||
/// version is below this was restored to a point before the bundle.
|
||||
pub current_version: u32,
|
||||
/// Operator-recorded immutable reference to the native Vault/HSM snapshot
|
||||
/// protecting this key. RustFS neither produces nor consumes that
|
||||
/// snapshot; the reference exists so restore evidence can name it.
|
||||
#[serde(default, skip_serializing_if = "Option::is_none")]
|
||||
pub native_snapshot_reference: Option<String>,
|
||||
}
|
||||
|
||||
/// KV generation of one key's record at snapshot time.
|
||||
#[derive(Debug, Clone, PartialEq, Eq, Serialize, Deserialize)]
|
||||
#[serde(deny_unknown_fields)]
|
||||
pub struct VaultKvRecordReference {
|
||||
/// Stable key id the record belongs to.
|
||||
pub key_id: String,
|
||||
/// KV2 secret version holding the record when the snapshot was taken.
|
||||
/// Restore compares the target's current version against it: a lower
|
||||
/// value means the target's Vault was restored to a point before the
|
||||
/// bundle, a higher value means the target moved ahead of it.
|
||||
pub kv_version: u64,
|
||||
}
|
||||
|
||||
/// Immutable references to the external Vault state a bundle depends on.
|
||||
///
|
||||
/// A Vault-backed bundle never carries the cryptographic root: Transit keys
|
||||
/// cannot be exported and KV2 records live in Vault's own storage. What the
|
||||
/// bundle can own is a precise description of *which* external state it was
|
||||
/// captured against, so a restore refuses to proceed when the operator's
|
||||
/// native Vault restore landed somewhere else. Credentials are structurally
|
||||
/// absent — no token, AppRole secret id, or TLS material is representable in
|
||||
/// this type.
|
||||
#[derive(Debug, Clone, PartialEq, Eq, Serialize, Deserialize)]
|
||||
#[serde(deny_unknown_fields)]
|
||||
pub struct VaultExternalReferences {
|
||||
/// Opaque identity of the Vault cluster the snapshot was taken from.
|
||||
pub cluster_id: String,
|
||||
/// Vault Enterprise namespace, when the deployment uses one.
|
||||
#[serde(default, skip_serializing_if = "Option::is_none")]
|
||||
pub namespace: Option<String>,
|
||||
/// KV v2 mount holding the RustFS records.
|
||||
pub kv_mount: String,
|
||||
/// Path prefix under `kv_mount` holding the RustFS records.
|
||||
pub kv_path_prefix: String,
|
||||
/// Transit engine reference; present exactly when the bundle's material
|
||||
/// is protected by a Transit key.
|
||||
#[serde(default, skip_serializing_if = "Option::is_none")]
|
||||
pub transit: Option<VaultTransitReference>,
|
||||
/// Per-key KV generation observed at snapshot time, one entry per key the
|
||||
/// bundle carries a record for.
|
||||
pub kv_records: Vec<VaultKvRecordReference>,
|
||||
}
|
||||
|
||||
impl VaultExternalReferences {
|
||||
fn validate(&self, backend: BackupBackendKind, protection: AtRestProtection) -> Result<(), BackupError> {
|
||||
require_non_empty("external_references.cluster_id", &self.cluster_id)?;
|
||||
require_non_empty("external_references.kv_mount", &self.kv_mount)?;
|
||||
require_non_empty("external_references.kv_path_prefix", &self.kv_path_prefix)?;
|
||||
if self.namespace.as_deref() == Some("") {
|
||||
return Err(BackupError::corrupted("external Vault namespace must not be empty when present"));
|
||||
}
|
||||
|
||||
// The Transit reference is required exactly where the cryptographic
|
||||
// root lives in Transit; a storage-only KV2 bundle that claims one
|
||||
// would misdescribe its own trust root.
|
||||
let transit_required = matches!(
|
||||
(backend, protection),
|
||||
(BackupBackendKind::VaultTransit, _) | (BackupBackendKind::VaultKv2, AtRestProtection::TransitWrapped)
|
||||
);
|
||||
match (&self.transit, transit_required) {
|
||||
(None, true) => {
|
||||
return Err(BackupError::corrupted(format!(
|
||||
"({backend:?}, {protection:?}) bundles must reference the Transit key protecting them"
|
||||
)));
|
||||
}
|
||||
(Some(_), false) => {
|
||||
return Err(BackupError::corrupted(format!(
|
||||
"({backend:?}, {protection:?}) bundles have no Transit trust root to reference"
|
||||
)));
|
||||
}
|
||||
(Some(transit), true) => {
|
||||
require_non_empty("external_references.transit.mount_path", &transit.mount_path)?;
|
||||
require_non_empty("external_references.transit.key_name", &transit.key_name)?;
|
||||
if transit.required_min_version == 0 {
|
||||
return Err(BackupError::corrupted("Transit reference must require at least key version 1"));
|
||||
}
|
||||
if transit.current_version < transit.required_min_version {
|
||||
return Err(BackupError::corrupted(format!(
|
||||
"Transit reference declares current version {} below the required minimum {}",
|
||||
transit.current_version, transit.required_min_version
|
||||
)));
|
||||
}
|
||||
if transit.native_snapshot_reference.as_deref() == Some("") {
|
||||
return Err(BackupError::corrupted("Transit native snapshot reference must not be empty when present"));
|
||||
}
|
||||
}
|
||||
(None, false) => {}
|
||||
}
|
||||
|
||||
let mut seen = BTreeSet::new();
|
||||
for record in &self.kv_records {
|
||||
require_non_empty("external_references.kv_records[].key_id", &record.key_id)?;
|
||||
if record.kv_version == 0 {
|
||||
return Err(BackupError::corrupted(format!(
|
||||
"KV generation for key '{}' must be at least 1",
|
||||
record.key_id
|
||||
)));
|
||||
}
|
||||
if !seen.insert(record.key_id.as_str()) {
|
||||
return Err(BackupError::corrupted(format!(
|
||||
"KV reference for key '{}' appears more than once",
|
||||
record.key_id
|
||||
)));
|
||||
}
|
||||
}
|
||||
Ok(())
|
||||
}
|
||||
}
|
||||
|
||||
/// Completeness marker of a bundle.
|
||||
///
|
||||
/// A producer writes `in-progress` state (or no manifest at all) until the
|
||||
@@ -355,8 +490,10 @@ struct ManifestProbe {
|
||||
///
|
||||
/// The manifest is the authoritative description of one backup bundle: what
|
||||
/// was captured, under which snapshot generation, protected by which backup
|
||||
/// KEK, and which restore responsibility applies. Field order is part of the
|
||||
/// canonical digest form and is frozen for this format version.
|
||||
/// KEK, and which restore responsibility applies. The digest's canonical
|
||||
/// form is the manifest's JSON value with the digest hex emptied (see
|
||||
/// [`Self::compute_digest`]); decoders verify it against the raw stored
|
||||
/// bytes and never re-serialize parsed fields.
|
||||
#[derive(Debug, Clone, PartialEq, Eq, Serialize, Deserialize)]
|
||||
#[serde(deny_unknown_fields)]
|
||||
pub struct BackupManifest {
|
||||
@@ -392,6 +529,10 @@ pub struct BackupManifest {
|
||||
/// Local backend section; required exactly when `backend` is `Local`.
|
||||
#[serde(default, skip_serializing_if = "Option::is_none")]
|
||||
pub local_kdf: Option<LocalKdfDescriptor>,
|
||||
/// External Vault references; required exactly when `backend` is a Vault
|
||||
/// backend. Never carries credentials — see [`VaultExternalReferences`].
|
||||
#[serde(default, skip_serializing_if = "Option::is_none")]
|
||||
pub external_references: Option<VaultExternalReferences>,
|
||||
/// Reserved for the per-key version/envelope inventory defined by the
|
||||
/// backlog#1565 contract. May not carry data in format version 1.
|
||||
#[serde(default, skip_serializing_if = "Option::is_none")]
|
||||
@@ -415,8 +556,14 @@ impl BackupManifest {
|
||||
///
|
||||
/// Fail-closed order: truncated or malformed input, then unknown format
|
||||
/// version, then a missing completeness marker, then schema decoding
|
||||
/// (unknown fields, duplicate fields, missing fields), then semantic
|
||||
/// validation including digest verification.
|
||||
/// (unknown fields, duplicate fields, missing fields), then digest
|
||||
/// verification against the raw input bytes, then semantic validation.
|
||||
///
|
||||
/// The digest is verified against the bytes as stored — parsed typed
|
||||
/// fields are never re-serialized for verification, so a field whose
|
||||
/// string form does not round-trip byte-identically through its parsed
|
||||
/// representation (timestamps in environment-dependent time zone
|
||||
/// spellings, for example) cannot produce a spurious mismatch.
|
||||
pub fn decode(bytes: &[u8]) -> Result<Self, BackupError> {
|
||||
let probe: ManifestProbe = serde_json::from_slice(bytes).map_err(map_serde_error)?;
|
||||
if probe.format_version != Self::FORMAT_VERSION {
|
||||
@@ -429,7 +576,9 @@ impl BackupManifest {
|
||||
return Err(BackupError::incomplete_bundle("manifest has no completeness marker"));
|
||||
}
|
||||
let manifest: Self = serde_json::from_slice(bytes).map_err(map_serde_error)?;
|
||||
manifest.validate()?;
|
||||
manifest.validate_pre_digest()?;
|
||||
Self::verify_digest_in_bytes(bytes, &manifest.manifest_digest)?;
|
||||
manifest.validate_content()?;
|
||||
Ok(manifest)
|
||||
}
|
||||
|
||||
@@ -449,23 +598,30 @@ impl BackupManifest {
|
||||
Ok(self)
|
||||
}
|
||||
|
||||
/// Compute the digest over the canonical manifest bytes.
|
||||
/// Compute the digest over the canonical manifest form.
|
||||
///
|
||||
/// Canonical form: compact JSON serialization of this manifest with the
|
||||
/// digest hex emptied. Field order is struct declaration order and is
|
||||
/// frozen for format version 1, so the same manifest content always
|
||||
/// hashes to the same value.
|
||||
/// Canonical form: the JSON *value* of the manifest with the digest hex
|
||||
/// emptied, object keys rebuilt in bytewise-sorted order at every level
|
||||
/// (see [`canonicalize_value`]), then serialized compactly. The value
|
||||
/// layer is what makes sealing and decoding agree byte-for-byte — a
|
||||
/// decoder recovers the identical value from the raw stored bytes
|
||||
/// without round-tripping any typed field through parse-and-reprint —
|
||||
/// and the explicit key sort makes the bytes independent of
|
||||
/// `serde_json`'s map implementation (`preserve_order` on or off).
|
||||
pub fn compute_digest(&self) -> Result<ContentDigest, BackupError> {
|
||||
let mut unsealed = self.clone();
|
||||
unsealed.manifest_digest = ContentDigest::placeholder(self.manifest_digest.algorithm);
|
||||
let canonical = serde_json::to_vec(&unsealed)
|
||||
let value = serde_json::to_value(&unsealed)
|
||||
.map_err(|error| BackupError::corrupted(format!("manifest canonicalization failed: {error}")))?;
|
||||
match self.manifest_digest.algorithm {
|
||||
DigestAlgorithm::Sha256 => Ok(ContentDigest::sha256_of(&canonical)),
|
||||
}
|
||||
Self::digest_of_canonical_value(value, self.manifest_digest.algorithm)
|
||||
}
|
||||
|
||||
/// Verify the sealed digest against the current manifest content.
|
||||
/// Verify the sealed digest against the current in-memory content.
|
||||
///
|
||||
/// This is the producer-side check (sealing and [`Self::encode`]).
|
||||
/// Decoders must use the raw stored bytes instead (see [`Self::decode`]):
|
||||
/// re-serializing parsed fields is not guaranteed to reproduce the
|
||||
/// stored spelling byte-for-byte.
|
||||
pub fn verify_digest(&self) -> Result<(), BackupError> {
|
||||
if !self.manifest_digest.is_well_formed() {
|
||||
return Err(BackupError::corrupted("manifest digest is not a well-formed digest value"));
|
||||
@@ -478,6 +634,33 @@ impl BackupManifest {
|
||||
Ok(())
|
||||
}
|
||||
|
||||
/// Verify a declared digest against raw manifest bytes, normalizing only
|
||||
/// through the JSON value layer and emptying the digest slot in place.
|
||||
fn verify_digest_in_bytes(bytes: &[u8], declared: &ContentDigest) -> Result<(), BackupError> {
|
||||
if !declared.is_well_formed() {
|
||||
return Err(BackupError::corrupted("manifest digest is not a well-formed digest value"));
|
||||
}
|
||||
let mut value: serde_json::Value = serde_json::from_slice(bytes).map_err(map_serde_error)?;
|
||||
let Some(slot) = value.get_mut("manifest_digest").and_then(|digest| digest.get_mut("hex")) else {
|
||||
return Err(BackupError::corrupted("manifest has no digest slot"));
|
||||
};
|
||||
*slot = serde_json::Value::String(String::new());
|
||||
if Self::digest_of_canonical_value(value, declared.algorithm)? != *declared {
|
||||
return Err(BackupError::corrupted(
|
||||
"manifest digest mismatch: content does not match the sealed digest",
|
||||
));
|
||||
}
|
||||
Ok(())
|
||||
}
|
||||
|
||||
fn digest_of_canonical_value(value: serde_json::Value, algorithm: DigestAlgorithm) -> Result<ContentDigest, BackupError> {
|
||||
let canonical = serde_json::to_vec(&canonicalize_value(value))
|
||||
.map_err(|error| BackupError::corrupted(format!("manifest canonicalization failed: {error}")))?;
|
||||
match algorithm {
|
||||
DigestAlgorithm::Sha256 => Ok(ContentDigest::sha256_of(&canonical)),
|
||||
}
|
||||
}
|
||||
|
||||
/// Look up a required artifact by kind, failing closed when absent.
|
||||
pub fn require_artifact(&self, kind: ArtifactKind) -> Result<&ArtifactDescriptor, BackupError> {
|
||||
self.artifacts
|
||||
@@ -486,11 +669,21 @@ impl BackupManifest {
|
||||
.ok_or_else(|| BackupError::missing_artifact(artifact_kind_name(kind)))
|
||||
}
|
||||
|
||||
/// Validate the full manifest contract.
|
||||
/// Validate the full manifest contract against the in-memory content.
|
||||
///
|
||||
/// This is decode-side validation and also guards [`Self::encode`], so a
|
||||
/// producer cannot publish a manifest a decoder would reject.
|
||||
/// This guards [`Self::encode`], so a producer cannot publish a manifest
|
||||
/// a decoder would reject. [`Self::decode`] runs the same checks but
|
||||
/// verifies the digest against the raw input bytes instead.
|
||||
pub fn validate(&self) -> Result<(), BackupError> {
|
||||
self.validate_pre_digest()?;
|
||||
self.verify_digest()?;
|
||||
self.validate_content()
|
||||
}
|
||||
|
||||
/// Checks that must run before any digest verification: an unknown
|
||||
/// version or an unsealed bundle is reported as its own typed error, not
|
||||
/// as a digest mismatch.
|
||||
fn validate_pre_digest(&self) -> Result<(), BackupError> {
|
||||
if self.format_version != Self::FORMAT_VERSION {
|
||||
return Err(BackupError::UnknownVersion {
|
||||
found: self.format_version,
|
||||
@@ -500,7 +693,12 @@ impl BackupManifest {
|
||||
if self.completeness != CompletenessState::Complete {
|
||||
return Err(BackupError::incomplete_bundle("completeness marker records an in-progress bundle"));
|
||||
}
|
||||
self.verify_digest()?;
|
||||
Ok(())
|
||||
}
|
||||
|
||||
/// Semantic validation of everything except version, completeness, and
|
||||
/// digest integrity.
|
||||
fn validate_content(&self) -> Result<(), BackupError> {
|
||||
require_non_empty("backup_id", &self.backup_id)?;
|
||||
require_non_empty("rustfs_version", &self.rustfs_version)?;
|
||||
require_non_empty("deployment_identity", &self.deployment_identity)?;
|
||||
@@ -513,6 +711,7 @@ impl BackupManifest {
|
||||
}
|
||||
self.validate_responsibility()?;
|
||||
self.validate_local_kdf()?;
|
||||
self.validate_external_references()?;
|
||||
self.validate_artifacts()?;
|
||||
Ok(())
|
||||
}
|
||||
@@ -542,6 +741,19 @@ impl BackupManifest {
|
||||
}
|
||||
}
|
||||
|
||||
fn validate_external_references(&self) -> Result<(), BackupError> {
|
||||
let vault_backend = matches!(self.backend, BackupBackendKind::VaultKv2 | BackupBackendKind::VaultTransit);
|
||||
match (&self.external_references, vault_backend) {
|
||||
(None, true) => Err(BackupError::corrupted("Vault backend manifest must carry external Vault references")),
|
||||
(Some(_), false) => Err(BackupError::corrupted(format!(
|
||||
"external Vault references are not valid for backend {:?}",
|
||||
self.backend
|
||||
))),
|
||||
(Some(references), true) => references.validate(self.backend, self.at_rest_protection),
|
||||
(None, false) => Ok(()),
|
||||
}
|
||||
}
|
||||
|
||||
fn validate_artifacts(&self) -> Result<(), BackupError> {
|
||||
let mut paths = BTreeSet::new();
|
||||
for artifact in &self.artifacts {
|
||||
@@ -645,6 +857,29 @@ fn map_serde_error(error: serde_json::Error) -> BackupError {
|
||||
}
|
||||
}
|
||||
|
||||
/// Rebuild a JSON value with object keys in bytewise-sorted order at every
|
||||
/// nesting level (array element order is preserved).
|
||||
///
|
||||
/// `serde_json`'s map keeps keys sorted by default but preserves insertion
|
||||
/// order when the `preserve_order` feature is unified into the build by any
|
||||
/// other crate. Digest bytes must not depend on that, so the ordering is
|
||||
/// imposed explicitly here instead of being inherited from the map type.
|
||||
fn canonicalize_value(value: serde_json::Value) -> serde_json::Value {
|
||||
match value {
|
||||
serde_json::Value::Object(map) => {
|
||||
let mut entries: Vec<(String, serde_json::Value)> = map.into_iter().collect();
|
||||
entries.sort_by(|a, b| a.0.cmp(&b.0));
|
||||
let mut sorted = serde_json::Map::with_capacity(entries.len());
|
||||
for (key, entry) in entries {
|
||||
sorted.insert(key, canonicalize_value(entry));
|
||||
}
|
||||
serde_json::Value::Object(sorted)
|
||||
}
|
||||
serde_json::Value::Array(items) => serde_json::Value::Array(items.into_iter().map(canonicalize_value).collect()),
|
||||
other => other,
|
||||
}
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
mod tests {
|
||||
use super::*;
|
||||
@@ -736,7 +971,7 @@ mod tests {
|
||||
/// canonical form and frozen. If serialization layout or field order
|
||||
/// changes, this value changes and the fixture test fails — which is the
|
||||
/// point: that is a format-version bump, not a patch.
|
||||
const FIXTURE_DIGEST_HEX: &str = "01accb3e2bc51e12d17d1efc52ad6ab9441c50bec52c3fb29e0a9f92b725cdaa";
|
||||
const FIXTURE_DIGEST_HEX: &str = "a8b104d61c358cd75cc9a7691d2f7d2d96f73c53653bef8940a49e96dbdf7775";
|
||||
|
||||
fn fixture() -> String {
|
||||
FIXTURE.replace("SEALED_DIGEST_HEX", FIXTURE_DIGEST_HEX)
|
||||
@@ -783,6 +1018,7 @@ mod tests {
|
||||
vec![AtRestProtection::EncryptedMasterKey],
|
||||
Some("verifier-opaque-1".to_string()),
|
||||
)),
|
||||
external_references: None,
|
||||
key_versions: None,
|
||||
capability_discovery: None,
|
||||
completeness: CompletenessState::InProgress,
|
||||
@@ -1010,12 +1246,39 @@ mod tests {
|
||||
let with_discovery = fixture().replace("\"completeness\"", "\"capability_discovery\": [1], \"completeness\"");
|
||||
expect_corrupted(BackupManifest::decode(with_discovery.as_bytes()), "reserved");
|
||||
|
||||
// Explicit null carries no data and is tolerated as absence; digest
|
||||
// verification still passes because null slots are skipped on
|
||||
// serialization.
|
||||
// An explicit null slot decodes as absence at the schema layer, but
|
||||
// sealed bundles never contain the key (`skip_serializing_if`), so
|
||||
// inserting one after sealing is a byte-level modification and the
|
||||
// raw-bytes digest check rejects it.
|
||||
let with_null = fixture().replace("\"completeness\"", "\"key_versions\": null, \"completeness\"");
|
||||
let decoded = BackupManifest::decode(with_null.as_bytes()).expect("null reserved slot should decode");
|
||||
assert_eq!(decoded.key_versions, None);
|
||||
expect_corrupted(BackupManifest::decode(with_null.as_bytes()), "digest mismatch");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn digest_verification_survives_non_round_tripping_timestamp_spellings() {
|
||||
// A legacy `created_at` spelling (no time zone annotation) parses via
|
||||
// the compat fallback and re-serializes differently ("+00:00[UTC]"),
|
||||
// and host-dependent zone spellings can do the same. Digest
|
||||
// verification therefore operates on the raw stored bytes and must
|
||||
// never re-serialize parsed fields.
|
||||
let sealed = seal(local_manifest_unsealed());
|
||||
let mut value = serde_json::to_value(&sealed).expect("manifest should convert to a value");
|
||||
value["created_at"] = serde_json::Value::String("2026-07-30T00:00:00+00:00".to_string());
|
||||
value["manifest_digest"]["hex"] = serde_json::Value::String(String::new());
|
||||
let digest = BackupManifest::digest_of_canonical_value(value.clone(), DigestAlgorithm::Sha256)
|
||||
.expect("canonical digest should compute");
|
||||
value["manifest_digest"]["hex"] = serde_json::Value::String(digest.hex);
|
||||
let bytes = serde_json::to_vec(&value).expect("manifest bytes");
|
||||
|
||||
let decoded = BackupManifest::decode(&bytes).expect("a non-round-tripping timestamp spelling must not break decoding");
|
||||
|
||||
// Precondition: the spelling really does not survive a typed
|
||||
// round-trip — otherwise this test is vacuous.
|
||||
let reserialized = serde_json::to_value(&decoded).expect("decoded manifest should convert to a value");
|
||||
assert_ne!(reserialized["created_at"], value["created_at"]);
|
||||
// Which is exactly why the producer-side (in-memory) digest check
|
||||
// cannot be used on decoded manifests.
|
||||
assert!(decoded.verify_digest().is_err());
|
||||
}
|
||||
|
||||
#[test]
|
||||
@@ -1162,4 +1425,161 @@ mod tests {
|
||||
assert_eq!(decoded, sealed);
|
||||
assert_eq!(decoded.responsibility, BackupResponsibility::ReferenceOnly);
|
||||
}
|
||||
|
||||
fn transit_references() -> VaultExternalReferences {
|
||||
VaultExternalReferences {
|
||||
cluster_id: "vault-cluster-a".to_string(),
|
||||
namespace: Some("tenant-a".to_string()),
|
||||
kv_mount: "secret".to_string(),
|
||||
kv_path_prefix: "rustfs/kms/transit-metadata".to_string(),
|
||||
transit: Some(VaultTransitReference {
|
||||
mount_path: "transit".to_string(),
|
||||
key_name: "rustfs-master".to_string(),
|
||||
required_min_version: 2,
|
||||
current_version: 5,
|
||||
native_snapshot_reference: Some("s3://dr/vault/2026-08-01.snap".to_string()),
|
||||
}),
|
||||
kv_records: vec![VaultKvRecordReference {
|
||||
key_id: "object-key".to_string(),
|
||||
kv_version: 4,
|
||||
}],
|
||||
}
|
||||
}
|
||||
|
||||
fn transit_manifest_unsealed() -> BackupManifest {
|
||||
BackupManifest {
|
||||
backend: BackupBackendKind::VaultTransit,
|
||||
at_rest_protection: AtRestProtection::ExternalNonExportable,
|
||||
responsibility: BackupResponsibility::MetadataPlusExternalRoot,
|
||||
artifacts: vec![artifact(ArtifactKind::KeyMetadata, "vault/records/object-key.json.enc")],
|
||||
local_kdf: None,
|
||||
external_references: Some(transit_references()),
|
||||
..local_manifest_unsealed()
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn vault_bundle_round_trips_with_external_references() {
|
||||
let sealed = seal(transit_manifest_unsealed());
|
||||
let bytes = sealed.encode().expect("encoding should succeed");
|
||||
let decoded = BackupManifest::decode(&bytes).expect("decoding should succeed");
|
||||
assert_eq!(decoded, sealed);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn vault_backend_requires_external_references() {
|
||||
let manifest = BackupManifest {
|
||||
external_references: None,
|
||||
..transit_manifest_unsealed()
|
||||
};
|
||||
expect_corrupted(
|
||||
BackupManifest::decode(&serde_json::to_vec(&seal(manifest)).expect("serialize")),
|
||||
"must carry external Vault references",
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn non_vault_backend_rejects_external_references() {
|
||||
let manifest = BackupManifest {
|
||||
external_references: Some(transit_references()),
|
||||
..local_manifest_unsealed()
|
||||
};
|
||||
expect_corrupted(
|
||||
BackupManifest::decode(&serde_json::to_vec(&seal(manifest)).expect("serialize")),
|
||||
"not valid for backend",
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn external_reference_contradictions_fail_closed() {
|
||||
let cases: [(&str, Box<dyn Fn(&mut VaultExternalReferences)>); 7] = [
|
||||
("cluster_id", Box::new(|refs| refs.cluster_id.clear())),
|
||||
("kv_mount", Box::new(|refs| refs.kv_mount.clear())),
|
||||
("namespace must not be empty", Box::new(|refs| refs.namespace = Some(String::new()))),
|
||||
(
|
||||
"must reference the Transit key",
|
||||
Box::new(|refs: &mut VaultExternalReferences| refs.transit = None),
|
||||
),
|
||||
(
|
||||
"at least key version 1",
|
||||
Box::new(|refs: &mut VaultExternalReferences| {
|
||||
if let Some(transit) = refs.transit.as_mut() {
|
||||
transit.required_min_version = 0;
|
||||
}
|
||||
}),
|
||||
),
|
||||
(
|
||||
"below the required minimum",
|
||||
Box::new(|refs: &mut VaultExternalReferences| {
|
||||
if let Some(transit) = refs.transit.as_mut() {
|
||||
transit.current_version = 1;
|
||||
}
|
||||
}),
|
||||
),
|
||||
(
|
||||
"appears more than once",
|
||||
Box::new(|refs: &mut VaultExternalReferences| {
|
||||
let first = refs.kv_records[0].clone();
|
||||
refs.kv_records.push(first);
|
||||
}),
|
||||
),
|
||||
];
|
||||
for (needle, mutate) in cases {
|
||||
let mut references = transit_references();
|
||||
mutate(&mut references);
|
||||
let manifest = BackupManifest {
|
||||
external_references: Some(references),
|
||||
..transit_manifest_unsealed()
|
||||
};
|
||||
expect_corrupted(BackupManifest::decode(&serde_json::to_vec(&seal(manifest)).expect("serialize")), needle);
|
||||
}
|
||||
}
|
||||
|
||||
/// The reference schema is the only place a Vault coordinate reaches a
|
||||
/// bundle, so its field set is pinned: a credential-carrying field could
|
||||
/// only appear by editing this assertion.
|
||||
#[test]
|
||||
fn external_reference_field_set_is_frozen() {
|
||||
let value = serde_json::to_value(transit_references()).expect("serialize");
|
||||
let mut top: Vec<&str> = value.as_object().expect("object").keys().map(String::as_str).collect();
|
||||
top.sort_unstable();
|
||||
assert_eq!(
|
||||
top,
|
||||
[
|
||||
"cluster_id",
|
||||
"kv_mount",
|
||||
"kv_path_prefix",
|
||||
"kv_records",
|
||||
"namespace",
|
||||
"transit"
|
||||
]
|
||||
);
|
||||
|
||||
let mut transit: Vec<&str> = value["transit"]
|
||||
.as_object()
|
||||
.expect("transit object")
|
||||
.keys()
|
||||
.map(String::as_str)
|
||||
.collect();
|
||||
transit.sort_unstable();
|
||||
assert_eq!(
|
||||
transit,
|
||||
[
|
||||
"current_version",
|
||||
"key_name",
|
||||
"mount_path",
|
||||
"native_snapshot_reference",
|
||||
"required_min_version"
|
||||
]
|
||||
);
|
||||
|
||||
let mut record: Vec<&str> = value["kv_records"][0]
|
||||
.as_object()
|
||||
.expect("record object")
|
||||
.keys()
|
||||
.map(String::as_str)
|
||||
.collect();
|
||||
record.sort_unstable();
|
||||
assert_eq!(record, ["key_id", "kv_version"]);
|
||||
}
|
||||
}
|
||||
|
||||
@@ -12,13 +12,16 @@
|
||||
// See the License for the specific language governing permissions and
|
||||
// limitations under the License.
|
||||
|
||||
//! Backup/restore contract types for KMS state.
|
||||
//! Backup/restore contracts and backup production for KMS state.
|
||||
//!
|
||||
//! This module is contract-only: it defines the versioned backup manifest,
|
||||
//! the per-backend responsibility matrix, typed failure modes, and the
|
||||
//! restore dry-run report. Nothing here is wired into handlers or backends;
|
||||
//! backup export, restore orchestration, and the admin API build on these
|
||||
//! types in follow-up changes.
|
||||
//! The contract side defines the versioned backup manifest, the per-backend
|
||||
//! responsibility matrix, typed failure modes, and the restore dry-run
|
||||
//! report. [`local_export`] implements the producer side and
|
||||
//! [`local_restore`] the consumer side for the Local backend;
|
||||
//! [`vault_restore`] orchestrates the consumer side for the Vault backends,
|
||||
//! whose cryptographic root is restored by Vault's own disaster-recovery
|
||||
//! flow. All are crate-internal APIs; the admin API builds on these pieces in
|
||||
//! follow-up changes.
|
||||
//!
|
||||
//! # Bundle model
|
||||
//!
|
||||
@@ -50,14 +53,30 @@
|
||||
mod capability;
|
||||
mod dry_run;
|
||||
mod error;
|
||||
pub mod local_export;
|
||||
pub mod local_restore;
|
||||
mod manifest;
|
||||
pub mod vault_restore;
|
||||
|
||||
pub use capability::{AtRestProtection, BackupBackendKind, BackupResponsibility};
|
||||
pub use dry_run::{
|
||||
ExternalDependencyMismatch, RestoreBlocker, RestoreBlockerCode, RestoreConflict, RestoreConflictKind, RestoreDryRunReport,
|
||||
};
|
||||
pub use error::BackupError;
|
||||
pub use local_export::{
|
||||
BackupKek, LOCAL_BUNDLE_MANIFEST_FILE, LocalBackupExportRequest, decrypt_bundle_artifact, export_local_backup,
|
||||
read_bundle_manifest, read_local_bundle_manifest,
|
||||
};
|
||||
pub use local_restore::{
|
||||
LocalRestoreReport, LocalRestoreRequest, RestoreConflictPolicy, abort_local_restore, dry_run_local_restore,
|
||||
restore_local_backup,
|
||||
};
|
||||
pub use manifest::{
|
||||
AeadAlgorithm, ArtifactDescriptor, ArtifactKind, BackupKekDescriptor, BackupManifest, CompletenessState, ContentDigest,
|
||||
DigestAlgorithm, LocalKdfDescriptor, LocalKeyDerivation, ReservedSlot,
|
||||
DigestAlgorithm, LocalKdfDescriptor, LocalKeyDerivation, ReservedSlot, VaultExternalReferences, VaultKvRecordReference,
|
||||
VaultTransitReference,
|
||||
};
|
||||
pub use vault_restore::{
|
||||
VaultRestoreClient, VaultRestoreMismatch, VaultRestoreReport, VaultRestoreRequest, VaultRestoreSequence, VaultRestoreStage,
|
||||
VaultRestoreTarget, abort_vault_restore, dry_run_vault_restore, restore_vault_backup,
|
||||
};
|
||||
|
||||
File diff suppressed because it is too large
Load Diff
@@ -14,8 +14,8 @@
|
||||
|
||||
//! Fault-injection matrix for the Vault backend operation policy.
|
||||
//!
|
||||
//! Offline cases run against locally injected transport faults (a closed
|
||||
//! port, a listener that never responds) — deterministic, no external
|
||||
//! Offline cases run against locally injected transport faults (a listener
|
||||
//! that never responds) — deterministic, no external
|
||||
//! dependencies. Real-Vault cases are `#[ignore]`d and need a dev Vault
|
||||
//! (default `http://127.0.0.1:8200`, override with `RUSTFS_KMS_VAULT_ADDR`).
|
||||
//!
|
||||
@@ -33,9 +33,11 @@ use std::time::Duration;
|
||||
|
||||
use metrics_util::MetricKind;
|
||||
use metrics_util::debugging::{DebugValue, DebuggingRecorder};
|
||||
use rustfs_kms::backends::KmsClient;
|
||||
use rustfs_kms::backends::vault::VaultKmsClient;
|
||||
use rustfs_kms::{KmsConfig, KmsError, VaultAuthMethod, VaultConfig};
|
||||
use rustfs_kms::backends::KmsBackend as KmsBackendTrait;
|
||||
use rustfs_kms::backends::vault::VaultKmsBackend;
|
||||
use rustfs_kms::{
|
||||
BackendConfig, DescribeKeyRequest, KmsBackend as KmsBackendKind, KmsConfig, KmsError, VaultAuthMethod, VaultConfig,
|
||||
};
|
||||
|
||||
const OPERATIONS_TOTAL: &str = "rustfs_kms_backend_operations_total";
|
||||
const ATTEMPT_FAILURES_TOTAL: &str = "rustfs_kms_backend_attempt_failures_total";
|
||||
@@ -54,14 +56,23 @@ fn vault_config(address: &str, token: &str) -> VaultConfig {
|
||||
}
|
||||
}
|
||||
|
||||
fn kms_config(attempt_timeout: Duration, retry_attempts: u32) -> KmsConfig {
|
||||
fn kms_config(vault_config: VaultConfig, attempt_timeout: Duration, retry_attempts: u32) -> KmsConfig {
|
||||
KmsConfig {
|
||||
backend: KmsBackendKind::VaultKv2,
|
||||
backend_config: BackendConfig::VaultKv2(Box::new(vault_config)),
|
||||
allow_insecure_dev_defaults: true,
|
||||
timeout: attempt_timeout,
|
||||
retry_attempts,
|
||||
..KmsConfig::default()
|
||||
}
|
||||
}
|
||||
|
||||
fn describe_key_request(key_id: &str) -> DescribeKeyRequest {
|
||||
DescribeKeyRequest {
|
||||
key_id: key_id.to_string(),
|
||||
}
|
||||
}
|
||||
|
||||
type MetricEntry = (
|
||||
metrics_util::CompositeKey,
|
||||
Option<metrics::Unit>,
|
||||
@@ -107,45 +118,6 @@ fn counter_value(snapshot: &[MetricEntry], name: &str, labels: &[(&str, &str)])
|
||||
.sum()
|
||||
}
|
||||
|
||||
/// Connection refused: connection-class failures are retried up to the
|
||||
/// configured budget, then surface as a backend error.
|
||||
#[test]
|
||||
fn connection_refused_is_retried_within_budget() {
|
||||
// Reserve a loopback port and release it so nothing is listening there.
|
||||
let listener = std::net::TcpListener::bind("127.0.0.1:0").expect("reserve a loopback port");
|
||||
let address = format!("http://{}", listener.local_addr().expect("reserved port addr"));
|
||||
drop(listener);
|
||||
|
||||
let snapshot = record_metrics(|| {
|
||||
Box::pin(async move {
|
||||
let client = VaultKmsClient::new(vault_config(&address, "unused"), &kms_config(Duration::from_secs(2), 2))
|
||||
.await
|
||||
.expect("client construction performs no network calls");
|
||||
let error = client
|
||||
.describe_key("fault-injection-refused", None)
|
||||
.await
|
||||
.expect_err("a refused connection must fail the operation");
|
||||
assert!(matches!(error, KmsError::BackendError { .. }), "got {error:?}");
|
||||
})
|
||||
});
|
||||
|
||||
assert_eq!(
|
||||
counter_value(&snapshot, ATTEMPT_FAILURES_TOTAL, &[("error_class", "retryable_conn")]),
|
||||
2,
|
||||
"both budgeted attempts must observe the refused connection"
|
||||
);
|
||||
assert_eq!(counter_value(&snapshot, OPERATIONS_TOTAL, &[("outcome", "budget_exhausted")]), 1);
|
||||
// The static-token login records its own success; the Vault read must not.
|
||||
assert_eq!(
|
||||
counter_value(
|
||||
&snapshot,
|
||||
OPERATIONS_TOTAL,
|
||||
&[("operation", "vault_kv2_read_key"), ("outcome", "success")]
|
||||
),
|
||||
0
|
||||
);
|
||||
}
|
||||
|
||||
/// Stalled connection: a server that accepts but never responds is cut off by
|
||||
/// the per-attempt timeout (either the policy timer or the equally sized HTTP
|
||||
/// client timeout, whichever fires first) instead of hanging forever.
|
||||
@@ -166,11 +138,10 @@ fn stalled_connection_is_cut_off_by_the_attempt_timeout() {
|
||||
}
|
||||
});
|
||||
|
||||
let client = VaultKmsClient::new(vault_config(&address, "unused"), &kms_config(Duration::from_millis(250), 1))
|
||||
let client = VaultKmsBackend::new(kms_config(vault_config(&address, "unused"), Duration::from_millis(250), 1))
|
||||
.await
|
||||
.expect("client construction performs no network calls");
|
||||
let error = client
|
||||
.describe_key("fault-injection-stalled", None)
|
||||
let error = KmsBackendTrait::describe_key(&client, describe_key_request("fault-injection-stalled"))
|
||||
.await
|
||||
.expect_err("a stalled request must be cut off by the attempt timeout");
|
||||
assert!(
|
||||
@@ -201,11 +172,10 @@ fn real_vault_invalid_token_is_fatal_and_never_retried() {
|
||||
let snapshot = record_metrics(|| {
|
||||
Box::pin(async {
|
||||
let config = vault_config(&real_vault_address(), "fault-injection-invalid-token");
|
||||
let client = VaultKmsClient::new(config, &kms_config(Duration::from_secs(5), 3))
|
||||
let client = VaultKmsBackend::new(kms_config(config, Duration::from_secs(5), 3))
|
||||
.await
|
||||
.expect("client construction performs no network calls");
|
||||
let error = client
|
||||
.describe_key("fault-injection-forbidden", None)
|
||||
let error = KmsBackendTrait::describe_key(&client, describe_key_request("fault-injection-forbidden"))
|
||||
.await
|
||||
.expect_err("an invalid token must be rejected");
|
||||
assert!(matches!(error, KmsError::BackendError { .. }), "got {error:?}");
|
||||
@@ -235,11 +205,10 @@ fn real_vault_missing_key_is_resolved_in_one_attempt() {
|
||||
let snapshot = record_metrics(|| {
|
||||
Box::pin(async move {
|
||||
let config = vault_config(&real_vault_address(), &token);
|
||||
let client = VaultKmsClient::new(config, &kms_config(Duration::from_secs(5), 3))
|
||||
let client = VaultKmsBackend::new(kms_config(config, Duration::from_secs(5), 3))
|
||||
.await
|
||||
.expect("client construction performs no network calls");
|
||||
let error = client
|
||||
.describe_key("fault-injection-definitely-missing", None)
|
||||
let error = KmsBackendTrait::describe_key(&client, describe_key_request("fault-injection-definitely-missing"))
|
||||
.await
|
||||
.expect_err("a missing key must resolve to key-not-found");
|
||||
assert!(matches!(error, KmsError::KeyNotFound { .. }), "got {error:?}");
|
||||
|
||||
@@ -19,7 +19,7 @@ use crate::{
|
||||
resolve_notify_object_store_handle,
|
||||
rule_engine::NotifyRuleEngine,
|
||||
runtime_facade::NotifyRuntimeFacade,
|
||||
with_notify_server_config_read_lock, with_notify_server_config_write_lock,
|
||||
with_notify_server_config_read_lock,
|
||||
};
|
||||
use rustfs_config::notify::{
|
||||
NOTIFY_AMQP_SUB_SYS, NOTIFY_KAFKA_SUB_SYS, NOTIFY_MQTT_SUB_SYS, NOTIFY_MYSQL_SUB_SYS, NOTIFY_NATS_SUB_SYS,
|
||||
@@ -41,12 +41,32 @@ enum NotifyConfigStoreError {
|
||||
StorageNotAvailable,
|
||||
Read(String),
|
||||
Save(String),
|
||||
Converge { persisted: bool, error: String },
|
||||
}
|
||||
|
||||
fn notification_convergence_error(persisted: bool, error: impl std::fmt::Display) -> NotificationError {
|
||||
let durable_state = if persisted { "persisted" } else { "unchanged" };
|
||||
NotificationError::Configuration(format!(
|
||||
"configuration was {durable_state} but notification runtime convergence failed: {error}"
|
||||
))
|
||||
}
|
||||
|
||||
async fn supervise_config_update<T>(
|
||||
mutation: impl std::future::Future<Output = Result<T, NotifyConfigStoreError>> + Send + 'static,
|
||||
) -> Result<T, NotifyConfigStoreError>
|
||||
where
|
||||
T: Send + 'static,
|
||||
{
|
||||
tokio::spawn(mutation).await.map_err(|error| {
|
||||
let outcome = if error.is_cancelled() { "cancelled" } else { "panicked" };
|
||||
NotifyConfigStoreError::Lock(format!("notify config update task {outcome}"))
|
||||
})?
|
||||
}
|
||||
|
||||
async fn update_server_config<F>(
|
||||
modifier: F,
|
||||
mut modifier: F,
|
||||
lifecycle: NotifyLifecycleCoordinator,
|
||||
) -> Result<Option<crate::lifecycle::NotificationLifecycleTransition>, NotifyConfigStoreError>
|
||||
) -> Result<Option<(crate::lifecycle::NotificationLifecycleTransition, bool)>, NotifyConfigStoreError>
|
||||
where
|
||||
F: FnMut(&mut Config) -> bool + Send + 'static,
|
||||
{
|
||||
@@ -54,51 +74,31 @@ where
|
||||
return Err(NotifyConfigStoreError::StorageNotAvailable);
|
||||
};
|
||||
|
||||
let store_for_read = store.clone();
|
||||
let store_for_save = store.clone();
|
||||
with_notify_server_config_write_lock(store, move || {
|
||||
read_modify_write(
|
||||
modifier,
|
||||
move || async move {
|
||||
crate::read_notify_server_config_without_migrate_no_lock(store_for_read)
|
||||
.await
|
||||
.map_err(NotifyConfigStoreError::Read)
|
||||
},
|
||||
move |config| async move {
|
||||
crate::save_notify_server_config_no_lock(store_for_save, &config)
|
||||
.await
|
||||
.map_err(NotifyConfigStoreError::Save)
|
||||
},
|
||||
move |config| lifecycle.update_config(config),
|
||||
)
|
||||
supervise_config_update(async move {
|
||||
let snapshot = crate::read_notify_server_config_snapshot(store.clone())
|
||||
.await
|
||||
.map_err(NotifyConfigStoreError::Read)?;
|
||||
let mut config = snapshot.config.clone();
|
||||
if !modifier(&mut config) {
|
||||
return Ok(None);
|
||||
}
|
||||
|
||||
let persisted = crate::save_notify_server_config_snapshot(store.clone(), &config, &snapshot)
|
||||
.await
|
||||
.map_err(NotifyConfigStoreError::Save)?;
|
||||
drop(snapshot);
|
||||
|
||||
let read_store = store.clone();
|
||||
with_notify_server_config_read_lock(store, move || async move {
|
||||
let latest = crate::read_existing_notify_server_config_no_lock(read_store)
|
||||
.await
|
||||
.map_err(|error| NotifyConfigStoreError::Converge { persisted, error })?;
|
||||
Ok::<_, NotifyConfigStoreError>(Some((lifecycle.update_config(latest), persisted)))
|
||||
})
|
||||
.await
|
||||
.map_err(|error| NotifyConfigStoreError::Converge { persisted, error })?
|
||||
})
|
||||
.await
|
||||
.map_err(NotifyConfigStoreError::Lock)?
|
||||
}
|
||||
|
||||
async fn read_modify_write<F, R, RFut, S, SFut, P, T>(
|
||||
mut modifier: F,
|
||||
read: R,
|
||||
save: S,
|
||||
publish: P,
|
||||
) -> Result<Option<T>, NotifyConfigStoreError>
|
||||
where
|
||||
F: FnMut(&mut Config) -> bool,
|
||||
R: FnOnce() -> RFut,
|
||||
RFut: std::future::Future<Output = Result<Config, NotifyConfigStoreError>>,
|
||||
S: FnOnce(Config) -> SFut,
|
||||
SFut: std::future::Future<Output = Result<(), NotifyConfigStoreError>>,
|
||||
P: FnOnce(Config) -> T,
|
||||
{
|
||||
let mut new_config = read().await?;
|
||||
|
||||
if !modifier(&mut new_config) {
|
||||
return Ok(None);
|
||||
}
|
||||
|
||||
save(new_config.clone()).await?;
|
||||
|
||||
Ok(Some(publish(new_config)))
|
||||
}
|
||||
|
||||
pub(crate) fn notify_configuration_hint() -> String {
|
||||
@@ -345,16 +345,18 @@ impl NotifyConfigManager {
|
||||
if self.lifecycle.state() == NotificationRuntimeState::Terminated {
|
||||
return Err(NotificationError::Initialization("Notification runtime has terminated".to_string()));
|
||||
}
|
||||
let Some(transition) = update_server_config(modifier, self.lifecycle.clone())
|
||||
.await
|
||||
.map_err(|err| match err {
|
||||
NotifyConfigStoreError::Lock(err) => NotificationError::StorageNotAvailable(err),
|
||||
NotifyConfigStoreError::StorageNotAvailable => NotificationError::StorageNotAvailable(
|
||||
"Failed to save target configuration: server storage not initialized".to_string(),
|
||||
),
|
||||
NotifyConfigStoreError::Read(err) => NotificationError::ReadConfig(err),
|
||||
NotifyConfigStoreError::Save(err) => NotificationError::SaveConfig(err),
|
||||
})?
|
||||
let Some((transition, persisted)) =
|
||||
update_server_config(modifier, self.lifecycle.clone())
|
||||
.await
|
||||
.map_err(|err| match err {
|
||||
NotifyConfigStoreError::Lock(err) => NotificationError::StorageNotAvailable(err),
|
||||
NotifyConfigStoreError::StorageNotAvailable => NotificationError::StorageNotAvailable(
|
||||
"Failed to save target configuration: server storage not initialized".to_string(),
|
||||
),
|
||||
NotifyConfigStoreError::Read(err) => NotificationError::ReadConfig(err),
|
||||
NotifyConfigStoreError::Save(err) => NotificationError::SaveConfig(err),
|
||||
NotifyConfigStoreError::Converge { persisted, error } => notification_convergence_error(persisted, error),
|
||||
})?
|
||||
else {
|
||||
debug!(
|
||||
event = EVENT_NOTIFY_CONFIG_UPDATE,
|
||||
@@ -367,23 +369,53 @@ impl NotifyConfigManager {
|
||||
return Ok(());
|
||||
};
|
||||
|
||||
let result = if persisted { "updated" } else { "unchanged" };
|
||||
info!(
|
||||
event = EVENT_NOTIFY_CONFIG_UPDATE,
|
||||
component = LOG_COMPONENT_NOTIFY,
|
||||
subsystem = LOG_SUBSYSTEM_CONFIG,
|
||||
action = "reload_if_changed",
|
||||
result = "updated",
|
||||
result,
|
||||
"notify config update"
|
||||
);
|
||||
transition.wait().await
|
||||
transition
|
||||
.wait()
|
||||
.await
|
||||
.map_err(|err| notification_convergence_error(persisted, err))
|
||||
}
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
mod tests {
|
||||
use super::{NotifyConfigManager, NotifyConfigStoreError, read_modify_write, runtime_target_id_for_subsystem};
|
||||
use super::{
|
||||
NotifyConfigManager, NotifyConfigStoreError, notification_convergence_error, runtime_target_id_for_subsystem,
|
||||
supervise_config_update,
|
||||
};
|
||||
use crate::rules::RulesMap;
|
||||
use crate::{NotificationError, NotificationRuntimeState};
|
||||
|
||||
#[tokio::test]
|
||||
async fn supervised_config_update_survives_waiter_cancellation() {
|
||||
let (started_tx, started_rx) = tokio::sync::oneshot::channel();
|
||||
let (release_tx, release_rx) = tokio::sync::oneshot::channel();
|
||||
let (completed_tx, completed_rx) = tokio::sync::oneshot::channel();
|
||||
|
||||
let waiter = tokio::spawn(async move {
|
||||
supervise_config_update(async move {
|
||||
let _ = started_tx.send(());
|
||||
let _ = release_rx.await;
|
||||
let _ = completed_tx.send(());
|
||||
Ok::<_, NotifyConfigStoreError>(())
|
||||
})
|
||||
.await
|
||||
});
|
||||
|
||||
started_rx.await.expect("config update should start");
|
||||
waiter.abort();
|
||||
release_tx.send(()).expect("config update should be released");
|
||||
let completed = tokio::time::timeout(std::time::Duration::from_secs(30), completed_rx).await;
|
||||
assert!(matches!(completed, Ok(Ok(()))), "detached config update should complete");
|
||||
}
|
||||
use crate::{
|
||||
integration::NotificationMetrics, notifier::EventNotifier, registry::TargetRegistry, rule_engine::NotifyRuleEngine,
|
||||
runtime_facade::NotifyRuntimeFacade,
|
||||
@@ -392,14 +424,11 @@ mod tests {
|
||||
NOTIFY_AMQP_SUB_SYS, NOTIFY_KAFKA_SUB_SYS, NOTIFY_MQTT_SUB_SYS, NOTIFY_NATS_SUB_SYS, NOTIFY_POSTGRES_SUB_SYS,
|
||||
NOTIFY_PULSAR_SUB_SYS, NOTIFY_REDIS_SUB_SYS, NOTIFY_WEBHOOK_SUB_SYS,
|
||||
};
|
||||
use rustfs_config::server_config::{Config, KVS};
|
||||
use rustfs_config::server_config::Config;
|
||||
use rustfs_s3_types::EventName;
|
||||
use rustfs_targets::ReplayWorkerManager;
|
||||
use rustfs_targets::arn::TargetID;
|
||||
use std::sync::{
|
||||
Arc,
|
||||
atomic::{AtomicBool, Ordering},
|
||||
};
|
||||
use std::sync::Arc;
|
||||
use tokio::sync::{RwLock, Semaphore};
|
||||
|
||||
fn build_manager() -> NotifyConfigManager {
|
||||
@@ -463,38 +492,6 @@ mod tests {
|
||||
assert!(matches!(manager.lifecycle().state(), NotificationRuntimeState::TargetsEnabled { .. }));
|
||||
}
|
||||
|
||||
#[tokio::test]
|
||||
async fn read_modify_write_publishes_only_after_save() {
|
||||
let saved = Arc::new(AtomicBool::new(false));
|
||||
let saved_by_writer = saved.clone();
|
||||
let observed_by_publisher = saved.clone();
|
||||
|
||||
read_modify_write(
|
||||
|config| {
|
||||
config
|
||||
.0
|
||||
.entry(NOTIFY_WEBHOOK_SUB_SYS.to_string())
|
||||
.or_default()
|
||||
.insert("primary".to_string(), KVS::default());
|
||||
true
|
||||
},
|
||||
|| async { Ok::<_, NotifyConfigStoreError>(Config::default()) },
|
||||
move |_config| async move {
|
||||
saved_by_writer.store(true, Ordering::Release);
|
||||
Ok::<_, NotifyConfigStoreError>(())
|
||||
},
|
||||
move |_config| {
|
||||
assert!(
|
||||
observed_by_publisher.load(Ordering::Acquire),
|
||||
"publication must observe the completed save"
|
||||
);
|
||||
},
|
||||
)
|
||||
.await
|
||||
.expect("read-modify-write should succeed")
|
||||
.expect("changed config should publish");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn runtime_target_id_for_subsystem_maps_notify_webhook_to_runtime_type() {
|
||||
let target_id = runtime_target_id_for_subsystem(NOTIFY_WEBHOOK_SUB_SYS, "Primary");
|
||||
@@ -550,4 +547,16 @@ mod tests {
|
||||
assert_eq!(target_id.id, "audittrail");
|
||||
assert_eq!(target_id.name, "postgres");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn convergence_error_reports_durable_write_state() {
|
||||
for (persisted, expected) in [(true, "configuration was persisted"), (false, "configuration was unchanged")] {
|
||||
let error = notification_convergence_error(persisted, "injected failure");
|
||||
let NotificationError::Configuration(message) = error else {
|
||||
panic!("convergence error should be a configuration error");
|
||||
};
|
||||
assert!(message.contains(expected), "unexpected convergence error: {message}");
|
||||
assert!(message.contains("injected failure"));
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
@@ -57,7 +57,6 @@ pub use services::NotifyServices;
|
||||
pub use status_view::NotifyStatusView;
|
||||
pub use storage_api::NotifyStore;
|
||||
pub(crate) use storage_api::crate_boundary::{
|
||||
read_existing_notify_server_config_no_lock, read_notify_server_config_without_migrate_no_lock,
|
||||
resolve_notify_object_store_handle, save_notify_server_config_no_lock, with_notify_server_config_read_lock,
|
||||
with_notify_server_config_write_lock,
|
||||
read_existing_notify_server_config_no_lock, read_notify_server_config_snapshot, resolve_notify_object_store_handle,
|
||||
save_notify_server_config_snapshot, with_notify_server_config_read_lock,
|
||||
};
|
||||
|
||||
@@ -15,11 +15,10 @@
|
||||
use std::sync::Arc;
|
||||
|
||||
use rustfs_ecstore::api::config::com::{
|
||||
read_config_without_migrate_no_lock as read_notify_config_without_migrate_from_backend_no_lock,
|
||||
read_existing_server_config_no_lock as read_existing_notify_config_from_backend_no_lock,
|
||||
save_server_config_no_lock as save_notify_server_config_to_backend_no_lock,
|
||||
read_server_config_snapshot as read_notify_server_config_snapshot_from_backend,
|
||||
save_server_config_snapshot as save_notify_server_config_snapshot_to_backend,
|
||||
with_server_config_read_lock as with_notify_server_config_read_lock_from_backend,
|
||||
with_server_config_write_lock as with_notify_server_config_write_lock_from_backend,
|
||||
};
|
||||
use rustfs_ecstore::api::runtime::object_store_handle as resolve_notify_object_store_handle_from_backend;
|
||||
pub use rustfs_ecstore::api::storage::ECStore as NotifyStore;
|
||||
@@ -28,10 +27,10 @@ pub(crate) fn resolve_notify_object_store_handle() -> Option<Arc<NotifyStore>> {
|
||||
resolve_notify_object_store_handle_from_backend()
|
||||
}
|
||||
|
||||
pub(crate) async fn read_notify_server_config_without_migrate_no_lock(
|
||||
store: Arc<NotifyStore>,
|
||||
) -> Result<rustfs_config::server_config::Config, String> {
|
||||
read_notify_config_without_migrate_from_backend_no_lock(store)
|
||||
pub(crate) type NotifyServerConfigSnapshot = rustfs_ecstore::api::config::com::ServerConfigSnapshot;
|
||||
|
||||
pub(crate) async fn read_notify_server_config_snapshot(store: Arc<NotifyStore>) -> Result<NotifyServerConfigSnapshot, String> {
|
||||
read_notify_server_config_snapshot_from_backend(store)
|
||||
.await
|
||||
.map_err(|err| err.to_string())
|
||||
}
|
||||
@@ -44,22 +43,12 @@ pub(crate) async fn read_existing_notify_server_config_no_lock(
|
||||
.map_err(|err| err.to_string())
|
||||
}
|
||||
|
||||
pub(crate) async fn save_notify_server_config_no_lock(
|
||||
pub(crate) async fn save_notify_server_config_snapshot(
|
||||
store: Arc<NotifyStore>,
|
||||
config: &rustfs_config::server_config::Config,
|
||||
) -> Result<(), String> {
|
||||
save_notify_server_config_to_backend_no_lock(store, config)
|
||||
.await
|
||||
.map_err(|err| err.to_string())
|
||||
}
|
||||
|
||||
pub(crate) async fn with_notify_server_config_write_lock<F, Fut, T>(store: Arc<NotifyStore>, operation: F) -> Result<T, String>
|
||||
where
|
||||
F: FnOnce() -> Fut + Send + 'static,
|
||||
Fut: std::future::Future<Output = T> + Send + 'static,
|
||||
T: Send + 'static,
|
||||
{
|
||||
with_notify_server_config_write_lock_from_backend(store, operation)
|
||||
snapshot: &NotifyServerConfigSnapshot,
|
||||
) -> Result<bool, String> {
|
||||
save_notify_server_config_snapshot_to_backend(store, config, snapshot)
|
||||
.await
|
||||
.map_err(|err| err.to_string())
|
||||
}
|
||||
@@ -77,8 +66,7 @@ where
|
||||
|
||||
pub(crate) mod crate_boundary {
|
||||
pub(crate) use super::{
|
||||
read_existing_notify_server_config_no_lock, read_notify_server_config_without_migrate_no_lock,
|
||||
resolve_notify_object_store_handle, save_notify_server_config_no_lock, with_notify_server_config_read_lock,
|
||||
with_notify_server_config_write_lock,
|
||||
read_existing_notify_server_config_no_lock, read_notify_server_config_snapshot, resolve_notify_object_store_handle,
|
||||
save_notify_server_config_snapshot, with_notify_server_config_read_lock,
|
||||
};
|
||||
}
|
||||
|
||||
@@ -72,4 +72,10 @@ pub enum Error {
|
||||
|
||||
#[error("invalid resource, type: '{0}', pattern: '{1}'")]
|
||||
InvalidResource(String, String),
|
||||
|
||||
#[error("KMS resources require a statement whose actions are all KMS actions")]
|
||||
KmsResourceWithNonKmsAction,
|
||||
|
||||
#[error("bucket policies do not support KMS actions or resources")]
|
||||
KmsUnsupportedInBucketPolicy,
|
||||
}
|
||||
|
||||
@@ -728,6 +728,8 @@ pub enum KmsAction {
|
||||
ListKeysAction,
|
||||
#[strum(serialize = "kms:DescribeKey")]
|
||||
DescribeKeyAction,
|
||||
#[strum(serialize = "kms:Decrypt")]
|
||||
DecryptAction,
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
@@ -764,6 +766,7 @@ mod tests {
|
||||
("kms:RotateKey", KmsAction::RotateKeyAction),
|
||||
("kms:ListKeys", KmsAction::ListKeysAction),
|
||||
("kms:DescribeKey", KmsAction::DescribeKeyAction),
|
||||
("kms:Decrypt", KmsAction::DecryptAction),
|
||||
] {
|
||||
let action = Action::try_from(raw).expect("Should parse KMS action");
|
||||
assert_eq!(action, Action::KmsAction(expected));
|
||||
|
||||
@@ -1760,6 +1760,391 @@ mod test {
|
||||
);
|
||||
}
|
||||
|
||||
fn kms_args<'a>(
|
||||
action: Action,
|
||||
key_id: &'a str,
|
||||
conditions: &'a HashMap<String, Vec<String>>,
|
||||
claims: &'a HashMap<String, Value>,
|
||||
) -> Args<'a> {
|
||||
Args {
|
||||
account: "testuser",
|
||||
groups: &None,
|
||||
action,
|
||||
bucket: "",
|
||||
conditions,
|
||||
is_owner: false,
|
||||
object: key_id,
|
||||
claims,
|
||||
deny_only: false,
|
||||
}
|
||||
}
|
||||
|
||||
#[tokio::test]
|
||||
async fn test_kms_statement_with_key_resource_scopes_by_key() -> Result<()> {
|
||||
use crate::policy::action::{Action, KmsAction};
|
||||
|
||||
let policy = Policy::parse_config(
|
||||
br#"{
|
||||
"Version": "2012-10-17",
|
||||
"Statement": [
|
||||
{
|
||||
"Effect": "Allow",
|
||||
"Action": ["kms:DisableKey"],
|
||||
"Resource": ["arn:aws:kms:::key/key-a"]
|
||||
}
|
||||
]
|
||||
}"#,
|
||||
)?;
|
||||
let conditions = HashMap::new();
|
||||
let claims = HashMap::new();
|
||||
let action = Action::KmsAction(KmsAction::DisableKeyAction);
|
||||
|
||||
assert!(
|
||||
policy.is_allowed(&kms_args(action, "key-a", &conditions, &claims)).await,
|
||||
"the granted key must be allowed"
|
||||
);
|
||||
assert!(
|
||||
!policy.is_allowed(&kms_args(action, "key-b", &conditions, &claims)).await,
|
||||
"a key outside the granted resource must be denied"
|
||||
);
|
||||
assert!(
|
||||
!policy
|
||||
.is_allowed(&kms_args(Action::KmsAction(KmsAction::EnableKeyAction), "key-a", &conditions, &claims))
|
||||
.await,
|
||||
"an action outside the grant must stay denied even for the granted key"
|
||||
);
|
||||
assert!(
|
||||
policy.is_allowed(&kms_args(action, "", &conditions, &claims)).await,
|
||||
"call sites that do not pass a key resource keep the legacy match-every-key behaviour"
|
||||
);
|
||||
|
||||
Ok(())
|
||||
}
|
||||
|
||||
#[tokio::test]
|
||||
async fn test_kms_statement_with_wildcard_key_resource() -> Result<()> {
|
||||
use crate::policy::action::{Action, KmsAction};
|
||||
|
||||
let policy = Policy::parse_config(
|
||||
br#"{
|
||||
"Version": "2012-10-17",
|
||||
"Statement": [
|
||||
{
|
||||
"Effect": "Allow",
|
||||
"Action": ["kms:GenerateDataKey"],
|
||||
"Resource": ["arn:aws:kms:::key/app-*"]
|
||||
}
|
||||
]
|
||||
}"#,
|
||||
)?;
|
||||
let conditions = HashMap::new();
|
||||
let claims = HashMap::new();
|
||||
let action = Action::KmsAction(KmsAction::GenerateDataKeyAction);
|
||||
|
||||
assert!(
|
||||
policy
|
||||
.is_allowed(&kms_args(action, "app-primary", &conditions, &claims))
|
||||
.await
|
||||
);
|
||||
assert!(
|
||||
!policy
|
||||
.is_allowed(&kms_args(action, "backup-primary", &conditions, &claims))
|
||||
.await
|
||||
);
|
||||
|
||||
Ok(())
|
||||
}
|
||||
|
||||
#[tokio::test]
|
||||
async fn test_kms_statement_without_resource_matches_every_key() -> Result<()> {
|
||||
use crate::policy::action::{Action, KmsAction};
|
||||
|
||||
let policy = Policy::parse_config(
|
||||
br#"{
|
||||
"Version": "2012-10-17",
|
||||
"Statement": [
|
||||
{
|
||||
"Effect": "Allow",
|
||||
"Action": ["kms:*"]
|
||||
}
|
||||
]
|
||||
}"#,
|
||||
)?;
|
||||
let conditions = HashMap::new();
|
||||
let claims = HashMap::new();
|
||||
|
||||
for key_id in ["", "key-a", "any-other-key"] {
|
||||
assert!(
|
||||
policy
|
||||
.is_allowed(&kms_args(Action::KmsAction(KmsAction::DisableKeyAction), key_id, &conditions, &claims))
|
||||
.await,
|
||||
"resource-less KMS statement must keep matching every key (key_id: {key_id:?})"
|
||||
);
|
||||
}
|
||||
|
||||
Ok(())
|
||||
}
|
||||
|
||||
#[tokio::test]
|
||||
async fn test_kms_deny_with_wildcard_resource_overrides_allow() -> Result<()> {
|
||||
use crate::policy::action::{Action, KmsAction};
|
||||
|
||||
let policy = Policy::parse_config(
|
||||
br#"{
|
||||
"Version": "2012-10-17",
|
||||
"Statement": [
|
||||
{
|
||||
"Effect": "Allow",
|
||||
"Action": ["kms:*"],
|
||||
"Resource": ["arn:aws:kms:::key/key-a"]
|
||||
},
|
||||
{
|
||||
"Effect": "Deny",
|
||||
"Action": ["kms:DisableKey"],
|
||||
"Resource": ["arn:aws:kms:::key/*"]
|
||||
}
|
||||
]
|
||||
}"#,
|
||||
)?;
|
||||
let conditions = HashMap::new();
|
||||
let claims = HashMap::new();
|
||||
|
||||
assert!(
|
||||
!policy
|
||||
.is_allowed(&kms_args(Action::KmsAction(KmsAction::DisableKeyAction), "key-a", &conditions, &claims))
|
||||
.await,
|
||||
"a wildcard Deny must override the narrower Allow"
|
||||
);
|
||||
assert!(
|
||||
policy
|
||||
.is_allowed(&kms_args(Action::KmsAction(KmsAction::RotateKeyAction), "key-a", &conditions, &claims))
|
||||
.await,
|
||||
"actions outside the Deny keep the Allow"
|
||||
);
|
||||
|
||||
Ok(())
|
||||
}
|
||||
|
||||
#[tokio::test]
|
||||
async fn test_kms_statement_with_not_resource_excludes_keys() -> Result<()> {
|
||||
use crate::policy::action::{Action, KmsAction};
|
||||
|
||||
let policy = Policy::parse_config(
|
||||
br#"{
|
||||
"Version": "2012-10-17",
|
||||
"Statement": [
|
||||
{
|
||||
"Effect": "Allow",
|
||||
"Action": ["kms:DescribeKey"],
|
||||
"NotResource": ["arn:aws:kms:::key/prod-*"]
|
||||
}
|
||||
]
|
||||
}"#,
|
||||
)?;
|
||||
let conditions = HashMap::new();
|
||||
let claims = HashMap::new();
|
||||
let action = Action::KmsAction(KmsAction::DescribeKeyAction);
|
||||
|
||||
assert!(policy.is_allowed(&kms_args(action, "dev-key", &conditions, &claims)).await);
|
||||
assert!(!policy.is_allowed(&kms_args(action, "prod-key", &conditions, &claims)).await);
|
||||
|
||||
Ok(())
|
||||
}
|
||||
|
||||
#[tokio::test]
|
||||
async fn test_kms_statement_with_s3_resource_is_treated_as_unscoped() -> Result<()> {
|
||||
use crate::policy::action::{Action, KmsAction};
|
||||
|
||||
// Statements combining KMS actions with S3 resources predate KMS resource
|
||||
// support and were always evaluated as if unscoped. Pin that they still
|
||||
// match every key (a warning is logged during evaluation).
|
||||
let policy = Policy::parse_config(
|
||||
br#"{
|
||||
"Version": "2012-10-17",
|
||||
"Statement": [
|
||||
{
|
||||
"Effect": "Allow",
|
||||
"Action": ["kms:DisableKey"],
|
||||
"Resource": ["arn:aws:s3:::somebucket/*"]
|
||||
}
|
||||
]
|
||||
}"#,
|
||||
)?;
|
||||
let conditions = HashMap::new();
|
||||
let claims = HashMap::new();
|
||||
let action = Action::KmsAction(KmsAction::DisableKeyAction);
|
||||
|
||||
for key_id in ["", "key-a"] {
|
||||
assert!(
|
||||
policy.is_allowed(&kms_args(action, key_id, &conditions, &claims)).await,
|
||||
"malformed KMS statement with S3 resources must keep matching every key (key_id: {key_id:?})"
|
||||
);
|
||||
}
|
||||
|
||||
Ok(())
|
||||
}
|
||||
|
||||
#[tokio::test]
|
||||
async fn test_kms_alias_resource_parses_but_matches_no_key() -> Result<()> {
|
||||
use crate::policy::action::{Action, KmsAction};
|
||||
|
||||
let policy = Policy::parse_config(
|
||||
br#"{
|
||||
"Version": "2012-10-17",
|
||||
"Statement": [
|
||||
{
|
||||
"Effect": "Allow",
|
||||
"Action": ["kms:DescribeKey"],
|
||||
"Resource": ["arn:aws:kms:::alias/app-alias"]
|
||||
}
|
||||
]
|
||||
}"#,
|
||||
)?;
|
||||
let conditions = HashMap::new();
|
||||
let claims = HashMap::new();
|
||||
let action = Action::KmsAction(KmsAction::DescribeKeyAction);
|
||||
|
||||
assert!(
|
||||
!policy.is_allowed(&kms_args(action, "app-alias", &conditions, &claims)).await,
|
||||
"alias patterns are parse-only until alias resolution lands and must not match key requests"
|
||||
);
|
||||
assert!(
|
||||
policy.is_allowed(&kms_args(action, "", &conditions, &claims)).await,
|
||||
"call sites without a key resource keep the legacy match-every-key behaviour"
|
||||
);
|
||||
|
||||
Ok(())
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_kms_resource_policy_round_trips() {
|
||||
let policy = Policy::parse_config(
|
||||
br#"{
|
||||
"Version": "2012-10-17",
|
||||
"Statement": [
|
||||
{
|
||||
"Effect": "Allow",
|
||||
"Action": ["kms:GenerateDataKey", "kms:Decrypt"],
|
||||
"Resource": ["arn:aws:kms:::key/app-*", "arn:aws:kms:::alias/app-alias"]
|
||||
}
|
||||
]
|
||||
}"#,
|
||||
)
|
||||
.expect("KMS resource policy should parse");
|
||||
|
||||
let json = serde_json::to_string(&policy).expect("policy should serialize");
|
||||
let round_trip = Policy::parse_config(json.as_bytes()).expect("serialized KMS resource policy should re-parse");
|
||||
assert_eq!(round_trip.statements[0].resources, policy.statements[0].resources);
|
||||
assert_eq!(round_trip.statements[0].actions, policy.statements[0].actions);
|
||||
|
||||
let value: serde_json::Value = serde_json::from_str(&json).expect("JSON valid");
|
||||
let resources: Vec<_> = value["Statement"][0]["Resource"]
|
||||
.as_array()
|
||||
.expect("Resource should serialize as an array")
|
||||
.iter()
|
||||
.map(|resource| resource.as_str().expect("Resource entries should be strings"))
|
||||
.collect();
|
||||
assert_eq!(resources, vec!["arn:aws:kms:::key/app-*", "arn:aws:kms:::alias/app-alias"]);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_kms_resource_with_non_kms_action_is_invalid() {
|
||||
for action in ["s3:GetObject", "admin:ServerInfo", "sts:AssumeRole"] {
|
||||
let data = format!(
|
||||
r#"{{
|
||||
"Version": "2012-10-17",
|
||||
"Statement": [
|
||||
{{
|
||||
"Effect": "Allow",
|
||||
"Action": ["{action}"],
|
||||
"Resource": ["arn:aws:kms:::key/key-a"]
|
||||
}}
|
||||
]
|
||||
}}"#
|
||||
);
|
||||
|
||||
let result = Policy::parse_config(data.as_bytes());
|
||||
assert!(
|
||||
matches!(result.as_ref().unwrap_err(), Error::PolicyError(IamError::KmsResourceWithNonKmsAction)),
|
||||
"{action} with a KMS resource should fail with KmsResourceWithNonKmsAction, got: {result:?}"
|
||||
);
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_bucket_policy_with_kms_statement_is_invalid() {
|
||||
for statement in [
|
||||
r#"{"Effect":"Allow","Principal":{"AWS":"*"},"Action":["kms:GenerateDataKey"],"Resource":["arn:aws:s3:::bucket/*"]}"#,
|
||||
r#"{"Effect":"Allow","Principal":{"AWS":"*"},"Action":["s3:GetObject"],"Resource":["arn:aws:kms:::key/key-a"]}"#,
|
||||
r#"{"Effect":"Allow","Principal":{"AWS":"*"},"NotAction":["kms:*"],"Resource":["arn:aws:s3:::bucket/*"]}"#,
|
||||
] {
|
||||
let data = format!(r#"{{"Version":"2012-10-17","Statement":[{statement}]}}"#);
|
||||
let policy: BucketPolicy =
|
||||
serde_json::from_str(&data).expect("bucket policy with KMS content should still deserialize");
|
||||
let result = policy.is_valid();
|
||||
assert!(
|
||||
matches!(result.as_ref().unwrap_err(), Error::PolicyError(IamError::KmsUnsupportedInBucketPolicy)),
|
||||
"bucket policy statement {statement} should fail with KmsUnsupportedInBucketPolicy, got: {result:?}"
|
||||
);
|
||||
}
|
||||
}
|
||||
|
||||
#[tokio::test]
|
||||
async fn test_stored_bucket_policy_with_kms_statement_loads_and_is_ignored() -> Result<()> {
|
||||
// Policies stored before KMS statements were rejected at validation must keep
|
||||
// deserializing, and their KMS statements must not affect bucket traffic.
|
||||
let bucket_policy: BucketPolicy = serde_json::from_str(
|
||||
r#"{
|
||||
"Version": "2012-10-17",
|
||||
"Statement": [
|
||||
{
|
||||
"Effect": "Allow",
|
||||
"Principal": {"AWS": "*"},
|
||||
"Action": ["kms:GenerateDataKey"],
|
||||
"Resource": ["arn:aws:s3:::bucket/*"]
|
||||
},
|
||||
{
|
||||
"Effect": "Allow",
|
||||
"Principal": {"AWS": "*"},
|
||||
"Action": ["s3:GetObject"],
|
||||
"Resource": ["arn:aws:s3:::bucket/*"]
|
||||
}
|
||||
]
|
||||
}"#,
|
||||
)?;
|
||||
|
||||
let conditions = HashMap::new();
|
||||
let args = BucketPolicyArgs {
|
||||
account: "testuser",
|
||||
groups: &None,
|
||||
action: Action::S3Action(crate::policy::action::S3Action::GetObjectAction),
|
||||
bucket: "bucket",
|
||||
conditions: &conditions,
|
||||
is_owner: false,
|
||||
object: "a.txt",
|
||||
};
|
||||
assert!(
|
||||
bucket_policy.is_allowed(&args).await,
|
||||
"the S3 statement must keep working alongside an ignored KMS statement"
|
||||
);
|
||||
|
||||
let put_args = BucketPolicyArgs {
|
||||
account: "testuser",
|
||||
groups: &None,
|
||||
action: Action::S3Action(crate::policy::action::S3Action::PutObjectAction),
|
||||
bucket: "bucket",
|
||||
conditions: &conditions,
|
||||
is_owner: false,
|
||||
object: "a.txt",
|
||||
};
|
||||
assert!(
|
||||
!bucket_policy.is_allowed(&put_args).await,
|
||||
"the ignored KMS statement must not grant anything"
|
||||
);
|
||||
|
||||
Ok(())
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_mixed_action_families_are_invalid_even_with_resource() {
|
||||
let data = r#"
|
||||
|
||||
@@ -40,7 +40,7 @@ impl Serialize for ResourceSet {
|
||||
for resource in &self.0 {
|
||||
let resource_str = match resource {
|
||||
Resource::S3(value) => format!("{}{}", Resource::S3_PREFIX, value),
|
||||
Resource::Kms(value) => value.clone(),
|
||||
Resource::Kms(value) => format!("{}{}", Resource::KMS_PREFIX, value),
|
||||
};
|
||||
seq.serialize_element(&resource_str)?;
|
||||
}
|
||||
@@ -171,6 +171,20 @@ pub enum Resource {
|
||||
|
||||
impl Resource {
|
||||
pub const S3_PREFIX: &'static str = "arn:aws:s3:::";
|
||||
/// KMS ARNs use the same empty-account form as [`Self::S3_PREFIX`]; the suffix is
|
||||
/// `key/<key_id>` (wildcards allowed in the id), `alias/<name>`, or a bare `*`.
|
||||
pub const KMS_PREFIX: &'static str = "arn:aws:kms:::";
|
||||
/// Resource-type segment for key ids; request-side KMS resource strings are
|
||||
/// `key/<key_id>` so they line up with these patterns.
|
||||
pub const KMS_KEY_SEGMENT: &'static str = "key/";
|
||||
/// Resource-type segment reserved for key aliases. Alias patterns parse and
|
||||
/// validate, but requests are always evaluated against `key/<key_id>` strings,
|
||||
/// so an alias pattern matches nothing until alias resolution lands.
|
||||
pub const KMS_ALIAS_SEGMENT: &'static str = "alias/";
|
||||
|
||||
pub fn is_kms(&self) -> bool {
|
||||
matches!(self, Resource::Kms(_))
|
||||
}
|
||||
|
||||
pub async fn is_match(&self, resource: &str, conditions: &HashMap<String, Vec<String>>) -> bool {
|
||||
self.is_match_with_resolver(resource, conditions, None).await
|
||||
@@ -228,10 +242,13 @@ impl Resource {
|
||||
impl TryFrom<&str> for Resource {
|
||||
type Error = Error;
|
||||
fn try_from(value: &str) -> std::result::Result<Self, Self::Error> {
|
||||
let Some(value) = value.strip_prefix(Self::S3_PREFIX) else {
|
||||
let resource = if let Some(suffix) = value.strip_prefix(Self::S3_PREFIX) {
|
||||
Resource::S3(suffix.into())
|
||||
} else if let Some(suffix) = value.strip_prefix(Self::KMS_PREFIX) {
|
||||
Resource::Kms(suffix.into())
|
||||
} else {
|
||||
return Err(IamError::InvalidResource("unknown".into(), value.into()).into());
|
||||
};
|
||||
let resource = Resource::S3(value.into());
|
||||
|
||||
resource.is_valid()?;
|
||||
Ok(resource)
|
||||
@@ -248,13 +265,17 @@ impl Validator for Resource {
|
||||
}
|
||||
}
|
||||
Self::Kms(pattern) => {
|
||||
if pattern.is_empty()
|
||||
// A bare `*` matches every key resource; anything else must carry a
|
||||
// resource-type segment. Key ids never contain separators (the KMS
|
||||
// backends reject them), while alias names may nest ("alias/aws/s3").
|
||||
let well_formed = pattern == "*"
|
||||
|| pattern
|
||||
.char_indices()
|
||||
.find(|&(_, c)| c == '/' || c == '\\' || c == '.')
|
||||
.map(|(i, _)| i)
|
||||
.is_some()
|
||||
{
|
||||
.strip_prefix(Self::KMS_KEY_SEGMENT)
|
||||
.is_some_and(|id| !id.is_empty() && !id.contains('/') && !id.contains('\\'))
|
||||
|| pattern
|
||||
.strip_prefix(Self::KMS_ALIAS_SEGMENT)
|
||||
.is_some_and(|name| !name.is_empty() && !name.contains('\\'));
|
||||
if !well_formed {
|
||||
return Err(IamError::InvalidResource("kms".into(), pattern.into()).into());
|
||||
}
|
||||
}
|
||||
@@ -270,7 +291,7 @@ impl Serialize for Resource {
|
||||
{
|
||||
match self {
|
||||
Resource::S3(s) => serializer.serialize_str(&format!("{}{}", Self::S3_PREFIX, s)),
|
||||
Resource::Kms(s) => serializer.serialize_str(s),
|
||||
Resource::Kms(s) => serializer.serialize_str(&format!("{}{}", Self::KMS_PREFIX, s)),
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -308,9 +329,71 @@ mod tests {
|
||||
#[test_case("arn:aws:s3:::mybucket","mybucket/myobject" => false; "15")]
|
||||
#[test_case("arn:aws:s3:::attacker-bucket/*","attacker-bucket/../victim-bucket/evil.txt" => false; "16")]
|
||||
#[test_case("arn:aws:s3:::attacker-bucket/*","attacker-bucket/safe/../../victim-bucket/evil.txt" => false; "17")]
|
||||
#[test_case("arn:aws:kms:::key/mykey","key/mykey" => true; "kms exact key")]
|
||||
#[test_case("arn:aws:kms:::key/mykey","key/otherkey" => false; "kms wrong key")]
|
||||
#[test_case("arn:aws:kms:::key/*","key/mykey" => true; "kms key wildcard")]
|
||||
#[test_case("arn:aws:kms:::key/app-*","key/app-primary" => true; "kms key prefix wildcard")]
|
||||
#[test_case("arn:aws:kms:::key/app-*","key/backup-primary" => false; "kms key prefix mismatch")]
|
||||
#[test_case("arn:aws:kms:::key/mykey?","key/mykey1" => true; "kms key question mark")]
|
||||
#[test_case("arn:aws:kms:::*","key/mykey" => true; "kms bare star")]
|
||||
#[test_case("arn:aws:kms:::key/mykey","key/mykey/../otherkey" => false; "kms traversal cleaned")]
|
||||
#[test_case("arn:aws:kms:::alias/myalias","key/myalias" => false; "kms alias never matches key")]
|
||||
#[test_case("arn:aws:kms:::alias/*","key/mykey" => false; "kms alias wildcard never matches key")]
|
||||
#[test_case("arn:aws:kms:::key/mykey","mykey" => false; "kms bare id lacks key segment")]
|
||||
fn test_resource_is_match(resource: &str, object: &str) -> bool {
|
||||
let resource: Resource = resource.try_into().unwrap();
|
||||
|
||||
pollster::block_on(resource.is_match(object, &HashMap::new()))
|
||||
}
|
||||
|
||||
#[test_case("arn:aws:kms:::key/mykey" => true; "key id parses")]
|
||||
#[test_case("arn:aws:kms:::key/*" => true; "key wildcard parses")]
|
||||
#[test_case("arn:aws:kms:::key/app-key.v2" => true; "key id with dot parses")]
|
||||
#[test_case("arn:aws:kms:::alias/myalias" => true; "alias parses")]
|
||||
#[test_case("arn:aws:kms:::alias/aws/s3" => true; "nested alias parses")]
|
||||
#[test_case("arn:aws:kms:::*" => true; "bare star parses")]
|
||||
#[test_case("arn:aws:kms:::" => false; "empty suffix rejected")]
|
||||
#[test_case("arn:aws:kms:::key/" => false; "empty key id rejected")]
|
||||
#[test_case("arn:aws:kms:::alias/" => false; "empty alias name rejected")]
|
||||
#[test_case("arn:aws:kms:::mykey" => false; "missing resource type segment rejected")]
|
||||
#[test_case("arn:aws:kms:::key/a/b" => false; "separator in key id rejected")]
|
||||
#[test_case("arn:aws:kms:::key/a\\b" => false; "backslash in key id rejected")]
|
||||
#[test_case("arn:aws:kms:us-east-1:123456789012:key/mykey" => false; "region and account form rejected")]
|
||||
fn test_kms_resource_parse(resource: &str) -> bool {
|
||||
Resource::try_from(resource).is_ok()
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_kms_resource_serialization_round_trip() {
|
||||
for raw in [
|
||||
"arn:aws:kms:::key/mykey",
|
||||
"arn:aws:kms:::key/*",
|
||||
"arn:aws:kms:::alias/myalias",
|
||||
"arn:aws:kms:::*",
|
||||
] {
|
||||
let resource = Resource::try_from(raw).expect("KMS resource should parse");
|
||||
assert!(resource.is_kms());
|
||||
|
||||
let json = serde_json::to_string(&resource).expect("KMS resource should serialize");
|
||||
assert_eq!(json, format!("\"{raw}\""), "serialization must write back the full ARN");
|
||||
|
||||
let round_trip: Resource = serde_json::from_str(&json).expect("serialized KMS resource should deserialize");
|
||||
assert_eq!(round_trip, resource);
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn test_kms_resource_set_serialization_round_trip() {
|
||||
use crate::policy::resource::ResourceSet;
|
||||
|
||||
let set: ResourceSet =
|
||||
serde_json::from_str(r#"["arn:aws:kms:::key/app-*","arn:aws:kms:::alias/myalias"]"#).expect("set should parse");
|
||||
assert_eq!(set.len(), 2);
|
||||
|
||||
let json = serde_json::to_string(&set).expect("set should serialize");
|
||||
assert_eq!(json, r#"["arn:aws:kms:::key/app-*","arn:aws:kms:::alias/myalias"]"#);
|
||||
|
||||
let round_trip: ResourceSet = serde_json::from_str(&json).expect("serialized set should deserialize");
|
||||
assert_eq!(round_trip, set);
|
||||
}
|
||||
}
|
||||
|
||||
@@ -16,6 +16,7 @@ use super::{
|
||||
ActionSet, Args, BucketPolicyArgs, Effect, Error as IamError, Functions, ID, Principal, ResourceSet, Validator,
|
||||
action::{Action, S3Action},
|
||||
function::key_name::{KeyName, S3KeyName},
|
||||
resource::Resource,
|
||||
variables::{VariableContext, VariableResolver},
|
||||
};
|
||||
use crate::error::{Error, Result};
|
||||
@@ -173,13 +174,79 @@ impl Statement {
|
||||
Some(ActionFamily::Mixed)
|
||||
}
|
||||
|
||||
/// Resource scope check for KMS statements, which match `arn:aws:kms:::key/<key_id>`
|
||||
/// patterns instead of the bucket/object path grammar used by S3 statements.
|
||||
///
|
||||
/// Call-site contract (wired up by the admin/SSE authorization paths): the requested
|
||||
/// key identifier (pre-alias-resolution) travels in `args.object` with `args.bucket`
|
||||
/// left empty. An empty `args.object` means the caller did not scope the request to a
|
||||
/// key, which preserves the legacy match-every-key behaviour.
|
||||
async fn kms_key_scope_matches(&self, args: &Args<'_>, resolver: &VariableResolver) -> bool {
|
||||
let kms_resources: Vec<&Resource> = self.resources.iter().filter(|resource| resource.is_kms()).collect();
|
||||
let kms_not_resources: Vec<&Resource> = self.not_resources.iter().filter(|resource| resource.is_kms()).collect();
|
||||
|
||||
if kms_resources.len() != self.resources.len() || kms_not_resources.len() != self.not_resources.len() {
|
||||
// Statements combining KMS actions with S3 resources predate KMS resource
|
||||
// support and were always evaluated as if unscoped; keep that behaviour
|
||||
// but surface it, since the S3 patterns never constrain key access.
|
||||
tracing::warn!(
|
||||
sid = %self.sid.0,
|
||||
"KMS statement carries non-KMS resources; they are ignored and the statement matches every key"
|
||||
);
|
||||
}
|
||||
|
||||
if kms_resources.is_empty() && kms_not_resources.is_empty() {
|
||||
// No KMS resources: the statement scopes by action only (legacy form).
|
||||
return true;
|
||||
}
|
||||
|
||||
if args.object.is_empty() {
|
||||
// Call sites that do not pass a key resource keep the pre-resource-scoping
|
||||
// behaviour where any key matches.
|
||||
return true;
|
||||
}
|
||||
|
||||
let requested = format!("{}{}", Resource::KMS_KEY_SEGMENT, args.object);
|
||||
|
||||
if !kms_resources.is_empty() {
|
||||
let mut matched = false;
|
||||
for resource in kms_resources {
|
||||
if resource
|
||||
.is_match_with_resolver(&requested, args.conditions, Some(resolver))
|
||||
.await
|
||||
{
|
||||
matched = true;
|
||||
break;
|
||||
}
|
||||
}
|
||||
if !matched {
|
||||
return false;
|
||||
}
|
||||
}
|
||||
|
||||
for resource in kms_not_resources {
|
||||
if resource
|
||||
.is_match_with_resolver(&requested, args.conditions, Some(resolver))
|
||||
.await
|
||||
{
|
||||
return false;
|
||||
}
|
||||
}
|
||||
|
||||
true
|
||||
}
|
||||
|
||||
/// Returns true when this statement would reach `conditions.evaluate_with_resolver` in
|
||||
/// [`Statement::is_allowed`] (including the KMS shortcut path). Does not evaluate conditions.
|
||||
/// [`Statement::is_allowed`] (including the KMS resource path). Does not evaluate conditions.
|
||||
pub(crate) async fn request_reaches_condition_eval(&self, args: &Args<'_>, resolver: &VariableResolver) -> bool {
|
||||
if (!self.actions.is_match(&args.action) && !self.actions.is_empty()) || self.not_actions.is_match(&args.action) {
|
||||
return false;
|
||||
}
|
||||
|
||||
if self.is_kms() {
|
||||
return self.kms_key_scope_matches(args, resolver).await;
|
||||
}
|
||||
|
||||
let resource = build_resource(
|
||||
&args.action,
|
||||
args.bucket,
|
||||
@@ -187,10 +254,6 @@ impl Statement {
|
||||
self.conditions.references_key_name(&KeyName::S3(S3KeyName::S3Prefix)),
|
||||
);
|
||||
|
||||
if self.is_kms() && (resource == "/" || self.resources.is_empty()) {
|
||||
return true;
|
||||
}
|
||||
|
||||
if self.resources.is_empty() && self.not_resources.is_empty() && !self.is_admin() && !self.is_sts() {
|
||||
return false;
|
||||
}
|
||||
@@ -273,6 +336,19 @@ impl Validator for Statement {
|
||||
return Err(IamError::BothResourceAndNotResource.into());
|
||||
}
|
||||
|
||||
// KMS resources only make sense on pure-KMS statements. The reverse
|
||||
// combination (KMS actions with S3 resources) predates KMS resources,
|
||||
// may already be stored, and stays loadable; evaluation treats it as
|
||||
// unscoped and warns.
|
||||
let has_kms_resource = self
|
||||
.resources
|
||||
.iter()
|
||||
.chain(self.not_resources.iter())
|
||||
.any(|resource| resource.is_kms());
|
||||
if has_kms_resource && !matches!(action_family, Some(ActionFamily::Kms)) {
|
||||
return Err(IamError::KmsResourceWithNonKmsAction.into());
|
||||
}
|
||||
|
||||
self.actions.is_valid()?;
|
||||
self.not_actions.is_valid()?;
|
||||
self.resources.is_valid()?;
|
||||
@@ -323,6 +399,18 @@ pub struct BPStatement {
|
||||
impl BPStatement {
|
||||
/// Returns true when this statement would reach `conditions.evaluate` in [`BPStatement::is_allowed`].
|
||||
pub(crate) async fn request_reaches_condition_eval(&self, args: &BucketPolicyArgs<'_>) -> bool {
|
||||
if !self.actions.is_empty() && self.actions.iter().all(|action| matches!(action, Action::KmsAction(_))) {
|
||||
// Bucket policies cannot grant or deny KMS access; such statements are
|
||||
// rejected at validation but may exist in policies stored before that
|
||||
// check. Skip them so they never influence bucket traffic. Statements
|
||||
// mixing KMS with S3 actions keep evaluating their S3 actions as before.
|
||||
tracing::warn!(
|
||||
sid = %self.sid.0,
|
||||
"ignoring bucket policy statement with KMS actions during evaluation"
|
||||
);
|
||||
return false;
|
||||
}
|
||||
|
||||
if !self.principal.is_match(args.account) {
|
||||
return false;
|
||||
}
|
||||
@@ -379,6 +467,24 @@ impl Validator for BPStatement {
|
||||
return Err(IamError::BothActionAndNotAction.into());
|
||||
}
|
||||
|
||||
// Bucket policies govern S3 access; KMS grants belong in identity policies.
|
||||
// Rejected here (PutBucketPolicy) only: deserialization stays permissive so
|
||||
// stored policies from before this check keep loading, and evaluation skips
|
||||
// pure-KMS statements with a warning.
|
||||
let has_kms_action = self
|
||||
.actions
|
||||
.iter()
|
||||
.chain(self.not_actions.iter())
|
||||
.any(|action| matches!(action, Action::KmsAction(_)));
|
||||
let has_kms_resource = self
|
||||
.resources
|
||||
.iter()
|
||||
.chain(self.not_resources.iter())
|
||||
.any(|resource| resource.is_kms());
|
||||
if has_kms_action || has_kms_resource {
|
||||
return Err(IamError::KmsUnsupportedInBucketPolicy.into());
|
||||
}
|
||||
|
||||
if self.resources.is_empty() && self.not_resources.is_empty() {
|
||||
return Err(IamError::NonResource.into());
|
||||
}
|
||||
|
||||
@@ -26,6 +26,12 @@
|
||||
//! policy is also allowed by the widened policy;
|
||||
//! (c) an empty policy denies every non-owner request (default deny).
|
||||
//!
|
||||
//! Properties (d)-(g) extend the same invariants to KMS key resources
|
||||
//! (`arn:aws:kms:::key/<key_id>`, backlog#1582): Deny-first over key scopes,
|
||||
//! wildcard/resource-less supersets implying concrete key grants, exact key
|
||||
//! scoping without cross-key leaks, and the legacy resource-less
|
||||
//! match-every-key compatibility pin.
|
||||
//!
|
||||
//! Pure evaluation: no IO, no global state, parallel-safe. Statements are built
|
||||
//! from JSON exactly like production policies arriving via PutPolicy. Generated
|
||||
//! bucket/key/action pools avoid wildcard metacharacters so resource patterns
|
||||
@@ -47,10 +53,23 @@ const OBJECT_ACTIONS: &[&str] = &[
|
||||
"s3:PutObjectTagging",
|
||||
];
|
||||
|
||||
/// Key-scoped KMS actions safe to pair with an `arn:aws:kms:::key/<id>` resource.
|
||||
const KMS_KEY_ACTIONS: &[&str] = &[
|
||||
"kms:GenerateDataKey",
|
||||
"kms:Decrypt",
|
||||
"kms:DisableKey",
|
||||
"kms:RotateKey",
|
||||
"kms:DescribeKey",
|
||||
];
|
||||
|
||||
fn statement_json(effect: &str, action: &str, resource: &str) -> String {
|
||||
format!(r#"{{"Effect":"{effect}","Action":["{action}"],"Resource":["{resource}"]}}"#)
|
||||
}
|
||||
|
||||
fn resourceless_statement_json(effect: &str, action: &str) -> String {
|
||||
format!(r#"{{"Effect":"{effect}","Action":["{action}"]}}"#)
|
||||
}
|
||||
|
||||
fn policy_from_statements(statements: &[String]) -> Policy {
|
||||
let json = format!(r#"{{"Version":"2012-10-17","Statement":[{}]}}"#, statements.join(","));
|
||||
serde_json::from_str(&json).expect("generated policy JSON should parse")
|
||||
@@ -88,6 +107,22 @@ fn action_strategy() -> impl Strategy<Value = &'static str> {
|
||||
proptest::sample::select(OBJECT_ACTIONS)
|
||||
}
|
||||
|
||||
/// Strategy: a KMS key id without wildcard metacharacters or separators.
|
||||
fn key_id_strategy() -> impl Strategy<Value = String> {
|
||||
"[a-z][a-z0-9-]{2,11}"
|
||||
}
|
||||
|
||||
/// Strategy: one action name from the key-scoped KMS action pool.
|
||||
fn kms_action_strategy() -> impl Strategy<Value = &'static str> {
|
||||
proptest::sample::select(KMS_KEY_ACTIONS)
|
||||
}
|
||||
|
||||
/// KMS evaluation contract: the requested key id travels in `args.object` with an
|
||||
/// empty bucket (see `Statement::kms_key_scope_matches`).
|
||||
fn is_allowed_for_key(policy: &Policy, action: &str, key_id: &str) -> bool {
|
||||
is_allowed(policy, action, "", key_id)
|
||||
}
|
||||
|
||||
proptest! {
|
||||
/// (a) Deny anywhere wins: a Deny statement matching the request denies it,
|
||||
/// no matter how many broad Allow statements surround it or at which index
|
||||
@@ -205,4 +240,122 @@ proptest! {
|
||||
"Policy::default() must deny {action} on {bucket}/{key}"
|
||||
);
|
||||
}
|
||||
|
||||
/// (d) KMS Deny anywhere wins: a Deny scoped to the exact key (or `key/*`)
|
||||
/// denies the request no matter how many broad KMS Allow statements
|
||||
/// (resource-less or `key/*`-scoped) surround it.
|
||||
#[test]
|
||||
fn kms_explicit_deny_anywhere_denies(
|
||||
key_id in key_id_strategy(),
|
||||
action in kms_action_strategy(),
|
||||
allow_count in 0usize..4,
|
||||
deny_pos_seed in 0usize..16,
|
||||
broad in proptest::bool::ANY,
|
||||
wildcard_deny in proptest::bool::ANY,
|
||||
) {
|
||||
let mut statements: Vec<String> = (0..allow_count)
|
||||
.map(|_| {
|
||||
if broad {
|
||||
resourceless_statement_json("Allow", "kms:*")
|
||||
} else {
|
||||
statement_json("Allow", "kms:*", "arn:aws:kms:::key/*")
|
||||
}
|
||||
})
|
||||
.collect();
|
||||
|
||||
let deny_resource = if wildcard_deny {
|
||||
"arn:aws:kms:::key/*".to_string()
|
||||
} else {
|
||||
format!("arn:aws:kms:::key/{key_id}")
|
||||
};
|
||||
let deny = statement_json("Deny", action, &deny_resource);
|
||||
let deny_pos = deny_pos_seed % (statements.len() + 1);
|
||||
statements.insert(deny_pos, deny);
|
||||
|
||||
let policy = policy_from_statements(&statements);
|
||||
|
||||
if allow_count > 0 {
|
||||
let mut allows_only = statements.clone();
|
||||
allows_only.remove(deny_pos);
|
||||
let allow_policy = policy_from_statements(&allows_only);
|
||||
prop_assert!(
|
||||
is_allowed_for_key(&allow_policy, action, &key_id),
|
||||
"sanity: the KMS Allow statements alone should permit {action} on key {key_id}"
|
||||
);
|
||||
}
|
||||
|
||||
prop_assert!(
|
||||
!is_allowed_for_key(&policy, action, &key_id),
|
||||
"explicit KMS Deny at index {deny_pos} of {} statements must deny {action} on key {key_id}",
|
||||
statements.len()
|
||||
);
|
||||
}
|
||||
|
||||
/// (e) KMS wildcard superset implies the concrete key grant: whatever an
|
||||
/// exact `key/<id>` Allow permits is also permitted by `key/*`, by a bare
|
||||
/// `arn:aws:kms:::*`, and by the legacy resource-less statement form.
|
||||
#[test]
|
||||
fn kms_wildcard_superset_implies_concrete_match(
|
||||
key_id in key_id_strategy(),
|
||||
action in kms_action_strategy(),
|
||||
) {
|
||||
let narrow = policy_from_statements(&[statement_json(
|
||||
"Allow",
|
||||
action,
|
||||
&format!("arn:aws:kms:::key/{key_id}"),
|
||||
)]);
|
||||
let widened = policy_from_statements(&[statement_json("Allow", "kms:*", "arn:aws:kms:::key/*")]);
|
||||
let star = policy_from_statements(&[statement_json("Allow", "kms:*", "arn:aws:kms:::*")]);
|
||||
let resourceless = policy_from_statements(&[resourceless_statement_json("Allow", "kms:*")]);
|
||||
|
||||
prop_assert!(is_allowed_for_key(&narrow, action, &key_id), "narrow KMS policy must allow its own grant");
|
||||
prop_assert!(is_allowed_for_key(&widened, action, &key_id), "kms:* on key/* must imply the concrete grant");
|
||||
prop_assert!(is_allowed_for_key(&star, action, &key_id), "kms:* on arn:aws:kms:::* must imply the concrete grant");
|
||||
prop_assert!(
|
||||
is_allowed_for_key(&resourceless, action, &key_id),
|
||||
"the legacy resource-less KMS statement must imply the concrete grant"
|
||||
);
|
||||
}
|
||||
|
||||
/// (f) Key scoping is exact: an Allow on `key/<a>` never leaks to a
|
||||
/// different key id, while the compatibility contract keeps unscoped
|
||||
/// requests (no key id passed) matching.
|
||||
#[test]
|
||||
fn kms_key_scope_does_not_leak_across_keys(
|
||||
key_a in key_id_strategy(),
|
||||
key_b in key_id_strategy(),
|
||||
action in kms_action_strategy(),
|
||||
) {
|
||||
prop_assume!(key_a != key_b);
|
||||
|
||||
let policy = policy_from_statements(&[statement_json(
|
||||
"Allow",
|
||||
action,
|
||||
&format!("arn:aws:kms:::key/{key_a}"),
|
||||
)]);
|
||||
|
||||
prop_assert!(is_allowed_for_key(&policy, action, &key_a), "the granted key must be allowed");
|
||||
prop_assert!(
|
||||
!is_allowed_for_key(&policy, action, &key_b),
|
||||
"an Allow scoped to key {key_a} must not leak to key {key_b}"
|
||||
);
|
||||
prop_assert!(
|
||||
is_allowed_for_key(&policy, action, ""),
|
||||
"call sites that pass no key resource keep the legacy match-every-key behaviour"
|
||||
);
|
||||
}
|
||||
|
||||
/// (g) Legacy compatibility pin: resource-less KMS statements match every
|
||||
/// generated key id, exactly as before KMS resources existed.
|
||||
#[test]
|
||||
fn kms_resourceless_statement_matches_every_key(
|
||||
key_id in key_id_strategy(),
|
||||
action in kms_action_strategy(),
|
||||
) {
|
||||
let policy = policy_from_statements(&[resourceless_statement_json("Allow", action)]);
|
||||
prop_assert!(
|
||||
is_allowed_for_key(&policy, action, &key_id),
|
||||
"resource-less KMS statement must keep matching {action} on key {key_id}"
|
||||
);
|
||||
}
|
||||
}
|
||||
|
||||
@@ -88,6 +88,7 @@ rmp-serde = { workspace = true }
|
||||
hmac = { workspace = true }
|
||||
sha2 = { workspace = true }
|
||||
rustfs-filemeta = { workspace = true }
|
||||
rustfs-lock.workspace = true
|
||||
tokio-util = { workspace = true, features = ["io", "compat", "rt"] }
|
||||
rustfs-ecstore = { workspace = true }
|
||||
rustfs-storage-api = { workspace = true }
|
||||
|
||||
@@ -12,17 +12,18 @@
|
||||
// See the License for the specific language governing permissions and
|
||||
// limitations under the License.
|
||||
|
||||
#[cfg(test)]
|
||||
use crate::RUSTFS_META_BUCKET;
|
||||
use crate::scanner_budget::{ScannerCycleBudget, ScannerCycleBudgetConfig};
|
||||
use crate::scanner_io::{
|
||||
DataUsageCacheScanState, ScannerDiskScanOutcome, ScannerIODisk, cache_root_entry_info, current_cache_root_or_prepare,
|
||||
scanner_cache_lock_resource, scanner_cache_lock_timeout, scanner_set_disk_inventory,
|
||||
DataUsageCacheScanState, ScannerDiskScanOutcome, ScannerIODisk, acquire_scanner_cache_locks, cache_root_entry_info,
|
||||
current_cache_root_or_prepare, scanner_set_disk_inventory,
|
||||
};
|
||||
use crate::storage_api::owner::NS_SCANNER_PROTOCOL_VERSION;
|
||||
use crate::storage_api::scan::NamespaceLocking as _;
|
||||
use crate::{
|
||||
DATA_USAGE_BLOOM_NAME_PATH, DATA_USAGE_CACHE_NAME, DataUsageCache, DataUsageCachePrepareOutcome, DataUsageCacheSource,
|
||||
DataUsageEntryInfo, DataUsageScanPlanDigest, Disk, EcstoreError, RUSTFS_META_BUCKET, ScannerDiskExt as _, ScannerError,
|
||||
ScannerObjectIO, StorageError, is_reserved_or_invalid_bucket, read_config, resolve_scanner_object_store_handle,
|
||||
DataUsageEntryInfo, DataUsageScanPlanDigest, Disk, EcstoreError, ScannerDiskExt as _, ScannerError, ScannerObjectIO,
|
||||
StorageError, is_reserved_or_invalid_bucket, read_config, resolve_scanner_object_store_handle,
|
||||
};
|
||||
use hmac::{Hmac, KeyInit, Mac};
|
||||
use rustfs_common::heal_channel::HealScanMode;
|
||||
@@ -63,6 +64,7 @@ const NS_SCANNER_DISK_HEALTH_TIMEOUT: Duration = Duration::from_secs(5);
|
||||
const NS_SCANNER_VALIDATED_CYCLE_TTL: Duration = Duration::from_secs(1);
|
||||
const NS_SCANNER_MAX_REPLAY_SESSIONS: usize = 65_536;
|
||||
const NS_SCANNER_MAX_ERROR_CHARS: usize = 4096;
|
||||
const NS_SCANNER_RETRY_BUCKET_ERROR_PREFIX: &str = "retry_bucket:";
|
||||
const NS_SCANNER_FRAME_AUTH_DOMAIN: &[u8] = b"rustfs-ns-scanner-frame-v3";
|
||||
|
||||
pub const NS_SCANNER_MAX_REQUEST_BODY_SIZE: usize = 16 * 1024;
|
||||
@@ -317,6 +319,13 @@ impl RemoteScannerServerError {
|
||||
}
|
||||
}
|
||||
|
||||
fn retry_bucket(message: impl Into<String>) -> Self {
|
||||
Self {
|
||||
scope: RemoteScannerErrorScope::Bucket,
|
||||
error: ScannerError::Other(format!("{}{}", NS_SCANNER_RETRY_BUCKET_ERROR_PREFIX, message.into())),
|
||||
}
|
||||
}
|
||||
|
||||
fn into_frame(self) -> RemoteScannerErrorFrame {
|
||||
RemoteScannerErrorFrame {
|
||||
scope: self.scope,
|
||||
@@ -336,6 +345,7 @@ struct RemoteScannerStreamError {
|
||||
error: StorageError,
|
||||
progress_fully_reported: bool,
|
||||
retire_worker: bool,
|
||||
retry_bucket: bool,
|
||||
}
|
||||
|
||||
impl RemoteScannerStreamError {
|
||||
@@ -344,6 +354,7 @@ impl RemoteScannerStreamError {
|
||||
error,
|
||||
progress_fully_reported: false,
|
||||
retire_worker: true,
|
||||
retry_bucket: false,
|
||||
}
|
||||
}
|
||||
|
||||
@@ -352,6 +363,7 @@ impl RemoteScannerStreamError {
|
||||
error,
|
||||
progress_fully_reported: true,
|
||||
retire_worker: true,
|
||||
retry_bucket: false,
|
||||
}
|
||||
}
|
||||
|
||||
@@ -360,6 +372,16 @@ impl RemoteScannerStreamError {
|
||||
error,
|
||||
progress_fully_reported: true,
|
||||
retire_worker: false,
|
||||
retry_bucket: false,
|
||||
}
|
||||
}
|
||||
|
||||
fn retry_bucket(error: StorageError) -> Self {
|
||||
Self {
|
||||
error,
|
||||
progress_fully_reported: true,
|
||||
retire_worker: false,
|
||||
retry_bucket: true,
|
||||
}
|
||||
}
|
||||
|
||||
@@ -951,16 +973,14 @@ async fn scan_and_persist_local_bucket(
|
||||
))
|
||||
})?;
|
||||
let cache_name = path_join_buf(&[&bucket, DATA_USAGE_CACHE_NAME]);
|
||||
let lock_resource = scanner_cache_lock_resource(&cache_name);
|
||||
let ns_lock = set
|
||||
.new_ns_lock(RUSTFS_META_BUCKET, &lock_resource)
|
||||
.await
|
||||
.map_err(|err| RemoteScannerServerError::worker(format!("remote namespace scanner cache lock creation failed: {err}")))?;
|
||||
let guard = ns_lock
|
||||
.get_write_lock_quiet(scanner_cache_lock_timeout())
|
||||
let guard = acquire_scanner_cache_locks(set.as_ref(), &cache_name, source)
|
||||
.await
|
||||
.map_err(|err| {
|
||||
RemoteScannerServerError::worker(format!("remote namespace scanner cache lock acquisition failed: {err}"))
|
||||
if err.is_contention() {
|
||||
RemoteScannerServerError::retry_bucket(format!("remote namespace scanner cache lock contention: {err}"))
|
||||
} else {
|
||||
RemoteScannerServerError::worker(format!("remote namespace scanner cache lock acquisition failed: {err}"))
|
||||
}
|
||||
})?;
|
||||
let mut cache = DataUsageCache::default();
|
||||
let revisions = cache.load_with_revisions(set.clone(), &cache_name).await.map_err(|err| {
|
||||
@@ -1216,7 +1236,9 @@ fn finish_remote_scanner_stream(
|
||||
if !error.progress_fully_reported {
|
||||
budget.cancel_after_unreported_remote_progress();
|
||||
}
|
||||
let failure = if error.retire_worker {
|
||||
let failure = if error.retry_bucket {
|
||||
RemoteScannerFailure::retry_bucket(error.error)
|
||||
} else if error.retire_worker {
|
||||
RemoteScannerFailure::transport(error.error)
|
||||
} else {
|
||||
RemoteScannerFailure::bucket(error.error)
|
||||
@@ -1374,9 +1396,16 @@ where
|
||||
return Ok(RemoteScannerOutcome::CycleAhead(required_cycle));
|
||||
}
|
||||
RemoteScannerFrameResult::Error(error_frame) => {
|
||||
let retry_bucket = error_frame.message.starts_with(NS_SCANNER_RETRY_BUCKET_ERROR_PREFIX);
|
||||
let message = error_frame
|
||||
.message
|
||||
.strip_prefix(NS_SCANNER_RETRY_BUCKET_ERROR_PREFIX)
|
||||
.map(str::trim_start)
|
||||
.unwrap_or(error_frame.message.as_str());
|
||||
let error =
|
||||
StorageError::other(format!("remote namespace scanner failed: {}", limit_error_message(error_frame.message)));
|
||||
StorageError::other(format!("remote namespace scanner failed: {}", limit_error_message(message.to_string())));
|
||||
return Err(match error_frame.scope {
|
||||
RemoteScannerErrorScope::Bucket if retry_bucket => RemoteScannerStreamError::retry_bucket(error),
|
||||
RemoteScannerErrorScope::Bucket => RemoteScannerStreamError::bucket(error),
|
||||
RemoteScannerErrorScope::Worker => RemoteScannerStreamError::reconciled(error),
|
||||
});
|
||||
@@ -2491,6 +2520,42 @@ mod tests {
|
||||
assert!(!budget.budget_elapsed());
|
||||
}
|
||||
|
||||
#[tokio::test]
|
||||
async fn retry_bucket_error_terminal_frame_requeues_bucket_work() {
|
||||
let request_id = Uuid::new_v4();
|
||||
let writer_auth = FrameAuthenticator::for_test(request_id);
|
||||
let reader_auth = FrameAuthenticator::for_test(request_id);
|
||||
let (mut writer, reader) = tokio::io::duplex(4096);
|
||||
tokio::spawn(async move {
|
||||
let mut sequence = 0;
|
||||
write_frame(
|
||||
&mut writer,
|
||||
&writer_auth,
|
||||
&mut sequence,
|
||||
&RemoteScannerFrame::terminal(
|
||||
RemoteScannerProgress::default(),
|
||||
RemoteScannerFrameResult::Error(RemoteScannerErrorFrame {
|
||||
scope: RemoteScannerErrorScope::Bucket,
|
||||
message: format!("{NS_SCANNER_RETRY_BUCKET_ERROR_PREFIX} cache lock contention"),
|
||||
}),
|
||||
),
|
||||
)
|
||||
.await
|
||||
.expect("retry-bucket error terminal frame should write");
|
||||
});
|
||||
|
||||
let parent = CancellationToken::new();
|
||||
let budget = ScannerCycleBudget::new(&parent, ScannerCycleBudgetConfig::default());
|
||||
let stream_result =
|
||||
consume_remote_scanner_stream(reader, parent, budget.clone(), "bucket", TEST_SOURCE, TEST_PLAN_DIGEST, reader_auth)
|
||||
.await;
|
||||
let error = finish_remote_scanner_stream(stream_result, budget.as_ref()).expect_err("retry-bucket frame must fail");
|
||||
|
||||
assert!(error.retry_bucket_work());
|
||||
assert!(!error.retire_worker());
|
||||
assert!(error.to_string().contains("cache lock contention"));
|
||||
}
|
||||
|
||||
#[tokio::test]
|
||||
async fn worker_error_terminal_frame_retires_worker_without_cancelling_reported_budget() {
|
||||
let request_id = Uuid::new_v4();
|
||||
|
||||
@@ -949,7 +949,7 @@ pub(crate) async fn probe_scanner_activity(storeapi: &ECStore, distributed: bool
|
||||
}
|
||||
SCANNER_ACTIVITY_PREVIOUS_PROTOCOL_VERSION => {
|
||||
return Err(format!(
|
||||
"scanner activity peer {host} cannot acknowledge distributed dirty usage with protocol {}",
|
||||
"scanner activity peer {host} cannot safely share scanner cache locks with protocol {}",
|
||||
SCANNER_ACTIVITY_PREVIOUS_PROTOCOL_VERSION
|
||||
));
|
||||
}
|
||||
@@ -7114,9 +7114,17 @@ mod tests {
|
||||
..scanner_node_activity("epoch-a", 7, 3)
|
||||
},
|
||||
)]);
|
||||
let previous = BTreeMap::from([(
|
||||
"node-2".to_string(),
|
||||
ScannerNodeActivity {
|
||||
protocol_version: SCANNER_ACTIVITY_PREVIOUS_PROTOCOL_VERSION,
|
||||
..scanner_node_activity("epoch-a", 7, 3)
|
||||
},
|
||||
)]);
|
||||
let current = BTreeMap::from([("node-2".to_string(), scanner_node_activity("epoch-a", 7, 3))]);
|
||||
|
||||
assert_ne!(scanner_activity_snapshot_digest(&legacy), scanner_activity_snapshot_digest(¤t));
|
||||
assert_ne!(scanner_activity_snapshot_digest(&previous), scanner_activity_snapshot_digest(¤t));
|
||||
}
|
||||
|
||||
#[test]
|
||||
|
||||
@@ -30,6 +30,7 @@ use rustfs_common::metrics::{Metric, Metrics, emit_scan_bucket_drive_complete, e
|
||||
use rustfs_config::{ENV_SCANNER_MAX_CONCURRENT_DISK_SCANS, ENV_SCANNER_MAX_CONCURRENT_SET_SCANS};
|
||||
use rustfs_data_usage::{BucketTargetUsageInfo, BucketUsageInfo};
|
||||
use rustfs_filemeta::FileMeta;
|
||||
use rustfs_lock::{LockError, NamespaceLockGuard};
|
||||
use rustfs_utils::path::path_join_buf;
|
||||
use s3s::dto::{BucketLifecycleConfiguration, ObjectLockConfiguration, ObjectLockEnabled, ReplicationConfiguration};
|
||||
use sha2::{Digest as _, Sha256};
|
||||
@@ -944,14 +945,85 @@ fn checked_bucket_usage_info(entry: &DataUsageEntry) -> Option<BucketUsageInfo>
|
||||
Some(usage)
|
||||
}
|
||||
|
||||
pub(crate) fn scanner_cache_lock_resource(cache_name: &str) -> String {
|
||||
path_join_buf(&[crate::BUCKET_META_PREFIX, cache_name, SCANNER_CACHE_LOCK_SUFFIX])
|
||||
pub(crate) fn scanner_cache_lock_resource(cache_name: &str, source: DataUsageCacheSource) -> String {
|
||||
let lock_name = format!("{SCANNER_CACHE_LOCK_SUFFIX}.pool-{}.set-{}", source.pool_index, source.set_index);
|
||||
path_join_buf(&[crate::BUCKET_META_PREFIX, cache_name, &lock_name])
|
||||
}
|
||||
|
||||
pub(crate) fn scanner_cache_lock_timeout() -> Duration {
|
||||
Duration::from_secs(rustfs_utils::get_env_u64("RUSTFS_LOCK_ACQUIRE_TIMEOUT", 5))
|
||||
}
|
||||
|
||||
#[derive(Debug)]
|
||||
pub(crate) struct ScannerCacheLockGuards {
|
||||
scoped: NamespaceLockGuard,
|
||||
}
|
||||
|
||||
impl ScannerCacheLockGuards {
|
||||
pub(crate) fn is_lock_lost(&self) -> bool {
|
||||
self.scoped.is_lock_lost()
|
||||
}
|
||||
}
|
||||
|
||||
#[derive(Debug)]
|
||||
pub(crate) enum ScannerCacheLockError {
|
||||
Create { resource: String, source: Error },
|
||||
Acquire { resource: String, source: LockError },
|
||||
}
|
||||
|
||||
impl ScannerCacheLockError {
|
||||
pub(crate) fn state(&self) -> &'static str {
|
||||
match self {
|
||||
Self::Create { .. } => "lock_create_failed",
|
||||
Self::Acquire { .. } => "lock_acquire_failed",
|
||||
}
|
||||
}
|
||||
|
||||
pub(crate) fn is_contention(&self) -> bool {
|
||||
matches!(
|
||||
self,
|
||||
Self::Acquire {
|
||||
source: LockError::Timeout { .. } | LockError::AlreadyLocked { .. },
|
||||
..
|
||||
}
|
||||
)
|
||||
}
|
||||
}
|
||||
|
||||
impl std::fmt::Display for ScannerCacheLockError {
|
||||
fn fmt(&self, formatter: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
|
||||
match self {
|
||||
Self::Create { resource, source } => write!(formatter, "create scanner cache lock {resource}: {source}"),
|
||||
Self::Acquire { resource, source } => write!(formatter, "acquire scanner cache lock {resource}: {source}"),
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
pub(crate) async fn acquire_scanner_cache_locks(
|
||||
store: &SetDisks,
|
||||
cache_name: &str,
|
||||
source: DataUsageCacheSource,
|
||||
) -> std::result::Result<ScannerCacheLockGuards, ScannerCacheLockError> {
|
||||
let timeout = scanner_cache_lock_timeout();
|
||||
let scoped_resource = scanner_cache_lock_resource(cache_name, source);
|
||||
let scoped_lock = store
|
||||
.new_ns_lock(RUSTFS_META_BUCKET, &scoped_resource)
|
||||
.await
|
||||
.map_err(|source| ScannerCacheLockError::Create {
|
||||
resource: scoped_resource.clone(),
|
||||
source,
|
||||
})?;
|
||||
let scoped = scoped_lock
|
||||
.get_write_lock_quiet(timeout)
|
||||
.await
|
||||
.map_err(|source| ScannerCacheLockError::Acquire {
|
||||
resource: scoped_resource,
|
||||
source,
|
||||
})?;
|
||||
|
||||
Ok(ScannerCacheLockGuards { scoped })
|
||||
}
|
||||
|
||||
async fn await_scanner_disk_shutdown<F>(scan: Pin<&mut F>)
|
||||
where
|
||||
F: Future,
|
||||
@@ -1697,6 +1769,19 @@ mod publish_gate_tests {
|
||||
assert!(!cache_snapshot_is_current(&cache, "photos", source, 11, 0, second));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn scanner_cache_lock_resource_is_scoped_to_cache_source() {
|
||||
let cache_name = "photos/.usage-cache.bin";
|
||||
let first_source = DataUsageCacheSource::new(0, 1);
|
||||
let same_source = DataUsageCacheSource::new(0, 1);
|
||||
let other_source = DataUsageCacheSource::new(1, 0);
|
||||
|
||||
let first = scanner_cache_lock_resource(cache_name, first_source);
|
||||
assert_eq!(first, scanner_cache_lock_resource(cache_name, same_source));
|
||||
assert_ne!(first, scanner_cache_lock_resource(cache_name, other_source));
|
||||
assert!(first.ends_with(".scanner-cycle.lock.pool-0.set-1"));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn count_budget_serializes_set_and_disk_work() {
|
||||
assert_eq!(scanner_budgeted_concurrency_limit(8, true), 1);
|
||||
@@ -1797,24 +1882,7 @@ async fn persist_and_publish_cache_snapshot(
|
||||
cache_cycle_floor: &AtomicU64,
|
||||
) -> Option<SystemTime> {
|
||||
let source = cache_snapshot.info.source?;
|
||||
let lock_resource = scanner_cache_lock_resource(DATA_USAGE_CACHE_NAME);
|
||||
let ns_lock = match store.new_ns_lock(RUSTFS_META_BUCKET, &lock_resource).await {
|
||||
Ok(lock) => lock,
|
||||
Err(err) => {
|
||||
error!(
|
||||
target: "rustfs::scanner::io",
|
||||
event = EVENT_SCANNER_CACHE_PERSIST_STATE,
|
||||
component = LOG_COMPONENT_SCANNER,
|
||||
subsystem = LOG_SUBSYSTEM_IO,
|
||||
cache_name = DATA_USAGE_CACHE_NAME,
|
||||
state = "lock_create_failed",
|
||||
error = %err,
|
||||
"Scanner cache snapshot lock creation failed"
|
||||
);
|
||||
return None;
|
||||
}
|
||||
};
|
||||
let guard = match ns_lock.get_write_lock_quiet(scanner_cache_lock_timeout()).await {
|
||||
let guard = match acquire_scanner_cache_locks(store.as_ref(), DATA_USAGE_CACHE_NAME, source).await {
|
||||
Ok(guard) => guard,
|
||||
Err(err) => {
|
||||
error!(
|
||||
@@ -1823,7 +1891,7 @@ async fn persist_and_publish_cache_snapshot(
|
||||
component = LOG_COMPONENT_SCANNER,
|
||||
subsystem = LOG_SUBSYSTEM_IO,
|
||||
cache_name = DATA_USAGE_CACHE_NAME,
|
||||
state = "lock_acquire_failed",
|
||||
state = err.state(),
|
||||
error = %err,
|
||||
"Scanner cache snapshot lock acquisition failed"
|
||||
);
|
||||
@@ -3080,32 +3148,35 @@ impl ScannerIOCache for SetDisks {
|
||||
None
|
||||
};
|
||||
|
||||
// Lock order: scanner leader fence -> per-bucket cache lock ->
|
||||
// cache object read/write. The cache lock stays outermost for
|
||||
// local and rolling-upgrade workers so leader failover cannot
|
||||
// execute the same bucket concurrently.
|
||||
let lock_resource = scanner_cache_lock_resource(&cache_name);
|
||||
let ns_lock = match store_clone_clone.new_ns_lock(RUSTFS_META_BUCKET, &lock_resource).await {
|
||||
Ok(lock) => lock,
|
||||
Err(e) => {
|
||||
record_failed_dirty_bucket(&failed_dirty_buckets_clone, &bucket.name).await;
|
||||
error!(
|
||||
target: "rustfs::scanner::io",
|
||||
event = EVENT_SCANNER_CACHE_PERSIST_STATE,
|
||||
component = LOG_COMPONENT_SCANNER,
|
||||
subsystem = LOG_SUBSYSTEM_IO,
|
||||
bucket = %bucket.name,
|
||||
cache_name = %cache_name,
|
||||
state = "lock_create_failed",
|
||||
error = %e,
|
||||
"Scanner bucket cache lock creation failed"
|
||||
);
|
||||
continue;
|
||||
}
|
||||
};
|
||||
let cache_guard = match ns_lock.get_write_lock_quiet(scanner_cache_lock_timeout()).await {
|
||||
// Lock order: scanner leader fence -> set-scoped per-bucket cache lock ->
|
||||
// cache object read/write.
|
||||
let cache_guard = match acquire_scanner_cache_locks(store_clone_clone.as_ref(), &cache_name, source).await {
|
||||
Ok(guard) => guard,
|
||||
Err(e) => {
|
||||
if e.is_contention() {
|
||||
if requeue_bucket_work(&bucket_tx_clone, &bucket, &mut work_guard).await {
|
||||
increment_disk_bucket_scans_queued(
|
||||
&queued_disk_bucket_scans_clone,
|
||||
&pool_label_clone,
|
||||
&set_label_clone,
|
||||
);
|
||||
} else {
|
||||
record_failed_dirty_bucket(&failed_dirty_buckets_clone, &bucket.name).await;
|
||||
}
|
||||
debug!(
|
||||
target: "rustfs::scanner::io",
|
||||
event = EVENT_SCANNER_CACHE_PERSIST_STATE,
|
||||
component = LOG_COMPONENT_SCANNER,
|
||||
subsystem = LOG_SUBSYSTEM_IO,
|
||||
bucket = %bucket.name,
|
||||
cache_name = %cache_name,
|
||||
state = "lock_contention_requeued",
|
||||
error = %e,
|
||||
"Scanner bucket cache lock contention requeued bucket work"
|
||||
);
|
||||
break;
|
||||
}
|
||||
|
||||
record_failed_dirty_bucket(&failed_dirty_buckets_clone, &bucket.name).await;
|
||||
error!(
|
||||
target: "rustfs::scanner::io",
|
||||
@@ -3114,7 +3185,7 @@ impl ScannerIOCache for SetDisks {
|
||||
subsystem = LOG_SUBSYSTEM_IO,
|
||||
bucket = %bucket.name,
|
||||
cache_name = %cache_name,
|
||||
state = "lock_acquire_failed",
|
||||
state = e.state(),
|
||||
error = %e,
|
||||
"Scanner bucket cache lock acquisition failed"
|
||||
);
|
||||
@@ -3894,6 +3965,52 @@ mod tests {
|
||||
(temp_dir, store)
|
||||
}
|
||||
|
||||
#[tokio::test]
|
||||
#[serial]
|
||||
async fn scanner_cache_locks_block_same_source_workers() {
|
||||
let (_temp_dir, store) = setup_two_pool_scanner_store().await;
|
||||
let set = &store.pools[0].disk_set[0];
|
||||
let source = DataUsageCacheSource::new(0, 0);
|
||||
let cache_name = "photos/.usage-cache.bin";
|
||||
|
||||
let guards = acquire_scanner_cache_locks(set.as_ref(), cache_name, source)
|
||||
.await
|
||||
.expect("scanner cache locks should be acquired");
|
||||
let scoped_lock = set
|
||||
.new_ns_lock(RUSTFS_META_BUCKET, &scanner_cache_lock_resource(cache_name, source))
|
||||
.await
|
||||
.expect("scoped scanner cache lock should be created");
|
||||
let scoped_err = scoped_lock
|
||||
.get_write_lock_quiet(Duration::from_millis(100))
|
||||
.await
|
||||
.expect_err("same-source workers must be blocked while scanner cache lock is held");
|
||||
assert!(matches!(scoped_err, LockError::Timeout { .. } | LockError::AlreadyLocked { .. }));
|
||||
|
||||
drop(guards);
|
||||
acquire_scanner_cache_locks(set.as_ref(), cache_name, source)
|
||||
.await
|
||||
.expect("scanner cache locks should be released when guards drop");
|
||||
}
|
||||
|
||||
#[tokio::test]
|
||||
#[serial]
|
||||
async fn scanner_cache_locks_allow_cross_source_workers() {
|
||||
let (_temp_dir, store) = setup_two_pool_scanner_store().await;
|
||||
let first_set = &store.pools[0].disk_set[0];
|
||||
let second_set = &store.pools[1].disk_set[0];
|
||||
let cache_name = "photos/.usage-cache.bin";
|
||||
|
||||
let first = acquire_scanner_cache_locks(first_set.as_ref(), cache_name, DataUsageCacheSource::new(0, 0))
|
||||
.await
|
||||
.expect("first source scanner cache locks should be acquired");
|
||||
let second = acquire_scanner_cache_locks(second_set.as_ref(), cache_name, DataUsageCacheSource::new(1, 0))
|
||||
.await
|
||||
.expect("different source scanner cache locks should not contend");
|
||||
|
||||
assert!(!first.is_lock_lost());
|
||||
assert!(!second.is_lock_lost());
|
||||
}
|
||||
|
||||
#[tokio::test]
|
||||
async fn data_usage_publish_fails_when_receiver_is_closed() {
|
||||
let (updates, receiver) = mpsc::channel(1);
|
||||
|
||||
@@ -28,8 +28,8 @@ pub const NS_SCANNER_SESSION_SEQUENCE_QUERY: &str = "ns_scanner_session_sequence
|
||||
pub const NS_SCANNER_PROTOCOL_VERSION_QUERY: &str = "ns_scanner_protocol";
|
||||
pub const NS_SCANNER_PROTOCOL_VERSION: u16 = 3;
|
||||
pub const SCANNER_ACTIVITY_LEGACY_PROTOCOL_VERSION: u32 = 0;
|
||||
pub const SCANNER_ACTIVITY_PREVIOUS_PROTOCOL_VERSION: u32 = 4;
|
||||
pub const SCANNER_ACTIVITY_PROTOCOL_VERSION: u32 = 5;
|
||||
pub const SCANNER_ACTIVITY_PREVIOUS_PROTOCOL_VERSION: u32 = 5;
|
||||
pub const SCANNER_ACTIVITY_PROTOCOL_VERSION: u32 = 6;
|
||||
|
||||
#[derive(Debug, serde::Deserialize, serde::Serialize)]
|
||||
#[serde(deny_unknown_fields)]
|
||||
|
||||
@@ -9,3 +9,5 @@ Import `grafana/rustfs-node-observability.json` into Grafana and select a Promet
|
||||
The dashboard uses the RustFS `server` metric label introduced with the node-local observability updates. `server` represents the RustFS node identity and is preferred for RustFS node comparisons. Prometheus `instance` still identifies the scrape target and remains useful for scrape/debugging views, but dashboards that compare RustFS nodes should group and filter by `server`.
|
||||
|
||||
During a rolling upgrade, older nodes may still emit metrics without the `server` label. Complete the rollout before using this dashboard for node-by-node comparisons.
|
||||
|
||||
`grafana/rustfs-kms-observability.json` covers the KMS backend operation metrics emitted at the operation-policy choke point. Unlike the node dashboard, KMS metrics do not carry the `server` label — use `job`/`instance` or promoted OTel resource attributes to split by node. Matching Prometheus alert rules live in `.docker/observability/prometheus-rules/rustfs-kms-alerts.yml`, and the alert response procedures are documented in `docs/operations/kms-observability-runbook.md`.
|
||||
|
||||
@@ -0,0 +1,793 @@
|
||||
{
|
||||
"annotations": {
|
||||
"list": [
|
||||
{
|
||||
"builtIn": 1,
|
||||
"datasource": {
|
||||
"type": "grafana",
|
||||
"uid": "-- Grafana --"
|
||||
},
|
||||
"enable": true,
|
||||
"hide": true,
|
||||
"iconColor": "rgba(0, 211, 255, 1)",
|
||||
"name": "Annotations & Alerts",
|
||||
"target": {
|
||||
"limit": 100,
|
||||
"matchAny": false,
|
||||
"tags": [],
|
||||
"type": "dashboard"
|
||||
},
|
||||
"type": "dashboard"
|
||||
}
|
||||
]
|
||||
},
|
||||
"description": "KMS backend operation metrics emitted at the operation-policy choke point (crates/kms/src/policy.rs). All label values are static enum strings; key identifiers, key material, and tokens never appear in labels. These metrics do not carry the RustFS `server` label — use your scrape topology (job/instance or promoted OTel resource attributes) to split by node. Alert response procedures: docs/operations/kms-observability-runbook.md.",
|
||||
"editable": true,
|
||||
"fiscalYearStartMonth": 0,
|
||||
"graphTooltip": 1,
|
||||
"id": null,
|
||||
"links": [],
|
||||
"liveNow": false,
|
||||
"panels": [
|
||||
{
|
||||
"datasource": {
|
||||
"type": "prometheus",
|
||||
"uid": "${datasource}"
|
||||
},
|
||||
"description": "Terminal outcomes of KMS backend operations. `fatal` means a non-retryable failure ended the operation on first observation; `budget_exhausted` and `deadline_exceeded` mean retries ran out; `cancelled` is normal during shutdown.",
|
||||
"fieldConfig": {
|
||||
"defaults": {
|
||||
"color": {
|
||||
"mode": "palette-classic"
|
||||
},
|
||||
"custom": {
|
||||
"axisCenteredZero": false,
|
||||
"axisColorMode": "text",
|
||||
"axisLabel": "",
|
||||
"axisPlacement": "auto",
|
||||
"barAlignment": 0,
|
||||
"drawStyle": "line",
|
||||
"fillOpacity": 12,
|
||||
"gradientMode": "none",
|
||||
"hideFrom": {
|
||||
"legend": false,
|
||||
"tooltip": false,
|
||||
"viz": false
|
||||
},
|
||||
"lineInterpolation": "linear",
|
||||
"lineWidth": 1,
|
||||
"pointSize": 4,
|
||||
"scaleDistribution": {
|
||||
"type": "linear"
|
||||
},
|
||||
"showPoints": "never",
|
||||
"spanNulls": false,
|
||||
"stacking": {
|
||||
"group": "A",
|
||||
"mode": "none"
|
||||
},
|
||||
"thresholdsStyle": {
|
||||
"mode": "off"
|
||||
}
|
||||
},
|
||||
"mappings": [],
|
||||
"thresholds": {
|
||||
"mode": "absolute",
|
||||
"steps": [
|
||||
{
|
||||
"color": "green",
|
||||
"value": null
|
||||
}
|
||||
]
|
||||
},
|
||||
"unit": "ops"
|
||||
},
|
||||
"overrides": []
|
||||
},
|
||||
"gridPos": {
|
||||
"h": 8,
|
||||
"w": 12,
|
||||
"x": 0,
|
||||
"y": 0
|
||||
},
|
||||
"id": 1,
|
||||
"options": {
|
||||
"legend": {
|
||||
"calcs": [
|
||||
"lastNotNull"
|
||||
],
|
||||
"displayMode": "table",
|
||||
"placement": "bottom",
|
||||
"showLegend": true
|
||||
},
|
||||
"tooltip": {
|
||||
"mode": "multi",
|
||||
"sort": "desc"
|
||||
}
|
||||
},
|
||||
"targets": [
|
||||
{
|
||||
"datasource": {
|
||||
"type": "prometheus",
|
||||
"uid": "${datasource}"
|
||||
},
|
||||
"editorMode": "code",
|
||||
"expr": "sum by (outcome) (rate(rustfs_kms_backend_operations_total{operation=~\"$operation\"}[$__rate_interval]))",
|
||||
"legendFormat": "{{outcome}}",
|
||||
"range": true,
|
||||
"refId": "A"
|
||||
}
|
||||
],
|
||||
"title": "Backend Operation Rate by Outcome",
|
||||
"type": "timeseries"
|
||||
},
|
||||
{
|
||||
"datasource": {
|
||||
"type": "prometheus",
|
||||
"uid": "${datasource}"
|
||||
},
|
||||
"description": "Operation names are static per-call-site identifiers (e.g. vault_kv2_read_key_version, vault_transit_encrypt, vault_login). `op_class` distinguishes read_idempotent, mutating_non_idempotent, and auth operations.",
|
||||
"fieldConfig": {
|
||||
"defaults": {
|
||||
"color": {
|
||||
"mode": "palette-classic"
|
||||
},
|
||||
"custom": {
|
||||
"axisCenteredZero": false,
|
||||
"axisColorMode": "text",
|
||||
"axisLabel": "",
|
||||
"axisPlacement": "auto",
|
||||
"barAlignment": 0,
|
||||
"drawStyle": "line",
|
||||
"fillOpacity": 12,
|
||||
"gradientMode": "none",
|
||||
"hideFrom": {
|
||||
"legend": false,
|
||||
"tooltip": false,
|
||||
"viz": false
|
||||
},
|
||||
"lineInterpolation": "linear",
|
||||
"lineWidth": 1,
|
||||
"pointSize": 4,
|
||||
"scaleDistribution": {
|
||||
"type": "linear"
|
||||
},
|
||||
"showPoints": "never",
|
||||
"spanNulls": false,
|
||||
"stacking": {
|
||||
"group": "A",
|
||||
"mode": "none"
|
||||
},
|
||||
"thresholdsStyle": {
|
||||
"mode": "off"
|
||||
}
|
||||
},
|
||||
"mappings": [],
|
||||
"thresholds": {
|
||||
"mode": "absolute",
|
||||
"steps": [
|
||||
{
|
||||
"color": "green",
|
||||
"value": null
|
||||
}
|
||||
]
|
||||
},
|
||||
"unit": "ops"
|
||||
},
|
||||
"overrides": []
|
||||
},
|
||||
"gridPos": {
|
||||
"h": 8,
|
||||
"w": 12,
|
||||
"x": 12,
|
||||
"y": 0
|
||||
},
|
||||
"id": 2,
|
||||
"options": {
|
||||
"legend": {
|
||||
"calcs": [
|
||||
"lastNotNull"
|
||||
],
|
||||
"displayMode": "table",
|
||||
"placement": "bottom",
|
||||
"showLegend": true
|
||||
},
|
||||
"tooltip": {
|
||||
"mode": "multi",
|
||||
"sort": "desc"
|
||||
}
|
||||
},
|
||||
"targets": [
|
||||
{
|
||||
"datasource": {
|
||||
"type": "prometheus",
|
||||
"uid": "${datasource}"
|
||||
},
|
||||
"editorMode": "code",
|
||||
"expr": "sum by (operation, op_class) (rate(rustfs_kms_backend_operations_total{operation=~\"$operation\"}[$__rate_interval]))",
|
||||
"legendFormat": "{{operation}} ({{op_class}})",
|
||||
"range": true,
|
||||
"refId": "A"
|
||||
}
|
||||
],
|
||||
"title": "Backend Operation Rate by Operation",
|
||||
"type": "timeseries"
|
||||
},
|
||||
{
|
||||
"datasource": {
|
||||
"type": "prometheus",
|
||||
"uid": "${datasource}"
|
||||
},
|
||||
"description": "Share of operations that terminated in fatal, budget_exhausted, or deadline_exceeded. The cancelled outcome is plotted separately because shutdown windows legitimately spike it. The ratio is meaningless at near-zero traffic — read it together with the operation rate panels.",
|
||||
"fieldConfig": {
|
||||
"defaults": {
|
||||
"color": {
|
||||
"mode": "palette-classic"
|
||||
},
|
||||
"custom": {
|
||||
"axisCenteredZero": false,
|
||||
"axisColorMode": "text",
|
||||
"axisLabel": "",
|
||||
"axisPlacement": "auto",
|
||||
"barAlignment": 0,
|
||||
"drawStyle": "line",
|
||||
"fillOpacity": 12,
|
||||
"gradientMode": "none",
|
||||
"hideFrom": {
|
||||
"legend": false,
|
||||
"tooltip": false,
|
||||
"viz": false
|
||||
},
|
||||
"lineInterpolation": "linear",
|
||||
"lineWidth": 1,
|
||||
"pointSize": 4,
|
||||
"scaleDistribution": {
|
||||
"type": "linear"
|
||||
},
|
||||
"showPoints": "never",
|
||||
"spanNulls": false,
|
||||
"stacking": {
|
||||
"group": "A",
|
||||
"mode": "none"
|
||||
},
|
||||
"thresholdsStyle": {
|
||||
"mode": "off"
|
||||
}
|
||||
},
|
||||
"mappings": [],
|
||||
"max": 1,
|
||||
"min": 0,
|
||||
"thresholds": {
|
||||
"mode": "absolute",
|
||||
"steps": [
|
||||
{
|
||||
"color": "green",
|
||||
"value": null
|
||||
}
|
||||
]
|
||||
},
|
||||
"unit": "percentunit"
|
||||
},
|
||||
"overrides": []
|
||||
},
|
||||
"gridPos": {
|
||||
"h": 8,
|
||||
"w": 12,
|
||||
"x": 0,
|
||||
"y": 8
|
||||
},
|
||||
"id": 3,
|
||||
"options": {
|
||||
"legend": {
|
||||
"calcs": [
|
||||
"lastNotNull"
|
||||
],
|
||||
"displayMode": "table",
|
||||
"placement": "bottom",
|
||||
"showLegend": true
|
||||
},
|
||||
"tooltip": {
|
||||
"mode": "multi",
|
||||
"sort": "desc"
|
||||
}
|
||||
},
|
||||
"targets": [
|
||||
{
|
||||
"datasource": {
|
||||
"type": "prometheus",
|
||||
"uid": "${datasource}"
|
||||
},
|
||||
"editorMode": "code",
|
||||
"expr": "sum(rate(rustfs_kms_backend_operations_total{outcome!~\"success|cancelled\",operation=~\"$operation\"}[$__rate_interval])) / clamp_min(sum(rate(rustfs_kms_backend_operations_total{operation=~\"$operation\"}[$__rate_interval])), 1e-9)",
|
||||
"legendFormat": "non-success (excl. cancelled)",
|
||||
"range": true,
|
||||
"refId": "A"
|
||||
},
|
||||
{
|
||||
"datasource": {
|
||||
"type": "prometheus",
|
||||
"uid": "${datasource}"
|
||||
},
|
||||
"editorMode": "code",
|
||||
"expr": "sum(rate(rustfs_kms_backend_operations_total{outcome=\"cancelled\",operation=~\"$operation\"}[$__rate_interval])) / clamp_min(sum(rate(rustfs_kms_backend_operations_total{operation=~\"$operation\"}[$__rate_interval])), 1e-9)",
|
||||
"legendFormat": "cancelled",
|
||||
"range": true,
|
||||
"refId": "B"
|
||||
}
|
||||
],
|
||||
"title": "Non-Success Outcome Ratio",
|
||||
"type": "timeseries"
|
||||
},
|
||||
{
|
||||
"datasource": {
|
||||
"type": "prometheus",
|
||||
"uid": "${datasource}"
|
||||
},
|
||||
"description": "Per-attempt failures by retry classification. retryable_conn covers connection-level failures, retryable_status covers retryable backend status codes, attempt_timeout covers attempts cut off by the per-attempt timeout, and fatal covers non-retryable failures (auth, permission, malformed request). Sustained fatal traffic is always actionable.",
|
||||
"fieldConfig": {
|
||||
"defaults": {
|
||||
"color": {
|
||||
"mode": "palette-classic"
|
||||
},
|
||||
"custom": {
|
||||
"axisCenteredZero": false,
|
||||
"axisColorMode": "text",
|
||||
"axisLabel": "",
|
||||
"axisPlacement": "auto",
|
||||
"barAlignment": 0,
|
||||
"drawStyle": "line",
|
||||
"fillOpacity": 12,
|
||||
"gradientMode": "none",
|
||||
"hideFrom": {
|
||||
"legend": false,
|
||||
"tooltip": false,
|
||||
"viz": false
|
||||
},
|
||||
"lineInterpolation": "linear",
|
||||
"lineWidth": 1,
|
||||
"pointSize": 4,
|
||||
"scaleDistribution": {
|
||||
"type": "linear"
|
||||
},
|
||||
"showPoints": "never",
|
||||
"spanNulls": false,
|
||||
"stacking": {
|
||||
"group": "A",
|
||||
"mode": "none"
|
||||
},
|
||||
"thresholdsStyle": {
|
||||
"mode": "off"
|
||||
}
|
||||
},
|
||||
"mappings": [],
|
||||
"thresholds": {
|
||||
"mode": "absolute",
|
||||
"steps": [
|
||||
{
|
||||
"color": "green",
|
||||
"value": null
|
||||
}
|
||||
]
|
||||
},
|
||||
"unit": "ops"
|
||||
},
|
||||
"overrides": []
|
||||
},
|
||||
"gridPos": {
|
||||
"h": 8,
|
||||
"w": 12,
|
||||
"x": 12,
|
||||
"y": 8
|
||||
},
|
||||
"id": 4,
|
||||
"options": {
|
||||
"legend": {
|
||||
"calcs": [
|
||||
"lastNotNull"
|
||||
],
|
||||
"displayMode": "table",
|
||||
"placement": "bottom",
|
||||
"showLegend": true
|
||||
},
|
||||
"tooltip": {
|
||||
"mode": "multi",
|
||||
"sort": "desc"
|
||||
}
|
||||
},
|
||||
"targets": [
|
||||
{
|
||||
"datasource": {
|
||||
"type": "prometheus",
|
||||
"uid": "${datasource}"
|
||||
},
|
||||
"editorMode": "code",
|
||||
"expr": "sum by (error_class) (rate(rustfs_kms_backend_attempt_failures_total{operation=~\"$operation\"}[$__rate_interval]))",
|
||||
"legendFormat": "{{error_class}}",
|
||||
"range": true,
|
||||
"refId": "A"
|
||||
}
|
||||
],
|
||||
"title": "Attempt Failure Rate by Error Class",
|
||||
"type": "timeseries"
|
||||
},
|
||||
{
|
||||
"datasource": {
|
||||
"type": "prometheus",
|
||||
"uid": "${datasource}"
|
||||
},
|
||||
"description": "Wall-clock duration of whole operations, including retries and backoff sleeps. A rising p99 with a flat p50 usually means a slow retry tail (backend degradation), not a uniform slowdown.",
|
||||
"fieldConfig": {
|
||||
"defaults": {
|
||||
"color": {
|
||||
"mode": "palette-classic"
|
||||
},
|
||||
"custom": {
|
||||
"axisCenteredZero": false,
|
||||
"axisColorMode": "text",
|
||||
"axisLabel": "",
|
||||
"axisPlacement": "auto",
|
||||
"barAlignment": 0,
|
||||
"drawStyle": "line",
|
||||
"fillOpacity": 12,
|
||||
"gradientMode": "none",
|
||||
"hideFrom": {
|
||||
"legend": false,
|
||||
"tooltip": false,
|
||||
"viz": false
|
||||
},
|
||||
"lineInterpolation": "linear",
|
||||
"lineWidth": 1,
|
||||
"pointSize": 4,
|
||||
"scaleDistribution": {
|
||||
"type": "linear"
|
||||
},
|
||||
"showPoints": "never",
|
||||
"spanNulls": false,
|
||||
"stacking": {
|
||||
"group": "A",
|
||||
"mode": "none"
|
||||
},
|
||||
"thresholdsStyle": {
|
||||
"mode": "off"
|
||||
}
|
||||
},
|
||||
"mappings": [],
|
||||
"thresholds": {
|
||||
"mode": "absolute",
|
||||
"steps": [
|
||||
{
|
||||
"color": "green",
|
||||
"value": null
|
||||
}
|
||||
]
|
||||
},
|
||||
"unit": "s"
|
||||
},
|
||||
"overrides": []
|
||||
},
|
||||
"gridPos": {
|
||||
"h": 8,
|
||||
"w": 12,
|
||||
"x": 0,
|
||||
"y": 16
|
||||
},
|
||||
"id": 5,
|
||||
"options": {
|
||||
"legend": {
|
||||
"calcs": [
|
||||
"lastNotNull"
|
||||
],
|
||||
"displayMode": "table",
|
||||
"placement": "bottom",
|
||||
"showLegend": true
|
||||
},
|
||||
"tooltip": {
|
||||
"mode": "multi",
|
||||
"sort": "desc"
|
||||
}
|
||||
},
|
||||
"targets": [
|
||||
{
|
||||
"datasource": {
|
||||
"type": "prometheus",
|
||||
"uid": "${datasource}"
|
||||
},
|
||||
"editorMode": "code",
|
||||
"expr": "histogram_quantile(0.50, sum by (le) (rate(rustfs_kms_backend_operation_duration_seconds_bucket{operation=~\"$operation\"}[$__rate_interval])))",
|
||||
"legendFormat": "p50",
|
||||
"range": true,
|
||||
"refId": "A"
|
||||
},
|
||||
{
|
||||
"datasource": {
|
||||
"type": "prometheus",
|
||||
"uid": "${datasource}"
|
||||
},
|
||||
"editorMode": "code",
|
||||
"expr": "histogram_quantile(0.99, sum by (le) (rate(rustfs_kms_backend_operation_duration_seconds_bucket{operation=~\"$operation\"}[$__rate_interval])))",
|
||||
"legendFormat": "p99",
|
||||
"range": true,
|
||||
"refId": "B"
|
||||
}
|
||||
],
|
||||
"title": "Operation Duration p50 / p99",
|
||||
"type": "timeseries"
|
||||
},
|
||||
{
|
||||
"datasource": {
|
||||
"type": "prometheus",
|
||||
"uid": "${datasource}"
|
||||
},
|
||||
"description": "p99 duration split by operation. Auth operations (vault_login, vault_token_renew) and mutating writes are expected to sit higher than idempotent reads.",
|
||||
"fieldConfig": {
|
||||
"defaults": {
|
||||
"color": {
|
||||
"mode": "palette-classic"
|
||||
},
|
||||
"custom": {
|
||||
"axisCenteredZero": false,
|
||||
"axisColorMode": "text",
|
||||
"axisLabel": "",
|
||||
"axisPlacement": "auto",
|
||||
"barAlignment": 0,
|
||||
"drawStyle": "line",
|
||||
"fillOpacity": 12,
|
||||
"gradientMode": "none",
|
||||
"hideFrom": {
|
||||
"legend": false,
|
||||
"tooltip": false,
|
||||
"viz": false
|
||||
},
|
||||
"lineInterpolation": "linear",
|
||||
"lineWidth": 1,
|
||||
"pointSize": 4,
|
||||
"scaleDistribution": {
|
||||
"type": "linear"
|
||||
},
|
||||
"showPoints": "never",
|
||||
"spanNulls": false,
|
||||
"stacking": {
|
||||
"group": "A",
|
||||
"mode": "none"
|
||||
},
|
||||
"thresholdsStyle": {
|
||||
"mode": "off"
|
||||
}
|
||||
},
|
||||
"mappings": [],
|
||||
"thresholds": {
|
||||
"mode": "absolute",
|
||||
"steps": [
|
||||
{
|
||||
"color": "green",
|
||||
"value": null
|
||||
}
|
||||
]
|
||||
},
|
||||
"unit": "s"
|
||||
},
|
||||
"overrides": []
|
||||
},
|
||||
"gridPos": {
|
||||
"h": 8,
|
||||
"w": 12,
|
||||
"x": 12,
|
||||
"y": 16
|
||||
},
|
||||
"id": 6,
|
||||
"options": {
|
||||
"legend": {
|
||||
"calcs": [
|
||||
"lastNotNull"
|
||||
],
|
||||
"displayMode": "table",
|
||||
"placement": "bottom",
|
||||
"showLegend": true
|
||||
},
|
||||
"tooltip": {
|
||||
"mode": "multi",
|
||||
"sort": "desc"
|
||||
}
|
||||
},
|
||||
"targets": [
|
||||
{
|
||||
"datasource": {
|
||||
"type": "prometheus",
|
||||
"uid": "${datasource}"
|
||||
},
|
||||
"editorMode": "code",
|
||||
"expr": "histogram_quantile(0.99, sum by (le, operation) (rate(rustfs_kms_backend_operation_duration_seconds_bucket{operation=~\"$operation\"}[$__rate_interval])))",
|
||||
"legendFormat": "{{operation}}",
|
||||
"range": true,
|
||||
"refId": "A"
|
||||
}
|
||||
],
|
||||
"title": "Operation Duration p99 by Operation",
|
||||
"type": "timeseries"
|
||||
},
|
||||
{
|
||||
"datasource": {
|
||||
"type": "prometheus",
|
||||
"uid": "${datasource}"
|
||||
},
|
||||
"description": "Attempts one operation used before completing. Healthy operations average close to 1. A rising average or p99 means the retry policy is absorbing backend failures — read together with the attempt-failure panel to see the failure class.",
|
||||
"fieldConfig": {
|
||||
"defaults": {
|
||||
"color": {
|
||||
"mode": "palette-classic"
|
||||
},
|
||||
"custom": {
|
||||
"axisCenteredZero": false,
|
||||
"axisColorMode": "text",
|
||||
"axisLabel": "",
|
||||
"axisPlacement": "auto",
|
||||
"barAlignment": 0,
|
||||
"drawStyle": "line",
|
||||
"fillOpacity": 12,
|
||||
"gradientMode": "none",
|
||||
"hideFrom": {
|
||||
"legend": false,
|
||||
"tooltip": false,
|
||||
"viz": false
|
||||
},
|
||||
"lineInterpolation": "linear",
|
||||
"lineWidth": 1,
|
||||
"pointSize": 4,
|
||||
"scaleDistribution": {
|
||||
"type": "linear"
|
||||
},
|
||||
"showPoints": "never",
|
||||
"spanNulls": false,
|
||||
"stacking": {
|
||||
"group": "A",
|
||||
"mode": "none"
|
||||
},
|
||||
"thresholdsStyle": {
|
||||
"mode": "off"
|
||||
}
|
||||
},
|
||||
"mappings": [],
|
||||
"thresholds": {
|
||||
"mode": "absolute",
|
||||
"steps": [
|
||||
{
|
||||
"color": "green",
|
||||
"value": null
|
||||
}
|
||||
]
|
||||
},
|
||||
"unit": "short"
|
||||
},
|
||||
"overrides": []
|
||||
},
|
||||
"gridPos": {
|
||||
"h": 8,
|
||||
"w": 12,
|
||||
"x": 0,
|
||||
"y": 24
|
||||
},
|
||||
"id": 7,
|
||||
"options": {
|
||||
"legend": {
|
||||
"calcs": [
|
||||
"lastNotNull"
|
||||
],
|
||||
"displayMode": "table",
|
||||
"placement": "bottom",
|
||||
"showLegend": true
|
||||
},
|
||||
"tooltip": {
|
||||
"mode": "multi",
|
||||
"sort": "desc"
|
||||
}
|
||||
},
|
||||
"targets": [
|
||||
{
|
||||
"datasource": {
|
||||
"type": "prometheus",
|
||||
"uid": "${datasource}"
|
||||
},
|
||||
"editorMode": "code",
|
||||
"expr": "sum by (operation) (rate(rustfs_kms_backend_operation_attempts_sum{operation=~\"$operation\"}[$__rate_interval])) / clamp_min(sum by (operation) (rate(rustfs_kms_backend_operation_attempts_count{operation=~\"$operation\"}[$__rate_interval])), 1e-9)",
|
||||
"legendFormat": "{{operation}} avg",
|
||||
"range": true,
|
||||
"refId": "A"
|
||||
},
|
||||
{
|
||||
"datasource": {
|
||||
"type": "prometheus",
|
||||
"uid": "${datasource}"
|
||||
},
|
||||
"editorMode": "code",
|
||||
"expr": "histogram_quantile(0.99, sum by (le) (rate(rustfs_kms_backend_operation_attempts_bucket{operation=~\"$operation\"}[$__rate_interval])))",
|
||||
"legendFormat": "p99 (all operations)",
|
||||
"range": true,
|
||||
"refId": "B"
|
||||
}
|
||||
],
|
||||
"title": "Operation Attempts Distribution",
|
||||
"type": "timeseries"
|
||||
},
|
||||
{
|
||||
"gridPos": {
|
||||
"h": 8,
|
||||
"w": 12,
|
||||
"x": 12,
|
||||
"y": 24
|
||||
},
|
||||
"id": 8,
|
||||
"options": {
|
||||
"code": {
|
||||
"language": "plaintext",
|
||||
"showLineNumbers": false,
|
||||
"showMiniMap": false
|
||||
},
|
||||
"content": "Placeholder for KMS metrics planned by sibling changes of rustfs/backlog#1584 that have **not landed yet**. Do not add panels for these names until the emitting code is merged, or the panels will render empty and mask real gaps.\n\n- TODO: key-cache effectiveness — hit/miss counters, entry gauge, eviction counter (replaces the former hardcoded-zero miss stat).\n- TODO: key lifecycle — pending-deletion/tombstone gauges from the deletion worker sweep, aggregate rotation-age gauge.\n- TODO: Vault credentials — token TTL remaining and fail-closed state gauges.\n- TODO: synthetic backend probe — probe outcome/failure-class metrics feeding the readiness cache.\n\nWhen a family lands, replace one bullet with a real panel and keep this list in sync with docs/operations/kms-observability-runbook.md (Coverage gaps).",
|
||||
"mode": "markdown"
|
||||
},
|
||||
"title": "Planned Panels (TODO — metrics not landed yet)",
|
||||
"type": "text"
|
||||
}
|
||||
],
|
||||
"refresh": "30s",
|
||||
"schemaVersion": 39,
|
||||
"style": "dark",
|
||||
"tags": [
|
||||
"rustfs",
|
||||
"observability",
|
||||
"kms"
|
||||
],
|
||||
"templating": {
|
||||
"list": [
|
||||
{
|
||||
"current": {},
|
||||
"hide": 0,
|
||||
"includeAll": false,
|
||||
"label": "Prometheus",
|
||||
"multi": false,
|
||||
"name": "datasource",
|
||||
"options": [],
|
||||
"query": "prometheus",
|
||||
"refresh": 1,
|
||||
"regex": "",
|
||||
"skipUrlSync": false,
|
||||
"type": "datasource"
|
||||
},
|
||||
{
|
||||
"allValue": ".*",
|
||||
"current": {
|
||||
"selected": true,
|
||||
"text": "All",
|
||||
"value": "$__all"
|
||||
},
|
||||
"datasource": {
|
||||
"type": "prometheus",
|
||||
"uid": "${datasource}"
|
||||
},
|
||||
"definition": "label_values(rustfs_kms_backend_operations_total, operation)",
|
||||
"hide": 0,
|
||||
"includeAll": true,
|
||||
"label": "KMS operation",
|
||||
"multi": true,
|
||||
"name": "operation",
|
||||
"options": [],
|
||||
"query": "label_values(rustfs_kms_backend_operations_total, operation)",
|
||||
"refresh": 1,
|
||||
"regex": "",
|
||||
"skipUrlSync": false,
|
||||
"sort": 1,
|
||||
"type": "query"
|
||||
}
|
||||
]
|
||||
},
|
||||
"time": {
|
||||
"from": "now-6h",
|
||||
"to": "now"
|
||||
},
|
||||
"timepicker": {},
|
||||
"timezone": "",
|
||||
"title": "RustFS KMS Observability",
|
||||
"uid": "rustfs-kms-observability",
|
||||
"version": 1,
|
||||
"weekStart": ""
|
||||
}
|
||||
@@ -17,7 +17,7 @@ for later deletion.
|
||||
- `rustfs-5416-kubernetes-alias-dns` Kubernetes endpoint identity fallback: deployments created before explicit local endpoint identity may use resolvable aliases that do not match the Pod hostname. An implicit auto-mode zero match retains legacy DNS locality with a bounded deadline, while ambiguous matches and invalid explicit anchors still fail closed. Remove the fallback after every supported direct-upgrade chart and deployment manifest provides a canonical RUSTFS_LOCAL_ENDPOINT_HOST for domain-based distributed topologies.
|
||||
- `rustfs-5416-zero-retry-delay` startup retry-delay validation: releases before bounded topology convergence accept RUSTFS_STARTUP_TOPOLOGY_RETRY_MAX_DELAY values of 0 or 0ms. New servers replace those values with the safe nonzero default so a direct upgrade neither fails startup nor enters a busy loop. Reject zero after the minimum supported direct-upgrade release validates or rewrites this setting before rollout.
|
||||
- `scanner-usage-v2` persisted scanner usage migration: pre-v2 scanners write `.usage.json`, so upgraded clusters read that primary/backup pair only while `.usage.v2.json` is absent and continue removing deleted buckets from legacy copies that still exist. The additive usage_snapshot_complete field in `.usage.v2.json` must remain optional while mixed-version clusters are supported; a missing field means the snapshot is not authoritative. Remove the legacy object fallback and cleanup only after every supported direct-upgrade source writes `.usage.v2.json`.
|
||||
- `ns-scanner-rpc-v3` namespace scanner capability and activity handshake: old peers and legacy internode transports lack the authenticated startup-epoch handshake. The oldest peers send an empty activity request and receive a field-empty protocol-0 response. Protocol v4 binds the challenge and response topology but cannot authenticate distributed dirty-usage state. Current protocol v5 binds the request version, acknowledgement target and generation, and the response dirty-usage state. Servers retain protocol-0 and protocol-v4 codecs for rolling upgrades, while the distributed scanner publishes usage only after every peer returns authenticated protocol v5 state. Scanner selection treats HTTP 404/405/426 and the legacy MethodNotAllowed default as an explicit lack of remote scanner v3 support and assigns those disks to coordinator-driven workers; transient capability failures remain incomplete and do not activate the fallback. Remove the coordinator fallback after the minimum supported RustFS peer version implements namespace scanner protocol v3, remove protocol-0 activity requests and responses after every supported peer implements authenticated scanner activity protocol v4, and remove the protocol-v4 activity codec after every supported peer implements protocol v5; future protocol revisions must keep the same dual-version server/codec window before changing the advertised version.
|
||||
- `ns-scanner-rpc-v3` namespace scanner capability and activity handshake: old peers and legacy internode transports lack the authenticated startup-epoch handshake. The oldest peers send an empty activity request and receive a field-empty protocol-0 response. Protocol v4 binds the challenge and response topology but cannot authenticate distributed dirty-usage state. Protocol v5 binds the request version, acknowledgement target and generation, and the response dirty-usage state, but predates set-scoped scanner cache locks. Current protocol v6 additionally fences scanner cache lock-domain changes, so distributed scanner cycles publish usage only after every peer reports protocol v6 state. Servers retain protocol-0 and protocol-v4 codecs for rolling upgrades, while protocol-v5 peers are treated as previous-version peers that cannot safely participate in the new cache lock domain. Scanner selection treats HTTP 404/405/426 and the legacy MethodNotAllowed default as an explicit lack of remote scanner v3 support and assigns those disks to coordinator-driven workers; transient capability failures remain incomplete and do not activate the fallback. Remove the coordinator fallback after the minimum supported RustFS peer version implements namespace scanner protocol v3, remove protocol-0 activity requests and responses after every supported peer implements authenticated scanner activity protocol v4, remove the protocol-v4 activity codec after every supported peer implements protocol v5, and remove protocol-v5 previous-version rejection after every supported peer implements protocol v6; future protocol revisions must keep the same dual-version server/codec window before changing the advertised version.
|
||||
- `#4648` walk-dir stream completion capability: old clients can append fallback output to an already-used metacache writer after a terminal body error, so servers emit terminal walk errors only to clients that sign the `walk_dir_stream_completion=error-v1` query capability and its request-body digest. Remove the legacy clean-EOF path after the minimum supported RustFS peer version always advertises this capability.
|
||||
- `heal-rpc-auth-v2` internode gRPC authentication: servers temporarily accept legacy prefix signatures so old peers remain available during rolling upgrades. Remove the legacy fallback after the minimum supported RustFS peer version sends v2 authentication on every internode gRPC request.
|
||||
- `disk-mutation-body-digest` internode mutating disk RPCs: servers temporarily accept mutating disk RPCs (RenameData, DeleteVersion, DeleteVersions, WriteMetadata, UpdateMetadata, WriteAll, Delete, DeletePaths, RenameFile, RenamePart, DeleteVolume, MakeVolume, MakeVolumes) that carry no signature-bound canonical body digest, so peers from releases that predate body-digest signing remain available during rolling upgrades. Accepted digestless mutations increment the internode body-digest fallback counter; that counter must read zero fleet-wide across a release window before RUSTFS_INTERNODE_RPC_BODY_DIGEST_STRICT is enabled. Because body-bound requests now consume replay-cache nonces on the receiver, deploy the raised RUSTFS_INTERNODE_RPC_REPLAY_CACHE_CAPACITY default fleet-wide before enabling strict mode, and watch the internode replay-cache overflow counter for undersized capacity during the rollout. Remove the digestless fallback after the minimum supported RustFS peer version body-binds every mutating disk RPC.
|
||||
|
||||
@@ -0,0 +1,277 @@
|
||||
# Hotpath warp ABBA validation runbook
|
||||
|
||||
This runbook describes how to collect formal Linux or production-cluster
|
||||
evidence for hotpath performance changes. Use it when a short local A/B smoke
|
||||
run is too noisy to decide whether a regression is real.
|
||||
|
||||
The ABBA runner executes each workload and drive-sync cell as:
|
||||
|
||||
```text
|
||||
A1 baseline -> B1 candidate -> B2 candidate -> A2 baseline
|
||||
```
|
||||
|
||||
`B1` and `B2` are compared with `A1` to measure the candidate delta. `A2` is
|
||||
also compared with `A1` to measure baseline drift. Treat a candidate regression
|
||||
as actionable only when the `A2` drift is passing or materially smaller than
|
||||
the `B1` and `B2` delta for the same workload.
|
||||
|
||||
## Scope
|
||||
|
||||
Use this runbook for hotpath profiling and performance validation of RustFS
|
||||
object I/O changes, especially when CPU, memory allocation, lock/channel wait
|
||||
time, request throughput, or tail latency is the review question.
|
||||
|
||||
The script validates the same workload matrix as the hotpath warp A/B gate:
|
||||
|
||||
| Workload | mode | size |
|
||||
| --- | --- | --- |
|
||||
| `put-4kib` | put | 4KiB |
|
||||
| `put-4mib` | put | 4MiB |
|
||||
| `get-4kib` | get | 4KiB |
|
||||
| `get-4mib` | get | 4MiB |
|
||||
| `get-10mib` | get | 10MiB |
|
||||
| `mixed-256k` | mixed | 256KiB |
|
||||
|
||||
Each workload runs with `RUSTFS_DRIVE_SYNC_ENABLE=true` and
|
||||
`RUSTFS_DRIVE_SYNC_ENABLE=false`, so a full ABBA pass produces 48 measurement
|
||||
cells: 6 workloads x 2 drive-sync modes x 4 ABBA legs.
|
||||
|
||||
## Prerequisites
|
||||
|
||||
Run the formal pass on Linux, not on a laptop smoke environment.
|
||||
|
||||
Required tools on the bench host:
|
||||
|
||||
- `bash`, `curl`, `git`, and core GNU userland.
|
||||
- `warp` on `PATH`, or pass `--warp-bin`.
|
||||
- Two RustFS Linux binaries: one baseline and one candidate.
|
||||
- Enough isolated disks or directories for the local runner, or an externally
|
||||
managed RustFS cluster for production-like validation.
|
||||
- Stable host telemetry collection such as `pidstat`, `mpstat`, `iostat`,
|
||||
`sar`, `perf`, `heaptrack`, or the platform's equivalent observability stack.
|
||||
|
||||
Cluster-mode requirements:
|
||||
|
||||
- A deploy hook that can replace the RustFS binary on every node.
|
||||
- The hook must apply `RUSTFS_DRIVE_SYNC_ENABLE` for the current ABBA leg.
|
||||
- The hook must restart RustFS and return only after the rollout command has
|
||||
been accepted. The ABBA script performs the HTTP readiness wait.
|
||||
- The benchmark client should run outside the RustFS nodes when possible.
|
||||
- Do not run against a production data set unless the workload bucket and test
|
||||
credentials are isolated and approved for destructive benchmark traffic.
|
||||
|
||||
## Build the binaries
|
||||
|
||||
Build the baseline from the comparison commit, usually `origin/main` or the
|
||||
previous accepted release:
|
||||
|
||||
```bash
|
||||
git fetch origin main
|
||||
git switch --detach origin/main
|
||||
cargo build --release -p rustfs --bins
|
||||
cp target/release/rustfs /tmp/rustfs-baseline
|
||||
```
|
||||
|
||||
Build the candidate from the PR commit:
|
||||
|
||||
```bash
|
||||
git switch <candidate-branch>
|
||||
cargo build --release -p rustfs --bins
|
||||
cp target/release/rustfs /tmp/rustfs-candidate
|
||||
```
|
||||
|
||||
For cross-compiled cluster binaries, keep both outputs on the bench host and
|
||||
make the deploy hook copy the selected binary to the cluster. The ABBA runner
|
||||
passes the selected binary path through `HOTPATH_ABBA_BINARY`.
|
||||
|
||||
## Local Linux runner
|
||||
|
||||
Use local mode for a dedicated Linux runner with disposable data paths. This is
|
||||
not a substitute for a production-like cluster, but it is useful before spending
|
||||
cluster time.
|
||||
|
||||
```bash
|
||||
scripts/run_hotpath_warp_abba.sh \
|
||||
--baseline-bin /tmp/rustfs-baseline \
|
||||
--candidate-bin /tmp/rustfs-candidate \
|
||||
--address 127.0.0.1:9000 \
|
||||
--data-root /var/tmp/rustfs-hotpath-abba \
|
||||
--disks 4 \
|
||||
--duration 120s \
|
||||
--rounds 3 \
|
||||
--cooldown 30 \
|
||||
--concurrency 16 \
|
||||
--out-dir target/hotpath-abba/linux-local
|
||||
```
|
||||
|
||||
The script starts and stops RustFS for each ABBA leg. The data root is
|
||||
throwaway and should not contain important data.
|
||||
|
||||
## Production-like cluster runner
|
||||
|
||||
Use external mode when RustFS lifecycle is managed by ansible, systemd, a
|
||||
cluster scheduler, or a dedicated deployment harness. In this mode the ABBA
|
||||
script does not start RustFS directly; it calls `--deploy-hook` before each leg
|
||||
and then waits for `http://<endpoint><health-path>`.
|
||||
|
||||
The deploy hook receives:
|
||||
|
||||
| Environment variable | Value |
|
||||
| --- | --- |
|
||||
| `HOTPATH_ABBA_LEG` | `A1`, `B1`, `B2`, or `A2` |
|
||||
| `HOTPATH_ABBA_PHASE` | `baseline` or `candidate` |
|
||||
| `HOTPATH_ABBA_BINARY` | selected baseline or candidate binary path |
|
||||
| `HOTPATH_ABBA_DRIVE_SYNC` | `true` or `false` |
|
||||
|
||||
Example ansible-shaped command:
|
||||
|
||||
```bash
|
||||
scripts/run_hotpath_warp_abba.sh \
|
||||
--baseline-bin /srv/rustfs-binaries/rustfs-baseline \
|
||||
--candidate-bin /srv/rustfs-binaries/rustfs-candidate \
|
||||
--endpoint rustfs-bench.example.internal:9000 \
|
||||
--deploy-hook '
|
||||
set -euo pipefail
|
||||
cd /srv/rustfs-ansible
|
||||
cp "${HOTPATH_ABBA_BINARY:?}" roles/rustfs/files/rustfs
|
||||
export RUSTFS_DRIVE_SYNC_ENABLE="${HOTPATH_ABBA_DRIVE_SYNC:?}"
|
||||
ansible-playbook -f 4 -l bench rustfs-manage.yml --tags stop
|
||||
ansible-playbook -f 4 -l bench rustfs-manage.yml --tags config
|
||||
ansible-playbook -f 4 -l bench rustfs-manage.yml --tags binary-copy
|
||||
ansible-playbook -f 4 -l bench rustfs-manage.yml --tags start
|
||||
' \
|
||||
--duration 180s \
|
||||
--rounds 5 \
|
||||
--cooldown 45 \
|
||||
--concurrency 32 \
|
||||
--out-dir target/hotpath-abba/cluster-pr-XXXX
|
||||
```
|
||||
|
||||
For formal evidence, prefer `--rounds 5` or higher when the cluster budget
|
||||
allows it. The script enforces `--rounds >= 3`.
|
||||
|
||||
## CPU and memory evidence
|
||||
|
||||
ABBA warp output answers whether the candidate changed throughput or latency.
|
||||
Collect host telemetry at the same time to explain why.
|
||||
|
||||
Recommended minimum:
|
||||
|
||||
```bash
|
||||
mkdir -p target/hotpath-abba/cluster-pr-XXXX/telemetry
|
||||
|
||||
pidstat -durh 5 > target/hotpath-abba/cluster-pr-XXXX/telemetry/pidstat.txt &
|
||||
PIDSTAT_PID=$!
|
||||
|
||||
mpstat 5 > target/hotpath-abba/cluster-pr-XXXX/telemetry/mpstat.txt &
|
||||
MPSTAT_PID=$!
|
||||
|
||||
iostat -xz 5 > target/hotpath-abba/cluster-pr-XXXX/telemetry/iostat.txt &
|
||||
IOSTAT_PID=$!
|
||||
```
|
||||
|
||||
Stop the collectors after the ABBA script exits:
|
||||
|
||||
```bash
|
||||
kill "$PIDSTAT_PID" "$MPSTAT_PID" "$IOSTAT_PID"
|
||||
```
|
||||
|
||||
For deeper CPU attribution, run `perf record` around one representative
|
||||
workload after the ABBA gate identifies a candidate regression or improvement:
|
||||
|
||||
```bash
|
||||
perf record -F 99 -g -- sleep 180
|
||||
perf report --stdio > target/hotpath-abba/cluster-pr-XXXX/telemetry/perf-report.txt
|
||||
```
|
||||
|
||||
For allocation profiling, build the candidate with:
|
||||
|
||||
```bash
|
||||
cargo build --release -p rustfs --bins --features hotpath-alloc
|
||||
```
|
||||
|
||||
Then run the same ABBA command with that binary. Compare allocation-heavy
|
||||
function sections only within the same build mode. Do not compare
|
||||
`hotpath-alloc` binaries directly with default release binaries for throughput
|
||||
acceptance, because allocation instrumentation intentionally changes what is
|
||||
measured.
|
||||
|
||||
For CPU hotpath sections emitted by hotpath, build with:
|
||||
|
||||
```bash
|
||||
cargo build --release -p rustfs --bins --features hotpath-cpu
|
||||
```
|
||||
|
||||
Use the CPU-enabled report to explain hotspots after the default or plain
|
||||
`hotpath` ABBA gate shows a real effect.
|
||||
|
||||
## Output layout
|
||||
|
||||
The ABBA runner writes:
|
||||
|
||||
```text
|
||||
<out-dir>/
|
||||
manifest.env
|
||||
abba_schedule.csv
|
||||
candidate_gate.md
|
||||
baseline_drift_gate.md
|
||||
summary.md
|
||||
<workload>/<sync>/<leg>/median_summary.csv
|
||||
<workload>/<sync>/<leg>/baseline_compare.csv
|
||||
```
|
||||
|
||||
Attach or link at least these files in the issue or PR:
|
||||
|
||||
- `summary.md`
|
||||
- `candidate_gate.md`
|
||||
- `baseline_drift_gate.md`
|
||||
- `abba_schedule.csv`
|
||||
- every `median_summary.csv` and `baseline_compare.csv` for a failed or
|
||||
borderline workload
|
||||
- host telemetry files used to explain CPU, memory, or disk saturation
|
||||
|
||||
## Interpretation
|
||||
|
||||
Use this decision table:
|
||||
|
||||
| Candidate gate | A2 drift gate | Interpretation |
|
||||
| --- | --- | --- |
|
||||
| PASS | PASS | Candidate is acceptable for the measured matrix. |
|
||||
| WARN | PASS | Candidate has a small measurable signal; inspect telemetry and decide if it is expected. |
|
||||
| FAIL | PASS | Candidate likely regressed the affected workload; investigate before merge. |
|
||||
| FAIL | FAIL on the same workload | Environment drift is high; rerun on a quieter runner or increase duration and rounds. |
|
||||
| PASS | FAIL | Candidate did not exceed the budget, but the rig was unstable; avoid using the numbers as proof of improvement. |
|
||||
|
||||
When `B1` and `B2` disagree, treat the result as inconclusive even if the gate
|
||||
passes. Increase duration, rounds, cooldown, or runner isolation before drawing
|
||||
a conclusion.
|
||||
|
||||
## AI execution checklist
|
||||
|
||||
When delegating the run to an AI agent or an automation runner, provide these
|
||||
inputs explicitly:
|
||||
|
||||
- repository checkout and candidate branch or commit;
|
||||
- baseline commit or binary path;
|
||||
- candidate binary path;
|
||||
- runner type: local Linux or external cluster;
|
||||
- endpoint, access key, secret key source, and region;
|
||||
- deploy hook path or exact command for cluster mode;
|
||||
- output directory;
|
||||
- required duration, rounds, cooldown, concurrency, and fail/warn budgets;
|
||||
- where to upload artifacts after the run.
|
||||
|
||||
The AI agent should execute this sequence:
|
||||
|
||||
1. Confirm `uname -a`, RustFS commits, binary SHA256 sums, `warp --version`,
|
||||
CPU model, memory size, disk layout, and whether the run is local or cluster.
|
||||
2. Run `scripts/run_hotpath_warp_abba.sh --dry-run` with the final arguments.
|
||||
3. Run the real ABBA command with `--rounds >= 3`.
|
||||
4. Preserve the full output directory without editing generated CSV files.
|
||||
5. Read `summary.md`, `candidate_gate.md`, and `baseline_drift_gate.md`.
|
||||
6. Summarize only measured facts: candidate deltas, baseline drift, CPU or
|
||||
memory saturation, and any failed workloads.
|
||||
7. Post the summary and artifact location to the tracking issue or PR.
|
||||
|
||||
Do not report a performance win or loss when the baseline drift gate failed on
|
||||
the same workload and no rerun was collected.
|
||||
@@ -2,7 +2,7 @@
|
||||
|
||||
RustFS ships several KMS backends. They differ not only in deployment effort but in **where master key material lives and who can read it**. Pick a backend based on the confidentiality boundary you need, not on the name alone.
|
||||
|
||||
For how the Vault backends authenticate (static token, AppRole, Vault Agent token file) and how credential refresh and the fail-closed window behave, see the [Vault KMS authentication runbook](vault-kms-authentication.md).
|
||||
For how the Vault backends authenticate (static token, AppRole, Vault Agent token file) and how credential refresh and the fail-closed window behave, see the [Vault KMS authentication runbook](vault-kms-authentication.md). For what may be claimed about the cryptographic implementations themselves, see [Cryptographic compliance positioning](kms-cryptographic-compliance.md).
|
||||
|
||||
## Backend comparison
|
||||
|
||||
@@ -69,6 +69,77 @@ Decryption loads exactly the version recorded in the envelope and fails closed w
|
||||
|
||||
Do not rotate any key until **every** RustFS node in the cluster runs a build that understands the `master_key_version` envelope field. Older binaries ignore the field and always decrypt with the current material: harmless while nothing has been rotated, but after a rotation they will fail to decrypt every object wrapped by an earlier key version. Complete the rolling upgrade of the entire cluster first, then rotate.
|
||||
|
||||
This is the sharpest instance of a broader class of constraints; the rest are collected in [Mixed-version clusters during a rolling upgrade](#mixed-version-clusters-during-a-rolling-upgrade).
|
||||
|
||||
## Mixed-version clusters during a rolling upgrade
|
||||
|
||||
During a rolling upgrade the cluster runs two RustFS builds at once. That window matters more for KMS than for most subsystems, because KMS state is shared three ways: **Vault** holds the key records and Transit metadata, **cluster storage** holds the persisted KMS configuration, and **each node's process memory** holds caches and the live backend instance. Nodes on different builds agree on the first, may disagree on the third, and — for configuration — can disagree for as long as the operator leaves them running.
|
||||
|
||||
This section states only what is true of the current implementation. It is written for the KV2 and Transit backends; the Local backend is unsupported for multi-node deployments regardless of version (see the [deployment support matrix](#deployment-support-matrix)).
|
||||
|
||||
### Persisted formats are backward compatible in both directions
|
||||
|
||||
Nothing in this list requires a coordinated format cutover. The compatibility is deliberate and is covered by decode tests.
|
||||
|
||||
- **DEK envelopes.** `DataKeyEnvelope::master_key_version` is optional and omitted when absent, so envelopes written by non-rotating backends stay byte-identical to the historical seven-field JSON shape. An upgraded node reading a pre-versioning envelope resolves `None` to the key's recorded baseline version, or — for a key that was never rotated, and so has no baseline — to the current version, which is exactly the pre-versioning behavior.
|
||||
- **KV2 key records.** `baseline_version` is read with a serde default, so records written by older builds deserialize unchanged, and `None` correctly means "never rotated".
|
||||
- **Transit metadata records.** Metadata persisted in KV v2 by either build decodes on the other.
|
||||
|
||||
The one-way hazard is the rotation constraint above: an older binary reading a *new* envelope silently ignores the version field and decrypts with the current material.
|
||||
|
||||
### Guarantees that hold only once every node is upgraded
|
||||
|
||||
These are properties of the upgraded code, so a single node left behind removes them for the whole cluster.
|
||||
|
||||
- **Check-and-set lifecycle writes.** Upgraded builds write every KV2 lifecycle mutation — create, enable, disable, tag metadata, schedule deletion, cancel deletion — as a versioned read followed by a check-and-set write, retrying on conflict by re-reading and re-validating the state gate (rustfs/rustfs#5518). Transit metadata writes got the same treatment (rustfs/rustfs#5520). Builds older than those write blind. A blind write from an old node can overwrite a check-and-set commit from an upgraded node without any conflict being reported, which is precisely the lost update the change was made to eliminate.
|
||||
- **`baseline_version` survives a write-back.** The KV2 key record does not deny unknown fields, so an old build reads a new record without error — and drops `baseline_version` when it writes that record back for any reason. A key that loses its baseline resolves pre-versioning envelopes to the current version again, which after a rotation means the wrong master key material. Any lifecycle operation issued to an old node is enough to trigger this.
|
||||
- **Version-record awareness.** Rotation stores each historical version under `{prefix}/{key_id}/versions/{N}` as a create-only record (check-and-set of 0), so two nodes racing the same version number produce exactly one creator; the loser adopts the persisted, never-current material or fails without touching the current pointer. Old builds have no concept of that sub-path: they never read or write it, and their key listing reports the KV2 directory entry (`my-key/`) as though it were a key, because the directory filter only exists in upgraded builds.
|
||||
|
||||
### Windows in which nodes can legitimately disagree
|
||||
|
||||
Even with every node on the same build, some state is process-local. These windows are bounded by design, except the last one.
|
||||
|
||||
| What can diverge | Bound | Mechanism |
|
||||
| --- | --- | --- |
|
||||
| Transit key lifecycle state used by the `encrypt` and `generate_data_key` gates | ≤ 300 s (`METADATA_CACHE_TTL`) | Each node caches Transit metadata in process, TTL- and capacity-bounded, with targeted invalidation when a data-path call reports the key is gone server-side. A disable or schedule-deletion performed on one node is enforced on the others within one TTL at the latest, sooner if they hit that signal. |
|
||||
| `describe_key` output | ≤ 300 s | The manager-level key metadata cache. This is a reporting cache; the KV2 state gates do not read it. |
|
||||
| KV2 key lifecycle state | None | The KV2 backend re-reads the key record from Vault for every lifecycle and data-key operation, so a committed disable is effective on every upgraded node immediately. |
|
||||
| Active KMS configuration | Until the remaining nodes are restarted | See below. |
|
||||
|
||||
Builds older than rustfs/rustfs#5520 held the Transit metadata cache with no TTL and no capacity bound. On such a node the divergence window is not 300 seconds but "until the process restarts": it can keep encrypting under a key that another node disabled, indefinitely.
|
||||
|
||||
### Configuration is persisted cluster-wide but applied per node
|
||||
|
||||
`POST /rustfs/admin/v3/kms/reconfigure` currently does two things: it switches the KMS service **on the node that handled the request**, and it persists the new configuration to cluster storage at `config/kms_config.json`. It does not notify peers. Every other node keeps running the configuration it started with until it is restarted, at which point it loads the persisted configuration during startup.
|
||||
|
||||
The practical consequences today:
|
||||
|
||||
- The configuration-split window has no upper bound other than the operator restarting the remaining nodes. Convergence is a restart, not a timeout.
|
||||
- During the window both configurations are live. If the reconfiguration changed backends, or changed the Vault mount or key prefix, different nodes write new key material to different places, and a key created through one node is invisible to the others.
|
||||
- `kms status` reflects the node that answered the request, so a single successful status response is not evidence that the cluster is consistent. Query every node.
|
||||
|
||||
Treat `reconfigure` as the first step of a cluster-wide operation, not as the operation itself.
|
||||
|
||||
### Recommended rolling upgrade order
|
||||
|
||||
Follow the node-at-a-time procedure in the [multi-node restart runbook](rolling-restart.md); this adds the KMS-specific sequencing around it.
|
||||
|
||||
1. **Freeze KMS administrative traffic** for the duration: no key creation, enable, disable, tagging, schedule-deletion, cancel-deletion, rotation, or reconfiguration. Object read and write traffic continues normally.
|
||||
2. **Upgrade one node at a time**, waiting for each to report ready before starting the next.
|
||||
3. **Verify no node is left behind** before unfreezing. A single old node is enough to reintroduce blind writes and to strip `baseline_version` on its next lifecycle write.
|
||||
4. **Resume administrative traffic.**
|
||||
5. **Only then perform the first rotation of any key.** Once the whole cluster understands `master_key_version`, rotation is safe; before that it is not.
|
||||
6. **If the KMS configuration was changed at any point**, restart the nodes that did not handle the request so they reload the persisted configuration, and confirm each one reports the intended backend.
|
||||
|
||||
### Do not do these during a mixed-version window
|
||||
|
||||
- **Rotate any key.** This is the hard constraint stated above; a rotation is unrecoverable for objects an old node must read.
|
||||
- **Issue any KV2 lifecycle write to an old node.** Its blind write can clobber a concurrent check-and-set commit and will drop `baseline_version` from the record.
|
||||
- **Create the same key ID from two nodes.** The create path is create-only on upgraded builds, but an old node's blind write does not honor that: the later writer's material wins and every DEK already wrapped with the earlier material becomes permanently unwrappable.
|
||||
- **Assume a disable or schedule-deletion took effect cluster-wide.** Old Transit nodes cache lifecycle state without expiry; confirm per node, or restart the old nodes, before treating a key as no longer in use.
|
||||
- **Reconfigure the KMS backend and consider it done.** The change applies to one node and is persisted; the rest need a restart.
|
||||
- **Delete or prune version records** under `{prefix}/{key_id}/versions/*` for any reason. This is never safe, mixed-version or not; see [Retention and destruction preconditions](#retention-and-destruction-preconditions).
|
||||
|
||||
## Choosing between Vault KV2 and Vault Transit
|
||||
|
||||
Use **Vault Transit** (`VaultTransit`) when key material must be cryptographically isolated from anyone holding storage-level read access: Transit keeps key-encryption keys inside Vault and only ever returns ciphertext, and supports server-side key versioning/rotation.
|
||||
|
||||
@@ -0,0 +1,148 @@
|
||||
# Cryptographic compliance positioning
|
||||
|
||||
This document records where RustFS stands on cryptographic module validation, what may and may not be said about it in external material, and what each possible route to a stronger position would actually cost. It exists so that the question is answered once, from the code, instead of being re-litigated from assumptions about crate names and feature flags.
|
||||
|
||||
For where master key material lives per backend and how rotation retention works, see [KMS backend security properties](kms-backend-security.md).
|
||||
|
||||
## Status: not FIPS 140-3 validated
|
||||
|
||||
**RustFS is not FIPS 140-3 (or 140-2) validated, and no component it links is running as a validated cryptographic module.** There is no CMVP certificate covering RustFS or the libraries it uses in the shipped configuration.
|
||||
|
||||
This is a deliberate position, not an oversight. It is also not a statement about algorithm strength: the algorithms in use are standard, well-reviewed AEADs. Validation is a property of a specific module build, its documented boundary, and a certificate — none of which RustFS has or currently pursues.
|
||||
|
||||
### What the process actually links
|
||||
|
||||
The table below is the audited inventory as of this document's writing. "Validated module" asks only whether the code performing the operation is a FIPS-validated cryptographic module; the answer is uniformly no.
|
||||
|
||||
| Layer | Where | Implementation | Primitives | Validated module |
|
||||
| --- | --- | --- | --- | --- |
|
||||
| TLS (S3 server, internode, outbound clients) | Process-wide default provider installed by `install_default_crypto_provider` in `rustfs/src/startup_runtime_hooks.rs` | `rustls` with the `aws-lc-rs` provider | TLS 1.2/1.3 suites, `prefer-post-quantum` hybrid key exchange | No — this is the ordinary `aws-lc-rs` build, not the `aws-lc-fips-sys`-backed FIPS variant |
|
||||
| Object data path AEAD (SSE) | `crates/kms/src/encryption/ciphers.rs`, `crates/rio/src/encrypt_reader.rs`, `crates/rio-v2/src/encrypt_reader.rs` | RustCrypto `aes-gcm`, `chacha20poly1305` | AES-256-GCM, ChaCha20-Poly1305 | No |
|
||||
| DEK wrapping | `crates/kms/src/encryption/dek.rs` | RustCrypto `aes-gcm` | AES-256-GCM | No |
|
||||
| Local KMS backend master key | `crates/kms/src/backends/local.rs` | RustCrypto `argon2`, `aes-gcm` | Argon2id KDF, AES-256-GCM | No |
|
||||
| Config and IAM blobs at rest | `crates/crypto/src/encdec/` (`rustfs-crypto`) | RustCrypto `pbkdf2`/`argon2`, `aes-gcm`, `chacha20poly1305`, `sha2` | see [the `fips` feature](#the-rustfs-crypto-fips-feature-what-it-actually-does) | No |
|
||||
| JWT signing and verification | `jsonwebtoken` with the `aws_lc_rs` feature (`crates/crypto`, `crates/iam`, `crates/policy`) | AWS-LC through `aws-lc-rs` | Non-FIPS build | No |
|
||||
|
||||
Two consequences follow directly from the table and are worth stating explicitly, because both are commonly assumed the other way:
|
||||
|
||||
- **AWS-LC being present does not imply FIPS.** `aws-lc-rs` has a FIPS variant; the workspace does not enable it. Every `aws-lc-rs` dependency in the workspace is the default, non-FIPS build.
|
||||
- **The data path never touches AWS-LC.** Every byte of object plaintext is encrypted by RustCrypto software implementations. Swapping the TLS provider would not change that; see [route 1](#route-1-adopt-the-aws-lc-rs-fips-variant) for what would.
|
||||
|
||||
## Terminology red lines for external material
|
||||
|
||||
These rules apply to the README, CHANGELOG, release notes, marketing pages, sales decks, RFP responses, and security questionnaires. Claiming validation RustFS does not have is a false statement of fact with regulatory and contractual consequences, not a marketing overreach.
|
||||
|
||||
### Never use
|
||||
|
||||
- "FIPS validated", "FIPS certified", "FIPS 140-2/140-3 compliant", "FIPS compliant"
|
||||
- "FIPS mode", "runs in FIPS mode", "FIPS-enabled"
|
||||
- "NIST certified", "NIST approved", "CMVP certificate", any certificate number
|
||||
- "meets FIPS requirements", "satisfies FIPS", or any phrasing a reader would reasonably read as validation
|
||||
- The internal Cargo feature name `fips` as a product capability. It is a build-time algorithm selector (see below), and surfacing it as a feature name invites exactly the misreading this section exists to prevent.
|
||||
|
||||
### Permitted, with the qualifier attached
|
||||
|
||||
- **"FIPS-preferred algorithms"** — permitted only when accompanied, in the same paragraph or table cell, by an explicit non-validation statement. The defined meaning is: *the default algorithm selection is restricted to algorithms on the FIPS 140-3 approved list, implemented by software that has not been validated as a cryptographic module.*
|
||||
- Naming specific primitives factually ("AES-256-GCM", "ChaCha20-Poly1305", "PBKDF2-HMAC-SHA256") is always fine. Algorithm names carry no validation claim.
|
||||
|
||||
Suggested boilerplate when the topic cannot be avoided:
|
||||
|
||||
> RustFS encrypts object data with AES-256-GCM and supports ChaCha20-Poly1305. These are FIPS-approved algorithms, but the implementations are not FIPS 140-3 validated cryptographic modules and RustFS makes no FIPS validation claim.
|
||||
|
||||
### Guard
|
||||
|
||||
There is currently no FIPS-related wording anywhere in the repository's Markdown; that clean baseline is what a grep anchor test protects. Any future occurrence of the banned strings in shipped documentation should be treated as a defect and either removed or brought under the qualifier rule above.
|
||||
|
||||
## The `rustfs-crypto` `fips` feature: what it actually does
|
||||
|
||||
`crates/crypto/Cargo.toml` declares `default = ["crypto", "fips"]`, so the feature is on in every normal build. Its entire effect is **which algorithm the write path selects**; the implementation is RustCrypto either way.
|
||||
|
||||
| `fips` | Algorithm ID written | KDF | AEAD |
|
||||
| --- | --- | --- | --- |
|
||||
| enabled (default) | `ID::Pbkdf2AESGCM` (`0x02`) | PBKDF2-HMAC-SHA256, 8192 iterations | AES-256-GCM |
|
||||
| disabled | `ID::Argon2idAESGCM` (`0x00`) or `ID::Argon2idChaCHa20Poly1305` (`0x01`), chosen at runtime by CPU AES support | Argon2id (64 MiB, t=1, p=4) | AES-256-GCM or ChaCha20-Poly1305 |
|
||||
|
||||
The selection sites are `crates/crypto/src/encdec/encrypt.rs` and `crates/crypto/src/encdec/stream_io.rs`; the algorithm identifiers and their KDF parameters live in `crates/crypto/src/encdec/id.rs`.
|
||||
|
||||
Three properties matter for anyone reasoning about this feature:
|
||||
|
||||
- **It affects writes only.** The decrypt path in `crates/crypto/src/encdec/id.rs` accepts all three identifiers unconditionally, and every ciphertext carries its identifier byte. Toggling the feature therefore never orphans existing data in either direction.
|
||||
- **It does not select a different implementation.** Both branches call RustCrypto. There is no validated module on either side of the switch, so the feature cannot move RustFS toward or away from validation.
|
||||
- **It is a trade-off, not an upgrade.** The FIPS-preferred branch uses PBKDF2-HMAC-SHA256 at 8192 iterations, a work factor well below current password-hashing guidance, whereas the non-FIPS branch uses memory-hard Argon2id. Against an attacker who has obtained an encrypted config or IAM blob and is attacking the passphrase offline, the default branch is the weaker of the two. Enabling the feature buys approved-algorithm alignment, not more resistance.
|
||||
|
||||
### Rename recommendation
|
||||
|
||||
The name `fips` states a compliance property the feature does not provide, and `rustfs-crypto` is published, so the name is visible to downstream consumers. Recommended direction:
|
||||
|
||||
1. Introduce `fips-preferred-algs` as the real feature name, carrying the current behavior.
|
||||
2. Redefine `fips = ["fips-preferred-algs"]` so existing consumers keep building, and mark it deprecated in the crate documentation with a pointer to this document.
|
||||
3. Drop the `fips` alias after one release cycle.
|
||||
4. While renaming, raise the PBKDF2 iteration count or document the trade-off above at the feature definition, so the choice is explicit rather than inherited.
|
||||
|
||||
This is a naming and documentation change only; no ciphertext format changes, because the identifier bytes stay as they are.
|
||||
|
||||
## Routes to a stronger position, and what each costs
|
||||
|
||||
### Route 1: adopt the `aws-lc-rs` FIPS variant
|
||||
|
||||
Switch the whole process to `aws-lc-rs`'s FIPS build (backed by `aws-lc-fips-sys`) so cryptographic operations run inside a validated module boundary.
|
||||
|
||||
**Scope.** The TLS provider swap is the small part — one feature flag plus the provider install sites. The substantial work is the data path: every AEAD call in `crates/kms/src/encryption/ciphers.rs`, `crates/kms/src/encryption/dek.rs`, `crates/rio/src/encrypt_reader.rs`, `crates/rio-v2/src/encrypt_reader.rs`, `crates/kms/src/backends/local.rs`, and `crates/crypto/src/encdec/` would have to be re-implemented against `aws-lc-rs` primitives. Anything the validated module does not expose has to be dropped or moved out of the boundary: Argon2id has no FIPS status, so the Local backend's KDF and the non-FIPS branch of `rustfs-crypto` would need a compatibility story (read-only support for existing records, PBKDF2 for new ones), and ChaCha20-Poly1305 would become non-approved for new writes.
|
||||
|
||||
**Build and platform cost.** `aws-lc-fips-sys` builds a pinned, validated source release and needs CMake, a C toolchain, and Go at build time; it supports a narrower target set than the ordinary crate. The platform matrix cost of plain AWS-LC is already documented and non-hypothetical: rustfs/backlog#883 records that the static musl release build compiles AWS-LC's `getentropy` entropy backend, which aborts on Linux kernels older than 3.17 (the Synology class of device), and that upstream considers this by design with no plan to fix it. The FIPS variant constrains the buildable matrix strictly harder than that, and pins upgrades to whatever the certified source revision allows.
|
||||
|
||||
**What it would and would not buy.** Linking the validated module makes the accurate claim "cryptographic operations are performed by a FIPS 140-3 validated module", not "RustFS is FIPS validated". A product-level claim additionally requires a documented module boundary, approved-mode enforcement, power-on self-tests, key zeroization, and entropy-source documentation, plus the operational procedures to keep them true across releases.
|
||||
|
||||
**Verdict.** Heavy, and it re-opens a platform-support question that is already an open problem. Justified only by a concrete customer or regulatory commitment that names FIPS as a requirement.
|
||||
|
||||
### Route 2: let an externally validated KMS carry key operations
|
||||
|
||||
Keep RustFS as-is and place key management inside someone else's validated boundary: the Vault Transit backend against a Vault deployment whose seal/HSM is validated, or an equivalent managed KMS.
|
||||
|
||||
**Scope.** Mostly already built. The Transit backend (`VaultTransit`) never lets key-encryption key material leave Vault; RustFS only ever holds Transit ciphertext. What remains is configuration guidance, a supported-deployment statement, and the operational documentation that says which parts of the system are covered.
|
||||
|
||||
**What it buys.** Master key generation, wrapping, unwrapping, and rotation happen inside the external module. That is a real, defensible partial answer to "where do keys live and who validated that": it covers the key operations, which is often the part an auditor actually asks about.
|
||||
|
||||
**What it does not buy.** The object data path is untouched. DEKs are used for bulk AEAD by RustCrypto inside the RustFS process, and TLS still runs the non-FIPS AWS-LC build. The honest formulation is "key management operations are performed by an externally validated module; the object data path is not validated".
|
||||
|
||||
**Verdict.** The nearest partial step, with no code rewrite and no platform-matrix risk. This is the route to point customers at when the requirement is about key custody rather than about a certificate covering the storage layer.
|
||||
|
||||
### Route 3: make no validation claim (current default)
|
||||
|
||||
Document the position, hold the terminology line, and revisit only when a requirement with a name attached shows up.
|
||||
|
||||
**Cost.** This document plus the grep guard. Nothing else.
|
||||
|
||||
**Verdict.** The current decision. FIPS 140-3 validation is explicitly not a roadmap target, and adjacent items (PKCS#11, KMIP, BYOK, signing keys) are deferred for lack of demand and because HSM-dependent paths cannot be exercised in CI.
|
||||
|
||||
## Algorithm disablement and migration policy
|
||||
|
||||
Retiring an algorithm from a storage system is not a code change; it is a data migration with a code change at each end. This section fixes the sequence so that no future deprecation removes a decrypt path while data still depends on it.
|
||||
|
||||
### Every persisted artifact is self-describing
|
||||
|
||||
The precondition for safe migration already holds: nothing relies on a global "current algorithm" setting to be decodable.
|
||||
|
||||
- `rustfs-crypto` blobs carry the `ID` byte (`crates/crypto/src/encdec/id.rs`) immediately after the salt.
|
||||
- KMS ciphers are selected from the recorded `EncryptionAlgorithm` (`crates/kms/src/types.rs`).
|
||||
- DEK envelopes record which master key version wrapped them in `DataKeyEnvelope::master_key_version` (`crates/kms/src/encryption/dek.rs`).
|
||||
|
||||
So for any stored object it is decidable, from the object alone, which algorithm and which key version it needs.
|
||||
|
||||
### Deprecation classes
|
||||
|
||||
Retirement moves an algorithm through these states, never skipping one:
|
||||
|
||||
1. **Write-disabled, read-supported.** New writes select a replacement; existing data decrypts unchanged. This is the only step that is cheap and reversible.
|
||||
2. **Read-deprecated.** Reads still work but are counted and warned on, so the remaining population is measurable.
|
||||
3. **Read-removed.** The decrypt path is deleted. Permitted only once the remaining population is provably zero.
|
||||
|
||||
### Sequencing rules
|
||||
|
||||
- Never advance to read-removed on the strength of an argument that data "should have been" migrated. Removal requires evidence that nothing references the algorithm, not an elapsed-time policy.
|
||||
- A change to default algorithm selection is a compatibility event: it changes what new nodes write, which matters in a mixed-version cluster. Record it in the release notes and in the relevant crate's feature documentation, and check it against the [mixed-version constraints](kms-backend-security.md#mixed-version-clusters-during-a-rolling-upgrade).
|
||||
- Roll out write-disablement before the corresponding read change, and let the cluster fully converge in between. A build that cannot read what a peer is still writing is the failure mode to avoid.
|
||||
|
||||
### Known gap
|
||||
|
||||
Step 3 is currently unreachable for object data. There is no object rewrap or re-encryption capability, so there is no supported way to migrate already-written objects off an algorithm or off a master key version — the same gap that forces the rotation retention rule in [KMS backend security properties](kms-backend-security.md#retention-and-destruction-preconditions). Until a rewrap capability exists, treat every algorithm that has ever been written as permanently read-required, and confine deprecation to step 1.
|
||||
@@ -0,0 +1,130 @@
|
||||
# KMS observability runbook
|
||||
|
||||
This runbook covers the KMS backend operation metrics, the Grafana dashboard that visualizes them, and the response procedure for each Prometheus alert shipped in `.docker/observability/prometheus-rules/rustfs-kms-alerts.yml`. It is the `runbook_url` target for those alerts. For what each KMS backend protects and how Vault authentication behaves, see the [KMS backend security properties](kms-backend-security.md) and the [Vault KMS authentication runbook](vault-kms-authentication.md).
|
||||
|
||||
## Metric reference
|
||||
|
||||
All four metrics are emitted at the single operation-policy choke point (`crates/kms/src/policy.rs`) that every instrumented KMS backend call flows through. Label values are exclusively static enum strings — key identifiers, key material, ciphertext, paths, and tokens never appear in metric labels, and any change that would add such a label is a regression.
|
||||
|
||||
| Metric | Type | Labels | Meaning |
|
||||
| --- | --- | --- | --- |
|
||||
| `rustfs_kms_backend_operations_total` | counter | `operation`, `op_class`, `outcome` | Operations executed under the operation policy, counted once per terminal outcome |
|
||||
| `rustfs_kms_backend_attempt_failures_total` | counter | `operation`, `error_class` | Individual failed attempts, including attempts the retry policy later absorbed |
|
||||
| `rustfs_kms_backend_operation_duration_seconds` | histogram | `operation`, `outcome` | Wall-clock duration of a whole operation, including retries and backoff sleeps |
|
||||
| `rustfs_kms_backend_operation_attempts` | histogram | `operation`, `outcome` | Number of attempts one operation used before completing |
|
||||
|
||||
Label values:
|
||||
|
||||
- `outcome`: `success`, `fatal` (a non-retryable failure ended the operation on first observation), `budget_exhausted` (the attempt budget ran out on retryable failures), `deadline_exceeded` (the operation deadline ran out before another attempt could complete), `cancelled` (shutdown or caller cancellation).
|
||||
- `op_class`: `read_idempotent` (safe to retry), `mutating_non_idempotent` (never replayed — a retryable failure terminates after a single attempt because the server may have processed the request), `auth` (login and token renewal).
|
||||
- `error_class`: `retryable_conn` (connection-level failure: dial, TLS, broken connection), `retryable_status` (retryable backend status, e.g. Vault 5xx or a sealed Vault's 503), `attempt_timeout` (the per-attempt timeout cut the attempt off; retried like a connection failure because the server may still have processed the request), `fatal` (non-retryable: authentication, permissions, malformed request, missing key or version).
|
||||
- `operation`: static per-call-site names, e.g. `vault_kv2_read_key_version`, `vault_kv2_cas_write_key`, `vault_transit_encrypt`, `vault_transit_decrypt`, `vault_login`, `vault_token_renew`.
|
||||
|
||||
Export path: the `metrics` facade feeds the OTel recorder in `crates/obs`, which exports over OTLP to the collector scraped by Prometheus. Histograms therefore appear in Prometheus as `_bucket`/`_sum`/`_count` series. These metrics do not carry the RustFS `server` label used by the node observability dashboard — distinguish nodes through your scrape topology (`job`/`instance` or promoted OTel resource attributes such as `service_instance_id`).
|
||||
|
||||
Instrumentation boundary: the Local and Static backends do not flow through the choke point and emit no operation metrics; bringing them under the same instrumentation is tracked separately (rustfs/backlog#1569). Absence of KMS series on a cluster using those backends is expected, not an outage.
|
||||
|
||||
## Dashboard
|
||||
|
||||
Import `deploy/observability/grafana/rustfs-kms-observability.json` into Grafana and select a Prometheus data source that scrapes RustFS metrics. The dashboard has two variables: `datasource` (Prometheus data source) and `operation` (multi-select over the `operation` label). In the docker-compose observability stack (`.docker/observability/`), dashboards are provisioned from a directory (`grafana/provisioning/dashboards/dashboard.yml` points at `/etc/grafana/dashboards`), so no per-file registration is needed there.
|
||||
|
||||
The dashboard's "Planned Panels (TODO)" text panel lists metric families designed under rustfs/backlog#1584 but not yet landed (key-cache effectiveness, lifecycle gauges, token TTL, synthetic probe). Do not add panels or alerts for those names until the emitting code is merged; keep that panel and the [Coverage gaps](#coverage-gaps-and-planned-metrics) section below in sync as they land.
|
||||
|
||||
## Alert rules
|
||||
|
||||
The rules live in `.docker/observability/prometheus-rules/rustfs-kms-alerts.yml`. The docker-compose Prometheus loads `/etc/prometheus/rules/*.yml`, so the `.yml` extension is load-bearing. Validate edits with `promtool check rules rustfs-kms-alerts.yml`.
|
||||
|
||||
Every threshold in that file is a conservative default chosen without a production baseline; see [Threshold calibration](#threshold-calibration) before treating a firing alert as an SLO breach or a quiet one as health.
|
||||
|
||||
## Alert response procedures
|
||||
|
||||
### KmsBackendFatalErrors
|
||||
|
||||
Meaning: attempts are failing with `error_class="fatal"` — failures the policy never retries. Each one is a KMS backend call that failed permanently (authentication, permissions, malformed request, or a missing key/version), so callers are seeing errors right now. This is the highest-signal KMS alert: fatal failures do not appear as background noise in a healthy system.
|
||||
|
||||
Investigation:
|
||||
|
||||
1. Break the rate down by operation: `sum by (operation) (rate(rustfs_kms_backend_attempt_failures_total{error_class="fatal"}[5m]))`.
|
||||
2. If the failing operations are `vault_login` or `vault_token_renew` (`op_class="auth"`), the Vault credentials are invalid or expired. Follow the [Vault KMS authentication runbook](vault-kms-authentication.md) — note that credential refresh is fail-closed, so a broken credential eventually takes down all Vault-backed operations, not just auth. Look for the `Vault token renewal failed; falling back to a fresh login` and `Vault credential refresh failed; retrying until the credentials recover` warnings in the RustFS logs.
|
||||
3. If the failing operations are `vault_kv2_*` or `vault_transit_*`, check for Vault permission denials: compare the token's policy against the minimal policy in [KMS backend security properties](kms-backend-security.md) (a policy that drifted or was re-scoped produces 403s that classify as fatal), and check the Vault audit log for the corresponding denied requests.
|
||||
4. A fatal `KeyVersionNotFound` on decrypt-path operations means a DEK envelope references a key version whose record is missing. Decryption deliberately fails closed with no fallback — see the rotation retention preconditions in [KMS backend security properties](kms-backend-security.md) and verify nobody destroyed version records under the key subtree.
|
||||
5. Confirm blast radius with the outcome view: `sum by (operation) (rate(rustfs_kms_backend_operations_total{outcome="fatal"}[5m]))`.
|
||||
|
||||
Related signals: the "Attempt Failure Rate by Error Class" and "Backend Operation Rate by Outcome" dashboard panels; Vault server audit and server logs; S3-level 5xx on encrypted buckets.
|
||||
|
||||
### KmsBackendHighErrorRate
|
||||
|
||||
Meaning: more than 5% of KMS operations are terminating without success (`fatal`, `budget_exhausted`, or `deadline_exceeded`; `cancelled` is excluded because shutdown windows legitimately produce it). A traffic guard suppresses the alert below ~0.02 ops/s so a single failure on a near-idle cluster does not page.
|
||||
|
||||
Investigation:
|
||||
|
||||
1. Break the failures down by outcome: `sum by (outcome) (rate(rustfs_kms_backend_operations_total{outcome!~"success|cancelled"}[5m]))`.
|
||||
2. If `fatal` dominates, follow [KmsBackendFatalErrors](#kmsbackendfatalerrors).
|
||||
3. If `budget_exhausted` or `deadline_exceeded` dominates, follow [KmsBackendRetryBudgetExhausted](#kmsbackendretrybudgetexhausted) — the backend is unavailable or too slow for longer than the retry policy can bridge.
|
||||
4. Correlate with client impact: encrypted-object PUT/GET failures and S3 error rates on buckets with encryption configured.
|
||||
|
||||
Related signals: the "Non-Success Outcome Ratio" dashboard panel; the KMS-related warnings listed under the other alerts in this runbook.
|
||||
|
||||
### KmsBackendP99LatencyHigh
|
||||
|
||||
Meaning: the p99 wall-clock duration of KMS operations is sustained above 2s. The histogram includes retries and backoff sleeps, so a high p99 with a healthy p50 usually means a slow retry tail (a subset of calls failing and being retried), not a uniform slowdown.
|
||||
|
||||
Investigation:
|
||||
|
||||
1. Compare p50 and p99 on the "Operation Duration p50 / p99" panel. Flat p50 with elevated p99 points at retries; both elevated points at the backend or the network path being uniformly slow.
|
||||
2. Split by operation with `histogram_quantile(0.99, sum by (le, operation) (rate(rustfs_kms_backend_operation_duration_seconds_bucket[5m])))` to see whether one backend call or all of them regressed.
|
||||
3. Check the attempts histogram: an average meaningfully above 1 confirms the latency is retry-driven; follow [KmsBackendAttemptFailureSpike](#kmsbackendattemptfailurespike) for the failure classes.
|
||||
4. If latency is not retry-driven, check the network path to Vault (TLS handshakes, DNS, proxies) and Vault's own telemetry (storage backend latency, load).
|
||||
5. Remember that this latency sits inside S3 request latency for encrypted objects: sustained p99 near the operation deadline will start converting into `deadline_exceeded` outcomes.
|
||||
|
||||
Related signals: the "Operation Duration p99 by Operation" and "Operation Attempts Distribution" panels; `KMS backend attempt failed with a retryable error; backing off before retry` warnings (fields: `operation`, `attempt`, `error_class`, `backoff`).
|
||||
|
||||
### KmsBackendAttemptFailureSpike
|
||||
|
||||
Meaning: individual attempts are failing at a sustained rate across all error classes. The retry policy may still be absorbing these — operations can keep succeeding while this alert fires — but the system is burning retry budget and running degraded, and a small further degradation will surface to callers.
|
||||
|
||||
Investigation:
|
||||
|
||||
1. Break the rate down by class: `sum by (error_class) (rate(rustfs_kms_backend_attempt_failures_total[5m]))`.
|
||||
2. `retryable_conn`: network-level failures — check connectivity, TLS, DNS, and whether Vault is down or restarting.
|
||||
3. `retryable_status`: the backend answered with a retryable error — check Vault health and seal status (a sealed Vault returns 503, which lands here), and Vault-side rate limiting.
|
||||
4. `attempt_timeout`: attempts are being cut off by the per-attempt timeout — either the backend is slow (correlate with [KmsBackendP99LatencyHigh](#kmsbackendp99latencyhigh)) or the configured attempt timeout is too tight for the deployment's network path.
|
||||
5. `fatal`: follow [KmsBackendFatalErrors](#kmsbackendfatalerrors).
|
||||
6. Grep RustFS logs for `KMS backend attempt failed with a retryable error; backing off before retry` — the structured fields (`operation`, `attempt`, `error_class`, `backoff`) identify which call sites are cycling.
|
||||
|
||||
Related signals: the "Attempt Failure Rate by Error Class" panel; the attempts histogram average rising above 1.
|
||||
|
||||
### KmsBackendRetryBudgetExhausted
|
||||
|
||||
Meaning: operations are terminating as `budget_exhausted` or `deadline_exceeded` — every individual failure was retryable, but the backend stayed unhealthy for longer than the retry policy could bridge, so callers received hard failures.
|
||||
|
||||
Investigation:
|
||||
|
||||
1. Identify the failing operations: `sum by (operation) (rate(rustfs_kms_backend_operations_total{outcome=~"budget_exhausted|deadline_exceeded"}[5m]))`.
|
||||
2. Establish how long the underlying failure has persisted from the attempt-failure rate history; follow [KmsBackendAttemptFailureSpike](#kmsbackendattemptfailurespike) for the class-specific diagnosis.
|
||||
3. Note the by-design case: `mutating_non_idempotent` operations (e.g. `vault_kv2_cas_write_key`, `vault_transit_create_key`) are never replayed, so a single retryable failure terminates them as `budget_exhausted` after one attempt. A spike confined to mutating operations means write-path failures, not an exhausted retry loop.
|
||||
4. `deadline_exceeded` clustering with duration p99 near the operation deadline means the budget is being spent on slow attempts rather than fast failures — treat as a latency problem first.
|
||||
5. Confirm client impact and, if the backend outage is confirmed external (Vault down), coordinate recovery there; RustFS will resume without intervention once the backend recovers.
|
||||
|
||||
Related signals: the "Backend Operation Rate by Outcome" panel; retry-backoff warnings in RustFS logs; Vault availability monitoring.
|
||||
|
||||
## Threshold calibration
|
||||
|
||||
Every numeric threshold in `rustfs-kms-alerts.yml` (5% error ratio, 2s p99, 0.5/s attempt failures, 0.05/s budget exhaustion) is a conservative default chosen without a production baseline, biased toward not paging on healthy-but-busy systems. Before relying on these alerts for paging: run the workload in staging for at least a week, record the steady-state values of the expressions above, then tighten thresholds to sit clearly above observed peaks. Once a stable baseline exists, consider converting `KmsBackendAttemptFailureSpike` to a baseline-relative form (`offset 1d` ratio, see `.docker/observability/prometheus-rules/rustfs-get-optimization-alerts.yaml` for the pattern). Formal SLO targets for KMS operations are deliberately out of scope until that baseline exists (rustfs/backlog#1584).
|
||||
|
||||
## Coverage gaps and planned metrics
|
||||
|
||||
The following families are designed under rustfs/backlog#1584 but land in separate changes; they are intentionally absent from the dashboard and alert rules, and referencing their names before the emitting code merges would produce permanently-empty panels and never-firing alerts:
|
||||
|
||||
- Key-cache effectiveness: hit/miss counters, entry gauge, eviction counter (replacing the former hardcoded-zero miss statistic).
|
||||
- Key lifecycle: pending-deletion/tombstone gauges from the deletion-worker sweep and an aggregate rotation-age gauge.
|
||||
- Vault credentials: token TTL remaining and fail-closed state gauges (today only the log warnings quoted above exist).
|
||||
- Synthetic backend probe: probe outcome and failure-class metrics feeding the readiness cache.
|
||||
|
||||
When one of these lands, add its panels, extend the alert rules, replace the corresponding TODO bullet in the dashboard's "Planned Panels" text panel, and update this section.
|
||||
|
||||
## Related documents
|
||||
|
||||
- [KMS backend security properties](kms-backend-security.md) — backend trust boundaries, minimal Vault policies, rotation retention preconditions.
|
||||
- [Vault KMS authentication runbook](vault-kms-authentication.md) — credential sources, refresh behavior, and the fail-closed window.
|
||||
- `deploy/observability/README.md` — dashboard import notes for all RustFS dashboards.
|
||||
@@ -15,13 +15,15 @@
|
||||
use crate::admin::handlers::supervise_admin_mutation;
|
||||
use crate::admin::handlers::target_descriptor::AdminTargetSpec;
|
||||
use crate::admin::runtime_sources::{AppContext, current_app_context, current_object_store_handle_for_context};
|
||||
use crate::admin::service::config::with_runtime_config_reload_lock;
|
||||
use crate::admin::storage_api::config::{
|
||||
read_admin_config_without_migrate, read_admin_server_config_snapshot, save_admin_server_config_snapshot,
|
||||
read_admin_config_without_migrate, read_admin_server_config_snapshot, read_existing_admin_server_config_no_lock,
|
||||
save_admin_server_config_snapshot, with_admin_server_config_read_lock,
|
||||
};
|
||||
use rustfs_audit::{audit_system, start_audit_system as start_global_audit_system, system::AuditSystemState};
|
||||
use rustfs_config::DEFAULT_DELIMITER;
|
||||
use rustfs_config::server_config::Config;
|
||||
use s3s::{S3Result, s3_error};
|
||||
use s3s::{S3Error, S3Result, s3_error};
|
||||
use tracing::warn;
|
||||
|
||||
pub(crate) async fn load_server_config_from_store_for_context(context: Option<&AppContext>) -> S3Result<Config> {
|
||||
@@ -48,6 +50,14 @@ fn has_any_audit_targets(specs: &[AdminTargetSpec], config: &Config) -> bool {
|
||||
})
|
||||
}
|
||||
|
||||
fn audit_config_convergence_error(persisted: bool, error: impl std::fmt::Display) -> S3Error {
|
||||
if persisted {
|
||||
s3_error!(InternalError, "audit config persisted but runtime convergence failed: {}", error)
|
||||
} else {
|
||||
s3_error!(InternalError, "audit config unchanged but runtime convergence failed: {}", error)
|
||||
}
|
||||
}
|
||||
|
||||
pub(crate) async fn apply_audit_runtime_config(specs: &[AdminTargetSpec], config: Config) -> S3Result<()> {
|
||||
let has_targets = has_any_audit_targets(specs, &config);
|
||||
|
||||
@@ -107,15 +117,23 @@ where
|
||||
return Ok(());
|
||||
}
|
||||
|
||||
save_admin_server_config_snapshot(store, &config, &snapshot)
|
||||
let persisted = save_admin_server_config_snapshot(store.clone(), &config, &snapshot)
|
||||
.await
|
||||
.map(|_| ())
|
||||
.map_err(|e| s3_error!(InternalError, "failed to save audit config: {}", e))?;
|
||||
.map_err(|e| s3_error!(InternalError, "failed to save audit config: {}", e))?
|
||||
.persisted();
|
||||
drop(snapshot);
|
||||
|
||||
// Keep persistence and runtime publication in one detached, serialized
|
||||
// mutation. Otherwise a cancelled caller or two concurrent updates can
|
||||
// leave the persisted config and active audit generation disagreeing.
|
||||
apply_audit_runtime_config(&specs, config).await
|
||||
let read_store = store.clone();
|
||||
with_runtime_config_reload_lock(async move {
|
||||
let latest = with_admin_server_config_read_lock(store, move || read_existing_admin_server_config_no_lock(read_store))
|
||||
.await
|
||||
.map_err(|e| s3_error!(InternalError, "failed to lock server config for audit reload: {}", e))?
|
||||
.map_err(|e| s3_error!(InternalError, "failed to read latest server config for audit reload: {}", e))?;
|
||||
|
||||
apply_audit_runtime_config(&specs, latest).await
|
||||
})
|
||||
.await
|
||||
.map_err(|e| audit_config_convergence_error(persisted, e))
|
||||
})
|
||||
.await
|
||||
}
|
||||
@@ -164,3 +182,154 @@ pub(crate) async fn remove_audit_target_config(specs: &[AdminTargetSpec], subsys
|
||||
})
|
||||
.await
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
mod tests {
|
||||
use super::*;
|
||||
use crate::admin::handlers::target_descriptor::admin_target_spec_from_builtin;
|
||||
use crate::admin::runtime_sources::{IamInterface, KmsInterface};
|
||||
use crate::admin::storage_api::config::save_admin_server_config;
|
||||
use rustfs_config::audit::AUDIT_WEBHOOK_SUB_SYS;
|
||||
use rustfs_config::server_config::KVS;
|
||||
use rustfs_config::{ENABLE_KEY, EnableState, SCANNER_CYCLE, SCANNER_SUB_SYS, WEBHOOK_ENDPOINT, WEBHOOK_QUEUE_DIR};
|
||||
use rustfs_iam::{store::object::ObjectStore, sys::IamSys};
|
||||
use rustfs_kms::KmsServiceManager;
|
||||
use rustfs_targets::catalog::builtin::builtin_audit_target_admin_descriptors;
|
||||
use std::sync::Arc;
|
||||
use std::time::Duration;
|
||||
use tempfile::TempDir;
|
||||
|
||||
struct TestIam;
|
||||
|
||||
impl IamInterface for TestIam {
|
||||
fn handle(&self) -> Arc<IamSys<ObjectStore>> {
|
||||
unreachable!("audit config tests do not use IAM")
|
||||
}
|
||||
|
||||
fn is_ready(&self) -> bool {
|
||||
false
|
||||
}
|
||||
}
|
||||
|
||||
struct TestKms;
|
||||
|
||||
impl KmsInterface for TestKms {
|
||||
fn handle(&self) -> Arc<KmsServiceManager> {
|
||||
Arc::new(KmsServiceManager::new())
|
||||
}
|
||||
}
|
||||
|
||||
fn audit_specs() -> Vec<AdminTargetSpec> {
|
||||
builtin_audit_target_admin_descriptors()
|
||||
.into_iter()
|
||||
.map(|descriptor| admin_target_spec_from_builtin(&descriptor))
|
||||
.collect()
|
||||
}
|
||||
|
||||
async fn wait_for_persisted_target(store: Arc<crate::admin::storage_api::runtime::ECStore>, subsystem: &str, target: &str) {
|
||||
tokio::time::timeout(Duration::from_secs(30), async {
|
||||
let mut poll = tokio::time::interval(Duration::from_millis(10));
|
||||
loop {
|
||||
poll.tick().await;
|
||||
let config = read_admin_config_without_migrate(store.clone())
|
||||
.await
|
||||
.expect("read persisted server config");
|
||||
if config.0.get(subsystem).is_some_and(|targets| targets.contains_key(target)) {
|
||||
return;
|
||||
}
|
||||
}
|
||||
})
|
||||
.await
|
||||
.expect("config mutation should become durable");
|
||||
}
|
||||
|
||||
#[tokio::test(flavor = "multi_thread", worker_threads = 4)]
|
||||
#[serial_test::serial]
|
||||
async fn audit_reload_reads_latest_durable_config_after_releasing_write_snapshot() {
|
||||
let temp_dir = TempDir::new().expect("audit config temp dir");
|
||||
let env = rustfs_test_utils::TestECStoreEnv::builder()
|
||||
.base_dir(temp_dir.path())
|
||||
.disk_count(1)
|
||||
.init_bucket_metadata(false)
|
||||
.build()
|
||||
.await;
|
||||
save_admin_server_config(env.ecstore.clone(), &Config::new())
|
||||
.await
|
||||
.expect("persist baseline server config");
|
||||
let context = Arc::new(AppContext::new(env.ecstore.clone(), Arc::new(TestIam), Arc::new(TestKms)));
|
||||
|
||||
let (locked_tx, locked_rx) = tokio::sync::oneshot::channel();
|
||||
let (release_tx, release_rx) = tokio::sync::oneshot::channel();
|
||||
let blocker = tokio::spawn(async move {
|
||||
with_runtime_config_reload_lock(async move {
|
||||
locked_tx.send(()).expect("signal runtime reload lock acquisition");
|
||||
release_rx.await.expect("release runtime reload lock");
|
||||
Ok(())
|
||||
})
|
||||
.await
|
||||
.expect("runtime reload lock blocker");
|
||||
});
|
||||
locked_rx.await.expect("runtime reload lock should be held");
|
||||
|
||||
let older_context = context.clone();
|
||||
let older = tokio::spawn(async move {
|
||||
let specs = audit_specs();
|
||||
update_audit_config_and_reload_for_context(Some(older_context.as_ref()), &specs, |config| {
|
||||
config
|
||||
.0
|
||||
.entry(SCANNER_SUB_SYS.to_string())
|
||||
.or_default()
|
||||
.entry(DEFAULT_DELIMITER.to_string())
|
||||
.or_insert_with(KVS::new)
|
||||
.insert(SCANNER_CYCLE.to_string(), "15s".to_string());
|
||||
true
|
||||
})
|
||||
.await
|
||||
});
|
||||
wait_for_persisted_target(env.ecstore.clone(), SCANNER_SUB_SYS, DEFAULT_DELIMITER).await;
|
||||
assert!(!older.is_finished(), "audit runtime publication must wait for the shared reload lock");
|
||||
|
||||
let snapshot = read_admin_server_config_snapshot(env.ecstore.clone())
|
||||
.await
|
||||
.expect("read newer server config snapshot");
|
||||
let mut latest = snapshot.config.clone();
|
||||
let mut latest_target = KVS::new();
|
||||
latest_target.insert(ENABLE_KEY.to_string(), EnableState::On.to_string());
|
||||
latest_target.insert(WEBHOOK_ENDPOINT.to_string(), "https://audit.invalid/hook".to_string());
|
||||
latest_target.insert(
|
||||
WEBHOOK_QUEUE_DIR.to_string(),
|
||||
temp_dir.path().join("audit-queue").to_string_lossy().into_owned(),
|
||||
);
|
||||
latest
|
||||
.0
|
||||
.entry(AUDIT_WEBHOOK_SUB_SYS.to_string())
|
||||
.or_default()
|
||||
.insert("latest".to_string(), latest_target);
|
||||
save_admin_server_config_snapshot(env.ecstore.clone(), &latest, &snapshot)
|
||||
.await
|
||||
.expect("persist newer audit config");
|
||||
drop(snapshot);
|
||||
|
||||
release_tx.send(()).expect("release runtime reload blocker");
|
||||
blocker.await.expect("runtime reload blocker task");
|
||||
tokio::time::timeout(Duration::from_secs(30), older)
|
||||
.await
|
||||
.expect("older audit update should complete")
|
||||
.expect("older audit update task should not panic")
|
||||
.expect("older audit update should converge from the latest durable config");
|
||||
|
||||
let system = audit_system().expect("latest durable audit target should start the audit system");
|
||||
let targets = system.list_targets().await;
|
||||
assert!(targets.iter().any(|target| target.contains("latest")), "active targets: {targets:?}");
|
||||
system.close().await.expect("audit system should stop after the test");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn audit_convergence_error_reports_durable_write_state() {
|
||||
for (persisted, expected) in [(true, "audit config persisted"), (false, "audit config unchanged")] {
|
||||
let error = audit_config_convergence_error(persisted, "injected failure");
|
||||
assert!(error.to_string().contains(expected), "unexpected convergence error: {error}");
|
||||
assert!(error.to_string().contains("injected failure"));
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
@@ -17,18 +17,17 @@ use crate::admin::handlers::supervise_admin_mutation;
|
||||
use crate::admin::router::{AdminOperation, Operation, S3Router};
|
||||
use crate::admin::runtime_sources::{
|
||||
current_action_credentials, current_app_context, current_object_store_handle_for_context, current_server_config_for_context,
|
||||
publish_server_config,
|
||||
};
|
||||
use crate::admin::service::config::{
|
||||
CONFIG_WORKER_RELOAD_FAILURE_STATE, EVENT_CONFIG_WORKER_RELOAD_FAILED, FULL_CONFIG_WORKER_SUBSYSTEMS, LOG_COMPONENT_ADMIN,
|
||||
LOG_SUBSYSTEM_CONFIG, PreparedRuntimeConfig, is_dynamic_config_subsystem, preflight_dynamic_config_reload,
|
||||
prepare_server_config, reload_dynamic_config_runtime_state, reload_runtime_config_snapshot, signal_config_snapshot_reload,
|
||||
signal_config_snapshot_reload_checked, signal_dynamic_config_reload_checked,
|
||||
FULL_CONFIG_WORKER_SUBSYSTEMS, is_dynamic_config_subsystem, preflight_dynamic_config_reload, prepare_server_config,
|
||||
publish_latest_runtime_config_snapshot, reload_dynamic_config_runtime_state, reload_runtime_config_snapshot,
|
||||
signal_config_snapshot_reload, signal_config_snapshot_reload_checked, signal_dynamic_config_reload_checked,
|
||||
};
|
||||
use crate::admin::storage_api::config::storageclass::{INLINE_BLOCK_ENV, OPTIMIZE_ENV, RRS_ENV, STANDARD_ENV};
|
||||
use crate::admin::storage_api::config::{
|
||||
AdminServerConfigSnapshot, RUSTFS_META_BUCKET, STORAGE_CLASS_SUB_SYS, delete_admin_config, read_admin_config,
|
||||
read_admin_config_without_migrate, read_admin_server_config_snapshot, save_admin_config, save_admin_server_config_snapshot,
|
||||
AdminServerConfigSaveResult, AdminServerConfigSnapshot, RUSTFS_META_BUCKET, STORAGE_CLASS_SUB_SYS, delete_admin_config,
|
||||
read_admin_config, read_admin_config_without_migrate, read_admin_server_config_snapshot, save_admin_config,
|
||||
save_admin_server_config_snapshot,
|
||||
};
|
||||
use crate::admin::storage_api::contract::list::ListOperations as _;
|
||||
use crate::admin::utils::{encode_compatible_admin_payload, is_compat_admin_request, read_compatible_admin_body};
|
||||
@@ -766,7 +765,10 @@ async fn load_server_config_snapshot_from_store() -> S3Result<AdminServerConfigS
|
||||
.map_err(Into::into)
|
||||
}
|
||||
|
||||
async fn save_server_config_to_store(config: &ServerConfig, snapshot: &AdminServerConfigSnapshot) -> S3Result<bool> {
|
||||
async fn save_server_config_to_store(
|
||||
config: &ServerConfig,
|
||||
snapshot: &AdminServerConfigSnapshot,
|
||||
) -> S3Result<AdminServerConfigSaveResult> {
|
||||
let store = object_store()?;
|
||||
save_admin_server_config_snapshot(store, config, snapshot)
|
||||
.await
|
||||
@@ -987,8 +989,8 @@ fn decode_config_history_snapshot(data: &[u8]) -> S3Result<ServerConfig> {
|
||||
Ok(config)
|
||||
}
|
||||
|
||||
fn validate_restore_rollback_generation(current: &ServerConfig, restored: &ServerConfig) -> S3Result<()> {
|
||||
if current == restored {
|
||||
fn validate_restore_rollback_generation(current: Option<Uuid>, committed: Option<Uuid>) -> S3Result<()> {
|
||||
if current.is_some() && current == committed {
|
||||
Ok(())
|
||||
} else {
|
||||
Err(s3_error!(
|
||||
@@ -1004,12 +1006,12 @@ async fn save_server_config_history_snapshot(config: &ServerConfig) -> S3Result<
|
||||
save_server_config_history(&sealed).await
|
||||
}
|
||||
|
||||
async fn cleanup_failed_config_history_snapshot(restore_id: &str) {
|
||||
async fn cleanup_config_history_snapshot(restore_id: &str) {
|
||||
if let Err(err) = delete_server_config_history(restore_id).await {
|
||||
warn!(
|
||||
restore_id,
|
||||
error = %err,
|
||||
"Failed to remove config history snapshot for an uncommitted mutation"
|
||||
"Failed to remove config history snapshot"
|
||||
);
|
||||
}
|
||||
}
|
||||
@@ -1843,23 +1845,6 @@ async fn reject_lost_config_transaction<T>(snapshot: AdminServerConfigSnapshot,
|
||||
))
|
||||
}
|
||||
|
||||
fn publish_prepared_config_snapshots(config: ServerConfig, prepared: PreparedRuntimeConfig) -> S3Result<()> {
|
||||
prepared.publish_storage_class()?;
|
||||
publish_server_config(config);
|
||||
Ok(())
|
||||
}
|
||||
|
||||
/// Re-apply local mutable worker families after a full-config replacement.
|
||||
/// Peers receive one full-snapshot signal after this returns; signaling each
|
||||
/// family here as well would recreate audit/scanner targets twice per peer.
|
||||
fn publish_notify_config_intent(
|
||||
config: &ServerConfig,
|
||||
sub_system: Option<&str>,
|
||||
) -> Option<rustfs_notify::NotificationLifecycleTransition> {
|
||||
(sub_system.is_none() || sub_system.is_some_and(|sub_system| NOTIFY_SUB_SYSTEMS.contains(&sub_system)))
|
||||
.then(|| rustfs_notify::ensure_live_events().publish_config(config.clone()))
|
||||
}
|
||||
|
||||
fn config_preflight_subsystems(sub_system: Option<&str>) -> Vec<&str> {
|
||||
if let Some(sub_system) = sub_system {
|
||||
return is_dynamic_config_subsystem(sub_system)
|
||||
@@ -1884,78 +1869,26 @@ async fn preflight_config_intent(sub_system: Option<&str>) -> S3Result<()> {
|
||||
Ok(())
|
||||
}
|
||||
|
||||
async fn wait_notify_config_intent(transition: Option<rustfs_notify::NotificationLifecycleTransition>) -> S3Result<bool> {
|
||||
let Some(transition) = transition else {
|
||||
return Ok(false);
|
||||
};
|
||||
transition.wait().await.map_err(|err| {
|
||||
warn!(error = %err, "Failed to apply local notification config");
|
||||
s3_error!(InternalError, "failed to apply notification config")
|
||||
})?;
|
||||
Ok(true)
|
||||
}
|
||||
|
||||
async fn reload_non_notify_dynamic_subsystems() -> Vec<String> {
|
||||
let mut failures = Vec::new();
|
||||
for sub_system in FULL_CONFIG_WORKER_SUBSYSTEMS {
|
||||
if NOTIFY_SUB_SYSTEMS.contains(&sub_system) {
|
||||
continue;
|
||||
}
|
||||
if reload_dynamic_config_runtime_state(sub_system).await.is_err() {
|
||||
failures.push(format!("local {sub_system}"));
|
||||
warn!(
|
||||
event = EVENT_CONFIG_WORKER_RELOAD_FAILED,
|
||||
component = LOG_COMPONENT_ADMIN,
|
||||
subsystem = LOG_SUBSYSTEM_CONFIG,
|
||||
config_subsystem = sub_system,
|
||||
state = CONFIG_WORKER_RELOAD_FAILURE_STATE,
|
||||
reason = "apply_failed",
|
||||
"Published server config but failed to reload a local worker subsystem"
|
||||
);
|
||||
}
|
||||
}
|
||||
failures
|
||||
}
|
||||
|
||||
fn finish_config_reconciliation(errors: Vec<String>) -> S3Result<()> {
|
||||
fn finish_config_reconciliation(errors: Vec<String>, persisted: bool) -> S3Result<()> {
|
||||
if errors.is_empty() {
|
||||
Ok(())
|
||||
} else {
|
||||
} else if persisted {
|
||||
Err(s3_error!(
|
||||
InternalError,
|
||||
"server config persisted but runtime convergence failed: {}",
|
||||
errors.join("; ")
|
||||
))
|
||||
} else {
|
||||
Err(s3_error!(
|
||||
InternalError,
|
||||
"server config was unchanged but runtime convergence failed: {}",
|
||||
errors.join("; ")
|
||||
))
|
||||
}
|
||||
}
|
||||
|
||||
async fn reconcile_targeted_config(
|
||||
sub_system: Option<String>,
|
||||
storage_class_applied: bool,
|
||||
notify_transition: Option<rustfs_notify::NotificationLifecycleTransition>,
|
||||
) -> S3Result<bool> {
|
||||
let mut errors = Vec::new();
|
||||
let notify_applied = notify_transition.is_some();
|
||||
if let Err(err) = wait_notify_config_intent(notify_transition).await {
|
||||
warn!(error = %err, "Local notification config failed to converge");
|
||||
errors.push("local notify".to_string());
|
||||
}
|
||||
|
||||
let config_applied = if notify_applied {
|
||||
if let Some(sub_system) = sub_system.as_deref()
|
||||
&& let Err(err) = signal_dynamic_config_reload_checked(sub_system).await
|
||||
{
|
||||
warn!(config_subsystem = sub_system, error = %err, "Peer config reload failed");
|
||||
errors.push(format!("peer {sub_system}"));
|
||||
}
|
||||
true
|
||||
} else if storage_class_applied {
|
||||
if let Err(err) = signal_dynamic_config_reload_checked(STORAGE_CLASS_SUB_SYS).await {
|
||||
warn!(error = %err, "Peer storage-class reload failed");
|
||||
errors.push(format!("peer {STORAGE_CLASS_SUB_SYS}"));
|
||||
}
|
||||
true
|
||||
} else if let Some(sub_system) = sub_system.as_deref()
|
||||
async fn reconcile_targeted_config(sub_system: Option<String>, persisted: bool, mut errors: Vec<String>) -> S3Result<bool> {
|
||||
let config_applied = if let Some(sub_system) = sub_system.as_deref()
|
||||
&& is_dynamic_config_subsystem(sub_system)
|
||||
{
|
||||
let config_applied = match reload_dynamic_config_runtime_state(sub_system).await {
|
||||
@@ -1979,21 +1912,19 @@ async fn reconcile_targeted_config(
|
||||
false
|
||||
};
|
||||
|
||||
finish_config_reconciliation(errors)?;
|
||||
finish_config_reconciliation(errors, persisted)?;
|
||||
Ok(config_applied)
|
||||
}
|
||||
|
||||
async fn reconcile_full_config(notify_transition: Option<rustfs_notify::NotificationLifecycleTransition>) -> S3Result<()> {
|
||||
let mut errors = Vec::new();
|
||||
if let Err(err) = wait_notify_config_intent(notify_transition).await {
|
||||
warn!(error = %err, "Local notification config failed to converge");
|
||||
async fn reconcile_full_config(persisted: bool, mut errors: Vec<String>) -> S3Result<()> {
|
||||
if let Err(err) = reload_dynamic_config_runtime_state(NOTIFY_WEBHOOK_SUB_SYS).await {
|
||||
warn!(error = %err, "Local notification config failed to converge from durable config");
|
||||
errors.push("local notify".to_string());
|
||||
}
|
||||
if let Err(err) = signal_dynamic_config_reload_checked(STORAGE_CLASS_SUB_SYS).await {
|
||||
warn!(error = %err, "Peer storage-class reload failed");
|
||||
errors.push(format!("peer {STORAGE_CLASS_SUB_SYS}"));
|
||||
}
|
||||
errors.extend(reload_non_notify_dynamic_subsystems().await);
|
||||
if let Err(err) = signal_dynamic_config_reload_checked(NOTIFY_WEBHOOK_SUB_SYS).await {
|
||||
warn!(error = %err, "Peer notification config reload failed");
|
||||
errors.push("peer notify".to_string());
|
||||
@@ -2002,91 +1933,82 @@ async fn reconcile_full_config(notify_transition: Option<rustfs_notify::Notifica
|
||||
warn!(error = %err, "Peer config snapshot reload failed");
|
||||
errors.push("peer config snapshot".to_string());
|
||||
}
|
||||
finish_config_reconciliation(errors)
|
||||
finish_config_reconciliation(errors, persisted)
|
||||
}
|
||||
|
||||
struct PersistedConfigTransaction {
|
||||
previous_config: ServerConfig,
|
||||
history_restore_id: Option<String>,
|
||||
storage_class_applied: bool,
|
||||
notify_transition: Option<rustfs_notify::NotificationLifecycleTransition>,
|
||||
persisted: bool,
|
||||
committed_generation: Option<Uuid>,
|
||||
}
|
||||
|
||||
async fn persist_server_config_transaction(
|
||||
config: ServerConfig,
|
||||
prepared: PreparedRuntimeConfig,
|
||||
snapshot: AdminServerConfigSnapshot,
|
||||
sub_system: Option<&str>,
|
||||
) -> S3Result<PersistedConfigTransaction> {
|
||||
snapshot.ensure_lock_held().map_err(ApiError::from).map_err(S3Error::from)?;
|
||||
let previous_config = snapshot.config.clone();
|
||||
let history_restore_id = save_server_config_history_snapshot(&previous_config).await?;
|
||||
if snapshot.is_lock_lost() {
|
||||
cleanup_failed_config_history_snapshot(&history_restore_id).await;
|
||||
cleanup_config_history_snapshot(&history_restore_id).await;
|
||||
return reject_lost_config_transaction(snapshot, "before persistence").await;
|
||||
}
|
||||
|
||||
let persisted = match save_server_config_to_store(&config, &snapshot).await {
|
||||
Ok(persisted) => persisted,
|
||||
let save_result = match save_server_config_to_store(&config, &snapshot).await {
|
||||
Ok(result) => result,
|
||||
Err(err) => {
|
||||
cleanup_failed_config_history_snapshot(&history_restore_id).await;
|
||||
cleanup_config_history_snapshot(&history_restore_id).await;
|
||||
return Err(err);
|
||||
}
|
||||
};
|
||||
let persisted = save_result.persisted();
|
||||
let committed_generation = save_result.generation();
|
||||
let history_restore_id = if persisted {
|
||||
Some(history_restore_id)
|
||||
} else {
|
||||
cleanup_failed_config_history_snapshot(&history_restore_id).await;
|
||||
cleanup_config_history_snapshot(&history_restore_id).await;
|
||||
None
|
||||
};
|
||||
if snapshot.is_lock_lost() {
|
||||
return reject_lost_config_transaction(snapshot, "after persistence").await;
|
||||
}
|
||||
|
||||
let storage_class_applied = sub_system.is_none_or(|value| value == STORAGE_CLASS_SUB_SYS);
|
||||
if storage_class_applied {
|
||||
if let Err(err) = publish_prepared_config_snapshots(config.clone(), prepared) {
|
||||
return Err(match history_restore_id.as_deref() {
|
||||
Some(restore_id) => s3_error!(
|
||||
InternalError,
|
||||
"config persisted but runtime publish failed; recovery snapshot restoreId={}: {}",
|
||||
restore_id,
|
||||
err
|
||||
),
|
||||
None => err,
|
||||
});
|
||||
}
|
||||
} else {
|
||||
publish_server_config(config.clone());
|
||||
}
|
||||
let notify_transition = publish_notify_config_intent(&config, sub_system);
|
||||
if snapshot.is_lock_lost() {
|
||||
return reject_lost_config_transaction(snapshot, "while publishing runtime snapshots").await;
|
||||
}
|
||||
|
||||
drop(snapshot);
|
||||
Ok(PersistedConfigTransaction {
|
||||
previous_config,
|
||||
history_restore_id,
|
||||
storage_class_applied,
|
||||
notify_transition,
|
||||
persisted,
|
||||
committed_generation,
|
||||
})
|
||||
}
|
||||
|
||||
async fn reconcile_committed_config(sub_system: Option<String>, persisted: bool) -> S3Result<bool> {
|
||||
let mut errors = Vec::new();
|
||||
let publication_result = match sub_system.as_deref() {
|
||||
Some(sub_system) => publish_latest_runtime_config_snapshot(sub_system).await,
|
||||
None => reload_runtime_config_snapshot().await,
|
||||
};
|
||||
if let Err(err) = publication_result {
|
||||
warn!(error = %err, "Failed to publish the latest durable server config");
|
||||
errors.push("local config snapshot".to_string());
|
||||
}
|
||||
|
||||
if sub_system.is_none() {
|
||||
reconcile_full_config(persisted, errors).await?;
|
||||
return Ok(false);
|
||||
}
|
||||
|
||||
reconcile_targeted_config(sub_system, persisted, errors).await
|
||||
}
|
||||
|
||||
async fn commit_server_config_transaction(
|
||||
config: ServerConfig,
|
||||
prepared: PreparedRuntimeConfig,
|
||||
snapshot: AdminServerConfigSnapshot,
|
||||
sub_system: Option<String>,
|
||||
) -> S3Result<bool> {
|
||||
let transaction = persist_server_config_transaction(config, prepared, snapshot, sub_system.as_deref()).await?;
|
||||
|
||||
if sub_system.is_none() {
|
||||
reconcile_full_config(transaction.notify_transition).await?;
|
||||
return Ok(false);
|
||||
let transaction = persist_server_config_transaction(config, snapshot).await?;
|
||||
let result = reconcile_committed_config(sub_system, transaction.persisted).await;
|
||||
match (result, transaction.history_restore_id.as_deref()) {
|
||||
(Err(err), Some(restore_id)) => Err(s3_error!(InternalError, "{}; recovery snapshot restoreId={}", err, restore_id)),
|
||||
(result, _) => result,
|
||||
}
|
||||
|
||||
reconcile_targeted_config(sub_system, transaction.storage_class_applied, transaction.notify_transition).await
|
||||
}
|
||||
|
||||
pub struct GetConfigKVHandler {}
|
||||
@@ -2127,8 +2049,8 @@ impl Operation for SetConfigKVHandler {
|
||||
let snapshot = load_server_config_snapshot_from_store().await?;
|
||||
let mut config = snapshot.config.clone();
|
||||
apply_set_directives(&mut config, &directives)?;
|
||||
let prepared = prepare_server_config(&config, sub_system.as_deref()).await?;
|
||||
commit_server_config_transaction(config, prepared, snapshot, sub_system).await
|
||||
prepare_server_config(&config, sub_system.as_deref()).await?;
|
||||
commit_server_config_transaction(config, snapshot, sub_system).await
|
||||
})
|
||||
.await?;
|
||||
|
||||
@@ -2155,8 +2077,8 @@ impl Operation for DelConfigKVHandler {
|
||||
let snapshot = load_server_config_snapshot_from_store().await?;
|
||||
let mut config = snapshot.config.clone();
|
||||
apply_delete_directives(&mut config, &directives);
|
||||
let prepared = prepare_server_config(&config, sub_system.as_deref()).await?;
|
||||
commit_server_config_transaction(config, prepared, snapshot, sub_system).await
|
||||
prepare_server_config(&config, sub_system.as_deref()).await?;
|
||||
commit_server_config_transaction(config, snapshot, sub_system).await
|
||||
})
|
||||
.await?;
|
||||
|
||||
@@ -2239,15 +2161,19 @@ impl Operation for RestoreConfigHistoryKVHandler {
|
||||
supervise_admin_mutation("config mutation", async move {
|
||||
preflight_config_intent(None).await?;
|
||||
let snapshot = load_server_config_snapshot_from_store().await?;
|
||||
let prepared = prepare_server_config(&config, None).await?;
|
||||
let restored_config = config.clone();
|
||||
let transaction = persist_server_config_transaction(config, prepared, snapshot, None).await?;
|
||||
let Err(restore_error) = reconcile_full_config(transaction.notify_transition).await else {
|
||||
prepare_server_config(&config, None).await?;
|
||||
let transaction = persist_server_config_transaction(config, snapshot).await?;
|
||||
let Err(restore_error) = reconcile_committed_config(None, transaction.persisted).await else {
|
||||
return Ok(());
|
||||
};
|
||||
|
||||
if !transaction.persisted {
|
||||
return Err(restore_error);
|
||||
}
|
||||
|
||||
let previous_config = transaction.previous_config;
|
||||
let recovery_restore_id = transaction.history_restore_id;
|
||||
let committed_generation = transaction.committed_generation;
|
||||
let recovery_reference = recovery_restore_id.as_deref().unwrap_or("not-created");
|
||||
let rollback_snapshot = match load_server_config_snapshot_from_store().await {
|
||||
Ok(snapshot) => snapshot,
|
||||
@@ -2261,7 +2187,9 @@ impl Operation for RestoreConfigHistoryKVHandler {
|
||||
));
|
||||
}
|
||||
};
|
||||
if let Err(rollback_error) = validate_restore_rollback_generation(&rollback_snapshot.config, &restored_config) {
|
||||
if let Err(rollback_error) =
|
||||
validate_restore_rollback_generation(rollback_snapshot.generation(), committed_generation)
|
||||
{
|
||||
return Err(s3_error!(
|
||||
InternalError,
|
||||
"config restore failed: {}; automatic rollback failed: {}; recovery snapshot restoreId={}",
|
||||
@@ -2271,12 +2199,20 @@ impl Operation for RestoreConfigHistoryKVHandler {
|
||||
));
|
||||
}
|
||||
|
||||
let rollback_prepared = prepare_server_config(&previous_config, None).await?;
|
||||
rollback_snapshot
|
||||
if let Err(rollback_error) = prepare_server_config(&previous_config, None).await {
|
||||
return Err(s3_error!(
|
||||
InternalError,
|
||||
"config restore failed: {}; automatic rollback failed: {}; recovery snapshot restoreId={}",
|
||||
restore_error,
|
||||
rollback_error,
|
||||
recovery_reference
|
||||
));
|
||||
}
|
||||
if let Err(rollback_error) = rollback_snapshot
|
||||
.ensure_lock_held()
|
||||
.map_err(ApiError::from)
|
||||
.map_err(S3Error::from)?;
|
||||
if let Err(rollback_error) = save_server_config_to_store(&previous_config, &rollback_snapshot).await {
|
||||
.map_err(S3Error::from)
|
||||
{
|
||||
return Err(s3_error!(
|
||||
InternalError,
|
||||
"config restore failed: {}; automatic rollback failed: {}; recovery snapshot restoreId={}",
|
||||
@@ -2285,10 +2221,9 @@ impl Operation for RestoreConfigHistoryKVHandler {
|
||||
recovery_reference
|
||||
));
|
||||
}
|
||||
if rollback_snapshot.is_lock_lost() {
|
||||
if let Err(rollback_error) =
|
||||
reject_lost_config_transaction::<()>(rollback_snapshot, "after restore rollback persistence").await
|
||||
{
|
||||
let rollback_persisted = match save_server_config_to_store(&previous_config, &rollback_snapshot).await {
|
||||
Ok(result) => result.persisted(),
|
||||
Err(rollback_error) => {
|
||||
return Err(s3_error!(
|
||||
InternalError,
|
||||
"config restore failed: {}; automatic rollback failed: {}; recovery snapshot restoreId={}",
|
||||
@@ -2297,46 +2232,20 @@ impl Operation for RestoreConfigHistoryKVHandler {
|
||||
recovery_reference
|
||||
));
|
||||
}
|
||||
return Err(s3_error!(InternalError, "restore rollback lock-loss handling returned unexpectedly"));
|
||||
}
|
||||
|
||||
if let Err(rollback_error) = publish_prepared_config_snapshots(previous_config.clone(), rollback_prepared) {
|
||||
return Err(s3_error!(
|
||||
InternalError,
|
||||
"config restore failed: {}; automatic rollback publish failed: {}; recovery snapshot restoreId={}",
|
||||
restore_error,
|
||||
rollback_error,
|
||||
recovery_reference
|
||||
));
|
||||
}
|
||||
let rollback_transition = publish_notify_config_intent(&previous_config, None);
|
||||
if rollback_snapshot.is_lock_lost() {
|
||||
if let Err(rollback_error) =
|
||||
reject_lost_config_transaction::<()>(rollback_snapshot, "while publishing restore rollback").await
|
||||
{
|
||||
return Err(s3_error!(
|
||||
InternalError,
|
||||
"config restore failed: {}; automatic rollback failed: {}; recovery snapshot restoreId={}",
|
||||
restore_error,
|
||||
rollback_error,
|
||||
recovery_reference
|
||||
));
|
||||
}
|
||||
return Err(s3_error!(InternalError, "restore rollback lock-loss handling returned unexpectedly"));
|
||||
}
|
||||
};
|
||||
drop(rollback_snapshot);
|
||||
|
||||
if let Err(rollback_error) = reconcile_full_config(rollback_transition).await {
|
||||
if let Err(rollback_error) = reconcile_committed_config(None, rollback_persisted).await {
|
||||
return Err(s3_error!(
|
||||
InternalError,
|
||||
"config restore failed: {}; persisted rollback did not converge: {}; recovery snapshot restoreId={}",
|
||||
"config restore failed: {}; automatic rollback convergence failed: {}; recovery snapshot restoreId={}",
|
||||
restore_error,
|
||||
rollback_error,
|
||||
recovery_reference
|
||||
));
|
||||
}
|
||||
if let Some(recovery_restore_id) = recovery_restore_id.as_deref() {
|
||||
cleanup_failed_config_history_snapshot(recovery_restore_id).await;
|
||||
if rollback_persisted && let Some(recovery_restore_id) = recovery_restore_id.as_deref() {
|
||||
cleanup_config_history_snapshot(recovery_restore_id).await;
|
||||
}
|
||||
Err(s3_error!(InternalError, "config restore failed and was rolled back: {}", restore_error))
|
||||
})
|
||||
@@ -2377,8 +2286,8 @@ impl Operation for SetConfigHandler {
|
||||
let snapshot = load_server_config_snapshot_from_store().await?;
|
||||
let mut config = ServerConfig::new();
|
||||
apply_set_directives(&mut config, &directives)?;
|
||||
let prepared = prepare_server_config(&config, None).await?;
|
||||
commit_server_config_transaction(config, prepared, snapshot, None).await?;
|
||||
prepare_server_config(&config, None).await?;
|
||||
commit_server_config_transaction(config, snapshot, None).await?;
|
||||
Ok(())
|
||||
})
|
||||
.await?;
|
||||
@@ -2408,6 +2317,23 @@ mod tests {
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn reconciliation_error_reports_whether_config_was_persisted() {
|
||||
let persisted = finish_config_reconciliation(vec!["local config snapshot".to_string()], true)
|
||||
.expect_err("committed config convergence failure must be reported");
|
||||
let unchanged = finish_config_reconciliation(vec!["local config snapshot".to_string()], false)
|
||||
.expect_err("unchanged config convergence failure must be reported");
|
||||
|
||||
assert_eq!(
|
||||
persisted.message(),
|
||||
Some("server config persisted but runtime convergence failed: local config snapshot")
|
||||
);
|
||||
assert_eq!(
|
||||
unchanged.message(),
|
||||
Some("server config was unchanged but runtime convergence failed: local config snapshot")
|
||||
);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn tokenize_config_line_handles_quotes_and_escapes() {
|
||||
let tokens = tokenize_config_line(r#"identity_openid client_id="console app" client_secret="s3cr\"et" enable=on"#)
|
||||
@@ -3191,23 +3117,32 @@ notify_webhook:secondary endpoint="https://secondary.example" auth_token="second
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn restore_rollback_generation_rejects_concurrent_config_change() {
|
||||
crate::admin::storage_api::config::init_admin_config_defaults();
|
||||
let restored = ServerConfig::new();
|
||||
let mut concurrent = restored.clone();
|
||||
apply_set_directives(
|
||||
&mut concurrent,
|
||||
&parse_config_directives(r#"identity_openid client_id="concurrent-client""#, false).expect("parse concurrent"),
|
||||
)
|
||||
.expect("apply concurrent");
|
||||
fn restore_rollback_generation_accepts_committed_write_identity() {
|
||||
let committed = Uuid::from_u128(1);
|
||||
validate_restore_rollback_generation(Some(committed), Some(committed)).expect("matching committed generation");
|
||||
}
|
||||
|
||||
validate_restore_rollback_generation(&restored, &restored).expect("unchanged restore generation");
|
||||
let error = validate_restore_rollback_generation(&concurrent, &restored)
|
||||
.expect_err("concurrent change must fence automatic rollback");
|
||||
#[test]
|
||||
fn restore_rollback_generation_rejects_aba_with_matching_config_content() {
|
||||
let error = validate_restore_rollback_generation(Some(Uuid::from_u128(2)), Some(Uuid::from_u128(1)))
|
||||
.expect_err("a later generation must fence automatic rollback even when config content matches");
|
||||
assert_eq!(error.code(), &S3ErrorCode::InvalidRequest);
|
||||
assert!(error.to_string().contains("concurrent configuration change"));
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn restore_rollback_generation_rejects_missing_write_identity() {
|
||||
for (current, committed) in [
|
||||
(None, Some(Uuid::from_u128(1))),
|
||||
(Some(Uuid::from_u128(1)), None),
|
||||
(None, None),
|
||||
] {
|
||||
let error = validate_restore_rollback_generation(current, committed)
|
||||
.expect_err("missing generation metadata must fence automatic rollback");
|
||||
assert_eq!(error.code(), &S3ErrorCode::InvalidRequest);
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn legacy_directive_history_is_rejected_without_being_applied_as_a_snapshot() {
|
||||
let error = decode_config_history_snapshot(b"identity_openid client_id=\"legacy\"")
|
||||
|
||||
@@ -2708,18 +2708,16 @@ impl<T: Operation> S3Router<T> {
|
||||
|
||||
pub fn insert(&mut self, method: Method, path: &str, operation: T) -> std::io::Result<()> {
|
||||
let path = Self::make_route_str(method, path);
|
||||
#[cfg(test)]
|
||||
let registered_path = path.clone();
|
||||
|
||||
// warn!("set uri {}", &path);
|
||||
|
||||
#[cfg(test)]
|
||||
{
|
||||
self.router.insert(path.clone(), operation).map_err(std::io::Error::other)?;
|
||||
self.registered_routes.push(path);
|
||||
}
|
||||
|
||||
#[cfg(not(test))]
|
||||
self.router.insert(path, operation).map_err(std::io::Error::other)?;
|
||||
|
||||
#[cfg(test)]
|
||||
self.registered_routes.push(registered_path);
|
||||
|
||||
Ok(())
|
||||
}
|
||||
|
||||
|
||||
@@ -16,7 +16,9 @@ use crate::admin::runtime_sources::{
|
||||
AppContext, current_app_context, current_notification_system_for_context, current_object_store_handle_for_context,
|
||||
publish_server_config, publish_storage_class_config,
|
||||
};
|
||||
use crate::admin::storage_api::config::{STORAGE_CLASS_SUB_SYS, read_admin_config_without_migrate, storageclass};
|
||||
use crate::admin::storage_api::config::{
|
||||
STORAGE_CLASS_SUB_SYS, read_existing_admin_server_config_no_lock, storageclass, with_admin_server_config_read_lock,
|
||||
};
|
||||
use crate::admin::storage_api::contract::admin::StorageAdminApi;
|
||||
use crate::admin::storage_api::runtime::ECStore;
|
||||
use crate::server::{
|
||||
@@ -43,6 +45,25 @@ use url::Url;
|
||||
|
||||
static RUNTIME_CONFIG_RELOAD_MUTEX: AsyncMutex<()> = AsyncMutex::const_new(());
|
||||
|
||||
// Runtime publication lock order: reload mutex -> server-config local lock ->
|
||||
// transaction lock -> config object lock.
|
||||
pub(crate) async fn with_runtime_config_reload_lock<Fut, T>(operation: Fut) -> S3Result<T>
|
||||
where
|
||||
Fut: Future<Output = S3Result<T>> + Send + 'static,
|
||||
T: Send + 'static,
|
||||
{
|
||||
let reload_guard = RUNTIME_CONFIG_RELOAD_MUTEX.lock().await;
|
||||
tokio::spawn(async move {
|
||||
let _reload_guard = reload_guard;
|
||||
operation.await
|
||||
})
|
||||
.await
|
||||
.map_err(|err| {
|
||||
let outcome = if err.is_cancelled() { "cancelled" } else { "panicked" };
|
||||
internal_error(format!("runtime config reload task {outcome}"))
|
||||
})?
|
||||
}
|
||||
|
||||
pub fn is_dynamic_config_subsystem(sub_system: &str) -> bool {
|
||||
NOTIFY_SUB_SYSTEMS.contains(&sub_system)
|
||||
|| matches!(
|
||||
@@ -121,11 +142,6 @@ impl PreparedRuntimeConfig {
|
||||
fn publish_storage_class_for_context(self, context: Option<&AppContext>) -> S3Result<()> {
|
||||
self.publish_storage_class_for_context_with(context, publish_storage_class_config)
|
||||
}
|
||||
|
||||
pub(crate) fn publish_storage_class(self) -> S3Result<()> {
|
||||
let context = current_app_context();
|
||||
self.publish_storage_class_for_context(context.as_deref())
|
||||
}
|
||||
}
|
||||
|
||||
fn publish_server_config_for_context(context: Option<&AppContext>, config: ServerConfig) {
|
||||
@@ -396,7 +412,18 @@ pub async fn apply_dynamic_config_for_subsystem(config: &ServerConfig, sub_syste
|
||||
}
|
||||
|
||||
pub async fn reload_dynamic_config_runtime_state_for_context(context: Option<&AppContext>, sub_system: &str) -> S3Result<()> {
|
||||
let _reload_guard = RUNTIME_CONFIG_RELOAD_MUTEX.lock().await;
|
||||
let context = context.cloned();
|
||||
let sub_system = sub_system.to_owned();
|
||||
with_runtime_config_reload_lock(async move {
|
||||
reload_dynamic_config_runtime_state_under_reload_lock_for_context(context.as_ref(), &sub_system).await
|
||||
})
|
||||
.await
|
||||
}
|
||||
|
||||
async fn reload_dynamic_config_runtime_state_under_reload_lock_for_context(
|
||||
context: Option<&AppContext>,
|
||||
sub_system: &str,
|
||||
) -> S3Result<()> {
|
||||
if sub_system == MODULE_SWITCHES_SIGNAL_SUBSYSTEM {
|
||||
let store = resolve_runtime_config_store_for_context(context)?;
|
||||
let notify_result = reconcile_event_notifier_from_store(store).await;
|
||||
@@ -414,18 +441,54 @@ pub async fn reload_dynamic_config_runtime_state_for_context(context: Option<&Ap
|
||||
}
|
||||
|
||||
let store = resolve_runtime_config_store_for_context(context)?;
|
||||
let config = read_admin_config_without_migrate(store).await.map_err(|err| {
|
||||
warn!("peer reload_dynamic_config: failed to load server config for {sub_system}: {err}");
|
||||
internal_error(format!("failed to load server config: {err}"))
|
||||
})?;
|
||||
let read_store = store.clone();
|
||||
let publication_context = context.cloned();
|
||||
let publication_sub_system = sub_system.to_owned();
|
||||
let notify_config = with_admin_server_config_read_lock(store, move || async move {
|
||||
let config = read_existing_admin_server_config_no_lock(read_store).await.map_err(|err| {
|
||||
warn!("peer reload_dynamic_config: failed to load server config for {publication_sub_system}: {err}");
|
||||
internal_error(format!("failed to load server config: {err}"))
|
||||
})?;
|
||||
let prepared = prepare_server_config_for_context(publication_context.as_ref(), &config, Some(&publication_sub_system))
|
||||
.await
|
||||
.map_err(|err| {
|
||||
if publication_sub_system == STORAGE_CLASS_SUB_SYS {
|
||||
internal_error(format!("failed to apply storage class config: {err}"))
|
||||
} else {
|
||||
err
|
||||
}
|
||||
})?;
|
||||
|
||||
if matches!(sub_system, SCANNER_SUB_SYS | HEAL_SUB_SYS) {
|
||||
validate_server_config_for_context(context, &config, Some(sub_system)).await?;
|
||||
// Scanner cycles refresh from the process-wide server config before
|
||||
// each pass. Publish the same validated snapshot first so that refresh
|
||||
// cannot overwrite this peer's dynamic scanner update with stale data.
|
||||
publish_server_config_for_context(context, config.clone());
|
||||
}
|
||||
if publication_sub_system == STORAGE_CLASS_SUB_SYS {
|
||||
prepared.publish_storage_class_for_context(publication_context.as_ref())?;
|
||||
return Ok::<Option<ServerConfig>, S3Error>(None);
|
||||
}
|
||||
if NOTIFY_SUB_SYSTEMS.contains(&publication_sub_system.as_str()) {
|
||||
return Ok(Some(config));
|
||||
}
|
||||
if matches!(publication_sub_system.as_str(), SCANNER_SUB_SYS | HEAL_SUB_SYS) {
|
||||
publish_server_config_for_context(publication_context.as_ref(), config.clone());
|
||||
}
|
||||
apply_dynamic_config_for_subsystem_for_context(publication_context.as_ref(), &config, &publication_sub_system)
|
||||
.await
|
||||
.inspect_err(|_| {
|
||||
warn!(
|
||||
config_subsystem = publication_sub_system,
|
||||
reason = "apply_failed",
|
||||
"Peer dynamic config apply failed"
|
||||
);
|
||||
})?;
|
||||
Ok(None)
|
||||
})
|
||||
.await
|
||||
.map_err(|err| {
|
||||
warn!("peer reload_dynamic_config: failed to acquire server config publication fence for {sub_system}: {err}");
|
||||
internal_error(format!("failed to lock server config: {err}"))
|
||||
})??;
|
||||
|
||||
let Some(config) = notify_config else {
|
||||
return Ok(());
|
||||
};
|
||||
apply_dynamic_config_for_subsystem_for_context(context, &config, sub_system)
|
||||
.await
|
||||
.inspect_err(|_| {
|
||||
@@ -439,23 +502,16 @@ pub async fn reload_dynamic_config_runtime_state(sub_system: &str) -> S3Result<(
|
||||
reload_dynamic_config_runtime_state_for_context(context.as_deref(), sub_system).await
|
||||
}
|
||||
|
||||
async fn reload_runtime_config_snapshot_with<ReadFuture, Prepare, PrepareFuture, Publish, ApplyWorkers, ApplyWorkersFuture>(
|
||||
read: ReadFuture,
|
||||
prepare: Prepare,
|
||||
publish: Publish,
|
||||
async fn reload_runtime_config_snapshot_with<PublishFuture, ApplyWorkers, ApplyWorkersFuture>(
|
||||
publish_snapshot: PublishFuture,
|
||||
apply_workers: ApplyWorkers,
|
||||
) -> S3Result<()>
|
||||
where
|
||||
ReadFuture: Future<Output = S3Result<ServerConfig>>,
|
||||
Prepare: FnOnce(ServerConfig) -> PrepareFuture,
|
||||
PrepareFuture: Future<Output = S3Result<(ServerConfig, PreparedRuntimeConfig)>>,
|
||||
Publish: FnOnce(&ServerConfig, PreparedRuntimeConfig) -> S3Result<()>,
|
||||
PublishFuture: Future<Output = S3Result<ServerConfig>>,
|
||||
ApplyWorkers: FnOnce(ServerConfig) -> ApplyWorkersFuture,
|
||||
ApplyWorkersFuture: Future<Output = S3Result<()>>,
|
||||
{
|
||||
let config = read.await?;
|
||||
let (config, prepared) = prepare(config).await?;
|
||||
publish(&config, prepared)?;
|
||||
let config = publish_snapshot.await?;
|
||||
|
||||
// Worker reloads mutate live state and have no rollback contract, so the
|
||||
// validated snapshots stay published. The RPC still reports convergence
|
||||
@@ -474,55 +530,105 @@ where
|
||||
Ok(())
|
||||
}
|
||||
|
||||
pub async fn reload_runtime_config_snapshot_for_context(context: Option<&AppContext>) -> S3Result<()> {
|
||||
let _reload_guard = RUNTIME_CONFIG_RELOAD_MUTEX.lock().await;
|
||||
async fn publish_latest_runtime_config_snapshot_under_reload_lock_for_context<ApplyWorkers, ApplyWorkersFuture>(
|
||||
context: Option<&AppContext>,
|
||||
sub_system: Option<&str>,
|
||||
apply_workers: ApplyWorkers,
|
||||
) -> S3Result<()>
|
||||
where
|
||||
ApplyWorkers: FnOnce(ServerConfig) -> ApplyWorkersFuture + Send + 'static,
|
||||
ApplyWorkersFuture: Future<Output = S3Result<()>> + Send + 'static,
|
||||
{
|
||||
let store = resolve_runtime_config_store_for_context(context)?;
|
||||
let read_store = store.clone();
|
||||
let publication_context = context.cloned();
|
||||
let publication_sub_system = sub_system.map(str::to_owned);
|
||||
|
||||
reload_runtime_config_snapshot_with(
|
||||
async move {
|
||||
read_admin_config_without_migrate(store).await.map_err(|err| {
|
||||
warn!("peer reload_runtime_config_snapshot: failed to load server config: {err}");
|
||||
internal_error(format!("failed to load server config: {err}"))
|
||||
})
|
||||
},
|
||||
|config| async move {
|
||||
let prepared = prepare_server_config_for_context(context, &config, None).await.map_err(|_| {
|
||||
warn!("peer reload_runtime_config_snapshot: failed to prepare server config");
|
||||
internal_error("failed to prepare server config")
|
||||
})?;
|
||||
Ok((config, prepared))
|
||||
},
|
||||
|config, prepared| {
|
||||
prepared.publish_storage_class_for_context(context)?;
|
||||
publish_server_config_for_context(context, config.clone());
|
||||
Ok(())
|
||||
},
|
||||
|config| async move {
|
||||
let mut failures = Vec::new();
|
||||
for sub_system in FULL_CONFIG_WORKER_SUBSYSTEMS {
|
||||
if apply_dynamic_config_for_subsystem_for_context(context, &config, sub_system)
|
||||
.await
|
||||
.is_err()
|
||||
with_admin_server_config_read_lock(store, move || async move {
|
||||
let config = read_existing_admin_server_config_no_lock(read_store).await.map_err(|err| {
|
||||
warn!("runtime config publication failed to load server config: {err}");
|
||||
internal_error(format!("failed to load server config: {err}"))
|
||||
})?;
|
||||
let prepared =
|
||||
prepare_server_config_for_context(publication_context.as_ref(), &config, publication_sub_system.as_deref())
|
||||
.await
|
||||
.map_err(|err| match publication_sub_system.as_deref() {
|
||||
None => {
|
||||
warn!("peer reload_runtime_config_snapshot: failed to prepare server config");
|
||||
internal_error("failed to prepare server config")
|
||||
}
|
||||
Some(STORAGE_CLASS_SUB_SYS) => internal_error(format!("failed to apply storage class config: {err}")),
|
||||
Some(_) => err,
|
||||
})?;
|
||||
reload_runtime_config_snapshot_with(
|
||||
async move {
|
||||
if publication_sub_system
|
||||
.as_deref()
|
||||
.is_none_or(|sub_system| sub_system == STORAGE_CLASS_SUB_SYS)
|
||||
{
|
||||
failures.push(sub_system);
|
||||
warn!(
|
||||
event = EVENT_CONFIG_WORKER_RELOAD_FAILED,
|
||||
component = LOG_COMPONENT_ADMIN,
|
||||
subsystem = LOG_SUBSYSTEM_CONFIG,
|
||||
config_subsystem = sub_system,
|
||||
state = CONFIG_WORKER_RELOAD_FAILURE_STATE,
|
||||
reason = "apply_failed",
|
||||
"Peer runtime config snapshot was published but a subsystem worker reload failed"
|
||||
);
|
||||
prepared.publish_storage_class_for_context(publication_context.as_ref())?;
|
||||
}
|
||||
publish_server_config_for_context(publication_context.as_ref(), config.clone());
|
||||
Ok(config)
|
||||
},
|
||||
apply_workers,
|
||||
)
|
||||
.await
|
||||
})
|
||||
.await
|
||||
.map_err(|err| {
|
||||
warn!("runtime config publication failed to acquire server config publication fence: {err}");
|
||||
internal_error(format!("failed to lock server config: {err}"))
|
||||
})?
|
||||
}
|
||||
|
||||
pub(crate) async fn publish_latest_runtime_config_snapshot(sub_system: &str) -> S3Result<()> {
|
||||
let context = current_app_context();
|
||||
let sub_system = sub_system.to_owned();
|
||||
with_runtime_config_reload_lock(async move {
|
||||
publish_latest_runtime_config_snapshot_under_reload_lock_for_context(context.as_deref(), Some(&sub_system), |_| async {
|
||||
Ok(())
|
||||
})
|
||||
.await
|
||||
})
|
||||
.await
|
||||
}
|
||||
|
||||
pub async fn reload_runtime_config_snapshot_for_context(context: Option<&AppContext>) -> S3Result<()> {
|
||||
let context = context.cloned();
|
||||
with_runtime_config_reload_lock(async move {
|
||||
reload_runtime_config_snapshot_under_reload_lock_for_context(context.as_ref()).await
|
||||
})
|
||||
.await
|
||||
}
|
||||
|
||||
async fn reload_runtime_config_snapshot_under_reload_lock_for_context(context: Option<&AppContext>) -> S3Result<()> {
|
||||
let worker_context = context.cloned();
|
||||
publish_latest_runtime_config_snapshot_under_reload_lock_for_context(context, None, move |config| async move {
|
||||
let mut failures = Vec::new();
|
||||
for sub_system in FULL_CONFIG_WORKER_SUBSYSTEMS {
|
||||
if apply_dynamic_config_for_subsystem_for_context(worker_context.as_ref(), &config, sub_system)
|
||||
.await
|
||||
.is_err()
|
||||
{
|
||||
failures.push(sub_system);
|
||||
warn!(
|
||||
event = EVENT_CONFIG_WORKER_RELOAD_FAILED,
|
||||
component = LOG_COMPONENT_ADMIN,
|
||||
subsystem = LOG_SUBSYSTEM_CONFIG,
|
||||
config_subsystem = sub_system,
|
||||
state = CONFIG_WORKER_RELOAD_FAILURE_STATE,
|
||||
reason = "apply_failed",
|
||||
"Peer runtime config snapshot was published but a subsystem worker reload failed"
|
||||
);
|
||||
}
|
||||
if failures.is_empty() {
|
||||
Ok(())
|
||||
} else {
|
||||
Err(internal_error(format!("runtime worker reload failed: {}", failures.join("; "))))
|
||||
}
|
||||
},
|
||||
)
|
||||
}
|
||||
if failures.is_empty() {
|
||||
Ok(())
|
||||
} else {
|
||||
Err(internal_error(format!("runtime worker reload failed: {}", failures.join("; "))))
|
||||
}
|
||||
})
|
||||
.await
|
||||
}
|
||||
|
||||
@@ -699,6 +805,77 @@ mod tests {
|
||||
}
|
||||
}
|
||||
|
||||
#[tokio::test]
|
||||
async fn runtime_reload_waiter_cancellation_does_not_leave_queued_work() {
|
||||
let _blocker = RUNTIME_CONFIG_RELOAD_MUTEX.lock().await;
|
||||
let (operation_started_tx, mut operation_started_rx) = tokio::sync::oneshot::channel();
|
||||
let mut reload = Box::pin(with_runtime_config_reload_lock(async move {
|
||||
let _ = operation_started_tx.send(());
|
||||
Ok(())
|
||||
}));
|
||||
|
||||
poll_fn(|cx| match reload.as_mut().poll(cx) {
|
||||
Poll::Pending => Poll::Ready(()),
|
||||
Poll::Ready(_) => panic!("reload must wait while the mutex is held"),
|
||||
})
|
||||
.await;
|
||||
drop(reload);
|
||||
|
||||
assert!(
|
||||
matches!(operation_started_rx.try_recv(), Err(tokio::sync::oneshot::error::TryRecvError::Closed)),
|
||||
"cancelling a queued reload must drop its work instead of detaching it"
|
||||
);
|
||||
}
|
||||
|
||||
#[tokio::test]
|
||||
async fn runtime_reload_lock_survives_waiter_cancellation() {
|
||||
let (started_tx, started_rx) = tokio::sync::oneshot::channel();
|
||||
let (release_tx, release_rx) = tokio::sync::oneshot::channel();
|
||||
let (completed_tx, completed_rx) = tokio::sync::oneshot::channel();
|
||||
|
||||
let waiter = tokio::spawn(async move {
|
||||
with_runtime_config_reload_lock(async move {
|
||||
started_tx.send(()).expect("signal runtime reload start");
|
||||
release_rx.await.expect("release supervised runtime reload");
|
||||
completed_tx.send(()).expect("signal runtime reload completion");
|
||||
Ok(())
|
||||
})
|
||||
.await
|
||||
});
|
||||
|
||||
started_rx.await.expect("supervised runtime reload started");
|
||||
waiter.abort();
|
||||
assert!(waiter.await.expect_err("reload waiter should be cancelled").is_cancelled());
|
||||
|
||||
let (contender_polled_tx, contender_polled_rx) = tokio::sync::oneshot::channel();
|
||||
let (contender_entered_tx, mut contender_entered_rx) = tokio::sync::oneshot::channel();
|
||||
let contender = tokio::spawn(async move {
|
||||
let mut lock = Box::pin(RUNTIME_CONFIG_RELOAD_MUTEX.lock());
|
||||
let mut contender_polled_tx = Some(contender_polled_tx);
|
||||
let _guard = poll_fn(|cx| {
|
||||
if let Some(tx) = contender_polled_tx.take() {
|
||||
let _ = tx.send(());
|
||||
}
|
||||
lock.as_mut().poll(cx)
|
||||
})
|
||||
.await;
|
||||
let _ = contender_entered_tx.send(());
|
||||
});
|
||||
|
||||
contender_polled_rx.await.expect("contending reload should be polled");
|
||||
assert!(
|
||||
matches!(contender_entered_rx.try_recv(), Err(tokio::sync::oneshot::error::TryRecvError::Empty)),
|
||||
"a cancelled waiter must not release the reload mutex while its detached reload is still running"
|
||||
);
|
||||
|
||||
release_tx.send(()).expect("release detached runtime reload");
|
||||
completed_rx.await.expect("detached runtime reload should complete");
|
||||
contender_entered_rx
|
||||
.await
|
||||
.expect("contending reload should enter after completion");
|
||||
contender.await.expect("contending reload task should not panic");
|
||||
}
|
||||
|
||||
#[tokio::test]
|
||||
async fn checked_scanner_reload_reports_unreachable_peer() {
|
||||
let temp_dir = TempDir::new().expect("scanner reload temp dir");
|
||||
@@ -1221,6 +1398,123 @@ mod tests {
|
||||
.await;
|
||||
}
|
||||
|
||||
#[tokio::test(flavor = "multi_thread", worker_threads = 4)]
|
||||
#[serial_test::serial(storage_class_env)]
|
||||
async fn full_reload_keeps_worker_apply_inside_durable_read_fence() {
|
||||
temp_env::async_with_vars(
|
||||
[
|
||||
(storageclass::STANDARD_ENV, None::<&str>),
|
||||
(storageclass::RRS_ENV, None::<&str>),
|
||||
(storageclass::OPTIMIZE_ENV, None::<&str>),
|
||||
(storageclass::INLINE_BLOCK_ENV, None::<&str>),
|
||||
],
|
||||
async {
|
||||
let fixture = runtime_config_reload_fixture().await;
|
||||
let older = scanner_server_config("41");
|
||||
save_admin_server_config(fixture.context.object_store(), &older)
|
||||
.await
|
||||
.expect("persist older scanner config");
|
||||
|
||||
let (worker_entered_tx, worker_entered_rx) = tokio::sync::oneshot::channel();
|
||||
let (release_worker_tx, release_worker_rx) = tokio::sync::oneshot::channel();
|
||||
let reload_context = fixture.context.clone();
|
||||
let reload = tokio::spawn(async move {
|
||||
publish_latest_runtime_config_snapshot_under_reload_lock_for_context(
|
||||
Some(&reload_context),
|
||||
None,
|
||||
move |config| async move {
|
||||
worker_entered_tx.send(config).expect("signal runtime config worker entry");
|
||||
release_worker_rx.await.expect("release runtime config worker");
|
||||
Ok(())
|
||||
},
|
||||
)
|
||||
.await
|
||||
});
|
||||
|
||||
let worker_config = tokio::time::timeout(REAL_STORE_TEST_TIMEOUT, worker_entered_rx)
|
||||
.await
|
||||
.expect("worker apply should enter")
|
||||
.expect("worker apply entry signal should be delivered");
|
||||
assert_eq!(
|
||||
worker_config
|
||||
.get_value(SCANNER_SUB_SYS, DEFAULT_DELIMITER)
|
||||
.expect("older scanner config should be loaded")
|
||||
.get(SCANNER_CYCLE),
|
||||
"41"
|
||||
);
|
||||
|
||||
let latest = scanner_server_config("71");
|
||||
let writer_store = fixture.context.object_store();
|
||||
let (writer_polled_tx, writer_polled_rx) = tokio::sync::oneshot::channel();
|
||||
let (writer_entered_tx, mut writer_entered_rx) = tokio::sync::oneshot::channel();
|
||||
let writer = tokio::spawn(async move {
|
||||
let transaction_store = writer_store.clone();
|
||||
let transaction = with_admin_server_config_write_lock(writer_store, move || async move {
|
||||
writer_entered_tx.send(()).expect("signal writer entry");
|
||||
save_admin_server_config_no_lock(transaction_store, &latest).await
|
||||
});
|
||||
tokio::pin!(transaction);
|
||||
let mut writer_polled_tx = Some(writer_polled_tx);
|
||||
poll_fn(|cx| match transaction.as_mut().poll(cx) {
|
||||
Poll::Pending => {
|
||||
if let Some(tx) = writer_polled_tx.take() {
|
||||
let _ = tx.send(());
|
||||
}
|
||||
Poll::Ready(())
|
||||
}
|
||||
Poll::Ready(_) => panic!("writer entered before the worker apply released its read fence"),
|
||||
})
|
||||
.await;
|
||||
transaction
|
||||
.await
|
||||
.expect("writer should acquire the server-config locks")
|
||||
.expect("writer should persist the latest config");
|
||||
});
|
||||
|
||||
tokio::time::timeout(REAL_STORE_TEST_TIMEOUT, writer_polled_rx)
|
||||
.await
|
||||
.expect("writer should be polled")
|
||||
.expect("writer poll signal should be delivered");
|
||||
assert!(
|
||||
matches!(writer_entered_rx.try_recv(), Err(tokio::sync::oneshot::error::TryRecvError::Empty)),
|
||||
"a server-config writer must wait until the old worker apply completes"
|
||||
);
|
||||
|
||||
release_worker_tx.send(()).expect("release worker apply");
|
||||
tokio::time::timeout(REAL_STORE_TEST_TIMEOUT, async {
|
||||
reload
|
||||
.await
|
||||
.expect("full reload task should not panic")
|
||||
.expect("full reload should finish");
|
||||
writer.await.expect("writer task should not panic");
|
||||
})
|
||||
.await
|
||||
.expect("reload and writer should finish");
|
||||
|
||||
reload_runtime_config_snapshot_for_context(Some(&fixture.context))
|
||||
.await
|
||||
.expect("a later full reload should publish the latest durable config");
|
||||
let snapshot = fixture
|
||||
.server_snapshot
|
||||
.lock()
|
||||
.expect("server config result lock")
|
||||
.clone()
|
||||
.expect("latest server config should be published");
|
||||
assert_eq!(
|
||||
snapshot
|
||||
.get_value(SCANNER_SUB_SYS, DEFAULT_DELIMITER)
|
||||
.expect("latest scanner config should be present")
|
||||
.get(SCANNER_CYCLE),
|
||||
"71"
|
||||
);
|
||||
assert_eq!(rustfs_scanner::scanner_runtime_config_status().cycle_interval_seconds.value, 71);
|
||||
|
||||
rustfs_scanner::apply_scanner_runtime_config(&ServerConfig::new()).expect("restore scanner runtime defaults");
|
||||
},
|
||||
)
|
||||
.await;
|
||||
}
|
||||
|
||||
#[tokio::test]
|
||||
#[serial_test::serial(storage_class_env)]
|
||||
async fn peer_full_reload_rejects_later_pool_without_publishing() {
|
||||
@@ -1251,11 +1545,9 @@ mod tests {
|
||||
let worker_events = events.clone();
|
||||
|
||||
let err = reload_runtime_config_snapshot_with(
|
||||
async { Ok(ServerConfig::new()) },
|
||||
|config| async { Ok((config, PreparedRuntimeConfig::default())) },
|
||||
move |_config, _prepared| {
|
||||
async move {
|
||||
publish_events.lock().expect("reload event lock").push("publish");
|
||||
Ok(())
|
||||
Ok(ServerConfig::new())
|
||||
},
|
||||
move |_config| async move {
|
||||
let mut events = worker_events.lock().expect("reload event lock");
|
||||
|
||||
@@ -644,12 +644,17 @@ pub(crate) async fn read_admin_config_without_migrate(api: Arc<ECStore>) -> Resu
|
||||
ecstore_config::com::read_config_without_migrate(api).await
|
||||
}
|
||||
|
||||
pub(crate) async fn read_existing_admin_server_config_no_lock(api: Arc<ECStore>) -> Result<rustfs_config::server_config::Config> {
|
||||
ecstore_config::com::read_existing_server_config_no_lock(api).await
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
pub(crate) async fn read_admin_config_without_migrate_no_lock(api: Arc<ECStore>) -> Result<rustfs_config::server_config::Config> {
|
||||
ecstore_config::com::read_config_without_migrate_no_lock(api).await
|
||||
}
|
||||
|
||||
pub(crate) type AdminServerConfigSnapshot = ecstore_config::com::ServerConfigSnapshot;
|
||||
pub(crate) type AdminServerConfigSaveResult = ecstore_config::com::ServerConfigSaveResult;
|
||||
|
||||
pub(crate) async fn save_admin_config(api: Arc<ECStore>, file: &str, data: Vec<u8>) -> Result<()> {
|
||||
ecstore_config::com::save_config(api, file, data).await
|
||||
@@ -690,8 +695,17 @@ pub(crate) async fn save_admin_server_config_snapshot(
|
||||
api: Arc<ECStore>,
|
||||
cfg: &rustfs_config::server_config::Config,
|
||||
snapshot: &AdminServerConfigSnapshot,
|
||||
) -> Result<bool> {
|
||||
ecstore_config::com::save_server_config_snapshot(api, cfg, snapshot).await
|
||||
) -> Result<AdminServerConfigSaveResult> {
|
||||
ecstore_config::com::save_server_config_snapshot_with_generation(api, cfg, snapshot).await
|
||||
}
|
||||
|
||||
pub(crate) async fn with_admin_server_config_read_lock<F, Fut, T>(api: Arc<ECStore>, operation: F) -> Result<T>
|
||||
where
|
||||
F: FnOnce() -> Fut + Send + 'static,
|
||||
Fut: std::future::Future<Output = T> + Send + 'static,
|
||||
T: Send + 'static,
|
||||
{
|
||||
ecstore_config::com::with_server_config_read_lock(api, operation).await
|
||||
}
|
||||
|
||||
pub(crate) fn init_admin_config_defaults() {
|
||||
@@ -793,9 +807,10 @@ pub(crate) mod cluster {
|
||||
pub(crate) mod config {
|
||||
pub(crate) use super::storageclass;
|
||||
pub(crate) use super::{
|
||||
AdminServerConfigSnapshot, RUSTFS_META_BUCKET, STORAGE_CLASS_SUB_SYS, delete_admin_config, init_admin_config_defaults,
|
||||
read_admin_config, read_admin_config_without_migrate, read_admin_server_config_snapshot, save_admin_config,
|
||||
save_admin_server_config_snapshot,
|
||||
AdminServerConfigSaveResult, AdminServerConfigSnapshot, RUSTFS_META_BUCKET, STORAGE_CLASS_SUB_SYS, delete_admin_config,
|
||||
init_admin_config_defaults, read_admin_config, read_admin_config_without_migrate, read_admin_server_config_snapshot,
|
||||
read_existing_admin_server_config_no_lock, save_admin_config, save_admin_server_config_snapshot,
|
||||
with_admin_server_config_read_lock,
|
||||
};
|
||||
#[cfg(test)]
|
||||
pub(crate) use super::{
|
||||
|
||||
@@ -6072,19 +6072,6 @@ impl DefaultObjectUsecase {
|
||||
));
|
||||
}
|
||||
};
|
||||
let has_replacement_metadata = metadata.is_some()
|
||||
|| cache_control.is_some()
|
||||
|| content_disposition.is_some()
|
||||
|| content_encoding.is_some()
|
||||
|| content_language.is_some()
|
||||
|| content_type.is_some()
|
||||
|| expires.is_some();
|
||||
if has_replacement_metadata && !replaces_metadata {
|
||||
return Err(S3Error::with_message(
|
||||
S3ErrorCode::InvalidRequest,
|
||||
"Replacement metadata requires the REPLACE metadata directive".to_string(),
|
||||
));
|
||||
}
|
||||
let replacement_metadata = if replaces_metadata {
|
||||
validate_archive_content_encoding(&key, content_type.as_deref(), content_encoding.as_deref())?;
|
||||
let mut replacement_metadata = metadata.unwrap_or_default();
|
||||
|
||||
+36
-1
@@ -12,9 +12,32 @@
|
||||
// See the License for the specific language governing permissions and
|
||||
// limitations under the License.
|
||||
|
||||
#[cfg(all(feature = "hotpath", feature = "hotpath-alloc"))]
|
||||
use std::alloc::{GlobalAlloc, Layout};
|
||||
|
||||
#[cfg(all(feature = "hotpath", feature = "hotpath-alloc"))]
|
||||
#[derive(Default)]
|
||||
struct DefaultMiMalloc;
|
||||
|
||||
#[cfg(all(feature = "hotpath", feature = "hotpath-alloc"))]
|
||||
// SAFETY: allocation and deallocation are forwarded unchanged to MiMalloc, so
|
||||
// MiMalloc's GlobalAlloc guarantees apply to every returned pointer and layout.
|
||||
#[allow(unsafe_code)]
|
||||
unsafe impl GlobalAlloc for DefaultMiMalloc {
|
||||
unsafe fn alloc(&self, layout: Layout) -> *mut u8 {
|
||||
// SAFETY: the caller upholds GlobalAlloc's contract for layout.
|
||||
unsafe { mimalloc::MiMalloc.alloc(layout) }
|
||||
}
|
||||
|
||||
unsafe fn dealloc(&self, ptr: *mut u8, layout: Layout) {
|
||||
// SAFETY: ptr and layout came from this allocator and are forwarded unchanged.
|
||||
unsafe { mimalloc::MiMalloc.dealloc(ptr, layout) }
|
||||
}
|
||||
}
|
||||
|
||||
#[cfg(all(feature = "hotpath", feature = "hotpath-alloc"))]
|
||||
#[global_allocator]
|
||||
static GLOBAL: hotpath::CountingAllocator = hotpath::CountingAllocator::new();
|
||||
static GLOBAL: hotpath::CountingAllocator<DefaultMiMalloc> = hotpath::CountingAllocator::new();
|
||||
|
||||
#[cfg(not(all(feature = "hotpath", feature = "hotpath-alloc")))]
|
||||
#[global_allocator]
|
||||
@@ -25,3 +48,15 @@ fn main() {
|
||||
|
||||
rustfs::startup_entrypoint::run_process();
|
||||
}
|
||||
|
||||
#[cfg(all(test, feature = "hotpath", feature = "hotpath-alloc"))]
|
||||
mod tests {
|
||||
#[test]
|
||||
#[allow(unsafe_code)]
|
||||
fn hotpath_allocator_uses_mimalloc() {
|
||||
let allocation = Box::new([0_u8; 64]);
|
||||
|
||||
// SAFETY: the live Box pointer is valid to inspect for heap ownership.
|
||||
assert!(unsafe { libmimalloc_sys::mi_is_in_heap_region(allocation.as_ptr().cast()) });
|
||||
}
|
||||
}
|
||||
|
||||
Executable
+83
@@ -0,0 +1,83 @@
|
||||
#!/usr/bin/env bash
|
||||
# ci.yml's pull_request paths-ignore and ci-docs-only.yml's paths must be equal.
|
||||
#
|
||||
# ci-docs-only.yml exists to report the required checks for pull requests that
|
||||
# ci.yml skips. The two lists are the complement of each other, so any drift
|
||||
# breaks one of two ways, both silent:
|
||||
#
|
||||
# - an entry only in ci.yml's paths-ignore: a PR touching only those files
|
||||
# triggers neither workflow, nobody reports "Test and Lint" or "Quick
|
||||
# Checks", and the PR waits on a required check forever;
|
||||
# - an entry only in ci-docs-only.yml's paths: both workflows run, which is
|
||||
# merely wasteful — but it also means the lists no longer describe the same
|
||||
# intent, and the next edit is made against a wrong assumption.
|
||||
#
|
||||
# The push paths-ignore in ci.yml is deliberately NOT compared: no required
|
||||
# check is reported for push events, so it does not have to pair with anything.
|
||||
#
|
||||
# Also asserts ci-docs-only.yml still declares both companion job names, since a
|
||||
# rename there produces exactly the permanent-pending failure above.
|
||||
#
|
||||
# Usage: scripts/check_ci_paths_sync.sh
|
||||
set -euo pipefail
|
||||
|
||||
cd "$(dirname "$0")/.."
|
||||
|
||||
CI=".github/workflows/ci.yml"
|
||||
DOCS=".github/workflows/ci-docs-only.yml"
|
||||
|
||||
# Print the quoted list items that follow $2 within the block introduced by $1.
|
||||
# Both files keep these as a flat list of quoted scalars, so no YAML parser is
|
||||
# needed and the script stays dependency-free like its check_* siblings.
|
||||
extract() {
|
||||
local file="$1" event="$2" key="$3"
|
||||
awk -v event="$event" -v key="$key" '
|
||||
$0 ~ "^ " event ":[[:space:]]*$" { in_event = 1; next }
|
||||
in_event && /^ [a-z_]+:[[:space:]]*$/ { in_event = 0 }
|
||||
in_event && $0 ~ "^ " key ":[[:space:]]*$" { in_list = 1; next }
|
||||
in_list {
|
||||
if ($0 ~ /^ - /) {
|
||||
item = $0
|
||||
sub(/^ - /, "", item)
|
||||
gsub(/^"|"$/, "", item)
|
||||
print item
|
||||
next
|
||||
}
|
||||
if ($0 !~ /^[[:space:]]*#/ && $0 !~ /^[[:space:]]*$/) in_list = 0
|
||||
}
|
||||
' "$file" | sort
|
||||
}
|
||||
|
||||
ci_list="$(extract "$CI" "pull_request" "paths-ignore")"
|
||||
docs_list="$(extract "$DOCS" "pull_request" "paths")"
|
||||
|
||||
if [ -z "$ci_list" ] || [ -z "$docs_list" ]; then
|
||||
echo "ERROR: could not read one of the path lists — did the file structure change?" >&2
|
||||
echo " $CI pull_request.paths-ignore: $(printf '%s' "$ci_list" | grep -c . || true) entries" >&2
|
||||
echo " $DOCS pull_request.paths: $(printf '%s' "$docs_list" | grep -c . || true) entries" >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
status=0
|
||||
|
||||
if ! diff_out="$(diff <(printf '%s\n' "$ci_list") <(printf '%s\n' "$docs_list"))"; then
|
||||
echo "ERROR: $CI pull_request paths-ignore and $DOCS paths have drifted." >&2
|
||||
echo " '<' is only in $CI, '>' is only in $DOCS:" >&2
|
||||
printf '%s\n' "$diff_out" | sed 's/^/ /' >&2
|
||||
status=1
|
||||
fi
|
||||
|
||||
for job_name in "Test and Lint" "Quick Checks"; do
|
||||
if ! grep -q "name: ${job_name}\$" "$DOCS"; then
|
||||
echo "ERROR: $DOCS no longer declares a job named '${job_name}'." >&2
|
||||
echo " It is a required status check; without a companion job here, a" >&2
|
||||
echo " docs-only PR waits on it forever." >&2
|
||||
status=1
|
||||
fi
|
||||
done
|
||||
|
||||
if [ "$status" -ne 0 ]; then
|
||||
exit 1
|
||||
fi
|
||||
|
||||
echo "OK: ci.yml and ci-docs-only.yml path lists agree ($(printf '%s\n' "$ci_list" | wc -l | tr -d ' ') entries)"
|
||||
Executable
+41
@@ -0,0 +1,41 @@
|
||||
#!/usr/bin/env bash
|
||||
# The io_uring CI lane runs `cargo test -p rustfs-ecstore --lib uring_`.
|
||||
#
|
||||
# `--lib` is there to avoid compiling and linking the 7 integration binaries
|
||||
# under crates/ecstore/tests/, which selected zero tests for this filter. That is
|
||||
# only safe while it stays true: add a matching test under tests/ and `--lib`
|
||||
# would skip it silently, with CI staying green — the failure mode is invisible,
|
||||
# so it is asserted here instead.
|
||||
#
|
||||
# The pattern must use the same `uring_` substring semantics as libtest, which
|
||||
# also matches names containing `during_`. Narrowing it to `io_uring` would make
|
||||
# this guard disagree with the filter it is protecting.
|
||||
#
|
||||
# Usage: scripts/check_uring_lane_lib_only.sh
|
||||
set -euo pipefail
|
||||
|
||||
cd "$(dirname "$0")/.."
|
||||
|
||||
TESTS_DIR="crates/ecstore/tests"
|
||||
|
||||
if [ ! -d "$TESTS_DIR" ]; then
|
||||
echo "OK: $TESTS_DIR does not exist; nothing for --lib to miss"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
hits="$(grep -rEn '^[[:space:]]*(pub[[:space:]]+)?(async[[:space:]]+)?fn[[:space:]]+[A-Za-z0-9_]*uring_' \
|
||||
"$TESTS_DIR" || true)"
|
||||
|
||||
if [ -n "$hits" ]; then
|
||||
echo "ERROR: test functions matching the 'uring_' substring found under $TESTS_DIR:" >&2
|
||||
printf '%s\n' "$hits" | sed 's/^/ /' >&2
|
||||
echo "" >&2
|
||||
echo "The io_uring lane in .github/workflows/ci.yml uses --lib, so these would" >&2
|
||||
echo "never run and CI would stay green. Pick one:" >&2
|
||||
echo " - move them into a #[cfg(test)] mod under crates/ecstore/src, or" >&2
|
||||
echo " - drop --lib from that step (and pay the link cost for all 7 binaries), or" >&2
|
||||
echo " - add an explicit --test <name> for the binary and update this script." >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
echo "OK: no 'uring_'-matching test functions under $TESTS_DIR; --lib is safe"
|
||||
Executable
+107
@@ -0,0 +1,107 @@
|
||||
#!/usr/bin/env bash
|
||||
# Background resource sampler for the Test and Lint lane (issue #5394).
|
||||
#
|
||||
# The post-mortem `pgrep` in ci.yml runs only after `timeout` has already TERM'd
|
||||
# the whole cargo process group, so it can never name a wedged process. This
|
||||
# samples every 60s instead; the last samples before a timeout show what was
|
||||
# stuck — rustc, the linker, a build script, memory pressure.
|
||||
#
|
||||
# WHY THE CGROUP READINGS MATTER
|
||||
#
|
||||
# The runners are ARC (Actions Runner Controller) pods, so /proc/loadavg,
|
||||
# /proc/pressure/*, `free` and `df` are all *node*-level: they include every
|
||||
# other runner pod scheduled on the same Kubernetes node. A sample showing high
|
||||
# I/O pressure says nothing about whether *this* job caused it. Measured
|
||||
# evidence: one sample reported loadavg 6.67 with 832 threads node-wide while
|
||||
# `ps` inside the pod showed about 10 processes.
|
||||
#
|
||||
# Only the /sys/fs/cgroup/* values and `ps` are attributable to this job, so
|
||||
# those are what any CARGO_BUILD_JOBS experiment must be judged on. The
|
||||
# node-level readings are kept — co-tenancy is itself a real cause of stalls,
|
||||
# and a 9m57s `git checkout` was traced to it — but labelled NODE-LEVEL so
|
||||
# nobody reads them as this pod's own load.
|
||||
#
|
||||
# Usage:
|
||||
# scripts/ci/resource_sampler.sh start <phase> # e.g. clippy, nextest, doctest
|
||||
# scripts/ci/resource_sampler.sh stop
|
||||
#
|
||||
# Safe to call `stop` when nothing is running, and safe to `start` a new phase
|
||||
# while another is sampling — the previous one is stopped first. Never fails the
|
||||
# calling step: this is diagnostics, not a gate.
|
||||
set -uo pipefail
|
||||
|
||||
# GNU date on the runners; the fallback keeps local smoke tests on macOS quiet.
|
||||
now() { date --utc --iso-8601=seconds 2>/dev/null || date -u +%Y-%m-%dT%H:%M:%S+00:00; }
|
||||
|
||||
LOG_DIR="artifacts/test-and-lint"
|
||||
LOG_FILE="${LOG_DIR}/sampler.log"
|
||||
PID_FILE="${LOG_DIR}/.sampler.pid"
|
||||
INTERVAL="${SAMPLER_INTERVAL_SECS:-60}"
|
||||
|
||||
stop_sampler() {
|
||||
[ -f "$PID_FILE" ] || return 0
|
||||
local pid
|
||||
pid="$(cat "$PID_FILE" 2>/dev/null || true)"
|
||||
if [ -n "${pid:-}" ]; then
|
||||
kill "$pid" 2>/dev/null || true
|
||||
fi
|
||||
rm -f "$PID_FILE"
|
||||
}
|
||||
|
||||
sample_once() {
|
||||
echo "=== $(now)"
|
||||
|
||||
# Attributable to this pod: cgroup v2 accounting and our own process table.
|
||||
echo "--- POD cpu/mem limits"
|
||||
grep -H . /sys/fs/cgroup/cpu.max /sys/fs/cgroup/memory.max \
|
||||
/sys/fs/cgroup/memory.peak /sys/fs/cgroup/memory.current 2>/dev/null || true
|
||||
echo "--- POD pressure (cgroup v2, this pod only)"
|
||||
grep -H . /sys/fs/cgroup/cpu.pressure /sys/fs/cgroup/io.pressure \
|
||||
/sys/fs/cgroup/memory.pressure 2>/dev/null || true
|
||||
echo "--- POD nproc"; nproc 2>/dev/null || true
|
||||
|
||||
# NODE-LEVEL: shared with every other runner pod on this Kubernetes node.
|
||||
# Useful for spotting co-tenancy, useless for attributing load to this job.
|
||||
echo "--- NODE-LEVEL load (shared with co-tenant runners)"; cat /proc/loadavg 2>/dev/null || true
|
||||
echo "--- NODE-LEVEL psi (shared with co-tenant runners)"
|
||||
grep -H . /proc/pressure/* 2>/dev/null || true
|
||||
echo "--- NODE-LEVEL mem (shared with co-tenant runners)"; free -m 2>/dev/null || true
|
||||
echo "--- NODE-LEVEL disk (shared with co-tenant runners)"
|
||||
df -h / /home/runner 2>/dev/null || df -h / 2>/dev/null || true
|
||||
|
||||
echo "--- top-rss"
|
||||
ps -eo pid,ppid,stat,etime,rss,pcpu,args --sort=-rss 2>/dev/null | head -15 || true
|
||||
echo "--- build/test processes"
|
||||
ps -eo pid,ppid,stat,etime,rss,pcpu,args 2>/dev/null \
|
||||
| grep -E '[c]argo|[r]ustc|[n]extest|[c]ollect2|rust-ll[d]|[b]uild-script|deps[/]' || true
|
||||
echo "--- d-state (uninterruptible IO)"
|
||||
ps -eo pid,stat,etime,args 2>/dev/null | awk 'NR > 1 && $2 ~ /D/' || true
|
||||
echo
|
||||
}
|
||||
|
||||
case "${1:-}" in
|
||||
start)
|
||||
phase="${2:-unknown}"
|
||||
mkdir -p "$LOG_DIR"
|
||||
stop_sampler
|
||||
{
|
||||
echo "--- phase=${phase} started_at=$(now) runner=${RUNNER_NAME:-unknown}"
|
||||
} >> "$LOG_FILE" 2>&1 || true
|
||||
(
|
||||
while true; do
|
||||
sample_once >> "$LOG_FILE" 2>&1 || true
|
||||
sleep "$INTERVAL"
|
||||
done
|
||||
) &
|
||||
echo $! > "$PID_FILE"
|
||||
;;
|
||||
stop)
|
||||
stop_sampler
|
||||
;;
|
||||
*)
|
||||
echo "usage: $0 start <phase> | stop" >&2
|
||||
exit 2
|
||||
;;
|
||||
esac
|
||||
|
||||
exit 0
|
||||
+13
-1
@@ -40,7 +40,19 @@ export RUST_BACKTRACE=full
|
||||
"$BIN" --address "127.0.0.1:${RUSTFS_TEST_PORT}" "$VOLUME" > "${RUSTFS_TEST_LOG}" 2>&1 &
|
||||
RUSTFS_PID=$!
|
||||
|
||||
sleep 10
|
||||
for _ in {1..60}; do
|
||||
if curl --noproxy '*' --silent --fail "http://127.0.0.1:${RUSTFS_TEST_PORT}/health/ready" >/dev/null; then
|
||||
break
|
||||
fi
|
||||
|
||||
if ! kill -0 "${RUSTFS_PID}" 2>/dev/null; then
|
||||
wait "${RUSTFS_PID}"
|
||||
fi
|
||||
|
||||
sleep 1
|
||||
done
|
||||
|
||||
curl --noproxy '*' --silent --show-error --fail "http://127.0.0.1:${RUSTFS_TEST_PORT}/health/ready" >/dev/null
|
||||
|
||||
export AWS_ACCESS_KEY_ID="${RUSTFS_ACCESS_KEY:-rustfsadmin}"
|
||||
export AWS_SECRET_ACCESS_KEY="${RUSTFS_SECRET_KEY:-rustfsadmin}"
|
||||
|
||||
Executable
+428
@@ -0,0 +1,428 @@
|
||||
#!/usr/bin/env bash
|
||||
# Formal Linux / production-cluster ABBA runner for the hotpath warp matrix.
|
||||
#
|
||||
# This script is intentionally a thin orchestrator around the existing
|
||||
# run_object_batch_bench_enhanced.sh load driver and hotpath_warp_ab_gate.sh
|
||||
# relative-budget gate. It runs each durability/workload cell as:
|
||||
#
|
||||
# A1 baseline -> B1 candidate -> B2 candidate -> A2 baseline
|
||||
#
|
||||
# Candidate legs are compared against A1. The final A2 leg is also compared
|
||||
# against A1 to quantify baseline drift separately from candidate deltas.
|
||||
|
||||
set -euo pipefail
|
||||
|
||||
PROJECT_ROOT="$(git rev-parse --show-toplevel)"
|
||||
ENHANCED_BENCH="${PROJECT_ROOT}/scripts/run_object_batch_bench_enhanced.sh"
|
||||
GATE="${PROJECT_ROOT}/scripts/hotpath_warp_ab_gate.sh"
|
||||
|
||||
BASELINE_BIN=""
|
||||
CANDIDATE_BIN=""
|
||||
ENDPOINT=""
|
||||
DEPLOY_HOOK=""
|
||||
HEALTH_PATH="/health"
|
||||
ADDRESS="127.0.0.1:9000"
|
||||
DATA_ROOT="/tmp/rustfs-hotpath-abba"
|
||||
DISKS=4
|
||||
ACCESS_KEY="rustfsadmin"
|
||||
SECRET_KEY="rustfsadmin"
|
||||
REGION="us-east-1"
|
||||
WARP_BIN="warp"
|
||||
CONCURRENCY=8
|
||||
DURATION="60s"
|
||||
ROUNDS=3
|
||||
COOLDOWN_SECS=20
|
||||
HEALTH_TIMEOUT_SECS=180
|
||||
FAIL_PCT=10
|
||||
WARN_PCT=5
|
||||
ALLOW_REGRESSION=false
|
||||
EXEMPTION_REASON="deliberate correctness tradeoff"
|
||||
OUT_DIR="${PROJECT_ROOT}/target/hotpath-abba/$(date -u +%Y%m%dT%H%M%SZ 2>/dev/null || echo run)"
|
||||
DRY_RUN=false
|
||||
|
||||
WORKLOADS=(
|
||||
"put-4kib|put|4KiB"
|
||||
"put-4mib|put|4MiB"
|
||||
"get-4kib|get|4KiB"
|
||||
"get-4mib|get|4MiB"
|
||||
"get-10mib|get|10MiB"
|
||||
"mixed-256k|mixed|256KiB"
|
||||
)
|
||||
DRIVE_SYNC_MATRIX=("sync-on|true" "sync-off|false")
|
||||
|
||||
usage() {
|
||||
cat <<'USAGE'
|
||||
Usage: scripts/run_hotpath_warp_abba.sh --baseline-bin <path> --candidate-bin <path> [options]
|
||||
|
||||
Formal ABBA mode for Linux runners or production-like clusters. The schedule is
|
||||
A1 baseline -> B1 candidate -> B2 candidate -> A2 baseline for every workload
|
||||
and drive-sync cell.
|
||||
|
||||
Required:
|
||||
--baseline-bin <path> Baseline RustFS binary.
|
||||
--candidate-bin <path> Candidate RustFS binary.
|
||||
|
||||
Local Linux runner mode:
|
||||
--address <host:port> Local RustFS address (default 127.0.0.1:9000).
|
||||
--disks <n> Throwaway local disks per node (default 4).
|
||||
--data-root <path> Local disk root (default /tmp/rustfs-hotpath-abba).
|
||||
|
||||
Production / cluster mode:
|
||||
--endpoint <host:port> Existing cluster endpoint. Enables external mode.
|
||||
--deploy-hook <cmd> Command run before each ABBA leg. It receives:
|
||||
HOTPATH_ABBA_LEG=A1|B1|B2|A2
|
||||
HOTPATH_ABBA_PHASE=baseline|candidate
|
||||
HOTPATH_ABBA_BINARY=<baseline/candidate binary>
|
||||
HOTPATH_ABBA_DRIVE_SYNC=true|false
|
||||
--health-path <path> Readiness path (default /health).
|
||||
|
||||
Benchmark:
|
||||
--duration <dur> warp duration per cell (default 60s).
|
||||
--rounds <n> rounds per cell; must be >= 3 (default 3).
|
||||
--cooldown <n> cooldown seconds between rounds/sizes (default 20).
|
||||
--concurrency <n> warp concurrency (default 8).
|
||||
--warp-bin <path> warp binary (default warp).
|
||||
|
||||
Credentials:
|
||||
--access-key <value> S3 access key (default rustfsadmin).
|
||||
--secret-key <value> S3 secret key (default rustfsadmin).
|
||||
--region <value> S3 region (default us-east-1).
|
||||
|
||||
Gate:
|
||||
--fail-pct <n> Regression budget that fails gate (default 10).
|
||||
--warn-pct <n> Regression budget that warns (default 5).
|
||||
--allow-regression Downgrade candidate gate FAIL to WARN.
|
||||
--exemption-reason <s> Reason recorded when allow-regression is used.
|
||||
|
||||
Output:
|
||||
--out-dir <path> Output dir (default target/hotpath-abba/<ts>).
|
||||
--dry-run Print commands without starting servers or warp.
|
||||
-h, --help
|
||||
|
||||
Outputs:
|
||||
<out-dir>/abba_schedule.csv
|
||||
<out-dir>/candidate_gate.md
|
||||
<out-dir>/baseline_drift_gate.md
|
||||
<out-dir>/summary.md
|
||||
<out-dir>/<workload>/<sync>/<leg>/{median_summary.csv,baseline_compare.csv}
|
||||
USAGE
|
||||
}
|
||||
|
||||
die() {
|
||||
echo "error: $*" >&2
|
||||
exit 2
|
||||
}
|
||||
|
||||
log() {
|
||||
printf '[hotpath-abba] %s\n' "$*" >&2
|
||||
}
|
||||
|
||||
run() {
|
||||
if [[ "$DRY_RUN" == "true" ]]; then
|
||||
{ printf 'DRY-RUN:'; printf ' %q' "$@"; printf '\n'; } >&2
|
||||
return 0
|
||||
fi
|
||||
"$@"
|
||||
}
|
||||
|
||||
validate_positive_int() {
|
||||
local value="$1" name="$2"
|
||||
[[ "$value" =~ ^[0-9]+$ && "$value" -gt 0 ]] || die "$name must be a positive integer"
|
||||
}
|
||||
|
||||
while [[ $# -gt 0 ]]; do
|
||||
case "$1" in
|
||||
--baseline-bin) BASELINE_BIN="$2"; shift 2 ;;
|
||||
--candidate-bin) CANDIDATE_BIN="$2"; shift 2 ;;
|
||||
--endpoint) ENDPOINT="$2"; shift 2 ;;
|
||||
--deploy-hook) DEPLOY_HOOK="$2"; shift 2 ;;
|
||||
--health-path) HEALTH_PATH="$2"; shift 2 ;;
|
||||
--address) ADDRESS="$2"; shift 2 ;;
|
||||
--data-root) DATA_ROOT="$2"; shift 2 ;;
|
||||
--disks) DISKS="$2"; shift 2 ;;
|
||||
--access-key) ACCESS_KEY="$2"; shift 2 ;;
|
||||
--secret-key) SECRET_KEY="$2"; shift 2 ;;
|
||||
--region) REGION="$2"; shift 2 ;;
|
||||
--warp-bin) WARP_BIN="$2"; shift 2 ;;
|
||||
--concurrency) CONCURRENCY="$2"; shift 2 ;;
|
||||
--duration) DURATION="$2"; shift 2 ;;
|
||||
--rounds) ROUNDS="$2"; shift 2 ;;
|
||||
--cooldown) COOLDOWN_SECS="$2"; shift 2 ;;
|
||||
--health-timeout) HEALTH_TIMEOUT_SECS="$2"; shift 2 ;;
|
||||
--fail-pct) FAIL_PCT="$2"; shift 2 ;;
|
||||
--warn-pct) WARN_PCT="$2"; shift 2 ;;
|
||||
--allow-regression) ALLOW_REGRESSION=true; shift ;;
|
||||
--exemption-reason) EXEMPTION_REASON="$2"; shift 2 ;;
|
||||
--out-dir) OUT_DIR="$2"; shift 2 ;;
|
||||
--dry-run) DRY_RUN=true; shift ;;
|
||||
-h|--help) usage; exit 0 ;;
|
||||
*) die "unknown argument: $1" ;;
|
||||
esac
|
||||
done
|
||||
|
||||
validate_positive_int "$DISKS" "--disks"
|
||||
validate_positive_int "$CONCURRENCY" "--concurrency"
|
||||
validate_positive_int "$ROUNDS" "--rounds"
|
||||
validate_positive_int "$COOLDOWN_SECS" "--cooldown"
|
||||
validate_positive_int "$HEALTH_TIMEOUT_SECS" "--health-timeout"
|
||||
validate_positive_int "$FAIL_PCT" "--fail-pct"
|
||||
validate_positive_int "$WARN_PCT" "--warn-pct"
|
||||
[[ "$ROUNDS" -ge 3 ]] || die "--rounds must be >= 3 for formal ABBA evidence"
|
||||
[[ -n "$BASELINE_BIN" ]] || die "--baseline-bin is required"
|
||||
[[ -n "$CANDIDATE_BIN" ]] || die "--candidate-bin is required"
|
||||
[[ "$DRY_RUN" == "true" || -x "$BASELINE_BIN" ]] || die "baseline binary is not executable: $BASELINE_BIN"
|
||||
[[ "$DRY_RUN" == "true" || -x "$CANDIDATE_BIN" ]] || die "candidate binary is not executable: $CANDIDATE_BIN"
|
||||
[[ -x "$ENHANCED_BENCH" ]] || die "missing load driver: $ENHANCED_BENCH"
|
||||
[[ -x "$GATE" ]] || die "missing gate: $GATE"
|
||||
if [[ "$DRY_RUN" != "true" ]] && ! command -v "$WARP_BIN" >/dev/null 2>&1; then
|
||||
die "warp not found on PATH; install warp or pass --warp-bin"
|
||||
fi
|
||||
|
||||
EXTERNAL=false
|
||||
if [[ -n "$ENDPOINT" ]]; then
|
||||
EXTERNAL=true
|
||||
ADDRESS="$ENDPOINT"
|
||||
[[ -n "$DEPLOY_HOOK" ]] || log "warning: external mode without --deploy-hook; binaries must be swapped out of band"
|
||||
fi
|
||||
|
||||
mkdir -p "$OUT_DIR"
|
||||
SERVER_LOG_DIR="$OUT_DIR/server-logs"
|
||||
run mkdir -p "$SERVER_LOG_DIR"
|
||||
|
||||
SERVER_PID=""
|
||||
SERVER_LOG=""
|
||||
tear_down() {
|
||||
[[ "$EXTERNAL" == "true" ]] && return 0
|
||||
[[ -n "$SERVER_PID" ]] || return 0
|
||||
run kill "$SERVER_PID" 2>/dev/null || true
|
||||
SERVER_PID=""
|
||||
}
|
||||
trap tear_down EXIT INT TERM
|
||||
|
||||
dump_server_log() {
|
||||
[[ -n "$SERVER_LOG" && -f "$SERVER_LOG" ]] || return 0
|
||||
echo "----- last 80 lines of $SERVER_LOG -----" >&2
|
||||
tail -n 80 "$SERVER_LOG" >&2 || true
|
||||
echo "----------------------------------------" >&2
|
||||
}
|
||||
|
||||
wait_health() {
|
||||
[[ "$DRY_RUN" == "true" ]] && return 0
|
||||
local i
|
||||
for ((i = 0; i < HEALTH_TIMEOUT_SECS; i++)); do
|
||||
if [[ "$EXTERNAL" != "true" && -n "$SERVER_PID" ]] && ! kill -0 "$SERVER_PID" 2>/dev/null; then
|
||||
echo "error: rustfs server (pid $SERVER_PID) exited before becoming healthy after ${i}s" >&2
|
||||
dump_server_log
|
||||
return 1
|
||||
fi
|
||||
if curl -fsS "http://${ADDRESS}${HEALTH_PATH}" >/dev/null 2>&1; then
|
||||
return 0
|
||||
fi
|
||||
sleep 1
|
||||
done
|
||||
echo "error: endpoint http://${ADDRESS}${HEALTH_PATH} did not become healthy within ${HEALTH_TIMEOUT_SECS}s" >&2
|
||||
dump_server_log
|
||||
return 1
|
||||
}
|
||||
|
||||
binary_for_leg() {
|
||||
case "$1" in
|
||||
A1|A2) echo "$BASELINE_BIN" ;;
|
||||
B1|B2) echo "$CANDIDATE_BIN" ;;
|
||||
*) die "unknown ABBA leg: $1" ;;
|
||||
esac
|
||||
}
|
||||
|
||||
phase_for_leg() {
|
||||
case "$1" in
|
||||
A1|A2) echo "baseline" ;;
|
||||
B1|B2) echo "candidate" ;;
|
||||
*) die "unknown ABBA leg: $1" ;;
|
||||
esac
|
||||
}
|
||||
|
||||
bring_up() {
|
||||
local leg="$1" drive_sync="$2"
|
||||
local phase bin
|
||||
phase="$(phase_for_leg "$leg")"
|
||||
bin="$(binary_for_leg "$leg")"
|
||||
|
||||
if [[ "$EXTERNAL" == "true" ]]; then
|
||||
if [[ -n "$DEPLOY_HOOK" ]]; then
|
||||
log "deploy hook: leg=$leg phase=$phase drive_sync=$drive_sync"
|
||||
HOTPATH_ABBA_LEG="$leg" HOTPATH_ABBA_PHASE="$phase" HOTPATH_ABBA_BINARY="$bin" HOTPATH_ABBA_DRIVE_SYNC="$drive_sync" \
|
||||
run bash -c "$DEPLOY_HOOK"
|
||||
fi
|
||||
wait_health
|
||||
return 0
|
||||
fi
|
||||
|
||||
local node_dir="$DATA_ROOT/$leg-sync-$drive_sync"
|
||||
local disks=() d
|
||||
for ((d = 1; d <= DISKS; d++)); do
|
||||
disks+=("$node_dir/d$d")
|
||||
done
|
||||
run mkdir -p "${disks[@]}"
|
||||
SERVER_LOG="$SERVER_LOG_DIR/$leg-sync-$drive_sync.log"
|
||||
|
||||
if [[ "$DRY_RUN" == "true" ]]; then
|
||||
{ printf 'DRY-RUN: RUSTFS_DRIVE_SYNC_ENABLE=%s %q server' "$drive_sync" "$bin"
|
||||
printf ' %q' "${disks[@]}"; printf '\n'; } >&2
|
||||
SERVER_PID="dry-run"
|
||||
return 0
|
||||
fi
|
||||
|
||||
cat >"$SERVER_LOG_DIR/$leg-sync-$drive_sync.env" <<EOF
|
||||
leg=$leg
|
||||
phase=$phase
|
||||
drive_sync=$drive_sync
|
||||
binary=$bin
|
||||
address=$ADDRESS
|
||||
disks=${disks[*]}
|
||||
health_url=http://${ADDRESS}${HEALTH_PATH}
|
||||
health_timeout_secs=$HEALTH_TIMEOUT_SECS
|
||||
uname=$(uname -a 2>/dev/null || echo unknown)
|
||||
warp_version=$("$WARP_BIN" --version 2>/dev/null | head -n1 || echo unknown)
|
||||
EOF
|
||||
RUSTFS_UNSAFE_BYPASS_DISK_CHECK=true \
|
||||
RUSTFS_ADDRESS="$ADDRESS" \
|
||||
RUSTFS_ACCESS_KEY="$ACCESS_KEY" \
|
||||
RUSTFS_SECRET_KEY="$SECRET_KEY" \
|
||||
RUSTFS_REGION="$REGION" \
|
||||
RUSTFS_CONSOLE_ENABLE=false \
|
||||
RUSTFS_DRIVE_SYNC_ENABLE="$drive_sync" \
|
||||
"$bin" server "${disks[@]}" >"$SERVER_LOG" 2>&1 &
|
||||
SERVER_PID=$!
|
||||
wait_health
|
||||
}
|
||||
|
||||
measure() {
|
||||
local leg="$1" workload="$2" mode="$3" size="$4" sync_label="$5" baseline_csv="${6:-}"
|
||||
local cell="$OUT_DIR/$workload/$sync_label/$leg"
|
||||
local args=(
|
||||
--tool warp --warp-bin "$WARP_BIN" --warp-mode "$mode"
|
||||
--endpoint "$ADDRESS" --access-key "$ACCESS_KEY" --secret-key "$SECRET_KEY"
|
||||
--region "$REGION" --sizes "$size" --concurrency "$CONCURRENCY"
|
||||
--duration "$DURATION" --rounds "$ROUNDS" --cooldown-secs "$COOLDOWN_SECS"
|
||||
--out-dir "$cell"
|
||||
)
|
||||
[[ -n "$baseline_csv" ]] && args+=(--baseline-csv "$baseline_csv")
|
||||
run "$ENHANCED_BENCH" "${args[@]}" >&2
|
||||
echo "$cell"
|
||||
}
|
||||
|
||||
write_schedule_header() {
|
||||
echo "sync_label,drive_sync,workload,mode,size,leg,phase,binary,out_dir" >"$OUT_DIR/abba_schedule.csv"
|
||||
}
|
||||
|
||||
append_schedule() {
|
||||
local sync_label="$1" drive_sync="$2" workload="$3" mode="$4" size="$5" leg="$6"
|
||||
local phase bin
|
||||
phase="$(phase_for_leg "$leg")"
|
||||
bin="$(binary_for_leg "$leg")"
|
||||
echo "$sync_label,$drive_sync,$workload,$mode,$size,$leg,$phase,$bin,$OUT_DIR/$workload/$sync_label/$leg" >>"$OUT_DIR/abba_schedule.csv"
|
||||
}
|
||||
|
||||
write_manifest() {
|
||||
cat >"$OUT_DIR/manifest.env" <<EOF
|
||||
generated_at_utc=$(date -u +%Y-%m-%dT%H:%M:%SZ 2>/dev/null || echo unknown)
|
||||
runner=$(uname -srm 2>/dev/null || echo unknown)
|
||||
schedule=ABBA
|
||||
rounds=$ROUNDS
|
||||
duration=$DURATION
|
||||
cooldown_secs=$COOLDOWN_SECS
|
||||
concurrency=$CONCURRENCY
|
||||
baseline_bin=$BASELINE_BIN
|
||||
candidate_bin=$CANDIDATE_BIN
|
||||
external=$EXTERNAL
|
||||
endpoint=$ADDRESS
|
||||
warp_version=$("$WARP_BIN" --version 2>/dev/null | head -n1 || echo unknown)
|
||||
EOF
|
||||
}
|
||||
|
||||
declare -a CANDIDATE_COMPARE_CSVS=()
|
||||
declare -a DRIFT_COMPARE_CSVS=()
|
||||
|
||||
write_manifest
|
||||
write_schedule_header
|
||||
|
||||
for ds_spec in "${DRIVE_SYNC_MATRIX[@]}"; do
|
||||
IFS='|' read -r sync_label drive_sync <<<"$ds_spec"
|
||||
|
||||
for leg in A1 B1 B2 A2; do
|
||||
log "=== $sync_label leg $leg ($(phase_for_leg "$leg")) ==="
|
||||
bring_up "$leg" "$drive_sync"
|
||||
|
||||
for wl_spec in "${WORKLOADS[@]}"; do
|
||||
IFS='|' read -r workload mode size <<<"$wl_spec"
|
||||
append_schedule "$sync_label" "$drive_sync" "$workload" "$mode" "$size" "$leg"
|
||||
|
||||
baseline_csv=""
|
||||
if [[ "$leg" != "A1" ]]; then
|
||||
baseline_csv="$OUT_DIR/$workload/$sync_label/A1/median_summary.csv"
|
||||
fi
|
||||
|
||||
cell="$(measure "$leg" "$workload" "$mode" "$size" "$sync_label" "$baseline_csv")"
|
||||
case "$leg" in
|
||||
B1|B2) CANDIDATE_COMPARE_CSVS+=("$cell/baseline_compare.csv") ;;
|
||||
A2) DRIFT_COMPARE_CSVS+=("$cell/baseline_compare.csv") ;;
|
||||
esac
|
||||
done
|
||||
|
||||
tear_down
|
||||
done
|
||||
done
|
||||
|
||||
gate_args=(--fail-pct "$FAIL_PCT" --warn-pct "$WARN_PCT" --markdown "$OUT_DIR/candidate_gate.md")
|
||||
for csv in "${CANDIDATE_COMPARE_CSVS[@]}"; do
|
||||
gate_args+=(--compare-csv "$csv")
|
||||
done
|
||||
[[ "$ALLOW_REGRESSION" == "true" ]] && gate_args+=(--allow-regression --exemption-reason "$EXEMPTION_REASON")
|
||||
|
||||
drift_gate_args=(--fail-pct "$FAIL_PCT" --warn-pct "$WARN_PCT" --markdown "$OUT_DIR/baseline_drift_gate.md")
|
||||
for csv in "${DRIFT_COMPARE_CSVS[@]}"; do
|
||||
drift_gate_args+=(--compare-csv "$csv")
|
||||
done
|
||||
|
||||
if [[ "$DRY_RUN" == "true" ]]; then
|
||||
log "dry-run complete; candidate compare CSVs=${#CANDIDATE_COMPARE_CSVS[@]} baseline drift CSVs=${#DRIFT_COMPARE_CSVS[@]}"
|
||||
{ printf 'DRY-RUN:'; printf ' %q' "$GATE" "${gate_args[@]}"; printf '\n'; } >&2
|
||||
{ printf 'DRY-RUN:'; printf ' %q' "$GATE" "${drift_gate_args[@]}"; printf '\n'; } >&2
|
||||
exit 0
|
||||
fi
|
||||
|
||||
log "applying candidate relative-budget gate"
|
||||
set +e
|
||||
"$GATE" "${gate_args[@]}"
|
||||
candidate_status=$?
|
||||
set -e
|
||||
|
||||
log "applying A2-vs-A1 baseline drift gate"
|
||||
set +e
|
||||
"$GATE" "${drift_gate_args[@]}"
|
||||
drift_status=$?
|
||||
set -e
|
||||
|
||||
cat >"$OUT_DIR/summary.md" <<EOF
|
||||
# Hotpath Warp ABBA Summary
|
||||
|
||||
- schedule: A1 baseline -> B1 candidate -> B2 candidate -> A2 baseline
|
||||
- runner: $(uname -srm 2>/dev/null || echo unknown)
|
||||
- warp: $("$WARP_BIN" --version 2>/dev/null | head -n1 || echo unknown)
|
||||
- matrix: duration=$DURATION rounds=$ROUNDS cooldown=$COOLDOWN_SECS disks=$DISKS concurrency=$CONCURRENCY
|
||||
- endpoint: $ADDRESS
|
||||
- baseline binary: $BASELINE_BIN
|
||||
- candidate binary: $CANDIDATE_BIN
|
||||
- candidate gate: $OUT_DIR/candidate_gate.md (exit $candidate_status)
|
||||
- baseline drift gate: $OUT_DIR/baseline_drift_gate.md (exit $drift_status)
|
||||
|
||||
Interpretation:
|
||||
|
||||
- Treat candidate gate failures as actionable only when the A2-vs-A1 drift gate is PASS or the affected workload's A2 drift is materially smaller than the B1/B2 candidate delta.
|
||||
- If both candidate and baseline drift fail on the same workload, rerun with longer duration, more rounds, or a quieter runner before assigning causality.
|
||||
EOF
|
||||
|
||||
log "summary written to $OUT_DIR/summary.md"
|
||||
if [[ "$candidate_status" -ne 0 || "$drift_status" -ne 0 ]]; then
|
||||
exit 1
|
||||
fi
|
||||
Executable
+66
@@ -0,0 +1,66 @@
|
||||
#!/usr/bin/env bash
|
||||
# Every `uses: ./.github/actions/setup` must state `cache-save-if` explicitly.
|
||||
#
|
||||
# The composite action's default is "false" (fail-safe: a forgotten input costs
|
||||
# a cold cache, not a slice of the repository-wide 10GB Actions cache quota).
|
||||
# But a silent default in either direction is how audit.yml ended up saving a
|
||||
# ~843MB PR-scoped entry on every dependency change, evicting the main-scoped
|
||||
# lanes that the expensive jobs restore from. Requiring the input to be spelled
|
||||
# out keeps the decision visible in review.
|
||||
#
|
||||
# Usage: scripts/security/check_cache_save_if.sh
|
||||
set -euo pipefail
|
||||
|
||||
cd "$(dirname "$0")/../.."
|
||||
|
||||
status=0
|
||||
|
||||
for file in .github/workflows/*.yml .github/workflows/*.yaml; do
|
||||
[ -e "$file" ] || continue
|
||||
|
||||
# Walk the file, and for every `uses: ./.github/actions/setup` scan the rest
|
||||
# of that step for cache-save-if. `uses:` and `with:` are siblings at the
|
||||
# same indentation, so the step ends only at a line indented *less* than the
|
||||
# `uses:` line (the next `- name:` item, or a dedent out of the steps list).
|
||||
awk -v file="$file" '
|
||||
function flush() {
|
||||
if (pending && !found) {
|
||||
printf "%s:%d: `uses: ./.github/actions/setup` without an explicit `cache-save-if:`\n", file, uses_line > "/dev/stderr"
|
||||
bad++
|
||||
}
|
||||
pending = 0; found = 0
|
||||
}
|
||||
{
|
||||
line = $0
|
||||
sub(/[[:space:]]+$/, "", line)
|
||||
if (line ~ /^[[:space:]]*#/ || line ~ /^[[:space:]]*$/) next
|
||||
|
||||
match(line, /^[[:space:]]*/)
|
||||
indent = RLENGTH
|
||||
|
||||
if (pending && indent < uses_indent) flush()
|
||||
|
||||
if (line ~ /uses:[[:space:]]*\.\/\.github\/actions\/setup[[:space:]]*$/) {
|
||||
flush()
|
||||
pending = 1; found = 0
|
||||
uses_indent = indent
|
||||
uses_line = NR
|
||||
next
|
||||
}
|
||||
|
||||
if (pending && line ~ /^[[:space:]]*cache-save-if:/) found = 1
|
||||
}
|
||||
END { flush(); exit (bad > 0) }
|
||||
' "$file" || status=1
|
||||
done
|
||||
|
||||
if [ "$status" -ne 0 ]; then
|
||||
echo "" >&2
|
||||
echo "Add an explicit cache-save-if to each call above, e.g." >&2
|
||||
echo " cache-save-if: \${{ github.ref == 'refs/heads/main' }} # this lane owns the key" >&2
|
||||
echo " cache-save-if: 'false' # this lane only reads it" >&2
|
||||
echo "Exactly one lane per shared-key may save; see rustfs/backlog#1600." >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
echo "OK: every ./.github/actions/setup call states cache-save-if explicitly"
|
||||
Executable
+65
@@ -0,0 +1,65 @@
|
||||
#!/usr/bin/env bash
|
||||
# Every job that occupies a runner must declare timeout-minutes.
|
||||
#
|
||||
# GitHub's default is 360 minutes. This repository has a history of runners
|
||||
# stalling intermittently (#5394) and of a plain `git checkout` taking 9m57s
|
||||
# under node-level I/O contention, so an undeclared timeout means one wedged job
|
||||
# can hold a runner for six hours out of a pool of roughly 15-21.
|
||||
#
|
||||
# Only jobs with `runs-on` are checked: a job that calls a reusable workflow has
|
||||
# no runner of its own and cannot declare a timeout.
|
||||
#
|
||||
# Usage: scripts/security/check_job_timeouts.sh
|
||||
set -euo pipefail
|
||||
|
||||
cd "$(dirname "$0")/../.."
|
||||
|
||||
status=0
|
||||
|
||||
for file in .github/workflows/*.yml .github/workflows/*.yaml; do
|
||||
[ -e "$file" ] || continue
|
||||
|
||||
awk -v file="$file" '
|
||||
function flush() {
|
||||
if (job != "" && has_runs_on && !has_timeout) {
|
||||
printf "%s:%d: job `%s` has runs-on but no timeout-minutes\n", file, job_line, job > "/dev/stderr"
|
||||
bad++
|
||||
}
|
||||
job = ""; has_runs_on = 0; has_timeout = 0
|
||||
}
|
||||
/^jobs:[[:space:]]*$/ { in_jobs = 1; next }
|
||||
{
|
||||
line = $0
|
||||
sub(/[[:space:]]+$/, "", line)
|
||||
if (line ~ /^[[:space:]]*#/ || line ~ /^[[:space:]]*$/) next
|
||||
|
||||
# A non-indented key ends the jobs mapping.
|
||||
if (in_jobs && line !~ /^[[:space:]]/) { flush(); in_jobs = 0 }
|
||||
if (!in_jobs) next
|
||||
|
||||
if (line ~ /^ [a-zA-Z0-9_-]+:[[:space:]]*$/) {
|
||||
flush()
|
||||
job = line
|
||||
sub(/^[[:space:]]*/, "", job)
|
||||
sub(/:.*$/, "", job)
|
||||
job_line = NR
|
||||
next
|
||||
}
|
||||
if (job == "") next
|
||||
if (line ~ /^ runs-on:/) has_runs_on = 1
|
||||
if (line ~ /^ timeout-minutes:/) has_timeout = 1
|
||||
}
|
||||
END { flush(); exit (bad > 0) }
|
||||
' "$file" || status=1
|
||||
done
|
||||
|
||||
if [ "$status" -ne 0 ]; then
|
||||
echo "" >&2
|
||||
echo "Add timeout-minutes to each job above. Rough budgets used in this repo:" >&2
|
||||
echo " 10 echo-only and guard-script jobs" >&2
|
||||
echo " 30 jobs that call the GitHub API, upload assets, or push over the network" >&2
|
||||
echo " 90+ anything using ./.github/actions/setup (cold cache restore alone is 11-21 min)" >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
echo "OK: every job with runs-on declares timeout-minutes"
|
||||
Reference in New Issue
Block a user