mirror of
https://github.com/rustfs/rustfs.git
synced 2026-09-09 21:56:03 +00:00
ae9fe62fb1
* fix(sse): resolve bucket default encryption per request PUT and the POST-object/extract path resolved a bucket's default encryption with a hard-coded "no explicit SSE-C" flag, so the default was layered onto a request that already carried an SSE-C header triple and then tripped that request's own mutual-exclusion check. Every bucket with default encryption refused SSE-C single PUTs with 400 InvalidArgument, while CreateMultipartUpload on the same bucket succeeded because it resolves SSE elsewhere. Both call sites now derive the flag from the request headers, as COPY already did. The bucket default's KMS key id was also inherited independently of the effective algorithm, so an explicit AES256 request against an aws:kms default bucket produced a self-contradictory algorithm/key-id pair and was rejected. The key id is now inherited only when the effective algorithm is aws:kms, matching the storage-layer resolver. Refs backlog#2368 B1, B2. * fix(sse): refuse SSE-KMS without a running KMS service A write requesting aws:kms on a node with no KMS service fell back to the node-local SSE-S3 provider: the data key was wrapped with RUSTFS_SSE_S3_MASTER_KEY while the object metadata still recorded aws:kms and the requested KMS key id. The stored object claimed a KMS protection it never had, under a key that was never consulted, and no signal distinguished it from a genuine SSE-KMS object. The managed-encryption path now asks the resolved DEK provider whether it wraps with a node-local master key and refuses SSE-KMS in that case: InvalidRequest when KMS was never configured, ServiceUnavailable when a configured service is not running. The check sits after the per-key authorization gate so an unauthorized caller still receives AccessDenied whatever the KMS runtime state is, and asks the provider rather than a parallel availability signal because the provider is what actually wraps the key. A missing master key no longer answers an SSE-KMS request with an SSE-S3-worded configuration error. The SSE-S3 local fallback is unchanged. Refs backlog#2368 B4. * fix(ecstore): restore and archive tiers in stored coordinates Multipart restore addressed the remote tier in plaintext coordinates while the copy-back reads the stored representation. Each part received a misaligned slice of the remote object whose length still satisfied the range, the hash reader and the completion size check, so the restore reported success and silently replaced the object's bytes. Encrypted and compressed multipart objects were both affected. Restore now accumulates stored part sizes, passes the stored length to the hash reader alongside the plaintext length, and validates against the stored size. The copy-back digests stored bytes, so its computed MD5 is not the object's public ETag. Restore now preserves the object ETag on both the single-part and multipart paths, and gives each restored part its own recorded part ETag rather than the object-level value. Transition also handed the tier the object's SSE headers and its RustFS-wrapped data key as request headers. Any S3 target rejected an SSE-C archive outright, an SSE-KMS archive asked the target to encrypt a second time under a key id it does not own, and the wrapped DEK left the cluster. The archive request now strips every SSE header and encryption marker with the predicate the replication path already uses; the local xl.meta keeps all of it, so read-through and restore are unaffected. Objects restored by an affected release are not detected or repaired retroactively and must be re-restored from the tier. Refs backlog#2368 B3, B5; backlog#2369 P7.1. * fix(rio): lock the v1 nonce layout within a segment Decrypting a v1 segment tried three historical nonce layouts per frame, independently for every frame. The last of them exists for streams written before 1.0.0-alpha.91, which reused a segment's part nonce for every block in it; because block zero's derived nonce equals that base nonce, a frame encrypted at index zero authenticated at any position. An attacker able to rewrite the underlying shards could replay it and have the forged plaintext returned with 200 and an unchanged length. Shard integrity uses a keyed-hash-free checksum, which such an attacker can recompute, so it is not a barrier. A segment now locks onto whichever layout decoded its first non-zero-index frame and rejects any later frame needing a different one. That leaves one residual shape: a stream built purely from repeats of frame zero has no later frame to disagree. New RUSTFS_ENCRYPTION_LEGACY_NONCE_FALLBACK (default true, so pre-alpha.91 objects keep decrypting) drops the third layout entirely when set to false, which closes it. Turning it off refuses pre-alpha.91 objects, so migrate them first by rewriting in place. Refs backlog#2369 P2. * fix(kms): reload a service that failed to start POST /rustfs/admin/v3/kms/reload short-circuited whenever the persisted configuration matched the in-memory one byte for byte. A node whose KMS failed to start keeps that configuration and sits in Error, so the documented recovery call returned "reloaded successfully" while leaving the node down. Peers reached the same path through the reload broadcast, so a cluster that lost Vault during a rolling restart had no working recovery route other than the node-local start endpoint. Reload now short-circuits only for a service that is actually running, and otherwise reconfigures, which starts a service that is not running. The AWS backend also advertised key-version enumeration through kms/status, which its own documentation says it cannot do; the capability and its golden snapshot now say false. Refs backlog#2369 P1, P7.3. * docs: record the SSE and KMS changes for 1.0.0 The Unreleased changelog section carried no entry for any encryption work merged since 1.0.0-rc.5, including three items with operational impact: the config-secret variable whose absence persists secrets in cleartext with only a warning, the v2 frame write switch and its rolling-upgrade constraint, and per-key authorization making a public bucket incompatible with SSE-KMS objects. Adds those plus this batch, including the SSE-KMS refusal as a breaking change with both routes out. Also corrects four places where documentation contradicted the code: the cleanup register still called encrypted range seek opt-in after its default flipped, the Helm README claimed vault_mount_path only applies to Transit while the template also feeds the KV2 mount, the disaster-recovery drill listed bundle contents for backends whose export is refused with 501, and the Chinese README capability table predated most of the feature set. Documents the SSE-S3 local master key as a first-class operational mode with its rotation dead end, and what the v1 frame layout does and does not authenticate. Refs backlog#2369 P5. * fix(kms): classify data-path KMS failures by what the caller can do Only "key not found" and a backend outage were classified; every other KMS failure that reached the S3 data path fell through to 500 InternalError with a generic message. A disabled or pending-deletion key, a denied KMS grant, an encryption-context mismatch, an unsupported algorithm, a credential or timeout failure, and a capability the configured backend does not have all looked identical to a server fault. SDKs therefore applied exponential backoff to configuration errors no retry can fix, and monitoring counted every one of them against the server's own error rate. Unusable-key and request-side failures now answer 400, a denied grant 403, transient backend failures 503, and a missing backend capability 501. Damaged, unreadable, or unknown-format key material keeps its 500: it is a server-side integrity fault, and existing tests pin it. The classifier is deliberately separate from the admin lifecycle mapping, which answers 404 for a missing key because there a key id is the resource being addressed; on the data path it arrives inside a request header or a bucket default. Messages either name what the caller asked for or stay generic, with deployment-side detail left on the error source the way the storage-IO mapping already does. Refs backlog#2368 B6. * fix(kms): track and renew static Vault tokens Token authentication hard-coded "this token carries no lease", so the renewal task never started, the remaining-TTL gauge was never published, and nothing looked wrong. `vault token create` grants a 768-hour TTL by default, so a cluster that had been healthy for a month turned every KMS call into a 403 and could not recover without a restart or a reconfigure. Production configuration validation only rejects the literal dev-token, so an ordinary expiring token reaches a whole cluster. The source now reads `auth/token/lookup-self` at login and adopts what Vault reports. A token with no expiry behaves exactly as before. An expiring renewable one is picked up by the existing renewal loop and renewed at half TTL like every other auth method. An expiring non-renewable one warns with its remaining lifetime and publishes the gauge, so the fail-closed window is visible before it arrives. The probe never fails the login: a policy that omits lookup-self, or a Vault that is briefly unreachable, warns and falls back to exactly the previous behaviour rather than taking down a deployment that works today. The scripted Vault test double answers the lookup out of band so existing scripts keep describing only the protocol under test. Refs backlog#2369 P3. * feat(sse): report SSE-C requests that arrive without TLS An SSE-C request carries the customer's AES key in a request header, so AWS S3 and MinIO both refuse one that did not arrive over TLS. RustFS accepted them on any transport: a plaintext hop hands the key to anyone on the path, and since the object cannot be read without that same key, the exposure lasts as long as the object does. Refusing outright is the correct end state but not a safe default to adopt inside a release window, because the project's own s3-tests and e2e lanes and most staging deployments speak plain HTTP. This release reports instead: each such request increments rustfs_ssec_plaintext_requests_total and logs one warning per process, so an operator can confirm nothing would break before the default flips. RUSTFS_SSE_C_REQUIRE_TLS=true opts into the AWS 400 now. The verdict is per connection rather than per deployment: the layer is built with whether this listener terminated TLS, and additionally accepts an https protocol forwarded by a proxy the trusted-proxy configuration already vetted. It sits beside the rate limiter, after the layer that makes a forwarded protocol trustworthy and after the request context, so a rejection can echo the request id. Refs backlog#2369 P7.2. * fix(kms): say what a node-local backend means for a cluster The Local backend keeps key material on each node's own disk and generates its Argon2id salt per node, so two nodes derive different keys from the same master_key and an object encrypted on one node cannot be decrypted on another. Behind a load balancer that surfaces as intermittent 500s on reads that succeeded moments earlier, with nothing tying the symptom to the cause: the only signal was a generic "development, testing and demos only" positioning warning that says nothing about what actually breaks. Configuring or reconfiguring Local while the deployment is distributed now logs a dedicated event and appends the consequence to the configure response, so the operator who made the change sees it. The product decision to warn rather than refuse is unchanged. Refs backlog#2369 P7.4. * docs: record the remaining SSE and KMS changes for 1.0.0 Adds changelog entries for the KMS data-path status classification, the Vault static-token lease probe, the SSE-C plaintext-transport report and its switch, and the node-local backend warning. Documents two things the backend security guide never stated: that SSE-C belongs on a secure transport, with the counter and switch to plan the change around, and that the Local backend cannot be shared by a multi-node deployment because each node derives different keys from the same master key. Refs backlog#2368 B6; backlog#2369 P3, P5, P7.2, P7.4. * fix(kms): report an unreadable key store as an outage on the S3 path A backend now distinguishes a key store it could not read from a key that is genuinely absent, but the S3 boundary collapsed the first one back onto 500 InternalError through the fallthrough for integrity faults. The distinction was therefore invisible to the client: a temporary key-directory outage looked exactly like a permanently damaged key record, and neither the status nor the metric said the request was worth retrying. An unreadable key store joins the retryable class and answers 503, next to a backend error and a credential failure. Damaged, unreadable or unknown-format key material keeps its 500. Refs backlog#2368 B6; builds on rustfs/rustfs#7470.
37 KiB
37 KiB
Changelog
All notable changes to this project will be documented in this file.
The format is based on Keep a Changelog, and this project adheres to Semantic Versioning.
[Unreleased]
Security
- Presigned URLs honour only signed headers (GHSA-g8w9-qw9q-fghr): a SigV4 presigned request that carries an
x-amz-*request header not listed inX-Amz-SignedHeadersis now rejected with403 AccessDenied("There were headers present in the request which were not signed"), matching AWS S3. Previously the holder of a presignedPutObjectURL could add unsignedx-amz-tagging,x-amz-storage-class,x-amz-website-redirect-location, ACL, metadata, Object Lock or SSE headers and have them applied. Presigners that intend a property must set it before signing so the SDK lists the header inSignedHeaders;x-amz-cf-id(CloudFront) remains tolerated unsigned. Header-signed SigV4 and SigV2 requests are unchanged.
Fixed
- Fresh multi-pool bootstrap with distinct format creators: a new deployment whose pools have their first endpoint on different nodes (for example two single-node pools) could never publish its initial
pool.bin: each node held fresh-bootstrap proof only for the pool it formatted, the deployment-wide proof collapsed to none, and every node died withpool metadata recovery required: no durable bootstrap identity or pool.bin replica is availableafter the startup retry budget. The first pool's creator now mints the pending cluster identity on its own pool, every other creator copies that nonce-bound identity onto the pool it formatted first-hand, and the elected writer publishespool.binonce every pool replica carries the same pending identity. Corrupt or disagreeing replicas, pools that merely have a format, expansion pools joining an initialized deployment, and restarts without first-hand proof still fail closed. Non-elected nodes that start beforepool.binexists, and the elected writer while it waits for the other creators, no longer latch their pool-metadata write gate for the life of the process. Refs rustfs/backlog#2338, rustfs/backlog#2375. - Lock RPC timeout storms (#7363): the remote lock client no longer evicts and re-dials the shared internode HTTP/2 channel on every request deadline. A timeout evicts only when the peer has not completed any lock RPC for two deadlines, evictions and transport-failure re-dials are rate limited per peer (
RUSTFS_OBJECT_LOCK_RPC_EVICTION_COOLDOWN_MS, default 5 s), and a timed-out request is left running instead of being reset (bounded per peer byRUSTFS_OBJECT_LOCK_RPC_DETACHED_LIMIT, default 256), so a slow lock endpoint can no longer drive theRST_STREAM/GOAWAY too_many_resets/reconnect loop. A lock granted after its caller timed out is released immediately, and unlocks that fail the quick retries continue on a deferred 1/2/4/8/16 s schedule before the server lease reclaims them. Newrustfs_remote_lock_*metrics cover timeouts, evictions, suppressed evictions, detached streams, late completions and late releases per peer. Operator guide atdocs/operations/lock-rpc-storm-protection.md. - KMS failures on the S3 data path carry an actionable status: only "key not found" and a backend outage were classified; every other KMS failure — a disabled or pending-deletion key, a denied KMS grant, an encryption-context mismatch, an unsupported algorithm, a credential or timeout failure, a capability the backend does not have — collapsed onto
500 InternalError. SDKs therefore applied exponential backoff to configuration errors that no retry can fix, and monitoring filed every one of them as a server fault. Unusable-key and request-side failures now return400, a denied grant403, transient backend failures503— including a key store the backend could not read, so an outage stays distinguishable from a missing key all the way to the client — and a missing backend capability501. Damaged or unreadable key material still returns500, which is what it is. - SSE-C on buckets with default encryption: a
PutObjectcarrying a valid SSE-C header triple on a bucket that has default encryption configured no longer fails with400 InvalidArgument("The SSE-C and managed server-side encryption headers cannot be used together"). PUT and the POST-object/extract path resolved the bucket default with a hard-coded "no explicit SSE-C" flag, so the default was layered onto the request and then tripped the request's own mutual-exclusion check; an SSE-C request now suppresses the bucket default on all three write paths, matching COPY and AWS S3. Every bucket with default encryption previously refused SSE-C single PUTs outright, whileCreateMultipartUploadon the same bucket succeeded. - Explicit SSE-S3 on SSE-KMS-default buckets:
x-amz-server-side-encryption: AES256against a bucket whose default isaws:kmsno longer fails with400 InvalidArgument. The bucket default's KMS key id was inherited independently of the effective algorithm, producing a self-contradictoryAES256+ key-id pair; the key id is now inherited only when the effective algorithm isaws:kms.PutBucketEncryptionfills in a default key id automatically, so this affected nearly every SSE-KMS-default bucket. - Restore of encrypted or compressed multipart objects (silent data corruption): restoring a multipart object from a remote tier addressed the tier in plaintext coordinates while the copy-back reads the stored representation. Every part received a misaligned slice of the remote object whose length still satisfied the range, the hash reader and the completion size check, so the restore reported success and replaced the object's bytes. Restore now accumulates stored part sizes, passes the stored length to the hash reader alongside the plaintext length, and validates against the stored size. Objects restored by an affected release must be re-restored from the tier or recovered from a backup — this release does not detect or repair them retroactively.
- Restore no longer drifts the object ETag: the copy-back digests stored (encrypted or compressed) bytes, so the recomputed MD5 is not the object's public ETag. Single-part and multipart restores now preserve the original object ETag, and each restored part keeps its own recorded part ETag.
- ILM archive no longer forwards encryption metadata to the tier: transition requests carried the object's SSE headers and the RustFS-wrapped data key as request headers. Any S3 target rejected an SSE-C archive outright (
400, no key supplied), an SSE-KMS archive asked the target to encrypt a second time under a key id it does not own, and the wrapped DEK left the cluster. The archive request now strips every SSE header and encryption marker using the same predicate the replication path uses; the localxl.metakeeps all of it, so read-through and restore are unaffected. - KMS reload is no longer a no-op on a node whose KMS failed to start:
POST /rustfs/admin/v3/kms/reloadshort-circuited whenever the persisted configuration matched the in-memory one byte for byte. A node whose KMS failed to start (for example Vault briefly unreachable during a rolling restart) keeps that configuration and sits inError, so the documented recovery call returned "reloaded successfully" while leaving the node down — and did the same on every peer through the reload broadcast. Reload now short-circuits only for a service that is actually running, and otherwise reconfigures, which starts the service. - AWS KMS capability reporting: the AWS backend no longer advertises
versioningsupport throughGET /rustfs/admin/v3/kms/status. AWS KMS key versions are not enumerable through this backend, as the backend documentation already stated. - Multipart admission queue: an
UploadPartwaiting for a foreground write permit now waits at most 10 s by default (RUSTFS_PUT_MULTIPART_FOREGROUND_ADMISSION_WAIT_TIMEOUT_MS, previously 30 s), so a queued part returns S3SlowDownbefore the client's socket write timeout drops the connection. Separately, the API listener no longer forces a 4 MiBSO_RCVBUFon every accepted socket (kernel autotuning applies;RUSTFS_HTTP_SOCKET_RECV_BUFFER_BYTESrestores a fixed size), so a queued part no longer lets up to 8 MiB of unread body accumulate in kernel memory per connection, which is what throttled whole nodes under SDK-default multipart concurrency. Fixes #7385. - Helm Ingress:
customAnnotationsare now merged with class-specific annotations (nginx/traefik) instead of being ignored wheningress.classNameis set. - Per-pool erasure parity: Erasure parity (STANDARD and reduced-redundancy) is now resolved independently for every pool instead of reusing the first pool's value. A heterogeneous topology — for example a 4-drive pool plus a 2-drive pool created during expansion — previously inherited the first pool's parity and could resolve to zero data shards in the smaller pool, panicking Reed-Solomon construction on write. Automatic parity now resolves per pool (for example
2+2in the 4-drive pool and1+1in the 2-drive pool). Fixes #4801.
Added
- On-Demand Migration: Lazy, pull-style migration of an existing S3-compatible bucket into RustFS. A local bucket is attached to an external source bucket; a GET for a key that does not exist locally fetches it from the source, streams it to the client, and stores it locally in the same pass, so every later read is served locally. The module is on by default; set
RUSTFS_ON_DEMAND_MIGRATION_ENABLED=falseon every node to turn it off. A bucket with no source configured behaves exactly as before — the runtime never intervenes on its reads and makes no outbound call. Operator guide atdocs/operations/on-demand-migration.md.- Per-bucket configuration persisted as
on-demand-migration.jsonin the bucket metadata: source provider (s3,aws,minio,rustfs,r2,gcs), endpoint, region, addressing style, credentials and TLS material, an optional key-prefix filter and source-prefix rewrite, and a policy block covering the inline size threshold, multipart part size, concurrency, queue capacity, timeouts, bandwidth limit and negative-cache TTL - Admin routes under
/rustfs/admin/v3/on-demand-migration/{bucket}:PUT(with?dry-run=trueto validate and probe the source without saving),GET,DELETE,GET .../status, plusPOST .../backfill?op=start|cancelandGET .../backfillfor the background full-backfill job with its resumable checkpoint. Authorized by the newadmin:GetBucketOnDemandMigrationandadmin:SetBucketOnDemandMigrationactions; every response redactssecret_keyandsession_token - Read paths: an object at or below
policy.inline_max_bytes(16 MiB by default) is teed to the client and to the local store in a single source read; a larger object or a Range read streams through and a background pull stores the whole object. A HEAD miss is proxied to the source and stores nothing (policy.head = local_onlydisables it). Every source-backed response carriesx-rustfs-on-demand-migration: source - Protections: a per-source circuit breaker, a per-key negative cache, singleflight per key, a concurrency limit and a bounded pull queue shared by the inline and background paths, an optional bandwidth limit, an anti-loop request marker, and the shared outbound-endpoint (SSRF) policy
- Metrics under
rustfs_on_demand_migration_*(requests_total,pulled_bytes_total,pulled_objects_total,pull_failures_total,inflight_pulls,queue_depth,source_latency_seconds_*,breaker_state), mirrored per node by the admin status route - Listings:
ListObjectsv1 remains local with ordinary key markers.ListObjectsV2can merge source objects whenpolicy.list_through = true; this is off by default - Upgrade and rollback: finish upgrading every node before enabling ODM. An rc.5 node that writes bucket configuration drops the ODM fields from metadata; neither a later restart nor moving the service out of ECStore recovers them. Before rollback, disable ODM and securely retain the original full configuration and credentials. After every node returns to a compatible version, restore and validate that configuration. Redacted exports cannot replace the credential backup; source-only objects are unavailable through RustFS while ODM is disabled. See the upgrade and rollback section of
docs/operations/on-demand-migration.md - Optional Google dependencies: default and
fullserver builds retain native GCS support.cargo build -p rustfs --no-default-features --features ftps,webdavexcludes Google SDKs while preserving configuration decoding and redaction; native GCS ODM and tier operations require thegcsfeature. Do not use that build with existing GCS-tiered data - Limitations: PUT and DELETE never reach the source; a source object updated after it was pulled is not re-fetched; SSE-C source objects are unsupported and answer 424;
Last-Modifiedon a pulled object is the local write time, with the source timestamp kept in metadata
- Per-bucket configuration persisted as
- NATS JetStream Publish Path: Opt-in at-least-once delivery for the NATS notify and audit targets. A NATS Core publish flushes to the connection without awaiting a broker acknowledgement, so an event can be lost across a broker restart or a reconnect after the send queue has already cleared it. A queued event now clears only after the JetStream
PublishAck, so bucket notifications survive those interruptions. Off by default and byte-identical to the NATS Core path when disabled.- Three configuration keys per target:
JETSTREAM_ENABLE,JETSTREAM_STREAM_NAME, andJETSTREAM_ACK_TIMEOUT_SECS, under theRUSTFS_NOTIFY_NATS_andRUSTFS_AUDIT_NATS_prefixes - Durable store-and-forward with a stable dedup id sent as the
Nats-Msg-Idheader, so a replay after a crash is collapsed by the server duplicate window - Pre-flight stream validation, and a bounded failed-events store (count and TTL). Only a non-retryable rejection is recorded in the failed-events store. A retryable condition keeps the entry on the live queue until it is delivered
- Operator guide at
docs/operations/nats-jetstream.md
- Three configuration keys per target:
- OpenStack Keystone Authentication Integration: Full support for OpenStack Keystone authentication via X-Auth-Token headers
- Tower-based middleware (
KeystoneAuthLayer) self-contained withinrustfs-keystonecrate - Task-local storage for async-safe credential passing between middleware and auth handlers
- Automatic detection of Keystone credentials (access keys prefixed with
keystone:) - Role-based permission mapping (admin/reseller_admin roles grant owner permissions)
- Token caching for high-performance validation with configurable cache size and TTL
- Dual authentication support: Keystone and standard AWS Signature v4 work simultaneously
- Immediate 401 response for invalid tokens (no fallback to local auth)
- XML-formatted error responses compatible with S3 API
- Comprehensive integration documentation with manual testing guide
- 32 unit and integration tests covering middleware, auth handlers, task-local storage, and role detection
- Tower-based middleware (
- SFTPv3 Protocol Support: SSH-hosted SFTPv3 subsystem that translates each file operation into S3 calls against the local object store. Authentication uses IAM credentials (SSH username = access key, SSH password = secret key).
- Full SFTPv3 packet coverage: open, read, write, stat, lstat, fstat, mkdir, rmdir, rename, remove, opendir, readdir, realpath, close, plus the rest of the 21-packet specification
- Streaming multipart write up to the part size times 10000 parts (156.25 GiB at the default part size)
- Per-handle read-ahead cache with configurable window size and process-wide memory ceiling
- Per-session liveness watchdog: Linux probes
/proc/net/tcpand cancels wedged sessions on the order of 45 seconds; non-Linux falls back to an inactivity ceiling on the order of 30 minutes - 30-second SSH handshake deadline, per-call backend operation timeout, bounded multipart-abort fan-out, graceful-shutdown cascade
- 34 SFTPv3 compliance test cases under
crates/e2e_test/src/protocols/sftp_compliance.rsspread across three entry points:test_sftp_compliance_suite(shared session),test_sftp_compliance_readonly(read-only mode), andtest_sftp_compliance_standalone(one rustfs spawn per case) - Four-layer regression-prevention tests guard against silent feature deletion: compile-time module assertion, module-presence unit test, cross-module
Protocolenum assertion, end-to-end SSH banner test against the running binary
Changed
- Encryption and KMS work merged since
1.0.0-rc.5(entries were missing from this section):- Persisted KMS configuration secrets are sealed field-by-field with
RUSTFS_KMS_CONFIG_SECRET. When the variable is unset the secrets are persisted in cleartext and the server only warns (persisted KMS configuration carries cleartext secrets); it never refuses the write. Set it, identically, on every node, and re-save the configuration to seal an existing one. - New v2 ciphertext frame format with per-frame index binding and final-frame authentication. Its write switch
RUSTFS_ENCRYPTION_FRAME_V2is off by default: v2 frames are unreadable by nodes without v2 read support, and encrypted ciphertext travels verbatim through transition, decommission and SSE-C replication passthrough, so turn it on only after every node — and every RustFS warm/replication target that receives raw ciphertext — runs a release with v2 read support. Reading v2 objects needs no switch. - Per-key SSE-KMS authorization (
RUSTFS_KMS_ENFORCE_SSE_KEY_POLICY, defaultfalse). With it on, anonymous callers hold no KMS grants, so a public bucket serving SSE-KMS objects is an incompatible combination and those reads returnAccessDenied. - Envelope context binding as KMS AAD (
ENV_KMS_ENVELOPE_AAD, off by default; a node that predates the field cannot open bound envelopes). - Vault custom CA and mutual TLS; object-level DEK rewrap plus a batch rekey admin API; a backend-locality runtime signal on
kms/status. - Single-pass decryption for encrypted GET, and encrypted single-part closed-range seek — the latter is now on by default (
RUSTFS_ENCRYPTED_RANGE_SEEK, defaulttrue; the switch remains as a kill switch).
- Persisted KMS configuration secrets are sealed field-by-field with
- Vault static tokens are now tracked and renewed: with
Tokenauthentication RustFS hard-coded "this token has no lease", so the renewal task never started and no remaining-TTL gauge was published.vault token creategrants a 768-hour TTL by default, which turned a healthy-looking cluster into one where every KMS call returned 403 about a month later, with no self-healing short of a restart or reconfigure. RustFS now callsauth/token/lookup-selfat login and adopts what Vault reports: a non-expiring token behaves exactly as before, an expiring renewable one is renewed at half TTL like the other auth methods, and an expiring non-renewable one logsvault_static_token_not_renewableand publishes its remaining TTL. The probe never fails the login: a token whose policy omitslookup-self(Vault'sdefaultpolicy grants it), or a Vault that is unreachable at that moment, logsvault_static_token_lookup_failedand falls back to the previous no-lease behaviour, so no deployment that works today stops working. - SSE-C over a plaintext transport is reported: AWS S3 and MinIO refuse an SSE-C request that did not arrive over TLS, because the customer key travels in a request header. RustFS accepted them on any transport and still does by default — flipping to a rejection inside a release window would break plaintext staging and test deployments. Each such request now increments
rustfs_ssec_plaintext_requests_totaland logs onessec_request_without_tlswarning per process, andRUSTFS_SSE_C_REQUIRE_TLS=trueopts into the AWS400now. The default is expected to flip in a later release; confirm the counter reads zero first. The verdict is per connection: a TLS listener satisfies it, and so does anhttpsprotocol forwarded by a proxy the trusted-proxy configuration accepts. - Local KMS backend on a distributed deployment says what actually breaks: the backend keeps key material and its Argon2id salt on each node's own disk, so two nodes derive different keys from the same
master_keyand an object encrypted on one node cannot be decrypted on another — intermittent 500s behind a load balancer. Configuring it while the deployment is distributed now logskms_node_local_backend_in_distributed_deploymentand appends that consequence to thekms/configureresponse, instead of only the generic "development only" positioning warning. It remains a warning, not a gate. - SSE-KMS is refused when no KMS is running (breaking): a write requesting
x-amz-server-side-encryption: aws:kmson a node with no KMS service no longer succeeds. Earlier releases wrapped the data key with the node-localRUSTFS_SSE_S3_MASTER_KEYwhile still writingaws:kmsand the requested key id into the object metadata — metadata that claimed a KMS protection the object never had, under a key that was never consulted. Such a request now returns400 InvalidRequestwhen KMS was never configured and503when a configured service is not running; the refusal is evaluated after the per-key authorization gate, so an unauthorized caller still receives403 AccessDenied. Upgrade note: a deployment that relied on this write succeeding will start receiving 4xx/503. Either configure a KMS, or requestAES256and keep the documented SSE-S3 local-master-key fallback, which is unchanged. Objects already written this way remain readable. - Legacy ciphertext nonce layouts are now locked per segment: while decrypting a v1 segment, the reader locks onto whichever of the three historical nonce layouts decoded the segment's first non-zero-index frame and rejects any later frame that needs a different one. Because a frame encrypted at block index zero authenticates under the pre-
1.0.0-alpha.91reused-part-nonce layout at any position, an attacker able to rewrite the underlying shards could previously replay it and have the forged plaintext returned with200. NewRUSTFS_ENCRYPTION_LEGACY_NONCE_FALLBACK(defaulttrue) drops that third layout entirely when set tofalse, which closes the residual case of a stream built purely from repeats of frame zero. Turn it off only after migrating pre-alpha.91 encrypted objects (rewrite in place with CopyObject); see KMS backend security properties for what the v1 frame layout does and does not authenticate. - HTTP Server Stack: Integrated
KeystoneAuthLayermiddleware fromrustfs-keystonecrate into service stack (positioned after ReadinessGateLayer) - Storage-class validation on startup (upgrade note): A persisted explicit storage class (
RUSTFS_STORAGE_CLASS_STANDARD/RUSTFS_STORAGE_CLASS_RRS, for exampleEC:2) is now validated against the actual per-pool drive counts at startup and rejected when a pool cannot satisfy it. This is fail-closed and correct, but a cluster that persisted a storage class larger than a small or heterogeneous pool can hold (for exampleEC:2alongside a 2-drive pool), which earlier releases accepted and silently resolved to an invalid layout, will now refuse to start after upgrade. To recover, unsetRUSTFS_STORAGE_CLASS_STANDARDso the server derives a valid per-pool default automatically, or set it to a value every pool can satisfy. - IAMAuth: Enhanced
get_secret_key()to return empty secret for Keystone credentials (bypasses signature validation) - Auth Module: Modified
check_key_valid()to retrieve Keystone credentials from task-local storage and determine admin status StorageBackendtrait: extended with multipart upload methods (create_multipart_upload,upload_part,complete_multipart_upload,abort_multipart_upload) plusupload_part_copy. Streaming-upload code path is now available to FTPS, WebDAV, and Swift drivers as well.Protocolenum: newProtocol::Sftpvariant with correspondingS3Actionmappings. Every match arm onProtocolupdated to handle the new variant exhaustively.
Technical Details
- Middleware is self-contained in
rustfs-keystonecrate following the trusted-proxies pattern for integration-specific middleware - Uses
BoxBodypattern for Hyper 1.x compatibility - Task-local storage provides request-scoped credential passing without modifying HTTP request/response types
- Integration preserves existing S3 authentication flow while adding Keystone support
- Zero breaking changes to existing functionality
- No new top-level directories in main binary crate (middleware lives in integration crate)
- SSH/SFTP wire handling via the
russhandrussh-sftpcrates. SFTPv3 framing is implemented byrussh-sftp; the rustfs-sideSftpDriverimplementsrussh_sftp::server::Handlerand dispatches to the storage backend - Drop-time abort for in-flight multipart uploads honours IAM Deny on
AbortMultipartUpload.start_multipart_uploadcaches the authorisation decision so the synchronousDroppath can honour Allow / Deny policies without re-querying IAM - Per-handle read cache uses an
Arc<AtomicU64>shared across everySftpDriverinstance to enforce a process-wide memory ceiling. On ceiling breach the populate is skipped and the read serves correctly via a single-call backend fetch - Per-session liveness watchdog runs as a tokio task per accepted connection. Reads
/proc/net/tcpand/proc/net/tcp6to look up the (local, peer) tuple's TCP state and cancels viatokio_util::sync::CancellationTokenwhen wedge conditions are confirmed across two consecutive ticks - Path canonicalisation rejects paths containing
\0,\r, or\nand resolves traversal viapath::clean()before any backend dispatch - Cipher / KEX / MAC / host-key algorithm allowlists are hardcoded with no environment override. Strict-KEX (CVE-2023-48795 / Terrapin) marker presence asserted by unit test
- Per-session handle cap (default 64, configurable 8 to 1024) with UUID-generated handle ids
- Crate-level
#![deny(unsafe_code)]is in force acrosscrates/protocols. Socket fd duplication for the watchdog uses the safeAsFd::try_clone_to_ownedpath (Linux). Non-Linux targets use the inactivity-ceiling watchdog - Platform-specific imports are cfg-gated. Unix enforces owner-only host-key mode bits (no group or other permission bits). Windows loads host keys without a mode check and trusts operator-managed NTFS ACLs. Targets that are neither Unix nor Windows fail SFTP at config-load with SftpInitError::UnsupportedPlatform
Documentation
- Updated
crates/keystone/README.mdwith complete integration architecture and workflow - Added detailed manual testing guide with 10 test scenarios
- Updated main
README.mdto list Keystone authentication as available feature - Added troubleshooting section for common integration issues
- Module-level rustdoc on
crates/protocols/src/sftp/mod.rsdescribing the public API surface, configuration contract, and the architecture of the read cache and the wedge watchdog
Configuration
New environment variables:
RUSTFS_KEYSTONE_ENABLE- Enable/disable Keystone authentication (default: false)RUSTFS_KEYSTONE_AUTH_URL- Keystone API endpoint URLRUSTFS_KEYSTONE_VERSION- Keystone API version (v3)RUSTFS_KEYSTONE_ADMIN_USER- Admin username for privileged operationsRUSTFS_KEYSTONE_ADMIN_PASSWORD- Admin passwordRUSTFS_KEYSTONE_ADMIN_PROJECT- Admin project nameRUSTFS_KEYSTONE_ADMIN_DOMAIN- Admin domain name (default: Default)RUSTFS_KEYSTONE_CACHE_SIZE- Token cache size (default: 10000)RUSTFS_KEYSTONE_CACHE_TTL- Token cache TTL in seconds (default: 300)RUSTFS_KEYSTONE_VERIFY_SSL- Verify SSL certificates (default: true)RUSTFS_SFTP_ENABLE- Enable/disable SFTP (default: false)RUSTFS_SFTP_ADDRESS- Listen address (default: 0.0.0.0:2222)RUSTFS_SFTP_HOST_KEY_DIR- Directory containing host key files (must exist). On Unix each file must grant no group or other permission bits (owner access only). On Windows the files load without a mode check and rustfs trusts the directory NTFS ACLRUSTFS_SFTP_HOST_KEY_RELOAD_ENABLE- Rescan the host-key directory without a restart (default: false)RUSTFS_SFTP_HOST_KEY_RELOAD_INTERVAL- Host-key rescan interval in seconds, minimum 5 (default: 30)RUSTFS_SFTP_IDLE_TIMEOUT- Session idle timeout in seconds (default: 600)RUSTFS_SFTP_PART_SIZE- Multipart part size in bytes (default: 16 MiB)RUSTFS_SFTP_READ_ONLY- Reject write packets at the protocol layer (default: false)RUSTFS_SFTP_BANNER- SSH protocol identification string, must begin withSSH-2.0-(default:SSH-2.0-RustFS)RUSTFS_SFTP_HANDLES_PER_SESSION- Per-session open-handle cap, 8 to 1024 (default: 64)RUSTFS_SFTP_BACKEND_OP_TIMEOUT_SECS- Per-call backend deadline in seconds, 5 to 600 (default: 60)RUSTFS_SFTP_READ_CACHE_WINDOW_BYTES- Per-handle read-cache window in bytes, 256 KiB to 64 MiB or 0 to disable (default: 4 MiB)RUSTFS_SFTP_READ_CACHE_TOTAL_MEM_BYTES- Process-wide read-cache memory ceiling in bytes, 16 MiB minimum (default: 256 MiB)
Files Added
crates/protocols/src/sftp/mod.rs- SFTP module entry point, public API surface, crate-level rustdoc, regression-prevention testcrates/protocols/src/sftp/config.rs-SftpConfigandSftpInitErrortypes, env-var resolvers, host-key directory loader with permission enforcementcrates/protocols/src/sftp/constants.rs- Named constants grouped by purpose: S3 error codes, HTTP error codes, POSIX mode bits, protocol identifiers, operational limitscrates/protocols/src/sftp/server.rs-SftpServerSSH server, russh handler, password authentication against IAM, accept loop, per-session task spawncrates/protocols/src/sftp/driver.rs-SftpDriverper-session SFTPv3 handler dispatching each operation onto theStorageBackendcrates/protocols/src/sftp/state.rs-HandleStatevariants for read, write-buffering, write-streaming, write-failed handlescrates/protocols/src/sftp/lifecycle.rs- Per-session activity stamp, weak-ref registry,/proc/net/tcpprobe for the wedge watchdogcrates/protocols/src/sftp/wedge_watchdog.rs- Per-session liveness watchdog cancelling sessions silent at the SFTP layer while the kernel reports CLOSE_WAITcrates/protocols/src/sftp/fallback_watchdog.rs- Per-session silence-only liveness backstop for non-Linux targets, cancelling sessions only at the fallback idle ceilingcrates/protocols/src/sftp/read_cache.rs- Per-handle in-memory read-ahead cache with shared atomic accumulator for the process-wide memory ceilingcrates/protocols/src/sftp/attrs.rs- SFTPv3FileAttributesmapping for objects and directories, longname formatting, mtime clampingcrates/protocols/src/sftp/dir.rs- OPENDIR / READDIR pagination, root-bucket listing, sub-directory listing under a prefixcrates/protocols/src/sftp/errors.rs-SftpErrorthiserror enum and S3-error classification into SFTPv3 status codescrates/protocols/src/sftp/paths.rs- Path canonicalisation, traversal rejection,\0/\r/\nrejection, bucket+key decompositioncrates/protocols/src/sftp/read.rs- READ packet handler, EOF semantics,MAX_READ_LENbound, integration with the read cachecrates/protocols/src/sftp/write.rs- WRITE packet handler, in-memory buffering up to part size, transition to streaming multipart, CLOSE finalisationcrates/protocols/src/sftp/test_support.rs- Test fixtures and helper builders for SFTP unit testscrates/protocols/src/common/dummy_storage.rs- In-memoryStorageBackendtest backend covering every method, used by SFTP unit tests and the FTPS / Swift / WebDAV test suitescrates/e2e_test/src/protocols/sftp_core.rs- End-to-end regressions for the handshake deadline, idle-timeout disconnect, and the wedge watchdogcrates/e2e_test/src/protocols/sftp_compliance.rs- SFTPv3 compliance suite entry points (test_sftp_compliance_suite,test_sftp_compliance_readonly,test_sftp_compliance_standalone)crates/e2e_test/src/protocols/sftp_compliance_tests.rs- Per-case test bodies (CMPTST-01..34), shared fixture helpers, lifecycle counterscrates/e2e_test/src/protocols/sftp_helpers.rs- SFTP-specific test helpers and fixture seeders
Files Modified
crates/keystone/src/middleware.rs- Created Keystone authentication middleware (self-contained in keystone crate)crates/keystone/src/lib.rs- Exported middleware module and KEYSTONE_CREDENTIALScrates/keystone/Cargo.toml- Added Tower/HTTP dependencies for middleware functionalityrustfs/src/server/http.rs- Integrated KeystoneAuthLayer from rustfs-keystone craterustfs/src/auth.rs- Enhanced IAMAuth and check_key_valid for Keystone support, imported KEYSTONE_CREDENTIALS from rustfs-keystonecrates/keystone/README.md- Comprehensive integration documentationREADME.md- Added Keystone as available featureCargo.toml- Added thesftpfeature alongside the existing protocol featuresCargo.lock- Updated to include the newrussh,russh-sftp,socket2,tokio-util,subtle,uuiddependencies and their transitive cratescrates/protocols/Cargo.toml- Declaredrussh,russh-sftp,socket2,tokio-util,subtle,uuidunder thesftpfeature flagcrates/protocols/src/lib.rs- Addedpub mod sftpbehind#[cfg(feature = "sftp")]plus the crate-level#![deny(unsafe_code)]lintcrates/protocols/src/common/client/s3.rs- Extended theStorageBackendtrait withcreate_multipart_upload,upload_part,complete_multipart_upload,abort_multipart_upload, andupload_part_copycrates/protocols/src/common/session.rs- Added theProtocol::Sftpvariant and itsS3Actionmappingscrates/protocols/src/common/gateway.rs- Handles the newProtocol::Sftpvariant exhaustivelycrates/protocols/src/common/mod.rs- Exposed the newdummy_storagemodulecrates/protocols/src/constants.rs- Added shared POSIX mode-bit constants used by SFTP and other protocolscrates/config/src/constants/protocols.rs-RUSTFS_SFTP_*environment variable names and defaultscrates/utils/src/retry.rs- Added the generic exponential-backoff retry helper used by the SFTP write pathcrates/e2e_test/Cargo.toml- Added the e2e test dependencies for SFTP (paramiko fixture, SSH keypair generation)crates/e2e_test/src/protocols/mod.rs- Registered the newsftp_core,sftp_compliance,sftp_compliance_tests, andsftp_helpersmodulescrates/e2e_test/src/protocols/README.md- Documented the SFTP test entry points and case indexcrates/e2e_test/src/protocols/test_env.rs- Added SFTP host-key directory provisioning to the shared protocol test environmentcrates/e2e_test/src/protocols/test_runner.rs- Wired the SFTP entry points into the runnerrustfs/Cargo.toml- Added thesftpfeature flagrustfs/src/lib.rs- One-line addition exporting the SFTP wiringrustfs/src/init.rs- Build and start theSftpServerwhenRUSTFS_SFTP_ENABLEis truerustfs/src/main.rs- Routed shutdown signals to the SFTP server alongside the other protocolsrustfs/src/protocols/client.rs- Client-builder support for the newProtocol::Sftpvariant
Testing
- 16 unit tests in rustfs-keystone crate (config, auth, middleware, identity)
- 10 integration tests in rustfs-keystone crate (task-local storage, middleware layer, scope isolation)
- 6 auth unit tests in rustfs crate (role detection, task-local storage, Keystone credential handling)
- Total: 32 tests passing with zero compilation errors
- Manual testing guide provided for end-to-end validation
- All Keystone tests passing with
cargo test --all --exclude e2e_test - 34 SFTPv3 compliance test cases (CMPTST-01..34) split across three entry points:
test_sftp_compliance_suite(shared session, cases 01-14),test_sftp_compliance_readonly(read-only mode, cases 15-23),test_sftp_compliance_standalone(one rustfs spawn per case, cases 24-34) - Regression-prevention tests at four layers: compile-time module assertion in
crates/protocols/src/lib.rs, module-presence unit test incrates/protocols/src/sftp/mod.rs, cross-moduleProtocolenum assertion, and end-to-end SSH banner test against the running binary - Standalone end-to-end regressions for the SSH handshake deadline, the idle-timeout disconnect path, and the wedge watchdog (Linux fast-kill and the cross-platform fallback path)
- Inline unit tests in every SFTP source file covering pure helpers (path canonicalisation, attribute mapping, S3-error classification, env-var bound resolvers)
- Strict-KEX (CVE-2023-48795) marker presence assertion as a unit test in
crates/protocols/src/sftp/server.rs - All tests passing with
cargo test --all --features sftpagainst a 64-bit Linux target
Previous Releases
See GitHub Releases for previous version history.