Files
rustfs/crates/protocols/src/sftp/server.rs
T
Zhengchao An b2a376c2d2 Merge commit from fork
* fix(admin): bound IAM import archive expansion

MAX_IAM_IMPORT_SIZE caps the compressed upload at 10 MB, but every member of the
archive was then read with read_to_end into an unbounded Vec. Deflate ratios well
above 100:1 are easy to construct, so a small authorized upload could expand
without limit across the seven members ImportIam reads.

Add a shared expansion budget (MAX_IAM_IMPORT_EXPANDED_SIZE, 10x the compressed
cap) drawn down by every member, and route all seven reads through one helper
that reads a byte past the remaining budget to detect overrun. Sharing the budget
bounds the archive as a whole rather than letting each member spend the full
limit independently.

Covers R03-CAN-024 through R03-CAN-030 plus R04-CAN-077 (backlog #1471) — one
fix rather than seven, since all seven call sites were byte-identical.

* fix(kms): confine local key paths and refuse silent key replacement

Local KMS key identifiers arrive from request input — the `name` tag on CreateKey,
the `keyId` body field or query parameter on DeleteKey — and were joined onto
`key_dir` with no validation. An identifier such as `../../tmp/evil` escaped the
configured directory, making key creation a constrained arbitrary-file write and
`DeleteKey` with `force_immediate` a cross-directory delete.

Validate in `master_key_path` and make it fallible, so every filesystem path in
this backend inherits the guard: decode_stored_key, load_master_key,
save_master_key, create_key and delete_key all derive their paths there. The rule
is containment rather than a character allowlist, so identifiers already in use
keep resolving; only separators, NUL, absolute paths and non-single-component
forms are refused. Note `.` and `..` are contained rather than refused — the
`.key` suffix turns them into the ordinary filenames `..key` and `...key`.

Separately, `LocalKmsBackend::create_key` had no existence check, while the
sibling `KmsClient::create_key` has always had one. Since `save_master_key`
renames over its destination, creating a key under an existing name silently
replaced its material and destroyed the ability to decrypt everything wrapped
under it — and the backend path is the one the admin API uses. It now returns
KeyAlreadyExists, matching StaticKmsBackend.

Covers R03-CAN-072, R03-CAN-073 and R07-CAN-103 (backlog #1475). R03-CAN-073
needed no separate change: delete_key routes both its load and its remove_file
through master_key_path.

* fix(swift): bound SLO manifest reads to the 2 MiB manifest limit

The three Swift SLO handlers that load a stored manifest (handle_slo_get,
handle_slo_get_manifest, handle_slo_delete) read the `<object>.slo-manifest`
object to EOF with AsyncReadExt::read_to_end. That key is predictable and
writable through the ordinary object PUT path, so a tenant can replace the
manifest with an arbitrarily large object and then make the server allocate
its full size on every SLO GET, multipart-manifest=get, or
multipart-manifest=delete request - a memory amplification bounded only by
the stored object size (CWE-400 / CWE-770). The 2 MiB manifest limit that
handle_slo_put enforces was not applied on the read side.

Introduce MAX_SLO_MANIFEST_SIZE (the existing 2 MiB PUT limit, now a named
constant) and a shared read_manifest_bytes helper that reads through a
`take(limit + 1)` and rejects anything larger, so an oversized manifest is
refused instead of being buffered first. All three call sites go through the
helper. handle_slo_put now checks the size before parsing the JSON.

Regression tests: test_read_manifest_bytes_rejects_oversized_manifest and
test_read_manifest_bytes_stops_reading_oversized_manifest (which asserts the
reader is not consumed past the limit), plus a boundary test that a manifest
at exactly 2 MiB is still accepted.

* fix(protocols): authorize every object in FTPS/WebDAV recursive deletes

The FTPS and WebDAV gateways authorized only the container before a
recursive delete and then destroyed everything inside it without a
further check:

- FTPS RMD (and DELE on a bucket path ending in '/') cleared
  s3:DeleteBucket, then delete_bucket_recursively listed the bucket and
  deleted every object.
- WebDAV DELETE on a bucket did the same via its own
  delete_bucket_recursively.
- WebDAV DELETE on a directory cleared s3:DeleteObject for the directory
  marker key ("dir/") only, then listed that prefix and deleted every
  child under it.

A principal holding s3:DeleteBucket (or s3:DeleteObject on a single
marker key) could therefore erase objects it had no s3:DeleteObject
permission for, and the operation reported success.

Deletion stays recursive - that is the expected behaviour for these
protocols - but each object now clears s3:DeleteObject on its own key
before it is removed, and the enumeration clears s3:ListBucket. A denial
aborts the whole operation with access denied rather than being skipped,
so the caller can never be told the delete succeeded while objects were
left behind or removed without authorization.

The test double gained shared-state cloning, delete_object/delete_bucket
call logs, and list/delete queue helpers so the regression tests can
observe that nothing is deleted once a deny lands.

* fix(server,ecstore): bound TLS handshakes and remote volume RPC waits

Three call sites let an unauthenticated client or a misbehaving peer hold
server resources with no deadline.

TLS listener (R03-CAN-035): process_connection awaited
`acceptor.accept(socket)` with no bound. A client that opens a TCP
connection and never finishes the handshake parks a Tokio task and a socket
forever, and the connection cap (RUSTFS_API_MAX_CONNECTIONS) is unlimited by
default, so nothing else sheds it. The handshake now runs under
accept_tls_with_deadline(), reusing the existing HTTP/1 header-read budget —
the established slow-client bound for the pre-request phase — and the
expiry is recorded through the same log/metric path as a handshake error,
under a new TIMEOUT failure kind.

Remote disk RPCs (R03-CAN-049, R03-CAN-050): list_volumes and delete_volume
passed Duration::ZERO, which execute_with_timeout treats as "no deadline",
so a peer that accepts the request and never answers stalls the coordinator
(and, for delete_volume, the bucket-deletion workflow). Both now pass
get_max_timeout_duration(), matching every sibling method in the file.

Regression tests: a silent TLS peer must be shed by the handshake deadline;
list_volumes/delete_volume against a peer that completes the TCP connect and
then goes silent must fail with DiskError::Timeout instead of hanging.

* fix(security): stop leaking signed headers and bound OIDC/KMS credentials

Three independent hygiene fixes found by the security review.

R03-CAN-018 (crates/signer): try_get_canonical_headers and get_signed_headers
logged the complete header map at DEBUG before signing. Runtime callers pass
session credentials and SSE-C key material through these headers, so anyone
able to raise the log level (or read DEBUG logs) recovered
X-Amz-Security-Token and SSE-C keys verbatim. The statements were debugging
leftovers with no operational value and are deleted rather than redacted.

R03-CAN-014 (crates/iam): the OIDC HTTP adapter buffered provider responses
with an unbounded Response::bytes(), so a configured, compromised or
attacker-pointed IdP endpoint could stream an arbitrarily large or endless
body into memory (the ValidateOidcConfig admin handler lets a ServerInfo
caller choose the endpoint). Responses are now read incrementally and fail
closed past MAX_OIDC_RESPONSE_SIZE, and the already SSRF-hardened client
builder gains request and connect timeouts so a stalled provider cannot pin
the calling task indefinitely.

R07-CAN-105 (helm): the Vault KMS token was serialized into the chart
ConfigMap, exposing it to every subject allowed to get ConfigMaps in the
namespace. It now renders into a dedicated Secret that the Deployment and
StatefulSet consume via envFrom; the Secret is separate from the main
credentials Secret so it also works when secret.existingSecret is set.

Regression tests:
- rustfs-signer: signing_never_logs_signed_header_material
- rustfs-iam: oidc_response_body_past_the_limit_is_rejected,
  oidc_response_body_at_the_limit_is_accepted
- scripts/test_helm_templates.sh: KMS token must never render in plaintext

* fix(webdav): enforce body limit, request timeout and connection cap

The configured WebDAV maximum body size was enforced from Content-Length, so a
chunked request declared no length and bypassed it entirely. The configured
request timeout was never applied to the connection at all, and the accept loop
spawned a task per connection with no bound, so an unauthenticated client could
hold resources indefinitely and in unbounded number.

Enforce the limit on bytes actually read rather than the declared length, apply
the configured timeout to the request, and bound accepted connections with a new
RUSTFS_WEBDAV_MAX_CONNECTIONS (default 1024) surfaced in the config report.

Covers R03-CAN-051, R03-CAN-052, R03-CAN-067, R04-CAN-089, R05-CAN-094 and
R05-CAN-097 (backlog #1471, #1474).

* fix(security): stop STS credentials from crossing the parent trust boundary

Two related credential-boundary holes let a short-lived STS credential act
with the full, unrestricted authority of the long-term user it was minted
from.

AddUser (R03-CAN-021, CWE-269/863): should_check_deny_only relaxes the admin
policy check to deny-only when a Console/STS session targets the IAM user it
represents. Nothing then stopped that session from calling AddUser with its
own parent's access key, so the handler wrote an attacker-chosen secret key
and status over the parent's stored Credentials via create_user ->
save_user_identity. A session that expires in minutes became permanent
control of the account. AddUser now rejects any temp or service-account
requester whose resolved parent equals the target access key, resolving the
parent the same way should_check_deny_only does (parent_user field, else the
JWT `parent` claim, since some stores persist the parent only in the token).

FTPS/SFTP/WebDAV password auth (R04-CAN-086, CWE-287/862): these protocols
looked the access key up with check_key, which falls back to the STS account
cache, and then compared only the stored secret. An STS access key plus
secret therefore authenticated with no session token presented and no
session-policy claims applied - the holder got the parent's full permissions.
Password authentication now rejects temporary credentials before the secret
comparison. The discriminator is is_temp() && !is_service_account(), the same
one IamCache::update_user_with_claims uses to route an identity into the STS
cache, so service accounts - which resolve policy from stored IAM state
rather than a client-presented token - keep working over these protocols.

Regression tests cover both predicates and pin the guards to their call
sites so neither can be dropped without a test failure.
2026-07-27 00:22:50 +08:00

1603 lines
64 KiB
Rust

// Copyright 2024 RustFS Team
//
// Licensed under the Apache License, Version 2.0 (the "License");
// you may not use this file except in compliance with the License.
// You may obtain a copy of the License at
//
// http://www.apache.org/licenses/LICENSE-2.0
//
// Unless required by applicable law or agreed to in writing, software
// distributed under the License is distributed on an "AS IS" BASIS,
// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
// See the License for the specific language governing permissions and
// limitations under the License.
//! SSH server entry point for the SFTP subsystem.
//!
//! Owns the russh server, accepts incoming SSH connections, performs
//! password authentication against IAM, and dispatches the SFTP subsystem
//! request to a per-session driver instance.
//!
//! Cipher, KEX, MAC, and host key algorithm lists are compile-time constants.
//! There are no env var overrides for the crypto allowlist.
use super::config::{SftpConfig, SftpInitError};
use super::constants::limits::{
DEFAULT_BACKEND_OP_TIMEOUT_SECS, DEFAULT_HANDLES_PER_SESSION, HANDSHAKE_DEADLINE_SECS, KEEPALIVE_INTERVAL_SECS,
KEEPALIVE_MAX, READ_CACHE_TOTAL_MEM_DEFAULT, READ_CACHE_WINDOW_DEFAULT, SSH_CHANNEL_BUFFER_SIZE, SSH_EVENT_BUFFER_SIZE,
SSH_MAXIMUM_PACKET_SIZE,
};
use super::constants::protocol::SFTP_SUBSYSTEM_NAME;
#[cfg(not(target_os = "linux"))]
use super::fallback_watchdog;
use super::lifecycle::{SessionDiag, SessionRegistry, new_session_registry};
#[cfg(target_os = "linux")]
use super::wedge_watchdog;
use crate::common::client::s3::StorageBackend;
use crate::common::session::{Protocol, ProtocolPrincipal, SessionContext, is_temporary_credential};
use russh::keys::{self, PrivateKey, PublicKeyBase64};
use russh::server::{Auth, ChannelOpenHandle, Msg, Session};
use russh::{Channel, ChannelId, ChannelOpenFailure, MethodKind, MethodSet, Pty, Sig};
use rustfs_config::{
DEFAULT_SFTP_HOST_KEY_RELOAD_ENABLE, DEFAULT_SFTP_HOST_KEY_RELOAD_INTERVAL, ENV_SFTP_HOST_KEY_RELOAD_ENABLE,
ENV_SFTP_HOST_KEY_RELOAD_INTERVAL,
};
use rustfs_utils::MaskedAccessKey;
use std::borrow::Cow;
use std::collections::HashMap;
use std::collections::hash_map::DefaultHasher;
use std::fmt::Debug;
use std::hash::Hasher;
use std::net::SocketAddr;
use std::sync::Arc;
use std::sync::RwLock;
use std::sync::atomic::AtomicU64;
use tokio::net::{TcpListener, TcpStream};
use tokio::sync::broadcast;
use tokio::task::JoinSet;
use tokio::time::{Duration, MissedTickBehavior, timeout};
use tokio_util::sync::CancellationToken;
use tracing::{debug, error, info, warn};
use crate::sftp::constants::limits::SHUTDOWN_DRAIN_TIMEOUT_SECS;
const LOG_COMPONENT_PROTOCOLS: &str = "protocols";
const LOG_SUBSYSTEM_SFTP_SERVER: &str = "sftp_server";
const LOG_SUBSYSTEM_SFTP_AUTH: &str = "sftp_auth";
const EVENT_SFTP_SERVER_STATE: &str = "sftp_server_state";
const EVENT_SFTP_HOST_KEY_RELOAD_STATE: &str = "sftp_host_key_reload_state";
const EVENT_SFTP_SESSION_STATE: &str = "sftp_session_state";
const EVENT_SFTP_AUTH_STATE: &str = "sftp_auth_state";
const EVENT_SFTP_SUBSYSTEM_STATE: &str = "sftp_subsystem_state";
// Cipher, KEX, MAC, and host-key algorithm lists. All four are compile-time
// constants with no environment-variable override, so operators cannot
// accidentally downgrade to weak ciphers.
/// AEAD ciphers only. When an AEAD cipher is negotiated the MAC is implicit.
const SFTP_CIPHERS: &[russh::cipher::Name] = &[
russh::cipher::CHACHA20_POLY1305,
russh::cipher::AES_256_GCM,
russh::cipher::AES_128_GCM,
];
/// Key exchange algorithms in preference order.
/// Post-quantum hybrid first, then modern elliptic curve, then FIPS DH and
/// ECDH-NIST, then the two mandatory extension markers.
const SFTP_KEX: &[russh::kex::Name] = &[
russh::kex::MLKEM768X25519_SHA256,
russh::kex::CURVE25519,
russh::kex::CURVE25519_PRE_RFC_8731,
russh::kex::DH_G16_SHA512,
russh::kex::ECDH_SHA2_NISTP256,
russh::kex::ECDH_SHA2_NISTP384,
russh::kex::EXTENSION_SUPPORT_AS_SERVER,
russh::kex::EXTENSION_OPENSSH_STRICT_KEX_AS_SERVER,
];
/// ETM-only MACs as defence-in-depth. The cipher list above is
/// AEAD-only so these MACs are unused under the current configuration.
/// They guard against a future cipher list change that adds a
/// non-AEAD cipher.
const SFTP_MACS: &[russh::mac::Name] = &[russh::mac::HMAC_SHA512_ETM, russh::mac::HMAC_SHA256_ETM];
/// Host key signature algorithms. ssh-rsa (SHA-1) is explicitly absent.
const SFTP_HOST_KEY_ALGORITHMS: &[keys::Algorithm] = &[
keys::Algorithm::Ed25519,
keys::Algorithm::Ecdsa {
curve: keys::EcdsaCurve::NistP256,
},
keys::Algorithm::Ecdsa {
curve: keys::EcdsaCurve::NistP384,
},
keys::Algorithm::Rsa {
hash: Some(keys::HashAlg::Sha512),
},
keys::Algorithm::Rsa {
hash: Some(keys::HashAlg::Sha256),
},
];
/// Compression is disabled. SSH requires the "none" method, and all clients
/// support it. zlib compression adds CPU cost, has been a historical source
/// of vulnerabilities, and provides minimal benefit for SFTP workloads
/// where payloads are typically already compressed (images, archives, etc).
const SFTP_COMPRESSION: &[russh::compression::Name] = &[russh::compression::NONE];
fn build_preferred() -> russh::Preferred {
russh::Preferred {
kex: Cow::Borrowed(SFTP_KEX),
key: Cow::Borrowed(SFTP_HOST_KEY_ALGORITHMS),
cipher: Cow::Borrowed(SFTP_CIPHERS),
mac: Cow::Borrowed(SFTP_MACS),
compression: Cow::Borrowed(SFTP_COMPRESSION),
}
}
fn build_ssh_config(host_keys: Vec<PrivateKey>, idle_timeout_secs: u64, banner: &str) -> Arc<russh::server::Config> {
Arc::new(russh::server::Config {
server_id: russh::SshId::Standard(Cow::from(banner.to_owned())),
methods: MethodSet::from(&[MethodKind::Password][..]),
// No artificial delay on auth failure. Matches the S3 and FTPS
// baseline where auth failures return immediately.
auth_rejection_time: std::time::Duration::from_secs(0),
auth_rejection_time_initial: None,
keys: host_keys,
preferred: build_preferred(),
inactivity_timeout: Some(std::time::Duration::from_secs(idle_timeout_secs)),
keepalive_interval: Some(std::time::Duration::from_secs(KEEPALIVE_INTERVAL_SECS)),
keepalive_max: KEEPALIVE_MAX,
nodelay: true,
// Rationale for the three values below lives on the constants.
maximum_packet_size: SSH_MAXIMUM_PACKET_SIZE,
channel_buffer_size: SSH_CHANNEL_BUFFER_SIZE,
event_buffer_size: SSH_EVENT_BUFFER_SIZE,
..Default::default()
})
}
#[derive(Debug)]
struct SshConfigHolder {
current: RwLock<Arc<russh::server::Config>>,
fingerprint: RwLock<u64>,
}
impl SshConfigHolder {
fn new(config: Arc<russh::server::Config>) -> Self {
let fingerprint = fingerprint_host_keys(&config.keys);
Self {
current: RwLock::new(config),
fingerprint: RwLock::new(fingerprint),
}
}
fn get(&self) -> Arc<russh::server::Config> {
match self.current.read() {
Ok(guard) => Arc::clone(&guard),
Err(poisoned) => Arc::clone(&poisoned.into_inner()),
}
}
async fn reload_from_config(&self, config: &SftpConfig) -> Result<Option<usize>, SftpInitError> {
let host_keys = SftpConfig::load_host_keys(&config.host_key_dir).await?;
let host_key_count = host_keys.len();
let fingerprint = fingerprint_host_keys(&host_keys);
let ssh_config = build_ssh_config(host_keys, config.idle_timeout_secs, &config.banner);
let mut fingerprint_guard = match self.fingerprint.write() {
Ok(guard) => guard,
Err(poisoned) => poisoned.into_inner(),
};
if *fingerprint_guard == fingerprint {
return Ok(None);
}
match self.current.write() {
Ok(mut guard) => *guard = ssh_config,
Err(poisoned) => {
let mut guard = poisoned.into_inner();
*guard = ssh_config;
}
}
*fingerprint_guard = fingerprint;
Ok(Some(host_key_count))
}
}
fn fingerprint_host_keys(host_keys: &[PrivateKey]) -> u64 {
let mut hasher = DefaultHasher::new();
let mut public_keys: Vec<(u8, String)> = host_keys
.iter()
.map(|key| (host_key_algorithm_rank(key.algorithm()), key.public_key_base64()))
.collect();
public_keys.sort_unstable();
for (algorithm_rank, public_key_base64) in public_keys {
hasher.write_u8(algorithm_rank);
hasher.write_usize(public_key_base64.len());
hasher.write(public_key_base64.as_bytes());
}
hasher.finish()
}
fn host_key_algorithm_rank(algorithm: keys::Algorithm) -> u8 {
match algorithm {
keys::Algorithm::Ed25519 => 0,
keys::Algorithm::Ecdsa { .. } => 1,
keys::Algorithm::Rsa { .. } => 2,
_ => 3,
}
}
fn spawn_host_key_reload_loop(config: SftpConfig, holder: Arc<SshConfigHolder>, shutdown_token: CancellationToken) {
let enabled = rustfs_utils::get_env_bool(ENV_SFTP_HOST_KEY_RELOAD_ENABLE, DEFAULT_SFTP_HOST_KEY_RELOAD_ENABLE);
if !enabled {
debug!(
event = EVENT_SFTP_HOST_KEY_RELOAD_STATE,
component = LOG_COMPONENT_PROTOCOLS,
subsystem = LOG_SUBSYSTEM_SFTP_SERVER,
state = "disabled",
env_var = ENV_SFTP_HOST_KEY_RELOAD_ENABLE,
"sftp host key reload state changed"
);
return;
}
let interval_secs =
rustfs_utils::get_env_u64(ENV_SFTP_HOST_KEY_RELOAD_INTERVAL, DEFAULT_SFTP_HOST_KEY_RELOAD_INTERVAL).max(5);
info!(
event = EVENT_SFTP_HOST_KEY_RELOAD_STATE,
component = LOG_COMPONENT_PROTOCOLS,
subsystem = LOG_SUBSYSTEM_SFTP_SERVER,
state = "enabled",
host_key_dir = %config.host_key_dir.display(),
interval_secs,
"SFTP host-key reload enabled"
);
tokio::spawn(async move {
let mut interval = tokio::time::interval(Duration::from_secs(interval_secs));
interval.set_missed_tick_behavior(MissedTickBehavior::Delay);
interval.tick().await;
loop {
tokio::select! {
_ = shutdown_token.cancelled() => {
debug!(
event = EVENT_SFTP_HOST_KEY_RELOAD_STATE,
component = LOG_COMPONENT_PROTOCOLS,
subsystem = LOG_SUBSYSTEM_SFTP_SERVER,
state = "stopped",
host_key_dir = %config.host_key_dir.display(),
"SFTP host-key reload stopped"
);
break;
}
_ = interval.tick() => {}
}
match holder.reload_from_config(&config).await {
Ok(Some(host_key_count)) => {
debug!(
event = EVENT_SFTP_HOST_KEY_RELOAD_STATE,
component = LOG_COMPONENT_PROTOCOLS,
subsystem = LOG_SUBSYSTEM_SFTP_SERVER,
result = "reloaded",
host_key_dir = %config.host_key_dir.display(),
host_key_count,
"SFTP host keys reloaded"
);
}
Ok(None) => {
debug!(
event = EVENT_SFTP_HOST_KEY_RELOAD_STATE,
component = LOG_COMPONENT_PROTOCOLS,
subsystem = LOG_SUBSYSTEM_SFTP_SERVER,
result = "unchanged",
host_key_dir = %config.host_key_dir.display(),
"SFTP host keys unchanged"
);
}
Err(err) => {
warn!(
event = EVENT_SFTP_HOST_KEY_RELOAD_STATE,
component = LOG_COMPONENT_PROTOCOLS,
subsystem = LOG_SUBSYSTEM_SFTP_SERVER,
result = "reload_failed",
host_key_dir = %config.host_key_dir.display(),
err = %err,
"SFTP host-key reload failed"
);
}
}
}
});
}
/// SSH server hosting the SFTP subsystem.
pub struct SftpServer<S: StorageBackend> {
config: SftpConfig,
ssh_config: Arc<SshConfigHolder>,
storage: S,
/// Weak refs to live per-session activity records. Walked by the
/// per-session liveness watchdog and by external observers that
/// enumerate live sessions.
session_registry: Arc<SessionRegistry>,
/// Process-wide accumulator of live read cache memory in bytes,
/// shared across every per-session SftpDriver. The Arc is cloned
/// into each driver and from there into every per-handle
/// ReadCache. Calls to the populate method on any cache, and the
/// Drop impl on any cache, update this one global total. The
/// total is enforced against config.read_cache_total_mem_bytes
/// by the read_inner pre-populate check.
read_cache_in_use: Arc<AtomicU64>,
}
impl<S: StorageBackend + Debug> Debug for SftpServer<S> {
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
f.debug_struct("SftpServer").field("config", &self.config).finish()
}
}
impl<S> SftpServer<S>
where
S: StorageBackend + Clone + Send + Sync + 'static + Debug,
{
/// Build a new server from validated configuration and loaded host keys.
pub fn new(config: SftpConfig, storage: S, host_keys: Vec<PrivateKey>) -> Result<Self, SftpInitError> {
let ssh_config = Arc::new(SshConfigHolder::new(build_ssh_config(
host_keys,
config.idle_timeout_secs,
&config.banner,
)));
Ok(Self {
config,
ssh_config,
storage,
session_registry: Arc::new(new_session_registry()),
read_cache_in_use: Arc::new(AtomicU64::new(0)),
})
}
/// Borrow the configuration the server was built with.
pub fn config(&self) -> &SftpConfig {
&self.config
}
/// Start accepting SSH connections until a shutdown signal is received.
///
/// Each accepted TCP stream is driven on a task tracked in a JoinSet.
/// On shutdown the accept loop exits, then waits up to
/// SHUTDOWN_DRAIN_TIMEOUT_SECS for the live session tasks to finish.
/// When a session task ends, the Drop impl on SftpDriver runs and
/// issues AbortMultipartUpload for every live upload_id where the
/// cached abort_authorized is set to true. Sessions still running
/// after SHUTDOWN_DRAIN_TIMEOUT_SECS are cancelled when the JoinSet
/// is dropped. Stale upload_ids are reclaimed by the bucket
/// AbortIncompleteMultipartUpload lifecycle rule.
///
/// Hot-path design: completed sessions are drained at the top of
/// every loop iteration via JoinSet::try_join_next, which never
/// blocks. The select! below uses biased selection so accept is
/// polled first on every iteration. This combination prevents any
/// interaction between the session drain and the accept-of-next-
/// connection. A second select! arm for join_next under tokio's
/// unbiased random pick could delay accept for closely spaced
/// connections where one session finishes as another arrives.
/// Draining at loop top is synchronous and cannot preempt accept.
pub async fn start(&self, mut shutdown_rx: broadcast::Receiver<()>) -> Result<(), SftpInitError> {
let listener = TcpListener::bind(self.config.bind_addr)
.await
.map_err(|e| SftpInitError::Server(format!("failed to bind {}: {}", self.config.bind_addr, e)))?;
info!(
event = EVENT_SFTP_SERVER_STATE,
component = LOG_COMPONENT_PROTOCOLS,
subsystem = LOG_SUBSYSTEM_SFTP_SERVER,
state = "listening",
bind_addr = %self.config.bind_addr,
read_only = self.config.read_only,
"SFTP server listening"
);
let mut sessions: JoinSet<()> = JoinSet::new();
// Parent cancellation token for the lifetime of this listener.
// Each per-session cancel_token is a child via child_token(),
// and the per-session watchdog selects on the same child. On
// shutdown_rx fire below this token is cancelled before the
// accept loop breaks, which cascades to every live session
// and every watchdog. On Linux the watchdog tasks shut their
// dup'd sockets so russh's inner tasks unblock at the next
// read. On non-Linux targets there is no dup'd socket and the
// session tasks observe the cancellation directly. Either way
// the session tasks drop the RunningSession futures and
// return. drain_sessions then catches up much faster than
// the SHUTDOWN_DRAIN_TIMEOUT_SECS ceiling because no session
// has to wait for the watchdog's natural tick to fire.
let server_shutdown_token = CancellationToken::new();
spawn_host_key_reload_loop(self.config.clone(), Arc::clone(&self.ssh_config), server_shutdown_token.child_token());
loop {
self.drain_finished_tasks(&mut sessions);
tokio::select! {
// Accept has explicit priority. The biased pick plus the
// synchronous drain above keep connection handoff deterministic.
biased;
accept_result = listener.accept() => {
self.handle_accept(accept_result, &mut sessions, &server_shutdown_token);
}
_ = shutdown_rx.recv() => {
info!(
event = EVENT_SFTP_SERVER_STATE,
component = LOG_COMPONENT_PROTOCOLS,
subsystem = LOG_SUBSYSTEM_SFTP_SERVER,
state = "shutdown_requested",
live_sessions = sessions.len(),
"sftp server state changed",
);
// Cascade cancellation to every live session and
// its watchdog before the drain loop runs. Wedged
// sessions need their watchdog to call shutdown on
// the dup'd socket so russh's inner task can end
// by EOF. Otherwise drain_sessions would block on
// those sessions for the full
// SHUTDOWN_DRAIN_TIMEOUT_SECS.
server_shutdown_token.cancel();
break;
}
}
}
drain_sessions(sessions).await;
Ok(())
}
/// Drain finished session tasks from the JoinSet. Synchronous
/// (try_join_next never blocks) and idempotent. Logs one event
/// per drained task: debug for clean ends and cancellations,
/// error for panics.
fn drain_finished_tasks(&self, sessions: &mut JoinSet<()>) {
while let Some(res) = sessions.try_join_next() {
let live = sessions.len();
match res {
Ok(()) => debug!(
event = EVENT_SFTP_SESSION_STATE,
component = LOG_COMPONENT_PROTOCOLS,
subsystem = LOG_SUBSYSTEM_SFTP_SERVER,
state = "finished",
live_sessions = live,
"sftp session state changed"
),
Err(e) if e.is_panic() => {
error!(
event = EVENT_SFTP_SESSION_STATE,
component = LOG_COMPONENT_PROTOCOLS,
subsystem = LOG_SUBSYSTEM_SFTP_SERVER,
result = "task_panicked",
err = %e,
live_sessions = live,
"sftp session state changed"
)
}
Err(e) => debug!(
event = EVENT_SFTP_SESSION_STATE,
component = LOG_COMPONENT_PROTOCOLS,
subsystem = LOG_SUBSYSTEM_SFTP_SERVER,
state = "cancelled",
err = %e,
live_sessions = live,
"sftp session state changed"
),
}
}
}
/// Process one accept-loop result. On Ok, build the per-session
/// state (SessionDiag, watchdog dup socket, child cancel token,
/// SshSessionHandler) and spawn run_session into the JoinSet. On
/// Err, log and return without spawning.
fn handle_accept(
&self,
accept_result: std::io::Result<(TcpStream, SocketAddr)>,
sessions: &mut JoinSet<()>,
server_shutdown_token: &CancellationToken,
) {
let (stream, peer_addr) = match accept_result {
Ok(v) => v,
Err(e) => {
warn!(
event = EVENT_SFTP_SESSION_STATE,
component = LOG_COMPONENT_PROTOCOLS,
subsystem = LOG_SUBSYSTEM_SFTP_SERVER,
result = "accept_failed",
err = %e,
"sftp session state changed"
);
return;
}
};
let ssh_config = self.ssh_config.get();
// Capture local_addr for the wedge watchdog's TCP-state probe.
// Failure here only happens if the kernel can no longer name
// the accepted socket. Fall back to an unspecified address
// that will not match any /proc/net/tcp row, so the probe
// returns None and the watchdog uses its fallback silence
// threshold rather than refusing to spawn.
let local_addr = stream.local_addr().unwrap_or_else(|_| SocketAddr::from(([0u8; 4], 0)));
let session_diag = Arc::new(SessionDiag::new(local_addr, peer_addr));
{
// The registry holds a Vec<Weak<SessionDiag>>. A poisoned
// lock is recovered with PoisonError::into_inner: the Vec
// is still consistent across panics, and the accept loop
// must keep running.
let mut reg = self.session_registry.lock().unwrap_or_else(|poisoned| {
warn!(
event = EVENT_SFTP_SESSION_STATE,
component = LOG_COMPONENT_PROTOCOLS,
subsystem = LOG_SUBSYSTEM_SFTP_SERVER,
result = "registry_poisoned",
"sftp session state changed"
);
poisoned.into_inner()
});
reg.push(Arc::downgrade(&session_diag));
}
let handler_session_diag = Arc::clone(&session_diag);
let watchdog_session_diag = Arc::clone(&session_diag);
// Duplicate the socket via the safe AsFd path before
// run_stream consumes the TcpStream, so the watchdog can
// shut the socket down on wedge detection without racing
// russh for the original fd. The wedge probe itself reads
// /proc/net/tcp[6] in lifecycle::probe_tcp_state and does
// not touch this dup. Non-Linux targets run the silence-only
// fallback watchdog, which needs no socket dup.
#[cfg(target_os = "linux")]
let watchdog_socket = {
let socket = wedge_watchdog::dup_socket(&stream);
if socket.is_none() {
warn!(
event = EVENT_SFTP_SESSION_STATE,
component = LOG_COMPONENT_PROTOCOLS,
subsystem = LOG_SUBSYSTEM_SFTP_SERVER,
result = "watchdog_socket_unavailable",
peer = %peer_addr,
session_id = session_diag.session_id,
"sftp session state changed",
);
}
socket
};
// Per-session cancellation token. Cascades from the
// listener-wide server_shutdown_token so a graceful server
// shutdown ends every live session promptly without waiting
// on the watchdog's natural tick cadence.
let session_shutdown_token = server_shutdown_token.child_token();
let handler = SshSessionHandler {
storage: Arc::new(self.storage.clone()),
peer_addr,
session_context: None,
channels: HashMap::new(),
read_only: self.config.read_only,
part_size: self.config.part_size,
handles_per_session: self.config.handles_per_session.unwrap_or(DEFAULT_HANDLES_PER_SESSION),
backend_op_timeout_secs: self.config.backend_op_timeout_secs.unwrap_or(DEFAULT_BACKEND_OP_TIMEOUT_SECS),
read_cache_window: self.config.read_cache_window_bytes.unwrap_or(READ_CACHE_WINDOW_DEFAULT),
read_cache_total_mem_limit: self.config.read_cache_total_mem_bytes.unwrap_or(READ_CACHE_TOTAL_MEM_DEFAULT),
read_cache_in_use: Arc::clone(&self.read_cache_in_use),
session_diag: handler_session_diag,
};
debug!(
event = EVENT_SFTP_SESSION_STATE,
component = LOG_COMPONENT_PROTOCOLS,
subsystem = LOG_SUBSYSTEM_SFTP_SERVER,
state = "spawned",
peer = %peer_addr,
// sessions.len() reads pre-spawn, so add one to include
// the session about to be inserted.
live_sessions = sessions.len() + 1,
session_id = session_diag.session_id,
"sftp session state changed",
);
#[cfg(target_os = "linux")]
sessions.spawn(run_session(
ssh_config,
stream,
handler,
watchdog_socket,
watchdog_session_diag,
session_shutdown_token,
peer_addr,
));
#[cfg(not(target_os = "linux"))]
sessions.spawn(run_session(
ssh_config,
stream,
handler,
watchdog_session_diag,
session_shutdown_token,
peer_addr,
));
}
}
#[cfg(all(test, unix))]
mod hot_reload_tests {
use super::*;
use std::os::unix::fs::OpenOptionsExt;
use std::path::Path;
use tempfile::TempDir;
const PEM_BOUNDARY_DASHES: &str = "-----";
const PEM_OPENSSH_LABEL: &str = "OPENSSH PRIVATE KEY";
fn build_pem_block(body: &str) -> String {
format!("{d}BEGIN {l}{d}\n{body}\n{d}END {l}{d}\n", d = PEM_BOUNDARY_DASHES, l = PEM_OPENSSH_LABEL,)
}
fn test_ed25519_pem() -> String {
build_pem_block(
"b3BlbnNzaC1rZXktdjEAAAAABG5vbmUAAAAEbm9uZQAAAAAAAAABAAAAMwAAAAtzc2gtZW\n\
QyNTUxOQAAACCkeMEUpnJEbOMBXiQfjZcHZMEbHW3DlNRL+Jbi1cIqMgAAAKDviRiQ74kY\n\
kAAAAAtzc2gtZWQyNTUxOQAAACCkeMEUpnJEbOMBXiQfjZcHZMEbHW3DlNRL+Jbi1cIqMg\n\
AAAEBb5q0DpuL1Rbx4CHUEaRQRSVn1xS2SF+A+qES7OkhrOKR4wRSmckRs4wFeJB+Nlwdk\n\
wRsdbcOU1Ev4luLVwioyAAAAGHNpbW9uc0B1YnVudHUtbGludXgtMjQwNAECAwQF",
)
}
fn test_ecdsa_pem() -> String {
build_pem_block(
"b3BlbnNzaC1rZXktdjEAAAAABG5vbmUAAAAEbm9uZQAAAAAAAAABAAAAaAAAABNlY2RzYS\n\
1zaGEyLW5pc3RwMjU2AAAACG5pc3RwMjU2AAAAQQSBp+cYoqTsQzIF+eQS23gIOBFkIqhi\n\
M8u54NeDrEyxKSewEHP+5i6/+1HURUWDnW+YfS6nbfGb8GxBkJ2ghVvZAAAAqPpS97P6Uv\n\
ezAAAAE2VjZHNhLXNoYTItbmlzdHAyNTYAAAAIbmlzdHAyNTYAAABBBIGn5xiipOxDMgX5\n\
5BLbeAg4EWQiqGIzy7ng14OsTLEpJ7AQc/7mLr/7UdRFRYOdb5h9Lqdt8ZvwbEGQnaCFW9\n\
kAAAAgBdQn3JuP2lSrY3082L+jmYvESyPu9bSmzUe8yMuILzIAAAALdGVzdC12ZWN0b3IB\n\
AgMEBQ==",
)
}
fn write_file_with_mode(path: &Path, content: &str, mode: u32) {
let mut opts = std::fs::OpenOptions::new();
opts.write(true).create(true).truncate(true).mode(mode);
let mut file = opts.open(path).expect("open file");
std::io::Write::write_all(&mut file, content.as_bytes()).expect("write file");
}
fn test_config(host_key_dir: &Path) -> SftpConfig {
SftpConfig {
bind_addr: "0.0.0.0:2222".parse().unwrap(),
host_key_dir: host_key_dir.to_path_buf(),
idle_timeout_secs: 600,
part_size: 16 * 1024 * 1024,
handles_per_session: None,
backend_op_timeout_secs: None,
read_cache_window_bytes: None,
read_cache_total_mem_bytes: None,
read_only: false,
banner: "SSH-2.0-RustFS".to_string(),
}
}
#[tokio::test]
async fn ssh_config_holder_reload_replaces_host_keys_for_new_sessions() {
let dir = TempDir::new().expect("tempdir");
write_file_with_mode(&dir.path().join("ssh_host_ed25519_key"), &test_ed25519_pem(), 0o600);
let config = test_config(dir.path());
let initial_keys = SftpConfig::load_host_keys(dir.path()).await.expect("initial key load");
let holder = SshConfigHolder::new(build_ssh_config(initial_keys, config.idle_timeout_secs, &config.banner));
assert!(matches!(holder.get().keys[0].algorithm(), russh::keys::Algorithm::Ed25519));
std::fs::remove_file(dir.path().join("ssh_host_ed25519_key")).expect("remove old key");
write_file_with_mode(&dir.path().join("ssh_host_ecdsa_key"), &test_ecdsa_pem(), 0o600);
let reloaded = holder.reload_from_config(&config).await.expect("reload host keys");
assert_eq!(reloaded, Some(1));
assert!(matches!(holder.get().keys[0].algorithm(), russh::keys::Algorithm::Ecdsa { .. }));
}
#[tokio::test]
async fn ssh_config_holder_reload_skips_when_host_keys_are_unchanged() {
let dir = TempDir::new().expect("tempdir");
write_file_with_mode(&dir.path().join("ssh_host_ed25519_key"), &test_ed25519_pem(), 0o600);
let config = test_config(dir.path());
let initial_keys = SftpConfig::load_host_keys(dir.path()).await.expect("initial key load");
let holder = SshConfigHolder::new(build_ssh_config(initial_keys, config.idle_timeout_secs, &config.banner));
let reloaded = holder.reload_from_config(&config).await.expect("reload host keys");
assert_eq!(reloaded, None);
}
#[tokio::test]
async fn ssh_config_holder_reload_failure_keeps_current_host_keys() {
let dir = TempDir::new().expect("tempdir");
write_file_with_mode(&dir.path().join("ssh_host_ed25519_key"), &test_ed25519_pem(), 0o600);
let config = test_config(dir.path());
let initial_keys = SftpConfig::load_host_keys(dir.path()).await.expect("initial key load");
let holder = SshConfigHolder::new(build_ssh_config(initial_keys, config.idle_timeout_secs, &config.banner));
assert!(matches!(holder.get().keys[0].algorithm(), russh::keys::Algorithm::Ed25519));
std::fs::remove_file(dir.path().join("ssh_host_ed25519_key")).expect("remove old key");
let err = holder
.reload_from_config(&config)
.await
.expect_err("empty host key directory must fail reload");
assert!(matches!(err, SftpInitError::NoHostKeysFound { .. }));
assert!(matches!(holder.get().keys[0].algorithm(), russh::keys::Algorithm::Ed25519));
}
#[tokio::test]
async fn fingerprint_host_keys_is_order_independent() {
let dir = TempDir::new().expect("tempdir");
write_file_with_mode(&dir.path().join("ssh_host_ed25519_key"), &test_ed25519_pem(), 0o600);
write_file_with_mode(&dir.path().join("ssh_host_ecdsa_key"), &test_ecdsa_pem(), 0o600);
let keys = SftpConfig::load_host_keys(dir.path()).await.expect("load keys");
let forward = fingerprint_host_keys(&keys);
let reversed_keys: Vec<_> = keys.into_iter().rev().collect();
let reversed = fingerprint_host_keys(&reversed_keys);
assert_eq!(forward, reversed);
}
}
/// Drive one accepted SSH session through handshake, optional
/// watchdog spawn, the post-handshake session loop, and cleanup.
/// Free function (not a method) so the spawn closure on the JoinSet
/// does not have to satisfy a 'static bound on a borrow of &self.
#[allow(clippy::too_many_arguments)]
async fn run_session<S>(
ssh_config: Arc<russh::server::Config>,
stream: tokio::net::TcpStream,
handler: SshSessionHandler<S>,
#[cfg(target_os = "linux")] watchdog_socket: Option<socket2::Socket>,
watchdog_session_diag: Arc<SessionDiag>,
cancel_token: CancellationToken,
peer_addr: SocketAddr,
) where
S: StorageBackend + Send + Sync + 'static,
{
debug!(
event = EVENT_SFTP_SESSION_STATE,
component = LOG_COMPONENT_PROTOCOLS,
subsystem = LOG_SUBSYSTEM_SFTP_SERVER,
state = "task_entered",
peer = %peer_addr,
"sftp session state changed"
);
// run_stream covers SSH KEX and password auth. Cap with a
// wallclock deadline so a peer that completes TCP but stalls
// before KEXINIT (or that drives KEX or auth so slowly that no
// SSH-layer timer fires) cannot pin a spawn-task slot forever.
// The post-handshake session loop has its own inactivity and
// keepalive timers.
let handshake_deadline = Duration::from_secs(HANDSHAKE_DEADLINE_SECS);
let session = match timeout(handshake_deadline, russh::server::run_stream(ssh_config, stream, handler)).await {
Ok(Ok(s)) => s,
Ok(Err(e)) => {
debug!(
event = EVENT_SFTP_SESSION_STATE,
component = LOG_COMPONENT_PROTOCOLS,
subsystem = LOG_SUBSYSTEM_SFTP_SERVER,
result = "setup_failed",
peer = %peer_addr,
err = %e,
"sftp session state changed"
);
return;
}
Err(_elapsed) => {
warn!(
event = EVENT_SFTP_SESSION_STATE,
component = LOG_COMPONENT_PROTOCOLS,
subsystem = LOG_SUBSYSTEM_SFTP_SERVER,
result = "handshake_timeout",
peer = %peer_addr,
deadline_secs = HANDSHAKE_DEADLINE_SECS,
"sftp session state changed",
);
return;
}
};
debug!(
event = EVENT_SFTP_SESSION_STATE,
component = LOG_COMPONENT_PROTOCOLS,
subsystem = LOG_SUBSYSTEM_SFTP_SERVER,
state = "handshake_completed",
peer = %peer_addr,
"sftp session state changed"
);
// Spawn the per-session watchdog. On Linux the wedge watchdog
// observes the SFTP-handler activity stamp and the kernel TCP
// state. On wedge detection it shuts the duplicated socket down
// (so russh's inner task unwedges via EOF propagation) and
// cancels the shared CancellationToken (so this task drops
// RunningSession and ends). On non-Linux targets the fallback
// watchdog cancels only after silence past the fallback
// threshold, with no socket introspection.
#[cfg(target_os = "linux")]
if let Some(socket) = watchdog_socket {
wedge_watchdog::spawn_for_session(watchdog_session_diag, socket, cancel_token.clone());
}
#[cfg(not(target_os = "linux"))]
fallback_watchdog::spawn_for_session(watchdog_session_diag, cancel_token.clone());
// Await the RunningSession until the client disconnects, until
// russh's inactivity_timeout closes a healthy idle peer, or until
// the watchdog (or the listener-wide shutdown cascade) cancels.
let session_result = tokio::select! {
res = session => Some(res),
_ = cancel_token.cancelled() => None,
};
// Either branch ends the watchdog: the cancel arm because
// cancel already fired, the await arm by signalling cancel here
// so the watchdog task exits its loop.
cancel_token.cancel();
let session_result = match session_result {
Some(r) => r,
None => {
warn!(
event = EVENT_SFTP_SESSION_STATE,
component = LOG_COMPONENT_PROTOCOLS,
subsystem = LOG_SUBSYSTEM_SFTP_SERVER,
state = "cancelled",
peer = %peer_addr,
"sftp session state changed"
);
return;
}
};
match session_result {
Ok(()) => {
debug!(
event = EVENT_SFTP_SESSION_STATE,
component = LOG_COMPONENT_PROTOCOLS,
subsystem = LOG_SUBSYSTEM_SFTP_SERVER,
state = "ended_cleanly",
peer = %peer_addr,
"sftp session state changed"
);
}
Err(e) => {
debug!(
event = EVENT_SFTP_SESSION_STATE,
component = LOG_COMPONENT_PROTOCOLS,
subsystem = LOG_SUBSYSTEM_SFTP_SERVER,
result = "ended_with_error",
peer = %peer_addr,
err = %e,
"sftp session state changed"
);
}
}
}
/// Wait up to SHUTDOWN_DRAIN_TIMEOUT_SECS for every live session task to
/// finish. Sessions that do not return within the window are cancelled
/// when drain_sessions returns and the JoinSet is dropped.
///
/// Scope: this drain covers the per-connection session tasks only. The
/// AbortMultipartUpload tasks that SftpDriver::Drop spawns via
/// tokio::spawn are fire-and-forget and are NOT tracked here, because
/// Drop is synchronous and cannot hand back a JoinHandle. Abort tasks
/// that do not complete before the runtime shuts down fall to the
/// bucket AbortIncompleteMultipartUpload lifecycle rule. Tracking them
/// would require Drop to own a shared JoinSet, which contradicts the
/// per-session ownership model.
async fn drain_sessions(mut sessions: JoinSet<()>) {
if sessions.is_empty() {
return;
}
let drain = async {
while let Some(res) = sessions.join_next().await {
if let Err(e) = res
&& e.is_panic()
{
error!(
event = EVENT_SFTP_SESSION_STATE,
component = LOG_COMPONENT_PROTOCOLS,
subsystem = LOG_SUBSYSTEM_SFTP_SERVER,
result = "drain_task_panicked",
err = %e,
"sftp session state changed"
);
}
}
};
match timeout(Duration::from_secs(SHUTDOWN_DRAIN_TIMEOUT_SECS), drain).await {
Ok(()) => info!(
event = EVENT_SFTP_SERVER_STATE,
component = LOG_COMPONENT_PROTOCOLS,
subsystem = LOG_SUBSYSTEM_SFTP_SERVER,
state = "drain_completed",
"sftp server state changed"
),
Err(_) => warn!(
event = EVENT_SFTP_SERVER_STATE,
component = LOG_COMPONENT_PROTOCOLS,
subsystem = LOG_SUBSYSTEM_SFTP_SERVER,
result = "drain_timeout",
timeout_secs = SHUTDOWN_DRAIN_TIMEOUT_SECS,
live = sessions.len(),
"sftp server state changed",
),
}
}
/// Per-connection SSH handler. Implements russh::server::Handler.
///
/// Handles authentication against IAM and dispatches the SFTP subsystem.
/// All non-SFTP channel types are rejected.
struct SshSessionHandler<S: StorageBackend> {
/// S3 storage backend shared across all sessions.
storage: Arc<S>,
/// Client IP from the TCP connection. Used for logging and for
/// building SessionContext.
peer_addr: SocketAddr,
/// Session context built after successful auth_password.
/// None before authentication.
session_context: Option<SessionContext>,
/// Open channels indexed by ChannelId. A HashMap rather than
/// Option because SSH permits multiple concurrent channels per
/// connection (RFC 4254 section 5.1).
channels: HashMap<ChannelId, Channel<Msg>>,
/// Whether write operations are rejected.
read_only: bool,
/// S3 multipart part size in bytes, forwarded to every per-session
/// SftpDriver at subsystem_request time.
part_size: u64,
/// Maximum number of simultaneously-open SFTP handles per session,
/// forwarded to every per-session SftpDriver at subsystem_request
/// time.
handles_per_session: usize,
/// Per-call deadline applied to every StorageBackend invocation,
/// forwarded to every per-session SftpDriver at subsystem_request
/// time. Resolved from RUSTFS_SFTP_BACKEND_OP_TIMEOUT_SECS or
/// DEFAULT_BACKEND_OP_TIMEOUT_SECS at server-build time.
backend_op_timeout_secs: u64,
/// Per-handle read cache window size in bytes, forwarded to every
/// per-session SftpDriver at subsystem_request time. Resolved
/// from RUSTFS_SFTP_READ_CACHE_WINDOW_BYTES or
/// READ_CACHE_WINDOW_DEFAULT at server-build time.
read_cache_window: u64,
/// Process-wide cumulative read cache memory ceiling in bytes,
/// forwarded to every per-session SftpDriver at subsystem_request
/// time. Resolved from RUSTFS_SFTP_READ_CACHE_TOTAL_MEM_BYTES or
/// READ_CACHE_TOTAL_MEM_DEFAULT at server-build time.
read_cache_total_mem_limit: u64,
/// Process-wide accumulator of live read cache memory in bytes,
/// shared across every per-session SftpDriver. The Arc is cloned
/// from SftpServer.read_cache_in_use so all drivers contribute to
/// one global total.
read_cache_in_use: Arc<AtomicU64>,
/// Per-session activity record. Stamped from auth_password and
/// subsystem_request at session level, then handed to the per-session
/// SftpDriver where the SFTP-handler stamps live.
session_diag: Arc<SessionDiag>,
}
impl<S: StorageBackend + Send + Sync + 'static> russh::server::Handler for SshSessionHandler<S> {
type Error = russh::Error;
#[tracing::instrument(level = "warn", skip(self), fields(user = %user, peer = %self.peer_addr))]
fn auth_none(&mut self, user: &str) -> impl std::future::Future<Output = Result<Auth, Self::Error>> + Send {
async { Ok(Auth::reject()) }
}
// NOTE: the russh 0.60 default for auth_publickey_offered is
// Auth::Accept, so this override is mandatory to prevent
// signature verification running for an auth method that is
// not offered.
#[tracing::instrument(level = "debug", skip(self, _key), fields(user = %user, peer = %self.peer_addr))]
fn auth_publickey_offered(
&mut self,
user: &str,
_key: &keys::PublicKey,
) -> impl std::future::Future<Output = Result<Auth, Self::Error>> + Send {
async { Ok(Auth::reject()) }
}
#[tracing::instrument(level = "warn", skip(self, _key), fields(user = %user, peer = %self.peer_addr))]
fn auth_publickey(
&mut self,
user: &str,
_key: &keys::PublicKey,
) -> impl std::future::Future<Output = Result<Auth, Self::Error>> + Send {
async { Ok(Auth::reject()) }
}
#[tracing::instrument(level = "info", skip(self, password), fields(user = %user, peer = %self.peer_addr))]
fn auth_password(
&mut self,
user: &str,
password: &str,
) -> impl std::future::Future<Output = Result<Auth, Self::Error>> + Send {
let user = user.to_owned();
let password = password.to_owned();
let masked_user = MaskedAccessKey(&user).to_string();
let peer_addr = self.peer_addr;
let session_diag = Arc::clone(&self.session_diag);
session_diag.stamp();
async move {
let iam_sys = match rustfs_iam::get() {
Ok(sys) => sys,
Err(e) => {
error!(
event = EVENT_SFTP_AUTH_STATE,
component = LOG_COMPONENT_PROTOCOLS,
subsystem = LOG_SUBSYSTEM_SFTP_AUTH,
result = "iam_unavailable",
peer = %peer_addr,
err = %e,
"sftp auth state changed"
);
return Ok(Auth::reject());
}
};
let (identity_opt, is_valid) = match iam_sys.check_key(&user).await {
Ok(result) => result,
Err(e) => {
error!(
event = EVENT_SFTP_AUTH_STATE,
component = LOG_COMPONENT_PROTOCOLS,
subsystem = LOG_SUBSYSTEM_SFTP_AUTH,
result = "check_key_failed",
user = %masked_user,
peer = %peer_addr,
err = %e,
"sftp auth state changed"
);
return Ok(Auth::reject());
}
};
let identity = match identity_opt {
Some(id) => id,
None => {
warn!(
event = EVENT_SFTP_AUTH_STATE,
component = LOG_COMPONENT_PROTOCOLS,
subsystem = LOG_SUBSYSTEM_SFTP_AUTH,
result = "unknown_access_key",
user = %masked_user,
peer = %peer_addr,
"sftp auth state changed"
);
return Ok(Auth::reject());
}
};
// Reject disabled or expired accounts. FTPS checks this at
// ftps/server.rs:286. SFTP must do the same.
if !is_valid {
warn!(
event = EVENT_SFTP_AUTH_STATE,
component = LOG_COMPONENT_PROTOCOLS,
subsystem = LOG_SUBSYSTEM_SFTP_AUTH,
result = "account_inactive",
user = %masked_user,
peer = %peer_addr,
"sftp auth state changed"
);
return Ok(Auth::reject());
}
if is_temporary_credential(&identity.credentials) {
warn!(
event = EVENT_SFTP_AUTH_STATE,
component = LOG_COMPONENT_PROTOCOLS,
subsystem = LOG_SUBSYSTEM_SFTP_AUTH,
result = "temporary_credential_rejected",
user = %masked_user,
peer = %peer_addr,
"sftp auth state changed"
);
return Ok(Auth::reject());
}
// Constant-time secret comparison to prevent timing side-channel
// attacks. Same primitive used by rustfs/src/auth.rs.
use subtle::ConstantTimeEq;
let secret_matches: bool = identity.credentials.secret_key.as_bytes().ct_eq(password.as_bytes()).into();
if !secret_matches {
warn!(
event = EVENT_SFTP_AUTH_STATE,
component = LOG_COMPONENT_PROTOCOLS,
subsystem = LOG_SUBSYSTEM_SFTP_AUTH,
result = "invalid_secret_key",
user = %masked_user,
peer = %peer_addr,
"sftp auth state changed"
);
return Ok(Auth::reject());
}
let principal = ProtocolPrincipal::new(Arc::new(identity));
self.session_context = Some(SessionContext::new(principal, Protocol::Sftp, peer_addr.ip()));
debug!(
event = EVENT_SFTP_AUTH_STATE,
component = LOG_COMPONENT_PROTOCOLS,
subsystem = LOG_SUBSYSTEM_SFTP_AUTH,
result = "authenticated",
user = %masked_user,
peer = %peer_addr,
"sftp auth state changed"
);
Ok(Auth::Accept)
}
}
#[tracing::instrument(level = "debug", skip(self, channel, reply, _session), fields(peer = %self.peer_addr))]
fn channel_open_session(
&mut self,
channel: Channel<Msg>,
reply: ChannelOpenHandle,
_session: &mut Session,
) -> impl std::future::Future<Output = Result<(), Self::Error>> + Send {
let id = channel.id();
self.channels.insert(id, channel);
async move {
reply.accept().await;
Ok(())
}
}
#[tracing::instrument(level = "debug", skip(self, _session), fields(peer = %self.peer_addr, channel = ?channel))]
fn channel_close(
&mut self,
channel: ChannelId,
_session: &mut Session,
) -> impl std::future::Future<Output = Result<(), Self::Error>> + Send {
self.channels.remove(&channel);
async { Ok(()) }
}
fn subsystem_request(
&mut self,
channel_id: ChannelId,
name: &str,
session: &mut Session,
) -> impl std::future::Future<Output = Result<(), Self::Error>> + Send {
self.session_diag.stamp();
// Inputs the future needs once we decide to actually run the SFTP
// driver. None means "rejected synchronously. Just return Ok".
struct RunInputs<S: StorageBackend> {
stream: russh::ChannelStream<Msg>,
storage: Arc<S>,
session_context: SessionContext,
read_only: bool,
part_size: u64,
handles_per_session: usize,
backend_op_timeout_secs: u64,
read_cache_window: u64,
read_cache_total_mem_limit: u64,
read_cache_in_use: Arc<AtomicU64>,
session_diag: Arc<SessionDiag>,
}
let inputs: Result<Option<RunInputs<S>>, Self::Error> = (|| {
if name != SFTP_SUBSYSTEM_NAME {
warn!(
event = EVENT_SFTP_SUBSYSTEM_STATE,
component = LOG_COMPONENT_PROTOCOLS,
subsystem = LOG_SUBSYSTEM_SFTP_SERVER,
result = "unsupported_subsystem",
subsystem = %name,
peer = %self.peer_addr,
"sftp subsystem state changed"
);
session.channel_failure(channel_id)?;
return Ok(None);
}
let channel = match self.channels.remove(&channel_id) {
Some(ch) => ch,
None => {
error!(
event = EVENT_SFTP_SUBSYSTEM_STATE,
component = LOG_COMPONENT_PROTOCOLS,
subsystem = LOG_SUBSYSTEM_SFTP_SERVER,
result = "channel_missing",
channel = ?channel_id,
peer = %self.peer_addr,
"sftp subsystem state changed"
);
session.channel_failure(channel_id)?;
return Ok(None);
}
};
let session_context = match self.session_context.clone() {
Some(ctx) => ctx,
None => {
error!(
event = EVENT_SFTP_SUBSYSTEM_STATE,
component = LOG_COMPONENT_PROTOCOLS,
subsystem = LOG_SUBSYSTEM_SFTP_SERVER,
result = "unauthenticated_request",
channel = ?channel_id,
peer = %self.peer_addr,
"sftp subsystem state changed"
);
session.channel_failure(channel_id)?;
return Ok(None);
}
};
session.channel_success(channel_id)?;
debug!(
event = EVENT_SFTP_SUBSYSTEM_STATE,
component = LOG_COMPONENT_PROTOCOLS,
subsystem = LOG_SUBSYSTEM_SFTP_SERVER,
state = "accepted",
channel = ?channel_id,
peer = %self.peer_addr,
"sftp subsystem state changed"
);
Ok(Some(RunInputs {
stream: channel.into_stream(),
storage: Arc::clone(&self.storage),
session_context,
read_only: self.read_only,
part_size: self.part_size,
handles_per_session: self.handles_per_session,
backend_op_timeout_secs: self.backend_op_timeout_secs,
read_cache_window: self.read_cache_window,
read_cache_total_mem_limit: self.read_cache_total_mem_limit,
read_cache_in_use: Arc::clone(&self.read_cache_in_use),
session_diag: Arc::clone(&self.session_diag),
}))
})();
async move {
if let Some(inputs) = inputs? {
// russh_sftp::server::run is async and spawns its own
// task internally. Awaiting the returned future ensures
// the spawn completes before subsystem_request returns.
let driver = super::driver::SftpDriver::new(
inputs.storage,
inputs.session_context,
inputs.read_only,
inputs.part_size,
inputs.handles_per_session,
inputs.backend_op_timeout_secs,
inputs.read_cache_window,
inputs.read_cache_total_mem_limit,
inputs.read_cache_in_use,
inputs.session_diag,
);
russh_sftp::server::run(inputs.stream, driver).await;
}
Ok(())
}
}
// Every channel-type method russh exposes is overridden explicitly so
// a russh default flip from "reject" to "accept" cannot silently turn
// this into a general SSH host. SFTP subsystem only.
#[tracing::instrument(level = "warn", skip(self, _modes, session), fields(peer = %self.peer_addr, channel = ?channel))]
fn pty_request(
&mut self,
channel: ChannelId,
_term: &str,
_col_width: u32,
_row_height: u32,
_pix_width: u32,
_pix_height: u32,
_modes: &[(Pty, u32)],
session: &mut Session,
) -> impl std::future::Future<Output = Result<(), Self::Error>> + Send {
let result = session.channel_failure(channel);
async move {
result?;
Ok(())
}
}
#[tracing::instrument(level = "warn", skip(self, session), fields(peer = %self.peer_addr, channel = ?channel))]
fn shell_request(
&mut self,
channel: ChannelId,
session: &mut Session,
) -> impl std::future::Future<Output = Result<(), Self::Error>> + Send {
let result = session.channel_failure(channel);
async move {
result?;
Ok(())
}
}
#[tracing::instrument(level = "warn", skip(self, _data, session), fields(peer = %self.peer_addr, channel = ?channel))]
fn exec_request(
&mut self,
channel: ChannelId,
_data: &[u8],
session: &mut Session,
) -> impl std::future::Future<Output = Result<(), Self::Error>> + Send {
let result = session.channel_failure(channel);
async move {
result?;
Ok(())
}
}
// Signal requests carry want_reply = false (RFC 4254
// section 6.9), so there is no channel_failure to send. The
// override exists to log probe attempts and to keep the SFTP
// server SFTP-only by code rather than by relying on the russh
// default.
#[tracing::instrument(level = "warn", skip(self, _session), fields(peer = %self.peer_addr, channel = ?channel))]
fn signal(
&mut self,
channel: ChannelId,
signal: Sig,
_session: &mut Session,
) -> impl std::future::Future<Output = Result<(), Self::Error>> + Send {
tracing::warn!(channel = ?channel, signal = ?signal, "rejecting SFTP signal request");
async { Ok(()) }
}
#[tracing::instrument(level = "warn", skip(self, session), fields(peer = %self.peer_addr, channel = ?channel))]
fn env_request(
&mut self,
channel: ChannelId,
_variable_name: &str,
_variable_value: &str,
session: &mut Session,
) -> impl std::future::Future<Output = Result<(), Self::Error>> + Send {
let result = session.channel_failure(channel);
async move {
result?;
Ok(())
}
}
#[tracing::instrument(level = "warn", skip(self, session), fields(peer = %self.peer_addr, channel = ?channel))]
fn x11_request(
&mut self,
channel: ChannelId,
_single_connection: bool,
_x11_auth_protocol: &str,
_x11_auth_cookie: &str,
_x11_screen_number: u32,
session: &mut Session,
) -> impl std::future::Future<Output = Result<(), Self::Error>> + Send {
let result = session.channel_failure(channel);
async move {
result?;
Ok(())
}
}
#[tracing::instrument(level = "warn", skip(self, _session), fields(peer = %self.peer_addr, address = %address))]
fn tcpip_forward(
&mut self,
address: &str,
_port: &mut u32,
_session: &mut Session,
) -> impl std::future::Future<Output = Result<bool, Self::Error>> + Send {
async { Ok(false) }
}
#[tracing::instrument(level = "warn", skip(self, _session), fields(peer = %self.peer_addr, address = %address))]
fn cancel_tcpip_forward(
&mut self,
address: &str,
_port: u32,
_session: &mut Session,
) -> impl std::future::Future<Output = Result<bool, Self::Error>> + Send {
async { Ok(false) }
}
#[tracing::instrument(level = "warn", skip(self, _session), fields(peer = %self.peer_addr, channel = ?_channel))]
fn agent_request(
&mut self,
_channel: ChannelId,
_session: &mut Session,
) -> impl std::future::Future<Output = Result<bool, Self::Error>> + Send {
async { Ok(false) }
}
// Channel-open rejections. russh defaults to rejecting dropped channel-open
// handles, but we override them explicitly with a warn log so
// (a) probe attempts are visible in operator logs and
// (b) a future russh default flip cannot silently allow these
// channel types.
#[tracing::instrument(level = "warn", skip(self, _channel, reply, _session), fields(peer = %self.peer_addr, host = %host_to_connect, port = port_to_connect))]
fn channel_open_direct_tcpip(
&mut self,
_channel: Channel<Msg>,
host_to_connect: &str,
port_to_connect: u32,
_originator_address: &str,
_originator_port: u32,
reply: ChannelOpenHandle,
_session: &mut Session,
) -> impl std::future::Future<Output = Result<(), Self::Error>> + Send {
async move {
reply.reject(ChannelOpenFailure::AdministrativelyProhibited).await;
Ok(())
}
}
#[tracing::instrument(level = "warn", skip(self, _channel, reply, _session), fields(peer = %self.peer_addr, host = %host_to_connect, port = port_to_connect))]
fn channel_open_forwarded_tcpip(
&mut self,
_channel: Channel<Msg>,
host_to_connect: &str,
port_to_connect: u32,
_originator_address: &str,
_originator_port: u32,
reply: ChannelOpenHandle,
_session: &mut Session,
) -> impl std::future::Future<Output = Result<(), Self::Error>> + Send {
async move {
reply.reject(ChannelOpenFailure::AdministrativelyProhibited).await;
Ok(())
}
}
#[tracing::instrument(level = "warn", skip(self, _channel, reply, _session), fields(peer = %self.peer_addr))]
fn channel_open_x11(
&mut self,
_channel: Channel<Msg>,
_originator_address: &str,
_originator_port: u32,
reply: ChannelOpenHandle,
_session: &mut Session,
) -> impl std::future::Future<Output = Result<(), Self::Error>> + Send {
async move {
reply.reject(ChannelOpenFailure::AdministrativelyProhibited).await;
Ok(())
}
}
#[tracing::instrument(level = "warn", skip(self, _channel, reply, _session), fields(peer = %self.peer_addr, socket = %socket_path))]
fn channel_open_direct_streamlocal(
&mut self,
_channel: Channel<Msg>,
socket_path: &str,
reply: ChannelOpenHandle,
_session: &mut Session,
) -> impl std::future::Future<Output = Result<(), Self::Error>> + Send {
async move {
reply.reject(ChannelOpenFailure::AdministrativelyProhibited).await;
Ok(())
}
}
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn no_sha1_rsa_in_host_key_algorithms() {
let preferred = build_preferred();
for algorithm in preferred.key.iter() {
if let keys::Algorithm::Rsa { hash } = algorithm {
assert!(hash.is_some(), "ssh-rsa SHA-1 must not appear in host key list");
}
}
}
#[test]
fn strict_kex_server_marker_present() {
let preferred = build_preferred();
assert!(
preferred.kex.contains(&russh::kex::EXTENSION_OPENSSH_STRICT_KEX_AS_SERVER),
"Terrapin strict-KEX server marker must be in KEX list"
);
}
#[test]
fn ext_info_server_marker_present() {
let preferred = build_preferred();
assert!(
preferred.kex.contains(&russh::kex::EXTENSION_SUPPORT_AS_SERVER),
"ext-info-s must be in KEX list for server-sig-algs extension"
);
}
#[test]
fn all_ciphers_are_aead() {
// If a non-AEAD cipher is added the MAC list must be reviewed.
let cipher_names: Vec<&str> = SFTP_CIPHERS.iter().map(|c| c.as_ref()).collect();
for name in &cipher_names {
assert!(
name.contains("poly1305") || name.contains("gcm"),
"cipher {} is not AEAD; review MAC list if adding non-AEAD ciphers",
name
);
}
}
#[test]
fn cipher_preference_order() {
// ChaCha20-Poly1305 must be first: constant-time on all hardware.
assert_eq!(SFTP_CIPHERS[0], russh::cipher::CHACHA20_POLY1305);
// AES-256-GCM before AES-128-GCM: prefer larger key size.
assert_eq!(SFTP_CIPHERS[1], russh::cipher::AES_256_GCM);
assert_eq!(SFTP_CIPHERS[2], russh::cipher::AES_128_GCM);
}
#[test]
fn kex_preference_order() {
// Post-quantum hybrid must be first for forward secrecy.
assert_eq!(SFTP_KEX[0], russh::kex::MLKEM768X25519_SHA256);
// curve25519 (RFC 8731) must come before pre-RFC variant.
assert_eq!(SFTP_KEX[1], russh::kex::CURVE25519);
assert_eq!(SFTP_KEX[2], russh::kex::CURVE25519_PRE_RFC_8731);
}
#[test]
fn host_key_algorithm_preference_order() {
// Ed25519 must be first: strongest, fastest, no nonce pitfalls.
assert_eq!(SFTP_HOST_KEY_ALGORITHMS[0], keys::Algorithm::Ed25519);
// RSA must come after ECDSA (ECDSA is smaller and faster).
let first_rsa = SFTP_HOST_KEY_ALGORITHMS
.iter()
.position(|a| matches!(a, keys::Algorithm::Rsa { .. }))
.expect("RSA must be in host key list");
let first_ecdsa = SFTP_HOST_KEY_ALGORITHMS
.iter()
.position(|a| matches!(a, keys::Algorithm::Ecdsa { .. }))
.expect("ECDSA must be in host key list");
assert!(first_ecdsa < first_rsa, "ECDSA must appear before RSA in preference order");
}
#[test]
fn ssh_config_zombie_connection_protection() {
let config = build_ssh_config(Vec::new(), 600, "SSH-2.0-RustFS");
// Idle timeout kills connections with no activity.
assert_eq!(config.inactivity_timeout, Some(std::time::Duration::from_secs(600)),);
// Keepalive probes detect dead TCP connections where the client
// disappeared without sending FIN.
assert_eq!(config.keepalive_interval, Some(std::time::Duration::from_secs(KEEPALIVE_INTERVAL_SECS)),);
assert_eq!(config.keepalive_max, KEEPALIVE_MAX);
}
#[test]
fn ssh_config_has_zero_auth_rejection_delay() {
let config = build_ssh_config(Vec::new(), 600, "SSH-2.0-RustFS");
assert_eq!(
config.auth_rejection_time,
std::time::Duration::from_secs(0),
"auth rejection time must be zero to match S3/FTPS baseline"
);
}
#[test]
fn ssh_config_advertises_only_password() {
let config = build_ssh_config(Vec::new(), 600, "SSH-2.0-RustFS");
assert!(config.methods.contains(&MethodKind::Password));
assert!(!config.methods.contains(&MethodKind::PublicKey));
}
}