* feat(madmin): add account and two-factor wire contract Defines the self-service account and MFA API shapes in one place so the console and the `rc` CLI decode identical payloads instead of each carrying its own copy of the contract. `AccountMutability` is part of the contract on purpose: a client needs to know whether the server will accept a password change for this identity before offering the control, rather than discovering it from a rejected request. * feat(s3-types): add IAM identity audit events Adds `iam:Identity:CredentialChanged` and `iam:Identity:AuthChallenge` so account and authentication activity reaches the audit pipeline in its own namespace, the way the KMS events already do. Neither is reachable from a bucket notification config. Two variants for the whole surface rather than one per operation: `mask()` gives every variant its own bit in a `u64`, and the budget is nearly spent (63 of 64 used after this). The per-operation detail lives in `AuditEntry::api.name` and the `iamOperation` tag, which is what a SIEM filters on anyway. Splitting these further needs `mask()` widened first. * feat(iam): add two-factor authentication primitives Implements the state machine behind TOTP enrollment and verification in the IAM domain, so the admin handlers stay HTTP plumbing and the console and CLI drive identical logic. * `totp`: RFC 6238 over the workspace's existing hmac/sha1, pinned to the published Appendix B vectors. SHA-1, 6 digits, 30s: the parameters every mainstream authenticator app implements. Verification returns the matched time step so the caller can burn it. * `recovery`: ten single-use codes, 100 bits each, in a Crockford base32 alphabet without I/L/O/U. Stored as domain-separated SHA-256 digests — a password KDF would have to run once per stored code on every attempt, turning each guess into an attacker-controlled cost, and with uniform 100-bit input there is no dictionary for it to defend against. * `challenge`: stateless HMAC tokens. A TTL cache would be node-local, so a cluster without session affinity would issue on one node and verify on another; nothing here needs replicating. * `record`: two-phase enrollment, replay high-water mark, and lockout. Pending enrollment never gates a login, so a mis-scanned QR cannot lock an operator out, and re-configuring keeps the old factor working until the new one is confirmed. * `store`: one object per identity under `config/mfa/`, a sibling of `config/iam/` so the IAM cache loader's startup walk does not sweep it up. Optimistic `If-Match` writes; deliberately uncached, because a cache would need cluster-wide invalidation to keep the replay mark and the lockout counter honest. * `qr`: server-side rendering, so neither client needs a QR encoder. Enrollment is refused without `RUSTFS_IAM_MASTER_KEY`. A TOTP secret is credential-equivalent, and one written in plaintext could be lifted off a disk — worse than no second factor, because the user believes they have one. IAM identities tolerate a missing master key for backward compatibility; a new feature has no such history to honour. Also adds `IamSys::revoke_sts_sessions_for_parent`, so a credential rotation can invalidate the sessions minted under the old secret. * feat(admin): add self-service account endpoints and the two-factor login gate Adds the account surface (`/v3/account/*`), the second-factor endpoints, the administrative reset (`/v3/user/mfa`), and `PUT /v3/set-user-secret-key`, plus the gate on `AssumeRole`. What the gate covers, and what it deliberately does not: * `AssumeRole` is the only interactive login RustFS has, so it is where a second factor can be enforced. With one enrolled it requires `TokenCode`; without an enrollment the code path is unchanged, so existing deployments are untouched. * A request signed directly with a long-term access key stays ungated. Gating it would break every script and CLI the moment a human enabled 2FA on their own account, and would add no protection: whoever holds the secret key already has full access without presenting a code. This is the division AWS draws; making 2FA meaningful for API access needs an `aws:MultiFactorAuthPresent` policy condition, tracked separately. `SerialNumber`/`TokenCode` are STS's own parameters, so an SDK or script authenticates the same way the console does. `caller_identity` resolves who a request acts as. The console signs with a short-lived STS session, so "the caller" is almost never the key that signed. It reports two separate capabilities: root cannot rotate its secret (a process-wide `OnceLock` that also derives the internode RPC secret) but *can* enroll a second factor — conflating the two would leave the default deployment's console login unprotectable. The self-service routes carry no admin action. Giving them one would be wrong in both directions: it would stop an ordinary user from changing their own password, and let any holder of that action change someone else's. They gate on possession of the credential plus, for the mutations, knowledge of the current secret — a signature only proves a credential was used, so without that a hijacked tab could rewrite the account's credentials or strip its second factor. `set-user-secret-key` exists because the only prior way to change a password was to re-POST the whole user through `add-user`, which rewrote `status` and dropped the policy field — a password reset that silently re-enabled a disabled account. Wrong, replayed and malformed codes are indistinguishable on the wire; the distinction survives only in the audit trail, where no submitted value, secret or code is ever recorded. * test(e2e): cover the two-factor lifecycle and its regressions Unit tests cover the state machine at its edges; only an end-to-end test proves the pieces are wired together and that the existing authentication paths still behave. Asserts, against a real server: enrollment is refused without a master key; the full enroll/activate flow works with a genuine RFC 6238 code; `AssumeRole` refuses without a factor and accepts a valid one; a recovery code works exactly once; a direct SigV4 admin request keeps working with a factor enrolled; `AssumeRole` for an unenrolled identity is unchanged; and a password rotation invalidates the old secret. The test computes TOTP codes itself rather than calling the server's implementation — a shared helper could agree with a bug on both sides. This suite caught a real defect during development: enrollment was refused for root because its *password* is immutable, which would have left the default deployment — an administrator signing into the console as root — unable to protect the one login the feature exists for. * docs(operations): document the two-factor authentication model Records what the second factor protects and what it deliberately does not, because several of the boundaries look like gaps until the alternative is spelled out: why direct SigV4 access stays ungated, why root credentials cannot be rotated at runtime, why secret keys cannot be hashed in an S3 server, and why at-rest protection is mandatory for a TOTP secret but optional for an IAM identity. Also states the limitations plainly, including that GHSA-m77q-r63m-pj89 is unaffected: a holder of the root secret can still forge a session token, 2FA claim included. Placed alongside the other authentication and KMS security documents rather than under a new `docs/security/`, which `.gitignore` excludes. * fix(admin): route the new account handlers through the admin s3 facade Two of the guardrails in the CI "Quick Checks" job rejected the previous commits, so the required check would have gone red as soon as a maintainer approved the workflow run. `check_architecture_migration_rules.sh` requires everything under `rustfs/src/admin` to reach `ECStore` through a domain module rather than the root of `storage_api`. The MFA handler and the two `AssumeRole` signatures now use `storage_api::runtime::ECStore`, which is where the other ten admin handlers already take it from. `check_s3s_footprint.sh` ratchets two counters that new code may not grow: files referencing `s3s` and error-macro invocation lines. This branch added four files and thirty-two lines to them. The ratchet is lower-only and its header forbids raising a baseline to get green, so the construction moves behind the facade instead: `storage_api::s3` now re-exports the request and body types these handlers need and gains an `error` constructor over `S3Error::with_message`. That is the same constructor the macro expands to and the one `handlers/mod.rs`, `rebalance_internal_error` and `invalid_object_lock_configuration` already call, so this is the existing practice rather than a new one, and it keeps the `s3s` dependency in the boundary file the s3gate migration replaces. Every error code and message is carried over unchanged. In `sts.rs` only the call site this branch added is converted; the sixteen that predate it are left alone, because rewriting them would put unrelated churn in a feature PR and push the counter below the baseline it is meant to hold.
e2e_test
End-to-end test suite for RustFS. Each test spawns a real rustfs binary
(built on demand from the workspace) and drives it over the network with the
AWS SDK (aws-sdk-s3), raw HTTP (reqwest / awscurl), or a protocol client
(FTPS / WebDAV / SFTP). This is the black-box integration layer: exhaustive
end-to-end behavior lives here, unit behavior stays in the source crates
(see AGENTS.md).
The harness lives in src/common.rs (single-node +
cluster environments, S3 client construction, awscurl helpers) and
src/chaos.rs (in-process disk fault injection). Crate-wide
test conventions and environment-safety rules are in
AGENTS.md; this file is the contributor guide.
Module map (~50 modules)
Registered in src/lib.rs. Grouped by concern:
| Group | Location | What it covers |
|---|---|---|
| functional | top-level *_test.rs |
S3 data plane: list_objects_*, copy_object_*, delete_objects_versioning, head_object_*, checksum_upload, compression, content_encoding, special_chars, leading_slash_key, create_bucket_region, quota, data_usage, snowball_auto_extract, mc_mirror_small_bucket, archive_download_integrity, version_id_regression, delete_marker_migration_semantics |
| object_lock | src/object_lock/ |
Retention / legal-hold / WORM semantics |
| kms | src/kms/ |
SSE-S3 / SSE-KMS / SSE-C, local + Vault backends, multipart encryption. Own guide: src/kms/README.md |
| policy | src/policy/, existing_object_tag_policy_test, bucket_policy_check_test, anonymous_access_test, security_boundary_test, multipart_auth_test |
IAM / bucket-policy / STS session policy, policy variables, anonymous access, DoS/SSRF boundaries. Own guide: src/policy/README.md |
| protocols | src/protocols/ |
FTPS, WebDAV, SFTP compliance. Fixed ports, own guide: src/protocols/README.md |
| reliant | src/reliant/ |
Tests that reuse an externally started server (SQL/select, conditional writes, lifecycle, deleted-object reads, node-interact). Run via scripts/run_e2e_tests.sh; see src/reliant/README.md |
| cluster | cluster_concurrency_test, stale_multipart_cleanup_cluster_test, namespace_lock_quorum_test, admin_timeout_regression_test, object_lambda_test, replication_extension_test |
Multi-node scenarios via RustFSTestClusterEnvironment |
| chaos / reliability | src/chaos.rs, reliability_disk_fault_test, heal_erasure_disk_rebuild_test, server_startup_failfast_test |
Disk offline/replace/corrupt, EC rebuild, heal, fail-fast startup |
| upgrade compatibility | upgrade_compatibility_test |
Pinned previous-release writes followed by current-build reads on the same data directory |
How to run
All commands assume repo root. cargo test triggers an on-demand build of the
rustfs binary from src/common.rs (rustfs_binary_path) on
first use — the first invocation is slow, later ones reuse the binary.
# Whole crate (default = ignored tests skipped)
cargo nextest run -p e2e_test
# One module
cargo nextest run -p e2e_test -E 'test(list_objects_v2_pagination_test)'
# PR smoke subset (see "CI smoke subset" below)
cargo nextest run --profile e2e-smoke -p e2e_test
# ILM serial lane — ignored lifecycle tests, single-threaded (mirrors CI)
cargo nextest run -j1 --run-ignored ignored-only -p rustfs-scanner -p rustfs \
-E 'binary(lifecycle_integration_test) or (package(rustfs) and test(lifecycle_transition_api_test))'
The protocols suite has its own contract (fixed bind ports 9022–9301,
single-worker execution, feature-gated scheduling) documented in
src/protocols/README.md. RUSTFS_BUILD_FEATURES
selects which features the spawned binary is built with; leave it unset to run
every protocol entry. Use the exact profile command under
Troubleshooting for CI-equivalent execution.
#[ignore] semantics
Ignored tests are excluded from the default cargo nextest run pass because
they need something the default runner does not provide. Do not maintain a
static count here — it rots (the set shrinks as ci-13 / ilm-3 activate
suites). Read the live sources instead:
rg -n '#\[ignore' crates/e2e_test/src # every ignore + its reason string
The reason string on each attribute is the classifier. Current classes:
- Needs a pre-started server —
"requires running RustFS server at localhost:9000"/"Connects to existing rustfs server". These are thereliant/*tests; start a server first (e.g.scripts/run_e2e_tests.sh) or use--run-ignored. - Heavy / external tool —
"Starts a rustfs server; enable when running full E2E","requires awscurl and spawns a real RustFS server". Spawn their own server and/or needawscurlonPATH. - Serial / global-state (ILM lane) — lifecycle tests bind fixed ports and share process-global singletons; run via the ILM serial lane above.
How to add a test
Single-node (the common case)
Use RustFSTestEnvironment from src/common.rs. It picks a
random free port and a unique temp dir per instance, so tests are
parallel-safe by construction and clean up on Drop:
use crate::common::{RustFSTestEnvironment, TEST_BUCKET};
#[tokio::test]
async fn my_case() -> Result<(), Box<dyn std::error::Error + Send + Sync>> {
let mut env = RustFSTestEnvironment::new().await?;
env.start_rustfs_server(vec![]).await?; // waits for readiness
let client = env.create_s3_client(); // aws-sdk-s3 Client
env.create_test_bucket(TEST_BUCKET).await?;
// ... drive `client` ...
Ok(())
}
Register the module in src/lib.rs under #[cfg(test)].
Cluster
Use RustFSTestClusterEnvironment::new(node_count) then .start(); it spawns
node_count servers over a shared erasure set and hands out per-node S3 clients
via create_s3_client(idx) / create_all_clients(). See
cluster_concurrency_test.rs and namespace_lock_quorum_test.rs for patterns.
Fixture / helper inventory (src/common.rs)
| Helper | Purpose |
|---|---|
RustFSTestEnvironment::new / with_address |
Single-node env; random or fixed address |
start_rustfs_server / _with_env / _without_cleanup |
Spawn the server (optional extra args / env vars / no pre-cleanup) |
wait_for_server_ready |
Poll readiness before issuing requests |
create_s3_client / create_test_bucket / delete_test_bucket |
aws-sdk-s3 client + bucket lifecycle |
find_available_port |
Random free port (isolation primitive) |
rustfs_binary_path / _with_features |
Locate/build the binary; honors RUSTFS_BUILD_FEATURES |
requested_rustfs_build_features / rustfs_build_feature_enabled |
Feature-gate a test to what the binary was built with |
execute_awscurl / awscurl_post / _get / _put / _delete / awscurl_post_sts_form_urlencoded |
Admin/STS API calls via awscurl; missing binaries are test failures |
replication_fast_env |
Env vars that shrink replication timers (from repl-4); pass to start_rustfs_server_with_env |
local_http_client / init_logging |
Loopback HTTP client; idempotent tracing init |
RustFSTestClusterEnvironment (new/start/start_node/stop_node/create_all_clients) |
Multi-node harness |
Constants: DEFAULT_ACCESS_KEY, DEFAULT_SECRET_KEY, TEST_BUCKET, ENV_RUSTFS_BUILD_FEATURES |
Shared credentials / bucket name / env-var name |
Fault injectors live in src/chaos.rs: DiskFaultHarness
(take_disk_offline, bring_disk_online, replace_disk_with_empty,
corrupt_object_shard, object_metadata_exists_on_disk, kill_server /
restart_server) plus signed_admin_post.
Isolation rules
- Port: never hard-code a port for single-node tests —
new()allocates a random one. Fixed ports (protocols, ILM lane) force--test-threads=1/ a serial CI lane. - Temp dir: each env owns a temp dir cleaned on
Drop; do not write under a shared path. - Orphans:
RustFSTestEnvironmentkills its child onDrop, but a panicked orkill -9'd run can leak arustfsprocess holding a port — see Troubleshooting.
#[serial] vs nextest reality
serial_test's #[serial] uses an in-process mutex. Under nextest each
test runs in its own process, so #[serial] does not serialize across
tests there — see the header of .config/nextest.toml.
Real cross-test serialization comes from a nextest test-group (max-threads = 1) or a -j1 CI lane. Single-node e2e tests should instead be parallel-safe by
construction (random port + isolated temp dir) and need no serialization.
CI map
e2e_test is excluded from the main cargo nextest run --profile ci --all
pass (--exclude e2e_test) — the whole crate is too slow to gate every PR.
Subsets join CI through nextest profiles; the fixed-port protocol suite uses
the same profile for membership and execution with one nightly worker.
| Suite | Runs where | Status |
|---|---|---|
Smoke subset (e2e-smoke profile) |
e2e-tests job, every PR |
Active (backlog#1149 ci-4) |
Full single-node suite (e2e-full profile) |
e2e-full job, merge queue + main |
Active (backlog#1149 ci-5) |
s3s-e2e black-box |
e2e-tests + e2e-tests-rio-v2 jobs |
Active (external conformance tool) |
| ILM / lifecycle (ignored) | test-ilm-integration-serial lane, -j1 |
Active (backlog#1148 ilm-1) |
| KMS suite | e2e-full job, merge queue + main |
Active |
| Direct upgrade from pinned previous release | e2e-upgrade.yml, storage-sensitive PRs + release tags + weekly |
Active |
Cluster faults (e2e-nightly profile) |
consolidated nightly workflow | Active (backlog#1149 ci-7) |
| Protocols (FTPS/WebDAV/SFTP) | consolidated nightly workflow, serial | Active (backlog#1149 ci-7) |
| Replication (fast subset) | e2e-smoke profile, e2e-tests job, every PR |
Active (backlog#1147 repl-1) |
| Replication (slow + multi-node) | e2e-repl-nightly profile, consolidated nightly workflow |
Active (backlog#1147 repl-1) |
reliant/* |
19 tests in PR smoke; remaining default tests in e2e-full |
Active except #[ignore] |
The profile filters in .config/nextest.toml are
the wiring source of truth. Committed test-ID digests under
.config/e2e-*-selection.txt make every membership change explicit.
Troubleshooting
Reproduce a CI failure locally — run the exact profile/lane:
# Smoke (e2e-tests job) — includes the 20 fast replication tests
cargo nextest run --profile e2e-smoke -p e2e_test
# Full single-node merge/main lane
cargo nextest run --profile e2e-full -p e2e_test
# Cluster fault nightly lane
cargo nextest run --profile e2e-nightly -p e2e_test
# Replication nightly lane; awscurl is required for STS paths
cargo nextest run --profile e2e-repl-nightly -p e2e_test
# Fixed-port protocol nightly lane
RUSTFS_BUILD_FEATURES=ftps,webdav,sftp \
cargo nextest run -j 1 --profile e2e-protocols -p e2e_test --no-capture
# ILM serial lane
cargo nextest run -j1 --run-ignored ignored-only -p rustfs-scanner -p rustfs \
-E 'binary(lifecycle_integration_test) or (package(rustfs) and test(lifecycle_transition_api_test))'
# s3s-e2e black box
./scripts/e2e-run.sh ./target/debug/rustfs /tmp/rustfs-e2e-data
Stale binary. Tests build the rustfs binary once and reuse it. To avoid
rebuilding while iterating on tests, common.rs reuses an existing binary when
running inside the e2e test process even if sources changed
(can_reuse_inside_e2e, src/common.rs line 98). Downside: if
you changed server code, force a rebuild with
cargo build -p rustfs (or touch a source file outside the reuse window)
before re-running, or CI's freshly built artifact will diverge from your local
one.
Port already in use / orphan processes. A hard-killed run can leak a
rustfs child holding its port. Find and kill it:
pkill -f 'target/debug/rustfs' ; pkill -f 'target/release/rustfs'
The s3s-e2e CI job selects a random RUSTFS_TEST_PORT (see the e2e-tests
job) to dodge this; local single-node tests already use random ports, so a
lingering orphan is usually the cause of a spurious bind failure.
awscurl not found. awscurl-dependent tests fail closed with a process
spawn error. Install the pinned CI version before running their profiles.
Related
- Crate rules & environment safety:
AGENTS.md - Sub-suite guides:
src/kms/README.md,src/policy/README.md,src/protocols/README.md,src/reliant/README.md - Authoritative per-module counts:
docs/testing/e2e-suite-inventory.md - Test pyramid & flake policy:
docs/testing/README.md
CI smoke subset (--profile e2e-smoke)
A subset of this crate runs on every PR via the e2e-tests job:
cargo nextest run --profile e2e-smoke -p e2e_test
The selection lives in .config/nextest.toml under [profile.e2e-smoke]
(default-filter). That filter is the single wiring mechanism for e2e
tests in CI — extend it (or add a sibling profile) instead of adding new e2e
jobs to ci.yml.
Admission criteria for the smoke subset
A test module may join the smoke filter only if every test in it is:
- Fast — single-digit seconds per test; the whole subset must keep the
e2e-testsjob ≤ 20 minutes. - Single-node — spawns its own server via
RustFSTestEnvironment/start_rustfs_serveron a random port with an isolated temp dir. NoRustFSTestClusterEnvironment, no fixed ports. - Hermetic dependencies — no pre-started server at
localhost:9000, no Vault, and no fixed protocol ports. Any required CLI must be pinned and installed by the workflow; a missing CLI must fail the test. - Not
#[ignore]— ignored tests are activation work (backlog#1149 ci-13 / backlog#1148 ilm-3), not smoke candidates.
Note on #[serial]: nextest runs each test in its own process, so
serial_test's in-process mutex does not serialize across tests there
(see the header of .config/nextest.toml). Smoke tests must therefore be
parallel-safe by construction (random port + isolated temp dir), which the
current subset is.
Authoritative test inventory
docs/testing/e2e-suite-inventory.md records the per-module test counts as
listed by cargo nextest list -p e2e_test. Regenerate it when adding or
moving e2e tests so acceptance numbers in the test-strategy issues
(backlog#1147–#1155) stay auditable. When a profile membership change is
intentional, review its JSON listing before updating the matching
.config/e2e-*-selection.txt test-ID digest. Update only the platform that
produced the listing:
python3 scripts/check_test_wiring.py --update-profile e2e-full /path/to/listing.json linux