* fix(iam): require an explicit permission for force-delete
A force-delete header no longer inherits s3:* or consoleAdmin. Bucket
force-delete requires s3:ForceDeleteBucket whenever the header is present,
and recursive object force-delete requires s3:ForceDeleteObject. A plain
delete keeps the existing checks.
Co-authored-by: RustFS <hello@rustfs.com>
Signed-off-by: loverustfs <155562731+loverustfs@users.noreply.github.com>
* test(e2e): keep force-delete header names static
The bucket force-delete helper must pass a static header name into the
SDK request mutator.
Co-authored-by: RustFS <hello@rustfs.com>
Signed-off-by: loverustfs <155562731+loverustfs@users.noreply.github.com>
* test(e2e): move the force-delete header into the request mutator
The SDK request customizer requires a static header name owned by the
closure.
Co-authored-by: RustFS <hello@rustfs.com>
Signed-off-by: loverustfs <155562731+loverustfs@users.noreply.github.com>
* fix(iam): keep force-delete out of NotAction grants
NotAction now uses plain wildcard matching, so NotAction "s3:*" still
excludes force-delete. An Allow statement grants s3:ForceDeleteObject or
s3:ForceDeleteBucket only when its Action list names the action; a
NotAction-only Allow never does. The rule applies to both IAM and bucket
policy statements.
Also build the invalid-header errors with S3Error::with_message to keep
the s3s footprint at its baseline, and fix a clippy single_match.
---------
Signed-off-by: loverustfs <155562731+loverustfs@users.noreply.github.com>
Co-authored-by: Hauser <housemecn@gmail.com>
Co-authored-by: overtrue <anzhengchao@gmail.com>
Record the upload id on the completed object so a lost-response retry
returns that object's ETag instead of NoSuchUpload, while a different
part list or a replaced object keeps the existing errors.
Signed-off-by: loverustfs <155562731+loverustfs@users.noreply.github.com>
Co-authored-by: Hauser <housemecn@gmail.com>
* fix(ci): restore mainline and scheduled test reliability
* fix(ci): provide GitHub CLI for CPU acceptance
* ci: provide Docker for CPU service acceptance
* ci: provision Python and Docker for OIDC validation
* ci: restore hosted runners for Docker validation
* ci: use verified MinIO release packages for interop
* ci: preserve host ownership of MinIO fixtures
* fix(ci): correct diagnostic limits and isolate startup checks
* test(connect): include object CLI failure details
* test(readiness): initialize unavailable drive diagnostics
Bound the release catalog to 32 MiB and 200,000 symbols based on the GNU build measurement; keep complete names and reject catalogs beyond either limit.
A restarted peer can accept a pooled connection and never send response
headers, so HttpReader::open waited past the client body timeout before
the body-stall timer or erasure hedge could run. Bound that header wait
by the stall timeout and retry the open once on a fresh connection.
After write quorum, MultiWriter still waited out the full disk stall
for a silent peer, which matches the client timeout. Give remaining
writers one second, then drop them so the caller returns.
Signed-off-by: loverustfs <155562731+loverustfs@users.noreply.github.com>
EC 2+2 still meets read quorum with two of four nodes up. A survivor
that had not cached a bucket returned 503 on GET, HEAD, and List, and
/health/ready left the Service once write quorum was lost. Reads and
Service membership now follow read quorum and shared locks. Writes and
/minio/health/cluster still require write quorum and exclusive locks.
Signed-off-by: loverustfs <155562731+loverustfs@users.noreply.github.com>
Co-authored-by: Hauser <housemecn@gmail.com>
* fix(storage): weight automatic multipart admission by part size
* fix(ci): restore filesystem runner capabilities and typos dependency
* fix(ci): make release guard portable and spell out part variables
* ci: restore sm-standard-2 runners for io_uring and distributed e2e jobs
---------
Co-authored-by: hector <42570491+majinghe@users.noreply.github.com>
Co-authored-by: Hauser <housemecn@gmail.com>
* fix(ecstore): admit cold bucket metadata at read quorum
Encode a durable creation-commit state in bucket metadata so committed Object Lock buckets can use read quorum for cold loads while pre-physical creation intents remain fail-closed. Drain commit-write fan-out, keep write quorum for uncommitted intents, and add regression coverage for exact, below, migrated, and partial-create quorum boundaries.
* fix(ecstore): fence bucket creation-commit persistence
Address review findings on the creation-commit proof.
Persist the proof only under the bucket metadata transaction fence: at bucket creation, and through a fenced migration that re-reads the authoritative metadata and revalidates physical presence at write quorum before writing. This stops a stale snapshot from reverting an acknowledged configuration update or outliving a delete/recreate.
Establish commitment when Object Lock is enabled on an existing bucket, inside the same configuration mutation, so a healthy cluster no longer rejects object operations with ErasureWriteQuorum.
* fix(ecstore): keep commit fence error message stable
The error(format!) ratchet requires a stable Display for quorum bucketing; use a fixed message instead of embedding the bucket name.
* test(ecstore): reuse canonical Object Lock fixture in regression tests
The s3s footprint ratchet is shrink-only. Use the existing ENABLED_OBJECT_LOCK_CONFIG static instead of naming s3s DTO types in store tests.
---------
Co-authored-by: hector <42570491+majinghe@users.noreply.github.com>
Co-authored-by: Hauser <housemecn@gmail.com>
* fix(scanner): bound SNSD deep scans and checkpoint cloning
* ci: run mount-dependent jobs on hosted VMs
---------
Co-authored-by: hector <42570491+majinghe@users.noreply.github.com>
A flat (recursive) listing resumed inside "s/" re-emitted everything in
its sibling "s-x/". scan_dir dropped the entries before forward_to by
comparing directory names without their trailing slash, where "s" sorts
before "s-x", while the keys they stand for sort the other way round:
'-' (0x2d) is below '/' (0x2f), so all of "s-x/..." precedes "s/...".
The drain stopped at "s" and kept "s-x".
The next page then started with keys at or before the marker; the
listing layer filtered all of them out, found no more candidates and
answered IsTruncated=false. A bucket of 104,137 objects with backup
directories named "<id>" and "<id>-rollbacks" listed as 5,000; two such
directories of 1,200 keys each listed as 2,000.
Compare every remaining entry as the key prefix it stands for, slash
included, and keep it only when it sorts at or after forward_to or
contains it. The remainder of forward_to is taken before `current` is
trimmed, so this also holds below the bucket root; retain also drops
entries when all of them precede forward_to, which the drain never did.
* fix(startup): retry transient bucket metadata quorum failures
* fix(ecstore): retain recovered shard damage for read repair
* fix(ecstore): reconnect missing read sources before GET
* test(heal): prove exact marker replay preserves history
* fix(ecstore): fence reconnect retries at the deadline
Reject expired reconnect waiters before dispatch and start cooldown on sweep completion, including timeout and cancellation. Cover exact-deadline admission independently of cooldown and preserve the original strict RPC-count regression.
* fix(ecstore): bound decommission target gate contention
Target capacity gate contention during pool decommission escalated a
per-object transient into a durable bucket pause: each contended object
failed its bucket entry, which re-ran the whole bucket listing and amplified
attempts on the same objects.
- centralize the decommission capacity failure classification so gate
contention, benign contention and fatal failures are decided once
- retry target-gate contention inline (bounded, jittered) before a mutation
is admitted, covering put, part, complete, new-multipart and abort
- defer contended entries to the end of the round instead of failing the
bucket entry, and require the deferred set to drain before a set completes
- treat a missing object or version, an overwrite and a superseded upload id
as benign contention that is neither counted as a failure nor escalated
- use full-jitter exponential backoff, capped, for decommission retries
- expose the capacity pause reason, the waiting reason and a cumulative
pause count in the admin pool status, plus gate-retry and per-object
attempt metrics
Related: rustfs/backlog#2644
* fix(ecstore): correct deferred replay and metadata compatibility
* fix(ecstore): bound decommission capacity gate waits
* fix(ecstore): keep rebalance meta save lock errors retryable
The rebalance metadata retry policy only retries when the source chain still
carries the typed lock error, but resolve_rebalance_meta_save_result collapsed
every failure into a plain string. A transient rebalance.bin write-lock timeout
therefore bypassed retry_rebalance_metadata_access entirely and surfaced as a
hard failure of the periodic and terminal rebalance metadata saves.
Wrap the failure with the data-movement stage context instead, which keeps the
original lock error reachable through rebalance_error_source so the existing
transient lock policy applies. The rendered message is unchanged.
Add a regression test that feeds the wrapped meta save lock timeout through the
retry helper and asserts the retry engages.
* chore(ecstore): refresh error format ratchet baseline
Removing the string-wrapped meta save failure drops one `::other(format!`
call site in the rebalance worker, so the shrink-only baseline must be
regenerated in the same change.
JSON log layers synthesized span lifecycle events for every span and
embedded the full ancestor span list on every record. With INFO spans on
object hot paths that produced multi-megabyte lines, which journald
truncated at 48 KiB: the payload was silently dropped and the volume
itself became the dominant load.
Add RUSTFS_OBS_LOG_SPAN_EVENTS (none|close|full) with a default of none,
and keep with_span_list enabled only for verbose (debug/trace) levels so
the ancestor chain stops being re-serialized per record. The request_id
promotion path reads the span scope, not the span list, so it is
unaffected.
Refs rustfs/backlog#2642