mirror of
https://github.com/rustfs/rustfs.git
synced 2026-10-04 12:31:36 +00:00
7192eecbfbcc90bd7bcfd857a53c73cb9f3d9d41
14 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
4e16705873 |
fix(release): backport main fixes and stabilize tier cleanup tests (#7793)
* fix(ecstore): stop pruning at nonempty directories (#7616) * fix(ecstore): stop pruning at nonempty directories * test(ecstore): release pruning fixtures before temp cleanup (cherry picked from commit |
||
|
|
7c85c72fd1 | docs: scope agent guidance and consolidate review workflows (#7419) | ||
|
|
37bda24e1c |
fix: reject unsupported pool expansion with actionable errors (#7360)
* fix: reject unsupported pool expansion with actionable errors Report singleton-pool and persisted-topology constraints before misleading startup retries. Preserve single-node multi-drive admission and existing parity policies, and cover format preservation plus operator recovery guidance for issue #6186. * fix: keep pool layout errors typed Preserve actionable pool layout diagnostics without adding generic formatted errors. Tighten the shrink-only baseline and assert that both typed payloads survive the I/O boundary. --------- Co-authored-by: houseme <housemecn@gmail.com> |
||
|
|
1dddf357cd | test(scanner): add bounded cache cost microprofile (#7261) | ||
|
|
07833379b4 |
test(e2e): add distributed cluster regression coverage (#7158)
* test(e2e): add distributed 4x4 validation * test(e2e): prove operations overlap data movement --------- Co-authored-by: Zhengchao An <anzhengchao@gmail.com> |
||
|
|
2e4ab045b6 |
test(scanner): add durable checkpoint diagnostics (#7175)
* chore(deps): refresh SDKs and pin clock skew regression coverage Refresh compatible dependencies for Scanner/Heal V2 batch 1 and verify the production S3 retry/signing path with a deterministic clock. Co-Authored-By: heihutu <heihutu@gmail.com> Co-Authored-By: zhi22915 <qiuzgang@gmail.com> * test(scanner): add durable checkpoint diagnostics Refs rustfs/backlog#2260 and rustfs/backlog#2240. Co-Authored-By: heihutu <heihutu@gmail.com> Co-Authored-By: zhi22915 <qiuzgang@gmail.com> --------- Co-authored-by: heihutu <heihutu@gmail.com> Co-authored-by: zhi22915 <qiuzgang@gmail.com> |
||
|
|
8023cf3e26 |
test(e2e): add the outbound target matrix and #7082 postmortem (#7092)
test(e2e): add the outbound target matrix and the replication checksum postmortem Defense work for rustfs#7082, the regression rustfs#6895 introduced while fixing rustfs#6853: a fix for one target class changed a client default for every target class and nothing in tree modeled the other classes. - docs/postmortems: timeline, root cause, why four defense layers missed it, and the SOP for changing any outbound client default; AGENTS.md and the adversarial compatibility lens point at it; the two env knobs from rustfs#6895 are documented in docs/operations. - fake_s3_target: reject_aws_chunked_uploads, require_checksum_for_object_lock (Content-MD5 always verified), create_bucket_with_object_lock with a GetObjectLockConfiguration handler, and a TransportSnapshot on every journal record. - replication_target_matrix_test: six object shapes against four target modes with an explicit expectation table; the two rustfs#7082 cells are pinned KnownFailing and fail with an XPASS message once the fix lands. Wired into e2e-repl-nightly, excluded from e2e-full. |
||
|
|
0a975f2fe2 | docs(knowledge-base): prune stale content and add agent-facing index (#7035) | ||
|
|
6f6dd19cc7 | ci(coverage): add security ratchet calibration (#6388) | ||
|
|
eb392f24d6 |
chore(scripts): add scripts index and archive one-shot scripts (#4822)
chore(scripts): index scripts/ and archive 29 one-shot scripts backlog#1153 infra-13. scripts/ had 80+ unlabelled top-level entries mixing CI gates with finished one-shot issue-validation scripts. - scripts/README.md — one index row per entry with status (ci-gate / dev-tool / archived), purpose, and wiring; subdirectories get one row each. run_scanner_benchmarks.sh is annotated "disposition owned by backlog perf-10" and deliberately untouched. - git mv 29 confirmed-stale one-shot entries to scripts/archive/: 11 issue-scoped validation/perf-capture scripts, the 5-script backlog#706 large-PUT breakdown family, the 4-file GET-optimization stress suite, 2 gt1g one-shots, and 7 other orphaned one-shots. Evidence: a whole-tree boundary-aware reference census showed zero references from CI/Makefiles/docs/code for every moved entry (or references only from other scripts inside the same archived set); re-run after the move shows zero dangling references. - docs/testing/README.md links the index. Moves only — no script content changed. |
||
|
|
5fd2e8b6a1 |
ci(coverage): add weekly cargo-llvm-cov baseline workflow (#4820)
ci(coverage): weekly cargo-llvm-cov workspace baseline (non-blocking) Add the weekly line-coverage report (backlog#1153 infra-5): - .github/workflows/coverage.yml — Sunday 07:00 UTC + workflow_dispatch; runs `cargo llvm-cov nextest --workspace --exclude e2e_test` under NEXTEST_PROFILE=ci (same scope and profile as the ci.yml test gate), writes a per-crate line-coverage table to the job summary, uploads lcov + JSON as a 90-day artifact, and routes scheduled failures through the shared schedule-failure-issue action (ci-8). Runs only on schedule/dispatch, so it can never become a required PR check. - scripts/coverage_per_crate.py — stdlib-only aggregation of the llvm-cov JSON export into the per-crate markdown table (worst-first, TOTAL row), shared by the workflow and the local target. - make coverage (.config/make/coverage.mak) — local equivalent with the same command sequence; fails with install hints when cargo-llvm-cov or cargo-nextest are missing. Listed in make help. - docs/testing/README.md — Coverage section: cadence, where the table and artifacts live, the trend-comparison method, and what is not measured (doctests, e2e_test). |
||
|
|
cd51d66321 |
docs(testing): populate the testing pyramid overview (backlog#1153 infra-11) (#4813)
docs(testing): populate testing pyramid, naming, serial/nextest rules Fill the docs/testing/README.md skeleton (backlog#1153 infra-11): - Test taxonomy table for all eight layers (unit / ecstore black-box / e2e / s3s-e2e / S3 compatibility / chaos / fuzz / bench) with a verified entry command and a qualitative "when it runs" per layer; the event x timeout x required-status matrix stays owned by docs/testing/ci-gates.md (ci-15), linked not duplicated. - Naming conventions section with the migration-gate reserved substrings (data_movement / rebalance / decommission / source_cleanup / delete_marker) linking the infra-12 count-floor guard, closing that task's docs cross-link. - Serial execution & nextest rules: nextest as a hard dependency with the RUSTFS_ALLOW_CARGO_TEST_FALLBACK escape hatch and the runner-semantics difference (folds the infra-14 README half), why #[serial] is a no-op under nextest, and the default/ci/e2e-smoke/e2e-repl-nightly profiles. - A time-control placeholder for infra-4 to fill. The pre-existing flake-policy section (ci-10) is preserved verbatim. Add pointers from CLAUDE.md and CONTRIBUTING.md. |
||
|
|
846aa95c32 |
test(security): GHSA-named regression tests for 3p3x and r5qv (backlog#1151 sec-6) (#4707)
test(security): add GHSA-named regression tests for 3p3x and r5qv (backlog#1151 sec-6) Anchor the two fixed advisories to discoverable, named regression tests so `rg -i "ghsa|3p3x|r5qv"` finds a guard for each, and future fixes are forced to update pinned behavior (red -> green). GHSA-3p3x-734c-h5vx (constant-time WebDAV/FTPS secret comparison, rustfs#4403): - ftps_core.rs: new `assert_ftps_ghsa_3p3x_wrong_credentials_rejected` drives the `ct_eq` reject branch in FtpsAuthenticator::authenticate; asserts wrong password and unknown user are both rejected (530) and indistinguishable. - webdav_core.rs: the auth-failure block now sends a valid access key with a wrong secret (exercising the `ct_eq` branch, not just the unknown-access-key path) plus an unknown user, asserting both 401 and indistinguishable. - Module doc comments map advisory -> tests -> fix PR on both files. GHSA-r5qv-rc46-hv8q (internode RPC fail-closed, rustfs#4402): - http_auth.rs: renamed the default-fallback rejection test to `ghsa_r5qv_resolve_shared_secret_rejects_default_fallback` (and broadened it to cover default env secret + blank secrets), and added `ghsa_r5qv_verify_rpc_signature_fails_closed_on_missing_or_invalid_auth` pinning the exact advisory scenario (missing/forged/cross-URL signature is rejected; a correctly signed request still passes). File-level doc maps the advisory. Docs: new docs/testing/security-regressions.md with the advisory -> test mapping table and where each layer runs; linked from docs/testing/README.md. sec-14 will formalize the written policy in AGENTS.md. The unit-level ghsa_r5qv_* tests run in the default CI pass. The WebDAV/FTPS e2e live in the protocols suite (fixed ports, --test-threads=1); they cannot join the e2e-smoke profile and are wired into CI by sec-5. Refs: rustfs/backlog#1151 (sec-6), rustfs/backlog#1155 |
||
|
|
89ea931ee1 |
test(ci): strict nextest ci profile with quarantine + flake policy (#4666)
test(ci): add strict nextest ci profile with quarantine + flake policy Formalize the existing ecstore-serial-flaky mechanism into a strict CI gate (ci-10, absorbs infra-15; backlog#1149). - .config/nextest.toml: add [profile.ci] with global retries=0 (never mask a new race's first occurrence), fail-fast=false, and JUnit output at target/nextest/ci/junit.xml. Add a quarantine section where flaky tests get retries=2 under the ci profile only; each entry links one OPEN issue. First members are the two backlog#937 ecstore groups (concurrent_resend_same_part_commits_one_generation and store::bucket::tests::bucket_delete_*), which keep their existing ecstore-serial-flaky test-group serialization. Local default profile still never retries. - .github/workflows/ci.yml: run the main test step with --profile ci and upload the JUnit report (if: always(), 3-day retention, run-number in name). The migration-proof step stays on the default profile to avoid clobbering the ci JUnit artifact (its tests are not quarantined). - docs/testing/README.md: new skeleton (owned by backlog#1153 infra-11) holding the flake policy: discover -> open issue within 24h -> quarantine with issue link -> fix or delete within 30 days. AGENTS.md points to it. Refs: rustfs/backlog#1149, rustfs/backlog#937, rustfs/backlog#1155 |