Reworks the S3 compatibility test harness for reproducibility and faster feedback: - Pin ceph/s3-tests to a fixed commit (S3TESTS_REV, fetch-by-SHA) so the 449-test PR gate is reproducible; previously every run cloned upstream master, letting test renames or assertion changes break CI silently. - PR gate (ci.yml s3-implemented-tests): MAXFAIL=0 + XDIST=4 so a single CI round reports every failure in parallel instead of stopping at the first one serially. - Add TEST_SCOPE=all to run.sh to run the entire upstream suite, and report_compat.py to diff junit results against the classification lists (regressions, promotion candidates, unclassified tests). - Rewrite e2e-s3tests.yml: delegate execution to run.sh (single source of truth; also fixes the broken config generation that left S3_PORT empty), add a weekly scheduled full sweep that fails only on whitelist regressions, and fix the multi-node topology to a real distributed cluster (endpoint-style RUSTFS_VOLUMES) instead of four independent single-node stores behind a load balancer. - Docs: rewrite stale .github/s3tests/README.md (marker-era strategy), update scripts/s3-tests/README.md, fix dead build_testexpr.sh reference in S3_COMPAT_WORKFLOW.md, drop legacy non_standard_tests.txt. All 747 classified test names verified present at the pinned revision; 13 upstream tests are currently unclassified and will surface in the first full-sweep report. Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
5.2 KiB
S3 Compatibility Fix Workflow
Step-by-step guide for identifying and fixing S3 API compatibility issues in RustFS.
Prerequisites
- Rust toolchain installed (
cargo,rustc) - Python 3 for running ceph s3-tests
- Access to upstream remote (
git remote add upstream https://github.com/rustfs/rustfs.git) - Familiarity with the run.sh test runner
Workflow
Step 1: Sync with upstream
git checkout main
git fetch upstream
git merge upstream/main
Step 2: Select candidate tests
Pick candidate tests from unimplemented_tests.txt and run them via a pytest -k expression:
TESTEXPR="test_foo or test_bar" DEPLOY_MODE=build MAXFAIL=0 ./scripts/s3-tests/run.sh
Alternatively, check the latest weekly full-sweep report (the e2e-s3tests workflow's job summary, or artifacts/s3tests-single/compat-report.md) — its "promotion candidates" section lists tests that already pass and only need reclassification.
Review artifacts/s3tests-single/pytest.log for results.
Step 3: Triage failures
From the test report, classify each failure:
| Category | Action |
|---|---|
Feature not implemented (e.g., returns 501 Not Implemented) |
Skip — do not attempt |
| Incorrect HTTP status code or response body | Good candidate for a fix |
| Teardown/cleanup issue (test passes but cleanup fails) | Good candidate — usually simple |
| Complex multi-feature dependency | Defer to a later iteration |
Pick the simplest failure to fix. "Simplest" means: fewest code paths affected, clearest expected behavior, closest to existing implementation.
Step 4: Deep analysis
Before writing any code:
- Read the test source in
s3-tests/s3tests/functional/test_s3.pyto understand exactly what the test expects. - Read the S3 API specification for the operation being tested.
- Search the RustFS codebase for the handler that serves this operation.
- Compare with MinIO (
github.com/minio/minio) — find the equivalent handler and see how it handles the same edge case. This is critical because RustFS was ported from MinIO's Go code to Rust. - Identify the root cause — is it a missing header, wrong status code, incorrect XML response, logic bug, etc.?
- Document your findings before making changes.
Step 5: Create a fix branch
git checkout main
git checkout -b fix/s3-compat-<short-description>
One branch per fix. Never combine unrelated fixes in a single branch.
Step 6: Write tests first (when applicable)
If the fix involves logic changes, add unit tests before modifying the production code:
- Co-locate tests with their module (
#[cfg(test)] mod tests { ... }) - Use descriptive test names:
test_<operation>_<scenario>_<expected_outcome> - Cover both the happy path and the edge case being fixed
Step 7: Implement the fix
Guidelines:
- Reuse existing abstractions — do not duplicate logic that already exists in helper functions or shared modules.
- Follow s3s conventions — RustFS's S3 layer is built on the
s3scrate. Respect its traits, error types, and request/response patterns. - Name things clearly — variable names, function names, and types should be self-explanatory.
- Add comments only where intent is non-obvious — explain why, not what.
- Do not introduce security holes — validate inputs, check permissions, handle errors properly.
- Do not use
unwrap()orexpect()in production code — use proper error handling withResultand?.
Step 8: Verify the fix
Run the specific test(s) you fixed:
TESTEXPR="test_the_fixed_test" DEPLOY_MODE=build ./scripts/s3-tests/run.sh
Then run the full implemented test suite to confirm no regressions:
./scripts/s3-tests/run.sh
Step 9: Run quality checks
make pre-commit
This runs:
cargo fmt --all --checkcargo clippy --all-targets --all-features -- -D warningscargo test --workspace --exclude e2e_test
All three must pass before committing.
Step 10: Commit and prepare PR
git add -A
git commit -m "fix(s3): <concise description of the fix>"
Write a PR description following .github/pull_request_template.md. The description must:
- Be written in English
- Use plain, natural language (no emoji, no marketing speak)
- Explain what was wrong, why, and how it was fixed
- Reference the specific s3-tests that now pass
Step 11: Update test lists
Move the now-passing test(s) from unimplemented_tests.txt to implemented_tests.txt. Update the test count comment in implemented_tests.txt.
Important Rules
- One branch, one fix — never mix unrelated changes.
- Analyze before coding — understand the root cause thoroughly before writing a fix.
- No shotgun debugging — do not blindly try different return codes or response shapes hoping to pass the test.
- Every change must be justified — if you cannot explain why a line changed, do not change it.
- Compare with MinIO — when in doubt about the correct behavior, check MinIO's source code for the equivalent logic.
- Security first — never skip input validation, permission checks, or error handling for the sake of passing a test.