mirror of
https://github.com/rustfs/rustfs.git
synced 2026-09-08 21:25:59 +00:00
8e987ce0a6
* fix(replication): close the GA blocker set from backlog#2366 (#7503) * fix(replication): close GA blockers from backlog#2366 Implements the P1 set from the pre-GA replication audit: - Replication rule tag filters now require every And.Tag to match, replacing the s3s OR semantics with a local AND matcher that fails closed on a malformed tag. - A replicated group membership change no longer writes the group status, so a membership update carrying the default Enabled status cannot silently re-enable a disabled group on the peer. - A successful IAM import schedules one collapsed full-IAM snapshot per remote peer instead of leaving the imported entities local-only. - A pending endpoint refresh is redriven by the heavyweight reconcile tick, carries its own ilm-expiry override, and no longer blocks a remove that drops every unacknowledged peer. - Site metrics expose local replication failure totals and rolling windows; node-level counters no longer report a constructed zero. - set/remove-remote-target notify peer metadata caches before returning, so a follow-up put-bucket-replication on another node sees the target. - Adds the site-replication operations runbook, a docs index, a replication support boundary section, and the Replication changelog section. * fix(site-replication): resume only a locally driven endpoint refresh The peer-side edit handler journals a pending endpoint refresh with an empty `remote_peers` map and commits it inside the same request through `apply_internal_peer_edit`. The reconcile tick could not tell that journal from the coordinator's own: with no required peers it reads as complete on sight, so the tick committed it with `edit_state` - losing the local-name sync - and cleared it under the request that owned it, whose commit then reported the refresh as changed and denied the coordinator the peer acknowledgement it was waiting for. Resume now runs only for a journal that carries the fan-out topology. A receiver's journal stays for the coordinator to redrive with the same refresh id, which is the path that already recovers it. * fix(site-replication): keep an explicit disabled group status on a snapshot Skipping the group-status write whenever an item carries members stopped a membership change from re-enabling a disabled group, but it also silenced the full-IAM snapshot, which always sends members together with the sender's real status. A peer that did not have the group yet created it through `GroupInfo::new` - enabled - so a bootstrap, a repair, or the snapshot an IAM import now schedules handed every member of a frozen group live access there. The madmin wire maps an unset `groupStatus` to Enabled, so only Enabled can be a default. Disabled is always explicit and is applied again. * fix(site-replication): schedule the import snapshot without recording a failure `import-iam` reused the failure-recording path to queue its full-IAM snapshot. That raises `retry_count` on every call, so three imports - the normal shape of a bulk migration done one archive at a time - escalated a healthy peer to `retryStats.failed` with the scheduling note shown as `lastError`, which is exactly the signal the runbook tells operators to repair. A full retry queue also turned a completed import into a 503. Scheduling now only ensures the collapsed entry exists, and a failure to schedule is logged instead of failing the request: the entities are already imported and the reconcile pass still closes the gap. * fix(admin): stop reporting replication failures as retries `retries` is the minio-go counter for redeliveries, and mc prints it as such. Filling it with the failure count claimed a redelivery that never happens: a failed object is not retried by an event today, it waits for the scanner heal pass. `errors` keeps the failure counters; `retries` stays zero until there is a real redelivery to count, and the runbook now says so. * perf(site-replication): aggregate failure windows without cloning bucket stats `site_metrics_snapshot` went through `get_all`, which clones every bucket's stats, and then scanned each target's sample deque twice. That deque is bounded only by the one-hour window, so an unreachable target under load - the case an operator polls this endpoint for - made every `mc admin replicate status` copy the whole backlog and hold the read lock against the failure path while doing it. It now folds under the read lock and takes both windows in one walk. The `max` against the serialized `last_minute` / `last_hour` snapshots is dropped: those are stamped onto per-bucket clones elsewhere and are always zero in this node-local cache. * fix(site-replication): reject a conflicting ilm-expiry override on a re-run The commit now reads the ilm-expiry override back out of the pending refresh journal, so a second edit that asks for a different value had it dropped while the request still reported success. Re-running without the flag keeps pinning the recorded value - that is the documented way to redrive a stuck refresh - but an explicit different value is now rejected instead of ignored. * fix(admin): do not fail a remote-target write on a peer reload error set/remove-remote-target propagated the peer metadata reload error, so a target that was already persisted and live on this node reported a 5xx to the client whenever one peer could not be reached. Every S3 bucket-config write path treats that reload as best effort and only warns; these two admin handlers now do the same, and the reason is logged with the bucket and action. * fix(site-replication): undo every bucket a cut-short refresh rewrote When a remove accepted on another node clears the refresh journal mid-pass, only the bucket holding the lock at that moment had its restored target undone. The buckets rewritten earlier in the same pass kept a target pointing at the removed peer whenever the remove's own cleanup had already walked past them. The undo now covers every bucket this pass rewrote, attempting all of them so one failure does not strand the rest. * fix(site-replication): keep replay running while an endpoint refresh is pending A pending endpoint refresh took the whole heavyweight pass with it, so a peer that never came back froze IAM and bucket replay to every healthy peer too - the stall this journal's resume path was meant to end. The refresh arm now drains the retry queue before returning; it replays per-peer deliveries against the endpoints currently committed in state, so it is unaffected by the edit in flight. Bucket wiring reconciliation still waits, because it rewrites the very targets the refresh is changing, and the runbook now says so. * test(e2e): cover the AND semantics of a two-tag replication filter The acceptance matrix only had a single-tag rule, which matches under both AND and OR semantics and therefore proved nothing about the filter this fix changed. It now also carries a two-tag `And` rule - the shape `mc replicate add --tags "k1=v1&k2=v2"` writes - and asserts that an object with one of the two tags is not admitted while an object with both is. No new test function, so the nightly selection digest is unchanged. * refactor(site-replication): fold the refresh state-change error into one constructor The endpoint-refresh work added three `s3_error!` invocation lines, which the s3s footprint ratchet is meant to prevent. Five copies of the same concurrent-change error now share one constructor, so the surface nets one line smaller than main; the baseline is retightened to match. * fix(site-replication): report a peer whose IAM snapshot waits for a repair An escalated snapshot entry records a deletion a snapshot cannot replay, so only a repair settles it and the marker must survive. Scheduling an import snapshot therefore leaves that peer's entry alone - and now says so, instead of returning success while nothing was scheduled for it. * docs(operations): state the group-status and escalation convergence limits Two boundaries the fixes in this branch make load-bearing: a membership change never carries an enable, so a group disabled on one site only has to be re-enabled there explicitly; and a peer holding an escalated IAM entry does not receive a scheduled snapshot, including the one a bulk import schedules, until a repair settles it. * fix(ci): bind performance runs to selected inputs (#7512) * test(scanner): add G09 upgrade evidence runner Add a reusable Linux x86_64 runner for the Scanner/Heal G09 mixed-version and rollback upgrade evidence lanes. The helper reads the pinned previous-release asset metadata from the upgrade workflow, verifies the downloaded binary, builds the current head, runs both ignored E2E tests, and fails unless the expected G09 JSON artifacts exist. Co-Authored-By: heihutu <heihutu@gmail.com> Co-Authored-By: zhi22915 <qiuzgang@gmail.com> --------- Co-authored-by: 唐小鸭 <tangtang1251@qq.com> Co-authored-by: Zhengchao An <anzhengchao@gmail.com> Co-authored-by: zhi22915 <qiuzgang@gmail.com>
892 lines
52 KiB
Python
892 lines
52 KiB
Python
#!/usr/bin/env python3
|
|
"""Exercise functional failures, chain dispatch, and security evidence without remote VMs."""
|
|
|
|
from __future__ import annotations
|
|
|
|
import glob
|
|
import json
|
|
import os
|
|
import re
|
|
import subprocess
|
|
import sys
|
|
import tempfile
|
|
import unittest
|
|
from pathlib import Path
|
|
|
|
from check_test_wiring import yaml_block
|
|
from functional_case_report import generate_report
|
|
|
|
|
|
ROOT = Path(__file__).resolve().parents[1]
|
|
WORKFLOW = ROOT / ".github/workflows/rustfs-security-test.yml"
|
|
CASE_ROW = "| IAM-101 | user CRUD lifecycle | PASS |"
|
|
|
|
|
|
def named_steps(job: list[str]) -> dict[str, list[str]]:
|
|
starts = [i for i, line in enumerate(job) if line.startswith(" - name: ")]
|
|
return {
|
|
job[start].split(": ", 1)[1].strip('"'): job[start:end]
|
|
for start, end in zip(starts, starts[1:] + [len(job)])
|
|
}
|
|
|
|
|
|
def shell_body(lines: list[str]) -> str:
|
|
start = lines.index(" run: |") + 1
|
|
shell_lines = []
|
|
for line in lines[start:]:
|
|
if line.strip() and not line.startswith(" "):
|
|
break
|
|
shell_lines.append(line[10:])
|
|
if not shell_lines:
|
|
raise ValueError("missing literal shell body")
|
|
return "\n".join(shell_lines)
|
|
|
|
|
|
class WorkflowSteps:
|
|
def uploaded_files(self) -> set[Path]:
|
|
upload = next(lines for lines in self.steps.values() if any("uses: actions/upload-artifact@" in line for line in lines))
|
|
start = upload.index(" path: |") + 1
|
|
paths = []
|
|
for line in upload[start:]:
|
|
if not line.startswith(" "):
|
|
break
|
|
paths.extend(Path(path) for path in glob.glob(self.render(line.strip())))
|
|
return {file for path in paths for file in (path.rglob("*") if path.is_dir() else [path]) if file.is_file()}
|
|
|
|
def render(self, value: str) -> str:
|
|
return re.sub(r"\$\{\{\s*(.*?)\s*\}\}", lambda match: self.context[match[1]], value)
|
|
|
|
def step_env(self, lines: list[str], indent: int = 8) -> dict[str, str]:
|
|
result = {}
|
|
for line in yaml_block(lines, "env", indent) or []:
|
|
if line.strip() and not line.lstrip().startswith("#"):
|
|
key, value = line.strip().split(": ", 1)
|
|
result[key] = self.render(value.strip("'\""))
|
|
return result
|
|
|
|
def run_step(self, name: str) -> subprocess.CompletedProcess[str]:
|
|
lines = self.steps[name]
|
|
result = subprocess.run(
|
|
["bash", "--noprofile", "--norc", "-e", "-o", "pipefail", "-c", self.render(shell_body(lines))],
|
|
cwd=self.directory, env={**self.env, **self.step_env(lines)}, capture_output=True, text=True,
|
|
)
|
|
for line in lines:
|
|
if line.startswith(" id: "):
|
|
self.context[f"steps.{line.split(': ', 1)[1]}.outcome"] = "failure" if result.returncode else "success"
|
|
if Path(self.env["GITHUB_ENV"]).exists():
|
|
for line in Path(self.env["GITHUB_ENV"]).read_text().splitlines():
|
|
key, value = line.split("=", 1)
|
|
self.env[key] = value
|
|
self.context[f"env.{key}"] = value
|
|
return result
|
|
|
|
|
|
class SecurityWorkflowTests(WorkflowSteps, unittest.TestCase):
|
|
def setUp(self) -> None:
|
|
self.source = WORKFLOW.read_text()
|
|
self.job = yaml_block(self.source.splitlines(), "security-test", 2)
|
|
self.assertIsNotNone(self.job)
|
|
self.steps = named_steps(self.job)
|
|
self.temp = tempfile.TemporaryDirectory()
|
|
self.addCleanup(self.temp.cleanup)
|
|
self.directory = Path(self.temp.name)
|
|
self.context = {
|
|
"runner.temp": self.temp.name,
|
|
"github.server_url": "https://github.com",
|
|
"github.repository": "rustfs/rustfs",
|
|
"github.run_id": "314159",
|
|
"github.run_attempt": "2",
|
|
"github.sha": "0123456789abcdef0123456789abcdef01234567",
|
|
"github.event_name": "workflow_dispatch",
|
|
"github.workspace": self.temp.name,
|
|
"inputs.package_url": "",
|
|
"inputs.rustfs_version": "test-version",
|
|
"inputs.topology": "all",
|
|
"inputs.oidc_live": "false",
|
|
"steps.evidence.outcome": "skipped",
|
|
"steps.test.outcome": "skipped",
|
|
"steps.report.outcome": "skipped",
|
|
}
|
|
self.env = {
|
|
**os.environ, "GITHUB_STEP_SUMMARY": str(self.directory / "summary.md"),
|
|
"GITHUB_ENV": str(self.directory / "github-env"), "RUNNER_TEMP": self.temp.name, "TMPDIR": self.temp.name,
|
|
"LOG_FILE": str(self.directory / "suite.log"),
|
|
}
|
|
for key in ("server_url", "repository", "run_id", "run_attempt", "sha", "event_name"):
|
|
self.env[f"GITHUB_{key.upper()}"] = self.context[f"github.{key}"]
|
|
self.context["env.SECURITY_ARTIFACTS_DIR"] = ""
|
|
self.artifacts = self.directory / "rustfs-security-314159-2"
|
|
suite = self.directory / "auto-testing/rustfs-security-test.sh"
|
|
suite.parent.mkdir()
|
|
suite.write_text(
|
|
'#!/usr/bin/env bash\nset -euo pipefail\n'
|
|
'log_dir=$(mktemp -d "$TMPDIR/rustfs-security.XXXXXX")\n'
|
|
'echo "CURRENT SUITE LOG" > "$log_dir/suite.log"\n'
|
|
'echo "CURRENT SUITE STDOUT"; echo "CURRENT SUITE STDERR" >&2\n'
|
|
'case "$FAKE_REPORT" in\n'
|
|
f' present) printf "%s\\n" "CURRENT SUITE DIAGNOSTIC" "{CASE_ROW}" > "$REPORT_FILE" ;;\n'
|
|
' empty) : > "$REPORT_FILE" ;;\n'
|
|
'esac\n'
|
|
'echo "UNWRAPPED SUITE SUMMARY" >> "$GITHUB_STEP_SUMMARY"\n'
|
|
'exit "$FAKE_EXIT"\n'
|
|
)
|
|
|
|
def test_workflow_wiring(self) -> None:
|
|
names = list(self.steps)
|
|
self.assertLess(names.index("Checkout repository (for the OIDC live gate script)"), names.index("Checkout auto-testing scripts (with retry)"))
|
|
self.assertNotIn(" continue-on-error: true", self.job)
|
|
self.assertIn(" continue-on-error: true", self.steps["Run security suite"])
|
|
for name in ("Initialize security evidence", "Generate report"):
|
|
self.assertNotIn(" continue-on-error: true", self.steps[name])
|
|
self.assertIn(" if: ${{ always() && steps.evidence.outcome == 'success' }}", self.steps["Generate report"])
|
|
self.assertNotIn("/tmp/rustfs-security", self.source)
|
|
for name in ("Upload functional report to dashboard", "File failure issue in rustfs/backlog"):
|
|
report = next(line for line in self.steps[name] if line.strip().startswith("REPORT_FILE:"))
|
|
self.assertIn("${{ env.SECURITY_ARTIFACTS_DIR }}/report.md", report)
|
|
for name in ("Upload functional report to dashboard", "Upload report and logs"):
|
|
self.assertIn(" if: ${{ always() && steps.evidence.outcome == 'success' }}", self.steps[name])
|
|
artifact_settings = yaml_block(self.steps["Upload report and logs"], "with", 8)
|
|
self.assertIn(" path: |", artifact_settings)
|
|
self.assertIn(" if-no-files-found: error", artifact_settings)
|
|
|
|
def test_suite_report_and_result_matrix(self) -> None:
|
|
for outcome, mode, exit_code in (
|
|
("success", "present", 0), ("failure", "present", 7), ("failure", "missing", 7),
|
|
("success", "missing", 0), ("success", "empty", 0),
|
|
("skipped", "missing", 0), ("skipped", "present", 0),
|
|
("cancelled", "missing", 0), ("cancelled", "present", 0),
|
|
):
|
|
with self.subTest(outcome=outcome, report=mode):
|
|
self.setUp()
|
|
initialized = self.run_step("Initialize security evidence")
|
|
self.assertEqual(initialized.returncode, 0, initialized.stderr)
|
|
self.assertEqual(self.env["SECURITY_ARTIFACTS_DIR"], str(self.artifacts))
|
|
self.env.update(FAKE_REPORT=mode, FAKE_EXIT=str(exit_code))
|
|
if outcome != "skipped" or mode == "present":
|
|
suite = self.run_step("Run security suite")
|
|
self.assertEqual(suite.returncode, exit_code, suite.stderr)
|
|
logs = list(Path(str(self.artifacts) + "-scratch").glob("rustfs-security.*/suite.log"))
|
|
self.assertEqual(len(logs), 1)
|
|
self.assertEqual(logs[0].read_text(), "CURRENT SUITE LOG\n")
|
|
self.context["steps.test.outcome"] = outcome
|
|
report = self.run_step("Generate report")
|
|
success = outcome == "success" and mode == "present"
|
|
self.assertEqual(report.returncode == 0, success, report.stderr)
|
|
contents = (self.artifacts / "report.md").read_text()
|
|
for expected in (
|
|
"https://github.com/rustfs/rustfs/actions/runs/314159", "Attempt: 2",
|
|
f"Workflow Commit: {self.context['github.sha']}", "Trigger: workflow_dispatch",
|
|
f"Test Step Outcome: {'success' if success else 'failure'}", f"Suite Step Outcome: {outcome}",
|
|
):
|
|
self.assertIn(expected, contents)
|
|
self.assertEqual(CASE_ROW in contents, success)
|
|
self.assertEqual("CURRENT SUITE DIAGNOSTIC" in contents, success)
|
|
if mode == "present":
|
|
raw = (self.artifacts / "suite-report.md").read_text()
|
|
self.assertEqual(raw, f"CURRENT SUITE DIAGNOSTIC\n{CASE_ROW}\n")
|
|
summary = Path(self.env["GITHUB_STEP_SUMMARY"]).read_text()
|
|
self.assertEqual(summary, contents)
|
|
self.assertNotIn("UNWRAPPED SUITE SUMMARY", summary)
|
|
expected = {self.artifacts / "report.md"}
|
|
if outcome != "skipped" or mode == "present":
|
|
expected.add(self.artifacts / "suite.log")
|
|
self.assertEqual((self.artifacts / "suite.log").read_text(), "CURRENT SUITE STDOUT\nCURRENT SUITE STDERR\n")
|
|
if mode in ("present", "empty"):
|
|
expected.add(self.artifacts / "suite-report.md")
|
|
(self.artifacts / "unexpected-token.json").write_text("FAKE-SECRET-CANARY")
|
|
scratch = Path(str(self.artifacts) + "-scratch")
|
|
(scratch / "case.out").write_text("FAKE-SECRET-CANARY")
|
|
self.assertEqual(self.uploaded_files(), expected)
|
|
|
|
|
|
def test_oidc_negative_responses_never_print_issued_credentials(self):
|
|
source = (ROOT / "scripts/test/oidc_keycloak_live.sh").read_text()
|
|
for variable, filename in (("TAMPERED_STATUS", "sts-tampered.xml"), ("BAD_STATUS", "sts-bad.xml")):
|
|
start = source.index('[[ "${' + variable + '}" == 403 ]]')
|
|
end = source.index("\n", source.index("grep -q '<Code>AccessDenied</Code>'", start))
|
|
guard = source[start:end]
|
|
credential_xml = "<Credentials><AccessKeyId>FAKE-ACCESS-CANARY</AccessKeyId><SecretAccessKey>FAKE-SECRET-CANARY</SecretAccessKey><SessionToken>FAKE-SESSION-CANARY</SessionToken></Credentials>"
|
|
for status, body, expected in (("200", credential_xml, 1), ("403", credential_xml, 1),
|
|
("403", "<Code>AccessDenied</Code>", 0)):
|
|
with self.subTest(variable=variable, status=status, expected=expected):
|
|
(self.directory / filename).write_text(body)
|
|
result = subprocess.run(["bash", "-e", "-o", "pipefail", "-c", guard],
|
|
env={**self.env, variable: status, "WORK_DIR": str(self.directory)},
|
|
capture_output=True, text=True)
|
|
self.assertEqual(result.returncode, expected, result.stderr)
|
|
for canary in ("FAKE-ACCESS-CANARY", "FAKE-SECRET-CANARY", "FAKE-SESSION-CANARY"):
|
|
self.assertNotIn(canary, result.stdout + result.stderr)
|
|
if status == "200":
|
|
self.assertIn("HTTP 403, got 200", result.stderr)
|
|
|
|
def test_oidc_incomplete_credentials_report_only_the_missing_field(self):
|
|
source = (ROOT / "scripts/test/oidc_keycloak_live.sh").read_text()
|
|
start = source.index("IFS=$'\\t' read -r STS_ACCESS_KEY")
|
|
end = source.index("\n)\n", start) + 3
|
|
extract = source[start:end]
|
|
values = {"AccessKeyId": "FAKE-ACCESS-CANARY", "SecretAccessKey": "FAKE-SECRET-CANARY",
|
|
"SessionToken": "FAKE-SESSION-CANARY", "Expiration": "2099-01-01T00:00:00Z",
|
|
"SubjectFromWebIdentityToken": "alice"}
|
|
for missing in (None, "Expiration", "SubjectFromWebIdentityToken"):
|
|
with self.subTest(missing=missing):
|
|
xml = "<Credentials>" + "".join(f"<{key}>{value}</{key}>" for key, value in values.items() if key != missing) + "</Credentials>"
|
|
(self.directory / "sts-good.xml").write_text(xml)
|
|
result = subprocess.run(["bash", "-e", "-o", "pipefail", "-c", extract],
|
|
env={**self.env, "WORK_DIR": str(self.directory)}, capture_output=True, text=True)
|
|
self.assertEqual(result.returncode, 1 if missing else 0, result.stderr)
|
|
for canary in ("FAKE-ACCESS-CANARY", "FAKE-SECRET-CANARY", "FAKE-SESSION-CANARY"):
|
|
self.assertNotIn(canary, result.stdout + result.stderr)
|
|
if missing:
|
|
self.assertIn(f"missing required STS field: {missing}", result.stderr)
|
|
|
|
def test_existing_evidence_directory_is_rejected(self) -> None:
|
|
for suffix in ("", "-scratch"):
|
|
self.setUp()
|
|
existing = Path(str(self.artifacts) + suffix)
|
|
existing.mkdir()
|
|
stale = existing / "suite-report.md"
|
|
stale.write_text("OLD RUN REPORT")
|
|
self.assertNotEqual(self.run_step("Initialize security evidence").returncode, 0)
|
|
self.assertEqual(stale.read_text(), "OLD RUN REPORT")
|
|
self.assertFalse(Path(self.env["GITHUB_ENV"]).exists())
|
|
(self.artifacts / "report.md").write_text("OLD RUN REPORT")
|
|
self.context.update({
|
|
"env.SECURITY_ARTIFACTS_DIR": str(self.artifacts), "secrets.PF_TESTING_GH_TOKEN": "fake-local-token",
|
|
})
|
|
fake_bin = self.directory / "bin"
|
|
fake_bin.mkdir()
|
|
gh = fake_bin / "gh"
|
|
gh.write_text(
|
|
'#!/usr/bin/env bash\nset -euo pipefail\n'
|
|
'if [ "$1 $2" = "issue create" ]; then\n'
|
|
' while [ "$#" -gt 0 ]; do\n'
|
|
' if [ "$1" = "--body-file" ]; then cat "$2" > "$CAPTURE_BODY"; fi\n'
|
|
' shift\n'
|
|
' done\n'
|
|
'fi\n'
|
|
)
|
|
gh.chmod(0o755)
|
|
body = self.directory / "issue-body.md"
|
|
self.env.update(PATH=f"{fake_bin}{os.pathsep}{os.environ['PATH']}", CAPTURE_BODY=str(body))
|
|
result = self.run_step("File failure issue in rustfs/backlog")
|
|
self.assertEqual(result.returncode, 0, result.stderr)
|
|
self.assertNotIn("OLD RUN REPORT", body.read_text())
|
|
self.assertIn("https://github.com/rustfs/rustfs/actions/runs/314159", body.read_text())
|
|
|
|
def test_all_ten_suites_hold_the_shared_lock_for_manual_and_chain_runs(self) -> None:
|
|
for suite in ("upgrade", "s3-compat", "kms", "tier", "storage", "heal", "pool-expand", "security", "replication", "performance"):
|
|
with self.subTest(suite=suite):
|
|
source = (ROOT / f".github/workflows/rustfs-{suite}-test.yml").read_text().splitlines()
|
|
# Workflow-level concurrency covers every job, including cleanup,
|
|
# regardless of trigger or the runner hosting the job.
|
|
self.assertEqual([
|
|
line.strip() for line in yaml_block(source, "concurrency", 0)
|
|
if line.strip() and not line.lstrip().startswith("#")
|
|
], [
|
|
"group: rustfs-shared-functional-tests", "cancel-in-progress: false",
|
|
])
|
|
self.assertIsNotNone(yaml_block(source, "workflow_dispatch", 2))
|
|
self.assertIsNotNone(yaml_block(source, "repository_dispatch", 2))
|
|
cleanup_name = "Reset test environment (after)" if suite == "performance" else "Cleanup environment (after)"
|
|
cleanup = named_steps(yaml_block(source, "jobs", 0))[cleanup_name]
|
|
self.assertTrue(any(line.startswith(" if:") and "always()" in line for line in cleanup))
|
|
|
|
def test_root_dispatches_only_upgrade_and_replication_hands_off_after_failure(self) -> None:
|
|
for failed_attempts, issue_exit, token in ((0, 0, "fixture"), (2, 0, "fixture"), (3, 0, "fixture"), (3, 7, "fixture"), (0, 0, "")):
|
|
with self.subTest(failed_attempts=failed_attempts, issue_exit=issue_exit, token=bool(token)):
|
|
self.setUp()
|
|
fake_bin = self.directory / "bin"
|
|
fake_bin.mkdir()
|
|
commands = {
|
|
"gh": '''#!/usr/bin/env bash
|
|
set -euo pipefail
|
|
if [ "$1" = api ]; then
|
|
printf '%s\\n' "$*" >> "$DISPATCHES"
|
|
attempt=$(wc -l < "$DISPATCHES")
|
|
[ "$attempt" -gt "$FAILED_ATTEMPTS" ]
|
|
elif [ "$1 $2" = 'issue create' ]; then
|
|
printf 'issue\\n' >> "$EXECUTED"
|
|
while [ "$#" -gt 0 ]; do
|
|
if [ "$1" = --body-file ]; then
|
|
cat "$2" > "$CAPTURE_BODY"
|
|
printf '%s\\n' "$2" > "$CAPTURE_BODY_PATH"
|
|
fi
|
|
shift
|
|
done
|
|
exit "$ISSUE_EXIT"
|
|
else
|
|
exit 99
|
|
fi
|
|
''',
|
|
"sleep": '#!/bin/sh\nprintf "sleep %s\\n" "$1" >> "$EXECUTED"\n',
|
|
"ssh": '#!/bin/sh\nprintf "cleanup\\n" >> "$EXECUTED"\n',
|
|
}
|
|
for name, contents in commands.items():
|
|
command = fake_bin / name
|
|
command.write_text(contents)
|
|
command.chmod(0o755)
|
|
dispatches = self.directory / "dispatches"
|
|
executed = self.directory / "executed"
|
|
body = self.directory / "issue-body.md"
|
|
body_path = self.directory / "issue-body-path"
|
|
self.env.update(
|
|
PATH=f"{fake_bin}{os.pathsep}{os.environ['PATH']}", DISPATCHES=str(dispatches),
|
|
EXECUTED=str(executed), CAPTURE_BODY=str(body), CAPTURE_BODY_PATH=str(body_path),
|
|
FAILED_ATTEMPTS="0", ISSUE_EXIT=str(issue_exit),
|
|
RUSTFS_NODES="fixture-node", RUSTFS_SSH_USER="fixture-user",
|
|
RUSTFS_NIGHTLY_PACKAGE_URL="https://example.invalid/package.deb",
|
|
)
|
|
self.context.update({"secrets.PF_TESTING_GH_TOKEN": "fixture", "inputs.suite": "all"})
|
|
driver = (ROOT / ".github/workflows/rustfs-functional-chain.yml").read_text()
|
|
self.steps = named_steps(yaml_block(driver.splitlines(), "start-chain", 2))
|
|
self.assertEqual(list(self.steps), ["Dispatch first suite (upgrade)"])
|
|
started = self.run_step("Dispatch first suite (upgrade)")
|
|
self.assertEqual(started.returncode, 0, started.stderr)
|
|
self.assertEqual(dispatches.read_text().splitlines(), [
|
|
"api --method POST repos/rustfs/rustfs/dispatches -f event_type=rustfs-chain-upgrade -F client_payload[from_suite]=nightly-build",
|
|
])
|
|
dispatches.unlink()
|
|
|
|
replication = (ROOT / ".github/workflows/rustfs-replication-test.yml").read_text()
|
|
job = yaml_block(replication.splitlines(), "replication-test", 2)
|
|
self.assertFalse(any(line.startswith(" continue-on-error:") for line in job))
|
|
self.steps = named_steps(job)
|
|
handoff = "Continue functional chain (next: Performance)"
|
|
self.assertIn(" if: ${{ always() && github.event_name == 'repository_dispatch' }}", self.steps[handoff])
|
|
self.assertFalse(any(line.strip().startswith("continue-on-error:") for line in self.steps[handoff]))
|
|
self.assertIn(" if: always()", self.steps["Cleanup environment (after)"])
|
|
self.assertLess(list(self.steps).index("Cleanup environment (after)"), list(self.steps).index(handoff))
|
|
initialized = self.run_step("Initialize functional evidence")
|
|
self.assertEqual(initialized.returncode, 0, initialized.stderr)
|
|
suite = self.directory / "auto-testing/rustfs-replication-test.sh"
|
|
suite.write_text('#!/bin/sh\nprintf "suite failed\\n" >> "$EXECUTED"\nexit 17\n')
|
|
failed = self.run_step("Run replication suite")
|
|
self.assertEqual(failed.returncode, 17, failed.stderr)
|
|
cleaned = self.run_step("Cleanup environment (after)")
|
|
self.assertEqual(cleaned.returncode, 0, cleaned.stderr)
|
|
self.assertEqual(executed.read_text().splitlines(), ["suite failed", "cleanup"])
|
|
self.env["FAILED_ATTEMPTS"] = str(failed_attempts)
|
|
self.context["secrets.PF_TESTING_GH_TOKEN"] = token
|
|
forwarded = self.run_step(handoff)
|
|
self.assertEqual(forwarded.returncode == 0, bool(token) and failed_attempts < 3, forwarded.stderr)
|
|
calls = dispatches.read_text().splitlines() if dispatches.exists() else []
|
|
self.assertEqual(calls, [
|
|
"api --method POST repos/rustfs/rustfs/dispatches -f event_type=rustfs-chain-performance -F client_payload[from_suite]=replication",
|
|
] * (min(failed_attempts + 1, 3) if token else 0))
|
|
if failed_attempts == 3:
|
|
self.assertIn("could not hand off from **replication** to **Performance**", body.read_text())
|
|
self.assertIn("rustfs-chain-performance", body.read_text())
|
|
self.assertEqual(executed.read_text().splitlines().count("issue"), 2 if issue_exit else 1)
|
|
self.assertFalse(Path(body_path.read_text().strip()).exists())
|
|
|
|
|
|
class FunctionalWorkflowTests(unittest.TestCase):
|
|
JOBS = {
|
|
"kms": "kms-test", "storage": "storage-test", "s3-compat": "s3-compat-test",
|
|
"upgrade": "upgrade-test", "replication": "replication-test", "heal": "heal-test",
|
|
"tier": "tier-test", "pool-expand": "pool-expansion-test", "performance": "performance-test",
|
|
}
|
|
DIRECT_TESTS = {
|
|
"kms": "Run KMS suite", "storage": "Run storage engine suite",
|
|
"s3-compat": "Run S3 compatibility suite", "upgrade": "Run upgrade compatibility suite",
|
|
"replication": "Run replication suite",
|
|
}
|
|
|
|
def test_failure_and_always_step_wiring(self) -> None:
|
|
for suite, job_id in self.JOBS.items():
|
|
with self.subTest(suite=suite):
|
|
source = (ROOT / f".github/workflows/rustfs-{suite}-test.yml").read_text()
|
|
job = yaml_block(source.splitlines(), job_id, 2)
|
|
self.assertIsNotNone(job)
|
|
self.assertNotRegex("\n".join(job), r'''(?m)^ ["']?continue-on-error["']?\s*:''')
|
|
steps = named_steps(job)
|
|
if suite in self.DIRECT_TESTS:
|
|
test = steps[self.DIRECT_TESTS[suite]]
|
|
self.assertNotRegex("\n".join(test), r'''(?m)^ ["']?continue-on-error["']?\s*:''')
|
|
self.assertIn(" if: ${{ always() && steps.evidence.outcome == 'success' }}", steps["Generate report"])
|
|
cleanup = steps["Reset test environment (after)" if suite == "performance" else "Cleanup environment (after)"]
|
|
condition = next(line.strip() for line in cleanup if line.startswith(" if:"))
|
|
self.assertIn(condition, (
|
|
"if: always()",
|
|
"if: ${{ always() && inputs.cleanup_after != 'false' }}",
|
|
"if: ${{ always() && (inputs.cleanup_after != 'false' || github.event_name != 'workflow_dispatch') }}",
|
|
))
|
|
if suite != "performance":
|
|
handoff = next(
|
|
value for name, value in steps.items() if name.startswith("Continue functional chain")
|
|
)
|
|
self.assertIn(" if: ${{ always() && github.event_name == 'repository_dispatch' }}", handoff)
|
|
|
|
def test_failed_suite_preserves_exit_and_cleanup_and_dispatch_execute(self) -> None:
|
|
for suite, test_name in self.DIRECT_TESTS.items():
|
|
with self.subTest(suite=suite), tempfile.TemporaryDirectory() as directory:
|
|
root = Path(directory)
|
|
(root / "auto-testing").mkdir()
|
|
script = root / f"auto-testing/rustfs-{suite}-test.sh"
|
|
script.write_text('#!/bin/sh\nprintf "partial suite diagnostics\\n"\nexit 17\n')
|
|
script.chmod(0o755)
|
|
fake_bin = root / "bin"
|
|
fake_bin.mkdir()
|
|
for command, marker in (("ssh", "cleanup"), ("gh", "dispatch")):
|
|
fake = fake_bin / command
|
|
fake.write_text(f'#!/bin/sh\nprintf "{marker}\\n" >> "$EXECUTED"\n')
|
|
fake.chmod(0o755)
|
|
env = {
|
|
**os.environ, "PATH": f"{fake_bin}{os.pathsep}{os.environ['PATH']}",
|
|
"EXECUTED": str(root / "executed"), "RUSTFS_NODES": "fixture-node",
|
|
"RUSTFS_SSH_USER": "fixture-user", "RUSTFS_NIGHTLY_PACKAGE_URL": "https://example.invalid/package.deb",
|
|
"GH_TOKEN": "local-fixture", "GITHUB_EVENT_NAME": "repository_dispatch", "GITHUB_RUN_ID": "314159",
|
|
}
|
|
source = (ROOT / f".github/workflows/rustfs-{suite}-test.yml").read_text()
|
|
steps = named_steps(yaml_block(source.splitlines(), self.JOBS[suite], 2))
|
|
context = {"github.event_name": "repository_dispatch", "steps.test.outcome": "failure"}
|
|
for expression in re.findall(r"\$\{\{\s*(.*?)\s*\}\}", source):
|
|
if expression.startswith("inputs.") and re.fullmatch(r"inputs\.\w+", expression):
|
|
context[expression] = ""
|
|
def execute(name):
|
|
lines = steps[name]
|
|
rendered = re.sub(r"\$\{\{\s*(.*?)\s*\}\}", lambda match: context[match[1]], shell_body(lines))
|
|
return subprocess.run(
|
|
["bash", "--noprofile", "--norc", "-e", "-o", "pipefail", "-c", rendered],
|
|
cwd=root, env={**env, "LOG_FILE": str(root / "suite.log")}, capture_output=True, text=True,
|
|
)
|
|
failed = execute(test_name)
|
|
self.assertEqual(failed.returncode, 17, failed.stderr)
|
|
self.assertIn("partial suite diagnostics", failed.stdout)
|
|
cleanup = execute("Cleanup environment (after)")
|
|
self.assertEqual(cleanup.returncode, 0, cleanup.stderr)
|
|
handoff_name = next(
|
|
name for name in steps if name.startswith("Continue functional chain")
|
|
)
|
|
handoff = execute(handoff_name)
|
|
self.assertEqual(handoff.returncode, 0, handoff.stderr)
|
|
markers = (root / "executed").read_text().splitlines()
|
|
self.assertEqual(markers, ["cleanup", "dispatch"])
|
|
|
|
|
|
class FunctionalCaseReportTests(unittest.TestCase):
|
|
def report(self, text: str | None, matrix: bool = False) -> tuple[bool, str, str]:
|
|
with tempfile.TemporaryDirectory() as directory:
|
|
root = Path(directory)
|
|
log = root / "suite.log"
|
|
if text is not None:
|
|
log.write_text(text)
|
|
valid = generate_report(log, root / "cases.md", root / "matrix.md" if matrix else None)
|
|
return valid, (root / "cases.md").read_text(), (root / "matrix.md").read_text() if matrix else ""
|
|
|
|
def test_repeated_case_executions_preserve_failure_and_context(self):
|
|
# log() from rustfs/auto-testing@6120aa0a76de, rustfs-kms-test.sh:131.
|
|
log = subprocess.check_output(["bash", "-c", r'''
|
|
log() { printf '\033[1;36m[INFO]\033[0m %s\n' "$*"; }
|
|
log '== topology: single-single kms-backend: local =='
|
|
printf '\033[32m--- KMS-101 roundtrip ---\033[0m\n[FAIL] KMS-101\n'
|
|
log '== topology: single-multi kms-backend: vault-kv2 =='
|
|
printf '%s\n' '--- KMS-101 roundtrip ---' '[PASS] KMS-101'
|
|
printf '%s\n' '--- KMS-101 roundtrip ---' '[UNSUPPORTED] KMS-101'
|
|
'''], text=True)
|
|
valid, cases, _ = self.report(log)
|
|
self.assertFalse(valid)
|
|
self.assertEqual(cases.count("| KMS-101 |"), 3)
|
|
self.assertIn("- Total: 3\n- PASS: 1\n- FAIL: 1\n- UNSUPPORTED: 1\n- RUNNING: 0\n", cases)
|
|
self.assertIn("roundtrip (topology: single-single kms-backend: local) | FAIL |", cases)
|
|
self.assertIn("roundtrip (topology: single-multi kms-backend: vault-kv2) | PASS |", cases)
|
|
self.assertNotIn("\\n", cases)
|
|
|
|
def test_missing_empty_unfinished_and_orphan_results_are_not_success(self):
|
|
for text in (None, "", "setup failed\n", "--- KMS-101 roundtrip ---\n", "[PASS] KMS-101\n",
|
|
"--- KMS-101 first ---\n--- KMS-101 second ---\n[PASS] KMS-101\n",
|
|
"--- KMS-101 first ---\n[FAIL] KMS-101\n[PASS] KMS-101\n"):
|
|
with self.subTest(log=text):
|
|
valid, cases, _ = self.report(text)
|
|
self.assertFalse(valid)
|
|
self.assertIn("## Case Summary", cases)
|
|
valid, cases, _ = self.report("[INFO] == suite: bucket replication (REP-*) ==\n--- REP-101 unsupported ---\n[UNSUPPORTED] REP-101\n")
|
|
self.assertTrue(valid)
|
|
self.assertIn("suite: bucket replication", cases)
|
|
self.assertIn("- UNSUPPORTED: 1\n", cases)
|
|
|
|
def test_upgrade_matrix_is_preserved_and_required_for_complete_report(self):
|
|
case = "--- UPG-101 upgrade ---\n[PASS] UPG-101\n"
|
|
for suffix, expected in (("", False), ("[UPG-TOPO] single-single local v1 v2 PASS=1 FAIL=0\n", True),
|
|
("[UPG-TOPO] single-single local v1 v2 PASS=1 FAIL=1\n", False)):
|
|
with self.subTest(matrix=suffix):
|
|
valid, _, matrix = self.report(case + suffix, matrix=True)
|
|
self.assertEqual(valid, expected)
|
|
self.assertIn("| Topology | KMS Backend | From Version | To Version | Result |", matrix)
|
|
self.assertIn("| single-single | local | v1 | v2 |" if suffix else "NOT RUN", matrix)
|
|
|
|
def test_s3_case_identifiers_include_digits(self):
|
|
valid, cases, _ = self.report("--- S3C-101 CreateBucket ---\n[PASS] S3C-101\n")
|
|
self.assertTrue(valid)
|
|
self.assertIn("| S3C-101 | CreateBucket (context not recorded) | PASS |", cases)
|
|
|
|
|
|
class FunctionalEvidenceTests(WorkflowSteps, unittest.TestCase):
|
|
SUITES = (*FunctionalWorkflowTests.DIRECT_TESTS, "heal", "performance")
|
|
|
|
def prepare(self, suite: str) -> None:
|
|
self.temp = tempfile.TemporaryDirectory()
|
|
self.addCleanup(self.temp.cleanup)
|
|
self.directory = Path(self.temp.name)
|
|
self.source = (ROOT / f".github/workflows/rustfs-{suite}-test.yml").read_text()
|
|
self.steps = named_steps(yaml_block(self.source.splitlines(), FunctionalWorkflowTests.JOBS[suite], 2))
|
|
self.context = {expression: "" for expression in re.findall(r"\$\{\{\s*(.*?)\s*\}\}", self.source)}
|
|
self.context.update({
|
|
"github.server_url": "https://github.com", "github.repository": "rustfs/rustfs",
|
|
"github.run_id": "314159", "github.run_attempt": "2", "github.sha": "0123456789abcdef0123456789abcdef01234567",
|
|
"github.event_name": "repository_dispatch", "steps.test.outcome": "success",
|
|
"secrets.PF_TESTING_GH_TOKEN": "local-fixture", "env.PF_TESTING_GH_TOKEN": "local-fixture",
|
|
})
|
|
self.artifacts = self.directory / f"rustfs-{suite}-314159-2"
|
|
self.env = {
|
|
**os.environ, "GITHUB_ENV": str(self.directory / "github-env"), "RUNNER_TEMP": self.temp.name,
|
|
"GITHUB_STEP_SUMMARY": str(self.directory / "summary.md"), "RUSTFS_NODES": "fixture-node",
|
|
"RUSTFS_NIGHTLY_PACKAGE_URL": "https://example.invalid/package.deb", "CAPTURE_BODY": str(self.directory / "issue.md"),
|
|
}
|
|
for key in ("server_url", "repository", "run_id", "run_attempt", "sha", "event_name"):
|
|
self.env[f"GITHUB_{key.upper()}"] = self.context[f"github.{key}"]
|
|
(self.directory / "scripts").mkdir()
|
|
(self.directory / "scripts/functional_case_report.py").symlink_to(ROOT / "scripts/functional_case_report.py")
|
|
fake_bin = self.directory / "bin"
|
|
fake_bin.mkdir()
|
|
(fake_bin / "python3").symlink_to(sys.executable)
|
|
for command, body in (
|
|
("ssh", 'printf "fixture-version\\n"\n'),
|
|
("gh", 'if [ "$1 $2" = "issue create" ]; then\n'
|
|
' while [ "$#" -gt 0 ]; do\n'
|
|
' if [ "$1" = "--body-file" ]; then cat "$2" > "$CAPTURE_BODY"; fi\n'
|
|
' shift\n'
|
|
' done\n'
|
|
'elif [ "$1 $2" = "api --method" ]; then cat >/dev/null; fi\n'),
|
|
):
|
|
script = fake_bin / command
|
|
script.write_text("#!/bin/sh\n" + body)
|
|
script.chmod(0o755)
|
|
self.env["PATH"] = f"{fake_bin}{os.pathsep}{os.environ['PATH']}"
|
|
|
|
def test_evidence_wiring_and_failed_initialization_cannot_publish_stale_files(self):
|
|
for suite, suffix in ((suite, suffix) for suite in self.SUITES for suffix in ("", "-scratch")):
|
|
with self.subTest(suite=suite, collision=suffix or "artifact"):
|
|
self.prepare(suite)
|
|
self.assertNotIn("/tmp/rustfs-", self.source)
|
|
names = list(self.steps)
|
|
self.assertLess(names.index("Initialize functional evidence"), names.index("Checkout auto-testing scripts (with retry)"))
|
|
if suite in FunctionalWorkflowTests.DIRECT_TESTS:
|
|
self.assertLess(names.index("Checkout repository (for report parser)"), names.index("Checkout auto-testing scripts (with retry)"))
|
|
for name, lines in self.steps.items():
|
|
if name in ("Generate report", "Upload functional report to dashboard") or any("uses: actions/upload-artifact@" in line for line in lines):
|
|
self.assertIn(" if: ${{ always() && steps.evidence.outcome == 'success' }}", lines)
|
|
if any("uses: actions/upload-artifact@" in line for line in lines):
|
|
self.assertIn(" path: |", lines)
|
|
self.assertIn(" if-no-files-found: error", lines)
|
|
existing = Path(str(self.artifacts) + suffix)
|
|
existing.mkdir()
|
|
for filename in ("report.md", "suite.log"):
|
|
(existing / filename).write_text("OLD RUN EVIDENCE")
|
|
self.env.update(REPORT_FILE=str(existing / "report.md"), LOG_FILE=str(existing / "suite.log"))
|
|
initialized = self.run_step("Initialize functional evidence")
|
|
self.assertNotEqual(initialized.returncode, 0)
|
|
self.assertFalse(Path(self.env["GITHUB_ENV"]).exists())
|
|
issue = self.run_step("File failure issue in rustfs/backlog")
|
|
self.assertEqual(issue.returncode, 0, issue.stderr)
|
|
body = Path(self.env["CAPTURE_BODY"]).read_text()
|
|
self.assertNotIn("OLD RUN EVIDENCE", body)
|
|
self.assertIn("no report or log file was produced", body)
|
|
self.assertEqual((existing / "report.md").read_text(), "OLD RUN EVIDENCE")
|
|
|
|
def test_reports_use_only_current_complete_suite_evidence(self):
|
|
for suite in self.SUITES[:-1]:
|
|
good = "--- KMS-101 roundtrip ---\n[PASS] KMS-101\n"
|
|
partial = "--- KMS-101 roundtrip ---\n[PASS] KMS-101\n--- KMS-102 unfinished ---\n"
|
|
if suite == "s3-compat":
|
|
good, partial = good.replace("KMS-", "S3C-"), partial.replace("KMS-", "S3C-")
|
|
if suite == "upgrade":
|
|
good += "[UPG-TOPO] single-single local v1 v2 PASS=1 FAIL=0\n"
|
|
if suite == "heal":
|
|
good = "".join(f"[HEAL-STEP] {step} fixture PASS\n" for step in range(1, 8))
|
|
partial = "[HEAL-STEP] 1 fixture PASS\n"
|
|
for outcome, log in (("success", good), ("failure", good), ("success", partial), ("success", ""),
|
|
("success", None), ("skipped", None), ("cancelled", good)):
|
|
with self.subTest(suite=suite, outcome=outcome, log=log):
|
|
self.prepare(suite)
|
|
stale = self.directory / "old-suite.log"
|
|
stale.write_text("OLD RUN EVIDENCE\n" + good)
|
|
self.env.update(LOG_FILE=str(stale), REPORT_FILE=str(stale))
|
|
self.assertEqual(self.run_step("Initialize functional evidence").returncode, 0)
|
|
self.assertEqual(self.env["LOG_FILE"], str(self.artifacts / "suite.log"))
|
|
self.assertEqual(self.env["TMPDIR"], str(self.artifacts) + "-scratch")
|
|
if log is not None:
|
|
Path(self.env["LOG_FILE"]).write_text(log)
|
|
self.context["steps.test.outcome"] = outcome
|
|
report = self.run_step("Generate report")
|
|
success = outcome == "success" and log == good
|
|
self.assertEqual(report.returncode == 0, success, report.stderr)
|
|
contents = Path(self.env["REPORT_FILE"]).read_text()
|
|
self.assertNotIn("OLD RUN EVIDENCE", contents)
|
|
self.assertEqual("| PASS |" in contents, success)
|
|
for value in ("actions/runs/314159", "Attempt: 2", "Workflow Commit: " + self.context["github.sha"],
|
|
f"Test Step Outcome: {'success' if success else 'failure'}", f"Suite Step Outcome: {outcome}"):
|
|
self.assertIn(value, contents)
|
|
self.assertEqual(Path(self.env["GITHUB_STEP_SUMMARY"]).read_text(), contents)
|
|
evidence = (self.artifacts / ("steps.md" if suite == "heal" else "cases.md")).read_text()
|
|
if log in (good, partial):
|
|
self.assertIn("| PASS |", evidence)
|
|
self.assertNotIn("OLD RUN EVIDENCE", evidence)
|
|
|
|
def test_actual_suite_commands_pass_the_current_log_and_scratch_paths(self):
|
|
for suite in self.SUITES:
|
|
with self.subTest(suite=suite):
|
|
self.prepare(suite)
|
|
self.assertEqual(self.run_step("Initialize functional evidence").returncode, 0)
|
|
if suite == "heal":
|
|
self.assertEqual(self.env["RUSTFS_WARP_LOG_FILE"], str(self.artifacts / "warp.log"))
|
|
scripts = self.directory / "auto-testing"
|
|
scripts.mkdir()
|
|
filename = f"rustfs_{suite}_test.sh" if suite in ("heal", "performance") else f"rustfs-{suite}-test.sh"
|
|
script = scripts / filename
|
|
script.write_text(
|
|
'#!/bin/bash\nset -euo pipefail\nlog=""\n'
|
|
'while [ "$#" -gt 0 ]; do\n'
|
|
' if [ "$1" = "--log-file" ]; then log="$2"; shift; fi\n'
|
|
' shift\n'
|
|
'done\n'
|
|
'[ "$log" = "$LOG_FILE" ] || exit 31\n'
|
|
'printf "CURRENT SUITE LOG\\n" > "$log"\n'
|
|
'scratch=$(mktemp -d "$TMPDIR/fixture.XXXXXX")\n'
|
|
'printf "CURRENT SCRATCH\\n" > "$scratch/trace.log"\n'
|
|
'if [ -n "${RUSTFS_RESULT_DIR:-}" ]; then\n'
|
|
' mkdir -p "$RUSTFS_RESULT_DIR"\n'
|
|
' printf "CURRENT RESULTS\\n" > "$RUSTFS_RESULT_DIR/summary.md"\n'
|
|
'fi\n'
|
|
)
|
|
script.chmod(0o755)
|
|
name = FunctionalWorkflowTests.DIRECT_TESTS.get(suite) or (
|
|
"Run benchmark (GET/PUT/MIXED)" if suite == "performance" else "Run heal test (write -> outage -> heal -> verify)"
|
|
)
|
|
result = self.run_step(name)
|
|
self.assertEqual(result.returncode, 0, result.stderr)
|
|
self.assertEqual((self.artifacts / "suite.log").read_text(), "CURRENT SUITE LOG\n")
|
|
self.assertEqual(len(list(Path(self.env["TMPDIR"]).glob("fixture.*/trace.log"))), 1)
|
|
self.assertEqual(list(self.artifacts.glob("fixture.*")), [])
|
|
if suite == "performance":
|
|
self.assertEqual((self.artifacts / "results/summary.md").read_text(), "CURRENT RESULTS\n")
|
|
|
|
def test_upload_allowlist_preserves_diagnostics_without_scratch(self):
|
|
extra = {
|
|
"kms": ["cases.md"], "storage": ["cases.md"], "s3-compat": ["cases.md"],
|
|
"upgrade": ["cases.md", "matrix.md"], "replication": ["cases.md"], "heal": ["steps.md", "warp.log"],
|
|
"performance": ["version.txt", "results/master.log", "results/summary.md", "results/summary.tsv",
|
|
"results/get_1KiB.txt", "results/put_1MiB.txt", "results/mixed_4MiB.txt"],
|
|
}
|
|
for suite in self.SUITES:
|
|
with self.subTest(suite=suite):
|
|
self.prepare(suite)
|
|
self.assertEqual(self.run_step("Initialize functional evidence").returncode, 0)
|
|
expected = {self.artifacts / name for name in ["report.md", "suite.log", *extra[suite]]}
|
|
for path in expected:
|
|
path.parent.mkdir(parents=True, exist_ok=True)
|
|
path.write_text("PARTIAL FAILURE DIAGNOSTIC")
|
|
for directory in (self.artifacts, Path(self.env["TMPDIR"]), self.artifacts / "results"):
|
|
directory.mkdir(exist_ok=True)
|
|
(directory / "init.json").write_text('{"root_token":"FAKE-SECRET-CANARY"}')
|
|
self.assertEqual(self.uploaded_files(), expected)
|
|
self.assertTrue(all("FAKE-SECRET-CANARY" not in path.read_text() for path in self.uploaded_files()))
|
|
|
|
def test_kms_failure_after_vault_init_keeps_only_failure_evidence(self):
|
|
self.prepare("kms")
|
|
self.assertEqual(self.run_step("Initialize functional evidence").returncode, 0)
|
|
scripts = self.directory / "auto-testing"
|
|
scripts.mkdir()
|
|
# Vault file writes and EXIT cleanup from auto-testing@06cd3c097350:23-24,57,479-487.
|
|
# The Docker boundary returns synthetic credentials; setup fails before vault_stop.
|
|
(scripts / "rustfs-kms-test.sh").write_text(r'''#!/bin/bash
|
|
set -Eeuo pipefail
|
|
TEST_TMP="$(mktemp -d "${TMPDIR:-/tmp}/rustfs-test.XXXXXX")"
|
|
trap 'rm -rf "${TEST_TMP}"' EXIT
|
|
VAULT_DATA_DIR="${RUSTFS_VAULT_DATA_DIR:-${TMPDIR:-/tmp}/rustfs-vault-data}"
|
|
VAULT_CONTAINER="rustfs-vault"
|
|
heal_run() { "$@"; }
|
|
while [ "$#" -gt 0 ]; do
|
|
if [ "$1" = "--log-file" ]; then LOG_FILE="$2"; shift; fi
|
|
shift
|
|
done
|
|
mkdir -p "${VAULT_DATA_DIR}"
|
|
tmp_init="${TEST_TMP}/vault-init.json"
|
|
tmp_err="${TEST_TMP}/vault-init.stderr"
|
|
heal_run docker exec -e "VAULT_ADDR=http://127.0.0.1:8200" "${VAULT_CONTAINER}" \
|
|
vault operator init -key-shares=1 -key-threshold=1 -format=json \
|
|
> "${tmp_init}" 2>"${tmp_err}"
|
|
cat "${tmp_init}" | heal_run tee "${VAULT_DATA_DIR}/init.json" >/dev/null
|
|
printf '%s\n' 'vault initialized; root token acquired' 'fixture setup failed after init' > "${LOG_FILE}"
|
|
exit 42
|
|
''')
|
|
docker = self.directory / "bin/docker"
|
|
docker.write_text("""#!/bin/sh
|
|
[ "$1" = exec ] || exit 99
|
|
printf '%s\\n' '{"root_token":"FAKE-ROOT-CANARY","unseal_keys_b64":["FAKE-UNSEAL-CANARY"]}'
|
|
""")
|
|
docker.chmod(0o755)
|
|
result = self.run_step("Run KMS suite")
|
|
self.assertEqual(result.returncode, 42, result.stderr)
|
|
vault_init = Path(self.env["TMPDIR"]) / "rustfs-vault-data/init.json"
|
|
self.assertEqual(json.loads(vault_init.read_text()), {"root_token": "FAKE-ROOT-CANARY", "unseal_keys_b64": ["FAKE-UNSEAL-CANARY"]})
|
|
self.assertNotIn(self.artifacts, vault_init.parents)
|
|
self.assertEqual(list(Path(self.env["TMPDIR"]).glob("rustfs-test.*")), [])
|
|
self.assertNotEqual(self.run_step("Generate report").returncode, 0)
|
|
self.assertEqual(self.uploaded_files(), {self.artifacts / name for name in ("suite.log", "cases.md", "report.md")})
|
|
self.assertIn("fixture setup failed after init", (self.artifacts / "suite.log").read_text())
|
|
for path in self.uploaded_files():
|
|
self.assertNotIn("FAKE-ROOT-CANARY", path.read_text())
|
|
self.assertNotIn("FAKE-UNSEAL-CANARY", path.read_text())
|
|
|
|
def test_heal_accumulates_actual_staged_steps_without_overwriting_failures(self):
|
|
self.prepare("heal")
|
|
self.assertEqual(self.run_step("Initialize functional evidence").returncode, 0)
|
|
script = self.directory / "auto-testing/rustfs_heal_test.sh"
|
|
script.parent.mkdir()
|
|
# Result printf and full-run condition from auto-testing@6120aa0a76de:143,1163-1168.
|
|
script.write_text(r'''#!/bin/bash
|
|
set -euo pipefail
|
|
SELECTED_STEPS=()
|
|
PREFLIGHT=0
|
|
while [ "$#" -gt 0 ]; do
|
|
case "$1" in
|
|
--steps) IFS=',' read -ra SELECTED_STEPS <<< "$2"; shift ;;
|
|
--log-file) LOG_FILE="$2"; shift ;;
|
|
--preflight) PREFLIGHT=1 ;;
|
|
esac
|
|
shift
|
|
done
|
|
if [ "$PREFLIGHT" -eq 1 ]; then
|
|
printf '\n' >> "$INVOKED_STEPS"
|
|
exit 0
|
|
fi
|
|
printf '%s\n' "${SELECTED_STEPS[*]}" >> "$INVOKED_STEPS"
|
|
emit_step_result() {
|
|
local n="$1" desc="$2" status="$3"
|
|
printf '[HEAL-STEP] %s %s %s\n' "${n}" "${desc}" "${status}"
|
|
}
|
|
{
|
|
for step in "${SELECTED_STEPS[@]}"; do
|
|
emit_step_result "$step" "fixture step $step" PASS
|
|
done
|
|
want_all=1
|
|
for s in 1 2 3 4 5 6 7; do
|
|
[[ " ${SELECTED_STEPS[*]} " == *" ${s} "* ]] || want_all=0
|
|
done
|
|
if [ "${want_all}" -eq 1 ]; then
|
|
printf '[HEAL-RESULT] PASS all steps passed\n'
|
|
fi
|
|
} >> "$LOG_FILE"
|
|
''')
|
|
script.chmod(0o755)
|
|
self.env["INVOKED_STEPS"] = str(self.directory / "invoked-steps")
|
|
for name in ("Install RustFS package & start cluster", "Preflight checks", "Run heal test (write -> outage -> heal -> verify)"):
|
|
result = self.run_step(name)
|
|
self.assertEqual(result.returncode, 0, result.stderr)
|
|
self.assertEqual(Path(self.env["INVOKED_STEPS"]).read_text().splitlines(), ["1 2", "", "3 4 5 6 7"])
|
|
log = Path(self.env["LOG_FILE"]).read_text()
|
|
self.assertNotIn("[HEAL-RESULT]", log)
|
|
self.assertEqual(log.count("[HEAL-STEP]"), 7)
|
|
report = self.run_step("Generate report")
|
|
self.assertEqual(report.returncode, 0, report.stderr)
|
|
failed_logs = ["\n".join(line for line in log.splitlines() if not line.startswith(f"[HEAL-STEP] {step} ")) + "\n"
|
|
for step in range(1, 8)]
|
|
failed_logs += [
|
|
log.replace("[HEAL-STEP] 3", "[HEAL-STEP] 3 original failure FAIL\n[HEAL-STEP] 3"),
|
|
log + "[HEAL-STEP] 3 later step failure FAIL\n",
|
|
log + "[HEAL-RESULT] FAIL earlier failure\n[HEAL-RESULT] PASS later success\n",
|
|
log.replace("[HEAL-STEP] 4 fixture step 4 PASS", "[HEAL-STEP] 4 fixture step 4 SKIP"),
|
|
]
|
|
for failed_log in failed_logs:
|
|
with self.subTest(log=failed_log):
|
|
Path(self.env["LOG_FILE"]).write_text(failed_log)
|
|
report = self.run_step("Generate report")
|
|
self.assertNotEqual(report.returncode, 0, report.stderr)
|
|
contents = Path(self.env["REPORT_FILE"]).read_text()
|
|
self.assertIn("Test Step Outcome: failure", contents)
|
|
self.assertNotIn("| PASS |", contents)
|
|
if "original failure" in failed_log:
|
|
self.assertIn("| 3 | original failure | FAIL |", (self.artifacts / "steps.md").read_text())
|
|
if "later step failure" in failed_log:
|
|
self.assertIn("| 3 | later step failure | FAIL |", (self.artifacts / "steps.md").read_text())
|
|
|
|
def test_performance_results_version_and_report_are_bound_to_the_run(self):
|
|
self.prepare("performance")
|
|
initialized = self.run_step("Initialize functional evidence")
|
|
self.assertEqual(initialized.returncode, 0, initialized.stderr)
|
|
self.assertEqual(self.env["RUSTFS_RESULT_DIR"], str(self.artifacts / "results"))
|
|
self.assertEqual(self.env["VERSION_FILE"], str(self.artifacts / "version.txt"))
|
|
version = self.run_step("Collect RustFS version info")
|
|
self.assertEqual(version.returncode, 0, version.stderr)
|
|
self.assertIn("fixture-version", Path(self.env["VERSION_FILE"]).read_text())
|
|
old_summary = self.directory / "old-results/summary.md"
|
|
old_summary.parent.mkdir()
|
|
old_summary.write_text("OLD RUN EVIDENCE")
|
|
upload = "Upload report to dashboard (reports/YYYY-MM-DD.md)"
|
|
self.assertNotEqual(self.run_step(upload).returncode, 0)
|
|
self.assertFalse(Path(self.env["REPORT_FILE"]).exists())
|
|
results = Path(self.env["RUSTFS_RESULT_DIR"])
|
|
results.mkdir()
|
|
(results / "summary.md").write_text("CURRENT PERFORMANCE RESULTS\n")
|
|
report = self.run_step(upload)
|
|
self.assertEqual(report.returncode, 0, report.stderr)
|
|
contents = Path(self.env["REPORT_FILE"]).read_text()
|
|
for value in ("actions/runs/314159", "**Attempt**: 2", "**Workflow Commit**: " + self.context["github.sha"],
|
|
"CURRENT PERFORMANCE RESULTS", "fixture-version"):
|
|
self.assertIn(value, contents)
|
|
self.assertNotIn("OLD RUN EVIDENCE", contents)
|
|
|
|
def test_performance_commands_bind_runner_selection_and_preserve_failures(self) -> None:
|
|
self.prepare("performance")
|
|
source = self.source.splitlines()
|
|
job = yaml_block(source, "performance-test", 2)
|
|
runner = WorkflowSteps()
|
|
runner.directory = self.directory / "workspace with spaces"
|
|
scripts = runner.directory / "auto-testing"
|
|
scripts.mkdir(parents=True)
|
|
wrapper = scripts / "rustfs_performance_test.sh"
|
|
wrapper.write_text(f"#!{sys.executable}\nimport json, os, sys\n" +
|
|
"print(json.dumps({'args': sys.argv[1:], 'env': {key: os.environ.get(key) for key in " +
|
|
"('RUSTFS_BENCH_SCRIPT', 'RUSTFS_WARP_METHODS', 'RUSTFS_WARP_SIZES', " +
|
|
"'RUSTFS_WARP_DURATION', 'RUSTFS_WARP_CONCURRENCY', 'WARP_METHODS', " +
|
|
"'WARP_SIZES', 'WARP_DURATION', 'WARP_CONCURRENCY')}}))\n" +
|
|
"sys.exit(int(os.environ['FAKE_BENCH_EXIT']))\n")
|
|
wrapper.chmod(0o755)
|
|
runner.steps = named_steps(job)
|
|
for methods, sizes, duration, concurrency in (
|
|
("get", "1KiB", "1s", "7"), ("all", "all", "5m", "64"), ("", "", "5m", "64")
|
|
):
|
|
runner.context = {"github.workspace": str(runner.directory), "inputs.test_method": methods,
|
|
"inputs.object_size": sizes, "inputs.warp_duration || '5m'": duration,
|
|
"inputs.warp_concurrency || '64'": concurrency}
|
|
runner.env = {**self.env, "RUSTFS_BENCH_SCRIPT": "/unverified/home-script.sh",
|
|
"RUSTFS_WARP_METHODS": "put", "RUSTFS_WARP_SIZES": "64MiB",
|
|
"RUSTFS_WARP_DURATION": "99h", "RUSTFS_WARP_CONCURRENCY": "2",
|
|
"WARP_DURATION": "88h", "WARP_CONCURRENCY": "3", "WARP_METHODS": "mixed", "WARP_SIZES": "32MiB",
|
|
"LOG_FILE": str(self.directory / "suite.log")}
|
|
runner.env.update(runner.step_env(job, indent=4))
|
|
for step, number in (("Run benchmark (GET/PUT/MIXED)", "5"), ("Analyze results", "6")):
|
|
for code in (0, 42):
|
|
with self.subTest(methods=methods, sizes=sizes, step=step, exit=code):
|
|
runner.env["FAKE_BENCH_EXIT"] = str(code)
|
|
result = runner.run_step(step)
|
|
self.assertEqual(result.returncode, code, result.stderr)
|
|
invocation = json.loads(result.stdout)
|
|
expected = ["--step", number, "-y", "--log-file", runner.env["LOG_FILE"]]
|
|
self.assertEqual(invocation["args"], expected)
|
|
self.assertEqual(invocation["env"]["RUSTFS_BENCH_SCRIPT"], str(scripts / "rustfs_performance_testing.sh"))
|
|
self.assertEqual(invocation["env"]["RUSTFS_WARP_METHODS"], methods)
|
|
self.assertEqual(invocation["env"]["RUSTFS_WARP_SIZES"], sizes)
|
|
self.assertEqual(invocation["env"]["RUSTFS_WARP_DURATION"], duration)
|
|
self.assertEqual(invocation["env"]["RUSTFS_WARP_CONCURRENCY"], concurrency)
|
|
if number == "6":
|
|
self.assertEqual(invocation["env"]["WARP_METHODS"], methods)
|
|
self.assertEqual(invocation["env"]["WARP_SIZES"], sizes)
|
|
self.assertEqual(invocation["env"]["WARP_DURATION"], duration)
|
|
self.assertEqual(invocation["env"]["WARP_CONCURRENCY"], concurrency)
|
|
|
|
|
|
if __name__ == "__main__":
|
|
unittest.main()
|