mirror of
https://github.com/rustfs/rustfs.git
synced 2026-10-04 04:21:35 +00:00
docs(storage): verify object generation authority contract (#8322)
This commit is contained in:
@@ -1,7 +1,7 @@
|
||||
# Object Generation Authority And Recovery Contract
|
||||
|
||||
**Use this when:** changing object commit fencing, rollback, old-directory cleanup, prepared reads, quota settlement, or the metadata and RPC fields used by those operations.
|
||||
**Source of truth:** `crates/ecstore/src/set_disk/ops/object.rs` (`assign_object_transaction_epoch`, `verify_object_transaction_epoch_fence`); `crates/ecstore/src/set_disk/core/io_primitives.rs` (`rename_data_owned_with_fence`, `commit_rename_data_dir`); `crates/ecstore/src/disk/local.rs` (`rename_data`, `write_all_meta`); `crates/lock/src/distributed_lock.rs` (`DistributedLockGuard`, `LockLostSignal`). The implementation boundary below distinguishes existing behavior from the selected design.
|
||||
**Source of truth:** `crates/ecstore/src/set_disk/ops/object.rs` (`assign_object_transaction_epoch`, `verify_object_transaction_epoch_fence`); `crates/ecstore/src/set_disk/core/io_primitives.rs` (`rename_data_owned_with_fence`, `commit_rename_data_dir`); `crates/ecstore/src/disk/local.rs` (`rename_data`, `write_all_meta`), with publication and rollback helpers in `crates/ecstore/src/disk/local/commit.rs`; `crates/lock/src/distributed_lock.rs` (`DistributedLockGuard`, `LockLostSignal`). The implementation boundary below distinguishes existing behavior from the selected design.
|
||||
|
||||
## Decision And Implementation Boundary
|
||||
|
||||
@@ -82,7 +82,7 @@ A single filesystem rename does not atomically commit a sidecar plus `xl.meta`.
|
||||
|
||||
A laggard need not contain the predecessor. Recovery fetches the chosen decision and its verified metadata, reconstructs or validates its shards under [erasure-coding.md](erasure-coding.md), and installs that state. An empty replacement uses a fresh disk incarnation and cannot reuse old preparation receipts. A disk whose durable state claims a later decision than the supplied one rejects the operation; a conflicting same-revision digest is quarantined. Never erase a divergent disk merely because it is in the minority.
|
||||
|
||||
`rollback_committed_rename_std`, `rollback_inline_metadata_commit_std`, and `restore_metadata_backup` in `crates/ecstore/src/disk/local.rs` must become decision-aware before strict mode includes them. The permitted rollback is limited to a transaction's unchosen private preparation, or restoration proven by recovery to be necessary before any newer local effect. A chosen operation is repaired forward. Neither a client timeout nor a rename-tail error permits reverting an acknowledged decision.
|
||||
`rollback_committed_rename_std` and `rollback_inline_metadata_commit_std` in `crates/ecstore/src/disk/local/commit.rs`, and `restore_metadata_backup` in `crates/ecstore/src/disk/local.rs`, must become decision-aware before strict mode includes them. The permitted rollback is limited to a transaction's unchosen private preparation, or restoration proven by recovery to be necessary before any newer local effect. A chosen operation is repaired forward. Neither a client timeout nor a rename-tail error permits reverting an acknowledged decision.
|
||||
|
||||
`RenameConvergence` remains a post-publication repair signal. `PartialCommit`, `SignatureDivergent`, and `Unknown` do not decide which transaction won. Keep their diagnostics and quorum accounting; resolve authority first. Early ACK may still precede minority-tail completion after the two quorum conditions hold. Tests that inspect all disks must synchronize the tail or assert the permitted minority residue separately.
|
||||
|
||||
@@ -104,6 +104,10 @@ Every semantic metadata change advances the whole-object generation, including m
|
||||
|
||||
Strict capability is withheld until every raw writer in `DiskAPI`, its local implementation, `DiskStore`, remote adapters, and server handlers has an enforced path. Internal authority persistence must use its own narrow storage primitive, not evade this rule by recursively calling generic object metadata writes.
|
||||
|
||||
The participation table includes semantic entry points, not just final PUT rename. Its low-level audit must also include `write_all`, `compare_and_update_file`, `rename_file`, `delete`, `delete_paths`, bucket creation/deletion, and streaming writers whenever their resolved source or destination can affect an enrolled object's live identity. An internal-volume name or `undo_write` flag is not an exemption. Upload-private staging and proven unreferenced retirement may be exempt only at a validated boundary; the request cannot declare its own exemption.
|
||||
|
||||
Existing `NamespaceCommitGuard` accounting is a separate scanner observation mechanism. PUT, MPU complete, single DELETE, and instance-bound target rename/undo retain some physical owners, but that count neither chooses an object decision nor excludes an admitted scanner publication. `write_unique_file_info` and `update_object_meta_with_opts` still fan out ordinary metadata methods. A writer roster is not implementation evidence: each row requires a strict admission path and a real test before the capability can be advertised.
|
||||
|
||||
## Reads, Garbage Collection, And Accounting
|
||||
|
||||
Strict reads need the chosen head, not a majority of arbitrary prepared/live UUIDs. Resolve the decision under the namespace read lock before accepting object metadata; validate the selected current or explicit version against it, and wait for or repair missing materialization. HEAD, GET, ListObjects/ListObjectVersions, scanner reads used for deletion, and prepared pool reads all need this distinction. A query may return an error while a chosen write is recovering; it must not expose an unchosen candidate or resurrect a retired version. This read-decision adapter is part of the strict-mode scope and is a reason E03 is larger than disk CAS.
|
||||
@@ -124,12 +128,16 @@ For strict mode, add the decision identity to the reservation/settlement binding
|
||||
|
||||
Authority configuration is durable and includes participant identities, routing, bucket incarnation, protocol version, and quorum rules. Changing storage pool placement must not remap the authority key. A replaced voter starts as a non-voter, catches up durable promise/accepted/chosen state, and only joins through a quorum-approved configuration transition. Configuration change requires intersecting old/new decision quorums; losing the old quorum is a recovery incident, not permission to bootstrap a new empty authority. Offline data disks rejoin through generation-aware catch-up, independent of voter admission.
|
||||
|
||||
The first implementation uses a fixed authority voter configuration. Online voter-set changes are unsupported and fail closed until a separately verified joint-quorum transition exists; simply taking a majority in the new set is forbidden. A new or replaced voter cannot count toward Qlock before catch-up, and a lost/corrupt persistence store cannot resume under its previous voter identity. Restart of an intact voter preserves promises and accepts, changes its transport boot epoch, and remains unavailable until recovery finishes. Data-disk replacement can proceed under the unchanged voter configuration with a fresh disk incarnation, verified catch-up, and a renewed complete capability proof. Unavailability does not authorize discarding the old configuration or its decisions.
|
||||
|
||||
A live capability proof must bind the authority configuration, topology, every participant process boot epoch, disk incarnations, writer/read/recovery protocol support, encoding version, and RPC signature/body/replay strictness. Restart, membership change, disk replacement, protocol downgrade, or a strict-transport setting change revokes it. Receivers revalidate before entering the publication critical section; accepted durable decisions survive proof revocation and are recovered under a fresh valid proof, never replayed as unvalidated requests. The remote-version-state fleet proof does not prove any of these generation capabilities.
|
||||
|
||||
The existing `RUSTFS_OBJECT_TRANSACTION_FENCING_WRITE` and `RUSTFS_OBJECT_TRANSACTION_FENCING_FLEET_CONFIRMED` flags in `crates/config/src/constants/object.rs` retain their current default-off behavior. They do not become a claim that the new protocol exists. If generation strictness is explicitly selected, missing capability is an error, never silent downgrade. No new environment variable is introduced by this document; a production gate must be documented with its implementation.
|
||||
|
||||
Strict enrollment requires quiescing old writers and readers for the enrolled namespace, recovering ambiguous operations, validating/importing legacy heads, persisting a strict-format/protocol marker, and enabling the complete fleet. New disks reject unbound legacy mutation RPCs for that namespace. Old binaries must be prevented from opening a strict-enrolled drive by a startup compatibility gate they understand before enrollment; an environment flag known only to new binaries is insufficient. Until that prerequisite is deployed, do not activate strict mode in a mixed fleet. Disabling flags after enrollment cannot drop durable authority; downgrade requires a separately verified quiescent materialization/export operation. Ordinary un-enrolled compatibility deployments keep their current behavior.
|
||||
|
||||
During a revoked or incomplete proof, strict reads and writes fail before returning a head or admitting publication. They cannot fall back to metadata-majority selection, UUID equality, an old scanner token, or the remote-version-state proof. Readiness requires decoded/reconciled voter and disk records, not just a new boot ID or pending count of zero. A chosen operation remains recoverable forward under a fresh proof; revocation does not change it into an abort. Supporting old binaries that lack the startup barrier is a deployment prerequisite, not permission to enroll them and rely on server-side checks.
|
||||
|
||||
## Encoding And Transport
|
||||
|
||||
- Do not bump `XL_META_VERSION` or `XL_HEADER_VERSION` in `crates/filemeta/src/filemeta.rs`. Do not add fields to positional-msgpack `FileInfo`; carry the UUID through the metadata map and protocol records through explicit versioned envelopes.
|
||||
@@ -168,6 +176,23 @@ These are durable ownership and acceptance boundaries, not permission to close t
|
||||
|
||||
The conservative immediate action is to keep the existing compatibility behavior and improve its local convergence/recovery independently. Those fixes must describe their smaller guarantee and must not advertise E03's distributed safety. A strict-only local CAS helper can be built behind the inactive capability boundary, but it cannot enable the feature or close the authority work.
|
||||
|
||||
## Executable Design Model And Acceptance Mapping
|
||||
|
||||
Run `python3 scripts/test_object_generation_protocol_model.py`. This standalone standard-library model implements the selected promise/accept choice rule for one successor slot with four durable voters, Qlock=3, two distinct ordered ballots, and two candidate values. It enumerates all enabled atomic prepare/reply and accept transitions within that bound, checks that both candidates can win, and rejects any execution choosing two values. Removing highest-accepted adoption must produce a double-choice counterexample; replacing voters with empty state also demonstrates why catch-up is mandatory.
|
||||
|
||||
Directed cases cover lost ACK followed by a fresh higher ballot, minority accepted-value recovery, durable promises across restart, same-ballot byte disagreement, unavailable decision quorum, and an isolated voter that has not learned the newer promise. The latter is intentionally allowed to retain older work: it cannot turn that work into another chosen value. Publication/rollback/retirement and null/delete-marker examples are abstract contract checks, not calls into RustFS. Durable records are modeled as atomic; reply delays/loss are not exhaustively enumerated, proposer identity allocation and Wdata receipts are assumed valid, and multiple slots, reconfiguration, codecs, real filesystems, transport, and performance are outside the exhaustive bound. Passing this model is design evidence only.
|
||||
|
||||
| Acceptance boundary | Design obligation | Required implementation evidence |
|
||||
|---|---|---|
|
||||
| Stale rename after a quorum commit | One chosen successor; no applied local revision regression; isolated older materialization has no current-authority vote. Immediate rejection on an uncontacted disk is not required. | Real lock-RPC partition with data RPC reachable, decoded decisions/metadata, and complete GET bytes. |
|
||||
| ACK loss, competing proposals, minority recovery | Preserve chosen values and adopt the highest accepted value; resolve the prior slot before choosing a successor. | Actual persisted voter transitions, restart/recovery before readiness, exact operation/digest retry results, and multi-slot conflict tests. |
|
||||
| Rename/rollback/cleanup and process death | Chosen work repairs forward; stale rollback cannot restore a winner's backup; GC requires retirement plus current reference/reader checks. | Inline/non-inline and null/version/marker disk tests at every sync/rename/replace/crash point, with directory trees and decoded records. |
|
||||
| All writers and strict reads | Every table row consumes authority; raw mutations cannot bypass enrollment; readers resolve the chosen head before serving metadata/bytes. | PUT/MPU/delete/heal/tier/replication/movement/raw-RPC coverage, prepared-read and quota bindings, and proof of no unguarded live entry point. |
|
||||
| Membership, mixed fleet, proof revocation | Fixed voters initially; no empty replacement vote, metadata fallback, or silent strict downgrade; startup barrier precedes enrollment. | Real old/new binaries, restart/rejoin and disk replacement, authenticated JSON/msgpack parity, body/method/replay failures, and recovery readiness checks. |
|
||||
| Availability and performance | Preserve Qlock and Wdata independently; do not replace quorum availability with N-disk waits. | N=4/W=3 at W-1/W/W+1 and reproducible 4 KiB/1 MiB/MPU measurements against the activation budgets below. |
|
||||
|
||||
The model does not close any row's implementation evidence. In particular, scanner M/N baseline checks and physical pending owners do not close the final admission or remote attempt/expiry/drain protocol. That scanner boundary remains separate from the object decision protocol, and both must be enforced if a publication participates in both domains. A design review may accept this availability contract while strict activation remains blocked on every unfinished implementation row.
|
||||
|
||||
## Performance And Activation Criteria
|
||||
|
||||
Measure the existing implementation and the full proposed path on identical machines, disk/filesystem, durability settings, network, object population, concurrency, and warmup. Include single hot-key and many-key 4 KiB PUT, 1 MiB PUT, metadata-only writes, and CompleteMultipartUpload with fixed part counts. Report throughput, p50/p95/p99, peak retained preparation/recovery bytes, recovery time, per-operation RPCs/fsyncs, and the object mutation critical-section duration. Include a slow minority disk, one voter loss, and restart recovery; a throughput result alone is insufficient.
|
||||
|
||||
@@ -50,6 +50,7 @@ their issue closes.
|
||||
| `prepare_replacement_migration.py` | dev-tool | Prepares digest-bound schema 5/6 replacement maintenance approvals | [Replacement recovery](../docs/operations/replacement-generation-recovery.md) |
|
||||
| `test_prepare_replacement_migration.py` | dev-tool | Verifies maintenance approval scope, publication, and stopped-writer assertion | Python unittest; same runbook |
|
||||
| `test_diagnose_scanner_enumeration_restart.py` | dev-tool | Driver report validation and positive convergence oracle tests | Python unittest; same guide |
|
||||
| `test_object_generation_protocol_model.py` | dev-tool | Bounded single-slot, four-voter promise/accept design model and recovery/retirement examples; no production runtime evidence | `python3 scripts/test_object_generation_protocol_model.py`; [Generation contract](../docs/architecture/unified-object-generation.md#executable-design-model-and-acceptance-mapping) |
|
||||
| `e2e-run.sh` | ci-gate | Boots a rustfs server and runs the `s3s-e2e` black-box conformance tool against it | ci.yml `e2e-tests` jobs; `docs/testing/README.md` |
|
||||
| `run_ecstore_validation_suite.sh` | dev-tool | ecstore black-box validation suite (`quick`/`full`/`destructive`/`fuzz` profiles) | `docs/testing/README.md`, `docs/testing/ecstore-validation-suite-design.md` |
|
||||
| `run_e2e_tests.sh` | dev-tool | Local `e2e_test` crate runner (starts a server, applies filters, cleans up) | `crates/e2e_test/README.md` |
|
||||
|
||||
@@ -0,0 +1,267 @@
|
||||
#!/usr/bin/env python3
|
||||
# Copyright 2026 RustFS Team
|
||||
# SPDX-License-Identifier: Apache-2.0
|
||||
"""Bounded design model, not RustFS's lock, disk, or RPC implementation.
|
||||
|
||||
Run with ``python3 scripts/test_object_generation_protocol_model.py``. The
|
||||
exhaustive check covers one successor slot, four durable voters, two distinct
|
||||
ballots/candidates, and every enabled prepare/accept ordering in that bound.
|
||||
Durable transitions are atomic; filesystem crash atomicity is an implementation
|
||||
obligation, not something this model can establish.
|
||||
"""
|
||||
|
||||
from collections import deque
|
||||
from dataclasses import dataclass, replace
|
||||
import unittest
|
||||
|
||||
|
||||
VOTERS = 4
|
||||
QUORUM = 3
|
||||
EMPTY = -1
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class Voter:
|
||||
promise: int = 0
|
||||
accepted_ballot: int = 0
|
||||
accepted_value: int = EMPTY
|
||||
|
||||
def prepare(self, ballot):
|
||||
if ballot < self.promise:
|
||||
return self, None
|
||||
persisted = replace(self, promise=ballot)
|
||||
return persisted, (persisted.accepted_ballot, persisted.accepted_value)
|
||||
|
||||
def accept(self, ballot, value):
|
||||
if ballot < self.promise:
|
||||
return self, False
|
||||
if self.accepted_ballot == ballot and self.accepted_value != value:
|
||||
return self, False
|
||||
return Voter(ballot, ballot, value), True
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class Proposer:
|
||||
prepare_attempts: int = 0
|
||||
replies: tuple = ()
|
||||
value: int = EMPTY
|
||||
accept_attempts: int = 0
|
||||
accepts: int = 0
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class State:
|
||||
voters: tuple = (Voter(),) * VOTERS
|
||||
proposers: tuple = (Proposer(), Proposer())
|
||||
chosen: int = 0
|
||||
|
||||
|
||||
def successors(state, adopt_accepted=True):
|
||||
"""A reply is sent only after persistence; arbitrary losses are tested below.
|
||||
|
||||
The first successful promise quorum fixes the candidate. A failed request
|
||||
need not retry in this bounded run. Preparing or accepting at a higher
|
||||
ballot after a restart is a separate recovery run.
|
||||
"""
|
||||
for index, proposer in enumerate(state.proposers):
|
||||
ballot = index + 1
|
||||
for disk in range(VOTERS):
|
||||
bit = 1 << disk
|
||||
voters = list(state.voters)
|
||||
if proposer.value == EMPTY:
|
||||
if proposer.prepare_attempts & bit:
|
||||
continue
|
||||
voters[disk], reply = voters[disk].prepare(ballot)
|
||||
replies = proposer.replies + ((reply,) if reply is not None else ())
|
||||
value = EMPTY
|
||||
if len(replies) == QUORUM:
|
||||
accepted = max(replies)
|
||||
value = accepted[1] if adopt_accepted and accepted[0] else index
|
||||
updated = replace(proposer, prepare_attempts=proposer.prepare_attempts | bit, replies=replies, value=value)
|
||||
event = (index, "prepare", disk)
|
||||
chosen = state.chosen
|
||||
else:
|
||||
if proposer.accept_attempts & bit:
|
||||
continue
|
||||
voters[disk], accepted = voters[disk].accept(ballot, proposer.value)
|
||||
accepts = proposer.accepts | bit if accepted else proposer.accepts
|
||||
updated = replace(proposer, accept_attempts=proposer.accept_attempts | bit, accepts=accepts)
|
||||
chosen = state.chosen
|
||||
if accepts.bit_count() >= QUORUM:
|
||||
chosen |= 1 << proposer.value
|
||||
event = (index, "accept", disk)
|
||||
proposers = list(state.proposers)
|
||||
proposers[index] = updated
|
||||
yield State(tuple(voters), tuple(proposers), chosen), event
|
||||
|
||||
|
||||
def explore(adopt_accepted=True):
|
||||
initial = State()
|
||||
queue = deque([initial])
|
||||
parents = {initial: None}
|
||||
transitions = 0
|
||||
terminal = 0
|
||||
chosen_values = set()
|
||||
while queue:
|
||||
state = queue.popleft()
|
||||
if state.chosen.bit_count() > 1:
|
||||
trace = []
|
||||
cursor = state
|
||||
while parents[cursor] is not None:
|
||||
previous, event = parents[cursor]
|
||||
trace.append(event)
|
||||
cursor = previous
|
||||
return len(parents), transitions, terminal, chosen_values, tuple(reversed(trace))
|
||||
if state.chosen:
|
||||
chosen_values.add(state.chosen)
|
||||
enabled = 0
|
||||
for successor, event in successors(state, adopt_accepted):
|
||||
enabled += 1
|
||||
transitions += 1
|
||||
if successor not in parents:
|
||||
parents[successor] = (state, event)
|
||||
queue.append(successor)
|
||||
terminal += enabled == 0
|
||||
return len(parents), transitions, terminal, chosen_values, None
|
||||
|
||||
|
||||
def choose(voters, ballot, candidate, prepare_ids, accept_ids, adopt=True):
|
||||
"""One recovery round; callers may deliberately lose every response."""
|
||||
for ids in (prepare_ids, accept_ids):
|
||||
if len(set(ids)) != len(ids) or any(disk < 0 or disk >= VOTERS for disk in ids):
|
||||
raise ValueError("a quorum contains distinct configured voters only")
|
||||
voters = list(voters)
|
||||
replies = []
|
||||
for disk in prepare_ids:
|
||||
voters[disk], reply = voters[disk].prepare(ballot)
|
||||
if reply is not None:
|
||||
replies.append(reply)
|
||||
if len(replies) < QUORUM:
|
||||
return tuple(voters), None, False
|
||||
accepted = max(replies)
|
||||
value = accepted[1] if adopt and accepted[0] else candidate
|
||||
count = 0
|
||||
for disk in accept_ids:
|
||||
voters[disk], success = voters[disk].accept(ballot, value)
|
||||
count += success
|
||||
return tuple(voters), value, count >= QUORUM
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class Disk:
|
||||
applied_revision: int = 0
|
||||
operation: str = "bootstrap"
|
||||
referenced: frozenset = frozenset()
|
||||
|
||||
def publish_chosen(self, revision, operation, directories):
|
||||
if revision < self.applied_revision:
|
||||
return self, False
|
||||
next_state = Disk(revision, operation, frozenset(directories))
|
||||
if revision == self.applied_revision and next_state != self:
|
||||
return self, False
|
||||
return next_state, True
|
||||
|
||||
def rollback_unchosen(self, operation, chosen):
|
||||
return not chosen and operation == self.operation
|
||||
|
||||
def may_retire(self, directory, authorized, reader_protected, accepted_references):
|
||||
return authorized and directory not in self.referenced and directory not in accepted_references and not reader_protected
|
||||
|
||||
|
||||
class ObjectGenerationProtocolModel(unittest.TestCase):
|
||||
def test_all_bounded_two_proposer_interleavings_choose_at_most_one_value(self):
|
||||
states, transitions, terminal, values, violation = explore()
|
||||
self.assertIsNone(violation, violation)
|
||||
self.assertEqual(values, {1, 2}, "both original candidates must be reachable")
|
||||
self.assertGreater(terminal, 0)
|
||||
print(f"model: states={states}, transitions={transitions}, terminal={terminal}, chosen_candidates={len(values)}")
|
||||
|
||||
def test_omitting_highest_accepted_adoption_has_a_double_choice_counterexample(self):
|
||||
_, _, _, _, violation = explore(adopt_accepted=False)
|
||||
self.assertIsNotNone(violation, "the model must detect removing the safety rule")
|
||||
print(f"negative control: double-choice trace={violation}")
|
||||
|
||||
def test_lost_ack_and_restart_preserve_the_chosen_operation(self):
|
||||
voters, value, chosen = choose((Voter(),) * VOTERS, 1, 0, (0, 1, 2), (0, 1, 2))
|
||||
self.assertTrue(chosen)
|
||||
self.assertEqual(value, 0)
|
||||
# A reboot drops proposer memory/replies, never the voter's persisted state.
|
||||
restarted = tuple(Voter(v.promise, v.accepted_ballot, v.accepted_value) for v in voters)
|
||||
_, recovered, chosen = choose(restarted, 2, 1, (1, 2, 3), (1, 2, 3))
|
||||
self.assertTrue(chosen)
|
||||
self.assertEqual(recovered, 0)
|
||||
|
||||
def test_minority_accepted_work_is_adopted_and_finished(self):
|
||||
voters, _, chosen = choose((Voter(),) * VOTERS, 1, 0, (0, 1, 2), (0, 1))
|
||||
self.assertFalse(chosen)
|
||||
_, recovered, chosen = choose(voters, 2, 1, (1, 2, 3), (1, 2, 3))
|
||||
self.assertTrue(chosen)
|
||||
self.assertEqual(recovered, 0)
|
||||
|
||||
def test_equal_ballot_cannot_accept_different_bytes_and_lower_ballot_is_rejected(self):
|
||||
voter, accepted = Voter().accept(2, 0)
|
||||
self.assertTrue(accepted)
|
||||
self.assertEqual(voter.accept(2, 1), (voter, False))
|
||||
self.assertEqual(voter.accept(1, 0), (voter, False))
|
||||
|
||||
def test_crash_after_promise_before_reply_does_not_restore_an_older_ballot(self):
|
||||
persisted, _ = Voter().prepare(2)
|
||||
rebooted = Voter(persisted.promise, persisted.accepted_ballot, persisted.accepted_value)
|
||||
self.assertEqual(rebooted.accept(1, 0), (rebooted, False))
|
||||
|
||||
def test_replacing_voters_with_empty_state_can_choose_a_conflicting_value(self):
|
||||
voters, original, chosen = choose((Voter(),) * VOTERS, 1, 0, (0, 1, 2), (0, 1, 2))
|
||||
self.assertTrue(chosen)
|
||||
# This deliberately violates enrollment: restored voters must catch up,
|
||||
# rather than reuse the old identity with an empty promise/accepted log.
|
||||
replaced = (Voter(), Voter(), voters[2], voters[3])
|
||||
_, replacement, chosen = choose(replaced, 2, 1, (0, 1, 3), (0, 1, 3))
|
||||
self.assertTrue(chosen)
|
||||
self.assertNotEqual(original, replacement)
|
||||
|
||||
def test_an_unavailable_decision_quorum_cannot_choose_a_candidate(self):
|
||||
_, candidate, chosen = choose((Voter(),) * VOTERS, 1, 0, (0, 1), (0, 1, 2, 3))
|
||||
self.assertIsNone(candidate)
|
||||
self.assertFalse(chosen)
|
||||
|
||||
def test_duplicate_or_unconfigured_voters_cannot_form_a_quorum(self):
|
||||
for invalid_ids in ((0, 0, 0), (-1, 0, 1), (0, 1, VOTERS)):
|
||||
for prepare_ids, accept_ids in ((invalid_ids, (0, 1, 2)), ((0, 1, 2), invalid_ids)):
|
||||
with self.subTest(prepare=prepare_ids, accept=accept_ids):
|
||||
with self.assertRaisesRegex(ValueError, "distinct configured voters"):
|
||||
choose((Voter(),) * VOTERS, 1, 0, prepare_ids, accept_ids)
|
||||
|
||||
def test_quorum_choice_does_not_require_an_uncontacted_disk_to_learn_it(self):
|
||||
voters, _, chosen = choose((Voter(),) * VOTERS, 2, 1, (0, 1, 2), (0, 1, 2))
|
||||
self.assertTrue(chosen)
|
||||
isolated, accepted = voters[3].accept(1, 0)
|
||||
self.assertTrue(accepted, "an isolated voter has not learned the newer promise")
|
||||
recovered = list(voters)
|
||||
recovered[3] = isolated
|
||||
_, value, chosen = choose(recovered, 3, 0, (1, 2, 3), (1, 2, 3))
|
||||
self.assertTrue(chosen)
|
||||
self.assertEqual(value, 1, "isolated older work cannot replace the chosen value")
|
||||
|
||||
def test_stale_publication_rollback_and_gc_cannot_change_the_winner(self):
|
||||
winner, published = Disk().publish_chosen(2, "B", {"B-data"})
|
||||
self.assertTrue(published)
|
||||
self.assertEqual(winner.publish_chosen(1, "A", {"A-data"}), (winner, False))
|
||||
self.assertEqual(winner.publish_chosen(2, "A", {"A-data"}), (winner, False))
|
||||
self.assertFalse(winner.rollback_unchosen("A", chosen=False))
|
||||
self.assertFalse(winner.rollback_unchosen("B", chosen=True))
|
||||
self.assertFalse(winner.may_retire("B-data", True, False, frozenset()))
|
||||
self.assertFalse(winner.may_retire("old-data", True, True, frozenset()))
|
||||
self.assertFalse(winner.may_retire("accepted-data", True, False, {"accepted-data"}))
|
||||
self.assertTrue(winner.may_retire("old-data", True, False, frozenset()))
|
||||
|
||||
def test_tombstone_and_null_marker_payloads_do_not_reset_authority_revision(self):
|
||||
current = Disk()
|
||||
for revision, operation, directories in ((1, "null", {"null-data"}), (2, "marker", set()), (3, "last-version-delete", set())):
|
||||
current, published = current.publish_chosen(revision, operation, directories)
|
||||
self.assertTrue(published)
|
||||
self.assertEqual(current.applied_revision, revision)
|
||||
self.assertEqual(current.publish_chosen(0, "bootstrap", set()), (current, False))
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
unittest.main(verbosity=2)
|
||||
Reference in New Issue
Block a user