mirror of
https://github.com/rustfs/rustfs.git
synced 2026-08-30 08:49:26 +00:00
fix(table-catalog): harden strong backing compatibility (#5941)
* fix(table-catalog): harden strong backing compatibility * fix(table-catalog): close strong backing recovery gaps * fix(table-catalog): harden strong backing recovery * fix(table-catalog): repair strong backing CI failures * fix(table-catalog): satisfy test clippy lint --------- Co-authored-by: Henry Guo <marshawcoco@users.noreply.github.com>
This commit is contained in:
@@ -13,6 +13,9 @@ for later deletion.
|
|||||||
## Open Items
|
## Open Items
|
||||||
|
|
||||||
- `table-publication-fence-v1` S3 Tables publication fencing: nodes that predate table and table-bucket publication fences can mutate live files while a new node is publishing a catalog pointer. New nodes retain exact object guards until the operator confirms that every serving node uses the new fences. Fleet confirmation also requires non-overlapping active warehouse prefixes and lifecycle workers that exclude table buckets. Remove the exact live-file fallback and the fleet-confirmation gate after the minimum supported RustFS release acquires table fences for registered-table mutations and table-bucket fences for unresolved-prefix mutations.
|
- `table-publication-fence-v1` S3 Tables publication fencing: nodes that predate table and table-bucket publication fences can mutate live files while a new node is publishing a catalog pointer. New nodes retain exact object guards until the operator confirms that every serving node uses the new fences. Fleet confirmation also requires non-overlapping active warehouse prefixes and lifecycle workers that exclude table buckets. Remove the exact live-file fallback and the fleet-confirmation gate after the minimum supported RustFS release acquires table fences for registered-table mutations and table-bucket fences for unresolved-prefix mutations.
|
||||||
|
- `table-catalog-strong-snapshot-v1` durable strong catalog snapshot compatibility: mixed-version deployments continue writing version 1 snapshots until operators confirm that every serving node reads version 2, and version 1 table/view identifier collisions remain available only for cleanup. Remove version 1 writes and collision cleanup after the minimum supported RustFS release reads version 2 and every retained durable strong snapshot is collision-free and has been upgraded to version 2.
|
||||||
|
- `table-catalog-migration-fence-v1` durable strong migration fence compatibility: version 1 "PREPARING" fences did not distinguish a known-absent global strong snapshot from an unknown baseline, so retries read them but fail closed if the global snapshot is missing. Version 2 preserves the same JSON shape and records the pre-migration global snapshot ETag in the existing target_snapshot_etag field while the fence is "PREPARING". Remove version 1 reads after every supported direct-upgrade source writes version 2 fences and operators have completed or cancelled every older in-progress backing migration.
|
||||||
|
- `table-catalog-backing-manifest-v1-wire-labels` durable strong backing manifest labels: version 1 published "STRONG_KV_WAL" and "CUT_OVER_LINEARIZABLE_READS" before the implementation was narrowed to the ETag-CAS durable snapshot backing. Internal names and operator documentation describe the implemented semantics, while version 1 responses retain those labels for existing clients. Replace the labels only in a new manifest version with an explicit client migration contract.
|
||||||
- `cross-pool-fence-v1` authenticated unsupported advertisement: predeployment servers recognize the versioned cross-pool fence capability probe but report support version 0, allowing a later all-peer probe to distinguish predeployment nodes without activating a second lock domain. Replace the unsupported advertisement only when composite lock acquisition, a cluster-wide activation fence, complete fleet proof, commit-time proof revalidation, and fail-closed revocation ship together.
|
- `cross-pool-fence-v1` authenticated unsupported advertisement: predeployment servers recognize the versioned cross-pool fence capability probe but report support version 0, allowing a later all-peer probe to distinguish predeployment nodes without activating a second lock domain. Replace the unsupported advertisement only when composite lock acquisition, a cluster-wide activation fence, complete fleet proof, commit-time proof revalidation, and fail-closed revocation ship together.
|
||||||
- `table-catalog-dotted-namespace` Iceberg REST namespace path compatibility: existing RustFS clients use dotted namespace paths, while the standard multi-level contract uses the URL-encoded unit separator `%1F`. New servers accept both forms so a rolling upgrade does not invalidate existing catalog configuration. Remove the dotted fallback after the minimum supported RustFS release advertises `%1F` and all supported clients have refreshed their catalog configuration.
|
- `table-catalog-dotted-namespace` Iceberg REST namespace path compatibility: existing RustFS clients use dotted namespace paths, while the standard multi-level contract uses the URL-encoded unit separator `%1F`. New servers accept both forms so a rolling upgrade does not invalidate existing catalog configuration. Remove the dotted fallback after the minimum supported RustFS release advertises `%1F` and all supported clients have refreshed their catalog configuration.
|
||||||
- `rustfs-5509` FileInfo positional MessagePack decoding: beta.11 serialized 28 fields, while beta.12 inserted transition-version fields in the middle and serialized an incompatible 30-field array. New releases write named maps and retain readers for both shipped array layouts so direct and rolling upgrades can read either release. Remove the positional-array readers after every supported direct-upgrade release writes named maps and no retained RPC payload can contain a pre-map FileInfo array.
|
- `rustfs-5509` FileInfo positional MessagePack decoding: beta.11 serialized 28 fields, while beta.12 inserted transition-version fields in the middle and serialized an incompatible 30-field array. New releases write named maps and retain readers for both shipped array layouts so direct and rolling upgrades can read either release. Remove the positional-array readers after every supported direct-upgrade release writes named maps and no retained RPC payload can contain a pre-map FileInfo array.
|
||||||
|
|||||||
@@ -116,12 +116,13 @@ catalog extension.
|
|||||||
| Commit publication fencing | Supported with rolling-upgrade gate | Existing deployments retain exact object guards so older writers cannot mutate referenced files during publication. Set `RUSTFS_TABLE_CATALOG_PUBLICATION_FENCE_FLEET_CONFIRMED=true` only after every serving node supports table and table-bucket publication fences. In scalable mode, active table warehouse prefixes must not overlap, ordinary lifecycle expiry remains disabled for table buckets, and first enablement, first publication, drop, and warehouse relocation are serialized by the table-bucket fence. |
|
| Commit publication fencing | Supported with rolling-upgrade gate | Existing deployments retain exact object guards so older writers cannot mutate referenced files during publication. Set `RUSTFS_TABLE_CATALOG_PUBLICATION_FENCE_FLEET_CONFIRMED=true` only after every serving node supports table and table-bucket publication fences. In scalable mode, active table warehouse prefixes must not overlap, ordinary lifecycle expiry remains disabled for table buckets, and first enablement, first publication, drop, and warehouse relocation are serialized by the table-bucket fence. |
|
||||||
| Post-CAS finalization recovery | Supported | Diagnostics and recovery can repair stale or missing idempotency indexes without changing the current table pointer. |
|
| Post-CAS finalization recovery | Supported | Diagnostics and recovery can repair stale or missing idempotency indexes without changing the current table pointer. |
|
||||||
| Catalog export | Supported | Exposes table state, commit recovery state, and backing migration information for operator inspection. |
|
| Catalog export | Supported | Exposes table state, commit recovery state, and backing migration information for operator inspection. |
|
||||||
| Strong backing state transfer | Supported | Object-backed table bucket, namespace, table, view, commit-log, and idempotency state can be materialized into the durable strong snapshot. The transfer is deterministic, ETag-CAS protected, idempotent after an interrupted finalization, and fails closed when a table or view has no owning namespace entry. |
|
| Strong backing state transfer | Supported | Object-backed table bucket, namespace, table, view, commit-log, and idempotency state can be materialized into the durable strong snapshot. The transfer is deterministic, ETag-CAS protected, idempotent after an interrupted finalization, validates candidate state through the restart decoder before publication, preserves resource-backed implicit namespaces, and fails closed when an inactive explicit namespace conflicts with active descendants or resources. Snapshot hydration requires a stable non-empty ETag, caps the encoded snapshot at 64 MiB, shares state and reload serialization across requests in one server context, and rejects disappearance or format-version rollback after observation. Configured durable-strong mode rejects a missing snapshot on its first catalog access after startup; only object-backed migration may initialize an empty target. |
|
||||||
| Durable backing migration preflight | Supported | `GET /iceberg/v1/{warehouse}/catalog/migration` and the `/_iceberg/v1` alias inspect object-backed catalog inventory, recovery blockers, warehouse prefix index readiness, persistent write-fence state, target snapshot agreement, and whether every table bucket is ready for cutover. |
|
| Durable backing migration preflight | Supported | `GET /iceberg/v1/{warehouse}/catalog/migration` and the `/_iceberg/v1` alias inspect object-backed catalog inventory, recovery blockers, warehouse prefix index readiness, active table/view identifier collisions, persistent write-fence state, target snapshot agreement, and whether every table bucket is ready for cutover. |
|
||||||
| Durable backing migration execution | Preview / controlled | `POST /iceberg/v1/{warehouse}/catalog/migration` fences table-bucket registry changes, acquires a persistent per-bucket write fence, drains in-flight catalog mutations, materializes the target snapshot, and reports `ready_to_enable_durable_strong`. `DELETE` safely releases the bucket fence only while its target state has not advanced, and releases the registry fence after the last bucket is cancelled. Both mutations require `admin:MigrateTableCatalog`. |
|
| Durable backing migration execution | Preview / controlled | `POST /iceberg/v1/{warehouse}/catalog/migration` fences table-bucket registry changes, acquires a persistent per-bucket write fence, records whether a global strong snapshot existed before publication, drains in-flight catalog mutations, materializes the target snapshot, and reports `ready_to_enable_durable_strong`. Retries and `DELETE` may restore a known-absent initial target after an ambiguous first write, but fail closed if a previously existing or materialized global snapshot disappears. `DELETE` releases the bucket fence only while its target state has not advanced, and releases the registry fence after the last bucket is cancelled. Both mutations require `admin:MigrateTableCatalog`. |
|
||||||
|
| Strong snapshot rolling compatibility | Supported | Durable strong control-plane reads snapshot versions 1 and 2, writes version 1 by default, and writes version 2 only after both the requested and fleet-confirmed gates are enabled. A running process rejects any lower-format snapshot after observing a higher format. Once version 2 is fleet-confirmed, table data-plane resolution fails closed until the persisted snapshot is version 2; a missing table-bucket entry also fails closed instead of bypassing table-aware authorization. |
|
||||||
| Disaster recovery rehearsal | Manual/live harness | `failure_coverage.py --print-disaster-recovery-rehearsal` generates an operator runbook covering catalog export, diagnostics, safe recovery repair, rollback/import, durable backing migration dry-run, post-recovery loadTable, and table data-plane policy probes. |
|
| Disaster recovery rehearsal | Manual/live harness | `failure_coverage.py --print-disaster-recovery-rehearsal` generates an operator runbook covering catalog export, diagnostics, safe recovery repair, rollback/import, durable backing migration dry-run, post-recovery loadTable, and table data-plane policy probes. |
|
||||||
| Scale and fault rehearsal | Manual/live harness | `failure_coverage.py --print-scale-fault-rehearsal` generates an opt-in runbook for concurrent writer stress, maintenance scheduler lease recovery, durable backing cutover preflight, recovery/rollback/import under load, and post-run evidence capture. |
|
| Scale and fault rehearsal | Manual/live harness | `failure_coverage.py --print-scale-fault-rehearsal` generates an opt-in runbook for concurrent writer stress, maintenance scheduler lease recovery, durable backing cutover preflight, recovery/rollback/import under load, and post-run evidence capture. |
|
||||||
| Strong KV/WAL backing cutover | Preview / controlled | Operators can select durable strong backing with `RUSTFS_TABLE_CATALOG_BACKING=durable-strong` only after every table bucket reports `SNAPSHOT_MATERIALIZED` and `ready_to_enable_durable_strong: true`. Object-only advanced operations fail closed in durable strong mode. |
|
| Durable strong snapshot backing cutover | Preview / controlled | Operators can select the ETag-CAS snapshot backing with `RUSTFS_TABLE_CATALOG_BACKING=durable-strong` only after every table bucket reports `SNAPSHOT_MATERIALIZED` and `ready_to_enable_durable_strong: true`. This mode does not claim a separate external KV/WAL service, and object-only advanced operations fail closed. Version 1 backing manifests retain the legacy `STRONG_KV_WAL` and `CUT_OVER_LINEARIZABLE_READS` wire labels for client compatibility; those labels do not expand the implementation claim. |
|
||||||
| Single active writer region | Supported policy | Diagnostics publish single-active-writer semantics and read-only replica limits. |
|
| Single active writer region | Supported policy | Diagnostics publish single-active-writer semantics and read-only replica limits. |
|
||||||
| Active-active multi-region writes | Not claimed | A table must not accept independent concurrent writers in multiple active regions. |
|
| Active-active multi-region writes | Not claimed | A table must not accept independent concurrent writers in multiple active regions. |
|
||||||
|
|
||||||
@@ -136,23 +137,60 @@ warehouse:
|
|||||||
`GetTableCatalogAction` on each table bucket. Treat every `blockers` entry as
|
`GetTableCatalogAction` on each table bucket. Treat every `blockers` entry as
|
||||||
fail-closed; repair commit recovery state and backfill the warehouse prefix
|
fail-closed; repair commit recovery state and backfill the warehouse prefix
|
||||||
index before continuing.
|
index before continuing.
|
||||||
3. Run `POST /iceberg/v1/{warehouse}/catalog/migration` with
|
3. Before the migration `POST`, drain every catalog writer that predates the
|
||||||
`admin:MigrateTableCatalog`. This persists the source write fence before it
|
durable-backing migration fence and restart it on a fence-aware release. An
|
||||||
drains in-flight mutations and copies the catalog state.
|
older writer does not recognize the persisted fence and can otherwise
|
||||||
4. Repeat the preflight and materialization for every table bucket. Do not set
|
mutate the object-backed source after the snapshot inventory is captured.
|
||||||
|
Keep all catalog writers on the fence-aware release until cutover completes.
|
||||||
|
4. Inventory object-only advanced operations, including maintenance workers,
|
||||||
|
catalog recovery, export, diagnostics, and external catalog bridge writes.
|
||||||
|
Quiesce mutating operations before cutover and confirm that each required
|
||||||
|
operation is supported by durable-strong mode; unsupported operations fail
|
||||||
|
closed after cutover rather than continuing against object-backed state.
|
||||||
|
5. Run `POST /iceberg/v1/{warehouse}/catalog/migration` with
|
||||||
|
`admin:MigrateTableCatalog`. This acquires the exclusive migration fence to
|
||||||
|
drain in-flight fence-aware mutations, persists the source fence while
|
||||||
|
exclusivity is held, and then copies the catalog state.
|
||||||
|
6. Repeat the preflight and materialization for every table bucket. Do not set
|
||||||
`RUSTFS_TABLE_CATALOG_BACKING=durable-strong` until the preflight reports
|
`RUSTFS_TABLE_CATALOG_BACKING=durable-strong` until the preflight reports
|
||||||
`SNAPSHOT_MATERIALIZED`, no blockers, and
|
`SNAPSHOT_MATERIALIZED`, no blockers, and
|
||||||
`ready_to_enable_durable_strong: true`.
|
`ready_to_enable_durable_strong: true`.
|
||||||
5. Restart with durable strong backing enabled, then verify catalog config,
|
7. Restart with durable strong backing enabled, then verify catalog config,
|
||||||
table and view loads, commit idempotency, and table data-plane policy
|
table and view loads, commit idempotency, and table data-plane policy
|
||||||
resolution before admitting writers.
|
resolution before admitting writers.
|
||||||
6. Before restarting into durable-strong mode, `DELETE` on the migration
|
8. Before restarting into durable-strong mode, `DELETE` on the migration
|
||||||
endpoint can remove a migration-created target bucket snapshot and release
|
endpoint can remove a migration-created target bucket snapshot and release
|
||||||
the source fence. After the durable-strong state advances, cancellation
|
the source fence. After the durable-strong state advances, cancellation
|
||||||
fails closed; recovery requires an operator-selected restore or reverse
|
fails closed; recovery requires an operator-selected restore or reverse
|
||||||
migration instead of restarting against the stale object-backed pointer.
|
migration instead of restarting against the stale object-backed pointer.
|
||||||
7. Preserve the object-backed catalog backup until durable strong backing has
|
9. Preserve the object-backed catalog backup until durable strong backing has
|
||||||
passed the operator's retention window.
|
passed the operator's retention window.
|
||||||
|
10. Keep strong snapshot writes on version 1 during a rolling binary upgrade.
|
||||||
|
After every catalog writer can read version 2, set both
|
||||||
|
`RUSTFS_TABLE_CATALOG_STRONG_SNAPSHOT_V2=true` and
|
||||||
|
`RUSTFS_TABLE_CATALOG_STRONG_SNAPSHOT_V2_FLEET_CONFIRMED=true`, then restart
|
||||||
|
the catalog writers. Perform a controlled catalog write or migration
|
||||||
|
materialization and confirm that the persisted snapshot is version 2 before
|
||||||
|
serving table data-plane traffic. Setting only one gate does not change the
|
||||||
|
write format.
|
||||||
|
11. After any version 2 snapshot is persisted, do not roll catalog writers back
|
||||||
|
to a binary that only reads version 1. Current binaries preserve version 2
|
||||||
|
even when the gates are later disabled. A running process rejects restored
|
||||||
|
version 1 content after observing version 2, but cannot distinguish an older
|
||||||
|
snapshot with the same format version from a deliberate restore. The format
|
||||||
|
high-water mark is process-local: restoring any older snapshot and restarting
|
||||||
|
every writer is a privileged disaster-recovery rollback that cannot be
|
||||||
|
inferred from the restored object alone. Recovery must restore a compatible
|
||||||
|
binary and a snapshot selected through the operator recovery procedure.
|
||||||
|
12. Migration preflight rejects an active table/view identifier collision before
|
||||||
|
it writes a migration fence. A pre-existing version 1 strong snapshot with
|
||||||
|
such a collision is loaded in cleanup-only quarantine. Ambiguous reads fail
|
||||||
|
closed; each cleanup mutation must reduce the collision set, and unrelated
|
||||||
|
writes remain blocked until all collisions are removed. Drain catalog
|
||||||
|
writers that predate cleanup quarantine before starting this repair, and
|
||||||
|
complete cleanup before the first version 2 write. Restoring any version 1
|
||||||
|
snapshot after a writer has observed version 2 fails closed instead of
|
||||||
|
replacing the in-process catalog state.
|
||||||
|
|
||||||
## Production Failure Coverage
|
## Production Failure Coverage
|
||||||
|
|
||||||
|
|||||||
@@ -2067,9 +2067,12 @@ fn job_id_from_params(params: &Params<'_, '_>) -> S3Result<String> {
|
|||||||
fn table_catalog_backend_from_extensions(
|
fn table_catalog_backend_from_extensions(
|
||||||
extensions: &http::Extensions,
|
extensions: &http::Extensions,
|
||||||
) -> S3Result<crate::table_catalog::EcStoreTableCatalogObjectBackend<ECStore>> {
|
) -> S3Result<crate::table_catalog::EcStoreTableCatalogObjectBackend<ECStore>> {
|
||||||
let store = runtime_sources::object_store_from_extensions(extensions)
|
let context = runtime_sources::app_context_from_extensions(extensions)
|
||||||
.ok_or_else(|| table_catalog_internal_error("request object store is not initialized"))?;
|
.ok_or_else(|| table_catalog_internal_error("request application context is not initialized"))?;
|
||||||
Ok(crate::table_catalog::EcStoreTableCatalogObjectBackend::new(store))
|
Ok(crate::table_catalog::EcStoreTableCatalogObjectBackend::new_with_strong_runtime(
|
||||||
|
context.object_store(),
|
||||||
|
context.table_catalog_strong_runtime(),
|
||||||
|
))
|
||||||
}
|
}
|
||||||
|
|
||||||
type EcStoreObjectTableCatalogStore =
|
type EcStoreObjectTableCatalogStore =
|
||||||
|
|||||||
@@ -2512,7 +2512,7 @@ async fn commit_publication_replays_historical_standard_commit_across_backings()
|
|||||||
crate::table_catalog::TableCatalogBackingMode::DurableStrong,
|
crate::table_catalog::TableCatalogBackingMode::DurableStrong,
|
||||||
] {
|
] {
|
||||||
let metadata_backend = TestTableCatalogObjectBackend::default();
|
let metadata_backend = TestTableCatalogObjectBackend::default();
|
||||||
let store = crate::table_catalog::ConfiguredTableCatalogStore::new(metadata_backend.clone(), mode);
|
let store = crate::table_catalog::ConfiguredTableCatalogStore::new_for_test(metadata_backend.clone(), mode);
|
||||||
let namespace = crate::table_catalog::Namespace::parse("analytics").expect("namespace should parse");
|
let namespace = crate::table_catalog::Namespace::parse("analytics").expect("namespace should parse");
|
||||||
create_standard_events_table(&store, &metadata_backend, &namespace).await;
|
create_standard_events_table(&store, &metadata_backend, &namespace).await;
|
||||||
let first_request = serde_json::json!({
|
let first_request = serde_json::json!({
|
||||||
@@ -2559,7 +2559,7 @@ async fn commit_publication_replays_historical_standard_commit_across_backings()
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
let store = crate::table_catalog::ConfiguredTableCatalogStore::new(metadata_backend.clone(), mode);
|
let store = crate::table_catalog::ConfiguredTableCatalogStore::new_for_test(metadata_backend.clone(), mode);
|
||||||
let second = standard_commit_table_response(
|
let second = standard_commit_table_response(
|
||||||
&store,
|
&store,
|
||||||
&trusted_table_commit_backend(&metadata_backend),
|
&trusted_table_commit_backend(&metadata_backend),
|
||||||
|
|||||||
@@ -76,6 +76,7 @@ pub struct AppContext {
|
|||||||
buffer_config: Arc<dyn BufferConfigInterface>,
|
buffer_config: Arc<dyn BufferConfigInterface>,
|
||||||
object_data_cache: Arc<ObjectDataCacheAdapter>,
|
object_data_cache: Arc<ObjectDataCacheAdapter>,
|
||||||
object_traffic_health: Arc<ObjectTrafficHealth>,
|
object_traffic_health: Arc<ObjectTrafficHealth>,
|
||||||
|
table_catalog_strong_runtime: crate::table_catalog::StrongTableCatalogRuntime,
|
||||||
}
|
}
|
||||||
|
|
||||||
impl AppContext {
|
impl AppContext {
|
||||||
@@ -125,6 +126,7 @@ impl AppContext {
|
|||||||
buffer_config: default_buffer_config_interface(),
|
buffer_config: default_buffer_config_interface(),
|
||||||
object_data_cache,
|
object_data_cache,
|
||||||
object_traffic_health: Arc::new(ObjectTrafficHealth::from_env()),
|
object_traffic_health: Arc::new(ObjectTrafficHealth::from_env()),
|
||||||
|
table_catalog_strong_runtime: crate::table_catalog::StrongTableCatalogRuntime::default(),
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -144,6 +146,10 @@ impl AppContext {
|
|||||||
Arc::clone(&self.object_traffic_health)
|
Arc::clone(&self.object_traffic_health)
|
||||||
}
|
}
|
||||||
|
|
||||||
|
pub(crate) fn table_catalog_strong_runtime(&self) -> crate::table_catalog::StrongTableCatalogRuntime {
|
||||||
|
self.table_catalog_strong_runtime.clone()
|
||||||
|
}
|
||||||
|
|
||||||
pub fn iam(&self) -> Arc<dyn IamInterface> {
|
pub fn iam(&self) -> Arc<dyn IamInterface> {
|
||||||
self.iam.clone()
|
self.iam.clone()
|
||||||
}
|
}
|
||||||
@@ -350,6 +356,7 @@ impl AppContext {
|
|||||||
buffer_config: interfaces.buffer_config,
|
buffer_config: interfaces.buffer_config,
|
||||||
object_data_cache: ObjectDataCacheAdapter::disabled_arc(),
|
object_data_cache: ObjectDataCacheAdapter::disabled_arc(),
|
||||||
object_traffic_health: Arc::new(ObjectTrafficHealth::from_env()),
|
object_traffic_health: Arc::new(ObjectTrafficHealth::from_env()),
|
||||||
|
table_catalog_strong_runtime: crate::table_catalog::StrongTableCatalogRuntime::default(),
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|||||||
@@ -1438,8 +1438,15 @@ fn table_data_plane_content_mutation(action: Action) -> bool {
|
|||||||
fn table_catalog_backend_for_data_plane<T>(
|
fn table_catalog_backend_for_data_plane<T>(
|
||||||
req: &S3Request<T>,
|
req: &S3Request<T>,
|
||||||
) -> S3Result<crate::table_catalog::EcStoreTableCatalogObjectBackend<ECStore>> {
|
) -> S3Result<crate::table_catalog::EcStoreTableCatalogObjectBackend<ECStore>> {
|
||||||
let store = request_object_store(req)?;
|
let context = match req.extensions.get::<Arc<ServerContextSlot>>() {
|
||||||
Ok(crate::table_catalog::EcStoreTableCatalogObjectBackend::new(store))
|
Some(server_ctx) => server_ctx.installed_app_context(),
|
||||||
|
None => runtime_sources::current_app_context(),
|
||||||
|
}
|
||||||
|
.ok_or_else(object_store_not_initialized_error)?;
|
||||||
|
Ok(crate::table_catalog::EcStoreTableCatalogObjectBackend::new_with_strong_runtime(
|
||||||
|
context.object_store(),
|
||||||
|
context.table_catalog_strong_runtime(),
|
||||||
|
))
|
||||||
}
|
}
|
||||||
|
|
||||||
fn table_catalog_store_for_data_plane<T>(
|
fn table_catalog_store_for_data_plane<T>(
|
||||||
|
|||||||
@@ -44,11 +44,14 @@ pub(crate) fn table_matches_staged_base(table: &TableEntry, commit_log: &CommitL
|
|||||||
pub(crate) struct TableCommitHistoryIndex<'a> {
|
pub(crate) struct TableCommitHistoryIndex<'a> {
|
||||||
table_id: &'a str,
|
table_id: &'a str,
|
||||||
reachable_states: BTreeSet<(&'a str, &'a str)>,
|
reachable_states: BTreeSet<(&'a str, &'a str)>,
|
||||||
|
ambiguous_states: BTreeSet<(&'a str, &'a str)>,
|
||||||
|
cycle_detected: bool,
|
||||||
}
|
}
|
||||||
|
|
||||||
impl<'a> TableCommitHistoryIndex<'a> {
|
impl<'a> TableCommitHistoryIndex<'a> {
|
||||||
pub(crate) fn new(table: &'a TableEntry, commits: impl IntoIterator<Item = &'a CommitLogEntry>) -> Self {
|
pub(crate) fn new(table: &'a TableEntry, commits: impl IntoIterator<Item = &'a CommitLogEntry>) -> Self {
|
||||||
let mut by_new_state = BTreeMap::<(&str, &str), Option<(&str, &str)>>::new();
|
let mut by_new_state = BTreeMap::<(&str, &str), Option<(&str, &str)>>::new();
|
||||||
|
let mut ambiguous_states = BTreeSet::new();
|
||||||
for commit in commits
|
for commit in commits
|
||||||
.into_iter()
|
.into_iter()
|
||||||
.filter(|commit| commit.table_id == table.table_id && !matches!(commit.status, CommitLogStatus::Failed))
|
.filter(|commit| commit.table_id == table.table_id && !matches!(commit.status, CommitLogStatus::Failed))
|
||||||
@@ -57,27 +60,39 @@ impl<'a> TableCommitHistoryIndex<'a> {
|
|||||||
let previous = (commit.previous_metadata_location.as_str(), commit.expected_version_token.as_str());
|
let previous = (commit.previous_metadata_location.as_str(), commit.expected_version_token.as_str());
|
||||||
by_new_state
|
by_new_state
|
||||||
.entry(key)
|
.entry(key)
|
||||||
.and_modify(|candidate| *candidate = None)
|
.and_modify(|candidate| {
|
||||||
|
*candidate = None;
|
||||||
|
ambiguous_states.insert(key);
|
||||||
|
})
|
||||||
.or_insert(Some(previous));
|
.or_insert(Some(previous));
|
||||||
}
|
}
|
||||||
|
|
||||||
let mut reachable_states = BTreeSet::new();
|
let mut reachable_states = BTreeSet::new();
|
||||||
let mut state = (table.metadata_location.as_str(), table.version_token.as_str());
|
let mut state = (table.metadata_location.as_str(), table.version_token.as_str());
|
||||||
while reachable_states.insert(state) {
|
let cycle_detected = loop {
|
||||||
|
if !reachable_states.insert(state) {
|
||||||
|
break true;
|
||||||
|
}
|
||||||
let Some(Some(previous)) = by_new_state.get(&state) else {
|
let Some(Some(previous)) = by_new_state.get(&state) else {
|
||||||
break;
|
break false;
|
||||||
};
|
};
|
||||||
state = *previous;
|
state = *previous;
|
||||||
}
|
};
|
||||||
Self {
|
Self {
|
||||||
table_id: &table.table_id,
|
table_id: &table.table_id,
|
||||||
reachable_states,
|
reachable_states,
|
||||||
|
ambiguous_states,
|
||||||
|
cycle_detected,
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
pub(crate) fn proves_committed(&self, target: &CommitLogEntry) -> bool {
|
pub(crate) fn proves_committed(&self, target: &CommitLogEntry) -> bool {
|
||||||
self.table_id == target.table_id.as_str()
|
!self.cycle_detected
|
||||||
|
&& self.table_id == target.table_id.as_str()
|
||||||
&& !matches!(target.status, CommitLogStatus::Failed)
|
&& !matches!(target.status, CommitLogStatus::Failed)
|
||||||
|
&& !self
|
||||||
|
.ambiguous_states
|
||||||
|
.contains(&(target.new_metadata_location.as_str(), target.new_version_token.as_str()))
|
||||||
&& self
|
&& self
|
||||||
.reachable_states
|
.reachable_states
|
||||||
.contains(&(target.new_metadata_location.as_str(), target.new_version_token.as_str()))
|
.contains(&(target.new_metadata_location.as_str(), target.new_version_token.as_str()))
|
||||||
@@ -202,7 +217,7 @@ pub(crate) fn table_commit_recovery_entry(
|
|||||||
TableCommitRecoveryState::FinalizationRequired,
|
TableCommitRecoveryState::FinalizationRequired,
|
||||||
"a later committed pointer proves this staged commit is part of table history".to_string(),
|
"a later committed pointer proves this staged commit is part of table history".to_string(),
|
||||||
)
|
)
|
||||||
} else if matches!(commit_log.status, CommitLogStatus::Committed) {
|
} else if matches!(commit_log.status, CommitLogStatus::Committed) && historically_committed {
|
||||||
if idempotency_index_repair_required {
|
if idempotency_index_repair_required {
|
||||||
(
|
(
|
||||||
TableCommitRecoveryState::IdempotencyIndexRepairRequired,
|
TableCommitRecoveryState::IdempotencyIndexRepairRequired,
|
||||||
@@ -214,6 +229,11 @@ pub(crate) fn table_commit_recovery_entry(
|
|||||||
"commit is finalized and may be older than the current table pointer".to_string(),
|
"commit is finalized and may be older than the current table pointer".to_string(),
|
||||||
)
|
)
|
||||||
}
|
}
|
||||||
|
} else if matches!(commit_log.status, CommitLogStatus::Committed) {
|
||||||
|
(
|
||||||
|
TableCommitRecoveryState::ManualReview,
|
||||||
|
"committed log is not reachable from the current table pointer".to_string(),
|
||||||
|
)
|
||||||
} else if table_matches_staged_base(table, commit_log) {
|
} else if table_matches_staged_base(table, commit_log) {
|
||||||
(
|
(
|
||||||
TableCommitRecoveryState::StagedBeforeTableUpdate,
|
TableCommitRecoveryState::StagedBeforeTableUpdate,
|
||||||
|
|||||||
@@ -286,17 +286,15 @@ where
|
|||||||
if table.state != TableCatalogEntryState::Active {
|
if table.state != TableCatalogEntryState::Active {
|
||||||
continue;
|
continue;
|
||||||
}
|
}
|
||||||
let Ok(warehouse_object_prefix) = table_warehouse_object_prefix(&table) else {
|
let warehouse_object_prefix = table_warehouse_object_prefix(&table)?;
|
||||||
continue;
|
|
||||||
};
|
|
||||||
if !object.starts_with(&warehouse_object_prefix) {
|
if !object.starts_with(&warehouse_object_prefix) {
|
||||||
continue;
|
continue;
|
||||||
}
|
}
|
||||||
if matched
|
if let Some(current) = matched.as_ref() {
|
||||||
.as_ref()
|
return Err(TableCatalogStoreError::Invalid(format!(
|
||||||
.is_some_and(|current| current.warehouse_object_prefix.len() >= warehouse_object_prefix.len())
|
"object {object} matches overlapping active table warehouse prefixes {} and {warehouse_object_prefix}",
|
||||||
{
|
current.warehouse_object_prefix
|
||||||
continue;
|
)));
|
||||||
}
|
}
|
||||||
matched = Some(table_data_plane_resource_from_entry(table, warehouse_object_prefix));
|
matched = Some(table_data_plane_resource_from_entry(table, warehouse_object_prefix));
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -97,6 +97,9 @@ pub(crate) const TABLE_CATALOG_BACKING_MANIFEST_VERSION: u16 = 1;
|
|||||||
pub(crate) const ENV_TABLE_CATALOG_BACKING: &str = "RUSTFS_TABLE_CATALOG_BACKING";
|
pub(crate) const ENV_TABLE_CATALOG_BACKING: &str = "RUSTFS_TABLE_CATALOG_BACKING";
|
||||||
pub(crate) const ENV_TABLE_CATALOG_PUBLICATION_FENCE_FLEET_CONFIRMED: &str =
|
pub(crate) const ENV_TABLE_CATALOG_PUBLICATION_FENCE_FLEET_CONFIRMED: &str =
|
||||||
"RUSTFS_TABLE_CATALOG_PUBLICATION_FENCE_FLEET_CONFIRMED";
|
"RUSTFS_TABLE_CATALOG_PUBLICATION_FENCE_FLEET_CONFIRMED";
|
||||||
|
pub(crate) const ENV_TABLE_CATALOG_STRONG_SNAPSHOT_V2: &str = "RUSTFS_TABLE_CATALOG_STRONG_SNAPSHOT_V2";
|
||||||
|
pub(crate) const ENV_TABLE_CATALOG_STRONG_SNAPSHOT_V2_FLEET_CONFIRMED: &str =
|
||||||
|
"RUSTFS_TABLE_CATALOG_STRONG_SNAPSHOT_V2_FLEET_CONFIRMED";
|
||||||
pub(crate) const TABLE_CATALOG_BACKING_OBJECT: &str = "object";
|
pub(crate) const TABLE_CATALOG_BACKING_OBJECT: &str = "object";
|
||||||
pub(crate) const TABLE_CATALOG_BACKING_DURABLE_STRONG: &str = "durable-strong";
|
pub(crate) const TABLE_CATALOG_BACKING_DURABLE_STRONG: &str = "durable-strong";
|
||||||
pub(crate) const TABLE_METADATA_DIGEST_REQUIREMENT_TYPE: &str = "assert-rustfs-metadata-sha256";
|
pub(crate) const TABLE_METADATA_DIGEST_REQUIREMENT_TYPE: &str = "assert-rustfs-metadata-sha256";
|
||||||
@@ -159,10 +162,12 @@ const ICEBERG_MAX_REF_AGE_MS_PROPERTY: &str = "history.expire.max-ref-age-ms";
|
|||||||
const ICEBERG_REF_MIN_SNAPSHOTS_TO_KEEP_FIELD: &str = "min-snapshots-to-keep";
|
const ICEBERG_REF_MIN_SNAPSHOTS_TO_KEEP_FIELD: &str = "min-snapshots-to-keep";
|
||||||
const ICEBERG_REF_MAX_SNAPSHOT_AGE_MS_FIELD: &str = "max-snapshot-age-ms";
|
const ICEBERG_REF_MAX_SNAPSHOT_AGE_MS_FIELD: &str = "max-snapshot-age-ms";
|
||||||
const ICEBERG_REF_MAX_REF_AGE_MS_FIELD: &str = "max-ref-age-ms";
|
const ICEBERG_REF_MAX_REF_AGE_MS_FIELD: &str = "max-ref-age-ms";
|
||||||
const STRONG_TABLE_CATALOG_SNAPSHOT_VERSION: u16 = 1;
|
const STRONG_TABLE_CATALOG_SNAPSHOT_MIN_READ_VERSION: u16 = 1;
|
||||||
|
const STRONG_TABLE_CATALOG_SNAPSHOT_VERSION: u16 = 2;
|
||||||
const STRONG_TABLE_CATALOG_BACKING_ROOT: &str = "strong-backing";
|
const STRONG_TABLE_CATALOG_BACKING_ROOT: &str = "strong-backing";
|
||||||
const STRONG_TABLE_CATALOG_SNAPSHOT_FILE: &str = "snapshot.json";
|
const STRONG_TABLE_CATALOG_SNAPSHOT_FILE: &str = "snapshot.json";
|
||||||
const TABLE_CATALOG_MIGRATION_VERSION: u16 = 1;
|
const TABLE_CATALOG_MIGRATION_MIN_READ_VERSION: u16 = 1;
|
||||||
|
const TABLE_CATALOG_MIGRATION_VERSION: u16 = 2;
|
||||||
const TABLE_CATALOG_MIGRATION_ROOT: &str = "backing-migration";
|
const TABLE_CATALOG_MIGRATION_ROOT: &str = "backing-migration";
|
||||||
const TABLE_CATALOG_MIGRATION_FENCE_FILE: &str = "durable-strong-fence.json";
|
const TABLE_CATALOG_MIGRATION_FENCE_FILE: &str = "durable-strong-fence.json";
|
||||||
const TABLE_CATALOG_MIGRATION_FENCE_LOCK: &str = "durable-strong-fence.lock";
|
const TABLE_CATALOG_MIGRATION_FENCE_LOCK: &str = "durable-strong-fence.lock";
|
||||||
|
|||||||
@@ -1067,7 +1067,9 @@ pub(crate) struct TableCatalogBackingProfile {
|
|||||||
#[serde(rename_all = "SCREAMING_SNAKE_CASE")]
|
#[serde(rename_all = "SCREAMING_SNAKE_CASE")]
|
||||||
pub(crate) enum TableCatalogBackingKind {
|
pub(crate) enum TableCatalogBackingKind {
|
||||||
ObjectBacked,
|
ObjectBacked,
|
||||||
StrongKvWal,
|
// RUSTFS_COMPAT_TODO(table-catalog-backing-manifest-v1-wire-labels): Keep the version 1 wire label for existing clients. Remove after a versioned manifest with an explicit client migration contract replaces it.
|
||||||
|
#[serde(rename = "STRONG_KV_WAL")]
|
||||||
|
DurableStrongSnapshot,
|
||||||
}
|
}
|
||||||
|
|
||||||
#[derive(Debug, Clone, PartialEq, Eq, Serialize)]
|
#[derive(Debug, Clone, PartialEq, Eq, Serialize)]
|
||||||
@@ -1203,7 +1205,9 @@ pub(crate) enum TableCatalogBackingMigrationStep {
|
|||||||
ReplayCommitLog,
|
ReplayCommitLog,
|
||||||
VerifyCurrentPointer,
|
VerifyCurrentPointer,
|
||||||
EnableSingleWriterFencing,
|
EnableSingleWriterFencing,
|
||||||
CutOverLinearizableReads,
|
// RUSTFS_COMPAT_TODO(table-catalog-backing-manifest-v1-wire-labels): Keep the version 1 wire label for existing clients. Remove after a versioned manifest with an explicit client migration contract replaces it.
|
||||||
|
#[serde(rename = "CUT_OVER_LINEARIZABLE_READS")]
|
||||||
|
CutOverDurableSnapshotReads,
|
||||||
}
|
}
|
||||||
|
|
||||||
#[derive(Debug, Clone, PartialEq, Eq, Serialize)]
|
#[derive(Debug, Clone, PartialEq, Eq, Serialize)]
|
||||||
@@ -1213,6 +1217,8 @@ pub(crate) enum TableCatalogBackingMigrationBlocker {
|
|||||||
CommitManualReviewRequired,
|
CommitManualReviewRequired,
|
||||||
WarehouseIndexBackfillRequired,
|
WarehouseIndexBackfillRequired,
|
||||||
DuplicateWarehousePrefix,
|
DuplicateWarehousePrefix,
|
||||||
|
DuplicateTableIdentity,
|
||||||
|
TableViewIdentifierCollision,
|
||||||
DurableStrongSnapshotChanged,
|
DurableStrongSnapshotChanged,
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -1222,6 +1228,8 @@ pub(crate) enum TableCatalogBackingMigrationAction {
|
|||||||
RunCatalogRecovery,
|
RunCatalogRecovery,
|
||||||
BackfillWarehouseIndex,
|
BackfillWarehouseIndex,
|
||||||
ReviewDuplicateWarehousePrefixes,
|
ReviewDuplicateWarehousePrefixes,
|
||||||
|
ReviewDuplicateTableIdentities,
|
||||||
|
ReviewTableViewIdentifierCollisions,
|
||||||
SnapshotObjectBackedCatalog,
|
SnapshotObjectBackedCatalog,
|
||||||
EnableDurableStrongBacking,
|
EnableDurableStrongBacking,
|
||||||
VerifyDurableStrongSnapshot,
|
VerifyDurableStrongSnapshot,
|
||||||
|
|||||||
@@ -13,11 +13,13 @@
|
|||||||
// limitations under the License.
|
// limitations under the License.
|
||||||
|
|
||||||
use super::object::{
|
use super::object::{
|
||||||
ObjectTableCatalogStore, validate_namespace_entry_object, validate_table_entry_object, validate_view_entry_object,
|
ObjectTableCatalogStore, validate_commit_idempotency_entry_object, validate_commit_log_entry_object,
|
||||||
|
validate_namespace_entry_object, validate_table_bucket_entry_object, validate_table_entry_object, validate_view_entry_object,
|
||||||
};
|
};
|
||||||
use super::strong::{
|
use super::strong::{
|
||||||
StrongCommitSnapshotRecord, StrongTableCatalogBucketSnapshot, StrongTableCatalogState, TableCatalogBackingMigrationFence,
|
StrongCommitSnapshotRecord, StrongTableCatalogBucketSnapshot, StrongTableCatalogState, TableCatalogBackingMigrationFence,
|
||||||
TableCatalogBackingMigrationFenceStatus, TableCatalogBackingMigrationGlobalFence, table_catalog_bucket_snapshot_fingerprint,
|
TableCatalogBackingMigrationFenceStatus, TableCatalogBackingMigrationGlobalFence,
|
||||||
|
TableCatalogBackingMigrationTargetSnapshotState, table_catalog_bucket_snapshot_fingerprint,
|
||||||
};
|
};
|
||||||
use super::*;
|
use super::*;
|
||||||
|
|
||||||
@@ -82,14 +84,14 @@ pub(super) fn table_catalog_backing_manifest(
|
|||||||
},
|
},
|
||||||
migration: TableCatalogBackingMigrationPlan {
|
migration: TableCatalogBackingMigrationPlan {
|
||||||
source_kind: TableCatalogBackingKind::ObjectBacked,
|
source_kind: TableCatalogBackingKind::ObjectBacked,
|
||||||
target_kind: TableCatalogBackingKind::StrongKvWal,
|
target_kind: TableCatalogBackingKind::DurableStrongSnapshot,
|
||||||
status: migration_status,
|
status: migration_status,
|
||||||
required_steps: vec![
|
required_steps: vec![
|
||||||
TableCatalogBackingMigrationStep::SnapshotCatalogExport,
|
TableCatalogBackingMigrationStep::SnapshotCatalogExport,
|
||||||
TableCatalogBackingMigrationStep::ReplayCommitLog,
|
TableCatalogBackingMigrationStep::ReplayCommitLog,
|
||||||
TableCatalogBackingMigrationStep::VerifyCurrentPointer,
|
TableCatalogBackingMigrationStep::VerifyCurrentPointer,
|
||||||
TableCatalogBackingMigrationStep::EnableSingleWriterFencing,
|
TableCatalogBackingMigrationStep::EnableSingleWriterFencing,
|
||||||
TableCatalogBackingMigrationStep::CutOverLinearizableReads,
|
TableCatalogBackingMigrationStep::CutOverDurableSnapshotReads,
|
||||||
],
|
],
|
||||||
blockers,
|
blockers,
|
||||||
},
|
},
|
||||||
@@ -118,12 +120,100 @@ impl<B> ObjectTableCatalogStore<B>
|
|||||||
where
|
where
|
||||||
B: TableCatalogObjectBackend,
|
B: TableCatalogObjectBackend,
|
||||||
{
|
{
|
||||||
|
fn migration_target_snapshot_state(
|
||||||
|
fence: &TableCatalogBackingMigrationFence,
|
||||||
|
) -> TableCatalogBackingMigrationTargetSnapshotState {
|
||||||
|
// RUSTFS_COMPAT_TODO(table-catalog-migration-fence-v1): Version 1 PREPARING fences have no durable baseline. Remove after all supported upgrade sources write version 2 fences and all version 1 migrations are completed or cancelled.
|
||||||
|
if fence.version < TABLE_CATALOG_MIGRATION_VERSION {
|
||||||
|
TableCatalogBackingMigrationTargetSnapshotState::Unknown
|
||||||
|
} else if fence.target_snapshot_etag.is_some() {
|
||||||
|
TableCatalogBackingMigrationTargetSnapshotState::Present
|
||||||
|
} else {
|
||||||
|
TableCatalogBackingMigrationTargetSnapshotState::Absent
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
fn validate_backing_migration_fence(
|
||||||
|
table_bucket: &str,
|
||||||
|
fence: &TableCatalogBackingMigrationFence,
|
||||||
|
) -> TableCatalogStoreResult<()> {
|
||||||
|
if !(TABLE_CATALOG_MIGRATION_MIN_READ_VERSION..=TABLE_CATALOG_MIGRATION_VERSION).contains(&fence.version)
|
||||||
|
|| fence.table_bucket != table_bucket
|
||||||
|
|| fence.migration_id.is_empty()
|
||||||
|
{
|
||||||
|
return Err(TableCatalogStoreError::Invalid(format!(
|
||||||
|
"invalid durable strong migration fence for table bucket {table_bucket}"
|
||||||
|
)));
|
||||||
|
}
|
||||||
|
let target_snapshot_state = Self::migration_target_snapshot_state(fence);
|
||||||
|
if fence.target_bucket_existed && target_snapshot_state == TableCatalogBackingMigrationTargetSnapshotState::Absent {
|
||||||
|
return Err(TableCatalogStoreError::Invalid(format!(
|
||||||
|
"durable strong migration fence for table bucket {table_bucket} has an inconsistent target snapshot baseline"
|
||||||
|
)));
|
||||||
|
}
|
||||||
|
match fence.status {
|
||||||
|
TableCatalogBackingMigrationFenceStatus::Preparing if fence.source_fingerprint.is_some() => {
|
||||||
|
return Err(TableCatalogStoreError::Invalid(format!(
|
||||||
|
"preparing durable strong migration fence for table bucket {table_bucket} has materialized state"
|
||||||
|
)));
|
||||||
|
}
|
||||||
|
TableCatalogBackingMigrationFenceStatus::Materialized
|
||||||
|
if fence.source_fingerprint.is_none() || fence.target_snapshot_etag.is_none() =>
|
||||||
|
{
|
||||||
|
return Err(TableCatalogStoreError::Invalid(format!(
|
||||||
|
"materialized durable strong migration fence for table bucket {table_bucket} is incomplete"
|
||||||
|
)));
|
||||||
|
}
|
||||||
|
_ => {}
|
||||||
|
}
|
||||||
|
Ok(())
|
||||||
|
}
|
||||||
|
|
||||||
|
fn validate_global_backing_migration_fence(fence: &TableCatalogBackingMigrationGlobalFence) -> TableCatalogStoreResult<()> {
|
||||||
|
if !(TABLE_CATALOG_MIGRATION_MIN_READ_VERSION..=TABLE_CATALOG_MIGRATION_VERSION).contains(&fence.version)
|
||||||
|
|| fence.migration_id.is_empty()
|
||||||
|
{
|
||||||
|
return Err(TableCatalogStoreError::Invalid(
|
||||||
|
"invalid durable strong global migration fence".to_string(),
|
||||||
|
));
|
||||||
|
}
|
||||||
|
Ok(())
|
||||||
|
}
|
||||||
|
|
||||||
|
async fn observe_durable_strong_migration_target(
|
||||||
|
strong_store: &StrongTableCatalogStore<B>,
|
||||||
|
table_bucket: &str,
|
||||||
|
fence: Option<&TableCatalogBackingMigrationFence>,
|
||||||
|
) -> TableCatalogStoreResult<(Option<String>, Option<String>)> {
|
||||||
|
let target_snapshot_state = fence.map(Self::migration_target_snapshot_state);
|
||||||
|
let permits_absent_snapshot = fence.is_some_and(|fence| {
|
||||||
|
fence.status == TableCatalogBackingMigrationFenceStatus::Preparing
|
||||||
|
&& !fence.target_bucket_existed
|
||||||
|
&& target_snapshot_state == Some(TableCatalogBackingMigrationTargetSnapshotState::Absent)
|
||||||
|
});
|
||||||
|
if permits_absent_snapshot {
|
||||||
|
strong_store.restore_absent_migration_snapshot_baseline().await?;
|
||||||
|
}
|
||||||
|
let observation = strong_store.bucket_snapshot_observation(table_bucket).await?;
|
||||||
|
if fence.is_some() && observation.1.is_none() && !permits_absent_snapshot {
|
||||||
|
return Err(TableCatalogStoreError::Conflict(
|
||||||
|
"durable strong catalog snapshot is missing for an in-progress backing migration".to_string(),
|
||||||
|
));
|
||||||
|
}
|
||||||
|
Ok(observation)
|
||||||
|
}
|
||||||
|
|
||||||
async fn read_backing_migration_fence(
|
async fn read_backing_migration_fence(
|
||||||
&self,
|
&self,
|
||||||
table_bucket: &str,
|
table_bucket: &str,
|
||||||
) -> TableCatalogStoreResult<Option<(TableCatalogBackingMigrationFence, Option<String>)>> {
|
) -> TableCatalogStoreResult<Option<(TableCatalogBackingMigrationFence, Option<String>)>> {
|
||||||
self.read_entry(self.catalog_bucket(), &self.paths.backing_migration_fence_path(table_bucket))
|
let fence = self
|
||||||
.await
|
.read_entry(self.catalog_bucket(), &self.paths.backing_migration_fence_path(table_bucket))
|
||||||
|
.await?;
|
||||||
|
if let Some((fence, _)) = fence.as_ref() {
|
||||||
|
Self::validate_backing_migration_fence(table_bucket, fence)?;
|
||||||
|
}
|
||||||
|
Ok(fence)
|
||||||
}
|
}
|
||||||
|
|
||||||
pub(super) async fn acquire_table_bucket_registry_write_permit(&self) -> TableCatalogStoreResult<Box<dyn Send>> {
|
pub(super) async fn acquire_table_bucket_registry_write_permit(&self) -> TableCatalogStoreResult<Box<dyn Send>> {
|
||||||
@@ -164,11 +254,7 @@ where
|
|||||||
.read_entry::<TableCatalogBackingMigrationGlobalFence>(self.catalog_bucket(), fence_path)
|
.read_entry::<TableCatalogBackingMigrationGlobalFence>(self.catalog_bucket(), fence_path)
|
||||||
.await?
|
.await?
|
||||||
{
|
{
|
||||||
if fence.version != TABLE_CATALOG_MIGRATION_VERSION {
|
Self::validate_global_backing_migration_fence(&fence)?;
|
||||||
return Err(TableCatalogStoreError::Invalid(
|
|
||||||
"invalid durable strong global migration fence".to_string(),
|
|
||||||
));
|
|
||||||
}
|
|
||||||
return Ok(fence);
|
return Ok(fence);
|
||||||
}
|
}
|
||||||
let fence = TableCatalogBackingMigrationGlobalFence {
|
let fence = TableCatalogBackingMigrationGlobalFence {
|
||||||
@@ -181,6 +267,13 @@ where
|
|||||||
}
|
}
|
||||||
|
|
||||||
async fn clear_global_backing_migration_fence_if_unused(&self, fence_path: &str) -> TableCatalogStoreResult<()> {
|
async fn clear_global_backing_migration_fence_if_unused(&self, fence_path: &str) -> TableCatalogStoreResult<()> {
|
||||||
|
let Some((fence, _)) = self
|
||||||
|
.read_entry::<TableCatalogBackingMigrationGlobalFence>(self.catalog_bucket(), fence_path)
|
||||||
|
.await?
|
||||||
|
else {
|
||||||
|
return Ok(());
|
||||||
|
};
|
||||||
|
Self::validate_global_backing_migration_fence(&fence)?;
|
||||||
let bucket_objects = self
|
let bucket_objects = self
|
||||||
.backend
|
.backend
|
||||||
.list_objects(self.catalog_bucket(), &self.paths.table_bucket_entries_prefix())
|
.list_objects(self.catalog_bucket(), &self.paths.table_bucket_entries_prefix())
|
||||||
@@ -194,19 +287,6 @@ where
|
|||||||
self.backend.delete_object(self.catalog_bucket(), fence_path).await
|
self.backend.delete_object(self.catalog_bucket(), fence_path).await
|
||||||
}
|
}
|
||||||
|
|
||||||
pub(super) async fn ensure_object_backed_writes_allowed(&self, table_bucket: &str) -> TableCatalogStoreResult<()> {
|
|
||||||
if self
|
|
||||||
.backend
|
|
||||||
.object_exists(self.catalog_bucket(), &self.paths.backing_migration_fence_path(table_bucket))
|
|
||||||
.await?
|
|
||||||
{
|
|
||||||
return Err(TableCatalogStoreError::Conflict(format!(
|
|
||||||
"object-backed catalog writes are fenced while table bucket {table_bucket} is prepared for durable strong cutover"
|
|
||||||
)));
|
|
||||||
}
|
|
||||||
Ok(())
|
|
||||||
}
|
|
||||||
|
|
||||||
async fn collect_bucket_snapshot_with_locks(
|
async fn collect_bucket_snapshot_with_locks(
|
||||||
&self,
|
&self,
|
||||||
table_bucket: &str,
|
table_bucket: &str,
|
||||||
@@ -220,11 +300,7 @@ where
|
|||||||
else {
|
else {
|
||||||
return Err(TableCatalogStoreError::NotFound(format!("table bucket {table_bucket}")));
|
return Err(TableCatalogStoreError::NotFound(format!("table bucket {table_bucket}")));
|
||||||
};
|
};
|
||||||
if table_bucket_entry.table_bucket != table_bucket {
|
validate_table_bucket_entry_object(&self.paths, &bucket_path, &table_bucket_entry)?;
|
||||||
return Err(TableCatalogStoreError::Invalid(format!(
|
|
||||||
"table bucket entry does not match migration target {table_bucket}"
|
|
||||||
)));
|
|
||||||
}
|
|
||||||
|
|
||||||
let mut namespaces = Vec::new();
|
let mut namespaces = Vec::new();
|
||||||
let mut tables = Vec::new();
|
let mut tables = Vec::new();
|
||||||
@@ -286,6 +362,7 @@ where
|
|||||||
"commit log changed while preparing durable strong snapshot: {commit_object}"
|
"commit log changed while preparing durable strong snapshot: {commit_object}"
|
||||||
)));
|
)));
|
||||||
};
|
};
|
||||||
|
validate_commit_log_entry_object(&self.paths, &commit_object, table_bucket, &table_entry.table_id, &commit)?;
|
||||||
commits.push(StrongCommitSnapshotRecord {
|
commits.push(StrongCommitSnapshotRecord {
|
||||||
table_bucket: table_bucket.to_string(),
|
table_bucket: table_bucket.to_string(),
|
||||||
table_id: table_entry.table_id.clone(),
|
table_id: table_entry.table_id.clone(),
|
||||||
@@ -313,6 +390,13 @@ where
|
|||||||
"idempotency index changed while preparing durable strong snapshot: {idempotency_object}"
|
"idempotency index changed while preparing durable strong snapshot: {idempotency_object}"
|
||||||
)));
|
)));
|
||||||
};
|
};
|
||||||
|
validate_commit_idempotency_entry_object(
|
||||||
|
&self.paths,
|
||||||
|
&idempotency_object,
|
||||||
|
table_bucket,
|
||||||
|
&table_entry.table_id,
|
||||||
|
&commit,
|
||||||
|
)?;
|
||||||
let lookup_key = commit.idempotency_key.clone().ok_or_else(|| {
|
let lookup_key = commit.idempotency_key.clone().ok_or_else(|| {
|
||||||
TableCatalogStoreError::Invalid(format!("idempotency index {idempotency_object} has no idempotency key"))
|
TableCatalogStoreError::Invalid(format!("idempotency index {idempotency_object} has no idempotency key"))
|
||||||
})?;
|
})?;
|
||||||
@@ -360,6 +444,22 @@ where
|
|||||||
|
|
||||||
fn validate_bucket_snapshot_for_migration(&self, snapshot: &StrongTableCatalogBucketSnapshot) -> TableCatalogStoreResult<()> {
|
fn validate_bucket_snapshot_for_migration(&self, snapshot: &StrongTableCatalogBucketSnapshot) -> TableCatalogStoreResult<()> {
|
||||||
let table_bucket = &snapshot.table_bucket.table_bucket;
|
let table_bucket = &snapshot.table_bucket.table_bucket;
|
||||||
|
let active_table_identifiers = snapshot
|
||||||
|
.tables
|
||||||
|
.iter()
|
||||||
|
.filter(|table| table.state == TableCatalogEntryState::Active)
|
||||||
|
.map(|table| (&table.namespace, &table.table))
|
||||||
|
.collect::<BTreeSet<_>>();
|
||||||
|
if snapshot
|
||||||
|
.views
|
||||||
|
.iter()
|
||||||
|
.filter(|view| view.state == TableCatalogEntryState::Active)
|
||||||
|
.any(|view| active_table_identifiers.contains(&(&view.namespace, &view.view)))
|
||||||
|
{
|
||||||
|
return Err(TableCatalogStoreError::Conflict(format!(
|
||||||
|
"table bucket {table_bucket} contains an active table/view identifier collision"
|
||||||
|
)));
|
||||||
|
}
|
||||||
let tables_by_id = snapshot
|
let tables_by_id = snapshot
|
||||||
.tables
|
.tables
|
||||||
.iter()
|
.iter()
|
||||||
@@ -498,6 +598,15 @@ where
|
|||||||
if self.get_table_bucket(table_bucket).await?.is_none() {
|
if self.get_table_bucket(table_bucket).await?.is_none() {
|
||||||
return Err(TableCatalogStoreError::NotFound(format!("table bucket {table_bucket}")));
|
return Err(TableCatalogStoreError::NotFound(format!("table bucket {table_bucket}")));
|
||||||
}
|
}
|
||||||
|
if let Some((global_fence, _)) = self
|
||||||
|
.read_entry::<TableCatalogBackingMigrationGlobalFence>(
|
||||||
|
self.catalog_bucket(),
|
||||||
|
&self.paths.backing_migration_global_fence_path(),
|
||||||
|
)
|
||||||
|
.await?
|
||||||
|
{
|
||||||
|
Self::validate_global_backing_migration_fence(&global_fence)?;
|
||||||
|
}
|
||||||
|
|
||||||
let namespace_objects = self
|
let namespace_objects = self
|
||||||
.backend
|
.backend
|
||||||
@@ -511,6 +620,10 @@ where
|
|||||||
let mut recovery_required_count: usize = 0;
|
let mut recovery_required_count: usize = 0;
|
||||||
let mut manual_review_count: usize = 0;
|
let mut manual_review_count: usize = 0;
|
||||||
let mut warehouse_prefix_owners = BTreeMap::<String, usize>::new();
|
let mut warehouse_prefix_owners = BTreeMap::<String, usize>::new();
|
||||||
|
let mut table_ids = BTreeSet::<String>::new();
|
||||||
|
let mut duplicate_table_identity = false;
|
||||||
|
let mut active_table_identifiers = BTreeSet::<(String, String)>::new();
|
||||||
|
let mut active_view_identifiers = BTreeSet::<(String, String)>::new();
|
||||||
|
|
||||||
for object in namespace_objects {
|
for object in namespace_objects {
|
||||||
if object.ends_with(NAMESPACE_ENTRY_FILE) {
|
if object.ends_with(NAMESPACE_ENTRY_FILE) {
|
||||||
@@ -526,6 +639,9 @@ where
|
|||||||
continue;
|
continue;
|
||||||
};
|
};
|
||||||
validate_view_entry_object(&self.paths, &object, &entry)?;
|
validate_view_entry_object(&self.paths, &object, &entry)?;
|
||||||
|
if entry.state == TableCatalogEntryState::Active {
|
||||||
|
active_view_identifiers.insert((entry.namespace.clone(), entry.view.clone()));
|
||||||
|
}
|
||||||
view_count = view_count.saturating_add(1);
|
view_count = view_count.saturating_add(1);
|
||||||
continue;
|
continue;
|
||||||
}
|
}
|
||||||
@@ -538,7 +654,11 @@ where
|
|||||||
};
|
};
|
||||||
validate_table_entry_object(&self.paths, &object, &table)?;
|
validate_table_entry_object(&self.paths, &object, &table)?;
|
||||||
table_count = table_count.saturating_add(1);
|
table_count = table_count.saturating_add(1);
|
||||||
|
if !table_ids.insert(table.table_id.clone()) {
|
||||||
|
duplicate_table_identity = true;
|
||||||
|
}
|
||||||
if table.state == TableCatalogEntryState::Active {
|
if table.state == TableCatalogEntryState::Active {
|
||||||
|
active_table_identifiers.insert((table.namespace.clone(), table.table.clone()));
|
||||||
let warehouse_prefix = table_warehouse_object_prefix(&table)?;
|
let warehouse_prefix = table_warehouse_object_prefix(&table)?;
|
||||||
warehouse_prefix_owners
|
warehouse_prefix_owners
|
||||||
.entry(warehouse_prefix)
|
.entry(warehouse_prefix)
|
||||||
@@ -548,17 +668,31 @@ where
|
|||||||
|
|
||||||
let recovery = self.table_commit_recovery_report_for_entry(&table, 0).await?;
|
let recovery = self.table_commit_recovery_report_for_entry(&table, 0).await?;
|
||||||
commit_log_count = commit_log_count.saturating_add(recovery.commits.len());
|
commit_log_count = commit_log_count.saturating_add(recovery.commits.len());
|
||||||
idempotency_index_count = idempotency_index_count.saturating_add(
|
for idempotency_object in self
|
||||||
self.backend
|
.backend
|
||||||
.list_objects(
|
.list_objects(
|
||||||
self.catalog_bucket(),
|
self.catalog_bucket(),
|
||||||
&self.paths.commit_idempotency_entries_prefix(table_bucket, &table.table_id),
|
&self.paths.commit_idempotency_entries_prefix(table_bucket, &table.table_id),
|
||||||
)
|
)
|
||||||
|
.await?
|
||||||
|
.into_iter()
|
||||||
|
.filter(|object| object.ends_with(".json"))
|
||||||
|
{
|
||||||
|
let Some((commit, _)) = self
|
||||||
|
.read_entry::<CommitLogEntry>(self.catalog_bucket(), &idempotency_object)
|
||||||
.await?
|
.await?
|
||||||
.into_iter()
|
else {
|
||||||
.filter(|object| object.ends_with(".json"))
|
continue;
|
||||||
.count(),
|
};
|
||||||
);
|
validate_commit_idempotency_entry_object(
|
||||||
|
&self.paths,
|
||||||
|
&idempotency_object,
|
||||||
|
table_bucket,
|
||||||
|
&table.table_id,
|
||||||
|
&commit,
|
||||||
|
)?;
|
||||||
|
idempotency_index_count = idempotency_index_count.saturating_add(1);
|
||||||
|
}
|
||||||
recovery_required_count = recovery_required_count
|
recovery_required_count = recovery_required_count
|
||||||
.saturating_add(recovery.staged_before_table_update_count)
|
.saturating_add(recovery.staged_before_table_update_count)
|
||||||
.saturating_add(recovery.finalization_required_count)
|
.saturating_add(recovery.finalization_required_count)
|
||||||
@@ -568,6 +702,13 @@ where
|
|||||||
|
|
||||||
let warehouse_index_ready = self.warehouse_index_ready(table_bucket).await?;
|
let warehouse_index_ready = self.warehouse_index_ready(table_bucket).await?;
|
||||||
let duplicate_warehouse_prefix_count = warehouse_prefix_owners.values().filter(|count| **count > 1).count();
|
let duplicate_warehouse_prefix_count = warehouse_prefix_owners.values().filter(|count| **count > 1).count();
|
||||||
|
let overlapping_warehouse_prefix = warehouse_prefix_owners
|
||||||
|
.keys()
|
||||||
|
.collect::<Vec<_>>()
|
||||||
|
.windows(2)
|
||||||
|
.any(|window| warehouse_object_prefixes_overlap(window[0], window[1]));
|
||||||
|
let conflicting_warehouse_prefix = duplicate_warehouse_prefix_count > 0 || overlapping_warehouse_prefix;
|
||||||
|
let table_view_identifier_collision_count = active_table_identifiers.intersection(&active_view_identifiers).count();
|
||||||
let mut blockers = Vec::new();
|
let mut blockers = Vec::new();
|
||||||
let mut recommended_actions = Vec::new();
|
let mut recommended_actions = Vec::new();
|
||||||
if recovery_required_count > 0 {
|
if recovery_required_count > 0 {
|
||||||
@@ -583,12 +724,24 @@ where
|
|||||||
blockers.push(TableCatalogBackingMigrationBlocker::WarehouseIndexBackfillRequired);
|
blockers.push(TableCatalogBackingMigrationBlocker::WarehouseIndexBackfillRequired);
|
||||||
recommended_actions.push(TableCatalogBackingMigrationAction::BackfillWarehouseIndex);
|
recommended_actions.push(TableCatalogBackingMigrationAction::BackfillWarehouseIndex);
|
||||||
}
|
}
|
||||||
if duplicate_warehouse_prefix_count > 0 {
|
if conflicting_warehouse_prefix {
|
||||||
blockers.push(TableCatalogBackingMigrationBlocker::DuplicateWarehousePrefix);
|
blockers.push(TableCatalogBackingMigrationBlocker::DuplicateWarehousePrefix);
|
||||||
recommended_actions.push(TableCatalogBackingMigrationAction::ReviewDuplicateWarehousePrefixes);
|
recommended_actions.push(TableCatalogBackingMigrationAction::ReviewDuplicateWarehousePrefixes);
|
||||||
}
|
}
|
||||||
|
if duplicate_table_identity {
|
||||||
|
blockers.push(TableCatalogBackingMigrationBlocker::DuplicateTableIdentity);
|
||||||
|
recommended_actions.push(TableCatalogBackingMigrationAction::ReviewDuplicateTableIdentities);
|
||||||
|
}
|
||||||
|
if table_view_identifier_collision_count > 0 {
|
||||||
|
blockers.push(TableCatalogBackingMigrationBlocker::TableViewIdentifierCollision);
|
||||||
|
recommended_actions.push(TableCatalogBackingMigrationAction::ReviewTableViewIdentifierCollisions);
|
||||||
|
}
|
||||||
|
|
||||||
let mut status = if manual_review_count > 0 || duplicate_warehouse_prefix_count > 0 {
|
let mut status = if manual_review_count > 0
|
||||||
|
|| conflicting_warehouse_prefix
|
||||||
|
|| duplicate_table_identity
|
||||||
|
|| table_view_identifier_collision_count > 0
|
||||||
|
{
|
||||||
TableCatalogBackingMigrationStatus::ManualReviewRequired
|
TableCatalogBackingMigrationStatus::ManualReviewRequired
|
||||||
} else if recovery_required_count > 0 || !warehouse_index_ready {
|
} else if recovery_required_count > 0 || !warehouse_index_ready {
|
||||||
TableCatalogBackingMigrationStatus::RecoveryRequired
|
TableCatalogBackingMigrationStatus::RecoveryRequired
|
||||||
@@ -597,6 +750,14 @@ where
|
|||||||
};
|
};
|
||||||
|
|
||||||
let strong_store = StrongTableCatalogStore::new(self.backend.clone());
|
let strong_store = StrongTableCatalogStore::new(self.backend.clone());
|
||||||
|
let source_table_buckets = self.object_backed_table_buckets().await?;
|
||||||
|
let source_table_bucket_names = source_table_buckets.keys().cloned().collect::<BTreeSet<_>>();
|
||||||
|
let target_table_buckets = strong_store.table_bucket_names().await?;
|
||||||
|
if !target_table_buckets.is_subset(&source_table_bucket_names) {
|
||||||
|
status = TableCatalogBackingMigrationStatus::ManualReviewRequired;
|
||||||
|
blockers.push(TableCatalogBackingMigrationBlocker::DurableStrongSnapshotChanged);
|
||||||
|
recommended_actions.push(TableCatalogBackingMigrationAction::ReviewDurableStrongSnapshot);
|
||||||
|
}
|
||||||
let migration_fence = self.read_backing_migration_fence(table_bucket).await?.map(|(fence, _)| fence);
|
let migration_fence = self.read_backing_migration_fence(table_bucket).await?.map(|(fence, _)| fence);
|
||||||
let object_backed_writes_fenced = migration_fence.is_some();
|
let object_backed_writes_fenced = migration_fence.is_some();
|
||||||
if status == TableCatalogBackingMigrationStatus::ReadyToSnapshot
|
if status == TableCatalogBackingMigrationStatus::ReadyToSnapshot
|
||||||
@@ -633,7 +794,7 @@ where
|
|||||||
Ok(TableCatalogBackingMigrationDryRunReport {
|
Ok(TableCatalogBackingMigrationDryRunReport {
|
||||||
table_bucket: table_bucket.to_string(),
|
table_bucket: table_bucket.to_string(),
|
||||||
source_kind: TableCatalogBackingKind::ObjectBacked,
|
source_kind: TableCatalogBackingKind::ObjectBacked,
|
||||||
target_kind: TableCatalogBackingKind::StrongKvWal,
|
target_kind: TableCatalogBackingKind::DurableStrongSnapshot,
|
||||||
status,
|
status,
|
||||||
namespace_count,
|
namespace_count,
|
||||||
table_count,
|
table_count,
|
||||||
@@ -684,11 +845,17 @@ where
|
|||||||
let existing_fence = self.read_backing_migration_fence(table_bucket).await?;
|
let existing_fence = self.read_backing_migration_fence(table_bucket).await?;
|
||||||
let strong_store = StrongTableCatalogStore::new(self.backend.clone());
|
let strong_store = StrongTableCatalogStore::new(self.backend.clone());
|
||||||
if let Some((fence, _)) = existing_fence.as_ref()
|
if let Some((fence, _)) = existing_fence.as_ref()
|
||||||
&& (fence.version != TABLE_CATALOG_MIGRATION_VERSION || fence.table_bucket != table_bucket)
|
&& fence.status == TableCatalogBackingMigrationFenceStatus::Preparing
|
||||||
|
&& !fence.target_bucket_existed
|
||||||
|
&& Self::migration_target_snapshot_state(fence) == TableCatalogBackingMigrationTargetSnapshotState::Absent
|
||||||
{
|
{
|
||||||
return Err(TableCatalogStoreError::Invalid(format!(
|
strong_store.restore_absent_migration_snapshot_baseline().await?;
|
||||||
"invalid durable strong migration fence for table bucket {table_bucket}"
|
}
|
||||||
)));
|
let source_table_bucket_names = self.object_backed_table_buckets().await?.into_keys().collect::<BTreeSet<_>>();
|
||||||
|
if !strong_store.table_bucket_names().await?.is_subset(&source_table_bucket_names) {
|
||||||
|
return Err(TableCatalogStoreError::Conflict(
|
||||||
|
"durable strong snapshot contains table buckets outside the object-backed catalog inventory".to_string(),
|
||||||
|
));
|
||||||
}
|
}
|
||||||
|
|
||||||
if !self.warehouse_index_ready(table_bucket).await? {
|
if !self.warehouse_index_ready(table_bucket).await? {
|
||||||
@@ -712,10 +879,16 @@ where
|
|||||||
}
|
}
|
||||||
|
|
||||||
self.ensure_global_backing_migration_fence(&global_fence_path).await?;
|
self.ensure_global_backing_migration_fence(&global_fence_path).await?;
|
||||||
|
let (target_fingerprint, target_snapshot_etag) = Self::observe_durable_strong_migration_target(
|
||||||
|
&strong_store,
|
||||||
|
table_bucket,
|
||||||
|
existing_fence.as_ref().map(|(fence, _)| fence),
|
||||||
|
)
|
||||||
|
.await?;
|
||||||
let (migration_id, target_bucket_existed) = if let Some((fence, _)) = existing_fence.as_ref() {
|
let (migration_id, target_bucket_existed) = if let Some((fence, _)) = existing_fence.as_ref() {
|
||||||
(fence.migration_id.clone(), fence.target_bucket_existed)
|
(fence.migration_id.clone(), fence.target_bucket_existed)
|
||||||
} else {
|
} else {
|
||||||
let target_bucket_existed = strong_store.bucket_snapshot_fingerprint(table_bucket).await?.is_some();
|
let target_bucket_existed = target_fingerprint.is_some();
|
||||||
let fence = TableCatalogBackingMigrationFence {
|
let fence = TableCatalogBackingMigrationFence {
|
||||||
version: TABLE_CATALOG_MIGRATION_VERSION,
|
version: TABLE_CATALOG_MIGRATION_VERSION,
|
||||||
table_bucket: table_bucket.to_string(),
|
table_bucket: table_bucket.to_string(),
|
||||||
@@ -723,7 +896,7 @@ where
|
|||||||
status: TableCatalogBackingMigrationFenceStatus::Preparing,
|
status: TableCatalogBackingMigrationFenceStatus::Preparing,
|
||||||
target_bucket_existed,
|
target_bucket_existed,
|
||||||
source_fingerprint: None,
|
source_fingerprint: None,
|
||||||
target_snapshot_etag: None,
|
target_snapshot_etag,
|
||||||
};
|
};
|
||||||
self.write_entry(self.catalog_bucket(), &fence_path, &fence, TableCatalogPutPrecondition::IfAbsent)
|
self.write_entry(self.catalog_bucket(), &fence_path, &fence, TableCatalogPutPrecondition::IfAbsent)
|
||||||
.await?;
|
.await?;
|
||||||
@@ -749,7 +922,7 @@ where
|
|||||||
Ok(TableCatalogBackingMigrationExecutionReport {
|
Ok(TableCatalogBackingMigrationExecutionReport {
|
||||||
table_bucket: table_bucket.to_string(),
|
table_bucket: table_bucket.to_string(),
|
||||||
source_kind: TableCatalogBackingKind::ObjectBacked,
|
source_kind: TableCatalogBackingKind::ObjectBacked,
|
||||||
target_kind: TableCatalogBackingKind::StrongKvWal,
|
target_kind: TableCatalogBackingKind::DurableStrongSnapshot,
|
||||||
status: if created {
|
status: if created {
|
||||||
TableCatalogBackingMigrationExecutionStatus::SnapshotMaterialized
|
TableCatalogBackingMigrationExecutionStatus::SnapshotMaterialized
|
||||||
} else {
|
} else {
|
||||||
@@ -793,11 +966,6 @@ where
|
|||||||
});
|
});
|
||||||
};
|
};
|
||||||
self.ensure_global_backing_migration_fence(&global_fence_path).await?;
|
self.ensure_global_backing_migration_fence(&global_fence_path).await?;
|
||||||
if fence.version != TABLE_CATALOG_MIGRATION_VERSION || fence.table_bucket != table_bucket {
|
|
||||||
return Err(TableCatalogStoreError::Invalid(format!(
|
|
||||||
"invalid durable strong migration fence for table bucket {table_bucket}"
|
|
||||||
)));
|
|
||||||
}
|
|
||||||
|
|
||||||
let mut source_guards = Vec::new();
|
let mut source_guards = Vec::new();
|
||||||
let source = self
|
let source = self
|
||||||
@@ -813,12 +981,21 @@ where
|
|||||||
}
|
}
|
||||||
|
|
||||||
let strong_store = StrongTableCatalogStore::new(self.backend.clone());
|
let strong_store = StrongTableCatalogStore::new(self.backend.clone());
|
||||||
if fence.status == TableCatalogBackingMigrationFenceStatus::Materialized
|
let (target_fingerprint, target_snapshot_etag) =
|
||||||
&& strong_store.bucket_snapshot_fingerprint(table_bucket).await?.as_deref() != Some(&source_fingerprint)
|
Self::observe_durable_strong_migration_target(&strong_store, table_bucket, Some(&fence)).await?;
|
||||||
{
|
if fence.status == TableCatalogBackingMigrationFenceStatus::Materialized {
|
||||||
return Err(TableCatalogStoreError::Conflict(format!(
|
let target_matches_source = target_fingerprint.as_deref() == Some(&source_fingerprint);
|
||||||
"durable strong catalog state changed after materializing table bucket {table_bucket}"
|
let target_was_already_removed = !fence.target_bucket_existed && target_fingerprint.is_none();
|
||||||
)));
|
if !target_matches_source && !target_was_already_removed {
|
||||||
|
return Err(TableCatalogStoreError::Conflict(format!(
|
||||||
|
"durable strong catalog state changed after materializing table bucket {table_bucket}"
|
||||||
|
)));
|
||||||
|
}
|
||||||
|
if target_matches_source && target_snapshot_etag != fence.target_snapshot_etag {
|
||||||
|
return Err(TableCatalogStoreError::Conflict(
|
||||||
|
"durable strong catalog snapshot advanced after materialization".to_string(),
|
||||||
|
));
|
||||||
|
}
|
||||||
}
|
}
|
||||||
if !fence.target_bucket_existed {
|
if !fence.target_bucket_existed {
|
||||||
strong_store
|
strong_store
|
||||||
@@ -836,30 +1013,19 @@ where
|
|||||||
}
|
}
|
||||||
|
|
||||||
async fn all_table_buckets_materialized(&self, strong_store: &StrongTableCatalogStore<B>) -> TableCatalogStoreResult<bool> {
|
async fn all_table_buckets_materialized(&self, strong_store: &StrongTableCatalogStore<B>) -> TableCatalogStoreResult<bool> {
|
||||||
if self
|
let Some((global_fence, _)) = self
|
||||||
.read_entry::<TableCatalogBackingMigrationGlobalFence>(
|
.read_entry::<TableCatalogBackingMigrationGlobalFence>(
|
||||||
self.catalog_bucket(),
|
self.catalog_bucket(),
|
||||||
&self.paths.backing_migration_global_fence_path(),
|
&self.paths.backing_migration_global_fence_path(),
|
||||||
)
|
)
|
||||||
.await?
|
.await?
|
||||||
.is_none()
|
else {
|
||||||
{
|
|
||||||
return Ok(false);
|
return Ok(false);
|
||||||
}
|
};
|
||||||
let table_bucket_objects = self
|
Self::validate_global_backing_migration_fence(&global_fence)?;
|
||||||
.backend
|
let source_table_buckets = self.object_backed_table_buckets().await?;
|
||||||
.list_objects(self.catalog_bucket(), &self.paths.table_bucket_entries_prefix())
|
let source_table_bucket_names = source_table_buckets.keys().cloned().collect::<BTreeSet<_>>();
|
||||||
.await?;
|
for entry in source_table_buckets.values() {
|
||||||
for table_bucket_object in table_bucket_objects
|
|
||||||
.iter()
|
|
||||||
.filter(|object| object.ends_with(TABLE_BUCKET_ENTRY_FILE))
|
|
||||||
{
|
|
||||||
let Some((entry, _)) = self
|
|
||||||
.read_entry::<TableBucketEntry>(self.catalog_bucket(), table_bucket_object)
|
|
||||||
.await?
|
|
||||||
else {
|
|
||||||
return Ok(false);
|
|
||||||
};
|
|
||||||
let Some((fence, _)) = self.read_backing_migration_fence(&entry.table_bucket).await? else {
|
let Some((fence, _)) = self.read_backing_migration_fence(&entry.table_bucket).await? else {
|
||||||
return Ok(false);
|
return Ok(false);
|
||||||
};
|
};
|
||||||
@@ -878,6 +1044,26 @@ where
|
|||||||
return Ok(false);
|
return Ok(false);
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
Ok(true)
|
Ok(strong_store.table_bucket_names().await? == source_table_bucket_names)
|
||||||
|
}
|
||||||
|
|
||||||
|
async fn object_backed_table_buckets(&self) -> TableCatalogStoreResult<BTreeMap<String, TableBucketEntry>> {
|
||||||
|
let mut table_buckets = BTreeMap::new();
|
||||||
|
for object in self
|
||||||
|
.backend
|
||||||
|
.list_objects(self.catalog_bucket(), &self.paths.table_bucket_entries_prefix())
|
||||||
|
.await?
|
||||||
|
.into_iter()
|
||||||
|
.filter(|object| object.ends_with(TABLE_BUCKET_ENTRY_FILE))
|
||||||
|
{
|
||||||
|
let Some((entry, _)) = self.read_entry::<TableBucketEntry>(self.catalog_bucket(), &object).await? else {
|
||||||
|
return Err(TableCatalogStoreError::Conflict(format!(
|
||||||
|
"table bucket changed while reading durable strong migration inventory: {object}"
|
||||||
|
)));
|
||||||
|
};
|
||||||
|
validate_table_bucket_entry_object(&self.paths, &object, &entry)?;
|
||||||
|
table_buckets.insert(entry.table_bucket.clone(), entry);
|
||||||
|
}
|
||||||
|
Ok(table_buckets)
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -21,8 +21,39 @@ mod strong;
|
|||||||
use migration::table_catalog_backing_manifest;
|
use migration::table_catalog_backing_manifest;
|
||||||
pub(crate) use object::ObjectTableCatalogStore;
|
pub(crate) use object::ObjectTableCatalogStore;
|
||||||
#[cfg(test)]
|
#[cfg(test)]
|
||||||
pub(super) use strong::StrongTableCatalogSnapshot;
|
pub(super) use strong::{
|
||||||
pub(crate) use strong::StrongTableCatalogStore;
|
STRONG_TABLE_CATALOG_RELOAD_MAX_ATTEMPTS, STRONG_TABLE_CATALOG_SNAPSHOT_MAX_SIZE, StrongCommitSnapshotRecord,
|
||||||
|
StrongTableCatalogBucketSnapshot, StrongTableCatalogSnapshot, strong_snapshot_write_version,
|
||||||
|
table_catalog_bucket_snapshot_fingerprint,
|
||||||
|
};
|
||||||
|
pub(crate) use strong::{StrongTableCatalogRuntime, StrongTableCatalogStore};
|
||||||
|
|
||||||
|
fn validate_table_bucket_entry(entry: &TableBucketEntry) -> TableCatalogStoreResult<()> {
|
||||||
|
validate_catalog_entry_version("table bucket", entry.version)?;
|
||||||
|
if entry.table_bucket.is_empty() {
|
||||||
|
return Err(TableCatalogStoreError::Invalid("table bucket name cannot be empty".to_string()));
|
||||||
|
}
|
||||||
|
if entry.catalog_type != TABLE_BUCKET_CATALOG_TYPE {
|
||||||
|
return Err(TableCatalogStoreError::Invalid("unsupported table bucket catalog type".to_string()));
|
||||||
|
}
|
||||||
|
Ok(())
|
||||||
|
}
|
||||||
|
|
||||||
|
fn validate_table_entry_version_and_id(entry: &TableEntry) -> TableCatalogStoreResult<()> {
|
||||||
|
validate_catalog_entry_version("table", entry.version)?;
|
||||||
|
if entry.table_id.is_empty() {
|
||||||
|
return Err(TableCatalogStoreError::Invalid("table id cannot be empty".to_string()));
|
||||||
|
}
|
||||||
|
Ok(())
|
||||||
|
}
|
||||||
|
|
||||||
|
fn validate_view_entry_version_and_id(entry: &ViewEntry) -> TableCatalogStoreResult<()> {
|
||||||
|
validate_catalog_entry_version("view", entry.version)?;
|
||||||
|
if entry.view_id.is_empty() {
|
||||||
|
return Err(TableCatalogStoreError::Invalid("view id cannot be empty".to_string()));
|
||||||
|
}
|
||||||
|
Ok(())
|
||||||
|
}
|
||||||
|
|
||||||
fn validate_namespace_entry_identity(entry: &NamespaceEntry) -> TableCatalogStoreResult<Namespace> {
|
fn validate_namespace_entry_identity(entry: &NamespaceEntry) -> TableCatalogStoreResult<Namespace> {
|
||||||
validate_catalog_entry_version("namespace", entry.version)?;
|
validate_catalog_entry_version("namespace", entry.version)?;
|
||||||
@@ -406,8 +437,31 @@ pub(crate) enum TableCatalogPutPrecondition {
|
|||||||
IfMatch(String),
|
IfMatch(String),
|
||||||
}
|
}
|
||||||
|
|
||||||
|
pub(in crate::table_catalog) fn catalog_list_next_continuation(
|
||||||
|
seen: &mut BTreeSet<String>,
|
||||||
|
is_truncated: bool,
|
||||||
|
next: Option<String>,
|
||||||
|
) -> TableCatalogStoreResult<Option<String>> {
|
||||||
|
if !is_truncated {
|
||||||
|
return Ok(None);
|
||||||
|
}
|
||||||
|
let next = next.filter(|next| !next.is_empty()).ok_or_else(|| {
|
||||||
|
TableCatalogStoreError::Internal("truncated catalog object listing has no continuation token".to_string())
|
||||||
|
})?;
|
||||||
|
if !seen.insert(next.clone()) {
|
||||||
|
return Err(TableCatalogStoreError::Internal(
|
||||||
|
"catalog object listing continuation token did not advance".to_string(),
|
||||||
|
));
|
||||||
|
}
|
||||||
|
Ok(Some(next))
|
||||||
|
}
|
||||||
|
|
||||||
#[async_trait::async_trait]
|
#[async_trait::async_trait]
|
||||||
pub(crate) trait TableCatalogObjectBackend: Clone + Send + Sync + 'static {
|
pub(crate) trait TableCatalogObjectBackend: Clone + Send + Sync + 'static {
|
||||||
|
fn strong_catalog_runtime(&self) -> Option<StrongTableCatalogRuntime> {
|
||||||
|
None
|
||||||
|
}
|
||||||
|
|
||||||
async fn read_object(&self, bucket: &str, object: &str) -> TableCatalogStoreResult<Option<TableCatalogObject>>;
|
async fn read_object(&self, bucket: &str, object: &str) -> TableCatalogStoreResult<Option<TableCatalogObject>>;
|
||||||
|
|
||||||
async fn read_object_limited(
|
async fn read_object_limited(
|
||||||
@@ -845,10 +899,16 @@ where
|
|||||||
B: TableCatalogObjectBackend,
|
B: TableCatalogObjectBackend,
|
||||||
{
|
{
|
||||||
pub(crate) fn from_env(backend: B) -> TableCatalogStoreResult<Self> {
|
pub(crate) fn from_env(backend: B) -> TableCatalogStoreResult<Self> {
|
||||||
Ok(Self::new(backend, TableCatalogBackingMode::from_env()?))
|
Ok(match TableCatalogBackingMode::from_env()? {
|
||||||
|
TableCatalogBackingMode::ObjectBacked => Self::ObjectBacked(ObjectTableCatalogStore::new(backend)),
|
||||||
|
TableCatalogBackingMode::DurableStrong => {
|
||||||
|
Self::DurableStrong(StrongTableCatalogStore::new_requiring_snapshot(backend))
|
||||||
|
}
|
||||||
|
})
|
||||||
}
|
}
|
||||||
|
|
||||||
pub(crate) fn new(backend: B, mode: TableCatalogBackingMode) -> Self {
|
#[cfg(test)]
|
||||||
|
pub(crate) fn new_for_test(backend: B, mode: TableCatalogBackingMode) -> Self {
|
||||||
match mode {
|
match mode {
|
||||||
TableCatalogBackingMode::ObjectBacked => Self::ObjectBacked(ObjectTableCatalogStore::new(backend)),
|
TableCatalogBackingMode::ObjectBacked => Self::ObjectBacked(ObjectTableCatalogStore::new(backend)),
|
||||||
TableCatalogBackingMode::DurableStrong => Self::DurableStrong(StrongTableCatalogStore::new(backend)),
|
TableCatalogBackingMode::DurableStrong => Self::DurableStrong(StrongTableCatalogStore::new(backend)),
|
||||||
@@ -1352,12 +1412,14 @@ where
|
|||||||
|
|
||||||
pub(crate) struct EcStoreTableCatalogObjectBackend<S> {
|
pub(crate) struct EcStoreTableCatalogObjectBackend<S> {
|
||||||
store: Arc<S>,
|
store: Arc<S>,
|
||||||
|
strong_runtime: StrongTableCatalogRuntime,
|
||||||
}
|
}
|
||||||
|
|
||||||
impl<S> Clone for EcStoreTableCatalogObjectBackend<S> {
|
impl<S> Clone for EcStoreTableCatalogObjectBackend<S> {
|
||||||
fn clone(&self) -> Self {
|
fn clone(&self) -> Self {
|
||||||
Self {
|
Self {
|
||||||
store: self.store.clone(),
|
store: self.store.clone(),
|
||||||
|
strong_runtime: self.strong_runtime.clone(),
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
@@ -1366,8 +1428,8 @@ impl<S> EcStoreTableCatalogObjectBackend<S>
|
|||||||
where
|
where
|
||||||
S: TableCatalogStorage,
|
S: TableCatalogStorage,
|
||||||
{
|
{
|
||||||
pub fn new(store: Arc<S>) -> Self {
|
pub fn new_with_strong_runtime(store: Arc<S>, strong_runtime: StrongTableCatalogRuntime) -> Self {
|
||||||
Self { store }
|
Self { store, strong_runtime }
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -1378,6 +1440,10 @@ impl<S> TableCatalogObjectBackend for EcStoreTableCatalogObjectBackend<S>
|
|||||||
where
|
where
|
||||||
S: TableCatalogStorage,
|
S: TableCatalogStorage,
|
||||||
{
|
{
|
||||||
|
fn strong_catalog_runtime(&self) -> Option<StrongTableCatalogRuntime> {
|
||||||
|
Some(self.strong_runtime.clone())
|
||||||
|
}
|
||||||
|
|
||||||
async fn read_object(&self, bucket: &str, object: &str) -> TableCatalogStoreResult<Option<TableCatalogObject>> {
|
async fn read_object(&self, bucket: &str, object: &str) -> TableCatalogStoreResult<Option<TableCatalogObject>> {
|
||||||
self.read_object_with_options(bucket, object, ObjectOptions::default(), None)
|
self.read_object_with_options(bucket, object, ObjectOptions::default(), None)
|
||||||
.await
|
.await
|
||||||
@@ -1537,6 +1603,7 @@ where
|
|||||||
|
|
||||||
async fn list_objects(&self, bucket: &str, prefix: &str) -> TableCatalogStoreResult<Vec<String>> {
|
async fn list_objects(&self, bucket: &str, prefix: &str) -> TableCatalogStoreResult<Vec<String>> {
|
||||||
let mut continuation = None;
|
let mut continuation = None;
|
||||||
|
let mut seen_continuations = BTreeSet::new();
|
||||||
let mut objects = BTreeSet::new();
|
let mut objects = BTreeSet::new();
|
||||||
let max_keys = i32::try_from(TABLE_CATALOG_LIST_MAX_KEYS)
|
let max_keys = i32::try_from(TABLE_CATALOG_LIST_MAX_KEYS)
|
||||||
.map_err(|_| TableCatalogStoreError::Internal("catalog list limit exceeds storage API range".to_string()))?;
|
.map_err(|_| TableCatalogStoreError::Internal("catalog list limit exceeds storage API range".to_string()))?;
|
||||||
@@ -1553,14 +1620,10 @@ where
|
|||||||
objects.insert(object.name);
|
objects.insert(object.name);
|
||||||
}
|
}
|
||||||
|
|
||||||
if !result.is_truncated {
|
match catalog_list_next_continuation(&mut seen_continuations, result.is_truncated, result.next_continuation_token)? {
|
||||||
break;
|
Some(next) => continuation = Some(next),
|
||||||
}
|
None => break,
|
||||||
|
|
||||||
let Some(next) = result.next_continuation_token else {
|
|
||||||
break;
|
|
||||||
};
|
};
|
||||||
continuation = Some(next);
|
|
||||||
}
|
}
|
||||||
|
|
||||||
Ok(objects.into_iter().collect())
|
Ok(objects.into_iter().collect())
|
||||||
@@ -1627,20 +1690,6 @@ where
|
|||||||
opts: ObjectOptions,
|
opts: ObjectOptions,
|
||||||
max_size: Option<usize>,
|
max_size: Option<usize>,
|
||||||
) -> TableCatalogStoreResult<Option<TableCatalogObject>> {
|
) -> TableCatalogStoreResult<Option<TableCatalogObject>> {
|
||||||
let info = match self.store.get_object_info(bucket, object, &opts).await {
|
|
||||||
Ok(info) => info,
|
|
||||||
Err(err) if is_missing_storage_error(&err) => return Ok(None),
|
|
||||||
Err(err) => return Err(storage_error_to_catalog("read catalog object info", err)),
|
|
||||||
};
|
|
||||||
if let Some(max_size) = max_size {
|
|
||||||
let object_size = usize::try_from(info.size)
|
|
||||||
.map_err(|_| TableCatalogStoreError::Invalid(format!("catalog object {bucket}/{object} has an invalid size")))?;
|
|
||||||
if object_size > max_size {
|
|
||||||
return Err(TableCatalogStoreError::Invalid(format!(
|
|
||||||
"catalog object {bucket}/{object} exceeds the maximum size of {max_size} bytes"
|
|
||||||
)));
|
|
||||||
}
|
|
||||||
}
|
|
||||||
let mut reader = match self
|
let mut reader = match self
|
||||||
.store
|
.store
|
||||||
.get_object_reader(bucket, object, None, HeaderMap::new(), &opts)
|
.get_object_reader(bucket, object, None, HeaderMap::new(), &opts)
|
||||||
@@ -1650,6 +1699,17 @@ where
|
|||||||
Err(err) if is_missing_storage_error(&err) => return Ok(None),
|
Err(err) if is_missing_storage_error(&err) => return Ok(None),
|
||||||
Err(err) => return Err(storage_error_to_catalog("read catalog object", err)),
|
Err(err) => return Err(storage_error_to_catalog("read catalog object", err)),
|
||||||
};
|
};
|
||||||
|
if let Some(max_size) = max_size {
|
||||||
|
let object_size = usize::try_from(reader.object_info.size)
|
||||||
|
.map_err(|_| TableCatalogStoreError::Invalid(format!("catalog object {bucket}/{object} has an invalid size")))?;
|
||||||
|
if object_size > max_size {
|
||||||
|
return Err(TableCatalogStoreError::Invalid(format!(
|
||||||
|
"catalog object {bucket}/{object} exceeds the maximum size of {max_size} bytes"
|
||||||
|
)));
|
||||||
|
}
|
||||||
|
}
|
||||||
|
let etag = reader.object_info.etag.clone();
|
||||||
|
let mod_time = reader.object_info.mod_time;
|
||||||
let mut data = Vec::new();
|
let mut data = Vec::new();
|
||||||
if let Some(max_size) = max_size {
|
if let Some(max_size) = max_size {
|
||||||
let read_limit = u64::try_from(max_size.saturating_add(1)).unwrap_or(u64::MAX);
|
let read_limit = u64::try_from(max_size.saturating_add(1)).unwrap_or(u64::MAX);
|
||||||
@@ -1666,11 +1726,7 @@ where
|
|||||||
TableCatalogStoreError::Internal(format!("failed to read catalog object {bucket}/{object}: {err}"))
|
TableCatalogStoreError::Internal(format!("failed to read catalog object {bucket}/{object}: {err}"))
|
||||||
})?;
|
})?;
|
||||||
}
|
}
|
||||||
Ok(Some(TableCatalogObject {
|
Ok(Some(TableCatalogObject { data, etag, mod_time }))
|
||||||
data,
|
|
||||||
etag: info.etag,
|
|
||||||
mod_time: info.mod_time,
|
|
||||||
}))
|
|
||||||
}
|
}
|
||||||
|
|
||||||
async fn put_object_with_options(
|
async fn put_object_with_options(
|
||||||
|
|||||||
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
+5559
-95
File diff suppressed because it is too large
Load Diff
Reference in New Issue
Block a user