mirror of
https://github.com/rustfs/rustfs.git
synced 2026-08-16 18:08:21 +00:00
cfa9276fad
* test(admin): pin minio-go Metrics/MetricsV2 wire contract for replication metrics Red-light evidence for backlog#1675 P1-11: ?replication-metrics[=2] serializes the internal snake_case BucketStats family straight onto the wire, while minio-go's replication.Metrics/MetricsV2 expect camelCase tags (currStats/queueStats/replicaCount/queued/...). Go's decoder is case-insensitive but does not ignore underscores, so 'mc replicate status' shows all zeros without any error. The rewritten snapshot tests assert the minio-go tags (plus a synthesized queueStats node — the aggregation path leaves queue_stats.nodes empty today) and fail against the current pass-through serialization. * fix(admin): serialize replication metrics in minio-go wire shapes ?replication-metrics[=2] and the admin replicationmetrics endpoint serialized the internal snake_case BucketStats family straight onto the wire, so 'mc replicate status' decoded all zeros without any error (backlog#1675 P1-11). The internal structs cannot be renamed: they are the intra-cluster peer-RPC wire format (rmp_serde to_vec_named in node_service.rs), pinned by a new regression test. - New admin/replication_metrics_wire.rs: Serialize-only projections onto minio-go replication.Metrics (v1 body, currStats) and MetricsV2 (uptime/currStats/queueStats/downtimeInfo) with the exact json tags; per-target failed becomes the TimedErrStats envelope fed from the FailStats rolling window; the queue peak is dual-emitted as max (MinIO server tag) and peak (minio-go tag). - queueStats synthesizes one node from the bucket queue snapshot — the aggregation path leaves queue_stats.nodes empty, and mc treats an empty node list as 'no data' — and carries transfer summaries (Large/Small/Total) derived from the per-target xfer rates. - Both endpoints share the DTOs; source-health extension keys (provider_available/cluster_complete/...) ride along and are ignored by Go decoders. - Widen the ecstore replication_stats_boundary re-exports (BucketReplicationStat/InQueueMetric/XferStats) so the admin facade chain can name the projected types. * fix(replication): carry failure rolling windows through cluster aggregation Review: both metrics endpoints aggregate first, and FailStats::merge dropped the process-local samples (which also never cross the peer-RPC wire — serde-skipped), so lastMinute/lastHour serialized as zero right after a failure while totals was nonzero. - FailStats gains serializable last_minute/last_hour window snapshots (serde default: old nodes read zeros, new fields are ignored by old decoders), recomputed on every add_size and re-stamped at the per-node collection point (get_latest_replication_stats), and summed by merge. - The wire DTO takes the component-wise max of the live samples and the snapshot, so both the single-node and the aggregated path report the window. - Regression test drives a stat through rmp round trip + merge before serialization, as requested. Also restore the #[allow(dead_code)] attribute to route_policy — the new module declaration had been inserted between the attribute and its item, which broke the -D warnings CI lanes. * fix(replication): bin transfer summaries at 128 MiB and keep window refresh off the hot path Second review round: - update_xfer_rate split at 1 MiB while the minio-go transferSummary labels (and RustFS's own worker-pool split) mean >= 128 MiB for Large, so a 2 MiB replication reported under Large with Small stuck at zero. The producer now bins on MIN_LARGE_OBJ_SIZE; a MetricsV2 assertion covers 2 MiB / 127 MiB / exactly 128 MiB. - add_size no longer recomputes the rolling windows: two full one-hour-deque scans per failure under the bucket-stats write lock made failure bursts quadratic (30k events ~2.1s). The windows are stamped only at the collection point (get_latest_replication_stats, which serves both the local leg and the peer RPC); the aggregation regression now drives that path explicitly before the RPC round trip and merge. * fix(replication): average transfer summaries --------- Co-authored-by: overtrue <anzhengchao@gmail.com>
87 lines
4.4 KiB
Rust
87 lines
4.4 KiB
Rust
// Copyright 2024 RustFS Team
|
|
//
|
|
// Licensed under the Apache License, Version 2.0 (the "License");
|
|
// you may not use this file except in compliance with the License.
|
|
// You may obtain a copy of the License at
|
|
//
|
|
// http://www.apache.org/licenses/LICENSE-2.0
|
|
//
|
|
// Unless required by applicable law or agreed to in writing, software
|
|
// distributed under the License is distributed on an "AS IS" BASIS,
|
|
// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
|
// See the License for the specific language governing permissions and
|
|
// limitations under the License.
|
|
|
|
pub mod datatypes;
|
|
mod replication_bandwidth_boundary;
|
|
mod replication_config_boundary;
|
|
mod replication_config_store;
|
|
mod replication_error_boundary;
|
|
mod replication_event_sink;
|
|
mod replication_filemeta_boundary;
|
|
mod replication_lifecycle_bridge;
|
|
mod replication_lock_boundary;
|
|
mod replication_logging;
|
|
mod replication_metadata_boundary;
|
|
mod replication_migration_bridge;
|
|
mod replication_msgp_boundary;
|
|
mod replication_object_bridge;
|
|
mod replication_object_config;
|
|
mod replication_object_decision_boundary;
|
|
pub(crate) mod replication_pool;
|
|
mod replication_queue_boundary;
|
|
mod replication_resync_boundary;
|
|
mod replication_resyncer;
|
|
mod replication_scanner_bridge;
|
|
mod replication_state;
|
|
mod replication_stats_boundary;
|
|
mod replication_storage_boundary;
|
|
mod replication_tagging_boundary;
|
|
mod replication_target_boundary;
|
|
mod replication_target_config_bridge;
|
|
pub(crate) mod replication_timing;
|
|
mod replication_versioning_boundary;
|
|
mod runtime_boundary;
|
|
|
|
pub use datatypes::ResyncStatusType;
|
|
pub use replication_config_boundary::{
|
|
ObjectOpts, REMOTE_TARGET_CAPABILITY_CONTRACT_VERSION, REMOTE_TARGET_UNSUPPORTED_FIELDS, REMOTE_TARGET_WRITABLE_FIELDS,
|
|
REPLICATION_CAPABILITY_CONTRACT_VERSION, REPLICATION_READ_ONLY_HISTORICAL_FIELDS, REPLICATION_WRITABLE_FIELDS,
|
|
ReplicationConfigStructureError, ReplicationConfigurationExt, ReplicationTargetValidationError,
|
|
invalid_replication_config_status_field, replication_target_arns, should_remove_replication_target,
|
|
unsupported_replication_config_field, validate_replication_config_structure, validate_replication_config_target_arns,
|
|
};
|
|
pub(crate) use replication_filemeta_boundary::version_purge_statuses_map;
|
|
pub use replication_filemeta_boundary::{
|
|
MrfOpKind, MrfReplicateEntry, REPLICATE_INCOMING_DELETE, ReplicateDecision, ReplicateObjectInfo, ReplicationState,
|
|
ReplicationStatusType, ReplicationType, VersionPurgeStatusType, replication_state_to_filemeta,
|
|
replication_status_to_filemeta, replication_statuses_map, version_purge_status_to_filemeta,
|
|
};
|
|
pub(crate) use replication_filemeta_boundary::{
|
|
replication_state_from_filemeta, replication_status_from_filemeta, version_purge_status_from_filemeta,
|
|
};
|
|
pub(crate) use replication_lifecycle_bridge::{ReplicationLifecycleBridge, ReplicationLifecycleConfig};
|
|
pub(crate) use replication_migration_bridge::ReplicationMigrationBridge;
|
|
pub use replication_object_bridge::ReplicationObjectBridge;
|
|
pub use replication_object_config::{DeleteReplicationConfigSnapshot, ReplicationConfig};
|
|
pub use replication_object_decision_boundary::{
|
|
MustReplicateOptions, ReplicationDeleteScheduleInput, ReplicationDeleteStateSource, delete_replication_state_from_config,
|
|
delete_replication_version_id, should_schedule_delete_replication, should_use_existing_delete_replication_info,
|
|
should_use_existing_delete_replication_source,
|
|
};
|
|
pub use replication_pool::{
|
|
DurableMrfBacklog, DynReplicationPool, ReplicationPoolTrait, commit_force_delete_intent, complete_force_delete_intent,
|
|
get_global_replication_pool, get_global_replication_stats, init_background_replication, persist_force_delete_intent,
|
|
read_durable_mrf_backlog, resync_start_conflict_id,
|
|
};
|
|
pub use replication_queue_boundary::{
|
|
DeletedObjectReplicationInfo, ReplicationBatchAdmission, ReplicationHealQueueResult, ReplicationOperation,
|
|
ReplicationPriority, ReplicationQueueAdmission,
|
|
};
|
|
pub use replication_resync_boundary::{BucketReplicationResyncStatus, ResyncOpts, TargetReplicationResyncStatus};
|
|
pub use replication_scanner_bridge::ReplicationScannerBridge;
|
|
pub use replication_state::{ReplicationStats, RuntimeReplicationTargetBacklog};
|
|
pub use replication_stats_boundary::{BucketReplicationStat, BucketReplicationStats, BucketStats, InQueueMetric, XferStats};
|
|
pub use replication_storage_boundary::{ReplicationObjectIO, ReplicationStorage};
|
|
pub(crate) use replication_target_config_bridge::ReplicationTargetConfigBridge;
|