mirror of
https://github.com/rustfs/rustfs.git
synced 2026-08-04 12:27:43 +00:00
fix(ecstore): exclude failed-shard disks from commit so write quorum isn't inflated (backlog#852) (#4309)
`MultiWriter` recorded per-writer failures only in a private `errs` vector and nulled the failed writer locally, but `put_object` never used that to prune the disk set: `rename_data` ran over the full `shuffle_disks`, so a disk that took a short/failed write still had its truncated shard renamed into place and counted as an online disk. The object then claimed N good shards while one was short/corrupt, so a single later disk failure could drop it below reconstructable quorum — silent data loss. - `write_shard`: null the writer on a generic write error too (not only on `ShortWrite`), so a failed writer is uniformly represented as `None` — matching `shutdown_writer`, which already does this. - `put_object`: after encode, drop every disk whose writer failed (`drop_failed_writer_disks`) from `shuffle_disks` before `rename_data`, and re-check write quorum over the survivors. `rename_data` already re-checks quorum and rolls back if too few disks remain, so excluded disks are neither committed nor counted (MinIO sets failed writers to nil before `renameData`). Adds unit tests for the exclusion/quorum accounting. Refs backlog#799 (B3).
This commit is contained in:
@@ -159,6 +159,7 @@ impl<'a> MultiWriter<'a> {
|
||||
}
|
||||
Err(e) => {
|
||||
*err = Some(Error::from(e));
|
||||
*writer_opt = None; // Mark as failed so the caller drops this disk before commit
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user