mirror of
https://github.com/rustfs/rustfs.git
synced 2026-08-12 08:06:54 +00:00
feat(helm): support multiple server pools for capacity expansion (#3325)
* feat(helm): support multiple server pools for capacity expansion
The server already understands multiple pools (RUSTFS_VOLUMES is split on
spaces, one pool expression each; rc admin pool/expand/rebalance/
decommission exist), but the chart could only render a single StatefulSet.
Add an opt-in pools mode:
- pools.enabled=false (default) renders byte-identical output to the
current chart (verified against five values variants)
- pools.list renders one StatefulSet per entry; entries may set
replicaCount (4 or 16) and a storageclass block, everything else is
inherited from the top-level values
- pool 0 keeps the legacy StatefulSet/pod/PVC names and its (immutable)
selector, so existing single-pool deployments can be expanded in place
without renaming or data loss; additional pools render as
<fullname>-poolN with a rustfs.com/pool label in their selector
- RUSTFS_VOLUMES and RUSTFS_SERVER_DOMAINS enumerate every pool
- pod anti-affinity is scoped per pool so pools may share nodes while
each pool still spreads its own pods across distinct nodes
- template validation fails fast on unsupported replicaCount, an empty
pools.list, or pools in standalone mode
- pools are append-only by design (list index = identity), documented in
values.yaml and the README
* fix(helm): truncate pool names DNS-safely, document PDB pool scope
Truncate <fullname>-poolN to 63 chars by shortening the base name rather
than the suffix, so the pool index always survives and two pools of a
max-length release cannot collide. Document that the single
PodDisruptionBudget deliberately spans all pools.
* fix(helm): pool-mode anti-affinity must be preferred, not required
Found by live-testing in-place expansion on a 5-node cluster (4-replica
pool 0 + 4-replica pool 1): with requiredDuringScheduling anti-affinity
the expansion deadlocks in a cycle that cannot self-heal -
1. the not-yet-rolled pool-0 pods still carry the unscoped required
rule and repel every rustfs pod from their nodes, so part of pool 1
stays Pending (no IP, no headless-DNS record);
2. every pod that already has the new RUSTFS_VOLUMES exits fatally on
'failed to lookup address information' for those Pending peers;
3. the StatefulSet rolling update is gated on the crashing pod going
Ready, so the remaining pool-0 pods never roll and keep repelling.
Because StatefulSets assign revisions per ordinal, even a subsequent
template fix cannot rescue a wedged rollout (Pending ordinals are only
recreated with the new template after the higher ordinals go Ready) -
the new pool's StatefulSet has to be recreated. Shipping the soft
affinity from the start avoids the state entirely; expansion was
re-tested end-to-end with it and converges.
Pool-mode rendering only - pools.enabled=false still renders the
existing required anti-affinity byte-identically.
---------
Co-authored-by: cxymds <Cxymds@qq.com>
Co-authored-by: houseme <housemecn@gmail.com>
Co-authored-by: cxymds <cxymds@gmail.com>
This commit is contained in:
@@ -133,6 +133,8 @@ uer. `ClusterIssuer` or `Issuer`. |
|
||||
| pdb.maxUnavailable | string | `1` | |
|
||||
| pdb.minAvailable | string | `""` | |
|
||||
| podAnnotations | object | `{}` | |
|
||||
| pools.enabled | bool | `false` | Enable multiple server pools (capacity expansion, distributed mode only). |
|
||||
| pools.list | list | `[]` | One entry per pool; entries may set `replicaCount` (4 or 16) and `storageclass`, omitted fields inherit top-level values. Append-only. |
|
||||
| podLabels | object | `{}` | |
|
||||
| podSecurityContext.fsGroup | int | `10001` | |
|
||||
| podSecurityContext.runAsGroup | int | `10001` | |
|
||||
@@ -216,6 +218,59 @@ Both approaches support pulling from private registries seamlessly and you can a
|
||||
|
||||
- The default size for data and logs dir is **256Mi** which must satisfy the production usage,you should specify `storageclass.dataStorageSize` and `storageclass.logStorageSize` to change the size, for example, 1Ti for data and 1Gi for logs.
|
||||
|
||||
# Server pools (capacity expansion)
|
||||
|
||||
In distributed mode the chart can run multiple **server pools** — independent
|
||||
StatefulSets whose drives together form one cluster, the same expansion model
|
||||
the RustFS server already supports via space-separated `RUSTFS_VOLUMES`
|
||||
expressions (`rc admin pool ls` / `expand` / `rebalance` / `decommission`).
|
||||
|
||||
With `pools.enabled=false` (default) the chart behaves exactly as before:
|
||||
one StatefulSet driven by the top-level `replicaCount`/`storageclass`.
|
||||
|
||||
To expand an existing deployment, enable pools and describe the current
|
||||
layout as pool 0 plus your new capacity:
|
||||
|
||||
```yaml
|
||||
pools:
|
||||
enabled: true
|
||||
list:
|
||||
- {} # pool 0: inherits top-level values and keeps the
|
||||
# existing StatefulSet/pod/PVC names and data
|
||||
- replicaCount: 4 # pool 1: new capacity (4 or 16)
|
||||
storageclass:
|
||||
dataStorageSize: 10Gi
|
||||
```
|
||||
|
||||
Each entry may set `replicaCount` (4 or 16) and/or a `storageclass` block;
|
||||
omitted fields inherit the top-level values. Additional pools render as
|
||||
`<fullname>-pool<N>` StatefulSets; all pools share the headless service,
|
||||
the main service, the configuration and the credentials.
|
||||
|
||||
Notes:
|
||||
|
||||
* **Pools are append-only.** The list index determines the StatefulSet name —
|
||||
never remove or reorder entries. Retire a pool with
|
||||
`rc admin decommission` before removing it from the list.
|
||||
* During the expansion rollout, pods restart until every pod of every pool is
|
||||
resolvable — the server refuses to start with unresolvable peers, so expect
|
||||
a few crash/restart cycles before the cluster converges. This is harmless.
|
||||
* After the cluster converges, run `rc admin rebalance start <alias>` to
|
||||
spread existing objects across the new pool.
|
||||
* Pod anti-affinity in pool mode is scoped per pool and **preferred**
|
||||
(soft), not required: two pools can share nodes, and each pool's own pods
|
||||
spread across distinct nodes when capacity allows. Soft affinity is
|
||||
load-bearing for in-place expansion — with required rules, the
|
||||
not-yet-rolled pods of the existing pool block the new pool's pods from
|
||||
their nodes while the rolled pods crash on the unresolvable (Pending)
|
||||
peers, deadlocking the rollout on any cluster with fewer nodes than the
|
||||
total pod count. Single-pool deployments (`pools.enabled=false`) keep the
|
||||
chart's existing required anti-affinity unchanged.
|
||||
* The PodDisruptionBudget spans all pools: with the default
|
||||
`pdb.maxUnavailable: 1`, at most one pod of the whole cluster may be
|
||||
evicted at a time. This is deliberately conservative — quorum safety
|
||||
matters across the union of all pools.
|
||||
|
||||
# Installation
|
||||
|
||||
## Requirement
|
||||
|
||||
Reference in New Issue
Block a user