ValkeyCluster deploys Valkey in Cluster mode, handling:
- Topology scheduling
- Slot allocation
- Failovers
- Rolling updates
- ACLs
- Config
- Containers
- Metrics
- Persistence
- Pod disruption budget
- Private image registries
- Scheduling
- TLS
- Users
- Workload type
config:
io-threads: 4
maxmemory-policy: noevictionUse config to pass Valkey configuration to all nodes in the cluster.
Listed below are configurations can be applied live without rolling pods. We are adopting configs that can be applied live on a case-by-case basis. For any requests please raise an issue.
maxclients
maxmemory # There are no safeguards, ensure you do not exceed your container capacity
maxmemory-policy
- Cluster management settings owned by the operator cannot be overwritten
- Operator validates configs before they are applied to the server
containers:
- name: server
env:
- name: MY_VAR
value: "example"
- name: my-sidecar
image: busybox:latest
command: ["sh", "-c", "sleep infinity"]containers patches the pod's container list using strategic merge patch. Containers named server or metrics-exporter are merged by name; anything else is appended as a sidecar.
exporter:
enabled: true # default
image: oliver006/redis_exporter:v1.80.0
resources:
requests:
memory: "64Mi"
cpu: "50m"Each pod runs a metrics-exporter sidecar by default, exposing Prometheus metrics on port 9121. To disable it:
exporter:
enabled: falsepersistence:
size: 10Gi
storageClassName: gp3
reclaimPolicy: RetainWhen persistence is set, the operator manages a PVC for each ValkeyNode. With the save config option, memory state survives pod rolls and partial resyncs are possible.
Retain keeps the PVC when a ValkeyNode is deleted; Delete removes it.
- Only supported with
workloadType: StatefulSet - Cannot be added or removed after creation
- Size can only grow
storageClassNameis immutable
- Live volume expansion
- Automated volume expansion
podDisruptionBudget:
mode: Cluster # defaultThe operator creates a PodDisruptionBudget with maxUnavailable: 1 selecting all pods in the cluster. Set mode: Disabled when the PDB is managed externally or is not required. Omitting podDisruptionBudget entirely is equivalent to mode: Cluster.
| Mode | Behaviour |
|---|---|
Cluster |
Operator creates and owns a single cluster-wide PDB |
Disabled |
Operator deletes the PDB if it exists and does not recreate it |
On SIGTERM (a node drain, eviction, or preemption), a cluster primary fails its slots over to a replica before exiting, so descheduling a primary the operator did not initiate does not leave the shard without a writer. This is enabled by default through the shutdown-on-sigterm failover server config and requires Valkey 9.0+.
The handoff runs inside the pod's termination grace period. With defaults there is comfortable margin: the Kubernetes default terminationGracePeriodSeconds is 30s and the Valkey default cluster-manual-failover-timeout is 5s, so the failover completes well before SIGKILL. If you raise cluster-manual-failover-timeout, the operator raises the derived terminationGracePeriodSeconds to match; see Termination grace period.
terminationGracePeriodSeconds: 60terminationGracePeriodSeconds sets the pod termination grace period for the Valkey nodes. On SIGTERM a primary gracefully fails its slots over to a replica, and that handover has to finish before Kubernetes sends SIGKILL, so the grace period must be at least cluster-manual-failover-timeout plus some headroom.
When omitted, the operator picks a safe value: the larger of the Kubernetes default (30s) and cluster-manual-failover-timeout (default 5s) plus a 10s buffer. With defaults that stays at 30s. Raising cluster-manual-failover-timeout pulls the derived grace period up with it.
An explicit value is honoured as-is, even if it is below the recommended minimum. In that case the operator sets a ConfigurationWarning condition (reason GracePeriodTooShort) on the ValkeyCluster and emits an event when the cluster first enters that state, rather than silently overriding the value. The value must be a positive integer; the CRD rejects zero or negative values.
image: registry.example.com/valkey/valkey:9.0.0
imagePullSecrets:
- name: registrycredentialimagePullSecrets is a list of Secret references (in the cluster's namespace) used to pull images from private registries. It is applied at the pod level, so a single list covers every image in the pod - the Valkey server, the metrics exporter sidecar, and any additional containers. It is optional and has no default; omit it when the nodes already authenticate to the registry.
scheduling:
tolerations:
- key: "dedicated"
operator: "Equal"
value: "valkey"
effect: "NoSchedule"
nodeSelector:
kubernetes.io/arch: amd64
affinity:
podAntiAffinity:
requiredDuringSchedulingIgnoredDuringExecution:
- labelSelector:
matchLabels:
app.kubernetes.io/name: valkey
topologyKey: kubernetes.io/hostname
priorityClassName: high-priorityscheduling.tolerations, scheduling.nodeSelector, scheduling.affinity, and scheduling.priorityClassName are passed through to every pod in the cluster. priorityClassName must reference an existing PriorityClass and protects the Valkey pods from eviction under resource pressure.
scheduling:
topologySpreadConstraints:
- maxSkew: 1
topologyKey: kubernetes.io/hostname
whenUnsatisfiable: DoNotScheduletopologySpreadConstraints uses Kubernetes' native pod topology spread constraints and applies them to every Valkey pod in the cluster.
By default, the operator does not add any topology spread constraints. If topologySpreadConstraints is omitted or empty, pods are scheduled normally using the other scheduling fields such as nodeSelector, affinity, tolerations, and resource requests.
For ValkeyCluster, when a topology spread constraint is configured, the operator makes it shard-aware. It scopes the constraint to the current cluster and shard, so pods from the same shard, for example a primary and its replica, are spread across the configured topology domain.
Each constraint must include:
| Field | Meaning |
|---|---|
maxSkew |
Maximum allowed difference in matching pod count between topology domains. 1 means Kubernetes keeps the matching pods as evenly spread as possible. |
topologyKey |
Node label used as the spread domain. Use kubernetes.io/hostname for worker-node spreading, or labels such as topology.kubernetes.io/zone for zone spreading. |
whenUnsatisfiable |
What Kubernetes should do when the constraint cannot be satisfied. |
whenUnsatisfiable supports:
| Value | Behaviour | Impact |
|---|---|---|
DoNotSchedule |
Hard rule. Kubernetes will not schedule the pod if placement would violate the constraint. | Stronger HA placement, but pods may remain Pending when there are not enough eligible nodes or topology domains. The operator reports PodUnschedulable on the ValkeyCluster. |
ScheduleAnyway |
Soft rule. Kubernetes prefers satisfying the constraint, but can still schedule the pod if it cannot. | Better scheduling availability in constrained clusters, but pods from the same shard may still land in the same topology domain. |
Example strict node-level spreading:
scheduling:
topologySpreadConstraints:
- maxSkew: 1
topologyKey: kubernetes.io/hostname
whenUnsatisfiable: DoNotScheduleThis tries to keep pods from the same shard on different worker nodes. If the cluster does not have enough eligible worker nodes, affected pods stay Pending.
Example preferred node-level spreading:
scheduling:
topologySpreadConstraints:
- maxSkew: 1
topologyKey: kubernetes.io/hostname
whenUnsatisfiable: ScheduleAnywayThis still prefers spreading pods from the same shard across worker nodes, but allows scheduling to continue if the constraint cannot be satisfied.
tls:
certificate:
secretName: valkey-tlstls enables TLS for all cluster communication. The Secret must contain:
| Key | Description |
|---|---|
ca.crt |
Certificate authority |
tls.crt |
Server certificate (or chain) |
tls.key |
Private key for the certificate |
users:
- name: alice
passwordSecret:
name: my-users-secret
keys: [alicepw]
commands:
allow: ["@read", "@write", "@connection"]
deny: ["@admin", "@dangerous"]
keys:
readWrite: ["app:*"]
readOnly: ["shared:*"]
channels:
patterns: ["notifications:*"]
- name: bob
nopass: true
permissions: "+@all ~* &*"users defines per-user ACL rules distributed to every node via a Secret mounted into each pod.
passwordSecret— one or more password keys from a Secret (multiple keys supported for rotation)commands— command categories (@read,@write,@admin, etc.), individual commands, and subcommands to allow or denykeys— key patterns by access type:readWrite,readOnly,writeOnlychannels— pub/sub channel patternspermissions— raw ACL string appended after any generated rules
- Usernames cannot start with
_(reserved for operator-managed system users)
workloadType: StatefulSet # defaultworkloadType controls whether ValkeyNodes use a StatefulSet or a Deployment. Use Deployment for cache-only clusters where you don't need persistent storage or stable pod identity.
- Immutable after creation
persistencerequiresworkloadType: StatefulSet
ValkeyCluster creates a ValkeyNode for each shard/replica position. The ValkeyNode controller owns the underlying StatefulSet or Deployment and its single pod.
graph TD
VC[ValkeyCluster]
VC -->|"1 per shard × node"| VN[ValkeyNode]
VN -->|creates| WL["StatefulSet / Deployment\n(single replica)"]
WL -->|manages| P[Pod]
P --> S[server container]
P --> E[metrics-exporter container]
ValkeyNode is an internal CRD — do not create or modify ValkeyNodes directly. All configuration goes through ValkeyCluster. See ValkeyNode design for why this abstraction exists.
For status conditions and events, see status-conditions.md.