A company is migrating a stateful application to Kubernetes. The application requires persistent storage that is 'zone-aware' to survive a single zone failure and must provide the highest possible I/O performance. Which storage solution best meets these requirements?
Trap 1: Use a network filesystem (NFS) server running as a single pod with…
Running an NFS server as a single pod introduces a critical single point of failure for the entire storage solution. While the underlying regional Persistent Disk offers data redundancy, the NFS server pod itself is not highly available; if it fails, all applications relying on it lose access to storage until it recovers. Furthermore, network file systems inherently add latency compared to block storage directly attached to a node.
Trap 2: Create a StorageClass with WaitForFirstConsumer binding and…
Configuring a StorageClass with volumeBindingMode: WaitForFirstConsumer ensures that a PersistentVolume is provisioned in the same zone where the consuming pod is scheduled. While this helps with zone affinity and prevents premature volume binding in an unsuitable zone, it does not inherently provide zone resilience for the data itself or guarantee high performance. The underlying storage type, not the binding mode, determines data replication and performance characteristics.
Trap 3: Deploy a StatefulSet with a local SSD on each node and use a…
Utilizing local SSDs on each node with a DaemonSet for replication is problematic because local SSDs are ephemeral and tied directly to a specific node; if that node fails, the data is lost unless the custom replication mechanism has successfully synchronized it. Managing data replication across nodes and zones with a DaemonSet is extremely complex, error-prone, and typically less reliable or performant than platform-managed regional storage solutions. This approach lacks inherent zone resilience and adds significant operational overhead.
- A
Use a network filesystem (NFS) server running as a single pod with a PersistentVolume backed by a regional Persistent Disk
Why it fails: Running an NFS server as a single pod introduces a critical single point of failure for the entire storage solution. While the underlying regional Persistent Disk offers data redundancy, the NFS server pod itself is not highly available; if it fails, all applications relying on it lose access to storage until it recovers. Furthermore, network file systems inherently add latency compared to block storage directly attached to a node.
- B
Create a StorageClass with WaitForFirstConsumer binding and volumeBindingMode: WaitForFirstConsumer
Why it fails: Configuring a StorageClass with volumeBindingMode: WaitForFirstConsumer ensures that a PersistentVolume is provisioned in the same zone where the consuming pod is scheduled. While this helps with zone affinity and prevents premature volume binding in an unsuitable zone, it does not inherently provide zone resilience for the data itself or guarantee high performance. The underlying storage type, not the binding mode, determines data replication and performance characteristics.
- C
Use a StorageClass that provisions regional Persistent Disks with replication across two zones
A StorageClass provisioning regional Persistent Disks with replication across two zones directly addresses the requirements for a stateful application. Regional PDs automatically replicate data synchronously between two zones within a region, ensuring high availability and resilience against single zone failures. This design provides both robust data durability and high performance, making it ideal for critical stateful workloads that demand continuous operation and data integrity.
- D
Deploy a StatefulSet with a local SSD on each node and use a DaemonSet to manage replication
Why it fails: Utilizing local SSDs on each node with a DaemonSet for replication is problematic because local SSDs are ephemeral and tied directly to a specific node; if that node fails, the data is lost unless the custom replication mechanism has successfully synchronized it. Managing data replication across nodes and zones with a DaemonSet is extremely complex, error-prone, and typically less reliable or performant than platform-managed regional storage solutions. This approach lacks inherent zone resilience and adds significant operational overhead.