CNCF · Free Practice Questions · Last reviewed May 2026
30real exam-style questions organised by domain, each with the correct answer highlighted and a plain-English explanation of why it's right — and why the others are wrong.
44% of exam · 6 sample questions below
A developer deploys a pod that continuously restarts. 'kubectl describe pod' shows the container exits with code 137. What is the most likely cause?
The container is exceeding its memory limit and being OOM-killed.
Exit code 137 equals 128 plus signal 9 (SIGKILL), which the kernel sends when a container exceeds its memory limit. The kubelet then reports the OOMKilled reason. Persistent restarts with this code therefore indicate the pod's memory limit is too low for its workload, not an application crash.
The liveness probe is failing and restarting the container.
The init container is failing and blocking the main container.
The pod is hitting a resource quota limit at the namespace level.
A pod is stuck in 'Pending' state. 'kubectl describe pod' shows '0/4 nodes are available: 4 node(s) had taint {node.kubernetes.io/unreachable: }, that the pod didn't tolerate.' What is the most likely cause?
All nodes have disk pressure.
All nodes are unreachable or have been cordoned.
The taint indicates nodes are unreachable.
The pod has a toleration that matches the taint.
The nodes do not have enough CPU or memory.
A team wants to minimize downtime during a Deployment rollout. Which strategy ensures that new pods are created before old pods are terminated?
Set strategy type to 'Recreate'.
Set strategy type to 'RollingUpdate' with maxSurge=0, maxUnavailable=1.
Set strategy type to 'RollingUpdate' with maxSurge=1, maxUnavailable=0.
RollingUpdate replaces pods incrementally; maxUnavailable=0 forbids any capacity drop, so the controller must create new pods before terminating old ones. maxSurge=1 permits one extra pod above the desired replica count, guaranteeing the zero-downtime constraint the team requires.
Set strategy type to 'RollingUpdate' with maxSurge=1, maxUnavailable=1.
A developer creates a pod that needs to securely access a database password stored in the cluster. Which Kubernetes resource should be used to inject the password as an environment variable?
Secret
A Secret stores sensitive values such as the database password and can be injected into a pod as an environment variable, keeping credentials out of the image and manifest while granting the pod secure access.
ServiceAccount
ConfigMap
PersistentVolumeClaim
An administrator runs 'kubectl get pods' and sees that a pod is in 'Pending' state. 'kubectl describe pod' shows the event: '0/4 nodes are available: 1 node had taints that the pod didn't tolerate, 3 nodes had insufficient memory'. What is the most likely issue?
The node with the taint has a toleration mismatch.
The pod's image pull is failing.
The pod's resource requests exceed available memory on three nodes.
The scheduler reports insufficient memory on three nodes, meaning each lacks the capacity to satisfy the pod's declared resource requests. Requests, not actual usage, drive scheduling decisions, so the pod remains Pending until a node with adequate allocatable memory becomes available.
The pod was evicted due to resource pressure.
A cluster has a node with the taint 'node-role.kubernetes.io/control-plane:NoSchedule'. A pod must be scheduled on this node for a special workload. Which action is required?
Use a nodeSelector to select the node.
Remove the taint from the node.
Add a toleration to the pod spec.
The control-plane taint repels pods lacking a matching toleration, so the scheduler will not place the workload there. Adding a toleration for node-role.kubernetes.io/control-plane with effect NoSchedule to the pod spec permits scheduling onto that node.
Use podAffinity to attract the pod to the node.
Want more Kubernetes Fundamentals practice?
Practice this domain28% of exam · 6 sample questions below
A team deploys a microservice that requires sticky sessions. The service runs on Kubernetes with multiple replicas. Which Kubernetes resource should be used to ensure requests from a client are consistently routed to the same pod?
Headless Service
Service with sessionAffinity: ClientIP
Setting sessionAffinity: ClientIP on the Service makes kube-proxy route repeated requests from the same client IP to the same backend pod, satisfying the sticky session requirement across multiple replicas without an Ingress or external load balancer.
Ingress with default settings
Deployment with hostNetwork: true
A Kubernetes cluster is experiencing network latency. The team suspects that the number of services and endpoints is causing iptables performance degradation. Which CNI plugin or network policy approach is most likely to improve performance?
Switch to Flannel with host-gw backend
Use Calico with iptables mode
Use an eBPF-based CNI plugin like Cilium
Cilium replaces iptables with eBPF programs attached in the kernel, giving O(1) service lookup instead of sequential rule traversal. As services and endpoints grow, iptables latency scales linearly; eBPF hash-map lookups stay constant, directly resolving the degradation described.
Apply a default-deny NetworkPolicy
A developer wants to ensure that a pod runs only on nodes with SSDs. Which mechanism should be used?
Apply a taint to nodes without SSDs and add tolerations to the pod
Use pod anti-affinity
Add a nodeSelector with disktype: ssd
A nodeSelector matches pod scheduling to node labels, so labelling SSD nodes with disktype: ssd and declaring that selector in the pod spec confines the workload to those nodes. This directly satisfies the stem's requirement that the pod run only on SSD-backed nodes, using Kubernetes' built-in label-based scheduling constraint.
Define a ResourceQuota
Which THREE of the following are correct statements about Kubernetes Deployments?
Deployments support canary deployments natively
A Deployment manages ReplicaSets
A Deployment owns and manages ReplicaSets, creating a new ReplicaSet for each revision while retaining older ones. This ownership hierarchy is what enables declarative scaling and rollouts, since the Deployment controller reconciles desired replica counts through the ReplicaSet it governs.
A Deployment directly manages Pods
The default update strategy is RollingUpdate
RollingUpdate is the default strategy, satisfying the stem's requirement for a correct Deployment statement. It replaces pods incrementally, keeping a configurable maximum unavailable and maximum surge, so a new ReplicaSet scales up while the old scales down. This preserves availability during updates without manual intervention.
Deployment supports rolling back to an earlier revision
Deployments retain a configurable revision history, so `kubectl rollout undo` reverts pods to a prior ReplicaSet spec, satisfying the requirement to restore an earlier state. This rollback capability distinguishes Deployments from standalone pods, which keep no revision record and therefore cannot be reverted.
A DevOps team notices that a new deployment of a web application is not receiving traffic even though the pods are running. The deployment has a selector matching the pod labels, and a Service of type ClusterIP exists. What is the most likely cause?
The Service's targetPort does not match the container's containerPort.
The Service routes traffic to the targetPort, which must match the port the container listens on.
The pods do not have a readiness probe defined.
The Service type should be NodePort to receive traffic.
The Service is not exposed via an Ingress.
An administrator needs to ensure that Pods from two different Deployments cannot communicate with each other. Which Kubernetes resource should be used?
NetworkPolicy
NetworkPolicy selects Pods by label and defines ingress and egress rules, so denying traffic between the two Deployments' label sets isolates them. It operates at layer 3/4 within the cluster, which RBAC or namespaces alone cannot enforce.
RBAC Role
PodSecurityPolicy
ResourceQuota
Want more Container Orchestration practice?
Practice this domain12% of exam · 6 sample questions below
A cloud-native application is designed with multiple microservices that need to handle a sudden spike in traffic without manual intervention. Which Kubernetes feature best enables this?
VerticalPodAutoscaler
Cluster Autoscaler
HorizontalPodAutoscaler
HorizontalPodAutoscaler adjusts the replica count of a workload based on observed metrics such as CPU utilisation, scaling pods out during traffic spikes and in afterwards. This satisfies the stem's requirement for automatic, intervention-free handling of sudden load.
PodDisruptionBudget
A team is designing a cloud-native system that must maintain high availability across multiple cloud regions. The application uses Kubernetes clusters in each region. Which approach best ensures that the system can tolerate a full region failure while minimizing complexity?
Deploy a single Kubernetes cluster spanning all regions
Use a global load balancer with active-passive regional failover
A global load balancer with active-passive regional failover keeps one region serving traffic and redirects to the standby on failure, tolerating a full region outage with far less operational complexity than active-active multi-region state replication.
Run active-active in all regions with synchronous data replication
Implement manual failover procedures documented in runbooks
A microservice logs errors when connecting to the database. The logs show 'connection refused'. Which troubleshooting step should be taken first?
Verify the database Service and Endpoints in Kubernetes
'Connection refused' means the client reached a host but nothing accepted the connection, so the Service may have no matching Endpoints. Verifying the Service selector and its Endpoints confirms whether pods are actually registered before investigating DNS, network policy or the database itself.
Scale up the microservice deployment
Restart the microservice pod
Check the logs of other microservices
Which practice is a key principle of cloud-native architecture?
Automated CI/CD pipelines
Automated CI/CD pipelines embody the cloud-native principle of continuous delivery, enabling rapid, reliable and frequent releases through repeatable automation. This satisfies the stem's requirement for a key cloud-native practice by removing manual deployment bottlenecks, supporting small incremental changes, and allowing teams to recover quickly from failures without disrupting running services.
Manual configuration management
Tight coupling of services
Preferring stateful applications over stateless
A cloud-native application uses a service mesh (Istio) for traffic management. The team notices increased latency in inter-service communication. Which likely cause should be investigated first?
Kubernetes Network Policies blocking traffic
Misconfigured sidecar proxy settings
Each Istio sidecar proxies every inbound and outbound request, so an incorrect proxy setting—such as a misconfigured concurrency limit, buffer size or timeout—directly adds per-hop latency to inter-service calls. Inspecting sidecar configuration and Envoy stats isolates this overhead before blaming the application or network.
Application code is not optimized for the mesh
mTLS encryption overhead
A team is designing a cloud-native application that requires each microservice to have its own database. This pattern is known as:
Saga pattern
Database-per-service pattern
Database-per-service gives each microservice a private datastore, accessed only through its own API. This enforces loose coupling and independent schema evolution, avoiding the shared-database integration that would couple services and undermine the autonomy the stem's design requires.
Shared database pattern
CQRS pattern
Want more Cloud Native Architecture practice?
Practice this domain8% of exam · 6 sample questions below
Which TWO of the following are best practices for structuring log output in cloud-native applications to maximize observability?
Include verbose debug-level information in every log line
Use multi-line log entries for detailed error information
Output logs in structured format such as JSON
JSON output gives each log entry discrete, machine-parseable fields rather than an opaque string, so aggregators such as Loki, Elasticsearch or Cloud Logging can index and query individual attributes. This satisfies the observability requirement by enabling filtering, correlation and alerting without brittle regex parsing.
Include a unique request or correlation ID in each log entry
A correlation or request ID propagated across service boundaries lets operators stitch together every log line belonging to one transaction, even when services log independently. This directly satisfies the observability goal by making distributed request tracing possible without relying on timestamps alone.
Avoid timestamps to reduce log size
Which THREE of the following are valid use cases for distributed tracing in a microservices architecture?
Monitoring CPU and memory usage of each service instance
Understanding the dependency graph between microservices
Distributed tracing records spans across service calls, letting teams reconstruct which services invoke which, including latency and error propagation. That reconstructed call graph directly reveals the dependency graph between microservices, a core tracing use case.
Pinpointing the root cause of an error in a distributed transaction
Distributed tracing stitches spans across service boundaries using a shared trace ID, so a failing span's parent chain exposes exactly which service and operation broke a distributed transaction. This satisfies the stem's requirement to pinpoint an error's root cause across microservices rather than within one process.
Identifying which service contributes the most latency to an end-user request
Span durations recorded per service reveal where time is spent across the request path, letting you rank services by cumulative latency contribution. This directly satisfies the stem's need to attribute end-user request latency to a specific service rather than guessing from host metrics.
Capturing detailed error messages and stack traces
A platform team wants to implement observability for a Kubernetes cluster running 500+ microservices. They need to reduce the cost of storing logs while retaining the ability to search for specific error patterns. Which strategy best achieves this?
Increase log retention to one year for compliance
Store all logs in a centralized Elasticsearch cluster with high retention
Aggregate logs into a single pod for easier indexing
Use structured logging and sample debug logs, retaining error logs fully
Structured logging plus sampling debug output slashes stored log volume, directly cutting storage cost, while retaining error logs in full preserves searchability for specific error patterns. This balances the cost constraint against the required troubleshooting capability.
A developer wants to monitor the health of a Kubernetes deployment by checking if the number of ready replicas matches the desired replicas. Which metric from kube-state-metrics should they query?
kube_deployment_status_replicas_ready
`kube_deployment_status_replicas_ready` exposes the count of ready replicas directly from the Deployment's status, satisfying the requirement to compare ready against desired replicas. Pairing it with `kube_deployment_spec_replicas` gives the desired count, enabling an alert when readiness diverges from the specification.
kube_deployment_spec_replicas
kube_node_status_condition
kube_pod_container_status_running
Which TWO of the following are best practices for implementing observability in a cloud-native environment?
Store all raw observability data indefinitely for forensic analysis
Use only metrics and avoid logs to reduce complexity
Add unique request IDs to logs for end-to-end tracing correlation
Unique request IDs let a single transaction be followed across distributed services, satisfying the stem's cloud-native observability requirement for end-to-end tracing. Because each request carries its own identifier, logs from separate pods and microservices can be correlated without relying on hostnames or timestamps, which are unreliable in ephemeral, dynamically scheduled workloads.
Randomly sample all traces and logs to reduce storage
Use structured logging (e.g., JSON format) for easier automated parsing
Structured logging emits machine-readable JSON with consistent field names, so log aggregators can parse, filter and correlate events without brittle regex. This directly satisfies the stem's automated-parsing requirement, unlike unstructured text, and scales across ephemeral containers where per-instance manual inspection is impractical.
A company is running a microservices application on a Kubernetes cluster. They have noticed that one of the services, 'payment-api', is experiencing intermittent high latency. The team wants to identify the root cause without modifying the application code. Which approach should they take?
Monitor CPU and memory metrics from kube-state-metrics and correlate with latency.
Increase log verbosity for all services and search for error messages.
Implement distributed tracing using tools like Jaeger or Zipkin to trace requests across services.
Distributed tracing tracks request flow and identifies slow components.
Check node-level metrics using Prometheus Node Exporter.
Want more Cloud Native Observability practice?
Practice this domain8% of exam · 6 sample questions below
A startup wants to minimize downtime during application updates in Kubernetes. Which deployment strategy should they use?
RollingUpdate
RollingUpdate replaces pods incrementally, keeping a proportion of replicas available via maxUnavailable and maxSurge, so the Service always has ready endpoints. This satisfies the requirement to minimise downtime during updates without extra tooling or duplicate environments.
Canary
Blue/Green
Recreate
A DevOps engineer notices that after a Helm upgrade, the new pods are crash looping with 'ImagePullBackOff'. What is the most likely cause?
The pod's liveness probe is misconfigured
The Helm chart has a wrong image tag
ImagePullBackOff means the kubelet cannot pull the specified image, most commonly because the tag referenced in the chart does not exist in the registry. A Helm upgrade that changed the image tag would produce exactly this crash-looping symptom.
The service account lacks permissions
The deployment's resource requests exceed node capacity
A team wants to implement GitOps for their Kubernetes workloads using Argo CD. They have multiple environments (dev, staging, prod) in separate clusters. What is the best practice for structuring the Git repository?
A single branch with all environment manifests in the same folder
Separate repositories per environment
Store all manifests in a single file with environment labels
A monorepo with a directory per environment and overlays for differences
A monorepo with a directory per environment, using overlays (for example Kustomize) to express per-environment differences, keeps shared manifests DRY while allowing dev, staging and prod to diverge safely. This satisfies the constraint of managing multiple separate clusters from one auditable source of truth, with Argo CD Applications targeting each overlay.
Which TWO of the following are benefits of using Helm for application delivery?
Automatic scaling based on CPU usage
Ability to roll back to previous releases
Helm tracks each release as a revision, so helm rollback restores the previous manifest and Kubernetes objects. This gives deterministic recovery from a failed upgrade, satisfying the benefit of reverting to an earlier working release without manual reapplication.
Automatic canary deployments
Simplified packaging and templating of Kubernetes resources
Helm charts package related Kubernetes manifests with templating and values files, so parameters vary per environment while structure stays consistent. This removes duplicated YAML and simplifies deploying complex applications, satisfying the benefit of simplified packaging and templating.
Built-in monitoring and alerting
Which THREE of the following practices are essential for a secure cloud native CI/CD pipeline?
Sign container images and verify signatures during deployment
Signing images at build time and verifying signatures at admission ensures only artefacts produced by trusted pipeline identities reach the cluster, blocking tampered or substituted images. This satisfies the supply-chain integrity requirement of a secure cloud-native CI/CD pipeline.
Store secrets in plain text in the pipeline configuration
Use a single long-lived service account for all pipeline steps
Scan container images for vulnerabilities before deployment
Scanning images for known CVEs before deployment catches vulnerable base images and dependencies while they are still cheap to fix, preventing flawed artefacts from reaching production. This satisfies the pipeline's requirement to gate on vulnerability findings before release.
Apply least-privilege IAM roles to pipeline components
Least-privilege IAM roles bound to pipeline service accounts limit blast radius if a build agent or credential is compromised, restricting what an attacker can read or deploy. This satisfies the pipeline's requirement to constrain permissions for its own components.
A team is deploying a new microservice that processes sensitive user data. They want to ensure that secrets such as database passwords are not exposed in the container image or environment variables. Which approach should they use?
Embed the secret directly in the Docker image and use it via environment variables
Store the secret in a ConfigMap and reference it in the pod spec
Store the secret in a Kubernetes Secret and mount it as a volume in the pod
Kubernetes Secrets store credentials separately from the image and pod specification, and mounting them as volumes injects the data as files at runtime, so passwords never appear in container layers or environment variables. This satisfies the requirement to avoid exposing sensitive data.
Use a PersistentVolumeClaim to store the secret and mount it into the pod
Want more Cloud Native Application Delivery practice?
Practice this domainThe KCNA exam is performance-based — there are no multiple-choice questions. It is a hands-on lab exam completed within 90 minutes. You complete practical tasks in a live or simulated environment. Courseiva practice questions cover the underlying concepts.
Hands-on labs and command-line tasks in a live Kubernetes cluster.
The exam covers 5 domains: Kubernetes Fundamentals, Container Orchestration, Cloud Native Architecture, Cloud Native Observability, Cloud Native Application Delivery. Questions are weighted by domain — higher-weight domains appear more on your actual exam.
No. These are original exam-style practice questions written against the official CNCF KCNA exam objectives. They are not copied from the real exam. Courseiva focuses on genuine understanding, not memorisation of braindumps.
Courseiva tracks your accuracy per domain and routes you toward weak areas automatically. Free, no account required.