Be able to map each observability pillar to the right tool and command: Prometheus for metrics, Grafana for dashboards, Fluent Bit for log forwarding, Loki for log aggregation, and kubectl logs for single-pod log retrieval. The key skill is distinguishing collection from storage and visualization.
Start practicing
Cloud Native Observability — choose a session length
Free · No account required
Domain overview
Cloud Native Observability covers how Kubernetes clusters are monitored, logged, and traced. For KCNA, questions focus on the three pillars (metrics, logs, traces), the tools that implement them, and the kubectl commands used to inspect workloads. Expect scenario-style questions matching tools like Prometheus, Grafana Loki, and Fluent Bit to their specific roles in a cluster.
Exam objectives
The three pillars of observability: metrics, logs, and traces, and what each answers
kubectl logs behavior, including container selection and previous-container flags
Prometheus for metrics collection and Grafana for visualization and dashboards
Fluent Bit and Loki as lightweight log collection, forwarding, and aggregation tools
Confusing metrics, logs, and traces, or assuming one tool covers all three pillars without integration.
Believing kubectl logs shows cluster-wide logs; it only retrieves logs from a single pod or container.
Mixing up Fluent Bit (collection and forwarding) with Loki (aggregation and storage) when asked which does what.
Click any question to see the full explanation and answer options, or start a focused practice session above.
Which TWO of the following are best practices for structuring log output in cloud-native applications to maximize observability?
2Which THREE of the following are valid use cases for distributed tracing in a microservices architecture?
3A platform team wants to implement observability for a Kubernetes cluster running 500+ microservices. They need to reduce the cost of storing logs while retaining the ability to search for specific error patterns. Which strategy best achieves this?
4A developer wants to monitor the health of a Kubernetes deployment by checking if the number of ready replicas matches the desired replicas. Which metric from kube-state-metrics should they query?
5Which TWO of the following are best practices for implementing observability in a cloud-native environment?
6A company is running a microservices application on a Kubernetes cluster. They have noticed that one of the services, 'payment-api', is experiencing intermittent high latency. The team wants to identify the root cause without modifying the application code. Which approach should they take?
7Which TWO of the following are recommended practices for achieving observability in a Kubernetes cluster?
8Drag and drop the steps to perform a backup of etcd in a Kubernetes cluster into the correct order.
9Which of the following is a core component of the three pillars of observability?
10What is the primary purpose of Prometheus in cloud native observability?
11In OpenTelemetry, what is the purpose of the Collector component?
12Which Prometheus metric type is best suited to count the number of HTTP requests received?
13What is the purpose of Alertmanager in Prometheus?
14Which tool is primarily used for distributed tracing in cloud native environments?
15What is context propagation in distributed tracing?
16Which component of the metrics-server provides resource metrics like CPU and memory usage?
17What does SLA stand for in the context of service reliability?
18Which log aggregation tool is designed specifically for Kubernetes and is often used as a lightweight alternative to Fluentd?
19A team wants to implement cost monitoring for their Kubernetes clusters. Which approach is most effective?
20Which TWO are pillars of observability? (Select two.)
21Which THREE are responsibilities of the OpenTelemetry project? (Select three.)
22Which of the following is the correct definition of a Service Level Indicator (SLI)?
23What is the primary purpose of structured logging?
24Which Prometheus metric type is best suited for counting the total number of HTTP requests received by a service?
25Which component is responsible for aggregating metrics from Kubernetes nodes and exposing them to the metrics API?
26In the context of distributed tracing, what is a 'span'?
27Which tool is specifically designed for log aggregation and is built by Grafana Labs as a lightweight, cost-effective alternative to traditional log systems?
28A team wants to ensure that at least 99.9% of all requests to their application complete within 500ms over a 30-day window. How should this requirement be classified?
29Which open-source project provides a unified standard for collecting and exporting telemetry data (metrics, logs, and traces) from applications?
30In Prometheus, what is the purpose of the Alertmanager component?
31When using OpenTelemetry, what is the role of the 'Collector'?
32Which TWO of the following are valid Prometheus metric types?
33Which THREE of the following are components of the OpenTelemetry project?
34Which TWO of the following are examples of context propagation mechanisms used in distributed tracing?
35Which of the following is considered one of the three pillars of observability?
36A DevOps team wants to collect logs from all Kubernetes nodes and forward them to a central log storage system. Which tool is specifically designed for lightweight log aggregation and forwarding on Kubernetes nodes?
37A company uses OpenTelemetry to instrument their microservices. They want to ensure that traces from one service can be correlated with those from another service across network calls. Which OpenTelemetry concept enables this correlation?
38Which Prometheus metric type is used to represent a value that can increase or decrease over time, such as memory usage?
39An application is instrumented with OpenTelemetry to export traces to Jaeger. The team notices that some traces are incomplete. What is the most likely cause?
40Which component of the OpenTelemetry architecture is responsible for receiving data from instrumented applications and processing it before export?
41What is the purpose of the metrics-server in Kubernetes?
42A team wants to visualize metrics from Prometheus in a dashboard. Which tool is commonly used for this purpose?
43A company defines an SLO that 99.9% of requests to a service should complete in under 200ms. Which metric type is used to measure this SLO?
44Which tool is specifically designed for distributed tracing and is a Cloud Native Computing Foundation (CNCF) graduated project?
45What does the 'kubectl logs' command retrieve?
46Which TWO of the following are valid Prometheus metric types? (Select two)
47Which THREE of the following are components of the OpenTelemetry project? (Select three)
48Which of the following is NOT one of the three pillars of observability?
49Which tool is commonly used for log aggregation in Kubernetes and is designed to be lightweight?
50A developer wants to view the logs of a specific container named 'sidecar' inside a pod named 'app-pod'. Which command should they use?
51What type of Prometheus metric is best suited to count the total number of HTTP requests received by a service?
52Which of the following is true about Prometheus's pull-based model for collecting metrics?
53What is the primary role of the OpenTelemetry Collector?
54In distributed tracing, what is a 'span'?
55Which tool is specifically designed for distributed tracing and was originally developed by Uber?
56An SRE team defines an SLO that 99.9% of requests to a service should complete in under 500ms over a 30-day rolling window. If the service receives 10 million requests in a month, what is the maximum number of requests that can exceed the latency threshold while still meeting the SLO?
57In PromQL, which function would you use to calculate the per-second rate of increase of a counter over a specified time window?
58What is the main advantage of using OpenTelemetry over vendor-specific instrumentation libraries?
59Which TWO of the following are valid Prometheus metric types? (Select two.)
60Which of the following is NOT one of the three pillars of observability in cloud-native environments?
61A DevOps team wants to collect and forward logs from all nodes in a Kubernetes cluster to a centralized logging backend. Which component is specifically designed for lightweight log collection and forwarding?
62A Prometheus alert rule fires when the error rate exceeds 5% for 5 minutes. The alert is sent to Alertmanager. What must be configured in Alertmanager to ensure the alert is deduplicated, grouped, and routed to the correct team?
63In OpenTelemetry, which component is responsible for receiving, processing, and exporting telemetry data from multiple sources?
64Which TWO of the following are Prometheus metric types? (Select two.)
65Which TWO of the following tools are commonly used for distributed tracing in cloud-native environments? (Select two.)
66Which TWO of the following are valid PromQL functions? (Select two.)
67Which THREE of the following are important considerations when defining SLOs (Service Level Objectives)? (Select three.)
68Which THREE of the following are benefits of structured logging? (Select three.)
69A platform team runs a 12-node Kubernetes cluster where each node hosts roughly 30 pods. They deployed Prometheus with a ServiceMonitor that scrapes every pod's /metrics endpoint every 15 seconds, but now the Prometheus pod is frequently OOMKilled and scrape targets intermittently report 'context deadline exceeded'. Which change best addresses the root cause while preserving observability?
70A platform team runs a Kubernetes cluster with the OpenTelemetry Collector deployed as a DaemonSet. They want to collect node-level metrics such as CPU and memory usage from every node without modifying application code. Which receiver should they configure in the Collector's pipeline?
71A platform team runs a Kubernetes cluster where the kubelet and container runtime expose metrics on each node. The team wants to collect node-level CPU and memory metrics into Prometheus without deploying a separate exporter on every node. Which component should they configure Prometheus to scrape?
72A team runs a Kubernetes cluster with Prometheus scraping application metrics. They want Prometheus to automatically discover new pods and scrape their metrics endpoints as pods are created and destroyed. Which Kubernetes resource should they configure Prometheus to use for this dynamic discovery?
73A site reliability engineer is investigating high latency in a microservices application. They have distributed tracing enabled with OpenTelemetry and traces exported to a backend. They want to correlate a specific slow trace with the logs generated by the involved services. Which practice best enables this correlation?
74A developer instruments a service using the OpenTelemetry SDK and wants to export traces to a backend. The backend expects data in the OTLP format. Which OpenTelemetry component is responsible for receiving, processing, and exporting telemetry data to the backend?
75A developer is investigating a performance issue in a microservices application. They want to trace a single request as it flows through multiple services, including a database call. Which OpenTelemetry concept allows them to correlate all these operations under a single logical unit?
76A Kubernetes operator wants to view real-time CPU and memory usage of pods and nodes in a cluster using kubectl top. Which component must be installed and running in the cluster for this command to work?
77A team is designing a Kubernetes observability stack and wants to use Prometheus for metrics. They need to understand how Prometheus collects and stores time series data. Which TWO of the following statements accurately describe Prometheus? (Choose two.)
78A Kubernetes cluster runs Prometheus for monitoring. The operations team wants to receive alerts when the 99th percentile latency of an HTTP service exceeds 500ms for 5 minutes. They have a metric http_request_duration_seconds histogram. Which PromQL expression should they use to calculate the 99th percentile latency?
79A security team wants to audit which users performed privileged actions in a Kubernetes cluster. They need a record of API requests, including the user, verb, resource, and response status. Which Kubernetes feature should they enable?
80A platform team is deploying a logging stack in Kubernetes. They want to collect logs from all pods and nodes, store them centrally, and provide a query interface. Which combination of tools is commonly used to achieve this?
81A site reliability engineer is investigating high latency in a microservices application. They have distributed traces but need to understand how a single trace spans multiple services and where time is spent. Which OpenTelemetry concept allows correlating spans across service boundaries into a single trace?
82A team is implementing observability for a cloud native application. They want to adopt OpenTelemetry to instrument their code and collect telemetry data. Which TWO of the following are core components of the OpenTelemetry project? (Choose two.)
83A platform engineer is configuring an OpenTelemetry Collector pipeline. They want to ensure that sensitive data such as credit card numbers is not exported to the observability backend. Which component of the Collector should they use to modify or drop attributes before export?
Be able to map each observability pillar to the right tool and command: Prometheus for metrics, Grafana for dashboards, Fluent Bit for log forwarding, Loki for log aggregation, and kubectl logs for single-pod log retrieval. The key skill is distinguishing collection from storage and visualization.
The Courseiva KCNA question bank contains 83 questions in the Cloud Native Observability domain, covering the 8% of the exam attributed to this domain in the official CNCF blueprint. Click any question to see the full explanation and answer breakdown.
Start with a 10-question focused session to identify your baseline accuracy in this domain. Read every explanation — even for questions you answer correctly — to understand the reasoning. Once you score consistently above 80%, move to a 20–30 question session to confirm depth before moving to the next domain.
Yes — the session launcher on this page draws questions exclusively from the Cloud Native Observability domain. Choose 10, 20, 30, or 50 questions for a focused session, or click individual questions to review them one by one.
Save your results, see per-domain analytics, and get readiness scores — free, for every certification.
Sign Up FreeFree forever · Every certification included