Courseiva

KCNA · domain

Cloud Native Observability

Cloud Native Observability covers how Kubernetes clusters are monitored, logged, and traced. For KCNA, questions focus on the three pillars (metrics, logs, traces), the tools that implement them, and the kubectl commands used to inspect workloads. Expect scenario-style questions matching tools like Prometheus, Grafana Loki, and Fluent Bit to their specific roles in a cluster.

83 questions18 easy40 medium25 hard

Focused practice

Practice Cloud Native Observability questions

Scored sessions drawing only from this domain — pick a length below.

Start 20-question practice test →

What this domain covers

What to know about Cloud Native Observability

Be able to map each observability pillar to the right tool and command: Prometheus for metrics, Grafana for dashboards, Fluent Bit for log forwarding, Loki for log aggregation, and kubectl logs for single-pod log retrieval. The key skill is distinguishing collection from storage and visualization.

The three pillars of observability: metrics, logs, and traces, and what each answers

kubectl logs behavior, including container selection and previous-container flags

Prometheus for metrics collection and Grafana for visualization and dashboards

Fluent Bit and Loki as lightweight log collection, forwarding, and aggregation tools

Why learners struggle

Why Cloud Native Observability questions are commonly missed

NAT questions are missed when learners confuse the four address types (inside local, inside global, outside local, outside global) or misapply the interface direction. A translation rule can look correct but still fail if the ACL, interface, or direction is wrong.

  • ·Inside local vs inside global — inside local is the private source, inside global is the translated public address
  • ·PAT overloads — many sources share one public IP using unique port numbers
  • ·Interface direction — ip nat inside and ip nat outside must be on the correct interfaces
  • ·Static NAT vs dynamic NAT vs PAT — each serves a different use case
  • ·The NAT ACL identifies traffic to translate, not traffic to permit or deny
  • ·A missing translation can look like a routing problem if the interfaces are misconfigured

Watch out for

Common Cloud Native Observability exam traps

  • ▸Confusing metrics, logs, and traces, or assuming one tool covers all three pillars without integration.
  • ▸Believing kubectl logs shows cluster-wide logs; it only retrieves logs from a single pod or container.
  • ▸Mixing up Fluent Bit (collection and forwarding) with Loki (aggregation and storage) when asked which does what.

Question index

All Cloud Native Observability questions (83)

Click any question to see the full explanation, or start a practice session above.

1

A security team wants to audit which users performed privileged actions in a Kubernetes cluster. They need a record of API requests, including the user, verb, resource, and response status. Which Kubernetes feature should they enable?

Easy
2

When using OpenTelemetry, what is the role of the 'Collector'?

Hard
3

Which TWO of the following are valid PromQL functions? (Select two.)

Medium
4

A site reliability engineer is investigating high latency in a microservices application. They have distributed traces but need to understand how a single trace spans multiple services and where time is spent. Which OpenTelemetry concept allows correlating spans across service boundaries into a single trace?

Hard
5

Which Prometheus metric type is used to represent a value that can increase or decrease over time, such as memory usage?

Easy
6

Which THREE of the following are components of the OpenTelemetry project?

Hard
7

What is the main advantage of using OpenTelemetry over vendor-specific instrumentation libraries?

Hard
8

Which Prometheus metric type is best suited for counting the total number of HTTP requests received by a service?

Easy
9

What is the purpose of the metrics-server in Kubernetes?

Easy
10

Which TWO of the following are valid Prometheus metric types?

Medium
11

In OpenTelemetry, which component is responsible for receiving, processing, and exporting telemetry data from multiple sources?

Medium
12

In Prometheus, what is the purpose of the Alertmanager component?

Hard
13

Drag and drop the steps to perform a backup of etcd in a Kubernetes cluster into the correct order.

Medium
14

Which of the following is a core component of the three pillars of observability?

Medium
15

Which Prometheus metric type is best suited to count the number of HTTP requests received?

Medium
16

Which of the following is NOT one of the three pillars of observability?

Easy
17

An application is instrumented with OpenTelemetry to export traces to Jaeger. The team notices that some traces are incomplete. What is the most likely cause?

Medium
18

What type of Prometheus metric is best suited to count the total number of HTTP requests received by a service?

Medium
19

Which tool is specifically designed for log aggregation and is built by Grafana Labs as a lightweight, cost-effective alternative to traditional log systems?

Medium
20

Which TWO of the following are Prometheus metric types? (Select two.)

Medium
21

A site reliability engineer is investigating high latency in a microservices application. They have distributed tracing enabled with OpenTelemetry and traces exported to a backend. They want to correlate a specific slow trace with the logs generated by the involved services. Which practice best enables this correlation?

Hard
22

Which of the following is considered one of the three pillars of observability?

Easy
23

A developer instruments a service using the OpenTelemetry SDK and wants to export traces to a backend. The backend expects data in the OTLP format. Which OpenTelemetry component is responsible for receiving, processing, and exporting telemetry data to the backend?

Medium
24

Which log aggregation tool is designed specifically for Kubernetes and is often used as a lightweight alternative to Fluentd?

Medium
25

What is the primary purpose of Prometheus in cloud native observability?

Easy
26

An SRE team defines an SLO that 99.9% of requests to a service should complete in under 500ms over a 30-day rolling window. If the service receives 10 million requests in a month, what is the maximum number of requests that can exceed the latency threshold while still meeting the SLO?

Hard
27

A team is implementing observability for a cloud native application. They want to adopt OpenTelemetry to instrument their code and collect telemetry data. Which TWO of the following are core components of the OpenTelemetry project? (Choose two.)

Medium
28

A platform team wants to implement observability for a Kubernetes cluster running 500+ microservices. They need to reduce the cost of storing logs while retaining the ability to search for specific error patterns. Which strategy best achieves this?

Hard
29

A DevOps team wants to collect logs from all Kubernetes nodes and forward them to a central log storage system. Which tool is specifically designed for lightweight log aggregation and forwarding on Kubernetes nodes?

Medium
30

A team wants to visualize metrics from Prometheus in a dashboard. Which tool is commonly used for this purpose?

Medium
31

A team wants to implement cost monitoring for their Kubernetes clusters. Which approach is most effective?

Hard
32

Which tool is specifically designed for distributed tracing and is a Cloud Native Computing Foundation (CNCF) graduated project?

Medium
33

Which component of the OpenTelemetry architecture is responsible for receiving data from instrumented applications and processing it before export?

Medium
34

Which TWO of the following tools are commonly used for distributed tracing in cloud-native environments? (Select two.)

Easy
35

In PromQL, which function would you use to calculate the per-second rate of increase of a counter over a specified time window?

Hard
36

In OpenTelemetry, what is the purpose of the Collector component?

Hard
37

Which of the following is the correct definition of a Service Level Indicator (SLI)?

Easy
38

Which TWO of the following are recommended practices for achieving observability in a Kubernetes cluster?

Hard
39

In distributed tracing, what is a 'span'?

Medium
40

Which component is responsible for aggregating metrics from Kubernetes nodes and exposing them to the metrics API?

Medium
41

Which of the following is NOT one of the three pillars of observability in cloud-native environments?

Easy
42

Which THREE of the following are important considerations when defining SLOs (Service Level Objectives)? (Select three.)

Hard
43

A Kubernetes cluster runs Prometheus for monitoring. The operations team wants to receive alerts when the 99th percentile latency of an HTTP service exceeds 500ms for 5 minutes. They have a metric http_request_duration_seconds histogram. Which PromQL expression should they use to calculate the 99th percentile latency?

Hard
44

In the context of distributed tracing, what is a 'span'?

Medium
45

What is context propagation in distributed tracing?

Hard
46

What is the primary purpose of structured logging?

Easy
47

Which tool is primarily used for distributed tracing in cloud native environments?

Medium
48

Which THREE of the following are benefits of structured logging? (Select three.)

Medium
49

What is the purpose of Alertmanager in Prometheus?

Easy
50

Which THREE are responsibilities of the OpenTelemetry project? (Select three.)

Hard
51

Which TWO of the following are best practices for implementing observability in a cloud-native environment?

Medium
52

A platform team runs a Kubernetes cluster where the kubelet and container runtime expose metrics on each node. The team wants to collect node-level CPU and memory metrics into Prometheus without deploying a separate exporter on every node. Which component should they configure Prometheus to scrape?

Medium
53

A developer wants to monitor the health of a Kubernetes deployment by checking if the number of ready replicas matches the desired replicas. Which metric from kube-state-metrics should they query?

Easy
54

A platform team runs a Kubernetes cluster with the OpenTelemetry Collector deployed as a DaemonSet. They want to collect node-level metrics such as CPU and memory usage from every node without modifying application code. Which receiver should they configure in the Collector's pipeline?

Medium
55

A team wants to ensure that at least 99.9% of all requests to their application complete within 500ms over a 30-day window. How should this requirement be classified?

Medium
56

Which of the following is true about Prometheus's pull-based model for collecting metrics?

Medium
57

A Prometheus alert rule fires when the error rate exceeds 5% for 5 minutes. The alert is sent to Alertmanager. What must be configured in Alertmanager to ensure the alert is deduplicated, grouped, and routed to the correct team?

Hard
58

A platform team runs a 12-node Kubernetes cluster where each node hosts roughly 30 pods. They deployed Prometheus with a ServiceMonitor that scrapes every pod's /metrics endpoint every 15 seconds, but now the Prometheus pod is frequently OOMKilled and scrape targets intermittently report 'context deadline exceeded'. Which change best addresses the root cause while preserving observability?

Hard
59

A developer is investigating a performance issue in a microservices application. They want to trace a single request as it flows through multiple services, including a database call. Which OpenTelemetry concept allows them to correlate all these operations under a single logical unit?

Medium
60

A platform team is deploying a logging stack in Kubernetes. They want to collect logs from all pods and nodes, store them centrally, and provide a query interface. Which combination of tools is commonly used to achieve this?

Easy
61

A company uses OpenTelemetry to instrument their microservices. They want to ensure that traces from one service can be correlated with those from another service across network calls. Which OpenTelemetry concept enables this correlation?

Hard
62

What does the 'kubectl logs' command retrieve?

Easy
63

Which tool is commonly used for log aggregation in Kubernetes and is designed to be lightweight?

Easy
64

Which TWO are pillars of observability? (Select two.)

Medium
65

Which open-source project provides a unified standard for collecting and exporting telemetry data (metrics, logs, and traces) from applications?

Medium
66

Which THREE of the following are valid use cases for distributed tracing in a microservices architecture?

Medium
67

What does SLA stand for in the context of service reliability?

Easy
68

Which TWO of the following are best practices for structuring log output in cloud-native applications to maximize observability?

Hard
69

A developer wants to view the logs of a specific container named 'sidecar' inside a pod named 'app-pod'. Which command should they use?

Medium
70

A platform engineer is configuring an OpenTelemetry Collector pipeline. They want to ensure that sensitive data such as credit card numbers is not exported to the observability backend. Which component of the Collector should they use to modify or drop attributes before export?

Hard
71

A company is running a microservices application on a Kubernetes cluster. They have noticed that one of the services, 'payment-api', is experiencing intermittent high latency. The team wants to identify the root cause without modifying the application code. Which approach should they take?

Medium
72

Which THREE of the following are components of the OpenTelemetry project? (Select three)

Hard
73

A team runs a Kubernetes cluster with Prometheus scraping application metrics. They want Prometheus to automatically discover new pods and scrape their metrics endpoints as pods are created and destroyed. Which Kubernetes resource should they configure Prometheus to use for this dynamic discovery?

Medium
74

Which TWO of the following are examples of context propagation mechanisms used in distributed tracing?

Hard
75

Which TWO of the following are valid Prometheus metric types? (Select two)

Medium
76

Which TWO of the following are valid Prometheus metric types? (Select two.)

Medium
77

Which component of the metrics-server provides resource metrics like CPU and memory usage?

Medium
78

Which tool is specifically designed for distributed tracing and was originally developed by Uber?

Medium
79

What is the primary role of the OpenTelemetry Collector?

Medium
80

A team is designing a Kubernetes observability stack and wants to use Prometheus for metrics. They need to understand how Prometheus collects and stores time series data. Which TWO of the following statements accurately describe Prometheus? (Choose two.)

Hard
81

A Kubernetes operator wants to view real-time CPU and memory usage of pods and nodes in a cluster using kubectl top. Which component must be installed and running in the cluster for this command to work?

Easy
82

A company defines an SLO that 99.9% of requests to a service should complete in under 200ms. Which metric type is used to measure this SLO?

Hard
83

A DevOps team wants to collect and forward logs from all nodes in a Kubernetes cluster to a centralized logging backend. Which component is specifically designed for lightweight log collection and forwarding?

Medium

Frequently asked questions

What does the Cloud Native Observability domain cover on the KCNA exam?
Be able to map each observability pillar to the right tool and command: Prometheus for metrics, Grafana for dashboards, Fluent Bit for log forwarding, Loki for log aggregation, and kubectl logs for single-pod log retrieval. The key skill is distinguishing collection from storage and visualization.
How many questions are in this domain?
This page lists all 83 Cloud Native Observability questions in the KCNA question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
What is the best way to practise this domain?
Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
Can I practise only Cloud Native Observability questions?
Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.
cncf-kcna CNCF-KCNA kcna observability Practice Questions