Courseiva

CCNA Cloud Native Observability Questions

75 of 83 questions · Page 1/2 · Cloud Native Observability · Answers revealed

1
MCQeasy

A security team wants to audit which users performed privileged actions in a Kubernetes cluster. They need a record of API requests, including the user, verb, resource, and response status. Which Kubernetes feature should they enable?

A.Metrics-server
B.Pod Security Admission
C.Event recording
D.Audit logging
AnswerD

Kubernetes audit logging records API server requests, capturing details such as the authenticated user, the verb, the resource, and the response status. Enabling audit policies and log backends provides the security team with the requested record of privileged actions. This directly satisfies the requirement to audit who did what in the cluster, making it the correct feature.

Why this answer

Audit logging in Kubernetes records API server requests with user, verb, resource, and response details, which is exactly what the security team needs to trace privileged actions. Event recording is limited to cluster events, metrics-server provides resource metrics, and Pod Security Admission enforces policies but does not create an audit trail. Only audit logging delivers the required request-level records.

Exam trap

The trap here is confusing event recording with audit logging, assuming that cluster events provide a complete record of API requests and user actions.

2
MCQhard

When using OpenTelemetry, what is the role of the 'Collector'?

A.To alert on abnormal metrics
B.To store traces for long-term retention
C.To receive, process, and export telemetry data in a vendor-neutral way
D.To instrument application code manually
AnswerC

The Collector's pipeline receives telemetry via receivers, optionally processes it, then exports through exporters, decoupling applications from backends. This vendor-neutral mediation satisfies the stem's requirement that data be handled independently of any specific monitoring vendor.

Why this answer

The OpenTelemetry Collector is a vendor-agnostic proxy that receives telemetry data (traces, metrics, logs) from instrumented applications, processes it (e.g., batching, filtering, enrichment), and exports it to one or more backends (e.g., Jaeger, Prometheus, or any OTLP-compatible system). It decouples data generation from data storage, enabling flexible, scalable observability pipelines without vendor lock-in.

Exam trap

CNCF often tests the misconception that the Collector is a storage or alerting system, when in fact it is a stateless pipeline component that only receives, processes, and exports telemetry data.

How to eliminate wrong answers

Option A is wrong because alerting on abnormal metrics is the responsibility of monitoring systems like Prometheus with Alertmanager, not the OpenTelemetry Collector, which focuses on data ingestion, processing, and export. Option B is wrong because long-term storage of traces is handled by backend systems (e.g., Jaeger, Tempo) or databases; the Collector is a pipeline component that forwards data, not a persistent store. Option D is wrong because manual instrumentation of application code is done using OpenTelemetry SDKs and APIs (e.g., for traces, metrics), while the Collector operates as a separate infrastructure component that receives already-instrumented telemetry.

3
Multi-Selectmedium

Which TWO of the following are valid PromQL functions? (Select two.)

Select 2 answers
A.topk()
B.histogram_quantile()
C.rate()
D.avg()
E.sum()
AnswersB, C

histogram_quantile() calculates quantiles from histogram metrics.

Why this answer

rate() and histogram_quantile() are common PromQL functions. avg_over_time() is also valid but avg is not a function, it's an aggregation operator.

4
MCQhard

A site reliability engineer is investigating high latency in a microservices application. They have distributed traces but need to understand how a single trace spans multiple services and where time is spent. Which OpenTelemetry concept allows correlating spans across service boundaries into a single trace?

A.Metric exemplars
B.Resource attributes
C.Span status
D.Context propagation
AnswerD

Context propagation is the mechanism that passes trace context, including trace ID and span ID, across service boundaries via headers or other carriers. It enables spans created in different services to be linked into a single distributed trace. Without it, each service would create isolated traces, so this concept directly addresses the need to correlate spans across microservices and understand end-to-end latency.

Why this answer

Context propagation carries trace and span identifiers across service calls, allowing spans from different services to be stitched into one distributed trace. Resource attributes describe the producer, span status indicates success or failure, and metric exemplars link metrics to traces, but none of these provide the cross-service linkage required. Therefore, context propagation is the concept that enables end-to-end trace correlation.

Exam trap

The trap here is assuming that attaching trace IDs to metrics or adding resource attributes is enough to correlate spans, when only context propagation actually carries trace context between services.

5
MCQeasy

Which Prometheus metric type is used to represent a value that can increase or decrease over time, such as memory usage?

A.Gauge
B.Histogram
C.Summary
D.Counter
AnswerA

A gauge holds a snapshot value that can rise or fall arbitrarily, matching memory usage sampled at scrape time. Unlike counters, which only increase (or reset), gauges support set, increment and decrement operations, making them the correct type for fluctuating measurements.

Why this answer

A gauge is the Prometheus metric type designed to represent a value that can arbitrarily go up or down, such as memory usage, temperature, or current connections. Unlike counters, gauges are not monotonic and can be set to any value, making them ideal for instantaneous measurements. Memory usage fluctuates as processes allocate and free memory, so a gauge accurately captures this behavior.

Exam trap

KCNA often tests the distinction between metric types by presenting a scenario that seems like a counter (e.g., 'number of requests') but actually requires a gauge (e.g., 'current memory usage'), so candidates must carefully read whether the value can decrease.

How to eliminate wrong answers

Option B is wrong because a histogram samples observations (usually request durations or response sizes) and counts them in configurable buckets, providing a distribution, not a single fluctuating value. Option C is wrong because a summary also samples observations and calculates configurable quantiles over a sliding time window, which is not suitable for representing a simple value like memory usage. Option D is wrong because a counter is a cumulative metric that only increases (or resets to zero on restart), so it cannot represent a value that decreases, such as memory usage.

6
Multi-Selecthard

Which THREE of the following are components of the OpenTelemetry project?

Select 3 answers
A.Prometheus
B.Collector
C.Specification
D.Jaeger
E.SDKs (Software Development Kits)
AnswersB, C, E

The Collector is a vendor-agnostic pipeline component that receives, processes and exports telemetry, satisfying OpenTelemetry's requirement for a decoupled data path between instrumented applications and backends. It ships as a standalone binary or agent, letting you centralise batching, filtering and retries without altering application code.

Why this answer

The OpenTelemetry project is composed of several core deliverables, and the Collector (B) is one of them: it is a vendor-neutral agent/gateway that receives, processes, and exports telemetry (traces, metrics, logs) via receivers, processors, and exporters. The Specification (C) is also a core component, defining the cross-language API, SDK, protocol (OTLP), and semantic conventions that all implementations must follow. SDKs (E) are likewise a core component, providing language-specific implementations (e.g., Java, Go, Python, .NET) that generate and export telemetry according to the specification.

Prometheus (A) is a separate CNCF monitoring system and time-series database, not a component of OpenTelemetry, though OTel can export metrics to it. Jaeger (D) is a separate distributed tracing backend (also CNCF) that predates and is independent of OpenTelemetry, even though it can receive OTLP data.

7
MCQhard

What is the main advantage of using OpenTelemetry over vendor-specific instrumentation libraries?

A.It provides a single, vendor-agnostic instrumentation standard
B.It eliminates the need for logging
C.It automatically reduces latency
D.It is the only tool that supports traces
AnswerA

OpenTelemetry supplies one vendor-neutral instrumentation standard, so traces, metrics and logs are emitted once and exported to any compatible backend. This avoids rewriting instrumentation when changing vendors, which is the core advantage over proprietary libraries.

Why this answer

OpenTelemetry provides a unified standard that avoids vendor lock-in, allowing data to be sent to any backend.

8
MCQeasy

Which Prometheus metric type is best suited for counting the total number of HTTP requests received by a service?

A.Summary
B.Counter
C.Gauge
D.Histogram
AnswerB

A Counter is a monotonically increasing metric that only goes up, making it ideal for cumulative totals such as HTTP requests received. Rate or increase functions applied to the counter yield per-second request rates, whereas Gauge and Histogram serve different measurement purposes.

Why this answer

A counter is a cumulative metric that only increases (or resets to zero). It is ideal for counting events like HTTP requests.

9
MCQeasy

What is the purpose of the metrics-server in Kubernetes?

A.To provide resource usage metrics for pods and nodes
B.To manage service meshes
C.To collect application logs
D.To store historical metrics
AnswerA

The metrics-server aggregates kubelet-reported resource consumption and exposes it through the Metrics API, letting Horizontal Pod Autoscalers and kubectl top read live CPU and memory figures for pods and nodes. It satisfies the requirement for cluster-wide resource usage visibility, unlike full monitoring stacks that persist historical data.

Why this answer

The metrics-server is a cluster-wide aggregator of resource usage data. It collects CPU and memory metrics from each node's kubelet (via the Summary API) and exposes them through the Kubernetes API server using the Metrics API (metrics.k8s.io). This enables core Kubernetes components like the Horizontal Pod Autoscaler (HPA) and Vertical Pod Autoscaler (VPA) to make scaling decisions, and allows users to view resource usage with `kubectl top`.

Exam trap

KCNA often tests the distinction between metrics-server (real-time resource metrics for HPA) and full monitoring solutions like Prometheus (historical metrics and long-term storage), so candidates may incorrectly assume metrics-server stores historical data.

How to eliminate wrong answers

Option B is wrong because service mesh management is handled by dedicated tools like Istio, Linkerd, or Consul, not by metrics-server. Option C is wrong because log collection is performed by logging agents such as Fluentd, Filebeat, or Loki, not by metrics-server. Option D is wrong because metrics-server only provides real-time, in-memory metrics with a short retention period (typically 15 minutes); it does not store historical metrics—that requires a full monitoring solution like Prometheus with long-term storage.

10
Multi-Selectmedium

Which TWO of the following are valid Prometheus metric types?

Select 2 answers
A.Quantile
B.Counter
C.Timer
D.Meter
E.Gauge
AnswersB, E

Counter is a valid Prometheus metric type representing a monotonically increasing cumulative value that only resets on restart. It is one of the four core types alongside gauge, histogram and summary, satisfying the question's requirement for valid metric types.

Why this answer

Option B (Counter) is correct because Prometheus defines the counter metric type as a cumulative value that only increases or resets to zero on restart, typically used for counts of events or errors. Option E (Gauge) is correct because Prometheus defines the gauge metric type as a value that can arbitrarily go up and down, used for measurements like temperature, memory usage, or current queue size. The other options are not Prometheus metric types: Quantile (A) is a concept used in summary/histogram quantile calculations rather than a metric type itself, while Timer (C) and Meter (D) are metric types from other monitoring libraries such as Dropwizard Metrics or Micrometer, not native Prometheus types.

Exam trap

KCNA often tests Prometheus metric types by including terms from other monitoring ecosystems — candidates pick Timer or Meter (StatsD/Dropwizard concepts) or Quantile (a computed value, not a type) instead of the four actual Prometheus types.

11
MCQmedium

In OpenTelemetry, which component is responsible for receiving, processing, and exporting telemetry data from multiple sources?

A.OpenTelemetry Collector
B.OpenTelemetry SDK
C.OpenTelemetry Exporter
D.OpenTelemetry API
AnswerA

The OpenTelemetry Collector receives telemetry via receivers, processes it through processors such as batching, then exports it to backends. This satisfies the stem's requirement to handle data from multiple sources, acting as a vendor-neutral pipeline between instrumented services and observability platforms.

Why this answer

The OpenTelemetry Collector is a vendor-agnostic proxy that can receive, process, and export telemetry data from multiple sources. It supports various receivers (e.g., OTLP, Jaeger, Prometheus) to ingest data, processors to modify or filter it, and exporters to send it to backends. This central component decouples instrumentation from backends, enabling flexible pipelines.

Exam trap

KCNA often tests the distinction between the Collector and other OpenTelemetry components, and candidates may confuse the Collector with the SDK or exporter, forgetting that the Collector is the only component designed to receive, process, and export data from multiple sources.

How to eliminate wrong answers

Option B is wrong because the OpenTelemetry SDK is a library used within applications to generate and emit telemetry data, not a standalone service for receiving from multiple sources. Option C is wrong because an exporter is a component within the Collector or SDK that sends data to a backend; it does not receive or process data from multiple sources. Option D is wrong because the OpenTelemetry API is a set of interfaces for instrumenting code, not a component that receives, processes, or exports telemetry.

12
MCQhard

In Prometheus, what is the purpose of the Alertmanager component?

A.To scrape metrics from targets
B.To provide a graphical dashboard for metrics
C.To manage, group, and route alerts to notification channels like email or Slack
D.To store historical metrics data long-term
AnswerC

Alertmanager receives alerts fired by Prometheus, then deduplicates, groups, and routes them to receivers such as email, Slack, or PagerDuty. It also handles silencing and inhibition, preventing notification floods during outages. This satisfies the stem's requirement to manage, group, and route alerts to notification channels.

Why this answer

Alertmanager is a separate component from the Prometheus server that receives alerts generated by Prometheus alerting rules. Its core purpose is to deduplicate, group, and route those alerts to the correct receiver—such as email, Slack, PagerDuty, or a webhook—based on routing rules and labels. It also handles silencing and inhibition to reduce noise during incidents, which is essential for effective on-call alerting.

Exam trap

KCNA often tests the misconception that Alertmanager is responsible for generating alerts or storing metrics, when in fact it only handles routing and notification after Prometheus evaluates alerting rules.

How to eliminate wrong answers

Option A is wrong because scraping metrics from targets is the job of the Prometheus server itself, which pulls metrics over HTTP from configured scrape targets; Alertmanager does not scrape. Option B is wrong because graphical dashboards are provided by tools like Grafana or Prometheus's own expression browser, not by Alertmanager, which has only a basic UI for viewing and silencing alerts. Option D is wrong because long-term storage of historical metrics is handled by Prometheus's local time-series database (TSDB) or remote storage integrations like Thanos, Cortex, or Mimir; Alertmanager does not store metrics.

13
Drag & Dropmedium

Drag and drop the steps to perform a backup of etcd in a Kubernetes cluster into the correct order.

Drag or tap steps into the slots.

Steps
Order
1Step 1
2Step 2
3Step 3
4Step 4

Why this order

Access the node, save snapshot, verify, store securely, and restore when necessary.

14
MCQmedium

Which of the following is a core component of the three pillars of observability?

A.Alerting
B.SLIs
C.Logs
D.Dashboards
AnswerC

Logs are one of the three pillars of observability, alongside metrics and traces. They record discrete, timestamped events with contextual detail, enabling retrospective debugging of why a system behaved as it did, which metrics and traces alone cannot fully explain.

Why this answer

The three pillars of observability are logs, metrics, and traces. Logs are discrete, timestamped records of events that provide detailed context about what happened in a system, making them a foundational pillar. Alerting, SLIs, and dashboards are downstream consumers or derived artifacts built on top of these pillars, not pillars themselves.

Exam trap

KCNA often tests whether candidates confuse observability data sources (logs, metrics, traces) with observability tooling and derived signals (dashboards, alerts, SLIs), so candidates who pick 'Alerting' or 'Dashboards' mistake consumption layers for foundational pillars.

How to eliminate wrong answers

Option A is wrong because alerting is a notification mechanism triggered by thresholds on metrics or logs — it is a consumer of observability data, not a source pillar. Option B is wrong because SLIs (Service Level Indicators) are quantitative measures of service performance derived from metrics, not a primary observability data type. Option D is wrong because dashboards are visualization layers that aggregate and display metrics, logs, and traces — they are presentation tooling, not a data pillar.

15
MCQmedium

Which Prometheus metric type is best suited to count the number of HTTP requests received?

A.Gauge
B.Histogram
C.Summary
D.Counter
AnswerD

Counters are monotonically increasing metrics that only go up or reset to zero on restart, making them ideal for cumulative counts such as total HTTP requests received. Use rate() or increase() to derive per-second request rates from the counter.

Why this answer

A Counter is a Prometheus metric type that represents a cumulative value that only increases (or resets to zero on restart). It is the correct choice for counting total HTTP requests received because the value monotonically increases over time. PromQL's rate() and increase() functions are designed to work with counters to compute per-second rates.

Exam trap

KCNA often tests the misconception that Histogram or Summary is needed for request counting because they also expose a _count — candidates overlook that Counter is the purpose-built type for monotonically increasing totals.

How to eliminate wrong answers

Option A is wrong because a Gauge represents a value that can go up or down arbitrarily (e.g., current memory usage or temperature) — it is not suitable for cumulative counts. Option B is wrong because a Histogram samples observations (like request durations) into configurable buckets and provides sum and count, but it is used for distributions, not simple request counting. Option C is wrong because a Summary is similar to a Histogram but calculates quantiles over a sliding time window on the client side — it is for latency/percentile tracking, not raw request counts.

16
MCQeasy

Which of the following is NOT one of the three pillars of observability?

A.Metrics
B.Logs
C.Traces
D.Alerts
AnswerD

Alerts are a downstream response mechanism built on observability data, not a source of it. The three pillars are logs, metrics, and traces; alerting consumes these signals through rules, so it does not itself constitute a pillar.

Why this answer

The three pillars of observability are metrics, logs, and traces. Metrics are numeric measurements aggregated over time, logs are timestamped records of discrete events, and traces follow a request's path through distributed systems. Alerts are not a pillar; they are a downstream action or notification triggered by conditions derived from the pillars.

Therefore, alerts is the correct answer as the item that is NOT a pillar.

Exam trap

The trap is that alerts feel like a core part of observability because they are prominent in monitoring tools, but the exam expects you to recognize the three data types that provide observability, not the actions taken on them.

How to eliminate wrong answers

Option A is wrong because metrics are indeed one of the three pillars, providing quantitative data such as CPU usage or request rates. Option B is wrong because logs are a core pillar, capturing detailed event data for debugging and auditing. Option C is wrong because traces are the third pillar, essential for understanding request flow and latency in microservices.

Option D is correct because alerts are a response mechanism built on top of observability data, not a fundamental data type.

17
MCQmedium

An application is instrumented with OpenTelemetry to export traces to Jaeger. The team notices that some traces are incomplete. What is the most likely cause?

A.Context propagation is not correctly implemented
B.Span attributes are missing
C.Sampling rate is too high
D.Jaeger database is full
AnswerA

Incomplete traces typically arise when trace context is not propagated across service boundaries, so downstream spans are created without the parent trace ID and appear as separate traces rather than linked spans within one trace.

Why this answer

Incomplete traces often occur when context propagation is not implemented correctly, causing spans to be disconnected.

18
MCQmedium

What type of Prometheus metric is best suited to count the total number of HTTP requests received by a service?

A.Histogram
B.Gauge
C.Summary
D.Counter
AnswerD

Counters are monotonically increasing values, so each HTTP request increments the total without resetting. This satisfies the stem's requirement to count cumulative requests; the rate() function then derives requests per second from that running total.

Why this answer

A Counter is a cumulative metric that only increases (or resets to zero on restart), making it ideal for tracking the total number of events like HTTP requests. Since the total request count is monotonically increasing, a Counter directly represents this without additional processing. Other metric types like Gauge, Histogram, and Summary serve different purposes: Gauge can go up and down, Histogram and Summary are for distributions (e.g., request durations).

Exam trap

KCNA often tests the distinction between metric types, and a common trap is confusing Counter with Gauge or Histogram, especially when the question mentions 'total number'—candidates might incorrectly think a Gauge can track totals or that a Histogram is needed for counting, but the key is that a Counter is specifically for cumulative counts.

How to eliminate wrong answers

Option A is wrong because a Histogram is used to track the distribution of values (e.g., request durations or sizes) across predefined buckets, not for simple counting. Option B is wrong because a Gauge represents a value that can arbitrarily go up and down, such as current memory usage or temperature, and is not suitable for a cumulative count. Option C is wrong because a Summary, like a Histogram, is designed to capture distributions and calculate quantiles over a sliding time window, not to count total events.

19
MCQmedium

Which tool is specifically designed for log aggregation and is built by Grafana Labs as a lightweight, cost-effective alternative to traditional log systems?

A.Loki
B.Zipkin
C.Prometheus
D.Jaeger
AnswerA

Loki indexes only metadata labels rather than full log text, storing compressed chunks in object storage, which satisfies the lightweight, cost-effective constraint. Built by Grafana Labs, it aggregates logs alongside Prometheus-style metrics, making it the purpose-built alternative to traditional full-text indexing systems such as Elasticsearch.

Why this answer

Loki is a log aggregation system optimized for Kubernetes, designed to be cost-effective and easy to operate.

20
Multi-Selectmedium

Which TWO of the following are Prometheus metric types? (Select two.)

Select 2 answers
A.Event
B.Gauge
C.Set
D.Counter
E.Timer
AnswersB, D

Gauge is a Prometheus metric type representing a value that can arbitrarily increase or decrease, such as current memory usage or temperature. It is one of the four core types alongside counter, histogram and summary.

Why this answer

Prometheus defines exactly four core metric types, and Gauge (B) is one of them: it represents a value that can arbitrarily go up or down, such as current temperature or memory usage. Counter (D) is also a core Prometheus metric type: it is a cumulative value that only increases (or resets to zero on restart), used for things like total requests served. The other options are not Prometheus metric types: Event (A) is not a Prometheus concept, Set (C) is a StatsD metric type, and Timer (E) is also a StatsD metric type, so none of them belong to Prometheus's metric model.

Exam trap

KCNA often tests the confusion between Prometheus metric types and general monitoring concepts like timers or events, which belong to other systems such as StatsD.

21
MCQhard

A site reliability engineer is investigating high latency in a microservices application. They have distributed tracing enabled with OpenTelemetry and traces exported to a backend. They want to correlate a specific slow trace with the logs generated by the involved services. Which practice best enables this correlation?

A.Enable debug logging on all services and increase log verbosity.
B.Configure the OpenTelemetry Collector to sample all traces at 100%.
C.Use a separate logging backend that is different from the tracing backend.
D.Include the trace ID and span ID in the log entries of each service.
AnswerD

Including the trace ID and span ID in log entries allows logs to be directly linked to the corresponding trace. When investigating a slow trace, the engineer can search logs for that trace ID and see all related log lines across services. This is a standard practice in cloud native observability and is supported by OpenTelemetry's logging integration, enabling efficient root cause analysis.

Why this answer

Correlating traces and logs requires a shared identifier. By injecting the trace ID and span ID into log records, each log line can be tied to the exact trace and span that produced it. This allows an engineer to take a slow trace, extract its trace ID, and query logs for that ID to see all related events across services.

Other options do not provide this direct linkage.

Exam trap

The trap here is thinking that increasing log detail or sampling rates enables correlation, when the essential requirement is propagating trace context into logs.

22
MCQeasy

Which of the following is considered one of the three pillars of observability?

A.Events
B.Metrics
C.Alerts
D.Profiles
AnswerB

Metrics are numerical measurements aggregated over time, forming one of observability's three pillars alongside logs and traces. They satisfy the stem's requirement by enabling efficient, low-cost monitoring of system health and alerting on thresholds, since pre-aggregated time-series data scales far better than inspecting individual events.

Why this answer

The three pillars of observability in cloud native environments are logs, metrics, and traces.

23
MCQmedium

A developer instruments a service using the OpenTelemetry SDK and wants to export traces to a backend. The backend expects data in the OTLP format. Which OpenTelemetry component is responsible for receiving, processing, and exporting telemetry data to the backend?

A.OpenTelemetry Protocol (OTLP)
B.OpenTelemetry SDK
C.OpenTelemetry Collector
D.OpenTelemetry API
AnswerC

The OpenTelemetry Collector is a vendor-neutral agent that receives telemetry in various formats, processes it, and exports it to one or more backends. It supports OTLP natively and can batch, filter, and transform data before export. In this scenario, it acts as the intermediary that takes the SDK's OTLP output and delivers it to the backend, making it the correct component for receiving, processing, and exporting telemetry.

Why this answer

The OpenTelemetry Collector is designed to receive telemetry in OTLP and other formats, apply processors such as batching or filtering, and export to one or more backends. The API and SDK are for instrumenting and generating data inside applications, while OTLP is only the transport protocol. Therefore, the Collector is the component that receives, processes, and exports telemetry in this pipeline.

Exam trap

The trap here is confusing the OTLP protocol with the Collector, treating the transport format as if it were the runtime component that receives and routes telemetry.

24
MCQmedium

Which log aggregation tool is designed specifically for Kubernetes and is often used as a lightweight alternative to Fluentd?

A.Logstash
B.Fluent Bit
C.Loki
D.Elasticsearch
AnswerB

Fluent Bit is a CNCF-hosted, C-based log processor and forwarder built for containerised environments, with a far smaller memory and CPU footprint than Fluentd. That lightweight design makes it the common Kubernetes alternative to Fluentd for node-level log collection.

Why this answer

Fluent Bit is a CNCF-graduated, lightweight log processor and forwarder written in C with a small memory footprint, explicitly designed for containerized and Kubernetes environments. It is commonly deployed as a DaemonSet to collect node and pod logs and forward them to backends like Elasticsearch, Loki, or S3. Its low resource usage makes it the standard 'lightweight alternative to Fluentd' in Kubernetes logging stacks.

Exam trap

The trap here is confusing the collector/forwarder role with the storage/backend role; candidates who know Loki or Elasticsearch from dashboards may pick them, forgetting the question asks for a lightweight log shipper designed for Kubernetes.

How to eliminate wrong answers

Option A is wrong because Logstash is a JVM-based, heavyweight log pipeline from the Elastic stack — it is not Kubernetes-specific and consumes far more memory than Fluent Bit, so it is not the lightweight alternative. Option C is wrong because Loki is a log aggregation/storage backend (like Elasticsearch) from Grafana Labs, not a log forwarder/collector — it receives logs rather than being the shipper. Option D is wrong because Elasticsearch is a search and analytics store used as a log backend, not a collector, and it is not lightweight nor Kubernetes-specific.

25
MCQeasy

What is the primary purpose of Prometheus in cloud native observability?

A.Provide distributed tracing
B.Visualize data
C.Collect and store logs
D.Collect and store metrics
AnswerD

Prometheus scrapes time-series metrics from instrumented targets and stores them in its local TSDB, directly fulfilling observability's metrics pillar. This satisfies the stem's requirement for the primary purpose: pulling and retaining numeric measurements for querying and alerting, rather than handling logs or distributed traces.

Why this answer

Prometheus is an open-source monitoring and alerting toolkit designed specifically for collecting and storing time-series metrics. It scrapes metrics from instrumented targets via HTTP endpoints, stores them in a time-series database, and supports powerful querying with PromQL. Its primary purpose in cloud native observability is metrics collection and storage, not tracing, visualization, or log aggregation.

Exam trap

KCNA often tests the distinction between the three pillars of observability — metrics (Prometheus), logs (Loki/Elasticsearch), and traces (Jaeger) — and candidates may confuse Prometheus with visualization tools like Grafana.

How to eliminate wrong answers

Option A is wrong because distributed tracing is handled by tools like Jaeger, Zipkin, or OpenTelemetry, not Prometheus. Option B is wrong because visualization is typically done by Grafana or the Prometheus expression browser, not by Prometheus itself as its primary purpose. Option C is wrong because log collection and storage is the domain of tools like Fluentd, Loki, or Elasticsearch, not Prometheus.

26
MCQhard

An SRE team defines an SLO that 99.9% of requests to a service should complete in under 500ms over a 30-day rolling window. If the service receives 10 million requests in a month, what is the maximum number of requests that can exceed the latency threshold while still meeting the SLO?

A.10,000
B.5,000
C.1,000
D.100,000
AnswerA

A 99.9% SLO permits 0.1% of requests to breach the 500ms latency threshold. With 10 million monthly requests, that error budget equals 10,000,000 × 0.001 = 10,000 requests. This satisfies the 30-day rolling window constraint exactly, leaving no headroom beyond the defined tolerance.

Why this answer

An SLO of 99.9% means 0.1% of requests may violate the latency threshold. With 10,000,000 requests, 0.1% equals 10,000 requests. This is the error budget for the 30-day rolling window.

Exam trap

KCNA often tests SLO math — candidates miscalculate the error budget by misplacing the decimal (e.g., treating 99.9% as 0.01% instead of 0.1%).

How to eliminate wrong answers

Option B (5,000) is wrong because it corresponds to 99.95%, not 99.9%. Option C (1,000) is wrong because it corresponds to 99.99%, a stricter SLO. Option D (100,000) is wrong because it corresponds to 99%, a looser SLO.

Only 10,000 matches the 0.1% error budget for 99.9%.

27
Multi-Selectmedium

A team is implementing observability for a cloud native application. They want to adopt OpenTelemetry to instrument their code and collect telemetry data. Which TWO of the following are core components of the OpenTelemetry project? (Choose two.)

Select 2 answers
A.Fluentd
B.OpenTelemetry API
C.Jaeger Backend
D.Prometheus Server
E.OpenTelemetry Collector
AnswersB, E

The OpenTelemetry API provides the interfaces and data types for instrumenting code to generate telemetry. It defines how to create spans, metrics, and logs, and is implemented by SDKs. It is a fundamental building block, as it allows developers to instrument their applications in a vendor-neutral way. Without the API, there would be no standard way to emit telemetry.

Why this answer

OpenTelemetry consists of several core components, including the API, SDK, and Collector. The API defines the instrumentation interfaces, while the SDK provides the implementation. The Collector is a standalone service that receives, processes, and exports telemetry.

Together, these components enable vendor-neutral collection of metrics, logs, and traces. Prometheus, Jaeger, and Fluentd are separate projects that can integrate with OpenTelemetry but are not part of it.

Exam trap

The trap here is assuming that any popular observability tool is part of OpenTelemetry; instead, OpenTelemetry defines its own API, SDK, and Collector, while backends like Jaeger and Prometheus are external.

28
MCQhard

A platform team wants to implement observability for a Kubernetes cluster running 500+ microservices. They need to reduce the cost of storing logs while retaining the ability to search for specific error patterns. Which strategy best achieves this?

A.Increase log retention to one year for compliance
B.Store all logs in a centralized Elasticsearch cluster with high retention
C.Aggregate logs into a single pod for easier indexing
D.Use structured logging and sample debug logs, retaining error logs fully
AnswerD

Structured logging plus sampling debug output slashes stored log volume, directly cutting storage cost, while retaining error logs in full preserves searchability for specific error patterns. This balances the cost constraint against the required troubleshooting capability.

Why this answer

Structured logging (e.g., JSON format) enables efficient indexing and querying of logs, while sampling debug logs and retaining error logs fully reduces storage costs without losing critical error patterns. This approach balances observability needs with cost optimization, a key principle in cloud-native environments.

Exam trap

The trap here is that candidates may assume centralized storage (Elasticsearch) or longer retention always improves observability, ignoring the cost and scalability constraints of 500+ microservices in a cloud-native environment.

How to eliminate wrong answers

Option A is wrong because increasing log retention to one year for compliance does not address cost reduction; it increases storage costs and may violate data minimization principles. Option B is wrong because storing all logs in a centralized Elasticsearch cluster with high retention is expensive and inefficient, as it retains unnecessary debug logs and scales poorly for 500+ microservices. Option C is wrong because aggregating logs into a single pod creates a single point of failure, violates pod isolation, and does not reduce storage costs or improve searchability.

29
MCQmedium

A DevOps team wants to collect logs from all Kubernetes nodes and forward them to a central log storage system. Which tool is specifically designed for lightweight log aggregation and forwarding on Kubernetes nodes?

A.Elasticsearch
B.Prometheus
C.Fluent Bit
D.Grafana
AnswerC

Fluent Bit is a lightweight, low-footprint log processor and forwarder designed for constrained environments, typically deployed as a DaemonSet so one pod runs per node. It tails node and container logs and ships them to central storage, matching the aggregation requirement.

Why this answer

Fluent Bit is a lightweight, high-performance log processor and forwarder designed specifically for resource-constrained environments like Kubernetes nodes. It runs as a DaemonSet, tails container log files from /var/log/containers, and forwards them to a central backend such as Elasticsearch, Loki, or Splunk. Its low CPU and memory footprint make it the standard choice for node-level log aggregation.

Exam trap

The trap is conflating log storage/search tools (Elasticsearch, Kibana) with log collection/forwarding tools — the question asks specifically about the node-level agent, which is Fluent Bit.

How to eliminate wrong answers

Option A is wrong because Elasticsearch is a storage and search engine for logs, not a node-level collector/forwarder. Option B is wrong because Prometheus is a metrics monitoring system using a pull model, not a log aggregation tool. Option D is wrong because Grafana is a visualization and dashboarding layer that queries data sources — it does not collect or forward logs from nodes.

30
MCQmedium

A team wants to visualize metrics from Prometheus in a dashboard. Which tool is commonly used for this purpose?

A.Grafana
B.Alertmanager
C.Jaeger UI
D.Kibana
AnswerA

Grafana queries Prometheus directly through its native data source, rendering the time-series metrics as dashboards without altering the stored data. This satisfies the stem's requirement to visualise Prometheus metrics, since Grafana is purpose-built for dashboarding while Prometheus itself only offers basic expression-browser graphs.

Why this answer

Grafana is the de facto standard visualization tool for Prometheus metrics, connecting to Prometheus as a data source and rendering dashboards, graphs, and alerts. It supports PromQL queries directly and is widely deployed alongside Prometheus in Kubernetes environments.

Exam trap

The trap is confusing the observability stack roles — candidates may pick Kibana (visualization) without noting it pairs with Elasticsearch, not Prometheus.

How to eliminate wrong answers

Option B is wrong because Alertmanager handles routing and deduplication of alerts fired by Prometheus — it does not visualize metrics. Option C is wrong because Jaeger UI is a distributed tracing interface, not a metrics dashboard. Option D is wrong because Kibana visualizes Elasticsearch data (logs and documents), not Prometheus time-series metrics.

31
MCQhard

A team wants to implement cost monitoring for their Kubernetes clusters. Which approach is most effective?

A.Use cloud provider billing APIs combined with resource utilization data
B.Use kubectl top to get resource usage
C.Estimate costs based on node count
D.Monitor CPU and memory usage with Prometheus
AnswerA

Provider billing APIs expose actual spend per cluster and namespace, while utilisation data attributes that cost to workloads, revealing idle or over-provisioned resources. Correlating the two axes is what enables meaningful cost monitoring rather than raw invoice totals alone.

Why this answer

Cloud provider billing APIs provide actual cost data per resource (e.g., per node, per persistent volume, per network egress), and combining this with resource utilization data (e.g., CPU/memory requests and actual usage from metrics) enables accurate cost allocation per namespace, pod, or workload. This approach directly maps infrastructure spend to Kubernetes abstractions, which is essential for chargeback or showback in multi-tenant clusters.

Exam trap

The trap here is that candidates confuse resource monitoring (CPU/memory) with cost monitoring, assuming that tracking utilization alone (e.g., with Prometheus or kubectl top) is sufficient to understand spending, when in fact cost data requires explicit billing integration.

How to eliminate wrong answers

Option B is wrong because 'kubectl top' only shows current resource usage (CPU/memory) for nodes and pods, not cost data; it lacks any billing context or historical aggregation needed for cost monitoring. Option C is wrong because estimating costs based solely on node count ignores variable costs like storage, network egress, and managed services (e.g., load balancers), leading to inaccurate cost attribution. Option D is wrong because Prometheus monitors resource utilization metrics (CPU, memory, disk I/O) but does not inherently provide cost data; it would need to be combined with pricing information from cloud provider APIs to calculate costs.

32
MCQmedium

Which tool is specifically designed for distributed tracing and is a Cloud Native Computing Foundation (CNCF) graduated project?

A.Grafana
B.Fluentd
C.Jaeger
D.Prometheus
AnswerC

Jaeger provides distributed tracing by propagating context across service boundaries, letting you visualise request flows and pinpoint latency in microservices. It satisfies the CNCF graduated constraint, unlike Prometheus (metrics) or Fluentd (logs), making it the tool specifically designed for tracing in cloud-native environments.

Why this answer

Jaeger is a CNCF graduated project focused on distributed tracing.

33
MCQmedium

Which component of the OpenTelemetry architecture is responsible for receiving data from instrumented applications and processing it before export?

A.OpenTelemetry SDK
B.OpenTelemetry API
C.OpenTelemetry Collector
D.OpenTelemetry exporter
AnswerC

The OpenTelemetry Collector receives telemetry from instrumented applications via OTLP receivers, then processes it through configurable processors before exporting to backends. This satisfies the stem's requirement for a component that both receives and processes data, unlike SDKs, which only generate and emit telemetry from within the application process itself.

Why this answer

The OpenTelemetry Collector is a vendor-agnostic component that receives telemetry (traces, metrics, logs) from instrumented applications via receivers, processes it (batching, filtering, sampling, attribute enrichment), and exports it to one or more backends. It decouples applications from backend-specific exporters.

Exam trap

The trap is confusing the SDK (in-app instrumentation) with the Collector (out-of-process receiver/processor/exporter) — the question specifically asks about receiving and processing data from applications.

How to eliminate wrong answers

Option A is wrong because the SDK is the in-process library that instruments application code and generates telemetry — it does not receive and process data from other applications. Option B is wrong because the API defines the interfaces and no-op implementations used by instrumentation libraries; it does not process or export data. Option D is wrong because an exporter is only the final stage of the Collector (or SDK) pipeline that sends data to a backend — it does not receive or process data on its own.

34
Multi-Selecteasy

Which TWO of the following tools are commonly used for distributed tracing in cloud-native environments? (Select two.)

Select 2 answers
A.Zipkin
B.Grafana
C.Jaeger
D.Fluentd
E.Prometheus
AnswersA, C

Zipkin satisfies the distributed tracing requirement by collecting and correlating span data across microservices, using trace and span IDs propagated through request headers to reconstruct end-to-end request flows. Its lightweight, vendor-neutral design suits cloud-native architectures, where it commonly pairs with instrumentation libraries to visualise latency across service boundaries.

Why this answer

Zipkin (A) is correct because it is a dedicated distributed tracing system that collects and visualizes latency data across microservice call chains using trace and span IDs propagated via headers such as B3. Jaeger (C) is also correct because it is a CNCF-graduated distributed tracing platform that instruments requests across services, supports OpenTracing/OpenTelemetry, and provides trace storage and a UI for span analysis. Grafana (B) is primarily a visualization and dashboarding layer that can query tracing backends but is not itself a tracing system.

Fluentd (D) is a log collection and forwarding agent, and Prometheus (E) is a metrics monitoring and time-series database, neither of which performs distributed request tracing.

Exam trap

The trap is selecting Grafana or Prometheus because they appear in observability stacks — but tracing requires a dedicated tracing backend, and Grafana is only a frontend that can query one.

35
MCQhard

In PromQL, which function would you use to calculate the per-second rate of increase of a counter over a specified time window?

A.rate()
B.delta()
C.avg_over_time()
D.increase()
AnswerA

rate() computes the per-second average rate of increase of a counter over the given range window, automatically handling counter resets. It directly satisfies the stem's requirement to derive per-second increase from a monotonically increasing counter, unlike irate() or increase().

Why this answer

The rate() function calculates the per-second average rate of increase of a counter over a time range.

36
MCQhard

In OpenTelemetry, what is the purpose of the Collector component?

A.Instrument code automatically
B.Receive, process, and export telemetry data
C.Visualize traces and metrics
D.Aggregate logs from multiple sources
AnswerB

The Collector is a vendor-agnostic pipeline that receives telemetry via OTLP or other receivers, processes it through processors such as batching and filtering, then exports it to backends. This receive-process-export flow satisfies the stem's requirement for a central telemetry handling component.

Why this answer

The OpenTelemetry Collector is a vendor-agnostic agent or gateway that receives telemetry data (traces, metrics, logs) from instrumented applications, processes it (e.g., batching, filtering, sampling), and exports it to one or more backends (e.g., Jaeger, Prometheus, or any OTLP-compatible system). It decouples data generation from data export, enabling flexible pipeline management without modifying application code.

Exam trap

CNCF often tests the distinction between the Collector's role (data pipeline) and other components like SDKs (instrumentation) or backends (visualization/storage), so candidates mistakenly associate the Collector with auto-instrumentation or visualization.

How to eliminate wrong answers

Option A is wrong because automatic code instrumentation is the role of OpenTelemetry SDKs and auto-instrumentation agents (e.g., Java agent), not the Collector; the Collector does not instrument code. Option C is wrong because visualization of traces and metrics is the responsibility of backend tools like Jaeger UI, Grafana, or Prometheus, not the Collector, which only processes and forwards data. Option D is wrong because while the Collector can handle logs, its primary purpose is not limited to log aggregation; it is a unified pipeline for traces, metrics, and logs, and log aggregation alone is a narrower function often served by tools like Fluentd or Logstash.

37
MCQeasy

Which of the following is the correct definition of a Service Level Indicator (SLI)?

A.A formal contract between a service provider and a customer
B.A target value or range for a metric, agreed upon with stakeholders
C.A quantitative measure of a specific aspect of the service's reliability
D.A tool for aggregating logs from multiple sources
AnswerC

An SLI is the quantitative measurement itself — for example request latency or error rate — expressed as a number, which is then compared against an SLO target. It measures a specific reliability aspect of the service rather than defining a threshold or agreement.

Why this answer

An SLI is a quantitative measure of a specific aspect of a service's reliability — for example, request latency, error rate, or availability — expressed as a ratio of good events to total events. It is the raw measurement that feeds into SLOs and error budgets, not a contract or a target.

Exam trap

KCNA often tests confusion between SLI, SLO, and SLA, tricking candidates into selecting the SLO definition (target value) when the question asks for the SLI (the measurement itself).

How to eliminate wrong answers

Option A is wrong because a formal contract between provider and customer is a Service Level Agreement (SLA), which is a legal/business document, not a measurement. Option B is wrong because a target value or range agreed with stakeholders is a Service Level Objective (SLO), which is built on top of SLIs. Option D is wrong because log aggregation is a logging/observability function (e.g., Fluentd, Loki), not an SLI definition.

38
Multi-Selecthard

Which TWO of the following are recommended practices for achieving observability in a Kubernetes cluster?

Select 2 answers
A.Use a single centralized logging solution to aggregate logs from all components.
B.Store all debug logs for a minimum of 90 days for compliance.
C.Include correlation IDs in structured logs to enable tracing across services.
D.Disable leader election for monitoring components to reduce complexity.
E.Use Prometheus with a pull-based model to scrape metrics from pods.
AnswersC, E

Correlation IDs help trace requests across microservices.

Why this answer

Including correlation IDs in structured logs is a key observability practice that enables distributed tracing across microservices. In Kubernetes, where requests often traverse multiple pods and services, correlation IDs allow you to link logs from different components into a single transaction flow, which is essential for debugging and understanding system behavior.

Exam trap

CNCF often tests the misconception that centralized logging is always best, but the trap here is that observability emphasizes distributed, resilient data collection over a single monolithic log sink, and that debug logs are not subject to long-term compliance retention like audit logs.

39
MCQmedium

In distributed tracing, what is a 'span'?

A.A metric measuring request latency
B.A single logical operation within a trace
C.A collection of related traces
D.A log entry with trace context
AnswerB

A span represents one logical operation within a trace, carrying its own operation name, start and end timestamps, duration, and parent-child relationships. This granular unit lets you break a distributed request into discrete steps, satisfying the stem's requirement for the building block that composes a trace.

Why this answer

A span represents a single logical operation within a trace, such as an HTTP request, a database query, or a function call. It has a start time, duration, and metadata, and spans are nested to form the trace tree. This is the fundamental unit of work in distributed tracing.

Exam trap

KCNA often tests the distinction between a span and a trace, catching candidates who think a span is a collection of traces or a metric rather than a single operation.

How to eliminate wrong answers

Option A is wrong because a metric is a numeric measurement aggregated over time (e.g., latency percentiles), not a discrete operation with parent-child relationships. Option C is wrong because a collection of related traces is a trace or a trace group, not a span; spans are the components of a trace, not the other way around. Option D is wrong because a log entry with trace context is a log correlated to a trace, not a span itself; spans are structured timing records, not log messages.

40
MCQmedium

Which component is responsible for aggregating metrics from Kubernetes nodes and exposing them to the metrics API?

A.Prometheus Server
B.Grafana
C.metrics-server
D.Fluentd
AnswerC

metrics-server collects resource metrics from kubelets on each node, aggregates them, and serves them through the metrics API, which Horizontal Pod Autoscalers and kubectl top consume. It satisfies the requirement for node metric aggregation and API exposure.

Why this answer

The metrics-server is the correct component because it is specifically designed to collect resource metrics (CPU and memory) from the kubelet on each node via the Summary API and expose them through the Kubernetes Metrics API. This allows tools like `kubectl top` and the Horizontal Pod Autoscaler to access real-time resource usage without requiring a full monitoring stack.

Exam trap

The trap here is that candidates often confuse Prometheus (a full monitoring system) with the metrics-server (a lightweight, Kubernetes-native component for the Metrics API), assuming Prometheus is required for `kubectl top` or HPA when in fact the metrics-server is the dedicated and simpler solution.

How to eliminate wrong answers

Option A is wrong because Prometheus Server is a full monitoring and alerting system that scrapes metrics from various endpoints, but it is not the component responsible for aggregating metrics from nodes and exposing them to the Kubernetes Metrics API; it typically scrapes the metrics-server or kubelet directly. Option B is wrong because Grafana is a visualization and dashboarding tool that queries data sources like Prometheus or metrics-server, but it does not aggregate or expose metrics to the Metrics API. Option D is wrong because Fluentd is a log collector and forwarder used for log aggregation, not for collecting or exposing resource metrics to the Kubernetes Metrics API.

41
MCQeasy

Which of the following is NOT one of the three pillars of observability in cloud-native environments?

A.Metrics
B.Traces
C.Security
D.Logs
AnswerC

The three pillars are metrics, logs and traces, which together describe system state and behaviour. Security is a separate discipline applied across those signals, not a pillar itself, so it is the correct exclusion. This satisfies the question's NOT requirement.

Why this answer

The three pillars are logs, metrics, and traces. Security is not one of them, though it is important.

42
Multi-Selecthard

Which THREE of the following are important considerations when defining SLOs (Service Level Objectives)? (Select three.)

Select 3 answers
A.They should include an error budget
B.They must be aligned with business impact
C.They must be based on measurable SLIs
D.They define a target percentage over a time window
E.They should minimize infrastructure cost
AnswersB, C, D

An SLO only matters if it reflects what users and the business actually care about, so targets must trace to business impact rather than arbitrary infrastructure metrics. This alignment ensures reliability investment addresses genuine customer-facing risk instead of technical vanity measures.

Why this answer

Option B is correct because SLOs must be aligned with business impact: an SLO only makes sense if the reliability target reflects what users and the business actually need, so the chosen SLI and threshold are tied to a meaningful customer or revenue outcome rather than an arbitrary technical metric. Option C is correct because SLOs must be based on measurable SLIs: an SLO is a target applied to a Service Level Indicator, so the underlying measurement (for example, request latency or success rate) must be quantifiable and observable via monitoring. Option D is correct because SLOs define a target percentage over a time window: a valid SLO takes the form of a target such as 99.9% of successful requests measured over a rolling 30-day window, giving a concrete, time-bounded reliability goal.

Option A is not correct as a defining consideration because the error budget is derived from the SLO (100% minus the SLO target) rather than being a required component when defining the SLO itself. Option E is not correct because minimizing infrastructure cost is a cost-optimization concern, not a criterion for setting reliability objectives, and can even conflict with achieving the intended service level.

Exam trap

KCNA often tests confusion between SLOs and error budgets, tricking candidates into selecting 'include an error budget' as a defining consideration when error budgets are actually derived from SLOs.

43
MCQhard

A Kubernetes cluster runs Prometheus for monitoring. The operations team wants to receive alerts when the 99th percentile latency of an HTTP service exceeds 500ms for 5 minutes. They have a metric http_request_duration_seconds histogram. Which PromQL expression should they use to calculate the 99th percentile latency?

A.http_request_duration_seconds_bucket{le="0.5"} > 500
B.quantile(0.99, http_request_duration_seconds)
C.rate(http_request_duration_seconds_sum[5m]) / rate(http_request_duration_seconds_count[5m])
D.histogram_quantile(0.99, rate(http_request_duration_seconds_bucket[5m]))
AnswerD

The histogram_quantile function calculates quantiles from histogram buckets. Using rate on the bucket metric over a 5-minute window accounts for counter resets and provides per-second rate of increase. This is the standard way to compute latency percentiles from histograms. The 0.99 quantile gives the 99th percentile, and the 5-minute window aligns with the alerting requirement.

Why this answer

To compute the 99th percentile from a Prometheus histogram, the histogram_quantile function is used with the desired quantile and a rate over the bucket metric. The rate function calculates the per-second increase of the cumulative bucket counters over a 5-minute window, which is necessary for accurate quantile estimation. This expression correctly yields the 99th percentile latency, enabling the alerting condition to be evaluated.

Exam trap

The trap here is using the quantile function directly on a histogram metric instead of histogram_quantile on the _bucket series, which is required for histograms.

44
MCQmedium

In the context of distributed tracing, what is a 'span'?

A.A metric that measures request latency
B.A tool for collecting logs from containers
C.The entire end-to-end transaction across services
D.A single logical operation within a service, with a start and end time
AnswerD

A span represents one logical operation inside a service, bounded by start and end timestamps, and carries trace context linking it to parent and child spans. This satisfies the stem's distributed-tracing definition, distinct from a trace, which is the whole request path.

Why this answer

A span represents a single logical operation within a service, with a start time and end time, forming the basic unit of work in distributed tracing. Spans are nested to form a trace, where the root span represents the overall request and child spans represent downstream calls, database queries, or internal operations.

Exam trap

KCNA often tests the distinction between a span (one operation) and a trace (the full end-to-end transaction), so candidates who pick the 'entire transaction' option confuse the container with its component.

How to eliminate wrong answers

Option A is wrong because a metric that measures request latency is a numeric time-series measurement (e.g., Prometheus histogram), not a span — spans carry contextual metadata (operation name, tags, parent ID) and are the building blocks of traces, not metrics. Option B is wrong because a tool for collecting logs from containers describes log aggregation agents like Fluentd or Fluent Bit, which handle logs, not tracing spans. Option C is wrong because the entire end-to-end transaction across services is a trace, not a span — a trace is composed of multiple spans linked by parent-child relationships, so this option confuses the whole with the part.

45
MCQhard

What is context propagation in distributed tracing?

A.Sampling traces to reduce data volume
B.Visualizing traces in a user interface
C.Carrying trace context (trace ID, span ID) across services
D.Storing trace data in a centralized database
AnswerC

Context propagation is the mechanism that carries trace context, specifically the trace ID and span ID, across service boundaries. Each downstream service continues the same trace rather than starting a new one, which is what makes a distributed trace reconstructable end to end.

Why this answer

Context propagation is the mechanism by which trace context — primarily the trace ID and the current span ID — is passed from one service to the next across network calls, so all spans belonging to the same request are stitched into a single distributed trace. This is typically done via HTTP headers (e.g., W3C `traceparent`) or message metadata. Without it, each service would create an isolated trace and the end-to-end picture would be lost.

Exam trap

The trap is conflating the different pillars of tracing — sampling, storage, visualization, and propagation — and picking a plausible-sounding but functionally distinct option like sampling or storage.

How to eliminate wrong answers

Option A is wrong because sampling is a separate concern — it decides which traces to record to control volume, not how context moves between services. Option B is wrong because visualizing traces in a UI is the job of a tracing backend/frontend (e.g., Jaeger UI, Grafana Tempo), not propagation. Option D is wrong because storing trace data in a centralized database is the storage/backend layer of a tracing system, unrelated to carrying context across service boundaries.

46
MCQeasy

What is the primary purpose of structured logging?

A.To format logs in a consistent, machine-readable way for easier processing
B.To compress log files and reduce storage usage
C.To encrypt log data for security purposes
D.To send logs directly to the user's terminal
AnswerA

Structured logging emits events as key-value pairs, typically JSON, so fields like timestamp, level and service are parsed programmatically without regex. This consistent, machine-readable schema satisfies the stem's requirement for easier automated processing, filtering and aggregation across distributed systems.

Why this answer

Structured logging emits log entries as machine-readable data — typically JSON with consistent key-value fields (timestamp, level, service, message, request_id) — rather than free-form text. This makes logs trivially parseable by log processors and query engines, enabling reliable filtering, aggregation, and alerting. The primary purpose is consistent, machine-readable formatting for easier downstream processing.

Exam trap

The trap is confusing the format of logs (structured vs unstructured) with operational concerns like compression, encryption, or output destination, which are handled by separate tooling layers.

How to eliminate wrong answers

Option B is wrong because compression is a storage/transport optimization applied to log files or streams, not the purpose of structured logging — structured logs are often larger than plain text. Option C is wrong because encryption is a security control applied to log data at rest or in transit, orthogonal to log format. Option D is wrong because sending logs to a terminal is a display/output concern; structured logging is about the format of the emitted record, not where it is rendered.

47
MCQmedium

Which tool is primarily used for distributed tracing in cloud native environments?

A.Grafana
B.Fluentd
C.Jaeger
D.Prometheus
AnswerC

Jaeger is a CNCF distributed tracing system that records and visualises request flows across microservices, using spans and trace context propagation. This makes it the primary tool for diagnosing latency and dependency issues in cloud native environments.

Why this answer

Jaeger is an open-source, CNCF-graduated distributed tracing system originally built by Uber, designed to collect, store, and visualize traces across microservices. It implements the OpenTracing/OpenTelemetry data model and provides a UI for trace analysis, making it the canonical tracing tool in cloud native environments. Its architecture (agents, collectors, query service, storage backends) is purpose-built for distributed tracing.

Exam trap

The trap is mixing up the three observability pillars — metrics (Prometheus), logs (Fluentd), and traces (Jaeger) — and picking a tool from the wrong pillar because it is familiar from dashboards.

How to eliminate wrong answers

Option A is wrong because Grafana is a visualization/dashboarding platform — it can display tracing data (e.g., from Tempo or Jaeger) but is not itself a tracing system. Option B is wrong because Fluentd is a log collection and forwarding tool, part of the logging pipeline, not tracing. Option D is wrong because Prometheus is a metrics monitoring system with a time-series database and pull-based scraping — it handles metrics, not distributed traces.

48
Multi-Selectmedium

Which THREE of the following are benefits of structured logging? (Select three.)

Select 3 answers
A.Easier querying and filtering
B.More human-readable than plain text
C.Reduced storage requirements
D.Machine-parseable output
E.Consistent field names across services
AnswersA, D, E

Structured logs store fields as discrete key-value pairs rather than free text, so queries can filter on specific attributes without regex parsing. This directly satisfies the stem's benefit of easier querying and filtering across large log volumes.

Why this answer

Structured logging emits events as key-value pairs (typically JSON), so option A is correct because fields like level, service, or trace_id can be queried and filtered directly in tools such as Elasticsearch, Loki, or CloudWatch Logs Insights without fragile regex parsing. Option D is correct because that same machine-parseable format lets log processors, SIEMs, and observability pipelines ingest and index records programmatically and reliably. Option E is correct because adopting a shared schema with consistent field names across services enables correlation and aggregation across distributed systems, which is a core goal of structured logging.

Option B is not a benefit: structured logs are usually less human-readable than free-form plain text, which is why pretty-printers exist. Option C is not a benefit either: JSON key-value output typically increases log size versus terse plain-text lines due to repeated field names and delimiters, so storage requirements usually grow rather than shrink.

Exam trap

KCNA often tests whether candidates understand that structured logging trades human readability and storage efficiency for machine parseability and consistency — the trap is selecting 'more human-readable' or 'reduced storage' as benefits when they are actually drawbacks or unrelated.

49
MCQeasy

What is the purpose of Alertmanager in Prometheus?

A.Handle alert notifications
B.Visualize metrics
C.Store long-term metrics
D.Collect metrics from targets
AnswerA

Alertmanager receives alerts fired by Prometheus, then deduplicates, groups, and routes them to receivers such as email, Slack, or PagerDuty. It also handles silencing and inhibition, controlling notification delivery rather than evaluating alert rules or storing metrics.

Why this answer

Alertmanager is the component in the Prometheus ecosystem responsible for handling alerts fired by the Prometheus server. It deduplicates, groups, and routes alerts to configured notification channels such as email, PagerDuty, or Slack, ensuring that operators receive actionable notifications without alert fatigue.

Exam trap

The trap here is that candidates confuse Alertmanager with Prometheus itself, thinking it collects or stores metrics, when in fact it is solely a notification routing and deduplication engine.

How to eliminate wrong answers

Option B is wrong because visualizing metrics is the role of Grafana or the Prometheus expression browser, not Alertmanager. Option C is wrong because long-term metrics storage is handled by remote storage integrations (e.g., Thanos, Cortex) or the Prometheus TSDB itself, not Alertmanager. Option D is wrong because collecting metrics from targets is the function of the Prometheus server via its scrape mechanism, not Alertmanager.

50
Multi-Selecthard

Which THREE are responsibilities of the OpenTelemetry project? (Select three.)

Select 3 answers
A.Visualize telemetry data
B.Store long-term telemetry data
C.Provide instrumentation libraries
D.Define a standard for telemetry data
E.Provide a vendor-agnostic Collector
AnswersC, D, E

OpenTelemetry ships language-specific instrumentation libraries that generate telemetry for common frameworks without manual coding, satisfying the stem's requirement for project responsibilities. These libraries emit spans, metrics and logs via the OpenTelemetry API, complementing the specification and collector rather than replacing them.

Why this answer

Option C is correct because OpenTelemetry provides language-specific instrumentation libraries (for Java, Python, Go, JavaScript, .NET, etc.) and APIs/SDKs that generate traces, metrics, and logs. Option D is correct because OpenTelemetry defines vendor-neutral specifications and semantic conventions for telemetry data, including the OTLP wire protocol, so signals are portable across tools. Option E is correct because the OpenTelemetry Collector is a vendor-agnostic component that receives, processes, and exports telemetry via receivers, processors, and exporters.

Option A is not a responsibility of OpenTelemetry itself; visualization is handled by backends such as Jaeger, Prometheus, Grafana, or commercial APM UIs. Option B is also not a responsibility of OpenTelemetry; long-term storage is provided by backend observability platforms, not by the OpenTelemetry project.

Exam trap

KCNA often tests the misconception that OpenTelemetry stores or visualizes data, when it is strictly a collection, standardization, and transport layer.

51
Multi-Selectmedium

Which TWO of the following are best practices for implementing observability in a cloud-native environment?

Select 2 answers
A.Store all raw observability data indefinitely for forensic analysis
B.Use only metrics and avoid logs to reduce complexity
C.Add unique request IDs to logs for end-to-end tracing correlation
D.Randomly sample all traces and logs to reduce storage
E.Use structured logging (e.g., JSON format) for easier automated parsing
AnswersC, E

Unique request IDs let a single transaction be followed across distributed services, satisfying the stem's cloud-native observability requirement for end-to-end tracing. Because each request carries its own identifier, logs from separate pods and microservices can be correlated without relying on hostnames or timestamps, which are unreliable in ephemeral, dynamically scheduled workloads.

Why this answer

Option C is correct because injecting a unique request ID (correlation ID) into every log entry lets you stitch together events across distributed microservices, load balancers, and queues, enabling true end-to-end tracing correlation in a cloud-native environment where a single request fans out across many ephemeral components. Option E is correct because structured logging in JSON (or similar key-value formats) makes logs machine-readable, so log aggregators like Fluentd, Loki, or Elasticsearch can parse fields automatically, enabling reliable filtering, alerting, and correlation without brittle regex parsing of free-text lines. Option A is wrong because retaining all raw observability data indefinitely is costly and unsustainable; best practice is tiered retention with sampling, aggregation, and lifecycle policies.

Option B is wrong because metrics alone lack the contextual detail needed to diagnose root causes; logs, traces, and metrics are complementary pillars of observability. Option D is wrong because randomly sampling all traces and logs indiscriminately can discard critical error and latency data; sampling should be intelligent (e.g., tail-based, error-biased) rather than uniform across everything.

Exam trap

The KCNA exam often tests the misconception that 'more data is always better' (Option A) or that 'simplifying to one data type is efficient' (Option B), while the correct approach balances cost, performance, and diagnostic value through structured logging and correlation IDs.

52
MCQmedium

A platform team runs a Kubernetes cluster where the kubelet and container runtime expose metrics on each node. The team wants to collect node-level CPU and memory metrics into Prometheus without deploying a separate exporter on every node. Which component should they configure Prometheus to scrape?

A.The kubelet's built-in cAdvisor endpoint
B.kube-state-metrics
C.The Kubernetes API server's /metrics endpoint
D.The container runtime's CRI logging socket
AnswerA

The kubelet exposes cAdvisor metrics on its /metrics endpoint, which includes node-level and container-level CPU and memory usage. Prometheus can scrape this endpoint on each node without deploying an additional exporter. This directly meets the requirement because the kubelet is already present on every node and surfaces the needed resource metrics through its built-in cAdvisor integration.

Why this answer

The kubelet already runs on every node and exposes cAdvisor metrics, including node and container CPU and memory usage, through its /metrics endpoint. Prometheus can scrape this endpoint directly, so no separate per-node exporter is required. Other listed components either report object state, control-plane request metrics, or logs, none of which provide the node resource metrics needed here.

Exam trap

The trap here is assuming that kube-state-metrics reports resource usage, when it actually reports the state of Kubernetes API objects rather than node or container CPU and memory consumption.

53
MCQeasy

A developer wants to monitor the health of a Kubernetes deployment by checking if the number of ready replicas matches the desired replicas. Which metric from kube-state-metrics should they query?

A.kube_deployment_status_replicas_ready
B.kube_deployment_spec_replicas
C.kube_node_status_condition
D.kube_pod_container_status_running
AnswerA

`kube_deployment_status_replicas_ready` exposes the count of ready replicas directly from the Deployment's status, satisfying the requirement to compare ready against desired replicas. Pairing it with `kube_deployment_spec_replicas` gives the desired count, enabling an alert when readiness diverges from the specification.

Why this answer

`kube_deployment_status_replicas_ready` directly exposes the number of ready replicas for a Deployment, which can be compared against `kube_deployment_spec_replicas` to determine if the desired state matches the actual healthy state. This metric is emitted by kube-state-metrics, which generates Prometheus-compatible metrics from Kubernetes API objects, making it the standard choice for monitoring Deployment health.

Exam trap

The trap here is that candidates might confuse metrics that show pod state (like `kube_pod_container_status_running`) with Deployment-level readiness, not realizing that a pod can be running but not ready, and that the correct metric must reflect the Deployment's own status field.

How to eliminate wrong answers

Option B is wrong because `kube_deployment_spec_replicas` only shows the desired number of replicas as defined in the Deployment spec, not the actual ready count, so it cannot alone indicate health. Option C is wrong because `kube_node_status_condition` tracks node-level conditions (e.g., Ready, DiskPressure) and has no relation to Deployment replica health. Option D is wrong because `kube_pod_container_status_running` counts containers in Running state, not ready replicas of a Deployment, and does not account for readiness probes or desired replica counts.

54
MCQmedium

A platform team runs a Kubernetes cluster with the OpenTelemetry Collector deployed as a DaemonSet. They want to collect node-level metrics such as CPU and memory usage from every node without modifying application code. Which receiver should they configure in the Collector's pipeline?

A.otlp receiver
B.prometheus receiver
C.hostmetrics receiver
D.jaeger receiver
AnswerC

The hostmetrics receiver is designed to collect host-level metrics such as CPU, memory, disk, and network usage from the node where the Collector runs. Since the Collector is deployed as a DaemonSet, one instance runs on each node, allowing hostmetrics to gather node-level metrics without any application instrumentation or code changes, exactly matching the requirement.

Why this answer

The hostmetrics receiver is purpose-built for collecting host-level metrics such as CPU, memory, disk, and network usage. When the OpenTelemetry Collector is deployed as a DaemonSet, each node runs a Collector instance, so hostmetrics can gather metrics from every node without application changes. Other receivers either require additional exporters, handle different telemetry types, or do not actively scrape node metrics.

Exam trap

The trap here is assuming that any receiver can collect node metrics, when only hostmetrics is designed for that specific purpose.

55
MCQmedium

A team wants to ensure that at least 99.9% of all requests to their application complete within 500ms over a 30-day window. How should this requirement be classified?

A.Service Level Agreement (SLA)
B.Service Level Objective (SLO)
C.Service Level Indicator (SLI)
D.Key Performance Indicator (KPI)
AnswerB

An SLO is a measurable reliability target, such as 99.9% of requests completing within 500ms over 30 days, that a service aims to achieve. This satisfies the stem's latency and availability threshold, whereas an SLA is a contractual commitment and an SLI is the raw measurement.

Why this answer

An SLO is a target value or range for a service level measured by an SLI — here, '99.9% of requests under 500ms over 30 days' is exactly a target threshold on a measured metric. It is the internal reliability goal that teams commit to and track. An SLA, by contrast, is a contractual agreement with consequences (credits, penalties) if the SLO is missed.

Exam trap

The trap here is conflating SLO with SLA — candidates often pick SLA because it sounds like the formal requirement, but the question describes an internal measurable target, which is an SLO.

How to eliminate wrong answers

Option A is wrong because an SLA is a formal, often contractual commitment between a provider and customer that specifies consequences for missing targets — the question describes an internal target, not a contract. Option C is wrong because an SLI is the actual measured value (e.g., the percentage of fast requests), not the target threshold itself. Option D is wrong because a KPI is a broader business or operational metric used to gauge success, not the specific reliability target tied to an SLI.

56
MCQmedium

Which of the following is true about Prometheus's pull-based model for collecting metrics?

A.Targets push metrics to Prometheus
B.Prometheus only collects metrics from Kubernetes API server
C.Prometheus scrapes metrics from HTTP endpoints
D.Prometheus stores metrics in a relational database
AnswerC

Prometheus actively initiates scrapes against configured HTTP endpoints, retrieving metrics at fixed intervals. This pull architecture satisfies the stem's requirement by having the server discover and fetch data itself, rather than passively receiving pushed measurements. Targets expose a `/metrics` path that Prometheus polls, giving it control over collection timing and target health.

Why this answer

Prometheus uses a pull-based model where it periodically scrapes metrics from HTTP endpoints exposed by targets, typically at /metrics. This means targets do not push data; instead, Prometheus initiates the HTTP GET requests on a schedule defined by scrape_interval. This design simplifies target discovery and allows Prometheus to control scrape frequency and timeout.

Exam trap

KCNA often tests the misconception that Prometheus receives pushed metrics; candidates confuse the pull model with push-based systems like StatsD or Graphite.

How to eliminate wrong answers

Option A is wrong because targets do not push metrics to Prometheus in the standard pull model; push is only used via the Pushgateway for short-lived jobs. Option B is wrong because Prometheus can scrape any HTTP endpoint, not just the Kubernetes API server; it commonly scrapes node exporters, kubelets, and application pods. Option D is wrong because Prometheus stores metrics in its own time-series database (TSDB), not a relational database.

57
MCQhard

A Prometheus alert rule fires when the error rate exceeds 5% for 5 minutes. The alert is sent to Alertmanager. What must be configured in Alertmanager to ensure the alert is deduplicated, grouped, and routed to the correct team?

A.An inhibition rule
B.A recording rule
C.A silence rule
D.A route configuration
AnswerD

Routes in Alertmanager match incoming alerts by label and direct them to the appropriate receiver, while grouping and deduplication happen within that routing tree. Configuring a route therefore satisfies the requirement to deduplicate, group and deliver the alert to the correct team.

Why this answer

A route configuration in Alertmanager defines how alerts are grouped, deduplicated, and routed to specific receivers (e.g., email, Slack, PagerDuty) based on label matchers. The route tree determines which team receives which alerts, and grouping parameters control deduplication and batching. Without a route, Alertmanager cannot direct alerts to the correct team.

Exam trap

KCNA often tests the confusion between Alertmanager routing and Prometheus recording rules; candidates may pick recording rule thinking it handles alert delivery.

How to eliminate wrong answers

Option A is wrong because an inhibition rule suppresses alerts when a related alert is already firing; it does not handle routing or deduplication. Option B is wrong because a recording rule is a Prometheus feature that precomputes frequently used queries, not an Alertmanager configuration. Option C is wrong because a silence rule mutes alerts for a specified time window; it does not route or deduplicate them.

58
MCQhard

A platform team runs a 12-node Kubernetes cluster where each node hosts roughly 30 pods. They deployed Prometheus with a ServiceMonitor that scrapes every pod's /metrics endpoint every 15 seconds, but now the Prometheus pod is frequently OOMKilled and scrape targets intermittently report 'context deadline exceeded'. Which change best addresses the root cause while preserving observability?

A.Increase the Prometheus container's memory limit and CPU request so it can hold all active time series in memory.
B.Shorten the scrape interval to 5 seconds so each scrape collects fewer samples and finishes before the deadline.
C.Enable the Prometheus remote write receiver and forward all samples to a long-term storage backend such as Thanos or Cortex.
D.Reduce scrape scope using relabeling and metricRelabelings to drop unused targets and high-cardinality metrics, and split scraping across multiple Prometheus shards.
AnswerD

The cluster-wide scrape of every pod creates excessive active time series and concurrent target load, which drives memory usage and scrape deadline errors. Dropping unneeded targets and metrics through relabeling lowers cardinality, while sharding distributes scrape work across multiple Prometheus instances so no single server is overwhelmed, preserving observability for the metrics that matter.

Why this answer

The symptoms point to scrape and cardinality overload: too many targets and time series for a single Prometheus instance. Trimming what is scraped through relabeling and metricRelabelings reduces active series, and sharding spreads the remaining scrape load across multiple servers. Together these address both the memory exhaustion and the scrape deadline errors without abandoning observability of essential metrics.

Exam trap

The trap here is assuming that more memory, a shorter scrape interval, or remote write storage solves Prometheus overload, when the actual driver is uncontrolled cardinality and target count.

59
MCQmedium

A developer is investigating a performance issue in a microservices application. They want to trace a single request as it flows through multiple services, including a database call. Which OpenTelemetry concept allows them to correlate all these operations under a single logical unit?

A.Trace
B.Span
C.Metric
D.Log record
AnswerA

A trace represents the entire journey of a request through a distributed system, composed of multiple spans that share a common trace ID. It allows developers to see the full path and timing of a request across services. This is exactly what is needed to correlate operations from different services and a database call under a single logical unit.

Why this answer

In OpenTelemetry, a trace is the overarching structure that ties together all spans associated with a single request. Each span represents a unit of work, and spans within the same trace share a trace ID, enabling correlation across service boundaries. This allows developers to visualize the entire request flow, including database calls, and identify bottlenecks or failures.

Metrics and logs serve different purposes and do not provide this end-to-end request context.

Exam trap

The trap here is confusing a span with a trace; a span is a single operation, while a trace is the complete end-to-end journey of a request.

60
MCQeasy

A platform team is deploying a logging stack in Kubernetes. They want to collect logs from all pods and nodes, store them centrally, and provide a query interface. Which combination of tools is commonly used to achieve this?

A.Prometheus for collection, Grafana for storage, Loki for visualization
B.OpenTelemetry Collector for collection, Jaeger for storage, Kibana for visualization
C.Jaeger for collection, Prometheus for storage, Grafana for visualization
D.Fluent Bit for collection, Elasticsearch for storage, Kibana for visualization
AnswerD

Fluent Bit is a lightweight log processor and forwarder commonly used as a DaemonSet in Kubernetes to collect logs from nodes and pods. Elasticsearch provides scalable storage and search capabilities, while Kibana offers visualization and querying. This combination, often called the EFK stack, is a standard solution for centralized logging in cloud native environments, directly addressing the requirements.

Why this answer

The EFK stack—Elasticsearch, Fluent Bit (or Fluentd), and Kibana—is a widely adopted solution for Kubernetes logging. Fluent Bit runs as a DaemonSet to collect logs from each node, Elasticsearch indexes and stores them, and Kibana provides a UI for searching and visualizing. This architecture meets the needs of collecting, storing, and querying logs centrally, making it the correct choice among the options.

Exam trap

The trap here is mixing components from different observability pillars; for example, using Prometheus (metrics) for logs or Jaeger (tracing) for storage.

61
MCQhard

A company uses OpenTelemetry to instrument their microservices. They want to ensure that traces from one service can be correlated with those from another service across network calls. Which OpenTelemetry concept enables this correlation?

A.Exporter configuration
B.Span attributes
C.Context propagation
D.Sampling
AnswerC

Context propagation transmits trace context (trace ID, span ID, sampling flags) across service boundaries via protocols such as W3C Trace Context headers. This satisfies the stem's requirement to correlate spans from separate microservices across network calls, letting each downstream service join the same distributed trace rather than starting an unrelated one.

Why this answer

Context propagation is the OpenTelemetry mechanism that carries trace context (trace ID, span ID, sampling flags) across service boundaries via headers such as W3C Trace Context. This allows spans from different services to be stitched into a single distributed trace.

Exam trap

The trap is confusing span attributes (metadata on a span) with context propagation (the mechanism that links spans across services), leading candidates to pick attributes for correlation.

How to eliminate wrong answers

Option A is wrong because exporter configuration determines where telemetry data is sent (e.g., OTLP endpoint), not how traces are correlated. Option B is wrong because span attributes are key-value metadata attached to a span; they enrich a span but do not link spans across services. Option D is wrong because sampling decides which traces are recorded, not how they are correlated.

62
MCQeasy

What does the 'kubectl logs' command retrieve?

A.Audit logs
B.Cluster events
C.Container logs
D.Node logs
AnswerC

`kubectl logs` streams stdout and stderr captured by the container runtime for a specific container in a pod, satisfying the stem's request for container-level output. It reads from the node's log files rather than cluster events or API audit records, so it returns application output only.

Why this answer

The 'kubectl logs' command retrieves the standard output (stdout) and standard error (stderr) streams from a container running in a pod. It directly accesses the container's log files on the node (e.g., /var/log/pods/...) or streams them from the container runtime. This is the primary way to debug application issues in Kubernetes.

Exam trap

KCNA often tests the distinction between different types of logs and events, and candidates may confuse 'kubectl logs' with 'kubectl get events' or assume it retrieves node-level logs.

How to eliminate wrong answers

Option A is wrong because audit logs are records of API server requests, typically stored in files or sent to a backend, and are not retrieved via 'kubectl logs'. Option B is wrong because cluster events are objects in the Kubernetes API (viewed with 'kubectl get events') that report state changes, not container output. Option D is wrong because node logs (e.g., kubelet or system logs) are not accessible through 'kubectl logs'; they require direct node access or a logging agent.

63
MCQeasy

Which tool is commonly used for log aggregation in Kubernetes and is designed to be lightweight?

A.Fluent Bit
B.Jaeger
C.Prometheus
D.Grafana
AnswerA

Fluent Bit is a CNCF log processor and forwarder written in C with a small memory and CPU footprint, designed for constrained environments. It collects container logs and ships them to backends, satisfying the lightweight log aggregation requirement in Kubernetes.

Why this answer

Fluent Bit is a CNCF-graduated, lightweight log processor and forwarder written in C with a tiny memory footprint (often under 1 MB), specifically designed for high-performance log aggregation in containerized and edge environments. It is commonly deployed as a DaemonSet in Kubernetes to collect container logs from /var/log/containers and ship them to backends like Elasticsearch, Loki, or CloudWatch. Its low resource usage distinguishes it from heavier aggregators like Fluentd.

Exam trap

KCNA often tests the observability triad confusion — candidates see 'logs' and pick Fluentd or Grafana, or see 'lightweight' and pick Prometheus, forgetting that Fluent Bit is the purpose-built lightweight log forwarder.

How to eliminate wrong answers

Option B is wrong because Jaeger is a distributed tracing system for capturing and visualizing request flows across microservices, not a log aggregation tool. Option C is wrong because Prometheus is a time-series metrics monitoring and alerting system that scrapes numeric metrics, not a log collector. Option D is wrong because Grafana is a visualization and dashboarding layer that queries data sources like Prometheus or Loki; it does not itself aggregate logs.

64
Multi-Selectmedium

Which TWO are pillars of observability? (Select two.)

Select 2 answers
A.SLIs
B.Alerting
C.Logs
D.Metrics
E.Dashboards
AnswersC, D

Logs are one of observability's pillars, recording discrete timestamped events that reveal what happened within a system. This satisfies the stem's requirement by providing the detailed, high-cardinality event data that metrics and traces alone cannot capture for debugging specific failures.

Why this answer

Logs and Metrics are two of the three pillars of observability (alongside Traces). Logs provide immutable, timestamped records of discrete events, while Metrics are numeric aggregations of data over time (e.g., Prometheus counters, histograms). Together they form the foundation for understanding system behavior in cloud-native environments.

Exam trap

CNCF often tests the distinction between the pillars of observability (Logs, Metrics, Traces) and the tools or outputs derived from them (e.g., SLIs, Alerting, Dashboards), leading candidates to confuse operational practices with foundational data types.

65
MCQmedium

Which open-source project provides a unified standard for collecting and exporting telemetry data (metrics, logs, and traces) from applications?

A.Prometheus
B.OpenTelemetry
C.Jaeger
D.Fluentd
AnswerB

OpenTelemetry is the CNCF project supplying vendor-neutral APIs, SDKs and specifications for metrics, logs and traces together, so it satisfies the stem's requirement for one unified standard covering all three telemetry signals rather than separate per-signal agents.

Why this answer

OpenTelemetry (OTel) is a CNCF project that provides a vendor-neutral, unified standard — APIs, SDKs, and the Collector — for instrumenting applications to emit metrics, logs, and traces. It solves the fragmentation problem where each backend required its own agent and format. OTel is the answer when the question asks for a unified telemetry collection and export standard.

Exam trap

KCNA often tests the difference between a telemetry standard and a backend — candidates pick Prometheus or Jaeger because they recognize the names, missing that the question asks for the unified collection/export standard, which is OpenTelemetry.

How to eliminate wrong answers

Option A is wrong because Prometheus is a metrics-only monitoring system and time-series database; it does not unify logs and traces. Option C is wrong because Jaeger is a distributed tracing backend focused solely on traces, not a unified telemetry standard. Option D is wrong because Fluentd is a log collection and forwarding tool; it handles logs only and is not a cross-signal standard.

66
Multi-Selectmedium

Which THREE of the following are valid use cases for distributed tracing in a microservices architecture?

Select 3 answers
A.Monitoring CPU and memory usage of each service instance
B.Understanding the dependency graph between microservices
C.Pinpointing the root cause of an error in a distributed transaction
D.Identifying which service contributes the most latency to an end-user request
E.Capturing detailed error messages and stack traces
AnswersB, C, D

Distributed tracing records spans across service calls, letting teams reconstruct which services invoke which, including latency and error propagation. That reconstructed call graph directly reveals the dependency graph between microservices, a core tracing use case.

Why this answer

Distributed tracing is designed to follow a single request as it propagates across service boundaries, so option B is correct because the spans and parent-child relationships recorded in a trace directly reveal the dependency graph and call topology between microservices. Option C is correct because a trace correlates spans across services with shared trace IDs, letting you locate the exact failing span and thus the root cause of an error in a distributed transaction. Option D is correct because span durations and timestamps expose per-service latency contributions, making it possible to identify which service adds the most latency to an end-user request.

Option A is not a tracing use case; CPU and memory usage per instance are collected by metrics/monitoring tools such as Prometheus or CloudWatch, not by distributed tracing. Option E is also not specific to tracing; detailed error messages and stack traces are typically captured by centralized logging or error-tracking systems, even though traces may carry limited error tags.

Exam trap

The KCNA exam often tests the distinction between observability pillars (metrics, logs, traces) and expects candidates to recognize that distributed tracing is not a catch-all for monitoring or logging tasks, so the trap is confusing request-level tracing with infrastructure metrics or detailed error logging.

67
MCQeasy

What does SLA stand for in the context of service reliability?

A.Service Level Agreement
B.Service Level Indicator
C.Service Level Availability
D.Service Level Objective
AnswerA

SLA denotes Service Level Agreement, a documented commitment defining expected service levels such as availability, latency and error budgets. In service reliability it formalises the target against which reliability is measured and error budgets are derived.

Why this answer

SLA stands for Service Level Agreement, a contract specifying expected service level.

68
Multi-Selecthard

Which TWO of the following are best practices for structuring log output in cloud-native applications to maximize observability?

Select 2 answers
A.Include verbose debug-level information in every log line
B.Use multi-line log entries for detailed error information
C.Output logs in structured format such as JSON
D.Include a unique request or correlation ID in each log entry
E.Avoid timestamps to reduce log size
AnswersC, D

JSON output gives each log entry discrete, machine-parseable fields rather than an opaque string, so aggregators such as Loki, Elasticsearch or Cloud Logging can index and query individual attributes. This satisfies the observability requirement by enabling filtering, correlation and alerting without brittle regex parsing.

Why this answer

Option C is correct because emitting logs as structured data such as JSON lets log processors and observability backends parse fields programmatically, enabling reliable filtering, aggregation, and indexing by attributes like severity, service, and trace context rather than relying on fragile regex over free text. Option D is correct because including a unique request or correlation ID (for example a W3C traceparent trace ID or an application-generated correlation ID) in every entry allows distributed traces and logs to be stitched together across microservices, which is essential for root-cause analysis in cloud-native systems. The remaining options are not best practices: A floods storage and cost with low-value debug noise and can obscure real signals, B breaks line-oriented log collectors and parsers that expect one event per line, and E removes timestamps that are required to order events and correlate them across services and time zones.

Exam trap

CNCF often tests the misconception that 'more detail is better' (Option A) or that 'human readability' (Option B) is the priority, when in cloud-native observability, machine-parseable, single-line structured logs are the standard for scalability and automation.

69
MCQmedium

A developer wants to view the logs of a specific container named 'sidecar' inside a pod named 'app-pod'. Which command should they use?

A.kubectl log app-pod --container sidecar
B.kubectl logs app-pod sidecar
C.kubectl logs -c sidecar app-pod
D.kubectl logs app-pod -c sidecar
AnswerD

The -c flag selects a specific container within a multi-container pod, so kubectl logs app-pod -c sidecar returns only the sidecar container's output. Without it, kubectl logs on a multi-container pod fails or prompts for a container name, so this syntax satisfies the requirement precisely.

Why this answer

The -c flag specifies the container name. The correct command is 'kubectl logs app-pod -c sidecar'.

70
MCQhard

A platform engineer is configuring an OpenTelemetry Collector pipeline. They want to ensure that sensitive data such as credit card numbers is not exported to the observability backend. Which component of the Collector should they use to modify or drop attributes before export?

A.A processor in the pipeline
B.An exporter in the pipeline
C.A receiver in the pipeline
D.A connector in the pipeline
AnswerA

Processors in the OpenTelemetry Collector operate on telemetry data between receivers and exporters. They can modify, filter, or drop attributes. For example, the attributes processor can delete or hash sensitive fields like credit card numbers. This makes processors the correct component to sanitize data before it leaves the cluster, ensuring compliance and privacy.

Why this answer

Processors are the components in an OpenTelemetry Collector pipeline that transform telemetry data. They can add, remove, or modify attributes, including dropping sensitive information. By placing a processor such as the attributes or filter processor before the exporter, the engineer can ensure that credit card numbers are removed or obfuscated before data is sent to the backend.

Receivers, exporters, and connectors do not serve this purpose.

Exam trap

The trap here is assuming that any component in the pipeline can modify data, when only processors are designed for that transformation step.

71
MCQmedium

A company is running a microservices application on a Kubernetes cluster. They have noticed that one of the services, 'payment-api', is experiencing intermittent high latency. The team wants to identify the root cause without modifying the application code. Which approach should they take?

A.Monitor CPU and memory metrics from kube-state-metrics and correlate with latency.
B.Increase log verbosity for all services and search for error messages.
C.Implement distributed tracing using tools like Jaeger or Zipkin to trace requests across services.
D.Check node-level metrics using Prometheus Node Exporter.
AnswerC

Distributed tracing tracks request flow and identifies slow components.

Why this answer

Distributed tracing with tools like Jaeger or Zipkin allows you to follow a single request as it traverses multiple microservices, identifying exactly which service or call introduces latency. This approach does not require code changes (if the service mesh or sidecar proxy handles instrumentation) and is specifically designed to pinpoint performance bottlenecks in distributed systems, unlike CPU/memory metrics or log analysis which cannot trace a request's end-to-end path.

Exam trap

CNCF often tests the distinction between observability tools that provide request-level context (distributed tracing) versus aggregate resource metrics (kube-state-metrics, Node Exporter) or unstructured logs, leading candidates to mistakenly choose CPU/memory correlation or log analysis for pinpointing intermittent latency in a microservices architecture.

How to eliminate wrong answers

Option A is wrong because kube-state-metrics provides resource utilization data (CPU, memory) per pod or container, but high latency in a microservice is often caused by network delays, database contention, or upstream service failures—not necessarily correlated with local resource usage; correlation does not imply causation and cannot trace the request path. Option B is wrong because increasing log verbosity for all services generates massive volumes of unstructured data and relies on error messages that may not appear during intermittent latency spikes; logs lack the context of a specific request's journey across services, making root cause identification inefficient and often impossible. Option D is wrong because node-level metrics from Prometheus Node Exporter only show host-level resource usage (e.g., disk I/O, network bandwidth) and cannot reveal which microservice or request is causing latency within the cluster; they are useful for infrastructure troubleshooting but not for application-level distributed tracing.

72
Multi-Selecthard

Which THREE of the following are components of the OpenTelemetry project? (Select three)

Select 3 answers
A.OpenTelemetry Agent
B.OpenTelemetry API
C.OpenTelemetry SDK
D.OpenTelemetry Collector
E.OpenTelemetry Exporter
AnswersB, C, D

The OpenTelemetry API defines vendor-neutral interfaces for instrumenting code to emit traces, metrics and logs, decoupling telemetry generation from any backend. It satisfies the stem's requirement for a genuine project component, sitting alongside the SDK and Collector as one of OpenTelemetry's core building blocks.

Why this answer

The OpenTelemetry API (option B) is a core component that defines the interfaces and instrumentation primitives (Tracer, Meter, Logger) that application code uses to generate telemetry, independent of any implementation. The OpenTelemetry SDK (option C) is the concrete implementation of that API, providing the processing pipeline (span processors, samplers, exporters, resource detection) that turns API calls into exportable telemetry. The OpenTelemetry Collector (option D) is a standalone, vendor-agnostic service that receives, processes, and exports telemetry data via receivers, processors, and exporters, and is a distinct official component of the project.

Option A is incorrect because there is no product called the 'OpenTelemetry Agent'; the project ships an SDK and the Collector, not an agent by that name. Option E is incorrect because an exporter is a subcomponent within the SDK or Collector pipeline, not a top-level OpenTelemetry project component.

Exam trap

KCNA often tests whether candidates confuse the OpenTelemetry Collector with a non-existent 'Agent' component, or mistake pluggable exporters for a core project pillar.

73
MCQmedium

A team runs a Kubernetes cluster with Prometheus scraping application metrics. They want Prometheus to automatically discover new pods and scrape their metrics endpoints as pods are created and destroyed. Which Kubernetes resource should they configure Prometheus to use for this dynamic discovery?

A.The kubelet's embedded cAdvisor endpoint on each node
B.A static list of pod IPs in the Prometheus configuration file
C.Kubernetes API server via Prometheus's kubernetes_sd_configs
D.A ConfigMap that lists all pod names and namespaces
AnswerC

Prometheus's kubernetes_sd_configs allows it to query the Kubernetes API server to discover targets such as pods, services, and endpoints. This enables automatic scraping of new pods as they appear, without manual configuration. It is the standard method for dynamic service discovery in Kubernetes environments, making it the correct choice for this scenario.

Why this answer

Prometheus can dynamically discover targets in Kubernetes by integrating with the Kubernetes API server through kubernetes_sd_configs. This allows it to watch for pod creation and deletion events and automatically add or remove scrape targets. Static configurations or ConfigMaps require manual intervention and do not adapt to the dynamic nature of Kubernetes workloads, so they are unsuitable for this use case.

Exam trap

The trap here is assuming that any Kubernetes resource can be watched by Prometheus for discovery, when only specific service discovery mechanisms like kubernetes_sd_configs are designed for that purpose.

74
Multi-Selecthard

Which TWO of the following are examples of context propagation mechanisms used in distributed tracing?

Select 2 answers
A.HTTP headers
B.Environment variables
C.Database queries
D.Shared filesystem
E.gRPC metadata
AnswersA, E

Headers like traceparent are used to propagate trace context across HTTP calls.

Why this answer

HTTP headers, such as the `traceparent` and `tracestate` headers defined in the W3C Trace Context specification, are the standard mechanism for propagating trace context across service boundaries in distributed tracing. When a service receives an incoming HTTP request, it extracts the trace ID and span ID from these headers to continue the same trace. This allows trace data to be correlated across multiple microservices as the request flows through the system.

Exam trap

CNCF often tests the distinction between static configuration mechanisms (like environment variables or shared filesystems) and dynamic, in-band propagation mechanisms (like HTTP headers and gRPC metadata) that travel with each request.

75
Multi-Selectmedium

Which TWO of the following are valid Prometheus metric types? (Select two)

Select 2 answers
A.Set
B.Counter
C.Timer
D.Meter
E.Gauge
AnswersB, E

A Counter is a cumulative Prometheus metric that only increases or resets to zero on restart, suitable for totals such as requests served. PromQL's rate() and increase() functions operate on counters, distinguishing them from gauges, histograms, and summaries.

Why this answer

Option B (Counter) is correct because Prometheus defines Counter as a core metric type: a cumulative value that only increases (or resets to zero on restart), typically used for counts like total requests or errors, and queried with rate()/increase(). Option E (Gauge) is also correct because Prometheus defines Gauge as a metric type representing a value that can arbitrarily go up and down, such as temperature, memory usage, or current queue depth. The other options are not Prometheus metric types: Set, Timer, and Meter are metric abstractions found in other monitoring libraries/systems (for example, Dropwizard Metrics or Micrometer), not in the Prometheus client data model, whose four types are Counter, Gauge, Histogram, and Summary.

Exam trap

KCNA often tests whether candidates confuse Prometheus metric types with StatsD (Set, Timer) or OpenTelemetry (Meter) terminology, so memorizing the exact four Prometheus types is essential.

Page 1 of 2 · 83 questions totalNext →

Ready to test yourself?

Try a timed practice session using only Cloud Native Observability questions.