Courseiva
easyMultiple Choice

PDE Practice Question: Based on the exhibit, what is the most likely…

Exhibit

Refer to the exhibit.

```
# Dataflow pipeline error log:
Workflow failed. Causes: S02:ReadPubSub/Read+Transform/ParDo(ExtractTimestamps)+ ... (4b9c3d2e)
The job failed because a worker experienced a "out of memory" error.
```

Pipeline configuration:
- Streaming engine: disabled
- Worker machine type: n1-standard-4 (4 vCPU, 15 GB memory)
- Number of workers: 2 (autoscaling enabled, max 10)
- Input: Pub/Sub topic with 1000 messages/sec, each message ~50 KB
- Transform: Parse JSON, enrich with external API call, window into 1-minute fixed windows, write to BigQuery

Based on the exhibit, what is the most likely cause of the out-of-memory error?

⚠ Common exam trap

Google Cloud often tests the misconception that OOM errors are caused by schema mismatches or Pub/Sub backlogs, but the real cause is almost always insufficient worker memory for the data volume.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

The worker machine type has insufficient memory for the message size and throughput.

The out-of-memory error in a Dataflow pipeline is most likely caused by the worker machine type having insufficient memory for the message size and throughput. When messages are large or the throughput is high, each worker must hold data in memory for processing, windowing, and shuffling. If the worker's memory is too small, the JVM heap runs out of memory, leading to an OOM error.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    The BigQuery output table schema does not match the transformed data, causing write failures.

    Why it's wrong here

    A schema mismatch causes write failures or rejected rows, not an out-of-memory error in the worker. It is tempting because BigQuery schema alignment is a common pipeline fault, and would be the correct diagnosis when jobs fail with type or field errors rather than memory exhaustion.

  • ✗

    The Pub/Sub subscription is not acknowledging messages quickly enough, causing a backlog.

    Why it's wrong here

    A slow-acknowledging subscription creates a message backlog and redelivery, not worker heap exhaustion. It is tempting because Pub/Sub backlog is a frequent streaming problem, and would be the correct cause when the symptom is growing subscription backlog or repeated message delivery.

  • ✓

    The worker machine type has insufficient memory for the message size and throughput.

    Why this is correct

    Each Dataflow worker buffers elements in memory during processing and shuffling, so message size multiplied by throughput determines the memory footprint. If that exceeds the worker machine type's available RAM, the worker throws an out-of-memory error, making insufficient worker memory the direct cause.

  • ✗

    The fixed window duration of 1 minute is too short, causing excessive state overhead.

    Why it's wrong here

    Short fixed windows increase state churn but do not by themselves exhaust worker memory; the exhibit points to a different cause. It is tempting because windowing configuration genuinely affects resource usage, and shortening windows would be the right fix for excessive per-window state in a streaming pipeline.

About these practice questions

Courseiva writes every PDE question from scratch — 747 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.