easyMultiple Choice
PDE Practice Question: Based on the exhibit, what is the most likely…
Exhibit
Refer to the exhibit. ``` # Dataflow pipeline error log: Workflow failed. Causes: S02:ReadPubSub/Read+Transform/ParDo(ExtractTimestamps)+ ... (4b9c3d2e) The job failed because a worker experienced a "out of memory" error. ``` Pipeline configuration: - Streaming engine: disabled - Worker machine type: n1-standard-4 (4 vCPU, 15 GB memory) - Number of workers: 2 (autoscaling enabled, max 10) - Input: Pub/Sub topic with 1000 messages/sec, each message ~50 KB - Transform: Parse JSON, enrich with external API call, window into 1-minute fixed windows, write to BigQuery
Based on the exhibit, what is the most likely cause of the out-of-memory error?
⚠ Common exam trap
Google Cloud often tests the misconception that OOM errors are caused by schema mismatches or Pub/Sub backlogs, but the real cause is almost always insufficient worker memory for the data volume.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
The worker machine type has insufficient memory for the message size and throughput.
The out-of-memory error in a Dataflow pipeline is most likely caused by the worker machine type having insufficient memory for the message size and throughput. When messages are large or the throughput is high, each worker must hold data in memory for processing, windowing, and shuffling. If the worker's memory is too small, the JVM heap runs out of memory, leading to an OOM error.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
The BigQuery output table schema does not match the transformed data, causing write failures.
Why it's wrong here
A schema mismatch causes write failures or rejected rows, not an out-of-memory error in the worker. It is tempting because BigQuery schema alignment is a common pipeline fault, and would be the correct diagnosis when jobs fail with type or field errors rather than memory exhaustion.
- ✗
The Pub/Sub subscription is not acknowledging messages quickly enough, causing a backlog.
Why it's wrong here
A slow-acknowledging subscription creates a message backlog and redelivery, not worker heap exhaustion. It is tempting because Pub/Sub backlog is a frequent streaming problem, and would be the correct cause when the symptom is growing subscription backlog or repeated message delivery.
- ✓
The worker machine type has insufficient memory for the message size and throughput.
Why this is correct
Each Dataflow worker buffers elements in memory during processing and shuffling, so message size multiplied by throughput determines the memory footprint. If that exceeds the worker machine type's available RAM, the worker throws an out-of-memory error, making insufficient worker memory the direct cause.
- ✗
The fixed window duration of 1 minute is too short, causing excessive state overhead.
Why it's wrong here
Short fixed windows increase state churn but do not by themselves exhaust worker memory; the exhibit points to a different cause. It is tempting because windowing configuration genuinely affects resource usage, and shortening windows would be the right fix for excessive per-window state in a streaming pipeline.
Go deeper
Related to this question
About these practice questions
Courseiva writes every PDE question from scratch — 747 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.