Google PCA Design and plan a cloud solution architecture Practice Question
Your organization runs a batch analytics platform that ingests data from a Pub/Sub topic into Cloud Storage, then loads it into BigQuery using a Dataflow streaming pipeline. The pipeline must handle sudden bursty traffic during month-end reporting, and you want to minimize operational overhead while ensuring the pipeline scales automatically. Which architectural approach should you choose?
⚠ Common exam trap
The trap here is assuming that any autoscaling compute service (like managed instance groups or Dataproc) is equally low-overhead, when managed Dataflow specifically abstracts infrastructure and integrates with the data services.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use a Dataflow streaming pipeline with autoscaling enabled, reading from Pub/Sub and writing to Cloud Storage and BigQuery.
A Dataflow streaming pipeline with autoscaling is the most suitable choice because it automatically adjusts worker count to handle bursty Pub/Sub traffic, integrates natively with Cloud Storage and BigQuery, and minimizes operational overhead. It provides exactly-once processing and handles backpressure, ensuring reliable and scalable analytics during month-end peaks.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Deploy a Compute Engine managed instance group running a custom ingestion script, and configure an autoscaler based on Pub/Sub queue depth.
Why it's wrong here
A managed instance group can scale, but you must build and maintain the ingestion logic, handle Pub/Sub message acknowledgement, and manage VM images and patching. This increases operational overhead compared to a managed Dataflow service. It also lacks built-in connectors for Cloud Storage and BigQuery, requiring custom code, which is error-prone and not ideal for bursty month-end traffic.
- ✓
Use a Dataflow streaming pipeline with autoscaling enabled, reading from Pub/Sub and writing to Cloud Storage and BigQuery.
Why this is correct
Dataflow streaming with autoscaling dynamically adjusts the number of workers based on backlog and CPU utilization, matching bursty Pub/Sub traffic without manual intervention. It natively integrates with Pub/Sub, Cloud Storage, and BigQuery, so you avoid managing infrastructure. This directly satisfies the requirement for automatic scaling and minimal operational overhead for a streaming analytics pipeline.
- ✗
Use a Dataproc cluster with autoscaling to run a Spark Streaming job that reads from Pub/Sub and writes to BigQuery.
Why it's wrong here
Dataproc with Spark Streaming requires managing a cluster, tuning Spark, and handling Pub/Sub connectors, which adds operational overhead. While it can scale, it is not as tightly integrated or as simple as Dataflow for this use case. The requirement emphasizes minimal operational overhead, and Dataproc demands more cluster and job management than a serverless Dataflow pipeline.
- ✗
Create a Cloud Function that triggers on each Pub/Sub message and writes directly to BigQuery and Cloud Storage.
Why it's wrong here
Cloud Functions are event-driven and can trigger per message, but they are not designed for high-throughput streaming pipelines with bursty traffic. Each invocation has limits on execution time and concurrency, and writing to both Cloud Storage and BigQuery per message would be inefficient. This approach lacks the scalable, stateful processing that Dataflow provides for month-end spikes.
Go deeper
Related to this question
Learn chapter
Cloud SQL and Managed Data Stores
Key term
Batch
Batch is a cloud computing service that runs large numbers of computing jobs as a group, or batch, without needing to manage individual servers.
Key term
Cloud storage
Cloud storage is a service that lets you save data on remote servers accessed over the internet instead of on your computer's hard drive.
About these practice questions
This PCA question is part of Courseiva's 807-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PCA practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PCA exam.