Courseiva

Google PCA Design and plan a cloud solution architecture Practice Question

A logistics company is planning to migrate a batch ETL pipeline from on-premises Hadoop to Google Cloud. The pipeline processes several terabytes nightly, and the team wants to minimize infrastructure management while keeping the ability to tune the cluster for cost and performance. The data currently resides in an on-premises HDFS cluster. Which combination of services should the architect recommend?

⚠ Common exam trap

The trap here is assuming a persistent Dataproc cluster is needed for a recurring batch job, when ephemeral per-job clusters are cheaper and require less management.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Cloud Storage for durable object storage and Dataproc clusters created per job for processing.

Separating storage from compute with Cloud Storage and using ephemeral Dataproc clusters lets the team process nightly data without paying for idle infrastructure, while still selecting machine types and cluster sizes per job. Persistent clusters, self-managed Hadoop, and a full BigQuery rewrite each fail at least one requirement around management overhead or preserving existing Spark logic.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    BigQuery for storage and analysis, replacing the Spark jobs entirely with SQL.

    Why it's wrong here

    BigQuery is a strong analytics engine, but rewriting an existing Spark or Hadoop pipeline into SQL is a large migration effort that may not preserve custom transformations or libraries. The question asks for a migration path that minimizes management while retaining tuning control, and a full rewrite goes beyond that scope and risks changing business logic.

  • ✓

    Cloud Storage for durable object storage and Dataproc clusters created per job for processing.

    Why this is correct

    Cloud Storage provides cheap, durable storage that decouples data from compute, and Dataproc offers managed Spark and Hadoop clusters that can be created for each job and deleted afterward. This pattern eliminates idle cluster costs and reduces operational burden while still allowing the team to choose machine types and cluster sizes per run. It fits the migration and tuning requirements well.

  • ✗

    Persistent Dataproc clusters with local HDFS storage that run continuously and process jobs on a schedule.

    Why it's wrong here

    Always-on clusters with local HDFS incur cost even when idle and require patching, scaling, and HDFS capacity management. This replicates the operational burden the team wants to eliminate. While it works functionally, it does not minimize infrastructure management or cost for a nightly batch workload, making it a poor architectural fit.

  • ✗

    Compute Engine VMs running a self-managed Hadoop distribution with Cloud Storage FUSE for input and output.

    Why it's wrong here

    Self-managed Hadoop on Compute Engine requires the team to install, patch, and scale the cluster themselves, which is exactly the operational overhead the company wants to avoid. Cloud Storage FUSE adds its own performance and consistency considerations. This approach offers tuning control but fails the goal of minimizing infrastructure management.

About these practice questions

Courseiva writes every PCA question from scratch — 807 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PCA practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PCA exam.