Courseiva

PDE Designing Data Processing Systems Practice Question

A data pipeline uses Cloud Data Fusion to perform ETL jobs. The pipeline reads from BigQuery, transforms data using Wrangler, and writes to Cloud Storage. The team notices that the pipeline runs slower than expected. They suspect the Data Fusion instance is under-provisioned. Which action should be taken to improve performance?

⚠ Common exam trap

PDE often tests the confusion that adding metadata or auxiliary services (like Dataproc Metastore) improves ETL performance, when the real lever is the Data Fusion instance type and Dataproc cluster sizing.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Change the Data Fusion instance type from Basic to Enterprise

Cloud Data Fusion instance types (Basic, Enterprise, Developer) determine the compute resources available to run pipelines. Upgrading from Basic to Enterprise increases the number of Dataproc worker nodes and resources available to the CDAP runtime, improving pipeline throughput and reducing runtime. This directly addresses the under-provisioned instance suspicion.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Add more Dataproc Metastore instances

    Why it's wrong here

    Dataproc Metastore is a managed Hive metastore service for cataloguing table metadata; it does not add execution capacity to a Data Fusion instance. The tempting fit is that Data Fusion pipelines can use it for schema metadata, so it appears pipeline-related, but it would be the right choice only when pipelines need a shared, persistent metastore rather than more compute.

  • ✓

    Change the Data Fusion instance type from Basic to Enterprise

    Why this is correct

    Data Fusion instance type determines the available compute and memory for pipeline execution. Basic instances are limited to a small, fixed profile, so upgrading to Enterprise provides greater resources and the ability to scale executors, directly addressing the suspected under-provisioning causing slow ETL runs.

  • ✗

    Enable Data Fusion accelerator for BigQuery

    Why it's wrong here

    The BigQuery accelerator is a plugin that optimises reading from BigQuery, but the stem's bottleneck is the under-provisioned Data Fusion instance itself, so throughput remains capped by instance resources. It is tempting because the pipeline reads from BigQuery, and it would be correct if BigQuery read latency, not instance capacity, were the constraint.

  • ✗

    Rewrite the pipeline using Cloud Dataprep instead

    Why it's wrong here

    Cloud Dataprep is a separate, interactive data-preparation service; rewriting the pipeline there does not add executor capacity to the Data Fusion instance, so the bottleneck remains. It is tempting because Dataprep shares the Wrangler transformation experience, and would be correct if the goal were ad-hoc data exploration rather than production ETL throughput.

About these practice questions

This PDE question is part of Courseiva's 747-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.