PDE Designing Data Processing Systems Practice Question
A data pipeline uses Cloud Data Fusion to perform ETL jobs. The pipeline reads from BigQuery, transforms data using Wrangler, and writes to Cloud Storage. The team notices that the pipeline runs slower than expected. They suspect the Data Fusion instance is under-provisioned. Which action should be taken to improve performance?
⚠ Common exam trap
PDE often tests the confusion that adding metadata or auxiliary services (like Dataproc Metastore) improves ETL performance, when the real lever is the Data Fusion instance type and Dataproc cluster sizing.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Change the Data Fusion instance type from Basic to Enterprise
Cloud Data Fusion instance types (Basic, Enterprise, Developer) determine the compute resources available to run pipelines. Upgrading from Basic to Enterprise increases the number of Dataproc worker nodes and resources available to the CDAP runtime, improving pipeline throughput and reducing runtime. This directly addresses the under-provisioned instance suspicion.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Add more Dataproc Metastore instances
Why it's wrong here
Dataproc Metastore is a managed Hive metastore service for cataloguing table metadata; it does not add execution capacity to a Data Fusion instance. The tempting fit is that Data Fusion pipelines can use it for schema metadata, so it appears pipeline-related, but it would be the right choice only when pipelines need a shared, persistent metastore rather than more compute.
- ✓
Change the Data Fusion instance type from Basic to Enterprise
Why this is correct
Data Fusion instance type determines the available compute and memory for pipeline execution. Basic instances are limited to a small, fixed profile, so upgrading to Enterprise provides greater resources and the ability to scale executors, directly addressing the suspected under-provisioning causing slow ETL runs.
- ✗
Enable Data Fusion accelerator for BigQuery
Why it's wrong here
The BigQuery accelerator is a plugin that optimises reading from BigQuery, but the stem's bottleneck is the under-provisioned Data Fusion instance itself, so throughput remains capped by instance resources. It is tempting because the pipeline reads from BigQuery, and it would be correct if BigQuery read latency, not instance capacity, were the constraint.
- ✗
Rewrite the pipeline using Cloud Dataprep instead
Why it's wrong here
Cloud Dataprep is a separate, interactive data-preparation service; rewriting the pipeline there does not add executor capacity to the Data Fusion instance, so the bottleneck remains. It is tempting because Dataprep shares the Wrangler transformation experience, and would be correct if the goal were ad-hoc data exploration rather than production ETL throughput.
Go deeper
Related to this question
About these practice questions
This PDE question is part of Courseiva's 747-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.