Courseiva
mediumMultiple Choice

PDE Practice Question: An e-commerce company runs a daily batch pipeline…

An e-commerce company runs a daily batch pipeline that processes clickstream data from Cloud Storage using Cloud Dataproc with Spark. The pipeline includes a join between a large fact table and a small dimension table. The dimension table is stored in Cloud Storage as a CSV file. The join is slow due to shuffling. The data engineer considers broadcasting the dimension table. However, the dimension table is updated daily and the pipeline reads the latest version. What is the best approach to implement this optimization?

⚠ Common exam trap

A common mix-up: candidates think increasing `spark.sql.autoBroadcastJoinThreshold` is a safe global fix, but it can cause memory pressure and does not guarantee a broadcast join if the table size fluctuates, whereas the explicit broadcast hint provides deterministic behavior.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use DataFrame.join with broadcast hint on the dimension DataFrame

Broadcasting the small dimension table using the broadcast hint (e.g., `broadcast(dimensionDF)`) forces Spark to replicate the dimension data to all executor nodes, eliminating the need for a shuffle during the join. This is ideal when the dimension table is small enough to fit in executor memory, and since the pipeline reads the latest CSV daily, the broadcast will automatically use the updated data without additional code changes.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Use DataFrame.join with broadcast hint on the dimension DataFrame

    Why this is correct

    Broadcast hints avoid the shuffle by replicating the small dimension table to each executor, eliminating the expensive join shuffle. Reading the CSV fresh each run keeps the daily-updated dimension current, satisfying the freshness constraint without caching stale data.

  • ✗

    Read the fact table and dimension table into separate DataFrames and use standard join

    Why it's wrong here

    A standard shuffle join repartitions both datasets by join key across the network, so the dimension table's small size delivers no benefit and the shuffle bottleneck persists. This is the default approach when both tables are large or the dimension exceeds the broadcast threshold.

  • ✗

    Read the dimension table as an RDD and collect as a map, then use map-side join

    Why it's wrong here

    Collecting the dimension table as an RDD into a map on the driver for a map-side join risks out-of-memory errors if the data exceeds driver memory, even for tables considered 'small' for distributed broadcasting. Spark's native broadcast mechanism efficiently distributes the table to executor memory, bypassing this driver bottleneck and scaling robustly. This option is tempting as it mimics a map-side join, suitable for extremely tiny, static lookup tables where the entire dataset is guaranteed to fit comfortably within the driver's allocated memory for direct driver-side processing.

  • ✗

    Increase the spark.sql.autoBroadcastJoinThreshold to a large value

    Why it's wrong here

    Raising autoBroadcastJoinThreshold lets Spark broadcast tables up to that size automatically, but the dimension is read as a CSV DataFrame whose size Spark estimates unreliably, so the join may still shuffle. The threshold setting suits known, accurately sized tables.

About these practice questions

Courseiva writes every PDE question from scratch — 747 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.