Courseiva
Using Spark SQL →hardMultiple Select

Databricks-Spark-Assoc Using Spark SQL Practice Question

You are using Spark SQL to join two large Delta tables, orders and customers, on a common column customer_id. The orders table is partitioned by order_date, and the customers table is not partitioned. You need to ensure the join is efficient and minimizes shuffling. Which TWO actions should you take? (Choose two.)

⚠ Common exam trap

The trap here is assuming that repartitioning both tables is always beneficial, or that partitioning the smaller table by the join key will help, when in fact it can cause more shuffling or small file problems.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Enable adaptive query execution (AQE) and set spark.sql.adaptive.enabled to true.

Broadcasting the smaller table eliminates shuffling of the larger table. Enabling AQE allows Spark to dynamically optimize the join, potentially converting to a broadcast join or adjusting partitions. These two actions together minimize shuffling and improve efficiency.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Enable adaptive query execution (AQE) and set spark.sql.adaptive.enabled to true.

    Why this is correct

    AQE can dynamically optimize joins by converting sort-merge joins to broadcast joins if one side is small after initial stages, and by coalescing partitions. Enabling AQE helps minimize shuffling and improves performance. It is a best practice for large joins in Spark SQL, especially when statistics are outdated or data skew exists.

  • ✗

    Use the MERGE command to combine the tables.

    Why it's wrong here

    MERGE is used for upserts into a Delta table, not for joining two tables for analysis. It is not a join operation and would not produce a joined result set. It is used for transactional writes, not for query optimization. Therefore, it is irrelevant to the scenario.

  • ✗

    Partition the customers table by customer_id to match the orders table.

    Why it's wrong here

    Partitioning the customers table by customer_id would not help the join because the orders table is partitioned by order_date, not customer_id. Partitioning by customer_id would create many small partitions and might not align with the join key distribution. Moreover, partitioning a large table by a high-cardinality column like customer_id is generally not recommended due to small file issues.

  • ✓

    Broadcast the customers table if it is small enough to fit in memory.

    Why this is correct

    Broadcasting a small table avoids shuffling the larger table. In Spark SQL, you can use the BROADCAST hint or rely on automatic broadcast join if the table size is below spark.sql.autoBroadcastJoinThreshold. This is effective when one side is small, as it replicates the small table to all executors, eliminating a shuffle of the large table.

  • ✗

    Repartition both tables by customer_id before the join.

    Why it's wrong here

    Repartitioning both tables by customer_id would ensure that matching keys are in the same partition, but it incurs a full shuffle of both large tables. This is expensive and not the most efficient approach when one table can be broadcast. It might be necessary if both tables are large, but the question asks for actions to minimize shuffling, and broadcasting is better when applicable.

About these practice questions

One of 295 original Databricks-Spark-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.