Courseiva
Develop data processing →hardMultiple Choice

DP-203 Develop data processing Practice Question

You are implementing a mapping data flow in Azure Data Factory that joins a large fact table in Azure Synapse Analytics with a slowly changing dimension (SCD) table in Azure SQL Database. The fact table has 500 million rows and the dimension has 2 million rows. You need to optimize the join performance and minimize data movement. The dimension table is small enough to fit in memory. Which join type should you configure in the data flow?

⚠ Common exam trap

The trap here is assuming that a hash join always avoids shuffling, when in mapping data flows only a broadcast join explicitly sends the small table to all nodes to avoid moving the large table.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Broadcast join

Broadcast join is designed for scenarios where one dataset is small enough to fit in memory. By broadcasting the dimension table, the large fact table is not shuffled, minimizing data movement and improving performance. Other join types like sort-merge or hash without broadcast would require shuffling the large fact table, leading to unnecessary data movement and slower execution.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Sort-merge join

    Why it's wrong here

    Sort-merge join requires both datasets to be sorted and then merged, which involves shuffling and sorting both sides of the join. For a 500-million-row fact table, this would cause massive data movement and high latency. While sort-merge can be efficient for large-to-large joins, it is not ideal when one side is small enough to broadcast, as it unnecessarily processes the large dataset through sorting and shuffling.

  • ✗

    Cross join

    Why it's wrong here

    Cross join produces a Cartesian product of both tables, which would generate an enormous number of rows (500 million times 2 million) and is not suitable for a join with a join condition. It would cause extreme resource consumption and is semantically incorrect for joining a fact table with a dimension based on a key relationship.

  • ✓

    Broadcast join

    Why this is correct

    A broadcast join sends the smaller dataset to all compute nodes, allowing the larger dataset to be partitioned and processed locally without shuffling. Since the dimension table has only 2 million rows and fits in memory, broadcasting it eliminates the need to shuffle the 500-million-row fact table, drastically reducing data movement and improving performance. This is the optimal choice for joining a large fact with a small dimension.

  • ✗

    Hash join

    Why it's wrong here

    Hash join builds a hash table on one side and probes it with the other, but in mapping data flows, the default join type may still require partitioning both sides unless broadcast is specified. Since the dimension is small, broadcasting is more efficient. A generic hash join without broadcast would still shuffle the fact table, increasing data movement and reducing performance compared to broadcasting the smaller dimension.

About these practice questions

This DP-203 question is part of Courseiva's 509-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Microsoft exam blueprint

This DP-203 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DP-203 exam.