Courseiva
Develop data processing →mediumMultiple Choice

DP-203 Develop data processing Practice Question

You have an Azure Synapse Analytics dedicated SQL pool. A nightly ELT process loads data into a staging table using PolyBase, then transforms and inserts it into a large fact table. You need to minimize data movement during the transformation step and ensure the fact table is optimized for large range scans. Which table distribution and index should you choose for the fact table?

⚠ Common exam trap

The trap here is choosing distribution or indexing based on load convenience rather than on the join keys and scan patterns that dominate the transformation and reporting workload.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use a hash-distributed table on the primary join key with a clustered columnstore index.

For a large fact table in a dedicated SQL pool, hash distribution on the most frequently joined column minimizes data movement during transformations and joins. Pairing it with a clustered columnstore index provides columnar compression and segment elimination, which accelerates the large range scans and aggregations typical of fact-table queries. Together they satisfy both stated requirements.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Use a round-robin distributed table with a clustered index on the date column.

    Why it's wrong here

    Round-robin distribution spreads rows evenly but provides no co-location for joins, forcing costly data movement during the transformation step. A clustered rowstore index on the date column is less efficient than columnstore for large range scans and consumes more space. This combination fails both the minimize-data-movement and optimize-for-range-scans requirements.

  • ✗

    Use a replicated table with a clustered columnstore index.

    Why it's wrong here

    Replication copies the full table to every compute node, which is suitable only for small dimension tables, not a large fact table. Replicating a large fact table would consume excessive storage and slow down loads and transformations. Although columnstore is appropriate for range scans, the replicated distribution is the wrong choice for a large fact table.

  • ✗

    Use a hash-distributed table on the date column with a heap index.

    Why it's wrong here

    Hash distribution on the date column can create data skew if dates cluster, and it does not co-locate rows for joins on other keys, causing movement during transformations. A heap has no index structure, so large range scans require full table scans, which is inefficient. This combination does not meet either the data-movement or range-scan optimization goals.

  • ✓

    Use a hash-distributed table on the primary join key with a clustered columnstore index.

    Why this is correct

    Hash distribution on the primary join key co-locates matching rows on the same distribution, minimizing data movement during joins and transformations. A clustered columnstore index compresses data and is optimized for large range scans and aggregations, which is exactly the workload described. This combination is the recommended pattern for large fact tables in a dedicated SQL pool.

About these practice questions

One of 509 original DP-203 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Microsoft exam blueprint

This DP-203 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DP-203 exam.