Courseiva
Develop data processing →hardMultiple Select

DP-203 Develop data processing Practice Question

Which THREE of the following are best practices for optimizing performance of Delta Lake tables in Azure Synapse Analytics? (Choose three.)

⚠ Common exam trap

Many candidates confuse high-cardinality partitioning with parallelism, not realizing that excessive partitions cause metadata bloat and slow down queries, while Z-order is a complementary technique for non-partition columns.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Run the OPTIMIZE command to compact small files and improve read performance.

Option A is correct because the OPTIMIZE command compacts many small Parquet files into fewer, larger files, reducing per-file overhead and metadata work so that Delta Lake scans and reads run faster. Option C is correct because partitioning on columns commonly used in WHERE clauses lets the engine perform partition pruning, skipping entire directories of data that cannot match the filter and cutting I/O. Option E is correct because Z-ordering (for example, OPTIMIZE ... ZORDER BY (col)) co-locates related data within files, so data-skipping statistics let queries with frequent filter predicates read far fewer rows. Option B is not a performance optimization: VACUUM deletes unreferenced old files to reclaim storage and control cost, and running it too aggressively can break time travel and concurrent readers. Option D is wrong because partitioning on high-cardinality columns such as UserID creates a huge number of tiny partitions and files, which increases metadata overhead and degrades performance rather than maximizing parallelism.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Run the OPTIMIZE command to compact small files and improve read performance.

    Why this is correct

    OPTIMIZE compacts many small files into larger ones, reducing per-file overhead and metadata pressure during reads. This satisfies the small-file performance constraint, since Delta Lake read latency grows with file count; bin compaction via OPTIMIZE restores efficient scan throughput in Synapse.

  • ✗

    Periodically run VACUUM to remove old versions of files that are no longer needed.

    Why it's wrong here

    VACUUM permanently deletes old file versions, destroying time-travel history and breaking concurrent readers; it is a maintenance task, not a performance optimisation. It tempts because removing stale files reclaims storage, which suits retention enforcement, but the question asks about query performance.

  • ✓

    Partition the table on columns that are used in WHERE clauses to enable partition pruning.

    Why this is correct

    Partitioning on frequently filtered columns lets Synapse's Delta reader skip entire partitions via partition pruning, cutting scanned data and I/O. This directly satisfies the WHERE-clause access pattern constraint, since pruning only works when predicates align with the partition key columns.

  • ✗

    Partition on high-cardinality columns like UserID to maximize parallelism.

    Why it's wrong here

    High-cardinality columns such as UserID produce thousands of tiny partitions, creating excessive file metadata overhead and slowing queries. It appeals because partitioning aids pruning, but that benefit applies to low-cardinality columns like date or region, not unique identifiers.

  • ✓

    Use Z-order on columns that are frequently used in filter predicates.

    Why this is correct

    Z-ordering co-locates related data within Delta files, so filter predicates on those columns skip irrelevant row groups via data skipping. This directly reduces bytes scanned for frequently filtered columns, satisfying the Delta Lake performance optimisation requirement.

About these practice questions

One of 509 original DP-203 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This DP-203 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DP-203 exam.