Courseiva
hardMultiple Select

Optimizing ADLS Gen2 Query Performance

You are monitoring an Azure Data Lake Storage Gen2 account that stores streaming data from IoT devices. You notice that query performance on the data in Parquet format is degrading over time. You need to improve query performance for both current and future data. Which TWO actions should you take?

Quick Answer

The answer is to convert the Parquet files to Delta Lake format and enable file compaction. This combination directly addresses query performance degradation by leveraging Delta Lake’s ACID transactions and built-in optimization features, specifically file compaction which reduces the number of small files that accumulate over time from streaming data, thereby minimizing metadata overhead and improving scan efficiency. On the Microsoft Azure Data Engineer Associate DP-203 exam, this scenario tests your understanding of how to optimize Azure Data Lake Storage Gen2 query performance using partitioning and Delta Lake, a common pattern for handling streaming IoT workloads. A frequent trap is to focus only on partitioning by a filter column like date—while that helps with predicate pushdown, it does not solve the small-file problem that degrades performance in streaming scenarios. Memory tip: think “compact and partition” like packing a suitcase—fewer, larger items (compacted files) are easier to carry, and grouping them by destination (partitioning) lets you skip unnecessary bags entirely.

⚠ Common exam trap

Candidates often confuse data protection features (like soft delete) or storage migration options (like Azure SQL or NetApp Files) with performance optimization techniques, failing to recognize that partitioning and file format optimization are the standard solutions for improving query performance on large-scale Parquet data in a data lake.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Partition the data by a column commonly used in filter conditions.

Option B is correct because partitioning the data by a column commonly used in filter conditions (for example, device ID or event date) enables partition pruning, so queries scan only the relevant folders instead of the entire dataset, which directly improves performance for both existing and newly arriving streaming data. Option C is correct because converting Parquet files to Delta Lake format and enabling file compaction addresses the small-file problem typical of streaming ingestion; OPTIMIZE compaction merges many small files into larger ones, reducing per-file overhead and metadata/listing costs, while Delta Lake adds a transaction log and data-skipping statistics that further accelerate queries. Option A is not appropriate because moving data to Azure SQL Database changes the storage platform rather than optimizing the Data Lake Gen2 Parquet data, and it is not a scalable fit for streaming IoT data. Option D is incorrect because soft delete is a data-protection feature for recovering deleted blobs and has no effect on read/query performance. Option E is incorrect because Azure NetApp Files is a high-performance NFS/SMB file service, not a query-optimization solution for Parquet data in Data Lake Storage Gen2, and migrating would not address the small-file and partitioning issues.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Move frequently accessed data to Azure SQL Database.

    Why it's wrong here

    Relocating data to Azure SQL Database abandons the Data Lake storage tier the streaming pipeline writes to, so it cannot improve Parquet query performance in place. It is tempting because SQL Database offers indexed, low-latency querying, and would be correct if the requirement were serving structured relational workloads rather than optimising files in the lake.

  • ✓

    Partition the data by a column commonly used in filter conditions.

    Why this is correct

    Partitioning by a frequently filtered column enables partition pruning, so queries scan only relevant folders rather than the whole dataset. This reduces I/O for both existing and newly arriving IoT data, directly addressing the degrading Parquet query performance described in the stem.

  • ✓

    Convert the Parquet files to Delta Lake format and enable file compaction.

    Why this is correct

    Delta Lake adds a transaction log and supports OPTIMIZE compaction, merging many small streaming files into larger ones. This cuts per-file overhead and metadata cost, improving query performance for current and future data, which is the stem's stated requirement.

  • ✗

    Enable soft delete on the storage account to optimize read performance.

    Why it's wrong here

    Soft delete is a data-protection feature that retains deleted blobs for recovery; it adds no read-path optimisation. It would be the right choice when guarding against accidental deletion, but it cannot address degrading Parquet query performance.

  • ✗

    Migrate the data to Azure NetApp Files for lower latency.

    Why it's wrong here

    Azure NetApp Files provides high-performance NFS and SMB file shares, not a replacement for Data Lake Storage Gen2's hierarchical namespace and Parquet analytics. Migration suits latency-sensitive file workloads, but the stem requires optimising existing and future Parquet data in place.

About these practice questions

Courseiva writes every DP-203 question from scratch — 509 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

Same concept, more angles

1 more way this is tested on DP-203

These questions test the same concept from different angles. Work through them to make sure you can recognise it however the exam phrases it.

Variation 1. You have an Azure Data Lake Storage Gen2 account that stores large volumes of parquet files. A reporting application frequently queries a specific subset of data filtered by a 'region' column. To minimize query latency and cost, which optimization should you implement?

medium
  • ✓ A.Partition the data by region in the folder structure.
  • B.Create a clustered index on the region column.
  • C.Compress the parquet files using gzip.
  • D.Enable hierarchical namespace on the storage account.

Why A: Partitioning the data by region in the folder structure (e.g., /region=NorthAmerica/...) enables Azure Data Lake Storage Gen2 and query engines like Azure Synapse or PolyBase to perform partition pruning. This skips scanning irrelevant files entirely, reducing I/O and query latency while lowering cost by minimizing data processed.

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This DP-203 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DP-203 exam.