DP-203 Develop data processing Practice Question
You are designing a data pipeline in Azure Synapse Analytics to ingest data from Azure Blob Storage into a dedicated SQL pool. The source files are CSV with varying row lengths, and you need to ensure optimal performance for reads. Which file format and compression should you recommend?
⚠ Common exam trap
Microsoft often tests the misconception that row-based formats like Avro or CSV are suitable for analytical workloads, but the trap here is that columnar formats (Parquet/ORC) are required for optimal read performance in Synapse dedicated SQL pools, and Snappy is preferred over Zlib for speed-critical pipelines.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Parquet with Snappy compression
Parquet with Snappy compression is optimal for dedicated SQL pools in Azure Synapse Analytics because Parquet is a columnar format that enables efficient predicate pushdown and column pruning, reducing I/O. Snappy provides fast compression/decompression with minimal CPU overhead, which is critical for high-throughput reads in a distributed MPP environment.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Avro with Deflate compression
Why it's wrong here
Avro is a row-based format, so reading selected columns still deserialises every field, and Deflate decompression adds CPU cost without column pruning. It suits whole-record streaming writes into Event Hubs or Data Lake, not analytical reads from a dedicated SQL pool where Parquet with Snappy or columnstore delivers the required performance.
- ✗
CSV with Gzip compression
Why it's wrong here
Gzip is not splittable, so a single compressed CSV cannot be read in parallel across Synapse distributions, throttling throughput. It tempts because CSV preserves the source layout and Gzip shrinks files cheaply, but row-variable CSV forces full-file parsing rather than columnar reads.
- ✓
Parquet with Snappy compression
Why this is correct
Parquet is columnar, so Synapse reads only referenced columns and skips irrelevant data, while Snappy decompresses quickly at low CPU cost. This combination delivers the optimal read performance the variable-length CSV source rows cannot provide.
- ✗
ORC with Zlib compression
Why it's wrong here
ORC is columnar and Zlib-compressed, which suits Synapse reads, but the source files are CSV with varying row lengths, so conversion is required and ragged rows break ORC's typed schema. It tempts because columnar storage accelerates dedicated SQL pool queries once data is already structured.
Quick reference
Azure Blob Storage Tier Comparison
| Tier | Storage Cost | Retrieval Cost | Latency | Use Case |
|---|---|---|---|---|
| Hot | Highest | Lowest | Immediate | Active data, frequent reads |
| Cool | Lower | Higher | Immediate | Data accessed < once / month |
| Cold | Lower still | Higher | Immediate | Data accessed < once / quarter |
| Archive | Lowest | Highest + rehydration delay | Hours | Long-term compliance retention |
Go deeper
Related to this question
About these practice questions
This DP-203 question is part of Courseiva's 509-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This DP-203 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DP-203 exam.