Courseiva

DP-900 Describe core data concepts Practice Question

Which THREE of the following are benefits of using a columnar storage format like Parquet for analytical workloads?

⚠ Common exam trap

A common mix-up: candidates confuse the benefits of columnar storage (optimized for read-heavy, aggregate queries) with row-oriented storage benefits (optimized for frequent updates and transactional integrity), leading them to incorrectly select Option C.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Reduced I/O when querying a subset of columns

Option B is correct because columnar formats like Parquet store each column contiguously, so a query that references only a subset of columns reads only those column chunks rather than full rows, dramatically reducing I/O. Option D is correct because values within a single column share the same data type and tend to be similar, which allows efficient encoding schemes (e.g., dictionary, run-length, delta encoding) and yields much better compression ratios than row-oriented storage. Option E is correct because Parquet stores per-row-group and per-column-chunk statistics (min/max, null counts), enabling predicate pushdown so the engine can skip row groups and column chunks that cannot satisfy the filter. Option A is not a benefit of Parquet: it is a file format, not a database, and it does not enforce referential integrity constraints such as foreign keys. Option C is also wrong because columnar layouts are optimized for bulk scans and aggregations, not frequent single-row updates; row-oriented stores handle point updates far better.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Enforced referential integrity constraints

    Why it's wrong here

    Columnar storage is a physical file layout for tables, not a set of logical rules; referential integrity constraints are enforced by a relational database management system using primary keys and foreign keys. Because formats like Parquet or ORC fundamentally lack an engine to validate relationships between tables, they cannot enforce referential integrity on their own. This option confuses a database server's data-governance capability with the on-disk organization of data, so it is not a benefit of columnar storage.

  • ✓

    Reduced I/O when querying a subset of columns

    Why this is correct

    With a columnar layout, each column is stored contiguously, so a query that needs only a few of a wide table's many columns can read exactly those column segments from disk. This drastically lowers I/O traffic and memory consumption compared to row-oriented storage, where the whole row must be pulled even when only one attribute is needed. Because analytics queries tend to aggregate a small subset of columns over huge numbers of rows, this column pruning is a major performance advantage.

  • ✗

    Optimized for frequent row updates

    Why it's wrong here

    A columnar format scatters the fields of a logical row across many different files or row-group sections, so updating one row means locating and rewriting data in every column involved. Most columnar files are also immutable, forcing a full rewrite of affected row groups or an append with a delete marker, which makes frequent transactional row updates costly. Row-oriented databases are designed for fast single-row changes, whereas columnar storage is optimized for bulk appends and scan-heavy workloads.

  • ✓

    Better compression ratios due to similar data types in columns

    Why this is correct

    Because all values in a column share the same data type and often similar distributions or repeated values, compression schemes like dictionary encoding, run-length encoding, and delta encoding work especially well. This yields compression ratios far higher than those seen in row layouts, where heterogeneous values from multiple columns are packed together. Better compression shrinks storage footprints and also reduces the number of bytes pulled from disk when scanning data.

  • ✓

    Support for predicate pushdown to skip irrelevant data

    Why this is correct

    Columnar engines often store per-file or per-row-group metadata such as minimum and maximum values, so a query engine can use a WHERE clause to skip entire chunks of data that cannot contain matching values. This predicate pushdown is filter-based data skipping, complementary to column pruning: pruning avoids unneeded columns, while predicate pushdown deselects row groups based on value ranges. Formats like Parquet and ORC support this natively, making analytical scans significantly faster.

About these practice questions

This DP-900 question is part of Courseiva's 851-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This DP-900 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DP-900 exam.