Courseiva

DA0-002 Data Concepts and Environments Practice Question

A data engineer is choosing a storage approach for a new analytics platform. The workload consists of wide, denormalized event tables with dozens of attributes, queries that scan a few columns across billions of rows, and heavy aggregation rather than single-row lookups. Which storage structure is best suited to this workload?

⚠ Common exam trap

The trap here is equating fast primary-key lookups with fast analytical scans, when indexing a row store does not reduce the columns read during a wide aggregation.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

A columnar storage format that stores values of each column contiguously and supports compression

The described workload reads a small number of columns from very wide tables and performs large aggregations, which is exactly what columnar storage optimizes: reading only the needed column segments and compressing similar values. Row-oriented tables with B-tree indexes favor point lookups, key-value stores favor document retrieval by key, and heavily normalized schemas add join overhead. Columnar layout therefore fits the analytical scan pattern best.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    A normalized schema with many small tables joined at query time

    Why it's wrong here

    Normalization reduces redundancy for transactional writes, but analytical queries then require many joins across the small tables, adding shuffle and lookup cost. The scenario describes denormalized wide event tables, so re-normalizing them would fight the existing design. Joins also prevent the engine from reading only the needed columns from a single wide table.

  • ✓

    A columnar storage format that stores values of each column contiguously and supports compression

    Why this is correct

    Columnar storage lays each column's values together, so a query that reads only a few attributes touches only those column segments and skips the rest. Contiguous values of the same type also compress far better, reducing I/O for the large aggregations described. This matches the analytical pattern of scanning few columns across many rows.

  • ✗

    A row-oriented relational table with a B-tree index on the primary key

    Why it's wrong here

    Row-oriented storage keeps all columns of a row physically together, so a query touching only a few attributes still reads the entire row from disk. B-tree indexes accelerate point lookups and range scans on the indexed key but do not help wide scans that aggregate a few columns across billions of rows. This design suits transactional workloads, not the described analytical scan pattern.

  • ✗

    A key-value store that maps each event ID to a serialized JSON document

    Why it's wrong here

    A key-value store is optimized for retrieving a document by its key, which suits point lookups rather than scanning and aggregating across billions of rows. To compute an aggregate, the engine would have to fetch and deserialize every document, discarding the wide attributes it does not need. That is the opposite of the efficient column pruning the workload requires.

About these practice questions

This DA0-002 question is part of Courseiva's 1,004-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official CompTIA exam blueprint

This DA0-002 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DA0-002 exam.