DA0-002 Data Concepts and Environments Practice Question
A data architect is designing a system for a subscription streaming service. The service must record every play, pause, and skip event from millions of concurrent viewers with very low write latency, and it must later support analytical queries over months of event history. The architect wants a single storage layer that handles both needs without a separate transformation pipeline. Which data architecture should the architect choose?
⚠ Common exam trap
The trap here is treating a message queue as a storage layer for long-term analytics, when its retention limits mean historical queries still require another system.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
A lakehouse that combines open table formats with ACID transactions and query engines over the same storage
A lakehouse unifies streaming ingestion and analytical querying on one storage layer by adding ACID transactions and schema management to open table formats on object storage. This satisfies the requirement for low-latency event writes plus months of queryable history without a separate transformation pipeline.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
A message queue that retains events for a fixed retention period and serves analytical queries directly
Why it's wrong here
A message queue such as Apache Kafka is designed for durable event transport and short-term retention, not for serving months of analytical queries. Query engines over queues are limited and retention windows are usually short. It would need to feed another storage layer for history, which is the separate pipeline the architect wants to avoid.
- ✗
A traditional data warehouse that ingests events only through nightly ETL batches
Why it's wrong here
Nightly ETL batches cannot capture millions of concurrent events with low write latency, and they delay analytical availability by up to a day. The scenario requires near-real-time ingestion, which a batch-only warehouse does not provide. It also typically imposes a rigid schema that complicates evolving event types.
- ✓
A lakehouse that combines open table formats with ACID transactions and query engines over the same storage
Why this is correct
A lakehouse uses open table formats such as Delta Lake or Apache Iceberg to add ACID transactions and schema enforcement directly on object storage, so streaming writes and analytical reads share one layer. It removes the need for a separate transformation pipeline while supporting low-latency ingestion and historical queries over the same data.
- ✗
A data lake that stores raw event files and requires a separate batch job to load them into a warehouse
Why it's wrong here
A data lake with a separate batch load introduces a transformation pipeline and latency between event capture and analytical availability. The architect explicitly wants a single layer without a separate pipeline. A lake alone also does not provide low-latency point writes for millions of concurrent events, so it does not meet the stated requirement.
Go deeper
Related to this question
About these practice questions
Courseiva writes every DA0-002 question from scratch — 1,004 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official CompTIA exam blueprint
This DA0-002 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DA0-002 exam.