DA0-002 Data Concepts and Environments Practice Question
A data engineer is designing a system to store raw sensor data from thousands of IoT devices. The data is expected to be used for exploratory analytics and machine learning. Which storage solution is most appropriate?
⚠ Common exam trap
A common mix-up: candidates confuse a data lake with a data warehouse or relational database, assuming raw data must be structured immediately, when in fact a data lake's schema-on-read approach is specifically designed for exploratory and machine learning use cases.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Data lake
A data lake is the most appropriate choice because it can store raw, unprocessed sensor data in its native format (e.g., JSON, Parquet, or binary) without requiring a predefined schema. This flexibility supports exploratory analytics and machine learning workflows where data schemas may evolve or be unknown at ingestion time. Data lakes also scale horizontally to handle the high volume and velocity of data from thousands of IoT devices, unlike traditional storage systems that impose rigid structures or size limits.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Data lake
Why this is correct
A data lake stores raw, schema-on-read data at scale without upfront transformation, suiting thousands of IoT devices and varied formats. Exploratory analytics and machine learning need that flexible, unmodelled raw data, which a data warehouse's structured schema would constrain.
- ✗
Relational database
Why it's wrong here
A relational database enforces a fixed schema and row-based storage, which struggles with the volume, velocity and schema variability of thousands of raw sensor streams. It is tempting because SQL querying suits analytics. A relational database would be correct for highly structured transactional data needing ACID guarantees and complex joins.
- ✗
Data mart
Why it's wrong here
A data mart is a curated subset serving one department's reporting, so it cannot ingest thousands of raw device streams or support broad exploratory analytics. It is tempting because marts hold analytics-ready data. A data mart would be correct when a defined business unit needs governed, pre-modelled reporting rather than raw ingestion.
- ✗
Key-value store
Why it's wrong here
A key-value store retrieves records by key alone, offering no columnar scan, aggregation or schema-on-read needed for exploratory analytics and machine learning over raw sensor data. It is tempting because it scales writes cheaply. A key-value store would be correct for session state, caching or simple lookup workloads.
Go deeper
Related to this question
About these practice questions
One of 1,004 original DA0-002 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This DA0-002 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DA0-002 exam.