DA0-002 Data Concepts and Environments Practice Question
An organization is implementing a data lake to store raw data from various sources. Which THREE characteristics are typically associated with a data lake compared to a data warehouse?
⚠ Common exam trap
CompTIA often tests the misconception that data lakes require data transformation before loading (schema-on-write), when in fact they use schema-on-read, allowing raw data storage without upfront transformation.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Supports batch and real-time processing
A is correct because a data lake is designed to handle both batch and real-time/streaming ingestion and processing, unlike a traditional data warehouse that is primarily optimized for batch ETL workloads. B is correct because a data lake stores data in its native/raw format (e.g., JSON, Parquet, CSV, images, logs) without forcing an upfront conversion. C is correct because a data lake applies schema-on-read, meaning the schema is defined when the data is queried rather than when it is written. D is incorrect because data lakes support structured, semi-structured, and unstructured data, not only structured data. E is incorrect because a data lake typically loads raw data as-is and defers transformation until read/query time, whereas a data warehouse usually requires transformation before loading.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Supports batch and real-time processing
Why this is correct
A data lake ingests streams and files through the same storage layer, so it handles both batch loads and real-time processing. A data warehouse typically relies on scheduled ETL batches, making this a defining capability of the lake architecture.
- ✓
Stores data in its native format
Why this is correct
A data lake retains source files exactly as received, whether CSV, JSON, Parquet or log data, without forcing transformation on ingestion. A data warehouse instead requires data to be modelled and loaded into its predefined relational schema before it can be queried.
- ✓
Schema-on-read approach
Why this is correct
Schema-on-read defers structure until query time, letting the lake ingest raw, unmodelled data from varied sources without upfront transformation. This satisfies the stem's requirement to store raw data, unlike a warehouse's schema-on-write, which demands modelling before loading and would block the ingestion of undefined formats.
- ✗
Supports only structured data
Why it's wrong here
A data lake stores structured, semi-structured and unstructured data alike, so restricting it to structured data contradicts its purpose. It is tempting because structured-only storage is a genuine characteristic of a data warehouse, and would be correct when describing warehouse capabilities rather than lake capabilities.
- ✗
Requires data transformation before loading
Why it's wrong here
A data lake loads raw data without requiring prior transformation, applying schema on read instead. It is tempting because transformation before loading is the defining trait of a data warehouse ETL process, and would be correct when describing warehouse ingestion rather than lake ingestion.
Go deeper
Related to this question
About these practice questions
One of 1,004 original DA0-002 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This DA0-002 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DA0-002 exam.