Courseiva

DA0-002 Data Concepts and Environments Practice Question

A company is designing a data lake to store raw sensor data from IoT devices. The data arrives as JSON objects with varying schemas. Which storage approach is most appropriate?

⚠ Common exam trap

Test-takers frequently confuse 'schema-on-read' with 'schema-on-write' and assume that converting to a structured format like Avro or columnar storage is always better for performance, ignoring the requirement to store raw, varying-schema data as-is.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Store raw JSON files in a distributed file system and apply schema-on-read

A data lake is designed to store raw data in its native format, and IoT sensor data with varying schemas is best handled by storing raw JSON files in a distributed file system (e.g., HDFS or Amazon S3). This approach leverages schema-on-read, where the schema is applied at query time rather than at write time, allowing flexibility for heterogeneous JSON objects without data loss or transformation overhead.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Ingest into a relational database with a predefined schema

    Why it's wrong here

    A relational database enforces a fixed schema at write time, so varying JSON fields would be rejected or require constant migrations. It suits structured transactional data with stable columns, not schema-on-read raw ingestion; a data lake stores JSON as-is until query time.

  • ✗

    Store each JSON object as a separate file in a compressed columnar format

    Why it's wrong here

    Columnar formats such as Parquet require a defined schema and organise data by column, so one JSON object per file yields tiny files with no columnar benefit and heavy metadata overhead. Columnar storage suits large, uniformly structured analytical datasets, not raw variable-schema JSON.

  • ✗

    Convert all JSON to Avro with a fixed schema before storing

    Why it's wrong here

    Converting to Avro with a fixed schema imposes schema-on-write, discarding or failing on fields outside that schema, which defeats the lake's schema-on-read purpose. Avro suits pipelines with known, evolving-but-governed schemas, not raw IoT JSON whose fields vary per device.

  • ✓

    Store raw JSON files in a distributed file system and apply schema-on-read

    Why this is correct

    Schema-on-read defers parsing until query time, so each JSON object's varying structure is interpreted individually rather than forced into a fixed table. This directly satisfies the stem's requirement to store raw sensor data whose schemas differ across IoT devices, avoiding ingestion-time transformation failures.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

One of 1,004 original DA0-002 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This DA0-002 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DA0-002 exam.