Courseiva

AI0-001 AI Models and Data Engineering Practice Question

A data engineer is designing a pipeline to ingest high-velocity clickstream events from a web application into a data lake. The events must be queryable within minutes of arrival, and the schema evolves frequently as new fields are added. Which storage approach best meets these requirements?

⚠ Common exam trap

The trap here is assuming that a relational database or message queue can serve as a scalable, queryable store for evolving, high-velocity data without significant drawbacks.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Store raw JSON files in an object storage bucket partitioned by date, and query them directly using a serverless SQL engine.

Storing raw JSON in object storage and querying with a serverless SQL engine provides schema-on-read flexibility and near-real-time access. It accommodates evolving schemas without costly migrations and scales to high volumes. The other options impose rigid schemas, high latency, or lack analytical query capabilities, making them unsuitable for this scenario.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Batch load events every 24 hours into a data warehouse using a rigid ETL process with predefined transformations.

    Why it's wrong here

    Daily batch loading introduces latency of up to 24 hours, failing the requirement for minute-level queryability. Rigid ETL with predefined transformations cannot easily accommodate frequent schema changes. This approach is better suited for stable, low-velocity data and does not support real-time analytics on evolving clickstream events.

  • ✓

    Store raw JSON files in an object storage bucket partitioned by date, and query them directly using a serverless SQL engine.

    Why this is correct

    This approach supports schema-on-read, allowing new fields to appear without rewriting existing data. Partitioning by date improves query performance and reduces scanned data. Serverless SQL engines can query JSON directly, enabling near-real-time analysis within minutes. It is cost-effective and scales with event volume, making it ideal for evolving clickstream data.

  • ✗

    Load events into a relational database with a fixed schema, using ALTER TABLE to add columns as new fields appear.

    Why it's wrong here

    A relational database with a fixed schema requires migrations for every new field, which is slow and disruptive for high-velocity, evolving data. Frequent ALTER TABLE operations can lock tables and degrade performance. This approach does not scale well for clickstream volumes and cannot provide the flexibility needed for rapid schema changes.

  • ✗

    Write events to a distributed message queue and retain them for 7 days, querying the queue directly for analytics.

    Why it's wrong here

    Message queues are designed for transient buffering, not long-term storage or analytical queries. Querying a queue directly is inefficient and lacks indexing and SQL support. Data older than the retention period is lost. This does not meet the requirement for minute-level queryability over historical data.

Quick reference

Cloud Service Model Comparison

ModelYou ManageProvider ManagesExamples
IaaSOS, runtime, apps, dataHardware, hypervisor, networkingEC2, Azure VMs, GCP Compute Engine
PaaSApps and dataOS, runtime, middleware, hardwareElastic Beanstalk, Azure App Service
SaaSData and settings onlyEverything elseMicrosoft 365, Salesforce, Workday
FaaS / ServerlessFunction code onlyInfra, scaling, runtimeLambda, Azure Functions, Cloud Run
CaaSContainers and appsKubernetes, OS, hardwareEKS, AKS, GKE

About these practice questions

Courseiva writes every AI0-001 question from scratch — 962 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official CompTIA exam blueprint

This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.