AI0-001 AI Models and Data Engineering Practice Question
A data engineer is designing a pipeline to ingest high-velocity clickstream events from a web application into a data lake. The events must be queryable within minutes of arrival, and the schema evolves frequently as new fields are added. Which storage approach best meets these requirements?
⚠ Common exam trap
The trap here is assuming that a relational database or message queue can serve as a scalable, queryable store for evolving, high-velocity data without significant drawbacks.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Store raw JSON files in an object storage bucket partitioned by date, and query them directly using a serverless SQL engine.
Storing raw JSON in object storage and querying with a serverless SQL engine provides schema-on-read flexibility and near-real-time access. It accommodates evolving schemas without costly migrations and scales to high volumes. The other options impose rigid schemas, high latency, or lack analytical query capabilities, making them unsuitable for this scenario.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Batch load events every 24 hours into a data warehouse using a rigid ETL process with predefined transformations.
Why it's wrong here
Daily batch loading introduces latency of up to 24 hours, failing the requirement for minute-level queryability. Rigid ETL with predefined transformations cannot easily accommodate frequent schema changes. This approach is better suited for stable, low-velocity data and does not support real-time analytics on evolving clickstream events.
- ✓
Store raw JSON files in an object storage bucket partitioned by date, and query them directly using a serverless SQL engine.
Why this is correct
This approach supports schema-on-read, allowing new fields to appear without rewriting existing data. Partitioning by date improves query performance and reduces scanned data. Serverless SQL engines can query JSON directly, enabling near-real-time analysis within minutes. It is cost-effective and scales with event volume, making it ideal for evolving clickstream data.
- ✗
Load events into a relational database with a fixed schema, using ALTER TABLE to add columns as new fields appear.
Why it's wrong here
A relational database with a fixed schema requires migrations for every new field, which is slow and disruptive for high-velocity, evolving data. Frequent ALTER TABLE operations can lock tables and degrade performance. This approach does not scale well for clickstream volumes and cannot provide the flexibility needed for rapid schema changes.
- ✗
Write events to a distributed message queue and retain them for 7 days, querying the queue directly for analytics.
Why it's wrong here
Message queues are designed for transient buffering, not long-term storage or analytical queries. Querying a queue directly is inefficient and lacks indexing and SQL support. Data older than the retention period is lost. This does not meet the requirement for minute-level queryability over historical data.
Quick reference
Cloud Service Model Comparison
| Model | You Manage | Provider Manages | Examples |
|---|---|---|---|
| IaaS | OS, runtime, apps, data | Hardware, hypervisor, networking | EC2, Azure VMs, GCP Compute Engine |
| PaaS | Apps and data | OS, runtime, middleware, hardware | Elastic Beanstalk, Azure App Service |
| SaaS | Data and settings only | Everything else | Microsoft 365, Salesforce, Workday |
| FaaS / Serverless | Function code only | Infra, scaling, runtime | Lambda, Azure Functions, Cloud Run |
| CaaS | Containers and apps | Kubernetes, OS, hardware | EKS, AKS, GKE |
About these practice questions
Courseiva writes every AI0-001 question from scratch — 962 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official CompTIA exam blueprint
This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.