Courseiva

AI0-001 AI Models and Data Engineering Practice Question

A data engineer is designing a pipeline for a streaming data application that uses a machine learning model to detect anomalies in real time. Which TWO practices should the engineer implement to ensure data quality and model reliability?

⚠ Common exam trap

CompTIA often tests the misconception that batch processing or fixed retraining schedules are sufficient for real-time streaming applications, when in fact sliding windows and continuous validation are required to maintain low latency and model accuracy.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use a sliding window for feature computation

Option C is correct because a sliding window computes features over a continuously advancing time range, which is essential for real-time anomaly detection since it captures recent, temporally relevant data points while maintaining feature freshness as the stream evolves. Option D is correct because implementing data validation checks at the ingestion point catches malformed, missing, or out-of-range records before they reach the model, preventing garbage-in-garbage-out failures and preserving both data quality and model reliability in a streaming pipeline. Option A is not appropriate because batch processing in fixed intervals introduces latency and defeats the real-time requirement of the streaming application. Option B is not required for data quality or model reliability; storing all raw data indefinitely is a retention/archival decision and does not itself validate or improve streaming data. Option E is not ideal because a fixed 24-hour retraining schedule cannot adapt to concept drift or anomalies that emerge within the stream, and retraining cadence should be driven by drift detection rather than an arbitrary fixed interval.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Use batch processing to transform data in fixed intervals

    Why it's wrong here

    Batch transformation in fixed intervals introduces latency, so anomalies are detected after the window closes rather than in real time. Batch suits periodic reporting; streaming pipelines require continuous per-event processing to preserve timeliness and quality.

  • ✗

    Store all raw data indefinitely for future analysis

    Why it's wrong here

    Storing raw data indefinitely addresses retention and auditability, not the streaming pipeline's need for schema validation and drift detection on incoming records. It is tempting because immutable raw storage genuinely supports later retraining and forensic replay, and would be the right choice when regulatory retention or reproducible model rebuilds are the stated requirement.

  • ✓

    Use a sliding window for feature computation

    Why this is correct

    A sliding window recomputes features over the most recent events, keeping anomaly scores aligned with current behaviour. Fixed or expanding windows let stale data dilute the signal, degrading real-time detection accuracy as stream characteristics drift.

  • ✓

    Implement data validation checks at the ingestion point

    Why this is correct

    Validating records at ingestion stops malformed or out-of-range values before they reach the model, satisfying the real-time constraint where downstream batch cleansing is impossible. This preserves anomaly-detection accuracy, since corrupt features would otherwise produce false alerts. It also supports model reliability by keeping the streaming input distribution consistent with training data.

  • ✗

    Retrain the model on a fixed schedule every 24 hours

    Why it's wrong here

    A fixed 24-hour retrain cannot react to concept drift or sudden anomalies within the streaming window, so the model stays stale between cycles. Scheduled retraining suits stable batch workloads; real-time pipelines need drift detection triggering retraining on data change.

About these practice questions

One of 962 original AI0-001 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.