Courseiva

AI0-001 AI Implementation and Operations Practice Question

A team deploys a real-time fraud detection model on a streaming platform. The model must produce predictions within 100 milliseconds per event. Initial latency is 150 ms. Which optimization is most likely to meet the latency requirement?

⚠ Common exam trap

CompTIA often tests the misconception that increasing batch size or model complexity improves throughput for real-time systems, but candidates must recognize that real-time streaming requires low per-event latency, not high aggregate throughput.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Apply model quantization to reduce precision from FP32 to INT8.

Model quantization reduces the numerical precision of the model's weights and activations from FP32 to INT8, which decreases memory footprint and speeds up inference. This optimization directly addresses the 150 ms latency by enabling faster arithmetic operations on modern hardware, often cutting inference time by 2-4x, which can bring latency below the 100 ms requirement.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Apply model quantization to reduce precision from FP32 to INT8.

    Why this is correct

    Quantising weights and activations from FP32 to INT8 shrinks memory footprint and enables faster integer arithmetic, typically cutting inference latency substantially. This brings per-event prediction time under the 100 ms streaming requirement without redesigning the model architecture.

  • ✗

    Increase the batch size to process more events simultaneously.

    Why it's wrong here

    Increasing batch size raises throughput, not per-event latency: queuing delays grow because each event waits for a full batch to accumulate before inference, pushing the 150 ms figure further past the 100 ms target. Batching suits offline or high-throughput pipelines where aggregate processing cost matters and individual response deadlines do not apply.

  • ✗

    Add more feature engineering steps to improve model accuracy.

    Why it's wrong here

    Adding feature engineering steps increases per-event computation, pushing latency further above the 100 ms budget rather than reducing it. Feature engineering targets accuracy, not throughput. It would be the right choice when predictions are batch-scored offline and richer features improve precision without any real-time constraint.

  • ✗

    Migrate from a decision tree ensemble to a deep neural network.

    Why it's wrong here

    Deep neural networks add inference overhead, raising latency above the 150 ms baseline rather than cutting it below 100 ms. They suit accuracy-critical batch or GPU-backed workloads; here the fix is reducing per-event compute, such as a smaller model or optimised runtime.

About these practice questions

This AI0-001 question is part of Courseiva's 962-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.