AI0-001 AI Implementation and Operations Practice Question
A company is deploying a fraud detection model that must return predictions within 100ms to avoid transaction delays. The team is deciding between batch and real-time inference. Which factor most strongly supports a real-time inference architecture?
⚠ Common exam trap
CompTIA often tests the misconception that batch inference is always cheaper or more efficient, but the trap here is that latency requirements (under 100ms) force a real-time architecture regardless of cost or data volume.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
The application requires immediate feedback for each transaction
Real-time inference is required when the application must return predictions within strict latency bounds (e.g., 100ms) to avoid transaction delays. The need for immediate feedback per transaction directly aligns with a real-time architecture, where each request is processed individually as it arrives, rather than waiting for a batch window. Batch inference would introduce unacceptable latency because it processes groups of records on a schedule, not on-demand.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
The model requires large amounts of historical data for each prediction
Why it's wrong here
Historical data volume concerns feature engineering and training, not serving latency; batch jobs handle large lookbacks comfortably. Real-time inference is justified when each request must be scored within a strict deadline, as with the 100ms transaction requirement here.
- ✓
The application requires immediate feedback for each transaction
Why this is correct
Immediate per-transaction feedback demands synchronous scoring, which only real-time inference provides; batch processing defers predictions until a scheduled job runs, breaching the 100ms constraint. Real-time endpoints score each transaction on arrival, satisfying the latency requirement that the fraud detection scenario specifies.
- ✗
The infrastructure budget is limited and must be optimized
Why it's wrong here
Real-time serving typically costs more, since always-on endpoints and autoscaling consume compute continuously, so budget pressure argues against it. Cost optimisation favours scheduled batch runs; real-time is chosen when per-request latency deadlines, such as 100ms, must be met.
- ✗
The model can be retrained weekly using gathered data
Why it's wrong here
Weekly retraining is a training-cadence concern and is orthogonal to inference mode; batch pipelines retrain just as readily. Real-time inference is required by the 100ms per-transaction deadline, not by how often the model is refreshed.
About these practice questions
One of 962 original AI0-001 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.