AI0-001 AI Infrastructure and Technologies Practice Question
A data science team is deploying a real-time fraud detection model on edge devices in retail stores. The model must infer under 10ms and fit within 50MB memory. Which combination of techniques should the team apply?
⚠ Common exam trap
AI0-001 often tests the misconception that more hardware parallelism or larger batches solve latency problems, when edge constraints actually demand model compression techniques like quantization and pruning.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Model quantization to INT8 and pruning of low-weight connections
INT8 quantization reduces each weight from 32-bit float to 8-bit integer, cutting model size ~4x and enabling faster integer arithmetic that meets the sub-10ms latency target. Pruning removes low-magnitude weight connections, further shrinking the 50MB footprint and reducing compute. Together they are the standard edge-optimization pairing for latency- and memory-constrained inference.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Model parallelism and distributed inference
Why it's wrong here
Model parallelism splits one model across multiple devices and distributed inference spreads requests across a cluster, both targeting throughput and large-model capacity, not single-device latency or footprint. They suit data-centre serving of models too large for one accelerator, whereas this scenario needs compression and quantisation on one constrained device.
- ✗
Increase batch size and use FP16 precision
Why it's wrong here
Larger batches raise throughput but increase per-request latency, and FP16 halves weight size without shrinking the architecture enough to meet 50MB. Batching and half precision belong in server-side throughput tuning; the edge constraint demands pruning, distillation or quantisation to cut parameters and memory.
- ✗
Train a larger model and use distillation to transfer knowledge
Why it's wrong here
Distillation compresses a trained model, but the stem demands under 10ms inference within 50MB; distillation alone does not guarantee that footprint or latency, since the student model's size and speed depend on architecture choices. It is tempting because distillation genuinely transfers knowledge from a large teacher to a smaller student, and would be correct when accuracy must be preserved while shrinking an already-trained model.
- ✓
Model quantization to INT8 and pruning of low-weight connections
Why this is correct
Quantisation to INT8 shrinks weights from 32-bit to 8-bit, cutting memory roughly fourfold, while pruning removes low-weight connections to reduce computation. Together they meet the 50MB footprint and sub-10ms inference latency constraints on constrained edge hardware.
About these practice questions
This AI0-001 question is part of Courseiva's 962-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official CompTIA exam blueprint
This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.