AI0-001 AI Infrastructure and Technologies Practice Question
A data science team is deploying a deep learning model for real-time inference on edge devices with limited power and memory. Which model optimisation technique would be MOST effective for reducing latency and memory footprint while maintaining acceptable accuracy?
⚠ Common exam trap
CompTIA often tests the misconception that increasing model complexity (more layers or epochs) improves deployment performance, when in fact the opposite is true for edge inference; candidates may confuse training optimization with inference optimization.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Apply quantisation to convert weights from FP32 to INT8
Quantization reduces the precision of model weights from 32-bit floating point (FP32) to 8-bit integer (INT8), which directly cuts memory usage by 75% and accelerates inference on edge devices by leveraging integer arithmetic. This technique is specifically designed for resource-constrained environments where power and memory are limited, and it typically preserves accuracy within 1-2% of the original model.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use a larger batch size during inference
Why it's wrong here
Batching raises peak memory and does not reduce per-inference latency on a single edge request; it improves throughput on server GPUs where requests queue. Edge real-time inference processes one sample at a time, so quantisation or pruning shrinks the model itself. Larger batches also add buffering delay.
- ✗
Train the model for more epochs to improve convergence
Why it's wrong here
Extra epochs alter training convergence only; the deployed model's parameter count, memory footprint and inference latency stay identical. Longer training suits improving accuracy when underfitting, not meeting edge resource limits. The stem asks for optimisation of the deployed artefact, so quantisation or pruning applies.
- ✓
Apply quantisation to convert weights from FP32 to INT8
Why this is correct
Quantisation converts FP32 weights to INT8, cutting memory footprint roughly fourfold and enabling integer arithmetic that executes faster on edge hardware, directly satisfying the limited power and memory constraint. Accuracy loss stays acceptable because scaling factors preserve the weight distribution, making it ideal for real-time inference on constrained devices.
- ✗
Increase the number of layers to improve feature extraction
Why it's wrong here
Adding layers increases parameter count, memory footprint and inference latency, directly worsening the edge constraints. Depth is chosen when accuracy on complex features is the priority and compute is plentiful, such as server-side training. The scenario demands pruning, quantisation or distillation instead.
About these practice questions
Courseiva writes every AI0-001 question from scratch — 962 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.