AI0-001 Machine Learning and Deep Learning Practice Question
A team is deploying a deep learning model for real-time image classification on edge devices with limited computational resources. Which technique would best help reduce model size and inference time without significant accuracy loss?
⚠ Common exam trap
Candidates often mistakenly believe that transfer learning alone reduces model size, but it only reuses weights—the architecture remains unchanged. For resource-constrained edge devices, pruning and quantization are the direct methods for compression and speed optimization.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Model pruning and quantization
Model pruning and quantization directly reduce the number of parameters and the precision of weights (e.g., from 32-bit floats to 8-bit integers), which shrinks the model size and speeds up inference on edge devices. This technique is specifically designed to minimize computational load while preserving accuracy, making it ideal for resource-constrained environments like real-time image classification on edge hardware.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Data augmentation
Why it's wrong here
Data augmentation expands training data to improve generalisation; it changes neither parameter count nor inference latency. It is tempting because it does raise accuracy, and would be the right choice when the problem is overfitting or insufficient labelled training samples rather than constrained edge compute.
- ✓
Model pruning and quantization
Why this is correct
Pruning removes redundant weights and neurons, while quantization reduces numeric precision (for example FP32 to INT8). Combined, they shrink model size and cut inference latency substantially on resource-constrained edge hardware, with accuracy loss typically recoverable through fine-tuning.
- ✗
Transfer learning
Why it's wrong here
Transfer learning reuses a pre-trained network's weights, cutting training data and time, but the deployed model's parameter count and per-inference cost stay unchanged. It suits scarce labelled data or a new domain, not shrinking an already-trained model for constrained edge hardware; pruning or quantisation addresses size and latency directly.
- ✗
Ensemble learning
Why it's wrong here
Ensemble learning combines multiple models, increasing parameter count, memory footprint and inference latency. It is tempting because ensembles reliably raise accuracy, and would be the correct choice when accuracy is the priority and compute, memory and latency budgets are unconstrained.
About these practice questions
Courseiva writes every AI0-001 question from scratch — 962 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.