Courseiva

AI0-001 AI Infrastructure and Technologies Practice Question

A data science team is deploying a real-time fraud detection model on edge devices in retail stores. The model must infer under 10ms and fit within 50MB memory. Which combination of techniques should the team apply?

⚠ Common exam trap

AI0-001 often tests the misconception that more hardware parallelism or larger batches solve latency problems, when edge constraints actually demand model compression techniques like quantization and pruning.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Model quantization to INT8 and pruning of low-weight connections

INT8 quantization reduces each weight from 32-bit float to 8-bit integer, cutting model size ~4x and enabling faster integer arithmetic that meets the sub-10ms latency target. Pruning removes low-magnitude weight connections, further shrinking the 50MB footprint and reducing compute. Together they are the standard edge-optimization pairing for latency- and memory-constrained inference.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Model parallelism and distributed inference

    Why it's wrong here

    Model parallelism splits one model across multiple devices and distributed inference spreads requests across a cluster, both targeting throughput and large-model capacity, not single-device latency or footprint. They suit data-centre serving of models too large for one accelerator, whereas this scenario needs compression and quantisation on one constrained device.

  • ✗

    Increase batch size and use FP16 precision

    Why it's wrong here

    Larger batches raise throughput but increase per-request latency, and FP16 halves weight size without shrinking the architecture enough to meet 50MB. Batching and half precision belong in server-side throughput tuning; the edge constraint demands pruning, distillation or quantisation to cut parameters and memory.

  • ✗

    Train a larger model and use distillation to transfer knowledge

    Why it's wrong here

    Distillation compresses a trained model, but the stem demands under 10ms inference within 50MB; distillation alone does not guarantee that footprint or latency, since the student model's size and speed depend on architecture choices. It is tempting because distillation genuinely transfers knowledge from a large teacher to a smaller student, and would be correct when accuracy must be preserved while shrinking an already-trained model.

  • ✓

    Model quantization to INT8 and pruning of low-weight connections

    Why this is correct

    Quantisation to INT8 shrinks weights from 32-bit to 8-bit, cutting memory roughly fourfold, while pruning removes low-weight connections to reduce computation. Together they meet the 50MB footprint and sub-10ms inference latency constraints on constrained edge hardware.

About these practice questions

This AI0-001 question is part of Courseiva's 962-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official CompTIA exam blueprint

This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.