AI0-001 AI Infrastructure and Technologies Practice Question
A hospital's AI team is deploying a real-time patient deterioration prediction model on bedside monitoring devices. The devices have limited RAM (512 MB) and no GPU, and the model must perform inference within 50 ms. The team has a trained TensorFlow model saved as a SavedModel. Which deployment approach best meets these constraints?
⚠ Common exam trap
The trap here is assuming that any optimized runtime like ONNX Runtime automatically solves memory and latency issues without quantization.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Convert the SavedModel to TensorFlow Lite and apply post-training quantization to INT8.
TensorFlow Lite is optimized for edge devices with limited compute and memory. Post-training INT8 quantization reduces the model size by up to 4x and speeds up CPU inference, directly addressing the 512 MB RAM and 50 ms latency constraints. Other options either rely on external servers or fail to optimize the model for the device's limitations.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Convert the SavedModel to TensorFlow Lite and apply post-training quantization to INT8.
Why this is correct
TensorFlow Lite is designed for resource-constrained edge devices, and post-training INT8 quantization reduces model size and memory usage while accelerating inference on CPUs. This directly addresses the 512 MB RAM limit and 50 ms latency requirement without needing a GPU, making it the most suitable approach for bedside monitors.
- ✗
Deploy the SavedModel using TensorFlow Serving on a central GPU server and stream patient data to it.
Why it's wrong here
TensorFlow Serving on a central GPU server introduces network latency and dependency on connectivity, which may violate the 50 ms inference requirement. It also does not leverage the limited on-device resources and adds infrastructure complexity. For bedside monitoring with strict latency and no GPU, a server-based approach is less appropriate.
- ✗
Use the SavedModel directly in a Python script with TensorFlow's default runtime on the bedside device.
Why it's wrong here
The full TensorFlow runtime is heavy and typically requires more than 512 MB of RAM, and inference without optimization may exceed 50 ms on a CPU-only device. This approach ignores the need for model compression and optimized execution, making it unsuitable for the constrained bedside hardware.
- ✗
Convert the model to ONNX and run it with ONNX Runtime on the bedside device without quantization.
Why it's wrong here
ONNX Runtime can be efficient, but without quantization the model size and memory footprint remain large, likely exceeding the 512 MB RAM limit. Additionally, ONNX Runtime's CPU performance may not meet the 50 ms target for a complex model. Quantization or a lighter runtime like TensorFlow Lite is needed.
About these practice questions
One of 962 original AI0-001 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official CompTIA exam blueprint
This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.