Courseiva

AI0-001 AI Infrastructure and Technologies Practice Question

A hospital's AI team is deploying a real-time patient deterioration prediction model on bedside monitoring devices. The devices have limited RAM (512 MB) and no GPU, and the model must perform inference within 50 ms. The team has a trained TensorFlow model saved as a SavedModel. Which deployment approach best meets these constraints?

⚠ Common exam trap

The trap here is assuming that any optimized runtime like ONNX Runtime automatically solves memory and latency issues without quantization.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Convert the SavedModel to TensorFlow Lite and apply post-training quantization to INT8.

TensorFlow Lite is optimized for edge devices with limited compute and memory. Post-training INT8 quantization reduces the model size by up to 4x and speeds up CPU inference, directly addressing the 512 MB RAM and 50 ms latency constraints. Other options either rely on external servers or fail to optimize the model for the device's limitations.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Convert the SavedModel to TensorFlow Lite and apply post-training quantization to INT8.

    Why this is correct

    TensorFlow Lite is designed for resource-constrained edge devices, and post-training INT8 quantization reduces model size and memory usage while accelerating inference on CPUs. This directly addresses the 512 MB RAM limit and 50 ms latency requirement without needing a GPU, making it the most suitable approach for bedside monitors.

  • ✗

    Deploy the SavedModel using TensorFlow Serving on a central GPU server and stream patient data to it.

    Why it's wrong here

    TensorFlow Serving on a central GPU server introduces network latency and dependency on connectivity, which may violate the 50 ms inference requirement. It also does not leverage the limited on-device resources and adds infrastructure complexity. For bedside monitoring with strict latency and no GPU, a server-based approach is less appropriate.

  • ✗

    Use the SavedModel directly in a Python script with TensorFlow's default runtime on the bedside device.

    Why it's wrong here

    The full TensorFlow runtime is heavy and typically requires more than 512 MB of RAM, and inference without optimization may exceed 50 ms on a CPU-only device. This approach ignores the need for model compression and optimized execution, making it unsuitable for the constrained bedside hardware.

  • ✗

    Convert the model to ONNX and run it with ONNX Runtime on the bedside device without quantization.

    Why it's wrong here

    ONNX Runtime can be efficient, but without quantization the model size and memory footprint remain large, likely exceeding the 512 MB RAM limit. Additionally, ONNX Runtime's CPU performance may not meet the 50 ms target for a complex model. Quantization or a lighter runtime like TensorFlow Lite is needed.

About these practice questions

One of 962 original AI0-001 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official CompTIA exam blueprint

This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.