Courseiva

AI0-001 AI Infrastructure and Technologies Practice Question

An MLOps team wants to deploy a trained PyTorch model to production with low latency inference. The model must be interoperable across different frameworks and runtimes. Which approach is BEST?

⚠ Common exam trap

CompTIA often tests the misconception that framework-native serving (TorchServe, TensorFlow Serving) is the best path for low latency, ignoring the explicit requirement for cross-framework interoperability that ONNX uniquely satisfies.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Export the model to ONNX format and deploy using ONNX Runtime

ONNX (Open Neural Network Exchange) provides a standardized, framework-agnostic format that ensures interoperability across different runtimes and hardware accelerators. By exporting the PyTorch model to ONNX and deploying with ONNX Runtime, the team achieves low-latency inference through graph optimizations and hardware-specific execution providers, while avoiding vendor lock-in.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Deploy the native PyTorch model using TorchServe

    Why it's wrong here

    TorchServe serves PyTorch's native format, so the artefact stays tied to the PyTorch runtime and cannot load in other frameworks. It is tempting because TorchServe gives low-latency PyTorch inference with minimal conversion effort, but the interoperability requirement demands a framework-neutral format such as ONNX.

  • ✗

    Quantize the model to INT8 and deploy as a TensorFlow Lite model

    Why it's wrong here

    Quantising to INT8 and converting to TensorFlow Lite changes the runtime away from PyTorch and targets mobile/edge delegates, not cross-framework server interoperability. It is tempting because INT8 cuts inference latency, but the requirement is framework-agnostic portability, which ONNX provides.

  • ✗

    Convert the model to TensorFlow SavedModel and deploy using TensorFlow Serving

    Why it's wrong here

    Converting to TensorFlow SavedModel ties the artefact to TensorFlow Serving, and PyTorch-to-TensorFlow conversion is lossy for many operators. It is tempting when the estate already runs TensorFlow Serving, but interoperability across frameworks and runtimes requires the ONNX open standard instead.

  • ✓

    Export the model to ONNX format and deploy using ONNX Runtime

    Why this is correct

    ONNX provides a framework-agnostic serialisation format, so the exported graph runs identically under ONNX Runtime regardless of the original PyTorch training stack. This satisfies the interoperability constraint across frameworks and runtimes while ONNX Runtime's optimised execution kernels deliver the required low-latency inference.

About these practice questions

This AI0-001 question is part of Courseiva's 962-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.