Courseiva

AI0-001 AI Infrastructure and Technologies Practice Question

An AI platform team runs inference for an image classifier on a shared GPU node. Multiple model replicas currently load the full model weights into GPU memory independently, and the node runs out of GPU memory when a third replica starts. The team wants to serve more replicas per GPU without changing model accuracy. Which approach best addresses the constraint?

⚠ Common exam trap

The trap here is reaching for a precision change or memory paging to save space, when the actual waste is duplicated weight copies that should be shared.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Serve the replicas from shared GPU memory using a model server that supports weight sharing or a runtime that loads the model once.

The exhaustion comes from duplicate copies of identical weights, so the fix is to load the weights once and let multiple replicas reference them. Model servers and optimized runtimes that support shared weights or a single loaded model instance cut per-replica memory to activations and cache, increasing density without touching precision or accuracy.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Pin each replica to a separate CUDA stream and rely on the driver's scheduler to time-slice GPU memory.

    Why it's wrong here

    CUDA streams control execution ordering of kernels, not memory ownership. Two replicas on different streams still each allocate their own weight buffers, so the out-of-memory condition persists. Streams improve concurrency and overlap but provide no mechanism for sharing a single resident copy of model parameters across replicas.

  • ✓

    Serve the replicas from shared GPU memory using a model server that supports weight sharing or a runtime that loads the model once.

    Why this is correct

    Frameworks such as NVIDIA Triton Inference Server and runtimes like NVIDIA TensorRT-LLM or vLLM allow multiple model instances to share a single copy of the weights in GPU memory, so additional replicas consume only activation and KV-cache space. This raises replica density on the same GPU while leaving the model's numerical precision and accuracy untouched, exactly matching the requirement.

  • ✗

    Convert the model to FP16 and run the replicas with a smaller batch size to fit more instances.

    Why it's wrong here

    FP16 conversion does shrink weight memory and is a legitimate optimization, but it changes numerics and requires validation, and the scenario explicitly forbids altering accuracy. It also does not eliminate the duplicate weight copies that cause the exhaustion. Reducing batch size lowers activation memory but leaves each replica still holding its own full set of weights.

  • ✗

    Enable CUDA Unified Memory so the driver pages weights between host RAM and GPU memory on demand.

    Why it's wrong here

    Unified Memory eases allocation pressure by migrating pages, but inference workloads touch weights constantly, so paging over PCIe introduces severe latency and can thrash. It does not reduce the resident footprint in a way that reliably supports more replicas at production throughput. The scenario requires more replicas without degrading serving performance, which on-demand paging undermines.

About these practice questions

One of 962 original AI0-001 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official CompTIA exam blueprint

This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.