Courseiva

AI0-001 AI Infrastructure and Technologies Practice Question

A logistics company runs a route-optimization model on a Kubernetes cluster. During peak hours the inference pods are frequently evicted because the nodes run out of memory, even though average GPU utilization stays below 40 percent. The team wants to reduce evictions without changing the model or adding nodes. Which action best addresses the root cause?

⚠ Common exam trap

The trap here is chasing GPU symptoms when the eviction signal points to host memory pressure, so adding accelerators or replicas leaves the actual cause untouched.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Set explicit CPU and memory requests and limits on the inference containers so the scheduler can place and protect them accurately.

Memory-pressure eviction happens when the kubelet must reclaim node memory, and pods without declared memory requests are the first to be sacrificed because the scheduler never reserved capacity for them. Setting realistic requests and limits makes placement accurate and bounds each pod's consumption, so the node stays below its eviction threshold without changing the model or adding hardware.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Lower the container image size by switching to a distroless base image for the inference service.

    Why it's wrong here

    Image size affects pull time and disk footprint, not steady-state host memory usage during inference. Once the container is running, the resident memory consumed by model weights and request buffers is what stresses the node, so shrinking the image does nothing to prevent the kubelet from evicting pods when memory pressure crosses its threshold.

  • ✗

    Enable horizontal pod autoscaling on GPU utilization so additional replicas start during peak hours.

    Why it's wrong here

    Autoscaling on GPU utilization adds replicas, but each new replica consumes host memory on nodes that are already under pressure, which can accelerate rather than relieve evictions. The scenario also notes that GPU utilization stays low, so a GPU-based scaling metric would not even trigger during the memory-constrained period.

  • ✓

    Set explicit CPU and memory requests and limits on the inference containers so the scheduler can place and protect them accurately.

    Why this is correct

    Evictions driven by node memory pressure usually mean the pods have no memory requests, so the scheduler overcommits the node and the kubelet later reclaims memory by evicting workloads. Declaring realistic requests lets the scheduler reserve capacity and declaring limits bounds each pod, which stops one inference process from consuming memory that other pods depend on and prevents pressure-driven eviction.

  • ✗

    Increase the GPU memory allocation per pod by requesting additional nvidia.com/gpu resources.

    Why it's wrong here

    GPU memory is allocated in whole-device or partitioned units and has no bearing on host memory pressure, which is what triggers eviction. Requesting more GPU resources would reduce the number of schedulable pods and could worsen placement, while leaving the actual cause, unbounded host memory consumption by the inference processes, completely unaddressed.

About these practice questions

This AI0-001 question is part of Courseiva's 962-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official CompTIA exam blueprint

This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.