Courseiva

AI0-001 AI Implementation and Operations Practice Question

An e-commerce company operates an AI recommendation service. After a marketing campaign, the operations team notices that inference costs have tripled while request volume has only doubled. They need to reduce cost per inference without degrading recommendation quality. Which two actions should the team take? (Choose two.)

⚠ Common exam trap

The trap here is equating more capacity with lower cost, when the objective is specifically cost per inference rather than raw throughput.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Apply model quantization or a distilled smaller model for the recommendation ranking stage, validating quality against offline metrics.

Cost per inference falls when each execution does more useful work or requires fewer resources. Dynamic batching amortizes execution overhead across concurrent requests, and quantization or distillation reduces the compute needed per prediction. Adding replicas, disabling caching, or upgrading instance size raises capacity or work without improving efficiency, so they do not meet the goal.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Disable caching of recommendation results so every request is computed fresh from the model.

    Why it's wrong here

    Disabling caching increases the number of model executions, raising cost and latency. Caching popular recommendations is a common cost-reduction technique because it avoids recomputation. Turning it off directly contradicts the goal of reducing cost per inference and would likely degrade response times during the campaign surge.

  • ✗

    Move the recommendation service to a larger GPU instance type with more memory.

    Why it's wrong here

    A larger instance may improve throughput for a single node, but it typically carries a higher hourly price and does not by itself reduce cost per inference. Without batching or model optimization, the extra capacity is often underutilized, so this change can increase overall spend rather than lower it.

  • ✓

    Apply model quantization or a distilled smaller model for the recommendation ranking stage, validating quality against offline metrics.

    Why this is correct

    Quantization or distillation reduces compute and memory per inference, directly lowering cost. Validating against offline metrics such as recall at k or NDCG ensures quality stays within acceptable bounds. This is a standard optimization when cost grows faster than traffic, and it complements batching by reducing the work per execution.

  • ✓

    Enable dynamic batching at the inference server so concurrent requests are grouped into a single model execution.

    Why this is correct

    Dynamic batching groups concurrent requests into one model execution, which raises GPU or CPU utilization and reduces cost per request. It is especially effective when request volume grows, because larger batches amortize fixed overhead. Recommendation quality is preserved because each request still receives its own prediction; only the execution is shared.

  • ✗

    Increase the number of replicas behind the load balancer so each instance handles fewer requests.

    Why it's wrong here

    Adding replicas improves latency and throughput headroom but increases total infrastructure cost rather than reducing cost per inference. Since the team's goal is to lower cost while handling more traffic, horizontal scaling in this manner works against the objective unless the underlying inefficiency is fixed first.

About these practice questions

One of 962 original AI0-001 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official CompTIA exam blueprint

This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.