An e-commerce company operates an AI recommendation service. After a marketing campaign, the operations team notices that inference costs have tripled while request volume has only doubled. They need to reduce cost per inference without degrading recommendation quality. Which two actions should the team take? (Choose two.)
Quantization or distillation reduces compute and memory per inference, directly lowering cost. Validating against offline metrics such as recall at k or NDCG ensures quality stays within acceptable bounds. This is a standard optimization when cost grows faster than traffic, and it complements batching by reducing the work per execution.
Why this answer
Cost per inference falls when each execution does more useful work or requires fewer resources. Dynamic batching amortizes execution overhead across concurrent requests, and quantization or distillation reduces the compute needed per prediction. Adding replicas, disabling caching, or upgrading instance size raises capacity or work without improving efficiency, so they do not meet the goal.
Exam trap
The trap here is equating more capacity with lower cost, when the objective is specifically cost per inference rather than raw throughput.