PMLE Serving and Scaling Models • Set 11
PMLE Serving and Scaling Models Practice Test 11 — 15 questions with explanations. Free, no signup.
Your team is serving a large language model on a Vertex AI endpoint using a custom container. You need to reduce inference latency for long prompts while keeping the deployment cost reasonable. The model uses an autoregressive decoder. Which optimization should you implement?
Choose an answer to begin — your selection is scored in the full session.
15 questions · instant feedback and full explanations after every question.