Courseiva

Generative AI Leader Fundamentals of Generative AI Practice Question

Which TWO options are best practices for deploying generative AI models on Vertex AI? (Choose two.)

⚠ Common exam trap

Google Cloud often tests the misconception that manual scaling is more reliable or cost-effective than automatic scaling, but in cloud-native environments, automatic scaling is the standard best practice for variable workloads.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Enable automatic scaling to handle variable traffic

Option B is correct because enabling automatic scaling on a Vertex AI Endpoint lets the deployed model add or remove replicas based on traffic and utilization, so it handles variable load efficiently while controlling cost during idle periods. Option C is correct because Vertex AI Model Monitoring continuously tracks prediction inputs and outputs against a baseline and alerts on training-serving skew and prediction drift, which is essential for maintaining generative AI model quality over time. Option A is wrong because disabling logging removes the audit trail and prediction telemetry needed for debugging, compliance, and monitoring, and the cost savings do not justify losing observability. Option D is wrong because manually scaling instances is error-prone and cannot react quickly to variable traffic the way autoscaling can. Option E is wrong because serving a model without optimization (for example, without appropriate machine types, batching, or quantization) wastes resources and degrades latency and throughput.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Disable logging to reduce cost

    Why it's wrong here

    Disabling logging removes the audit trail and monitoring data needed to trace model inputs, outputs and abuse, which Vertex AI deployments require for governance and troubleshooting. It is tempting because log ingestion carries cost, and logging would be the right lever in a purely experimental sandbox where no compliance or observability obligations apply.

  • ✓

    Enable automatic scaling to handle variable traffic

    Why this is correct

    Automatic scaling adjusts the number of serving replicas to match incoming request volume, so generative AI endpoints stay responsive during traffic spikes without over-provisioning. This satisfies the best-practise requirement to handle variable traffic efficiently when deploying models on Vertex AI.

  • ✓

    Use Vertex AI Model Monitoring to detect drift

    Why this is correct

    Vertex AI Model Monitoring tracks prediction inputs and outputs against baselines, alerting when feature distributions or performance drift over time. This detects drift for deployed generative AI models, satisfying the best-practise requirement to maintain reliability after deployment on Vertex AI.

  • ✗

    Manually scale instances based on expected load

    Why it's wrong here

    Manual scaling cannot react to real-time request volume, so the endpoint either over-provisions or drops traffic during spikes; Vertex AI autoscaling adjusts replicas from live metrics. Manual sizing is tempting for predictable, steady workloads, but generative AI traffic is bursty, making metric-driven autoscaling the required approach.

  • ✗

    Serve the model directly without optimization

    Why it's wrong here

    Serving without optimisation ignores Vertex AI's required model deployment configuration, so the endpoint cannot meet latency or throughput targets for production traffic. It is tempting when validating a model quickly, where raw serving is acceptable, but production deployment demands tuned machine types, accelerators and autoscaling before traffic is routed.

About these practice questions

Courseiva writes every Generative AI Leader question from scratch — 1,008 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.