Courseiva

Generative AI Leader Fundamentals of Generative AI Practice Question

Network Topology
gcloud ai endpoints listregion=us-central1Output:ENDPOINT_ID: 123456DISPLAY_NAME: my-endpointMODEL: projects/123/locations/us-central1/models/789DEPLOYED_MODELS:MACHINE_TYPE: n1-standard-2ACCELERATOR_TYPE: NVIDIA_TESLA_T4

Refer to the exhibit. A developer executed the command to list endpoints. They notice that two models are deployed to the same endpoint. What is the most likely reason for this configuration?

⚠ Common exam trap

Google Cloud often tests the misconception that deploying two models to the same endpoint is always an error, when in fact it is a deliberate pattern for canary testing or A/B testing with traffic splitting.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

It is a canary deployment with traffic splitting

A is correct because deploying two models to the same endpoint with traffic splitting is a standard canary deployment strategy. In this configuration, a small percentage of inference requests are routed to the new model while the majority go to the stable model, allowing validation of the new model's performance before full rollout. This is commonly supported by Google Cloud's Vertex AI, where you can deploy multiple models to an endpoint and assign traffic percentages to each model variant (e.g., 90% to the stable model and 10% to the canary model).

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    It is a canary deployment with traffic splitting

    Why this is correct

    A canary deployment routes a small slice of live traffic to a new model version while the rest continues to the stable one, so both models must sit behind the same endpoint for traffic splitting to work. This matches the observed configuration without implying an A/B test or blue-green swap.

  • ✗

    The endpoint is misconfigured and will cause conflicts

    Why it's wrong here

    A single endpoint can host multiple model deployments by design, so no conflict arises; the exhibit shows intentional co-location. Misconfiguration is tempting because duplicate deployments look accidental, yet it would be the answer only if the endpoint returned routing errors or failed health checks.

  • ✗

    The models are from different frameworks

    Why it's wrong here

    Endpoints route by deployment name, not framework, so framework origin does not explain co-location. This is tempting because multi-framework serving exists, but that scenario involves separate endpoints or containers, not two models sharing one endpoint's routing configuration.

  • ✗

    It is a batch prediction endpoint

    Why it's wrong here

    Batch endpoints process jobs against stored data, not real-time inference traffic, so they would not appear as a standard online endpoint listing. Batch is tempting because it also serves models, but it is correct only when the workload is asynchronous scoring of large datasets.

About these practice questions

Courseiva writes every Generative AI Leader question from scratch — 1,008 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.