Courseiva

Generative AI Leader Google Cloud's Generative AI Offerings Practice Question

Network Topology
region=us-central1display-name=my-modelcontainer-image-uri=us-docker.pkg.dev/cloud-aiplatform/prediction/tf2-cpu.2-6:latestartifact-uri=gs://my-bucket/model"

A data scientist runs the above command to upload a model to Vertex AI Model Registry. The model is a TensorFlow 2.6 model trained on tabular data. After deployment to an endpoint, the prediction latency is higher than expected. What is the most likely cause?

⚠ Common exam trap

Candidates often assume GPU acceleration is always the fix for latency issues, but for tabular models, correct model packaging and artifact URI structure are more common pitfalls. Always verify that the artifact URI points to a directory, not a single file.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

The artifact URI points to a single file instead of a directory

When uploading a model to Vertex AI Model Registry, the artifact URI must point to a directory containing the SavedModel or model artifacts, not a single file. If the URI points to a single file, the model may not load properly or may serve with higher latency. For a TensorFlow 2.6 tabular model, GPU acceleration is generally not the primary latency bottleneck; tabular models are typically small and run efficiently on CPU. The container image choice is less likely to be the root cause than an incorrect artifact URI.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    The artifact URI points to a single file instead of a directory

    Why this is correct

    The URI likely points to a directory with SavedModel, which is correct.

  • ✗

    The model should be uploaded with a different display name

    Why it's wrong here

    Display name does not affect latency.

  • ✗

    The container image used is CPU-only, but a GPU-accelerated image would improve latency

    Why it's wrong here

    TensorFlow models served on CPU-only containers execute inference without GPU acceleration, so matrix operations run slower, inflating prediction latency. Switching to a GPU-accelerated container image provides the parallel compute that reduces that latency at the endpoint.

  • ✗

    The model is uploaded to the wrong region

    Why it's wrong here

    Region is specified as us-central1, which is correct.

About these practice questions

This Generative AI Leader question is part of Courseiva's 1,008-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.