Courseiva

Generative AI Leader · topic practice

Google Cloud's Generative AI Offerings practice questions

This domain covers Google Cloud's generative AI product surface: Vertex AI, Gemini models, Model Garden, and tuning options. Questions are scenario-based, asking you to pick the right model, feature, or configuration for summarization, long-context video analysis, cost-efficient fine-tuning, and securing deployed applications against prompt injection.

Courseiva uses original exam-style practice questions designed for learning and revision. The goal is to understand the concepts, recognise exam patterns, and improve through explanations — not memorise copied exam dumps.

Editorial oversight:Johnson Ajibi· MSc IT Security, IEEE Senior Member
20 questionsDomain: Google Cloud's Generative AI Offerings

What the exam tests

What to know about Google Cloud's Generative AI Offerings

Be able to match a business scenario to the right Vertex AI service, Gemini model capability, or tuning approach, and explain how to secure and improve outputs. The most important thing is choosing the correct feature for the stated constraint, not the biggest model.

Selecting Gemini model variants and context windows for long video or document inputs

Using Vertex AI tuning options like supervised fine-tuning and parameter-efficient methods

Applying Vertex AI grounding, safety filters, and settings to improve output quality

Implementing prompt injection defenses such as input validation and separation of instructions

Watch out for

Common Google Cloud's Generative AI Offerings exam traps

  • ▸Assuming a larger model always fixes incomplete summaries instead of adjusting prompts or grounding
  • ▸Confusing context window size with output quality when choosing a model for long inputs
  • ▸Treating prompt injection as a model bug rather than an application-layer security concern

Practice set

Google Cloud's Generative AI Offerings questions

20 questions · select your answer, then reveal the explanation

A financial services firm uses a fine-tuned Gemini model in Vertex AI for regulatory compliance checks. They notice that token usage is high, increasing costs. They want to reduce costs without sacrificing accuracy. Which approach should they take?

What is the most likely cause of the error?

Network Topology
gcloud ai models uploadregion=us-central1 \display-name=my-model \container-image-uri=gcr.io/cloud-aiplatform/prediction/tf2-cpu.2-12:latest \artifact-uri=gs://my-bucket/model/ \predict-schemata=gs://my-bucket/schema/predict_schema.yamlRefer to the exhibit.```

Why is the model responding in English despite the prompt asking for French translation?

Exhibit

Refer to the exhibit.

```json
{
  "instances": [
    {"content": "Translate to French: Hello, how are you?"}
  ],
  "parameters": {
    "temperature": 0.7,
    "maxOutputTokens": 100,
    "topP": 0.9
  }
}
```

A data scientist sends this request to a Gemini model endpoint and receives a response in English. What is the most likely reason?

A financial services firm needs to deploy a large language model (LLM) for analyzing sensitive client documents. They require the model to run within their Virtual Private Cloud (VPC) with no internet access and must comply with data residency regulations. Which Google Cloud generative AI offering should they use?

A company is using Vertex AI for multimodal generative AI to analyze images and text. They need to ensure that the model's outputs are auditable and can be traced back to the input data. Which feature should they enable?

A team deployed a custom generative AI model using KServe on Google Kubernetes Engine (GKE) with the above configuration. They notice that the model is taking longer than expected to respond. What is the most likely cause?

Exhibit

Refer to the exhibit.

```
# deployment.yaml
apiVersion: serving.kserve.io/v1beta1
kind: InferenceService
metadata:
  name: my-model
spec:
  predictor:
    containers:
      - name: model-container
        image: us-central1-docker.pkg.dev/my-project/my-repo/my-model:latest
        resources:
          limits:
            nvidia.com/gpu: 1
```

You are a machine learning engineer at a healthcare startup. Your team has developed a generative AI model that summarizes patient medical records. The model is deployed on Vertex AI Endpoints using a custom container. You have configured the endpoint with a single n1-standard-4 machine (4 vCPUs, 15 GB memory) without accelerators. The model uses a small transformer architecture. During load testing with 50 concurrent requests, you observe that the average latency is 8 seconds, which exceeds the requirement of 2 seconds. Additionally, some requests time out after 10 seconds. You suspect the CPU is the bottleneck. You also notice that the model inference code uses TensorFlow but is not optimized for inference. Which action should you take to reduce latency?

A data scientist is using Vertex AI Model-as-a-Service (MaaS) to deploy a fine-tuned open-source model. They notice high latency during inference. What is the most likely cause?

A financial services company wants to use Vertex AI Grounding with enterprise data to power a regulatory compliance chatbot. They have strict data residency requirements: data must remain in the EU. What should they do?

You are using Vertex AI Model Garden to deploy a Llama model. Which deployment option provides the best latency for real-time inference?

A team is deploying a real-time chat application using Gemini. They need to ensure the model does not generate harmful content. Which safety filter configuration should they use?

Which TWO components are essential for building a multi-turn conversational agent using Vertex AI Agent Builder? (Choose two.)

Refer to the exhibit. You ran the gcloud command to list a model, but received this error. What is the most likely issue?

Network Topology
gcloud ai models listfilter='name:my_model'Output:

Refer to the exhibit. A developer has defined a dynamic action in the Vertex AI Agent Builder agent YAML. The agent is not triggering the action. What is the most likely issue?

Exhibit

agent:
  display_name: travel_agent
  dynamic_actions:
    - action_name: book_flight
      http_endpoint: https://api.example.com/flights

A company is deploying a chatbot that must ensure customer data remains within the European Union. Which approach should they take?

During a load test, a Vertex AI endpoint serving a large language model experiences high latency and increased error rates. The endpoint is configured with autoscaling. What is the most likely cause?

A data scientist is using Vertex AI Model Registry to manage multiple versions of a custom text classification model. They need to ensure that only the version that passes all evaluation metrics can be deployed to a Vertex AI Endpoint for online predictions. What deployment strategy should they use?

Which TWO actions can help reduce latency for an online prediction endpoint served by a large language model on Vertex AI? (Select TWO.)

Which THREE capabilities does Vertex AI Agent Builder provide out of the box? (Select THREE.)

A machine learning engineer submits the above batch prediction job for a large language model. The job is expected to process 100,000 instances. The job takes much longer than expected. Which change would most likely reduce the execution time?

Exhibit

Refer to the exhibit. {
  "name": "projects/123/locations/us-central1/batchPredictionJobs/bpj456",
  "model": "projects/123/locations/us-central1/models/789",
  "inputConfig": {
    "instancesFormat": "jsonl",
    "gcsSource": {"uris": ["gs://bucket/input.jsonl"]}
  },
  "outputConfig": {
    "predictionsFormat": "jsonl",
    "gcsDestination": {"outputUriPrefix": "gs://bucket/output/"}
  },
  "machineType": "n1-standard-4",
  "batchSize": 64,
  "startingReplicaCount": 1,
  "maxReplicaCount": 1
}

Free account

Track your progress over time

Create a free account to save your results and see which topics improve across sessions.

Focused Google Cloud's Generative AI Offerings sessions

Start a Google Cloud's Generative AI Offerings only practice session

Every question in these sessions is drawn from the Google Cloud's Generative AI Offerings domain — nothing else.

Related practice questions

Related Generative AI Leader topic practice pages

Move into related areas when this topic feels solid.

Frequently asked questions

What does the Generative AI Leader exam test about Google Cloud's Generative AI Offerings?
Be able to match a business scenario to the right Vertex AI service, Gemini model capability, or tuning approach, and explain how to secure and improve outputs. The most important thing is choosing the correct feature for the stated constraint, not the biggest model.
How should I use these practice questions?
Select your answer before revealing the explanation. Then read why each option is right or wrong — this active recall approach builds retention far faster than re-reading notes.
Can I practise just Google Cloud's Generative AI Offerings questions in a focused session?
Yes — the session launcher on this page draws every question from the Google Cloud's Generative AI Offerings domain. Use a 10-question session first to gauge your baseline, then move to 20 or 30 once the weak spots are clear.
Where can I practise other Generative AI Leader topics?
Use the topic links above to move to related areas, or go back to the Generative AI Leader question bank to see all topics.
Are these real exam questions or dumps?
These are original practice questions written to test the same concepts the Generative AI Leader exam covers. They are not copied from any real exam or dump site.