Courseiva

Generative AI Leader Google Cloud's Generative AI Offerings Practice Question

An organization is using Vertex AI to fine-tune a large language model. They notice training is taking longer than expected and cost is increasing. Which action is most likely to reduce training time and cost without significantly impacting model quality?

⚠ Common exam trap

A common pitfall is assuming increasing batch size or learning rate will speed up training, but this can destabilize training or degrade quality. Mixed-precision training (bfloat16) directly reduces computation without the risk of underflow, making it the optimal choice for Vertex AI.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Enable mixed-precision training (bfloat16)

Mixed-precision training with bfloat16 reduces memory usage and accelerates computation by using half the bits of standard float32, which directly decreases training time and cost on TPUs and modern GPUs. Vertex AI supports bfloat16 natively on TPU v3+ and A100 GPUs, and for many large language models, this precision preserves model quality because bfloat16 retains the same exponent range as float32, avoiding underflow issues common with float16.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Increase the number of training steps

    Why it's wrong here

    More training steps lengthen the run and increase compute charges, worsening both symptoms. It is tempting because extra steps can improve convergence when a model is undertrained, but the stem describes an already costly run, so the lever needed is reducing iterations, not adding them.

  • ✗

    Increase the batch size

    Why it's wrong here

    A larger batch size reduces step count but raises per-step memory, often forcing smaller models or gradient accumulation, which can slow convergence and worsen quality. It suits throughput-bound training on abundant accelerator memory, not fine-tuning where quality must hold.

  • ✗

    Use a higher learning rate

    Why it's wrong here

    Raising the learning rate risks divergence or degraded quality, and does not itself cut the number of passes over data. It is tempting because a larger step size can shorten convergence in some runs, but that trades stability for speed rather than reducing the training workload.

  • ✓

    Enable mixed-precision training (bfloat16)

    Why this is correct

    bfloat16 mixed-precision training halves memory bandwidth and uses tensor cores for faster matrix operations, cutting training time and compute cost. It preserves numerical range well enough that model quality remains largely unaffected, unlike aggressive precision reduction.

About these practice questions

This Generative AI Leader question is part of Courseiva's 1,008-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.