Courseiva

Generative AI Leader Fundamentals of Generative AI Practice Question

A multimodal generative AI system processes both image and text inputs to produce captions. During inference, the image encoder sometimes produces noisy or missing features. Which architectural design decision best handles such input degradation without retraining?

⚠ Common exam trap

Google Cloud often tests the misconception that preprocessing or model capacity adjustments are the only ways to handle input noise, but the key insight is that architectural mechanisms like gating can adaptively handle degradation at inference time without retraining.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Introduce a gating mechanism that learns to weigh image features based on confidence scores from the encoder.

A gating mechanism dynamically adjusts the contribution of image features based on confidence scores from the encoder, allowing the model to gracefully handle noisy or missing features without retraining. This architectural design learns to suppress unreliable image inputs and rely more on text or other modalities, ensuring robust caption generation under input degradation.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Train a separate variational autoencoder to produce a clean latent representation from the noisy image.

    Why it's wrong here

    Training a separate variational autoencoder requires a training phase, so it cannot address degradation at inference without retraining. It is tempting because VAEs genuinely learn clean latent representations, making them the right choice when you can pre-train a denoising model on paired noisy and clean data.

  • ✗

    Increase the image encoder’s capacity to better extract robust features.

    Why it's wrong here

    Increasing encoder capacity requires architectural change and retraining, and larger encoders do not inherently tolerate missing features at inference. It tempts because capacity genuinely improves feature extraction, making it the right choice when accuracy is limited by an underfitting encoder during initial model design.

  • ✗

    Apply standard image preprocessing (e.g., denoising) to all inputs before feeding to the encoder.

    Why it's wrong here

    Preprocessing operates on raw pixel inputs, not on the encoder's internal feature vectors, so it cannot repair noisy or missing features produced downstream. It tempts because denoising is genuinely effective when degradation occurs in the input image itself before encoding.

  • ✓

    Introduce a gating mechanism that learns to weigh image features based on confidence scores from the encoder.

    Why this is correct

    A learned gating mechanism dynamically weights image features by encoder confidence, letting the decoder down-weight noisy or missing inputs at inference. This handles degradation without retraining, satisfying the stem's constraint of robustness to unreliable visual features.

About these practice questions

One of 1,008 original Generative AI Leader practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.