Courseiva
mediumMultiple Select

Generative AI Leader Practice Question: A data scientist wants to improve the performance…

A data scientist wants to improve the performance of a text classification model for customer feedback. They have a small labeled dataset of 500 examples and a large unlabeled corpus of 100,000 feedback messages. Which TWO strategies would be most effective? (Choose 2)

⚠ Common exam trap

A common misconception is that RAG can improve model training, but RAG retrieves information during inference and does not augment the training data. In contrast, pseudo-labeling directly expands the training set.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Apply semi-supervised learning by pseudo-labeling the unlabeled data

Option B is correct because semi-supervised learning via pseudo-labeling leverages the large unlabeled corpus: the model is first trained on the 500 labeled examples, then used to predict labels on the 100,000 unlabeled messages, and high-confidence predictions are added back as training data, effectively expanding the labeled set and improving classification performance. Option E is correct because using a pre-trained LLM such as Gemini and fine-tuning it on the small labeled dataset exploits transfer learning: the model already encodes broad language understanding from large-scale pre-training, so only a small amount of task-specific labeled data is needed to adapt it to customer feedback classification, which is far more sample-efficient than training from scratch. Option A is not appropriate because increasing the context window affects how much input text the model can attend to, not the model's ability to learn from limited labels, and does not address the small-labeled-data problem. Option C is not appropriate because RAG retrieves relevant documents to augment generation at inference time; it does not train or improve the classification model's parameters and is not a training strategy for a text classifier. Option D is not appropriate because training a model from scratch on only 500 labeled examples will almost certainly overfit and underperform compared to transfer learning or semi-supervised approaches.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Increase the context window of the model

    Why it's wrong here

    Increasing the context window changes how much text the model attends to per input; it does not exploit the 100,000 unlabeled messages to improve classification accuracy. Context expansion suits tasks needing longer documents, not a small labelled set with abundant unlabelled data.

  • ✓

    Apply semi-supervised learning by pseudo-labeling the unlabeled data

    Why this is correct

    Pseudo-labeling assigns predicted labels to the 100,000 unlabeled messages, then retrains on the combined set. This exploits the large unlabeled corpus to compensate for the 500-example labelled set, expanding effective training signal without new annotation.

  • ✗

    Use RAG to retrieve similar examples from the unlabeled corpus during inference

    Why it's wrong here

    RAG retrieves passages to ground generation at inference; it does not update model weights, so classification accuracy on customer feedback does not improve. RAG suits knowledge-intensive question answering over a document corpus, not fine-tuning a classifier with unlabelled data.

  • ✗

    Train a model from scratch on the labeled data only

    Why it's wrong here

    Training from scratch on 500 labelled examples yields poor generalisation and discards the 100,000 unlabeled messages entirely. This approach suits large, task-specific labelled corpora where pretrained representations offer no advantage, not a data-scarce classification problem.

  • ✓

    Use a pre-trained LLM (e.g., Gemini) and fine-tune on the labeled data

    Why this is correct

    Fine-tuning a pre-trained LLM leverages representations already learned from vast text corpora, so the 500 labelled examples need only adapt the model rather than teach language from scratch. This directly addresses the stem's constraint of a small labelled dataset paired with abundant unlabelled text, where training from scratch would overfit.

About these practice questions

Courseiva writes every Generative AI Leader question from scratch — 1,008 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.