Courseiva

PDE Preparing and Using Data for Analysis Practice Question

A data scientist is training a binary classification model on an imbalanced dataset (95% negative, 5% positive) using AutoML Tables. Which strategy should they use to handle the class imbalance?

⚠ Common exam trap

PDE often tests whether candidates know that AutoML Tables has a built-in weight column for class imbalance, rather than assuming external techniques like SMOTE or duplication are required.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Specify a weight column with higher weights for positive examples in the dataset.

AutoML Tables supports a weight column that lets you assign higher importance to specific rows during training. By giving positive examples (the 5% minority class) higher weights, the model's loss function penalizes misclassification of the minority class more heavily, effectively rebalancing the learning signal without altering the dataset. This is the native, supported mechanism in AutoML Tables for handling class imbalance.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Set the budget to a higher value to allow more training on minority class.

    Why it's wrong here

    A larger budget only extends the search over architectures and hyperparameters; AutoML Tables still optimises on the skewed distribution, so the minority class remains underweighted. Raising the budget is tempting when validation loss is still improving, but it addresses training time, not the 95/5 class ratio.

  • ✗

    Use SMOTE in a Dataflow pipeline before importing the data to AutoML Tables.

    Why it's wrong here

    SMOTE generates synthetic minority samples in a Dataflow pipeline, but AutoML Tables applies its own internal weighting and preprocessing, so the injected distribution is not preserved as intended. SMOTE is tempting because it is the standard remedy for imbalance in custom training pipelines, not managed AutoML.

  • ✓

    Specify a weight column with higher weights for positive examples in the dataset.

    Why this is correct

    A weight column lets AutoML Tables apply per-row loss multipliers, so positive examples at 5% prevalence can contribute proportionally more during training. This directly addresses the stem's imbalance constraint without resampling, preserving all 95% negative rows while preventing the model from defaulting to the majority class.

  • ✗

    Create duplicate copies of the positive class rows to balance the dataset.

    Why it's wrong here

    Duplicating positive rows inflates the same records, encouraging overfitting to those exact examples rather than teaching the model the minority pattern. Duplication is tempting because it is a quick way to equalise counts, but AutoML Tables provides a built-in weight column that reweights classes without altering the underlying data.

About these practice questions

This PDE question is part of Courseiva's 747-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.