Courseiva
mediumMultiple Choice

PMLE Practice Question: A retailer uses BigQuery ML to build a linear…

A retailer uses BigQuery ML to build a linear regression model for sales forecasting. The model's evaluation shows high RMSE. Which step should they take first?

⚠ Common exam trap

Google Cloud often tests the misconception that high RMSE is always a model complexity issue, leading candidates to jump to advanced algorithms or feature engineering without considering fundamental data quality checks.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Examine the data for outliers and missing values

High RMSE in a linear regression model often indicates issues with data quality, such as outliers or missing values, which can disproportionately skew the model's predictions. BigQuery ML's linear regression is sensitive to such anomalies, so examining and cleaning the data is the most appropriate first step before considering model complexity or feature engineering.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Use a more complex model like XGBoost

    Why it's wrong here

    Switching to XGBoost changes the algorithm before diagnosing why linear regression underfits or misfits, and adds tuning burden without evidence. XGBoost suits non-linear relationships and mixed feature types; it is the right choice once diagnostics show the linear form cannot capture the pattern.

  • ✗

    Increase the number of features

    Why it's wrong here

    Adding features before diagnosing the error inflates dimensionality and risks overfitting without addressing the underlying cause. Feature expansion is correct when residual analysis shows the current inputs omit a genuine predictor, not as a first response to high RMSE.

  • ✗

    Set a larger training budget

    Why it's wrong here

    Increasing the training budget extends iteration counts but cannot reduce RMSE caused by an inappropriate model class: linear regression cannot capture the non-linear, seasonal patterns in retail sales. It is tempting because larger budgets genuinely help when a suitable model is underfitted, such as tuning boosted trees or neural networks that already fit the data's structure.

  • ✓

    Examine the data for outliers and missing values

    Why this is correct

    High RMSE often stems from data quality problems, so inspecting for outliers and missing values addresses the root cause before tuning the model or adding features. This diagnostic step is the logical first action when evaluation error is unexpectedly large.

About these practice questions

One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.