mediumMultiple Choice
PMLE Practice Question: A retailer uses BigQuery ML to build a linear…
A retailer uses BigQuery ML to build a linear regression model for sales forecasting. The model's evaluation shows high RMSE. Which step should they take first?
⚠ Common exam trap
Google Cloud often tests the misconception that high RMSE is always a model complexity issue, leading candidates to jump to advanced algorithms or feature engineering without considering fundamental data quality checks.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Examine the data for outliers and missing values
High RMSE in a linear regression model often indicates issues with data quality, such as outliers or missing values, which can disproportionately skew the model's predictions. BigQuery ML's linear regression is sensitive to such anomalies, so examining and cleaning the data is the most appropriate first step before considering model complexity or feature engineering.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use a more complex model like XGBoost
Why it's wrong here
Switching to XGBoost changes the algorithm before diagnosing why linear regression underfits or misfits, and adds tuning burden without evidence. XGBoost suits non-linear relationships and mixed feature types; it is the right choice once diagnostics show the linear form cannot capture the pattern.
- ✗
Increase the number of features
Why it's wrong here
Adding features before diagnosing the error inflates dimensionality and risks overfitting without addressing the underlying cause. Feature expansion is correct when residual analysis shows the current inputs omit a genuine predictor, not as a first response to high RMSE.
- ✗
Set a larger training budget
Why it's wrong here
Increasing the training budget extends iteration counts but cannot reduce RMSE caused by an inappropriate model class: linear regression cannot capture the non-linear, seasonal patterns in retail sales. It is tempting because larger budgets genuinely help when a suitable model is underfitted, such as tuning boosted trees or neural networks that already fit the data's structure.
- ✓
Examine the data for outliers and missing values
Why this is correct
High RMSE often stems from data quality problems, so inspecting for outliers and missing values addresses the root cause before tuning the model or adding features. This diagnostic step is the logical first action when evaluation error is unexpectedly large.
Go deeper
Related to this question
About these practice questions
One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.