easyMultiple Choice
PMLE Practice Question: A data scientist wants to perform feature…
A data scientist wants to perform feature engineering on a large dataset stored in BigQuery before training a model. Which feature engineering tool is most appropriate?
⚠ Common exam trap
PMLE often tests the misconception that external tools like Dataproc or Dataflow are required for feature engineering on BigQuery data, when in fact BigQuery ML's TRANSFORM clause is designed for this purpose.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use BigQuery ML TRANSFORM clause
BigQuery ML's TRANSFORM clause allows feature engineering to be performed directly within BigQuery using SQL, eliminating the need to export data or build separate pipelines. It automatically applies transformations during model training and prediction, ensuring consistency and efficiency for large datasets. This is the most integrated and appropriate tool for feature engineering on BigQuery data.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use Vertex AI Feature Store to store engineered features
Why it's wrong here
Vertex AI Feature Store serves and shares already-engineered features for training and online serving; it does not perform the transformation itself. BigQuery ML or Dataflow would be chosen when the engineering must run directly over the large BigQuery dataset before training.
- ✗
Export data to Cloud Dataproc for feature engineering
Why it's wrong here
Exporting data to Cloud Dataproc introduces unnecessary I/O overhead and cluster provisioning latency, whereas BigQuery’s native SQL-based feature engineering functions (e.g., ML.FEATURE_CROSS, ML.ONE_HOT_ENCODER) operate directly on the data without moving it. This option is tempting because Dataproc excels at distributed, custom preprocessing with Spark or Hadoop when transformations require non-SQL logic or external libraries, making it correct for complex, iterative feature pipelines that cannot be expressed in BigQuery SQL.
- ✗
Create a Dataflow pipeline to compute features
Why it's wrong here
A Dataflow pipeline is designed for stream or batch processing of data at scale, but the question specifically requires feature engineering *before* training a model, which is a one-time or iterative transformation task best handled by BigQuery ML’s `TRANSFORM` clause or a `CREATE MODEL` statement. It is tempting because Dataflow excels at complex, stateful data transformations for production pipelines, and would be correct if the features needed continuous, low-latency updates from streaming data rather than static preprocessing for a single training run.
- ✓
Use BigQuery ML TRANSFORM clause
Why this is correct
BigQuery ML's TRANSFORM clause applies feature engineering expressions inside BigQuery, so transformations run where the large dataset already resides. This satisfies the constraint of processing data in place without exporting it, and the transformations are automatically applied during training and prediction.
Go deeper
Related to this question
About these practice questions
One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.