20+ practice questions focused on Preparing and Using Data for Analysis — one of the most tested topics on the Google Professional Data Engineer exam. Each question includes a detailed explanation so you learn why the right answer is correct.
Start Preparing and Using Data for Analysis PracticeA data engineer wants to train a linear regression model in BigQuery ML to predict sales. The training data includes a categorical feature with 1000+ unique values. Which method is most appropriate to handle this feature in the CREATE MODEL statement?
Explanation: In BigQuery ML, high-cardinality categorical features should be handled in a TRANSFORM clause using ML.FEATURE_CROSS (to combine features) or manual hashing (e.g., ML.HASH_BUCKETIZE) to reduce dimensionality before training. This avoids the explosion of one-hot encoded columns that would occur with 1000+ unique values, keeping the model efficient and preventing overfitting.
A company uses Looker to define business logic in LookML. They need to create a new measure that calculates the average order value, defined as total revenue divided by number of orders. Which LookML syntax should they use?
Explanation: Measures in LookML are defined with type and sql expression. The correct syntax for a calculated measure is: measure: avg_order_value { type: average; sql: ${revenue} / ${order_count} ;; }
A data engineer needs to build a feature engineering pipeline using Vertex AI Pipelines. The pipeline should preprocess data, train a model, and deploy it. Which two components are required to define the pipeline? (Choose 2)
Explanation: The Kubeflow Pipelines SDK (option A) is required because Vertex AI Pipelines is built on Kubeflow Pipelines, and you use the KFP SDK (for example, the @dsl.pipeline decorator and dsl.component) to author, compile, and submit the pipeline definition to Vertex AI. TensorFlow Extended (TFX) (option B) is also required here because TFX provides the components and orchestration libraries (such as ExampleGen, Transform, Trainer, and Pusher) used to implement the preprocessing, training, and deployment steps of a feature engineering pipeline. Together, KFP defines the pipeline graph and TFX supplies the ML components that run within it. Vertex AI Feature Store (option C) is an optional managed feature storage/serving service, not a required component for defining a pipeline. Dataflow (option D) is a managed Apache Beam runner for data processing and is not needed to define the pipeline. Cloud Composer (option E) is a managed Apache Airflow orchestration service and is not required to define a Vertex AI Pipeline.
A company uses AutoML Tables to train a classification model. They want to improve model performance by engineering new features from existing timestamp columns. Which three techniques can they apply within AutoML Tables? (Choose 3)
Explanation: Option B is correct because AutoML Tables supports creating derived numeric columns such as the time difference between two timestamp columns, which can expose useful duration or recency signals to the model. Option C is correct because AutoML Tables provides a feature engineering capability that can automatically generate transformations like polynomial features from existing numeric columns, helping the model capture nonlinear relationships. Option E is correct because AutoML Tables allows extracting date/time components such as day of week from timestamp columns through its UI, producing a categorical feature that can improve classification performance. Option A is not correct because manually adding a weekend boolean column is not a built-in AutoML Tables feature-engineering technique; it would require preprocessing outside AutoML Tables. Option D is not correct because AutoML Tables does not support applying SQL UDFs within its training configuration; custom SQL logic is not part of the AutoML Tables feature engineering interface.
A data engineer needs to create a BigQuery ML model for predicting customer churn using a dataset with 10 million rows and 50 features. The dataset is highly imbalanced (5% churn). Which approach should the engineer use to handle class imbalance during model training?
Explanation: BigQuery ML's CREATE MODEL statement supports the CLASS_WEIGHTS option in the training options, which assigns higher weight to the minority class (churn = 1) so the model penalizes misclassification of churners more heavily. This is the native, scalable way to handle class imbalance in BigQuery ML without preprocessing the data. The other approaches either aren't supported natively or require exporting data outside BigQuery ML.
+15 more Preparing and Using Data for Analysis questions available
Practice all Preparing and Using Data for Analysis questions1. Baseline your knowledge
Start with 10 questions to gauge your current understanding of Preparing and Using Data for Analysis. This tells you whether you need a concept refresher or just practice.
2. Review every explanation
For each question — right or wrong — read the full explanation. Understanding why an answer is correct is more valuable than knowing the answer itself.
3. Focus on exam traps
Preparing and Using Data for Analysis questions on the PDE frequently use trap wording. Look for subtle differences in answers that test your precision, not just general knowledge.
4. Reach 80% consistently
Do repeated sessions until you score 80%+ three times in a row. Then move to mixed-mode practice to test cross-topic recall under realistic conditions.
The exact number varies per candidate. Preparing and Using Data for Analysis is tested as part of the Google Professional Data Engineer blueprint. Practicing with targeted Preparing and Using Data for Analysis questions ensures you can handle any format or difficulty that appears.
Yes. Courseiva provides free PDE practice questions across all exam topics and domains. The platform includes topic-based practice, mock exams, missed-question review, bookmarked questions, and readiness tracking — no account required.
Difficulty is subjective, but Preparing and Using Data for Analysis is a high-priority exam concept tested in multiple ways — direct recall, scenario analysis, and command-output interpretation. Consistent practice is the best way to build confidence.
Launch a full Preparing and Using Data for Analysis practice session with instant scoring and detailed explanations.
Start Preparing and Using Data for Analysis Practice →