Courseiva

PDE · topic practice

Preparing and Using Data for Analysis practice questions

This domain covers ingesting, cleaning, transforming, and preparing data for ML and analytics on Google Cloud. Expect questions on BigQuery, Dataflow, Dataprep, Dataproc, Vertex AI Feature Store, and LookML modeling, testing whether you can choose the right tool and configure pipelines for scalable, correct feature engineering.

Courseiva uses original exam-style practice questions designed for learning and revision. The goal is to understand the concepts, recognise exam patterns, and improve through explanations — not memorise copied exam dumps.

Editorial oversight:Johnson Ajibi· MSc IT Security, IEEE Senior Member
20 questionsDomain: Preparing and Using Data for Analysis

What the exam tests

What to know about Preparing and Using Data for Analysis

Be able to choose the right Google Cloud service for each data preparation task and configure it correctly. The most important thing is matching the tool to the requirement: BigQuery for SQL, Dataflow for pipelines, Dataprep for visual cleaning, Vertex AI for ML features, and LookML for semantic models.

Selecting BigQuery, Dataflow, Dataproc, or Dataprep for ingestion and transformation at scale

Building Vertex AI preprocessing pipelines with scaling, one-hot encoding, and missing-value handling

Using Vertex AI Workbench notebooks and AutoML Tables for model training and low-latency deployment

Defining LookML views, dimensions, and measures over BigQuery tables for analytics

Watch out for

Common Preparing and Using Data for Analysis exam traps

  • ▸Confusing Dataflow (streaming/batch pipelines) with Dataprep (visual, serverless data wrangling) when the question emphasizes code-free preparation.
  • ▸Assuming AutoML Tables handles all preprocessing automatically instead of recognizing when custom TensorFlow preprocessing or Vertex AI Feature Store is needed.
  • ▸Mixing up LookML view files (table and dimensions) with model files or explores when asked which object defines a table.

Practice set

Preparing and Using Data for Analysis questions

20 questions · select your answer, then reveal the explanation

A data engineer wants to train a linear regression model in BigQuery ML to predict sales. The training data includes a categorical feature with 1000+ unique values. Which method is most appropriate to handle this feature in the CREATE MODEL statement?

A company uses Looker to define business logic in LookML. They need to create a new measure that calculates the average order value, defined as total revenue divided by number of orders. Which LookML syntax should they use?

A data engineer needs to build a feature engineering pipeline using Vertex AI Pipelines. The pipeline should preprocess data, train a model, and deploy it. Which two components are required to define the pipeline? (Choose 2)

A company uses AutoML Tables to train a classification model. They want to improve model performance by engineering new features from existing timestamp columns. Which three techniques can they apply within AutoML Tables? (Choose 3)

A data engineer needs to create a BigQuery ML model for predicting customer churn using a dataset with 10 million rows and 50 features. The dataset is highly imbalanced (5% churn). Which approach should the engineer use to handle class imbalance during model training?

A data engineer is building a production ML pipeline on Vertex AI. The pipeline must preprocess features (e.g., scaling, encoding) and then train a model. The preprocessing logic must be reusable for serving predictions. Which Vertex AI component should they use?

A data engineer needs to implement data quality rules and governance policies across multiple data lakes in GCP. They want to automatically discover and catalog data assets, and enforce row-level security. Which two services should they use? (Select TWO)

A company wants to use BigQuery ML to build a recommendation system for movies. The data includes user IDs, movie IDs, and ratings. Which BigQuery ML model types are suitable for this? (Select TWO)

You are building a real-time fraud detection system using BigQuery streaming and a BQML logistic regression model. The model must be retrained every hour with new labeled data. What is the MOST cost-effective approach to serve predictions with low latency?

You need to analyze customer churn and want to understand the rank of each customer's churn probability within their subscription plan. Which BigQuery window function computes the relative ranking from 1 (highest probability) to N?

You are building a machine learning pipeline for credit risk assessment. The dataset has a severe class imbalance (1% default rate). You want to use AutoML Tables on Vertex AI. Which strategy should you incorporate to handle imbalance?

You need to build a Looker model that joins multiple tables from BigQuery. Which LookML object defines the relationship between tables?

You want to query data across Google Cloud and AWS using a single SQL interface without moving data. Which TWO services can you use?

You have a BigQuery table 'orders' with columns order_id, customer_id, order_amount, and order_date. You need to rank customers by total spend per month, assigning the rank 1 to the highest spender. Which SQL function should you use in a window clause?

You need to preprocess tabular data for training a classification model using Vertex AI. The dataset has missing values in numerical columns and categorical columns with high cardinality. Which Vertex AI service provides automated feature engineering and preprocessing as part of the pipeline?

A retailer wants to use machine learning to predict customer churn based on transaction history and demographic data. The dataset has 500 features, many of which are correlated. The data is highly imbalanced: only 2% churn. They need to deploy a model that provides feature importance and is interpretable. Which model type should they use in BigQuery ML?

Your team uses Looker Studio to build dashboards on top of BigQuery. The dashboards are slow when filtering on a high-cardinality dimension (e.g., user ID). You want to improve performance without changing the underlying BigQuery table design. Which action should you take?

You need to track data lineage from a BigQuery table through a series of transformations and into a Vertex AI model training pipeline. Which Google Cloud service provides automated data lineage tracking?

You need to select two BigQuery features that improve query performance by reducing the amount of data read. Which two options accomplish this? (Choose TWO)

You are building a time-series forecasting model with BigQuery ML. Which three steps should you perform to properly split the data and evaluate the model? (Choose THREE)

Free account

Track your progress over time

Create a free account to save your results and see which topics improve across sessions.

Focused Preparing and Using Data for Analysis sessions

Start a Preparing and Using Data for Analysis only practice session

Every question in these sessions is drawn from the Preparing and Using Data for Analysis domain — nothing else.

Related practice questions

Related PDE topic practice pages

Move into related areas when this topic feels solid.

Frequently asked questions

What does the PDE exam test about Preparing and Using Data for Analysis?
Be able to choose the right Google Cloud service for each data preparation task and configure it correctly. The most important thing is matching the tool to the requirement: BigQuery for SQL, Dataflow for pipelines, Dataprep for visual cleaning, Vertex AI for ML features, and LookML for semantic models.
How should I use these practice questions?
Select your answer before revealing the explanation. Then read why each option is right or wrong — this active recall approach builds retention far faster than re-reading notes.
Can I practise just Preparing and Using Data for Analysis questions in a focused session?
Yes — the session launcher on this page draws every question from the Preparing and Using Data for Analysis domain. Use a 10-question session first to gauge your baseline, then move to 20 or 30 once the weak spots are clear.
Where can I practise other PDE topics?
Use the topic links above to move to related areas, or go back to the PDE question bank to see all topics.
Are these real exam questions or dumps?
These are original practice questions written to test the same concepts the PDE exam covers. They are not copied from any real exam or dump site.