Courseiva

PDE · topic practice

Scenario practice questions

Practise Google Professional Data Engineer Scenario practice questions — original exam-style scenarios with answer choices, explanations, and analysis of common mistakes.

Courseiva uses original exam-style practice questions designed for learning and revision. The goal is to understand the concepts, recognise exam patterns, and improve through explanations — not memorise copied exam dumps.

Editorial oversight:Johnson Ajibi· MSc IT Security, IEEE Senior Member
17 questionsDomain: Scenario

What the exam tests

What to know about Scenario

Scenario questions test whether you can apply the concept in context, not just recognise a definition.

How the topic appears in realistic exam-style scenarios.

Which detail in the question changes the correct answer.

How to eliminate plausible but wrong options.

How to connect the question back to the wider exam objective.

Watch out for

Common Scenario exam traps

  • ▸Answering from memory before reading the full scenario.
  • ▸Missing a constraint such as cost, availability, security, scope or command context.
  • ▸Choosing a broad answer when the question asks for the most specific fix.
  • ▸Ignoring why the wrong options are tempting.

Practice set

Scenario questions

17 questions · select your answer, then reveal the explanation

Question 1hardmultiple choice
Read the full Scenario explanation →

You are optimizing a BigQuery query that scans 1 TB of data every day. The query joins a large fact table (partitioned by date) with a small dimension table. You notice that the query always scans the entire fact table, even though you only need the last 7 days of data. Which optimization will MOST reduce the bytes scanned?

Question 2mediummulti select
Read the full Scenario explanation →

You are designing a Dataflow pipeline for processing real-time clickstream data. The pipeline must group events into 30-second windows and handle late data up to 5 minutes. You want to output partial results every 10 seconds for low-latency monitoring. Which THREE configurations should you use? (Choose three.)

Question 3mediummulti select
Read the full Scenario explanation →

A company is migrating an on-premises PostgreSQL database to Google Cloud. They need a fully managed database that is compatible with PostgreSQL and can handle both transactional and analytical workloads with high performance. Which two database services meet these requirements? (Choose TWO.)

Question 4mediummulti select
Read the full Scenario explanation →

You are building a data pipeline that ingests data from on-premises into Cloud Storage, then processes it with Dataproc, and finally loads into BigQuery. You need to schedule the pipeline to run daily. The pipeline must handle occasional failures gracefully. Which THREE Google Cloud services should you use together to achieve this? (Choose 3)

Question 5hardmultiple choice
Read the full Scenario explanation →

A retail company ingests point-of-sale clickstream events into Cloud Pub/Sub at roughly 200,000 messages per second during flash sales. Analysts need near-real-time dashboards that aggregate revenue by product category over sliding 5-minute windows, with results visible in BigQuery within 30 seconds of the event. The pipeline must handle occasional bursts up to 3x the normal rate without dropping messages, and the team wants to minimize operational overhead. Which design should the data engineer use?

Question 6hardmultiple choice
Read the full Scenario explanation →

A logistics company ingests GPS pings from delivery vans into Pub/Sub, and a Dataflow streaming pipeline writes them to BigQuery. Latency requirements are lenient (about 5 minutes), but the finance team needs the pipeline's cost to be predictable and low, and the data volume fluctuates by a factor of ten between day and night. The team wants to minimize per-element cost without losing data. Which configuration should the data engineer choose?

Question 7easymultiple choice
Read the full Scenario explanation →

A company is designing a streaming data pipeline to process real-time clickstream events. They need to aggregate events by session window with a 5-minute gap and enable exactly-once processing semantics. Which Google Cloud service should they use?

Question 8hardmultiple choice
Read the full Scenario explanation →

A data engineer manages a Cloud Composer 2 environment. A DAG that downloads a large reference dataset each night occasionally exceeds the default task timeout because the source API is slow. The engineer wants the task to fail fast and be retried automatically rather than hanging for hours, and wants failed runs to be visible for alerting. Which configuration should be applied to the task?

Question 9easymultiple choice
Read the full Scenario explanation →

You need to schedule a Dataproc Spark job to run at 2 AM every day, and upon completion, trigger a BigQuery load job. Which Cloud Composer operator should you use to run the Spark job?

Question 10mediummultiple choice
Read the full Scenario explanation →

A company uses Cloud Spanner and needs to store a parent-child relationship where the child table is frequently queried together with the parent. The parent has millions of rows and the child billions. Which Spanner feature optimizes performance for this pattern?

Question 11mediummulti select
Read the full Scenario explanation →

A company needs to stream real-time user activity data from their application into BigQuery for immediate dashboarding. They want to minimize latency (under 5 seconds) and ensure exactly-once delivery. Which TWO options should they consider? (Choose 2)

Question 12mediummulti select
Read the full Scenario explanation →

You want to optimize BigQuery costs for a large dataset that is frequently queried by time range. You also need to ensure that predictable workloads have dedicated slot capacity. Which TWO strategies should you combine? (Choose 2)

Question 13hardmulti select
Read the full Scenario explanation →

You are designing a data pipeline for ML training with Vertex AI. You need to split time-series data into train/validation/test sets without leaking future data. Which THREE practices should you follow?

Question 14mediummultiple choice
Read the full Scenario explanation →

You are building a binary classification model using AutoML Tables on Vertex AI. The dataset has a severe class imbalance (1% positive class). Which strategy should you use to handle the imbalance?

Question 15hardmulti select
Read the full Scenario explanation →

You are building a Dataflow pipeline that reads from Pub/Sub, applies transformations, and writes to BigQuery. The pipeline must handle late-arriving data and ensure that the windowing and triggering are correct. Which THREE configurations should you consider? (Choose 3)

Question 16mediummulti select
Read the full Scenario explanation →

A company wants to build a reporting pipeline where data is collected from IoT devices, stored raw in Cloud Storage, and then processed into BigQuery for analytics. They need to ensure data is encrypted at rest using customer-managed keys. Which THREE steps should they take? (Choose 3 correct options)

Question 17hardmultiple choice
Read the full Scenario explanation →

A company's Dataflow pipeline uses the PubSubIO source to read messages and writes to BigQuery via the BigQueryIO sink. The pipeline is running in Streaming mode with exactly-once semantics enabled. Occasionally, duplicate rows appear in BigQuery. What is the most likely reason?

Free account

Track your progress over time

Create a free account to save your results and see which topics improve across sessions.

Focused Scenario sessions

Start a Scenario only practice session

Every question in these sessions is drawn from the Scenario domain — nothing else.

Related practice questions

Related PDE topic practice pages

Move into related areas when this topic feels solid.

Frequently asked questions

What does the PDE exam test about Scenario?
Scenario questions test whether you can apply the concept in context, not just recognise a definition.
How should I use these practice questions?
Select your answer before revealing the explanation. Then read why each option is right or wrong — this active recall approach builds retention far faster than re-reading notes.
Can I practise just Scenario questions in a focused session?
Yes — the session launcher on this page draws every question from the Scenario domain. Use a 10-question session first to gauge your baseline, then move to 20 or 30 once the weak spots are clear.
Where can I practise other PDE topics?
Use the topic links above to move to related areas, or go back to the PDE question bank to see all topics.
Are these real exam questions or dumps?
These are original practice questions written to test the same concepts the PDE exam covers. They are not copied from any real exam or dump site.