Reinforce PDE concepts with active-recall study cards covering all 5 blueprint domains. Each card shows the question on the front and the correct answer with a full explanation on the back.
Flashcards work through active recall — the process of retrieving information from memory rather than passively re-reading it. Research consistently shows that active recall produces stronger, longer-lasting memory than re-reading study guides. For PDE preparation, this means flashcards are one of the highest-return study tools available.
Attempt recall first
Read the PDE question on each card, pause, and attempt to formulate the answer in your own words before revealing. This retrieval attempt — even if wrong — dramatically strengthens memory compared to immediately reading the answer.
Review wrong cards again
When you get a card wrong, note it and add it back to your review pile. Spaced repetition — seeing difficult cards more frequently — is the mechanism that makes flashcard study far more efficient than linear reading.
Study by domain
Group your PDE flashcard sessions by domain for the first 3–4 weeks. Master one domain before moving to the next. In the final week, shuffle all cards together to test cross-domain recall — which is what the real PDE exam requires.
Short sessions beat marathon reviews
20–30 flashcard cards per session, done daily, produces better retention than a single 200-card marathon session. Five short daily sessions per week over 4 weeks gives you over 400 total card reviews — enough to reliably pass PDE.
Sample cards from the PDE flashcard bank. Read the question, think of the answer, then read the explanation below.
A company needs a fully managed, globally distributed relational database with strong consistency, external consistency, and 99.999% SLA for a financial transaction processing system. Which Google Cloud service should they use?
Cloud Spanner
Cloud Spanner is the correct choice because it is a fully managed, globally distributed relational database service that provides strong consistency, external consistency (true serializable transactions across regions), and a 99.999% SLA. These features are essential for a financial transaction processing system that requires ACID compliance and global scalability without sacrificing consistency.
A data engineer needs to store raw sensor data in Cloud Storage and automatically transition it to a lower-cost storage class after 30 days, then delete it after 365 days. What should they configure?
Configure a lifecycle rule with SetStorageClass to Nearline after 30 days and Delete after 365 days.
Cloud Storage lifecycle management rules allow you to automatically transition objects to a lower-cost storage class (such as Nearline) after a specified number of days and then delete them after another period. This is the native, serverless way to manage object lifecycle without external scripts or compute resources.
An e-commerce company uses Cloud Spanner for order processing. They need to query orders by customer ID and retrieve all order items. Which schema design pattern should they use for optimal performance?
Denormalize by storing order items as a repeated field in the orders table.
For this access pattern, the optimal Spanner design is to denormalize order items as a repeated field in the Orders table. A repeated field stores child rows inline with the parent row, so a query by customer_id (with a secondary index on customer_id) retrieves the order and all its items in a single lookup without a join or extra network round-trip. Interleaved tables co-locate child rows with the parent, but they only help when the query uses the parent's primary key prefix; querying by customer_id would still require a secondary index and then a join-like lookup to fetch interleaved children, so option A is not the best fit here.
A company uses Dataproc to run daily Spark ML jobs. The jobs run for 2 hours each day. The team wants to reduce costs without changing job characteristics. Which strategy is MOST cost-effective?
Use preemptible instances for worker nodes
Preemptible (Spot) VMs cost up to 80% less than standard VMs and are ideal for fault-tolerant, batch-oriented workloads like Spark ML jobs that can tolerate occasional preemption. Since the jobs run only 2 hours daily and Dataproc automatically handles node replacement when preemptible instances are reclaimed, using preemptible workers delivers the largest cost reduction without changing job characteristics.
A data engineer needs to create a BigQuery table that is partitioned by ingestion time and clustered by customer_id and transaction_date. They also want to limit access so that only users from a specific domain can query the table. Which approach should they use?
Create the table with partitioning and clustering, then create an authorized view on the table and grant the view access to the domain users
Authorized views allow sharing query results with specific users/groups without giving direct table access. Clustering and partitioning are defined at table creation. IAM roles at dataset level are too broad. Row-level security filters rows but doesn't restrict domain.
A startup needs a fully managed, serverless Spark service to run occasional data processing jobs without managing clusters. They want to pay only for the resources used during job execution. Which Google Cloud service should they use?
Dataproc Serverless
Dataproc Serverless provides a serverless Spark environment where you pay per job execution. Cloud Data Fusion is for visual ETL. Dataproc is managed but not serverless. Dataflow is serverless for Beam, not Spark.
You need to stream real-time user click events from your application into BigQuery for immediate analysis. The events must be available for query within seconds. Which approach is recommended?
Use Pub/Sub to Dataflow to BigQuery with the Storage Write API for high-throughput streaming.
The recommended approach is to use Pub/Sub to Dataflow to BigQuery with the Storage Write API. Dataflow provides a managed stream processing service that can handle high-throughput, low-latency ingestion, and the Storage Write API offers exactly-once semantics and is optimized for streaming inserts into BigQuery. This combination ensures events are available for query within seconds and scales well.
Your company is migrating an on-premises Hadoop cluster to Google Cloud. You need to transform large datasets using Spark SQL. Which Google Cloud service should you use?
Dataproc
Dataproc is the managed Spark and Hadoop service on Google Cloud, purpose-built for running existing Spark SQL workloads with minimal changes. It allows you to spin up a cluster, run your Spark SQL transformations on large datasets stored in Cloud Storage or BigQuery, and then tear it down, making it the direct equivalent of an on-premises Hadoop cluster in the cloud.
A data engineer needs to transfer 500 TB of on-premises data to Google Cloud Storage. The data is stored on NAS devices and the network bandwidth is limited to 100 Mbps. What is the most cost-effective and timely transfer method?
Use Transfer Appliance
Transfer Appliance is Google's physical data transfer service that ships a rackable appliance to the customer site, where data is copied locally and the appliance is shipped back to Google for upload to Cloud Storage. For 500 TB over a 100 Mbps link, the network transfer time would be roughly 1.5 years, making online transfer impractical. Transfer Appliance is the most cost-effective and timely option for this volume and bandwidth.
You are building a Dataflow pipeline in Python that reads messages from Pub/Sub, enriches them with data from a BigQuery table, and writes the results to BigQuery. The enrichment lookup table is large and changes infrequently. Which approach minimizes cost and latency?
Use a side input that reads the BigQuery table periodically and caches it.
Using a side input that periodically reads the BigQuery table and caches it avoids querying BigQuery for every incoming message, which would be prohibitively expensive and high-latency. The side input is refreshed at a configurable interval (e.g., every 10 minutes) via a pipeline option, and the cached data is broadcast to all workers, enabling fast, in-memory lookups without per-element I/O. This approach minimizes cost by reducing BigQuery API calls and minimizes latency by avoiding synchronous queries for each message.
You need to create a Looker model that defines a 'sales' view based on a BigQuery table, with a measure for total revenue. Which LookML object defines the table and dimensions?
view
In LookML, a view defines the underlying database table and contains its dimensions (columns) and measures (aggregations like total revenue). The view is the fundamental building block that maps a physical table to a logical set of fields, and it is referenced by explores to make those fields queryable.
A company uses Looker Studio to build dashboards from BigQuery data. They notice that queries take several seconds to return. They want to improve performance without changing the schema or adding materialized views. Which option should they use?
Enable BigQuery BI Engine on the relevant project.
BI Engine accelerates sub-second query response times in Looker Studio by caching data in memory within the BigQuery region.
A data engineer uses Cloud Composer to orchestrate a daily batch pipeline. A downstream task should only start after an upstream BigQuery load job finishes successfully and a specific file appears in Cloud Storage. Which combination of operators should the engineer use in the Airflow DAG?
BigQueryInsertJobOperator and GCSObjectExistenceSensor with upstream dependency
The engineer needs a BigQuery load job to finish successfully and a specific file to appear in Cloud Storage before a downstream task starts. The correct combination is BigQueryInsertJobOperator to run and wait for the BigQuery job, and GCSObjectExistenceSensor to check for the file, with the sensor set as an upstream dependency of the downstream task. This ensures both conditions are met before proceeding.
An engineer needs to create a reusable Dataflow pipeline that can be executed with different parameters without modifying code. Which Dataflow feature should they use?
Dataflow Flex Templates
Dataflow Flex Templates allow packaging a pipeline as a Docker image with a metadata file, enabling reuse with different parameters at runtime without code changes. They support dynamic parameters and are the recommended approach for reusable pipelines.
A data engineer needs to alert when Pub/Sub subscription has messages older than 1 hour. Which Cloud Monitoring metric and filter should they use?
Metric: subscription/oldest_unacked_message_age; filter: subscription_id
The metric subscription/oldest_unacked_message_age directly reports the age of the oldest unacknowledged message in a subscription, which is exactly what's needed to alert when messages are older than 1 hour. The filter subscription_id scopes the metric to the specific subscription. This metric is designed for monitoring message backlog age and is the correct choice for this requirement.
The PDE flashcard bank covers all 5 official blueprint domains published by Google Cloud. Cards are distributed proportionally, so domains with higher exam weight have more cards.
Domain Coverage
Storing the Data
Designing Data Processing Systems
Ingesting and Processing the Data
Preparing and Using Data for Analysis
Maintaining and Automating Data Workloads
Both flashcards and practice questions are evidence-based study tools. The difference is in what they train:
Flashcards — concept retention
Best for memorising definitions, acronyms, protocol behaviours, command syntax, and conceptual distinctions. Use flashcards to build the foundational vocabulary that PDE questions assume you know.
Best in: weeks 1–3
Practice tests — application
Best for applying concepts to realistic scenarios, eliminating distractors, and building exam stamina.PDE questions test scenario reasoning — not just recall — so practice tests are essential.
Best in: weeks 3–6
The most effective PDE study plan combines both: use flashcards for the first 2–3 weeks to build conceptual foundations, then shift to practice tests and mock exams in the final 2–3 weeks to apply and benchmark that knowledge. Most candidates who pass on their first attempt use both tools.
Yes. Courseiva provides free PDE flashcards across all official exam domains. Every card includes the correct answer and a full explanation of why it is right and why the distractors are wrong. The platform also includes topic-based practice, mock exams, and readiness tracking — no account required.
Courseiva has 747+ original PDE flashcards across all 5 exam blueprint domains. New cards are added regularly as the question bank grows. All cards are checked against the official Google Cloud exam objectives, with editorial oversight from an experienced network and security engineer.
Courseiva flashcards are purpose-built for IT certification exams. Unlike generic flashcard platforms where content quality varies, every Courseiva card is mapped to the official PDE exam blueprint, written by engineers who hold the certification, and includes a full explanation of the correct answer and why the distractors are wrong. This explanation quality is what separates genuine learning from rote memorisation.
Courseiva is a web platform — an internet connection is required. For offline study, we recommend creating free Courseiva account, using the platform in your browser, and using your device's offline capabilities if your browser supports offline web apps.
Save your results, see which domains need more work, and get spaced repetition recommendations — all free.
Sign Up FreeFree forever · Every certification included