This guide covers all official exam objectives for the Google Professional Data Engineer certification, focusing on designing, building, maintaining, and optimizing data processing systems on Google Cloud.
This guide works best as a loop: read a chapter, test yourself with practice questions, look up unfamiliar terms in the glossary, then move to the next chapter.
15 chapters covering every exam objective. Each chapter includes key concepts, exam tips, common traps, comparison tables, and a 5-question quiz at the end.
Start Chapter 1Free timed and untimed practice with instant feedback and full explanations. Pick 10–120 questions per session. Filter by domain to drill your weak areas.
Go to practice testEvery PDEterm defined and searchable. Use it when a chapter mentions a concept you haven't seen before or want a quick refresher on.
Browse glossaryExam blueprint, domain weights, passing score, duration, cost, and registration links. Start here if you're new to this certification.
View exam guideOverview of Data Engineering on Google Cloud Platform
Objective 1.1 · Design data processing systems (e.g., batch, streaming, real-time, big data, and machine learning) on GCP.
Designing Scalable and Reliable Data Storage
Objective 1.2 · Design data storage systems (e.g., Cloud Storage, Bigtable, BigQuery, Spanner, Cloud SQL) for scalability, reliability, and performance.
Data Modeling and Schema Design
Objective 1.3 · Design data models and schemas (relational, NoSQL, and BigQuery) to support analytical and operational workloads.
Designing Data Pipelines and Orchestration
Objective 1.4 · Design data pipelines (batch, streaming, and hybrid) using tools like Cloud Dataflow, Cloud Composer, and Cloud Dataproc.
Building and Operationalizing Data Ingestion
Objective 2.1 · Design and implement data ingestion solutions (e.g., Cloud Pub/Sub, Cloud Storage, Cloud Dataflow, and API-based ingestion).
Processing Batch Data with Cloud Dataproc
Objective 2.2 · Design and implement batch data processing using Cloud Dataproc and Apache Spark/Hadoop ecosystem.
Processing Streaming Data with Cloud Dataflow
Objective 2.3 · Design and implement stream processing pipelines using Cloud Dataflow and Apache Beam.
Managing Data Lakes and Warehouses
Objective 2.4 · Design and implement data lakes (Cloud Storage) and data warehouses (BigQuery) including partitioning, clustering, and data lifecycle.
Storing Relational Data with Cloud SQL and Spanner
Objective 3.1 · Implement relational databases (Cloud SQL, Cloud Spanner) for transactional workloads, including high availability and scaling.
Storing NoSQL Data with Cloud Bigtable and Firestore
Objective 3.2 · Implement NoSQL databases (Bigtable, Firestore) for low-latency, high-throughput workloads.
Securing and Governing Data on GCP
Objective 3.3 · Implement data security, access control, encryption, and data governance using IAM, Cloud KMS, and Data Loss Prevention (DLP).
Machine Learning on GCP with Vertex AI
Objective 4.1 · Design and implement machine learning models and pipelines using Vertex AI, including AutoML and custom training.
ML Model Deployment and Monitoring
Objective 4.2 · Deploy, monitor, and manage machine learning models in production using Vertex AI endpoints, model registry, and monitoring tools.
Automating Infrastructure with Deployment Manager and Terraform
Objective 5.1 · Automate the provisioning and management of data infrastructure using Infrastructure as Code (Deployment Manager, Terraform).
Cost Optimization and Performance Tuning
Objective 5.3 · Optimize data processing and storage costs and performance (BigQuery slot management, Cloud Storage lifecycle, Dataflow tuning).
Free PDE practice questions with full explanations. Test what you learn chapter by chapter.
PDE Practice Questions