PDE Storing the Data Practice Question
A healthcare company stores patient records in Cloud SQL for PostgreSQL and needs to run analytical queries on the same data without impacting the production database. The analytics team requires near-real-time replication and wants to use BigQuery for querying. The data must be kept in sync with minimal latency and without writing custom ETL code. Which solution should the data engineer implement?
⚠ Common exam trap
The trap here is assuming that BigQuery federated queries or read replicas provide a scalable analytical solution, when they actually query the transactional database and do not offload analytics.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use Datastream to replicate from Cloud SQL for PostgreSQL to BigQuery with a dataset-level replication
Datastream provides serverless change data capture from Cloud SQL for PostgreSQL to BigQuery. It reads the database's write-ahead log and applies changes to BigQuery with low latency, keeping the analytical store in sync without custom code. This offloads analytics from the production database and meets the near-real-time and no-ETL requirements. Other options either introduce batch latency, require pipeline development, or do not create a separate analytical store.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Create a read replica of the Cloud SQL instance and point BigQuery federated queries at the replica
Why it's wrong here
BigQuery federated queries can read from Cloud SQL, but they are not designed for high-volume analytical workloads and can impact the source instance. A read replica reduces impact on the primary but still runs queries on a Cloud SQL instance, which is not optimized for large analytical scans. This does not provide a separate analytical store and may still compete for resources.
- ✓
Use Datastream to replicate from Cloud SQL for PostgreSQL to BigQuery with a dataset-level replication
Why this is correct
Datastream is a serverless change data capture service that can replicate from Cloud SQL for PostgreSQL to BigQuery with low latency. It reads the write-ahead log and applies changes to BigQuery, keeping the data in sync without custom ETL. This meets the near-real-time and no-code requirements while offloading analytics from the production database.
- ✗
Use Cloud Data Fusion to build a pipeline that reads from Cloud SQL and writes to BigQuery on a schedule
Why it's wrong here
Cloud Data Fusion is a code-free ETL tool, but it still requires building and maintaining a pipeline, which is custom ETL in practice. It can run on a schedule, but near-real-time replication would require a streaming pipeline that is more complex. Datastream is purpose-built for change data capture and requires less configuration for this scenario.
- ✗
Schedule a nightly export from Cloud SQL to Cloud Storage and load the files into BigQuery
Why it's wrong here
A nightly batch export does not provide near-real-time replication; the analytics team would see data up to 24 hours old. It also requires custom scripting to export, transfer, and load, which violates the no-custom-ETL requirement. This approach adds latency and operational overhead, making it unsuitable for the stated near-real-time need.
Quick reference
Cloud Service Model Comparison
| Model | You Manage | Provider Manages | Examples |
|---|---|---|---|
| IaaS | OS, runtime, apps, data | Hardware, hypervisor, networking | EC2, Azure VMs, GCP Compute Engine |
| PaaS | Apps and data | OS, runtime, middleware, hardware | Elastic Beanstalk, Azure App Service |
| SaaS | Data and settings only | Everything else | Microsoft 365, Salesforce, Workday |
| FaaS / Serverless | Function code only | Infra, scaling, runtime | Lambda, Azure Functions, Cloud Run |
| CaaS | Containers and apps | Kubernetes, OS, hardware | EKS, AKS, GKE |
Go deeper
Related to this question
About these practice questions
One of 747 original PDE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.