PDE Designing Data Processing Systems Practice Question
A healthcare analytics group must build a pipeline that ingests HL7 messages from an on-premises interface engine, must retain raw messages for seven years for compliance, and must expose de-identified aggregates to analysts. The security team requires that protected health information never be written to a dataset analysts can query, and that all data be encrypted with keys the organization manages and can revoke. Which two design choices satisfy these requirements? (Choose two.)
⚠ Common exam trap
The trap here is treating masking or authorized views over a raw BigQuery table as equivalent to never writing PHI into an analyst-queryable dataset, when the requirement demands physical separation, not query-time filtering.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
De-identify the messages in Dataflow, write only the de-identified aggregates into a separate BigQuery dataset secured with CMEK and policy tags, and grant analysts access only to that dataset.
The requirements demand a hard separation between raw PHI and anything analysts can query, plus organization-managed revocable encryption and seven-year retention. Landing raw messages in Cloud Storage with CMEK and a retention lock, accessible only to the ingestion service account, satisfies retention and key control. De-identifying in Dataflow and publishing only aggregates to a separate CMEK-protected BigQuery dataset keeps PHI out of analyst-facing storage entirely, so both boundaries hold independently.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Load the raw messages directly into a BigQuery table that analysts can query, and rely on column-level security to mask the PHI columns.
Why it's wrong here
Column-level security and policy tags can restrict field access, but the requirement states that PHI must never be written to a dataset analysts can query. Placing raw messages in an analyst-accessible table violates that boundary even if masking is applied, because policy tags protect columns rather than preventing the data from residing there. This fails the explicit architectural constraint.
- ✓
De-identify the messages in Dataflow, write only the de-identified aggregates into a separate BigQuery dataset secured with CMEK and policy tags, and grant analysts access only to that dataset.
Why this is correct
Performing de-identification in Dataflow before the data reaches BigQuery ensures PHI never enters an analyst-queryable dataset, satisfying the hard boundary. Writing aggregates to a distinct dataset protected by CMEK gives the organization revocable key control, and policy tags add fine-grained access restrictions. Analysts are granted access only to this curated dataset, keeping the raw zone entirely separate.
- ✓
Land raw HL7 messages in a Cloud Storage bucket configured with a customer-managed encryption key (CMEK) and a retention lock, and restrict access with IAM to the ingestion service account only.
Why this is correct
Cloud Storage with CMEK lets the organization control and revoke the key, and a bucket retention lock enforces the seven-year immutability requirement for compliance. Restricting access to the ingestion service account keeps raw PHI out of analyst reach. This satisfies the encryption-control and raw-retention requirements while establishing a clear trust boundary between the raw zone and anything analysts can query.
- ✗
Store the raw messages in BigQuery and use authorized views so analysts query a view that filters out PHI columns instead of the base table.
Why it's wrong here
Authorized views let analysts read a curated view without direct base-table access, which is a useful pattern, but the raw PHI still resides in a BigQuery dataset and the boundary depends on view correctness rather than physical separation. The requirement is that PHI never be written to an analyst-queryable dataset, and a single misconfigured view or grant would expose it. This is a weaker control than keeping raw data outside BigQuery's analyst-facing datasets.
- ✗
Encrypt the raw messages with a customer-supplied key in the on-premises interface engine and upload only the ciphertext to Cloud Storage, decrypting in Dataflow for de-identification and discarding the key afterward.
Why it's wrong here
This provides encryption at rest before upload, but discarding the key after de-identification makes the seven-year raw retention impossible to satisfy because the ciphertext can never be decrypted again for compliance review. The organization also loses the ability to manage a revocable key over the retained data. The approach conflicts with the retention requirement rather than meeting it.
Go deeper
Related to this question
About these practice questions
One of 747 original PDE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.