PDE Ingesting and Processing the Data Practice Question
A healthcare company needs to process HL7 messages containing sensitive patient data. The messages arrive in Cloud Storage as JSON files. The pipeline must de-identify the data using the Cloud Healthcare API DLP de-identification, then load the results into BigQuery. The pipeline must ensure that no unredacted data is ever written to BigQuery, and that processing is fault-tolerant. The Dataflow pipeline reads from Cloud Storage, calls the DLP API for de-identification, and writes to BigQuery. Which additional configuration ensures that only de-identified data reaches BigQuery?
⚠ Common exam trap
The trap here is assuming that encryption or access controls on original data in BigQuery are sufficient, but the requirement explicitly forbids any unredacted data in BigQuery.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use a DoFn that calls the DLP API and only emits the de-identified record; then use a BigQueryIO write with a dead-letter queue for failed DLP calls.
The pipeline must ensure that only de-identified data is written to BigQuery. Emitting only de-identified records from the DoFn guarantees that unredacted data never reaches the sink. A dead-letter queue for failed DLP calls maintains fault tolerance by isolating problematic records for later analysis or reprocessing, rather than blocking the pipeline or writing original data. This design meets both privacy and reliability requirements.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use a ParDo that calls the DLP API and writes the de-identified data to BigQuery; if the DLP call fails, retry indefinitely until success.
Why it's wrong here
Infinite retries can cause the pipeline to stall indefinitely if the DLP API is unavailable or the request is invalid. This reduces fault tolerance and may block processing of other records. A dead-letter queue is a better pattern for handling failures without halting the pipeline. Additionally, indefinite retries do not ensure that unredacted data is never written, as the record is held but not written; however, the pipeline may never progress, which is operationally unacceptable.
- ✓
Use a DoFn that calls the DLP API and only emits the de-identified record; then use a BigQueryIO write with a dead-letter queue for failed DLP calls.
Why this is correct
By only emitting de-identified records, the pipeline ensures that unredacted data never reaches BigQuery. Using a dead-letter queue for failed DLP calls prevents data loss and allows reprocessing. This design is fault-tolerant and meets the strict requirement that no original data is written to BigQuery. It also handles API errors gracefully without compromising data privacy.
- ✗
Use a ParDo transform that calls the DLP API and then writes the original and de-identified data to separate BigQuery tables, with access controls on the original table.
Why it's wrong here
Writing the original data to BigQuery, even with access controls, violates the requirement that no unredacted data is ever written to BigQuery. Access controls are not sufficient because the data still resides in BigQuery and could be exposed through misconfiguration. The pipeline must ensure that only de-identified data is written, so this approach is fundamentally flawed.
- ✗
Use a GroupByKey to batch records, then call the DLP API in a batch request; write both original and de-identified data to BigQuery with column-level encryption.
Why it's wrong here
Writing original data to BigQuery, even with column-level encryption, does not satisfy the requirement that no unredacted data is ever written. Encryption protects data at rest but does not prevent access by authorized users or applications. The pipeline must avoid writing original data entirely. Batching DLP calls is efficient but does not address the core privacy requirement.
Go deeper
Related to this question
About these practice questions
One of 747 original PDE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.