PDE Maintaining and Automating Data Workloads Practice Question
You are designing a data quality pipeline that must inspect PII in BigQuery tables and de-identify sensitive columns before sharing with analysts. Which GCP service should you use?
⚠ Common exam trap
It's easy for candidates to confuse Dataplex's data governance features (like policy tags and metadata) with actual de-identification, but Dataplex cannot transform data—it only applies access controls, whereas Cloud DLP performs the actual masking or tokenization of sensitive values.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Cloud DLP
Cloud DLP (Data Loss Prevention) is the correct choice because it is purpose-built for inspecting, classifying, and de-identifying sensitive data such as PII. It integrates natively with BigQuery via inspection jobs and de-identification templates, allowing you to scan tables for over 150 built-in infoTypes (e.g., email, SSN) and apply transformations like masking, tokenization, or encryption before sharing data with analysts.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Dataplex
Why it's wrong here
Dataplex provides data discovery, profiling and governance across lakes, but its DLP-based de-identification runs as discovery scans rather than transforming columns in place before sharing. It is tempting because Dataplex catalogues and classifies PII, yet the stem requires masking sensitive BigQuery columns, which Sensitive Data Protection does directly.
- ✗
Cloud Data Catalog
Why it's wrong here
Cloud Data Catalog is a metadata discovery and governance service; it inventories and tags assets but does not inspect row contents or de-identify values. It suits cataloguing and lineage. De-identification of sensitive columns requires a service that performs actual DLP inspection and transformation.
- ✓
Cloud DLP
Why this is correct
Cloud DLP natively scans BigQuery tables, using infoType detectors to identify PII and de-identification transforms such as masking, tokenisation, and format-preserving encryption to protect sensitive columns. This directly satisfies the stem's requirement to inspect and de-identify PII in place before analysts access the shared data.
- ✗
Dataflow
Why it's wrong here
Dataflow is a stream and batch processing runner, not a PII inspection or de-identification engine; it lacks built-in DLP transforms unless calling the DLP API. Dataflow suits custom pipelines moving and transforming data. The scenario needs native sensitive-data discovery and masking across BigQuery columns.
Go deeper
Related to this question
About these practice questions
One of 747 original PDE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.