PMLE Collaborating to manage data and models Practice Question
A healthcare organization is building a machine learning model to predict patient readmission risk. They have sensitive data stored in BigQuery that includes protected health information (PHI). The data science team uses Vertex AI Workbench notebooks to explore the data and develop models. The organization's security policy requires that all PHI data must be encrypted at rest and in transit, and that access to the data is logged and audited. They also need to ensure that the data used for model training is de-identified to remove direct identifiers such as patient names and SSNs. The team wants to automate the de-identification process as part of the data pipeline. Which approach meets these requirements?
⚠ Common exam trap
Google Cloud often tests the distinction between data masking/encryption (which still exposes PHI to authorized users) and true de-identification (which removes or transforms PHI so it is no longer considered protected health information).
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Create a Dataflow pipeline that reads from the original BigQuery table, applies Cloud DLP de-identification transforms, and writes to a new BigQuery table. Grant the data science team access to the de-identified table.
It uses Cloud DLP within a Dataflow pipeline to automatically de-identify PHI data as it is read from the original BigQuery table and written to a new, de-identified table. This satisfies the requirement for automated de-identification, while the original table remains encrypted at rest (BigQuery default) and in transit (TLS), and access to the original data can be logged via Cloud Audit Logs. The data science team only gets access to the de-identified table, ensuring PHI is not exposed during model development.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Create a Dataflow pipeline that reads from the original BigQuery table, applies Cloud DLP de-identification transforms, and writes to a new BigQuery table. Grant the data science team access to the de-identified table.
Why this is correct
Cloud DLP de-identification transforms applied in a Dataflow pipeline remove direct identifiers before the data lands in a separate BigQuery table, satisfying the de-identification requirement while BigQuery's default encryption at rest and TLS in transit plus audit logging cover the remaining policy constraints.
- ✗
Enable Shielded VM on Vertex AI Workbench notebooks and use VPC-SC to restrict data access.
Why it's wrong here
Shielded VM verifies boot integrity and VPC Service Controls perimeter network boundaries; neither de-identifies PHI nor removes names and SSNs, so the automation requirement stays unmet. Shielded VM suits hardening notebook instances against rootkit and boot-level tampering, not data transformation.
- ✗
Use Cloud Key Management Service to encrypt the PHI columns in BigQuery, and share the encryption key with the data science team.
Why it's wrong here
CMEK encrypts column data at rest but leaves identifiers intact and readable by anyone holding the key, so de-identification never occurs; sharing the key also widens PHI exposure. CMEK is the right control when the requirement is customer-managed encryption keys rather than removing direct identifiers.
- ✗
Use BigQuery row-level security to mask PHI columns for the data science team, and train the model directly on the original table.
Why it's wrong here
Row-level security filters which rows a user sees, not which columns; masking policies hide values at query time yet the model still trains on the original table containing identifiers. Row-level security fits restricting access by row attributes such as tenant or region, not de-identification for training.
Go deeper
Related to this question
About these practice questions
One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.