PMLE Collaborating to manage data and models Practice Question
A company uses a Cloud Composer DAG to run a daily ML pipeline that includes Dataflow jobs and model training on Vertex AI. The pipeline frequently fails due to insufficient permissions when the Dataflow worker accesses data in Cloud Storage. What is the most efficient way to resolve this issue?
⚠ Common exam trap
Google Cloud often tests the misconception that using a single service account for all components (like the Composer environment's service account) is simpler and sufficient, but this ignores the principle of least privilege and can cause security vulnerabilities or permission conflicts in distributed pipelines.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Create a custom service account with required permissions and assign it to the Dataflow job.
The most efficient way to resolve insufficient permissions for Dataflow workers accessing Cloud Storage is to create a custom service account with the required roles (e.g., roles/storage.objectViewer) and assign it to the Dataflow job via the --serviceAccount option. This follows the principle of least privilege and ensures that only the Dataflow workers have the necessary permissions, without affecting other pipeline components or exposing the bucket publicly.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Create a custom service account with required permissions and assign it to the Dataflow job.
Why this is correct
Dataflow workers execute under a service account, so granting the required Cloud Storage permissions to a dedicated custom service account and assigning it to the job fixes the access failures directly. This is more efficient than broadly widening project-level roles.
- ✗
Grant the 'roles/storage.objectViewer' role to 'allUsers' on the Cloud Storage bucket.
Why it's wrong here
Granting objectViewer to allUsers makes the bucket publicly readable, exposing data rather than fixing the worker's identity, and violates least privilege. It is tempting because it removes the permission error immediately, but the correct approach grants the Dataflow worker service account the required bucket role.
- ✗
Use the Composer environment's service account for all pipeline components.
Why it's wrong here
The Composer environment's service account is not automatically the Dataflow worker service account; workers use a separate controller service account, so Cloud Storage access still fails. It is tempting because reusing one identity across pipeline components simplifies IAM management, but that identity is not the one Dataflow workers assume.
- ✗
Move the Dataflow job to run after the pipeline so that data is already processed.
Why it's wrong here
Reordering the DAG changes execution sequence, not the Dataflow worker's IAM permissions, so the Cloud Storage access denial persists. It is tempting because sequencing Dataflow after upstream processing can reduce redundant work, but the failure is authorisation, not data readiness.
Visual reference
Go deeper
Related to this question
About these practice questions
One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.