PMLE Collaborating to manage data and models Practice Question
When distributing training across multiple workers using Vertex AI Training, how should the team share the training dataset?
⚠ Common exam trap
Many exam-takers confuse 'shared storage' with 'local copies' or 'user-friendly sync tools,' assuming NFS or Drive are viable for distributed ML, when Vertex AI explicitly requires a cloud-native object store like GCS for scalability and fault tolerance.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use Cloud Storage
Vertex AI Training workers need shared, concurrent read access to the training dataset without manual replication. Cloud Storage (GCS) is the recommended and fully integrated solution because it provides a distributed, highly available object store that all workers can read from in parallel via the `tf.io.gfile` API or GCS connector, eliminating data duplication and ensuring consistency across the cluster.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Copy the dataset to each worker's local disk
Why it's wrong here
Copying to each worker's local disk gives every worker an isolated, divergent copy, so distributed training cannot shard or synchronise the dataset and disk capacity limits scale. This suits single-node jobs reading small local files, whereas Vertex AI expects a shared source such as Cloud Storage that all workers read consistently.
- ✗
Use NFS
Why it's wrong here
NFS provides a shared POSIX filesystem, but Vertex AI Training workers run in separate containers without a pre-mounted NFS server, so the dataset is unreachable unless you build and secure that infrastructure yourself. NFS suits on-premises clusters sharing files across nodes, not managed Vertex AI jobs where Cloud Storage or a shared Filestore mount is configured.
- ✓
Use Cloud Storage
Why this is correct
Cloud Storage provides a shared, region-agnostic object store that every worker can read concurrently, satisfying the multi-worker distribution requirement without duplicating data. Vertex AI Training mounts or streams GCS paths directly, so each worker accesses the same dataset shards. Persistent disks or local storage would tie data to a single VM, breaking parallel access.
- ✗
Use Google Drive
Why it's wrong here
Google Drive provides no POSIX filesystem semantics or high-throughput parallel reads, so workers cannot reliably stream shards during distributed training. Drive suits small manual datasets and notebooks; Vertex AI Training expects data in Cloud Storage or a mounted Filestore.
Go deeper
Related to this question
About these practice questions
One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.