Courseiva

PMLE Collaborating to manage data and models Practice Question

When distributing training across multiple workers using Vertex AI Training, how should the team share the training dataset?

⚠ Common exam trap

Many exam-takers confuse 'shared storage' with 'local copies' or 'user-friendly sync tools,' assuming NFS or Drive are viable for distributed ML, when Vertex AI explicitly requires a cloud-native object store like GCS for scalability and fault tolerance.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use Cloud Storage

Vertex AI Training workers need shared, concurrent read access to the training dataset without manual replication. Cloud Storage (GCS) is the recommended and fully integrated solution because it provides a distributed, highly available object store that all workers can read from in parallel via the `tf.io.gfile` API or GCS connector, eliminating data duplication and ensuring consistency across the cluster.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Copy the dataset to each worker's local disk

    Why it's wrong here

    Copying to each worker's local disk gives every worker an isolated, divergent copy, so distributed training cannot shard or synchronise the dataset and disk capacity limits scale. This suits single-node jobs reading small local files, whereas Vertex AI expects a shared source such as Cloud Storage that all workers read consistently.

  • ✗

    Use NFS

    Why it's wrong here

    NFS provides a shared POSIX filesystem, but Vertex AI Training workers run in separate containers without a pre-mounted NFS server, so the dataset is unreachable unless you build and secure that infrastructure yourself. NFS suits on-premises clusters sharing files across nodes, not managed Vertex AI jobs where Cloud Storage or a shared Filestore mount is configured.

  • ✓

    Use Cloud Storage

    Why this is correct

    Cloud Storage provides a shared, region-agnostic object store that every worker can read concurrently, satisfying the multi-worker distribution requirement without duplicating data. Vertex AI Training mounts or streams GCS paths directly, so each worker accesses the same dataset shards. Persistent disks or local storage would tie data to a single VM, breaking parallel access.

  • ✗

    Use Google Drive

    Why it's wrong here

    Google Drive provides no POSIX filesystem semantics or high-throughput parallel reads, so workers cannot reliably stream shards during distributed training. Drive suits small manual datasets and notebooks; Vertex AI Training expects data in Cloud Storage or a mounted Filestore.

About these practice questions

One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.