Courseiva

PMLE Scaling Prototypes into ML Models Practice Question

A machine learning engineer is training a model on Vertex AI using a custom container. The training job uses a large dataset stored in BigQuery. The engineer wants to minimize data transfer costs and maximize training speed. Which of the following approaches is most efficient?

⚠ Common exam trap

The trap here is assuming that exporting to Cloud Storage is always necessary, but the BigQuery Storage API allows direct, efficient access.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use the BigQuery Storage API to read data directly into the training program with parallel streams.

The BigQuery Storage API is designed for high-performance data reading, offering parallel streams and efficient data transfer. It avoids the overhead of exporting data to Cloud Storage or loading into memory all at once. This makes it the most efficient and cost-effective approach for training on large BigQuery datasets.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Use the BigQuery Storage API to read data directly into the training program with parallel streams.

    Why this is correct

    The BigQuery Storage API provides high-throughput, parallel access to BigQuery data directly from the training program. It avoids exporting data to Cloud Storage, reducing costs and latency. It also supports column filtering and predicate pushdown, which can reduce the amount of data read. This is the most efficient method for training on large BigQuery datasets.

  • ✗

    Run a query in BigQuery to extract the data, then save it to a local SSD on the training VM.

    Why it's wrong here

    Saving to local SSD requires first extracting the data, which involves network transfer and storage costs. It also limits the dataset size to the SSD capacity. This approach is not scalable and adds unnecessary steps. The BigQuery Storage API is designed for this purpose.

  • ✗

    Export the BigQuery table to Cloud Storage in TFRecord format, then read it using the tf.data API.

    Why it's wrong here

    Exporting to Cloud Storage and reading with tf.data is a common pattern, but it involves an extra export step and data duplication. It can be efficient if the data is reused multiple times, but for a single training job, it may not be the most cost-effective due to storage and egress costs. It also adds latency for the export.

  • ✗

    Use the BigQuery client library to run a query and fetch results into memory using the to_dataframe() method.

    Why it's wrong here

    The to_dataframe() method loads all results into memory, which may not be feasible for large datasets. It also does not provide parallel streaming and can be slow. The BigQuery Storage API is optimized for large-scale data reading with parallel streams.

About these practice questions

One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.