Courseiva

PMLE Vertex AI Batch Prediction Practice Question

You are designing a batch prediction pipeline using Vertex AI. The input data is 50 TB in CSV format on GCS. The model requires feature engineering that involves complex transformations (e.g., datetime parsing, one-hot encoding). Which TWO services or steps should you include in your pipeline?

⚠ Common exam trap

Candidates often assume that any scalable service can handle batch processing, but Cloud Functions and Cloud SQL are unsuitable for 50TB. The trap is to think that more than two services are needed, but a Dataflow pipeline and Vertex AI batch prediction are sufficient.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Run Vertex AI batch prediction job with GCS source pointing to the processed TFRecord files.

Option D is correct because Dataflow is the managed, horizontally scalable Apache Beam service designed for large-scale ETL like 50 TB of CSV, and it can perform the required complex transformations (datetime parsing, one-hot encoding) before writing the engineered features to GCS in TFRecord format, which is an efficient binary format for ML training and prediction. Option C is correct because a Vertex AI batch prediction job can then consume those processed TFRecord files directly from GCS as its input source, letting the model run inference at scale without re-doing feature engineering. Option A is not appropriate because Cloud Functions is an event-driven, short-lived serverless compute service with limited memory and execution time, making it unsuitable for transforming 50 TB of data file-by-file. Option B is not appropriate because Cloud SQL is a relational OLTP database, not a scalable store for massive intermediate ML feature data. Option E is not appropriate because writing engineered features to BigQuery does not produce the TFRecord input that Vertex AI batch prediction expects, and BigQuery is not the target format for this pipeline's model input.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Use Cloud Functions to transform each file individually.

    Why it's wrong here

    Cloud Functions have a 9-minute timeout and 32 GB memory ceiling, making them incapable of processing 50 TB of CSV data or executing complex transformations like datetime parsing and one-hot encoding across such volume. This option is tempting because Cloud Functions excel at lightweight, event-driven file processing, such as triggering a single-file validation or enrichment when a new CSV lands in GCS, where latency and scale are minimal.

  • ✗

    Use Cloud SQL to store intermediate results.

    Why it's wrong here

    Cloud SQL is a relational database designed for transactional workloads with structured, row-level access, not for storing 50 TB of intermediate batch data; its storage and throughput limits would create a bottleneck in a high-volume pipeline. It is tempting because Cloud SQL offers managed persistence for interim results in smaller-scale ETL jobs, and would be correct if the pipeline required low-latency querying of a few gigabytes of transformed records for iterative debugging or real-time lookups.

  • ✓

    Run Vertex AI batch prediction job with GCS source pointing to the processed TFRecord files.

    Why this is correct

    Running a Vertex AI batch prediction job against GCS-hosted TFRecord files lets the pipeline consume the preprocessed, engineered features at scale, avoiding repeated transformation of 50 TB of raw CSV. TFRecord input suits large batch workloads efficiently.

  • ✓

    Use Dataflow to read CSV, perform feature engineering, and write to GCS in TFRecord format.

    Why this is correct

    Dataflow handles the 50 TB scale with distributed processing, performing the required datetime parsing and one-hot encoding, then writing TFRecord for efficient training ingestion. This satisfies the complex feature engineering constraint that BigQuery or client-side preprocessing cannot handle at this volume.

  • ✗

    Use Dataflow to read CSV, perform feature engineering, and write to BigQuery.

    Why it's wrong here

    Dataflow reading CSV, engineering features and writing to BigQuery is a valid pipeline component, so this option is not incorrect on technical grounds; the stem asks for two services or steps, and this is one legitimate pairing element. It tempts as a complete answer, yet selection requires a second complementary step.

Quick reference

Cloud Service Model Comparison

ModelYou ManageProvider ManagesExamples
IaaSOS, runtime, apps, dataHardware, hypervisor, networkingEC2, Azure VMs, GCP Compute Engine
PaaSApps and dataOS, runtime, middleware, hardwareElastic Beanstalk, Azure App Service
SaaSData and settings onlyEverything elseMicrosoft 365, Salesforce, Workday
FaaS / ServerlessFunction code onlyInfra, scaling, runtimeLambda, Azure Functions, Cloud Run
CaaSContainers and appsKubernetes, OS, hardwareEKS, AKS, GKE

About these practice questions

This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.