Courseiva

CCNA Automating and Orchestrating ML Pipelines Questions

75 of 81 questions · Page 1/2 · Automating and Orchestrating ML Pipelines · Answers revealed

1
Multi-Selecthard

You are tasked with building a robust ML pipeline that must be idempotent and handle data skew between training and serving. Which three practices should you implement?

Select 3 answers
A.Store intermediate data in Cloud Storage with unique run IDs.
B.Pass large datasets between components as serialized in-memory objects.
C.Monitor feature distributions in training data vs. serving data to detect skew.
D.Use the same random seed for every run to ensure reproducibility.
E.Ensure each component produces deterministic outputs given the same inputs.
AnswersA, C, E

Unique run IDs in Cloud Storage paths make each pipeline execution write to distinct locations, so reruns neither overwrite nor reuse stale intermediates. This directly satisfies the idempotency requirement, since identical inputs produce isolated, reproducible outputs rather than colliding with prior runs.

Why this answer

Option A is correct because writing intermediate artifacts to Cloud Storage under a unique run ID (e.g., gs://bucket/run_id/...) isolates each pipeline execution, so reruns don't overwrite or collide with prior outputs — a key requirement for idempotency. Option C is correct because comparing training feature distributions against serving feature distributions (e.g., via statistics like mean, variance, or histogram distance) is the standard way to detect training/serving skew and trigger remediation. Option E is correct because deterministic components that yield identical outputs for identical inputs make reruns safe and reproducible, which is the essence of an idempotent pipeline.

Option B is wrong because passing large datasets as serialized in-memory objects is fragile, memory-bound, and non-idempotent; components should exchange data via durable storage references instead. Option D is wrong because a fixed random seed only aids reproducibility of stochastic steps and does nothing to guarantee idempotency or address training/serving skew.

Exam trap

PMLE often tests the difference between reproducibility (same seed) and idempotency (safe reruns) — candidates pick the seed option because it 'sounds like best practice' but it does not address pipeline idempotency or skew.

2
Multi-Selectmedium

A pipeline uses the Google Cloud Pipeline Components to perform AutoML training and batch prediction. Which two components from the GCPC library should they use? (Choose two.)

Select 2 answers
A.CustomJobRunOp
B.DataflowPythonOp
C.AutoMLTabularTrainingJobRunOp
D.BatchPredictOp
E.EndpointPredictOp
AnswersC, D

AutoMLTabularTrainingJobRunOp submits a tabular AutoML training job directly from the pipeline, satisfying the AutoML training requirement. It wraps the Vertex AI training API as a pipeline component, so the pipeline orchestrates training without custom code. Batch prediction is handled separately by a prediction component, making this one of the two required GCPC components.

Why this answer

Option C, AutoMLTabularTrainingJobRunOp, is correct because it is the Google Cloud Pipeline Components (GCPC) operator that wraps the Vertex AI AutoML tabular training job, allowing the pipeline to launch an AutoMLTabularTrainingJob and produce a model artifact. Option D, BatchPredictOp, is correct because it is the GCPC component that submits a Vertex AI batch prediction job against a trained model and a specified input data source, which is exactly the batch prediction step described in the scenario. Option A, CustomJobRunOp, is not appropriate here because it runs a custom training container rather than an AutoML training job.

Option B, DataflowPythonOp, is a Dataflow-based Python execution component and does not perform AutoML training or batch prediction. Option E, EndpointPredictOp, performs online prediction against a deployed Vertex AI endpoint, not batch prediction, so it does not fit the scenario.

Exam trap

PMLE often tests the confusion between online prediction (EndpointPredictOp) and batch prediction (BatchPredictOp), and between custom training (CustomJobRunOp) and AutoML training (AutoMLTabularTrainingJobRunOp) — candidates pick the wrong op for the stated workload.

3
MCQmedium

An ML engineer has a Vertex AI pipeline that trains a model and then evaluates it. The engineer wants the pipeline to automatically deploy the model to an endpoint only if the evaluation metric (e.g., accuracy) exceeds a threshold. The pipeline is defined using the Kubeflow Pipelines SDK. Which approach should the engineer use to implement this conditional deployment?

A.Set the deployment component's 'condition' parameter to the metric value, and rely on the component to skip itself if the value is below threshold.
B.Configure the pipeline to always deploy, and use a Cloud Function triggered by a Pub/Sub message from the evaluation step to roll back if the metric is below threshold.
C.Use a Vertex AI Model Evaluation component and set its 'deploy' flag to true, which automatically deploys if the metric meets the threshold.
D.Use a Vertex AI Pipelines condition with a comparison of the metric output to the threshold, and place the deployment component inside the condition's 'then' branch.
AnswerD

Vertex AI Pipelines supports dsl.Condition to control execution flow based on pipeline parameters or component outputs. By comparing the evaluation metric to a threshold, the deployment component runs only when the condition is true. This is the standard approach for conditional steps in Kubeflow Pipelines.

Why this answer

Conditional deployment in Vertex AI Pipelines is achieved by using dsl.Condition to evaluate a metric against a threshold. The deployment component is placed inside the condition's true branch, ensuring it only executes when the metric meets the requirement. This provides a clean, orchestrated way to gate deployment without external services or manual intervention.

Exam trap

The trap here is assuming that evaluation components or deployment components have built-in conditional flags, when actually the condition must be explicitly defined in the pipeline graph.

4
MCQmedium

An ML engineer is designing a Vertex AI Pipeline that includes a custom training component. The component must read a dataset from a Cloud Storage bucket and write the trained model to another Cloud Storage location. The engineer wants the component to be reusable across pipelines and to ensure that the pipeline tracks the exact dataset and model artifacts. Which approach should the engineer take?

A.Pass the Cloud Storage URIs as component inputs and outputs, and declare them as artifacts of type Dataset and Model.
B.Hardcode the Cloud Storage URIs inside the component's container code and use environment variables for configuration.
C.Use a single string parameter for the dataset URI and a single string parameter for the model URI, without declaring artifact types.
D.Mount a Cloud Storage bucket as a volume in the component container and write the model directly to the mounted path.
AnswerA

Declaring inputs and outputs as artifacts of type Dataset and Model allows Vertex AI Pipelines to track lineage and metadata automatically. The component remains reusable because the URIs are parameterized. This approach also enables the pipeline to visualize the artifacts and their relationships in the Vertex AI Pipelines UI, and supports artifact-based triggering and caching.

Why this answer

Using artifact inputs and outputs with types Dataset and Model ensures that Vertex AI Pipelines tracks lineage and metadata. It keeps the component reusable because the actual URIs are passed at runtime. Declaring artifacts also enables the pipeline to display them in the UI and to use them for caching and conditional execution.

Exam trap

The trap here is assuming that passing URIs as plain string parameters is sufficient for artifact tracking, when only declared artifacts provide lineage and metadata.

5
Multi-Selecteasy

A data scientist is creating a Vertex AI pipeline using the Kubeflow Pipelines SDK v2. Which TWO statements about pipeline parameters are correct? (Choose two.)

Select 2 answers
A.Pipeline parameters are defined as inputs to the pipeline function decorated with @dsl.pipeline.
B.Pipeline parameters must be serialized to JSON before use.
C.Pipeline parameters can only be of type str.
D.Pipeline parameters can be overridden at pipeline run time.
E.Pipeline parameters can be used to pass large datasets between components.
AnswersA, D

Declaring parameters as typed arguments on the @dsl.pipeline-decorated function is how KFP v2 compiles them into pipeline-level inputs, letting callers pass values per run. This satisfies the requirement that parameters be defined at the pipeline level rather than hard-coded inside individual components.

Why this answer

Option A is correct because in the Kubeflow Pipelines SDK v2, pipeline parameters are declared as typed arguments of the function decorated with @dsl.pipeline, which defines the pipeline's input interface. Option D is correct because these parameters are runtime inputs: when submitting a run (e.g., via the Vertex AI Pipelines API or the SDK's create_run_from_pipeline_func), you can supply new values that override the defaults specified in the pipeline definition. Option B is wrong because KFP v2 handles parameter serialization automatically; you do not manually JSON-serialize parameters before use.

Option C is wrong because KFP v2 parameters support multiple types such as int, float, bool, str, list, and dict, not only str. Option E is wrong because parameters are intended for small scalar/structured configuration values; large datasets should be passed between components as artifacts (e.g., Dataset, Model), not as parameters.

Exam trap

The trap is that candidates often assume pipeline parameters must be JSON-serialized or limited to strings due to older Kubeflow v1 conventions, but Vertex AI's Kubeflow Pipelines SDK v2 natively supports multiple Python types and automatic serialization.

6
Multi-Selecthard

A team is using Vertex AI Pipelines to orchestrate a training workflow. They want to ensure that the pipeline can be reproduced exactly six months later for auditing purposes. They need to capture all necessary information to rerun the pipeline and obtain identical results. (Choose two.)

Select 2 answers
A.Rely on the default caching behavior of Vertex AI Pipelines to reuse previous step outputs.
B.Use the same pipeline name and run it again with the same parameters.
C.Use a fixed pipeline template with pinned component versions and container image digests.
D.Store the pipeline parameters and input artifacts in a versioned location, such as a Cloud Storage bucket with versioning enabled.
E.Enable Vertex AI Experiments to log metrics and parameters for each run.
AnswersC, D

Pinning component versions and container image digests ensures that the exact same code and dependencies are used when the pipeline is rerun. This is critical for reproducibility because container images can be updated, and using digests guarantees immutability. Pinned versions also prevent unexpected changes in component behavior.

Why this answer

To reproduce a pipeline run exactly, you must pin the pipeline template and container images to immutable versions, and store the exact parameters and input artifacts in a versioned location. These two practices ensure that both the code and data are preserved, allowing an identical rerun. Other options like caching or experiment tracking do not provide the necessary immutability.

Exam trap

The trap here is assuming that caching or experiment tracking alone provides reproducibility, but they only log or reuse outputs without guaranteeing the exact code and data are preserved.

7
Multi-Selectmedium

Your organization wants to automate the retraining of a model when new data is available and also on a weekly schedule. Which TWO services would you use together to achieve this? (Choose two.)

Select 2 answers
A.Cloud Functions
B.Cloud Composer
C.Cloud Tasks
D.Dataflow
E.Cloud Scheduler
AnswersA, E

Cloud Functions provides the event-driven compute: a trigger fires when new data lands, invoking code that submits the retraining pipeline. This satisfies the stem's requirement to automate retraining on data arrival, complementing a scheduler for the weekly cadence.

Why this answer

Cloud Scheduler (E) is the correct service for the time-based trigger, since it can invoke a target on a cron schedule such as weekly, which satisfies the 'weekly schedule' requirement. Cloud Functions (A) is correct because it provides the serverless compute that actually executes the retraining logic, and it can be triggered both by Cloud Scheduler for the weekly run and by an event (for example, a Cloud Storage object finalize event when new data lands) for the data-availability requirement. Together, Cloud Scheduler fires on the cron schedule and Cloud Functions runs the retraining code, covering both triggers.

Cloud Composer (B) is a managed Airflow workflow orchestrator and could schedule DAGs, but it is heavier than needed and is not the marked pairing for this scenario. Cloud Tasks (C) is a queue for asynchronous task dispatch, not a cron scheduler or compute runtime, and Dataflow (D) is a managed Apache Beam service for data processing pipelines, not for triggering or hosting model-retraining logic.

Exam trap

Google often tests the distinction between orchestration (Cloud Composer) and simple scheduling/event-driven triggers (Cloud Scheduler + Cloud Functions), leading candidates to over-engineer the solution by choosing Cloud Composer when a lightweight combination suffices.

8
MCQeasy

A company wants to automatically retrain their model every night at 2 AM using Vertex AI Pipelines. Which approach should they use to trigger the pipeline on a schedule?

A.Use Cloud Scheduler to call the Vertex AI pipeline creation API
B.Deploy the pipeline as a Cloud Run job with a cron trigger
C.Use Vertex AI Experiments to schedule runs
D.Configure a cron job inside the pipeline definition
AnswerA

Cloud Scheduler provides cron-based triggering, satisfying the nightly 2 AM requirement. It invokes the Vertex AI Pipelines API endpoint directly, which compiles and runs the pipeline on the specified schedule without manual intervention. This decouples scheduling from pipeline logic, letting Vertex AI handle orchestration and execution.

Why this answer

Cloud Scheduler is the correct approach because it can directly invoke the Vertex AI Pipeline creation API via an HTTP trigger at a specified cron schedule (e.g., 2 AM daily). This integrates natively with Vertex AI's pipeline orchestration, allowing the scheduler to submit a pipeline run without additional infrastructure. The other options either lack native Vertex AI pipeline support or introduce unnecessary complexity.

Exam trap

A common mistake is confusing scheduling a pipeline run (using Cloud Scheduler + Vertex AI API) with scheduling tasks inside a pipeline (using cron within the pipeline definition). Neither Vertex AI Experiments nor Cloud Run jobs are designed for scheduled pipeline orchestration.

How to eliminate wrong answers

Option B is wrong because Cloud Run jobs are designed for stateless container execution and do not natively support Vertex AI Pipelines; they would require custom code to call the API, adding overhead and breaking the managed pipeline lifecycle. Option C is wrong because Vertex AI Experiments is used for tracking and comparing model training runs, not for scheduling or triggering pipeline executions. Option D is wrong because a cron job inside the pipeline definition would only schedule tasks within a single pipeline run, not trigger the pipeline itself on a recurring schedule.

9
Multi-Selecthard

A company is using Vertex AI Pipelines for ML workflows. They want to implement best practices for idempotent components and data passing. Which THREE practices should they adopt?

Select 3 answers
A.Pass large datasets between components using GCS URIs instead of in-memory values.
B.Avoid hard-coding file paths; use pipeline parameters to pass URIs.
C.Read data into memory in the first component and pass the in-memory object to subsequent components.
D.Use global variables in the pipeline code to store intermediate results.
E.Design components to be idempotent so that the same input always produces the same output.
AnswersA, B, E

Passing GCS URIs keeps components stateless and idempotent, since each run reads from an immutable location rather than relying on in-memory state that cannot survive retries or cross-component boundaries. This satisfies the stem's data-passing best practice for large datasets.

Why this answer

Option A is correct because Vertex AI Pipelines components exchange data through artifacts, and passing large datasets as GCS URIs (e.g., gs://bucket/path) avoids serializing huge in-memory objects into the pipeline's metadata/execution store, keeping runs efficient and reproducible. Option B is correct because hard-coded paths break portability and reproducibility across environments; using pipeline parameters (or input artifacts) to supply URIs lets the same pipeline definition run against dev, test, or prod buckets without code changes. Option E is correct because idempotent components—where identical inputs always yield identical outputs and re-execution has no additional side effects—are a core Vertex AI Pipelines best practice, enabling safe retries and caching.

Option C is not appropriate because passing in-memory objects between components requires serialization, can exceed metadata limits, and undermines the artifact-based data-passing model. Option D is not appropriate because global variables introduce hidden state, are not tracked as pipeline parameters or artifacts, and break reproducibility and caching across pipeline runs.

Exam trap

A common misconception in Vertex AI Pipelines is that in-memory data passing is acceptable in containerized components, but the correct pattern is to use GCS URIs and artifact references to ensure idempotency and scalability.

10
MCQmedium

An ML engineer is using Vertex AI Pipelines to orchestrate a workflow that includes a data preprocessing step, a training step, and an evaluation step. The evaluation step must run only if the training step succeeds and the evaluation metric meets a threshold. The engineer wants to define this logic natively in the pipeline without writing a custom component that exits with a specific code. Which Vertex AI Pipelines feature should they use?

A.Use a pipeline parameter of type 'bool' to control whether the evaluation step runs, and set it based on the metric before submitting the pipeline.
B.Configure the pipeline's 'exit handler' to check the metric and either continue or fail the pipeline.
C.Use the kfp.dsl.Condition context manager to conditionally execute the evaluation step based on the metric value.
D.Set the evaluation step's 'trigger' policy to 'on_success' and pass the metric threshold as a pipeline parameter.
AnswerC

The kfp.dsl.Condition context manager allows you to define conditional execution in a pipeline based on the value of a pipeline parameter or a task output. You can compare the evaluation metric to a threshold and conditionally run subsequent steps. This is a native KFP feature and does not require custom exit-code logic, making it the correct approach for conditional execution based on a metric.

Why this answer

Vertex AI Pipelines supports conditional execution through the kfp.dsl.Condition context manager. This allows you to branch based on the value of a task output, such as an evaluation metric. You can wrap the evaluation step in a Condition that checks if the metric exceeds a threshold.

The other options either misuse trigger policies, misunderstand exit handlers, or rely on static pipeline parameters, none of which provide runtime conditional execution based on a metric.

Exam trap

The trap here is confusing task trigger policies (which are about upstream success/failure) with conditional execution based on data or metrics.

11
Multi-Selecthard

An ML engineer is designing a Vertex AI pipeline that includes a custom training component. The component must read training data from a Cloud Storage bucket and write the trained model to a Vertex AI Model Registry. The engineer wants to ensure the component can access these resources securely. Which two configurations should the engineer implement? (Choose two.)

Select 2 answers
A.Use the default Compute Engine service account for all pipeline runs.
B.Set the component's environment variable GOOGLE_APPLICATION_CREDENTIALS to point to a mounted secret.
C.Specify the pipeline's service account when submitting the pipeline run.
D.Embed a service account key file in the component's container image.
E.Grant the Vertex AI Pipelines service account the necessary IAM roles for Cloud Storage and Vertex AI.
AnswersC, E

When submitting a pipeline run, you can specify a service account that the pipeline will use. This service account's permissions determine what resources the components can access. By specifying a service account with the right roles for Cloud Storage and Vertex AI, the engineer ensures secure access. This is the correct way to manage authentication for pipeline components in Vertex AI.

Why this answer

To securely access Cloud Storage and Vertex AI Model Registry, the pipeline's service account must have the required IAM roles, and that service account must be specified when submitting the pipeline run. This ensures the component authenticates with the correct permissions without embedding credentials.

Exam trap

The trap here is thinking that embedding credentials or using environment variables is necessary, when in fact Vertex AI Pipelines uses the pipeline's service account for authentication.

12
MCQeasy

A machine learning engineer is using Vertex AI Pipelines and wants to run a custom Python function as a component. They need to pass a dataset artifact from a previous component and output a model artifact. Which decorator should they use to define the component in the Kubeflow Pipelines SDK v2?

A.@dsl.task
B.@dsl.pipeline
C.@dsl.component
D.@dsl.container
AnswerC

The @dsl.component decorator converts a plain Python function into a KFP v2 pipeline component, with typed parameters and artifacts. It supports declaring an input dataset artifact and an output model artifact, satisfying the stem's requirement for a custom Python component.

Why this answer

The correct decorator is @dsl.component because in Kubeflow Pipelines SDK v2, this decorator is used to define a custom Python function as a reusable pipeline component. It automatically handles input and output artifact serialization, such as passing a dataset artifact from a previous component and outputting a model artifact, by leveraging the component's type annotations and the KFP artifact system.

Exam trap

Candidates often confuse @dsl.component (for custom Python functions with artifact I/O) with @dsl.container (for pre-built container images) when the question emphasizes running a custom Python function.

How to eliminate wrong answers

Option A is wrong because @dsl.task is not a valid decorator in Kubeflow Pipelines SDK v2; it is a concept from Vertex AI custom jobs, not for defining pipeline components. Option B is wrong because @dsl.pipeline is used to define the entire pipeline graph, not an individual component function. Option D is wrong because @dsl.container is used to define a component that runs a container image directly, not a custom Python function with artifact handling.

13
Multi-Selecthard

A company wants to implement a CI/CD pipeline for their ML models using Vertex AI. They need to automatically retrain the model when new data arrives, but only if the model performance on a validation set has degraded by more than 5% compared to the current production model. Which three services or components should they incorporate into the automated pipeline? (Choose three.)

Select 3 answers
A.Dataflow pipeline to clean the new data before training
B.Vertex AI Evaluation component to compute model performance metrics on the validation set
C.Cloud Functions to trigger the pipeline when new data arrives in Cloud Storage
D.Vertex AI Model Registry alias update to promote the model if performance passes the threshold
E.Cloud Scheduler to run the pipeline on a fixed schedule
AnswersB, C, D

The Vertex AI Evaluation component computes metrics on the validation set, producing the performance figures needed to compare against the production model. This satisfies the stem's 5% degradation gate by supplying the quantitative basis for the retrain decision.

Why this answer

Option B is correct because the Vertex AI Evaluation component is the mechanism that computes model performance metrics (such as accuracy, AUC, or RMSE) on a validation set, which is exactly what is needed to compare the newly trained model against the current production model and detect a degradation greater than 5%. Option C is correct because Cloud Functions can be configured with a Cloud Storage trigger (via Eventarc) to fire when new data objects arrive in a bucket, thereby automatically kicking off the retraining pipeline without manual intervention. Option D is correct because the Vertex AI Model Registry uses aliases (for example, a 'production' alias) to point to a specific model version, and updating that alias is how a newly validated model gets promoted to production once it passes the performance threshold.

Option A is not required by the scenario, since the question focuses on triggering, evaluating, and promoting models rather than on data cleaning, and no data quality issue is stated. Option E is incorrect because Cloud Scheduler runs jobs on a fixed time-based schedule, whereas the requirement is event-driven retraining triggered by the arrival of new data.

Exam trap

In Google Cloud, the distinction between event-driven triggers (Cloud Functions/Eventarc) and scheduled triggers (Cloud Scheduler) is commonly tested. Candidates often mistakenly choose Cloud Scheduler when the requirement is for an event-driven retraining pipeline triggered by new data arrival.

14
MCQmedium

An ML engineer is building a pipeline component that takes a dataset URI and a model URI as inputs, and outputs a classification metrics artifact. Which KFP SDK v2 type should the output artifact be annotated with?

A.Dataset
B.Metrics
C.ClassificationMetrics
D.Model
AnswerC

ClassificationMetrics is the KFP v2 artifact type purpose-built for classification evaluation output, carrying fields such as confusion matrix, precision and recall. Annotating the output with it satisfies the stem's requirement for a classification metrics artifact.

Why this answer

In KFP SDK v2, the `ClassificationMetrics` type is specifically designed to output classification metrics such as confusion matrix, ROC curve, and AUC. The question asks for a component that outputs classification metrics, so `ClassificationMetrics` is the correct artifact type. Using `Metrics` would be too generic and not provide the structured schema needed for classification-specific visualizations in the KFP UI.

Exam trap

The trap here is that candidates confuse the generic `Metrics` type (which handles scalar values) with the specialized `ClassificationMetrics` type, not realizing that KFP SDK v2 requires the specific artifact type to enable proper UI rendering and schema validation for classification outputs.

How to eliminate wrong answers

Option A is wrong because `Dataset` is used for input or output of tabular data, not for metrics artifacts. Option B is wrong because `Metrics` is a generic artifact for scalar metrics (e.g., accuracy, loss) but lacks the structured fields (e.g., confusion matrix, ROC) required for classification metrics; it would not render classification-specific visualizations in the KFP UI. Option D is wrong because `Model` is used for serialized model artifacts, not for evaluation metrics.

15
MCQmedium

An organization runs a Vertex AI pipeline that includes a model evaluation step. Team members want to reuse previously computed evaluation metrics when re-running the pipeline with unchanged code and hyperparameters. Which feature should they enable?

A.Manually store outputs in Cloud Storage and check for existence
B.Enable pipeline caching (default behavior)
C.Use the importer component to fetch previous results
D.Disable caching for the evaluation component
AnswerB

Pipeline caching reuses outputs from previously executed components when the pipeline definition, code, and inputs are unchanged, so the evaluation step's metrics are not recomputed. Enabling it satisfies the stem's requirement to reuse prior evaluation metrics across identical re-runs.

Why this answer

Vertex AI Pipelines has caching enabled by default. When a pipeline is re-run with unchanged code, hyperparameters, and inputs, the evaluation step will reuse the cached output from the previous run, saving time and cost. Team members do not need to manually store outputs or disable caching.

Exam trap

The trap is overcomplicating the solution by suggesting manual storage or importer components; candidates often forget that Vertex AI Pipelines caching is enabled by default and automatically reuses outputs when inputs and code are unchanged.

How to eliminate wrong answers

Option A is wrong because manually storing outputs in Cloud Storage and checking for existence is a custom workaround that is unnecessary when pipeline caching is already available and enabled by default. Option C is wrong because the importer component is used to import external artifacts into a pipeline, not to reuse previously computed evaluation metrics from a prior run. Option D is wrong because disabling caching would force the evaluation step to re-run, which is the opposite of what they want.

16
MCQhard

A team has a pipeline that trains a model and then evaluates it. They want to conditionally deploy the model to a staging endpoint only if evaluation metrics exceed a threshold. Which KFP feature should they use?

A.Use dsl.Condition (deprecated) or dsl.If to check metrics and conditionally run deployment.
B.Use dsl.ParallelFor to evaluate and deploy in parallel.
C.Use an exit handler to deploy regardless of metrics.
D.Split the pipeline into two separate pipelines and run the second only if metrics are good.
AnswerA

dsl.If evaluates the evaluation task's output metric at runtime and gates the deployment task on the threshold, so deployment runs only when metrics pass. dsl.Condition is the deprecated predecessor; dsl.If is the current KFP SDK v2 construct satisfying this conditional-deployment requirement.

Why this answer

KFP provides `dsl.Condition` (deprecated) and `dsl.If` as first-class pipeline constructs to conditionally execute pipeline components based on runtime metrics or other pipeline outputs. By wrapping the deployment step inside a `dsl.If` block that checks whether evaluation metrics exceed a threshold, the pipeline can deploy the model to a staging endpoint only when the condition is met, avoiding unnecessary deployments for underperforming models.

Exam trap

Google often tests the distinction between conditional execution (`dsl.If`) and unconditional execution patterns (exit handlers, parallel loops), tempting candidates to choose a pattern that always runs the deployment step or runs it in parallel without any gate.

How to eliminate wrong answers

Option B is wrong because `dsl.ParallelFor` is designed for iterating over a collection of items to execute the same component in parallel, not for conditionally executing a component based on a runtime evaluation result. Option C is wrong because an exit handler (e.g., `dsl.ExitHandler`) always runs a specified component when the pipeline exits, regardless of success or failure, so it would deploy the model even if metrics are poor, which contradicts the requirement. Option D is wrong because splitting the pipeline into two separate pipelines loses the benefit of a single orchestrated workflow; it introduces manual coordination, external state management, and additional operational complexity, whereas KFP’s conditional constructs handle this natively within one pipeline.

17
MCQmedium

An ML engineer is authoring a Vertex AI Pipelines component that runs a custom Python script. The component must accept a GCS path to training data and output a model artifact. The engineer wants the component interface to be strongly typed and to automatically generate the component specification from the Python function. Which approach should the engineer use?

A.Define the component using the @component decorator from the google.cloud.aiplatform.v1alpha1 package, specifying the input and output types as function annotations.
B.Use the google.cloud.aiplatform.CustomContainerTrainingJob class to package the script and run it as a pipeline step.
C.Write a Dockerfile that installs the required dependencies and exposes the script as an entrypoint, then build and push the image to Artifact Registry.
D.Define the component using the @dsl.component decorator from the kfp package, annotating the function parameters with types like str and Output[Model].
AnswerD

The kfp.dsl.component decorator (Kubeflow Pipelines SDK v2) generates a component specification from the Python function's type annotations. It supports InputPath, OutputPath, and artifact types like Model, ensuring strong typing. This is the standard method for authoring lightweight Python components in Vertex AI Pipelines.

Why this answer

The Kubeflow Pipelines SDK v2 @dsl.component decorator introspects Python type annotations to build a component specification, including input and output artifacts. This provides strong typing and automatic generation, which is exactly what the engineer needs. Other approaches either require manual YAML definition or are intended for different use cases like custom training jobs.

Exam trap

The trap here is assuming that any decorator or containerization automatically generates a typed component interface, when only the KFP v2 @dsl.component decorator does so from annotations.

18
Multi-Selecthard

An ML team is using Vertex AI Pipelines to orchestrate a training workflow. They need to pass a large dataset (500 GB) between two components. The first component preprocesses the data and writes the output to Cloud Storage. The second component trains a model using that preprocessed data. The team wants to minimize pipeline execution time and cost. Which two strategies should they use? (Choose two.)

Select 2 answers
A.Mount a Cloud Storage bucket as a Persistent Disk to both components so they can share the data via a common file system.
B.Serialize the preprocessed data into a single TFRecord file and pass it as a pipeline parameter to the second component.
C.Have the first component write the preprocessed data to a Cloud Storage location and output the URI as a string parameter to the second component.
D.Use a Dataset artifact to represent the preprocessed data and pass it as an output of the first component and input to the second.
E.Use a Model artifact to pass the preprocessed data to the training component, and have the training component load the data from the artifact's URI.
AnswersC, D

Passing the Cloud Storage URI as a string parameter is lightweight and allows the second component to read the data directly from Cloud Storage. This avoids data duplication and keeps the pipeline spec small. It is a common pattern when a dedicated artifact type is not necessary, though using a Dataset artifact is more semantically rich.

Why this answer

For large data, the pipeline should pass references (URIs) rather than the data itself. Using a Dataset artifact or a string parameter to convey the Cloud Storage location lets each component access the data independently, minimizing data movement and pipeline overhead. Serializing or sharing via disks is not feasible at this scale and would increase cost and time.

Exam trap

The trap here is thinking that large data can be passed directly between components as parameters or that shared storage can be mounted, when in fact only lightweight references should be passed.

19
Multi-Selecthard

A team wants to implement CI/CD for their ML pipeline using Cloud Build. They want to automatically compile and deploy the pipeline when code is pushed to the main branch. Which three steps should they include in the Cloud Build configuration? (Choose three.)

Select 3 answers
A.Create or update the pipeline in Vertex AI using the compiled file
B.Upload the compiled pipeline to Cloud Storage
C.Run the pipeline immediately after deployment
D.Install KFP SDK and compile the pipeline
E.Configure Cloud Scheduler to trigger on push
AnswersA, B, D

Creating or updating the pipeline in Vertex AI registers the compiled pipeline definition, making it runnable and schedulable. This satisfies the deployment half of the CI/CD requirement, since compilation alone leaves the pipeline unregistered and unable to execute on Vertex AI.

Why this answer

Option D is correct because the Cloud Build configuration must first set up the environment by installing the Kubeflow Pipelines (KFP) SDK and then compile the pipeline definition into a compiled YAML/JSON artifact, which is the essential build step for a CI/CD ML pipeline. Option B is correct because the compiled pipeline artifact needs to be uploaded to a Cloud Storage bucket so it can be referenced and used by Vertex AI when creating or updating the pipeline. Option A is correct because the final deployment step is to create or update the pipeline in Vertex AI using the compiled file, which is the actual CD action that makes the pipeline available in Vertex AI Pipelines.

Option C is not correct because running the pipeline immediately after deployment is not a required CI/CD configuration step; deployment and execution are separate concerns, and the scenario only asks to compile and deploy on push. Option E is not correct because Cloud Scheduler is used for time-based triggering, whereas the scenario requires triggering on code push to the main branch, which is handled by Cloud Build triggers, not Cloud Scheduler.

Exam trap

A common trap is confusing build-time actions (compilation, upload, registration) with runtime actions (execution, scheduling). Candidates often mistakenly include immediate pipeline execution as a CI/CD step instead of focusing on deploying and registering the pipeline artifact.

20
Multi-Selecthard

A team is designing a ML pipeline that includes training, evaluation, and conditional deployment. They want to use Vertex AI Pipelines. Which THREE concepts should they use? (Choose three.)

Select 3 answers
A.Artifact types (e.g., Model, Metrics) for passing outputs
B.Manual approval via Cloud Console
C.Cloud SQL for storing intermediate results
D.Pre-built Google Cloud Pipeline Components for training and evaluation
E.dsl.If for conditional execution
AnswersA, D, E

Artifact types such as Model and Metrics carry typed outputs between pipeline steps, letting the evaluation component consume the trained model and emit metrics that the conditional deployment step reads. This satisfies the stem's need to pass outputs across training, evaluation and conditional deployment.

Why this answer

Option A is correct because Vertex AI Pipelines is built on ML Metadata, and typed artifacts such as Model, Metrics, Dataset, and Artifact let components pass structured outputs between training and evaluation steps so downstream steps and lineage tracking work correctly. Option D is correct because pre-built Google Cloud Pipeline Components (e.g., CustomTrainingJobOp, ModelEvaluationOp) provide ready-made, versioned steps for training and evaluation, reducing boilerplate and integrating natively with Vertex AI services. Option E is correct because conditional deployment requires branching logic in the pipeline graph, which is expressed with the Kubeflow Pipelines DSL construct dsl.If (or dsl.Condition) to run a deployment component only when evaluation metrics meet a threshold.

Option B is not appropriate because manual approval via the Cloud Console is not a Vertex AI Pipelines concept for conditional deployment; gating is done programmatically in the pipeline DAG. Option C is not appropriate because intermediate results in Vertex AI Pipelines are passed as artifacts and metadata, not stored in Cloud SQL, which is a relational database service unrelated to pipeline data flow.

Exam trap

PMLE often tests whether candidates confuse pipeline orchestration concepts with general GCP services, so distractors like Cloud SQL or manual approval must be recognized as non-pipeline constructs.

21
MCQmedium

An ML engineer is building a Vertex AI pipeline that must run a custom training component for each of 12 hyperparameter combinations. The component is defined as a custom Python function (Lightweight Python component). The engineer wants each combination to run as a separate parallel task so the pipeline completes faster, and wants the pipeline to fail fast if any single trial fails. Which approach should the engineer take?

A.Wrap the training component in a single Vertex AI CustomJob with a hyperparameter tuning job specification and let the service manage trials.
B.Create 12 separate pipelines, one per hyperparameter combination, and schedule them with Cloud Scheduler at the same time.
C.Use a ParallelFor loop over the hyperparameter list and set the component's retry policy to 0.
D.Define a sequential for-loop in the pipeline function that calls the training component 12 times in order.
AnswerC

A ParallelFor loop in Vertex AI Pipelines (using dsl.ParallelFor) unrolls the loop into independent parallel tasks, one per hyperparameter combination, which is exactly the fan-out pattern needed for 12 trials. Setting the retry policy to 0 ensures that when any trial fails, the pipeline does not silently retry and mask the failure, so it fails fast as required.

Why this answer

Vertex AI Pipelines supports dsl.ParallelFor to fan out a list of values into parallel component tasks in one DAG, which is the idiomatic way to run 12 hyperparameter combinations concurrently. Setting retries to 0 on the component ensures a failed trial stops the pipeline rather than being retried and hidden. The other approaches either collapse trials into a single job, serialize them, or split them across pipelines, none of which meet the parallel fan-out and fail-fast requirements.

Exam trap

The trap here is assuming that a Vertex AI Hyperparameter Tuning CustomJob is the same as a pipeline ParallelFor fan-out, when the former is a single managed job and the latter is a pipeline-level DAG construct.

22
MCQmedium

An ML team wants to run a hyperparameter tuning job on Vertex AI using a pre-built pipeline component. Which component should they use?

A.AutoMLTabularTrainingJobRunOp
B.CustomTrainingJobRunOp with hyperparameter arguments.
C.ModelTrainComponent
D.HyperparameterTuningJobRunOp
AnswerD

HyperparameterTuningJobRunOp is the pre-built pipeline component that wraps Vertex AI's HyperparameterTuningJob, launching a tuning job with the specified search space and metrics. It satisfies the stem's requirement to run tuning via a pre-built component rather than custom code.

Why this answer

The HyperparameterTuningJobRunOp is the correct pre-built Vertex AI pipeline component specifically designed to launch a hyperparameter tuning job. It wraps the Vertex AI HyperparameterTuningJob API, allowing you to specify the worker pool spec, metric target, and parameter specifications directly within a Kubeflow Pipelines (KFP) or Vertex AI Pipelines orchestration context.

Exam trap

A common mistake on the Google PMLE exam is confusing the pre-built HyperparameterTuningJobRunOp with CustomTrainingJobRunOp that accepts hyperparameter arguments, but the latter requires manual tuning logic rather than leveraging the built-in hyperparameter tuning service.

How to eliminate wrong answers

Option A is wrong because AutoMLTabularTrainingJobRunOp is used to launch an AutoML training job for tabular data, which does not support custom hyperparameter tuning; it uses AutoML's own search. Option B is wrong because CustomTrainingJobRunOp with hyperparameter arguments is not a pre-built component for tuning; it launches a single custom training job and would require manual orchestration to implement a tuning loop, whereas the question asks for a pre-built component. Option C is wrong because ModelTrainComponent is not a standard pre-built Vertex AI pipeline component; it is a generic name that does not correspond to any official Vertex AI component, and using it would require custom implementation.

23
MCQeasy

A machine learning engineer is building a Vertex AI pipeline that uses a pre-built Google Cloud Pipeline Components (GCPC) to train a custom model. Which component should the engineer use to submit a custom training job to Vertex AI?

A.HyperparameterTuningJob
B.CustomJob
C.BatchPredictionJob
D.ModelDeploy
AnswerB

The CustomJob component submits a custom training job to Vertex AI, letting the pipeline run bespoke training code within a managed job rather than a pre-built trainer. It satisfies the requirement to launch a custom training workload as a pipeline step.

Why this answer

The CustomJob component is the correct choice because it is the pre-built GCPC component specifically designed to submit a custom training job to Vertex AI. It allows the engineer to specify a custom container image or a Python training script, along with machine configuration and hyperparameters, directly within a Vertex AI pipeline. Other components serve different purposes, such as hyperparameter tuning, batch predictions, or model deployment.

Exam trap

The trap here is that candidates may confuse HyperparameterTuningJob with CustomJob because both involve training, but HyperparameterTuningJob is for multi-trial optimization, not a single training run, and ModelDeploy is a distractor that does not exist as a GCPC component.

How to eliminate wrong answers

Option A is wrong because HyperparameterTuningJob is used for optimizing hyperparameters across multiple trials, not for submitting a single custom training job. Option C is wrong because BatchPredictionJob is for running batch predictions on a trained model, not for training. Option D is wrong because ModelDeploy is not a standard GCPC component; the correct component for deploying a model to an endpoint is ModelDeployer or a similar deployment component, and ModelDeploy does not exist in the GCPC library.

24
MCQmedium

An ML engineer is designing a Vertex AI pipeline that trains a model, evaluates it, and conditionally deploys it only if the evaluation metric meets a threshold. The pipeline must pass the evaluation metric from the evaluation component to a downstream conditional. Which Vertex AI Pipelines feature should the engineer use to implement this flow?

A.Use the Vertex AI Model Registry to compare the new model's metric with the deployed model's metric and let the registry automatically deploy if better.
B.Write the metric to a Cloud Storage file and use a Cloud Function to trigger a separate pipeline for deployment.
C.Use the output of the evaluation component as an input parameter to a condition in the pipeline definition, referencing the component's output parameter.
D.Store the metric in a Vertex AI Experiment run and then use the Experiment's metric in the condition.
AnswerC

Vertex AI Pipelines allows a component to expose outputs as parameters that can be consumed by downstream conditions. The evaluation component can output a metric value as a pipeline parameter, and the condition can compare this value to a threshold. This is the standard way to implement conditional logic based on runtime values in a pipeline.

Why this answer

In Vertex AI Pipelines, component outputs can be promoted to pipeline parameters, which can then be used in conditions to control downstream execution. The evaluation component should output the metric as a parameter, and the condition compares it to the threshold. This maintains a single, traceable pipeline run and avoids external services.

Exam trap

The trap here is assuming that Vertex AI Experiments or Model Registry can directly control pipeline flow based on metrics, when in fact they are for tracking and storage, not runtime decision-making.

25
MCQmedium

An ML engineer has a Vertex AI pipeline that trains a model and then evaluates it. The engineer wants the pipeline to deploy the model to an endpoint only if the evaluation metric exceeds a threshold defined at pipeline submission time. The threshold must be changeable without recompiling the pipeline. Which mechanism should the engineer use?

A.Set the threshold as an environment variable in the pipeline's service account and read it inside the deployment component.
B.Hardcode the threshold in the evaluation component and use a dsl.Condition on the evaluation component's output.
C.Define the threshold as a pipeline input parameter and use a dsl.Condition on that parameter to gate the deployment component.
D.Use a Vertex AI Model Registry alias and configure the endpoint to only serve models whose alias matches the threshold value.
AnswerC

A pipeline input parameter is part of the pipeline's runtime interface, so the threshold can be supplied at submission time without recompiling. Wrapping the deployment component in dsl.Condition on that parameter creates the conditional gate. This is exactly the pattern Vertex AI Pipelines supports for runtime-configurable branching.

Why this answer

Vertex AI Pipelines distinguishes compile-time structure from runtime parameters. Making the threshold a pipeline input parameter allows it to be supplied at submission time without recompiling, and dsl.Condition on that parameter creates the deployment gate. The other options either bake the threshold into the pipeline definition, misuse Model Registry aliases, or rely on an unsupported environment variable mechanism that the compiler cannot see.

Exam trap

The trap here is confusing a dsl.Condition on a component output with a condition on a runtime parameter, when only the latter allows the threshold to change without recompiling.

26
MCQmedium

You are using KFP SDK v2 to define a pipeline. You need to pass a large dataset between components. What is the best practice for passing data?

A.Use the component's temporary directory to share data between containers.
B.Pass the data as a serialized Python object in memory.
C.Write the data to Cloud Storage and pass the GCS URI as an artifact.
D.Store the data in a BigQuery table and pass the table reference.
AnswerC

Writing data to Cloud Storage and passing the GCS URI as an artifact avoids embedding large payloads in pipeline metadata, satisfying the large-dataset constraint. KFP passes lightweight artifact references between components rather than the data itself.

Why this answer

In KFP SDK v2, passing large datasets between components is best done by writing the data to Cloud Storage and passing the GCS URI as an artifact. This approach leverages KFP's built-in artifact tracking, ensures data persistence across container restarts, and avoids memory or disk limitations of ephemeral containers. The artifact is automatically serialized and passed as an input/output parameter, enabling efficient, scalable data exchange.

Exam trap

Google often tests the misconception that temporary directories are shared between containers in a pod, but in KFP each component runs in its own container with isolated storage, making Cloud Storage the correct choice for durable, cross-component data sharing.

How to eliminate wrong answers

Option A is wrong because a component's temporary directory is ephemeral and not shared between containers; each container runs in its own isolated filesystem, so data written there is lost after the component finishes. Option B is wrong because passing a serialized Python object in memory is limited by the container's memory capacity and cannot handle large datasets; KFP does not support in-memory object passing between components. Option D is wrong because storing data in a BigQuery table and passing the table reference is overkill for intermediate pipeline data; it introduces unnecessary latency, cost, and complexity compared to using Cloud Storage artifacts, which are the standard for KFP artifact passing.

27
Multi-Selecthard

A company runs a Vertex AI pipeline that uses a container component to preprocess data. The component downloads a large file from a public URL and saves the output to Cloud Storage. The pipeline fails intermittently with a 'timeout' error. Which THREE steps should the team take to improve reliability? (Choose three.)

Select 3 answers
A.Make the component idempotent by checking for existing output before processing.
B.Increase the component's timeout setting.
C.Reduce the size of the file being downloaded.
D.Implement retries with exponential backoff in the component.
E.Increase the machine type for the component.
AnswersA, B, D

Idempotency lets a retried component detect that output already exists in Cloud Storage and skip re-downloading, so repeated attempts after a timeout do not duplicate work or waste the timeout window. This directly addresses the intermittent timeout failures.

Why this answer

Option A is correct because making the component idempotent (e.g., checking whether the output already exists in Cloud Storage before re-downloading and reprocessing) prevents redundant work on retries and ensures a re-run after a transient timeout does not corrupt or duplicate results. Option B is correct because the intermittent 'timeout' error indicates the component's configured timeout is too short for the large download plus preprocessing; raising the timeout setting gives the component enough time to finish on slow runs. Option D is correct because implementing retries with exponential backoff inside the component lets it recover from transient network or service failures when downloading from the public URL, which is the typical cause of intermittent timeouts.

Option C is not appropriate because reducing the file size changes the data being processed and is not a reliability fix the team can simply apply to the existing pipeline. Option E is not appropriate because increasing the machine type addresses CPU or memory constraints, not a timeout caused by download duration or transient network failures.

Exam trap

Google exams often test the distinction between fixing the symptom (increasing timeout) and addressing the root cause (idempotency and retries), leading candidates to overlook that idempotency and retries together with a reasonable timeout form the most robust solution.

28
MCQhard

An ML engineer is using Vertex AI Pipelines to orchestrate a workflow that includes a custom component for data validation. The component takes a dataset URI and outputs a validation report. The engineer wants to fail the pipeline immediately if the validation report indicates that the data is invalid, without running subsequent steps. How should the engineer implement this?

A.Use a conditional step after validation that checks the report and only proceeds if valid, otherwise skips subsequent steps.
B.Have the data validation component raise an exception if the data is invalid, causing the task to fail and the pipeline to stop.
C.Write the validation result to a Cloud Storage bucket and use a Cloud Function to monitor the bucket and cancel the pipeline if invalid.
D.Configure the pipeline to use a 'fail_fast' execution option that stops the pipeline if any component outputs a failure status.
AnswerB

If the data validation component raises an exception when data is invalid, the task fails, and Vertex AI Pipelines will not execute downstream steps that depend on its output. This is the simplest and most direct way to halt the pipeline upon validation failure. The pipeline status will reflect the failure, allowing for alerting and debugging.

Why this answer

The most effective way to fail the pipeline immediately upon invalid data is to have the data validation component raise an exception. This causes the task to fail, and Vertex AI Pipelines will not run any downstream tasks that depend on its output. The pipeline will be marked as failed, which can trigger alerts.

This approach is simple, reliable, and leverages the native failure handling of the pipeline orchestrator.

Exam trap

The trap here is thinking that a conditional step or an external monitor is needed to stop the pipeline, when simply failing the component achieves the goal.

29
MCQeasy

What is the purpose of the 'importer' component in Vertex AI Pipelines?

A.To import data from external sources into the pipeline.
B.To import existing ML artifacts (e.g., models, datasets) into a pipeline as inputs.
C.To import Python libraries into the pipeline environment.
D.To import pipeline definitions from other projects.
AnswerB

The importer component registers an existing artefact, such as a model or dataset already in Cloud Storage or Artifact Registry, as a pipeline input without rerunning upstream steps. This lets pipelines consume pre-existing ML artefacts as declared inputs.

Why this answer

The Importer component in Vertex AI Pipelines (a Kubeflow Pipelines component) brings existing ML artifacts — such as pre-trained models, datasets, or other resources already registered in Vertex AI — into a pipeline as inputs. It creates an artifact reference so downstream components can consume it without regenerating it. This is essential for reusing models or datasets across pipeline runs.

Exam trap

PMLE often tests whether candidates confuse the Importer (existing Vertex AI artifacts) with data ingestion components (external raw data) — the word 'import' is deliberately ambiguous.

How to eliminate wrong answers

Option A is wrong because importing raw data from external sources is done by data ingestion components (e.g., custom Python components or BigQuery loaders), not the Importer — the Importer works with existing Vertex AI artifacts, not arbitrary external data. Option C is wrong because Python library imports are handled by container images and requirements files, not by a pipeline component. Option D is wrong because pipeline definitions are compiled and submitted via the Vertex AI SDK or templates; the Importer does not import pipeline definitions from other projects.

30
MCQhard

In a Vertex AI Pipeline, a component produces a Metrics artifact that includes an evaluation metric. The engineer wants to use this metric value as a condition to decide whether to deploy the model. However, the metric value is stored in the artifact's metadata and not directly as a pipeline parameter. How can the engineer pass the metric value to a downstream conditional task?

A.Configure the component that produces the Metrics artifact to also output the metric as a pipeline parameter.
B.Use the importer component to convert the artifact into a parameter.
C.Add a component that reads the artifact's metadata and outputs the metric as a parameter, then use that parameter in the condition.
D.Use the artifact directly in the dsl.If condition, as artifacts are comparable.
AnswerC

Metrics stored in artifact metadata are not pipeline parameters, so dsl.If cannot consume them directly. A component that reads the artifact's metadata and emits the metric as an output parameter makes the value usable in the downstream condition, satisfying the stem's constraint.

Why this answer

Vertex AI Pipeline conditions require pipeline parameters (typed values) to evaluate expressions like `dsl.If`. A Metrics artifact's metadata is stored as an artifact property, not a pipeline parameter, so a custom component must read that metadata and output the metric as a parameter. This parameter can then be used in the `dsl.If` condition to control downstream deployment.

Exam trap

A common trap in this Google exam is the misconception that artifact metadata can be directly used in pipeline conditions, but conditions require typed parameters, not artifact objects or their metadata fields.

How to eliminate wrong answers

Option A is wrong because modifying the upstream component to output the metric as a pipeline parameter would require changing the component's implementation, which may not be feasible if the component is from a shared library or third-party. Option B is wrong because the importer component is designed to bring external artifacts into the pipeline, not to extract metadata from an existing artifact and convert it into a parameter. Option D is wrong because artifacts are not directly comparable in `dsl.If` conditions; conditions only work with pipeline parameters (e.g., integers, strings), not with artifact objects or their metadata.

31
MCQeasy

Which of the following is a best practice when designing idempotent pipeline components in Vertex AI?

A.Use global variables to share state between components.
B.Pass data through Cloud Storage URIs rather than in-memory.
C.Write component outputs to a database with timestamps.
D.Use the same output name for all runs to avoid duplication.
AnswerB

Cloud Storage URIs persist data outside component memory, so a retried or rerun component reads the same immutable input and writes the same output, satisfying idempotency. In-memory passing loses state between container executions and cannot guarantee reproducible reruns.

Why this answer

Passing data through Cloud Storage URIs ensures that component outputs are stored persistently and can be retrieved by downstream components, even if the original component instance is terminated or scaled down. This aligns with the principle of idempotency because the same input will always produce the same output stored at the same URI, and re-running the component will not cause side effects or data loss. In contrast, in-memory data is ephemeral and tied to a specific runtime instance, breaking idempotency across retries or parallel executions.

Exam trap

The Google PMLE exam often tests the misconception that idempotency is about avoiding duplication of output names or using timestamps for uniqueness, when in fact idempotency requires that repeated executions produce the same result without side effects, which is achieved by using immutable, deterministic storage like Cloud Storage URIs rather than mutable state or time-dependent writes.

How to eliminate wrong answers

Option A is wrong because using global variables to share state between components introduces mutable shared state that can cause non-deterministic behavior across retries or parallel runs, violating idempotency. Option C is wrong because writing component outputs to a database with timestamps introduces a side effect that changes with each run (different timestamps), making the component non-idempotent; idempotent components should produce the same output regardless of how many times they are executed. Option D is wrong because using the same output name for all runs does not guarantee idempotency; it can lead to overwriting or collision of outputs, and idempotency requires that repeated executions produce the same result without unintended side effects, not just the same output name.

32
MCQeasy

A data scientist is defining a Vertex AI pipeline and needs to include a step that imports a pre-existing model from Cloud Storage into the pipeline as an artifact. Which Kubeflow Pipelines SDK v2 component should they use?

A.dsl.Collected
B.dsl.importer
C.dsl.Importer
D.dsl.Artifact
AnswerB

dsl.importer registers an existing Cloud Storage artefact, such as a trained model, as a pipeline artefact without executing a container to recreate it. This satisfies the requirement to bring a pre-existing model into the pipeline graph as a usable input.

Why this answer

The `dsl.importer` component in Kubeflow Pipelines SDK v2 is specifically designed to import existing artifacts (such as models, datasets, or metrics) from external storage (e.g., Cloud Storage) into a pipeline as a pipeline artifact. It allows you to reference a pre-existing model without retraining or re-uploading, making it the correct choice for this use case.

Exam trap

The trap here is that candidates may confuse the Python class naming convention (capitalized `Importer`) with the actual SDK v2 function name (lowercase `importer`), or mistakenly think `dsl.Artifact` can import artifacts when it only defines the artifact schema.

How to eliminate wrong answers

Option A is wrong because `dsl.Collected` is not a valid Kubeflow Pipelines SDK v2 component; it does not exist in the API. Option C is wrong because `dsl.Importer` (capital 'I') is not a valid class or function in the SDK v2; the correct name is all lowercase `dsl.importer`. Option D is wrong because `dsl.Artifact` is a base class for defining custom artifact types, not a component for importing artifacts into a pipeline.

33
MCQhard

An ML engineer is running a Vertex AI pipeline that includes a data validation component and a training component. The engineer wants the pipeline to stop before training if data validation fails, but wants the validation component to record its result as an output artifact for later inspection. Which combination of pipeline features should the engineer use?

A.Configure the validation component with a retry policy of 3 so transient validation errors are retried, then let training proceed.
B.Use a dsl.Condition to skip training when validation fails, and have the validation component always succeed while writing the report.
C.Have the validation component write the report artifact and return a success status, then check the artifact contents in the training component before training.
D.Have the validation component raise an exception on failure and write a validation report artifact before raising.
AnswerD

Raising an exception causes the pipeline task to fail, which stops downstream training tasks because they depend on the validation task. Writing the validation report artifact before raising ensures the artifact is persisted and available for inspection even though the task failed. This combination satisfies both the stop-before-training and record-for-inspection requirements.

Why this answer

In Vertex AI Pipelines, a task that raises an exception fails, and dependent tasks are not executed. Writing the validation report artifact before raising ensures the artifact is captured for later inspection. The other options either let training proceed, rely on retries that do not address deterministic data failures, or move the validation check into the training component, none of which stop the pipeline before training while preserving the report.

Exam trap

The trap here is assuming that a failed component cannot produce artifacts, when artifacts written before the exception are still persisted and available for inspection.

34
Multi-Selectmedium

An ML engineer is building a continuous training pipeline that retrains a model when new data arrives. The pipeline should also detect skew between training and serving data. Which TWO Google Cloud services should they use? (Choose two.)

Select 2 answers
A.Cloud Logging
B.Vertex AI Model Monitoring
C.Cloud Functions
D.Vertex AI Pipelines
E.Cloud Monitoring
AnswersB, D

Vertex AI Model Monitoring detects training-serving skew by comparing live prediction traffic against a baseline dataset, directly satisfying the pipeline's skew-detection requirement. It computes distribution drift metrics on features and predictions, alerting when serving data diverges from training data, so retraining triggers fire on genuine data drift rather than a fixed schedule.

Why this answer

Vertex AI Pipelines (D) is correct because it is the managed service for orchestrating and automating ML workflows, allowing the engineer to define a pipeline that triggers retraining when new data arrives. Vertex AI Model Monitoring (B) is correct because it is specifically designed to detect training-serving skew and drift by comparing the statistical distribution of incoming prediction requests against the training data baseline. Together, these two services directly satisfy both requirements: continuous training orchestration and skew detection.

Cloud Logging (A) only stores and queries log entries and does not orchestrate pipelines or compute skew. Cloud Functions (C) is a lightweight event-driven compute service that could trigger code but lacks native ML pipeline orchestration and skew detection. Cloud Monitoring (E) tracks operational metrics and alerts but does not perform training-serving skew analysis for ML models.

Exam trap

The Google PMLE exam often tests the distinction between monitoring for infrastructure health (Cloud Monitoring) versus monitoring for ML-specific data skew (Vertex AI Model Monitoring), leading candidates to confuse general observability with ML-specific drift detection.

35
MCQhard

A company runs a Vertex AI Pipeline that includes a custom component for hyperparameter tuning. The component uses a large search space and runs many trials. The pipeline is taking too long to complete, and the team wants to reduce the execution time without sacrificing model quality. They have already optimized the training code. Which Vertex AI Pipelines feature should they use to speed up the tuning component?

A.Enable pipeline caching for the tuning component so that repeated runs with identical parameters are skipped.
B.Use the ParallelFor loop in Vertex AI Pipelines to run multiple trials concurrently as separate tasks.
C.Reduce the number of trials by using a random search instead of a grid search.
D.Increase the machine type of the tuning component to a higher CPU or GPU configuration.
AnswerB

The ParallelFor loop allows you to execute multiple tasks in parallel, which is ideal for hyperparameter tuning where trials are independent. By running trials concurrently, the overall tuning time is reduced. This approach leverages Vertex AI Pipelines' orchestration to manage parallel execution and resource allocation, and it integrates with the pipeline's tracking and caching mechanisms.

Why this answer

ParallelFor in Vertex AI Pipelines enables concurrent execution of independent tasks, which is well-suited for hyperparameter tuning trials. By running multiple trials in parallel, the overall tuning time is significantly reduced while maintaining the same search space and model quality. This leverages the pipeline's orchestration capabilities and is more effective than caching or increasing machine type.

Exam trap

The trap here is assuming that caching or larger machines solve the runtime problem, when the real bottleneck is the sequential execution of independent trials.

36
MCQeasy

A machine learning engineer needs to pass a large dataset between two components in a Vertex AI pipeline. What is the recommended way to pass this data?

A.Store the dataset as a Dataset artifact and pass the artifact between components.
B.Write the dataset to a temporary BigQuery table and pass the table name.
C.Serialize the dataset to a string and pass it as a pipeline parameter.
D.Use a Cloud Storage bucket and pass the bucket name as a parameter.
AnswerA

Dataset artifacts are references to Cloud Storage or BigQuery URIs, so components exchange a lightweight metadata pointer rather than the bytes themselves. This satisfies the stem's large-dataset constraint, since passing raw data through component inputs would exceed the metadata limits.

Why this answer

In Vertex AI Pipelines, the recommended way to pass large datasets between components is to use a `Dataset` artifact. Artifacts are metadata references that point to the underlying data stored in Cloud Storage, enabling efficient, scalable, and type-safe data passing without serialization overhead or size limits. This approach leverages the Kubeflow Pipelines SDK's artifact tracking, which automatically handles lineage and versioning.

Exam trap

The trap here is that candidates often assume passing a Cloud Storage bucket name (Option D) is sufficient, but they miss that artifacts provide automatic metadata tracking, type safety, and integration with Vertex AI's lineage system, which is required for production ML pipelines.

How to eliminate wrong answers

Option B is wrong because writing a large dataset to a temporary BigQuery table introduces unnecessary latency, cost, and complexity; BigQuery is designed for analytical queries, not as an intermediate data transfer mechanism for pipeline components. Option C is wrong because serializing a large dataset to a string and passing it as a pipeline parameter violates the parameter size limit (typically 64KB in Kubeflow Pipelines) and would cause out-of-memory errors or pipeline failures. Option D is wrong because passing only the bucket name as a parameter lacks the structured metadata and type safety that artifacts provide; it forces components to independently resolve file paths and does not automatically track lineage or versioning.

37
MCQeasy

An organization wants to use Cloud Composer (Airflow) to orchestrate a machine learning workflow that includes running a Vertex AI Pipeline, followed by a BigQuery job, and then a Dataflow pipeline. What is the primary advantage of using Cloud Composer for this orchestration?

A.It allows orchestrating heterogeneous workflows across multiple GCP services with dependencies and retries.
B.It automatically caches the outputs of each step to avoid recomputation.
C.It integrates natively with the Vertex AI Model Registry for model versioning.
D.It provides a serverless execution environment for ML pipelines.
AnswerA

Cloud Composer runs Apache Airflow DAGs, whose operators and sensors coordinate tasks across disparate GCP services with explicit dependencies and retry policies. This satisfies the stem's need to sequence a Vertex AI Pipeline, BigQuery job and Dataflow pipeline in one workflow.

Why this answer

Cloud Composer (Apache Airflow) is designed to orchestrate heterogeneous workflows across multiple GCP services. In this scenario, it can define a Directed Acyclic Graph (DAG) that runs a Vertex AI Pipeline, then a BigQuery job, and finally a Dataflow pipeline, with built-in support for dependency management, retries, and failure handling. This is the primary advantage because it allows you to coordinate disparate services in a single, reliable workflow.

Exam trap

The trap here is that candidates may confuse Cloud Composer's orchestration capabilities with features specific to individual GCP services (like caching, model registry, or serverless execution), leading them to pick options that describe those services' features rather than the primary advantage of using an orchestrator.

How to eliminate wrong answers

Option B is wrong because Cloud Composer does not automatically cache outputs of each step; caching is a feature of specific services like Vertex AI Pipelines or Dataflow, not Airflow itself. Option C is wrong because Cloud Composer does not natively integrate with the Vertex AI Model Registry; that integration is handled by Vertex AI Pipelines or custom operators, not by Airflow's core orchestration. Option D is wrong because Cloud Composer is not serverless; it runs on a managed GKE cluster, and serverless ML pipeline execution is provided by Vertex AI Pipelines, not Cloud Composer.

38
MCQeasy

A data engineer wants to orchestrate a complex workflow that includes running a Vertex AI pipeline, then a BigQuery job, and finally a Dataflow pipeline. The workflow must handle dependencies, retries, and monitoring. Which Google Cloud service is most suitable for this orchestration?

A.Cloud Tasks
B.Cloud Composer
C.Cloud Scheduler
D.Workflows
AnswerB

Cloud Composer is managed Apache Airflow, whose DAGs natively orchestrate heterogeneous tasks across Vertex AI, BigQuery and Dataflow with dependency handling, retries and monitoring. This satisfies the stem's requirement for cross-service workflow orchestration with dependencies and retries.

Why this answer

Cloud Composer (based on Apache Airflow) is the most suitable service for orchestrating a complex workflow with dependencies, retries, and monitoring across Vertex AI, BigQuery, and Dataflow. It provides a managed Airflow environment that natively supports DAG-based orchestration, built-in retry logic, and integration with Google Cloud services via operators like VertexAIPipelineOperator, BigQueryOperator, and DataflowTemplatedJobStartOperator.

Exam trap

A common misconception is that Workflows is sufficient for complex ML orchestration, but it lacks the built-in operator integrations and retry semantics that Cloud Composer provides for multi-service pipelines.

How to eliminate wrong answers

Option A is wrong because Cloud Tasks is a distributed task queue for executing discrete, short-lived tasks with HTTP endpoints, not for orchestrating multi-step workflows with complex dependencies and retries across different services. Option C is wrong because Cloud Scheduler is a cron-based job scheduler that triggers single events at specified times, lacking the ability to manage dependencies between multiple pipeline stages or handle retries. Option D is wrong because Workflows is a low-code orchestration service for sequential or parallel steps, but it does not natively support the rich operator ecosystem, retry policies, or monitoring capabilities that Cloud Composer provides for ML pipelines involving Vertex AI, BigQuery, and Dataflow.

39
MCQmedium

A team uses Cloud Build to automatically trigger a Vertex AI pipeline when changes are pushed to the model code repository. They have a cloudbuild.yaml file that builds a container image and submits the pipeline. However, they want to run the pipeline only if the commit includes changes to the 'training/' directory. Which Cloud Build configuration option should be used to filter the trigger?

A.Add a 'ignoreFiles' field with 'training/**' to the trigger.
B.Use a 'substitutions' field with a regex pattern to filter commits.
C.Configure a Cloud Function to check the commit diff and call Cloud Build API conditionally.
D.Set the 'includedFiles' field to 'training/**' in the trigger configuration.
AnswerD

The includedFiles field with 'training/**' restricts the trigger to commits touching that directory, directly satisfying the requirement to run the pipeline only for training-code changes. Cloud Build evaluates this glob against changed paths before executing cloudbuild.yaml, avoiding unnecessary builds.

Why this answer

Cloud Build triggers support an `includedFiles` field that specifies a glob pattern. When set to `training/**`, the trigger will only fire if the commit includes changes to files under the `training/` directory. This is the native, declarative way to filter triggers based on changed file paths without additional infrastructure.

Exam trap

The trap here is that candidates confuse `ignoreFiles` with `includedFiles`, or assume that a custom solution like Cloud Functions is required when Cloud Build already provides a native, simpler mechanism for path-based filtering.

How to eliminate wrong answers

Option A is wrong because `ignoreFiles` excludes commits that match the pattern, but the requirement is to run the pipeline only when changes occur in `training/`, not to ignore them. Option B is wrong because `substitutions` are used for variable replacement in build configuration, not for filtering trigger conditions based on file changes. Option C is wrong because while a Cloud Function could achieve this, it introduces unnecessary complexity and cost; Cloud Build triggers natively support file path filtering via `includedFiles`, making a separate function an anti-pattern.

40
MCQmedium

A company wants to implement continuous delivery (CD) for ML models, where a model is automatically deployed to a staging environment and only promoted to production after passing an evaluation gate. Which combination of GCP services is BEST suited for orchestrating this CD pipeline?

A.Cloud Scheduler and Pub/Sub
B.Cloud Composer (Airflow) with Cloud Functions
C.Cloud Build with Vertex AI Pipelines and Cloud Deploy
D.Vertex AI Pipelines with Cloud Run
AnswerC

Cloud Build handles CI image and pipeline builds, Vertex AI Pipelines runs the training and evaluation gate, and Cloud Deploy manages progressive promotion to staging then production. Together they satisfy the stem's requirement that promotion occur only after the evaluation gate passes.

Why this answer

Cloud Build can trigger on code/model changes and run a pipeline that deploys to staging. After evaluation, if successful, it can promote to production using Cloud Deploy or directly update Vertex AI endpoints. Cloud Composer (Airflow) is also a good option for complex orchestration, but for CI/CD, Cloud Build is a natural fit.

The combination of Cloud Build, Cloud Deploy, and Vertex AI provides a robust CD pipeline.

41
MCQmedium

A team runs a Vertex AI pipeline that includes a component which downloads a large dataset from BigQuery and writes it to Cloud Storage. The pipeline's caching is enabled by default. During iterative development, the engineer modifies the SQL query inside the component to include an additional feature column, but the pipeline still uses the previously cached output because the component's input parameters and code hash are unchanged. The engineer needs the component to re-execute with the updated query without disabling caching for the entire pipeline. What should the engineer do?

A.Add the SQL query as a pipeline parameter and pass it to the component as an input.
B.Set the component's `enable_caching` argument to False in the pipeline definition.
C.Modify the component's container image tag to a new version and rebuild the pipeline.
D.Clear the pipeline's cache by deleting the pipeline run's metadata from Vertex ML Metadata.
AnswerA

Vertex AI Pipelines computes the cache key from the component's inputs (including code and parameters). By promoting the SQL query from hardcoded code to an input parameter, any change to the query alters the input, invalidating the cache and triggering re-execution. This maintains caching benefits for unchanged queries and aligns with pipeline best practices of parameterizing variable logic. It directly solves the problem without globally disabling caching.

Why this answer

Vertex AI Pipelines caching uses a fingerprint derived from the component's code and input parameters. When a component's internal logic changes but its inputs and code hash remain the same, the cache is considered valid. To force re-execution without disabling caching entirely, the engineer should expose the changing element—the SQL query—as an input parameter.

This changes the cache key only when the query changes, preserving caching for other runs. Rebuilding the image or disabling caching are less precise and more disruptive.

Exam trap

The trap here is assuming that modifying the component's source code automatically changes the cache key when the component is not rebuilt or its inputs are not changed.

42
MCQmedium

An organization wants to trigger a Vertex AI pipeline whenever new data arrives in a Cloud Storage bucket. Which approach should they use?

A.Configure a Vertex AI pipeline trigger directly on the bucket using the GCP Console.
B.Use Pub/Sub notifications from the bucket and a Dataflow job to start the pipeline.
C.Set up a Cloud Scheduler job that runs every minute and checks for new files in the bucket.
D.Use Cloud Functions triggered by Cloud Storage events to call the Vertex AI pipeline API.
AnswerD

A Cloud Storage-triggered Cloud Function receives the object-finalise event and calls the Vertex AI Pipelines API to launch a run, giving event-driven execution without polling. This satisfies the stem's requirement to trigger the pipeline whenever new data lands in the bucket.

Why this answer

Cloud Functions can be directly triggered by Cloud Storage events (e.g., `google.storage.object.finalize`) and can then call the Vertex AI pipeline API using the Cloud SDK or client libraries. This provides a serverless, event-driven architecture that reacts immediately to new data without polling or additional infrastructure.

Exam trap

The trap here is that candidates may assume Vertex AI pipelines have a built-in Cloud Storage trigger (Option A) or over-engineer the solution with Dataflow (Option B), when the simplest and most native serverless approach is Cloud Functions.

How to eliminate wrong answers

Option A is wrong because Vertex AI pipelines do not support configuring a trigger directly on a Cloud Storage bucket via the GCP Console; there is no native bucket-to-pipeline trigger. Option B is wrong because while Pub/Sub notifications from the bucket are possible, adding a Dataflow job introduces unnecessary complexity and latency—Dataflow is a batch/stream processing engine, not a lightweight event router. Option C is wrong because a Cloud Scheduler job that runs every minute and checks for new files is inefficient (polling), introduces up to 60 seconds of latency, and does not scale well; it also requires custom code to track file states.

43
Multi-Selectmedium

An ML pipeline must run a set of preprocessing tasks for each data shard in parallel. Which KFP SDK features should they use to implement this? (Choose two.)

Select 2 answers
A.dsl.ParallelFor
B.dsl.PipelineParam
C.dsl.Collected
D.dsl.Condition
E.dsl.ExitHandler
AnswersA, C

dsl.ParallelFor iterates over the shard list at pipeline-compile time, generating one preprocessing task instance per shard that KFP schedules concurrently. This directly satisfies the requirement to run preprocessing tasks for each data shard in parallel.

Why this answer

dsl.ParallelFor (A) is correct because it is the KFP SDK construct that iterates over a list (such as the set of data shards) and fans out the loop body into parallel task executions, which is exactly what is needed to run preprocessing per shard concurrently. dsl.Collected (C) is correct because it is used with dsl.ParallelFor to gather the outputs of all parallel iterations into a single list, allowing downstream pipeline steps to consume the aggregated results of the per-shard preprocessing. dsl.PipelineParam (B) is not the mechanism for parallel iteration; it only represents a runtime parameter passed into a pipeline or component. dsl.Condition (D) implements conditional branching (if/else) rather than parallel fan-out, and dsl.ExitHandler (E) defines cleanup logic that runs when a scope exits, neither of which provides the required parallel execution over shards.

Exam trap

In the Google PMLE exam, note that dsl.ParallelFor is used for parallel iteration over data shards, and dsl.Collected gathers outputs from all iterations. Candidates often confuse these with dsl.Condition (for branching) or dsl.PipelineParam (for parameters), so carefully read whether the question asks for parallel processing or conditional logic.

44
MCQmedium

An ML engineer is designing a pipeline that should run only when new training data arrives in a Cloud Storage bucket. Which event-driven approach should they use to trigger the Vertex AI Pipeline?

A.Use Cloud Storage Pub/Sub notifications to send events to a Cloud Function that triggers the pipeline.
B.Use Cloud Tasks to queue a pipeline run whenever a new file is uploaded.
C.Configure the pipeline to run on a schedule and check for new data inside the pipeline.
D.Set up a Cloud Scheduler job that runs every minute to check for new files.
AnswerA

Cloud Storage Pub/Sub notifications emit an event whenever an object lands in the bucket, and a Cloud Function subscribed to that topic invokes the Vertex AI Pipeline. This satisfies the stem's requirement to run only when new training data actually arrives, avoiding polling.

Why this answer

The best approach is to use Cloud Storage notifications via Pub/Sub, then a Cloud Function that receives the event and calls the Vertex AI API to create a pipeline job. This is a common event-driven pattern. Cloud Scheduler is for scheduled triggers, not event-driven.

Cloud Tasks and Cloud Run are not typically used for this purpose.

45
MCQmedium

An ML engineer wants to containerize a custom training script and use it as a component in a Vertex AI Pipeline. The component should accept a dataset URI and a learning rate parameter, and output a trained model artifact. Which approach should the engineer use to define the component?

A.Use a pre-built Google Cloud Pipeline Component for Vertex AI Training with custom container configuration.
B.Use ContainerComponent from kfp.v2.components to define the container, its inputs, and outputs.
C.Define a Python function component with @dsl.component and include the container code inline.
D.Use the importer component to import the script and then run it as a task.
AnswerB

ContainerComponent from kfp.v2.components lets the engineer declare a custom container image plus typed inputs (dataset URI, learning rate) and outputs (model artifact), satisfying the stem's requirement to containerise a custom training script as a pipeline component.

Why this answer

ContainerComponent from kfp.v2.components allows you to define a custom container component by specifying the container image, command, inputs, and outputs directly. This is the appropriate approach when you have a custom training script that you want to containerize and use as a component in a Vertex AI Pipeline, as it gives you full control over the container configuration and artifact handling.

Exam trap

Candidates often confuse Python function components (@dsl.component) with container components. The trap here is that they may think a Python function component can containerize a custom script, but it cannot directly specify a container image and artifact outputs like ContainerComponent does.

How to eliminate wrong answers

Option A is wrong because pre-built Google Cloud Pipeline Components for Vertex AI Training are designed for standard training jobs with built-in algorithms or custom containers, but they do not allow you to define custom inputs and outputs as artifacts in the same declarative way as ContainerComponent; they are more rigid and less suited for a fully custom component with a dataset URI and learning rate parameter. Option C is wrong because @dsl.component is used for Python function components that run Python code directly, not for containerized components; including container code inline would mix the container definition with Python function logic, which is not the intended use and would not properly handle container image specification and artifact outputs. Option D is wrong because the importer component is used to import existing artifacts (like models or datasets) into the pipeline, not to run a training script; it cannot execute a custom training script or produce a trained model artifact from scratch.

46
Multi-Selecthard

A company is deploying a Vertex AI pipeline that trains a model and then runs a custom evaluation component. The evaluation component must only run if the training component succeeds and the model's accuracy exceeds a threshold. The pipeline must also support retries for transient errors in the training component. The engineer needs to configure the pipeline to meet these requirements. Which two actions should the engineer take? (Choose two.)

Select 2 answers
A.Configure the pipeline to run the evaluation component in parallel with training to reduce latency.
B.Set `enable_caching=False` on the evaluation component to ensure it always runs after training.
C.Set the `retry` policy on the training component to retry on specific exit codes or exceptions.
D.Use a `dsl.Condition` to wrap the evaluation component and check the accuracy metric against the threshold.
E.Use a `dsl.ExitHandler` to catch failures in the training component and retry it manually.
AnswersC, D

Vertex AI Pipelines supports retry policies on individual components. By configuring a retry policy with a maximum retry count and backoff, the training component can automatically retry on transient failures such as resource exhaustion or network timeouts. This satisfies the requirement to support retries for transient errors. The retry policy can be specified using the `retry` argument when defining the component or task, and it applies only to that task.

Why this answer

The evaluation component must run only when training succeeds and accuracy exceeds a threshold, so a dsl.Condition is used to gate its execution based on the accuracy metric. To handle transient errors in training, a retry policy on the training component is configured. These two actions together meet the requirements.

Other options either do not enforce the condition, do not provide retries, or are not applicable.

Exam trap

The trap here is confusing caching with conditional execution, or assuming that an ExitHandler can serve as a retry mechanism.

47
MCQhard

An ML engineer is authoring a Vertex AI pipeline where a custom training component must read a dataset from a BigQuery table and write the trained model to a Cloud Storage bucket. The engineer wants the component to be reusable across projects and environments without hardcoding project IDs or bucket names. Which design should the engineer use?

A.Use the Vertex AI SDK's aiplatform.init() with no arguments inside the component and rely on the default project and bucket.
B.Read the project ID and bucket name from environment variables set inside the component's container image at build time.
C.Hardcode the production project ID and bucket name in the component, then override them with a pipeline-level parameter only when running in non-production.
D.Pass the BigQuery table URI and the Cloud Storage output URI as component input parameters, and let the pipeline caller supply them at runtime.
AnswerD

Making the dataset URI and output URI component inputs parameterizes the component so the same definition can be reused across projects and environments. The pipeline caller supplies the concrete values at runtime, which is the standard Vertex AI Pipelines pattern for portability. Hardcoding or deriving them inside the component would tie the component to one project and defeat reusability.

Why this answer

Component reusability in Vertex AI Pipelines comes from parameterizing all environment-specific values as inputs. Passing the BigQuery table URI and Cloud Storage output URI as component inputs lets the same component definition run in any project or environment, with the pipeline caller providing concrete values. The other options embed environment-specific values in the container or rely on implicit defaults, which breaks portability and can cause cross-environment mistakes.

Exam trap

The trap here is thinking that relying on default credentials and default project inside a component is equivalent to passing explicit inputs, when implicit defaults make the component environment-dependent.

48
MCQmedium

An ML engineer is building a Vertex AI Pipeline that includes a data validation component. The component should fail the pipeline if the input data does not meet certain statistical thresholds. The engineer wants to ensure that the pipeline stops immediately and does not proceed to training if validation fails. Which mechanism should the engineer use in the component?

A.Raise an exception in the component code and catch it in the pipeline definition.
B.Return a non-zero exit code from the component's container.
C.Write a warning to the logs and continue execution.
D.Use a conditional branch in the pipeline to skip training if validation fails.
AnswerB

In Vertex AI Pipelines, a component failure is indicated by a non-zero exit code from the container. This causes the pipeline task to fail, and by default, the pipeline stops executing subsequent tasks. This is the standard way to enforce validation gates and prevent downstream tasks from running with invalid data.

Why this answer

A non-zero exit code from a component's container is the standard way to signal failure in Vertex AI Pipelines. When a component fails, the pipeline task fails, and the pipeline stops by default, preventing downstream tasks from executing. This enforces the validation gate effectively.

Exam trap

The trap here is thinking that logging a warning or using a conditional branch is sufficient, but only a non-zero exit code causes the pipeline to fail and stop immediately.

49
Multi-Selectmedium

A company is using Vertex AI Pipelines to orchestrate a training workflow. They want to implement a CI/CD process where the pipeline is automatically triggered when a new version of the training code is pushed to a GitHub repository. They also want to ensure that the pipeline uses the latest code. Which two actions should they take? (Choose two.)

Select 2 answers
A.Use a Cloud Function to monitor the GitHub repository and trigger the pipeline directly when a push occurs.
B.Configure the Vertex AI Pipeline to pull the training code from GitHub at runtime using a git clone step.
C.Create a Cloud Build trigger that builds a new container image from the GitHub repository and pushes it to Artifact Registry.
D.Set up a Cloud Build trigger that submits the Vertex AI Pipeline job, using the newly built image as a parameter.
E.Store the training code in a Cloud Storage bucket and have the pipeline download it at the start of each run.
AnswersC, D

A Cloud Build trigger can be configured to watch the GitHub repository and automatically build a container image whenever code is pushed. This image contains the latest training code and is pushed to Artifact Registry. The pipeline can then reference this image, ensuring the latest code is used. This is a standard CI step for ML pipelines.

Why this answer

The two correct actions are to create a Cloud Build trigger that builds and pushes a container image from the GitHub repository, and to set up a Cloud Build trigger that submits the Vertex AI Pipeline job using that image. This creates a CI/CD pipeline where code changes automatically trigger a build and then a pipeline run with the latest code, ensuring reproducibility and automation.

Exam trap

The trap here is thinking that the pipeline can directly pull code from GitHub or that a Cloud Function is needed, when Cloud Build triggers are the native integration for CI/CD with Vertex AI Pipelines.

50
MCQmedium

A team wants to implement continuous training for their ML model. The pipeline should be triggered when new training data arrives in a Cloud Storage bucket. Which combination of services should they use?

A.Cloud Storage → BigQuery → Vertex AI Pipelines
B.Cloud Storage → Cloud Functions → Vertex AI Pipelines
C.Cloud Storage → Cloud Build → Vertex AI Pipelines
D.Cloud Storage → Cloud Scheduler → Vertex AI Pipelines
AnswerB

A Cloud Storage object-change event triggers a Cloud Functions function, which calls the Vertex AI API to launch the training pipeline. This event-driven chain satisfies the stem's requirement to retrain automatically when new training data lands in the bucket.

Why this answer

The correct combination is Cloud Storage → Cloud Functions → Vertex AI Pipelines. Cloud Storage hosts the new training data; a Cloud Function triggered by an object-finalize event (or Pub/Sub notification) detects the new data and invokes a Vertex AI Pipeline to retrain the model. This creates a fully event-driven continuous training pipeline without polling or manual intervention.

Exam trap

PMLE often tests the difference between event-driven triggers (Cloud Functions) and time-based triggers (Cloud Scheduler), catching candidates who pick Scheduler for data-arrival events.

How to eliminate wrong answers

Option A is wrong because BigQuery is a data warehouse, not an event trigger—it does not natively detect new Cloud Storage objects and trigger pipelines. Option C is wrong because Cloud Build is a CI/CD service for building and deploying code/containers, not for reacting to data arrival events. Option D is wrong because Cloud Scheduler is time-based (cron), not event-driven; it cannot trigger on new data arrival, only on a schedule.

51
MCQhard

An ML engineer is designing a Vertex AI Pipeline that includes a hyperparameter tuning step. The tuning step runs multiple trials and outputs the best model. The engineer wants to ensure that the pipeline can resume from the tuning step if it fails later, without re-running all the tuning trials. What should the engineer do?

A.Configure the tuning step to use a Vertex AI HyperparameterTuningJob with a persistent resource pool.
B.Store the tuning results in a Cloud Storage bucket and add a conditional step that checks for existing results before running tuning.
C.Enable caching for the tuning step by setting the component's 'enable_caching' attribute to True.
D.Set the pipeline's 'resume' option to True when submitting the pipeline.
AnswerC

Vertex AI Pipelines supports caching of component executions. When caching is enabled, the pipeline stores the outputs of a component based on its inputs and code. If the pipeline is re-run and the tuning step's inputs have not changed, the cached outputs (including the best model) are reused, avoiding re-execution of all trials. This is the correct way to resume without re-running tuning.

Why this answer

The correct method is to enable caching for the tuning component. Vertex AI Pipelines caching stores the outputs of a component based on a hash of its inputs, code, and environment. When the pipeline is re-run, if the tuning step's inputs are unchanged, the cached outputs are used, and the tuning trials are not re-executed.

This saves time and cost, and allows the pipeline to resume from the tuning step.

Exam trap

The trap here is assuming that Vertex AI Pipelines has a built-in 'resume' feature or that manual result storage is needed, when caching is the intended mechanism.

52
MCQeasy

An ML engineer is creating a Vertex AI Pipeline that includes a component to train a model. The component requires a machine type with a GPU. The engineer wants to specify the machine type and GPU type for the component. Which approach should the engineer use?

A.Configure the GPU in the pipeline's runtime configuration, which applies to all components.
B.Use a pre-built training container from Vertex AI that automatically selects the appropriate GPU.
C.Set the machine type and accelerator in the component's container specification using the Vertex AI SDK.
D.Pass the machine type and GPU type as command-line arguments to the component's container.
AnswerC

In Vertex AI Pipelines, you can specify the machine type and accelerator (GPU) in the component's container specification when defining the component. This is done using the Vertex AI SDK's component decorator or by setting the resource requirements in the pipeline task. This ensures that the component runs on the desired hardware.

Why this answer

To specify a machine type and GPU for a Vertex AI Pipelines component, you set the resource requirements in the component's container specification. This ensures the component runs on the desired hardware. Other approaches like command-line arguments or global configuration do not correctly allocate the resources for the specific component.

Exam trap

The trap here is thinking that machine type and GPU can be passed as runtime arguments or set globally, but they must be declared in the component's resource specification.

53
MCQhard

A company has a CI/CD pipeline that retrains a model every time new training data is available. They want to automatically deploy the new model to production only if it passes a set of evaluation tests on a staging environment. Which approach best implements this?

A.Implement a two-stage pipeline: train and deploy to staging, run evaluation tests, and if passed, deploy to production using conditional logic.
B.Use Cloud Build to trigger a training job and then a separate deployment job without evaluation.
C.Use a single Vertex AI pipeline that trains and deploys to staging, then manually promote.
D.Train and deploy directly to production in one pipeline.
AnswerA

A two-stage pipeline trains and deploys to staging, runs evaluation tests, then uses conditional logic to promote to production only on passing results. This gates deployment on measured quality, satisfying the requirement to release solely when evaluation succeeds.

Why this answer

The correct approach is a two-stage pipeline: train and deploy to a staging environment, run automated evaluation tests against the staged model, and use conditional logic to promote to production only if the tests pass. This implements a proper ML CI/CD gate that prevents regressions from reaching production.

Exam trap

PMLE often tests whether candidates confuse 'automated deployment' with 'automated promotion' — the trap is choosing an option that deploys automatically but omits the evaluation gate, or one that evaluates but requires manual promotion.

How to eliminate wrong answers

Option B is wrong because a deployment job without evaluation skips the quality gate entirely, allowing untested models into production. Option C is wrong because manual promotion does not satisfy the requirement for automatic deployment based on evaluation results — it introduces a human step that the question explicitly wants to avoid. Option D is wrong because training and deploying directly to production in one pipeline bypasses staging and evaluation, which is exactly the risk the company wants to mitigate.

54
MCQhard

A machine learning engineer is building a Vertex AI pipeline that uses a pre-built AutoML Tables component to train a classification model. The pipeline also includes a conditional step that deploys the model to an endpoint only if the evaluation metrics exceed a threshold. Which KFP feature should be used to implement the conditional deployment?

A.dsl.ParallelFor
B.dsl.ExitHandler
C.dsl.Condition
D.dsl.Collected
AnswerC

dsl.Condition wraps pipeline steps in a conditional branch whose predicate is evaluated at runtime, so the deployment step executes only when the evaluation metrics exceed the threshold. This directly implements the stem's conditional deployment requirement within the KFP pipeline definition.

Why this answer

The `dsl.Condition` feature from KFP (Kubeflow Pipelines) is specifically designed to conditionally execute pipeline steps based on the output of a previous component. In this scenario, the AutoML Tables component produces evaluation metrics; `dsl.Condition` allows the pipeline to check whether those metrics exceed a threshold and, if true, run the deployment step. This is the correct, native KFP construct for implementing branching logic within a pipeline.

Exam trap

The trap here is that candidates often confuse `dsl.Condition` with `dsl.ExitHandler` because both involve decision-making, but `ExitHandler` is only for post-exit cleanup, not for branching based on step outputs.

How to eliminate wrong answers

Option A is wrong because `dsl.ParallelFor` is used for iterating over a collection of items and executing steps in parallel, not for conditional branching based on a single metric threshold. Option B is wrong because `dsl.ExitHandler` is a mechanism to run a cleanup or notification step when a pipeline exits (successfully or with failure), not for conditionally deploying a model based on evaluation results. Option D is wrong because `dsl.Collected` is a function used to gather outputs from parallel iterations (e.g., from `dsl.ParallelFor`) into a single list, not a control flow construct for conditional execution.

55
MCQmedium

A team has a Vertex AI pipeline that includes a container component for data preprocessing. The team notices that the component is re-executed every time the pipeline runs, even when the inputs and code haven't changed. They want to leverage pipeline caching to avoid redundant executions. What should they do to enable caching for this component?

A.Set the 'caching' flag to 'True' in the pipeline definition using 'pipeline.caching = True'.
B.Set the environment variable 'ENABLE_CACHE' to 'true' on the pipeline run request.
C.Re-compile the pipeline with the '--enable-cache' flag.
D.Ensure that the component does not have 'dsl.cache_options(enable_cache=False)' set.
AnswerD

Caching is enabled by default in Kubeflow Pipelines; the component re-executes because cache_options(enable_cache=False) was explicitly set, disabling it. Removing that setting restores default caching behaviour, so unchanged inputs and code reuse the cached execution instead of rerunning.

Why this answer

Vertex AI pipeline caching is enabled by default for all components unless explicitly disabled using `dsl.cache_options(enable_cache=False)`. The component re-executing every time indicates that caching was likely disabled on that specific component. Removing or ensuring this setting is not present will allow the pipeline to reuse cached outputs when inputs and code have not changed.

Exam trap

The trap here is that candidates assume caching must be explicitly enabled (like in some other cloud platforms), but Vertex AI caches by default, so the issue is usually that caching was explicitly disabled on the component.

How to eliminate wrong answers

Option A is wrong because Vertex AI pipelines do not have a global `pipeline.caching` attribute; caching is controlled per component via the `@component` decorator or `ContainerComponent` definition, not at the pipeline level. Option B is wrong because there is no environment variable `ENABLE_CACHE` for Vertex AI pipeline runs; caching is configured in the pipeline definition, not via runtime environment variables. Option C is wrong because Vertex AI pipelines do not use a `--enable-cache` compilation flag; caching is a runtime feature controlled by component-level settings, not a compile-time option.

56
MCQmedium

An engineer needs to compile a Kubeflow Pipeline defined in Python to a JSON format that can be run on Vertex AI Pipelines. Which command should they use?

A.kfp.compiler.Compiler().compile(pipeline_func, 'pipeline.json')
B.gcloud ai pipelines compile command.
C.kfp.Client().upload_pipeline()
D.dsl.pipeline decorator automatically compiles at runtime.
AnswerA

The KFP compiler's compile method converts a Python pipeline function into an IR YAML or JSON specification that Vertex AI Pipelines can submit and execute. Passing the pipeline function and output filename produces the required JSON artefact directly.

Why this answer

The Kubeflow Pipelines SDK provides the `kfp.compiler.Compiler().compile()` method to convert a Python-based pipeline function into a JSON or YAML format that is compatible with Vertex AI Pipelines. This JSON representation defines the pipeline's components, dependencies, and execution graph, enabling it to be submitted to Vertex AI for orchestration. The `compile()` method is the standard way to produce a portable pipeline specification from Python code.

Exam trap

The trap here is that candidates may mistakenly believe that the `gcloud ai pipelines` command or the `dsl.pipeline` decorator directly compiles the pipeline, but in Google's Vertex AI Pipelines, you must use the KFP SDK's `Compiler().compile()` method to generate the pipeline JSON specification.

How to eliminate wrong answers

Option B is wrong because `gcloud ai pipelines compile` is not a valid gcloud command; the gcloud CLI for Vertex AI uses `gcloud ai pipelines run` to submit a pre-compiled pipeline, but compilation must be done separately using the KFP SDK. Option C is wrong because `kfp.Client().upload_pipeline()` uploads a compiled pipeline package to a Kubeflow Pipelines instance, but it does not perform the compilation step itself; the pipeline must already be compiled into a JSON or YAML file before uploading. Option D is wrong because the `dsl.pipeline` decorator defines the pipeline structure and components but does not automatically compile it at runtime; explicit invocation of `Compiler().compile()` is required to generate the JSON artifact.

57
MCQmedium

A data science team wants to build a machine learning pipeline on Vertex AI Pipelines that preprocesses data, trains a model, and evaluates it. They need to ensure that components can be reused across multiple pipelines and that outputs from one component can be passed as inputs to another. Which approach should they take?

A.Write each component as a Cloud Composer DAG task using Python operators and manage dependencies via Airflow.
B.Use Vertex AI pre-built components exclusively and chain them using the Vertex AI SDK without a pipeline definition.
C.Define each step as a separate Cloud Build step and chain them via build triggers.
D.Use Kubeflow Pipelines SDK v2 to create Python function components decorated with @dsl.component and compose them into a pipeline using @dsl.pipeline.
AnswerD

Decorating Python functions with @dsl.component produces self-contained, reusable components whose typed inputs and outputs Vertex AI Pipelines resolves automatically, so one component's output artefact can be wired directly into the next component's input. The @dsl.pipeline decorator then composes these into a directed acyclic graph, satisfying the reuse and data-passing constraints.

Why this answer

Kubeflow Pipelines SDK v2 with @dsl.component and @dsl.pipeline decorators is the native way to define reusable, composable components in Vertex AI Pipelines. This approach allows each component to be a self-contained Python function that can be independently versioned and reused across multiple pipelines, with outputs automatically serialized and passed as inputs to downstream components via the pipeline graph.

Exam trap

Google PMLE often tests the misconception that any orchestration tool (Airflow, Cloud Build) can substitute for a purpose-built ML pipeline framework, but the key differentiator is Vertex AI Pipelines' native support for reusable components with typed artifact passing and managed execution.

How to eliminate wrong answers

Option A is wrong because Cloud Composer (Airflow) is a workflow orchestrator for general DAGs, not a purpose-built ML pipeline framework; it lacks native support for Vertex AI Pipelines' artifact tracking, component reuse, and ML-specific I/O handling. Option B is wrong because Vertex AI pre-built components cannot be chained without a pipeline definition; the Vertex AI SDK requires a pipeline specification (e.g., via Kubeflow Pipelines) to define the execution graph and pass outputs between steps. Option C is wrong because Cloud Build is a CI/CD service for building and testing code, not for orchestrating ML pipelines; it does not provide managed artifact passing, caching, or the runtime environment needed for ML training and evaluation steps.

58
MCQmedium

A machine learning engineer needs to create a pipeline that runs a custom container component on Vertex AI. The container expects a Cloud Storage path as input and outputs a model artifact. Which component type should they define using the Kubeflow Pipelines SDK v2?

A.Google Cloud Pipeline Components (GCPC) for custom containers
B.Python function component using @dsl.component
C.Importer component to load the container as an artifact
D.Container component using @dsl.container_component
AnswerD

The @dsl.container_component decorator lets you define a component from a custom container image, with typed inputs and outputs declared in the function signature. This satisfies the stem's need to pass a Cloud Storage path in and emit a model artifact out.

Why this answer

The Kubeflow Pipelines SDK v2 provides the @dsl.container_component decorator specifically for defining components that wrap custom container images. This allows the engineer to specify the container image, input/output paths (like a Cloud Storage path), and artifact metadata, enabling Vertex AI to execute the container as a pipeline step and capture the model artifact.

Exam trap

The trap here is that candidates confuse the Importer component (which only imports existing artifacts) with a component that runs a container to produce an artifact, or they mistakenly think Google Cloud Pipeline Components can wrap any custom container when they only provide pre-built Google service integrations.

How to eliminate wrong answers

Option A is wrong because Google Cloud Pipeline Components (GCPC) are pre-built components for Google Cloud services (e.g., AI Platform, BigQuery), not for wrapping arbitrary custom containers. Option B is wrong because @dsl.component is used for Python function components that execute inline Python code, not for running a custom container image. Option C is wrong because the Importer component is used to import existing artifacts (like a pre-trained model) into the pipeline's metadata store, not to run a container that produces an artifact.

59
MCQhard

A company runs a Vertex AI Pipeline that trains a model and then deploys it to a Vertex AI Endpoint. The pipeline uses a conditional deployment step based on the model's evaluation metric. The team wants to ensure that if the evaluation metric falls below a threshold, the pipeline fails and no deployment occurs. Which approach should they use?

A.Add a custom component that evaluates the metric and, if below threshold, calls the Vertex AI API to cancel the pipeline.
B.Use a Condition to check the metric, and in the else branch, include a component that raises an exception or returns a failure status.
C.Configure the pipeline to use a fail-fast strategy by setting the pipeline's failure_policy to FAIL_FAST.
D.Use a Condition in the pipeline that checks if the metric is greater than the threshold, and only then execute the deployment component.
AnswerB

This approach ensures that if the metric is below threshold, the else branch executes a component that fails the pipeline. If the metric is above threshold, the deployment component runs. This provides both conditional deployment and a clear failure signal when the metric does not meet requirements.

Why this answer

To conditionally deploy and fail the pipeline when the metric is unsatisfactory, the pipeline should use a Condition with an else branch that triggers a failing component. This ensures that deployment only happens when the metric passes, and the pipeline fails otherwise. Other methods either do not cause failure or rely on unsupported features.

Exam trap

The trap here is believing that a Condition alone can fail the pipeline or that a built-in fail-fast policy exists, when in fact you must explicitly cause a failure in the else branch.

60
MCQmedium

A data science team wants to build a Vertex AI pipeline that trains a model, evaluates it, and conditionally deploys it if the accuracy exceeds 0.9. They want to use the Kubeflow Pipelines SDK v2. Which construct allows them to conditionally execute the deployment step based on the evaluation metric?

A.dsl.If
B.dsl.Conditional
C.dsl.ExitHandler
D.dsl.Collected
AnswerA

dsl.If is the KFP SDK v2 conditional construct that evaluates a pipeline parameter or task output at runtime and executes its enclosed tasks only when the boolean condition holds. It satisfies the stem's requirement to deploy only when evaluation accuracy exceeds 0.9.

Why this answer

In Kubeflow Pipelines SDK v2, `dsl.If` is the correct construct for conditionally executing pipeline steps based on runtime metrics or parameters. It allows you to define a condition that, when evaluated to true, triggers the deployment step only if the model accuracy exceeds 0.9. This is the standard way to implement branching logic in v2 pipelines.

Exam trap

Candidates often confuse `dsl.If` with `dsl.Conditional`, but in Kubeflow Pipelines SDK v2 for Vertex AI, `dsl.If` is the correct construct for conditional execution based on runtime metrics.

How to eliminate wrong answers

Option B is wrong because `dsl.Conditional` is not a valid construct in Kubeflow Pipelines SDK v2; the correct name is `dsl.If`. Option C is wrong because `dsl.ExitHandler` is used to execute cleanup or notification steps when a pipeline or component exits, not for conditional branching based on evaluation metrics. Option D is wrong because `dsl.Collected` is used to gather outputs from parallel tasks (e.g., from a loop or fan-out), not to conditionally execute a step.

61
MCQeasy

What is the primary benefit of using pipeline caching in Vertex AI Pipelines?

A.It reduces execution time and cost by reusing unchanged component outputs.
B.It encrypts data at rest.
C.It automatically scales the pipeline resources.
D.It enables parallel execution of components.
AnswerA

Pipeline caching stores each component's outputs keyed by its inputs and code version, so unchanged steps skip re-execution entirely. This directly satisfies the stem's constraint of reducing execution time and cost, since Vertex AI Pipelines avoids recomputing identical artefacts and only reruns components whose inputs or definitions changed.

Why this answer

Pipeline caching in Vertex AI Pipelines automatically detects when a component's inputs and code have not changed from a previous execution and reuses the cached output artifacts. This avoids redundant computation, directly reducing both execution time and cost by skipping re-execution of unchanged steps.

Exam trap

The exam often tests the distinction between caching (reusing outputs) and parallelization (running components concurrently), so candidates may confuse the two and incorrectly select parallel execution as the primary benefit.

How to eliminate wrong answers

Option B is wrong because encryption at rest is a data security feature managed by Cloud KMS or default Google Cloud encryption, not a benefit of pipeline caching. Option C is wrong because automatic scaling of pipeline resources is handled by Vertex AI's underlying infrastructure (e.g., node auto-scaling) or custom configuration, not by caching. Option D is wrong because parallel execution of components is achieved through pipeline design (e.g., using `dsl.ParallelFor` or independent component dependencies), not through caching; caching can actually reduce the need for parallel execution by reusing results.

62
MCQeasy

An ML engineer is using Vertex AI Pipelines to orchestrate a training workflow. The pipeline must run on a schedule every day at 2:00 AM UTC. The engineer wants to use a fully managed Google Cloud service to trigger the pipeline. Which service should the engineer use?

A.Cloud Composer
B.Cloud Scheduler
C.Cloud Functions
D.Cloud Tasks
AnswerB

Cloud Scheduler is a fully managed cron job service that can trigger Vertex AI Pipelines via HTTP requests to the Vertex AI API. It supports cron expressions and time zones, making it ideal for scheduled pipeline runs. The engineer can create a job that invokes the pipeline at 2:00 AM UTC daily without managing any infrastructure.

Why this answer

Cloud Scheduler is the fully managed service for cron-based scheduling in Google Cloud. It can directly invoke the Vertex AI Pipelines API endpoint on a schedule, such as daily at 2:00 AM UTC. Other services either lack native scheduling or require additional components to achieve the same result.

Exam trap

The trap here is assuming that Cloud Composer is the default scheduler for Vertex AI Pipelines, when Cloud Scheduler is the simpler, fully managed option for basic cron triggers.

63
MCQeasy

You are defining a Python function component in KFP SDK v2. Which decorator should you use?

A.@dsl.task
B.@component
C.@dsl.pipeline
D.@dsl.component
AnswerD

@dsl.component converts a plain Python function into a lightweight KFP v2 component, automatically inferring its inputs, outputs and container image. This satisfies the requirement to define a function-based component without authoring a separate component YAML specification.

Why this answer

In KFP SDK v2, the `@dsl.component` decorator is used to define a Python function as a lightweight, reusable pipeline component that can be executed independently. This decorator automatically generates a containerized component from the function's signature and type annotations, enabling type-safe inputs and outputs without requiring a separate component YAML specification.

Exam trap

The exam often tests the distinction between v1 and v2 decorators, so the trap here is that candidates familiar with KFP SDK v1 may incorrectly choose `@component` (option B) instead of the v2-specific `@dsl.component`.

How to eliminate wrong answers

Option A is wrong because `@dsl.task` is not a valid decorator in KFP SDK v2; tasks are implicitly created when a component is called within a pipeline, not via a decorator. Option B is wrong because `@component` is the decorator from KFP SDK v1 (the older `kfp.components` module) and is not used in v2, which requires the `dsl` namespace. Option C is wrong because `@dsl.pipeline` is used to define a pipeline (a DAG of components), not a single component function.

64
MCQmedium

A team develops a pipeline that trains a model and evaluates it. They want to pass the test accuracy (a float) from the evaluation component to a subsequent deployment component. Which KFP SDK type should the evaluation component output be annotated with?

A.Output[float]
B.Output[Metrics]
C.Output[Artifact]
D.Output[ClassificationMetrics]
AnswerA

Output[float] annotates the component's return value as a scalar float, which is exactly the test accuracy type the deployment component consumes. KFP serialises this primitive and passes it as a typed input, avoiding string parsing or artifact handling for a simple numeric metric.

Why this answer

In the KFP SDK, a component output annotated as Output[float] is treated as a lightweight scalar parameter that is passed directly between components via the pipeline's execution graph. Since test accuracy is a single floating-point value (not a file, dataset, or structured metric object), Output[float] is the correct type to declare so the downstream deployment component can consume it as a typed input. KFP serializes this primitive and makes it available for parameter passing without needing artifact storage.

Exam trap

PMLE often tests the distinction between KFP parameter types (Output[float], Output[int], Output[str]) and artifact types (Output[Metrics], Output[Artifact], Output[ClassificationMetrics]), catching candidates who assume any numeric output should be a Metrics artifact.

How to eliminate wrong answers

Option B is wrong because Output[Metrics] is used to emit a dictionary of named scalar metrics for visualization in the KFP UI, not to pass a single value as a typed parameter to a downstream component. Option C is wrong because Output[Artifact] declares a file- or directory-based artifact (e.g., a model, dataset, or plot) that is stored in the artifact repository, which is overkill and semantically incorrect for a single float. Option D is wrong because Output[ClassificationMetrics] is a specialized artifact type for confusion matrices, ROC curves, and calibration plots, not for passing a scalar accuracy value between components.

65
MCQmedium

An ML engineer is building a Vertex AI pipeline that includes a component to train a model. The component takes a long time to run and occasionally fails due to transient errors (e.g., network timeouts). The engineer wants to automatically retry the component a few times if it fails. How should the engineer configure this?

A.Use a Cloud Scheduler job to trigger the pipeline again if it fails, and rely on the pipeline to resume from the failed step.
B.Set the pipeline's execution timeout to a high value and hope that the component succeeds on its own.
C.Set the retry policy on the pipeline task using the 'retry' parameter in the component's task definition, specifying the number of retries and backoff.
D.Configure the component to catch exceptions and loop internally until it succeeds, with no limit on retries.
AnswerC

Vertex AI Pipelines allows you to set a retry policy on individual tasks. By specifying the retry parameter with max_retries and backoff settings, the pipeline will automatically retry the component if it fails, which is ideal for transient errors.

Why this answer

To automatically retry a component on failure, the engineer should set a retry policy on the task. This is done by specifying the retry parameter in the task definition, including max_retries and backoff settings. This leverages Vertex AI Pipelines' built-in retry mechanism, which is designed for transient failures.

Exam trap

The trap here is thinking that re-triggering the entire pipeline or increasing timeouts solves transient failures, when a task-level retry policy is the correct mechanism.

66
Multi-Selectmedium

A machine learning team uses Vertex AI Pipelines for model training. They want to implement a conditional step that runs additional evaluation if the model accuracy exceeds 0.9, otherwise it runs a data augmentation component. Which two Kubeflow Pipelines SDK v2 constructs can they use to achieve this? (Choose two.)

Select 2 answers
A.dsl.ParallelFor
B.dsl.Else
C.dsl.ExitHandler
D.dsl.Collected
E.dsl.If
AnswersB, E

dsl.Else provides the alternative branch executed when the dsl.If condition evaluates false, so the data augmentation component runs whenever accuracy does not exceed 0.9. It is the required companion construct to dsl.If for building if/else conditionals in Kubeflow Pipelines SDK v2.

Why this answer

The team needs branching logic in a Vertex AI Pipelines (Kubeflow Pipelines SDK v2) graph, and dsl.If (option E) is the SDK v2 construct that creates a conditional branch whose condition is evaluated at runtime, so it correctly expresses 'if accuracy > 0.9, run the evaluation component.' dsl.Else (option B) is the companion construct used together with dsl.If to define the alternative branch, so it correctly expresses 'otherwise run the data augmentation component.' Together, dsl.If and dsl.Else form the if/else control-flow pattern required for this scenario. dsl.ParallelFor (option A) is for iterating over a list and launching parallel tasks, not for conditional branching. dsl.ExitHandler (option C) is for defining cleanup or exit tasks that run when a scope completes or fails, not for choosing between two branches. dsl.Collected (option D) is a type annotation used to collect outputs from loops, not a conditional construct.

Exam trap

The trap here is that candidates may confuse `dsl.ParallelFor` or `dsl.ExitHandler` with conditional constructs, but `dsl.If` and `dsl.Else` are the only SDK v2 constructs specifically designed for branching based on runtime conditions.

67
MCQeasy

A team is using Vertex AI Pipelines to orchestrate a machine learning workflow. They want to ensure that the pipeline can be reproduced with the same results even if the underlying data changes. Which practice should they follow?

A.Use a fixed random seed in the training component and snapshot the training data.
B.Run the pipeline on a dedicated Vertex AI training cluster with the same machine type and accelerators.
C.Version all pipeline components and pin their dependencies to specific versions.
D.Store the pipeline definition in a Git repository and use the same pipeline template for all runs.
AnswerA

Reproducibility requires controlling both code and data. A fixed random seed ensures deterministic behavior in stochastic processes like weight initialization. Snapshotting the training data (e.g., using a BigQuery snapshot or copying to a versioned Cloud Storage bucket) ensures the same data is used. Together, these practices enable reproducible results even if the source data changes later.

Why this answer

Reproducibility in ML pipelines requires controlling both the code and the data. A fixed random seed makes training deterministic, and snapshotting the training data ensures the same input is used across runs. Without data versioning, changes in the source data would lead to different results.

Therefore, combining a fixed seed with data snapshots is the correct approach.

Exam trap

The trap here is focusing only on code versioning or hardware consistency, while overlooking the critical role of data versioning and random seed control in achieving reproducible ML results.

68
MCQeasy

An ML engineer needs to trigger a Vertex AI Pipeline on a recurring schedule, every 24 hours, to retrain a model with the latest data. Which approach should they use to set up this schedule?

A.Use Cloud Tasks to queue pipeline runs daily.
B.Create a Cloud Scheduler job that calls the Vertex AI API to create a pipeline job.
C.Set a cron expression in the pipeline definition file using the 'schedule' parameter.
D.Use the Vertex AI Pipelines UI to set a schedule directly on the pipeline.
AnswerB

Cloud Scheduler provides cron-based recurring triggers, satisfying the 24-hour retraining requirement. Its HTTP target can authenticate to the Vertex AI API and invoke pipelines.create, launching a pipeline run on each fire. This avoids maintaining external orchestration infrastructure, unlike Cloud Functions or manual invocation, and natively supports the fixed daily cadence the stem demands.

Why this answer

Cloud Scheduler is the native Google Cloud service for cron-based job scheduling. By configuring a Cloud Scheduler job to call the Vertex AI API (e.g., via a HTTP POST to the projects.locations.pipelineJobs.create endpoint), the engineer can trigger a pipeline run every 24 hours. This approach is reliable, supports authentication via OAuth, and integrates directly with Vertex AI Pipelines without requiring additional orchestration code.

Exam trap

The PMLE exam often tests the misconception that Vertex AI Pipelines has a built-in scheduling feature (like a cron parameter in the pipeline definition or a UI schedule button), when in fact scheduling must be implemented using Cloud Scheduler or similar external services.

How to eliminate wrong answers

Option A is wrong because Cloud Tasks is a distributed task queue designed for asynchronous message delivery and retries, not for recurring cron-based scheduling; it would require an additional scheduler to enqueue tasks daily. Option C is wrong because Vertex AI Pipeline definitions do not support a 'schedule' parameter; scheduling is handled externally, not within the pipeline YAML or JSON definition. Option D is wrong because the Vertex AI Pipelines UI does not provide a built-in recurring schedule feature; schedules must be created using Cloud Scheduler or other external tools.

69
MCQmedium

A team is implementing CI/CD for ML using Cloud Build. They want to trigger a training pipeline in Vertex AI whenever a new model code is pushed to the main branch of the repository. Which Cloud Build configuration should they use to achieve this?

A.Set up a Cloud Build trigger that runs on push to any branch, and in the build step, use gcloud to submit a Vertex AI Pipeline job.
B.Use a Cloud Scheduler job to periodically check for new commits on main and trigger Cloud Build.
C.Use Cloud Functions to watch the repository and call Cloud Build on push to main.
D.Set up a Cloud Build trigger that runs on push to main branch, and in the build step, use gcloud to submit a Vertex AI Pipeline job.
AnswerD

A Cloud Build trigger scoped to pushes on the main branch fires the build, and a gcloud step submits the Vertex AI Pipeline job. This satisfies the stem's constraint that training be triggered specifically by new model code pushed to main.

Why this answer

Cloud Build triggers can be configured to fire specifically on pushes to the main branch. The build step then uses the gcloud command to submit a Vertex AI Pipeline job, which directly integrates the CI/CD pipeline with Vertex AI's orchestration. This approach is event-driven, immediate, and requires no additional services or polling.

Exam trap

Google often tests the candidate's understanding that Cloud Build triggers can be scoped to specific branches and that using gcloud directly in a build step is the simplest and most efficient way to invoke Vertex AI Pipelines, rather than introducing unnecessary intermediate services like Cloud Functions or Scheduler.

How to eliminate wrong answers

Option A is wrong because triggering on push to any branch would cause the pipeline to run on feature branches, pull requests, and other non-main branches, leading to unnecessary executions and potential conflicts. Option B is wrong because Cloud Scheduler polling is inefficient, introduces latency, and is not the intended event-driven mechanism; Cloud Build triggers are designed to react to repository events directly. Option C is wrong because using Cloud Functions as an intermediary adds unnecessary complexity and cost; Cloud Build natively supports repository event triggers without requiring a separate compute service.

70
Multi-Selectmedium

You are building a CI/CD pipeline for an ML model using Cloud Build. When code is pushed to the main branch, you want to automatically build a training image, run a Vertex AI pipeline, and if the model evaluation passes, deploy it to a staging endpoint. Which two components are essential for this CI/CD pipeline?

Select 2 answers
A.Cloud Scheduler to trigger the pipeline on a schedule.
B.Cloud Functions to deploy the model.
C.Vertex AI Pipelines to orchestrate training and evaluation.
D.Cloud Build trigger configured to respond to push events to the main branch.
E.Vertex AI Continuous Training service.
AnswersC, D

Vertex AI Pipelines orchestrates the training and evaluation steps as a managed DAG, executing each component container in sequence and surfacing the evaluation metrics that gate deployment. This satisfies the stem's requirement to run a pipeline and conditionally promote the model only when evaluation passes.

Why this answer

Option D is correct because a Cloud Build trigger is the native mechanism that responds to push events on the main branch, automatically starting the build of the training image and initiating the pipeline workflow. Option C is correct because Vertex AI Pipelines orchestrates the training and evaluation steps, and its evaluation component determines whether the model passes the quality gate before deployment to the staging endpoint. Option A is incorrect because Cloud Scheduler triggers on a time-based schedule, not on code push events, so it does not satisfy the push-to-main requirement.

Option B is incorrect because Cloud Functions is not the deployment mechanism for Vertex AI models; deployment is handled through Vertex AI endpoints or pipeline components. Option E is incorrect because Vertex AI Continuous Training is a managed retraining feature, not the CI/CD orchestration component needed to trigger and gate deployments on code changes.

Exam trap

Google often tests the distinction between event-driven triggers (Cloud Build trigger on push) and schedule-based triggers (Cloud Scheduler), so candidates mistakenly pick Cloud Scheduler when the requirement is for a code-push event.

71
MCQmedium

An organization wants to trigger a Vertex AI pipeline whenever a new commit is pushed to the main branch of their Cloud Source Repository. The pipeline should retrain and evaluate the model. Which service should they use to detect the push event and start the pipeline?

A.Vertex AI Pipeline schedule
B.Cloud Build trigger
C.Cloud Scheduler on a short interval
D.Pub/Sub with push subscription to a Cloud Function
AnswerB

Cloud Build triggers natively watch Cloud Source Repository branches and fire on push events, satisfying the commit-detection constraint. The trigger then invokes the Vertex AI pipeline, so no polling or custom webhook infrastructure is needed to bridge the repository event to pipeline execution.

Why this answer

Cloud Build triggers are designed to automatically invoke a build pipeline in response to events from Cloud Source Repository, such as a push to a specific branch. This allows the organization to directly start a Vertex AI pipeline for retraining and evaluation without additional infrastructure. Cloud Build triggers natively integrate with Cloud Source Repository, making them the simplest and most reliable choice for this event-driven workflow.

Exam trap

This question tests the distinction between event-driven triggers (Cloud Build) and time-based schedulers (Cloud Scheduler, Vertex AI Pipeline schedule), leading candidates to mistakenly choose a polling or cron-based option for an event-driven requirement.

How to eliminate wrong answers

Option A is wrong because Vertex AI Pipeline schedules are time-based (cron) triggers, not event-driven; they cannot detect a Git push event. Option C is wrong because Cloud Scheduler on a short interval would poll for changes, introducing latency and inefficiency, and it does not natively detect push events from Cloud Source Repository. Option D is wrong because while Pub/Sub with a push subscription to a Cloud Function could technically work, it adds unnecessary complexity and an extra compute layer; Cloud Build triggers provide a direct, managed integration without the need for custom code or additional services.

72
MCQhard

A company runs a Vertex AI Pipeline that includes a hyperparameter tuning step followed by a training step. The tuning step outputs the best hyperparameters. The engineer wants the training step to use these hyperparameters and to ensure that the training step only runs if tuning succeeds. Which approach should the engineer take?

A.Use a conditional expression in the training step to check if the tuning step succeeded, and only then proceed.
B.Configure the tuning step to write hyperparameters to a Cloud Storage bucket, and have the training step read from that bucket at runtime.
C.Run the tuning and training steps in parallel to save time, and use a shared memory store to exchange hyperparameters.
D.Pass the output of the tuning step as an input parameter to the training step and set the `after` parameter of the training step to include the tuning step.
AnswerD

In Vertex AI Pipelines, data dependencies are established by passing outputs as inputs. The tuning step's output (best hyperparameters) can be passed as an input to the training step. Additionally, the `after` parameter ensures the training step runs only after the tuning step completes successfully. This combination enforces both data flow and execution order.

Why this answer

To pass data between tasks in Vertex AI Pipelines, the output of one task must be provided as an input to another. The tuning step's output (best hyperparameters) is passed as an input to the training step, creating a data dependency. Additionally, the `after` parameter ensures the training step runs only after the tuning step succeeds.

This dual approach guarantees both correct data flow and execution order, and if tuning fails, the pipeline stops before training.

Exam trap

The trap here is relying on external storage or parallel execution to share data between tasks, which does not enforce execution order or failure propagation; Vertex AI Pipelines requires explicit input/output connections and the `after` parameter for dependencies.

73
MCQmedium

An ML engineer is building a Vertex AI pipeline that trains a model and then evaluates it. The evaluation component must compare the new model's accuracy against a fixed threshold and, if the model passes, trigger a downstream deployment component. The engineer wants to avoid running the deployment component when the model fails. Which approach should the engineer use?

A.Define a condition on the deployment component using the evaluation metric output from the evaluation component.
B.Use a pipeline parameter to pass the evaluation result to the deployment component and let the deployment component decide whether to run.
C.Configure the evaluation component to raise an exception if the model fails, which will stop the entire pipeline before reaching the deployment component.
D.Create a separate pipeline that only contains the deployment component and trigger it manually after reviewing the evaluation results.
AnswerA

Vertex AI Pipelines supports conditional execution using the dsl.Condition construct. By referencing the evaluation metric output from the evaluation component, the engineer can conditionally execute the deployment component only when the metric meets the threshold. This ensures the deployment component is not run when the model fails, saving resources and aligning with the pipeline's logic.

Why this answer

Conditional execution in Vertex AI Pipelines is implemented using the dsl.Condition construct, which allows a component to run only if a specified condition is met. By referencing the evaluation metric output, the pipeline can automatically decide whether to deploy the model. This approach avoids running unnecessary components and keeps the pipeline logic self-contained.

Exam trap

The trap here is assuming that passing evaluation results as parameters or raising exceptions can control component execution, when in fact Vertex AI Pipelines requires explicit conditional constructs for branching.

74
MCQeasy

A machine learning engineer is building a pipeline with Vertex AI Pipelines and wants to pass a large dataset between components without copying it to the container's memory. What is the best practice for passing data between pipeline components?

A.Mount an NFS volume to all containers and share data via the filesystem.
B.Use Cloud Storage URIs (gs://) to point to the data location.
C.Serialize the dataset to JSON and include it as a pipeline parameter.
D.Use the importer component to load the data into the pipeline as an in-memory artifact.
AnswerB

Passing gs:// URIs means components exchange only small string references, so the large dataset stays in Cloud Storage and each container reads it directly. This satisfies the constraint of avoiding copying data into container memory between pipeline components.

Why this answer

Vertex AI Pipelines natively supports passing Cloud Storage URIs (gs://) as artifact references between components, allowing components to read the dataset directly from GCS without copying it into container memory. This avoids memory limits and enables efficient handling of large datasets by leveraging GCS's scalable object storage.

Exam trap

A common mistake in this exam is to think that large data must be passed as in-memory artifacts or serialized parameters, when the correct Vertex AI Pipelines pattern is to pass a Cloud Storage URI and let components read data lazily from GCS.

How to eliminate wrong answers

Option A is wrong because mounting an NFS volume introduces network filesystem latency, requires additional infrastructure setup, and is not a native or recommended pattern in Vertex AI Pipelines, which is designed for serverless, cloud-native artifact passing. Option C is wrong because serializing a large dataset to JSON and including it as a pipeline parameter would exceed the maximum parameter size limit (typically 512 KB in Vertex AI Pipelines) and would force the entire dataset into memory, defeating the purpose of avoiding memory copies. Option D is wrong because the importer component registers an external artifact (e.g., a GCS URI) into the pipeline's metadata store but does not load data into memory; the misconception is that it creates an in-memory artifact, whereas it merely creates a metadata reference.

75
MCQhard

A team is using Vertex AI Pipelines to orchestrate a multi-step ML workflow. They need to pass a large dataset (several terabytes) between two components: a preprocessing component and a training component. The preprocessing component outputs a preprocessed dataset that the training component consumes. The team wants to minimize data transfer time and cost. What is the most efficient way to pass the data between these components?

A.Mount a Filestore instance as a shared volume to both components and write the data there.
B.Use a `Dataset` artifact to pass the data, and configure the artifact to store the data in the pipeline's metadata store.
C.Have the preprocessing component output the data as a base64-encoded string parameter, and pass it to the training component.
D.Write the preprocessed data to a Cloud Storage bucket and pass the bucket URI as an output parameter to the training component.
AnswerD

Passing a Cloud Storage URI as an output parameter is the standard and most efficient way to share large datasets between components in Vertex AI Pipelines. The training component can directly read from Cloud Storage, avoiding unnecessary data movement and leveraging Google's high-bandwidth network.

Why this answer

For large datasets, the best practice is to write data to Cloud Storage and pass the URI as a parameter. This avoids moving data through the pipeline's control plane and leverages Cloud Storage's scalability and high throughput. The training component can then read directly from the bucket.

Exam trap

The trap here is thinking that the pipeline must pass the actual data between components, when it should only pass references to data stored in Cloud Storage.

Page 1 of 2 · 81 questions totalNext →

Ready to test yourself?

Try a timed practice session using only Automating and Orchestrating ML Pipelines questions.