Be able to author and compile a KFP v2 pipeline that passes artifacts and parameters correctly, extracts a metric for conditional deployment, and fans out shards with ParallelFor. The key is knowing when data belongs in an artifact versus a parameter.
Start practicing
Automating and Orchestrating ML Pipelines — choose a session length
Free · No account required
Domain overview
This domain covers building, running, and scheduling ML workflows on Vertex AI Pipelines and the KFP SDK v2. Questions test component authoring, artifact and parameter passing, conditional and parallel control flow, and pipeline orchestration choices rather than model quality itself.
Exam objectives
Authoring Vertex AI Pipeline components with the KFP SDK v2 and compiling pipelines to YAML
Passing data between components via parameters versus artifacts, including large datasets and Metrics artifacts
Using lightweight Python-function components and prebuilt Google Cloud pipeline components without custom containers
Implementing control flow: conditionals, ParallelFor loops, and scheduling recurring pipeline runs
Passing large datasets as component parameters instead of artifacts, which bloats metadata and hits parameter size limits
Comparing a Metrics artifact's value directly in a pipeline condition without first extracting it into a parameter
Rebuilding a container image for every small code change instead of using lightweight Python components or prebuilt components
Click any question to see the full explanation and answer options, or start a focused practice session above.
A machine learning engineer is building a Vertex AI pipeline that uses a pre-built Google Cloud Pipeline Components (GCPC) to train a custom model. Which component should the engineer use to submit a custom training job to Vertex AI?
2A team has a Vertex AI pipeline that includes a container component for data preprocessing. The team notices that the component is re-executed every time the pipeline runs, even when the inputs and code haven't changed. They want to leverage pipeline caching to avoid redundant executions. What should they do to enable caching for this component?
3A data engineer wants to orchestrate a complex workflow that includes running a Vertex AI pipeline, then a BigQuery job, and finally a Dataflow pipeline. The workflow must handle dependencies, retries, and monitoring. Which Google Cloud service is most suitable for this orchestration?
4A machine learning engineer is building a Vertex AI pipeline that uses a pre-built AutoML Tables component to train a classification model. The pipeline also includes a conditional step that deploys the model to an endpoint only if the evaluation metrics exceed a threshold. Which KFP feature should be used to implement the conditional deployment?
5A team uses Cloud Build to automatically trigger a Vertex AI pipeline when changes are pushed to the model code repository. They have a cloudbuild.yaml file that builds a container image and submits the pipeline. However, they want to run the pipeline only if the commit includes changes to the 'training/' directory. Which Cloud Build configuration option should be used to filter the trigger?
6A machine learning engineer needs to pass a large dataset between two components in a Vertex AI pipeline. What is the recommended way to pass this data?
7A company runs a Vertex AI pipeline that uses a container component to preprocess data. The component downloads a large file from a public URL and saves the output to Cloud Storage. The pipeline fails intermittently with a 'timeout' error. Which THREE steps should the team take to improve reliability? (Choose three.)
8A data scientist is creating a Vertex AI pipeline using the Kubeflow Pipelines SDK v2. Which TWO statements about pipeline parameters are correct? (Choose two.)
9A data science team wants to build a Vertex AI pipeline that trains a model, evaluates it, and conditionally deploys it if the accuracy exceeds 0.9. They want to use the Kubeflow Pipelines SDK v2. Which construct allows them to conditionally execute the deployment step based on the evaluation metric?
10A team wants to implement continuous training for their ML model. The pipeline should be triggered when new training data arrives in a Cloud Storage bucket. Which combination of services should they use?
11You are using KFP SDK v2 to define a pipeline. You need to pass a large dataset between components. What is the best practice for passing data?
12You are defining a Python function component in KFP SDK v2. Which decorator should you use?
13You are building a CI/CD pipeline for an ML model using Cloud Build. When code is pushed to the main branch, you want to automatically build a training image, run a Vertex AI pipeline, and if the model evaluation passes, deploy it to a staging endpoint. Which two components are essential for this CI/CD pipeline?
14You are tasked with building a robust ML pipeline that must be idempotent and handle data skew between training and serving. Which three practices should you implement?
15An organization wants to trigger a Vertex AI pipeline whenever new data arrives in a Cloud Storage bucket. Which approach should they use?
16A team has a pipeline that trains a model and then evaluates it. They want to conditionally deploy the model to a staging endpoint only if evaluation metrics exceed a threshold. Which KFP feature should they use?
17What is the primary benefit of using pipeline caching in Vertex AI Pipelines?
18An engineer needs to compile a Kubeflow Pipeline defined in Python to a JSON format that can be run on Vertex AI Pipelines. Which command should they use?
19A company has a CI/CD pipeline that retrains a model every time new training data is available. They want to automatically deploy the new model to production only if it passes a set of evaluation tests on a staging environment. Which approach best implements this?
20Which of the following is a best practice when designing idempotent pipeline components in Vertex AI?
21An ML team wants to run a hyperparameter tuning job on Vertex AI using a pre-built pipeline component. Which component should they use?
22What is the purpose of the 'importer' component in Vertex AI Pipelines?
23Your organization wants to automate the retraining of a model when new data is available and also on a weekly schedule. Which TWO services would you use together to achieve this? (Choose two.)
24A team is designing a ML pipeline that includes training, evaluation, and conditional deployment. They want to use Vertex AI Pipelines. Which THREE concepts should they use? (Choose three.)
25An ML engineer is building a continuous training pipeline that retrains a model when new data arrives. The pipeline should also detect skew between training and serving data. Which TWO Google Cloud services should they use? (Choose two.)
26A data scientist wants to define a lightweight Python function component in Vertex AI Pipelines using Kubeflow Pipelines SDK v2. Which decorator should be applied to the function to make it a pipeline component?
27An ML engineer wants to containerize a custom training script and use it as a component in a Vertex AI Pipeline. The component should accept a dataset URI and a learning rate parameter, and output a trained model artifact. Which approach should the engineer use to define the component?
28An ML engineer needs to trigger a Vertex AI Pipeline on a recurring schedule, every 24 hours, to retrain a model with the latest data. Which approach should they use to set up this schedule?
29A company wants to implement continuous delivery (CD) for ML models, where a model is automatically deployed to a staging environment and only promoted to production after passing an evaluation gate. Which combination of GCP services is BEST suited for orchestrating this CD pipeline?
30An organization wants to use Cloud Composer (Airflow) to orchestrate a machine learning workflow that includes running a Vertex AI Pipeline, followed by a BigQuery job, and then a Dataflow pipeline. What is the primary advantage of using Cloud Composer for this orchestration?
31An ML engineer is designing a pipeline that should run only when new training data arrives in a Cloud Storage bucket. Which event-driven approach should they use to trigger the Vertex AI Pipeline?
32A team is implementing CI/CD for ML using Cloud Build. They want to trigger a training pipeline in Vertex AI whenever a new model code is pushed to the main branch of the repository. Which Cloud Build configuration should they use to achieve this?
33In a Vertex AI Pipeline, a component produces a Metrics artifact that includes an evaluation metric. The engineer wants to use this metric value as a condition to decide whether to deploy the model. However, the metric value is stored in the artifact's metadata and not directly as a pipeline parameter. How can the engineer pass the metric value to a downstream conditional task?
34An ML engineer is building a pipeline component that takes a dataset URI and a model URI as inputs, and outputs a classification metrics artifact. Which KFP SDK v2 type should the output artifact be annotated with?
35A company is using Vertex AI Pipelines for ML workflows. They want to implement best practices for idempotent components and data passing. Which THREE practices should they adopt?
36A data science team wants to build a machine learning pipeline on Vertex AI Pipelines that preprocesses data, trains a model, and evaluates it. They need to ensure that components can be reused across multiple pipelines and that outputs from one component can be passed as inputs to another. Which approach should they take?
37A team is building a CI/CD pipeline for an ML model. They want to automatically trigger a Vertex AI pipeline for retraining whenever new training data arrives in a Cloud Storage bucket, but only if a specific Pub/Sub notification is published by a data ingestion process. Which approach meets these requirements with minimal operational overhead?
38A machine learning engineer is using Vertex AI Pipelines and wants to run a custom Python function as a component. They need to pass a dataset artifact from a previous component and output a model artifact. Which decorator should they use to define the component in the Kubeflow Pipelines SDK v2?
39A data scientist is defining a Vertex AI pipeline and needs to include a step that imports a pre-existing model from Cloud Storage into the pipeline as an artifact. Which Kubeflow Pipelines SDK v2 component should they use?
40A company uses Vertex AI Pipelines to train models on a daily schedule. The pipeline includes a component that runs a BigQuery query to extract features. The team wants to ensure that if the BigQuery component fails due to transient network errors, the pipeline automatically retries it. How can they configure retries in Vertex AI Pipelines?
41A machine learning engineer needs to create a pipeline that runs a custom container component on Vertex AI. The container expects a Cloud Storage path as input and outputs a model artifact. Which component type should they define using the Kubeflow Pipelines SDK v2?
42An organization wants to trigger a Vertex AI pipeline whenever a new commit is pushed to the main branch of their Cloud Source Repository. The pipeline should retrain and evaluate the model. Which service should they use to detect the push event and start the pipeline?
43A machine learning engineer is building a pipeline with Vertex AI Pipelines and wants to pass a large dataset between components without copying it to the container's memory. What is the best practice for passing data between pipeline components?
44A company wants to implement a CI/CD pipeline for their ML models using Vertex AI. They need to automatically retrain the model when new data arrives, but only if the model performance on a validation set has degraded by more than 5% compared to the current production model. Which three services or components should they incorporate into the automated pipeline? (Choose three.)
45A machine learning team uses Vertex AI Pipelines for model training. They want to implement a conditional step that runs additional evaluation if the model accuracy exceeds 0.9, otherwise it runs a data augmentation component. Which two Kubeflow Pipelines SDK v2 constructs can they use to achieve this? (Choose two.)
46A machine learning engineer wants to define a lightweight pipeline component that runs custom Python code without building a container image. Which KFP SDK feature should they use?
47An organization runs a Vertex AI pipeline that includes a model evaluation step. Team members want to reuse previously computed evaluation metrics when re-running the pipeline with unchanged code and hyperparameters. Which feature should they enable?
48A company wants to automatically retrain their model every night at 2 AM using Vertex AI Pipelines. Which approach should they use to trigger the pipeline on a schedule?
49A team develops a pipeline that trains a model and evaluates it. They want to pass the test accuracy (a float) from the evaluation component to a subsequent deployment component. Which KFP SDK type should the evaluation component output be annotated with?
50An ML pipeline runs on Vertex AI and includes a component that uses a third-party library not available in the default Python environment. The team wants to avoid building a custom container image. Which approach should they use?
51A pipeline uses the Google Cloud Pipeline Components to perform AutoML training and batch prediction. Which two components from the GCPC library should they use? (Choose two.)
52An ML pipeline must run a set of preprocessing tasks for each data shard in parallel. Which KFP SDK features should they use to implement this? (Choose two.)
53A team wants to implement CI/CD for their ML pipeline using Cloud Build. They want to automatically compile and deploy the pipeline when code is pushed to the main branch. Which three steps should they include in the Cloud Build configuration? (Choose three.)
54A pipeline includes a component that produces a model artifact. The team wants to automatically detect skew between the training data distribution and the serving data distribution. Which three best practices should they implement? (Choose three.)
55A company runs a Vertex AI Pipeline that includes a hyperparameter tuning step followed by a training step. The tuning step outputs the best hyperparameters. The engineer wants the training step to use these hyperparameters and to ensure that the training step only runs if tuning succeeds. Which approach should the engineer take?
56An ML engineer has a Vertex AI pipeline that trains a model and then evaluates it. The engineer wants the pipeline to automatically deploy the model to an endpoint only if the evaluation metric (e.g., accuracy) exceeds a threshold. The pipeline is defined using the Kubeflow Pipelines SDK. Which approach should the engineer use to implement this conditional deployment?
57An ML engineer is building a Vertex AI pipeline that trains a model and then evaluates it. The evaluation component must compare the new model's accuracy against a fixed threshold and, if the model passes, trigger a downstream deployment component. The engineer wants to avoid running the deployment component when the model fails. Which approach should the engineer use?
58An ML engineer is building a Vertex AI pipeline that includes a component to train a model. The component takes a long time to run and occasionally fails due to transient errors (e.g., network timeouts). The engineer wants to automatically retry the component a few times if it fails. How should the engineer configure this?
59An ML engineer is designing a Vertex AI Pipeline that includes a custom training component. The component must read a dataset from a Cloud Storage bucket and write the trained model to another Cloud Storage location. The engineer wants the component to be reusable across pipelines and to ensure that the pipeline tracks the exact dataset and model artifacts. Which approach should the engineer take?
60A team is using Vertex AI Pipelines to orchestrate a multi-step ML workflow. They need to pass a large dataset (several terabytes) between two components: a preprocessing component and a training component. The preprocessing component outputs a preprocessed dataset that the training component consumes. The team wants to minimize data transfer time and cost. What is the most efficient way to pass the data between these components?
61An ML engineer is authoring a Vertex AI Pipelines component that runs a custom Python script. The component must accept a GCS path to training data and output a model artifact. The engineer wants the component interface to be strongly typed and to automatically generate the component specification from the Python function. Which approach should the engineer use?
62A company runs a Vertex AI Pipeline that includes a custom component for hyperparameter tuning. The component uses a large search space and runs many trials. The pipeline is taking too long to complete, and the team wants to reduce the execution time without sacrificing model quality. They have already optimized the training code. Which Vertex AI Pipelines feature should they use to speed up the tuning component?
63An ML team is using Vertex AI Pipelines to orchestrate a training workflow. They need to pass a large dataset (500 GB) between two components. The first component preprocesses the data and writes the output to Cloud Storage. The second component trains a model using that preprocessed data. The team wants to minimize pipeline execution time and cost. Which two strategies should they use? (Choose two.)
64An ML engineer is building a Vertex AI Pipeline that includes a data validation component. The component should fail the pipeline if the input data does not meet certain statistical thresholds. The engineer wants to ensure that the pipeline stops immediately and does not proceed to training if validation fails. Which mechanism should the engineer use in the component?
65A team runs a Vertex AI pipeline that includes a component which downloads a large dataset from BigQuery and writes it to Cloud Storage. The pipeline's caching is enabled by default. During iterative development, the engineer modifies the SQL query inside the component to include an additional feature column, but the pipeline still uses the previously cached output because the component's input parameters and code hash are unchanged. The engineer needs the component to re-execute with the updated query without disabling caching for the entire pipeline. What should the engineer do?
66An ML engineer is designing a Vertex AI pipeline that trains a model, evaluates it, and conditionally deploys it only if the evaluation metric meets a threshold. The pipeline must pass the evaluation metric from the evaluation component to a downstream conditional. Which Vertex AI Pipelines feature should the engineer use to implement this flow?
67A company is deploying a Vertex AI pipeline that trains a model and then runs a custom evaluation component. The evaluation component must only run if the training component succeeds and the model's accuracy exceeds a threshold. The pipeline must also support retries for transient errors in the training component. The engineer needs to configure the pipeline to meet these requirements. Which two actions should the engineer take? (Choose two.)
68A team is using Vertex AI Pipelines to orchestrate a training workflow. They want to ensure that the pipeline can be reproduced exactly six months later for auditing purposes. They need to capture all necessary information to rerun the pipeline and obtain identical results. (Choose two.)
69An ML engineer is building a Vertex AI pipeline that must run a custom training component for each of 12 hyperparameter combinations. The component is defined as a custom Python function (Lightweight Python component). The engineer wants each combination to run as a separate parallel task so the pipeline completes faster, and wants the pipeline to fail fast if any single trial fails. Which approach should the engineer take?
70An ML engineer is creating a Vertex AI Pipeline that includes a component to train a model. The component requires a machine type with a GPU. The engineer wants to specify the machine type and GPU type for the component. Which approach should the engineer use?
71An ML engineer is designing a Vertex AI pipeline that includes a custom training component. The component must read training data from a Cloud Storage bucket and write the trained model to a Vertex AI Model Registry. The engineer wants to ensure the component can access these resources securely. Which two configurations should the engineer implement? (Choose two.)
72An ML engineer is authoring a Vertex AI pipeline where a custom training component must read a dataset from a BigQuery table and write the trained model to a Cloud Storage bucket. The engineer wants the component to be reusable across projects and environments without hardcoding project IDs or bucket names. Which design should the engineer use?
73An ML engineer is using Vertex AI Pipelines to orchestrate a workflow that includes a data preprocessing step, a training step, and an evaluation step. The evaluation step must run only if the training step succeeds and the evaluation metric meets a threshold. The engineer wants to define this logic natively in the pipeline without writing a custom component that exits with a specific code. Which Vertex AI Pipelines feature should they use?
74A company runs a Vertex AI Pipeline that trains a model and then deploys it to a Vertex AI Endpoint. The pipeline uses a conditional deployment step based on the model's evaluation metric. The team wants to ensure that if the evaluation metric falls below a threshold, the pipeline fails and no deployment occurs. Which approach should they use?
75An ML engineer has a Vertex AI pipeline that trains a model and then evaluates it. The engineer wants the pipeline to deploy the model to an endpoint only if the evaluation metric exceeds a threshold defined at pipeline submission time. The threshold must be changeable without recompiling the pipeline. Which mechanism should the engineer use?
76An ML engineer is running a Vertex AI pipeline that includes a data validation component and a training component. The engineer wants the pipeline to stop before training if data validation fails, but wants the validation component to record its result as an output artifact for later inspection. Which combination of pipeline features should the engineer use?
77An ML engineer is designing a Vertex AI Pipeline that includes a hyperparameter tuning step. The tuning step runs multiple trials and outputs the best model. The engineer wants to ensure that the pipeline can resume from the tuning step if it fails later, without re-running all the tuning trials. What should the engineer do?
78A team is using Vertex AI Pipelines to orchestrate a machine learning workflow. They want to ensure that the pipeline can be reproduced with the same results even if the underlying data changes. Which practice should they follow?
79A company is using Vertex AI Pipelines to orchestrate a training workflow. They want to implement a CI/CD process where the pipeline is automatically triggered when a new version of the training code is pushed to a GitHub repository. They also want to ensure that the pipeline uses the latest code. Which two actions should they take? (Choose two.)
80An ML engineer is using Vertex AI Pipelines to orchestrate a workflow that includes a custom component for data validation. The component takes a dataset URI and outputs a validation report. The engineer wants to fail the pipeline immediately if the validation report indicates that the data is invalid, without running subsequent steps. How should the engineer implement this?
81An ML engineer is using Vertex AI Pipelines to orchestrate a training workflow. The pipeline must run on a schedule every day at 2:00 AM UTC. The engineer wants to use a fully managed Google Cloud service to trigger the pipeline. Which service should the engineer use?
Be able to author and compile a KFP v2 pipeline that passes artifacts and parameters correctly, extracts a metric for conditional deployment, and fans out shards with ParallelFor. The key is knowing when data belongs in an artifact versus a parameter.
The Courseiva PMLE question bank contains 81 questions in the Automating and Orchestrating ML Pipelines domain, covering the 18% of the exam attributed to this domain in the official Google Cloud blueprint. Click any question to see the full explanation and answer breakdown.
Start with a 10-question focused session to identify your baseline accuracy in this domain. Read every explanation — even for questions you answer correctly — to understand the reasoning. Once you score consistently above 80%, move to a 20–30 question session to confirm depth before moving to the next domain.
Yes — the session launcher on this page draws questions exclusively from the Automating and Orchestrating ML Pipelines domain. Choose 10, 20, 30, or 50 questions for a focused session, or click individual questions to review them one by one.
Save your results, see per-domain analytics, and get readiness scores — free, for every certification.
Sign Up FreeFree forever · Every certification included