PMLE · domain
Automating and Orchestrating ML Pipelines
This domain covers building, running, and scheduling ML workflows on Vertex AI Pipelines and the KFP SDK v2. Questions test component authoring, artifact and parameter passing, conditional and parallel control flow, and pipeline orchestration choices rather than model quality itself.
Focused practice
Practice Automating and Orchestrating ML Pipelines questions
Scored sessions drawing only from this domain — pick a length below.
Start 20-question practice test →What this domain covers
What to know about Automating and Orchestrating ML Pipelines
Be able to author and compile a KFP v2 pipeline that passes artifacts and parameters correctly, extracts a metric for conditional deployment, and fans out shards with ParallelFor. The key is knowing when data belongs in an artifact versus a parameter.
Authoring Vertex AI Pipeline components with the KFP SDK v2 and compiling pipelines to YAML
Passing data between components via parameters versus artifacts, including large datasets and Metrics artifacts
Using lightweight Python-function components and prebuilt Google Cloud pipeline components without custom containers
Implementing control flow: conditionals, ParallelFor loops, and scheduling recurring pipeline runs
Watch out for
Common Automating and Orchestrating ML Pipelines exam traps
- ▸Passing large datasets as component parameters instead of artifacts, which bloats metadata and hits parameter size limits
- ▸Comparing a Metrics artifact's value directly in a pipeline condition without first extracting it into a parameter
- ▸Rebuilding a container image for every small code change instead of using lightweight Python components or prebuilt components
Question index
All Automating and Orchestrating ML Pipelines questions (81)
Click any question to see the full explanation, or start a practice session above.
You are tasked with building a robust ML pipeline that must be idempotent and handle data skew between training and serving. Which three practices should you implement?
Hard2A pipeline uses the Google Cloud Pipeline Components to perform AutoML training and batch prediction. Which two components from the GCPC library should they use? (Choose two.)
Medium3An ML engineer has a Vertex AI pipeline that trains a model and then evaluates it. The engineer wants the pipeline to automatically deploy the model to an endpoint only if the evaluation metric (e.g., accuracy) exceeds a threshold. The pipeline is defined using the Kubeflow Pipelines SDK. Which approach should the engineer use to implement this conditional deployment?
Medium4An ML engineer is designing a Vertex AI Pipeline that includes a custom training component. The component must read a dataset from a Cloud Storage bucket and write the trained model to another Cloud Storage location. The engineer wants the component to be reusable across pipelines and to ensure that the pipeline tracks the exact dataset and model artifacts. Which approach should the engineer take?
Medium5A data scientist is creating a Vertex AI pipeline using the Kubeflow Pipelines SDK v2. Which TWO statements about pipeline parameters are correct? (Choose two.)
Easy6A team is using Vertex AI Pipelines to orchestrate a training workflow. They want to ensure that the pipeline can be reproduced exactly six months later for auditing purposes. They need to capture all necessary information to rerun the pipeline and obtain identical results. (Choose two.)
Hard7Your organization wants to automate the retraining of a model when new data is available and also on a weekly schedule. Which TWO services would you use together to achieve this? (Choose two.)
Medium8A company wants to automatically retrain their model every night at 2 AM using Vertex AI Pipelines. Which approach should they use to trigger the pipeline on a schedule?
Easy9A company is using Vertex AI Pipelines for ML workflows. They want to implement best practices for idempotent components and data passing. Which THREE practices should they adopt?
Hard10An ML engineer is using Vertex AI Pipelines to orchestrate a workflow that includes a data preprocessing step, a training step, and an evaluation step. The evaluation step must run only if the training step succeeds and the evaluation metric meets a threshold. The engineer wants to define this logic natively in the pipeline without writing a custom component that exits with a specific code. Which Vertex AI Pipelines feature should they use?
Medium11An ML engineer is designing a Vertex AI pipeline that includes a custom training component. The component must read training data from a Cloud Storage bucket and write the trained model to a Vertex AI Model Registry. The engineer wants to ensure the component can access these resources securely. Which two configurations should the engineer implement? (Choose two.)
Hard12A machine learning engineer is using Vertex AI Pipelines and wants to run a custom Python function as a component. They need to pass a dataset artifact from a previous component and output a model artifact. Which decorator should they use to define the component in the Kubeflow Pipelines SDK v2?
Easy13A company wants to implement a CI/CD pipeline for their ML models using Vertex AI. They need to automatically retrain the model when new data arrives, but only if the model performance on a validation set has degraded by more than 5% compared to the current production model. Which three services or components should they incorporate into the automated pipeline? (Choose three.)
Hard14An ML engineer is building a pipeline component that takes a dataset URI and a model URI as inputs, and outputs a classification metrics artifact. Which KFP SDK v2 type should the output artifact be annotated with?
Medium15An organization runs a Vertex AI pipeline that includes a model evaluation step. Team members want to reuse previously computed evaluation metrics when re-running the pipeline with unchanged code and hyperparameters. Which feature should they enable?
Medium16A team has a pipeline that trains a model and then evaluates it. They want to conditionally deploy the model to a staging endpoint only if evaluation metrics exceed a threshold. Which KFP feature should they use?
Hard17An ML engineer is authoring a Vertex AI Pipelines component that runs a custom Python script. The component must accept a GCS path to training data and output a model artifact. The engineer wants the component interface to be strongly typed and to automatically generate the component specification from the Python function. Which approach should the engineer use?
Medium18An ML team is using Vertex AI Pipelines to orchestrate a training workflow. They need to pass a large dataset (500 GB) between two components. The first component preprocesses the data and writes the output to Cloud Storage. The second component trains a model using that preprocessed data. The team wants to minimize pipeline execution time and cost. Which two strategies should they use? (Choose two.)
Hard19A team wants to implement CI/CD for their ML pipeline using Cloud Build. They want to automatically compile and deploy the pipeline when code is pushed to the main branch. Which three steps should they include in the Cloud Build configuration? (Choose three.)
Hard20A team is designing a ML pipeline that includes training, evaluation, and conditional deployment. They want to use Vertex AI Pipelines. Which THREE concepts should they use? (Choose three.)
Hard21An ML engineer is building a Vertex AI pipeline that must run a custom training component for each of 12 hyperparameter combinations. The component is defined as a custom Python function (Lightweight Python component). The engineer wants each combination to run as a separate parallel task so the pipeline completes faster, and wants the pipeline to fail fast if any single trial fails. Which approach should the engineer take?
Medium22An ML team wants to run a hyperparameter tuning job on Vertex AI using a pre-built pipeline component. Which component should they use?
Medium23A machine learning engineer is building a Vertex AI pipeline that uses a pre-built Google Cloud Pipeline Components (GCPC) to train a custom model. Which component should the engineer use to submit a custom training job to Vertex AI?
Easy24An ML engineer is designing a Vertex AI pipeline that trains a model, evaluates it, and conditionally deploys it only if the evaluation metric meets a threshold. The pipeline must pass the evaluation metric from the evaluation component to a downstream conditional. Which Vertex AI Pipelines feature should the engineer use to implement this flow?
Medium25An ML engineer has a Vertex AI pipeline that trains a model and then evaluates it. The engineer wants the pipeline to deploy the model to an endpoint only if the evaluation metric exceeds a threshold defined at pipeline submission time. The threshold must be changeable without recompiling the pipeline. Which mechanism should the engineer use?
Medium26You are using KFP SDK v2 to define a pipeline. You need to pass a large dataset between components. What is the best practice for passing data?
Medium27A company runs a Vertex AI pipeline that uses a container component to preprocess data. The component downloads a large file from a public URL and saves the output to Cloud Storage. The pipeline fails intermittently with a 'timeout' error. Which THREE steps should the team take to improve reliability? (Choose three.)
Hard28An ML engineer is using Vertex AI Pipelines to orchestrate a workflow that includes a custom component for data validation. The component takes a dataset URI and outputs a validation report. The engineer wants to fail the pipeline immediately if the validation report indicates that the data is invalid, without running subsequent steps. How should the engineer implement this?
Hard29What is the purpose of the 'importer' component in Vertex AI Pipelines?
Easy30In a Vertex AI Pipeline, a component produces a Metrics artifact that includes an evaluation metric. The engineer wants to use this metric value as a condition to decide whether to deploy the model. However, the metric value is stored in the artifact's metadata and not directly as a pipeline parameter. How can the engineer pass the metric value to a downstream conditional task?
Hard31Which of the following is a best practice when designing idempotent pipeline components in Vertex AI?
Easy32A data scientist is defining a Vertex AI pipeline and needs to include a step that imports a pre-existing model from Cloud Storage into the pipeline as an artifact. Which Kubeflow Pipelines SDK v2 component should they use?
Easy33An ML engineer is running a Vertex AI pipeline that includes a data validation component and a training component. The engineer wants the pipeline to stop before training if data validation fails, but wants the validation component to record its result as an output artifact for later inspection. Which combination of pipeline features should the engineer use?
Hard34An ML engineer is building a continuous training pipeline that retrains a model when new data arrives. The pipeline should also detect skew between training and serving data. Which TWO Google Cloud services should they use? (Choose two.)
Medium35A company runs a Vertex AI Pipeline that includes a custom component for hyperparameter tuning. The component uses a large search space and runs many trials. The pipeline is taking too long to complete, and the team wants to reduce the execution time without sacrificing model quality. They have already optimized the training code. Which Vertex AI Pipelines feature should they use to speed up the tuning component?
Hard36A machine learning engineer needs to pass a large dataset between two components in a Vertex AI pipeline. What is the recommended way to pass this data?
Easy37An organization wants to use Cloud Composer (Airflow) to orchestrate a machine learning workflow that includes running a Vertex AI Pipeline, followed by a BigQuery job, and then a Dataflow pipeline. What is the primary advantage of using Cloud Composer for this orchestration?
Easy38A data engineer wants to orchestrate a complex workflow that includes running a Vertex AI pipeline, then a BigQuery job, and finally a Dataflow pipeline. The workflow must handle dependencies, retries, and monitoring. Which Google Cloud service is most suitable for this orchestration?
Easy39A team uses Cloud Build to automatically trigger a Vertex AI pipeline when changes are pushed to the model code repository. They have a cloudbuild.yaml file that builds a container image and submits the pipeline. However, they want to run the pipeline only if the commit includes changes to the 'training/' directory. Which Cloud Build configuration option should be used to filter the trigger?
Medium40A company wants to implement continuous delivery (CD) for ML models, where a model is automatically deployed to a staging environment and only promoted to production after passing an evaluation gate. Which combination of GCP services is BEST suited for orchestrating this CD pipeline?
Medium41A team runs a Vertex AI pipeline that includes a component which downloads a large dataset from BigQuery and writes it to Cloud Storage. The pipeline's caching is enabled by default. During iterative development, the engineer modifies the SQL query inside the component to include an additional feature column, but the pipeline still uses the previously cached output because the component's input parameters and code hash are unchanged. The engineer needs the component to re-execute with the updated query without disabling caching for the entire pipeline. What should the engineer do?
Medium42An organization wants to trigger a Vertex AI pipeline whenever new data arrives in a Cloud Storage bucket. Which approach should they use?
Medium43An ML pipeline must run a set of preprocessing tasks for each data shard in parallel. Which KFP SDK features should they use to implement this? (Choose two.)
Medium44An ML engineer is designing a pipeline that should run only when new training data arrives in a Cloud Storage bucket. Which event-driven approach should they use to trigger the Vertex AI Pipeline?
Medium45An ML engineer wants to containerize a custom training script and use it as a component in a Vertex AI Pipeline. The component should accept a dataset URI and a learning rate parameter, and output a trained model artifact. Which approach should the engineer use to define the component?
Medium46A company is deploying a Vertex AI pipeline that trains a model and then runs a custom evaluation component. The evaluation component must only run if the training component succeeds and the model's accuracy exceeds a threshold. The pipeline must also support retries for transient errors in the training component. The engineer needs to configure the pipeline to meet these requirements. Which two actions should the engineer take? (Choose two.)
Hard47An ML engineer is authoring a Vertex AI pipeline where a custom training component must read a dataset from a BigQuery table and write the trained model to a Cloud Storage bucket. The engineer wants the component to be reusable across projects and environments without hardcoding project IDs or bucket names. Which design should the engineer use?
Hard48An ML engineer is building a Vertex AI Pipeline that includes a data validation component. The component should fail the pipeline if the input data does not meet certain statistical thresholds. The engineer wants to ensure that the pipeline stops immediately and does not proceed to training if validation fails. Which mechanism should the engineer use in the component?
Medium49A company is using Vertex AI Pipelines to orchestrate a training workflow. They want to implement a CI/CD process where the pipeline is automatically triggered when a new version of the training code is pushed to a GitHub repository. They also want to ensure that the pipeline uses the latest code. Which two actions should they take? (Choose two.)
Medium50A team wants to implement continuous training for their ML model. The pipeline should be triggered when new training data arrives in a Cloud Storage bucket. Which combination of services should they use?
Medium51An ML engineer is designing a Vertex AI Pipeline that includes a hyperparameter tuning step. The tuning step runs multiple trials and outputs the best model. The engineer wants to ensure that the pipeline can resume from the tuning step if it fails later, without re-running all the tuning trials. What should the engineer do?
Hard52An ML engineer is creating a Vertex AI Pipeline that includes a component to train a model. The component requires a machine type with a GPU. The engineer wants to specify the machine type and GPU type for the component. Which approach should the engineer use?
Easy53A company has a CI/CD pipeline that retrains a model every time new training data is available. They want to automatically deploy the new model to production only if it passes a set of evaluation tests on a staging environment. Which approach best implements this?
Hard54A machine learning engineer is building a Vertex AI pipeline that uses a pre-built AutoML Tables component to train a classification model. The pipeline also includes a conditional step that deploys the model to an endpoint only if the evaluation metrics exceed a threshold. Which KFP feature should be used to implement the conditional deployment?
Hard55A team has a Vertex AI pipeline that includes a container component for data preprocessing. The team notices that the component is re-executed every time the pipeline runs, even when the inputs and code haven't changed. They want to leverage pipeline caching to avoid redundant executions. What should they do to enable caching for this component?
Medium56An engineer needs to compile a Kubeflow Pipeline defined in Python to a JSON format that can be run on Vertex AI Pipelines. Which command should they use?
Medium57A data science team wants to build a machine learning pipeline on Vertex AI Pipelines that preprocesses data, trains a model, and evaluates it. They need to ensure that components can be reused across multiple pipelines and that outputs from one component can be passed as inputs to another. Which approach should they take?
Medium58A machine learning engineer needs to create a pipeline that runs a custom container component on Vertex AI. The container expects a Cloud Storage path as input and outputs a model artifact. Which component type should they define using the Kubeflow Pipelines SDK v2?
Medium59A company runs a Vertex AI Pipeline that trains a model and then deploys it to a Vertex AI Endpoint. The pipeline uses a conditional deployment step based on the model's evaluation metric. The team wants to ensure that if the evaluation metric falls below a threshold, the pipeline fails and no deployment occurs. Which approach should they use?
Hard60A data science team wants to build a Vertex AI pipeline that trains a model, evaluates it, and conditionally deploys it if the accuracy exceeds 0.9. They want to use the Kubeflow Pipelines SDK v2. Which construct allows them to conditionally execute the deployment step based on the evaluation metric?
Medium61What is the primary benefit of using pipeline caching in Vertex AI Pipelines?
Easy62An ML engineer is using Vertex AI Pipelines to orchestrate a training workflow. The pipeline must run on a schedule every day at 2:00 AM UTC. The engineer wants to use a fully managed Google Cloud service to trigger the pipeline. Which service should the engineer use?
Easy63You are defining a Python function component in KFP SDK v2. Which decorator should you use?
Easy64A team develops a pipeline that trains a model and evaluates it. They want to pass the test accuracy (a float) from the evaluation component to a subsequent deployment component. Which KFP SDK type should the evaluation component output be annotated with?
Medium65An ML engineer is building a Vertex AI pipeline that includes a component to train a model. The component takes a long time to run and occasionally fails due to transient errors (e.g., network timeouts). The engineer wants to automatically retry the component a few times if it fails. How should the engineer configure this?
Medium66A machine learning team uses Vertex AI Pipelines for model training. They want to implement a conditional step that runs additional evaluation if the model accuracy exceeds 0.9, otherwise it runs a data augmentation component. Which two Kubeflow Pipelines SDK v2 constructs can they use to achieve this? (Choose two.)
Medium67A team is using Vertex AI Pipelines to orchestrate a machine learning workflow. They want to ensure that the pipeline can be reproduced with the same results even if the underlying data changes. Which practice should they follow?
Easy68An ML engineer needs to trigger a Vertex AI Pipeline on a recurring schedule, every 24 hours, to retrain a model with the latest data. Which approach should they use to set up this schedule?
Easy69A team is implementing CI/CD for ML using Cloud Build. They want to trigger a training pipeline in Vertex AI whenever a new model code is pushed to the main branch of the repository. Which Cloud Build configuration should they use to achieve this?
Medium70You are building a CI/CD pipeline for an ML model using Cloud Build. When code is pushed to the main branch, you want to automatically build a training image, run a Vertex AI pipeline, and if the model evaluation passes, deploy it to a staging endpoint. Which two components are essential for this CI/CD pipeline?
Medium71An organization wants to trigger a Vertex AI pipeline whenever a new commit is pushed to the main branch of their Cloud Source Repository. The pipeline should retrain and evaluate the model. Which service should they use to detect the push event and start the pipeline?
Medium72A company runs a Vertex AI Pipeline that includes a hyperparameter tuning step followed by a training step. The tuning step outputs the best hyperparameters. The engineer wants the training step to use these hyperparameters and to ensure that the training step only runs if tuning succeeds. Which approach should the engineer take?
Hard73An ML engineer is building a Vertex AI pipeline that trains a model and then evaluates it. The evaluation component must compare the new model's accuracy against a fixed threshold and, if the model passes, trigger a downstream deployment component. The engineer wants to avoid running the deployment component when the model fails. Which approach should the engineer use?
Medium74A machine learning engineer is building a pipeline with Vertex AI Pipelines and wants to pass a large dataset between components without copying it to the container's memory. What is the best practice for passing data between pipeline components?
Easy75A team is using Vertex AI Pipelines to orchestrate a multi-step ML workflow. They need to pass a large dataset (several terabytes) between two components: a preprocessing component and a training component. The preprocessing component outputs a preprocessed dataset that the training component consumes. The team wants to minimize data transfer time and cost. What is the most efficient way to pass the data between these components?
Hard76A data scientist wants to define a lightweight Python function component in Vertex AI Pipelines using Kubeflow Pipelines SDK v2. Which decorator should be applied to the function to make it a pipeline component?
Easy77A company uses Vertex AI Pipelines to train models on a daily schedule. The pipeline includes a component that runs a BigQuery query to extract features. The team wants to ensure that if the BigQuery component fails due to transient network errors, the pipeline automatically retries it. How can they configure retries in Vertex AI Pipelines?
Medium78A pipeline includes a component that produces a model artifact. The team wants to automatically detect skew between the training data distribution and the serving data distribution. Which three best practices should they implement? (Choose three.)
Hard79An ML pipeline runs on Vertex AI and includes a component that uses a third-party library not available in the default Python environment. The team wants to avoid building a custom container image. Which approach should they use?
Hard80A team is building a CI/CD pipeline for an ML model. They want to automatically trigger a Vertex AI pipeline for retraining whenever new training data arrives in a Cloud Storage bucket, but only if a specific Pub/Sub notification is published by a data ingestion process. Which approach meets these requirements with minimal operational overhead?
Hard81A machine learning engineer wants to define a lightweight pipeline component that runs custom Python code without building a container image. Which KFP SDK feature should they use?
EasyOther domains
All PMLE exam domains
Frequently asked questions
- What does the Automating and Orchestrating ML Pipelines domain cover on the PMLE exam?
- Be able to author and compile a KFP v2 pipeline that passes artifacts and parameters correctly, extracts a metric for conditional deployment, and fans out shards with ParallelFor. The key is knowing when data belongs in an artifact versus a parameter.
- How many questions are in this domain?
- This page lists all 81 Automating and Orchestrating ML Pipelines questions in the PMLE question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
- What is the best way to practise this domain?
- Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
- Can I practise only Automating and Orchestrating ML Pipelines questions?
- Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.