Be able to wire an automatic retraining trigger to Vertex AI Pipelines, share Workbench notebooks with Git versioning, and standardize preprocessing as reusable components. The single most important thing: choose the managed Google Cloud service that enforces shared, versioned artifacts rather than manual copying.
Start practicing
Collaborating to manage data and models — choose a session length
Free · No account required
Domain overview
This domain covers how ML teams share code, data, and models on Google Cloud. Expect questions on Vertex AI Pipelines triggers, Vertex AI Workbench notebook sharing and versioning, enforcing shared preprocessing, Cloud Composer collaboration, and Artifact Registry or Model Registry usage for reusable, governed assets across teams.
Exam objectives
Triggering Vertex AI Pipelines automatically when new data lands, using Cloud Storage events or Cloud Scheduler
Sharing Vertex AI Workbench notebooks with version history via Git integration in the Workbench instance
Enforcing consistent preprocessing across teams with containerized components or shared Vertex AI Pipelines templates
Using Cloud Composer DAGs, shared repositories, and IAM roles to coordinate multi-team ML workflows
Assuming Vertex AI Pipelines retrain themselves on new data; you must configure an explicit trigger such as a Cloud Storage event or scheduled run.
Sharing notebooks by copying files instead of using Git-backed version history in Vertex AI Workbench, losing provenance and collaboration.
Letting each team write its own preprocessing code rather than packaging shared steps as reusable pipeline components or containers.
Click any question to see the full explanation and answer options, or start a focused practice session above.
Which TWO statements about Vertex AI Feature Store are correct? (Choose 2)
2Which THREE actions are best practices for managing ML models in production on Google Cloud? (Choose 3)
3A healthcare organization is building a machine learning model to predict patient readmission risk. They have sensitive data stored in BigQuery that includes protected health information (PHI). The data science team uses Vertex AI Workbench notebooks to explore the data and develop models. The organization's security policy requires that all PHI data must be encrypted at rest and in transit, and that access to the data is logged and audited. They also need to ensure that the data used for model training is de-identified to remove direct identifiers such as patient names and SSNs. The team wants to automate the de-identification process as part of the data pipeline. Which approach meets these requirements?
4Drag and drop the steps to deploy a trained TensorFlow model to Vertex AI Prediction in the correct order.
5A team of ML engineers is collaborating on a project using Vertex AI. They want to ensure that only approved models are deployed to production. Which approach should they use?
6A company uses a Cloud Composer DAG to run a daily ML pipeline that includes Dataflow jobs and model training on Vertex AI. The pipeline frequently fails due to insufficient permissions when the Dataflow worker accesses data in Cloud Storage. What is the most efficient way to resolve this issue?
7A data scientist wants to share a trained model with the team for review before deployment. The model is stored in Vertex AI Model Registry. What is the recommended way to grant the team read access to the model?
8Which TWO actions are recommended for collaborating on machine learning models using Vertex AI Model Registry?
9Which TWO strategies help ensure data consistency when multiple teams are contributing features to a shared Vertex AI Feature Store?
10Which THREE practices improve collaboration when using Cloud Composer for ML pipelines?
11A data science team uses Vertex AI Workbench and wants to share notebooks with version history. Which service should they use?
12A team uses Vertex AI Pipelines. They need to ensure that only certain team members can deploy models to production. What is the best approach?
13A company has multiple teams working on different models. They want to enforce consistent data preprocessing steps across all teams. Which approach should they take?
14A data scientist wants to track the lineage of a dataset used in a training run. Which Vertex AI feature should they use?
15An MLOps team needs to automatically retrain a model when new training data becomes available. They use Vertex AI Pipelines. What is the recommended way to trigger the pipeline?
16A large organization uses a multi-project setup with a central data lake. Different teams manage their own models. To enable cross-team sharing of features, they want to use Vertex AI Feature Store. What is the best practice to manage access?
17When distributing training across multiple workers using Vertex AI Training, how should the team share the training dataset?
18A team is using Vertex AI Experiments to compare different hyperparameters. They want to automatically record the hyperparameters. What is the correct way?
19A data engineering team uses Dataflow for preprocessing and wants to integrate with Vertex AI Pipelines. They need to pass the preprocessed data location to the training step. What is the best practice?
20Which TWO practices help ensure reproducible ML experiments?
21Which THREE actions should be taken to manage model versions effectively?
22Refer to the exhibit. The team wants to automatically deploy the best-performing model version to production. They have set up a Cloud Function triggered by Model Registry events. Which alias should they use in the function to get the latest champion?
23Refer to the exhibit. The team notices that the pipeline fails to read data from the specified Cloud Storage path. What is the most likely issue?
24A team uses Vertex AI Feature Store for storing features. They want to share feature definitions with other teams in a collaborative manner. What is the best way to collaborate on feature definitions?
25A company uses BigQuery to store feature data for ML training. A data engineer notices that a Vertex AI Training job is failing with 'Access Denied' errors when reading from a BigQuery table. The training job uses a custom service account that has been granted the 'bigquery.dataViewer' role on the dataset. What is the most likely cause of the failure?
26An organization uses Cloud Dataflow to preprocess training data. Dataflow jobs are often failing because of insufficient quota for certain resources. The team has requested a quota increase, but the jobs still fail with 'quota exceeded' errors for a different resource. They want to proactively monitor and manage quotas to avoid failures. What is the best approach?
27Which TWO of the following are best practices for managing data in a collaborative machine learning environment on Google Cloud?
28Which THREE of the following are recommended practices for model governance and lineage in Vertex AI?
29A retail company uses Vertex AI AutoML to train a product recommendation model. They have a dataset of past purchases stored in BigQuery. The data science team wants to iteratively train and improve the model. They need to track which dataset version was used for each model and preserve the exact data for reproducibility. They currently export data to CSV files and store them in Cloud Storage. However, the dataset is updated daily, and they want to ensure that models are trained on a consistent snapshot. What should they do?
30A data science team is using Vertex AI Pipelines to orchestrate their ML workflows. They want to ensure that each pipeline run is reproducible and that artifacts are versioned. Which Vertex AI feature should they use to track and manage pipeline artifacts?
31You are a machine learning engineer at a retail company. Your team uses Vertex AI Pipelines to train a model that predicts customer churn. The pipeline reads training data from a BigQuery table that is updated daily by an external marketing analytics team. You need to ensure that every pipeline run uses a consistent snapshot of the data and that you can reproduce any past run for auditing. What should you do?
32A team is collaborating on a Vertex AI model using Vertex AI Model Registry. They need to ensure that model versions are properly managed and that deployments are reproducible. Which TWO practices should they follow? (Choose two.)
33Your organization uses Vertex AI Feature Store to serve features for a real-time fraud detection model. Multiple teams contribute features, and you need to ensure that feature values are consistent between training and serving. Which practice should you implement to prevent training-serving skew?
34Your organization uses Vertex AI Model Registry to manage models. A data scientist has trained a new model version and wants to ensure that only approved models are deployed to production. You need to implement a workflow where a model must be reviewed and approved by a designated approver before it can be deployed. What should you do?
35A machine learning team uses Vertex AI Pipelines to orchestrate training workflows. They want to share pipeline runs and artifacts with stakeholders who do not have Google Cloud accounts. What should they do?
36You are collaborating with a team of data scientists on a Vertex AI Workbench notebook that preprocesses data for a machine learning model. You need to ensure that all team members can work on the notebook simultaneously without overwriting each other's changes, and that the notebook's execution environment remains consistent across the team. (Choose two.)
37You are a machine learning engineer working on a team that uses Vertex AI Feature Store. A colleague has created a new feature and wants to make it available to other teams for training and serving. You need to ensure that the feature can be discovered and reused across projects. What should you do?
38Your team uses Vertex AI Pipelines to automate the training and deployment of a recommendation model. The pipeline includes a step that evaluates the model and only deploys it if the evaluation metric exceeds a threshold. You need to ensure that the pipeline's artifacts, including the evaluation metrics and the deployed model, are tracked and can be traced back to the pipeline run for auditing. What should you do?
Deep-dive questions
The most-searched questions in this domain — detailed explanations, worked examples, full answer breakdowns.
Be able to wire an automatic retraining trigger to Vertex AI Pipelines, share Workbench notebooks with Git versioning, and standardize preprocessing as reusable components. The single most important thing: choose the managed Google Cloud service that enforces shared, versioned artifacts rather than manual copying.
The Courseiva PMLE question bank contains 38 questions in the Collaborating to manage data and models domain, covering the 5% of the exam attributed to this domain in the official Google Cloud blueprint. Click any question to see the full explanation and answer breakdown.
Start with a 10-question focused session to identify your baseline accuracy in this domain. Read every explanation — even for questions you answer correctly — to understand the reasoning. Once you score consistently above 80%, move to a 20–30 question session to confirm depth before moving to the next domain.
Yes — the session launcher on this page draws questions exclusively from the Collaborating to manage data and models domain. Choose 10, 20, 30, or 50 questions for a focused session, or click individual questions to review them one by one.
Save your results, see per-domain analytics, and get readiness scores — free, for every certification.
Sign Up FreeFree forever · Every certification included