Courseiva

PMLE · domain

Collaborating to manage data and models

This domain covers how ML teams share code, data, and models on Google Cloud. Expect questions on Vertex AI Pipelines triggers, Vertex AI Workbench notebook sharing and versioning, enforcing shared preprocessing, Cloud Composer collaboration, and Artifact Registry or Model Registry usage for reusable, governed assets across teams.

38 questions12 easy13 medium13 hard

Focused practice

Practice Collaborating to manage data and models questions

Scored sessions drawing only from this domain — pick a length below.

Start 20-question practice test →

What this domain covers

What to know about Collaborating to manage data and models

Be able to wire an automatic retraining trigger to Vertex AI Pipelines, share Workbench notebooks with Git versioning, and standardize preprocessing as reusable components. The single most important thing: choose the managed Google Cloud service that enforces shared, versioned artifacts rather than manual copying.

Triggering Vertex AI Pipelines automatically when new data lands, using Cloud Storage events or Cloud Scheduler

Sharing Vertex AI Workbench notebooks with version history via Git integration in the Workbench instance

Enforcing consistent preprocessing across teams with containerized components or shared Vertex AI Pipelines templates

Using Cloud Composer DAGs, shared repositories, and IAM roles to coordinate multi-team ML workflows

Watch out for

Common Collaborating to manage data and models exam traps

  • ▸Assuming Vertex AI Pipelines retrain themselves on new data; you must configure an explicit trigger such as a Cloud Storage event or scheduled run.
  • ▸Sharing notebooks by copying files instead of using Git-backed version history in Vertex AI Workbench, losing provenance and collaboration.
  • ▸Letting each team write its own preprocessing code rather than packaging shared steps as reusable pipeline components or containers.

Question index

All Collaborating to manage data and models questions (38)

Click any question to see the full explanation, or start a practice session above.

1

An organization uses Cloud Dataflow to preprocess training data. Dataflow jobs are often failing because of insufficient quota for certain resources. The team has requested a quota increase, but the jobs still fail with 'quota exceeded' errors for a different resource. They want to proactively monitor and manage quotas to avoid failures. What is the best approach?

Hard
2

You are a machine learning engineer working on a team that uses Vertex AI Feature Store. A colleague has created a new feature and wants to make it available to other teams for training and serving. You need to ensure that the feature can be discovered and reused across projects. What should you do?

Easy
3

A retail company uses Vertex AI AutoML to train a product recommendation model. They have a dataset of past purchases stored in BigQuery. The data science team wants to iteratively train and improve the model. They need to track which dataset version was used for each model and preserve the exact data for reproducibility. They currently export data to CSV files and store them in Cloud Storage. However, the dataset is updated daily, and they want to ensure that models are trained on a consistent snapshot. What should they do?

Easy
4

Which TWO strategies help ensure data consistency when multiple teams are contributing features to a shared Vertex AI Feature Store?

Hard
5

Drag and drop the steps to deploy a trained TensorFlow model to Vertex AI Prediction in the correct order.

Medium
6

Your organization uses Vertex AI Feature Store to serve features for a real-time fraud detection model. Multiple teams contribute features, and you need to ensure that feature values are consistent between training and serving. Which practice should you implement to prevent training-serving skew?

Hard
7

A data engineering team uses Dataflow for preprocessing and wants to integrate with Vertex AI Pipelines. They need to pass the preprocessed data location to the training step. What is the best practice?

Hard
8

Refer to the exhibit. The team wants to automatically deploy the best-performing model version to production. They have set up a Cloud Function triggered by Model Registry events. Which alias should they use in the function to get the latest champion?

Hard
9

A data scientist wants to share a trained model with the team for review before deployment. The model is stored in Vertex AI Model Registry. What is the recommended way to grant the team read access to the model?

Easy
10

You are collaborating with a team of data scientists on a Vertex AI Workbench notebook that preprocesses data for a machine learning model. You need to ensure that all team members can work on the notebook simultaneously without overwriting each other's changes, and that the notebook's execution environment remains consistent across the team. (Choose two.)

Medium
11

A large organization uses a multi-project setup with a central data lake. Different teams manage their own models. To enable cross-team sharing of features, they want to use Vertex AI Feature Store. What is the best practice to manage access?

Hard
12

You are a machine learning engineer at a retail company. Your team uses Vertex AI Pipelines to train a model that predicts customer churn. The pipeline reads training data from a BigQuery table that is updated daily by an external marketing analytics team. You need to ensure that every pipeline run uses a consistent snapshot of the data and that you can reproduce any past run for auditing. What should you do?

Medium
13

A data science team uses Vertex AI Workbench and wants to share notebooks with version history. Which service should they use?

Easy
14

A machine learning team uses Vertex AI Pipelines to orchestrate training workflows. They want to share pipeline runs and artifacts with stakeholders who do not have Google Cloud accounts. What should they do?

Easy
15

Which THREE actions are best practices for managing ML models in production on Google Cloud? (Choose 3)

Medium
16

A team is using Vertex AI Experiments to compare different hyperparameters. They want to automatically record the hyperparameters. What is the correct way?

Medium
17

A company uses a Cloud Composer DAG to run a daily ML pipeline that includes Dataflow jobs and model training on Vertex AI. The pipeline frequently fails due to insufficient permissions when the Dataflow worker accesses data in Cloud Storage. What is the most efficient way to resolve this issue?

Hard
18

A team uses Vertex AI Feature Store for storing features. They want to share feature definitions with other teams in a collaborative manner. What is the best way to collaborate on feature definitions?

Easy
19

A data scientist wants to track the lineage of a dataset used in a training run. Which Vertex AI feature should they use?

Easy
20

A company has multiple teams working on different models. They want to enforce consistent data preprocessing steps across all teams. Which approach should they take?

Hard
21

When distributing training across multiple workers using Vertex AI Training, how should the team share the training dataset?

Easy
22

Which TWO actions are recommended for collaborating on machine learning models using Vertex AI Model Registry?

Medium
23

Which THREE practices improve collaboration when using Cloud Composer for ML pipelines?

Medium
24

A team is collaborating on a Vertex AI model using Vertex AI Model Registry. They need to ensure that model versions are properly managed and that deployments are reproducible. Which TWO practices should they follow? (Choose two.)

Hard
25

A data science team is using Vertex AI Pipelines to orchestrate their ML workflows. They want to ensure that each pipeline run is reproducible and that artifacts are versioned. Which Vertex AI feature should they use to track and manage pipeline artifacts?

Easy
26

Which THREE actions should be taken to manage model versions effectively?

Hard
27

A team of ML engineers is collaborating on a project using Vertex AI. They want to ensure that only approved models are deployed to production. Which approach should they use?

Medium
28

Which TWO of the following are best practices for managing data in a collaborative machine learning environment on Google Cloud?

Medium
29

An MLOps team needs to automatically retrain a model when new training data becomes available. They use Vertex AI Pipelines. What is the recommended way to trigger the pipeline?

Medium
30

Which TWO statements about Vertex AI Feature Store are correct? (Choose 2)

Easy
31

A healthcare organization is building a machine learning model to predict patient readmission risk. They have sensitive data stored in BigQuery that includes protected health information (PHI). The data science team uses Vertex AI Workbench notebooks to explore the data and develop models. The organization's security policy requires that all PHI data must be encrypted at rest and in transit, and that access to the data is logged and audited. They also need to ensure that the data used for model training is de-identified to remove direct identifiers such as patient names and SSNs. The team wants to automate the de-identification process as part of the data pipeline. Which approach meets these requirements?

Medium
32

Refer to the exhibit. The team notices that the pipeline fails to read data from the specified Cloud Storage path. What is the most likely issue?

Easy
33

Which THREE of the following are recommended practices for model governance and lineage in Vertex AI?

Hard
34

Your team uses Vertex AI Pipelines to automate the training and deployment of a recommendation model. The pipeline includes a step that evaluates the model and only deploys it if the evaluation metric exceeds a threshold. You need to ensure that the pipeline's artifacts, including the evaluation metrics and the deployed model, are tracked and can be traced back to the pipeline run for auditing. What should you do?

Hard
35

A team uses Vertex AI Pipelines. They need to ensure that only certain team members can deploy models to production. What is the best approach?

Medium
36

Your organization uses Vertex AI Model Registry to manage models. A data scientist has trained a new model version and wants to ensure that only approved models are deployed to production. You need to implement a workflow where a model must be reviewed and approved by a designated approver before it can be deployed. What should you do?

Hard
37

Which TWO practices help ensure reproducible ML experiments?

Easy
38

A company uses BigQuery to store feature data for ML training. A data engineer notices that a Vertex AI Training job is failing with 'Access Denied' errors when reading from a BigQuery table. The training job uses a custom service account that has been granted the 'bigquery.dataViewer' role on the dataset. What is the most likely cause of the failure?

Medium

Frequently asked questions

What does the Collaborating to manage data and models domain cover on the PMLE exam?
Be able to wire an automatic retraining trigger to Vertex AI Pipelines, share Workbench notebooks with Git versioning, and standardize preprocessing as reusable components. The single most important thing: choose the managed Google Cloud service that enforces shared, versioned artifacts rather than manual copying.
How many questions are in this domain?
This page lists all 38 Collaborating to manage data and models questions in the PMLE question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
What is the best way to practise this domain?
Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
Can I practise only Collaborating to manage data and models questions?
Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.
google-pmle GOOGLE-PMLE data model mgmt Practice Questions