PMLE · domain
Collaborating Within and Across Teams to Manage Data and Models
This domain covers how ML work is shared, reproduced, and governed across teams on Google Cloud. It is tested through scenario questions that ask you to pick the right Vertex AI or data service for tracking runs, versioning datasets, monitoring feature drift, and tracing lineage across pipeline executions.
Focused practice
Practice Collaborating Within and Across Teams to Manage Data and Models questions
Scored sessions drawing only from this domain — pick a length below.
Start 20-question practice test →What this domain covers
What to know about Collaborating Within and Across Teams to Manage Data and Models
Be able to choose the correct Google Cloud service for tracking, versioning, lineage, and monitoring, and explain why it fits the scenario. The most important thing is distinguishing Vertex AI Experiments from Vertex ML Metadata and knowing when ACID transactions are required.
Selecting Vertex AI Experiments to log hyperparameters, metrics, and artifacts with minimal code changes
Using Vertex ML Metadata to track lineage of datasets, parameters, and models across pipeline runs
Configuring Vertex AI Feature Store drift and skew monitoring with alerting thresholds
Applying BigQuery or Cloud Storage versioning and ACID guarantees for concurrent dataset access
Watch out for
Common Collaborating Within and Across Teams to Manage Data and Models exam traps
- ▸Confusing Vertex AI Experiments with Vertex ML Metadata: Experiments logs runs and metrics, while ML Metadata stores lineage and artifacts
- ▸Assuming Feature Store drift detection is automatic without configuring a monitoring schedule, baseline, and alert threshold
- ▸Ignoring that concurrent multi-job writes need ACID transactions, and choosing plain Cloud Storage objects instead of BigQuery or a transactional layer
Question index
All Collaborating Within and Across Teams to Manage Data and Models questions (75)
Click any question to see the full explanation, or start a practice session above.
A team wants to track the lineage of ML pipeline runs, including which datasets, parameters, and models were used in each execution. Which Vertex AI service should they use?
Easy2An ML team uses Delta Lake on Dataproc for data versioning. Which THREE benefits does Delta Lake provide?
Medium3A data scientist finishes training a model in a Vertex AI Workbench notebook and wants to save it so that the deployment team can later deploy it to an endpoint without re-running the notebook. The deployment team needs to see the model's version history and assign a production alias. Which action should the data scientist take?
Easy4A company uses Vertex AI Model Registry to manage multiple model versions. They want to designate a model version as 'champion' for production deployment and another as 'challenger' for A/B testing. Which feature of the registry should they use?
Easy5A team wants to share feature definitions across multiple projects in their organization using Vertex AI Feature Store. What is the recommended approach?
Medium6A machine learning pipeline in Vertex AI produces a dataset artifact, a trained model, and evaluation metrics. The team wants to query the lineage to find all downstream artifacts that depend on a particular dataset. Which Vertex AI service should they use?
Hard7An ML team wants to automatically track training runs, including hyperparameters and metrics, with minimal code changes. Which Vertex AI service should they use?
Easy8An organization uses Vertex AI Pipelines and wants to track the lineage of datasets, models, and metrics across pipeline runs. They need to query upstream and downstream dependencies of an artifact. Which service should they use?
Medium9An ML team uses Vertex AI Pipelines to train and evaluate models. They want to ensure that only models meeting a minimum accuracy threshold are registered in Vertex AI Model Registry. Which approach should they take?
Medium10A company wants to implement a centralized model registry for governance. Which two features should they use? (Choose two.)
Medium11A company uses Vertex AI Feature Store with an online store for low-latency serving. They observe high latency during peak hours. The feature values are small (< 1 KB each) and the workload is read-heavy. Which change would most effectively reduce latency?
Hard12A junior engineer on your team has a trained scikit-learn model saved as a local joblib file and wants other teams to be able to discover it, view its evaluation metrics, and deploy it to a Vertex AI Endpoint. Which action should they take first?
Easy13A team wants to use Vertex AI Workbench for collaborative notebook development. They need a persistent environment that can be stopped and restarted without losing installed packages and data. Which instance type should they choose?
Easy14A machine learning team wants to share features across multiple models to reduce training-serving skew and ensure consistency. Which Vertex AI service should they use?
Easy15A machine learning team needs to ensure that the same features used for training are used for serving in production to avoid training-serving skew. They use Vertex AI Feature Store. Which THREE actions should they take?
Hard16A data science team needs to share features across multiple ML models while ensuring consistency between training and serving. Which approach best achieves this?
Medium17A company wants to use DVC for data versioning alongside their ML code in Git. Which TWO statements about DVC are correct? (Select 2)
Easy18A regulated enterprise must prove to auditors that a specific production prediction can be traced to the exact model version, training data, and pipeline execution that produced it. Their ML workflows run on Vertex AI Pipelines. Which two practices should the team adopt? (Choose two.)
Medium19A team is training a model using historical data and wants to avoid data leakage when joining feature values from a feature store. The features include time-varying data like user activity counts. Which retrieval method should they use when creating a training dataset?
Hard20A company wants to monitor features in Vertex AI Feature Store for drift over time. Which two services should they use? (Choose two.)
Medium21An ML team wants to share feature definitions across multiple projects to reduce training-serving skew and ensure consistency. They currently store features in Cloud Storage and manually coordinate updates, leading to errors. Which Google Cloud service should they use to centrally manage and serve features for both training and online inference?
Medium22A team uses Vertex AI Feature Store with an online store for low-latency serving. They need to support frequent updates to features (e.g., every minute) and require high write throughput (thousands of writes per second). Which online store type should they choose?
Hard23An ML team wants to monitor feature drift in their production model. Which Vertex AI Feature Store capability should they use?
Easy24An ML team uses Vertex AI Pipelines and wants to automatically generate model cards documenting model purpose, evaluation results, and intended use. Which approach should they take?
Easy25Two teams train models in separate Vertex AI projects but must share the same curated feature set. The platform team wants a single authoritative definition of each feature so that online serving and offline training always return consistent values, while each team keeps its own model training pipeline. Which approach should the platform team implement?
Hard26A data scientist wants to automatically generate model documentation that includes model purpose, training data, evaluation results, and intended use. Which tool should they use?
Easy27Your team trains models in a shared Vertex AI project. A data engineer accidentally overwrites a BigQuery training table that three production pipelines depend on, and nobody can tell which pipeline used which version of the data. You need to make dataset versions immutable and traceable so that any training run can be reproduced. What should you do?
Medium28A data scientist wants to track machine learning experiments, including parameters, metrics, and artifacts, and compare runs. Which Vertex AI service should they use?
Easy29A data science team wants to share a set of engineered features across multiple projects and teams to reduce training-serving skew and ensure consistency. They need low-latency serving (single-digit milliseconds) for online predictions and also need to retrieve historical feature values for training. Which approach should they take?
Medium30Which Vertex AI service is used to track the lineage of ML pipeline components, artefacts, and executions?
Easy31A data engineer needs to version large datasets (multiple TB) in a Data Lake on Google Cloud. They require ACID transactions to ensure consistency when multiple jobs read/write concurrently. Which solution should they use?
Medium32A data science team collaborates using Vertex AI Workbench user-managed notebooks. They want to version control their notebook code and share it with team members. Which TWO tools should they use? (Choose 2)
Medium33Your team trains a scikit-learn model locally and uploads it to Vertex AI Model Registry. A colleague needs to deploy it to a Vertex AI Endpoint for online prediction with a prebuilt container. The model artifacts are stored in a Cloud Storage bucket. Which deployment approach should they use?
Medium34A team uses Vertex AI Pipelines with a custom training component that reads data from a BigQuery table. They need to ensure that a new pipeline run uses a specific snapshot of the training data for reproducibility. Which approach should they take?
Medium35Your team is preparing to hand a trained model to a separate operations team that will deploy it to a Vertex AI endpoint. The operations team needs to understand the model's input schema, the training run that produced it, and which alias currently points to production. Which two Vertex AI resources should you share with them to provide this information? (Choose two.)
Medium36You are using DVC for data versioning in an ML project on Google Cloud. Your training data is stored in Cloud Storage. You want to track a new version of the dataset after preprocessing. Which DVC command should you use to register the changes?
Medium37A data science team uses Vertex AI Experiments to compare multiple model training runs. They want to capture and compare hyperparameters, metrics, and code versions for each run. Which TWO steps should they take?
Medium38An ML engineer needs to deploy a model to an endpoint and gradually shift traffic from the previous version (champion) to a new version (challenger) for A/B testing. How should they configure the endpoint?
Medium39A company uses Vertex AI Feature Store for feature engineering. They need to ensure point-in-time correctness to avoid data leakage during training. Which feature retrieval method should they use?
Hard40A team monitors features in Vertex AI Feature Store for drift. They want to set up automated alerts when a feature's distribution deviates significantly from the baseline. Which feature monitoring configuration should they use?
Hard41Your team owns a Vertex AI Model Registry entry that other teams depend on for production serving. A new retrained model shows better offline metrics, and you need to roll it out gradually to a small percentage of live traffic while keeping the ability to revert instantly if quality degrades. What should you do?
Medium42A team is using Vertex AI Model Registry to manage models. They need to ensure that when a new model version is registered, it is automatically evaluated for fairness and bias before being deployed. Which two Google Cloud services should they integrate to achieve this? (Choose two.)
Medium43You need to create a reproducible snapshot of a BigQuery table as of a specific timestamp for ML model training. The snapshot should be queryable without copying the entire dataset. Which BigQuery feature should you use?
Hard44You are setting up feature monitoring in Vertex AI Feature Store to detect drift in a numerical feature. The monitoring job should run daily and alert if the Jensen-Shannon divergence exceeds 0.1. Which configuration should you use?
Medium45A team uses Vertex AI Workbench managed notebooks. They want to version control their notebook files and collaborate using Git. What is the best way to integrate Git?
Medium46Your team uses a Vertex AI Pipeline that reads from a BigQuery table, trains a model, and registers it. A teammate wants to know which BigQuery table snapshot was used for a specific registered model version so they can reproduce the training data exactly. Which action should they take?
Hard47A data scientist needs to retrieve training data from Vertex AI Feature Store that exactly matches the feature values as they were at a specific historical timestamp to avoid label leakage. Which feature view configuration should they use?
Medium48A team is using Delta Lake on Dataproc for their data lake with ACID transactions. They want to version data for ML experiments and roll back to a previous version if needed. Which Delta Lake feature should they use?
Medium49What is the primary benefit of using a centralised model registry in MLOps?
Easy50A data science team wants to version control their datasets along with code using Git. They need a tool that integrates with Git and tracks changes to large data files. Which tool should they use?
Medium51An organization needs to implement MLOps with standardized pipeline templates across multiple teams. Which Vertex AI feature should they use to create reusable pipeline components?
Hard52A team uses Vertex AI Workbench notebooks for collaborative model development. They want to ensure that code changes are version-controlled, that multiple data scientists can work on the same notebook without conflicts, and that the environment is reproducible across team members. Which approach should they take?
Medium53An ML team wants to implement data versioning for large datasets stored in Google Cloud Storage. They need to track changes over time and reproduce previous data states. Which tool is most appropriate?
Medium54A team is building a fraud detection model that requires joining real-time transaction features with historical user features. They need to ensure that the training data does not use future information (data leakage). Which Vertex AI Feature Store capability should they use?
Medium55A team uses Vertex AI Pipelines and wants to track lineage of artifacts and executions. Which three resources should they use? (Choose three.)
Hard56A fraud detection team trains a model on data stored in BigQuery. They want to ensure that the model can be reproduced exactly one year later, including the specific data version and training code. They use Vertex AI Pipelines for orchestration. Which practice should they implement?
Medium57An ML engineer trained a model and registered it in Vertex AI Model Registry. They want to assign the alias 'champion' to the best-performing version for production deployment. Which gcloud command should they use?
Hard58A company uses Vertex AI Pipelines to orchestrate ML workflows. After a pipeline run, they want to query the lineage of a particular model artifact to find out which dataset and hyperparameters were used to produce it. Which API method should they use?
Hard59A team is operationalizing a machine learning pipeline using Vertex AI. They want to automatically track experiment runs, log model parameters and metrics, and store model artifacts for reproducibility. They also need to capture lineage between pipeline components (e.g., which dataset and hyperparameter tuning job produced a model). Which TWO services should they use together to achieve this? (Choose two.)
Hard60An ML engineer has a model trained in Vertex AI and wants to deploy it to an endpoint with autoscaling and traffic splitting for canary testing. They have the model artifact stored in Vertex AI Model Registry with alias 'champion'. What is the correct sequence of steps?
Medium61A team is building ML pipelines with Vertex AI. They want to reuse standard pipeline components across teams and enforce governance. What approach should they take?
Medium62Your organization uses Vertex AI Pipelines for training. A compliance auditor asks you to prove which dataset version and which preprocessing code commit produced a model that is currently deployed. You need to retrieve this information programmatically for a specific model version. Which approach should you use?
Hard63A company trains a model using features from Vertex AI Feature Store. They notice training-serving skew because the feature values used at training time differ from those served online. How should they address this?
Hard64Your team trains a model on a Vertex AI Workbench notebook and logs hyperparameters, metrics, and a confusion matrix. Your manager asks you to ensure that anyone in the organization can reproduce the exact training run and compare it with other runs without manually digging through notebook cells. Which Vertex AI component should you use to record this information?
Medium65You are collaborating on a Vertex AI Feature Store implementation. A data engineer updates a feature's values in the offline store, but the online store still serves the old values for several hours. The online store is configured with a feature value TTL of 24 hours and uses batch ingestion. What is the most likely cause of the stale online values?
Hard66A company wants to implement a central model governance strategy using Vertex AI. They need to track model lineage, store evaluation metrics, and manage model versions across teams. Which THREE Vertex AI services should they use? (Choose 3)
Medium67A team uses Vertex AI Metadata to track pipeline runs. They need to identify all artifacts that were generated by a particular pipeline execution. Which API method should they use?
Hard68Your team trains models in a shared Vertex AI project, and multiple engineers run pipelines against the same BigQuery training tables. A reviewer needs to reproduce the exact dataset used to train a model six weeks ago, but the source tables have been overwritten many times since. Which BigQuery capability should you have used to make each training snapshot reproducible?
Medium69An organisation uses Delta Lake on Dataproc to manage a data lake for ML training. They need ACID transactions for concurrent reads and writes. Which file format does Delta Lake use as the underlying storage?
Medium70A team wants to enforce governance and compliance for all ML models across the organisation. They need a centralised repository that tracks model versions, deployment history, and evaluation metrics. Which service should they use?
Easy71A company uses Vertex AI Pipelines to train and deploy models. They want to automatically generate model documentation that includes model details, intended use, and evaluation results. What should they use?
Medium72A data science team wants to share engineered features across multiple projects while ensuring low-latency serving for online predictions. Which Google Cloud service should they use to store and serve these features?
Easy73An organization wants to implement central governance for ML models across teams. Which TWO services should they use together to achieve model versioning, lineage, and deployment management? (Select 2)
Medium74An ML team trains a model using a dataset stored in a BigQuery table. They want to ensure that the exact data snapshot used for training is recorded and can be reproduced later for auditing. Which approach should they take?
Medium75A team wants to implement automated model documentation that captures training data, feature importance, evaluation metrics, and intended use. Which Vertex AI feature supports this?
HardOther domains
All PMLE exam domains
Frequently asked questions
- What does the Collaborating Within and Across Teams to Manage Data and Models domain cover on the PMLE exam?
- Be able to choose the correct Google Cloud service for tracking, versioning, lineage, and monitoring, and explain why it fits the scenario. The most important thing is distinguishing Vertex AI Experiments from Vertex ML Metadata and knowing when ACID transactions are required.
- How many questions are in this domain?
- This page lists all 75 Collaborating Within and Across Teams to Manage Data and Models questions in the PMLE question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
- What is the best way to practise this domain?
- Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
- Can I practise only Collaborating Within and Across Teams to Manage Data and Models questions?
- Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.