Courseiva

PMLE · topic practice

Collaborating to manage data and models practice questions

This domain covers how ML teams share code, data, and models on Google Cloud. Expect questions on Vertex AI Pipelines triggers, Vertex AI Workbench notebook sharing and versioning, enforcing shared preprocessing, Cloud Composer collaboration, and Artifact Registry or Model Registry usage for reusable, governed assets across teams.

Courseiva uses original exam-style practice questions designed for learning and revision. The goal is to understand the concepts, recognise exam patterns, and improve through explanations — not memorise copied exam dumps.

Editorial oversight:Johnson Ajibi· MSc IT Security, IEEE Senior Member
20 questionsDomain: Collaborating to manage data and models

What the exam tests

What to know about Collaborating to manage data and models

Be able to wire an automatic retraining trigger to Vertex AI Pipelines, share Workbench notebooks with Git versioning, and standardize preprocessing as reusable components. The single most important thing: choose the managed Google Cloud service that enforces shared, versioned artifacts rather than manual copying.

Triggering Vertex AI Pipelines automatically when new data lands, using Cloud Storage events or Cloud Scheduler

Sharing Vertex AI Workbench notebooks with version history via Git integration in the Workbench instance

Enforcing consistent preprocessing across teams with containerized components or shared Vertex AI Pipelines templates

Using Cloud Composer DAGs, shared repositories, and IAM roles to coordinate multi-team ML workflows

Watch out for

Common Collaborating to manage data and models exam traps

  • ▸Assuming Vertex AI Pipelines retrain themselves on new data; you must configure an explicit trigger such as a Cloud Storage event or scheduled run.
  • ▸Sharing notebooks by copying files instead of using Git-backed version history in Vertex AI Workbench, losing provenance and collaboration.
  • ▸Letting each team write its own preprocessing code rather than packaging shared steps as reusable pipeline components or containers.

Practice set

Collaborating to manage data and models questions

20 questions · select your answer, then reveal the explanation

A data science team uses BigQuery to store raw data and Vertex AI for model training. They want to ensure that only authorized users can access training data, and that model artifacts are automatically versioned and tracked. Which combination of Google Cloud services should they use?

An ML team uses Vertex AI Pipelines to automate model retraining. The pipeline includes a step that queries BigQuery to create a training dataset. The team notices that the pipeline fails intermittently with a '403 Exceeded rate limits' error. What is the most likely cause and solution?

A company stores training data in Cloud Storage and uses Vertex AI Training for model training. They want to implement a data validation pipeline to detect data drift before retraining. Which service should they use?

A team uses Vertex AI Feature Store to serve features for real-time predictions. They notice that feature values are frequently updated from multiple source systems, leading to inconsistencies. They need to ensure that feature values are consistent across all serving endpoints. What should they do?

An organization uses Cloud Composer to orchestrate ML workflows. A DAG that triggers Vertex AI training jobs fails because the training job exceeds the 7-day maximum runtime. What is the best way to handle long-running training jobs in Cloud Composer?

A team wants to share a trained model with other teams within the organization. They need to provide access to the model artifact in Vertex AI Model Registry and ensure that only authorized teams can deploy the model. What should they do?

A data scientist is using Vertex AI Workbench user-managed notebooks. They need to collaborate with a colleague on the same notebook. The colleague should be able to edit the notebook simultaneously. What should they do?

A team uses Vertex AI Pipelines with CustomJob components that pull training code from a Cloud Source Repository. The pipeline fails with a 'Permission denied' error when trying to access the repository. The service account used by the pipeline has the 'Source Repository Viewer' role. What is the likely issue?

Which TWO factors should you consider when choosing between BigQuery and Cloud Storage for storing training data? (Choose 2)

A financial services company uses Vertex AI to deploy multiple models for fraud detection. The ML team has set up a CI/CD pipeline using Cloud Build and Cloud Deploy. The pipeline builds a custom container with the trained model, pushes it to Artifact Registry, and deploys it to a Vertex AI Endpoint. Recently, a new regulation requires that all model deployments be audited and approved by the compliance team before going live. The compliance team wants to review the model's evaluation metrics and approve the deployment via a ticketing system. Currently, the CI/CD pipeline automatically deploys after the container is built. The team needs to implement a gating process without slowing down the development cycle. What should they do?

Match each regularization technique to its effect.

Drag a concept onto its matching description — or click a concept then click the description.

Concepts
Matches

Adds absolute value of weights to loss, induces sparsity

Adds squared magnitude of weights to loss, prevents overfitting

Randomly drops units during training to prevent co-adaptation

Stops training when validation performance stops improving

Increases training data diversity through transformations

Your team is using Vertex AI Feature Store for online predictions. You notice that feature values for some entities are missing in production, leading to failed predictions. Upon investigation, you find that the ingestion pipeline has been failing intermittently. What is the best immediate course of action to prevent prediction failures?

A team of ML engineers is building a real-time fraud detection system. They use Cloud Pub/Sub to stream transactions, Dataflow for feature engineering, and Vertex AI to get predictions. They want to ensure that the data used for training matches the data used for serving to avoid training-serving skew. Which approach should they take?

You are using Cloud Datalab for collaborative data exploration with your team. However, some team members cannot access the Datalab instances. What is the most likely issue?

A company trains models using Vertex AI Training and wants to share the resulting model artifacts with a different team in another Google Cloud project. What is the most secure way to grant access?

A company uses Cloud Composer to orchestrate an ML pipeline. They notice that the pipeline occasionally fails because the Composer environment runs out of disk space on the worker nodes. The pipeline uses many large dependencies. What is the most effective long-term solution?

Your team is using Vertex AI Pipelines to build an automated training pipeline. You need to share the pipeline definition with another team so they can run it in their own project. Which format should you use?

Refer to the exhibit. A user receives the error shown when trying to upload a model to Vertex AI. What is the most likely cause?

Network Topology
gcloud ai models uploadregion=us-central1 \display-name=fraud-detection-v2 \container-image-uri=gcr.io/cloud-aiplatform/prediction/tf2-cpu.2-12:latest \artifact-uri=gs://my-model-artifacts/fraud-detection/v2/ \version-aliases=championRefer to the exhibit.

Refer to the exhibit. A user is trying to upload a Vertex AI pipeline definition. The error indicates an invalid dependency order. What should the user do to fix this?

Exhibit

Refer to the exhibit.

# pipeline.yaml
pipeline:
  name: training-pipeline
  description: End-to-end ML pipeline
  params:
    project_id: {type: String}
    dataset_id: {type: String}
  tasks:
    - task1:
        component: preprocessing
        inputs:
          project_id: {inputValue: project_id}
          dataset_id: {inputValue: dataset_id}
    - task2:
        component: training
        inputs:
          data: {taskOutputs: task1.output}
        dependentTasks: [task1]

Error: (gsutil cp pipeline.yaml gs://my-bucket/pipelines/): RuntimeException: Failed to compile pipeline. Invalid pipeline definition: task 'task2' depends on 'task1' but 'task1' is defined after 'task2' in YAML ordering.

Refer to the exhibit. An ML engineer in the team needs to deploy the model to an endpoint. The engineer is assigned the 'roles/aiplatform.user' role at the project level but still cannot deploy. What is the most likely reason?

Exhibit

Refer to the exhibit.

{
  "bindings": [
    {
      "role": "roles/aiplatform.user",
      "members": [
        "user:alice@example.com",
        "serviceAccount:sa-training@my-project.iam.gserviceaccount.com"
      ]
    }
  ]
}

This IAM policy is attached to a Vertex AI model resource. Alice can view the model but cannot deploy it to an endpoint. The service account can use the model for training.

Free account

Track your progress over time

Create a free account to save your results and see which topics improve across sessions.

Focused Collaborating to manage data and models sessions

Start a Collaborating to manage data and models only practice session

Every question in these sessions is drawn from the Collaborating to manage data and models domain — nothing else.

Related practice questions

Related PMLE topic practice pages

Move into related areas when this topic feels solid.

Frequently asked questions

What does the PMLE exam test about Collaborating to manage data and models?
Be able to wire an automatic retraining trigger to Vertex AI Pipelines, share Workbench notebooks with Git versioning, and standardize preprocessing as reusable components. The single most important thing: choose the managed Google Cloud service that enforces shared, versioned artifacts rather than manual copying.
How should I use these practice questions?
Select your answer before revealing the explanation. Then read why each option is right or wrong — this active recall approach builds retention far faster than re-reading notes.
Can I practise just Collaborating to manage data and models questions in a focused session?
Yes — the session launcher on this page draws every question from the Collaborating to manage data and models domain. Use a 10-question session first to gauge your baseline, then move to 20 or 30 once the weak spots are clear.
Where can I practise other PMLE topics?
Use the topic links above to move to related areas, or go back to the PMLE question bank to see all topics.
Are these real exam questions or dumps?
These are original practice questions written to test the same concepts the PMLE exam covers. They are not copied from any real exam or dump site.