Courseiva

CCNA Data Model Mgmt Questions

38 questions · Data Model Mgmt topic · All types, answers revealed

1
MCQhard

An organization uses Cloud Dataflow to preprocess training data. Dataflow jobs are often failing because of insufficient quota for certain resources. The team has requested a quota increase, but the jobs still fail with 'quota exceeded' errors for a different resource. They want to proactively monitor and manage quotas to avoid failures. What is the best approach?

A.Set up Cloud Monitoring alerts for quota usage and automate quota increase requests.
B.Configure Dataflow to use a different pipeline type that avoids the quota.
C.Use Dataflow's autoscaling feature to reduce resource usage.
D.Increase the maximum number of workers in the Dataflow job.
AnswerA

Proactive monitoring and automation allow scaling quotas as needed.

Why this answer

The best approach is to set up Cloud Monitoring alerts for quota usage and automate quota increase requests. This provides proactive visibility into all resource quotas (not just the one initially increased) and enables automated remediation before jobs fail. Cloud Monitoring can track quota metrics for services like Compute Engine, and you can use Cloud Functions or Pub/Sub to trigger quota increase requests via the Service Usage API.

Exam trap

PMLE often tests the difference between reactive fixes (increasing workers) and proactive monitoring, and candidates may choose autoscaling or pipeline changes instead of addressing quota management directly.

How to eliminate wrong answers

Option B is wrong because changing the pipeline type does not address the root cause — quota limits apply regardless of pipeline type, and you may still hit them. Option C is wrong because autoscaling reduces resource usage but does not eliminate the need to monitor and manage quotas; it may even mask the problem temporarily. Option D is wrong because increasing the maximum number of workers would consume more quota, potentially worsening the issue, and does not provide proactive monitoring.

2
MCQeasy

You are a machine learning engineer working on a team that uses Vertex AI Feature Store. A colleague has created a new feature and wants to make it available to other teams for training and serving. You need to ensure that the feature can be discovered and reused across projects. What should you do?

A.Use Vertex AI Feature Store's built-in feature sharing capability by setting the feature's visibility to 'Public' within the organization.
B.Create the feature in Vertex AI Feature Store and use Dataplex to catalog it, making it searchable and accessible to other teams.
C.Create the feature in a Vertex AI Feature Store instance and share the instance's project ID with other teams.
D.Register the feature in Vertex AI Feature Store and assign it a label that other teams can search for in the Vertex AI console.
AnswerB

Dataplex provides a centralized data catalog that can include Vertex AI Feature Store features. By cataloging the feature, other teams can discover it through search, understand its metadata, and request access. This promotes reuse and collaboration across projects while maintaining proper governance.

Why this answer

Using Dataplex to catalog Vertex AI Feature Store features enables centralized discovery and metadata management. Other teams can search the catalog, find the feature, and request access, which fosters reuse and collaboration. This approach integrates with existing governance and access controls, unlike ad-hoc sharing methods.

Exam trap

The trap here is believing that Vertex AI Feature Store has a native public sharing or cross-project visibility toggle, when in fact you need an external catalog like Dataplex.

3
MCQeasy

A retail company uses Vertex AI AutoML to train a product recommendation model. They have a dataset of past purchases stored in BigQuery. The data science team wants to iteratively train and improve the model. They need to track which dataset version was used for each model and preserve the exact data for reproducibility. They currently export data to CSV files and store them in Cloud Storage. However, the dataset is updated daily, and they want to ensure that models are trained on a consistent snapshot. What should they do?

A.Use Vertex AI Dataset service to create a dataset and export it to BigQuery.
B.Use BigQuery snapshots to capture a versioned dataset and reference the snapshot in the training pipeline.
C.Train the model directly on the BigQuery table and let AutoML handle versioning.
D.Export the data to a timestamped CSV file and store it in Cloud Storage before each training run.
AnswerB

BigQuery snapshots preserve table data at a point in time and are immutable, so each training run references a fixed snapshot rather than the daily-changing source table. This guarantees a consistent dataset version and reproducibility without manual CSV exports.

Why this answer

BigQuery snapshots provide a consistent, versioned view of the dataset at a specific point in time, ensuring reproducibility without duplicating data. By referencing the snapshot in the Vertex AI training pipeline, the team can train models on the exact same data snapshot, even as the source table is updated daily. This approach avoids the overhead of exporting to CSV and Cloud Storage while maintaining data integrity and lineage.

Exam trap

Google Cloud often tests the misconception that exporting to CSV or using Vertex AI Dataset is sufficient for versioning, when in fact BigQuery snapshots provide the native, scalable, and auditable mechanism for point-in-time data consistency without data duplication.

How to eliminate wrong answers

Option A is wrong because the Vertex AI Dataset service is designed for managing training data within Vertex AI, but exporting to BigQuery does not inherently create a versioned snapshot; it simply moves data back to BigQuery without preserving a consistent point-in-time copy. Option C is wrong because training directly on a live BigQuery table does not guarantee a consistent snapshot; AutoML does not handle versioning, and the table may change between training runs, breaking reproducibility. Option D is wrong because exporting to a timestamped CSV file in Cloud Storage is a manual workaround that introduces storage overhead, potential data drift from export timing, and lacks the built-in versioning and query capabilities of BigQuery snapshots.

4
Multi-Selecthard

Which TWO strategies help ensure data consistency when multiple teams are contributing features to a shared Vertex AI Feature Store?

Select 2 answers
A.Each team should create their own feature store to avoid conflicts.
B.Use only batch ingestion to keep features synchronized.
C.Define and enforce feature schemas using the Feature Store API.
D.Allow each team to independently define feature engineering logic.
E.Set up monitoring and alerting on feature value distributions to detect drift.
AnswersC, E

Schemas ensure consistent data types and values.

Why this answer

Defining and enforcing feature schemas using the Vertex AI Feature Store API ensures that all teams adhere to a consistent data structure (e.g., fixed feature names, data types, and value ranges). This prevents schema drift and ingestion conflicts, which are common when multiple teams independently push features to the same feature store. Without schema enforcement, one team might inadvertently change a feature's data type or add unexpected values, breaking downstream models.

Exam trap

Google Cloud often tests the misconception that 'separate stores' or 'batch-only ingestion' are valid consistency strategies, when in fact the correct approach is centralized schema governance with monitoring to detect drift.

5
Drag & Dropmedium

Drag and drop the steps to deploy a trained TensorFlow model to Vertex AI Prediction in the correct order.

Drag or tap steps into the slots.

Steps
Order
1Step 1
2Step 2
3Step 3
4Step 4

Why this order

Export the model, upload to GCS, register as a model, deploy to endpoint, then test.

6
MCQhard

Your organization uses Vertex AI Feature Store to serve features for a real-time fraud detection model. Multiple teams contribute features, and you need to ensure that feature values are consistent between training and serving. Which practice should you implement to prevent training-serving skew?

A.Implement a separate feature engineering pipeline for training and another for serving to optimize each for its specific needs.
B.Export features from the featurestore to BigQuery for training, and use the featurestore online serving for predictions.
C.Use a single featurestore for both training and serving, and ingest features via the same pipeline.
D.Use Vertex AI Feature Store's built-in monitoring to detect skew, and retrain the model when skew is detected.
AnswerC

Using a single featurestore and a unified ingestion pipeline ensures that the same transformation logic and data sources are used for both training and serving, eliminating discrepancies. This practice directly addresses training-serving skew by maintaining consistency in feature computation and storage, which is essential for model reliability.

Why this answer

Training-serving skew arises when feature values differ between training and serving due to disparate data processing. Using a single featurestore and a unified ingestion pipeline ensures that features are computed and stored consistently, so the model sees identical feature distributions in both phases. This is a fundamental best practice in ML engineering.

Exam trap

The trap here is thinking that monitoring and retraining can solve training-serving skew, but they only detect and react to it; the only way to prevent it is to ensure identical feature computation for training and serving.

7
MCQhard

A data engineering team uses Dataflow for preprocessing and wants to integrate with Vertex AI Pipelines. They need to pass the preprocessed data location to the training step. What is the best practice?

A.Store the path in Data Catalog
B.Use Cloud Pub/Sub
C.Use PipelineParam to pass the output path
D.Write the output to a fixed Cloud Storage path and hardcode it in the pipeline
AnswerC

PipelineParam passes the Dataflow output path as a runtime parameter between pipeline components, satisfying the requirement to hand the preprocessed data location to the training step. Vertex AI Pipelines resolves this value at execution, decoupling the training component from hard-coded paths and enabling reuse across runs.

Why this answer

PipelineParam is the native mechanism in Vertex AI Pipelines (Kubeflow Pipelines SDK) to pass runtime outputs—such as a Cloud Storage path—between components. It creates a dependency graph that ensures the training step receives the exact output path from the preprocessing step, enabling dynamic, reproducible pipelines without hardcoding.

Exam trap

The trap here is that candidates confuse metadata services (Data Catalog) or messaging systems (Pub/Sub) with pipeline parameter passing, overlooking that Vertex AI Pipelines uses Kubeflow Pipelines' built-in component I/O for deterministic, graph-based data flow.

How to eliminate wrong answers

Option A is wrong because Data Catalog is a metadata management service for discovering and tagging assets, not designed to pass runtime pipeline parameters between steps; it would introduce unnecessary latency and coupling. Option B is wrong because Cloud Pub/Sub is an asynchronous messaging service for event-driven architectures, not a direct parameter-passing mechanism within a single pipeline execution; it would add complexity and potential ordering issues. Option D is wrong because hardcoding a fixed Cloud Storage path defeats pipeline reproducibility and scalability—if the preprocessing step changes its output location (e.g., due to timestamped folders), the training step would fail or use stale data.

8
MCQhard

Refer to the exhibit. The team wants to automatically deploy the best-performing model version to production. They have set up a Cloud Function triggered by Model Registry events. Which alias should they use in the function to get the latest champion?

A.'champion'
B.''
C.'experiment'
D.'latest'
AnswerA

The 'champion' alias conventionally indicates the best-performing production version.

Why this answer

The 'champion' alias is specifically reserved in Vertex AI Model Registry to denote the best-performing model version in production. By configuring the Cloud Function to trigger on the assignment of the 'champion' alias, the team ensures that only the model version promoted as the production champion is automatically deployed, aligning with MLOps best practices for staged model promotion.

Exam trap

Google Cloud often tests the distinction between 'champion' (a production alias) and 'latest' (a version number concept), leading candidates to incorrectly choose 'latest' because they confuse chronological recency with performance-based promotion.

How to eliminate wrong answers

Option B is wrong because an empty string is not a valid alias in MLflow; aliases must be non-empty strings, and using an empty string would cause the function to fail or match no events. Option C is wrong because 'experiment' is not a predefined alias in MLflow Model Registry; it refers to an MLflow Experiment, not a model version alias, and would not trigger on model promotion events. Option D is wrong because 'latest' is not a standard alias in MLflow; while MLflow can retrieve the latest model version by version number, the 'latest' alias does not exist, and using it would not capture the champion promotion event.

9
MCQeasy

A data scientist wants to share a trained model with the team for review before deployment. The model is stored in Vertex AI Model Registry. What is the recommended way to grant the team read access to the model?

A.Grant the IAM role 'roles/aiplatform.admin' to the team members.
B.Export the model as a local file and share it via a shared drive.
C.Grant the IAM role 'roles/aiplatform.viewer' to the team members on the project.
D.Add the team members to the Cloud Storage bucket ACL with 'READER' access.
AnswerC

Granting `roles/aiplatform.viewer` at project level gives team members read-only access to all Vertex AI resources, including registered models in Model Registry, satisfying the review-before-deployment requirement. It is the least-privilege predefined role that permits viewing models without granting deploy or edit permissions.

Why this answer

The 'roles/aiplatform.viewer' IAM role grants read-only access to Vertex AI resources, including models in the Model Registry. Option A is incorrect because 'roles/aiplatform.admin' grants full administrative access, which is too broad for read-only needs. Option B is wrong because exporting the model and sharing via a shared drive bypasses version control and security best practices.

Option D is incorrect because Cloud Storage bucket ACLs control access to the underlying bucket, not to the Vertex AI Model Registry; the model is managed through Vertex AI IAM.

10
Multi-Selectmedium

You are collaborating with a team of data scientists on a Vertex AI Workbench notebook that preprocesses data for a machine learning model. You need to ensure that all team members can work on the notebook simultaneously without overwriting each other's changes, and that the notebook's execution environment remains consistent across the team. (Choose two.)

Select 2 answers
A.Store the notebook in a Cloud Source Repositories repository and use Git for version control.
B.Use Vertex AI Workbench's managed notebooks, which automatically sync all user changes to a central repository without requiring Git.
C.Enable the notebook's built-in real-time collaboration feature, which allows multiple users to edit the same notebook simultaneously.
D.Schedule regular exports of the notebook to a shared Cloud Storage bucket, and have team members import the latest version before making changes.
E.Configure the notebook to use a custom container image stored in Artifact Registry, and share the image with the team.
AnswersA, E

Using Cloud Source Repositories with Git provides version control, allowing multiple team members to work on the notebook concurrently by branching and merging. It tracks changes and resolves conflicts, ensuring that no work is overwritten. This is a standard collaborative practice for code and notebooks, and it integrates with Vertex AI Workbench.

Why this answer

Git-based version control via Cloud Source Repositories enables concurrent work with conflict resolution, while a shared custom container image ensures a consistent environment. Together, they allow team members to collaborate effectively without overwriting changes and with identical dependencies. Other options either lack robust version control or misstate Vertex AI Workbench capabilities.

Exam trap

The trap here is assuming that Vertex AI Workbench has built-in real-time collaboration or automatic syncing, when in fact you must use external tools like Git and Artifact Registry for effective teamwork.

11
MCQhard

A large organization uses a multi-project setup with a central data lake. Different teams manage their own models. To enable cross-team sharing of features, they want to use Vertex AI Feature Store. What is the best practice to manage access?

A.Create a single Feature Store in a central project and grant fine-grained IAM roles
B.Export features to Cloud Storage
C.Create separate Feature Stores per team project
D.Use BigQuery authorized views
AnswerA

A single central Feature Store lets teams share features across projects, while fine-grained IAM roles restrict each team to only the features and operations it needs. This satisfies the cross-team sharing requirement without granting broad project-level access, which would violate least privilege.

Why this answer

Creating a single Feature Store in a central project with fine-grained IAM roles is the best practice because it centralizes feature management while allowing cross-team access control at the feature group or feature level. Vertex AI Feature Store supports IAM roles like `aiplatform.featureStoreAdmin` and `aiplatform.featureStoreDataViewer` to grant granular permissions, enabling teams to share features without duplicating data or exposing sensitive information. This approach avoids data silos and ensures consistent governance across the organization.

Exam trap

Google Cloud often tests the misconception that separate Feature Stores per team are needed for isolation, but the correct approach is to use a single Feature Store with fine-grained IAM to enable sharing while maintaining security.

How to eliminate wrong answers

Option B is wrong because exporting features to Cloud Storage introduces data duplication, latency, and manual synchronization overhead, defeating the purpose of a centralized feature store for real-time serving. Option C is wrong because creating separate Feature Stores per team project creates data silos, preventing cross-team sharing and requiring complex cross-project networking or data replication. Option D is wrong because BigQuery authorized views are designed for table-level access control in BigQuery, not for managing access to Vertex AI Feature Store entities like feature groups or online/offline stores, and they lack the low-latency serving capabilities of Feature Store.

12
MCQmedium

You are a machine learning engineer at a retail company. Your team uses Vertex AI Pipelines to train a model that predicts customer churn. The pipeline reads training data from a BigQuery table that is updated daily by an external marketing analytics team. You need to ensure that every pipeline run uses a consistent snapshot of the data and that you can reproduce any past run for auditing. What should you do?

A.Configure the pipeline's training component to query the BigQuery table using a parameter that specifies the current date, and log that date in the pipeline run metadata.
B.Before each pipeline run, create a BigQuery table snapshot of the source table and configure the pipeline to read from that snapshot.
C.Set up a Cloud Scheduler job that copies the BigQuery table to a Cloud Storage bucket as CSV files, and have the pipeline read from that bucket.
D.Use BigQuery's streaming inserts to push data into a new table that the pipeline reads, ensuring that the data is frozen at the time of insertion.
AnswerB

BigQuery table snapshots are immutable and preserve the table's data at the time of creation. By reading from a snapshot, each pipeline run uses a consistent, reproducible dataset, and the snapshot can be referenced later for auditing. This directly addresses both consistency and reproducibility without altering the source table.

Why this answer

Using BigQuery table snapshots ensures that each pipeline run reads an immutable, point-in-time copy of the source data. This guarantees consistency across runs and enables exact reproducibility for audits, because the snapshot preserves the data as it existed when created. Other methods either do not freeze the data or introduce complexity without the same guarantees.

Exam trap

The trap here is assuming that querying a table with a date filter or exporting data provides a consistent snapshot, when only a native BigQuery snapshot guarantees immutability.

13
MCQeasy

A data science team uses Vertex AI Workbench and wants to share notebooks with version history. Which service should they use?

A.Artifact Registry
B.Cloud Storage
C.Data Catalog
D.Cloud Source Repositories
AnswerD

Cloud Source Repositories provides Git-based version control, so notebooks committed from Vertex AI Workbench retain full commit history and can be shared and cloned by teammates. Workbench itself stores notebooks locally without versioning, and Cloud Storage offers object storage rather than revision tracking.

Why this answer

Cloud Source Repositories (CSR) is the correct choice because it provides Git-based version control for notebooks, enabling teams to track changes, collaborate, and maintain a full version history. Vertex AI Workbench integrates natively with CSR, allowing users to clone, commit, and push notebook files directly from the JupyterLab interface, which is essential for collaborative development with revision tracking.

Exam trap

Google Cloud often tests the distinction between storage services (Cloud Storage) and version control services (Cloud Source Repositories), leading candidates to choose Cloud Storage because it has object versioning, but it lacks the collaborative Git workflow required for notebook version history.

How to eliminate wrong answers

Option A is wrong because Artifact Registry is designed for storing and managing container images and ML artifacts (e.g., models, packages), not for version-controlling notebook files or providing a Git-based history. Option B is wrong because Cloud Storage is an object store for unstructured data; it supports object versioning but lacks the branching, merging, and collaborative workflow features of a Git repository, making it unsuitable for notebook version history. Option C is wrong because Data Catalog is a metadata management service for discovering and tagging assets (e.g., datasets, models), not a version control system for code or notebooks.

14
MCQeasy

A machine learning team uses Vertex AI Pipelines to orchestrate training workflows. They want to share pipeline runs and artifacts with stakeholders who do not have Google Cloud accounts. What should they do?

A.Generate a pipeline run report using the Vertex AI Pipelines SDK and export it as a static HTML file to share.
B.Use Vertex AI Pipelines' built-in sharing feature to generate a public URL for the pipeline run.
C.Export the pipeline run's metadata and artifacts to a Cloud Storage bucket and grant public access.
D.Create a custom dashboard in Looker Studio that reads from Vertex ML Metadata and share it with the stakeholders.
AnswerA

The Vertex AI Pipelines SDK allows generating a detailed report of a pipeline run, including artifacts and metrics, which can be exported as a static HTML file. This file can be shared with anyone, regardless of Google Cloud account, providing a snapshot of the run. It is a secure and straightforward way to share information with external stakeholders.

Why this answer

Exporting a pipeline run report as a static HTML file using the Vertex AI Pipelines SDK enables sharing with stakeholders who lack Google Cloud accounts. This method provides a comprehensive view of the run's artifacts and metrics without requiring access to the cloud environment, ensuring both security and accessibility.

Exam trap

The trap here is assuming that Vertex AI Pipelines has a built-in public sharing URL or that public bucket access is acceptable, when the correct approach is to generate a static report for external sharing.

15
Multi-Selectmedium

Which THREE actions are best practices for managing ML models in production on Google Cloud? (Choose 3)

Select 3 answers
A.Manually tune hyperparameters for each retraining run.
B.Monitor model performance and data drift continuously.
C.Use a central model registry for model governance.
D.Version all model artifacts and training datasets.
E.Store all raw training data indefinitely for auditability.
AnswersB, C, D

Continuous monitoring of model performance and data drift detects degradation after deployment, satisfying the production oversight requirement. Unlike static validation, it tracks live inference distributions against training baselines, triggering retraining when drift exceeds thresholds. This directly addresses the stem's demand for ongoing operational management of deployed ML models.

Why this answer

Option B is correct because continuous monitoring of model performance and data drift is essential in production to detect degradation and trigger retraining before business impact occurs. Option C is correct because a central model registry (such as Vertex AI Model Registry) provides governance, lineage, and controlled promotion of models across environments. Option D is correct because versioning model artifacts and training datasets ensures reproducibility, traceability, and rollback capability for every deployed model.

Option A is not a best practice because manual hyperparameter tuning does not scale and should be automated with tools like Vertex AI Vizier. Option E is not a best practice because retaining all raw training data indefinitely increases cost and compliance risk; retention should follow defined policies and lifecycle rules.

Exam trap

Google Cloud often tests the misconception that manual hyperparameter tuning is acceptable for production, when in fact automation (e.g., Vertex AI Vizier) is the recommended practice to ensure reproducibility and efficiency.

16
MCQmedium

A team is using Vertex AI Experiments to compare different hyperparameters. They want to automatically record the hyperparameters. What is the correct way?

A.Manually log to console
B.Use the `aiplatform.start_run()` context manager
C.Write to a CSV file
D.Use BigQuery
AnswerB

This context manager automatically logs hyperparameters and metrics to Vertex AI Experiments.

Why this answer

Vertex AI Experiments provides a native `aiplatform.start_run()` context manager that automatically captures hyperparameters passed as key-value arguments, logging them to the experiment run metadata without manual intervention. This integrates directly with the Vertex AI SDK, ensuring consistency and traceability across runs.

Exam trap

Google Cloud often tests the misconception that any logging method (console, CSV, BigQuery) is equivalent to native SDK integration, but the key requirement is automatic, structured recording tied to the experiment run, which only the SDK's context manager provides.

How to eliminate wrong answers

Option A is wrong because manually logging to console only outputs data to stdout, which is not persisted in Vertex AI Experiments and cannot be queried or compared programmatically. Option C is wrong because writing to a CSV file requires custom I/O code, lacks integration with Vertex AI's experiment tracking, and does not associate the hyperparameters with a specific experiment run. Option D is wrong because BigQuery is a data warehouse for analytics, not a mechanism for automatically recording hyperparameters during model training; it would require additional infrastructure to capture and store the parameters.

17
MCQhard

A company uses a Cloud Composer DAG to run a daily ML pipeline that includes Dataflow jobs and model training on Vertex AI. The pipeline frequently fails due to insufficient permissions when the Dataflow worker accesses data in Cloud Storage. What is the most efficient way to resolve this issue?

A.Create a custom service account with required permissions and assign it to the Dataflow job.
B.Grant the 'roles/storage.objectViewer' role to 'allUsers' on the Cloud Storage bucket.
C.Use the Composer environment's service account for all pipeline components.
D.Move the Dataflow job to run after the pipeline so that data is already processed.
AnswerA

Dataflow workers execute under a service account, so granting the required Cloud Storage permissions to a dedicated custom service account and assigning it to the job fixes the access failures directly. This is more efficient than broadly widening project-level roles.

Why this answer

The most efficient way to resolve insufficient permissions for Dataflow workers accessing Cloud Storage is to create a custom service account with the required roles (e.g., roles/storage.objectViewer) and assign it to the Dataflow job via the --serviceAccount option. This follows the principle of least privilege and ensures that only the Dataflow workers have the necessary permissions, without affecting other pipeline components or exposing the bucket publicly.

Exam trap

Google Cloud often tests the misconception that using a single service account for all components (like the Composer environment's service account) is simpler and sufficient, but this ignores the principle of least privilege and can cause security vulnerabilities or permission conflicts in distributed pipelines.

How to eliminate wrong answers

Option B is wrong because granting roles/storage.objectViewer to 'allUsers' makes the Cloud Storage bucket publicly readable, which is a severe security risk and violates least privilege principles. Option C is wrong because the Composer environment's service account typically has broader permissions than needed for Dataflow workers, and using it for all components can lead to over-privileging and potential security issues; moreover, Dataflow workers require a separate identity to access resources independently. Option D is wrong because moving the Dataflow job to run after the pipeline does not address the root cause of insufficient permissions; the Dataflow job will still fail when it tries to access Cloud Storage data, regardless of when it runs.

18
MCQeasy

A team uses Vertex AI Feature Store for storing features. They want to share feature definitions with other teams in a collaborative manner. What is the best way to collaborate on feature definitions?

A.Use a shared repository with feature definition files and CI/CD to update the feature store.
B.Grant all teams write access to the same feature store so they can modify definitions directly.
C.Export the feature definitions as CSV and email them to the other teams.
D.Use a wiki page to document feature definitions and update it manually.
AnswerA

Feature definitions stored as version-controlled files in a shared repository let multiple teams review, reuse and propose changes collaboratively. CI/CD then applies those definitions to Vertex AI Feature Store consistently, satisfying the requirement for collaborative sharing rather than ad-hoc manual updates.

Why this answer

Using a shared repository with feature definition files and CI/CD pipelines enables version control, peer review, and automated deployment to Vertex AI Feature Store. This approach ensures consistency, traceability, and collaboration without risking direct, uncoordinated changes to the production feature store.

Exam trap

The trap here is that candidates may assume direct write access (Option B) is efficient for collaboration, but the exam tests understanding that feature stores require controlled, versioned updates to maintain data integrity and avoid breaking downstream models.

How to eliminate wrong answers

Option B is wrong because granting all teams write access to the same feature store allows uncoordinated, direct modifications to feature definitions, which can lead to conflicts, data corruption, and lack of version control. Option C is wrong because exporting feature definitions as CSV and emailing them is error-prone, lacks versioning, and does not provide a single source of truth for collaboration. Option D is wrong because using a wiki page for manual documentation is static, easily outdated, and does not integrate with the feature store's actual schema or deployment process.

19
MCQeasy

A data scientist wants to track the lineage of a dataset used in a training run. Which Vertex AI feature should they use?

A.Vertex ML Metadata
B.Vertex AI Feature Store
C.Vertex AI Experiments
D.Vertex AI Model Registry
AnswerA

Vertex ML Metadata records artefacts, executions and contexts, automatically capturing dataset inputs and training runs as lineage graphs. This directly satisfies the stem's requirement to track which dataset fed a training run, letting the scientist query provenance rather than reconstruct it manually from logs.

Why this answer

Vertex ML Metadata is the correct choice because it is specifically designed to track the lineage of datasets, models, and other artifacts throughout the ML lifecycle. It records metadata about each step in a pipeline, including the source dataset used for a training run, enabling full provenance tracking. This allows data scientists to trace back which data was used, how it was transformed, and which model version it produced.

Exam trap

The trap here is that candidates often confuse Vertex AI Experiments (which tracks run metrics and parameters) with lineage tracking, but Experiments does not capture the full artifact-to-execution graph that ML Metadata provides for dataset provenance.

How to eliminate wrong answers

Option B is wrong because Vertex AI Feature Store is a centralized repository for storing, serving, and sharing feature values for ML models, not for tracking dataset lineage. Option C is wrong because Vertex AI Experiments is used to track and compare model training runs, hyperparameters, and metrics, but it does not natively capture the lineage of the dataset itself beyond run-level parameters. Option D is wrong because Vertex AI Model Registry is a version control system for trained models, managing model deployments and versions, but it does not track the provenance of the training data used to create those models.

20
MCQhard

A company has multiple teams working on different models. They want to enforce consistent data preprocessing steps across all teams. Which approach should they take?

A.Use Cloud Composer to orchestrate preprocessing
B.Write shared Python packages in Artifact Registry
C.Use Cloud Dataflow templates
D.Create shared Vertex AI Pipelines components
AnswerD

Shared Vertex AI Pipelines components package preprocessing logic as reusable, versioned artefacts that every team imports into its own pipeline. This enforces identical preprocessing steps across teams, satisfying the consistency constraint without duplicating code per model.

Why this answer

Vertex AI Pipelines components allow teams to define reusable, versioned, and parameterized preprocessing steps that can be shared across models and pipelines. This ensures consistent execution of data transformations because each component encapsulates the exact code and environment, and pipelines enforce the same DAG of steps regardless of which team triggers them.

Exam trap

Google Cloud often tests the distinction between 'sharing code' (e.g., packages) and 'sharing executable, environment-encapsulated pipeline steps' (e.g., components), leading candidates to choose a code-sharing option like Artifact Registry instead of the pipeline component approach that enforces consistency.

How to eliminate wrong answers

Option A is wrong because Cloud Composer is an orchestration service for workflows (based on Apache Airflow) and does not inherently enforce consistent preprocessing logic across teams; it only schedules and monitors tasks, leaving the actual preprocessing code to be defined separately and potentially inconsistently. Option B is wrong because writing shared Python packages in Artifact Registry provides a way to distribute code, but it does not enforce a standardized execution environment or pipeline structure; teams could still call the packages with different parameters or in different orders, leading to inconsistency. Option C is wrong because Cloud Dataflow templates are used for batch and stream data processing jobs (based on Apache Beam), but they are not designed to be shared as reusable, composable steps across multiple ML pipelines; they lack the pipeline-level DAG enforcement and versioning that Vertex AI Pipelines components provide.

21
MCQeasy

When distributing training across multiple workers using Vertex AI Training, how should the team share the training dataset?

A.Copy the dataset to each worker's local disk
B.Use NFS
C.Use Cloud Storage
D.Use Google Drive
AnswerC

Cloud Storage provides a shared, region-agnostic object store that every worker can read concurrently, satisfying the multi-worker distribution requirement without duplicating data. Vertex AI Training mounts or streams GCS paths directly, so each worker accesses the same dataset shards. Persistent disks or local storage would tie data to a single VM, breaking parallel access.

Why this answer

Vertex AI Training workers need shared, concurrent read access to the training dataset without manual replication. Cloud Storage (GCS) is the recommended and fully integrated solution because it provides a distributed, highly available object store that all workers can read from in parallel via the `tf.io.gfile` API or GCS connector, eliminating data duplication and ensuring consistency across the cluster.

Exam trap

The trap here is that candidates confuse 'shared storage' with 'local copies' or 'user-friendly sync tools,' assuming NFS or Drive are viable for distributed ML, when Vertex AI explicitly requires a cloud-native object store like GCS for scalability and fault tolerance.

How to eliminate wrong answers

Option A is wrong because copying the dataset to each worker's local disk introduces data duplication, increases startup latency, and risks inconsistency if workers are preempted or auto-scaled; Vertex AI does not manage local disk replication. Option B is wrong because NFS (Network File System) is not natively supported in Vertex AI Training; it would require manual setup of an NFS server, introduces a single point of failure, and adds network latency that GCS avoids with its native parallel read capabilities. Option D is wrong because Google Drive is a user-facing file sync service, not designed for high-throughput, concurrent access by distributed training jobs; it lacks the necessary IAM integration, access controls, and performance guarantees for ML workloads.

22
Multi-Selectmedium

Which TWO actions are recommended for collaborating on machine learning models using Vertex AI Model Registry?

Select 2 answers
A.Use Cloud Storage object labels to store model descriptions.
B.Use version aliases such as 'champion' and 'challenger' to manage model lifecycle.
C.Deploy all model versions to a single endpoint for comparison.
D.Attach custom metadata (e.g., training dataset, hyperparameters) to each model version.
E.Create a separate model entry for each training run.
AnswersB, D

Aliases enable controlled promotion of models.

Why this answer

Vertex AI Model Registry supports version aliases like 'champion' and 'challenger' to designate which model version should serve as the production candidate and which is under evaluation, enabling controlled lifecycle management and A/B testing without manual version tracking.

Exam trap

Google Cloud often tests the distinction between a single model entry with multiple versions versus separate model entries per run, and candidates mistakenly think separate entries provide better traceability, but the registry's versioning and alias system is specifically designed to avoid that fragmentation.

23
Multi-Selectmedium

Which THREE practices improve collaboration when using Cloud Composer for ML pipelines?

Select 3 answers
A.Keep all pipeline logic in a single large DAG for simplicity.
B.Use a shared Cloud Storage bucket for intermediate artifacts with appropriate permissions.
C.Store DAGs in a version-controlled repository and use CI/CD to deploy them.
D.Embed service account keys directly in DAG code for authentication.
E.Use Airflow variables and connections to parameterize DAGs.
AnswersB, C, E

Facilitates handoff between pipeline steps and teams.

Why this answer

Cloud Composer workflows often require sharing intermediate data (e.g., transformed datasets, model checkpoints) across multiple DAGs or team members. A shared Cloud Storage bucket with fine-grained IAM permissions enables secure, centralized artifact exchange without duplicating data or exposing it to unauthorized users. This practice avoids hard-coded paths and ensures that all pipeline stages can reliably access the same artifacts, which is critical for reproducibility and collaboration in ML pipelines.

Exam trap

Google Cloud often tests the misconception that a single monolithic DAG simplifies collaboration, when in fact it creates bottlenecks and merge conflicts; the trap is that candidates confuse 'simplicity' with 'ease of collaboration' without considering modularity and CI/CD practices.

24
Multi-Selecthard

A team is collaborating on a Vertex AI model using Vertex AI Model Registry. They need to ensure that model versions are properly managed and that deployments are reproducible. Which TWO practices should they follow? (Choose two.)

Select 2 answers
A.Use a single default version for all deployments to simplify endpoint configuration.
B.Rely on the model's auto-generated resource name to identify versions, as it includes a timestamp.
C.Grant all team members the Vertex AI User role on the project to allow unrestricted model uploads.
D.Assign each model version a unique alias that maps to a specific model artifact and container image.
E.Store the training pipeline's parameters and data snapshot URI in the model's description or labels.
AnswersD, E

Aliases in Vertex AI Model Registry provide a stable reference to a specific model version, enabling reproducible deployments. By mapping an alias to a model artifact and its container image, teams can deploy the exact version regardless of later updates. This supports collaboration because team members can refer to the alias without tracking version numbers manually.

Why this answer

Using aliases to map model versions to specific artifacts and container images, and recording pipeline parameters and data snapshot URIs in metadata, together ensure that deployments are reproducible and that team members can collaborate effectively. Aliases provide stable references, while metadata captures the training context for audits and retraining.

Exam trap

The trap here is assuming that auto-generated resource names or broad permissions are sufficient for version management, when in fact deliberate aliasing and metadata capture are required for reproducibility.

25
MCQeasy

A data science team is using Vertex AI Pipelines to orchestrate their ML workflows. They want to ensure that each pipeline run is reproducible and that artifacts are versioned. Which Vertex AI feature should they use to track and manage pipeline artifacts?

A.Vertex AI Model Registry
B.Vertex ML Metadata
C.Vertex AI Experiments
D.Vertex AI Feature Store
AnswerB

Vertex ML Metadata automatically tracks artifacts, executions, and contexts produced by Vertex AI Pipelines runs. It provides lineage and versioning, enabling reproducibility by recording parameters and artifacts. This is the native service for managing pipeline artifacts and their relationships, making it the correct choice for the team's requirement.

Why this answer

Vertex ML Metadata is the dedicated service for capturing and managing metadata and artifacts from Vertex AI Pipelines. It records executions, artifacts, and their relationships, enabling lineage tracking and reproducibility. The other options serve different purposes: Experiments for run comparison, Model Registry for models, and Feature Store for feature serving.

Exam trap

The trap here is confusing Vertex AI Experiments with Vertex ML Metadata; while both track metadata, only Vertex ML Metadata provides comprehensive artifact lineage and versioning for pipeline executions.

26
Multi-Selecthard

Which THREE actions should be taken to manage model versions effectively?

Select 3 answers
A.Delete old versions immediately
B.Use Vertex AI Model Registry
C.Set up model evaluation alerts
D.Use the same model name for all versions
E.Assign version aliases like 'champion' and 'experiment'
AnswersB, C, E

Model Registry provides versioning and deployment control.

Why this answer

Vertex AI Model Registry is a centralized repository that tracks, versions, and manages ML models. It enables you to organize models, assign aliases (like 'champion' or 'experiment'), and control deployment, ensuring reproducibility and governance across the model lifecycle.

Exam trap

Google Cloud often tests the misconception that deleting old versions is a best practice for storage optimization, when in reality versioning requires retaining history for reproducibility and rollback, and that aliases are the correct mechanism for labeling model stages.

27
MCQmedium

A team of ML engineers is collaborating on a project using Vertex AI. They want to ensure that only approved models are deployed to production. Which approach should they use?

A.Store all models in a Cloud Storage bucket and manually control access via IAM permissions.
B.Deploy models directly from training jobs to an endpoint without version tracking.
C.Use Vertex AI Model Registry with version aliases to manage model versions and promote them after approval.
D.Use Cloud Dataflow to transform raw predictions and then store them in BigQuery for analysis.
AnswerC

Vertex AI Model Registry with version aliases lets the team track model versions and control which alias points to an approved artefact, so only vetted models reach production. Promotion after approval enforces the governance gate the scenario requires.

Why this answer

Vertex AI Model Registry provides a centralized repository for managing model versions, with support for version aliases (e.g., 'champion', 'challenger') that allow teams to promote models to production only after approval. This ensures governance and traceability, meeting the requirement that only approved models are deployed.

Exam trap

The trap here is that candidates may confuse storage access control (IAM) with model lifecycle governance, or assume that any data pipeline tool (Dataflow) can manage model approvals, when in fact only a dedicated model registry with version aliases provides the required approval workflow and traceability.

How to eliminate wrong answers

Option A is wrong because storing models in Cloud Storage with manual IAM control lacks version tracking, approval workflows, and integration with Vertex AI's deployment services, making it error-prone and unscalable for production governance. Option B is wrong because deploying directly from training jobs without version tracking bypasses model validation, approval gates, and rollback capabilities, violating the requirement for controlled production deployments. Option D is wrong because Cloud Dataflow is a data processing service for stream/batch pipelines, not a model management or approval mechanism; it is irrelevant to controlling which models are deployed.

28
Multi-Selectmedium

Which TWO of the following are best practices for managing data in a collaborative machine learning environment on Google Cloud?

Select 2 answers
A.Always replicate data across multiple regions to ensure low latency.
B.Implement fine-grained access control using IAM conditions.
C.Use Cloud Data Catalog to discover and annotate datasets.
D.Store all raw data in a single Cloud Storage bucket for easy access.
E.Use data versioning with tools like DVC or Dataflow to track changes.
AnswersC, E

Data Catalog aids in data governance and collaboration.

Why this answer

Cloud Data Catalog provides a managed metadata management service that allows teams to discover, annotate, and manage datasets across Google Cloud. It enables data scientists to search for datasets by tags, descriptions, and schema, which is essential for collaboration and data governance in a multi-user ML environment.

Exam trap

Google Cloud often tests the misconception that 'replication equals performance' or that 'single bucket simplicity is best,' when in reality collaborative ML requires discoverability (Data Catalog) and reproducibility (versioning) over raw storage or access control alone.

29
MCQmedium

An MLOps team needs to automatically retrain a model when new training data becomes available. They use Vertex AI Pipelines. What is the recommended way to trigger the pipeline?

A.Use Model Evaluation to decide
B.Set up a trigger in Vertex AI Pipelines
C.Cloud Functions triggered by Cloud Storage events
D.Cloud Scheduler on a daily basis
AnswerC

Cloud Storage object-finalise events routed through Eventarc invoke a Cloud Function, which then submits the Vertex AI Pipeline job. This satisfies the stem's requirement to retrain automatically when new training data lands, since the trigger fires on data arrival rather than on a schedule or manual invocation.

Why this answer

Vertex AI Pipelines does not natively support event-driven triggers. The recommended pattern is to use Cloud Functions, which can be triggered by Cloud Storage events (e.g., object finalize/create) when new training data is uploaded. The Cloud Function then programmatically submits the pipeline run via the Vertex AI Pipelines client library or REST API, enabling an automated retraining workflow.

Exam trap

The trap here is that candidates assume Vertex AI Pipelines has a built-in trigger mechanism (Option B) because many CI/CD tools do, but Google Cloud's recommended pattern relies on external event-driven services like Cloud Functions.

How to eliminate wrong answers

Option A is wrong because Model Evaluation is a post-training assessment step, not a trigger mechanism; it cannot initiate pipeline execution. Option B is wrong because Vertex AI Pipelines itself does not provide a built-in trigger; triggers must be implemented externally via Cloud Functions, Cloud Scheduler, or similar services. Option D is wrong because Cloud Scheduler on a daily basis is a time-based trigger, not an event-driven one; it would retrain on a fixed schedule regardless of whether new data has arrived, leading to unnecessary runs or missed retraining opportunities.

30
Multi-Selecteasy

Which TWO statements about Vertex AI Feature Store are correct? (Choose 2)

Select 2 answers
A.Feature Store automatically applies feature engineering transformations.
B.Feature Store can only store numerical features.
C.Feature Store can only be used with Vertex AI models.
D.Feature Store provides a centralized repository for feature data.
E.Feature Store supports both online and offline serving.
AnswersD, E

Vertex AI Feature Store acts as a centralised repository, letting teams register, version and share feature data across projects and models instead of duplicating feature engineering per pipeline, which is the defining architectural property this statement asserts.

Why this answer

Option D is correct because Vertex AI Feature Store acts as a centralized repository where feature data is registered, versioned, and managed as feature groups and features, so multiple teams and models can share a single consistent source of features. Option E is correct because Feature Store supports both online serving, which returns low-latency feature values for real-time predictions, and offline serving, which reads historical feature values in bulk for training and batch scoring. Options A, B, and C are incorrect: Feature Store does not automatically perform feature engineering transformations (you define and ingest the transformed values yourself), it is not limited to numerical features (it supports types such as strings, booleans, and arrays as well), and it is not restricted to Vertex AI models since features can be served to any application or model via the online/offline APIs.

Exam trap

Google Cloud often tests the misconception that Vertex AI Feature Store is tightly coupled to Vertex AI models or that it performs automatic feature engineering, when in fact it is a decoupled storage and serving layer that supports any ML framework and requires explicit feature engineering steps.

31
MCQmedium

A healthcare organization is building a machine learning model to predict patient readmission risk. They have sensitive data stored in BigQuery that includes protected health information (PHI). The data science team uses Vertex AI Workbench notebooks to explore the data and develop models. The organization's security policy requires that all PHI data must be encrypted at rest and in transit, and that access to the data is logged and audited. They also need to ensure that the data used for model training is de-identified to remove direct identifiers such as patient names and SSNs. The team wants to automate the de-identification process as part of the data pipeline. Which approach meets these requirements?

A.Create a Dataflow pipeline that reads from the original BigQuery table, applies Cloud DLP de-identification transforms, and writes to a new BigQuery table. Grant the data science team access to the de-identified table.
B.Enable Shielded VM on Vertex AI Workbench notebooks and use VPC-SC to restrict data access.
C.Use Cloud Key Management Service to encrypt the PHI columns in BigQuery, and share the encryption key with the data science team.
D.Use BigQuery row-level security to mask PHI columns for the data science team, and train the model directly on the original table.
AnswerA

Cloud DLP de-identification transforms applied in a Dataflow pipeline remove direct identifiers before the data lands in a separate BigQuery table, satisfying the de-identification requirement while BigQuery's default encryption at rest and TLS in transit plus audit logging cover the remaining policy constraints.

Why this answer

It uses Cloud DLP within a Dataflow pipeline to automatically de-identify PHI data as it is read from the original BigQuery table and written to a new, de-identified table. This satisfies the requirement for automated de-identification, while the original table remains encrypted at rest (BigQuery default) and in transit (TLS), and access to the original data can be logged via Cloud Audit Logs. The data science team only gets access to the de-identified table, ensuring PHI is not exposed during model development.

Exam trap

Google Cloud often tests the distinction between data masking/encryption (which still exposes PHI to authorized users) and true de-identification (which removes or transforms PHI so it is no longer considered protected health information).

How to eliminate wrong answers

Option B is wrong because Shielded VM and VPC-SC provide infrastructure security (integrity, network perimeter) but do not de-identify PHI data; the data science team would still see raw PHI in the notebooks. Option C is wrong because Cloud KMS encryption protects data at rest but does not remove or mask PHI columns; sharing the encryption key with the data science team would give them access to the raw PHI, violating the de-identification requirement. Option D is wrong because BigQuery row-level security masks columns at query time but does not de-identify the underlying data; the model training would still use the original table with PHI present in the masked columns, and the masking is not a permanent de-identification suitable for an automated pipeline.

32
MCQeasy

Refer to the exhibit. The team notices that the pipeline fails to read data from the specified Cloud Storage path. What is the most likely issue?

A.The bucket does not exist
B.The pipeline runner is incorrect
C.The region is mismatched
D.The service account lacks `storage.objectViewer` permission
AnswerD

A missing `storage.objectViewer` role on the service account directly prevents the pipeline from listing and reading objects at the Cloud Storage path, satisfying the stem's read-failure constraint. Without this permission, requests return 403 errors, so the pipeline cannot access the data regardless of path correctness.

Why this answer

The pipeline fails to read data from Cloud Storage because the service account lacks the `storage.objectViewer` IAM role, which grants the `storage.objects.get` and `storage.objects.list` permissions required to read objects. Without this role, the pipeline cannot authenticate or authorize the read operation, even if the bucket and path are correct.

Exam trap

Google Cloud often tests the distinction between bucket-level permissions (like `storage.objectViewer`) and project-level roles, leading candidates to overlook that the service account must have the specific IAM role on the bucket or project, not just any storage role.

How to eliminate wrong answers

Option A is wrong because if the bucket did not exist, the error would typically be a 404 'Bucket not found' or a similar explicit message, not a generic read failure. Option B is wrong because the pipeline runner (e.g., Dataflow, Apache Beam) is responsible for executing the pipeline logic, not for authenticating to Cloud Storage; a runner mismatch would cause execution errors, not permission-related read failures. Option C is wrong because Cloud Storage bucket access is global and region-mismatch errors occur only for specific operations like writing to a regional bucket from a different region, but reading is allowed across regions; a region mismatch would not block read access.

33
Multi-Selecthard

Which THREE of the following are recommended practices for model governance and lineage in Vertex AI?

Select 3 answers
A.Enable Vertex AI ML Metadata to track artifacts, executions, and contexts.
B.Use Vertex AI Experiments to log parameters and metrics.
C.Store model artifacts in Cloud Storage with metadata in a database.
D.Manually record model lineage in a spreadsheet.
E.Use Vertex AI Model Registry to manage model versions and stages.
AnswersA, B, E

ML Metadata provides automated lineage tracking.

Why this answer

Vertex AI ML Metadata is a fully managed service that automatically tracks artifacts, executions, and contexts across the ML workflow. By enabling it, you create a lineage graph that records every step from data preparation to model deployment, which is essential for auditability and reproducibility. This is a core recommended practice for model governance because it provides an immutable, queryable history of all model-related activities.

Exam trap

Google Cloud often tests the distinction between using native Vertex AI services (like ML Metadata, Experiments, and Model Registry) versus ad-hoc or manual methods (like spreadsheets or custom databases) that lack automated governance and audit trails.

34
MCQhard

Your team uses Vertex AI Pipelines to automate the training and deployment of a recommendation model. The pipeline includes a step that evaluates the model and only deploys it if the evaluation metric exceeds a threshold. You need to ensure that the pipeline's artifacts, including the evaluation metrics and the deployed model, are tracked and can be traced back to the pipeline run for auditing. What should you do?

A.Configure the pipeline to send an email with the evaluation metrics and model details to a distribution list for record-keeping.
B.Enable Cloud Logging for the pipeline and rely on the logs to capture the evaluation metrics and model deployment events.
C.Store the evaluation metrics in a BigQuery table and the model in Cloud Storage, and record the pipeline run ID in both locations.
D.Use Vertex AI ML Metadata to log the evaluation metrics and model artifacts, and associate them with the pipeline run.
AnswerD

Vertex AI ML Metadata automatically tracks artifacts, executions, and contexts for Vertex AI Pipelines runs. By logging metrics and model artifacts, you create a lineage that links them to the pipeline run. This enables auditing and reproducibility, as you can trace which run produced which model and its metrics.

Why this answer

Vertex AI ML Metadata provides automatic lineage tracking for pipeline artifacts, including metrics and models. It records relationships between executions and artifacts, enabling auditing and reproducibility. Other methods lack the structured, integrated metadata store that ML Metadata offers, making them less reliable for tracing and compliance.

Exam trap

The trap here is assuming that manual logging or Cloud Logging can substitute for ML Metadata's automatic lineage tracking, when only ML Metadata provides a queryable artifact graph.

35
MCQmedium

A team uses Vertex AI Pipelines. They need to ensure that only certain team members can deploy models to production. What is the best approach?

A.Use Vertex AI Experiments to track models
B.Store model artifacts in a bucket with bucket-level permissions
C.Use IAM roles with custom permissions on the Vertex AI Model Registry
D.Create separate projects for dev and prod
AnswerC

Model Registry integrates with IAM to grant specific deployment permissions.

Why this answer

Vertex AI Model Registry supports IAM roles with custom permissions, allowing fine-grained access control over who can promote or deploy models to production. By assigning specific roles (e.g., `roles/aiplatform.modelDeployer`) to only authorized team members, you can restrict deployment actions while still permitting others to view or register models. This approach directly addresses the need to control production deployments without affecting other pipeline stages.

Exam trap

The trap here is that candidates often confuse artifact storage permissions (bucket-level IAM) with deployment permissions (model registry IAM), leading them to choose Option B, even though bucket permissions do not control the Vertex AI deployment API call.

How to eliminate wrong answers

Option A is wrong because Vertex AI Experiments is designed for tracking and comparing model training runs (e.g., hyperparameters, metrics), not for controlling access or permissions to deploy models. Option B is wrong because bucket-level permissions control access to the storage location of model artifacts, but they do not govern the deployment action itself within Vertex AI Pipelines; a user with bucket access could still lack deployment permissions, or vice versa. Option D is wrong because creating separate projects for dev and prod is an organizational boundary that can help with isolation, but it does not provide granular control over which specific team members can deploy within the same project; it also introduces overhead in managing multiple projects and does not leverage Vertex AI's native IAM capabilities for model registry operations.

36
MCQhard

Your organization uses Vertex AI Model Registry to manage models. A data scientist has trained a new model version and wants to ensure that only approved models are deployed to production. You need to implement a workflow where a model must be reviewed and approved by a designated approver before it can be deployed. What should you do?

A.Implement a CI/CD pipeline using Cloud Build that, upon a model version being tagged as 'approved', deploys it to an endpoint, and restrict the tagging permission to the approver.
B.Use Vertex AI Model Registry's built-in approval workflow by assigning the approver role to a user and setting the model version's state to 'Pending Approval'.
C.Create a Cloud Function that listens for new model versions, triggers a manual review via email, and upon approval, updates the model version's labels to indicate approval.
D.Restrict deployment permissions using IAM so that only the approver has the Vertex AI User role, and require the approver to manually deploy the model after review.
AnswerA

This approach uses a CI/CD pipeline to enforce that only models tagged as 'approved' are deployed. By restricting the tagging permission to the approver via IAM, you ensure that only authorized personnel can trigger deployment. Cloud Build automates the deployment, providing an auditable and repeatable process that integrates with Vertex AI Model Registry.

Why this answer

Enforcing an approval workflow requires a combination of access control and automation. By using a CI/CD pipeline that reacts to an 'approved' tag, and limiting who can apply that tag, you create a secure gate. This ensures that only reviewed models are deployed and provides an audit trail.

Native Model Registry lacks built-in approval, so custom logic is necessary.

Exam trap

The trap here is assuming that Vertex AI Model Registry includes a native approval workflow or that labels alone can securely gate deployments, when in fact you must combine IAM and external automation.

37
Multi-Selecteasy

Which TWO practices help ensure reproducible ML experiments?

Select 2 answers
A.Store all artifacts in a temporary bucket
B.Use a random seed for each run
C.Use Vertex AI Experiments to track parameters and metrics
D.Version control training code with Cloud Source Repositories
E.Use preemptible VMs
AnswersC, D

Experiments record the exact configuration and results.

Why this answer

Vertex AI Experiments automatically logs parameters, metrics, and artifacts for each run, creating a complete lineage that enables exact reproduction of results. By tracking these details alongside the code version, you can recreate the exact environment and configuration that produced a given outcome, which is essential for reproducibility.

Exam trap

Google Cloud often tests the distinction between practices that improve reproducibility (like tracking parameters and versioning code) versus practices that improve cost efficiency or speed (like using preemptible VMs or temporary storage), leading candidates to conflate operational convenience with scientific reproducibility.

38
MCQmedium

A company uses BigQuery to store feature data for ML training. A data engineer notices that a Vertex AI Training job is failing with 'Access Denied' errors when reading from a BigQuery table. The training job uses a custom service account that has been granted the 'bigquery.dataViewer' role on the dataset. What is the most likely cause of the failure?

A.The service account is not in the same project as the BigQuery dataset.
B.The BigQuery table is partitioned and requires row-level access.
C.The service account lacks the 'bigquery.jobs.create' permission in the project.
D.The training job does not have the required network access to BigQuery.
AnswerC

Reading a BigQuery table requires a query job, which needs bigquery.jobs.create in the project. The bigquery.dataViewer role on the dataset grants table-level read access only, so the custom service account cannot start the job and Vertex AI returns Access Denied.

Why this answer

The 'bigquery.dataViewer' role grants permissions to read BigQuery data (e.g., bigquery.tables.getData), but it does not include the 'bigquery.jobs.create' permission. When a Vertex AI training job reads from BigQuery, it must first create a BigQuery job (a query job) to retrieve the data. Without 'bigquery.jobs.create' at the project level, the service account cannot initiate the read operation, resulting in an 'Access Denied' error even though it has data-level access.

Exam trap

The trap here is that candidates often assume 'bigquery.dataViewer' is sufficient for all read operations, overlooking the requirement for 'bigquery.jobs.create' to initiate the query job that actually reads the data.

How to eliminate wrong answers

Option A is wrong because the service account does not need to be in the same project as the BigQuery dataset; cross-project access is supported as long as IAM permissions are granted at the dataset or table level. Option B is wrong because partitioned tables do not require row-level access by default; row-level access is controlled via BigQuery row-level security policies, which are not automatically required for partitioned tables. Option D is wrong because Vertex AI training jobs run within Google Cloud's internal network and have built-in access to BigQuery via the Cloud API; network access is not a common cause of 'Access Denied' errors for BigQuery reads.

Ready to test yourself?

Try a timed practice session using only Data Model Mgmt questions.