Courseiva

CCNA Collaborating Within and Across Teams to Manage Data and Models Questions

75 questions · Collaborating Within and Across Teams to Manage Data and Models · All types, answers revealed

1
MCQeasy

A team wants to track the lineage of ML pipeline runs, including which datasets, parameters, and models were used in each execution. Which Vertex AI service should they use?

A.Vertex AI Metadata
B.Vertex AI Feature Store
C.Vertex AI Model Registry
D.Vertex AI Experiments
AnswerA

Vertex AI Metadata stores artefacts and executions in a lineage graph, recording which datasets, parameters and models each pipeline run consumed and produced. This directly satisfies the requirement to track lineage across ML pipeline executions.

Why this answer

Vertex AI Metadata (part of Vertex ML Metadata) is the service designed to record and query ML metadata, including artifacts (datasets, models), executions (pipeline runs), and events, forming a lineage graph. It captures which datasets, parameters, and models were used in each pipeline execution, enabling reproducibility and auditability. This directly matches the requirement to track lineage of ML pipeline runs.

Exam trap

PMLE often tests the confusion between Vertex AI Experiments (run tracking and comparison) and Vertex ML Metadata (lineage and artifact relationships), causing candidates to pick Experiments when lineage is the requirement.

How to eliminate wrong answers

Option B (Vertex AI Feature Store) is wrong because it manages and serves feature values for training and online serving, not pipeline run lineage or artifact tracking. Option C (Vertex AI Model Registry) is wrong because it stores and versions trained models and their metadata, but it does not track the full pipeline execution lineage including datasets and parameters. Option D (Vertex AI Experiments) is wrong because it tracks experiment runs, metrics, and parameters for comparison, but it is not the underlying lineage/metadata service that records artifact relationships across pipeline executions.

2
Multi-Selectmedium

An ML team uses Delta Lake on Dataproc for data versioning. Which THREE benefits does Delta Lake provide?

Select 3 answers
A.Automatic data encryption at rest
B.Time travel for accessing previous versions
C.Schema enforcement and evolution
D.ACID transactions on data lakes
E.Built-in real-time streaming
AnswersB, C, D

Delta Lake's transaction log records every write as a numbered commit, so querying a prior version by timestamp or version number retrieves the exact historical snapshot. This directly satisfies the team's data versioning requirement on Dataproc, enabling reproducible ML training and rollback without duplicating datasets.

Why this answer

Delta Lake provides time travel (B), allowing queries against earlier snapshots of a table via version numbers or timestamps, which directly supports the team's data versioning requirement. It also delivers schema enforcement and evolution (C), rejecting writes that violate the table schema while permitting controlled schema changes such as adding columns. Additionally, Delta Lake brings ACID transactions (D) to data lakes, ensuring atomic, consistent, isolated, and durable writes on top of object storage like GCS.

Options A and E are not Delta Lake benefits: encryption at rest is handled by the underlying storage service (e.g., Google Cloud Storage), not Delta Lake itself, and Delta Lake is a storage/transaction layer rather than a built-in real-time streaming engine.

3
MCQeasy

A data scientist finishes training a model in a Vertex AI Workbench notebook and wants to save it so that the deployment team can later deploy it to an endpoint without re-running the notebook. The deployment team needs to see the model's version history and assign a production alias. Which action should the data scientist take?

A.Upload the model to a Vertex AI Feature Store entity type so the deployment team can retrieve it.
B.Create a Vertex AI Pipeline that retrains the model and send the pipeline template to the deployment team.
C.Register the model in Vertex AI Model Registry and share the model resource name with the deployment team.
D.Export the model artifacts to a Cloud Storage bucket and share the bucket path with the deployment team.
AnswerC

Vertex AI Model Registry stores model versions with metadata and supports aliases, so the deployment team can review history and assign a production alias before deploying to an endpoint. It is the intended handoff mechanism between experimentation and serving, satisfying both version tracking and alias requirements.

Why this answer

The handoff from experimentation to deployment is standardized through Vertex AI Model Registry, which versions models and supports aliases. Sharing raw artifacts or retraining pipelines does not provide the deployment team with a governed, aliasable model resource, and Feature Store is unrelated to model storage.

Exam trap

The trap here is assuming that a Cloud Storage path is an acceptable substitute for a registered model resource.

4
MCQeasy

A company uses Vertex AI Model Registry to manage multiple model versions. They want to designate a model version as 'champion' for production deployment and another as 'challenger' for A/B testing. Which feature of the registry should they use?

A.Model version labels
B.Model lineage
C.Model aliases
D.Model evaluation metrics
AnswerC

Model aliases are mutable, named pointers (for example 'champion', 'challenger') that reference a specific version within a registered model, so traffic can be switched between versions without redeploying or changing version IDs. This directly supports designating production and A/B testing versions.

Why this answer

Model aliases in Vertex AI Model Registry let you assign a mutable, named reference (e.g., 'champion', 'challenger') to a specific model version, so deployment endpoints can point to the alias rather than a fixed version ID. This enables seamless promotion or rollback by reassigning the alias, and supports A/B testing by directing traffic to different aliased versions. It is the intended feature for designating champion/challenger roles.

Exam trap

PMLE often tests the confusion between labels (static tags for filtering) and aliases (mutable pointers for deployment), leading candidates to choose labels for champion/challenger designation.

How to eliminate wrong answers

Option A (Model version labels) is wrong because labels are key-value tags for organization and filtering, not mutable pointers that endpoints can resolve to a specific version for deployment. Option B (Model lineage) is wrong because lineage tracks the provenance and relationships of artifacts, not deployment role assignment. Option D (Model evaluation metrics) is wrong because metrics describe model performance (e.g., AUC, accuracy) and do not control which version is served in production or testing.

5
MCQmedium

A team wants to share feature definitions across multiple projects in their organization using Vertex AI Feature Store. What is the recommended approach?

A.Export features to BigQuery datasets in each project
B.Use Vertex AI Feature Store's feature view for cross-project access
C.Create separate feature stores in each project and synchronize them with Dataflow
D.Use a centralized feature store in a shared project and grant access to other projects via IAM
AnswerD

A centralized feature store hosted in one shared project lets multiple projects consume identical feature definitions, with IAM policies granting cross-project access. This satisfies the requirement to share definitions organisation-wide while avoiding duplicated, divergent feature stores per project.

Why this answer

Vertex AI Feature Store is project-scoped, so to share features across projects, the recommended pattern is to create a centralized feature store in a shared (host) project and grant IAM roles to users/service accounts in other projects. This avoids duplication and ensures a single source of truth for feature definitions and online serving.

Exam trap

The trap is thinking feature views or BigQuery exports provide cross-project sharing; the correct pattern is a centralized feature store with IAM grants, which candidates often overlook in favor of data duplication.

How to eliminate wrong answers

Option A is wrong because exporting features to BigQuery datasets in each project duplicates data and does not provide online serving or feature freshness management — it is a batch workaround, not a feature store sharing mechanism. Option B is wrong because a feature view is a logical grouping within a feature store, not a cross-project access mechanism; feature views do not grant cross-project permissions by themselves. Option C is wrong because creating separate feature stores and synchronizing with Dataflow introduces complexity, latency, and consistency issues — it is an anti-pattern for sharing definitions.

6
MCQhard

A machine learning pipeline in Vertex AI produces a dataset artifact, a trained model, and evaluation metrics. The team wants to query the lineage to find all downstream artifacts that depend on a particular dataset. Which Vertex AI service should they use?

A.Vertex AI Feature Store
B.Vertex AI Experiments
C.Vertex AI Model Registry
D.Vertex AI Metadata
AnswerD

Vertex AI Metadata stores artefacts, executions and contexts in a lineage graph, so querying it returns every downstream artefact derived from a given dataset. This satisfies the stem's need to trace dependencies from a specific dataset artefact.

Why this answer

Vertex AI Metadata is the service that records and stores ML metadata — artifacts, executions, and contexts — and their relationships, forming the lineage graph. It lets you query upstream and downstream dependencies of any artifact, such as finding all models and metrics derived from a dataset. This is exactly the lineage-query capability the team needs.

Exam trap

PMLE often tests the overlap between Vertex AI Experiments (run tracking) and Vertex AI Metadata (lineage graph), so candidates pick Experiments when the question explicitly asks for upstream/downstream dependency queries.

How to eliminate wrong answers

Option A is wrong because Vertex AI Feature Store manages and serves feature values, not artifact lineage across pipeline runs. Option B is wrong because Vertex AI Experiments tracks runs, parameters, and metrics for comparison, but does not expose a queryable lineage graph of artifact dependencies. Option C is wrong because Vertex AI Model Registry catalogs and versions models for deployment, not the full dataset-to-model lineage graph.

7
MCQeasy

An ML team wants to automatically track training runs, including hyperparameters and metrics, with minimal code changes. Which Vertex AI service should they use?

A.Vertex AI Prediction
B.Vertex AI Workbench
C.Vertex AI Metadata
D.Vertex AI Experiments with autologging
AnswerD

Autologging hooks into supported frameworks (for example scikit-learn, TensorFlow, XGBoost) and records parameters, metrics and artifacts to an Experiment run automatically, satisfying the minimal-code-change constraint. Manual logging via the SDK would require explicit calls in every training script.

Why this answer

Vertex AI Experiments with autologging automatically captures parameters, metrics, and artifacts from training runs with minimal code changes—often just a single call to initialize and start a run. This directly satisfies the requirement to track training runs with minimal code. It integrates with the Vertex AI SDK to log framework-specific details (e.g., TensorFlow, PyTorch) automatically.

Exam trap

PMLE often tests the confusion between Vertex AI Experiments (high-level run tracking with autologging) and Vertex ML Metadata (low-level lineage store), causing candidates to choose Metadata when minimal-code tracking is required.

How to eliminate wrong answers

Option A (Vertex AI Prediction) is wrong because it is for serving predictions from deployed models, not tracking training runs. Option B (Vertex AI Workbench) is wrong because it is a managed notebook environment for development, not a run-tracking service. Option C (Vertex AI Metadata) is wrong because it is the underlying lineage store, but it does not provide the high-level autologging convenience for experiments; you would have to log metadata manually.

8
MCQmedium

An organization uses Vertex AI Pipelines and wants to track the lineage of datasets, models, and metrics across pipeline runs. They need to query upstream and downstream dependencies of an artifact. Which service should they use?

A.Vertex AI Feature Store
B.Vertex AI Experiments
C.Vertex AI Model Registry
D.Vertex AI Metadata
AnswerD

Vertex AI Metadata stores pipeline resources as a lineage graph of executions, artifacts and contexts, so you can query upstream and downstream dependencies of any artifact. Cloud Logging records events but holds no typed lineage relationships between datasets, models and metrics.

Why this answer

Vertex AI Metadata stores the ML metadata graph produced by Vertex AI Pipelines, including artifacts (datasets, models, metrics), executions, and their input/output relationships. It exposes APIs to traverse this graph in both directions, so you can query upstream sources and downstream dependents of any artifact. This is the correct service for cross-run lineage queries.

Exam trap

PMLE often tests the confusion between Experiments (which run produced these metrics) and Metadata (which artifacts depend on which), so candidates choose Experiments when the question asks for dependency traversal.

How to eliminate wrong answers

Option A is wrong because Vertex AI Feature Store serves and monitors feature values, not pipeline artifact lineage. Option B is wrong because Vertex AI Experiments compares training runs and their metrics but does not provide a queryable artifact dependency graph. Option C is wrong because Vertex AI Model Registry manages model versions and deployment stages, not the full dataset-to-metric lineage across pipeline runs.

9
MCQmedium

An ML team uses Vertex AI Pipelines to train and evaluate models. They want to ensure that only models meeting a minimum accuracy threshold are registered in Vertex AI Model Registry. Which approach should they take?

A.Register all models in Model Registry and use a separate Cloud Function triggered by a Model Registry event to delete models that do not meet the threshold.
B.Add a condition in the pipeline that checks the evaluation metric and only executes the ModelUploadOp if the metric meets the threshold.
C.Configure the ModelUploadOp to include the evaluation metric as a label, and rely on downstream consumers to filter models by label.
D.Use a Vertex AI Model Evaluation component to compute metrics, then manually review the metrics and register the model only if it passes.
AnswerB

Vertex AI Pipelines supports conditional execution using the condition parameter on tasks. By adding a condition that evaluates the model's accuracy metric, you can gate the ModelUploadOp so that only models meeting the threshold are uploaded to Model Registry. This enforces the quality gate directly in the pipeline without manual intervention.

Why this answer

The most reliable way to enforce a quality gate is to use a conditional execution in Vertex AI Pipelines that checks the evaluation metric and only runs the ModelUploadOp when the threshold is met. This prevents unqualified models from being registered and automates governance.

Exam trap

The trap here is assuming that post-registration filtering or manual review can enforce quality, when the pipeline itself should prevent registration of subpar models.

10
Multi-Selectmedium

A company wants to implement a centralized model registry for governance. Which two features should they use? (Choose two.)

Select 2 answers
A.Vertex AI Feature Store
B.Vertex AI Model Registry
C.Vertex AI Experiments
D.Model versioning and aliases
E.Vertex AI Metadata
AnswersB, D

Vertex AI Model Registry is the centralised catalogue that tracks model artefacts, versions and lineage across projects. It satisfies the governance requirement by giving one authoritative location to register, discover and manage models throughout their lifecycle.

Why this answer

Vertex AI Model Registry (B) is the centralized repository for managing the lifecycle of ML models, providing a single source of truth for governance across an organization, which directly matches the requirement for a centralized model registry. Model versioning and aliases (D) are core capabilities of the Model Registry that let teams track multiple model versions, assign meaningful aliases (e.g., 'production', 'staging'), and control which version is deployed, which is essential for governance and reproducibility. Together, B and D provide the registry plus the version/alias controls needed for centralized model governance.

Vertex AI Feature Store (A) manages feature storage and serving, not model registration, so it does not fulfill the model registry requirement. Vertex AI Experiments (C) tracks experiment runs, parameters, and metrics for training, not model governance or registration. Vertex AI Metadata (E) stores metadata about artifacts and executions in a lineage graph, but it is not the centralized model registry itself.

Exam trap

The trap here is confusing Vertex AI's data/experiment tracking services (Feature Store, Experiments, Metadata) with the model governance service (Model Registry) — candidates often pick Metadata because it sounds like the 'registry' layer.

11
MCQhard

A company uses Vertex AI Feature Store with an online store for low-latency serving. They observe high latency during peak hours. The feature values are small (< 1 KB each) and the workload is read-heavy. Which change would most effectively reduce latency?

A.Enable caching on the client side
B.Switch from Bigtable online store to Optimized online store
C.Use a larger machine type for Bigtable
D.Increase the number of Bigtable nodes
AnswerB

The optimized online store is purpose-built for low-latency, high-throughput serving of small feature values, so migrating from the Bigtable online store addresses the peak-hour latency directly. Increasing node count or reducing feature size does not change the underlying serving architecture.

Why this answer

Switching from Bigtable online store to Optimized online store is recommended for read-heavy workloads with small feature values, offering lower latency at high QPS.

12
MCQeasy

A junior engineer on your team has a trained scikit-learn model saved as a local joblib file and wants other teams to be able to discover it, view its evaluation metrics, and deploy it to a Vertex AI Endpoint. Which action should they take first?

A.Create a Vertex AI Endpoint and manually copy the joblib file onto the underlying prediction nodes.
B.Upload the model artifact to Vertex AI Model Registry with its serving container and metadata.
C.Copy the joblib file into a shared Cloud Storage bucket and email the object path.
D.Commit the joblib file to the team's Git repository and share the repository link.
AnswerB

Vertex AI Model Registry is the central catalog where models become discoverable, versioned, and deployable to endpoints. Importing the joblib artifact with the appropriate pre-built scikit-learn serving container and attaching evaluation metrics gives other teams visibility and a one-click path to deployment. This is the correct first step for sharing a model across teams.

Why this answer

To let other teams discover, evaluate, and deploy a model, it must be registered in a shared catalog. Vertex AI Model Registry stores the artifact, its serving container, and metadata such as evaluation metrics, and it integrates directly with endpoints for deployment. Git, raw Cloud Storage objects, and manual node manipulation all lack the cataloging, versioning, and deployment integration the scenario requires.

Exam trap

The trap here is treating a shared file location as sufficient for model sharing, when cross-team discovery and deployment require registration in a model catalog with serving metadata.

13
MCQeasy

A team wants to use Vertex AI Workbench for collaborative notebook development. They need a persistent environment that can be stopped and restarted without losing installed packages and data. Which instance type should they choose?

A.User-managed notebooks
B.Managed notebooks
C.Colab Enterprise notebooks
D.Vertex AI Pipelines
AnswerA

User-managed notebooks give a persistent Compute Engine instance with a customisable environment, so installed packages and data survive stopping and restarting. This satisfies the stem's requirement for persistence, unlike managed notebooks with ephemeral or containerised runtimes.

Why this answer

User-managed notebooks are Compute Engine instances where the user controls the environment, including installed packages, kernel state, and persistent disk contents. Because the underlying VM and its boot/data disks persist across stop and start operations, installed packages and data remain intact. This makes them the right choice for a persistent, customizable environment.

Exam trap

The trap is assuming 'managed' means 'better persistence' — in fact, user-managed notebooks give the user direct control over the persistent VM and disk, which is what preserves custom packages and data.

How to eliminate wrong answers

Option B is wrong because managed notebooks are designed for ephemeral, auto-managed environments with automatic idle shutdown and dependency management; while they persist some state, they are optimized for reproducibility rather than user-controlled persistence of arbitrary installed packages. Option C is wrong because Colab Enterprise notebooks are a managed, browser-based collaborative environment where the runtime is typically ephemeral and not intended for long-lived custom package persistence. Option D is wrong because Vertex AI Pipelines is an orchestration service for ML workflows, not a notebook environment at all.

14
MCQeasy

A machine learning team wants to share features across multiple models to reduce training-serving skew and ensure consistency. Which Vertex AI service should they use?

A.Vertex AI Workbench
B.Vertex AI Model Registry
C.Vertex AI Feature Store
D.Vertex AI Experiments
AnswerC

Vertex AI Feature Store provides a centralised repository where features are computed once and served identically to both training and prediction pipelines, directly eliminating training-serving skew. Sharing one feature definition across multiple models satisfies the consistency requirement, since online and batch serving draw from the same managed source rather than duplicated transformations.

Why this answer

Vertex AI Feature Store centralizes feature storage, ensuring the same features are used for training and serving, reducing training-serving skew.

15
Multi-Selecthard

A machine learning team needs to ensure that the same features used for training are used for serving in production to avoid training-serving skew. They use Vertex AI Feature Store. Which THREE actions should they take?

Select 3 answers
A.Enable point-in-time correct retrieval when creating training datasets
B.Use different feature views for training and serving to compare performance
C.Use the same feature view for both training data export and online serving
D.Export training data from the online store directly
E.Set up feature monitoring to detect drift in feature distributions
AnswersA, C, E

Point-in-time correct retrieval reconstructs the feature values that existed at each training label's timestamp, preventing future data from leaking into training. This directly satisfies the stem's requirement that training features match those available at serving time, eliminating training-serving skew.

Why this answer

Option A is correct because point-in-time correct retrieval ensures training datasets are built from the exact feature values that were available at the time of each event, preventing label leakage and keeping training data consistent with what would have been served historically. Option C is correct because using the same feature view for both training data export and online serving guarantees that the identical feature transformation and source are used in both paths, which is the core defense against training-serving skew. Option E is correct because feature monitoring detects drift in feature distributions between training and serving, surfacing skew or data quality issues so the team can remediate them.

Option B is incorrect because deliberately using different feature views for training and serving introduces skew rather than preventing it. Option D is incorrect because exporting training data directly from the online store is not the recommended pattern; training datasets should come from the offline store with point-in-time correctness, while the online store serves low-latency predictions.

16
MCQmedium

A data science team needs to share features across multiple ML models while ensuring consistency between training and serving. Which approach best achieves this?

A.Store features in a shared BigQuery dataset without versioning
B.Export features to CSV files shared via Cloud Storage
C.Use Vertex AI Feature Store to define and serve features for both training and online prediction
D.Each team maintains its own feature engineering code in separate pipelines
AnswerC

Vertex AI Feature Store provides a centralised repository where feature values are defined once and served consistently to both training jobs and online prediction, eliminating training-serving skew. This shared definition satisfies the consistency requirement while allowing multiple models to reuse the same features.

Why this answer

Vertex AI Feature Store provides a central repository where features are defined once and reused across models, reducing training-serving skew.

17
Multi-Selecteasy

A company wants to use DVC for data versioning alongside their ML code in Git. Which TWO statements about DVC are correct? (Select 2)

Select 2 answers
A.DVC uses a separate .dvc file to track data versions.
B.DVC can push data to remote storage like Google Cloud Storage.
C.DVC stores the actual data files in Git.
D.DVC only works with AWS S3 as remote storage.
E.DVC replaces Git for code versioning.
AnswersA, B

DVC stores a small .dvc metafile in Git containing the data's hash and path, while the actual dataset lives outside the repository. This lightweight pointer lets Git track dataset versions without bloating the repo, satisfying the requirement to version data alongside ML code.

Why this answer

Option A is correct because DVC creates small .dvc metafiles (e.g., data.dvc) that record the MD5 hash and path of the tracked data, allowing Git to version the pointer while the actual data lives in the DVC cache. Option B is correct because DVC supports many remote storage backends, including Google Cloud Storage, via `dvc remote add` and `dvc push`, so data can be stored outside the Git repository. Option C is wrong because DVC deliberately keeps large data files out of Git, storing them in its cache and remotes instead.

Option D is wrong because DVC supports S3, GCS, Azure Blob Storage, SSH, HDFS, and local remotes, not only S3. Option E is wrong because DVC complements Git for data versioning; it does not replace Git for source code version control.

Exam trap

The trap is assuming DVC stores data in Git or is AWS-only — candidates must remember DVC uses pointer files and supports multiple remotes, and that it augments rather than replaces Git.

18
Multi-Selectmedium

A regulated enterprise must prove to auditors that a specific production prediction can be traced to the exact model version, training data, and pipeline execution that produced it. Their ML workflows run on Vertex AI Pipelines. Which two practices should the team adopt? (Choose two.)

Select 2 answers
A.Log each prediction request and response to Cloud Logging and include the deployed model resource ID in the log entry.
B.Enable VPC Service Controls around the Vertex AI project to restrict data exfiltration.
C.Manually maintain a spreadsheet mapping model names to training dates and update it after each release.
D.Emit pipeline parameters, artifact URIs, and execution metadata so Vertex ML Metadata records the lineage graph automatically.
E.Store periodic screenshots of the Vertex AI console showing the deployed model and its metrics.
AnswersA, D

Capturing the model resource ID with each prediction creates the link from a specific prediction to a specific registered model version. Combined with lineage in Vertex ML Metadata, an auditor can follow that model version back to its training artifacts and pipeline execution, completing the end-to-end trace the regulation demands.

Why this answer

End-to-end traceability requires two linked records: automatic lineage of pipeline executions, artifacts, and models captured by Vertex ML Metadata, and a runtime record tying each prediction to the deployed model resource. Together they let an auditor walk from a prediction to a model version and onward to the training data and execution that created it. Screenshots, network controls, and manual spreadsheets provide no verifiable lineage.

Exam trap

The trap here is assuming that security controls or manual records constitute an audit trail, when traceability specifically requires machine-recorded lineage plus a prediction-to-model link.

19
MCQhard

A team is training a model using historical data and wants to avoid data leakage when joining feature values from a feature store. The features include time-varying data like user activity counts. Which retrieval method should they use when creating a training dataset?

A.Retrieve the latest feature values for each entity
B.Aggregate features over all historical data
C.Use random sampling of feature values
D.Use point-in-time correct retrieval with timestamp matching
AnswerD

Point-in-time correct retrieval with timestamp matching returns feature values as they existed at each training example's timestamp, preventing leakage from future user activity counts. Naive latest-value retrieval would expose post-event data, inflating offline metrics relative to production.

Why this answer

Point-in-time correct retrieval with timestamp matching ensures that for each training row, the feature values used are the ones that were actually available at the time of the label event, preventing future information from leaking into the training set. This is critical for time-varying features like user activity counts, where using the latest value would introduce look-ahead bias. By joining on entity ID and event timestamp, the feature store returns the feature value as of that timestamp, mimicking the production inference environment.

Exam trap

The trap here is confusing 'latest feature values' with 'correct feature values'—candidates often assume that using the most recent data is always best, but in training, it causes data leakage and inflated offline metrics that fail in production.

How to eliminate wrong answers

Option A is wrong because retrieving the latest feature values ignores the temporal relationship between features and labels, causing data leakage (future data used to predict past events). Option B is wrong because aggregating features over all historical data also incorporates future values relative to the label timestamp, and it destroys the point-in-time semantics needed for training. Option C is wrong because random sampling of feature values breaks the entity-timestamp association and introduces noise, not a valid retrieval strategy for feature stores.

20
Multi-Selectmedium

A company wants to monitor features in Vertex AI Feature Store for drift over time. Which two services should they use? (Choose two.)

Select 2 answers
A.Vertex AI Feature Store monitoring
B.Cloud Logging
C.Vertex AI Model Monitoring
D.Vertex AI Experiments
E.Cloud Monitoring
AnswersA, E

Vertex AI Feature Store monitoring computes drift and skew metrics natively against feature values, detecting distribution shifts over time without custom code. It satisfies the requirement to monitor features for drift directly within the feature store, providing scheduled drift detection and alerting on the stored feature data.

Why this answer

Option A, Vertex AI Feature Store monitoring, is correct because it is the native capability that computes feature-level drift and skew statistics for features served from a Vertex AI Feature Store, generating metrics and alerts when distributions shift over time. Option E, Cloud Monitoring, is correct because the drift metrics produced by Feature Store monitoring are exported to Cloud Monitoring, where they can be visualized on dashboards and used to create alerting policies. Option B, Cloud Logging, is not the right choice because it captures log entries rather than computing or charting feature drift metrics.

Option C, Vertex AI Model Monitoring, targets model prediction drift/skew on deployed endpoints, not feature drift within a Feature Store. Option D, Vertex AI Experiments, is for tracking training runs, parameters, and metrics, not for ongoing drift monitoring of stored features.

Exam trap

PMLE often tests the distinction between Feature Store monitoring (feature-level drift) and Model Monitoring (prediction-level drift/skew), so candidates who conflate the two pick Vertex AI Model Monitoring.

21
MCQmedium

An ML team wants to share feature definitions across multiple projects to reduce training-serving skew and ensure consistency. They currently store features in Cloud Storage and manually coordinate updates, leading to errors. Which Google Cloud service should they use to centrally manage and serve features for both training and online inference?

A.Cloud Data Catalog
B.Vertex AI Model Registry
C.Vertex AI Feature Store
D.Cloud Storage with versioning
AnswerC

Vertex AI Feature Store provides a central registry where feature definitions are authored once and served consistently to both training and online inference, removing manual Cloud Storage coordination errors. It also supplies point-in-time correctness, which prevents training-serving skew.

Why this answer

Vertex AI Feature Store is purpose-built to centrally define, store, and serve ML features for both training (batch) and online inference (low-latency) with a consistent feature definition, which directly eliminates training-serving skew. It provides a single source of truth so teams across projects can share features without manual coordination.

Exam trap

PMLE often tests whether candidates confuse metadata (Data Catalog), model (Model Registry), and feature (Feature Store) services — the trap is picking a storage or catalog service when the requirement is centralized feature serving with consistency.

How to eliminate wrong answers

Option A is wrong because Cloud Data Catalog is a metadata management service for discovery and governance, not a feature serving system — it cannot serve features for online inference. Option B is wrong because Vertex AI Model Registry stores and versions models, not features; it has no feature-serving capability. Option D is wrong because Cloud Storage with versioning is just object storage — it can hold feature data but provides no online serving, no point-in-time correctness, and no centralized feature definition, so it does not solve the coordination problem.

22
MCQhard

A team uses Vertex AI Feature Store with an online store for low-latency serving. They need to support frequent updates to features (e.g., every minute) and require high write throughput (thousands of writes per second). Which online store type should they choose?

A.Optimized online store
B.Firestore online store
C.Bigtable online store
D.Cloud SQL online store
AnswerC

Bigtable online store supports high write throughput and frequent feature updates, scaling to thousands of writes per second with low-latency reads. This satisfies the stated requirement for minute-level updates and high-throughput ingestion that the default online store cannot match.

Why this answer

Bigtable online store is designed for high-throughput, low-latency serving with frequent updates, making it suitable for thousands of writes per second and minute-level feature refreshes. It leverages Bigtable's scalable NoSQL architecture, which handles high write loads efficiently. Optimized online store is for low-latency but may not sustain such high write throughput, while Firestore and Cloud SQL have lower write limits.

Exam trap

The trap is assuming that any low-latency online store can handle high write throughput; candidates must distinguish between read-optimized and write-optimized stores, with Bigtable being the only one designed for massive write scalability.

How to eliminate wrong answers

Option A is wrong because the optimized online store, while low-latency, is not optimized for high write throughput and may throttle under thousands of writes per second. Option B is wrong because Firestore online store is designed for smaller-scale, lower-throughput applications and has write limits that cannot handle thousands of writes per second. Option D is wrong because Cloud SQL is a relational database with lower write scalability and is not intended for high-throughput feature serving.

23
MCQeasy

An ML team wants to monitor feature drift in their production model. Which Vertex AI Feature Store capability should they use?

A.Feature views
B.Online store
C.Point-in-time retrieval
D.Feature monitoring (drift detection)
AnswerD

Feature monitoring in Vertex AI Feature Store continuously computes drift metrics by comparing production feature distributions against a baseline snapshot, directly satisfying the requirement to monitor feature drift. It detects skew and drift per feature, emitting alerts without retraining, which is precisely the capability the stem requests.

Why this answer

Feature monitoring (drift detection) is the Vertex AI Feature Store capability that computes drift and skew metrics for feature values against a baseline and emits them for alerting. It is purpose-built to detect when production feature distributions diverge from training or reference distributions. This directly answers the requirement to monitor feature drift.

Exam trap

PMLE often tests the confusion between feature views (serving constructs) and feature monitoring (drift detection), so candidates pick feature views when the question asks specifically about drift.

How to eliminate wrong answers

Option A is wrong because feature views are the logical groupings that serve features for online/offline use, not the drift-detection mechanism itself. Option B is wrong because the online store is the low-latency serving layer for real-time feature retrieval, not a monitoring component. Option C is wrong because point-in-time retrieval ensures training data uses historically correct feature values to avoid leakage, which is a correctness feature, not drift monitoring.

24
MCQeasy

An ML team uses Vertex AI Pipelines and wants to automatically generate model cards documenting model purpose, evaluation results, and intended use. Which approach should they take?

A.Manually create a Google Doc and share it with the team.
B.Use Cloud Data Catalog to annotate the model artifact.
C.Write a custom Kubeflow Pipelines component that creates a BigQuery table with model metadata.
D.Use Vertex AI Model Registry to generate model cards automatically from model metadata.
AnswerD

Vertex AI Model Registry generates model cards from registered model metadata, including purpose, evaluation results and intended use, when the model is uploaded with those details. This automates documentation from the pipeline's existing metadata rather than requiring manual authoring.

Why this answer

Vertex AI Model Registry automatically generates model cards from model metadata, including purpose, evaluation results, and intended use, when models are registered and their metadata is populated. This provides a native, integrated way to document models without manual effort or custom components. It aligns with Vertex AI's governance and documentation features.

Exam trap

PMLE often tests the distinction between metadata storage and model card generation — candidates pick Cloud Data Catalog or custom BigQuery tables because they store metadata, but only Vertex AI Model Registry generates the structured model card.

How to eliminate wrong answers

Option A is wrong because manually creating a Google Doc is not automated, not integrated with Vertex AI Pipelines, and does not generate model cards from model metadata. Option B is wrong because Cloud Data Catalog is a metadata management service for discovery and governance, not a model card generator; annotating a model artifact does not produce a model card. Option C is wrong because writing a custom Kubeflow component to create a BigQuery table stores metadata but does not generate the structured model card format that Vertex AI Model Registry provides.

25
MCQhard

Two teams train models in separate Vertex AI projects but must share the same curated feature set. The platform team wants a single authoritative definition of each feature so that online serving and offline training always return consistent values, while each team keeps its own model training pipeline. Which approach should the platform team implement?

A.Publish the features as a BigQuery authorized view and instruct teams to call it from their pipelines.
B.Create the features once in a Vertex AI Feature Group with a BigQuery source and register them in a shared Feature Online Store, then grant both teams read access.
C.Have each team copy the feature engineering SQL into its own project and schedule it independently.
D.Export the feature table to a shared Cloud Storage bucket as Parquet and let both teams read the files.
AnswerB

Vertex AI Feature Registry and Feature Groups let the platform team define each feature once over a BigQuery source, and a Feature Online Store serves the same definitions online. Granting both teams read access means their separate pipelines consume identical feature definitions, which is exactly the consistency and single-source-of-truth requirement described.

Why this answer

A shared, authoritative feature definition is what prevents training-serving skew across teams. Vertex AI Feature Groups in Feature Registry define features once over a BigQuery source, and a Feature Online Store serves those same definitions at low latency, so both teams' pipelines consume identical feature semantics while retaining their own training pipelines. Copies, static Parquet exports, and authorized views all lack the managed online serving and registry linkage.

Exam trap

The trap here is equating shared data access with shared feature definitions, when access control mechanisms like authorized views do not provide a managed online store or consistent training-serving values.

26
MCQeasy

A data scientist wants to automatically generate model documentation that includes model purpose, training data, evaluation results, and intended use. Which tool should they use?

A.Vertex AI Workbench
B.Vertex AI Experiment
C.Cloud Datalab
D.Model Cards in Vertex AI
AnswerD

Model Cards in Vertex AI automatically capture and display model purpose, training data, evaluation metrics, and intended use, directly satisfying the documentation requirement. Unlike generic metadata or manual reports, Model Cards generate this structured governance artefact from the model's own evaluation results, giving the data scientist the exact fields the stem demands.

Why this answer

Model Cards in Vertex AI is a feature specifically designed to generate structured documentation for models, including purpose, training data, evaluation results, and intended use. It provides a standardized format for model transparency and governance. The other tools are for development, experimentation, or notebooks, not documentation.

Exam trap

The trap is confusing experiment tracking with model documentation; candidates might pick Vertex AI Experiments because it deals with models, but Model Cards is the dedicated tool for documentation.

How to eliminate wrong answers

Option A is wrong because Vertex AI Workbench is a Jupyter notebook environment for development, not a documentation tool. Option B is wrong because Vertex AI Experiment is for tracking and comparing ML experiments, not for generating model documentation. Option C is wrong because Cloud Datalab is an interactive data exploration tool, not a model documentation service.

27
MCQmedium

Your team trains models in a shared Vertex AI project. A data engineer accidentally overwrites a BigQuery training table that three production pipelines depend on, and nobody can tell which pipeline used which version of the data. You need to make dataset versions immutable and traceable so that any training run can be reproduced. What should you do?

A.Copy the table nightly into a Cloud Storage bucket using a scheduled Dataproc job.
B.Enable BigQuery table snapshots and record the snapshot ID in each pipeline's run metadata.
C.Enable BigQuery time travel and query the table with a FOR SYSTEM_TIME AS OF clause during retraining.
D.Grant the data engineer only bigquery.dataViewer on the table so they can no longer modify it.
AnswerB

BigQuery table snapshots create immutable, point-in-time copies of a table that persist independently of later writes, so a training job can always be reproduced from the exact snapshot it consumed. Recording the snapshot ID in pipeline metadata ties the artifact to the run, giving the traceability the team currently lacks without duplicating data or blocking the engineer's writes.

Why this answer

Immutable dataset versions are needed so any training run can be reproduced exactly, even after the source table changes. BigQuery table snapshots provide durable, read-only point-in-time copies, and storing the snapshot identifier alongside the pipeline run metadata creates the audit trail the team is missing. Access control, nightly copies, and time travel all fail to combine immutability with explicit per-run traceability.

Exam trap

The trap here is assuming that restricting write permissions or copying data on a schedule provides reproducibility, when only an immutable version identifier recorded per run actually ties a training job to its exact input data.

28
MCQeasy

A data scientist wants to track machine learning experiments, including parameters, metrics, and artifacts, and compare runs. Which Vertex AI service should they use?

A.Vertex AI Metadata
B.Vertex AI Experiments
C.Vertex AI Feature Store
D.Vertex AI Model Registry
AnswerB

Vertex AI Experiments logs parameters, metrics, and artifacts for each run and provides comparison across runs within a experiment. This directly satisfies the need to track training metadata and evaluate multiple runs side by side.

Why this answer

Vertex AI Experiments is the service designed to track ML experiments, including parameters, metrics, and artifacts, and to compare runs. It integrates with Vertex AI Metadata to log experiment lineage and provides a UI for comparing runs. This directly matches the data scientist's requirement.

Exam trap

PMLE often tests the overlap between Vertex AI Experiments and Vertex AI Metadata — candidates may pick Metadata, but Experiments is the user-facing service for tracking and comparing runs.

How to eliminate wrong answers

Option A is wrong because Vertex AI Metadata stores metadata about artifacts and executions but is not the primary service for tracking and comparing experiment runs. Option C is wrong because Vertex AI Feature Store manages and serves ML features, not experiment tracking. Option D is wrong because Vertex AI Model Registry manages model versions and deployment, not experiment parameters and metrics.

29
MCQmedium

A data science team wants to share a set of engineered features across multiple projects and teams to reduce training-serving skew and ensure consistency. They need low-latency serving (single-digit milliseconds) for online predictions and also need to retrieve historical feature values for training. Which approach should they take?

A.Use Vertex AI Feature Store to define features once, serve online predictions from the online store, and retrieve historical features from the offline store for training.
B.Create a shared BigQuery dataset where each team writes features; serve predictions by querying BigQuery synchronously.
C.Store features in Cloud Storage Parquet files and load them into BigQuery for training; serve predictions from a custom microservice that reads from Cloud Storage.
D.Use Cloud Memorystore (Redis) to store the latest feature values for low-latency serving; each team independently computes features and pushes to Redis.
AnswerA

Feature Store decouples feature engineering from consumption: one definition feeds both the low-latency online store for serving and the offline store for point-in-time historical retrieval, eliminating training-serving skew. Reusing a Cloud Storage copy cannot meet single-digit millisecond online latency.

Why this answer

Vertex AI Feature Store is a fully managed, purpose-built feature store that centralizes feature definitions and provides both an online store for low-latency serving and an offline store for historical retrieval. The online store is optimized for single-digit millisecond reads, while the offline store (backed by BigQuery) allows point-in-time correct feature retrieval for training, which directly addresses training-serving skew. By defining features once and reusing them across projects, it ensures consistency between training and serving.

Exam trap

PMLE often tests the misconception that a data warehouse like BigQuery can serve online predictions with low latency, or that a simple key-value store like Redis suffices for both online and offline needs, ignoring the need for point-in-time correct historical retrieval and centralized feature definitions.

How to eliminate wrong answers

Option B is wrong because BigQuery is not designed for synchronous low-latency online serving; querying it per prediction would introduce hundreds of milliseconds to seconds of latency and lacks a built-in online store. Option C is wrong because Cloud Storage is not a low-latency serving layer, and a custom microservice reading Parquet files would not provide the required single-digit millisecond performance or point-in-time correctness for training. Option D is wrong because Redis alone does not provide historical feature retrieval for training, and independent feature computation by each team reintroduces training-serving skew and lacks centralized governance.

30
MCQeasy

Which Vertex AI service is used to track the lineage of ML pipeline components, artefacts, and executions?

A.Vertex AI Metadata
B.Vertex AI Model Registry
C.Vertex AI Feature Store
D.Vertex AI Experiments
AnswerA

Vertex AI Metadata provides the lineage tracking required, storing artefacts, executions and contexts as nodes within a managed metadata graph. It records relationships between pipeline components and their outputs, satisfying the stem's demand for tracing component, artefact and execution provenance across ML workflows.

Why this answer

Vertex AI Metadata is the service that tracks the lineage of ML pipeline components, artefacts, and executions. It provides a centralized repository to store and manage metadata about ML workflows, enabling reproducibility, auditing, and collaboration. By recording relationships between components, artefacts, and executions, it allows you to trace the provenance of models and data.

Exam trap

PMLE often tests the distinction between Vertex AI services that sound similar, such as confusing Metadata with Experiments or Model Registry, because all deal with tracking aspects of ML workflows.

How to eliminate wrong answers

Option B is wrong because Vertex AI Model Registry is a central repository for managing the lifecycle of ML models, including versioning and deployment, but it does not track pipeline component lineage. Option C is wrong because Vertex AI Feature Store is a centralized repository for organizing, storing, and serving ML features, focusing on feature management rather than lineage tracking. Option D is wrong because Vertex AI Experiments helps track and compare experiment runs, including parameters and metrics, but it does not provide comprehensive lineage tracking for pipeline components and artefacts.

31
MCQmedium

A data engineer needs to version large datasets (multiple TB) in a Data Lake on Google Cloud. They require ACID transactions to ensure consistency when multiple jobs read/write concurrently. Which solution should they use?

A.Delta Lake on Dataproc
B.BigQuery table snapshots
C.DVC (Data Version Control)
D.Vertex AI Feature Store
AnswerA

Delta Lake provides ACID transactions and scalable metadata handling over Parquet files in Cloud Storage, letting concurrent Dataproc jobs read and write multi-terabyte datasets consistently. Its transaction log delivers the snapshot isolation and versioning the stem demands.

Why this answer

Delta Lake on Dataproc provides ACID transactions on cloud storage data lakes, enabling concurrent reads/writes with consistency.

32
Multi-Selectmedium

A data science team collaborates using Vertex AI Workbench user-managed notebooks. They want to version control their notebook code and share it with team members. Which TWO tools should they use? (Choose 2)

Select 2 answers
A.Git integration in Vertex AI Workbench
B.Cloud Functions
C.Vertex AI Model Registry
D.Vertex AI Experiments
E.Cloud Source Repositories
AnswersA, E

Git integration in Vertex AI Workbench lets users commit notebook code to a remote repository, track revisions and share branches with teammates. This satisfies the version-control and collaboration requirement without leaving the managed notebook environment.

Why this answer

Option A, Git integration in Vertex AI Workbench, is correct because user-managed notebooks include built-in Git support, allowing data scientists to clone repositories, commit, push, and pull notebook code directly from the JupyterLab interface for version control and collaboration. Option E, Cloud Source Repositories, is correct because it is a fully managed private Git repository service on Google Cloud that can host the team's notebook code and integrate with Workbench's Git tooling for sharing and versioning. Option B, Cloud Functions, is incorrect because it is a serverless compute service for event-driven functions, not a version control or code-sharing tool.

Option C, Vertex AI Model Registry, is incorrect because it manages trained model versions and their metadata, not notebook source code. Option D, Vertex AI Experiments, is incorrect because it tracks experiment runs, parameters, and metrics, not Git-based notebook version control.

Exam trap

The trap is selecting other Vertex AI components like Model Registry or Experiments, which sound related to ML workflows but are not for source code version control — candidates must distinguish between code versioning and model/experiment tracking.

33
MCQmedium

Your team trains a scikit-learn model locally and uploads it to Vertex AI Model Registry. A colleague needs to deploy it to a Vertex AI Endpoint for online prediction with a prebuilt container. The model artifacts are stored in a Cloud Storage bucket. Which deployment approach should they use?

A.Create a custom prediction routine in a Python source distribution, upload it as a model artifact, and deploy it without specifying a serving container.
B.Export the model to a SavedModel format, import it with a TensorFlow prebuilt container, and deploy it to an Endpoint.
C.Import the model with the prebuilt scikit-learn container image and deploy it to an Endpoint using the Model Registry UI or the gcloud ai endpoints deploy-model command.
D.Upload the model artifact to Vertex AI Model Registry as a BigQuery ML model and deploy it using the BigQuery ML serving container.
AnswerC

This is correct because Vertex AI Model Registry supports importing scikit-learn models with a prebuilt serving container, and the imported model can be deployed directly to an Endpoint. The prebuilt scikit-learn container handles prediction requests without requiring custom inference code, and deployment can be done through the console or gcloud CLI. This matches the standard workflow for deploying a locally trained scikit-learn model.

Why this answer

The scikit-learn model was trained locally and stored in Cloud Storage, so it should be imported into Vertex AI Model Registry using the prebuilt scikit-learn container. This allows direct deployment to a Vertex AI Endpoint for online prediction. The prebuilt container removes the need for custom inference code and supports the standard deployment workflow via console or gcloud CLI.

Exam trap

The trap here is assuming that a locally trained scikit-learn model must be converted to TensorFlow SavedModel or packaged with custom code before it can be deployed to a Vertex AI Endpoint.

34
MCQmedium

A team uses Vertex AI Pipelines with a custom training component that reads data from a BigQuery table. They need to ensure that a new pipeline run uses a specific snapshot of the training data for reproducibility. Which approach should they take?

A.Create a BigQuery table snapshot of the training data and pass the snapshot name to the training component as an input artifact.
B.Pass the BigQuery table name and a WHERE clause that filters on a timestamp column to the training component, and record the query in the pipeline parameters.
C.Export the BigQuery table to a Cloud Storage bucket in CSV format and pass the Cloud Storage URI to the training component.
D.Use BigQuery's time travel feature by adding FOR SYSTEM_TIME AS OF a fixed timestamp to the query in the training component.
AnswerA

A BigQuery table snapshot is an immutable, point-in-time copy of the table that persists even if the source table changes. By creating a snapshot and passing its name to the training component, you ensure the pipeline run always reads the same data. This provides reproducibility and aligns with Vertex AI Pipelines artifact tracking, as the snapshot can be recorded as an input artifact.

Why this answer

To ensure reproducibility, the training data must be immutable. A BigQuery table snapshot provides a point-in-time copy that does not change even if the source table is updated. Passing the snapshot name to the training component as an input artifact also integrates with Vertex AI Pipelines lineage tracking, making the data reference explicit and stable.

Exam trap

The trap here is relying on time travel or filtered queries for reproducibility, which are not durable or immutable beyond their retention limits.

35
Multi-Selectmedium

Your team is preparing to hand a trained model to a separate operations team that will deploy it to a Vertex AI endpoint. The operations team needs to understand the model's input schema, the training run that produced it, and which alias currently points to production. Which two Vertex AI resources should you share with them to provide this information? (Choose two.)

Select 2 answers
A.The Vertex AI Pipeline job resource that was used to schedule nightly retraining.
B.The Cloud Storage bucket containing the raw training data.
C.The registered model resource in Vertex AI Model Registry, including its version and alias.
D.The Vertex AI Feature Store online store resource used during training.
E.The Vertex AI Experiments run that logged the training parameters, metrics, and input schema.
AnswersC, E

The Model Registry resource holds the model version, its metadata, and aliases such as 'production'. Sharing this resource name lets the operations team see exactly which version is aliased for production and deploy it, satisfying the alias and versioning part of the handoff.

Why this answer

A complete handoff to operations requires the registered model resource for version and alias information, and the training run record for parameters, metrics, and input schema. Feature stores, pipeline jobs, and raw data buckets serve different purposes and do not provide the deployment team with the model-specific governance details they need.

Exam trap

The trap here is equating data access with model handoff, assuming the operations team needs the raw training data rather than the model's schema and version metadata.

36
MCQmedium

You are using DVC for data versioning in an ML project on Google Cloud. Your training data is stored in Cloud Storage. You want to track a new version of the dataset after preprocessing. Which DVC command should you use to register the changes?

A.dvc add data/processed
B.dvc push
C.dvc run -n preprocess
D.dvc commit
AnswerA

dvc add computes the hash of the processed directory, writes a .dvc file and updates .gitignore, registering the new dataset version for tracking. dvc push only uploads already-tracked data to remote storage; it does not register changes.

Why this answer

The `dvc add` command registers a file or directory with DVC, creating a .dvc metafile that captures the hash and metadata of the data, and adds the actual data to the DVC cache. After preprocessing produces a new dataset in data/processed, running `dvc add data/processed` creates the versioned pointer that DVC tracks in Git, which is exactly what 'registering the changes' means.

Exam trap

PMLE often tests the confusion between `dvc add` (register a new dataset/file) and `dvc commit` (update cache for existing stage outputs) — candidates pick `dvc commit` when the question asks about registering a brand-new dataset.

How to eliminate wrong answers

Option B is wrong because `dvc push` uploads cached data to remote storage (e.g., Cloud Storage) — it transfers data but does not create or update the .dvc metafile that registers a new version. Option C is wrong because `dvc run -n preprocess` (or `dvc stage add`) defines a pipeline stage that executes a command and tracks its outputs; it is used to create the preprocessing step, not to register an already-produced dataset. Option D is wrong because `dvc commit` writes the current state of tracked outputs to the DVC cache without re-running the pipeline — it is used after manual edits to outputs of an existing stage, not for adding a brand-new dataset to DVC tracking.

37
Multi-Selectmedium

A data science team uses Vertex AI Experiments to compare multiple model training runs. They want to capture and compare hyperparameters, metrics, and code versions for each run. Which TWO steps should they take?

Select 2 answers
A.Use Cloud Logging to capture all training outputs
B.Store code versions in Cloud Storage and link them to experiments manually
C.Log hyperparameters and metrics using the Vertex AI SDK's experiment logging functions
D.Export experiment data to BigQuery for comparison
E.Integrate the training code with Git and use the commit hash as a run parameter
AnswersC, E

The Vertex AI SDK's experiment logging functions (aiplatform.log_params and log_metrics) attach hyperparameters and metrics to a named experiment run, which is exactly what enables side-by-side comparison of runs in the Vertex AI Experiments console. Code versions are captured separately via Git.

Why this answer

Option C is correct because the Vertex AI SDK provides dedicated experiment logging functions (e.g., aiplatform.log_params() and aiplatform.log_metrics()) that record hyperparameters and metrics directly into a Vertex AI Experiment run, which is exactly what the team needs to capture and compare across runs. Option E is correct because integrating training code with Git and passing the commit hash as a run parameter ties each experiment run to a specific, reproducible code version, satisfying the requirement to capture code versions alongside hyperparameters and metrics. Option A is not appropriate because Cloud Logging captures log output, not structured experiment parameters or metrics for comparison in Vertex AI Experiments.

Option B is unnecessary and less precise than using Git commit hashes, since manually linking Cloud Storage artifacts does not automatically associate code versions with runs. Option D is not required because Vertex AI Experiments already provides comparison capabilities natively, so exporting to BigQuery is an extra step not needed for the stated goal.

Exam trap

PMLE often tests the misconception that Cloud Logging or BigQuery export are part of the experiment tracking workflow, when in fact Vertex AI Experiments requires explicit SDK logging and manual code version parameterization.

38
MCQmedium

An ML engineer needs to deploy a model to an endpoint and gradually shift traffic from the previous version (champion) to a new version (challenger) for A/B testing. How should they configure the endpoint?

A.Use a canary deployment with Cloud Run
B.Manually update the endpoint to point to the challenger after testing
C.Create a new endpoint for the challenger and route traffic via load balancer
D.Deploy both versions to the same endpoint and set traffic splitting
AnswerD

Deploying both versions to one endpoint with traffic splitting lets the engineer route a defined percentage to the challenger while the champion serves the remainder, enabling gradual A/B comparison. Separate endpoints would not permit proportional traffic distribution between versions.

Why this answer

To gradually shift traffic between a champion and challenger model for A/B testing, the engineer should deploy both versions to the same endpoint and configure traffic splitting. This allows the endpoint to route a percentage of requests to each version, enabling controlled experimentation and rollback. This is the standard pattern for A/B testing on managed ML platforms.

Exam trap

PMLE often tests deployment strategies, and candidates may confuse infrastructure-level canary deployments (e.g., Cloud Run) with model-level traffic splitting, or they may choose manual switching, missing the need for gradual, controlled experimentation.

How to eliminate wrong answers

Option A is wrong because Cloud Run is a container platform, not an ML endpoint service, and canary deployment there does not provide model-level traffic splitting. Option B is wrong because manually updating the endpoint to point to the challenger after testing does not allow gradual traffic shift or A/B testing; it is an all-or-nothing switch. Option C is wrong because creating a new endpoint and routing via load balancer adds complexity and does not use the native traffic splitting feature of the ML platform.

39
MCQhard

A company uses Vertex AI Feature Store for feature engineering. They need to ensure point-in-time correctness to avoid data leakage during training. Which feature retrieval method should they use?

A.Use the `get_features` API without specifying a timestamp.
B.Use BigQuery to manually join features with a sliding window.
C.Use the offline store with point-in-time join using the `feature_view` with a timestamp column.
D.Use the online store to retrieve the latest feature values.
AnswerC

The offline store performs point-in-time joins, matching each training example's timestamp against feature values valid at that moment. This guarantees the model only sees historically accurate features, eliminating the temporal leakage the scenario requires avoiding during training.

Why this answer

Point-in-time correctness requires retrieving feature values as they existed at the timestamp of each training example, which is exactly what the offline store's point-in-time join does when a feature_view is configured with an event/timestamp column. Vertex AI Feature Store uses this timestamp column to perform an as-of join, preventing future data from leaking into the training row. The online store and timestamp-less get_features calls return only the latest values, which is the classic source of label leakage.

Exam trap

PMLE often tests the distinction between online (latest-value, low-latency) and offline (historical, point-in-time) feature retrieval, tricking candidates into choosing the online store because it sounds more 'real-time' and therefore more accurate.

How to eliminate wrong answers

Option A is wrong because calling get_features without a timestamp returns the most recent feature values, which leaks future information into historical training rows. Option B is wrong because manually joining in BigQuery with a sliding window is error-prone, not the Feature Store-native mechanism, and does not leverage the feature_view's point-in-time semantics. Option D is wrong because the online store is optimized for low-latency serving of the latest feature values, not for historical as-of retrieval during training.

40
MCQhard

A team monitors features in Vertex AI Feature Store for drift. They want to set up automated alerts when a feature's distribution deviates significantly from the baseline. Which feature monitoring configuration should they use?

A.Enable feature monitoring on the feature group with drift threshold and notification channel.
B.Use Cloud Monitoring custom metrics and log-based alerts manually.
C.Use Vertex AI Experiments to compare distributions.
D.Export features to BigQuery and set up scheduled queries with alerts.
AnswerA

Feature group monitoring computes drift statistics against a baseline distribution and triggers alerts via a configured notification channel when the drift threshold is exceeded. Enabling it at the feature group level with both parameters satisfies the automated-alert requirement without custom pipeline code.

Why this answer

Feature monitoring in Vertex AI Feature Store allows defining drift thresholds and alerting via Cloud Monitoring.

41
MCQmedium

Your team owns a Vertex AI Model Registry entry that other teams depend on for production serving. A new retrained model shows better offline metrics, and you need to roll it out gradually to a small percentage of live traffic while keeping the ability to revert instantly if quality degrades. What should you do?

A.Create a second endpoint for the new model and update the client application to call both endpoints.
B.Register the new model under a new model resource and delete the previous version to avoid ambiguity.
C.Assign the new model version the default alias in Model Registry and redeploy the endpoint from that alias.
D.Deploy the new model version to the existing endpoint with a traffic split, then shift the split percentage as confidence grows.
AnswerD

Vertex AI Endpoints support deploying multiple model versions and splitting prediction traffic by percentage. Deploying the new version alongside the current one and starting with a small split enables a controlled canary rollout, and because the old version remains deployed, reverting is simply resetting the split back to the previous version.

Why this answer

Gradual rollout with instant rollback on Vertex AI means deploying both model versions to the same endpoint and controlling the percentage of traffic each receives. Starting with a small split limits blast radius, and because the prior version stays deployed, reverting is a split change rather than a redeployment. Alias promotion, duplicate endpoints, and deleting the old version all fail to provide controlled, reversible traffic shifting.

Exam trap

The trap here is confusing version promotion via aliases with traffic management, when only an endpoint traffic split gives gradual exposure and instant rollback.

42
Multi-Selectmedium

A team is using Vertex AI Model Registry to manage models. They need to ensure that when a new model version is registered, it is automatically evaluated for fairness and bias before being deployed. Which two Google Cloud services should they integrate to achieve this? (Choose two.)

Select 2 answers
A.Vertex AI Pipelines
B.Vertex AI Feature Store
C.Vertex AI Model Evaluation
D.Cloud Build
E.Vertex AI Model Monitoring
AnswersA, C

Vertex AI Pipelines can orchestrate a workflow that triggers upon model registration, running evaluation steps such as fairness and bias checks. It allows integration with other services and custom code, enabling automated pre-deployment assessment. By using pipelines, the team can enforce that no model is deployed without passing the fairness evaluation.

Why this answer

Vertex AI Pipelines can orchestrate an automated workflow triggered by model registration, and Vertex AI Model Evaluation provides the fairness and bias metrics. Together, they enable pre-deployment evaluation. Other services like Model Monitoring and Feature Store focus on different aspects and cannot perform fairness assessment at registration time.

Exam trap

The trap here is confusing model monitoring with model evaluation; monitoring detects drift in deployed models, while evaluation computes fairness metrics for a model version.

43
MCQhard

You need to create a reproducible snapshot of a BigQuery table as of a specific timestamp for ML model training. The snapshot should be queryable without copying the entire dataset. Which BigQuery feature should you use?

A.BigQuery export to Cloud Storage as Parquet
B.BigQuery time travel (FOR SYSTEM_TIME AS OF)
C.CREATE TABLE AS SELECT with WHERE clause
D.BigQuery table snapshots
AnswerD

BigQuery table snapshots capture a table's contents at a specified timestamp as a lightweight, queryable reference that shares storage with the base table rather than duplicating data. This satisfies the stem's reproducibility and no-full-copy constraints.

Why this answer

BigQuery table snapshots create a lightweight, queryable copy of a table at a point in time using copy-on-write storage — only changed data consumes additional bytes, and the snapshot is a first-class table you can query directly. This satisfies the requirement for a reproducible, timestamped, queryable snapshot without duplicating the full dataset. Time travel, by contrast, is a query-time window (default 7 days) and does not persist as a durable object.

Exam trap

The trap here is conflating time travel (a transient query window) with table snapshots (a durable, queryable object) — candidates often pick FOR SYSTEM_TIME AS OF because it sounds like a snapshot, but it does not persist beyond the time-travel window.

How to eliminate wrong answers

Option A is wrong because exporting to Cloud Storage as Parquet creates a physical copy of the data outside BigQuery, which is neither queryable in-place nor storage-efficient. Option B is wrong because FOR SYSTEM_TIME AS OF is a query modifier limited to the time-travel window (default 7 days, max 7 days unless configured), and it does not create a persistent snapshot object. Option C is wrong because CREATE TABLE AS SELECT physically materializes a full copy of the data, incurring full storage cost and defeating the 'without copying the entire dataset' requirement.

44
MCQmedium

You are setting up feature monitoring in Vertex AI Feature Store to detect drift in a numerical feature. The monitoring job should run daily and alert if the Jensen-Shannon divergence exceeds 0.1. Which configuration should you use?

A.Configure feature monitoring in the feature view with a drift threshold of 0.1 using Jensen-Shannon divergence
B.Use BigQuery scheduled queries to compare distributions and send alerts
C.Set up a Cloud Composer DAG to compute drift and publish to Cloud Monitoring
D.Enable model monitoring on the Vertex AI endpoint to detect drift
AnswerA

Configuring monitoring on the feature view applies drift detection directly to the served feature data, satisfying the daily Jensen-Shannon divergence threshold of 0.1. Feature-view-level monitoring evaluates the numerical feature against its baseline distribution, so alerts trigger precisely when divergence exceeds the specified limit.

Why this answer

Vertex AI Feature Store supports native feature monitoring configured at the feature view level, where you specify the drift detection method (including Jensen-Shannon divergence) and a threshold. Setting the threshold to 0.1 with a daily schedule directly satisfies the requirement without building custom infrastructure. This is the first-party, managed approach for detecting feature drift.

Exam trap

The trap is assuming model monitoring on the endpoint covers feature drift — PMLE candidates often pick endpoint monitoring when the question is specifically about Feature Store feature-level drift detection.

How to eliminate wrong answers

Option B is wrong because BigQuery scheduled queries require you to implement drift math manually and do not integrate with Feature Store's native monitoring or alerting. Option C is wrong because a Cloud Composer DAG is a custom orchestration workaround that duplicates functionality already provided by Feature Store monitoring. Option D is wrong because model monitoring on a Vertex AI endpoint detects prediction drift and skew at serving time, not feature-level drift in the Feature Store itself.

45
MCQmedium

A team uses Vertex AI Workbench managed notebooks. They want to version control their notebook files and collaborate using Git. What is the best way to integrate Git?

A.Use Cloud Source Repositories only
B.Use the built-in Git integration in Vertex AI Workbench managed notebooks
C.Use gcloud source repos clone inside the terminal
D.Manually download notebooks and upload to GitHub via browser
AnswerB

Managed notebooks include native Git integration, exposing repository cloning, credential handling and commit or push controls through the interface. This satisfies version control and collaboration requirements without manual command-line setup, unlike unmanaged instances where git must be configured by hand.

Why this answer

Vertex AI Workbench managed notebooks include native Git integration through the JupyterLab Git extension, allowing users to clone repositories, commit, push, and pull directly from the notebook UI without leaving the environment. This is the officially supported and most seamless method for version-controlling notebooks in managed notebooks, supporting GitHub, Cloud Source Repositories, and any Git-compatible remote.

Exam trap

The trap here is confusing Cloud Storage auto-persistence of the notebook VM with actual Git version control, leading candidates to pick Cloud Source Repositories or manual uploads instead of the built-in integration.

How to eliminate wrong answers

Option A is wrong because Cloud Source Repositories is only one possible Git remote and is not required — the built-in integration supports any Git provider, so restricting to CSR is unnecessarily narrow. Option C is wrong because running 'gcloud source repos clone' in the terminal only works with Cloud Source Repositories and bypasses the integrated UI workflow, making it less flexible than the built-in Git integration. Option D is wrong because manually downloading and re-uploading notebooks via a browser is error-prone, breaks commit history, and defeats the purpose of proper version control.

46
MCQhard

Your team uses a Vertex AI Pipeline that reads from a BigQuery table, trains a model, and registers it. A teammate wants to know which BigQuery table snapshot was used for a specific registered model version so they can reproduce the training data exactly. Which action should they take?

A.Query Vertex ML Metadata for the dataset artifact linked as an input to the execution that produced the model version.
B.List the model's versions in Vertex AI Model Registry and inspect the version's description field for the table name.
C.Check the Vertex AI Feature Store for the entity type that corresponds to the BigQuery table.
D.Open the pipeline run in the Vertex AI console and read the BigQuery table name from the run's parameters.
AnswerA

Vertex ML Metadata records the dataset artifact as an input to the training execution, and that artifact can carry a URI or metadata identifying the BigQuery snapshot. Traversing from the model artifact backward to this input artifact gives the exact data reference needed for reproduction.

Why this answer

Reproducing training data requires the exact dataset reference, which is captured as an input artifact in Vertex ML Metadata. Traversing the lineage from the model artifact to the training execution and its input dataset artifact yields that reference, whereas table names, descriptions, or feature stores do not guarantee snapshot-level accuracy.

Exam trap

The trap here is assuming that a table name in pipeline parameters or a description field is equivalent to an immutable dataset snapshot.

47
MCQmedium

A data scientist needs to retrieve training data from Vertex AI Feature Store that exactly matches the feature values as they were at a specific historical timestamp to avoid label leakage. Which feature view configuration should they use?

A.Enable point-in-time retrieval on the feature view.
B.Use the offline store without point-in-time and rely on data ordering.
C.Use the online store with a timestamp filter.
D.Create a new feature view with only historical data.
AnswerA

Point-in-time retrieval returns feature values as they existed at the supplied timestamp, joining each training row to its historical feature state. This directly prevents label leakage, satisfying the requirement that retrieved features exactly match values at a specific historical timestamp rather than current values.

Why this answer

Point-in-time retrieval is a feature of Vertex AI Feature Store that returns feature values as of a specified timestamp.

48
MCQmedium

A team is using Delta Lake on Dataproc for their data lake with ACID transactions. They want to version data for ML experiments and roll back to a previous version if needed. Which Delta Lake feature should they use?

A.Delta Lake streaming
B.Delta Lake schema enforcement
C.Delta Lake time travel
D.Delta Lake optimization (Z-order)
AnswerC

Time travel queries Delta table snapshots by version or timestamp, letting the team reproduce an exact training dataset and restore an earlier state via RESTORE. This directly satisfies the versioning and rollback requirement, unlike schema evolution or Z-ordering, which address structure and query performance.

Why this answer

Delta Lake time travel allows querying previous versions of a Delta table using either a version number or a timestamp, enabling reproducibility for ML experiments and rollback to a prior state. This is the native Delta Lake feature designed for versioning and auditing data changes. It directly satisfies the requirement to version data and roll back.

Exam trap

PMLE often tests confusion between Delta Lake features — candidates may pick schema enforcement or Z-order thinking they provide versioning, when only time travel enables historical queries and rollback.

How to eliminate wrong answers

Option A is wrong because Delta Lake streaming is about ingesting and processing streaming data, not versioning or rollback. Option B is wrong because schema enforcement prevents incompatible writes but does not provide historical version access. Option D is wrong because Z-order optimization improves query performance by co-locating related data, not versioning or rollback.

49
MCQeasy

What is the primary benefit of using a centralised model registry in MLOps?

A.Governance and version control of models
B.Better hyperparameter tuning
C.Faster model training
D.Automatic model deployment
AnswerA

A centralised model registry provides governance and version control, tracking each model's lineage, approvals, and deployment stage. This gives teams a single authoritative source for model artefacts, satisfying auditability and reproducibility requirements across the MLOps lifecycle.

Why this answer

A centralised model registry provides governance, versioning, and lineage tracking, enabling collaboration and auditability.

50
MCQmedium

A data science team wants to version control their datasets along with code using Git. They need a tool that integrates with Git and tracks changes to large data files. Which tool should they use?

A.BigQuery table snapshots
B.Delta Lake
C.Git LFS
D.DVC
AnswerD

DVC stores large dataset files outside Git while committing small metafiles that capture content hashes and versions, so Git tracks dataset changes without bloating the repository. This pointer-based mechanism satisfies the requirement to version data alongside code using Git.

Why this answer

DVC (Data Version Control) is purpose-built to integrate with Git and version large datasets and ML models by storing metadata and pointers in Git while keeping the actual data in remote storage. It provides commands like dvc add, dvc push, and dvc pull that mirror Git workflows, making it the right choice for versioning datasets alongside code.

Exam trap

The trap is confusing Git LFS with DVC — both handle large files, but only DVC provides dataset versioning, pipeline reproducibility, and remote storage abstraction for ML workflows.

How to eliminate wrong answers

Option A is wrong because BigQuery table snapshots version BigQuery tables for time travel and recovery, not Git-integrated dataset versioning for arbitrary files. Option B is wrong because Delta Lake provides ACID transactions and time travel on data lakes but is not a Git-integrated version control tool for datasets. Option C is wrong because Git LFS tracks large files via pointers but is designed for binary assets, not dataset pipelines with reproducibility, metrics, and experiment tracking like DVC.

51
MCQhard

An organization needs to implement MLOps with standardized pipeline templates across multiple teams. Which Vertex AI feature should they use to create reusable pipeline components?

A.Vertex AI Pipelines
B.Vertex AI Experiments
C.Vertex AI Metadata
D.Vertex AI Workbench
AnswerA

Vertex AI Pipelines lets teams author reusable components and pipeline templates with the Kubeflow Pipelines DSL, then share and parameterise them across projects. This directly satisfies the requirement for standardised, reusable pipeline templates across multiple teams, since components are versioned artefacts rather than one-off scripts.

Why this answer

Vertex AI Pipelines is the managed orchestration service that lets teams define, run, and monitor ML workflows as directed acyclic graphs (DAGs) using Kubeflow Pipelines or TFX. It supports reusable pipeline components and templates, enabling standardization across teams by packaging preprocessing, training, evaluation, and deployment steps into versioned, shareable artifacts.

Exam trap

The trap is confusing Vertex AI Experiments (run tracking) with Vertex AI Pipelines (workflow orchestration) — both are MLOps tools, but only Pipelines provides reusable component templates for standardized workflows.

How to eliminate wrong answers

Option B is wrong because Vertex AI Experiments is designed for tracking and comparing experiment runs (metrics, parameters, artifacts), not for building reusable pipeline components or orchestrating multi-step workflows. Option C is wrong because Vertex AI Metadata is the underlying artifact and lineage tracking service — it records what happened but does not define or execute pipelines. Option D is wrong because Vertex AI Workbench is an interactive notebook environment for development, not a pipeline orchestration or component-reuse system.

52
MCQmedium

A team uses Vertex AI Workbench notebooks for collaborative model development. They want to ensure that code changes are version-controlled, that multiple data scientists can work on the same notebook without conflicts, and that the environment is reproducible across team members. Which approach should they take?

A.Use a shared JupyterLab instance launched on a single VM; data scientists connect simultaneously.
B.Use Vertex AI Workbench managed notebooks with Git integration and a custom container image for environment reproducibility.
C.Store notebooks in Cloud Storage and share the bucket; each user edits their own copy.
D.Use Vertex AI Pipelines to run all code as pipelines; data scientists only view results in notebooks.
AnswerB

Git integration in managed notebooks provides branch-based version control and conflict resolution for concurrent editors, while a custom container image pins library versions so every data scientist gets an identical runtime. This satisfies the stem's three constraints: version-controlled changes, conflict-free collaboration, and reproducible environments across the team.

Why this answer

Vertex AI Workbench managed notebooks support Git integration for version control and custom container images for reproducible environments. This combination allows multiple data scientists to collaborate on notebooks with version history and ensures the runtime environment is consistent across team members.

Exam trap

The trap is thinking that sharing a VM or Cloud Storage bucket enables collaboration — the exam expects you to recognize that Git integration and custom containers are required for version control and reproducibility.

How to eliminate wrong answers

Option A is wrong because a shared JupyterLab instance on a single VM does not provide version control, conflict resolution, or environment reproducibility, and simultaneous editing can cause conflicts. Option C is wrong because storing notebooks in Cloud Storage and sharing copies lacks Git-based version control and does not prevent conflicts or ensure reproducibility. Option D is wrong because Vertex AI Pipelines is for orchestrating ML workflows, not for interactive collaborative notebook development with Git integration.

53
MCQmedium

An ML team wants to implement data versioning for large datasets stored in Google Cloud Storage. They need to track changes over time and reproduce previous data states. Which tool is most appropriate?

A.Cloud Storage Object Versioning
B.BigQuery table snapshots
C.Git LFS
D.DVC
AnswerD

DVC versions large datasets held in Google Cloud Storage by storing content hashes in small metafiles committed to Git, leaving the data in place. This delivers change tracking and reproducible checkouts of previous data states, exactly the capability the team requires.

Why this answer

DVC (Data Version Control) is purpose-built for versioning large datasets and ML models alongside Git. It stores small metadata files in Git while pushing the actual large data to remote storage like Google Cloud Storage, enabling teams to track dataset changes over time and reproduce exact previous data states via commits or tags.

Exam trap

The trap is assuming native GCS Object Versioning is sufficient for ML data versioning — candidates overlook that object-level versioning lacks dataset-level snapshots, lineage, and Git-integrated reproducibility that DVC provides.

How to eliminate wrong answers

Option A is wrong because Cloud Storage Object Versioning only retains prior versions of individual objects — it does not provide dataset-level versioning, lineage, or the ability to check out a coherent snapshot of an entire dataset at a point in time. Option B is wrong because BigQuery table snapshots version BigQuery tables, not datasets stored in GCS, and the question specifies data in Cloud Storage. Option C is wrong because Git LFS is designed for large binary files in Git repos but is not optimized for ML dataset workflows, lacks data-specific features like pipeline stage caching, and struggles with very large datasets.

54
MCQmedium

A team is building a fraud detection model that requires joining real-time transaction features with historical user features. They need to ensure that the training data does not use future information (data leakage). Which Vertex AI Feature Store capability should they use?

A.Online store serving with Bigtable
B.Feature store time travel
C.Point-in-time correct join
D.Feature monitoring for drift
AnswerC

Point-in-time correct join retrieves feature values as they existed at each training example's timestamp, preventing future information from leaking into training. This satisfies the stem's requirement that training data avoid data leakage when joining real-time and historical features.

Why this answer

Point-in-time correct joins in Vertex AI Feature Store ensure that when training examples are generated, each row uses only feature values that were valid as of the event timestamp of the label — preventing future data from leaking into training. This is the specific capability designed to avoid label leakage in time-series or event-driven ML. Time travel and online serving do not by themselves guarantee temporal correctness of the join.

Exam trap

PMLE often tests the confusion between time travel (retrieving historical values) and point-in-time correct joins (aligning features to label timestamps to prevent leakage) — candidates pick time travel thinking it solves leakage.

How to eliminate wrong answers

Option A is wrong because online store serving with Bigtable is about low-latency retrieval of the latest feature values for real-time inference — it does not enforce temporal alignment between features and labels during training. Option B is wrong because feature store time travel lets you retrieve historical feature values as of a past timestamp, but it does not automatically align each training row's features to the label's event time; you still need a point-in-time join to prevent leakage. Option D is wrong because feature monitoring for drift detects distribution changes in served features over time — it is an observability capability, not a training-data correctness mechanism.

55
Multi-Selecthard

A team uses Vertex AI Pipelines and wants to track lineage of artifacts and executions. Which three resources should they use? (Choose three.)

Select 3 answers
A.Artifacts
B.Vertex AI Experiments
C.Vertex AI Metadata
D.Model Registry
E.Executions
AnswersA, C, E

Artifacts are the versioned, typed outputs and inputs that Vertex AI ML Metadata records, such as datasets, models and metrics. Naming them satisfies the lineage requirement because each Artifact node links to the Executions and Contexts that produced or consumed it, forming the traceable graph.

Why this answer

Vertex AI Metadata is the core lineage service that stores and connects metadata about ML resources, so option C is correct because it provides the underlying metadata store for tracking lineage. Within that metadata store, Artifacts (option A) represent the inputs and outputs of pipeline steps—such as datasets, models, and metrics—and are the nodes whose lineage is tracked. Executions (option E) represent a single run of a pipeline step or component and record the events that consume and produce artifacts, which is exactly what links artifacts together into a lineage graph.

Vertex AI Experiments (option B) is for tracking and comparing experiment runs and metrics, not for artifact/execution lineage, and Model Registry (option D) is for managing model versions and deployment, not for general lineage tracking.

Exam trap

PMLE often tests the distinction between Metadata (lineage: artifacts/executions) and Experiments (runs/metrics) — candidates pick Experiments because it sounds like the tracking service.

56
MCQmedium

A fraud detection team trains a model on data stored in BigQuery. They want to ensure that the model can be reproduced exactly one year later, including the specific data version and training code. They use Vertex AI Pipelines for orchestration. Which practice should they implement?

A.Export the training data to a Cloud Storage bucket and enable Object Versioning on the bucket.
B.Use BigQuery table snapshots and commit the pipeline definition to a Git repository, then log the snapshot ID as a pipeline parameter.
C.Enable Vertex AI Model Registry's automatic versioning and rely on the model's artifact URI to retrieve the training data.
D.Schedule a daily BigQuery export to Cloud Storage and use the latest export for training.
AnswerB

BigQuery table snapshots provide a point-in-time immutable copy of the training data. Storing the pipeline code in Git and passing the snapshot ID as a parameter ensures both data and code are versioned. This combination enables exact reproduction of the training run.

Why this answer

Reproducibility requires capturing both the code and the exact data state. BigQuery table snapshots provide an immutable copy of the training data at a point in time. By committing the pipeline definition to Git and logging the snapshot ID as a parameter, the team can rerun the pipeline with the same code and data, ensuring identical results.

Exam trap

The trap here is assuming that model versioning alone ensures reproducibility, when in fact the training data and code must also be versioned.

57
MCQhard

An ML engineer trained a model and registered it in Vertex AI Model Registry. They want to assign the alias 'champion' to the best-performing version for production deployment. Which gcloud command should they use?

A.gcloud ai models versions describe --model=MODEL_ID --version=VERSION_ID
B.gcloud ai models upload --model-id=MODEL_ID --display-name=champion
C.gcloud ai endpoints deploy-model --model=MODEL_ID --alias=champion
D.gcloud ai models versions update --model=MODEL_ID --version=VERSION_ID --update-aliases=champion
AnswerD

The versions update subcommand with --update-aliases assigns the 'champion' alias to a specific registered model version, enabling production deployment by alias rather than version number. Other commands set IAM policy or labels, which do not create the deployment alias the scenario requires.

Why this answer

The correct command is 'gcloud ai models versions update' with the --update-aliases flag. This command updates a specific model version's aliases, allowing you to assign or remove aliases like 'champion' to denote the best-performing version for production. Option A only describes a version, option B uploads a new model, and option C deploys a model to an endpoint, none of which assign aliases to a version.

58
MCQhard

A company uses Vertex AI Pipelines to orchestrate ML workflows. After a pipeline run, they want to query the lineage of a particular model artifact to find out which dataset and hyperparameters were used to produce it. Which API method should they use?

A.projects.locations.metadataStores.artifacts.queryArtifactLineageSubgraph
B.projects.locations.metadataStores.artifacts.get
C.projects.locations.metadataStores.contexts.addContextArtifactsAndExecutions
D.projects.locations.metadataStores.executions.queryExecutionInputsAndOutputs
AnswerA

queryArtifactLineageSubgraph traverses the metadata store's lineage graph in both directions from a given artifact, returning the executions and artifacts that produced or consumed it. This directly satisfies the requirement to trace a model artifact back to its source dataset and hyperparameters.

Why this answer

The queryArtifactLineageSubgraph method returns the lineage subgraph for a given artifact, showing the executions, contexts, and other artifacts connected to it — exactly what is needed to trace which dataset and hyperparameters produced a model. It traverses both upstream (inputs) and downstream (outputs) relationships in the metadata store.

Exam trap

PMLE often tests the difference between a single-node lookup (artifacts.get) and a graph traversal (queryArtifactLineageSubgraph) — candidates pick the simpler get method and miss that lineage requires traversing relationships.

How to eliminate wrong answers

Option B is wrong because artifacts.get only retrieves the artifact's own metadata (name, URI, properties) and does not traverse lineage relationships. Option C is wrong because contexts.addContextArtifactsAndExecutions is a write operation that associates artifacts and executions with a context — it does not query lineage. Option D is wrong because executions.queryExecutionInputsAndOutputs returns the inputs and outputs of a single execution, which is only one hop and does not give the full lineage subgraph across the pipeline.

59
Multi-Selecthard

A team is operationalizing a machine learning pipeline using Vertex AI. They want to automatically track experiment runs, log model parameters and metrics, and store model artifacts for reproducibility. They also need to capture lineage between pipeline components (e.g., which dataset and hyperparameter tuning job produced a model). Which TWO services should they use together to achieve this? (Choose two.)

Select 2 answers
A.Vertex AI Model Registry
B.Vertex AI Feature Store
C.Vertex AI Metadata
D.Vertex AI Experiments
E.Vertex AI Workbench
AnswersC, D

Vertex AI Metadata records and stores lineage artefacts, capturing which dataset, pipeline component and hyperparameter tuning job produced each model. This satisfies the lineage requirement, letting teams trace model provenance across pipeline components for reproducibility and audit.

Why this answer

Vertex AI Experiments (D) is correct because it automatically tracks and compares experiment runs, logging parameters, metrics, and artifacts so results are reproducible and comparable across training runs. Vertex AI Metadata (C) is correct because it records lineage and context for ML artifacts, capturing relationships such as which dataset and hyperparameter tuning job produced a given model, which is exactly the lineage requirement. Together they cover both experiment tracking and artifact lineage.

Vertex AI Model Registry (A) manages model versions and deployment lifecycle but does not itself track experiment runs or component lineage. Vertex AI Feature Store (B) serves and manages feature values for training/serving, not experiment tracking or lineage. Vertex AI Workbench (E) is a notebook development environment and does not provide the required automatic tracking or lineage services.

Exam trap

PMLE often tests the overlap between Experiments and Metadata — candidates pick one service thinking it covers both experiment tracking and lineage, when the scenario explicitly requires both capabilities together.

60
MCQmedium

An ML engineer has a model trained in Vertex AI and wants to deploy it to an endpoint with autoscaling and traffic splitting for canary testing. They have the model artifact stored in Vertex AI Model Registry with alias 'champion'. What is the correct sequence of steps?

A.Upload model to registry, create endpoint, then deploy model to endpoint with traffic split.
B.Create endpoint, upload model to registry, then deploy model to endpoint with traffic split.
C.Create endpoint, deploy model directly from Cloud Storage, then add traffic split.
D.Upload model to registry, then create endpoint and deploy in one command using gcloud ai endpoints deploy-model.
AnswerA

Deployment requires the model in the registry first, then an endpoint, then deploying that model to the endpoint where traffic splits are configured. This sequence satisfies the canary traffic-splitting constraint, since traffic allocation is set at deploy time on the endpoint, not during upload.

Why this answer

The correct sequence is: upload the model to Vertex AI Model Registry (creating a model resource with the 'champion' alias), create an endpoint, then deploy the model to the endpoint with traffic split for canary testing. The model must exist in the registry before it can be deployed, and the endpoint must exist before deployment.

Exam trap

PMLE often tests the deployment sequence — candidates reverse the order (endpoint before model) or assume models can be deployed directly from GCS, missing the mandatory Model Registry upload step.

How to eliminate wrong answers

Option B is wrong because the model must be uploaded to the registry before it can be deployed — creating the endpoint first does not change the dependency that the model resource must exist. Option C is wrong because Vertex AI does not deploy models directly from Cloud Storage to an endpoint; the model must first be imported into Model Registry as a model resource. Option D is wrong because although gcloud can create an endpoint and deploy in sequence, the model must still be uploaded to the registry first, and the 'one command' framing omits the required upload step and the explicit traffic split configuration.

61
MCQmedium

A team is building ML pipelines with Vertex AI. They want to reuse standard pipeline components across teams and enforce governance. What approach should they take?

A.Use Vertex AI Pipelines with pre-built and custom components organized in a component registry.
B.Store pipeline definitions in a shared Cloud Storage bucket and copy them manually.
C.Use Cloud Composer to orchestrate ad-hoc scripts.
D.Have each team build their own pipelines independently.
AnswerA

Vertex AI Pipelines executes containerised components, and storing pre-built and custom components in a component registry lets teams share and version them centrally, enforcing governance and reuse across teams. Components are defined by YAML specs, so standard interfaces are preserved.

Why this answer

Vertex AI Pipelines lets teams define ML workflows as DAGs of containerized components, and organizing those components in a component registry (e.g., Vertex AI's component registry or Artifact Registry) enables reuse across teams. This approach enforces governance through versioning, access control, and standardized interfaces, so teams share vetted components rather than duplicating pipeline logic.

Exam trap

PMLE often tests whether candidates choose ad-hoc or manual approaches (shared buckets, independent pipelines) over the standardized, governed Vertex AI Pipelines + component registry pattern, so picking 'shared Cloud Storage bucket' is the common trap.

How to eliminate wrong answers

Option B is wrong because storing pipeline definitions in a Cloud Storage bucket and copying them manually provides no versioning, access control, or discoverability — it is an anti-pattern that leads to drift and governance gaps. Option C is wrong because Cloud Composer orchestrates general-purpose workflows (often data/ETL) and is not the standard for reusable ML pipeline components with governance; ad-hoc scripts lack the structure and lineage of Vertex AI Pipelines. Option D is wrong because each team building pipelines independently defeats reuse and governance, causing duplicated effort and inconsistent standards.

62
MCQhard

Your organization uses Vertex AI Pipelines for training. A compliance auditor asks you to prove which dataset version and which preprocessing code commit produced a model that is currently deployed. You need to retrieve this information programmatically for a specific model version. Which approach should you use?

A.List the endpoint's deployed models and read the model description field, which automatically records the dataset version and code commit.
B.Use Vertex ML Metadata to traverse the lineage from the model artifact to its parent execution and input artifacts.
C.Inspect the model's Cloud Storage directory for a metadata.json file that lists the dataset and code commit.
D.Query Vertex AI Experiments for the run that has the same display name as the model version.
AnswerB

Vertex ML Metadata stores artifacts, executions, and events, so you can start from the model artifact and walk backward to the training execution, then to the dataset artifact and code artifact that were inputs. This provides an auditable, programmatic chain of provenance for the deployed model version.

Why this answer

Vertex ML Metadata is the system of record for lineage, linking artifacts such as models and datasets through executions. Traversing lineages from the model artifact to its parent execution and input artifacts yields the dataset version and code commit programmatically, which is exactly what an auditor needs.

Exam trap

The trap here is trusting a human-readable description or display name instead of the machine-recorded lineage graph.

63
MCQhard

A company trains a model using features from Vertex AI Feature Store. They notice training-serving skew because the feature values used at training time differ from those served online. How should they address this?

A.Use the same online store for both training and serving
B.Disable caching in the online store
C.Enable feature monitoring to detect drift
D.Use point-in-time correct retrieval from the offline store for training data
AnswerD

Point-in-time correct retrieval joins each training label to feature values as they existed at that timestamp, preventing future data leaking into training. This aligns offline training vectors with the online serving values, directly eliminating the training-serving skew caused by mismatched feature snapshots.

Why this answer

Training-serving skew occurs when the feature values used during training differ from those served at inference time. Point-in-time correct retrieval from the offline store ensures that, for each training example, the feature values used are exactly those that would have been available at that timestamp — matching what the online store would have served. This eliminates the temporal mismatch that causes skew.

Exam trap

PMLE often tests the misconception that training and serving can share the same online store for consistency — candidates pick 'use the same online store' thinking it guarantees identical values, when in fact the online store only holds the latest value and cannot provide historical point-in-time correctness.

How to eliminate wrong answers

Option A is wrong because using the online store for training is not how Vertex AI Feature Store is designed — the online store serves low-latency latest values, not historical point-in-time values, so it cannot reproduce training-time conditions. Option B is wrong because disabling caching only affects latency and freshness of online reads; it does nothing to reconcile the historical training data with serving-time values. Option C is wrong because feature monitoring detects drift after the fact but does not fix the root cause of skew — it is a detection tool, not a correction mechanism.

64
MCQmedium

Your team trains a model on a Vertex AI Workbench notebook and logs hyperparameters, metrics, and a confusion matrix. Your manager asks you to ensure that anyone in the organization can reproduce the exact training run and compare it with other runs without manually digging through notebook cells. Which Vertex AI component should you use to record this information?

A.Vertex AI Experiments
B.Vertex AI Pipelines
C.Vertex AI Model Registry
D.Vertex AI TensorBoard
AnswerA

Vertex AI Experiments records parameters, metrics, and artifacts from training runs, and each run is associated with a lineage context. It provides a searchable UI and API so teammates can compare runs and reproduce them without inspecting notebook code, which satisfies the requirement for organization-wide tracking.

Why this answer

A system of record for training runs must capture parameters, metrics, and artifacts in a searchable way so teammates can reproduce and compare results. Vertex AI Experiments provides exactly this by linking each run to metadata and lineage, while visualization, deployment, and orchestration tools address different stages of the lifecycle.

Exam trap

The trap here is confusing visualization with tracking, assuming that a dashboard of metric curves is sufficient to reproduce a run.

65
MCQhard

You are collaborating on a Vertex AI Feature Store implementation. A data engineer updates a feature's values in the offline store, but the online store still serves the old values for several hours. The online store is configured with a feature value TTL of 24 hours and uses batch ingestion. What is the most likely cause of the stale online values?

A.The online store is configured to read from the offline store at query time, and the delay is due to eventual consistency between the two stores.
B.The feature values are ingested into the offline store only, and the online store is not being updated because batch ingestion does not automatically sync to the online store unless a separate ingestion job is run.
C.The online store's feature value TTL is set too high, causing it to serve cached values until the TTL expires.
D.The feature's online store TTL has expired, causing the online store to fall back to the offline store for the latest values.
AnswerB

In Vertex AI Feature Store, offline and online stores are separate. Batch ingestion writes to the offline store, and you must explicitly run an ingestion job to update the online store, or use streaming ingestion for real-time updates. If the data engineer only updated the offline store, the online store will not reflect the changes until a sync job is executed.

Why this answer

Vertex AI Feature Store maintains separate offline and online stores. Batch ingestion updates the offline store, but the online store requires its own ingestion job to sync values. Without running that job, the online store continues to serve the previous values.

The TTL affects validity, not propagation, so the missing sync job is the most likely cause.

Exam trap

The trap here is assuming that updating the offline store automatically propagates to the online store, or that TTL controls synchronization rather than validity.

66
Multi-Selectmedium

A company wants to implement a central model governance strategy using Vertex AI. They need to track model lineage, store evaluation metrics, and manage model versions across teams. Which THREE Vertex AI services should they use? (Choose 3)

Select 3 answers
A.Vertex AI Metadata
B.Vertex AI Model Registry
C.Vertex AI Workbench
D.Vertex AI Experiments
E.Vertex AI Feature Store
AnswersA, B, D

Vertex AI Metadata stores artefacts, executions and contexts in a managed ML metadata store, recording lineage as models train and deploy. It satisfies the lineage-tracking constraint directly, letting teams trace which datasets and training runs produced each model version. Evaluation metrics attach as metadata to those artefacts, supporting governance across teams.

Why this answer

Vertex AI Metadata (A) is correct because it provides the managed ML metadata store that automatically tracks artifacts, executions, and contexts, enabling model lineage tracking across teams. Vertex AI Model Registry (B) is correct because it is the central repository for managing model versions, including versioning, aliasing, and deployment tracking across teams. Vertex AI Experiments (D) is correct because it records and compares experiment runs, storing evaluation metrics and parameters so teams can track model performance over time.

Vertex AI Workbench (C) is not correct because it is an interactive notebook development environment, not a governance or lineage tracking service. Vertex AI Feature Store (E) is not correct because it manages feature storage and serving for training and prediction, not model lineage, metrics, or version management.

Exam trap

PMLE often tests which Vertex AI service does what; the trap is confusing Feature Store or Workbench with governance services, when the question specifically asks for lineage, versioning, and metrics.

67
MCQhard

A team uses Vertex AI Metadata to track pipeline runs. They need to identify all artifacts that were generated by a particular pipeline execution. Which API method should they use?

A.List executions and then list artifacts separately
B.Use the lineage query API with the execution ID
C.Create a context and query executions
D.Query artifacts by filter on execution ID
AnswerB

The lineage query API accepts an execution ID and returns the subgraph of artifacts produced by that execution, directly answering which artifacts a given pipeline run generated. Listing executions alone returns run metadata without the produced artifact relationships.

Why this answer

The Vertex AI Metadata lineage query API accepts an execution ID and returns all artifacts, contexts, and events connected to that execution, giving the full provenance graph in one call. This is the purpose-built method for tracing which artifacts a pipeline run produced.

Exam trap

PMLE often tests the assumption that artifacts can be filtered directly by execution ID — candidates pick the 'filter artifacts' option because it sounds precise, missing that Vertex AI Metadata expresses execution-artifact relationships through lineage events, not a direct foreign key.

How to eliminate wrong answers

Option A is wrong because listing executions and artifacts separately returns flat, unlinked lists — you would have to manually correlate them, and there is no guarantee of a clean join without lineage data. Option C is wrong because creating a context and querying executions inverts the relationship — contexts group related executions, but the question asks for artifacts generated by a specific execution, not executions within a context. Option D is wrong because artifacts do not carry a direct 'execution ID' filter field in the Metadata API; the linkage is expressed through lineage events, so filtering artifacts by execution ID is not a supported query pattern.

68
MCQmedium

Your team trains models in a shared Vertex AI project, and multiple engineers run pipelines against the same BigQuery training tables. A reviewer needs to reproduce the exact dataset used to train a model six weeks ago, but the source tables have been overwritten many times since. Which BigQuery capability should you have used to make each training snapshot reproducible?

A.BigQuery streaming inserts into a dedicated audit table
B.BigQuery time travel queries against the live table
C.BigQuery materialized views over the training tables
D.BigQuery table snapshots created before each training run
AnswerD

Table snapshots capture the table state at a point in time and are retained as independent, read-only copies that keep working even after the source table is overwritten. Taking a snapshot before each training run gives the reviewer a stable, addressable dataset to reproduce the exact training input without duplicating full storage cost.

Why this answer

Reproducing a past training run requires an immutable copy of the data as it existed then. Table snapshots provide exactly that: a point-in-time, read-only reference retained independently of later writes, so the reviewer can rebuild the training dataset even after many overwrites. Time travel, materialized views, and streaming audit tables all reflect or lose the historical state.

Exam trap

The trap here is assuming BigQuery time travel retains history indefinitely, when its window is limited and expires long before many reproducibility requirements.

69
MCQmedium

An organisation uses Delta Lake on Dataproc to manage a data lake for ML training. They need ACID transactions for concurrent reads and writes. Which file format does Delta Lake use as the underlying storage?

A.Apache Parquet
B.Apache ORC
C.CSV
D.Apache Avro
AnswerA

Delta Lake stores its data files as Apache Parquet, layering a transaction log over them to provide ACID guarantees. Parquet supplies the columnar, compressed storage; the Delta log records commits, enabling concurrent reads and writes without corrupting the underlying Parquet files.

Why this answer

Delta Lake stores its data files in Apache Parquet format and layers a transaction log (the _delta_log directory) on top to provide ACID guarantees, schema enforcement, and time travel. Parquet's columnar layout is ideal for ML workloads because it enables efficient column pruning and predicate pushdown during feature engineering. The Delta transaction log, not the file format itself, is what delivers the ACID properties.

Exam trap

PMLE often tests whether candidates know that Delta Lake's ACID properties come from the transaction log layer, not from Parquet itself — candidates may pick ORC or Avro thinking the format provides ACID, when in fact Parquet is the storage and the log provides the guarantees.

How to eliminate wrong answers

Option B is wrong because Apache ORC is a columnar format used primarily in the Hive/Spark ecosystem but is not the underlying storage format for Delta Lake — Delta is built on Parquet. Option C is wrong because CSV is a row-based, schema-less text format with no columnar compression or predicate pushdown, and it cannot support ACID transactions. Option D is wrong because Apache Avro is a row-based serialization format used for streaming and schema evolution in Kafka/Hadoop pipelines, not the columnar storage layer Delta Lake uses.

70
MCQeasy

A team wants to enforce governance and compliance for all ML models across the organisation. They need a centralised repository that tracks model versions, deployment history, and evaluation metrics. Which service should they use?

A.Cloud Storage
B.Vertex AI Feature Store
C.Vertex AI Experiments
D.Vertex AI Model Registry
AnswerD

Vertex AI Model Registry provides a centralised, versioned catalogue tracking each model's versions, deployment history and evaluation metrics, giving the organisation-wide governance and compliance visibility the stem requires. It integrates with Vertex AI Pipelines and endpoints, so lineage is captured automatically.

Why this answer

Vertex AI Model Registry is a centralized repository that tracks model versions, lineage, deployment history, and evaluation metrics across the organization, making it the correct choice for governance and compliance. It integrates with Vertex AI Pipelines and Model Monitoring so that every model artifact has an auditable lifecycle. This directly addresses the requirement for a single source of truth for ML models.

Exam trap

PMLE often tests the distinction between experimentation tracking (Vertex AI Experiments) and production model governance (Model Registry) — candidates pick Experiments because it also stores metrics, but it lacks deployment history and versioning for production models.

How to eliminate wrong answers

Option A is wrong because Cloud Storage is a generic object store — it can hold model artifacts but provides no versioning semantics, deployment tracking, or evaluation metric storage specific to ML governance. Option B is wrong because Vertex AI Feature Store manages feature values and serving for training/inference, not model versions or deployment history. Option C is wrong because Vertex AI Experiments tracks training runs, parameters, and metrics for experimentation, but it does not manage deployed model versions or serve as a governance repository for production models.

71
MCQmedium

A company uses Vertex AI Pipelines to train and deploy models. They want to automatically generate model documentation that includes model details, intended use, and evaluation results. What should they use?

A.Vertex AI Explanations
B.Vertex AI Metadata
C.Model Cards
D.Vertex AI Model Registry with custom metadata
AnswerC

Model Cards generate structured documentation covering model details, intended use and evaluation results, directly satisfying the documentation requirement. They attach to registered models in Vertex AI Model Registry, pulling evaluation metrics automatically rather than requiring manual authoring.

Why this answer

Model Cards are a standardized format for model documentation, and Vertex AI supports automated generation of model cards.

72
MCQeasy

A data science team wants to share engineered features across multiple projects while ensuring low-latency serving for online predictions. Which Google Cloud service should they use to store and serve these features?

A.Vertex AI Model Registry
B.Cloud Storage
C.BigQuery
D.Vertex AI Feature Store
AnswerD

Vertex AI Feature Store provides a centralised repository for engineered features with low-latency online serving, letting multiple projects reuse the same feature definitions. This satisfies both the sharing requirement and the online prediction latency constraint.

Why this answer

Vertex AI Feature Store is purpose-built to store, share, and serve machine learning features with low-latency online serving and consistent offline serving for training. It lets a data science team centralize engineered features so multiple projects reuse them, while providing an online serving endpoint that returns feature values in milliseconds for real-time predictions.

Exam trap

PMLE often tests the distinction between storing models (Model Registry) and storing features (Feature Store) — candidates conflate the two because both are 'Vertex AI' artifacts.

How to eliminate wrong answers

Option A is wrong because Vertex AI Model Registry stores and versions trained models, not feature data, and provides no low-latency feature-serving endpoint. Option B is wrong because Cloud Storage is object storage with high-latency access patterns, unsuitable for millisecond online feature retrieval and lacking feature-versioning or point-in-time correctness. Option C is wrong because BigQuery is an analytical warehouse optimized for large-scale SQL scans, not sub-millisecond per-entity online lookups, and it lacks native feature-serving semantics.

73
Multi-Selectmedium

An organization wants to implement central governance for ML models across teams. Which TWO services should they use together to achieve model versioning, lineage, and deployment management? (Select 2)

Select 2 answers
A.Vertex AI Feature Store
B.Vertex AI Model Registry
C.Vertex AI Metadata
D.Vertex AI Experiments
E.Cloud Data Catalog
AnswersB, C

Vertex AI Model Registry provides central model versioning, lineage tracking and deployment management across teams, satisfying the governance requirement. It organises models into versioned lineages and integrates with endpoints, giving one governed catalogue for all teams' models.

Why this answer

Vertex AI Model Registry (B) is correct because it provides a central repository for managing the lifecycle of ML models, including versioning, tracking model artifacts, and managing deployment to endpoints, which directly satisfies the model versioning and deployment management requirements. Vertex AI Metadata (C) is correct because it captures and stores metadata about artifacts, executions, and contexts, enabling lineage tracking across the ML workflow so teams can trace how models were produced and what data or training runs contributed to them. Together, Model Registry handles versioning and deployment while Metadata provides the lineage and governance context needed for central oversight.

Vertex AI Feature Store (A) is for organizing, serving, and sharing feature data, not for model versioning or deployment management. Vertex AI Experiments (D) is for tracking and comparing training experiment runs and metrics, which supports experimentation but not centralized model deployment governance. Cloud Data Catalog (E) is a data discovery and metadata management service for data assets, not a model registry or ML lineage/deployment tool.

74
MCQmedium

An ML team trains a model using a dataset stored in a BigQuery table. They want to ensure that the exact data snapshot used for training is recorded and can be reproduced later for auditing. Which approach should they take?

A.Use BigQuery table snapshots and log the snapshot ID as a parameter in the Vertex AI Pipeline run.
B.Schedule a daily export of the BigQuery table to Cloud Storage and use the latest export for training.
C.Export the BigQuery table to a Cloud Storage bucket and record the URI in Vertex AI Experiments.
D.Enable BigQuery audit logging and rely on the logs to reconstruct the data state.
AnswerA

BigQuery table snapshots provide a point-in-time, immutable copy of the table. By capturing the snapshot ID as a pipeline parameter, the training run is linked to the exact data version. This ensures reproducibility and auditability, as the snapshot can be restored or queried later. It directly addresses the need to record the data snapshot.

Why this answer

Using BigQuery table snapshots captures an immutable, point-in-time copy of the training data. Recording the snapshot ID in the pipeline run creates a direct link between the model and the exact data version, enabling reproducibility and audit compliance. Other methods either do not capture the precise data state or lack automatic linkage to the training run.

Exam trap

The trap here is assuming that exporting data or enabling audit logs provides a reproducible snapshot, when only a table snapshot guarantees an immutable, point-in-time copy.

75
MCQhard

A team wants to implement automated model documentation that captures training data, feature importance, evaluation metrics, and intended use. Which Vertex AI feature supports this?

A.Vertex AI Model Registry with model cards
B.Vertex AI Metadata
C.Vertex AI Explainable AI
D.Vertex AI Pipelines
AnswerA

Vertex AI Model Registry stores model cards alongside each version, capturing training data, feature importance, evaluation metrics and intended use. This satisfies the automated documentation requirement by centralising governance artefacts with the model lineage rather than in separate manual documents.

Why this answer

Vertex AI Model Registry supports model cards, which are structured documents that capture training data, feature importance, evaluation metrics, intended use, and limitations. Model cards are designed specifically for automated, standardized model documentation and governance. This makes Model Registry with model cards the correct feature for the stated requirement.

Exam trap

The trap is confusing metadata/lineage tracking (Vertex AI Metadata) or explainability (Explainable AI) with formal model documentation, which is specifically the model card feature in Model Registry.

How to eliminate wrong answers

Option B is wrong because Vertex AI Metadata stores lineage and artifact metadata for ML workflows (experiments, runs, artifacts), not human-readable model documentation like intended use and feature importance. Option C is wrong because Explainable AI provides feature attributions for predictions, not documentation of training data, metrics, and intended use. Option D is wrong because Vertex AI Pipelines orchestrates ML workflows; it does not itself produce model documentation artifacts.

Ready to test yourself?

Try a timed practice session using only Collaborating Within and Across Teams to Manage Data and Models questions.