Courseiva

PMLE Practice Question: Collaborating Within and Across Teams to Manage Data and Models

You are using DVC for data versioning in an ML project on Google Cloud. Your training data is stored in Cloud Storage. You want to track a new version of the dataset after preprocessing. Which DVC command should you use to register the changes?

⚠ Common exam trap

PMLE often tests the confusion between `dvc add` (register a new dataset/file) and `dvc commit` (update cache for existing stage outputs) — candidates pick `dvc commit` when the question asks about registering a brand-new dataset.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

dvc add data/processed

The `dvc add` command registers a file or directory with DVC, creating a .dvc metafile that captures the hash and metadata of the data, and adds the actual data to the DVC cache. After preprocessing produces a new dataset in data/processed, running `dvc add data/processed` creates the versioned pointer that DVC tracks in Git, which is exactly what 'registering the changes' means.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    dvc add data/processed

    Why this is correct

    dvc add computes the hash of the processed directory, writes a .dvc file and updates .gitignore, registering the new dataset version for tracking. dvc push only uploads already-tracked data to remote storage; it does not register changes.

  • ✗

    dvc push

    Why it's wrong here

    dvc push uploads cached data to remote storage; it does not record the new dataset version in the DVC metafiles. It is tempting because push transfers data to Cloud Storage, which is needed eventually, but registering changes first requires updating the .dvc file.

  • ✗

    dvc run -n preprocess

    Why it's wrong here

    dvc run defines a pipeline stage and its dependencies, not a dataset version; it records outputs of a command rather than registering changed data. It is tempting because preprocessing is a pipeline step, so dvc run suits building reproducible stages. Registering the new dataset version requires dvc add, which updates the .dvc file and cache.

  • ✗

    dvc commit

    Why it's wrong here

    dvc commit records the current data state into existing DVC metafiles without rerunning the pipeline stage that produced it. It is tempting as a way to save changes, but it skips the stage execution and dependency tracking that dvc repro or dvc add would register.

About these practice questions

One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

Same concept, more angles

1 more way this is tested on PMLE

These questions test the same concept from different angles. Work through them to make sure you can recognise it however the exam phrases it.

Variation 1. A company wants to use DVC for data versioning alongside their ML code in Git. Which TWO statements about DVC are correct? (Select 2)

easy
  • ✓ A.DVC uses a separate .dvc file to track data versions.
  • ✓ B.DVC can push data to remote storage like Google Cloud Storage.
  • C.DVC stores the actual data files in Git.
  • D.DVC only works with AWS S3 as remote storage.
  • E.DVC replaces Git for code versioning.

Why A: Option A is correct because DVC creates small .dvc metafiles (e.g., data.dvc) that record the MD5 hash and path of the tracked data, allowing Git to version the pointer while the actual data lives in the DVC cache. Option B is correct because DVC supports many remote storage backends, including Google Cloud Storage, via `dvc remote add` and `dvc push`, so data can be stored outside the Git repository. Option C is wrong because DVC deliberately keeps large data files out of Git, storing them in its cache and remotes instead. Option D is wrong because DVC supports S3, GCS, Azure Blob Storage, SSH, HDFS, and local remotes, not only S3. Option E is wrong because DVC complements Git for data versioning; it does not replace Git for source code version control.

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.