Courseiva

PMLE Practice Question: Collaborating Within and Across Teams to Manage Data and Models

A company wants to use DVC for data versioning alongside their ML code in Git. Which TWO statements about DVC are correct? (Select 2)

⚠ Common exam trap

The trap is assuming DVC stores data in Git or is AWS-only — candidates must remember DVC uses pointer files and supports multiple remotes, and that it augments rather than replaces Git.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

DVC uses a separate .dvc file to track data versions.

Option A is correct because DVC creates small .dvc metafiles (e.g., data.dvc) that record the MD5 hash and path of the tracked data, allowing Git to version the pointer while the actual data lives in the DVC cache. Option B is correct because DVC supports many remote storage backends, including Google Cloud Storage, via `dvc remote add` and `dvc push`, so data can be stored outside the Git repository. Option C is wrong because DVC deliberately keeps large data files out of Git, storing them in its cache and remotes instead. Option D is wrong because DVC supports S3, GCS, Azure Blob Storage, SSH, HDFS, and local remotes, not only S3. Option E is wrong because DVC complements Git for data versioning; it does not replace Git for source code version control.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    DVC uses a separate .dvc file to track data versions.

    Why this is correct

    DVC stores a small .dvc metafile in Git containing the data's hash and path, while the actual dataset lives outside the repository. This lightweight pointer lets Git track dataset versions without bloating the repo, satisfying the requirement to version data alongside ML code.

  • ✓

    DVC can push data to remote storage like Google Cloud Storage.

    Why this is correct

    DVC decouples storage from Git by pushing cached data to configurable remotes, including Google Cloud Storage, S3 and Azure Blob Storage. This satisfies the scenario's need to version large datasets without committing them to Git, keeping only lightweight metafiles in the repository.

  • ✗

    DVC stores the actual data files in Git.

    Why it's wrong here

    DVC deliberately keeps data out of Git, committing only small .dvc pointer files that reference content-addressed objects in remote storage; Git's blob store cannot handle large binaries. Storing data in Git is what Git LFS does, which suits modest binary assets versioned directly alongside code rather than multi-gigabyte datasets.

  • ✗

    DVC only works with AWS S3 as remote storage.

    Why it's wrong here

    DVC supports many remote backends, including Google Cloud Storage, Azure Blob Storage, SSH servers and local directories, so S3 is not the only option. The claim tempts because S3 is the most commonly documented remote, but DVC's remote configuration is storage-agnostic.

  • ✗

    DVC replaces Git for code versioning.

    Why it's wrong here

    DVC is designed to complement Git, storing data and model artefacts while Git tracks code and small .dvc pointer files; it does not replace Git's code versioning. The claim tempts because DVC adds versioning for large files, but Git remains the underlying version control system for source code.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

Courseiva writes every PMLE question from scratch — 775 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.