Courseiva

PMLE Practice Question: Collaborating Within and Across Teams to Manage Data and Models

A data engineer needs to version large datasets (multiple TB) in a Data Lake on Google Cloud. They require ACID transactions to ensure consistency when multiple jobs read/write concurrently. Which solution should they use?

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Delta Lake on Dataproc

Delta Lake on Dataproc provides ACID transactions on cloud storage data lakes, enabling concurrent reads/writes with consistency.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Delta Lake on Dataproc

    Why this is correct

    Delta Lake provides ACID transactions and scalable metadata handling over Parquet files in Cloud Storage, letting concurrent Dataproc jobs read and write multi-terabyte datasets consistently. Its transaction log delivers the snapshot isolation and versioning the stem demands.

  • ✗

    BigQuery table snapshots

    Why it's wrong here

    BigQuery table snapshots preserve a table's state at a point in time for querying or restore, but they do not provide ACID transactions across concurrent writers in a data lake. They suit auditing or rollback of BigQuery tables, not multi-job consistency on TB-scale lake data.

  • ✗

    DVC (Data Version Control)

    Why it's wrong here

    DVC versions datasets by tracking files and metadata in Git, typically against object storage; it offers no transactional locking or ACID guarantees for concurrent writers. It suits reproducible ML pipelines and experiment tracking, not the multi-job consistency the stem demands.

  • ✗

    Vertex AI Feature Store

    Why it's wrong here

    Vertex AI Feature Store serves and versions online feature values for model training and serving; it does not provide ACID transactions over multi-terabyte data lake files. It would be correct when managing reusable ML features, but here concurrent read/write consistency across large datasets is required.

About these practice questions

This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.