PMLE Practice Question: Collaborating Within and Across Teams to Manage Data and Models
A data engineer needs to version large datasets (multiple TB) in a Data Lake on Google Cloud. They require ACID transactions to ensure consistency when multiple jobs read/write concurrently. Which solution should they use?
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Delta Lake on Dataproc
Delta Lake on Dataproc provides ACID transactions on cloud storage data lakes, enabling concurrent reads/writes with consistency.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Delta Lake on Dataproc
Why this is correct
Delta Lake provides ACID transactions and scalable metadata handling over Parquet files in Cloud Storage, letting concurrent Dataproc jobs read and write multi-terabyte datasets consistently. Its transaction log delivers the snapshot isolation and versioning the stem demands.
- ✗
BigQuery table snapshots
Why it's wrong here
BigQuery table snapshots preserve a table's state at a point in time for querying or restore, but they do not provide ACID transactions across concurrent writers in a data lake. They suit auditing or rollback of BigQuery tables, not multi-job consistency on TB-scale lake data.
- ✗
DVC (Data Version Control)
Why it's wrong here
DVC versions datasets by tracking files and metadata in Git, typically against object storage; it offers no transactional locking or ACID guarantees for concurrent writers. It suits reproducible ML pipelines and experiment tracking, not the multi-job consistency the stem demands.
- ✗
Vertex AI Feature Store
Why it's wrong here
Vertex AI Feature Store serves and versions online feature values for model training and serving; it does not provide ACID transactions over multi-terabyte data lake files. It would be correct when managing reusable ML features, but here concurrent read/write consistency across large datasets is required.
Go deeper
Related to this question
About these practice questions
This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.