PMLE Practice Question: Collaborating Within and Across Teams to Manage Data and Models
An organisation uses Delta Lake on Dataproc to manage a data lake for ML training. They need ACID transactions for concurrent reads and writes. Which file format does Delta Lake use as the underlying storage?
⚠ Common exam trap
PMLE often tests whether candidates know that Delta Lake's ACID properties come from the transaction log layer, not from Parquet itself — candidates may pick ORC or Avro thinking the format provides ACID, when in fact Parquet is the storage and the log provides the guarantees.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Apache Parquet
Delta Lake stores its data files in Apache Parquet format and layers a transaction log (the _delta_log directory) on top to provide ACID guarantees, schema enforcement, and time travel. Parquet's columnar layout is ideal for ML workloads because it enables efficient column pruning and predicate pushdown during feature engineering. The Delta transaction log, not the file format itself, is what delivers the ACID properties.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Apache Parquet
Why this is correct
Delta Lake stores its data files as Apache Parquet, layering a transaction log over them to provide ACID guarantees. Parquet supplies the columnar, compressed storage; the Delta log records commits, enabling concurrent reads and writes without corrupting the underlying Parquet files.
- ✗
Apache ORC
Why it's wrong here
Delta Lake does not use ORC.
- ✗
CSV
Why it's wrong here
CSV is a plain text row format with no schema, statistics, or transactional metadata, so it cannot support Delta's ACID guarantees. It tempts because CSV is simple to write and read, but Delta Lake's underlying storage is Parquet, which provides the columnar structure the transaction log relies on.
- ✗
Apache Avro
Why it's wrong here
Avro is a row-based serialisation format used for streaming and schema evolution, not Delta's columnar storage. It tempts because Avro supports schema evolution and is common in data pipelines, but Delta Lake's ACID transactions are implemented over Parquet files, not Avro.
Go deeper
Related to this question
About these practice questions
One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.