Databricks-DE-Assoc Databricks Intelligence Platform Practice Question
A data engineer needs to run a nightly transformation that reads a large Parquet dataset, writes a curated Delta table, and then immediately runs OPTIMIZE and VACUUM on that table. The engineer wants each step to be observable, retryable, and to avoid data loss if VACUUM fails. Which orchestration approach best meets these requirements?
⚠ Common exam trap
The trap here is treating VACUUM as just another statement in the same notebook, which hides the fact that it is a separate, retryable unit whose failure should not force a full recomputation.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
A Databricks Job with separate tasks for the write, OPTIMIZE, and VACUUM, chained with task dependencies and configured retries.
Modeling the pipeline as a Databricks Job with discrete tasks lets each stage be independently logged, retried, and ordered through dependencies. The write commits first, OPTIMIZE then compacts files, and VACUUM runs only after the others succeed, so a VACUUM failure can be retried without recomputing the curated table. Single-cell notebooks and externally triggered jobs sacrifice that granular control and coordinated retry behavior.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
A single Databricks notebook that performs the read, write, OPTIMIZE, and VACUUM in one cell so the steps always run together.
Why it's wrong here
Combining all steps in one cell makes the unit atomic only in appearance: if VACUUM fails or the cluster dies mid-cell, the entire notebook task is marked failed and there is no per-step retry or visibility. The write may have committed while OPTIMIZE did not, and rerunning the whole cell repeats expensive work. It also prevents independent scheduling and monitoring of each stage.
- ✓
A Databricks Job with separate tasks for the write, OPTIMIZE, and VACUUM, chained with task dependencies and configured retries.
Why this is correct
A multi-task Databricks Job gives each step its own task, so the write, OPTIMIZE, and VACUUM run as distinct units with individual logs, durations, and retry policies. Dependencies ensure ordering, and a failed VACUUM can be retried without rerunning the write, preserving the committed Delta table. This directly satisfies observability, retryability, and data-loss avoidance.
- ✗
An external cron job on a VM that calls the Databricks REST API to submit each step as a separate one-time run.
Why it's wrong here
Submitting separate one-time runs from an external cron loses Databricks-native dependency tracking, retry semantics, and unified run history. The engineer must implement polling, error handling, and ordering in the external script, and a VACUUM failure would not automatically block or retry correctly. It also introduces infrastructure outside Databricks that must be maintained and secured.
- ✗
A Databricks Job that runs one notebook task containing the write and OPTIMIZE, followed by a separate job that runs VACUUM on a trigger.
Why it's wrong here
Splitting VACUUM into a second job decouples it from the pipeline, so there is no guarantee it runs after the write completes or that it sees the same table version. Failure of the second job is invisible to the first, and retry logic cannot coordinate the two. This adds operational complexity without delivering the required per-step observability and dependency control.
About these practice questions
One of 276 original Databricks-DE-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-DE-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DE-Assoc exam.