Databricks-DE-Assoc Data Transformation and Modeling Practice Question
Exhibit
--- Config Snippet --- spark.databricks.delta.retentionDurationCheck.enabled = true spark.databricks.delta.logRetentionDuration = interval 30 days
Refer to the exhibit. A data engineer is attempting to run a VACUUM command on a table, but the command fails with an error indicating that the retention period is too short. Given the configuration, what is the most appropriate action the engineer should take to safely remove files older than 7 days?
⚠ Common exam trap
Candidates mistakenly try to bypass safety checks or use incorrect time units, forgetting that 7 days equals 168 hours when configuring retention parameters.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Execute VACUUM table_name RETAIN 168 HOURS.
The configuration restricts the vacuuming of files to ensure Time Travel remains available. Databricks prevents setting a retention period shorter than the log retention to avoid data loss. To safely vacuum, the engineer must temporarily adjust the configuration or ensure the retention period respects the safety threshold. Understanding these safety constraints is critical for data lifecycle management and cost control in cloud storage environments.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Set 'spark.databricks.delta.retentionDurationCheck.enabled' to false.
Why it's wrong here
Disabling the retention check is dangerous and is not the recommended path. It removes the safety guardrail that prevents users from accidentally deleting data needed for recent Time Travel, leading to potential data loss if the table needs to be rolled back.
- ✓
Execute VACUUM table_name RETAIN 168 HOURS.
Why this is correct
Setting the retention period to 168 hours (7 days) is valid as long as it exceeds the minimum safety threshold. This command will effectively remove files that are no longer referenced by the Delta log and are older than seven days.
- ✗
Delete the underlying parquet files directly from cloud storage.
Why it's wrong here
Directly deleting files from cloud storage corrupts the Delta Lake table metadata. Delta Lake expects the log to match the physical files. Manual deletion results in inconsistent state, 'file not found' errors, and the inability to recover data using time travel.
- ✗
Run the OPTIMIZE command followed by REFRESH TABLE.
Why it's wrong here
OPTIMIZE performs file compaction but does not remove old files from storage. REFRESH TABLE only invalidates cached metadata. Neither command performs the cleanup required to reclaim storage space by removing orphaned files that are no longer part of the current table state.
About these practice questions
This Databricks-DE-Assoc question is part of Courseiva's 276-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-DE-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DE-Assoc exam.