Courseiva

Databricks-DE-Assoc Data Transformation and Modeling Practice Question

Exhibit

--- Config Snippet ---
spark.databricks.delta.retentionDurationCheck.enabled = true
spark.databricks.delta.logRetentionDuration = interval 30 days

Refer to the exhibit. A data engineer is attempting to run a VACUUM command on a table, but the command fails with an error indicating that the retention period is too short. Given the configuration, what is the most appropriate action the engineer should take to safely remove files older than 7 days?

⚠ Common exam trap

Candidates mistakenly try to bypass safety checks or use incorrect time units, forgetting that 7 days equals 168 hours when configuring retention parameters.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Execute VACUUM table_name RETAIN 168 HOURS.

The configuration restricts the vacuuming of files to ensure Time Travel remains available. Databricks prevents setting a retention period shorter than the log retention to avoid data loss. To safely vacuum, the engineer must temporarily adjust the configuration or ensure the retention period respects the safety threshold. Understanding these safety constraints is critical for data lifecycle management and cost control in cloud storage environments.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Set 'spark.databricks.delta.retentionDurationCheck.enabled' to false.

    Why it's wrong here

    Disabling the retention check is dangerous and is not the recommended path. It removes the safety guardrail that prevents users from accidentally deleting data needed for recent Time Travel, leading to potential data loss if the table needs to be rolled back.

  • ✓

    Execute VACUUM table_name RETAIN 168 HOURS.

    Why this is correct

    Setting the retention period to 168 hours (7 days) is valid as long as it exceeds the minimum safety threshold. This command will effectively remove files that are no longer referenced by the Delta log and are older than seven days.

  • ✗

    Delete the underlying parquet files directly from cloud storage.

    Why it's wrong here

    Directly deleting files from cloud storage corrupts the Delta Lake table metadata. Delta Lake expects the log to match the physical files. Manual deletion results in inconsistent state, 'file not found' errors, and the inability to recover data using time travel.

  • ✗

    Run the OPTIMIZE command followed by REFRESH TABLE.

    Why it's wrong here

    OPTIMIZE performs file compaction but does not remove old files from storage. REFRESH TABLE only invalidates cached metadata. Neither command performs the cleanup required to reclaim storage space by removing orphaned files that are no longer part of the current table state.

About these practice questions

This Databricks-DE-Assoc question is part of Courseiva's 276-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-DE-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DE-Assoc exam.