Courseiva

Databricks-DE-Pro Cost and Performance Optimization Practice Question

A data engineer maintains a Delta Lake table that stores 5 years of order data. Analysts frequently query the most recent 90 days, but compliance requires that older data remain queryable. The table is currently partitioned by order_date and has 200,000 small files because data arrives continuously via Structured Streaming. Queries on the last 90 days are slow and expensive. Which combination of actions will most effectively reduce query cost and improve performance for the recent-data queries?

⚠ Common exam trap

The trap here is reaching for Z-ORDER on the partition column, which is already used for partitioning and therefore provides minimal additional data-skipping benefit.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Run OPTIMIZE on the partitions covering the last 90 days to compact small files, and configure the streaming writer with a longer trigger interval and optimized writes to produce larger files.

The dominant problem is the 200,000 small files created by continuous streaming ingestion. Queries on the last 90 days must open many tiny files, which drives up overhead and cost. Compacting only the hot partitions with OPTIMIZE reduces file count where it matters most, and adjusting the streaming writer to produce larger files prevents the problem from recurring. Z-ORDER on a partition key that is already the partition column adds little value, so focusing on compaction and ingestion file sizing is the correct optimization.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Run OPTIMIZE with Z-ORDER BY order_date on the entire table and set delta.autoOptimize.optimizeWrite to true.

    Why it's wrong here

    Z-ORDER on order_date can improve data skipping within each partition, but applying it to the entire 5-year table rewrites all data, which is expensive and unnecessary when only recent data is queried frequently. It also does not reduce file count for the streaming-ingested partitions unless compaction is combined. This is a costly operation that does not focus on the hot recent-data region.

  • ✗

    Enable Delta Lake column mapping and rewrite the table with generated columns for order month.

    Why it's wrong here

    Column mapping allows schema evolution without rewriting files and is useful for renaming or dropping columns, but it does not reduce the number of files or improve data skipping for date-range predicates. Generated columns can help derive partition values, but the table is already partitioned by order_date, so this adds complexity without addressing the small-file problem. It does not target the recent-90-day query pattern.

  • ✓

    Run OPTIMIZE on the partitions covering the last 90 days to compact small files, and configure the streaming writer with a longer trigger interval and optimized writes to produce larger files.

    Why this is correct

    Compacting the hot partitions reduces the number of files that queries must open, which directly lowers scan cost and improves performance for recent-data queries. Increasing the streaming trigger interval and enabling optimized writes reduces the rate of small-file creation going forward. This targets both the existing small-file problem and its root cause while leaving cold historical data untouched, which is the most cost-effective strategy.

  • ✗

    Run OPTIMIZE on partitions covering the last 90 days, then use Z-ORDER BY order_date on those partitions, and enable optimized writes for the streaming ingestion.

    Why it's wrong here

    This is close to the right approach, but Z-ORDER BY order_date inside partitions that are already partitioned by order_date provides little additional benefit because each partition already contains a narrow date range. The main gain comes from compaction, not Z-ordering. This option overstates the value of Z-ORDER and does not mention the more impactful lever of reducing file count for the hot partitions.

About these practice questions

One of 267 original Databricks-DE-Pro practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-DE-Pro practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DE-Pro exam.