Courseiva

Databricks-DE-Pro · topic practice

Cost and Performance Optimization practice questions

This domain covers tuning Databricks workloads for speed and cost: Delta Lake file layout, partitioning and compaction, Photon and cluster sizing, Databricks SQL warehouse configuration, and cost visibility. Questions are scenario-based, asking you to pick the right command, feature, or configuration for a described latency, memory, or spend problem.

Courseiva uses original exam-style practice questions designed for learning and revision. The goal is to understand the concepts, recognise exam patterns, and improve through explanations — not memorise copied exam dumps.

Editorial oversight:Johnson Ajibi· MSc IT Security, IEEE Senior Member
20 questionsDomain: Cost and Performance Optimization

What the exam tests

What to know about Cost and Performance Optimization

Be able to select the correct Delta command, DLT setting, or SQL warehouse configuration for a stated performance or cost symptom, and know which Databricks feature reports or limits spend. The key is matching the symptom (small files, cold starts, OOM, overspend) to the precise tool rather than a generic tuning step.

Using OPTIMIZE with Z-ORDER and liquid clustering to compact small files and improve Delta read latency

Configuring Delta Live Tables pipeline modes, serverless compute, and enhanced autoscaling to control streaming cost

Choosing Databricks SQL warehouse sizing, auto-stop, and serverless options for sporadic unpredictable queries

Applying system tables, budgets, and SQL warehouse alerts to monitor and cap Databricks spend

Watch out for

Common Cost and Performance Optimization exam traps

  • ▸Confusing OPTIMIZE (compaction and layout) with VACUUM (removing old files) when the stated problem is read latency from small files
  • ▸Assuming partitioning always helps, when over-partitioning creates small files and hurts performance; liquid clustering is often the better answer
  • ▸Treating cluster auto-termination or auto-scaling as the fix for SQL warehouse cold starts, when serverless or a running warehouse addresses that

Practice set

Cost and Performance Optimization questions

20 questions · select your answer, then reveal the explanation

A data engineering team manages a Delta Lake table that is frequently updated and queried by multiple downstream jobs. Users report that queries are increasingly slow over time. Inspection shows thousands of small JSON files in the storage location due to streaming appends. Which optimization technique should the data engineer apply to restore query performance cost-effectively?

To optimize performance for a table that is frequently filtered by a date column, what is the best strategy?

A data engineer observes that a Delta table containing 50TB of data is experiencing slow scan performance. The table is partitioned by 'date', but many queries filter by 'region' and 'customer_id'. Which optimization strategy should the engineer implement to improve query performance with minimal overhead?

Refer to the exhibit. A data engineer is running a heavy join operation on a cluster with the provided configuration. Despite the auto-scaling being set to a maximum of 8 workers, the join operation is consistently spilling to disk. What is the most likely cause?

Exhibit

{
  "cluster_config": {
    "autoscale": {"min_workers": 2, "max_workers": 8},
    "spark_conf": {
      "spark.databricks.io.cache.enabled": "true",
      "spark.sql.shuffle.partitions": "auto"
    }
  }
}

Which THREE actions should be taken to minimize the cost of running a Databricks Job that processes large volumes of intermittent data?

A data engineering team runs a nightly batch job that performs a full outer join between a 5 TB fact table and a 200 GB dimension table. The job runs on a job cluster with Photon enabled. The team wants to reduce the amount of data shuffled across the network during the join. Which configuration change should they apply to the fact table DataFrame before the join?

A data engineer manages a Databricks job that runs daily on a job cluster and processes a 2 TB Delta table. The job performs a large shuffle and writes the result back to the same table using MERGE. Analysis shows the shuffle read is 1.5 TB and the shuffle write is 1.5 TB. The engineer wants to reduce shuffle overhead and improve performance. Which configuration change should the engineer make?

A data engineer is using Databricks SQL to query a large Delta table that is frequently updated with small appends. The engineer notices that queries are slow and the table has many small files. The engineer wants to improve query performance and reduce storage costs without changing the table schema. Which action should the engineer take?

A data engineer is optimizing a Databricks SQL warehouse that runs a mix of short ad-hoc queries and long-running analytical queries. The warehouse is currently configured with a fixed size and auto-stop after 10 minutes. Users report that short queries sometimes wait for cluster startup, and long queries compete for resources, causing delays. The engineer wants to improve performance and reduce cost. Which two actions should the engineer take? (Choose two.)

A data engineer is analyzing a Databricks SQL warehouse that runs a dashboard with multiple queries. The warehouse is configured with auto-stop after 10 minutes of inactivity. The dashboard is accessed sporadically throughout the day, with periods of high activity and long idle times. The engineer wants to reduce cost without impacting user experience. Which action should the engineer take?

A data engineer is optimizing a Databricks job that reads from a large Parquet table stored in cloud object storage. The job performs a full scan of the table for each run, and the engineer wants to reduce the amount of data read and improve performance. The engineer decides to convert the table to Delta Lake. Which two actions should the engineer take to further reduce data read and improve performance? (Choose two.)

A data engineer is tuning a large Spark SQL join on Databricks that processes 20 TB of fact data joined to a 50 GB dimension table. The job uses a broadcast hash join because the dimension table is under the default auto-broadcast threshold, but the driver is running out of memory and the job is slow. The engineer wants to improve performance and stability without increasing driver size. Which two actions should the data engineer take? (Choose two.)

A data engineer is configuring a Databricks job cluster for a nightly ETL job that processes data in a single pass. The job reads from cloud storage, transforms the data, and writes to a Delta table. The engineer wants to minimize cost while ensuring the job completes within its SLA. Which cluster configuration is most appropriate?

A data engineer is tasked with reducing compute costs for an interactive SQL analytics workspace that runs sporadic, highly unpredictable queries. The jobs experience cold start delays and occasional out-of-memory errors due to sudden concurrency spikes. Which TWO strategies should the engineer implement to balance cost efficiency and performance?

A data engineer is designing an ETL pipeline processing high-frequency streaming data into Delta tables on Databricks. The pipeline experiences frequent small file creation and high metadata overhead, degrading query performance. Which optimization technique should the engineer implement to resolve this issue?

An enterprise data team runs a large nightly batch job using a standard all-purpose cluster. The job frequently fails due to cloud provider spot instance pre-emptions and takes over four hours to complete. How should the engineer refactor this architecture for maximum cost efficiency and reliability?

Refer to the exhibit. A data engineer creates an instance pool to reduce cluster startup times for development teams. However, finance reports indicate unexpected cloud infrastructure charges. Based on the configuration shown in the exhibit, what is the primary driver of these unexpected costs?

Exhibit

{
  "instance_pool_id": "pool-0412-182230-cried5",
  "min_idle_instances": 2,
  "max_capacity": 10,
  "node_type_id": "i3.xlarge"
}

A data engineer is optimizing a Delta Lake table that experiences high read latency due to many small files. Which command should be executed to physically reorganize the data layout to improve query performance?

A Databricks SQL warehouse is experiencing high costs due to idle resources. Which TWO configurations should be implemented to effectively manage and reduce warehouse costs?

Refer to the exhibit. An engineer has configured the cluster settings as shown. What is the expected impact on the Delta table's performance and write operations?

Exhibit

{"cluster_config": {"spark.databricks.delta.optimizeWrite.enabled": "true", "spark.databricks.delta.autoCompact.enabled": "true"}}

Free account

Track your progress over time

Create a free account to save your results and see which topics improve across sessions.

Focused Cost and Performance Optimization sessions

Start a Cost and Performance Optimization only practice session

Every question in these sessions is drawn from the Cost and Performance Optimization domain — nothing else.

Related practice questions

Related Databricks-DE-Pro topic practice pages

Move into related areas when this topic feels solid.

Frequently asked questions

What does the Databricks-DE-Pro exam test about Cost and Performance Optimization?
Be able to select the correct Delta command, DLT setting, or SQL warehouse configuration for a stated performance or cost symptom, and know which Databricks feature reports or limits spend. The key is matching the symptom (small files, cold starts, OOM, overspend) to the precise tool rather than a generic tuning step.
How should I use these practice questions?
Select your answer before revealing the explanation. Then read why each option is right or wrong — this active recall approach builds retention far faster than re-reading notes.
Can I practise just Cost and Performance Optimization questions in a focused session?
Yes — the session launcher on this page draws every question from the Cost and Performance Optimization domain. Use a 10-question session first to gauge your baseline, then move to 20 or 30 once the weak spots are clear.
Where can I practise other Databricks-DE-Pro topics?
Use the topic links above to move to related areas, or go back to the Databricks-DE-Pro question bank to see all topics.
Are these real exam questions or dumps?
These are original practice questions written to test the same concepts the Databricks-DE-Pro exam covers. They are not copied from any real exam or dump site.