Databricks-DE-Pro · domain
Cost and Performance Optimization
This domain covers tuning Databricks workloads for speed and cost: Delta Lake file layout, partitioning and compaction, Photon and cluster sizing, Databricks SQL warehouse configuration, and cost visibility. Questions are scenario-based, asking you to pick the right command, feature, or configuration for a described latency, memory, or spend problem.
Focused practice
Practice Cost and Performance Optimization questions
Scored sessions drawing only from this domain — pick a length below.
Start 20-question practice test →What this domain covers
What to know about Cost and Performance Optimization
Be able to select the correct Delta command, DLT setting, or SQL warehouse configuration for a stated performance or cost symptom, and know which Databricks feature reports or limits spend. The key is matching the symptom (small files, cold starts, OOM, overspend) to the precise tool rather than a generic tuning step.
Using OPTIMIZE with Z-ORDER and liquid clustering to compact small files and improve Delta read latency
Configuring Delta Live Tables pipeline modes, serverless compute, and enhanced autoscaling to control streaming cost
Choosing Databricks SQL warehouse sizing, auto-stop, and serverless options for sporadic unpredictable queries
Applying system tables, budgets, and SQL warehouse alerts to monitor and cap Databricks spend
Watch out for
Common Cost and Performance Optimization exam traps
- ▸Confusing OPTIMIZE (compaction and layout) with VACUUM (removing old files) when the stated problem is read latency from small files
- ▸Assuming partitioning always helps, when over-partitioning creates small files and hurts performance; liquid clustering is often the better answer
- ▸Treating cluster auto-termination or auto-scaling as the fix for SQL warehouse cold starts, when serverless or a running warehouse addresses that
Question index
All Cost and Performance Optimization questions (33)
Click any question to see the full explanation, or start a practice session above.
A data engineer is tasked with reducing compute costs for an interactive SQL analytics workspace that runs sporadic, highly unpredictable queries. The jobs experience cold start delays and occasional out-of-memory errors due to sudden concurrency spikes. Which TWO strategies should the engineer implement to balance cost efficiency and performance?
Hard2A data engineering team runs a nightly batch job on a Databricks job cluster. The job reads a large Parquet dataset, performs transformations, and writes the result to a Delta table. The cluster is configured with autoscaling from 4 to 16 workers and uses the default Spark configuration. The team observes that the job runs for 2 hours, but the cluster's CPU utilization is only around 30% throughout the run. They want to reduce cost without increasing runtime. Which action is most likely to achieve this?
Medium3A data engineer manages a Delta table that is used for both batch analytics and frequent small updates from a streaming job. The table is not partitioned, and the engineer notices that queries are slowing down as the table grows. The engineer wants to improve query performance without changing the table schema or partitioning strategy. Which action should the engineer take?
Medium4Refer to the exhibit. A data engineer creates an instance pool to reduce cluster startup times for development teams. However, finance reports indicate unexpected cloud infrastructure charges. Based on the configuration shown in the exhibit, what is the primary driver of these unexpected costs?
Hard5A data team is using Liquid Clustering on a Delta table. How does this feature improve performance compared to traditional Z-Ordering or Partitioning?
Hard6A data engineer is designing an ETL pipeline processing high-frequency streaming data into Delta tables on Databricks. The pipeline experiences frequent small file creation and high metadata overhead, degrading query performance. Which optimization technique should the engineer implement to resolve this issue?
Medium7A data engineer is reviewing a Databricks job that runs on a job cluster and reads a large Delta table. The engineer notices that the job takes a long time to start because the cluster is provisioned from scratch each time. The engineer wants to reduce the startup time without increasing cost significantly. Which action should the engineer take?
Easy8Your organization runs numerous batch data engineering pipelines using standard Databricks jobs. Finance reports indicate that compute costs are inflated due to cluster startup times and rigid over-provisioning. Which optimization approach provides the best balance of cost savings and execution reliability for scheduled production batch jobs?
Medium9Which THREE actions can help reduce the 'shuffle' operations in a Spark job?
Hard10Refer to the exhibit. An administrator reviews the cluster configuration JSON for an all-purpose interactive development cluster used by data engineers. Based on Databricks cost and performance optimization best practices, which specific parameter in this configuration represents the highest risk for unnecessary financial expenditure?
Medium11A data engineer maintains a Delta Lake table that stores 5 years of order data. Analysts frequently query the most recent 90 days, but compliance requires that older data remain queryable. The table is currently partitioned by order_date and has 200,000 small files because data arrives continuously via Structured Streaming. Queries on the last 90 days are slow and expensive. Which combination of actions will most effectively reduce query cost and improve performance for the recent-data queries?
Hard12Which THREE techniques are recommended for improving the performance of Spark SQL joins on large Databricks tables?
Medium13A data scientist reports that their notebook takes 20 minutes to initialize, even when the cluster is running. What is the most likely reason for this high initialization time?
Medium14A data engineer notices that a Databricks SQL warehouse used for executive dashboards runs 24/7 but is only actively queried during business hours. The warehouse is a Pro-sized warehouse with auto-stop set to 10 minutes. The team wants to reduce cost without affecting dashboard availability during business hours. Which action should the data engineer take?
Easy15A team is building a streaming pipeline that processes millions of events per second. They are using Structured Streaming with a Delta Lake sink. What is the most effective way to optimize the performance and cost of this write-heavy workload?
Medium16A data engineer is tuning a Spark job that reads from a Delta table and writes to another Delta table. The job uses a groupByKey operation followed by an aggregation. The engineer notices that the job is spilling to disk during the shuffle and taking a long time. The engineer wants to reduce shuffle spill and improve performance. Which action is most likely to help?
Hard17A data engineer is designing a pipeline and notices that the cost of processing is unexpectedly high during development. Which action provides the most immediate cost reduction when using Databricks?
Easy18Which property should be configured to allow Databricks to automatically optimize the size of files during write operations in Delta Lake?
Hard19A Databricks SQL warehouse is experiencing high costs due to idle resources. Which TWO configurations should be implemented to effectively manage and reduce warehouse costs?
Hard20A data engineering team experiences massive compute waste because interactive development notebooks are frequently left running overnight by engineers. As a Databricks administrator, which configuration should you implement at the cluster policy level to automatically mitigate this financial exposure without disrupting ongoing development work?
Medium21A data engineering team is running a nightly batch job on a Databricks job cluster that processes a 10 TB Delta table. The job reads the entire table, performs transformations, and writes results to another Delta table. The team notices that the job takes 4 hours and consumes significant DBUs. They want to reduce runtime and cost without changing the business logic. The table is partitioned by ingestion date, but queries often filter on a high-cardinality column 'customer_id'. Which optimization technique is most appropriate to improve performance and reduce cost?
Medium22Which metric should a data engineer prioritize when investigating a slow-running query in the Databricks SQL query history?
Easy23A data engineer is optimizing a Delta Lake table that experiences high read latency due to many small files. Which command should be executed to physically reorganize the data layout to improve query performance?
Medium24A data engineer is configuring a Delta Live Tables (DLT) pipeline that processes streaming data from Apache Kafka. The pipeline performs a series of transformations and writes to a Delta table. The engineer notices that the pipeline is experiencing high latency and wants to optimize it for cost and performance. The pipeline is set to continuous mode. Which configuration change is most effective to reduce cost while maintaining acceptable latency?
Hard25A data engineer is reviewing a Databricks job that runs a notebook to process a large Delta table. The job takes 45 minutes, and the engineer notices that the cluster spends a significant amount of time in the 'Pending' state before execution begins. The cluster is a job cluster with autoscaling enabled and no cluster pool. The engineer wants to reduce the overall job duration and cost. Which action should the data engineer take?
Medium26A data engineer is configuring a Databricks SQL warehouse to handle a workload that consists of many concurrent short queries during business hours and almost no queries at night. The engineer wants to minimize cost while ensuring low latency during peak hours. Which configuration should the engineer use?
Medium27Refer to the exhibit. An engineer has configured the cluster settings as shown. What is the expected impact on the Delta table's performance and write operations?
Medium28A data engineer is reviewing the cost of a Databricks job that runs on a daily basis. The job uses an all-purpose cluster that is manually started and stopped by the engineer. The job typically runs for 30 minutes, but the engineer often forgets to stop the cluster, leading to hours of idle time. Which action should the engineer take to reduce cost?
Easy29An organization wants to monitor and limit the spend of their Databricks SQL warehouses. Which feature is most appropriate for setting alerts when costs exceed a certain threshold?
Medium30An enterprise data team runs a large nightly batch job using a standard all-purpose cluster. The job frequently fails due to cloud provider spot instance pre-emptions and takes over four hours to complete. How should the engineer refactor this architecture for maximum cost efficiency and reliability?
Medium31A team has a large Delta table that is rarely updated. What is the most cost-effective way to store this data while maintaining the ability to query it with Databricks SQL?
Easy32A data engineer is optimizing a Spark job that reads from a large Delta table and performs a join with a smaller dimension table. The job is running slowly, and the engineer suspects data skew and shuffle overhead are the main issues. Which two techniques should the engineer apply to improve performance and reduce cost? (Choose two.)
Hard33A Spark job is failing with an OutOfMemoryError (OOM) during a group-by operation on a skewed key. What is the most effective way to resolve this?
MediumOther domains
All Databricks-DE-Pro exam domains
Frequently asked questions
- What does the Cost and Performance Optimization domain cover on the Databricks-DE-Pro exam?
- Be able to select the correct Delta command, DLT setting, or SQL warehouse configuration for a stated performance or cost symptom, and know which Databricks feature reports or limits spend. The key is matching the symptom (small files, cold starts, OOM, overspend) to the precise tool rather than a generic tuning step.
- How many questions are in this domain?
- This page lists all 33 Cost and Performance Optimization questions in the Databricks-DE-Pro question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
- What is the best way to practise this domain?
- Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
- Can I practise only Cost and Performance Optimization questions?
- Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.