Databricks-DE-Pro Cost and Performance Optimization Practice Question
Your organization runs numerous batch data engineering pipelines using standard Databricks jobs. Finance reports indicate that compute costs are inflated due to cluster startup times and rigid over-provisioning. Which optimization approach provides the best balance of cost savings and execution reliability for scheduled production batch jobs?
⚠ Common exam trap
Candidates often suggest interactive clusters to avoid configuration complexity. They ignore that interactive clusters bill at higher rates and stay running, failing to leverage the cost-effective nature of job-specific compute.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Migrate scheduled pipeline execution to Databricks Jobs compute leveraging job clusters configured with spot instances for workers.
Transitioning scheduled production batch pipelines from interactive all-purpose clusters to Databricks Jobs compute with spot instance integration delivers massive financial savings. Jobs compute provides lower compute unit pricing compared to interactive workspaces, while spot instances discount infrastructure further, and robust retry mechanisms handle any cloud-level pre-emptions gracefully.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Keep using interactive all-purpose clusters but implement notebook-level python scripts to manually stop clusters when jobs finish.
Why it's wrong here
All-purpose clusters bill interactive DBUs and cannot autoscale per job run, so manual stop scripts leave idle-time charges and add no startup reduction. Job clusters with autoscaling exist precisely for scheduled production batch workloads, terminating on completion and sizing to the task.
- ✗
Refactor all batch jobs to execute exclusively on Serverless SQL Warehouses regardless of workload dependencies.
Why it's wrong here
Serverless SQL Warehouses execute SQL only, so pipelines with Python, Spark or library dependencies cannot run there. They suit BI query workloads needing elastic SQL concurrency, not general batch ETL jobs requiring DataFrame and notebook execution.
- ✓
Migrate scheduled pipeline execution to Databricks Jobs compute leveraging job clusters configured with spot instances for workers.
Why this is correct
Databricks Jobs compute bills at a significantly lower rate than all-purpose compute. Configuring job clusters with spot instance workers leverages spare cloud capacity at steep discounts, while job orchestration automatically provisions and terminates clusters per run.
- ✗
Reduce the executor memory allocation below default recommendations to force Spark to spill data to disk more frequently.
Why it's wrong here
Reducing executor memory below recommended levels causes excessive disk spilling, lengthening runtimes and destabilising jobs rather than cutting cost reliably. It is tempting as a way to shrink cluster footprint, but Spark's memory guidance exists precisely to avoid spill-driven slowdowns. Right-sizing clusters and pools addresses startup and over-provisioning instead.
About these practice questions
This Databricks-DE-Pro question is part of Courseiva's 267-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-DE-Pro practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DE-Pro exam.