Databricks-DE-Pro Data Ingestion and Acquisition Practice Question
You are designing an ingestion pipeline that must handle massive bursts of data at irregular intervals. Which feature should you prioritize to ensure the ingestion process remains cost-effective?
⚠ Common exam trap
Candidates often suggest fixed-size clusters to save money, ignoring that auto-scaling is essential for cost-effectively managing the irregular, bursty nature of the described workload.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use auto-scaling clusters with a minimum of zero workers.
Using Photon-accelerated clusters with auto-scaling is the most effective way to handle bursty workloads. By configuring the cluster to scale out when the queue of files is large and scale down during idle periods, you maximize resource utilization. This approach ensures you have enough compute to meet latency SLAs during bursts while minimizing costs by scaling to zero or minimal size when no data is being ingested, optimizing for both performance and budget.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use a fixed-size cluster with a large number of nodes.
Why it's wrong here
A fixed-size cluster is highly inefficient for bursty workloads. It will either be too small to handle the peak bursts, leading to ingestion lag, or it will be too large during quiet periods, wasting money on idle compute. Dynamic resource allocation is required for optimal cost management in cloud environments.
- ✓
Use auto-scaling clusters with a minimum of zero workers.
Why this is correct
Auto-scaling is the primary tool for managing bursty workloads. Setting the minimum to zero allows the cluster to shut down completely when no ingestion tasks are pending, effectively reducing costs to near zero. When data arrives, the cluster scales up automatically, ensuring throughput requirements are met during high-traffic periods.
- ✗
Run the pipeline continuously on a single-node cluster.
Why it's wrong here
Single-node clusters are not designed for high-throughput ingestion and will quickly hit resource constraints during data bursts. Furthermore, running continuously means you are paying for the compute even when no data is arriving, which is the opposite of a cost-effective strategy for irregular, bursty ingestion workloads.
- ✗
Increase the 'spark.sql.shuffle.partitions' to 10000.
Why it's wrong here
Increasing shuffle partitions is an optimization for query performance, not a cost management strategy for ingestion. Excessive partitioning can actually increase overhead and latency, making the pipeline less efficient. It does nothing to scale the underlying compute capacity in response to fluctuating data volumes from source systems.
About these practice questions
Courseiva writes every Databricks-DE-Pro question from scratch — 267 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-DE-Pro practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DE-Pro exam.