Courseiva

Databricks-DE-Assoc Troubleshooting, Monitoring, and Optimization Practice Question

A junior data engineer notices that a scheduled Databricks job running a heavy ETL notebook is failing intermittently due to cluster driver out-of-memory errors. Which TWO configuration changes or architectural adjustments should be implemented to resolve this issue? (Select exactly TWO)

⚠ Common exam trap

Many candidates assume scaling out worker nodes fixes driver memory issues. However, worker nodes process distributed data partitions independently, whereas the driver coordinates execution and collects results, meaning worker scaling has no direct impact on driver memory pressure.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Select a larger driver node instance type with more memory capacity.

Driver out-of-memory errors typically occur when the driver node collects too much data into local memory using actions like collect() or handles excessive broadcast joins. Upgrading to a driver node with more RAM provides immediate headroom, while refactoring code to avoid pulling massive datasets to the driver prevents memory exhaustion fundamentally.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Increase the maximum number of worker nodes in the autoscaling cluster configuration.

    Why it's wrong here

    Autoscaling adds executors to process partitions, but driver memory is fixed by the driver node size; more workers cannot relieve driver heap pressure from collect actions or large broadcast variables. Increasing worker count is correct when executor-side parallelism, not driver memory, is the bottleneck.

  • ✓

    Select a larger driver node instance type with more memory capacity.

    Why this is correct

    Driver out-of-memory errors stem from insufficient memory on the driver node. Selecting a larger driver instance type with more memory capacity directly addresses that constraint, giving the notebook enough headroom to complete its ETL workload.

  • ✓

    Refactor the PySpark notebook code to avoid using collect() on large DataFrames.

    Why this is correct

    Calling collect() pulls the entire DataFrame into the driver's memory, which causes driver out-of-memory failures on large datasets. Refactoring to avoid collect() — using distributed writes or aggregations instead — removes that memory pressure at its source.

  • ✗

    Enable Delta caching on the cluster workers to offload data from the driver.

    Why it's wrong here

    Delta caching stores remote Parquet data on worker local disks, which accelerates repeated reads but does nothing to reduce driver heap consumption, since the driver still collects results and holds job state. It would be the right choice when repeated scans of the same Delta table dominate runtime.

  • ✗

    Decrease the Spark SQL shuffle partitions default setting to a lower number.

    Why it's wrong here

    Lowering spark.sql.shuffle.partitions reduces task count, which can worsen skew per partition and increase per-task memory, while driver OOM stems from driver-side collection and state. Raising this value is the correct tuning when shuffle partitions are too few and tasks are oversized.

Visual reference

Client Recursive Resolver Root DNS (13 root servers) TLD DNS (.com, .org, …) Authoritative example.com query IP addr answer

About these practice questions

This Databricks-DE-Assoc question is part of Courseiva's 276-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-DE-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DE-Assoc exam.