Courseiva

Databricks-Spark-Assoc Troubleshooting and Tuning DataFrame Apps Practice Question

A Databricks notebook job calls a PySpark UDF built on a Python function that performs string parsing. The job completes successfully on a small sample, but on the full production dataset it fails with a PythonException and the executor logs show high garbage collection time. Which change is the most appropriate first step to make the job reliable without changing business logic?

⚠ Common exam trap

The trap here is assuming that more executor memory or cores will fix a Python UDF hotspot, when the real cost is the per-row Python worker and serialization overhead.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Replace the Python UDF with a Spark SQL built-in function such as regexp_extract or split, which runs inside the JVM and avoids Python serialization and per-row interpreter overhead.

A scalar Python UDF forces each row through a Python worker, incurring serialization, interpreter startup, and per-row overhead that shows up as GC pressure at scale. Rewriting the logic with Spark SQL built-in functions keeps execution in the JVM with whole-stage code generation, eliminating that overhead and typically resolving both the reliability and performance symptoms. It preserves the business logic while removing the root cause rather than masking it with more resources.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Set spark.sql.execution.arrow.pyspark.enabled to true so the UDF uses Arrow for every row it processes.

    Why it's wrong here

    Arrow-based conversion accelerates data transfer between the JVM and Python, but it applies to createDataFrame, toPandas, and Pandas UDFs, not scalar Python UDFs. A plain Python UDF still runs one row at a time through the Python worker regardless of Arrow. Enabling the Arrow setting here would not remove the per-row interpreter cost or the PythonException. A Pandas UDF would be the Arrow-relevant change.

  • ✗

    Mark the UDF with the deterministic flag and register it with spark.udf.register so Spark can cache its results automatically.

    Why it's wrong here

    Marking a UDF deterministic only tells Spark the function returns the same output for the same input, which enables certain optimizations but does not cache results. Spark does not automatically cache UDF outputs; caching must be requested explicitly with cache or persist. It also does nothing to reduce Python worker overhead. This is a misunderstanding of what the deterministic flag provides.

  • ✓

    Replace the Python UDF with a Spark SQL built-in function such as regexp_extract or split, which runs inside the JVM and avoids Python serialization and per-row interpreter overhead.

    Why this is correct

    Built-in Spark SQL functions execute in the JVM with code generation, so they avoid Python worker startup, serialization, and interpreter overhead per row. For string parsing, regexp_extract, split, or transform are usually equivalent to the Python logic. This directly reduces executor memory pressure and GC time while preserving the same results. It is the least invasive change and the standard tuning step for Python UDF hotspots.

  • ✗

    Increase spark.executor.memory and spark.executor.cores so each executor has more headroom to run the Python workers.

    Why it's wrong here

    Adding executor memory or cores does not remove the per-row Python serialization and interpreter cost; it only postpones the failure and increases cluster cost. The PythonException indicates the UDF itself is throwing on unexpected input, which more memory cannot fix. This change also risks longer GC pauses because larger heaps take longer to collect. It is not the appropriate first step.

Visual reference

Client Recursive Resolver Root DNS (13 root servers) TLD DNS (.com, .org, …) Authoritative example.com query IP addr answer

About these practice questions

Courseiva writes every Databricks-Spark-Assoc question from scratch — 295 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.