A developer is writing a Spark Connect application that must run a pandas UDF on a remote Databricks cluster. The code uses the spark.conf.set() API to pass a custom Python module to the executors. Which statement describes the correct behavior?
spark.conf.set() is strictly for Spark configuration properties. It has no mechanism to ship a Python module to the remote executors, so any pandas UDF that imports that module will raise an ImportError. To distribute code with Spark Connect, the module must be installed as a cluster library or uploaded to a workspace path.
Why this answer
spark.conf.set() is designed exclusively for Spark configuration properties and cannot distribute Python code. In a Spark Connect architecture, the client and executors are decoupled, so any custom Python dependencies used by pandas UDFs must be installed on the cluster as libraries or uploaded to a path that the cluster can access. Setting a configuration value will not make the module available.
Exam trap
The trap here is assuming that spark.conf.set() can be used as a generic code-distribution mechanism because it accepts arbitrary key-value pairs.