Databricks-Spark-Assoc Using Spark Connect Practice Question
A data engineer is using Spark Connect from a remote Python client to interact with a Databricks cluster. The engineer wants to understand which operations are executed on the server side versus the client side. Which two statements correctly describe this behavior? (Choose two.)
⚠ Common exam trap
The trap here is assuming that client-side Python code, such as a UDF body, executes locally, when Spark Connect actually ships it to the server for execution.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Calls such as df.filter() and df.select() build a logical plan on the client and are sent to the server for execution.
The correct statements highlight the division of labor in Spark Connect: transformations build a logical plan on the client, while actions like collect trigger server-side execution and return results to the client. UDFs are shipped to the server, and there is no local SparkContext or client-side data cache. Understanding this split is essential for predicting where code runs and where memory is consumed.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Calls such as df.filter() and df.select() build a logical plan on the client and are sent to the server for execution.
Why this is correct
In Spark Connect, DataFrame transformations like filter and select are lazy and only construct an unresolved logical plan on the client. That plan is serialized and sent to the server, where it is analyzed, optimized, and executed. The client does not process data locally for these operations, so the heavy lifting occurs on the remote Spark server.
- ✗
The client maintains a local SparkContext that runs tasks in parallel with the remote server to reduce latency.
Why it's wrong here
Spark Connect clients do not create or maintain a local SparkContext. The architecture intentionally removes the driver from the client. All task scheduling and execution occur on the server. The client only builds logical plans and sends them over gRPC, so there is no local task parallelism to reduce latency.
- ✗
User-defined functions defined with @udf are executed in the client Python process to avoid shipping code to the server.
Why it's wrong here
In Spark Connect, Python UDFs are serialized and sent to the server, where they are executed in Python worker processes alongside the Spark executors. They are not executed in the client process. This ensures the UDF runs where the data resides, but it also means the UDF code must be serializable and available to the server environment.
- ✓
df.collect() brings the result set to the client, so the returned data is materialized in the client process memory.
Why this is correct
When collect() is invoked, Spark Connect sends the plan to the server, executes it, and streams the resulting rows back over gRPC to the client. The client then materializes those rows in its local process memory as Python objects. This is why collect() on a large result set can exhaust client memory even though execution happened remotely.
- ✗
df.show() executes entirely on the client by sampling data from a local cache maintained by Spark Connect.
Why it's wrong here
Spark Connect does not maintain a local data cache on the client for show(). The show() action sends the plan to the server, which computes and returns a formatted preview. The client only receives and displays the small result. There is no client-side cache of DataFrame contents, so the operation depends on the remote server.
About these practice questions
This Databricks-Spark-Assoc question is part of Courseiva's 295-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.