Databricks-Spark-Assoc Using Spark Connect Practice Question
A developer writes a Spark Connect client that creates a DataFrame, calls `df.collect()`, and then reuses the same DataFrame for a second `df.count()`. The cluster is remote. What happens on the second action?
⚠ Common exam trap
The trap here is assuming that reusing the same DataFrame object across actions reuses results, when results are only reused if the DataFrame was explicitly cached or persisted on the server.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
The client sends the logical plan again and the server re-executes the computation unless the DataFrame was cached.
In Spark Connect, a DataFrame is a client-side logical plan builder, not a materialized dataset. Each action serializes the current plan and sends it to the server, which plans and executes it. Because no cache or persist was invoked, the second action recomputes the result from the source. Reusing the DataFrame object is valid; it simply does not reuse prior results.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
The server automatically detects the repeated plan and returns a cached count without recomputation.
Why it's wrong here
Spark does not automatically cache results of prior actions. While the optimizer may reuse exchanges within a single job, separate actions are separate jobs and do not share materialized results unless a cache or persist was requested. The server therefore recomputes the count unless the DataFrame was explicitly cached, which was not the case here.
- ✓
The client sends the logical plan again and the server re-executes the computation unless the DataFrame was cached.
Why this is correct
Spark Connect is lazy: each action sends the accumulated logical plan to the server, which plans and executes it. Without an explicit `cache()` or `persist()`, the second action re-runs the computation from source. The DataFrame object on the client is only a plan builder; it holds no materialized data, so the server must recompute the result for `count()`.
- ✗
The client raises an error because a DataFrame cannot be used for multiple actions in Spark Connect.
Why it's wrong here
Spark Connect DataFrames are reusable for multiple actions; they are immutable plan descriptions. There is no restriction preventing a second action on the same DataFrame. The developer can call `collect()` and then `count()` on the same object, and each action triggers a separate server-side execution, so no error is raised.
- ✗
The client reuses cached results from the first action because the DataFrame is immutable.
Why it's wrong here
Immutability of the DataFrame definition does not imply result caching. Spark Connect does not cache results on the client, and unless the DataFrame was explicitly cached or persisted, the second action triggers a new execution on the server. The first `collect()` materialized rows only in the client process for that call; those rows are not reused by later actions.
About these practice questions
Courseiva writes every Databricks-Spark-Assoc question from scratch — 295 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.