Databricks-Spark-Assoc Using Spark Connect Practice Question
A developer is building a Spark Connect client application in Python that runs on a local workstation and connects to a remote Databricks cluster. The application must construct a DataFrame from a list of Python dictionaries without requiring the data to be uploaded to cloud storage first. Which approach should the developer use?
⚠ Common exam trap
The trap here is assuming Spark Connect requires data to be staged in remote storage before a DataFrame can be built, when local Python collections are actually serialized and sent directly.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use spark.createDataFrame(list_of_dicts) directly, because Spark Connect serializes local Python collections into Arrow batches and sends them to the server.
Spark Connect allows client-side Python collections to be converted into DataFrames through the standard createDataFrame API. The client serializes the data into Arrow format and streams it to the server over gRPC, so the data need not be staged in cloud storage. Approaches relying on SparkContext, RDD parallelization, or broadcasting fail because Spark Connect deliberately omits the low-level RDD and context APIs.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Write the list of dictionaries to a local JSON file, then call spark.read.json() with the local file path.
Why it's wrong here
Spark Connect resolves file paths on the remote cluster's filesystem, not the client's local filesystem. A local path such as file:/home/user/data.json would not exist on the server, so the read would fail. Writing locally and reading remotely does not bridge the two environments without an explicit upload to a path visible to the cluster.
- ✗
Broadcast the list of dictionaries with spark.sparkContext.broadcast() and then call spark.createDataFrame() on the broadcast variable.
Why it's wrong here
Broadcasting is an RDD-level mechanism that depends on SparkContext, which Spark Connect does not expose to client applications. Moreover, createDataFrame does not accept a broadcast variable as input. This approach misuses both the broadcast API and the DataFrame constructor, so it cannot create the intended DataFrame.
- ✓
Use spark.createDataFrame(list_of_dicts) directly, because Spark Connect serializes local Python collections into Arrow batches and sends them to the server.
Why this is correct
Spark Connect supports creating DataFrames from local Python collections. The client serializes the list of dictionaries into Apache Arrow record batches and transmits them over gRPC to the server, which reconstructs the DataFrame. No intermediate cloud storage or manual upload step is required, making this the correct approach for the described scenario.
- ✗
Use spark.sparkContext.parallelize(list_of_dicts) and convert the resulting RDD to a DataFrame with spark.createDataFrame().
Why it's wrong here
Spark Connect does not expose a usable SparkContext for user code; the sparkContext attribute is not supported for RDD operations. Attempting to call parallelize on it raises an error or is unavailable. This legacy RDD-based pattern is incompatible with the Spark Connect architecture and cannot satisfy the requirement.
About these practice questions
Courseiva writes every Databricks-Spark-Assoc question from scratch — 295 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.