Courseiva
Using Spark Connect →mediumMultiple Choice

Databricks-Spark-Assoc Using Spark Connect Practice Question

A developer is building a Spark Connect client application in Python that runs on a local workstation and connects to a remote Databricks cluster. The application must construct a DataFrame from a list of Python dictionaries without requiring the data to be uploaded to cloud storage first. Which approach should the developer use?

⚠ Common exam trap

The trap here is assuming Spark Connect requires data to be staged in remote storage before a DataFrame can be built, when local Python collections are actually serialized and sent directly.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use spark.createDataFrame(list_of_dicts) directly, because Spark Connect serializes local Python collections into Arrow batches and sends them to the server.

Spark Connect allows client-side Python collections to be converted into DataFrames through the standard createDataFrame API. The client serializes the data into Arrow format and streams it to the server over gRPC, so the data need not be staged in cloud storage. Approaches relying on SparkContext, RDD parallelization, or broadcasting fail because Spark Connect deliberately omits the low-level RDD and context APIs.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Write the list of dictionaries to a local JSON file, then call spark.read.json() with the local file path.

    Why it's wrong here

    Spark Connect resolves file paths on the remote cluster's filesystem, not the client's local filesystem. A local path such as file:/home/user/data.json would not exist on the server, so the read would fail. Writing locally and reading remotely does not bridge the two environments without an explicit upload to a path visible to the cluster.

  • ✗

    Broadcast the list of dictionaries with spark.sparkContext.broadcast() and then call spark.createDataFrame() on the broadcast variable.

    Why it's wrong here

    Broadcasting is an RDD-level mechanism that depends on SparkContext, which Spark Connect does not expose to client applications. Moreover, createDataFrame does not accept a broadcast variable as input. This approach misuses both the broadcast API and the DataFrame constructor, so it cannot create the intended DataFrame.

  • ✓

    Use spark.createDataFrame(list_of_dicts) directly, because Spark Connect serializes local Python collections into Arrow batches and sends them to the server.

    Why this is correct

    Spark Connect supports creating DataFrames from local Python collections. The client serializes the list of dictionaries into Apache Arrow record batches and transmits them over gRPC to the server, which reconstructs the DataFrame. No intermediate cloud storage or manual upload step is required, making this the correct approach for the described scenario.

  • ✗

    Use spark.sparkContext.parallelize(list_of_dicts) and convert the resulting RDD to a DataFrame with spark.createDataFrame().

    Why it's wrong here

    Spark Connect does not expose a usable SparkContext for user code; the sparkContext attribute is not supported for RDD operations. Attempting to call parallelize on it raises an error or is unavailable. This legacy RDD-based pattern is incompatible with the Spark Connect architecture and cannot satisfy the requirement.

About these practice questions

Courseiva writes every Databricks-Spark-Assoc question from scratch — 295 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.