Databricks-Spark-Assoc Using Spark SQL Practice Question
A developer is writing a PySpark script that must run a SQL statement against DataFrames already registered as temporary views named orders and returns. The developer wants the query to use Spark SQL syntax while returning a DataFrame that can be further transformed with the DataFrame API. Which call accomplishes this?
⚠ Common exam trap
Watch out — candidates often confuse the catalog interface, which manages metadata such as tables and functions, with the SQL execution entry point on the SparkSession.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
spark.sql("SELECT o.order_id, r.amount FROM orders o JOIN returns r ON o.order_id = r.order_id")
SparkSession.sql runs a SQL string against views registered in the session catalog and returns a DataFrame, allowing the result to be combined with DataFrame transformations. Because orders and returns are temporary views in the same session, the query resolves and returns a join result that can be further processed. The other calls reference methods that do not exist on those objects.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
spark.sql("SELECT o.order_id, r.amount FROM orders o JOIN returns r ON o.order_id = r.order_id")
Why this is correct
spark.sql executes a SQL string against the current catalog and returns a DataFrame, so the result can be chained with DataFrame transformations such as filter or withColumn. Because orders and returns are registered as temporary views, they resolve in the session catalog. This is the standard way to mix SQL and the DataFrame API in one pipeline.
- ✗
spark.catalog.sql("SELECT o.order_id, r.amount FROM orders o JOIN returns r ON o.order_id = r.order_id")
Why it's wrong here
There is no spark.catalog.sql method. The catalog object exposes operations such as listTables, tableExists, and refreshTable, but it does not execute arbitrary SQL. Attempting to call it raises an attribute error rather than returning a DataFrame. SQL execution belongs to the SparkSession, not the catalog interface.
- ✗
spark.executeSql("SELECT o.order_id, r.amount FROM orders o JOIN returns r ON o.order_id = r.order_id")
Why it's wrong here
SparkSession does not expose an executeSql method. Methods such as executeSql exist in other database drivers, but the SparkSession API uses spark.sql for SQL execution. Calling spark.executeSql would raise an attribute error. The developer must use the documented entry point to run SQL against registered views.
- ✗
spark.sqlContext.runSql("SELECT o.order_id, r.amount FROM orders o JOIN returns r ON o.order_id = r.order_id")
Why it's wrong here
The SQLContext object does not provide a runSql method. SQL execution in modern Spark goes through spark.sql on the SparkSession. SQLContext is largely a legacy compatibility layer, and its available methods do not include runSql, so this call would fail rather than return a DataFrame. The correct entry point is the SparkSession method.
Visual reference
About these practice questions
Courseiva writes every Databricks-Spark-Assoc question from scratch — 295 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.