Databricks-Spark-Assoc Using Spark SQL Practice Question
A developer runs `spark.sql("SELECT * FROM sales")` and receives an error stating the table or view cannot be found, even though a Parquet directory exists at `dbfs:/mnt/raw/sales/`. The developer wants to query that Parquet data using Spark SQL without moving or copying the files. Which action should the developer take?
⚠ Common exam trap
Many exam-takers confuse metadata-maintenance commands such as `MSCK REPAIR TABLE` or `REFRESH TABLE` with table creation, when those commands only operate on tables that already exist in the catalog.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Create a table using `CREATE TABLE sales USING parquet LOCATION 'dbfs:/mnt/raw/sales/'` and then query it.
Spark SQL resolves table names through the metastore catalog, so a directory on storage is invisible until a table definition points at it. Using `CREATE TABLE ... USING parquet LOCATION ...` registers the path without moving data, enabling immediate querying by name. Commands that repair or refresh metadata presuppose an existing table and cannot create one.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Create a table using `CREATE TABLE sales USING parquet LOCATION 'dbfs:/mnt/raw/sales/'` and then query it.
Why this is correct
Defining a table with `USING parquet` and a `LOCATION` clause registers the existing directory in the metastore so Spark SQL can resolve the name `sales`. Because the data stays in place, no copying occurs, and subsequent queries read the Parquet files directly. This matches the requirement to query existing files without relocation.
- ✗
Set `spark.sql.warehouse.dir` to `dbfs:/mnt/raw/sales/` and rerun the query.
Why it's wrong here
`spark.sql.warehouse.dir` controls where managed tables store their data; it does not make arbitrary directories queryable by name. Changing it would affect future managed-table writes and not register `sales` in the catalog. The query would still fail with a table-not-found error, so this setting does not address the requirement.
- ✗
Execute `REFRESH TABLE sales` so Spark reloads the metadata from disk.
Why it's wrong here
`REFRESH TABLE` invalidates cached metadata for an existing table so new files are visible, but it requires the table to already exist in the catalog. It cannot bootstrap a table definition from a bare directory. Because the developer receives a 'table not found' error, refreshing has no target and will not resolve the issue.
- ✗
Run `MSCK REPAIR TABLE sales` to register the directory with the metastore.
Why it's wrong here
`MSCK REPAIR TABLE` only adds partition metadata to an already-defined table; it cannot create a table from scratch or discover a raw directory that has never been registered. Since the error is that the table does not exist, there is nothing for the repair command to act on. This command is inapplicable to the scenario.
Visual reference
About these practice questions
One of 295 original Databricks-Spark-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.