Databricks-Spark-Assoc Using Spark SQL Practice Question
A developer is using Spark SQL to analyze a DataFrame that contains a column named tags, which holds an array of strings for each row. The developer needs to filter rows where the array contains the string 'urgent' and also produce a new column with the number of elements in the array. Which TWO Spark SQL expressions should be used in the query? (Choose two.)
⚠ Common exam trap
Many exam-takers confuse array inspection functions with generator or aggregate functions, leading to incorrect row-level results or unnecessary query complexity.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
ARRAY_CONTAINS(tags, 'urgent')
ARRAY_CONTAINS provides a direct boolean test for whether an array includes a given value, making it ideal for filtering rows without altering their structure. SIZE returns the element count of an array, which satisfies the need for a new column with the number of tags. Together, they allow the query to filter and augment the DataFrame efficiently. Other functions like EXPLODE change row cardinality, COLLECT_LIST aggregates, and ARRAY_MAX finds a maximum, none of which meet the specific goals.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
ARRAY_MAX(tags)
Why it's wrong here
ARRAY_MAX returns the maximum value in an array, which for strings would be the lexicographically largest element. It does not check for a specific value like 'urgent' nor count elements. This function is useful for numeric arrays to find the peak value, but it is irrelevant to the requirements of filtering by membership or counting array length.
- ✓
ARRAY_CONTAINS(tags, 'urgent')
Why this is correct
ARRAY_CONTAINS is the correct function to test whether an array column includes a specific value. It returns a boolean and works directly on array-typed columns in Spark SQL. In a WHERE clause, it filters rows where the tags array contains 'urgent'. This is the idiomatic and efficient way to perform membership tests on arrays without exploding them first, preserving row granularity.
- ✗
EXPLODE(tags)
Why it's wrong here
EXPLODE is a generator function that transforms an array column into multiple rows, one per element. While it can be used to find arrays containing a value via a filter, it changes the row granularity and requires aggregation to restore the original rows. It is not suitable for simply filtering rows or producing a count column. Using EXPLODE here would complicate the query unnecessarily.
- ✗
COLLECT_LIST(tags)
Why it's wrong here
COLLECT_LIST is an aggregate function that gathers values into an array, typically after a GROUP BY. It does not test for membership or count elements within an existing array. Applying it to an array column would create nested arrays and is not appropriate for filtering or counting. It serves a different purpose: aggregating rows, not inspecting array contents.
- ✓
SIZE(tags)
Why this is correct
SIZE returns the number of elements in an array or map column. For the tags array, it produces an integer count per row. This is the correct function to create a new column with the array length. It handles null arrays by returning -1 in some Spark versions, so developers should be aware of null handling. It directly addresses the requirement for counting elements.
About these practice questions
This Databricks-Spark-Assoc question is part of Courseiva's 295-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.