Courseiva
Analyzing Queries →hardMultiple Choice

Databricks-DA-Assoc Analyzing Queries Practice Question

An analyst is reviewing a slow query and identifies that the 'FileScan' stage is taking most of the time. Which THREE factors could be causing this inefficiency?

⚠ Common exam trap

Students often overlook metadata overhead caused by tiny files, focusing only on compute sizing instead of file management and partition pruning factors.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

The table contains a high number of small files.

Slow FileScans usually point to I/O-related issues. If too many small files are present, the metadata overhead becomes significant. Without proper pruning, the system scans unnecessary data. Z-Ordering or partitioning issues cause the engine to read more data than required. Addressing these factors is vital for analysts, as they directly impact the 'Data Skipping' efficiency of the Delta Lake engine, which is the cornerstone of high-performance analytics in Databricks.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    The table contains a high number of small files.

    Why this is correct

    A high number of small files increases metadata processing time, as the system must open, read, and close many files to retrieve even a small amount of data. This 'small file problem' is a common cause of slow query performance and can be mitigated by using the OPTIMIZE command.

  • ✓

    The query lacks filters on partition columns.

    Why this is correct

    Without filters on partition columns, the engine must perform a full scan of the table instead of reading only relevant directories. This drastically increases the volume of data retrieved from storage, leading to longer execution times, especially on large datasets that would benefit from efficient partition pruning strategies.

  • ✗

    The query uses a cross join.

    Why it's wrong here

    While a cross join is extremely inefficient, it affects the 'Join' stage of the physical plan, not the 'FileScan' stage. The FileScan stage is strictly related to data retrieval from storage, whereas a cross join is a compute-heavy operation that occurs after the data has been read into memory.

  • ✓

    The data is not Z-Ordered on filtered columns.

    Why this is correct

    Z-Ordering organizes data within files to ensure that related values are physically co-located. If a table is not Z-Ordered on columns used in WHERE clauses, the engine cannot skip files effectively, forcing it to read significantly more data than necessary, which results in a slower FileScan process.

  • ✗

    The query is using a User-Defined Function (UDF).

    Why it's wrong here

    UDFs are executed during the row-processing stages of the query. They do not influence the efficiency of reading data from the underlying Parquet/Delta files. While UDFs can be slow, they would manifest as bottlenecks in the evaluation stages, not within the initial FileScan operation of the execution plan.

About these practice questions

One of 291 original Databricks-DA-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-DA-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DA-Assoc exam.