An analyst is reviewing a slow query and identifies that the 'FileScan' stage is taking most of the time. Which THREE factors could be causing this inefficiency?
A high number of small files increases metadata processing time, as the system must open, read, and close many files to retrieve even a small amount of data. This 'small file problem' is a common cause of slow query performance and can be mitigated by using the OPTIMIZE command.
Why this answer
Slow FileScans usually point to I/O-related issues. If too many small files are present, the metadata overhead becomes significant. Without proper pruning, the system scans unnecessary data.
Z-Ordering or partitioning issues cause the engine to read more data than required. Addressing these factors is vital for analysts, as they directly impact the 'Data Skipping' efficiency of the Delta Lake engine, which is the cornerstone of high-performance analytics in Databricks.
Exam trap
Students often overlook metadata overhead caused by tiny files, focusing only on compute sizing instead of file management and partition pruning factors.