A data engineer observes that a transformation query run on a MEDIUM warehouse spends most of its time in the 'Remote Disk Spilling' phase of the Query Profile. The query joins two large tables and performs a large sort. Memory usage shows the warehouse consistently near its limit. The engineer wants the most direct fix that addresses the root cause. Which action should be taken?
Remote disk spilling means the operation exceeded the memory available on its warehouse and had to write intermediate data to remote storage, which is far slower. Scaling up the warehouse adds memory capacity to handle the join and sort working set. This directly addresses the memory shortfall that the Query Profile is showing.
Why this answer
Remote disk spilling indicates that the join and sort working set exceeded warehouse memory and intermediate data was written to remote storage. The most direct remedy is to give the operation more memory by scaling up the virtual warehouse, which increases the memory available per node. Query rewriting, query acceleration, and clustering all fail to address the memory shortfall that the Query Profile is reporting.
Exam trap
The trap here is treating remote spilling as a scan-efficiency problem and reaching for clustering or query acceleration, when the spill happens in memory-intensive join and sort operators.