A company runs an Amazon Redshift cluster with 10 RA3 nodes. The data warehouse stores 50 TB of data. The company notices that queries are slow and the cluster's storage utilization is high. The data engineer needs to improve query performance and reduce storage costs without changing the cluster's node count. Which action should the engineer take?
Redshift Spectrum queries external tables in Amazon S3 directly, so historical data leaves RA3 managed storage without altering node count. This lowers storage utilisation and cost while letting the cluster focus compute on hot data, addressing both stated constraints.
Why this answer
Redshift Spectrum allows you to query data directly from Amazon S3 without loading it into the cluster. By offloading historical or less-frequently accessed data to S3, you reduce the storage utilization on the RA3 nodes, which frees up managed storage and can improve query performance. This approach also lowers storage costs because S3 is cheaper than Redshift managed storage, and it does not change the node count.
Exam trap
The trap here is that candidates often confuse concurrency scaling (which improves query throughput) with storage optimization, or they assume that changing distribution styles (like DISTSTYLE ALL) will always improve performance, ignoring the storage cost impact in a high-utilization scenario.
How to eliminate wrong answers
Option B is wrong because changing large tables to DISTSTYLE ALL replicates the entire table to every node, which increases storage utilization and can worsen the high storage issue, not reduce it. Option C is wrong because migrating to Dense Compute nodes would change the node type, which violates the constraint of not changing the cluster's node count; also, Dense Compute nodes use local SSD storage and are not designed for the same storage-to-compute ratio as RA3 nodes. Option D is wrong because concurrency scaling adds additional compute capacity to handle more concurrent queries but does not reduce storage utilization or costs; it addresses throughput, not the underlying storage pressure.