A company runs a data lake on Amazon S3 with AWS Glue and Amazon Athena. The data engineer notices that queries are slow and scanning large amounts of data. Which THREE actions should the engineer take to optimize query performance and reduce costs?
Compression shrinks file sizes on S3, so Athena scans fewer bytes per query, cutting both runtime and per-terabyte scan costs. This directly addresses the large-data-scanning problem in the stem while remaining readable by Glue and Athena.
Why this answer
Option C is correct because compressing data files with gzip or snappy reduces the total bytes stored and scanned, and Athena charges and performs based on data scanned, so smaller files lower both query latency and cost. Option D is correct because partitioning the S3 data by frequently filtered columns such as date or region enables partition pruning, so Athena reads only the relevant prefixes instead of scanning the entire table. Option E is correct because columnar formats like Parquet or ORC let Athena read only the columns referenced in the query and provide better compression and predicate pushdown, dramatically reducing scanned data compared to row-based formats like CSV or JSON.
Option A is not appropriate because increasing the Athena query timeout only allows long-running queries to finish; it does not reduce the amount of data scanned or improve performance. Option B is not appropriate because adding DPUs to the Glue job speeds up ETL processing, not Athena query performance or the volume of data scanned at query time.
Exam trap
The trap is selecting scaling options (timeout, DPUs) instead of data organization techniques; the exam tests that Athena performance is primarily about reducing data scanned via partitioning, columnar formats, and compression.