A company is using AWS Glue to prepare data for a machine learning pipeline. The source data is in an Amazon S3 bucket in CSV format. The data scientist wants to convert the data to Parquet format and partition it by date. Which AWS Glue feature should be used to optimize the data for query performance and reduce storage costs?
Trap 1: Use Amazon Athena to convert the data to JSON format and store it…
Athena is a query service, not a data transformation service.
Trap 2: Use AWS Glue to convert the data to Apache Hive format.
Hive format is not a standard file format; Parquet is columnar and efficient.
Trap 3: Use Apache Spark DataFrame to write the data as CSV with Snappy…
CSV is not columnar and does not optimize query performance as well as Parquet.
- A
Use Amazon Athena to convert the data to JSON format and store it in S3.
Why it fails: Athena is a query service, not a data transformation service.
- B
Use AWS Glue DynamicFrame to repartition the data and write it as Parquet.
AWS Glue DynamicFrames extend Spark DataFrames with schema flexibility and explicit write options, letting you convert CSV to columnar Parquet while partitioning by date in a single job. Parquet's columnar compression cuts storage costs, and date partitioning enables partition pruning, satisfying the query-performance and cost-reduction requirements.
- C
Use AWS Glue to convert the data to Apache Hive format.
Why it fails: Hive format is not a standard file format; Parquet is columnar and efficient.
- D
Use Apache Spark DataFrame to write the data as CSV with Snappy compression.
Why it fails: CSV is not columnar and does not optimize query performance as well as Parquet.