CLF-C02 Cloud Technology and Services Practice Question
A company stores large amounts of data in Amazon S3 and wants to query it using standard SQL without loading it into a database. They need queries to complete in seconds. Which query optimization technique should they apply?
⚠ Common exam trap
Many candidates confuse data transfer optimization (S3 Transfer Acceleration) with query optimization, or assume that CSV's universal compatibility makes it the best choice for performance, ignoring the critical role of columnar formats and partitioning in reducing scan volume.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Convert data to columnar Parquet format and implement partitioning
B is correct because converting data to columnar Parquet format reduces the amount of data scanned by only reading the columns needed for the query, and partitioning further limits the data scanned by filtering on partition keys. This combination enables queries to complete in seconds on Amazon S3 using services like Amazon Athena or Amazon Redshift Spectrum, without loading data into a database.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Store data in CSV format for maximum compatibility
Why it's wrong here
CSV stores all columns together in a row-wise format, so Athena must read the entire file even for queries that reference only a few columns. This maximizes the bytes scanned, directly increasing query runtime and cost. Additionally, CSV lacks native compression and predicate pushdown optimizations, making it a poor choice for large-scale analytics in Athena.
- ✓
Convert data to columnar Parquet format and implement partitioning
Why this is correct
Parquet is a columnar file format that groups columns into separate data blocks, enabling Athena to perform column pruning and skip irrelevant data entirely. Partitioning organizes data into Hive-style directories (e.g., by date or region), allowing Athena to exclude non-matching partitions via partition pruning. Combining these practices drastically reduces the volume of data scanned, often cutting query time from minutes to seconds and lowering costs by over 90%.
- ✗
Enable S3 Transfer Acceleration on the bucket
Why it's wrong here
Amazon S3 Transfer Acceleration uses AWS edge locations to speed up uploads over the internet to S3 buckets, but it has no effect on how Athena reads data from S3. Athena executes queries over S3 through AWS internal network paths, and query performance is determined by the amount of data scanned, not transfer speed. This option addresses the wrong bottleneck and does nothing to optimize the underlying file layout or data partitioning.
- ✗
Move all data to Amazon RDS for faster SQL queries
Why it's wrong here
Moving data to Amazon RDS introduces a fully managed relational database, requiring schema design, data loading, and persistent infra-management (provisioning, scaling, patching). Athena is serverless and queries S3 directly without requiring you to load data into a database, so this change adds operational overhead and compromises the architecture's scalability. For analytical queries, RDS is also often slower than columnar formats in Athena because RDS is optimized for transaction processing, not large-scale scans.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
This CLF-C02 question is part of Courseiva's 993-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This CLF-C02 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the CLF-C02 exam.