Courseiva

CLF-C02 Cloud Technology and Services Practice Question

A company stores large amounts of data in Amazon S3 and wants to query it using standard SQL without loading it into a database. They need queries to complete in seconds. Which query optimization technique should they apply?

⚠ Common exam trap

Many candidates confuse data transfer optimization (S3 Transfer Acceleration) with query optimization, or assume that CSV's universal compatibility makes it the best choice for performance, ignoring the critical role of columnar formats and partitioning in reducing scan volume.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Convert data to columnar Parquet format and implement partitioning

B is correct because converting data to columnar Parquet format reduces the amount of data scanned by only reading the columns needed for the query, and partitioning further limits the data scanned by filtering on partition keys. This combination enables queries to complete in seconds on Amazon S3 using services like Amazon Athena or Amazon Redshift Spectrum, without loading data into a database.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Store data in CSV format for maximum compatibility

    Why it's wrong here

    CSV stores all columns together in a row-wise format, so Athena must read the entire file even for queries that reference only a few columns. This maximizes the bytes scanned, directly increasing query runtime and cost. Additionally, CSV lacks native compression and predicate pushdown optimizations, making it a poor choice for large-scale analytics in Athena.

  • ✓

    Convert data to columnar Parquet format and implement partitioning

    Why this is correct

    Parquet is a columnar file format that groups columns into separate data blocks, enabling Athena to perform column pruning and skip irrelevant data entirely. Partitioning organizes data into Hive-style directories (e.g., by date or region), allowing Athena to exclude non-matching partitions via partition pruning. Combining these practices drastically reduces the volume of data scanned, often cutting query time from minutes to seconds and lowering costs by over 90%.

  • ✗

    Enable S3 Transfer Acceleration on the bucket

    Why it's wrong here

    Amazon S3 Transfer Acceleration uses AWS edge locations to speed up uploads over the internet to S3 buckets, but it has no effect on how Athena reads data from S3. Athena executes queries over S3 through AWS internal network paths, and query performance is determined by the amount of data scanned, not transfer speed. This option addresses the wrong bottleneck and does nothing to optimize the underlying file layout or data partitioning.

  • ✗

    Move all data to Amazon RDS for faster SQL queries

    Why it's wrong here

    Moving data to Amazon RDS introduces a fully managed relational database, requiring schema design, data loading, and persistent infra-management (provisioning, scaling, patching). Athena is serverless and queries S3 directly without requiring you to load data into a database, so this change adds operational overhead and compromises the architecture's scalability. For analytical queries, RDS is also often slower than columnar formats in Athena because RDS is optimized for transaction processing, not large-scale scans.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

This CLF-C02 question is part of Courseiva's 993-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This CLF-C02 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the CLF-C02 exam.