SAP-C02 Design for New Solutions Practice Question
A company is designing a data lake on AWS using Amazon S3. They need to run SQL queries on the data without moving it to a separate database. Which AWS service should they use?
⚠ Common exam trap
A common mix-up: candidates confuse AWS Glue (which catalogs and transforms data but does not run SQL queries) with Athena, or they assume Amazon EMR is required for SQL-on-S3, overlooking Athena's serverless and direct-query capability.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Amazon Athena
Amazon Athena is a serverless, interactive query service that allows you to run standard SQL queries directly against data stored in Amazon S3 without needing to move or transform the data. It uses Presto under the hood and charges only for the data scanned per query, making it ideal for ad-hoc SQL analysis on a data lake.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Amazon EMR
Why it's wrong here
Amazon EMR runs distributed processing frameworks such as Spark and Hive on provisioned clusters; it does not offer serverless SQL querying directly against S3 without cluster management. It is tempting because EMR can query S3 data, and would be correct for large-scale custom transformation jobs rather than ad hoc SQL.
- ✓
Amazon Athena
Why this is correct
Amazon Athena is a serverless query service that runs standard SQL directly against data in Amazon S3 using the Glue Data Catalog, so no data movement or separate database is needed. This satisfies the requirement to query in place without loading.
- ✗
Amazon Redshift
Why it's wrong here
Amazon Redshift requires data to be loaded into its own cluster storage before querying, which contradicts the requirement to query in place without moving data. It is tempting because Redshift delivers high-performance SQL analytics, and would be correct once data is deliberately loaded into a warehouse.
- ✗
AWS Glue
Why it's wrong here
AWS Glue is a serverless ETL service that catalogs and transforms data; it does not itself provide an interactive SQL query engine over S3. It is tempting because Glue crawls S3 and populates the Data Catalog, and would be correct for building ETL jobs rather than running analytical queries.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
This SAP-C02 question is part of Courseiva's 984-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This SAP-C02 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the SAP-C02 exam.