Static IP for Whitelisting with Multi-AZ Deployment Using NLB and ALB
A company wants to implement a data lake on AWS with data from multiple sources. They need to store data in its raw format and allow multiple teams to query it using different tools. Which service should be used as the central storage layer?
⚠ Common exam trap
Candidates often confuse a data lake's raw storage layer with a data warehouse (Redshift) or a transactional database (RDS, DynamoDB), failing to recognize that a data lake requires schema-on-read, object storage, and multi-engine query support, which only S3 provides.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Amazon S3
Amazon S3 is the correct choice because it provides a highly durable, scalable, and cost-effective object storage service that can store data in its raw, native format (e.g., CSV, JSON, Parquet, images). It supports multiple query engines like Amazon Athena, Amazon Redshift Spectrum, and AWS Glue, allowing diverse teams to query the same data using different tools without data movement.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Amazon DynamoDB
Why it's wrong here
DynamoDB is a key-value and document store accessed via APIs, so it cannot hold raw files in native format for arbitrary query engines. It is tempting because DynamoDB scales massively for low-latency application lookups, and would be correct for a high-throughput operational store rather than a data lake.
- ✗
Amazon Redshift
Why it's wrong here
Redshift is a columnar warehouse requiring loaded, schema-modelled data; it cannot serve as the raw-format landing layer queried by many independent tools. It is tempting because Redshift excels at large-scale SQL analytics, and would be correct for the curated query layer sitting on top of a data lake.
- ✓
Amazon S3
Why this is correct
Amazon S3 provides durable, schema-on-read object storage that keeps data in its native raw format, satisfying the requirement to preserve source fidelity. Its decoupling from compute lets Athena, Redshift Spectrum and EMR query the same objects independently, which no single-purpose analytics engine allows.
- ✗
Amazon RDS
Why it's wrong here
Amazon RDS is a managed relational engine for transactional workloads, not a repository for raw, multi-format files queried by diverse analytics tools. It is tempting because RDS handles structured OLTP data well, and would suit a scenario needing a managed PostgreSQL or MySQL database rather than a data lake storage layer.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
This SAP-C02 question is part of Courseiva's 984-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This SAP-C02 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the SAP-C02 exam.