Courseiva
Design for New Solutions →hardMultiple Choice

SAP-C02 Design for New Solutions Practice Question

A healthcare company is designing a new system on AWS to store and analyze patient health records. The system must comply with HIPAA regulations. Data includes structured lab results and unstructured clinical notes. The company needs to run complex SQL queries on the structured data and perform natural language processing (NLP) on the unstructured data. The solution should be cost-effective and minimize administrative overhead. Which solution should a Solutions Architect recommend?

⚠ Common exam trap

SAP-C02 often tests whether candidates confuse general-purpose AI/ML services (Textract, Glue, SageMaker) with purpose-built, HIPAA-eligible medical NLP services like Comprehend Medical, leading them to pick a solution that technically 'processes text' but lacks medical entity extraction and compliance alignment.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Store structured data in Amazon Redshift, store unstructured data in S3, and use Amazon Comprehend Medical for NLP.

Amazon Redshift is a petabyte-scale, columnar data warehouse optimized for complex analytical SQL queries over structured data, which fits the lab results workload. Amazon Comprehend Medical is a HIPAA-eligible, purpose-built NLP service that extracts medical entities (medications, conditions, PHI) from unstructured clinical notes without requiring custom model training. Both services are fully managed, minimizing administrative overhead, and the combination is cost-effective for this mixed workload. S3 provides durable, low-cost storage for the unstructured notes that Comprehend Medical reads directly.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Store structured data in Amazon RDS for PostgreSQL, store unstructured data in S3, and use AWS Glue to run NLP jobs.

    Why it's wrong here

    AWS Glue runs Spark ETL jobs and provides no NLP capability, so clinical notes would need custom model code. Glue is the right choice for cataloguing and transforming data at scale, but the scenario's NLP requirement calls for a managed AI service offering pre-trained language processing.

  • ✗

    Store all data in S3, use Amazon Athena for SQL queries and Amazon Textract for NLP.

    Why it's wrong here

    Amazon Textract extracts text and forms from documents; it performs no natural language processing such as entity or sentiment analysis. Textract is correct for OCR-style document digitisation, but the scenario's NLP requirement needs a managed language AI service, and Athena alone cannot serve both workloads.

  • ✗

    Store structured data in DynamoDB, store unstructured data in S3, use Amazon SageMaker to build custom NLP models.

    Why it's wrong here

    DynamoDB is a key-value store that cannot run the complex SQL queries the lab results demand, and SageMaker requires building and training custom models, adding administrative overhead. DynamoDB suits high-scale key-based access, while SageMaker fits bespoke model development, not pre-trained NLP.

  • ✓

    Store structured data in Amazon Redshift, store unstructured data in S3, and use Amazon Comprehend Medical for NLP.

    Why this is correct

    Redshift handles complex SQL analytics on structured lab results, S3 stores unstructured clinical notes cost-effectively, and Amazon Comprehend Medical provides HIPAA-eligible NLP for extracting medical entities. This combination meets the query, NLP, cost and low-administrative-overhead constraints within a compliant AWS architecture.

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

Courseiva writes every SAP-C02 question from scratch — 984 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This SAP-C02 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the SAP-C02 exam.