Courseiva

AWS Certified Data Engineer Associate DEA-C01 (DEA-C01) — Questions 1201–1275

1321 questions total · 18pages · All types, answers revealed

Page 16

Page 17 of 18

Page 18
1201
MCQmedium

A data engineer notices that an Amazon Redshift cluster is running low on disk space. The cluster has three nodes of type dc2.large. Which action will increase the available storage capacity?

A.Increase the number of nodes in the cluster.
B.Mount an Amazon S3 bucket as a file system to store data.
C.Change the volume type to Provisioned IOPS SSD (io1) to increase capacity.
D.Enable automatic compression on the tables.
AnswerA

Adding nodes to the dc2.large cluster increases total storage because each node contributes its own local disk capacity; Redshift distributes data across all nodes. Resizing node type or other settings does not add storage as directly.

Why this answer

Amazon Redshift stores data on the local instance store volumes attached to each node in the cluster. With a dc2.large cluster, each node provides approximately 160 GB of SSD storage. Adding nodes increases the total available storage linearly because each new node contributes its local storage to the cluster.

Therefore, increasing the number of nodes is the correct way to expand disk capacity.

Exam trap

The trap here is that candidates may confuse Redshift's local storage model with EBS-backed storage, leading them to think they can change volume types or mount external storage like S3, when in fact Redshift relies solely on the aggregate of each node's local instance store for persistent data.

How to eliminate wrong answers

Option B is wrong because Amazon S3 is an object storage service and cannot be mounted as a file system directly to Redshift; Redshift can only load data from S3 via COPY commands or external tables using Redshift Spectrum, but S3 does not expand the local disk space of the cluster. Option C is wrong because Provisioned IOPS SSD (io1) is an EBS volume type used for Amazon EC2 instances, not for Redshift nodes; Redshift dc2.large nodes use local instance store SSDs, and volume type cannot be changed. Option D is wrong because enabling automatic compression on tables optimizes storage efficiency by reducing the size of data on disk, but it does not increase the total available storage capacity of the cluster; it only helps use existing space more efficiently.

1202
MCQeasy

A company's security policy states that no S3 bucket in the data platform account may ever be made public, even accidentally. A data engineer must implement a guardrail that blocks any attempt to set a public bucket ACL or public bucket policy, regardless of who makes the change. Which solution enforces this requirement?

A.Configure AWS Config with the s3-bucket-public-read-prohibited and s3-bucket-public-write-prohibited managed rules and rely on their evaluations.
B.Turn on S3 server access logging and create an Amazon EventBridge rule that notifies the security team when a public ACL is applied.
C.Create an S3 bucket policy that denies s3:PutBucketAcl and s3:PutBucketPolicy for all principals in the account.
D.Enable S3 Block Public Access at the account level and ensure the setting remains on for all existing and future buckets.
AnswerD

Account-level S3 Block Public Access applies to every bucket in the account, including buckets created later, and overrides any public ACL or policy that would otherwise grant public access. It can be enforced through AWS Organizations so member accounts cannot turn it off. This directly satisfies the requirement to block public access regardless of which principal attempts the change, with no per-bucket maintenance.

Why this answer

S3 Block Public Access at the account level prevents public ACLs and policies for every bucket, including future ones, and can be locked down via AWS Organizations. Detective controls such as AWS Config rules or logging only report exposure after the fact. The requirement is preventive, account-wide, and independent of who attempts the change, which the account-level block setting delivers.

Exam trap

The trap here is choosing a detective control such as an AWS Config rule or logging, when the requirement demands prevention before public access is granted.

1203
MCQhard

A data engineer is building an AWS Glue ETL job that reads from an Amazon S3 bucket containing nested JSON files. The job must flatten the nested structure and write the output to Amazon Redshift. The engineer notices that some JSON records have missing fields and inconsistent schemas. Which AWS Glue feature should be used to handle these inconsistencies and ensure the job does not fail?

A.Use a Glue DataFrame with a predefined schema and set the job to ignore malformed records.
B.Use Amazon Kinesis Data Firehose to transform the JSON before the Glue job reads it.
C.Use AWS Glue crawlers to infer the schema and then manually edit the Data Catalog table to add missing columns.
D.Use a Glue DynamicFrame with the resolveChoice transform to handle schema inconsistencies.
AnswerD

DynamicFrame and resolveChoice are designed to handle schema variability in semi-structured data. They allow you to specify how to resolve ambiguous types or missing fields, such as casting to a specific type or dropping nulls. This prevents job failures due to inconsistent schemas and is the recommended approach for nested JSON with varying structures.

Why this answer

AWS Glue DynamicFrames are specifically built to handle semi-structured data with evolving schemas. The resolveChoice transform lets you resolve ambiguous or conflicting schema types, such as choosing a specific data type or dropping nulls. This ensures the ETL job can process nested JSON with missing fields without failing, making it the correct choice for this scenario.

Exam trap

The trap here is thinking that a predefined schema or manual catalog edits can handle runtime schema variations, but they do not provide the flexibility needed for inconsistent JSON.

1204
MCQmedium

A company is ingesting streaming data into Kinesis Data Streams. The consumer application experiences high latency due to a single shard bottleneck. What is the most effective way to reduce latency?

A.Increase the number of shards in the data stream.
B.Wait for automatic scaling to add shards.
C.Use the Kinesis Client Library (KCL) to process records.
D.Switch to Amazon Kinesis Data Firehose.
AnswerA

A single shard caps throughput and processing parallelism, creating the bottleneck. Adding shards spreads records across more consumers, raising aggregate throughput and reducing per-record latency, directly addressing the stem's single-shard constraint. Resharding via UpdateShardCount is the standard remedy.

Why this answer

Increasing the number of shards in the Kinesis data stream distributes the incoming records across more shards, allowing the consumer application to process them in parallel and reducing the per-shard bottleneck that causes high latency. Each shard supports up to 1 MB/s or 1,000 records/s ingress, so adding shards directly increases throughput capacity. This is the most direct and effective way to address a single-shard bottleneck.

Exam trap

The trap is confusing KCL (a consumer-side library for coordination) with a throughput solution — KCL does not add capacity, only shard count does.

How to eliminate wrong answers

Option B is wrong because Kinesis Data Streams does not automatically scale shards; on-demand mode can scale automatically, but the question describes a provisioned stream with a shard bottleneck, and waiting is not an effective action. Option C is wrong because the Kinesis Client Library helps with consumer coordination and checkpointing but does not increase the stream's throughput capacity; it cannot fix a shard-level bottleneck. Option D is wrong because Kinesis Data Firehose is a delivery service for loading data into destinations, not a replacement for a low-latency consumer application, and it does not solve the shard bottleneck.

1205
MCQmedium

A data engineer is designing a data lake on Amazon S3. Data is ingested from multiple sources in JSON format. The engineer needs to optimize query performance for Amazon Athena while minimizing storage costs. Which storage strategy should the engineer use?

A.Store data as CSV files in a single S3 bucket without prefixes.
B.Convert data to Parquet format and partition by date.
C.Store data as JSON files in a single prefix without partitioning.
D.Store compressed JSON files in Amazon S3 Glacier.
AnswerB

Parquet is columnar and compressed, so Athena scans and bills far less data than JSON. Partitioning by date enables partition pruning, restricting each query to relevant prefixes. Together these cut query latency and S3 storage cost.

Why this answer

Parquet is a columnar storage format that significantly reduces data scan volume in Amazon Athena, which charges per byte scanned. Partitioning by date further limits the data scanned to only relevant partitions, optimizing both query performance and cost. JSON and CSV are row-based formats that require full scans, and Glacier is unsuitable for interactive querying.

Exam trap

The trap here is that candidates assume JSON or CSV are acceptable for Athena due to their simplicity, overlooking that columnar formats like Parquet are required for cost-efficient querying in AWS's pay-per-scan model.

How to eliminate wrong answers

Option A is wrong because CSV files are row-based and lack compression, leading to higher storage costs and larger data scans in Athena, and storing them without prefixes prevents partition pruning. Option C is wrong because JSON files are also row-based and verbose, resulting in inefficient queries and higher costs, and a single prefix without partitioning forces full table scans. Option D is wrong because Amazon S3 Glacier is designed for archival storage with retrieval times of minutes to hours, making it incompatible with Athena's requirement for immediate data access.

1206
MCQmedium

A data engineer maintains an AWS Glue ETL job that reads from an Amazon Kinesis Data Streams stream and writes to Amazon S3 in Parquet format. The job has been running successfully, but after the data volume increased threefold, the job now fails with an error stating that the Glue job's bookmarks are not advancing and the job is reprocessing old data. The engineer has enabled job bookmarks with the default settings. Which action should the engineer take to resolve the issue?

A.Disable job bookmarks and instead use an AWS Lambda function to track the last processed Kinesis sequence number in Amazon DynamoDB.
B.Verify that the Kinesis Data Streams source is configured with the correct stream name and that the job's transformation does not include an unsupported operation that blocks bookmark propagation, such as a custom transformation that does not preserve the bookmark state.
C.Increase the number of AWS Glue DPUs allocated to the job and enable auto-scaling.
D.Change the output format from Parquet to JSON, as Parquet does not support job bookmarks with Kinesis sources.
AnswerB

Job bookmarks rely on the source and transformations to propagate state. If a transformation, such as a custom code node or a filter that changes the record order, does not properly pass along the bookmark, the job will reprocess data. Ensuring the source configuration is correct and that transformations are bookmark-compatible is essential to resolve the issue.

Why this answer

Job bookmarks in AWS Glue track the last processed data to avoid reprocessing. When bookmarks fail to advance, it is typically due to an unsupported transformation or a misconfigured source. Ensuring the Kinesis source is correctly configured and that all transformations preserve bookmark state will allow the job to process only new data and advance the bookmark.

Exam trap

The trap here is assuming that scaling resources or changing the output format will fix bookmark issues, when the root cause is usually a transformation that breaks bookmark propagation.

1207
MCQhard

A company is ingesting streaming data from social media feeds using Amazon Kinesis Data Streams. The data volume peaks at 10,000 records per second, and each record is up to 1 KB. The company needs to archive the raw data in Amazon S3 in near real-time and also make it available for real-time analytics using Amazon Kinesis Data Analytics. What is the MOST efficient architecture to meet these requirements?

A.Use Kinesis Data Streams as the ingestion point. Use Kinesis Data Firehose to read from the stream, convert to Parquet, and write to S3. Use a Lambda function to send data to Kinesis Data Analytics.
B.Use Kinesis Data Streams as the ingestion point. Use a Lambda function to read from the stream, write to S3, and send data to Kinesis Data Analytics.
C.Use two Kinesis Data Streams: one for S3 delivery and one for Kinesis Data Analytics.
D.Use Kinesis Data Streams as the ingestion point. Use Kinesis Data Firehose to read from the stream and write to S3. Use Kinesis Data Analytics to read directly from the same stream.
AnswerD

Kinesis Data Firehose consumes directly from the stream and buffers records before delivering them to Amazon S3, satisfying the near real-time archival requirement without custom consumer code. Kinesis Data Analytics reads the same stream independently, so both consumers run in parallel. This decouples archival from analytics, avoiding duplicate ingestion and extra compute.

Why this answer

Kinesis Data Streams can serve as a single ingestion point, with Kinesis Data Firehose reading from the stream to deliver data to S3 (with optional transformation) and Kinesis Data Analytics reading directly from the same stream for real-time analytics. This avoids unnecessary duplication of streams or Lambda-based processing, which would add latency and complexity. The architecture is the most efficient as it leverages native integrations without intermediate compute.

Exam trap

The trap here is that candidates often overcomplicate the architecture by adding unnecessary Lambda functions or duplicate streams, not realizing that Kinesis Data Firehose and Kinesis Data Analytics can both consume from the same Kinesis Data Stream natively.

How to eliminate wrong answers

Option A is wrong because it suggests using a Lambda function to send data to Kinesis Data Analytics, which is unnecessary and introduces additional cost and latency; Kinesis Data Analytics can read directly from the Kinesis Data Stream. Option B is wrong because using a Lambda function to write to S3 and send data to Kinesis Data Analytics adds processing overhead and potential throughput limitations, whereas Kinesis Data Firehose is purpose-built for streaming to S3 with near-real-time delivery. Option C is wrong because using two separate Kinesis Data Streams is redundant and increases cost and management overhead; a single stream can be consumed by both Firehose and Kinesis Data Analytics simultaneously.

1208
Multi-Selecteasy

Which TWO AWS services can be used to automatically back up an Amazon RDS for SQL Server DB instance? (Choose TWO.)

Select 2 answers
A.AWS Database Migration Service (DMS)
B.AWS Data Pipeline
C.Amazon RDS automated backups
D.Amazon S3
E.AWS Backup
AnswersC, E

RDS automated backups capture daily snapshots plus transaction logs for SQL Server, enabling point-in-time recovery within the retention window. This is native, automatic and requires no additional service, directly satisfying the requirement to back up the instance without manual intervention.

Why this answer

Amazon RDS automated backups (C) are a native RDS feature that automatically takes daily snapshots of the DB instance during the backup window and continuously archives transaction logs, enabling point-in-time recovery for RDS for SQL Server. AWS Backup (E) is a fully managed, policy-driven backup service that supports RDS as a protected resource, allowing you to define backup plans with schedules and retention rules that trigger automated RDS snapshots. AWS DMS (A) is a replication and migration service, not a backup mechanism, so it does not create restorable backups.

AWS Data Pipeline (B) is an orchestration service for data movement and transformation, not a backup tool for RDS. Amazon S3 (D) is object storage that can hold exported backups, but it does not itself automatically back up an RDS for SQL Server DB instance.

Exam trap

The trap here is that candidates often confuse AWS Backup with a service that only works for on-premises or EC2 backups, or mistakenly think DMS or Data Pipeline can handle automated backups, when in fact only RDS automated backups and AWS Backup provide native, automatic backup capabilities for RDS for SQL Server.

1209
MCQmedium

A company uses Amazon Kinesis Data Streams to ingest real-time financial data. The security team requires that all data be encrypted at rest using a customer-managed AWS KMS key, and that the key be rotated annually. The data engineer needs to configure the Kinesis stream to meet these requirements. Which combination of actions should the data engineer take?

A.Enable server-side encryption on the Kinesis stream using the StartStreamEncryption API with the customer-managed KMS key, and enable automatic key rotation on the KMS key.
B.Enable encryption at rest by setting the Kinesis stream's EncryptionType to KMS and specifying a customer-managed key, then manually rotate the key by creating a new key and updating the stream configuration every year.
C.Use AWS CloudFormation to deploy the Kinesis stream with the KmsKeyId property set to a customer-managed key, and set the EnableKeyRotation property to true on the key resource.
D.Configure the Kinesis stream to use AWS-managed KMS keys for encryption, and create a custom AWS Lambda function that rotates the key every 365 days.
AnswerA

The StartStreamEncryption API enables server-side encryption for a Kinesis stream using a specified KMS key. By specifying a customer-managed key, the stream data is encrypted at rest with that key. Enabling automatic key rotation on the KMS key ensures the key is rotated annually, meeting the requirement. This is the correct and direct way to achieve both encryption and rotation.

Why this answer

To encrypt a Kinesis data stream at rest with a customer-managed KMS key, you use the StartStreamEncryption API, specifying the stream and the KMS key. This enables server-side encryption. To rotate the key annually, you enable automatic key rotation on the customer-managed KMS key.

AWS KMS automatically rotates the key material every year when automatic rotation is enabled. This combination meets both requirements with minimal effort.

Exam trap

The trap here is thinking that AWS-managed KMS keys can be rotated by the customer or that manual rotation is necessary, when automatic rotation is available for customer-managed keys.

1210
MCQmedium

A data engineer has an AWS Glue job that reads JSON files from Amazon S3, applies transformations, and writes Parquet files to another S3 location. The job runs daily and takes about 2 hours. Recently, the job has been failing intermittently with an error indicating that the job bookmark is not being updated correctly, causing duplicate processing of some files. The engineer needs to ensure that only new files are processed on each run. Which action should the engineer take to resolve this issue?

A.Modify the AWS Glue job to use a custom transformation that filters files based on their last modified timestamp.
B.Enable job bookmarks in the AWS Glue job configuration and ensure that the input source is an S3 data store with the correct path.
C.Change the output format to Parquet and enable compression to reduce the chance of bookmark errors.
D.Increase the number of AWS Glue DPUs to improve job performance and avoid timeouts that cause bookmark failures.
AnswerB

Job bookmarks are the AWS Glue feature that tracks previously processed data. Enabling them in the job configuration and specifying the correct S3 path allows Glue to persist state and skip already-processed files on subsequent runs. This directly addresses duplicate processing by maintaining a bookmark for the S3 source.

Why this answer

AWS Glue job bookmarks maintain state to track data already processed. When enabled with an S3 source, Glue uses the bookmark to process only new files. The intermittent failures and duplicate processing indicate the bookmark is not being updated, likely because it is not enabled or misconfigured.

Enabling and correctly configuring job bookmarks resolves the issue.

Exam trap

The trap here is assuming that performance tuning or output format changes can resolve data duplication, when the root cause is the missing or misconfigured job bookmark feature.

1211
MCQmedium

A data engineer is using AWS Glue to catalog data stored in Amazon S3. The data is in Parquet format and partitioned by year, month, and day. The engineer needs to ensure that AWS Glue crawlers correctly identify the partitions and that Amazon Athena queries can efficiently prune partitions. Which action should the engineer take?

A.Store the data in a directory structure like s3://bucket/year=2023/month=01/day=01/, and run the crawler with the default settings to automatically detect partitions.
B.Enable AWS Glue Data Catalog encryption and set the 'classification' property to 'parquet' in the table definition.
C.Configure the crawler to use a custom classifier that recognizes the partition structure, and set the table property 'partition_filtering.enabled' to true.
D.Manually create the table in the AWS Glue Data Catalog with partition keys, and use AWS Glue ETL jobs to add partitions as new data arrives.
AnswerA

AWS Glue crawlers automatically detect Hive-style partitions when the S3 path follows the key=value format, such as year=2023/month=01/day=01. This allows the crawler to populate the AWS Glue Data Catalog with partition metadata. Athena can then use this metadata to prune partitions during queries, improving performance and reducing cost by scanning only relevant data.

Why this answer

The correct action is to store data using Hive-style partition paths (e.g., year=2023/month=01/day=01/) and run the AWS Glue crawler with default settings. The crawler will automatically detect the partitions and update the Data Catalog. Athena can then use partition pruning to optimize queries.

Other options either use unnecessary custom classifiers, manual partition management, or irrelevant settings.

Exam trap

The trap here is assuming that custom classifiers or manual partition creation are needed for standard Hive-style partitions, when the crawler handles them automatically.

1212
MCQeasy

A data engineer needs to move data from an Amazon S3 bucket to an Amazon Redshift cluster on a daily schedule. The data is in CSV format and the target table already exists. Which AWS service should the engineer use to automate this task?

A.AWS Glue
B.Amazon Athena
C.Amazon EMR
D.Amazon Kinesis Data Analytics
AnswerA

AWS Glue provides managed, serverless ETL with built-in schedulers and crawlers, so the engineer can automate the daily S3-to-Redshift load without provisioning servers. It reads CSV directly and writes to the existing Redshift table via JDBC, satisfying the daily schedule and existing-target constraints.

Why this answer

AWS Glue is a fully managed extract, transform, and load (ETL) service that can schedule and run jobs to move data from S3 to Redshift. It provides built-in connectors for both S3 and Redshift, and you can define a crawler to infer the schema and a job to load the data into the existing Redshift table. Glue handles the scheduling, retries, and scaling, making it the ideal choice for this daily automated task.

Exam trap

DEA-C01 often tests the distinction between services that query data in place (Athena) versus those that move data (Glue). Candidates may pick Athena because it works with S3 and Redshift, but Athena does not load data into Redshift.

How to eliminate wrong answers

Option B is wrong because Amazon Athena is an interactive query service that analyzes data in S3 using SQL; it does not load data into Redshift. Option C is wrong because Amazon EMR is a big data platform for processing large datasets using frameworks like Hadoop and Spark; while it can be used for ETL, it requires more management and is not the simplest choice for a scheduled S3-to-Redshift load. Option D is wrong because Amazon Kinesis Data Analytics is for real-time stream processing, not for batch data movement from S3 to Redshift.

1213
MCQeasy

A data pipeline ingests streaming data from thousands of IoT devices into Kinesis Data Streams. The data must be transformed using a simple field mapping before being stored in S3. Which service should be used to perform the transformation with minimal operational overhead?

A.AWS Lambda function invoked by the Kinesis stream
B.AWS Glue ETL job
C.Kinesis Data Analytics
D.Kinesis Data Firehose with a Lambda transformation
AnswerD

Firehose performs the field mapping via an attached Lambda function and delivers to S3, so no clusters or application code need managing. This satisfies the minimal operational overhead constraint, unlike Kinesis Data Analytics or a custom consumer on EC2.

Why this answer

Kinesis Data Firehose can invoke a Lambda function to perform simple field mapping transformations before delivering data to S3, minimizing operational overhead. Option A is wrong because AWS Lambda invoked directly by the Kinesis stream requires custom logic for S3 delivery and stream management, increasing overhead. Option B is wrong because AWS Glue ETL jobs are designed for batch processing and are more complex to set up for streaming transformations.

Option C is wrong because Kinesis Data Analytics is used for real-time analytics with SQL or Flink, not simple field mapping transformations.

1214
MCQeasy

A company is using Amazon S3 to store critical data and needs to ensure that objects are automatically transitioned to S3 Glacier Deep Archive after 180 days to reduce costs. Which S3 lifecycle action should be configured?

A.Expiration
B.Transition
C.AbortIncompleteMultipartUpload
D.NoncurrentVersionTransition
AnswerB

The Transition lifecycle action moves objects to another storage class after a specified age, so configuring it with a 180-day threshold and the Glacier Deep Archive target automates the cost reduction. Expiration would delete objects instead of archiving them.

Why this answer

The S3 lifecycle 'Transition' action is specifically designed to move objects between storage classes after a specified number of days. To reduce costs by moving objects to S3 Glacier Deep Archive after 180 days, you configure a lifecycle rule with a Transition action that targets the 'DEEP_ARCHIVE' storage class at the 180-day mark.

Exam trap

The trap here is that candidates often confuse 'Expiration' (deletion) with 'Transition' (storage class change), or incorrectly apply 'NoncurrentVersionTransition' when the question does not mention versioning or noncurrent versions.

How to eliminate wrong answers

Option A is wrong because 'Expiration' is used to permanently delete objects after a set period, not to transition them to a different storage class. Option C is wrong because 'AbortIncompleteMultipartUpload' is used to clean up incomplete multipart uploads after a specified number of days, not to transition objects between storage classes. Option D is wrong because 'NoncurrentVersionTransition' applies only to noncurrent versions of versioned objects, not to current versions, and the question does not specify versioning or noncurrent versions.

1215
MCQeasy

A company needs to store archival logs that must be retained for 10 years. The logs are accessed infrequently, but when accessed, retrieval must occur within 12 hours. Which storage class is MOST cost-effective?

A.Amazon S3 Glacier Deep Archive
B.Amazon S3 Intelligent-Tiering
C.Amazon S3 Standard
D.Amazon S3 One Zone-Infrequent Access
AnswerA

S3 Glacier Deep Archive is the lowest-cost archival class, with a standard retrieval time within 12 hours, matching the stated access requirement. It suits 10-year retention of infrequently accessed logs where cost minimisation outweighs retrieval speed.

Why this answer

Amazon S3 Glacier Deep Archive is the most cost-effective storage class for archival logs that must be retained for 10 years with infrequent access and a 12-hour retrieval window. It offers the lowest storage cost among S3 classes, with retrieval times typically within 12 hours for standard retrievals, making it ideal for long-term archival data that is rarely accessed.

Exam trap

The trap here is that candidates often confuse retrieval time with cost, assuming that any class with faster retrieval is better, but the 12-hour retrieval window explicitly allows the use of the lowest-cost archival tier, making Glacier Deep Archive the correct choice despite its slower retrieval speed.

How to eliminate wrong answers

Option B (S3 Intelligent-Tiering) is wrong because it is designed for data with unknown or changing access patterns and automatically moves objects between tiers, but it incurs monitoring and automation fees that make it less cost-effective than Glacier Deep Archive for purely archival data with a known 10-year retention. Option C (S3 Standard) is wrong because it is optimized for frequently accessed data with millisecond retrieval times and has a much higher storage cost, making it prohibitively expensive for 10 years of archival logs. Option D (S3 One Zone-Infrequent Access) is wrong because it is intended for infrequently accessed data that can be recreated if lost, but it does not provide the durability or cost savings of Glacier Deep Archive for long-term archival, and its retrieval times are faster than needed, leading to unnecessary cost.

1216
MCQmedium

A company wants to ingest data from SaaS applications (e.g., Salesforce, Marketo) into Amazon S3 for analytics. The data volume is moderate and updates occur frequently. Which AWS service is BEST suited for this task?

A.Amazon Kinesis Data Streams
B.Amazon AppFlow
C.AWS Database Migration Service (DMS)
D.AWS Glue
AnswerB

Amazon AppFlow provides managed, bidirectional connectors to SaaS applications such as Salesforce and Marketo, writing directly into Amazon S3 without custom code. It supports both scheduled and event-driven flows, satisfying the stem's requirement for frequent updates at moderate volume, unlike AWS DataSync or Storage Gateway, which target file or on-premises transfer rather than SaaS APIs.

Why this answer

Amazon AppFlow is purpose-built for ingesting data from SaaS applications such as Salesforce, Marketo, Slack, and Google Analytics into AWS services like Amazon S3. It provides native connectors, supports scheduled or event-driven flows, and handles authentication and incremental transfers without custom code. For moderate-volume, frequently updated SaaS data landing in S3, AppFlow is the most direct and operationally simple fit.

Exam trap

The trap here is confusing 'streaming' with 'SaaS ingestion' — candidates see 'frequent updates' and jump to Kinesis, but Kinesis has no native SaaS connectors, so AppFlow is the correct managed integration service.

How to eliminate wrong answers

Option A is wrong because Kinesis Data Streams is a low-level streaming ingestion service for custom producers (SDKs, KPL, Kinesis Agent) and has no native SaaS connectors — you would have to build and maintain the Salesforce/Marketo integration yourself. Option C is wrong because AWS DMS is designed to migrate databases (homogeneous/heterogeneous) into AWS, not to pull object or record data from SaaS CRM/marketing platforms. Option D is wrong because AWS Glue is an ETL/catalog service that transforms and moves data between sources; it does not natively authenticate to and extract from SaaS apps like Salesforce or Marketo without custom connectors or AppFlow as the source.

1217
MCQeasy

A data pipeline ingests daily CSV files from an FTP server into an Amazon S3 bucket. The files must be converted to Parquet format and partitioned by date for efficient querying using Amazon Athena. Which AWS service is most suitable for this transformation?

A.Amazon Kinesis Data Firehose
B.Amazon EMR
C.AWS Glue
D.AWS Lambda
AnswerC

Glue provides a serverless Spark environment that can transform CSV to Parquet and partition data efficiently.

Why this answer

AWS Glue is the most suitable service because it provides a fully managed ETL (Extract, Transform, Load) capability that can natively read CSV files from S3, convert them to Parquet format, and write the output partitioned by date. Glue's built-in transform 'ConvertToParquet' and dynamic frame partitioning make this a straightforward, serverless solution without needing to manage infrastructure.

Exam trap

The trap here is that candidates often confuse AWS Glue with Amazon EMR, thinking EMR is always needed for Parquet conversion, but Glue's serverless ETL is more appropriate for scheduled batch jobs without cluster management overhead.

How to eliminate wrong answers

Option A is wrong because Amazon Kinesis Data Firehose is designed for streaming data ingestion, not batch processing of daily CSV files from an FTP server; it lacks native support for reading from S3 as a source and performing complex transformations like CSV-to-Parquet conversion with custom partitioning. Option B is wrong because Amazon EMR is a managed Hadoop cluster that can perform this transformation, but it requires provisioning and managing EC2 instances, which is overkill for a simple daily batch job and not the most suitable service for a serverless, cost-effective solution. Option D is wrong because AWS Lambda has a maximum execution time of 15 minutes and a limited memory capacity (up to 10 GB), which is insufficient for processing large daily CSV files (e.g., gigabytes in size) and performing efficient Parquet conversion with partitioning.

1218
Multi-Selectmedium

A data engineer is designing a pipeline to ingest data from an Amazon RDS for PostgreSQL database into Amazon S3 using AWS Database Migration Service (AWS DMS). The source database has a high volume of transactions and the engineer needs to capture ongoing changes with minimal impact on the source. The target S3 bucket must store the data in Parquet format for querying with Amazon Athena. Which two actions should the engineer take to meet these requirements? (Choose two.)

Select 2 answers
A.Enable multi-AZ on the replication instance and use a single large replication instance to handle all tables.
B.Use AWS DMS to replicate directly to Amazon S3 in Parquet format by selecting the Parquet output option in the target endpoint settings.
C.Set up a full load only and schedule it to run every hour to capture changes.
D.Configure AWS DMS to write to Amazon S3 in CSV format and then use AWS Glue to convert the data to Parquet.
E.Configure AWS DMS to use change data capture (CDC) with a replication instance that has sufficient resources and enable logical replication on the source PostgreSQL database.
AnswersB, E

AWS DMS supports Amazon S3 as a target and can write data in Parquet format when configured in the endpoint settings. This eliminates the need for a separate conversion step and allows Athena to query the data directly. The engineer must specify the appropriate serialization format and ensure the S3 bucket has the necessary permissions.

Why this answer

To capture ongoing changes from PostgreSQL with minimal impact, AWS DMS must use CDC, which requires logical replication on the source and a replication instance sized appropriately. To store data in Parquet for Athena, the DMS target endpoint for S3 must be configured to output Parquet. These two actions together satisfy the requirements: near-real-time replication with low source overhead and a columnar format optimized for query performance.

Exam trap

The trap here is assuming that a periodic full load is sufficient for ongoing changes, when it actually causes high source load and misses intermediate changes; and overlooking that DMS can natively output Parquet to S3.

1219
MCQmedium

A data engineer manages an Amazon S3 data lake that holds sensitive customer transaction logs. Compliance requires that all objects be encrypted at rest with keys that the company rotates every 90 days and fully controls, including the ability to immediately revoke access and audit key usage separately from other AWS accounts. The engineer must choose an encryption method that meets these requirements with minimal operational overhead. Which solution should the engineer implement?

A.Use server-side encryption with Amazon S3 managed keys (SSE-S3)
B.Use client-side encryption with a customer-provided key stored in AWS Secrets Manager
C.Use server-side encryption with AWS KMS customer managed keys (SSE-KMS)
D.Use server-side encryption with customer-provided keys (SSE-C)
AnswerC

SSE-KMS with customer managed keys gives the organization full control over the key, allows a custom rotation period (including 90 days), supports immediate revocation via key policy changes, and logs every key use in AWS CloudTrail. This directly satisfies the compliance requirements with minimal operational overhead because S3 handles encryption transparently.

Why this answer

The requirement for customer-controlled keys, custom 90-day rotation, immediate revocation, and separate key usage audit points directly to AWS KMS customer managed keys with SSE-KMS. S3 manages the encryption process, so operational overhead remains low, while the key policy and CloudTrail integration provide the necessary control and visibility. Other encryption options either lack customer control or shift too much operational responsibility to the application.

Exam trap

The trap here is assuming that any server-side encryption option provides the same level of key control and auditability, when only SSE-KMS with customer managed keys meets the specific rotation and revocation requirements.

1220
MCQeasy

A data engineer needs to monitor the number of records processed by an AWS Glue ETL job and send an alert if the count drops below a threshold. Which AWS service should be used to create this custom metric?

A.Amazon S3
B.AWS Config
C.Amazon CloudWatch
D.AWS CloudTrail
AnswerC

Amazon CloudWatch accepts custom metrics published from the Glue job and evaluates them against metric alarms, triggering Amazon SNS notifications when the record count falls below the threshold. This provides the monitoring and alerting the scenario requires.

Why this answer

Amazon CloudWatch is the correct service for creating custom metrics because it allows you to publish your own data points, such as the number of records processed by an AWS Glue ETL job. You can use the CloudWatch PutMetricData API or the AWS Glue job script to emit a custom metric, then set an alarm on that metric to trigger an alert when the count drops below a threshold.

Exam trap

The trap here is that candidates often confuse AWS CloudTrail with CloudWatch because both are monitoring-related, but CloudTrail is for auditing API calls, not for ingesting custom numerical metrics or setting alarms on them.

How to eliminate wrong answers

Option A is wrong because Amazon S3 is an object storage service and does not provide a mechanism to create or monitor custom metrics; it only stores data and logs access via server access logs or AWS CloudTrail. Option B is wrong because AWS Config is a service for evaluating and auditing resource configurations against rules, not for ingesting or alerting on custom operational metrics like record counts. Option D is wrong because AWS CloudTrail records API activity for auditing and governance, but it cannot be used to create custom metrics or set threshold-based alarms; it captures events, not numerical data points.

1221
Multi-Selecthard

A data engineer is using Amazon Athena to query data stored in an S3 bucket. The queries are running slowly. Which THREE actions can improve query performance?

Select 3 answers
A.Partition the data on commonly filtered columns.
B.Convert the data to JSON format for better schema evolution.
C.Move the data to S3 Standard-IA storage class.
D.Convert the data to a columnar format such as Parquet or ORC.
E.Use compression (e.g., Snappy, Gzip) on the data files.
AnswersA, D, E

Partition pruning reduces amount of data scanned.

Why this answer

Partitioning data on commonly filtered columns (Option A) improves Athena query performance by reducing the amount of data scanned. Athena uses Hive-style partitioning (e.g., `s3://bucket/table/year=2023/month=01/`), and when a query includes a filter on the partition column, Athena prunes partitions and only reads the relevant S3 prefixes. This directly reduces I/O and query cost, as Athena charges per TB of data scanned.

Exam trap

The trap here is that candidates may think S3 storage class (Standard-IA) affects query performance, but Athena's performance is independent of storage class; the key levers are data format, partitioning, and compression.

1222
MCQeasy

A data engineer needs to monitor the number of records processed by a Kinesis Data Firehose delivery stream and set an alarm if the count drops below a threshold. Which CloudWatch metric should be used?

A.IncomingRecords
B.PutRecord.Success
C.DeliveryToS3.Success
D.IncomingBytes
AnswerA

IncomingRecords counts records successfully put into the Firehose delivery stream, directly satisfying the need to monitor processed record volume and alarm on a low threshold. Other metrics track delivery success, latency, or byte size rather than record count, so they cannot detect a drop in incoming records.

Why this answer

IncomingRecords is the CloudWatch metric that specifically counts the number of records successfully ingested into a Kinesis Data Firehose delivery stream. It is emitted per delivery stream and directly reflects the volume of data being processed, making it the correct choice for monitoring record count and setting threshold alarms. Other metrics either measure API call success, delivery outcomes, or byte volume, not record count.

Exam trap

DEA-C01 often tests the distinction between ingestion metrics (like IncomingRecords) and delivery metrics (like DeliveryToS3.Success), causing candidates to confuse record count with byte count or API success.

How to eliminate wrong answers

Option B is wrong because PutRecord.Success is a metric for the Kinesis Data Streams PutRecord API, not for Firehose; Firehose uses PutRecord and PutRecordBatch APIs but the success metric is not the primary record count metric. Option C is wrong because DeliveryToS3.Success tracks the success of delivering data to Amazon S3, not the number of records ingested; it is a delivery metric, not an ingestion count. Option D is wrong because IncomingBytes measures the volume of data in bytes, not the number of records, so it cannot be used to alarm on record count.

1223
MCQeasy

A data engineer is designing a data lake on Amazon S3. Which feature should be used to manage the lifecycle of objects and move them to cheaper storage classes automatically?

A.S3 Lifecycle policies
B.S3 Object Lock
C.S3 Storage Class Analysis
D.S3 Inventory
AnswerA

S3 Lifecycle policies define rules that transition objects to cheaper storage classes such as Standard-IA, Intelligent-Tiering or Glacier, and expire them, based on age or prefixes. This automates cost optimisation across the data lake without manual intervention.

Why this answer

S3 Lifecycle policies are the correct feature for automatically managing object lifecycles and transitioning objects to cheaper storage classes (e.g., from S3 Standard to S3 Glacier Deep Archive) based on age or other rules. This directly meets the requirement to move objects to cost-optimized storage without manual intervention.

Exam trap

The trap here is confusing S3 Storage Class Analysis (which only recommends transitions) with S3 Lifecycle policies (which actually execute them), leading candidates to pick Option C thinking it automates the move.

How to eliminate wrong answers

Option B is wrong because S3 Object Lock is designed to prevent object deletion or overwrites for compliance or retention purposes, not to automate storage class transitions. Option C is wrong because S3 Storage Class Analysis provides recommendations and visibility into access patterns to help decide when to transition objects, but it does not automatically move objects—it only generates reports. Option D is wrong because S3 Inventory provides a flat-file list of objects and their metadata for auditing or sync, but it has no capability to trigger lifecycle actions.

1224
MCQmedium

Refer to the exhibit. A data engineer needs to connect to the Redshift cluster from an EC2 instance in the same VPC. The engineer can ping the EC2 instance but cannot connect to Redshift using the endpoint address and port 5439. What is the most likely cause?

A.The security group for the Redshift cluster does not allow inbound traffic on port 5439 from the EC2 instance.
B.The Redshift cluster is in a different VPC.
C.The Redshift cluster is not in an available state.
D.The Redshift cluster is publicly accessible and requires an internet gateway.
AnswerA

Reachability via ping only proves ICMP is permitted; Redshift listens on TCP 5439, so the cluster's security group must allow inbound 5439 from the EC2 instance's security group or private IP. Without that rule, the endpoint connection times out.

Why this answer

The most likely cause is that the security group associated with the Redshift cluster does not have an inbound rule allowing TCP traffic on port 5439 from the security group or IP address of the EC2 instance. Since the engineer can ping the EC2 instance (ICMP works), but cannot connect to Redshift on port 5439, this points to a firewall or security group rule blocking the specific port, not a network reachability issue.

Exam trap

AWS often tests the distinction between ICMP reachability (ping) and TCP port-level connectivity, leading candidates to overlook security group rules when they see successful ping results.

How to eliminate wrong answers

Option B is wrong because if the Redshift cluster were in a different VPC, the engineer would not be able to ping the EC2 instance from the same VPC context, and VPC peering or transit gateway would be required; the question states they are in the same VPC. Option C is wrong because if the cluster were not in an available state, the engineer would likely receive a different error (e.g., 'cluster not found' or connection timeout), and the question does not indicate any cluster status issues. Option D is wrong because the Redshift cluster is in the same VPC as the EC2 instance, so public accessibility and an internet gateway are not required; traffic stays within the VPC and uses private IPs.

1225
Multi-Selecthard

A data engineer is building a pipeline to ingest data from an on-premises Oracle database into Amazon S3. The pipeline must capture change data (CDC) in near real-time and handle schema changes. Which TWO AWS services should the engineer use?

Select 2 answers
A.AWS Glue Schema Registry
B.AWS Snowball Edge
C.Amazon AppFlow
D.Amazon Kinesis Data Streams with Kinesis Agent
E.AWS Database Migration Service (DMS) with CDC
AnswersA, E

Manages schema evolution for streaming data.

Why this answer

AWS Glue Schema Registry (A) is correct because it enables schema discovery, validation, and evolution for streaming data, allowing the pipeline to handle schema changes from the Oracle CDC source. It integrates with Apache Kafka and Amazon Kinesis Data Streams to enforce schema compatibility rules (e.g., backward, forward, full) as data arrives, ensuring downstream consumers can adapt to evolving schemas without breaking.

Exam trap

The DEA-C01 exam often tests the misconception that Amazon Kinesis Data Streams alone can perform CDC from a database, but Kinesis requires a separate agent or connector (like Debezium or DMS) to read database logs, making DMS the correct CDC service.

1226
MCQeasy

A company stores raw event files in an Amazon S3 bucket that receives thousands of small objects per hour. An AWS Glue job reads the prefix and writes a compacted Parquet dataset to a curated bucket. Operations reports that the Glue job's runtime keeps growing even though the hourly data volume is constant. Which change is MOST likely to reduce runtime?

A.Increase the Glue job's timeout value so long-running reads are not terminated before completion.
B.Enable S3 Transfer Acceleration on the raw bucket so Glue can download objects faster.
C.Add a Glue job bookmark to the raw prefix and schedule the compaction job to run hourly instead of daily.
D.Run an S3 compaction or grouping step that merges small objects into larger files before the Glue job reads them.
AnswerD

Each small object incurs per-file overhead for listing, opening, and task scheduling, so thousands of tiny files dominate runtime regardless of total bytes. Merging them into larger objects, for example with an S3 Batch Operations copy or a preceding compaction job, cuts the number of files Glue must open and lets Spark read efficiently, which directly reduces runtime.

Why this answer

Glue and Spark incur fixed overhead per input file for listing, opening, and scheduling tasks. Thousands of small objects make that overhead dominate total runtime even when the byte volume is unchanged. Consolidating small files into larger objects before the ETL read reduces the number of files processed and shortens the job, which is the standard remedy for the small-file problem.

Exam trap

The trap here is assuming runtime scales only with data volume, when per-file overhead from many small objects is usually the dominant cost.

1227
MCQmedium

A data engineer is using AWS Glue to transform data from Amazon S3. The source data is in CSV format with inconsistent date formats across files (e.g., 'MM/DD/YYYY' and 'YYYY-MM-DD'). The engineer needs to standardize all dates to 'YYYY-MM-DD' format in the output. Which AWS Glue transform should the engineer use to achieve this?

A.ApplyMapping
B.Filter
C.ResolveChoice
D.Map
AnswerD

The `Map` transform allows applying a custom function to each record in a DynamicFrame. The engineer can write a Python function that attempts to parse the date string using multiple formats (e.g., using `datetime.strptime` with different format strings) and then outputs the standardized 'YYYY-MM-DD' string. This provides the flexibility needed to handle inconsistent date formats across files.

Why this answer

The `Map` transform in AWS Glue enables custom record-level transformations. By writing a Python function that parses various date formats and outputs a standardized string, the engineer can handle inconsistent date formats. This approach is flexible and can be applied across all records, ensuring uniform output.

Exam trap

The trap here is thinking that `ApplyMapping` can handle date format conversion, when it only handles type casting and renaming.

1228
MCQhard

A data engineer is designing a solution to securely store and rotate database credentials used by an application. The credentials should be automatically rotated every 90 days. Which AWS service should be used?

A.AWS Secrets Manager
B.AWS Systems Manager Parameter Store
C.AWS Key Management Service (KMS)
D.AWS Identity and Access Management (IAM)
AnswerA

AWS Secrets Manager natively stores database credentials and performs scheduled rotation via Lambda functions, satisfying the 90-day automatic rotation constraint. Unlike Parameter Store, which lacks built-in rotation, it directly manages credential lifecycle for RDS and other databases, meeting the security requirement without custom orchestration.

Why this answer

AWS Secrets Manager is designed specifically for storing, managing, and automatically rotating secrets such as database credentials. It natively supports rotation schedules (e.g., every 90 days) using Lambda rotation functions, and integrates with RDS, Redshift, and DocumentDB for managed rotation. This directly meets the requirement for automatic rotation every 90 days.

Exam trap

DEA-C01 often tests the confusion between Secrets Manager and Parameter Store, where candidates pick Parameter Store for secret rotation even though it lacks built-in automatic rotation.

How to eliminate wrong answers

Option B is wrong because AWS Systems Manager Parameter Store can store secrets securely (as SecureString), but it does not provide built-in automatic rotation; you would need to implement custom rotation logic. Option C is wrong because AWS KMS is a key management service for encryption keys, not for storing and rotating database credentials; it can encrypt secrets but does not manage their lifecycle. Option D is wrong because IAM is for identity and access management, not for storing or rotating secrets; it manages permissions, not credentials storage.

1229
MCQmedium

A data engineer is troubleshooting an AWS Glue job that writes data to an S3 bucket. The IAM role attached to the Glue job has the policy shown in the exhibit. The job fails when writing to the 'secrets/' prefix but succeeds when writing to other prefixes. What is the reason for the failure?

A.The job does not have permission to write to the bucket at all.
B.The resource ARN in the Allow statement does not include the bucket itself.
C.The Deny statement is not effective because it is placed after the Allow.
D.The Deny statement explicitly denies PutObject to the secrets/ prefix.
AnswerD

The bucket policy's explicit Deny for `s3:PutObject` on the `secrets/` prefix overrides the role's Allow, because an explicit Deny in IAM always wins. This satisfies the stem's constraint that writes to other prefixes succeed while `secrets/` fails, confirming the Deny is scoped precisely to that prefix.

Why this answer

The IAM policy attached to the Glue job contains an explicit Deny statement for the 'secrets/' prefix, which overrides any Allow. In AWS IAM, an explicit Deny always takes precedence over Allow. Therefore, the job fails when writing to 'secrets/' but succeeds for other prefixes where no Deny exists.

Exam trap

DEA-C01 often tests the IAM policy evaluation logic, and candidates frequently forget that an explicit Deny always overrides an Allow, regardless of statement order.

How to eliminate wrong answers

Option A is wrong because the job succeeds when writing to other prefixes, so it clearly has write permission to the bucket. Option B is wrong because the resource ARN in the Allow statement typically includes the bucket and objects; if it did not include the bucket itself, the job would fail for all prefixes, not just 'secrets/'. Option C is wrong because the order of statements in an IAM policy does not matter; explicit Deny always wins regardless of placement.

1230
Multi-Selecteasy

A data engineer is monitoring an Amazon RDS for PostgreSQL instance. The engineer wants to set up alerts for high CPU utilization and low free storage space. Which AWS services can be used together to achieve this? (Choose TWO.)

Select 2 answers
A.Amazon Simple Notification Service (SNS)
B.Amazon CloudWatch
C.AWS CloudTrail
D.AWS Config
E.Amazon Route 53
AnswersA, B

Amazon SNS delivers the notifications that CloudWatch alarms trigger when CPU utilisation or free storage crosses a defined threshold on the RDS for PostgreSQL instance. It satisfies the alerting requirement by fanning out those alarm state changes to email, SMS or HTTP subscribers, completing the monitoring pipeline alongside CloudWatch.

Why this answer

Amazon CloudWatch [CORRECT] is the AWS monitoring service that collects RDS for PostgreSQL metrics such as CPUUtilization and FreeStorageSpace, and it lets you create alarms that trigger when those metrics cross defined thresholds. Amazon Simple Notification Service (SNS) [CORRECT] is the notification service that CloudWatch alarms publish to, delivering email, SMS, or other messages to the engineer when high CPU or low free storage is detected. Together, CloudWatch alarms plus an SNS topic form the standard alerting pipeline for RDS metric thresholds.

AWS CloudTrail records API activity and audit events, not performance metrics, so it cannot detect high CPU or low storage. AWS Config evaluates resource configuration compliance and does not monitor runtime utilization metrics. Amazon Route 53 is a DNS and traffic-routing service and has no role in RDS metric alerting.

Exam trap

The trap here is that candidates often confuse AWS CloudTrail (audit logging) with CloudWatch (monitoring), or think AWS Config can monitor performance metrics instead of just configuration compliance.

1231
MCQeasy

A data engineer is using AWS Glue to transform data stored in Amazon S3. The security team requires that data in transit between AWS Glue and Amazon S3 be encrypted. The engineer wants to ensure that all connections use TLS. Which action should the engineer take to enforce encryption in transit for AWS Glue jobs accessing S3?

A.Enable SSL/TLS for the Glue connection and set the require_ssl parameter to true in the JDBC URL.
B.Attach a bucket policy to the S3 bucket that denies requests where aws:SecureTransport is false.
C.Configure the Glue job to use a VPC endpoint for S3 and enable encryption on the endpoint.
D.Set the Glue job parameter --encryption-mode to SSE-S3.
AnswerB

To enforce encryption in transit for S3, a bucket policy can deny any requests that do not use TLS, by checking the aws:SecureTransport condition key. This ensures that all access, including from AWS Glue, must use HTTPS. This is the standard method to require encryption in transit for S3.

Why this answer

Enforcing encryption in transit for S3 is done by adding a bucket policy that denies requests when aws:SecureTransport is false. This ensures that all clients, including AWS Glue, must use HTTPS/TLS when accessing the bucket. Other options either address encryption at rest or do not enforce TLS.

Exam trap

The trap here is confusing encryption at rest settings like SSE-S3 or VPC endpoints with encryption in transit, which requires TLS enforcement via bucket policy.

1232
MCQmedium

A data engineer is using Amazon Kinesis Data Streams to ingest clickstream data. The stream has 10 shards and each record is 50 KB. The engineer notices that the PutRecords API is frequently returning ProvisionedThroughputExceededException errors, even though the total incoming data rate is below the stream's overall capacity. What is the MOST likely cause?

A.The partition key is not uniformly distributed across shards, causing hot shards.
B.The stream's retention period is set too low, causing data loss.
C.The PutRecords API call is using too many records per request, exceeding the 500-record limit.
D.The stream's shards are not encrypted, causing throttling.
AnswerA

Kinesis Data Streams distributes records to shards based on the partition key hash. If the partition key is skewed, some shards receive more traffic than others, exceeding their individual limits (1 MB/s or 1000 records/s per shard) even if total stream capacity is sufficient. This leads to ProvisionedThroughputExceededException on the hot shards. Ensuring a uniform partition key distribution resolves the issue.

Why this answer

ProvisionedThroughputExceededException occurs when a specific shard exceeds its 1 MB/s or 1000 records/s limit. If total incoming rate is below stream capacity but partition keys are not uniformly distributed, some shards become hot and throttle. Using a well-distributed partition key or resharding can mitigate this.

Exam trap

The trap here is assuming that overall stream capacity is the only factor, while ignoring per-shard limits and partition key distribution.

1233
MCQhard

A data engineer is designing a real-time analytics pipeline that ingests clickstream data into Amazon Kinesis Data Streams. The data must be stored in Amazon S3 for later analysis with Amazon Athena. The engineer needs the data to be queryable with minimal latency and wants to avoid managing complex ETL jobs. Which solution should the engineer use?

A.Use Amazon Kinesis Data Analytics to run SQL queries on the stream and write results to Amazon S3.
B.Use Amazon Kinesis Data Firehose to deliver data to Amazon S3 with record format conversion to Parquet using an AWS Glue table.
C.Use AWS Lambda to read from Kinesis Data Streams and write JSON files to Amazon S3.
D.Use AWS Glue streaming ETL jobs to read from Kinesis Data Streams and write Parquet to Amazon S3.
AnswerB

Kinesis Data Firehose can deliver streaming data to Amazon S3 and perform record format conversion to Parquet using a schema from the AWS Glue Data Catalog. This eliminates the need for custom ETL jobs and makes the data immediately queryable by Athena with columnar performance. It provides near-real-time delivery and automatic partitioning, meeting the latency and simplicity requirements.

Why this answer

Amazon Kinesis Data Firehose can ingest streaming data and deliver it to Amazon S3 with automatic record format conversion to Parquet using an AWS Glue table. This provides near-real-time, queryable data in a columnar format without custom ETL code. Other options require writing and managing Lambda functions, Kinesis Data Analytics applications, or Glue streaming jobs, which add complexity and do not offer the same level of managed simplicity.

Exam trap

The trap here is overlooking that Kinesis Data Firehose can perform schema-based record format conversion to Parquet, which is often assumed to require a separate ETL tool.

1234
MCQmedium

A data engineer is running a Spark job on Amazon EMR. The job reads from S3, processes data, and writes to S3. The job is taking longer than expected. The engineer notices that the job is spending a lot of time in the 'GC' (garbage collection) phase. Which configuration change is most likely to improve performance?

A.Increase the spark.executor.memory setting.
B.Increase the spark.sql.shuffle.partitions.
C.Decrease the number of executor cores.
D.Decrease the spark.executor.memoryOverhead.
AnswerA

Excessive garbage collection means each executor's heap is too small for the working set, so the JVM constantly reclaims objects. Raising spark.executor.memory enlarges the heap, reducing GC frequency and duration, which directly addresses the bottleneck the engineer observed in the Spark job.

Why this answer

Garbage collection overhead in Spark is directly tied to the size of the JVM heap on each executor. When executors have too little heap, the JVM runs GC far more frequently and for longer pauses, stalling task execution. Increasing spark.executor.memory enlarges the heap, reducing GC frequency and duration, which is the standard remedy when the Spark UI shows high GC time relative to task time.

Exam trap

The trap here is confusing shuffle/parallelism tuning knobs (shuffle.partitions, cores) with memory tuning; candidates who see 'slow job' reflexively reach for partition counts instead of recognizing GC as a heap-size symptom.

How to eliminate wrong answers

Option B is wrong because spark.sql.shuffle.partitions controls the number of partitions after a shuffle; increasing it changes parallelism and task granularity but does nothing to reduce GC pressure on executors. Option C is wrong because decreasing executor cores reduces parallelism and can actually worsen GC behavior by concentrating more data per core, and it does not address heap size. Option D is wrong because spark.executor.memoryOverhead governs off-heap memory for things like Python workers and native allocations; decreasing it risks OOM kills and has no positive effect on JVM garbage collection.

1235
MCQeasy

A company has CSV files in an S3 bucket that need to be converted to Parquet and loaded into a Redshift table daily. The transformation is a simple schema mapping without joins. Which AWS Glue feature is BEST suited for this task?

A.AWS Glue ETL job
B.AWS Glue DataBrew
C.AWS Glue Workflow
D.AWS Glue Crawler
AnswerA

AWS Glue ETL jobs handle schema mapping and format conversion natively, reading CSV from S3 and writing Parquet to Redshift through the Glue Data Catalog and JDBC/Redshift connections. No joins are required, so a straightforward ETL job satisfies the daily conversion and load requirement.

Why this answer

AWS Glue ETL jobs are purpose-built for extracting data from sources like S3, applying transformations (including schema mapping and format conversion), and loading the results into targets like Redshift. For a simple CSV-to-Parquet conversion with schema mapping and no joins, a Glue ETL job using the built-in transforms or a Spark script is the correct tool. It handles the conversion and load in a single managed job.

Exam trap

DEA-C01 often tests Glue service selection by describing a transformation task — candidates confuse Glue Workflow (orchestration) or Glue Crawler (cataloging) with the actual ETL engine, picking a service that schedules or catalogs rather than transforms.

How to eliminate wrong answers

Option B is wrong because AWS Glue DataBrew is a visual data preparation tool aimed at analysts for cleaning and normalizing data interactively, not for scheduled ETL pipelines that load into Redshift. Option C is wrong because AWS Glue Workflow is an orchestration service that chains crawlers and jobs together — it schedules and coordinates but does not itself perform the transformation. Option D is wrong because a Glue Crawler only discovers schema and populates the Data Catalog; it does not transform or load data into Redshift.

1236
MCQeasy

A data engineer is building a data lake on Amazon S3. The engineer needs to store structured data that will be queried by Amazon Athena. The data is currently in CSV format and is partitioned by date. The engineer wants to improve query performance and reduce the amount of data scanned. Which action should the engineer take?

A.Use Amazon Redshift Spectrum to query the CSV data directly from S3.
B.Compress the CSV files using gzip and update the table metadata in the AWS Glue Data Catalog.
C.Convert the data to Apache Parquet format and use AWS Glue to update the table metadata in the AWS Glue Data Catalog.
D.Increase the number of partitions by adding a partition for each hour of the day.
AnswerC

Parquet is a columnar format that enables Athena to read only the columns needed, reducing data scanned and improving performance. Updating the AWS Glue Data Catalog ensures Athena can query the new format. This is a best practice for optimizing Athena queries on S3 data lakes.

Why this answer

Converting data to a columnar format like Parquet significantly reduces the amount of data scanned by Athena because only the required columns are read. Updating the AWS Glue Data Catalog ensures the table metadata reflects the new format. This combination improves query performance and lowers cost.

Exam trap

The trap here is assuming that compression alone will provide the same performance benefits as converting to a columnar format like Parquet.

1237
MCQmedium

A company uses AWS Glue to process data stored in Amazon S3. The security team mandates that all data in transit between AWS Glue and Amazon S3 must be encrypted with TLS. The Glue job connects to S3 using the AWS SDK. Which configuration should the data engineer implement to enforce TLS encryption for the Glue job's S3 connections?

A.Configure the Glue job to use a VPC endpoint for S3 and enable AWS PrivateLink.
B.Attach an S3 bucket policy that denies requests where aws:SecureTransport is false.
C.Set the Glue job parameter --encryption-mode to TLS.
D.Enable default encryption on the S3 bucket with SSE-KMS.
AnswerB

An S3 bucket policy with a condition that denies access when aws:SecureTransport is false enforces that all requests to the bucket use TLS. This policy applies to any client, including AWS Glue, ensuring data in transit is encrypted. This is the standard AWS method to enforce TLS for S3 access.

Why this answer

To enforce TLS for all connections to an S3 bucket, including from AWS Glue, an S3 bucket policy that denies requests when aws:SecureTransport is false is the correct approach. This condition evaluates the transport protocol and blocks non-TLS requests. It is a best practice recommended by AWS for ensuring data in transit encryption.

Exam trap

The trap here is confusing encryption at rest with encryption in transit, or assuming Glue has a built-in TLS enforcement parameter.

1238
MCQmedium

A company uses AWS Glue to process data in Amazon S3. The Glue job fails with an error indicating that the partition keys in the catalog do not match the actual S3 partition structure. What is the most likely cause?

A.The IAM role does not have permissions to read the S3 data
B.The data files are encrypted with SSE-KMS
C.The table name in the catalog is different from the one used in the job
D.The Glue Data Catalog partition metadata is outdated after the S3 structure changed
AnswerD

Glue compares the partition keys registered in the Data Catalog against the actual S3 prefix layout. When the S3 structure changes without a corresponding crawler run or catalog update, the stored metadata no longer matches, so the job fails on the mismatch.

Why this answer

The Glue Data Catalog stores partition metadata separately from the actual S3 partition layout. When the S3 partition structure changes (e.g., new partitions are added or existing ones are renamed) without updating the catalog, the Glue job reads stale partition metadata, leading to a mismatch error. The job fails because it expects partitions based on the catalog, not the live S3 structure.

Exam trap

The trap here is that candidates confuse a partition metadata mismatch with other common Glue errors like IAM permissions or encryption issues, but the error message explicitly references partition keys, not access or decryption problems.

How to eliminate wrong answers

Option A is wrong because an IAM permissions issue would typically cause an Access Denied error, not a partition key mismatch error. Option B is wrong because SSE-KMS encryption affects data decryption, not the structure or metadata of partitions in the catalog. Option C is wrong because a table name mismatch would cause a 'table not found' error, not a partition key mismatch; the error specifically points to partition keys, not table identifiers.

1239
MCQeasy

A data engineer needs to store semi-structured JSON log files from multiple sources and query them using SQL. The data is rarely updated and access frequency is low. Which storage solution is MOST cost-effective?

A.Amazon Redshift with JSON ingestion and compression.
B.Amazon DynamoDB with JSON documents.
C.Amazon S3 with Amazon Athena for querying.
D.Amazon RDS for PostgreSQL with JSONB columns.
AnswerC

Amazon S3 provides low-cost durable storage for semi-structured JSON, and Athena queries it in place using SQL without loading or provisioning servers. For rarely updated, infrequently accessed logs, this serverless combination avoids the ongoing cost of a data warehouse or database.

Why this answer

Amazon S3 with Athena is the most cost-effective solution because the data is semi-structured JSON, rarely updated, and accessed infrequently. S3 provides low-cost storage for static data, and Athena uses a serverless, pay-per-query model, eliminating the need for a running cluster or provisioned capacity. This combination avoids the fixed costs of Redshift, DynamoDB, or RDS, making it ideal for low-frequency SQL querying of archival logs.

Exam trap

The trap here is that candidates often choose Redshift or RDS because they associate SQL querying with traditional databases, overlooking that Athena's serverless, pay-per-query model is far more cost-effective for infrequent access to static data stored in S3.

How to eliminate wrong answers

Option A is wrong because Amazon Redshift requires a provisioned cluster with ongoing compute costs, making it overkill and expensive for rarely accessed data; its JSON ingestion and compression do not offset the fixed infrastructure cost. Option B is wrong because Amazon DynamoDB is a NoSQL key-value store optimized for high-frequency, low-latency reads/writes, not for SQL-based ad-hoc querying of large JSON logs; its on-demand capacity mode still incurs per-request charges that are wasteful for infrequent access. Option D is wrong because Amazon RDS for PostgreSQL with JSONB columns requires a provisioned database instance with continuous compute and storage costs, and while JSONB supports indexing, it is not cost-effective for rarely queried, static log data compared to S3's pay-per-byte storage and Athena's pay-per-query model.

1240
MCQeasy

A data engineer is designing a data lake on Amazon S3. The data includes sensitive personally identifiable information (PII). Which combination of services would provide the most comprehensive data protection?

A.Use S3 Transfer Acceleration and enable versioning
B.Enable S3 server-side encryption with AWS KMS
C.Use Amazon CloudWatch Logs to monitor access and enable MFA Delete
D.Enable S3 Block Public Access and use Amazon Macie to discover and classify PII
AnswerD

S3 Block Public Access prevents accidental public exposure of the data lake buckets, while Amazon Macie uses machine learning to discover and classify PII stored in S3. Together they satisfy the stem's requirement for comprehensive protection of sensitive PII.

Why this answer

S3 Block Public Access prevents accidental public exposure of buckets and objects, which is the leading cause of S3 data leaks, while Amazon Macie uses machine learning and pattern matching to discover, classify, and alert on PII stored in S3. Together they address both the access-control and data-discovery dimensions of protecting sensitive PII, making this the most comprehensive combination.

Exam trap

DEA-C01 often tests the misconception that encryption alone equals comprehensive data protection — candidates pick SSE-KMS because it sounds strongest, but the exam expects you to combine preventive access controls (Block Public Access) with detective classification (Macie) for PII.

How to eliminate wrong answers

Option A is wrong because Transfer Acceleration only speeds up uploads over long distances and versioning only protects against overwrites/deletes — neither addresses PII exposure or classification. Option B is wrong because SSE-KMS encrypts data at rest but does nothing to prevent misconfigured public access or to discover where PII actually lives; encryption alone is not comprehensive protection. Option C is wrong because CloudWatch Logs monitors API activity and MFA Delete protects against accidental/malicious deletion, but neither discovers or classifies PII, and MFA Delete is a deletion safeguard, not a data-protection control for confidentiality.

1241
MCQhard

A data engineer is using AWS Step Functions to orchestrate a series of AWS Glue jobs. The engineer notices that one of the Glue jobs occasionally fails due to a transient network error. The engineer wants to implement a retry mechanism that retries the failed job up to three times with an exponential backoff. Which Step Functions state configuration should the engineer use?

A.Configure the Glue job itself to retry on failure by setting the 'MaxRetries' parameter in the job definition.
B.Add a 'Retry' field with 'MaxAttempts': 3, 'IntervalSeconds': 2, and 'BackoffRate': 2.0 to the task state.
C.Use a 'Parallel' state to run the Glue job multiple times simultaneously and take the first successful result.
D.Add a 'Catch' field with 'ErrorEquals': ['States.TaskFailed'] and a 'Next' state that increments a counter and loops back.
AnswerB

'Retry' with 'MaxAttempts': 3, 'IntervalSeconds': 2, and 'BackoffRate': 2.0 implements exponential backoff retries. The task will retry up to three times, with delays of 2, 4, and 8 seconds. This matches the requirement for exponential backoff and is the correct configuration for Step Functions.

Why this answer

The 'Retry' field in Step Functions is designed to handle transient errors by retrying the task with optional exponential backoff. Specifying 'MaxAttempts': 3, 'IntervalSeconds': 2, and 'BackoffRate': 2.0 achieves the desired retry behavior. Other options either require complex workarounds or do not support exponential backoff, making the 'Retry' field the correct choice.

Exam trap

The trap here is assuming that AWS Glue's built-in retry mechanism supports exponential backoff, when it actually uses a fixed interval.

1242
MCQhard

A company is using AWS DMS to replicate data from an on-premises Oracle database to Amazon RDS for MySQL. The replication is working, but the target table has a different schema. Which DMS feature should be used to transform the source schema to match the target?

A.Use AWS Schema Conversion Tool (SCT)
B.Use AWS Glue ETL jobs
C.Use DMS transformation rules
D.Use AWS Lambda triggers
AnswerC

DMS transformation rules apply schema and table mappings during replication, letting you rename schemas, tables and columns, change data types, and filter rows so the Oracle source matches the RDS for MySQL target schema. This directly satisfies the stem's requirement to transform a differing schema without altering the source database.

Why this answer

AWS DMS transformation rules allow you to modify the schema, table, or column names and data types during the migration process. This feature is specifically designed to handle schema transformations within the DMS task itself, enabling you to map the source Oracle schema to the target MySQL schema without external tools or services.

Exam trap

The trap here is that candidates often confuse AWS Schema Conversion Tool (SCT) with DMS transformation rules, assuming SCT handles runtime schema mapping, whereas SCT is a separate pre-migration assessment and conversion tool, not a DMS feature for ongoing replication transformations.

How to eliminate wrong answers

Option A is wrong because AWS Schema Conversion Tool (SCT) is used for heterogeneous database migrations to convert the entire database schema and code objects, but it is not a feature of DMS for runtime schema transformation during ongoing replication. Option B is wrong because AWS Glue ETL jobs are for batch data processing and transformation in a data lake or warehouse, not for real-time schema mapping within a DMS replication task. Option D is wrong because AWS Lambda triggers can be used for custom post-processing or validation, but they are not a built-in DMS feature for transforming source schemas to match target schemas during replication.

1243
MCQeasy

A data engineer must load a 500 MB CSV file from Amazon S3 into an existing Amazon Redshift table once per day. The file has a header row and uses a pipe delimiter. The engineer wants the fastest load and the least operational overhead. Which approach should the engineer use?

A.Use the Redshift Data API to run individual INSERT statements for each row in the CSV file.
B.Use an AWS Glue ETL job to read the CSV from S3 and write to Redshift through a JDBC connection.
C.Use AWS Database Migration Service with a change data capture task to replicate the S3 file into Redshift.
D.Use the Redshift COPY command with the IAM role, CSV, HEADER, and DELIMITER options specified.
AnswerD

COPY is Redshift's native bulk loader and reads directly from S3 in parallel, which is the fastest option for a 500 MB file and needs no extra infrastructure. Specifying CSV, HEADER, and the pipe DELIMITER matches the file layout, so the load succeeds without preprocessing and with minimal ongoing operational effort.

Why this answer

The Redshift COPY command is purpose-built for bulk loading from S3 and parallelizes the read, making it the fastest and simplest choice for a single delimited file. The CSV, HEADER, and DELIMITER parameters align the load with the file format, so no staging or transformation layer is needed.

Exam trap

The trap here is reaching for a managed ETL service by default, when Redshift's own COPY command is both faster and simpler for a straightforward file load.

1244
MCQeasy

A company stores raw clickstream data in an Amazon S3 bucket and uses AWS Glue crawlers to populate the AWS Glue Data Catalog. Analysts report that new partitions are not appearing in Amazon Athena queries even though new date-based folders exist in S3. The crawler runs successfully each night. What is the MOST likely cause?

A.The new folders do not follow the Hive-style partition naming convention (for example, year=/month=/day=) that the crawler and Athena expect.
B.The crawler's 'Update the table definition in the Data Catalog' behavior is set to 'Create new partitions only' but the table was created manually without a partition key.
C.The S3 bucket has versioning enabled, which prevents the crawler from detecting new prefixes.
D.The crawler's output is configured to a different database, or the crawler has an exclude pattern that filters out the new date folders.
AnswerA

AWS Glue crawlers and Athena rely on Hive-style partition paths such as year=2024/month=01/day=15 to detect and register partitions automatically. If new folders are named plainly like 2024/01/15 without the key=value structure, the crawler will not recognize them as partitions, so they will not appear in the Data Catalog or in Athena query results.

Why this answer

Glue crawlers infer partitions from Hive-style key=value directory structures. When new folders use a plain date hierarchy without partition keys, the crawler treats them as ordinary prefixes and does not register partitions, so Athena cannot see the new data until the folder layout is corrected or partitions are added manually.

Exam trap

The trap here is assuming any nested folder structure is automatically treated as partitions, when only Hive-style key=value paths are recognized.

1245
MCQeasy

A company is migrating its on-premises MySQL database to Amazon RDS for MySQL. They want to minimize downtime and ensure data consistency. Which AWS service should be used for the migration?

A.AWS S3 Transfer Acceleration
B.AWS Glue
C.AWS Database Migration Service (DMS)
D.AWS Snowball Edge
AnswerC

AWS Database Migration Service performs continuous change data capture from the on-premises MySQL source, replicating ongoing transactions to Amazon RDS for MySQL until cutover. This satisfies the minimal-downtime requirement, while its validation feature confirms row-level data consistency between source and target before the switch.

Why this answer

AWS Database Migration Service (DMS) is the correct choice because it is specifically designed for migrating databases to AWS with minimal downtime. DMS supports homogeneous migrations like MySQL to Amazon RDS for MySQL, and it uses ongoing replication (change data capture) to keep the source and target databases in sync during the migration, ensuring data consistency and allowing a cutover with only seconds of downtime.

Exam trap

The trap here is that candidates may confuse AWS DMS with AWS Glue, thinking both are for data migration, but Glue is for batch ETL and cannot perform live database replication with minimal downtime, while DMS is purpose-built for that task.

How to eliminate wrong answers

Option A is wrong because AWS S3 Transfer Acceleration is a service that speeds up uploads to Amazon S3 by using optimized network paths and edge locations; it has no capability to migrate or replicate a live MySQL database to RDS. Option B is wrong because AWS Glue is a serverless data integration service for ETL (extract, transform, load) jobs, primarily used for preparing and transforming data for analytics, not for ongoing database replication or minimizing downtime during a live database migration. Option D is wrong because AWS Snowball Edge is a physical data transport device used for large-scale data transfers over slow or unreliable networks, but it is not suitable for minimizing downtime in a live database migration as it involves shipping hardware and cannot perform continuous replication.

1246
MCQhard

A data engineer is monitoring an Amazon Redshift cluster and notices that queries are taking longer than expected. The engineer checks the system tables and sees that many queries are waiting for 'WLM' resources. What is the most likely cause and recommended fix?

A.The table sort keys are poorly designed; recreate tables with better sort keys.
B.The distribution style is set to ALL; change to KEY distribution.
C.The WLM queue concurrency is set too low; increase the concurrency level.
D.The cluster is running low on disk space; resize the cluster.
AnswerC

WLM queue waits occur when queries queue behind the concurrency limit. Raising the queue's concurrency lets more queries run simultaneously, clearing the WLM wait state identified in the system tables and reducing overall query latency on the cluster.

Why this answer

When queries wait on WLM resources, it means they are queued because the workload management (WLM) queue does not have enough concurrency slots to run them. Increasing the concurrency level of the WLM queue allows more queries to run simultaneously, reducing wait time. This is the most direct fix for WLM queue waits.

Exam trap

DEA-C01 often tests the misconception that WLM waits are caused by sort keys or distribution styles, when they are actually caused by insufficient queue concurrency slots.

How to eliminate wrong answers

Option A is wrong because poor sort keys cause slow scans and disk I/O, not WLM queue waits; sort key issues would show up as high scan times, not WLM resource contention. Option B is wrong because an ALL distribution style causes data duplication and slower loads, but it does not directly cause WLM queue waits; distribution style affects query execution, not queue concurrency. Option D is wrong because low disk space would trigger disk-based operations or errors, not WLM queue waits; resizing addresses capacity, not concurrency.

1247
Multi-Selectmedium

A company is using AWS Glue to run ETL jobs that transform data from Amazon S3 to Amazon Redshift. The jobs are failing intermittently with 'Out of Memory' errors. The team wants to resolve this issue without increasing costs significantly. Which TWO actions should the team take?

Select 1 answer
A.Increase the Spark memory overhead parameter in the Glue job configuration.
B.Use DynamicFrame instead of Spark DataFrame for transformations.
C.Increase the number of workers to maximum allowed.
D.Switch from a Spark job to a Python shell job.
E.Change the worker type from 'G.1x' to 'G.2x' to double memory per worker.
AnswersE

G.2x workers provide double the memory and vCPU of G.1x at a higher hourly rate, so each executor handles larger partitions and shuffles without spilling to disk, eliminating the out-of-memory failures. This directly addresses the memory constraint while scaling cost proportionally rather than adding workers.

Why this answer

Option E is correct because switching from G.1x to G.2x workers doubles the memory (and vCPU) per worker while keeping the same worker count, which directly addresses OOM conditions in memory-intensive transformations and is a targeted, cost-controlled change rather than scaling out to maximum workers. Option A is not correct because AWS Glue does not provide a 'Spark memory overhead parameter' in the job configuration; memory overhead is not a user-configurable Glue job setting, so this action cannot be taken as described. Option B is not correct because DynamicFrame vs Spark DataFrame is an API choice for schema handling and does not by itself reduce memory pressure enough to fix OOM errors.

Option C is not correct because increasing workers to the maximum allowed raises cost significantly and does not fix per-executor memory limits. Option D is not correct because a Python shell job cannot run distributed Spark ETL at the scale needed for S3-to-Redshift transformations and would not resolve Spark executor OOM issues.

Exam trap

The trap here is that candidates may assume increasing the number of workers (Option C) is the only way to fix OOM errors, but this ignores the cost-effective tuning of worker type upgrades. Another trap is selecting a nonexistent 'memory overhead' parameter (Option A) that sounds plausible but is not a real Glue job configuration option.

1248
Multi-Selecthard

A data engineer is designing a data lake on Amazon S3. The data is ingested from multiple sources and must be queryable using Amazon Athena. The engineer needs to optimize query performance and reduce costs. Which THREE actions would achieve this?

Select 3 answers
A.Store data in many small files to increase parallelism.
B.Partition the data by a commonly used filter column.
C.Use S3 Select instead of Athena for queries.
D.Compress data with a splittable compression format like Snappy.
E.Convert data to Apache Parquet or ORC format.
AnswersB, D, E

Partitioning by a frequently filtered column lets Athena prune irrelevant S3 prefixes via partition metadata, scanning far less data. This directly reduces bytes scanned, cutting both query latency and per-query cost, satisfying the optimisation and cost-reduction requirements.

Why this answer

Option B is correct because partitioning the S3 data by a commonly used filter column (e.g., date or region) lets Athena prune irrelevant partitions via partition projection or the AWS Glue Data Catalog, drastically reducing the bytes scanned and thus query latency and cost. Option D is correct because splittable compression formats like Snappy (and also LZO/BZip2) allow Athena to split large objects across multiple readers, preserving parallelism while lowering storage and scan costs; non-splittable formats like Gzip force a single reader per file. Option E is correct because columnar formats such as Apache Parquet or ORC enable column pruning and predicate pushdown, so Athena reads only the needed columns and row groups, cutting both I/O and cost compared to row-based CSV/JSON.

Option A is incorrect because many small files create excessive per-file overhead in S3 listing, Glue catalog, and Athena planning, hurting performance and raising costs rather than increasing useful parallelism. Option C is incorrect because S3 Select is a lightweight object-level filtering feature, not a replacement for Athena's distributed SQL engine over a data lake, and it cannot perform joins, aggregations, or catalog-based queries.

Exam trap

The trap here is that candidates confuse 'more files = more parallelism' with Athena's actual recommendation of fewer, larger files to minimize the overhead of S3 list and get operations, and they may also mistake S3 Select as a viable alternative to Athena for full SQL querying.

1249
MCQhard

A company uses Amazon EMR to process large datasets stored in Amazon S3. The data is in Parquet format and partitioned by date. The EMR cluster uses Spark SQL for transformations. Recently, the job has been slow and some tasks are failing due to 'java.lang.OutOfMemoryError'. The cluster has 10 core nodes of type m5.xlarge. Which configuration change would MOST improve performance and stability?

A.Increase the number of Spark partitions using repartition(), but keep the same nodes.
B.Change the core node instance type to r5.xlarge (memory-optimized).
C.Increase the number of executor cores in the Spark configuration.
D.Enable Kryo serialization in the Spark configuration.
AnswerB

OutOfMemoryError during Spark SQL tasks indicates executor heap exhaustion, not CPU shortage. Switching core nodes to r5.xlarge increases RAM per node, giving executors more memory for shuffles and aggregations, which addresses the memory constraint causing task failures and slowness.

Why this answer

The error 'java.lang.OutOfMemoryError' indicates that the Spark executors are running out of memory during processing. The m5.xlarge instance type provides 16 GiB of memory, but the workload likely requires more memory per task. Switching to r5.xlarge (32 GiB of memory) doubles the available memory per node, reducing memory pressure and preventing task failures, which directly improves stability and performance for memory-intensive transformations.

Exam trap

The trap here is that candidates often focus on tuning Spark configurations (partitions, cores, serialization) to fix OutOfMemoryErrors, but the real issue is insufficient physical memory per node, which requires a change in instance family rather than software settings.

How to eliminate wrong answers

Option A is wrong because increasing the number of partitions with repartition() can actually increase memory overhead due to shuffle operations and does not address the root cause of insufficient memory per executor; it may even worsen the OutOfMemoryError by creating more tasks that compete for the same limited memory. Option C is wrong because increasing the number of executor cores without increasing memory per core will cause more concurrent tasks to share the same fixed heap, exacerbating memory contention and making OutOfMemoryErrors more likely. Option D is wrong because enabling Kryo serialization reduces the size of serialized objects and improves CPU efficiency, but it does not increase the available heap memory; it cannot prevent OutOfMemoryErrors caused by insufficient memory for data processing.

1250
MCQmedium

A data engineer is managing an Amazon Redshift cluster that experiences performance degradation during peak query hours. The engineer notices that many queries are waiting in the queue, and the WLM query queue wait time is high. The cluster uses automatic WLM. Which action should the engineer take to improve query throughput?

A.Change the cluster's distribution style to EVEN to balance data across nodes.
B.Enable concurrency scaling to automatically add transient clusters for bursts of queries.
C.Increase the number of nodes in the cluster to add more compute resources.
D.Modify the WLM configuration to increase the concurrency level for the queue.
AnswerB

Concurrency scaling in Amazon Redshift automatically adds transient clusters to handle bursts of queries, reducing queue wait times. It is designed for workloads with unpredictable query spikes. When queries queue, concurrency scaling kicks in to provide additional capacity. This directly addresses the high queue wait time during peak hours without manual intervention. It is a feature of both automatic and manual WLM.

Why this answer

Concurrency scaling automatically adds transient clusters to handle query bursts, reducing queue wait times. It is specifically designed to improve throughput during peak periods. Manually increasing concurrency is not possible with automatic WLM.

Adding nodes or changing distribution style does not directly address queue wait times caused by concurrency limits.

Exam trap

The trap here is assuming that automatic WLM allows manual concurrency adjustments, or that adding nodes is the first step for queue wait times.

1251
Multi-Selectmedium

A data engineer manages an AWS Lake Formation governed data lake. Analysts in the finance department must query only the rows in a shared Amazon S3 table where the region column equals 'EMEA', while analysts in the marketing department must see all rows but must not see the customer_email column. Which TWO Lake Formation configurations should the data engineer implement to meet these requirements? (Choose two.)

Select 2 answers
A.Create a data filter on the table that excludes the customer_email column and grant SELECT on the table with that data filter to the marketing analyst role.
B.Define an AWS Glue Data Catalog table property named column.filter with the value customer_email and a table property row.filter with the value region='EMEA', then grant DESCRIBE to both roles.
C.Create a data filter on the table that includes a row filter expression of region = 'EMEA' and grant SELECT on the table with that data filter to the finance analyst role.
D.Register the S3 location with Lake Formation in hybrid access mode and grant the analysts ALL permissions so that IAM policies alone control row and column visibility.
E.Attach an IAM policy to the finance analyst role that denies s3:GetObject on any S3 prefix whose object metadata does not contain a region tag of EMEA.
AnswersA, C

A Lake Formation data filter can specify an included column list that omits customer_email. Granting SELECT with that column-scoped data filter to the marketing role means queries through integrated engines return all rows but cannot project the excluded column. This implements the column-level restriction without duplicating the underlying S3 data.

Why this answer

Lake Formation data filters are the native mechanism for row-level and column-level security. A data filter with a row expression scopes which rows a principal can read, and a data filter with an included column list scopes which columns are visible. Granting SELECT with the appropriate data filter to each analyst role enforces both requirements across integrated engines without copying or duplicating the underlying S3 objects.

Exam trap

The trap here is reaching for IAM or S3 object metadata to express row predicates, when those layers cannot filter individual rows inside a data file and Lake Formation data filters are the intended control.

1252
Multi-Selecthard

A data engineer is troubleshooting an Amazon Redshift cluster where nightly COPY loads from Amazon S3 are intermittently slow and sometimes fail with 'S3ServiceException' errors. The engineer suspects the cluster's network configuration and load design are contributing. Which TWO actions should the engineer take to improve load performance and reliability? (Choose two.)

Select 2 answers
A.Convert all input files to a single large gzip file to minimize the number of S3 GET requests during COPY.
B.Increase the cluster's node count and change the distribution style of all tables to EVEN before every load.
C.Disable compression on the input files so Redshift can parse records faster.
D.Split large input files into multiple files sized roughly between 1 MB and 1 GB and load them in parallel.
E.Enable Amazon Redshift enhanced VPC routing so COPY traffic to S3 stays within the VPC and avoids public internet paths.
AnswersD, E

Redshift parallelizes COPY across slices, and the number of files should be a multiple of the slice count for even distribution. Splitting large files into many smaller files within the 1 MB to 1 GB range lets each slice read its own file, dramatically improving load performance and reducing the chance that a single large file bottlenecks the load.

Why this answer

Redshift COPY performance depends on parallelizing reads across slices and on a stable network path to S3. Splitting input into many files sized between 1 MB and 1 GB lets slices read concurrently, while enhanced VPC routing keeps S3 traffic inside the VPC and avoids public internet variability that produces intermittent S3ServiceException errors.

Exam trap

The trap here is assuming that fewer, larger compressed files reduce overhead, when Redshift COPY actually requires many files to parallelize across slices.

1253
MCQhard

A company uses Amazon Redshift for data warehousing. They notice that queries are running slowly, and the STL_LOAD_ERRORS table shows many 'Parse error' entries. The data is loaded from Amazon S3 using COPY commands. What is the MOST likely cause of the parse errors?

A.The source data files have a different schema or delimiter than what is specified in the COPY command.
B.The Redshift cluster does not have enough compute nodes to process the data.
C.The IAM role used by Redshift does not have permission to decrypt the S3 objects.
D.The source data files are compressed using an unsupported compression format.
AnswerA

COPY parses each S3 file against the declared column layout and delimiter. A mismatch between the file's actual schema or delimiter and the COPY parameters causes field parsing to fail, producing the Parse error entries in STL_LOAD_ERRORS. Aligning the COPY specification with the source files resolves it.

Why this answer

COPY loads data by parsing each source record according to the FORMAT, DELIMITER, and column definitions supplied in the command. When the actual file layout (column count, order, or delimiter) does not match those specifications, Redshift cannot map fields to columns and writes a 'Parse error' row into STL_LOAD_ERRORS. This is the classic signature of a schema/delimiter mismatch, not a resource or permission problem.

Exam trap

The trap here is conflating any COPY failure with infrastructure or permissions; candidates must recognize that 'Parse error' in STL_LOAD_ERRORS specifically points to data-format/schema mismatch, not cluster capacity or IAM.

How to eliminate wrong answers

Option B is wrong because insufficient compute nodes cause slow query execution or queuing, not parse errors during COPY; node count has no bearing on record parsing. Option C is wrong because a missing IAM decrypt permission produces an S3 access-denied or 'S3ServiceException' error, not a parse error. Option D is wrong because an unsupported compression format fails at the file/format detection stage with a compression-related error, not a per-record parse error.

1254
Multi-Selectmedium

Which TWO actions can help improve the read performance of an Amazon DynamoDB table that is experiencing throttling? (Choose two.)

Select 2 answers
A.Enable DynamoDB Accelerator (DAX) for write-heavy workloads.
B.Increase the provisioned read capacity units (RCUs) for the table.
C.Use eventually consistent reads instead of strongly consistent reads.
D.Add a global secondary index (GSI) with a different partition key.
E.Enable auto-scaling with a lower minimum capacity.
AnswersB, C

Raising provisioned RCUs directly lifts the table's read throughput ceiling, so sustained request rates that previously exceeded capacity no longer trigger ProvisionedThroughputExceededException throttling. This satisfies the stem's throttling constraint by matching allocated capacity to actual read demand, provided the workload is read-capacity-bound rather than partition-hot.

Why this answer

Option B is correct because DynamoDB throttling on reads occurs when request rate exceeds the table's provisioned read capacity units (RCUs); increasing the provisioned RCUs directly raises the available read throughput and relieves the throttling. Option C is correct because eventually consistent reads consume only half the RCUs of strongly consistent reads (1 RCU per 4 KB vs. 1 RCU per 2 KB), so switching to eventually consistent reads effectively doubles the read throughput available for the same provisioned capacity. Option A is wrong because DAX is designed to accelerate read-heavy workloads by caching, not write-heavy ones, and it does not address the stated read throttling.

Option D is wrong because adding a GSI does not increase the base table's read capacity and actually adds its own separate capacity costs and potential throttling. Option E is wrong because auto-scaling with a lower minimum capacity would reduce available RCUs and worsen, not improve, read throttling.

Exam trap

The trap is that candidates often overlook eventually consistent reads as a cost-effective way to reduce throttling, focusing instead on adding GSIs or enabling DAX, which address latency or query patterns but not the root cause of exceeding provisioned capacity.

1255
MCQeasy

A company uses Amazon EMR to run Spark jobs on a transient cluster. The jobs process data from S3 and write results back to S3. The team wants to reduce costs by optimizing the cluster. Which action should the team take?

A.Use Spot Instances for the task nodes.
B.Increase the number of core nodes and use larger instance types.
C.Enable EMRFS consistent view.
D.Terminate the cluster after each job and manually restart it for the next job.
AnswerA

Task nodes perform no HDFS storage, so their loss does not corrupt data, making them safe for Spot capacity. Spot Instances cost substantially less than On-Demand, cutting transient cluster expense while core nodes retain On-Demand reliability.

Why this answer

Using Spot Instances for task nodes in a transient EMR cluster significantly reduces compute costs because Spot Instances are spare AWS EC2 capacity offered at up to 90% discount compared to On-Demand. Since transient clusters are terminated after job completion, the risk of Spot Instance interruptions is mitigated—the job can simply be retried on a new cluster if needed. This directly addresses the cost optimization goal without sacrificing job functionality.

Exam trap

The trap here is that candidates confuse cost optimization with performance tuning, leading them to choose larger instances (Option B) or consistency features (Option C), when the real cost lever for transient workloads is leveraging Spot pricing for ephemeral compute.

How to eliminate wrong answers

Option B is wrong because increasing the number of core nodes and using larger instance types increases costs, not reduces them, and core nodes host HDFS which is unnecessary for a transient cluster that reads/writes directly to S3. Option C is wrong because EMRFS consistent view is a feature to handle S3 eventual consistency for listing and renaming, not a cost optimization mechanism—it adds overhead without reducing spend. Option D is wrong because EMR transient clusters already terminate automatically after the job completes; manually restarting is redundant and does not further reduce costs, and it introduces operational overhead.

1256
MCQeasy

A data engineer runs the above SQL commands on an Amazon Redshift cluster. The table 'users' is created with DISTSTYLE EVEN. What is the effect of the DISTSTYLE EVEN on query performance?

A.It stores all data on a single node for fast local queries.
B.It ensures data is evenly distributed across all nodes to prevent data skew.
C.It reduces data movement during queries by co-locating data based on user_id.
D.It improves join performance when joining on user_id.
AnswerB

DISTSTYLE EVEN spreads rows round-robin across every compute node slice, guaranteeing uniform row counts regardless of column values. This directly satisfies the stem's concern by preventing data skew, which would otherwise concentrate rows on some slices and force other nodes to idle during scans and joins.

Why this answer

DISTSTYLE EVEN in Amazon Redshift distributes rows across all nodes in a round-robin fashion, ensuring each node holds approximately the same amount of data. This prevents data skew, which can cause some nodes to become bottlenecks, and improves overall query performance for workloads that do not benefit from key-based distribution. It is the correct choice because it directly addresses the goal of balanced data distribution.

Exam trap

The trap here is that candidates often confuse DISTSTYLE EVEN with DISTSTYLE ALL, thinking EVEN improves performance by keeping data local, when in fact EVEN distributes data to prevent skew, not to localize it.

How to eliminate wrong answers

Option A is wrong because DISTSTYLE EVEN does not store all data on a single node; that describes DISTSTYLE ALL, which replicates the entire table to every node. Option C is wrong because EVEN distribution does not co-locate data based on user_id; that behavior is achieved with DISTSTYLE KEY, which distributes rows by the hash of a specified column. Option D is wrong because EVEN distribution does not improve join performance on user_id; joins on user_id benefit from DISTSTYLE KEY on user_id to enable collocated joins, while EVEN may require data redistribution across nodes during query execution.

1257
MCQmedium

A data engineering team uses AWS Glue ETL jobs to process data daily. They notice that job run times are increasing as data volume grows. Which action will most effectively improve performance without changing the code?

A.Use a smaller instance type for the Glue job.
B.Enable job bookmark to skip previously processed data.
C.Split the data into more files in S3.
D.Increase the number of DPUs for the Glue job.
AnswerD

Increasing DPUs adds more Apache Spark executors, raising parallel task capacity so each stage processes partitions faster. This directly addresses the growing data volume constraint while requiring no code changes, since Glue scales worker allocation automatically. It outperforms alternatives that alter code or partitioning logic.

Why this answer

Increasing the number of DPUs (Data Processing Units) for the Glue job directly scales the compute resources allocated to the job, which can improve performance and reduce run times as data volume grows. This does not require code changes. Job bookmarks help skip already processed data but do not address increasing data volume.

Smaller instance types and more files in S3 may not effectively improve performance.

Exam trap

The trap is thinking that job bookmarks or file splitting will improve performance for growing data; candidates might overlook that scaling DPUs directly addresses compute needs.

How to eliminate wrong answers

Option A is wrong because using a smaller instance type reduces compute resources, which would likely worsen performance. Option B is wrong because enabling job bookmark only skips previously processed data; it does not improve performance for new data and may not be applicable if all data is new. Option C is wrong because splitting data into more files can sometimes improve parallelism, but it is not the most effective action and may not help if the job is already parallel; increasing DPUs directly scales compute.

1258
Multi-Selectmedium

A company uses AWS Glue to catalog data stored in Amazon S3. The data is in Parquet format and partitioned by date. The company wants to improve query performance in Amazon Athena and reduce costs. Which THREE actions should the company take? (Choose THREE.)

Select 3 answers
A.Convert the data to JSON format for better schema evolution.
B.Use Glue DataBrew to clean the data before querying.
C.Partition the data by date so Athena can use partition pruning.
D.Ensure the data is in a columnar format like Parquet or ORC.
E.Compress the data using a codec like Snappy or Gzip.
AnswersC, D, E

Partition pruning limits the amount of data scanned per query.

Why this answer

Partitioning the data by date allows Athena to use partition pruning, which limits the amount of data scanned by only reading the partitions that match the query's WHERE clause. This directly reduces both query cost (since Athena charges per byte scanned) and query latency, especially for date-range queries on large datasets.

Exam trap

The trap here is that candidates may confuse data preparation tools (like Glue DataBrew) with query optimization techniques, or mistakenly think that converting to a non-columnar format like JSON improves schema evolution, when in fact columnar formats with compression and partitioning are the standard best practices for Athena performance and cost efficiency.

1259
MCQhard

A company is designing a data pipeline using Amazon Kinesis Data Streams. The data includes personally identifiable information (PII). The security team requires that data be encrypted at rest using a customer-managed KMS key. How should the data engineer configure the Kinesis stream?

A.Configure the Kinesis stream to use AWS CloudHSM for encryption.
B.Enable server-side encryption on the Kinesis stream and specify the customer-managed KMS key.
C.Store the encrypted data in S3 and use Kinesis to stream the S3 object keys.
D.Use client-side encryption in the producer application to encrypt data before sending to Kinesis.
AnswerB

Server-side encryption with a customer-managed KMS key encrypts data at rest within the stream's storage layer, satisfying the requirement for customer-controlled key management. Kinesis Data Streams supports specifying a customer-managed key rather than the AWS-owned default, giving the security team the key control and rotation they demanded.

Why this answer

Amazon Kinesis Data Streams supports server-side encryption (SSE) with AWS KMS, and you can choose either an AWS-managed key (aws/kinesis) or a customer-managed KMS key. To meet the requirement of encryption at rest with a customer-managed key, you enable SSE on the stream and specify the customer-managed KMS key, which Kinesis uses to encrypt data as it is written to storage.

Exam trap

DEA-C01 often tests the confusion between client-side encryption (producer encrypts before sending) and server-side encryption with a customer-managed KMS key, which is what the requirement explicitly asks for.

How to eliminate wrong answers

Option A is wrong because CloudHSM is not an encryption option for Kinesis Data Streams; Kinesis SSE integrates with AWS KMS, not CloudHSM directly, so this configuration is not supported. Option C is wrong because storing data in S3 and streaming only object keys does not encrypt the Kinesis stream itself and changes the architecture rather than satisfying the requirement that the stream be encrypted at rest with a customer-managed KMS key. Option D is wrong because client-side encryption encrypts data before it reaches Kinesis, but it does not provide server-side encryption at rest with a KMS key and does not meet the stated requirement of using a customer-managed KMS key for the stream.

1260
MCQhard

A data engineer is building an AWS Glue ETL job that reads from an Amazon S3 bucket with millions of small JSON files. The job is running slowly and often fails with out-of-memory errors. The engineer needs to improve performance and reliability. What should the engineer do?

A.Enable job bookmarks to process only new files and reduce the workload.
B.Use AWS Glue's groupFiles and groupSize parameters to combine small files into larger groups.
C.Increase the number of DPUs for the Glue job to handle the large number of files.
D.Convert the JSON files to Parquet using an AWS Glue crawler before running the ETL job.
AnswerB

AWS Glue provides groupFiles and groupSize to coalesce multiple small files into a single partition read. This reduces the number of tasks and metadata overhead, improving performance and reducing memory pressure. It is the recommended approach for large numbers of small files in S3 when using Glue ETL.

Why this answer

When an AWS Glue job reads millions of small files, the overhead of listing and opening each file causes performance degradation and memory issues. The groupFiles and groupSize parameters allow Glue to combine multiple small files into larger read groups, reducing task count and improving throughput. This is the most direct and effective solution for this scenario.

Exam trap

The trap here is assuming that simply adding more DPUs will solve performance problems caused by many small files, when the real fix is to reduce file count through grouping.

1261
MCQmedium

A company is using AWS Lake Formation to manage access to a data lake in S3. They want to grant a data analyst access to specific columns in a table, but not to the entire table. Which Lake Formation feature should be used?

A.Row-level security (cell-level filtering)
B.IAM policies on the S3 bucket
C.Column-level filtering
D.Tag-based access control (TBAC)
AnswerC

Column-level filtering in Lake Formation applies column-level permissions on a table, letting the analyst query only the granted columns while excluded columns are hidden. This satisfies the stem's requirement to restrict access to specific columns rather than the entire table.

Why this answer

Lake Formation column-level filtering allows granting access to specific columns in a table without granting access to the entire table. Option A (row-level security) controls access to rows, not columns. Option B (IAM policies on the S3 bucket) would grant access to the entire dataset or bucket, not specific columns.

Option D (tag-based access control) uses tags to manage permissions but does not provide column-level granularity.

1262
Multi-Selecthard

A data engineer is designing an AWS Glue ETL job that reads data from an Amazon S3 bucket containing many small JSON files and writes the output to Amazon S3 in Parquet format. The engineer wants to improve read performance and reduce the number of output files. Which two actions should the engineer take? (Choose two.)

Select 2 answers
A.Increase the number of DPUs to allow more executors to read the small files in parallel.
B.Call repartition or coalesce on the DynamicFrame before writing to control the number of output files.
C.Enable job bookmarks to skip previously processed files and reduce the number of files read.
D.Convert the source JSON files to Parquet before the Glue job runs using an AWS Lambda function triggered on each S3 object.
E.Use the AWS Glue groupFiles option to group multiple small files into a single partition during the read.
AnswersB, E

Repartitioning or coalescing before the write step controls how many partitions exist and therefore how many output files are produced. Coalesce reduces the number of partitions without a full shuffle, which is efficient when reducing file count. This action directly reduces the number of output files written to S3, complementing the read-side grouping.

Why this answer

Grouping small files during the read reduces file-open overhead and improves read performance, while repartitioning or coalescing before writing controls the number of output files. Job bookmarks only help across runs, adding DPUs does not fix per-file overhead or output count, and pre-converting files with Lambda does not address the small-file or output-file problems within the Glue job.

Exam trap

The trap here is thinking that adding more compute or enabling bookmarks will solve the small-file problem, when the real fixes are read-side file grouping and write-side partition control.

1263
MCQhard

A data engineer needs to grant a data scientist access to query a Glue Data Catalog database but must prevent the data scientist from seeing the underlying S3 data locations. Which approach should be used?

A.Use a Glue resource policy to restrict access to the database
B.Grant the data scientist IAM permissions to access the Glue Data Catalog and the underlying S3 data
C.Create a VPC endpoint for Glue and S3 to restrict network access
D.Use AWS Lake Formation to grant SELECT permission on the database and tables without granting S3 access
AnswerD

Lake Formation enforces table-level permissions through its own access layer, so the data scientist queries via Athena without IAM or S3 bucket policies exposing the storage location. Granting SELECT on the database and tables satisfies the query requirement while the underlying S3 paths remain hidden, because Lake Formation mediates access rather than S3 directly.

Why this answer

AWS Lake Formation allows fine-grained access control at the database, table, and column level without requiring direct S3 permissions. By granting SELECT on the Glue Data Catalog database and tables through Lake Formation, the data scientist can query the data via Athena or Redshift Spectrum while Lake Formation handles credential vending to access S3 on their behalf. This meets the requirement of preventing the data scientist from seeing the underlying S3 locations.

Exam trap

DEA-C01 often tests the misconception that Glue resource policies alone can restrict S3 visibility — candidates must recognize that only Lake Formation's credential vending hides S3 locations while still enabling queries.

How to eliminate wrong answers

Option A is wrong because a Glue resource policy controls access to the Data Catalog API itself but does not provide a mechanism to query data without S3 permissions, nor does it hide S3 locations from users who also have S3 access. Option B is wrong because granting IAM permissions to both Glue and S3 would allow the data scientist to see and access the S3 bucket directly, violating the requirement. Option C is wrong because a VPC endpoint only controls network routing and does not provide authorization or hide S3 locations from IAM principals.

1264
MCQeasy

A company uses AWS Glue to run ETL jobs daily. The data is stored in S3 as Parquet files partitioned by date. Recently, jobs have failed with the error 'No such file or directory' for certain partitions. What is the MOST likely cause?

A.The schema has changed and Glue cannot parse the data.
B.A partition folder was deleted or not created by the upstream process.
C.The files are compressed with an unsupported codec.
D.The IAM role does not have s3:GetObject permission.
AnswerB

Glue resolves each partition path from the catalog at run time; if the upstream job failed to write a partition folder or it was deleted, the read fails with 'No such file or directory'. This matches the stem's missing-partition error rather than a schema or permissions issue.

Why this answer

The error 'No such file or directory' indicates that the Glue ETL job is attempting to read a specific S3 partition path that does not exist. Since the data is partitioned by date and the job runs daily, the most likely cause is that the upstream process failed to create or accidentally deleted the partition folder for that date. Glue's dynamic frame or Spark DataFrame will throw this error when it tries to list or read files from a missing prefix.

Exam trap

The trap here is that candidates confuse file-level permission errors (Option D) with missing directory errors, but S3 returns distinct HTTP status codes (403 vs 404) that map to different error messages in Spark/Glue.

How to eliminate wrong answers

Option A is wrong because a schema change would typically cause a parsing or schema mismatch error (e.g., 'Schema mismatch' or 'Cannot convert type'), not a 'No such file or directory' error. Option C is wrong because unsupported compression codecs (e.g., LZO without proper libraries) would cause a 'Codec not found' or 'Compression error', not a missing file error. Option D is wrong because missing s3:GetObject permission would result in an 'Access Denied' (403) error, not a 'No such file or directory' error.

1265
Multi-Selecteasy

A data engineer is setting up a data pipeline using AWS DMS to migrate data from an on-premises database to Amazon RDS for MySQL. The data must be encrypted in transit. Which TWO options can the engineer use? (Choose TWO.)

Select 2 answers
A.Use VPC peering between on-premises and AWS
B.Enable SSL encryption on the DMS endpoint
C.Set up a VPN connection between on-premises and AWS
D.Use KMS to encrypt the DMS connection
E.Use a VPC endpoint for DMS
AnswersB, C

AWS DMS endpoints expose an SSL mode setting; enabling it encrypts replication traffic between the source, replication instance and target. This directly satisfies the in-transit encryption requirement for the on-premises to Amazon RDS for MySQL migration.

Why this answer

Option B is correct because AWS DMS endpoints for MySQL support SSL/TLS, and enabling the SSL encryption setting on the source and target endpoints causes DMS to negotiate an encrypted connection to the database, satisfying the in-transit encryption requirement. Option C is correct because a VPN connection (AWS Site-to-Site VPN) creates an IPsec-encrypted tunnel between the on-premises network and the AWS VPC, protecting data in transit as it crosses the public internet to reach Amazon RDS for MySQL. Option A is not correct because VPC peering only connects AWS VPCs to each other and does not extend to an on-premises network, so it cannot secure this migration path.

Option D is not correct because AWS KMS provides encryption at rest for stored data and keys, not encryption of the DMS network connection in transit. Option E is not correct because a VPC endpoint (AWS PrivateLink) provides private connectivity to AWS services within AWS, not an encrypted path from an on-premises database to RDS.

1266
Multi-Selectmedium

A data engineer is designing a data ingestion pipeline for real-time clickstream data from a website. The data must be ingested with low latency (seconds) and made available for multiple consumer applications, including a dashboard that refreshes every minute and a machine learning model that processes data in near-real-time. The engineer needs to choose a streaming ingestion service. Which TWO services meet these requirements? (Select TWO.)

Select 2 answers
A.Amazon Kinesis Data Firehose
B.Amazon Managed Streaming for Apache Kafka (Amazon MSK)
C.Amazon Kinesis Data Streams
D.Amazon Simple Queue Service (SQS)
E.Amazon S3
AnswersB, C

MSK is a fully managed Kafka service that provides low-latency streaming and supports multiple consumer groups.

Why this answer

Amazon Kinesis Data Streams (C) provides sub-second ingestion latency and supports multiple consumer applications via its enhanced fan-out feature, enabling a dashboard and ML model to consume data concurrently with low latency. Amazon MSK (B) offers similar real-time capabilities with Apache Kafka's native pub/sub model, allowing multiple consumers to process the same stream independently and with low latency, meeting the near-real-time requirements.

Exam trap

The trap here is that candidates often confuse Amazon Kinesis Data Firehose with Kinesis Data Streams, assuming Firehose provides real-time ingestion, but Firehose buffers data for at least 60 seconds before delivery, making it unsuitable for sub-second latency requirements.

1267
MCQmedium

A company uses Amazon Kinesis Data Streams to ingest real-time logs from thousands of applications. The data must be transformed and enriched with reference data from Amazon S3 before being stored in Amazon S3 in Parquet format. The transformation logic is stateful and requires exactly-once processing. Which AWS service should the data engineer use to perform the transformation?

A.Amazon Kinesis Data Firehose with AWS Lambda transformation.
B.AWS Glue streaming ETL job.
C.Amazon Managed Service for Apache Flink (formerly Kinesis Data Analytics) with a Flink application.
D.AWS Lambda with Kinesis Data Streams event source mapping.
AnswerC

Amazon Managed Service for Apache Flink supports stateful stream processing with exactly-once semantics. It can read from Kinesis Data Streams, enrich data by joining with reference data from S3 (using Flink's rich functions or async I/O), and write to S3 in Parquet format using the FileSystem connector. This service is designed for complex, stateful transformations and ensures exactly-once processing.

Why this answer

The need for stateful processing, exactly-once semantics, and enrichment with reference data from S3 points to a stream processing framework. Amazon Managed Service for Apache Flink provides these capabilities with native support for state, exactly-once processing, and connectors to Kinesis and S3. Other services either lack stateful processing, exactly-once guarantees, or both.

Exam trap

The trap here is assuming that AWS Glue streaming ETL or Lambda can handle stateful, exactly-once processing, when they are better suited for simpler transformations.

1268
MCQeasy

A company wants to enforce that all data in an S3 bucket is encrypted at rest using AWS KMS. Which bucket policy condition key should be used?

A.s3:x-amz-acl with value bucket-owner-full-control
B.s3:x-amz-server-side-encryption with value aws:kms
C.s3:x-amz-server-side-encryption with value AES256
D.aws:SourceIp with value 10.0.0.0/8
AnswerB

The s3:x-amz-server-side-encryption condition key with value aws:kms enforces that any PutObject request specifies AWS KMS encryption. This denies uploads lacking the required header, satisfying the mandate that all bucket data is encrypted at rest using KMS.

Why this answer

The condition key `s3:x-amz-server-side-encryption` with value `aws:kms` enforces that objects uploaded to the S3 bucket must be encrypted using AWS KMS (SSE-KMS). This bucket policy condition ensures that any PUT request includes the `x-amz-server-side-encryption` header set to `aws:kms`, thereby enforcing encryption at rest with KMS-managed keys.

Exam trap

The trap here is that candidates confuse `aws:kms` with `AES256` (SSE-S3), thinking both enforce KMS encryption, but only `aws:kms` enforces AWS KMS, while `AES256` enforces S3-managed keys.

How to eliminate wrong answers

Option A is wrong because `s3:x-amz-acl` with value `bucket-owner-full-control` enforces access control list ownership, not encryption. Option C is wrong because `s3:x-amz-server-side-encryption` with value `AES256` enforces SSE-S3 (Amazon S3-managed keys), not AWS KMS. Option D is wrong because `aws:SourceIp` restricts requests based on IP address, which is unrelated to encryption enforcement.

1269
Matchingmedium

Match each AWS monitoring tool to its primary use.

Drag a concept onto its matching description — or click a concept then click the description.

Concepts
Matches

Metrics, logs, and alarms

API call history and auditing

Trace and analyze distributed applications

Event-driven automation

Resource configuration tracking

Why these pairings

CloudWatch monitors metrics and logs, CloudTrail audits API calls, X-Ray traces requests, and Trusted Advisor provides optimization recommendations. Common confusions include swapping the roles of CloudWatch and CloudTrail.

1270
MCQeasy

A company uses Amazon S3 as a data lake. A data engineer needs to ensure that all objects uploaded to the 'incoming' prefix are automatically encrypted at rest using AWS KMS with a specific customer managed key. What is the simplest way to enforce this?

A.Enable S3 Transfer Acceleration to force encryption in transit.
B.Use a bucket policy that denies PutObject requests without the required encryption header.
C.Configure S3 Inventory to report on encryption status and alert on non-compliance.
D.Enable default encryption on the bucket with SSE-S3.
AnswerB

A bucket policy denying PutObject requests lacking the required encryption header enforces AWS KMS encryption server-side at upload time. It applies to every principal without application changes, satisfying the requirement to use a specific customer managed key.

Why this answer

A bucket policy that denies s3:PutObject unless the request includes the required encryption header (e.g., x-amz-server-side-encryption: aws:kms and x-amz-server-side-encryption-aws-kms-key-id) enforces encryption at upload time with the specified customer managed key. This is the simplest enforcement mechanism because it operates at the bucket level and rejects non-compliant uploads before the object is written. It also works regardless of which client or SDK performs the upload.

Exam trap

DEA-C01 often tests the confusion between detective controls (S3 Inventory, Config rules) and preventive controls (bucket policies with deny conditions), tempting candidates to pick a reporting tool when enforcement is required.

How to eliminate wrong answers

Option A is wrong because S3 Transfer Acceleration only speeds up uploads over long distances using AWS edge locations; it has nothing to do with encryption at rest. Option C is wrong because S3 Inventory generates reports on object metadata including encryption status, but it is a detective control that reports after the fact — it cannot prevent unencrypted uploads. Option D is wrong because default bucket encryption with SSE-S3 applies AES-256 with S3-managed keys, not the customer managed KMS key required, and default encryption does not enforce a specific key on PutObject requests.

1271
MCQhard

A data engineer is using AWS Database Migration Service (AWS DMS) to replicate ongoing changes from an Amazon RDS for PostgreSQL database to an Amazon S3 bucket. The source table has a primary key and the engineer needs near-real-time change data capture (CDC) with minimal impact on the source. Which DMS task setting should the engineer configure to meet these requirements?

A.Set the migration type to full load plus CDC and enable logical replication on the source
B.Set the migration type to full load only and schedule hourly re-runs
C.Set the migration type to CDC only and configure a change table on the source
D.Set the migration type to full load only and enable binary logging on the source
AnswerA

Full load plus CDC performs an initial load and then continuously reads changes from the PostgreSQL write-ahead log using logical replication. This provides near-real-time replication of inserts, updates, and deletes while reading from the log rather than querying tables, which minimizes impact on the source database. It is the correct DMS configuration for ongoing CDC.

Why this answer

AWS DMS uses PostgreSQL logical replication to read changes from the write-ahead log for CDC. Configuring the task as full load plus CDC performs the initial copy and then continuously applies changes, giving near-real-time replication with minimal source impact because changes are read from the log rather than by polling tables. Enabling logical replication on the source is a prerequisite for this mode.

Exam trap

The trap here is confusing full load only with CDC or assuming that any logging feature (such as MySQL binary logging) applies to PostgreSQL; the correct CDC mechanism for PostgreSQL in DMS is logical replication from the WAL.

1272
Drag & Dropmedium

Arrange the steps to implement data encryption at rest for an Amazon Redshift cluster using AWS KMS.

Drag or tap steps into the slots.

Steps
Order
1Step 1
2Step 2
3Step 3
4Step 4

Why this order

First, create the KMS key. Then launch a new encrypted cluster, specify the key, configure, and verify encryption.

1273
MCQmedium

A company uses Amazon Redshift for data warehousing. The security team requires that all data loading into Redshift be encrypted in transit. Which configuration ensures this requirement is met?

A.Use a VPC security group to restrict access
B.Configure the Redshift cluster to require SSL connections
C.Use client-side encryption before loading data
D.Enable server-side encryption on the Redshift cluster
AnswerB

Setting the cluster parameter require_ssl to true rejects any connection or COPY/UNLOAD operation that does not use TLS, so data in transit is encrypted. This directly satisfies the mandate that all loading into Redshift be encrypted in transit, without altering at-rest encryption or IAM permissions.

Why this answer

Configuring the Redshift cluster to require SSL connections (require_ssl = true in the parameter group) enforces TLS for all client connections, ensuring data is encrypted in transit during loads. This is the direct configuration that satisfies the in-transit encryption requirement.

Exam trap

The trap is confusing encryption at rest (SSE on the cluster) with encryption in transit (SSL required) — candidates pick server-side encryption thinking it covers network traffic.

How to eliminate wrong answers

Option A is wrong because VPC security groups control network reachability, not encryption; they do not encrypt traffic. Option C is wrong because client-side encryption protects data at rest before upload, not the in-transit channel to Redshift. Option D is wrong because server-side encryption on Redshift encrypts data at rest on disk, not data moving over the network.

1274
MCQmedium

A data engineer stores Apache Parquet files in an Amazon S3 data lake partitioned by dt=YYYY-MM-DD. Analysts query the data with Amazon Athena, and monthly reports that scan one month of data are slow and expensive. The engineer confirms that queries filter on the dt column. Which action will MOST effectively reduce the amount of data scanned by these reports?

A.Convert the Parquet files to CSV so that Athena can read them faster.
B.Enable Amazon S3 Transfer Acceleration on the data lake bucket.
C.Run MSCK REPAIR TABLE on the table, or enable partition projection for the dt partition column.
D.Increase the Athena per-query data limit in the workgroup settings.
AnswerC

Athena prunes partitions only when the table's partition metadata matches the S3 prefixes. If new dt= prefixes were written without registering partitions in the AWS Glue Data Catalog, every query falls back to scanning the whole table. Running MSCK REPAIR TABLE, or configuring partition projection so Athena derives partitions from a pattern, restores pruning so a one-month filter reads only that month's prefixes.

Why this answer

Athena reduces scan size through partition pruning, which requires the table's partition metadata to reflect the actual S3 prefixes. When files are written as dt=YYYY-MM-DD without catalog registration, filters on dt cannot eliminate prefixes, so monthly reports scan the entire table. Registering partitions with MSCK REPAIR TABLE or defining partition projection restores pruning, cutting both bytes scanned and cost.

Exam trap

The trap here is assuming that switching file formats or tuning the workgroup fixes slow Athena queries, when the real issue is unregistered partitions that disable pruning.

1275
MCQeasy

A data engineer needs to ingest data from an Amazon S3 bucket into Amazon Redshift for analytics. The data is in CSV format and the Redshift table already exists. Which service can be used to perform this ingestion with minimal configuration?

A.AWS Glue
B.Amazon Kinesis Data Firehose
C.Amazon Redshift COPY command
D.AWS Database Migration Service (DMS)
AnswerC

The COPY command loads CSV files directly from Amazon S3 into an existing Redshift table using parallel processing, requiring only an IAM role for authorisation. It satisfies the minimal-configuration constraint because no external ETL service, cluster provisioning, or intermediate staging is needed.

Why this answer

The Amazon Redshift COPY command is the most direct and minimal-configuration method to load data from an S3 bucket into an existing Redshift table. It is purpose-built for bulk data ingestion from S3, supports CSV format natively, and requires only the table name, S3 path, IAM role, and format options — no additional services or pipelines needed.

Exam trap

The trap here is that candidates often overthink and choose AWS Glue or Kinesis Firehose because they assume a managed service is always required, but the COPY command is the simplest and most efficient native tool for batch loading from S3 into an existing Redshift table.

How to eliminate wrong answers

Option A is wrong because AWS Glue is an ETL service that requires creating crawlers, jobs, and triggers, which adds unnecessary complexity for a simple S3-to-Redshift load; it is overkill when the COPY command suffices. Option B is wrong because Amazon Kinesis Data Firehose is designed for streaming data ingestion into Redshift, not for batch loading from static CSV files in S3, and it requires configuring a delivery stream, buffering, and transformation. Option D is wrong because AWS DMS is used for continuous database migration and replication between databases, not for one-time bulk loading of CSV files from S3 into an existing Redshift table.

Page 16

Page 17 of 18

Page 18