Courseiva

AWS Certified Data Engineer Associate DEA-C01 (DEA-C01) — Questions 76–150

1321 questions total · 18pages · All types, answers revealed

Page 1

Page 2 of 18

Page 3
76
MCQhard

A data pipeline uses Amazon Kinesis Data Firehose to deliver data to an S3 bucket. The delivery stream is configured with a buffer interval of 60 seconds and a buffer size of 5 MB. The data arrives at an average rate of 2 MB per second. What is the expected time interval between S3 writes?

A.Approximately 2.5 seconds
B.Approximately 30 seconds
C.Approximately 60 seconds
D.Approximately 10 seconds
AnswerA

Firehose flushes when either the 5 MB buffer fills or 60 seconds elapses. At 2 MB/s the size threshold is reached first, after roughly 2.5 seconds, so writes occur far more frequently than the configured interval.

Why this answer

Amazon Kinesis Data Firehose writes to S3 when either the buffer interval (60 seconds) or buffer size (5 MB) is reached first. With data arriving at 2 MB/s, the 5 MB buffer fills in 2.5 seconds (5 MB / 2 MB/s), triggering a write before the 60-second interval expires. Thus, the expected time between S3 writes is approximately 2.5 seconds.

Exam trap

The trap here is that candidates assume the buffer interval (60 seconds) is the primary determinant of write frequency, ignoring that the buffer size threshold triggers writes much earlier when data arrival rates are high.

How to eliminate wrong answers

Option B is wrong because 30 seconds would imply a buffer fill rate of ~0.167 MB/s, which does not match the given 2 MB/s arrival rate. Option C is wrong because 60 seconds is the buffer interval, but the buffer size threshold is reached much sooner at 2.5 seconds, making the interval the active trigger only if data arrival is slower. Option D is wrong because 10 seconds would correspond to a buffer size of 20 MB (2 MB/s * 10 s), which is not the configured 5 MB buffer size.

77
MCQhard

A healthcare company is building a data pipeline to ingest electronic health records (EHR) from hospitals. The data is sent as JSON files via SFTP to an on-premises server. The company wants to move this data to AWS using AWS Transfer Family (SFTP) and then process it with AWS Glue. Data sovereignty regulations require that all data remain within the EU (Frankfurt) region. The pipeline must detect when a new file arrives and start the Glue job automatically. The engineer has set up an AWS Transfer Family server in Frankfurt, and files are uploaded to an S3 bucket in the same region. However, the Glue job is not triggering automatically. The engineer needs to implement automated triggering. What should the engineer do?

A.Configure AWS Step Functions to poll the S3 bucket every minute and start the Glue job if new files exist.
B.Configure Amazon CloudWatch Events to trigger the Glue job on a schedule that checks for new files.
C.Use Amazon Simple Queue Service (SQS) to queue file metadata and have a Lambda function poll the queue to start the Glue job.
D.Set up an S3 event notification on the bucket to invoke an AWS Lambda function that starts the Glue job.
AnswerD

S3 event notifications fire on ObjectCreated events, invoking a Lambda function that calls StartJobRun on the Glue job. This provides the event-driven triggering the pipeline lacks, and keeps processing within the Frankfurt region, satisfying the EU data sovereignty constraint.

Why this answer

S3 event notifications can be configured to invoke an AWS Lambda function when objects are created, and the Lambda function can call the Glue StartJobRun API to launch the job. This provides event-driven, near-real-time triggering without polling, and keeps all processing within the Frankfurt region to satisfy data sovereignty requirements.

Exam trap

DEA-C01 often tests the misconception that polling or scheduled checks are needed to trigger Glue jobs, when S3 event notifications with Lambda provide the native, event-driven solution.

How to eliminate wrong answers

Option A is wrong because polling S3 every minute with Step Functions is inefficient, adds latency, and incurs unnecessary cost compared to native event notifications. Option B is wrong because scheduled CloudWatch Events checks for new files rather than reacting to arrivals, introducing delay and wasted executions. Option C is wrong because SQS with Lambda polling adds unnecessary complexity — S3 event notifications can directly invoke Lambda without an intermediary queue for this use case.

78
Multi-Selecteasy

A data engineer needs to enforce that all data in an Amazon S3 bucket is encrypted at rest. Which of the following can be used to achieve this? (Choose TWO.)

Select 2 answers
A.Use AWS CloudTrail to monitor for unencrypted objects
B.Use VPC endpoints to restrict access
C.Configure a bucket policy to deny PutObject if encryption headers are missing
D.Enable default encryption on the S3 bucket using SSE-S3
E.Use AWS KMS to generate encryption keys for the bucket
AnswersC, D

A bucket policy with a Deny effect on s3:PutObject conditioned on the absence of encryption headers (for example, s3:x-amz-server-side-encryption) rejects unencrypted uploads at the API layer, satisfying the requirement to enforce encryption at rest for every object written to the bucket.

Why this answer

Option C is correct because an S3 bucket policy can include a Deny statement on s3:PutObject that uses a condition such as StringNotEquals on s3:x-amz-server-side-encryption (or its aws:kms variant), which blocks any upload request that does not carry the required encryption header, thereby enforcing encryption at rest for newly written objects. Option D is correct because enabling default bucket encryption with SSE-S3 (AES-256) causes Amazon S3 to automatically encrypt every object at rest on write, even when the request specifies no encryption headers, satisfying the requirement without relying on the client. Option A is not correct because AWS CloudTrail only records and logs API activity for auditing; it cannot prevent or enforce encryption of objects.

Option B is not correct because VPC endpoints only control network access paths to S3 and have no bearing on whether objects are encrypted at rest. Option E is not correct because AWS KMS generates and manages keys, but merely having KMS keys available does not by itself enforce encryption on the bucket; enforcement requires default encryption or a bucket policy condition.

79
MCQeasy

A data engineer needs to store large volumes of infrequently accessed compliance data in Amazon S3 for 10 years. The data must be retrievable within 12 hours if required for audits. The engineer wants the most cost-effective storage solution. Which S3 storage class should be used?

A.S3 Glacier Instant Retrieval
B.S3 Standard-Infrequent Access (S3 Standard-IA)
C.S3 Glacier Deep Archive
D.S3 Glacier Flexible Retrieval
AnswerC

S3 Glacier Deep Archive is the lowest-cost storage class in Amazon S3, designed for long-term retention of data that is rarely accessed. It provides standard retrieval within 12 hours, which meets the audit requirement. For compliance data retained for 10 years with infrequent access, it offers the most cost-effective solution.

Why this answer

The scenario requires long-term retention (10 years) with infrequent access and retrieval within 12 hours. S3 Glacier Deep Archive is specifically designed for this use case, offering the lowest storage cost among S3 classes while providing standard retrieval within 12 hours. It is the most cost-effective choice for compliance data that is rarely accessed.

Exam trap

The trap here is assuming that any Glacier class would be equally cost-effective, but S3 Glacier Deep Archive is the cheapest and still meets the 12-hour retrieval requirement.

80
MCQmedium

A company ingests streaming data into Amazon Kinesis Data Streams. Producers write records using the PutRecords API with explicit partition keys based on customer ID. A data engineer observes that a few shards are consistently at 100 percent write throughput while others are underutilized, causing throttling. Which action should the engineer take to distribute the load more evenly?

A.Change the partition key to include a random suffix or use a higher-cardinality key.
B.Enable enhanced fan-out on the stream to give consumers dedicated throughput.
C.Switch the producers to use the PutRecord API instead of PutRecords.
D.Increase the number of shards in the stream to match the number of partition keys.
AnswerA

Kinesis Data Streams maps records to shards by hashing the partition key. If a small number of customer IDs generate most of the traffic, those keys hash to a few shards and create hot spots. Adding a random suffix or using a higher-cardinality key spreads records across more shards, balancing write throughput and eliminating throttling on the hot shards.

Why this answer

Hot shards occur when partition keys are skewed, causing the hash function to map most traffic to a few shards. Including a random suffix or using a higher-cardinality key distributes records more evenly across all shards. Adding shards or changing the write API does not alter the key-to-shard mapping, and enhanced fan-out only affects consumer reads, not producer write distribution.

Exam trap

The trap here is assuming that adding shards or enabling enhanced fan-out will fix write throttling, when the actual cause is a skewed partition key that keeps sending traffic to the same shards.

81
MCQmedium

A data engineer is using Amazon Managed Workflows for Apache Airflow (Amazon MWAA) to orchestrate a data pipeline. The engineer needs to ensure that the Airflow environment can access an Amazon S3 bucket to read and write data. The S3 bucket is in the same AWS account and Region as the MWAA environment. Which configuration is required to allow MWAA to access the S3 bucket?

A.Attach an IAM policy to the MWAA execution role that grants the necessary S3 permissions.
B.Configure the S3 bucket policy to allow the MWAA service principal.
C.Enable public access on the S3 bucket and use the S3 REST API with presigned URLs.
D.Store S3 credentials in AWS Secrets Manager and reference them in the Airflow connection.
AnswerA

Amazon MWAA uses an execution role that grants the Airflow environment permissions to access AWS services. To allow MWAA to read and write to S3, you must attach an IAM policy to this execution role with the appropriate s3:GetObject, s3:PutObject, and s3:ListBucket permissions. This is the standard method for granting MWAA access to S3.

Why this answer

MWAA environments use an execution role that is assumed by the Airflow components. To grant access to S3, you attach an IAM policy to this role with the necessary S3 permissions. This is the standard and recommended approach.

The execution role must also have a trust policy that allows MWAA to assume it, but that is created automatically when you create the environment.

Exam trap

The trap here is thinking that a bucket policy or Secrets Manager is needed, when the MWAA execution role's IAM policy is the primary mechanism.

82
MCQmedium

A data engineer is responsible for an Amazon Redshift cluster that ingests data continuously from Amazon Kinesis Data Streams. The engineer needs to ensure that the raw streaming data is immediately queryable in Redshift with minimal latency. Which approach should the engineer take?

A.Use the Amazon Kinesis Data Firehose delivery stream to load data into Redshift.
B.Use Amazon Redshift Spectrum to query the Kinesis stream directly.
C.Use Amazon Redshift federated query to query Kinesis Data Streams.
D.Use AWS Database Migration Service (AWS DMS) to replicate from Kinesis to Redshift.
AnswerA

Kinesis Data Firehose can deliver streaming data directly to Amazon Redshift, using an intermediate S3 bucket and issuing COPY commands. This provides near-real-time ingestion with minimal latency, making the data immediately queryable. Firehose handles buffering, compression, and retry logic, simplifying the pipeline.

Why this answer

Kinesis Data Firehose is the managed service for loading streaming data into Amazon Redshift. It buffers, compresses, and delivers data to an intermediate S3 bucket, then issues COPY commands to load into Redshift. This provides low-latency, near-real-time ingestion, making the data immediately available for querying.

Other options either do not support Kinesis as a source or are not optimized for this scenario.

Exam trap

The trap here is assuming that Redshift Spectrum or federated query can directly access streaming data, but they are designed for querying external data in S3 or relational databases, not Kinesis streams.

83
MCQeasy

A company needs to ingest data from multiple on-premises databases into Amazon S3 for analytics. The databases include Oracle, MySQL, and PostgreSQL. The data must be continuously replicated with minimal latency. Which AWS service should be used?

A.AWS Database Migration Service (AWS DMS)
B.Amazon Kinesis Data Streams
C.AWS Snowball
D.AWS Glue
AnswerA

DMS can continuously replicate from multiple source databases to S3.

Why this answer

AWS DMS supports continuous replication (change data capture, CDC) from Oracle, MySQL, and PostgreSQL to S3 as a target, enabling near-real-time data ingestion with minimal latency. It handles schema conversion and can replicate ongoing changes without interrupting source databases, making it the correct choice for this use case.

Exam trap

The trap here is that candidates may confuse Kinesis Data Streams as a general-purpose ingestion service for databases, but it lacks native CDC connectors for relational databases and is optimized for streaming data from applications, not for replicating transactional changes from databases to S3.

How to eliminate wrong answers

Option B is wrong because Amazon Kinesis Data Streams is a real-time streaming service for ingesting high-throughput data from applications or devices, not designed for continuous database replication with CDC from relational databases. Option C is wrong because AWS Snowball is a physical data transfer device for offline, bulk data migration, not suitable for continuous, low-latency replication. Option D is wrong because AWS Glue is a serverless ETL service primarily for batch data transformation and cataloging, not for continuous replication with minimal latency from live databases.

84
MCQeasy

A data engineer is configuring an Amazon S3 lifecycle policy to transition objects to S3 Glacier Deep Archive after 90 days. The bucket receives new objects daily. The engineer wants to ensure that objects are not deleted before 90 days. Which lifecycle action should be used?

A.Expiration
B.Transition
C.NoncurrentVersionTransition
D.AbortIncompleteMultipartUpload
AnswerB

Transition moves objects to another storage class without deleting them, so objects remain retrievable in Glacier Deep Archive after 90 days. Expiration would delete them, and the engineer explicitly wants no deletion before 90 days, making Transition the appropriate lifecycle action.

Why this answer

(Transition) is correct because the S3 Lifecycle Transition action moves objects between storage classes over time. To ensure objects are moved to S3 Glacier Deep Archive after 90 days without deletion, a Transition rule is configured to specify the target storage class and the number of days from object creation.

Exam trap

The trap here is confusing Expiration (which deletes objects) with Transition (which moves objects to another storage class), leading candidates to select Expiration when the goal is to retain objects for a minimum period before moving them to archival storage.

How to eliminate wrong answers

Option A (Expiration) is wrong because it permanently deletes objects after a specified number of days, which would remove them before they could be transitioned to Glacier Deep Archive. Option C (NoncurrentVersionTransition) is wrong because it applies only to noncurrent versions of versioned objects, not to current objects in a non-versioned or versioned bucket. Option D (AbortIncompleteMultipartUpload) is wrong because it only aborts incomplete multipart uploads after a specified number of days, not transitioning or deleting complete objects.

85
MCQeasy

A company wants to securely store database credentials used by a Lambda function. Which AWS service should be used to store and rotate the credentials automatically?

A.AWS CloudHSM
B.AWS Secrets Manager
C.AWS Key Management Service (KMS)
D.AWS Systems Manager Parameter Store
AnswerB

Secrets Manager natively stores credentials and performs scheduled rotation via Lambda, satisfying the automatic rotation requirement. Unlike Systems Manager Parameter Store, it includes built-in rotation for RDS, Redshift and DocumentDB, so no custom rotation logic is needed.

Why this answer

AWS Secrets Manager is purpose-built for storing, retrieving, and automatically rotating secrets such as database credentials, API keys, and OAuth tokens. It natively integrates with Amazon RDS, Redshift, and DocumentDB to rotate credentials on a schedule using Lambda rotation functions. Lambda functions can retrieve secrets at runtime via the Secrets Manager API or the AWS Parameters and Secrets Lambda Extension, avoiding hardcoded credentials.

Exam trap

DEA-C01 often tests the confusion between Secrets Manager and SSM Parameter Store — candidates pick Parameter Store because it is cheaper, forgetting that only Secrets Manager provides native automatic rotation for database credentials.

How to eliminate wrong answers

Option A is wrong because AWS CloudHSM is a dedicated hardware security module for cryptographic key operations and PKI, not a secrets store with built-in rotation for database credentials. Option C is wrong because KMS manages encryption keys used to encrypt data — it does not store or rotate application secrets like database passwords. Option D is wrong because Systems Manager Parameter Store can store SecureString parameters, but it lacks native automatic rotation for RDS/database credentials (rotation must be custom-built with Lambda and EventBridge).

86
MCQmedium

A data engineer has an AWS Glue ETL job that processes JSON files from Amazon S3. The job currently uses the DynamicFrame method to write output to Amazon Redshift. The engineer needs to improve write performance by using a staging Amazon S3 bucket and parallel COPY operations. Which AWS Glue connection option should the engineer configure?

A.Increase the number of Glue DPUs allocated to the job to enable parallel writes.
B.Set the 'useConnectionProperties' flag to true in the Glue job script.
C.Enable the 'redshiftTmpDir' connection property to specify an Amazon S3 temporary directory.
D.Configure the job to write to Amazon Redshift Spectrum using an external schema.
AnswerC

Setting the redshiftTmpDir connection property directs AWS Glue to stage data in Amazon S3 before issuing a COPY command into Amazon Redshift. This enables parallel loading and significantly improves write performance for large datasets. Without this property, Glue may fall back to slower JDBC-based inserts, which are not optimized for bulk transfers.

Why this answer

The redshiftTmpDir connection property tells AWS Glue to stage data in Amazon S3 and then run a COPY command into Redshift, which supports parallel loading and is far faster than JDBC inserts. The other options either reference non-existent properties or confuse read-side services like Redshift Spectrum with write-side optimizations.

Exam trap

The trap here is assuming that adding more DPUs or enabling a generic flag will automatically optimize Redshift writes, when the key is the staging directory property.

87
MCQmedium

A company uses AWS Glue to run ETL jobs on data stored in S3. The data is encrypted with SSE-KMS. The Glue job fails with an 'AccessDenied' error when trying to read the data. What is the MOST likely cause?

A.The S3 bucket policy denies access to the Glue service role.
B.The AWS Glue Data Catalog does not have permission to the table.
C.The IAM role used by Glue does not have kms:Decrypt permission on the KMS key.
D.The Glue job's connection does not have the necessary permissions.
AnswerC

SSE-KMS encrypts objects with a KMS key, so reading them requires both S3 access and kms:Decrypt on that key. The Glue job's IAM role lacks this permission, producing AccessDenied despite valid S3 permissions. Granting kms:Decrypt on the key to the Glue role resolves the failure.

Why this answer

When S3 data is encrypted with SSE-KMS, any principal reading the object must have both s3:GetObject on the bucket and kms:Decrypt on the KMS key. The Glue job's IAM role is the principal making the read, so if it lacks kms:Decrypt, S3 returns AccessDenied even though the bucket policy may allow access. This is the most common cause of this specific error pattern.

Exam trap

DEA-C01 often tests the layered permission model of SSE-KMS, so candidates who focus only on S3 bucket policies or IAM S3 actions miss that kms:Decrypt is a separate, required permission on the KMS key.

How to eliminate wrong answers

Option A is wrong because a bucket policy denying the Glue role would also produce AccessDenied, but the question specifies SSE-KMS encryption, which points to the KMS key policy as the more likely missing permission. Option B is wrong because Data Catalog permissions affect metadata access, not the ability to read encrypted S3 objects. Option D is wrong because Glue connections are used for JDBC/network sources, not for S3 reads, so connection permissions are irrelevant here.

88
MCQhard

Refer to the exhibit. A data engineer runs a Glue job manually and receives a ThrottlingException. The engineer checks the job run history and sees a previous failure with the same error. What is the MOST likely cause of the throttling, and which solution is MOST appropriate?

A.Increase the number of DPUs for the job to reduce runtime.
B.Implement retry logic with exponential backoff in the script that calls start-job-run.
C.Use AWS Glue reserved capacity to guarantee API throughput.
D.Delete old job runs to reduce the number of entries in the job run history.
AnswerB

ThrottlingException arises when start-job-run calls exceed the Glue API request rate. Retry logic with exponential backoff spaces repeated calls progressively, allowing the request rate to fall within limits and the job to start successfully without manual intervention.

Why this answer

A ThrottlingException from AWS Glue typically occurs when the StartJobRun API is called too frequently, exceeding the service's request rate limits. Implementing retry logic with exponential backoff in the calling script is the standard AWS-recommended solution to handle transient throttling gracefully. This allows the job to start successfully after a short delay without manual intervention.

Exam trap

DEA-C01 often tests the confusion between performance tuning (DPUs) and API throttling, leading candidates to choose DPU increases when the real issue is request rate limiting.

How to eliminate wrong answers

Option A is wrong because increasing DPUs addresses job performance and runtime, not API throttling on StartJobRun. Option C is wrong because AWS Glue does not offer reserved capacity to guarantee API throughput; Glue is serverless and throttling is managed by AWS service quotas. Option D is wrong because deleting old job runs reduces history storage but does not affect the API request rate that causes throttling.

89
MCQmedium

A data engineer maintains an AWS Glue ETL job that processes millions of small JSON files stored in Amazon S3. The job's runtime has increased significantly, and CloudWatch logs show many small executor tasks and frequent garbage collection. The engineer wants to improve job performance by reducing the number of small files processed per task. Which action should the engineer take?

A.Convert the source files to Parquet format using an AWS Glue crawler before running the ETL job.
B.Use the Glue ETL job's 'groupFiles' option to group multiple small files into a single partition.
C.Enable job bookmarks to track previously processed files and skip them on subsequent runs.
D.Increase the number of DPUs allocated to the job to provide more memory per executor.
AnswerB

The 'groupFiles' option in AWS Glue ETL allows grouping multiple small files into a single partition, reducing the number of tasks and improving read efficiency. This directly addresses the issue of many small files causing overhead. By setting groupFiles to 'inPartition' or 'inPartitionAcrossBuckets', the job can process larger chunks of data per task, reducing garbage collection and improving performance.

Why this answer

Grouping small files into larger partitions reduces the number of tasks and the overhead of task scheduling and garbage collection. This is a specific optimization for AWS Glue ETL jobs that read many small files. Enabling job bookmarks or increasing DPUs does not address the root cause.

Converting formats via a crawler is not a valid approach because crawlers do not transform data.

Exam trap

The trap here is assuming that adding more resources or enabling bookmarks will fix small-file inefficiency, when the real solution is to group files.

90
MCQeasy

A data engineer needs to schedule an AWS Glue extract, transform, and load job to run every day at 02:00 UTC and trigger a dependent Amazon Redshift stored procedure only after the Glue job succeeds. The engineer wants a managed orchestration option that avoids provisioning servers. Which approach should the engineer use?

A.Run an Amazon EC2 instance with a cron entry that calls the Glue StartJobRun API and then the Redshift API
B.Configure an Amazon EventBridge schedule rule that invokes the Glue job and rely on the job to call Redshift on completion
C.Use AWS Step Functions Express Workflows with a cron expression to call Glue and then Redshift
D.Create an AWS Glue workflow with a daily trigger that starts the job and a conditional trigger that fires the Redshift step on success
AnswerD

AWS Glue workflows provide serverless orchestration with schedule triggers for the daily 02:00 UTC start and conditional triggers that fire downstream actions only when the job reaches a SUCCEEDED state. This satisfies the dependency requirement without managing any servers, and it keeps orchestration within the Glue service already used for the job.

Why this answer

Glue workflows are serverless and purpose-built for chaining Glue jobs and dependent actions. A schedule trigger handles the daily 02:00 UTC start, and a conditional trigger keyed to the job's success state ensures the Redshift procedure runs only after the job completes successfully, with no servers to manage.

Exam trap

The trap here is assuming any scheduler that can start a Glue job also enforces success-based dependencies, when only a conditional trigger provides that guarantee natively.

91
Multi-Selectmedium

A data engineer is troubleshooting a slow-running Amazon Athena query on a large dataset stored in S3. The query scans many small files. Which TWO actions can improve query performance?

Select 2 answers
A.Increase the number of files to increase parallelism
B.Disable S3 server-side encryption
C.Concatenate small files into larger files
D.Partition the data by a frequently filtered column
E.Convert files from CSV to JSON
AnswersC, D

Merging many small files reduces per-file overhead: Athena must open, list and read metadata for each object, so fewer larger files cut S3 request costs and task scheduling overhead. This directly addresses the small-file scanning inefficiency described in the stem.

Why this answer

Option C is correct because Athena performance is heavily degraded by many small files, since each file incurs overhead for opening, listing, and reading metadata; concatenating small files into larger files (typically 128 MB or more) reduces this per-file overhead and lets Athena scan data more efficiently. Option D is correct because partitioning the data by a frequently filtered column allows Athena to use partition pruning, so it reads only the relevant S3 prefixes instead of scanning the entire dataset, dramatically reducing the amount of data scanned and query time. Option A is incorrect because adding more small files increases metadata and open/close overhead rather than improving performance, even though parallelism is a factor.

Option B is incorrect because S3 server-side encryption is transparent to Athena and disabling it does not affect query performance. Option E is incorrect because converting CSV to JSON does not inherently reduce scanned data or file count and JSON is typically more verbose, so it would not improve performance.

92
MCQeasy

A company needs to store application log files for 90 days for compliance. The logs are generated continuously and are rarely accessed after 30 days. The data engineer must minimize storage costs. Which storage solution should the engineer choose?

A.Amazon CloudWatch Logs with a retention policy of 90 days
B.Amazon S3 Glacier Deep Archive
C.Amazon EBS gp3 volumes attached to an EC2 instance
D.Amazon S3 Standard with a lifecycle policy to transition to S3 Standard-IA after 30 days and expire after 90 days
AnswerD

S3 Standard with a lifecycle policy transitioning to Standard-IA at 30 days and expiring at 90 days matches the access pattern: frequent early access, rare later access, and deletion at the compliance boundary. This minimises cost while meeting the 90-day retention requirement.

Why this answer

Amazon S3 Standard with a lifecycle policy to transition to S3 Standard-IA after 30 days and expire after 90 days is correct because it aligns with the access pattern: logs are frequently accessed only in the first 30 days, then rarely accessed for the remaining 60 days. S3 Standard-IA offers lower storage costs for infrequently accessed data while still providing millisecond retrieval, and the lifecycle policy automates the transition and eventual deletion, minimizing costs without sacrificing availability.

Exam trap

The trap here is that candidates often choose CloudWatch Logs (Option A) because it is a familiar logging service, but they overlook that its cost model (per GB ingested, per GB stored, and per GB archived) can be significantly higher than S3 for long-term retention of large log volumes, and it lacks the automated tiering to lower-cost storage classes.

How to eliminate wrong answers

Option A is wrong because Amazon CloudWatch Logs is designed for real-time monitoring and log ingestion, not for long-term, cost-optimized archival storage; its retention policy only controls deletion, not tiered storage transitions, and costs can be higher than S3 for large volumes of rarely accessed logs. Option B is wrong because S3 Glacier Deep Archive is intended for data that is accessed at most once or twice a year and has retrieval times of 12 hours or more, making it unsuitable for logs that may need occasional access within 90 days; it also incurs minimum storage charges that make it cost-ineffective for short retention periods. Option C is wrong because EBS gp3 volumes attached to an EC2 instance incur compute costs even when idle, and managing log storage on block storage requires manual lifecycle management, leading to higher operational overhead and cost compared to a fully managed object storage solution.

93
MCQmedium

A company is using Amazon Kinesis Data Streams with a Lambda consumer to process clickstream data. The data rate is high and the Lambda function is falling behind, resulting in increased processing latency. What is the MOST effective way to improve throughput?

A.Increase the memory allocated to the Lambda function.
B.Increase the Lambda function timeout.
C.Use Kinesis Data Firehose instead of Lambda.
D.Increase the number of shards in the Kinesis stream.
AnswerD

Adding shards raises the stream's parallel capacity, since each shard supports one Lambda invocation per batch and caps ingestion at 1 MB/s or 1,000 records/s. With the Lambda consumer throttled by shard-level concurrency, horizontal scaling of the stream directly relieves the backlog causing the latency.

Why this answer

Increasing the number of shards in the Kinesis stream increases the stream's total read capacity, allowing more concurrent Lambda invocations to process records in parallel. Since each shard supports up to 5 read transactions per second and a maximum of 2 MB/s read throughput, adding shards directly raises the aggregate throughput, enabling the Lambda consumer to keep up with the high data rate.

Exam trap

The DEA-C01 exam often tests the misconception that Lambda performance tuning (memory/timeout) is the primary solution for stream processing backpressure, when in fact the shard count is the fundamental parallelism bottleneck in Kinesis Data Streams with a Lambda consumer.

How to eliminate wrong answers

Option A is wrong because increasing Lambda memory also increases CPU allocation, which can speed up individual function execution, but it does not address the bottleneck of limited shard-level parallelism; the function is falling behind due to insufficient concurrent processing capacity, not per-invocation performance. Option B is wrong because increasing the Lambda timeout only allows the function to run longer before being terminated, but it does not improve throughput; if the function is already timing out, extending the timeout may mask the issue but does not increase the rate at which records are consumed. Option C is wrong because Kinesis Data Firehose is a delivery stream that buffers and loads data to destinations like S3 or Redshift; it does not support real-time per-record processing with custom logic like Lambda, and switching to Firehose would lose the ability to transform or react to each record individually, which is likely required for clickstream processing.

94
Multi-Selecteasy

A company wants to enforce encryption in transit for data moving between an EC2 instance and an S3 bucket. Which TWO methods can achieve this? (Choose 2)

Select 2 answers
A.Add a bucket policy that denies requests without the aws:SecureTransport condition.
B.Use a VPC endpoint for S3.
C.Enable default SSE-S3 encryption on the bucket.
D.Use the HTTPS endpoint for S3 API calls.
E.Enable CloudTrail to monitor for non-encrypted requests.
AnswersA, D

The aws:SecureTransport condition key evaluates whether the request arrived over TLS; denying when it is false blocks plain HTTP calls to S3. This enforces encryption in transit at the bucket policy layer, satisfying the requirement that EC2-to-S3 traffic never travel unencrypted.

Why this answer

Option A is correct because a bucket policy with a Deny effect on s3:* conditioned on "aws:SecureTransport": "false" explicitly rejects any request made over plain HTTP, thereby enforcing encryption in transit for all access to the bucket, including from EC2. Option D is correct because S3's HTTPS endpoint (https://bucket.s3.amazonaws.com or the regional equivalent) uses TLS to encrypt data in transit between the EC2 instance and S3, satisfying the encryption-in-transit requirement directly. Option B is not correct because a VPC endpoint (Gateway or Interface) only changes the network path and can be used with either HTTP or HTTPS; it does not by itself guarantee encryption in transit.

Option C is not correct because SSE-S3 provides encryption at rest for objects stored in the bucket, not encryption of data moving over the network. Option E is not correct because CloudTrail only logs and monitors API activity; it is a detective control that does not encrypt traffic or prevent unencrypted requests.

Exam trap

DEA-C01 often tests the confusion between encryption at rest (SSE-S3) and encryption in transit (HTTPS/TLS), causing candidates to select SSE-S3 or CloudTrail as methods for enforcing in-transit encryption.

95
Multi-Selectmedium

Which TWO actions are recommended for securing data at rest in Amazon S3? (Choose two.)

Select 2 answers
A.Enable default encryption on the S3 bucket using SSE-S3 or SSE-KMS.
B.Use S3 Bucket Key to reduce KMS request costs.
C.Enable S3 Versioning to protect against accidental deletions.
D.Apply a bucket policy that denies PutObject requests without the x-amz-server-side-encryption header.
E.Configure cross-region replication to replicate data to another bucket.
AnswersA, D

Enabling default encryption ensures every object written to the bucket is encrypted at rest automatically, satisfying the data-at-rest requirement without relying on per-request headers. SSE-S3 provides AES-256 with Amazon-managed keys, while SSE-KMS adds customer-managed key control and audit trails through CloudTrail, meeting compliance needs.

Why this answer

Option A is correct because enabling default bucket encryption with SSE-S3 or SSE-KMS ensures that every object written to the bucket is automatically encrypted at rest, satisfying the core requirement for data-at-rest protection in Amazon S3. Option D is correct because a bucket policy that denies PutObject requests lacking the x-amz-server-side-encryption header enforces encryption on upload, preventing unencrypted objects from being stored even if default encryption is bypassed or misconfigured. Option B is not a security control; S3 Bucket Key reduces AWS KMS request costs and CloudTrail event volume but does not itself secure data at rest.

Option C addresses availability and data durability against accidental deletion through versioning, not encryption at rest. Option E provides geographic redundancy and disaster recovery via cross-region replication, but replication does not encrypt data at rest by itself.

Exam trap

The trap here is that candidates often confuse data protection features like Versioning or replication with encryption controls, but the question specifically asks for securing data at rest, which requires encryption mechanisms such as default encryption or policy-enforced encryption headers.

96
MCQmedium

A data engineer needs to set up a data catalog for a new data lake in AWS Glue. The data resides in S3 in Parquet format. The engineer wants to ensure that the schema is automatically detected and updated when new columns are added to the data. Which configuration should the engineer use?

A.Add a partition index to the Glue Data Catalog table.
B.Configure the crawler's 'Schema updates' option to 'Update the table schema'.
C.Set the crawler's 'Database' output to a new database.
D.Enable partition indexing on the table.
AnswerB

The Glue crawler's schema-update behaviour controls whether re-crawls revise the Data Catalog table definition. Setting it to update the table schema lets the crawler add newly detected columns automatically, satisfying the requirement that schema changes in the Parquet data be picked up without manual intervention.

Why this answer

The Glue crawler's 'Schema updates' option controls what happens when the crawler detects schema changes in the data. Setting it to 'Update the table schema' allows the crawler to add new columns and modify the existing table definition in the Data Catalog when new columns appear in the Parquet files. This is the correct configuration for automatic schema evolution.

Exam trap

DEA-C01 often tests the confusion between crawler schema update behaviour and partition indexing, tricking candidates into picking an optimization feature when the question is about schema evolution.

How to eliminate wrong answers

Option A is wrong because a partition index improves query performance on partitioned tables but does not handle schema detection or updates. Option C is wrong because setting the crawler's database output only determines where the table metadata is stored, not how schema changes are handled. Option D is wrong because partition indexing, like option A, is a query optimization feature and has no role in schema evolution.

97
MCQmedium

A data engineer manages an AWS Glue ETL job that processes millions of small JSON files in Amazon S3. The job is slow and often fails with an OutOfMemory error on the driver. The engineer wants to improve performance without changing the output format. Which solution should the engineer implement?

A.Enable job bookmarks to track processed files and skip already-processed data.
B.Use AWS Glue's groupFiles and groupSize options to combine small files into larger groups.
C.Convert the JSON files to Parquet format using an AWS Glue crawler before running the ETL job.
D.Increase the number of DPUs allocated to the job to provide more memory per executor.
AnswerB

The groupFiles and groupSize options in AWS Glue allow the job to coalesce many small files into larger groups before processing. This reduces the number of input partitions and the memory overhead on the driver, directly addressing the OutOfMemory error and improving performance. It is the recommended approach for large numbers of small files without changing the output format.

Why this answer

Grouping small files with groupFiles and groupSize is the most effective way to handle millions of small files in AWS Glue. It reduces the number of input splits, lowers driver memory pressure, and improves overall job performance. Other options either do not address the driver memory issue or require changing the output format.

Exam trap

The trap here is assuming that adding more DPUs will always resolve OutOfMemory errors, but driver memory is not proportional to the number of DPUs.

98
MCQhard

A data team runs a daily AWS Glue ETL job that processes data from an Amazon Redshift cluster and writes results to Amazon S3. The job completes successfully but takes 2 hours longer than expected. The job uses the JDBC connection to Redshift. The Redshift cluster is 4 dc2.large nodes. The Glue job has 10 workers of type G.1X. Which change would MOST likely reduce the job duration?

A.Use Redshift Spectrum to query data directly from S3
B.Use the S3 staging option in the Glue connection to unload data from Redshift to S3 first
C.Increase the Redshift cluster size to 8 nodes
D.Increase the number of Glue workers to 20
AnswerB

The JDBC connector reads Redshift row-by-row through the driver, which is slow at this scale. Using the S3 staging option runs a Redshift UNLOAD to S3 in parallel, then Glue reads S3 directly, removing the JDBC bottleneck and cutting job duration.

Why this answer

The JDBC connection in AWS Glue reads data row-by-row from Redshift, which is slow for large datasets. By enabling the S3 staging option in the Glue connection, the job uses Redshift's UNLOAD command to export data to S3 in parallel, then Glue reads from S3. This bypasses the JDBC bottleneck and leverages Redshift's massively parallel processing (MPP) to export data much faster.

Exam trap

The trap here is that candidates assume the bottleneck is either Redshift compute (C) or Glue parallelism (D), when in fact the JDBC driver's single-threaded row-by-row fetch is the primary performance limiter.

How to eliminate wrong answers

Option A is wrong because Redshift Spectrum queries data directly from S3, but the source data is in Redshift, not S3; Spectrum does not help extract data from Redshift. Option C is wrong because the bottleneck is the JDBC connection, not Redshift compute capacity; adding more Redshift nodes would not speed up a single-threaded JDBC read. Option D is wrong because increasing Glue workers only helps if the job is CPU-bound or parallelizable; the JDBC read is I/O-bound and limited by the single connection, so more workers would not reduce the 2-hour delay.

99
MCQeasy

A data engineer is responsible for monitoring an AWS Glue ETL job that runs daily. The job reads data from an Amazon S3 bucket and writes to an Amazon Redshift table. The engineer wants to receive an alert if the job fails or if it takes longer than expected to complete. Which AWS service should the engineer use to set up these alerts with the LEAST operational overhead?

A.Amazon CloudWatch Events (EventBridge) to trigger an AWS Lambda function that sends an email via Amazon SES.
B.AWS Glue job bookmarks to track job progress and send notifications on failure.
C.AWS CloudTrail to log AWS Glue API calls and analyze logs for failures.
D.Amazon CloudWatch Alarms based on AWS Glue job metrics, with Amazon SNS notifications.
AnswerD

CloudWatch Alarms can monitor AWS Glue job metrics such as glue.driver.aggregate.elapsedTime and glue.driver.aggregate.numFailedTasks. When a threshold is breached, the alarm can publish to an SNS topic, sending notifications. This requires minimal setup and no custom code, making it the least operational overhead solution for alerting on job failures and durations.

Why this answer

Amazon CloudWatch Alarms can monitor AWS Glue job metrics and trigger Amazon SNS notifications when thresholds are exceeded. This setup requires no custom code and provides timely alerts on job failures or prolonged execution, making it the most efficient solution with minimal operational overhead.

Exam trap

The trap here is overcomplicating the solution with Lambda and SES, when CloudWatch Alarms with SNS directly provide the needed alerting.

100
Multi-Selecteasy

A company is using AWS Glue to catalog data in Amazon S3. The data is stored in CSV format, but the schema is not consistent across all files. Which TWO actions can the company take to handle schema evolution and ensure the Glue Data Catalog is up to date? (Choose TWO.)

Select 2 answers
A.Configure the Glue crawler to update the table schema on each run.
B.Manually update the Glue Data Catalog tables whenever the schema changes.
C.Disable schema update in the crawler and add partitions manually.
D.Schedule the Glue crawler to run periodically to detect changes.
E.Require all data producers to use a single fixed schema.
AnswersA, D

Configuring the crawler to update the table schema on each run lets it revise column definitions in the Glue Data Catalog when it encounters differing CSV structures, directly handling schema evolution rather than leaving stale definitions in place.

Why this answer

Option A is correct because configuring the Glue crawler to update the table schema on each run allows it to detect new columns or changed data types in the CSV files and automatically revise the Data Catalog table definition, which is essential when schemas vary across files. Option D is correct because scheduling the crawler to run periodically ensures that any schema changes introduced by new or modified files are detected and reflected in the Glue Data Catalog in a timely, automated manner. Together, these two actions provide automated, recurring schema evolution handling.

Option B is not ideal because manual updates are error-prone and do not scale, and the question asks for actions the company can take to handle schema evolution automatically. Option C is incorrect because disabling schema update prevents the crawler from adapting to schema changes, and manual partition addition does not address evolving column structures. Option E is incorrect because enforcing a single fixed schema contradicts the scenario where schemas are already inconsistent and does not help update the catalog for existing variation.

Exam trap

DEA-C01 often tests the features of Glue crawlers for schema evolution, and candidates may overlook the need for both enabling schema update and scheduling the crawler.

101
Multi-Selecthard

A company uses AWS DMS to continuously replicate data from an on-premises SQL Server to Amazon Aurora MySQL. The replication lag is increasing. Which THREE actions can reduce the lag? (Choose three.)

Select 3 answers
A.Use parallel apply on the target endpoint.
B.Filter out unnecessary tables from replication.
C.Enable DMS validation.
D.Enable Multi-AZ for the DMS replication instance.
E.Increase the DMS replication instance size.
AnswersA, B, E

Parallel apply speeds up writes on the target.

102
MCQeasy

A data analyst needs to query a large Amazon S3 bucket containing CSV files using Amazon Athena. The bucket has millions of small files (less than 1 MB each). The analyst reports that queries are very slow and often time out. The data is partitioned by date and the partition columns are defined in the table. What is the most effective way to improve query performance?

A.Convert the files to Apache Parquet format using an AWS Glue ETL job.
B.Run a compaction job to consolidate small files into fewer larger files (e.g., 128 MB each).
C.Add more partitions by including hour and minute as partition keys.
D.Use S3 Select to push down filtering to S3 before Athena processes the data.
AnswerB

Millions of sub-1 MB files force Athena to open each object individually, so per-file overhead dominates and queries time out. Compacting into roughly 128 MB objects cuts the number of GET requests and partition listings, letting Athena scan far fewer, larger files.

Why this answer

Many small files (under 1 MB) cause high overhead because each file requires a separate read operation and metadata call. Consolidating them into fewer larger files (e.g., 128 MB each) reduces the number of read operations and improves I/O efficiency, directly addressing the root cause of slowdowns and timeouts. Option A (converting to Parquet) improves storage efficiency and query performance but does not reduce the file count; it is a beneficial addition but not the most effective standalone fix for the small file problem.

Option C (adding more partitions) would increase overhead by creating even more directories/files to scan. Option D (S3 Select) applies within individual files and does not mitigate overhead from file quantity.

103
MCQmedium

A company runs a data warehouse on Amazon Redshift. Queries are slow, and the team suspects data distribution is skewed. Which approach would best help identify distribution skew?

A.Check the STL_LOAD_ERRORS table for load failures
B.Query the SVV_TABLE_INFO table to see table size
C.Query the SVV_DISKUSAGE table to examine data distribution across slices
D.Review the WLM configuration in the parameter group
AnswerC

SVV_DISKUSAGE reports per-slice row counts and disk usage, directly exposing uneven distribution across slices that causes skew. Unlike planner-oriented views such as SVV_TABLE_INFO, it shows physical storage per slice, satisfying the need to confirm skew empirically before choosing a distribution key.

Why this answer

The SVV_DISKUSAGE table provides per-slice data distribution information, allowing you to identify skew by comparing the number of blocks allocated to each slice for a given table. In Amazon Redshift, data is distributed across slices based on the distribution key, and significant variation in block counts across slices indicates distribution skew, which can cause query performance degradation due to uneven workload distribution.

Exam trap

The trap here is that candidates confuse table-level metadata (SVV_TABLE_INFO) with slice-level distribution data (SVV_DISKUSAGE), assuming overall table size alone can reveal skew, when in fact only per-slice block counts expose uneven data distribution.

How to eliminate wrong answers

Option A is wrong because STL_LOAD_ERRORS records errors during COPY or INSERT operations, such as data type mismatches or malformed data, and has no relation to data distribution skew. Option B is wrong because SVV_TABLE_INFO shows overall table size, row count, and compression ratios, but it does not provide per-slice data distribution details needed to identify skew. Option D is wrong because WLM configuration in the parameter group manages query concurrency and memory allocation, not data distribution or skew detection.

104
MCQhard

An organization is using AWS Glue to process sensitive data. The data is stored in S3 with server-side encryption using AWS KMS (SSE-KMS). The Glue job fails with an error indicating that it cannot read the data. The IAM role used by Glue has the following policy. What is missing?

A.The s3:GetObject permission on the bucket
B.The kms:Decrypt permission on the KMS key
C.The kms:GenerateDataKey permission on the KMS key
D.The kms:ReEncrypt permission on the KMS key
AnswerB

Reading SSE-KMS encrypted S3 objects requires both S3 GetObject and kms:Decrypt on the customer managed key. The Glue execution role's policy grants S3 access but omits the KMS decrypt action, so the job cannot unwrap the data key and fails with an access-denied error.

Why this answer

The Glue job fails because the IAM role lacks the kms:Decrypt permission on the KMS key used for SSE-KMS encryption. When S3 objects are encrypted with SSE-KMS, any principal reading the object must have both s3:GetObject on the object and kms:Decrypt on the KMS key. Without kms:Decrypt, S3 returns an AccessDenied error even though the s3:GetObject permission is present.

Exam trap

DEA-C01 often tests the misconception that s3:GetObject alone is sufficient to read SSE-KMS encrypted objects, when in fact kms:Decrypt on the KMS key is also mandatory.

How to eliminate wrong answers

Option A is wrong because the question states the IAM role already has a policy (implied to include S3 read access); the missing piece is KMS decryption, not s3:GetObject. Option C is wrong because kms:GenerateDataKey is required for writing (encrypting) objects with SSE-KMS, not for reading them. Option D is wrong because kms:ReEncrypt is only needed when changing the encryption key or re-encrypting data, not for simple decryption during a read.

105
MCQmedium

A company uses Amazon Kinesis Data Streams to ingest real-time data. The compliance team requires that all data in the stream be encrypted at rest. Which configuration should be enabled?

A.Enable TLS encryption on the Kinesis stream
B.Enable server-side encryption using an AWS KMS key
C.Use client-side encryption in the producer application
D.Store the data in Amazon CloudWatch Logs instead
AnswerB

Server-side encryption with an AWS KMS key encrypts stream data at rest, meeting the compliance mandate. Kinesis encrypts using the specified customer managed key before writing to storage and decrypts on retrieval, so producers and consumers need no changes beyond KMS permissions.

Why this answer

Server-side encryption (SSE) for Amazon Kinesis Data Streams uses an AWS KMS key to automatically encrypt data at rest as it is written to the stream and decrypt it when read. This meets the compliance requirement for encryption at rest without requiring any changes to the producer or consumer applications.

Exam trap

The trap here is confusing encryption in transit (TLS) with encryption at rest (SSE), leading candidates to select TLS as the solution for at-rest compliance.

How to eliminate wrong answers

Option A is wrong because TLS encryption protects data in transit between clients and the Kinesis endpoint, not data at rest within the stream. Option C is wrong because client-side encryption encrypts data before it is sent to Kinesis, but this is a client-managed approach that does not leverage Kinesis's native at-rest encryption and adds complexity; the question asks which configuration should be enabled on the stream itself. Option D is wrong because storing data in CloudWatch Logs does not encrypt the Kinesis stream data at rest and is a different service entirely, not a configuration for Kinesis Data Streams.

106
MCQeasy

A data engineer needs to ingest JSON data from an on-premises relational database into Amazon S3 every hour. Which AWS service should be used to set up a scheduled, incremental data transfer?

A.Amazon S3 Transfer Acceleration with a cron job.
B.AWS Database Migration Service (DMS) with S3 as target.
C.AWS Glue with a JDBC connection and a scheduled crawler.
D.Amazon Kinesis Data Firehose with a database source.
AnswerB

AWS DMS performs ongoing, incremental replication from on-premises relational databases to Amazon S3, tracking changes via CDC rather than full reloads. Scheduled hourly tasks satisfy the stem's requirement for recurring incremental transfer, which SCT or DataSync alone cannot provide from a database.

Why this answer

AWS DMS is purpose-built for migrating databases to AWS targets, including Amazon S3. It supports ongoing replication (change data capture) and scheduled full-load tasks, making it ideal for hourly incremental transfers from an on-premises relational database to S3 without custom scripting.

Exam trap

The trap here is that candidates confuse AWS Glue's ETL capabilities with DMS's managed database migration, assuming Glue's JDBC connections can handle incremental transfers, but Glue lacks built-in change data capture and requires custom logic for scheduled incremental loads.

How to eliminate wrong answers

Option A is wrong because S3 Transfer Acceleration only speeds up uploads over long distances via edge locations; it does not provide scheduling, incremental data capture, or database connectivity. Option C is wrong because AWS Glue crawlers are designed for schema discovery and metadata cataloging, not for scheduled incremental data transfer from a database to S3; Glue ETL jobs can do this but require custom code, whereas DMS is the managed service for database migration. Option D is wrong because Kinesis Data Firehose ingests streaming data from producers like Kinesis streams or direct PUT, not from a relational database via JDBC; it lacks built-in change data capture for incremental database loads.

107
MCQhard

A data engineer is troubleshooting an Amazon DynamoDB table that has frequent throttling exceptions for write requests. The table has auto scaling enabled. What is the most likely cause?

A.The partition key is causing a hot partition
B.The table's read capacity is set too low
C.The table's auto scaling is disabled
D.The table is using global tables without conflict resolution
AnswerA

Auto scaling adjusts capacity for uniform load, but a hot partition concentrates writes on one partition key value, exceeding that partition's throughput ceiling regardless of table-level capacity. This is the most likely cause of persistent write throttling.

Why this answer

Auto scaling adjusts capacity based on utilization, but it cannot prevent throttling caused by a hot partition. If a single partition key value receives a disproportionate share of write traffic, that partition's throughput limit (3,000 WCU or 10 MB per partition) is exceeded, triggering ProvisionedThroughputExceededException. Auto scaling operates at the table level, not per partition, so it cannot resolve this imbalance.

Exam trap

The trap here is that candidates assume auto scaling automatically prevents all throttling, but it only adjusts table-level capacity and cannot fix uneven data access patterns like a hot partition.

How to eliminate wrong answers

Option B is wrong because write throttling is unrelated to read capacity; the question specifies write request throttling, so read capacity settings are irrelevant. Option C is wrong because the question states auto scaling is enabled, so this option describes a scenario that does not match the given condition. Option D is wrong because global tables with conflict resolution handle eventual consistency and replication conflicts, not throughput throttling on write requests.

108
Matchingmedium

Match each AWS data analytics service to its primary function.

Drag a concept onto its matching description — or click a concept then click the description.

Concepts
Matches

Serverless SQL query on S3

Business intelligence and dashboards

Data lake setup and access control

Real-time SQL on streaming data

Query data in S3 from Redshift

Why these pairings

The correct matches are: Amazon Athena for serverless SQL querying on S3, Amazon Redshift for data warehousing, Amazon EMR for big data processing, and Amazon Kinesis for real-time streaming. Common confusions include swapping the roles of Athena and Redshift.

109
MCQmedium

A company is running a data warehouse on Amazon Redshift. The data engineering team notices that query performance has degraded over time. They suspect that data distribution is causing excessive data movement between nodes. The table is joined frequently on the customer_id column. Which column should be chosen as the distribution key to optimize join performance?

A.AUTO distribution
B.customer_id
C.order_date
D.EVEN distribution
AnswerB

Choosing customer_id as the distribution key colocates rows with identical customer_id values on the same node, so frequent joins on that column execute locally without network redistribution. This directly addresses the excessive data movement degrading query performance.

Why this answer

(customer_id) because Redshift distributes data across nodes based on the distribution key. When two tables are joined on customer_id, using it as the distribution key ensures that matching rows from both tables are co-located on the same node, eliminating the need for data redistribution (broadcast or shuffle) during the join. This minimizes network traffic and reduces query latency, directly addressing the performance degradation caused by excessive data movement.

Exam trap

The trap here is that candidates may choose EVEN distribution (D) thinking it balances data evenly, but they overlook that it causes maximum data movement for joins, while AUTO distribution (A) seems safe but does not guarantee co-location for the specific join column.

How to eliminate wrong answers

Option A (AUTO distribution) is wrong because AUTO lets Redshift choose the distribution style based on table size and usage patterns, but it may not guarantee co-location for frequent joins on customer_id, potentially still causing data movement. Option C (order_date) is wrong because it is not the join column; using it as the distribution key would scatter customer_id values across nodes, forcing redistribution for every join on customer_id. Option D (EVEN distribution) is wrong because it distributes rows round-robin across nodes without considering join keys, which maximizes data movement during joins on customer_id and degrades performance.

110
Multi-Selecteasy

A company uses Kinesis Data Firehose to deliver streaming data to S3. They need to transform the data by adding a timestamp and removing sensitive fields. Which TWO approaches can achieve this?

Select 2 answers
A.Use Kinesis Data Analytics to transform the stream
B.Use S3 Select to transform data at rest
C.Use AWS Glue ETL to process data after delivery to S3
D.Use Amazon Redshift Spectrum to transform data
E.Configure a Lambda function as a data transformation in Firehose
AnswersC, E

Running AWS Glue ETL over the delivered S3 objects lets you add the timestamp and drop sensitive fields after Firehose lands the data. This satisfies the transformation requirement using a managed, serverless job rather than altering the delivery stream itself.

Why this answer

Option E is correct because Kinesis Data Firehose natively supports Lambda-based data transformation: you attach a Lambda function to the delivery stream, and Firehose invokes it synchronously on each record (or batch) before delivery, allowing the function to add a timestamp and strip sensitive fields. Option C is correct because AWS Glue ETL can process the data after it lands in S3, using Spark-based jobs to enrich records with timestamps and drop sensitive columns, which is a valid post-delivery transformation approach. Option A is not correct because Kinesis Data Analytics is for real-time SQL/Flink analytics on streams, not for modifying records delivered by Firehose.

Option B is not correct because S3 Select only filters and projects data at rest using SQL on individual objects; it cannot add timestamps or rewrite/remove fields. Option D is not correct because Redshift Spectrum queries data in S3 for analytics and does not transform or modify the delivered objects.

111
MCQhard

A data engineer created the IAM policy shown in the exhibit. The engineer then attempts to upload an object to 'my-bucket' using the AWS CLI with the command: aws s3 cp file.txt s3://my-bucket/ --sse aws:kms. The upload fails with an 'AccessDenied' error. What is the most likely cause?

A.The policy resource is incorrect
B.The policy requires SSE-S3 (AES256), but the command uses SSE-KMS
C.The policy does not allow the s3:PutObject action
D.The command is missing the --sse-customer-algorithm parameter
AnswerB

The policy's condition permits only AES256 (SSE-S3) encryption, but the CLI command requests aws:kms, so the encryption header mismatches the allowed value and S3 denies the upload. Aligning the command with SSE-S3 or amending the policy resolves it.

Why this answer

The IAM policy in the exhibit requires the `s3:x-amz-server-side-encryption` header to be set to `AES256`, which corresponds to SSE-S3. The AWS CLI command uses `--sse aws:kms`, which sets the header to `aws:kms` for SSE-KMS. This mismatch causes the request to fail the `s3:PutObject` condition check in the policy, resulting in an 'AccessDenied' error.

Exam trap

The trap here is that candidates may overlook the condition key in the policy and assume the error is due to a missing action or incorrect resource, rather than recognizing that the encryption header value must exactly match the policy's requirement.

How to eliminate wrong answers

Option A is wrong because the policy resource `arn:aws:s3:::my-bucket/*` correctly specifies the bucket and its objects, so the resource is not the issue. Option B is wrong because the policy explicitly requires SSE-S3 (AES256), but the command uses SSE-KMS, which is the direct cause of the failure. Option C is wrong because the policy does allow `s3:PutObject` via the `Effect: Allow` statement; the failure is due to the condition key mismatch, not a missing action.

Option D is wrong because `--sse-customer-algorithm` is used for SSE-C, not SSE-KMS or SSE-S3, and the command already specifies `--sse aws:kms` correctly for SSE-KMS.

112
MCQmedium

A data engineer needs to audit all access to an Amazon S3 bucket containing sensitive data. The audit must capture who accessed the bucket, from which IP address, and what actions were performed. Which AWS service should be enabled?

A.Enable S3 server access logging for the bucket.
B.Enable AWS CloudTrail with data events for the S3 bucket.
C.Use AWS Config to record S3 bucket-level changes.
D.Configure Amazon CloudWatch Logs to monitor S3 access.
AnswerB

CloudTrail data events capture object-level S3 operations, recording the caller identity, source IP address and action performed. Management events alone omit GetObject and PutObject, so enabling data events for the bucket is necessary to audit who accessed the sensitive data.

Why this answer

AWS CloudTrail data events capture object-level API activity on S3 buckets, including the identity of the caller, source IP address, timestamp, and the specific action (GetObject, PutObject, DeleteObject). Enabling data events on the bucket produces an auditable record of who accessed sensitive objects and from where, which is exactly what the requirement demands. Management events alone would only show bucket-level configuration changes, not object reads/writes.

Exam trap

DEA-C01 often tests the confusion between S3 server access logging and CloudTrail data events — candidates pick access logging because it sounds like an audit log, but only CloudTrail captures the full identity and IP context required for compliance audits.

How to eliminate wrong answers

Option A is wrong because S3 server access logs record requests to the bucket but do not natively capture the IAM identity or full request context (they show requester, but not the rich CloudTrail fields like userIdentity and full event JSON), and they are delivered best-effort with delays. Option C is wrong because AWS Config records resource configuration changes and compliance state — it does not log data-plane access to objects. Option D is wrong because CloudWatch Logs is a log aggregation service; S3 does not natively stream access events to CloudWatch Logs without CloudTrail or access logging already in place, so it is not the audit source itself.

113
Multi-Selectmedium

A data engineer needs to ensure that data stored in Amazon S3 is protected against accidental deletion and that the data cannot be altered for a period of 7 years to meet compliance requirements. The engineer must also ensure that the data remains encrypted at rest using SSE-KMS. Which two actions should the engineer take? (Choose two.)

Select 2 answers
A.Enable S3 Object Lock in compliance mode on the bucket and set a retention period of 7 years for all objects.
B.Use AWS Backup to create a vault lock with a 7-year retention policy and schedule daily backups of the S3 bucket.
C.Configure default encryption on the bucket to use SSE-KMS with a customer managed key.
D.Enable versioning on the bucket and configure a lifecycle rule to transition objects to S3 Glacier Deep Archive after 30 days.
E.Create an S3 bucket policy that denies s3:DeleteObject and s3:PutObject for all principals except a specific IAM role.
AnswersA, C

S3 Object Lock in compliance mode prevents objects from being overwritten or deleted for the specified retention period, even by the root user. This meets the requirement that data cannot be altered for 7 years. It also provides immutability, which is essential for compliance. Enabling Object Lock requires versioning to be enabled on the bucket.

Why this answer

To protect data from deletion and alteration for 7 years, S3 Object Lock in compliance mode must be enabled on the bucket with a retention period of 7 years. This ensures objects are immutable. Additionally, default encryption with SSE-KMS ensures data is encrypted at rest.

Other options like versioning, lifecycle rules, bucket policies, or AWS Backup do not provide the required immutability for the primary data in S3.

Exam trap

The trap here is thinking that a bucket policy denying delete operations or versioning alone can provide the same immutability as S3 Object Lock in compliance mode, which is specifically designed for regulatory retention.

114
MCQeasy

A data engineer needs to share a dataset from an S3 bucket in Account A with another AWS account (Account B). The data must remain encrypted at rest with KMS. Which steps are required?

A.Update the KMS key policy to allow Account B's root user
B.Create an IAM role in Account A and grant cross-account access
C.Update the S3 bucket policy and the KMS key policy to allow Account B
D.Update the S3 bucket policy to allow Account B's root user
AnswerC

Correct: Both the S3 bucket policy and the KMS key policy must be updated to allow Account B.

Why this answer

The correct steps are to update both the S3 bucket policy to grant Account B access to the objects and the KMS key policy to grant Account B the necessary decrypt permissions. Option A is insufficient because it only updates the KMS key policy; the bucket policy must also be updated to allow Account B's access to the S3 objects. Option B is also insufficient because creating an IAM role alone does not grant KMS decrypt permissions; the KMS key policy must be updated.

Option C is correct as it includes both. Option D is insufficient because it only updates the bucket policy, not the KMS key policy.

115
MCQeasy

A company wants to ingest real-time clickstream data from a website into Amazon S3 with minimal code. The data should be delivered within 60 seconds of generation. Which AWS service should be used?

A.Amazon Kinesis Data Firehose
B.AWS Database Migration Service (DMS)
C.Amazon Kinesis Data Streams
D.Amazon S3 Transfer Acceleration
AnswerA

Amazon Kinesis Data Firehose buffers incoming records and delivers them to Amazon S3 automatically, with a configurable buffer interval as low as 60 seconds, satisfying the near-real-time constraint. It requires no custom consumer code, meeting the minimal-code requirement, unlike Kinesis Data Streams, which needs a separate delivery application.

Why this answer

Amazon Kinesis Data Firehose is a fully managed service that can ingest real-time streaming data and deliver it to Amazon S3 with minimal code. It can buffer data and deliver within 60 seconds, making it ideal for this use case. Firehose requires no custom code for delivery and can scale automatically.

Exam trap

The trap is confusing Kinesis Data Streams with Firehose; candidates may think Data Streams can directly deliver to S3, but it requires custom code, whereas Firehose is the managed solution for minimal code.

How to eliminate wrong answers

Option B is wrong because AWS DMS is designed for database migration, not for ingesting clickstream data into S3. Option C is wrong because Kinesis Data Streams requires custom code to read from the stream and write to S3; it does not natively deliver to S3. Option D is wrong because S3 Transfer Acceleration speeds up uploads to S3 over long distances but does not provide a streaming ingestion pipeline for real-time data.

116
Multi-Selecteasy

A data engineer is setting up a data pipeline using AWS Glue. The engineer wants to monitor job failures and receive notifications. Which TWO services can be used together for this purpose?

Select 2 answers
A.AWS Step Functions
B.Amazon CloudWatch
C.Amazon SNS
D.Amazon Kinesis Data Streams
E.Amazon SQS
AnswersB, C

Amazon CloudWatch captures Glue job metrics and emits events on state changes such as FAILED, satisfying the requirement to detect job failures. Combined with Amazon SNS for notification delivery, it forms the monitoring half of the pair. Glue natively publishes logs and metrics to CloudWatch, so no custom instrumentation is needed.

Why this answer

Amazon CloudWatch (B) is correct because AWS Glue automatically publishes job metrics and logs to CloudWatch, and you can create CloudWatch alarms that trigger when a Glue job fails (for example, on the glue.driver.aggregate.numFailedTasks metric or a FAILED job run state). Amazon SNS (C) is correct because it is the notification service that CloudWatch alarms target via an SNS topic, delivering email, SMS, or HTTP notifications to subscribers when a Glue job failure alarm fires. Together, CloudWatch detects the failure and SNS delivers the alert, which is the standard AWS pattern for Glue job failure notifications.

AWS Step Functions (A) can orchestrate Glue jobs but is not a monitoring/notification service, Amazon Kinesis Data Streams (D) is for real-time streaming ingestion, and Amazon SQS (E) is a message queue for decoupling applications, so none of these provide the alarm-and-notify capability required here.

Exam trap

The trap here is that candidates may confuse AWS Step Functions (A) as a monitoring tool because it can orchestrate retries, but it does not natively send notifications and is not the primary service for monitoring Glue job failures.

117
MCQmedium

A company is using Amazon Redshift for analytics and needs to ensure that all data is encrypted at rest. The current cluster does not have encryption enabled. What is the most efficient way to enable encryption?

A.Change the cluster parameter group to enable encryption
B.Modify the cluster configuration to enable encryption
C.Use AWS DMS to migrate data to a new encrypted cluster
D.Create a snapshot of the cluster and restore it to a new cluster with encryption enabled
AnswerD

Redshift cannot enable encryption in place on an existing unencrypted cluster. Snapshotting and restoring into a new cluster with encryption enabled is the supported, least-effort migration path, preserving data while applying KMS encryption at rest.

Why this answer

Redshift does not support enabling encryption on an existing cluster; a new encrypted cluster must be created and data migrated. Modifying the cluster configuration or parameter groups does not enable encryption. Creating a snapshot and restoring it to a new cluster with encryption enabled is the standard approach.

118
MCQeasy

A data engineer is designing a data lake on AWS using Amazon S3. The data consists of CSV files generated by IoT devices. The data is accessed by multiple analytics jobs, and the engineer needs to ensure that new files are immediately visible to all consumers after writing. What S3 consistency model applies?

A.Consistent reads require S3 Object Lock.
B.Strong consistency for all operations.
C.Eventual consistency for all operations.
D.Read-after-write consistency for new object PUTS.
AnswerB

Amazon S3 now provides strong read-after-write consistency for all PUT and DELETE operations across all Regions, so newly written IoT CSV objects are immediately visible to every analytics job. This satisfies the stem's requirement that new files be instantly visible to all consumers.

Why this answer

Amazon S3 now provides strong consistency for all operations. After a successful write of a new object (or overwrite of an existing object), any subsequent read request immediately receives the latest version of the object, and list operations are also strongly consistent. Therefore, for new CSV files written to S3, the applicable model is strong consistency for all operations.

Exam trap

Candidates may incorrectly choose 'Read-after-write consistency for new object PUTS' (option D) because the question mentions new files after writing. While read-after-write behavior for new object PUTs is true, the current S3 consistency model is strong consistency for all operations. Another pitfall is selecting 'Eventual consistency for all operations' (option C) due to outdated knowledge of S3's previous eventual consistency model.

How to eliminate wrong answers

Option A is wrong because S3 Object Lock is a feature for preventing object deletion or overwrites for compliance or retention, not for ensuring consistency. Option B is wrong because while S3 now offers strong consistency for all operations (including overwrites and deletes), this was not always the case; historically S3 offered eventual consistency for overwrites, and the question's phrasing about 'new files' specifically tests the read-after-write consistency model for new PUTS. Option C is wrong because S3 no longer provides eventual consistency for new object PUTS; it guarantees strong read-after-write consistency for new objects since December 2020.

119
MCQmedium

A data engineer reviewed the S3 lifecycle policy shown in the exhibit. The engineer notices that objects under the 'logs/' prefix are being deleted after 365 days. The business requirement is to retain logs for at least 5 years. What should the engineer change in the lifecycle policy?

A.Change the prefix to 'logs/archive/'
B.Set the expiration days to 1825
C.Change the transition to GLACIER on day 365
D.Remove the expiration action
AnswerB

The lifecycle rule expires objects after 365 days, but the business needs five years of retention. Setting expiration to 1825 days (5 × 365) aligns deletion with the retention requirement, correcting the premature removal of logs under the 'logs/' prefix.

Why this answer

The business requirement is to retain logs for at least 5 years, which is 1,825 days (5 × 365). The current lifecycle policy sets expiration to 365 days, causing premature deletion. By setting the expiration days to 1,825, the S3 lifecycle policy will delete objects under the 'logs/' prefix only after 5 years, meeting the retention requirement.

Exam trap

The trap here is that candidates may confuse transition actions (which change storage class) with expiration actions (which delete objects), or incorrectly assume that changing the prefix or removing expiration will meet the retention requirement without adjusting the day count.

How to eliminate wrong answers

Option A is wrong because changing the prefix to 'logs/archive/' would only apply the lifecycle rules to a different subset of objects, not fix the retention period for the original 'logs/' prefix. Option C is wrong because transitioning to GLACIER on day 365 only changes the storage class for cost optimization; it does not extend the deletion timeline, so objects would still be deleted after 365 days. Option D is wrong because removing the expiration action entirely would mean objects are never automatically deleted, which may lead to indefinite storage and increased costs, not a 5-year retention.

120
Multi-Selecthard

A company uses Amazon Kinesis Data Analytics for Apache Flink to process streaming data. The Flink application reads from a Kinesis Data Streams source, performs aggregations, and writes results to Amazon S3. The application is experiencing high checkpoint failures, and the processing lag is increasing. The data volume is 50 MB/s with an average record size of 1 KB. Which TWO actions would improve checkpoint reliability and reduce lag? (Choose TWO.)

Select 2 answers
A.Decrease the checkpoint interval to complete checkpoints faster.
B.Replace the S3 sink with Kinesis Data Firehose.
C.Decrease the parallelism of the Flink application.
D.Increase the checkpoint interval in the Flink configuration.
E.Increase the number of Kinesis Processing Units (KPUs) for the application.
AnswersD, E

Less frequent checkpoints reduce overhead.

Why this answer

Increasing the checkpoint interval (Option D) reduces the frequency of checkpoint operations, which decreases the overhead on the Flink application and allows it to dedicate more resources to processing data, thereby reducing lag. This is especially effective when checkpoint failures are caused by the system being unable to complete checkpoints within the current interval due to high throughput (50 MB/s).

Exam trap

The trap here is that candidates often think decreasing the checkpoint interval will speed up checkpoints, but in reality, it increases overhead and failure rates, while increasing parallelism (Option C) seems intuitive but actually reduces per-task resources and can worsen backpressure.

121
Multi-Selectmedium

A data engineer is configuring an AWS Glue ETL job that reads from an Amazon S3 bucket and writes to another S3 bucket. The job must be able to handle a large number of small files efficiently and minimize the number of output files to improve downstream query performance. Which two actions should the engineer take? (Choose two.)

Select 2 answers
A.Use the AWS Glue groupFiles option to group multiple small input files into a single partition for processing.
B.Enable the job bookmark feature to track previously processed files and avoid reprocessing.
C.Increase the number of DPUs allocated to the job to allow more parallel processing of the small files.
D.Use the AWS Glue Data Catalog to store metadata about the small files and let the job read from the catalog.
E.Set the job's output to use the coalesce transformation to reduce the number of partitions before writing.
AnswersA, E

The groupFiles option in AWS Glue allows the job to combine multiple small input files into a single partition based on a target size. This reduces the overhead of opening and processing many small files, improving read efficiency. It directly addresses the large number of small files problem by consolidating them during the read phase, which is a recommended practice for Glue ETL jobs.

Why this answer

To handle many small input files, AWS Glue provides the groupFiles option, which combines small files into larger partitions for processing. To minimize output files, the coalesce transformation reduces the number of partitions before writing, resulting in fewer, larger files. Together, these actions improve both read efficiency and downstream query performance.

Increasing DPUs, enabling bookmarks, or using the Data Catalog do not directly address these file-count issues.

Exam trap

The trap here is thinking that adding more DPUs will solve the small file problem, when in fact it can increase the number of output files and does not consolidate them.

122
MCQmedium

A data engineer is building a data lake on Amazon S3 and needs to catalog metadata for a large number of CSV files stored in a nested folder structure. The engineer wants to automatically discover the schema and update the AWS Glue Data Catalog as new files are added. Which solution should the engineer use?

A.Run an AWS Glue ETL job that reads the CSV files and writes to a new location.
B.Create an AWS Glue crawler that points to the S3 bucket and schedule it to run periodically.
C.Use AWS Lake Formation to define a data lake and register the S3 location.
D.Manually define tables in the AWS Glue Data Catalog using the AWS Management Console.
AnswerB

AWS Glue crawlers automatically scan data sources, infer schemas, and populate the Data Catalog. They can be scheduled to run periodically to detect new files and update table definitions. This meets the requirement for automatic schema discovery and catalog updates with minimal manual effort. The crawler can handle nested folder structures and CSV format.

Why this answer

AWS Glue crawlers are designed to automatically scan data stores, infer schemas, and create or update tables in the Data Catalog. Scheduling a crawler ensures that new files are detected and the catalog stays current. This is the standard, low-effort solution for the described scenario.

Other options either require manual intervention or do not provide automatic schema discovery.

Exam trap

The trap here is confusing AWS Lake Formation with AWS Glue crawlers; Lake Formation manages permissions and governance but does not automatically discover schemas.

123
MCQeasy

A data engineer is monitoring an Amazon EMR cluster and notices that the cluster is running out of disk space on the core nodes. Which action can be taken to resolve this issue?

A.Reduce the retention period of data stored on HDFS
B.Change the core node instance type to a compute-optimized type
C.Increase the EBS volume size attached to core nodes
D.Use Spot Instances for core nodes
AnswerC

Core nodes hold HDFS data and shuffle output, so exhausted local disk blocks task execution. Expanding each core node's attached EBS volume increases available block storage, directly resolving the disk-space constraint without altering cluster topology or losing HDFS data.

Why this answer

Increasing the EBS volume size attached to core nodes directly adds storage capacity, resolving the disk space issue. Option A is wrong because reducing HDFS data retention may free space but does not increase available disk space; it could result in data loss. Option B is wrong because changing to a compute-optimized instance type affects CPU and memory, not storage.

Option D is wrong because Spot Instances are a pricing model and do not add disk space.

124
MCQmedium

A data engineer must ensure that all objects written to an S3 bucket by an AWS Glue ETL job are encrypted with a customer-managed AWS KMS key, even if the job does not explicitly specify encryption parameters. The bucket policy already denies unencrypted PUT requests. Which configuration will enforce the required encryption with the LEAST operational overhead?

A.Enable S3 Bucket Keys for the bucket to reduce KMS costs.
B.Set the default encryption on the S3 bucket to SSE-KMS with the customer-managed key.
C.Attach an IAM policy to the Glue job role that allows only s3:PutObject with the s3:x-amz-server-side-encryption condition.
D.Modify the Glue job to set the --encryption-type parameter to SSE-KMS for all writes.
AnswerB

Configuring default bucket encryption with SSE-KMS and the customer-managed key ensures that any object PUT without explicit encryption headers is automatically encrypted with that key. This meets the requirement with no changes to the Glue job and works even if the job omits encryption parameters. The bucket policy can still deny non-compliant requests as an additional layer.

Why this answer

Default bucket encryption with SSE-KMS using the customer-managed key automatically encrypts all objects written without explicit encryption headers, satisfying the requirement without modifying the Glue job. It also works with the existing bucket policy denial for non-compliant requests. This approach minimizes operational effort and ensures consistent encryption across all writers.

Exam trap

The trap here is assuming that a bucket policy denying unencrypted PUTs is sufficient, but it only blocks requests and does not apply encryption automatically.

125
Multi-Selecthard

A data engineer is using AWS Step Functions to orchestrate a daily ETL workflow that includes an AWS Glue job, an Amazon EMR step, and an Amazon Redshift stored procedure. The workflow occasionally fails and the engineer needs to troubleshoot and recover. Which TWO actions should the engineer take to identify the failure and resume from the failed step? (Choose two.)

Select 2 answers
A.Delete the state machine and recreate it with a retry policy on every state.
B.Use the Step Functions 'Redrive' capability to restart the execution from the failed state after fixing the underlying issue.
C.Enable AWS Step Functions logging to CloudWatch Logs and inspect the execution history for the failed state.
D.Configure the state machine to use an Amazon SQS dead-letter queue for failed states and reprocess messages manually.
E.Manually trigger the AWS Glue job and Amazon EMR step from the console, then restart the Step Functions execution from the beginning.
AnswersB, C

Redrive allows resuming a failed execution from the point of failure, preserving previous successful steps and their outputs. This avoids re-running the entire workflow and is the recommended way to recover after correcting the root cause. It requires the state machine to have the necessary IAM permissions and the execution to be in a failed state.

Why this answer

To troubleshoot a failed Step Functions execution, enabling logging and reviewing execution history provides the necessary error details. Once the root cause is fixed, the Redrive feature allows resuming from the failed state without re-running successful steps. These two actions together enable efficient diagnosis and recovery.

Exam trap

The trap here is assuming that restarting the entire workflow or manually rerunning individual jobs is the only recovery path, while overlooking Step Functions' native logging and Redrive capabilities.

126
MCQhard

A financial services company processes real-time stock trade data. They use Amazon Kinesis Data Streams with a shard count of 5, each shard receiving about 500 records per second. The consumer application uses the Kinesis Client Library (KCL) with DynamoDB for checkpointing. Lately, some records are being processed multiple times. What is the most likely cause?

A.The consumer application is crashing and restarting, causing re-processing of records.
B.The Kinesis stream's iterator age is exceeding the retention period.
C.The DynamoDB table used for checkpointing is throttling write requests.
D.The record size exceeds the 1 MB API limit, causing retries.
AnswerA

Frequent crashes force the KCL to resume from the last DynamoDB checkpoint, replaying every record consumed after it — at-least-once delivery guarantees duplicates on restart. With five shards at 500 records per second, each restart re-processes a substantial backlog, matching the observed duplicate processing.

Why this answer

The Kinesis Client Library (KCL) uses DynamoDB to track checkpoint progress for each shard. If the consumer application crashes and restarts, the KCL will resume processing from the last committed checkpoint, which may be behind the actual processing point. This causes records that were already processed (but not yet checkpointed) to be re-processed, leading to duplicate processing.

Exam trap

The trap here is that candidates often confuse checkpoint throttling (Option C) with duplicate processing, but throttling would cause checkpoint failures and potential re-processing only if the application cannot recover, whereas the direct cause of duplicates is the gap between processing and checkpointing after a crash.

How to eliminate wrong answers

Option B is wrong because iterator age exceeding the retention period would cause data to expire and become unavailable, not cause duplicate processing. Option C is wrong because DynamoDB throttling on checkpoint writes would cause checkpoint failures and potential re-processing, but the question states checkpointing is occurring and the issue is duplicate processing, not checkpoint failures. Option D is wrong because the 1 MB API limit applies to the total payload per PutRecords request, not per record, and exceeding it would cause write failures or retries, not duplicate processing of already-successful records.

127
MCQhard

A data engineer is using AWS Step Functions to orchestrate a daily ETL pipeline that includes an AWS Glue job, an Amazon EMR step, and an Amazon Redshift stored procedure. The pipeline occasionally fails with the error 'States.TaskFailed' from the Glue job, but the Glue job's own logs show that it completed successfully. The Step Functions execution history shows that the Glue job task timed out after 15 minutes, while the Glue job actually ran for 18 minutes. The Step Functions state machine uses the optimized Glue service integration with a TaskTimeout of 900 seconds. Which change will allow the pipeline to complete successfully without reducing the Glue job's runtime?

A.Change the Step Functions integration from the optimized Glue service integration to the AWS SDK integration with a longer timeout.
B.Add a Retry policy with a maximum of 3 attempts and an interval of 60 seconds to the Glue job task.
C.Reduce the Glue job's runtime by increasing the number of DPUs so it finishes within 15 minutes.
D.Increase the Step Functions TaskTimeout to a value greater than the Glue job's maximum expected runtime, such as 3600 seconds.
AnswerD

The Step Functions task timed out because its TaskTimeout was set to 900 seconds, but the Glue job took 18 minutes (1080 seconds). The optimized Glue service integration waits for the job to finish, and if the task exceeds the TaskTimeout, Step Functions fails the task even if the Glue job continues running. Increasing the TaskTimeout to a value larger than the job's maximum runtime, such as 3600 seconds, allows Step Functions to wait for completion without timing out.

Why this answer

The Step Functions task failed because its TaskTimeout was shorter than the Glue job's actual runtime. The optimized Glue service integration waits for the job to complete, but if the task exceeds the TaskTimeout, Step Functions marks it as failed even though the Glue job continues. Setting the TaskTimeout to a value greater than the job's maximum expected runtime ensures the state machine waits for the job to finish successfully.

Exam trap

The trap here is believing that a Retry policy or more DPUs will solve a timeout, when the real issue is that the Step Functions TaskTimeout is shorter than the job's runtime.

128
Multi-Selectmedium

A company needs to ingest streaming data from thousands of IoT devices. The data must be processed in real-time and stored in Amazon S3. Which TWO services should be used together?

Select 2 answers
A.Amazon Kinesis Data Streams
B.Amazon Kinesis Data Firehose
C.AWS Glue
D.Amazon Simple Queue Service (SQS)
E.AWS Direct Connect
AnswersA, B

Kinesis Data Streams ingests the high-volume IoT telemetry with low latency and durable, replicated storage, buffering thousands of device writes for downstream consumers. It satisfies the real-time processing requirement by feeding analytics or Lambda consumers before data lands in S3.

Why this answer

Amazon Kinesis Data Streams (A) is correct because it provides a highly scalable, real-time streaming ingestion layer that can continuously capture data from thousands of IoT devices with low latency and durable, ordered record storage across shards. Amazon Kinesis Data Firehose (B) is correct because it is the fully managed delivery service that can consume that streaming data and automatically batch, transform, and load it directly into Amazon S3 without writing custom consumer applications. Together they satisfy the requirement for real-time processing plus reliable storage in S3.

AWS Glue (C) is a serverless ETL and data catalog service, not a streaming ingestion or delivery mechanism, so it does not fit this pipeline. Amazon SQS (D) is a message queue for decoupling applications, not a real-time streaming service designed for high-throughput IoT telemetry or direct S3 delivery. AWS Direct Connect (E) is a dedicated network connection from on-premises to AWS, not a data streaming or processing service.

Exam trap

The trap here is that candidates often confuse Kinesis Data Firehose with Kinesis Data Streams, thinking only one is needed, but the question requires both: Data Streams for real-time ingestion from devices and Data Firehose for automated delivery to S3.

129
MCQeasy

A data engineer needs to securely store database credentials used by a Lambda function. The solution must automatically rotate the credentials every 90 days. Which AWS service should the engineer use?

A.AWS CloudHSM
B.AWS Systems Manager Parameter Store
C.IAM Roles for Lambda
D.AWS Secrets Manager
AnswerD

AWS Secrets Manager natively stores and rotates credentials on a schedule, satisfying the 90-day rotation requirement without custom code. Lambda retrieves secrets via the SDK at runtime, so database passwords never persist in environment variables or code.

Why this answer

AWS Secrets Manager is purpose-built for storing and managing secrets such as database credentials, API keys, and tokens. It natively supports automatic rotation via Lambda rotation functions on a schedule (e.g., every 90 days), and integrates with RDS, Redshift, and DocumentDB for managed rotation. This directly satisfies both the secure storage and automatic 90-day rotation requirements.

Exam trap

The trap here is confusing Parameter Store SecureString with Secrets Manager — both store secrets, but only Secrets Manager provides native scheduled rotation, which is the deciding requirement.

How to eliminate wrong answers

Option A is wrong because AWS CloudHSM is a dedicated hardware security module for cryptographic key operations and does not provide secret storage with built-in rotation workflows for database credentials. Option B is wrong because Systems Manager Parameter Store can store SecureString parameters but does not natively rotate database credentials on a schedule — rotation must be custom-built. Option C is wrong because IAM Roles for Lambda provide temporary AWS credentials to the function itself, not storage or rotation of external database credentials.

130
Multi-Selecteasy

A company is building a data pipeline that ingests streaming data from IoT devices. The data must be stored in a durable, scalable, and cost-effective manner for batch processing. Which TWO AWS services should be used together?

Select 2 answers
A.Amazon ElastiCache
B.Amazon Kinesis Data Streams
C.Amazon Redshift
D.Amazon DynamoDB
E.Amazon S3
AnswersB, E

Amazon Kinesis Data Streams ingests the IoT telemetry durably and scales by shard, satisfying the streaming ingestion requirement. It decouples producers from consumers, buffering records for the companion storage service that handles batch processing, so ingestion spikes never overwhelm downstream analytics.

Why this answer

Amazon Kinesis Data Streams (B) is the correct ingestion service for streaming IoT data because it provides a durable, scalable, and real-time data streaming platform that can capture and store data records for up to 365 days. Amazon S3 (E) is the correct storage service for batch processing because it offers virtually unlimited durability (99.999999999%), cost-effective tiered storage, and native integration with batch processing frameworks like Amazon EMR and AWS Glue. Together, they form a classic streaming-to-batch pipeline: Kinesis ingests and buffers the streaming data, which is then persisted in S3 for downstream batch analytics.

Exam trap

The trap here is that candidates often confuse Amazon Kinesis Data Streams with Amazon Kinesis Data Firehose (which directly writes to S3) or mistakenly choose Amazon Redshift for storage, overlooking that S3 is the correct durable and cost-effective storage layer for raw streaming data before any warehousing.

131
MCQeasy

A data engineer is running an AWS Glue ETL job that reads from an Amazon RDS MySQL database and writes to Amazon S3. The job fails with a 'Communications link failure' error. The security group for the RDS instance allows inbound traffic from the Glue job's security group. What is the most likely cause of the failure?

A.The JDBC connection string in the Glue job does not include the database name.
B.The Glue job is using the wrong JDBC driver.
C.The Glue job's security group does not allow outbound traffic to the RDS security group on port 3306.
D.The IAM role used by the Glue job does not have rds:Connect permission.
AnswerC

Security groups are stateful, but the Glue job's own outbound rules must still permit traffic to the RDS security group on MySQL port 3306. Inbound access alone is insufficient, so the missing egress rule explains the communications link failure.

Why this answer

AWS Glue ETL jobs run in a VPC that requires outbound security group rules to initiate connections to RDS. Even if the RDS security group allows inbound traffic from the Glue security group, the Glue security group must also have an outbound rule allowing traffic to the RDS security group on port 3306 (MySQL default port). Without this outbound rule, the TCP handshake from Glue to RDS fails, causing a 'Communications link failure'.

Exam trap

The trap here is that candidates assume only inbound rules matter for security groups, but outbound rules are equally critical for initiating connections from the client (Glue) to the server (RDS).

How to eliminate wrong answers

Option A is wrong because omitting the database name from the JDBC connection string would cause a different error (e.g., 'Unknown database' or connection rejection), not a 'Communications link failure', which indicates a network-level issue. Option B is wrong because AWS Glue automatically includes the correct JDBC driver for MySQL (compatible with Amazon RDS MySQL) when using the Glue connection type 'MySQL'; using the wrong driver would typically produce a class-not-found or driver-incompatibility error, not a communications link failure. Option D is wrong because IAM permissions for Glue jobs use actions like 'glue:GetConnection' and 'rds:DescribeDBInstances' to retrieve connection metadata, but there is no 'rds:Connect' IAM action; database authentication is handled via username/password in the Glue connection, not IAM.

132
MCQmedium

A company is ingesting data from multiple sources into S3 using AWS Glue. The data engineer notices that the Glue job is failing with an OutOfMemory error. Which step should be taken to resolve this issue?

A.Reduce the volume of incoming data
B.Configure the job to use a larger memory setting
C.Use a smaller file size for input
D.Increase the number of DPUs allocated to the Glue job
AnswerD

Glue allocates memory per executor across its DPUs; OutOfMemory errors during ingestion indicate insufficient memory for the workload's partitions. Increasing DPU count adds executors and memory capacity, directly resolving the resource constraint causing the job failure.

Why this answer

AWS Glue jobs run on Apache Spark, which distributes data processing across multiple executors. An OutOfMemory error typically indicates that the data being processed exceeds the memory available to the executors. Increasing the number of DPUs (Data Processing Units) allocates more memory and compute resources to the job, allowing it to handle larger datasets without running out of memory.

Exam trap

The trap here is that candidates may think they can directly increase memory settings (Option B) or reduce data volume (Option A), but AWS Glue abstracts memory management through DPUs, and the correct approach is to increase DPU allocation to provide more resources.

How to eliminate wrong answers

Option A is wrong because reducing the volume of incoming data is not a scalable solution and may not be feasible; the job should be able to handle the required data volume. Option B is wrong because AWS Glue does not allow direct configuration of memory settings per executor; memory is tied to DPU allocation, and increasing DPUs is the correct way to increase total memory. Option C is wrong because using a smaller file size for input does not address the root cause of memory exhaustion; Glue can process many small files efficiently, but the issue is the total data volume or skewed partitions, not file size.

133
Multi-Selectmedium

A data engineering team is using AWS DMS to migrate a 2 TB Oracle database to Amazon RDS for PostgreSQL. The migration must have minimal downtime and needs to capture ongoing changes after the full load. Which THREE resources are required for this task? (Choose three.)

Select 3 answers
A.A DMS source endpoint configured for Oracle.
B.An AWS DMS replication instance.
C.An AWS Snowball Edge device for initial data transfer.
D.An Amazon S3 bucket for staging the data.
E.A DMS target endpoint configured for Amazon RDS PostgreSQL.
AnswersA, B, E

Oracle is the migration source, so a DMS source endpoint must define connection details and credentials for the Oracle database. This endpoint lets the replication instance read the full load and, with change data capture enabled, ongoing redo log changes, satisfying the minimal-downtime requirement.

Why this answer

Option A is correct because AWS DMS requires a source endpoint that defines the Oracle database connection details, including server name, port, SID/service name, and credentials, so the replication instance can read the full load data and ongoing redo/archive log changes. Option B is correct because the AWS DMS replication instance is the managed compute resource that actually runs the migration tasks, performs the full load, and applies change data capture (CDC) from Oracle to the target. Option E is correct because a target endpoint pointing to the Amazon RDS for PostgreSQL instance is mandatory so DMS knows where to write the migrated tables and replicated changes.

Option C is not needed because Snowball Edge is an offline physical transfer device, whereas DMS performs the migration over the network and supports ongoing replication. Option D is not required because DMS does not need an Amazon S3 staging bucket for an Oracle-to-RDS PostgreSQL migration; S3 is only used for certain source/target types such as S3 endpoints or specific native backup/restore workflows.

Exam trap

DEA-C01 often tests whether candidates confuse DMS's core three-component architecture (source endpoint, target endpoint, replication instance) with optional services like S3 staging or Snowball that are only relevant for specific migration patterns.

134
Multi-Selecthard

A data engineer is using AWS Glue to process data from an Amazon Kinesis Data Stream. The Glue job is configured to run every 15 minutes and uses job bookmarks to track processed data. Recently, the job started reprocessing old data, leading to duplicate records in the target Amazon S3 bucket. The engineer verifies that the job bookmark is enabled and the job is not being run manually. Which TWO actions should the engineer take to resolve the duplicate processing issue? (Choose two.)

Select 2 answers
A.Check that the Glue job's bookmark key is not being changed between runs, as a changed bookmark key resets the bookmark state.
B.Increase the Kinesis stream's retention period to 168 hours to allow the Glue job more time to process data.
C.Configure the Glue job to use the '--job-bookmark-option' parameter set to 'job-bookmark-enable'.
D.Ensure that the Glue job is not being run concurrently with the same job bookmark, as concurrent runs can cause duplicate processing.
E.Verify that the Kinesis stream's shard iterator type is set to LATEST in the Glue job's connection options.
AnswersA, D

The job bookmark key is used to identify the bookmark state. If the key changes, the job treats it as a new job and starts processing from the beginning, causing duplicates. Ensuring the bookmark key remains consistent is crucial for proper bookmark functionality.

Why this answer

Duplicate processing with AWS Glue job bookmarks often occurs when the bookmark key changes or when concurrent runs interfere with bookmark updates. Verifying the bookmark key remains constant and preventing concurrent runs with the same bookmark are key steps. These actions ensure the bookmark state is preserved and correctly updated, preventing reprocessing of already-processed data.

Exam trap

The trap here is assuming that enabling job bookmarks is sufficient to prevent duplicates, when in fact bookmark key changes or concurrent runs can still cause reprocessing.

135
MCQmedium

A company uses AWS Glue ETL to process data from Amazon S3 and write results to Amazon Redshift. The job fails with a memory error when processing large files. Which action should the data engineer take to resolve this issue?

A.Reduce the number of partitions in the Glue job.
B.Increase the number of DPUs allocated to the Glue job.
C.Switch to a smaller instance type in the Glue job configuration.
D.Use S3 Select to filter columns before reading into Glue.
AnswerB

Increasing DPUs adds more executors and memory per worker, directly addressing the out-of-memory failure during large S3 file processing. Glue distributes partitions across these additional workers, so each handles a smaller share. This satisfies the stem's constraint: the job fails with a memory error, which horizontal scaling of compute resolves.

Why this answer

Increasing the number of DPUs (Data Processing Units) allocated to the AWS Glue job provides more memory and compute resources, which directly addresses the out-of-memory error when processing large files. Glue jobs run on Apache Spark, and insufficient DPUs can cause executors to run out of memory during shuffle or aggregation operations on large datasets.

Exam trap

The trap here is that candidates may confuse memory errors with I/O bottlenecks and incorrectly choose S3 Select (Option D) to reduce data volume, when the real issue is insufficient compute memory for Spark transformations.

How to eliminate wrong answers

Option A is wrong because reducing the number of partitions would increase the data size per partition, worsening memory pressure and likely causing the same or a more severe memory error. Option C is wrong because switching to a smaller instance type would reduce available memory per executor, directly contradicting the need to resolve a memory error. Option D is wrong because S3 Select can reduce the amount of data read from S3, but it does not increase the memory available to the Glue job's Spark executors; the memory error occurs during processing, not during data ingestion.

136
MCQmedium

A company uses AWS Glue to run ETL jobs on a schedule. Recently, a job failed with the error: 'AnalysisException: cannot resolve '`column_name`' given input columns: ...'. The job reads from an Amazon S3 source that has a schema defined in the AWS Glue Data Catalog. What is the MOST likely cause?

A.The schema of the source data has changed and is not reflected in the Data Catalog.
B.The source data file is corrupted and cannot be parsed.
C.The IAM role associated with the Glue job does not have permissions to read the S3 bucket.
D.The data type of the column in the source does not match the Data Catalog definition.
AnswerA

The AnalysisException arises because Spark resolves column names against the schema registered in the Data Catalog, not the actual S3 objects. If the source's schema evolved—columns renamed, added or dropped—without a crawler or manual update refreshing the catalog table, the job's query references columns the catalog no longer exposes.

Why this answer

The AnalysisException 'cannot resolve column_name given input columns' means the Spark/Glue job's code references a column that does not exist in the DataFrame schema derived from the Data Catalog. The most likely cause is that the source data's schema has evolved (new/renamed/dropped columns) but the Glue Data Catalog table definition was not updated, so the job's schema and the actual data are out of sync.

Exam trap

DEA-C01 often tests whether candidates can distinguish schema-resolution errors from permission or corruption errors, so the trap is picking IAM or file-corruption options for a clearly schema-related exception.

How to eliminate wrong answers

Option B is wrong because a corrupted file typically produces a parse/format error (e.g., 'MalformedInputException'), not a column-resolution error — Spark would fail reading the file, not resolving a column name. Option C is wrong because missing IAM permissions produce an 'Access Denied' error at the S3 read stage, not an AnalysisException about column resolution. Option D is wrong because a data-type mismatch would produce a cast/conversion error, not a 'cannot resolve column' error — the column would be found but with the wrong type.

137
MCQmedium

A company is using Amazon S3 to store large amounts of archival data. The data is accessed infrequently but must be immediately retrievable when needed. Which storage class is the most cost-effective choice?

A.S3 Standard
B.S3 Standard-IA
C.S3 Glacier Deep Archive
D.S3 Intelligent-Tiering
AnswerB

S3 Standard-IA matches the stem's dual constraint: infrequent access plus immediate retrieval. Its lower per-GB storage price than S3 Standard suits archival data, while millisecond retrieval preserves instant availability. Glacier classes would fail the immediate-retrieval requirement despite lower cost.

Why this answer

S3 Standard-IA (Infrequent Access) is the most cost-effective choice because it offers lower storage costs than S3 Standard while still providing millisecond first-byte latency for immediate retrieval. The data is accessed infrequently but requires instant availability, which matches the IA use case exactly.

Exam trap

The trap here is that candidates confuse 'immediately retrievable' with 'lowest cost' and choose Glacier Deep Archive, overlooking the critical requirement for instant access versus the 12-48 hour retrieval time of Deep Archive.

How to eliminate wrong answers

Option A is wrong because S3 Standard is designed for frequently accessed data and has higher storage costs than Standard-IA, making it less cost-effective for archival data with infrequent access. Option C is wrong because S3 Glacier Deep Archive has the lowest storage cost but retrieval times range from 12 to 48 hours, failing the 'immediately retrievable' requirement. Option D is wrong because S3 Intelligent-Tiering automatically moves data between tiers based on access patterns but incurs a monthly monitoring and automation fee per object, making it less cost-effective than Standard-IA for a predictable infrequent access pattern.

138
MCQhard

A data engineer is setting up an Amazon Kinesis Data Analytics application to process streaming data from a Kinesis data stream named "input-stream". The application uses a reference data source from an S3 bucket. The engineer has attached the IAM policy shown in the exhibit to the application's IAM role. When starting the application, the engineer receives an 'AccessDeniedException' error. Which additional permission is required?

A.kinesis:PutRecord on the input stream
B.s3:GetObject on the S3 bucket containing the reference data
C.kinesis:CreateStream on the input stream
D.kinesis:PutRecords on the input stream
AnswerB

Reading reference data from S3 requires object-level read access, which the attached policy omits. Granting s3:GetObject on the bucket holding the reference data supplies the missing action, resolving the AccessDeniedException thrown when Kinesis Data Analytics loads the reference source at application start.

Why this answer

The Kinesis Data Analytics application needs to read reference data from the S3 bucket, which requires the s3:GetObject permission on the bucket and its objects. The error 'AccessDeniedException' indicates the IAM role lacks this specific permission to retrieve the reference data file. Option B correctly adds the missing s3:GetObject action to allow the application to fetch the reference data from S3.

Exam trap

The trap here is that candidates often confuse the direction of data flow and assume the application needs write permissions (PutRecord/PutRecords) to the input stream, when in fact it only needs read permissions (kinesis:DescribeStream, kinesis:GetShardIterator, kinesis:GetRecords) and the missing permission is for the separate S3 reference data source.

How to eliminate wrong answers

Option A is wrong because kinesis:PutRecord is used to write data to a Kinesis stream, but the application reads from the input stream as a source, not writes to it; the error is not about writing. Option C is wrong because kinesis:CreateStream is an administrative action to create a new stream, which is irrelevant to an existing stream used as input. Option D is wrong because kinesis:PutRecords is for batch writing to a stream, not for reading or for accessing reference data from S3.

139
MCQmedium

A data engineer needs to run an AWS Glue for Apache Spark ETL job that joins a 40 GB Amazon S3 Parquet dataset with a small 8 MB reference lookup table stored as CSV in Amazon S3. The reference table is read on every join and the job's executors are spending a large amount of shuffle time on the join. The reference table changes only once per month. Which approach MOST efficiently reduces shuffle overhead in the Glue job?

A.Partition both datasets by the join key and write them to S3 before the join so Spark can perform a partition-wise merge join.
B.Increase the number of Glue DPUs so that more executors are available to parallelize the shuffle stage of the join.
C.Broadcast the reference table by reading it with the Spark DataFrame API and applying broadcast() before the join.
D.Convert the reference CSV to Parquet and enable Glue job bookmarks so the lookup is only read once per month.
AnswerC

Broadcasting the small reference DataFrame replicates it to each executor, so the large Parquet dataset never needs to be shuffled. This eliminates the expensive shuffle of the 40 GB dataset that the join currently causes. Because the table is only 8 MB and changes monthly, broadcasting is safe and fits comfortably in executor memory, making it the most efficient fix for shuffle overhead in the Glue job.

Why this answer

The small, slowly changing reference table is an ideal broadcast candidate. Replicating it to every executor lets Spark perform a broadcast hash join where the large Parquet side stays in place and is never redistributed. Increasing DPUs, converting formats, or pre-partitioning all leave the shuffle of the large dataset intact, so they fail to address the actual bottleneck.

Exam trap

The trap here is assuming that adding compute resources or changing file formats eliminates shuffle cost, when only broadcasting the small side removes the redistribution of the large dataset.

140
Multi-Selectmedium

A company is designing a data lake on Amazon S3. Which TWO strategies improve query performance for Amazon Athena?

Select 2 answers
A.Enable S3 Versioning on the bucket.
B.Use server-side encryption with AWS KMS (SSE-KMS).
C.Partition the data by frequently queried columns such as date or region.
D.Use columnar file formats like Parquet or ORC.
E.Store data in CSV format with header rows.
AnswersC, D

Partitioning by frequently filtered columns lets Athena prune entire S3 prefixes using partition metadata, so it scans only relevant data rather than the whole table. This directly cuts bytes scanned and query runtime for date- or region-filtered workloads.

Why this answer

Option C is correct because partitioning the S3 data by frequently filtered columns such as date or region lets Athena prune partitions and scan only the relevant prefixes, drastically reducing the amount of data read and thus improving query speed and lowering cost. Option D is correct because columnar formats like Parquet or ORC allow Athena to read only the columns referenced in the query and benefit from compression and predicate pushdown, minimizing I/O compared to row-based formats. Options A and B do not improve query performance: S3 Versioning only retains object versions for recovery, and SSE-KMS only encrypts data at rest, adding no scan optimization (and KMS calls can even add latency).

Option E is incorrect because CSV is a row-based, uncompressed text format that forces Athena to scan entire rows and cannot skip unneeded columns, making queries slower and more expensive than Parquet/ORC.

Exam trap

The trap here is that candidates often confuse data management features (like versioning or encryption) with performance optimizations, or assume that simpler formats like CSV are sufficient for analytics, ignoring the significant performance benefits of partitioning and columnar storage.

141
MCQhard

A company uses AWS KMS to encrypt sensitive data stored in S3. To meet compliance requirements, they need to ensure that the encryption keys are automatically rotated every year. Which type of KMS key should they use?

A.Customer managed key with manual rotation
B.AWS managed key
C.Custom key store (CloudHSM) key
D.Customer managed key with automatic rotation enabled
AnswerD

Customer managed keys in AWS KMS support automatic annual rotation, which AWS-managed keys handle on a fixed three-year cycle you cannot change. Enabling rotation on a customer managed key satisfies the yearly compliance requirement, since the key material is replaced without altering the key ID or ARN.

Why this answer

Customer managed keys with automatic rotation enabled support automatic annual rotation, meeting the compliance requirement. AWS managed keys rotate automatically every year, but they cannot be controlled or customized by the customer, so they are not the best choice when the customer needs to manage the key policy or rotation schedule. Custom key stores (CloudHSM) do not support automatic rotation.

Option A (customer managed key with manual rotation) requires manual intervention to rotate, not automatic. Therefore, D is the correct answer.

142
Multi-Selectmedium

A company uses AWS Glue to run ETL jobs daily. The jobs consume data from an Amazon RDS for MySQL database and write results to Amazon S3. The company wants to minimize the impact on the source database during extraction. Which THREE actions should the data engineer take to achieve this? (Choose THREE.)

Select 3 answers
A.Schedule the Glue job to run during off-peak hours.
B.Configure the Glue job to connect to a read replica of the RDS instance.
C.Increase the number of Glue DPUs to process data faster.
D.Disable Glue job bookmarks to force full refresh.
E.Use a JDBC connection with a WHERE clause to extract only incremental data.
AnswersA, B, E

Runs when database load is naturally low.

Why this answer

Scheduling the Glue job to run during off-peak hours minimizes the load on the source RDS for MySQL database by avoiding high-traffic periods, reducing contention for CPU, memory, and I/O resources. This is a straightforward operational practice to reduce impact on production databases during extraction.

Exam trap

The trap here is that candidates often assume increasing DPUs (parallelism) always improves performance without realizing it can amplify the load on the source database, and they may overlook that disabling bookmarks forces full refreshes, which is the opposite of minimizing impact.

143
MCQhard

A data engineer is troubleshooting an AWS Glue ETL job that fails with the error 'java.lang.OutOfMemoryError: Java heap space'. The job processes a large number of small files in Amazon S3. Which action would MOST effectively resolve the issue?

A.Enable S3 groupFiles option in the Glue job
B.Change the worker type to G.1X
C.Increase the number of workers in the Glue job
D.Use a G.2X worker type with more memory
AnswerA

Grouping coalesces many small S3 objects into larger input partitions, so Glue reads far fewer files and holds less per-file metadata and buffer overhead in the driver and executors. This directly relieves the Java heap exhaustion caused by the large number of small files described in the stem.

Why this answer

The 'java.lang.OutOfMemoryError: Java heap space' in AWS Glue when processing many small files is typically caused by the driver or executor accumulating too many file metadata objects. Enabling the S3 groupFiles option (with groupSize and groupFiles parameters) consolidates small files into larger groups, reducing the number of objects processed and alleviating heap pressure. This directly addresses the root cause of the memory issue.

Exam trap

DEA-C01 often tests the misconception that scaling up worker type or count solves OutOfMemory errors, when the root cause (many small files) requires the groupFiles optimization instead.

How to eliminate wrong answers

Option B is wrong because changing to G.1X workers increases memory per worker but does not reduce the number of small files being processed, so the heap pressure from file metadata remains. Option C is wrong because increasing the number of workers adds parallelism but does not solve the per-executor memory issue caused by many small files; it may even worsen driver memory pressure. Option D is wrong because G.2X workers provide more memory per worker, but again, the fundamental problem of too many small files persists, and the driver may still run out of heap.

144
MCQmedium

A data engineer is troubleshooting a failed AWS Glue job that reads from an Amazon RDS for MySQL table. The error message indicates 'java.sql.SQLException: No suitable driver'. What is the most likely cause?

A.The MySQL JDBC driver JAR is not included in the Glue job's dependencies.
B.The Glue job is using the wrong JDBC driver class name.
C.The Glue job's VPC subnet does not have a route to the RDS instance.
D.The RDS instance is not publicly accessible.
AnswerA

Glue's JDBC connection requires the MySQL driver JAR to be supplied via the job's dependent JARs path; the bundled environment lacks it, so the driver class cannot load and the connection throws 'No suitable driver'. Adding the MySQL Connector/J JAR resolves the missing driver.

Why this answer

The error 'java.sql.SQLException: No suitable driver' indicates that the Java application (the Glue job) cannot find a suitable JDBC driver for the database URL. In AWS Glue, you must include the JDBC driver JAR as a dependency for the job. If the MySQL JDBC driver JAR is not included, the job cannot load the driver class, resulting in this error.

Exam trap

DEA-C01 often tests the confusion between connectivity errors (timeouts, network) and driver errors, where 'No suitable driver' specifically points to a missing or misconfigured JDBC driver JAR.

How to eliminate wrong answers

Option B is wrong because if the wrong JDBC driver class name were used, the error would typically be 'ClassNotFoundException' or 'No suitable driver' only if the class is not found, but the more common cause is the missing JAR. Option C is wrong because a network routing issue would result in a connection timeout or 'Communications link failure', not 'No suitable driver'. Option D is wrong because if the RDS instance is not publicly accessible, the error would be a connection timeout or 'Unknown host', not a driver error.

145
Multi-Selecthard

A company is using Amazon Redshift for its data warehouse. The data engineering team needs to improve query performance for a large fact table that is frequently joined with multiple dimension tables. Which THREE strategies should be considered?

Select 3 answers
A.Define sort keys on columns used in WHERE clauses.
B.Use DISTSTYLE EVEN to distribute data evenly.
C.Increase the number of nodes in the cluster.
D.Choose an appropriate distribution key based on join columns.
E.Apply columnar compression to reduce storage and I/O.
AnswersA, D, E

Sort keys physically order rows on disk by the chosen columns, so range and equality predicates in WHERE clauses skip irrelevant blocks via zone maps. This reduces the rows scanned for the large fact table, directly improving the filtered queries feeding the joins.

Why this answer

Option A is correct because defining sort keys on columns frequently used in WHERE clauses allows Redshift to use zone maps and block-level metadata to skip irrelevant blocks, dramatically reducing I/O for filtered queries on the large fact table. Option D is correct because choosing a distribution key based on the fact table's join columns co-locates matching rows on the same node slice, enabling collocated joins that avoid costly data redistribution (broadcast or shuffle) across the cluster. Option E is correct because columnar compression reduces the amount of data read from disk and improves I/O efficiency, which is especially beneficial for large fact tables scanned by analytical queries.

Option B is not ideal here because DISTSTYLE EVEN spreads rows uniformly but does not co-locate join keys, so joins with dimension tables require network redistribution, hurting performance for frequent joins. Option C is not a targeted strategy for join performance; adding nodes increases compute and storage capacity but does not by itself address sort keys, distribution keys, or compression, and may not resolve join-related bottlenecks.

Exam trap

The trap here is that candidates often assume DISTSTYLE EVEN is always the best choice for performance, but for frequently joined fact tables, a distribution key aligned with the join columns is critical to avoid network-heavy data shuffling.

146
MCQmedium

A data engineer needs to store sensitive data in Amazon S3 and automatically classify the data using a managed service. The data is uploaded via an S3 bucket. Which AWS service can automatically detect and classify sensitive data?

A.Amazon Macie
B.AWS WAF
C.Amazon Inspector
D.AWS Shield
AnswerA

Amazon Macie uses machine learning and pattern matching to automatically discover, classify, and alert on sensitive data such as personally identifiable information in S3. It is the managed service that fulfils the automatic classification requirement without custom code.

Why this answer

Amazon Macie is a fully managed data security and data privacy service that uses machine learning and pattern matching to automatically discover, classify, and protect sensitive data stored in Amazon S3. It continuously evaluates S3 buckets and identifies sensitive content such as PII, credentials, and financial data, generating findings that can be routed to EventBridge or Security Hub. This directly matches the requirement to automatically detect and classify sensitive data in S3.

Exam trap

The trap here is confusing security services that operate at different layers — candidates often pick AWS WAF or Shield because they associate 'security' with network protection, but the question specifically asks for data classification in S3, which only Macie provides.

How to eliminate wrong answers

Option B is wrong because AWS WAF is a web application firewall that filters HTTP/HTTPS traffic to protect web apps from exploits like SQL injection and XSS — it does not inspect or classify data at rest in S3. Option C is wrong because Amazon Inspector is a vulnerability management service that scans EC2 instances, container images, and Lambda functions for software vulnerabilities and network exposure, not S3 data classification. Option D is wrong because AWS Shield is a managed DDoS protection service for applications running on AWS, with no data classification capability.

147
Multi-Selecthard

A company stores sensitive data in Amazon S3. The security team requires encryption at rest and that the encryption keys are managed by the company using AWS KMS. The data is frequently accessed by multiple AWS services. Which THREE steps should be taken to meet these requirements?

Select 3 answers
A.Use client-side encryption with the KMS key before uploading to S3
B.Configure the KMS key policy to allow the necessary AWS services to use the key for decryption
C.Enable default encryption on the S3 bucket using SSE-S3
D.Create a bucket policy that denies s3:PutObject if the object is not encrypted with SSE-KMS
E.Enable default encryption on the S3 bucket using SSE-KMS
AnswersB, D, E

Services must have decrypt permissions to access the encrypted objects.

Why this answer

The security team requires that encryption keys be managed by the company using AWS KMS, and that multiple AWS services can access the data. To allow those services to decrypt objects encrypted with a customer-managed KMS key, the KMS key policy must explicitly grant the necessary AWS services (e.g., AWS Lambda, Amazon Athena) permission to use the key for decryption (kms:Decrypt). Without this policy, even if the bucket is configured for SSE-KMS, the services will fail to read the encrypted objects.

Exam trap

AWS often tests the distinction between enforcing encryption (bucket policy) and enabling access to encrypted data (KMS key policy), leading candidates to overlook the KMS key policy step when multiple services need to decrypt objects.

148
MCQmedium

An AWS Glue job that performs data transformation on large Parquet files in Amazon S3 is taking a long time to complete. The job uses the default number of DPUs. Which change would most likely improve the job's performance?

A.Increase 'Max capacity' (number of DPUs) for the job.
B.Use 'coalesce' to reduce the number of output files.
C.Reduce the number of partitions in the source data.
D.Change the input format from Parquet to CSV.
AnswerA

The job is bottlenecked by compute, not partitioning or file size. Raising Max capacity adds more DPUs, distributing the Parquet transformation across additional executors and reducing wall-clock time, directly addressing the stem's constraint that the job runs on the default DPU allocation.

Why this answer

AWS Glue allocates DPUs (Data Processing Units) to run the job, and increasing Max capacity adds more compute and memory resources that can parallelize the transformation across partitions. For large Parquet files with the default DPU count, the job is likely CPU- or memory-bound, so scaling DPUs is the most direct performance improvement. Parquet is already a columnar, compressed format, so the bottleneck is compute, not I/O format.

Exam trap

DEA-C01 often tests the misconception that reducing output files or changing file formats improves Glue performance, when the real lever for compute-bound transformations is scaling DPUs and ensuring adequate partitioning.

How to eliminate wrong answers

Option B is wrong because coalesce reduces the number of output files, which can actually hurt parallelism and does not address the compute bottleneck during transformation. Option C is wrong because reducing source partitions decreases parallelism, making the job slower, not faster. Option D is wrong because switching from Parquet to CSV increases I/O volume and loses columnar pushdown benefits, degrading performance rather than improving it.

149
Multi-Selecteasy

Which THREE AWS services can be used to centrally manage and govern data across multiple AWS accounts? (Select THREE.)

Select 3 answers
A.Amazon S3
B.AWS Control Tower
C.AWS Organizations
D.Amazon Redshift
E.AWS Lake Formation
AnswersB, C, E

AWS Control Tower applies guardrails and account baselines across an organisation, enforcing governance at scale. It satisfies the multi-account central governance requirement by orchestrating AWS Organizations, IAM Identity Center and CloudTrail into a managed landing zone.

Why this answer

AWS Control Tower (B) is correct because it provides a centralized landing zone with guardrails and account provisioning that enforces governance and data management policies across multiple AWS accounts. AWS Organizations (C) is correct because it enables central management of multiple accounts through organizational units, service control policies (SCPs), and consolidated billing, which are foundational for cross-account governance. AWS Lake Formation (E) is correct because it centrally manages data lake permissions, cataloging, and fine-grained access control across accounts using AWS RAM resource shares and the Glue Data Catalog.

Amazon S3 (A) is incorrect because it is an object storage service, not a cross-account governance or central management service, even though it can host data governed by other services. Amazon Redshift (D) is incorrect because it is a data warehouse service for analytics, not a service for centrally managing and governing data across multiple accounts.

150
MCQmedium

A company is using Amazon Kinesis Data Firehose to deliver streaming data to Amazon S3. The data must be transformed from JSON to Parquet format before landing in S3. The transformation logic is simple: convert the JSON schema to Parquet. Which approach meets the requirements with the least operational overhead?

A.Use the built-in data format conversion feature of Firehose with an AWS Glue Data Catalog table
B.Use an AWS Lambda function to transform records to Parquet before sending to Firehose
C.Use Amazon Kinesis Data Analytics to convert the stream to Parquet
D.Provision an Amazon EMR cluster to convert the data in micro-batches
AnswerA

Firehose's built-in data format conversion uses an AWS Glue Data Catalog table as the schema reference to transform JSON records into Parquet before delivery. This is serverless and requires no custom code, minimising operational overhead.

Why this answer

Amazon Kinesis Data Firehose provides a built-in data format conversion feature that can automatically convert incoming JSON data to Parquet format using an AWS Glue Data Catalog table as the schema reference. This approach requires no custom code, no additional infrastructure, and no manual transformation logic, making it the simplest solution with the least operational overhead for a straightforward JSON-to-Parquet conversion.

Exam trap

The trap here is that candidates often overcomplicate the solution by choosing Lambda or EMR, not realizing that Firehose's built-in format conversion with Glue Data Catalog is the simplest, fully managed option for JSON-to-Parquet conversion without any custom code.

How to eliminate wrong answers

Option B is wrong because using an AWS Lambda function to transform records to Parquet before sending to Firehose introduces unnecessary complexity, additional cost, and operational overhead (e.g., managing Lambda concurrency, packaging Parquet libraries, handling record size limits), whereas Firehose's built-in conversion handles this natively. Option C is wrong because Amazon Kinesis Data Analytics is designed for real-time analytics and stream processing using SQL or Flink, not for simple format conversion; it adds latency, complexity, and cost without benefit for a straightforward schema conversion. Option D is wrong because provisioning an Amazon EMR cluster to convert data in micro-batches is a heavy, over-engineered solution that requires cluster management, scaling, and job orchestration, far exceeding the operational overhead needed for a simple format conversion that Firehose can perform automatically.

Page 1

Page 2 of 18

Page 3