Courseiva

AWS Certified Data Engineer Associate DEA-C01 (DEA-C01) — Questions 376–450

1321 questions total · 18pages · All types, answers revealed

Page 5

Page 6 of 18

Page 7
376
MCQeasy

A company wants to store data from thousands of IoT devices with varying data rates. The data must be stored in a schema-on-read fashion and support SQL queries. Which AWS service should be used?

A.Amazon RDS for MySQL
B.Amazon S3 with Amazon Athena
C.Amazon DynamoDB
D.Amazon Redshift
AnswerB

S3's schema-on-read storage decouples ingestion from structure, letting thousands of IoT devices write at varying rates without transformation. Athena then queries that data in place using standard SQL, satisfying both the schema-on-read and SQL query constraints without managing servers or loading into a warehouse.

Why this answer

Amazon S3 stores data in its native format (e.g., JSON, Parquet) without requiring a predefined schema, enabling schema-on-read. Amazon Athena uses Presto-based SQL to query data directly from S3, making it ideal for IoT data with varying rates and ad-hoc SQL analysis without provisioning servers.

Exam trap

The trap here is that candidates confuse schema-on-read with schema-on-write, assuming DynamoDB's flexible schema or Redshift's SQL support fits, but they miss that DynamoDB lacks native SQL and Redshift requires upfront table definitions, while Athena directly queries raw files in S3 with SQL.

How to eliminate wrong answers

Option A is wrong because Amazon RDS for MySQL requires a fixed schema defined before writing data, which contradicts the schema-on-read requirement and cannot handle the high write throughput of thousands of IoT devices without scaling limitations. Option C is wrong because Amazon DynamoDB is a NoSQL key-value and document database that does not support SQL queries natively (it uses PartiQL with limited SQL compatibility) and is not designed for schema-on-read. Option D is wrong because Amazon Redshift is a columnar data warehouse that requires schema-on-write (tables must be defined before loading data) and is optimized for structured, batch-loaded analytics rather than streaming IoT ingestion with varying data rates.

377
MCQhard

A data engineer is building an AWS Glue job that reads from a large Parquet dataset in Amazon S3 partitioned by year/month/day and writes aggregated results to Amazon Redshift. The job currently reads all partitions and takes several hours. The engineer wants the job to process only partitions from the last seven days and reduce runtime. Which change should the engineer make?

A.Pass a pushdown predicate in the from_catalog call that filters on the year, month, and day partition columns for the last seven days.
B.Convert the source data to ORC format to reduce the bytes read per partition.
C.Increase the number of DPUs allocated to the Glue job and enable auto scaling.
D.Enable job bookmarks so the job processes only new files since the last successful run.
AnswerA

A pushdown predicate on partition columns is evaluated by the Glue Data Catalog and S3 reader before data is loaded, so only the matching partitions are listed and read. Restricting to the last seven days dramatically reduces the number of files scanned, cutting runtime and cost while satisfying the requirement.

Why this answer

Partition pruning is the key optimization for partitioned datasets on S3. Supplying a pushdown predicate that references the year, month, and day partition columns lets the Glue reader skip non-matching partitions entirely, so only seven days of data are listed and read. More workers, format conversion, and bookmarks do not restrict the read window and therefore do not solve the runtime problem.

Exam trap

The trap here is reaching for more DPUs or bookmarks, when the actual bottleneck is scanning partitions that fall outside the desired date range.

378
Multi-Selecteasy

Which TWO of the following are features of Amazon RDS Multi-AZ deployments? (Choose 2.)

Select 2 answers
A.Read replicas in the same region for offloading read traffic.
B.Automatic failover to a standby instance in case of an AZ failure.
C.A standby instance that is not accessible for reads or writes.
D.Automatic storage scaling based on usage.
E.Synchronous replication across AWS Regions.
AnswersB, C

Multi-AZ automatically fails over to the standby in another AZ.

Why this answer

Amazon RDS Multi-AZ deployments automatically handle failover to a standby instance in a different Availability Zone when the primary instance fails or the AZ becomes unavailable. This synchronous replication ensures zero data loss and minimal downtime, with the standby instance automatically promoted to primary without manual intervention.

Exam trap

The DEA-C01 exam often tests the distinction between Multi-AZ (high availability with a passive standby) and read replicas (scaling reads with active replicas), leading candidates to incorrectly associate read offloading with Multi-AZ deployments.

379
MCQmedium

A data engineer is using AWS Step Functions to orchestrate a series of AWS Glue jobs. One of the Glue jobs occasionally fails due to transient issues, such as a temporary network glitch. The engineer wants the Step Functions state machine to automatically retry that specific Glue job up to three times before moving to a failure state. Which Step Functions feature should be used to implement this?

A.Configure a Catch field to transition to a retry state on failure.
B.Use a Parallel state to run the Glue job multiple times concurrently.
C.Increase the Glue job's timeout and memory allocation.
D.Add a Retry field to the state that invokes the Glue job.
AnswerD

Step Functions states support a Retry field where you can specify the error types to retry, the interval between retries, the maximum number of attempts, and a backoff rate. This is exactly designed for transient failures. By configuring MaxAttempts to 3, the state will automatically retry the Glue job invocation three times before failing.

Why this answer

The Retry field in Step Functions is specifically designed to handle transient errors by automatically retrying a state. It allows you to specify the error types to retry (e.g., States.TaskFailed), the maximum number of attempts, the interval between retries, and a backoff rate. This is the standard way to implement retry logic for a Glue job invocation within a state machine.

Exam trap

The trap here is confusing error handling with retry logic; Catch handles errors by branching, while Retry actually re-executes the state.

380
Multi-Selecthard

A data pipeline uses AWS Glue to process large CSV files. The team notices that some jobs fail with out-of-memory errors. Which TWO configuration changes can help mitigate this issue?

Select 2 answers
A.Reduce the number of DPUs to limit concurrency.
B.Increase the number of DPUs for the Glue job.
C.Enable Glue job autoscaling.
D.Convert input files from CSV to Parquet.
E.Enable job bookmarks.
AnswersB, C

Adding DPUs allocates more executors and memory per worker, directly relieving the heap pressure causing out-of-memory failures when processing large CSV files. Horizontal scaling suits Glue's distributed Spark engine, satisfying the stem's memory constraint without altering the transformation logic itself.

Why this answer

Options B and C are correct: increasing the number of DPUs provides more memory, and enabling autoscaling allows the job to automatically scale resources as needed. Option A (reducing DPUs) would worsen the problem by limiting resources. Option D (converting to Parquet) can improve performance but is not a direct configuration change for the Glue job itself.

Option E (job bookmarks) is for incremental processing and does not affect memory.

381
MCQmedium

A data engineer is designing a data store for a real-time leaderboard application that requires sub-millisecond read and write latency. The leaderboard stores scores for millions of users and needs to be sorted by score. Which AWS service should the engineer use?

A.Amazon RDS for PostgreSQL with an index on score
B.Amazon DynamoDB with a global secondary index on score
C.Amazon ElastiCache for Redis with a sorted set
D.Amazon Neptune with a graph model
AnswerC

Amazon ElastiCache for Redis sorted sets store members ordered by score, giving O(log N) inserts and rank queries entirely in memory, which delivers the sub-millisecond latency the leaderboard demands. This satisfies the stem's requirement for millions of users sorted by score, unlike disk-based stores such as DynamoDB.

Why this answer

Amazon ElastiCache for Redis provides a sorted set data structure (ZADD/ZRANGE commands) that maintains elements ordered by a numeric score with O(log N) complexity for both writes and reads, enabling sub-millisecond latency for real-time leaderboard updates and queries. This is the only option purpose-built for in-memory, sorted, real-time leaderboards at scale.

Exam trap

The trap here is that candidates often choose DynamoDB (Option B) because it is a common NoSQL choice for high-performance applications, but they overlook that DynamoDB lacks a native sorted data structure and requires costly scan operations to retrieve a globally sorted leaderboard, whereas Redis sorted sets are purpose-built for this exact use case.

How to eliminate wrong answers

Option A is wrong because Amazon RDS for PostgreSQL, even with an index on score, is a disk-based relational database that cannot guarantee sub-millisecond read/write latency under high concurrency due to disk I/O and transaction overhead. Option B is wrong because Amazon DynamoDB with a global secondary index on score does not natively maintain a globally sorted order; it requires expensive scan operations to retrieve the top scores, and write latency can exceed sub-millisecond under heavy load due to throughput limits and index propagation. Option D is wrong because Amazon Neptune is a graph database designed for traversing relationships, not for sorted, real-time score retrieval, and its query latency is not optimized for sub-millisecond leaderboard operations.

382
Multi-Selecthard

Which THREE steps are recommended for migrating an on-premises Oracle database to Amazon RDS for Oracle with minimal downtime? (Choose 3.)

Select 3 answers
A.Set up a VPN or Direct Connect between on-premises and AWS
B.Disable archiving on the source database
C.Use AWS Schema Conversion Tool (SCT) to convert the schema
D.Perform a full load migration without change data capture
E.Use AWS Database Migration Service (DMS) for ongoing replication
AnswersA, C, E

Secure connectivity is essential.

Why this answer

Establishing a VPN or Direct Connect provides a secure, private, and low-latency network connection between the on-premises environment and AWS. This is essential for minimizing downtime during a migration, as it ensures reliable and fast data transfer for both the initial full load and ongoing replication, reducing the risk of network interruptions that could extend the migration window.

Exam trap

The trap here is that candidates often think disabling archiving simplifies the migration, but they miss that CDC requires archived logs for minimal downtime, and they may also assume a full load alone is sufficient without realizing it forces a longer outage to ensure data consistency.

383
MCQmedium

A data engineer applies the bucket policy shown in the exhibit to an S3 bucket. The bucket contains sensitive data that must be encrypted at rest and accessed only over HTTPS. Which of the following statements is true?

A.The policy allows both HTTP and HTTPS access.
B.The policy allows anonymous access to list objects in the bucket.
C.The policy enforces that all PutObject requests must include the x-amz-server-side-encryption header with value AES256.
D.The policy requires the use of AWS KMS for server-side encryption.
AnswerC

The policy's `StringNotEquals` condition on `s3:x-amz-server-side-encryption` with value `AES256` denies any PutObject lacking that header, so uploads without SSE-AES256 encryption are rejected. This directly satisfies the stem's encryption-at-rest constraint, though it does not address the HTTPS-only requirement.

Why this answer

The bucket policy includes a condition that denies PutObject requests unless the `s3:x-amz-server-side-encryption` header is present and set to `AES256`. This enforces server-side encryption with S3-managed keys (SSE-S3) for all uploads, ensuring data at rest is encrypted.

Exam trap

AWS often tests the distinction between SSE-S3 (`AES256`) and SSE-KMS (`aws:kms`) in bucket policy conditions, and candidates may mistakenly think the policy requires KMS when it actually specifies AES256.

How to eliminate wrong answers

Option A is wrong because the policy includes a `Deny` statement that blocks requests when `aws:SecureTransport` is `false`, which effectively denies HTTP access and allows only HTTPS. Option B is wrong because the policy does not grant any `s3:ListBucket` permission to anonymous principals; it only denies requests that fail encryption or transport conditions, but does not allow anonymous listing. Option D is wrong because the policy requires the `x-amz-server-side-encryption` header with value `AES256`, which corresponds to SSE-S3, not AWS KMS (which would require `aws:kms`).

384
MCQmedium

A data engineer needs to keep a near-real-time copy of an Amazon DynamoDB table in Amazon S3 for analytics, capturing every item-level change with the before and after images and no impact on table write latency. Which approach meets these requirements with the LEAST operational effort?

A.Enable DynamoDB point-in-time recovery and export the recovery window to S3 on a schedule.
B.Schedule an AWS Glue job every five minutes to run a full table scan and overwrite the S3 objects.
C.Enable DynamoDB Streams on the table and write an AWS Lambda function to batch stream records into S3.
D.Use AWS Database Migration Service with change data capture from DynamoDB to Amazon S3.
AnswerC

DynamoDB Streams captures an ordered log of item-level modifications, and configuring the stream view type to include both new and old images provides the before and after values. A Lambda function subscribed to the stream can batch records into S3 with no servers to manage and no added latency to table writes, since streams are written asynchronously.

Why this answer

DynamoDB Streams records every item-level change in order and can be configured to include both the old and new images of each item. A Lambda function consuming the stream writes batched records to S3 asynchronously, so table write latency is unaffected, and the serverless subscription keeps operational effort low. The alternative approaches are periodic, snapshot-based, or heavier to operate.

Exam trap

The trap here is treating point-in-time recovery or scheduled scans as change data capture, when only a stream carries every item modification with before and after images in near real time.

385
MCQmedium

A data engineer is designing a data ingestion pipeline to load millions of small JSON files from an on-premises FTP server into Amazon S3. The pipeline should minimize cost and operational overhead. Which approach is most suitable?

A.Use S3 Transfer Acceleration to upload files directly from the FTP server
B.Deploy AWS DataSync to transfer files from the FTP server to S3
C.Use AWS Snowball Edge to ship the data to AWS
D.Set up an AWS Direct Connect connection and use AWS CLI to copy files
AnswerB

AWS DataSync is designed for efficient data transfer from on-premises to AWS, handling small files well with minimal operational overhead.

Why this answer

AWS DataSync is the most suitable option because it is designed to efficiently transfer large volumes of data from on-premises storage (including FTP servers) to AWS, handling millions of small files with minimal operational overhead. It automates data transfer, retries, and validation, and it is cost-effective as you pay only for the data transferred, with no need for additional infrastructure or complex scripting.

Exam trap

The trap here is that candidates often assume S3 Transfer Acceleration is a general-purpose acceleration tool for any source, but it only accelerates the upload leg from the client to AWS and does not address the FTP-to-S3 protocol conversion or the orchestration of millions of small files.

How to eliminate wrong answers

Option A is wrong because S3 Transfer Acceleration is a feature that speeds up uploads over the internet by using AWS edge locations, but it does not handle the protocol mismatch between FTP and S3; you would still need a client to read from FTP and write to S3, and it adds cost per GB transferred without solving the file ingestion logic. Option C is wrong because AWS Snowball Edge is designed for large-scale, offline data transfers (typically terabytes to petabytes) and is overkill and costly for millions of small files; it also introduces significant latency for shipping and manual handling. Option D is wrong because AWS Direct Connect provides a dedicated network connection but does not automate the transfer of files from an FTP server; you would still need to write custom scripts using AWS CLI to copy files, increasing operational overhead and complexity.

386
MCQmedium

A data engineer is using AWS Database Migration Service (AWS DMS) to migrate an on-premises Oracle database to Amazon Aurora PostgreSQL. The migration is ongoing, and the engineer needs to ensure that changes made to the source database during the migration are replicated to the target. The engineer has set up a full load plus change data capture (CDC) task. However, after the full load completes, the CDC task fails with an error indicating that it cannot find the archive log files. What should the engineer do to resolve this issue?

A.Increase the number of AWS DMS replication instances to handle the CDC load.
B.Restart the AWS DMS task with the 'Stop task after full load completes' option enabled.
C.Configure the source Oracle database to retain archive log files for a longer period and ensure they are accessible to AWS DMS.
D.Change the target endpoint to Amazon RDS for Oracle to match the source database engine.
AnswerC

AWS DMS CDC requires access to Oracle archive log files to capture changes. If the logs are purged too quickly or not available, the CDC task fails. Increasing the retention period and ensuring accessibility allows DMS to read the logs and continue replication without interruption.

Why this answer

AWS DMS CDC for Oracle relies on reading archive log files. If these logs are not retained or accessible, the CDC task fails. The engineer must ensure the source database retains archive logs for a sufficient period and that DMS can access them, typically by configuring Oracle to keep logs and granting necessary permissions.

Exam trap

The trap here is thinking that scaling DMS resources or changing the target engine would fix a CDC error, when the error clearly points to missing archive logs on the source.

387
MCQhard

A data engineer is troubleshooting a daily batch ingestion pipeline that uses AWS Glue to read CSV files from Amazon S3 and write Parquet files to another S3 bucket. The job runs successfully but takes significantly longer than expected. The engineer notices that the input data is highly skewed with many small files. Which is the most effective optimization to reduce job duration?

A.Change the output format to JSON
B.Enable the 'groupFiles' option in the S3 source configuration
C.Increase the number of DPUs allocated to the job
D.Enable the 'use_glue_schema_registry' option
AnswerB

The 'groupFiles' option coalesces multiple small CSV files into larger partitions during the S3 read, reducing per-file overhead and task count. This directly addresses the many-small-files skew identified in the stem, cutting the job's overall duration.

Why this answer

The 'groupFiles' option in AWS Glue's S3 source configuration allows Glue to combine small files into larger splits, reducing the number of tasks and the overhead associated with processing many small files. This directly addresses the high skew and small file issue, significantly reducing job duration. Option A is incorrect because changing the output format from Parquet to JSON would likely increase file size and processing time, not reduce it.

Option C is incorrect because increasing DPUs may improve parallelism but does not solve the file-level overhead; it could even be wasteful if the bottleneck is task scheduling. Option D is incorrect because enabling the schema registry is unrelated to file grouping; it is used for schema management and validation.

388
MCQhard

A data engineer is troubleshooting a failed AWS Glue ETL job that reads from a JDBC source. The error log shows 'java.sql.SQLException: Connection timed out'. The job previously ran successfully. Which of the following is the MOST likely cause?

A.The JDBC connection string has incorrect credentials.
B.The source database schema has changed.
C.The Glue job's timeout setting is too low.
D.The security group for the source database no longer allows traffic from the Glue job's IP range.
AnswerD

A security group rule change blocking the Glue job's ENI traffic causes the JDBC connection to time out, matching the sudden failure after prior success. Network ACLs or credentials would produce different errors, so the revoked inbound rule is the likely cause.

Why this answer

The error 'Connection timed out' indicates a network-level failure, not an authentication or schema issue. Since the job previously ran successfully, the most likely cause is that the security group for the source database no longer allows inbound traffic from the Glue job's IP range. AWS Glue ETL jobs run in a VPC with elastic network interfaces, and the security group rules must permit traffic on the JDBC port (e.g., 5432 for PostgreSQL, 3306 for MySQL).

Exam trap

AWS often tests the distinction between authentication errors (wrong credentials) and network connectivity errors (timeout), and candidates may confuse the Glue job timeout setting with a network timeout.

How to eliminate wrong answers

Option A is wrong because incorrect credentials would produce an authentication error (e.g., 'Access denied for user'), not a timeout. Option B is wrong because a schema change would cause a data type mismatch or column-not-found error, not a connection timeout. Option C is wrong because the Glue job's timeout setting controls how long the job can run before being terminated, not the network connection timeout to the JDBC source.

389
MCQmedium

A data engineer must transform data in Amazon S3 using Apache Spark. The transformation logic needs to be reused across multiple AWS Glue jobs, and the engineer wants to version-control the code and run it in a serverless environment without managing clusters. Which approach should the engineer take?

A.Create an AWS Lambda function containing the transformation logic and invoke it from each AWS Glue job using boto3.
B.Package the transformation logic as a Python wheel file stored in Amazon S3, reference it in AWS Glue job parameters, and import it as a module in the job script.
C.Provision an Amazon EMR cluster, store the transformation code in a Git repository, and run Spark jobs on the cluster.
D.Create an AWS Glue job with a script that contains all transformation logic, and copy-paste the code into each job that needs it.
AnswerB

AWS Glue supports referencing additional Python modules via the --extra-py-files job parameter, which can point to a wheel file in Amazon S3. This enables code reuse across jobs, allows version control of the wheel artifact, and keeps the execution serverless. The engineer writes the transformation once, packages it, and imports it in any job that needs it.

Why this answer

AWS Glue supports referencing external Python libraries through the --extra-py-files job parameter. Packaging transformation logic as a wheel file in Amazon S3 allows the code to be versioned, reused across multiple Glue jobs, and executed in a serverless Spark environment. This satisfies both the reusability and serverless requirements without duplicating code or introducing cluster management overhead.

Exam trap

The trap here is assuming that AWS Glue cannot import external code and that all transformation logic must reside in a single job script.

390
MCQmedium

A retail company uses Amazon Kinesis Data Firehose to ingest clickstream data from its website into an Amazon S3 bucket. The data includes fields: user_id, event_type, timestamp, page_url. Recently, the data engineering team noticed that some records have malformed JSON (missing commas, extra brackets) causing delivery failures to S3. The Firehose delivery stream is configured to retry failed records for 300 seconds, after which the records are sent to an S3 bucket for failed records. The team wants to transform the data to correct malformed JSON before delivery to the main S3 bucket. They need a solution that does not require managing servers and can handle high throughput. What should the team do?

A.Configure an AWS Lambda function as a data transformation in Kinesis Data Firehose to correct malformed JSON.
B.Set up an Amazon EMR cluster with Apache Spark to process the data in micro-batches and fix JSON errors.
C.Use an AWS Glue streaming ETL job to read from Firehose and write corrected data to S3.
D.Use Amazon Kinesis Data Analytics with a SQL application to parse and fix JSON.
AnswerA

Lambda transformation runs inline within Firehose, letting you repair malformed JSON before delivery to the main S3 bucket. It is fully serverless and scales with throughput, satisfying the no-server-management and high-volume constraints while preventing records reaching the failure bucket.

Why this answer

Kinesis Data Firehose supports Lambda-based data transformation natively: you attach a Lambda function to the delivery stream, and Firehose invokes it on each batch before delivery. This is serverless, scales with throughput, and lets you parse and repair malformed JSON before it reaches the main S3 bucket.

Exam trap

DEA-C01 often tests whether candidates know Firehose's built-in Lambda transformation versus external processing services — candidates pick Glue or EMR for a problem that Firehose solves natively.

How to eliminate wrong answers

Option B is wrong because EMR requires managing clusters (even with managed scaling) and is overkill for simple per-record JSON repair; it also adds latency and cost. Option C is wrong because Glue streaming ETL reads from Kinesis Data Streams, not directly from Firehose, and would require re-architecting the ingestion path. Option D is wrong because Kinesis Data Analytics (now Managed Service for Apache Flink) processes streams with SQL/Flink but is not the Firehose-native transformation mechanism and adds operational complexity.

391
Drag & Dropmedium

Arrange the steps to create an AWS Glue job that transforms data from Amazon S3 to Amazon Redshift in the correct order.

Drag or tap steps into the slots.

Steps
Order
1Step 1
2Step 2
3Step 3
4Step 4

Why this order

First, catalog the source data with a crawler. Then, prepare the ETL script. Configure the job with connections, run it, and finally verify the results in Redshift.

392
Multi-Selectmedium

A data engineer needs to ensure that an AWS Glue ETL job can access an Amazon S3 bucket that is encrypted with SSE-KMS. The Glue job runs with an IAM role. The KMS key policy grants access to the account root. Which TWO actions are required to allow the Glue job to read and write data in the bucket? (Choose two.)

Select 2 answers
A.Configure the S3 bucket to use SSE-S3 instead of SSE-KMS.
B.Add s3:GetObject and s3:PutObject permissions to the Glue job's IAM role.
C.Attach an S3 bucket policy that allows the Glue job role to perform s3:GetObject and s3:PutObject.
D.Add kms:Decrypt and kms:GenerateDataKey permissions to the Glue job's IAM role.
E.Add kms:CreateGrant permission to the Glue job's IAM role.
AnswersB, D

The Glue job's IAM role must have s3:GetObject and s3:PutObject permissions to read and write objects in the S3 bucket. These permissions allow the job to perform the necessary S3 operations. Without them, the job cannot access the bucket regardless of KMS permissions. Together with KMS permissions, they enable full access to the encrypted data.

Why this answer

To allow an AWS Glue job to access an S3 bucket encrypted with SSE-KMS, the job's IAM role must have both S3 permissions (s3:GetObject, s3:PutObject) and KMS permissions (kms:Decrypt, kms:GenerateDataKey). These permissions enable the job to read and write objects and to use the KMS key for encryption and decryption. Other options either do not provide the necessary permissions or would change the encryption method, which is not desired.

Exam trap

The trap here is assuming that S3 permissions alone are sufficient for accessing SSE-KMS encrypted objects, overlooking the need for KMS permissions.

393
MCQeasy

A company stores its application logs in an Amazon S3 bucket. The logs are accessed frequently for the first 30 days, after which they are rarely accessed but must be retained for 7 years for compliance. The company wants to optimize storage costs while maintaining immediate retrieval availability for the first 30 days and the ability to retrieve logs within 12 hours after that. Which lifecycle policy should the data engineer configure?

A.Delete objects after 30 days to minimize storage costs.
B.Transition objects to S3 Standard-IA after 30 days and then to S3 Glacier Deep Archive after 1 year.
C.Transition objects to S3 One Zone-IA after 30 days and delete after 7 years.
D.Transition objects to S3 Glacier Flexible Retrieval after 30 days and delete after 7 years.
AnswerB

Standard-IA provides immediate retrieval for the first 30 days, then Deep Archive for cost-effective long-term retention.

Why this answer

It uses S3 Standard-IA for the first 30 days (frequent access, immediate retrieval) and then transitions to S3 Glacier Deep Archive after 1 year, which provides retrieval within 12 hours at the lowest cost for long-term retention. This meets the compliance requirement of 7-year retention while optimizing costs by moving data to progressively cheaper storage classes based on access patterns.

Exam trap

The DEA-C01 exam often tests the misconception that S3 Glacier Flexible Retrieval is the cheapest option for long-term archival, but S3 Glacier Deep Archive is significantly cheaper for data that is rarely accessed and can tolerate a 12-hour retrieval time.

How to eliminate wrong answers

Option A is wrong because deleting objects after 30 days violates the 7-year compliance retention requirement. Option C is wrong because S3 One Zone-IA does not provide the durability (99.999999999% vs 99.9999999999%) or availability needed for compliance data, and it lacks the 12-hour retrieval capability required after 30 days. Option D is wrong because S3 Glacier Flexible Retrieval has a retrieval time of minutes to hours (typically 1-5 minutes for expedited, 3-5 hours for standard), but the requirement is for retrieval within 12 hours, which is met; however, transitioning directly to Glacier Flexible Retrieval after 30 days is more expensive than using Standard-IA first, and the option does not include a transition to Deep Archive for further cost optimization over 7 years.

394
Multi-Selectmedium

A company uses AWS Glue to transform data in S3. The Glue job fails with memory errors. Which THREE actions can help resolve this?

Select 3 answers
A.Optimize the transformation to use pushdown predicates.
B.Use a larger worker type (e.g., G.2X).
C.Increase the number of DPUs.
D.Increase the job timeout.
E.Decrease the number of DPUs.
AnswersA, B, C

Pushdown predicates filter rows in the source system before loading, so fewer records reach the Glue workers, directly easing the memory pressure causing the failures. This reduces the data volume shuffled and held in memory during transformation, satisfying the stem's constraint of resolving out-of-memory errors without simply adding capacity.

Why this answer

Options A, B, and C are correct. Using pushdown predicates (A) reduces the amount of data read by filtering at the data source, which can alleviate memory pressure. Using a larger worker type (B), such as G.2X, increases the memory available per worker, directly addressing out-of-memory errors.

Increasing the number of DPUs (C) adds more workers, distributing the memory load. Option D (increasing job timeout) does not solve memory issues, and Option E (decreasing DPUs) would reduce available memory, making the problem worse.

395
MCQeasy

A data engineer needs to grant an IAM user the ability to view Amazon CloudWatch Logs log groups and stream log events from a specific log group. Which IAM policy action should be used?

A.logs:DescribeLogGroups and logs:GetLogEvents
B.logs:PutLogEvents
C.logs:CreateLogGroup
D.logs:DeleteLogGroup
AnswerA

The `logs:DescribeLogGroups` action permits listing and viewing log groups, while `logs:GetLogEvents` retrieves individual log events from a specified log group. Together they satisfy both stated requirements: viewing CloudWatch Logs log groups and streaming log events from the specific log group named in the stem.

Why this answer

Viewing log groups requires logs:DescribeLogGroups, and reading/streaming log events from a specific group requires logs:GetLogEvents (or logs:FilterLogEvents for filtered reads). Together these two actions satisfy the read-only requirement. They map directly to the CloudWatch Logs API calls DescribeLogGroups and GetLogEvents.

Exam trap

DEA-C01 often tests the confusion between read actions (DescribeLogGroups, GetLogEvents, FilterLogEvents) and write/admin actions (PutLogEvents, CreateLogGroup, DeleteLogGroup), catching candidates who pick PutLogEvents thinking it 'streams' logs to a viewer.

How to eliminate wrong answers

Option B is wrong because logs:PutLogEvents is a write action that uploads log events to a stream — it grants no read capability and is used by agents/SDKs sending logs. Option C is wrong because logs:CreateLogGroup only creates new log groups and provides no visibility into existing groups or their events. Option D is wrong because logs:DeleteLogGroup is a destructive administrative action that removes log groups entirely, the opposite of read-only access.

396
MCQhard

A data engineer is designing a real-time analytics solution using Amazon DynamoDB. The workload requires capturing all changes to a DynamoDB table and processing them in near-real-time to update a materialized view in Amazon Redshift. Which approach should the engineer use to capture and process the changes?

A.Enable DynamoDB global tables and use Amazon Kinesis Data Streams to replicate changes to Redshift.
B.Use AWS Database Migration Service (AWS DMS) with ongoing replication from DynamoDB to Redshift.
C.Use DynamoDB Accelerator (DAX) to cache changes and periodically export to Redshift.
D.Enable DynamoDB Streams and use AWS Lambda to process the stream and write to Redshift.
AnswerD

DynamoDB Streams captures item-level changes in near-real-time. AWS Lambda can be triggered by the stream to process records and write to Redshift. This serverless approach is scalable and requires minimal operational overhead. It is the recommended pattern for real-time change data capture from DynamoDB.

Why this answer

DynamoDB Streams captures item-level changes in near-real-time, and AWS Lambda can process these changes and write to Amazon Redshift. This serverless architecture is scalable, cost-effective, and requires minimal operational overhead. It is the standard pattern for change data capture from DynamoDB to other data stores.

Exam trap

The trap here is confusing DynamoDB Streams with other features like global tables or DAX, which serve different purposes such as multi-region replication or caching, and do not provide change data capture for real-time processing.

397
MCQmedium

A data engineer is tasked with designing a disaster recovery solution for a data lake stored in Amazon S3. The data lake contains sensitive customer data that must be replicated to a different AWS Region. The engineer needs to ensure that all objects, including those with encryption using SSE-KMS, are replicated. Which solution meets the requirements?

A.Use S3 Batch Operations to copy objects to the destination bucket.
B.Enable S3 Cross-Region Replication (CRR) with the appropriate KMS key and IAM role.
C.Use S3 Transfer Acceleration to copy objects across regions.
D.Use the AWS CLI s3 sync command scheduled in a cron job.
AnswerB

CRR replicates objects across Regions, but SSE-KMS-encrypted objects require the destination bucket's KMS key permissions and an IAM role allowing decrypt and encrypt. Configuring both satisfies the requirement that all objects, including KMS-encrypted ones, replicate successfully.

Why this answer

S3 Cross-Region Replication (CRR) with the appropriate KMS key and IAM role is the correct solution because it automatically replicates objects, including those encrypted with SSE-KMS, to a destination bucket in another region. To replicate SSE-KMS encrypted objects, you must specify a KMS key in the destination region and grant the necessary IAM permissions to the replication role. This meets the requirement for disaster recovery of sensitive data.

Exam trap

DEA-C01 often tests the misconception that S3 Transfer Acceleration or Batch Operations can serve as replication solutions, when in fact CRR is the only automated, continuous replication feature.

How to eliminate wrong answers

Option A is wrong because S3 Batch Operations can copy objects but does not provide continuous replication and requires manual initiation; it is not a disaster recovery solution with automatic replication. Option C is wrong because S3 Transfer Acceleration speeds up uploads to S3 over long distances but does not replicate data across regions; it is for accelerating transfers, not for disaster recovery. Option D is wrong because using AWS CLI s3 sync in a cron job is a manual, scripted approach that lacks the automation, reliability, and metadata preservation of CRR; it also does not handle SSE-KMS encryption seamlessly.

398
MCQmedium

A financial services company stores transaction records in an Amazon DynamoDB table. An audit requires that all data older than 7 years be automatically and permanently deleted. The data engineering team must implement this with minimal operational overhead and no application code changes. What should the team do?

A.Configure a TTL attribute on the table with an expiry timestamp set to 7 years from the transaction date.
B.Enable point-in-time recovery and restore the table to a point before the 7-year window, then delete the original table.
C.Create a scheduled Amazon EventBridge rule that invokes an AWS Step Functions state machine to scan and delete old items.
D.Enable DynamoDB Streams and write an AWS Lambda function that deletes items older than 7 years.
AnswerA

DynamoDB Time to Live (TTL) lets you define an attribute holding an expiration timestamp; DynamoDB automatically deletes expired items at no extra cost, with no code or infrastructure to manage. Setting the attribute to transaction date plus 7 years satisfies the audit requirement and requires only a one-time schema and write-path change, not ongoing operations.

Why this answer

DynamoDB TTL is the native, serverless mechanism for automatic item expiration. By storing an epoch timestamp attribute set to the transaction date plus seven years, DynamoDB deletes expired items in the background without consuming write capacity or requiring custom code. It directly satisfies the audit's permanent deletion requirement with minimal operational effort, unlike stream-based, step-function, or restore-based approaches.

Exam trap

The trap here is assuming that DynamoDB Streams or point-in-time recovery can enforce time-based retention, when only TTL provides automatic, attribute-driven item expiration.

399
MCQhard

A company needs to ingest real-time clickstream data from a web application into Amazon Redshift with minimal latency. The data volume is high and requires processing before loading. Which architecture is MOST appropriate?

A.AWS Glue ETL jobs scheduled every 5 minutes -> Redshift
B.S3 -> Lambda -> Redshift
C.DynamoDB Streams -> Lambda -> Redshift
D.Kinesis Data Streams -> Kinesis Data Firehose -> Redshift
AnswerD

Provides real-time ingestion with transformation capability.

Why this answer

D is correct because Kinesis Data Streams captures high-volume clickstream data in real time, and Kinesis Data Firehose can buffer, transform (e.g., with Lambda), and load the data directly into Amazon Redshift with near-zero latency. This architecture is purpose-built for streaming ingestion with minimal overhead, unlike batch or intermediary storage approaches.

Exam trap

The trap here is that candidates often confuse 'real-time' with 'near-real-time' and choose a batch option like Glue (A) or an indirect streaming path like S3 -> Lambda (B), failing to recognize that Kinesis Data Firehose is the only AWS service that natively integrates streaming ingestion with Redshift without additional latency or complexity.

How to eliminate wrong answers

Option A is wrong because AWS Glue ETL jobs scheduled every 5 minutes introduce batch latency, which violates the 'minimal latency' requirement for real-time clickstream data. Option B is wrong because S3 -> Lambda -> Redshift requires Lambda to write to Redshift, which is inefficient for high-volume streaming data due to Lambda's invocation limits and lack of native streaming buffering, plus S3 adds an unnecessary intermediate storage hop. Option C is wrong because DynamoDB Streams are designed for change data capture from DynamoDB tables, not for ingesting raw clickstream data from a web application; it would require an additional service to capture the data into DynamoDB first, adding complexity and latency.

400
MCQmedium

A data engineer is configuring an AWS Glue crawler to catalog data stored in an Amazon S3 bucket. The data is partitioned by year, month, and day in a Hive-style structure (for example, s3://bucket/data/year=2023/month=01/day=15/). The engineer wants the crawler to recognize the partitions and add them to the AWS Glue Data Catalog. What should the engineer do?

A.Manually create a partitioned table in the AWS Glue Data Catalog using the AWS CLI before running the crawler.
B.Configure the crawler with the 'Add new columns only' option and enable partition detection.
C.Ensure the crawler's include path points to the bucket root and that partition detection is enabled in the crawler configuration.
D.Create a separate crawler for each partition prefix and run them individually.
AnswerC

AWS Glue crawlers automatically detect Hive-style partitions when the include path points to the root of the partitioned data and partition detection is enabled. The crawler parses the key=value structure in the S3 prefixes and adds partition metadata to the Data Catalog. This allows queries in Athena and Redshift Spectrum to use partition pruning for efficiency.

Why this answer

AWS Glue crawlers can automatically detect Hive-style partitions when the include path covers the partitioned data and partition detection is enabled. The crawler reads the key=value structure in the S3 prefixes, creates partition metadata, and updates the Data Catalog. This enables partition pruning in query engines like Amazon Athena and Redshift Spectrum.

Exam trap

The trap here is assuming that partition detection is automatic regardless of crawler configuration or include path scope.

401
Multi-Selectmedium

A data engineer is building an Amazon Redshift data warehouse that ingests large staged files from Amazon S3 using the COPY command. The team wants to maximize load performance and minimize the time spent on ingestion. Which TWO practices should the engineer apply? (Choose two.)

Select 2 answers
A.Load compressed columnar files such as Parquet in a single COPY statement
B.Run VACUUM SORT ONLY on the target table immediately after each COPY
C.Use a single large compressed file to reduce the number of S3 GET requests
D.Split the input data into multiple files sized roughly equal and load them in parallel
E.Add a sort key to the target table on the timestamp column before loading
AnswersA, D

COPY supports columnar formats like Parquet and ORC, and columnar compression reduces the bytes transferred from S3 while allowing column-level pruning during load. Loading multiple Parquet files in one COPY statement lets Redshift distribute them across slices in parallel. Combining columnar compression with parallel file distribution reduces I/O and CPU work, directly improving ingestion throughput and shortening load windows.

Why this answer

COPY performance in Redshift depends on parallelism and data volume. Splitting input into multiple similarly sized files lets every slice read concurrently, and using compressed columnar formats such as Parquet reduces bytes moved and enables column pruning. Together these practices shorten load time.

Sort keys, single-file loads, and post-load VACUUM operations affect query performance or add overhead rather than accelerating ingestion, so they do not belong in this optimization.

Exam trap

The trap here is equating query-performance tuning such as sort keys and VACUUM with load-performance tuning, when COPY speed is driven primarily by parallel file distribution and compressed columnar formats.

402
MCQmedium

A company stores sensitive data in an Amazon S3 bucket. To comply with regulations, all data must be encrypted at rest using server-side encryption. The security team wants to ensure that any attempt to upload an unencrypted object is automatically denied. Which S3 bucket policy condition should be used?

A.s3:x-amz-server-side-encryption-aws-kms-key-id
B.s3:x-amz-acl
C.s3:x-amz-server-side-encryption
D.s3:x-amz-storage-class
AnswerC

The `s3:x-amz-server-side-encryption` condition key inspects the `x-amz-server-side-encryption` request header, letting the bucket policy deny any PutObject lacking a server-side encryption algorithm. This directly enforces the stem's requirement that unencrypted uploads be automatically rejected, since requests without that header fail the condition.

Why this answer

The bucket policy condition key s3:x-amz-server-side-encryption matches the x-amz-server-side-encryption request header that S3 requires on PUT requests when server-side encryption is specified. By using a Deny effect with a StringNotEquals condition on this key (or Null check), any PutObject request that does not include the encryption header is rejected before the object is written. This enforces encryption at rest at the API layer, regardless of client behavior.

Exam trap

DEA-C01 often tests the confusion between default bucket encryption (which silently encrypts) and a bucket policy condition that explicitly denies unencrypted PUTs — candidates pick the KMS key ID condition thinking it enforces encryption, but only the base x-amz-server-side-encryption key does.

How to eliminate wrong answers

Option A is wrong because s3:x-amz-server-side-encryption-aws-kms-key-id only validates which KMS key is used when SSE-KMS is specified — it does not require that any encryption header be present, so an unencrypted PUT would still pass. Option B is wrong because s3:x-amz-acl controls canned ACLs (private, public-read, etc.) and has nothing to do with encryption enforcement. Option D is wrong because s3:x-amz-storage-class governs storage tier selection (STANDARD, INTELLIGENT_TIERING, GLACIER) and does not enforce encryption.

403
MCQeasy

A data engineer is migrating an on-premises PostgreSQL database to Amazon RDS for PostgreSQL. The database is 2 TB in size. The engineer needs to minimize downtime. Which AWS service should be used for the migration?

A.AWS Data Pipeline
B.AWS Database Migration Service (DMS)
C.AWS Snowball
D.Amazon S3
AnswerB

DMS supports continuous replication with minimal downtime.

Why this answer

AWS Database Migration Service (DMS) is the correct choice because it supports continuous replication (change data capture) from an on-premises PostgreSQL source to Amazon RDS for PostgreSQL, enabling near-zero downtime migration. DMS can handle a 2 TB database by using a large replication instance and tuning task settings, and it automatically converts the source schema to the target RDS engine.

Exam trap

The trap here is that candidates often choose AWS Snowball for large databases, mistakenly thinking physical transfer is faster, but they overlook that Snowball requires stopping writes to the source database during the export and shipping process, causing unacceptable downtime for a live migration.

How to eliminate wrong answers

Option A is wrong because AWS Data Pipeline is a batch-oriented workflow orchestration service for moving and transforming data between AWS services, but it does not support live, ongoing replication or schema conversion for database migrations, making it unsuitable for minimizing downtime. Option C is wrong because AWS Snowball is a physical data transfer device designed for large-scale offline data movement (e.g., petabyte-scale), but it introduces significant downtime due to shipping and manual transfer, and it cannot perform continuous replication for a live migration. Option D is wrong because Amazon S3 is an object storage service and cannot directly migrate a live PostgreSQL database to RDS; it would require an intermediate export/import process that causes extended downtime and lacks native change data capture.

404
Matchingmedium

Match each AWS database service to its primary use case.

Drag a concept onto its matching description — or click a concept then click the description.

Concepts
Matches

Relational database with managed operations

NoSQL key-value and document database

In-memory caching for low latency

Graph database for connected data

Time-series data for IoT and analytics

Why these pairings

Correct matches: RDS -> OLTP relational, DynamoDB -> NoSQL low-latency, Redshift -> data warehousing, ElastiCache -> in-memory caching. Common confusions include swapping RDS and DynamoDB use cases.

405
MCQmedium

A team is designing a data lake on S3 and needs to enforce encryption at rest. They want to use server-side encryption with a KMS key that they manage. Which encryption option should they configure on the S3 bucket?

A.SSE-KMS
B.Client-side encryption
C.SSE-S3
D.SSE-C
AnswerA

SSE-KMS encrypts objects at rest using keys held in AWS KMS, and the customer controls those keys through key policies and grants. This satisfies the requirement for server-side encryption with a customer-managed KMS key on the S3 bucket.

Why this answer

SSE-KMS is the correct choice because it provides server-side encryption using a customer-managed KMS key. This allows the team to enforce encryption at rest with their own key, giving them control over key rotation, access policies, and audit trails via AWS CloudTrail, which aligns with the requirement to manage the encryption key themselves.

Exam trap

The trap here is that candidates often confuse SSE-S3 with SSE-KMS, assuming both use customer-managed keys, but SSE-S3 uses AWS-managed keys and does not provide the customer with key management control or audit capabilities.

How to eliminate wrong answers

Option B (Client-side encryption) is wrong because it encrypts data before it is sent to S3, not at rest on the server side, and does not involve configuring encryption on the S3 bucket itself. Option C (SSE-S3) is wrong because it uses an AWS-managed key, not a customer-managed KMS key, so the team would not have control over key management. Option D (SSE-C) is wrong because it requires the customer to provide their own encryption keys in each request, and the bucket configuration does not manage the key; instead, the key is supplied per-object, which is not a bucket-level encryption setting.

406
MCQhard

A company is using AWS Database Migration Service (DMS) to migrate a 2 TB MySQL database to Amazon Aurora MySQL. The migration is taking longer than expected. The source database is in a different AWS region. Which change would MOST likely improve the migration speed?

A.Use a smaller DMS replication instance to reduce costs.
B.Use a Multi-AZ deployment for the DMS replication instance in the target region.
C.Increase the number of parallel tables being migrated.
D.Disable binary logging on the source MySQL database.
AnswerC

DMS parallel load splits large tables into concurrent threads, so raising the number of tables migrated simultaneously uses more of the available bandwidth and replication capacity. This directly addresses the slow 2 TB transfer across regions, where serial table processing underuses throughput.

Why this answer

DMS migrates tables in parallel, and the number of tables loaded concurrently is controlled by the replication task's parallel load settings. Increasing the number of parallel tables (or enabling parallel load with a higher thread count) directly increases throughput when the bottleneck is per-table serialization, which is the most likely cause of a slow 2 TB MySQL-to-Aurora migration. This is the change most likely to improve speed without altering the source or target architecture.

Exam trap

DEA-C01 often tests whether candidates blame network or instance size for slow DMS migrations when the real bottleneck is DMS's default serialized table loading, leading them to pick instance or Multi-AZ changes instead of parallel load tuning.

How to eliminate wrong answers

Option A is wrong because a smaller replication instance reduces CPU, memory, and network capacity, which would slow the migration further rather than speed it up. Option B is wrong because Multi-AZ on the replication instance improves availability, not throughput; the standby does not participate in the migration workload. Option D is wrong because disabling binary logging on the source MySQL database would break ongoing replication (CDC) and is not a supported or safe way to accelerate a DMS migration; binary logs are required for change data capture.

407
Multi-Selectmedium

A company uses AWS CloudTrail to log all API calls. The security team wants to ensure that log files are tamper-proof and cannot be deleted. Which TWO actions should the data engineer take? (Choose TWO.)

Select 2 answers
A.Enable CloudTrail log file validation
B.Enable S3 Object Lock on the S3 bucket
C.Enable MFA Delete on the S3 bucket
D.Enable S3 Versioning on the S3 bucket
E.Enable SSE-KMS encryption on the S3 bucket
AnswersA, B

Enabling CloudTrail log file validation generates a digest file for each log, letting you verify integrity via the AWS CLI. This satisfies the tamper-proof constraint: any modification or deletion of delivered log files is detectable through hash comparison, though it does not itself prevent deletion.

Why this answer

Option A is correct because enabling CloudTrail log file validation generates a digitally signed digest file for each log file, allowing you to verify that logs have not been tampered with or modified after delivery. Option B is correct because S3 Object Lock enforces WORM (Write Once Read Many) protection, preventing log files from being deleted or overwritten for a specified retention period, which directly satisfies the tamper-proof and non-deletable requirement. Option C is not correct because MFA Delete only requires multi-factor authentication for deleting objects or changing versioning state; it adds a control but does not make logs inherently tamper-proof or prevent deletion by an authorized MFA holder.

Option D is not correct because S3 Versioning merely preserves prior versions of objects, so deleted or altered logs can still be removed and versioning alone does not prevent deletion. Option E is not correct because SSE-KMS encryption protects data confidentiality at rest but does not prevent log files from being modified or deleted.

408
MCQhard

A data engineer is building an AWS Glue Data Catalog table that references an Amazon S3 bucket containing CSV files. The security team requires that column-level access be restricted so that only specific IAM principals can view the column containing personally identifiable information (PII). The engineer needs to implement this restriction without modifying the underlying data. Which combination of actions should the engineer take?

A.Create an IAM policy that denies access to the S3 prefix containing the PII column.
B.Use AWS Glue Studio to create a transform that masks the PII column and write a new table.
C.Apply an S3 bucket policy that restricts access to objects based on object tags.
D.Use AWS Lake Formation to define column-level permissions on the table and grant access to the specific IAM principals.
AnswerD

AWS Lake Formation provides fine-grained access control, including column-level permissions, on Data Catalog tables. By registering the S3 location with Lake Formation and defining column-level grants, the engineer can restrict access to the PII column for specific principals. This meets the requirement without altering the data and is the recommended approach for column-level security in Glue Data Catalog.

Why this answer

AWS Lake Formation is designed for fine-grained access control on Data Catalog tables, including column-level permissions. It allows granting or denying access to specific columns for specific IAM principals without changing the data. IAM policies, S3 bucket policies, and Glue transforms operate at different levels and cannot achieve column-level restriction on the existing table.

Exam trap

The trap here is confusing object-level S3 permissions or data masking with column-level access control, which requires Lake Formation.

409
MCQeasy

A company needs to ingest data from a relational database into Amazon S3 for analytics. The database is an Amazon RDS MySQL instance. Which AWS service should be used for a one-time historical data load?

A.AWS Database Migration Service (DMS)
B.AWS Glue ETL
C.Amazon Athena
D.Amazon Kinesis Data Firehose
AnswerA

AWS DMS performs bulk migrations and can replicate existing table data from Amazon RDS MySQL into Amazon S3, making it suited to a one-time historical load. It reads source data directly without custom extraction code.

Why this answer

AWS Database Migration Service (DMS) is designed for migrating data from relational databases like Amazon RDS MySQL into targets such as Amazon S3, supporting one-time full loads as well as ongoing replication. For a one-time historical data load, DMS can extract the existing data and land it in S3 efficiently. This makes DMS the correct service for the described scenario.

Exam trap

DEA-C01 often tests the distinction between DMS (database migration/ingestion) and Glue (ETL transformation), catching candidates who pick Glue for a straightforward one-time database-to-S3 load.

How to eliminate wrong answers

Option B is wrong because AWS Glue ETL is a serverless ETL service for transforming data, but it is not the primary tool for bulk database migration; it would require custom connectors and more setup for a one-time RDS-to-S3 load. Option C is wrong because Amazon Athena is a query service that reads data in S3, not a data ingestion or migration tool. Option D is wrong because Kinesis Data Firehose is for streaming data delivery to destinations like S3, not for extracting historical data from a relational database.

410
Multi-Selecthard

A data engineer is optimizing an Amazon Redshift cluster for a workload that includes frequent complex queries with multiple joins and aggregations. The engineer wants to improve query performance by using appropriate distribution styles and sort keys. Which TWO actions should the engineer take? (Choose two.)

Select 2 answers
A.Create a materialized view for each complex query.
B.Set the distribution style of small dimension tables to ALL.
C.Use EVEN distribution for all tables to ensure uniform data distribution.
D.Apply a sort key on the column used in the WHERE clause of frequent queries.
E.Set the distribution style of large fact tables to KEY on the join column.
AnswersB, E

ALL distribution replicates small dimension tables to every node, eliminating the need to redistribute data during joins. This is effective for small tables that are frequently joined with large fact tables. It reduces network overhead and improves join performance, but should only be used for tables that are small enough to fit comfortably in node memory and storage.

Why this answer

Setting the distribution style of large fact tables to KEY on the join column collocates matching rows, reducing data movement during joins. Setting small dimension tables to ALL replicates them to all nodes, eliminating redistribution. Together, these actions minimize network traffic and accelerate complex join queries.

They are fundamental best practices for optimizing Amazon Redshift performance in join-heavy workloads.

Exam trap

The trap here is assuming that sort keys or materialized views are the primary optimizations for join performance, when distribution styles have a more direct impact on data movement during joins.

411
MCQhard

A data engineer needs to share a dataset stored in an Amazon S3 bucket with another AWS account. The dataset must remain encrypted at rest using AWS KMS. The data engineer creates a bucket policy that grants the other account access to the bucket. However, the other account reports that objects appear encrypted and they cannot decrypt them. What is the most likely cause?

A.The KMS key policy does not grant the other account the kms:Decrypt permission
B.The bucket policy does not grant the s3:GetObject permission
C.The other account must use the same KMS key to upload objects
D.The objects are encrypted with SSE-S3, which is not supported for cross-account access
AnswerA

Cross-account access requires permissions on both the S3 bucket policy and the KMS key policy. The bucket policy alone grants object access, but decryption fails because the KMS key policy does not authorise the other account's principals to call kms:Decrypt on that customer-managed key.

Why this answer

For cross-account access to KMS-encrypted S3 objects, both the S3 bucket policy and the KMS key policy must grant the necessary permissions. The bucket policy grants s3:GetObject, but decryption requires kms:Decrypt on the KMS key. If the key policy does not allow the other account, the objects cannot be decrypted, even with S3 access.

Exam trap

DEA-C01 often tests the misconception that S3 bucket policies alone are sufficient for cross-account access to KMS-encrypted objects, ignoring the need for KMS key policy permissions.

How to eliminate wrong answers

Option B is wrong because the scenario states the bucket policy grants access, and the objects appear encrypted, indicating S3 access is working but decryption is failing. Option C is wrong because the other account does not need to use the same KMS key to upload; they need decrypt permission to read. Option D is wrong because SSE-S3 does not use KMS and is not the issue here; the objects are encrypted with KMS (SSE-KMS), and cross-account access is supported with proper key policy.

412
MCQhard

A company has an Amazon Redshift cluster with a mix of frequently accessed hot data and rarely accessed cold data. They want to reduce storage costs without affecting query performance for the hot data. Which strategy is MOST effective?

A.Use RA3 nodes with managed storage to automatically offload cold data to Amazon S3.
B.Reduce the number of nodes and increase the number of slices.
C.Create external tables in Redshift Spectrum to query cold data in S3.
D.Use Dense Compute nodes and unload cold data to Amazon S3 manually.
AnswerA

RA3 nodes separate compute from Redshift Managed Storage, automatically moving cold blocks to Amazon S3 while hot data stays on local SSD. This lowers storage cost without degrading hot-data query performance, and needs no manual data movement.

Why this answer

RA3 nodes with managed storage automatically separate compute and storage, offloading cold data to Amazon S3 while keeping hot data on local SSD for fast queries. This reduces storage costs without manual intervention or affecting hot data performance.

Exam trap

The trap here is that candidates may choose Redshift Spectrum (Option C) thinking it automatically offloads cold data, but Spectrum requires manual external table creation and does not integrate with the cluster's automatic storage tiering.

How to eliminate wrong answers

Option B is wrong because reducing nodes and increasing slices does not address cold data storage; it changes cluster configuration without reducing storage costs for cold data. Option C is wrong because creating external tables in Redshift Spectrum allows querying cold data in S3 but does not automatically offload cold data from the cluster; it requires manual data movement and schema management. Option D is wrong because Dense Compute nodes are compute-optimized and do not support managed storage offloading; manually unloading cold data to S3 adds operational overhead and does not leverage automatic tiering.

413
MCQmedium

A data engineer is using AWS Glue ETL to transform data from an S3 data lake. The job fails with a memory error. Which approach should be used to resolve this issue without major code changes?

A.Rewrite the ETL script in PySpark instead of Scala
B.Change the input file format from CSV to Parquet
C.Increase the number of DPUs allocated to the Glue job
D.Use Amazon EMR instead of AWS Glue
AnswerC

Increasing the number of DPUs allocated to the Glue job directly increases memory and parallelism, which helps resolve memory errors without major code changes.

Why this answer

Increasing the number of DPUs (Data Processing Units) allocated to the Glue job provides more memory and parallelism. Option A is wrong because rewriting in PySpark is a major code change. Option B is wrong because using a smaller file format may not address memory issues.

Option D is wrong because using a different service is unnecessary.

414
MCQeasy

A company uses Amazon Kinesis Data Firehose to deliver data to an Amazon S3 bucket. The data delivery is delayed by up to 5 minutes. The engineer wants to reduce the delay to under 1 minute. Which parameter should be adjusted?

A.Enable error logging to CloudWatch.
B.Increase the buffer size in Kinesis Data Firehose.
C.Enable data compression.
D.Decrease the buffer interval in Kinesis Data Firehose.
AnswerD

Firehose buffers incoming records and delivers them when the buffer interval elapses or the buffer size fills. Lowering the buffer interval shortens the maximum wait, reducing delivery latency below one minute, whereas the size parameter alone cannot guarantee that timing.

Why this answer

Decreasing the buffer interval reduces the time Kinesis Data Firehose waits before delivering a batch, thus lowering latency to under 1 minute. Option A is incorrect because error logging to CloudWatch does not affect delivery timing. Option B is incorrect because increasing the buffer size would actually increase the delay as Firehose waits for more data to accumulate.

Option C is incorrect because enabling data compression reduces storage size but has no impact on delivery frequency.

415
MCQmedium

A company uses Amazon Kinesis Data Firehose to deliver streaming data to an Amazon S3 bucket. The data is then processed by a scheduled AWS Glue ETL job that loads it into an Amazon Redshift table. Recently, the Glue job has been failing with the error: 'S3ServiceException: Access Denied'. The Firehose delivery stream is configured with a prefix and error logging to the same S3 bucket. The Glue job uses the same IAM role that has s3:GetObject and s3:ListBucket permissions on the bucket. What is the most likely cause?

A.The Glue job expects a different data format than what Firehose writes.
B.The Glue job's IAM role does not have s3:GetObjectVersion permission.
C.The Glue job is using the wrong IAM role that does not have permissions to the S3 bucket.
D.The S3 bucket has default encryption enabled with AWS KMS (SSE-KMS), and the Glue job's IAM role lacks kms:Decrypt permission.
AnswerD

With SSE-KMS default encryption, reading objects requires kms:Decrypt in addition to s3:GetObject. The Glue role holds only S3 permissions, so decryption fails and surfaces as Access Denied, despite the bucket and prefix configuration being correct.

Why this answer

The Glue job's IAM role has s3:GetObject and s3:ListBucket, which are sufficient for reading objects from S3 in the absence of encryption. However, when the bucket uses SSE-KMS default encryption, reading an object requires the kms:Decrypt permission on the KMS key in addition to s3:GetObject. The 'Access Denied' error is the classic symptom of a missing kms:Decrypt grant, making option D the most likely cause.

Exam trap

DEA-C01 often tests the misconception that s3:GetObject alone is sufficient to read encrypted objects, ignoring the separate KMS permission requirement for SSE-KMS.

How to eliminate wrong answers

Option A is wrong because a data format mismatch would produce parsing or schema errors (e.g., Glue DataFormatException), not an S3ServiceException: Access Denied. Option B is wrong because s3:GetObjectVersion is only needed when accessing specific object versions; the Glue job reads the current version, so its absence would not cause Access Denied. Option C is wrong because the scenario explicitly states the Glue job uses the same IAM role that has s3:GetObject and s3:ListBucket on the bucket, so the role is not the issue.

416
MCQeasy

A data pipeline uses Amazon Kinesis Data Firehose to deliver data to Amazon S3. The delivery occasionally fails with 'Firehose is throttled'. What should be done to reduce throttling?

A.Enable compression on the Firehose delivery stream
B.Increase the buffer size and buffer interval
C.Decrease the buffer size to flush more frequently
D.Increase the number of shards in the Kinesis stream
AnswerB

Larger buffer reduces the number of write requests.

Why this answer

Increasing the buffer size and buffer interval gives Kinesis Data Firehose more time and data volume to accumulate before delivering to S3, reducing the frequency of PutRecord.Batch calls to the underlying Kinesis stream. This directly mitigates throttling by lowering the request rate, as Firehose throttling typically occurs when the per-shard write throughput limit (1,000 records/second or 1 MB/second) is exceeded.

Exam trap

The DEA-C01 exam often tests the misconception that Firehose throttling is resolved by scaling shards (like in Kinesis Data Streams), but Firehose manages its own internal shards and the correct fix is to adjust buffer settings to reduce API call frequency.

How to eliminate wrong answers

Option A is wrong because enabling compression reduces the data size sent to S3 but does not reduce the number of API calls or the request rate to the Kinesis stream, so it does not address throttling at the stream level. Option C is wrong because decreasing the buffer size causes more frequent flushes, which increases the request rate and exacerbates throttling rather than reducing it. Option D is wrong because Kinesis Data Firehose does not use a Kinesis data stream as its source by default; it uses its own internal stream with a fixed number of shards (default 1), and increasing shards is not a configurable option for Firehose—this option confuses Firehose with Kinesis Data Streams.

417
Multi-Selectmedium

A data engineer is troubleshooting an AWS Glue job that fails with 'java.lang.OutOfMemoryError: Java heap space'. The job processes a large dataset. Which TWO configuration changes should the engineer consider to resolve this issue? (Choose TWO.)

Select 2 answers
A.Change the output format from Parquet to CSV.
B.Increase the Spark shuffle partitions configuration (spark.sql.shuffle.partitions).
C.Reduce the number of partitions in the source data.
D.Increase the number of DPUs allocated to the Glue job.
E.Disable job bookmarks to avoid incremental processing.
AnswersB, D

Raising spark.sql.shuffle.partitions splits shuffle data into more, smaller partitions, so each task's working set fits within executor heap. This directly addresses the Java heap space exhaustion caused by oversized partitions during wide transformations on the large dataset.

Why this answer

Option B is correct because increasing spark.sql.shuffle.partitions creates more, smaller partitions during shuffles, which reduces the amount of data each executor task must hold in memory and helps prevent Java heap space exhaustion on large datasets. Option D is correct because allocating more DPUs adds more executors and memory to the Glue job, giving Spark more heap capacity to process the large dataset without running out of memory. Option A is incorrect because switching from Parquet to CSV increases data size and I/O, worsening memory pressure rather than relieving it.

Option C is incorrect because reducing source partitions concentrates more data per partition, increasing per-task memory usage and the risk of OOM. Option E is incorrect because disabling job bookmarks only affects incremental processing state and does not address heap memory consumption.

Exam trap

The trap is that candidates might think reducing partitions or changing output format would help, but these can worsen memory issues; the key is to increase resources and parallelism.

418
Multi-Selectmedium

A data engineer is designing a pipeline that ingests streaming data into Amazon S3 using Amazon Kinesis Data Firehose. The data must be delivered to S3 in Parquet format and partitioned by date. The engineer needs to configure the Firehose delivery stream. Which two actions are required to meet these requirements? (Choose two.)

Select 2 answers
A.Enable Amazon CloudWatch Logs for the Firehose delivery stream.
B.Set the Firehose buffer size to 128 MB and buffer interval to 900 seconds.
C.Enable record format conversion in the Firehose delivery stream and select Apache Parquet as the output format.
D.Configure dynamic partitioning with a JQ expression that extracts the date from each record.
E.Attach an IAM role to Firehose that allows s3:PutObject and glue:GetTable.
AnswersC, D

Firehose supports record format conversion from JSON to Parquet using an AWS Glue Data Catalog table for the schema. Enabling this feature and selecting Parquet as the output format is necessary to deliver data in columnar format, which is the requirement. Without it, Firehose delivers raw JSON.

Why this answer

To deliver Parquet, Firehose must use record format conversion, which requires selecting Parquet and referencing a Glue Data Catalog table for the schema. To partition by date, Firehose must use dynamic partitioning with a JQ expression that extracts the date from each record and includes it in the S3 prefix. Together, these two configurations satisfy the format and partitioning requirements.

Exam trap

The trap here is confusing supporting IAM or monitoring configurations with the specific features that enable Parquet conversion and dynamic partitioning.

419
MCQmedium

A data engineer is configuring an AWS Glue ETL job that reads from an Amazon S3 bucket containing sensitive customer records. The security team requires that the job's data be encrypted at rest using a customer managed AWS KMS key, and that the key policy restrict usage to the specific IAM role used by the Glue job. The engineer has already created the KMS key and attached the necessary IAM policy to the Glue job role. What additional step is required to ensure the Glue job can decrypt the S3 data using the customer managed key?

A.Add a KMS key policy statement granting kms:Decrypt and kms:GenerateDataKey permissions to the Glue job's IAM role.
B.Configure the S3 bucket to use SSE-S3 instead of SSE-KMS, because Glue cannot use customer managed keys.
C.Enable AWS CloudTrail logging for the KMS key to allow Glue to decrypt the data.
D.Attach an S3 bucket policy that grants the Glue job role s3:GetObject and s3:PutObject permissions.
AnswerA

The KMS key policy must explicitly allow the Glue job's IAM role to perform kms:Decrypt and kms:GenerateDataKey. Even if the IAM policy on the role grants these actions, the key policy is the primary access control for KMS keys and must also grant them. Without this, the Glue job will receive an AccessDenied error when attempting to read the encrypted S3 objects.

Why this answer

A customer managed KMS key requires that both the IAM policy of the calling principal and the KMS key policy grant the necessary permissions. The IAM policy alone is insufficient; the key policy must explicitly allow the Glue job's IAM role to use the key for decryption and data key generation. Without this key policy statement, the Glue job cannot decrypt the S3 objects.

Exam trap

The trap here is assuming that an IAM policy granting KMS actions is sufficient, when the KMS key policy must also grant those actions for a customer managed key.

420
MCQmedium

A data engineer is building an AWS Glue job that reads a table from the AWS Glue Data Catalog. The table contains columns with customer names, email addresses, and account numbers. The security team wants the job to mask the last four digits of account numbers in the output while leaving other columns unchanged. Which approach should the data engineer use?

A.Use the AWS Glue DataBrew masking recipe and apply it as a transformation step in the Glue job.
B.Apply a DynamicFrame map transformation and call the PII detection transform on the account number column.
C.Use a DynamicFrame map transformation with a custom function that replaces the account number value with a masked version.
D.Enable column-level encryption on the Data Catalog table definition for the account number column.
AnswerC

A DynamicFrame map transformation applies a user-defined function to each record, allowing the job to rewrite only the account number column while leaving other columns untouched. This gives precise control over the masking logic, such as keeping the last four digits. It runs natively inside the Glue job and produces the required masked output.

Why this answer

Masking specific column values during a Glue ETL run is best done with a map transformation that applies a custom function to each record. This keeps the transformation inside the job, gives full control over the masking logic, and leaves other columns unchanged. PII detection, DataBrew recipes, and column-level encryption do not perform value-level masking in this context.

Exam trap

The trap here is confusing PII detection with PII masking; detection identifies sensitive data but does not alter it.

421
Multi-Selecthard

A data engineer is designing an ingestion pipeline where AWS Lambda processes records from an Amazon Kinesis Data Stream. During peak traffic, records are being reprocessed and some are lost. The engineer needs to make the consumer resilient to failures and avoid duplicate processing. (Choose two.)

Select 2 answers
A.Enable a dead-letter queue or on-failure destination for the event source mapping.
B.Configure the event source mapping with a bisect-on-error function response and a maximum retry count.
C.Increase the Lambda function's reserved concurrency to match the number of shards.
D.Reduce the batch size to one record to eliminate any possibility of duplicates.
E.Set the starting position to LATEST so the consumer only reads new records.
AnswersA, B

Configuring an on-failure destination, such as an SQS queue or SNS topic, captures records that exceed the maximum retry count so they are not silently lost. This gives operators visibility and a path to reprocess failed records after fixing the root cause, which supports resilience.

Why this answer

Bisect-on-error with a bounded retry count isolates and stops poison records from blocking a shard, while an on-failure destination captures records that exhaust retries so they are not lost. Together they make the Kinesis consumer resilient and give operators a recovery path. Concurrency, starting position, and batch size changes do not address the underlying failure-handling gaps.

Exam trap

The trap here is assuming higher concurrency or smaller batches fix Kinesis consumer reliability, when the real issue is bounded retries and a failure destination for undeliverable records.

422
MCQhard

A data engineer is using AWS Lake Formation to manage permissions on a Data Catalog table backed by Amazon S3. Analysts query the table with Amazon Athena. The security team wants analysts to see only rows where the region column equals 'EU' and to prevent them from viewing the customer_id column entirely. Which combination of Lake Formation features should the engineer implement?

A.Use AWS Glue DataBrew to create a project that removes the customer_id column and filters to EU rows, then publish the result as a new table for analysts.
B.Define a Lake Formation data filter with a row filter expression region='EU' and exclude the customer_id column, then grant the analysts SELECT on the table with that filter.
C.Attach a tag-based access control policy to the table that grants access only when the analyst's IAM principal carries a tag matching the region value.
D.Create an IAM policy that allows Athena access only to the S3 prefix containing EU data and deny access to the customer_id column in the Data Catalog.
AnswerB

Lake Formation data filters support both row filter expressions and column inclusion or exclusion. A filter with region='EU' limits visible rows, excluding customer_id hides that column, and granting SELECT with the filter enforces both restrictions for Athena queries. This directly matches the stated row and column requirements.

Why this answer

Lake Formation data filters combine row filter expressions with column selection. A row filter of region='EU' restricts visible rows, and excluding customer_id from the filter removes that column from query results. Granting analysts SELECT on the table through the filter applies both restrictions automatically for Athena and other integrated engines without duplicating or transforming the underlying data.

Exam trap

The trap here is assuming that IAM policies or Lake Formation LF-Tags can filter individual data rows, when row-value filtering requires a data filter with a row expression.

423
MCQmedium

A company uses AWS Glue ETL jobs to transform data from Amazon RDS to Amazon S3 daily. The job recently started failing with memory errors. The data volume has grown 3x in the past month. Which change should the data engineer make to resolve the issue?

A.Increase the size of the Amazon RDS instance
B.Switch the Glue job type from Python Shell to Spark
C.Partition the output data in Amazon S3 by date
D.Increase the number of DPUs allocated to the Glue job
AnswerD

Glue allocates memory per DPU, so tripling data volume exhausts the current worker capacity. Increasing DPUs scales the compute and memory available to each task, resolving the out-of-memory failures. This directly addresses the resource constraint created by the threefold data growth.

Why this answer

The Glue job is failing with memory errors due to a 3x increase in data volume. Increasing the number of DPUs (Data Processing Units) allocated to the job provides more memory and compute resources, directly addressing the out-of-memory condition without changing the job logic or architecture.

Exam trap

The trap here is that candidates may confuse scaling the source database (RDS) with scaling the ETL compute (Glue), or assume that output partitioning (S3) will fix an in-memory processing error, when the actual solution is to increase the compute resources allocated to the Glue job.

How to eliminate wrong answers

Option A is wrong because increasing the RDS instance size does not affect the memory available to the Glue ETL job; the bottleneck is in the Glue execution environment, not the source database. Option B is wrong because switching from Python Shell to Spark would change the execution model but does not inherently resolve memory errors; Python Shell jobs are limited to a single executor with fixed memory, while Spark jobs distribute work but still require sufficient DPUs to handle the data volume. Option C is wrong because partitioning output data in S3 by date improves query performance and cost but does not reduce the memory footprint of the Glue job during the transformation phase; the memory error occurs during processing, not during writing.

424
MCQhard

A data engineer is using AWS Lake Formation to manage permissions on a data lake in Amazon S3. The engineer grants SELECT permission on a table to an IAM role used by an Amazon Athena user. However, the user reports that queries against the table return an 'Access Denied' error. The engineer verifies that the IAM role has the necessary S3 permissions and that Lake Formation permissions are correctly set. What is the most likely cause of the error?

A.The Athena workgroup has a query result location that the IAM role cannot access.
B.The AWS Glue Data Catalog is not integrated with Lake Formation.
C.The S3 bucket policy does not allow access from the Athena service.
D.The IAM role lacks the lakeformation:GetDataAccess permission.
AnswerD

AWS Lake Formation requires the IAM principal to have the lakeformation:GetDataAccess permission to obtain temporary credentials for accessing data. Even with SELECT permission on the table, without this permission, Athena cannot retrieve data from S3. This is a common oversight when configuring Lake Formation permissions.

Why this answer

When Lake Formation manages access, IAM principals need the lakeformation:GetDataAccess permission to obtain temporary credentials for data access. Without it, Athena cannot read data even if table permissions are granted. This permission is often overlooked because it is not part of standard S3 or Glue permissions.

Exam trap

The trap here is focusing on S3 bucket policies or Data Catalog integration, while the missing lakeformation:GetDataAccess permission is the subtle requirement for data access with Lake Formation.

425
Multi-Selecthard

A data engineer needs to transform data in Amazon S3 using AWS Glue. The job must handle schema evolution and partition pruning. Which THREE features should be used?

Select 3 answers
A.AWS Glue Data Catalog
B.AWS Glue job bookmarks
C.AWS Glue FindMatches transform
D.AWS Glue crawlers
E.Partition indexes
AnswersA, D, E

The AWS Glue Data Catalog stores table definitions, schemas and partition metadata centrally, so Glue jobs read current schema versions and prune partitions using that metadata. It is the component that lets the job handle evolving schemas and avoid scanning irrelevant partitions.

Why this answer

AWS Glue Data Catalog (A) is correct because it stores the table definitions, schemas, and partition metadata that Glue ETL jobs reference, and it is the component that must be updated when schema evolution occurs so the job can read new columns. AWS Glue crawlers (D) are correct because they automatically scan S3 data, detect new or changed columns, and update the Data Catalog with revised schemas and newly discovered partitions, which is exactly how schema evolution is handled. Partition indexes (E) are correct because they accelerate partition pruning by allowing Glue and Athena to look up partitions without listing the entire catalog, dramatically reducing planning time for highly partitioned tables.

AWS Glue job bookmarks (B) only track previously processed data to support incremental loads and do not address schema evolution or partition pruning. AWS Glue FindMatches transform (C) is a machine-learning transform for deduplicating and matching records, which is unrelated to schema evolution or partition pruning.

Exam trap

The DEA-C01 exam often tests the distinction between 'incremental processing' (job bookmarks) and 'schema evolution' (Data Catalog + crawlers), leading candidates to incorrectly select job bookmarks for schema changes.

426
MCQhard

A data engineer is designing a data ingestion pipeline for a social media analytics platform. The pipeline must ingest tweets in real-time, perform sentiment analysis, and store results in Amazon S3. The sentiment analysis is compute-intensive and must be done as the data arrives. The estimated throughput is 10,000 tweets per second. Which architecture is most suitable?

A.Amazon SQS with AWS Lambda pollers to process tweets and store in S3.
B.Amazon EMR with Spark Streaming to process tweets and write to S3.
C.Amazon Kinesis Data Streams with Amazon Kinesis Data Analytics for sentiment analysis, then Kinesis Data Firehose to S3.
D.Amazon API Gateway with AWS Lambda to process each tweet and store in S3.
AnswerC

Scalable real-time stream processing.

Why this answer

The most suitable because Amazon Kinesis Data Streams can ingest up to 10,000 records per second per shard (with shard-level scaling), and Kinesis Data Analytics provides built-in, low-latency stream processing for compute-intensive sentiment analysis using SQL or Apache Flink. Kinesis Data Firehose then reliably buffers and writes the processed results to Amazon S3 without custom code, ensuring near-real-time delivery.

Exam trap

The trap here is that candidates often choose SQS+Lambda (Option A) for simplicity, underestimating the throughput ceiling and polling overhead, while overlooking Kinesis Data Analytics as the only AWS-managed service that natively supports real-time, compute-intensive stream processing without custom infrastructure.

How to eliminate wrong answers

Option A is wrong because Amazon SQS with Lambda pollers introduces polling latency and cannot efficiently handle 10,000 tweets per second; Lambda has a maximum concurrency limit and SQS batch sizes are capped at 10 messages, leading to throttling and backpressure. Option B is wrong because Amazon EMR with Spark Streaming is designed for large-scale batch and micro-batch processing, not for true real-time, per-record sentiment analysis at 10,000 TPS; it incurs startup overhead and is better suited for historical analysis. Option D is wrong because Amazon API Gateway with Lambda processes each tweet synchronously, which cannot sustain 10,000 requests per second without aggressive throttling and cold starts; it also lacks built-in stream buffering and ordering for real-time ingestion.

427
MCQeasy

A data engineer receives an alert that an AWS KMS key has been scheduled for deletion by mistake. What is the immediate action to prevent the key from being deleted?

A.Cancel the key deletion from the KMS console or API.
B.Create a new KMS key and re-encrypt the data.
C.Wait for the key to be deleted and restore it from backup.
D.Disable the key immediately to stop usage.
AnswerA

Cancelling the scheduled deletion via the KMS console or API immediately halts the pending deletion window, satisfying the stem's requirement to prevent the key from being destroyed. AWS KMS permits cancellation at any point during the mandatory waiting period, restoring the key to a usable state before permanent deletion occurs.

Why this answer

When a KMS key is scheduled for deletion, the deletion can be canceled from the AWS KMS console or via the CancelKeyDeletion API during the pending deletion period. This immediate action restores the key to its previous state and prevents it from being deleted. Option B is incorrect because creating a new key does not cancel the deletion of the existing key.

Option C is incorrect because deleted KMS keys cannot be restored; the default waiting period of 7–30 days exists specifically to allow cancellation. Option D is incorrect because disabling the key does not affect the deletion schedule.

428
MCQhard

A data engineer is managing an Amazon DynamoDB table that stores user session data. The table has a partition key of user_id and a sort key of session_start_time. The workload includes frequent queries that retrieve all sessions for a user within a specific time range. The engineer notices that some queries are slow and wants to optimize the table design. Which action should the engineer take to improve query performance?

A.Ensure that queries use the partition key and sort key condition to retrieve data efficiently, and consider enabling adaptive capacity.
B.Increase the provisioned read capacity units (RCUs) for the table to handle the query load.
C.Create a global secondary index (GSI) with session_start_time as the partition key and user_id as the sort key.
D.Enable DynamoDB Streams to capture changes and use AWS Lambda to maintain a separate queryable table.
AnswerA

The base table's key schema (user_id as partition key, session_start_time as sort key) is ideal for the query pattern. Using both keys in queries allows DynamoDB to efficiently locate the relevant items. Enabling adaptive capacity helps handle uneven access patterns by automatically adjusting throughput. This approach directly addresses performance by leveraging the existing design and DynamoDB's built-in optimizations.

Why this answer

The correct action is to ensure queries use the partition key and sort key condition to retrieve data efficiently, and to consider enabling adaptive capacity. The table's key schema is already optimized for the query pattern, so leveraging it properly is key. Other options introduce unnecessary complexity, do not address the query pattern, or simply scale capacity without fixing efficiency.

Exam trap

The trap here is assuming that adding a GSI or increasing capacity will solve slow queries, when the base table design is already correct and the issue may be query patterns.

429
MCQhard

A data engineer is using AWS Step Functions to orchestrate a series of AWS Glue jobs and Amazon EMR steps as part of a nightly ETL pipeline. The pipeline occasionally fails due to transient issues such as Amazon S3 throttling or temporary network errors. The engineer wants to make the workflow more resilient without duplicating the entire state machine. Which Step Functions feature should be used to automatically retry failed states?

A.Configure Retry and Catch fields on individual states.
B.Enable X-Ray tracing on the state machine.
C.Set the state machine's TimeoutSeconds to a large value.
D.Use a Map state to iterate over each Glue job.
AnswerA

The Retry field on a state allows you to define automatic retries with backoff and max attempts for specific error types, while Catch can route to a fallback state after retries are exhausted. This directly addresses transient failures without modifying the state machine structure. It is the intended mechanism for handling intermittent errors such as S3 throttling in Step Functions workflows.

Why this answer

Step Functions provides Retry and Catch fields at the state level to handle transient errors. By configuring Retry with a backoff strategy and maximum attempts, the workflow can automatically re-execute a failed state, such as a Glue job start or EMR step, without manual intervention. Catch can then define a fallback if retries are exhausted.

This is the standard approach for making orchestration resilient to intermittent service issues.

Exam trap

The trap here is confusing observability features like X-Ray tracing or execution timeouts with actual error-handling and retry capabilities.

430
MCQeasy

Refer to the exhibit. An IAM policy is attached to a user. What is the security implication of this policy?

A.The policy only allows read access.
B.The policy is invalid because it uses asterisks.
C.The policy is too restrictive.
D.The policy grants excessive permissions, violating least privilege.
AnswerD

The policy's Action element includes wildcard permissions such as "s3:*" across all resources, so the attached user can perform any S3 operation on every bucket in the account. This exceeds the task's requirement, directly breaching the least-privilege constraint by granting far broader access than necessary.

Why this answer

The policy, which grants full S3 access to all resources, violates the principle of least privilege by providing excessive permissions. Option A is incorrect because the policy does not only allow read access; it allows all actions. Option B is incorrect because the use of asterisks is valid syntax in IAM policies.

Option C is incorrect because the policy is overly permissive, not restrictive.

431
MCQmedium

A company uses AWS Glue to process streaming data from Amazon Kinesis Data Streams. The data is JSON formatted and includes a timestamp field. The company wants to partition the output in Amazon S3 by date and hour, and ensure exactly-once processing semantics. Which combination of configurations should be used?

A.Disable checkpointing and use the 'exactly_once' delivery option in Kinesis Data Streams.
B.Enable checkpointing in the AWS Glue streaming job and specify an S3 location for checkpoint data.
C.Use Amazon DynamoDB as a checkpoint store by configuring the Glue job with a DynamoDB connection.
D.Use Kinesis Client Library (KCL) checkpointing with a DynamoDB table.
AnswerB

Glue streaming jobs support checkpointing to S3 for exactly-once processing.

Why this answer

AWS Glue streaming jobs require checkpointing to track the progress of data consumption from Kinesis Data Streams and to ensure exactly-once processing semantics. By enabling checkpointing and specifying an S3 location, Glue periodically saves the state of processed records, allowing it to resume from the last committed offset in case of failures, thus preventing duplicates or data loss.

Exam trap

The trap here is that candidates confuse the checkpointing mechanism of AWS Glue (which uses S3) with the Kinesis Client Library (KCL) pattern (which uses DynamoDB), leading them to select option D or C, even though Glue streaming jobs do not support DynamoDB for checkpointing.

How to eliminate wrong answers

Option A is wrong because disabling checkpointing removes the mechanism for tracking processed records, making exactly-once semantics impossible; the 'exactly_once' delivery option in Kinesis Data Streams refers to producer-side delivery guarantees, not consumer-side processing semantics. Option C is wrong because AWS Glue streaming jobs do not support DynamoDB as a checkpoint store; they only support S3 for checkpoint data. Option D is wrong because Kinesis Client Library (KCL) checkpointing with DynamoDB is a pattern for custom applications, not for AWS Glue streaming jobs, which manage checkpointing internally via S3.

432
MCQhard

A company uses Amazon Redshift for a data warehouse. They notice that queries are slow due to heavy data skew. Which optimization technique should be applied first?

A.Configure workload management (WLM) queues
B.Define sort keys on frequently filtered columns
C.Set an appropriate distribution style
D.Apply compression encodings to columns
AnswerC

Data skew concentrates rows on some slices, so those slices do redundant work during joins and aggregations. Choosing a distribution style that spreads rows evenly, such as KEY on a high-cardinality column or ALL for small tables, addresses the root cause first.

Why this answer

Data skew occurs when rows are distributed unevenly across Redshift slices, causing some nodes to process far more data than others. Setting an appropriate distribution style (e.g., KEY, EVEN, or ALL) redistributes the data to balance the workload, directly addressing the root cause of the slowness. This is the first optimization to apply because skew is a fundamental distribution issue that other tuning steps cannot fix.

Exam trap

The trap here is that candidates often confuse distribution skew with sort key optimization or compression, mistakenly believing that improving data organization on disk (sort keys) or reducing I/O (compression) will fix uneven data distribution across nodes.

How to eliminate wrong answers

Option A is wrong because WLM queues manage concurrency and memory allocation for query slots, not the physical distribution of data across nodes; they cannot fix performance degradation caused by data skew. Option B is wrong because sort keys optimize the order of data on disk to improve range-restricted scans and merge joins, but they do not redistribute data or alleviate skew across slices. Option D is wrong because compression encodings reduce storage footprint and I/O by compressing column data, but they have no effect on how rows are distributed across nodes or on query parallelism.

433
MCQmedium

A data engineer is designing a data lake on Amazon S3 to store JSON logs from an application. The logs are written once and never modified. The engineer needs to query the data using Amazon Athena with the best performance and lowest cost. The engineer wants to partition the data by year, month, and day based on the log timestamp. Which approach should the engineer use to organize the S3 objects?

A.Enable S3 Inventory to generate a manifest of objects and use it to create a partitioned table in Athena.
B.Use a Hive-style prefix structure such as s3://bucket/year=YYYY/month=MM/day=DD/ and define partitions in the AWS Glue Data Catalog.
C.Store objects in a flat structure and create an AWS Glue crawler to automatically add partitions based on file metadata.
D.Store all objects in a single prefix and create a table with a partition projection based on the timestamp column.
AnswerB

This approach creates a hierarchical prefix that Athena and AWS Glue recognize as partitions. When queries filter on year, month, or day, Athena prunes irrelevant partitions, scanning less data and lowering cost. It is the standard, efficient way to organize time-series data in S3 for query engines like Athena and Redshift Spectrum.

Why this answer

Organizing data in a Hive-style prefix hierarchy allows Athena and AWS Glue to recognize partitions. When queries filter on partition columns, only relevant prefixes are scanned, reducing data scanned and cost. This is the recommended practice for time-series data in S3.

Other options either do not create real partitions or fail to provide efficient pruning.

Exam trap

The trap here is assuming that any S3 organization can be partitioned by simply defining a table schema, without the physical prefix structure that enables partition pruning.

434
MCQmedium

A company uses AWS Glue to process sensitive data stored in S3. The security team requires that all data be encrypted at rest using customer-managed KMS keys. The data engineers are encountering 'Access Denied' errors when running Glue ETL jobs. What is the most likely cause?

A.The Glue service role does not have kms:Decrypt and kms:Encrypt permissions for the KMS key.
B.The KMS key policy does not allow the AWS Glue service to use the key.
C.The Glue Data Catalog is encrypted with a different KMS key.
D.The S3 bucket policy denies access to the Glue service role.
AnswerA

Correct. The Glue service role must have kms:Decrypt and kms:Encrypt permissions for the KMS key to read/write encrypted data from S3.

Why this answer

To decrypt S3 objects encrypted with a customer-managed KMS key, the AWS Glue job's service role must have kms:Decrypt and kms:Encrypt permissions in its IAM policy, and the KMS key policy must grant that role access. The most likely cause of the Access Denied error is missing KMS permissions on the Glue service role (A). Option B is a possible KMS key policy denial, but it is less likely than missing IAM KMS permissions in this scenario and is not as direct.

Option C is irrelevant because Data Catalog encryption uses a different key and does not prevent S3 object access. Option D would be an S3 bucket policy denial, not a KMS 'Access Denied'; the issue here is KMS permissions.

435
Multi-Selecthard

A company is building a data lake on S3. They have a large volume of CSV files (hundreds of GB) in a source bucket. They need to convert them to Parquet, partition by date, and ensure the data is encrypted at rest with SSE-KMS. The pipeline must be triggered automatically when new files arrive. Which THREE steps should be part of the solution? (Choose THREE.)

Select 3 answers
A.Configure S3 Event Notification to send events to an SQS queue
B.Use Amazon Kinesis Data Firehose to ingest new files
C.Create an AWS Glue ETL job that converts to Parquet and partitions by date
D.Use Amazon Athena CTAS query to convert files in batch
E.Configure the Glue job to use a KMS key for server-side encryption in S3
AnswersA, C, E

SQS can buffer events and trigger a Lambda or Step Functions workflow.

Why this answer

S3 Event Notifications can be configured to send events to an SQS queue when new CSV files arrive. This decouples the ingestion pipeline, allowing the Glue job to poll the queue for new file notifications and trigger processing without tight coupling or polling the S3 bucket directly. SQS provides reliable, scalable message delivery that can trigger downstream ETL workflows.

Exam trap

The trap here is that candidates often confuse batch conversion tools like Athena CTAS with event-driven pipelines, or assume Kinesis Firehose can process existing S3 files, when in fact Firehose only ingests streaming data and cannot read from S3 buckets.

436
MCQhard

A data engineer is monitoring an Amazon Redshift cluster and notices that the 'WLM query wait time' metric is consistently high during peak hours. The cluster uses automatic WLM. The engineer wants to reduce query wait times without changing the cluster size. Which action is MOST effective?

A.Enable concurrency scaling.
B.Change WLM to manual mode and increase the number of queues.
C.Increase the maximum number of queries per queue.
D.Enable short query acceleration (SQA).
AnswerA

Concurrency scaling adds transient Redshift compute capacity automatically when queries queue, directly cutting WLM query wait time during peak hours. It absorbs burst concurrency without resizing the cluster, satisfying the constraint of reducing wait times at unchanged cluster size.

Why this answer

Enabling concurrency scaling (Option A) is the most effective action because it automatically adds transient cluster capacity during peak loads, allowing more queries to run concurrently without increasing wait times. This is specifically designed to reduce WLM query wait time. Option B (manual WLM) requires tuning and does not add capacity.

Option C (increasing max queries per queue) could increase concurrency but may lead to resource contention and longer wait times if the cluster is already saturated. Option D (short query acceleration) prioritizes short queries, which does not address overall wait times for all queries. Therefore, A is correct.

437
MCQhard

A data engineer is troubleshooting a DMS task that is replicating data from an on-premises Oracle database to an RDS for MySQL instance. The task is failing with 'ORA-1555: snapshot too old' error. What is the best course of action?

A.Disable full supplemental logging on the source tables.
B.Increase the size of the redo logs on the source database.
C.Enable batch optimized apply on the DMS task.
D.Increase the UNDO tablespace size and set UNDO_RETENTION to a higher value.
AnswerD

ORA-1555 occurs when Oracle cannot reconstruct a consistent read image because undo data was overwritten before the DMS read completed. Enlarging the UNDO tablespace and raising UNDO_RETENTION preserves that undo long enough for the long-running extraction.

Why this answer

ORA-1555 'snapshot too old' occurs when Oracle cannot reconstruct a consistent read image because the required undo data has been overwritten. Increasing the UNDO tablespace size and raising UNDO_RETENTION gives Oracle more undo history to retain, allowing the long-running DMS full-load read to complete without losing its snapshot. This directly addresses the root cause of insufficient undo retention on the source.

Exam trap

The trap is confusing redo logs (crash recovery) with undo tablespace (read consistency), leading candidates to increase redo log size when the actual fix is undo retention.

How to eliminate wrong answers

Option A is wrong because disabling supplemental logging would break change data capture for ongoing replication and does nothing to fix undo retention; supplemental logging is required for DMS CDC. Option B is wrong because redo logs protect against instance failure and support recovery, not consistent-read reconstruction — the snapshot-too-old error is an undo problem, not a redo problem. Option C is wrong because batch optimized apply is a target-side apply optimization for MySQL and does not affect the source Oracle undo retention that causes ORA-1555.

438
MCQeasy

A data engineer is using AWS Step Functions to orchestrate a series of AWS Glue jobs that process data in Amazon S3. The engineer needs to ensure that if a Glue job fails, the entire workflow stops and an Amazon SNS notification is sent. Which Step Functions state type should be used to handle the error and send the notification?

A.A Choice state that evaluates the output of the Glue job and branches based on success or failure.
B.A Task state with a Catch field that transitions to a state that publishes to an Amazon SNS topic.
C.A Parallel state that runs the Glue job and an SNS notification task simultaneously.
D.A Task state with a Retry field that retries the Glue job up to three times.
AnswerB

The Catch field in a Task state is used to catch errors and transition to a fallback state. By configuring a Catch that leads to a state which publishes to an SNS topic, the engineer can stop the workflow and send a notification. This is the standard way to handle errors in Step Functions and meet the requirement.

Why this answer

In AWS Step Functions, the Catch field on a Task state allows you to define a fallback state when an error occurs. By transitioning to a state that publishes to an Amazon SNS topic, the workflow can stop and send a notification. This is the correct pattern for error handling and alerting in Step Functions.

Exam trap

The trap here is confusing error handling with retry logic or conditional branching; only the Catch field provides the mechanism to transition to a notification state on failure.

439
MCQmedium

A data engineer needs to transfer 10 TB of data from an on-premises Hadoop cluster to Amazon S3. The network bandwidth is limited to 100 Mbps, and the transfer must be completed within 48 hours. Which solution meets the requirements?

A.Use AWS DataSync to transfer data online
B.Use AWS Snowball Edge device to transfer data offline
C.Use S3 Transfer Acceleration over the internet
D.Set up AWS Direct Connect to increase bandwidth
AnswerB

At 100 Mbps, transferring 10 TB over the network would take roughly ten days, far exceeding the 48-hour deadline. AWS Snowball Edge ships the data physically, sidestepping the bandwidth constraint entirely and meeting the required completion window.

Why this answer

The on-premises Hadoop cluster has 10 TB of data to transfer, but the network bandwidth is only 100 Mbps. At 100 Mbps, the theoretical maximum transfer rate is about 12.5 MB/s, which would take approximately 10 TB / 12.5 MB/s ≈ 800,000 seconds ≈ 222 hours — far exceeding the 48-hour window. AWS Snowball Edge is an offline, physical device that bypasses network constraints entirely, allowing you to transfer the data by shipping the device, which completes within days regardless of bandwidth.

Exam trap

The trap here is that candidates may assume S3 Transfer Acceleration or Direct Connect can magically overcome a hard bandwidth cap, but neither increases the last-mile bandwidth; the only way to transfer 10 TB in under 48 hours with a 100 Mbps link is to use an offline physical device like Snowball Edge.

How to eliminate wrong answers

Option A is wrong because AWS DataSync is an online data transfer service that still relies on network bandwidth; at 100 Mbps, it cannot transfer 10 TB within 48 hours due to the same bandwidth limitation. Option C is wrong because S3 Transfer Acceleration only optimizes routing over the internet using AWS edge locations, but it does not increase the underlying 100 Mbps bandwidth; the transfer would still take far longer than 48 hours. Option D is wrong because AWS Direct Connect provides a dedicated network connection, but it does not inherently increase bandwidth beyond the 100 Mbps limit unless you provision a higher-capacity circuit, which is not specified and would still require time to set up; the question assumes the bandwidth is fixed at 100 Mbps.

440
MCQmedium

A data engineer manages an AWS Glue ETL job that reads CSV files from Amazon S3 and writes Parquet to another S3 prefix. The job recently started failing with the error 'AnalysisException: Unable to infer schema for CSV.' The engineer confirms the S3 path contains files and the IAM role has s3:GetObject permissions. The job's script uses glueContext.create_dynamic_frame.from_catalog with a database and table name. What is the MOST likely cause?

A.The AWS Glue job's bookmarks are enabled, causing it to skip files that were already processed and leaving no data to infer schema from.
B.The IAM role lacks s3:ListBucket permission on the source bucket, so the job cannot list the CSV files.
C.The CSV files are compressed with gzip, and AWS Glue cannot infer schema from compressed files.
D.The AWS Glue Data Catalog table's SerDe parameters are misconfigured, so the crawler did not correctly identify the CSV delimiter or header.
AnswerD

The error 'Unable to infer schema for CSV' occurs when the Glue job reads from the Data Catalog and the table metadata lacks a valid SerDe or has incorrect parameters such as separatorType or header. The job relies on the catalog table definition, not direct file inspection, so misconfigured metadata prevents schema inference.

Why this answer

When an AWS Glue job uses create_dynamic_frame.from_catalog, it depends entirely on the AWS Glue Data Catalog table definition to determine the schema. If the table's SerDe parameters are incorrect—such as a wrong delimiter, missing header, or incorrect classification—the job cannot infer the schema and throws the AnalysisException. Verifying and repairing the catalog table resolves the failure.

Exam trap

The trap here is assuming that direct S3 file access or compression causes the schema inference error, when the job is actually reading metadata from the Data Catalog.

441
Multi-Selecthard

Which THREE of the following are benefits of using Amazon DynamoDB Accelerator (DAX)? (Choose three.)

Select 3 answers
A.Offloads read traffic from the DynamoDB table.
B.Improves write throughput by batching writes.
C.Reduces read latency from single-digit milliseconds to microseconds.
D.Supports write-through caching to improve write performance.
E.Provides in-memory caching for DynamoDB tables.
AnswersA, C, E

DAX handles read requests, reducing load on the table.

Why this answer

DAX acts as a read-through cache that offloads read traffic from the DynamoDB table, reducing the number of read requests that hit the underlying table and thus lowering the consumed read capacity units (RCUs). This allows the table to handle more concurrent reads without scaling up provisioned capacity.

Exam trap

The trap here is that candidates often assume DAX improves write performance or supports write-through caching, but DAX is strictly a read cache and does not accelerate or batch writes.

442
MCQeasy

A data engineer needs to transform a large dataset stored in Amazon S3 using Apache Spark. The engineer wants to minimize startup time and use a serverless approach. Which AWS service should the engineer use?

A.Amazon Redshift
B.Amazon EMR
C.AWS Glue
D.Amazon Athena
AnswerC

AWS Glue provides a serverless Apache Spark environment, so the engineer submits Spark jobs without provisioning or waiting for cluster startup. This directly satisfies the requirement to minimise startup time while avoiding infrastructure management, unlike Amazon EMR, which requires cluster provisioning.

Why this answer

AWS Glue provides a serverless Spark environment with fast startup. Option A is wrong because Amazon Redshift is a data warehouse, not a Spark environment. Option B is wrong because Amazon EMR requires cluster provisioning, which increases startup time.

Option D is wrong because Amazon Athena is for querying data, not for transforming with Spark.

443
MCQeasy

A data engineering team needs to ingest streaming data from an application into Amazon S3 for analytics. The data volume is moderate and the team wants the lowest operational overhead. Which AWS service should they use?

A.Amazon SQS
B.AWS Glue
C.Amazon Kinesis Data Streams
D.Amazon Kinesis Data Firehose
AnswerD

Kinesis Data Firehose is fully managed, automatically scaling and delivering streaming data to Amazon S3 without managing clusters or consumers, meeting the lowest operational overhead constraint. Kinesis Data Streams would require provisioning shards and custom consumers, adding operational burden for moderate volume.

Why this answer

Amazon Kinesis Data Firehose is a fully managed service for loading streaming data into S3 with no code required and minimal operational overhead. Option A is incorrect because Amazon SQS is a message queue service, not designed for streaming data ingestion into S3. Option B is incorrect because AWS Glue is primarily a batch ETL service, not suitable for real-time streaming.

Option C is incorrect because Amazon Kinesis Data Streams requires custom consumers and more management, increasing operational overhead.

444
MCQeasy

A company wants to migrate its on-premises MySQL database to Amazon RDS for MySQL with minimal downtime. Which AWS service should be used for the migration?

A.AWS Database Migration Service (DMS)
B.AWS Schema Conversion Tool (SCT)
C.AWS DataSync
D.AWS Direct Connect
AnswerA

AWS Database Migration Service performs ongoing replication from the on-premises MySQL source to Amazon RDS for MySQL, keeping both in sync until cutover. This continuous change data capture satisfies the minimal-downtime constraint, unlike a one-off dump and restore, which would require halting writes for the full load duration.

Why this answer

AWS Database Migration Service (DMS) is purpose-built for migrating databases to AWS with minimal downtime by using ongoing replication (change data capture, CDC) from the source MySQL database to the target Amazon RDS for MySQL instance. This allows the source to remain fully operational during the migration, meeting the minimal-downtime requirement.

Exam trap

The trap here is that candidates confuse AWS DMS with AWS DataSync or SCT, assuming any data transfer tool works for database migration, but DMS is the only service that supports ongoing replication for minimal downtime database migrations.

How to eliminate wrong answers

Option B (AWS Schema Conversion Tool) is wrong because SCT is used for converting database schemas from one engine to another (e.g., Oracle to Aurora), not for migrating data with minimal downtime; it does not handle ongoing replication. Option C (AWS DataSync) is wrong because DataSync is designed for moving large volumes of file data (e.g., NFS, SMB) to Amazon S3 or EFS, not for database migrations or CDC replication. Option D (AWS Direct Connect) is wrong because Direct Connect establishes a dedicated network connection between on-premises and AWS, but it is a connectivity service, not a migration tool; it does not perform data migration or replication.

445
MCQhard

A data engineer is migrating an on-premises Apache HBase workload to Amazon DynamoDB. The HBase table has a row key with composite structure: customer_id (10 chars) + timestamp (10 digits). The access pattern is to query by customer_id and retrieve the latest entries. How should the DynamoDB table be designed to optimize performance?

A.Create a table with partition key = customer_id and sort key = timestamp.
B.Use Amazon S3 with customer_id as prefix and timestamp as object name.
C.Create a table with partition key = concatenated customer_id and timestamp.
D.Create a table with partition key = timestamp and sort key = customer_id.
AnswerA

Splitting the composite row key into partition key customer_id and sort key timestamp lets DynamoDB distribute items across partitions by customer, while the sort key orders entries chronologically. Querying a single customer_id with ScanIndexForward=false returns the latest entries efficiently, satisfying the query-by-customer, retrieve-latest access pattern without full table scans.

Why this answer

DynamoDB's partition key (customer_id) evenly distributes data across partitions, while the sort key (timestamp) enables efficient range queries using Query with ScanIndexForward=false to retrieve the latest entries. This design directly maps the HBase composite row key pattern to DynamoDB's primary key structure, optimizing for the described access pattern.

Exam trap

The trap here is that candidates may think concatenating the row key into a single partition key (Option C) preserves the query pattern, but DynamoDB requires the partition key to be known exactly for queries, making it impossible to query by customer_id alone without a full scan.

How to eliminate wrong answers

Option B is wrong because Amazon S3 is an object store, not a low-latency NoSQL database; it lacks native support for range queries and cannot efficiently retrieve the latest entries by timestamp without scanning all objects. Option C is wrong because using a concatenated partition key (customer_id + timestamp) prevents querying by customer_id alone, as DynamoDB requires the exact partition key value for queries, forcing a full scan. Option D is wrong because using timestamp as the partition key leads to hot partitions (e.g., all writes for the same second hit one partition) and does not allow efficient retrieval by customer_id without a scan.

446
Multi-Selecteasy

Which TWO actions are effective ways to monitor the health of an Amazon DynamoDB table? (Choose two.)

Select 2 answers
A.Use AWS S3 inventory to track table size.
B.Use EC2 instance status checks.
C.Enable DynamoDB Streams and process with Lambda to detect failures.
D.Set up Amazon CloudWatch alarms on ConsumedReadCapacityUnits.
E.Monitor the 'TableHealth' metric in CloudWatch.
AnswersC, D

DynamoDB Streams capture item-level changes in near real time; processing them with Lambda lets you detect write failures, anomalies or unexpected patterns. This satisfies the monitoring requirement by providing change-level visibility that capacity metrics alone cannot reveal.

Why this answer

Option C is correct because DynamoDB Streams captures item-level changes in near real time, and a Lambda consumer can inspect those records to detect anomalies or failed writes, providing an event-driven health signal for the table. Option D is correct because CloudWatch publishes DynamoDB metrics such as ConsumedReadCapacityUnits (and ConsumedWriteCapacityUnits), and alarms on these metrics reveal throttling or unexpected capacity consumption that indicates table health issues. Option A is wrong because S3 Inventory reports on S3 bucket objects, not DynamoDB table size or health.

Option B is wrong because EC2 instance status checks only reflect the health of EC2 instances, not a managed DynamoDB table. Option E is wrong because DynamoDB does not expose a 'TableHealth' metric in CloudWatch; health must be inferred from real metrics like throttled requests, consumed capacity, and latency.

Exam trap

DEA-C01 often tests the difference between monitoring tools for DynamoDB; candidates might incorrectly assume that EC2 status checks or S3 inventory apply, or invent non-existent metrics like 'TableHealth'.

447
MCQmedium

A data engineer manages an AWS Glue Data Catalog used by Amazon Athena analysts. The security team wants column-level restrictions so that analysts querying a specific table cannot view the values in a cardholder_name column, while still being able to query all other columns. The analysts connect through Athena using an IAM role. Which approach meets this requirement with the LEAST operational overhead?

A.Enable AWS CloudTrail data events on the table and alert when cardholder_name is queried.
B.Create a separate Athena view that omits cardholder_name and grant the analysts access only to that view.
C.Define an IAM policy that denies the glue:GetTable action for the cardholder_name column and attach it to the analysts' role.
D.Create an AWS Lake Formation data filter that excludes the cardholder_name column, then grant the analysts' IAM role SELECT on the table with that data filter applied.
AnswerD

Lake Formation data filters restrict column visibility at the catalog level, so Athena queries referencing the excluded column fail while all other columns remain readable. Granting SELECT with the filter attached enforces the restriction centrally without duplicating tables or rewriting queries, which is the least-overhead method for column-level governance.

Why this answer

Column-level access control in the Glue Data Catalog is provided by Lake Formation data filters, which can include or exclude specific columns and are attached to grants. Granting the analysts' role SELECT with a filter that excludes cardholder_name lets all other columns remain queryable through Athena while the restricted column is inaccessible. IAM policies, views, and CloudTrail do not enforce column-level prevention centrally.

Exam trap

The trap here is assuming IAM policies can restrict access at the column level, when Glue Catalog column restrictions are enforced by Lake Formation data filters attached to grants.

448
MCQmedium

A data engineer manages an AWS Glue ETL job that processes sensitive customer records stored in Amazon S3. The security team mandates that all data at rest in the S3 bucket be encrypted with AWS KMS keys, and that the Glue job have the minimum permissions necessary to read and write data. The engineer creates a KMS key and configures the S3 bucket to use SSE-KMS with that key. Which additional step is required to allow the Glue job to access the encrypted data?

A.Attach an IAM policy to the Glue job's execution role that allows kms:Decrypt and kms:GenerateDataKey on the KMS key.
B.Modify the S3 bucket policy to grant the Glue service principal kms:Decrypt and kms:GenerateDataKey permissions.
C.Enable default encryption on the S3 bucket using SSE-S3 instead of SSE-KMS to avoid KMS permission complexity.
D.Create a VPC endpoint for KMS and attach it to the Glue job's subnet to allow KMS traffic.
AnswerA

The Glue job's execution role must have explicit permissions to use the KMS key for both reading (kms:Decrypt) and writing (kms:GenerateDataKey) encrypted objects. Without these permissions, the job cannot decrypt source data or encrypt output. This is the minimum required step to enable access while adhering to least privilege.

Why this answer

For an AWS Glue job to read and write SSE-KMS encrypted S3 objects, its execution role must have IAM permissions for kms:Decrypt and kms:GenerateDataKey on the specific KMS key. S3 bucket policies do not control KMS key access, and switching encryption types would violate requirements. VPC endpoints address network connectivity, not authorization.

Exam trap

The trap here is assuming that S3 bucket policies or network configurations can grant KMS permissions, when actually KMS key access requires IAM policies or KMS key policies.

449
MCQeasy

A marketing analytics team needs to ingest customer transaction data from an on-premises PostgreSQL database into Amazon S3 for analysis. The data volume is about 10 GB daily, and the team wants to perform full refresh daily (truncate and load) into S3 as Parquet files. The company has a Direct Connect connection to AWS. The team needs a simple, managed solution that minimizes operational overhead. What should the team use?

A.Set up AWS Database Migration Service (DMS) to continuously replicate data to S3 in Parquet format.
B.Use Amazon EMR with a Spark job that reads from PostgreSQL and writes to S3.
C.Use an AWS Glue ETL job with a JDBC connection to the PostgreSQL database, extract data, and write to S3 in Parquet format.
D.Use AWS Data Pipeline with a SQLActivity to extract data and copy to S3.
AnswerC

AWS Glue is serverless and managed, using a JDBC connection to read PostgreSQL and writing Parquet to S3, which minimises operational overhead for the daily 10 GB full refresh. It satisfies the stem's simplicity and managed-solution constraints without managing servers.

Why this answer

AWS Glue ETL jobs provide a serverless, managed environment to extract data from JDBC sources like PostgreSQL, transform it, and write to S3 in Parquet format. It minimizes operational overhead as it handles provisioning, scaling, and job execution. For a daily full refresh, a Glue job can be scheduled to truncate and load data.

Exam trap

DEA-C01 often tests the choice between DMS and Glue for batch ingestion; DMS is for continuous replication, while Glue is for batch ETL with transformations.

How to eliminate wrong answers

Option A is wrong because DMS is designed for continuous replication, not full refresh truncate-and-load; it would require additional steps to truncate and may not produce Parquet directly without transformation. Option B is wrong because Amazon EMR requires managing clusters, which increases operational overhead. Option D is wrong because AWS Data Pipeline is a legacy service and less managed than Glue; it also requires more configuration.

450
Multi-Selectmedium

A data engineer is configuring an AWS Glue job bookmark on a job that reads partitioned Parquet data from Amazon S3 and writes to another S3 location. The engineer notices that reprocessing keeps occurring and wants the bookmark to correctly skip already-processed data. Which two actions should the engineer take? (Choose two.)

Select 2 answers
A.Ensure the job reads from a source that supports bookmarks, such as S3 or the Glue Data Catalog, rather than an unsupported source
B.Increase the number of worker nodes so the bookmark commits faster between runs
C.Enable the Glue Data Catalog schema evolution setting on the job
D.Set the job's bookmark option to enable and confirm no manual state reset was performed between runs
E.Convert the output to a single unpartitioned file so the bookmark can track it
AnswersA, D

Job bookmarks only track state for supported sources, primarily Amazon S3 and JDBC/Glue Data Catalog tables. If the transform reads from an unsupported origin, Glue cannot persist progress and the job reprocesses everything each run. Confirming the source is a bookmark-capable type is a prerequisite for the feature to function at all in this pipeline.

Why this answer

Reprocessing under a bookmark usually means the feature is not actually tracking the source. The bookmark must be enabled on the job and the source must be a supported type such as S3 or the Glue Data Catalog, with no manual reset in between. Confirming both restores incremental processing so already-consumed partitions are skipped.

Exam trap

The trap here is treating job bookmarks as automatic, when they must be explicitly enabled and only work against supported sources with intact state.

Page 5

Page 6 of 18

Page 7