Courseiva

AWS Certified Data Engineer Associate DEA-C01 (DEA-C01) — Questions 601–675

1321 questions total · 18pages · All types, answers revealed

Page 8

Page 9 of 18

Page 10
601
MCQmedium

A data engineer is responsible for an Amazon Redshift cluster that stores financial data. The security team requires that all connections to the cluster from outside the VPC use SSL, and that the cluster's audit logs capture connection and user activity. The engineer has already enabled audit logging to Amazon S3. Which additional configuration should the engineer apply to meet the SSL requirement?

A.Set the Redshift cluster parameter 'require_ssl' to true in the parameter group associated with the cluster.
B.Attach an IAM policy to the Redshift cluster that denies connections without the 'aws:SecureTransport' condition.
C.Create a VPC endpoint for Redshift and require that all clients use the endpoint's private IP address.
D.Enable Redshift Spectrum and configure it to use SSL for all external table access.
AnswerA

The require_ssl parameter in a Redshift parameter group enforces SSL for all connections to the cluster. When set to true, clients that do not use SSL are rejected. This is the standard cluster-level setting for meeting encryption-in-transit requirements, and it applies to connections from outside the VPC as well as inside, ensuring consistent enforcement.

Why this answer

The require_ssl parameter in the Redshift parameter group is the correct way to enforce SSL for all connections to the cluster. When set to true, any client that attempts to connect without SSL is rejected. This directly meets the security team's requirement for encrypted connections from outside the VPC and works alongside audit logging, which the engineer has already enabled.

Exam trap

The trap here is confusing IAM policies and VPC endpoints with database-level SSL enforcement, when Redshift uses a cluster parameter for this purpose.

602
MCQmedium

A data engineer runs the AWS CLI command to retrieve the lifecycle configuration of the 'my-data-lake' bucket. The output is shown in the exhibit. What is the effect of this lifecycle policy?

A.Objects in the 'logs/' prefix are deleted after 365 days and their delete markers are removed.
B.All objects in the bucket are moved to STANDARD_IA after 30 days.
C.Objects in the 'logs/' prefix are moved to S3 Standard-IA after 30 days, to Glacier after 90 days, and deleted after 365 days.
D.Objects in the 'logs/' prefix are moved to Glacier after 90 days and expired after 90 days.
AnswerC

The lifecycle rule scopes transitions to the logs/ prefix: objects transition to S3 Standard-IA at day 30, then to Glacier at day 90, and expire at day 365. This matches the exhibit's prefix filter and staged transition actions, satisfying the retention requirement stated in the policy.

Why this answer

The lifecycle policy explicitly applies to the 'logs/' prefix, transitioning objects to S3 Standard-IA after 30 days, then to Glacier after 90 days, and finally expiring (deleting) them after 365 days. The 'Expiration' action with 'Days: 365' permanently removes the objects, while the 'Transitions' define the storage class changes at the specified intervals.

Exam trap

The trap here is that candidates often overlook the prefix filter and assume the policy applies to the entire bucket, or they misread the expiration as occurring at 90 days instead of 365 days, leading to incorrect answers like B or D.

How to eliminate wrong answers

Option A is wrong because the lifecycle policy does not include any action to remove delete markers; the 'Expiration' action simply deletes the objects after 365 days, and delete marker removal would require a separate 'ExpiredObjectDeleteMarker' setting. Option B is wrong because the policy only applies to objects under the 'logs/' prefix, not to all objects in the bucket, and the transition to STANDARD_IA occurs after 30 days, not immediately. Option D is wrong because it omits the initial transition to S3 Standard-IA after 30 days and incorrectly states that objects are expired after 90 days, whereas the actual expiration is after 365 days.

603
Multi-Selectmedium

A data engineer is optimizing an Amazon S3 data lake that stores large volumes of JSON logs. The engineer wants to reduce storage costs and improve query performance in Amazon Athena. Which TWO actions should the engineer take? (Choose two.)

Select 2 answers
A.Enable S3 Transfer Acceleration for the bucket.
B.Convert the JSON logs to Apache Parquet format.
C.Use S3 Standard-Infrequent Access (S3 Standard-IA) storage class for all log data.
D.Enable S3 server access logging on the bucket.
E.Partition the data by date and other commonly filtered columns.
AnswersB, E

Parquet is a columnar format that compresses data more efficiently than JSON and allows Athena to read only the columns needed for a query. This reduces storage costs and improves query performance by minimizing data scanned. Converting to Parquet is a best practice for optimizing Athena queries and reducing S3 storage costs.

Why this answer

Converting JSON logs to Parquet reduces storage size and improves Athena query performance by enabling columnar reads and better compression. Partitioning the data by commonly filtered columns like date allows Athena to skip irrelevant data, reducing the amount scanned and lowering costs. These two actions together address both storage cost reduction and query performance improvement.

Exam trap

The trap here is selecting actions that improve data transfer or auditing but do not affect storage costs or query performance in Athena.

604
MCQeasy

A data engineer needs to grant an IAM user access to query a specific table in Amazon Athena, but the user should not be able to view other tables in the same database. Which method should the engineer use?

A.Attach an IAM policy that allows athena:StartQueryExecution and restrict the query by table name
B.Use AWS Lake Formation to grant SELECT permission on the specific table to the user
C.Apply an S3 bucket policy that restricts access to the table's underlying data
D.Create a separate Athena workgroup with a query limit that only allows queries on that table
AnswerB

Lake Formation provides table-level grants within a database, so granting SELECT on the single table lets the user query it while other tables remain inaccessible. Athena alone cannot enforce such granular per-table restrictions through IAM policies.

Why this answer

AWS Lake Formation provides fine-grained access control at the table and column level, so granting SELECT permission on the specific table to the IAM user restricts them to that table while denying access to others in the same database. This is the intended service for table-level Athena permissions. It integrates with the Glue Data Catalog to enforce permissions at query time.

Exam trap

DEA-C01 often tests whether candidates know that IAM policies cannot filter by table name in Athena, so they pick an IAM policy or S3 bucket policy instead of Lake Formation for table-level access control.

How to eliminate wrong answers

Option A is wrong because an IAM policy allowing athena:StartQueryExecution cannot restrict queries by table name — IAM does not evaluate SQL content, so the user could query any table in the database. Option C is wrong because an S3 bucket policy restricts access to the underlying data objects, but Athena queries run under the user's or workgroup's role and the policy cannot express 'this table only' at the SQL level; it also does not prevent metadata access to other tables. Option D is wrong because an Athena workgroup controls query limits, encryption, and output location, not table-level permissions, so it cannot restrict which tables a user can query.

605
MCQeasy

A data engineer must load a 250 GB uncompressed CSV dataset from Amazon S3 into Amazon Redshift. The data must be loaded daily, and the engineer wants to minimize load time and avoid saturating the cluster's leader node. Which approach meets these requirements?

A.Use the COPY command with a single large file and the COMPUPDATE OFF option.
B.Use Amazon Redshift Spectrum to query the CSV files directly from Amazon S3 without loading them.
C.Use AWS Database Migration Service (AWS DMS) with a full load task to copy the CSV data into Amazon Redshift.
D.Use the COPY command with a manifest file that references multiple files in Amazon S3.
AnswerD

The COPY command can load multiple files in parallel when a manifest is provided. Each slice on the compute nodes reads a portion of the data, distributing the workload and avoiding leader node saturation. This is the recommended way to load large datasets from Amazon S3 into Amazon Redshift efficiently.

Why this answer

The COPY command is the most efficient way to load large datasets from Amazon S3 into Amazon Redshift. By using a manifest file that lists multiple S3 objects, the load is parallelized across all slices, reducing load time and preventing the leader node from becoming a bottleneck. This is the standard best practice for bulk data ingestion into Redshift.

Exam trap

The trap here is assuming that a single large file or an ETL service like AWS DMS is required for bulk loading, when the COPY command with a manifest is designed specifically for parallel, high-performance loads from Amazon S3.

606
MCQeasy

A data pipeline uses AWS Glue to process data from an S3 data lake. The pipeline fails intermittently with a 'ThrottlingException' when writing to a DynamoDB table. What is the MOST likely cause?

A.The DynamoDB table's write capacity is insufficient for the workload.
B.The network connection between Glue and DynamoDB is unstable.
C.The Glue job's timeout setting is too low.
D.The Glue job does not have sufficient IAM permissions to write to DynamoDB.
AnswerA

DynamoDB returns ThrottlingException when write requests exceed the table's provisioned write capacity units. Intermittent failures during Glue writes indicate the table's write capacity is insufficient for the pipeline's throughput, not a code or schema fault.

Why this answer

A ThrottlingException from DynamoDB indicates that the request rate to the table has exceeded the provisioned write capacity. AWS Glue jobs can generate high-throughput writes, and if the DynamoDB table's write capacity units (WCUs) are not sufficient to handle the burst, DynamoDB will throttle the requests. This is the most direct cause of the intermittent failure described.

Exam trap

The trap here is that candidates may confuse ThrottlingException with permission errors (Option D) or network issues (Option B), but AWS specifically tests the understanding that DynamoDB throttling is a capacity management mechanism, not a connectivity or authorization problem.

How to eliminate wrong answers

Option B is wrong because network instability between Glue and DynamoDB would typically result in connection timeouts or retryable network errors, not a specific ThrottlingException which is an application-level error from DynamoDB's API. Option C is wrong because a Glue job's timeout setting controls how long the job can run before being terminated, not how it handles individual API throttling errors; a timeout would cause a different error (e.g., 'Timeout exceeded'). Option D is wrong because insufficient IAM permissions would result in an AccessDeniedException, not a ThrottlingException; the error message directly indicates capacity limits, not authorization failures.

607
MCQeasy

A company stores IoT sensor data in S3 as JSON files. They need to convert the data to Parquet format for efficient querying with Amazon Athena. Which AWS service can perform this transformation with minimal effort?

A.Kinesis Data Firehose
B.Amazon Athena
C.AWS Glue ETL job
D.AWS Lambda
AnswerC

AWS Glue ETL jobs read JSON from S3, apply schema and transformation logic, and write Parquet back to S3 with minimal coding via the visual editor or generated scripts. This satisfies the conversion requirement while remaining serverless and low-effort for Athena querying.

Why this answer

AWS Glue ETL jobs can easily convert JSON to Parquet with built-in transforms. Option A is wrong because Kinesis Data Firehose is for streaming data ingestion, not batch transformations. Option B is wrong because Amazon Athena is a query engine, not a transformation service.

Option D is wrong because AWS Lambda is for small, event-driven transformations and is not ideal for large-scale batch conversion.

608
MCQhard

A data engineer is building an Amazon DynamoDB table that will store IoT sensor readings. Each reading has a deviceId (partition key) and a timestamp (sort key). The team wants to retrieve all readings for a device within the last 24 hours, and they also want to minimize the number of read capacity units consumed. Which access pattern should the engineer implement?

A.Create a global secondary index on timestamp and Query it with a key condition on the timestamp range.
B.Issue a Query operation with the partition key equal to the deviceId and a key condition expression on the sort key using a BETWEEN clause for the timestamp range.
C.Use a parallel Scan with multiple segments and a FilterExpression on deviceId.
D.Issue a Scan operation on the table with a FilterExpression on deviceId and timestamp.
AnswerB

Query targets a single partition, and adding a key condition on the sort key restricts the read to the requested timestamp range. DynamoDB reads only the items within that range, so consumed read capacity is proportional to the matching items, which is the most efficient way to satisfy the access pattern described.

Why this answer

The table already has deviceId as the partition key and timestamp as the sort key, which is the ideal design for retrieving a device's readings over a time range. Querying that partition with a BETWEEN condition on the sort key reads only the matching items, minimizing consumed read capacity. Scan, parallel Scan, and a GSI on timestamp all read or return excessive data.

Exam trap

The trap here is reaching for a secondary index or a Scan when the base table's composite key already supports the access pattern, wasting capacity and money.

609
MCQhard

A data engineer is using AWS Lake Formation to manage access to a data lake stored in Amazon S3. The engineer grants SELECT permission on a table in the AWS Glue Data Catalog to an IAM role used by an Amazon Athena user. However, the user still cannot query the table and receives an 'Access Denied' error. The engineer verifies that the IAM role has the necessary AWS Lake Formation permissions and that the S3 bucket policy allows access. What is the most likely cause of the issue?

A.The IAM role lacks permissions to the AWS Glue Data Catalog.
B.The S3 bucket policy does not grant access to the IAM role.
C.The IAM role requires an explicit DENY in the S3 bucket policy.
D.The AWS Lake Formation data lake location is not registered.
AnswerD

For Lake Formation to enforce and grant access to data in S3, the S3 location must be registered with Lake Formation. If the location is not registered, Lake Formation cannot manage permissions on the underlying data, and queries will fail with 'Access Denied' even if table-level permissions are granted. Registering the location allows Lake Formation to assume control of S3 access for that path.

Why this answer

The root cause is that the S3 data location is not registered with AWS Lake Formation. Even with Lake Formation table permissions, if the underlying S3 path is not registered, Lake Formation cannot manage access, and Athena queries fail with 'Access Denied'. Registering the location enables Lake Formation to control S3 permissions and grant access based on its policies.

Exam trap

The trap here is assuming that granting Lake Formation table permissions alone is sufficient, overlooking the need to register the S3 location with Lake Formation for it to manage data access.

610
MCQmedium

A company is using AWS Glue to process streaming data from Amazon Kinesis Data Streams. The job fails intermittently with a 'MemoryError' when the stream has a sudden spike in data volume. Which configuration change would best prevent this error?

A.Increase the number of DPUs (Data Processing Units) for the Glue job.
B.Store intermediate results in Amazon RDS.
C.Use a batch transformation instead of streaming.
D.Increase the number of shards in the Kinesis data stream.
AnswerA

MemoryError arises when executors exhaust heap during volume spikes. Adding DPUs increases the number of workers and total memory available, spreading the load across more executors. This directly addresses the memory constraint rather than tuning batch size or windowing.

Why this answer

Increasing the number of DPUs for the AWS Glue job provides more compute and memory resources to the job, which directly addresses the MemoryError caused by sudden data volume spikes. Glue DPUs are the unit of processing capacity, and scaling them up allows the job to handle larger volumes without running out of memory. This is the most direct configuration change to prevent memory-related failures.

Exam trap

The trap is confusing stream-side scaling (adding Kinesis shards) with job-side scaling (adding Glue DPUs), leading candidates to pick D when the error is in the Glue job's memory.

How to eliminate wrong answers

Option B is wrong because storing intermediate results in Amazon RDS does not increase the Glue job's memory and adds latency and complexity; it does not address the root cause of the MemoryError. Option C is wrong because switching to batch transformation changes the processing model and may not be feasible for a streaming use case; it also does not directly solve memory pressure during spikes. Option D is wrong because increasing Kinesis shards increases stream throughput but does not give the Glue job more memory; it could actually worsen the spike by delivering more data faster to an under-provisioned job.

611
MCQeasy

A company uses AWS Glue ETL jobs to transform data in Amazon S3. The data arrives in JSON format but needs to be converted to Parquet for efficient querying. Which AWS Glue feature should be used to infer the schema and generate transformation code?

A.Amazon S3 Select
B.Amazon Athena
C.Amazon Kinesis Data Analytics
D.AWS Glue crawlers
AnswerD

Crawlers populate the Data Catalog with schema information used by Glue ETL jobs.

Why this answer

AWS Glue crawlers are the correct feature because they automatically connect to data stores (like S3), infer the schema of JSON data by sampling it, and populate the AWS Glue Data Catalog with table definitions. This catalog schema can then be used by AWS Glue ETL jobs to generate transformation code (e.g., converting JSON to Parquet) without manual schema definition.

Exam trap

The trap here is that candidates confuse AWS Glue crawlers with Amazon Athena or S3 Select, assuming any query or analysis tool can infer schemas for ETL, but only crawlers are designed to automatically discover and catalog schemas for Glue ETL jobs.

How to eliminate wrong answers

Option A is wrong because Amazon S3 Select is a query-in-place service that retrieves subsets of data from S3 objects using SQL, but it does not infer schemas or generate ETL transformation code. Option B is wrong because Amazon Athena is an interactive query service that uses SQL to analyze data directly in S3, but it does not generate ETL transformation code or automatically infer schemas for Glue ETL jobs (though it can query Glue Data Catalog tables). Option C is wrong because Amazon Kinesis Data Analytics processes streaming data in real time using SQL or Apache Flink, not batch transformation of JSON to Parquet in S3, and it does not infer schemas for Glue ETL jobs.

612
MCQeasy

A company wants to transform data in Amazon S3 using SQL queries without provisioning servers. The transformations are ad-hoc and run occasionally. Which service should be used?

A.AWS Glue
B.Amazon Redshift Spectrum
C.Amazon EMR
D.Amazon Athena
AnswerD

Athena is serverless and supports SQL queries directly on S3 data.

Why this answer

Amazon Athena is the correct choice because it enables serverless, ad-hoc SQL querying directly on data stored in Amazon S3 without requiring any infrastructure provisioning. Since the transformations are occasional and ad-hoc, Athena's pay-per-query model and zero setup overhead align perfectly with the requirement.

Exam trap

The trap here is that candidates often confuse AWS Glue's ability to run SQL via Spark SQL or Athena as a transformation engine, but Glue requires provisioning resources and is not designed for ad-hoc serverless SQL queries.

How to eliminate wrong answers

Option A is wrong because AWS Glue is primarily an ETL service that requires provisioning and managing crawlers, jobs, and triggers, and is designed for scheduled or event-driven batch transformations, not lightweight ad-hoc SQL queries. Option B is wrong because Amazon Redshift Spectrum extends Redshift's query capabilities to S3 but still requires a provisioned Redshift cluster to run, which violates the 'without provisioning servers' constraint. Option C is wrong because Amazon EMR is a managed big data platform that requires provisioning EC2 instances and configuring clusters, making it unsuitable for occasional, serverless SQL queries.

613
MCQmedium

A data engineer is setting up an AWS Glue job that reads from an Amazon Kinesis Data Stream and writes to an Amazon S3 bucket. The security team requires that all data in transit be encrypted using TLS, and that the Glue job must use a VPC endpoint to access Kinesis and S3. Which configuration ensures compliance?

A.Attach a bucket policy to the S3 bucket that denies requests where aws:SecureTransport is false, and enable Kinesis Data Stream encryption in transit.
B.Enable default encryption on the S3 bucket and enable Kinesis Data Stream encryption at rest.
C.Configure the Glue job to use a VPC connection and create VPC endpoints for Kinesis and S3 with policies that enforce TLS.
D.Use AWS Glue job bookmarks and enable S3 Transfer Acceleration.
AnswerC

Using a VPC connection for the Glue job and creating VPC endpoints for Kinesis and S3 ensures that traffic goes through the endpoints. Endpoint policies can enforce TLS by denying requests where aws:SecureTransport is false. This meets the requirement for TLS and use of VPC endpoints.

Why this answer

To meet both TLS encryption in transit and VPC endpoint usage, the Glue job must run within a VPC using a VPC connection, and VPC endpoints for Kinesis and S3 must be created with policies that enforce TLS. This ensures all traffic goes through the endpoints and is encrypted. Other options either address only encryption at rest or do not include VPC endpoints.

Exam trap

The trap here is focusing only on encryption at rest or TLS without ensuring the use of VPC endpoints, which is an explicit requirement in the scenario.

614
MCQmedium

Refer to the exhibit. A data engineer runs the above CLI command and sees the output. The security team requires that the RDS instance not be accessible from the internet. Which change should the engineer make?

A.Change the storage type to io1 for better performance.
B.Modify the DB instance to set PubliclyAccessible to false.
C.Enable Multi-AZ deployment to improve security.
D.Update the VPC security group to deny inbound traffic from 0.0.0.0/0.
AnswerB

Setting PubliclyAccessible to false removes the instance's public IP and prevents internet routing, directly satisfying the security team's requirement. Security group rules alone do not make an RDS instance private, so this attribute change is the correct control.

Why this answer

Setting the `PubliclyAccessible` attribute to `false` ensures that the RDS instance is not assigned a public IP address and is not reachable from the internet. This directly satisfies the security team's requirement, as the instance will only be accessible from within the VPC. The CLI command shown modifies the DB instance, and this parameter is the standard AWS mechanism to control internet accessibility for RDS.

Exam trap

The DEA-C01 exam often tests the misconception that modifying a security group to block all inbound traffic is sufficient to make an RDS instance private, but the trap here is that the instance can still have a public IP address and be reachable from the internet if the security group rule is later removed or if the instance is in a public subnet.

How to eliminate wrong answers

Option A is wrong because changing the storage type to io1 (provisioned IOPS) improves I/O performance, not security or internet accessibility. Option C is wrong because enabling Multi-AZ deployment provides high availability and failover support, but does not restrict internet access; it can still leave the instance publicly accessible. Option D is wrong because updating the VPC security group to deny inbound traffic from 0.0.0.0/0 is a valid security measure, but it does not prevent the RDS instance from having a public IP address; the instance could still be assigned a public IP and be reachable if the security group rule is misconfigured or overridden, making this an incomplete solution compared to directly setting PubliclyAccessible to false.

615
MCQmedium

A data engineer is designing a data lake on Amazon S3 and needs to ensure that objects are automatically encrypted at rest using server-side encryption with AWS KMS. Which bucket policy statement achieves this?

A.Deny PutObject requests where the x-amz-server-side-encryption header is not set to aws:kms.
B.Deny PutObject requests that do not include the x-amz-server-side-encryption header.
C.Deny PutObject requests where the x-amz-server-side-encryption header is not set to AES256.
D.Allow PutObject requests only if the x-amz-server-side-encryption header is set to AES256.
AnswerA

A bucket policy condition on `s3:PutObject` can require the `x-amz-server-side-encryption` header to equal `aws:kms`, denying any upload that omits SSE-KMS. This enforces encryption at rest with AWS KMS at the API layer, satisfying the stem's automatic SSE-KMS requirement regardless of client defaults.

Why this answer

It enforces server-side encryption with AWS KMS (SSE-KMS) by denying any PutObject request that does not include the `x-amz-server-side-encryption` header set to `aws:kms`. This bucket policy ensures that all objects written to the S3 bucket are automatically encrypted at rest using AWS KMS, meeting the requirement for mandatory encryption with a specific key management service.

Exam trap

The trap here is that candidates often confuse the encryption header values (`aws:kms` vs `AES256`) and mistakenly choose an option that enforces SSE-S3 (AES256) instead of SSE-KMS, or they pick a Deny statement that only checks for the presence of the header without validating its specific value.

How to eliminate wrong answers

Option B is wrong because it denies PutObject requests that do not include the `x-amz-server-side-encryption` header at all, but it does not enforce the use of `aws:kms`; a request with the header set to `AES256` (SSE-S3) would still be denied, which is overly restrictive and not aligned with the requirement for KMS encryption. Option C is wrong because it denies PutObject requests where the header is not set to `AES256`, which would enforce SSE-S3 instead of SSE-KMS, directly contradicting the requirement for AWS KMS encryption. Option D is wrong because it allows PutObject requests only if the header is set to `AES256`, which again enforces SSE-S3, not SSE-KMS, and an Allow statement alone does not block requests that omit the header entirely, leaving a gap for unencrypted uploads.

616
MCQmedium

A data engineer manages an Amazon DynamoDB table that stores IoT sensor readings. Each item has a partition key of deviceId and a sort key of timestamp. The table is configured with on-demand capacity mode. The engineer needs to retrieve all readings for a specific device within the last 24 hours, and the query must return results sorted by timestamp in ascending order. Which operation should the engineer use?

A.Call the Query API with the partition key value and a condition on the sort key using the BETWEEN operator.
B.Call the GetItem API with the partition key and sort key values.
C.Call the Scan API with a FilterExpression on deviceId and timestamp.
D.Call the BatchGetItem API with a list of partition and sort key pairs.
AnswerA

The Query API requires a partition key value and allows an optional sort key condition. Using BETWEEN on the timestamp sort key efficiently retrieves the desired range. Results are returned in sorted order by the sort key, satisfying the ascending requirement. This is the most efficient and appropriate operation for this access pattern.

Why this answer

The Query API is designed for efficient retrieval of items sharing a partition key, with optional sort key conditions. Using BETWEEN on the timestamp sort key retrieves exactly the last 24 hours of readings for a device. Results are automatically sorted by the sort key, meeting the ascending order requirement.

Scan, GetItem, and BatchGetItem do not provide this combination of range filtering and sorted output.

Exam trap

The trap here is assuming that Scan with a filter is equivalent to Query for range-based access patterns, but Scan is inefficient and does not return sorted results.

617
MCQhard

A data engineer is managing an Amazon DynamoDB table that stores user session data. The table has a partition key of user_id and a sort key of session_start_time. The engineer needs to retrieve all sessions for a specific user that started within the last 30 days. Which operation should the engineer use to achieve the lowest latency?

A.Use the Scan operation with a filter expression on user_id and session_start_time.
B.Use the GetItem operation with the user_id and a session_start_time range.
C.Use the Query operation with the partition key set to the user_id and a condition on the sort key to filter sessions from the last 30 days.
D.Create a global secondary index (GSI) on user_id and session_start_time, then query the index.
AnswerC

The Query operation retrieves items with a specific partition key and can apply a condition on the sort key. Since the table's sort key is session_start_time, the engineer can query for a specific user_id and use a condition like session_start_time >= <30 days ago>. This efficiently retrieves only the relevant items and provides low latency by avoiding a full table scan.

Why this answer

The Query operation is designed to retrieve multiple items with the same partition key and can filter on the sort key. Since the table's schema matches the access pattern, querying the base table with a sort key condition yields the lowest latency and consumes minimal read capacity. Other operations either scan the entire table or are not suited for range retrievals.

Exam trap

The trap here is thinking that a GSI is needed for efficient queries, but the base table already has the right key schema, so a GSI would add unnecessary overhead.

618
MCQeasy

An e-commerce company wants to capture clickstream data from its website and store it in Amazon S3 for analytics. The data arrives continuously and the company needs near-real-time processing. Which solution is most appropriate?

A.AWS Data Pipeline
B.AWS Snowball Edge
C.Amazon Kinesis Data Firehose
D.Amazon S3 Transfer Acceleration
AnswerC

Amazon Kinesis Data Firehose ingests streaming clickstream data continuously and delivers it into Amazon S3 with near-real-time buffering, satisfying the stem's continuous-arrival and near-real-time requirements without custom consumer code. It handles scaling and delivery automatically, unlike batch uploads or query-based approaches.

Why this answer

Amazon Kinesis Data Firehose is the most appropriate solution because it is a fully managed service designed to ingest streaming data and deliver it to destinations like Amazon S3 with near-real-time latency. The company needs continuous clickstream capture and near-real-time processing, which Firehose provides. Option A (AWS Data Pipeline) is for batch processing, not streaming.

Option B (AWS Snowball Edge) is for offline data transfer, not real-time. Option D (S3 Transfer Acceleration) improves upload speed but is not a streaming ingestion service.

619
MCQmedium

A data engineer maintains an Amazon Kinesis Data Firehose delivery stream that writes JSON records to Amazon S3 and then invokes an AWS Lambda function for transformation. The Lambda function occasionally times out, causing records to be delivered untransformed. The engineer must ensure failed records are captured for later reprocessing without blocking delivery. What should the engineer do?

A.Enable the Kinesis Data Firehose data transformation failure option to send failed records to a separate S3 bucket for later reprocessing.
B.Increase the Lambda function's timeout to 15 minutes and memory to 10 GB so transformations always complete.
C.Configure the delivery stream to use Amazon Redshift as the destination and enable S3 backup for all records.
D.Attach a dead-letter queue to the Lambda function and configure Kinesis Data Firehose to read from that queue.
AnswerA

Kinesis Data Firehose supports a processing configuration with a Lambda function and a failure destination for records that fail transformation. Enabling that option writes failed records to a designated S3 bucket so they can be reprocessed later, while successful records continue to the main destination. This meets the requirement without blocking delivery.

Why this answer

Kinesis Data Firehose transformation supports a failure destination for records that the Lambda function cannot process. Enabling it sends failed records to a separate S3 bucket for later reprocessing while successful records continue to the primary destination. The other options either mask timeouts, change the destination without capturing failures, or use a dead-letter queue mechanism that does not apply to synchronous Firehose invocations.

Exam trap

The trap here is assuming a Lambda dead-letter queue captures Kinesis Data Firehose transformation failures, when Firehose invokes the function synchronously and cannot consume from an SQS queue.

620
MCQhard

Refer to the exhibit. A data engineer is configuring an IAM policy for a Lambda function that writes transformed data to S3. The function writes to both 'example-bucket/data/' and 'example-bucket/public/'. The policy is intended to enforce server-side encryption with SSE-S3 for all objects written to the 'public/' prefix, while allowing all operations on other prefixes. However, the Lambda function is failing with an AccessDenied error when writing to 'example-bucket/public/'. What is the most likely cause?

A.The policy denies DeleteObject on 'public/'.
B.The policy denies PutObject on 'public/' unconditionally.
C.The policy does not allow GetObject for 'public/'.
D.The Lambda function is not setting the 'x-amz-server-side-encryption' header to 'AES256' when writing to 'public/'.
AnswerD

The policy's `s3:PutObject` condition requires the `x-amz-server-side-encryption` header to equal `AES256` for the `public/` prefix. Because the Lambda function omits this header, the request fails the condition and S3 returns AccessDenied, satisfying the stem's encryption-enforcement constraint.

Why this answer

The policy enforces SSE-S3 encryption for objects written to the 'public/' prefix. When a Lambda function writes to S3 without setting the 'x-amz-server-side-encryption' header to 'AES256', the request fails with an AccessDenied error if the bucket policy requires SSE-S3. The policy explicitly denies PutObject unless the encryption header is present, so the function must include this header to succeed.

Exam trap

The DEA-C01 exam often tests the nuance that bucket policies can conditionally deny operations based on request headers, and candidates mistakenly think the error is due to missing IAM permissions rather than a missing encryption header.

How to eliminate wrong answers

Option A is wrong because the error is about writing (PutObject), not deleting (DeleteObject), and the policy focuses on encryption enforcement, not delete permissions. Option B is wrong because the policy does not unconditionally deny PutObject; it denies PutObject only when the required SSE-S3 encryption header is missing, which is a conditional denial. Option C is wrong because GetObject is not relevant to the write operation failing; the error occurs during PutObject, and the policy does not restrict read access for 'public/'.

621
MCQmedium

A company is using an Amazon RDS for MySQL database for an e-commerce application. During a sales event, the database experiences high read traffic, causing slow query performance. The company wants to reduce the read load on the primary database without changing the application code. Which solution meets these requirements?

A.Enable Multi-AZ on the RDS instance.
B.Create an Amazon RDS read replica and direct read traffic to it.
C.Increase the instance size of the RDS database.
D.Deploy Amazon ElastiCache to cache query results.
AnswerB

Read replicas handle read-only traffic, reducing load on the primary.

Why this answer

An Amazon RDS read replica is a read-only copy of the primary database that can offload read traffic without requiring any application code changes. By directing read queries to the replica, the primary database's load is reduced, improving performance during high-read events. This solution is specifically designed for read-heavy workloads and integrates seamlessly with existing MySQL connections.

Exam trap

The trap here is that candidates often confuse Multi-AZ with read replicas, assuming Multi-AZ provides read scaling, when in fact Multi-AZ only ensures failover redundancy and does not serve read traffic.

How to eliminate wrong answers

Option A is wrong because enabling Multi-AZ provides high availability through a standby replica in a different Availability Zone, but it does not offload read traffic; the standby is not used for reads unless a failover occurs. Option C is wrong because increasing the instance size scales the primary database vertically, which can improve performance but does not reduce read load on the primary instance and may incur higher costs without addressing the read traffic distribution. Option D is wrong because deploying Amazon ElastiCache caches query results in memory, which can reduce database load, but it requires application code changes to implement caching logic, violating the requirement of no code changes.

622
Multi-Selecteasy

A data engineer is designing a data ingestion pipeline for real-time clickstream data. Which TWO services can be used to ingest the data into Amazon Kinesis Data Streams?

Select 2 answers
A.Amazon S3
B.Kinesis Producer Library (KPL)
C.Kinesis Data Firehose
D.AWS SDK
E.AWS Glue
AnswersB, D

The Kinesis Producer Library aggregates and batches records, then sends them to Kinesis Data Streams using the PutRecords API. It handles retries and throughput optimisation, satisfying the real-time clickstream ingestion requirement without custom serialisation or buffering code.

Why this answer

Options B and D are correct. The Kinesis Producer Library (KPL) is a library for producers to send data to Kinesis Data Streams. AWS SDK can also be used directly.

Option A is wrong because Amazon S3 is a storage service, not a producer. Option C is wrong because Kinesis Data Firehose is a downstream consumer or delivery service, not a producer. Option E is wrong because AWS Glue is an ETL service, not a producer.

623
MCQmedium

A data engineer is building an AWS Lambda function that processes records from an Amazon Kinesis data stream. The Lambda function needs to read from the stream and write processed data to an Amazon S3 bucket. The security team requires that all data in transit be encrypted using TLS, and that the Lambda function authenticate to Kinesis and S3 using temporary credentials. Which combination of configurations should the engineer use?

A.Store long-term IAM user access keys in Lambda environment variables, and use them to sign requests to Kinesis and S3 over HTTPS.
B.Assign an IAM role to the Lambda function with permissions for Kinesis and S3, and use the AWS SDK for Python (Boto3) to read and write data, relying on the SDK's default TLS.
C.Configure the Lambda function to use AWS KMS to encrypt the data before sending it to Kinesis and S3, and use an IAM role for authentication.
D.Use Amazon Cognito identity pools to provide temporary credentials to the Lambda function, and enable TLS on the Kinesis and S3 clients.
AnswerB

Lambda functions assume an IAM role that provides temporary credentials automatically. The AWS SDKs use TLS by default for all service communications, including Kinesis and S3. This setup meets the requirements for temporary credentials and encryption in transit. No additional configuration is needed for TLS; it is enforced by the SDK endpoints.

Why this answer

AWS Lambda functions use an IAM role to obtain temporary credentials automatically. The AWS SDKs, such as Boto3, use TLS for all API calls by default, ensuring encryption in transit. This combination satisfies both the authentication and encryption requirements without manual key management.

Other options either use long-term credentials, confuse encryption at rest with in transit, or use an inappropriate identity service.

Exam trap

The trap here is thinking that TLS must be explicitly configured or that KMS encryption is needed for data in transit.

624
MCQhard

A company has multiple AWS accounts and wants to centrally manage permissions and access to data lakes. They have enabled AWS Organizations and want to use a single set of policies that apply to all accounts. Which policy type should be used at the organization level?

A.IAM policies
B.KMS key policies
C.S3 bucket policies
D.Service control policies (SCPs)
AnswerD

SCPs are the only AWS Organizations policy type that sets permission guardrails inherited by every member account, centrally restricting what principals can do. Applied at the organisation or OU level, they satisfy the requirement for one policy set governing all accounts.

Why this answer

Service Control Policies (SCPs) are used in AWS Organizations to centrally manage permissions across accounts. Option A (IAM policies) are attached to IAM users/roles within an account, not across accounts. Option B (KMS key policies) control access to KMS keys.

Option C (S3 bucket policies) are specific to S3 buckets.

625
Multi-Selectmedium

Which TWO actions should a data engineer take to protect sensitive data in an Amazon S3 bucket from being accessed by unauthorized users? (Select TWO.)

Select 2 answers
A.Create a VPC endpoint for S3
B.Enable S3 server access logging
C.Add a bucket policy with a Deny effect for unauthorized principals
D.Enable AWS CloudTrail for the bucket
E.Enable S3 Block Public Access
AnswersC, E

A bucket policy with a Deny effect explicitly overrides any Allow granted elsewhere, including IAM policies, so unauthorised principals are blocked regardless of other permissions. This directly satisfies the requirement to prevent access by unauthorised users.

Why this answer

Option C is correct because an S3 bucket policy with an explicit Deny effect overrides any Allow in IAM or bucket policies, so it reliably blocks the specified unauthorized principals from accessing the sensitive objects. Option E is correct because S3 Block Public Access settings prevent public ACLs and public bucket policies from exposing the bucket, closing off the most common accidental data-exposure path. Options A, B, and D do not restrict access: a VPC endpoint only controls how traffic reaches S3 (not who is authorized), server access logging only records requests after they occur, and CloudTrail only provides audit logging of API activity rather than preventing unauthorized access.

626
MCQeasy

A data engineer is setting up an Amazon S3 bucket to store large CSV files that will be queried using Amazon Athena. The engineer wants to minimize query costs and improve performance. The files are currently stored in a single prefix without any partitioning. The most common queries filter data by `year` and `month`. What should the engineer do to optimize the Athena queries?

A.Compress the CSV files using gzip and keep them in a single prefix.
B.Convert the CSV files to Apache Parquet format and partition the data by `year` and `month`.
C.Create an AWS Glue Data Catalog table with partition keys `year` and `month`, and use Amazon Athena to query the CSV files directly.
D.Use Amazon Redshift Spectrum to query the CSV files in S3 instead of Athena.
AnswerB

Converting to Parquet reduces the amount of data scanned because Parquet is columnar and compressed, and partitioning by `year` and `month` allows Athena to prune partitions based on query filters. This significantly lowers query cost and improves performance. Athena charges based on data scanned, so both optimizations directly reduce cost. This is a best practice for optimizing Athena queries on large datasets.

Why this answer

Converting data to Apache Parquet and partitioning by commonly filtered columns like `year` and `month` are the most effective ways to reduce data scanned by Athena. Parquet's columnar format allows Athena to read only the columns needed, and partitioning enables partition pruning. Together, they minimize query cost and improve performance.

Other options either do not address the core optimizations or introduce unnecessary complexity.

Exam trap

The trap here is thinking that simply defining partition keys in the Glue Data Catalog without reorganizing the data will enable partition pruning.

627
MCQmedium

Refer to the exhibit. A data engineer is troubleshooting a Kinesis Data Streams consumer that is falling behind. The stream has 2 shards and is receiving data at a rate of 2 MB/s. The consumer is an AWS Lambda function with a batch size of 100 records. What should the engineer do to improve consumer throughput?

A.Decrease the Lambda batch size to 10 records
B.Increase the retention period of the stream to 168 hours
C.Increase the number of shards in the stream to 4
D.Increase the memory allocation of the Lambda function
AnswerC

More shards increase parallelism and throughput for both producers and consumers.

Why this answer

Increasing the number of shards to 4 doubles the stream's total ingestion capacity to 4 MB/s, which directly increases the number of concurrent Lambda invocations and thus consumer throughput. With 2 shards, each shard can support up to 1 MB/s input and 2 MB/s output, so the current 2 MB/s load is at the shard-level output limit, causing the consumer to fall behind.

Exam trap

The DEA-C01 exam often tests the misconception that Lambda memory or batch size adjustments are the primary levers for throughput, when in fact the shard count directly controls the parallelism and read capacity of a Kinesis stream consumer.

How to eliminate wrong answers

Option A is wrong because decreasing the batch size reduces the number of records processed per invocation, which increases the number of Lambda invocations and overhead, potentially worsening throughput rather than improving it. Option B is wrong because increasing the retention period (up to 365 days) only affects how long records are stored in the stream, not the rate at which the consumer can read or process data. Option D is wrong because while increasing Lambda memory can improve CPU performance, the bottleneck here is the shard-level read throughput limit (2 MB/s per shard for the consumer), not Lambda compute capacity.

628
MCQeasy

A data engineer needs to transform JSON data from an S3 bucket into Parquet format and load it into Amazon Redshift. The transformation must be performed incrementally as new data arrives. Which AWS service is BEST suited for this task?

A.Use AWS Lambda to transform the data on the fly and write to Redshift.
B.Use Amazon EMR with Apache Spark to transform the data and load it into Redshift.
C.Use AWS Glue to create an ETL job that runs on a schedule or trigger.
D.Use Amazon Kinesis Data Firehose to transform and load data into Redshift in real time.
AnswerC

AWS Glue provides managed, serverless Spark ETL with job bookmarks, enabling incremental processing of only new S3 objects. Its native Redshift connector and Parquet conversion satisfy the transformation and load requirements, while schedule or event triggers handle the incremental arrival constraint.

Why this answer

AWS Glue provides a serverless ETL service that can run jobs triggered by S3 events to transform data incrementally and load into Redshift. Option A (AWS Lambda) can be used for simple transformations but may hit time limits for complex transformations. Option B (Amazon EMR) is more suited for large-scale big data processing but requires cluster management.

Option D (Amazon Kinesis Data Firehose) is for streaming data, not for batch transformation of existing S3 objects.

629
MCQhard

A data engineer is using Amazon Managed Workflows for Apache Airflow (MWAA) to orchestrate a data pipeline. The pipeline includes a task that runs an AWS Glue job. The engineer notices that the Glue job occasionally fails due to transient issues, and the Airflow task fails immediately without retrying. The engineer wants to configure the Airflow task to retry the Glue job up to 2 times with a 5-minute delay between retries. Which configuration in the Airflow DAG should the engineer use?

A.Set retries=2 and retry_delay=timedelta(minutes=5) in the default_args dictionary.
B.Set max_retries=2 and retry_interval=300 in the Glue job's default arguments.
C.Use the retry_exponential_backoff=True parameter in the DAG definition.
D.Configure the Airflow task to use the GlueJobOperator with retry_limit=2 and retry_delay=300.
AnswerA

In Apache Airflow, the retries and retry_delay parameters in default_args control the number of retries and the delay between them for all tasks in the DAG. Setting retries to 2 and retry_delay to 5 minutes will automatically retry the failed Glue job task up to two times with the specified delay. This is the standard way to configure retry behavior in Airflow.

Why this answer

In Amazon MWAA, retry behavior for tasks is controlled by Airflow's built-in retry parameters. Setting retries and retry_delay in default_args applies to all tasks, including those running Glue jobs. This ensures that transient failures are retried with the specified delay, improving pipeline resilience.

Exam trap

The trap here is confusing retry configuration at the orchestration layer (Airflow) with retry settings on the Glue job itself, leading to attempts to set non-existent Glue parameters.

630
MCQhard

A company is ingesting streaming data from multiple sources using Amazon Kinesis Data Streams. The data is then processed by an AWS Lambda function that transforms the records and writes them to an Amazon S3 bucket. The Lambda function is failing intermittently with timeout errors. The average record size is 5 KB, and the shard count is 2. What is the MOST likely cause of the timeout errors?

A.The Lambda function timeout is set too low for the processing time required.
B.The Kinesis data retention period is too short, causing data to be lost before processing.
C.The Lambda function's reserved concurrency is set too low, causing throttling.
D.The Lambda function is receiving too many records per invocation, exceeding the 6 MB payload limit.
AnswerA

Lambda timeouts occur when execution exceeds the configured limit; intermittent failures with small 5 KB records and two shards point to processing duration, not throughput. Raising the function timeout addresses the stem's constraint directly, unlike shard or batch-size changes.

Why this answer

The Lambda function is timing out, indicating that the configured timeout is insufficient for the actual processing time. Lambda has a default timeout of 3 seconds, but it can be set from 1 second to 15 minutes. If the transformation logic or S3 write operation takes longer than the configured timeout, the function will fail with a timeout error.

Option B is incorrect because the Kinesis data retention period (default 24 hours) affects data availability, not Lambda execution time. Option C is incorrect because reserved concurrency controls the number of concurrent invocations and can cause throttling, not timeouts. Option D is incorrect because with an average record size of 5 KB, even the default batch size of 100 records results in only 500 KB per invocation, well below the 6 MB payload limit.

631
MCQmedium

A data engineer is troubleshooting a slow-running query on an Amazon Redshift cluster. The query involves joining two large tables. The engineer notices that the query plan shows a large number of distribution and broadcast operations. Which design change would most likely improve query performance?

A.Change the distribution style of both tables to ALL
B.Change the distribution style of both tables to KEY on the join column
C.Change the distribution style of both tables to EVEN
D.Add a sort key on the join column
AnswerB

Distributing both tables on the join column co-locates matching rows on the same slice, so Redshift performs local joins instead of redistributing or broadcasting data across nodes. This directly eliminates the distribution and broadcast operations the plan revealed, satisfying the stem's requirement to reduce network overhead during large-table joins.

Why this answer

Changing the distribution style of both tables to KEY on the join column ensures that rows with the same join key value are co-located on the same node. This eliminates the need for expensive broadcast or redistribution operations during the join, as Redshift can perform the join locally on each slice without moving data across the network.

Exam trap

The trap here is that candidates often confuse distribution and sort keys, thinking a sort key on the join column will reduce data movement, when in fact only distribution key alignment eliminates broadcast/redistribution operations in the query plan.

How to eliminate wrong answers

Option A is wrong because setting both tables to ALL distribution replicates the entire table to every node, which increases storage and maintenance overhead, and does not address the root cause of excessive data movement during joins; it can also degrade performance for large tables due to increased load and memory pressure. Option C is wrong because EVEN distribution distributes rows round-robin across nodes, which does not co-locate join keys and forces Redshift to redistribute or broadcast rows during the join, exacerbating the problem. Option D is wrong because adding a sort key on the join column improves the efficiency of range-restricted scans and merge joins but does not reduce the number of distribution or broadcast operations; the query plan's large number of such operations indicates a distribution mismatch, not a sorting issue.

632
MCQeasy

Refer to the exhibit. A data engineer runs this AWS Glue Data Catalog DDL statement to create a table. The CSV files in 's3://my-bucket/sales/' use a pipe delimiter (|) instead of a comma. What change is needed to correctly read the data?

A.Change the 'field.delim' property to '|'.
B.Change the LOCATION to read from a subfolder.
C.Add a partition projection configuration.
D.Run a crawler to detect the schema automatically.
AnswerA

Setting `field.delim` to `|` overrides the default comma separator, so the Glue SerDe splits each CSV row on the pipe character rather than commas. This satisfies the stem's constraint that the source files in `s3://my-bucket/sales/` are pipe-delimited, allowing columns to be parsed correctly.

Why this answer

The AWS Glue Data Catalog DDL statement uses the default 'field.delim' property, which expects comma-separated values. Since the CSV files use a pipe delimiter (|), the table will not parse rows correctly. Setting 'field.delim' to '|' in the SerDe properties tells the Hive-compatible SerDe to split on pipes instead of commas, enabling correct data ingestion.

Exam trap

The DEA-C01 exam often tests the misconception that changing the LOCATION or adding partition projection will fix parsing issues, when in fact the core problem is the SerDe delimiter property not matching the actual file format.

How to eliminate wrong answers

Option B is wrong because changing the LOCATION to a subfolder does not alter the delimiter interpretation; it only changes the source path, leaving the parsing issue unresolved. Option C is wrong because partition projection configuration optimizes partition pruning for partitioned tables, but it does not affect how individual records are parsed within files. Option D is wrong because running a crawler would detect the schema and delimiter automatically, but the question explicitly asks what change is needed to the given DDL statement, and a crawler is an alternative approach, not a modification to the existing DDL.

633
MCQeasy

A data engineer needs to transform JSON data into CSV format using AWS Glue. The transformation is simple and must be executed on a schedule. Which Glue component is MOST suitable?

A.Glue Crawler
B.Glue Data Catalog
C.Glue Development Endpoint
D.Glue ETL job
AnswerD

A Glue ETL job runs Apache Spark transformations, converting JSON to CSV, and supports scheduling via triggers. Glue crawlers only catalogue data and DataBrew is interactive, so the ETL job matches the simple scheduled transformation requirement.

Why this answer

A Glue ETL job is the component designed to run transformation scripts (Spark or Python shell) on data, and it can be scheduled via triggers. For a simple JSON-to-CSV transformation on a schedule, an ETL job is the correct choice.

Exam trap

DEA-C01 often tests the misconception that a Crawler performs transformation, when in fact Crawlers only catalog schema and ETL jobs do the actual data processing.

How to eliminate wrong answers

Option A is wrong because a Glue Crawler only discovers schema and populates the Data Catalog; it does not transform data. Option B is wrong because the Glue Data Catalog is a metadata repository, not an execution engine. Option C is wrong because a Development Endpoint is an interactive environment for authoring and testing ETL scripts, not a scheduled production job.

634
Multi-Selectmedium

An e-commerce company is building a near-real-time dashboard to monitor customer clickstream data. The data is ingested via Amazon Kinesis Data Streams, transformed using AWS Lambda, and stored in Amazon S3. The team needs to query the data using Amazon Athena. Which THREE steps should be taken to optimize cost and performance? (Choose three.)

Select 3 answers
A.Use AWS Glue Data Catalog to store the table metadata.
B.Store the data in JSON format for flexibility.
C.Convert the data to Apache Parquet or ORC format.
D.Compress the data using gzip or snappy.
E.Partition the data by date in S3 (e.g., year/month/day).
AnswersC, D, E

Columnar Parquet or ORC lets Athena read only the columns each query references rather than every field, and columnar encoding compresses better. This directly reduces bytes scanned, which is the metric Athena bills on, improving both cost and latency.

Why this answer

Option C is correct because converting the clickstream data to columnar formats like Apache Parquet or ORC lets Athena read only the columns referenced in each query and applies columnar compression and predicate pushdown, dramatically reducing bytes scanned and therefore cost and latency. Option D is correct because compressing data with gzip or snappy reduces the amount of data Athena must read from S3, lowering query cost and improving performance; snappy is especially effective with Parquet/ORC since it is splittable and column-oriented. Option E is correct because partitioning the S3 data by date (e.g., year/month/day) enables Athena partition pruning so queries that filter on time only scan the relevant prefixes instead of the entire dataset.

Option A is not among the marked answers: while the AWS Glue Data Catalog is the standard metastore Athena uses, defining table metadata is a prerequisite for querying rather than a cost/performance optimization step for the data itself. Option B is not correct because storing data as JSON is row-oriented and non-columnar, which forces Athena to scan and parse entire records, increasing bytes scanned and query cost compared with Parquet or ORC.

Exam trap

A common trap is to consider AWS Glue Data Catalog as an optimization step, but it is merely a requirement; the actual optimizations are compression, partitioning, and columnar formats.

635
MCQeasy

A data engineer must load data from an Amazon S3 bucket into an Amazon Redshift cluster as part of a nightly batch pipeline. The source files are already in Parquet format and include columns that map directly to the target table. The engineer wants the fastest, most cost-effective load method that avoids staging the data through an external service. Which command should the engineer use?

A.Redshift Spectrum external table query that selects from S3 into the target table.
B.INSERT statements generated from a Python script that reads the Parquet files.
C.COPY command referencing the S3 bucket with an IAM role.
D.AWS Glue ETL job that reads from S3 and writes to Redshift using the Spark connector.
AnswerC

The COPY command loads data directly from Amazon S3 into Redshift in parallel across slices, which is the fastest and most cost-effective path for Parquet files that already match the target schema. Using an IAM role grants secure access without embedding credentials. It avoids staging through any intermediate service and is the native, recommended bulk-load mechanism for this scenario.

Why this answer

The COPY command performs a parallel bulk load directly from Amazon S3 into Redshift, using an IAM role for secure access. Because the Parquet files already match the target table, no intermediate transformation is required, making COPY the fastest and most economical option. Scripted inserts, Glue, and Spectrum add overhead or serve different purposes.

Exam trap

The trap here is reaching for a managed ETL service or scripted inserts when the native parallel COPY command is both faster and cheaper for schema-aligned S3 data.

636
MCQeasy

A small startup is building a data pipeline to ingest customer orders from a web application into Amazon Redshift for analytics. The orders are written to an Amazon RDS MySQL database. The startup wants to replicate the orders to Redshift in near-real time (within 5 minutes) with minimal operational overhead. The data volume is low, averaging 100 new orders per minute. The startup has a single data engineer who is also responsible for other tasks. What is the simplest solution?

A.Use AWS Glue with a scheduled job every 5 minutes to copy data from MySQL to Redshift
B.Use Amazon EMR with Spark streaming to read from MySQL and write to Redshift
C.Use an AWS Lambda function to query MySQL every minute and insert into Redshift
D.Use AWS Database Migration Service (DMS) with continuous replication
AnswerD

AWS DMS with continuous replication (change data capture) streams ongoing inserts from RDS MySQL to Redshift within minutes, satisfying the near-real-time requirement. It is fully managed, so the lone data engineer avoids building and operating custom replication infrastructure.

Why this answer

AWS DMS can continuously replicate from MySQL to Redshift with minimal setup and low overhead. Option A (AWS Glue) is batch-oriented and may not meet the 5-minute latency. Option B (Amazon EMR) is overkill for low data volumes.

Option C (AWS Lambda) requires custom code and may not efficiently handle the replication.

637
Multi-Selecteasy

A data engineer needs to securely store database credentials for an RDS instance. Which TWO AWS services can be used?

Select 2 answers
A.AWS KMS
B.AWS Secrets Manager
C.AWS IAM
D.AWS CloudFormation
E.AWS Systems Manager Parameter Store
AnswersB, E

AWS Secrets Manager stores and rotates RDS credentials natively, with built-in rotation via Lambda and fine-grained IAM and KMS encryption. It satisfies the requirement to securely store database credentials rather than embedding them in code or configuration files.

Why this answer

AWS Secrets Manager (B) is correct because it is purpose-built to store, rotate, and retrieve secrets such as RDS database credentials, and it natively integrates with RDS for automatic credential rotation. AWS Systems Manager Parameter Store (E) is also correct because it can store database credentials as SecureString parameters, which are encrypted with AWS KMS, allowing the data engineer to retrieve them securely at runtime. AWS KMS (A) only provides encryption keys and cryptographic operations; it does not itself store credentials or secrets.

AWS IAM (C) manages identities, roles, and permissions, not secret values. AWS CloudFormation (D) is an infrastructure-as-code provisioning service and is not designed to store or retrieve database credentials.

Exam trap

DEA-C01 often tests the confusion between KMS (encryption keys) and Secrets Manager/Parameter Store (credential storage), and between IAM (permissions) and actual secret storage services.

638
Matchingmedium

Match each AWS Glue component to its role.

Drag a concept onto its matching description — or click a concept then click the description.

Concepts
Matches

Scans data sources and populates catalog

Central metadata repository

Transform and load data

Orchestrates multiple jobs and crawlers

Interactive development environment

Why these pairings

AWS Glue components work together for ETL. The Data Catalog stores metadata, Crawlers populate the catalog, Classifiers infer schema, and ETL Jobs execute transformations. Connections provide connectivity but are not listed here.

639
MCQmedium

A data engineer is using AWS Database Migration Service (AWS DMS) to replicate ongoing changes from an Amazon RDS for MySQL database to an Amazon S3 bucket in Parquet format. The replication task is configured with full load plus change data capture (CDC). After several hours, the engineer notices that the S3 bucket contains only the full load data and no incremental changes. Which action should the engineer take to ensure CDC changes are captured?

A.Increase the DMS replication instance size to handle the change volume.
B.Enable Multi-AZ on the source RDS for MySQL instance.
C.Modify the DMS task to use a different target endpoint for CDC data.
D.Enable binary logging on the source RDS for MySQL instance and set the binlog_format parameter to ROW.
AnswerD

AWS DMS requires binary logging enabled on the source MySQL instance for CDC, and the binlog format must be ROW to capture row-level changes. Without this, DMS cannot read the transaction log and will only perform the initial full load. Setting binlog_format to ROW is a prerequisite for ongoing replication, so this action directly enables CDC.

Why this answer

For AWS DMS to perform ongoing replication from MySQL, the source must have binary logging enabled and binlog_format set to ROW. Without these settings, DMS cannot read the transaction log to capture changes. The engineer must enable binary logging and set the format to ROW, then restart the task.

Other actions like scaling the instance or enabling Multi-AZ do not enable CDC.

Exam trap

The trap here is assuming that CDC will work automatically once the DMS task is set to full load plus CDC, without verifying source database prerequisites like binary logging.

640
MCQmedium

A company uses AWS Glue to catalog data in Amazon S3. The data includes personally identifiable information (PII). The security team requires that PII be masked when queried by users who are not data owners. Which AWS service should be used to enforce this requirement?

A.Use Amazon Macie to automatically redact PII from S3 objects.
B.Use IAM policies with condition keys to restrict access based on tags.
C.Use AWS Lake Formation to define column-level security and data masking.
D.Use Amazon S3 Object Lambda to transform data on the fly.
AnswerC

AWS Lake Formation enforces column-level security and data masking through its permissions model, filtering PII columns for non-owner users at query time. This satisfies the requirement that PII be masked for users who are not data owners, without duplicating data or altering the underlying S3 objects.

Why this answer

AWS Lake Formation allows you to define column-level security and data masking policies on tables cataloged in the AWS Glue Data Catalog. You can grant or deny access to specific columns and apply masking to sensitive columns like PII for users who are not data owners, enforcing the requirement at query time. This is the native AWS service for fine-grained access control on data lakes.

Exam trap

DEA-C01 often tests the misconception that Macie or S3 Object Lambda can enforce masking, when Lake Formation is the service designed for column-level security and data masking in a Glue catalog.

How to eliminate wrong answers

Option A is wrong because Amazon Macie identifies and alerts on sensitive data but does not redact or mask PII in S3 objects; it is a discovery and classification service, not an enforcement mechanism. Option B is wrong because IAM policies with condition keys can restrict access to S3 objects based on tags, but they cannot perform column-level masking or redact PII within query results. Option D is wrong because S3 Object Lambda can transform data on the fly, but it requires custom code and does not provide a declarative, centrally managed masking policy like Lake Formation; it is more complex and less integrated with the Glue catalog.

641
MCQhard

A data engineer is responsible for a data pipeline that uses Amazon S3 as a data lake, AWS Glue for ETL, and Amazon Athena for ad-hoc queries. The pipeline ingests CSV files from an external partner via SFTP into an S3 bucket. The files are then processed by a Glue job that converts them to Parquet and writes to a separate S3 bucket partitioned by date. The Glue job runs daily and is triggered by a scheduled CloudWatch Events rule. Recently, the data engineer noticed that some days the Glue job fails because of memory errors, and on those days the Athena queries that rely on the data return incomplete results. The engineer needs to ensure that the pipeline is resilient and that Athena queries always see a complete view of the data, even if the Glue job fails mid-run. The engineer also needs to minimize re-processing of data. Which course of action should the engineer take?

A.Increase the number of workers and the worker type to G.2X to handle the memory errors, and enable job retries.
B.Replace the Glue job with an AWS Lambda function that processes the CSV files and writes Parquet to S3, and use S3 Event Notifications to trigger the function.
C.Modify the Glue job to use job bookmarks for incremental processing and write the Parquet output to a temporary location, then use an S3 copy operation to move the data into the final partitioned location only after the job completes successfully.
D.Use Athena partition projection to automatically discover partitions and set up a retry mechanism using AWS Step Functions.
AnswerC

Writing Parquet to a temporary prefix and copying only after successful completion keeps the final partitioned location free of partial output, so Athena never reads incomplete data. Glue job bookmarks track previously processed S3 objects, satisfying the minimise re-processing constraint by skipping already-ingested files on retry.

Why this answer

Writing Parquet output to a temporary location and only copying it into the final partitioned path after the job succeeds ensures Athena never sees partial data from a failed run. Job bookmarks enable incremental processing so already-processed files aren't reprocessed, minimizing re-work. This combination directly addresses both resilience (atomic visibility) and efficiency (no duplicate processing).

Exam trap

DEA-C01 often tests whether candidates conflate 'fix the failure' (more workers, retries) with 'ensure atomic visibility' — the trap is choosing a scaling fix that doesn't address partial-data exposure to Athena.

How to eliminate wrong answers

Option A is wrong because increasing workers and retries only mitigates memory errors — it doesn't prevent Athena from seeing partial output if a run fails mid-write. Option B is wrong because Lambda has a 15-minute timeout and limited memory, making it unsuitable for large CSV-to-Parquet ETL, and it doesn't solve atomic visibility. Option D is wrong because partition projection only speeds up Athena partition discovery — it does nothing to prevent partial data visibility or reduce Glue reprocessing.

642
Multi-Selecthard

A data engineer is optimizing an AWS Glue ETL job that processes large Parquet files in Amazon S3. The job currently takes several hours to complete. The engineer wants to improve performance by tuning the job's execution parameters. Which TWO actions will MOST effectively reduce the job's runtime? (Choose two.)

Select 2 answers
A.Partition the input data in Amazon S3 and use predicate pushdown in the Glue job.
B.Use the Glue DynamicFrame instead of Spark DataFrame for all transformations.
C.Increase the number of DPUs allocated to the Glue job.
D.Increase the Glue job's timeout value to allow more time for completion.
E.Enable job bookmarks to track processed data.
AnswersA, C

Partitioning the input data (e.g., by date or category) allows Glue to read only relevant partitions, reducing I/O. Predicate pushdown pushes filter conditions to the data source, so only matching data is read. Together, they minimize the amount of data scanned and processed, significantly improving runtime for large datasets. This is a best practice for optimizing Glue ETL jobs.

Why this answer

To reduce the runtime of a Glue ETL job processing large Parquet files, increasing DPUs provides more compute resources for parallel processing, and partitioning the input data with predicate pushdown reduces the amount of data read. These two actions directly address performance bottlenecks. Job bookmarks, DynamicFrames, and timeout adjustments do not effectively reduce runtime for a single large job run.

Exam trap

The trap here is confusing features that improve incremental processing (like job bookmarks) with those that improve single-run performance, and assuming that DynamicFrames are always faster.

643
MCQmedium

A data engineer is configuring cross-account access so that an analytics AWS account can read objects from a data-lake S3 bucket in a producer account. The objects are encrypted with SSE-KMS using a customer managed key in the producer account. The engineer has already added a bucket policy granting s3:GetObject to the analytics account's IAM role. Reads still fail with AccessDenied. Which additional change is required?

A.Add s3:GetObjectVersion to the analytics role's IAM policy and include the s3:ExistingObjectTag condition in the bucket policy.
B.Enable S3 Block Public Access on the data-lake bucket and re-run the cross-account read.
C.Change the bucket's default encryption to SSE-S3 so that cross-account readers no longer need KMS permissions.
D.Attach an IAM policy to the analytics role allowing kms:Decrypt and kms:DescribeKey, and update the KMS key policy to allow the analytics account role to use the key.
AnswerD

SSE-KMS decryption requires two independent grants: the caller's identity policy must allow kms:Decrypt (and typically kms:DescribeKey), and the KMS key policy must permit that principal to use the key. The bucket policy alone only authorizes the S3 object read. Because the key lives in the producer account, its key policy must explicitly trust the analytics account role, otherwise KMS denies the data key unwrap and S3 returns AccessDenied.

Why this answer

Reading SSE-KMS objects across accounts needs authorization on both sides: the identity must be allowed kms:Decrypt and the KMS key policy must trust that identity. A bucket policy granting s3:GetObject is necessary but not sufficient, because S3 calls KMS to unwrap the object's data key, and KMS independently evaluates its key policy. Both the identity policy and key policy must permit the analytics role.

Exam trap

The trap here is assuming that a bucket policy granting s3:GetObject is enough for cross-account reads, when SSE-KMS also requires identity permissions and key-policy trust.

644
MCQmedium

A data engineer is designing a data lake on Amazon S3 and needs to store data in a format that supports schema evolution and efficient columnar storage. The data will be queried using Amazon Athena and Amazon Redshift Spectrum. The engineer wants to minimize storage costs and improve query performance. Which storage format should the engineer choose?

A.Apache Avro
B.CSV
C.Apache Parquet
D.JSON
AnswerC

Parquet is a columnar storage format that provides efficient compression and encoding schemes, reducing storage costs and improving query performance by allowing column pruning and predicate pushdown. It supports schema evolution and is widely used with Athena and Redshift Spectrum. This makes it ideal for the described requirements.

Why this answer

Apache Parquet is a columnar format that offers high compression and efficient encoding, reducing storage costs and enabling fast query performance through column pruning and predicate pushdown. It supports schema evolution and is well-integrated with Athena and Redshift Spectrum, making it the best choice.

Exam trap

The trap here is assuming that any format supporting schema evolution is sufficient, but columnar storage is critical for analytical query performance and cost reduction.

645
MCQmedium

A company uses Amazon S3 to store historical stock market data as CSV files. They run daily Amazon Athena queries to generate reports. Recently, the finance team reported that queries are timing out and costs have increased significantly. The data engineering team notices that the S3 bucket contains thousands of small files (average 100 KB) due to a misconfigured ingestion pipeline. They need to improve query performance and reduce costs without changing the existing reporting schedule. The team has access to AWS Glue and can create new tables. Which solution should they implement?

A.Partition the data by date and create a new Athena table with partitions.
B.Use S3 Select to filter rows within each file before Athena processes them.
C.Increase the Athena query timeout to 30 minutes.
D.Use AWS Glue ETL to read the CSV files, convert them to Parquet, and write them back to S3 in fewer, larger files.
AnswerD

Parquet is columnar, so Athena reads only referenced columns, and consolidating thousands of 100 KB files into fewer larger objects cuts per-file overhead and request costs. Glue ETL performs the conversion, improving performance while keeping the daily schedule unchanged.

Why this answer

Converting the thousands of small CSV files into fewer, larger Parquet files using AWS Glue ETL directly addresses the root cause of poor Athena performance and high costs. Parquet is a columnar format that reduces the amount of data scanned per query, and larger files minimize the overhead of S3 LIST and GET operations, improving throughput. This solution does not change the reporting schedule and leverages existing Glue capabilities to create new optimized tables.

Exam trap

The trap here is that candidates often assume partitioning (Option A) is a universal performance fix, but they overlook that partitioning does not address the 'small files problem' which is a distinct performance killer in Athena due to S3 request overhead and file open costs.

How to eliminate wrong answers

Option A is wrong because partitioning by date does not solve the problem of thousands of tiny files; while partitioning can help prune scanned data, the overhead of reading many small files per partition still causes high latency and cost due to excessive S3 API calls. Option B is wrong because S3 Select operates at the object level to filter rows within a single file, but it does not consolidate files or change the file format; Athena would still need to process thousands of small files, and S3 Select cannot be used directly within Athena queries to replace table scans. Option C is wrong because increasing the query timeout does not reduce the amount of data scanned or the number of S3 requests; it merely allows the query to run longer without addressing the performance bottleneck or cost issue.

646
MCQmedium

A company is migrating an on-premises Hadoop cluster to AWS. The cluster processes large files in CSV format using Apache Spark. Which data store should be used as the primary storage for the data lake to optimize cost and performance?

A.Amazon EMR File System (EMRFS) backed by HDFS
B.Amazon RDS for MySQL
C.Amazon EBS volumes attached to the EMR cluster
D.Amazon S3
AnswerD

Amazon S3 provides durable, virtually unlimited object storage with separate compute and storage scaling, so Spark reads CSV files directly and cost stays low. HDFS on Amazon EMR couples storage to cluster lifetime, raising cost and complicating elasticity.

Why this answer

Amazon S3 is the optimal primary storage for a data lake on AWS because it offers high durability, scalability, and cost-effectiveness. It integrates seamlessly with Apache Spark on Amazon EMR, allowing direct access to data without moving it. S3 also supports various file formats and decouples storage from compute, enabling independent scaling.

Exam trap

DEA-C01 often tests the misconception that HDFS or EBS is suitable for a data lake, but they are not cost-effective or scalable for long-term storage. Candidates might choose EMRFS with HDFS due to familiarity with Hadoop.

How to eliminate wrong answers

Option A is wrong because EMRFS backed by HDFS is not cost-effective for a data lake; HDFS is designed for temporary storage on cluster nodes and is not durable or scalable for long-term data. Option B is wrong because Amazon RDS for MySQL is a relational database, not suitable for storing large CSV files for big data processing. Option C is wrong because EBS volumes are block storage attached to EC2 instances, not a shared data lake storage; they are limited in size and not durable across cluster termination.

647
MCQhard

A financial services company stores transactional records in Amazon DynamoDB. Auditors require that any item be recoverable to its exact state from any point within the last 30 days, including after an accidental delete or overwrite caused by a faulty deployment. The table uses on-demand capacity and must remain highly available during recovery. Which feature should the data engineer enable?

A.DynamoDB global tables with multi-Region replication
B.DynamoDB on-demand backups
C.DynamoDB Streams with an AWS Lambda consumer writing to Amazon S3
D.DynamoDB point-in-time recovery (PITR)
AnswerD

PITR continuously backs up the table with per-second granularity and lets you restore to any point in the preceding 35 days, which covers the 30-day audit requirement. It protects against accidental writes and deletes because the restore uses the continuous backup stream, not a periodic snapshot. Restores create a new table with the chosen timestamp, preserving availability of the source table.

Why this answer

DynamoDB point-in-time recovery maintains a continuous backup with per-second granularity and supports restore to any moment within the preceding 35 days, which exceeds the 30-day audit window. Because it captures every write and delete, a faulty deployment that corrupts or removes items can be undone by restoring to a timestamp just before the incident. The restore produces a new table, so the production table stays available throughout.

Exam trap

The trap here is confusing periodic on-demand backups or multi-Region replication with continuous point-in-time recovery, when only PITR can restore to an arbitrary second after an accidental delete.

648
Multi-Selecthard

Which THREE factors should a data engineer consider when choosing between Amazon RDS and Amazon DynamoDB for a new application? (Choose three.)

Select 3 answers
A.Whether the workload requires serverless scaling.
B.Whether the data model is relational or key-value.
C.Whether the data must be encrypted at rest by default.
D.Whether the application requires VPC isolation.
E.Whether the application needs to scale horizontally for high throughput.
AnswersA, B, E

DynamoDB is serverless; RDS requires manual scaling.

Why this answer

Amazon RDS is a relational database service that requires provisioning and managing server capacity, while DynamoDB is a fully managed NoSQL key-value and document database that supports serverless scaling. Option A is correct because DynamoDB can automatically scale throughput capacity up or down based on traffic patterns, making it suitable for unpredictable workloads, whereas RDS requires manual scaling or the use of Auto Scaling with predefined policies.

Exam trap

The trap here is that candidates mistakenly think encryption at rest or VPC isolation are exclusive to one service, when in fact both RDS and DynamoDB support these features, making them irrelevant as differentiators.

649
MCQeasy

A data engineer needs to set up a new Amazon RDS for MySQL database for a web application. The application experiences variable read traffic and requires low read latency. The engineer needs to minimize downtime during maintenance and provide read scalability. Which configuration meets these requirements?

A.Multi-AZ db.r5.large instance with two Read Replicas
B.Multi-AZ db.r5.large instance
C.Single-AZ db.r5.large instance
D.Single-AZ db.r5.xlarge instance
AnswerA

Multi-AZ deployment provides automatic failover to a standby in a second Availability Zone, minimising downtime during maintenance, while Read Replicas offload read traffic to scale reads and reduce latency. This satisfies the variable read traffic, low read latency, and read scalability requirements.

Why this answer

A Multi-AZ deployment provides high availability and automatic failover to minimize downtime during maintenance, while adding two Read Replicas offloads read traffic from the primary instance, reducing read latency and enabling read scalability. The db.r5.large instance size is sufficient for the variable read workload, and Read Replicas can be promoted to standalone instances if needed.

Exam trap

The trap here is that candidates often assume Multi-AZ alone provides read scalability, but Multi-AZ only provides high availability and failover, not read offloading—Read Replicas are required for read scaling.

How to eliminate wrong answers

Option B is wrong because a Multi-AZ instance alone provides high availability and failover but does not offer read scalability or reduce read latency for variable read traffic, as all reads still hit the primary instance. Option C is wrong because a Single-AZ instance lacks high availability, meaning any maintenance or failure causes downtime, and it provides no read scalability. Option D is wrong because a Single-AZ db.r5.xlarge instance, while larger, still lacks high availability and read scalability; scaling vertically does not address variable read traffic efficiently and does not minimize downtime during maintenance.

650
MCQeasy

A data engineer is designing a data lake on Amazon S3. The data is ingested from multiple sources and needs to be partitioned by year, month, day, and event type for efficient querying with Amazon Athena. Which S3 key prefix structure is most appropriate?

A.s3://bucket/events/2024-01-01/event_type=data.parquet
B.s3://bucket/2024/01/01/event_type/events/data.parquet
C.s3://bucket/event_type=events/year=2024/month=01/day=01/data.parquet
D.s3://bucket/day=01/month=01/year=2024/event_type=events/data.parquet
AnswerC

Correct Hive-style partitioning with logical key order (year, month, day) and event type, enabling efficient partition pruning.

Why this answer

Uses Hive-style partitioning (event_type=events/year=2024/month=01/day=01), which Athena and other query engines natively support. This structure allows Athena to perform partition pruning, reading only the relevant directories based on WHERE clause filters, significantly reducing data scanned and improving query performance. Option D also uses Hive-style partitioning but with a different order of partition keys (day, month, year).

While still valid, this non-standard order may cause issues with automatic partition discovery when using MSCK REPAIR TABLE, which expects the partition order to match the table definition. Therefore, option C is the most appropriate because it follows the common convention of listing partitions from coarse to fine granularity (year > month > day) and can be easily loaded into Athena without additional configuration.

Exam trap

AWS often tests the distinction between Hive-style partitioning (key=value) and flat or date-only prefixes, where candidates mistakenly choose a structure that does not support partition pruning or is incompatible with Athena's partition discovery.

How to eliminate wrong answers

Option A is wrong because it embeds the date as a single prefix (2024-01-01) and places event_type as a filename suffix, which does not create separate partition directories; Athena cannot prune partitions efficiently without explicit partition columns. Option B is wrong because it uses a date-only hierarchy (year/month/day) but does not include event_type as a partition column, forcing full scans when filtering by event type. Option D is identical to C and is also correct, but the question expects the most appropriate structure; since both C and D are the same, the intended correct answer is C (the first occurrence).

651
MCQmedium

A data engineer is troubleshooting an AWS Glue ETL job that fails intermittently with the error 'Rate exceeded.' The job reads from an Amazon RDS for MySQL source and writes to Amazon S3. What is the MOST likely cause of this error?

A.The Glue job is using Amazon Kinesis Data Streams as a source, which has a shard throughput limit.
B.The number of Glue job workers or parallel queries is exceeding the maximum connections or IOPS of the RDS instance.
C.The Amazon S3 bucket has a bucket policy that limits the number of objects written per second.
D.The IAM role attached to the Glue job does not have sufficient permissions to read from RDS.
AnswerB

Glue workers open concurrent JDBC connections to the RDS for MySQL source. When worker count or parallel read queries exceed the instance's maximum connections or provisioned IOPS, the source rejects requests, surfacing as 'Rate exceeded' — an RDS-side throttling limit, not a Glue service quota.

Why this answer

The 'Rate exceeded' error in AWS Glue typically occurs when the job opens too many concurrent connections to the source database. With RDS for MySQL, each Glue worker or parallel read task establishes a connection, and exceeding the instance's max_connections or IOPS limits triggers throttling. This is the most likely cause when reading from RDS and writing to S3.

Exam trap

The trap is assuming the error comes from the target (S3) or from a misconfigured IAM role; candidates must recognize that 'Rate exceeded' often points to source database connection limits when using JDBC sources like RDS.

How to eliminate wrong answers

Option A is wrong because the source is explicitly Amazon RDS for MySQL, not Kinesis Data Streams; Kinesis shard limits would produce a different error and are not relevant here. Option C is wrong because S3 does not have a bucket policy that limits objects per second; S3 scales automatically and throttling would be reported as 503 Slow Down, not 'Rate exceeded' from RDS. Option D is wrong because insufficient IAM permissions would cause an AccessDenied error, not a rate limit error.

652
MCQmedium

A data engineer needs to share a dataset stored in Amazon S3 with another AWS account. The bucket policy currently grants access only to the owning account. What is the simplest way to grant cross-account access?

A.Add a bucket policy that grants access to the other account's IAM role
B.Set the object ACL to public-read
C.Use an S3 access control list (ACL) to grant access to the other account
D.Create an IAM role in the other account and attach a policy to it
AnswerA

Adding a bucket policy that names the other account's IAM role as principal grants cross-account access directly, satisfying the requirement to share the S3 dataset without extra infrastructure. S3 evaluates the resource-based policy against the requesting role, so no role switching, trust policy, or intermediate service is needed — the simplest mechanism available.

Why this answer

A bucket policy is the simplest and most direct way to grant cross-account access because it is a resource-based policy attached to the S3 bucket that can explicitly name the other account's IAM role or account ID as a principal. This avoids the need to create roles or modify ACLs, and it works even when ACLs are disabled (the default for new buckets).

Exam trap

DEA-C01 often tests the confusion between ACLs and bucket policies, and the misconception that creating a role in the other account is sufficient — candidates forget that the resource owner must also grant permission via a bucket policy.

How to eliminate wrong answers

Option B is wrong because setting an object ACL to public-read grants access to everyone on the internet, not just the other AWS account, creating a serious security exposure. Option C is wrong because S3 ACLs are legacy access controls that cannot grant permissions to IAM roles in another account — they only support canonical user IDs and predefined groups, and ACLs are disabled by default on modern buckets. Option D is wrong because creating an IAM role in the other account only defines an identity; without a bucket policy or cross-account trust, that role has no permission to access the bucket in the owning account.

653
MCQeasy

A data engineer is configuring an S3 bucket for a data lake. The engineer runs the command shown in the exhibit. What does the output indicate about the bucket?

A.Versioning is enabled on the bucket.
B.The bucket retains only the latest version of each object.
C.Versioning is suspended on the bucket.
D.MFA Delete is enabled for the bucket.
AnswerA

The command's output reports a Versioning configuration, and its Status field reads Enabled, confirming that object versions are retained rather than overwritten. This distinguishes versioning from bucket logging, encryption or lifecycle settings, which appear under different configuration keys.

Why this answer

The output of the command (likely 'aws s3api get-bucket-versioning') shows a Status of 'Enabled', indicating that versioning is enabled on the bucket. This means multiple versions of objects are retained.

Exam trap

DEA-C01 often tests the difference between versioning enabled and suspended; the output 'Enabled' means versioning is active, not suspended.

How to eliminate wrong answers

Option B is wrong because if versioning is enabled, all versions are retained, not just the latest. Option C is wrong because 'Suspended' would be the status if versioning were suspended. Option D is wrong because MFA Delete is a separate setting that requires MFA for deletion and is not indicated by the versioning status alone.

654
Multi-Selectmedium

A data engineer is designing a data pipeline that ingests streaming data from an IoT device fleet. The data must be processed in near real-time and stored in Amazon S3 for long-term analytics. Which TWO AWS services should the engineer use together to achieve this?

Select 2 answers
A.Amazon Athena
B.AWS Glue
C.Amazon Kinesis Data Firehose
D.Amazon Kinesis Data Streams
E.Amazon Simple Queue Service (SQS)
AnswersC, D

Kinesis Data Firehose provides a fully managed delivery stream that ingests streaming records and automatically batches, transforms and writes them into Amazon S3, satisfying the near real-time processing and long-term S3 storage requirements without custom consumer code.

Why this answer

Amazon Kinesis Data Streams (D) is correct because it provides a highly scalable, low-latency ingestion service for real-time streaming data from thousands of IoT devices, allowing custom consumers to process records in near real-time. Amazon Kinesis Data Firehose (C) is correct because it can consume that stream (or receive data directly) and reliably deliver it in near real-time to Amazon S3 for long-term analytics, handling batching, compression, and format conversion. Together they form the canonical AWS pattern for real-time IoT ingestion plus durable S3 storage.

Amazon Athena (A) is only a query service over S3 and does not ingest or process streaming data. AWS Glue (B) is a serverless ETL/catalog service, not a real-time streaming ingestion or delivery mechanism. Amazon SQS (E) is a message queue for decoupling applications, not designed for high-throughput real-time streaming ingestion into S3.

Exam trap

DEA-C01 often tests whether candidates can distinguish ingestion services (Kinesis) from storage/query services (Athena, Glue) and decoupling services (SQS), so picking Athena or Glue for the ingestion leg is the common mistake.

655
MCQmedium

A data engineer must give an AWS Lambda function temporary credentials to read objects from a specific Amazon S3 prefix. The function runs in a VPC and accesses S3 through a gateway VPC endpoint. The security team forbids long-lived access keys on the function. What should the data engineer configure?

A.Store an IAM user's access key ID and secret access key in Lambda environment variables and rotate them monthly.
B.Attach an IAM role to the Lambda function with an S3 read policy scoped to the prefix, and add an S3 bucket policy restricting access to the gateway VPC endpoint.
C.Use the Lambda function's reserved concurrency setting to limit access, and allow S3 access through the VPC endpoint policy only.
D.Create an IAM user and store the credentials in AWS Secrets Manager, then call the secret at runtime from Lambda.
AnswerB

Lambda assumes an execution role and receives temporary credentials through the environment, so no static keys are stored. Scoping the policy to the prefix and adding a bucket policy that allows access only through the specified VPC endpoint enforces least privilege and ensures traffic stays on the private path.

Why this answer

Lambda execution roles supply temporary credentials automatically, eliminating static keys while allowing a policy scoped to the required prefix. Combining that with a bucket policy that permits access only through the gateway VPC endpoint keeps the data path private and prevents access from outside the approved network route.

Exam trap

The trap here is assuming that storing credentials in Secrets Manager or environment variables removes the risk of long-lived keys, when only role-based temporary credentials satisfy the requirement.

656
MCQhard

A data engineer is troubleshooting an AWS Glue ETL job that fails with an 'Access Denied' error when trying to write to an S3 bucket. The IAM role used by the job has the policy shown in the exhibit. The bucket 'my-bucket' uses S3 default encryption with AWS KMS. What is the most likely missing permission?

A.s3:GetObjectVersion
B.glue:GetObject
C.s3:ListBucketMultipartUploads
D.s3:PutObjectAcl
E.kms:GenerateDataKey and kms:Decrypt
AnswerE

Writing to an SSE-KMS encrypted bucket requires the role to call KMS GenerateDataKey to obtain a data key for encrypting each object, plus Decrypt for reads and multipart uploads. The IAM policy grants only S3 actions, so the missing kms:GenerateDataKey and kms:Decrypt permissions cause the Access Denied.

Why this answer

When an S3 bucket uses AWS KMS for default encryption, any write operation requires the IAM role to have kms:GenerateDataKey and kms:Decrypt permissions on the KMS key. The policy in the exhibit grants s3:PutObject but does not include any KMS actions, resulting in an 'Access Denied' error. Option A (s3:GetObjectVersion) is not needed for writing.

Option B (glue:GetObject) is not a valid AWS action. Option C (s3:ListBucketMultipartUploads) is not required for a write operation. Option D (s3:PutObjectAcl) is unnecessary unless the job explicitly sets ACLs.

657
MCQhard

A data engineer is using AWS Step Functions to orchestrate a series of AWS Glue jobs. The engineer notices that when a Glue job fails, the Step Functions execution also fails, but the engineer wants to retry the failed job up to three times with exponential backoff before failing the entire workflow. The engineer needs to implement this with minimal changes to the state machine. What should the engineer do?

A.Configure the Glue job to automatically retry on failure by setting the MaxRetries parameter in the job's default arguments.
B.Create an Amazon CloudWatch alarm that triggers an AWS Lambda function to restart the Glue job on failure.
C.Modify the state machine to use a Map state that iterates over the Glue job three times.
D.Add a Retry field with MaxAttempts: 3 and BackoffRate: 2.0 to the state that invokes the Glue job.
AnswerD

The Retry field in Amazon States Language allows specifying retry behavior for a state, including maximum attempts and exponential backoff. Setting MaxAttempts to 3 and BackoffRate to 2.0 will retry the Glue job up to three times with exponential backoff. This is the standard and minimal change to achieve the requirement without altering the job itself.

Why this answer

AWS Step Functions provides native retry capabilities via the Retry field in the state definition. Specifying MaxAttempts and BackoffRate on the state that calls the Glue job enables automatic retries with exponential backoff. This is the simplest and most direct method to meet the requirement without altering the Glue job or adding external components.

Exam trap

The trap here is confusing retry mechanisms at the job level with those at the orchestration level, leading to attempts to configure retries in Glue instead of Step Functions.

658
MCQhard

A data engineer is using AWS Step Functions to orchestrate a series of AWS Glue jobs. During a recent run, a Glue job failed due to a transient network issue, and the Step Functions execution stopped. The engineer needs to ensure that the workflow can automatically retry the failed Glue job up to three times before considering the step failed. Which Step Functions state configuration should the engineer implement?

A.Set the TimeoutSeconds and HeartbeatSeconds fields in the Task state to trigger a retry on failure.
B.Add a Catch field to the Task state with ErrorEquals set to States.ALL and Next set to a retry state.
C.Add a Retry field to the Task state with ErrorEquals set to States.ALL and MaxAttempts set to 3.
D.Configure the Glue job itself to have a maximum retry count of 3 in its job definition.
AnswerC

The Retry field in a Step Functions Task state allows automatic retries on specified errors. Setting ErrorEquals to States.ALL catches all errors, and MaxAttempts to 3 retries the task up to three times. This is the correct way to handle transient failures in Glue jobs orchestrated by Step Functions. It ensures the workflow does not fail immediately and provides resilience.

Why this answer

The Retry field in Step Functions is designed to automatically retry a task when specified errors occur. By setting ErrorEquals to States.ALL and MaxAttempts to 3, the Glue job will be retried up to three times on any error, including transient network issues. The Catch field is for fallback, not retry.

Glue job definitions do not have a retry count, and timeout/heartbeat fields are for monitoring, not retry logic.

Exam trap

The trap here is confusing the Catch field with the Retry field, or assuming Glue jobs have built-in retry settings.

659
MCQhard

A data engineer is using AWS Lake Formation to manage fine-grained access control on an Amazon S3 data lake. The engineer has registered the S3 bucket as a Lake Formation data location and created a table in the AWS Glue Data Catalog. The engineer needs to grant a data analyst permission to query only specific columns (customer_id, order_date) in the sales table using Amazon Athena, while hiding other columns (credit_card_number, address). The analyst uses an IAM role that has no direct S3 permissions. Which action should the engineer take?

A.Grant the analyst's IAM role SELECT permission on the sales table in Lake Formation, and then use Lake Formation column-level security to include only the customer_id and order_date columns.
B.Create an IAM policy that allows s3:GetObject on the sales table's S3 prefix, and attach it to the analyst's role. Then use Athena to restrict columns via a view.
C.Use AWS Glue DataBrew to create a dataset that includes only the allowed columns, and grant the analyst access to that dataset.
D.Grant the analyst's IAM role DESCRIBE and SELECT on the sales table in Lake Formation, and then create an Athena view that selects only the allowed columns.
AnswerA

Lake Formation column-level security allows granting SELECT on specific columns. By granting SELECT on the table and then applying column filters, the analyst can query only the allowed columns. Since the analyst's role has no direct S3 permissions, Lake Formation provides temporary credentials for data access, enforcing the column-level restrictions.

Why this answer

Lake Formation column-level security is designed to grant SELECT on specific columns while hiding others. By granting SELECT on the table and then specifying included columns, the analyst can query only those columns. Because the analyst's IAM role lacks direct S3 permissions, all data access goes through Lake Formation, which enforces the column restrictions and provides temporary credentials.

This meets the requirement without granting broad S3 access.

Exam trap

The trap here is assuming that an Athena view or IAM policy can enforce column-level security when Lake Formation is managing the data lake. Only Lake Formation's column-level permissions prevent direct access to hidden columns.

660
MCQeasy

A data engineer is configuring an Amazon Redshift cluster for a reporting workload. The team needs to load data from Amazon S3 into a Redshift table and wants the fastest possible load while keeping the data compressed. Which approach should the engineer use?

A.Run individual INSERT statements for each row from an AWS Lambda function that reads the S3 objects.
B.Create an AWS Glue job that writes to Redshift using JDBC in small batches and commits after each batch.
C.Use Redshift Spectrum to create an external table over the S3 data and run CREATE TABLE AS SELECT to materialize it.
D.Use the COPY command to load data from S3 into the Redshift table, specifying the appropriate compression and format options.
AnswerD

COPY is Redshift's native parallel load utility. It reads from Amazon S3 using multiple slices in parallel, applies compression and format options such as GZIP, PARQUET, or ORC, and loads directly into the table. This is the recommended, most efficient path for bulk loads and avoids intermediate staging.

Why this answer

Amazon Redshift COPY is purpose-built for bulk ingestion from Amazon S3. It distributes the read across all compute slices in parallel, supports compressed and columnar formats, and applies options that validate and transform data during load. Alternatives such as row inserts, JDBC batching, or Spectrum materialization add layers that reduce throughput, so COPY remains the fastest supported method.

Exam trap

The trap here is assuming that an orchestration service such as AWS Glue must be faster because it is serverless, when the actual load mechanism, COPY, determines throughput.

661
MCQhard

A company uses Amazon Kinesis Data Streams to ingest IoT sensor data. The data is processed by an AWS Lambda function that transforms the records and writes to an Amazon S3 bucket. Recently, the Lambda function has been failing with 'Rate exceeded' errors for the S3 PUT API calls. The data volume is 10 MB/s with average record size 2 KB. What should be done to resolve this issue?

A.Add a random prefix to the S3 object key to distribute writes across multiple prefixes
B.Switch to Amazon Kinesis Data Firehose to write to S3
C.Increase the Lambda function's reserved concurrency
D.Increase the number of Kinesis shards
AnswerA

Random prefixes increase the number of S3 partitions, raising the PUT request limit.

Why this answer

The 'Rate exceeded' error for S3 PUT API calls indicates that the Lambda function is hitting S3 request rate limits. S3 buckets have a default limit of 3,500 PUT requests per second per prefix. With a data volume of 10 MB/s and an average record size of 2 KB, the Lambda function is generating approximately 5,000 PUT requests per second (10 MB/s ÷ 2 KB), which exceeds the per-prefix limit.

Adding a random prefix to the S3 object key distributes writes across multiple prefixes, effectively increasing the aggregate request rate limit to 3,500 PUT requests per second per prefix, thereby resolving the throttling issue.

Exam trap

The trap here is that candidates often confuse Kinesis shard scaling (Option D) or Lambda concurrency (Option C) with S3 rate limits, not realizing that the bottleneck is the S3 API request rate per prefix, not the data ingestion pipeline throughput.

How to eliminate wrong answers

Option B is wrong because Amazon Kinesis Data Firehose writes to S3 in batches and can also encounter S3 rate limits if the underlying prefix is not partitioned; it does not inherently solve the per-prefix request rate limit issue. Option C is wrong because increasing the Lambda function's reserved concurrency would increase the number of concurrent invocations, which would generate even more S3 PUT requests per second, exacerbating the rate limiting problem. Option D is wrong because increasing the number of Kinesis shards increases the data ingestion parallelism but does not affect the S3 PUT request rate limit; the Lambda function would still write to the same S3 prefix at the same rate.

662
MCQmedium

A company stores customer transaction data in an Amazon DynamoDB table. The table has a partition key of CustomerID and a sort key of TransactionDate. The data engineering team needs to retrieve all transactions for a specific customer within a date range, and the queries must be efficient. Which DynamoDB operation should the team use?

A.Scan the table with a FilterExpression on CustomerID and TransactionDate.
B.Use a BatchGetItem request with multiple CustomerID and TransactionDate pairs.
C.Query the table with a KeyConditionExpression on CustomerID and a condition on TransactionDate.
D.Use a GetItem request with the CustomerID and TransactionDate as the primary key.
AnswerC

A Query operation retrieves items based on the primary key. By specifying the partition key value and a condition on the sort key, DynamoDB can efficiently locate the relevant items without scanning the entire table. This is the most efficient way to retrieve a range of transactions for a specific customer.

Why this answer

The Query operation is designed to retrieve items sharing the same partition key value, with optional conditions on the sort key. By specifying CustomerID and a condition on TransactionDate, the team can efficiently fetch the desired transactions. Other operations either scan the entire table or retrieve single items, which are less efficient for this use case.

Exam trap

The trap here is confusing Query with Scan or GetItem; Query is the only operation that efficiently retrieves a range of items based on the sort key.

663
MCQeasy

A data engineer notices that an Amazon Kinesis Data Firehose delivery stream is failing to deliver data to an Amazon S3 bucket. The CloudWatch metrics show 'DeliveryToS3.Success' is 0 and 'S3.BucketExists' is 1. What is the MOST likely cause?

A.The S3 bucket has an ACL that denies access to Firehose.
B.The Firehose delivery stream Lambda transformation function is failing.
C.The IAM role for Firehose lacks s3:PutObject permission.
D.The S3 bucket does not exist.
AnswerC

S3.BucketExists being 1 confirms the bucket is reachable, so the failure lies in authorisation. Firehose requires s3:PutObject in its IAM role to write objects; without it, every delivery attempt is denied and DeliveryToS3.Success stays at 0.

Why this answer

The metric 'S3.BucketExists' is 1, confirming the S3 bucket exists, so the issue is not bucket existence. With 'DeliveryToS3.Success' at 0, the failure is in the write operation. The IAM role assumed by Firehose must have the s3:PutObject permission to deliver data; lacking it would cause all delivery attempts to fail silently, matching the observed metrics.

Exam trap

The trap here is that candidates may confuse 'S3.BucketExists' with successful delivery, or assume a missing bucket is the issue when the metric clearly shows the bucket exists, leading them to overlook the IAM permission gap.

How to eliminate wrong answers

Option A is wrong because S3 bucket ACLs are not evaluated when the IAM role grants the s3:PutObject permission via a bucket policy or identity-based policy; ACLs are legacy and Firehose uses IAM for authorization. Option B is wrong because a failing Lambda transformation function would cause 'DeliveryToS3.Success' to be 0 only if the transformation is mandatory, but the metric 'S3.BucketExists' would still be 1, and the failure would be logged as 'Lambda.ExecutionErrors' or similar, not directly as a delivery failure. Option D is wrong because 'S3.BucketExists' is 1, which explicitly indicates the bucket exists, so the bucket not existing cannot be the cause.

664
Multi-Selecthard

A company uses AWS Glue to run ETL jobs that transform data from Amazon S3 (Parquet) into a denormalized format for Amazon Redshift. The Glue job uses the DynamicFrame API. The job is failing with a 'MemoryError' when performing a join operation. The data is skewed on the join key. Which THREE actions can reduce memory usage and improve job stability? (Choose THREE.)

Select 3 answers
A.Use a broadcast join if one of the tables is small enough.
B.Use a salted join key to distribute skewed keys across partitions.
C.Increase the number of DPUs for the Glue job.
D.Repartition the data on the join key before the join operation.
E.Split the transformation into multiple Glue job steps to reduce per-step memory.
AnswersA, B, E

Avoids shuffling small table.

Why this answer

A broadcast join (using `join` with `broadcast` hint or `DynamicFrame.join(..., transformation_ctx='...')` with broadcast enabled) avoids shuffling the larger table across the cluster by copying the small table to every executor. This eliminates the memory pressure from skewed key distribution during the shuffle phase, reducing the risk of a MemoryError.

Exam trap

The trap here is that candidates often assume increasing resources (DPUs) or repartitioning will fix memory issues, but they fail to recognize that data skew on the join key is the root cause, which requires skew-aware techniques like salting or broadcast joins.

665
MCQhard

A data engineer is using Amazon Athena to query data stored in Amazon S3 in Parquet format. The engineer notices that a specific query is scanning much more data than expected, resulting in high costs and slow performance. The query filters on a column named 'event_date' which is a string in 'YYYY-MM-DD' format. The table is partitioned by 'year', 'month', and 'day' as separate string columns. The engineer wants to reduce the amount of data scanned. Which action should the engineer take?

A.Compress the Parquet files using Snappy compression.
B.Enable Amazon Athena query result reuse to cache the results.
C.Convert the 'event_date' column to a date type and use it in the WHERE clause.
D.Use the partition columns 'year', 'month', and 'day' in the WHERE clause instead of 'event_date'.
AnswerD

Athena uses partition pruning to limit the data scanned based on partition columns in the WHERE clause. Since the table is partitioned by year, month, and day, filtering on these columns allows Athena to scan only the relevant partitions, drastically reducing data scanned. This is the most effective way to optimize the query and reduce costs.

Why this answer

The table is partitioned by year, month, and day, so filtering on those partition columns in the WHERE clause enables Athena's partition pruning. This limits the data scanned to only the relevant partitions, significantly reducing cost and improving performance. Filtering on a non-partition column like event_date does not provide the same benefit.

Exam trap

The trap here is assuming that converting a column to date type or enabling caching will reduce data scanned, when the key optimization is to use the existing partition columns in the filter.

666
MCQeasy

A data engineer needs to transform JSON data from Amazon S3 into Parquet format using AWS Glue. The source files are in a bucket with thousands of small files. What is the best practice to optimize the Glue job performance?

A.Convert the JSON files to CSV before processing with Glue.
B.Enable 'Group small files' in the Glue job or use a DynamicFrame with coalesce.
C.Use an AWS Lambda function to pre-process the files.
D.Increase the number of DPUs to the maximum.
AnswerB

Coalescing merges thousands of small S3 objects into fewer, larger partitions before processing, cutting per-file overhead and task scheduling latency. This satisfies the small-files constraint by reducing the number of Spark tasks and improving Parquet write efficiency.

Why this answer

Enabling 'Group small files' in AWS Glue automatically coalesces thousands of small input files into larger partitions, reducing the number of tasks and minimizing overhead from task scheduling and S3 list operations. This is the recommended best practice for handling small files in Glue ETL jobs, as it optimizes read performance without requiring manual coalesce or repartitioning.

Exam trap

The trap here is that candidates assume more DPUs always improve performance, but for small files the bottleneck is metadata overhead, not compute capacity, so increasing DPUs without addressing file grouping leads to wasted resources and no speedup.

How to eliminate wrong answers

Option A is wrong because converting JSON to CSV adds an unnecessary preprocessing step and does not address the root cause of small file overhead; Glue can read JSON directly and convert to Parquet efficiently. Option C is wrong because using Lambda to pre-process files introduces additional cost, complexity, and potential timeout issues for large numbers of files, and does not leverage Glue's built-in optimization for small files. Option D is wrong because simply increasing DPUs does not solve the small file problem; it may even worsen performance by creating more task slots that compete for the same small files, leading to inefficient resource utilization.

667
MCQeasy

A data engineer needs to run a daily AWS Glue ETL job that transforms data in Amazon S3. The job must start at 2:00 AM UTC every day. The engineer wants to minimize operational overhead and ensure the job runs reliably. Which approach should the engineer use?

A.Configure an AWS Lambda function with a CloudWatch Events rule to start the Glue job daily at 2:00 AM UTC.
B.Use Amazon EventBridge to create a rule with a cron expression '0 2 * * ? *' that targets the AWS Glue job.
C.Use AWS Step Functions with a Wait state to delay execution until 2:00 AM UTC and then start the Glue job.
D.Create an AWS Glue trigger of type 'Schedule' with a cron expression '0 2 * * ? *' and attach it to the job.
AnswerD

AWS Glue triggers of type 'Schedule' allow defining a cron expression to run jobs on a schedule. The expression '0 2 * * ? *' runs daily at 2:00 AM UTC. This is a native Glue feature that requires no external services, minimizing operational overhead and ensuring reliable execution within the Glue environment.

Why this answer

AWS Glue provides built-in schedule triggers that use cron expressions to start jobs at specified times. Creating a trigger with the cron expression '0 2 * * ? *' and attaching it to the job is the simplest, most reliable method. It requires no additional services, reducing operational overhead and potential failure points, while ensuring the job runs daily at 2:00 AM UTC.

Exam trap

The trap here is overcomplicating the scheduling by involving external services like EventBridge or Lambda, when AWS Glue has a native trigger mechanism designed for this exact purpose.

668
MCQhard

A data engineer manages an Amazon S3 data lake with a bucket that has S3 Versioning enabled. A downstream analytics job accidentally overwrites thousands of current objects with corrupted data. The engineer must restore the previous good versions quickly and prevent the corrupted versions from being served. Which action should the engineer take?

A.Restore the previous good versions by copying them to overwrite the corrupted current versions
B.Delete the corrupted object versions so the previous versions become current
C.Remove the delete markers and re-upload the original files from a local backup
D.Enable S3 Object Lock in governance mode on the bucket to roll back the changes
AnswerA

With versioning enabled, each overwrite creates a new version while retaining the prior one. Copying the previous good version back onto the same key creates a new current version containing the good data, so downstream readers immediately see correct content without losing the audit trail. This can be scripted across thousands of objects using S3 Batch Operations with a version-aware copy, restoring the data lake quickly and safely.

Why this answer

S3 Versioning retains the prior object version on every overwrite, so the good data still exists as a noncurrent version. Copying that version back to the same key creates a new current version with the correct content, immediately restoring what consumers read while preserving history. Deleting versions, removing delete markers, or enabling Object Lock do not perform this rollback and may cause data loss or fail to address overwrites.

Exam trap

The trap here is confusing overwrite recovery with delete-marker handling or treating Object Lock as a rollback feature, when restoring a prior version requires copying the noncurrent version back to the key.

669
MCQeasy

A company needs to ingest data from an on-premises database to Amazon S3 with minimal impact on the source database. The data volume is several TB. Which AWS service is best suited for this task?

A.AWS Direct Connect
B.AWS Snowball Edge
C.AWS Database Migration Service (DMS)
D.Amazon S3 Transfer Acceleration
AnswerC

AWS DMS performs a one-time full load plus ongoing replication by reading from the source via a replication instance, using log-based change capture rather than querying application tables. This minimises load on the on-premises database, satisfying the minimal-impact constraint for multi-terabyte ingestion into Amazon S3.

Why this answer

AWS Database Migration Service (DMS) is best suited because it can continuously replicate data from an on-premises database to Amazon S3 with minimal impact on the source. DMS uses change data capture (CDC) to capture only incremental changes after an initial full load, avoiding heavy read loads on the source database. This makes it ideal for migrating several TB of data while keeping the source operational.

Exam trap

The trap here is that candidates confuse network acceleration services (Direct Connect, Transfer Acceleration) or offline transfer devices (Snowball) with database-specific migration tools, overlooking that DMS is the only option that directly reads from a database with minimal impact via CDC.

How to eliminate wrong answers

Option A is wrong because AWS Direct Connect provides a dedicated network connection for consistent bandwidth, but it does not perform data ingestion or migration itself; it is a transport layer, not a service that reads from a database. Option B is wrong because AWS Snowball Edge is a physical device for offline data transfer, which is suitable for very large datasets (petabytes) but introduces significant latency and is not designed for minimal impact on a live database during continuous ingestion. Option D is wrong because Amazon S3 Transfer Acceleration speeds up uploads to S3 over the internet using optimized network paths, but it does not interact with the source database or handle database-specific data extraction and transformation.

670
Multi-Selectmedium

A data engineer is designing an AWS Glue ETL job that reads from an Amazon S3 bucket containing nested JSON files. The job must flatten the nested structures and write the output to Amazon Redshift. The engineer needs to ensure the job can handle schema evolution and efficiently process only new data on subsequent runs. (Choose two.)

Select 2 answers
A.Use the AWS Glue Relationalize transform to flatten nested JSON structures.
B.Configure the job to use the 'ApplyMapping' transform to rename nested fields.
C.Enable bookmarking in the AWS Glue job to track previously processed data.
D.Use the 'ResolveChoice' transform to handle data type conflicts in nested fields.
E.Set the job's maximum concurrency to 1 to prevent schema conflicts.
AnswersA, C

The Relationalize transform in AWS Glue is specifically designed to flatten nested JSON and semi-structured data into relational tables. It produces multiple tables that can be joined, which is necessary when writing to Amazon Redshift. This transform handles arrays and nested objects, making it the appropriate choice for converting nested JSON into a format suitable for a relational data warehouse.

Why this answer

Enabling job bookmarks ensures that only new data is processed on subsequent runs, which is essential for incremental ETL. The Relationalize transform flattens nested JSON into relational tables, making the data compatible with Amazon Redshift. Together, these two features address the requirements for incremental processing and flattening nested structures.

Exam trap

The trap here is confusing transforms that adjust schema (ApplyMapping, ResolveChoice) with transforms that flatten nested data (Relationalize).

671
MCQhard

A healthcare company processes patient records in near-real-time using Amazon Kinesis Data Streams. Each record contains sensitive personal health information (PHI). The data must be encrypted at rest and in transit. The company also needs to audit access to the data. The data engineer is designing the ingestion pipeline. Which combination of services and configurations meets these requirements?

A.Use Kinesis Data Firehose to deliver data to S3 with SSE-S3, and enable CloudTrail for S3.
B.Use Kinesis Data Streams with TLS and enable CloudTrail for auditing. Do not enable SSE.
C.Use Kinesis Data Streams with SSE-KMS and TLS, and enable CloudTrail for data events.
D.Use Kinesis Data Streams with SSE-KMS and TLS. Do not enable any auditing.
AnswerC

SSE-KMS encrypts stream data at rest, TLS secures it in transit, and CloudTrail data events log every API-level access to the stream. Together these satisfy the PHI encryption and access-auditing requirements for the ingestion pipeline.

Why this answer

Kinesis Data Streams supports server-side encryption with AWS KMS (SSE-KMS) for encryption at rest, and TLS is used for encryption in transit by default. Enabling CloudTrail data events captures API-level access to Kinesis streams (e.g., PutRecord, GetRecords), satisfying the audit requirement. This combination meets all three requirements: encryption at rest, encryption in transit, and auditable access.

Exam trap

DEA-C01 often tests the distinction between encryption at rest (SSE-KMS on the stream) and encryption in transit (TLS), plus the difference between CloudTrail management events (automatic) and data events (opt-in) — candidates who assume CloudTrail logs all access by default will pick the wrong answer.

How to eliminate wrong answers

Option A is wrong because it uses Kinesis Data Firehose (a delivery service, not a stream-processing service) and SSE-S3, which encrypts only the S3 destination, not the data in the Firehose buffer or during ingestion; also, CloudTrail for S3 does not audit Kinesis access. Option B is wrong because it explicitly disables SSE, failing the encryption-at-rest requirement for PHI, which is a compliance violation (HIPAA). Option D is wrong because it omits auditing entirely, failing the requirement to audit access to the data, which is mandatory for PHI under HIPAA.

672
Multi-Selecthard

A data engineer is configuring an AWS Glue ETL job to read data from an Amazon S3 bucket that contains nested JSON files. The job must flatten the nested structure and write the output to Amazon Redshift. The engineer wants to optimize the job for performance and cost. Which two actions should the engineer take? (Choose two.)

Select 2 answers
A.Increase the number of DPUs to the maximum allowed to ensure faster execution.
B.Enable job bookmarks to track processed data and avoid reprocessing.
C.Convert the JSON files to Parquet using an AWS Glue crawler before running the ETL job.
D.Use the AWS Glue DynamicFrame Relationalize transform to flatten nested JSON.
E.Use the AWS Glue ResolveChoice transform to handle data type conflicts.
AnswersB, D

Job bookmarks in AWS Glue keep track of data that has already been processed in previous job runs. When reading from S3, bookmarks use the last modified time of objects to determine new data. Enabling bookmarks prevents reprocessing of unchanged data, reducing runtime and cost, especially for incremental loads. This is a best practice for ETL jobs that run repeatedly.

Why this answer

The Relationalize transform is purpose-built to flatten nested JSON, simplifying the ETL code and improving performance. Enabling job bookmarks ensures that only new or changed data is processed on subsequent runs, reducing cost and runtime. Together, these actions optimize the job for both performance and cost.

The other options either do not address flattening, are not cost-effective, or misunderstand service capabilities.

Exam trap

The trap here is assuming that a Glue crawler can convert file formats or that simply adding more DPUs is always the best way to optimize performance without considering cost.

673
MCQeasy

A data engineer needs to ingest on-premises CSV files into Amazon S3 every hour. The files are less than 1 GB each. Which service is the most cost-effective and requires the least operational overhead?

A.AWS DataSync
B.Amazon Kinesis Data Firehose
C.AWS Snowball Edge
D.AWS Database Migration Service (DMS)
AnswerA

AWS DataSync provides a managed agent that incrementally transfers on-premises files to Amazon S3 on a schedule, meeting the hourly, sub-1 GB, low-overhead constraint. It handles scheduling, integrity verification and retries natively, unlike scripting the AWS CLI or running bespoke cron jobs.

Why this answer

AWS DataSync is the most cost-effective and least overhead option for scheduled, recurring transfers of on-premises CSV files to S3. It provides a simple agent-based setup, supports hourly scheduling, and handles files under 1GB efficiently without complex configuration. In contrast, Amazon Kinesis Data Firehose is designed for streaming data ingestion, not batch file transfers; AWS Snowball Edge is intended for large-scale offline data migrations, not hourly incremental transfers; and AWS Database Migration Service (DMS) is specialized for migrating databases, not file transfers.

674
MCQmedium

A data engineer needs to store JSON documents that are accessed by a key-value pattern. The workload requires single-digit millisecond latency at any scale. Which AWS service is most appropriate?

A.Amazon DocumentDB (with MongoDB compatibility)
B.Amazon RDS for PostgreSQL
C.Amazon DynamoDB
D.Amazon Neptune
AnswerC

Amazon DynamoDB stores JSON as native items and retrieves them by primary key, delivering consistent single-digit millisecond latency regardless of table size. Its partition-based architecture scales horizontally without the query-engine overhead of Amazon Athena or the fixed schema of Amazon RDS, satisfying both the key-value access pattern and the any-scale latency constraint.

Why this answer

Amazon DynamoDB is a fully managed NoSQL key-value and document database that delivers single-digit millisecond latency at any scale. It is optimized for key-value access patterns, making it the ideal choice for storing and retrieving JSON documents by a primary key with consistent low latency.

Exam trap

The trap here is that candidates may confuse DocumentDB's document storage capability with DynamoDB's key-value performance, overlooking that DocumentDB is not designed for single-digit millisecond latency at any scale, especially under high throughput.

How to eliminate wrong answers

Option A is wrong because Amazon DocumentDB is a document database designed for MongoDB workloads, but it does not guarantee single-digit millisecond latency at any scale; its performance can vary with query complexity and indexing. Option B is wrong because Amazon RDS for PostgreSQL is a relational database that requires schema definition and is not optimized for key-value access patterns; it incurs higher latency due to SQL parsing and ACID overhead. Option D is wrong because Amazon Neptune is a graph database built for highly connected data and graph queries (e.g., using Gremlin or SPARQL), not for simple key-value lookups, and its latency profile is not designed for single-digit millisecond key-value access at scale.

675
MCQeasy

A team uses Amazon Kinesis Data Analytics to process streaming data. They notice that the application's output is delayed. Which AWS service can be used to monitor the application's performance and identify bottlenecks?

A.AWS CloudTrail
B.Amazon CloudWatch
C.Amazon Athena
D.AWS X-Ray
AnswerB

Amazon CloudWatch collects Kinesis Data Analytics metrics such as millisBehindLatest, input/output records and processing time, exposing exactly where the pipeline stalls. Alarms and dashboards surface the bottleneck causing the output delay, satisfying the stem's requirement to monitor application performance.

Why this answer

Amazon CloudWatch is the native monitoring service for AWS resources and applications, and Kinesis Data Analytics (now Amazon Managed Service for Apache Flink) publishes metrics such as millisBehindLatest, KPU utilization, and input/output throughput directly to CloudWatch. These metrics let the team pinpoint whether the delay stems from input backlog, insufficient KPUs, or downstream sink throttling. CloudWatch alarms and dashboards then provide the operational visibility needed to remediate the bottleneck.

Exam trap

The trap here is confusing observability services: candidates see 'monitor performance' and pick X-Ray or CloudTrail, but DEA-C01 expects you to know CloudWatch is the metrics/alarms backbone for Kinesis Data Analytics, while X-Ray is for request tracing and CloudTrail is for API auditing.

How to eliminate wrong answers

Option A is wrong because AWS CloudTrail records API activity and audit events (who called what, when), not runtime performance metrics, so it cannot reveal processing latency or throughput bottlenecks. Option C is wrong because Amazon Athena is a serverless query service for data in S3 — it analyzes historical data at rest, not live streaming application performance. Option D is wrong because AWS X-Ray traces requests through distributed applications (e.g., Lambda, API Gateway, ECS) and, while useful for latency analysis in some architectures, it is not the primary monitoring surface for Kinesis Data Analytics application metrics like millisBehindLatest or KPU usage.

Page 8

Page 9 of 18

Page 10