Courseiva

AWS Certified Data Engineer Associate DEA-C01 (DEA-C01) — Questions 451–525

1321 questions total · 18pages · All types, answers revealed

Page 6

Page 7 of 18

Page 8
451
MCQmedium

An organization wants to audit all API calls made to AWS services for compliance. Which AWS service should be used to capture and store these API calls?

A.AWS CloudTrail
B.AWS Config
C.Amazon VPC Flow Logs
D.Amazon CloudWatch Logs
AnswerA

AWS CloudTrail records every API call to AWS services as an event, capturing the caller identity, timestamp, source IP and request details, then delivering logs to Amazon S3 for retention. This directly satisfies the compliance audit requirement to capture and store all API activity across the account.

Why this answer

AWS CloudTrail records API activity across AWS services, capturing who made each call, from which IP, when, and with what parameters, and delivers the events to S3 and/or CloudWatch Logs for auditing. It is the canonical service for compliance auditing of API calls. AWS Config, VPC Flow Logs, and CloudWatch Logs serve different observability purposes.

Exam trap

The trap is confusing 'audit API calls' with 'track resource configuration changes' (AWS Config) or 'capture network traffic' (VPC Flow Logs); candidates who do not distinguish control-plane API auditing from configuration history or flow telemetry pick the wrong service.

How to eliminate wrong answers

Option B is wrong because AWS Config records resource configuration state and changes over time, not the API call history itself; it answers 'what does this resource look like now and how did it change,' not 'who called this API.' Option C is wrong because VPC Flow Logs capture IP traffic metadata (source/destination, ports, bytes, accept/reject) at the ENI, subnet, or VPC level, which is network telemetry, not AWS API auditing. Option D is wrong because CloudWatch Logs is a log storage and analysis service; it can receive CloudTrail events but is not itself the audit capture mechanism, and it does not natively record API calls without CloudTrail.

452
MCQhard

A data engineer is reviewing an IAM policy that controls access to an S3 bucket. The policy is attached to a user group. The policy includes a condition that explicitly requires server-side encryption with SSE-S3 for all GetObject requests. The engineer notices that users are unable to download objects from the bucket. What is the likely cause?

A.The policy is attached to a user group instead of an IAM role.
B.The policy does not specify the correct bucket ARN.
C.The policy does not allow the s3:GetObject action.
D.The objects are encrypted using SSE-KMS, not SSE-S3.
AnswerD

The condition requires SSE-S3 (AES256), so SSE-KMS objects are denied.

Why this answer

The IAM policy includes a condition that allows downloads only if the object is encrypted with SSE-S3. However, the objects in the bucket are encrypted using SSE-KMS, which does not satisfy the condition. As a result, the s3:GetObject request is denied.

This is a common scenario where a specific encryption condition in the policy blocks access when the actual encryption method differs. The other options are less likely: attaching the policy to a user group is valid and does not cause download issues; an incorrect bucket ARN would affect all operations, not just downloads; and missing s3:GetObject would be a straightforward policy error that would be easily identified.

Exam trap

The trap here is that candidates often focus only on S3 actions (like `s3:GetObject`) and overlook the required KMS permissions when SSE-KMS is involved, assuming SSE-S3 or no encryption is the default.

How to eliminate wrong answers

Option A is wrong because attaching a policy to a user group is a valid and common practice for granting permissions to multiple users; the issue is not about the attachment target but the permissions themselves. Option B is wrong because an incorrect bucket ARN would typically cause all actions to fail, not just downloads, and the question implies other operations might work. Option C is wrong because if the policy did not allow `s3:GetObject`, users would likely receive an Access Denied error for any read operation, but the question specifically mentions download failures, which can occur even with `s3:GetObject` allowed if KMS decrypt is missing.

453
MCQhard

A company has an Amazon Redshift cluster that stores petabytes of data. Queries are experiencing high disk usage due to large intermediate results. The data engineer needs to improve query performance without adding more nodes. Which action should the engineer take?

A.Set appropriate distribution keys to minimize data movement.
B.Configure workload management (WLM) queues to limit concurrency.
C.Apply column compression encoding to reduce data size.
D.Define sort keys on all columns used in WHERE clauses.
AnswerA

Distribution keys control which node stores each row, so choosing them well colocates joined rows and reduces data redistribution across nodes. That cuts the disk spill from large intermediate results without adding nodes, directly addressing the stated constraint.

Why this answer

In Amazon Redshift, distribution keys determine how table rows are spread across compute nodes. When joined tables share the same distribution key (or use ALL distribution for small dimension tables), matching rows are co-located on the same node slice, eliminating the broadcast or redistribution steps that spill large intermediate result sets to disk. Since the question specifies high disk usage from large intermediate results and prohibits adding nodes, fixing distribution keys directly reduces the data movement that generates those spills.

Exam trap

The trap here is conflating storage/scan optimizations (compression, sort keys) with join-time data movement; candidates pick compression or sort keys because they sound like general performance fixes, but only distribution keys address the intermediate-result disk spill caused by row redistribution.

How to eliminate wrong answers

Option B is wrong because WLM queues only manage concurrency and memory allocation for query slots; they throttle or queue queries but do not reduce the intermediate result size or disk spill caused by poor data co-location. Option C is wrong because column compression encoding reduces on-disk storage footprint and I/O for scans, but it does not address the redistribution/broadcast of rows during joins that produces large intermediate results. Option D is wrong because sort keys optimize range-restricted scans and merge joins by enabling zone-map block skipping, but defining sort keys on every WHERE column is impractical (Redshift allows only one sort key per table) and does not fix the data-movement problem driving disk usage.

454
Multi-Selectmedium

A data engineer is designing a data ingestion pipeline for IoT sensor data. The data is generated at a high velocity and must be processed in near real-time. The pipeline must also handle bursty traffic. Which TWO AWS services should be combined to achieve this? (Choose TWO.)

Select 2 answers
A.Amazon S3
B.Amazon Kinesis Data Analytics
C.Amazon Simple Queue Service (SQS)
D.AWS Glue
E.Amazon Kinesis Data Streams
AnswersB, E

Kinesis Data Analytics runs SQL or Apache Flink over streaming data, delivering the near real-time transformation the pipeline demands. Paired with a stream, it consumes bursty IoT input continuously, satisfying the high-velocity processing constraint without batch delays.

Why this answer

Amazon Kinesis Data Streams is designed for real-time, high-velocity data ingestion, providing durable, ordered data streams that can handle bursty traffic by scaling shard capacity. Amazon Kinesis Data Analytics can process these streams in near real-time using SQL or Apache Flink, enabling immediate transformations and analytics without needing to store data first.

Exam trap

The DEA-C01 exam often tests the distinction between streaming services (Kinesis Data Streams) and batch/queue services (SQS, S3), so the trap here is assuming SQS can handle real-time streaming or that S3 can serve as a primary ingestion point for high-velocity data.

455
MCQhard

A company runs a real-time analytics platform using Amazon Kinesis Data Streams with a shard count of 10. The data is consumed by an AWS Lambda function that writes to an Amazon DynamoDB table. The DynamoDB table has a partition key of 'user_id' and a sort key of 'timestamp'. The table is provisioned with 5000 RCUs and 5000 WCUs. Recently, the application experienced increased write latency and throttling errors (ProvisionedThroughputExceededException) on the DynamoDB table. The CloudWatch metrics show that ConsumedWriteCapacityUnits averages 4500 with occasional spikes to 6000. The Lambda function’s concurrency is set to 1000. The data engineer suspects the issue is due to hot partitions. Upon investigation, the engineer finds that a small number of users generate a disproportionately large amount of data. Which course of action would best resolve the throttling while minimizing cost?

A.Enable DynamoDB adaptive capacity and implement write sharding by adding a suffix to the partition key for high-volume users
B.Increase the provisioned WCUs to 10000 to handle spikes
C.Switch the DynamoDB table to on-demand capacity mode
D.Reduce the Lambda function concurrency to 100 to limit write requests
AnswerA

Write sharding splits a hot user's items across multiple synthetic partitions, letting DynamoDB spread the load rather than concentrating it on one partition key. Adaptive capacity alone only isolates frequently accessed items; it cannot fix a single partition key exceeding its per-partition write limit. Together they resolve the ProvisionedThroughputExceededException spikes without raising the 5000 WCUs, minimising cost.

Why this answer

The root cause is hot partitions caused by a small number of high-volume users. Enabling DynamoDB adaptive capacity allows the table to automatically adjust throughput to accommodate uneven access patterns, but the key fix is write sharding — adding a random or calculated suffix to the partition key for those high-volume users. This distributes writes across multiple physical partitions, eliminating the hot spot without requiring a global increase in provisioned capacity, thus resolving throttling while minimizing cost.

Exam trap

AWS often tests the misconception that throttling is always solved by increasing total provisioned capacity or switching to on-demand, when the real issue is partition-level hot spots that require key design changes like write sharding.

How to eliminate wrong answers

Option B is wrong because simply increasing provisioned WCUs to 10000 does not address the hot partition issue; the throttling is due to uneven distribution of writes across partitions, not a lack of total capacity, so this would waste money without fixing the root cause. Option C is wrong because switching to on-demand capacity mode would handle spikes but at a significantly higher cost for sustained high write volumes, and it still does not solve the hot partition problem — on-demand tables can still throttle individual partitions if a single partition exceeds 1,000 WCUs (the per-partition throughput limit). Option D is wrong because reducing Lambda concurrency to 100 would limit the total write throughput, potentially causing data backlogs in Kinesis, and it does not address the uneven distribution of writes across DynamoDB partitions; the hot partition would still be throttled even with fewer concurrent writers.

456
Multi-Selecthard

A data engineer is building an AWS Glue ETL job that must read from an Amazon S3 bucket in the same account and write to an Amazon Redshift cluster in a private VPC. The job must not traverse the public internet and must use least-privilege credentials. (Choose two.)

Select 2 answers
A.Store the Redshift user password in the job script as a plaintext variable so the connection can authenticate.
B.Set the Glue job's connection type to JDBC and rely on the default Glue service role for all S3 and Redshift permissions.
C.Create an internet gateway and a NAT gateway in the VPC so the Glue job can reach Redshift endpoints.
D.Grant the Glue job an IAM role with an S3 bucket policy that allows access only to the required prefixes.
E.Attach the Glue job to a VPC connection that includes a subnet with a route to the Redshift cluster, and configure the required Glue security group rules.
AnswersD, E

Least privilege is enforced through the IAM role attached to the job plus a bucket policy scoped to the needed prefixes. This limits what the job can read and write even if the code is changed later, and it is the credential control the scenario asks for rather than broad account-wide S3 permissions.

Why this answer

Private connectivity to Redshift requires the Glue job to run inside the VPC through a VPC connection with correct subnet routing and security group rules. Least privilege is then enforced by a scoped IAM role and bucket policy, so the two together satisfy both the no-public-internet and least-privilege requirements.

Exam trap

The trap here is assuming a JDBC connection by itself keeps traffic private, when the Glue job must actually be attached to a VPC connection with proper routing and security groups.

457
MCQmedium

A company is using AWS Database Migration Service (DMS) to migrate a 2 TB Oracle database to Amazon Aurora PostgreSQL. The migration must have minimal downtime. The source database is highly active with continuous writes. Which DMS migration type and additional configuration should the engineer use?

A.Use a CDC-only migration task to capture changes from the source.
B.Use a full load migration task and stop the source database before starting.
C.Use a full load migration task with task restart enabled.
D.Use a full load migration task followed by ongoing replication (CDC).
AnswerD

A full load plus change data capture (CDC) copies existing rows then continuously applies ongoing source changes, keeping Aurora PostgreSQL synchronised until cutover. This satisfies the minimal-downtime constraint for a highly active 2 TB Oracle source, unlike full load alone.

Why this answer

A full load migration followed by ongoing replication (CDC) allows the initial data copy to complete while continuously capturing and applying incremental changes from the highly active source. This minimizes downtime by keeping the target database nearly synchronized, requiring only a brief cutover window to stop writes and finalize replication.

Exam trap

The trap here is that candidates often assume a CDC-only task can handle both initial load and ongoing changes, but DMS requires a full load phase to populate the target before CDC can start, and they overlook that a full load alone cannot capture writes occurring during the migration.

How to eliminate wrong answers

Option A is wrong because a CDC-only task cannot migrate the existing 2 TB of data; it only captures ongoing changes, so the target would lack the initial dataset. Option B is wrong because stopping the source database before starting the migration would cause significant downtime, which contradicts the requirement for minimal downtime. Option C is wrong because a full load migration task with task restart enabled only retries the full load on failure; it does not capture ongoing writes after the initial load, so changes made during the migration would be lost, leading to data inconsistency.

458
MCQeasy

A company is migrating an on-premises MySQL database to Amazon RDS for MySQL. The database is 500 GB in size. The migration must have minimal downtime and must be completed within a week. Which AWS service should the data engineer use to perform the migration?

A.Amazon S3 Transfer Acceleration.
B.AWS Snowball Edge.
C.AWS DataSync.
D.AWS Database Migration Service (DMS).
AnswerD

AWS DMS performs continuous replication from the on-premises MySQL source to RDS for MySQL, keeping the target current until cutover. This satisfies the minimal-downtime constraint while completing the 500 GB migration well within a week.

Why this answer

AWS Database Migration Service (DMS) is the correct choice because it is specifically designed for migrating databases to AWS with minimal downtime. DMS can perform a full load of the 500 GB MySQL database and then continuously replicate ongoing changes from the on-premises source to the Amazon RDS target using change data capture (CDC), allowing the migration to complete within a week with near-zero downtime.

Exam trap

The trap here is that candidates may confuse AWS DataSync (a file-transfer service) with database migration, or assume Snowball Edge is required for any migration over a few hundred gigabytes, when in fact DMS with CDC is the appropriate service for minimizing downtime in database migrations even with moderate data sizes.

How to eliminate wrong answers

Option A is wrong because Amazon S3 Transfer Acceleration is a feature for speeding up uploads to S3 buckets over long distances using optimized network paths and edge locations; it is not a database migration service and cannot handle schema conversion or ongoing replication. Option B is wrong because AWS Snowball Edge is a physical data transport device intended for large-scale data transfers (typically tens of TB or more) when network bandwidth is limited; for a 500 GB database with a one-week timeline and minimal downtime requirement, the network transfer is feasible and Snowball Edge would introduce unnecessary latency and downtime for shipping. Option C is wrong because AWS DataSync is designed for moving large amounts of file data (e.g., NFS, SMB shares) to or from Amazon S3, Amazon EFS, or Amazon FSx; it does not support database engines like MySQL and cannot perform schema conversion or ongoing replication of transactional changes.

459
MCQmedium

A data engineer is building a pipeline that ingests records from an Amazon Kinesis data stream and writes them to Amazon S3 in Parquet format. The engineer wants to use AWS Glue to perform the transformation and needs the pipeline to handle records that arrive out of order and to deduplicate based on a record ID. Which combination of features should the engineer use?

A.Use Amazon Kinesis Data Firehose with a Lambda function for record transformation, and rely on Firehose's built-in deduplication to handle duplicates.
B.Use AWS Glue batch ETL with an S3 source, schedule the job every minute, and use the DropDuplicates transform on the record ID.
C.Use AWS Glue streaming ETL with a Kinesis source, set the job to process records in order, and use the ResolveChoice transform to merge duplicate records.
D.Use AWS Glue streaming ETL with a Kinesis source, apply a windowed deduplication using the record ID and event timestamp, and write the results to S3 in Parquet.
AnswerD

Glue streaming ETL can read from Kinesis and process micro-batches. Applying a windowed deduplication on the record ID and event timestamp handles out-of-order arrival and removes duplicates within the window. Writing the output as Parquet to S3 satisfies the format requirement. This combination directly addresses both the ordering and deduplication needs while leveraging Glue's streaming capabilities.

Why this answer

Glue streaming ETL reading from Kinesis, combined with windowed deduplication on record ID and event timestamp, handles both out-of-order arrival and duplicates. Writing Parquet to S3 meets the format goal. Batch ETL from S3 adds latency and misses cross-batch duplicates, Firehose lacks built-in deduplication and does not use Glue for transformation, and ResolveChoice is not a deduplication tool.

Exam trap

The trap here is assuming that Kinesis Data Firehose or Glue's ResolveChoice transform can deduplicate records by ID, when deduplication requires explicit windowed logic on a key and timestamp.

460
MCQmedium

A company runs a critical application on Amazon RDS for MySQL. To ensure high availability and automatic failover, the database is deployed as a Multi-AZ DB instance. The application uses read-heavy workloads. Which additional configuration should be used to offload read traffic without impacting write performance?

A.Use the Multi-AZ standby instance for read queries.
B.Create one or more Read Replicas in different AZs.
C.Use Amazon ElastiCache to cache read queries.
D.Change to a Single-AZ deployment with a larger instance size.
AnswerB

Read Replicas asynchronously replicate from the primary and serve read-only traffic, offloading SELECT queries while writes continue unaffected on the Multi-AZ primary. Placing them in different AZs also improves resilience, complementing the existing automatic failover configuration.

Why this answer

Amazon RDS Read Replicas are designed to offload read traffic from the primary DB instance without affecting write performance. Unlike the Multi-AZ standby instance, which is not accessible for reads, Read Replicas can be placed in different Availability Zones and serve read queries independently, improving read scalability while the primary handles writes.

Exam trap

The trap here is that candidates often assume the Multi-AZ standby instance can be used for read queries, but AWS explicitly prevents this to ensure failover consistency, making Read Replicas the correct choice for offloading reads.

How to eliminate wrong answers

Option A is wrong because the Multi-AZ standby instance is a passive replica that cannot serve read traffic; it only provides automatic failover and is not accessible for read queries. Option C is wrong because while Amazon ElastiCache can cache read queries, it does not offload read traffic from the database itself—it only reduces the number of queries hitting the database by caching results, but the question specifically asks for offloading read traffic without impacting write performance, and Read Replicas directly address this by providing additional database endpoints for reads. Option D is wrong because switching to a Single-AZ deployment with a larger instance size eliminates high availability and does not offload read traffic; it only increases capacity for both reads and writes, which can still impact write performance under heavy read load.

461
Multi-Selectmedium

A company is using AWS Glue ETL to transform and load data from Amazon S3 to Amazon Redshift. The data engineer notices that the job is taking longer than expected. Which TWO actions can improve the job performance?

Select 2 answers
A.Use Amazon Redshift Spectrum to query data directly.
B.Partition the source data in S3.
C.Increase the number of DPUs for the Glue job.
D.Enable S3 Transfer Acceleration.
E.Use a larger Redshift node type.
AnswersB, C

Partitioning the S3 source data lets AWS Glue read only the relevant partitions rather than scanning the entire dataset, cutting I/O and shuffle volume during the transform stage. This directly addresses the stem's slow job by reducing the bytes read before loading into Amazon Redshift.

Why this answer

Options B and C are correct because partitioning the source data in S3 reduces the amount of data scanned by Glue, improving I/O efficiency, and increasing the number of DPUs adds more parallelism for transformations. Option A is incorrect because Redshift Spectrum is for querying data in S3 directly from Redshift, not for Glue ETL jobs. Option D is incorrect because S3 Transfer Acceleration speeds up uploads to S3 but does not affect Glue job performance during ETL processing.

Option E is incorrect because larger Redshift node types do not impact Glue job execution; they only affect Redshift query performance.

462
MCQhard

A data engineer is using AWS Glue to transform data from Amazon S3 and load it into Amazon Redshift. The job runs daily and processes 500 GB of data. The engineer notices that the job takes several hours and wants to optimize performance. The data is stored in Parquet format and partitioned by date. Which optimization should the engineer implement to improve the job's performance?

A.Convert the Parquet data to CSV to improve read performance.
B.Use predicate pushdown to filter data at the source and partition pruning to read only necessary partitions.
C.Increase the number of DPUs for the Glue job to scale horizontally.
D.Enable job bookmarks to track processed data and avoid reprocessing.
AnswerB

Predicate pushdown and partition pruning allow Glue to read only the relevant partitions and rows from S3, reducing I/O and the amount of data processed. Since the data is partitioned by date, the job can target specific partitions. This significantly speeds up the job and lowers cost. It is a best practice for large datasets in Parquet.

Why this answer

Predicate pushdown and partition pruning are key optimizations for AWS Glue jobs reading partitioned Parquet data. They minimize the amount of data scanned by pushing filters to the source and skipping irrelevant partitions. This reduces I/O and compute time, directly addressing the performance issue without unnecessary cost increases.

Exam trap

The trap here is assuming that simply adding more DPUs will solve performance problems, when actually data layout optimizations like predicate pushdown and partition pruning often yield greater benefits for partitioned Parquet data.

463
Multi-Selectmedium

A company runs a data processing pipeline on Amazon EMR. The pipeline reads data from S3, processes it with Spark, and writes results back to S3. The engineer notices that the cluster is underutilized and wants to reduce costs. Which TWO actions should the engineer take? (Choose TWO.)

Select 2 answers
A.Use Spot instances for task nodes.
B.Configure the cluster to terminate after the job completes.
C.Change the master node to a larger instance type.
D.Enable EMRFS consistent view.
E.Increase the number of core nodes to improve parallelism.
AnswersA, B

Spot instances suit task nodes because EMR task nodes hold no HDFS data, so interruption only loses in-flight Spark tasks, which rerun. This directly addresses the underutilised cluster's cost problem by cutting compute spend on the elastic portion of the cluster.

Why this answer

Option A is correct because using Spot instances for task nodes is a standard EMR cost-optimization technique: task nodes perform only HDFS-non-persistent work (Spark executors), so they can be interrupted without data loss, and Spot capacity typically costs significantly less than On-Demand. Option B is correct because configuring the cluster to terminate after the job completes (a transient cluster, e.g., via --auto-terminate or a termination-protected=false setting) stops paying for idle EC2 instances once the batch pipeline finishes, directly addressing the underutilization. Option C is wrong because enlarging the master node does not improve processing throughput and increases cost, since the master only manages the cluster.

Option D is wrong because EMRFS consistent view is a data-consistency feature for S3 reads/writes, not a cost-reduction mechanism. Option E is wrong because adding core nodes increases cost and capacity rather than reducing spend, and the cluster is already underutilized.

Exam trap

The trap here is that candidates may confuse cost optimization features like Spot instances and auto-termination with performance improvements or data consistency settings, leading them to select options that increase resources or enable features unrelated to cost reduction.

464
MCQeasy

A data engineer uses AWS CloudTrail to investigate a security incident. The engineer runs the command shown in the exhibit. What does the output indicate?

A.A file was downloaded from the S3 bucket.
B.A file was deleted from the S3 bucket.
C.A batch of files was listed from the S3 bucket.
D.A file was uploaded to the S3 bucket.
AnswerD

The CloudTrail event records an S3 PutObject API call, indicating that a file was uploaded to the bucket. The event source and action name confirm the upload operation rather than a download, deletion or bucket-level configuration change.

Why this answer

The CloudTrail event shown is an S3 PutObject API call, which is the operation used to upload an object to an S3 bucket. The event name 'PutObject' and the presence of request parameters like 'bucketName' and 'key' confirm that a file was successfully written to the bucket. Therefore, the output indicates that a file was uploaded to the S3 bucket.

Exam trap

DEA-C01 often tests the ability to differentiate between S3 API operations based on CloudTrail event names, so candidates must memorize that PutObject corresponds to upload, GetObject to download, and DeleteObject to deletion.

How to eliminate wrong answers

Option A is wrong because downloading a file from S3 would generate a GetObject event, not PutObject. Option B is wrong because deleting a file from S3 would generate a DeleteObject event, not PutObject. Option C is wrong because listing files in an S3 bucket would generate a ListObjects or ListObjectsV2 event, not PutObject.

465
MCQeasy

A company needs to encrypt data in transit between an EC2 instance and an S3 bucket. Which method should be used?

A.Use HTTPS endpoints
B.Use plain HTTP
C.Use an IPsec VPN
D.Server-side encryption (SSE)
AnswerA

HTTPS endpoints encrypt traffic using TLS between the EC2 instance and S3, directly satisfying the in-transit encryption requirement. AWS supports HTTPS for all S3 API operations, so requests to the REST endpoint are secured without extra configuration. This protects data crossing the network, unlike at-rest options such as SSE-KMS.

Why this answer

HTTPS endpoints encrypt data in transit between EC2 and S3 using TLS/SSL, ensuring confidentiality and integrity over the public internet. S3 supports HTTPS natively on its REST endpoints, and the AWS SDKs default to HTTPS, making this the simplest and most secure method for encrypting data in motion.

Exam trap

The trap here is confusing encryption in transit (HTTPS) with encryption at rest (SSE), leading candidates to select server-side encryption even though it does not protect data during network transfer.

How to eliminate wrong answers

Option B is wrong because plain HTTP transmits data in cleartext, exposing it to interception and tampering, which violates encryption-in-transit requirements. Option C is wrong because an IPsec VPN encrypts traffic between networks but is unnecessary and overly complex for direct EC2-to-S3 communication, which can be secured via HTTPS without additional infrastructure. Option D is wrong because server-side encryption (SSE) protects data at rest within S3, not data in transit between EC2 and S3.

466
MCQeasy

A company needs to ingest data from multiple SaaS applications into Amazon S3. The data sources provide REST APIs. Which AWS service can be used to build a fully managed data ingestion pipeline without writing custom code?

A.Amazon AppFlow
B.Amazon Kinesis Data Streams
C.AWS Lambda with custom code
D.AWS Glue with Python shell
AnswerA

Amazon AppFlow provides managed connectors for SaaS applications and can transfer data into Amazon S3 on a schedule or event trigger, requiring no custom code. Glue or Lambda would demand development effort the stem explicitly excludes.

Why this answer

Amazon AppFlow is a fully managed service designed to transfer data from SaaS applications to AWS services like Amazon S3 without writing custom code. Option B is wrong because Amazon Kinesis Data Streams is primarily for real-time streaming data, not for directly ingesting from SaaS APIs. Option C is wrong because AWS Lambda requires custom code to integrate with SaaS APIs.

Option D is wrong because AWS Glue with Python shell is for ETL transformations and also requires custom scripting.

467
Multi-Selecthard

A data engineer manages an Amazon Redshift provisioned cluster that serves a nightly ELT workload. The cluster's largest fact table is loaded with new rows each night, and queries frequently filter on a date column and join to a customer dimension. The engineer wants to improve query performance and reduce the time spent vacuuming. Which TWO actions should the engineer take? (Choose two.)

Select 2 answers
A.Enable automatic compression on all columns and disable the sort key to reduce storage footprint.
B.Convert the fact table to a view over the raw staging table to avoid storing duplicate data.
C.Set the customer join column as the distribution key so matching rows of the fact and dimension tables co-locate on the same slice.
D.Increase the cluster's number of nodes without changing the table design or sort keys.
E.Define the date column as the sort key so range-filtered queries skip irrelevant blocks during scans.
AnswersC, E

Distributing both the fact table and the customer dimension on the join column places matching rows on the same slice, enabling collocated joins that avoid broadcasting or shuffling data across the network. This lowers join cost and improves overall query performance for the frequent fact-to-dimension joins in this workload.

Why this answer

Performance and vacuum efficiency in Redshift are driven primarily by table design. Sorting on the frequently filtered date column enables zone-map pruning and makes VACUUM's work easier, while distributing the fact table and customer dimension on the join column enables collocated joins that avoid network-heavy shuffles. Together they reduce scanned blocks and join cost without adding hardware.

Exam trap

The trap here is believing that adding cluster capacity or compression alone fixes slow filtered joins, when the decisive levers are the sort key and distribution key choices.

468
MCQhard

Refer to the exhibit. A data engineer ran the CLI command to check the configuration of an RDS instance named 'mydb'. Which statement accurately describes the current configuration?

A.The database is in the 'stopped' state
B.The database is a Single-AZ deployment and is not a read replica
C.The database is a read replica of another instance
D.The database is a Multi-AZ deployment
AnswerB

The CLI output shows no MultiAZ or ReadReplica fields, confirming a standalone Single-AZ instance. This satisfies the stem's requirement to describe the current configuration accurately: the database runs in one Availability Zone with no read replica, so it offers no automatic failover or read scaling.

Why this answer

The CLI command output shows 'DBInstanceStatus: available' and 'MultiAZ: False', indicating the database is running as a Single-AZ deployment. Additionally, the absence of a 'SourceDBInstanceIdentifier' field confirms it is not a read replica. The 'ReadReplicaSourceDBInstanceIdentifier' is not present, which would be required if it were a read replica.

Exam trap

AWS often tests the misconception that a database with 'available' status must be a read replica or Multi-AZ, but the absence of the 'ReadReplicaSourceDBInstanceIdentifier' and 'MultiAZ: False' clearly identify it as a standalone Single-AZ instance.

How to eliminate wrong answers

Option A is wrong because the CLI output shows 'DBInstanceStatus: available', not 'stopped', so the database is running, not stopped. Option C is wrong because the output lacks a 'ReadReplicaSourceDBInstanceIdentifier' field, which is mandatory for a read replica; without it, the instance is a primary database, not a replica. Option D is wrong because the output explicitly shows 'MultiAZ: False', which means it is a Single-AZ deployment, not Multi-AZ.

469
MCQeasy

A data engineer needs to run an AWS Glue ETL job that reads from an Amazon S3 bucket in another AWS account. The bucket owner has granted cross-account access, and the Glue job runs with an IAM role in the engineer's account. The job fails with an access denied error when reading the source objects. Which change is required to allow the Glue job to read the cross-account S3 data?

A.Add an S3 bucket policy in the source account that grants the Glue job's IAM role s3:GetObject and s3:ListBucket on the bucket.
B.Recreate the Glue job in the source account so it uses a role owned by the bucket owner.
C.Attach the AWSGlueServiceRole managed policy to the Glue job's IAM role in the engineer's account.
D.Enable AWS Glue Data Catalog cross-account sharing by creating a resource link to the source account's catalog.
AnswerA

Cross-account S3 access requires both the caller's IAM role to allow the S3 actions and the bucket owner to grant those actions through a bucket policy. Since the role already runs in the engineer's account, the missing piece is the resource-based policy in the source account. Adding s3:GetObject and s3:ListBucket for that role resolves the denial.

Why this answer

S3 authorization evaluates both identity-based policies in the caller's account and resource-based policies on the bucket. Because the Glue job's role lives in a different account than the bucket, the bucket owner must attach a bucket policy granting s3:GetObject and s3:ListBucket to that role. Catalog sharing and service-role policies do not extend to cross-account object reads.

Exam trap

The trap here is assuming a Glue catalog resource link also grants data access, when it only shares metadata and never authorizes S3 object reads.

470
Multi-Selecthard

A data engineer is using AWS Step Functions to orchestrate an ETL workflow that includes an AWS Glue job, an Amazon EMR step, and an Amazon Redshift stored procedure. The workflow sometimes fails due to transient errors, and the engineer wants to implement a retry strategy that avoids duplicate data processing. Which TWO actions should the engineer take? (Choose two.)

Select 2 answers
A.Implement idempotency in the ETL tasks so that retries do not produce duplicate data.
B.Configure the Step Functions execution to run in Express mode for faster retries.
C.Use a Catch block to redirect to a cleanup state that deletes partially written data before retrying the entire workflow.
D.Add a Retry policy on the individual Task states with a backoff rate and max attempts to handle transient failures.
E.Increase the timeout of the entire state machine to allow more time for tasks to complete.
AnswersA, D

Idempotency ensures that re-executing a task produces the same result without side effects, such as duplicate records. By designing Glue jobs, EMR steps, and Redshift procedures to be idempotent, retries from Step Functions will not cause data duplication. This is essential when combined with retry policies to safely handle transient failures without compromising data integrity.

Why this answer

To handle transient errors without duplicating data, the engineer should add Retry policies to Task states and ensure ETL tasks are idempotent. Retry policies automatically re-execute failed tasks, while idempotency guarantees that repeated executions do not create duplicate records. Together, they provide robust error handling and data integrity.

Exam trap

The trap here is thinking that simply increasing timeouts or switching to Express workflows will solve transient errors, when the real solution requires retries and idempotency.

471
MCQhard

Refer to the exhibit. A data engineer runs the AWS CLI command to describe a Glue job. The job is expected to process new data incrementally using job bookmarks. However, the job reprocesses all data every time it runs. What is the MOST likely reason?

A.The job bookmark option is set to 'job-bookmark-enable' but should be 'job-bookmark-disable'.
B.The job's MaxRetries is set to 0, which disables bookmarks.
C.The ETL script does not use the 'transformation_ctx' parameter in its DynamicFrame transformations.
D.The Glue job's command name is 'glueetl', which does not support job bookmarks.
AnswerC

Job bookmarks rely on the `transformation_ctx` argument to persist state per transformation; without it, Glue cannot track which data each DynamicFrame has already processed, so every run reads the full dataset. Supplying a unique `transformation_ctx` for each source and transformation satisfies the incremental-processing requirement.

Why this answer

AWS Glue job bookmarks rely on the `transformation_ctx` parameter to track state. Without it, Glue cannot identify which data has already been processed, causing the job to reprocess all data on every run. The `transformation_ctx` must be passed to each DynamicFrame transformation (e.g., `apply_mapping`, `filter`, `join`) to enable bookmark-based incremental processing.

Exam trap

The trap here is that candidates often assume bookmarks are controlled only by the job configuration setting (`job-bookmark-enable`) and overlook the critical role of `transformation_ctx` in the ETL script, which is a common oversight in AWS Glue exam questions.

How to eliminate wrong answers

Option A is wrong because `job-bookmark-enable` is the correct setting to enable bookmarks; setting it to `job-bookmark-disable` would disable them, not fix the reprocessing issue. Option B is wrong because `MaxRetries` controls the number of retry attempts on failure and has no effect on job bookmark behavior. Option D is wrong because `glueetl` is the standard command name for ETL jobs and fully supports job bookmarks; the command name does not disable bookmarks.

472
MCQhard

Refer to the exhibit. An IAM policy is attached to an IAM role used by an application. The application needs to read objects from 'my-bucket' that have the tag 'classification=public'. The application account is 123456789012. However, the application is getting 'Access Denied' errors. What is the most likely reason?

A.The Deny statement uses StringNotEquals, which incorrectly denies the application account.
B.The policy does not grant s3:ListBucket permission, so the application cannot list objects.
C.The object being accessed does not have the tag 'classification=public'.
D.The Deny statement blocks all access from accounts other than 123456789012, but the application is in that account.
AnswerC

Tag-based access control evaluates the tag attached to the object itself, not the bucket. If the requested object lacks `classification=public`, the condition in the IAM policy fails and S3 returns Access Denied, even though the role's permissions and bucket policy are otherwise valid.

Why this answer

The most likely reason for Access Denied is that the object being accessed does not have the tag 'classification=public'. The policy's Allow statement is conditioned on s3:ExistingObjectTag/classification equaling 'public', so if the tag is missing or has a different value, the Allow does not match and the request is denied. The Deny statement in the policy is a separate guardrail and is not the cause here.

Exam trap

DEA-C01 often tests whether candidates read the policy conditions carefully — the trap is blaming the Deny statement or missing permissions when the real issue is that the object lacks the required tag, making the Allow condition false.

How to eliminate wrong answers

Option A is wrong because StringNotEquals in a Deny statement is a common pattern to block access from accounts other than the trusted one — it does not incorrectly deny the application account if the condition is structured correctly (e.g., denying when aws:PrincipalAccount != 123456789012). Option B is wrong because s3:ListBucket is only needed for listing operations; GetObject on a specific key does not require ListBucket, so its absence would not cause Access Denied for a read. Option D is wrong because the Deny statement is designed to block other accounts, not the application account — if the application is in 123456789012, the Deny does not apply to it.

473
MCQmedium

A data engineer manages an AWS Glue ETL job that processes JSON files from an S3 bucket and writes Parquet to another bucket. The job uses a Glue DynamicFrame with a specified schema. During execution, the job fails with the error: 'AnalysisException: cannot resolve column 'transaction_id' given input columns: [txn_id, amount, timestamp]'. The source data has a column named 'txn_id', but the Glue job's script references 'transaction_id'. The job's catalog table for the source points to the correct S3 location and has the correct schema. What is the most likely cause of this error?

A.The Glue job's IAM role lacks permission to read the S3 bucket, resulting in an empty DynamicFrame.
B.The Glue Data Catalog table for the source has an incorrect schema definition.
C.The Glue job's script uses a hardcoded schema that does not match the actual data columns.
D.The S3 bucket contains files with inconsistent schemas, causing Glue to infer a different column name.
AnswerC

The error indicates that the column 'transaction_id' is not present in the input data, while 'txn_id' exists. If the script applies a hardcoded schema or a mapping that expects 'transaction_id', the Spark engine cannot resolve it. This is a common issue when the script is written with a static schema that diverges from the actual source structure, even if the catalog is correct.

Why this answer

The error occurs because the Glue script references a column that does not exist in the source data. Even though the Glue Data Catalog is correct, the script may have been written with a hardcoded schema or mapping that expects 'transaction_id' instead of 'txn_id'. To resolve, the engineer should update the script to use the correct column name or apply a mapping to rename 'txn_id' to 'transaction_id'.

Exam trap

The trap here is assuming that the Glue Data Catalog schema is always used by the job, but the script may define its own schema or mappings that override the catalog.

474
MCQeasy

A data engineer needs to store a large volume of time-series data from IoT sensors. The data will be queried by timestamp and sensor ID, and the engineer wants to use a managed AWS database that can handle high write throughput and provide fast queries on recent data. The engineer also wants to automatically expire old data after 90 days to reduce storage costs. Which AWS service should the engineer use?

A.Amazon DynamoDB with TTL
B.Amazon Timestream
C.Amazon RDS for PostgreSQL with table partitioning
D.Amazon S3 with lifecycle policies
AnswerB

Amazon Timestream is a purpose-built time-series database that automatically scales to handle high write throughput and provides fast queries on recent data. It has built-in time-series functions and supports automatic data retention policies to expire data after a specified period, such as 90 days. This directly matches the requirements for IoT sensor data with timestamp and sensor ID queries and automatic data expiration.

Why this answer

Amazon Timestream is a fully managed time-series database that handles high write throughput, provides fast queries on recent data, and supports automatic data retention policies. It is designed specifically for time-series data like IoT sensor readings, with built-in functions for time-based queries. Other options either lack native time-series optimization or require manual management of retention and querying.

Exam trap

The trap here is assuming that DynamoDB with TTL is sufficient for time-series data, but Timestream offers purpose-built time-series capabilities and automatic retention that simplify management.

475
MCQeasy

A company needs to ingest data from an on-premises Oracle database into Amazon S3 on a daily basis. The data volume is about 100 GB per day. Which AWS service is BEST suited for this task?

A.Use AWS DataSync to copy the database files to S3.
B.Use Amazon Kinesis Data Firehose with a database connector.
C.Use AWS Database Migration Service (DMS) to replicate data to S3.
D.Use AWS Glue to extract data from Oracle and write to S3.
AnswerC

AWS DMS performs continuous change data capture from Oracle and writes directly to Amazon S3, handling the 100 GB daily volume without custom extraction code. Its native S3 target endpoint satisfies the daily ingestion requirement, unlike batch-only tools that lack Oracle CDC support.

Why this answer

AWS Database Migration Service (DMS) can continuously replicate data from Oracle to S3, and it supports full load and change data capture (CDC). Option A (AWS DataSync) is for file-based transfers, not database replication. Option B (Amazon Kinesis Data Firehose) is for streaming data, not database pull.

Option D (AWS Glue) is for ETL but does not natively support continuous CDC from Oracle.

476
MCQmedium

A data engineer is configuring an AWS Glue crawler to catalog data stored in Amazon S3. The data is organized as Parquet files under prefixes named by year, month, and day, such as s3://analytics/events/year=2024/month=05/day=17/. Queries in Amazon Athena must use partition pruning to limit scanned data. Which crawler configuration should the engineer choose?

A.Create a crawler for each individual day prefix, such as s3://analytics/events/year=2024/month=05/day=17/, so each day becomes its own table.
B.Create a crawler with the S3 path pointing to s3://analytics/events/ and enable 'Update all new and existing partitions with metadata from the table' so partition metadata stays current.
C.Create a crawler with the S3 path pointing to s3://analytics/events/ and set the 'Schema change policy' to 'Delete tables and columns' so stale partitions are removed.
D.Create a crawler with the S3 path pointing to s3://analytics/events/ and enable the 'Add new columns only' option so new partitions are detected automatically.
AnswerB

The Hive-style year=/month=/day= prefixes are automatically recognized as partitions when the crawler targets the parent prefix. The 'Update all new and existing partitions with metadata from the table' setting ensures that as new date prefixes appear, the crawler adds them and refreshes partition metadata. Athena can then use the partition keys for pruning, scanning only the relevant date ranges instead of the whole bucket.

Why this answer

Hive-style prefixes such as year=/month=/day= are recognized as partition keys when an AWS Glue crawler targets the parent S3 prefix. To keep the catalog current as new date prefixes appear, the crawler should be configured to update all new and existing partitions with metadata from the table. Athena then reads the partition keys from the Data Catalog and prunes to only the relevant prefixes, reducing scanned bytes and query cost.

Exam trap

The trap here is confusing schema-update options such as 'Add new columns only' with partition-update behavior, when partition registration is controlled by the crawler's partition update setting.

477
Multi-Selectmedium

A data engineering team is designing a data lake on Amazon S3 for storing sensor data from IoT devices. The data is written in near real-time and needs to be queried using Amazon Athena. Which TWO configurations should the team implement to optimize query performance and minimize costs?

Select 2 answers
A.Compress the data using GZIP.
B.Use S3 Standard-IA storage class.
C.Store the data in Apache Parquet format.
D.Partition the data by date and sensor ID.
E.Enable Requester Pays on the S3 bucket.
AnswersC, D

Parquet is columnar, so Athena reads only the columns referenced in each query and scans far less data than row-based CSV or JSON. This directly satisfies the stem's requirement to optimise query performance and minimise cost, since Athena charges per terabyte scanned.

Why this answer

Option C is correct because storing the data in Apache Parquet, a columnar format, allows Athena to read only the columns referenced in a query and leverages columnar compression and predicate pushdown, which dramatically reduces the amount of data scanned and therefore query cost and latency. Option D is correct because partitioning the data by date and sensor ID lets Athena prune irrelevant partitions using the partition metadata in the AWS Glue Data Catalog, so queries only scan the specific date/sensor combinations needed instead of the entire dataset. Together, Parquet plus partitioning is the standard best practice for cost-efficient, high-performance Athena queries over S3 data lakes.

Option A is not correct because GZIP is a row-based compression format that cannot be split for parallel reads and does not provide column pruning, so it is inferior to Parquet's built-in columnar compression for this use case. Option B is not correct because S3 Standard-IA is a storage-class cost optimization for infrequently accessed data, not a query-performance optimization, and it can even add retrieval costs for frequently queried near real-time data. Option E is not correct because Requester Pays shifts data-transfer costs to the requester and does nothing to improve Athena query performance or reduce the team's own query scanning costs.

Exam trap

AWS often tests the misconception that any compression (like GZIP alone) is sufficient for Athena optimization, but the trap is that without a columnar format like Parquet or ORC, compression alone does not enable column pruning or predicate pushdown, leading to higher scan costs and slower queries.

478
MCQmedium

A data engineer is using AWS Glue to run a nightly ETL job that reads from an Amazon DynamoDB table and writes to Amazon S3 in Parquet format. The DynamoDB table is large and has a high read capacity. The engineer wants to minimize the impact on the DynamoDB table's performance and reduce the ETL job's runtime. Which approach should the engineer take?

A.Use AWS Glue with DynamoDB connector and enable DynamoDB Accelerator (DAX) for the table.
B.Use AWS Glue with DynamoDB connector and increase the job's number of workers to read the table faster.
C.Use AWS Glue with DynamoDB connector and configure the job to use parallel scans with a segment ratio.
D.Export the DynamoDB table to Amazon S3 using DynamoDB's native export feature, then use AWS Glue to read from S3 and write to Parquet.
AnswerD

DynamoDB's native export to S3 does not consume read capacity units and does not affect table performance. The export creates data in DynamoDB JSON format in S3. AWS Glue can then read from S3, transform as needed, and write to Parquet. This approach minimizes impact on the table and can be faster for large tables since it uses a separate export process. It also reduces the ETL job's runtime because Glue reads from S3 instead of DynamoDB.

Why this answer

To minimize impact on DynamoDB performance and reduce ETL runtime, exporting the table to S3 is ideal because it does not consume read capacity. AWS Glue can then process the exported data from S3, which is faster and does not affect the source table. Other methods like parallel scans or increasing workers still consume capacity and can impact performance.

Exam trap

The trap here is assuming that using DAX or parallel scans will avoid impacting DynamoDB when they still consume read capacity or add complexity.

479
Drag & Dropmedium

Order the steps to migrate an on-premises database to Amazon RDS using AWS DMS.

Drag or tap steps into the slots.

Steps
Order
1Step 1
2Step 2
3Step 3
4Step 4

Why this order

First, create the replication instance. Then configure endpoints, create the migration task, start it, and finally validate the migrated data.

480
MCQmedium

A data engineer is building an AWS Glue ETL job that reads a large JDBC table from Amazon RDS for PostgreSQL. The job must read the table in parallel to reduce runtime, but the table has no numeric primary key or monotonically increasing column. Which AWS Glue connection property should the engineer configure to enable parallel reads?

A.partitionColumn with lowerBound and upperBound
B.hashfield
C.hashpartitions and hashfield
D.enableParallelRead with numPartitions
AnswerC

AWS Glue supports hashpartitions and hashfield connection options specifically for JDBC sources that lack a suitable numeric partitioning column. hashpartitions defines the number of parallel read partitions, and hashfield specifies a column (often a string or UUID) used to hash rows into those partitions. This enables parallel reads without requiring a numeric key, directly addressing the scenario's constraint.

Why this answer

For JDBC sources without a numeric or monotonic column, AWS Glue provides hashpartitions and hashfield to enable parallel reads. hashpartitions sets the number of partitions, and hashfield names a column used for hashing rows across those partitions. This avoids the need for a numeric partitionColumn and is the documented approach for this exact situation.

Exam trap

The trap here is assuming that partitionColumn with lowerBound and upperBound always works for parallel JDBC reads, when it actually requires a numeric, evenly distributed column.

481
MCQmedium

A data engineer is using Amazon Redshift and needs to load data from Amazon S3 into a Redshift table. The data is in Parquet format and is partitioned by date in S3. The engineer wants to load only the data for the last 7 days to reduce load time and cost. Which Redshift command should the engineer use?

A.COPY command with a manifest file that lists only the files for the last 7 days.
B.UNLOAD command with the PARTITION BY clause.
C.INSERT INTO ... SELECT FROM spectrum ... WHERE date >= ...
D.COPY command with the PARTITION clause.
AnswerA

The COPY command can use a manifest file to specify exactly which files to load. By creating a manifest that lists only the Parquet files for the last 7 days, the engineer can load only the required data. This reduces load time and cost by avoiding unnecessary data transfer.

Why this answer

The COPY command with a manifest file allows precise selection of files to load, enabling the engineer to load only the Parquet files for the last 7 days. This minimizes data transfer and load time. Other options either do not exist (PARTITION clause in COPY) or are for unloading data (UNLOAD) or are less efficient for bulk loading (INSERT from Spectrum).

Exam trap

The trap here is assuming that the COPY command can filter by partition automatically, but it requires a manifest or specific prefix to limit the files loaded.

482
MCQmedium

A company stores sensitive customer data in an S3 bucket. The security team requires that all data be encrypted at rest using a customer-managed AWS KMS key. However, when a data engineer attempts to upload an object using the AWS CLI, the upload fails with an access denied error. The engineer has s3:PutObject permission on the bucket. Which additional permission is most likely missing?

A.kms:CreateKey
B.kms:Decrypt
C.s3:PutObjectAcl
D.kms:GenerateDataKey
AnswerD

Uploading with a customer-managed KMS key requires kms:GenerateDataKey to obtain a data key for envelope encryption. s3:PutObject alone is insufficient; without that KMS permission, the request fails with AccessDenied even though the S3 action is allowed.

Why this answer

When S3 encrypts an object with SSE-KMS using a customer-managed key, the caller must have kms:GenerateDataKey permission on that KMS key so S3 can obtain a data key to encrypt the object. The s3:PutObject permission alone authorizes the S3 API call but does not grant the KMS operation needed to produce the encryption key, so the upload fails with AccessDenied. Granting kms:GenerateDataKey on the CMK resolves the failure.

Exam trap

DEA-C01 often tests the confusion between encrypt-time and decrypt-time KMS permissions, so the trap is selecting kms:Decrypt for an upload failure when the actual missing permission is kms:GenerateDataKey.

How to eliminate wrong answers

Option A is wrong because kms:CreateKey is only needed to create a new KMS key, not to use an existing customer-managed key for encryption. Option B is wrong because kms:Decrypt is required for reading/decrypting objects, not for uploading/encrypting them — it would be the missing permission for a download failure, not an upload failure. Option C is wrong because s3:PutObjectAcl only controls the ability to set an object's ACL, which is unrelated to KMS encryption and would not cause an encryption-related AccessDenied.

483
MCQeasy

An e-commerce application uses Amazon ElastiCache for Redis to cache product catalog data. The cache currently uses lazy loading. The team wants to ensure that frequently accessed product data is always fresh. Which caching strategy should they implement?

A.Write-through caching
B.Set a TTL of 5 minutes for all cached items
C.Use database read replicas to serve data
D.Lazy loading with TTL
AnswerA

Write-through caching updates the Redis cache synchronously on every database write, so cached product data never goes stale. This directly satisfies the freshness constraint that lazy loading cannot guarantee, since lazy loading only populates entries on a miss and leaves existing values stale until eviction or expiry.

Why this answer

Write-through caching ensures that data is written to the cache simultaneously with the database, guaranteeing that frequently accessed product data is always fresh. This strategy eliminates stale reads by synchronously updating the cache on every write, which directly addresses the requirement for freshness without relying on expiration or lazy population.

Exam trap

The trap here is that candidates often assume lazy loading with a short TTL is sufficient for freshness, but the exam tests the understanding that only write-through (or write-behind) strategies guarantee synchronous cache updates without relying on expiration windows.

How to eliminate wrong answers

Option B is wrong because setting a TTL of 5 minutes does not guarantee freshness; data can still become stale within the TTL window, and frequently accessed items may be served from the cache even after they have been updated in the database. Option C is wrong because database read replicas serve stale data asynchronously and do not cache product data in ElastiCache, failing to meet the caching freshness requirement. Option D is wrong because lazy loading with TTL still allows stale data to be served until the TTL expires or a cache miss triggers a refresh, which does not ensure that frequently accessed data is always fresh.

484
MCQhard

A company uses AWS Glue to process JSON logs from S3. The logs have a nested structure and the schema evolves over time. The data engineer needs to ensure the Glue job can handle schema changes without failing. Which configuration should be used?

A.Manually update the table schema in the Glue Data Catalog before each run
B.Use Spark SQL with a static schema definition in the script
C.Set the job parameter '--enable-glue-datacatalog' and '--mergeDynamicColumns' to true
D.Enable AWS Glue Schema Registry and define a schema version
AnswerC

This allows Glue DynamicFrame to merge schema variations automatically.

Why this answer

Setting '--enable-glue-datacatalog' allows the Glue job to use the Data Catalog as the metastore, and '--mergeDynamicColumns' (or the equivalent '--enable-schema-evolution' in newer Glue versions) instructs the job to dynamically merge new columns from the evolving JSON schema into the existing table schema during runtime, preventing job failures due to schema mismatches. This is specifically designed for nested, schema-evolving data like JSON logs, as it automatically reconciles differences between the source data and the catalog definition.

Exam trap

The trap here is that candidates often confuse the AWS Glue Schema Registry (which enforces schema compatibility and versioning) with the schema evolution capabilities of the Glue DynamicFrame, leading them to choose Option D even though it would reject schema changes rather than adapt to them.

How to eliminate wrong answers

Option A is wrong because manually updating the table schema before each run is not scalable, error-prone, and defeats the purpose of automated schema evolution; it also introduces operational overhead and potential downtime. Option B is wrong because using Spark SQL with a static schema definition in the script will cause the job to fail when new fields appear in the JSON logs, as Spark's static schema cannot adapt to dynamic changes without manual code modifications. Option D is wrong because the AWS Glue Schema Registry is designed for schema validation and serialization/deserialization (e.g., Avro, Protobuf) to enforce compatibility rules, not for dynamically merging evolving schemas during ETL processing; it would reject records that don't conform to the registered schema version, causing job failures instead of handling changes gracefully.

485
MCQmedium

A company runs a SQL Server transactional database on Amazon RDS. They need to capture change data (inserts, updates, deletes) in near real-time and replicate them to an Amazon S3 data lake. Which AWS service is most suitable?

A.AWS Database Migration Service (DMS) with change data capture
B.AWS Glue DataBrew
C.Amazon Kinesis Data Streams with Kinesis Client Library
D.Amazon Redshift Spectrum
AnswerA

DMS change data capture reads the database transaction log continuously, streaming inserts, updates and deletes to S3 with minimal source impact. This satisfies the near real-time replication requirement, unlike batch exports or snapshot-only tools that cannot capture ongoing changes.

Why this answer

AWS DMS with change data capture (CDC) is the most suitable service because it can continuously capture and replicate incremental changes (inserts, updates, deletes) from a SQL Server transactional database on Amazon RDS to an S3 data lake in near real-time. DMS uses native SQL Server transaction logs (e.g., MS-CDC or log-based replication) to read changes without impacting source performance, and it supports target S3 in formats like Parquet or CSV. This directly meets the requirement for near-real-time CDC replication to a data lake.

Exam trap

The trap here is that candidates may confuse Kinesis Data Streams as a general-purpose streaming solution for any real-time data, but it lacks native CDC capabilities for relational databases without additional custom code or connectors, making DMS the correct choice for database-to-S3 replication.

How to eliminate wrong answers

Option B is wrong because AWS Glue DataBrew is a visual data preparation tool for cleaning and normalizing data, not a service for capturing and replicating change data from a live database. Option C is wrong because Amazon Kinesis Data Streams is a real-time streaming service that requires custom producers and consumers (e.g., KCL) to ingest and process data, but it cannot natively capture CDC from a SQL Server database without additional middleware like Debezium or a custom application. Option D is wrong because Amazon Redshift Spectrum is a query engine that allows running SQL queries directly against data in S3, not a service for ingesting or replicating change data from a source database.

486
Multi-Selectmedium

Which TWO practices improve the performance of AWS Glue ETL jobs? (Choose two.)

Select 2 answers
A.Use pushdown predicates to filter data at the source
B.Increase the number of DPUs to the maximum allowed
C.Use the smallest possible file size for input data
D.Enable AWS Glue job metrics and debug logging
E.Use column pruning to select only required columns
AnswersA, E

Filters data early, reducing data scanned.

Why this answer

Pushdown predicates (Option A) improve AWS Glue ETL performance by filtering data at the source before it is read into the job. This reduces the volume of data transferred and processed, which is especially effective when using formats like Parquet or ORC that support predicate pushdown natively. By applying filters early, Glue avoids scanning unnecessary partitions or rows, leading to faster execution and lower costs.

Exam trap

The trap here is that candidates often confuse monitoring features (like enabling metrics and logging) with performance optimizations, or mistakenly believe that maximizing resources (DPUs) always improves speed, ignoring the overhead of small files and the benefits of early filtering and column selection.

487
MCQhard

A company is using Amazon S3 to store sensitive customer data. The security team requires that all data be encrypted in transit and at rest. Additionally, they want to prevent any accidental public access. Which combination of actions should the data engineer take?

A.Enable default encryption with SSE-S3, enforce HTTPS only via bucket policy, and enable S3 Block Public Access.
B.Enable default encryption with SSE-KMS, allow both HTTP and HTTPS, and set bucket ACLs to private.
C.Use client-side encryption, enforce HTTPS via bucket policy, and enable S3 Block Public Access.
D.Enable default encryption with SSE-S3, allow HTTP and HTTPS, and use bucket ACLs to block public access.
AnswerA

SSE-S3 default encryption protects data at rest, a bucket policy denying non-TLS requests enforces encryption in transit, and S3 Block Public Access prevents accidental exposure. Together these satisfy all three stated requirements: encryption at rest, encryption in transit, and no public access.

Why this answer

Enabling default encryption with SSE-S3 ensures data is encrypted at rest automatically, enforcing HTTPS only via bucket policy ensures encryption in transit by rejecting HTTP requests, and enabling S3 Block Public Access prevents any accidental public exposure regardless of bucket policies or ACLs. This combination satisfies all security requirements: encryption in transit, encryption at rest, and prevention of public access.

Exam trap

The trap here is that candidates may think bucket ACLs or client-side encryption alone satisfy the requirements, but the exam tests that S3 Block Public Access is needed to fully prevent accidental public access and that HTTPS enforcement is mandatory for encryption in transit.

How to eliminate wrong answers

Option B is wrong because allowing both HTTP and HTTPS violates the encryption-in-transit requirement; HTTPS-only must be enforced. Option C is wrong because client-side encryption does not guarantee encryption at rest on the server side (the data may be decrypted before upload), and the security team requires encryption at rest managed by AWS. Option D is wrong because allowing HTTP and HTTPS fails encryption-in-transit, and bucket ACLs alone are insufficient to block all public access (e.g., bucket policies can still grant public access).

488
MCQeasy

A company uses AWS KMS to encrypt data in Amazon S3. The security team wants to ensure that the KMS key can only be used from within the company's VPC. Which policy element should be added to the KMS key policy?

A.Set the Principal element to restrict access to the VPC.
B.Add a condition using aws:SourceIp to allow only IP addresses from the VPC.
C.Add a condition using aws:SourceVpc to allow only requests from the VPC.
D.Add a condition using kms:ViaService to allow only via VPC endpoints.
AnswerC

Adding an `aws:SourceVpc` condition to the KMS key policy restricts key usage to requests originating from the specified VPC endpoint, satisfying the requirement that the key only be usable from within the company's VPC. This condition evaluates the VPC ID of the source, blocking access from outside the VPC.

Why this answer

The correct approach is to add a condition using the aws:SourceVpc condition key in the KMS key policy. This condition evaluates the VPC ID from which the request originates, allowing you to restrict key usage to requests that come through a VPC endpoint within your VPC. When a request is made to KMS through a VPC endpoint, the aws:SourceVpc condition key is populated with the VPC ID, enabling precise control.

This ensures that the KMS key can only be used from within the specified VPC, meeting the security team's requirement.

Exam trap

The trap here is confusing the Principal element with condition keys, or mistakenly using aws:SourceIp or kms:ViaService instead of aws:SourceVpc. Candidates often think that restricting by IP or by service is sufficient to restrict to a VPC, but only aws:SourceVpc directly validates the VPC origin.

How to eliminate wrong answers

Option A is wrong because the Principal element in a policy specifies who (which AWS identities) can access the resource, not from where the request originates; it cannot restrict access based on VPC. Option B is wrong because aws:SourceIp checks the source IP address of the request, which for requests through a VPC endpoint is the private IP of the endpoint, not the VPC itself; moreover, IP addresses can change and are not a reliable way to restrict to a VPC. Option D is wrong because kms:ViaService is used to restrict usage to requests made through specific AWS services (e.g., s3.amazonaws.com), not to restrict to a VPC; it does not validate the VPC origin.

489
MCQhard

A company uses Amazon Kinesis Data Analytics for Apache Flink to process streaming data. The application reads from a Kinesis data stream, performs a 1-minute tumbling window aggregation, and writes results to an S3 bucket. Recently, the application started experiencing checkpoint failures and increasing processing delay. Which action should the engineer take FIRST to diagnose the issue?

A.Increase the parallelism of the Flink application.
B.Monitor CPU and memory utilization of the Flink application using Amazon CloudWatch metrics.
C.Switch to the Kinesis Client Library (KCL) for checkpointing.
D.Increase the checkpoint interval to reduce checkpoint frequency.
AnswerB

Checkpoint failures and rising processing delay typically stem from resource exhaustion, such as CPU or memory pressure, or backpressure within the Flink application. Reviewing CloudWatch CPU and memory metrics first identifies whether the task managers are saturated, directly addressing the stem's diagnostic constraint before deeper tuning.

Why this answer

When a Flink application shows checkpoint failures and growing processing delay, the first diagnostic step is to check resource utilization via CloudWatch metrics (CPU, memory, heap, checkpoint duration). Checkpoint failures often stem from backpressure or resource exhaustion, so understanding whether the application is CPU- or memory-bound guides the next action. This is the least invasive, most informative first step.

Exam trap

DEA-C01 often tests the urge to jump to a fix (increase parallelism, change checkpoint interval) instead of first gathering diagnostic data, so candidates pick a remediation rather than a monitoring step.

How to eliminate wrong answers

Option A is wrong because increasing parallelism before diagnosing the root cause can worsen the problem (e.g., if the bottleneck is the source shard count or downstream sink). Option C is wrong because KCL is for Kinesis consumer applications, not Flink checkpointing — Flink manages its own checkpointing via the Flink runtime. Option D is wrong because increasing the checkpoint interval masks symptoms rather than diagnosing them, and can increase recovery time.

490
Multi-Selecthard

A company is ingesting real-time financial transactions into Amazon Kinesis Data Streams. The data is then consumed by a Kinesis Data Analytics for Apache Flink application that calculates running totals. The application is experiencing high latency and checkpoint failures. Which TWO steps should the engineer take to improve performance and reliability? (Select TWO.)

Select 2 answers
A.Enable enhanced fan-out for the Flink application.
B.Reduce the batch size of records processed per checkpoint.
C.Increase the number of shards in the Kinesis data stream.
D.Increase the number of KPUs (Kinesis Processing Units) for the Flink application.
E.Decrease the checkpoint interval to reduce state size.
AnswersC, D

Adding shards raises the stream's ingest and read throughput ceiling, relieving the backpressure that stalls the Flink consumer and triggers checkpoint timeouts. This directly addresses the latency and checkpoint failures caused by insufficient parallelism in the source.

Why this answer

Option C is correct because increasing the number of shards in the Kinesis data stream raises the stream's read throughput capacity (each shard supports up to 2 MB/sec read and 1 MB/sec write, or 1,000 records/sec), which relieves the ingestion bottleneck that causes high latency in the Flink application. Option D is correct because increasing KPUs adds parallel processing resources (each KPU provides 1 vCPU and 4 GB memory) to the Kinesis Data Analytics for Apache Flink application, allowing it to consume and process records faster and complete checkpoints more reliably. Option A is not correct because enhanced fan-out is a Kinesis Data Streams consumer feature that gives dedicated 2 MB/sec read throughput per consumer via HTTP/2, but it does not address the Flink application's internal processing or checkpointing bottlenecks.

Option B is not correct because reducing the batch size per checkpoint does not improve throughput and can actually increase checkpoint overhead and latency. Option E is not correct because decreasing the checkpoint interval makes checkpoints more frequent, increasing overhead and state I/O rather than reducing state size or improving reliability.

Exam trap

DEA-C01 often tests the misconception that reducing checkpoint interval or batch size improves reliability, when in reality checkpoint failures under load are solved by increasing parallelism (KPUs) and stream capacity (shards), not by making checkpoints more frequent or smaller.

491
MCQeasy

A data engineer needs to schedule a daily AWS Glue ETL job that transforms data in Amazon S3. The job must run at 2:00 AM UTC every day. What is the simplest way to achieve this?

A.Use AWS Step Functions with a Wait state and a Glue task to run the job daily.
B.Create an AWS Glue trigger of type SCHEDULED with a cron expression.
C.Use Amazon EventBridge to schedule a Lambda function that calls the Glue StartJobRun API.
D.Configure an AWS Glue workflow with a schedule and add the job as a starting trigger.
AnswerB

AWS Glue triggers can be scheduled with a cron expression. You can create a trigger that fires at 2:00 AM UTC daily and starts the Glue job. This is the native and simplest method within AWS Glue, requiring no external services. It integrates directly with the job and provides status monitoring.

Why this answer

The simplest way to schedule a Glue job is to use an AWS Glue scheduled trigger with a cron expression. It is a built-in feature that directly starts the job at the specified time. Other options involve additional services like Lambda, EventBridge, or Step Functions, which add complexity without benefit for a single scheduled job.

Exam trap

The trap here is overcomplicating the solution by using external orchestration services when AWS Glue provides a native scheduling mechanism.

492
MCQhard

A data engineer is building an AWS Glue ETL job that reads JSON files from Amazon S3, flattens nested arrays, and writes Parquet to another S3 bucket. The job runs daily and processes about 2 TB. The engineer notices that job runs are failing intermittently with OutOfMemory errors during the shuffle phase. The job uses 10 G.1X workers. Which change should the engineer make to resolve the memory failures while minimizing cost?

A.Reduce the number of workers from 10 to 5 to lower concurrency.
B.Convert the source JSON files to CSV before running the Glue job.
C.Enable AWS Glue job bookmarks to skip previously processed files.
D.Switch the worker type to G.2X to provide more memory and disk per worker.
AnswerD

G.2X workers provide twice the memory and disk of G.1X workers, which directly addresses OutOfMemory failures during shuffle-intensive operations like flattening nested arrays. Because the job is memory-bound rather than CPU-bound, increasing per-worker memory is the most targeted fix. Keeping the worker count the same with larger workers often costs less than adding many small workers for shuffle-heavy workloads.

Why this answer

OutOfMemory errors during the shuffle phase of a Glue job indicate insufficient per-worker memory. G.2X workers provide double the memory and disk of G.1X, which is the most direct and cost-effective fix for shuffle-heavy transformations such as flattening nested arrays. Reducing workers or changing the source format does not address the memory bottleneck and can worsen the problem.

Exam trap

The trap here is assuming that adding more small workers or changing the input format will fix shuffle memory errors, when the issue is per-executor memory that only a larger worker type resolves.

493
MCQhard

A data engineer manages an Amazon Kinesis Data Stream with multiple shards. The stream is experiencing high throughput, and the engineer notices that some shards are throttling while others are underutilized. The engineer needs to redistribute the data evenly across shards to avoid throttling. Which action should the engineer take?

A.Use the MergeShards API to combine underutilized shards and reduce the total number of shards.
B.Modify the partition key strategy in the producer to ensure even distribution across shards.
C.Increase the number of shards using the UpdateShardCount API to add more capacity.
D.Enable enhanced fan-out for the stream to provide dedicated throughput for consumers.
AnswerB

Throttling on specific shards while others are underutilized indicates a skewed partition key. By changing the partition key strategy—such as adding a random suffix or using a more uniform key—the producer can distribute records more evenly across all shards, resolving the throttling without changing shard count.

Why this answer

Uneven shard utilization with throttling on some shards is typically caused by a partition key that does not distribute records uniformly. Changing the partition key strategy—for example, by incorporating a random value or using a composite key—ensures records are spread evenly across all shards. This addresses the root cause without altering shard count or read capacity.

Exam trap

The trap here is thinking that adding more shards or enabling enhanced fan-out will fix throttling, when the real issue is data distribution due to partition key skew.

494
MCQhard

A data engineer is using AWS Step Functions to orchestrate a daily ingestion workflow. The workflow calls an AWS Glue job, then runs an AWS Lambda function to validate output, and finally starts an Amazon Redshift stored procedure. The engineer needs to ensure that if the Glue job fails, the workflow retries the Glue job up to three times with exponential backoff before failing the entire execution. Which Step Functions feature should the engineer configure?

A.Enable idempotency on the Glue job so repeated invocations succeed.
B.Configure a Retry policy on the Glue job task with MaxAttempts and BackoffRate.
C.Increase the Step Functions execution timeout to allow the Glue job to rerun.
D.Configure a Catch block on the Glue job task to transition to a failure state.
AnswerB

Step Functions Retry policies are applied per state and support MaxAttempts, IntervalSeconds, MaxDelaySeconds, and BackoffRate. Setting MaxAttempts to three with a BackoffRate greater than one implements exponential backoff for the Glue job task specifically, which is exactly the required behavior for retrying only that step before the execution fails.

Why this answer

Step Functions Retry policies are defined on individual states and accept MaxAttempts, IntervalSeconds, MaxDelaySeconds, and BackoffRate. Applying a Retry to the Glue job task with MaxAttempts of three and a BackoffRate above one produces exponential backoff retries scoped to that task. If all attempts fail, the state then fails and can be handled by a Catch block.

Exam trap

The trap here is conflating Catch, which redirects execution on error, with Retry, which re-invokes the failed state and is the only feature that produces the required retry-with-backoff behavior.

495
MCQmedium

A data engineer is building a streaming ingestion pipeline using Amazon Kinesis Data Streams. The producer application writes records with an explicit partition key derived from the device ID, and there are approximately 2,000 active devices. The engineer needs to ensure that records for the same device are processed in order by a downstream consumer. Which configuration should the engineer verify to guarantee per-device ordering?

A.Confirm that the partition key is stable per device so records map to the same shard.
B.Increase the stream's retention period to 365 days to preserve ordering.
C.Ensure all records use the same partition key so they land in a single shard.
D.Enable enhanced fan-out on the stream so each consumer gets a dedicated read throughput.
AnswerA

Kinesis Data Streams preserves order within a shard, and records with the same partition key are consistently routed to the same shard via the MD5 hash of the key. By keeping the partition key stable per device, all records from a given device land in the same shard and are delivered in order, satisfying the requirement while retaining parallelism across shards.

Why this answer

Kinesis Data Streams guarantees ordering only within an individual shard. Because the partition key is hashed to select a shard, any record sharing the same partition key is consistently routed to the same shard. Keeping the device identifier as the partition key ensures all records for a device are serialized in one shard, delivering the required per-device ordering without sacrificing horizontal scale across the fleet.

Exam trap

The trap here is assuming that any stream-level feature such as enhanced fan-out or extended retention can enforce ordering, when ordering in Kinesis Data Streams is strictly a per-shard property driven by the partition key.

496
MCQhard

A data engineer is configuring an Amazon S3 bucket for a data lake. The bucket must store sensitive financial data and comply with a regulation that requires all data to be encrypted at rest with keys that are automatically rotated every year. The engineer also needs to audit key usage and control access to the keys separately from other AWS services. Which encryption option should the engineer choose?

A.Server-side encryption with Amazon S3 managed keys (SSE-S3)
B.Server-side encryption with AWS KMS keys (SSE-KMS)
C.Server-side encryption with customer-provided keys (SSE-C)
D.Client-side encryption with an AWS KMS customer managed key
AnswerB

SSE-KMS uses AWS KMS customer managed keys, which can be configured for automatic annual rotation. KMS provides separate access control through key policies and IAM, and it logs key usage in AWS CloudTrail for auditing. This meets the compliance requirements for encryption, rotation, auditing, and separate key management. SSE-KMS is the appropriate choice for sensitive data with regulatory mandates.

Why this answer

SSE-KMS with customer managed keys provides automatic key rotation, separate access control via key policies and IAM, and detailed auditing through CloudTrail. These features directly address the compliance requirements for encryption at rest, annual key rotation, and auditing key usage. SSE-S3 and SSE-C lack the necessary auditing and separate access control, while client-side encryption shifts management burden to the application.

Exam trap

The trap here is assuming that any encryption option with automatic rotation meets all compliance needs, ignoring the requirement for separate key access control and auditing.

497
MCQhard

A data engineer is using AWS Glue to process a large dataset in Amazon S3. The dataset consists of many small JSON files (average 100 KB each) stored in a single prefix. The Glue job reads these files, performs transformations, and writes the output to Parquet in another S3 location. The job is running slowly and consuming many DPUs. Which action should the data engineer take to improve performance?

A.Use AWS Glue's `groupFiles` option to group multiple small files into larger chunks for processing.
B.Convert the JSON files to Parquet using an AWS Glue crawler before processing.
C.Enable AWS Glue job bookmarks to skip already processed files.
D.Increase the number of DPUs allocated to the job to enable more parallel tasks.
AnswerA

AWS Glue provides the `groupFiles` option (e.g., 'inPartition') to group small files into larger groups, reducing the number of tasks and improving read throughput. This is specifically designed to handle many small files by coalescing them, which reduces overhead and improves job performance without excessive DPUs.

Why this answer

The `groupFiles` option in AWS Glue allows the job to group multiple small files into a single partition, reducing the number of tasks and the overhead of reading many small files. This directly addresses the performance bottleneck caused by many small JSON files, improving job speed and reducing DPU consumption. Other options do not solve the small file issue.

Exam trap

The trap here is assuming that adding more DPUs or using bookmarks will solve performance issues with many small files, when the real solution is to group files to reduce overhead.

498
MCQmedium

A company wants to ingest streaming data from thousands of IoT devices into Amazon S3 with minimal latency and then transform the data using Spark SQL. Which AWS service should be used for data ingestion?

A.Amazon EMR
B.AWS Glue
C.Amazon Athena
D.Amazon Kinesis Data Firehose
AnswerD

Kinesis Data Firehose ingests streaming data and delivers it directly into Amazon S3 with minimal latency, requiring no custom consumer code. It satisfies the ingestion requirement, after which Spark SQL can transform the landed data.

Why this answer

Amazon Kinesis Data Firehose is the correct choice because it is a fully managed service designed for ingesting streaming data into Amazon S3 with near-real-time latency (typically 60 seconds or less). It can directly write data to S3 without requiring custom code or additional infrastructure, and it supports optional transformations via AWS Lambda, making it ideal for the described use case of streaming IoT data ingestion.

Exam trap

The trap here is confusing data ingestion services (Kinesis Data Firehose) with data processing or query services (EMR, Glue, Athena), leading candidates to pick EMR for its Spark SQL capability instead of recognizing that Firehose handles the ingestion step before transformation.

How to eliminate wrong answers

Option A is wrong because Amazon EMR is a big data processing service for running frameworks like Spark and Hadoop, not a streaming ingestion service; it would require additional setup (e.g., Kinesis or Kafka) to ingest data into S3. Option B is wrong because AWS Glue is a serverless ETL service primarily for batch data transformation and cataloging, not designed for real-time streaming ingestion into S3. Option C is wrong because Amazon Athena is an interactive query service for analyzing data in S3 using SQL, not an ingestion tool; it cannot ingest streaming data.

499
MCQeasy

A data engineer needs to store large volumes of semi-structured JSON data in Amazon S3 and query it using Amazon Athena. The data is generated continuously and appended to S3 in small files. The engineer wants to optimize query performance and reduce costs. Which action should the engineer take?

A.Convert the JSON data to Apache Parquet and partition it by date.
B.Enable S3 Transfer Acceleration for faster data uploads.
C.Use Amazon S3 Select to query the JSON data directly.
D.Store the JSON data in a single large file per day.
AnswerA

Converting JSON to Parquet provides columnar storage, enabling Athena to read only needed columns and reduce scan volume. Partitioning by date allows Athena to prune partitions based on query filters, further reducing data scanned. This combination significantly improves query performance and lowers costs, making it the optimal solution for large-scale semi-structured data queried by Athena.

Why this answer

Converting JSON to Parquet and partitioning by date leverages columnar storage and partition pruning, which are key optimizations for Athena. Parquet reduces the amount of data scanned by reading only required columns, while partitioning allows Athena to skip irrelevant partitions based on query filters. Together, they minimize cost and improve performance for large-scale semi-structured data.

Exam trap

The trap here is focusing on file size or upload speed instead of the format and partitioning strategy that directly impact Athena query cost and performance.

500
MCQhard

A large e-commerce company uses Amazon DynamoDB to store shopping cart data. The table has a partition key of 'user_id' and a sort key of 'item_id'. The application performs frequent updates to the 'quantity' attribute for items in a user's cart. Recently, the operations team noticed that write requests are being throttled during peak shopping hours. The table is provisioned with 10,000 write capacity units (WCUs) and uses DynamoDB Accelerator (DAX) for read caching. The data engineer suspects that the throttling is due to hot partitions. The application uses a single AWS SDK client configured with retries. After reviewing the Amazon CloudWatch metrics, the engineer sees that the WriteThrottleEvents metric spikes for a few partition keys. The table has a high number of partitions. What should the data engineer do to resolve the throttling issue with minimal application changes?

A.Increase the provisioned write capacity to 20,000 WCUs permanently.
B.Enable DynamoDB Global Tables to distribute writes across regions.
C.Add more nodes to the DAX cluster to offload write traffic.
D.Configure DynamoDB Auto Scaling with a maximum WCU setting of 20,000 and a target utilization of 70%.
AnswerD

Auto Scaling dynamically adjusts capacity based on traffic, reducing throttling without permanent overprovisioning.

Why this answer

DynamoDB Auto Scaling can dynamically adjust write capacity in response to traffic patterns, reducing throttling on hot partitions without requiring application changes. Option A is incorrect because permanently increasing WCUs does not adapt to variable demand and may lead to over-provisioning. Option B (Global Tables) replicates data across regions but does not increase write capacity for a single table, so it does not resolve hot partition throttling.

Option C (DAX) is a read cache and does not offload write traffic; it only improves read performance.

501
MCQhard

A financial services company stores sensitive transaction data in an Amazon S3 bucket. The security team requires that all data be encrypted at rest using AWS Key Management Service (AWS KMS) customer managed keys, and that the encryption keys be automatically rotated every year. A data engineer needs to configure the S3 bucket to meet these requirements with minimal ongoing operational effort. Which solution should the engineer implement?

A.Enable default encryption on the S3 bucket using SSE-S3 (AES-256).
B.Enable default encryption on the S3 bucket using SSE-KMS with a customer managed key that has automatic key rotation enabled.
C.Use S3 client-side encryption with a customer-provided key stored in AWS Secrets Manager, and rotate the key annually using a Lambda function.
D.Enable default encryption on the S3 bucket using SSE-KMS with an AWS managed key (aws/s3).
AnswerB

SSE-KMS with a customer managed key allows the use of AWS KMS keys that the customer controls. Enabling automatic key rotation on the customer managed key rotates the key material every year by default, meeting the annual rotation requirement. This configuration requires minimal ongoing effort because S3 automatically encrypts new objects with the specified key.

Why this answer

To meet the requirements, the engineer should use server-side encryption with AWS KMS customer managed keys (SSE-KMS) and enable automatic key rotation on the key. This ensures data is encrypted at rest with a key the customer controls, and rotation occurs annually without manual intervention. The bucket default encryption setting ensures all new objects are encrypted automatically.

Exam trap

The trap here is assuming that SSE-S3 or AWS managed keys provide customer control over rotation, when only customer managed keys allow configuring annual automatic rotation.

502
MCQeasy

A data engineer needs to grant an AWS Lambda function permission to read objects from a specific Amazon S3 bucket. The Lambda function assumes an IAM role. Which policy should the engineer attach to the IAM role to allow the Lambda function to read objects from the bucket?

A.An IAM policy with s3:PutObject and s3:GetObject actions, with resource set to the bucket ARN and objects ARN.
B.An IAM policy with s3:ListBucket action and resource set to the objects ARN only.
C.An IAM policy with s3:GetObject and s3:ListBucket actions, with the resource set to the bucket ARN and objects ARN.
D.An IAM policy with s3:GetObject action and resource set to the bucket ARN only.
AnswerC

To read objects from an S3 bucket, the Lambda function needs s3:GetObject on the objects and s3:ListBucket on the bucket. The resource for s3:GetObject should be the object ARN (e.g., arn:aws:s3:::bucket/*) and for s3:ListBucket the bucket ARN (arn:aws:s3:::bucket). This policy grants the necessary permissions correctly.

Why this answer

To read objects from S3, the Lambda function requires s3:GetObject on the object ARN and s3:ListBucket on the bucket ARN. This combination allows listing the bucket and retrieving objects. Policies that omit the object ARN for GetObject or include unnecessary write permissions are incorrect.

The correct policy follows the principle of least privilege.

Exam trap

The trap here is confusing the resource ARN requirements for bucket-level versus object-level S3 actions, leading to policies that either do not work or grant excessive permissions.

503
Multi-Selectmedium

Which THREE are best practices for managing data in Amazon S3 for a data lake? (Choose three.)

Select 3 answers
A.Enable S3 Versioning to protect against accidental deletions.
B.Configure lifecycle policies to transition data to colder storage tiers.
C.Enable S3 Snapshot for point-in-time recovery.
D.Disable S3 server access logging to reduce costs.
E.Use bucket policies to restrict access based on IAM roles.
AnswersA, B, E

Versioning provides data protection.

Why this answer

Enabling S3 Versioning is a best practice for data lakes because it protects against accidental deletions or overwrites by preserving all versions of an object, including deletions (which are recorded as delete markers). This allows you to recover previous object states and is essential for data governance and auditability in a data lake environment.

Exam trap

The trap here is that candidates may confuse S3 Versioning with a non-existent 'S3 Snapshot' feature, or mistakenly think disabling server access logging is a cost-saving best practice, when in fact it undermines security auditing.

504
Drag & Dropmedium

Arrange the steps to set up a streaming ETL pipeline using Amazon Kinesis Data Firehose to Amazon S3.

Drag or tap steps into the slots.

Steps
Order
1Step 1
2Step 2
3Step 3
4Step 4

Why this order

First, create the Firehose stream, configure source, set S3 destination, enable optional Lambda transformation, and test.

505
MCQhard

A data engineer is using AWS Glue DataBrew to clean a dataset stored in Amazon S3. The dataset contains a column with inconsistent date formats and another with trailing whitespace. The engineer wants to apply these transformations reproducibly and schedule the recipe to run daily. Which combination of steps should the engineer take?

A.Use the DataBrew console to apply transformations, then export the results manually each day without creating a recipe or job.
B.Create a recipe, then attach it to a Glue ETL job as a transform step and schedule the Glue job using a trigger.
C.Create a Glue crawler to catalog the dataset, then use a Glue ETL job with custom PySpark code to apply the transformations on a schedule.
D.Create a project, add the dataset, build a recipe with the appropriate transformation steps, publish the recipe, and create a job that runs the recipe on a schedule.
AnswerD

DataBrew projects are used to explore and build recipes interactively. Publishing the recipe captures the transformation steps, and a DataBrew job applies the recipe to the dataset on a defined schedule. This workflow is the intended way to perform reproducible, scheduled cleaning tasks in DataBrew.

Why this answer

DataBrew workflows use projects to build recipes, which are published and then executed by jobs on a schedule. This captures the date-format and whitespace transformations reproducibly and runs them daily, matching the requirement exactly.

Exam trap

The trap here is assuming a DataBrew recipe can be attached to a Glue ETL job, when recipes are executed only through DataBrew jobs.

506
MCQhard

A data streaming application uses Kinesis Data Streams with 10 shards. The data producer is throttled frequently. Which action should be taken to resolve this issue?

A.Decrease the data retention period
B.Use enhanced fan-out for consumers
C.Enable server-side encryption
D.Increase the number of shards
AnswerD

Each Kinesis shard supports a fixed ingest ceiling of 1 MB/s or 1,000 records/s, so ten shards cap throughput and cause throttling once exceeded. Adding shards raises that aggregate write capacity, directly resolving the producer throttling described in the stem.

Why this answer

Throttling in Kinesis Data Streams occurs when the write throughput exceeds the shard limits. Each shard supports up to 1 MB/s or 1,000 records/s for writes. With 10 shards, the total write capacity is 10 MB/s or 10,000 records/s.

Increasing the number of shards (Option D) directly increases the write capacity, resolving the throttling issue by distributing the load across more shards.

Exam trap

The trap here is that candidates confuse consumer-side features (like enhanced fan-out or retention period) with producer-side capacity issues, leading them to pick options that do not address the root cause of write throttling.

How to eliminate wrong answers

Option A is wrong because decreasing the data retention period (default 24 hours, up to 365 days) does not affect write throughput or throttling; it only controls how long records are stored. Option B is wrong because enhanced fan-out is a consumer-side feature that provides dedicated read throughput (2 MB/s per consumer per shard) and does not address producer-side write throttling. Option C is wrong because enabling server-side encryption (SSE-S3 or SSE-KMS) secures data at rest but has no impact on write throughput or throttling.

507
MCQhard

Refer to the exhibit. A data engineer has attached this bucket policy to an S3 bucket. What is the effect of this policy?

A.It enforces server-side encryption for all objects written to the bucket.
B.It allows the DataLakeRole to read and write objects, but only over HTTPS.
C.It allows anonymous access to the bucket for HTTPS requests.
D.It denies all access to the bucket except for requests from the DataLakeRole.
AnswerB

The policy grants DataLakeRole s3:GetObject and s3:PutObject permissions, but the condition restricts access to requests made over TLS. Plain HTTP requests are denied, so the role can read and write objects only when the connection is encrypted via HTTPS.

Why this answer

The bucket policy uses a condition key `aws:SecureTransport` set to `true`, which restricts access to HTTPS (TLS) connections only. The `Principal` is `DataLakeRole`, and the `Action` includes `s3:GetObject` and `s3:PutObject`, so the policy allows that role to read and write objects exclusively over HTTPS, enforcing encrypted data in transit.

Exam trap

AWS often tests the distinction between encryption in transit (HTTPS/TLS) and encryption at rest (SSE), leading candidates to confuse the `aws:SecureTransport` condition with server-side encryption requirements.

How to eliminate wrong answers

Option A is wrong because the policy does not reference `s3:x-amz-server-side-encryption` or any condition enforcing server-side encryption (SSE) at rest; it only enforces encryption in transit via `aws:SecureTransport`. Option C is wrong because the `Principal` is explicitly set to `DataLakeRole` (an IAM role ARN), not `"*"` or `{"AWS": "*"}`, so anonymous access is not granted. Option D is wrong because the policy includes an `Allow` effect for `DataLakeRole` under the HTTPS condition, but it does not contain a `Deny` statement for other principals or conditions; without an explicit `Deny`, other access may still be allowed by other policies (e.g., bucket ACLs or IAM policies), so it does not deny all other access.

508
Multi-Selecthard

A company wants to implement least privilege access for its data lake on S3. Which THREE practices should be followed? (Choose THREE.)

Select 3 answers
A.Grant s3:* to all users for simplicity
B.Use S3 bucket policies for cross-account access
C.Use S3 access points to enforce network policies
D.Disable S3 Block Public Access to allow flexibility
E.Use IAM policies to grant specific permissions to users and roles
AnswersB, C, E

S3 bucket policies are resource-based and evaluate the bucket owner's permissions, so they grant cross-account principals access without sharing long-lived IAM credentials. This satisfies least privilege by scoping access to specific buckets or prefixes, and by letting the data lake owner retain control over who may read or write.

Why this answer

Option B is correct because S3 bucket policies are resource-based policies that explicitly define which principals (including cross-account identities) may perform which s3: actions on the bucket, allowing tightly scoped cross-account access instead of broad grants. Option C is correct because S3 access points provide dedicated endpoints with their own access point policies and can be restricted to a VPC via network origin controls, enforcing network-level least privilege for data lake access. Option E is correct because IAM policies attached to users and roles grant only the specific S3 actions and resources needed, which is the core mechanism for implementing least privilege.

Option A is incorrect because granting s3:* violates least privilege by allowing all S3 operations. Option D is incorrect because disabling S3 Block Public Access increases exposure risk and does not support least privilege.

509
MCQeasy

A company uses Amazon S3 to store sensitive data. The security team requires that all data be encrypted at rest using a customer-managed key that is rotated annually. Which encryption option should be used?

A.SSE-KMS (Server-Side Encryption with AWS KMS).
B.SSE-S3 (Server-Side Encryption with S3-managed keys).
C.Client-side encryption.
D.SSE-C (Server-Side Encryption with Customer-Provided keys).
AnswerA

SSE-KMS encrypts objects at rest using keys held in AWS KMS, and supports customer-managed keys with automatic annual rotation. This satisfies the requirement for a customer-managed, annually rotated key, unlike SSE-S3, which uses AWS-managed keys that cannot be rotated on demand.

Why this answer

SSE-KMS is the correct choice because it allows you to use a customer-managed key (CMK) in AWS KMS, which you can configure to rotate automatically on an annual schedule. This satisfies the security team's requirement for encryption at rest with a key you control and rotate yearly, while still leveraging server-side encryption that integrates with S3's existing infrastructure.

Exam trap

The trap here is that candidates often confuse SSE-C with customer-managed keys, but SSE-C requires you to supply the key on every operation and does not support AWS-managed rotation, making it unsuitable for the 'rotated annually' requirement.

How to eliminate wrong answers

Option B (SSE-S3) is wrong because it uses S3-managed keys that are automatically rotated by AWS, not customer-managed keys, so you cannot control the rotation schedule or manage the key yourself. Option C (Client-side encryption) is wrong because it encrypts data before it reaches S3, which does not meet the requirement for server-side encryption at rest managed by AWS; it also places the key management burden entirely on the client, not the customer-managed key service. Option D (SSE-C) is wrong because it requires you to provide your own encryption key with each request, and AWS does not manage or rotate the key—you must handle key storage and rotation entirely outside of AWS, which contradicts the requirement for a customer-managed key that is rotated annually within AWS.

510
MCQhard

Refer to the exhibit. A data engineer has attached this bucket policy to an S3 bucket named data-lake-bucket. The engineer wants to allow only GET requests from the corporate network (10.0.0.0/16) over HTTPS. However, users report that they cannot access objects even when connected to the corporate network. What is the issue?

A.The Deny statement should include a condition on the source IP.
B.The Allow statement should include a condition for SecureTransport.
C.The Allow statement should specify s3:GetObject instead of s3:GetObject.
D.The Deny statement blocks all requests that are not using HTTPS, including those from the corporate network.
AnswerD

Deny overrides Allow when condition is met.

Why this answer

The Deny statement with `aws:SecureTransport` set to `false` blocks all HTTP requests. Since the Allow statement only permits GET requests from the corporate network (10.0.0.0/16) but does not require HTTPS, any request from that network that uses HTTP is denied by the explicit Deny. The Deny statement overrides the Allow, so even legitimate corporate users are blocked if they use HTTP.

Exam trap

AWS often tests the principle that an explicit Deny overrides any Allow, leading candidates to focus on fixing the Allow statement rather than recognizing that the Deny unconditionally blocks HTTP traffic from all sources, including the corporate network.

How to eliminate wrong answers

Option A is wrong because the Deny statement already includes a condition on `aws:SecureTransport`, not on source IP; adding a source IP condition would not fix the HTTPS enforcement issue. Option B is wrong because the Allow statement already includes a condition for `aws:SecureTransport` equal to `true` in the Deny, but the Allow itself lacks a SecureTransport condition, so it permits both HTTP and HTTPS; adding SecureTransport to the Allow would not resolve the Deny blocking HTTP. Option C is wrong because `s3:GetObject` is the correct action for GET requests; the typo 's3:GetObject' in the question is a red herring, and the actual policy uses the correct action.

511
MCQmedium

A company is using Amazon RDS for MySQL and needs to automate backups with a retention period of 35 days. They also want to be able to restore to any point within the retention period. Which configuration should be used?

A.Enable manual snapshots daily and retain for 35 days.
B.Set the backup retention period to 35 days and enable automatic backups.
C.Set the backup retention period to 7 days and create daily manual snapshots.
D.Disable automated backups and rely on Multi-AZ for recovery.
AnswerB

Automatic backups capture daily snapshots plus continuous transaction logs, enabling point-in-time recovery to any second within the retention window. Setting retention to 35 days satisfies the stated requirement exactly, since RDS permits 0–35 days for MySQL. Manual snapshots alone cannot deliver point-in-time restore.

Why this answer

Amazon RDS for MySQL supports automated backups with a configurable retention period of up to 35 days. By setting the backup retention period to 35 days and enabling automatic backups, RDS automatically performs daily snapshots and transaction log backups, enabling point-in-time recovery (PITR) to any second within the retention window. This meets the requirement for both a 35-day retention and full PITR capability without manual intervention.

Exam trap

The trap here is that candidates often confuse manual snapshots (which are retained indefinitely but do not support PITR) with automated backups (which support PITR but have a maximum retention of 35 days), leading them to choose Option A or C, thinking manual snapshots can extend the PITR window.

How to eliminate wrong answers

Option A is wrong because manual snapshots are not automatically taken daily and do not support point-in-time recovery; they only provide a single point-in-time restore, not continuous PITR. Option C is wrong because setting the backup retention period to 7 days limits automated backups and PITR to only 7 days, and adding daily manual snapshots does not extend the PITR window beyond 7 days. Option D is wrong because disabling automated backups eliminates both automated snapshots and transaction log backups, making PITR impossible; Multi-AZ provides high availability but does not create backups or enable recovery to any point in time.

512
MCQmedium

A data engineer needs to implement encryption for data at rest in an Amazon S3 bucket that stores sensitive financial records. The company's security policy requires that the encryption keys be managed by the company and rotated annually. The data engineer wants to use AWS Key Management Service (AWS KMS) to meet these requirements. Which solution should the data engineer implement?

A.Use client-side encryption with a customer-provided key stored in AWS Secrets Manager, and manually rotate the key annually.
B.Use server-side encryption with Amazon S3 managed keys (SSE-S3) and enable automatic key rotation.
C.Use server-side encryption with customer-provided keys (SSE-C) and rotate the keys by re-uploading objects with new keys annually.
D.Use server-side encryption with AWS KMS customer managed keys (SSE-KMS) and configure automatic key rotation for the KMS key.
AnswerD

SSE-KMS with customer managed keys allows the company to create and manage KMS keys. AWS KMS supports automatic annual rotation of customer managed keys, which aligns with the security policy. This solution provides the required control over key management and rotation, and it encrypts data at rest in S3 using KMS keys.

Why this answer

SSE-KMS with customer managed keys enables the company to control KMS keys and configure automatic annual rotation. This meets the security policy for company-managed keys and rotation, while providing seamless encryption for S3 objects.

Exam trap

The trap here is confusing SSE-S3 with SSE-KMS; SSE-S3 uses Amazon-managed keys and does not provide customer control over key management or rotation.

513
MCQeasy

A data engineer needs to monitor the number of records processed by an AWS Glue ETL job. Which CloudWatch metric should the engineer use?

A.glue.driver.aggregate.elapsedTime
B.glue.driver.aggregate.numRecords
C.glue.driver.aggregate.bytesRead
D.glue.driver.aggregate.recordsRead
AnswerB

The glue.driver.aggregate.numRecords metric reports the total record count processed by an AWS Glue ETL job, aggregated at the driver. It directly answers the requirement to monitor how many records the job handles, unlike per-executor or duration metrics.

Why this answer

Glue emits a 'glue.driver.aggregate.numRecords' metric for the number of records processed. Option A is wrong because 'glue.driver.aggregate.elapsedTime' is for time. Option C is wrong because 'glue.driver.aggregate.bytesRead' is for bytes.

Option D is wrong because 'glue.driver.aggregate.recordsRead' is not a standard metric.

514
MCQmedium

A data engineer needs to load data from an Amazon DynamoDB table into an Amazon S3 bucket for analytics. The table is approximately 500 GB and has a high volume of write traffic. The engineer must minimize the impact on the table's read capacity and avoid consuming provisioned throughput. What is the MOST appropriate method to export the data?

A.Use the DynamoDB export to Amazon S3 feature to export the table to an S3 bucket.
B.Use AWS Database Migration Service (AWS DMS) with a DynamoDB source and an S3 target.
C.Use AWS Glue to read the DynamoDB table with the dynamodb connector and write to Amazon S3.
D.Enable DynamoDB Streams and use an AWS Lambda function to write changed records to Amazon S3.
AnswerA

DynamoDB's native export to Amazon S3 uses a consistent point-in-time snapshot and does not consume any read capacity units on the table. It is designed for large tables and high-traffic workloads, making it the ideal choice to avoid impacting provisioned throughput while exporting 500 GB.

Why this answer

The DynamoDB export to Amazon S3 feature creates a point-in-time export without consuming any read capacity, which is critical for a high-traffic 500 GB table. Other methods such as Glue or DMS perform scans that consume RCUs, and Streams only provides incremental changes. Therefore, the native export is the only option that fully satisfies the requirement.

Exam trap

The trap here is assuming that any export method must scan the table and consume read capacity, overlooking the zero-RCU native export feature.

515
MCQhard

A data engineer is designing a data lake on Amazon S3 that contains personally identifiable information (PII). The compliance team requires that all access to the data be logged and that any attempt to delete or modify data be detected and alerted. The engineer enables AWS CloudTrail data events for the S3 bucket and configures Amazon CloudWatch alarms. Which additional AWS service should the engineer use to automatically detect and remediate unauthorized changes to the S3 bucket's ACLs or policies?

A.AWS Trusted Advisor with S3 bucket permissions checks.
B.AWS Config with managed rules and automatic remediation.
C.Amazon Macie with custom data identifiers.
D.Amazon GuardDuty with S3 protection.
AnswerB

AWS Config can monitor S3 bucket ACLs and policies for changes, evaluate them against desired configurations, and trigger automatic remediation using AWS Systems Manager Automation documents. This provides continuous detection and enforcement, aligning with the requirement to detect and remediate unauthorized changes.

Why this answer

AWS Config is designed to assess, audit, and evaluate configurations of AWS resources. With managed rules for S3, it can detect changes to ACLs and policies, and trigger automatic remediation via SSM Automation. Other services like Macie, Trusted Advisor, and GuardDuty focus on data classification, recommendations, or threat detection, not configuration compliance and remediation.

Exam trap

The trap here is confusing threat detection or data classification services with configuration compliance and remediation, which are the core capabilities of AWS Config.

516
Multi-Selectmedium

A company is designing a data store for IoT sensor data that is written once and never updated. The data must be stored with high durability and low cost. Which TWO AWS storage services are most suitable? (Choose TWO.)

Select 2 answers
A.Amazon ElastiCache
B.Amazon EBS
C.Amazon S3
D.Amazon DynamoDB
E.Amazon S3 Glacier Deep Archive
AnswersC, E

S3 provides 99.999999999% durability and low cost for infrequently accessed data.

Why this answer

Amazon S3 is correct because it provides 99.999999999% (11 9's) durability, is designed for write-once-read-many (WORM) workloads, and offers low-cost storage tiers suitable for IoT sensor data that is never updated. S3's object storage model and lifecycle policies allow automatic transition to colder storage, making it ideal for immutable data at scale.

Exam trap

The trap here is that candidates often choose DynamoDB (D) for its scalability and low latency, overlooking that the question emphasizes low cost and write-once immutability, where S3 and Glacier Deep Archive are orders of magnitude cheaper per GB stored.

517
MCQhard

A data engineer is investigating intermittent failures in an AWS Step Functions state machine that orchestrates a nightly ETL workflow. The state machine invokes an AWS Glue job, then an Amazon EMR step, then an AWS Lambda function. Occasionally a task fails transiently and the entire workflow stops instead of retrying. The engineer needs the workflow to automatically retry failed tasks with exponential backoff before alerting. What should the engineer do?

A.Set the state machine's execution role to include the 'states:Retry' IAM action and re-run the workflow.
B.Configure the Glue job, EMR step, and Lambda function each with their own internal retry logic and remove error handling from the state machine.
C.Add a Retry field to the relevant state with ErrorEquals, IntervalSeconds, MaxAttempts, and BackoffRate values.
D.Enable the state machine's 'Retry on failure' setting in the Amazon CloudWatch console.
AnswerC

Step Functions supports a Retry field on any state, where you specify which errors to match and how many times to retry with an interval and a backoff multiplier. Adding Retry with ErrorEquals, IntervalSeconds, MaxAttempts, and BackoffRate implements automatic exponential-backoff retries within the state machine, satisfying the requirement without external tooling.

Why this answer

Step Functions state machines are defined in Amazon States Language, and each state can include a Retry field that matches specific error names and retries with an interval that grows by a BackoffRate multiplier. This declarative approach centralizes retry logic in the orchestrator, so transient failures in Glue, EMR, or Lambda are retried automatically before any alert is raised.

Exam trap

The trap here is assuming the retry policy lives in IAM or CloudWatch rather than being declared inside the state machine definition itself.

518
MCQhard

A data engineer needs to grant an IAM role used by an AWS Lambda function permission to read encrypted data from an Amazon S3 bucket. The data is encrypted with a customer-managed AWS KMS key. The engineer wants to follow the principle of least privilege. Which combination of actions should the engineer include in the IAM policy for the Lambda execution role?

A.s3:GetObject on the bucket and kms:Decrypt on the KMS key.
B.s3:GetObject on the bucket and kms:GenerateDataKey on the KMS key.
C.s3:GetObject on the bucket and kms:CreateGrant on the KMS key.
D.s3:GetObject on the bucket and kms:Encrypt on the KMS key.
AnswerA

To read an encrypted object from S3, the Lambda function needs s3:GetObject permission on the object and kms:Decrypt permission on the KMS key used to encrypt the object. Without kms:Decrypt, S3 cannot decrypt the object for the caller. This combination follows least privilege by granting only the necessary actions for reading the specific data.

Why this answer

The Lambda function requires s3:GetObject to retrieve the object and kms:Decrypt to decrypt the data key used by S3 server-side encryption with KMS. These permissions are the minimal set needed to read the encrypted object. Other KMS actions like Encrypt, GenerateDataKey, or CreateGrant do not provide decryption capability and would not allow the function to read the data.

Exam trap

The trap here is confusing KMS actions, such as using kms:Encrypt or kms:GenerateDataKey instead of kms:Decrypt for reading encrypted data.

519
MCQhard

A company uses Amazon Kinesis Data Firehose to deliver data to an Amazon S3 bucket. The data is in JSON format and contains a 'timestamp' field with a Unix epoch value. The company wants to partition the S3 objects by year, month, day, and hour based on the timestamp. What is the MOST efficient method to achieve this?

A.Use the dynamic partitioning feature of Kinesis Data Firehose with inline parsing to extract the timestamp and create the S3 prefix.
B.Configure a custom S3 prefix in Firehose using the 'YYYY/MM/dd/HH' format based on the current time.
C.Use an AWS Glue ETL job to read from Firehose, partition, and write to S3.
D.Use Amazon Athena to run a CTAS query that partitions the data by timestamp.
AnswerA

Dynamic partitioning with inline parsing extracts the timestamp field directly within Firehose and derives year/month/day/hour prefixes automatically, avoiding custom Lambda transformation or downstream reprocessing. This satisfies the requirement for the most efficient, serverless partitioning method without managing extra compute.

Why this answer

Kinesis Data Firehose dynamic partitioning with inline parsing extracts the 'timestamp' field from each JSON record and uses it to build the S3 prefix (year/month/day/hour) automatically as records are delivered. This is the native, serverless, most efficient approach because Firehose handles the partitioning logic per-record without any additional compute services. It also supports JQ expressions for extracting and formatting the timestamp into the desired prefix pattern.

Exam trap

The trap here is confusing Firehose's static custom prefix (which uses delivery time) with dynamic partitioning (which uses record content), causing candidates to pick option B as a simpler solution.

How to eliminate wrong answers

Option B is wrong because a custom S3 prefix in Firehose uses the delivery time (current time), not the timestamp inside each record, so records would be grouped by arrival time rather than event time. Option C is wrong because an AWS Glue ETL job adds unnecessary cost, latency, and operational overhead — Firehose can partition natively without a separate ETL pipeline. Option D is wrong because Athena CTAS is a query-time operation that reads existing data and writes new partitioned data; it does not partition data as it is streamed into S3 and would require re-running on new data.

520
MCQeasy

A company wants to ingest real-time data from a social media API into Amazon S3 for analysis. The API provides data as JSON records. Which AWS service is best suited for this ingestion?

A.AWS Glue
B.Amazon Kinesis Data Firehose
C.Amazon Simple Queue Service (SQS)
D.Amazon DataZone
AnswerB

Amazon Kinesis Data Firehose is a fully managed delivery service that ingests streaming JSON records and loads them directly into Amazon S3, handling buffering, batching and scaling. It satisfies the real-time ingestion-to-S3 requirement without custom consumers, unlike Kinesis Data Streams, which requires separate delivery logic.

Why this answer

Amazon Kinesis Data Firehose is the best choice because it is a fully managed service designed to ingest real-time streaming data, such as JSON records from a social media API, and automatically load it into Amazon S3 with optional data transformation and compression. It handles scaling, buffering, and delivery without requiring custom code or infrastructure management, making it ideal for this use case.

Exam trap

The trap here is that candidates often confuse Amazon Kinesis Data Streams (which requires custom consumers) with Kinesis Data Firehose (which is serverless and directly writes to S3), or they incorrectly assume SQS can directly deliver to S3 without additional processing.

How to eliminate wrong answers

Option A is wrong because AWS Glue is a serverless data integration service for batch ETL (extract, transform, load) jobs and cataloging, not designed for real-time streaming ingestion from an API. Option C is wrong because Amazon Simple Queue Service (SQS) is a message queue for decoupling application components, but it does not natively write data to S3; you would need additional compute to poll and deliver messages, adding complexity. Option D is wrong because Amazon DataZone is a data governance and catalog service for managing data assets across an organization, not a data ingestion service for real-time streaming.

521
MCQeasy

A data engineer is using AWS Glue to process data stored in Amazon S3. The security team requires that all data in transit between AWS Glue and Amazon S3 be encrypted using TLS. The engineer has verified that the Glue job uses the AWS Glue Data Catalog and that the S3 bucket policy allows access. What should the engineer do to ensure that the Glue job enforces TLS when reading from and writing to S3?

A.Attach an IAM policy to the Glue job role that denies requests where the aws:SecureTransport condition is false.
B.Configure the Glue job to use the --enable-tls parameter.
C.Enable S3 default encryption on the bucket to enforce TLS.
D.Use an S3 VPC endpoint with a policy that allows only TLS connections.
AnswerA

By attaching an IAM policy that denies S3 actions when aws:SecureTransport is false, any request made over HTTP (non-TLS) will be denied. This enforces TLS for all S3 access, including from AWS Glue. This is a best practice for ensuring encryption in transit. The Glue job will then only be able to access S3 over HTTPS, meeting the security requirement.

Why this answer

To enforce TLS for AWS Glue access to S3, an IAM policy denying requests when aws:SecureTransport is false is the most direct method. This condition key evaluates to true for HTTPS requests and false for HTTP. By denying when false, all non-TLS requests are blocked.

This applies to any principal using the role, including AWS Glue, ensuring encryption in transit.

Exam trap

The trap here is thinking that S3 default encryption or a non-existent Glue parameter enforces TLS, when the correct approach is an IAM policy condition on aws:SecureTransport.

522
MCQeasy

A data engineer is using Amazon Athena to query data stored in Amazon S3. The engineer notices that queries are returning incorrect results, specifically missing some rows that are known to exist in the underlying data. The data is stored in Parquet format and is partitioned by date. The engineer runs a query with a WHERE clause on the date partition and finds that some dates are missing from the results. The S3 bucket contains folders for each date, but some folders are empty. What is the MOST likely cause of the missing rows?

A.Athena is using a stale metadata cache and needs to be refreshed.
B.The Parquet files are corrupted, causing Athena to skip them.
C.The Athena table is not configured with the correct partition projection settings.
D.The empty folders in S3 indicate that the data for those dates was never written or was deleted, so Athena correctly returns no rows for those dates.
AnswerD

If the S3 folders for certain dates are empty, there is no data for Athena to query. Athena reads files from S3; empty folders contain no files, so queries for those dates return no rows. This is expected behavior. The missing rows are due to missing data files, not an Athena configuration issue. The engineer should investigate why those folders are empty, possibly due to upstream ETL failures.

Why this answer

The missing rows correspond to dates for which the S3 folders are empty. Athena queries data directly from S3, so if there are no files in a partition folder, no rows will be returned for that partition. This is not an Athena misconfiguration; it is a data availability issue.

The engineer should check the upstream processes that write data to those partitions.

Exam trap

The trap here is assuming that Athena is malfunctioning when the underlying data is simply absent, leading to unnecessary troubleshooting of Athena settings.

523
MCQeasy

A company needs to migrate an on-premises 10 TB PostgreSQL database to Amazon RDS for PostgreSQL with minimal downtime. Which AWS service should be used for the migration?

A.AWS Storage Gateway
B.AWS Snowball Edge
C.AWS DataSync
D.AWS Database Migration Service (DMS)
AnswerD

DMS supports continuous replication.

Why this answer

AWS Database Migration Service (DMS) is the correct choice because it is specifically designed for migrating databases to AWS with minimal downtime. It supports continuous replication from an on-premises PostgreSQL source to Amazon RDS for PostgreSQL using change data capture (CDC), allowing the source database to remain operational during the migration.

Exam trap

The trap here is that candidates may confuse data transfer services (like Snowball Edge or DataSync) with database migration tools, overlooking that DMS is the only AWS service that supports live, ongoing replication and schema conversion for relational databases like PostgreSQL.

How to eliminate wrong answers

Option A is wrong because AWS Storage Gateway is a hybrid storage service for on-premises access to cloud storage, not a database migration tool; it cannot perform schema conversion or ongoing replication for a PostgreSQL database. Option B is wrong because AWS Snowball Edge is a physical data transport device for large-scale data transfers, but it does not support live database replication or CDC, making it unsuitable for minimal-downtime migrations of a live 10 TB database. Option C is wrong because AWS DataSync is designed for moving large amounts of file data over the network, not for database-level migrations; it lacks the ability to handle PostgreSQL-specific objects, transactions, or ongoing replication.

524
Multi-Selecthard

A company runs a data processing pipeline using Amazon EMR with Spark. The pipeline reads from S3, processes data, and writes to S3. Recently, the job started failing with 'S3AccessDeniedException' even though the EMR role has appropriate S3 permissions. Which TWO actions should the data engineer take to resolve this issue? (Choose TWO.)

Select 2 answers
A.Enable S3 versioning on the bucket to allow multiple access methods.
B.Verify that the EMR service role has the necessary S3 permissions in IAM.
C.Disable S3 Block Public Access settings on the bucket.
D.Check the S3 bucket policy for explicit deny statements that may override the IAM role.
E.Ensure the EMR cluster is launched in a VPC with an S3 VPC endpoint.
AnswersB, D

The EMR service role is the identity Spark uses to call S3, so its attached IAM policy must explicitly allow the required s3:GetObject and s3:PutObject actions on the bucket and prefix. Confirming these permissions rules out an identity-based policy gap before investigating bucket-level controls.

Why this answer

Option B is correct because even when the EMR role appears to have S3 permissions, the actual attached IAM service role (EMR_DefaultRole or a custom EC2 instance profile for EMR) must explicitly allow the required s3:GetObject, s3:PutObject, s3:ListBucket, and related actions on the specific bucket and object ARNs; a missing or mis-scoped policy on the role produces exactly the S3AccessDeniedException seen here. Option D is correct because an explicit Deny in the S3 bucket policy always overrides any Allow granted through the IAM role, so a deny statement (for example, one enforcing TLS, VPC endpoint, or encryption conditions) would cause access failures despite the role having appropriate permissions. Option A is incorrect because S3 versioning only affects object version retention and does not grant or alter access permissions.

Option C is incorrect because disabling Block Public Access would only matter for public/anonymous access and would not fix a role-based access denial; it would also weaken security. Option E is incorrect because a missing S3 VPC endpoint causes connectivity/timeout issues, not an S3AccessDeniedException, and the error indicates the request reached S3 and was explicitly denied by permissions.

525
MCQmedium

A data engineer is using AWS Step Functions to orchestrate a complex ETL workflow that includes multiple AWS Glue jobs, Amazon EMR steps, and AWS Lambda functions. The engineer notices that on rare occasions, the entire workflow fails due to a transient error in one of the Lambda functions. The engineer wants to make the workflow more resilient without changing the overall architecture. Which approach is the MOST effective?

A.Enable AWS X-Ray tracing for the Lambda function to identify the root cause of the errors.
B.Modify the Lambda function code to catch exceptions and return a success status to avoid workflow failure.
C.Add a Retry state in the Step Functions state machine for the Lambda task with exponential backoff and a maximum number of attempts.
D.Configure the Lambda function to have a longer timeout and increase its memory size.
AnswerC

Adding a Retry state in Step Functions allows the workflow to automatically retry the Lambda function when it fails due to transient errors. Configuring exponential backoff and a maximum attempt count ensures that temporary issues are handled gracefully without manual intervention. This directly improves the workflow's resilience and is a best practice for handling transient failures.

Why this answer

Adding a Retry state in Step Functions with exponential backoff and a maximum attempts limit is the most effective way to handle transient errors in Lambda functions. It automatically retries the failed task, increasing the likelihood of success without manual intervention, and is a standard resilience pattern in Step Functions.

Exam trap

The trap here is thinking that increasing Lambda resources or enabling tracing will resolve transient errors, when the real solution is to implement automated retries.

Page 6

Page 7 of 18

Page 8