Courseiva

AWS Certified Data Engineer Associate DEA-C01 (DEA-C01) — Questions 826–900

1321 questions total · 18pages · All types, answers revealed

Page 11

Page 12 of 18

Page 13
826
Multi-Selecthard

A company runs a data lake on Amazon S3 with AWS Glue and Amazon Athena. The data engineer notices that queries are slow and scanning large amounts of data. Which THREE actions should the engineer take to optimize query performance and reduce costs?

Select 3 answers
A.Increase the query timeout in Athena.
B.Increase the number of DPUs in the Glue job.
C.Compress data files using gzip or snappy.
D.Partition the data by frequently filtered columns (e.g., date, region).
E.Use columnar data formats like Parquet or ORC.
AnswersC, D, E

Compression shrinks file sizes on S3, so Athena scans fewer bytes per query, cutting both runtime and per-terabyte scan costs. This directly addresses the large-data-scanning problem in the stem while remaining readable by Glue and Athena.

Why this answer

Option C is correct because compressing data files with gzip or snappy reduces the total bytes stored and scanned, and Athena charges and performs based on data scanned, so smaller files lower both query latency and cost. Option D is correct because partitioning the S3 data by frequently filtered columns such as date or region enables partition pruning, so Athena reads only the relevant prefixes instead of scanning the entire table. Option E is correct because columnar formats like Parquet or ORC let Athena read only the columns referenced in the query and provide better compression and predicate pushdown, dramatically reducing scanned data compared to row-based formats like CSV or JSON.

Option A is not appropriate because increasing the Athena query timeout only allows long-running queries to finish; it does not reduce the amount of data scanned or improve performance. Option B is not appropriate because adding DPUs to the Glue job speeds up ETL processing, not Athena query performance or the volume of data scanned at query time.

Exam trap

The trap is selecting scaling options (timeout, DPUs) instead of data organization techniques; the exam tests that Athena performance is primarily about reducing data scanned via partitioning, columnar formats, and compression.

827
MCQeasy

A company runs a data pipeline that uses AWS Lambda to process files uploaded to an S3 bucket. Recently, some files have been processed multiple times. The Lambda function is triggered by S3 event notifications. What is the MOST likely cause of duplicate processing?

A.The Lambda function has a high error rate and retries.
B.The Lambda function is not idempotent.
C.The Lambda function has a reserved concurrency setting.
D.S3 event notifications are delivered at least once.
AnswerD

Amazon S3 event notifications provide at-least-once delivery, so the same event can be published more than once, invoking Lambda repeatedly for one uploaded object. This inherent duplication, not Lambda retries or concurrency, explains the repeated file processing.

Why this answer

S3 event notifications are delivered using an at-least-once delivery model, meaning the same event can be delivered more than once, which causes Lambda to be invoked multiple times for the same object. This is the most likely cause of duplicate processing when the function itself is not idempotent. The root cause is the delivery guarantee, not the function's error handling.

Exam trap

The trap is choosing 'the function is not idempotent' as the cause — that is the reason duplicates cause damage, but the actual cause of duplicate invocations is S3's at-least-once event delivery.

How to eliminate wrong answers

Option A is wrong because a high error rate with retries would cause retries only on failures, but the question states files are processed multiple times — this can happen even on successful invocations due to at-least-once delivery. Option B is wrong because non-idempotency is a contributing factor that makes duplicates harmful, but it is not the cause of duplicate invocations; the cause is the delivery model. Option C is wrong because reserved concurrency controls how many concurrent executions are allowed, not whether the same event is delivered more than once.

828
MCQeasy

A company needs to ingest data from multiple SaaS sources (e.g., Salesforce, Marketo) into Amazon S3 for analytics. Which AWS service is designed for this purpose?

A.AWS Transfer Family
B.AWS Glue
C.Amazon AppFlow
D.AWS DataSync
AnswerC

Amazon AppFlow provides managed connectors for SaaS sources such as Salesforce and Marketo, transferring data directly into Amazon S3 without custom code. This satisfies the stem's requirement to ingest from multiple SaaS applications into S3 for analytics.

Why this answer

Amazon AppFlow is a fully managed integration service specifically designed to securely transfer data between SaaS applications (like Salesforce, Marketo, Slack, and Zendesk) and AWS services such as Amazon S3 and Amazon Redshift. It supports scheduled, event-driven, or on-demand data ingestion with built-in transformations, filtering, and validation, making it the ideal choice for this use case.

Exam trap

The trap here is that candidates often confuse AWS Glue's ETL capabilities with native SaaS connectivity, but Glue requires custom connectors or AWS Glue Studio's visual ETL jobs to connect to SaaS sources, whereas AppFlow is purpose-built for this task.

How to eliminate wrong answers

Option A is wrong because AWS Transfer Family is used for transferring files into and out of Amazon S3 or Amazon EFS using SFTP, FTPS, or FTP protocols, not for integrating with SaaS APIs. Option B is wrong because AWS Glue is a serverless data integration service for ETL (extract, transform, load) jobs, but it does not natively connect to SaaS sources like Salesforce or Marketo without custom connectors or third-party libraries. Option D is wrong because AWS DataSync is designed for moving large volumes of data between on-premises storage and AWS (e.g., NFS, SMB, S3) or between AWS storage services, not for ingesting data from SaaS applications.

829
MCQeasy

A data engineer needs to ensure that an Amazon Redshift cluster only accepts encrypted connections. Which parameter should be modified?

A.enable_user_activity_logging
B.max_concurrency_scaling_clusters
C.require_SSL
D.wlm_json_configuration
AnswerC

The require_SSL parameter in the Redshift cluster's parameter group enforces TLS, rejecting unencrypted client connections. Modifying it satisfies the constraint that the cluster accept only encrypted connections, whereas other parameters govern query behaviour or logging rather than transport encryption.

Why this answer

Setting the `require_SSL` parameter to `true` forces all connections to the Amazon Redshift cluster to use SSL/TLS encryption, ensuring that data in transit is encrypted. This parameter is modified in the cluster's parameter group and applies to both JDBC and ODBC connections, as well as the Redshift Query Editor.

Exam trap

The trap here is that candidates may confuse `require_SSL` with other security-related parameters like `enable_user_activity_logging` (auditing) or assume that encryption is handled by a different mechanism (e.g., WLM or concurrency scaling), leading them to pick a wrong option that sounds security-adjacent but is technically unrelated.

How to eliminate wrong answers

Option A is wrong because `enable_user_activity_logging` controls the logging of user activity (e.g., queries run by users) for auditing purposes, not connection encryption. Option B is wrong because `max_concurrency_scaling_clusters` defines the maximum number of concurrency scaling clusters that can be used to handle spikes in concurrent queries, unrelated to encryption. Option D is wrong because `wlm_json_configuration` defines workload management (WLM) queue configurations (e.g., concurrency, memory allocation) and has no effect on SSL/TLS enforcement.

830
MCQhard

A data engineer is managing an Amazon Redshift cluster that experiences performance degradation during peak hours. The cluster uses a single node type and has no concurrency scaling enabled. Analysis shows that many long-running queries are queued behind short ad-hoc queries, causing delays for critical reports. The engineer needs to ensure that critical reports run promptly without affecting ad-hoc queries. Which solution meets these requirements?

A.Create a workload management (WLM) queue with a higher priority for critical reports and assign them to that queue.
B.Use short query acceleration (SQA) to run short queries in a dedicated space.
C.Increase the cluster size by adding more nodes to reduce overall query time.
D.Enable concurrency scaling to automatically add transient clusters for queued queries.
AnswerA

Amazon Redshift WLM allows you to create separate queues with different priorities. By assigning critical reports to a high-priority queue, they are scheduled before queries in lower-priority queues. This ensures that critical reports run promptly even when ad-hoc queries are present. This is the most direct way to prioritize specific workloads without affecting others.

Why this answer

Amazon Redshift workload management (WLM) enables you to define queues with different priorities. By routing critical reports to a high-priority queue, they are scheduled ahead of ad-hoc queries, ensuring timely execution. This approach directly addresses the need to prioritize specific workloads without impacting others.

Exam trap

The trap here is confusing concurrency scaling or SQA with workload prioritization; these features improve throughput or short query performance but do not enforce priority for specific queries.

831
MCQeasy

A data engineer needs to set up a disaster recovery solution for an Amazon RDS for MySQL database. The database must be available in another AWS Region with minimal data loss. What is the simplest approach?

A.Enable Multi-AZ deployment in the same Region.
B.Set up AWS Database Migration Service (DMS) for continuous replication.
C.Take a manual snapshot and copy it to the other Region daily.
D.Create a cross-Region read replica of the database.
AnswerD

A cross-Region read replica asynchronously replicates the primary RDS for MySQL database to another Region, providing a standby that can be promoted during disaster. This delivers cross-Region availability with minimal data loss and is simpler than manual snapshot copying or restore procedures.

Why this answer

A cross-Region read replica of an RDS for MySQL database asynchronously replicates data to another AWS Region, providing a warm standby with minimal data loss (typically seconds of lag) and the ability to promote the replica during a regional failure. It is the simplest native RDS feature for cross-Region DR without building custom replication pipelines.

Exam trap

DEA-C01 often tests whether candidates pick Multi-AZ for cross-Region DR or daily snapshots for minimal data loss, confusing availability within a Region with disaster recovery across Regions.

How to eliminate wrong answers

Option A is wrong because Multi-AZ keeps the standby in the same Region, so it does not protect against a regional outage and does not meet the cross-Region requirement. Option B is wrong because DMS continuous replication is more complex to set up and manage than a native cross-Region read replica, and it is typically used for heterogeneous migrations or ongoing replication into analytics systems. Option C is wrong because daily manual snapshots copied to another Region have an RPO of up to 24 hours, which does not meet the 'minimal data loss' requirement.

832
MCQhard

A data engineer is using Amazon Redshift and needs to load data from Amazon S3 into a table. The data is in Parquet format and is partitioned by date. The engineer wants to load only the latest partition and ensure that the load is efficient and cost-effective. Which method should the engineer use?

A.Use AWS Glue to read the Parquet data, filter by the latest date, and write to Redshift.
B.Use the COPY command with the FROM 's3://bucket/prefix/date=2023-10-01/' option to load only that partition.
C.Use Amazon Redshift Spectrum to create an external table and query the latest partition.
D.Use the COPY command with the FROM 's3://bucket/prefix' option and specify the partition column in the WHERE clause.
AnswerB

The COPY command can load data from a specific S3 prefix. By specifying the exact prefix for the desired partition, the engineer loads only that partition's files. This is efficient and cost-effective because it avoids scanning and loading unnecessary data. Parquet format is natively supported, and the load will be fast and compressed.

Why this answer

The COPY command can load data directly from a specific S3 prefix, so specifying the exact partition prefix loads only that partition efficiently. This avoids unnecessary data transfer and cost. Using a WHERE clause with COPY is invalid, and alternatives like AWS Glue or Redshift Spectrum add complexity or do not load data into Redshift tables as required.

Exam trap

The trap here is assuming that the COPY command supports filtering with a WHERE clause, when in fact you must specify the exact S3 prefix to load a subset of data.

833
Multi-Selecteasy

A company is building a data lake on Amazon S3 and needs to ingest data from various on-premises sources. Which TWO AWS services can be used to transfer data securely over the internet?

Select 2 answers
A.AWS Snowcone
B.AWS DataSync
C.Amazon Kinesis Data Firehose
D.AWS Direct Connect
E.AWS CLI
AnswersB, E

AWS DataSync transfers files over the internet using TLS encryption, satisfying the secure on-premises ingestion requirement. It supports NFS, SMB, and object storage sources, with bandwidth throttling and integrity verification. Agents deploy on-premises, replicating data directly to Amazon S3 without VPN or AWS Direct Connect, which the stem does not mandate.

Why this answer

AWS DataSync (B) is correct because it is a purpose-built data transfer service that moves data between on-premises storage and Amazon S3 over the internet using TLS encryption, and it supports scheduled, incremental, and bandwidth-limited transfers. AWS CLI (E) is correct because commands such as aws s3 cp or aws s3 sync transfer data from on-premises systems to S3 over the public internet using HTTPS, making it a valid secure internet-based ingestion method. AWS Snowcone (A) is not correct here because it is an edge computing and offline/physical data transfer device, not an internet-based transfer service.

Amazon Kinesis Data Firehose (C) is not correct because it delivers streaming data to destinations like S3 but is not designed to pull files from on-premises sources over the internet. AWS Direct Connect (D) is not correct because it uses a dedicated private network connection rather than the public internet.

Exam trap

The trap here is that candidates often confuse AWS DataSync with AWS Direct Connect, thinking both are required for secure transfers, but Direct Connect is a network service that bypasses the internet entirely, while DataSync is a data transfer service that works over the internet or Direct Connect.

834
Multi-Selectmedium

A company is building a data lake on Amazon S3. They need to ingest data from multiple sources, including relational databases, streaming data, and log files. Which THREE AWS services can be used to ingest data into the data lake?

Select 3 answers
A.Amazon Kinesis Data Firehose
B.AWS Database Migration Service (DMS)
C.Amazon Athena
D.Amazon Redshift Spectrum
E.AWS Glue
AnswersA, B, E

Ingests streaming data into S3.

Why this answer

Amazon Kinesis Data Firehose is a fully managed service for streaming data ingestion that can capture, transform, and load streaming data into Amazon S3 in near real-time. It supports sources like Amazon CloudWatch Logs, AWS IoT, and custom producers via the Kinesis Agent, making it ideal for log files and streaming data.

Exam trap

The trap here is confusing query engines (Athena, Redshift Spectrum) with ingestion services, as candidates often assume any service that touches S3 can be used for data loading.

835
MCQeasy

A data engineer is setting up an Amazon RDS for MySQL database. The compliance team requires that all data at rest be encrypted. What must the engineer do to enable encryption for this database?

A.Specify an AWS KMS key when launching the DB instance
B.Enable encryption after the DB instance is created by modifying the DB instance
C.Use AWS Secrets Manager to store the encryption key and attach it to the DB instance
D.Encrypt the underlying EBS volumes after the instance is created
AnswerA

Specifying a customer-managed AWS KMS key at DB instance creation enables encryption at rest for the underlying storage, volumes, snapshots and read replicas, satisfying the compliance requirement. Encryption cannot be enabled retroactively on an existing unencrypted instance; it must be set at launch.

Why this answer

For Amazon RDS, encryption at rest must be enabled at the time of DB instance creation by specifying an AWS KMS key. Once the DB instance is created, you cannot enable encryption by simply modifying the instance; you would need to create an encrypted snapshot and restore it to a new encrypted instance. Therefore, the engineer must specify a KMS key when launching the DB instance to meet the compliance requirement.

Exam trap

DEA-C01 often tests the misconception that encryption can be enabled on an existing RDS instance by modifying it, when in fact it must be set at creation time or via snapshot restore.

How to eliminate wrong answers

Option B is wrong because RDS does not allow enabling encryption on an existing unencrypted DB instance via modification; the only way is to create a snapshot, encrypt it, and restore to a new instance. Option C is wrong because AWS Secrets Manager is used for storing and rotating credentials, not for managing encryption keys for RDS storage encryption. Option D is wrong because you cannot directly encrypt the underlying EBS volumes of an RDS instance; RDS manages the storage and encryption is handled at the service level, not by manual EBS encryption.

836
MCQhard

A data engineer needs to set up a new Amazon RDS for PostgreSQL database for a production workload. The database must be highly available and resilient to a single Availability Zone failure. Which configuration should the engineer choose?

A.Single-AZ with automated backups
B.Multi-AZ deployment with one standby in a different AZ
C.Multi-AZ with two readable standbys
D.Single-AZ with a read replica
AnswerB

Multi-AZ maintains a synchronous standby replica in a separate Availability Zone, so RDS automatically fails over to it if the primary AZ fails. This directly satisfies the resilience-to-single-AZ-failure constraint, unlike a single-AZ instance or read replicas, which use asynchronous replication and do not provide automatic failover.

Why this answer

A Multi-AZ deployment for Amazon RDS PostgreSQL automatically provisions and maintains a synchronous standby replica in a different Availability Zone. This configuration provides automatic failover in the event of an AZ failure, ensuring high availability and resilience without manual intervention. The synchronous replication ensures zero data loss during failover, which is critical for production workloads.

Exam trap

The trap here is that candidates often confuse Multi-AZ with read replicas, assuming that a read replica can serve as a failover target, but in RDS PostgreSQL, read replicas are asynchronous and require manual promotion, making them unsuitable for automatic high availability against AZ failures.

How to eliminate wrong answers

Option A is wrong because a Single-AZ deployment with automated backups only protects against data loss via point-in-time recovery, but does not provide automatic failover or resilience to an AZ failure; the database becomes unavailable if the AZ goes down. Option C is wrong because Amazon RDS for PostgreSQL does not support Multi-AZ with two readable standbys; that feature is specific to Amazon RDS for Oracle and SQL Server Enterprise Edition, and PostgreSQL Multi-AZ only provides a single standby that is not readable. Option D is wrong because a Single-AZ with a read replica provides read scaling and some disaster recovery capability, but the read replica is asynchronous and does not provide automatic failover; a manual promotion is required, and the primary remains vulnerable to AZ failure.

837
MCQeasy

A data engineer must load a 50 GB uncompressed CSV file from Amazon S3 into an Amazon Redshift cluster using the COPY command. The load is taking a long time and the engineer wants to improve performance. Which action should the engineer take?

A.Convert the CSV file to JSON and load it with the COPY command using the JSON 'auto' option.
B.Compress the file with gzip and split it into multiple smaller files, then run the COPY command with the GZIP option.
C.Use the COPY command with the PARALLEL OFF option to force Redshift to distribute the load across all slices.
D.Increase the cluster's node count and rerun the COPY command on the same single uncompressed file.
AnswerB

Compressing the data reduces the bytes transferred from Amazon S3 and splitting the file into multiple parts lets Redshift load slices in parallel across the cluster. The COPY command supports the GZIP parameter to decompress on ingest. Together these changes dramatically reduce load time compared with a single large uncompressed file.

Why this answer

Redshift COPY performance depends on parallelizing reads across slices and minimizing bytes transferred. Compressing the CSV with gzip reduces transfer volume, and splitting it into multiple files lets each slice load a portion concurrently. The GZIP parameter handles decompression during ingest.

A single large uncompressed file limits parallelism regardless of cluster size.

Exam trap

The trap here is believing that adding cluster nodes automatically speeds up loading a single large file.

838
MCQeasy

A data engineer needs to transfer 10 TB of data from an on-premises data center to Amazon S3. The network bandwidth is limited to 100 Mbps, and the data transfer must be completed within 5 days. What is the most cost-effective solution?

A.Use AWS Snowball Edge to physically ship the data.
B.Use S3 Transfer Acceleration to speed up the transfer over the internet.
C.Use AWS DataSync over the internet to transfer the data.
D.Set up an AWS Direct Connect connection to increase bandwidth.
AnswerA

At 100 Mbps, transferring 10 TB over the network would take roughly ten days, exceeding the five-day deadline. AWS Snowball Edge ships data physically, bypassing bandwidth limits and completing well within the window at lower cost.

Why this answer

With 10 TB of data and a 100 Mbps link, the theoretical transfer time over the internet is approximately 10 days (10 TB * 8 / 100 Mbps = 800,000 seconds ≈ 9.26 days), which exceeds the 5-day requirement. AWS Snowball Edge is the most cost-effective solution because it bypasses the network bottleneck entirely by physically shipping the data, and it is designed for large-scale data transfers where network constraints make online transfer impractical.

Exam trap

The trap here is that candidates assume S3 Transfer Acceleration or DataSync can magically overcome bandwidth limitations, but they only optimize the path, not increase the pipe size, so the math of bandwidth vs. data volume always dictates the minimum transfer time.

How to eliminate wrong answers

Option B is wrong because S3 Transfer Acceleration only optimizes the network path using AWS edge locations and does not increase the available bandwidth; it cannot overcome the fundamental 100 Mbps bottleneck, so the transfer would still take over 9 days. Option C is wrong because AWS DataSync over the internet is still limited by the 100 Mbps bandwidth, and even with optimization, it cannot complete 10 TB within 5 days. Option D is wrong because setting up AWS Direct Connect requires significant upfront cost and provisioning time (often weeks), making it neither cost-effective nor timely for a one-time transfer within 5 days.

839
MCQmedium

A data engineer is using Amazon Kinesis Data Firehose to deliver streaming data to an Amazon S3 bucket. The engineer notices that some records are being delivered to S3 with a delay of several minutes, and sometimes records are missing. The Firehose stream is configured with a buffer size of 5 MB and a buffer interval of 300 seconds. The engineer wants to reduce latency and ensure all records are delivered. Which action should the engineer take?

A.Reduce the buffer interval to 60 seconds and enable Amazon CloudWatch Logs for error logging.
B.Increase the buffer size to 128 MB to allow more data to be batched.
C.Enable Amazon S3 versioning on the destination bucket.
D.Configure the Firehose stream to use AWS Lambda for data transformation.
AnswerA

Reducing the buffer interval lowers the maximum time records wait before delivery, decreasing latency. Enabling CloudWatch Logs allows you to capture delivery errors and troubleshoot missing records. Firehose can log errors to CloudWatch, which helps identify why records are not delivered, such as permission issues or malformed records.

Why this answer

The buffer interval determines how long Firehose waits before delivering data. A 300-second interval means records can be delayed up to 5 minutes. Reducing it to 60 seconds lowers latency.

Missing records often indicate delivery errors, which can be diagnosed by enabling CloudWatch Logs for the Firehose stream. Firehose logs errors related to S3 delivery, Lambda processing, and format conversion, helping identify and fix issues.

Exam trap

The trap here is thinking that increasing buffer size improves latency, when actually it increases the time data waits to be flushed.

840
MCQhard

A company is using Amazon S3 to store sensitive data. The security team requires that all data be encrypted at rest using a customer-managed AWS KMS key. The data engineer must ensure that only a specific IAM role can decrypt the data. Which policy should the data engineer attach to the KMS key?

A.A KMS key policy that allows the IAM role to perform kms:Decrypt
B.An IAM user policy that allows kms:Decrypt for the specific key
C.An IAM policy attached to the role that allows kms:Decrypt
D.An S3 bucket policy that denies access unless encryption is used
AnswerA

KMS key policies grant permissions to use the key.

Why this answer

KMS key policies are the primary mechanism for controlling access to a customer-managed KMS key. By specifying the IAM role as a principal in the key policy and granting kms:Decrypt, you ensure that only that role can decrypt data encrypted with this key, regardless of any IAM policies that might otherwise allow broader access.

Exam trap

The DEA-C01 exam often tests the misconception that IAM policies alone can control KMS key access, but the correct approach is to use a KMS key policy that explicitly grants the required action to the specific principal.

How to eliminate wrong answers

Option B is wrong because an IAM user policy alone is insufficient; KMS key access requires either a key policy that explicitly grants permissions to the user/role or a grant, and IAM policies only take effect if the key policy allows IAM policy-based access (via a root principal). Option C is wrong because while an IAM policy attached to the role can allow kms:Decrypt, it will only work if the KMS key policy also permits IAM policy-based access (e.g., by allowing the root account), which is not guaranteed and does not restrict decryption to that specific role as tightly as a key policy. Option D is wrong because an S3 bucket policy that denies access unless encryption is used does not control who can decrypt data; it only enforces encryption in transit or at rest, and does not restrict decryption permissions to a specific IAM role.

841
MCQhard

A data engineer is troubleshooting a slow Amazon Redshift query that joins a large fact table with several dimension tables. The EXPLAIN plan shows a hash join on the distribution key, but the query still runs slowly. The fact table is distributed by KEY(column_x) and the dimension tables are distributed ALL. The engineer notices that the fact table has a high number of rows with the same value in column_x. What is the most likely cause of the slow performance?

A.The fact table's distribution key column has data skew, causing uneven data distribution across nodes.
B.The dimension tables should be distributed by KEY instead of ALL.
C.The Redshift cluster does not have enough disk space.
D.The fact table does not have a sort key.
AnswerA

Distribution by KEY(column_x) with many duplicate values concentrates those rows on few slices, so some nodes process far more data than others. That skew makes the hash join uneven and dominates query runtime despite correct distribution style.

Why this answer

Data skew in the distribution key column_x causes some slices to hold a disproportionate number of rows, leading to uneven workload distribution during the hash join. The EXPLAIN plan shows a hash join on the distribution key, which should be efficient if data is evenly distributed, but skew forces the node with the most rows to become a bottleneck, slowing the entire query.

Exam trap

The trap here is that candidates often assume a hash join on the distribution key is always optimal, overlooking that data skew in the distribution key itself can negate the benefit and cause severe performance degradation.

How to eliminate wrong answers

Option B is wrong because distributing dimension tables by KEY would likely worsen performance by requiring redistribution or broadcasting during joins, whereas ALL distribution is optimal for small dimension tables to avoid data movement. Option C is wrong because insufficient disk space would manifest as disk-full errors or failed writes, not as slow query performance with a hash join plan. Option D is wrong because while a sort key can improve query performance for range-restricted scans, the EXPLAIN plan indicates the bottleneck is the hash join on the distribution key, not a missing sort key.

842
MCQhard

A company uses AWS Glue to run ETL jobs that process data from Amazon RDS to Amazon S3. The jobs run nightly and take 3 hours to complete. The data volume is growing by 20% each month. The engineer needs to reduce job runtime and cost. The source RDS is a db.r5.large instance. Which approach would be MOST effective?

A.Reduce the number of DPUs to lower cost and accept longer runtime.
B.Increase the number of Glue workers and choose a G.1X or G.2X worker type.
C.Create a read replica of the RDS instance and point the Glue job to the replica.
D.Enable S3 Transfer Acceleration on the destination bucket.
AnswerB

Glue scales horizontally, so adding workers and moving to G.1X or G.2X types increases per-worker compute and memory, cutting the three-hour runtime. This addresses the growing data volume without the cost of resizing the source RDS instance, which is not the bottleneck.

Why this answer

Increasing the number of Glue workers and choosing a G.1X or G.2X worker type directly increases parallelism and provides more memory per worker, reducing job runtime at a manageable cost increase. Option A is wrong; reducing DPUs would increase runtime, not reduce it. Option C is wrong because a read replica does not improve Glue processing speed; the bottleneck is Glue's processing capacity, not the source database's read capacity.

Option D is wrong because S3 Transfer Acceleration improves upload speed to S3, not Glue job processing.

843
MCQmedium

Refer to the exhibit. An IAM policy is attached to a user who needs to read objects from the 'example-bucket' S3 bucket. The user reports being unable to read any object under the 'confidential/' prefix. What is the reason for this access issue?

A.The allow statement is evaluated before the deny statement
B.The deny statement is missing an explicit allow for the confidential prefix
C.The explicit deny statement overrides the allow statement
D.The resource ARN in the deny statement is incorrect
AnswerC

An explicit Deny in an IAM policy always wins over any Allow, regardless of order or specificity. The deny statement covering the confidential/ prefix therefore blocks those read requests even though the allow statement grants s3:GetObject on the bucket.

Why this answer

An explicit deny statement overrides any allow statement, regardless of the order in which they appear. In this policy, there is an allow for GetObject on all objects in example-bucket, but there is an explicit deny for GetObject on the 'confidential/' prefix. Since explicit deny takes precedence, the user cannot read objects under that prefix.

Option A is incorrect because the order of evaluation does not matter; explicit deny always wins. Option B is incorrect because the deny statement does not need an explicit allow; the deny itself is effective. Option D is incorrect because the resource ARN in the deny statement is correctly specified as 'arn:aws:s3:::example-bucket/confidential/*'.

844
MCQmedium

A data engineer manages an AWS Glue job that processes JSON files from Amazon S3 and writes to Amazon Redshift. The job fails with the error "Unable to find a suitable JDBC driver". The engineer has verified that the Glue connection to Redshift is configured correctly and the IAM role has the necessary permissions. What is the most likely cause of this error?

A.The Redshift cluster is in a private subnet, and the Glue job cannot reach it due to a missing NAT gateway.
B.The Glue job's script is missing the required JDBC driver dependency, and the job's dependent jars path is not set.
C.The Glue job's IAM role lacks permissions to access the Redshift cluster's metadata.
D.The Redshift cluster's security group does not allow inbound traffic from the Glue job's ENI.
AnswerB

AWS Glue does not include the Redshift JDBC driver by default. The engineer must specify the driver's S3 path in the job's dependent jars path or include it as a job parameter. Without it, the job cannot load the driver class, causing this error. The connection and IAM role are unrelated to driver availability.

Why this answer

AWS Glue jobs require the appropriate JDBC driver to connect to Amazon Redshift. The driver is not included in the Glue environment by default, so the engineer must provide it via the dependent jars path. The error indicates the driver class is missing, which is resolved by adding the driver JAR.

Network and IAM issues would produce different errors.

Exam trap

The trap here is assuming that a configured Glue connection automatically provides the JDBC driver, when in fact the driver must be supplied separately.

845
Multi-Selectmedium

A company uses Amazon Kinesis Data Streams to ingest clickstream data. The data must be transformed and stored in Amazon S3 for batch analytics. The engineer wants to use AWS Lambda for transformation. Which TWO configurations are required? (Choose two.)

Select 2 answers
A.Configure the Lambda function to write to S3 via Kinesis Data Firehose.
B.Configure the Kinesis stream to send records directly to S3.
C.Set up an SQS queue as a destination for Lambda errors.
D.Create an event source mapping from the Kinesis stream to the Lambda function.
E.Assign an IAM role to Lambda with permissions to read from Kinesis and write to S3.
AnswersD, E

An event source mapping registers the Kinesis stream as a Lambda trigger, letting the service poll shards, batch records, and invoke the function with checkpointing. Without it, Lambda cannot consume the stream for transformation before writing to S3.

Why this answer

Option D is correct because AWS Lambda consumes records from Kinesis Data Streams through an event source mapping, which polls the stream using the enhanced fan-out or standard iterator and invokes the function with batches of records. Option E is correct because the Lambda execution role must include IAM permissions for the required actions, such as kinesis:GetRecords, kinesis:GetShardIterator, kinesis:DescribeStream, and kinesis:ListShards to read the stream, plus s3:PutObject to store the transformed data in Amazon S3. Option A is not required because Kinesis Data Firehose is a separate delivery service; the Lambda function can write directly to S3 using the AWS SDK, and Firehose is not a mandatory intermediary.

Option B is incorrect because a Kinesis data stream cannot send records directly to Amazon S3; it requires a consumer such as Lambda or Firehose. Option C is incorrect because an SQS dead-letter queue for Lambda errors is an optional reliability configuration, not a required configuration for transforming and storing the data.

Exam trap

DEA-C01 often tests the required components for Lambda-Kinesis integration, and candidates may confuse optional error handling (SQS) or alternative services (Firehose) with mandatory configurations.

846
MCQhard

A data engineer is configuring AWS Glue to crawl a dataset stored in Amazon S3 and populate the AWS Glue Data Catalog. The security team requires that all data in transit between AWS Glue and Amazon S3 be encrypted using TLS. The engineer has already configured the Glue crawler to use a connection with the appropriate VPC settings. What additional step must the engineer take to enforce encryption in transit?

A.Attach a bucket policy to the S3 bucket that denies requests where aws:SecureTransport is false.
B.Configure the Glue crawler to use an S3 endpoint with SSL enabled in the connection options.
C.Set the Glue crawler's security configuration to require SSL for S3 connections.
D.Enable server-side encryption with AWS KMS (SSE-KMS) on the S3 bucket, which automatically enforces TLS for all requests.
AnswerA

To enforce encryption in transit, the S3 bucket policy must include a condition that denies requests when aws:SecureTransport is false. This ensures that any request to the bucket, including from AWS Glue, must use TLS. AWS Glue uses HTTPS by default when accessing S3, so the policy will not block legitimate traffic but will reject unencrypted requests, satisfying the security requirement.

Why this answer

To enforce encryption in transit between AWS Glue and Amazon S3, the engineer must attach a bucket policy that denies requests when aws:SecureTransport is false. This ensures all requests, including those from Glue, use TLS. AWS Glue already uses HTTPS by default, but the bucket policy provides a hard enforcement.

Glue security configurations and connection options do not control TLS for S3, and SSE-KMS is for encryption at rest.

Exam trap

The trap here is confusing encryption at rest (SSE-KMS) with encryption in transit (TLS), and assuming that enabling SSE-KMS or a Glue security configuration will enforce TLS for S3 requests.

847
MCQeasy

A data engineer is configuring an Amazon S3 bucket that will receive raw clickstream files from a mobile application. The engineer must ensure that the objects are protected against accidental overwrites and deletions for a defined retention period, and that the protection cannot be removed or shortened by any user, including the account root user. Which S3 feature should the engineer use?

A.S3 Object Lock in compliance mode with a retention period matching the required window.
B.S3 Versioning with a lifecycle rule that transitions objects to S3 Standard-IA after 30 days.
C.A bucket policy that denies s3:DeleteObject and s3:PutObject to all principals except the ingestion role.
D.S3 default encryption with AWS KMS customer managed keys and a restrictive key policy.
AnswerA

S3 Object Lock in compliance mode prevents any user, including the account root user, from overwriting or deleting a protected object version until the retention date passes, and the retention cannot be shortened. That exactly matches the requirement for enforceable, tamper-proof retention on the incoming clickstream objects.

Why this answer

S3 Object Lock provides write-once-read-many protection by binding a retention period to object versions. In compliance mode, no principal, including the account root user, can shorten or remove that retention, which is the only option that satisfies an absolute immutability requirement. Governance mode, by contrast, allows privileged users to bypass retention.

Exam trap

The trap here is assuming that a deny-based bucket policy or versioning provides the same guarantee as Object Lock, when both can be changed or bypassed by privileged identities.

848
MCQmedium

A data engineer is using AWS Glue DataBrew to clean a dataset stored in Amazon S3. The recipe must replace all null values in a specific column with the string 'UNKNOWN' and then convert the column to uppercase. The engineer wants to apply these steps in a repeatable recipe. Which combination of DataBrew transforms should be used?

A.Use the 'Fill missing values' transform with a custom value, then the 'Upper case' transform on the column.
B.Use the 'Remove nulls' transform, then the 'Upper case' transform.
C.Use the 'Replace value' transform with a regex for null, then the 'Upper case' transform.
D.Use the 'Custom formula' transform with an IFNULL expression, then the 'Upper case' transform.
AnswerA

The 'Fill missing values' transform replaces nulls with a specified value such as 'UNKNOWN', and the 'Upper case' transform converts the column's text to uppercase. Applied in sequence within a recipe, they achieve the required cleaning and normalization. DataBrew recipes are repeatable and can be scheduled or run as jobs.

Why this answer

DataBrew provides a dedicated 'Fill missing values' transform for substituting nulls with a constant, and an 'Upper case' transform for case normalization. Together they form a repeatable recipe that preserves rows and produces the required output. Removing nulls or using regex replacement does not achieve the same result.

Exam trap

The trap here is assuming that nulls can be handled with a regex-based replace transform, when DataBrew requires the dedicated fill-missing-values transform for null substitution.

849
MCQmedium

A data engineer is building an AWS Glue ETL job that reads from an Amazon S3 bucket containing CSV files and writes to an Amazon Redshift table. The job runs successfully but the Redshift table ends up empty. The engineer checks the AWS Glue job run metrics and sees that the job processed 0 rows. The S3 bucket contains files under the prefix 'data/'. The Glue crawler created a table with the correct schema. What is the MOST likely cause of the empty output?

A.The Redshift table is in a different AWS Region than the S3 bucket.
B.The Glue job's output is written to a temporary S3 location instead of Redshift because the connection is misconfigured.
C.The Glue job's IAM role lacks s3:GetObject permission on the bucket.
D.The Glue job's script is filtering out all records because of an incorrect partition predicate or a misconfigured pushdown predicate.
AnswerD

A common cause of a Glue job processing zero rows is that the job's script applies a filter that excludes all data, such as a pushdown predicate on a partition column that does not match any partition values. Since the job completed successfully but processed no rows, this is the most likely cause.

Why this answer

When an AWS Glue job completes successfully but processes zero rows, the issue is typically within the job's transformation logic rather than permissions or connectivity. A pushdown predicate or filter that does not match any data will cause the job to read no records. Checking the script for such filters is the first step.

Exam trap

The trap here is assuming that a successful job run means data was processed, when in fact a filter can exclude all records.

850
MCQmedium

Refer to the exhibit. A data engineer runs the above AWS CLI command to view the table metadata in the AWS Glue Data Catalog. The data is stored as CSV in S3 with partitions by year and month. When querying the table using Amazon Athena, no data is returned. What is the most likely cause?

A.The partitions have not been added to the Glue Data Catalog.
B.The SerDe is not compatible with CSV files.
C.The S3 location points to a file instead of a folder.
D.The column data types are incorrect for the CSV data.
AnswerA

Athena reads partition locations from the AWS Glue Data Catalog. If year and month partitions were never registered, the table metadata exposes no partition paths, so queries against those partitions scan nothing and return zero rows despite the CSV objects existing in S3.

Why this answer

The AWS CLI command shown only retrieves table metadata, not partition metadata. In AWS Glue, partitions must be explicitly added to the Data Catalog via `MSCK REPAIR TABLE`, `ALTER TABLE ADD PARTITION`, or a Glue crawler. Without partition metadata, Athena cannot locate the data files under the partitioned S3 paths (e.g., `s3://bucket/year=2024/month=01/`), resulting in zero rows returned even though the table schema is defined.

Exam trap

The trap here is that candidates assume the `PARTITIONED BY` clause in the table definition automatically registers the partitions in the Glue Data Catalog, but it only defines the schema; partition metadata must be added separately.

How to eliminate wrong answers

Option B is wrong because the default SerDe for CSV in Athena (`LazySimpleSerDe`) is fully compatible with standard CSV files; no SerDe mismatch would cause zero rows. Option C is wrong because the `LOCATION` in the Glue table points to a folder (the base path), not a file; Athena expects a folder and would fail with an error if a file were specified, not silently return no data. Option D is wrong because incorrect column data types would cause query failures or data conversion errors, not an empty result set; Athena would still attempt to read the data and return rows with nulls or errors.

851
MCQmedium

A data engineer is designing a data ingestion pipeline to load data from an on-premises Oracle database into Amazon Redshift. The pipeline must capture changes (inserts, updates, deletes) with low latency and minimal impact on the source database. Which combination of AWS services should the engineer use?

A.AWS Database Migration Service (DMS) to Amazon S3, then COPY into Redshift
B.Amazon Kinesis Data Streams with a custom producer on Oracle
C.AWS Glue with JDBC connection to Oracle, writing to Redshift
D.AWS Lambda reading from Oracle logs and writing to Redshift
AnswerA

AWS DMS reads Oracle redo logs via change data capture, so inserts, updates and deletes replicate continuously without querying source tables, satisfying the low-latency and minimal-impact constraints. Landing changes in Amazon S3 lets Redshift COPY them in, though deletes require merge handling.

Why this answer

AWS DMS with ongoing replication (change data capture) can capture changes from Oracle with minimal impact and replicate to S3, then COPY into Redshift. Option B is wrong because Kinesis Data Streams requires a custom producer to capture Oracle changes and does not provide native CDC from Oracle databases. Option C is wrong because AWS Glue with JDBC is batch-oriented and does not support real-time change data capture natively.

Option D is wrong because AWS Lambda can process events but is not designed for continuous, low-latency CDC from a database and would require polling or triggers.

852
MCQmedium

Refer to the exhibit. A data engineer notices that the Redshift cluster 'mycluster' does not have automated backups beyond 7 days. However, the compliance team requires a minimum of 35 days of backup retention. What should the engineer do?

A.Change the node type to ra3.xlplus to enable automatic backups for 35 days.
B.Enable audit logging to capture changes for recovery.
C.Take manual snapshots every day and retain them for 35 days.
D.Modify the cluster's automated snapshot retention period to 35 days.
AnswerD

Redshift's automated snapshot retention is a cluster-level setting adjustable from 1 to 35 days. Modifying it to 35 days satisfies the compliance requirement directly, without manual snapshots or a new cluster, since the current 7-day default falls short.

Why this answer

Amazon Redshift allows you to modify the automated snapshot retention period for a cluster up to 35 days. The engineer can use the AWS Management Console, CLI, or API to change the `automated_snapshot_retention_period` parameter from the current 7 days to 35 days, meeting the compliance requirement without additional manual intervention.

Exam trap

The trap here is that candidates may confuse backup retention with node type capabilities or audit logging, assuming that hardware or logging features inherently extend backup duration, when in fact the retention period is a simple configuration parameter.

How to eliminate wrong answers

Option A is wrong because changing the node type to ra3.xlplus does not affect the automated backup retention period; retention is configured independently of node type. Option B is wrong because audit logging captures user activity and SQL queries for security and compliance, not for point-in-time recovery of data; it does not replace backup retention. Option C is wrong because while manual snapshots can be retained for 35 days, this approach requires daily manual effort and does not leverage the automated backup feature that is already available; modifying the automated retention period is simpler and more reliable.

853
MCQeasy

A data engineer needs to store streaming data from IoT devices for real-time analytics. The data has a fixed schema and requires low-latency queries. Which AWS service should be used?

A.Amazon DynamoDB
B.Amazon Redshift
C.Amazon S3
D.Amazon Timestream
AnswerD

Amazon Timestream is purpose-built for time-series IoT data, offering serverless auto-scaling and its memory store for low-latency queries on recent data, with magnetic store for historical retention. It satisfies the fixed-schema, real-time analytics constraint directly, unlike general-purpose databases that require manual provisioning and tuning for streaming ingest.

Why this answer

Amazon Timestream is a time-series database purpose-built for IoT and operational applications that generate large volumes of time-stamped data. It automatically manages data retention and storage tiers (memory and magnetic) to provide fast query performance for recent data and cost-effective storage for historical data, making it ideal for real-time analytics on streaming IoT data with a fixed schema.

Exam trap

AWS often tests the misconception that any database can handle time-series data equally well, but the trap here is that candidates choose DynamoDB for its low-latency reads, overlooking that Timestream is the only AWS service purpose-built for time-series workloads with native support for time-based partitioning, retention policies, and analytical functions.

How to eliminate wrong answers

Option A is wrong because Amazon DynamoDB is a NoSQL key-value and document database optimized for high-throughput, low-latency read/write operations on individual items, but it lacks native time-series optimizations such as automatic downsampling, interpolation, and time-based partitioning, making it less efficient for time-series queries like aggregations over time windows. Option B is wrong because Amazon Redshift is a petabyte-scale data warehouse designed for complex analytical queries on structured and semi-structured data using SQL, but it is not optimized for real-time streaming ingestion or low-latency queries on high-frequency time-series data; its batch-oriented architecture introduces higher latency for streaming use cases. Option C is wrong because Amazon S3 is an object storage service that provides durable, scalable storage for any type of data, but it does not support real-time querying directly; querying S3 requires services like Athena or S3 Select, which add latency and are not designed for sub-second, low-latency queries on streaming data.

854
MCQhard

A data engineer is designing a data lake on Amazon S3. The data is frequently accessed by multiple analytics services, and the company needs to enforce fine-grained access control based on data tags. Which combination of AWS services should be used?

A.S3 Block Public Access settings
B.AWS Lake Formation with tag-based access control
C.S3 Access Points with bucket policies
D.S3 Object Lambda with IAM policies
AnswerB

Lake Formation tag-based access control enforces column, row and table permissions using LF-Tags across analytics services, satisfying the fine-grained, tag-driven requirement. Plain S3 bucket policies or IAM alone cannot express tag-based governance at that granularity.

Why this answer

AWS Lake Formation with tag-based access control (TBAC) is the correct choice because it provides fine-grained, attribute-based access control (ABAC) at the column, row, and cell level across a data lake on S3. By assigning LF-tags to Data Catalog resources and defining permissions based on those tags, you can enforce granular access policies that scale without managing individual user-to-resource mappings. This directly meets the requirement for tag-driven, fine-grained access for multiple analytics services.

Exam trap

The trap here is that candidates often confuse S3 Access Points (which provide network-level or prefix-level restrictions) with the fine-grained, tag-driven access control that Lake Formation TBAC uniquely offers, leading them to pick Option C despite its inability to enforce column- or row-level security based on tags.

How to eliminate wrong answers

Option A is wrong because S3 Block Public Access settings only prevent public exposure of S3 objects and do not provide any fine-grained, tag-based access control for internal users or services. Option C is wrong because S3 Access Points with bucket policies can restrict access based on VPC or IP, but they do not natively support tag-based access control at the column or row level; they operate at the bucket or prefix level only. Option D is wrong because S3 Object Lambda transforms data on read but does not enforce access control based on data tags; IAM policies attached to it cannot dynamically filter data by tags without custom code, and it lacks the centralized governance Lake Formation provides.

855
Multi-Selecthard

A company uses a Kinesis Data Firehose delivery stream to load data into an S3 bucket. The data is in JSON format and must be converted to Parquet before landing in S3. Which steps are required to achieve this? (Choose THREE.)

Select 3 answers
A.Configure the Firehose delivery stream to enable data format conversion to Parquet.
B.Create a table in the AWS Glue Data Catalog with the schema.
C.Store the schema in Amazon DynamoDB.
D.Set the Firehose's schema mapping to reference the Glue table.
E.Use Kinesis Data Analytics to convert the data.
AnswersA, B, D

Firehose has built-in conversion capability.

Why this answer

Kinesis Data Firehose natively supports converting incoming data from JSON to Parquet format. This conversion is enabled directly in the delivery stream configuration, eliminating the need for separate processing steps.

Exam trap

The trap here is that candidates may think DynamoDB is needed for schema storage or that Kinesis Data Analytics is required for the conversion, but Firehose's built-in Parquet conversion with Glue schema support is the correct and simpler approach.

856
MCQhard

A company runs a Redshift cluster for analytics. The data engineering team notices that COPY commands from S3 are failing for large files (>1 GB) with the error 'S3ServiceException: SlowDown'. What is the most effective solution?

A.Use Redshift Spectrum to query the data directly in S3.
B.Enable automatic compression on the target tables.
C.Increase the number of Redshift nodes to distribute the load.
D.Split the large files into smaller parts (e.g., 100 MB each) and use parallel COPY.
AnswerD

S3 SlowDown signals request-rate throttling when Redshift reads very large objects. Splitting files into roughly 100 MB parts and issuing parallel COPY lets Redshift distribute the load across slices, avoiding the throttling that stalls single large-file reads.

Why this answer

The 'SlowDown' error from S3 occurs when too many requests are made to the same S3 prefix, causing throttling. Splitting large files into smaller parts (e.g., 100 MB) and using parallel COPY distributes the load across multiple slices in Redshift, reducing the request rate per prefix and avoiding throttling.

Exam trap

DEA-C01 often tests the misconception that increasing cluster size or using Spectrum solves COPY performance issues, when the root cause is often S3 request throttling due to large files or hot prefixes.

How to eliminate wrong answers

Option A is wrong because Redshift Spectrum is for querying data directly in S3, not for fixing COPY errors. Option B is wrong because enabling compression on target tables does not address the S3 request throttling. Option C is wrong because increasing the number of Redshift nodes does not reduce the request rate to S3; it may even increase it.

857
MCQeasy

A data engineer needs to store JSON documents that are frequently accessed by a low-latency web application. The data does not require complex queries, and the access pattern is primarily by a key. Which AWS service is most appropriate?

A.Amazon ElastiCache for Redis
B.Amazon S3
C.Amazon RDS for MySQL
D.Amazon DynamoDB
AnswerD

Amazon DynamoDB stores JSON as native items and retrieves them by primary key in single-digit milliseconds, satisfying the low-latency, key-based access pattern. Its schemaless design handles document data without complex query requirements, unlike relational engines or object storage, which add overhead unnecessary for this workload.

Why this answer

Amazon DynamoDB is the most appropriate service because it is a fully managed NoSQL key-value and document database designed for single-digit millisecond latency at any scale. It natively supports JSON documents and provides fast, consistent access by primary key without requiring complex query capabilities, making it ideal for low-latency web applications.

Exam trap

The trap here is that candidates often confuse ElastiCache for Redis as a persistent data store for JSON documents, but it is primarily an in-memory cache with optional persistence, not a durable, low-latency database designed for primary key access patterns.

How to eliminate wrong answers

Option A is wrong because Amazon ElastiCache for Redis is an in-memory data store primarily used for caching, session management, and real-time analytics, not for persistent storage of JSON documents that need durability and key-based access with low latency. Option B is wrong because Amazon S3 is an object storage service with higher latency (typically tens to hundreds of milliseconds) and is not optimized for frequent, low-latency key-based lookups required by a web application. Option C is wrong because Amazon RDS for MySQL is a relational database that requires predefined schemas and supports complex queries via SQL, which is overkill and adds unnecessary overhead for simple key-based access to JSON documents.

858
MCQhard

A data engineer is managing an AWS Glue Data Catalog that contains metadata for tables in Amazon S3. The security team requires that access to the Data Catalog be restricted based on the user's department, and that users can only see tables that belong to their department. The Data Catalog tables are tagged with a 'Department' key. Which AWS feature should the engineer use to enforce this requirement?

A.Amazon S3 bucket policies that restrict access to objects based on the 'Department' tag on the S3 objects.
B.AWS Lake Formation tag-based access control (LF-TBAC) with tags on Data Catalog resources and matching IAM principals.
C.AWS Glue resource policies that allow or deny access based on the 'Department' tag.
D.IAM policies with condition keys that match the 'Department' tag on the Data Catalog tables.
AnswerB

AWS Lake Formation supports tag-based access control, which allows you to define permissions based on tags attached to Data Catalog resources and IAM principals. By tagging tables with a Department key and assigning matching tags to users, you can grant or deny access dynamically. This meets the requirement for department-based access without managing individual table permissions.

Why this answer

AWS Lake Formation tag-based access control (LF-TBAC) allows you to define permissions using tags on Data Catalog resources and IAM principals. This enables attribute-based access control, so users can only access tables with matching tags. This is the recommended way to implement fine-grained, scalable access control for the Glue Data Catalog.

Other options either do not support tag-based filtering for the catalog or address the wrong resource.

Exam trap

The trap here is assuming that IAM policies alone can enforce tag-based access to Glue Data Catalog tables, when Lake Formation is required.

859
MCQeasy

A data engineer is building a data lake on Amazon S3 and needs to catalog metadata for a large number of CSV files stored in a folder structure. The engineer wants to use AWS Glue crawlers to automatically infer schemas and create tables in the AWS Glue Data Catalog. The crawler should run daily to detect new files and schema changes. Which configuration should the engineer use for the crawler?

A.Set the crawler's data source to the specific S3 folder and disable the crawler schedule, running it manually when needed.
B.Set the crawler's data source to the specific S3 folder containing the CSV files, configure it to create a single schema for each S3 path, and set a daily schedule.
C.Set the crawler's data source to the specific S3 folder and configure it to create a separate table for each file, with a daily schedule.
D.Set the crawler's data source to the S3 bucket root and enable 'Update the table definition in the Data Catalog' with a schedule of daily.
AnswerB

Pointing the crawler to the specific folder limits its scope to relevant data. Configuring it to create a single schema for each S3 path groups files with the same schema into one table, which is efficient for partitioned data. A daily schedule ensures new files and schema changes are detected automatically.

Why this answer

The crawler should be pointed to the specific S3 folder to avoid scanning irrelevant data. Configuring it to create a single schema for each S3 path groups files with the same schema into one table, which is ideal for a data lake with partitioned folders. A daily schedule ensures the catalog stays up to date with new files and schema changes.

Exam trap

The trap here is choosing to create a separate table for each file, which seems granular but leads to catalog sprawl and is not how crawlers are typically used for partitioned data.

860
MCQhard

A company uses Amazon DynamoDB with on-demand capacity for a gaming application that experiences unpredictable traffic spikes. The application reads the same set of 'hot' items frequently. Users report high latency during peak hours. Which action would MOST effectively reduce read latency for the hot items?

A.Enable DynamoDB Accelerator (DAX) for the table.
B.Switch to provisioned capacity with auto-scaling.
C.Increase the read capacity units for the table.
D.Enable DynamoDB Global Tables for multi-region replication.
AnswerA

DAX provides an in-memory write-through cache for eventually consistent reads, absorbing repeated access to the same hot items and cutting microsecond-level latency. This directly addresses the unpredictable spikes and repeated hot-item reads that overwhelm on-demand capacity.

Why this answer

DynamoDB Accelerator (DAX) is an in-memory cache that sits between the application and DynamoDB, providing microsecond read latency for frequently accessed items. Since the application reads the same set of 'hot' items repeatedly, DAX can serve these reads from its cache, bypassing the storage layer and reducing latency during traffic spikes without requiring any table schema changes.

Exam trap

The trap here is that candidates often confuse throughput capacity (RCUs/WCUs) with latency, assuming that increasing capacity will speed up individual reads, when in fact capacity only controls the rate of requests, not the response time per request.

How to eliminate wrong answers

Option B is wrong because switching to provisioned capacity with auto-scaling does not reduce read latency; it only manages throughput capacity based on load, but the underlying read latency from DynamoDB remains the same. Option C is wrong because increasing read capacity units (RCUs) is only applicable to provisioned capacity mode, not on-demand capacity, and even if it were, it would not reduce latency for hot items—it only increases the maximum throughput. Option D is wrong because DynamoDB Global Tables replicate data across regions for disaster recovery and low-latency reads from distant regions, but it does not reduce latency for reads within the same region; it adds complexity and cost without addressing the hot-item caching issue.

861
Multi-Selecteasy

A data engineer needs to monitor the performance of an Amazon Redshift cluster. Which Amazon CloudWatch metric should the engineer monitor to detect disk space issues?

Select 1 answer
A.ReadIOPS
B.WriteIOPS
C.PercentageDiskSpace
D.NetworkThroughput
E.CPUUtilization
AnswersC

PercentageDiskSpace reports the proportion of each node's disk already consumed, so it surfaces storage exhaustion directly. Monitoring it lets the engineer detect when Redshift nodes approach capacity, which is the disk-space issue the stem asks about, unlike CPU or query-latency metrics.

Why this answer

PercentageDiskSpace is the correct metric because it directly reports the percentage of disk space used in the Redshift cluster. Monitoring this metric allows the engineer to detect when disk space is running low, which can lead to performance degradation or query failures. Unlike I/O or CPU metrics, it specifically measures storage utilization, making it the most relevant for disk space issues.

Exam trap

DEA-C01 often tests the distinction between performance metrics (e.g., IOPS, CPU) and capacity metrics (e.g., disk space), causing candidates to overlook PercentageDiskSpace as the direct indicator of disk space issues.

862
MCQmedium

The exhibit shows an S3 bucket policy. What is the effect of this policy?

A.Allows all S3 actions over HTTPS only.
B.Allows all S3 actions to the bucket over any protocol.
C.Denies all S3 actions to the bucket.
D.Allows only GetObject and PutObject over HTTPS.
AnswerD

Explicit allow for those actions over HTTPS; deny for HTTP.

Why this answer

The S3 bucket policy in the exhibit uses a condition key `aws:SecureTransport` set to `true`, which restricts access to HTTPS only. The `Effect` is `Allow` for `s3:GetObject` and `s3:PutObject` actions, meaning only these two actions are permitted over HTTPS. This matches option D.

Exam trap

The trap here is that candidates see the `Deny` statement and assume the entire policy denies all actions, overlooking the `Allow` statement that permits specific actions over HTTPS.

How to eliminate wrong answers

Option A is wrong because the policy does not allow all S3 actions; it explicitly allows only `s3:GetObject` and `s3:PutObject`. Option B is wrong because the policy denies all actions over non-HTTPS protocols via the `Deny` statement with `aws:SecureTransport=false`, and the `Allow` statement only permits HTTPS. Option C is wrong because the policy does not deny all S3 actions; it allows `GetObject` and `PutObject` over HTTPS, while only denying actions that do not use HTTPS.

863
Multi-Selecthard

A data engineer is configuring a VPC for an Amazon Redshift cluster. The cluster must be accessible only from a specific on-premises network via a Direct Connect connection. Which TWO actions should the engineer take to meet this requirement? (Choose TWO.)

Select 2 answers
A.Enable Redshift Enhanced VPC Routing.
B.Configure a security group to allow inbound traffic from the on-premises CIDR block.
C.Configure a network ACL to allow inbound traffic from the on-premises CIDR block.
D.Create a VPC endpoint for Redshift.
E.Make the Redshift cluster publicly accessible.
AnswersB, C

A security group is stateful and instance-level, so allowing inbound traffic from the on-premises CIDR block restricts Redshift access to that network only, satisfying the requirement that the cluster be reachable solely via Direct Connect.

Why this answer

Option B is correct because a security group acts as the stateful firewall for the Redshift cluster, and adding an inbound rule that permits the on-premises CIDR block on the Redshift port (5439) is required to allow that specific network to reach the cluster. Option C is correct because the network ACL is the stateless subnet-level control, so it must also include an inbound rule allowing the on-premises CIDR block (and corresponding outbound return traffic) for the connection to succeed. Option A is not needed because Enhanced VPC Routing only affects how COPY/UNLOAD traffic is routed to S3 or other services, not client access from on-premises.

Option D is wrong because a VPC endpoint is for private access to AWS services like S3, not for enabling on-premises clients to reach a Redshift cluster. Option E is wrong because making the cluster publicly accessible would expose it to the internet, violating the requirement to restrict access to the on-premises network only.

864
MCQhard

A data engineer manages an AWS Glue ETL job that reads from an Amazon S3 bucket and writes to another bucket. The security team mandates that all data at rest in both buckets be encrypted with customer-managed AWS KMS keys, and that each service use a distinct key. The Glue job's IAM role currently has s3:GetObject and s3:PutObject permissions but jobs fail with an access denied error when writing output. What is the MOST likely cause?

A.The output bucket uses SSE-S3 by default, which overrides the SSE-KMS setting and prevents the customer-managed key from being used.
B.The KMS key policy does not allow the AWS Glue service principal to use the key for cryptographic operations.
C.The S3 bucket policy on the output bucket does not include a statement allowing the Glue job role to perform s3:PutObject.
D.The IAM role lacks kms:GenerateDataKey and kms:Decrypt permissions on the customer-managed KMS key used by the output bucket.
AnswerD

When S3 uses SSE-KMS with a customer-managed key, writing an object requires the caller to have kms:GenerateDataKey on that key, and reading requires kms:Decrypt. The Glue role has S3 permissions but no KMS permissions, so the PutObject call fails with access denied. Adding the required KMS actions on the output key resolves the failure.

Why this answer

Writing to an S3 bucket encrypted with SSE-KMS using a customer-managed key requires both the S3 PutObject permission and KMS permissions on that key, specifically kms:GenerateDataKey for uploads and kms:Decrypt for downloads. The Glue execution role only has S3 permissions, so the KMS authorization check fails and S3 returns access denied. Granting the role the needed KMS actions on the output key fixes the job.

Exam trap

The trap here is focusing only on S3 bucket policies and IAM S3 actions while overlooking that SSE-KMS adds a second authorization layer requiring kms:GenerateDataKey and kms:Decrypt on the key itself.

865
MCQeasy

A data engineer is configuring an AWS Glue crawler to catalog data stored in an Amazon S3 bucket. The security team requires that all data in transit between the crawler and S3 be encrypted using TLS. Which configuration should the engineer implement to meet this requirement?

A.Enable default encryption on the S3 bucket with SSE-S3, which automatically encrypts data in transit.
B.Enable SSL/TLS for the Glue crawler by setting the --enable-ssl parameter in the crawler configuration.
C.Attach a bucket policy to the S3 bucket that denies requests that do not use the aws:SecureTransport condition.
D.Configure the Glue crawler to use a VPC endpoint for S3 and enable encryption in transit on the endpoint.
AnswerC

Enforcing TLS for data in transit to S3 is done by adding a bucket policy that denies requests where aws:SecureTransport is false. This ensures that any request, including from Glue crawlers, must use HTTPS. Glue crawlers use the AWS SDK, which uses HTTPS by default, so they will comply. This is the standard method to enforce encryption in transit for S3.

Why this answer

To enforce encryption in transit for S3, the most direct method is to use a bucket policy that denies requests where aws:SecureTransport is false. This forces all clients, including AWS Glue crawlers, to use HTTPS. Glue crawlers use the AWS SDK, which defaults to HTTPS, so they will continue to work.

This approach is recommended by AWS for compliance with encryption-in-transit requirements.

Exam trap

The trap here is confusing encryption at rest with encryption in transit, and assuming that S3 default encryption covers data in transit.

866
MCQmedium

A company uses Amazon S3 to store raw data and needs to transform it into Parquet format for analytics. The transformation job runs daily on a schedule. Which AWS service is BEST suited for this task?

A.Amazon Redshift
B.Amazon EMR
C.AWS Lambda
D.AWS Glue
AnswerD

AWS Glue provides serverless Spark-based ETL with a scheduler and built-in Parquet conversion, so the daily transformation runs without managing infrastructure. This satisfies the requirement to convert S3 raw data to Parquet on a recurring schedule.

Why this answer

AWS Glue is a fully managed, serverless ETL service that can automatically convert data formats (e.g., from CSV to Parquet) and run on a schedule (e.g., daily). It is ideal for this use case because it is purpose-built for ETL transformations and handles schema discovery, data cataloging, and job scheduling without managing infrastructure. Option A (Amazon Redshift) is wrong because Redshift is a data warehouse for querying, not a transformation service; it could load Parquet but not convert raw data to Parquet directly.

Option B (Amazon EMR) is wrong because EMR requires provisioning and managing clusters, adding operational overhead. Option C (AWS Lambda) is wrong because Lambda has a maximum execution timeout of 15 minutes, which is too short for daily large-scale data transformation jobs.

867
MCQmedium

A data engineer runs an AWS Glue job that reads from a JDBC connection to a PostgreSQL database. The job fails with a 'Connection timed out' error. The Glue job runs in a VPC with the appropriate security group. What is the most likely cause?

A.The network ACL associated with the Glue job's subnet is blocking outbound traffic.
B.The Glue job does not have permission to access the database.
C.The security group does not allow inbound traffic from the Glue job.
D.The database credentials are incorrect.
AnswerA

Security groups are stateful and already correct, so the remaining subnet-level control is the network ACL. A stateless NACL blocking outbound traffic to the PostgreSQL port or return ephemeral range would cause the JDBC connection to time out, matching the reported error.

Why this answer

A network ACL (NACL) is a stateless subnet-level firewall that can block outbound traffic. If the NACL associated with the Glue job's subnet does not allow outbound traffic to the PostgreSQL database's port (e.g., 5432), the connection times out. Since the security group is stated to be appropriate, the NACL is the most likely remaining network-layer cause.

Exam trap

The trap is confusing security groups with NACLs—candidates often blame the security group, but when the SG is stated as correct, the stateless NACL is the next network-layer suspect for timeouts.

How to eliminate wrong answers

Option B is wrong because a permissions issue would typically produce an authentication or authorization error (e.g., 'access denied'), not a connection timeout—timeouts indicate network reachability problems. Option C is wrong because security groups are stateful and the question says the security group is appropriate; also, the Glue job initiates outbound connections, so inbound rules on the Glue side are not the issue. Option D is wrong because incorrect credentials would yield an authentication failure, not a timeout.

868
Multi-Selecthard

Which THREE factors should be considered when choosing between AWS Glue and Amazon EMR for data transformation? (Choose three.)

Select 3 answers
A.Glue automatically stores data in S3 after transformation.
B.EMR allows fine-grained control over cluster configuration and software.
C.EMR supports real-time stream processing with Spark Streaming.
D.Glue is serverless, reducing operational overhead.
E.Glue integrates natively with the Glue Data Catalog for schema management.
AnswersB, D, E

EMR provides flexibility to install custom software and tune clusters.

Why this answer

Amazon EMR provides full control over cluster configuration, including the ability to customize software, install libraries, and tune Spark, Hadoop, or Hive parameters. This fine-grained control is essential for complex or specialized data transformation pipelines that require specific versions or custom configurations.

Exam trap

The trap here is that candidates may confuse Glue's automatic schema discovery with automatic data storage, or assume EMR is the only option for streaming, when in fact both services support streaming but with different levels of control and operational overhead.

869
MCQmedium

A company uses Amazon Kinesis Data Streams to ingest clickstream data from a website. The data is consumed by an AWS Lambda function that writes to Amazon DynamoDB. The Lambda function is seeing high error rates due to DynamoDB write throttling. Which action should be taken to reduce throttling?

A.Use Amazon Kinesis Data Firehose instead of Kinesis Data Streams
B.Add an Amazon SQS queue between Lambda and DynamoDB
C.Increase the Lambda function memory
D.Enable auto scaling on the DynamoDB table
AnswerD

Enabling DynamoDB auto scaling adjusts provisioned write capacity automatically as consumed capacity rises, directly relieving the write throttling that Lambda is hitting. Since the stem identifies DynamoDB write throttling as the error source, scaling the table's write capacity units satisfies that constraint without changing the Kinesis or Lambda components.

Why this answer

Enabling DynamoDB auto scaling increases write capacity automatically when needed. Using Kinesis Data Firehose would change the architecture but does not address throttling directly. Increasing Lambda memory does not help with DynamoDB throttling.

Using SQS would add a queue but does not increase DynamoDB capacity.

870
MCQeasy

A company is using AWS Glue to run ETL jobs that transform data from Amazon DynamoDB to Amazon S3. The DynamoDB table has a large number of items (over 10 million) and is heavily used by production applications. The Glue job reads the entire DynamoDB table each time it runs, causing increased read capacity consumption and affecting production performance. The team wants to reduce the impact on the source DynamoDB table while still keeping the S3 data up-to-date. What should the team do?

A.Use DynamoDB Streams and AWS Lambda to capture changes and write them to S3, then run incremental Glue jobs.
B.Increase the DynamoDB read capacity units to handle the Glue job's read load.
C.Use the DynamoDB console to export the table to S3 in Parquet format.
D.Reduce the parallelism of the Glue job to lower the read throughput.
AnswerA

DynamoDB Streams capture item-level changes without consuming provisioned read capacity, so the production table is no longer scanned in full. Lambda writes those changes to S3, and Glue then processes only the incremental delta, satisfying the requirement to keep S3 current while removing the read-capacity impact on the heavily used source table.

Why this answer

Using DynamoDB Streams and AWS Lambda to capture changes and write them to S3, then running incremental Glue jobs, reduces the read load on the DynamoDB table because only changed data is processed. This approach keeps S3 data up-to-date without scanning the entire table, minimizing impact on production performance.

Exam trap

DEA-C01 often tests the misconception that increasing read capacity or reducing parallelism solves the impact on DynamoDB, but the real solution is to avoid full table scans by using change data capture with Streams and Lambda.

How to eliminate wrong answers

Option B is wrong because increasing read capacity units would allow the Glue job to read more but does not reduce the impact on the table; it just provides more capacity, which may be costly and still affects performance if the table is heavily used. Option C is wrong because exporting the table to S3 in Parquet format is a full export and does not provide incremental updates, and it still consumes read capacity. Option D is wrong because reducing parallelism lowers read throughput but also slows down the job, and it does not eliminate the full table scan; it just makes it slower, which may not keep S3 up-to-date efficiently.

871
MCQhard

A data engineer is building a data pipeline that ingests sensitive data into Amazon S3 and then processes it with AWS Glue. The security team requires that the data be encrypted at rest using a customer managed key in AWS KMS, and that the engineer be able to audit all key usage. The engineer creates a KMS customer managed key and configures the S3 bucket to use SSE-KMS with that key. The Glue job's IAM role has been granted kms:Decrypt and kms:GenerateDataKey permissions on the key. However, when the Glue job runs, it fails with an access denied error related to KMS. Which additional action should the engineer take to resolve the error?

A.Enable automatic key rotation on the KMS key to ensure the Glue job can retrieve the latest key material.
B.Update the KMS key policy to allow the Glue job's IAM role to use the key for cryptographic operations.
C.Change the S3 bucket encryption to SSE-S3 so that KMS permissions are no longer required.
D.Grant the Glue job's IAM role kms:CreateGrant permission on the KMS key.
AnswerB

KMS key policies are the primary access control for KMS keys. Even if an IAM policy grants kms:Decrypt and kms:GenerateDataKey, the key policy must also allow the principal to use the key. If the key policy does not explicitly grant access to the Glue job's role, access is denied. Updating the key policy to allow the role resolves the error while maintaining least privilege.

Why this answer

KMS key policies must explicitly allow the principal to use the key for cryptographic operations. Even with IAM permissions granting kms:Decrypt and kms:GenerateDataKey, the key policy is the ultimate gatekeeper. The Glue job's role must be listed in the key policy with the necessary permissions.

Updating the key policy resolves the access denied error while adhering to the requirement for customer managed keys and auditable key usage.

Exam trap

The trap here is assuming that IAM permissions alone are sufficient for KMS access, when the key policy must also grant access.

872
MCQeasy

A data engineer is troubleshooting a Kinesis Data Firehose delivery stream that ingests JSON log data from web servers. The stream is configured to transform records with an AWS Lambda function and deliver to an Amazon S3 bucket. Recently, the stream has been failing with 'InvalidData' errors. Which action should the engineer take to resolve the issue?

A.Verify the S3 bucket policy allows Firehose to write.
B.Increase the buffer size and interval in the Firehose delivery stream.
C.Change the data format to CSV in the Firehose configuration.
D.Check the CloudWatch Logs for the Lambda function to identify transformation errors.
AnswerD

Firehose returns InvalidData when the Lambda transform function throws an error or returns a malformed record. Inspecting the function's CloudWatch Logs reveals the exact transformation exception, letting the engineer correct the code rather than guessing at stream configuration.

Why this answer

The 'InvalidData' error in Kinesis Data Firehose typically indicates that the Lambda function used for data transformation is failing or returning malformed records. By checking the CloudWatch Logs for the Lambda function, the engineer can identify specific transformation errors, such as incorrect JSON parsing, missing fields, or exceptions, which cause Firehose to reject the records. This is the most direct troubleshooting step because Firehose relies on the Lambda function to return valid transformed data in the expected format.

Exam trap

The trap here is that candidates often confuse 'InvalidData' errors with S3 permission issues or buffer configuration problems, but the error specifically points to a failure in the data transformation step, not the delivery destination or batching settings.

How to eliminate wrong answers

Option A is wrong because if the S3 bucket policy were the issue, the error would be a permission or access denied error, not 'InvalidData'. Option B is wrong because increasing buffer size or interval would not resolve data transformation errors; it only affects how data is batched before delivery. Option C is wrong because changing the data format to CSV would not fix transformation errors; Firehose expects the Lambda function to return data in the same format as the input (JSON) unless explicitly configured otherwise, and the 'InvalidData' error is unrelated to the output format.

873
MCQhard

A data engineer is using AWS Glue to run an ETL job that reads from Amazon S3, performs a join between two large datasets, and writes the result to Amazon Redshift. The job is taking longer than expected, and the engineer suspects data skew. Which technique can help mitigate data skew in the join?

A.Enable the 'spark.sql.adaptive.enabled' configuration and set 'spark.sql.adaptive.skewJoin.enabled' to true.
B.Use AWS Glue's 'groupFiles' option to combine small files before the join.
C.Increase the number of DPUs to provide more memory for the join operation.
D.Use a broadcast join to replicate the smaller dataset across all nodes.
AnswerA

Adaptive Query Execution (AQE) in Spark dynamically optimizes query plans at runtime. Enabling skew join handling allows AQE to detect skewed partitions and split them into smaller sub-partitions, balancing the workload. This significantly reduces the impact of data skew during joins, improving job performance.

Why this answer

Enabling Adaptive Query Execution with skew join handling allows Spark to dynamically detect and split skewed partitions during the join. This redistributes the workload more evenly across executors, mitigating the performance bottleneck caused by data skew.

Exam trap

The trap here is assuming that adding more resources or using broadcast join will solve skew, when in fact skew requires specific runtime optimization techniques like AQE.

874
MCQmedium

A data engineer is using AWS Step Functions to orchestrate a data pipeline that includes an AWS Glue job, an Amazon EMR step, and an Amazon Redshift stored procedure. The engineer needs to ensure that if the AWS Glue job fails, the pipeline retries the job up to three times before failing the entire execution. Which Step Functions state should the engineer use to implement this retry logic?

A.Wait state with a Retry field to pause and then retry.
B.Map state with a Retry field to iterate over the Glue job.
C.Task state with a Retry field configured with MaxAttempts set to 3.
D.Parallel state with a Catch field to handle failures.
AnswerC

In AWS Step Functions, a Task state can include a Retry field that specifies the number of retry attempts and the interval between them. Setting MaxAttempts to 3 will retry the task up to three times upon failure. This is the correct way to implement retry logic for a specific task, such as the AWS Glue job, within the state machine.

Why this answer

AWS Step Functions Task states support a Retry field that allows you to specify retry behavior for a task, including the maximum number of attempts. By setting MaxAttempts to 3, the state machine will automatically retry the AWS Glue job up to three times if it fails. This is the standard and recommended way to implement retry logic for individual tasks.

Exam trap

The trap here is confusing retry mechanisms with error handling states like Catch or Parallel, which serve different purposes and do not provide per-task retry counts.

875
Multi-Selecthard

A data engineer is using AWS Glue to transform data from Amazon S3 and load it into Amazon Redshift. The job must handle upserts (inserts and updates) into an existing Redshift table based on a primary key. The engineer needs to ensure that the job efficiently processes only changed records and minimizes data movement. Which two AWS Glue features or techniques should be used to achieve this? (Choose two.)

Select 2 answers
A.Enable job bookmarks to track previously processed data and only process new or changed files in Amazon S3.
B.Use the 'write_dynamic_frame.from_jdbc_conf' method with the 'preactions' and 'postactions' parameters to run SQL commands before and after writing.
C.Stage the transformed data in Amazon S3 and use a Redshift COPY command to load into a staging table, then perform a MERGE operation.
D.Use the 'glueContext.write_dynamic_frame.from_options' with 'upsert' mode and specify the primary key.
E.Configure the Glue job to use the 'redshift-upsert' connection type and set the 'mergeKey' parameter.
AnswersA, C

This option is correct because AWS Glue job bookmarks maintain state information about data that has already been processed. When reading from Amazon S3, bookmarks allow the job to process only new files or files that have changed since the last run. This reduces the amount of data read and transformed, improving efficiency and minimizing data movement, which is essential for incremental upserts.

Why this answer

To efficiently upsert data into Redshift, job bookmarks reduce processing to only new or changed S3 files, and staging the data in S3 followed by a COPY into a staging table and a MERGE operation applies changes without full reloads. These techniques together minimize data movement and ensure only changed records are processed.

Exam trap

The trap here is assuming AWS Glue has native upsert capabilities for Redshift, when in fact upserts require manual staging and merge logic.

876
Multi-Selectmedium

A company is designing a data ingestion pipeline for real-time sensor data from thousands of devices. The data must be processed with low latency and stored in Amazon S3. Which TWO services would be appropriate for this use case? (Choose TWO.)

Select 2 answers
A.AWS Glue
B.AWS DataSync
C.Amazon Athena
D.Amazon Kinesis Data Firehose
E.Amazon Kinesis Data Streams
AnswersD, E

Firehose can deliver streaming data to S3 with buffering.

Why this answer

Amazon Kinesis Data Firehose is the correct choice because it is a fully managed service designed to ingest real-time streaming data, transform it on the fly (e.g., convert to Parquet/ORC), and deliver it directly to Amazon S3 with low latency. It handles buffering, compression, and partitioning automatically, making it ideal for the described sensor data pipeline.

Exam trap

The trap here is that candidates may choose only one streaming service, but the question requires two services, and the correct pairing is Kinesis Data Streams for real-time ingestion and Kinesis Data Firehose for delivery to S3, as Firehose alone cannot provide sub-second latency.

877
MCQmedium

A company uses Amazon S3 to store log files. The security team notices that some objects are being accessed from an unexpected AWS account. The data engineer needs to identify which specific IAM user or role is accessing the objects. Which AWS service should be used to get this information?

A.AWS Trusted Advisor
B.Amazon S3 server access logs
C.AWS CloudTrail
D.AWS Config
AnswerC

CloudTrail records the identity of the principal making each S3 data-plane API call, including the IAM user or role ARN and the source account. This satisfies the stem's need to attribute unexpected cross-account object access to a specific identity.

Why this answer

AWS CloudTrail records API activity in your AWS account, including S3 data events such as GetObject, PutObject, and DeleteObject. To identify the specific IAM user or role accessing S3 objects, you need to enable S3 data events in CloudTrail, which will log the identity of the caller. CloudTrail is the correct service for auditing and tracking user activity.

Exam trap

DEA-C01 often tests the difference between S3 server access logs (which lack IAM identity) and CloudTrail data events (which include the IAM principal), leading candidates to choose the wrong service.

How to eliminate wrong answers

Option A is wrong because AWS Trusted Advisor provides recommendations and best practices checks, but it does not provide detailed access logs or identify specific IAM principals. Option B is wrong because Amazon S3 server access logs record requests to a bucket but do not include the IAM user or role identity; they only show the bucket owner, requester, and other details, but not the IAM principal. Option D is wrong because AWS Config is used for assessing, auditing, and evaluating the configurations of AWS resources, not for tracking API access.

878
MCQeasy

A data engineer needs to restrict access to an Amazon S3 bucket so that only objects encrypted with a specific AWS KMS key can be uploaded. Which S3 bucket policy condition should be used?

A.s3:x-amz-server-side-encryption-aws-kms-key-id
B.kms:ViaService
C.s3:x-amz-server-side-encryption
D.kms:EncryptionContext
AnswerA

The s3:x-amz-server-side-encryption-aws-kms-key-id condition key inspects the KMS key ID supplied in the upload request, so the bucket policy denies any PutObject not encrypted with the specified key. This directly enforces the stem's single-key upload restriction.

Why this answer

The correct condition is s3:x-amz-server-side-encryption-aws-kms-key-id (option A). This condition key allows you to require that objects uploaded to the S3 bucket are encrypted with a specific AWS KMS key by checking the key ID used in the encryption header. Option B (kms:ViaService) restricts KMS key usage to specific AWS services but does not enforce a key ID on S3 objects.

Option C (s3:x-amz-server-side-encryption) only checks whether server-side encryption is enabled, not the specific key. Option D (kms:EncryptionContext) is used to enforce encryption context, not the key ID.

879
MCQhard

Refer to the exhibit. A data engineer runs the AWS CLI command shown to encrypt a file using AWS KMS. The command succeeds. Later, the engineer tries to decrypt the file using the same key but without providing an encryption context. The decryption fails. What is the most likely reason?

A.The KMS key policy does not allow decryption.
B.The KMS key has been disabled.
C.The plaintext file was corrupted.
D.The encryption context must be provided during decryption.
AnswerD

The encryption context supplied during encryption is cryptographically bound to the ciphertext, so KMS requires the identical key-value pairs to authorise decryption. Omitting them causes the request to fail, satisfying the stem's constraint that the same key alone is insufficient.

Why this answer

When encrypting data with AWS KMS, an encryption context can be provided as additional authenticated data (AAD). This context must be supplied during decryption; otherwise, decryption fails. The command succeeded during encryption, so the key is not disabled (B is wrong) and the key policy is not the issue (A is wrong).

File corruption would cause a different error (C is wrong).

880
MCQeasy

A data engineer needs to troubleshoot why an AWS Glue job is failing with a 'Insufficient Memory' error. The job processes a 10 GB dataset. Which step should the engineer take FIRST?

A.Switch from using Apache Spark to Python shell.
B.Repartition the data into more partitions within the job.
C.Change the job type from Python to Java.
D.Increase the number of DPUs allocated to the job.
AnswerD

Insufficient memory in AWS Glue typically stems from insufficient worker capacity for the 10 GB dataset. Increasing the number of DPUs adds executors and memory, directly addressing the constraint before considering other tuning such as partitioning or worker type changes.

Why this answer

The FIRST step when a Glue job fails with 'Insufficient Memory' is to increase the number of DPUs (Data Processing Units) allocated to the job. More DPUs provide more executors and memory, directly addressing the memory constraint. This is the most direct, low-risk remediation before considering code or job-type changes.

Exam trap

DEA-C01 often tests the order of troubleshooting steps — candidates jump to code changes (repartition, language switch) when the question asks for the FIRST step, which is the simplest resource fix: increase DPUs.

How to eliminate wrong answers

Option A is wrong because switching to a Python shell job is for lightweight, single-node scripts — it cannot handle a 10 GB distributed dataset and would likely fail or be far slower. Option B is wrong because repartitioning can help with skew but does not increase total memory; if the job is genuinely memory-constrained, repartitioning alone will not resolve it. Option C is wrong because changing the job language from Python to Java does not inherently increase memory — it changes runtime characteristics but not the fundamental resource allocation.

881
MCQmedium

A data engineer needs to design a data ingestion pipeline that ingests CSV files from an Amazon S3 bucket, transforms the data by adding a timestamp column, and loads it into an Amazon Redshift table. The pipeline should run automatically whenever a new file is uploaded to the S3 bucket. Which AWS service should be used to trigger the transformation?

A.AWS Lambda
B.AWS Step Functions
C.Amazon EventBridge
D.Amazon Simple Queue Service (SQS)
AnswerA

Lambda integrates natively with S3 event notifications, so an object-created event invokes the function immediately. The function runs the transformation code, adding the timestamp column, then loads results into Redshift, meeting the automatic per-upload trigger requirement.

Why this answer

Amazon S3 can be configured to send events directly to AWS Lambda when a new CSV file is uploaded. Lambda then executes the transformation (adding a timestamp column) and loads the data into Redshift. Option B (Step Functions) is not triggered directly by S3 events without an intermediate service like Lambda.

Option C (EventBridge) can route S3 events to Lambda, but a direct S3 event notification to Lambda is simpler and more common; however, the key point is that Lambda is the direct trigger. Option D (SQS) requires a separate process to poll the queue and invoke Lambda; it is not a direct trigger.

882
MCQeasy

A data engineer is using Amazon Athena to query data stored in Amazon S3. The engineer notices that queries are slow and scan large amounts of data. The data is stored in CSV format without compression. Which action should the engineer take to improve query performance and reduce cost?

A.Use Amazon Redshift Spectrum to query the S3 data instead of Athena.
B.Convert the data to Parquet format and compress it with Snappy.
C.Increase the number of Athena query executions by using a larger workgroup.
D.Partition the data by date and store it in CSV format.
AnswerB

Converting data to a columnar format like Parquet allows Athena to read only the columns needed for a query, reducing data scanned. Compressing with Snappy further reduces storage size and I/O. This combination significantly improves query performance and lowers cost because Athena charges based on data scanned.

Why this answer

Athena queries are optimized by using columnar formats like Parquet, which allow column pruning, and compression, which reduces data scanned. Converting from CSV to Parquet with Snappy compression will dramatically reduce the amount of data read, improving performance and lowering cost. Partitioning can further help but is secondary to format and compression.

Exam trap

The trap here is focusing on partitioning alone while ignoring the larger gains from columnar storage and compression, which directly reduce data scanned.

883
MCQhard

A data engineer is using AWS Glue to process a large dataset stored in Amazon S3. The dataset is partitioned by year/month/day and consists of Parquet files. The engineer notices that the Glue job is running slowly and consuming excessive DPU hours. The job performs a join between two large tables and writes the output back to S3. Which optimization technique should the engineer implement to improve performance and reduce cost?

A.Increase the number of DPUs allocated to the job to speed up execution.
B.Use predicate pushdown to filter data before the join and partition the output by a common column.
C.Convert the Parquet files to CSV to reduce storage size and improve read performance.
D.Enable AWS Glue job bookmarks to avoid reprocessing old data.
AnswerB

Predicate pushdown filters data at the source, reducing the amount of data read and shuffled during the join. Partitioning the output by a common column improves downstream query performance and reduces write overhead. These techniques directly address the slow join and high DPU usage by minimizing data movement and processing. This is a standard optimization for Glue ETL jobs dealing with large partitioned datasets.

Why this answer

Predicate pushdown reduces the amount of data read and shuffled by applying filters early in the execution plan. Partitioning the output by a common column improves data layout for subsequent queries and reduces write overhead. Together, these techniques optimize join performance and lower DPU consumption, directly addressing the slow job and high cost.

Other options either do not target the join inefficiency or would worsen performance.

Exam trap

The trap here is thinking that simply adding more DPUs will solve performance issues without addressing the underlying data movement and join inefficiencies.

884
MCQmedium

Refer to the exhibit. A data engineer runs two queries on an Athena table partitioned by 'ds'. Both queries scan the same amount of data. What does this indicate?

A.The partition column is not being used as a filter
B.The table does not have any partitions defined
C.The table is not partitioned
D.Partition pruning is working correctly
AnswerA

Athena bills and scans by partition. Identical bytes scanned across both queries means neither predicate pruned partitions, so the ds column was not applied as a filter and the engine read the full table each time.

Why this answer

If two queries that filter on different partition values scan the same amount of data, it means the partition column is not being used as a filter in the WHERE clause — Athena is scanning all partitions instead of pruning to the relevant ones. Partition pruning only occurs when the query includes a predicate on the partition column (e.g., WHERE ds = '2024-01-01'), so equal scan sizes across different filters indicate the predicate is missing or not applied.

Exam trap

The trap is assuming that simply having a partitioned table guarantees pruning — candidates forget that pruning requires an explicit predicate on the partition column in the WHERE clause, and equal scan sizes are the telltale sign it's missing.

How to eliminate wrong answers

Option B is wrong because if the table had no partitions defined, the question's premise of a partitioned table would be false, and the symptom would be a full-table scan regardless — but the question states the table is partitioned by 'ds'. Option C is wrong for the same reason: the table is explicitly partitioned, so 'not partitioned' contradicts the scenario. Option D is wrong because if partition pruning were working correctly, the two queries would scan different amounts of data corresponding to their respective partitions, not the same amount.

885
MCQhard

The exhibit shows an IAM policy attached to a role used by an AWS Glue ETL job. The job reads from an S3 bucket and writes to another S3 bucket. However, the job fails with an access denied error when trying to write to the output bucket. What is the most likely cause?

A.The policy is missing permissions for AWS KMS to decrypt/encrypt objects
B.The policy does not allow glue:StartJobRun on the specific job
C.The policy only allows PutObject on the my-data-lake bucket, but the job writes to a different bucket
D.The policy does not include s3:ListBucket permission
AnswerC

IAM evaluates the role's policy against the actual resource ARN in the request. The statement permits s3:PutObject only on the my-data-lake bucket, so writing to the separate output bucket matches no Allow statement and is implicitly denied, producing the access denied error.

Why this answer

The IAM policy attached to the Glue job's role grants s3:PutObject only on the my-data-lake bucket ARN, but the ETL job's output path points to a different bucket. S3 authorization is evaluated per-resource ARN, so a PutObject call to any bucket not listed in the Resource element is implicitly denied. The fix is to add the output bucket's ARN (or a wildcard) to the policy's Resource list.

Exam trap

DEA-C01 often tests the misconception that a role with any S3 write permission can write to any bucket, when in fact IAM Resource ARNs are bucket-specific and a missing ARN produces an AccessDenied even when the action itself is granted.

How to eliminate wrong answers

Option A is wrong because KMS permissions are only required when objects are encrypted with SSE-KMS; the symptom here is a write denial, not a decrypt/encrypt failure, and the question gives no evidence of KMS encryption. Option B is wrong because glue:StartJobRun controls whether the job can be launched at all — if it were missing, the job would never start, not fail mid-execution on an S3 write. Option D is wrong because s3:ListBucket is needed for listing/head operations, not for PutObject; a missing ListBucket would cause listing errors, not an access-denied on write.

886
MCQeasy

A data engineer is building a data lake on Amazon S3. The engineer needs to catalog metadata for data stored in Parquet format and make it queryable by Amazon Athena. The data is partitioned by year, month, and day in the S3 path. Which AWS service should the engineer use to create and manage the table definitions and partitions?

A.Amazon Athena Data Catalog
B.AWS Glue Data Catalog
C.AWS Lake Formation
D.Amazon S3 Inventory
AnswerB

AWS Glue Data Catalog is a centralized metadata repository that stores table definitions, schemas, and partition information. It integrates natively with Amazon Athena, allowing the engineer to define tables and partitions once and query them directly from Athena without additional configuration.

Why this answer

AWS Glue Data Catalog is the metadata store for Amazon Athena. It allows the data engineer to define table schemas and partitions that Athena can query directly. Using the Glue Data Catalog ensures seamless integration and simplifies data discovery and querying.

Exam trap

The trap here is thinking that Athena has its own separate data catalog, when in fact it relies exclusively on the AWS Glue Data Catalog for table metadata.

887
MCQeasy

A data engineer is setting up an AWS Glue ETL job that reads data from an Amazon S3 bucket and writes to another S3 bucket. The security team requires that all data in transit be encrypted using TLS. The engineer has configured the job to use the appropriate S3 endpoints. Which additional configuration is necessary to enforce TLS for data in transit between AWS Glue and Amazon S3?

A.Set the AWS Glue job's security configuration to enable S3 encryption in transit.
B.Attach an S3 bucket policy that denies requests where aws:SecureTransport is false.
C.Enable default encryption on the S3 bucket using SSE-KMS.
D.Configure the AWS Glue job to use a VPC endpoint for S3 and enable private DNS.
AnswerB

An S3 bucket policy with a condition that denies requests when aws:SecureTransport is false enforces that all requests to the bucket must use TLS. This applies to AWS Glue and any other client, ensuring data in transit is encrypted. This is a standard method to enforce TLS for S3 access.

Why this answer

To enforce TLS for data in transit to S3, you must use a bucket policy that denies requests when aws:SecureTransport is false. This ensures that all access, including from AWS Glue, uses HTTPS. Other options address encryption at rest or network routing, not TLS enforcement.

Exam trap

The trap here is confusing encryption at rest with encryption in transit, or assuming that VPC endpoints automatically enforce TLS. The condition aws:SecureTransport is the key to enforcing TLS.

888
MCQmedium

A data engineer needs to ingest data from an Amazon S3 bucket into an Amazon Redshift cluster. The data is in Parquet format and is partitioned by date. The engineer wants to load only the latest partition into Redshift and ensure the load is efficient. Which method should the engineer use?

A.Use the Redshift COPY command specifying the S3 path for the latest partition.
B.Use Amazon Redshift Spectrum to query the latest partition and insert into a Redshift table.
C.Use AWS Glue to read the entire S3 bucket and write to Redshift, filtering by date.
D.Use the Redshift COPY command with the PARTITION option.
AnswerA

The COPY command can load data directly from a specific S3 prefix. By specifying the path to the latest partition (e.g., s3://bucket/data/date=2023-10-01/), you load only that partition efficiently. This is the most direct and performant method for Parquet data.

Why this answer

The Redshift COPY command is the most efficient way to load Parquet data from S3. By specifying the S3 path to the latest partition, you load only the required data. Other methods like using Glue or Spectrum involve extra processing or are less efficient for this specific task.

Exam trap

The trap here is thinking there is a PARTITION parameter in the COPY command, or that you must process the entire bucket to filter later.

889
MCQmedium

Refer to the exhibit. A data engineer applies this S3 bucket policy to an S3 bucket. What is the effect of this policy?

A.Allows access only from specific IP addresses.
B.Allows only HTTPS requests to get and put objects, and denies HTTP requests.
C.Allows only GetObject actions over HTTPS.
D.Allows anonymous access to get and put objects over HTTP.
AnswerB

The policy's `aws:SecureTransport` condition evaluates the request protocol, so `"aws:SecureTransport": "false"` in a Deny statement blocks any non-TLS call. This satisfies the stem's requirement to enforce encryption in transit: plain HTTP `GetObject` and `PutObject` attempts are rejected, while HTTPS equivalents remain permitted.

Why this answer

The bucket policy allows GetObject and PutObject actions only when the request uses HTTPS, and explicitly denies all S3 actions when the request uses HTTP due to the condition `aws:SecureTransport=false`. Therefore, only HTTPS requests for Get and Put are permitted. Option A is incorrect because the policy does not restrict by IP addresses.

Option C is incorrect because both Get and Put are allowed over HTTPS, not just Get. Option D is incorrect because the policy does not grant anonymous access; it requires secure transport and does not allow HTTP.

890
MCQhard

A company runs an AWS Glue ETL job that reads data from Amazon S3, transforms it, and writes back to S3 in a different partition structure. The job uses the 'spark.sql.shuffle.partitions' option set to 200. After the job completes, the output has many small files. The data engineer wants to minimize the number of output files while maintaining job performance. Which action should the engineer take?

A.Use 'coalesce(n)' with n based on target file size (e.g., 128 MB) before writing.
B.Enable S3 multipart upload for the Glue job.
C.Increase the 'spark.sql.shuffle.partitions' to 500.
D.Reduce the 'spark.sql.shuffle.partitions' to 50.
AnswerA

Using `coalesce(n)` merges partitions without a full shuffle, so the 200 shuffle partitions collapse into roughly 128 MB output files, satisfying the requirement to minimise small files while preserving job performance. Unlike `repartition`, it avoids the network overhead of redistributing data across executors, keeping the write efficient.

Why this answer

'coalesce(n)' reduces the number of partitions without triggering a full shuffle, allowing you to control the number of output files based on a target file size (e.g., 128 MB). This minimizes small files while preserving job performance, as coalesce is a narrow transformation that avoids the overhead of a shuffle. In contrast, 'repartition(n)' would cause a full shuffle, degrading performance.

Exam trap

The trap here is that candidates often confuse 'coalesce' with 'repartition' or assume that adjusting 'spark.sql.shuffle.partitions' directly controls output file count, when in fact it only controls the number of partitions during shuffle operations, not the final write partition count.

How to eliminate wrong answers

Option B is wrong because enabling S3 multipart upload does not reduce the number of output files; it only improves upload reliability and throughput for large objects, but the job still writes the same number of small files. Option C is wrong because increasing 'spark.sql.shuffle.partitions' to 500 would increase the number of shuffle partitions, leading to even more small output files and potentially worse performance due to higher task overhead. Option D is wrong because reducing 'spark.sql.shuffle.partitions' to 50 would reduce the number of shuffle partitions, but it does not directly control the number of output files written; it may cause data skew and memory pressure, and the output file count still depends on the final partition count, which may remain high if the job uses repartition or other transformations.

891
MCQhard

A company uses Amazon Kinesis Data Streams to ingest clickstream data. The data is consumed by an AWS Lambda function that processes each record and writes to an Amazon DynamoDB table. Recently, the Lambda function has been failing with 'ProvisionedThroughputExceededException' from DynamoDB. The Lambda function uses the AWS SDK to batch write items in batches of 25. The DynamoDB table has on-demand capacity mode. The stream has 10 shards, and the Lambda function is configured with a batch size of 100 and 5 concurrent invocations per shard. What step should the team take to resolve the issue?

A.Switch the DynamoDB table from on-demand to provisioned capacity with a high write capacity unit (WCU) value.
B.Reduce the Lambda batch size to 25 and implement exponential backoff with jitter in the Lambda code.
C.Increase the number of Kinesis shards to 20 to reduce the load per shard.
D.Increase the Lambda function's reserved concurrency to allow more parallel executions.
AnswerB

Reducing batch size and adding exponential backoff directly reduces the write rate and adds retry logic, mitigating throttling. This is the best approach.

Why this answer

The issue is DynamoDB throttling due to high write traffic from Lambda. The DynamoDB table is on-demand, which can throttle if bursts exceed sustained limits. Reducing the Lambda batch size from 100 to 25 decreases the number of records processed per invocation, lowering the instantaneous write rate.

Implementing exponential backoff with jitter in the Lambda code allows retries on throttled requests, making the system more resilient. Option A is not required because on-demand mode automatically scales, and switching to provisioned can be costly. Option C is incorrect because increasing shards would increase parallelism and worsen throttling.

Option D is incorrect because increasing concurrency would also increase write pressure on DynamoDB.

892
Multi-Selecthard

A data engineer is troubleshooting an AWS Glue job that reads from Amazon RDS MySQL and writes to Amazon S3. The job runs successfully but takes longer than expected. The engineer wants to optimize performance. Which THREE actions would improve job performance?

Select 3 answers
A.Increase the number of DPUs allocated to the Glue job.
B.Use a single JDBC connection per partition.
C.Increase the JDBC fetch size parameter.
D.Convert the output format from Parquet to CSV.
E.Use a pushdown predicate to filter data at the source.
AnswersA, C, E

More DPUs provide more parallelism.

Why this answer

Increasing the number of DPUs (Data Processing Units) allocated to the AWS Glue job provides more parallel processing capacity, allowing the job to process data faster. This is a direct way to improve performance when the job is CPU or memory-bound, as Glue distributes the workload across the allocated DPUs.

Exam trap

The trap here is that candidates might think converting to CSV improves performance due to simplicity, but in reality, Parquet's columnar storage and compression provide significant performance benefits for analytics workloads on S3.

893
Multi-Selectmedium

Which TWO AWS services can be used to ingest streaming data from a mobile application into Amazon S3 for near-real-time analytics? (Choose 2.)

Select 2 answers
A.Amazon Kinesis Data Firehose
B.Amazon Kinesis Data Streams
C.Amazon DynamoDB Streams
D.AWS Glue
E.Amazon SQS
AnswersA, B

Firehose can ingest streaming data and deliver to S3 near real-time.

Why this answer

Amazon Kinesis Data Firehose is the correct choice because it is a fully managed service designed to load streaming data directly into Amazon S3, Redshift, Elasticsearch, or Splunk without requiring custom code. It can capture and transform streaming data from mobile applications in near-real-time and automatically deliver it to S3, making it ideal for near-real-time analytics pipelines.

Exam trap

The DEA-C01 exam often tests the distinction between managed ingestion (Firehose) and raw stream processing (Data Streams), and the trap here is that candidates may incorrectly choose DynamoDB Streams or SQS because they associate 'streaming' with any service containing 'stream' or 'queue', without understanding that DynamoDB Streams only captures internal table changes and SQS requires custom code to write to S3.

894
Multi-Selecthard

A data engineer is designing a streaming ingestion pipeline using Amazon Kinesis Data Streams. The stream receives records from thousands of IoT devices, and the engineer must ensure that records from the same device are processed in order. The engineer also needs to scale the stream to handle peak loads without manual intervention. Which two actions should the engineer take? (Choose two.)

Select 2 answers
A.Enable enhanced fan-out for consumers to increase read throughput.
B.Use the device ID as the partition key when putting records into the stream.
C.Use a random partition key to evenly distribute records across shards.
D.Configure the stream to use on-demand capacity mode.
E.Increase the retention period of the stream to 365 days.
AnswersB, D

Using the device ID as the partition key ensures that all records from the same device are routed to the same shard. Within a shard, records are processed in the order they arrive. This guarantees per-device ordering, which is a key requirement. It also distributes the load across shards based on the number of devices, enabling horizontal scaling.

Why this answer

To ensure per-device ordering, records from the same device must go to the same shard, which is achieved by using the device ID as the partition key. To handle peak loads without manual intervention, the stream should use on-demand capacity mode, which automatically scales shards. Together, these actions meet both ordering and scaling requirements for the IoT streaming pipeline.

Exam trap

The trap here is assuming that enhanced fan-out or random partitioning improves ordering or scaling; enhanced fan-out only affects read throughput, and random partitioning breaks ordering.

895
MCQeasy

A retail company uses Amazon Redshift for its data warehouse. The security team requires that all data in the cluster be encrypted at rest using a hardware security module (HSM) to manage the encryption keys. The data engineer needs to configure the Redshift cluster accordingly. Which action should the data engineer take?

A.Use Redshift Spectrum to query data in Amazon S3 that is encrypted with an HSM, and enable encryption for the Redshift cluster with AWS KMS.
B.Enable Redshift encryption at rest with an AWS owned key and use AWS CloudHSM to store the key.
C.Enable Redshift encryption at rest using AWS KMS with a customer managed key, and configure the cluster to use an HSM for key storage.
D.Configure the Redshift cluster to use an HSM for encryption at rest by specifying the HSM connection details and enabling encryption when creating the cluster.
AnswerD

Amazon Redshift supports encryption at rest using a hardware security module (HSM) for key management. When creating or modifying a cluster, you can enable encryption and specify an HSM connection. This meets the requirement for HSM-based encryption. The HSM must be configured with the appropriate keys and network access.

Why this answer

Amazon Redshift supports encryption at rest using an HSM. When creating a cluster, you can choose HSM encryption and provide the HSM connection details. This allows the cluster to use keys stored in a hardware security module, satisfying the security team's requirement.

Other options either use KMS or do not encrypt the cluster itself.

Exam trap

The trap here is assuming that AWS KMS or CloudHSM can be used directly for Redshift HSM encryption, when Redshift requires a specific HSM connection configuration.

896
MCQeasy

A company stores time-series sensor data in Amazon S3. They need to query the data using SQL with minimal latency and no infrastructure management. Which service should they use?

A.Amazon Kinesis Data Analytics
B.Amazon Athena
C.Amazon Redshift
D.Amazon DynamoDB
AnswerB

Amazon Athena queries S3 data directly using standard SQL, with no servers to provision or manage. It satisfies both stated constraints: minimal latency through parallel query execution, and zero infrastructure management. Unlike Amazon Redshift, which requires cluster provisioning, Athena is serverless and pay-per-query, making it ideal for ad hoc time-series analysis.

Why this answer

Amazon Athena is the correct choice because it is a serverless interactive query service that allows you to analyze data directly in Amazon S3 using standard SQL without any infrastructure to manage. It is optimized for querying structured, semi-structured, and unstructured data stored in S3, making it ideal for time-series sensor data with minimal latency requirements.

Exam trap

The trap here is that candidates often confuse Amazon Athena with Amazon Redshift Spectrum, but the question explicitly requires 'no infrastructure management,' which eliminates Redshift; also, Kinesis Data Analytics is mistakenly chosen by those who think it can query static S3 data, but it is strictly for real-time streams.

How to eliminate wrong answers

Option A is wrong because Amazon Kinesis Data Analytics is designed for real-time stream processing using SQL or Apache Flink, not for querying static data already stored in S3; it requires a streaming data source and incurs ongoing processing costs. Option C is wrong because Amazon Redshift is a fully managed data warehouse that requires provisioning and managing clusters, which contradicts the 'no infrastructure management' requirement; it is also overkill for simple SQL queries on S3 data and incurs higher costs for idle compute. Option D is wrong because Amazon DynamoDB is a NoSQL key-value and document database, not designed for SQL queries on S3 data; it requires data to be loaded into tables and does not support direct querying of S3 objects.

897
MCQmedium

A data engineer is troubleshooting an AWS Glue job that is reading from an Amazon Kinesis Data Stream. The job is configured with a 1-minute window and is supposed to process the latest records. However, the engineer notices that the job is reprocessing old data from the stream. What is the most likely cause of this issue?

A.The Glue job's checkpointing is disabled, so it restarts from the beginning each time it runs.
B.The Glue job is configured with the starting position 'TRIM_HORIZON', which reads all available records from the beginning of the stream.
C.The Kinesis stream has multiple shards, and the Glue job is not parallelizing reads across all shards, causing it to reread from one shard.
D.The Glue job's bookmarking is enabled, causing it to track processed records and reprocess them if the job fails.
AnswerB

In AWS Glue, when reading from Kinesis, you can specify the starting position. 'TRIM_HORIZON' means the job starts reading from the oldest record in the stream. If the job is reprocessing old data, this setting is likely the cause. Changing to 'LATEST' would start from new records.

Why this answer

When an AWS Glue job reads from Kinesis, the starting position determines where reading begins. 'TRIM_HORIZON' reads from the oldest record, while 'LATEST' reads only new records. If the job is reprocessing old data, the starting position is likely set to 'TRIM_HORIZON'. Adjusting this to 'LATEST' will make the job process only new data going forward.

Exam trap

The trap here is confusing job bookmarks with Kinesis starting position; bookmarks are not used for Kinesis streams.

898
MCQmedium

A data engineer is designing an Amazon S3 data lake and needs to enforce schema-on-read for a dataset that is queried by Amazon Athena. The data is stored as Parquet files partitioned by year, month, and day. The engineer wants to minimize the amount of data scanned by queries that filter on a specific date range. Which approach should the engineer take?

A.Create an AWS Glue Data Catalog table with partition keys year, month, and day, and run MSCK REPAIR TABLE or use partition projection before querying in Athena.
B.Store all Parquet files in a single prefix and rely on Parquet column pruning to skip dates.
C.Convert the dataset to CSV and add a WHERE clause on the date column in every query.
D.Create an Amazon Redshift Spectrum external table and query it from Redshift instead of Athena.
AnswerA

Registering the partitions in the Data Catalog lets Athena prune partitions that do not match the date filter, so only the relevant S3 prefixes are read. Partition projection can compute partition locations from the table properties without running a repair step, which is more scalable. Either way, partition pruning is what minimizes scanned bytes.

Why this answer

Athena uses the AWS Glue Data Catalog to resolve table schemas and partition locations. When the year, month, and day columns are declared as partition keys and the partitions are either registered or projected, date filters cause Athena to read only the matching S3 prefixes. Storing everything in one prefix, converting to CSV, or moving to Redshift Spectrum does not achieve the same partition-pruning benefit.

Exam trap

The trap here is confusing Parquet column pruning and predicate pushdown with partition pruning, assuming file format alone eliminates scanning of non-matching dates.

899
MCQeasy

A data engineer needs to store semi-structured JSON transaction logs for analytics. The logs are written once and rarely accessed. The storage must be cost-effective. Which AWS service should be used?

A.Amazon S3
B.Amazon DynamoDB
C.Amazon RDS
D.Amazon Redshift
AnswerA

Amazon S3 provides durable object storage with tiered classes such as S3 Standard-IA and Glacier, matching the write-once, rarely accessed, cost-sensitive requirement. It natively holds semi-structured JSON, and analytics tools query it directly without provisioning servers.

Why this answer

Amazon S3 is the correct choice because it provides highly durable, cost-effective object storage ideal for semi-structured JSON transaction logs that are written once and rarely accessed. S3's lifecycle policies can automatically transition such infrequently accessed data to S3 Glacier or S3 Glacier Deep Archive for even lower storage costs, making it the most economical option for this use case.

Exam trap

The trap here is that candidates may choose DynamoDB or Redshift because they support JSON natively, but they overlook the core requirement of cost-effective storage for rarely accessed data, which is best met by S3's low-cost object storage and lifecycle management features.

How to eliminate wrong answers

Option B (Amazon DynamoDB) is wrong because it is a NoSQL key-value and document database optimized for low-latency, high-throughput read/write operations, not for cost-effective archival storage of rarely accessed logs; storing large volumes of infrequently accessed JSON logs in DynamoDB would incur significant costs for provisioned throughput and storage. Option C (Amazon RDS) is wrong because it is a relational database service designed for transactional workloads with structured data and frequent queries, not for storing semi-structured JSON logs at low cost; it would require schema management and incur higher per-GB storage costs compared to S3. Option D (Amazon Redshift) is wrong because it is a petabyte-scale data warehouse optimized for complex analytical queries on structured and semi-structured data, not for simple, cost-effective archival storage; using Redshift for rarely accessed logs would be over-provisioned and expensive due to its compute and storage costs.

900
MCQmedium

A data engineer is designing a data lake on Amazon S3. The data is ingested from multiple sources in Parquet format, and the schema evolves over time. Which approach allows querying the data with Amazon Athena while supporting schema evolution?

A.Use AWS Glue Data Catalog with crawlers to automatically update the table schema.
B.Define Hive-style partitions in Athena and manually update the schema.
C.Use S3 Select to query the data directly without a schema.
D.Use Amazon Redshift Spectrum with external tables and update the schema manually.
AnswerA

AWS Glue crawlers inspect Parquet data in Amazon S3 and populate the AWS Glue Data Catalog, which Athena queries. When new columns appear, re-running the crawler updates the table definition, satisfying the schema evolution requirement without manual DDL.

Why this answer

AWS Glue Data Catalog with crawlers automatically infers and updates the table schema as new Parquet files with evolving schemas are ingested into S3. This allows Athena to query the data using the latest schema without manual intervention, making it the ideal solution for schema evolution in a data lake.

Exam trap

The trap here is that candidates may think S3 Select or Redshift Spectrum can handle schema evolution automatically, but they lack the schema inference and versioning capabilities that AWS Glue Data Catalog provides for Athena.

How to eliminate wrong answers

Option B is wrong because manually updating the schema in Athena is error-prone and does not scale with frequent schema changes; Hive-style partitions alone do not handle schema evolution. Option C is wrong because S3 Select operates on individual objects and returns data in CSV/JSON format, not Parquet, and it does not support schema evolution or table-level queries across multiple files. Option D is wrong because Redshift Spectrum requires manual schema updates for external tables and is not designed for automatic schema evolution like AWS Glue Data Catalog.

Page 11

Page 12 of 18

Page 13