Courseiva

AWS Certified Data Engineer Associate DEA-C01 (DEA-C01) — Questions 151–225

1321 questions total · 18pages · All types, answers revealed

Page 2

Page 3 of 18

Page 4
151
MCQeasy

Refer to the exhibit. A Lambda function named 'IngestionProcessor' is failing. The engineer checks CloudWatch Logs and sees the log group exists but storedBytes is 0. Why might the logs show no data?

A.The Lambda execution role does not have permission to write logs to CloudWatch
B.The Lambda function is configured with a dead letter queue
C.The Lambda function has not been invoked yet
D.The log group is encrypted with a KMS key and the Lambda function lacks decrypt permission
AnswerA

Without logs:write permission on its execution role, Lambda cannot create log streams or put events, so the log group exists but storedBytes stays 0. The stem's constraint is an empty log group despite invocations, which points to missing CloudWatch Logs write authorisation rather than retention or filtering.

Why this answer

The Lambda execution role must have the `logs:CreateLogStream` and `logs:PutLogEvents` permissions to write logs to CloudWatch Logs. If the role lacks these permissions, the log group will be created (if it doesn't exist) but no log events will be written, resulting in `storedBytes` being 0. This is a common misconfiguration when the IAM policy does not include the necessary CloudWatch Logs actions.

Exam trap

The DEA-C01 exam often tests the distinction between log group creation (which requires `logs:CreateLogGroup`) and log writing (which requires `logs:CreateLogStream` and `logs:PutLogEvents`), leading candidates to confuse the existence of a log group with successful log delivery.

How to eliminate wrong answers

Option B is wrong because a dead letter queue (DLQ) is used to capture failed events for asynchronous invocations, not to prevent logs from being written; it does not affect CloudWatch Logs permissions. Option C is wrong because if the Lambda function had not been invoked, the log group would not exist at all; the presence of the log group with `storedBytes` of 0 indicates the function was invoked but failed to write logs. Option D is wrong because if the log group were encrypted with a KMS key and the Lambda function lacked decrypt permission, the function would fail with an access denied error when trying to write logs, but the log group would still show `storedBytes` as 0; however, the question states the log group exists and `storedBytes` is 0, which is consistent with missing write permissions, not KMS decryption issues (KMS errors would typically produce a different error message in CloudWatch).

152
Multi-Selectmedium

A data engineer is designing a data pipeline that processes PII data using AWS Glue and stores results in S3. Which TWO actions should be taken to protect the data? (Choose 2)

Select 2 answers
A.Use S3 default encryption with SSE-S3 for the output bucket.
B.Store database credentials in AWS Secrets Manager and reference them in Glue connections.
C.Enable S3 object deletion protection by setting a retention policy.
D.Configure AWS Glue to use a KMS key for encrypting data written to S3.
E.Use HTTPS for all data transfer between Glue and S3.
AnswersB, D

AWS Secrets Manager stores the database credentials encrypted and rotates them, and Glue connections reference the secret rather than embedding plaintext passwords in job scripts or catalog properties. This removes hard-coded credentials from the PII pipeline, satisfying the protection requirement.

Why this answer

Option B is correct because storing database credentials in AWS Secrets Manager and referencing them from Glue connections avoids hardcoding secrets in scripts or job parameters, enabling secure, auditable, and rotatable credential management for PII pipelines. Option D is correct because configuring AWS Glue to use a customer-managed KMS key for encrypting data written to S3 provides encryption at rest with control over key policies, rotation, and access auditing, which is appropriate for sensitive PII. Option A is not the best choice because SSE-S3 uses AWS-managed keys with less control and auditability than a KMS key, and the question asks for protective actions beyond default encryption.

Option C is incorrect because S3 object deletion protection via retention policies (Object Lock) addresses immutability/deletion, not the confidentiality of PII data being processed. Option E is not selected because HTTPS protects data in transit, but Glue-to-S3 traffic within AWS is already encrypted in transit by default, so it is not one of the two required protective actions for this scenario.

Exam trap

DEA-C01 often tests the distinction between encryption at rest with AWS-managed keys (SSE-S3) versus customer-managed KMS keys (SSE-KMS), and candidates frequently pick SSE-S3 as 'good enough' for PII, missing that compliance and key control requirements demand KMS; similarly, they may choose HTTPS (already default) as a security measure, overlooking that it only covers transit, not at-rest protection or credential management.

153
MCQhard

A data engineer is designing a data pipeline that ingests data from an on-premises system into Amazon S3 using AWS Transfer Family. The data must be encrypted at rest using a customer-managed key in AWS KMS. The S3 bucket policy must allow only encrypted connections. Which policy condition should be used?

A.aws:SecureTransport
B.kms:EncryptionContext
C.s3:x-amz-server-side-encryption-aws-kms-key-id
D.s3:x-amz-server-side-encryption
AnswerA

aws:SecureTransport is a boolean condition key that evaluates whether the request arrived over HTTPS. Setting it to false in a Deny statement blocks unencrypted connections, satisfying the bucket policy requirement that only encrypted transport be permitted to the Transfer Family endpoint.

Why this answer

The correct condition is 'aws:SecureTransport' because it ensures that requests to the S3 bucket are made over HTTPS (TLS), thereby enforcing encrypted connections. This condition key is a boolean that indicates whether the request used SSL/TLS. Setting it to 'true' in a bucket policy denies any non-encrypted (HTTP) requests.

Exam trap

DEA-C01 often tests the confusion between encryption in transit and encryption at rest, and candidates may incorrectly choose a condition related to server-side encryption headers instead of 'aws:SecureTransport' for enforcing encrypted connections.

How to eliminate wrong answers

Option B is wrong because 'kms:EncryptionContext' is used to enforce specific encryption context values for KMS operations, not to enforce encrypted connections to S3. Option C is wrong because 's3:x-amz-server-side-encryption-aws-kms-key-id' is used to require a specific KMS key for server-side encryption, but it does not enforce that the connection itself is encrypted; it only checks the header for the key ID. Option D is wrong because 's3:x-amz-server-side-encryption' checks for the presence of the encryption header (e.g., 'aws:kms'), but again, it does not enforce that the connection is encrypted; a request could be made over HTTP with the header set.

154
MCQeasy

A company is using Amazon Kinesis Data Firehose to deliver data to an Amazon S3 bucket. The data is delivered in 5-minute intervals. The company wants to reduce the delivery frequency to 1 minute to get data faster. Which parameter should be changed in the Firehose delivery stream configuration?

A.Reduce the buffer interval from 300 seconds to 60 seconds.
B.Increase the buffer size to trigger delivery sooner.
C.Enable dynamic partitioning to deliver data more frequently.
D.Enable compression to reduce data size and speed up delivery.
AnswerA

Firehose buffers incoming records and delivers when either the buffer size or buffer interval is reached. Lowering the interval from 300 to 60 seconds forces delivery every minute, satisfying the requirement for faster one-minute delivery.

Why this answer

Kinesis Data Firehose buffers incoming records and delivers them when either the buffer size (e.g., 5 MB) or the buffer interval (e.g., 300 seconds) is reached, whichever comes first. Reducing the buffer interval from 300 seconds to 60 seconds makes Firehose flush every minute, directly achieving the 1-minute delivery frequency the company wants.

Exam trap

The trap is confusing buffer size with buffer interval; candidates often think increasing buffer size speeds delivery, but only the interval controls the time-based flush, and the exam tests whether you know the two thresholds and their defaults.

How to eliminate wrong answers

Option B is wrong because increasing the buffer size makes Firehose wait for more data before delivering, which delays delivery rather than accelerating it. Option C is wrong because dynamic partitioning changes how records are grouped into S3 prefixes based on partition keys; it does not change the buffering interval and therefore does not control delivery frequency. Option D is wrong because compression reduces the size of delivered data and can affect how quickly the buffer size threshold is reached, but it does not set a time-based delivery cadence and is not the parameter for frequency.

155
MCQhard

A company uses AWS Lake Formation to manage access to a data lake in Amazon S3. A data engineer needs to grant a specific IAM role access to only the columns containing non-sensitive data in a table, while hiding columns with personally identifiable information (PII). The engineer has already registered the S3 bucket and the table in Lake Formation. What should the engineer do to meet this requirement?

A.Create a Lake Formation data filter that excludes the PII columns and grant SELECT permission on the filtered table to the IAM role.
B.Attach a bucket policy to the S3 bucket that allows the IAM role to read only specific objects.
C.Use an IAM policy that denies access to the S3 prefixes containing the PII columns.
D.Configure an AWS Glue job to create a new table that omits the PII columns and grant access to that table.
AnswerA

Lake Formation data filters allow column-level and row-level access control. By creating a data filter that excludes PII columns, the engineer can grant SELECT permission on that filtered view to the IAM role. This ensures the role can query only the non-sensitive columns. This is the correct approach because Lake Formation enforces these permissions at the table level without modifying the underlying data.

Why this answer

Lake Formation data filters enable column-level security by allowing you to exclude specific columns from a table. When you grant SELECT permission on a filtered table, the user can only access the included columns. This meets the requirement of hiding PII columns while providing access to non-sensitive data, without duplicating data or altering S3 objects.

Exam trap

The trap here is confusing S3 object-level permissions with column-level permissions, which are enforced by Lake Formation data filters.

156
MCQeasy

A company uses Amazon RDS for PostgreSQL. The data engineer needs to ensure that the database is automatically backed up and that backups are retained for 35 days. What is the simplest way to achieve this?

A.Use AWS Backup to schedule daily backups with a 35-day retention.
B.Enable automated backups with a retention period of 35 days in the RDS instance configuration.
C.Create a manual snapshot every day and delete them after 35 days using a script.
D.Enable automatic export of transaction logs to Amazon S3 and use S3 lifecycle policies.
AnswerB

Configuring automated backups on the RDS instance with a 35-day retention period uses RDS's native backup mechanism, which takes daily snapshots and retains them for the specified window. This is the simplest approach, requiring no custom scripting or external tooling.

Why this answer

Amazon RDS for PostgreSQL allows you to enable automated backups directly in the instance configuration. By setting the backup retention period to 35 days, RDS automatically performs daily snapshots and retains transaction logs for point-in-time recovery within that window. This is the simplest method because it requires no external services or custom scripting.

Exam trap

The trap here is that candidates may overcomplicate the solution by choosing AWS Backup (Option A) or manual scripting (Option C), not realizing that RDS native automated backups already provide the simplest, fully managed way to achieve the required retention period.

How to eliminate wrong answers

Option A is wrong because AWS Backup is an additional service that adds complexity and cost; RDS native automated backups already support retention up to 35 days without needing AWS Backup. Option C is wrong because manual snapshots require custom scripting to create and delete daily, which is not the simplest approach and does not provide automated point-in-time recovery. Option D is wrong because automatic export of transaction logs to S3 is not a native RDS feature for PostgreSQL; RDS handles transaction logs internally for point-in-time recovery, and using S3 lifecycle policies would not replace the need for automated backups.

157
MCQhard

A data engineer is designing a streaming pipeline using Amazon Kinesis Data Streams with a shard count of 10. The incoming data rate is 1 MB/second. The consuming application uses the Kinesis Client Library (KCL) with a single worker. What is the most likely performance bottleneck?

A.The Lambda function invoked by the stream has a cold start issue
B.The data stream has insufficient write capacity
C.The single KCL worker cannot process all shards in parallel
D.The shard count is too low to handle the data rate
AnswerC

KCL workers should be scaled to match shard count for parallel processing.

Why this answer

The Kinesis Client Library (KCL) uses a 1:1 mapping between shards and record processors by default. With 10 shards and only a single KCL worker, that worker must run all 10 record processors sequentially on a single host, creating a bottleneck. The worker cannot process records from multiple shards in parallel, so the throughput is limited by the single worker's processing capacity, not the stream's write capacity.

Exam trap

The trap here is that candidates often assume the bottleneck is on the write side (insufficient shards or write capacity) because they focus on the incoming data rate, but the question specifically tests the consumer-side limitation of a single KCL worker unable to parallelize across multiple shards.

How to eliminate wrong answers

Option A is wrong because Lambda cold starts are a potential issue only if the consuming application uses Lambda as a consumer, but the question specifies a KCL worker, not a Lambda function. Option B is wrong because the incoming data rate is 1 MB/second, and a single Kinesis shard supports up to 1 MB/second write capacity, so 10 shards provide 10 MB/second—far more than needed. Option D is wrong because the shard count of 10 is more than sufficient to handle the 1 MB/second data rate; the bottleneck is on the consumer side, not the stream's capacity.

158
MCQeasy

A company wants to ingest real-time clickstream data from a website into Amazon S3 with a maximum latency of 60 seconds. The data volume peaks at 500 MB/s. Which service should they use to buffer and deliver the data to S3?

A.Amazon Kinesis Data Firehose
B.Amazon Simple Queue Service (SQS)
C.Amazon Kinesis Data Streams
D.AWS Lambda
AnswerA

Kinesis Data Firehose buffers streaming records and delivers them to Amazon S3 in batches, with configurable buffer intervals that keep latency well under 60 seconds. It scales to handle 500 MB/s peaks without custom consumers, unlike Kinesis Data Streams, which requires separate delivery code.

Why this answer

Amazon Kinesis Data Firehose is the correct choice because it is designed to ingest streaming data, buffer it, and deliver it to destinations like Amazon S3 with configurable buffer intervals (e.g., 60 seconds) and buffer sizes (e.g., up to 128 MB). It can handle the peak throughput of 500 MB/s by automatically scaling, and it meets the maximum latency requirement of 60 seconds by flushing data to S3 based on time or size thresholds.

Exam trap

The trap here is that candidates often confuse Kinesis Data Streams (a real-time processing stream requiring custom consumers) with Kinesis Data Firehose (a fully managed delivery service), leading them to pick Data Streams for its real-time capabilities, even though Firehose is the correct choice for direct S3 delivery with minimal latency.

How to eliminate wrong answers

Option B (Amazon Simple Queue Service) is wrong because SQS is a message queue for decoupling application components, not a streaming buffer designed for high-throughput data delivery to S3; it lacks native integration to automatically write data to S3 with configurable latency. Option C (Amazon Kinesis Data Streams) is wrong because it is a real-time data streaming service that requires custom consumers (e.g., Lambda or Kinesis Client Library) to read and write data to S3, adding complexity and latency beyond the 60-second requirement; it does not natively buffer and deliver to S3. Option D (AWS Lambda) is wrong because Lambda is a serverless compute service for running code in response to events, not a buffer or delivery mechanism; it cannot handle sustained 500 MB/s ingestion without additional services and would require custom orchestration to meet latency goals.

159
MCQmedium

A data engineer needs to transform data in an AWS Glue job using a custom Python library that is not available by default. The library is packaged as a .whl file stored in Amazon S3. The engineer wants the Glue job to use this library without modifying the job script to install it at runtime. What should the engineer do?

A.Add the library to the Glue Data Catalog as a custom classifier.
B.Use a Glue development endpoint to install the library, then run the job there.
C.Place the .whl file in the Glue job's Python library path parameter (--extra-py-files).
D.Upload the .whl to the Glue job's script bucket and import it directly by filename.
AnswerC

The --extra-py-files job parameter allows specifying additional Python files or wheels stored in S3, which Glue makes available to the job without runtime installation. This directly satisfies the requirement to use a custom library without modifying the script to install it.

Why this answer

Specifying the wheel via the --extra-py-files job parameter makes the custom Python library available to the AWS Glue job without script changes, which is the correct way to include additional dependencies stored in Amazon S3.

Exam trap

The trap here is confusing catalog classifiers or dev endpoints with dependency injection, when only job parameters like --extra-py-files add libraries.

160
MCQmedium

A company is using Kinesis Data Firehose to deliver data to an S3 bucket. The delivery stream is failing with 'S3 bucket access denied' errors. The bucket policy allows the Firehose service principal. What could be the issue?

A.The S3 bucket is in a different VPC
B.The S3 bucket uses SSE-KMS and Firehose does not have KMS permissions
C.The S3 bucket name contains invalid characters
D.The IAM role assigned to Firehose lacks s3:PutObject permission
AnswerD

Correct. The IAM role assigned to the Firehose delivery stream must have the s3:PutObject permission to write objects to the S3 bucket. The bucket policy allowing the service principal is not sufficient; the role also needs the appropriate S3 action.

Why this answer

Even though the S3 bucket policy allows the Firehose service principal, Kinesis Data Firehose uses an IAM role to write data. This role must have the s3:PutObject permission. Without it, Firehose will receive an access denied error.

Option A is incorrect because VPC differences affect network connectivity, not IAM permissions. Option B is incorrect because SSE-KMS requires KMS permissions, but the error here is specifically about S3 access. Option C is incorrect because bucket name validation occurs during stream creation, not during data delivery.

Exam trap

Candidates often confuse the bucket policy and the IAM role permissions. The bucket policy allowing the service principal is necessary but not sufficient; the delivery role must also have s3:PutObject.

161
Multi-Selecthard

A data engineer is designing an Amazon Redshift data warehouse for a high-traffic analytics workload. The engineer needs to ensure fast query performance and minimize data movement. Which THREE design decisions should be made? (Choose THREE.)

Select 3 answers
A.Choose DISTSTYLE KEY for tables that are frequently joined.
B.Use the default distribution style for all tables.
C.Use DISTSTYLE ALL for all large fact tables.
D.Apply appropriate compression encodings to columns.
E.Define SORT KEYs on columns used in WHERE clauses.
AnswersA, D, E

DISTSTYLE KEY distributes rows according to a chosen column's values, so matching join keys co-locate on the same slice. For frequently joined tables, this eliminates cross-node data movement during joins, satisfying the requirement to minimise data movement and speed queries.

Why this answer

Option A is correct because choosing DISTSTYLE KEY on the columns used in frequent joins colocates matching rows on the same slice, so join processing avoids broadcasting or redistributing data across nodes. Option D is correct because applying appropriate compression encodings (for example AZ64, ZSTD, or LZO) reduces the amount of data read from disk and moved across the network, directly improving query performance. Option E is correct because defining SORT KEYs on columns used in WHERE clauses enables zone-map block skipping, so the query reads only the relevant blocks instead of scanning the whole table.

Option B is not appropriate because the default distribution style (AUTO) may not optimally colocate join keys, and relying on it for all tables can cause unnecessary data movement. Option C is wrong because DISTSTYLE ALL replicates the entire table to every node, which is suitable only for small dimension tables, not large fact tables, since it would multiply storage and load/redistribution overhead.

162
MCQmedium

A data engineer is monitoring an AWS Glue ETL job that processes data from an S3 bucket and writes to a Redshift table. The job completes successfully but takes longer than expected. The engineer notices that the job uses 10 DPUs and the data size is 500 GB. The job runs in standard mode. Which change would MOST reduce job duration?

A.Increase the number of DPUs to 20.
B.Use a smaller worker type like G.1X.
C.Change the output format from Parquet to CSV.
D.Reduce the number of partitions in the data.
AnswerA

Scaling DPUs adds parallel executors, so the 500 GB shuffle and Redshift write are processed concurrently rather than serially. With 10 DPUs the job is compute-bound in standard mode; doubling to 20 directly halves wall-clock duration, satisfying the requirement to most reduce job duration.

Why this answer

Increasing the number of DPUs from 10 to 20 allows the job to process data in parallel, reducing execution time. AWS Glue standard mode scales linearly with DPUs for ETL jobs that are not I/O bound. In this case, with 500 GB of data and 10 DPUs, the job is likely CPU-bound and can benefit from additional parallelism.

Option B is incorrect because using a smaller worker type (G.1X) reduces available memory and CPU, worsening performance. Option C is incorrect because changing the output from Parquet (columnar, compressed) to CSV (row-based, uncompressed) increases data size and I/O, slowing the job. Option D is incorrect because reducing partitions can cause data skew and reduce parallelism, increasing runtime.

Exam trap

Candidates may assume that increasing DPUs always helps, but for small datasets or I/O-bound jobs, diminishing returns occur. However, for large datasets like 500 GB in standard mode, increasing DPUs typically reduces duration linearly up to a point.

163
MCQhard

A company uses Amazon Kinesis Data Firehose to ingest log data from web servers into Amazon S3. The data is in JSON format and each record is approximately 2 KB. The delivery stream is configured to buffer incoming records for 60 seconds or 5 MB, whichever comes first. The company notices that the data in S3 is delayed by up to 5 minutes during peak hours. Which action would most effectively reduce the delivery latency?

A.Increase the buffer size to 10 MB to allow more records per delivery.
B.Decrease the buffer interval to 15 seconds.
C.Enable compression (GZIP) on the delivery stream.
D.Enable data transformation with AWS Lambda to convert JSON to Parquet.
AnswerB

Shorter buffer interval triggers more frequent deliveries, reducing latency.

Why this answer

The observed delay of up to 5 minutes during peak hours indicates that the buffer size threshold (5 MB) is rarely reached because each record is only ~2 KB, so the delivery stream relies on the buffer interval (60 seconds) to trigger delivery. By decreasing the buffer interval to 15 seconds, Kinesis Data Firehose will push data to S3 more frequently, directly reducing the maximum latency from 60 seconds to 15 seconds per batch, which eliminates the compounding delays caused by queuing during high-throughput periods.

Exam trap

The trap here is that candidates assume increasing buffer size or enabling compression will speed up delivery, but they fail to recognize that with small records, the buffer interval is the bottleneck, and only reducing that interval directly lowers latency.

How to eliminate wrong answers

Option A is wrong because increasing the buffer size to 10 MB would actually increase the time needed to fill the buffer, worsening the latency issue during peak hours when records are small and the buffer interval is the primary trigger. Option C is wrong because enabling GZIP compression reduces storage size and cost but does not affect the delivery frequency or buffer flush timing, so it has no impact on latency. Option D is wrong because converting JSON to Parquet via Lambda adds processing overhead and introduces additional latency from the transformation invocation, which would increase rather than reduce delivery delay.

164
MCQmedium

A company uses Amazon Redshift for analytics. The data engineer notices that queries are slow due to many small inserts. Which technique would improve write performance?

A.Use the COPY command to load data from Amazon S3.
B.Define DISTKEY and SORTKEY on the table.
C.Increase the number of nodes in the cluster.
D.Configure workload management (WLM) queues.
AnswerA

Row-by-row inserts force Redshift to commit many small transactions, which is inefficient. COPY loads large batches in parallel from Amazon S3 directly into the cluster, dramatically improving write throughput and reducing the overhead of many small inserts.

Why this answer

The COPY command is the recommended way to load data into Amazon Redshift because it performs bulk inserts in parallel across all nodes, leveraging the cluster's distributed architecture. Small individual INSERT statements cause high overhead due to transaction logging and commit processing, leading to slow write performance. By loading data from Amazon S3 using COPY, you bypass these per-row overheads and achieve optimal throughput.

Exam trap

The trap here is that candidates often confuse performance tuning for reads (DISTKEY/SORTKEY) or general scaling (adding nodes) with the specific write performance bottleneck caused by many small inserts, overlooking the COPY command as the primary solution for bulk data loading.

How to eliminate wrong answers

Option B is wrong because defining DISTKEY and SORTKEY improves query read performance by optimizing data distribution and sort order, but does not directly address the write performance issue caused by many small inserts. Option C is wrong because increasing the number of nodes adds compute and storage capacity, but does not solve the fundamental problem of per-insert overhead; small inserts will still be slow on a larger cluster. Option D is wrong because configuring workload management (WLM) queues manages concurrency and prioritizes queries, but does not reduce the overhead of individual small INSERT statements.

165
MCQhard

A company uses AWS Glue to process data from multiple S3 buckets. The Glue job runs daily and reads data from a bucket that contains millions of small files (each < 1 MB). The job has been running for hours and is often close to the 8-hour timeout limit. Which optimization would MOST reduce the job's runtime?

A.Pre-process the data to consolidate small files into larger files before the Glue job.
B.Convert the source data from CSV to Parquet format.
C.Increase the number of DPUs allocated to the Glue job.
D.Use a larger Spark shuffle partition size.
AnswerA

Millions of sub-1 MB files force Glue to open each object individually, so per-file overhead dominates runtime. Consolidating them into larger files before the job drastically cuts the number of read operations and task overhead.

Why this answer

Consolidating millions of small files (<1 MB each) into larger files before the Glue job is the most impactful optimization because Spark and Glue incur massive per-file overhead: each file requires a separate S3 GET, task scheduling, and metadata operation. With millions of tiny files, the job spends most of its time on I/O and task overhead rather than actual processing. Merging them into a few hundred MB files dramatically reduces task count and runtime.

Exam trap

The trap is reaching for the 'more resources' answer (more DPUs) — candidates assume scaling compute fixes slow jobs, but the exam tests whether you recognize that small-file I/O overhead is the true bottleneck.

How to eliminate wrong answers

Option B is wrong because converting CSV to Parquet improves compression and scan efficiency but does not solve the small-file problem — millions of tiny Parquet files still cause the same per-file overhead. Option C is wrong because adding DPUs increases parallelism but cannot overcome the serialization bottleneck of millions of file-open operations; it also increases cost without addressing the root cause. Option D is wrong because a larger shuffle partition size affects post-shuffle aggregation, not the initial read of millions of small files, so it does not reduce the dominant cost.

166
MCQeasy

A data engineer is configuring an AWS Glue ETL job that processes data from an Amazon S3 bucket. The security team requires that all data in transit between AWS Glue and Amazon S3 be encrypted using TLS. The engineer needs to ensure that the Glue job uses HTTPS endpoints when reading from and writing to S3. Which action should the engineer take?

A.Ensure that the Glue job's IAM role has permissions to use HTTPS, and no further configuration is needed.
B.No action is needed; AWS Glue uses HTTPS for S3 by default.
C.Add the --enable-s3-ssl parameter to the Glue job's job parameters.
D.Configure the Glue connection to use an S3 endpoint with the https:// prefix.
AnswerB

AWS Glue uses the AWS SDK for S3 operations, and the SDK defaults to HTTPS endpoints. Therefore, data in transit between Glue and S3 is encrypted with TLS by default. No additional configuration is required to meet the requirement. This is the correct and simplest answer.

Why this answer

AWS Glue uses the AWS SDK to interact with Amazon S3, and the SDK uses HTTPS endpoints by default. This means data in transit is encrypted with TLS without any additional configuration. There is no need for special job parameters, connection settings, or IAM permissions to enable TLS.

The correct action is to recognize that no action is needed.

Exam trap

The trap here is assuming that you must enable SSL/TLS explicitly for Glue-S3 communication, when it is already enabled by default.

167
MCQeasy

A company uses Amazon DynamoDB as its primary data store for a web application. The application experiences high latency during peak hours. The data engineer notices that the table has a large number of items with the same partition key. Which DynamoDB feature should the engineer use to improve performance?

A.Redesign the partition key to use a composite key that includes a timestamp or random suffix.
B.Enable DynamoDB Accelerator (DAX) to cache read requests.
C.Create a global table to replicate data across multiple Regions.
D.Enable auto scaling on the table to increase write capacity.
AnswerA

Many items sharing one partition key concentrate reads on a single partition, creating a hot partition that throttles throughput and raises latency. A composite key adding a timestamp or random suffix spreads items across partitions, distributing load and restoring performance.

Why this answer

The high latency is caused by a hot partition, where many items share the same partition key, overwhelming a single DynamoDB partition. Redesigning the partition key to include a timestamp or random suffix distributes the workload evenly across partitions, improving throughput and reducing latency. This directly addresses the root cause of the performance issue.

Exam trap

The trap here is that candidates often confuse caching solutions (DAX) or scaling mechanisms (auto scaling) with the need to fix the data model itself, which is the only way to resolve a hot partition caused by a skewed partition key.

How to eliminate wrong answers

Option B is wrong because DynamoDB Accelerator (DAX) caches read requests, which can reduce read latency but does not solve the underlying hot partition issue caused by skewed write or read traffic on a single partition key. Option C is wrong because creating a global table replicates data across multiple Regions for disaster recovery or low-latency global access, but it does not distribute load within a single table's partitions. Option D is wrong because enabling auto scaling increases the table's provisioned capacity, but if the workload is concentrated on one partition, the partition's throughput limit (3000 RCU or 1000 WCU) will still be exceeded, causing throttling and high latency.

168
Multi-Selecthard

A data engineer is designing a data pipeline that ingests data from multiple sources into Amazon S3, then processes it with AWS Glue and loads it into Amazon Redshift. Which THREE practices should be implemented to ensure data quality?

Select 3 answers
A.Implement data validation checks at the ingestion stage
B.Use AWS Glue DataBrew for data profiling and schema enforcement
C.Compress data files to reduce storage costs
D.Use manual sampling to check data quality periodically
E.Set up Amazon CloudWatch alarms for pipeline failures and data anomalies
AnswersA, B, E

Validating records as they land in Amazon S3 catches malformed, incomplete or out-of-range data before Glue transforms it, preventing corrupt rows from propagating into Redshift. This satisfies the requirement to ensure data quality across the multi-source ingestion pipeline at its earliest stage.

Why this answer

Option A is correct because validating data at ingestion (for example, checking types, nulls, ranges, and referential integrity before writing to Amazon S3) prevents corrupt or malformed records from propagating downstream into Glue ETL jobs and Redshift tables. Option B is correct because AWS Glue DataBrew provides visual data profiling, column statistics, and built-in transformations/rules that can enforce schema consistency and detect anomalies before loading into Redshift. Option E is correct because Amazon CloudWatch alarms on Glue job metrics, Redshift load events, and custom anomaly metrics enable proactive detection of pipeline failures and unexpected data patterns, which is essential for ongoing data quality assurance.

Option C is not correct because compressing files reduces storage and I/O costs but does not validate or improve data accuracy, completeness, or consistency. Option D is not correct because manual sampling is ad hoc, non-scalable, and not a reliable or automated data quality control compared with systematic validation, profiling, and monitoring.

Exam trap

DEA-C01 often tests the confusion between cost optimisation (compression) and data quality practices, so candidates pick compression or manual sampling instead of automated validation, profiling, and monitoring.

169
MCQeasy

A company needs to store JSON documents that are accessed by a key-value pattern. The data is 500 GB and requires single-digit millisecond latency. Which AWS database is most suitable?

A.Amazon Redshift
B.Amazon DynamoDB
C.Amazon Neptune
D.Amazon RDS for MySQL
AnswerB

DynamoDB is a fully managed key-value and document store delivering consistent single-digit millisecond latency at any scale, comfortably handling 500 GB of JSON documents accessed by primary key. Relational or object stores cannot meet that latency and access pattern.

Why this answer

Amazon DynamoDB is the most suitable choice because it is a fully managed NoSQL key-value and document database that delivers single-digit millisecond latency at any scale, making it ideal for storing and retrieving JSON documents via a key-value access pattern. It supports document data types natively and can handle 500 GB of data efficiently with consistent low-latency performance.

Exam trap

The trap here is that candidates may choose Amazon RDS for MySQL because they associate JSON documents with relational databases, overlooking that DynamoDB is purpose-built for key-value and document workloads with guaranteed single-digit millisecond latency, while RDS introduces schema rigidity and higher latency for this pattern.

How to eliminate wrong answers

Option A is wrong because Amazon Redshift is a petabyte-scale data warehouse optimized for complex analytical queries using SQL, not for low-latency key-value lookups on JSON documents; it incurs higher latency and is not designed for single-digit millisecond access patterns. Option C is wrong because Amazon Neptune is a graph database optimized for highly connected data and graph queries (e.g., using Gremlin or SPARQL), not for simple key-value access to JSON documents; it would add unnecessary complexity and cost. Option D is wrong because Amazon RDS for MySQL is a relational database that requires predefined schemas and is not optimized for key-value access patterns on JSON documents; while it can store JSON, it lacks the native partitioning and low-latency throughput of DynamoDB for this use case.

170
MCQmedium

A data engineer is ingesting records from Amazon Kinesis Data Streams into Amazon S3 using AWS Lambda as the consumer. Each stream shard delivers up to 1,000 records per second, and the Lambda function writes each record as an individual small object, causing many tiny S3 files and high PUT costs. The engineer wants fewer, larger objects while keeping near-real-time delivery. What should the engineer do?

A.Replace the Lambda consumer with Amazon Data Firehose, which buffers records and delivers batched objects to S3.
B.Enable enhanced fan-out on the stream so consumers get dedicated throughput per shard.
C.Increase the Lambda function memory so each invocation can write more records per second.
D.Add more shards to the Kinesis data stream to increase parallelism.
AnswerA

Amazon Data Firehose reads from Kinesis Data Streams and buffers incoming records by size and time before delivering consolidated objects to S3. This directly produces fewer, larger files and cuts PUT request costs, while its configurable buffer interval preserves near-real-time delivery. It removes the need for custom batching logic in Lambda.

Why this answer

The inefficiency comes from writing one S3 object per Kinesis record. Amazon Data Firehose natively buffers records by configurable size and time windows and delivers them as consolidated objects, which reduces file count and PUT costs while still meeting near-real-time needs. Tuning Lambda memory, adding shards, or enabling enhanced fan-out changes throughput and parallelism but not the per-record write pattern.

Exam trap

The trap here is tuning Lambda or stream capacity to fix a file-size problem, when the actual fix is introducing a buffering delivery layer that batches records.

171
MCQmedium

A company needs to monitor and record all changes to IAM policies in their AWS account. Which AWS service should be used?

A.Amazon CloudWatch Logs
B.Amazon GuardDuty
C.AWS CloudTrail
D.AWS IAM Access Analyzer
AnswerC

AWS CloudTrail captures API activity, including IAM policy creation, modification and deletion, recording each event with the caller identity, timestamp and request parameters. This satisfies the requirement to monitor and record all IAM policy changes, since CloudTrail logs management events by default and delivers them to S3 or CloudWatch for auditing.

Why this answer

AWS CloudTrail records API activity in an AWS account, including all IAM policy changes (CreatePolicy, AttachRolePolicy, PutRolePolicy, etc.), capturing who made the change, when, from which IP, and with what parameters. It is the authoritative audit service for governance, compliance, and security analysis of control-plane actions. CloudWatch Logs, GuardDuty, and IAM Access Analyzer serve monitoring, threat detection, and permission analysis—not comprehensive API change recording.

Exam trap

DEA-C01 often tests the overlap between CloudTrail, CloudWatch, and GuardDuty—candidates pick GuardDuty because it 'detects' IAM issues, but the question asks for the service that records all changes, which is CloudTrail.

How to eliminate wrong answers

Option A is wrong because Amazon CloudWatch Logs stores and analyzes log data (including CloudTrail logs if forwarded), but it is not the service that natively records IAM API changes. Option B is wrong because Amazon GuardDuty is a threat-detection service that analyzes CloudTrail, VPC Flow Logs, and DNS logs for malicious activity; it does not itself record all IAM policy changes. Option D is wrong because IAM Access Analyzer identifies resources shared with external entities and validates policies against best practices; it does not provide a complete audit trail of IAM policy modifications.

172
Multi-Selectmedium

A company uses Amazon Kinesis Data Firehose to ingest data into an S3 bucket. The data is in JSON format and the team wants to convert it to Parquet before storage. Which TWO configurations are required?

Select 2 answers
A.Use Kinesis Data Analytics to transform data to Parquet.
B.Create a Glue table with the schema of the data.
C.Configure a Lambda function to convert data on the fly.
D.Set up an Athena table to read the data.
E.Enable data format conversion in Firehose and set Output format to Parquet.
AnswersB, E

Firehose's Parquet conversion relies on the AWS Glue Data Catalog to resolve the source schema, so a Glue table describing the JSON structure is mandatory. Without it, Firehose cannot map incoming records to Parquet columns, and the conversion configuration fails.

Why this answer

Option B is correct because Firehose data format conversion relies on the AWS Glue Data Catalog: you must create a Glue table (with the appropriate schema and SerDe) that Firehose references so it knows how to interpret the incoming JSON records. Option E is correct because the actual conversion is enabled in the Firehose delivery stream configuration by turning on data format conversion and setting the output format to Parquet (with the Glue table as the schema source). Option A is not required because Kinesis Data Analytics is for SQL/Flink stream processing, not for Firehose's built-in format conversion.

Option C is not required because Firehose performs the JSON-to-Parquet conversion natively via Glue, so a custom Lambda transformation is unnecessary. Option D is not required because Athena is a query service for reading data in S3, not a prerequisite for converting it during ingestion.

Exam trap

DEA-C01 often tests whether candidates know that Firehose Parquet conversion is a native two-part configuration (Glue table + Firehose setting) rather than something requiring Lambda, Glue ETL, or Kinesis Data Analytics.

173
MCQhard

A company uses AWS Glue DataBrew for data preparation. The data source is an S3 bucket with millions of small CSV files (each < 1 MB). The DataBrew project takes a long time to load the sample data. What is the most likely cause and solution?

A.Use Amazon Athena to query the data instead of DataBrew
B.The DataBrew job is under-provisioned; increase the number of DPUs
C.The large number of small files causes S3 LIST overhead; concatenate files into larger files
D.Use AWS Glue ETL instead of DataBrew for this volume
AnswerC

DataBrew must LIST and open each object individually, so millions of sub-1 MB files impose heavy S3 LIST and per-request overhead during sampling. Concatenating them into larger files reduces request count and dramatically speeds up sampling.

Why this answer

DataBrew loads a sample of the data by listing objects in the S3 bucket. With millions of small CSV files, the S3 LIST API call becomes a bottleneck because each list operation has a 1000-object limit per response, requiring multiple paginated requests. Concatenating the small files into larger files reduces the number of objects, dramatically decreasing LIST overhead and speeding up sample loading.

Exam trap

The DEA-C01 exam often tests the misconception that increasing DPUs or switching to a different AWS service will fix performance issues, when the real root cause is S3's small-file overhead and the LIST API's pagination limit.

How to eliminate wrong answers

Option A is wrong because Athena is a query engine, not a data preparation tool; it would still suffer from the same small-file overhead when reading data, and it does not solve the DataBrew sample loading issue. Option B is wrong because DataBrew projects do not use DPUs for sample loading; DPUs are only relevant for running DataBrew jobs (recipes), and the bottleneck here is S3 LIST latency, not compute capacity. Option D is wrong because switching to Glue ETL would not inherently solve the small-file problem; Glue ETL also incurs overhead from listing and processing many small files, and the question specifically asks about DataBrew sample loading, not ETL job performance.

174
Multi-Selecthard

A company uses Amazon Redshift for its data warehouse and needs to enforce column-level security on sensitive columns. Which TWO approaches can achieve this?

Select 2 answers
A.Apply an S3 bucket policy to the underlying data files.
B.Create views that expose only non-sensitive columns and grant access to the views.
C.Use Redshift Spectrum to query external tables and restrict columns via the external schema.
D.Use Redshift column-level security to grant or revoke permissions on specific columns.
E.Use Redshift row-level security policies to restrict column access.
AnswersB, D

Views restrict the projection to non-sensitive columns, so grantees query only exposed attributes; the underlying sensitive columns remain inaccessible through the view. This delivers column-level security by omission, satisfying the requirement without granting base-table access.

Why this answer

Option B is correct because creating views that expose only non-sensitive columns and granting users access to those views (rather than the base tables) is a standard Redshift pattern for column-level security: users query the view and cannot see the restricted columns. Option D is correct because Amazon Redshift natively supports column-level access control via GRANT and REVOKE on individual columns (for example, GRANT SELECT(col1, col2) ON table TO user), which directly enforces column-level security. Option A is wrong because an S3 bucket policy governs access to objects in S3, not to columns within Redshift tables, and does not control SQL-level column visibility.

Option C is wrong because Redshift Spectrum external schemas and tables do not provide a column-restriction mechanism for enforcing column-level security on Redshift data. Option E is wrong because Redshift row-level security (RLS) policies filter which rows a user can see, not which columns they can access.

Exam trap

DEA-C01 often tests the confusion between row-level security (filters rows) and column-level security (restricts columns), and the misconception that S3 bucket policies or Spectrum external schemas can enforce column-level access inside Redshift tables.

175
Multi-Selecteasy

Which TWO methods can be used to enforce least-privilege access to an Amazon S3 bucket? (Choose two.)

Select 2 answers
A.Use IAM policies to grant specific permissions to users and roles.
B.Set bucket ACLs to allow full control to the bucket owner only.
C.Use an S3 bucket policy that explicitly denies actions not required.
D.Configure a VPC endpoint to restrict access to the bucket.
E.Generate pre-signed URLs for all access.
AnswersA, C

IAM policies allow granular permissions.

Why this answer

IAM policies allow you to grant granular, specific permissions to individual users and roles, adhering to the principle of least privilege by explicitly allowing only the actions required. This avoids granting broad or default permissions, ensuring that each identity has only the access necessary for its function.

Exam trap

The trap here is that candidates often confuse network-level controls (like VPC endpoints) with identity-based access controls, or they mistakenly think that granting full control to the owner is a form of least privilege, when in fact it violates the principle by providing excessive permissions.

176
MCQmedium

A company wants to centrally manage encryption keys for multiple AWS services and automatically rotate them every year. Which AWS service should be used?

A.AWS CloudHSM
B.AWS Certificate Manager (ACM)
C.AWS Secrets Manager
D.AWS Key Management Service (KMS)
AnswerD

AWS Key Management Service provides customer managed keys with automatic annual rotation, satisfying the yearly rotation requirement. It integrates natively across multiple AWS services, enabling centralised key management. Unlike CloudHSM, which offers dedicated hardware but no built-in automatic rotation, KMS delivers the managed rotation and multi-service integration the scenario demands.

Why this answer

AWS Key Management Service (KMS) is designed to centrally manage encryption keys for multiple AWS services and supports automatic annual rotation of customer master keys. It integrates with services like S3, EBS, and RDS, making it the correct choice.

Exam trap

DEA-C01 often tests the distinction between KMS and CloudHSM; candidates may choose CloudHSM for key management but overlook the automatic rotation and integration features of KMS.

How to eliminate wrong answers

Option A is wrong because AWS CloudHSM provides dedicated hardware security modules but does not offer automatic key rotation and is not centrally managed across services in the same way as KMS. Option B is wrong because AWS Certificate Manager (ACM) manages SSL/TLS certificates, not encryption keys for data at rest. Option C is wrong because AWS Secrets Manager is for managing secrets like database credentials, not for general encryption key management with rotation.

177
MCQmedium

A team uses Amazon Redshift for analytics. They notice that some queries are slow and the system shows high disk usage. The team wants to improve query performance without adding more nodes. Which action should they take first?

A.Run the VACUUM and ANALYZE commands on the tables.
B.Enable compression on all columns.
C.Redistribute the tables by changing the distribution key to a column with high cardinality.
D.Modify the workload management (WLM) queue to increase concurrency.
AnswerA

VACUUM reclaims space from deleted rows and re-sorts data, while ANALYZE updates table statistics used by the query planner. Both address the high disk usage and slow queries without adding nodes, making them the correct first action.

Why this answer

VACUUM and ANALYZE are the first actions to take because they reclaim disk space from deleted rows and update table statistics, which can significantly improve query performance and reduce disk usage. Option B is incorrect because enabling compression on all columns is not a first step; compression is typically set during table creation and may not address current disk usage issues. Option C is incorrect: redistributing tables by changing the distribution key requires recreating the table, which is an invasive operation and not the first troubleshooting step.

Option D is incorrect: modifying the WLM queue affects concurrency management, not disk usage or query performance directly related to high disk usage.

178
Multi-Selectmedium

Which TWO actions can help improve query performance in Amazon Redshift? (Choose two.)

Select 2 answers
A.Use appropriate sort keys for tables.
B.Disable SSL encryption for connections.
C.Use VARCHAR instead of CHAR for fixed-length strings.
D.Apply compression encodings to columns.
E.Increase the number of nodes in the cluster.
AnswersA, D

Sort keys help the query optimizer scan less data.

Why this answer

Defining appropriate sort keys in Amazon Redshift enables the query optimizer to use zone maps to skip irrelevant data blocks during table scans, significantly reducing the amount of data read from disk. Sort keys also improve the effectiveness of merge joins and the performance of range-restricted queries by physically co-locating rows with similar sort key values on disk.

Exam trap

The trap here is that candidates often assume scaling out (adding nodes) always speeds up individual queries, but in Redshift, query performance is more dependent on data layout (sort keys, distribution, compression) than on cluster size, and adding nodes primarily benefits concurrent workloads rather than single-query latency.

179
MCQhard

A data engineering team is designing a data lake on Amazon S3. They need to store raw data in its original format and transformed data in Parquet. The data is accessed by multiple analytics services, including Amazon Athena and Amazon Redshift Spectrum. Compliance requirements mandate that all data be encrypted at rest with AWS KMS and that the encryption keys be rotated every 90 days. Which S3 bucket configuration meets these requirements?

A.Use SSE-KMS with a customer-managed KMS key that has automatic key rotation enabled.
B.Use SSE-C with client-managed keys and rotate them manually.
C.Use a bucket policy to enforce encryption and rely on default S3 encryption.
D.Use SSE-S3 with default encryption enabled.
AnswerA

SSE-KMS with automatic rotation meets compliance requirements.

Why this answer

SSE-KMS with a customer-managed KMS key allows you to implement custom key rotation, such as every 90 days, by creating new keys and updating the bucket policy or key alias. AWS KMS automatic key rotation for customer-managed keys occurs yearly, not every 90 days, but you can achieve a 90-day rotation schedule manually or through automation (e.g., AWS Lambda). SSE-C requires manual key management and does not integrate with AWS services like Amazon Athena.

SSE-S3 does not support configurable rotation, and the default encryption option (C) does not meet compliance if rotation is required.

Exam trap

The trap is that candidates assume SSE-S3 or default encryption meets the 90-day rotation requirement because AWS rotates keys automatically, but they overlook that SSE-S3 key rotation is not configurable—it follows AWS-managed rotation, which is not guaranteed every 90 days. SSE-KMS with a customer-managed key is the only option that allows a custom rotation schedule, even though it may require additional automation beyond the automatic yearly rotation.

How to eliminate wrong answers

Option B is wrong because SSE-C requires you to manage and rotate encryption keys client-side, which adds operational overhead and does not integrate with AWS KMS for automated rotation; manual rotation every 90 days is possible but not automated, and it violates the requirement to use AWS KMS. Option C is wrong because relying on default S3 encryption (SSE-S3) uses S3-managed keys that cannot be rotated on a 90-day schedule; AWS rotates SSE-S3 keys annually, but you have no control over the rotation frequency. Option D is wrong because SSE-S3 does not support customer-controlled key rotation; it uses S3-managed keys with automatic rotation by AWS, but the rotation period is not configurable and does not meet the 90-day requirement.

180
MCQmedium

A data engineer needs to ensure that all data stored in an Amazon S3 bucket is encrypted at rest using a customer managed key in AWS KMS. The engineer also needs to enforce that any attempt to upload an object without specifying the correct KMS key is denied. Which combination of actions should the engineer take?

A.Configure the bucket to use SSE-KMS with the customer managed key and attach an IAM policy to all users that allows only s3:PutObject with the correct encryption header.
B.Enable default encryption with SSE-S3 and use AWS KMS grants to restrict access to the customer managed key.
C.Use an S3 Object Lambda access point to encrypt objects on upload with the customer managed key.
D.Enable default encryption on the bucket with SSE-KMS using the customer managed key, and add a bucket policy that denies s3:PutObject requests that do not include the correct x-amz-server-side-encryption header.
AnswerD

Default encryption ensures all objects are encrypted with the specified KMS key, but it does not prevent uploads with a different encryption method. A bucket policy that denies PutObject requests lacking the correct encryption header enforces the use of the specific KMS key. Together, they meet the requirement.

Why this answer

To enforce encryption with a specific customer managed KMS key, enable default SSE-KMS encryption on the bucket and add a bucket policy that denies PutObject requests unless they include the correct x-amz-server-side-encryption header. This ensures all objects are encrypted with the designated key and prevents uploads with other encryption methods.

Exam trap

The trap here is assuming that enabling default encryption alone enforces the use of a specific KMS key, when in fact uploads can override the default with a different encryption method unless a bucket policy explicitly denies it.

181
Multi-Selecthard

A company is using AWS KMS to encrypt data in Amazon S3. The security team wants to ensure that only specific IAM roles can decrypt the data. Which TWO steps should the data engineer take? (Choose two.)

Select 2 answers
A.Use the default AWS managed KMS key for S3 (aws/s3)
B.Use SSE-S3 encryption instead of KMS
C.Create a customer-managed KMS key with a key policy that grants kms:Decrypt only to the allowed IAM roles
D.Add an IAM policy to the role that requires MFA for kms:Decrypt
E.Configure the S3 bucket to use SSE-KMS with the customer-managed key
AnswersC, E

A customer-managed KMS key lets you attach a key policy restricting kms:Decrypt to named IAM roles, satisfying the stem's requirement that only specific roles decrypt. AWS-managed keys use fixed policies you cannot edit, so they cannot enforce this restriction.

Why this answer

Option C is correct because a customer-managed KMS key allows you to define a key policy that explicitly grants kms:Decrypt only to the specific IAM roles, which is the core mechanism for restricting decryption permissions in KMS. Option E is correct because the S3 bucket must be configured to use SSE-KMS with that customer-managed key; otherwise, S3 would use a different key (such as the default aws/s3 key) and the key policy restriction would not apply to the objects. Option A is incorrect because the default AWS managed key aws/s3 has a key policy managed by AWS that grants broad permissions to the account, so it cannot be scoped to only specific IAM roles.

Option B is incorrect because SSE-S3 uses AES-256 with S3-managed keys and does not involve KMS or IAM role-based decrypt permissions at all. Option D is incorrect because requiring MFA for kms:Decrypt adds an authentication condition but does not by itself limit decryption to specific IAM roles, and MFA is not the mechanism that enforces role-level restriction.

182
MCQeasy

A data engineer needs to run a transformation that processes semi-structured JSON records already stored in Amazon S3 and write the results back to S3 in Parquet format. The team prefers a serverless, Apache Spark-based approach with minimal infrastructure management and wants to use the AWS Glue Data Catalog for metadata. Which approach should the engineer use?

A.Run an AWS Glue crawler to convert the JSON files to Parquet automatically.
B.Use AWS Lambda with pandas to convert each JSON file to Parquet.
C.Launch an Amazon EMR cluster with Spark and submit the job manually.
D.Create an AWS Glue ETL job using the Spark engine and write output with the Glue Parquet writer.
AnswerD

AWS Glue ETL jobs run on a serverless Apache Spark environment, so the engineer avoids managing clusters while gaining Spark's transformation capabilities. Reading JSON from S3, transforming it, and writing Parquet through the Glue Parquet writer produces columnar output optimized for analytics, and the job integrates natively with the Glue Data Catalog for schema and table metadata.

Why this answer

AWS Glue ETL jobs provide a serverless Apache Spark runtime that reads JSON from Amazon S3, applies transformations, and writes Parquet using the Glue Parquet writer. This combination satisfies the serverless and Spark-based preferences while integrating with the AWS Glue Data Catalog for table metadata. EMR adds cluster management overhead, and Lambda or a crawler cannot perform the required rewrite.

Exam trap

The trap here is confusing the role of an AWS Glue crawler, which only catalogs schemas, with an AWS Glue ETL job, which actually transforms and rewrites data.

183
MCQeasy

A data engineer is using AWS Database Migration Service (AWS DMS) to migrate an on-premises Oracle database to Amazon Aurora PostgreSQL. The migration is running but the engineer observes that some tables are not being replicated. The DMS task logs show no errors, and the task status is 'Running'. Which action should the engineer take to identify the missing tables?

A.Check the DMS task's table mappings to ensure all required tables are included in the selection rules.
B.Enable CloudWatch Logs for the DMS task and look for warnings about table selection.
C.Increase the DMS replication instance size to handle more tables concurrently.
D.Restart the DMS task with the 'Reload target' option to force a full load of all tables.
AnswerA

In AWS DMS, table mappings define which tables are included or excluded from the migration. If some tables are missing, it is likely that they are not covered by the selection rules. Reviewing and updating the table mappings to include the missing tables will ensure they are replicated. This is a common oversight when setting up DMS tasks.

Why this answer

Table mappings in AWS DMS determine which tables are migrated. If tables are missing and no errors are present, the most likely cause is that the tables are not included in the selection rules. Reviewing and correcting the table mappings is the direct way to ensure all required tables are replicated.

Exam trap

The trap here is assuming that a running task with no errors is replicating all tables, but table mappings may exclude some tables by design.

184
MCQmedium

A company wants to monitor and alert on unauthorized API calls in their AWS account. Which AWS service should be used to detect and notify on such events?

A.Amazon GuardDuty and AWS Security Hub
B.Amazon VPC Flow Logs and Amazon CloudWatch Logs
C.AWS Config and AWS Systems Manager
D.AWS CloudTrail and Amazon CloudWatch Events
AnswerD

CloudTrail records every API call as a management event, capturing the caller identity and whether it was authorised. CloudWatch Events (EventBridge) then matches those unauthorised-call patterns via rules and triggers notifications, satisfying the requirement to both detect and alert on unauthorised API activity.

Why this answer

D is correct because AWS CloudTrail records all API calls in the AWS account, and Amazon CloudWatch Events (or EventBridge) can be configured with rules to detect specific API calls (e.g., unauthorized actions) and trigger notifications. Option A is incorrect because Amazon GuardDuty and AWS Security Hub are threat detection and security management services, not primarily for monitoring all API calls. Option B is incorrect because Amazon VPC Flow Logs capture network traffic metadata, not API calls.

Option C is incorrect because AWS Config monitors resource configuration changes, not API calls.

Exam trap

Candidates often assume GuardDuty is the go-to for API call monitoring, but GuardDuty focuses on threat detection, not comprehensive API logging. CloudTrail is the correct service for logging all API calls.

185
MCQhard

A data engineer runs the describe-stream command and sees the output above. The stream has a retention period of 24 hours. The engineer needs to ensure that consumers can replay data for up to 7 days. Which action is required?

A.Increase the number of shards to allow more data storage.
B.Delete the stream and recreate it with a longer retention period.
C.Use the IncreaseStreamRetentionPeriod API to set retention to 168 hours.
D.Create new consumer applications that read from the stream.
AnswerC

The API can increase retention up to 365 days.

Why this answer

The describe-stream output shows a retention period of 24 hours, but the requirement is to allow consumers to replay data for up to 7 days (168 hours). Amazon Kinesis Data Streams supports modifying the retention period dynamically without recreating the stream, using the IncreaseStreamRetentionPeriod API or the update-shard-count command. Option C correctly uses this API to set retention to 168 hours, which is the maximum supported retention period for Kinesis Data Streams.

Exam trap

The trap here is that candidates often confuse shard count with storage capacity, assuming that more shards allow more data to be stored, when in fact shards only control throughput and retention is a separate, configurable parameter.

How to eliminate wrong answers

Option A is wrong because increasing the number of shards increases the stream's throughput capacity (read/write operations per second), not the data retention period; shards do not affect how long data is stored. Option B is wrong because deleting and recreating the stream is unnecessary and disruptive; Kinesis allows you to modify the retention period on an existing stream without data loss or downtime. Option D is wrong because creating new consumer applications does not change the retention period; consumers can only replay data within the existing retention window, so they would still be limited to 24 hours of replay.

186
MCQeasy

A data engineer needs to ingest streaming data from a social media API into Amazon S3 for batch analytics. The data arrives at a rate of 500 records per second. Which service should be used to capture the stream?

A.Amazon Simple Notification Service (SNS)
B.Amazon Simple Queue Service (SQS)
C.Amazon Kinesis Data Streams
D.Amazon MQ
AnswerC

Kinesis Data Streams is designed for real-time streaming data ingestion.

Why this answer

Amazon Kinesis Data Streams is designed for real-time streaming data ingestion at scale, supporting throughput of up to 1 MB/s or 1,000 records per second per shard. With 500 records per second, Kinesis can reliably capture and store the social media API data for up to 365 days, enabling batch analytics via S3 delivery through Kinesis Firehose or custom consumers.

Exam trap

The trap here is that candidates confuse SQS's message queueing with Kinesis's stream processing, overlooking that SQS lacks ordered, replayable, and high-throughput streaming capabilities required for real-time data ingestion into S3.

How to eliminate wrong answers

Option A is wrong because Amazon SNS is a pub/sub messaging service for push notifications and fan-out, not designed for persistent, ordered streaming data ingestion or high-throughput record capture. Option B is wrong because Amazon SQS is a message queue for decoupling microservices with at-least-once delivery, but it lacks the shard-based parallelism, replay capability, and long-term retention needed for streaming data to S3. Option D is wrong because Amazon MQ is a managed message broker for ActiveMQ or RabbitMQ protocols, optimized for JMS and enterprise messaging, not for high-velocity stream ingestion or direct integration with S3 batch analytics.

187
MCQmedium

A company is streaming IoT data from thousands of devices into Amazon Kinesis Data Streams. The data must be transformed in real time before being stored in Amazon S3. Which service should be used to perform the transformation as the data streams through Kinesis?

A.AWS Glue
B.Amazon Kinesis Data Analytics for Apache Flink
C.Amazon EMR
D.AWS Lambda
AnswerB

Kinesis Data Analytics for Apache Flink runs continuous SQL or Flink applications directly against the stream, transforming records in flight before they land in Amazon S3. This satisfies the real-time transformation constraint, whereas AWS Glue and Lambda-based batch approaches operate after or outside the stream.

Why this answer

Amazon Kinesis Data Analytics for Apache Flink is the correct choice because it is purpose-built for running Apache Flink applications that can perform real-time transformations, filtering, and enrichment on data streaming through Kinesis Data Streams before outputting the results to destinations like Amazon S3. It integrates natively with Kinesis Data Streams as a source and can write transformed data directly to S3 using a Flink sink, making it ideal for this streaming ETL use case.

Exam trap

The trap here is that candidates often choose AWS Lambda because it is a familiar serverless option for event-driven processing, but they overlook its limitations in execution time, payload size, and lack of native state management for complex transformations, which makes Kinesis Data Analytics for Apache Flink the more robust and scalable choice for continuous streaming ETL.

How to eliminate wrong answers

Option A is wrong because AWS Glue is primarily a batch ETL service that processes data in job runs, not a real-time streaming transformation engine; while Glue Streaming exists, it is based on Spark Streaming and requires a separate Glue job with a streaming source, not a native Kinesis Data Streams integration for real-time transformations. Option C is wrong because Amazon EMR is a managed Hadoop/Spark cluster platform that can process streaming data but requires manual cluster management, provisioning, and configuration of Spark Streaming or Flink, adding operational overhead that is unnecessary for a simple transformation before S3 storage. Option D is wrong because AWS Lambda can process Kinesis Data Streams records in near real-time, but it has a maximum execution timeout of 15 minutes and a payload limit of 6 MB per invocation, making it unsuitable for high-throughput, continuous transformations of thousands of devices' data without risk of throttling or data loss.

188
Multi-Selecteasy

Which TWO AWS services can be used as sources for AWS Glue ETL jobs? (Choose two.)

Select 2 answers
A.Amazon Route 53
B.Amazon CloudFront
C.Amazon API Gateway
D.Amazon S3
E.Amazon RDS
AnswersD, E

S3 is a common source for Glue jobs.

Why this answer

Amazon S3 is a fully managed object storage service that serves as a common source for AWS Glue ETL jobs. Glue can read data from S3 using its built-in crawlers and connectors, supporting formats like Parquet, JSON, CSV, and Avro. The Glue Data Catalog can reference S3 locations, and ETL scripts can directly read from S3 buckets via the s3:// protocol.

Exam trap

The DEA-C01 exam often tests the misconception that any AWS service that stores or serves data (like Route 53 for DNS records or CloudFront for cached content) can be a Glue source, but Glue only supports sources that provide a direct data access interface (e.g., object storage, databases, or streaming services like Kinesis).

189
Multi-Selectmedium

A company is using Amazon Redshift for data warehousing. They need to ensure that data is encrypted at rest and in transit. Which TWO configurations are required to meet these requirements?

Select 2 answers
A.Enable encryption on the Redshift cluster using AWS KMS.
B.Configure the Redshift cluster to require SSL connections.
C.Use AWS CloudHSM to manage encryption keys for Redshift.
D.Enable VPC Flow Logs on the Redshift subnet.
E.Enable EBS encryption on the Redshift cluster nodes.
AnswersA, B

Enabling encryption on the Redshift cluster with AWS KMS satisfies the at-rest requirement, as it encrypts cluster data and snapshots using customer-managed or AWS-managed keys. This addresses the storage-layer constraint directly, though it does nothing for data in transit, which requires separate SSL/TLS configuration.

Why this answer

Option A is correct because enabling encryption on the Redshift cluster using AWS KMS provides encryption at rest — Redshift uses KMS customer master keys to encrypt the cluster's data blocks and system metadata on disk. Option B is correct because configuring the Redshift cluster to require SSL connections (via the require_ssl parameter set to true in the cluster's parameter group) enforces encryption in transit for all client and JDBC/ODBC connections to the cluster. Option C is not required because Redshift's at-rest encryption is natively managed through AWS KMS, not CloudHSM, and CloudHSM is not a prerequisite for meeting these requirements.

Option D is incorrect because VPC Flow Logs capture IP traffic metadata for network monitoring and do not encrypt data at rest or in transit. Option E is incorrect because Redshift manages its own storage encryption at the cluster level; enabling EBS encryption on cluster nodes is neither a supported nor a required configuration for Redshift data encryption.

190
MCQhard

A data engineer is troubleshooting an AWS Glue ETL job that fails with an access denied error when writing to an S3 bucket. The Glue job uses an IAM role that has an S3 bucket policy attached. The bucket policy denies access to any principal that does not use server-side encryption. What is the most likely cause of the failure?

A.The VPC endpoint policy for S3 is too restrictive.
B.The IAM role does not have s3:PutObject permission.
C.The Glue job is not using server-side encryption when writing to S3.
D.The S3 bucket uses S3 Block Public Access which denies all writes.
AnswerC

The bucket policy denies requests without encryption, causing access denied even if the role has PutObject permission.

Why this answer

If the Glue job does not set the encryption header (or the role does not have the kms:GenerateDataKey permission for SSE-KMS), the bucket policy will deny the request. Option A is wrong because Glue requires permissions on the S3 bucket and KMS key. Option B is wrong because VPC endpoints do not cause access denied errors for encryption.

Option D is wrong because S3 Block Public Access does not deny write access to authorized roles.

191
MCQmedium

A data engineer is designing a data lake on Amazon S3. The team wants to optimize query performance and reduce storage costs for a large dataset of JSON logs that are queried frequently by Amazon Athena. The logs are currently stored as uncompressed JSON files, each around 1 GB, in a single prefix. The engineer needs to improve query performance and reduce costs without changing the data format. Which action should the engineer take?

A.Compress the JSON files with gzip and split them into smaller files of approximately 128 MB.
B.Convert the JSON files to Apache Parquet and partition the data by date.
C.Move the data to Amazon Redshift and query it using Redshift Spectrum.
D.Enable S3 Transfer Acceleration on the bucket to speed up data retrieval.
AnswerA

Compressing with gzip reduces storage costs and the amount of data scanned by Athena. Splitting large files into smaller ones (around 128 MB) enables parallel processing and improves query performance. This approach maintains the JSON format and directly addresses the requirements without altering the data structure.

Why this answer

Compressing the JSON files with gzip reduces the storage footprint and the amount of data scanned by Athena, lowering costs. Splitting the large files into smaller, more manageable sizes allows Athena to process them in parallel, significantly improving query performance. This solution respects the requirement to keep the data in JSON format and directly addresses both performance and cost concerns.

Exam trap

The trap here is assuming that converting to a columnar format like Parquet is always the best optimization, but the scenario explicitly requires keeping the data in its original JSON format.

192
MCQeasy

A company is using AWS Glue to catalog data stored in Amazon S3. The data is partitioned by year, month, and day. A data analyst reports that new partitions are not automatically discovered by the Glue crawler. The crawler runs on a schedule every hour. What is the MOST likely reason for the missing partitions?

A.The IAM role used by the crawler does not have permission to list the S3 bucket.
B.The Glue Data Catalog is not configured to use a Hive metastore.
C.The number of partitions exceeds the Glue catalog limit of 100,000.
D.The crawler schedule is set to run too frequently.
AnswerA

Without s3:ListBucket permission, the crawler cannot enumerate prefixes under the bucket, so newly written year/month/day partitions remain invisible even though the hourly schedule runs. The missing IAM action is the specific cause of undiscovered partitions.

Why this answer

For a Glue crawler to discover new partitions in S3, its IAM role must have s3:ListBucket (and GetObject) permissions on the bucket and prefix. If the role lacks ListBucket, the crawler cannot enumerate the partition folders and will silently miss new partitions even though it runs on schedule. This is the most common cause of 'crawler runs but doesn't find new data' issues.

Exam trap

DEA-C01 often tests the misconception that crawler scheduling or catalog limits cause missing partitions, when the most common root cause is insufficient S3 ListBucket permission on the crawler's IAM role.

How to eliminate wrong answers

Option B is wrong because the Glue Data Catalog is a managed Hive-compatible metastore by default; it does not require an external Hive metastore to discover partitions. Option C is wrong because while Glue has soft limits on partitions per table (and 100,000 is not the relevant hard limit for this symptom), exceeding a limit would typically produce errors, not silent omission of new partitions. Option D is wrong because running the crawler more frequently does not prevent discovery — if anything, it would discover partitions sooner; schedule frequency is not the cause of missing partitions.

193
Matchingmedium

Match each AWS service to its primary purpose in data engineering.

Drag a concept onto its matching description — or click a concept then click the description.

Concepts
Matches

Serverless ETL and data catalog

Data warehousing and SQL analytics

Big data processing using Hadoop/Spark

Building and managing data lakes

Real-time streaming data ingestion

Why these pairings

The correct matches are: Amazon S3 → object storage for data lakes (A), AWS Glue → serverless ETL (B), Amazon Athena → interactive SQL queries on S3 (C), and Amazon Redshift → data warehousing (D). Common confusions involve mixing up services like Kinesis (streaming) with Glue (ETL) or Data Pipeline (orchestration) with Kinesis (ingestion).

194
MCQeasy

A data engineer is using AWS Step Functions to orchestrate a daily pipeline that runs several AWS Glue jobs in sequence. One Glue job intermittently fails due to a transient Amazon S3 503 error. The engineer wants the state machine to automatically retry only that Glue job up to three times with exponential backoff, without retrying the other jobs. What should the engineer do?

A.Set the Glue job's MaximumRetries property to 3 in the job definition, then add a Catch field in the state machine to handle final failures.
B.Add a Retry field to the Glue job task state with ErrorEquals set to States.TaskFailed, MaxAttempts set to 3, and an appropriate IntervalSeconds and BackoffRate.
C.Configure an Amazon CloudWatch alarm on Glue job failures and use an Application Auto Scaling policy to restart the state machine execution.
D.Wrap the entire state machine in a Map state and set MaxConcurrency to 1 so that failures are retried automatically by Step Functions.
AnswerB

Step Functions Retry is configured per state, so adding it to the specific Glue job task retries only that task up to three times. Exponential backoff is achieved with IntervalSeconds and BackoffRate, and ErrorEquals can include States.TaskFailed to catch transient failures. This precisely satisfies the requirement without affecting other jobs in the sequence.

Why this answer

Step Functions retry logic is declared inside the state definition, so attaching a Retry field to the Glue job task scopes retries to that task alone. Specifying MaxAttempts of 3 with IntervalSeconds and BackoffRate produces exponential backoff for transient failures. The other choices either apply retries at the wrong layer, change pipeline structure, or use services that cannot restart executions.

Exam trap

The trap here is conflating AWS Glue job-level MaximumRetries with Step Functions state-level Retry, which are separate mechanisms with different scope and backoff behavior.

195
MCQhard

A company runs a real-time analytics platform on Amazon ECS that ingests streaming data from Amazon Kinesis Data Streams, processes it, and stores results in Amazon DynamoDB. The data volume spikes unpredictably, causing DynamoDB to throttle write requests. The application uses on-demand capacity mode. The data engineer notices that the throttling occurs on a specific partition due to a hot key. The hot key is a customer ID that receives a disproportionate number of writes. The application cannot change the partition key design immediately. The engineer needs to reduce throttling while maintaining low latency. Which solution is most effective?

A.Switch to provisioned capacity with auto scaling and increase the write capacity units.
B.Implement a write buffer using Amazon SQS, and have consumers write to DynamoDB at a controlled rate.
C.Enable DynamoDB Accelerator (DAX) to cache the hot key writes.
D.Use DynamoDB Streams to trigger a Lambda function that retries throttled writes.
AnswerB

An SQS write buffer decouples ingestion from DynamoDB, letting consumers write at a controlled rate so the hot partition key is no longer overwhelmed. This absorbs unpredictable spikes while preserving low latency, and works without redesigning the partition key, which the stem forbids.

Why this answer

Buffering writes through Amazon SQS decouples the ingestion rate from DynamoDB's capacity, allowing consumers to write at a controlled pace. This directly mitigates throttling on the hot key without requiring a partition key redesign, and SQS provides low-latency, durable buffering suitable for real-time analytics.

Exam trap

The trap here is that candidates often assume on-demand capacity eliminates all throttling, but it does not protect against hot key skew; they may also confuse DAX's read caching with write buffering, or think retrying throttled writes is a viable solution rather than a reactive fix that increases latency.

How to eliminate wrong answers

Option A is wrong because switching to provisioned capacity with auto scaling does not solve the hot key issue; throttling occurs on a specific partition regardless of total capacity, and increasing write capacity units would not prevent a single partition from exceeding its 1,000 WCU limit. Option C is wrong because DAX is a caching layer for reads, not writes; it cannot buffer or absorb write throttling on a hot key. Option D is wrong because using DynamoDB Streams to retry throttled writes introduces latency and does not prevent throttling; it only retries failed writes, which can lead to backlog and increased latency, not a controlled rate.

196
MCQhard

A company is using an Amazon RDS for PostgreSQL database to store personally identifiable information (PII). The security team wants to ensure that database administrators cannot view the plaintext PII data. Which solution should a data engineer implement?

A.Use IAM policies to restrict DBA access to the RDS instance
B.Enable Dynamic Data Masking in RDS to obfuscate PII for all users
C.Enable encryption at rest for the RDS instance using AWS KMS
D.Use client-side encryption with AWS KMS to encrypt PII before inserting into the database
AnswerD

Client-side encryption with AWS KMS encrypts PII before it reaches PostgreSQL, so ciphertext alone is stored. Database administrators lack the KMS key permissions needed to decrypt, satisfying the requirement that they cannot view plaintext PII. Server-side options such as RDS encryption leave administrators able to read data through SQL queries.

Why this answer

Using AWS KMS with client-side encryption ensures that data is encrypted before being sent to RDS, so database administrators cannot read the plaintext. Dynamic data masking in RDS is not natively supported; application-level masking would be needed. RDS encryption at rest protects data on disk but DBAs with access can still query plaintext.

Using IAM policies to restrict access does not prevent DBAs with database credentials from viewing data.

197
MCQeasy

A data engineer needs to run a one-time transformation on a 500 GB dataset stored in Amazon S3. The transformation is written in Python and uses pandas, which cannot handle the full dataset in memory. The engineer wants a serverless option that can parallelize the work without managing servers. Which AWS service should the engineer use?

A.AWS Lambda with a function that streams the S3 objects and processes them in 15-minute increments.
B.AWS Glue with a Python shell job and the pandas library, splitting the data into chunks.
C.Amazon EMR Serverless with a Spark job that uses Spark DataFrame operations instead of pandas.
D.Amazon Athena with a CREATE TABLE AS SELECT statement that applies the transformation in SQL.
AnswerC

EMR Serverless runs Spark without cluster management and can distribute a 500 GB transformation across many workers. Spark DataFrames provide a distributed alternative to pandas, allowing the engineer to scale out and avoid the single-machine memory limit while remaining serverless.

Why this answer

EMR Serverless provides a serverless Spark environment that can distribute a large transformation across workers, removing the single-node memory limit of pandas. By converting the logic to Spark DataFrames, the engineer can process 500 GB in parallel without provisioning or managing servers, which matches the stated constraints.

Exam trap

The trap here is assuming that any serverless compute (Lambda or Glue Python shell) can handle large datasets, when only a distributed engine like Spark on EMR Serverless provides the needed parallelism and memory.

198
MCQeasy

A company uses Amazon DynamoDB as the primary data store for a web application. The application experiences occasional throttling on write requests. The data engineer needs to implement a solution that handles throttling gracefully without losing data. Which approach should the engineer use?

A.Increase the provisioned write capacity to a higher value
B.Use an Amazon SQS queue to buffer write requests before sending to DynamoDB
C.Implement exponential backoff in the application's write retry logic
D.Enable DynamoDB Accelerator (DAX) to cache writes
AnswerC

Exponential backoff retries throttled writes with progressively longer, randomised delays, absorbing transient capacity bursts without dropping requests. DynamoDB returns ProvisionedThroughputExceededException for throttled writes; retrying with backoff lets the request succeed once capacity frees, satisfying the no-data-loss constraint. It handles throttling gracefully rather than preventing it.

Why this answer

Implementing exponential backoff in the application's write retry logic is the standard AWS-recommended approach for handling DynamoDB throttling (ProvisionedThroughputExceededException). Exponential backoff gradually increases the wait time between retries, reducing the retry rate and allowing the throttling condition to subside, while ensuring no write data is lost as long as the retries eventually succeed. This approach is lightweight, requires no additional AWS services, and aligns with best practices for building resilient applications against DynamoDB throttling.

Exam trap

The trap here is that candidates often confuse DAX as a write cache or assume SQS is the only way to buffer writes, but the question specifically asks for handling throttling gracefully without losing data, and exponential backoff is the direct, built-in mechanism for retrying throttled requests in DynamoDB.

How to eliminate wrong answers

Option A is wrong because simply increasing provisioned write capacity may reduce throttling but does not handle throttling gracefully when it occurs; it also incurs higher costs and does not address the root cause of occasional spikes. Option B is wrong because using an SQS queue to buffer write requests introduces eventual consistency and potential data loss if the queue messages expire or are not processed before the DynamoDB write; it also adds complexity and latency, and is not the standard pattern for handling DynamoDB throttling directly. Option D is wrong because DynamoDB Accelerator (DAX) is an in-memory cache for reads only, not writes; it cannot cache write requests or mitigate write throttling.

199
MCQhard

A company uses AWS DMS to migrate data from an on-premises Oracle database to Amazon Redshift. The migration is successful, but after a few days, data in Redshift becomes inconsistent with the source due to ongoing changes. The company needs to keep Redshift synchronized with minimal latency. Which approach should the data engineer use?

A.Configure DMS with ongoing replication using change data capture (CDC).
B.Use Amazon Redshift COPY with S3 staging and AWS Lambda triggers.
C.Schedule a full DMS load every night.
D.Set up Amazon Redshift Spectrum to query the Oracle database directly.
AnswerA

Ongoing replication with CDC lets AWS DMS continuously apply source Oracle changes to Redshift, satisfying the requirement to keep Redshift synchronised with minimal latency after the initial full load. Without CDC, only a one-time migration occurs, so later source changes never propagate.

Why this answer

AWS DMS supports ongoing replication using change data capture (CDC), which captures incremental changes from the Oracle source (via Oracle LogMiner or binary logs) and applies them to Amazon Redshift in near real-time. This approach ensures that Redshift remains synchronized with the source database with minimal latency, meeting the requirement for ongoing consistency after the initial full load.

Exam trap

The trap here is that candidates may confuse Amazon Redshift Spectrum's federated querying capability with actual data replication, or assume that nightly batch loads (Option C) are sufficient for 'minimal latency' requirements, when DMS CDC is the only option that provides continuous, low-latency synchronization.

How to eliminate wrong answers

Option B is wrong because Amazon Redshift COPY with S3 staging and AWS Lambda triggers requires manual or event-driven extraction of data from Oracle, which introduces latency and complexity, and does not provide native CDC-based continuous replication. Option C is wrong because scheduling a full DMS load every night would result in significant data loss between loads (up to 24 hours of inconsistency) and does not achieve minimal latency. Option D is wrong because Amazon Redshift Spectrum queries external data directly from Oracle via federated querying, but it does not replicate or synchronize data into Redshift; it only provides a query-time view, which incurs high latency and does not maintain a consistent local copy.

200
MCQmedium

A company uses AWS Lake Formation to manage data lake permissions. The data lake contains sensitive customer data in the 'customer' database. The security team wants to ensure that only users with a specific tag 'access_level=analyst' can query the 'customer' table. Which combination of steps should the data engineer take to enforce this?

A.In Lake Formation, create an LF-tag 'access_level' with values 'analyst' and 'admin'. Grant 'SELECT' permission on the 'customer' table to the tag value 'analyst'. Associate the LF-tag with the 'customer' table.
B.Create an IAM policy that conditionally allows 'glue:GetTable' based on the tag 'access_level=analyst'.
C.Apply a bucket policy on the S3 location of the 'customer' table that allows access only if the request carries the tag 'access_level=analyst'.
D.Use Lake Formation column-level filters to restrict access to columns based on the tag 'access_level=analyst'.
AnswerA

This uses Lake Formation TBAC to restrict access based on the user's tag.

Why this answer

Lake Formation LF-tags allow you to define metadata tags (key-value pairs) and grant permissions to those tags. By creating an LF-tag 'access_level' with values 'analyst' and 'admin', granting SELECT on the 'customer' table to the tag value 'analyst', and associating that LF-tag with the table, only principals who have the tag 'access_level=analyst' (or are granted via the tag) can query the table. This enforces tag-based access control at the Lake Formation permission layer, which is the intended mechanism for fine-grained, attribute-based access control in Lake Formation.

Exam trap

The trap here is that candidates often confuse IAM tag-based policies (Option B) or S3 bucket policies (Option C) with Lake Formation's native LF-tag mechanism, not realizing that LF-tags are a Lake Formation-specific construct that must be managed within Lake Formation itself, not at the IAM or S3 level.

How to eliminate wrong answers

Option B is wrong because an IAM policy conditionally allowing 'glue:GetTable' based on a tag controls access to the Glue Data Catalog API, but it does not enforce Lake Formation permissions on the underlying data; Lake Formation permissions override IAM policies for registered locations, and this approach would not prevent a user with the tag from querying the table if Lake Formation grants are not also configured. Option C is wrong because S3 bucket policies operate at the object storage layer and cannot evaluate Lake Formation LF-tags; they can use IAM tags via the 'aws:RequestTag' condition key, but this would require the request to carry the tag, which is not how Lake Formation principals are identified, and it would bypass Lake Formation's centralized permission model. Option D is wrong because column-level filters in Lake Formation restrict access to specific columns based on a filter expression, not based on LF-tags; LF-tags are used for row-level or table-level permission grants, not for column-level filtering.

201
MCQhard

A company runs a time-series forecasting model that writes results to an S3 bucket every 5 minutes. A downstream ETL job reads this data, but sometimes fails because it encounters incomplete files (zero bytes). What is the MOST reliable way to ensure the ETL job only processes complete files?

A.Set an S3 Lifecycle policy to delete files smaller than 1 MB.
B.Use S3 Copy to move files to a 'processed' folder after the ETL job reads them.
C.Configure S3 Select to query the files and only return rows if the file is complete.
D.Use S3 Event Notifications to trigger a Lambda function that checks file size and then moves the file to a 'ready' prefix.
AnswerD

S3 Event Notifications fire on object creation, letting a Lambda verify the object is non-zero before relocating it to a ready prefix. The ETL job then reads only that prefix, so it never encounters the zero-byte partial writes.

Why this answer

S3 Event Notifications can trigger a Lambda function upon object creation, which can check the file size (e.g., > zero bytes) and then copy the file to a 'ready' prefix, ensuring the ETL job only processes complete files. Option A is wrong because an S3 Lifecycle policy can delete small files but does not prevent the ETL from reading incomplete files. Option B is wrong because S3 Copy does not verify completeness.

Option C is wrong because S3 Select still reads the file even if it is incomplete; it doesn't guarantee completeness.

202
MCQmedium

A data engineering team needs to ingest streaming data from thousands of IoT devices and store it in Amazon S3 for batch processing. The data arrives at a rate of 10 MB/s, with occasional spikes up to 50 MB/s. The data must be processed in near real-time with minimal latency. Which AWS service should be used for ingestion?

A.Amazon DynamoDB Streams
B.Amazon Kinesis Data Streams
C.Amazon SQS
D.Amazon S3
AnswerB

Kinesis Data Streams ingests high-throughput streaming data with sub-second latency and scales elastically to absorb the 50 MB/s spikes, unlike SQS or batch uploads. It satisfies the near-real-time, minimal-latency requirement while buffering records for downstream delivery into Amazon S3.

Why this answer

Amazon Kinesis Data Streams is designed for real-time streaming data ingestion at scale, handling throughput from megabytes to gigabytes per second with low latency. It can absorb the described 10 MB/s baseline and 50 MB/s spikes by sharding, and integrates directly with AWS Lambda or Kinesis Data Firehose to land data into Amazon S3 for batch processing.

Exam trap

The trap here is that candidates confuse Amazon SQS with a streaming service, but SQS is a pull-based queue with no ordering guarantees across multiple consumers, whereas Kinesis Data Streams provides ordered, replayable, and near-real-time data ingestion.

How to eliminate wrong answers

Option A is wrong because DynamoDB Streams captures changes to DynamoDB tables, not arbitrary streaming data from IoT devices, and its throughput is limited by the table's capacity, making it unsuitable for high-volume, low-latency ingestion. Option C is wrong because Amazon SQS is a message queue for decoupling components, not a streaming ingestion service; it does not support real-time processing with sub-second latency for continuous data streams and has a 256 KB message size limit. Option D is wrong because Amazon S3 is an object storage service, not a real-time ingestion endpoint; writing directly to S3 from thousands of devices would cause high latency due to HTTP overhead and lack of streaming semantics, and it cannot handle the required near-real-time processing.

203
MCQeasy

A data engineer needs to store semi-structured JSON logs from multiple microservices in a cost-effective manner for later analysis using Amazon Athena. The logs are generated continuously, and the total volume is about 1 TB per day. The data must be queryable within minutes of arrival. Which storage solution is most appropriate?

A.Amazon DynamoDB table with JSON attribute
B.Amazon RDS for PostgreSQL table with JSON column
C.Amazon S3 bucket with partitioned folders
D.Amazon Redshift cluster with JSON ingestion
AnswerC

Amazon S3 stores JSON at low cost per terabyte and integrates natively with Athena, which queries data in place. Partitioning folders by date or service prunes scanned data, cutting query cost and latency, so logs become queryable within minutes of arrival.

Why this answer

Amazon S3 with partitioned folders is the most appropriate solution because it provides a cost-effective, scalable storage layer for semi-structured JSON logs, and integrates natively with Amazon Athena for serverless querying. By partitioning the data by time (e.g., year/month/day/hour), Athena can use partition pruning to minimize scanned data, enabling queries within minutes of arrival. S3's low cost per GB and lifecycle policies further optimize storage for the 1 TB/day volume.

Exam trap

AWS often tests the misconception that a data warehouse (Redshift) or a NoSQL database (DynamoDB) is required for analytical queries on semi-structured data, when in fact S3 with Athena is the most cost-effective and scalable solution for serverless ad-hoc analysis on raw logs.

How to eliminate wrong answers

Option A is wrong because Amazon DynamoDB is optimized for key-value and document access patterns with low-latency reads/writes, not for ad-hoc analytical queries on large volumes of JSON logs; scanning 1 TB/day would be prohibitively expensive and slow, and it lacks native integration with Athena. Option B is wrong because Amazon RDS for PostgreSQL is a relational database designed for transactional workloads, not for storing and analyzing 1 TB/day of semi-structured logs; it would require manual partitioning, incur high storage costs, and cannot scale to petabyte-scale analytics efficiently. Option D is wrong because Amazon Redshift is a petabyte-scale data warehouse optimized for complex analytical queries, but it is overkill and more expensive than S3 for raw log storage; ingesting 1 TB/day of JSON logs into Redshift requires an ETL pipeline (e.g., COPY from S3) and incurs compute costs even when idle, whereas S3 with Athena is serverless and pay-per-query.

204
MCQmedium

A data engineer manages an AWS Glue Data Catalog table that contains sensitive customer PII. The table's underlying data is in Amazon S3 and is queried by several AWS analytics services. The security team wants to implement column-level access control so that only authorized principals can view the PII columns, while other principals can still query non-sensitive columns. The solution must integrate with AWS Lake Formation and be enforced consistently across all query engines. Which approach should the data engineer take?

A.Implement an S3 bucket policy that allows only certain prefixes and use separate buckets for PII and non-PII data.
B.Create an IAM policy that denies access to the PII columns and attach it to the roles used by the analytics services.
C.Use AWS Glue Data Catalog resource policies to restrict access to specific columns in the table.
D.Register the S3 bucket as a Lake Formation data location, then grant column-level permissions on the table using Lake Formation.
AnswerD

Lake Formation column-level permissions allow granular control over specific columns in a table. By registering the S3 location and granting column-level permissions, the engineer ensures that only authorized principals can access PII columns, and this enforcement is applied across integrated services like Athena, Redshift Spectrum, and Glue ETL. This directly meets the requirement for consistent column-level access control.

Why this answer

Lake Formation column-level permissions are designed to provide fine-grained access control on tabular data in the Data Catalog. By registering the S3 location and granting column-level permissions, the data engineer can restrict access to sensitive columns while allowing queries on other columns. This enforcement is consistent across integrated analytics services.

Exam trap

The trap here is assuming that IAM policies or Glue Data Catalog resource policies can enforce column-level access control, but they operate at different levels and cannot restrict specific columns within a table.

205
MCQhard

A financial services company is building a real-time fraud detection system. Transaction data is ingested via Amazon Kinesis Data Streams and processed by an Amazon Kinesis Data Analytics for Apache Flink application that runs sliding window aggregations. The output is written to an Amazon S3 bucket for downstream analysis. The Flink application is configured with parallelism of 4 and checkpointing every minute. The company has noticed that the application is experiencing high latency and the checkpointing is frequently failing. The CloudWatch metrics show that the Flink application's CPU utilization is near 100% and the checkpoint duration is spiking to over 5 minutes. The data engineer needs to improve performance. Which action should the data engineer take?

A.Increase the number of shards in the source Kinesis stream to improve throughput.
B.Increase the parallelism of the Flink application to distribute the workload across more resources.
C.Increase the heap memory of the Flink application to handle larger state.
D.Decrease the checkpoint interval to 30 seconds to reduce the amount of state being checkpointed.
AnswerB

Raising parallelism spreads the sliding-window aggregation across more task slots, directly relieving the near-100% CPU saturation that is stretching checkpoint duration beyond the one-minute interval. More parallel subtasks also reduce per-subtask state size, so checkpoint barriers complete faster, satisfying the stem's latency and checkpoint-failure constraint.

Why this answer

The Flink application is experiencing high CPU utilization and checkpointing failures, indicating that the current parallelism of 4 is insufficient to handle the workload. Increasing parallelism distributes the workload across more resources, reducing CPU load per task and allowing checkpointing to complete faster. This is a standard scaling approach for Flink applications in Kinesis Data Analytics.

Exam trap

DEA-C01 often tests the confusion between scaling the source (Kinesis shards) and scaling the processing application (Flink parallelism); candidates might think adding shards helps, but the bottleneck is in the Flink application's CPU.

How to eliminate wrong answers

Option A is wrong because increasing Kinesis shards improves source throughput but does not address the Flink application's CPU bottleneck; it could even worsen the problem by providing more data to process. Option C is wrong because increasing heap memory may help with state size but does not address CPU saturation; checkpoint duration is spiking due to CPU, not memory. Option D is wrong because decreasing the checkpoint interval would increase checkpointing frequency, potentially worsening the issue and not reducing state size.

206
Multi-Selectmedium

A company needs to protect sensitive data stored in Amazon S3 from unauthorized access. Which TWO actions should the data engineer take? (Choose two.)

Select 2 answers
A.Configure S3 bucket policies to require MFA for delete operations
B.Enable cross-region replication for all buckets
C.Set up an S3 Lifecycle policy to transition objects to Glacier
D.Enable S3 Block Public Access at the account level
E.Enable S3 Versioning on all buckets
AnswersA, D

Bucket policies requiring MFA for delete operations enforce multi-factor authentication before objects can be removed, guarding against compromised credentials and accidental deletion. This satisfies the protection requirement by adding a strong authentication control at the S3 API layer.

Why this answer

Option A is correct because an S3 bucket policy can include a condition such as aws:MultiFactorAuthPresent to deny s3:DeleteObject or s3:DeleteBucket unless the request is authenticated with MFA, adding a strong control against unauthorized or accidental deletion of sensitive data. Option D is correct because enabling S3 Block Public Access at the account level applies the four block-public-access settings to every bucket in the account, preventing bucket policies or ACLs from exposing objects publicly and thus blocking a major unauthorized-access vector. Option B is not correct because cross-region replication is a durability/availability and compliance feature that copies objects to another region; it does not by itself prevent unauthorized access.

Option C is not correct because a Lifecycle policy transitioning objects to Glacier is a cost/storage-class management action, not an access-control mechanism. Option E is not correct because S3 Versioning preserves prior object versions for recovery but does not restrict who can read or access the data.

207
MCQhard

A company uses Amazon DynamoDB as the primary data store for a high-traffic application. Recently, read latency has increased significantly. The DynamoDB table has on-demand capacity mode. Which action is MOST effective to reduce read latency?

A.Add a DynamoDB Accelerator (DAX) cluster in front of the table
B.Switch the table to provisioned capacity mode with higher read capacity
C.Increase the read capacity units in the table's auto scaling settings
D.Enable DynamoDB Global Tables to distribute reads across regions
AnswerA

DynamoDB Accelerator is an in-memory cache that serves eventually consistent reads in microseconds, absorbing the repeated read traffic that drives latency on the on-demand table. Placing DAX in front reduces read latency by orders of magnitude without changing the table's capacity mode.

Why this answer

DynamoDB Accelerator (DAX) is an in-memory cache purpose-built for DynamoDB that reduces read latency from single-digit milliseconds to microseconds for eventually consistent reads. For a high-traffic, read-heavy application on on-demand mode, adding a DAX cluster in front of the table is the most direct and effective way to cut read latency without changing capacity mode or table architecture.

Exam trap

DEA-C01 often tests the distinction between throughput scaling (RCUs, on-demand) and latency reduction (DAX, caching), so candidates who equate 'more capacity' with 'lower latency' pick the wrong option.

How to eliminate wrong answers

Option B is wrong because switching to provisioned mode with higher RCUs does not reduce per-request latency — it only changes capacity management and can even cause throttling if misconfigured; on-demand already scales automatically. Option C is wrong because auto scaling settings apply to provisioned mode, not on-demand, and increasing RCUs addresses throughput, not latency. Option D is wrong because Global Tables replicate data across regions for disaster recovery and local read performance, but they do not reduce latency for reads in the primary region and add replication cost and complexity.

208
Multi-Selectmedium

A data engineer is monitoring an AWS Glue ETL job that intermittently fails with 'Container killed by YARN for exceeding memory limits' during a large shuffle stage. The job reads from Amazon S3, performs a groupByKey aggregation, and writes to Amazon S3. The engineer wants to reduce the chance of executor memory exhaustion without changing the source data. (Choose two.)

Select 2 answers
A.Enable the Glue job bookmark to skip previously processed files.
B.Replace the groupByKey operation with reduceByKey or aggregateByKey to combine values before the shuffle.
C.Change the output write format from Parquet to uncompressed JSON.
D.Set the job's max concurrent runs to 1 to avoid overlapping executions.
E.Increase the number of workers allocated to the Glue job.
AnswersB, E

groupByKey shuffles all values for a key before aggregation, which can produce enormous intermediate data. reduceByKey and aggregateByKey perform map-side combiners that pre-aggregate values locally before the shuffle, dramatically shrinking the data moved across the network. Less shuffled data means smaller per-executor memory footprints during the aggregation, addressing the root cause of the container kills.

Why this answer

The container kill originates from a single executor holding too much data during the shuffle. Adding workers spreads the shuffle across more containers, reducing per-executor memory demand. Replacing groupByKey with reduceByKey or aggregateByKey introduces map-side combining so far less data crosses the network.

Both changes reduce peak memory in the shuffle stage without touching the source data.

Exam trap

The trap here is attributing an executor memory kill to total job volume rather than to how much data a single shuffle partition must hold.

209
MCQmedium

A company wants to grant cross-account access to an S3 bucket without using IAM roles. The data engineer needs to write a bucket policy that allows another AWS account to list objects. Which Principal should be specified in the bucket policy?

A.The AWS account ID that owns the bucket
B.The AWS account ID of the other account
C.The IAM user ARN in the other account
D.The root user of the other account
AnswerB

Specifying the other account's AWS account ID as the Principal delegates access at the account level, letting that account's administrators manage their own users without IAM role assumption. This satisfies the stem's constraint of cross-account access without roles, since S3 bucket policies accept account ARNs directly as valid principals.

Why this answer

Specifying the AWS account ID of the other account as the Principal in the bucket policy grants cross-account access to all users and roles in that account, allowing them to list objects. Option A is incorrect because the owning account's ID would grant access to itself, not the other account. Option C is incorrect because specifying an IAM user ARN would restrict access to only that user, not the entire account.

Option D is incorrect because the root user is a specific principal, not the account-wide access needed for cross-account delegation.

210
MCQmedium

A data engineer manages an Amazon S3 data lake where analytics queries run through Amazon Athena. Monthly partition folders hold Parquet files, and each partition contains tens of thousands of small files averaging 40 KB. Athena queries that scan a single month take much longer than expected and consume far more bytes scanned than the actual data volume. The engineer must improve query performance without changing the table schema or the folder layout. What should the engineer do?

A.Enable S3 Transfer Acceleration on the bucket so that Athena can retrieve the small objects with lower latency across edge locations.
B.Run an AWS Glue ETL job that reads each partition and rewrites it as fewer, larger Parquet files of roughly 128 MB, then update the partitions in the Data Catalog.
C.Attach an S3 Lifecycle policy that transitions the Parquet objects to S3 Glacier Instant Retrieval after 30 days to improve read throughput.
D.Increase the number of partitions by splitting each monthly folder into daily folders and repointing the table at the new prefixes.
AnswerB

Consolidating many tiny files into fewer large Parquet files drastically reduces the per-file overhead that Athena pays when listing and opening objects, so scans finish faster and less metadata work is repeated. The rewrite keeps the same partition structure and schema, satisfying the constraint that neither the table definition nor the folder layout may change.

Why this answer

Athena performance degrades when a partition holds a huge number of very small objects because each object requires a separate request and contributes metadata overhead, inflating both runtime and reported bytes scanned. Compacting the files with an AWS Glue job into larger Parquet objects preserves the schema and partition layout while cutting that overhead, which is the standard remedy for this pattern.

Exam trap

The trap here is assuming that a storage-class or network-acceleration change can fix a small-file problem, when the real fix is rewriting the objects into fewer larger files.

211
MCQhard

A data engineer is using AWS Database Migration Service (AWS DMS) to migrate an on-premises Oracle database to Amazon Aurora PostgreSQL. The migration is ongoing, and the engineer notices that some transactions are not being replicated to the target. The DMS task is configured for full load plus change data capture (CDC). Which action should the engineer take to troubleshoot the missing transactions?

A.Check the DMS task's CloudWatch logs for errors related to unsupported data types or DDL changes.
B.Enable AWS DMS validation to compare source and target rows and identify discrepancies.
C.Modify the DMS task to use 'Drop and create' target table preparation mode.
D.Increase the DMS replication instance size to handle the transaction volume.
AnswerA

AWS DMS logs errors for issues like unsupported data types, DDL changes, or transformation failures in CloudWatch Logs. These errors can cause specific transactions to be skipped during CDC. Reviewing the logs is the first step to identify why transactions are missing. The logs provide detailed messages that help pinpoint the root cause, such as a data type not supported by the target.

Why this answer

When AWS DMS CDC skips transactions, it is often due to errors such as unsupported data types or DDL changes. These errors are logged in CloudWatch Logs. Reviewing the logs allows the engineer to identify the specific cause and take corrective action, such as excluding the problematic table or adjusting the task settings.

This is the most direct troubleshooting step.

Exam trap

The trap here is assuming that performance tuning or validation will fix missing transactions, when the real issue is often logged errors that need to be addressed.

212
MCQhard

A data engineer is designing a data lake on Amazon S3 that must comply with a regulatory requirement to prevent any data from being overwritten or deleted for 7 years after creation. Which S3 feature should be used?

A.S3 bucket policy that denies s3:DeleteObject
B.S3 bucket versioning with MFA Delete
C.S3 Object Lock with retention mode set to COMPLIANCE
D.S3 bucket versioning only
AnswerC

S3 Object Lock in COMPLIANCE mode enforces a write-once-read-many (WORM) model, preventing any user, including the root account, from overwriting or deleting objects for the defined retention period. Setting a 7-year retention directly satisfies the regulatory immutability requirement, unlike GOVERNANCE mode, which privileged users can bypass.

Why this answer

S3 Object Lock with retention mode set to COMPLIANCE ensures that objects cannot be overwritten or deleted for the specified retention period (7 years). The retention period cannot be shortened or removed by any user, including the root user, making it suitable for regulatory compliance. Option A is incorrect because a bucket policy that denies s3:DeleteObject can be modified or removed, and it does not prevent overwrites.

Option B is incorrect because MFA Delete requires an additional authentication factor but can still be disabled by an authorized user, and it does not enforce a retention period. Option D is incorrect because bucket versioning alone does not prevent deletion; it only creates delete markers, and objects can still be permanently deleted.

213
Multi-Selecthard

A data engineer is troubleshooting a Kinesis Data Streams consumer application that is falling behind. The stream has 10 shards and is receiving 5 MB/s of data. The consumer uses the Kinesis Client Library (KCL) with a single worker. The worker is processing all 10 shards but is experiencing high latency and checkpointing delays. Which THREE actions should the engineer take to improve consumer performance? (Select THREE.)

Select 3 answers
A.Increase the number of KCL workers to match the number of shards.
B.Enable enhanced fan-out for the consumer.
C.Decrease the checkpoint interval to reduce checkpointing overhead.
D.Increase the KCL maxRecords parameter to process more records per call.
E.Increase the number of shards in the stream.
AnswersA, B, D

Multiple workers can process shards in parallel, reducing per-worker load.

Why this answer

The KCL worker is processing all 10 shards sequentially within a single worker, causing a bottleneck. By increasing the number of KCL workers to match the number of shards, each worker can process one shard in parallel, significantly improving throughput and reducing latency. This is a standard scaling pattern for KCL-based consumers.

Exam trap

The trap here is that candidates may think decreasing the checkpoint interval (Option C) reduces overhead, when in fact it increases the frequency of DynamoDB writes and can degrade performance; the correct approach is to increase the checkpoint interval or use asynchronous checkpointing.

214
MCQeasy

Refer to the exhibit. A data engineer checks the versioning status of an S3 bucket and sees the above output. The bucket contains critical logs that must not be permanently deleted. What should the engineer do to enhance protection against accidental or malicious deletion?

A.Enable MFA Delete on the bucket
B.Enable versioning on the bucket
C.Enable cross-region replication
D.Configure a lifecycle policy to expire noncurrent versions
AnswerA

MFA Delete requires multi-factor authentication for permanently deleting object versions or changing bucket versioning state, so a compromised credential alone cannot purge the critical logs. This directly satisfies the requirement that versioned data must not be permanently deleted.

Why this answer

Enabling MFA Delete on the bucket adds an extra layer of protection by requiring multi-factor authentication for permanently deleting object versions or changing the versioning state. Since the bucket already has versioning enabled (as shown in the exhibit), MFA Delete is the appropriate enhancement to prevent accidental or malicious permanent deletion.

Exam trap

DEA-C01 often tests the difference between versioning, MFA Delete, and Object Lock, so candidates may choose versioning again or lifecycle policies without realizing the exhibit already shows versioning enabled.

How to eliminate wrong answers

Option B is wrong because versioning is already enabled (the exhibit shows versioning status), so enabling it again is redundant. Option C is wrong because cross-region replication provides durability but does not prevent deletion; deleted objects can be replicated as deletions. Option D is wrong because a lifecycle policy to expire noncurrent versions would actually delete older versions, which contradicts the requirement to prevent permanent deletion.

215
MCQhard

A data engineer is troubleshooting a Kinesis Data Streams application that is experiencing high latency. The stream has 2 shards. The application is using a single Kinesis Client Library (KCL) worker to process all shards. Which change will MOST likely reduce latency?

A.Increase the number of shards to 4.
B.Deploy multiple KCL workers to process shards in parallel.
C.Use a larger instance type for the Kinesis stream.
D.Decrease the number of shards to 1.
AnswerB

Deploying multiple KCL workers lets each worker lease a distinct shard, so the two shards are processed concurrently rather than sequentially by one worker. This directly addresses the stem's constraint: a single worker cannot parallelise across shards, so adding workers raises aggregate throughput and reduces processing latency.

Why this answer

The application uses a single KCL worker to process all 2 shards, which processes records sequentially and causes high latency. Deploying multiple KCL workers (ideally one per shard) enables parallel processing of shards, significantly reducing latency. Option A is incorrect because increasing shard count to 4 adds more capacity but does not address the bottleneck of a single worker; the same worker would process all 4 shards sequentially, potentially worsening latency.

Option C is incorrect because Kinesis Data Streams is a managed service; there is no instance type to change for the stream itself. The KCL worker runs on your compute resources, not on the stream. Option D is incorrect because decreasing shards to 1 reduces the level of parallelism, increasing the workload per shard and likely increasing latency further.

216
MCQeasy

A data engineer needs to store semi-structured JSON files that are accessed infrequently but must be retrievable within minutes. The data is immutable and must be stored cost-effectively. Which AWS service should the engineer use?

A.Amazon DynamoDB with on-demand capacity
B.Amazon EBS with gp3 volume
C.Amazon S3 with S3 Standard-IA storage class
D.Amazon RDS for PostgreSQL with JSONB data type
AnswerC

S3 Standard-IA suits infrequent access with millisecond retrieval, satisfying the minutes-based constraint. Its lower storage cost than S3 Standard meets the cost-effectiveness requirement, while object immutability is preserved through versioning and Object Lock. Unlike Glacier tiers, no retrieval job or restore delay is needed, keeping access immediate.

Why this answer

Amazon S3 Standard-IA (Infrequent Access) is designed for data that is accessed less frequently but requires rapid retrieval when needed, with retrieval times in milliseconds. It offers lower storage costs than S3 Standard while maintaining high durability and availability, making it ideal for storing immutable semi-structured JSON files that must be retrievable within minutes. The service is cost-effective for infrequently accessed data because it charges a retrieval fee per GB, but the storage price is significantly lower than standard tiers.

Exam trap

The trap here is that candidates often confuse 'infrequently accessed' with 'archival' and choose Glacier or Deep Archive, but the requirement for retrieval within minutes eliminates those options, while DynamoDB or RDS seem plausible for JSON but are not cost-effective for immutable, infrequently accessed data.

How to eliminate wrong answers

Option A is wrong because Amazon DynamoDB with on-demand capacity is a NoSQL database optimized for high-frequency, low-latency queries and is not cost-effective for infrequently accessed, immutable JSON files; it charges per read/write request unit and storage, which would be wasteful for archival-like data. Option B is wrong because Amazon EBS with gp3 volume is a block storage service designed for EC2 instances and requires an attached compute instance to access data, adding unnecessary cost and complexity; it is not a standalone object storage solution for infrequently accessed files. Option D is wrong because Amazon RDS for PostgreSQL with JSONB data type is a relational database service that incurs ongoing compute and storage costs, even when idle, and is overkill for storing immutable JSON files that are only occasionally retrieved; it is designed for transactional workloads and complex queries, not cost-effective archival storage.

217
Multi-Selecteasy

A data engineer needs to transfer 50 TB of data from an on-premises data center to Amazon S3 over a 1 Gbps network. The transfer must be completed within one week. Which TWO AWS services can be used for this task? (Choose TWO.)

Select 2 answers
A.AWS Glue
B.AWS DataSync
C.AWS Snowball
D.Amazon S3 Transfer Acceleration
E.AWS Direct Connect
AnswersB, C

Designed for network-based bulk data transfer.

Why this answer

AWS DataSync is correct because it is designed to efficiently transfer large datasets over the network using a purpose-built agent that parallelizes data transfer and optimizes network utilization. With a 1 Gbps link, DataSync can transfer 50 TB within a week by leveraging its built-in compression, encryption, and incremental transfer capabilities, making it suitable for this time-constrained migration.

Exam trap

The trap here is that candidates assume S3 Transfer Acceleration can accelerate any transfer, but it only optimizes the last-mile upload to S3 and does not address the bottleneck of moving data from on-premises storage to the internet, nor does it provide a mechanism to pull data from on-premises systems.

218
MCQmedium

A data engineer manages an Amazon Redshift cluster that stores sales data in a table with a sort key on the sale_date column. The table is growing rapidly, and queries that filter by sale_date are becoming slower. The engineer notices that the table has a high percentage of unsorted rows. What should the engineer do to improve query performance with the least effort?

A.Run the VACUUM REINDEX command on the table.
B.Run the ANALYZE command on the table.
C.Run the VACUUM SORT ONLY command on the table.
D.Run the VACUUM FULL command on the table.
AnswerC

VACUUM SORT ONLY sorts the rows in the table according to the sort key without reclaiming space. This directly addresses the high percentage of unsorted rows and improves the efficiency of range-restricted scans on sale_date. It is a straightforward maintenance operation that requires minimal effort and targets the specific issue of unsorted data.

Why this answer

VACUUM SORT ONLY sorts the rows in a table according to the defined sort key, which directly improves the performance of queries that filter on that key. It is the least intrusive vacuum operation that addresses unsorted rows. ANALYZE only updates statistics, while VACUUM FULL and VACUUM REINDEX are more resource-intensive and not required for the described issue.

Exam trap

The trap here is confusing VACUUM SORT ONLY with VACUUM FULL; the former sorts without reclaiming space, while the latter does both and is overkill for just sorting.

219
MCQhard

A company is using Amazon Kinesis Data Streams to ingest real-time clickstream data. The data is consumed by a fleet of EC2 instances running a custom application that processes the records and writes to DynamoDB. The application is experiencing high latency and records are being processed slower than they are produced. The stream has 5 shards. Which action would MOST effectively improve processing speed?

A.Use the Kinesis Client Library (KCL) to automatically distribute shards among instances.
B.Increase the EC2 instance size to provide more CPU and memory.
C.Add more EC2 instances consuming from the same stream without changing shard count.
D.Increase the number of shards in the Kinesis stream.
AnswerD

Kinesis shard count sets the stream's total ingest and per-shard consumer throughput ceiling. With five shards saturated, adding shards raises aggregate capacity and lets the EC2 fleet's consumers process records faster, addressing the production-versus-processing imbalance directly.

Why this answer

The bottleneck is the number of shards in the Kinesis stream. Each shard provides a fixed read capacity of 2 MB/s and 5 read transactions per second. With only 5 shards, the total read throughput is limited regardless of how many EC2 instances consume the data.

Increasing the number of shards increases the total read capacity, allowing more records to be consumed in parallel and reducing processing latency.

Exam trap

The trap here is that candidates often think adding more consumers (EC2 instances) will automatically speed up processing, but they fail to recognize that each shard's read throughput is fixed, so without increasing shards, additional consumers cannot consume more data in parallel.

How to eliminate wrong answers

Option A is wrong because the Kinesis Client Library (KCL) manages shard-to-instance assignment and checkpointing, but it does not increase the total throughput of the stream; it only distributes existing shard capacity among consumers. Option B is wrong because increasing EC2 instance size improves compute resources but does not address the fundamental read throughput limit imposed by the number of shards; the application will still be throttled by the shard's 2 MB/s read limit. Option C is wrong because adding more EC2 instances without increasing the number of shards does not increase the total read capacity; each shard can only be consumed by one record processor at a time (within a single KCL application), so additional instances will remain idle or cause contention.

220
Multi-Selectmedium

Which THREE of the following are valid storage classes in Amazon S3? (Choose THREE.)

Select 3 answers
A.S3 Standard
B.S3 Archive
C.S3 Intelligent-Tiering
D.S3 Cold
E.S3 One Zone-IA
AnswersA, C, E

S3 Standard is a genuine storage class, designed for frequently accessed data with low latency and high throughput, replicated across a minimum of three Availability Zones. It satisfies the question's requirement for a valid Amazon S3 storage class.

Why this answer

S3 Standard (A) is a valid storage class designed for frequently accessed data with low latency and high throughput, and it is the default class for S3 objects. S3 Intelligent-Tiering (C) is a valid class that automatically moves objects between access tiers based on changing access patterns, charging a small monitoring and automation fee. S3 One Zone-IA (E) is a valid class that stores data in a single Availability Zone at lower cost, intended for infrequently accessed data that can be easily recreated.

The other options are not real S3 storage classes: S3 Archive (B) does not exist — the archival class is S3 Glacier, and S3 Cold (D) is not a class either, as the closest real classes are S3 Glacier Instant Retrieval and S3 Glacier Flexible Retrieval.

Exam trap

AWS often tests the distinction between valid S3 storage classes and fabricated names like 'S3 Archive' or 'S3 Cold', expecting candidates to recall the exact naming conventions (e.g., S3 Glacier, S3 Glacier Deep Archive) rather than generic terms.

221
Multi-Selectmedium

A data engineer is designing a disaster recovery strategy for an Amazon RDS for PostgreSQL database. The primary database is in us-east-1. Which TWO approaches provide cross-region disaster recovery?

Select 2 answers
A.Configure cross-region automated backups to copy to us-west-2.
B.Take a manual snapshot and copy it to us-west-2 daily.
C.Use Amazon S3 cross-region replication for the database export.
D.Enable Multi-AZ in us-east-1.
E.Create a cross-region read replica in us-west-2.
AnswersA, E

Backups are automatically copied and can be restored.

Why this answer

Amazon RDS supports cross-region automated backups, which automatically copy backup data (snapshots and transaction logs) from the primary region (us-east-1) to a secondary region (us-west-2). This provides a fully managed, automated disaster recovery solution that allows point-in-time recovery in the secondary region without manual intervention.

Exam trap

The trap here is that candidates often confuse Multi-AZ (which provides in-region high availability) with cross-region disaster recovery, or they assume manual snapshot copying is equivalent to automated cross-region backups, not realizing the significant difference in RPO and operational overhead.

222
Multi-Selectmedium

A data engineer needs to ensure that sensitive data stored in Amazon S3 is encrypted at rest. Which TWO options meet this requirement? (Choose TWO.)

Select 2 answers
A.Server-Side Encryption with AWS KMS-Managed Keys (SSE-KMS)
B.Server-Side Encryption with S3-Managed Keys (SSE-S3)
C.Using a VPC to restrict network access
D.Enabling MFA Delete on the S3 bucket
E.Client-Side Encryption with SSL/TLS
AnswersA, B

SSE-KMS encrypts objects at rest using keys managed in AWS KMS, satisfying the encryption-at-rest requirement. It provides envelope encryption with an independent data key per object, plus audit trails via CloudTrail and granular access control through key policies — unlike SSE-S3, which offers no separate key permissions or key-usage auditing.

Why this answer

Options A (SSE-KMS) and B (SSE-S3) are correct because both are server-side encryption mechanisms that encrypt S3 objects at rest: SSE-KMS uses AWS KMS customer master keys (CMKs) to generate and manage data keys, while SSE-S3 uses AES-256 keys fully managed by Amazon S3. Both satisfy the requirement that sensitive data stored in S3 be encrypted at rest, and each is applied per-object when the object is written to the bucket. Option C is incorrect because a VPC only controls network-level access to S3 (via endpoints and policies) and does not encrypt data at rest.

Option D is incorrect because MFA Delete only adds an authentication requirement for deleting objects or changing versioning state; it provides no encryption. Option E is incorrect because SSL/TLS encrypts data in transit, not at rest, and client-side encryption is a separate approach not represented by that option.

Exam trap

The trap here is that candidates often confuse encryption in transit (SSL/TLS) with encryption at rest, or they mistakenly think network controls like VPCs or access controls like MFA Delete provide data encryption, when they only address different security domains.

223
MCQhard

A company is using Amazon ElastiCache for Redis to cache frequently accessed data. The cache hit ratio is low, and the engineering team suspects that the eviction policy is causing important data to be removed. Which eviction policy should be used to minimize eviction of the most frequently accessed keys?

A.allkeys-lru
B.allkeys-lfu
C.noeviction
D.volatile-lru
AnswerB

allkeys-lfu evicts the least frequently used keys across the entire keyspace, so frequently accessed items survive even when newer keys arrive. This directly raises the hit ratio by retaining hot keys, unlike LRU or volatile-ttl policies that discard them based on recency or expiry.

Why this answer

The allkeys-lfu (Least Frequently Used) eviction policy is the correct choice because it explicitly tracks and retains keys that are accessed most frequently across the entire keyspace. Since the cache hit ratio is low due to eviction of important data, LFU ensures that frequently accessed keys are evicted last, directly addressing the problem of important data being removed.

Exam trap

The trap here is that candidates often confuse recency (LRU) with frequency (LFU), assuming that 'least recently used' also implies 'least frequently used,' but LRU can evict a frequently accessed key that hasn't been touched recently, which is exactly the problem described.

How to eliminate wrong answers

Option A is wrong because allkeys-lru (Least Recently Used) evicts keys based on recency of access, not frequency, so a frequently accessed key that hasn't been used recently could be evicted. Option C is wrong because noeviction returns errors for write operations when memory is full, which would cause application failures rather than solving the low hit ratio. Option D is wrong because volatile-lru only applies to keys with a TTL set, leaving keys without TTLs unprotected and potentially evicting important data that lacks an expiration.

224
MCQhard

A data engineer is designing an Amazon DynamoDB table for an order-processing application. The table uses a partition key of order_id and a sort key of order_date. The application needs to retrieve all orders for a specific customer within a date range, and the queries must be efficient at scale. The engineer must choose a design that supports these access patterns without full table scans. What should the engineer do?

A.Enable DynamoDB Streams and use a Lambda function to write customer-based query results to a separate table.
B.Create a local secondary index with customer_id as the partition key and order_date as the sort key.
C.Create a global secondary index with customer_id as the partition key and order_date as the sort key.
D.Use a Scan operation with a FilterExpression on customer_id and order_date.
AnswerC

A global secondary index with customer_id as the partition key and order_date as the sort key allows the application to query by customer and date range directly. DynamoDB can then use the index to retrieve only the relevant items, avoiding a full table scan. This design matches the required access pattern and scales horizontally with the table.

Why this answer

The application needs to query orders by customer and date range efficiently. A global secondary index can use a different partition key and sort key from the base table, so customer_id as the partition key and order_date as the sort key provides the required access path. Local secondary indexes must share the base table partition key, scans are inefficient, and stream-based copies do not replace a queryable index.

Exam trap

The trap here is confusing local secondary indexes with global secondary indexes, because a local secondary index cannot use a different partition key from the base table.

225
MCQhard

A data engineer is managing an Amazon Redshift cluster that experiences performance degradation during peak query loads. The cluster uses a dc2.large node type with 2 nodes. Analysis shows that the workload involves frequent large table scans and complex joins. The engineer wants to improve query performance without changing the overall data volume. Which action should the engineer take?

A.Change the distribution style of large tables to KEY on the join columns and enable sort keys on frequently filtered columns.
B.Increase the number of nodes in the cluster by resizing to a larger node type with more memory.
C.Enable concurrency scaling to automatically add transient clusters during peak loads.
D.Enable short query acceleration (SQA) to prioritize short-running queries.
AnswerA

Choosing an appropriate distribution key ensures that join data is collocated on the same nodes, reducing data movement during joins. Sort keys on filtered columns enable efficient range scans. Together, these optimizations significantly improve performance for large scans and complex joins on a small cluster.

Why this answer

Proper distribution and sort keys are fundamental Redshift tuning techniques. Distributing large tables on join keys collocates data, minimizing network traffic during joins. Sort keys on filter columns allow the query engine to skip blocks.

These changes directly address the performance bottlenecks without adding hardware, making them the most effective and cost-efficient solution.

Exam trap

The trap here is assuming that adding more nodes or enabling concurrency scaling will automatically fix performance issues, when the root cause is suboptimal data distribution and sorting.

Page 2

Page 3 of 18

Page 4