Courseiva

AWS Certified Data Engineer Associate DEA-C01 (DEA-C01) — Questions 226–300

1321 questions total · 18pages · All types, answers revealed

Page 3

Page 4 of 18

Page 5
226
Multi-Selecteasy

Which THREE actions can help improve read performance in Amazon DynamoDB? (Choose THREE.)

Select 3 answers
A.Use DynamoDB global tables to replicate data.
B.Use parallel scans to distribute read load across partitions.
C.Use strongly consistent reads for all queries.
D.Enable DynamoDB Accelerator (DAX) to cache reads.
E.Increase the read capacity units (RCU) for the table.
AnswersB, D, E

Parallel scans can improve scan performance.

Why this answer

Parallel scans in DynamoDB can improve read performance by dividing a scan operation into multiple segments that are processed concurrently across partitions. This reduces the overall latency of the scan by leveraging the distributed nature of DynamoDB's storage, though it consumes more read capacity units (RCUs) due to the parallel execution.

Exam trap

The DEA-C01 exam often tests the misconception that strongly consistent reads always improve performance, when in fact they increase latency and RCU consumption, making eventually consistent reads the better choice for read-heavy workloads.

227
MCQeasy

A data engineer is ingesting streaming data from thousands of IoT devices into AWS. The data is JSON-formatted and must be stored in Amazon S3 for long-term analytics. Which service is most appropriate for real-time ingestion and routing to S3?

A.Amazon SQS
B.Amazon Kinesis Data Firehose
C.Amazon Kinesis Data Streams
D.AWS Glue
AnswerB

Kinesis Data Firehose ingests streaming records and delivers them to Amazon S3 with built-in buffering, compression and format conversion, requiring no consumer code. That satisfies real-time ingestion from thousands of IoT devices plus durable S3 storage for later analytics.

Why this answer

Amazon Kinesis Data Firehose is the most appropriate service because it is designed for real-time ingestion of streaming data and can directly deliver data to Amazon S3 without requiring custom code. It automatically handles buffering, compression, and partitioning of JSON data, making it ideal for long-term analytics storage.

Exam trap

The trap here is that candidates often confuse Kinesis Data Streams with Kinesis Data Firehose, assuming both can directly write to S3, but Data Streams requires a downstream consumer to perform the write, making Firehose the correct choice for direct, managed ingestion to S3.

How to eliminate wrong answers

Option A is wrong because Amazon SQS is a message queue service for decoupling application components, not a streaming ingestion service; it lacks built-in data transformation and direct S3 delivery capabilities. Option C is wrong because Amazon Kinesis Data Streams is a real-time data streaming service that requires a separate consumer (e.g., Lambda or Firehose) to write data to S3, adding complexity and latency; it is not a direct ingestion-to-S3 solution. Option D is wrong because AWS Glue is a serverless ETL service for batch data processing and cataloging, not designed for real-time streaming ingestion or direct routing to S3.

228
MCQmedium

A data engineer is building a data lake on Amazon S3. The raw data arrives as JSON files, but the analytics team needs to query the data using standard SQL in Amazon Athena with optimal performance and minimal cost. The engineer wants to convert the JSON to a columnar format that supports predicate pushdown and is natively supported by Athena. Which storage format should the engineer choose?

A.Apache Parquet
B.CSV
C.JSON
D.Apache Avro
AnswerA

Apache Parquet is a columnar format that stores data by column, enabling Athena to read only the columns referenced in a query. This reduces I/O and cost. Parquet also supports predicate pushdown and is natively supported by Athena, making it ideal for this scenario where the goal is optimal query performance and minimal cost.

Why this answer

Apache Parquet is the best choice because it is a columnar format that allows Athena to read only the necessary columns, reducing data scanned and cost. It also supports predicate pushdown, which filters data at the storage level. Other formats like Avro, JSON, and CSV are row-based and less efficient for analytical queries in Athena.

Exam trap

The trap here is assuming that any format supported by Athena is equally performant, but columnar formats like Parquet provide significant advantages for query performance and cost.

229
MCQmedium

A data engineer manages an Amazon S3 data lake with millions of small JSON files under prefixes partitioned by year/month/day. Amazon Athena queries scan far more data than expected and return slowly. The engineer wants to reduce bytes scanned and improve query performance while keeping files in S3 and queryable by Athena. Which solution meets these requirements with the LEAST operational overhead?

A.Enable S3 Transfer Acceleration on the bucket and configure Athena workgroups to use it for all queries.
B.Use AWS Glue to convert the JSON files to Apache Parquet, write them to partitioned prefixes, and register the resulting tables in the AWS Glue Data Catalog.
C.Attach an S3 Lifecycle policy that transitions the JSON objects to S3 Glacier Instant Retrieval after 30 days and query them through Athena.
D.Create an Amazon S3 Inventory report for each partition and use it to rewrite the Athena queries so they reference only the inventory files.
AnswerB

Converting to columnar Parquet with partitioning lets Athena read only needed columns and partitions, cutting bytes scanned. Registering the tables in the Data Catalog makes them immediately queryable without managing a separate metastore, and Glue handles the conversion job, keeping operational overhead low compared with self-managed ETL.

Why this answer

Athena charges and performs based on bytes scanned, and row-oriented JSON forces reading entire objects. Converting to columnar Parquet lets the engine read only referenced columns, while partitioning by year/month/day enables partition pruning so irrelevant dates are skipped. Registering the converted data in the AWS Glue Data Catalog makes it available to Athena with minimal management, directly reducing scanned bytes and latency.

Exam trap

The trap here is assuming that changing storage class or accelerating transfer improves Athena query cost and speed, when the real driver is data format and partition pruning.

230
MCQhard

A healthcare company is ingesting patient data from a legacy system into an Amazon S3 data lake using AWS Glue. The legacy system produces CSV files with inconsistent schemas (columns may appear or disappear in different files). The data engineer needs to create a Glue ETL job that can handle schema evolution and transform the data into a standardized parquet format. The job should also be able to process new files as they arrive. Which approach should the data engineer use?

A.Use AWS Glue crawlers to create a schema in the Data Catalog and then use a standard Spark DataFrame for transformation.
B.Use AWS Glue DynamicFrames to read the CSV files and apply transformations using resolveChoice and applyMapping.
C.Use a Python shell job in Glue to manually parse each file and write to parquet.
D.Use a Glue ETL job with a static schema defined in the script and ignore files that don't match.
AnswerB

DynamicFrames natively accommodate schema evolution: resolveChoice reconciles columns that appear or disappear across files by casting or dropping them, while applyMapping standardises surviving fields before writing Parquet. This satisfies the inconsistent-schema constraint, and Glue job bookmarks let the same job process newly arrived files incrementally.

Why this answer

AWS Glue DynamicFrames are designed to handle schema evolution and inconsistent data. Using resolveChoice to handle columns that appear/disappear and applyMapping to standardize the schema allows the job to process files with varying schemas and output consistent Parquet.

Exam trap

The trap is assuming a static schema or standard Spark DataFrame can handle schema evolution; candidates must recognize DynamicFrames as the Glue-native solution for inconsistent schemas.

How to eliminate wrong answers

Option A is wrong because Glue crawlers create a static schema in the Data Catalog; a standard Spark DataFrame would fail on files with different schemas. Option C is wrong because a Python shell job lacks the distributed processing and built-in schema evolution capabilities of Glue ETL with DynamicFrames. Option D is wrong because ignoring files that don't match a static schema would drop data, which is unacceptable for patient data.

231
Multi-Selecteasy

A data engineer is monitoring Amazon CloudWatch metrics for an Amazon Redshift cluster and notices high CPU utilization. The engineer wants to reduce CPU usage. Which TWO actions should the engineer take?

Select 2 answers
A.Enable concurrency scaling to offload read queries to additional clusters.
B.Increase the number of nodes in the cluster.
C.Optimize the table design by using sort keys and compression.
D.Run the VACUUM command on all tables.
E.Enable audit logging to monitor queries.
AnswersA, C

Concurrency scaling adds transient Redshift clusters that serve read queries, so eligible read workloads move off the main cluster and its CPU utilisation drops. This directly addresses the high CPU constraint by offloading read query processing rather than resizing or tuning the existing cluster.

Why this answer

Option A is correct because concurrency scaling automatically adds transient cluster capacity to handle bursts of read (SELECT) queries, offloading that work from the main cluster and thereby reducing its CPU utilization. Option C is correct because choosing appropriate sort keys and compression encodings reduces the amount of data scanned and I/O performed, which directly lowers CPU work during query execution. Option B is not the best fit because adding nodes increases capacity but does not address inefficient queries or design, and it is a costlier, less targeted remedy for CPU pressure.

Option D is not correct because VACUUM reclaims space and re-sorts rows after deletes/updates; it is maintenance rather than a primary CPU-reduction action and can itself consume significant resources. Option E is not correct because audit logging only records connection and user activity for monitoring/compliance and does nothing to reduce CPU usage.

Exam trap

DEA-C01 often tests whether candidates confuse scaling out (adding nodes) with optimizing query and table design, so 'add more nodes' looks appealing but does not reduce CPU utilization per node.

232
MCQhard

A company stores sensitive customer data in an Amazon S3 bucket with versioning enabled. A data engineer accidentally deleted the current version of an object. What is the quickest way to restore the object to its previous state without additional data transfer costs?

A.Use S3 Batch Operations to restore the object from the Recycle Bin.
B.Delete the delete marker that was created by the deletion.
C.Copy the previous version from the bucket to itself.
D.Use the S3 sync command to restore the previous version.
AnswerB

Deleting an object in a versioned bucket inserts a delete marker rather than erasing prior versions. Removing that delete marker restores the object to its previous state instantly, with no data transfer and no re-upload, since the original version's bytes were never removed from the bucket.

Why this answer

With S3 versioning enabled, deleting an object does not permanently remove it; instead, a delete marker is placed. To restore the object to its previous state, you simply remove the delete marker, which makes the previous version the current version again. Option A is incorrect because S3 does not have a Recycle Bin; S3 Batch Operations are for bulk actions but not for restoring from a recycle bin.

Option C would not restore the object properly; copying the previous version to itself would create a new version, not restore the original. Option D is incorrect because the s3 sync command synchronizes objects between locations and does not restore previous versions.

233
MCQhard

A data pipeline uses AWS Glue ETL jobs to process data from Amazon RDS for MySQL to Amazon S3. Recently, the jobs have been failing with the error 'Communications link failure' during the connection phase. The RDS instance is in a private subnet, and the Glue job uses a VPC endpoint for S3. What is the most likely cause?

A.The RDS database has reached the maximum number of connections.
B.The Glue job does not have IAM permissions to decrypt the RDS database using AWS KMS.
C.The JDBC driver used by Glue is incompatible with the MySQL version.
D.The Glue job does not have a network path to the RDS instance because it is not attached to the same VPC subnet.
AnswerD

AWS Glue jobs run inside a VPC only when attached to a subnet; without that attachment they cannot route to a private RDS instance. The S3 VPC endpoint does not provide a path to RDS, so the connection fails during the handshake.

Why this answer

The most likely cause is that the AWS Glue job does not have a network path to the RDS instance because it is not attached to the same VPC subnet. AWS Glue jobs run in a serverless environment that, by default, is outside the customer's VPC. To connect to a private RDS instance, the Glue job must be configured with a VPC connection that specifies the VPC, subnet, and security group, allowing it to reach the RDS instance.

Without this, the JDBC connection fails with 'Communications link failure' during the connection phase.

Exam trap

DEA-C01 often tests the misconception that enabling a VPC endpoint for S3 automatically provides network access to other VPC resources like RDS, causing candidates to overlook the need for a separate VPC connection for Glue.

How to eliminate wrong answers

Option A is wrong because reaching the maximum number of connections would typically produce a 'Too many connections' error, not 'Communications link failure'; also, the error occurs during the connection phase, which suggests a network issue rather than connection exhaustion. Option B is wrong because IAM permissions to decrypt the RDS database using KMS are not required for a JDBC connection; KMS decryption is for encrypted data at rest, and the error would be an access denied or encryption-related error, not a communications link failure. Option C is wrong because JDBC driver incompatibility would typically manifest as a version mismatch error or a specific SQL exception, not a generic 'Communications link failure'; moreover, AWS Glue includes compatible JDBC drivers for MySQL.

234
MCQeasy

A company uses AWS Glue to process sensitive data stored in Amazon S3. The security team requires that all data in transit between AWS Glue and S3 be encrypted. Which configuration should be used to meet this requirement?

A.Use an S3 bucket policy that denies requests not using HTTPS.
B.Use an AWS KMS key to encrypt the data before uploading to S3.
C.Configure AWS Glue to use SSL by setting the 'ssl' parameter to 'true'.
D.Enable default encryption on the S3 bucket using SSE-S3.
AnswerA

An S3 bucket policy denying requests where aws:SecureTransport is false enforces TLS on every request, so Glue-to-S3 traffic cannot fall back to plain HTTP. This satisfies the in-transit encryption requirement at the bucket level, covering all clients including AWS Glue.

Why this answer

Requiring HTTPS for all requests to the S3 bucket ensures that data in transit between AWS Glue and S3 is encrypted using TLS. By using an S3 bucket policy with a condition that denies requests where `aws:SecureTransport` is false, the company enforces encryption for all connections, including those from AWS Glue. This meets the security requirement without needing to modify Glue or S3 configurations beyond the bucket policy.

Exam trap

The trap here is that candidates often confuse encryption at rest (SSE-S3, SSE-KMS, client-side encryption) with encryption in transit (TLS/HTTPS), and may incorrectly assume that enabling default encryption or using KMS keys secures the data during transfer.

How to eliminate wrong answers

Option B is wrong because encrypting data with an AWS KMS key before uploading to S3 (client-side encryption) protects data at rest, not data in transit; the security team specifically requires encryption in transit. Option C is wrong because AWS Glue does not have an 'ssl' parameter; Glue uses HTTPS by default when connecting to S3, and this setting is not configurable via a simple parameter. Option D is wrong because enabling default encryption on the S3 bucket (SSE-S3) only encrypts data at rest, not data in transit between Glue and S3.

235
MCQmedium

An e-commerce company ingests clickstream data from their website into Amazon S3. The data is in JSON format, and each file is about 10 MB. They need to transform the data into a columnar format for analytics and load it into Amazon Redshift nightly. The transformation should be cost-effective and require minimal operational overhead. Which approach meets these requirements?

A.Use AWS Glue ETL job to convert to Parquet and load into Redshift.
B.Use Amazon Redshift COPY command to load JSON directly.
C.Use Amazon EMR with Spark to transform and load data.
D.Use AWS Lambda to transform each file and write to Redshift.
AnswerA

AWS Glue ETL jobs convert JSON to Parquet on serverless Spark infrastructure, eliminating cluster provisioning and satisfying the minimal-operational-overhead constraint. Parquet's columnar layout and compression reduce Redshift storage and scan costs, meeting the cost-effectiveness requirement for the nightly 10 MB file transformations.

Why this answer

AWS Glue ETL is the correct choice because it is a serverless, managed service that can efficiently convert JSON to Parquet (a columnar format optimized for Redshift) and load the data into Redshift with minimal operational overhead. The nightly batch processing of 10 MB files is well-suited for Glue's pay-per-use pricing, making it cost-effective without requiring infrastructure management.

Exam trap

The trap here is that candidates may choose Amazon EMR or Lambda because they are familiar with Spark or serverless functions, but they overlook the operational overhead of EMR and the execution limits of Lambda for batch workloads, while Glue provides a balanced, managed solution for this specific use case.

How to eliminate wrong answers

Option B is wrong because the Redshift COPY command can load JSON directly, but it does not transform the data into a columnar format like Parquet; it loads JSON as-is, which is less efficient for analytics and may require additional schema handling. Option C is wrong because Amazon EMR with Spark introduces significant operational overhead for managing clusters, tuning, and monitoring, which is unnecessary for a simple nightly transformation of small 10 MB files. Option D is wrong because AWS Lambda has a maximum execution timeout of 15 minutes and limited memory (up to 10 GB), making it unsuitable for batch processing multiple files or handling large datasets; it is designed for event-driven, short-lived tasks, not nightly ETL workloads.

236
MCQhard

A data engineer is troubleshooting a failed AWS Glue job that reads from an Apache Hive metastore in an Amazon EMR cluster. The error message indicates 'ClassNotFoundException: org.apache.hadoop.hive.ql.metadata.HiveException'. The Glue job uses a custom Python shell script. What is the most likely cause of this error?

A.Check the network connectivity between Glue and the EMR cluster.
B.Include the Hive JAR files in the 'Python library path' or use a Glue version with Hive support.
C.Modify the Python script to import the Hive libraries manually.
D.Update the IAM role to allow 'hive:Describe*' actions.
AnswerB

The ClassNotFoundException for the Hive metastore class means the Hive client JARs are absent from the Python shell job's classpath. Adding the Hive JARs to the Python library path, or using a Glue version bundling Hive support, supplies the missing classes.

Why this answer

The 'ClassNotFoundException' for a Hive class indicates the Hive JARs are not on the classpath at runtime. AWS Glue's Python shell jobs run in an environment that does not include Hive libraries by default, so the engineer must either add the Hive JARs to the Python library path or use a Glue version/configuration that bundles Hive support. This is a classpath/dependency issue, not a network or IAM problem.

Exam trap

DEA-C01 often tests the confusion between dependency/classpath errors and network or IAM errors — candidates see 'Hive' and reach for IAM or connectivity fixes when the error is a missing JAR on the classpath.

How to eliminate wrong answers

Option A is wrong because a network connectivity failure would produce a connection timeout or 'UnknownHostException', not a 'ClassNotFoundException' — the JVM cannot even find the class locally. Option C is wrong because manually importing Hive libraries in Python does not resolve a missing JAR; the class must be present on the JVM classpath, and Python imports cannot load Java classes that are absent. Option D is wrong because 'hive:Describe*' is not a valid IAM action for Hive metastore access, and IAM permission errors manifest as 'AccessDeniedException', not 'ClassNotFoundException'.

237
MCQmedium

A data engineer maintains an AWS Glue job that reads JSON files from Amazon S3, applies a transform, and writes Parquet to a second bucket. The job's bookmark was enabled at creation, but each nightly run reprocesses all previously handled files, and downstream tables now contain duplicate rows. The job script has not been modified and the S3 prefix is unchanged. Which action will MOST directly resolve the duplicate processing?

A.Reprocess the prefix with a job that has job bookmarks disabled and rely on the S3 object LastModified timestamp to filter files.
B.Confirm the transformation_ctx parameter is passed to each source and sink call, then reset the job bookmark and rerun once to rebuild state.
C.Increase the number of AWS Glue DPUs allocated to the job so the run completes before the next scheduled trigger.
D.Change the job's output write mode to append and add a deduplication step that drops rows whose keys already exist.
AnswerB

Job bookmarks rely on the transformation_ctx value to namespace state per source and sink; when it is absent or has changed between runs, Glue cannot correlate prior state and falls back to reading everything. Supplying a stable transformation_ctx on the S3 source and sink, then resetting the bookmark to clear stale state, restores correct incremental processing.

Why this answer

AWS Glue job bookmarks persist per-source and per-sink state keyed by the transformation_ctx argument. If that context is missing or inconsistent, the job cannot determine which objects were already processed and re-reads the entire prefix, producing duplicates. Passing a stable transformation_ctx and resetting the bookmark to rebuild the state store fixes the incremental behavior without changing the transform logic.

Exam trap

The trap here is assuming duplicated output always means the transformation is non-deterministic, when the usual cause is bookmark state that cannot be matched to a source or sink.

238
MCQeasy

A data engineer is tasked with transforming JSON data from an S3 bucket into Parquet format for efficient querying. The transformation should run on a schedule every hour. Which AWS service is best suited for this task?

A.AWS Lambda
B.Amazon Athena
C.AWS Glue
D.Amazon EMR
AnswerC

AWS Glue satisfies the hourly schedule constraint through time-based triggers, and its ETL engine converts JSON to Parquet using built-in transforms that infer schemas automatically. Crawlers catalogue the S3 source, while the Parquet output is written directly to S3 for efficient downstream querying by Athena or Redshift Spectrum.

Why this answer

AWS Glue is the best choice because it is a fully managed ETL service designed specifically for transforming and cataloging data at scale. It can natively read JSON from S3, convert it to Parquet, and run on a scheduled hourly basis using a Glue job with a trigger, without requiring server management or custom infrastructure.

Exam trap

The trap here is that candidates often confuse Athena's ability to query Parquet with the ability to transform data into Parquet, but Athena is a query engine, not an ETL service, and cannot perform scheduled data format conversions.

How to eliminate wrong answers

Option A is wrong because AWS Lambda has a maximum execution timeout of 15 minutes and a 10 GB memory limit, making it unsuitable for processing large JSON datasets or running long-running hourly transformations. Option B is wrong because Amazon Athena is an interactive query service for analyzing data directly in S3, not a transformation engine; it cannot convert JSON to Parquet and write the output back to S3 in a scheduled, automated manner. Option D is wrong because Amazon EMR requires provisioning and managing a cluster of EC2 instances, which adds operational overhead and cost, whereas the task calls for a serverless, scheduled transformation with minimal management.

239
MCQmedium

A data engineer maintains an Amazon S3 data lake with millions of small JSON objects. The engineer needs to improve query performance by reducing the number of objects and compressing them into a columnar format that Amazon Athena can query efficiently. The data must remain partitioned by date. Which solution should the engineer use?

A.Enable S3 Transfer Acceleration on the bucket and run Amazon Athena queries with the existing JSON objects.
B.Use AWS Glue ETL to read the JSON objects, repartition by date, write them as Parquet with Snappy compression, and register the resulting tables in the AWS Glue Data Catalog.
C.Use Amazon S3 Lifecycle policies to transition the JSON objects to S3 Glacier Instant Retrieval and query them with Athena.
D.Create an Amazon EMR cluster and run a MapReduce job that merges the JSON files into larger JSON files without changing the format.
AnswerB

AWS Glue ETL can transform JSON to Parquet, repartition by date, and update the Data Catalog, which Athena uses for schema and partition metadata. Parquet with Snappy reduces storage and scan size, and fewer larger files improve query performance. This directly addresses the small-file problem and columnar requirement while preserving date partitioning.

Why this answer

Converting small JSON files to Parquet with Snappy compression and repartitioning by date reduces storage footprint and the amount of data scanned by Athena. AWS Glue ETL can perform this transformation and update the Data Catalog so Athena queries the new tables. This is the standard pattern for optimizing S3 data lakes for analytics.

Exam trap

The trap here is assuming that simply merging files or changing storage classes solves the small-file problem, when the key is converting to a columnar format and repartitioning.

240
MCQmedium

A data engineer notices that an AWS Glue ETL job processing data from Amazon S3 to Amazon Redshift has been failing intermittently with the error 'S3ServiceException: SlowDown'. Which action is MOST likely to resolve this issue?

A.Increase the number of partitions in the Glue job to parallelize reads.
B.Switch from a Standard to a G.2X large Glue worker type.
C.Implement exponential backoff and retry logic in the Glue job.
D.Enable S3 Transfer Acceleration on the source bucket.
AnswerC

S3 SlowDown is a throttling response signalling too many concurrent requests to a partition. Exponential backoff with retries spaces requests progressively, letting the Glue job ride out the throttle rather than failing, which resolves the intermittent S3ServiceException errors.

Why this answer

The 'S3ServiceException: SlowDown' error indicates that the AWS Glue job is making requests to Amazon S3 at a rate that exceeds the bucket's request rate limits. Implementing exponential backoff and retry logic (option C) is the most effective solution because it reduces the effective request rate by introducing delays between retries, allowing S3 to recover from throttling. Option A is incorrect because increasing partitions would likely increase the number of concurrent requests, exacerbating throttling.

Option B is incorrect because switching to a larger worker type does not affect the rate of S3 requests. Option D is incorrect because S3 Transfer Acceleration improves network transfer speed but does not reduce request throttling.

241
MCQmedium

An organization needs to audit all access to their S3 buckets for compliance purposes. They want to log both successful and failed API calls. Which AWS service should be used?

A.Amazon CloudWatch Logs
B.AWS Config
C.AWS CloudTrail
D.VPC Flow Logs
AnswerC

CloudTrail records S3 data-plane and management API activity, capturing both successful and failed calls with identity, timestamp and source IP. This satisfies the compliance requirement to audit all bucket access, which S3 server access logging cannot match for failed calls.

Why this answer

AWS CloudTrail is the correct service because it records all API calls made to S3, including both successful and failed requests, and delivers log files to an S3 bucket for auditing and compliance. CloudTrail captures management events (e.g., CreateBucket) and, when enabled, data events (e.g., GetObject, PutObject) for S3, providing a complete audit trail of access.

Exam trap

The trap here is that candidates often confuse AWS CloudTrail (API auditing) with AWS Config (configuration auditing) or VPC Flow Logs (network traffic logging), failing to recognize that only CloudTrail captures the specific API call details needed for access auditing.

How to eliminate wrong answers

Option A is wrong because Amazon CloudWatch Logs is used for monitoring, storing, and accessing log files from various AWS services (e.g., EC2, Lambda), but it does not natively capture S3 API calls; it can only receive logs forwarded from CloudTrail or other sources. Option B is wrong because AWS Config is a service for evaluating resource configurations against desired policies and tracking configuration changes, not for logging API calls or access events. Option D is wrong because VPC Flow Logs capture IP traffic metadata (source/destination IP, ports, protocol) at the network interface level, not S3 API-level operations or authentication details.

242
MCQhard

A company runs an Amazon Redshift cluster for analytics. During peak hours, query performance degrades significantly. The data engineer notices that disk space usage is above 80% on many nodes. Which of the following is the MOST effective long-term solution to improve query performance?

A.Increase the workload management (WLM) queue slots.
B.Resize the cluster to include additional nodes.
C.Apply compression encoding to all columns.
D.Run the VACUUM command to reclaim space.
AnswerB

Adding nodes increases both storage capacity and compute parallelism across the cluster, relieving the 80% disk pressure while distributing query workload. This addresses the root cause of degradation rather than temporarily masking it, providing a durable performance improvement for peak-hour analytics.

Why this answer

Resizing the cluster to include additional nodes increases both storage and compute capacity, directly addressing the high disk usage and improving query performance. Increasing WLM queue slots (Option A) only manages concurrency but does not add capacity. Compression encoding (Option C) reduces storage but may not alleviate immediate performance degradation, and is not a long-term solution for capacity.

Running VACUUM (Option D) reclaims space from deleted rows but does not add new capacity.

243
MCQeasy

A data engineer needs to store semi-structured JSON event logs in a data lake on Amazon S3 and query them with Amazon Athena using SQL, including filtering on individual JSON attributes. The team wants to avoid transforming the files before querying. Which approach should the engineer use?

A.Store the logs as CSV and use Athena to parse the columns
B.Use Amazon CloudWatch Logs Insights to query the JSON logs in place
C.Load the JSON into Amazon Redshift using COPY JSON and query it there
D.Define an Athena table with a struct or map column over the JSON and query nested fields
AnswerD

Athena, built on the AWS Glue Data Catalog, supports JSON SerDe and complex types such as struct, array, and map. Defining columns with these types lets SQL reference nested attributes directly, for example using dot notation on a struct field. No preprocessing is needed, which satisfies the requirement to query raw JSON in place while still allowing filters on individual attributes.

Why this answer

Athena queries data in place on S3 using table definitions in the AWS Glue Data Catalog. For semi-structured JSON, defining columns as struct, array, or map with the appropriate JSON SerDe lets SQL access nested attributes directly without preprocessing. Storing as CSV would require flattening, loading into Redshift adds a transformation step, and CloudWatch Logs Insights is scoped to log groups rather than the S3 data lake.

Exam trap

The trap here is assuming JSON must be flattened or loaded into a warehouse before SQL can filter individual attributes, when Athena supports complex types that expose nested fields directly.

244
MCQmedium

A company uses AWS Glue ETL jobs to process data from an S3 data lake. The job reads data in CSV format, transforms it, and writes to Parquet. The job runs daily and takes 2 hours to complete. The data volume is increasing by 20% each month. The engineer wants to reduce the job runtime. Which action is most effective?

A.Increase the number of DPUs for the Glue job
B.Enable compression on the input CSV files
C.Switch from Python Shell to Spark ETL
D.Partition the input data in S3 by date and use partition pruning in the job
AnswerD

Partitioning S3 input by date lets Glue read only the relevant partitions, cutting the bytes scanned each run. This addresses the growing data volume constraint far more effectively than scaling workers or changing formats alone.

Why this answer

Most effective because partitioning the input data by date and using partition pruning allows the Glue ETL job to read only the relevant partitions instead of scanning the entire S3 data lake. This drastically reduces the amount of data processed, which directly addresses the growing data volume and shortens job runtime. Partition pruning is a core optimization for Spark-based Glue jobs, as it leverages Hive-style partitioning to skip unnecessary files.

Exam trap

The trap here is that candidates often assume increasing DPUs or enabling compression is the universal fix, but they fail to recognize that reducing the data scanned via partition pruning is the most impactful optimization for growing datasets in S3-based Glue jobs.

How to eliminate wrong answers

Option A is wrong because increasing DPUs (Data Processing Units) adds more parallelism but does not reduce the volume of data read; it may help only if the job is CPU-bound, but the primary bottleneck here is the increasing data volume, not compute capacity. Option B is wrong because enabling compression on input CSV files reduces storage size and I/O overhead, but CSV is not splittable when compressed (e.g., Gzip), which can actually harm parallelism and increase runtime; moreover, the job still reads all data. Option C is wrong because the question states the job already uses AWS Glue ETL, which is Spark-based by default; switching from Python Shell to Spark ETL would be a regression, as Python Shell is single-node and slower for large datasets, but the current job is already using Spark (implied by Glue ETL), so this change is irrelevant or counterproductive.

245
MCQmedium

A streaming application sends data to Amazon Kinesis Data Streams. The data must be enriched with reference data from an Amazon DynamoDB table in real-time. Which AWS service can be used to perform this enrichment with minimal latency?

A.Amazon Kinesis Data Analytics for Apache Flink
B.Amazon Kinesis Data Firehose with Lambda transformation
C.AWS Lambda function triggered by Kinesis Data Streams
D.AWS Glue streaming ETL
AnswerA

Amazon Kinesis Data Analytics for Apache Flink runs continuous SQL or Flink applications directly over the stream, joining each record against the DynamoDB reference table in-flight. This satisfies the real-time enrichment requirement with minimal latency, avoiding the batching delay of Lambda-based or Firehose-based approaches.

Why this answer

Amazon Kinesis Data Analytics for Apache Flink is correct because it allows you to run Apache Flink applications that can read from a Kinesis data stream, perform stateful stream processing, and enrich records in real-time by joining with reference data stored in DynamoDB. Flink's asynchronous I/O and managed state enable sub-second enrichment latency without the cold-start delays or concurrency limits of Lambda-based approaches.

Exam trap

The DEA-C01 exam often tests the distinction between real-time stream processing (Kinesis Data Analytics for Flink) and near-real-time or batch-oriented services (Firehose, Glue ETL), leading candidates to choose Lambda because they assume serverless functions are always the lowest-latency option, ignoring concurrency and cold-start limitations in streaming contexts.

How to eliminate wrong answers

Option B is wrong because Amazon Kinesis Data Firehose is a near-real-time delivery service with a minimum buffer interval of 60 seconds, making it unsuitable for real-time enrichment with minimal latency. Option C is wrong because AWS Lambda triggered by Kinesis Data Streams has a maximum concurrency limit per shard (e.g., 10 concurrent invocations per shard) and incurs cold-start latency, which can cause backpressure and increased processing delays for high-throughput streaming workloads. Option D is wrong because AWS Glue streaming ETL is based on Apache Spark Structured Streaming, which introduces higher startup overhead and micro-batch latency (typically seconds), making it less optimal for sub-second real-time enrichment compared to Flink's event-at-a-time processing.

246
MCQhard

A data engineer is using AWS Glue Studio to build a job that reads from an Amazon S3 bucket containing nested JSON files. The job must flatten the nested structure and write the output to Amazon Redshift. The engineer notices that the job fails with an error indicating that the Redshift table does not exist. The engineer has already created the Redshift cluster and database. What should the engineer do to resolve the error?

A.Create the target table in Amazon Redshift with the appropriate schema before running the job.
B.Add a 'ResolveChoice' transform before the Redshift writer to handle nested types.
C.Use the 'ApplyMapping' transform to map columns to Redshift data types.
D.Configure the Redshift writer to use the 'CREATE TABLE' option and specify the table name.
AnswerA

AWS Glue jobs writing to Amazon Redshift require the target table to exist beforehand. The error indicates the table is missing. The engineer must create the table in Redshift with a schema that matches the output of the Glue job, including column names and data types. This is a common prerequisite when using the Redshift connector in Glue Studio.

Why this answer

When writing to Amazon Redshift from AWS Glue, the target table must already exist. Glue does not automatically create Redshift tables. The engineer should create the table with the correct schema in Redshift, then rerun the job.

The other transforms mentioned are for data manipulation within the job and do not address the missing table.

Exam trap

The trap here is thinking that AWS Glue can create the Redshift table automatically, similar to how it might create a Data Catalog table, when in fact Redshift tables must be pre-created.

247
MCQeasy

A company runs an Amazon RDS for PostgreSQL database and wants to capture change data (inserts, updates, deletes) to stream into Amazon Kinesis Data Streams for real-time processing. Which AWS service should be used to capture the changes directly from the database?

A.Amazon RDS automated snapshots
B.AWS Glue ETL job scheduled to run every minute
C.Amazon Kinesis Agent
D.AWS Database Migration Service (DMS) with ongoing replication
AnswerD

DMS ongoing replication reads the PostgreSQL write-ahead log via logical replication slots, capturing inserts, updates and deletes as they occur and streaming them to Kinesis. This satisfies the requirement to capture change data directly from the database without application changes.

Why this answer

AWS DMS with ongoing replication (change data capture) is the correct service because it can continuously capture insert, update, and delete operations from the PostgreSQL transaction logs (WAL) and stream them to a Kinesis Data Streams endpoint. This allows real-time processing without modifying the source database or requiring application-level triggers.

Exam trap

The trap here is that candidates confuse scheduled polling (Glue) or file-based agents (Kinesis Agent) with true CDC, failing to recognize that only DMS ongoing replication can stream row-level changes directly from the database transaction log in real time.

How to eliminate wrong answers

Option A is wrong because Amazon RDS automated snapshots are point-in-time backups of the entire database, not a mechanism to capture individual row-level changes in real time. Option B is wrong because an AWS Glue ETL job scheduled every minute introduces at least 60 seconds of latency and cannot capture every single change as it happens, making it unsuitable for true real-time streaming. Option C is wrong because Amazon Kinesis Agent is designed to stream log files (e.g., from EC2 instances) to Kinesis, not to connect directly to a database and read transactional changes from its WAL.

248
Multi-Selecteasy

A company is designing a data lake on Amazon S3. The data includes CSV files, Parquet files, and images. The data engineering team needs to catalog the metadata and enable SQL queries. Which TWO AWS services should be used together?

Select 2 answers
A.Amazon EMR
B.Amazon Redshift Spectrum
C.Amazon QuickSight
D.Amazon Athena
E.AWS Glue
AnswersD, E

Amazon Athena queries S3 data directly using standard SQL, reading the table and partition metadata held in the AWS Glue Data Catalog. It satisfies the SQL query requirement without loading data, complementing Glue's cataloguing role for the CSV, Parquet and image lake.

Why this answer

Amazon Athena is correct because it is a serverless interactive query service that can directly query data stored in Amazon S3 using standard SQL, without needing to load or transform data. AWS Glue is correct because it provides a fully managed data catalog (AWS Glue Data Catalog) that stores metadata about the data lake's schema, partitions, and locations, which Athena can use to discover and query the data efficiently.

Exam trap

The trap here is that candidates often confuse Amazon Redshift Spectrum (which requires a Redshift cluster) with Athena (which is serverless), or they think Amazon EMR is needed for SQL queries on S3, not realizing Athena provides a simpler, cluster-free solution.

249
MCQeasy

A data engineer runs a Spark job on Amazon EMR that reads data from Amazon S3 and writes results back to S3. The job fails with an 'S3AccessDenied' error. The engineer verifies that the IAM role attached to the EMR cluster has s3:GetObject and s3:PutObject permissions on the relevant buckets. What is the MOST likely cause of the error?

A.S3 Transfer Acceleration is not enabled on the bucket.
B.EMRFS consistent view is not configured.
C.The S3 bucket is in a different AWS Region than the EMR cluster.
D.The IAM role does not have s3:ListBucket permission on the bucket.
AnswerD

Spark's S3A filesystem lists the bucket or prefix before reading and writing objects, and that listing call requires s3:ListBucket on the bucket resource. GetObject and PutObject alone are insufficient, so the missing ListBucket permission causes the S3AccessDenied failure.

Why this answer

The IAM role attached to the EMR cluster must have the s3:ListBucket permission on the bucket to allow the Spark job to enumerate objects when reading from S3. Without this permission, even with s3:GetObject and s3:PutObject, the job fails with an 'S3AccessDenied' error because the S3 list operation is required for directory listing and file discovery.

Exam trap

The trap here is that candidates often assume GetObject and PutObject are sufficient for S3 read/write operations, overlooking that the ListBucket permission is required for directory listing and file discovery in Spark jobs.

How to eliminate wrong answers

Option A is wrong because S3 Transfer Acceleration is a feature for faster uploads over long distances and is not required for basic read/write operations; its absence does not cause an access denied error. Option B is wrong because EMRFS consistent view is a consistency mechanism for eventually consistent S3 buckets, not a permission or access control feature; its absence would not produce an S3AccessDenied error. Option C is wrong because while cross-region access can cause latency or additional costs, it does not inherently cause an access denied error as long as the IAM role has the correct permissions and the bucket policy allows cross-region access.

250
MCQmedium

A data engineer needs to ensure that an S3 bucket can only be accessed from a specific VPC. Which policy element should be used?

A.Use the condition key aws:VpcSourceIp in the bucket policy.
B.Use the condition key aws:SourceIp in the bucket policy.
C.Use the condition key aws:SourceVpce in the bucket policy.
D.Use the condition key aws:SourceVpc in the bucket policy.
AnswerD

The aws:SourceVpc condition key evaluates the VPC endpoint through which the request arrives, so the bucket policy denies any traffic not originating from the specified VPC. This directly satisfies the requirement that the bucket be accessible only from that VPC.

Why this answer

The condition key aws:SourceVpc restricts requests to originate from a specific VPC. Option C (aws:SourceVpce) limits access to a VPC endpoint, not the VPC itself. Option B (aws:SourceIp) restricts by IP address, not VPC.

Option A (aws:VpcSourceIp) is not a valid condition key.

251
MCQeasy

A data engineer receives an alert that a Kinesis Data Stream has a 'WriteProvisionedThroughputExceeded' error. The stream has 5 shards with 1 MB/s write capacity per shard. The producer application is sending data at 8 MB/s sustained. What should the engineer do to resolve the issue?

A.Reduce the record size to below 1 MB per record.
B.Enable enhanced fan-out on the stream.
C.Increase the number of shards from 5 to 10.
D.Use Kinesis Firehose as an intermediary to buffer data.
AnswerC

The producer writes 8 MB/s while five shards provide only 5 MB/s aggregate, causing WriteProvisionedThroughputExceeded. Doubling to ten shards raises capacity to 10 MB/s, comfortably absorbing the sustained 8 MB/s load and clearing the error.

Why this answer

The 'WriteProvisionedThroughputExceeded' error indicates that the total write throughput to the Kinesis Data Stream exceeds the provisioned capacity. With 5 shards, each offering 1 MB/s write capacity, the total write capacity is 5 MB/s. The producer is sending 8 MB/s, which is above this limit.

Increasing the number of shards to 10 raises the total write capacity to 10 MB/s, accommodating the sustained 8 MB/s throughput and resolving the throttling.

Exam trap

The trap here is that candidates confuse write-side throttling with read-side limitations, leading them to choose enhanced fan-out (a read-side optimization) instead of scaling shards to increase write capacity.

How to eliminate wrong answers

Option A is wrong because reducing record size below 1 MB does not address the throughput limit; the error is about aggregate write throughput exceeding shard capacity, not individual record size limits. Option B is wrong because enhanced fan-out is a feature for increasing read throughput (up to 2 MB/s per shard per consumer) and does not affect write capacity or resolve write-side throttling. Option D is wrong because Kinesis Firehose is a delivery service that reads from a Kinesis stream; it cannot buffer data before it is written to the stream, so it does not solve the write throughput exceedance at the producer side.

252
MCQeasy

A data engineer needs to grant an IAM role read-only access to Amazon DynamoDB tables in a specific AWS account. Which IAM policy element should be used to restrict access to only the 'GetItem' and 'Query' actions?

A.Resource
B.Action
C.Effect
D.Condition
AnswerB

The Action element lists the specific API operations a policy allows or denies, so specifying dynamodb:GetItem and dynamodb:Query grants exactly those read-only calls and nothing else. This directly satisfies the stem's constraint of restricting access to only those two DynamoDB actions within the account.

Why this answer

The Action element in an IAM policy specifies the AWS API operations that the policy allows or denies. To restrict access to only 'GetItem' and 'Query' on DynamoDB tables, you list these actions in the Action element, e.g., 'dynamodb:GetItem' and 'dynamodb:Query'. This directly controls which operations the role can perform, making it the correct choice for limiting permissions to specific actions.

Exam trap

DEA-C01 often tests confusion between IAM policy elements, particularly Action versus Resource, where candidates might mistakenly think Resource specifies the allowed operations instead of the target AWS resources.

How to eliminate wrong answers

Option A is wrong because the Resource element defines the AWS resources (e.g., specific DynamoDB table ARNs) to which the policy applies, not the allowed actions. Option C is wrong because the Effect element specifies whether the policy allows or denies access (Allow/Deny), not which actions are permitted. Option D is wrong because the Condition element adds constraints (e.g., IP range, time) under which the policy is in effect, but does not list the actions themselves.

253
MCQmedium

An e-commerce company uses Amazon DynamoDB as the primary data store for its product catalog. The table has a simple primary key (ProductID) and handles 10,000 writes per second during peak hours. Recently, the engineering team noticed increased write latency and throttled requests during peak times. The table's provisioned write capacity is set to 12,000 WCU. What is the most likely cause of the throttling?

A.The table has reached the maximum number of partitions
B.DynamoDB Accelerator (DAX) is not configured
C.Write traffic is unevenly distributed across partitions
D.A global secondary index is consuming write capacity
AnswerC

Unevenly distributed write traffic concentrates load on a subset of partitions, so individual partitions exceed their 1,000 WCU per-partition limit even though total provisioned capacity of 12,000 WCU is not exhausted. DynamoDB throttles at partition level, making hot partitions the likely cause of the latency and throttling during peak hours.

Why this answer

DynamoDB partitions data by the primary key's hash value. If write traffic is unevenly distributed across partitions (e.g., a few ProductIDs receive most writes), those hot partitions can exceed their individual throughput limits (3,000 WCU per partition for provisioned tables), causing throttling even when the table's total provisioned WCU of 12,000 is not fully utilized.

Exam trap

The trap here is that candidates assume throttling only occurs when total provisioned capacity is exceeded, overlooking the per-partition throughput limits that cause throttling on hot partitions even when the table's overall WCU is underutilized.

How to eliminate wrong answers

Option A is wrong because DynamoDB tables do not have a maximum number of partitions; partitions are automatically added or removed based on storage and throughput needs. Option B is wrong because DAX is an in-memory cache for reads, not writes; it does not affect write capacity or throttling. Option D is wrong because while a global secondary index (GSI) does consume write capacity from the table's WCU pool, the question states the table has 12,000 WCU provisioned, and throttling occurs during peak writes of 10,000 writes per second, so the GSI would only contribute to throttling if its own provisioned WCU were insufficient, but the scenario does not indicate that.

254
MCQhard

A data engineer manages an Amazon Redshift cluster that contains a table with credit card numbers. The security team requires that the credit card column be stored in encrypted form and that only users with a specific IAM role can see the full values. Other users should see a partially masked value when they query the table. Which Redshift feature should the data engineer use?

A.Redshift row-level security with a policy that filters rows containing credit card numbers.
B.Redshift column-level encryption with a customer managed key in AWS KMS.
C.Redshift Spectrum with an external table that uses a SerDe to mask the credit card column.
D.Redshift dynamic data masking with a masking policy attached to the credit card column.
AnswerD

Redshift dynamic data masking applies a masking policy to a column so that unauthorized users see a masked value at query time, while authorized roles see the full value. It is role-based and does not require changing the stored data. This directly satisfies the requirement to show partial values to most users and full values only to a specific role.

Why this answer

Redshift dynamic data masking attaches a masking policy to a column and evaluates the querying user's role at query time. Authorized roles see the original value, while others see a masked form such as the last four digits. Column-level encryption, Spectrum, and row-level security address different concerns and cannot produce role-based partial masking of a column value.

Exam trap

The trap here is assuming that column-level encryption also masks values for unauthorized users, when it only protects data at rest.

255
MCQeasy

A data engineer needs to ingest data from an external partner's FTP server to Amazon S3. The data arrives once daily as a CSV file. Which AWS service should be used for this ingestion?

A.AWS DataSync
B.Amazon Kinesis Data Firehose
C.Amazon AppFlow
D.AWS Transfer Family
AnswerD

AWS Transfer Family provides a fully managed FTP endpoint that writes directly into Amazon S3, satisfying the daily CSV ingestion requirement without custom code or servers. Unlike AWS DataSync, which synchronises existing data stores, Transfer Family natively speaks FTP, letting the external partner push files straight to S3.

Why this answer

AWS Transfer Family provides fully managed support for file transfers over SFTP, FTPS, and FTP protocols, making it the correct choice for ingesting CSV files from an external partner's FTP server. It integrates directly with Amazon S3 as a destination, enabling automated, secure, and scheduled transfers without custom infrastructure.

Exam trap

The trap here is that candidates often confuse AWS DataSync (which is for NFS/SMB, not FTP) with a general-purpose file transfer service, or they incorrectly assume Kinesis Data Firehose can handle batch file ingestion from external sources.

How to eliminate wrong answers

Option A is wrong because AWS DataSync is designed for high-speed, large-scale data transfers between on-premises storage and AWS, but it does not support the FTP protocol; it uses its own agent-based architecture over NFS/SMB. Option B is wrong because Amazon Kinesis Data Firehose is a streaming ingestion service for real-time data (e.g., logs, events) and cannot connect to an FTP server or handle scheduled batch file transfers. Option C is wrong because Amazon AppFlow supports SaaS application integrations (e.g., Salesforce, Slack) and does not support FTP as a source or destination.

256
MCQhard

A data engineer is designing a solution to move data from an on-premises Oracle database to Amazon S3 using AWS DMS. The engineer needs to ensure that data changes are replicated continuously with minimal latency. Which DMS configuration is most appropriate?

A.Use AWS SCT to convert the schema and then use DMS for full load
B.Use a full-load task with ongoing replication (CDC)
C.Use a full-load task that runs daily
D.Use Kinesis Data Streams to capture changes and write to S3
AnswerB

Full load plus ongoing replication (CDC) migrates existing rows and then continuously applies change data capture records from the Oracle redo logs, keeping S3 synchronised with minimal latency. A full-load-only task would leave subsequent changes unreplicated.

Why this answer

DMS with continuous replication (CDC) captures ongoing changes with low latency. Option A is wrong because AWS SCT is used for schema conversion, not data movement; DMS with full load only does an initial copy. Option C is wrong because a daily full-load task does not provide continuous replication or low latency.

Option D is wrong because it describes Kinesis Data Streams, which is not a DMS configuration—DMS itself supports CDC to S3.

257
MCQeasy

A company is using Amazon Kinesis Data Firehose to ingest data into Amazon S3. The data must be transformed from JSON to Parquet format before delivery. Which feature should be enabled on the Firehose delivery stream?

A.Amazon Kinesis Data Analytics
B.Amazon S3 event notifications
C.Format conversion (Parquet/ORC)
D.AWS Lambda transformation
AnswerC

Format conversion performs schema-aware JSON-to-Parquet transcoding inside the Firehose delivery stream, using a referenced AWS Glue Data Catalog table to map source fields to Parquet columns. This satisfies the stem's requirement that records be transformed before delivery, removing the need for downstream ETL or Lambda-based conversion.

Why this answer

Amazon Kinesis Data Firehose has a built-in format conversion feature that can automatically convert input data from JSON to Parquet or ORC format before delivery to Amazon S3. Option A (Amazon Kinesis Data Analytics) is for real-time stream processing, not format conversion within Firehose. Option B (Amazon S3 event notifications) triggers notifications on S3 events, not data transformation.

Option D (AWS Lambda transformation) allows custom code for data transformation but is not specifically for converting JSON to Parquet; the built-in format conversion is the appropriate feature for this task.

258
Multi-Selectmedium

A data engineer is configuring a data lake on Amazon S3 that contains sensitive customer information. The company requires that all access to this data be logged and monitored, and that any data shared with external partners must be anonymized before leaving the S3 bucket. Which combination of AWS services should the engineer use to meet these requirements? (Choose THREE.)

Select 3 answers
A.AWS WAF
B.AWS Lake Formation
C.AWS CloudTrail
D.AWS Direct Connect
E.Amazon Macie
AnswersB, C, E

Lake Formation provides fine-grained access control and can be used to enforce anonymization policies.

Why this answer

AWS Lake Formation (B) is correct because it provides fine-grained access control and data anonymization capabilities for data lakes on Amazon S3. It allows you to define column-level and row-level security policies, and can automatically anonymize sensitive data (e.g., via masking or tokenization) before it is shared with external partners, ensuring compliance with data governance requirements.

Exam trap

The trap here is that candidates often confuse AWS WAF (a web-layer security tool) with data-level security, or assume Direct Connect provides logging and monitoring, when in fact neither service addresses S3 data access logging or anonymization.

259
MCQhard

A media company ingests millions of small JSON files per day into an Amazon S3 bucket. Analysts run Amazon Athena queries over this data and report that each query scans far more data than the files matching their filters, resulting in high cost and slow performance. The files are partitioned by year/month/day in S3. What should a data engineer do to reduce the data scanned per query?

A.Move the files into an Amazon Redshift cluster using COPY and query them with Redshift Spectrum.
B.Create an Amazon CloudFront distribution in front of the S3 bucket and run Athena against the distribution.
C.Enable S3 Transfer Acceleration on the bucket and increase the Athena query timeout.
D.Convert the JSON files to Apache Parquet, store them in the existing partition structure, and define the table in the AWS Glue Data Catalog with the correct partition columns.
AnswerD

Athena charges by data scanned, and columnar Parquet with predicate pushdown lets the engine read only the columns and row groups needed. Combined with partition pruning on year/month/day, queries skip irrelevant prefixes entirely. This directly reduces bytes scanned, lowering cost and improving latency without changing the query interface.

Why this answer

Athena's cost model is based on bytes scanned, so the most effective optimization is to store data in a columnar format such as Parquet and rely on partition pruning. Parquet enables column projection and predicate pushdown, while year/month/day partitions let Athena skip entire prefixes. Together they sharply reduce scanned bytes, cutting cost and improving query speed for the analyst workload.

Exam trap

The trap here is focusing on ingestion or delivery speedups such as Transfer Acceleration or CloudFront, which do not change how many bytes Athena reads during a query.

260
Multi-Selectmedium

A data engineer needs to design a data ingestion pipeline that captures streaming data from mobile app events into Amazon S3 for analytics. The pipeline must support real-time processing of events and allow for schema evolution over time. Which AWS services should the engineer use? (Choose THREE.)

Select 3 answers
A.Amazon Kinesis Data Analytics
B.Amazon Kinesis Data Firehose
C.AWS Glue ETL jobs
D.Amazon Kinesis Data Streams
E.AWS AppFlow
AnswersA, B, D

Enables real-time processing and schema evolution.

Why this answer

Amazon Kinesis Data Analytics is correct because it enables real-time processing of streaming data using SQL or Apache Flink, allowing the engineer to analyze mobile app events as they arrive. This supports the requirement for real-time processing before the data is stored in Amazon S3 for analytics.

Exam trap

The trap here is that candidates often confuse AWS Glue ETL jobs as a streaming solution, but Glue is fundamentally batch-oriented and cannot meet real-time processing requirements, while AppFlow is mistakenly chosen for its integration capabilities despite lacking streaming ingestion support.

261
MCQhard

A company uses Amazon Kinesis Data Analytics for Apache Flink to process streaming data. The application reads from a Kinesis data stream and writes results to an S3 bucket. The application is consistently running out of memory and failing. The operator has already increased the Parallelism and TaskManager memory. What is the next BEST step to troubleshoot?

A.Change the processing mode from exactly-once to at-least-once
B.Reduce the number of shards in the source stream
C.Enable Apache Flink metrics in Amazon CloudWatch to monitor heap and checkpoint details
D.Increase the buffer timeout for the S3 sink
AnswerC

Enabling Flink metrics in CloudWatch exposes heap usage, garbage-collection pressure and checkpoint behaviour, revealing whether state growth or backpressure—not raw parallelism—causes the out-of-memory failures. Since TaskManager memory and parallelism were already raised without success, these per-operator metrics identify the actual bottleneck before further resource changes.

Why this answer

When a Flink application on Kinesis Data Analytics runs out of memory despite increased parallelism and TaskManager memory, the next best step is to enable Apache Flink metrics in CloudWatch to observe heap usage, garbage collection, and checkpoint behavior. These metrics reveal whether the issue is memory leaks, backpressure, or state growth. Without observability, further tuning is guesswork.

Exam trap

DEA-C01 often tests the impulse to keep scaling resources (parallelism, memory) without first enabling metrics, when the correct troubleshooting step is to gain visibility into heap and checkpoint behavior.

How to eliminate wrong answers

Option A is wrong because changing from exactly-once to at-least-once affects checkpointing semantics and may reduce overhead slightly, but it does not diagnose or resolve the root cause of memory exhaustion. Option B is wrong because reducing source shards lowers parallelism and throughput, potentially worsening the problem and not addressing memory pressure. Option D is wrong because increasing the S3 sink buffer timeout only affects write batching, not the application's memory footprint or the underlying cause of OOM failures.

262
MCQhard

A company uses Amazon Redshift for its data warehouse. The data engineer notices that queries are slow on a large table that is frequently filtered on a column 'transaction_date'. Which optimization technique best improves query performance?

A.Apply compression encoding to 'transaction_date'.
B.Set the sort key to 'transaction_date'.
C.Set the distribution key to 'transaction_date'.
D.Run VACUUM on the table.
AnswerB

Sort keys physically order rows on disk by transaction_date, so Redshift's zone maps skip irrelevant blocks when filtering that column, cutting scanned data. This targets the stem's frequently filtered column directly, unlike distribution keys, which address join and data-skew costs instead.

Why this answer

Setting the sort key to 'transaction_date' organizes the table data physically by that column, which allows Redshift to use zone maps to skip blocks that don't match query filters. This dramatically reduces the amount of data scanned for range-restricted queries on 'transaction_date', improving query performance.

Exam trap

The trap here is that candidates confuse distribution keys (which optimize joins) with sort keys (which optimize filtering and range scans), leading them to pick distribution key as the answer for a single-table filter performance issue.

How to eliminate wrong answers

Option A is wrong because compression encoding reduces storage size and I/O but does not directly optimize query filtering on a column; it can even slow down scans if the column is frequently used in predicates. Option C is wrong because setting the distribution key to 'transaction_date' distributes rows across nodes based on that column, which can help with joins but does not improve the efficiency of range-restricted scans on a single table. Option D is wrong because VACUUM reclaims space and re-sorts data but does not improve query performance unless the table is already sorted on a key; without a sort key on 'transaction_date', VACUUM has no effect on filter performance.

263
MCQhard

A data engineer uses AWS Glue DataBrew to clean a dataset stored in Amazon S3. The recipe includes a 'Filter' step that removes rows where the 'status' column equals 'INVALID'. After running the recipe, the engineer notices that the output still contains rows with status 'INVALID'. The recipe was published and the job ran successfully. What is the most likely cause?

A.The filter step was configured as a 'Remove' transformation instead of a 'Filter' transformation, so it only flagged rows without deleting them.
B.The recipe was applied to a sample of the data during job execution, so only a subset was filtered.
C.The DataBrew job was run in profile mode instead of recipe mode, so transformations were not applied.
D.The filter condition used a case-sensitive match and the actual values are 'invalid' in lowercase, so the condition did not match any rows.
AnswerD

DataBrew filter conditions are case-sensitive by default. If the source data contains 'invalid' rather than 'INVALID', a condition checking for 'INVALID' will not match, so those rows are retained. The engineer should either normalize case or adjust the condition. This is a common cause of filters appearing to have no effect when the job succeeds but output is unchanged.

Why this answer

DataBrew filter conditions are case-sensitive, so a condition matching 'INVALID' will not remove rows containing 'invalid'. The engineer should either standardize the case of the status column or adjust the filter condition to match the actual values. This explains why the job succeeded but the unwanted rows remained.

Exam trap

The trap here is assuming that DataBrew filter conditions are case-insensitive or that a successful job run guarantees the filter matched the intended rows.

264
MCQmedium

A company stores log files in Amazon S3. They want to automatically move logs older than 90 days to S3 Glacier Deep Archive to reduce costs. Which S3 feature should be used?

A.S3 Intelligent-Tiering
B.S3 Lifecycle configuration
C.S3 Replication
D.S3 Object Lock
AnswerB

S3 Lifecycle configuration applies transition rules that automatically move objects to Glacier Deep Archive once they cross the specified age threshold, here 90 days. It satisfies the stem's requirement for automatic, age-based archival without custom code, and transitions are billed per request rather than requiring retrieval or manual intervention.

Why this answer

S3 Lifecycle configuration allows you to define rules that automatically transition objects to colder storage classes, such as S3 Glacier Deep Archive, based on age. By setting a rule to move objects older than 90 days to S3 Glacier Deep Archive, you reduce storage costs without manual intervention. This is the correct feature for automating tier-based data lifecycle management.

Exam trap

The trap here is that candidates may confuse S3 Intelligent-Tiering with lifecycle policies, but Intelligent-Tiering does not support age-based transitions to Glacier Deep Archive and is designed for unpredictable access patterns, not fixed retention schedules.

How to eliminate wrong answers

Option A is wrong because S3 Intelligent-Tiering automatically moves data between access tiers based on changing access patterns, not on a fixed age-based schedule, and it does not support direct transition to S3 Glacier Deep Archive. Option C is wrong because S3 Replication is used to copy objects across buckets or regions for redundancy or compliance, not to transition objects to colder storage classes. Option D is wrong because S3 Object Lock is designed to prevent object deletion or overwrites for a specified retention period, not to manage storage tier transitions.

265
MCQeasy

A data engineer needs to grant an AWS Glue ETL job access to read data from an Amazon S3 bucket that is encrypted with SSE-KMS using a customer managed key. The Glue job runs with an IAM role. Which action must the engineer take to allow the Glue job to decrypt the data?

A.Enable default encryption on the S3 bucket using SSE-S3 instead of SSE-KMS, which eliminates the need for KMS permissions.
B.Attach a policy to the Glue job's IAM role that allows kms:Decrypt and kms:GenerateDataKey on the specific KMS key.
C.Use an S3 Access Point with a policy that allows the Glue job's IAM role to access objects, and configure the access point to use a different KMS key.
D.Modify the S3 bucket policy to allow the Glue job's IAM role to perform s3:GetObject, and rely on S3 to decrypt the data automatically.
AnswerB

For a Glue job to read SSE-KMS encrypted data, its IAM role must have permissions to use the KMS key for decryption. The required permissions are kms:Decrypt and kms:GenerateDataKey (for some operations). Attaching a policy with these actions on the specific key resource grants the necessary access. Without these permissions, the Glue job will fail with access denied errors.

Why this answer

To allow an AWS Glue job to read SSE-KMS encrypted data from S3, the IAM role associated with the Glue job must have permissions to use the KMS key for decryption. Specifically, the role needs kms:Decrypt and kms:GenerateDataKey permissions on the key. S3 bucket policies alone do not grant KMS access, and changing encryption to SSE-S3 or using access points does not satisfy the requirement to use SSE-KMS.

Exam trap

The trap here is assuming that granting s3:GetObject in a bucket policy is sufficient for reading SSE-KMS encrypted objects, when in fact the requester also needs explicit KMS permissions.

266
Multi-Selecteasy

A data engineer is designing a serverless data ingestion pipeline that uses Amazon Kinesis Data Firehose to deliver data to Amazon S3. The data must be transformed using AWS Lambda before being written to S3. Which two steps are required to enable this transformation? (Select TWO.)

Select 2 answers
A.Set up an S3 event notification to trigger the Lambda function on object creation.
B.Configure a Lambda function as a data transformation source in the Firehose delivery stream.
C.Ensure the Lambda function returns the transformed data in the format required by Firehose.
D.Subscribe the Lambda function to the CloudWatch Logs log group for the Firehose stream.
E.Have the Lambda function write the transformed data directly to the S3 bucket.
AnswersB, C

Firehose invokes a Lambda function only when it is attached as the delivery stream's data transformation source, which enables the buffered records to be processed before delivery. Without this configuration, Firehose writes raw records straight to S3, so transformation never occurs.

Why this answer

Option B is correct because Firehose supports Lambda-based data transformation by letting you specify a Lambda function as the processor in the delivery stream's transform configuration (via the console or the ProcessingConfiguration/TransformParameters in the API), which Firehose then invokes synchronously for each buffered batch. Option C is correct because the Lambda function must return the transformed records in the exact structure Firehose expects — a JSON object containing records with recordId, result (Ok, Dropped, or ProcessingFailed), and base64-encoded data — otherwise Firehose cannot continue delivery. Option A is wrong because S3 event notifications trigger actions on object creation and play no role in Firehose's inline transformation; Firehose itself invokes the Lambda function.

Option D is wrong because subscribing Lambda to the Firehose CloudWatch Logs log group is not a configuration step for transformation and would not enable it. Option E is wrong because the Lambda function must return transformed data to Firehose, not write directly to S3; Firehose remains responsible for delivering the records to the destination bucket.

Exam trap

The trap here is that candidates often confuse post-delivery transformations (using S3 event notifications) with in-stream transformations (using Firehose's built-in Lambda integration), leading them to select Option A instead of the correct Firehose-specific configuration.

267
MCQeasy

A data engineer notices that an Amazon RDS for PostgreSQL instance's CPU utilization is consistently above 90% during business hours. The database is used for reporting queries. Which action should be taken FIRST to improve performance?

A.Enable Multi-AZ deployment for automatic failover.
B.Enable Performance Insights and review slow queries.
C.Create a read replica to offload reporting queries.
D.Increase the instance size to a larger instance class.
AnswerB

Performance Insights identifies the specific SQL statements and wait events driving high CPU, so slow reporting queries can be targeted rather than guessed at. This diagnostic step precedes changes such as indexing or instance scaling, satisfying the requirement to act first on evidence.

Why this answer

The first step in diagnosing high CPU utilization on an RDS for PostgreSQL instance used for reporting queries is to identify the root cause. Enabling Performance Insights provides a detailed view of database load, wait events, and SQL query performance, allowing the data engineer to pinpoint slow or inefficient queries that are consuming CPU resources. Without this diagnostic data, any other action would be premature and could lead to unnecessary cost or complexity.

Exam trap

The trap here is that candidates often jump to scaling solutions (like increasing instance size or adding a read replica) without first diagnosing the root cause, but AWS emphasizes observability and optimization before capacity changes.

How to eliminate wrong answers

Option A is wrong because enabling Multi-AZ deployment improves availability and failover, not performance; it does not reduce CPU utilization or address query performance issues. Option C is wrong because creating a read replica offloads read traffic but does not fix the underlying inefficient queries that are causing high CPU on the source instance; the replica would also suffer from the same workload if queries are poorly optimized. Option D is wrong because increasing the instance size may temporarily mask the problem by providing more CPU capacity, but it does not resolve the root cause of inefficient queries and incurs higher costs without guaranteeing sustained performance improvement.

268
MCQeasy

A company wants to grant read-only access to an S3 bucket for a data analyst. The analyst should be able to list objects and read object content. Which IAM policy effect and action combination is correct?

A.Effect: Allow, Actions: s3:GetObject, s3:DeleteObject
B.Effect: Allow, Actions: s3:ListAllMyBuckets, s3:GetObject
C.Effect: Allow, Actions: s3:PutObject, s3:GetObject
D.Effect: Allow, Actions: s3:ListBucket, s3:GetObject
AnswerD

ListBucket operates on the bucket resource and GetObject on the object resource, so both actions are required for listing and reading. An Allow effect grants the analyst read-only access without write permissions, matching the least-privilege requirement for listing objects and retrieving their content.

Why this answer

To grant read-only access to an S3 bucket, the policy must allow s3:ListBucket to list objects in the bucket and s3:GetObject to read object content. These actions are the minimum required for read-only access to a specific bucket.

Exam trap

DEA-C01 often tests the distinction between s3:ListBucket and s3:ListAllMyBuckets; candidates might confuse the two, leading to incorrect policies.

How to eliminate wrong answers

Option A is wrong because s3:DeleteObject grants delete permissions, which is not read-only. Option B is wrong because s3:ListAllMyBuckets lists all buckets in the account, not just the specific bucket, and does not grant permission to list objects within the bucket. Option C is wrong because s3:PutObject grants write permissions, which is not read-only.

269
MCQeasy

A data engineer needs to store semi-structured JSON data that is accessed infrequently but requires immediate retrieval when needed. The data must be durable and cost-effective. Which Amazon S3 storage class should be used?

A.S3 Standard-IA
B.S3 Glacier
C.S3 Standard
D.S3 One Zone-IA
AnswerA

S3 Standard-IA suits infrequent access with immediate retrieval, satisfying the low-frequency, on-demand constraint. Its lower storage cost than S3 Standard meets cost-effectiveness, while replicating across a minimum of three Availability Zones preserves durability. Unlike Glacier classes, no retrieval delay applies, so JSON remains instantly available.

Why this answer

S3 Standard-IA is the correct choice because it offers the same durability and low-latency retrieval as S3 Standard but at a lower storage cost, making it ideal for infrequently accessed data that still needs immediate retrieval when requested. The scenario specifies 'infrequently accessed' and 'immediate retrieval,' which aligns with Standard-IA's design for data accessed less than once a month but with millisecond first-byte latency.

Exam trap

The trap here is that candidates often confuse 'infrequently accessed' with 'archival' and choose S3 Glacier, overlooking the 'immediate retrieval' requirement that rules out Glacier's multi-minute or multi-hour retrieval times.

How to eliminate wrong answers

Option B (S3 Glacier) is wrong because it is designed for archival data with retrieval times ranging from minutes to hours, not immediate retrieval. Option C (S3 Standard) is wrong because it is optimized for frequently accessed data and would be less cost-effective for infrequently accessed data, incurring higher storage costs without benefit. Option D (S3 One Zone-IA) is wrong because it stores data in a single Availability Zone, which does not meet the durability requirement of the scenario (data must be durable, implying multi-AZ resilience).

270
MCQeasy

A company is using AWS Lake Formation to manage permissions on a data lake. They want to grant a data scientist the ability to query tables in the 'analytics' database using Amazon Athena, but prevent them from accessing the underlying S3 data directly. What is the best way to achieve this?

A.Grant the data scientist an IAM policy with s3:GetObject on the S3 bucket.
B.Grant SELECT permission on the 'analytics' database tables in Lake Formation.
C.Create an IAM policy that allows Athena queries only.
D.Add the data scientist to a Lake Formation data lake location with read access.
AnswerB

Lake Formation grants SELECT on the analytics tables, letting the data scientist query them through Athena while Lake Formation mediates access to the underlying S3 objects. Direct S3 access stays blocked because permissions are enforced at the catalog and table level, not through S3 bucket policies.

Why this answer

Lake Formation grants SELECT permission on named database tables, which allows querying via Athena without granting direct S3 access. Option A is incorrect because granting s3:GetObject on the entire bucket would allow the data scientist to bypass Lake Formation and access the data directly. Option C is incorrect because a policy that allows Athena queries only does not grant the necessary permissions to access the database tables.

Option D is incorrect because adding the user to a data lake location with read access is too broad and would also grant direct S3 access, which does not meet the requirement of preventing direct S3 access.

271
MCQmedium

A data engineer is using AWS Glue to run an ETL job that reads from an Amazon S3 bucket and writes to another S3 bucket. The job processes data in CSV format and the engineer wants to ensure the output is partitioned by year, month, and day based on a timestamp column in the data. The engineer needs to optimize the job for performance and cost. Which approach should the engineer take?

A.Write the output to a single CSV file and then use Amazon Athena to create a partitioned table.
B.Use the AWS Glue job's bookmark feature to automatically partition the output by timestamp.
C.Use the AWS Glue ResolveChoice transformation to split the timestamp column into separate year, month, and day columns, then write to S3 with partitioning.
D.Use the AWS Glue DynamicFrame partitionColumns parameter to specify year, month, and day as partition keys.
AnswerD

The partitionColumns parameter in AWS Glue's write_dynamic_frame method allows you to specify columns to partition the output by. When writing to S3, Glue will create a directory structure based on these columns, such as year=2023/month=01/day=01. This is the standard and efficient way to partition output data in Glue, improving query performance and reducing costs for downstream analytics.

Why this answer

The partitionColumns parameter in AWS Glue's write operation is the correct way to partition output data by specified columns. It creates a hierarchical directory structure in S3, enabling efficient querying with services like Athena and Redshift Spectrum. Other options either misuse transformations or misunderstand the purpose of bookmarks and single-file output.

Exam trap

The trap here is confusing job bookmarks with partitioning, or thinking that ResolveChoice can derive new columns, when actually partitionColumns is the direct method for output partitioning.

272
MCQeasy

A data engineer has an AWS Glue job that processes data from an Amazon S3 bucket and writes to an Amazon Redshift cluster. The job is scheduled to run daily. Recently, the job started failing with the error: 'java.sql.SQLException: [Amazon](500310) Invalid operation: Spectrum Scan Error: S3 Access Denied'. The engineer verifies that the IAM role associated with the Glue job has full access to the S3 bucket. What is the most likely cause of this error?

A.The Glue job's script is using an incorrect S3 path for the Redshift COPY command.
B.The Amazon Redshift cluster's IAM role lacks permission to access the S3 bucket.
C.The Amazon Redshift cluster is in a different AWS Region than the S3 bucket, causing access latency.
D.The S3 bucket policy explicitly denies access to the Glue job's IAM role.
AnswerB

When AWS Glue writes to Amazon Redshift using the COPY command, Redshift itself must access the S3 bucket. If the Redshift cluster's IAM role does not have the necessary S3 permissions, Redshift cannot read the data, resulting in a Spectrum Scan Error with S3 Access Denied. The Glue job's role permissions are separate from Redshift's role.

Why this answer

The error occurs because Amazon Redshift, not AWS Glue, is attempting to access S3 during the COPY operation. The Redshift cluster's IAM role must have permissions to read the S3 bucket. Even if the Glue job's role has access, Redshift uses its own role.

The engineer should attach a policy to the Redshift cluster's IAM role granting S3 read access.

Exam trap

The trap here is assuming that the Glue job's IAM role is the only one involved, but Redshift uses its own IAM role for COPY commands.

273
MCQmedium

A data engineer manages an Amazon DynamoDB table that stores IoT sensor readings. The table uses a partition key of deviceId and a sort key of timestamp, with a provisioned read capacity of 100 RCUs. During a sudden spike in traffic, the engineer observes throttling on read operations even though the consumed read capacity is well below the provisioned limit. What is the MOST likely cause of the throttling?

A.The table uses eventually consistent reads, which are throttled more aggressively than strongly consistent reads.
B.The provisioned read capacity is set too low for the table's data volume.
C.The table lacks a global secondary index (GSI) on the timestamp attribute.
D.The table's partition key is a low-cardinality attribute, causing a hot partition.
AnswerD

A low-cardinality partition key, such as deviceId with few unique values, concentrates read traffic on a small number of partitions. DynamoDB limits each partition to 3,000 RCUs, so even if total consumed capacity is under 100 RCUs, a single partition can exceed its local limit and throttle requests. This matches the symptom of throttling despite low aggregate usage.

Why this answer

DynamoDB throttling can occur even when total consumed capacity is below the provisioned level if reads are concentrated on a few partitions. Each partition has a maximum throughput of 3,000 RCUs. A low-cardinality partition key creates hot partitions, causing localized throttling.

The solution is to use a higher-cardinality partition key or add a write sharding strategy.

Exam trap

The trap here is assuming that throttling always means total provisioned capacity is insufficient, overlooking partition-level limits.

274
MCQmedium

A data engineer is using AWS Glue Studio to create a visual ETL job that reads from an Amazon S3 bucket containing JSON files, applies a filter transformation, and writes the output to Amazon Redshift. The job must run daily. The engineer notices that the job is taking a long time to complete and wants to improve performance. Which action should the engineer take to optimize the job?

A.Convert the source JSON files to Parquet format before running the ETL job.
B.Use a larger number of smaller files instead of fewer large files.
C.Increase the number of AWS Glue DPUs allocated to the job.
D.Enable job bookmarks to track processed files and avoid reprocessing.
AnswerA

Parquet is a columnar format that offers better compression and faster read performance compared to JSON. Converting the source data to Parquet reduces I/O and CPU overhead during the ETL job, significantly improving performance. This is a common best practice for AWS Glue jobs reading from S3, especially when the data is large. It directly addresses the slow read and transform steps.

Why this answer

Converting JSON to Parquet improves ETL performance because Parquet is columnar, compressed, and requires less I/O and CPU to parse. AWS Glue jobs benefit significantly from columnar formats when reading from S3. While increasing DPUs or enabling bookmarks can help in some cases, the most direct and effective optimization for slow JSON processing is to use a more efficient file format.

Exam trap

The trap here is assuming that adding more DPUs is always the first step to improve AWS Glue job performance, when data format optimization often yields greater benefits at lower cost.

275
MCQhard

A healthcare company uses AWS Lake Formation to manage access to a data lake in Amazon S3. The data lake contains a table with patient records, and the company needs to ensure that only users in the 'Cardiology' department can query columns containing sensitive information such as patient name and diagnosis. Other departments should be able to query non-sensitive columns like patient ID and visit date. The company wants to implement this with the least operational overhead. What should the data engineer do?

A.Create an AWS Glue ETL job that reads the table, filters out sensitive columns based on the user's department, and writes the results to separate S3 buckets for each department. Grant each department access to its respective bucket.
B.Use AWS Identity and Access Management (IAM) policies to deny access to the sensitive columns for all departments except Cardiology. Attach these policies to the IAM roles used by each department.
C.Use Lake Formation column-level security to grant SELECT on the sensitive columns only to the Cardiology department role, and grant SELECT on non-sensitive columns to all other department roles.
D.Configure an Amazon Athena workgroup for each department and use Athena's column-level access control to restrict sensitive columns. Grant each department's users access to their workgroup.
AnswerC

Lake Formation supports column-level permissions, allowing fine-grained access control. By granting SELECT on sensitive columns only to the Cardiology role and on non-sensitive columns to other roles, the engineer enforces the requirement with minimal overhead. Lake Formation manages the permissions centrally, and no additional ETL or view creation is needed. This is the most efficient and secure approach.

Why this answer

Lake Formation column-level security allows granular access control at the column level, enabling the company to grant SELECT on sensitive columns only to the Cardiology department while allowing other departments to query non-sensitive columns. This approach centralizes permission management and requires no data duplication or custom ETL, minimizing operational overhead. It is the intended solution for fine-grained access control in a data lake.

Exam trap

The trap here is assuming that IAM policies can enforce column-level access, but IAM operates at the resource level and cannot restrict individual columns within a table.

276
Multi-Selecthard

A company uses Amazon Kinesis Data Firehose to deliver data to an S3 bucket. The data contains personally identifiable information (PII) that must be redacted before storage. Which TWO actions can achieve this requirement? (Choose TWO.)

Select 2 answers
A.Use AWS Glue ETL to read from the S3 bucket and write redacted data to another S3 bucket.
B.Use Amazon Athena to query the data and redact PII on the fly.
C.Use Amazon Macie to discover and automatically redact PII before storage.
D.Use AWS Database Migration Service (AWS DMS) to replicate data and apply transformations.
E.Use an AWS Lambda function as a transformation in the Firehose delivery stream.
AnswersA, E

Correct. AWS Glue ETL can perform schema-aware transformations, including redacting PII fields, and write to a target S3 bucket.

Why this answer

Option A is correct because AWS Glue ETL jobs can read the raw data from the S3 bucket, apply transformation logic (such as masking or removing PII fields), and write the redacted output to another S3 bucket, satisfying the requirement that PII be redacted before storage in the destination. Option E is correct because Kinesis Data Firehose supports Lambda-based record transformation, where a Lambda function is invoked on each incoming record to redact PII before Firehose delivers the data to S3, which is the most direct in-stream solution. Option B is not correct because Amazon Athena is a query service that reads data already stored in S3; it cannot redact PII before storage.

Option C is not correct because Amazon Macie discovers and classifies sensitive data but does not automatically redact or modify it. Option D is not correct because AWS DMS is designed for database migration and replication, not for transforming and redacting PII in a Firehose-to-S3 data pipeline.

Exam trap

The trap is that candidates may assume Macie can automatically redact PII, but it only discovers and classifies; it cannot modify data without additional services like Lambda. Thus, Option C is not a correct answer.

277
MCQhard

Refer to the exhibit. An IAM policy is attached to an IAM role used by an application. The application needs to decrypt objects in an S3 bucket using a customer managed KMS key. What is the effect of this policy?

A.The application cannot perform any KMS operations.
B.The application can decrypt objects from any service.
C.The application can decrypt objects only when accessing them through S3.
D.The application can encrypt but not decrypt objects.
AnswerC

The policy grants kms:Decrypt but conditions it on the request arriving via the S3 service, so the principal can decrypt only through S3 GetObject calls. Direct KMS decrypt requests, or access via other services, are denied by the condition.

Why this answer

The IAM policy grants the `kms:Decrypt` permission with a `kms:ViaService` condition key set to `s3.amazonaws.com`. This condition restricts the decryption operation to only when the request is made through the S3 service. Therefore, the application can decrypt objects only when accessing them through S3, not via direct KMS API calls or other services.

Exam trap

AWS often tests the `kms:ViaService` condition key to trap candidates who assume that granting `kms:Decrypt` alone allows decryption from any source, ignoring the service-specific restriction.

How to eliminate wrong answers

Option A is wrong because the policy explicitly allows `kms:Decrypt` under the condition, so the application can perform KMS decryption operations when invoked via S3. Option B is wrong because the `kms:ViaService` condition restricts decryption to S3 only, preventing decryption from any other service or direct KMS API calls. Option D is wrong because the policy grants `kms:Decrypt` permission, not `kms:Encrypt`, so the application can decrypt but not encrypt objects.

278
MCQeasy

A data engineer is using AWS Glue to run an ETL job that reads data from Amazon DynamoDB and writes to Amazon Redshift. The job fails with a 'ThroughputExceededException' error. What is the most likely cause?

A.The Glue job has a timeout setting that is too low
B.The Redshift cluster's concurrency scaling is insufficient
C.The DynamoDB table's read capacity is insufficient for the Glue job's read rate
D.The S3 bucket where Glue writes temporary data does not have proper permissions
AnswerC

DynamoDB returns ThroughputExceededException when read requests exceed the table's provisioned or on-demand read capacity. Glue's parallel executors consume read capacity units rapidly, exhausting the table's limits. This directly satisfies the stem's constraint: the read rate from DynamoDB surpasses available capacity, causing throttling rather than a Redshift or Glue configuration fault.

Why this answer

A ThroughputExceededException from DynamoDB indicates that the read or write request rate exceeded the provisioned throughput (RCUs/WCUs) or the burst capacity of the table or index. When AWS Glue reads from DynamoDB, it consumes read capacity units; if the table's read capacity is too low for the parallel read rate of the Glue job, DynamoDB throttles the requests and the job fails.

Exam trap

DEA-C01 often tests whether candidates correctly attribute throttling errors to the source database (DynamoDB) rather than the target (Redshift) or the Glue job configuration; the exception name itself points to DynamoDB throughput.

How to eliminate wrong answers

Option A is wrong because a Glue job timeout would produce a timeout error, not a ThroughputExceededException from DynamoDB. Option B is wrong because Redshift concurrency scaling affects query performance on the write side, not DynamoDB read throttling; the error originates from DynamoDB. Option D is wrong because S3 permissions issues would manifest as AccessDenied or similar errors when writing temporary data, not as a DynamoDB throughput exception.

279
MCQhard

A data engineer is troubleshooting a Kinesis Data Firehose delivery stream that is experiencing high error rates when writing to an S3 bucket. The error logs indicate 'AccessDenied' errors. The S3 bucket policy allows access from the Firehose service, but the errors persist. What is the most likely cause?

A.The S3 bucket has a lifecycle policy that is deleting objects too quickly
B.The IAM role assumed by Firehose does not have the s3:PutObject permission
C.The S3 bucket has default encryption enabled
D.The S3 bucket uses an AWS KMS key for encryption and Firehose does not have kms:Decrypt permission
AnswerB

Firehose requires both the bucket policy and its assumed IAM role to grant access. The bucket policy alone is insufficient; if the role lacks s3:PutObject, every delivery attempt returns AccessDenied, which matches the persistent errors despite the policy appearing correct.

Why this answer

The most likely cause is that the IAM role assumed by Kinesis Data Firehose lacks the `s3:PutObject` permission. Even if the S3 bucket policy allows access from the Firehose service, the IAM role must explicitly grant the necessary S3 write permissions for Firehose to deliver data. Without this permission, Firehose receives 'AccessDenied' errors when attempting to write objects to the bucket.

Exam trap

The trap here is that candidates assume a bucket policy allowing Firehose access is sufficient, but the IAM role assumed by Firehose must also explicitly grant the write permissions, as AWS evaluates both identity-based and resource-based policies.

How to eliminate wrong answers

Option A is wrong because a lifecycle policy that deletes objects too quickly would not cause 'AccessDenied' errors; it would cause data to be deleted after delivery, not prevent writes. Option C is wrong because default encryption on the S3 bucket does not block write access; Firehose can write encrypted objects as long as it has the necessary permissions. Option D is wrong because the error is 'AccessDenied', not a KMS-related error; if the issue were KMS permissions, the error would typically be 'KMS.AccessDeniedException' or similar, and Firehose would need `kms:GenerateDataKey` (not `kms:Decrypt`) to encrypt objects with SSE-KMS.

280
MCQmedium

A company uses Amazon Redshift for data warehousing. The security team requires that all data stored in Redshift be encrypted at rest using a customer-managed KMS key. How should the data engineer configure this?

A.Enable encryption using a KMS key when creating the Redshift cluster
B.Configure S3 SSE-KMS on the underlying S3 storage
C.Use the AWS KMS console to encrypt the Redshift cluster after creation
D.Set a cluster parameter group with encryption enabled
AnswerA

Specifying a customer-managed KMS key at cluster creation makes Redshift encrypt all data at rest with that key, satisfying the requirement for customer-managed encryption. Redshift cannot retroactively swap the key type, so this must be set during provisioning.

Why this answer

Redshift encryption at rest with a customer-managed KMS key must be specified at cluster creation time via the 'Encrypted' and 'KmsKeyId' parameters (console: 'KMS' encryption option). Once a cluster is created unencrypted, it cannot be converted in place — you must create a new encrypted cluster and migrate data. This satisfies the security team's requirement of a CMK rather than the default AWS-managed key.

Exam trap

DEA-C01 often tests the misconception that Redshift encryption can be enabled after cluster creation or via parameter groups, when in reality it is an immutable creation-time setting requiring cluster recreation or snapshot restore.

How to eliminate wrong answers

Option B is wrong because SSE-KMS on S3 is irrelevant — Redshift stores data on its own managed compute/storage nodes, not in a customer-visible S3 bucket, so S3 encryption settings have no effect on Redshift cluster data. Option C is wrong because the KMS console cannot encrypt an existing Redshift cluster; KMS only manages keys, and Redshift clusters cannot be retroactively encrypted in place. Option D is wrong because cluster parameter groups control engine/runtime settings (e.g., enable_user_activity_logging, max_cursor_result_set_size), not encryption, which is a cluster-level immutable property set at creation.

281
Multi-Selectmedium

A data engineer is troubleshooting an AWS Glue ETL job that fails with the error: 'An error occurred while calling o123.pyWriteDynamicFrame. Access Denied when writing to S3 bucket: my-bucket'. The job uses a Glue service role named 'GlueServiceRole'. Which TWO actions should the engineer take to resolve the issue? (Choose TWO.)

Select 2 answers
A.Disable S3 Block Public Access on the bucket.
B.Grant the GlueServiceRole permission to write to the AWS Glue Data Catalog.
C.Check if the S3 bucket policy denies access from the GlueServiceRole.
D.Verify that the IAM policy attached to GlueServiceRole includes s3:PutObject on the bucket.
E.Ensure the Glue job is in the same VPC as the S3 bucket.
AnswersC, D

An explicit Deny in the bucket policy overrides any Allow in the identity policy, producing Access Denied on s3:PutObject. Inspecting the bucket policy for a deny targeting GlueServiceRole identifies that blocking statement, satisfying the need to find the actual cause.

Why this answer

The error 'Access Denied when writing to S3 bucket' during pyWriteDynamicFrame indicates an S3 authorization failure for the Glue job's execution role, so option D is correct: the engineer must verify that the IAM policy attached to GlueServiceRole includes s3:PutObject (and typically s3:PutObjectAcl) on the target bucket/prefix, since Glue writes output objects to S3 using that role. Option C is also correct because an explicit Deny in the S3 bucket policy overrides any Allow in the IAM policy, so the engineer must check whether the bucket policy denies access from GlueServiceRole. Option A is wrong because disabling S3 Block Public Access is unrelated to a role-based write failure and would weaken security without fixing the permission issue.

Option B is wrong because the failure is writing to S3, not to the Glue Data Catalog, so Data Catalog permissions would not resolve the S3 Access Denied error. Option E is wrong because S3 is accessed via AWS public service endpoints and does not require the Glue job to be in the same VPC as the bucket; VPC placement only matters for private endpoint configurations, not for this authorization error.

Exam trap

The trap here is that candidates may confuse S3 access errors with network or VPC issues, but S3 is a global service and access is governed by IAM and bucket policies, not VPC placement.

282
Multi-Selectmedium

A data engineer needs to store event data from IoT devices that arrives in bursts. The data is key-value and requires single-digit millisecond read and write latency. The engineer also needs to run complex analytical queries on the data for reporting. Which TWO services should be used together? (Choose TWO.)

Select 2 answers
A.Amazon DynamoDB
B.Amazon ElastiCache for Redis
C.Amazon Redshift
D.Amazon S3
E.Amazon RDS for MySQL
AnswersA, C

Amazon DynamoDB delivers consistent single-digit millisecond latency for key-value workloads, satisfying the burst IoT ingestion requirement. Its partition-based architecture scales horizontally without manual sharding. However, DynamoDB alone cannot run complex analytical queries, so it must pair with a separate analytics service to meet the reporting constraint.

Why this answer

Amazon DynamoDB (A) is correct because it is a fully managed key-value and document NoSQL database that delivers consistent single-digit millisecond read and write latency at any scale, making it ideal for bursty IoT event ingestion. Amazon Redshift (C) is correct because it is a petabyte-scale, columnar data warehouse designed for complex analytical queries and reporting, so it complements DynamoDB by handling the analytics workload. Together, DynamoDB captures the high-throughput key-value events while Redshift serves the reporting and complex query needs.

Amazon ElastiCache for Redis (B) is an in-memory cache, not a durable primary store for event data, and it does not natively run complex analytical SQL queries. Amazon S3 (D) is object storage with higher latency and no native single-digit millisecond key-value access or complex query engine. Amazon RDS for MySQL (E) is a relational database that generally cannot match DynamoDB's single-digit millisecond latency at burst scale and is not optimized for large-scale complex analytical reporting like Redshift.

Exam trap

The trap here is that candidates often choose ElastiCache for Redis because of its low latency, forgetting that it is not a durable data store for analytical queries, or they pick S3 thinking it can serve as a primary database, ignoring its lack of single-digit millisecond latency for key-value access.

283
MCQmedium

A data engineering team needs to encrypt data at rest in an Amazon S3 bucket that stores sensitive customer information. The team must use an AWS Key Management Service (AWS KMS) customer managed key with automatic rotation enabled. Which configuration meets these requirements?

A.Use default encryption with SSE-KMS and specify the customer managed key ID.
B.Use default encryption with SSE-S3.
C.Use default encryption with SSE-C and provide a customer-provided key.
D.Use default encryption with SSE-KMS and leave the key ID empty to use the AWS managed key.
AnswerA

SSE-KMS with a customer managed key encrypts objects at rest under a key the team controls, and enabling automatic rotation on that key satisfies the rotation requirement. SSE-S3 and AWS managed keys cannot provide customer-controlled rotation.

Why this answer

It enables SSE-KMS with a customer managed key that supports automatic key rotation, meeting the requirement for a customer managed key with automatic rotation. Option B is incorrect because SSE-S3 uses AWS managed keys that are not customer-controlled and do not support automatic rotation. Option C is incorrect because SSE-C requires the customer to provide and manage their own keys, and does not support automatic rotation.

Option D is incorrect because leaving the key ID empty defaults to the AWS managed KMS key, which is not a customer managed key and does not support automatic rotation.

284
MCQhard

A media company ingests thousands of small JSON files per hour into an Amazon S3 bucket. A data engineer needs to convert these files into a compact, columnar format for efficient querying with Amazon Athena. The engineer wants to minimize storage costs and improve query performance. Which approach should the engineer take?

A.Use AWS Glue ETL to read the JSON files, convert them to Apache Parquet, and write the output to a new S3 prefix partitioned by date.
B.Use Amazon Athena to create a new table as Parquet using CREATE TABLE AS SELECT (CTAS) from the JSON table.
C.Use Amazon Kinesis Data Firehose to convert JSON to Parquet and deliver to S3.
D.Use AWS Lambda to read each JSON file, convert to Parquet, and write back to S3.
AnswerA

AWS Glue ETL can read JSON, transform to Parquet, and write partitioned data to S3. Parquet is columnar, reducing storage and improving Athena query performance. Partitioning by date further reduces data scanned. This directly addresses the need for compact columnar format and cost efficiency.

Why this answer

AWS Glue ETL is a fully managed extract, transform, and load service that can efficiently process large volumes of data. It can read JSON from S3, convert to Parquet, and write partitioned output. This reduces storage costs due to Parquet's compression and columnar format, and improves Athena query performance by reducing data scanned.

Partitioning by date further optimizes queries.

Exam trap

The trap here is assuming that streaming services like Kinesis Data Firehose or serverless functions like Lambda are the best fit for batch conversion of existing small files, when a managed ETL service like AWS Glue is more appropriate.

285
Multi-Selectmedium

Which TWO options are valid ways to reduce storage costs for an Amazon S3 data lake that stores historical data rarely accessed after 30 days? (Choose TWO.)

Select 2 answers
A.Enable S3 Transfer Acceleration for all uploads.
B.Create a lifecycle policy to transition objects to S3 Standard-IA after 30 days.
C.Create a lifecycle policy to delete objects after 30 days.
D.Create a lifecycle policy to transition objects to S3 Glacier Deep Archive after 90 days.
E.Enable S3 Versioning to preserve all object versions.
AnswersB, D

Standard-IA reduces storage cost for infrequent access.

Why this answer

S3 Standard-IA (Infrequent Access) is designed for data accessed less frequently but requires rapid access when needed. Transitioning objects to Standard-IA after 30 days reduces storage costs compared to S3 Standard while maintaining low-latency retrieval, making it ideal for a data lake where historical data is rarely accessed after the first month.

Exam trap

The trap here is that candidates often confuse data protection features (like Versioning) with cost optimization, or they assume that deleting data is the only way to reduce costs, overlooking lifecycle transitions to lower-cost storage classes that retain data accessibility.

286
MCQhard

A data engineer is using AWS Lake Formation to manage access to a data lake in Amazon S3. The engineer needs to grant a specific IAM role read access to only the columns 'customer_id' and 'purchase_amount' in a table stored in the AWS Glue Data Catalog. The table contains sensitive columns like 'credit_card_number'. Which Lake Formation permission model should the engineer use to achieve this?

A.Grant the IAM role DESCRIBE permission on the table, and then use AWS Glue to create a transformed dataset with only the required columns.
B.Grant the IAM role SELECT permission on the table, and then use an IAM policy to deny access to the sensitive columns.
C.Grant the IAM role SELECT permission on the table, and then create a data filter that includes only the required columns.
D.Create a view in Amazon Athena that selects only the required columns, and grant the IAM role access to the view.
AnswerC

Lake Formation supports column-level security through data filters. A data filter allows you to specify which columns are included or excluded. By granting SELECT on the table and then creating a data filter that includes only customer_id and purchase_amount, the role can access only those columns. This is the correct way to implement column-level access control in Lake Formation.

Why this answer

AWS Lake Formation provides fine-grained access control, including column-level security. To grant read access to specific columns, you grant SELECT permission on the table and then attach a data filter that includes only the allowed columns. Data filters are evaluated at query time, ensuring that the role can only access the specified columns.

This is the most direct and secure method.

Exam trap

The trap here is thinking that IAM policies can enforce column-level access within Lake Formation, when Lake Formation uses its own permission model with data filters.

287
MCQmedium

A company uses Amazon RDS for MySQL to store transactional data. The database contains sensitive financial information. The company's security policy requires that all data at rest be encrypted using a customer-managed KMS key. The database was originally launched without encryption at rest. The security team now needs to enable encryption without significant downtime. What should they do?

A.Create a snapshot of the database, copy the snapshot with encryption enabled, and restore a new DB instance from the encrypted snapshot.
B.Enable encryption by modifying the DB instance's storage type to 'encrypted'.
C.Use the AWS DMS (Database Migration Service) to migrate data to a new encrypted RDS instance.
D.Modify the DB instance and enable encryption under the 'Storage' settings.
AnswerA

RDS does not support encrypting an existing unencrypted instance in place. Snapshot-copy with encryption re-encrypts the data under a customer-managed KMS key, and restoring creates a new encrypted instance, meeting the policy with only the brief downtime of a snapshot restore.

Why this answer

RDS for MySQL does not support enabling encryption at rest on an existing unencrypted DB instance. The only supported method is to take a snapshot, copy the snapshot with encryption enabled using a customer-managed KMS key, and then restore a new DB instance from that encrypted snapshot. This approach requires a brief downtime during the switchover but avoids a full data migration.

Exam trap

DEA-C01 often tests the misconception that you can enable encryption on an existing RDS instance by modifying it, when in fact encryption must be applied at creation or via snapshot restore, leading candidates to choose the modify option.

How to eliminate wrong answers

Option B is wrong because modifying the storage type of an RDS instance does not enable encryption; encryption must be enabled at creation time or via snapshot restore. Option C is wrong because AWS DMS is used for migrating data between different database engines or platforms, and while it could be used, it is not the recommended or simplest method for enabling encryption on an existing RDS instance. Option D is wrong because the RDS modify operation does not provide an option to enable encryption on an existing unencrypted instance; encryption settings are immutable after creation unless you restore from an encrypted snapshot.

288
MCQeasy

A data engineer needs to monitor an AWS Glue ETL job that runs daily. The job sometimes fails due to missing partitions in the Data Catalog. The engineer wants to receive an alert when the job fails. What is the MOST operationally efficient way to achieve this?

A.Enable AWS Glue job bookmarks and configure an Amazon CloudWatch Logs subscription filter to email the engineer on error patterns.
B.Use AWS Glue job event notifications via Amazon EventBridge to trigger an AWS Lambda function that sends an Amazon SNS notification.
C.Create an Amazon CloudWatch alarm on the Glue job's 'glue.driver.aggregate.numFailedTasks' metric and notify an Amazon SNS topic.
D.Schedule a daily AWS Lambda function that calls the Glue GetJobRuns API and sends an email if the last run failed.
AnswerB

AWS Glue emits job state change events to Amazon EventBridge. You can create an EventBridge rule that matches Glue job state changes (e.g., FAILED, TIMEOUT) and targets a Lambda function or SNS topic directly. This is a serverless, operationally efficient approach that requires no polling and provides near-real-time alerts on job failures, including those caused by missing partitions.

Why this answer

Amazon EventBridge can capture AWS Glue job state change events and route them to targets like Lambda or SNS. This provides immediate, event-driven notifications on job failures without polling or custom log parsing. It is the most operationally efficient method for alerting on Glue job failures, including those due to missing partitions.

Exam trap

The trap here is assuming that CloudWatch metrics or log filters are the primary way to alert on Glue job failures, when EventBridge events are the most direct and efficient mechanism.

289
MCQeasy

A data engineer is using AWS Glue Studio to build a visual ETL job that reads JSON files from Amazon S3, applies a mapping transform, and writes Parquet to another S3 location. The engineer notices that the job is writing many small files, which hurts downstream query performance. Which action should the engineer take to reduce the number of output files without changing the source data?

A.Use the repartition or coalesce transform before the S3 target to reduce the number of output partitions.
B.Increase the number of DPUs for the Glue job so more workers write files in parallel.
C.Enable the Glue job bookmark so the job only processes new files on each run.
D.Change the output format from Parquet to CSV so the files are smaller and easier to manage.
AnswerA

Repartition or coalesce changes the number of partitions in the DataFrame, which directly controls how many files are written. Coalesce is especially useful when reducing partitions without a full shuffle. Applying it before the target write consolidates output into fewer, larger files and improves downstream query performance.

Why this answer

The number of files written by a Glue job corresponds to the number of partitions in the DataFrame at write time. Repartition or coalesce reduces that partition count, so the job writes fewer, larger files. This directly addresses the small-file issue without altering the source data or the output format.

Exam trap

The trap here is assuming that more compute resources will consolidate output, when adding workers usually increases the number of output files.

290
MCQeasy

A company stores sensitive data in Amazon S3. To meet compliance requirements, they need to ensure that any data older than 1 year is automatically moved to a lower-cost storage class. Which S3 feature should they use?

A.S3 Replication
B.S3 Lifecycle policies
C.S3 Glacier
D.S3 Intelligent-Tiering
AnswerB

S3 Lifecycle policies define transition rules that automatically move objects to a lower-cost storage class once they reach a specified age, such as 365 days. This satisfies the stem's compliance requirement that data older than one year be moved automatically without manual intervention.

Why this answer

S3 Lifecycle policies let you define rules that automatically transition objects to a lower-cost storage class (e.g., S3 Standard-IA, S3 Glacier Instant Retrieval, S3 Glacier Flexible Retrieval) after a specified age, such as 365 days. This is the native, declarative mechanism for age-based storage class transitions and satisfies the compliance requirement without custom code.

Exam trap

DEA-C01 often tests the difference between a storage class (Glacier) and the feature that moves data into it (Lifecycle policy) — the trap is choosing 'S3 Glacier' as if it were an automated tiering mechanism.

How to eliminate wrong answers

Option A is wrong because S3 Replication copies objects to another bucket or Region for durability/DR or latency — it does not change storage class based on age and does not reduce cost by itself. Option B is correct. Option C is wrong because S3 Glacier is a storage class, not a feature — you cannot 'use S3 Glacier' to move data automatically; you must use a lifecycle policy (or Intelligent-Tiering) to transition objects into Glacier.

Option D is wrong because S3 Intelligent-Tiering automatically moves objects between access tiers based on changing access patterns, not based on a fixed age threshold like 1 year, and it charges a monitoring fee — it does not meet a deterministic 'older than 1 year' compliance rule.

291
MCQmedium

A data engineer is configuring an S3 bucket to host sensitive data. The security policy requires that all objects be encrypted with a key that is generated and managed by the customer, and that the key be stored in AWS KMS. Which encryption option should be used?

A.Server-Side Encryption with S3-Managed Keys (SSE-S3)
B.Client-Side Encryption
C.Server-Side Encryption with Customer-Provided Keys (SSE-C)
D.Server-Side Encryption with AWS KMS-Managed Keys (SSE-KMS)
AnswerD

SSE-KMS encrypts objects using keys held in AWS KMS, and selecting a customer managed key means the customer generates and controls that key, including its policy and rotation. SSE-S3 uses AWS-owned keys the customer cannot manage.

Why this answer

SSE-KMS allows you to use customer-managed keys stored in AWS KMS, meeting the requirement for customer-generated and managed keys. Option A is incorrect because SSE-S3 uses AWS-managed keys, not customer-managed. Option B is incorrect because client-side encryption encrypts data outside S3, not using S3 server-side encryption.

Option C is incorrect because SSE-C uses customer-provided keys that you manage yourself, but they are not stored in AWS KMS—you must provide them with each request.

292
Multi-Selecteasy

A data engineer is setting up a data pipeline to ingest streaming data from an IoT fleet. The data must be processed in near real-time and stored in Amazon S3 for analytics. Which THREE AWS services should the engineer consider using?

Select 3 answers
A.Amazon EMR
B.Amazon Kinesis Data Firehose
C.AWS Lambda
D.AWS Glue
E.Amazon Kinesis Data Streams
AnswersB, C, E

Amazon Kinesis Data Firehose satisfies the near real-time ingestion and S3 delivery constraints by buffering streaming records and writing them directly to Amazon S3 without custom consumer code. It handles scaling, batching and format conversion automatically, so IoT telemetry lands in S3 for analytics with minimal operational overhead.

Why this answer

Amazon Kinesis Data Firehose (B) is correct because it is a fully managed service designed to reliably load streaming data directly into Amazon S3 (and other destinations) with near real-time delivery, requiring no server management. Amazon Kinesis Data Streams (E) is correct because it ingests and buffers high-throughput streaming data from IoT fleets in real time, allowing custom consumers to process the data before it is stored in S3. AWS Lambda (C) is correct because it can be invoked by Kinesis to process streaming records in near real-time, enabling serverless transformation or enrichment of the IoT data before it lands in S3.

Amazon EMR (A) is not the best fit here because it is a batch-oriented big data processing platform (Hadoop/Spark) rather than a streaming ingestion service, and AWS Glue (D) is primarily a serverless ETL and data catalog service for batch and some streaming jobs, not a dedicated real-time ingestion pipeline component for this scenario.

Exam trap

DEA-C01 often tests the distinction between real-time streaming ingestion (Kinesis Data Streams) and near-real-time delivery to S3 (Kinesis Data Firehose), causing candidates to overlook Lambda as the processing layer or incorrectly select batch-oriented services like EMR or Glue.

293
MCQmedium

A data engineer is setting up a data pipeline that ingests streaming data from Amazon Kinesis Data Streams into an S3 data lake using Amazon Kinesis Data Firehose. The data contains personally identifiable information (PII). The security team requires that all data be encrypted at rest in S3 using an AWS KMS customer managed key (CMK) that is specific to the application. Additionally, the data must be encrypted in transit between all services. The engineer creates the KMS key and configures Firehose to use server-side encryption with the key for the S3 destination. However, Firehose delivery fails with an error indicating that the KMS key is not accessible. What is the most likely cause?

A.The KMS key policy does not grant the firehose.amazonaws.com service principal the required permissions.
B.The Kinesis data stream is not encrypted at rest.
C.The Firehose delivery stream is not in the same region as the KMS key.
D.The S3 bucket policy does not grant the Firehose delivery stream access to write objects.
AnswerA

Firehose assumes the `firehose.amazonaws.com` service principal to call KMS GenerateDataKey and Decrypt on your behalf, so the CMK's key policy must explicitly grant that principal those actions. Without this grant, the key is inaccessible and delivery fails, even though the IAM role used for S3 access is correctly configured.

Why this answer

When Firehose writes to S3 using SSE-KMS with a customer managed key, the Firehose service principal (firehose.amazonaws.com) must be granted kms:GenerateDataKey and kms:Decrypt in the KMS key policy. Without that grant, Firehose cannot obtain the data key to encrypt objects, and delivery fails with an access-denied error referencing the KMS key. The key policy is the resource-based control that authorizes the service principal.

Exam trap

DEA-C01 often tests the two-layer KMS authorization model — candidates focus on IAM roles or S3 bucket policies and forget that the KMS key policy must explicitly grant the AWS service principal access.

How to eliminate wrong answers

Option B is wrong because Kinesis Data Streams encryption at rest is independent of Firehose's ability to write encrypted objects to S3; an unencrypted stream does not block KMS-based S3 encryption. Option C is wrong because KMS keys are regional and Firehose and the key must be in the same region, but the question states the engineer created the key and configured Firehose, and the error is specifically about key accessibility, not region mismatch. Option D is wrong because an S3 bucket policy denying write access would produce an S3 access-denied error, not a KMS key accessibility error; the failure occurs at the encryption step before the PutObject call.

294
MCQmedium

A data engineer is designing a data store for a real-time analytics application that requires low-latency reads and writes at scale. The data model includes time-series data with high ingest rates and queries that aggregate data over sliding time windows. The engineer needs a fully managed AWS service that supports automatic scaling and can handle millions of writes per second. Which service should the engineer choose?

A.Amazon S3
B.Amazon Redshift
C.Amazon RDS for PostgreSQL
D.Amazon DynamoDB
AnswerD

DynamoDB is a fully managed NoSQL database that supports high-throughput reads and writes with automatic scaling. It can handle millions of requests per second and provides low-latency performance. With features like time-to-live (TTL) and on-demand capacity mode, it is well-suited for time-series data and real-time analytics. Its ability to scale horizontally without downtime makes it the best fit for this scenario.

Why this answer

Amazon DynamoDB is a fully managed NoSQL database that provides low-latency performance at any scale, with automatic scaling and support for high-throughput workloads. It can handle millions of writes per second and is ideal for time-series data due to features like TTL and on-demand capacity. Its ability to scale horizontally without downtime makes it the best choice for real-time analytics applications requiring low-latency reads and writes.

Exam trap

The trap here is assuming that a relational database or data warehouse can handle the same scale and latency requirements as a purpose-built NoSQL service like DynamoDB.

295
MCQeasy

A data engineer needs to schedule a recurring AWS Glue ETL job that must run every night at 02:00 UTC and must not start a new run while a previous run is still executing. The engineer wants the simplest managed scheduling option that integrates natively with Glue job run state. Which approach should the engineer use?

A.Create an AWS Step Functions state machine with a Wait state that loops every 24 hours and starts the Glue job.
B.Use an AWS Lambda function invoked by a CloudWatch Events rule to call StartJobRun on a fixed schedule.
C.Configure a Glue trigger of type SCHEDULED with a cron expression and set the job's maximum concurrency to 1.
D.Create an Amazon EventBridge scheduled rule with a cron expression that invokes the Glue job via a target.
AnswerC

A SCHEDULED Glue trigger natively starts the job on a cron schedule, and setting the job's maximum concurrency to 1 prevents a new run from starting while a previous run is still active. Both settings live within Glue, so no external orchestration service is required. This is the simplest managed approach that meets the time-based schedule and the no-overlap requirement.

Why this answer

Glue scheduled triggers use cron expressions to start jobs at fixed times, and the job's maximum concurrency setting controls whether overlapping runs are allowed. Setting maximum concurrency to 1 ensures that if a run is still active when the next trigger fires, the new run does not start. This keeps scheduling and concurrency control entirely within Glue, which is the simplest managed solution.

Exam trap

The trap here is reaching for an external scheduler like EventBridge or Lambda when Glue already provides native cron triggers and concurrency limits.

296
MCQmedium

A data engineer uses Amazon EMR to run a Spark job that reads from S3 and writes to HDFS on the cluster. The job fails with an 'OutOfMemoryError: Java heap space' error in the executors. Which parameter adjustment should be made to resolve this?

A.Increase spark.default.parallelism
B.Increase spark.sql.shuffle.partitions
C.Increase spark.executor.memory
D.Increase spark.driver.memory
AnswerC

Raising spark.executor.memory enlarges each executor's JVM heap, directly addressing the 'OutOfMemoryError: Java heap space' thrown during Spark execution. Since the failure occurs in executors rather than the driver, this parameter targets the constrained component, giving shuffle and aggregation buffers sufficient headroom to complete the S3-to-HDFS job.

Why this answer

The OutOfMemoryError: Java heap space in Spark executors indicates that the executor JVM heap is insufficient to hold the data being processed. Increasing spark.executor.memory allocates more heap space to each executor, allowing it to handle larger partitions or aggregations without running out of memory. This directly addresses the root cause of the error.

Exam trap

DEA-C01 often tests the confusion between driver and executor memory, and the misconception that increasing parallelism or shuffle partitions directly solves heap memory errors.

How to eliminate wrong answers

Option A is wrong because increasing spark.default.parallelism changes the number of partitions for RDD operations, which can increase parallelism but does not increase the memory available to each executor; it may even create more tasks that compete for the same heap. Option B is wrong because increasing spark.sql.shuffle.partitions controls the number of partitions after a shuffle, which can reduce the size of each partition but does not increase executor heap; it might help if partitions are too large, but the error is specifically about heap space, so increasing memory is more direct. Option D is wrong because increasing spark.driver.memory affects the driver JVM, not the executors; the error is in the executors, so this would not resolve it.

297
Multi-Selecteasy

Which TWO AWS services can be used to transform data in transit during ingestion? (Choose 2.)

Select 2 answers
A.Amazon S3 Transfer Acceleration
B.Amazon Kinesis Data Firehose with Lambda transformation
C.AWS Glue ETL
D.Amazon Athena
E.AWS Data Pipeline
AnswersB, C

Firehose supports invoking a Lambda function to transform each record in flight before delivery, satisfying the stem's in-transit transformation requirement. The transformation occurs during ingestion rather than after landing, so downstream consumers receive already-processed data.

Why this answer

Amazon Kinesis Data Firehose can invoke an AWS Lambda function to transform streaming data in real time before delivering it to a destination, making it suitable for transforming data in transit during ingestion. AWS Glue ETL can be used for both batch and streaming transformations; specifically, AWS Glue supports streaming ETL jobs that can transform data in transit as it is ingested from sources like Amazon MSK or Kinesis Data Streams. Therefore, both services can transform data during ingestion, albeit with different use cases and latency characteristics.

Exam trap

Candidates often mistakenly think that AWS Glue ETL is only for batch processing and cannot transform data in transit, but AWS Glue supports streaming ETL jobs that can transform data during ingestion, making it a valid choice for in-transit transformation.

298
Multi-Selecthard

A data engineer is using AWS Glue to transform data from Amazon S3 and load it into Amazon Redshift. The engineer needs to improve the performance of the Glue job, which currently takes several hours to complete. The job reads large Parquet files, performs joins, and writes to Redshift. Which two actions should the engineer take to improve performance? (Choose two.)

Select 2 answers
A.Partition the source Parquet data in Amazon S3 by commonly filtered columns and use partition pruning in the Glue job.
B.Enable job bookmarks to avoid reprocessing previously processed data.
C.Increase the number of AWS Glue DPUs allocated to the job to add more Spark executors.
D.Convert the Parquet files to CSV to reduce storage size and speed up reading.
E.Use the 'Relationalize' transform to flatten the Parquet data before joining.
AnswersA, C

Partitioning the source data by columns used in filters allows Glue to read only relevant partitions, reducing I/O and the amount of data shuffled during joins. This can dramatically improve performance for large datasets. The engineer should ensure the Glue job uses predicate pushdown and that the partition columns are used in the query. This is a best practice for optimizing Glue ETL on S3.

Why this answer

Increasing DPUs adds compute resources to parallelize the job, and partitioning the source data enables partition pruning to reduce the amount of data read. Both actions directly address performance bottlenecks in a large Glue job. Other options either do not affect single-run performance or would degrade it.

Exam trap

The trap here is assuming that job bookmarks improve performance for a single large run, when they only help avoid reprocessing data across runs.

299
MCQhard

A company uses AWS DMS to migrate a 2 TB Oracle database to Amazon RDS for PostgreSQL. The migration completes successfully, but data validation shows some tables have missing rows. The task is configured for ongoing replication using change data capture (CDC). What is the MOST likely cause of the missing rows?

A.Source database archive log retention period too short
B.Large objects (LOBs) not supported by the target
C.Source tables missing primary keys
D.Insufficient storage on the DMS replication instance
AnswerC

DMS applies CDC changes using the source table's primary key to identify and update the correct target rows. Tables lacking a primary key cannot be matched reliably, so updates and deletes are dropped or misapplied, producing missing rows in the PostgreSQL target after replication.

Why this answer

AWS DMS requires a primary key or unique index on source tables to reliably identify and apply row changes during CDC. Without a primary key, DMS cannot uniquely match rows for updates and deletes, and it may also fail to capture all inserts during the initial load plus CDC handoff, resulting in missing rows.

Exam trap

DEA-C01 often tests the misconception that DMS CDC works on any table, when missing primary keys silently break change capture and cause missing rows.

How to eliminate wrong answers

Option A is wrong because insufficient archive log retention would cause the CDC task to fail with a specific error about missing log files, not silently drop rows; DMS would stop rather than skip. Option B is wrong because unsupported LOBs cause LOB truncation or task failure with explicit errors, not missing rows in non-LOB tables. Option D is wrong because insufficient replication instance storage causes the task to fail with a storage-full error, not selective row loss.

300
MCQmedium

A media company stores video metadata in an Amazon DynamoDB table. The security team requires that all data at rest in the table be encrypted with a customer managed key in AWS KMS, and that the key usage be auditable. The data engineer needs to configure encryption for the table. Which action should the data engineer take?

A.Use DynamoDB Accelerator (DAX) with encryption in transit enabled, and configure the DAX cluster to use a customer managed KMS key.
B.Enable DynamoDB encryption at rest using the default AWS owned key, and enable CloudTrail logging for DynamoDB.
C.Enable DynamoDB encryption at rest using an AWS managed key, and use AWS CloudTrail to monitor key usage.
D.Create a customer managed KMS key and specify it when creating the DynamoDB table, ensuring the key policy allows DynamoDB to use it for encryption and decryption.
AnswerD

DynamoDB supports customer managed KMS keys for encryption at rest. By specifying the key during table creation and granting DynamoDB permissions in the key policy, the table data is encrypted with that key. KMS key usage is logged in CloudTrail, providing the required auditability. This meets the security team's requirements.

Why this answer

DynamoDB encryption at rest can use customer managed KMS keys. Specifying such a key during table creation ensures the table is encrypted with that key. The key policy must grant DynamoDB permission to use the key.

KMS key usage is logged in CloudTrail, providing auditability. Other options use keys that are not customer managed or do not encrypt the table at rest.

Exam trap

The trap here is assuming that AWS managed keys or DAX provide the same control and auditability as customer managed keys for DynamoDB encryption at rest.

Page 3

Page 4 of 18

Page 5