Courseiva

AWS Certified Data Engineer Associate DEA-C01 (DEA-C01) — Questions 1126–1200

1321 questions total · 18pages · All types, answers revealed

Page 15

Page 16 of 18

Page 17
1126
MCQhard

A company runs an Amazon Redshift cluster with 10 RA3 nodes. The data warehouse stores 50 TB of data. The company notices that queries are slow and the cluster's storage utilization is high. The data engineer needs to improve query performance and reduce storage costs without changing the cluster's node count. Which action should the engineer take?

A.Use Redshift Spectrum to offload historical data to Amazon S3 and query it in place.
B.Change the distribution style of large tables to DISTSTYLE ALL.
C.Migrate the cluster to Dense Compute node types.
D.Enable concurrency scaling to handle more concurrent queries.
AnswerA

Redshift Spectrum queries external tables in Amazon S3 directly, so historical data leaves RA3 managed storage without altering node count. This lowers storage utilisation and cost while letting the cluster focus compute on hot data, addressing both stated constraints.

Why this answer

Redshift Spectrum allows you to query data directly from Amazon S3 without loading it into the cluster. By offloading historical or less-frequently accessed data to S3, you reduce the storage utilization on the RA3 nodes, which frees up managed storage and can improve query performance. This approach also lowers storage costs because S3 is cheaper than Redshift managed storage, and it does not change the node count.

Exam trap

The trap here is that candidates often confuse concurrency scaling (which improves query throughput) with storage optimization, or they assume that changing distribution styles (like DISTSTYLE ALL) will always improve performance, ignoring the storage cost impact in a high-utilization scenario.

How to eliminate wrong answers

Option B is wrong because changing large tables to DISTSTYLE ALL replicates the entire table to every node, which increases storage utilization and can worsen the high storage issue, not reduce it. Option C is wrong because migrating to Dense Compute nodes would change the node type, which violates the constraint of not changing the cluster's node count; also, Dense Compute nodes use local SSD storage and are not designed for the same storage-to-compute ratio as RA3 nodes. Option D is wrong because concurrency scaling adds additional compute capacity to handle more concurrent queries but does not reduce storage utilization or costs; it addresses throughput, not the underlying storage pressure.

1127
MCQeasy

A data engineer needs to run a one-time AWS Glue ETL job that reads data from an Amazon S3 bucket and writes transformed Parquet files to another S3 bucket. The job does not need a schedule and should be run immediately after creation. Which method should the engineer use to start the job?

A.Configure a scheduled trigger with a cron expression that runs once.
B.Create a trigger with type ON_DEMAND and activate it.
C.Create a workflow and add the job to it, then start the workflow.
D.Use the AWS Glue console or AWS CLI to start the job run directly.
AnswerD

AWS Glue jobs can be started immediately by using the StartJobRun API, the AWS CLI command start-job-run, or the console's Run job button. This is the most straightforward method for a one-time, unscheduled job run and does not require any trigger configuration.

Why this answer

For a one-time, unscheduled AWS Glue job, the simplest way to start it is to invoke StartJobRun through the console, AWS CLI, or SDK. Triggers and workflows are used for scheduling or orchestration and are not needed when the job should run immediately after creation.

Exam trap

The trap here is assuming that a trigger or workflow is always required to start an AWS Glue job, when a direct start is sufficient for on-demand runs.

1128
Multi-Selecthard

A data engineer is designing a data lake on Amazon S3 with sensitive data. The engineer needs to ensure that data at rest is encrypted and that access is logged for compliance. Which TWO actions should the engineer take? (Choose TWO.)

Select 2 answers
A.Enable S3 Select to filter data at read time.
B.Enable S3 Transfer Acceleration.
C.Enable CloudTrail data events for S3 object-level operations.
D.Enable default encryption on the S3 bucket using SSE-KMS.
E.Enable S3 Block Public Access on the account.
AnswersC, D

CloudTrail data events capture object-level S3 operations such as GetObject and PutObject, which management events omit entirely. This satisfies the stem's requirement that access be logged for compliance, since only data events record who read or wrote the sensitive objects.

Why this answer

Option D is correct because enabling default bucket encryption with SSE-KMS ensures all objects written to the S3 bucket are encrypted at rest using AWS KMS-managed keys, satisfying the data-at-rest encryption requirement. Option C is correct because CloudTrail data events capture object-level S3 operations such as GetObject and PutObject, providing the access logging needed for compliance auditing. Option A is incorrect because S3 Select only filters data at read time and does not provide encryption or access logging.

Option B is incorrect because S3 Transfer Acceleration only speeds up uploads and downloads over long distances, not encryption or logging. Option E is incorrect because S3 Block Public Access prevents public exposure but does not encrypt data at rest or log access.

Exam trap

The trap here is that candidates often confuse S3 Block Public Access (a security control) with encryption or logging, or they mistakenly think S3 Select or Transfer Acceleration contribute to compliance requirements, when they are unrelated to data-at-rest encryption and access logging.

1129
MCQmedium

A company needs to ingest data from multiple SaaS applications (e.g., Salesforce, Marketo) into Amazon S3 for analytics. The data sources have different schemas and update frequencies. Which AWS service should be used to build this ingestion pipeline with minimal code?

A.AWS Data Pipeline
B.AWS Glue
C.Amazon Kinesis Data Firehose
D.Amazon AppFlow
AnswerD

Amazon AppFlow provides managed, low-code connectors for SaaS sources such as Salesforce and Marketo, handling differing schemas and update schedules natively. It satisfies the minimal-code requirement by removing custom ingestion scripts, delivering data directly into Amazon S3 for analytics.

Why this answer

Amazon AppFlow (Option D) is the correct answer because it is purpose-built for ingesting data from SaaS applications like Salesforce and Marketo into Amazon S3 with minimal code. It supports various source connectors, handles schema variations, and allows scheduling based on update frequencies. AWS Data Pipeline (A) requires more manual configuration and code for connectors.

AWS Glue (B) has connectors but is more complex for this use case, often requiring additional ETL scripting. Amazon Kinesis Data Firehose (C) is designed for streaming data, not batch extraction from SaaS APIs.

1130
MCQeasy

A company uses an Amazon RDS for MySQL DB instance with Multi-AZ deployment. The primary DB instance fails unexpectedly. What happens to the database endpoint?

A.A new endpoint is created for the standby and the application must use the new endpoint.
B.The existing endpoint continues to work and automatically points to the standby DB instance.
C.The database becomes unavailable until the primary is restored from a snapshot.
D.The existing endpoint is deleted and a new endpoint is provided after manual DNS update.
AnswerB

The DNS endpoint is a stable CNAME that Multi-AZ failover repoints to the standby, which is promoted to primary. Applications therefore keep using the same endpoint without reconfiguration, satisfying the requirement that connectivity survives the unexpected primary failure.

Why this answer

In a Multi-AZ RDS deployment, the DNS endpoint remains unchanged during a failover. When the primary DB instance fails, Amazon RDS automatically updates the DNS record to point to the standby instance in the other Availability Zone. This ensures the application can continue using the same endpoint without any manual intervention, providing high availability.

Exam trap

The trap here is that candidates may think a new endpoint is created or that manual DNS changes are required, confusing Multi-AZ failover with a manual snapshot restore or a cross-region read replica promotion.

How to eliminate wrong answers

Option A is wrong because the DNS endpoint is not recreated; it remains the same and is automatically remapped to the standby instance. Option C is wrong because Multi-AZ failover is automatic and typically completes within 1-2 minutes, so the database does not become unavailable until a snapshot restore is performed. Option D is wrong because no manual DNS update is required; the existing endpoint is automatically updated by RDS to point to the standby instance.

1131
MCQeasy

A company wants to ingest data from multiple SaaS applications into Amazon S3 using a fully managed service that supports schema discovery and transformation. Which AWS service should they use?

A.Amazon Kinesis Data Firehose
B.Amazon AppFlow
C.AWS Glue
D.AWS Data Pipeline
AnswerB

Amazon AppFlow is the fully managed ingestion service that connects SaaS sources such as Salesforce and SAP to Amazon S3, performing schema discovery and optional transformation during the flow. It satisfies the no-code, managed requirement without building custom extraction pipelines.

Why this answer

Amazon AppFlow is a fully managed integration service that enables you to securely transfer data between SaaS applications and AWS services like Amazon S3. It supports schema discovery and transformation, making it ideal for ingesting data from multiple SaaS apps without writing code.

Exam trap

The trap is confusing AppFlow with AWS Glue, as both can move data, but AppFlow is specifically for SaaS integration with built-in connectors and transformations.

How to eliminate wrong answers

Option A is wrong because Kinesis Data Firehose is designed for streaming data and does not natively connect to SaaS applications or perform schema discovery. Option C is wrong because AWS Glue is a serverless ETL service that can connect to SaaS via custom connectors, but it is not fully managed for SaaS integration and requires more setup. Option D is wrong because AWS Data Pipeline is a legacy service for orchestrating data workflows, not specifically for SaaS ingestion.

1132
MCQeasy

A data engineer needs to move 50 TB of existing data from an on-premises data center into Amazon S3 as a one-time migration. The data center has a 1 Gbps internet connection that is shared with production traffic, and the migration must complete within two weeks without disrupting production. Which approach should the engineer use?

A.Use Amazon S3 Transfer Acceleration with multipart uploads from the data center.
B.Use AWS Snowball Edge Storage Optimized devices to ship the data to AWS.
C.Use AWS Transfer Family with SFTP to upload the files to Amazon S3.
D.Use AWS DataSync over the internet to copy the data directly to Amazon S3.
AnswerB

Snowball Edge Storage Optimized devices are designed for large one-time migrations where network transfer is impractical. Shipping devices avoids consuming the shared 1 Gbps link and can move 50 TB well within two weeks. The data is copied to the device on-premises and imported into S3 after AWS receives it, meeting both the timing and non-disruption requirements.

Why this answer

A 50 TB one-time migration over a shared 1 Gbps link cannot reliably finish in two weeks without harming production traffic. AWS Snowball Edge Storage Optimized devices move the data physically, sidestepping the bandwidth constraint entirely. Network-based options such as DataSync, Transfer Family, or Transfer Acceleration all depend on the same limited connection and therefore do not meet the stated requirements.

Exam trap

The trap here is focusing on transfer speed features like Transfer Acceleration while ignoring that the shared 1 Gbps link is the real bottleneck.

1133
MCQmedium

Refer to the exhibit. A data engineer configured the lifecycle policy shown. The 'logs/' prefix contains important audit logs. After 365 days, what happens to the objects?

A.Objects are permanently deleted.
B.Objects are transitioned to Glacier Deep Archive.
C.Objects are transitioned to Glacier.
D.Objects are transitioned to Standard-IA.
AnswerA

The lifecycle policy's expiration action permanently deletes objects once they reach 365 days; S3 lifecycle expiration removes the current version outright rather than transitioning it. Because the logs/ prefix falls under the rule, those audit logs are irrecoverably deleted unless versioning or Object Lock preserves them.

Why this answer

The lifecycle policy shown has a single rule that expires objects in the 'logs/' prefix after 365 days. In Amazon S3, an 'expiration' action permanently deletes the objects once the specified number of days has passed since object creation. There is no transition action configured, so objects are not moved to any storage class; they are simply deleted.

Exam trap

The DEA-C01 exam often tests the distinction between 'expiration' (permanent deletion) and 'transition' (moving to another storage class), and candidates mistakenly assume that expiration implies a transition to a cold storage class like Glacier.

How to eliminate wrong answers

Option B is wrong because the policy does not include a transition action to Glacier Deep Archive; expiration deletes objects, not transitions them. Option C is wrong because there is no transition rule to Glacier; the policy only specifies expiration. Option D is wrong because Standard-IA is a transition target, but the policy lacks any transition action and only has an expiration action.

1134
Multi-Selecteasy

Which TWO statements about Amazon Redshift data distribution are correct? (Choose two.)

Select 2 answers
A.DISTSTYLE is a distribution style option
B.AUTO distribution always chooses EVEN
C.KEY distribution places rows with the same distribution key on the same slice
D.EVEN distribution distributes rows across slices evenly
E.ALL distribution distributes data across all slices
AnswersC, D

KEY distribution colocates data by key.

Why this answer

In Amazon Redshift, KEY distribution places all rows with the same distribution key value on the same slice (compute node segment). This ensures that join operations on the distribution key are collocated, reducing data movement across the network and improving query performance.

Exam trap

The trap here is confusing distribution styles with distribution options (e.g., DISTSTYLE is a parameter, not a style) and misunderstanding that ALL distribution replicates the entire table to every node, not slices, while AUTO dynamically selects the best style rather than defaulting to EVEN.

1135
Multi-Selectmedium

A company's Amazon Redshift cluster is running slowly. The data engineer suspects that table design is the cause. Which TWO design practices can improve query performance? (Choose TWO.)

Select 2 answers
A.Define appropriate sort keys on frequently filtered columns.
B.Use GROUP BY instead of DISTINCT in queries.
C.Define appropriate distribution keys to collocate joins.
D.Increase the number of slices per node by resizing the cluster.
E.Use VARCHAR instead of CHAR for fixed-length strings.
AnswersA, C

Sort keys physically order rows on disk by the chosen columns, so range-restricted and equality filters on frequently filtered columns skip whole blocks via zone maps. This directly addresses the slow-query symptom attributed to table design, unlike compression or vacuum, which affect storage and maintenance rather than filter pruning.

Why this answer

Option A is correct because defining an appropriate sort key on columns that are frequently used in WHERE filters (especially range and equality predicates) lets Redshift use zone maps to skip scanning irrelevant blocks, dramatically reducing I/O and improving query performance. Option C is correct because choosing an appropriate distribution key (for example, DISTRIBUTING BY the join column) collocates matching rows on the same slice, enabling collocated joins that avoid expensive data redistribution (broadcast or shuffle) across nodes during query execution. Option B is not a table design practice and, while DISTINCT and GROUP BY can differ in execution, it does not address the suspected table-design cause.

Option D is incorrect because the number of slices per node is determined by the node type and cannot be changed by resizing; resizing changes node count or type, not slices per node. Option E is incorrect because CHAR and VARCHAR have essentially equivalent performance in Redshift, and CHAR is not a cause of the reported slowness.

Exam trap

DEA-C01 often tests the misconception that resizing a cluster increases slices per node or that query syntax changes (GROUP BY vs. DISTINCT) are table design practices — the exam wants you to focus on sort keys and distribution keys as the core performance levers.

1136
MCQmedium

A company stores application logs in Amazon S3 in JSON format. The logs are partitioned by year/month/day. A data engineer needs to create a table in the AWS Glue Data Catalog so that Amazon Athena can query the logs efficiently. The engineer wants to minimize query costs and ensure that new partitions are automatically recognized. Which combination of actions should the engineer take?

A.Create a table with partitions and configure an AWS Glue crawler to run on a schedule to discover new partitions.
B.Create a table with partitions defined manually and run ALTER TABLE ADD PARTITION for each new day.
C.Create a table with partition projection enabled for year/month/day and set the storage location to the S3 prefix.
D.Create a table without partitions and rely on Athena to scan the entire S3 prefix.
AnswerC

Partition projection in Athena allows the table to automatically infer partitions based on a defined pattern, eliminating the need to manually add partitions or run crawlers. By specifying year/month/day projection, Athena can generate partition locations on the fly, reducing metadata overhead and ensuring new partitions are immediately queryable. This minimizes query costs by enabling partition pruning.

Why this answer

Partition projection in Athena allows the table to compute partition locations dynamically based on a defined pattern, so new partitions are recognized immediately without manual intervention or crawlers. This reduces metadata operations and enables partition pruning, which lowers query costs. Manual partition management, non-partitioned tables, or scheduled crawlers either add overhead or fail to provide immediate partition recognition.

Exam trap

The trap here is assuming that a scheduled AWS Glue crawler is required to discover new partitions, when partition projection can eliminate that need entirely for a known partition scheme.

1137
MCQhard

A data engineer is using AWS Glue to transform data from Amazon S3 and load it into Amazon Redshift. The job uses the write_dynamic_frame.from_jdbc_conf method. The engineer notices that the job is slow and sometimes fails due to connection timeouts. Which action should the engineer take to improve performance and reliability?

A.Increase the number of DPUs for the Glue job to provide more parallelism.
B.Enable job bookmarks to avoid reprocessing previously loaded data.
C.Use the Amazon Redshift COPY command via a staging area in Amazon S3 instead of JDBC.
D.Increase the JDBC connection timeout and retry settings in the Glue job script.
AnswerC

The COPY command is the most efficient way to load large datasets into Redshift. It parallelizes the load and uses Redshift's massively parallel processing. Writing to S3 first and then using COPY reduces JDBC overhead and connection timeouts. This is the recommended best practice for bulk loads into Redshift from Glue.

Why this answer

For large data loads into Amazon Redshift, the COPY command is significantly faster and more reliable than JDBC. AWS Glue can write transformed data to Amazon S3 and then invoke the COPY command to load it into Redshift. This approach leverages Redshift's parallel processing, reduces connection overhead, and avoids timeouts.

Other options do not address the fundamental performance issue.

Exam trap

The trap here is assuming that increasing DPUs or tweaking timeouts will solve JDBC write performance issues, when the real solution is to use the Redshift COPY command.

1138
MCQmedium

A data engineer needs to ensure that all objects written to an S3 bucket are encrypted with SSE-KMS using a specific customer managed key, and that any upload without that encryption is rejected. The engineer has created the bucket and the KMS key. Which approach will enforce this requirement at the bucket level?

A.Enable S3 Block Public Access on the bucket.
B.Use an S3 Lifecycle rule to transition objects to S3 Glacier with encryption.
C.Add a bucket policy that denies s3:PutObject requests where the s3:x-amz-server-side-encryption header is not aws:kms or the s3:x-amz-server-side-encryption-aws-kms-key-id does not match the specified key.
D.Configure the bucket's default encryption to use SSE-KMS with the customer managed key.
AnswerC

A bucket policy with a Deny effect on s3:PutObject can evaluate conditions on request headers such as s3:x-amz-server-side-encryption and s3:x-amz-server-side-encryption-aws-kms-key-id. This explicitly blocks uploads that do not use the required encryption method and key. This is the only option that enforces rejection of non-compliant uploads at the bucket level, satisfying the requirement.

Why this answer

To enforce that every object is uploaded with a specific SSE-KMS key, a bucket policy must explicitly deny s3:PutObject requests that do not carry the required encryption headers. Default encryption only applies when no encryption is specified, and it cannot reject requests that specify a different method. The policy approach is the only one that guarantees non-compliant uploads are denied.

Exam trap

The trap here is assuming that setting default encryption on the bucket is sufficient to enforce a specific encryption method for all uploads.

1139
MCQeasy

A company uses AWS Glue ETL jobs to transform data stored in Amazon S3. The job reads data in Parquet format, applies transformations, and writes the output back to S3 in Parquet format. The team wants to improve the job's performance and reduce costs. Which action is MOST effective?

A.Change the input format from Parquet to CSV to simplify parsing.
B.Coalesce the input data into a single large file before processing.
C.Use column pruning and predicate pushdown to read only necessary columns and filter data early.
D.Increase the number of workers to maximum allowed.
AnswerC

Parquet is columnar, so column pruning and predicate pushdown let Glue read only required columns and filter rows at the storage layer, cutting bytes scanned. This directly reduces both shuffle volume and cost, the stated objectives.

Why this answer

Column pruning and predicate pushdown reduce the amount of data read from S3 by Spark-based AWS Glue ETL jobs. By reading only the necessary columns and filtering rows early in the scan, I/O and memory usage decrease, directly improving performance and reducing costs.

Exam trap

The trap here is that candidates often confuse 'coalesce' (reducing partitions) with 'repartition' (increasing parallelism) and assume fewer files always improve performance, ignoring that Glue ETL benefits from parallel reads across many small files when using columnar formats.

How to eliminate wrong answers

Option A is wrong because changing from Parquet to CSV would increase data size and parsing overhead, degrading performance and increasing costs. Option B is wrong because coalescing input into a single file eliminates parallelism, causing a single executor to process all data, which increases runtime and resource contention. Option D is wrong because increasing workers to the maximum allowed without addressing data skew or I/O bottlenecks can lead to excessive cost with diminishing returns, and may hit service limits or shuffle overhead.

1140
Multi-Selecteasy

A company must comply with a regulation that requires logging all access to sensitive data stored in Amazon S3. Which AWS services can be used to capture and store access logs? (Choose TWO.)

Select 2 answers
A.AWS Config
B.Amazon CloudWatch Logs
C.AWS CloudTrail
D.Amazon S3 server access logs
E.VPC Flow Logs
AnswersC, D

AWS CloudTrail records S3 API activity, including object-level data events when enabled, capturing who accessed sensitive objects and when. It satisfies the regulation by providing an auditable log of all access requests to the bucket.

Why this answer

AWS CloudTrail (C) is correct because it records S3 data events (GetObject, PutObject, DeleteObject) for object-level access to sensitive data, delivering an audit trail of every API call to the S3 bucket. Amazon S3 server access logs (D) are also correct because they capture detailed, object-level records of every request made to a bucket, including requester, operation, and response status, and can be stored in a target bucket for compliance. AWS Config (A) evaluates resource configuration and compliance over time but does not log individual data access requests.

Amazon CloudWatch Logs (B) stores and monitors log data but does not itself capture S3 access events. VPC Flow Logs (E) capture IP traffic metadata at the ENI/subnet level and cannot record S3 object-level access.

Exam trap

DEA-C01 often tests the confusion between configuration auditing (AWS Config), network traffic logging (VPC Flow Logs), and actual data access logging (CloudTrail data events and S3 server access logs), causing candidates to select services that do not capture S3 object access.

1141
MCQmedium

A data engineer is configuring an AWS Glue ETL job that reads from an Amazon S3 bucket and writes transformed records to an Amazon Redshift cluster. The job must load data in parallel and use an Amazon Redshift IAM role for authentication. Which connection option should the engineer configure in the Glue job to enable parallel loading and IAM-based authentication?

A.Use a JDBC connection with the Redshift cluster endpoint, database name, and a database user credential stored in AWS Secrets Manager.
B.Configure an AWS Glue connection of type JDBC with the Redshift JDBC URL and set the 'redshift_tmp_dir' parameter to an S3 path, using a database password stored in AWS Glue.
C.Create an AWS Glue connection of type Amazon Redshift, specify the cluster, database, and an IAM role with permissions to access both S3 and Redshift, and enable the 'Use Redshift COPY' option.
D.Use an AWS Glue connection of type Amazon S3 and configure the Redshift cluster endpoint in the job's script using boto3.
AnswerC

An AWS Glue connection of type Amazon Redshift with an attached IAM role allows Glue to use the Redshift COPY command with temporary S3 staging, which parallelizes data loading across Redshift slices. The IAM role provides authentication without embedding database credentials. Enabling the Redshift COPY option directs Glue to stage data in S3 and issue a COPY command, which is the documented way to achieve parallel, IAM-authenticated loads from Glue to Redshift.

Why this answer

The correct configuration uses an AWS Glue connection of type Amazon Redshift with an IAM role and the Redshift COPY option enabled. This lets Glue stage data in Amazon S3 and issue a COPY command, which parallelizes loading across Redshift slices and authenticates via IAM instead of database credentials. Other connection types may provide connectivity but do not combine parallel COPY-based loading with IAM role authentication in the way Glue's native Redshift integration does.

Exam trap

The trap here is assuming that any JDBC connection to Redshift enables parallel loading and IAM authentication, when only the native Amazon Redshift connection type with the COPY option provides both.

1142
MCQmedium

A data engineer manages an Amazon DynamoDB table used for a high-traffic gaming leaderboard. The table uses on-demand capacity mode and has a partition key of UserId (string) with no sort key. The leaderboard must retrieve the top 100 scores across all users. Currently, the engineer scans the entire table and sorts the results in application code, which takes several seconds and consumes large amounts of read capacity. What should the engineer do to improve the performance of retrieving the top scores?

A.Increase the table's read capacity by switching to provisioned mode with a high RCU value.
B.Add a local secondary index (LSI) on the Score attribute and query it with a limit of 100.
C.Enable DynamoDB Streams and use AWS Lambda to maintain a separate sorted list in Amazon ElastiCache for Redis.
D.Create a global secondary index (GSI) with a constant partition key and Score as the sort key, then query the index in descending order with a limit of 100.
AnswerD

A GSI with a constant partition key (e.g., 'leaderboard') and Score as sort key groups all items under one partition, allowing a Query with ScanIndexForward=false and Limit=100 to efficiently return the top scores. This avoids full table scans and leverages DynamoDB's sorted index for fast, low-cost retrieval.

Why this answer

The leaderboard requires a global, sorted view of scores. A GSI with a constant partition key and Score as sort key allows efficient Query operations that return the highest scores first using ScanIndexForward=false and a Limit. This design avoids full table scans and scales well.

Exam trap

The trap here is assuming that increasing read capacity or adding a local secondary index can solve a global sorting problem, when the core issue is the access pattern requiring a GSI with an appropriate sort key.

1143
MCQhard

A company uses Amazon DynamoDB for a gaming leaderboard. The table has a partition key of 'GameId' and a sort key of 'Score'. The application needs to query the top 10 scores for a given game. Which DynamoDB feature should be used for optimal performance?

A.Use a Query operation on the base table with ScanIndexForward set to false.
B.Enable DynamoDB Streams and use a Lambda function to compute the leaderboard.
C.Use DynamoDB Accelerator (DAX) to cache the results of a Scan operation.
D.Create a Global Secondary Index with the same partition key and sort key, then query with ScanIndexForward false.
AnswerA

A Query on the base table with ScanIndexForward false directly retrieves items in descending order by Score because the base table already has Score as sort key. This is the most efficient and cost-effective method.

Why this answer

The base table already has 'GameId' as partition key and 'Score' as sort key. A Query operation on the base table with ScanIndexForward set to false retrieves items in descending order by Score, so the first 10 items are the top scores. This is efficient without additional cost or complexity.

Option D is unnecessary and incurs extra storage cost.

Exam trap

Candidates often mistakenly think they need a GSI when the base table already has the optimal sort key. The trap is to assume that queries on the base table are inefficient, but DynamoDB can efficiently query with the sort key condition and ScanIndexForward.

How to eliminate wrong answers

Option A is wrong because a Query operation on the base table with `ScanIndexForward` set to false would require the sort key to be 'Score' (which it is), but the base table's sort key is 'Score' and the partition key is 'GameId', so a Query on the base table would work for retrieving top scores; however, the question asks for 'optimal performance' and the base table may have other attributes or be subject to throttling, but more critically, the base table's sort key is 'Score', so a Query with `ScanIndexForward` false is actually valid and efficient—this option is not incorrect in isolation, but the exam expects a GSI because the base table might have a different sort key or the question implies the need for a separate index for read-heavy workloads; however, the provided answer key marks D as correct, so A is considered wrong because the base table's sort key is 'Score', but the question's scenario likely intends that the base table has a different sort key or that a GSI is needed for optimal performance with large datasets, but technically A would work—this is a common exam trap where candidates overlook that the base table's sort key is already 'Score', making A a valid but less optimal choice due to potential hot partition issues or the need to isolate leaderboard reads. Option B is wrong because DynamoDB Streams with Lambda is an event-driven pattern for real-time processing, not for efficiently querying the top 10 scores on demand; it adds latency and complexity without improving query performance. Option C is wrong because DAX caches query results to reduce latency, but it does not eliminate the need for an efficient query pattern—using a Scan operation (even cached) is inherently inefficient for retrieving top scores, as it reads all items in the table rather than using an index to seek the highest values.

1144
MCQhard

Refer to the exhibit. A data engineer runs an AWS Glue ETL job that writes output to an S3 bucket. The job fails with the error shown. What is the most likely cause?

A.The IAM role used by the Glue job lacks the s3:PutObject permission for the output bucket
B.The Glue job attempted to write data in an unsupported format
C.The S3 bucket does not exist
D.The output file name contains invalid characters
AnswerA

Glue assumes the job's IAM role to write output, so a missing s3:PutObject permission on the target bucket produces an access-denied failure at the write stage. Granting that action to the role resolves it, satisfying the stem's constraint of the S3 write error shown in the exhibit.

Why this answer

The error shown in the exhibit indicates an access denied or permission failure when the AWS Glue ETL job attempts to write its output to the S3 bucket. The most likely cause is that the IAM role assigned to the Glue job does not include the s3:PutObject permission for the target bucket, which is required to upload objects. Without this permission, the job cannot complete the write operation, resulting in the failure.

Exam trap

The trap here is that candidates may focus on the data format or bucket existence, but the error message explicitly points to an access permission issue, which is a common misconfiguration in IAM roles for Glue jobs.

How to eliminate wrong answers

Option B is wrong because AWS Glue supports writing data in multiple formats (e.g., Parquet, ORC, JSON, CSV) and the error message does not indicate an unsupported format issue; such a problem would typically produce a different error related to format conversion. Option C is wrong because if the S3 bucket did not exist, the error would be a 'NoSuchBucket' or '404 Not Found' error, not an access denied error. Option D is wrong because while invalid characters in file names can cause errors, the error message shown is specifically about access permissions, not about invalid object key syntax.

1145
MCQeasy

A company uses AWS Glue to run ETL jobs that process data from an Amazon RDS for MySQL database and load it into an Amazon S3 data lake. The Glue job runs daily and processes incremental data. Recently, the job has been taking longer than expected. The engineer checks the CloudWatch logs and sees that the job is spending most of its time on the 'Reading from JDBC' phase. The MySQL table has 10 million rows and is indexed on the primary key. The Glue job uses a 'job bookmark' to track processed data. The engineer wants to improve the performance of the read phase. Which action is most likely to help?

A.Increase the JDBC 'fetchSize' parameter to 10000.
B.Disable job bookmark and perform a full refresh each time.
C.Increase the number of DPUs for the Glue job.
D.Modify the job to use a 'query' parameter that selects only the new or modified rows based on a timestamp column.
AnswerD

A timestamp-based query parameter pushes filtering to MySQL, so only new or modified rows cross JDBC rather than the full 10 million rows. This directly reduces time spent in the Reading from JDBC phase, satisfying the stem's constraint of incremental daily loads tracked by a job bookmark.

Why this answer

Using a 'query' parameter with a timestamp filter allows the Glue job to push down the filtering to the JDBC source, so only new or modified rows are read from MySQL. This drastically reduces the volume of data transferred and the time spent in the 'Reading from JDBC' phase, especially when combined with an index on the timestamp column.

Exam trap

The trap is thinking that adding more DPUs or increasing fetchSize will solve the problem; the real issue is reading too much data, so the fix is to reduce the data volume at the source with a filtered query.

How to eliminate wrong answers

Option A is wrong because increasing fetchSize may improve throughput slightly but does not reduce the total data read; it can also cause memory issues. Option B is wrong because disabling the job bookmark and doing a full refresh would read all 10 million rows every run, worsening performance. Option C is wrong because adding DPUs increases compute capacity but does not address the bottleneck of reading unnecessary data from the source.

1146
MCQeasy

A company uses Amazon QuickSight for data visualization. The data engineer needs to ensure that users can only see data relevant to their department. The data is stored in Amazon S3 and is accessed via SPICE. The engineer has created datasets in QuickSight and wants to implement row-level security (RLS). The dataset contains a column 'Department' that indicates which department a row belongs to. The engineer has configured RLS rules using a separate permissions dataset. However, users report that they can see all rows, not just their department's rows. What is the most likely reason?

A.The RLS permissions dataset is not correctly configured to map users to department values.
B.The 'Department' column is not included in the dataset.
C.The users have been granted admin access to the QuickSight dashboard.
D.The SPICE dataset does not support row-level security.
AnswerA

The permissions dataset must map each user or group to specific Department values; if its columns or entries are malformed, QuickSight applies no matching rule and defaults to showing all rows. Correctly mapping users to department values satisfies the stem's requirement that each user sees only their own department's data.

Why this answer

RLS in QuickSight works by matching the user's identity (via username or group) against rules in a permissions dataset that maps users/groups to allowed values of a column (e.g., Department). If users see all rows, the most likely cause is that the permissions dataset is not correctly mapping users to department values — for example, wrong usernames, missing group mappings, or a dataset that does not join properly to the main dataset.

Exam trap

DEA-C01 often tests the assumption that SPICE does not support RLS or that a missing column is the cause — the real trap is recognizing that RLS failures usually stem from identity-mapping mismatches in the permissions dataset, not from SPICE limitations.

How to eliminate wrong answers

Option B is wrong because if the 'Department' column were missing from the dataset, RLS rules referencing it would fail to apply or error out, but the symptom described (users see all rows) points to a mapping/configuration issue rather than a missing column, which would typically break the rule setup entirely. Option C is wrong because admin access to the dashboard would grant broad permissions, but the question states RLS rules were configured and users still see all rows — the issue is the RLS configuration itself, not dashboard-level access. Option D is wrong because SPICE datasets fully support row-level security; RLS is applied at query time regardless of whether data is in SPICE or direct query.

1147
MCQhard

Refer to the exhibit. An AWS Glue job is failing with 'AccessDenied' when trying to write to the 'data-lake-bucket' which is encrypted with an AWS KMS key. The IAM role used by the Glue job has the attached policy shown. What is the MOST likely cause of the failure?

A.The policy does not include s3:ListBucket permission.
B.The policy does not include s3:GetObject permission.
C.The KMS key ARN in the policy is incorrect.
D.The policy does not include kms:GenerateDataKey or kms:Encrypt permission.
AnswerD

Writing to a KMS-encrypted S3 bucket requires kms:GenerateDataKey and kms:Encrypt on the key. The attached policy grants only S3 actions, so the Glue role cannot obtain a data key, producing AccessDenied despite valid bucket permissions.

Why this answer

When an S3 bucket is encrypted with a customer-managed KMS key, any principal writing objects must have both S3 write permissions and KMS permissions to use the key. The Glue job's IAM policy lacks kms:GenerateDataKey and kms:Encrypt, so S3 rejects the PutObject with AccessDenied. Adding those KMS actions to the policy resolves the failure.

Exam trap

DEA-C01 often tests the misconception that S3 permissions alone are sufficient for encrypted buckets — candidates forget that SSE-KMS requires explicit kms:GenerateDataKey and kms:Encrypt permissions on the key.

How to eliminate wrong answers

Option A is wrong because s3:ListBucket is a read/list permission and is not required to write an object; its absence would cause a different error (e.g., 403 on ListObjects), not a write AccessDenied. Option B is wrong because s3:GetObject is for reading objects, not writing, so it is irrelevant to a PutObject failure. Option C is wrong because an incorrect KMS key ARN would typically produce a KMS-specific error like KMS.AccessDeniedException or InvalidArn, and the scenario implies the policy is otherwise correct — the missing actions are the real gap.

1148
MCQmedium

A company uses Amazon RDS for MySQL with Multi-AZ deployment. The primary instance fails, and automatic failover occurs. After failover, the application experiences higher latency. What is the most likely cause?

A.The read replica is now the primary and cannot handle write traffic.
B.The failover process disabled automatic backups.
C.The DNS endpoint did not update to point to the new primary.
D.The new primary instance is in a different Availability Zone, increasing network latency.
AnswerD

Cross-AZ latency can be higher than same-AZ.

Why this answer

After a Multi-AZ failover in Amazon RDS for MySQL, the new primary instance is launched in a different Availability Zone (AZ) than the original primary. If the application's compute resources (e.g., EC2 instances) remain in the original AZ, cross-AZ network traffic incurs additional latency due to the physical distance and the need to traverse the AZ boundary, which typically adds 1–2 ms of round-trip time. This increased network latency directly impacts application performance, especially for latency-sensitive queries.

Exam trap

The trap here is that candidates often assume the DNS endpoint fails to update (Option C) or that the standby cannot handle writes (Option A), but AWS explicitly ensures both are handled correctly, and the real issue is the unavoidable cross-AZ network latency introduced by the new primary's location.

How to eliminate wrong answers

Option A is wrong because in a Multi-AZ deployment, there is no read replica; the standby instance is a synchronous replica that is promoted to primary during failover, and it is fully capable of handling write traffic. Option B is wrong because the failover process does not disable automatic backups; automated backups continue to run on the new primary instance based on the same backup window and retention policy. Option C is wrong because the DNS endpoint (CNAME) for the RDS instance automatically updates to point to the new primary within 60–120 seconds after failover, so the application's connection string remains valid without manual intervention.

1149
MCQhard

A data engineer is using AWS Glue Studio to build a job that reads from Amazon S3, applies a filter, and writes to Amazon Redshift. The engineer needs to ensure the job can be rerun safely without creating duplicate rows in Redshift if a previous run partially succeeded. Which design choice best meets this requirement?

A.Enable AWS Glue job bookmarks and rely on them to prevent duplicate writes to Redshift on rerun.
B.Configure the Redshift write to use a pre-action and post-action that truncate the target table before each load.
C.Set the Glue job's write mode to append and add a unique constraint on the Redshift target table to reject duplicates.
D.Write the output to a staging table in Redshift and use a MERGE or upsert operation keyed on a business key to apply changes idempotently.
AnswerD

Staging plus MERGE keyed on a business key makes the load idempotent: rerunning the job re-applies the same changes without creating duplicates. This is a standard pattern for safe reruns in Redshift, and it isolates failures to the staging step before the final merge.

Why this answer

Loading into a staging table and then merging into the target on a business key is the canonical idempotent pattern for Redshift. Reruns re-execute the same merge deterministically, so partial failures do not leave duplicates. This decouples the risky load from the final apply step and supports safe retries.

Exam trap

The trap here is assuming that Glue job bookmarks or Redshift unique constraints guarantee idempotent writes, when only an explicit merge on a business key does.

1150
MCQeasy

A company wants to ingest streaming data from thousands of IoT devices into AWS for real-time analytics. Which AWS service is best suited for this purpose?

A.Amazon S3
B.AWS Lambda
C.Amazon RDS
D.Amazon Kinesis Data Streams
AnswerD

Amazon Kinesis Data Streams ingests high-throughput streaming data from thousands of concurrent producers with low latency, supporting real-time analytics downstream. Its shard-based scaling handles the device volume, unlike batch-oriented services such as AWS Glue or S3-based ingestion.

Why this answer

Amazon Kinesis Data Streams is purpose-built for ingesting and processing streaming data at scale from thousands of sources. It can capture and store terabytes of data per hour from IoT devices, enabling real-time analytics with millisecond latencies. The service provides durable, ordered data streams that can be consumed by multiple applications simultaneously.

Exam trap

The trap here is that candidates often confuse batch-oriented services like S3 or compute services like Lambda with the dedicated streaming ingestion layer required for real-time data, overlooking that Kinesis Data Streams provides the necessary buffering, ordering, and replay capabilities.

How to eliminate wrong answers

Option A is wrong because Amazon S3 is an object storage service designed for static data, not for real-time streaming ingestion; it lacks the low-latency, ordered delivery and concurrent consumer support required for streaming IoT data. Option B is wrong because AWS Lambda is a serverless compute service that can process events but is not designed as a primary ingestion buffer for high-throughput streaming data; it has a maximum invocation duration of 15 minutes and cannot natively store or replay streaming data. Option C is wrong because Amazon RDS is a relational database service for transactional workloads and structured queries, not for high-velocity, unbounded streaming data ingestion; it would create bottlenecks and cannot handle the throughput and ordering requirements of thousands of IoT devices.

1151
Drag & Dropmedium

Order the steps to set up a Kinesis Data Analytics application for real-time stream processing.

Drag or tap steps into the slots.

Steps
Order
1Step 1
2Step 2
3Step 3
4Step 4

Why this order

First, set up the source stream. Then create the analytics application, configure it with the source and logic, start it, and finally monitor performance.

1152
MCQeasy

A data engineer needs to ingest data from a relational database (MySQL) into Amazon S3 for analytics. The database is 500 GB and the job must run daily with incremental updates. Which AWS service is BEST suited for this task?

A.Amazon EMR with Apache Sqoop.
B.Amazon Kinesis Data Firehose with a database source.
C.AWS Database Migration Service (DMS) with a replication task.
D.AWS Glue ETL job with a JDBC connection.
AnswerC

AWS DMS with a replication task performs ongoing change data capture from MySQL to Amazon S3, handling the 500 GB initial load plus daily incremental updates. This satisfies the incremental requirement, whereas one-off tools like Snowball or Glue crawlers cannot continuously replicate ongoing database changes.

Why this answer

AWS DMS with a replication task is the best choice because it is specifically designed for continuous, incremental data replication from relational databases like MySQL to Amazon S3. DMS supports ongoing replication (change data capture) to capture incremental changes without custom scripting, and it can handle the initial 500 GB load efficiently. Other services either lack native incremental support or require additional configuration for this use case.

Exam trap

The trap here is that candidates often choose AWS Glue ETL (Option D) because it is familiar for data transformation, but they overlook that Glue lacks native incremental replication from databases, whereas DMS is purpose-built for this exact scenario with minimal overhead.

How to eliminate wrong answers

Option A is wrong because Amazon EMR with Apache Sqoop is a batch-oriented tool that does not natively support continuous incremental updates; it requires manual scripting for change data capture and is more complex to manage for daily incremental runs. Option B is wrong because Amazon Kinesis Data Firehose does not natively connect to a relational database as a source; it ingests streaming data from producers like Kinesis Data Streams or SDK, not directly from MySQL. Option D is wrong because AWS Glue ETL with a JDBC connection is designed for batch ETL jobs and does not have built-in change data capture for incremental updates; it would require custom logic to track changes, making it less suitable for daily incremental ingestion.

1153
MCQeasy

A company needs to store relational data that requires complex joins and transactional consistency. The workload is predictable and the data size is less than 500 GB. Which AWS service is MOST cost-effective for this use case?

A.Amazon Redshift
B.Amazon S3
C.Amazon RDS for PostgreSQL
D.Amazon DynamoDB
AnswerC

Amazon RDS for PostgreSQL provides full relational joins and ACID transactional consistency, which the workload demands. For predictable traffic under 500 GB, a single provisioned instance avoids the cost of distributed engines such as Aurora or DynamoDB, making it most cost-effective.

Why this answer

Amazon RDS for PostgreSQL is the most cost-effective choice because it provides a fully managed relational database service that supports complex joins and transactional consistency (ACID compliance) for predictable workloads under 500 GB. Unlike Redshift, which is optimized for petabyte-scale analytics, RDS offers lower cost for this data size and workload pattern, while DynamoDB lacks native SQL join capabilities and S3 is not a relational database.

Exam trap

The trap here is that candidates often choose Amazon Redshift for any data that involves joins or analytics, ignoring that it is cost-prohibitive and architecturally mismatched for transactional, sub-500 GB workloads, while RDS is the correct relational database service for this scale.

How to eliminate wrong answers

Option A is wrong because Amazon Redshift is a columnar data warehouse designed for large-scale analytical queries (petabytes), not for transactional workloads requiring complex joins and ACID compliance; it is over-provisioned and cost-inefficient for sub-500 GB data. Option B is wrong because Amazon S3 is an object store that does not support relational queries, joins, or transactional consistency natively; it lacks a SQL engine and ACID guarantees. Option D is wrong because Amazon DynamoDB is a NoSQL key-value and document database that does not support complex joins or relational operations; it is optimized for high-throughput, low-latency access patterns, not transactional consistency across multiple tables.

1154
MCQeasy

A data engineer must give an AWS Glue ETL job temporary access to data in an Amazon S3 bucket without creating long-term IAM user access keys. The job runs on a schedule and must retrieve credentials automatically. Which mechanism should the engineer use?

A.Attach an IAM role to the AWS Glue job and let the service assume it to obtain temporary credentials automatically.
B.Store an IAM user's access key and secret key in AWS Secrets Manager and have the Glue job retrieve them at runtime.
C.Generate a pre-signed URL for each S3 object and pass the URLs as job parameters to the Glue script.
D.Create an IAM user with programmatic access and embed the access key in the Glue job script as a hard-coded variable.
AnswerA

AWS Glue jobs run with an IAM role that the service assumes on your behalf, and the AWS SDK and Glue runtime retrieve temporary credentials automatically. No access keys are created or stored, and permissions are governed by the role's policies. This is the standard, secure way to grant a Glue job access to S3.

Why this answer

AWS Glue jobs are associated with an IAM role that the service assumes to call other AWS services. The Glue runtime and AWS SDK automatically obtain temporary credentials from that role, so the job can read S3 data without any stored access keys. This satisfies the no-long-term-credentials requirement and follows AWS best practice for service-to-service authorization.

Exam trap

The trap here is treating Secrets Manager or pre-signed URLs as the default way to give compute services credentials, when AWS compute services such as Glue should use an attached IAM role for automatic temporary credentials.

1155
MCQeasy

A company runs a daily batch processing job on Amazon EMR that reads data from Amazon S3 and writes results back to S3. The job takes longer than expected. The engineer wants to monitor the job's resource utilization. Which AWS service should be used to collect and visualize metrics such as CPU and memory usage of the EMR cluster's nodes?

A.AWS Config to record configuration changes in the EMR cluster.
B.Amazon Athena to query EMR job logs stored in S3.
C.Amazon CloudWatch with the CloudWatch Agent installed on the EMR nodes.
D.AWS CloudTrail to log API calls made by the EMR job.
AnswerC

CloudWatch alone reports EMR node metrics but not memory, which the hypervisor does not expose. Installing the CloudWatch Agent on the nodes publishes memory and disk utilisation as custom metrics, enabling the CPU and memory visualisation the engineer requires.

Why this answer

Amazon CloudWatch is the AWS service for monitoring and visualizing metrics. While EMR automatically sends some metrics (like CPU utilization) to CloudWatch, memory usage is not a default metric. To collect memory metrics (and other OS-level metrics) from EMR nodes, the CloudWatch Agent must be installed on each node.

This agent collects the desired metrics and publishes them to CloudWatch, where they can be graphed and alarmed. Thus, CloudWatch with the CloudWatch Agent is the correct solution for monitoring CPU and memory usage of EMR cluster nodes.

Exam trap

DEA-C01 often tests the misconception that EMR automatically sends memory metrics to CloudWatch, when in fact the CloudWatch Agent must be installed to collect memory and disk metrics.

How to eliminate wrong answers

Option A is wrong because AWS Config records configuration changes and evaluates compliance, not performance metrics like CPU or memory. Option B is wrong because Amazon Athena is a query service for data in S3, not a monitoring tool; it cannot collect or visualize resource utilization metrics. Option D is wrong because AWS CloudTrail logs API activity for auditing, not resource utilization metrics.

1156
MCQeasy

A data engineer is using AWS Step Functions to orchestrate a data pipeline that includes an AWS Glue job, an Amazon EMR step, and an Amazon Redshift stored procedure. The engineer needs to ensure that if the Glue job fails, the pipeline stops and does not proceed to the EMR step. Which Step Functions state type should be used to handle the error and stop the execution?

A.A Wait state that delays the next step until the Glue job is manually restarted.
B.A Fail state that transitions to a terminal state and stops the execution.
C.A Parallel state that runs the EMR step concurrently with the Glue job to reduce overall time.
D.A Choice state that evaluates the output of the Glue job and branches to an end state if it fails.
AnswerB

The Fail state in AWS Step Functions stops the execution and marks it as failed. It is specifically designed for error handling, allowing the pipeline to halt when a preceding task like the Glue job fails. This ensures no subsequent steps, such as the EMR step, are executed.

Why this answer

In AWS Step Functions, the Fail state is used to stop an execution and mark it as failed. When a Glue job fails, a Catch block can transition to a Fail state, ensuring the pipeline terminates without executing the EMR step. This is the standard way to handle errors and prevent downstream processing.

Exam trap

The trap here is confusing error handling with conditional branching; a Choice state can route based on conditions but does not inherently fail the execution, while a Fail state explicitly stops and marks failure.

1157
MCQhard

A company uses Amazon Redshift for its data warehouse. During a routine audit, the data engineer discovers that some queries are returning stale data even though the underlying source data has been updated. The engineer confirms that the COPY command completes successfully and that no errors are reported. Which action should the engineer take to ensure queries reflect the latest data?

A.Run the VACUUM command on the source tables.
B.Clear the Redshift result cache by running RESET ALL.
C.Run the ANALYZE command on the source tables.
D.Refresh the materialized views that the queries are using.
AnswerD

Materialized views store precomputed results and do not automatically reflect base-table changes, so queries reading them return stale data even after COPY succeeds. Refreshing the views re-executes their defining queries against updated source tables, satisfying the requirement for current results.

Why this answer

Materialized views in Amazon Redshift are not automatically refreshed when the underlying base tables are updated via COPY. To ensure queries return the latest data, the materialized views must be manually refreshed using the REFRESH MATERIALIZED VIEW command. Option A (VACUUM) reclaims disk space and re-sorts rows but does not update materialized views.

Option B (RESET ALL) clears the query result cache, but stale materialized views persist even if the cache is cleared. Option C (ANALYZE) updates table statistics for query planning, not the actual data content. Therefore, refreshing the materialized views is the necessary action.

1158
MCQhard

A data engineer runs an AWS Glue job that reads a large partitioned Parquet dataset from Amazon S3 and writes aggregated results to another S3 prefix. The job runs daily and currently reprocesses the entire dataset each time, which is becoming expensive. The engineer wants subsequent runs to process only data added since the last successful run, based on the job's state. Which feature should the engineer enable?

A.AWS Glue job bookmarks
B.AWS Glue job retry with a smaller max retry count
C.AWS Glue Data Catalog partition indexes
D.Amazon S3 Inventory reports on the source prefix
AnswerA

Job bookmarks persist state about which S3 objects or partitions have already been processed during prior runs. On subsequent runs, Glue reads only new or changed data, which reduces runtime and cost. This directly satisfies the requirement to process only data added since the last successful run, using the job's own state tracking.

Why this answer

AWS Glue job bookmarks maintain state across runs, recording which S3 objects, partitions, or JDBC rows were processed. When enabled, a subsequent run reads only new or modified data, avoiding full reprocessing. This is the built-in mechanism for incremental Glue jobs, unlike partition indexes, S3 Inventory, or retry settings, which do not track processed data across runs.

Exam trap

The trap here is confusing catalog performance features such as partition indexes with state tracking, when only job bookmarks persist which data a Glue job has already processed between runs.

1159
MCQeasy

A company is migrating an on-premises MongoDB database to Amazon DocumentDB. The data engineer needs to ensure minimal downtime during migration. Which AWS service should be used to facilitate the migration?

A.AWS Glue
B.AWS Snowball
C.AWS Database Migration Service (DMS)
D.Amazon S3 Transfer Acceleration
AnswerC

AWS Database Migration Service replicates ongoing changes from the on-premises MongoDB source to Amazon DocumentDB using change data capture, enabling near-zero-downtime cutover. Continuous replication keeps the target synchronised until you switch applications over, directly satisfying the stem's minimal-downtime constraint.

Why this answer

AWS Database Migration Service (DMS) supports continuous replication from MongoDB to Amazon DocumentDB using change data capture (CDC), enabling near-zero downtime migration. DMS can perform a full load followed by ongoing replication to keep the target synchronized until cutover.

Exam trap

The trap here is that candidates may confuse AWS Glue's ETL capabilities with database migration, overlooking that DMS is the only service purpose-built for live database migrations with minimal downtime and CDC support.

How to eliminate wrong answers

Option A is wrong because AWS Glue is a serverless data integration service for ETL (Extract, Transform, Load) jobs, not designed for live database migration with minimal downtime. Option B is wrong because AWS Snowball is a physical data transfer device for moving large volumes of data offline, which introduces significant downtime and is unsuitable for a live migration requiring minimal interruption. Option D is wrong because Amazon S3 Transfer Acceleration is a feature that speeds up uploads to S3 over the internet, not a database migration tool and cannot handle schema conversion or ongoing replication.

1160
MCQmedium

A company is ingesting streaming data from IoT devices into Amazon Kinesis Data Streams. The data must be transformed in real-time using custom Python code before being stored in Amazon S3. Which AWS service should be used to perform this transformation?

A.Amazon EMR with Spark Streaming
B.AWS Lambda function triggered by Kinesis Data Streams
C.Kinesis Data Analytics for Apache Flink
D.Kinesis Data Firehose with custom data transformation
AnswerC

Kinesis Data Analytics for Apache Flink is the correct choice because it allows running custom Apache Flink applications that support custom Python code via the Apache Flink Python API, enabling real-time data transformation.

Why this answer

Amazon EMR with Spark Streaming, which is optimized for large-scale batch and stream processing but is not the simplest or most direct service for this specific use case. Option B is AWS Lambda, which can be used for simple transformations but has limitations on execution time and complexity. Option D is Kinesis Data Firehose with custom data transformation, which supports only built-in transformations or Lambda functions, not arbitrary custom Python code directly.

Option C, Kinesis Data Analytics for Apache Flink, is correct because it allows running custom Apache Flink applications, which support custom Python code via the Apache Flink Python API, for real-time data transformation.

1161
Multi-Selecthard

A company needs to ingest data from multiple SaaS applications (Salesforce, Marketo) into Amazon S3 for analytics. The data volume is moderate (~100 GB per day). The pipeline must handle schema changes, deduplicate records, and provide low latency (under 1 hour). Which THREE services should be used? (Choose THREE.)

Select 3 answers
A.Amazon AppFlow
B.Amazon EventBridge
C.Amazon Kinesis Data Streams
D.AWS Glue DataBrew
E.AWS Database Migration Service (DMS)
AnswersA, B, D

AppFlow can ingest data from SaaS applications like Salesforce and Marketo.

Why this answer

Amazon AppFlow is the correct choice because it is a fully managed integration service specifically designed to transfer data from SaaS applications like Salesforce and Marketo to AWS services such as Amazon S3. It supports incremental transfers, handles schema changes automatically via its schema evolution feature, and can achieve sub-hour latency for moderate data volumes (~100 GB/day) without custom coding.

Exam trap

The trap here is that candidates often confuse Amazon Kinesis Data Streams as a universal ingestion service, but it lacks native SaaS connectors and schema evolution handling, making it unsuitable for this specific use case compared to AppFlow.

1162
Multi-Selecthard

A company is migrating a legacy on-premises ETL pipeline to AWS. The pipeline processes daily batch files from an FTP server. The data must be transformed using complex business logic before being loaded into Amazon Redshift. Which THREE AWS services should be used for this migration?

Select 3 answers
A.Amazon Athena
B.Amazon Redshift
C.AWS Glue
D.Amazon Kinesis Data Streams
E.AWS Transfer Family
AnswersB, C, E

Amazon Redshift serves as the target data warehouse, satisfying the requirement to load transformed data. Its massively parallel processing architecture handles analytical queries over the daily batch volumes, and native integration with AWS Glue and Amazon S3 enables efficient bulk loads via COPY commands after transformation completes.

Why this answer

AWS Transfer Family (E) is correct because it provides a fully managed FTP/FTPS/SFTP endpoint that can land the legacy daily batch files directly into Amazon S3, replacing the on-premises FTP server without changing the source's protocol. AWS Glue (C) is correct because it is a serverless ETL service whose Spark-based jobs and Glue Data Catalog can implement the complex business logic transformations on the batch data before loading. Amazon Redshift (B) is correct because it is the target data warehouse where the transformed data is loaded for analytics, and Glue can write to it via the Redshift JDBC/ODBC connection or the Redshift data source.

Amazon Athena (A) is not appropriate because it is an interactive query service over S3, not an ETL engine for applying complex transformation logic. Amazon Kinesis Data Streams (D) is not appropriate because it is designed for real-time streaming ingestion, whereas this pipeline processes daily batch files.

Exam trap

The trap is confusing batch ETL with streaming (Kinesis) or picking Athena as a transformation service; candidates may also forget Transfer Family for FTP ingestion and choose only Glue and Redshift.

1163
MCQeasy

A data engineer is designing a data lake on Amazon S3. The data lake will store raw data, transformed data, and curated datasets. The engineer needs to ensure that raw data is immutable (never overwritten or deleted) and that only authorized users can access the transformed data. Which combination of S3 features should the engineer use?

A.Use S3 Lifecycle policies to archive raw data to S3 Glacier and set bucket policies for transformed data.
B.Enable S3 Versioning and use S3 Access Points for each prefix.
C.Enable default encryption with SSE-KMS and use S3 bucket policies to restrict access.
D.Enable S3 Object Lock in compliance mode on the raw data prefix and use bucket policies to restrict access to transformed data prefix.
AnswerD

S3 Object Lock in compliance mode enforces WORM protection, preventing any user, including the root account, from overwriting or deleting objects for the retention period — satisfying the immutability constraint on raw data. Bucket policies then restrict access to the transformed prefix, meeting the authorisation requirement.

Why this answer

S3 Object Lock in compliance mode enforces a write-once-read-many (WORM) model, preventing any user—including the root user—from overwriting or deleting raw data. Bucket policies then provide granular access control to restrict the transformed data prefix to authorized users only, meeting both immutability and access control requirements.

Exam trap

The DEA-C01 exam often tests the distinction between versioning (which preserves history but allows overwrites) and Object Lock (which enforces immutability), leading candidates to choose versioning when immutability is explicitly required.

How to eliminate wrong answers

Option A is wrong because S3 Lifecycle policies only automate data transitions and deletions; they do not prevent overwrites or deletions, so raw data would not be immutable. Option B is wrong because S3 Versioning preserves previous versions but does not prevent deletion or overwrite of the current version; users can still delete or overwrite objects, and Access Points alone do not enforce immutability. Option C is wrong because SSE-KMS provides encryption at rest but does not prevent data from being overwritten or deleted; bucket policies control access but do not enforce immutability.

1164
MCQeasy

A data engineer needs to ingest data from multiple SaaS applications (Salesforce, Marketo) into Amazon S3 for a data lake. The data volumes are moderate and the sync needs to be scheduled daily. Which AWS service is most appropriate for this task?

A.AWS Glue
B.Amazon AppFlow
C.AWS Database Migration Service (DMS)
D.Amazon Kinesis Data Firehose
AnswerB

Amazon AppFlow provides managed connectors for SaaS sources such as Salesforce and Marketo, with scheduled flows delivering data into Amazon S3. It handles moderate volumes and daily synchronisation without custom ingestion code, matching the stated constraints.

Why this answer

Amazon AppFlow is purpose-built for securely transferring data between SaaS applications (like Salesforce and Marketo) and AWS services (like S3). It supports scheduled, incremental data syncs with built-in connectors, making it the most appropriate choice for moderate-volume daily ingestion into a data lake.

Exam trap

The trap here is that candidates often confuse AWS Glue's ETL capabilities with direct SaaS ingestion, overlooking that Glue requires a custom connector or script to pull from APIs, whereas AppFlow provides native, managed connectors.

How to eliminate wrong answers

Option A is wrong because AWS Glue is an ETL service designed for batch data transformation and cataloging, not for direct ingestion from SaaS applications; it lacks native connectors for Salesforce or Marketo. Option C is wrong because AWS DMS is intended for migrating databases (e.g., Oracle, MySQL) to AWS, not for pulling data from SaaS APIs. Option D is wrong because Amazon Kinesis Data Firehose is optimized for streaming data ingestion (e.g., from IoT or logs) and does not provide native SaaS connectors or scheduled sync capabilities.

1165
Matchingmedium

Match each AWS data migration tool to its primary function.

Drag a concept onto its matching description — or click a concept then click the description.

Concepts
Matches

Migrate databases with minimal downtime

Physical device for large data transfer

Online data transfer between on-prem and AWS

Fast uploads over long distances

Combine data across sources into views

Why these pairings

AWS DMS migrates databases with minimal downtime; AWS Snowball transfers large datasets physically; AWS DataSync automates online data transfer. Common confusions include swapping tool functions or attributing schema conversion to DataSync instead of SCT.

1166
MCQhard

A data engineer is designing a data warehouse on Amazon Redshift. The workload includes many ad-hoc queries that filter on a high-cardinality column, such as customer_id, and join large dimension tables. The engineer wants to improve query performance by choosing an appropriate distribution style and sort key. Which combination should the engineer use?

A.Use KEY distribution on customer_id and set the sort key to customer_id.
B.Use AUTO distribution and set the sort key to the date column with a compound sort key.
C.Use EVEN distribution and set the sort key to the date column.
D.Use ALL distribution on the fact table and set the sort key to the join key of the largest dimension.
AnswerA

KEY distribution on customer_id colocates matching rows on the same slice, minimizing data movement during joins on that column. Setting the sort key to customer_id also enables efficient range-restricted scans and merge joins on that column. This directly addresses the high-cardinality filter and join performance, making it the best choice for this workload.

Why this answer

For a workload that frequently filters and joins on a high-cardinality column like customer_id, distributing the fact table by that key colocates matching rows and minimizes network traffic during joins. Using the same column as the sort key further speeds up range scans and merge joins. Together, they reduce data movement and I/O, improving ad-hoc query performance.

Exam trap

The trap here is assuming that sorting by date or using AUTO distribution will optimize all queries, when the key is to align distribution and sort keys with the most common join and filter columns.

1167
MCQhard

A data engineer is building an AWS Glue ETL job that reads from an Amazon S3 bucket containing nested JSON files and must flatten the nested structures before writing to Amazon Redshift. The job uses the Glue DynamicFrame API. The engineer wants the transformation to run as a single pass without an intermediate shuffle. Which operation should the engineer use?

A.ResolveChoice
B.ApplyMapping
C.Relationalize
D.DropNullFields
AnswerC

Relationalize flattens nested DynamicFrames into a set of related flat DynamicFrames in a single pass, producing one frame per nested array or struct. It is designed exactly for unnesting JSON before loading into a relational target such as Redshift, and it avoids a full shuffle because it processes the frame as a single traversal.

Why this answer

Relationalize is the Glue DynamicFrame transform specifically designed to unnest nested JSON into multiple flat frames in one traversal. It produces a root frame and one frame per nested array, each with a join key back to the parent. This matches the requirement to flatten nested structures before loading into Redshift without an intermediate shuffle.

Exam trap

The trap here is confusing ResolveChoice, which resolves type ambiguity, with Relationalize, which unnest nested structures; only Relationalize performs the flattening required by the scenario.

1168
Multi-Selectmedium

A data engineer is configuring an Amazon Redshift cluster for a workload that runs large nightly ELT jobs loading data from Amazon S3 and then executes complex analytical queries. The team wants to improve query performance and reduce the time the cluster spends on data loading. Which TWO configuration choices should the engineer make? (Choose two.)

Select 2 answers
A.Use the COPY command with a manifest file and parallelism to load data from multiple S3 objects.
B.Enable automatic vacuum and analyze only during the nightly load window.
C.Define distribution keys on the largest fact tables to co-locate join rows on the same slice.
D.Load data using single-row INSERT statements in a loop from the application.
E.Set the cluster's distribution style to ALL for every table to avoid joins.
AnswersA, C

COPY is the native, massively parallel load path in Redshift and automatically splits work across slices when multiple files or a manifest are provided. Using a manifest ensures a consistent, complete set of files is loaded, and parallelism exploits the cluster's MPP architecture to shorten load time. This directly addresses the goal of reducing time spent loading data.

Why this answer

Bulk loading with COPY and a manifest, combined with parallelism, uses Redshift's MPP architecture to shorten load time. Choosing an appropriate distribution key on large fact tables reduces inter-slice data movement during complex joins, improving query performance. Row-by-row inserts, restricting maintenance to the load window, and applying ALL distribution to every table all degrade performance or increase load time.

Exam trap

The trap here is treating maintenance commands like VACUUM and broad distribution styles like ALL as universal performance fixes, when they can lengthen loads and inflate storage.

1169
MCQmedium

A data engineer runs an AWS Glue for Apache Spark job that writes partitioned Parquet output to Amazon S3. The engineer notices thousands of small output files in each partition, which slows downstream Amazon Athena queries. The job reads from a large S3 source and uses default partitioning. Which change should the engineer make to reduce the number of small files?

A.Set the Glue job's output to a single partition using coalesce(1) before writing.
B.Enable Glue job bookmarks and increase the number of DPUs allocated to the job.
C.Convert the output to CSV so fewer files are created per partition.
D.Use repartition or a controlled partition count before writing, or enable Glue's grouping of output files.
AnswerD

Repartitioning to a sensible number of partitions, or using Glue's built-in grouping of output files, consolidates many small files into fewer larger ones while preserving parallelism. This reduces per-file overhead and improves Athena query performance by lowering the number of objects scanned and listed. It scales with data volume, unlike forcing a single partition, and directly targets the small-file cause.

Why this answer

The number of output files equals the number of partitions written, so consolidating partitions with repartition or using Glue's output file grouping reduces small files without sacrificing parallelism. Coalescing to one partition kills scalability, bookmarks and DPUs do not affect file count, and CSV does not change partitioning while harming query efficiency.

Exam trap

The trap here is assuming that changing the output format or adding compute reduces file count, when file count is governed by the number of written partitions.

1170
MCQmedium

A company is streaming IoT sensor data to Amazon Kinesis Data Streams. The data is JSON with a schema that changes occasionally. They want to load the data into Amazon S3 in Parquet format partitioned by date and sensor_id. Which approach is MOST cost-effective and operationally efficient?

A.Use Amazon EMR to read from Kinesis Data Streams and write to S3 in Parquet format.
B.Use a Lambda function to transform records to Parquet and write to S3.
C.Use a custom Kinesis Client Library application on EC2 to buffer and write Parquet files to S3.
D.Use Amazon Kinesis Data Firehose with a schema from AWS Glue Data Catalog to convert to Parquet and enable dynamic partitioning by date and sensor_id.
AnswerD

Amazon Kinesis Data Firehose can directly convert incoming JSON data to Parquet using a schema from AWS Glue Data Catalog, and supports dynamic partitioning by date and sensor_id without custom code. It is fully managed, making it the most cost-effective and operationally efficient.

Why this answer

Amazon Kinesis Data Firehose is the most cost-effective and operationally efficient solution because it natively supports converting JSON to Parquet using a schema from the AWS Glue Data Catalog and can dynamically partition data by date and sensor_id. Firehose handles buffering, compression, and delivery to S3 without requiring custom code or infrastructure management.

Exam trap

DEA-C01 often tests the misconception that custom code (Lambda, KCL) is needed for Parquet conversion and partitioning, but Firehose provides built-in features for these tasks.

How to eliminate wrong answers

Option A is wrong because Amazon EMR requires provisioning and managing clusters, which is operationally heavy and less cost-effective for a simple streaming ETL task. Option B is wrong because a Lambda function would need custom code to buffer, transform to Parquet, and manage partitioning, increasing complexity and potential errors. Option C is wrong because a custom Kinesis Client Library application on EC2 requires managing instances, writing code for buffering and Parquet conversion, and handling scaling, which is not operationally efficient.

1171
MCQmedium

A data engineer is using Amazon Athena to query data stored in Amazon S3. The data is partitioned by date, but the engineer notices that queries are scanning the entire bucket instead of only the relevant partitions. The table is defined in the AWS Glue Data Catalog. Which action should the engineer take to ensure that Athena only scans the necessary partitions?

A.Run the `MSCK REPAIR TABLE` command in Athena to update the partition metadata in the AWS Glue Data Catalog.
B.Increase the query timeout setting for the Athena workgroup to allow more time for scanning.
C.Convert the data to a columnar format such as Parquet and enable predicate pushdown.
D.Use AWS Glue crawlers to re-crawl the data and update the table schema.
AnswerA

The `MSCK REPAIR TABLE` command scans the S3 location for partition folders and adds any missing partition metadata to the AWS Glue Data Catalog. If partitions are not registered, Athena cannot prune them and will scan the entire table. Running this command ensures that the partition metadata is up to date, enabling partition pruning.

Why this answer

When Athena scans the entire bucket despite partitioning, it is often because the partition metadata in the AWS Glue Data Catalog is not up to date. Running `MSCK REPAIR TABLE` in Athena adds missing partitions to the catalog, allowing Athena to prune partitions and scan only the relevant data.

Exam trap

The trap here is assuming that converting to a columnar format or increasing timeout will solve the issue, when the root cause is missing partition metadata in the Data Catalog.

1172
MCQmedium

Refer to the exhibit. A data engineer is troubleshooting an AWS Lambda function that reads from an S3 bucket and writes to a Kinesis Data Stream. The Lambda function fails with an AccessDeniedException when calling the kinesis:PutRecords API. Which change is needed to the IAM policy?

A.Add s3:PutObject permission to the policy
B.Change the resource ARN for Kinesis to a wildcard
C.Change the resource ARN for Kinesis to include the correct stream name
D.Add kinesis:PutRecords permission to the policy
AnswerC

The Lambda execution role's policy must scope the Kinesis action to the exact stream ARN, since `kinesis:PutRecords` is evaluated against the stream resource, not the account or wildcard. A mismatched stream name in the resource ARN yields AccessDeniedException even when the action itself is permitted, satisfying the stem's requirement to correct the policy.

Why this answer

The Lambda function's IAM policy grants kinesis:PutRecords permission but the resource ARN does not match the target Kinesis Data Stream. The AccessDeniedException occurs because IAM evaluates the resource ARN against the stream's ARN and denies access on mismatch. The correct fix is to change the resource ARN to include the correct stream name (option C).

While using a wildcard (option B) would also resolve the error, it is not the recommended approach because it violates the principle of least privilege. Option A is incorrect because the error is from Kinesis, not S3. Option D is incorrect because the kinesis:PutRecords action is already present in the policy.

Exam trap

The DEA-C01 exam often tests the misconception that adding the missing API action (kinesis:PutRecords) is the fix, but here the action is already present and the error stems from an incorrect resource ARN. The exam expects the specific stream name to be used, aligning with AWS best practices for least privilege.

How to eliminate wrong answers

Option A is wrong because the error is from Kinesis, not S3; adding s3:PutObject does not address the kinesis:PutRecords AccessDeniedException. Option C is wrong because the policy already includes a specific stream name, but it is incorrect; changing to a wildcard is the fix, not specifying a different name. Option D is wrong because the policy already includes kinesis:PutRecords permission; the issue is the resource ARN restriction, not a missing action.

1173
MCQeasy

A data engineer is using Amazon Kinesis Data Streams to ingest real-time data. The engineer needs to ensure that records are delivered to consumers in the same order they were written and that each record is processed exactly once. Which combination of Kinesis Data Streams features should the engineer use?

A.Use partition keys to distribute records across shards and enable AWS Lambda as a consumer with checkpointing.
B.Use a single shard and implement idempotent processing in the consumer application.
C.Use multiple shards with sequence numbers and enable Kinesis Client Library (KCL) with deduplication.
D.Use a single shard and enable enhanced fan-out for consumers.
AnswerB

A single shard ensures that records are stored and delivered in order. To achieve exactly-once processing, the consumer must be idempotent, meaning it can safely process the same record multiple times without side effects. Kinesis Data Streams itself provides at-least-once delivery, so idempotency is necessary for exactly-once semantics.

Why this answer

To guarantee order, all records must go to a single shard, as Kinesis preserves order only within a shard. For exactly-once processing, the consumer must be idempotent because Kinesis Data Streams provides at-least-once delivery. Combining a single shard with idempotent processing ensures both requirements are met.

Exam trap

The trap here is assuming that Kinesis provides exactly-once processing natively or that multiple shards can maintain global order, when actually ordering is per shard and exactly-once requires consumer-side idempotency.

1174
MCQeasy

A company wants to ensure that all S3 buckets are encrypted using server-side encryption. Which AWS service can be used to automatically remediate non-compliant buckets?

A.AWS CloudTrail
B.Amazon Inspector
C.AWS Trusted Advisor
D.AWS Config
AnswerD

AWS Config rules continuously evaluate S3 bucket encryption against your desired configuration and can trigger automatic remediation, such as invoking a Systems Manager automation to enable default encryption. This satisfies the requirement to remediate non-compliant buckets automatically, rather than merely detecting or reporting drift.

Why this answer

AWS Config can use managed rules like s3-bucket-server-side-encryption-enabled to check compliance and trigger auto-remediation via SSM Automation or Lambda. Option D is correct.

1175
MCQeasy

A data engineer is setting up an Amazon DynamoDB table to store user session data. The table must handle sudden spikes in read traffic during peak hours, and the engineer wants to minimize operational overhead while ensuring consistent performance. The table's read capacity mode should automatically adjust to traffic changes. Which capacity mode should the engineer choose?

A.On-demand capacity mode.
B.Provisioned capacity mode with global secondary indexes.
C.Provisioned capacity mode with auto scaling enabled.
D.Provisioned capacity mode with reserved capacity.
AnswerA

On-demand capacity mode automatically scales read and write capacity to handle traffic changes without requiring any capacity planning. It instantly accommodates sudden spikes in traffic and charges per request, making it ideal for unpredictable workloads with minimal operational overhead. This directly meets the requirement for automatic adjustment and consistent performance.

Why this answer

On-demand capacity mode is designed for workloads with unpredictable traffic patterns. It automatically scales to handle sudden spikes in read and write traffic, eliminating the need for capacity planning and reducing operational overhead. This makes it the best choice for session data with variable demand.

Exam trap

The trap here is assuming that provisioned mode with auto scaling is equivalent to on-demand; auto scaling has a delay and requires configuration, while on-demand provides immediate, automatic scaling.

1176
MCQeasy

A data engineer is configuring AWS Glue jobs to access data stored in Amazon S3. The data is encrypted using server-side encryption with AWS KMS (SSE-KMS). The Glue job needs to read and write data to the S3 bucket. Which IAM policy statement should be added to the Glue job's IAM role to allow it to use the KMS key?

A.{"Effect":"Allow","Action":["kms:Decrypt"],"Resource":"*"}
B.{"Effect":"Allow","Action":["kms:Decrypt","kms:GenerateDataKey"],"Resource":"*"}
C.{"Effect":"Allow","Action":["kms:Decrypt","kms:ReEncrypt"],"Resource":"*"}
D.{"Effect":"Allow","Action":["kms:Decrypt","kms:Encrypt"],"Resource":"*"}
AnswerB

SSE-KMS requires both kms:Decrypt to read encrypted objects and kms:GenerateDataKey to obtain a data key for writing. Granting these two actions on the key resource satisfies the stem's read-and-write requirement, since Glue cannot access KMS-encrypted S3 data with either permission missing.

Why this answer

To read and write data encrypted with SSE-KMS, AWS Glue needs both `kms:Decrypt` (to read existing encrypted data) and `kms:GenerateDataKey` (to create a new data key for writing encrypted data). `kms:GenerateDataKey` is required because S3 uses a data key to encrypt objects, and the caller must generate that key via KMS. Option B correctly includes both actions, allowing the Glue job to perform read and write operations on the SSE-KMS encrypted bucket.

Exam trap

The trap here is that candidates often assume `kms:Encrypt` is needed for writing encrypted data, but S3 SSE-KMS actually requires `kms:GenerateDataKey` because the encryption is done with a derived data key, not by calling `kms:Encrypt` directly.

How to eliminate wrong answers

Option A is wrong because it only grants `kms:Decrypt`, which allows reading encrypted data but not writing new encrypted objects; writing requires `kms:GenerateDataKey` to create the encryption key. Option C is wrong because `kms:ReEncrypt` is used for re-encrypting data under a different KMS key, which is not needed for standard S3 read/write operations with SSE-KMS. Option D is wrong because `kms:Encrypt` is used to encrypt plaintext data directly with a KMS key, but S3 SSE-KMS requires `kms:GenerateDataKey` (not `kms:Encrypt`) to obtain a data key for object-level encryption.

1177
Multi-Selectmedium

A company wants to use AWS Glue to transform data stored in Amazon S3. The data is partitioned by date and includes both CSV and Parquet files. The transformation should be optimized for cost and performance. Which THREE actions should the data engineer take? (Choose THREE.)

Select 3 answers
A.Run a crawler to update the schema before each job run.
B.Use partition pruning by filtering on the date column in the ETL script.
C.Use job bookmarks to process only new data.
D.Increase the number of DPUs to the maximum allowed.
E.Convert all files to Parquet format before processing.
AnswersB, C, E

Filtering on the date partition column lets Glue read only relevant S3 prefixes rather than the full dataset. This reduces bytes scanned and DPU hours, directly meeting the cost and performance optimisation constraint in the stem.

Why this answer

Option B is correct because partition pruning by filtering on the date column lets AWS Glue read only the relevant S3 partitions instead of scanning the entire dataset, which directly reduces I/O, runtime, and cost. Option C is correct because Glue job bookmarks persist state between runs so the job processes only new or changed data since the last run, avoiding reprocessing of already-transformed partitions and lowering both execution time and DPU consumption. Option E is correct because converting CSV files to Parquet enables columnar storage and compression, so Glue reads less data and benefits from predicate pushdown, improving performance and reducing cost compared with row-based CSV.

Option A is not appropriate because running a crawler before every job run adds unnecessary cost and time, and the schema can be supplied directly or updated only when it changes. Option D is not appropriate because maximizing DPUs increases cost rather than optimizing it, and the required performance can be achieved through pruning, bookmarks, and columnar formats.

Exam trap

DEA-C01 often tests cost optimization by tempting candidates with 'more DPUs = faster' (Option D) or 'crawl every run' (Option A), when the correct answers are the data-reduction techniques — partition pruning, job bookmarks, and columnar format conversion.

1178
Multi-Selectmedium

Which TWO statements are true about Amazon Redshift distribution styles? (Choose TWO.)

Select 2 answers
A.KEY distribution is always the best choice to minimize data skew.
B.AUTO distribution always selects EVEN distribution.
C.ALL distribution copies the entire table to every node.
D.Redshift automatically assigns a ROUND ROBIN distribution style by default.
E.EVEN distribution distributes rows across slices in a round-robin fashion.
AnswersC, E

ALL distribution is useful for small tables that are frequently joined.

Why this answer

The ALL distribution style in Amazon Redshift copies the entire table to every node in the cluster. This is ideal for small, slowly changing dimension tables (like date or location tables) that need to be joined with large fact tables, as it eliminates the need to redistribute data across nodes during query execution.

Exam trap

The trap here is that candidates often confuse AUTO distribution with ROUND ROBIN, or assume that KEY distribution is always optimal, when in fact poor key selection can lead to severe data skew and performance degradation.

1179
MCQeasy

A media company is building a data pipeline to ingest user activity logs from multiple sources into Amazon S3. The logs are JSON files generated every minute. The company wants to use Amazon Athena to query the logs with minimal latency and cost. The current approach is to use AWS Kinesis Data Firehose to deliver the logs to S3 with a prefix like 'logs/2024/01/01/00/file.json'. However, when running Athena queries, the team notices high query costs because Athena scans all files in the 'logs/' prefix even when querying for a specific date. What should the team do to reduce the amount of data scanned by Athena?

A.Create an Athena view that filters by date.
B.Increase the number of partitions by using a more granular prefix like 'logs/2024/01/01/00/00/'.
C.Convert the JSON files to Apache Parquet format using AWS Glue ETL jobs.
D.Create a Hive-style partition structure in S3 with keys like 'year=2024/month=01/day=01/hour=00/' and update the Glue Data Catalog accordingly.
AnswerD

Hive-style partitioning splits data into year/month/day/hour prefixes, letting Athena prune irrelevant partitions via partition projection or the Glue Data Catalog. Queries for one date then scan only that partition rather than every file under logs/, cutting both cost and latency.

Why this answer

Athena reduces data scanned by using partition pruning, which requires a Hive-style partition structure in S3 (e.g., year=2024/month=01/day=01/hour=00/) registered in the Glue Data Catalog. When queries filter on partition columns, Athena only reads the relevant partitions instead of scanning the entire logs/ prefix. This directly addresses the high query cost caused by scanning all files.

Exam trap

DEA-C01 often tests the difference between S3 prefix organization and true Hive-style partitioning — candidates pick 'more granular prefix' thinking it enables pruning, but without Glue Data Catalog partition registration, Athena still scans all files.

How to eliminate wrong answers

Option A is wrong because an Athena view is just a saved query — it does not change how data is stored or pruned, so it still scans all files. Option B is wrong because a more granular prefix like logs/2024/01/01/00/00/ is not a Hive-style partition; without partition metadata in the Glue Data Catalog, Athena cannot prune and will still scan everything. Option C is wrong because converting to Parquet reduces data scanned per file (columnar compression) but does not eliminate scanning irrelevant date partitions; it's a complementary optimization, not the primary fix for date-based pruning.

1180
MCQmedium

A data engineer maintains an AWS Glue Data Catalog table for an S3-based dataset. After new files were added, queries in Amazon Athena fail with the error 'HIVE_BAD_DATA: Error parsing field value for field 3: For input string: "N/A"'. The column is defined as bigint in the Data Catalog but contains the string 'N/A' in some records. Which action should the data engineer take to allow Athena to query the data without changing the underlying files?

A.Add a partition projection configuration for the table.
B.Run MSCK REPAIR TABLE on the table to refresh partitions.
C.Create a new table with the column as bigint and use a SerDe that skips malformed rows.
D.Alter the table to set the column type to string in the AWS Glue Data Catalog.
AnswerD

Changing the column type to string prevents Athena from attempting numeric conversion on values like 'N/A', allowing the query to succeed without modifying the files. The data engineer can update the table schema through the AWS Glue Data Catalog, and Athena will then treat the field as text. This is the least invasive way to handle mixed or invalid data when the source files cannot be rewritten.

Why this answer

The error is caused by a mismatch between the declared bigint type and actual string values in the files. The most direct fix that avoids rewriting data is to change the column type to string in the AWS Glue Data Catalog. Athena will then read the field as text, and the query will no longer attempt numeric conversion.

Partition repair, projection, and SerDe tweaks do not address the type conversion failure.

Exam trap

The trap here is assuming that partition-level operations like MSCK REPAIR TABLE or partition projection can resolve column-level data type errors.

1181
MCQmedium

A company runs a nightly ETL job using AWS Glue. The job reads data from a JDBC connection to an on-premises MySQL database. The job fails with an error indicating that the connection pool is exhausted. What is the most likely cause and solution?

A.The database is not reachable due to network issues. Check VPC and security groups.
B.The Glue job is hitting the AWS Glue connection pool limit. Increase the Glue connection pool size.
C.The database credentials are expired. Rotate the password in AWS Secrets Manager.
D.The Glue job is using too many executors, exhausting the database connections. Reduce the number of DPUs or increase the database max connections.
AnswerD

Each Glue executor opens its own JDBC connections, so a high DPU count multiplies concurrent sessions against the on-premises MySQL instance until its connection limit is hit. Reducing DPUs lowers that parallelism, directly addressing the exhausted connection pool constraint; raising MySQL's max_connections is the complementary fix.

Why this answer

AWS Glue jobs distribute work across multiple executors, each of which opens its own JDBC connection to the source database. When the number of executors (controlled by DPUs) exceeds the database's configured maximum connections, the database connection pool is exhausted, causing the error. Reducing the number of DPUs or increasing the database's max_connections setting resolves the issue.

Exam trap

The trap here is that candidates confuse a database-side connection pool exhaustion with an AWS Glue service limit or network issue, leading them to incorrectly choose options about Glue connection pools or VPC configurations.

How to eliminate wrong answers

Option A is wrong because a network connectivity issue (e.g., VPC or security group misconfiguration) would typically result in a timeout or 'cannot connect' error, not a 'connection pool exhausted' error. Option B is wrong because AWS Glue does not have a configurable 'connection pool size' for JDBC connections; the pool exhaustion is on the database side, not Glue's internal pool. Option C is wrong because expired credentials would cause an authentication failure (e.g., 'Access denied for user'), not a connection pool exhaustion error.

1182
MCQmedium

A company uses AWS Glue ETL jobs to transform data in S3. The job runs successfully but takes longer than expected. The data is in Parquet format and partitioned by date. Which change would most improve performance without increasing cost?

A.Repartition the data by a different column.
B.Convert Parquet to CSV for faster serialization.
C.Increase the number of DPUs for the job.
D.Enable pushdown predicates to filter partitions early.
AnswerD

Pushdown predicates let AWS Glue push partition and column filters to the S3 data source, so only matching partitions and row groups are read. With date partitioning, this prunes irrelevant partitions early, cutting I/O and runtime without adding compute capacity or cost.

Why this answer

Pushdown predicates allow AWS Glue to filter data at the storage layer (e.g., S3 partition pruning) before reading it into memory. Since the data is partitioned by date, enabling pushdown predicates reduces the amount of data scanned, which directly decreases job runtime without requiring additional DPUs or changing the data format.

Exam trap

The trap here is that candidates often assume performance issues are solved by adding more resources (DPUs) or changing file formats, when the real bottleneck is reading unnecessary data due to lack of partition pruning.

How to eliminate wrong answers

Option A is wrong because repartitioning by a different column would likely increase shuffle overhead and may not align with the existing partition structure, potentially worsening performance. Option B is wrong because converting Parquet to CSV would increase data size and I/O due to CSV's lack of compression and columnar storage, making the job slower and more expensive. Option C is wrong because increasing DPUs would raise cost without addressing the root cause (scanning unnecessary partitions), and the question explicitly asks for a change that does not increase cost.

1183
MCQeasy

A media company stores video files in an Amazon S3 bucket. The bucket policy allows access only from a specific VPC. The company has enabled S3 Server Access Logs to monitor access. Recently, the security team found that some requests were coming from an IP address outside the allowed VPC. They suspect that the bucket policy may have an incorrect condition. What should they check first?

A.Verify that the bucket policy uses the 'aws:SourceVpc' condition key with the correct VPC ID.
B.Review the S3 Server Access Logs to identify the source IP addresses.
C.Ensure that the IAM role used by the application has the correct permissions.
D.Check if the bucket policy allows public access.
AnswerA

The aws:SourceVpc condition key only matches when the request arrives through a VPC endpoint. A typo or wrong VPC ID in that condition lets external IP addresses bypass the intended restriction, so verify it first.

Why this answer

The first thing to check is the bucket policy condition key. The 'aws:SourceVpc' condition key is used to restrict access to a specific VPC, but if it is misconfigured (e.g., wrong VPC ID or incorrect condition operator), requests from outside the VPC might be allowed. Verifying this key ensures the policy is correctly enforcing the VPC restriction.

Exam trap

DEA-C01 often tests the confusion between diagnosing and fixing. Candidates might choose to review logs (Option B) as a first step, but the question asks what to check first to address the suspected policy condition error, making the policy condition the priority.

How to eliminate wrong answers

Option B is wrong because reviewing S3 Server Access Logs can identify source IPs but does not directly address the policy misconfiguration; it's a diagnostic step, not a fix. Option C is wrong because IAM role permissions are separate from bucket policy conditions; even with correct IAM, a misconfigured bucket policy could allow access. Option D is wrong because checking for public access is not specific to the VPC condition issue; the policy might not be public but still have an incorrect VPC condition.

1184
MCQmedium

A data engineer is configuring an AWS Glue ETL job to read from an Amazon S3 bucket that contains Apache Parquet files partitioned by year, month, and day. The engineer wants the job to only process data for the year 2023 and month 10, and to minimize the amount of data scanned. The Glue job uses the Glue Data Catalog table `sales_data` with the correct partition structure. What is the MOST efficient way to configure the job to read only the required partitions?

A.Read the entire `sales_data` table into a DynamicFrame and then apply a `Filter` transform with the condition `year=='2023' and month=='10'`.
B.Use the Glue `create_dynamic_frame.from_options` with the S3 path `s3://bucket/sales_data/year=2023/month=10/` and format `parquet`.
C.Use the Glue `create_dynamic_frame.from_catalog` with `push_down_predicate` set to `year=='2023' and month=='10'`.
D.Create a new Glue Data Catalog table that points only to the `year=2023/month=10` prefix, and then read from that table.
AnswerC

Using `push_down_predicate` with `create_dynamic_frame.from_catalog` allows Glue to filter partitions at the catalog level, so only the specified partitions are read from S3. This reduces data scanned and improves performance. The predicate syntax uses SQL-like expressions on partition columns, and Glue leverages the partition metadata to avoid listing and reading unnecessary partitions. This is the recommended approach for partitioned data in Glue ETL jobs.

Why this answer

The correct approach is to use `push_down_predicate` with `create_dynamic_frame.from_catalog`. This pushes the filter down to the Glue Data Catalog, so only the relevant partitions are read from Amazon S3. It minimizes data scanned, reduces cost, and improves job performance.

Other methods either read all data first or bypass the catalog, which is less efficient and harder to maintain.

Exam trap

The trap here is assuming that filtering after reading the data is equivalent to partition pruning, but it still scans all data.

1185
MCQmedium

A data engineer needs to catalog a growing S3 data lake. New CSV files land in s3://analytics/raw/orders/ with a partition structure year=YYYY/month=MM/day=DD/. The engineer must create an AWS Glue Data Catalog table that automatically recognizes these partitions and requires no crawler runs for future dates. Which approach meets these requirements?

A.Create a Glue crawler with a schedule of every hour to discover new partitions.
B.Create the table with CREATE EXTERNAL TABLE and specify PARTITIONED BY (year string, month string, day string), then enable partition projection in the table properties.
C.Create a table using an AWS Glue crawler once, then run MSCK REPAIR TABLE on the table after each new partition is written.
D.Create the table with CREATE EXTERNAL TABLE and add each partition manually with ALTER TABLE ADD PARTITION as files arrive.
AnswerB

Partition projection on an Athena/Glue table lets the engine compute partition locations mathematically from the defined pattern and date range instead of consulting the catalog for each partition. New day, month, or year folders are therefore queryable immediately with no crawler or repair step, which is exactly the automatic, low-maintenance behavior the scenario requires.

Why this answer

Partition projection makes Athena and Glue compute partition metadata from the table's own properties, so newly arriving year/month/day folders are queryable without a crawler or repair command. The other choices all require periodic human or scheduled action after each new partition lands, which violates the explicit no-crawler-runs requirement.

Exam trap

The trap here is assuming the Data Catalog must always be refreshed by a crawler or MSCK REPAIR TABLE before new partitions are visible, when partition projection removes that dependency entirely.

1186
MCQeasy

A data engineer needs to store JSON documents that are frequently updated and require ACID transactions. Which AWS database service is most appropriate?

A.Amazon Neptune
B.Amazon DocumentDB
C.Amazon DynamoDB
D.Amazon S3
AnswerB

DocumentDB is for MongoDB workloads, ACID transactions are not fully supported.

Why this answer

Amazon DocumentDB (with MongoDB compatibility) is the correct choice because it is a fully managed document database designed to store, query, and index JSON-like documents. It supports multi-document ACID transactions, providing atomicity, consistency, isolation, and durability for frequently updated JSON documents. DocumentDB also offers a MongoDB-compatible API, making it a natural fit for JSON document workloads.

While Amazon DynamoDB supports ACID transactions via TransactWriteItems and TransactGetItems and can store JSON-like items, it is primarily a key-value and document store with a different data model and is not the AWS service purpose-built for general-purpose JSON document storage with complex querying and multi-document transactions. Amazon Neptune is a graph database, and Amazon S3 is object storage without ACID transactional semantics, so neither is appropriate.

Exam trap

Candidates may choose DynamoDB because it supports ACID transactions and can store JSON-like items. However, DynamoDB is a key-value/document store optimized for single-digit millisecond performance at scale, not for general-purpose JSON document storage with rich querying and multi-document transactions. DocumentDB is the AWS service specifically designed for JSON document workloads with ACID transaction support.

How to eliminate wrong answers

Option A is wrong because Amazon Neptune is a graph database designed for highly connected data (e.g., social networks, recommendation engines) and does not support ACID transactions across multiple documents in the same way DynamoDB does; it uses a property graph model and SPARQL/Gremlin, not a document store. Option B is wrong because Amazon DocumentDB is a MongoDB-compatible document database that supports ACID transactions only at the document level (single-document atomicity), not multi-document transactions, and its JSON handling is optimized for MongoDB workloads, not the high-frequency updates with full ACID guarantees required here. Option D is wrong because Amazon S3 is an object storage service that does not support ACID transactions; it offers eventual consistency for overwrite PUTS and lacks atomic multi-key operations, making it unsuitable for frequently updated JSON documents requiring transactional integrity.

1187
Multi-Selectmedium

A data engineer is designing an ingestion pipeline that uses Amazon Kinesis Data Firehose to deliver streaming records into an Amazon S3 bucket. The records arrive as JSON, and downstream consumers require Parquet with a stable schema. The engineer must configure the Firehose delivery stream so records are converted to Parquet before landing in S3. (Choose two.)

Select 2 answers
A.Enable Amazon S3 Object Lambda access points on the destination bucket to convert objects to Parquet on read.
B.Enable record format conversion on the Firehose delivery stream and select Apache Parquet as the output format.
C.Attach an AWS Glue table as the schema reference for the record format conversion configuration.
D.Configure an AWS Lambda function as the delivery stream's transformation to rewrite each record into Parquet bytes.
E.Set the S3 destination prefix to include a .parquet file extension so Firehose writes columnar files.
AnswersB, C

Firehose supports record format conversion from JSON to Parquet. Enabling it and selecting Apache Parquet as the output format instructs Firehose to convert records in flight before writing to S3. Without this setting, Firehose writes the raw JSON payload, so downstream consumers would not receive Parquet as required.

Why this answer

Firehose record format conversion converts JSON to Parquet in flight, and it requires a schema reference. The supported schema source is an AWS Glue table in the Data Catalog that describes the incoming JSON structure. Enabling conversion with Apache Parquet output and pointing to a matching Glue table together produce Parquet files in S3 for downstream analytics.

Exam trap

The trap here is assuming a Lambda transform or a file extension change can produce Parquet, when Firehose requires its built-in record format conversion with a Glue table schema reference.

1188
MCQmedium

A company is migrating an on-premises MySQL database to Amazon RDS for MySQL. The database is 500 GB and has a 24/7 uptime requirement. The migration must minimize downtime. Which approach should be used?

A.Take a snapshot of the on-premises database, convert it to a volume, and restore to RDS.
B.Use AWS Database Migration Service (DMS) with ongoing replication to migrate the data.
C.Export the database using mysqldump and import it into RDS using mysql command.
D.Create an RDS MySQL read replica from the on-premises database using native replication.
AnswerB

AWS DMS with ongoing replication performs a full load then continuously applies change data capture from the source MySQL binlog, keeping the target synchronised so cutover downtime is limited to a brief switchover rather than the whole 500 GB transfer.

Why this answer

AWS DMS with ongoing replication (change data capture) allows you to perform a full load of the 500 GB database and then continuously replicate changes from the on-premises MySQL source to the Amazon RDS target. This minimizes downtime because you can cut over to RDS in seconds after the target is synchronized, rather than taking the source offline for an extended period.

Exam trap

The trap here is that candidates often choose mysqldump (Option C) because it is a familiar tool, but they overlook the requirement for minimal downtime and the fact that a 500 GB dump/import would take hours, violating the 24/7 uptime requirement.

How to eliminate wrong answers

Option A is wrong because taking a snapshot of an on-premises database and converting it to a volume is not a supported method for migrating to RDS; snapshots are native to AWS block storage and cannot be directly created from an on-premises database. Option C is wrong because using mysqldump and mysql import requires the source database to be read-locked or offline during the export/import process, causing significant downtime for a 500 GB database with a 24/7 uptime requirement. Option D is wrong because RDS cannot be configured as a read replica of an on-premises MySQL database using native replication; native MySQL replication requires the replica to have direct network access to the source, and RDS does not support being a replica of an external source—only the reverse (RDS as source to external replica) is possible.

1189
MCQmedium

A data engineer applies the following IAM policy to an IAM user: ```json { "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Action": "s3:GetObject", "Resource": "arn:aws:s3:::example-bucket/*", "Condition": { "StringEquals": { "s3:x-amz-server-side-encryption": "AES256" } } } ] } ``` The user attempts to download an object from the bucket 'example-bucket' that is encrypted with SSE-S3 (AES256). Will the request succeed?

A.Yes, but only if the user also has s3:ListBucket permission.
B.No, because the policy requires the encryption to be specified in the request.
C.Yes, because the object is encrypted with SSE-S3 which uses AES256.
D.No, because the policy does not allow the s3:GetObject action for encrypted objects.
AnswerB

The condition key `s3:x-amz-server-side-encryption` evaluates the request header, not the object's stored encryption state. SSE-S3 encrypts objects automatically without requiring that header, so a plain GetObject request carries no matching key and the condition fails. The download is denied despite the object being AES256-encrypted.

Why this answer

The IAM policy includes a condition that requires the request to include the `x-amz-server-side-encryption` header with value `AES256`. Even though the object is encrypted with SSE-S3, the policy condition evaluates the request headers, not the object's encryption state. Since the user does not specify the encryption header in the download request, the condition fails, and the request is denied.

Exam trap

The trap is that candidates assume SSE-S3 is transparent and always allows access, overlooking that the IAM policy condition explicitly requires the encryption header in the request. The condition applies to the request, not the object's encryption-at-rest.

How to eliminate wrong answers

Option A is wrong because s3:ListBucket permission is not required to download an object; s3:GetObject alone suffices, and the policy does not reference ListBucket. Option B is wrong because the policy requires encryption to be specified in the request only for objects encrypted with SSE-KMS or SSE-C, not for SSE-S3 objects, which are automatically handled by S3 without client-side encryption headers. Option D is wrong because the policy does allow s3:GetAction for encrypted objects; the condition only denies requests that fail to include encryption headers, and SSE-S3 objects do not require such headers.

1190
MCQhard

A company uses AWS Glue to process data from Amazon S3. The data contains personally identifiable information (PII). The data engineer needs to automatically detect and mask PII fields before the data is loaded into Amazon Redshift. Which combination of AWS services should be used?

A.Amazon Macie and AWS Glue
B.Amazon CloudWatch Logs and AWS Lambda
C.Amazon S3 Object Lambda and AWS Glue
D.AWS IAM Access Analyzer and AWS Glue
AnswerA

Amazon Macie uses managed and custom data identifiers to detect PII in S3, publishing findings that drive the masking logic. AWS Glue then applies those detections during ETL, masking fields before writing to Redshift. Together they deliver automated detection and masking without manual schema inspection.

Why this answer

Option A is correct because Amazon Macie uses machine learning to automatically discover and classify PII in S3 data, and AWS Glue can then apply transforms (e.g., via a Glue ETL job or Glue Studio) to mask or redact those fields before loading into Redshift. Macie identifies sensitive data locations, and Glue performs the masking during the ETL pipeline, satisfying the requirement to detect and mask PII automatically.

Exam trap

DEA-C01 often tests the assumption that any AWS service with 'access' or 'analyzer' in its name can detect PII — candidates pick IAM Access Analyzer or S3 Object Lambda instead of Macie, which is the only service purpose-built for PII discovery.

How to eliminate wrong answers

Option B is wrong because CloudWatch Logs and Lambda are for log monitoring and event-driven compute, not for PII detection or data masking in S3-to-Redshift pipelines. Option C is wrong because S3 Object Lambda is used to transform data on retrieval via a Lambda function, but it does not provide automatic PII detection/classification like Macie, and it is not the standard combination for Glue-based ETL masking. Option D is wrong because IAM Access Analyzer identifies resource policies that grant external access — it does not detect or mask PII content.

1191
MCQeasy

A company stores sensitive data in Amazon S3 and needs to ensure that data is encrypted at rest. The security team requires that the company manage its own encryption keys and have the ability to audit key usage. Which S3 encryption option should the data engineer choose?

A.Client-side encryption with an AWS KMS key
B.Server-side encryption with customer-provided keys (SSE-C)
C.Server-side encryption with Amazon S3 managed keys (SSE-S3)
D.Server-side encryption with AWS KMS keys (SSE-KMS)
AnswerD

SSE-KMS allows the company to use AWS KMS customer managed keys, giving them control over key management and the ability to audit key usage via AWS CloudTrail. This meets the requirements for self-managed keys and auditability. It also supports key rotation and granular access policies, making it the correct choice for sensitive data with strict security requirements.

Why this answer

SSE-KMS allows the use of AWS KMS customer managed keys, providing control over key management and auditability through CloudTrail. This satisfies the requirements for self-managed keys and auditing key usage. Other options either do not provide key control or lack auditability, making SSE-KMS the correct choice.

Exam trap

The trap here is confusing SSE-C with SSE-KMS; SSE-C lets you provide keys but does not offer built-in auditability of key usage.

1192
MCQeasy

A company needs to store JSON documents that are frequently read and written by a web application. The data must be highly available and durable across multiple Availability Zones. Which AWS database service meets these requirements?

A.Amazon RDS for PostgreSQL
B.Amazon S3
C.Amazon DynamoDB
D.Amazon ElastiCache for Redis
AnswerC

DynamoDB is a fully managed NoSQL store holding JSON as native items, replicating synchronously across multiple Availability Zones for durability and high availability. Its key-value and document model suits frequent reads and writes from a web application without schema migrations.

Why this answer

Amazon DynamoDB is a fully managed NoSQL key-value and document database that provides single-digit millisecond performance at any scale. It stores JSON documents natively, supports frequent reads and writes, and offers built-in high availability and durability by automatically replicating data across multiple Availability Zones (AZs) in an AWS Region. This makes it the ideal choice for the described web application workload.

Exam trap

The trap here is that candidates often confuse Amazon S3's high durability and availability with database capabilities, overlooking that S3 is an object store with higher latency and no native query support, while DynamoDB is purpose-built for low-latency, high-throughput document storage with ACID transactions via DynamoDB Transactions.

How to eliminate wrong answers

Option A is wrong because Amazon RDS for PostgreSQL is a relational database that stores data in tables with a fixed schema, not as JSON documents natively, and while it can be deployed in a Multi-AZ configuration for high availability, it does not provide the same level of automatic, seamless scaling and native JSON document support as DynamoDB. Option B is wrong because Amazon S3 is an object storage service, not a database; it can store JSON files but is not designed for frequent, low-latency read/write operations from a web application, and it lacks features like atomic transactions and query capabilities that a database provides. Option D is wrong because Amazon ElastiCache for Redis is an in-memory cache, not a durable database; while it can store JSON documents using the RedisJSON module, data is primarily stored in memory and is not durable by default across AZs, making it unsuitable for the durability and persistence requirements of a primary data store.

1193
MCQeasy

A company runs a MySQL database on Amazon RDS. The database size is 500 GB and is experiencing high read traffic. The team wants to improve read performance with minimal operational overhead. Which action should they take?

A.Create a read replica in the same region
B.Enable Multi-AZ deployment
C.Implement Amazon ElastiCache for caching
D.Upgrade to a larger instance class
AnswerA

A read replica offloads read traffic from the primary MySQL instance via asynchronous replication, directly addressing the high read volume. It requires no application rewrite and minimal operational effort, satisfying the minimal-overhead constraint. Multi-AZ standby would not serve reads, and sharding adds substantial complexity.

Why this answer

Creating a read replica in the same region offloads read traffic from the primary RDS instance by providing a separate, read-only copy of the database. This directly addresses high read traffic with minimal operational overhead, as RDS manages the asynchronous replication using MySQL's native binlog-based replication. The replica can serve SELECT queries, reducing load on the primary instance without requiring application changes.

Exam trap

The trap here is confusing Multi-AZ (high availability) with read scaling, leading candidates to select Multi-AZ deployment thinking it improves read performance, when in fact the standby is not accessible for reads.

How to eliminate wrong answers

Option B is wrong because Multi-AZ deployment provides high availability and automatic failover, not read scaling; the standby instance cannot serve read traffic. Option C is wrong because while ElastiCache can improve read performance for cached data, it requires application-level caching logic and does not offload database read queries for all data, adding operational complexity. Option D is wrong because upgrading to a larger instance class scales vertically, which can improve performance but incurs higher cost and downtime during scaling, and does not distribute read load as efficiently as a read replica.

1194
MCQmedium

A data engineer needs to transform JSON data from Amazon S3 into Parquet using AWS Glue. The JSON is nested and contains arrays. The engineer wants to flatten the nested structure and write the result to S3 partitioned by a 'region' field. Which combination of Glue transforms should the engineer use?

A.Use the 'ResolveChoice' transform to flatten nested data, then write with partition keys.
B.Use the 'Relationalize' transform to flatten the nested JSON, then write the resulting DynamicFrame with partition keys set to 'region'.
C.Use the 'DropNullFields' transform to remove nulls, then write with partition keys.
D.Use the 'Unbox' transform to extract nested fields, then use 'ApplyMapping' to rename columns, and write with partition keys.
AnswerB

Relationalize flattens nested structures and arrays into separate relational tables, producing a DynamicFrame collection. The engineer can select the desired table, set partition keys to 'region', and write to S3 in Parquet. This is the standard Glue approach for flattening nested JSON before writing partitioned output, and it avoids manual schema manipulation.

Why this answer

Relationalize is designed to flatten nested JSON and arrays into relational tables, producing a collection of DynamicFrames. The engineer can then select the relevant table, set the 'region' field as a partition key, and write Parquet to S3. This is the intended Glue transform for nested-to-relational conversion.

Exam trap

The trap here is confusing Unbox, which extracts a single nested field, with Relationalize, which fully flattens nested structures and arrays into multiple relational tables.

1195
MCQmedium

A company uses AWS DMS to migrate data from an on-premises Oracle database to Amazon Aurora MySQL. The migration is successful, but the ongoing replication task is experiencing high latency. Which configuration change is most likely to reduce latency?

A.Increase the size of the DMS replication instance.
B.Decrease the task's batch size and batch apply timeout.
C.Change the target endpoint to Amazon S3.
D.Enable Change Data Capture (CDC) from binary logs.
AnswerA

Replication latency often stems from insufficient CPU, memory, or I/O on the replication instance. Upsizing it gives more resources to apply changes and process transactions, reducing the lag between source Oracle and target Aurora MySQL.

Why this answer

Ongoing replication latency in AWS DMS is most commonly caused by insufficient compute, memory, or I/O capacity on the replication instance, especially when CDC processing, transformations, or high transaction volumes are involved. Increasing the replication instance size provides more CPU, memory, and network throughput to process change streams faster. This is the standard first remediation for sustained CDC latency before tuning task-level parameters.

Exam trap

The trap is reaching for task-level tuning parameters (batch size, timeout) when the root cause is instance capacity — candidates often assume configuration tweaks beat vertical scaling for latency.

How to eliminate wrong answers

Option B is wrong because decreasing batch size and batch apply timeout typically reduces throughput and can increase latency — larger batches with appropriate timeouts generally improve apply performance, though tuning must be balanced against target constraints. Option C is wrong because changing the target to Amazon S3 does not address latency; it changes the target type entirely and would not satisfy the Aurora MySQL migration requirement. Option D is wrong because CDC from binary logs is already the mechanism DMS uses for Oracle-to-Aurora MySQL ongoing replication — enabling it is not a new change and does not by itself reduce latency.

1196
MCQmedium

A data engineer is troubleshooting a Lambda function that reads from the Kinesis stream 'my-data-stream'. The Lambda function is able to read data but occasionally fails with 'KMS.AccessDeniedException'. What is the most likely cause?

A.The Lambda function's execution role does not have kms:Decrypt permission for the KMS key.
B.The retention period is too short; increase it.
C.The stream has too few shards; increase shard count.
D.The Lambda function is not authorized to consume from Kinesis streams.
AnswerA

Kinesis encrypts stream data with a customer managed KMS key, so every GetRecords call requires kms:Decrypt. The intermittent KMS.AccessDeniedException means the function's execution role lacks that permission on the key policy, blocking decryption of returned records even though the read itself succeeds.

Why this answer

The KMS.AccessDeniedException indicates that the Lambda function's execution role lacks the kms:Decrypt permission for the AWS KMS key used to encrypt the Kinesis stream. When a Kinesis stream is encrypted with a customer managed KMS key, the consumer (Lambda) must have explicit decrypt permissions on that key to read the data records.

Exam trap

The trap here is that candidates may confuse KMS permissions with Kinesis stream permissions, assuming the error is about stream consumption authorization rather than decryption of encrypted data.

How to eliminate wrong answers

Option B is wrong because a short retention period would cause data to expire, not produce a KMS.AccessDeniedException. Option C is wrong because insufficient shards would cause throttling or throughput issues, not a KMS access error. Option D is wrong because the Lambda function is already able to read data (as stated), so it has Kinesis consumption authorization; the error is specifically about KMS decryption, not stream-level permissions.

1197
MCQmedium

A data engineer is monitoring an AWS Glue job that reads from Amazon S3 and writes to Amazon Redshift. The job runs daily and recently started taking significantly longer to complete. The engineer checks the job metrics and notices that the number of DPUs used is consistently at the maximum allocated, and the job's Spark UI shows many tasks spilling to disk. Which action should the engineer take to improve performance?

A.Change the output format from Parquet to CSV to reduce write overhead.
B.Increase the number of DPUs allocated to the job.
C.Repartition the data or adjust the number of partitions to reduce skew and improve memory usage.
D.Enable job bookmarks to avoid reprocessing old data.
AnswerC

Disk spilling occurs when Spark executors run out of memory and write intermediate data to disk. This is often caused by data skew or too few partitions, leading to large tasks. Repartitioning the data or increasing the number of partitions distributes the workload more evenly, reducing memory pressure and spilling. This directly addresses the root cause shown in the Spark UI.

Why this answer

Disk spilling in Spark indicates that executors are running out of memory during processing, often due to data skew or insufficient partitions. Repartitioning the data or adjusting partition counts distributes the workload more evenly across executors, reducing memory pressure and eliminating spilling. This is a targeted fix for the observed symptom and improves job performance without unnecessary resource increases.

Exam trap

The trap here is assuming that adding more DPUs will always solve performance issues, when the real problem is often data distribution and memory management within Spark.

1198
MCQeasy

A data engineer must load a 2 GB uncompressed CSV file from Amazon S3 into Amazon Redshift using the COPY command. The cluster is a two-node ra3.xlplus cluster, and the load is running far slower than expected. The engineer wants the fastest reliable improvement without changing the cluster. What should the engineer do?

A.Convert the CSV to GZIP before loading and keep it as one file.
B.Add the COMPUPDATE OFF and STATUPDATE OFF parameters to the COPY command.
C.Increase the cluster to four nodes so more slices can read the same file simultaneously.
D.Split the single CSV file into multiple files and load them in parallel with a single COPY command.
AnswerD

Amazon Redshift parallelizes COPY across slices, but a single file can only be read by one slice at a time, so a lone 2 GB file serializes the load. Splitting it into multiple files lets each slice read its own file concurrently, dramatically improving throughput. This is the standard remedy and requires no cluster change.

Why this answer

Redshift distributes a COPY workload across slices, but each input file is read by only one slice. A single 2 GB CSV therefore runs on one slice while the rest idle. Splitting the file into several smaller files, ideally a multiple of the slice count, allows all slices to read concurrently, which is the fastest reliable fix that does not alter the cluster.

Exam trap

The trap here is thinking that more cluster nodes or compression will speed up a single-file COPY, when the real limiter is that one file maps to one slice.

1199
MCQeasy

A company uses Amazon Athena to query data in S3. Recently, queries have become slow. The data is stored as CSV files in a partitioned table. What is the most effective way to improve query performance?

A.Increase the number of nodes in the Athena query engine.
B.Convert the data to Parquet format and optimize partitioning.
C.Convert the data to JSON format.
D.Increase the size of the CSV files to reduce the number of files.
AnswerB

Parquet is columnar, so Athena reads only referenced columns and skips row-by-row parsing of CSV, cutting scanned bytes and cost. Combined with tighter partition pruning, this directly addresses the slow queries against the partitioned S3 table.

Why this answer

The correct answer. Parquet is a columnar storage format that allows Athena to read only the columns needed for a query, reducing I/O and improving performance. Combined with effective partitioning, it enables partition pruning, which further limits the data scanned.

CSV files are row-based and require full scans, even with partitioning. Option A is incorrect because Athena is serverless and users cannot increase nodes; resources are managed automatically. Option C is incorrect: JSON is also row-based and verbose, making it even slower than CSV.

Option D is incorrect because larger CSV files still lead to full scans; Parquet's columnar nature is more impactful than file size.

Exam trap

A common trap is to think that simply increasing file size or using a more popular format like JSON will help. However, the key is switching to a columnar format (Parquet or ORC) that minimizes data scanned.

1200
MCQhard

A data engineer is using AWS Glue DataBrew to profile a dataset in Amazon S3. The dataset contains a column 'customer_id' that should be unique. The engineer runs a profile job and notices that the 'customer_id' column has a uniqueness metric of 98%. The engineer needs to identify the duplicate values. Which DataBrew feature should the engineer use to display the duplicate values?

A.Use the 'Column profile' view to see a histogram of values.
B.Use the 'Value frequencies' view in the profile job results to see the most frequent values.
C.Use the 'Column statistics' view in the profile job results to see the count of duplicates.
D.Use the 'Data quality rules' feature to create a rule that flags duplicates.
AnswerB

In AWS Glue DataBrew profile results, the 'Value frequencies' view displays the most common values in a column along with their counts. For a column like 'customer_id' with duplicates, this view will show which IDs appear more than once, allowing the engineer to identify the duplicate values. It is the correct feature for this purpose.

Why this answer

AWS Glue DataBrew profile jobs include a 'Value frequencies' view that lists the most frequent values in a column with their counts. For a column that should be unique, this view directly reveals which customer_id values are duplicated. The other options provide summary statistics or rules but do not display the specific duplicate values.

Exam trap

The trap here is assuming that the uniqueness metric or column statistics will list the duplicate values, when in fact you need to examine the value frequencies to see the actual duplicates.

Page 15

Page 16 of 18

Page 17